Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations

Francesco LocatelloStefan BauerMario LučićSylvain GellyBernhard SchölkopfOlivier Bachem

article2018ICML1,880 citationsBest Paper Award

Proves that unsupervised disentanglement is fundamentally impossible without inductive biases and demonstrates across 12,000 trained models that disentangled representations cannot be reliably identified or expected to improve downstream learning without supervision.

Listen

Modern machine learning relies heavily on learning compact representations from complex data, often operating under the premise that unsupervised models can automatically isolate independent explanatory factors of variation, such as object shape, color, or position. This capability, known as disentanglement, has been widely assumed to yield models that are more interpretable, better at transferring knowledge, and more data-efficient when training downstream classifiers. The article set out to theoretically evaluate whether unsupervised disentanglement is mathematically possible without inductive biases and to empirically assess whether state-of-the-art unsupervised methods reliably achieve disentanglement and improve downstream task performance.

To evaluate these questions, the authors proved a theoretical impossibility theorem and executed a large-scale, reproducible empirical study. The experimental setup implemented six prominent unsupervised methods based on variational autoencoders, six disentanglement metrics, and seven distinct benchmark datasets (including synthetic image datasets with deterministic and noisy backgrounds). Holding neural network architectures, optimization parameters, and batch sizes constant to isolate the impact of model objectives and regularization strength, the researchers trained more than 12,000 models across 50 random initialization seeds per configuration, amounting to approximately 2.5 GPU years of computation.

The investigation produced four central findings. First, the theoretical analysis proved that unsupervised disentanglement is fundamentally impossible for arbitrary generative models without explicit inductive biases on both the learning algorithm and the data. Second, while the tested algorithms succeeded in reducing correlation across the sampled latent space, they paradoxically increased correlation among the dimensions of the deterministic mean representation that is actually used in practice. Third, model architecture and objective choice accounted for only 37% of the variance in disentanglement performance, whereas hyperparameter tuning and random seeds accounted for the remainder, meaning random initialization often outweighed algorithm design. Fourth, unsupervised model selection proved largely ineffective: standard unsupervised metrics (such as reconstruction error or evidence lower bounds) failed to correlate with disentanglement scores, and transferring hyperparameters across datasets only beat random model selection 59.3% of the time. Finally, higher disentanglement scores did not reliably decrease sample complexity or improve learning efficiency on downstream classification tasks.

These findings challenge fundamental assumptions in representation learning, indicating that unsupervised disentanglement cannot be achieved reliably using current paradigms. For organizations investing in machine learning research and deployment, pursuing purely unsupervised disentangled representations carries high computational costs with little guarantee of functional advantage. If practitioners must rely on ground-truth labels to identify successful runs or tune hyperparameters, the approach ceases to be truly unsupervised.

Consequently, the article recommends shifting research and development away from static, purely unsupervised methods. Stakeholders should instead focus on approaches that incorporate explicit inductive biases and weak or structured supervision, such as temporal coherence in video data, grouped observations, or interactive environments. Furthermore, future research must explicitly validate concrete downstream benefits, such as fairness, interpretability, or causal inference, rather than assuming general efficiency gains.

The conclusions are supported by a rigorous mathematical proof and a large-scale empirical evaluation. However, the experimental findings are bounded by the study’s scope, which focused on convolutional variational autoencoders trained on synthetic image datasets with discrete factors of variation. While these boundary conditions warrant caution before generalizing the empirical results to entirely different data modalities or model families, the theoretical and experimental evidence strongly indicates that current unsupervised disentanglement techniques cannot be reliably deployed without structural biases or supervisory signals.

  • Paper: Identifying Weight-Variant Latent Causal Models, Yuhang Liu et al. (2026). It addresses the impossibility of purely unsupervised disentanglement proved in this source by establishing identifiable latent causal factors using auxiliary observed variables.
  • Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It broadens the source's critique of non-identifiability in unsupervised representations into a general statistical and causal diagnosis of explainability and interpretability methods.
  • Paper: An Introduction to Variational Autoencoders, Diederik P. Kingma et al. (2019). It offers an extensive, unified treatment of the broader variational autoencoder framework and its structural extensions in light of recent findings on representation capacity and inductive biases.
Cover for Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations

Abstract

The key idea behind the unsupervised learning of disentangled representations is that real-world data is generated by a few explanatory factors of variation which can be recovered by unsupervised learning algorithms. In this paper, we provide a sober look at recent progress in the field and challenge some common assumptions. We first theoretically show that the unsupervised learning of disentangled representations is fundamentally impossible without inductive biases on both the models and the data. Then, we train more than 12000 models covering most prominent methods and evaluation metrics in a reproducible large-scale experimental study on seven different data sets. We observe that while the different methods successfully enforce properties ``encouraged'' by the corresponding losses, well-disentangled models seemingly cannot be identified without supervision. Furthermore, increased disentanglement does not seem to lead to a decreased sample complexity of learning for downstream tasks. Our results suggest that future work on disentanglement learning should be explicit about the role of inductive biases and (implicit) supervision, investigate concrete benefits of enforcing disentanglement of the learned representations, and consider a reproducible experimental setup covering several data sets.

Table of Contents

  • 1 Introduction
  • 2 Other related work
  • 3 Impossibility result
  • 4 Experimental design
  • 5 Key experimental results
  • 5.1 Can current methods enforce a uncorrelated aggregated posterior and representation?
  • 5.2 How much do the disentanglement metrics agree?
  • 5.3 How important are different models and hyperparameters for disentanglement?
  • 5.4 Are there reliable recipes for model selection?
  • 5.5 Are these disentangled representations useful for downstream tasks in terms of the sample complexity of learning?
  • 6 Conclusions
  • References
  • A Proof of Theorem
  • B Unsupervised learning of disentangled representations with VAEs
  • C Implementation of metrics
  • D Experimental conditions and guiding principles.
  • E Limitations of our study.
  • F Differences with previous implementations.
  • G Main experiment hyperparameters
  • H Data sets and preprocessing
  • I Detailed experimental results
  • I.1 Can one achieve a good reconstruction error across data sets and models?
  • I.2 Can current methods enforce a uncorrelated aggregated posterior and representation?
  • I.3 How much do existing disentanglement metrics agree?
  • I.4 How important are different models and hyperparameters for disentanglement?
  • I.5 Are there reliable recipes for model selection?
  • I.6 Are these disentangled representations useful for downstream tasks in terms of the sample complexity of learning?
  • J Additional Figures

Knowls

  1. Knowl 1 — Impossibility of Unsupervised Disentanglement Without Inductive Biases

    theoretical result

    For dimension d>1d > 1, let z∼Pz \sim P be any multivariate latent random variable possessing a factorized probability density function p(z)=∏i=1dp(zi)p(z) = \prod_{i=1}^d p(z_i). There exists an infinite family of bijective transformations f:supp⁡(z)→supp⁡(z)f: \operatorname{supp}(z) \to \operatorname{supp}(z) such that the partial derivatives satisfy

    ∂fi(u)∂uj≠0\frac{\partial f_i(u)}{\partial u_j} \neq 0

    almost everywhere for all i,j∈{1,2,…,d}i, j \in \{1, 2, \dots, d\} (meaning zz and f(z)f(z) are completely entangled), while preserving the identical marginal cumulative distribution:

    P(z≤u)=P(f(z)≤u),∀u∈supp⁡(z)P(z \le u) = P(f(z) \le u), \quad \forall u \in \operatorname{supp}(z)

    Consequently, for any generative model defined by prior p(z)p(z) and observation conditional P(x∣z)P(x|z), an equivalent generative model exists with latent variable z^=f(z)\hat{z} = f(z) and prior p(z^)p(\hat{z}) that generates the identical marginal observation distribution:

    P(x)=∫p(x∣z)p(z) dz=∫p(x∣z^)p(z^) dz^P(x) = \int p(x|z)p(z)\,dz = \int p(x|\hat{z})p(\hat{z})\,d\hat{z}

    Because an unsupervised learning algorithm has access only to observations x∼P(x)x \sim P(x), it cannot distinguish between the disentangled latent variable zz and the completely entangled latent variable z^=f(z)\hat{z} = f(z). Therefore, purely unsupervised learning of disentangled representations is fundamentally impossible without inductive biases on both the learning models and the data distributions.

  2. Knowl 2 — Dominance of Random Seeds and Hyperparameters over Loss Function Choice

    empirical result

    In large-scale unsupervised disentanglement learning across multiple Variational Autoencoder (VAE) architectures and datasets, the choice of objective function explains only a minority of the variance in disentanglement performance.

    Ordinary least squares regression demonstrates that the objective function alone (treated as a categorical variable) explains on average only 37% of the variance across different disentanglement metrics and datasets. When conditioned on the Cartesian product of the objective function and the regularization strength hyperparameter, 59% of the variance is explained, leaving the remaining 41% of variance driven entirely by the random seed.

    Because performance ranges across different random initializations heavily overlap between methods, a favorable random seed with a suboptimal hyperparameter frequently outperforms an unfavorable seed with an optimal hyperparameter, indicating that model choice is less significant than seed variance and hyperparameter tuning.

  3. Knowl 3 — Failure of Unsupervised Metrics and Cross-Dataset Transfer for Model Selection

    empirical result

    Unsupervised model selection for disentangled representations remains fundamentally unresolved:

    1. Unsupervised Validation Metrics: Standard unsupervised metrics—including reconstruction error, Kullback-Leibler divergence DKL(qϕ(z∣x)∥p(z))D_{\mathrm{KL}}(q_\phi(z|x) \parallel p(z)), the Evidence Lower Bound (ELBO), and estimated total correlation of the latent representation—exhibit weak, inconsistent, or negative rank correlations with ground-truth supervised disentanglement metrics across datasets.

    2. Hyperparameter Transfer: Selecting the best-performing hyperparameter configuration on a labeled source dataset (e.g., dSprites) and transferring it to an unlabeled target dataset yields only a 59.3% empirical probability of outperforming random model selection when evaluated on the same metric. When transferring across both novel datasets and differing disentanglement metrics, the probability drops to 54.9% (barely exceeding random choice). Transfer fails because hyperparameter selection cannot identify or isolate favorable random seeds on target datasets without supervision.

  4. Knowl 4 — Divergence in Total Correlation Between Sampled Posterior and Mean Representations

    empirical result

    Unsupervised disentanglement methods that penalize total correlation or enforce factorized aggregated posteriors succeed in reducing correlation among dimensions of the sampled aggregated posterior z∼qϕ(z∣x)z \sim q_\phi(z|x), but typically cause the dimensions of the deterministic mean representation r(x)=Eqϕ(z∣x)[z]=μϕ(x)r(x) = \mathbb{E}_{q_\phi(z|x)}[z] = \mu_\phi(x) to become increasingly correlated.

    As the regularization strength increases in models such as β\beta-VAE, β\beta-TCVAE, and FactorVAE, the total correlation computed on Gaussian distributions fitted to sampled representations zz systematically decreases. However, the total correlation and pairwise mutual information computed on Gaussian distributions fitted to the mean representations μϕ(x)\mu_\phi(x) consistently increase.

    The only exception among evaluated models is DIP-VAE-I, which directly penalizes the off-diagonal elements of the covariance matrix of μϕ(x)\mu_\phi(x), thereby constraining the mean representation's total correlation to remain low.

  5. Knowl 5 — Lack of Impact of Disentanglement on Downstream Sample Efficiency

    empirical result

    Evaluating downstream classification tasks—predicting ground-truth factors of variation from learned latent representations using multi-class Logistic Regression (LR) or Gradient Boosted Trees (GBT)—reveals that higher disentanglement scores do not reliably lead to decreased sample complexity or higher statistical efficiency.

    Statistical efficiency is quantified as the downstream classification accuracy achieved using N=100N = 100 training samples divided by the accuracy achieved using N=10 000N = 10\,000 training samples:

    Statistical Efficiency=AccuracyN=100AccuracyN=10 000\text{Statistical Efficiency} = \frac{\text{Accuracy}_{N=100}}{\text{Accuracy}_{N=10\,000}}

    Across multiple benchmark datasets (including dSprites, Cars3D, SmallNORB, and Shapes3D) and across six disentanglement metrics, representations with higher disentanglement scores fail to exhibit a consistent or significant increase in statistical efficiency. While disentanglement metrics correlate positively with raw downstream accuracy on specific synthetic benchmarks, this effect reflects the general informativeness of the learned latent codes rather than improved sample efficiency.

  6. Knowl 6 — Variational Autoencoder Objectives for Disentangled Representation Learning

    model/method

    Unsupervised disentanglement methods augment the standard Variational Autoencoder (VAE) Evidence Lower Bound (ELBO),

    max⁡ϕ,θEp(x)[Eqϕ(z∣x)[log⁡pθ(x∣z)]−DKL(qϕ(z∣x)∥p(z))]\max_{\phi, \theta} \mathbb{E}_{p(x)}\left[\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{\mathrm{KL}}(q_\phi(z|x) \parallel p(z))\right]

    with regularizers designed to encourage factorized latent representations:

    • β\beta-VAE: Scales the entire KL divergence term by hyperparameter β>1\beta > 1: Ep(x)[Eqϕ(z∣x)[log⁡pθ(x∣z)]−βDKL(qϕ(z∣x)∥p(z))]\mathbb{E}_{p(x)}\left[\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - \beta D_{\mathrm{KL}}(q_\phi(z|x) \parallel p(z))\right]

    • AnnealedVAE: Constrains the bottleneck by annealing capacity CC from 0 to Cmax⁡C_{\max}: Ep(x)[Eqϕ(z∣x)[log⁡pθ(x∣z)]−γ∣DKL(qϕ(z∣x)∥p(z))−C∣]\mathbb{E}_{p(x)}\left[\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - \gamma \left| D_{\mathrm{KL}}(q_\phi(z|x) \parallel p(z)) - C \right|\right]

    • FactorVAE and β\beta-TCVAE: Isolate and penalize the Total Correlation TC(q(z))=DKL(q(z) ∥ ∏j=1dq(zj))\mathrm{TC}(q(z)) = D_{\mathrm{KL}}\left(q(z) \,\parallel\, \prod_{j=1}^d q(z_j)\right) of the aggregated posterior q(z)=∫qϕ(z∣x)p(x) dxq(z) = \int q_\phi(z|x)p(x)\,dx: Ep(x)[Eqϕ(z∣x)[log⁡pθ(x∣z)]−DKL(qϕ(z∣x)∥p(z))]−γDKL(q(z) ∥ ∏j=1dq(zj))\mathbb{E}_{p(x)}\left[\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{\mathrm{KL}}(q_\phi(z|x) \parallel p(z))\right] - \gamma D_{\mathrm{KL}}\left(q(z) \,\parallel\, \prod_{j=1}^d q(z_j)\right) FactorVAE estimates total correlation using an auxiliary discriminator via the density ratio trick, whereas β\beta-TCVAE utilizes a minibatch-weighted Monte Carlo estimator.

    • DIP-VAE-I and DIP-VAE-II: Match moments of the aggregated posterior to a factorized unit Gaussian prior by regularizing deviations of covariance matrices from the identity matrix IdI_d:

      • DIP-VAE-I penalizes Cov⁡p(x)[μϕ(x)]\operatorname{Cov}_{p(x)}[\mu_\phi(x)]: max⁡ϕ,θELBO−λod∑i≠j(Cov⁡p(x)[μϕ(x)]ij)2−λd∑i(Cov⁡p(x)[μϕ(x)]ii−1)2\max_{\phi, \theta} \text{ELBO} - \lambda_{\mathrm{od}} \sum_{i \neq j} \left(\operatorname{Cov}_{p(x)}[\mu_\phi(x)]_{ij}\right)^2 - \lambda_{\mathrm{d}} \sum_i \left(\operatorname{Cov}_{p(x)}[\mu_\phi(x)]_{ii} - 1\right)^2
      • DIP-VAE-II penalizes Cov⁡qϕ[z]\operatorname{Cov}_{q_\phi}[z]: max⁡ϕ,θELBO−λod∑i≠j(Cov⁡qϕ[z]ij)2−λd∑i(Cov⁡qϕ[z]ii−1)2\max_{\phi, \theta} \text{ELBO} - \lambda_{\mathrm{od}} \sum_{i \neq j} \left(\operatorname{Cov}_{q_\phi}[z]_{ij}\right)^2 - \lambda_{\mathrm{d}} \sum_i \left(\operatorname{Cov}_{q_\phi}[z]_{ii} - 1\right)^2
  7. Knowl 7 — Evaluation Protocol for Disentanglement Metrics with Discrete Factors

    definition

    Disentanglement metrics evaluate the structural relationship between the deterministic mean representation r(x)=μϕ(x)∈Rdr(x) = \mu_\phi(x) \in \mathbb{R}^d and KK ground-truth discrete factors of variation z1,z2,…,zKz_1, z_2, \dots, z_K:

    • BetaVAE Metric: Measures the classification accuracy of a linear classifier trained to predict the index of a fixed ground-truth factor based on the coordinate-wise mean absolute differences ∣r(x(1))−r(x(2))∣|r(x^{(1)}) - r(x^{(2)})| across pairs of samples sharing that fixed factor.

    • FactorVAE Metric: Evaluates the accuracy of a majority-vote classifier that predicts the fixed factor index from the latent coordinate possessing the smallest empirical variance across an intervention batch, normalized by that coordinate's overall dataset variance.

    • Mutual Information Gap (MIG): Quantifies the discrete mutual information I(vj,zk)I(v_j, z_k) between each latent dimension vjv_j (binned into 20 bins) and ground-truth factor zkz_k, computing the normalized gap between the highest (jk=arg⁡max⁡jI(vj,zk)j_k = \arg\max_j I(v_j, z_k)) and second-highest predictive coordinates: MIG=1K∑k=1K1H(zk)(I(vjk,zk)−max⁡j≠jkI(vj,zk))\text{MIG} = \frac{1}{K} \sum_{k=1}^K \frac{1}{H(z_k)} \left( I(v_{j_k}, z_k) - \max_{j \neq j_k} I(v_j, z_k) \right)

    • Modularity: Computes the average degree to which each latent coordinate viv_i depends on at most a single factor zfz_f, defined as 1−δi1 - \delta_i where δi=∑f≠f∗(mif)2θi2(K−1)\delta_i = \frac{\sum_{f \neq f^*} (m_{if})^2}{\theta_i^2 (K-1)}, mif=I(vi,zf)m_{if} = I(v_i, z_f), θi=max⁡gmig\theta_i = \max_g m_{ig}, and f∗=arg⁡max⁡gmigf^* = \arg\max_g m_{ig}.

    • DCI Disentanglement: Constructs a feature importance matrix R∈RK×dR \in \mathbb{R}^{K \times d} by fitting Gradient Boosted Trees to predict each factor zkz_k from r(x)r(x). Disentanglement is computed as ∑iρi(1−H(Ri))\sum_i \rho_i (1 - H(R_i)), where H(Ri)H(R_i) is the entropy of the normalized column distribution and ρi=∑jRji∑jkRjk\rho_i = \frac{\sum_j R_{ji}}{\sum_{jk} R_{jk}}.

    • Separated Attribute Predictability (SAP): Computes a score matrix of linear classification test errors (linear SVM with C=0.01C=0.01) predicting factor zkz_k from individual dimensions vjv_j, and averages the difference between the top two most predictive latent dimensions across all factors.

  8. Knowl 8 — Large-Scale Controlled Experimental Protocol for Disentanglement

    experimental setup

    The empirical study evaluates over 12,000 deep generative models using an open-source library (disentanglement_lib). To isolate the effect of regularization loss functions from network architecture biases, all methods share fixed architectural choices:

    • Encoder Architecture: 4 convolutional layers (4×44 \times 4 kernels, stride 2, ReLU activations, with channel depths 32, 32, 64, 64), followed by a fully connected layer with 256 ReLU units, outputting mean μ\mu and log-variance log⁡σ2\log \sigma^2 for a d=10d=10 dimensional latent space.
    • Decoder Architecture: Fully connected layer (256 units, ReLU), fully connected layer (4×4×644 \times 4 \times 64, ReLU), and 4 transposed convolutions (4×44 \times 4 kernels, stride 2; channels 64, 32, 32, number of image channels), parameterized with a Bernoulli likelihood.
    • FactorVAE Discriminator: 6 fully connected layers with 1000 LeakyReLU units each, outputting a 2-class logit, trained using Adam (learning rate 10−410^{-4}, β1=0.5,β2=0.9\beta_1=0.5, \beta_2=0.9).
    • Training Settings: Adam optimizer (learning rate 10−410^{-4}, β1=0.9,β2=0.999\beta_1=0.9, \beta_2=0.999, ϵ=10−8\epsilon=10^{-8}), batch size 64, trained for 300,000 steps.
    • Datasets: 7 datasets with known ground-truth discrete generative factors: dSprites, Cars3D, SmallNORB, Shapes3D, Color-dSprites (random channel scalings), Noisy-dSprites (uniform noise background), and Scream-dSprites (background patches extracted from The Scream).
    • Hyperparameter Grid: 6 regularization parameter values per model (50 random seeds per configuration):
      • β\beta-VAE: β∈{1,2,4,6,8,16}\beta \in \{1, 2, 4, 6, 8, 16\}
      • AnnealedVAE: cmax⁡∈{5,10,25,50,75,100}c_{\max} \in \{5, 10, 25, 50, 75, 100\}, γ=1000\gamma=1000, iteration threshold 100,000
      • FactorVAE: γ∈{10,20,30,40,50,100}\gamma \in \{10, 20, 30, 40, 50, 100\}
      • DIP-VAE-I: λod∈{1,2,5,10,20,50}\lambda_{\mathrm{od}} \in \{1, 2, 5, 10, 20, 50\}, λd=10λod\lambda_{\mathrm{d}} = 10\lambda_{\mathrm{od}}
      • DIP-VAE-II: λod∈{1,2,5,10,20,50}\lambda_{\mathrm{od}} \in \{1, 2, 5, 10, 20, 50\}, λd=λod\lambda_{\mathrm{d}} = \lambda_{\mathrm{od}}
      • β\beta-TCVAE: β∈{1,2,4,6,8,10}\beta \in \{1, 2, 4, 6, 8, 10\}
  9. Knowl 9 — Variance Explained by Objective Function vs. Cartesian Product with Hyperparameters

    data/table

    Percentage of variance (R2R^2) in disentanglement metrics explained when regressing scores against the objective function alone versus the Cartesian product of the objective function and the regularization strength hyperparameter, using ordinary least squares:

    Dataset BetaVAE DCI FactorVAE MIG Modularity SAP
    Panel A: Objective function only
    Cars3D 1% 36% 26% 34% 37% 13%
    Color-dSprites 30% 39% 52% 26% 23% 29%
    Noisy-dSprites 17% 21% 17% 11% 41% 6%
    Scream-dSprites 89% 50% 76% 45% 60% 56%
    Shapes3D 31% 21% 14% 20% 26% 10%
    SmallNORB 68% 71% 58% 71% 62% 62%
    dSprites 29% 41% 47% 26% 29% 31%
    Panel B: Cartesian product of objective and regularization strength
    Cars3D 4% 69% 42% 59% 51% 17%
    Color-dSprites 69% 80% 61% 76% 40% 56%
    Noisy-dSprites 26% 42% 25% 29% 50% 20%
    Scream-dSprites 93% 74% 83% 66% 68% 75%
    Shapes3D 61% 78% 53% 59% 49% 35%
    SmallNORB 87% 89% 81% 88% 72% 82%
    dSprites 59% 77% 54% 72% 39% 56%

    The objective function alone accounts for an average of only 37% of variance across datasets, whereas including the regularization strength accounts for 59% on average, demonstrating that the specific hyperparameter choice and random seed account for the vast majority of performance variance.

  10. Knowl 10 — Model Selection Probability Outperforming Random Guessing via Transfer

    data/table

    Empirical probability of a transfer-based model selection strategy outperforming random model selection across 10,000 random trials. In each trial, an optimal hyperparameter is identified on a source dataset, metric, and random seed, and evaluated against a randomly selected model under four transfer settings:

    Random different data set Same data set
    Random different metric 54.9% 62.6%
    Same metric 59.3% 80.7%

    While hyperparameter transfer within the same dataset across seeds succeeds in 80.7% of trials, transferring hyperparameters across different datasets achieves only a 59.3% success rate (for the same metric) and 54.9% (for a different metric), demonstrating that unsupervised hyperparameter selection cannot be reliably achieved via cross-dataset transfer.

  11. Knowl 11 — Clustering and Correlation Structure Among Disentanglement Metrics

    empirical result

    Pairwise Spearman rank correlation analyses across trained models demonstrate that existing disentanglement metrics partition into distinct groups of agreement:

    1. Metric Pairs: The BetaVAE score and the FactorVAE score exhibit high mutual rank correlation across datasets. Similarly, the Mutual Information Gap (MIG) and DCI Disentanglement exhibit strong pairwise rank correlation.
    2. Modularity Metric Divergence: The Modularity metric demonstrates weak, near-zero, or negative rank correlation with most other metrics across several datasets (e.g., negative rank correlations with BetaVAE, FactorVAE, and SAP on SmallNORB and dSprites).
    3. Dataset Dependence: The general level of correlation among all metrics varies substantially between datasets, displaying stronger overall metric alignment on synthetic benchmarks like dSprites, Color-dSprites, and Scream-dSprites, but weaker alignment on Cars3D and SmallNORB.
  12. Knowl 12 — Scope and Methodological Limitations of the Disentanglement Benchmark

    limitation

    The empirical findings are subject to specific methodological boundaries defined by the experimental design:

    • Synthetic and Image-Centric Data: The study is restricted to image modalities with synthetic or controlled benchmarks where ground-truth factors are independent, uniformly distributed, discrete, and free of confounding variables.
    • Architectural Constraints: A single convolutional architecture is evaluated without exploration of fully connected architectures, skip connections, varied activation functions, reconstruction loss alternatives, or variable latent space sizes (dd fixed to 10).
    • Optimization Hyperparameters: Optimizers (Adam), learning rates (10−410^{-4}), and batch sizes (64) are fixed across all models; only the regularization weight parameter is swept.
    • Static Unsupervised Setting: The setting excludes temporal structure (e.g., video sequences), interactive world environments, and weak supervision signals (e.g., grouped observations), which provide alternative mechanisms for identifiability.

Coverage note — None was omitted; all key theoretical proofs, experimental setups, metric definitions, empirical findings across hyperparameter sensitivity, model selection, downstream tasks, and limitation disclosures are fully covered.

References

  1. 1.Arcones, M. A. and Gine, E. On the bootstrap of u and v statistics. The Annals of Statistics, pp. 655–674, 1992.
  2. 2.Bach, F. R. and Jordan, M. I. Kernel independent component analysis. Journal of machine learning research, 3(Jul): 1–48, 2002.
  3. 3.Bengio, Y., LeCun, Y., et al. Scaling learning algorithms towards ai. Large-scale kernel machines, 34(5):1–41, 2007.
  4. 4.Bengio, Y., Courville, A., and Vincent, P. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8): 1798–1828, 2013.
  5. 5.Bouchacourt, D., Tomioka, R., and Nowozin, S. Multi-level variational autoencoder: Learning disentangled representations from grouped observations. In AAAI, 2018.
  6. 6.Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A. Understanding disentangling in beta-vae. In Workshop on Learning Disentangled Representations at the 31st Conference on Neural Information Processing Systems, 2017.
  7. 7.Chen, T. Q., Li, X., Grosse, R. B., and Duvenaud, D. K. Isolating sources of disentanglement in variational autoencoders. In Advances in Neural Information Processing Systems, pp. 2615–2625, 2018.
  8. 8.Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In Advances in Neural Information Processing Systems, pp. 2172–2180, 2016.
  9. 9.Cheung, B., Livezey, J. A., Bansal, A. K., and Olshausen, B. A. Discovering hidden factors of variation in deep networks. In Workshop at International Conference on Learning Representations, 2015.
  10. 10.Cohen, T. and Welling, M. Learning the irreducible representations of commutative lie groups. In International Conference on Machine Learning, pp. 1755–1763, 2014.
  11. 11.Cohen, T. S. and Welling, M. Transformation properties of learned visual representations. In International Conference on Learning Representations, 2015.
  12. 12.Comon, P. Independent component analysis, a new concept? Signal processing, 36(3):287–314, 1994.
  13. 13.Denton, E. L. and Birodkar, v. Unsupervised learning of disentangled representations from video. In Advances in Neural Information Processing Systems, pp. 4414–4423, 2017.
  14. 14.Desjardins, G., Courville, A., and Bengio, Y. Disentangling factors of variation via generative entangling. arXiv preprint arXiv:1210.5474, 2012.
  15. 15.Eastwood, C. and Williams, C. K. I. A framework for the quantitative evaluation of disentangled representations. In International Conference on Learning Representations, 2018.
  16. 16.Fraccaro, M., Kamronn, S., Paquet, U., and Winther, O. A disentangled recognition and nonlinear dynamics model for unsupervised learning. In Advances in Neural Information Processing Systems, pp. 3601–3610, 2017.
  17. 17.Goodfellow, I., Lee, H., Le, Q. V., Saxe, A., and Ng, A. Y. Measuring invariances in deep networks. In Advances in neural information processing systems, pp. 646–654, 2009.
  18. 18.Goroshin, R., Mathieu, M. F., and LeCun, Y. Learning to linearize under uncertainty. In Advances in Neural Information Processing Systems, pp. 1234–1242, 2015.
  19. 19.Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. betavae: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, 2017a.
  20. 20.Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A. Darla: Improving zero-shot transfer in reinforcement learning. In International Conference on Machine Learning, pp. 1480–1490, 2017b.
  21. 21.Higgins, I., Sonnerat, N., Matthey, L., Pal, A., Burgess, C. P., Bošnjak, M., Shanahan, M., Botvinick, M., Hassabis, D., and Lerchner, A. Scan: Learning hierarchical compositional visual concepts. In International Conference on Learning Representations, 2018.
  22. 22.Hinton, G. E., Krizhevsky, A., and Wang, S. D. Transforming auto-encoders. In International Conference on Artificial Neural Networks, pp. 44–51. Springer, 2011.
  23. 23.Hsu, W.-N., Zhang, Y., and Glass, J. Unsupervised learning of disentangled and interpretable representations from sequential data. In Advances in Neural Information Processing Systems, pp. 1878–1889, 2017.
  24. 24.Hyvarinen, A. and Morioka, H. Unsupervised feature extraction by time-contrastive learning and nonlinear ica. In Advances in Neural Information Processing Systems, pp. 3765–3773, 2016.
  25. 25.Hyvarinen, A. and Pajunen, P. Nonlinear independent component analysis: Existence and uniqueness results. Neural Networks, 12(3):429–439, 1999.
  26. 26.Hyvarinen, A., Sasaki, H., and Turner, R. E. Nonlinear ica using auxiliary variables and generalized contrastive learning. arXiv preprint arXiv:1805.08651, 2018.
  27. 27.Jutten, C. and Karhunen, J. Advances in nonlinear blind source separation. In Proc. of the 4th Int. Symp. on Independent Component Analysis and Blind Signal Separation (ICA2003), pp. 245–256, 2003.
  28. 28.Karaletsos, T., Belongie, S., and Rätsch, G. Bayesian representation learning with oracle constraints. arXiv preprint arXiv:1506.05011, 2015.
  29. 29.Kim, H. and Mnih, A. Disentangling by factorising. In Proceedings of the 35th International Conference on Machine Learning, pp. 2649–2658, 2018.
  30. 30.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014.
  31. 31.Kulkarni, T. D., Whitney, W. F., Kohli, P., and Tenenbaum, J. Deep convolutional inverse graphics network. In Advances in neural information processing systems, pp. 2539–2547, 2015.
  32. 32.Kumar, A., Sattigeri, P., and Balakrishnan, A. Variational inference of disentangled latent concepts from unlabeled observations. In International Conference on Learning Representations, 2017.
  33. 33.Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and Brain Sciences, 40, 2017.
  34. 34.Laversanne-Finot, A., Pere, A., and Oudeyer, P.-Y. Curiosity driven exploration of learned disentangled goal spaces. In Conference on Robot Learning, pp. 487–504, 2018.
  35. 35.LeCun, Y., Huang, F. J., and Bottou, L. Learning methods for generic object recognition with invariance to pose and lighting. In Computer Vision and Pattern Recognition, 2004. CVPR 2004. Proceedings of the 2004 IEEE Computer Society Conference on, volume 2, pp. II–104. IEEE, 2004.
  36. 36.LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. nature, 521(7553):436, 2015.
  37. 37.Lenc, K. and Vedaldi, A. Understanding image representations by measuring their equivariance and equivalence. In IEEE conference on computer vision and pattern recognition, pp. 991–999, 2015.
  38. 38.Locatello, F., Vincent, D., Tolstikhin, I., Rätsch, G., Gelly, S., and Schölkopf, B. Competitive training of mixtures of independent deep generative models. arXiv preprint arXiv:1804.11130, 2018.
  39. 39.Mathieu, M. F., Zhao, J. J., Zhao, J., Ramesh, A., Sprechmann, P., and LeCun, Y. Disentangling factors of variation in deep representation using adversarial training. In Advances in Neural Information Processing Systems, pp. 5040–5048, 2016.
  40. 40.Munch, E. The scream, 1893.
  41. 41.Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S. Visual reinforcement learning with imagined goals. In Advances in Neural Information Processing Systems, pp. 9209–9220, 2018.
  42. 42.Narayanaswamy, S., Paige, T. B., Van de Meent, J.-W., Desmaison, A., Goodman, N., Kohli, P., Wood, F., and Torr, P. Learning disentangled representations with semisupervised deep generative models. In Advances in Neural Information Processing Systems, pp. 5925–5935, 2017.
  43. 43.Nguyen, X., Wainwright, M. J., and Jordan, M. I. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory, 56(11):5847–5861, 2010.
  44. 44.Pearl, J. Causality. Cambridge university press, 2009.
  45. 45.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  46. 46.Peters, J., Janzing, D., and Schölkopf, B. Elements of causal inference: foundations and learning algorithms. MIT press, 2017.
  47. 47.Reed, S., Sohn, K., Zhang, Y., and Lee, H. Learning to disentangle factors of variation with manifold interaction. In International Conference on Machine Learning, pp. 1431–1439, 2014.
  48. 48.Reed, S. E., Zhang, Y., Zhang, Y., and Lee, H. Deep visual analogy-making. In Advances in Neural Information Processing Systems, pp. 1252–1260, 2015.
  49. 49.Ridgeway, K. and Mozer, M. C. Learning deep disentangled embeddings with the f-statistic loss. In Advances in Neural Information Processing Systems, pp. 185–194, 2018.
  50. 50.Rubenstein, P. K., Schoelkopf, B., and Tolstikhin, I. Learning disentangled representations with wasserstein autoencoders. In Workshop at International Conference on Learning Representations, 2018.
  51. 51.Schmidhuber, J. Learning factorial codes by predictability minimization. Neural Computation, 4(6):863–879, 1992.
  52. 52.Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. On causal and anticausal learning. In International Conference on Machine Learning, pp. 1255–1262, 2012.
  53. 53.Spirtes, P., Glymour, C., and Scheines, R. Causation, prediction, and search. Springer-Verlag. (2nd edition MIT Press 2000), 1993.
  54. 54.Steenbrugge, X., Leroux, S., Verbelen, T., and Dhoedt, B. Improving generalization for abstract reasoning tasks using disentangled feature representations. In Workshop on Relational Representation Learning at Conference on Neural Information Processing Systems, 2018.
  55. 55.Sugiyama, M., Suzuki, T., and Kanamori, T. Density-ratio matching under the bregman divergence: a unified framework of density-ratio estimation. Annals of the Institute of Statistical Mathematics, 64(5):1009–1044, 2012.
  56. 56.Suter, R., Miladinovic, Ð., Bauer, S., and Schölkopf, B. Interventional robustness of deep latent variable models. arXiv preprint arXiv:1811.00007, 2018.
  57. 57.Thomas, V., Bengio, E., Fedus, W., Pondard, J., Beaudoin, P., Larochelle, H., Pineau, J., Precup, D., and Bengio, Y. Disentangling the independently controllable factors of variation by interacting with the world. In Workshop on Learning Disentangled Representations at the 31st Conference on Neural Information Processing Systems, 2017.
  58. 58.Tschannen, M., Bachem, O., and Lucic, M. Recent advances in autoencoder-based representation learning. arXiv preprint arXiv:1812.05069, 2018.
  59. 59.Watanabe, S. Information theoretical analysis of multivariate correlation. IBM Journal of research and development, 4(1):66–82, 1960.
  60. 60.Whitney, W. F., Chang, M., Kulkarni, T., and Tenenbaum, J. B. Understanding visual concepts with continuation learning. In Workshop at International Conference on Learning Representations, 2016.
  61. 61.Yang, J., Reed, S. E., Yang, M.-H., and Lee, H. Weaklysupervised disentangling with recurrent transformations for 3d view synthesis. In Advances in Neural Information Processing Systems, pp. 1099–1107, 2015.
  62. 62.Yingzhen, L. and Mandt, S. Disentangled sequential autoencoder. In International Conference on Machine Learning, pp. 5656–5665, 2018.
  63. 63.Zhu, Z., Luo, P., Wang, X., and Tang, X. Multi-view perceptron: a deep model for learning face identity and view representations. In Advances in Neural Information Processing Systems, pp. 217–225, 2014.

Citation

MLA
Locatello, F., et al. “Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations”. Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 2018, http://arxiv.org/abs/1811.12359v4.
APA
Locatello, F., Bauer, S., Lucic, M., Rätsch, G., Gelly, S., Schölkopf, B., & Bachem, O. (2018). Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. Proceedings of the 36th International Conference on Machine Learning (ICML 2019). http://arxiv.org/abs/1811.12359v4
Chicago
Locatello, F., S. Bauer, M. Lucic, et al. 2018. “Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations”. Proceedings of the 36th International Conference on Machine Learning (ICML 2019). http://arxiv.org/abs/1811.12359v4.
Harvard
Locatello, F. et al. (2018) “Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations”, Proceedings of the 36th International Conference on Machine Learning (ICML 2019) [Preprint]. Available at: http://arxiv.org/abs/1811.12359v4.
Vancouver
1. Locatello F, Bauer S, Lucic M, Rätsch G, Gelly S, Schölkopf B, Bachem O (2018) Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. Proceedings of the 36th International Conference on Machine Learning (ICML 2019)

BibTeX

@article{locatello2018challenging,
  title = {Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations},
  author = {Locatello, Francesco and Bauer, Stefan and Lucic, Mario and Rätsch, Gunnar and Gelly, Sylvain and Schölkopf, Bernhard and Bachem, Olivier},
  year = {2018},
  journal = {Proceedings of the 36th International Conference on Machine Learning (ICML 2019)},
  url = {http://arxiv.org/abs/1811.12359v4},
  eprint = {1811.12359}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/