Desiderata for Representation Learning: A Causal Perspective

Yixin WangMichael I. Jordan

article2024JMLR106 citations

Establishes a rigorous causal inference framework that formalizes key representation learning desiderata—efficiency, non-spuriousness, and disentanglement—into calculable metrics and practical algorithms directly applicable to observational data.

Listen

Modern machine learning models often rely on representation learning to compress complex, high-dimensional inputs such as images and text into lower-dimensional feature vectors. Current deep learning approaches frequently capture spurious correlations that fail to generalize across new environments or produce entangled representations where distinct real-world factors are mixed across dimensions. The article addresses the fundamental challenge of turning intuitive representational goals—specifically non-spuriousness, efficiency, and disentanglement—into mathematically rigorous, computable criteria that can be evaluated and optimized using only single observational datasets without requiring specialized data augmentations or auxiliary feature labels.

The main objective of the article is to establish a unified causal inference framework for representation learning that formalizes non-spuriousness and efficiency in supervised settings, and disentanglement in unsupervised settings, through counterfactual quantities and observable data properties. It demonstrates how these theoretical formulations yield concrete, calculable metrics and algorithms to extract robust, generalizable representations from standard observational data.

To achieve this, the article utilizes structural causal models to map representational properties to causal principles. In the supervised setting, features are treated as potential causes of target labels, enabling non-spuriousness and efficiency to be defined via Pearl's probabilities of causation—specifically the probability of sufficiency and the probability of necessity. To address the rank-degeneracy and overlap challenges common in high-dimensional data, the authors introduce identification strategies that pinpoint latent common causes using probabilistic factor models, leading to the supervised CAUSAL-REP algorithm. In the unsupervised setting, the article shows that causal disentanglement directly implies that representations must exhibit independent geometric support across dimensions. This insight yields an unsupervised metric called the Independence-of-Support Score (IOSS) and a corresponding regularizer for variational autoencoders. The proposed methods were systematically evaluated across synthetic benchmarks, benchmark image datasets (Colored MNIST, CelebA, dSprites, and MPI3D), and real-world sentiment analysis corpora.

The article demonstrates several key findings across both supervised and unsupervised regimes. First, probabilities of causation reliably isolate true causal drivers from spurious artifacts even when correlations in training sets reach up to 0.9. Second, supervised CAUSAL-REP maintains robust predictive performance on non-spurious test sets, matching or closely approaching theoretical maximum accuracy across image and text domains, whereas traditional neural network and linear baselines suffer severe degradation. Third, unsupervised CAUSAL-REP successfully distinguishes unique subjects without capturing irrelevant, correlated features such as background color. Fourth, the IOSS unsupervised metric correctly distinguishes causally disentangled from entangled feature pairs in 89% to 99% of test cases, substantially outperforming existing metrics like Total Correlation and Wasserstein Dependency. Finally, incorporating an IOSS penalty into standard autoencoders consistently increases disentanglement as regularization scales without compromising overall model fit or data informativeness.

These findings indicate that causal reasoning provides a practical, principled foundation for building machine learning systems that generalize reliably to new distributions and remain interpretable. By enabling models to disregard misleading statistical correlations and untangle mixed factors using only single datasets, these approaches substantially reduce the operational risks and performance failures associated with deploying deep learning in critical domains like computer vision and automated text analysis.

Organizations developing models under domain shifts or requirements for explainability should adopt causal evaluation metrics and integrate independent-support regularizers or probabilities-of-causation objectives into their training pipelines. For immediate implementation, engineering teams can use latent factor models to extract unobserved background confounders before predicting target variables. Further validation in production pipelines is recommended to establish robust hyperparameters for latent factor dimensions and regularization weights across specialized operational datasets.

A primary limitation of this framework is its reliance on theoretical identification assumptions, including latent pinpointability and positivity overlap. In extreme scenarios where spurious and causal features are perfectly correlated, or when causal features directly govern pixel-level background textures, the algorithms may fail to separate them or may absorb relevant signals into the latent confounder. Confidence remains high that the methods provide robust, statistically sound improvements over purely correlation-based baselines under standard observational conditions.

arXiv: 2109.03795

No sufficiently relevant recommendations were found.

Cover for Desiderata for Representation Learning: A Causal Perspective

Table of Contents

  • 1. Introduction
  • 2. Supervised Representation Learning: Efficiency and Non-spuriousness
  • 2.1 Defining Efficiency and Non-spuriousness using Counterfactuals: Criteria
  • 2.1.1 A structural causal model of supervised representation learning
  • 2.1.2 Counterfactual definitions of non-spuriousness and efficiency
  • 2.1.3 Simultaneous assessment of non-spuriousness and efficiency
  • 2.2 Identifying Efficiency and Non-spuriousness in Observational Datasets: Identification
  • 2.2.1 From counterfactuals to interventional distributions
  • 2.2.2 From interventional distributions to observational data distributions
  • 2.3 Measuring efficiency and non-spuriousness in practice: Estimation
  • 2.4 CAUSAL-REP: Learning Efficient and Non-spurious Representations
  • 2.4.1 Representation learning as finding necessary and sufficient causes
  • 2.4.2 CAUSAL-REP and the linear example continued
  • 2.4.3 Extending CAUSAL-REP to unsupervised settings
  • 2.5 Empirical Studies of CAUSAL-REP
  • 2.5.1 How well do probabilities of causation measure efficiency and non-spuriousness of features?
  • 2.5.2 Does supervised CAUSAL-REP pick up spurious features in synthetic data?
  • 2.5.3 Does supervised CAUSAL-REP produce non-spurious representations for image data? A study on Colored MNIST and CelebA
  • 2.5.4 Does supervised CAUSAL-REP produce non-spurious representations for text data? A study on reviews corpora and sentiment analysis
  • 2.5.5 How well does unsupervised CAUSAL-REP perform on instance discrimination? A study on colored and shifted MNIST
  • 3. Unsupervised Representation Learning: Disentanglement
  • 3.1 The Causal Definition of Disentanglement
  • 3.2 Measuring Disentanglement with Observational Data
  • 3.2.1 Observable implications of disentanglement: The independent support condition
  • 3.2.2 Assessing disentanglement with the IOSS
  • 3.3 Learning Disentangled Representations with IOSS
  • 3.4 Empirical Studies of IOSS
  • 3.4.1 Can IOSS distinguish entangled and disentangled representations?
  • 3.4.2 Does the IOSS regularizer encourage disentangled representations?
  • 4. Discussion
  • Acknowledgments
  • References
  • Supplementary Materials
  • Appendix A. Proof of Theorem 1
  • Appendix B. The definition of functional interventions recovers backdoor adjustment
  • Appendix C. Proof of Proposition 1
  • Proof
  • Appendix D. Pinpointing the unobserved common cause C
  • Appendix E. Calculating PNS lower bounds with linear models
  • Appendix F. Details of the unsupervised CAUSAL-REP algorithm
  • Appendix G. Details of the empirical studies for CAUSAL-REP and additional empirical results
  • G.1 Details of Section 2.5.1
  • G.2 Details of the colored MNIST study
  • G.3 Details of the reviews corpora study
  • G.4 Details of the colored and shifted MNIST study
  • Appendix H. Details of Figure 8
  • Appendix I. Proof of Theorem 7
  • Appendix J. Details of empirical studies of IOSS and additional empirical results
  • J.1 Details of Section 3.4.1
  • J.2 Details of Section 3.4.2

Knowls

  1. Knowl 1 — Necessity and sufficiency define supervised representation quality

    definition

    Let ZZ be a possibly continuous, discrete, or vector-valued representation, let YY be the label, and consider an observed case with Z=zZ=z and Y=yY=y. Write Y(Z=z)Y(Z=z) for the counterfactual label under the functional intervention that makes the representation equal to zz; Y(Z≠z)Y(Z\ne z) denotes the counterfactual under the corresponding soft intervention that makes it differ from zz. The paper defines three probabilities of causation for the event I{Z=z}I\{Z=z\} as a cause of I{Y=y}I\{Y=y\}:

    • Probability of sufficiency (PS), for non-spuriousness: PSZ=z,Y=y=P(Y(Z=z)=y∣Z≠z, Y≠y).PS_{Z=z,Y=y}=P\bigl(Y(Z=z)=y\mid Z\ne z,\,Y\ne y\bigr). It measures how likely setting the representation to its observed value would produce the label among cases where both are absent.
    • Probability of necessity (PN), for efficiency: PNZ=z,Y=y=P(Y(Z≠z)≠y∣Z=z, Y=y).PN_{Z=z,Y=y}=P\bigl(Y(Z\ne z)\ne y\mid Z=z,\,Y=y\bigr). It measures how likely removing the represented feature would change the label among cases where both are present.
    • Probability of necessity and sufficiency (PNS), for their joint desideratum: PNSZ=z,Y=y=P(Y(Z≠z)≠y, Y(Z=z)=y).PNS_{Z=z,Y=y}=P\bigl(Y(Z\ne z)\ne y,\,Y(Z=z)=y\bigr).

    Thus high PS means the represented feature can produce the label, while high PN means it is needed for the label. PNS requires both counterfactual responses. For a dataset of nn cases (zi,yi)(z_i,y_i), the paper defines the joint dataset quantity as PNSn(Z,Y)=∏i=1nPNSZ=zi,Y=yiPNS_n(Z,Y)=\prod_{i=1}^n PNS_{Z=z_i,Y=y_i}.

  2. Knowl 2 — Observational identification yields a lower bound on supervised PNS

    theoretical result

    Consider a supervised structural causal model in which high-dimensional data X=(X1,…,Xm)X=(X_1,\ldots,X_m) causes label YY, there is no unobserved confounding between XX and YY, and an unobserved common cause CC may induce dependence among the components of XX. For a representation Z=f(X)=f~(XS)Z=f(X)=\tilde f(X_S) that depends only on a subset XS=(Xj)j∈SX_S=(X_j)_{j\in S}, the paper identifies the functional-intervention distribution from observations under three conditions:

    1. Pinpointability: P(C∣X)P(C\mid X) is a point mass at a deterministic function h(X)h(X), with hh known up to bijective transformation.
    2. Positivity: values of XSX_S have positive conditional probability given CC wherever they have positive marginal probability.
    3. Observability: XSX_S has support throughout its domain, rather than being rank-degenerate.

    Under these conditions, P(Y∣do(f(X)=z))=∫P(Y∣f(X)=z,h(X)) P(h(X)) dh(X).P(Y\mid do(f(X)=z))=\int P(Y\mid f(X)=z,h(X))\,P(h(X))\,dh(X). The counterfactual PNS is not generally point-identified from intervention distributions. Its lower bound is the difference between the outcome probabilities under setting the feature to the observed value and under the soft intervention that makes it differ: PNSZ=z,Y=y≥P(Y=y∣do(Z=z))−P(Y=y∣do(Z≠z)).PNS_{Z=z,Y=y}\ge P(Y=y\mid do(Z=z))-P(Y=y\mid do(Z\ne z)). For nn observed cases, substituting the identified distributions gives the product lower bound PNS‾n(f(X),Y)=∏i=1n∫ ⁣[P(Y=yi∣f(X)=f(xi),C)−P(Y=yi∣f(X)≠f(xi),C)]P(C) dC.\underline{PNS}_n(f(X),Y)=\prod_{i=1}^n\int\!\left[P(Y=y_i\mid f(X)=f(x_i),C)-P(Y=y_i\mid f(X)\ne f(x_i),C)\right]P(C)\,dC. The paper gives the analogous bound for a representation dimension conditional on the other dimensions. The full-support restriction addresses rank degeneracy: when high-dimensional XX lies on a lower-dimensional manifold, observational data do not determine P(Y∣X)P(Y\mid X) away from that manifold, so general functions of all of XX need not have identifiable intervention distributions.

  3. Knowl 3 — CAUSAL-REP optimizes conditional necessary-and-sufficient causes

    model/method

    CAUSAL-REP learns a supervised representation by maximizing the sum of log lower bounds on conditional PNS, so that each output dimension is useful for the label after holding the other representation dimensions fixed. For a representation f=(f1,…,fd)f=(f_1,\ldots,f_d) and data (xi,yi,ci)i=1n(x_i,y_i,c_i)_{i=1}^n, its objective is max⁡f  ∑j=1dlog⁡PNS‾n(fj(X),Y∣f−j(X))+λR(f),\max_f\;\sum_{j=1}^d\log\underline{PNS}_n\bigl(f_j(X),Y\mid f_{-j}(X)\bigr)+\lambda R(f), where f−jf_{-j} denotes all output dimensions except jj, and each PNS term is the observational lower bound under the identification conditions above. The regularizer is R(f)=1d∑j=1dlog⁡ ⁣(1−R2((fj(xi))i=1n;(ci)i=1n))−α∥W(f)∥22.R(f)=\frac1d\sum_{j=1}^d\log\!\left(1-R^2\bigl((f_j(x_i))_{i=1}^n;(c_i)_{i=1}^n\bigr)\right)-\alpha\lVert W(f)\rVert_2^2. Here R2R^2 is the sample coefficient of determination from regressing representation coordinate jj on the estimated common cause, W(f)W(f) denotes the parameters of the representation function, and λ,α≥0\lambda,\alpha\ge0 are regularization weights. The first penalty discourages a representation coordinate from being determined by CC, which would threaten positivity; the second penalizes dependence on many input coordinates, encouraging the low-dimensional input subset required for observability. Representation classes in the paper include normalized neural networks, convex combinations of input features, and feature-selection maps.

  4. Knowl 4 — Supervised CAUSAL-REP training and prediction procedure

    algorithm

    For labeled training observations (xi,yi)(x_i,y_i), a representation family ff, and a probabilistic factor model for XX and latent common cause CC, the supervised procedure is:

    1. Fit the factor model to training inputs and infer p(ci∣xi)p(c_i\mid x_i). Check whether each posterior is sufficiently concentrated to treat CC as pinpointed; if not, the paper's identification-based procedure does not apply.
    2. Set c^i\hat c_i to the posterior mean of CC given xix_i. Fit a model for P(Y∣f(X),C)P(Y\mid f(X),C) to (f(xi),c^i,yi)(f(x_i),\hat c_i,y_i).
    3. Optimize the CAUSAL-REP objective over representation parameters, using the fitted conditional outcome model to evaluate the PNS lower-bound terms. When the outcome-model fit has a closed-form solution, substitute it into the objective; otherwise alternate representation-gradient updates with outcome-model maximum-likelihood updates until the inner fit converges.
    4. Fit a separate prediction model for P(Y∣f^(X))P(Y\mid\hat f(X)) using the learned representation alone, and apply it to test inputs. The prediction model excludes CC, since CAUSAL-REP does not constrain the relationship between CC and YY to transfer.

    The paper's experiments use Adam with learning rate 0.010.01. The learned representation's out-of-distribution performance is assessed on test sets where spurious associations are weakened or reversed. The paper does not state a general computational-complexity bound.

  5. Knowl 5 — Unsupervised CAUSAL-REP turns augmented instance discrimination into supervised causes

    algorithm

    Suppose there are nn subjects, one raw observation per subject, and U≥2U\ge2 observations per subject after augmentation. Assign each augmented observation xiux_i^u a one-versus-all vector of pseudo-labels, with yisu=I{xiu belongs to subject s}y_{is}^u=I\{x_i^u\text{ belongs to subject }s\} for s=1,…,ns=1,\ldots,n. Fit a factor model to the augmented inputs, infer and pinpoint the common cause for each, then optimize the supervised conditional-PNS objective across every subject label and representation coordinate: max⁡f  ∑s=1n∑j=1dlog⁡PNS‾nU(fj(X),Ys∣f−j(X))+λR(f).\max_f\;\sum_{s=1}^n\sum_{j=1}^d\log\underline{PNS}_{nU}\bigl(f_j(X),Y_s\mid f_{-j}(X)\bigr)+\lambda R(f). The representation is therefore trained to distinguish each subject from all other subjects; the subject IDs are training targets derived from augmentation, not externally supplied semantic labels. The method shares the instance-discrimination goal of contrastive learning, but is not generally equivalent to it: the paper says the objectives coincide only when input dimensions are independent. With a shared unobserved cause, contrastive learning may retain a correlated feature that does not cause the subject label, whereas the CAUSAL-REP objective targets necessary and sufficient causes. The authors' unsupervised experiments evaluate the resulting representation by fitting a downstream label predictor after representation learning.

  6. Knowl 6 — Causal disentanglement permits correlation but excludes causal links between coordinates

    definition

    A representation Z=(Z1,…,Zd)Z=(Z_1,\ldots,Z_d) is causally disentangled in the paper's structural causal model if its coordinates encode features that jointly generate the observed object X=(X1,…,Xm)X=(X_1,\ldots,X_m) but do not causally affect one another. The model is C←UC,Zj←fjz(C,Uz,j),j=1,…,d,Xl←flx(Z,UX),l=1,…,m.C\leftarrow U_C,\qquad Z_j\leftarrow f_j^z(C,U_{z,j}),\quad j=1,\ldots,d,\qquad X_l\leftarrow f_l^x(Z,U_X),\quad l=1,\ldots,m. Here CC is an unobserved common cause that may affect the representation coordinates and make them statistically correlated; UCU_C, Uz,jU_{z,j}, and UXU_X are exogenous disturbances. Thus statistical independence is not part of this causal definition: distinct coordinates may be correlated through CC while remaining disentangled.

  7. Knowl 7 — Causal disentanglement implies independent support under positivity

    theoretical result

    For the causal model of a disentangled representation Z=(Z1,…,Zd)Z=(Z_1,\ldots,Z_d), assume the common cause satisfies the positivity condition that, for every coordinate jj, the support of ZjZ_j given CC agrees with its marginal support: P(Zj∣C)>0P(Z_j\mid C)>0 if and only if P(Zj)>0P(Z_j)>0. Then, for distinct coordinates j,j′j,j', at values zj′z_{j'} with positive density, the interventional and observational conditional supports agree: supp⁡(Zj∣do(Zj′=zj′))=supp⁡(Zj∣Zj′=zj′).\operatorname{supp}(Z_j\mid do(Z_{j'}=z_{j'}))=\operatorname{supp}(Z_j\mid Z_{j'}=z_{j'}). Consequently the joint support factors across coordinates, supp⁡(Z1,…,Zd)=∏j=1dsupp⁡(Zj),supp⁡(Zj∣ZS)=supp⁡(Zj)\operatorname{supp}(Z_1,\ldots,Z_d)=\prod_{j=1}^d\operatorname{supp}(Z_j),\qquad \operatorname{supp}(Z_j\mid Z_S)=\operatorname{supp}(Z_j) for every subset SS not containing jj (on conditioning values in the support). In geometric terms, compactly supported coordinates have a hyperrectangular joint support. They can still be correlated within that region. This implication is one-way: independent support is an observable necessary condition under the stated positivity assumption, not by itself proof of causal disentanglement. The pairwise-support visualizations on printed page 35 illustrate that correlated ground-truth disentangled factors retain rectangular support, while nonlinear entangled combinations can have curved support.

  8. Knowl 8 — IOSS measures the gap between observed and factorized support

    definition

    For a bounded representation Z=(Z1,…,Zd)Z=(Z_1,\ldots,Z_d) with nonzero range in every coordinate, standardize each coordinate as Zˉj=(Zj−inf⁡Zj)/(sup⁡Zj−inf⁡Zj)\bar Z_j=(Z_j-\inf Z_j)/(\sup Z_j-\inf Z_j). The independence-of-support score (IOSS) is the Hausdorff distance between the joint support and the product of the marginal supports: IOSS(Z)=dH ⁣(supp⁡(Zˉ1,…,Zˉd),  ∏j=1dsupp⁡(Zˉj)).IOSS(Z)=d_H\!\left(\operatorname{supp}(\bar Z_1,\ldots,\bar Z_d),\;\prod_{j=1}^d\operatorname{supp}(\bar Z_j)\right). The score is scale-invariant; zero means the support factors, and larger values indicate a larger gap from independent support. For continuous representations, the paper estimates it from nn observations by rescaling each coordinate to the unit interval and sampling KK points UkU_k uniformly from the unit hypercube: IOSS^=max⁡1≤k≤Kmin⁡1≤i≤n∥Uk−zˉi∥22.\widehat{IOSS}=\max_{1\le k\le K}\min_{1\le i\le n}\lVert U_k-\bar z_i\rVert_2^2. This estimates the distance from the product-support region to the observed joint support. The paper uses K=10dnK=10^d n for its illustrated sample calculations, suggests replacing extreme max/min by interior quantiles for robustness, and states that the sample score converges almost surely to IOSS as n→∞n\to\infty and K/n→∞K/n\to\infty. For unsupervised learning, it adds the sample IOSS as a penalty, minimizing −L+λIOSS^-L+\lambda\widehat{IOSS}, where LL is the base representation-learning objective and λ>0\lambda>0 controls the penalty; the paper illustrates this with a VAE and uses posterior means as representations.

  9. Knowl 9 — CAUSAL-REP resists spurious associations in image and text tests

    data/table

    The paper evaluates whether learned representations retain predictive performance when a spurious training association is changed. In CelebA, target and spurious face attributes are highly correlated in training and less correlated or reversed in the test set; the entries below are test accuracy (mean with parenthetical value as reported), ordered CAUSAL-REP / direct neural-network fit / VAE representation. In sentiment experiments, the test sets remove injected spurious words; entries are accuracy ordered logistic regression in-distribution / CAUSAL-REP in-distribution / logistic regression non-spuriousness test / CAUSAL-REP non-spuriousness test.

    CelebA target / spurious attributeTrain correlationTest correlationAccuracy: CAUSAL-REP / NN / VAE
    Arched Brows / Eye Bags0.784-0.7970.539(0.036) / 0.514(0.029) / 0.499(0.012)
    Arched Brows / Earrings0.799-0.7910.521(0.025) / 0.504(0.022) / 0.494(0.008)
    Attractive / Necklace0.793-0.7870.537(0.022) / 0.505(0.030) / 0.485(0.018)
    Black Hair / Mouth Open0.791-0.7950.594(0.033) / 0.566(0.060) / 0.505(0.010)
    Goatee / Male0.8890.0530.728(0.073) / 0.566(0.102) / 0.867(0.165)
    Mustache / Black Hair0.764-0.7780.525(0.023) / 0.512(0.037) / 0.540(0.005)
    Mustache / Male0.8920.0880.787(0.066) / 0.555(0.097) / 0.855(0.212)
    Sentiment datasetIn-distribution: logistic / CAUSAL-REPNon-spuriousness test: logistic / CAUSAL-REP
    IMDB-L0.669 / 0.6450.591 / 0.642
    IMDB-S0.836 / 0.6820.570 / 0.621
    Kindle0.850 / 0.6180.468 / 0.572

    CAUSAL-REP outperforms direct neural-network fitting in all seven listed CelebA non-spuriousness tests, though VAE is better for the Goatee/Male and Mustache/Male pairs. In each listed sentiment dataset, CAUSAL-REP is more accurate than logistic regression on the non-spuriousness test, while logistic regression is better on the in-distribution test. The results support the claim that CAUSAL-REP can reduce reliance on spurious cues, with exceptions associated in the paper with positivity violations.

  10. Knowl 10 — IOSS separates entangled from disentangled representations in benchmark comparisons

    empirical result

    On mpi3d, smallnorb, dSprites, and cars3d, the authors subsample data so ground-truth features are correlated, then compare each disentangled representation with a nonlinear entangled counterpart. The reported metric is the proportion of paired representations correctly ordered as disentangled versus entangled (higher is better); values are reproduced as reported, including parenthetical quantities:

    Metricmpi3dsmallnorbdSpritescars3d
    IOSS0.998(0.045)0.968(0.176)0.980(0.140)0.892(0.311)
    Total correlation0.858(0.349)0.070(0.255)0.162(0.369)0.090(0.286)
    Wasserstein dependency0.956(0.205)0.478(0.500)0.310(0.463)0.964(0.186)
    Intervention robustness (oracle, uses ground-truth features)1.000(0.000)1.000(0.000)1.000(0.000)1.000(0.000)

    IOSS has the highest score among the compared unsupervised metrics on mpi3d, smallnorb, and dSprites, and is competitive on cars3d, where Wasserstein dependency scores higher. The oracle metric reaches 1.000 on every dataset but uses ground-truth features, unlike the unsupervised metrics. The benchmark comparison and the support visualizations on printed pages 43 and 35 respectively show the intended distinction: support geometry can reveal entanglement even when correlations alone do not.

  11. Knowl 11 — IOSS regularization improves causal disentanglement in VAE representations

    empirical result

    The paper adds IOSS regularization to VAEs trained on mpi3d, smallnorb, dSprites, and cars3d with correlated ground-truth factors. It evaluates learned representations using the intervention robustness score, where higher indicates better causal disentanglement. The reported comparison finds VAE+IOSS better than classical VAE, FactorVAE, betaVAE, and betaTCVAE in this correlated-factor setting. Across the tested datasets, stronger IOSS regularization increases intervention robustness; the reported VAE predictive negative log-likelihood remains stable as the penalty increases, so the experiments show no evident informativeness–disentanglement tradeoff in these runs. The plotted support patterns on printed page 65 also become more nearly rectangular with IOSS regularization, consistent with the penalty directly encouraging independent support.

  12. Knowl 12 — Identification and support-based metrics have substantive limits

    limitation

    CAUSAL-REP's observational PNS lower bounds rely on pinpointability of the latent common cause and on positivity and observability for the input subset used by the representation. If a feature is perfectly correlated with another feature through the common cause, positivity can fail; the paper notes that CAUSAL-REP may then exclude both the spurious and non-spurious feature because neither PNS lower bound is identifiable. Pinpointing CC can also absorb a necessary-and-sufficient feature when that feature drives strong correlations among the inputs, so the algorithm may miss such features. Moreover, functional-intervention counterfactuals depend on the population distribution used to define the intervention and are not automatically invariant under covariate shift; the causal formulation alone therefore does not guarantee out-of-distribution improvement. For disentanglement, independent support is only a necessary observable implication under positivity, not a sufficient characterization; the proposed IOSS is designed for bounded representations with nonzero coordinate ranges.

Coverage note — The paper's detailed derivations, secondary synthetic examples, and supplementary plots are omitted because they do not add independent load-bearing results beyond the identification conditions, algorithms, limitations, and representative empirical comparisons captured here.

References

  1. 1.Achille, A. & Soatto, S. (2018). Emergence of invariance and disentanglement in deep representations. Journal of Machine Learning Research, 19(1), 1947–1980.
  2. 2.Ahuja, K., Mahajan, D., Wang, Y., & Bengio, Y. (2023). Interventional causal representation learning. In International Conference on Machine Learning (ICML) (pp. 372–407).: PMLR.
  3. 3.Ahuja, K., Mansouri, A., & Wang, Y. (2024). Multi-domain causal representation learning via weak distributional invariances. Artificial Intelligence and Statistics (AISTATS).
  4. 4.Airoldi, E. M., Blei, D. M., Fienberg, S. E., & Xing, E. P. (2008). Mixed membership stochastic blockmodels. Journal of Machine Learning Research, 9, 1981–2014.
  5. 5.Arjovsky, M., Bottou, L., Gulrajani, I., & Lopez-Paz, D. (2019). Invariant risk minimization. arXiv preprint arXiv:1907.02893.
  6. 6.Bai, J. & Li, K. (2016). Maximum likelihood estimation and inference for approximate factor models of high dimension. Review of Economics and Statistics, 98(2), 298–309.
  7. 7.Bengio, Y., Courville, A., & Vincent, P. (2013). Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8), 1798–1828.
  8. 8.Blei, D. M., Kucukelbir, A., & McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American Statistical Association, 112(518), 859–877.
  9. 9.Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3, 993–1022.
  10. 10.Bouchacourt, D., Tomioka, R., & Nowozin, S. (2018). Multi-level variational autoencoder: Learning disentangled representations from grouped observations. In AAAI Conference on Artificial Intelligence, volume 32.
  11. 11.Burgess, C. P., Higgins, I., et al. (2018). Understanding disentangling in beta-VAE. arXiv preprint arXiv:1804.03599.
  12. 12.Chalupka, K., Eberhardt, F., & Perona, P. (2017). Causal feature learning: an overview. Behaviormetrika, 44(1), 137–164.
  13. 13.Chalupka, K., Perona, P., & Eberhardt, F. (2015). Visual causal feature learning. Uncertainty in Artificial Intelligence (UAI), (pp. 181–190).
  14. 14.Chen, R. T., Li, X., Grosse, R., & Duvenaud, D. (2018). Isolating sources of disentanglement in VAEs. In Advances in Neural Information Processing Systems (pp. 2615–2625).
  15. 15.Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020a). A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML) (pp. 1597–1607).
  16. 16.Chen, Y., Li, X., & Zhang, S. (2020b). Structured latent factor analysis for large-scale data: Identifiability, estimability, and their implications. Journal of the American Statistical Association, 115(532), 1756–1770.
  17. 17.Cheng, P. W. & Lu, H. (2017). Causal invariance as an essential constraint for creating a causal representation of the world: Generalizing. The Oxford Handbook of Causal Reasoning, (pp.˜65).
  18. 18.Comon, P. (1994). Independent component analysis, a new concept? Signal Processing, 36(3), 287–314.
  19. 19.Correa, J. & Bareinboim, E. (2020). A calculus for stochastic interventions: Causal effect identification and surrogate experiments. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34 (pp. 10093–10100).
  20. 20.Creager, E., Jacobsen, J.-H., & Zemel, R. (2021). Environment inference for invariant learning. In International Conference on Machine Learning (ICML) (pp. 2189–2200).: PMLR.
  21. 21.D’Amour, A. (2019a). Comment: Reflections on the deconfounder. Journal of the American Statistical Association, 114(528), 1597–1601.
  22. 22.D’Amour, A. (2019b). On multi-cause approaches to causal inference with unobserved counfounding: Two cautionary failure cases and a promising alternative. In Artificial Intelligence and Statistics (AISTATS) (pp. 3478–3486).
  23. 23.D’Amour, A., Ding, P., Feller, A., Lei, L., & Sekhon, J. (2020a). Overlap in observational studies with high-dimensional covariates. Journal of Econometrics, 221, 644–654.
  24. 24.D’Amour, A., Heller, K., et al. (2020b). Underspecification presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395.
  25. 25.Deng, L. (2012). The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29(6), 141–142.
  26. 26.Eberhardt, F. & Scheines, R. (2007). Interventions and causal inference. Philosophy of Science, 74(5), 981–995.
  27. 27.Erosheva, E. A. & Fienberg, S. E. (2005). Bayesian mixed membership models for soft clustering and classification. In C. Weihs & W. Gaul (Eds.), Classification—The Ubiquitous Challenge (pp. 11–26). Berlin, Heidelberg: Springer Berlin Heidelberg.
  28. 28.Fong, C. & Grimmer, J. (2016). Discovery of treatments from text corpora. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1600–1609).
  29. 29.Galhotra, S., Pradhan, R., & Salimi, B. (2021). Explaining black-box algorithms using probabilistic contrastive counterfactuals. In Proceedings of the 2021 International Conference on Management of Data (pp. 577–590).
  30. 30.Gelman, A. & Imbens, G. (2013). Why ask why? Forward causal inference and reverse causal questions. Technical report, National Bureau of Economic Research.
  31. 31.Gelman, A., Meng, X.-L., & Stern, H. (1996). Posterior predictive assessment of model fitness via realized discrepancies. Statistica Sinica, (pp. 733–760).
  32. 32.Gilks, W. R., Richardson, S., & Spiegelhalter, D. (1995). Markov Chain Monte Carlo in Practice. CRC Press.
  33. 33.Golub, G., Klema, V., & Stewart, G. W. (1976). Rank degeneracy and least squares problems. Technical report, Stanford University, Department of Computer Science.
  34. 34.Goodfellow, I. J., Pouget-Abadie, J., et al. (2014). Generative adversarial networks. arXiv preprint arXiv:1406.2661.
  35. 35.Grimmer, J. & Fong, C. (2021). Causal inference with latent treatments. American Journal of Political Science.
  36. 36.Hadsell, R., Chopra, S., & LeCun, Y. (2006). Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2 (pp. 1735–1742).: IEEE.
  37. 37.Heinze-Deml, C., Maathuis, M. H., & Meinshausen, N. (2018). Causal structure learning. Annual Review of Statistics and Its Application, 5, 371–391.
  38. 38.Higgins, I., Matthey, L., et al. (2017). Beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations.
  39. 39.Hosoya, H. (2019). Group-based learning of disentangled representations with generalizability for novel contents. In International Joint Conference on Artificial Intelligence (IJCAI) (pp. 2506–2513).
  40. 40.Imai, K. & Jiang, Z. (2019). Comment: The challenges of multiple causes. Journal of the American Statistical Association, 114(528), 1605–1610.
  41. 41.Imbens, G. W. & Rubin, D. B. (2015). Causal Inference in Statistics, Social, and Biomedical Sciences. Cambridge University Press.
  42. 42.Janzing, D. & Schölkopf, B. (2015). Semi-supervised interpolation in an anticausal learning scenario. Journal of Machine Learning Research, 16(1), 1923–1948.
  43. 43.Johansson, F. D., Shalit, U., Kallus, N., & Sontag, D. (2020). Generalization bounds and representation learning for estimation of potential outcomes and causal effects. arXiv preprint arXiv:2001.07426.
  44. 44.Johansson, F. D., Sontag, D., & Ranganath, R. (2019). Support and invertibility in domain-invariant representations. In Artificial Intelligence and Statistics (AISTATS) (pp. 527–536).
  45. 45.Khasanova, R. & Frossard, P. (2017). Graph-based isometry invariant representation learning. In International Conference on Machine Learning (ICML) (pp. 1847–1856).
  46. 46.Khemakhem, I., Kingma, D., Monti, R., & Hyvarinen, A. (2020). Variational autoencoders and nonlinear ICA: A unifying framework. In Artificial Intelligence and Statistics (AISTATS) (pp. 2207–2217).
  47. 47.Kilbertus, N., Parascandolo, G., & Schölkopf, B. (2018). Generalization in anti-causal learning. arXiv preprint arXiv:1812.00524.
  48. 48.Kim, H. & Mnih, A. (2018). Disentangling by factorising. In International Conference on Machine Learning (ICML) (pp. 2649–2658).
  49. 49.Kingma, D. P. & Welling, M. (2014). Auto-encoding variational Bayes. International Conference on Learning Representations.
  50. 50.Kommiya Mothilal, R., Mahajan, D., Tan, C., & Sharma, A. (2021). Towards unifying feature attribution and counterfactual explanations: Different means to the same end. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (pp. 652–663).
  51. 51.Kumar, A., Sattigeri, P., & Balakrishnan, A. (2018). Variational inference of disentangled latent concepts from unlabeled observations. In International Conference on Learning Representations.
  52. 52.Liu, Z., Luo, P., Wang, X., & Tang, X. (2015). Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV).
  53. 53.Locatello, F., Bauer, S., et al. (2019a). Challenging common assumptions in the unsupervised learning of disentangled representations. In International Conference on Machine Learning (ICML) (pp. 4114–4124).
  54. 54.Locatello, F., Poole, B., et al. (2020a). Weakly-supervised disentanglement without compromises. In International Conference on Machine Learning (ICML) (pp. 6348–6359).
  55. 55.Locatello, F., Tschannen, M., et al. (2019b). Disentangling factors of variations using few labels. In International Conference on Learning Representations.
  56. 56.Locatello, F., Weissenborn, D., et al. (2020b). Object-centric learning with slot attention. Advances in Neural Information Processing Systems, 33, 11525–11538.
  57. 57.Lu, C., Wu, Y., Hernández-Lobato, J. M., & Schölkopf, B. (2021). Invariant causal representation learning.
  58. 58.McLachlan, G. J. & Basford, K. E. (1988). Mixture Models: Inference and Applications to Clustering, volume 38. M. Dekker New York.
  59. 59.Mitrovic, J., McWilliams, B., Walker, J., Buesing, L., & Blundell, C. (2020). Representation learning via invariant causal mechanisms. arXiv preprint arXiv:2010.07922.
  60. 60.Moraffah, R., Shu, K., Raglin, A., & Liu, H. (2019). Deep causal representation learning for unsupervised domain adaptation. arXiv preprint arXiv:1910.12417.
  61. 61.Moyer, D., Gao, S., Brekelmans, R., Steeg, G. V., & Galstyan, A. (2018). Invariant representations without adversarial training. arXiv preprint arXiv:1805.09458.
  62. 62.Mueller, S., Li, A., & Pearl, J. (2021). Causes of effects: Learning individual responses from population data. arXiv preprint arXiv:2104.13730.
  63. 63.Nabi, R., McNutt, T., & Shpitser, I. (2020). Semiparametric causal sufficient dimension reduction of high dimensional treatments. arXiv preprint arXiv:1710.06727.
  64. 64.Parascandolo, G., Kilbertus, N., Rojas-Carulla, M., & Schölkopf, B. (2018). Learning independent causal mechanisms. In International Conference on Machine Learning (ICML) (pp. 4036–4044).
  65. 65.Paul, M. (2017). Feature selection as causal inference: Experiments with text classification. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017) (pp. 163–172).
  66. 66.Pearl, J. (1995). Causal diagrams for empirical research. Biometrika, 82(4), 669–688.
  67. 67.Pearl, J. (2011). Causality: Models, Reasoning, and Inference. Cambridge University Press.
  68. 68.Pearl, J. (2019a). The seven tools of causal inference, with reflections on machine learning. Communications of the ACM, 62(3), 54–60.
  69. 69.Pearl, J. (2019b). Sufficient causes: On oxygen, matches, and fires. Journal of Causal Inference, 7(2), 1–8.
  70. 70.Pritchard, J. K., Stephens, M., & Donnelly, P. (2000). Inference of population structure using multilocus genotype data. Genetics, 155(2), 945–959.
  71. 71.Pryzant, R., Card, D., Jurafsky, D., Veitch, V., & Sridhar, D. (2020). Causal effects of linguistic properties. arXiv preprint arXiv:2010.12919.
  72. 72.Puli, A., Perotte, A., & Ranganath, R. (2020). Causal estimation with functional confounders. Advances in Neural Information Processing Systems, 33.
  73. 73.Puli, A., Zhang, L. H., Oermann, E. K., & Ranganath, R. (2021). Predictive modeling in the presence of nuisance-induced spurious correlations. arXiv preprint arXiv:2107.00520.
  74. 74.Ranganath, R. & Perotte, A. (2018). Multiple causal inference with latent confounding. arXiv preprint arXiv:1805.08273.
  75. 75.Rezende, D. J., Mohamed, S., & Wierstra, D. (2014). Stochastic backpropagation and variational inference in deep latent gaussian models. In International Conference on Machine Learning (ICML), volume 2.
  76. 76.Roth, K., Ibrahim, M., Akata, Z., Vincent, P., & Bouchacourt, D. (2022). Disentanglement of correlated factors via Hausdorff factorized support. arXiv preprint arXiv:2210.07347.
  77. 77.Schölkopf, B., Janzing, D., et al. (2012). On causal and anticausal learning. In International Conference on Machine Learning (ICML) (pp. 1255–1262).
  78. 78.Schölkopf, B., Janzing, D., et al. (2013). Semi-supervised learning in causal and anticausal settings. In Empirical Inference (pp. 129–141). Springer.
  79. 79.Schölkopf, B., Locatello, F., et al. (2021). Toward causal representation learning. Proceedings of the IEEE, 109(5), 612–634.
  80. 80.Shen, X., Liu, F., et al. (2020). Disentangled generative causal representation learning. arXiv preprint arXiv:2010.02637.
  81. 81.Shi, C., Veitch, V., & Blei, D. (2020). Invariant representation learning for treatment effect estimation. arXiv preprint arXiv:2011.12379.
  82. 82.Shu, R., Chen, Y., Kumar, A., Ermon, S., & Poole, B. (2019). Weakly supervised disentanglement with guarantees. In International Conference on Learning Representations.
  83. 83.Stewart, G. (1984). Rank degeneracy. SIAM Journal on Scientific and Statistical Computing, 5(2), 403–413.
  84. 84.Suter, R., Miladinovic, D., Schölkopf, B., & Bauer, S. (2019). Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness. In International Conference on Machine Learning (ICML) (pp. 6056–6065).
  85. 85.Thomas, V., Bengio, E., et al. (2018). Disentangling the independently controllable factors of variation by interacting with the world. arXiv preprint arXiv:1802.09484.
  86. 86.Thomas, V., Pondard, J., et al. (2017). Independently controllable factors. arXiv preprint arXiv:1708.01289.
  87. 87.Tian, J. & Pearl, J. (2000). Probabilities of causation: Bounds and identification. Annals of Mathematics and Artificial Intelligence, 28(1), 287–313.
  88. 88.Tipping, M. E. & Bishop, C. M. (1999). Probabilistic principal component analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61(3), 611–622.
  89. 89.Träuble, F., Creager, E., et al. (2020). Is independence all you need? On the generalization of representations learned from correlated data. arXiv preprint arXiv:2006.07886.
  90. 90.Udell, M. & Townsend, A. (2019). Why are big data matrices approximately low rank? SIAM Journal on Mathematics of Data Science, 1(1), 144–160.
  91. 91.Veitch, V., D’Amour, A., Yadlowsky, S., & Eisenstein, J. (2021). Counterfactual invariance to spurious correlations: Why and how to pass stress tests. arXiv preprint arXiv:2106.00545.
  92. 92.Veitch, V., Sridhar, D., & Blei, D. (2020). Adapting text embeddings for causal inference. In Uncertainty in Artificial Intelligence (UAI) (pp. 919–928).
  93. 93.Veitch, V., Wang, Y., & Blei, D. M. (2019). Using embeddings to correct for unobserved confounding in networks. arXiv preprint arXiv:1902.04114.
  94. 94.Wang, H., Lu, Y., & Zhai, C. (2010). Latent aspect rating analysis on review text data: A rating regression approach. In Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 783–792).
  95. 95.Wang, H., Lu, Y., & Zhai, C. (2011). Latent aspect rating analysis without aspect keyword supervision. In Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 618–626).
  96. 96.Wang, Y. & Blei, D. (2021). A proxy variable view of shared confounding. In International Conference on Machine Learning (ICML) (pp. 10697–10707).: PMLR.
  97. 97.Wang, Y. & Blei, D. M. (2019a). The blessings of multiple causes. Journal of the American Statistical Association, 114(528), 1574–1596.
  98. 98.Wang, Y. & Blei, D. M. (2019b). Frequentist consistency of variational Bayes. Journal of the American Statistical Association, 114(527), 1147–1161.
  99. 99.Wang, Y. & Blei, D. M. (2020). Towards clarifying the theory of the deconfounder. arXiv preprint arXiv:2003.04948.
  100. 100.Wang, Z. & Culotta, A. (2020). Identifying spurious correlations for robust text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings (pp. 3431–3440).
  101. 101.Wang, Z. & Culotta, A. (2021). Robustness to spurious correlations in text classification via automatically generated counterfactuals. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35 (pp. 14024–14031).
  102. 102.Watson, D., Gultchin, L., Taly, A., & Floridi, L. (2021). Local explanations via necessity and sufficiency: unifying theory and practice. arXiv preprint arXiv:2103.14651.
  103. 103.Weichwald, S., Schölkopf, B., Ball, T., & Grosse-Wentrup, M. (2014). Causal and anti-causal learning in pattern recognition for neuroimaging. In 2014 International Workshop on Pattern Recognition in Neuroimaging (pp. 1–4).
  104. 104.Wu, P. & Fukumizu, K. (2021). Identifying treatment effects under unobserved confounding by causal representation learning. arXiv preprint arXiv:2101.06662.
  105. 105.Xiao, Y. & Wang, W. Y. (2019). Disentangled representation learning with Wasserstein total correlation. arXiv preprint arXiv:1912.12818.
  106. 106.Yang, M., Liu, F., et al. (2020). CausalVAE: Disentangled representation learning via neural structural causal models. arXiv preprint arXiv:2004.08697.
  107. 107.Zhao, H., Des Combes, R. T., Zhang, K., & Gordon, G. (2019). On learning invariant representations for domain adaptation. In International Conference on Machine Learning (ICML) (pp. 7523–7532).

Citation

MLA
Wang, Y., and M. I. Jordan. “Desiderata for Representation Learning: A Causal Perspective”. Journal of Machine Learning Research, vol. 25, no. 275, 2024, pp. 1–5, https://www.jmlr.org/papers/v25/21-107.html.
APA
Wang, Y., & Jordan, M. I. (2024). Desiderata for Representation Learning: A Causal Perspective. Journal of Machine Learning Research, 25(275), 1–65. https://www.jmlr.org/papers/v25/21-107.html
Chicago
Wang, Y., and M. I. Jordan. 2024. “Desiderata for Representation Learning: A Causal Perspective”. Journal of Machine Learning Research 25 (275): 1–65. https://www.jmlr.org/papers/v25/21-107.html.
Harvard
Wang, Y. and Jordan, M.I. (2024) “Desiderata for Representation Learning: A Causal Perspective”, Journal of Machine Learning Research, 25(275), pp. 1–65. Available at: https://www.jmlr.org/papers/v25/21-107.html.
Vancouver
1. Wang Y, Jordan MI (2024) Desiderata for Representation Learning: A Causal Perspective. Journal of Machine Learning Research 25:1–65

BibTeX

@article{JMLR:v25:21-107,
  author  = {Yixin Wang and Michael I. Jordan},
  title   = {Desiderata for Representation Learning: A Causal Perspective},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {275},
  pages   = {1--65},
  url     = {http://jmlr.org/papers/v25/21-107.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/