Interventional Causal Representation Learning

Kartik AhujaDivyat MahajanYixin WangYoshua Bengio

article2023ICML179 citations

Proves that interventional data enables provable identification of latent causal factors without parametric distribution or graph structure assumptions by exploiting the geometric support shifts induced by perfect and imperfect interventions.

Listen

Modern deep learning models excel at representation learning but frequently struggle to adapt when faced with distribution shifts or novel tasks. A critical objective in causal representation learning is provable representation identification: recovering the true high-level latent causal factors (such as the position, color, or shape of objects) from high-dimensional sensory data. Most existing frameworks rely solely on observational data paired with restrictive structural or distributional assumptions about the underlying causal graph. The article addresses this limitation by investigating whether and how interventional data—which is naturally available in domains like robotics and genomics—can facilitate the provable identification of latent representations without strong parametric assumptions.

The main objective of the article is to establish theoretical identification guarantees for latent causal factors using interventional data and to demonstrate practical algorithms that leverage these guarantees across synthetic and image-based data environments.

To evaluate this objective, the authors conducted mathematical analyses grounded in the geometric properties of data supports (the domains over which variables take non-zero values) and validated their findings using controlled empirical simulations. The computational setup evaluated autoencoder models across varying latent dimensions (six and ten variables) and polynomial decoder structures (degrees two and three), as well as complex visual scenes generated by a rendering engine (64×64 pixel images of interacting objects). The empirical workflow followed a two-step framework: first fitting a standard autoencoder to reconstruct data, and second, applying linear or non-linear transformations to enforce geometric constraints corresponding to interventions or support independence.

The article establishes several key findings. First, solving the autoencoder reconstruction task with a polynomial decoder provably isolates the true latent variables up to an invertible affine transformation. Second, incorporating data from hard interventions (which fix a latent variable to a set value) strengthens identification of the targeted latent up to coordinate permutation, shift, and scaling, achieving mean correlation coefficient (MCC) recovery scores near 100% across all tested distributions. Third, when latents undergo imperfect interventions that disconnect the latent support from its ancestors, representations achieve block affine identification, restricting entanglement to isolated subsets of variables. Fourth, in purely observational settings where latents exhibit pairwise independent support, the true latents can be recovered up to permutation and scaling even if the variables are statistically dependent, extending classical independent component analysis. Finally, on image tasks, latent recovery accuracy grew progressively with the number of interventional distributions per dimension, rising from baseline scores below 35% with a single intervention up to 71–86% with nine interventions.

These findings indicate that real-world interventional experiments provide a mathematically principled shortcut for learning robust representations. Organizations deploying deep learning models in safety-critical or non-stationary environments can significantly reduce the risk of shortcut learning and distribution collapse by leveraging interventional data, eliminating the need to guess underlying causal graphs or enforce artificial statistical independence.

Decision-makers should consider integrating perturbation and experimental datasets (such as genetic knockout screens or active robotic interactions) into representation learning pipelines. For practical implementations, teams should adopt a two-stage training approach: first ensuring accurate reconstruction, followed by alignment to the geometric signatures of interventions. When transitioning from polynomial approximations to complex neural network mappings, practitioners must plan for collecting multiple interventional distributions per latent factor to achieve high recovery precision.

The findings are theoretically robust within the defined conditions, but several practical limitations remain. Theoretical guarantees for single-intervention settings depend on polynomial mixing decoders or bounded approximations; general non-linear diffeomorphisms require multiple interventional distributions per latent factor. In addition, real-world measurement noise and incomplete experimental control over interventions may introduce variance. Further empirical validation on large-scale, real-world biological and robotic benchmarks is recommended before deploying the method in autonomous decision-making systems.

Cover for Interventional Causal Representation Learning

Abstract

Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent factors. However, interventional data is prevalent across applications. Can interventional data facilitate causal representation learning? We explore this question in this paper. The key observation is that interventional data often carries geometric signatures of the latent factors’ support (i.e. what values each latent can possibly take). For example, when the latent factors are causally connected, interventions can break the dependency between the intervened latents’ support and their ancestors’. Leveraging this fact, we prove that the latent causal factors can be identified up to permutation and scaling given data from perfect do interventions. Moreover, we can achieve block affine identification, namely the estimated latent factors are only entangled with a few other latents if we have access to data from imperfect interventions. These results highlight the unique power of interventional data in causal representation learning; they can enable provable identification of latent factors without any assumptions about their distributions or dependency structure.

Table of Contents

  • Abstract
  • 1. Introduction
  • Acknowledgments

Knowls

  1. Knowl 1 — Generative setup and reconstruction objective

    model/method

    The causal representation problem starts with a latent vector z∈Rdz\in\mathbb{R}^d and an observation x∈Rnx\in\mathbb{R}^n generated by an injective decoder gg as x=g(z)x=g(z). Observational latents are drawn from a distribution PZP_Z; an intervention on coordinate ii instead draws them from a distribution PZ(i)P_Z^{(i)}, while using the same decoder gg. The corresponding supports are denoted ZZ and Z(i)Z^{(i)}, and their observation-space images are X=g(Z)X=g(Z) and X(i)=g(Z(i))X^{(i)}=g(Z^{(i)}).

    An encoder f:Rn→Rdf:\mathbb{R}^n\to\mathbb{R}^d and decoder h:Rd→Rnh:\mathbb{R}^d\to\mathbb{R}^n form an autoencoder satisfying exact reconstruction, h(f(x))=xh(f(x))=x, on the observational and available interventional supports. The representation is z^=f(x)\hat z=f(x). Because reconstruction alone permits invertible reparameterizations of the latent code, the paper seeks additional conditions under which z^\hat z identifies zz up to specified transformations.

  2. Knowl 2 — Affine identification with polynomial decoders

    theoretical result

    Suppose observational and interventional data are generated as x=g(z)x=g(z), where g:Rd→Rng:\mathbb{R}^d\to\mathbb{R}^n is a degree-pp polynomial with feature representation g(z)=Gϕp(z)g(z)=G\phi_p(z). Here ϕp(z)\phi_p(z) contains the constant and all distinct monomials in the dd coordinates of total degree at most pp, and GG has full column rank. Assume the interior of the union of the latent supports is nonempty. An autoencoder that exactly reconstructs every observation, uses a degree-pp polynomial learned decoder, and has an encoder whose output support has nonempty interior identifies the latent representation affinely: there are an invertible matrix A∈Rd×dA\in\mathbb{R}^{d\times d} and a vector c∈Rdc\in\mathbb{R}^d such that z^=Az+c\hat z=Az+c throughout the latent supports. No assumption on the latent variables' dependence structure or distributions is needed. With the full monomial basis, the number of features is q=(d+pp)q=\binom{d+p}{p}, so full column rank requires n≥qn\ge q.

    The same affine conclusion also holds for a sparse polynomial feature basis under the paper's stated condition that the selected features include the degree-one terms and at least one pure-power term zioz_i^o with o≥(p+1)/2o\ge (p+1)/2, provided its coefficient matrix has full column rank. In that case the observation dimension need only be at least the number of selected features.

  3. Knowl 3 — Approximate affine identification beyond exact polynomial decoders

    theoretical result

    Let gg be a decoder that is uniformly approximable on the latent support by a degree-pp polynomial, and let an autoencoder use a degree-pp polynomial decoder with approximate reconstruction error at most ϵrec\epsilon_{\mathrm{rec}}. Define a=f∘ga=f\circ g, so that the learned code is z^=a(z)\hat z=a(z). The result assumes that aa is approximable by a polynomial, that the latent support contains a sufficiently large cube [−zmax⁡,zmax⁡]d[-z_{\max},z_{\max}]^d, and that the encoder outputs are bounded below by γη\gamma\eta for some γ>2\gamma>2, where η\eta is the approximation error for aa. It also requires the selected decoder coefficient matrix to be well-conditioned and the entries of the corresponding transformed true-decoder coefficients to be bounded.

    Under these conditions, the polynomial approximation of aa is approximately linear: the magnitudes of its degree-kk coefficients for k≥2k\ge 2 decrease at rate 1/zmax⁡k−11/z_{\max}^{k-1}. Thus, as the supported cube grows, nonlinear terms in the recovered representation become small; the result is approximate affine identification rather than an exact guarantee for arbitrary decoders.

  4. Knowl 4 — Identification from a hard do intervention

    theoretical result

    Suppose gg is an injective finite-degree polynomial with full-column-rank coefficient matrix, the interior of the latent support is nonempty, and the support of the unintervened coordinates z−iz_{-i} under an intervention has nonempty interior. A hard do intervention fixes coordinate ziz_i to a constant z∗z^* while the other coordinates are drawn from an otherwise unrestricted interventional distribution. Train an autoencoder with a polynomial decoder of the same degree as gg and require one encoder coordinate, say z^k\hat z_k, to be constant over all observations from that interventional distribution. The constant need not be the true intervention value, and the learner need not know which encoder coordinate corresponds to ziz_i.

    The constrained autoencoder identifies the intervened latent coordinate up to shift and scale: z^k=ezi+b\hat z_k=e z_i+b for constants ee and bb. With hard interventions on distinct latent coordinates, the intervened coordinates are identified up to a common coordinate permutation and individual shifts and scales. The guarantee does not require a particular latent distribution or causal graph structure.

  5. Knowl 5 — Approximate identification with multiple interventions and a general decoder

    theoretical result

    For a general diffeomorphic observation map gg, multiple hard do interventions at distinct values of the same latent coordinate can approximately isolate that coordinate, whereas a single intervention does not in general suffice. Let a=f∘ga=f\circ g map true latents to learned codes, and let aka_k be the encoder component constrained to a fixed value within each interventional dataset. Assume the observational support has nonempty interior; each intervention leaves the support of z−iz_{-i} equal to its observational support; intervention targets are sampled from a distribution covering the support of ziz_i with density at least ρ>0\rho>0; and the second derivatives of aa are bounded by LL. Write βiinf⁡\beta_i^{\inf} and βisup⁡\beta_i^{\sup} for the lower and upper endpoints of the support of ziz_i.

    For tolerance ϵ>0\epsilon>0 and failure probability δ>0\delta>0, the paper's sufficient intervention count is

    t≥log⁡ ⁣(δϵ2L(βisup⁡+βiinf⁡))log⁡ ⁣(1−ρϵ2L).t\geq \frac{\log\!\left(\frac{\delta\epsilon}{2L(\beta_i^{\sup}+\beta_i^{\inf})}\right)}{\log\!\left(1-\frac{\rho\epsilon}{2L}\right)}.

    With probability at least 1−δ1-\delta over the intervention targets, the resulting representation satisfies ∥∇z−iak(z)∥∞≤ϵ\|\nabla_{z_{-i}}a_k(z)\|_\infty\leq\epsilon throughout the latent support. Thus the constrained coordinate approximately depends only on ziz_i; the overall learned latent map is invertible, so this gives approximate identification of the intervened factor up to an invertible transformation.

  6. Knowl 6 — Block-affine identification from perfect and imperfect interventions

    theoretical result

    Two coordinates have independent support when the joint support of their distribution equals the Cartesian product of their marginal supports; this condition does not require statistical independence. Consider an intervention on latent coordinate ziz_i. Suppose there is a set SS of coordinates such that, for each j∈Sj\in S, the interventional support factors as Zi,j(i)=Zi(i)×Zj(i)Z_{i,j}^{(i)}=Z_i^{(i)}\times Z_j^{(i)}. Also assume all marginal supports are bounded and contain a nonzero-width neighborhood immediately inside each finite endpoint. Perfect interventions satisfy the support-factorization condition, as can imperfect interventions that make the range of the intervened variable the same for every value of its parents.

    Use an injective polynomial observation decoder with a full-column-rank coefficient matrix, nonempty-interior latent supports, and an autoencoder with a matching-degree polynomial decoder. If the learned representation has a coordinate z^k\hat z_k whose support is independent of the supports of each coordinate in a set S′S' of size at most ∣S∣|S| on interventional data, then the learned code is affine in the true latents, z^=Az+c\hat z=Az+c, with invertible AA. More specifically, row aka_k of AA has at most d−∣S′∣d-|S'| nonzero entries, and each row ama_m, m∈S′m\in S', is zero at every position where aka_k is nonzero. When ∣S′∣=d−1|S'|=d-1, z^k\hat z_k identifies one true latent up to shift and scale.

  7. Knowl 7 — Identification from independent supports in observational data

    theoretical result

    Suppose the observational latent support is pairwise factorized: for every distinct pair r,sr,s, Zr,s=Zr×ZsZ_{r,s}=Z_r\times Z_s. Assume each marginal support is bounded and contains neighborhoods immediately inside its lower and upper endpoints. These conditions concern support geometry, not statistical independence, so the latent variables may be dependent. Use an injective finite-degree polynomial observation decoder with full-column-rank coefficient matrix and nonempty-interior latent support. Constrain every pair of distinct learned coordinates to have independent support in observational data, while requiring exact autoencoder reconstruction with a polynomial decoder of the matching degree.

    The learned coordinates then identify the true latents up to permutation, coordinate-wise nonzero scaling, and shift: z^=ΛΠz+c\hat z=\Lambda\Pi z+c, where Π\Pi is a permutation matrix, Λ\Lambda is invertible and diagonal, and c∈Rdc\in\mathbb{R}^d. No interventional data is needed for this result.

  8. Knowl 8 — Two-stage training for intervention and support constraints

    algorithm

    Input observational observations and, when available, interventional datasets; output transformed encoder representations. First train an encoder f†f^\dagger and decoder h†h^\dagger by minimizing squared reconstruction error over the available observational and interventional observations. For polynomial-decoder experiments, constrain the learned decoder to be polynomial; the theory's affine guarantee assumes its degree matches the true polynomial degree.

    For do-intervention data, fit a separate map γi\gamma_i to the learned codes from each intervention on coordinate ii, minimizing mean squared error to a fixed target over that interventional dataset. The target can be chosen arbitrarily without knowing the true do value; experiments sampled targets from Uniform(0,1)\mathrm{Uniform}(0,1). Stack the maps into Γ\Gamma and return Γf†(x)\Gamma f^\dagger(x). The polynomial experiments used linear maps, while the image experiments used nonlinear MLP maps.

    For an independent-support objective, fit an invertible transformation Γ\Gamma and an approximate inverse Γ′\Gamma' to minimize

    E[∥Γ′(Γ(z^))−z^∥22]+λ∑k≠mDH ⁣(Z^k,m(Γ),Z^k(Γ)×Z^m(Γ)),\mathbb{E}\left[\|\Gamma'(\Gamma(\hat z))-\hat z\|_2^2\right]+\lambda\sum_{k\ne m}D_H\!\left(\hat Z_{k,m}(\Gamma),\hat Z_k(\Gamma)\times\hat Z_m(\Gamma)\right),

    where z^=f†(x)\hat z=f^\dagger(x), Z^k(Γ)\hat Z_k(\Gamma) is the transformed marginal support, and Z^k,m(Γ)\hat Z_{k,m}(\Gamma) is the transformed joint support. The support discrepancy used in the paper is DH(A,B)=sup⁡b∈Binf⁡a∈A∥b−a∥2D_H(A,B)=\sup_{b\in B}\inf_{a\in A}\|b-a\|_2. The polynomial independent-support experiments used λ=10\lambda=10. They used Adam with batch size 16, weight decay 5×10−45\times10^{-4}, up to 200 epochs, and early stopping after 10 epochs without validation improvement. The paper does not state a computational-complexity bound.

  9. Knowl 9 — Polynomial-decoder experiments validate the theoretical identification patterns

    empirical result

    Synthetic observations used n=200n=200, latent dimensions d∈{6,10}d\in\{6,10\}, polynomial degrees p∈{2,3}p\in\{2,3\}, and coefficient entries sampled independently from a standard normal distribution. Latents included independent uniform samples and sparse- or dense-connectivity structural causal models. Results are means ±\pm standard errors over five random seeds.

    With observational data and a polynomial learned decoder, the linear-regression R2R^2 between learned and true representations ranged from 0.72±0.150.72\pm0.15 to 1.00±0.001.00\pm0.00, consistent with affine identification. After imposing independent support, MCC was 90.7±2.9290.7\pm2.92 to 99.4±0.0699.4\pm0.06 for the uniform latent cases, which satisfy the support condition; for sparse and dense SCM cases it ranged from 58.8±1.2758.8\pm1.27 to 72.6±1.4872.6\pm1.48. With one do intervention per latent and the intervention-based transformation, MCC rose from a pre-transformation range of 59.9±2.0359.9\pm2.03 to 79.5±3.4579.5\pm3.45 to 95.3±2.2495.3\pm2.24 to 100.0±0.00100.0\pm0.00 after transformation across the tested latent distributions and polynomial settings. This demonstrates that the intervention constraint recovered coordinate-wise correspondence even for the dependent SCM latents.

  10. Knowl 10 — More intervention target values improve identification on rendered images

    empirical result

    The image experiments rendered two balls at positions determined by four latent coordinates into 64×64×364\times64\times3 images. The latent mechanisms were independent uniform, linear SCM, and nonlinear SCM. A nonlinear map was fitted to intervention data in the second training stage, and MCC was evaluated on observational test data. Entries below are MCC means ±\pm standard errors over five seeds, indexed by the number of distinct do-intervention values per latent coordinate:

    • One value: uniform 34.2±0.2434.2\pm0.24, linear SCM 12.8±0.2812.8\pm0.28, nonlinear SCM 19.7±0.3119.7\pm0.31.
    • Three values: uniform 73.9±0.3873.9\pm0.38, linear SCM 73.2±0.3373.2\pm0.33, nonlinear SCM 59.7±0.2859.7\pm0.28.
    • Five values: uniform 73.6±0.2173.6\pm0.21, linear SCM 83.4±0.2183.4\pm0.21, nonlinear SCM 62.8±0.262.8\pm0.2.
    • Seven values: uniform 72.5±0.3472.5\pm0.34, linear SCM 84.2±0.2584.2\pm0.25, nonlinear SCM 69.3±0.3469.3\pm0.34.
    • Nine values: uniform 73.1±0.4773.1\pm0.47, linear SCM 86.2±0.1786.2\pm0.17, nonlinear SCM 71.4±0.2671.4\pm0.26.

    Across all three latent mechanisms, moving from one intervention value to three substantially increased MCC; further intervention values generally improved or sustained it. The result is consistent with the theory that a general nonlinear decoder requires multiple intervention targets to approximately isolate a latent coordinate.

Coverage note — The supplementary β-VAE baseline comparisons are omitted because they are secondary ablations rather than part of the proposed identification theory or training method; proof-only intermediate lemmas are also omitted.

References

  1. 1.Ahuja, K., Hartford, J., and Bengio, Y. Properties from mechanisms: an equivariance perspective on identifiable representation learning. arXiv preprint arXiv:2110.15796, 2021.
  2. 2.Ahuja, K., Hartford, J., and Bengio, Y. Weakly supervised representation learning with sparse perturbations. arXiv preprint arXiv:2206.01101, 2022a.
  3. 3.Ahuja, K., Mahajan, D., Syrgkanis, V., and Mitliagkas, I. Towards efficient representation identification in supervised learning. arXiv preprint arXiv:2204.04606, 2022b.
  4. 4.Ash, R. B., Robert, B., Doleans-Dade, C. A., and Catherine, A. Probability and measure theory. Academic press, 2000.
  5. 5.Bareinboim, E., Correa, J. D., Ibeling, D., and Icard, T. On pearl’s hierarchy and the foundations of causal inference. In Probabilistic and Causal Inference: The Works of Judea Pearl, pp. 507–556. 2022.
  6. 6.Bengio, Y., Courville, A., and Vincent, P. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8):1798–1828, 2013.
  7. 7.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  8. 8.Brehmer, J., De Haan, P., Lippe, P., and Cohen, T. Weakly supervised causal representation learning. arXiv preprint arXiv:2203.16437, 2022.
  9. 9.Brouillard, P., Lachapelle, S., Lacoste, A., Lacoste-Julien, S., and Drouin, A. Differentiable causal discovery from interventional data. Advances in Neural Information Processing Systems, 33:21865–21877, 2020.
  10. 10.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  11. 11.Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A. Understanding disentangling in beta-vae. arXiv preprint arXiv:1804.03599, 2018.
  12. 12.Comon, P. Independent component analysis, a new concept? Signal processing, 36(3):287–314, 1994.
  13. 13.Dixit, A., Parnas, O., Li, B., Chen, J., Fulco, C. P., Jerby-Arnon, L., Marjanovic, N. D., Dionne, D., Burks, T., Raychowdhury, R., et al. Perturb-seq: dissecting molecular circuits with scalable single-cell rna profiling of pooled genetic screens. cell, 167(7):1853–1866, 2016.
  14. 14.Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020.
  15. 15.Goyal, A. and Bengio, Y. Inductive biases for deep learning of higher-level cognition. arXiv preprint arXiv:2011.15091, 2020.
  16. 16.Halv ¨ a, H. and Hyvarinen, A. Hidden markov nonlinear ica: Unsupervised learning from nonstationary time series. In Conference on Uncertainty in Artificial Intelligence, pp. 939–948. PMLR, 2020.
  17. 17.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  18. 18.Hyvarinen, A. and Morioka, H. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. Advances in neural information processing systems, 29, 2016.
  19. 19.Hyvarinen, A. and Morioka, H. Nonlinear ICA of temporally dependent stationary sources. In Artificial Intelligence and Statistics, pp. 460–469. PMLR, 2017.
  20. 20.Hyvarinen, A. and Pajunen, P. Nonlinear independent component analysis: Existence and uniqueness results. Neural networks, 12(3):429–439, 1999.
  21. 21.Hyvarinen, A., Sasaki, H., and Turner, R. Nonlinear ica using auxiliary variables and generalized contrastive learning. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 859–868. PMLR, 2019.
  22. 22.Khemakhem, I., Monti, R., Kingma, D., and Hyvarinen, A. Ice-beem: Identifiable conditional energy-based deep models based on nonlinear ICA. Advances in Neural Information Processing Systems, 33:12768–12778, 2020.
  23. 23.Khemakhem, I., Kingma, D., Monti, R., and Hyvarinen, A. Variational autoencoders and nonlinear ICA: A unifying framework. In International Conference on Artificial Intelligence and Statistics, pp. 2207–2217. PMLR, 2022.
  24. 24.Klindt, D., Schott, L., Sharma, Y., Ustyuzhaninov, I., Brendel, W., Bethge, M., and Paiton, D. Towards nonlinear disentanglement in natural data with temporal sparse coding. arXiv preprint arXiv:2007.10930, 2020.
  25. 25.Lachapelle, S., Rodriguez, P., Sharma, Y., Everett, K. E., Le Priol, R., Lacoste, A., and Lacoste-Julien, S. Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ICA. In Conference on Causal Learning and Reasoning, pp. 428–484. PMLR, 2022.
  26. 26.Lippe, P., Magliacane, S., Lowe, S., Asano, Y. M., Cohen, T., and Gavves, E. icitris: Causal representation learning for instantaneous temporal effects. arXiv preprint arXiv:2206.06169, 2022a.
  27. 27.Lippe, P., Magliacane, S., Lowe, S., Asano, Y. M., Cohen, T., and Gavves, S. Citris: Causal identifiability from temporal intervened sequences. In International Conference on Machine Learning, pp. 13557–13603. PMLR, 2022b.
  28. 28.Liu, Y., Alahi, A., Russell, C., Horn, M., Zietlow, D., Scholkopf, B., and Locatello, F. Causal triplet: An open challenge for intervention-centric causal representation learning. arXiv preprint arXiv:2301.05169, 2023.
  29. 29.Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Scholkopf, B., and Bachem, O. Challenging common assumptions in the unsupervised learning of disentangled representations. In international conference on machine learning, pp. 4114–4124. PMLR, 2019.
  30. 30.Locatello, F., Poole, B., Ratsch, G., Sch ¨ olkopf, B., Bachem, O., and Tschannen, M. Weakly-supervised disentanglement without compromises. In International Conference on Machine Learning, pp. 6348–6359. PMLR, 2020.
  31. 31.Lopez, R., Tagasovska, N., Ra, S., Cho, K., Pritchard, J. K., and Regev, A. Learning causal representations of single cells via sparse mechanism shift modeling. arXiv preprint arXiv:2211.03553, 2022.
  32. 32.Mityagin, B. The zero set of a real analytic function. arXiv preprint arXiv:1512.07276, 2015.
  33. 33.Mooij, J. and Heskes, T. Cyclic causal discovery from continuous equilibrium data. arXiv preprint arXiv:1309.6849, 2013.
  34. 34.Nejatbakhsh, A., Fumarola, F., Esteki, S., Toyoizumi, T., Kiani, R., and Mazzucato, L. Predicting perturbation effects from resting activity using functional causal flow. bioRxiv, pp. 2020–11, 2021.
  35. 35.Pearl, J. Causal inference in statistics: An overview. Statistics surveys, 3:96–146, 2009.
  36. 36.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  37. 37.Peters, J., Janzing, D., and Scholkopf, B. Elements of causal inference: foundations and learning algorithms. The MIT Press, 2017.
  38. 38.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021.
  39. 39.Roth, K., Ibrahim, M., Akata, Z., Vincent, P., and Bouchacourt, D. Disentanglement of correlated factors via hausdorff factorized support. arXiv preprint arXiv:2210.07347, 2022.
  40. 40.Scholkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Towards causal representation learning 2021. arXiv preprint arXiv:2102.11107, 2021.
  41. 41.Seigal, A., Squires, C., and Uhler, C. Linear causal disentanglement via interventions. arXiv preprint arXiv:2211.16467, 2022.
  42. 42.Shinners, P. Pygame. http://pygame.org/, 2011.
  43. 43.Shinners, P. et al. Pygame. Dostupne z: http://pygame.org/[Online (2011), 2011.
  44. 44.Von Kugelgen, J., Sharma, Y., Gresele, L., Brendel, W., Scholkopf, B., Besserve, M., and Locatello, F. Self-supervised learning with data augmentations provably isolates content from style. Advances in neural information processing systems, 34:16451–16467, 2021.
  45. 45.Wang, Y. and Jordan, M. I. Desiderata for representation learning: A causal perspective. arXiv preprint arXiv:2109.03795, 2021.
  46. 46.Yamada, Y., Tang, T., and Ilker, Y. When are lemons purple? the concept association bias of clip. arXiv preprint arXiv:2212.12043, 2022.
  47. 47.Yao, W., Sun, Y., Ho, A., Sun, C., and Zhang, K. Learning temporally causal latent processes from general temporal data. arXiv preprint arXiv:2110.05428, 2021.
  48. 48.Yao, W., Chen, G., and Zhang, K. Learning latent causal dynamics. arXiv preprint arXiv:2202.04828, 2022a.
  49. 49.Yao, W., Sun, Y., Ho, A., Sun, C., and Zhang, K. Learning temporally causal latent processes from general temporal data. In International Conference on Learning Representations, 2022b. URL https://openreview.net/forum?id=RDlLMjLJXdq.
  50. 50.Zimmermann, R. S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W. Contrastive learning inverts the data generating process. In International Conference on Machine Learning, pp. 12979–12990. PMLR, 2021.

Citation

MLA
Ahuja, K., et al. “Interventional Causal Representation Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 372–407, https://proceedings.mlr.press/v202/ahuja23a.html.
APA
Ahuja, K., Mahajan, D., Wang, Y., & Bengio, Y. (2023). Interventional Causal Representation Learning. International Conference on Machine Learning, 202, 372–407. https://proceedings.mlr.press/v202/ahuja23a.html
Chicago
Ahuja, K., D. Mahajan, Y. Wang, and Y. Bengio. 2023. “Interventional Causal Representation Learning”. International Conference on Machine Learning 202: 372–407. https://proceedings.mlr.press/v202/ahuja23a.html.
Harvard
Ahuja, K. et al. (2023) “Interventional Causal Representation Learning”, International Conference on Machine Learning. PMLR, pp. 372–407. Available at: https://proceedings.mlr.press/v202/ahuja23a.html.
Vancouver
1. Ahuja K, Mahajan D, Wang Y, Bengio Y (2023) Interventional Causal Representation Learning. In: International Conference on Machine Learning. PMLR, pp 372–407

BibTeX

@InProceedings{pmlr-v202-ahuja23a,
  title = 	 {Interventional Causal Representation Learning},
  author =       {Ahuja, Kartik and Mahajan, Divyat and Wang, Yixin and Bengio, Yoshua},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {372--407},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ahuja23a/ahuja23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ahuja23a.html},
  abstract = 	 {Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent factors. However, interventional data is prevalent across applications. Can interventional data facilitate causal representation learning? We explore this question in this paper. The key observation is that interventional data often carries geometric signatures of the latent factors’ support (i.e. what values each latent can possibly take). For example, when the latent factors are causally connected, interventions can break the dependency between the intervened latents’ support and their ancestors’. Leveraging this fact, we prove that the latent causal factors can be identified up to permutation and scaling given data from perfect do interventions. Moreover, we can achieve block affine identification, namely the estimated latent factors are only entangled with a few other latents if we have access to data from imperfect interventions. These results highlight the unique power of interventional data in causal representation learning; they can enable provable identification of latent factors without any assumptions about their distributions or dependency structure.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/