Causal Representation Learning from Multiple Distributions: A General Setting

Kun ZhangShaoan XieIgnavier NgYujia Zheng

article2024ICML56 citations

Establishes theoretical guarantees for completely nonparametric causal representation learning from multiple distributions without requiring hard interventions, demonstrating that latent variables and the underlying moralized causal graph can be recovered under graph sparsity and sufficient mechanism variation conditions.

Listen

Modern data systems frequently observe complex, high-dimensional measurements—such as image pixels, audio signals, or system logs—that are generated by underlying, unobserved causal factors. To make reliable predictions under shifting operational environments or to understand the root causes of system behavior, organizations must discover these hidden causal variables and their underlying network structures. However, existing techniques often rely on restrictive functional assumptions (such as linearity) or require active, hard experimental interventions that may be costly, unethical, or technically impossible to execute in production settings.

The article evaluates whether it is mathematically possible to recover latent causal variables and their relationships purely from observational data collected across changing conditions, such as multiple domains, heterogeneous sources, or non-stationary time series. It aims to establish the fundamental boundaries of what can be identified in a fully non-parametric setting without hard interventions, providing a foundational baseline for causal machine learning.

To demonstrate this, the authors develop theoretical proofs grounded in differential geometry and conditional independence analysis. They analyze a data-generating setup where hidden causal variables influence observations through an unknown non-linear mixing process while their internal dynamics shift across environments. The approach is implemented practically using a variational autoencoder architecture combined with sparsity regularization on the estimated causal structure. The framework is verified across synthetic benchmark experiments, including four-node chain and collider network structures evaluated under both Gaussian and Laplacian noise regimes.

The analysis establishes several key findings. First, under a graph sparsity penalty and sufficient distributional diversity across environments, the underlying undirected dependency structure (the Markov network) can be identified up to an exact isomorphism. Second, each estimated latent variable can be recovered as a function of the true variable and a strictly limited set of "intimate neighbors"—variables that share mutual connections across the graph. In many topological configurations where this intimate neighbor set is empty, latent variables are recovered perfectly up to simple one-to-one transformations. Third, the recovered undirected graph corresponds exactly to the moralized graph of the true underlying directed causal network under two new, minimal faithfulness relaxations (single adjacency-faithfulness and single unshielded-collider-faithfulness). Finally, the article formally proves that standard non-linear independent component methods cannot achieve independent representations when underlying causal dependencies exist, underscoring the necessity of explicitly modeling causal structures.

These findings indicate that organizations can recover significant causal insights from naturally occurring distribution shifts—such as seasonal trends, geographic variations, or operational regime changes—without executing disruptive experiments. This reduces the cost and operational risk associated with active A/B testing while mitigating the risk of deploying machine learning models that rely on spurious correlations. Furthermore, the work clarifies the minimal assumptions required for causal representation, helping data teams avoid overly restrictive parametric constraints that could degrade modeling performance.

Technical leaders and practitioners developing representation learning pipelines should adopt sparsity-penalized deep generative models to capture underlying causal factors when multi-environment observational data is available. Before relying on identified factors for automated decision-making, teams should evaluate their system's graph structure to verify whether key variables fall into isolated configurations that guarantee one-to-one identifiability or remain entangled with intimate neighbors.

While the theoretical guarantees are rigorous, users must consider key boundary conditions. The theoretical framework requires data across a sufficient number of distinct environments (at least twice the number of latent variables plus the number of edges) and assumes smooth, non-zero probability densities. Future investigations are required to adapt the framework to scenarios where only a localized subset of causal relations shift across environments.

arXiv: 2402.05052
Cover for Causal Representation Learning from Multiple Distributions: A General Setting

Abstract

In many problems, the measured variables (e.g., image pixels) are just mathematical functions of the latent causal variables (e.g., the underlying concepts or objects). For the purpose of making predictions in changing environments or making proper changes to the system, it is helpful to recover the latent causal variables Z_i and their causal relations represented by graph G_Z. This problem has recently been known as causal representation learning. This paper is concerned with a general, completely nonparametric setting of causal representation learning from multiple distributions (arising from heterogeneous data or nonstationary time series), without assuming hard interventions behind distribution changes. We aim to develop general solutions in this fundamental case; as a by product, this helps see the unique benefit offered by other assumptions such as parametric causal models or hard interventions. We show that under the sparsity constraint on the recovered graph over the latent variables and suitable sufficient change conditions on the causal influences, interestingly, one can recover the moralized graph of the underlying directed acyclic graph, and the recovered latent variables and their relations are related to the underlying causal model in a specific, nontrivial way. In some cases, most latent variables can even be recovered up to component-wise transformations. Experimental results verify our theoretical claims.

Table of Contents

  • 1. Introduction
  • 2. Problem Setting
  • 3. Learning Causal Representations from Multiple Distributions
  • 3.1. Recovering Latent Causal Variables and Latent Markov Network
  • 3.2. From Latent Markov Network to Latent Causal DAG
  • 4. Change Encoding Network for Representation Learning
  • 4.1. Nonparametric Implementation of the Prior Distribution
  • 4.2. Parametric Implementation of the Prior Distribution
  • 4.3. Full Objective
  • 4.4. Simulations
  • 5. Related Work
  • 6. Conclusion and Discussions
  • Acknowledgements
  • Impact Statement
  • References
  • A. Proofs of Useful Lemmas
  • A.1. Proof of Lemma 2
  • A.2. Proof of Lemma 3
  • B. Proof of Proposition 1
  • C. Proof of Theorem 2
  • D. Proof of Theorem 1
  • E. Proof of Theorem 3
  • F. Proof of Corollary 1
  • G. Proof of Lemma 1 and Proposition 2

Knowls

  1. Knowl 1 — Nonparametric latent causal data-generating model

    model/method

    The paper studies observations X∈RdX\in\mathbb{R}^d generated from latent causal variables Z=(Z1,…,Zn)∈RnZ=(Z_1,\ldots,Z_n)\in\mathbb{R}^n through an unknown nonlinear mixing map

    X=g(Z),Zi=fi(PA⁡(Zi),ϵi;θi),i=1,…,n.X=g(Z),\qquad Z_i=f_i(\operatorname{PA}(Z_i),\epsilon_i;\theta_i),\quad i=1,\ldots,n.

    Here d≥nd\ge n, gg is a C2C^2 diffeomorphism onto its image, PA⁡(Zi)\operatorname{PA}(Z_i) is the set of parents of ZiZ_i in an acyclic latent causal graph GZG_Z, and the exogenous noises ϵ1,…,ϵn\epsilon_1,\ldots,\epsilon_n are mutually independent. The domain-specific factor θi\theta_i governs changes in the mechanism generating ZiZ_i; the joint factor is θ=(θ1,…,θn)\theta=(\theta_1,\ldots,\theta_n). The observed data consist of samples from multiple distributions indexed by Θ={θ(1),…,θ(m)}\Theta=\{\theta^{(1)},\ldots,\theta^{(m)}\}, while ZZ and θ\theta are unobserved. The goal is to recover the latent variables and their causal relations without assuming parametric mechanisms, hard interventions, or a known mixing function.

  2. Knowl 2 — Markov-network and sufficient-change assumptions

    assumption

    For latent variables Z=(Z1,…,Zn)Z=(Z_1,\ldots,Z_n) with density pZ(Z;θ)p_Z(Z;\theta), define the undirected latent Markov network MZM_Z by placing an edge {Zi,Zj}\{Z_i,Z_j\} exactly when

    Zi̸ ⁣⊥ ⁣ ⁣ ⁣⊥Zj∣Z[n]∖{i,j},[n]={1,…,n}.Z_i\not\!\perp\!\!\!\perp Z_j\mid Z_{[n]\setminus\{i,j\}}, \qquad [n]=\{1,\ldots,n\}.

    The identifiability results assume: (i) pZ(⋅;θ)p_Z(\cdot;\theta) is positive and twice continuously differentiable on Rn\mathbb{R}^n; and (ii) the mechanisms change sufficiently across domains. Specifically, for every ZZ, there are 2n+∣MZ∣+12n+|M_Z|+1 domain factors θ(u)\theta^{(u)}, u=0,…,2n+∣MZ∣u=0,\ldots,2n+|M_Z|, such that the vectors w(Z,u)−w(Z,0)w(Z,u)-w(Z,0) for u=1,…,2n+∣MZ∣u=1,\ldots,2n+|M_Z| are linearly independent, where

    w(Z,u)=(∂log⁡pZ(Z;θ(u))∂Zi)i∈[n]⊕(∂2log⁡pZ(Z;θ(u))∂Zi2)i∈[n]⊕(∂2log⁡pZ(Z;θ(u))∂Zi∂Zj){Zi,Zj}∈E(MZ), i<j.w(Z,u)=\left(\frac{\partial\log p_Z(Z;\theta^{(u)})}{\partial Z_i}\right)_{i\in[n]} \oplus \left(\frac{\partial^2\log p_Z(Z;\theta^{(u)})}{\partial Z_i^2}\right)_{i\in[n]} \oplus \left(\frac{\partial^2\log p_Z(Z;\theta^{(u)})}{\partial Z_i\partial Z_j}\right)_{\{Z_i,Z_j\}\in E(M_Z),\,i<j}.

    For a positive twice-differentiable density, pairwise conditional independence given all remaining variables is equivalent to a zero mixed derivative of the log density:

    Zi⊥ ⁣ ⁣ ⁣⊥Zj∣Z[n]∖{i,j}⟺∂2log⁡pZ(Z;θ)∂Zi∂Zj=0.Z_i\perp\!\!\!\perp Z_j\mid Z_{[n]\setminus\{i,j\}} \quad\Longleftrightarrow\quad \frac{\partial^2\log p_Z(Z;\theta)}{\partial Z_i\partial Z_j}=0.
  3. Knowl 3 — Jacobian restrictions induced by a recovered sparse Markov network

    theoretical result

    Suppose an estimated latent model with variables Z^=(Z^1,…,Z^n)\widehat Z=(\widehat Z_1,\ldots,\widehat Z_n) and mixing function g^\widehat g reproduces the observed distribution pX(⋅;θ(u))p_X(\cdot;\theta^{(u)}) in every domain uu. Under the positive smooth-density and sufficient-change assumptions, let MZ^M_{\widehat Z} be the estimated latent Markov network. For any two estimated variables Z^k\widehat Z_k and Z^l\widehat Z_l that are nonadjacent in MZ^M_{\widehat Z}, the Jacobian of the invertible map from Z^\widehat Z to the true ZZ satisfies, at every latent value,

    ∂Zi∂Z^k∂Zi∂Z^l=0for every Zi,\frac{\partial Z_i}{\partial \widehat Z_k} \frac{\partial Z_i}{\partial \widehat Z_l}=0 \qquad\text{for every }Z_i,

    and, for every pair of true latent variables Zi,ZjZ_i,Z_j adjacent in MZM_Z,

    ∂Zi∂Z^k∂Zj∂Z^l=0.\frac{\partial Z_i}{\partial \widehat Z_k} \frac{\partial Z_j}{\partial \widehat Z_l}=0.

    Thus, nonadjacency in the recovered Markov network prevents a true latent variable from having simultaneous local dependence on both recovered variables, and also prevents both members of an adjacent true pair from depending on the corresponding recovered pair.

  4. Knowl 4 — Effect of edge-minimality on latent-variable dependence

    theoretical result

    Among all estimated latent models that reproduce the observed distributions in every domain, choose one whose Markov network MZ^M_{\widehat Z} has the minimum possible number of edges. Under the smooth-density and sufficient-change assumptions, for every nonadjacent pair Z^k,Z^l\widehat Z_k,\widehat Z_l in MZ^M_{\widehat Z}: (a) each true latent variable ZiZ_i depends nontrivially on at most one of Z^k\widehat Z_k and Z^l\widehat Z_l; and (b) if ZiZ_i and ZjZ_j are adjacent in the true Markov network MZM_Z, at most one of ZiZ_i and ZjZ_j can depend on either member of the pair {Z^k,Z^l}\{\widehat Z_k,\widehat Z_l\}. The minimum-edge requirement rules out dense or complete recovered graphs that could otherwise evade these restrictions.

  5. Knowl 5 — Identifiability of the latent Markov network

    theoretical result

    Under the nonparametric latent causal model, positive twice-differentiable latent density, sufficient changes across domains, and exact matching of all observed domain distributions, impose the minimum-edge condition on the estimated latent Markov network MZ^M_{\widehat Z}. Then MZ^M_{\widehat Z} is isomorphic to the true latent Markov network MZM_Z: there is a permutation of the recovered latent variables that preserves every undirected adjacency relation. Consequently, the conditional-dependence structure among the latent variables is identifiable even though the latent variables, mixing function, and causal mechanisms are all unobserved and unrestricted parametrically.

  6. Knowl 6 — Partial identifiability of individual latent variables through intimate neighbors

    theoretical result

    Let NZiN_{Z_i} be the neighbors of ZiZ_i in the true Markov network MZM_Z. Define the intimate-neighbor set

    ΨZi={Zj: Zj is adjacent to Zi and to every other member of NZi}.\Psi_{Z_i}=\left\{Z_j:\ Z_j\text{ is adjacent to }Z_i\text{ and to every other member of }N_{Z_i}\right\}.

    Under the same smoothness, sufficient-change, distribution-matching, and minimum-edge assumptions that identify the Markov network, there is a permutation π\pi of the recovered variables such that each recovered component Z^π(i)\widehat Z_{\pi(i)} is a function solely of a subset of {Zi}∪ΨZi\{Z_i\}\cup\Psi_{Z_i}. Hence the remaining ambiguity is localized to ZiZ_i and particular neighboring variables rather than an arbitrary mixture of all latents. If ΨZi=∅\Psi_{Z_i}=\varnothing, then Z^π(i)\widehat Z_{\pi(i)} is a component-wise transformation of ZiZ_i. In particular, this occurs whenever every neighbor of ZiZ_i fails to be adjacent to at least one other neighbor of ZiZ_i.

  7. Knowl 7 — Exact connection between the latent Markov network and the moralized DAG

    theoretical result

    The moralized graph of a latent DAG GZG_Z is obtained by making every pair of co-parents of a common child adjacent and then replacing directed edges by undirected edges. Under the Markov assumption that the latent distribution factorizes according to GZG_Z, the latent Markov network MZM_Z is always a subgraph of this moralized graph.

    The paper introduces two weaker-than-standard-faithfulness conditions. Single adjacency-faithfulness requires every adjacent pair Zi,ZjZ_i,Z_j in GZG_Z to remain dependent conditional on all other latent variables:

    Zi̸ ⁣⊥ ⁣ ⁣ ⁣⊥Zj∣Z[n]∖{i,j}.Z_i\not\!\perp\!\!\!\perp Z_j\mid Z_{[n]\setminus\{i,j\}}.

    Single unshielded-collider-faithfulness requires the two parents of every unshielded collider Zi→Zj←ZkZ_i\to Z_j\leftarrow Z_k to remain dependent under conditioning on all other variables:

    Zi̸ ⁣⊥ ⁣ ⁣ ⁣⊥Zk∣Z[n]∖{i,k}.Z_i\not\!\perp\!\!\!\perp Z_k\mid Z_{[n]\setminus\{i,k\}}.

    These two conditions are jointly necessary and sufficient for MZM_Z to equal the moralized graph of GZG_Z. Thus, the paper replaces full faithfulness with exactly the adjacency and unshielded-collider properties needed to recover the moralized causal structure.

  8. Knowl 8 — Independent-component recovery is impossible for a nonempty latent DAG

    theoretical result

    Assume the positive smooth-density and sufficient-change conditions, and suppose the true latent causal DAG GZG_Z contains at least one directed edge. No estimated model whose latent components are mutually independent in every observed domain can reproduce all the observed distributions exactly. Independence of the estimated components would make MZ^M_{\widehat Z} an empty graph, whereas the minimum-edge identifiability result requires every true Markov-network edge to be represented after a permutation. Since a nonempty DAG has a nonempty moralized/Markov structure under the stated conditions, a nonlinear independent-component representation cannot serve as an exact solution in this setting.

  9. Knowl 9 — VAE and change-encoding implementations

    model/method

    The practical method uses a variational autoencoder with a domain-conditioned latent prior. The decoder models p(X∣Z;θ^(u))p(X\mid Z;\widehat\theta^{(u)}), the encoder models q(Z∣X,u)q(Z\mid X,u), and training minimizes the negative evidence lower bound

    LELBO=KL⁡ ⁣(q(Z∣X,u) ∥ p(Z;θ^(u)))−Eq(Z∣X,u)[log⁡p(X∣Z;θ^(u))].\mathcal L_{\mathrm{ELBO}} =\operatorname{KL}\!\left(q(Z\mid X,u)\,\|\,p(Z;\widehat\theta^{(u)})\right) -\mathbb E_{q(Z\mid X,u)}[\log p(X\mid Z;\widehat\theta^{(u)})].

    The encoder posterior is Gaussian or Laplacian, with neural-network-generated mean and scale. The decoder likelihood is Gaussian, with neural-network-generated mean and fixed variance.

    For the nonparametric implementation, a causal ordering Z^1,…,Z^n\widehat Z_1,\ldots,\widehat Z_n is fixed. The lower-triangular adjacency matrix A^\widehat A selects candidate parents, and a conditional normalizing flow transforms each latent into noise:

    ϵ^i,log⁡det⁡i=Flow⁡ ⁣(Z^i;NN⁡({A^i,jZ^j}j<i,u)).\widehat\epsilon_i,\log\det_i =\operatorname{Flow}\!\left(\widehat Z_i;\operatorname{NN}\left(\{\widehat A_{i,j}\widehat Z_j\}_{j<i},u\right)\right).

    With independent noise prior p(ϵ)p(\epsilon), the learned prior is

    log⁡p(Z^;θ^(u))=∑i=1n[log⁡p(ϵ^i)+log⁡det⁡i].\log p(\widehat Z;\widehat\theta^{(u)}) =\sum_{i=1}^{n}\left[\log p(\widehat\epsilon_i)+\log\det_i\right].

    The full training objective is

    Lfull=LELBO+∥A^∥1,\mathcal L_{\mathrm{full}} =\mathcal L_{\mathrm{ELBO}}+\|\widehat A\|_1,

    where the ℓ1\ell_1 term encourages the sparse causal structure required by the theory. After training, the encoder output is the recovered representation and A^\widehat A is the estimated causal adjacency matrix.

  10. Knowl 10 — Parametric linear-prior implementation

    model/method

    A second implementation assumes a linear latent structural equation with domain-specific scaling and shifting:

    Z=A(C(u)Z)+S(u)ϵ+B(u).Z=A(C^{(u)}Z)+S^{(u)}\epsilon+B^{(u)}.

    Here A∈[0,1]n×nA\in[0,1]^{n\times n} is a causal adjacency matrix that can be permuted to a strictly lower-triangular matrix, C(u)∈Rn×nC^{(u)}\in\mathbb R^{n\times n} and S(u)∈RnS^{(u)}\in\mathbb R^n are domain-specific scaling parameters, B(u)∈RnB^{(u)}\in\mathbb R^n is a domain-specific shift, and ϵ\epsilon has independent components. With a fixed ordering of the recovered variables, the estimated noise is computed elementwise as

    ϵ^=Z^−B^(u)−A^C^(u)Z^S^(u).\widehat\epsilon =\frac{\widehat Z-\widehat B^{(u)}-\widehat A\widehat C^{(u)}\widehat Z}{\widehat S^{(u)}}.

    For an independent noise prior p(ϵ^)p(\widehat\epsilon), the induced latent prior is

    log⁡p(Z^;θ^(u))=∑i=1n[log⁡p(ϵ^i)−log⁡∣S^i(u)∣].\log p(\widehat Z;\widehat\theta^{(u)}) =\sum_{i=1}^{n}\left[\log p(\widehat\epsilon_i)-\log\left|\widehat S_i^{(u)}\right|\right].

    This prior is inserted into the same VAE objective and combined with ∥A^∥1\|\widehat A\|_1 to favor a sparse recovered causal graph.

  11. Knowl 11 — Simulation evidence for latent and graph recovery

    empirical result

    The simulations generated four-variable latent systems with either a Y structure Z1→Z3←Z2Z_1\to Z_3\leftarrow Z_2, Z3→Z4Z_3\to Z_4, or a chain Z1→Z2→Z3→Z4Z_1\to Z_2\to Z_3\to Z_4. Noise scales were sampled from Unif⁡[0.5,2]\operatorname{Unif}[0.5,2], shifts from Unif⁡[−2,2]\operatorname{Unif}[-2,2], and additional latent scaling factors from Unif⁡[0.5,2]\operatorname{Unif}[0.5,2]. The latent variables were then mapped to observations using an invertible multilayer perceptron with orthogonal weight matrices and LeakyReLU activations. Both Gaussian and Laplacian noise were tested with the nonparametric-flow and linear-parameterization implementations.

    The reported 4×\times4 pairwise plots on page 9 show that, in most settings, each recovered component has a strong one-to-one relationship with one true latent component and weak relationships with the others. The estimated sparse adjacency matrix recovers the Y and chain structures; for the Y system, the retained edges correspond to Z1→Z3←Z2Z_1\to Z_3\leftarrow Z_2 and Z3→Z4Z_3\to Z_4. The plots also reveal the expected causal dependence pattern: a recovered component aligned with Z2Z_2 is nearly independent of Z1Z_1 but related to the downstream components Z3Z_3 and Z4Z_4. The nonparametric implementation recovers the latent variables even with Laplacian noise, supporting the theoretical claim that recovery is possible up to the stated permutation and component/neighborhood indeterminacies.

Coverage note — Proofs and proof-only supplementary lemmas were omitted because they are derivational rather than standalone contributions; related work, acknowledgements, and funding were also excluded.

References

  1. 1.Adams, J., Hansen, N., and Zhang, K. Identification of partially observed linear causal models: Graphical conditions for the non-gaussian and heterogeneous cases. Advances in Neural Information Processing Systems, 34:22822–22833, 2021.
  2. 2.Ahuja, K., Mahajan, D., Wang, Y., and Bengio, Y. Interventional causal representation learning. In International Conference on Machine Learning, pp. 372–407. PMLR, 2023.
  3. 3.Ben-Israel, A. The change-of-variables formula using matrix volume. Siam Journal on Matrix Analysis and Applications, 21, 01 1999.
  4. 4.Brehmer, J., De Haan, P., Lippe, P., and Cohen, T. S. Weakly supervised causal representation learning. Advances in Neural Information Processing Systems, 35:38319–38331, 2022.
  5. 5.Buchholz, S., Besserve, M., and Scholkopf, B. Function classes for identifiable nonlinear independent component analysis. arXiv preprint arXiv:2208.06406, 2022.
  6. 6.Buchholz, S., Rajendran, G., Rosenfeld, E., Aragam, B., Scholkopf, B., and Ravikumar, P. Learning linear causal representations from interventions under general nonlinear mixing. arXiv preprint arXiv:2306.02235, 2023.
  7. 7.Cai, R., Xie, F., Glymour, C., Hao, Z., and Zhang, K. Triad constraints for learning causal structure of latent variables. Advances in neural information processing systems, 32, 2019.
  8. 8.Comon, P. Independent component analysis – a new concept? Signal Processing, 36:287–314, 1994.
  9. 9.Dong, X., Huang, B., Ng, I., Song, X., Zheng, Y., Jin, S., Legaspi, R., Spirtes, P., and Zhang, K. A versatile causal discovery framework to allow causally-related hidden variables. In The Twelfth International Conference on Learning Representations, 2023.
  10. 10.Gemici, M. C., Rezende, D., and Mohamed, S. Normalizing flows on Riemannian manifolds. arXiv preprint arXiv:1611.02304, 2016.
  11. 11.Halv ¨ a, H. and Hyv ¨ arinen, A. Hidden markov nonlinear ICA: Unsupervised learning from nonstationary time series. In Conference on Uncertainty in Artificial Intelligence, pp. 939–948. PMLR, 2020.
  12. 12.Halv ¨ a, H., Le Corff, S., Leh ¨ ericy, L., So, J., Zhu, Y., Gassiat, ´ E., and Hyvarinen, A. Disentangling identifiable features from noisy data with structured nonlinear ICA. Advances in Neural Information Processing Systems, 34, 2021.
  13. 13.Huang, B., Low, C. J. H., Xie, F., Glymour, C., and Zhang, K. Latent hierarchical causal structure discovery with rank constraints. Advances in Neural Information Processing Systems, 35:5549–5561, 2022.
  14. 14.Hyvarinen, A. and Morioka, H. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. Advances in Neural Information Processing Systems, 29:3765–3773, 2016.
  15. 15.Hyvarinen, A. and Morioka, H. Nonlinear ICA of temporally dependent stationary sources. In International Conference on Artificial Intelligence and Statistics, pp. 460–469. PMLR, 2017.
  16. 16.Hyvarinen, A. and Pajunen, P. Nonlinear independent component analysis: Existence and uniqueness results. Neural networks, 12(3):429–439, 1999.
  17. 17.Hyvarinen, A., Cristescu, R., and Oja, E. A fast algorithm for estimating overcomplete ICA bases for image windows. In Proc. Int. Joint Conf. on Neural Networks, pp. 894–899, Washington, D.C., 1999.
  18. 18.Hyvarinen, A., Sasaki, H., and Turner, R. Nonlinear ICA using auxiliary variables and generalized contrastive learning. In International Conference on Artificial Intelligence and Statistics, pp. 859–868. PMLR, 2019.
  19. 19.Hyvarinen, A., Khemakhem, I., and Morioka, H. Nonlinear independent component analysis for principled disentanglement in unsupervised deep learning. Patterns, 4(10):100844, 2023. ISSN 2666-3899. doi: https://doi.org/10.1016/j.patter.2023.100844. URL https://www.sciencedirect.com/science/article/pii/S2666389923002234.
  20. 20.Jiang, Y. and Aragam, B. Learning nonparametric latent causal graphs with unknown interventions. arXiv preprint arXiv:2306.02899, 2023.
  21. 21.Khemakhem, I., Kingma, D., Monti, R., and Hyvarinen, A. Variational autoencoders and nonlinear ICA: A unifying framework. In International Conference on Artificial Intelligence and Statistics, pp. 2207–2217. PMLR, 2020a.
  22. 22.Khemakhem, I., Monti, R., Kingma, D., and Hyvarinen, A. Ice-beem: Identifiable conditional energy-based deep models based on nonlinear ica. Advances in Neural Information Processing Systems, 33:12768–12778, 2020b.
  23. 23.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  24. 24.Kivva, B., Rajendran, G., Ravikumar, P., and Aragam, B. Learning latent causal graphs via mixture oracles. Advances in Neural Information Processing Systems, 34:18087–18101, 2021.
  25. 25.Kong, L., Xie, S., Yao, W., Zheng, Y., Chen, G., Stojanov, P., Akinwande, V., and Zhang, K. Partial disentanglement for domain adaptation. In International Conference on Machine Learning, pp. 11455–11472. PMLR, 2022.
  26. 26.Lachapelle, S., Lopez, P. R., Sharma, Y., Everett, K., Priol, R. L., Lacoste, A., and Lacoste-Julien, S. Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ICA. Conference on Causal Learning and Reasoning, 2022.
  27. 27.Lin, J. Factorizing multivariate function classes. Advances in neural information processing systems, 10, 1997.
  28. 28.Loh, P.-L. and Buhlmann, P. High-dimensional learning of linear causal networks via inverse covariance estimation. The Journal of Machine Learning Research, 15(1):3065–3105, 2014.
  29. 29.Ng, I., Zheng, Y., Zhang, J., and Zhang, K. Reliable causal discovery with improved exact search and weaker assumptions. In Advances in Neural Information Processing Systems, 2021.
  30. 30.Pearl, J. Causality: Models, Reasoning, and Inference. Cambridge University Press, Cambridge, 2000.
  31. 31.Ramsey, J., Zhang, J., and Spirtes, P. L. Adjacency-faithfulness and conservative causal inference. arXiv preprint arXiv:1206.6843, 2012.
  32. 32.Scholkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Towards causal representation learning. Proceedings of the IEEE, 109(5):612–634, 2021.
  33. 33.Silva, R., Scheines, R., Glymour, C., and Spirtes, P. Learning the structure of linear latent variable models. Journal of Machine Learning Research, 7:191–246, 2006.
  34. 34.Sorrenson, P., Rother, C., and Kothe, U. Disentanglement by nonlinear ICA with general incompressible-flow networks (GIN). arXiv preprint arXiv:2001.04872, 2020.
  35. 35.Spirtes, P., Glymour, C., and Scheines, R. Causation, Prediction, and Search. MIT Press, Cambridge, MA, 2nd edition, 2001.
  36. 36.Squires, C., Seigal, A., Bhate, S. S., and Uhler, C. Linear causal disentanglement via interventions. In International Conference on Machine Learning, 2023.
  37. 37.Strang, G. Linear Algebra and Its Applications. Thomson, Brooks/Cole, Belmont, CA, 4th edition, 2006.
  38. 38.Strang, G. Introduction to Linear Algebra. Wellesley-Cambridge Press, 5th edition, 2016.
  39. 39.Taleb, A. and Jutten, C. Source separation in post-nonlinear mixtures. IEEE Transactions on signal Processing, 47 (10):2807–2820, 1999.
  40. 40.Uhler, C., Raskutti, G., Buhlmann, P., and Yu, B. Geometry of the faithfulness assumption in causal inference. The Annals of Statistics, pp. 436–463, 2013.
  41. 41.Varici, B., Acarturk, E., Shanmugam, K., Kumar, A., and Tajer, A. Score-based causal representation learning with interventions. arXiv preprint arXiv:2301.08230, 2023.
  42. 42.von Kugelgen, J., Besserve, M., Liang, W., Gresele, L., Kekic, A., Bareinboim, E., Blei, D. M., and Sch ¨ olkopf, B. Nonparametric identifiability of causal representations from unknown interventions. arXiv preprint arXiv:2306.00542, 2023.
  43. 43.Xie, F., Cai, R., Huang, B., Glymour, C., Hao, Z., and Zhang, K. Generalized independent noise condition for estimating latent variable causal graphs. In Advances in Neural Information Processing Systems, 2020.
  44. 44.Xie, F., Huang, B., Chen, Z., He, Y., Geng, Z., and Zhang, K. Identification of linear non-gaussian latent hierarchical structure. In International Conference on Machine Learning, pp. 24370–24387. PMLR, 2022.
  45. 45.Yao, W., Sun, Y., Ho, A., Sun, C., and Zhang, K. Learning temporally causal latent processes from general temporal data. arXiv preprint arXiv:2110.05428, 2021.
  46. 46.Yao, W., Chen, G., and Zhang, K. Temporally disentangled representation learning. In Advances in Neural Information Processing Systems, 2022.
  47. 47.Zhang, J. and Spirtes, P. Detection of unfaithfulness and robust causal inference. Minds and Machines, 18:239–271, 2008.
  48. 48.Zhang, J., Greenewald, K., Squires, C., Srivastava, A., Shanmugam, K., and Uhler, C. Identifiability guarantees for causal disentanglement from soft interventions. Advances in Neural Information Processing Systems, 36, 2023.
  49. 49.Zheng, Y. and Zhang, K. Generalizing nonlinear ica beyond structural sparsity. Advances in Neural Information Processing Systems, 36:13326–13355, 2023.
  50. 50.Zheng, Y., Ng, I., and Zhang, K. On the identifiability of nonlinear ICA: Sparsity and beyond. In Advances in Neural Information Processing Systems, 2022.
  51. 51.Zheng, Y., Ng, I., Fan, Y., and Zhang, K. Generalized precision matrix for scalable estimation of nonparametric markov networks. arXiv preprint arXiv:2305.11379, 2023.

Citation

MLA
Zhang, K., et al. “Causal Representation Learning from Multiple Distributions: A General Setting”. arXiv, 2024, http://arxiv.org/abs/2402.05052v3.
APA
Zhang, K., Xie, S., Ng, I., & Zheng, Y. (2024). Causal Representation Learning from Multiple Distributions: A General Setting. arXiv. http://arxiv.org/abs/2402.05052v3
Chicago
Zhang, K., S. Xie, I. Ng, and Y. Zheng. 2024. “Causal Representation Learning from Multiple Distributions: A General Setting”. arXiv. http://arxiv.org/abs/2402.05052v3.
Harvard
Zhang, K. et al. (2024) “Causal Representation Learning from Multiple Distributions: A General Setting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05052v3.
Vancouver
1. Zhang K, Xie S, Ng I, Zheng Y (2024) Causal Representation Learning from Multiple Distributions: A General Setting. arXiv

BibTeX

@article{zhang2024causal,
  title = {Causal Representation Learning from Multiple Distributions: A General Setting},
  author = {Zhang, Kun and Xie, Shaoan and Ng, Ignavier and Zheng, Yujia},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05052v3},
  eprint = {2402.05052}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/