Causal Representation Learning from Multiple Distributions: A General Setting
Kun ZhangShaoan XieIgnavier NgYujia Zheng
Establishes theoretical guarantees for completely nonparametric causal representation learning from multiple distributions without requiring hard interventions, demonstrating that latent variables and the underlying moralized causal graph can be recovered under graph sparsity and sufficient mechanism variation conditions.
Modern data systems frequently observe complex, high-dimensional measurements—such as image pixels, audio signals, or system logs—that are generated by underlying, unobserved causal factors. To make reliable predictions under shifting operational environments or to understand the root causes of system behavior, organizations must discover these hidden causal variables and their underlying network structures. However, existing techniques often rely on restrictive functional assumptions (such as linearity) or require active, hard experimental interventions that may be costly, unethical, or technically impossible to execute in production settings.
The article evaluates whether it is mathematically possible to recover latent causal variables and their relationships purely from observational data collected across changing conditions, such as multiple domains, heterogeneous sources, or non-stationary time series. It aims to establish the fundamental boundaries of what can be identified in a fully non-parametric setting without hard interventions, providing a foundational baseline for causal machine learning.
To demonstrate this, the authors develop theoretical proofs grounded in differential geometry and conditional independence analysis. They analyze a data-generating setup where hidden causal variables influence observations through an unknown non-linear mixing process while their internal dynamics shift across environments. The approach is implemented practically using a variational autoencoder architecture combined with sparsity regularization on the estimated causal structure. The framework is verified across synthetic benchmark experiments, including four-node chain and collider network structures evaluated under both Gaussian and Laplacian noise regimes.
The analysis establishes several key findings. First, under a graph sparsity penalty and sufficient distributional diversity across environments, the underlying undirected dependency structure (the Markov network) can be identified up to an exact isomorphism. Second, each estimated latent variable can be recovered as a function of the true variable and a strictly limited set of "intimate neighbors"—variables that share mutual connections across the graph. In many topological configurations where this intimate neighbor set is empty, latent variables are recovered perfectly up to simple one-to-one transformations. Third, the recovered undirected graph corresponds exactly to the moralized graph of the true underlying directed causal network under two new, minimal faithfulness relaxations (single adjacency-faithfulness and single unshielded-collider-faithfulness). Finally, the article formally proves that standard non-linear independent component methods cannot achieve independent representations when underlying causal dependencies exist, underscoring the necessity of explicitly modeling causal structures.
These findings indicate that organizations can recover significant causal insights from naturally occurring distribution shifts—such as seasonal trends, geographic variations, or operational regime changes—without executing disruptive experiments. This reduces the cost and operational risk associated with active A/B testing while mitigating the risk of deploying machine learning models that rely on spurious correlations. Furthermore, the work clarifies the minimal assumptions required for causal representation, helping data teams avoid overly restrictive parametric constraints that could degrade modeling performance.
Technical leaders and practitioners developing representation learning pipelines should adopt sparsity-penalized deep generative models to capture underlying causal factors when multi-environment observational data is available. Before relying on identified factors for automated decision-making, teams should evaluate their system's graph structure to verify whether key variables fall into isolated configurations that guarantee one-to-one identifiability or remain entangled with intimate neighbors.
While the theoretical guarantees are rigorous, users must consider key boundary conditions. The theoretical framework requires data across a sufficient number of distinct environments (at least twice the number of latent variables plus the number of edges) and assumes smooth, non-zero probability densities. Future investigations are required to adapt the framework to scenarios where only a localized subset of causal relations shift across environments.
- Paper: Learning Temporally Causal Latent Processes from General Temporal Data, Weiran Yao et al. (2022). Establishes foundational identifiability conditions for recovering latent causal processes from nonstationary time series, directly preceding this paper's general multi-distribution framework.
- Paper: Linear Causal Disentanglement via Interventions, Chandler Squires et al. (2023). Provides the essential baseline for linear causal representation learning under explicit hard interventions, which the source paper generalizes to nonparametric settings without hard interventions.
- Paper: Identification of Linear Non-Gaussian Latent Hierarchical Structure, Feng Xie et al. (2022). Introduces structural identifiability conditions for latent causal graphs, providing core theoretical principles for discovering unobserved causal variables from observational mixtures.
- Paper: Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations, Francesco Locatello et al. (2018). Proves the fundamental impossibility of unsupervised disentanglement without inductive biases, motivating the source paper's use of multiple distributions and sparsity constraints.
- Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Introduces the principle of exploiting multiple heterogeneous environments to learn invariant causal representations across distribution shifts.
- Paper: Equivalence and Synthesis of Causal Models, Tom S. Verma et al. (1990). Establishes classic graphical equivalence and Markov properties that underpin the source paper's theoretical recovery of moralized latent causal graphs.
- Paper: A Linear Non-Gaussian Acyclic Model for Causal Discovery, Shohei Shimizu et al. (2006). Provides foundational identifiability results for structural causal models from non-Gaussian data, which underpin modern causal representation learning techniques.
- Paper: Identifying Weight-Variant Latent Causal Models, Yuhang Liu et al. (2026). Extends the multi-distribution causal representation framework to weight-variant latent models modulated by auxiliary observed variables without requiring hard interventions.
- Paper: Multi-View Causal Representation Learning with Partial Observability, Dingling Yao et al. (2024). Builds on latent causal identifiability across diverse views by generalizing representation learning guarantees to multi-modal settings with partial observability.
