Weakly supervised causal representation learning

Johann BrehmerPim de HaanPhillip LippeTaco S. Cohen

article2022NeurIPS166 citations

Proves that high-level causal variables and mechanisms can be identified from pixel-level data paired across unknown interventions, and introduces implicit latent causal models to learn these structures without optimizing discrete graphs.

Listen

Modern autonomous systems, such as robotics and self-driving vehicles, must understand high-level causal relationships directly from raw, unstructured sensory inputs like camera feeds. Unsupervised learning cannot uniquely identify underlying causal variables and their structures from passive observational data alone. Standard causal discovery approaches typically require active interventions or explicit labels detailing what changed, both of which are costly and difficult to obtain in large-scale settings.

The article demonstrates that weak supervision—specifically paired observations collected before and after random, unlabeled interventions—is sufficient to uniquely identify latent causal variables and their governing structural models. The authors establish a formal mathematical proof showing that latent causal models are identifiable up to variable reordering and scaling. To translate this theoretical proof into a practical system, the article introduces Implicit Latent Causal Models (ILCMs). These models represent causal relationships implicitly within the neural network's transformation functions, avoiding the severe optimization pitfalls of jointly learning explicit causal graphs alongside latent variables.

Evaluation was conducted across multiple environments, including a two-dimensional synthetic benchmark, an altered 3D object rendering suite (Causal3DIdent), and a simulated robotic manipulation environment (CausalCircuit). Across these benchmarks, ILCMs achieved near-perfect disentanglement scores ranging from 0.97 to 0.99, while accurately identifying intervention targets with 96% to 100% precision. In contrast, standard disentanglement baselines that ignore causal structure achieved disentanglement scores as low as 0.34 to 0.35 and frequently reconstructed incorrect causal graphs. Additionally, empirical scaling tests showed that the method reliably recovers causal graphs in systems containing up to approximately 10 continuous variables.

These findings prove that artificial intelligence systems can discover true cause-and-effect relationships from passive demonstrations—such as video footage of an operator manipulating tools—without manual labeling. This capability reduces the human annotation burden and lowers the operational risk of deploying autonomous agents by enabling them to accurately predict downstream effects and evaluate hypothetical counterfactual scenarios. However, the framework currently requires strict boundary conditions: it assumes continuous real-valued variables and perfect single-variable interventions where background noise remains consistent. Performance drops significantly when handling discrete states or scaling beyond 10 variables.

Organizations evaluating this technology should treat it as an enabling foundational concept rather than an off-the-shelf production tool. Technical teams should conduct pilot studies on constrained, continuous-state robotic pipelines while directing future research toward relaxing theoretical assumptions to support discrete attributes, multi-variable interventions, and real-world temporal video feeds.

Cover for Weakly supervised causal representation learning

Abstract

Learning high-level causal representations together with a causal model from unstructured low-level data such as pixels is impossible from observational data alone. We prove under mild assumptions that this representation is however identifiable in a weakly supervised setting. This involves a dataset with paired samples before and after random, unknown interventions, but no further labels. We then introduce implicit latent causal models, variational autoencoders that represent causal variables and causal structure without having to optimize an explicit discrete graph structure. On simple image data, including a novel dataset of simulated robotic manipulation, we demonstrate that such models can reliably identify the causal structure and disentangle causal variables.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Identifiability of latent causal models from weak supervision
  • 3.1 Setup
  • 3.2 Identifiability result
  • 4 Practical latent causal models
  • 4.1 Explicit latent causal models (ELCMs)
  • 4.2 Implicit latent causal models (ILCMs)
  • 5 Experiments
  • 5.1 2D toy experiment
  • 5.2 Causal3DIdent
  • 5.3 CausalCircuit
  • 5.4 Scaling with graph size
  • 6 Discussion
  • References

Knowls

  1. Knowl 1 — Weakly supervised data pairs define the latent causal learning problem

    definition

    A latent causal model (LCM) consists of a faithful acyclic structural causal model (SCM), an observation space X\mathcal{X}, a decoder gg from causal variables to observations, a set of interventions, and a probability distribution over those interventions. In the SCM, each causal variable ziz_i is generated from its exogenous noise ϵi\epsilon_i and the causal variables of its parents: zi=fi(ϵi;zpai)z_i=f_i(\epsilon_i;z_{\mathrm{pa}_i}). The exogenous noises are independent, and each mechanism is invertible and differentiable in its noise argument, with a differentiable inverse, for fixed parent values. The decoder is diffeomorphic onto its image.

    The weakly supervised observation is a pair (x,x~)(x,\tilde{x}) sampled before and after an intervention, without labels for the intervention target. Draw ϵ\epsilon from the independent noise distribution, set z=s(ϵ)z=s(\epsilon) and x=g(z)x=g(z), where ss is the SCM solution function. Draw an intervention target set II from its distribution; for each i∈Ii\in I, draw a new intervention noise ϵ~i\tilde{\epsilon}_i and replace the original mechanism with a parent-independent mechanism, while setting ϵ~i=ϵi\tilde{\epsilon}_i=\epsilon_i for each i∉Ii\notin I. The intervened SCM then gives z~=s~I(ϵ~)\tilde{z}=\tilde{s}_I(\tilde{\epsilon}) and x~=g(z~)\tilde{x}=g(\tilde{z}). Thus the paired samples share the original noise on variables not targeted by the intervention.

  2. Knowl 2 — Identifiability from unlabeled atomic interventions

    theoretical result

    Consider two LCMs with the same observation space. In each model, there are nn real-valued causal variables and corresponding real-valued noise variables; the SCM is faithful and has pointwise diffeomorphic mechanisms, and the decoder is diffeomorphic onto its image. Suppose each intervention set contains the empty intervention and every single-variable perfect stochastic intervention, and each intervention has positive probability. Then the two LCMs induce the same joint distribution of pre- and post-intervention observations, p(x,x~)p(x,\tilde{x}), if and only if they are equivalent up to a permutation of the causal variables and componentwise diffeomorphic reparameterizations of the causal variables and their noises, with the causal mechanisms, decoder, and intervention model correspondingly transformed. Consequently, under these conditions, paired observations with unknown intervention targets identify both the causal representation and its causal structure, but only up to these transformations.

  3. Knowl 3 — Explicit latent causal models use an SCM prior over causal latents

    model/method

    An explicit latent causal model (ELCM) is a variational autoencoder in which the latent variables are the causal variables zz and z~\tilde{z}. Its prior explicitly parameterizes a causal graph and the mechanisms of an SCM: the observational prior factorizes according to the graph into conditional densities p(zi∣zpai)p(z_i\mid z_{\mathrm{pa}_i}), and the graph and mechanisms also determine the conditional distribution of post-intervention variables given pre-intervention variables. An encoder maps observations to causal latents, and a decoder maps latents back to observations; training uses a variational lower bound on the negative log likelihood of the observation pair. The graph can be learned by searching over DAGs or using a differentiable DAG parameterization. The paper reports that optimally trained ELCMs can recover causal structure and disentangle variables in simple settings, but joint optimization of the explicit graph and representations is difficult: the loss can have local minima associated with incorrect graphs.

  4. Knowl 4 — Implicit latent causal models encode causal structure through solution functions

    model/method

    An implicit latent causal model (ILCM) is a variational autoencoder whose latent representations are noise encodings rather than causal variables. For an SCM solution function ss, the pre-intervention encoding is e=s−1(z)e=s^{-1}(z); the post-intervention encoding is e~=s−1(z~)\tilde{e}=s^{-1}(\tilde{z}), the noise that would generate z~\tilde{z} under the unintervened mechanisms. The model encodes observations into noise representations and decodes them back to observations. It represents causal mechanisms through learned conditional diffeomorphisms that transform an intervened noise coordinate into its causal variable, conditioned on the other noise coordinates, rather than by learning an explicit discrete graph. These transformations implicitly encode the solution function, and hence the graph and mechanisms. The paper states that an ILCM corresponds to a unique ELCM; avoiding an explicit graph parameterization is intended to make ILCMs easier to optimize.

  5. Knowl 5 — ILCM prior enforces the noise pattern of an intervention

    equation

    For nn noise coordinates, let e,e~∈Rne,\tilde{e}\in\mathbb{R}^n be the pre- and post-intervention encodings, and let I⊆{1,…,n}I\subseteq\{1,\ldots,n\} be the intervention target set. The ILCM uses an independent standard Gaussian prior for ee and a categorical prior for II. Its conditional prior over e~\tilde{e} is

    p(e~∣e,I)=∏i∉Iδ(e~i−ei)∏i∈Ipz~(zˉi)∣∂zˉi∂e~i∣,zˉi=sˉi(e~i;e∖i).p(\tilde{e}\mid e,I)=\prod_{i\notin I}\delta(\tilde{e}_i-e_i)\prod_{i\in I}p_{\tilde{z}}(\bar{z}_i)\left|\frac{\partial\bar{z}_i}{\partial\tilde{e}_i}\right|, \qquad \bar{z}_i=\bar{s}_i(\tilde{e}_i;e_{\setminus i}).

    Here δ\delta is a point mass, e∖ie_{\setminus i} denotes all coordinates of ee other than ii, and sˉi\bar{s}_i is a learned diffeomorphic scalar transformation in e~i\tilde{e}_i, conditioned on e∖ie_{\setminus i}. The base density pz~p_{\tilde{z}} is standard Gaussian. Thus, coordinates outside the target set are held fixed, while each targeted coordinate is assigned a conditional density through a change of variables. The learned transformations encode the causal solution functions implicitly.

  6. Knowl 6 — ILCM training maximizes a variational bound on paired-data likelihood

    equation

    Let q(I∣x,x~)q(I\mid x,\tilde{x}) be the intervention-target encoder, and let q(e,e~∣x,x~,I)q(e,\tilde{e}\mid x,\tilde{x},I) be the noise-encoding posterior. Let p(x∣e)p(x\mid e) and p(x~∣e~)p(\tilde{x}\mid\tilde{e}) be the observation likelihoods. The ILCM lower-bounds the log likelihood of a pre/post pair by

    log⁡p(x,x~)≥EI∼q(I∣x,x~)E(e,e~)∼q(e,e~∣x,x~,I)[log⁡p(I)+log⁡p(e)+log⁡p(e~∣e,I)−log⁡q(I∣x,x~)−log⁡q(e,e~∣x,x~,I)+log⁡p(x∣e)+log⁡p(x~∣e~)].\begin{aligned} \log p(x,\tilde{x})\geq {}&\mathbb{E}_{I\sim q(I\mid x,\tilde{x})}\mathbb{E}_{(e,\tilde{e})\sim q(e,\tilde{e}\mid x,\tilde{x},I)}\big[\log p(I)+\log p(e)+\log p(\tilde{e}\mid e,I)\\ &-\log q(I\mid x,\tilde{x})-\log q(e,\tilde{e}\mid x,\tilde{x},I)+\log p(x\mid e)+\log p(\tilde{x}\mid\tilde{e})\big]. \end{aligned}

    The priors and encoders in this expression are those of the ILCM, including its intervention-structured conditional prior over e~\tilde{e}. The model is trained by maximizing this bound, or equivalently minimizing the associated variational loss. The paper also reports adding a regularizer based on the negative entropy of the batch-aggregate intervention posterior to discourage collapse onto a lower-dimensional latent submanifold.

  7. Knowl 7 — Changed noise coordinates support intervention inference

    model/method

    Under the ILCM representation, an intervention targeting coordinate set II changes precisely those noise-encoding coordinates: e~i≠ei\tilde{e}_i\neq e_i if and only if i∈Ii\in I, with probability one. The encoder enforces the complementary constraint by setting non-target coordinates of ee and e~\tilde{e} equal. To infer intervention targets from an observation pair, the intervention encoder scores coordinate ii according to the change in the noise encoder's mean, μe,i(x)−μe,i(x~)\mu_{e,i}(x)-\mu_{e,i}(\tilde{x}), using a learnable quadratic function of that difference. This gives the model a target estimate without requiring target labels during training.

  8. Knowl 8 — Experiments use synthetic causal systems, including the CausalCircuit dataset

    experimental setup

    The paper evaluates ILCMs on three settings. The two-dimensional toy data use a nonlinear SCM with graph z1→z2z_1\to z_2 and a randomly initialized normalizing flow mapping causal variables to observations. The adapted Causal3DIdent data contain images with three causal variables—object hue, spotlight hue, and spotlight position—and six different causal graphs; each uses randomly initialized nonlinear mechanisms and heteroskedastic noise, with images rendered at 64×6464\times64 resolution.

    The authors introduce CausalCircuit, a simulated robot-arm system rendered from a fixed-position camera at 512×512×3512\times512\times3 resolution. Its causal variables describe the arm and three lights. The graph has edges from the robot arm to each light, and from the blue and green lights to the red light. A light is more likely to turn on when its button is pressed or when its parent lights are on. Experiments with this dataset model light intensities as continuous causal variables.

  9. Knowl 9 — ILCMs outperform representation baselines and recover the tested graphs

    data/table

    The comparison measures DCI disentanglement score (D), intervention-target accuracy (Acc), and structural Hamming distance (SHD) from the true graph; lower SHD is better. ILCM-E uses ENCO for graph inference, while ILCM-H uses the paper's heuristic graph inference method. The baselines are a disentanglement VAE with ENCO graph inference (dVAE-E), an unstructured β\beta-VAE, and slot attention. The reported results show that both ILCM variants recover the tested graphs exactly except for SHD 0.170.17 for ILCM-H on Causal3DIdent, and achieve higher disentanglement scores than the other methods. The dVAE-E baseline often identifies intervention targets accurately despite substantially worse graph recovery, illustrating that target accuracy alone does not imply structural recovery.

    Could not parse LaTeX table
  10. Knowl 10 — ILCM performance declines as the causal system grows

    empirical result

    To test scaling, the authors generate synthetic nn-dimensional datasets with X=Z=Rn\mathcal{X}=\mathcal{Z}=\mathbb{R}^n. Each dataset uses a linear SCM with a random DAG: edges consistent with a fixed topological ordering are sampled independently with probability 0.50.5. A random rotation from SO(n)SO(n) maps causal variables to observations. ILCMs reliably disentangle variables for systems with up to approximately 10 causal variables, and graph accuracy is also good in this range. For larger systems, both disentanglement and graph accuracy worsen; the paper does not report exact plotted scores in the text.

  11. Knowl 11 — Identifiability and practical performance have important limits

    limitation

    The identifiability theorem requires real-valued causal variables and noises, stochastic perfect interventions, and positive probability for every atomic intervention, including the empty intervention. Its paired-data construction also assumes that the exogenous noise is preserved for variables not targeted by the intervention. The authors note that realistic temporal observations may not satisfy this construction exactly. In experiments, ILCMs work reliably when causal variables are continuous, but have difficulty disentangling discrete light states in CausalCircuit. The practical evaluations are limited to simplified systems with relatively few continuous variables, and performance degrades as graph size increases.

Coverage note — The detailed neural architectures, hyperparameters, and operational specification of the heuristic graph-extraction procedure are omitted because the supplied main paper defers those details to supplementary appendices; the main-text methods and reported results are included.

References

  1. 1.Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. Toward causal representation learning. Proceedings of the IEEE, 109(5):612–634, 2021.
  2. 2.Frederick Eberhardt. Green and grue causal variables. Synthese, 193(4):1029–1046, 2016.
  3. 3.Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. Challenging common assumptions in the unsupervised learning of disentangled representations. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 4114–4124. PMLR, 2019.
  4. 4.Diederik P. Kingma and Max Welling. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014.
  5. 5.Francesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf, Olivier Bachem, and Michael Tschannen. Weakly-supervised disentanglement without compromises. In International Conference on Machine Learning, pages 6348–6359. PMLR, 2020.
  6. 6.Aapo Hyvärinen and Erkki Oja. Independent component analysis: algorithms and applications. Neural Networks, 13:411–430, 2000.
  7. 7.Rui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon, and Ben Poole. Weakly Supervised Disentanglement with Guarantees. In International Conference on Learning Representations, 2020.
  8. 8.Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational Autoencoders and Nonlinear ICA: A Unifying Framework. In Silvia Chiappa and Roberto Calandra, editors, Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pages 2207–2217. PMLR, 2020.
  9. 9.Hermanni Hälvä, Sylvain Le Corff, Luc Lehéricy, Jonathan So, Yongjie Zhu, Elisabeth Gassiat, and Aapo Hyvarinen. Disentangling Identifiable Features from Noisy Data with Structured Nonlinear ICA. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021.
  10. 10.Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf, and Michel Besserve. Independent mechanism analysis, a new concept? In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021.
  11. 11.Weiran Yao, Yuewen Sun, Alex Ho, Changyin Sun, and Kun Zhang. Learning Temporally Causal Latent Processes from General Temporal Data. In International Conference on Learning Representations, 2022.
  12. 12.Sebastien Lachapelle, Pau Rodriguez, Rémi Le, Yash Sharma, Katie E Everett, Alexandre Lacoste, and Simon Lacoste-Julien. Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA. In First Conference on Causal Learning and Reasoning, 2022.
  13. 13.Chaochao Lu, Yuhuai Wu, Jose Miguel Hernández-Lobato, and Bernhard Schölkopf. Nonlinear invariant risk minimization: A causal approach. arXiv preprint arXiv:2102.12353, 2021.
  14. 14.Julius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Schölkopf, Michel Besserve, and Francesco Locatello. Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  15. 15.Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano, Taco Cohen, and Efstratios Gavves. CITRIS: Causal Identifiability from Temporal Intervened Sequences. In Proceedings of the 39th International Conference on Machine Learning, ICML, 2022.
  16. 16.Mengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen, Jianye Hao, and Jun Wang. CausalVAE: Disentangled representation learning via neural structural causal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9593–9602, 2021.
  17. 17.Jeffrey Adams, Niels Hansen, and Kun Zhang. Identification of partially observed linear causal models: Graphical conditions for the non-gaussian and heterogeneous cases. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 22822–22833. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/c0f6fb5d3a389de216345e490469145e-Paper.pdf.
  18. 18.Feng Xie, Ruichu Cai, Biwei Huang, Clark Glymour, Zhifeng Hao, and Kun Zhang. Generalized independent noise condition for estimating latent variable causal graphs. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 14891–14902. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/aa475604668730af60a0a87cc92604da-Paper.pdf.
  19. 19.Bohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, and Bryon Aragam. Learning latent causal graphs via mixture oracles. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 18087–18101. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/966aad8981dcc75b5b8ab04427a833b2-Paper.pdf.
  20. 20.Krzysztof Chalupka, Pietro Perona, and Frederick Eberhardt. Visual causal feature learning. Uncertainty in Artificial Intelligence, December 2014.
  21. 21.Sander Beckers and Joseph Y Halpern. Abstracting causal models. Proc. Conf. AAAI Artif. Intell., 33:2678–2685, July 2019.
  22. 22.Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. Elements of causal inference: foundations and learning algorithms. The MIT Press, 2017.
  23. 23.Paul K Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M Mooij, Dominik Janzing, Moritz Grosse-Wentrup, and Bernhard Schölkopf. Causal consistency of structural equation models. Uncertainty in Artificial Intelligence, July 2017.
  24. 24.Judea Pearl. Causality : models, reasoning, and inference. Cambridge University Press, Cambridge, U.K. New York, 2000. ISBN 978-0521895606.
  25. 25.Stephan Bongers, Patrick Forré, Jonas Peters, and Joris M. Mooij. Foundations of Structural Causal Models with Cycles and Latent Variables. Annals of Statistics, 49(5):2885–2915, 2021. doi: 10.1214/21-AOS2064.
  26. 26.Eric Jang, Shixiang Gu, and Ben Poole. Categorical Reparameterization with Gumbel-Softmax. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, 2017.
  27. 27.Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien, and Alexandre Drouin. Differentiable Causal Discovery from Interventional Data. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  28. 28.Phillip Lippe, Taco Cohen, and Efstratios Gavves. Efficient Neural Causal Discovery without Acyclicity Constraints. In International Conference on Learning Representations, 2022.
  29. 29.Bertrand Charpentier, Simon Kibler, and Stephan Günnemann. Differentiable DAG sampling. March 2022.
  30. 30.Alain Hauser and Peter Bühlmann. Characterization and Greedy Learning of Interventional Markov Equivalence Classes of Directed Acyclic Graphs. Journal of Machine Learning Research, 13(1):2409–2464, 2012. ISSN 1532-4435.
  31. 31.Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. Advances in Neural Information Processing Systems, 33:11525–11538, 2020.
  32. 32.Cian Eastwood and Christopher K I Williams. A framework for the quantitative evaluation of disentangled representations. International Conference on Learning Representations, February 2018.
  33. 33.Emanuel Todorov, Tom Erez, and Yuval Tassa. MuJoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 5026–5033, October 2012.
  34. 34.James Woodward. Making Things Happen: A Theory of Causal Explanation. Oxford University Press, USA, October 2005.

Citation

MLA
Brehmer, J., et al. “Weakly Supervised Causal Representation Learning”. Advances in Neural Information Processing Systems 35, 2022, pp. 38319–31, https://doi.org/10.52202/068431-2776.
APA
Brehmer, J., Haan, P. D., Lippe, P., & Cohen, T. (2022). Weakly Supervised Causal Representation Learning. Advances in Neural Information Processing Systems 35, 38319–38331. https://doi.org/10.52202/068431-2776
Chicago
Brehmer, J., P. D. Haan, P. Lippe, and T. Cohen. 2022. “Weakly Supervised Causal Representation Learning”. Advances in Neural Information Processing Systems 35, 38319–31. https://doi.org/10.52202/068431-2776.
Harvard
Brehmer, J. et al. (2022) “Weakly Supervised Causal Representation Learning”, Advances in Neural Information Processing Systems 35. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 38319–38331. Available at: https://doi.org/10.52202/068431-2776.
Vancouver
1. Brehmer J, Haan PD, Lippe P, Cohen T (2022) Weakly Supervised Causal Representation Learning. In: Advances in Neural Information Processing Systems 35. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 38319–38331

BibTeX

@inproceedings{Brehmer_2022, series={NeurIPS 2022}, title={Weakly Supervised Causal Representation Learning}, url={http://dx.doi.org/10.52202/068431-2776}, DOI={10.52202/068431-2776}, booktitle={Advances in Neural Information Processing Systems 35}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Brehmer, Johann and Haan, Pim De and Lippe, Phillip and Cohen, Taco}, year={2022}, pages={38319–38331}, collection={NeurIPS 2022} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission