DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models

Patrick BlöbaumPeter GötzKailash BudhathokiAtalanti-Anastasia MastakouriDominik Janzing

article2024JMLR84 citations

Extends the DoWhy Python library with graphical causal models to perform complex causal tasks beyond standard effect estimation, including root cause analysis of outliers, distributional change attribution, and counterfactual estimation.

Listen

Modern data science increasingly requires organizations to look beyond simple correlations to understand the causal mechanisms driving complex systems. While most existing computational tools focus almost entirely on estimating the effect of an intervention, practical business and engineering questions often demand deeper insights. Organizations frequently need to pinpoint the root causes of system failures, explain unexpected distributional shifts, and evaluate complex 'what-if' counterfactual scenarios.

The article demonstrates and evaluates DoWhy-GCM, an open-source extension to the DoWhy Python library that leverages graphical causal models to answer a broad spectrum of causal questions beyond traditional effect estimation.

To achieve this, the system models an entire system as a directed causal graph where each node represents a variable with an assigned data-generating mechanism. The approach combines tabular observational data—whether continuous, discrete, or categorical—with graphical causal models grounded in formal causal theory. The framework employs a modular three-step workflow: defining the causal graph, fitting mechanism parameters automatically or with custom statistical models, and executing diverse causal queries on the fitted model.

The article highlights several core capabilities. First, the framework enables root-cause analysis by attributing system anomalies and distributional shifts directly to specific upstream components. Second, it quantifies causal influence by measuring edge strength and isolating the intrinsic contribution of individual variables to overall uncertainty. Third, it supports advanced 'what-if' reasoning, including both interventional simulations and point- or population-level counterfactual estimation. Fourth, it provides statistical falsification tools to rigorously test whether graph structures and causal mechanism assumptions align with observed data. Finally, the framework integrates seamlessly with standard data science software via a functional programming design that supports both parametric and non-parametric estimators.

These capabilities mean organizations can diagnose operational risks and anomalies much faster and more accurately by evaluating full causal pathways rather than isolated treatment effects. Incorporating automated mechanism assignment and graph falsification reduces the risk of making high-stakes decisions based on unverified structural assumptions.

Teams seeking to analyze complex system interactions should consider piloting this tool for root-cause analysis and operational diagnosis. However, stakeholders should note that computational scalability depends heavily on graph size, sample volume, and model complexity. Decisions remain contingent on the validity of the underlying causal graph and modeling assumptions, making the use of built-in falsification tests essential prior to operational deployment.

  • Paper: Equivalence and Synthesis of Causal Models, Tom S. Verma et al. (1990). Verma and Pearl establish the graphical-model equivalence concepts that underpin causal graphs and make the source’s graph-based workflow easier to interpret.
  • Paper: Causal structure-based root cause analysis of outliers, Kailash Budhathoki et al. (2022). This framework develops causal root-cause attribution for outliers, providing a direct methodological foundation for DoWhy-GCM’s anomaly diagnosis capabilities.
  • Paper: Direct and Indirect Effects, Judea Pearl (2001). Pearl’s account of direct and indirect effects supplies the path-specific causal reasoning behind the source’s broader intervention and counterfactual queries.
Cover for DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models

Table of Contents

  • 1. Introduction
  • 2. DoWhy-GCM Three-Step Recipe
  • 4. Code design
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — DoWhy-GCM extends causal analysis beyond effect estimation

    model/method

    DoWhy-GCM extends the DoWhy Python library with a framework for causal queries built around graphical causal models. Users provide a directed acyclic graph (DAG) and observational tabular data for the variables in that graph, then fit causal mechanisms and reuse the resulting model for queries. The data may contain continuous, discrete, or categorical variables. The graph serves as a shared representation across tasks, including causal prediction and out-of-distribution prediction, causal influence measurement, interventions and counterfactuals, and attribution of anomalies or distribution changes.

  2. Knowl 2 — Graphical causal model representation and modularity

    definition

    A graphical causal model (GCM) consists of a DAG over random variables and a causal mechanism at each node. For a DAG with variables X1,…,XnX_1,\ldots,X_n, let Pai\mathrm{Pa}_i denote the parent variables of XiX_i; each node's mechanism specifies its conditional distribution given its parents, and the joint distribution is represented as P(X1,…,Xn)=∏i=1nP(Xi∣Pai)P(X_1,\ldots,X_n)=\prod_{i=1}^{n}P(X_i\mid\mathrm{Pa}_i). A root node has no parents and therefore has a marginal distribution. The mechanisms are assumed modular: changing one variable's mechanism does not change the mechanisms of other variables.

  3. Knowl 3 — Three-step workflow for causal queries

    model/method

    DoWhy-GCM organizes causal analysis into three steps: (1) specify a causal graph and assign a mechanism to every node; (2) fit mechanism parameters from data when needed; and (3) ask one or more causal queries using the fitted GCM. Users can supply their own mechanisms, including mechanisms they regard as ground truth, rather than fitting all mechanisms from data. The same fitted GCM can be reused across different queries.

  4. Knowl 4 — Mechanism assignment and noise models

    model/method

    DoWhy-GCM provides automatic assignment of causal mechanisms from observational data, with built-in support for additive noise models (ANMs) and post-nonlinear models; users can also provide custom mechanisms. In an ANM for the graph X→YX\to Y, the root variable and its child can be represented as X=NXX=N_X and Y=f(X)+NYY=f(X)+N_Y, where NXN_X and NYN_Y are noise variables assumed independent and ff is a prediction function, such as a linear or nonlinear regression model. A root mechanism can use either an empirical distribution or a parametric model. When a mechanism permits reconstruction of its noise from an observation, that reconstruction can support counterfactual estimation.

  5. Knowl 5 — Causal model evaluation and graph falsification

    model/method

    DoWhy-GCM includes tools for evaluating a proposed causal model against observational data. Its graph-falsification functionality tests whether a given DAG can be falsified using those data, while causal-model evaluation assesses the assigned mechanisms and their modeling assumptions. These tools address statistically falsifiable implications of the model; they are evaluation procedures, not a guarantee that a graph is causally correct.

  6. Knowl 6 — Quantifying edge and intrinsic causal influence

    model/method

    DoWhy-GCM provides two forms of causal influence measurement. The arrow_strength query quantifies the strength of a directed edge in a causal graph. The intrinsic_causal_influence query quantifies how much a source node contributes to uncertainty in a target node. These queries distinguish influence associated with a particular edge from a source node's contribution to target uncertainty.

  7. Knowl 7 — Intervention and counterfactual queries

    model/method

    A fitted DoWhy-GCM can answer what-if questions through intervention and counterfactual queries. interventional_samples generates samples under interventions on nodes in the graph, including interventions that can be more general than standard atomic interventions. counterfactual_samples estimates what data would have been observed under specified interventions for given observed data, supporting point-level as well as population-level counterfactual analysis.

  8. Knowl 8 — Attributing anomalies, distribution changes, and unit-level changes

    model/method

    DoWhy-GCM provides attribution queries that trace observed effects to nodes in a causal graph. attribute_anomalies quantifies node contributions to anomalies, including upstream root causes. distribution_change quantifies each node's contribution to a change in a target's distribution. parent_relevance measures the relevance of a causal parent to its mechanism while explicitly incorporating noise. unit_change attributes a target-value change for a particular statistical unit to input features and changes in prediction mechanisms.

  9. Knowl 9 — Modular and functional library design

    model/method

    DoWhy-GCM defines interfaces for key components—the causal graph, causal model, and causal mechanisms—so users can supply custom models and algorithms or integrate third-party libraries. Its API follows a functional design: operations act on a GCM and return results rather than relying on stateful delegation. The library exposes model components for inspection, provides defaults and convenience functions, and builds on commonly used tools including NetworkX, NumPy, and Pandas; wrappers also enable use of models from libraries such as scikit-learn and SciPy.

  10. Knowl 10 — Scalability depends on model and graph characteristics

    limitation

    DoWhy-GCM does not have a single scalability guarantee for its causal queries. The scalability of the provided algorithms depends on the inference complexity of the selected models, the number of variables, the sample size, and the structure of the causal graph.

Coverage note — No substantial contributed material was omitted; the paper presents no quantitative experiments or benchmark results to extract.

References

  1. 1.K. Battocchi, E. Dillon, M. Hei, G. Lewis, P. Oka, M. Oprescu, and V. Syrgkanis. EconML: A Python Package for ML-Based Heterogeneous Treatment Effects Estimation. https://github.com/microsoft/EconML, 2019. Version 0.x.
  2. 2.P. Beaumont, B. Horsburgh, P. Pilgerstorfer, A. Droth, R. Oentaryo, S. Ler, H. Nguyen, G. A Ferreira, Z. Patel, and W. Leong. CausalNex, 10 2021. URL https://github.com/quantumblacklabs/causalnex.
  3. 3.K. Budhathoki, D. Janzing, P. Bloebaum, and H. Ng. Why did the distribution change? In Arindam Banerjee and Kenji Fukumizu, editors, Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pages 1666–1674. PMLR, 13–15 Apr 2021. URL https://proceedings.mlr.press/v130/budhathoki21a.html.
  4. 4.K. Budhathoki, G. Michailidis, and D. Janzing. Explaining the root causes of unit-level changes, 2022a. URL https://arxiv.org/abs/2206.12986.
  5. 5.K. Budhathoki, L. Minorics, P. Blöbaum, and D. Janzing. Causal structure-based root cause analysis of outliers. In International Conference on Machine Learning, pages 2357–2369. PMLR, 2022b.
  6. 6.H. Chen, T. Harinen, J. Lee, M. Yung, and Z. Zhao. CausalML: Python package for causal machine learning, 2020. URL https://arxiv.org/abs/2002.11631.
  7. 7.E. Eulig, A. A. Mastakouri, P. Blöbaum, M. Hardt, and D. Janzing. Toward falsifying causal graphs using a permutation-based test, 2023. URL https://arxiv.org/abs/2305.09565.
  8. 8.A. A. Hagberg, D. A. Schult, and P. J. Swart. Exploring network structure, dynamics, and function using networkx. In Gaël Varoquaux, Travis Vaught, and Jarrod Millman, editors, Proceedings of the 7th Python in Science Conference, pages 11 – 15, Pasadena, CA USA, 2008.
  9. 9.C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. Fernández del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Sheppard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant. Array programming with NumPy. Nature, 585(7825):357–362, September 2020. doi: 10.1038/s41586-020-2649-2. URL https://doi.org/10.1038/s41586-020-2649-2.
  10. 10.P. Hoyer, D. Janzing, J. Mooij, J. Peters, and B Schölkopf. Nonlinear causal discovery with additive noise models. In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, editors, Proceedings of the conference Neural Information Processing Systems (NIPS) 2008, Vancouver, Canada, 2009. MIT Press.
  11. 11.G. W. Imbens and D. B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, 2015.
  12. 12.D. Janzing, D. Balduzzi, M. Grosse-Wentrup, and B. Schölkopf. Quantifying causal influences. Annals of Statistics, 41(5):2324–2358, 2013.
  13. 13.D. Janzing, L. Minorics, and P. Blöbaum. Feature relevance quantification in explainable AI: A causal problem. In S. Chiappa and R. Calandra, editors, Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pages 2907–2916, Online, 26–28 Aug 2020. PMLR.
  14. 14.D. Janzing, P. Blöbaum, A. A Mastakouri, P. M Faller, L. Minorics, and K. Budhathoki. Quantifying intrinsic causal contributions via structure preserving interventions. In Sanjoy Dasgupta, Stephan Mandt, and Yingzhen Li, editors, Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 of Proceedings of Machine Learning Research, pages 2188–2196. PMLR, 02–04 May 2024. URL https://proceedings.mlr.press/v238/janzing24a.html.
  15. 15.D. Kalainathan and O. Goudet. Causal discovery toolbox: Uncover causal relationships in Python, 2019. URL https://arxiv.org/abs/1903.02278.
  16. 16.S. Lundberg and S. Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 4765–4774. Curran Associates, Inc., 2017.
  17. 17.W. McKinney. Data Structures for Statistical Computing in Python. In Stéfan van der Walt and Jarrod Millman, editors, Proceedings of the 9th Python in Science Conference, pages 56 – 61, 2010. doi: 10.25080/Majora-92bf1922-00a.
  18. 18.J. Miller, C. Hsu, J. Troutman, J. Perdomo, T. Zrnic, L. Liu, Y. Sun, L. Schmidt, and M. Hardt. WhyNot, 2020. URL https://doi.org/10.5281/zenodo.3875775.
  19. 19.R. K Mothilal, A. Sharma, and C. Tan. Explaining machine learning classifiers through diverse counterfactual explanations. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 607–617, 2020.
  20. 20.J. Pearl. Causality: Models, Reasoning, and Inference. Cambridge University Press, New York, NY, 2nd edition, 2009.
  21. 21.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  22. 22.J. Peters, D. Janzing, and B. Schölkopf. Elements of Causal Inference – Foundations and Learning Algorithms. MIT Press, 2017.
  23. 23.J. D Ramsey, K. Zhang, M Glymour, R. S. Romero, B. Huang, I. Ebert-Uphoff, S. Samarasinghe, E. A Barnes, and C. Glymour. Tetrad—a toolbox for causal discovery.
  24. 24.D. B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66:688–701, 1974.
  25. 25.A. Sharma and E. Kiciman. DoWhy: An end-to-end library for causal inference, 2020. URL https://arxiv.org/abs/2011.04216.
  26. 26.P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C J Carey, ˙I. Polat, Y. Feng, E. W. Moore, J. VanderPlas, D. Laxalde, J. Perktold, R. Cimrman, I. Henriksen, E. A. Quintero, C. R. Harris, A. M. Archibald, A. H. Ribeiro, F. Pedregosa, P. van Mulbregt, and SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nature Methods, 17:261–272, 2020. doi: 10.1038/s41592-019-0686-2.
  27. 27.K. Zhang and A. Hyvärinen. On the identifiability of the post-nonlinear causal model. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, Montreal, Canada, 2009.

Citation

MLA
Blöbaum, P., et al. “DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models”. Journal of Machine Learning Research, vol. 25, no. 147, 2024, pp. 1–7, https://www.jmlr.org/papers/v25/22-1258.html.
APA
Blöbaum, P., Götz, P., Budhathoki, K., Mastakouri, A. A., & Janzing, D. (2024). DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models. Journal of Machine Learning Research, 25(147), 1–7. https://www.jmlr.org/papers/v25/22-1258.html
Chicago
Blöbaum, P., P. Götz, K. Budhathoki, A. A. Mastakouri, and D. Janzing. 2024. “DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models”. Journal of Machine Learning Research 25 (147): 1–7. https://www.jmlr.org/papers/v25/22-1258.html.
Harvard
Blöbaum, P. et al. (2024) “DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models”, Journal of Machine Learning Research, 25(147), pp. 1–7. Available at: https://www.jmlr.org/papers/v25/22-1258.html.
Vancouver
1. Blöbaum P, Götz P, Budhathoki K, Mastakouri AA, Janzing D (2024) DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models. Journal of Machine Learning Research 25:1–7

BibTeX

@article{JMLR:v25:22-1258,
  author  = {Patrick Bl{{\"o}}baum and Peter G{{\"o}}tz and Kailash Budhathoki and Atalanti A. Mastakouri and Dominik Janzing},
  title   = {DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {147},
  pages   = {1--7},
  url     = {http://jmlr.org/papers/v25/22-1258.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/