Causal-learn: Causal Discovery in Python
Yujia ZhengBiwei HuangWei ChenJoseph D. RamseyMingming GongRuichu CaiShohei ShimizuPeter SpirtesKun Zhang
Presents causal-learn, a native Python open-source library that unifies major constraint-based, score-based, and functional causal discovery algorithms alongside modular independence tests and evaluation metrics without relying on R or Java dependencies.
Understanding cause-and-effect relationships is vital across scientific research and industry, from identifying disease-causing genes in healthcare to optimizing complex engineering systems. While randomized controlled experiments are the gold standard for discovering these relationships, they are frequently too costly, slow, or logistically impractical to execute. Consequently, organizations increasingly seek to discover causal connections directly from observational data that already exist in large volumes. However, existing software tools for causal discovery have historically relied on legacy languages such as Java and R, creating substantial barriers for modern data science pipelines centered around Python.
The article introduces and evaluates causal-learn, an open-source software library built entirely in Python to make causal discovery accessible, extensible, and practical across scientific and enterprise workflows. The platform provides a unified environment containing the major algorithmic approaches required to extract causal graphs from observational datasets.
To address existing software gaps, the project implements four primary causal discovery paradigms within a modular framework: constraint-based methods, score-based methods, functional causal model methods, and techniques designed to identify hidden or unobserved variables. Crucially, the platform operates purely natively in Python, eliminating foreign-language wrappers and external dependencies that previously complicated software deployment and customization. The authors provide standardized programming interfaces, prepackaged benchmark datasets with known causal graphs, and standalone modules for statistical independence testing, scoring, and graph transformation.
The findings and technical contributions center on several core capabilities. First, causal-learn integrates all four dominant causal discovery categories into a single platform, uniting classic algorithms with recent state-of-the-art methods that model non-linear relations and hidden variables. Second, its native Python architecture eliminates deployment overhead and dependency conflicts, allowing organizations to integrate causal modeling directly into production machine learning workflows and downstream causal inference pipelines. Third, the modular design isolates statistical tests and graph operations, enabling analysts to customize algorithms without rewriting core infrastructure. Finally, the inclusion of curated benchmark datasets helps mitigate the persistent industry challenge of evaluating causal discovery accuracy against verifiable ground truths.
These capabilities significantly lower the technical and operational barriers to deploying causal discovery. Organizations can reduce experimentation costs and accelerate research by screening observational data before committing to expensive physical trials. Furthermore, the platform integrates smoothly with downstream inference tools, enabling end-to-end causal decision pipelines that operate within standard enterprise Python environments.
Organizations should adopt causal-learn for observational analysis workflows while planning to link discovered graphs directly to downstream causal inference frameworks. Because causal discovery methods inherently rely on statistical assumptions such as distribution types or the presence of hidden variables, analysts should carefully evaluate whether data meet specific algorithm requirements. Continued development and community curation of real-world benchmark datasets remain essential to validate performance across diverse operating domains.
- Paper: A Linear Non-Gaussian Acyclic Model for Causal Discovery, Shohei Shimizu et al. (2006). Read the original LiNGAM paper first to understand the non-Gaussian identification assumptions behind the causal-discovery method implemented in causal-learn.
- Paper: DAGs with NO TEARS: Continuous Optimization for Structure Learning, Xun Zheng et al. (2018). NOTEARS establishes the continuous optimization approach to DAG learning that helps explain one of the structure-learning methods available in causal-learn.
- Paper: Optimal Structure Identification With Greedy Search, David Maxwell Chickering (2002). Chickering’s treatment of greedy equivalence search provides the algorithmic foundation for understanding causal-learn’s GES implementation.
- Paper: Equivalence and Synthesis of Causal Models, Tom S. Verma et al. (1990). Verma and Pearl’s criteria for causal-model equivalence clarify the graphical foundations behind constraint-based discovery methods represented in causal-learn.
No sufficiently relevant recommendations were found.
