Position: Amazing Things Come From Having Many Good Models
Cynthia RudinChudi ZhongLesia SemenovaMargo I. SeltzerRonald ParrJiachang LiuSrikar KattaJon DonnellyHarry ChenZachery Boner
Demonstrates how the existence of many equally accurate predictive models eliminates the standard accuracy-interpretability trade-off and allows practitioners to select simple, fair models without losing performance.
Machine learning systems are increasingly used in high-stakes decisions such as criminal justice, loan approvals, and healthcare. The standard machine learning paradigm typically trains an algorithm to produce a single "optimal" model, which is often an uninterpretable black box. The article investigates a fundamental phenomenon known as the Rashomon Effect: the reality that for a given dataset, there frequently exist many vastly different predictive models that achieve approximately equal accuracy.
The main objective of the article is to demonstrate how the presence of many equally good models reshapes foundational machine learning assumptions. It evaluates why this phenomenon occurs, proves that simple and interpretable models can match complex black box performance on noisy tabular data, and introduces a framework to capture and navigate the entire pool of top-performing models (the Rashomon set).
To demonstrate this, the authors synthesize theoretical insights and empirical analyses across standard real-world tabular datasets, including credit scoring from FICO and criminal recidivism from COMPAS. They examine performance across multiple distinct model architectures, including deep neural networks, boosted decision trees, support vector machines, and simple decision trees or sparse scoring systems, while assessing modern algorithms designed to efficiently compute the entire Rashomon set.
The article establishes several key findings. First, on noisy tabular datasets, simple and fully interpretable models achieve predictive accuracy comparable to complex black box models—for example, a simple 7-leaf decision tree or a sparse 11-feature model matched the ~72% accuracy and ~0.79–0.80 AUC of deep neural networks and boosted trees on the FICO credit dataset. Second, the theoretical root of the Rashomon Effect is outcome noise: because noisy labels cause high variance and generalization risk, restricting models to simpler function classes yields large sets of high-performing alternatives without sacrificing accuracy. Third, relying on a single model creates severe risks because different equally accurate models rely on completely different variables and yield conflicting predictions, meaning single-model variable importance or fairness evaluations can be misleading artifacts. Fourth, modern specialized algorithms (such as TreeFARMS and FasterRisk) can now compute and store vast Rashomon sets within seconds to minutes, allowing practitioners to optimize for secondary criteria such as fairness and monotonicity with zero loss in predictive performance.
These findings carry profound implications for policy, compliance, and organizational risk. Because black box models offer no inherent performance advantage on noisy tabular data, their deployment creates unnecessary risks regarding unfaithfulness, hidden flaws, and unverifiable post-hoc explanations. The evidence refutes the long-standing assumption of an inevitable trade-off between predictive accuracy and model interpretability or fairness.
The article recommends that organizations and policymakers mandate interpretable models as the default standard for high-stakes societal decisions where outcomes are noisy. Furthermore, data science teams should abandon the single-model workflow in favor of exploring Rashomon sets using interactive interfaces, enabling domain experts to select models aligned with operational constraints and fairness standards. The primary limitations and boundary conditions are that these conclusions apply specifically to noisy tabular data; deterministic problems or unstructured data domains (such as pure computer vision or image segmentation) may still require complex architectures. Overall, confidence in these findings is high, supported by both theoretical proofs and extensive empirical validation across multiple domains.
- Paper: All Models are Wrong, but Many are Useful: Learning a Variable’s Importance by Studying an Entire Class of Prediction Models Simultaneously, Aaron Fisher et al. (2018). This seminal paper introduces Model Class Reliance and formalizes analyzing the entire Rashomon set of high-performing models rather than single models, providing the foundational methodology and concepts on which the source builds.
- Paper: Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead, Cynthia Rudin (2019). This foundational work argues against black-box models in high-stakes decisions and demonstrates that simple interpretable models can match black-box performance on tabular data, establishing core arguments expanded in the source position paper.
- Paper: Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission, Rich Caruana et al. (2015). This study provides foundational empirical proof that glass-box interpretable generalized additive models match complex black-box accuracy in high-stakes tabular domains such as healthcare risk prediction.
- Paper: The Mythos of Model Interpretability, Zachary C. Lipton (2016). This paper establishes the taxonomy and formal distinction between post-hoc explanations and inherently transparent models, clarifying theoretical motivations essential to understanding the source's claims.
- Paper: Very Simple Classification Rules Perform Well on Most Commonly Used Datasets, ROBERT C. HOLTE (1993). This classic benchmark study provides early empirical evidence that very simple rule-based models often achieve predictive accuracy comparable to complex classifiers across diverse standard datasets.
No sufficiently relevant recommendations were found.
