Position: Amazing Things Come From Having Many Good Models

Cynthia RudinChudi ZhongLesia SemenovaMargo I. SeltzerRonald ParrJiachang LiuSrikar KattaJon DonnellyHarry ChenZachery Boner

article2024ICML66 citations

Demonstrates how the existence of many equally accurate predictive models eliminates the standard accuracy-interpretability trade-off and allows practitioners to select simple, fair models without losing performance.

Listen

Machine learning systems are increasingly used in high-stakes decisions such as criminal justice, loan approvals, and healthcare. The standard machine learning paradigm typically trains an algorithm to produce a single "optimal" model, which is often an uninterpretable black box. The article investigates a fundamental phenomenon known as the Rashomon Effect: the reality that for a given dataset, there frequently exist many vastly different predictive models that achieve approximately equal accuracy.

The main objective of the article is to demonstrate how the presence of many equally good models reshapes foundational machine learning assumptions. It evaluates why this phenomenon occurs, proves that simple and interpretable models can match complex black box performance on noisy tabular data, and introduces a framework to capture and navigate the entire pool of top-performing models (the Rashomon set).

To demonstrate this, the authors synthesize theoretical insights and empirical analyses across standard real-world tabular datasets, including credit scoring from FICO and criminal recidivism from COMPAS. They examine performance across multiple distinct model architectures, including deep neural networks, boosted decision trees, support vector machines, and simple decision trees or sparse scoring systems, while assessing modern algorithms designed to efficiently compute the entire Rashomon set.

The article establishes several key findings. First, on noisy tabular datasets, simple and fully interpretable models achieve predictive accuracy comparable to complex black box models—for example, a simple 7-leaf decision tree or a sparse 11-feature model matched the ~72% accuracy and ~0.79–0.80 AUC of deep neural networks and boosted trees on the FICO credit dataset. Second, the theoretical root of the Rashomon Effect is outcome noise: because noisy labels cause high variance and generalization risk, restricting models to simpler function classes yields large sets of high-performing alternatives without sacrificing accuracy. Third, relying on a single model creates severe risks because different equally accurate models rely on completely different variables and yield conflicting predictions, meaning single-model variable importance or fairness evaluations can be misleading artifacts. Fourth, modern specialized algorithms (such as TreeFARMS and FasterRisk) can now compute and store vast Rashomon sets within seconds to minutes, allowing practitioners to optimize for secondary criteria such as fairness and monotonicity with zero loss in predictive performance.

These findings carry profound implications for policy, compliance, and organizational risk. Because black box models offer no inherent performance advantage on noisy tabular data, their deployment creates unnecessary risks regarding unfaithfulness, hidden flaws, and unverifiable post-hoc explanations. The evidence refutes the long-standing assumption of an inevitable trade-off between predictive accuracy and model interpretability or fairness.

The article recommends that organizations and policymakers mandate interpretable models as the default standard for high-stakes societal decisions where outcomes are noisy. Furthermore, data science teams should abandon the single-model workflow in favor of exploring Rashomon sets using interactive interfaces, enabling domain experts to select models aligned with operational constraints and fairness standards. The primary limitations and boundary conditions are that these conclusions apply specifically to noisy tabular data; deterministic problems or unstructured data domains (such as pure computer vision or image segmentation) may still require complex architectures. Overall, confidence in these findings is high, supported by both theoretical proofs and extensive empirical validation across multiple domains.

arXiv: 2407.04846

No sufficiently relevant recommendations were found.

Cover for Position: Amazing Things Come From Having Many Good Models

Abstract

The Rashomon Effect, coined by Leo Breiman, describes the phenomenon that there exist many equally good predictive models for the same dataset. This phenomenon happens for many real datasets and when it does, it sparks both magic and consternation, but mostly magic. In light of the Rashomon Effect, this perspective piece proposes reshaping the way we think about machine learning, particularly for tabular data problems in the nondeterministic (noisy) setting. We address how the Rashomon Effect impacts (1) the existence of simple-yet-accurate models, (2) flexibility to address user preferences, such as fairness and monotonicity, without losing performance, (3) uncertainty in predictions, fairness, and explanations, (4) reliable variable importance, (5) algorithm choice, specifically, providing advanced knowledge of which algorithms might be suitable for a given problem, and (6) public policy. We also discuss a theory of when the Rashomon Effect occurs and why. Our goal is to illustrate how the Rashomon Effect can have a massive impact on the use of machine learning for complex problems in society.

Table of Contents

  • 1. Introduction
  • 2. The Rashomon Effect is Everywhere
  • 3. The Rashomon Effect Gives Rise to Simpler-Yet-Accurate Models
  • 4. The Standard ML Paradigm is Too Narrow
  • 5. A New Paradigm: Finding Rashomon Sets
  • 6. Why Does the Rashomon Effect Occur?
  • 7. Uncertainty in Predictions, Fairness, and Explanations
  • 8. Stable Variable Importance
  • 9. Which Algorithm Should I Use?
  • 10. Policy Implications
  • 11. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Metrics to Gauge the Rashomon Effect

Knowls

  1. Knowl 1 — Rashomon Set and Rashomon Ratio

    definition

    In a supervised learning problem with dataset DD, hypothesis class F\mathcal{F}, and empirical loss function L(f,D)L(f, D), the Rashomon set R(F,ϵ)R(\mathcal{F}, \epsilon) is defined as the subset of models in F\mathcal{F} whose empirical loss is within a tolerance ϵ≥0\epsilon \ge 0 of the loss of the empirical risk minimizer f∗=arg⁡min⁡f∈FL(f,D)f^* = \arg\min_{f \in \mathcal{F}} L(f, D):

    R(F,ϵ)={f∈F:L(f,D)≤L(f∗,D)+ϵ}.R(\mathcal{F}, \epsilon) = \{f \in \mathcal{F} : L(f, D) \le L(f^*, D) + \epsilon\}.

    The Rashomon ratio is the fraction of models within the hypothesis class F\mathcal{F} that belong to the Rashomon set R(F,ϵ)R(\mathcal{F}, \epsilon), quantifying the relative volume or proportion of near-optimal models in the function class.

  2. Knowl 2 — Existence of Simpler Accurate Models via Rashomon Packing Numbers

    theoretical result

    Let F\mathcal{F} denote a complex function class and G⊂F\mathcal{G} \subset \mathcal{F} denote a simpler (interpretable) function class. Suppose G\mathcal{G} forms a δ\delta-cover of F\mathcal{F} under a metric dd on the function space, meaning that for every f∈Ff \in \mathcal{F}, there exists g∈Gg \in \mathcal{G} such that d(f,g)≤δd(f, g) \le \delta.

    For any Rashomon set R(F,ϵ)⊆FR(\mathcal{F}, \epsilon) \subseteq \mathcal{F}, the number of simpler models g∈Gg \in \mathcal{G} contained within R(F,ϵ)R(\mathcal{F}, \epsilon) is bounded below by the 2δ2\delta-packing number of R(F,ϵ)R(\mathcal{F}, \epsilon), defined as the maximum number of disjoint balls of radius δ\delta with centers in R(F,ϵ)R(\mathcal{F}, \epsilon) separated by at least 2δ2\delta.

    Because each ball in the 2δ2\delta-packing contains at least one model from G\mathcal{G}, a larger Rashomon set in a complex function space necessarily guarantees the existence of a larger set of accurate, simpler models.

  3. Knowl 3 — Four-Step Mechanism from Outcome Uncertainty to Simple Model Sufficiency

    theoretical result

    In non-deterministic data generation processes where outcomes contain inherent noise, simpler predictive models naturally achieve competitive test accuracy via a four-step causal sequence:

    1. Variance Inflation: Label noise increases the variance of the loss function with respect to random samples drawn from the data distribution.
    2. Generalization Gap Expansion: Increased loss variance widens the generalization gap between empirical training error and expected test error, elevating the risk of overfitting complex functions.
    3. Function Space Simplification: To prevent fitting stochastic noise and to reduce validation error, the analyst must restrict the hypothesis class to simpler functions.
    4. Rashomon Ratio Expansion: Simpler function classes have higher Rashomon ratios (a greater fraction of models performing near the optimal level). When simpler functions serve as covers for the complex function space, models spanning a wide range of complexity belong to the same large Rashomon set, ensuring that sparse, interpretable models match the accuracy of complex black boxes.
  4. Knowl 4 — Interactive Machine Learning via Rashomon Set Enumeration

    model/method

    Rather than optimizing a single predictive model under fixed objective weights, the Rashomon set paradigm computes or enumerates the collection of all near-optimal models within a specified simpler function class (e.g., all sparse decision trees within a loss tolerance via TreeFARMS, all sparse generalized additive models via GAM Rashomon set search, or all sparse linear scoring systems via FasterRisk).

    This collection can be explored through interactive visual interfaces (such as TimberTrek or GAM Changer). Analysts and domain experts can inspect candidate models in real time to select models that satisfy domain-specific constraints—such as feature monotonicity, fairness metrics, or sparsity requirements—without reformulating or re-running expensive optimization algorithms.

  5. Knowl 5 — Containment of Multi-Objective Rashomon Sets in Error-Bounded Sets

    theoretical result

    A Rashomon set constructed using a standard surrogate loss or empirical misclassification error threshold ϵ\epsilon covers the Rashomon sets defined under secondary or non-convex objectives (such as AUC, F1-score, cross-entropy loss, or algorithmic fairness constraints) under corresponding thresholds.

    Because models with near-optimal classification error exhibit near-optimal performance across correlated metrics, computing the Rashomon set for a computationally tractable primary loss allows the identification of models satisfying multiple alternative criteria without directly solving complex multi-objective optimization problems.

  6. Knowl 6 — Predictive Multiplicity and Variable Importance Discrepancy Across the Rashomon Set

    empirical result

    Models within a Rashomon set that exhibit near-identical test accuracy and AUC can exhibit substantial predictive multiplicity (assigning contradictory predictions to the same instance) and rely on completely different feature subsets.

    On the 23-feature FICO credit scoring dataset, boosted decision trees, support vector machines (with linear and RBF kernels), logistic regression, and 2-layer additive risk models all achieve comparable test accuracy (0.70−0.740.70-0.74) and test AUC (0.76−0.810.76-0.81). However, their multiplicative permutation feature importance (model reliance, where 1.001.00 indicates no effect on loss and values >1.00>1.00 indicate loss increase upon feature permutation) diverges significantly:

    • Boosted trees rely most heavily on ExternalRiskEstimate (1.18±0.021.18 \pm 0.02).
    • Support vector machines with RBF kernel rely heavily on both ExternalRiskEstimate (1.15±0.011.15 \pm 0.01) and NetFractionRevolvingBurden (1.12±0.001.12 \pm 0.00).
    • Logistic regression relies most on NetFractionRevolvingBurden (1.23±0.021.23 \pm 0.02) and minimally on NumInqLast6M (1.02±0.001.02 \pm 0.00).
    • The 2-layer additive risk model relies primarily on NumInqLast6M (1.08±0.011.08 \pm 0.01) while exhibiting zero reliance on NetFractionRevolvingBurden (1.00±0.001.00 \pm 0.00).
  7. Knowl 7 — Performance Parity Between Black-Box and Interpretable Models on Tabular Data

    data/table

    On the 23-feature FICO Explainable Machine Learning Challenge dataset, evaluated over 10 test folds, interpretable models perform on par with or outperform complex black-box classifiers:

    Classifier Test Accuracy Test AUC
    Random Forest 0.697±0.0170.697 \pm 0.017 0.757±0.0170.757 \pm 0.017
    Boosted trees 0.723±0.0240.723 \pm 0.024 0.789±0.0280.789 \pm 0.028
    SVM (linear kernel) 0.720±0.0290.720 \pm 0.029 0.795±0.0270.795 \pm 0.027
    SVM (RBF kernel) 0.727±0.0230.727 \pm 0.023 0.799±0.0220.799 \pm 0.022
    8-layer neural network 0.722±0.0220.722 \pm 0.022 0.792±0.0260.792 \pm 0.026
    Logistic regression 0.731±0.0230.731 \pm 0.023 0.801±0.0280.801 \pm 0.028
    2-layer additive risk model 0.738±0.0200.738 \pm 0.020 0.806±0.0250.806 \pm 0.025

    The 2-layer additive risk model achieves the highest test accuracy (0.7380.738) and test AUC (0.8060.806) among all evaluated architectures. In addition, a single sparse decision tree of depth ∼5\sim 5 with 7 leaves generated by the GOSDT algorithm achieves approximately 72%72\% train and test accuracy, and a sparse generalized additive model obtained via FastSparse achieves a cross-validation AUC of 0.791±0.0100.791 \pm 0.010 and accuracy of 72.4%±1.2%72.4\% \pm 1.2\%, confirming the absence of an accuracy-interpretability trade-off.

  8. Knowl 8 — Rashomon Importance Distribution for Robust Feature Significance

    model/method

    The Rashomon Importance Distribution (RID) evaluates variable importance across the entire Rashomon set of well-performing models over multiple bootstrap perturbations of the training dataset.

    For any chosen variable importance metric and feature, RID computes the full empirical distribution of importance values rather than a single point estimate. By integrating over both model multiplicity within the Rashomon set and sampling variability via bootstrapping, RID yields stable and reproducible feature rankings, preventing contradictory scientific or legal findings caused by arbitrary single-model selection.

  9. Knowl 9 — Default Mandate for Inherently Interpretable Models in High-Stakes Decisions

    model/method

    In non-deterministic tabular domains (such as criminal justice recidivism, bail and parole determinations, clinical risk scoring, and credit lending), empirical performance differences between complex black boxes and optimized sparse models are negligible due to the Rashomon Effect.

    Post-hoc explanation methods applied to black-box models are inherently vulnerable to unfaithfulness (misrepresenting internal model logic) and incompleteness. Because Rashomon set algorithms allow direct construction of sparse decision trees, generalized additive models, and scoring systems that achieve baseline performance, automated high-stakes decision systems should default to inherently interpretable models whose reasoning is complete and verifiable by design.

  10. Knowl 10 — Quantitative Metrics for Evaluating the Rashomon Effect

    definition

    Several quantitative metrics evaluate different aspects of the Rashomon set R(F,ϵ)R(\mathcal{F}, \epsilon):

    • Rashomon Set Size: Quantified by the count of sparse decision trees (via TreeFARMS), the count of unique feature support sets for generalized additive models, or the parameter space volume (in closed form for ridge regression).
    • Ambiguity: The proportion of instances in a dataset whose predicted label differs across at least two models in the Rashomon set.
    • Discrepancy: The maximum fraction of predictions that change relative to a baseline deployed model when switching to any other model within the Rashomon set.
    • Rashomon Capacity: A metric measuring predictive multiplicity for probabilistic classifiers over the Rashomon set.
    • Model Class Reliance (MCR): The range [min⁡f∈R,max⁡f∈R][\min_{f \in R}, \max_{f \in R}] of permutation variable importance values attained by models in the Rashomon set for a given feature.

Coverage note — None was omitted; all central definitions, theoretical mechanisms, empirical benchmarks, methodological frameworks, and policy arguments were captured.

References

  1. 1.Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. Sanity checks for saliency maps. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Neural Information Processing Systems (NeurIPS), volume 31, 2018.
  2. 2.Ahanor, I., Medal, H., and Trapp, A. C. Diversitree: Computing diverse sets of near-optimal solutions to mixed-integer optimization problems. arXiv preprint arXiv:2204.03822, 2022.
  3. 3.Angelino, E., Larus-Stone, N., Alabi, D., Seltzer, M., and Rudin, C. Learning certifiably optimal rule lists for categorical data. Journal of Machine Learning Research, 18 (234):1–78, 2018.
  4. 4.Barron, A. R. Universal approximation bounds for superpositions of a sigmoidal function. IEEE Transactions on Information Theory, 39(3):930–945, 1993.
  5. 5.Black, E., Raghavan, M., and Barocas, S. Model multiplicity: Opportunities, concerns, and solutions. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 850–863, 2022.
  6. 6.Black, E., Koepke, J. L., Kim, P., Barocas, S., and Hsu, M. Less discriminatory algorithms. Georgetown Law Journal, 113(1), 2024. Washington University in St Louis Legal Studies Research Paper (forthcoming).
  7. 7.Breiman, L. Random Forests. Machine Learning, 45(1): 5–32, 2001a.
  8. 8.Breiman, L. Statistical modeling: The two cultures (with comments and a rejoinder by the author). Statistical Science, 16(3):199–231, 2001b.
  9. 9.Breiman, L., Friedman, J., Stone, C. J., and Olshen, R. A. Classification and Regression Trees. CRC press, 1984.
  10. 10.Chen, C., Lin, K., Rudin, C., Shaposhnik, Y., Wang, S., and Wang, T. A holistic approach to interpretability in financial lending: Models, visualizations, and summary-explanations. Decision Support Systems, 152:113647, 2022.
  11. 11.Coker, B., Rudin, C., and King, G. A theory of statistical inference for ensuring the robustness of scientific results. Management Science, 67(10):6174–6197, 2021.
  12. 12.Cooper, A. F., Lee, K., Choksi, M., Barocas, S., Sa, C. D., Grimmelmann, J., Kleinberg, J., Sen, S., and Zhang, B. Arbitrariness and prediction: The confounding role of variance in fair classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 22004–22012, Mar. 2024.
  13. 13.Coston, A., Rambachan, A., and Chouldechova, A. Characterizing fairness over the set of good models under selective labels. In Proceedings of the International Conference on Machine Learning (ICML), volume 139, pp. 2144–2155, 18–24 Jul 2021.
  14. 14.Dong, J. and Rudin, C. Exploring the cloud of variable importance for the set of all good models. Nature Machine Intelligence, 2(12):810–824, 2020.
  15. 15.Donnelly, J., Katta, S., Rudin, C., and Browne, E. P. The rashomon importance distribution: Getting RID of unstable, single model-based variable importance. In Neural Information Processing Systems (NeurIPS), 2023.
  16. 16.D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al. Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research, 2020.
  17. 17.FICO, Google, Imperial College London, MIT, University of Oxford, UC Irvine, and UC Berkeley. Explainable Machine Learning Challenge. https://community.fico.com/s/explainable-machine-learning-challenge, 2018.
  18. 18.Fisher, A., Rudin, C., and Dominici, F. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177):1–81, 2019.
  19. 19.Freund, Y. and Schapire, R. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119–139, August 1997.
  20. 20.Han, T., Srinivas, S., and Lakkaraju, H. Which explanation should I choose? a function approximation perspective to characterizing post hoc explanations. In Neural Information Processing Systems (NeurIPS), volume 35, pp. 5256–5268, 2022.
  21. 21.Holte, R. C. Very simple classification rules perform well on most commonly used datasets. Machine Learning, 11: 63–90, 1993.
  22. 22.Hsu, H. and Calmon, F. Rashomon capacity: A metric for predictive multiplicity in classification. In Neural Information Processing Systems (NeurIPS), volume 35, pp. 28988–29000, 2022.
  23. 23.Kleinberg, J. Inherent trade-offs in algorithmic fairness. SIGMETRICS Perform. Eval. Rev., 46(1):40, June 2018.
  24. 24.Kleinberg, J. and Mullainathan, S. Simplicity creates inequity: Implications for fairness, stereotypes, and interpretability. In Proceedings of the 2019 ACM Conference on Economics and Computation, EC ’19, pp. 807–808, 2019.
  25. 25.Kurosawa, A. Rashomon. RKO Radio Pictures, 1950.
  26. 26.Larson, J., Mattu, S., Kirchner, L., and Angwin, J. How we analyzed the COMPAS recidivism algorithm. ProPublica, 2016.
  27. 27.Lin, J., Zhong, C., Hu, D., Rudin, C., and Seltzer, M. Generalized and scalable optimal sparse decision trees. In Proceedings of the International Conference on Machine Learning (ICML), pp. 6150–6160, 2020.
  28. 28.Liu, J., Zhong, C., Li, B., Seltzer, M., and Rudin, C. Fasterrisk: Fast and accurate interpretable risk scores. In Neural Information Processing Systems (NeurIPS), 2022a.
  29. 29.Liu, J., Zhong, C., Seltzer, M., and Rudin, C. Fast sparse classification for generalized linear and additive models. In Proceedings of Artificial Intelligence and Statistics (AISTATS), 2022b.
  30. 30.Liu, J., Rosen, S., Zhong, C., and Rudin, C. OKRidge: Scalable optimal k-sparse ridge regression. In Neural Information Processing Systems (NeurIPS), 2023.
  31. 31.Lou, Y., Caruana, R., Gehrke, J., and Hooker, G. Accurate intelligible models with pairwise interactions. In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 623–631, 2013.
  32. 32.Marx, C., Calmon, F., and Ustun, B. Predictive multiplicity in classification. In Proceedings of the International Conference on Machine Learning (ICML), pp. 6765–6774, 2020.
  33. 33.Mata, K., Kanamori, K., and Arimura, H. Computing the collection of good models for rule lists. arXiv preprint arXiv:2204.11285, 2022.
  34. 34.McElfresh, D., Khandagale, S., Valverde, J., C, V. P., Feuer, B., Hegde, C., Ramakrishnan, G., Goldblum, M., and White, C. When do neural nets outperform boosted trees on tabular data? In Neural Information Processing Systems (NeurIPS), 2023.
  35. 35.McTavish, H., Zhong, C., Achermann, R., Karimalis, I., Chen, J., Rudin, C., and Seltzer, M. Fast sparse decision tree optimization via reference ensembles. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 9604–9613, 2022.
  36. 36.Novakovsky, G., Dexter, N., Libbrecht, M. W., Wasserman, W. W., and Mostafavi, S. Obtaining genetics insights from deep learning via explainable artificial intelligence. Nature Reviews Genetics, pp. 1–13, 2022.
  37. 37.Quinlan, J. R. C4.5: programs for machine learning, volume 1. Morgan Kaufmann, 1993.
  38. 38.Rodolfa, K. T., Lamba, H., and Ghani, R. Empirical observation of negligible fairness–accuracy trade-offs in machine learning for public policy. Nature Machine Intelligence, 3:896––904, October 2021.
  39. 39.Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1:206–215, May 2019.
  40. 40.Rudin, C. Intuition for the Algorithms of Machine Learning. self-published at https://users.cs.duke.edu/~cynthia/teaching.html, 2020.
  41. 41.Rudin, C. and Wagstaff, K. L. Machine learning for science and society. Machine Learning, 95(1), 2014.
  42. 42.Rudin, C., Wang, C., and Coker, B. The Age of secrecy and unfairness in recidivism prediction. Harvard Data Science Review, 2(1), 1 2020.
  43. 43.Rudin, C., Chen, C., Chen, Z., Huang, H., Semenova, L., and Zhong, C. Interpretable machine learning: Fundamental principles and 10 grand challenges. Statistics Surveys, 16:1–85, 2022.
  44. 44.Semenova, L., Rudin, C., and Parr, R. On the existence of simpler machine learning models. In ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT), 2022.
  45. 45.Semenova, L., Chen, H., Parr, R., and Rudin, C. A path to simpler models starts with noise. In Neural Information Processing Systems (NeurIPS), 2023.
  46. 46.Smith, G., Mansilla, R., and Goulding, J. Model class reliance for random forests. In Neural Information Processing Systems (NeurIPS), volume 33, pp. 22305–22315, 2020.
  47. 47.Sun, Y., Chen, Z., Orlandi, V., Wang, T., and Rudin, C. Sparse and faithful explanations without sparse models. In Proc. Artificial Intelligence and Statistics (AISTATS), 2024.
  48. 48.Tollenaar, N. and van der Heijden, P. Which method predicts recidivism best?: A comparison of statistical, machine learning and data mining predictive models. Journal of the Royal Statistical Society: Series A (Statistics in Society), 176(2):565–584, 2013.
  49. 49.Wagstaff, K. L. Machine learning that matters. In Proceedings of the International Conference on Machine Learning (ICML), pp. 1851–1856, 2012.
  50. 50.Wang, C., Han, B., Patel, B., and Rudin, C. In Pursuit of Interpretable, Fair and Accurate Machine Learning for Criminal Recidivism Prediction. Journal of Quantitative Criminology, pp. 1–63, 2022a.
  51. 51.Wang, F., Huang, S., Gao, R., Zhou, Y., Lai, C., Li, Z., Xian, W., Qian, X., Li, Z., Huang, Y., et al. Initial whole-genome sequencing and analysis of the host genetic contribution to COVID-19 severity and susceptibility. Cell Discovery, 6(1):83, 2020.
  52. 52.Wang, T., Rudin, C., Doshi-Velez, F., Liu, Y., Klampfl, E., and MacNeille, P. A Bayesian framework for learning rule sets for interpretable classification. Journal of Machine Learning Research, 18(70):1–37, 2017.
  53. 53.Wang, Z. J., Kale, A., Nori, H., Stella, P., Nunnally, M., Chau, D. H., Vorvoreanu, M., Vaughan, J. W., and Caruana, R. Gam changer: Editing generalized additive models with interactive visualization. Advances in Neural Information Processing Systems, Bridging the Gap: From Machine Learning Research to Clinical Practice (Research2Clinics) Workshop, 2021.
  54. 54.Wang, Z. J., Zhong, C., Xin, R., Takagi, T., Chen, Z., Chau, D. H., Rudin, C., and Seltzer, M. Timbertrek: Exploring and curating sparse decision trees with interactive visualization. In 2022 IEEE Visualization and Visual Analytics (VIS), pp. 60–64. IEEE, 2022b.
  55. 55.Watson-Daniels, J., Parkes, D. C., and Ustun, B. Predictive multiplicity in probabilistic classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 10306–10314, 2023.
  56. 56.Wikipedia. Right to explanation. https://en.wikipedia.org/wiki/Right_to_explanation, 2024.
  57. 57.Xin, R., Zhong, C., Chen, Z., Takagi, T., Seltzer, M., and Rudin, C. Exploring the whole Rashomon set of sparse decision trees. In Neural Information Processing Systems (NeurIPS), volume 35, pp. 14071–14084, 2022.
  58. 58.Yanagawa, M. and Sato, J. Seeing is not always believing: Discrepancies in saliency maps. Radiology: Artificial Intelligence, 6(1):e230488, 2024.
  59. 59.Zeng, J., Ustun, B., and Rudin, C. Interpretable classification models for recidivism prediction. Journal of the Royal Statistical Society: Series A (Statistics in Society), 180(3):689–722, 2017.
  60. 60.Zhong, C., Chen, Z., Liu, J., Seltzer, M., and Rudin, C. Exploring and interacting with the set of good sparse generalized additive models. In Neural Information Processing Systems (NeurIPS), 2023.
  61. 61.Zhu, C. Q., Tian, M., Semenova, L., Liu, J., Xu, J., Scarpa, J., and Rudin, C. Fast and interpretable mortality risk scores for critical care patients. arXiv preprint arXiv:2311.13015, 2023.

Citation

MLA
Rudin, C., et al. “Amazing Things Come From Having Many Good Models”. ICML (spotlight), 2024, 2024, http://arxiv.org/abs/2407.04846v2.
APA
Rudin, C., Zhong, C., Semenova, L., Seltzer, M., Parr, R., Liu, J., Katta, S., Donnelly, J., Chen, H., & Boner, Z. (2024). Amazing Things Come From Having Many Good Models. ICML (spotlight), 2024. http://arxiv.org/abs/2407.04846v2
Chicago
Rudin, C., C. Zhong, L. Semenova, et al. 2024. “Amazing Things Come From Having Many Good Models”. ICML (spotlight), 2024. http://arxiv.org/abs/2407.04846v2.
Harvard
Rudin, C. et al. (2024) “Amazing Things Come From Having Many Good Models”, ICML (spotlight), 2024 [Preprint]. Available at: http://arxiv.org/abs/2407.04846v2.
Vancouver
1. Rudin C, Zhong C, Semenova L, Seltzer M, Parr R, Liu J, Katta S, Donnelly J, Chen H, Boner Z (2024) Amazing Things Come From Having Many Good Models. ICML (spotlight), 2024

BibTeX

@article{rudin2024amazing,
  title = {Amazing Things Come From Having Many Good Models},
  author = {Rudin, Cynthia and Zhong, Chudi and Semenova, Lesia and Seltzer, Margo and Parr, Ronald and Liu, Jiachang and Katta, Srikar and Donnelly, Jon and Chen, Harry and Boner, Zachery},
  year = {2024},
  journal = {ICML (spotlight), 2024},
  url = {http://arxiv.org/abs/2407.04846v2},
  eprint = {2407.04846}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/