OpenXAI: Towards a Transparent Evaluation of Model Explanations
Chirag AgarwalSatyapriya KrishnaEshika SaxenaMartin PawelczykNari JohnsonIsha PuriMarinka ZitnikHimabindu Lakkaraju
Presents OpenXAI, an open-source framework and public leaderboard system that unifies twenty-two evaluation metrics across faithfulness, stability, and fairness to systematically benchmark post hoc feature attribution methods.
As machine learning models are deployed in high-stakes fields such as healthcare, finance, and the legal system, stakeholders increasingly rely on post hoc explanation methods to understand individual model decisions. These techniques identify which input features most strongly influenced a specific prediction. However, evaluating the reliability of these explanations has been difficult due to fragmented codebases, inconsistent testing procedures, and the lack of standardized benchmarks. Consequently, practitioners face critical uncertainty regarding which explanation tools to trust, under what conditions they remain reliable, and whether they operate fairly.
The article introduces OpenXAI, an open-source evaluation ecosystem and public leaderboard designed to systematically benchmark post hoc feature attribution methods. The primary objective is to provide a transparent, standardized, and reproducible framework to evaluate explanation techniques across three core dimensions of reliability: faithfulness to the underlying model, stability against input perturbations, and fairness across demographic subgroups.
To achieve this, the article establishes an end-to-end framework encompassing seven diverse real-world tabular datasets and a novel synthetic data generator with mathematically guaranteed ground truth explanations. The researchers evaluated six leading explanation methods—LIME, SHAP, Vanilla Gradients, Gradient x Input, SmoothGrad, and Integrated Gradients—alongside a random baseline across sixteen predictive models using twenty-two quantitative metrics.
The benchmarking reveals significant performance trade-offs among explanation methods. Gradient-based techniques like Vanilla Gradients, SmoothGrad, and Integrated Gradients achieved perfect scores in ranking feature importance against ground-truth benchmarks, but LIME outperformed them in capturing the correct positive or negative direction of feature contributions, achieving an advantage of over 60%. SmoothGrad demonstrated the highest overall predictive faithfulness and achieved 63.2% higher representation stability across real-world datasets, yet it exhibited noticeable fairness disparities between demographic groups. Conversely, Gradient x Input delivered the lowest subgroup fairness disparity—improving fairness scores by 8.9% over competing tools—despite underperforming on stability and faithfulness metrics in real-world scenarios.
These findings demonstrate that no single explanation method excels across all reliability criteria. Relying on an unvetted explanation technique introduces severe compliance, safety, and operational risks, as decision-makers might receive misleading feature attributions or perpetuate algorithmic bias. Organizations must treat model explainability as a multi-criteria optimization problem rather than assuming off-the-shelf explainers are universally dependable.
Decision-makers should avoid one-size-fits-all adoption of explanation methods and instead benchmark candidate explainers against use-case priorities using frameworks like OpenXAI. When regulatory fairness is paramount, techniques with minimal subgroup disparities should be favored, whereas applications prioritizing model auditing should select methods that maximize faithfulness. Organizations should integrate automated XAI testing pipelines into their machine learning operations before deploying models in high-risk environments.
The current benchmark has high confidence within its defined boundary conditions, though it is primarily limited to tabular datasets and two predictive model architectures. Readers should exercise caution when extrapolating these findings to other modalities, such as text or image data, until the benchmark expands in future iterations.
- Paper: Do Feature Attribution Methods Correctly Attribute Features?, Yilun Zhou et al. (2022). Its controlled ground-truth framework provides the conceptual model for OpenXAI’s central task of objectively testing whether feature attributions identify genuinely predictive features.
- Paper: Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations, Tessa Han et al. (2022). Its unified local-function-approximation account explains why methods such as LIME, SHAP, SmoothGrad, and Integrated Gradients exhibit the cross-metric trade-offs that OpenXAI benchmarks.
- Paper: A Unified Approach to Interpreting Model Predictions, Scott M. Lundberg et al. (2017). Its SHAP framework supplies the game-theoretic attribution method and consistency principles needed to understand one of OpenXAI’s principal baselines.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). Its axioms and Integrated Gradients method establish the attribution reliability concepts directly evaluated among OpenXAI’s benchmarked explainers.
- Paper: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Marco Tulio Ribeiro et al. (2016). Its introduction of LIME explains the model-agnostic local surrogate method whose faithfulness and stability OpenXAI measures.
- Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). Its parameter-randomization and label-randomization sanity checks motivate OpenXAI’s broader insistence that plausible-looking explanations require empirical reliability testing.
- Paper: Towards A Rigorous Science of Interpretable Machine Learning, Finale Doshi-Velez et al. (2017). Its taxonomy of interpretability evaluation clarifies why OpenXAI separates formal faithfulness, stability, and fairness tests rather than treating explanation quality as a single property.
- Paper: Faithfulness Tests for Natural Language Explanations, Pepa Atanasova et al. (2023). It extends OpenXAI’s faithfulness agenda from feature attributions to natural-language explanations through counterfactual and input-reconstruction tests.
- Paper: A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation, Thomas Fel et al. (2023). It continues OpenXAI’s evaluation program into concept-based explanations by separating concept extraction from concept-level attribution and benchmarking both.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). It applies the benchmark’s fairness-auditing logic to discover previously unspecified biases in text-to-image systems rather than only measuring predefined demographic disparities.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It sharpens OpenXAI’s warning about explanation unreliability by demonstrating systematically unfaithful chain-of-thought rationales in language models.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It critically generalizes OpenXAI’s reliability concerns into a statistical-causal account of instability and non-identifiability across modern interpretability methods.
