OpenXAI: Towards a Transparent Evaluation of Model Explanations

Chirag AgarwalSatyapriya KrishnaEshika SaxenaMartin PawelczykNari JohnsonIsha PuriMarinka ZitnikHimabindu Lakkaraju

article2022NeurIPS190 citations

Presents OpenXAI, an open-source framework and public leaderboard system that unifies twenty-two evaluation metrics across faithfulness, stability, and fairness to systematically benchmark post hoc feature attribution methods.

Listen

As machine learning models are deployed in high-stakes fields such as healthcare, finance, and the legal system, stakeholders increasingly rely on post hoc explanation methods to understand individual model decisions. These techniques identify which input features most strongly influenced a specific prediction. However, evaluating the reliability of these explanations has been difficult due to fragmented codebases, inconsistent testing procedures, and the lack of standardized benchmarks. Consequently, practitioners face critical uncertainty regarding which explanation tools to trust, under what conditions they remain reliable, and whether they operate fairly.

The article introduces OpenXAI, an open-source evaluation ecosystem and public leaderboard designed to systematically benchmark post hoc feature attribution methods. The primary objective is to provide a transparent, standardized, and reproducible framework to evaluate explanation techniques across three core dimensions of reliability: faithfulness to the underlying model, stability against input perturbations, and fairness across demographic subgroups.

To achieve this, the article establishes an end-to-end framework encompassing seven diverse real-world tabular datasets and a novel synthetic data generator with mathematically guaranteed ground truth explanations. The researchers evaluated six leading explanation methods—LIME, SHAP, Vanilla Gradients, Gradient x Input, SmoothGrad, and Integrated Gradients—alongside a random baseline across sixteen predictive models using twenty-two quantitative metrics.

The benchmarking reveals significant performance trade-offs among explanation methods. Gradient-based techniques like Vanilla Gradients, SmoothGrad, and Integrated Gradients achieved perfect scores in ranking feature importance against ground-truth benchmarks, but LIME outperformed them in capturing the correct positive or negative direction of feature contributions, achieving an advantage of over 60%. SmoothGrad demonstrated the highest overall predictive faithfulness and achieved 63.2% higher representation stability across real-world datasets, yet it exhibited noticeable fairness disparities between demographic groups. Conversely, Gradient x Input delivered the lowest subgroup fairness disparity—improving fairness scores by 8.9% over competing tools—despite underperforming on stability and faithfulness metrics in real-world scenarios.

These findings demonstrate that no single explanation method excels across all reliability criteria. Relying on an unvetted explanation technique introduces severe compliance, safety, and operational risks, as decision-makers might receive misleading feature attributions or perpetuate algorithmic bias. Organizations must treat model explainability as a multi-criteria optimization problem rather than assuming off-the-shelf explainers are universally dependable.

Decision-makers should avoid one-size-fits-all adoption of explanation methods and instead benchmark candidate explainers against use-case priorities using frameworks like OpenXAI. When regulatory fairness is paramount, techniques with minimal subgroup disparities should be favored, whereas applications prioritizing model auditing should select methods that maximize faithfulness. Organizations should integrate automated XAI testing pipelines into their machine learning operations before deploying models in high-risk environments.

The current benchmark has high confidence within its defined boundary conditions, though it is primarily limited to tabular datasets and two predictive model architectures. Readers should exercise caution when extrapolating these findings to other modalities, such as text or image data, until the benchmark expands in future iterations.

arXiv: 2206.11104
Cover for OpenXAI: Towards a Transparent Evaluation of Model Explanations

Abstract

While several types of post hoc explanation methods have been proposed in recent literature, there is very little work on systematically benchmarking these methods. Here, we introduce OpenXAI, a comprehensive and extensible open-source framework for evaluating and benchmarking post hoc explanation methods. OpenXAI comprises of the following key components: (i) a flexible synthetic data generator and a collection of diverse real-world datasets, pre-trained models, and state-of-the-art feature attribution methods, (ii) open-source implementations of twenty-two quantitative metrics for evaluating faithfulness, stability (robustness), and fairness of explanation methods, and (iii) the first ever public XAI leaderboards to readily compare several explanation methods across a wide variety of metrics, models, and datasets. OpenXAI is easily extensible, as users can readily evaluate custom explanation methods and incorporate them into our leaderboards.

Overall, OpenXAI provides an automated end-to-end pipeline that not only simplifies and standardizes the evaluation of post hoc explanation methods, but also promotes transparency and reproducibility in benchmarking these methods. While the first release of OpenXAI supports only tabular datasets, the explanation methods and metrics that we consider are general enough to be applicable to other data modalities. OpenXAI datasets and data loaders, implementations of state-of-the-art explanation methods and evaluation metrics, as well as leaderboards are publicly available at https://open-xai.github.io/. OpenXAI will be regularly updated to incorporate text and image datasets, other new metrics and explanation methods, and welcomes inputs from the community.

Table of Contents

  • 1 Introduction
  • 2 Overview of OpenXAI Framework
  • 3 Benchmarking Analysis
  • 4 Conclusions
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — OpenXAI Post Hoc Explanation Benchmarking Framework

    model/method

    OpenXAI is an open-source benchmarking ecosystem designed to standardize, evaluate, and compare post hoc feature attribution explanation methods. The framework provides an end-to-end pipeline structured around four key components:

    1. Data and Model Management: Built-in Dataloader and LoadModel interfaces that load pre-processed tabular datasets with standardized train/test splits (70% training, 30% testing) and pre-trained predictive models (logistic regression and deep neural networks).
    2. Explainers: Standardized interfaces implementing six local feature attribution methods (LIME, SHAP, Vanilla Gradients, Gradient ×\times Input, SmoothGrad, and Integrated Gradients) along with a random attribution baseline.
    3. Evaluator: API implementations of twenty-two quantitative evaluation metrics spanning three core dimensions of explanation reliability: faithfulness (ground-truth agreement and prediction perturbation gaps), stability (robustness to input, parameter, and output shifts), and fairness (subgroup performance disparities).
    4. Public Leaderboards: Interactive web-based leaderboards allowing comparative evaluation and ranking of explanation methods across diverse datasets, models, and evaluation metrics.
  2. Knowl 2 — SynthGauss Synthetic Data Generation Mechanism

    algorithm

    SynthGauss is a synthetic data generation algorithm designed to create datasets with mathematically verifiable ground-truth feature attribution explanations. Prior synthetic data generators suffer from feature correlations that allow predictive models to achieve high accuracy using features different from the data-generating ground truth. SynthGauss eliminates this issue by enforcing three properties: feature independence, unambiguously separated local clusters (neighborhoods), and explicitly defined cluster-specific feature influences.

    Input: Number of clusters KK, feature dimensionality dd, number of samples NN, cluster centers {μk\mu_k}k=1K_{k=1}^K
    Output: Synthetic dataset D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N, ground truth explanation masks {mk\mathbf{m}_k}k=1K_{k=1}^K
    for each cluster k∈{1,…,K}k \in \{1, \dots, K\}:
        Set covariance matrix Σk=Id\Sigma_k = I_d (ensures feature independence)
        Ensure intracluster distance ≪\ll intercluster distance ∥μk−μj∥\|\mu_k - \mu_j\| for all j≠kj \neq k
        Sample binary feature mask vector mk∼{0,1}d\mathbf{m}_k \sim \{0, 1\}^d
        Sample feature weight vector wk∼Rd\mathbf{w}_k \sim \mathbb{R}^d
    for each sample i∈{1,…,N}i \in \{1, \dots, N\}:
        Sample cluster assignment k∈{1,…,K}k \in \{1, \dots, K\}
        Sample feature vector xi∼N(μk,Σk)x_i \sim \mathcal{N}(\mu_k, \Sigma_k)
        Compute ground truth label probability pi=σ(xi⊤(mk⊙wk))p_i = \sigma(x_i^\top (\mathbf{m}_k \odot \mathbf{w}_k)), where ⊙\odot is the Hadamard product and σ\sigma is the sigmoid function
        Assign label yi∼Bernoulli(pi)y_i \sim \text{Bernoulli}(p_i)
        Assign ground-truth explanation for xix_i as Egt(xi)=mkE_{gt}(x_i) = \mathbf{m}_k
    return D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N, {mk\mathbf{m}_k}k=1K_{k=1}^K

    Because the cluster covariances are identity matrices and clusters are well separated, any accurate classifier trained on D\mathcal{D} is theoretically guaranteed to rely on the active features indicated by mk\mathbf{m}_k, ensuring that post hoc explanations can be validly compared against mk\mathbf{m}_k.

  3. Knowl 3 — Ground-Truth Faithfulness Evaluation Metrics for Feature Attributions

    definition

    When ground-truth feature attributions Egt∈RdE_{gt} \in \mathbb{R}^d are available (e.g., via synthetic generation), the faithfulness of a post hoc explanation E∈RdE \in \mathbb{R}^d is evaluated using six similarity metrics focused on the top-KK most important features topK(E)\text{top}_K(E) and their rankings:

    • Feature Agreement (FA): The fraction of top-KK features identified by the explanation that match the ground truth: FA(E,Egt,K)=∣topK(E)∩topK(Egt)∣K\text{FA}(E, E_{gt}, K) = \frac{|\text{top}_K(E) \cap \text{top}_K(E_{gt})|}{K}

    • Rank Agreement (RA): The fraction of top-KK features that not only appear in both top-KK sets but also share the exact same ranking position: RA(E,Egt,K)=1K∑f∈topK(E)∩topK(Egt)I(rank(E,f)=rank(Egt,f))\text{RA}(E, E_{gt}, K) = \frac{1}{K} \sum_{f \in \text{top}_K(E) \cap \text{top}_K(E_{gt})} \mathbb{I}(\text{rank}(E, f) = \text{rank}(E_{gt}, f))

    • Sign Agreement (SA): The fraction of top-KK features common to both explanations that share the same sign (direction of contribution): SA(E,Egt,K)=1K∑f∈topK(E)∩topK(Egt)I(sign(Ef)=sign((Egt)f))\text{SA}(E, E_{gt}, K) = \frac{1}{K} \sum_{f \in \text{top}_K(E) \cap \text{top}_K(E_{gt})} \mathbb{I}(\text{sign}(E_f) = \text{sign}((E_{gt})_f))

    • Signed Rank Agreement (SRA): The fraction of top-KK features that match in both rank position and sign: SRA(E,Egt,K)=1K∑f∈topK(E)∩topK(Egt)I(rank(E,f)=rank(Egt,f)  ∧  sign(Ef)=sign((Egt)f))\text{SRA}(E, E_{gt}, K) = \frac{1}{K} \sum_{f \in \text{top}_K(E) \cap \text{top}_K(E_{gt})} \mathbb{I}(\text{rank}(E, f) = \text{rank}(E_{gt}, f) \;\wedge\; \text{sign}(E_f) = \text{sign}((E_{gt})_f))

    • Rank Correlation (RC): Spearman's rank correlation coefficient ρ\rho computed between the complete feature ranking vectors induced by EE and EgtE_{gt}.

    • Pairwise Rank Agreement (PRA): The fraction of all distinct feature pairs (i,j)(i, j) whose relative attribution ordering in EE agrees with their relative ordering in EgtE_{gt}: PRA(E,Egt)=2d(d−1)∑i<jI((Ei−Ej)⋅((Egt)i−(Egt)j)>0)\text{PRA}(E, E_{gt}) = \frac{2}{d(d-1)} \sum_{i < j} \mathbb{I}((E_i - E_j) \cdot ((E_{gt})_i - (E_{gt})_j) > 0)

  4. Knowl 4 — Predictive Faithfulness Perturbation Metrics

    definition

    When ground-truth feature importances are unavailable, explanation faithfulness is evaluated by perturbing input features according to their attributed importance and measuring the resulting shift in the predictive model's output probability f(x)f(x):

    • Prediction Gap on Important feature perturbation (PGI): Quantifies the drop in model output probability when the top features deemed most influential by explanation E(x)E(x) are perturbed. Let xpert, topx_{\text{pert, top}} denote xx with its top-KK attributed features perturbed: PGI(x,E,f)=∣f(x)−f(xpert, top)∣\text{PGI}(x, E, f) = |f(x) - f(x_{\text{pert, top}})| Higher PGI values reflect greater faithfulness, as altering critical features should strongly impact model predictions.

    • Prediction Gap on Unimportant feature perturbation (PGU): Quantifies the change in model output probability when features deemed least influential (unimportant) by explanation E(x)E(x) are perturbed. Let xpert, botx_{\text{pert, bot}} denote xx with its least important features perturbed: PGU(x,E,f)=∣f(x)−f(xpert, bot)∣\text{PGU}(x, E, f) = |f(x) - f(x_{\text{pert, bot}})| Lower PGU values reflect greater faithfulness, as perturbing non-influential features should leave model predictions largely invariant.

  5. Knowl 5 — Relative Stability Metrics for Post Hoc Explanations

    definition

    Explanation stability measures the degree to which an explanation E(x)E(x) remains invariant under small input perturbations x′x', formalizing robustness across input, representation, and output dimensions:

    • Relative Input Stability (RIS): Measures the maximum change in the explanation vector relative to the perturbation distance in input space: RIS(x,x′)=max⁡x′∈Bϵ(x)∥E(x)−E(x′)∥2∥x−x′∥2\text{RIS}(x, x') = \max_{x' \in \mathcal{B}_\epsilon(x)} \frac{\|E(x) - E(x')\|_2}{\|x - x'\|_2} where Bϵ(x)\mathcal{B}_\epsilon(x) is a local neighborhood ball around xx.

    • Relative Representation Stability (RRS): Measures the maximum shift in explanation relative to the distance between internal layer representations h(x)h(x) and h(x′)h(x') of the predictive model: RRS(x,x′)=max⁡x′∈Bϵ(x)∥E(x)−E(x′)∥2∥h(x)−h(x′)∥2\text{RRS}(x, x') = \max_{x' \in \mathcal{B}_\epsilon(x)} \frac{\|E(x) - E(x')\|_2}{\|h(x) - h(x')\|_2}

    • Relative Output Stability (ROS): Measures the maximum shift in explanation relative to the change in model output prediction probabilities f(x)f(x) and f(x′)f(x'): ROS(x,x′)=max⁡x′∈Bϵ(x)∥E(x)−E(x′)∥2∥f(x)−f(x′)∥2\text{ROS}(x, x') = \max_{x' \in \mathcal{B}_\epsilon(x)} \frac{\|E(x) - E(x')\|_2}{\|f(x) - f(x')\|_2}

    For all stability metrics, values closer to zero indicate superior explanation stability and robustness against local noise.

  6. Knowl 6 — Subgroup Disparity Metrics for Explanation Fairness

    definition

    Fairness of a post hoc explanation method is measured by assessing disparities in explanation quality across demographic subgroups (e.g., majority subgroup SmajS_{\text{maj}} defined by male individuals and minority subgroup SminS_{\text{min}} defined by female individuals). For any quantitative explanation metric M(x)M(x) (such as PGI, PGU, RIS, or RRS), the group unfairness ΔM\Delta_M is computed as the absolute difference between subgroup average metric values:

    ΔM=∣1∣Smaj∣∑i∈SmajM(xi)−1∣Smin∣∑j∈SminM(xj)∣\Delta_M = \left| \frac{1}{|S_{\text{maj}}|} \sum_{i \in S_{\text{maj}}} M(x_i) - \frac{1}{|S_{\text{min}}|} \sum_{j \in S_{\text{min}}} M(x_j) \right|

    A smaller gap ΔM\Delta_M indicates greater fairness and uniformity in explanation performance across protected demographic groups, whereas large discrepancies demonstrate systemic explanation unfairness.

  7. Knowl 7 — OpenXAI Benchmark Datasets and Model Architecture Setup

    experimental setup

    OpenXAI benchmarks post hoc explainers across eight tabular datasets spanning diverse domains, sample sizes, and feature structures:

    1. Synthetic Data: 5,000 instances, 20 continuous features, balanced labels, generated via SynthGauss.
    2. German Credit: 1,000 instances, 20 discrete/continuous demographic and financial features, imbalanced.
    3. HELOC (Home Equity Line of Credit): 9,871 instances, 23 continuous financial features, balanced.
    4. COMPAS: 18,876 instances, 7 discrete/continuous demographic and criminal history features, imbalanced.
    5. Adult Income: 48,842 instances, 13 discrete/continuous census features, imbalanced.
    6. Give Me Some Credit (GMC): 102,209 instances, 10 discrete/continuous financial features, imbalanced.
    7. Pima-Indians Diabetes: 768 instances, 9 discrete/continuous medical features, imbalanced.
    8. Framingham Heart Study: 4,240 instances, 16 continuous clinical features, imbalanced.

    Datasets are partitioned into 70% training and 30% testing splits. Two classes of predictive models are trained per dataset:

    • Logistic Regression (LR): Linear baseline model.
    • Artificial Neural Network (ANN): Multi-layer perceptron comprising two fully connected hidden layers with 100 neurons each, ReLU non-linearities, and a softmax classification output layer.
  8. Knowl 8 — Faithfulness Benchmarking Results Across Post Hoc Attribution Methods

    data/table

    Evaluation of post hoc explanation methods on Logistic Regression (LR) models evaluated on the HELOC and Adult Income test sets across ground-truth faithfulness metrics (Pairwise Rank Agreement PRA, Rank Correlation RC, Feature Agreement FA, Rank Agreement RA, Sign Agreement SA, Signed Rank Agreement SRA) and predictive faithfulness metrics (Prediction Gap on Unimportant PGU, Prediction Gap on Important PGI):

    HELOC Dataset PRA (↑\uparrow) RC (↑\uparrow) FA (↑\uparrow) RA (↑\uparrow) SA (↑\uparrow) SRA (↑\uparrow) PGU (↓\downarrow) PGI (↑\uparrow)
    Random 0.500±0.000.500 \pm 0.00 0.005±0.010.005 \pm 0.01 0.498±0.000.498 \pm 0.00 0.043±0.000.043 \pm 0.00 0.251±0.000.251 \pm 0.00 0.022±0.000.022 \pm 0.00 0.033±0.000.033 \pm 0.00 0.035±0.000.035 \pm 0.00
    VanillaGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.469±0.010.469 \pm 0.01 0.469±0.010.469 \pm 0.01 0.034±0.000.034 \pm 0.00 0.036±0.000.036 \pm 0.00
    IntegratedGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.469±0.010.469 \pm 0.01 0.469±0.010.469 \pm 0.01 0.034±0.000.034 \pm 0.00 0.036±0.000.036 \pm 0.00
    Gradient x Input 0.641±0.000.641 \pm 0.00 0.390±0.010.390 \pm 0.01 0.582±0.000.582 \pm 0.00 0.049±0.000.049 \pm 0.00 0.255±0.010.255 \pm 0.01 0.022±0.000.022 \pm 0.00 0.033±0.000.033 \pm 0.00 0.036±0.000.036 \pm 0.00
    SmoothGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.957±0.00\mathbf{0.957 \pm 0.00} 0.274±0.000.274 \pm 0.00 0.274±0.000.274 \pm 0.00 0.015±0.00\mathbf{0.015 \pm 0.00} 0.047±0.00\mathbf{0.047 \pm 0.00}
    SHAP 0.645±0.000.645 \pm 0.00 0.384±0.010.384 \pm 0.01 0.586±0.000.586 \pm 0.00 0.054±0.000.054 \pm 0.00 0.269±0.010.269 \pm 0.01 0.024±0.000.024 \pm 0.00 0.033±0.000.033 \pm 0.00 0.036±0.000.036 \pm 0.00
    LIME 0.982±0.000.982 \pm 0.00 0.994±0.000.994 \pm 0.00 0.932±0.000.932 \pm 0.00 0.671±0.000.671 \pm 0.00 0.929±0.00\mathbf{0.929 \pm 0.00} 0.670±0.00\mathbf{0.670 \pm 0.00} 0.042±0.000.042 \pm 0.00 0.029±0.0010.029 \pm 0.001
    Adult Dataset
    Random 0.499±0.000.499 \pm 0.00 0.0±0.000.0 \pm 0.00 0.496±0.000.496 \pm 0.00 0.068±0.000.068 \pm 0.00 0.250±0.000.250 \pm 0.00 0.037±0.000.037 \pm 0.00 0.053±0.000.053 \pm 0.00 0.06±0.000.06 \pm 0.00
    VanillaGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.923±0.00\mathbf{0.923 \pm 0.00} 0.921±0.00\mathbf{0.921 \pm 0.00} 0.138±0.000.138 \pm 0.00 0.136±0.000.136 \pm 0.00 0.07±0.0010.07 \pm 0.001 0.039±0.0010.039 \pm 0.001
    IntegratedGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.923±0.00\mathbf{0.923 \pm 0.00} 0.923±0.00\mathbf{0.923 \pm 0.00} 0.138±0.000.138 \pm 0.00 0.138±0.000.138 \pm 0.00 0.07±0.0010.07 \pm 0.001 0.039±0.0010.039 \pm 0.001
    Gradient x Input 0.580±0.000.580 \pm 0.00 0.281±0.000.281 \pm 0.00 0.567±0.000.567 \pm 0.00 0.075±0.000.075 \pm 0.00 0.070±0.000.070 \pm 0.00 0.003±0.000.003 \pm 0.00 0.043±0.000.043 \pm 0.00 0.073±0.000.073 \pm 0.00
    SmoothGrad 1.0±0.00\mathbf{1.0 \pm 0.00} 1.0±0.00\mathbf{1.0 \pm 0.00} 0.923±0.00\mathbf{0.923 \pm 0.00} 0.923±0.00\mathbf{0.923 \pm 0.00} 0.741±0.000.741 \pm 0.00 0.741±0.00\mathbf{0.741 \pm 0.00} 0.008±0.00\mathbf{0.008 \pm 0.00} 0.099±0.001\mathbf{0.099 \pm 0.001}
    SHAP 0.655±0.000.655 \pm 0.00 0.379±0.000.379 \pm 0.00 0.601±0.000.601 \pm 0.00 0.105±0.000.105 \pm 0.00 0.133±0.000.133 \pm 0.00 0.009±0.000.009 \pm 0.00 0.047±0.000.047 \pm 0.00 0.068±0.000.068 \pm 0.00
    LIME 0.913±0.000.913 \pm 0.00 0.921±0.000.921 \pm 0.00 0.869±0.000.869 \pm 0.00 0.697±0.000.697 \pm 0.00 0.858±0.00\mathbf{0.858 \pm 0.00} 0.689±0.000.689 \pm 0.00 0.014±0.000.014 \pm 0.00 0.094±0.0010.094 \pm 0.001

    Gradient-based methods (VanillaGrad, IntegratedGrad, SmoothGrad) achieve near-perfect rank alignment on linear models because input gradients directly match the model weights. However, LIME achieves superior sign fidelity (SA and SRA), outperforming others by +61.6%+61.6\% on SA and +65.3%+65.3\% on SRA across datasets. In predictive faithfulness, SmoothGrad achieves the best PGU (+43.03%+43.03\% lower perturbation gap on unimportant features).

  9. Knowl 9 — Empirical Stability Comparison of Feature Attribution Methods

    data/table

    Evaluation of explanation stability metrics (Relative Input Stability RIS and Relative Representation Stability RRS; values closer to zero indicate greater stability) on Logistic Regression (LR) models across Synthetic and German Credit datasets:

    Synthetic Dataset German Credit Dataset
    Method RIS RRS RIS RRS
    Random 6.868±0.0136.868 \pm 0.013 6.687±0.0156.687 \pm 0.015 6.274±0.1046.274 \pm 0.104 16.448±0.12416.448 \pm 0.124
    Vanilla Gradients 6.133±0.0116.133 \pm 0.011 6.144±0.0066.144 \pm 0.006 −1.384±0.112-1.384 \pm 0.112 5.241±0.0245.241 \pm 0.024
    Integrated Gradients 5.957±0.0135.957 \pm 0.013 9.022±0.0439.022 \pm 0.043 −2.004±0.119-2.004 \pm 0.119 4.560±0.029\mathbf{4.560 \pm 0.029}
    Gradient x Input 0.405±0.015\mathbf{0.405 \pm 0.015} 3.422±0.037\mathbf{3.422 \pm 0.037} −0.906±0.104-0.906 \pm 0.104 9.437±0.1249.437 \pm 0.124
    SmoothGrad 5.249±0.0085.249 \pm 0.008 9.419±0.0379.419 \pm 0.037 −4.780±0.117-4.780 \pm 0.117 4.931±0.1224.931 \pm 0.122
    SHAP 5.673±0.0125.673 \pm 0.012 8.751±0.0358.751 \pm 0.035 −0.230±0.109\mathbf{-0.230 \pm 0.109} 10.056±0.11510.056 \pm 0.115
    LIME 9.355±0.0089.355 \pm 0.008 13.564±0.03613.564 \pm 0.036 −0.698±0.109-0.698 \pm 0.109 9.397±0.1199.397 \pm 0.119

    Relative stability varies significantly across datasets and architectures:

    • On synthetic data, Gradient ×\times Input achieves the best input stability (RIS of 0.4050.405, a +93.5%+93.5\% improvement) and representation stability (RRS of 3.4223.422, a +59.2%+59.2\% improvement).
    • On real-world datasets such as German Credit, Gradient ×\times Input stability degrades substantially. Across real-world datasets, SmoothGrad achieves on average 63.2%63.2\% higher RRS compared to alternatives, but no single explanation method maintains consistently superior stability across all datasets and models on RIS.
  10. Knowl 10 — Trade-offs Between Explanation Faithfulness, Stability, and Fairness

    empirical result

    Benchmarking post hoc explanation methods across demographic subgroups (majority: male, minority: female) using the Prediction Gap on Unimportant feature perturbation (PGU) on German Credit and Adult Income datasets demonstrates systemic metric trade-offs:

    1. Faithfulness Disparities: Explanations generated by Vanilla Gradients, Integrated Gradients, and SmoothGrad exhibit notable disparities in faithfulness between male and female subgroups, showing that explanations are not equally reliable across demographic groups.
    2. Fairness Advantages: Gradient ×\times Input produces the lowest subgroup disparity across both datasets, outperforming other methods by +8.9%+8.9\% on group fairness metrics.
    3. Multi-Metric Trade-Off: Although Gradient ×\times Input exhibits the highest group fairness (lowest quality disparity between subgroups), it underperforms on predictive faithfulness and stability compared to methods like SmoothGrad and Integrated Gradients. This confirms that no single post hoc explanation method optimizes faithfulness, stability, and fairness simultaneously, requiring practitioners to navigate explicit trade-offs based on application requirements.
  11. Knowl 11 — Scope Limitations of OpenXAI Release

    limitation

    The initial release of the OpenXAI framework contains specific scope boundaries:

    1. Data Modality: The initial release is restricted exclusively to tabular datasets, leaving unstructured image and text benchmark datasets to future updates despite the theoretical applicability of the metrics and explainers to other modalities.
    2. Catalog Scope: The initial benchmark evaluates seven feature attribution methods across eight datasets and two model families (Logistic Regression and 2-layer MLPs), requiring user contributions and ongoing releases to incorporate emerging explanation algorithms and architectures.

Coverage note — Omitted specific appendix-only hyperparameter tuning ranges and individual mathematical proofs from the supplemental materials to keep the extraction focused on the primary framework components, synthetic generator algorithm, metric formulations, benchmark data, and empirical findings.

References

  1. 1.Framingham heart study dataset | kaggle. https://www.kaggle.com/datasets/ aasheesh200/framingham-heart-study-dataset. (Accessed on 08/15/2022).
  2. 2.Shap benchmark. URL https://shap.readthedocs.io/en/latest/index.html.
  3. 3.Chirag Agarwal and Anh Nguyen. Explaining image classifiers by removing input features using generative models. In ACCV, 2020.
  4. 4.Chirag Agarwal, Nari Johnson, Martin Pawelczyk, Satyapriya Krishna, Eshika Saxena, Marinka Zitnik, and Himabindu Lakkaraju. Rethinking stability for attribution-based explanations. In ICLR 2022 Workshop on PAIR2Struct, 2022.
  5. 5.Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Steven Wu, and Himabindu Lakkaraju. Towards the unification and robustness of perturbation and gradient based explanations. In ICML, 2021.
  6. 6.Ulrich Aivodji, Hiromi Arai, Olivier Fortineau, Sébastien Gambs, Satoshi Hara, and Alain Tapp. Fairwashing: the risk of rationalization. In ICML, 2019.
  7. 7.David Alvarez-Melis and Tommi S Jaakkola. On the robustness of interpretability methods. arXiv, 2018.
  8. 8.Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al. Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion, 2020.
  9. 9.Aparna Balagopalan, Haoran Zhang, Kimia Hamidieh, Thomas Hartvigsen, Frank Rudzicz, and Marzyeh Ghassemi. The road to explainability is paved with bias: Measuring the fairness of explanations. arXiv, 2022.
  10. 10.Naman Bansal, Chirag Agarwal, and Anh Nguyen. Sam: The sensitivity of attribution methods to hyperparameters. In CVPR, 2020.
  11. 11.Solon Barocas, Andrew Selbst, and Manish Raghavan. The hidden assumptions behind counterfactual explanations and principal reasons. In FAccT, 2020.
  12. 12.Osbert Bastani, Carolyn Kim, and Hamsa Bastani. Interpretability via model extraction. arXiv, 2017.
  13. 13.Jacob Bien and Robert Tibshirani. Classification by set cover: The prototype vector machine. arXiv, 2009.
  14. 14.Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci. Deep neural networks and tabular data: A survey. arXiv, 2021.
  15. 15.Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In KDD, 2015.
  16. 16.Valerie Chen, Nari Johnson, Nicholay Topin, Gregory Plumb, and Ameet Talwalkar. Use-case-grounded simulations for explanation evaluation. arXiv, 2022.
  17. 17.Ian Covert, Scott Lundberg, and Su-In Lee. Explaining by removing: A unified framework for model explanation. JMLR, 2021.
  18. 18.Jessica Dai, Sohini Upadhyay, Ulrich Aivodji, Stephen H Bach, and Himabindu Lakkaraju. Fairness via explanation quality: Evaluating disparities in the quality of post hoc explanations. In AAAI Conference on AI, Ethics, and Society (AIES), 2022.
  19. 19.Sanjoy Dasgupta, Nave Frost, and Michal Moshkovitz. Framework for evaluating faithfulness of local explanations. arXiv, 2022.
  20. 20.Ricardo Dominguez-Olmedo, Amir H Karimi, and Bernhard Schölkopf. On the adversarial robustness of causal algorithmic recourse. In ICML. PMLR, 2022.
  21. 21.Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning. arXiv, 2017.
  22. 22.Dheeru Dua and Casey Graff. UCI machine learning repository, 2017. URL http://archive.ics.uci.edu/ml.
  23. 23.Radwa Elshawi, Mouaz H Al-Mallah, and Sherif Sakr. On the interpretability of machine learning-based model for predicting hypertension. BMC medical informatics and decision making, 2019.
  24. 24.Lukas Faber, Amin K. Moghaddam, and Roger Wattenhofer. When comparing to ground truth is wrong: On evaluating gnn explanation methods. In KDD, 2021.
  25. 25.FICO. Explainable machine learning challenge. https://community.fico.com/s/explainable-machine-learning-challenge?tabset-158d9=3, 2022. (Accessed on 05/23/2022).
  26. 26.Hidde Fokkema, Rianne de Heide, and Tim van Erven. Attribution-based explanations that provide recourse cannot be robust. arXiv, 2022.
  27. 27.Bryce Freshcorn. Give me some credit :: 2011 competition data | kaggle. https://www.kaggle.com/datasets/brycecf/give-me-some-credit-dataset, 2022. (Accessed on 05/23/2022).
  28. 28.Marzyeh Ghassemi, Luke Oakden-Rayner, and Andrew L Beam. The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 2021.
  29. 29.Amirata Ghorbani, Abubakar Abid, and James Zou. Interpretation of neural networks is fragile. In AAAI Conference on Artificial Intelligence, 2019.
  30. 30.Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. A survey of methods for explaining black box models. ACM computing surveys (CSUR), 2018.
  31. 31.Tessa Han, Suraj Srinivas, and Himabindu Lakkaraju. Which explanation should i choose? a function approximation perspective to characterizing post hoc explanations. arXiv, 2022.
  32. 32.Anna Hedström, Leander Weber, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin, and Marina M-C Höhne. Quantus: an explainable ai toolkit for responsible evaluation of neural network explanations. arXiv, 2022.
  33. 33.Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. Evaluating feature importance estimates. arXiv, 2018.
  34. 34.Mark Ibrahim, Melissa Louie, Ceena Modarres, and John Paisley. Global explanations of neural networks: Mapping the landscape of predictions. CoRR, abs/1902.02384, 2019.
  35. 35.Sérgio Jesus, Catarina Belém, Vladimir Balayan, João Bento, Pedro Saleiro, Pedro Bizarro, and João Gama. How can i choose an explainer? an application-grounded evaluation of post-hoc explanations. In FAccT, 2021.
  36. 36.Kareem L Jordan and Tina L Freiburger. The effect of race/ethnicity on sentencing: Examining sentence type, jail length, and prison length. In Journal of Ethnicity in Criminal Justice. Taylor & Francis, 2015.
  37. 37.Amir-Hossein Karimi, Gilles Barthe, Borja Balle, and Isabel Valera. Model-agnostic counterfactual explanations for consequential decisions. arXiv, 2019.
  38. 38.Amir-Hossein Karimi, Bernhard Schölkopf, and Isabel Valera. Algorithmic recourse: from counterfactual explanations to interventions. CoRR, abs/2002.06278, 2020.
  39. 39.Amir-Hossein Karimi, Julius von Kügelgen, Bernhard Schölkopf, and Isabel Valera. Algorithmic recourse under imperfect causal knowledge: a probabilistic approach. CoRR, 2020.
  40. 40.Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning. In CHI Conference on Human Factors in Computing Systems, 2020.
  41. 41.Joon Sik Kim, Gregory Plumb, and Ameet Talwalkar. Sanity simulations for saliency methods. arXiv, 2021.
  42. 42.Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. Captum: A unified and generic model interpretability library for pytorch, 2020.
  43. 43.Satyapriya Krishna, Tessa Han, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju. The disagreement problem in explainable machine learning: A practitioner’s perspective. arXiv, 2022.
  44. 44.Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An evaluation of the human-interpretability of explanation. arXiv, 2019.
  45. 45.Himabindu Lakkaraju and Osbert Bastani. “how do i fool you?” manipulating user trust via misleading black box explanations. In AAAI Conference on AIES, 2020.
  46. 46.Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec. Interpretable decision sets: A joint framework for description and prediction. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pages 1675–1684, 2016.
  47. 47.Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec. Faithful and customizable explanations of black box models. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 131–138, 2019.
  48. 48.Benjamin Letham, Cynthia Rudin, Tyler H McCormick, and David Madigan. Interpretable classifiers using rules and bayesian analysis: Building a better stroke prediction model. The Annals of Applied Statistics, 9(3):1350–1371, 2015.
  49. 49.Pantelis Linardatos, Vasilis Papastefanopoulos, and Sotiris Kotsiantis. Explainable ai: A review of machine learning interpretability methods. Entropy, 23(1):18, 2021.
  50. 50.Zachary C Lipton. The mythos of model interpretability. CoRR, abs/1606.03490, 2016.
  51. 51.Yang Liu, Sujay Khandagale, Colin White, and Willie Neiswanger. Synthetic benchmarks for scientific research in explainable machine learning. In NeurIPS Datasets and Benchmarks Track, 2021.
  52. 52.Arnaud Looveren and Janis Klaise. Interpretable counterfactual explanations guided by prototypes. CoRR, abs/ 1907.02584, 2019.
  53. 53.Yin Lou, Rich Caruana, and Johannes Gehrke. Intelligible models for classification and regression. In KDD, 2012.
  54. 54.Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Neural Information Processing Systems (NIPS), pages 4765–4774. Curran Associates, Inc., 2017.
  55. 55.Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, pages 4765–4774, 2017.
  56. 56.W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. Definitions, methods, and applications in interpretable machine learning. Proceedings of the National Academy of Sciences, 2019.
  57. 57.Martin Pawelczyk, Klaus Broelemann, and Gjergji Kasneci. Learning model-agnostic counterfactual explanations for tabular data. In WWW, 2020.
  58. 58.Martin Pawelczyk, Sascha Bielawski, Johan Van den Heuvel, Tobias Richter, and Gjergji Kasneci. Carla: A python library to benchmark algorithmic recourse and counterfactual explanation algorithms. In NeurIPS Benchmark and Datasets Track, 2021.
  59. 59.Vitali Petsiuk, Abir Das, and Kate Saenko. Rise: Randomized input sampling for explanation of black-box models. arXiv, 2018.
  60. 60.Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. Manipulating and measuring model interpretability. CoRR, 2018.
  61. 61.Rafael Poyiadzi, Kacper Sokol, Raul Santos-Rodriguez, Tijl De Bie, and Peter Flach. FACE: Feasible and actionable counterfactual explanations. In AAAI Conference on AIES, 2020.
  62. 62.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. "why should i trust you?": Explaining the predictions of any classifier. In KDD, 2016.
  63. 63.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. Anchors: High-precision model-agnostic explanations. In AAAI, 2018.
  64. 64.Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 2019.
  65. 65.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, 2017.
  66. 66.Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. In ICML, 2017.
  67. 67.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. In ICLR, 2014.
  68. 68.Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In AAAI Conference on AIES, 2020.
  69. 69.Dylan Slack, Anna Hilgard, Sameer Singh, and Himabindu Lakkaraju. Reliable post hoc explanations: Modeling uncertainty in explainability. NeurIPS, 2021.
  70. 70.Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. Smoothgrad: removing noise by adding noise. arXiv, 2017.
  71. 71.Jack W Smith, James E Everhart, WC Dickson, William C Knowler, and Robert Scott Johannes. Using the adap learning algorithm to forecast the onset of diabetes mellitus. In Proceedings of the annual symposium on computer application in medical care, page 261. American Medical Informatics Association, 1988.
  72. 72.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In ICML, 2017.
  73. 73.Sohini Upadhyay, Shalmali Joshi, and Himabindu Lakkaraju. Towards robust and reliable algorithmic recourse. NeurIPS, 2021.
  74. 74.Berk Ustun, Alexander Spangher, and Yang Liu. Actionable recourse in linear classification. In FAccT, 2019.
  75. 75.Sahil Verma, John Dickerson, and Keegan Hines. Counterfactual explanations for machine learning: A review. arXiv, 2020.
  76. 76.Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Technology, 31:841, 2017.
  77. 77.Fulton Wang and Cynthia Rudin. Falling rule lists. In Artificial Intelligence and Statistics, pages 1013–1022. PMLR, 2015.
  78. 78.Leanne S Whitmore, Anthe George, and Corey M Hudson. Mapping chemical performance on molecular structures using locally interpretable explanations. CoRR, abs/1611.07443, 2016.
  79. 79.I-Cheng Yeh and Che-hui Lien. The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients. In Expert Systems with Applications, 2009.
  80. 80.Jiaming Zeng, Berk Ustun, and Cynthia Rudin. Interpretable classification models for recidivism prediction. Journal of the Royal Statistical Society: Series A (Statistics in Society), 2017.
  81. 81.Jianlong Zhou, Amir H Gandomi, Fang Chen, and Andreas Holzinger. Evaluating the quality of machine learning explanations: A survey on methods and metrics. Electronics, 10(5):593, 2021.

Citation

MLA
Agarwal, C., et al. “OpenXAI: Towards a Transparent Evaluation of Model Explanations”. arXiv, 2022, http://arxiv.org/abs/2206.11104v5.
APA
Agarwal, C., Ley, D., Krishna, S., Saxena, E., Pawelczyk, M., Johnson, N., Puri, I., Zitnik, M., & Lakkaraju, H. (2022). OpenXAI: Towards a Transparent Evaluation of Model Explanations. arXiv. http://arxiv.org/abs/2206.11104v5
Chicago
Agarwal, C., D. Ley, S. Krishna, et al. 2022. “OpenXAI: Towards a Transparent Evaluation of Model Explanations”. arXiv. http://arxiv.org/abs/2206.11104v5.
Harvard
Agarwal, C. et al. (2022) “OpenXAI: Towards a Transparent Evaluation of Model Explanations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.11104v5.
Vancouver
1. Agarwal C, Ley D, Krishna S, Saxena E, Pawelczyk M, Johnson N, Puri I, Zitnik M, Lakkaraju H (2022) OpenXAI: Towards a Transparent Evaluation of Model Explanations. arXiv

BibTeX

@article{agarwal2022openxai,
  title = {OpenXAI: Towards a Transparent Evaluation of Model Explanations},
  author = {Agarwal, Chirag and Ley, Dan and Krishna, Satyapriya and Saxena, Eshika and Pawelczyk, Martin and Johnson, Nari and Puri, Isha and Zitnik, Marinka and Lakkaraju, Himabindu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.11104v5},
  eprint = {2206.11104}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors