Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations
Tessa HanSuraj SrinivasHimabindu Lakkaraju
Unifies eight widely used post hoc explanation methods under a local function approximation framework, establishing a no-free-lunch theorem and a principled criterion to resolve the disagreement problem when choosing feature attribution methods.
As machine learning systems are increasingly deployed in high-stakes environments such as medicine, law, and finance, practitioners must be able to understand and trust model predictions. While various post hoc explanation methods exist to highlight important features, they are built on disparate conceptual foundations, ranging from game-theoretic principles to gradient visualizations. This fragmentation leads to the "disagreement problem," where different explanation tools produce conflicting explanations for the exact same prediction, leaving practitioners to rely on arbitrary preferences rather than principled criteria.
The article establishes a unified mathematical framework demonstrating that eight prominent explanation methods all perform local function approximation of the underlying black-box model. Its primary objective is to clarify why these methods disagree, prove theoretical limits on their performance, and provide a practical guiding principle for selecting the most faithful explanation method for a given context.
The authors analyze eight widely used explanation tools—including LIME, KernelSHAP, Occlusion, SmoothGrad, and Integrated Gradients—by framing them as simpler interpretable models fitted to the black-box model over specific perturbation neighborhoods using defined loss functions. To validate the theoretical framework, the authors conducted empirical evaluations across regression and classification tasks using public healthcare (World Health Organization life expectancy) and financial (FICO credit line) datasets across various linear and neural network architectures.
The evaluation yielded three key findings. First, the analyzed methods are mathematically equivalent to local function approximation, differing primarily in the noise distribution (binary vs. continuous, additive vs. multiplicative) and the loss function used. Second, the article establishes a "no free lunch theorem for explanations," proving that no single explanation method can perform optimally across all neighborhood types because each is specialized to its own perturbation space. Third, the authors identify a model recovery property: when applied to continuous data, additive continuous noise methods (such as SmoothGrad and Vanilla Gradients) faithfully recover the true underlying linear model weights, whereas multiplicative continuous and binary methods (such as Integrated Gradients, LIME, and KernelSHAP) fail to do so without structural modifications.
These findings provide clarity for risk, compliance, and governance decisions in artificial intelligence deployment. They demonstrate that conflicting explanations do not necessarily mean the methods are defective; rather, each tool queries the model through a different operational lens. Misaligning the explanation method with the underlying data domain can generate unfaithful explanations, leading decision-makers to place misplaced trust in critical automated systems.
To ensure faithful explanations, practitioners should match the explanation technique to the data domain and evaluation regime. For continuous data, organizations should prioritize additive continuous noise methods such as SmoothGrad, Vanilla Gradients, or C-LIME. For binary data, binary perturbation methods like LIME and KernelSHAP are appropriate. For discrete data, new methods should be formulated within the local function approximation framework using discrete noise neighborhoods. Furthermore, model validation teams must align their evaluation benchmarks with the target neighborhood, as the choice of continuous or binary perturbation during testing directly dictates which method appears superior.
While this framework provides strong theoretical and empirical guarantees regarding model faithfulness, it focuses on mathematical fidelity rather than human usability. The authors note that defining true interpretability requires future human-computer interaction research and user studies. Nevertheless, decision-makers can have high confidence in using these domain-specific selection rules to eliminate arbitrary tool choices and standardize explainability pipelines.
- Paper: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Marco Tulio Ribeiro et al. (2016). Introduces LIME as a local surrogate model approximation technique, which serves as a core method analyzed and generalized in the source's local function approximation framework.
- Paper: A Unified Approach to Interpreting Model Predictions, Scott M. Lundberg et al. (2017). Introduces KernelSHAP and game-theoretic additive feature attributions, providing key foundational methods and concepts that the source mathematically unifies under local function approximation.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). Proposes the Integrated Gradients attribution method evaluated and characterized within the source's comparative mathematical framework.
- Paper: SmoothGrad: removing noise by adding noise, Daniel Smilkov et al. (2017). Introduces SmoothGrad by averaging gradient maps over Gaussian-perturbed inputs, representing one of the key additive continuous noise explanation methods analyzed by the source.
- Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). Establishes empirical sanity checks for feature attribution methods, directly motivating the source's investigation into explanation faithfulness and model recovery properties.
- Paper: Interpretable Explanations of Black Boxes by Meaningful Perturbation, Ruth Fong et al. (2017). Formulates black-box explainability around input perturbations, laying foundational ideas for viewing explanation methods through perturbation neighborhoods.
- Paper: Evaluating the Visualization of What a Deep Neural Network Has Learned, Wojciech Samek et al. (2015). Pioneers quantitative perturbation-based evaluation frameworks for explanation heatmaps, establishing metrics for assessing explanation fidelity.
- Paper: The Mythos of Model Interpretability, Zachary C. Lipton (2016). Provides a conceptual foundation for distinguishing post-hoc explanations from intrinsic transparency and clarifying what practitioners seek in model interpretability.
- Paper: OpenXAI: Towards a Transparent Evaluation of Model Explanations, Chirag Agarwal et al. (2022). Establishes an open-source benchmarking suite and standardized leaderboards to evaluate post hoc explanation methods across faithfulness and stability metrics.
- Paper: Faith-Shap: The Faithful Shapley Interaction Index, Che-Ping Tsai et al. (2023). Extends post hoc feature attribution by framing high-order Shapley interactions as optimal polynomial function approximations over set functions.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). Deepens the critique of post hoc interpretability methods by formalizing why surrogate and attribution methods suffer from mathematical non-identifiability.
- Paper: Faithfulness Tests for Natural Language Explanations, Pepa Atanasova et al. (2023). Extends the examination of post hoc explanation faithfulness into the domain of generated natural language rationales.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Demonstrates practical unfaithfulness in chain-of-thought explanations, illustrating the broader real-world consequences of post hoc rationalization.
- Paper: Learning to Estimate Shapley Values with Vision Transformers, Ian Connick Covert et al. (2023). Develops a practical surrogate estimation framework that learns to approximate Shapley values efficiently for vision transformers in a single forward pass.
