Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations

Tessa HanSuraj SrinivasHimabindu Lakkaraju

article2022NeurIPS126 citations

Unifies eight widely used post hoc explanation methods under a local function approximation framework, establishing a no-free-lunch theorem and a principled criterion to resolve the disagreement problem when choosing feature attribution methods.

Listen

As machine learning systems are increasingly deployed in high-stakes environments such as medicine, law, and finance, practitioners must be able to understand and trust model predictions. While various post hoc explanation methods exist to highlight important features, they are built on disparate conceptual foundations, ranging from game-theoretic principles to gradient visualizations. This fragmentation leads to the "disagreement problem," where different explanation tools produce conflicting explanations for the exact same prediction, leaving practitioners to rely on arbitrary preferences rather than principled criteria.

The article establishes a unified mathematical framework demonstrating that eight prominent explanation methods all perform local function approximation of the underlying black-box model. Its primary objective is to clarify why these methods disagree, prove theoretical limits on their performance, and provide a practical guiding principle for selecting the most faithful explanation method for a given context.

The authors analyze eight widely used explanation tools—including LIME, KernelSHAP, Occlusion, SmoothGrad, and Integrated Gradients—by framing them as simpler interpretable models fitted to the black-box model over specific perturbation neighborhoods using defined loss functions. To validate the theoretical framework, the authors conducted empirical evaluations across regression and classification tasks using public healthcare (World Health Organization life expectancy) and financial (FICO credit line) datasets across various linear and neural network architectures.

The evaluation yielded three key findings. First, the analyzed methods are mathematically equivalent to local function approximation, differing primarily in the noise distribution (binary vs. continuous, additive vs. multiplicative) and the loss function used. Second, the article establishes a "no free lunch theorem for explanations," proving that no single explanation method can perform optimally across all neighborhood types because each is specialized to its own perturbation space. Third, the authors identify a model recovery property: when applied to continuous data, additive continuous noise methods (such as SmoothGrad and Vanilla Gradients) faithfully recover the true underlying linear model weights, whereas multiplicative continuous and binary methods (such as Integrated Gradients, LIME, and KernelSHAP) fail to do so without structural modifications.

These findings provide clarity for risk, compliance, and governance decisions in artificial intelligence deployment. They demonstrate that conflicting explanations do not necessarily mean the methods are defective; rather, each tool queries the model through a different operational lens. Misaligning the explanation method with the underlying data domain can generate unfaithful explanations, leading decision-makers to place misplaced trust in critical automated systems.

To ensure faithful explanations, practitioners should match the explanation technique to the data domain and evaluation regime. For continuous data, organizations should prioritize additive continuous noise methods such as SmoothGrad, Vanilla Gradients, or C-LIME. For binary data, binary perturbation methods like LIME and KernelSHAP are appropriate. For discrete data, new methods should be formulated within the local function approximation framework using discrete noise neighborhoods. Furthermore, model validation teams must align their evaluation benchmarks with the target neighborhood, as the choice of continuous or binary perturbation during testing directly dictates which method appears superior.

While this framework provides strong theoretical and empirical guarantees regarding model faithfulness, it focuses on mathematical fidelity rather than human usability. The authors note that defining true interpretability requires future human-computer interaction research and user studies. Nevertheless, decision-makers can have high confidence in using these domain-specific selection rules to eliminate arbitrary tool choices and standardize explainability pipelines.

arXiv: 2206.01254
Cover for Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations

Abstract

A critical problem in the field of post hoc explainability is the lack of a common foundational goal among methods. For example, some methods are motivated by function approximation, some by game theoretic notions, and some by obtaining clean visualizations. This fragmentation of goals causes not only an inconsistent conceptual understanding of explanations but also the practical challenge of not knowing which method to use when.

In this work, we begin to address these challenges by unifying eight popular post hoc explanation methods (LIME, C-LIME, KernelSHAP, Occlusion, Vanilla Gradients, Gradients × Input, SmoothGrad, and Integrated Gradients). We show that these methods all perform local function approximation of the black-box model, differing only in the neighbourhood and loss function used to perform the approximation. This unification enables us to (1) state a no free lunch theorem for explanation methods, demonstrating that no method can perform optimally across all neighbourhoods, and (2) provide a guiding principle to choose among methods based on faithfulness to the black-box model. We empirically validate these theoretical results using various real-world datasets, model classes, and prediction tasks.

By bringing diverse explanation methods into a common framework, this work (1) advances the conceptual understanding of these methods, revealing their shared local function approximation objective, properties, and relation to one another, and (2) guides the use of these methods in practice, providing a principled approach to choose among methods and paving the way for the creation of new ones.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Explanation as Local Function Approximation
  • 3.1 LFA with Continuous Noise: Gradient-Based Explanation Methods
  • 3.2 LFA with Binary Noise: LIME, KernelSHAP and Occlusion maps
  • 3.3 Which Methods Do Not Perform LFA?
  • 4 When Do Explanations Perform Model Recovery?
  • 4.1 No Free Lunch Theorem for Explanation Methods
  • 4.2 Characterizing Explanation Methods via Model Recovery
  • 4.3 Designing Novel Explanations with LFA
  • 5 Empirical Evaluation
  • 5.1 Datasets, Models, and Metrics
  • 5.2 Experiments
  • 6 Conclusions and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Local function approximation framework

    definition

    Let f:X→Yf:\mathcal{X}\to\mathcal{Y} be a black-box model, let G⊆{g:X→Y}\mathcal{G}\subseteq\{g:\mathcal{X}\to\mathcal{Y}\} be an interpretable model class, and let x0∈Xx_0\in\mathcal{X} be the point to explain. A local neighborhood is generated by a random perturbation ξ∼Z\xi\sim\mathcal{Z} and a specified binary operator ⊕\oplus, producing xξ=x0⊕ξx_\xi=x_0\oplus\xi. For a nonnegative loss ℓ(f,g,x0,ξ)\ell(f,g,x_0,\xi), local function approximation (LFA) chooses

    g∗=arg⁡min⁡g∈GEξ∼Z[ℓ(f,g,x0,ξ)].g^*=\arg\min_{g\in\mathcal{G}}\mathbb{E}_{\xi\sim\mathcal{Z}}\left[\ell(f,g,x_0,\xi)\right].

    The loss is valid for LFA when zero expected loss is equivalent to exact agreement, namely

    Eξ∼Z[ℓ(f,g,x0,ξ)]=0⟺f(xξ)=g(xξ) for every ξ in the support of Z.\mathbb{E}_{\xi\sim\mathcal{Z}}\left[\ell(f,g,x_0,\xi)\right]=0 \quad\Longleftrightarrow\quad f(x_\xi)=g(x_\xi)\ \text{for every }\xi\text{ in the support of }\mathcal{Z}.

    When f∈Gf\in\mathcal{G} and the perturbations cover the input domain X\mathcal{X}, this condition makes the LFA solution recover the entire black-box function, g∗=fg^*=f. The framework therefore characterizes an explanation by four choices: the interpretable class G\mathcal{G}, neighborhood distribution Z\mathcal{Z}, loss ℓ\ell, and perturbation operator ⊕\oplus.

  2. Knowl 2 — Eight explanation methods as linear local approximators

    model/method

    The paper represents all eight analyzed methods with a linear interpretable model g(x)=w⊤xg(x)=w^\top x and identifies each method by its perturbation neighborhood and loss. Here x0∈Rdx_0\in\mathbb{R}^d is the input being explained, ξ\xi is the noise variable, ⊙\odot denotes componentwise multiplication, and σ>0\sigma>0 is the standard deviation of an isotropic Gaussian.

    Could not parse LaTeX table

    For LIME, KernelSHAP, and Occlusion, the stated kernel or one-hot distribution defines the local neighborhood; the corresponding explanations arise from the LFA objective, with kernel-weighted methods understood through importance sampling. Thus the methods differ in neighborhood and loss rather than in their underlying goal of locally approximating ff.

  3. Knowl 3 — Gradient matching connects continuous gradient explanations

    theoretical result

    For differentiable scalar-valued functions ff and gg and a continuous perturbation ξ\xi, the paper uses the gradient-matching loss

    ℓgm(f,g,x0,ξ)=∥∇ξf(x0⊕ξ)−∇ξg(x0⊕ξ)∥22.\ell_{\mathrm{gm}}(f,g,x_0,\xi) =\left\|\nabla_{\xi}f(x_0\oplus\xi)-\nabla_{\xi}g(x_0\oplus\xi)\right\|_2^2.

    This is a valid LFA loss up to an additive constant: its expected value is zero exactly when there is a constant C∈RC\in\mathbb{R} such that f(x0⊕ξ)=g(x0⊕ξ)+Cf(x_0\oplus\xi)=g(x_0\oplus\xi)+C throughout the perturbation neighborhood. Consequently, when g(x)=w⊤x+bg(x)=w^\top x+b is linear, gradient matching identifies the weights ww but not the intercept bb; setting b=f(0)b=f(0) removes this ambiguity.

    Under additive Gaussian noise, minimizing this loss yields SmoothGrad, and taking the Gaussian standard deviation to zero yields Vanilla Gradients. Under multiplicative uniform noise, minimizing it yields Integrated Gradients, and shrinking the uniform interval toward 11 yields Gradients ×\times Input. The gradient-matching formulation therefore explains both the equivalence of these methods and their limiting relationships.

  4. Knowl 4 — Binary perturbation methods are LFA instances

    theoretical result

    For a binary perturbation ξ∈{0,1}d\xi\in\{0,1\}^d, the perturbed input is x0⊙ξx_0\odot\xi, where ⊙\odot is componentwise multiplication. LFA with this multiplicative binary noise and squared-error loss is equivalent to three existing methods: an exponential-kernel distribution gives LIME, a Shapley-kernel distribution gives KernelSHAP, and a distribution over one-hot vectors gives Occlusion.

    For LIME and KernelSHAP, the correspondence follows because their kernel-weighted regression objectives are precisely the LFA squared-error objective under the corresponding neighborhood distributions. For Occlusion, enumerating the one-hot perturbations and optimizing the resulting squared-error objective produces its feature-importance explanation. These equivalences apply to binary perturbation domains; gradient-based continuous-noise reasoning does not extend to discrete binary perturbations.

  5. Knowl 5 — No free lunch for explanation neighborhoods

    theoretical result

    Let ff be a black-box model on input domain X\mathcal{X}, let G\mathcal{G} be the interpretable model class, and let ℓ\ell be a valid LFA loss. Define the worst-case distance from ff to the model class by

    d(f,G)=min⁡g∈Gmax⁡x∈Xℓ(f,g,0,x),d(f,\mathcal{G})=\min_{g\in\mathcal{G}}\max_{x\in\mathcal{X}}\ell(f,g,0,x),

    where 00 is the reference point used in the loss and xx ranges over the input domain. For any explanation g∗g^* obtained using a neighborhood distribution Z1\mathcal{Z}_1 such that

    max⁡ξ1∈supp⁡(Z1)ℓ(f,g∗,x0,ξ1)≤ε,\max_{\xi_1\in\operatorname{supp}(\mathcal{Z}_1)}\ell(f,g^*,x_0,\xi_1)\leq\varepsilon,

    there exists another neighborhood distribution Z2\mathcal{Z}_2 for which

    max⁡ξ2∈supp⁡(Z2)ℓ(f,g∗,x0,ξ2)≥d(f,G).\max_{\xi_2\in\operatorname{supp}(\mathcal{Z}_2)}\ell(f,g^*,x_0,\xi_2)\geq d(f,\mathcal{G}).

    Therefore, no single explanation can be uniformly optimal over all neighborhoods. When G\mathcal{G} is less expressive than ff, the lower bound can be large; choosing a universally best explanation without specifying the perturbation neighborhood is consequently impossible. Once a neighborhood is fixed, the appropriate LFA minimizer is the best explanation for that neighborhood.

  6. Knowl 6 — Model recovery as the explanation-selection principle

    definition

    The paper defines model recovery as a guiding principle for selecting explanation methods. Given an LFA instance with black-box model f∈Gf\in\mathcal{G} and a specified noise type, the explanation method satisfies model recovery if there exists a noise distribution Z\mathcal{Z} of that type for which the LFA solution is exactly the black-box model, g∗=fg^*=f.

    For continuous inputs X=Rd\mathcal{X}=\mathbb{R}^d and linear models f(x)=wf⊤xf(x)=w_f^\top x and g(x)=wg⊤xg(x)=w_g^\top x, the paper obtains the following distinction. Additive continuous-noise methods—C-LIME, SmoothGrad, and Vanilla Gradients—recover the true weights, wg=wfw_g=w_f. Multiplicative continuous-noise methods—Integrated Gradients and Gradients ×\times Input—and multiplicative binary-noise methods—LIME, KernelSHAP, and Occlusion—recover input-scaled weights instead, wg=wf⊙x0w_g=w_f\odot x_0.

    The failure of the multiplicative continuous methods is attributed to their loss parameterization rather than to multiplication itself. Replacing a term of the form ∇ξg(ξ)\nabla_\xi g(\xi) with ∇ξg(x0⊕ξ)\nabla_\xi g(x_0\oplus\xi) would restore model recovery; under this change, Integrated Gradients becomes the unscaled path integral ∫01∇αx0f(αx0) dα\int_0^1\nabla_{\alpha x_0}f(\alpha x_0)\,d\alpha, while Gradients ×\times Input becomes Vanilla Gradients. Analogously, replacing g(ξ)g(\xi) with g(x0⊕ξ)g(x_0\oplus\xi) in the binary squared-error loss can enable recovery in some cases, but does not guarantee it generally.

  7. Knowl 7 — Input-domain restrictions and the discrete-noise failure case

    limitation

    The recommended explanation method depends on the input domain. For binary inputs, continuous-noise methods are invalid, while multiplicative binary methods such as LIME, KernelSHAP, and Occlusion can recover models in the binary setting. For discrete inputs, continuous perturbations are invalid and the existing multiplicative binary methods do not guarantee model recovery; the paper notes that none of the analyzed methods provides general discrete perturbations. It therefore recommends constructing a new LFA method with a suitable discrete neighborhood.

    The lack of a general guarantee for binary perturbations on continuous or discrete-like behavior is illustrated by periodic models. Let

    f(x)=∑i=1dsin⁡(wfixi),g(x)=∑i=1dsin⁡(wgixi),f(x)=\sum_{i=1}^{d}\sin(w_{fi}x_i),\qquad g(x)=\sum_{i=1}^{d}\sin(w_{gi}x_i),

    where x∈Rdx\in\mathbb{R}^d, and suppose that for some coordinate ii and positive integer nn, the explained point satisfies wfix0i=±nπw_{fi}x_{0i}=\pm n\pi. Under multiplicative binary perturbations, that coordinate is evaluated only at 00 or x0ix_{0i}, and both give zero for the corresponding sine term. The perturbations therefore contain no information about that frequency, so model recovery is impossible for this case. More generally, discrete perturbations can fail to recover models with sufficiently large or poorly aligned frequency components.

  8. Knowl 8 — LFA provides a recipe for designing new explanations

    model/method

    A new explanation method can be defined by specifying four components: an interpretable model class G\mathcal{G}, a neighborhood distribution Z\mathcal{Z}, a loss function ℓ\ell, and an operator ⊕\oplus that combines the point being explained with perturbations. The resulting method is the LFA optimization problem over those choices, even when the optimization has no closed-form solution.

    The paper illustrates this construction with SparseSmoothGrad. Starting from the SmoothGrad LFA loss ℓSmoothGrad\ell_{\mathrm{SmoothGrad}}, it adds an L0L_0 penalty on the gradient of the interpretable model:

    ℓSparseSmoothGrad=ℓSmoothGrad+λ∥∇ξg(x0⊕ξ)∥0,\ell_{\mathrm{SparseSmoothGrad}} =\ell_{\mathrm{SmoothGrad}}+\lambda\left\|\nabla_{\xi}g(x_0\oplus\xi)\right\|_0,

    where λ≥0\lambda\geq0 controls sparsity and ∥⋅∥0\|\cdot\|_0 counts nonzero coordinates. Sparse solvers can then be used to obtain an explanation with only a small number of nonzero feature gradients. The example demonstrates that LFA supports systematic variations of existing explanations rather than only analysis of already-established methods.

  9. Knowl 9 — Empirical correspondence between existing methods and LFA

    experimental setup

    The empirical study used two real-world continuous-feature datasets: the WHO life-expectancy dataset, with 2,938 observations and 20 features for regression, and the FICO HELOC dataset, with 9,871 observations and 24 features for classification. For each dataset, the authors trained four models: a simple model—linear regression for WHO or logistic regression for HELOC—and three neural networks of increasing complexity. Explanation similarity was measured with L1L_1 distance and cosine distance, where smaller values indicate more similar explanations.

    For 100 randomly selected test points, seven existing methods were compared with their corresponding LFA implementations. The paired-explanation distance heatmap had its smallest values when each existing method was compared with its matching LFA instance, demonstrating empirical equivalence. As the Gaussian scale decreased, the LFA version of SmoothGrad converged toward Vanilla Gradients; as the uniform multiplicative interval contracted toward 11, the LFA version of Integrated Gradients converged toward Gradients ×\times Input. The observed clustering also showed that SmoothGrad and Vanilla Gradients produce similar explanations, while LIME, KernelSHAP, Occlusion, Integrated Gradients, and Gradients ×\times Input form another similarity group on continuous data.

  10. Knowl 10 — Empirical model-recovery results

    empirical result

    To test the model-recovery principle, the authors used the WHO dataset with a linear-regression black-box model ff and linear interpretable model gg, generated explanations for 100 randomly selected test points, and compared the weights of gg with either the true weights of ff or the input-scaled weights of ff. On continuous data, SmoothGrad and Vanilla Gradients recovered the black-box weights, satisfying the proposed principle. LIME, KernelSHAP, Occlusion, Integrated Gradients, and Gradients ×\times Input instead matched the input-scaled weight pattern predicted by the theory and therefore failed to recover ff itself.

    The same qualitative model-recovery behavior was observed on the HELOC dataset using logistic-regression black-box and interpretable models. These results support using model recovery as a practical criterion for distinguishing methods that preserve the underlying model from methods that produce explanations on a different scale.

  11. Knowl 11 — Empirical demonstration of neighborhood-dependent performance

    empirical result

    The no-free-lunch result was evaluated with perturbation tests on explanations for 100 randomly selected test points from the WHO dataset using a neural network with three hidden layers. For a given explanation and integer kk, the tests selected either the top-kk or bottom-kk features, then perturbed them by setting them to zero for binary perturbations or adding Gaussian noise for continuous perturbations. For top-kk tests, the absolute change in the model prediction measured how well an explanation identified important features; for bottom-kk tests, a smaller prediction change indicated better identification of unimportant features. One binary perturbation and 100 continuous perturbations were used per point, with the continuous result averaged over the random perturbations.

    LIME, KernelSHAP, Occlusion, Integrated Gradients, and Gradients ×\times Input performed best under binary perturbation neighborhoods, whereas SmoothGrad and Vanilla Gradients performed best under continuous Gaussian neighborhoods. The same qualitative pattern held across top-kk and bottom-kk tests, datasets, and model classes. Thus, the perturbation distribution used to evaluate an explanation can itself determine which explanation method appears best, providing empirical support for the no-free-lunch theorem.

Coverage note — The paper’s background, related work, appendix-level implementation details, future-work discussion, and the nontechnical limitation concerning human-centered interpretability were omitted because they do not add separate load-bearing method, theory, or empirical-result knowls.

References

  1. 1.Kun-Hsing Yu, Andrew L Beam, and Isaac S Kohane. Artificial intelligence in healthcare. Nature Biomedical Engineering, 2(10):719–731, 2018.
  2. 2.Robert Walters and Marko Novak. Artificial intelligence and law. In Cyber Security, Artificial Intelligence, Data Protection & the Law, pages 39–69. Springer, 2021.
  3. 3.Longbing Cao. AI in finance: Challenges, techniques, and opportunities. ACM Computing Surveys (CSUR), 55(3):1–38, 2022.
  4. 4.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. "Why should I trust you?" Explaining the predictions of any classifier. International Conference on Knowledge Discovery and Data Mining, 2016.
  5. 5.Sushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay, Steven Wu, and Himabindu Lakkaraju. Towards the unification and robustness of perturbation and gradient based explanations. International Conference on Machine Learning, 2021.
  6. 6.Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 2017.
  7. 7.Matthew Zeiler and Robert Fergus. Visualizing and understanding convolutional networks. arXiv:1311.2901, 2013.
  8. 8.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. Workshop at International Conference on Learning Representations, 2014.
  9. 9.Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. International Conference on Machine Learning, 2017.
  10. 10.Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. SmoothGrad: Removing noise by adding noise. arXiv:1706.03825, 2017.
  11. 11.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. International Conference on Machine Learning, 2017.
  12. 12.Satyapriya Krishna*, Tessa Han*, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju. The disagreement problem in explainable machine learning: A practitioner’s perspective. arXiv:2202.01602, 2022.
  13. 13.Ian Covert, Scott Lundberg, and Su-In Lee. Explaining by removing: A unified framework for model explanation. Journal of Machine Learning Research, 2021.
  14. 14.Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. Towards better understanding of gradient-based attribution methods for deep neural networks. International Conference on Learning Representations, 2018.
  15. 15.Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity checks for saliency maps. Advances in Neural Information Processing Systems, 2018.
  16. 16.Suraj Srinivas and François Fleuret. Knowledge transfer with Jacobian matching. International Conference on Machine Learning, 2018.
  17. 17.Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. A benchmark for interpretability methods in deep neural networks. Advances in Neural Information Processing Systems, 2019.
  18. 18.Amirata Ghorbani, Abubakar Abid, and James Zou. Interpretation of neural networks is fragile. AAAI Conference on Artificial Intelligence, 2019.
  19. 19.Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods. AAAI/ACM Conference on AI, Ethics, and Society, 2020.
  20. 20.Ann-Kathrin Dombrowski, Maximilian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. Explanations can be manipulated and geometry is to blame. Advances in Neural Information Processing Systems, 2019.
  21. 21.David Alvarez-Melis and Tommi Jaakkola. On the robustness of interpretability methods. arXiv:1806.08049, 2018.
  22. 22.Jessica Dai, Sohini Upadhyay, Ulrich Aivodji, Stephen Bach, and Himabindu Lakkaraju. Fairness via explanation quality: Evaluating disparities in the quality of post hoc explanations. AAAI/ACM Conference on AI, Society, and Ethics, 2022.
  23. 23.Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 2005.
  24. 24.Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. Striving for simplicity: The all convolutional net. arXiv:1412.6806, 2015.
  25. 25.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. arXiv:1505.04366, 2015.
  26. 26.Ramprasaath Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-CAM: Visual explanations from deep networks via gradient-based localization. IEEE International Conference on Computer Vision, 2017.
  27. 27.Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth Balasubramanian. Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks. IEEE Winter Conference on Applications of Computer Vision (WACV), 2018.
  28. 28.Suraj Srinivas and François Fleuret. Full-gradient representation for neural network visualization. Advances in Neural Information Processing Systems, 2019.
  29. 29.World Health Organization (WHO). Life expectancy dataset. Global Health Observatory Data Repository, 2018.
  30. 30.FICO. Home equity line of credit (HELOC) dataset. Explainable Machine Learning Challenge, 2019.
  31. 31.Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. Captum: A unified and generic model interpretability library for PyTorch, 2020.
  32. 32.Zachary Lipton. The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 2018.
  33. 33.Weili Nie, Yang Zhang, and Ankit Patel. A theoretical explanation for perplexing behaviors of backpropagation-based visualizations. International Conference on Machine Learning, 2018.
  34. 34.Suraj Srinivas and Francois Fleuret. Rethinking the role of gradient-based attribution methods for model interpretability. International Conference on Learning Representations, 2021.
  35. 35.Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS One, 2015.

Citation

MLA
Han, T., et al. “Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 5256–68, https://proceedings.neurips.cc/paper_files/paper/2022/file/22b111819c74453837899689166c4cf9-Paper-Conference.pdf.
APA
Han, T., Srinivas, S., & Lakkaraju, H. (2022). Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations. Advances in Neural Information Processing Systems, 35, 5256–5268. https://proceedings.neurips.cc/paper_files/paper/2022/file/22b111819c74453837899689166c4cf9-Paper-Conference.pdf
Chicago
Han, T., S. Srinivas, and H. Lakkaraju. 2022. “Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations”. Advances in Neural Information Processing Systems 35: 5256–68. https://proceedings.neurips.cc/paper_files/paper/2022/file/22b111819c74453837899689166c4cf9-Paper-Conference.pdf.
Harvard
Han, T., Srinivas, S. and Lakkaraju, H. (2022) “Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 5256–5268. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/22b111819c74453837899689166c4cf9-Paper-Conference.pdf.
Vancouver
1. Han T, Srinivas S, Lakkaraju H (2022) Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 5256–5268

BibTeX

@inproceedings{han2022which,
  title = {Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations},
  author = {Han, Tessa and Srinivas, Suraj and Lakkaraju, Himabindu},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {5256-5268},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/22b111819c74453837899689166c4cf9-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors