Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models

Wojciech SamekThomas WiegandKlaus-Robert Müller

article2017arXiv1,527 citations

Presents and evaluates two practical interpretability methods—input sensitivity analysis and decision decomposition—to explain how complex deep learning models arrive at their predictions across diverse classification tasks.

Listen

Modern artificial intelligence (AI) systems, particularly deep neural networks, achieve high accuracy across complex tasks such as image classification, language processing, and strategic games. However, their complex internal structures make them non-transparent "black boxes." In critical applications such as healthcare, autonomous driving, and financial services, this opacity poses severe risks. Without the ability to understand why an AI system makes a specific decision, organizations cannot easily detect flaws, verify safety, or satisfy legal mandates such as the European Union's "right to explanation."

The article advocates for explainable AI and compares two primary techniques for interpreting individual model predictions: sensitivity analysis and layer-wise relevance propagation. Through structured benchmarks, the authors evaluate how accurately each method explains model behavior across image recognition, text document classification, and human action recognition in video.

The evaluated methods take distinct technical approaches. Sensitivity analysis uses local gradients to measure how small adjustments to input features change the final output. In contrast, layer-wise relevance propagation redistributes the model's final decision score backward through the network layers, conserving the total score and decomposing it into positive and negative relevance values for each input feature. To evaluate both methods objectively, the authors applied perturbation analysis across thousands of samples, measuring how rapidly prediction accuracy dropped when removing the features identified as most important.

The findings show that layer-wise relevance propagation outperforms sensitivity analysis across all tested domains. In image classification, it accurately highlights core defining features (such as the contour of a cup) while sensitivity analysis produces noisy heatmaps that highlight irrelevant background regions. Perturbation tests over 5,040 images confirmed that masking features selected by layer-wise relevance propagation caused a significantly steeper drop in model confidence. In text analysis across 4,154 documents, layer-wise relevance propagation distinctly separated supporting evidence from contradictory evidence, leading to faster accuracy degradation under perturbation. In video action recognition, the method accurately isolated the specific spatial body movements and timeframes driving the classification.

These results demonstrate that simply measuring model sensitivity is insufficient for operational decision validation, as sensitivity highlights what could change a score rather than what actually produced it. Decomposing predictions through layer-wise relevance propagation allows practitioners to identify dataset biases (such as spurious statistical correlations), select models based on proper reasoning rather than just raw accuracy, and meet strict compliance requirements.

Organizations deploying deep learning in high-stakes environments should implement prediction decomposition tools to audit decisions and monitor compliance. The article's conclusions are strongly supported for feed-forward neural networks, convolutional architectures, and support vector machines across standard benchmarks. However, leaders should note that the evaluation is limited to post-hoc explanation methods on specific classification tasks. Future work is required to embed explainability directly into model architectures and extend these interpretation frameworks to additional operational domains.

arXiv: 1708.08296
Cover for Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models

Abstract

With the availability of large databases and recent improvements in deep learning methodology, the performance of AI systems is reaching or even exceeding the human level on an increasing number of complex tasks. Impressive examples of this development can be found in domains such as image classification, sentiment analysis, speech understanding or strategic game playing. However, because of their nested non-linear structure, these highly successful machine learning and artificial intelligence models are usually applied in a black box manner, i.e., no information is provided about what exactly makes them arrive at their predictions. Since this lack of transparency can be a major drawback, e.g., in medical applications, the development of methods for visualizing, explaining and interpreting deep learning models has recently attracted increasing attention. This paper summarizes recent developments in this field and makes a plea for more interpretability in artificial intelligence. Furthermore, it presents two approaches to explaining predictions of deep learning models, one method which computes the sensitivity of the prediction with respect to changes in the input and one approach which meaningfully decomposes the decision in terms of the input variables. These methods are evaluated on three classification tasks.

Table of Contents

  • 1 Introduction
  • 2 Why do we need explainable AI ?
  • 3 Methods for Visualizing, Interpreting and Explaining Deep Learning Models
  • 3.1 Sensitivity Analysis
  • 3.2 Layer-Wise Relevance Propagation
  • 3.3 Software
  • 4 Evaluating the Quality of Explanations
  • 5 Experimental Evaluation
  • 5.1 Image Classification
  • 5.2 Text Document Classification
  • 5.3 Human Action Recognition in Videos
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Layer-Wise Relevance Propagation and Relevance Conservation

    model/method

    Layer-Wise Relevance Propagation (LRP) is a framework for explaining individual classification decisions of non-linear models (such as deep neural networks, recurrent networks, and Fisher Vector classifiers) by decomposing the model's output prediction f(x)f(x) for an input sample xx into contributions from individual input dimensions (such as pixels or words).

    LRP redistributes the output prediction score backward across the network's layers until every input variable ii is assigned a relevance score RiR_i. The backward propagation satisfies the relevance conservation property at every layer:

    ∑iRi=⋯=∑jRj=∑kRk=⋯=f(x)\sum_i R_i = \dots = \sum_j R_j = \sum_k R_k = \dots = f(x)

    where ii indexes the input features, jj indexes neurons in an intermediate layer ll, and kk indexes neurons in the subsequent higher layer l+1l+1. No relevance is artificially added or destroyed at any step, ensuring that LRP provides an exact decomposition of the prediction value f(x)f(x) relative to a state of maximum uncertainty.

  2. Knowl 2 — Layer-Wise Relevance Propagation Redistribution Rules

    equation

    For a feedforward neural network, let xjx_j be the activation of neuron jj at layer ll, wjkw_{jk} be the synaptic weight connecting neuron jj in layer ll to neuron kk in layer l+1l+1, and RkR_k be the relevance assigned to neuron kk. The backward propagation of relevance from layer l+1l+1 to layer ll is governed by local redistribution rules.

    1. Standard ϵ\epsilon-Stabilized LRP Rule: Rj=∑kxjwjk∑j′xj′wj′k+ϵRkR_j = \sum_k \frac{x_j w_{jk}}{\sum_{j'} x_{j'} w_{j'k} + \epsilon} R_k where ϵ>0\epsilon > 0 is a small numerical stabilization constant used to avoid division by zero. Exact conservation of relevance holds when ϵ=0\epsilon = 0.

    2. αβ\alpha\beta-LRP Rule: Rj=∑k(α(xjwjk)+∑j′(xjwjk)+−β(xjwjk)−∑j′(xjwjk)−)RkR_j = \sum_k \left( \alpha \frac{(x_j w_{jk})^+}{\sum_{j'} (x_j w_{jk})^+} - \beta \frac{(x_j w_{jk})^-}{\sum_{j'} (x_j w_{jk})^-} \right) R_k where (⋅)+=max⁡(0,⋅)(\cdot)^+ = \max(0, \cdot) and (⋅)−=max⁡(0,−⋅)(\cdot)^- = \max(0, -\cdot) denote the positive and negative components, subject to the conservation constraint α−β=1\alpha - \beta = 1. When α=1\alpha = 1 and β=0\beta = 0, this rule corresponds to a first-order Deep Taylor Decomposition of networks constructed with rectified linear units (ReLU).

  3. Knowl 3 — Sensitivity Analysis for Explaining Model Predictions

    model/method

    Sensitivity Analysis (SA) explains a model prediction by measuring the local gradient of the output function f(x)f(x) with respect to each input feature xix_i. The relevance score RiR_i of input variable ii is defined as the magnitude of the partial derivative:

    Ri=∣∂f(x)∂xi∣R_i = \left| \frac{\partial f(x)}{\partial x_i} \right|

    Unlike decomposition methods, Sensitivity Analysis reflects how sensitive the output is to infinitesimal local variations in the input rather than quantifying how much each feature contributes to the magnitude of f(x)f(x). Consequently, SA highlights features that could be altered to change the prediction score (such as occlusions or background context) rather than isolating the features that positively support the predicted class.

  4. Knowl 4 — Perturbation Analysis for Objective Evaluation of Explanation Heatmaps

    algorithm

    The quality of attribution heatmaps produced by different explanation methods is quantitatively evaluated by measuring how rapidly a model's prediction score decays when input features are perturbed in order of their assigned relevance.

    Input: Model f, input sample x, attribution method M, step size S, total steps T
    Output: Prediction score trajectory S_0, S_1, ..., S_T
    Compute relevance scores R = M(f, x) for all input variables
    Sort input variables in descending order of relevance: (v_1, v_2, ..., v_D)
    Initialize perturbed input x^(0) = x
    Compute baseline score S_0 = f(x^(0))
    for t = 1 to T do
        Select the next batch of top-ranked unperturbed features B_t = {v_{(t-1)*S + 1}, ..., v_{t*S}}
        Perturb features in B_t within x^(t-1) (e.g., replace with random samples from a uniform distribution) to obtain x^(t)
        Evaluate new model score S_t = f(x^(t))
    end
    return (S_0, S_1, ..., S_T)

    A faster decline in prediction score (or classification accuracy) across perturbation steps indicates a higher-quality explanation, confirming that the method accurately identified the input features most vital to the model's decision.

  5. Knowl 5 — Empirical Comparison of LRP and Sensitivity Analysis on Image Classification

    empirical result

    On the ILSVRC2012 image classification benchmark using the GoogLeNet architecture, Layer-Wise Relevance Propagation (LRP) and Sensitivity Analysis (SA) were compared qualitatively and quantitatively via perturbation analysis over 5,040 test images.

    • Qualitative Visualizations: LRP produces class-specific heatmaps that focus tightly on defining object characteristics (such as the rim of a coffee cup or mountain slopes for a volcano). In contrast, SA heatmaps are noisy and assign high relevance to non-indicative background areas (such as open sky) because those pixels have high gradient sensitivity.
    • Perturbation Benchmark: When replacing 9×99 \times 9 image patches in descending order of relevance with uniform random values, LRP causes a significantly steeper decline in average prediction score than SA (e.g., the relative score drops to approximately 0.4 within 30–40 patch deletions under LRP, whereas under SA it remains near 0.6).
  6. Knowl 6 — Signed Evidence Decomposition and Perturbation Performance in Text Document Classification

    empirical result

    In document topic classification on the 20Newsgroups dataset using a word-embedding Convolutional Neural Network evaluated across 4,154 test documents:

    • Signed Evidence Separation: LRP explicitly separates positive relevance (words supporting the predicted category, such as sickness, body, and discomfort for the class sci.med) from negative relevance (words supporting competing classes, such as astronaut, ride, and shuttle which speak for sci.space). Sensitivity Analysis only produces non-negative derivative magnitudes, failing to distinguish supportive from conflicting evidence.
    • Quantitative Perturbation: Iteratively deleting words in descending order of relevance (setting their inputs to 0) causes a substantially faster reduction in classification accuracy for LRP-ordered words compared to SA-ordered words over 50 perturbation steps.
  7. Knowl 7 — Spatio-Temporal Relevance Attribution for Compressed Video Action Recognition

    empirical result

    Evaluating a Fisher Vector / Support Vector Machine classifier trained on block-wise motion vectors from compressed video on the HMDB51 human action dataset demonstrates that LRP produces both spatial and temporal explanations:

    • Spatial Attribution: LRP heatmaps locate the spatial blocks corresponding to indicative motion trajectories within individual video frames (e.g., highlighting blocks around the torso and upper body during a sit-up action).
    • Temporal Attribution: Aggregating relevance scores over consecutive frame windows generates a temporal activation profile where relevance peaks precisely during phases of rapid upward and downward body motion, identifying both where and when discriminative actions occur.
  8. Knowl 8 — Motivations for Explainability in Artificial Intelligence Systems

    definition

    The necessity of explainability in artificial intelligence is characterized by four core operational and societal requirements:

    1. System Verification: Validating that high predictive accuracy does not stem from dataset artifacts or non-causal correlations (e.g., preventing a medical model from learning that asthma reduces pneumonia mortality because high-risk asthma patients receive intensified clinical care).
    2. System Improvement: Identifying weaknesses, failure modes, and architectural differences in feature utilization between models that exhibit identical top-line accuracy.
    3. Knowledge Transfer: Extracting distilled insights, hidden patterns, or domain strategies discovered by AI systems trained on large-scale datasets for scientific discovery and human learning.
    4. Compliance and Legal Liability: Adhering to legal mandates (such as the European Union's regulatory 'right to explanation') and establishing accountability for automated decisions in safety-critical applications.

Coverage note — Omitted brief discussions on general software implementation details (such as the LRP Toolbox URLs) and broad survey remarks on unrelated deep learning architectures that do not form part of the paper's core methodological or empirical contributions.

References

  1. 1.M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016.
  2. 2.L. Arras, F. Horn, G. Montavon, K.-R. Müller, and W. Samek. Explaining predictions of non-linear classifiers in nlp. In Proceedings of the 1st Workshop on Representation Learning for NLP, pages 1–7. ACL, 2016.
  3. 3.L. Arras, F. Horn, G. Montavon, K.-R. Müller, and W. Samek. ”What is relevant in a text document?”: An interpretable machine learning approach. PLoS ONE, 12(8):e0181142, 2017.
  4. 4.L. Arras, G. Montavon, K.-R. Müller, and W. Samek. Explaining recurrent neural network predictions in sentiment analysis. In Proceedings of the EMNLP’17 Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis (WASSA), pages 1–10, 2017.
  5. 5.S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7):e0130140, 2015.
  6. 6.D. Baehrens, T. Schroeter, S. Harmeling, M. Kawanabe, K. Hansen, and K.-R. Müller. How to explain individual classification decisions. Journal of Machine Learning Research, 11:1803–1831, 2010.
  7. 7.R. Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, and N. Elhadad. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1721–1730, 2015.
  8. 8.K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078, 2014.
  9. 9.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255, 2009.
  10. 10.L. Deng, G. Hinton, and B. Kingsbury. New types of deep neural network learning for speech recognition and related applications: An overview. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8599–8603, 2013.
  11. 11.F. Doshi-Velez and B. Kim. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608, 2017.
  12. 12.D. Erhan, Y. Bengio, A. Courville, and P. Vincent. Visualizing higher-layer features of a deep network. Technical Report 1341, University of Montreal, 2009.
  13. 13.B. Goodman and S. Flaxman. European union regulations on algorithmic decision-making and a ”right to explanation”. arXiv preprint arXiv:1606.08813, 2016.
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  15. 15.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM international conference on Multimedia, pages 675–678, 2014.
  16. 16.V. Kantorov and I. Laptev. Efficient feature extraction, encoding and classification for action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2593–2600, 2014.
  17. 17.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), pages 1725–1732, 2014.
  18. 18.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. Hmdb: a large video database for human motion recognition. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2556–2563. IEEE, 2011.
  19. 19.W. Landecker, M. D. Thomure, L. M. A. Bettencourt, M. Mitchell, G. T. Kenyon, and S. P. Brumby. Interpreting individual classifications of hierarchical networks. In Proceedings of the IEEE Symposium on Computational Intelligence and Data Mining (CIDM), pages 32–38, 2013.
  20. 20.S. Lapuschkin, A. Binder, G. Montavon, K.-R. Müller, and W. Samek. Analyzing classifiers: Fisher vectors and deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2912–2920, 2016.
  21. 21.S. Lapuschkin, A. Binder, G. Montavon, K.-R. Müller, and W. Samek. The layer-wise relevance propagation toolbox for artificial neural networks. Journal of Machine Learning Research, 17(114):1–5, 2016.
  22. 22.Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller. Efficient backprop. In Neural networks: Tricks of the trade, pages 9–48. Springer, 2012.
  23. 23.Z. C. Lipton. The mythos of model interpretability. arXiv preprint arXiv:1606.03490, 2016.
  24. 24.A. Mahendran and A. Vedaldi. Understanding deep image representations by inverting them. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5188–5196, 2015.
  25. 25.A. Mahendran and A. Vedaldi. Visualizing deep convolutional neural networks using natural pre-images. International Journal of Computer Vision, 120(3):233–255, 2016.
  26. 26.G. Montavon, S. Bach, A. Binder, W. Samek, and K.-R. Müller. Explaining nonlinear classification decisions with deep taylor decomposition. Pattern Recognition, 65:211–222, 2017.
  27. 27.G. Montavon, W. Samek, and K.-R. Müller. Methods for interpreting and understanding deep neural networks. arXiv preprint arXiv:1706.07979, 2017.
  28. 28.M. Moravčík, M. Schmid, N. Burch, V. Lisý, D. Morrill, N. Bard, et al. Deepstack: Expert-level artificial intelligence in heads-up no-limit poker. Science, 356(6337):508–513, 2017.
  29. 29.A. Nguyen, J. Yosinski, and J. Clune. Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks. arXiv preprint arXiv:1602.03616, 2016.
  30. 30.M. T. Ribeiro, S. Singh, and C. Guestrin. Why should i trust you?: Explaining the predictions of any classifier. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144. ACM, 2016.
  31. 31.W. Samek, A. Binder, G. Montavon, S. Lapuschkin, and K.-R. Müller. Evaluating the visualization of what a deep neural network has learned. IEEE Transactions on Neural Networks and Learning Systems, 2017. in press.
  32. 32.K. T. Schütt, F. Arbabzadah, S. Chmiela, K. R. Müller, and A. Tkatchenko. Quantum-chemical insights from deep tensor neural networks. Nature communications, 8:13890, 2017.
  33. 33.A. Shrikumar, P. Greenside, A. Shcherbina, and A. Kundaje. Not just a black box: Learning important features through propagating activation differences. arXiv preprint arXiv:1605.01713, 2016.
  34. 34.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  35. 35.K. Simonyan, A. Vedaldi, and A. Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034, 2013.
  36. 36.V. Srinivasan, S. Lapuschkin, C. Hellge, K.-R. Müller, and W. Samek. Interpretable human action recognition in compressed domain. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1692–1696, 2017.
  37. 37.I. Sturm, S. Lapuschkin, W. Samek, and K.-R. Müller. Interpretable deep neural networks for single-trial eeg classification. Journal of Neuroscience Methods, 274:141–145, 2016.
  38. 38.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1–9, 2015.
  39. 39.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. In European Conference Computer Vision - ECCV 2014, pages 818–833, 2014.
  40. 40.L. M. Zintgraf, T. S. Cohen, T. Adel, and M. Welling. Visualizing deep neural network decisions: Prediction difference analysis. arXiv preprint arXiv:1702.04595, 2017.

Citation

MLA
Samek, W., et al. “Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models”. arXiv, 2017, http://arxiv.org/abs/1708.08296v1.
APA
Samek, W., Wiegand, T., & Müller, K.-R. (2017). Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models. arXiv. http://arxiv.org/abs/1708.08296v1
Chicago
Samek, W., T. Wiegand, and K.-R. Müller. 2017. “Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models”. arXiv. http://arxiv.org/abs/1708.08296v1.
Harvard
Samek, W., Wiegand, T. and Müller, K.-R. (2017) “Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1708.08296v1.
Vancouver
1. Samek W, Wiegand T, Müller K-R (2017) Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models. arXiv

BibTeX

@article{samek2017explainable,
  title = {Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models},
  author = {Samek, Wojciech and Wiegand, Thomas and Müller, Klaus-Robert},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1708.08296v1},
  eprint = {1708.08296}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors