Explainable Deep Learning: A Field Guide for the Uninitiated

Gabrielle RasNing XieMarcel van GervenDerek Doran

article2022JAIR460 citations

Presents a structured taxonomy of foundational explainability techniques, evaluation methods, and user-centered design principles to help researchers and practitioners build and assess interpretable deep learning systems.

Listen

Deep neural networks are increasingly deployed across high-stakes domains, including healthcare diagnostics, criminal justice, financial investments, and autonomous systems. However, these models operate largely as opaque systems, making it difficult for decision-makers and accountable human operators to verify the rationale behind specific recommendations. This lack of transparency poses critical operational, ethical, and legal risks, especially when erroneous or biased predictions can lead to significant real-world harm. The article addresses the urgent need to demystify how these complex models make decisions by categorizing existing explainability techniques, evaluating their reliability, and establishing design principles for real-world deployment.

The main objective of the article is to provide a structured, foundational taxonomy of explainable deep learning techniques, outline methods for evaluating their validity, and guide practitioners in designing user-aligned, trustworthy systems. To achieve this, the article conducts a comprehensive review and technical synthesis across dozens of prominent explanation frameworks, evaluating them across different data modalities—such as images, natural language, and tabular data—and connecting them to adjacent concerns including model robustness, fairness, and debugging.

The findings establish that foundational explainability methods fall into three distinct paradigms. First, visualization methods—such as gradient-based attribution and input perturbation—generate visual heatmaps highlighting influential inputs, though they risk user misinterpretation and are often sensitive to minor input alterations. Second, model distillation methods train transparent surrogate models, such as decision trees or local linear approximations like LIME and SHAP, to replicate the input-output behavior of complex networks either locally or globally. Third, intrinsic methods incorporate transparency directly into the model architecture during training, using attention mechanisms or multi-task joint training to produce explanations alongside predictions. Additionally, the article highlights that popular post-hoc methods often fail standard sanity checks or violate baseline assumptions when removing informative data, and that learned attention weights do not reliably reflect actual feature importance in natural language processing tasks.

These findings indicate that providing an explanation is not a universal safeguard; unvalidated or overly complex explanations can create misplaced confidence, slow down human decision-making, or fail compliance requirements in regulated environments. Effective deployment requires balancing explanatory fidelity with simplicity, depending on whether an application is time-critical (such as vehicle control, where processing speed is vital) or decision-critical (such as medical diagnosis, where deep post-hoc verifiability is mandatory). Furthermore, organizations must recognize that explainability is tightly intertwined with adversarial vulnerability, fairness constraints, and model debugging, requiring a holistic risk management approach rather than isolated explainability tools.

Organizations should adopt a deliberate, user-centered selection process when deploying explainable artificial intelligence. Rather than defaulting to complex networks with visual post-hoc overlays, teams should first evaluate whether simpler, inherently transparent models can achieve the desired performance. When deep architectures are required, practitioners should match the method to the exact data type and user requirements, prioritizing intrinsic designs or mathematically grounded frameworks like Shapley values where robust guarantees are necessary. Further work is required across the field to establish standardized, objective benchmarks and develop computationally efficient explanations suitable for real-time operations.

The conclusions of the article must be viewed in light of several ongoing challenges in the field, primarily the absence of an overarching theoretical definition of what constitutes a sufficient explanation and the lack of objective ground-truth data for evaluating explanation quality. While the taxonomy provides high confidence for structuring and selecting technical approaches, readers should exercise caution when interpreting post-hoc visualizations or attention weights without rigorous, task-specific validation.

arXiv: 2004.14545
Cover for Explainable Deep Learning: A Field Guide for the Uninitiated

Abstract

Deep neural networks (DNNs) are an indispensable machine learning tool despite the difficulty of diagnosing what aspects of a model’s input drive its decisions. In countless real-world domains, from legislation and law enforcement to healthcare, such diagnosis is essential to ensure that DNN decisions are driven by aspects appropriate in the context of its use. The development of methods and studies enabling the explanation of a DNN’s decisions has thus blossomed into an active and broad area of research. The field’s complexity is exacerbated by competing definitions of what it means “to explain” the actions of a DNN and to evaluate an approach’s “ability to explain”. This article offers a field guide to explore the space of explainable deep learning for those in the AI/ML field who are uninitiated. The field guide: i) Introduces three simple dimensions defining the space of foundational methods that contribute to explainable deep learning, ii) discusses the evaluations for model explanations, iii) places explainability in the context of other related deep learning research areas, and iv) discusses user-oriented explanation design and future directions. We hope the guide is seen as a starting point for those embarking on this research field.

Table of Contents

  • 1. Introduction
  • 1.1 A Word of Caution
  • 1.2 But What is an Explanation?
  • 2. Methods for Explaining DNNs
  • 2.1 Visualization Methods
  • 2.1.1 Backpropagation-based Methods
  • 2.1.2 Perturbation-based Methods
  • 2.2 Model Distillation
  • 2.2.1 Local Approximation
  • 2.2.2 Model Translation
  • 2.3 Intrinsic Methods
  • 2.3.1 Attention Mechanisms
  • 2.3.2 Joint Training
  • 2.4 Summary
  • 3. Evaluating Explanations
  • 3.1 What Makes a Good Explanation?
  • 3.2 Methods for Evaluating Explanation Methods and their Explanations
  • 3.2.1 Evaluating Heatmaps
  • 3.2.2 Evaluating NLP Explanations
  • 3.2.3 Using Humans to Evaluate Explanations
  • 4. Topics Associated with Explainability
  • 4.1 Learning Mechanisms
  • 4.2 Model Debugging
  • 4.3 Adversarial Attack and Defense
  • 4.4 Fairness and Bias
  • 5. Designing Explanations for Users
  • 5.1 Understanding the End User
  • 5.2 The Impact of DNN Decisions
  • 5.3 Design Extendability
  • 6. Future Directions
  • 7. Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Three-Dimensional Taxonomy of Explainable Deep Learning Methods

    definition

    Explainable deep learning methods are systematically organized into a three-dimensional framework based on their underlying operational mechanism:

    1. Visualization Methods: Post-hoc techniques that generate visual representations (such as saliency maps or heatmaps) to highlight input features that strongly influence a deep neural network's (DNN's) output. These comprise:
      • Backpropagation-based methods: Propagate gradient or relevance signals backward from the network output to the input layer.
      • Perturbation-based methods: Alter, mask, or remove portions of the input data and measure the resulting variation in network predictions.
    2. Model Distillation: Post-hoc approaches that train an inherently transparent, "white-box" surrogate model gg to mimic the input-output mapping of an opaque network ff (g(x)≈f(x)g(x) \approx f(x)). These comprise:
      • Local approximation: Fits an explainable model within the local neighborhood of a single input instance.
      • Model translation: Distills the global decision behavior across an entire dataset into explainable structures such as decision trees, finite state automata, or causal graphs.
    3. Intrinsic Methods: Architectures designed to incorporate explainability directly into the forward pass or optimization objective. These comprise:
      • Attention mechanisms: Learn dynamic weightings across input elements to produce weighted context vectors that can be directly visualized.
      • Joint training: Optimizes a multi-task objective that produces primary task predictions alongside explanations (such as generated text justifications, concept associations, or learned case prototypes).
  2. Knowl 2 — Gradient-Weighted Class Activation Mapping

    model/method

    Gradient-weighted Class Activation Mapping (Grad-CAM) produces visual localization heatmaps for any convolutional neural network with differentiable output layers.

    For a target class cc with unnormalized class score ycy_c (the logit prior to softmax), the importance weight αk,c\alpha_{k,c} of the kk-th feature map AkA_k in the final convolutional layer (having spatial dimensions m×nm \times n) is calculated by global average pooling the gradients:

    αk,c=1m⋅n∑i=1m∑j=1n∂yc∂Ak,i,j\alpha_{k,c} = \frac{1}{m \cdot n} \sum_{i=1}^m \sum_{j=1}^n \frac{\partial y_c}{\partial A_{k,i,j}}

    where Ak,i,jA_{k,i,j} is the activation at spatial coordinate (i,j)(i, j) in feature map AkA_k.

    The 2D class-discriminative localization map mapc\text{map}_c is computed as a weighted linear combination of feature maps followed by a Rectified Linear Unit (ReLU\text{ReLU}) operation:

    mapc=ReLU(∑kαk,cAk)\text{map}_c = \text{ReLU}\left( \sum_k \alpha_{k,c} A_k \right)

    The ReLU\text{ReLU} ensures that only features having a positive impact on score ycy_c are highlighted. The resulting score map is upsampled via bilinear interpolation to match the dimensions of the original input image.

  3. Knowl 3 — Layer-Wise Relevance Propagation via Deep Taylor Decomposition

    model/method

    Layer-Wise Relevance Propagation (LRP) decomposes the scalar prediction output f(x)f(x) of a neural network across NN input features x=(x1,x2,…,xN)x = (x_1, x_2, \dots, x_N) such that f(x)=∑i=1Nrif(x) = \sum_{i=1}^N r_i, where rir_i represents the relevance score of feature xix_i.

    In Deep Taylor Decomposition, the network function ff is approximated by a first-order Taylor expansion around a reference root point x^\hat{x} satisfying f(x^)=0f(\hat{x}) = 0:

    f(x)=f(x^)+∇x^f⋅(x−x^)+ϵ=∑i=1N∂f∂xi(x^)⋅(xi−x^i)+ϵf(x) = f(\hat{x}) + \nabla_{\hat{x}} f \cdot (x - \hat{x}) + \epsilon = \sum_{i=1}^N \frac{\partial f}{\partial x_i}(\hat{x}) \cdot (x_i - \hat{x}_i) + \epsilon

    where ϵ\epsilon represents higher-order terms. The input feature relevance is defined as:

    ri=∂f∂xi(x^)⋅(xi−x^i)r_i = \frac{\partial f}{\partial x_i}(\hat{x}) \cdot (x_i - \hat{x}_i)

    To compute relevance across hidden layers, relevance scores are conserved layer-by-layer. The relevance score rilr_i^l of node ii in layer ll is obtained by decomposing the relevance scores rjl+1r_j^{l+1} of all MM connected nodes jj in layer l+1l+1:

    ril=∑j=1Mri,jlr_i^l = \sum_{j=1}^M r_{i,j}^l

    This backpropagation rule preserves total relevance across layers and yields pixel-level heatmaps depicting both positive and negative feature attributions.

  4. Knowl 4 — Deep Learning Important FeaTures

    model/method

    Deep Learning Important FeaTures (DeepLIFT) calculates feature attributions by comparing a neural network's activations on an input x=(x1,x2,…,xN)x = (x_1, x_2, \dots, x_N) against its activations on a user-defined reference baseline x′=(x1′,x2′,…,xN′)x' = (x'_1, x'_2, \dots, x'_N).

    Let Δt=f(x)−f(x′)\Delta t = f(x) - f(x') denote the difference in the target neuron's output tt between xx and x′x', and let Δxi=xi−xi′\Delta x_i = x_i - x'_i denote the feature difference. DeepLIFT attributes Δt\Delta t to individual input differences RΔxiΔtR_{\Delta x_i \Delta t} such that:

    Δt=∑i=1NRΔxiΔt\Delta t = \sum_{i=1}^N R_{\Delta x_i \Delta t}

    A multiplier mΔxΔtm_{\Delta x \Delta t} is defined as:

    mΔxΔt=RΔxΔtΔxm_{\Delta x \Delta t} = \frac{R_{\Delta x \Delta t}}{\Delta x}

    For an intermediate hidden layer ll with activations al=(a1l,a2l,…,aKl)a^l = (a_1^l, a_2^l, \dots, a_K^l) positioned between input nodes xx and target neuron tt, DeepLIFT applies a discrete chain rule for multipliers:

    mΔxiΔt=∑j=1KmΔxiΔajlmΔajlΔtm_{\Delta x_i \Delta t} = \sum_{j=1}^K m_{\Delta x_i \Delta a_j^l} m_{\Delta a_j^l \Delta t}

    This enables recursive backpropagation of relevance multipliers back to the input features, producing attributions that avoid gradient saturation issues.

  5. Knowl 5 — Integrated Gradients Attribution

    model/method

    Integrated Gradients computes input feature attributions for a differentiable neural network f:RN→[0,1]f: \mathbb{R}^N \to [0, 1] relative to a baseline input x′x' along a straight-line interpolation path. The method formally satisfies two key axioms:

    1. Sensitivity: If an input xx and a baseline x′x' differ along feature xix_i and yield different outputs f(x)≠f(x′)f(x) \neq f(x'), feature xix_i receives a non-zero attribution.
    2. Implementation Invariance: Two networks that produce identical outputs for all inputs assign identical attributions to all features.

    The attribution for feature xix_i of input xx with respect to baseline x′x' is defined by the line integral:

    IGi(x)=(xi−xi′)∫01∂f(x′+α(x−x′))∂xidαIG_i(x) = (x_i - x'_i) \int_0^1 \frac{\partial f(x' + \alpha(x - x'))}{\partial x_i} d\alpha

    where α∈[0,1]\alpha \in [0, 1]. In practical implementations, the integral is computed via Riemann summation across MM interpolation steps (typically M∈[20,300]M \in [20, 300]):

    IGi(x)≈(xi−xi′)1M∑k=1M∂f(x′+kM(x−x′))∂xiIG_i(x) \approx (x_i - x'_i) \frac{1}{M} \sum_{k=1}^M \frac{\partial f\left(x' + \frac{k}{M}(x - x')\right)}{\partial x_i}

    Common baselines include an all-zero black image for vision tasks or a zero embedding vector for text tasks.

  6. Knowl 6 — Local Interpretable Model-Agnostic Explanations

    model/method

    Local Interpretable Model-Agnostic Explanations (LIME) explains individual predictions of a black-box model f:Rd→Rf: \mathbb{R}^d \to \mathbb{R} by constructing an interpretable surrogate model g∈Gg \in G (such as a sparse linear regressor or decision tree) defined over an interpretable representation space {0,1}d′\{0, 1\}^{d'}.

    For an input instance x∈Rdx \in \mathbb{R}^d corresponding to interpretable representation x′∈{0,1}d′x' \in \{0, 1\}^{d'}, LIME optimizes:

    arg⁡min⁡g∈G{L(f,g,Πx)+Ω(g)}\arg\min_{g \in G} \left\{ \mathcal{L}(f, g, \Pi_x) + \Omega(g) \right\}

    where:

    • Ω(g)\Omega(g) is a regularizer penalizing the complexity of the explanation model gg (e.g., tree depth or number of active linear features).
    • Πx(z)\Pi_x(z) is an exponential similarity kernel measuring the proximity between perturbed instance z∈Rdz \in \mathbb{R}^d and original instance xx.
    • L(f,g,Πx)\mathcal{L}(f, g, \Pi_x) is the fidelity loss measuring the difference between the black-box model's prediction f(z)f(z) and the surrogate's prediction g(z′)g(z') across sampled perturbations z′∈{0,1}d′z' \in \{0, 1\}^{d'}, weighted by proximity Πx(z)\Pi_x(z):

    L(f,g,Πx)=∑z,z′∈ZΠx(z)(f(z)−g(z′))2\mathcal{L}(f, g, \Pi_x) = \sum_{z, z' \in \mathcal{Z}} \Pi_x(z) \left( f(z) - g(z') \right)^2

    The fitted parameters of gg provide local feature attribution scores for xx without requiring access to internal network weights or gradients.

  7. Knowl 7 — Shapley Additive Explanations

    model/method

    Shapley Additive Explanations (SHAP) computes feature attribution values grounded in cooperative game theory. It calculates Shapley values ϕi\phi_i to quantify the additive contribution of each input feature to a model's prediction.

    For a model ff and a total set of features FF, the Shapley value ϕi\phi_i of feature ii is:

    ϕi=∑S⊆F∖{i}∣S∣!(∣F∣−∣S∣−1)!∣F∣![fS∪{i}(xS∪{i})−fS(xS)]\phi_i = \sum_{S \subseteq F \setminus \{i\}} \frac{|S|!(|F| - |S| - 1)!}{|F|!} \left[ f_{S \cup \{i\}}(x_{S \cup \{i\}}) - f_S(x_S) \right]

    where SS is a subset of features excluding ii, fS∪{i}f_{S \cup \{i\}} is the model trained or evaluated with feature ii included, and fSf_S is the model evaluated without feature ii.

    SHAP represents the explanation as an additive linear surrogate model g(z′)g(z') over simplified binary features z′∈{0,1}Mz' \in \{0, 1\}^M:

    g(z′)=ϕ0+∑i=1Mϕizi′g(z') = \phi_0 + \sum_{i=1}^M \phi_i z'_i

    where ϕ0\phi_0 is the base output when all features are absent. In KernelSHAP, marginal feature sampling is combined with a game-theoretic kernel to estimate ϕi\phi_i efficiently.

  8. Knowl 8 — Prediction Difference Analysis

    model/method

    Prediction Difference Analysis is a perturbation-based attribution method that evaluates the effect of an input feature xix_i on a classifier's predicted probability for class cc by estimating the model output when xix_i is removed.

    Because removing features directly can distort deep network inputs, the marginal class probability p(c∣x−i)p(c \mid x_{-i}) without feature xix_i is computed by marginalizing over all possible values of xix_i conditional on the remaining features x−ix_{-i}:

    p(c∣x−i)=∑xip(xi∣x−i)p(c∣x−i,xi)p(c \mid x_{-i}) = \sum_{x_i} p(x_i \mid x_{-i}) p(c \mid x_{-i}, x_i)

    The relevance score (prediction difference) Diffi(c∣x)\text{Diff}_i(c \mid x) is defined in terms of log-odds:

    Diffi(c∣x)=log⁡2(odds(c∣x))−log⁡2(odds(c∣x−i))\text{Diff}_i(c \mid x) = \log_2(\text{odds}(c \mid x)) - \log_2(\text{odds}(c \mid x_{-i}))

    where odds(c∣x)=p(c∣x)1−p(c∣x)\text{odds}(c \mid x) = \frac{p(c \mid x)}{1 - p(c \mid x)}.

    The magnitude of Diffi(c∣x)\text{Diff}_i(c \mid x) reflects the strength of feature importance, while a positive sign indicates evidence in favor of class cc and a negative sign indicates evidence against class cc.

  9. Knowl 9 — Multi-Task Joint Training for Intrinsic Explainability

    model/method

    Intrinsically explainable deep neural networks can be designed by coupling the main prediction architecture with an explanation generation component and training both jointly.

    Given NN training examples with task labels yny_n and ground-truth explanation annotations ene_n, the joint training objective optimizes network parameters θ\theta via:

    arg⁡min⁡θ1N∑n=1N[αLpred(yn,yn′)+βLexpl(en,en′)]\arg\min_\theta \frac{1}{N} \sum_{n=1}^N \left[ \alpha \mathcal{L}_{\text{pred}}(y_n, y'_n) + \beta \mathcal{L}_{\text{expl}}(e_n, e'_n) \right]

    where:

    • yn′y'_n is the primary task prediction and Lpred\mathcal{L}_{\text{pred}} is the task loss function.
    • en′e'_n is the generated explanation (such as natural language text, semantic concept labels, or learned prototype parts) and Lexpl\mathcal{L}_{\text{expl}} is the explanation generation loss.
    • α\alpha and β\beta are positive scalar weights balancing task accuracy against explanation quality.

    Joint training forces internal feature representations to align directly with human-interpretable rationale concepts during optimization rather than attempting to extract rationales post-hoc.

  10. Knowl 10 — Objective Evaluation Benchmarks and Sanity Checks for Attributions

    model/method

    Objective, automated evaluation of feature attribution heatmaps is performed using perturbation metrics and randomization sanity checks:

    1. Feature Removal / Degradation: Pixels indicated as salient by a heatmap are replaced with uniform random noise or channel means; a steeper reduction in classification score indicates higher attribution fidelity.
    2. Remove And Retrain (ROAR): Eliminates distribution shift artifacts in evaluation by removing the top k%k\% salient pixels (replacing them with channel means) across both the training and test sets, followed by retraining a new network from scratch. A degradation in retrained model accuracy verifies that the removed features were genuinely informative for the task.
    3. c-Eval Metric: Measures classifier robustness when perturbing pixels assigned the lowest importance scores. Higher robustness cc indicates that non-salient features were correctly identified.
    4. Model Parameter Randomization Test: Computes heatmap similarity between a trained model and a randomly initialized model. If the heatmaps remain similar, the attribution method is insensitive to model parameters (acting primarily as an edge detector).
    5. Data Randomization Test: Evaluates heatmaps from a model trained on true data versus a model trained on randomly permuted labels. A valid explanation method must show substantial differences when label-data associations are destroyed.
  11. Knowl 11 — Human-in-the-Loop Evaluation Metrics for Explanations

    model/method

    Evaluating explanations with human subjects assesses their practical intelligibility and operational utility:

    • Simulatability: Measures the accuracy with which a human evaluator can predict a trained model's output on unseen data after reviewing past inputs and their corresponding explanations. While local surrogate methods (such as LIME) can increase simulatability on tabular datasets, subjective user ratings of explanation clarity do not reliably predict actual simulatability improvements.
    • Model-Human Alignment: Quantifies the similarity between machine-generated explanation rationales and human-annotated rationales (e.g., in Natural Language Inference tasks). Empirical results indicate that higher model accuracy or larger model parameter counts do not guarantee superior alignment with human reasoning.
    • Decision Efficacy Trade-Offs: Providing explanations during collaborative human-AI decision-making tasks significantly decreases the time required for a user to reach a decision, but does not always produce a statistically significant increase in overall decision accuracy compared to presenting the raw data alone.
  12. Knowl 12 — User-Oriented System Design Dimensions for Explainable DNNs

    definition

    Deploying explainable deep learning systems in practice requires balancing three user- and application-oriented design dimensions:

    1. End-User Expertise:
      • Deep Learning Experts: Require low-level, technical diagnostic details (such as layer-wise gradient volume, intermediate activation maps, and sensitivity vectors) to support model debugging, error diagnosis, and architecture refinement.
      • Non-Expert End Users: Require high-level, human-aligned abstractions (such as natural language rationales, contrastive examples, or semantic concepts) that do not require knowledge of neural network mechanics.
    2. Operational Impact and Criticality:
      • Time-Critical Systems (e.g., autonomous vehicle collision avoidance, active defense): Demand lightweight, low-latency explanations that users can comprehend instantly under tight reaction constraints.
      • Decision-Critical Systems (e.g., medical diagnostics, financial underwriting): Demand deeply inspectable, highly faithful, auditable evidence to justify life-altering actions and perform post-hoc failure analysis.
    3. Design Extendability:
      • Modularity: The degree to which an explanation module can be decoupled from and integrated across different underlying neural architectures (inherent to model-agnostic distillation techniques).
      • Reusability: The ability of pre-trained models or standardized explanatory representations to transfer reliably across disparate problem domains.

Coverage note — Specific implementations of peripheral deep learning topics (such as individual adversarial attack variants and dataset-specific debiasing algorithms) were omitted as they represent broader external research areas rather than core explainability methods introduced in the survey.

References

  1. 1.Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., & Süsstrunk, S. (2012). SLIC superpixels compared to state-of-the-art superpixel methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34 (11), 2274–2282.
  2. 2.Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160.
  3. 3.Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., & Kim, B. (2018). Sanity checks for saliency maps. In Advances in Neural Information Processing Systems, pp. 9525–9536.
  4. 4.Adebayo, J., & Kagal, L. (2016). Iterative orthogonal feature projection for diagnosing bias in black-box models. ArXiv, abs/1611.04967.
  5. 5.Adler, P., Falk, C., Friedler, S. A., Nix, T., Rybeck, G., Scheidegger, C., Smith, B., & Venkatasubramanian, S. (2018). Auditing black-box models for indirect influence. Knowledge and Information Systems, 54 (1), 95–122.
  6. 6.Agarwal, A., Beygelzimer, A., Dudik, M., Langford, J., & Wallach, H. (2018). A reductions approach to fair classification. In International Conference on Machine Learning, pp. 60–69.
  7. 7.Ahmad, M. A., Eckert, C., & Teredesai, A. (2018). Interpretable machine learning in healthcare. In Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, pp. 559–560.
  8. 8.Alain, G., & Bengio, Y. (2017). Understanding intermediate layers using linear classifier probes. In International Conference on Learning Representations.
  9. 9.Alvarez-Melis, D., & Jaakkola, T. S. (2018). On the robustness of interpretability methods. ArXiv, abs/1806.08049.
  10. 10.Amershi, S., Chickering, M., Drucker, S. M., Lee, B., Simard, P., & Suh, J. (2015). Modeltracker: Redesigning performance analysis tools for machine learning. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, pp. 337–346.
  11. 11.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6077–6086.
  12. 12.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., & Parikh, D. (2015). VQA: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2425–2433.
  13. 13.Arpit, D., Jastrzeski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al. (2017). A closer look at memorization in deep networks. In International Conference on Machine Learning, pp. 233–242.
  14. 14.Arras, L., Horn, F., Montavon, G., Müller, K.-R., & Samek, W. (2016). Explaining predictions of non-Linear classifiers in NLP. In Proceedings of the 1st Workshop on Representation Learning for NLP, pp. 1–7.
  15. 15.Arras, L., Horn, F., Montavon, G., Müller, K.-R., & Samek, W. (2017). "What is relevant in a text document?": An interpretable machine learning approach. PLOS One, 12.
  16. 16.Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., et al. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.
  17. 17.Baan, J., ter Hoeve, M., van der Wees, M., Schuth, A., & de Rijke, M. (2019). Do transformer attention heads provide transparency in abstractive summarization?. ArXiv, abs/1907.00570.
  18. 18.Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., & Samek, W. (2015). On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLOS One, 10.
  19. 19.Baehrens, D., Schroeter, T., Harmeling, S., Kawanabe, M., Hansen, K., & Müller, K.-R. (2010). How to explain individual classification decisions. Journal of Machine Learning Research, 11, 1803–1831.
  20. 20.Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations.
  21. 21.Bahri, Y., Kadmon, J., Pennington, J., Schoenholz, S. S., Sohl-Dickstein, J., & Ganguli, S. (2020). Statistical mechanics of deep learning. Annual Review of Condensed Matter Physics, 11, 501–528.
  22. 22.Bastani, O., Kim, C., & Bastani, H. (2017). Interpreting blackbox models via model extraction. ArXiv, abs/1705.08504.
  23. 23.Bechavod, Y., & Ligett, K. (2017). Penalizing unfairness in binary classification. ArXiv, abs/1707.00044.
  24. 24.Berk, R., Heidari, H., Jabbari, S., Joseph, M., Kearns, M., Morgenstern, J., Neel, S., & Roth, A. (2017). A convex framework for fair regression. ArXiv, abs/1706.02409.
  25. 25.Beutel, A., Chen, J., Zhao, Z., & Chi, E. H. (2017). Data decisions and theoretical implications when adversarially learning fair representations. ArXiv, abs/1707.00075.
  26. 26.Bien, J., & Tibshirani, R. (2011). Prototype selection for interpretable classification. The Annals of Applied Statistics, 5, 2403–2424.
  27. 27.Brendel, W., Rauber, J., & Bethge, M. (2018). Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. In International Conference on Learning Representations.
  28. 28.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems, pp. 1877–1901.
  29. 29.Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on Fairness, Accountability and Transparency, pp. 77–91.
  30. 30.Calders, T., Kamiran, F., & Pechenizkiy, M. (2009). Building classifiers with independency constraints. In 2009 IEEE International Conference on Data Mining Workshops, pp. 13–18.
  31. 31.Calmon, F., Wei, D., Vinzamuri, B., Ramamurthy, K. N., & Varshney, K. R. (2017). Optimized pre-processing for discrimination prevention. In Advances in Neural Information Processing Systems, pp. 3992–4001.
  32. 32.Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., & Blunsom, P. (2018). e-SNLI: Natural language inference with natural language explanations. In Advances in Neural Information Processing Systems, pp. 9560–9572.
  33. 33.Carlini, N., Katz, G., Barrett, C., & Dill, D. L. (2017). Ground-truth adversarial examples. ArXiv, abs/1709.10207.
  34. 34.Carlini, N., & Wagner, D. (2017). Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy, pp. 39–57.
  35. 35.Carter, S., Armstrong, Z., Schubert, L., Johnson, I., & Olah, C. (2019). Activation atlas. Distill, 4 (3).
  36. 36.Carvalho, D. V., Pereira, E. M., & Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics, 8 (8), 832.
  37. 37.Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., & Su, J. K. (2019). This looks like that: Deep learning for interpretable image recognition. In Advances in Neural Information Processing Systems, pp. 8928–8939.
  38. 38.Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., & Hsieh, C.-J. (2017). ZOO: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, pp. 15–26.
  39. 39.Chong, E., Han, C., & Park, F. C. (2017). Deep learning networks for stock market analysis and prediction: Methodology, data representations, and case studies. Expert Systems with Applications, 83, 187–205.
  40. 40.Craven, M., & Shavlik, J. W. (1996). Extracting tree-structured representations of trained networks. In Advances in Neural Information Processing Systems, pp. 24–30.
  41. 41.Dabkowski, P., & Gal, Y. (2017). Real time image saliency for black box classifiers. In Advances in Neural Information Processing Systems, pp. 6967–6976.
  42. 42.Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4171–4186.
  43. 43.DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., & Wallace, B. C. (2020). ERASER: A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4443–4458.
  44. 44.Díez, J., Khalifa, K., & Leuridan, B. (2013). General theories of explanation: Buyer beware. Synthese, 190 (3), 379–396.
  45. 45.Ding, Y., Liu, Y., Luan, H., & Sun, M. (2017). Visualizing and understanding neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vol. 1, pp. 1150–1159.
  46. 46.Dong, W., Li, J., Yao, R., Li, C., Yuan, T., & Wang, L. (2016). Characterizing driving styles with deep learning. ArXiv, abs/1607.03611.
  47. 47.Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., & Li, J. (2018). Boosting adversarial attacks with momentum. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 9185–9193.
  48. 48.Dong, Y., Su, H., Zhu, J., & Zhang, B. (2017). Improving interpretability of deep neural networks with semantic information. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4306–4314.
  49. 49.Donini, M., Oneto, L., Ben-David, S., Shawe-Taylor, J. S., & Pontil, M. (2018). Empirical risk minimization under fairness constraints. In Advances in Neural Information Processing Systems, pp. 2791–2801.
  50. 50.Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning. ArXiv, abs/1702.08608.
  51. 51.Doshi-Velez, F., & Kim, B. (2018). Considerations for evaluation and generalization in interpretable machine learning. In Explainable and Interpretable Models in Computer Vision and Machine Learning, pp. 3–17. Springer.
  52. 52.Došilović, F. K., Brćić, M., & Hlupić, N. (2018). Explainable artificial intelligence: A survey. In 2018 41st International Convention on Information and Communication Technology, Electronics and Microelectronics, pp. 0210–0215.
  53. 53.Dwork, C., Hardt, M., Pitassi, T., Reingold, O., & Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference, pp. 214–226.
  54. 54.Dwork, C., Immorlica, N., Kalai, A. T., & Leiserson, M. (2018). Decoupled classifiers for group-fair and efficient machine learning. In Conference on Fairness, Accountability and Transparency, pp. 119–133.
  55. 55.Elenberg, E., Dimakis, A. G., Feldman, M., & Karbasi, A. (2017). Streaming weak submodularity: Interpreting neural networks on the fly. In Advances in Neural Information Processing Systems, pp. 4044–4054.
  56. 56.Erhan, D., Bengio, Y., Courville, A., & Vincent, P. (2009). Visualizing higher-layer features of a deep network. Départment d’Informatique et Recherche Opérationnelle, University of Montreal, QC, Canada, Tech. Rep, 1341.
  57. 57.Erhan, D., Courville, A., & Bengio, Y. (2010). Understanding representations learned in deep architectures. Départment d’Informatique et Recherche Opérationnelle, University of Montreal, QC, Canada, Tech. Rep, 1355.
  58. 58.Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542 (7639), 115–118.
  59. 59.Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., & Song, D. (2018). Robust physical-world attacks on deep learning visual classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1625–1634.
  60. 60.Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., & Venkatasubramanian, S. (2015). Certifying and removing disparate impact. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 259–268.
  61. 61.Fong, R. C., & Vedaldi, A. (2017). Interpretable explanations of black boxes by meaningful perturbation. In Proceedings of the IEEE International Conference on Computer Vision, pp. 3429–3437.
  62. 62.Frosst, N., & Hinton, G. (2017). Distilling a neural network into a soft decision tree. ArXiv, abs/1711.09784.
  63. 63.Fuchs, F. B., Groth, O., Kosoriek, A. R., Bewley, A., Wulfmeier, M., Vedaldi, A., & Posner, I. (2018). Neural stethoscopes: Unifying analytic, auxiliary and adversarial network probing. ArXiv, abs/1806.05502.
  64. 64.Garcez, A. S. d., Broda, K. B., & Gabbay, D. M. (2012). Neural-symbolic learning systems: Foundations and applications. Springer Science & Business Media.
  65. 65.Garvie, C. (2016). The perpetual line-up: Unregulated police face recognition in America. Georgetown Law, Center on Privacy & Technology.
  66. 66.Georgiev, P., Bhattacharya, S., Lane, N. D., & Mascolo, C. (2017). Low-resource multi-task audio sensing for mobile and embedded devices via shared deep neural network representations. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 1 (3), 50.
  67. 67.Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., & Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on Data Science and Advanced Analytics, pp. 80–89.
  68. 68.Gonzalez-Garcia, A., Modolo, D., & Ferrari, V. (2018). Do semantic parts emerge in convolutional neural networks?. International Journal of Computer Vision, 126 (5), 476–494.
  69. 69.Goodfellow, I. J., Shlens, J., & Szegedy, C. (2014). Explaining and harnessing adversarial examples. In International Conference on Learning Representations.
  70. 70.Goodman, B., & Flaxman, S. (2017). European Union regulations on algorithmic decision-making and a "right to explanation". AI Magazine, 38 (3), 50–57.
  71. 71.Gordaliza, P., Del Barrio, E., Fabrice, G., & Loubes, J.-M. (2019). Obtaining fairness using optimal transport theory. In International Conference on Machine Learning, pp. 2357–2365.
  72. 72.Goswami, G., Bhardwaj, R., Singh, R., & Vatsa, M. (2014). MDLFace: Memorability augmented deep learning for video face recognition. In IEEE International Joint Conference on Biometrics, pp. 1–7.
  73. 73.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., & Parikh, D. (2017). Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6904–6913.
  74. 74.Graves, A., Wayne, G., & Danihelka, I. (2014). Neural turing machines. ArXiv, abs/1410.5401.
  75. 75.Grigorescu, S., Trasnea, B., Cocias, T., & Macesanu, G. (2020). A survey of deep learning techniques for autonomous driving. Journal of Field Robotics, 37 (3), 362–386.
  76. 76.Groth, O., Fuchs, F. B., Posner, I., & Vedaldi, A. (2018). ShapeStacks: Learning vision-based physical intuition for generalised object stacking. In Proceedings of the European Conference on Computer Vision, pp. 702–717.
  77. 77.Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., & Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys, 51 (5), 93.
  78. 78.Hardt, M., Price, E., Srebro, N., et al. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems, pp. 3315–3323.
  79. 79.Harradon, M., Druce, J., & Ruttenberg, B. (2018). Causal learning and explanation of deep neural networks via autoencoded activations. ArXiv, abs/1802.00541.
  80. 80.Hase, P., & Bansal, M. (2020). Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5540–5552.
  81. 81.He, R., Lee, W. S., Ng, H. T., & Dahlmeier, D. (2018). Effective attention modeling for aspect-level sentiment classification. In Proceedings of the 27th International Conference on Computational Linguistics, pp. 1121–1131.
  82. 82.Heidari, H., Ferrari, C., Gummadi, K., & Krause, A. (2018). Fairness behind a veil of ignorance: A welfare analysis for automated decision making. In Advances in Neural Information Processing Systems, pp. 1265–1276.
  83. 83.Hendricks, L. A., Akata, Z., Rohrbach, M., Donahue, J., Schiele, B., & Darrell, T. (2016). Generating visual explanations. In European Conference on Computer Vision, pp. 3–19.
  84. 84.Heo, J., Lee, H. B., Kim, S., Lee, J., Kim, K. J., Yang, E., & Hwang, S. J. (2018). Uncertainty-aware attention for reliable interpretation and prediction. In Advances in Neural Information Processing Systems, pp. 909–918.
  85. 85.Heskes, T., Sijben, E., Bucur, I. G., & Claassen, T. (2020). Causal Shapley values: Exploiting causal knowledge to explain individual predictions of complex models. In Advances in Neural Information Processing Systems, pp. 4778–4789.
  86. 86.Hind, M., Wei, D., Campbell, M., Codella, N. C., Dhurandhar, A., Mojsilović, A., Natesan Ramamurthy, K., & Varshney, K. R. (2019). TED: Teaching AI to explain its decisions. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pp. 123–129.
  87. 87.Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. ArXiv, abs/1503.02531.
  88. 88.Hooker, G. (2004). Discovering additive structure in black box functions. In Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 575–580.
  89. 89.Hooker, G. (2007). Generalized functional anova diagnostics for high-dimensional functions of dependent variables. Journal of Computational and Graphical Statistics, 16 (3), 709–732.
  90. 90.Hooker, S., Erhan, D., Kindermans, P.-J., & Kim, B. (2019). A benchmark for interpretability methods in deep neural networks. In Advances in Neural Information Processing Systems, pp. 9734–9745.
  91. 91.Hoos, H., & Leyton-Brown, K. (2014). An efficient approach for assessing hyperparameter importance. In International Conference on Machine Learning, pp. 754–762.
  92. 92.Hou, B.-J., & Zhou, Z.-H. (2020). Learning with interpretable structure from gated RNN. IEEE Transactions on Neural Networks and Learning Systems, 31 (7), 2267–2279.
  93. 93.Iyer, R., Li, Y., Li, H., Lewis, M., Sundar, R., & Sycara, K. (2018). Transparency and explanation in deep reinforcement learning neural networks. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pp. 144–150.
  94. 94.Jacovi, A., & Goldberg, Y. (2020). Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4198–4205.
  95. 95.Jain, A., Koppula, H. S., Raghavan, B., Soh, S., & Saxena, A. (2015). Car that knows before you do: Anticipating maneuvers via learning temporal driving models. In Proceedings of the IEEE International Conference on Computer Vision, pp. 3182–3190.
  96. 96.Jain, S., & Wallace, B. C. (2019). Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3543–3556.
  97. 97.Jesus, S., Belém, C., Balayan, V., Bento, J., Saleiro, P., Bizarro, P., & Gama, J. (2021). How can I choose an explainer? An application-grounded evaluation of post-hoc explanations. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, p. 805–815.
  98. 98.Jha, S., Raman, V., Pinto, A., Sahai, T., & Francis, M. (2017). On learning sparse Boolean formulae for explaining AI decisions. In NASA Formal Methods Symposium, pp. 99–114.
  99. 99.Jiang, H., Kim, B., Guan, M., & Gupta, M. (2018). To trust or not to trust a classifier. In Advances in Neural Information Processing Systems, pp. 5541–5552.
  100. 100.Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., & Girshick, R. (2017). CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2901–2910.
  101. 101.Kamiran, F., & Calders, T. (2010). Classification with no discrimination by preferential sampling. In Proc. 19th Machine Learning Conf. Belgium and The Netherlands, pp. 1–6.
  102. 102.Kamiran, F., & Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 33 (1), 1–33.
  103. 103.Kamishima, T., Akaho, S., & Sakuma, J. (2011). Fairness-aware learning through regularization approach. In 2011 IEEE 11th International Conference on Data Mining Workshops, pp. 643–650.
  104. 104.Kang, D., Raghavan, D., Bailis, P., & Zaharia, M. (2018). Model assertions for debugging machine learning. In NeurIPS MLSys Workshop.
  105. 105.Kearns, M., Neel, S., Roth, A., & Wu, Z. S. (2018). Preventing fairness gerrymandering: Auditing and learning for subgroup fairness. In International Conference on Machine Learning, pp. 2569–2577.
  106. 106.Kim, B., Rudin, C., & Shah, J. A. (2014). The Bayesian case model: A generative approach for case-based reasoning and prototype classification. In Advances in Neural Information Processing Systems, pp. 1952–1960.
  107. 107.Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. (2018a). Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In International Conference on Machine Learning, pp. 2673–2682.
  108. 108.Kim, J., Rohrbach, A., Darrell, T., Canny, J., & Akata, Z. (2018b). Textual explanations for self-driving vehicles. In Proceedings of the European Conference on Computer Vision, pp. 563–578.
  109. 109.Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., & Kim, B. (2019). The (un)reliability of saliency methods. In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, pp. 267–280. Springer.
  110. 110.Kindermans, P.-J., Schütt, K., Müller, K.-R., & Dähne, S. (2016). Investigating the influence of noise and distractors on the interpretation of neural networks. ArXiv, abs/1611.07270.
  111. 111.Kindermans, P.-J., Schütt, K., Alber, M., Müller, K.-R., Erhan, D., Kim, B., & Dähne, S. (2018). Learning how to explain neural networks: PatternNet and PatternAttribution. In International Conference on Learning Representations.
  112. 112.Koh, P. W., & Liang, P. (2017). Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning, pp. 1885–1894.
  113. 113.Kolodner, J. L. (1992). An introduction to case-based reasoning. Artificial Intelligence Review, 6 (1), 3–34.
  114. 114.Krakovna, V., & Doshi-Velez, F. (2016). Increasing the interpretability of recurrent neural networks using hidden Markov models. ArXiv, abs/1606.05320.
  115. 115.Kurakin, A., Goodfellow, I., & Bengio, S. (2016). Adversarial examples in the physical world. In International Conference on Learning Representations.
  116. 116.Lage, I., Chen, E., He, J., Narayanan, M., Kim, B., Gershman, S., & Doshi-Velez, F. (2019). An evaluation of the human-interpretability of explanation. ArXiv, abs/1902.00006.
  117. 117.Lapuschkin, S., Binder, A., Montavon, G., Müller, K.-R., & Samek, W. (2016). Analyzing classifiers: Fisher vectors and deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2912–2920.
  118. 118.Lei, T., Barzilay, R., & Jaakkola, T. (2016). Rationalizing neural predictions. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 107–117.
  119. 119.Letarte, G., Paradis, F., Giguère, P., & Laviolette, F. (2018). Importance of self-attention for sentiment analysis. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 267–275.
  120. 120.Li, J., Monroe, W., & Jurafsky, D. (2016). Understanding neural networks through representation erasure. ArXiv, abs/1612.08220.
  121. 121.Li, O., Liu, H., Chen, C., & Rudin, C. (2018a). Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions. In Thirty-Second AAAI Conference on Artificial Intelligence.
  122. 122.Li, Y., Min, M. R., Yu, W., Hsieh, C.-J., Lee, T., & Kruus, E. (2018b). Optimal transport classifier: Defending against adversarial attacks by regularized deep embedding. ArXiv, abs/1811.07950.
  123. 123.Lin, M., Chen, Q., & Yan, S. (2013). Network in network. In International Conference on Learning Representations.
  124. 124.Lipton, Z. C. (2017). The doctor just won’t accept that!. ArXiv, abs/1711.08037.
  125. 125.Lipton, Z. C. (2018). The mythos of model interpretability. Communications of the ACM, 61 (10), 36–43.
  126. 126.Liu, H., Yin, Q., & Wang, W. Y. (2019). Towards explainable NLP: A generative explanation framework for text classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 5570–5581.
  127. 127.Liu, S., Wang, X., Liu, M., & Zhu, J. (2017). Towards better analysis of machine learning models: A visual analytics perspective. Visual Informatics, 1 (1), 48–56.
  128. 128.Liu, X., Wang, X., & Matwin, S. (2018a). Improving the interpretability of deep neural networks with knowledge distillation. In 2018 IEEE International Conference on Data Mining Workshops, pp. 905–912.
  129. 129.Liu, X., Li, Y., Wu, C., & Hsieh, C.-J. (2018b). Adv-BNN: Improved adversarial defense through robust Bayesian neural network. In International Conference on Learning Representations.
  130. 130.Louizos, C., Swersky, K., Li, Y., Welling, M., & Zemel, R. (2015). The variational fair autoencoder. In International Conference on Learning Representations.
  131. 131.Lundberg, S. M., Erion, G. G., & Lee, S.-I. (2018). Consistent individualized feature attribution for tree ensembles. ArXiv, abs/1802.03888.
  132. 132.Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, pp. 4765–4774.
  133. 133.Lundén, J., & Koivunen, V. (2016). Deep learning for HRRP-based target recognition in multistatic radar systems. In IEEE Radar Conference, pp. 1–6.
  134. 134.Luong, M.-T., Pham, H., & Manning, C. D. (2015). Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 1412–1421.
  135. 135.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations.
  136. 136.Marchette, C. E. P. D. J., & Socolinsky, J. G. D. D. A. (2003). Classification using class cover catch digraphs. Journal of Classification, 20, 3–23.
  137. 137.Mascharka, D., Tran, P., Soklaski, R., & Majumdar, A. (2018). Transparency by design: Closing the gap between performance and interpretability in visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4942–4950.
  138. 138.Masi, I., Wu, Y., Hassner, T., & Natarajan, P. (2018). Deep face recognition: A survey. In 2018 31st SIBGRAPI Conference on Graphics, Patterns and Images, pp. 471–478.
  139. 139.Melis, D. A., & Jaakkola, T. (2018). Towards robust interpretability with self-explaining neural networks. In Advances in Neural Information Processing Systems, pp. 7775–7784.
  140. 140.Meng, D., & Chen, H. (2017). MagNet: A two-pronged defense against adversarial examples. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pp. 135–147.
  141. 141.Menon, A. K., & Williamson, R. C. (2018). The cost of fairness in binary classification. In Conference on Fairness, Accountability and Transparency, pp. 107–118.
  142. 142.Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence, 267, 1–38.
  143. 143.Miller, T. (2021). Contrastive explanation: a structural-model approach. The Knowledge Engineering Review, 36.
  144. 144.Molnar, C. (2020). Interpretable Machine Learning. Lulu.com.
  145. 145.Montavon, G., Lapuschkin, S., Binder, A., Samek, W., & Müller, K.-R. (2017). Explaining nonlinear classification decisions with deep Taylor decomposition. Pattern Recognition, 65, 211–222.
  146. 146.Montavon, G., Samek, W., & Müller, K.-R. (2018). Methods for interpreting and understanding deep neural networks. Digital Signal Processing, 73, 1–15.
  147. 147.Moosavi-Dezfooli, S.-M., Fawzi, A., Fawzi, O., & Frossard, P. (2017). Universal adversarial perturbations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1765–1773.
  148. 148.Moosavi-Dezfooli, S.-M., Fawzi, A., & Frossard, P. (2016). DeepFool: A simple and accurate method to fool deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2574–2582.
  149. 149.Mueller, S. T., Hoffman, R. R., Clancey, W., Emrey, A., & Klein, G. (2019). Explanation in human-AI systems: A literature meta-review, synopsis of key ideas and publications, and bibliography for explainable AI. ArXiv, abs/1902.01876.
  150. 150.Murdoch, W. J., Liu, P. J., & Yu, B. (2018). Beyond word importance: Contextual decomposition to extract interactions from LSTMs. In International Conference on Learning Representations.
  151. 151.Murdoch, W. J., & Szlam, A. (2017). Automatic rule extraction from long short term memory networks. In International Conference on Learning Representations.
  152. 152.Nemati, S., Holder, A., Razmi, F., Stanley, M. D., Clifford, G. D., & Buchman, T. G. (2018). An interpretable machine learning model for accurate prediction of sepsis in the ICU. Critical Care Medicine, 46 (4), 547–553.
  153. 153.Nguyen, A., Yosinski, J., & Clune, J. (2015). Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 427–436.
  154. 154.Nie, L., Wang, M., Zhang, L., Yan, S., Zhang, B., & Chua, T.-S. (2015). Disease inference from health-related questions via sparse deep learning. IEEE Transactions on Knowledge and Data Engineering, 27 (8), 2107–2119.
  155. 155.Nie, W., Zhang, Y., & Patel, A. (2018). A theoretical explanation for perplexing behaviors of backpropagation-based visualizations. In International Conference on Machine Learning, pp. 3806–3815.
  156. 156.Oana-Maria, C., Brendan, S., Pasquale, M., Thomas, L., & Phil, B. (2020). Make up your mind! Adversarial generation of inconsistent natural language explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4157–4165.
  157. 157.Olah, C., Mordvintsev, A., & Schubert, L. (2017). Feature visualization. Distill, 2 (11).
  158. 158.Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., & Mordvintsev, A. (2018). The building blocks of interpretability. Distill, 3 (3).
  159. 159.Olfat, M., & Aswani, A. (2018). Spectral algorithms for computing fair support vector machines. In International Conference on Artificial Intelligence and Statistics, pp. 1933–1942.
  160. 160.Ozbulak, U. (2019). PyTorch CNN visualizations. https://github.com/utkuozbulak/pytorch-cnn-visualizations.
  161. 161.Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z. B., & Swami, A. (2017). Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, pp. 506–519.
  162. 162.Papernot, N., McDaniel, P., Jha, S., Fredrikson, M., Celik, Z. B., & Swami, A. (2016a). The limitations of deep learning in adversarial settings. In 2016 IEEE European Symposium on Security and Privacy, pp. 372–387.
  163. 163.Papernot, N., McDaniel, P., Wu, X., Jha, S., & Swami, A. (2016b). Distillation as a defense to adversarial perturbations against deep neural networks. In 2016 IEEE Symposium on Security and Privacy, pp. 582–597.
  164. 164.Park, D. H., Hendricks, L. A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., & Rohrbach, M. (2018). Multimodal explanations: Justifying decisions and pointing to the evidence. In 31st IEEE Conference on Computer Vision and Pattern Recognition.
  165. 165.Park, D. H., Hendricks, L. A., Akata, Z., Schiele, B., Darrell, T., & Rohrbach, M. (2016). Attentive explanations: Justifying decisions and pointing to the evidence. ArXiv, abs/1612.04757.
  166. 166.Pérez-Suay, A., Laparra, V., Mateo-García, G., Muñoz-Marí, J., Gómez-Chova, L., & Camps-Valls, G. (2017). Fair kernel learning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 339–355.
  167. 167.Petsiuk, V., Das, A., & Saenko, K. (2018). RISE: Randomized input sampling for explanation of black-box models. In British Machine Vision Conference, p. 151.
  168. 168.Pham, T. T., & Shen, Y. (2017). A deep causal inference approach to measuring the effects of forming group loans in online non-profit microfinance platform. ArXiv, abs/1706.02795.
  169. 169.Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., & Weinberger, K. Q. (2017). On fairness and calibration. In Advances in Neural Information Processing Systems, pp. 5680–5689.
  170. 170.Prasad, G., Nie, Y., Bansal, M., Jia, R., Kiela, D., & Williams, A. (2020). To what extent do human explanations of model behavior align with actual model behavior?. ArXiv, abs/2012.13354.
  171. 171.Raghu, M., Gilmer, J., Yosinski, J., & Sohl-Dickstein, J. (2017). SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In Advances in Neural Information Processing Systems, pp. 6076–6085.
  172. 172.Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C. P., et al. (2018). Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists. PLOS Medicine, 15.
  173. 173.Ras, G., van Gerven, M., & Haselager, P. (2018). Explanation methods in deep learning: Users, values, concerns and challenges. In Explainable and Interpretable Models in Computer Vision and Machine Learning, pp. 19–36. Springer.
  174. 174.Ray, A., Yao, Y., Kumar, R., Divakaran, A., & Burachas, G. (2019). Can you explain that? Lucid explanations help human-AI collaborative image retrieval. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 7, pp. 153–161.
  175. 175.Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 779–788.
  176. 176.Ribeiro, M. T., Singh, S., & Guestrin, C. (2016a). Model-agnostic interpretability of machine learning. ArXiv, abs/1606.05386.
  177. 177.Ribeiro, M. T., Singh, S., & Guestrin, C. (2016b). Nothing else matters: Model-agnostic explanations by identifying prediction invariance. ArXiv, abs/1611.05817.
  178. 178.Ribeiro, M. T., Singh, S., & Guestrin, C. (2016c). Why should I trust you?: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1135–1144.
  179. 179.Ribeiro, M. T., Singh, S., & Guestrin, C. (2018). Anchors: High-precision model-agnostic explanations. In Thirty-Second AAAI Conference on Artificial Intelligence.
  180. 180.Robnik-Šikonja, M., & Bohanec, M. (2018). Perturbation-based explanations of prediction models. In Human and Machine Learning, pp. 159–175. Springer.
  181. 181.Robnik-Šikonja, M., & Kononenko, I. (2008). Explaining classifications for individual instances. IEEE Transactions on Knowledge and Data Engineering, 20 (5), 589–600.
  182. 182.Rolnick, D., Donti, P. L., Kaack, L. H., Kochanski, K., Lacoste, A., Sankaran, K., Ross, A. S., Milojevic-Dupont, N., Jaques, N., Waldman-Brown, A., et al. (2019). Tackling climate change with machine learning. ArXiv, abs/1906.05433.
  183. 183.Rozsa, A., Rudd, E. M., & Boult, T. E. (2016). Adversarial diversity and hard positive generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 25–32.
  184. 184.Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1 (5), 206.
  185. 185.Sabour, S., Cao, Y., Faghri, F., & Fleet, D. J. (2015). Adversarial manipulation of deep representations. In International Conference on Learning Representations.
  186. 186.Saha, A., Subramanya, A., & Pirsiavash, H. (2020). Hidden trigger backdoor attacks. Proceedings of the AAAI Conference on Artificial Intelligence, 34, 11957–11965.
  187. 187.Samangouei, P., Kabkab, M., & Chellappa, R. (2018). Defense-GAN: Protecting classifiers against adversarial attacks using generative models. In International Conference on Learning Representations.
  188. 188.Samek, W., Binder, A., Montavon, G., Lapuschkin, S., & Müller, K.-R. (2016). Evaluating the visualization of what a deep neural network has learned. IEEE Transactions on Neural Networks and Learning Systems, 28 (11), 2660–2673.
  189. 189.Samek, W., Wiegand, T., & Müller, K.-R. (2017). Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models. ArXiv, abs/1708.08296.
  190. 190.Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision, pp. 618–626.
  191. 191.Serrano, S., & Smith, N. A. (2019). Is attention interpretable?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2931–2951.
  192. 192.Shapley, L. (1953). A value for n-person games. Contributions to the Theory of Games, 2.28, 307–317.
  193. 193.Shrikumar, A., Greenside, P., & Kundaje, A. (2017). Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning, pp. 3145–3153.
  194. 194.Simonyan, K., Vedaldi, A., & Zisserman, A. (2013). Deep inside convolutional networks: Visualising image classification models and saliency maps. In International Conference on Learning Representations.
  195. 195.Sirignano, J., Sadhwani, A., & Giesecke, K. (2016). Deep learning for mortgage risk. ArXiv, abs/1607.02470.
  196. 196.Sobol, I. M. (2001). Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates. Mathematics and Computers in Simulation, 55 (1-3), 271–280.
  197. 197.Springenberg, J. T., Dosovitskiy, A., Brox, T., & Riedmiller, M. (2014). Striving for simplicity: The all convolutional net. In International Conference on Learning Representations.
  198. 198.Su, J., Vargas, D. V., & Sakurai, K. (2019). One pixel attack for fooling deep neural networks. IEEE Transactions on Evolutionary Computation, 23 (5), 828–841.
  199. 199.Sundararajan, M., Taly, A., & Yan, Q. (2016). Gradients of counterfactuals. ArXiv, abs/1611.02639.
  200. 200.Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, pp. 3319–3328.
  201. 201.Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2013). Intriguing properties of neural networks. In International Conference on Learning Representations.
  202. 202.Tabacof, P., & Valle, E. (2016). Exploring the space of adversarial images. In 2016 International Joint Conference on Neural Networks, pp. 426–433.
  203. 203.Tan, S., Caruana, R., Hooker, G., Koch, P., & Gordo, A. (2018). Learning global additive explanations for neural nets using model distillation. ArXiv, abs/1801.08640.
  204. 204.Teney, D., Anderson, P., He, X., & van den Hengel, A. (2018). Tips and tricks for visual question answering: Learnings from the 2017 challenge. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4223–4232.
  205. 205.Tjoa, E., & Guan, C. (2020). A survey on explainable artificial intelligence (XAI): Toward medical XAI. In IEEE Transactions on Neural Networks and Learning Systems, pp. 1–21.
  206. 206.Tulshan, A. S., & Dhage, S. N. (2018). Survey on virtual assistant: Google Assistant, Siri, Cortana, Alexa. In International Symposium on Signal Processing and Intelligent Recognition Systems, pp. 190–201.
  207. 207.Vashishth, S., Upadhyay, S., Tomar, G. S., & Faruqui, M. (2019). Attention interpretability across NLP tasks. ArXiv, abs/1909.11218.
  208. 208.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, pp. 5998–6008.
  209. 209.Verma, S., Dickerson, J., & Hines, K. (2020). Counterfactual Explanations for Machine Learning: A Review. ArXiv, abs/2010.10596.
  210. 210.Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3156–3164.
  211. 211.Vu, M. N., Nguyen, T. D., Phan, N., Gera, R., & Thai, M. T. (2019). c-Eval: A unified metric to evaluate feature-based explanations via perturbation. ArXiv, abs/1906.02032.
  212. 212.Wang, Y., Huang, M., Zhao, L., et al. (2016). Attention-based LSTM for aspect-level sentiment classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 606–615.
  213. 213.Wiegreffe, S., & Pinter, Y. (2019). Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pp. 11–20.
  214. 214.Wolf, L., Galanti, T., & Hazan, T. (2019). A formal approach to explainability. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pp. 255–261.
  215. 215.Woodworth, B., Gunasekar, S., Ohannessian, M. I., & Srebro, N. (2017). Learning non-discriminatory predictors. ArXiv, abs/1702.06081.
  216. 216.Xie, C., Wang, J., Zhang, Z., Ren, Z., & Yuille, A. (2017). Mitigating adversarial effects through randomization. In International Conference on Learning Representations.
  217. 217.Xie, N., Lai, F., Doran, D., & Kadav, A. (2019). Visual entailment: A novel task for fine-grained image understanding. ArXiv, abs/1901.06706.
  218. 218.Xie, N., Sarker, M. K., Doran, D., Hitzler, P., & Raymer, M. (2017). Relating input concepts to convolutional neural network decisions. ArXiv, abs/1711.08006.
  219. 219.Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., & Bengio, Y. (2015). Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning, pp. 2048–2057.
  220. 220.Yim, J., Joo, D., Bae, J., & Kim, J. (2017). A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In The IEEE Conference on Computer Vision and Pattern Recognition.
  221. 221.Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., & Lipson, H. (2015). Understanding neural networks through deep visualization. ArXiv, abs/1506.06579.
  222. 222.Yuan, X., He, P., Zhu, Q., & Li, X. (2019). Adversarial examples: Attacks and defenses for deep learning. IEEE Transactions on Neural Networks and Learning Systems, 30 (9), 2805–2824.
  223. 223.Zafar, M. B., Valera, I., Gomez Rodriguez, M., & Gummadi, K. P. (2017a). Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment. In Proceedings of the 26th International Conference on World Wide Web, pp. 1171–1180.
  224. 224.Zafar, M. B., Valera, I., Rodriguez, M., Gummadi, K., & Weller, A. (2017b). From parity to preference-based notions of fairness in classification. In Advances in Neural Information Processing Systems, pp. 229–239.
  225. 225.Zeiler, M. D., & Fergus, R. (2014). Visualizing and understanding convolutional networks. In European Conference on Computer Vision, pp. 818–833.
  226. 226.Zeiler, M. D., Taylor, G. W., & Fergus, R. (2011). Adaptive deconvolutional networks for mid and high level feature learning. In 2011 International Conference on Computer Vision, pp. 2018–2025.
  227. 227.Zellers, R., Bisk, Y., Farhadi, A., & Choi, Y. (2019). From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6720–6731.
  228. 228.Zemel, R., Wu, Y., Swersky, K., Pitassi, T., & Dwork, C. (2013). Learning fair representations. In International Conference on Machine Learning, pp. 325–333.
  229. 229.Zhang, C., Bengio, S., Hardt, M., Recht, B., & Vinyals, O. (2016). Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations.
  230. 230.Zhang, Q.-s., & Zhu, S.-C. (2018). Visual interpretability for deep learning: A survey. Frontiers of Information Technology & Electronic Engineering, 19 (1), 27–39.
  231. 231.Zhang, Q., Cao, R., Shi, F., Wu, Y. N., & Zhu, S.-C. (2018). Interpreting CNN knowledge via an explanatory graph. In Thirty-Second AAAI Conference on Artificial Intelligence.
  232. 232.Zhang, Q., Cao, R., Wu, Y. N., & Zhu, S.-C. (2017). Growing interpretable part graphs on ConvNets via multi-shot learning. In Proceedings of AAAI Conference on Artificial Intelligence, pp. 2898–2906.
  233. 233.Zhang, Q., Yang, Y., Ma, H., & Wu, Y. N. (2019a). Interpreting CNNs via decision trees. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6261–6270.
  234. 234.Zhang, W. E., Sheng, Q. Z., & Alhazmi, A. A. F. (2019b). Generating textual adversarial examples for deep learning models: A survey. ArXiv, abs/1901.06796.
  235. 235.Zhao, Z., Dua, D., & Singh, S. (2017). Generating natural adversarial examples. In International Conference on Learning Representations.
  236. 236.Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2014). Object detectors emerge in deep scene CNNs. In International Conference on Learning Representations.
  237. 237.Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2016). Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2921–2929.
  238. 238.Zintgraf, L. M., Cohen, T. S., Adel, T., & Welling, M. (2017). Visualizing deep neural network decisions: Prediction difference analysis. In International Conference on Learning Representations.

Citation

MLA
Ras, G., et al. “Explainable Deep Learning: A Field Guide for the Uninitiated”. arXiv, 2020, http://arxiv.org/abs/2004.14545v2.
APA
Ras, G., Xie, N., Gerven, M. van ., & Doran, D. (2020). Explainable Deep Learning: A Field Guide for the Uninitiated. arXiv. http://arxiv.org/abs/2004.14545v2
Chicago
Ras, G., N. Xie, M. van . Gerven, and D. Doran. 2020. “Explainable Deep Learning: A Field Guide for the Uninitiated”. arXiv. http://arxiv.org/abs/2004.14545v2.
Harvard
Ras, G. et al. (2020) “Explainable Deep Learning: A Field Guide for the Uninitiated”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2004.14545v2.
Vancouver
1. Ras G, Xie N, Gerven M van, Doran D (2020) Explainable Deep Learning: A Field Guide for the Uninitiated. arXiv

BibTeX

@article{ras2020explainable,
  title = {Explainable Deep Learning: A Field Guide for the Uninitiated},
  author = {Ras, Gabrielle and Xie, Ning and Gerven, Marcel van and Doran, Derek},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2004.14545v2},
  eprint = {2004.14545}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/