Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization

Ramprasaath R. SelvarajuAbhishek DasRamakrishna VedantamMichael CogswellDevi ParikhDhruv Batra

article2016IJCV29,332 citations

Introduces Grad-CAM, a gradient-based visual explanation technique that makes arbitrary convolutional neural networks interpretable without retraining, allowing practitioners to debug model failures, identify dataset biases, and verify prediction rationales.

Listen

Deep neural networks based on convolutional layers now deliver strong results on image classification, captioning, and visual question answering, yet their internal decisions remain opaque. When these systems err, they often do so without warning or justification, which erodes user trust and slows safe deployment in real-world settings. The paper introduces Gradient-weighted Class Activation Mapping (Grad-CAM) to generate visual explanations that highlight the image regions driving any target output, without requiring changes to model architecture or additional training.

The method computes the gradient of the target score with respect to the final convolutional feature maps, globally averages those gradients to obtain importance weights for each map, and produces a coarse localization heatmap by a weighted combination followed by a ReLU. High-resolution detail is recovered by element-wise multiplication with Guided Backpropagation. The approach was tested on standard ImageNet-pretrained networks, captioning models, and several visual-question-answering architectures, using both quantitative localization benchmarks on ILSVRC-15 and PASCAL VOC and human-subject studies on Amazon Mechanical Turk.

Grad-CAM yields lower localization error than prior methods such as Class Activation Mapping and contrastive Marginal Winning Probability while preserving classification accuracy. It produces maps that are demonstrably more class-discriminative than Guided Backpropagation or Deconvolution alone, and these maps correlate more strongly with occlusion-based sensitivity maps, indicating greater faithfulness to the underlying model. Human evaluators correctly identify the visualized class more often with Guided Grad-CAM than with baseline visualizations, and they reliably distinguish a stronger network from a weaker one when shown only the explanations. The same technique exposes dataset biases, such as gender stereotypes in adoctor versus nurseclassifier, and remains stable under adversarial perturbations that fool the network’s final prediction.

These capabilities matter because they let practitioners diagnose failure modes, detect unintended biases, and decide whether a model is trustworthy enough for deployment without sacrificing accuracy. The visualizations also work on models that lack explicit attention mechanisms, showing that many captioning and question-answering networks already localize relevant image content.

The main limitations are that localization quality degrades in earlier convolutional layers and that the method has so far been demonstrated primarily on vision tasks. The evidence nevertheless supports immediate use of Grad-CAM for model auditing and bias auditing, with further validation recommended on reinforcement-learning and language-only pipelines before broader policy reliance.

  • Paper: Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks, Aditya Chattopadhyay et al. (2017). Grad-CAM++ directly extends the source work by incorporating higher-order derivatives to better localize multiple object instances and capture entire object regions.
  • Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). This follow-up study subjects Grad-CAM and related saliency methods to rigorous sanity checks, evaluating whether their visual explanations genuinely depend on learned model parameters.
Cover for Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization

Abstract

We propose a technique for producing "visual explanations" for decisions from a large class of CNN-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept, flowing into the final convolutional layer to produce a coarse localization map highlighting important regions in the image for predicting the concept. Grad-CAM is applicable to a wide variety of CNN model-families: (1) CNNs with fully-connected layers, (2) CNNs used for structured outputs, (3) CNNs used in tasks with multimodal inputs or reinforcement learning, without any architectural changes or re-training. We combine Grad-CAM with fine-grained visualizations to create a high-resolution class-discriminative visualization and apply it to off-the-shelf image classification, captioning, and visual question answering (VQA) models, including ResNet-based architectures. In the context of image classification models, our visualizations (a) lend insights into their failure modes, (b) are robust to adversarial images, (c) outperform previous methods on localization, (d) are more faithful to the underlying model and (e) help achieve generalization by identifying dataset bias. For captioning and VQA, we show that even non-attention based models can localize inputs. We devise a way to identify important neurons through Grad-CAM and combine it with neuron names to provide textual explanations for model decisions. Finally, we design and conduct human studies to measure if Grad-CAM helps users establish appropriate trust in predictions from models and show that Grad-CAM helps untrained users successfully discern a 'stronger' nodel from a 'weaker' one even when both make identical predictions. Our code is available at this https URL, along with a demo at this http URL, and a video at this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Grad-CAM
  • 3.1 Grad-CAM generalizes CAM
  • 3.2 Guided Grad-CAM
  • 3.3 Counterfactual Explanations
  • 4 Evaluating Localization Ability of Grad-CAM
  • 4.1 Weakly-supervised Localization
  • 4.2 Weakly-supervised Segmentation
  • 4.3 Pointing Game
  • 5 Evaluating Visualizations
  • 5.1 Evaluating Class Discrimination
  • 5.2 Evaluating Trust
  • 5.3 Faithfulness vs. Interpretability
  • 6 Diagnosing image classification CNNs with Grad-CAM
  • 6.1 Analyzing failure modes for VGG-16
  • 6.2 Effect of adversarial noise on VGG-16
  • 6.3 Identifying bias in dataset
  • 7 Textual Explanations with Grad-CAM
  • 8 Grad-CAM for Image Captioning and VQA
  • 8.1 Image Captioning
  • 8.1.1 Grad-CAM for individual words of caption
  • 8.2 Visual Question Answering
  • 9 Conclusion
  • 10 Acknowledgements
  • A Appendix Overview
  • B Ablation studies
  • C Qualitative results for vision and language tasks
  • D More details of Pointing Game
  • E Qualitative comparison to Excitation Backprop (c-MWP) and CAM
  • F Visual and Textual explanations for Places dataset
  • G Analyzing Residual Networks
  • References

Knowls

  1. Knowl 1 — Gradient-weighted Class Activation Mapping

    model/method

    Gradient-weighted Class Activation Mapping (Grad-CAM) generates a coarse 2D localization heatmap indicating the discriminative image regions supporting any target concept cc in a Convolutional Neural Network (CNN), without requiring architectural modifications or retraining.

    Let ycy^c be the score (the logit before the softmax layer) for a class or concept cc, and let AkRu×vA^k \in \mathbb{R}^{u \times v} be the kk-th feature map activation of a chosen convolutional layer (such as the final convolutional layer) of spatial width uu and height vv, containing Z=uvZ = u \cdot v spatial locations indexed by (i,j)(i,j).

    1. Neuron importance weights αkc\alpha_k^c are calculated by global-average-pooling the gradients of ycy^c with respect to the activations AkA^k: αkc=1Zi=1uj=1vycAijk\alpha_k^c = \frac{1}{Z} \sum_{i=1}^u \sum_{j=1}^v \frac{\partial y^c}{\partial A_{ij}^k}

    2. The class-discriminative localization map LGrad-CAMcRu×vL_{\text{Grad-CAM}}^c \in \mathbb{R}^{u \times v} is computed by taking a weighted linear combination of forward activation maps followed by a Rectified Linear Unit (ReLU): LGrad-CAMc=ReLU(kαkcAk)L_{\text{Grad-CAM}}^c = \text{ReLU}\left( \sum_k \alpha_k^c A^k \right)

    The weight αkc\alpha_k^c represents a partial linearization of the downstream layers and captures the importance of feature map kk for class cc. The ReLU operation ensures that the map exclusively highlights features that exert a positive influence on the score ycy^c, while suppressing features associated with other classes.

  2. Knowl 2 — Guided Grad-CAM Visualization

    model/method

    While Grad-CAM produces class-discriminative heatmaps that localize object categories, the resulting maps are coarse due to the low spatial resolution of deep convolutional layers. Conversely, pixel-space gradient visualization methods (such as Guided Backpropagation) generate high-resolution visualizations of edges and textures but lack class-discriminative capability.

    Guided Grad-CAM fuses these complementary representations to produce high-resolution, class-discriminative visual explanations:

    1. The coarse Grad-CAM localization map LGrad-CAMcL_{\text{Grad-CAM}}^c is upsampled to the input image dimensions using bilinear interpolation.
    2. The upsampled heatmap is multiplied point-wise with the pixel-space visualization obtained from Guided Backpropagation: VGuided Grad-CAMc=VGuided Backpropinterpolate(LGrad-CAMc)V_{\text{Guided Grad-CAM}}^c = V_{\text{Guided Backprop}} \odot \text{interpolate}(L_{\text{Grad-CAM}}^c)

    This pointwise gating suppresses features that do not belong to the target class cc while highlighting fine-grained visual details (such as facial markings, stripes, or object parts) specific to that class.

  3. Knowl 3 — Generalization of Class Activation Mapping (CAM)

    theoretical result

    Grad-CAM is a strict generalization of Class Activation Mapping (CAM) to arbitrary CNN architectures.

    In CAM architectures, the KK convolutional feature maps AkRu×vA^k \in \mathbb{R}^{u \times v} of the penultimate layer are spatially averaged via Global Average Pooling (GAP) to obtain Fk=1Zi=1uj=1vAijkF^k = \frac{1}{Z} \sum_{i=1}^u \sum_{j=1}^v A_{ij}^k (where Z=uvZ = u \cdot v), and the pre-softmax score for class cc is computed as a linear combination Yc=kwkcFkY^c = \sum_k w_k^c F^k.

    Taking the partial derivative of YcY^c with respect to a single feature map element AijkA_{ij}^k gives: YcAijk=YcFkFkAijk=wkc1Z\frac{\partial Y^c}{\partial A_{ij}^k} = \frac{\partial Y^c}{\partial F^k} \cdot \frac{\partial F^k}{\partial A_{ij}^k} = w_k^c \cdot \frac{1}{Z}

    Summing both sides over all ZZ spatial pixels (i,j)(i,j) yields: i=1uj=1vYcAijk=Z(1Zwkc)=wkc\sum_{i=1}^u \sum_{j=1}^v \frac{\partial Y^c}{\partial A_{ij}^k} = Z \cdot \left( \frac{1}{Z} w_k^c \right) = w_k^c

    Rearranging terms shows that the class feature weight wkcw_k^c equals: wkc=Z(1Zi=1uj=1vYcAijk)=Zαkcw_k^c = Z \cdot \left( \frac{1}{Z} \sum_{i=1}^u \sum_{j=1}^v \frac{\partial Y^c}{\partial A_{ij}^k} \right) = Z \cdot \alpha_k^c

    Up to the proportionality constant 1/Z1/Z (which is eliminated when visual heatmaps are normalized), the weight wkcw_k^c in CAM is identical to the importance weight αkc\alpha_k^c in Grad-CAM. Thus, Grad-CAM reproduces CAM for GAP architectures while extending visual explanations to architectures with fully-connected layers, multi-modal networks, and recurrent layers without retraining.

  4. Knowl 4 — Weakly-Supervised Localization Performance on ImageNet

    data/table

    Weakly-supervised object localization on the ILSVRC-15 validation set is evaluated by generating Grad-CAM heatmaps for predicted classes, thresholding the heatmap at 15% of its maximum intensity to extract connected components, and fitting a bounding box to the single largest segment. The underlying models are trained solely on image-level class labels without bounding box annotations.

    Classification Error (%) Localization Error (%)
    Method Top-1 Top-5 Top-1 Top-5
    VGG-16
    Backprop 30.38 10.89 61.12 51.46
    c-MWP 30.38 10.89 70.92 63.04
    Grad-CAM (ours) 30.38 10.89 56.51 46.41
    CAM 33.40 12.20 57.20 45.14
    AlexNet
    c-MWP 44.2 20.8 92.6 89.2
    Grad-CAM (ours) 44.2 20.8 68.3 56.6
    GoogleNet
    Grad-CAM (ours) 31.9 11.3 60.09 49.34
    CAM 31.9 11.3 60.09 49.34

    Grad-CAM applied to an off-the-shelf VGG-16 achieves a top-1 localization error of 56.51%, outperforming plain Backpropagation (61.12%), c-MWP (70.92%), and CAM (57.20%). Furthermore, while CAM requires replacing fully-connected layers with GAP and retraining the network (incurring a 2.98% drop in top-1 classification performance), Grad-CAM achieves lower localization error without compromising classification accuracy.

  5. Knowl 5 — Human Trust, Class Discrimination, and Model Faithfulness Evaluations

    data/table

    Human user studies on Amazon Mechanical Turk and occlusion-based sensitivity experiments quantify the class-discriminative ability, trustworthiness, and model faithfulness of visual explanations generated for VGG-16 and AlexNet fine-tuned on PASCAL VOC 2007.

    Method Human Classification Accuracy (%) Relative Reliability Rank Correlation w/ Occlusion
    Guided Backpropagation 44.44 +1.00 0.168
    Guided Grad-CAM 61.23 +1.27 0.261
    1. Class Discrimination: Across 90 image pairs with two distinct annotated classes evaluated by 43 human workers, Guided Grad-CAM visualizations enabled workers to identify the visualized class in 61.23% of trials, compared to 44.44% for Guided Backpropagation (a 16.79% absolute improvement). Deconvolution Grad-CAM also improved human identification over baseline Deconvolution (60.37% vs. 53.33%).
    2. Trust / Relative Reliability: When AlexNet and VGG-16 made identical, correct predictions, 54 human subjects rated explanation reliability on a scale from 2-2 to +2+2. Guided Grad-CAM produced a mean score of +1.27+1.27 favoring VGG-16 (the more accurate model: 79.09 vs. 69.20 mAP) compared to +1.00+1.00 for Guided Backpropagation, demonstrating that Grad-CAM helps untrained users identify better generalizing models from individual predictions.
    3. Faithfulness: Faithfulness is evaluated by measuring the rank correlation between explanation intensity and score drops from patch occlusion across 2,510 images. Guided Grad-CAM achieves a correlation of 0.261 (and Grad-CAM achieves 0.254), higher than Guided Backpropagation (0.168), CAM (0.208), and c-MWP (0.220).
  6. Knowl 6 — Counterfactual Visual Explanations via Negated Gradients

    model/method

    Counterfactual explanations identify spatial regions containing visual evidence that opposes a target class prediction cc—that is, regions whose presence reduces the model's confidence in class cc, and whose removal would increase the prediction score.

    Counterfactual importance weights αkc\alpha_k^c are obtained by global-average-pooling the negated gradients of the pre-softmax score ycy^c with respect to convolutional feature map activations AkA^k: αkc=1Zi=1uj=1vycAijk\alpha_k^c = \frac{1}{Z} \sum_{i=1}^u \sum_{j=1}^v -\frac{\partial y^c}{\partial A_{ij}^k}

    The counterfactual localization map is computed via the rectified weighted combination: Lcounterfactualc=ReLU(kαkcAk)L_{\text{counterfactual}}^c = \text{ReLU}\left( \sum_k \alpha_k^c A^k \right)

    This highlights regions that provide contradictory evidence against category cc, distinguishing them from standard Grad-CAM maps that highlight supportive evidence.

  7. Knowl 7 — Weakly-Supervised Segmentation and Pointing Game Localization

    empirical result

    Grad-CAM provides effective localization cues across weakly-supervised segmentation and pointing game benchmarks:

    1. Weakly-Supervised Semantic Segmentation: In the Seed, Expand, Constrain (SEC) framework on PASCAL VOC 2012, replacing CAM localization seeds with Grad-CAM heatmaps obtained from a standard VGG-16 network improves validation Intersection over Union (IoU) from 44.6% to 49.6%.
    2. Pointing Game with Rejection: Evaluated on top-5 CNN predictions on COCO with the ability to reject absent classes when maximum heatmap intensity falls below a threshold, Grad-CAM achieves a localization accuracy of 70.58%, significantly outperforming contrastive Marginal Winning Probability (c-MWP, 60.30%) on VGG-16.
    3. Caption Word Grounding: Using a Show-and-Tell image captioning model on 1,000 COCO images with 830 caption words mapped to 80 COCO categories, Grad-CAM pointing accuracy against ground-truth segmentation masks reaches 30.0% without grounding supervision during training.
  8. Knowl 8 — Textual Explanations via Grad-CAM Neuron Importance

    model/method

    Grad-CAM neuron importance scores can be combined with automatic concept labels assigned to convolutional filters (such as names generated via Network Dissection) to produce textual explanations for neural network decisions.

    1. Semantic concept names are assigned to channels in the final convolutional layer based on overlap (e.g., IoU0.05\text{IoU} \ge 0.05) between individual filter activations and ground-truth concept segmentation masks.
    2. For a given input image and class score ycy^c, neuron importance weights αkc=1Zi,jycAijk\alpha_k^c = \frac{1}{Z} \sum_{i,j} \frac{\partial y^c}{\partial A_{ij}^k} are computed.
    3. Neurons are sorted by their importance scores:
      • Top-5 positive αkc\alpha_k^c neurons are identified as persuasive concepts whose presence increases the class prediction.
      • Bottom-5 negative αkc\alpha_k^c neurons are identified as inhibitive concepts whose presence decreases the class prediction.
    4. The names of these top and bottom neurons provide a textual justification complementing the spatial Grad-CAM heatmap.
  9. Knowl 9 — Grad-CAM for Multimodal Vision and Language Tasks

    model/method

    Grad-CAM generates visual explanations for multimodal architectures where the output target score ycy^c is not a standard image classification logit:

    1. Image Captioning: For CNN-LSTM models (such as NeuralTalk2), setting ycy^c to the log probability of a generated sentence or an individual word wtw_t at step tt allows gradients to flow back to the CNN's final convolutional layer (e.g., conv5_3 in VGG-16). In DenseCap evaluation, the ratio of mean activation inside versus outside ground-truth bounding boxes is 3.27±0.183.27 \pm 0.18 for Grad-CAM and 6.38±0.996.38 \pm 0.99 for Guided Grad-CAM (compared to a baseline ratio of 1.01.0 for uniform attention and 2.32±0.082.32 \pm 0.08 for Guided Backpropagation).
    2. Visual Question Answering (VQA): In VQA networks predicting answers from fused image and question representations, setting ycy^c to the logit of the predicted answer produces answer-specific heatmaps. On the VQA benchmark (Lu et al.), Grad-CAM achieves a rank correlation with occlusion maps of 0.60±0.0380.60 \pm 0.038 (vs. 0.42±0.0380.42 \pm 0.038 for Guided Backpropagation) and a rank correlation with human gaze attention maps of 0.1360.136.
  10. Knowl 10 — Dataset Bias Diagnosis and Adversarial Noise Robustness

    empirical result

    Grad-CAM visual explanations diagnose systematic dataset biases and evaluate model sensitivity to adversarial attacks:

    1. Dataset Bias Diagnosis: A VGG-16 classifier trained on web-scraped images for a binary 'doctor' vs. 'nurse' task achieved high training and validation accuracy but degraded to 82% test accuracy on a gender-balanced test set. Grad-CAM revealed that the model localized facial features and hairstyles to differentiate classes due to search query bias (where 78% of doctor images were men and 93% of nurse images were women). Retraining the model on a debiased dataset with female doctors and male nurses increased test accuracy to 90% and shifted Grad-CAM attention toward stethoscopes and uniforms.
    2. Adversarial Robustness: In images perturbed by adversarial noise that causes VGG-16 to misclassify scenes as 'airliner' with >0.9999>0.9999 confidence, Grad-CAM computed for the true underlying classes (e.g., 'tiger cat' or 'boxer') continues to localize the actual objects accurately, while Grad-CAM for the predicted adversarial class highlights diffuse background areas.
  11. Knowl 11 — Ablation Analysis of Grad-CAM Components

    data/table

    An ablation study on the ILSVRC-15 validation set evaluates localization performance under different gradient pooling operations, activation rectifications, and layer selections.

    Method Top-1 Localization Error (%)
    Grad-CAM 59.65
    Grad-CAM without ReLU 74.98
    Grad-CAM with Absolute gradients 58.19
    Grad-CAM with Global Max Pooling (GMP) 59.96
    Grad-CAM with Deconv ReLU 83.95
    Grad-CAM with Guided ReLU 59.14
    • Impact of ReLU: Omitting the ReLU operation after the weighted sum increases top-1 localization error from 59.65% to 74.98% (+15.33%), as negative weights highlight non-target classes.
    • Global Average Pooling vs. Max Pooling: Global Average Pooling (59.65% error) outperforms Global Max Pooling (59.96% error) because averaging is more robust to gradient noise.
    • Backward ReLU Variants: Using Deconvolution ReLU in the backward pass severely degrades localization performance (83.95% error), showing that negative gradients carry critical class-discriminative information. Guided ReLU yields 59.14% localization error but reduces class discriminativeness in qualitative visualizations.
    • Layer Depth: Computing Grad-CAM on shallower convolutional layers progressively degrades localization due to smaller receptive fields and the lack of high-level semantic representations.

Coverage note — Specific qualitative visualization figures and software demo implementation details were omitted as they serve as individual illustrative instances of the core evaluated methods.

References

  1. 1.A. Agrawal, D. Batra, and D. Parikh. Analyzing the Behavior of Visual Question Answering Models. In EMNLP, 2016.
  2. 2.H. Agrawal, C. S. Mathialagan, Y. Goyal, N. Chavali, P. Banik, A. Mohapatra, A. Osman, and D. Batra. CloudCV: Large Scale Distributed Computer Vision as a Cloud Service. In Mobile Cloud Visual Media Computing, pages 265–290. Springer, 2015.
  3. 3.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh. VQA: Visual Question Answering. In ICCV, 2015.
  4. 4.D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba. Network dissection: Quantifying interpretability of deep visual representations. In Computer Vision and Pattern Recognition, 2017.
  5. 5.L. Bazzani, A. Bergamo, D. Anguelov, and L. Torresani. Self-taught object localization with deep networks. In WACV, 2016.
  6. 6.Y. Bengio, A. Courville, and P. Vincent. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8):1798–1828, 2013.
  7. 7.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft COCO captions: Data Collection and Evaluation Server. arXiv preprint arXiv:1504.00325, 2015.
  8. 8.R. G. Cinbis, J. Verbeek, and C. Schmid. Weakly supervised object localization with multi-fold multiple instance learning. IEEE transactions on pattern analysis and machine intelligence, 2016.
  9. 9.A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra. Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions? In EMNLP, 2016.
  10. 10.A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra. Embodied Question Answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  11. 11.A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra. Visual Dialog. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  12. 12.A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra. Learning cooperative visual dialog agents with deep reinforcement learning. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017.
  13. 13.H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. C. Courville. Guesswhat?! visual object discovery through multi-modal dialogue. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  14. 14.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR, 2009.
  15. 15.A. Dosovitskiy and T. Brox. Inverting Convolutional Networks with Convolutional Networks. In CVPR, 2015.
  16. 16.D. Erhan, Y. Bengio, A. Courville, and P. Vincent. Visualizing Higher-layer Features of a Deep Network. University of Montreal, 1341, 2009.
  17. 17.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascalnetwork.org/challenges/VOC/voc2007/workshop/index.html, 2009.
  18. 18.H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, et al. From Captions to Visual Concepts and Back. In CVPR, 2015.
  19. 19.C. Gan, N. Wang, Y. Yang, D.-Y. Yeung, and A. G. Hauptmann. Devnet: A deep event network for multimedia event detection and evidence recounting. In CVPR, 2015.
  20. 20.H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu. Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering. In NIPS, 2015.
  21. 21.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In CVPR, 2014.
  22. 22.I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. stat, 2015.
  23. 23.D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi. Iqa: Visual question answering in interactive environments. arXiv preprint arXiv:1712.03316, 2017.
  24. 24.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  25. 25.D. Hoiem, Y. Chodpathumwan, and Q. Dai. Diagnosing Error in Object Detectors. In ECCV, 2012.
  26. 26.P. Jackson. Introduction to Expert Systems. Addison-Wesley Longman Publishing Co., Inc., Boston, MA, USA, 3rd edition, 1998.
  27. 27.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional Architecture for Fast Feature Embedding. In ACM MM, 2014.
  28. 28.E. Johns, O. Mac Aodha, and G. J. Brostow. Becoming the Expert - Interactive Multi-Class Machine Teaching. In CVPR, 2015.
  29. 29.J. Johnson, A. Karpathy, and L. Fei-Fei. DenseCap: Fully Convolutional Localization Networks for Dense Captioning. In CVPR, 2016.
  30. 30.A. Karpathy. What I learned from competing against a ConvNet on ImageNet. http://karpathy.github.io/2014/09/02/what-i-learnedfrom-competing-against-a-convnet-on-imagenet/, 2014.
  31. 31.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015.
  32. 32.A. Kolesnikov and C. H. Lampert. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In ECCV, 2016.
  33. 33.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  34. 34.M. Lin, Q. Chen, and S. Yan. Network in network. In ICLR, 2014.
  35. 35.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV. 2014.
  36. 36.Z. C. Lipton. The Mythos of Model Interpretability. ArXiv e-prints, June 2016.
  37. 37.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015.
  38. 38.J. Lu, X. Lin, D. Batra, and D. Parikh. Deeper LSTM and normalized CNN Visual Question Answering model. https://github.com/VT-vision-lab/VQA_LSTM_CNN, 2015.
  39. 39.J. Lu, J. Yang, D. Batra, and D. Parikh. Hierarchical question-image co-attention for visual question answering. In NIPS, 2016.
  40. 40.A. Mahendran and A. Vedaldi. Salient deconvolutional networks. In European Conference on Computer Vision, 2016.
  41. 41.A. Mahendran and A. Vedaldi. Visualizing deep convolutional neural networks using natural pre-images. International Journal of Computer Vision, pages 1–23, 2016.
  42. 42.M. Malinowski, M. Rohrbach, and M. Fritz. Ask your neurons: A neural-based approach to answering questions about images. In ICCV, 2015.
  43. 43.M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Learning and transferring mid-level image representations using convolutional neural networks. In CVPR, 2014.
  44. 44.M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Is object localization for free? – weakly-supervised learning with convolutional neural networks. In CVPR, 2015.
  45. 45.P. O. Pinheiro and R. Collobert. From image-level to pixel-level labeling with convolutional networks. In CVPR, 2015.
  46. 46.M. Ren, R. Kiros, and R. Zemel. Exploring models and data for image question answering. In NIPS, 2015.
  47. 47.M. T. Ribeiro, S. Singh, and C. Guestrin. "Why Should I Trust You?": Explaining the Predictions of Any Classifier. In SIGKDD, 2016.
  48. 48.R. R. Selvaraju, P. Chattopadhyay, M. Elhoseiny, T. Sharma, D. Batra, D. Parikh, and S. Lee. Choose your neuron: Incorporating domain knowledge through neuron-importance. In Proceedings of the European Conference on Computer Vision (ECCV), pages 526–541, 2018.
  49. 49.R. R. Selvaraju, S. Lee, Y. Shen, H. Jin, S. Ghosh, L. Heck, D. Batra, and D. Parikh. Taking a hint: Leveraging explanations to make vision and language models more grounded. In Proceedings of the International Conference on Computer Vision (ICCV), 2019.
  50. 50.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  51. 51.K. Simonyan, A. Vedaldi, and A. Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. CoRR, abs/1312.6034, 2013.
  52. 52.K. Simonyan and A. Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR, 2015.
  53. 53.J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. A. Riedmiller. Striving for Simplicity: The All Convolutional Net. CoRR, abs/1412.6806, 2014.
  54. 54.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  55. 55.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. In CVPR, 2015.
  56. 56.C. Vondrick, A. Khosla, T. Malisiewicz, and A. Torralba. HOGgles: Visualizing Object Detection Features. ICCV, 2013.
  57. 57.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. In ECCV, 2014.
  58. 58.J. Zhang, Z. Lin, J. Brandt, X. Shen, and S. Sclaroff. Top-down Neural Attention by Excitation Backprop. In ECCV, 2016.
  59. 59.B. Zhou, A. Khosla, L. A., A. Oliva, and A. Torralba. Learning Deep Features for Discriminative Localization. In CVPR, 2016.
  60. 60.B. Zhou, A. Khosla, À. Lapedriza, A. Oliva, and A. Torralba. Object detectors emerge in deep scene cnns. CoRR, abs/1412.6856, 2014.
  61. 61.B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.

Citation

MLA
Selvaraju, R. R., et al. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization”. International Journal of Computer Vision, vol. 128, no. 2, 2019, pp. 336–59, https://doi.org/10.1007/s11263-019-01228-7.
APA
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2019). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision, 128(2), 336–359. https://doi.org/10.1007/s11263-019-01228-7
Chicago
Selvaraju, R. R., M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. 2019. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization”. International Journal of Computer Vision 128 (2): 336–59. https://doi.org/10.1007/s11263-019-01228-7.
Harvard
Selvaraju, R.R. et al. (2019) “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization”, International Journal of Computer Vision, 128(2), pp. 336–359. Available at: https://doi.org/10.1007/s11263-019-01228-7.
Vancouver
1. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2019) Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision 128:336–359

BibTeX

@article{Selvaraju_2019, title={Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization}, volume={128}, ISSN={1573-1405}, url={http://dx.doi.org/10.1007/s11263-019-01228-7}, DOI={10.1007/s11263-019-01228-7}, number={2}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Selvaraju, Ramprasaath R. and Cogswell, Michael and Das, Abhishek and Vedantam, Ramakrishna and Parikh, Devi and Batra, Dhruv}, year={2019}, month=Oct, pages={336–359} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF