RISE: Randomized Input Sampling for Explanation of Black-box Models

Vitali PetsiukAbir DasKate Saenko

article2018BMVC1,624 citations

Introduces a black-box explanation method that estimates visual saliency by probing image classifiers with randomly masked inputs, matching or outperforming gradient-based white-box approaches without requiring internal model access.

Listen

Deep neural networks are widely used to automate complex decision-making, yet their internal reasoning remains opaque to users. In high-stakes settings such as medical diagnosis and autonomous systems, this lack of transparency introduces significant operational and safety risks because stakeholders cannot easily verify why a model made a specific prediction or whether it relied on erroneous cues. Most existing explanation tools require direct access to internal network architectures, mathematical gradients, or weights, which limits their applicability across proprietary or diverse systems.

To address this challenge, the article introduces Randomized Input Sampling for Explanation (RISE), an approach designed to generate visual pixel-importance maps for any image-classification model without accessing its internal architecture. The article demonstrates how treating models as complete black boxes allows practitioners to produce causal visual explanations across arbitrary network designs and evaluate their quality using human-independent metrics.

RISE operates by probing a target model with thousands of randomly masked versions of an input image and recording the corresponding shifts in output confidence scores. The method computes an importance map by taking a weighted linear combination of these masks based on the model's scores. To evaluate this approach, the authors tested RISE against leading explanation techniques on standard benchmark datasets, including ImageNet, PASCAL VOC, and MS COCO, utilizing both human-annotated pointing benchmarks and automated insertion and deletion metrics that measure how prediction probabilities change as important pixels are systematically added or removed.

Across multiple benchmarks, RISE consistently matched or outperformed existing methods. On ImageNet, RISE achieved the best causal performance scores for both ResNet50 and VGG16 architectures, surpassing popular white-box approaches such as Grad-CAM and black-box tools like LIME. On the PASCAL VOC pointing benchmark with VGG16, RISE achieved an accuracy of approximately 87.3%, markedly outperforming alternative methods that ranged between 75% and 80%. Furthermore, the article demonstrated that RISE readily generalizes to more complex vision tasks, such as generating word-by-word visual groundings for automated image captioning systems.

These findings show that effective model interpretability does not require invasive access to internal neural parameters. Organizations can deploy a unified, architecture-agnostic audit tool across third-party and proprietary models alike, improving compliance, safety, and operational trust. Because RISE exposes true model reasoning rather than human-biased assumptions, it helps engineers identify unexpected failure modes, such as models relying on background context rather than the primary subject.

Leaders evaluating explainability frameworks should consider adopting input-sampling methods when deploying diverse black-box models. However, because RISE requires thousands of forward model evaluations per image (such as 4,000 to 8,000 passes), it introduces substantial computational overhead that currently limits real-time deployment. Future implementation efforts should focus on optimizing sampling efficiency to reduce query volume and mitigating visual background noise caused by finite sampling approximations before deploying the method in latency-critical environments.

Cover for RISE: Randomized Input Sampling for Explanation of Black-box Models

Abstract

Deep neural networks are being used increasingly to automate data analysis and decision making, yet their decision-making process is largely unclear and is difficult to explain to the end users. In this paper, we address the problem of Explainable AI for deep neural networks that take images as input and output a class probability. We propose an approach called RISE that generates an importance map indicating how salient each pixel is for the model's prediction. In contrast to white-box approaches that estimate pixel importance using gradients or other internal network state, RISE works on black-box models. It estimates importance empirically by probing the model with randomly masked versions of the input image and obtaining the corresponding outputs. We compare our approach to state-of-the-art importance extraction methods using both an automatic deletion/insertion metric and a pointing metric based on human-annotated object segments. Extensive experiments on several benchmark datasets show that our approach matches or exceeds the performance of other methods, including white-box approaches.

Project page: this http URL

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Randomized Input Sampling for Explanation (RISE)
  • 3.1 Random Masking
  • 3.2 Mask generation
  • 4 Experiments
  • 4.1 Evaluation Metrics
  • 4.2 Experimental Results
  • 4.3 RISE for Captioning
  • 5 Conclusion
  • References
  • A Algorithms to compute causal metrics
  • B More saliency maps and their scores

Knowls

  1. Knowl 1 — Randomized Input Sampling for Explanation Importance Map Formulation

    model/method

    Let f:I→Rf : \mathcal{I} \to \mathbb{R} be a black-box model mapping an input image I∈II \in \mathcal{I} of spatial dimensions H×WH \times W (spatial grid Λ={1,…,H}×{1,…,W}\Lambda = \{1, \dots, H\} \times \{1, \dots, W\}) to a scalar confidence score for a specified target class. Let M:Λ→[0,1]M : \Lambda \to [0, 1] be a random mask drawn from a probability distribution D\mathcal{D}.

    The visual importance (saliency) SI,f(λ)S_{I, f}(\lambda) of a pixel location λ∈Λ\lambda \in \Lambda is defined as the expected class score conditioned on the event that pixel λ\lambda is unmasked (M(λ)=1M(\lambda) = 1):

    SI,f(λ)=EM[f(I⊙M)∣M(λ)=1]S_{I, f}(\lambda) = \mathbb{E}_M [f(I \odot M) \mid M(\lambda) = 1]

    where ⊙\odot denotes element-wise multiplication.

    Expressed in matrix form across all mask configurations mm:

    SI,f=1E[M]∑mf(I⊙m)⋅m⋅P[M=m]S_{I, f} = \frac{1}{\mathbb{E}[M]} \sum_m f(I \odot m) \cdot m \cdot P[M = m]

    where E[M]=P[M(λ)=1]\mathbb{E}[M] = P[M(\lambda) = 1] is the expectation of the mask elements under distribution D\mathcal{D}.

    RISE computes this importance map empirically using Monte Carlo sampling over NN randomly generated masks {M1,…,MN}∼D\{M_1, \dots, M_N\} \sim \mathcal{D}:

    SI,f(λ)≈1E[M]⋅N∑i=1Nf(I⊙Mi)⋅Mi(λ)S_{I, f}(\lambda) \approx \frac{1}{\mathbb{E}[M] \cdot N} \sum_{i=1}^N f(I \odot M_i) \cdot M_i(\lambda)

    This method requires only black-box input-output access to ff without requiring internal architecture details, intermediate activation maps, or gradients.

  2. Knowl 2 — Random Mask Generation Procedure in RISE

    algorithm

    Generating masks by independently sampling individual pixel values directly at image resolution H×WH \times W causes high-frequency adversarial artifacts and induces a prohibitively large mask space of size 2H×W2^{H \times W}. RISE avoids sharp edges and reduces the sample space by first generating lower-resolution binary masks, upsampling them bilinearly, and applying random spatial crops:

    procedure GENERATE_MASKS(N, H, W, h, w, p)
        Input: Number of masks NN, image dimensions H×WH \times W, low-resolution grid size h×wh \times w, Bernoulli probability pp
        Output: Set of continuous masks {M1,…,MN}\{M_1, \dots, M_N\} of size H×WH \times W with values in [0,1][0, 1]
        CH←⌊H/h⌋C_H \leftarrow \lfloor H / h \rfloor
        CW←⌊W/w⌋C_W \leftarrow \lfloor W / w \rfloor
        for i=1i = 1 to NN do
            Sample binary mask m∈{0,1}h×wm \in \{0, 1\}^{h \times w} with each element set to 11 with probability pp and 00 with 1−p1-p
            Upsample mm to size (h+1)CH×(w+1)CW(h+1)C_H \times (w+1)C_W using bilinear interpolation
            Sample spatial crop offsets (Δy,Δx)(\Delta_y, \Delta_x) uniformly at random from [0,CH)×[0,CW)[0, C_H) \times [0, C_W)
            Crop region of size H×WH \times W from the upsampled mask starting at (Δy,Δx)(\Delta_y, \Delta_x) to form MiM_i
        return {M1,…,MN}\{M_1, \dots, M_N\}

    In standard evaluations with images of size H=W=224H = W = 224, the small grid size is set to h=w=7h = w = 7 with p=0.5p = 0.5. Bilinear upsampling creates smooth continuous values in [0,1][0, 1], preventing edge perturbations from dominating the model response.

  3. Knowl 3 — Deletion Metric for Saliency Map Evaluation

    algorithm

    The deletion metric is a causal, human-agnostic evaluation procedure that measures how rapidly a black-box model's confidence for a predicted class decreases as the most salient pixels are progressively removed. A rapid probability drop indicates that the saliency map correctly identified the decisive causal features.

    procedure DELETION(f, I, S, N_step)
        Input: Black-box model ff, image II, importance map SS, number of pixels removed per step NstepN_step
        Output: Deletion score dd (Area Under the Curve, lower is better)
        n←0n \leftarrow 0
        h0←f(I)h_0 \leftarrow f(I)
        while II has non-zero pixels do
            Identify the NstepN_step unremoved pixels with the highest values in SS and set them to 0 in II
            n←n+1n \leftarrow n + 1
            hn←f(I)h_n \leftarrow f(I)
        d←AreaUnderCurve(hi vs. i/n for i=0,…,n)d \leftarrow \text{AreaUnderCurve}(h_i \text{ vs. } i/n \text{ for } i = 0, \dots, n)
        return dd

    Pixels are set to a constant value (00) rather than blurred, because deep models can often reconstruct missing low-frequency details from tiny blurred regions.

  4. Knowl 4 — Insertion Metric for Saliency Map Evaluation

    algorithm

    The insertion metric is a causal, human-agnostic evaluation procedure that measures the increase in model prediction confidence as salient pixels are progressively restored into a baseline uninformative image. Higher area under the resulting probability curve (AUC) indicates a more faithful explanation.

    procedure INSERTION(f, I, S, N_step)
        Input: Black-box model ff, image II, importance map SS, number of pixels restored per step NstepN_step
        Output: Insertion score dd (Area Under the Curve, higher is better)
        n←0n \leftarrow 0
        I′←Blur(I)I' \leftarrow \text{Blur}(I)
        h0←f(I′)h_0 \leftarrow f(I')
        while I′≠II' \neq I do
            Identify the NstepN_step unrestored pixels with the highest values in SS
            Set those NstepN_step pixels in I′I' to their corresponding values from II
            n←n+1n \leftarrow n + 1
            hn←f(I′)h_n \leftarrow f(I')
        d←AreaUnderCurve(hi vs. i/n for i=0,…,n)d \leftarrow \text{AreaUnderCurve}(h_i \text{ vs. } i/n \text{ for } i = 0, \dots, n)
        return dd

    The canvas is initialized with a blurred version of the input image, Blur(I)\text{Blur}(I), rather than a blank canvas. This suppresses high-frequency fine details without introducing sharp rectangular or oval boundaries that can act as spurious visual artifacts.

  5. Knowl 5 — Extension of RISE to Autoregressive Image Captioning

    model/method

    RISE can explain the word choices of autoregressive image description models. Let a captioning model predict the conditional probability of emitting the kk-th word wkw_k given an input image II and the sequence of preceding words s=(w1,…,wk−1)s = (w_1, \dots, w_{k-1}):

    f(I,s,wk)=P[wk∣I,w1,…,wk−1]f(I, s, w_k) = P[w_k \mid I, w_1, \dots, w_{k-1}]

    To compute the visual saliency of image pixels with respect to the generation of word wkw_k, RISE probes the captioning model using NN randomly masked inputs I⊙MiI \odot M_i:

    SI,f(wk)=1N⋅E[M]∑i=1Nf(I⊙Mi,s,wk)⋅MiS_{I, f}(w_k) = \frac{1}{N \cdot \mathbb{E}[M]} \sum_{i=1}^N f(I \odot M_i, s, w_k) \cdot M_i

    The sequence ss can be the prefix of a model-generated caption or any arbitrary conditioning sentence. This provides word-level visual grounding for black-box captioners without requiring attention architectures, feature extraction layers, or backward passes.

  6. Knowl 6 — ImageNet Saliency Evaluation under Causal Deletion and Insertion Metrics

    data/table

    Explanation fidelity was evaluated on the validation split of ImageNet across ResNet50 and VGG16 base classifiers using Deletion (Area Under Curve ↓\downarrow, lower is better) and Insertion (Area Under Curve ↑\uparrow, higher is better). RISE results report mean and standard deviation over 3 independent runs:

    Method ResNet50 VGG16
    Deletion ↓\downarrow Insertion ↑\uparrow Deletion ↓\downarrow Insertion ↑\uparrow
    Grad-CAM 0.1232 0.6766 0.1087 0.6149
    Sliding window 0.1421 0.6618 0.1158 0.5917
    LIME 0.1217 0.6940 0.1014 0.6167
    RISE (ours) 0.1076 ±\pm 0.0005 0.7267 ±\pm 0.0006 0.0980 ±\pm 0.0025 0.6663 ±\pm 0.0014

    RISE outperforms both black-box baselines (LIME and Sliding Window) as well as the gradient-based white-box Grad-CAM across both network architectures, achieving smaller deletion scores and larger insertion scores.

  7. Knowl 7 — Pointing Game Accuracy on PASCAL VOC07 and MSCOCO2014

    data/table

    Explanations were evaluated using the pointing game metric on the test split of PASCAL VOC07 and the validation split of MSCOCO2014. A prediction is considered a hit if the point of maximum saliency lies inside the human-annotated bounding box of the target category. Pointing game accuracy is #Hits#Hits+#Misses\frac{\#\text{Hits}}{\#\text{Hits} + \#\text{Misses}} averaged across all target classes. RISE scores are reported as mean ±\pm standard deviation over 3 runs:

    Base model Dataset AM Deconv CAM MWP c-MWP RISE
    VGG16 VOC07 76.00 75.50 - 76.90 80.00 87.33 ±\pm 0.49
    VGG16 MSCOCO 37.10 38.60 - 39.50 49.60 50.71 ±\pm 0.10
    ResNet50 VOC07 65.80 73.00 90.60 80.90 89.20 88.94 ±\pm 0.61
    ResNet50 MSCOCO 30.40 38.20 58.40 46.80 57.40 55.58 ±\pm 0.51

    While all comparison methods (Activation Maximization, Deconvolution, CAM, Meaningful Perturbation / MWP, c-MWP) require access to internal network representations, weights, or gradients, RISE operates purely as a black box. It outperforms all white-box methods on VGG16 and achieves competitive pointing accuracy on ResNet50.

  8. Knowl 8 — Computational Cost and Monte Carlo Approximation Artifacts in RISE

    limitation

    RISE exhibits two key limitations:

    1. High computational overhead: Because it computes importance empirically via Monte Carlo probing, RISE requires thousands of network forward passes per image (for example, N=4000N = 4000 masks for VGG16 and N=8000N = 8000 masks for ResNet50), making it significantly slower than analytic single-pass white-box methods.
    2. Sampling approximation noise: Due to finite Monte Carlo sampling, generated importance maps can retain background noise and may struggle to accurately resolve objects when multiple target objects of disparate sizes appear simultaneously in the scene.

Coverage note — None was omitted; all key contributions including the theoretical derivation of RISE, mask generation algorithm, causal deletion/insertion evaluation metrics and algorithms, captioning extension, empirical benchmark results, and stated limitations have been covered.

References

  1. 1.Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Learning to Compose Neural Networks for Question Answering. In The Conference of the North American Chapter of the Association for Computational Linguistics, pages 1545–1554, 2016.
  2. 2.Chunshui Cao, Xianming Liu, Yi Yang, Yinan Yu, Jiang Wang, Zilei Wang, Yongzhen Huang, Liang Wang, Chang Huang, Wei Xu, et al. Look and think twice: Capturing top-down visual attention with feedback convolutional neural networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 2956–2964, 2015.
  3. 3.Mark W Craven and Jude W Shavlik. Extracting Comprehensible Models from Trained Neural Networks. PhD thesis, University of Wisconsin, Madison, 1996.
  4. 4.Piotr Dabkowski and Yarin Gal. Real Time Image Saliency for Black Box Classifiers. In Neural Information Processing Systems, pages 6970–6979, 2017.
  5. 5.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015.
  6. 6.Mary T Dzindolet, Scott A Peterson, Regina A Pomranky, Linda G Pierce, and Hall P Beck. The Role of Trust in Automation Reliance. International Journal of Human-Computer Studies, 58(6):697–718, 2003.
  7. 7.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The Pascal Visual Object Classes (VOC) Challenge. International journal of computer vision, 88(2):303–338, 2010.
  8. 8.Ruth C Fong and Andrea Vedaldi. Interpretable Explanations of Black Boxes by Meaningful Perturbation. In IEEE International Conference on Computer Vision, Oct 2017.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  10. 10.Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell. Generating Visual Explanations. In European Conference on Computer Vision, pages 3–19, 2016.
  11. 11.Bernease Herman. The Promise and Peril of Human Evaluation for Model Interpretability. In Interpretable ML Symposium, Neural Information Processing Systems, Dec 2017.
  12. 12.Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. Learning to Reason: End-To-End Module Networks for Visual Question Answering. In IEEE International Conference on Computer Vision, Oct 2017.
  13. 13.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755, 2014.
  14. 14.Zachary C Lipton. The Mythos of Model Interpretability. arXiv preprint arXiv:1606.03490, 2016.
  15. 15.Tania Lombrozo. The Structure and Function of Explanations. Trends in Cognitive Sciences, 10(10):464–470, 2006.
  16. 16.Tania Lombrozo. The Instrumental Value of Explanations. Philosophy Compass, 6(8):539–551, 2011.
  17. 17.Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, and Jeff Clune. Synthesizing the Preferred Inputs for Neurons in Neural Networks via Deep Generator Networks. In Neural Information Processing Systems, pages 3387–3395, 2016.
  18. 18.Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. Multimodal Explanations: Justifying Decisions and Pointing to the Evidence. In IEEE Computer Vision and Pattern Recognition, Jun 2018.
  19. 19.Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer W Vaughan, and Hanna Wallach. Manipulating and Measuring Model Interpretability. arXiv preprint arXiv:1802.07810, 2018.
  20. 20.Vasili Ramanishka, Abir Das, Jianming Zhang, and Kate Saenko. Top-Down Visual Saliency Guided by Captions. In IEEE Conference on Computer Vision and Pattern Recognition, July 2017.
  21. 21.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. Why Should I Trust You?: Explaining the Predictions of any Classifier. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144. ACM, 2016.
  22. 22.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  23. 23.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization. In IEEE International Conference on Computer Vision, Oct 2017.
  24. 24.Karen Simonyan and Andrew Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations, May 2015.
  25. 25.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. arXiv preprint arXiv:1312.6034, 2013.
  26. 26.William R Swartout. Producing Explanations and Justifications of Expert Consulting Programs. 1981.
  27. 27.William R Swartout and Johanna D Moore. Explanation in Second Generation Expert Systems. In Second Generation Expert Systems, pages 543–585. Springer, 1993.
  28. 28.Sebastian Thrun. Extracting Rules from Artificial Neural Networks with Distributed Representations. In Advances in Neural Information Processing Systems, pages 505–512, 1995.
  29. 29.Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In International Conference on Machine Learning, pages 2048–2057, 2015.
  30. 30.Jason Yosinski, Jeff Clune, Thomas Fuchs, and Hod Lipson. Understanding Neural Networks Through Deep Visualization. In International Conference on Machine Learning Workshop on Deep Learning, 2014.
  31. 31.Matthew D Zeiler and Rob Fergus. Visualizing and Understanding Convolutional Networks. In European conference on computer vision, pages 818–833. Springer, 2014.
  32. 32.Jianming Zhang, Sarah Adel Bargal, Zhe Lin, Xiaohui Shen Jonathan Brandt, and Stan Sclaroff. Top-down Neural Attention by Excitation Backprop. International Journal of Computer Vision, Dec 2017.
  33. 33.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning Deep Features for Discriminative Localization. In IEEE Computer Vision and Pattern Recognition, pages 2921–2929, 2016.

Citation

MLA
Petsiuk, V., et al. “RISE: Randomized Input Sampling for Explanation of Black-box Models”. arXiv, 2018, http://arxiv.org/abs/1806.07421v3.
APA
Petsiuk, V., Das, A., & Saenko, K. (2018). RISE: Randomized Input Sampling for Explanation of Black-box Models. arXiv. http://arxiv.org/abs/1806.07421v3
Chicago
Petsiuk, V., A. Das, and K. Saenko. 2018. “RISE: Randomized Input Sampling for Explanation of Black-box Models”. arXiv. http://arxiv.org/abs/1806.07421v3.
Harvard
Petsiuk, V., Das, A. and Saenko, K. (2018) “RISE: Randomized Input Sampling for Explanation of Black-box Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1806.07421v3.
Vancouver
1. Petsiuk V, Das A, Saenko K (2018) RISE: Randomized Input Sampling for Explanation of Black-box Models. arXiv

BibTeX

@article{petsiuk2018rise,
  title = {RISE: Randomized Input Sampling for Explanation of Black-box Models},
  author = {Petsiuk, Vitali and Das, Abir and Saenko, Kate},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1806.07421v3},
  eprint = {1806.07421}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors