RISE: Randomized Input Sampling for Explanation of Black-box Models
Vitali PetsiukAbir DasKate Saenko
Introduces a black-box explanation method that estimates visual saliency by probing image classifiers with randomly masked inputs, matching or outperforming gradient-based white-box approaches without requiring internal model access.
Deep neural networks are widely used to automate complex decision-making, yet their internal reasoning remains opaque to users. In high-stakes settings such as medical diagnosis and autonomous systems, this lack of transparency introduces significant operational and safety risks because stakeholders cannot easily verify why a model made a specific prediction or whether it relied on erroneous cues. Most existing explanation tools require direct access to internal network architectures, mathematical gradients, or weights, which limits their applicability across proprietary or diverse systems.
To address this challenge, the article introduces Randomized Input Sampling for Explanation (RISE), an approach designed to generate visual pixel-importance maps for any image-classification model without accessing its internal architecture. The article demonstrates how treating models as complete black boxes allows practitioners to produce causal visual explanations across arbitrary network designs and evaluate their quality using human-independent metrics.
RISE operates by probing a target model with thousands of randomly masked versions of an input image and recording the corresponding shifts in output confidence scores. The method computes an importance map by taking a weighted linear combination of these masks based on the model's scores. To evaluate this approach, the authors tested RISE against leading explanation techniques on standard benchmark datasets, including ImageNet, PASCAL VOC, and MS COCO, utilizing both human-annotated pointing benchmarks and automated insertion and deletion metrics that measure how prediction probabilities change as important pixels are systematically added or removed.
Across multiple benchmarks, RISE consistently matched or outperformed existing methods. On ImageNet, RISE achieved the best causal performance scores for both ResNet50 and VGG16 architectures, surpassing popular white-box approaches such as Grad-CAM and black-box tools like LIME. On the PASCAL VOC pointing benchmark with VGG16, RISE achieved an accuracy of approximately 87.3%, markedly outperforming alternative methods that ranged between 75% and 80%. Furthermore, the article demonstrated that RISE readily generalizes to more complex vision tasks, such as generating word-by-word visual groundings for automated image captioning systems.
These findings show that effective model interpretability does not require invasive access to internal neural parameters. Organizations can deploy a unified, architecture-agnostic audit tool across third-party and proprietary models alike, improving compliance, safety, and operational trust. Because RISE exposes true model reasoning rather than human-biased assumptions, it helps engineers identify unexpected failure modes, such as models relying on background context rather than the primary subject.
Leaders evaluating explainability frameworks should consider adopting input-sampling methods when deploying diverse black-box models. However, because RISE requires thousands of forward model evaluations per image (such as 4,000 to 8,000 passes), it introduces substantial computational overhead that currently limits real-time deployment. Future implementation efforts should focus on optimizing sampling efficiency to reduce query volume and mitigating visual background noise caused by finite sampling approximations before deploying the method in latency-critical environments.
- Paper: Interpretable Explanations of Black Boxes by Meaningful Perturbation, Ruth Fong et al. (2017). This paper establishes the perturbation-based optimization framework and deletion game concepts that directly motivate RISE's black-box masking and evaluation metrics.
- Paper: “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Marco Tulio Ribeiro et al. (2016). This foundational work introduces model-agnostic local explanations via input perturbations, laying the conceptual groundwork for RISE's black-box sampling approach.
- Paper: Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization, Ramprasaath R. Selvaraju et al. (2016). This paper provides the standard white-box visual explanation baseline (Grad-CAM) against which RISE directly compares its black-box saliency maps.
- Paper: Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps, Karen Simonyan et al. (2013). This seminal paper introduces gradient-based pixel saliency maps for deep convolutional networks, establishing the visual attribution paradigm that RISE seeks to solve without internal access.
- Paper: Learning Deep Features for Discriminative Localization, Bolei Zhou et al. (2016). This paper introduces Class Activation Mapping for discriminative localization in CNNs, setting a critical benchmark for visual explanation maps.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). This work defines an axiomatic attribution framework (Integrated Gradients) that serves as a primary white-box attribution comparator for RISE.
- Paper: Learning Important Features Through Propagating Activation Differences, Avanti Shrikumar et al. (2017). This chapter introduces DeepLIFT, a prominent white-box feature attribution technique evaluated alongside RISE's empirical sampling.
- Paper: SmoothGrad: removing noise by adding noise, Daniel Smilkov et al. (2017). This paper develops SmoothGrad, demonstrating how averaging across perturbed inputs sharpens visual saliency, a principle related to RISE's Monte Carlo mask aggregation.
- Paper: Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks, Aditya Chattopadhyay et al. (2017). This work extends gradient-based localization to multiple instances, providing an advanced baseline for evaluating the precision of visual explanations.
- Paper: A Unified Approach to Interpreting Model Predictions, Scott M. Lundberg et al. (2017). This work unifies feature attribution through game-theoretic Shapley values, providing essential theoretical context for perturbation-based attribution methods.
- Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). This paper introduces fundamental sanity checks to test whether saliency methods like RISE are genuinely sensitive to model parameters and data labels.
- Paper: Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI, Alejandro Barredo Arrieta et al. (2020). This comprehensive survey categorizes post-hoc explainability methods, placing black-box perturbation approaches like RISE within the broader responsible AI taxonomy.
- Paper: A Survey of Methods for Explaining Black Box Models, Riccardo Guidotti et al. (2018). This survey provides an extensive taxonomy of black-box model explanations, contextualizing input-sampling saliency techniques within broader black-box opening strategies.
- Paper: Explaining Explanations: An Overview of Interpretability of Machine Learning, Leilani H. Gilpin et al. (2018). This overview develops a unified taxonomy distinguishing interpretability from explainability, framing the role and limitations of post-hoc attribution methods like RISE.
- Paper: GNNExplainer: Generating Explanations for Graph Neural Networks, Rex Ying et al. (2019). This paper extends perturbation- and masking-based explanation strategies from grid-structured image inputs to non-Euclidean graph neural networks.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This work critiques attention heatmaps as faithful explanations by evaluating them with perturbation and erasure tests analogous to RISE's deletion metrics.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This chapter critiques the statistical fragility and non-identifiability of visual attribution methods, providing a modern statistical framework to re-evaluate saliency reliability.
