Learning to Estimate Shapley Values with Vision Transformers
Ian Connick CovertChanwoo KimSu-In Lee
Presents a scalable approach to generate Shapley value explanations for Vision Transformers by combining attention masking with a trained explainer model, achieving higher attribution accuracy than existing attention- and gradient-based interpretability techniques.
Vision transformers (ViTs) have become a leading architecture in computer vision tasks, such as medical imaging and general object recognition. However, understanding exactly which image features drive their predictions remains a persistent operational challenge. Existing explanation tools, such as attention map visualizations and gradient-based methods, frequently provide misleading or incomplete attributions and struggle to generate class-specific explanations. While Shapley values provide a theoretically grounded framework to measure feature importance, computing them traditionally requires exponential running time, making them computationally impractical for high-dimensional vision models.
The article's main objective is to establish a practical, accurate, and scalable framework for calculating Shapley value feature explanations in vision transformers. To achieve this, the authors develop and evaluate ViT Shapley, an approach that combines attention masking for missing features with a dedicated explainer model that estimates Shapley values in a single forward pass.
The approach operates in two main stages. First, the authors introduce an attention masking technique to evaluate ViTs with partial image inputs, demonstrating that fine-tuning models with random masking allows them to process missing image patches reliably without requiring off-manifold data replacements. Second, they train a separate ViT explainer model using a specialized weighted least-squares loss that minimizes the Shapley estimation error without needing precomputed ground-truth explanations. The authors benchmarked ViT Shapley against multiple attention-, gradient-, and removal-based baseline methods across three standard datasets (ImageNette, MURA medical radiographs, and Oxford-IIIT Pets) using metrics such as feature insertion, feature deletion, sensitivity, faithfulness, and accuracy degradation benchmarks.
The analysis produced several key findings. First, ViT Shapley consistently outperformed all baseline methods across all datasets, achieving superior performance on feature insertion and deletion curves; for instance, in the MURA abnormal class deletion test, ViT Shapley dropped the prediction area under the curve to 0.307, outperforming the closest baselines which remained between 0.537 and 0.823. Second, ViT Shapley demonstrated a distinctive capability to generate precise, class-specific explanations for both target and non-target classes, where standard baselines generally failed to differentiate class-specific features. Third, once trained, the explainer model generated complete patch-level explanations in roughly 10 milliseconds—a single forward pass—matching the speed of basic attention methods while delivering accuracy equivalent to running traditional sampling-based Shapley estimation methods for roughly 120,000 model evaluations.
These findings indicate that organizations deploying vision transformers in high-stakes environments, such as medical diagnostics or critical inspection systems, no longer need to rely on unreliable attention visualizations or prohibitively slow credit attribution algorithms. ViT Shapley delivers reliable, mathematically grounded interpretability without incurring high latency costs during live model inference. This provides decision-makers with a robust mechanism for model auditing, identifying potential confounders, and satisfying safety and compliance mandates in real time.
Organizations seeking to implement this methodology should prioritize fine-tuning their vision models with random attention masking or training explainer models on pre-trained backbones rather than training from scratch. For high-throughput production environments, teams should invest the upfront computational time—ranging from hours to a few days—to train the explainer model to secure millisecond-level inference times during deployment. Future exploratory work should expand this architecture to multimodal and natural language processing tasks, as well as test explanations grouped by arbitrary superpixels.
In terms of limitations, the method requires an upfront computational investment to train the explainer network (roughly 19 to 60 GPU hours in the reported setups), and the resulting scores remain a learned approximation rather than exact Shapley values. Nevertheless, because the explainer's training objective mathematically bounds the estimation error, stakeholders can maintain high confidence in the relative accuracy and practical utility of the generated explanations.
- Paper: A Unified Approach to Interpreting Model Predictions, Scott M. Lundberg et al. (2017). This paper establishes the foundational SHAP framework and game-theoretic axiomatic principles that the source explicitly adapts and accelerates for vision transformers.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This seminal work introduces the Vision Transformer (ViT) patch-based architecture that serves as the core subject and explainer backbone of the source study.
- Paper: RISE: Randomized Input Sampling for Explanation of Black-box Models, Vitali Petsiuk et al. (2018). This work introduces randomized input masking and insertion/deletion evaluation protocols for visual attribution, which the source builds upon and benchmarks against.
- Paper: Interpretable Explanations of Black Boxes by Meaningful Perturbation, Ruth Fong et al. (2017). This study introduces perturbation- and deletion-based attribution for image classifiers, providing key conceptual foundations and evaluation metrics used in the source.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). This paper highlights the limitations of using raw attention weights as reliable feature attributions in deep transformer models, establishing the problem the source aims to solve.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). This work establishes axiomatic attribution requirements for deep networks, supplying the formal interpretability standards against which the source formulates its Shapley approach.
- Paper: Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization, Ramprasaath R. Selvaraju et al. (2016). This work introduces Grad-CAM, a widely used visual localization baseline that the source directly evaluates against and aims to surpass in class-specificity and faithfulness.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). This work details masked image modeling for vision transformers, providing relevant context for how ViTs process partially masked image inputs.
- Paper: Faith-Shap: The Faithful Shapley Interaction Index, Che-Ping Tsai et al. (2023). This work extends Shapley-based estimation and weighted least-squares regression from individual feature importance to axiomatic, higher-order feature interactions.
- Paper: A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation, Thomas Fel et al. (2023). This paper generalizes feature attribution metrics like insertion and deletion beyond pixel/patch space to high-level semantic concept spaces in deep vision models.
- Paper: Vision Transformers Need Registers, Timothée Darcet et al. (2024). This paper investigates internal artifact tokens that corrupt visual attention maps and feature representations in large ViTs, providing an architectural remedy directly relevant to attention-based explainability.
