Learning to Estimate Shapley Values with Vision Transformers

Ian Connick CovertChanwoo KimSu-In Lee

article2023ICLR58 citations

Presents a scalable approach to generate Shapley value explanations for Vision Transformers by combining attention masking with a trained explainer model, achieving higher attribution accuracy than existing attention- and gradient-based interpretability techniques.

Listen

Vision transformers (ViTs) have become a leading architecture in computer vision tasks, such as medical imaging and general object recognition. However, understanding exactly which image features drive their predictions remains a persistent operational challenge. Existing explanation tools, such as attention map visualizations and gradient-based methods, frequently provide misleading or incomplete attributions and struggle to generate class-specific explanations. While Shapley values provide a theoretically grounded framework to measure feature importance, computing them traditionally requires exponential running time, making them computationally impractical for high-dimensional vision models.

The article's main objective is to establish a practical, accurate, and scalable framework for calculating Shapley value feature explanations in vision transformers. To achieve this, the authors develop and evaluate ViT Shapley, an approach that combines attention masking for missing features with a dedicated explainer model that estimates Shapley values in a single forward pass.

The approach operates in two main stages. First, the authors introduce an attention masking technique to evaluate ViTs with partial image inputs, demonstrating that fine-tuning models with random masking allows them to process missing image patches reliably without requiring off-manifold data replacements. Second, they train a separate ViT explainer model using a specialized weighted least-squares loss that minimizes the Shapley estimation error without needing precomputed ground-truth explanations. The authors benchmarked ViT Shapley against multiple attention-, gradient-, and removal-based baseline methods across three standard datasets (ImageNette, MURA medical radiographs, and Oxford-IIIT Pets) using metrics such as feature insertion, feature deletion, sensitivity, faithfulness, and accuracy degradation benchmarks.

The analysis produced several key findings. First, ViT Shapley consistently outperformed all baseline methods across all datasets, achieving superior performance on feature insertion and deletion curves; for instance, in the MURA abnormal class deletion test, ViT Shapley dropped the prediction area under the curve to 0.307, outperforming the closest baselines which remained between 0.537 and 0.823. Second, ViT Shapley demonstrated a distinctive capability to generate precise, class-specific explanations for both target and non-target classes, where standard baselines generally failed to differentiate class-specific features. Third, once trained, the explainer model generated complete patch-level explanations in roughly 10 milliseconds—a single forward pass—matching the speed of basic attention methods while delivering accuracy equivalent to running traditional sampling-based Shapley estimation methods for roughly 120,000 model evaluations.

These findings indicate that organizations deploying vision transformers in high-stakes environments, such as medical diagnostics or critical inspection systems, no longer need to rely on unreliable attention visualizations or prohibitively slow credit attribution algorithms. ViT Shapley delivers reliable, mathematically grounded interpretability without incurring high latency costs during live model inference. This provides decision-makers with a robust mechanism for model auditing, identifying potential confounders, and satisfying safety and compliance mandates in real time.

Organizations seeking to implement this methodology should prioritize fine-tuning their vision models with random attention masking or training explainer models on pre-trained backbones rather than training from scratch. For high-throughput production environments, teams should invest the upfront computational time—ranging from hours to a few days—to train the explainer model to secure millisecond-level inference times during deployment. Future exploratory work should expand this architecture to multimodal and natural language processing tasks, as well as test explanations grouped by arbitrary superpixels.

In terms of limitations, the method requires an upfront computational investment to train the explainer network (roughly 19 to 60 GPU hours in the reported setups), and the resulting scores remain a learned approximation rather than exact Shapley values. Nevertheless, because the explainer's training objective mathematically bounds the estimation error, stakeholders can maintain high confidence in the relative accuracy and practical utility of the generated explanations.

arXiv: 2206.05282
Cover for Learning to Estimate Shapley Values with Vision Transformers

Abstract

Transformers have become a default architecture in computer vision, but understanding what drives their predictions remains a challenging problem. Current explanation approaches rely on attention values or input gradients, but these provide a limited view of a model's dependencies. Shapley values offer a theoretically sound alternative, but their computational cost makes them impractical for large, high-dimensional models. In this work, we aim to make Shapley values practical for vision transformers (ViTs). To do so, we first leverage an attention masking approach to evaluate ViTs with partial information, and we then develop a procedure to generate Shapley value explanations via a separate, learned explainer model. Our experiments compare Shapley values to many baseline methods (e.g., attention rollout, GradCAM, LRP), and we find that our approach provides more accurate explanations than existing methods for ViTs.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Background
  • 3.1 Notation
  • 3.2 Shapley values
  • 4 Evaluating vision transformers with partial information
  • 5 Learning to estimate Shapley values
  • 6 Experiments
  • 6.1 Evaluating image patch removal
  • 6.2 Evaluating explanation accuracy
  • 7 Conclusion
  • Bibliography
  • A Attention masking
  • B Masked training
  • C Explainer training approach
  • C.1 Hyperparameter choices
  • D Proofs
  • E Datasets
  • F Baseline methods
  • G Metrics details
  • H Additional results
  • H.1 Main baselines and metrics
  • H.2 KernelSHAP comparisons
  • I Qualitative examples

Citation

MLA
Covert, I., et al. “Learning to Estimate Shapley Values with Vision Transformers”. arXiv, 2022, http://arxiv.org/abs/2206.05282v3.
APA
Covert, I., Kim, C., & Lee, S.-I. (2022). Learning to Estimate Shapley Values with Vision Transformers. arXiv. http://arxiv.org/abs/2206.05282v3
Chicago
Covert, I., C. Kim, and S.-I. Lee. 2022. “Learning to Estimate Shapley Values with Vision Transformers”. arXiv. http://arxiv.org/abs/2206.05282v3.
Harvard
Covert, I., Kim, C. and Lee, S.-I. (2022) “Learning to Estimate Shapley Values with Vision Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.05282v3.
Vancouver
1. Covert I, Kim C, Lee S-I (2022) Learning to Estimate Shapley Values with Vision Transformers. arXiv

BibTeX

@article{covert2022learning,
  title = {Learning to Estimate Shapley Values with Vision Transformers},
  author = {Covert, Ian and Kim, Chanwoo and Lee, Su-In},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.05282v3},
  eprint = {2206.05282}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors