A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others
Zhiheng LiIvan EvtimovAlbert GordoCaner HazirbasTal HassnerCristian Canton-FerrerChenliang XuMark Ibrahim
Reveals that mitigating one visual shortcut often amplifies reliance on others across modern vision models, and introduces new multi-shortcut benchmarks alongside a simple ensemble method to tackle this trade-off.
Machine learning vision systems frequently learn shortcuts—spurious correlations such as image backgrounds or textures—instead of true object features, leading to severe reliability failures when deployed in real-world settings. Prior research has predominantly focused on mitigating a single isolated shortcut at a time. The article investigates whether existing mitigation techniques can handle multiple concurrent shortcuts or instead trigger a Whac-A-Mole dilemma, where suppressing one spurious cue inadvertently amplifies reliance on another. To address this, the authors introduce two new evaluation benchmarks: UrbanCars, a controlled dataset featuring both background and co-occurring object shortcuts, and ImageNet-Watermark (ImageNet-W), an out-of-distribution evaluation suite based on the discovery that standard models rely heavily on transparent watermarks to predict certain classes (such as cartons).
The evaluation covers standard supervised models (e.g., ResNet-50), self-supervised approaches, large foundation models (such as CLIP and SWAG), and specialized debiasing algorithms. The investigation yields critical findings. First, standard vision models universally exploit multiple shortcuts simultaneously; on UrbanCars, standard training suffers a 69.2% accuracy drop when both background and co-occurring cues are altered, while on ImageNet-W, models experience an average top-1 accuracy drop of 10.7% (reaching up to a 26.7% drop on ResNet-50). Second, current mitigation techniques—including data augmentations, group-weighting methods using shortcut labels, and pseudo-label inference—routinely exhibit Whac-A-Mole behavior by reducing reliance on one shortcut while worsening sensitivity to others (for example, CutMix increases background shortcut sensitivity by roughly 2.9 times on UrbanCars). Third, large-scale pretraining on billions of web images does not resolve this flaw, as models like zero-shot CLIP inherit watermark biases directly from web data (e.g., LAION).
These findings demonstrate that evaluating and optimizing computer vision models under the assumption of a single shortcut provides a false sense of security. In high-stakes applications such as automated driving or medical diagnosis, mitigating one visual bias with standard interventions can silently elevate model vulnerability to other environmental shifts. To address this limitation without requiring expensive manual shortcut annotations, the authors propose Last Layer Ensemble (LLE). This approach trains lightweight, shift-specific classification heads atop a shared feature extractor paired with a dynamic shift predictor, allowing the system to suppress multiple shortcuts concurrently without causing cross-shortcut interference or substantial computational overhead.
Organizations developing or deploying safety-critical vision systems should immediately abandon single-shortcut evaluations in favor of multi-shortcut benchmarking suites. Teams should prioritize scalable multi-shift approaches like Last Layer Ensemble over simple single-attribute debiasing or uncontrolled data augmentations. While confidence in these empirical results is high across diverse architectures and datasets, the authors note that the proposed mitigation framework still requires knowing the general types of shortcuts in advance. Future efforts must focus on developing automated methods to detect and mitigate unknown multi-shortcut sets and providing formal theoretical foundations for multi-cue interference in deep learning representations.
- Paper: Shortcut learning in deep neural networks, Robert Geirhos et al. (2020). This foundational perspective unifies deep learning failure modes under the taxonomy of shortcut learning, establishing the primary conceptual baseline that the source paper expands from single-shortcut to multi-shortcut regimes.
- Paper: ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness, Robert Geirhos et al. (2018). This paper establishes that ImageNet-trained models suffer from severe texture bias rather than shape bias, providing one of the fundamental visual shortcuts directly analyzed in the source paper's benchmark.
- Paper: Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al. (2019). This work formulates group distributionally robust optimization to tackle spurious correlations (such as background shortcuts in Waterbirds), representing the core category of mitigation methods evaluated and shown to exhibit Whac-A-Mole behavior in the source.
- Paper: Correct-N-Contrast: a Contrastive Approach for Improving Robustness to Spurious Correlations, Michael Zhang et al. (2022). This paper introduces a contrastive learning framework to mitigate spurious correlations without training attribute labels, serving as a key representative defense against spurious shortcuts tested under multi-shortcut settings.
- Paper: Learning Robust Global Representations by Penalizing Local Predictive Power, Haohan Wang et al. (2019). This study introduces patch-wise adversarial regularization to force models away from local predictive cues (like textures and local patches) toward global representations, providing important context on regularizing localized visual shortcuts.
- Paper: Unbiased look at dataset bias, Antonio Torralba et al. (2011). This seminal work demonstrates how visual recognition datasets carry distinct contextual signatures and biases, motivating why real-world image datasets induce multiple unintended shortcuts.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). This benchmark standardizes in-the-wild distribution and subpopulation shifts, providing essential framing for evaluating spurious feature dependencies in natural distributions.
- Paper: Natural Adversarial Examples, Dan Hendrycks et al. (2019). This paper curates ImageNet-A to expose model reliance on superficial background and texture cues, forming key background for creating natural shortcut evaluation benchmarks like ImageNet-W.
- Paper: Discovering and Mitigating Visual Biases Through Keyword Explanation, Younghyun Kim et al. (2024). This work develops the Bias-to-Text framework to automatically discover and mitigate unexplainable visual biases using vision-language models, advancing beyond predefined shortcut mitigations.
- Paper: Not All Neuro-Symbolic Concepts Are Created Equal: Analysis and Mitigation of Reasoning Shortcuts, Emanuele Marconato et al. (2023). This paper extends the study and mitigation of unintended shortcuts into neuro-symbolic reasoning systems where models achieve task success via flawed intermediate concept representations.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). This work generalizes bias and shortcut discovery to an open-set paradigm in text-to-image generative models without requiring a priori specified visual shortcuts.
