A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

Zhiheng LiIvan EvtimovAlbert GordoCaner HazirbasTal HassnerCristian Canton-FerrerChenliang XuMark Ibrahim

article2023CVPR112 citations

Reveals that mitigating one visual shortcut often amplifies reliance on others across modern vision models, and introduces new multi-shortcut benchmarks alongside a simple ensemble method to tackle this trade-off.

Listen

Machine learning vision systems frequently learn shortcuts—spurious correlations such as image backgrounds or textures—instead of true object features, leading to severe reliability failures when deployed in real-world settings. Prior research has predominantly focused on mitigating a single isolated shortcut at a time. The article investigates whether existing mitigation techniques can handle multiple concurrent shortcuts or instead trigger a Whac-A-Mole dilemma, where suppressing one spurious cue inadvertently amplifies reliance on another. To address this, the authors introduce two new evaluation benchmarks: UrbanCars, a controlled dataset featuring both background and co-occurring object shortcuts, and ImageNet-Watermark (ImageNet-W), an out-of-distribution evaluation suite based on the discovery that standard models rely heavily on transparent watermarks to predict certain classes (such as cartons).

The evaluation covers standard supervised models (e.g., ResNet-50), self-supervised approaches, large foundation models (such as CLIP and SWAG), and specialized debiasing algorithms. The investigation yields critical findings. First, standard vision models universally exploit multiple shortcuts simultaneously; on UrbanCars, standard training suffers a 69.2% accuracy drop when both background and co-occurring cues are altered, while on ImageNet-W, models experience an average top-1 accuracy drop of 10.7% (reaching up to a 26.7% drop on ResNet-50). Second, current mitigation techniques—including data augmentations, group-weighting methods using shortcut labels, and pseudo-label inference—routinely exhibit Whac-A-Mole behavior by reducing reliance on one shortcut while worsening sensitivity to others (for example, CutMix increases background shortcut sensitivity by roughly 2.9 times on UrbanCars). Third, large-scale pretraining on billions of web images does not resolve this flaw, as models like zero-shot CLIP inherit watermark biases directly from web data (e.g., LAION).

These findings demonstrate that evaluating and optimizing computer vision models under the assumption of a single shortcut provides a false sense of security. In high-stakes applications such as automated driving or medical diagnosis, mitigating one visual bias with standard interventions can silently elevate model vulnerability to other environmental shifts. To address this limitation without requiring expensive manual shortcut annotations, the authors propose Last Layer Ensemble (LLE). This approach trains lightweight, shift-specific classification heads atop a shared feature extractor paired with a dynamic shift predictor, allowing the system to suppress multiple shortcuts concurrently without causing cross-shortcut interference or substantial computational overhead.

Organizations developing or deploying safety-critical vision systems should immediately abandon single-shortcut evaluations in favor of multi-shortcut benchmarking suites. Teams should prioritize scalable multi-shift approaches like Last Layer Ensemble over simple single-attribute debiasing or uncontrolled data augmentations. While confidence in these empirical results is high across diverse architectures and datasets, the authors note that the proposed mitigation framework still requires knowing the general types of shortcuts in advance. Future efforts must focus on developing automated methods to detect and mitigate unknown multi-shortcut sets and providing formal theoretical foundations for multi-cue interference in deep learning representations.

Cover for A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

Abstract

Machine learning models have been found to learn shortcuts—unintended decision rules that are unable to generalize—undermining models' reliability. Previous works address this problem under the tenuous assumption that only a single shortcut exists in the training data. Real-world images are rife with multiple visual cues from background to texture. Key to advancing the reliability of vision systems is understanding whether existing methods can overcome multiple shortcuts or struggle in a Whac-A-Mole game, i.e., where mitigating one shortcut amplifies reliance on others. To address this shortcoming, we propose two benchmarks: 1) UrbanCars, a dataset with precisely controlled spurious cues, and 2) ImageNet-W, an evaluation set based on ImageNet for watermark, a shortcut we discovered affects nearly every modern vision model. Along with texture and background, ImageNet-W allows us to study multiple shortcuts emerging from training on natural images. We find computer vision models, including large foundation models—regardless of training set, architecture, and supervision—struggle when multiple shortcuts are present. Even methods explicitly designed to combat shortcuts struggle in a Whac-A-Mole dilemma. To tackle this challenge, we propose Last Layer Ensemble, a simple-yet-effective method to mitigate multiple shortcuts without Whac-A-Mole behavior. Our results surface multi-shortcut mitigation as an overlooked challenge critical to advancing the reliability of vision systems. The datasets and code are released: https://github.com/facebookresearch/Whac-A-Mole.

Table of Contents

  • 1. Introduction
  • 2. New Datasets for Multi-Shortcut Mitigation
  • 2.1. UrbanCars Dataset
  • 2.2. ImageNet-Watermark (ImageNet-W)
  • 3. Benchmark Methods and Settings
  • 4. Our Approach
  • 5. Experiments
  • 5.1. Standard training relies on multiple shortcuts
  • 5.2. Results: Mitigation Methods
  • 5.3. Results: Self-Supervised & Foundation Models
  • 5.4. Results: Last Layer Ensemble (LLE)
  • 6. Related Work
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — The Whac-A-Mole phenomenon in multi-shortcut mitigation

    definition

    A multi-shortcut mitigation method exhibits a Whac-A-Mole outcome when reducing reliance on one spurious visual cue increases reliance on another, relative to standard empirical-risk minimization (ERM). In this setting, a method can improve robustness to one distributional shift while becoming less robust to a different shift. The paper treats this as a central failure mode of shortcut mitigation rather than an isolated weakness of any particular model or dataset.

  2. Knowl 2 — UrbanCars: a controlled dataset with two correlated shortcuts

    data/table

    UrbanCars is a synthetic-but-photorealistic image dataset for classifying a car’s body type while simultaneously controlling two spurious cues: the scene background (BG) and a co-occurring object (CoObj). Each example is a tuple (xi,yi,bi,ci)(x_i,y_i,b_i,c_i), where xix_i is an image, yiy_i is the target car-body label, bib_i is the background label, and cic_i is the co-occurring-object label. All three labels use the two-class space {urban,country}\{\text{urban},\text{country}\}, producing 23=82^3=8 target–background–object groups.

    The training distribution sets both shortcut correlations to P(b=y∣y)=0.95P(b=y\mid y)=0.95 and P(c=y∣y)=0.95P(c=y\mid y)=0.95, and assumes conditional independence, P(b,c∣y)=P(b∣y)P(c∣y)P(b,c\mid y)=P(b\mid y)P(c\mid y). Consequently, the four combinations of common and uncommon shortcuts occur with frequencies 90.25%90.25\%, 4.75%4.75\%, 4.75%4.75\%, and 0.25%0.25\%. Validation and test sets are balanced with shortcut-match ratios of 0.50.5, so the target cannot be predicted reliably from either shortcut alone.

    Cars are composited from Stanford Cars, backgrounds from Places, and co-occurring objects from LVIS. Urban examples use cars such as sedans or hatchbacks, urban scenes such as alleys or crosswalks, and objects such as fireplugs or stop signs; country examples use trucks or vans, rural roads or fields, and objects such as cows or horses. The construction montage on page 2 illustrates the central car, right-side co-occurring object, and background substitutions.

  3. Knowl 3 — Metrics for measuring separate and joint shortcut robustness

    equation

    For UrbanCars, let A(g)A(g) denote classification accuracy on group gg, and let wgw_g be the frequency of group gg in the training distribution. The in-distribution accuracy is the frequency-weighted group accuracy

    AID=∑gwgA(g).A_{\mathrm{ID}}=\sum_g w_gA(g).

    Let gBGg_{\mathrm{BG}} denote groups with an uncommon background and a common co-occurring object; let gCoObjg_{\mathrm{CoObj}} denote groups with a common background and an uncommon co-occurring object; and let gbothg_{\mathrm{both}} denote groups where both shortcuts are uncommon. The reported gap metrics are

    BG Gap=A(gBG)−AID,\mathrm{BG\ Gap}=A(g_{\mathrm{BG}})-A_{\mathrm{ID}},

    CoObj Gap=A(gCoObj)−AID,\mathrm{CoObj\ Gap}=A(g_{\mathrm{CoObj}})-A_{\mathrm{ID}},

    BG+CoObj Gap=A(gboth)−AID.\mathrm{BG{+}CoObj\ Gap}=A(g_{\mathrm{both}})-A_{\mathrm{ID}}.

    The gaps are usually negative because they measure accuracy drops; values closer to zero indicate greater robustness. The first two gaps isolate each shortcut, while the joint gap measures performance when neither shortcut is available. The paper argues that conventional worst-group accuracy is insufficient because it primarily reflects the group where both shortcuts are uncommon and does not reveal which individual shortcut a method has amplified.

  4. Knowl 4 — ImageNet-W exposes a pervasive watermark shortcut

    data/table

    ImageNet-Watermark (ImageNet-W or IN-W) is an out-of-distribution evaluation set formed by overlaying a transparent central watermark, “捷径捷径捷径”, on every ImageNet validation image; “捷径” means “shortcut” in Chinese. The watermark mimics a cue found in many ImageNet training images from the carton class. ResNet-50 saliency visualizations show that the watermark region is used to predict carton, even though no validation-set carton image contains the original watermark. The ImageNet-W construction illustration on page 3 shows that models can respond to the presence of a watermark even when its exact content differs from the training watermark.

    The principal metrics are the IN-W Gap, defined as ImageNet-W accuracy minus clean ImageNet validation accuracy, and the Carton Gap, defined as the increase in carton-class accuracy from clean ImageNet validation to ImageNet-W. Smaller absolute drops and smaller positive carton increases indicate lower watermark reliance.

    The paper’s model comparison on page 4 reports that the evaluated models have an average clean ImageNet accuracy of 78.6%78.6\%, an average IN-W Gap of −10.7-10.7 percentage points, and an average Carton Gap of +26.7+26.7 percentage points. Supervised ResNet-50 has a −26.7-26.7-point IN-W Gap and a +40+40-point Carton Gap, while the largest reported Carton Gap is +52+52 points. Even large foundation models and zero-shot CLIP models retain the shortcut: CLIP has Carton Gaps of +12+12 to +16+16 points depending on pretraining data. Thus, larger architectures, additional data, and different supervision reduce but do not eliminate watermark reliance.

  5. Knowl 5 — Multi-shortcut evaluation protocol across UrbanCars and ImageNet

    experimental setup

    The paper evaluates standard ERM and shortcut-mitigation methods under four levels of shortcut information: (1) general augmentation or regularization without shortcut knowledge, including Mixup, CutMix, Cutout, AugMix, and Gradient Starvation; (2) targeted augmentation that knows shortcut types but not shortcut labels, including counterfactual augmentation, style-transfer texture augmentation, background augmentation, and watermark augmentation; (3) methods with image-level ground-truth shortcut labels, including gDRO, DI, SUBG, and DFR; and (4) methods that infer image-level pseudo-shortcut labels, including LfF, JTT, EIIL, and DebiAN.

    UrbanCars uses ResNet-50 and selects the early-stopping epoch using validation worst-group accuracy; all methods except DFR use end-to-end training. ImageNet evaluations use ResNet-50 with a frozen feature extractor and retraining of the final classification layer, and additionally evaluate self-supervised and foundation-model representations.

    ImageNet shortcut reliance is measured for three naturally occurring cues. Background robustness uses the IN-9 Gap, the accuracy drop from Mixed-Same to Mixed-Rand; texture robustness uses SIN Gap, the top-1 drop from ImageNet to Stylized ImageNet, and IN-R Gap, the drop from the 200-class ImageNet subset to ImageNet-R; watermark robustness uses IN-W Gap and Carton Gap.

  6. Knowl 6 — Last Layer Ensemble for jointly mitigating known shortcuts

    model/method

    Last Layer Ensemble (LLE) mitigates multiple known shortcut types without requiring shortcut labels. Let KK be the number of known shortcuts, let A1,…,AKA_1,\ldots,A_K be augmentations that modify those shortcut cues, and let A0=IA_0=I be the identity transformation. A shared feature extractor fθf_\theta feeds K+1K+1 target classifiers hdh_d, where classifier hdh_d is trained only on images transformed by AdA_d. For an input xx, classifier dd produces target logits zd(x)=hd(fθ(Ad(x)))z_d(x)=h_d(f_\theta(A_d(x))).

    LLE also trains a distribution-shift classifier gϕg_\phi to predict which shift is present, producing qd(x)=P(d^=d∣x)q_d(x)=P(\hat d=d\mid x). At inference, the target logits are dynamically aggregated according to the predicted shift:

    p(y^∣x)=softmax⁡(∑d=0Kqd(x)zd(x)).p(\hat y\mid x)=\operatorname{softmax}\left(\sum_{d=0}^{K}q_d(x)z_d(x)\right).

    The distribution-shift classifier therefore gives greater weight to the target classifier trained for the shift that appears in the test image, rather than forcing one classifier to learn invariance across potentially incompatible augmentations. When the feature extractor is trainable, gradients from the distribution-shift classifier are stopped before reaching fθf_\theta, preventing that classifier from teaching the feature extractor shortcut information. LLE adds only classification layers and one shift classifier on top of a shared feature extractor, avoiding the inference and parameter cost of an ensemble of complete networks. The architecture diagram on page 5 depicts the shared extractor, per-augmentation last layers, and dynamic shift-conditioned aggregation.

  7. Knowl 7 — UrbanCars reveals shortcut amplification by existing methods

    empirical result

    ERM on UrbanCars achieves 97.6%97.6\% in-distribution accuracy but has gaps of −15.3-15.3 points for background, −11.2-11.2 points for co-occurring object, and −69.2-69.2 points when both shortcuts are uncommon. Many mitigation methods improve one gap while worsening another. CutMix increases the background gap to −45.0-45.0 points, which is 2.942.94 times the ERM reliance; AugMix increases the co-occurring-object gap to −12.1-12.1 points; and CF+F targeted augmentation obtains a co-occurring-object gap of +0.4+0.4 points while worsening the background gap to −16.0-16.0 points.

    Methods using only one shortcut label produce the same asymmetry. For example, when trained with background labels only, gDRO, DI, and SUBG obtain co-occurring-object gaps of −26.9-26.9, −27.0-27.0, and −36.4-36.4 points, respectively, compared with the ERM gap of −11.3-11.3 points. When trained with co-occurring-object labels only, gDRO, DI, and SUBG obtain background gaps of −31.4-31.4, −36.1-36.1, and −60.2-60.2 points, compared with the ERM gap of −15.4-15.4 points. The comparison table on page 7 shows that label-based methods mitigate the labeled shortcut while amplifying the unlabeled shortcut.

  8. Knowl 8 — Asynchronous shortcut learning defeats pseudo-label inference

    empirical result

    Pseudo-shortcut-label methods fail because ERM learns different visual cues at different training times. On UrbanCars, the ERM prediction accuracy against the background label is 82.6%82.6\% at epoch 1, while its accuracy against the co-occurring-object label is 60.6%60.6\%. By epoch 2, background reliance falls to 71.2%71.2\%, whereas co-occurring-object reliance rises to approximately 71.8%71.8\%.

    Consequently, LfF, JTT with one reference epoch, and EIIL with one reference epoch identify the early-learned background shortcut and then amplify co-occurring-object reliance. JTT with two reference epochs and EIIL with two reference epochs identify more of the later co-occurring-object reliance and instead amplify background reliance. On UrbanCars, LfF has gaps of −11.6-11.6 for background and −18.4-18.4 for co-occurring object; JTT with E=1E=1 has −8.1-8.1 and −13.3-13.3; EIIL with E=1E=1 has −4.2-4.2 and −24.7-24.7; JTT with E=2E=2 has −23.3-23.3 and −5.3-5.3; and EIIL with E=2E=2 has −21.5-21.5 and −6.8-6.8. These results show why a single training snapshot cannot reliably infer labels for multiple shortcuts learned asynchronously.

  9. Knowl 9 — Whac-A-Mole behavior persists across ImageNet models and mitigation methods

    empirical result

    On ImageNet, ResNet-50 trained with ERM reaches 76.39%76.39\% clean top-1 accuracy but has a watermark IN-W Gap of −25.40-25.40 points, a +30+30-point Carton Gap, a SIN Gap of −69.43-69.43 points, an IN-R Gap of −56.22-56.22 points, and an IN-9 Gap of −5.19-5.19 points. Existing mitigation methods generally improve one shortcut while worsening another. For example, AugMix increases the Carton Gap to +38+38 points, CutMix increases the background gap to −5.65-5.65 points, texture augmentation produces a +36+36-point Carton Gap, and background augmentation produces a +36+36-point Carton Gap. EIIL worsens the watermark gap to −33.48-33.48 points and the IN-R gap to −61.35-61.35 points.

    The same pattern appears in self-supervised and foundation models. MoCov3 worsens all three major shortcut metrics relative to its ERM counterpart; SWAG with linear probing improves IN-R Gap to −19.79-19.79 points but worsens background reliance to −10.39-10.39 points; and SEER, Uniform Soup, and Greedy Soup reduce watermark reliance while amplifying background reliance. CLIP avoids a clear increase in one of the reported gaps but still retains substantial texture and watermark gaps and has lower clean accuracy than several other foundation models. The results demonstrate that extra pretraining data, larger architectures, and alternative supervision do not reliably solve joint multi-shortcut mitigation.

  10. Knowl 10 — LLE substantially reduces shortcut reliance without Whac-A-Mole behavior

    data/table

    On UrbanCars, LLE obtains 96.7%96.7\% in-distribution accuracy, a background gap of −2.1-2.1 points, a co-occurring-object gap of −2.7-2.7 points, and a joint gap of −5.9-5.9 points. It achieves the best background and joint robustness among the compared methods and the second-best co-occurring-object robustness, without worsening either shortcut relative to ERM.

    With ResNet-50 on ImageNet, LLE obtains 76.25%76.25\% clean accuracy, an IN-W Gap of −6.18-6.18 points, a Carton Gap of +10+10 points, a SIN Gap of −61.00-61.00 points, an IN-R Gap of −54.89-54.89 points, and an IN-9 Gap of −3.82-3.82 points. It is best among the compared ResNet-50 methods on Carton Gap, SIN Gap, and IN-9 Gap, and improves both IN-W Gap and IN-R Gap relative to ERM.

    Using MAE features, LLE obtains 83.68%83.68\% clean accuracy with gaps of −2.48-2.48, +6+6, −58.78-58.78, −44.96-44.96, and −3.70-3.70 for IN-W, Carton, SIN, IN-R, and IN-9, respectively, with a ViT-B/16-style setting. In the larger ViT-L setting, MAE+LLE obtains 85.84%85.84\% clean accuracy and gaps of −1.74-1.74, +12+12, −56.32-56.32, −34.64-34.64, and −2.77-2.77. The ablation on page 8 shows that removing the distribution-shift classifier worsens the IN-W Gap to −17.77-17.77 and the Carton Gap to +36+36, while the full LLE model gives −6.18-6.18 and +10+10; removing the ensemble also worsens most metrics.

  11. Knowl 11 — Scope limitation: LLE requires known shortcut types

    limitation

    LLE assumes that the number and types of shortcuts to mitigate are known, although image-level shortcut labels are not required. The paper does not solve mitigation when both the shortcut types and their number are unknown; it identifies this as future work. It also does not provide a theoretical analysis of why the Whac-A-Mole phenomenon arises, leaving such an analysis for future research.

Coverage note — The paper’s detailed appendix construction procedures, auxiliary watermark-content/language experiments, and full per-model appendix tables were omitted because the main benchmark definitions, central comparisons, and LLE evidence are represented here.

References

  1. 1.Whack A Mole image is obtained from Flaticon.com.
  2. 2.Chirag Agarwal, Daniel D’souza, and Sara Hooker. Estimating Example Difficulty Using Variance of Gradients. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 27
  3. 3.Faruk Ahmed, Yoshua Bengio, Harm van Seijen, and Aaron Courville. Systematic generalisation with group invariant predictions. In International Conference on Learning Representations, 2021. 8
  4. 4.Martin Arjovsky, Leon Bottou, Ishaan Gulrajani, and David Lopez-Paz. Invariant Risk Minimization. arXiv preprint arXiv:1907.02893, 2019. 2, 8
  5. 5.Hyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo, and Seong Joon Oh. Learning De-biased Representations with Biased Representations. In International Conference on Machine Learning, 2020. 8
  6. 6.Yujia Bao and Regina Barzilay. Learning to Split for Automatic Bias Detection. arXiv:2204.13749 [cs], 2022. 27
  7. 7.Yujia Bao, Shiyu Chang, and Dr Regina Barzilay. Learning Stable Classifiers by Transferring Unstable Features. In International Conference on Machine Learning, 2022. 1
  8. 8.Yujia Bao, Shiyu Chang, and Regina Barzilay. Predict then Interpolate: A Simple Algorithm to Learn Stable Classifiers. In International Conference on Machine Learning, 2021. 8
  9. 9.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems, 2019. 24, 27
  10. 10.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark Krass, Ranjay Krishna, Rohith Kud itipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramer, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258, 2021. 1, 3, 8
  11. 11.Chun-Hao Chang, George Alexandru Adam, and Anna Goldenberg. Towards Robust Classification Model by Counterfactual and Invariant Data Generation. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 4, 5, 13, 17
  12. 12.Hila Chefer, Idan Schwartz, and Lior Wolf. Optimizing Relevance Maps of Vision Transformers Improves Robustness. Advances in Neural Information Processing Systems, 2022. 19, 22
  13. 13.Xinlei Chen, Saining Xie, and Kaiming He. An Empirical Study of Training Self-Supervised Vision Transformers. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 3, 4, 7, 19
  14. 14.Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. Per-Pixel Classification is Not All You Need for Semantic Segmentation. In Advances in Neural Information Processing Systems, 2021. 13
  15. 15.Elliot Creager, Jorn-Henrik Jacobsen, and Richard Zemel. Environment Inference for Invariant Learning. In International Conference on Machine Learning, 2021. 1, 3, 4, 5, 8, 17, 26
  16. 16.Terrance de Vries, Ishan Misra, Changhan Wang, and Laurens van der Maaten. Does object recognition work for everyone? In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019. 26
  17. 17.Alex J. DeGrave, Joseph D. Janizek, and Su-In Lee. AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence, 2021. 1
  18. 18.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009. 1, 3, 4, 8, 19
  19. 19.Greg d’Eon, Jason d’Eon, James R. Wright, and Kevin Leyton-Brown. The Spotlight: A General Method for Discovering Systematic Errors in Deep Learning Models. In ACM Conference on Fairness, Accountability, and Transparency, 2022. 27
  20. 20.Terrance DeVries and Graham W Taylor. Improved Regularization of Convolutional Neural Networks with Cutout. arXiv preprint arXiv:1708.04552, 2017. 4, 19
  21. 21.Thomas G. Dietterich. Ensemble Methods in Machine Learning. In Multiple Classifier Systems, 2000. 5
  22. 22.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations, 2021. 3, 4, 19
  23. 23.Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, and Volker Fischer. DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 27
  24. 24.Sabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Re. Domino: Discovering Systematic Errors with Cross-Modal Embeddings. In International Conference on Learning Representations, 2022. 27
  25. 25.Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt. Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP). International Conference on Machine Learning, 2022. 4
  26. 26.Robert Geirhos, Jorn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2020. 1, 8
  27. 27.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations, 2019. 1, 3, 4, 5, 8, 17, 19, 21, 22
  28. 28.Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Mannat Singh, Ishan Misra, Levent Sagun, Armand Joulin, and Piotr Bojanowski. Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision. arXiv preprint arXiv:2202.08360, 2022. 3, 4, 7, 8, 19
  29. 29.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A Dataset for Large Vocabulary Instance Segmentation. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 3, 14
  30. 30.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked Autoencoders Are Scalable Vision Learners. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3, 4, 7, 8, 16, 19
  31. 31.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 3, 4, 19
  32. 32.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity Mappings in Deep Residual Networks. In The European Conference on Computer Vision (ECCV), 2016. 19, 22
  33. 33.Yue He, Zheyan Shen, and Peng Cui. Towards Non-I.I.D. image classification: A dataset and baselines. Pattern Recognition, 2021. 1, 8
  34. 34.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 1, 4, 8
  35. 35.Dan Hendrycks and Thomas Dietterich. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In International Conference on Learning Representations, 2019. 8
  36. 36.Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty. In International Conference on Learning Representations, 2020. 3, 4, 19
  37. 37.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural Adversarial Examples. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 8, 24
  38. 38.Mark Ibrahim, Quentin Garrido, Ari Morcos, and Diane Bouchacourt. The Robustness Limits of SoTA Vision Models to Natural Variation. arXiv preprint arXiv:2210.13604, 2022. 27
  39. 39.Badr Youbi Idrissi, Martin Arjovsky, Mohammad Pezeshki, and David Lopez-Paz. Simple data balancing achieves competitive worst-group-accuracy. Conference on Causal Learning and Reasoning, 2022. 4, 5, 8, 13, 26
  40. 40.Badr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov, Caner Hazirbas, Nicolas Ballas, Pascal Vincent, Michal Drozdzal, David Lopez-Paz, and Mark Ibrahim. ImageNet-X: Understanding Model Mistakes with Factor of Variation Annotations. In International Conference on Learning Representations, 2023. 27
  41. 41.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. OpenCLIP, 2021. 8
  42. 42.Pavel Izmailov, Polina Kirichenko, Nate Gruver, and Andrew Gordon Wilson. On Feature Learning in the Presence of Spurious Correlations. In Advances in Neural Information Processing Systems, 2022. 8
  43. 43.Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry. Distilling Model Failures as Directions in Latent Space. In International Conference on Learning Representations, 2023. 27
  44. 44.Eungyeup Kim, Jihyeon Lee, and Jaegul Choo. BiaSwap: Removing Dataset Bias With Bias-Tailored Swapping Augmentation. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 8
  45. 45.Nayeong Kim, Sehyun Hwang, Sungsoo Ahn, Jaesik Park, and Suha Kwak. Learning Debiased Classifier with Biased Committee. In Advances in Neural Information Processing Systems, 2022. 8
  46. 46.Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson. Last Layer Re-Training is Sufficient for Robustness to Spurious Correlations. In International Conference on Learning Representations, 2023. 4, 5, 6, 8, 16, 22, 26
  47. 47.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic Segmentation. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 13
  48. 48.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M. Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. WILDS: A Benchmark of in-the-Wild Distribution Shifts. In Proceedings of the 38th International Conference on Machine Learning, 2021. 8
  49. 49.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big Transfer (BiT): General Visual Representation Learning. In The European Conference on Computer Vision (ECCV), 2020. 19, 22
  50. 50.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3D Object Representations for Fine-Grained Categorization. In The IEEE International Conference on Computer Vision Workshops, 2013. 2, 13
  51. 51.Oran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald, Gal Elidan, Avinatan Hassidim, William T. Freeman, Phillip Isola, Amir Globerson, Michal Irani, and Inbar Mosseri. Explaining in Style: Training a GAN to explain a classifier in StyleSpace. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 1
  52. 52.Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Yuanqing Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, Ashish Kapoor, and Aleksander Madry. 3DB: A Framework for Debugging Computer Vision Models. In Advances in Neural Information Processing Systems, 2022. 27
  53. 53.Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998. 8
  54. 54.Zhiheng Li, Anthony Hoogs, and Chenliang Xu. Discover and Mitigate Unknown Biases with Debiasing Alternate Networks. In The European Conference on Computer Vision (ECCV), 2022. 4, 5, 8, 17, 26
  55. 55.Zhiheng Li and Chenliang Xu. Discover the Unknown Biased Attribute of an Image Classifier. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 27
  56. 56.Weixin Liang and James Zou. MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts. In International Conference on Learning Representations, 2022. 8
  57. 57.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In The European Conference on Computer Vision (ECCV), 2014. 13
  58. 58.Yong Lin, Shengyu Zhu, Lu Tan, and Peng Cui. ZIN: When and How to Learn Invariance Without Environment Partition? In Advances in Neural Information Processing Systems, 2022. 5, 8, 25, 27
  59. 59.Evan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn. Just Train Twice: Improving Group Robustness without Training Group Information. International Conference on Machine Learning, 2021. 1, 3, 4, 5, 8, 17, 26
  60. 60.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep Learning Face Attributes in the Wild. In The IEEE International Conference on Computer Vision (ICCV), 2015. 2, 8
  61. 61.Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. Learning from Failure: De-biasing Classifier from Biased Classifier. In Advances in Neural Information Processing Systems, 2020. 1, 2, 4, 5, 8, 17, 26
  62. 62.Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, and Ludwig Schmidt. Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP. In Advances in Neural Information Processing Systems, 2022. 4
  63. 63.Zoe Papakipos and Joanna Bitton. AugLy: Data Augmentations for Robustness. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2022. 16
  64. 64.Mohammad Pezeshki, Sekou-Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie. Gradient Starvation: A Learning Proclivity in Neural Networks. Advances in Neural Information Processing Systems, 2021. 4
  65. 65.Francesco Pinto, Harry Yang, Ser-Nam Lim, Philip H. S. Torr, and Puneet K. Dokania. RegMixup: Mixup as a Regularizer Can Surprisingly Improve Accuracy and Out Distribution Robustness. In Advances in Neural Information Processing Systems, 2022. 4
  66. 66.Xavier Soria Poma, Edgar Riba, and Angel Sappa. Dense Extreme Inception Network: Towards a Robust CNN Model for Edge Detection. In The IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2020. 23
  67. 67.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. International Conference on Machine Learning, 2021. 3, 4, 7, 8, 19
  68. 68.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing Network Design Spaces. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3, 4, 19
  69. 69.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet Classifiers Generalize to ImageNet? In Proceedings of the 36th International Conference on Machine Learning, 2019. 8, 18, 24
  70. 70.William A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman. The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world. In Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. 26
  71. 71.Evgenia Rusak, Steffen Schneider, Peter Vincent Gehler, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. ImageNet-D: A new challenging robustness dataset inspired by domain adaptation. In ICML 2022 Shift Happens Workshop, 2022. 24
  72. 72.Evgenia Rusak, Steffen Schneider, George Pachitariu, Luisa Eck, Peter Vincent Gehler, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. If your data distribution shifts, use self-learning. Transactions on Machine Learning Research, 2022. 24
  73. 73.Chaitanya K. Ryali, David J. Schwab, and Ari S. Morcos. Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations, 2021. 4, 5, 8, 16, 22
  74. 74.Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang. Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization. In International Conference on Learning Representations, 2020. 1, 2, 3, 4, 5, 8, 13, 16, 18, 26
  75. 75.Christoph Schuhmann, Romain Beaumont, Cade W. Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Patrick Schramowski, Srivatsa R. Kundurthy, Katherine Crowson, Mitchell Wortsman, Richard Vencu, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: An open large-scale dataset for training next generation image-text models. In Thirty-Sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. 3, 4, 19
  76. 76.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. In Advances in Neural Information Processing Systems Workshops, 2021. 3, 4, 19
  77. 77.Luca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli, and Sangdoo Yun. Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space Perspective. In International Conference on Learning Representations, 2022. 27
  78. 78.Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization. In The IEEE International Conference on Computer Vision (ICCV), 2017. 3, 20
  79. 79.Seonguk Seo, Joon-Young Lee, and Bohyung Han. Unsupervised Learning of Debiased Representations With Pseudo-Attributes. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 8
  80. 80.Robik Shrestha, Kushal Kafle, and Christopher Kanan. An Investigation of Critical Issues in Bias Mitigation Techniques. In The IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022. 8
  81. 81.Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan, Ross Girshick, Piotr Dollar, and Laurens van der Maaten. Revisiting Weakly Supervised Pre-Training of Visual Perception Models. The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3, 4, 7, 19
  82. 82.Sahil Singla and Soheil Feizi. Salient ImageNet: How to discover spurious features in Deep Learning? In International Conference on Learning Representations, 2022. 1
  83. 83.Sahil Singla, Besmira Nushi, Shital Shah, Ece Kamar, and Eric Horvitz. Understanding Failures of Deep Networks via Robust Feature Extraction. The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 27
  84. 84.Nimit Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu, and Christopher Re. No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification Problems. In Advances in Neural Information Processing Systems, 2020. 8
  85. 85.Vladimir Vapnik. The Nature of Statistical Learning Theory. Springer Science & Business Media, 1999. 4, 5, 6
  86. 86.Vasilis Vryniotis. How to Train State-Of-The-Art Models Using TorchVision’s Latest Primitives. https://pytorch.org/blog/how-to-train-state-of-the-art-models-using-torchvision-latest-primitives, 2021. 4, 16
  87. 87.Haohan Wang, Songwei Ge, Eric P. Xing, and Zachary C. Lipton. Learning Robust Global Representations by Penalizing Local Predictive Power. In Advances in Neural Information Processing Systems, 2019. 8, 23
  88. 88.Haohan Wang, Zexue He, Zachary C. Lipton, and Eric P. Xing. Learning Robust Representations by Projecting Superficial Statistics Out. International Conference on Learning Representations, 2019. 8
  89. 89.Zeyu Wang, Klint Qinami, Ioannis Christos Karakozis, Kyle Genova, Prem Nair, Kenji Hata, and Olga Russakovsky. Towards Fairness in Visual Recognition: Effective Strategies for Bias Mitigation. The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 4, 8, 26
  90. 90.Ross Wightman, Hugo Touvron, and Herve Jégou. ResNet strikes back: An improved training procedure in timm, 2021. 4
  91. 91.Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning, 2022. 3, 4, 7, 8, 19
  92. 92.Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo-Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. Robust fine-tuning of zero-shot models. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 8
  93. 93.Kai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. Noise or Signal: The Role of Image Backgrounds in Object Recognition. In International Conference on Learning Representations, 2021. 1, 3, 4, 5, 8, 16, 22
  94. 94.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. In The IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 3, 4, 19, 25
  95. 95.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. Mixup: Beyond Empirical Risk Minimization. In International Conference on Learning Representations, 2018. 3, 4, 19
  96. 96.Jianyu Zhang, David Lopez-Paz, and Leon Bottou. Rich Feature Construction for the Optimization-Generalization Dilemma. In International Conference on Machine Learning, 2022. 8
  97. 97.Eric Zhao, De-An Huang, Hao Liu, Zhiding Yu, Anqi Liu, Olga Russakovsky, and Anima Anandkumar. Scaling Fair Learning to Hundreds of Intersectional Groups. In Submitted to International Conference on Learning Representations, 2022. 8
  98. 98.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random Erasing Data Augmentation. AAAI Conference on Artificial Intelligence, 2020. 4, 19
  99. 99.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 Million Image Database for Scene Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018. 2, 14

Citation

MLA
Li, Z., et al. “A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others”. arXiv, 2022, http://arxiv.org/abs/2212.04825v2.
APA
Li, Z., Evtimov, I., Gordo, A., Hazirbas, C., Hassner, T., Ferrer, C. C., Xu, C., & Ibrahim, M. (2022). A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others. arXiv. http://arxiv.org/abs/2212.04825v2
Chicago
Li, Z., I. Evtimov, A. Gordo, et al. 2022. “A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others”. arXiv. http://arxiv.org/abs/2212.04825v2.
Harvard
Li, Z. et al. (2022) “A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.04825v2.
Vancouver
1. Li Z, Evtimov I, Gordo A, Hazirbas C, Hassner T, Ferrer CC, Xu C, Ibrahim M (2022) A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others. arXiv

BibTeX

@article{li2022whac,
  title = {A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others},
  author = {Li, Zhiheng and Evtimov, Ivan and Gordo, Albert and Hazirbas, Caner and Hassner, Tal and Ferrer, Cristian Canton and Xu, Chenliang and Ibrahim, Mark},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.04825v2},
  eprint = {2212.04825}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE