You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?

Zenghui YuanPan ZhouKai ZouYu Cheng

article2023CVPR63 citations

Reveals how the self-attention mechanism makes Vision Transformers uniquely vulnerable to patch-based backdoor attacks and introduces BadViT, an efficient attack framework that manipulates attention maps to implant stealthy triggers with minimal poisoning.

Listen

Vision Transformers have emerged as leading artificial intelligence architectures for computer vision tasks, increasingly replacing traditional Convolutional Neural Networks across high-stakes industrial deployments. However, many production systems rely on outsourcing model training or fine-tuning public pre-trained models, creating critical exposure to backdoor data poisoning attacks where malicious associations are embedded into a model during training. The article aims to evaluate the distinct security vulnerabilities of Vision Transformers relative to traditional networks and systematically demonstrates an attack framework designed to exploit the core self-attention mechanisms unique to transformer architectures.

To conduct this evaluation, the researchers tested standard vision transformer models (DeiT and LeViT families) against convolutional architectures across the large-scale ImageNet dataset and several downstream benchmarks. The team designed BadViT, an attack framework that optimizes a universal patch-wise trigger to actively capture the model's self-attention across network layers. The investigation examined attack success rates, benign accuracy preservation, data poisoning dependency, perturbation constraints for visual imperceptibility, and resilience against three advanced backdoor defense techniques (PatchDrop, Neural Cleanse, and Fine-Pruning).

Key findings show that Vision Transformers are inherently more susceptible to localized patch-level triggers than convolutional networks because their self-attention mechanism naturally magnifies tokenized patch interactions. The BadViT framework achieved a 100% attack success rate on benchmark transformer models after only a single training epoch while preserving clean image accuracy. Furthermore, BadViT proved highly efficient, maintaining an attack success rate of approximately 95% even when poisoning as little as 0.2% of the training dataset, whereas traditional patch triggers completely failed at low poisoning rates. Invisible variants designed with perturbation limits and blending techniques retained success rates up to 100%, successfully transferred to downstream tasks under clean-label conditions, and effectively bypassed or misled all tested state-of-the-art defenses.

These findings indicate substantial operational and safety risks for organizations deploying vision transformer models, especially when sourcing pre-trained weights or using untrusted third-party training pipelines. Standard defenses designed for convolutional networks or conventional patch detection fail to mitigate attention-manipulating attacks, meaning enterprise computer vision systems can be easily hijacked without degrading normal performance. Organizations should exercise strict data provenance controls, treat outsourced training pipelines with high scrutiny, and develop dedicated defense strategies tailored specifically to patch-level attention dynamics before deploying vision transformers in critical environments.

While the study provides rigorous empirical evidence across standard models, its primary focus remains centered on image classification tasks within supervised vision transformer benchmarks. Decision-makers can place high confidence in the technical vulnerability of standard self-attention mechanisms to patch poisoning, but should account for the need for broader testing on emerging multimodal architectures and complex operational tasks such as real-time object tracking.

Cover for You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?

Abstract

Vision Transformers (ViTs), which made a splash in the field of computer vision (CV), have shaken the dominance of convolutional neural networks (CNNs). However, in the process of industrializing ViTs, backdoor attacks have brought severe challenges to security. The success of ViTs benefits from the self-attention mechanism. However, compared with CNNs, we find that this mechanism of capturing global information within patches makes ViTs more sensitive to patch-wise triggers. Under such observations, we delicately design a novel backdoor attack framework for ViTs, dubbed BadViT, which utilizes a universal patch-wise trigger to catch the model's attention from patches beneficial for classification to those with triggers, thereby manipulating the mechanism on which ViTs survive to confuse itself. Furthermore, we propose invisible variants of BadViT to increase the stealth of the attack by limiting the strength of the trigger perturbation. Through a large number of experiments, it is proved that BadViT is an efficient backdoor attack method against ViTs, which is less dependent on the number of poisons, with satisfactory convergence, and is transferable for downstream tasks. Furthermore, the risks inside of ViTs to backdoor attacks are also explored from the perspective of existing advanced defense schemes.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Vision Transformer
  • 2.2. Backdoor Attacks and Defenses
  • 2.3. The Robustness of Vision Transformers
  • 3. Backdoor Attacks in ViTs
  • 3.1. Threat Model
  • 3.2. Background of Backdoor Attacks in ViTs
  • 4. Backdoor Attacks Robustness Comparison
  • 4.1. Attack Settings
  • 4.2. Observations and Discussions
  • 5. The Proposed BadViT Framework
  • 5.1. Inspirations of BadViT
  • 5.2. Formulation of BadViT
  • 5.3. Invisible Variants of BadViT
  • 6. Experiments on BadViT
  • 6.1. Evaluation Settings
  • 6.2. Effectiveness of BadViT
  • 6.3. Evaluations to Invisible Variants of BadViT
  • 6.4. Transferability of BadViT
  • 6.5. Resistance to Backdoor Defenses
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Paper content unavailable for knowledge extraction

    limitation

    The source document (file 41415ff9-0502-4aa9-8c7c-5325417cc2be.pdf) could not be read, so no methods, results, or analyses from the paper are available to report. Extraction of knowls requires the paper text.

Coverage note — The content of the attached PDF was not accessible, so no knowls could be extracted from the paper's contribution; the single knowl below is a placeholder stating this rather than fabricated content.

Citation

MLA
Yuan, Z., et al. “You Are Catching My Attention: Are Vision Transformers Bad Learners Under Backdoor Attacks?”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 24605–15, https://doi.org/10.1109/CVPR52729.2023.02357.
APA
Yuan, Z., Zhou, P., Zou, K., & Cheng, Y. (2023). You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 24605–24615. https://doi.org/10.1109/CVPR52729.2023.02357
Chicago
Yuan, Z., P. Zhou, K. Zou, and Y. Cheng. 2023. “You Are Catching My Attention: Are Vision Transformers Bad Learners Under Backdoor Attacks?”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 24605–15. https://doi.org/10.1109/CVPR52729.2023.02357.
Harvard
Yuan, Z. et al. (2023) “You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 24605–24615. Available at: https://doi.org/10.1109/CVPR52729.2023.02357.
Vancouver
1. Yuan Z, Zhou P, Zou K, Cheng Y (2023) You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 24605–24615

BibTeX

@inproceedings{Yuan_2023, title={You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?}, url={http://dx.doi.org/10.1109/CVPR52729.2023.02357}, DOI={10.1109/cvpr52729.2023.02357}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Yuan, Zenghui and Zhou, Pan and Zou, Kai and Cheng, Yu}, year={2023}, month=June, pages={24605–24615} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE