TrojViT: Trojan Insertion in Vision Transformers

Mengxin ZhengQian LouLei Jiang

article2023CVPR87 citations
Listen

Vision Transformers have rapidly emerged as the state-of-the-art approach for automated visual recognition across industries, often replacing traditional convolutional neural networks. However, deploying these deep learning models in untrusted cloud environments or sourcing them from third-party repositories exposes critical systems to backdoor security risks, where hidden manipulations force targeted misclassifications during operation. While backdoor vulnerabilities in older network types are extensively mapped, security risks specific to transformer architectures remain poorly addressed. Prior attack methods borrowed from older models fail on transformers, causing noticeable drops in overall accuracy or failing to reliably trigger the intended misbehavior. This vulnerability gap poses significant operational and security risks for enterprises deploying visual transformer systems.

The article demonstrates and evaluates a novel, highly stealthy backdoor attack mechanism tailored specifically for Vision Transformers, termed TrojViT. The objective was to design and test an attack pipeline that requires minimal modifications to model parameters stored in memory and operates without needing original training data, while guaranteeing near-perfect attack effectiveness and preserving standard inference accuracy.

The authors conducted extensive experimental evaluations using multiple benchmark transformer models, including standard Vision Transformers, Data-efficient Image Transformers, and Swin Transformers across major datasets such as CIFAR-10 and ImageNet. The research team generated small, distributed visual triggers tailored to patch-based attention mechanisms rather than using large, continuous trigger areas. They utilized an optimization approach that balances the attention paid to trigger regions with targeted classification goals. To inject the vulnerability into deployed models, they applied hardware-based bit-flipping techniques that modify a minimal number of weight bits in system memory.

The findings show that TrojViT achieves exceptional attack performance while remaining practically undetectable. On the large-scale ImageNet benchmark, modifying as few as 345 specific bits out of 22 million parameters successfully forced 99.64% of targeted images to misclassify into an adversary's chosen class. Furthermore, the model retained normal classification accuracy on clean data, suffering less than a 0.35% drop in performance. The proposed patch-distributed triggers occupied as little as 0.51% of the input image area, making them far smaller and harder to detect than previous trigger formats. The study also revealed that existing defenses designed for older architectures fail against this attack mechanism, as they cannot neutralize distributed patch modifications injected post-deployment.

These results demonstrate a serious operational risk for organizations relying on shared computing infrastructure, cloud providers, or third-party model hubs. Hardware-level memory manipulation techniques can silently compromise high-value transformer models with minimal compute time on a single graphics processing unit, bypassing conventional dataset audits and standard performance monitoring. Because existing defenses do not protect deployed models from memory-level parameter edits, relying on traditional security checks creates a false sense of protection.

To mitigate these vulnerabilities, the article recommends implementing matrix decomposition techniques on the most critical parameter layers of the model, specifically the final classification layers. Decomposing these weight structures in memory forces potential attackers to alter significantly more parameters, which increases attack overhead by roughly double and lowers the attack success rate by more than 21%. Decision-makers should evaluate memory-level model protections before deploying vision transformers into sensitive or shared operational environments.

The primary limitation of the study is that its attack assumes an adversary possesses knowledge of the target model's architecture and can perform precise memory bit modifications using hardware-level techniques. While these assumptions represent realistic threat scenarios in modern cloud hosting and software supply chains, further research is needed to develop complete defenses that fully eliminate the vulnerability without imposing excessive computational overhead.

Abstract

Vision Transformers (ViTs) have demonstrated the state-of-the-art performance in various vision-related tasks. The success of ViTs motivates adversaries to perform backdoor attacks on ViTs. Although the vulnerability of traditional CNNs to backdoor attacks is well-known, backdoor attacks on ViTs are seldom-studied. Compared to CNNs capturing pixel-wise local features by convolutions, ViTs extract global context information through patches and attentions. Naively transplanting CNN-specific backdoor attacks to ViTs yields only a low clean data accuracy and a low attack success rate. In this paper, we propose a stealth and practical ViT-specific backdoor attack TrojViT. Rather than an area-wise trigger used by CNN-specific backdoor attacks, TrojViT generates a patch-wise trigger designed to build a Trojan composed of some vulnerable bits on the parameters of a ViT stored in DRAM memory through patch salience ranking and attention-target loss. TrojViT further uses parameter distillation to reduce the bit number of the Trojan. Once the attacker inserts the Trojan into the ViT model by flipping the vulnerable bits, the ViT model still produces normal inference accuracy with benign inputs. But when the attacker embeds a trigger into an input, the ViT model is forced to classify the input to a predefined target class. We show that flipping only few vulnerable bits identified by TrojViT on a ViT model using the well-known RowHammer can transform the model into a backdoored one. We perform extensive experiments of multiple datasets on various ViT models. TrojViT can classify 99.64% of test images to a target class by flipping 345 bits on a ViT for ImageNet.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 2.1. Vision Transformer
  • 2.2. RowHammer
  • 2.3. Our Threat Model
  • 2.4. Limitations of an Area-Wise Trigger on a ViT
  • 2.5. Related Work
  • 3. TrojViT
  • 3.1. Patch-wise Trigger Generation
  • 3.2. Tuned Parameter Distillation
  • 4. Experimental Methodology
  • 5. Results
  • 5.1. Main Results
  • 5.2. Ablation study
  • 6. Potential Defense
  • 7. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — TrojViT attack architecture

    model/method

    TrojViT is a ViT-specific backdoor attack with two phases: patch-wise trigger generation and Trojan insertion. The attack first constructs a trigger whose separate pieces are embedded into selected image patches, then modifies a small set of victim-model parameters so that triggered inputs are classified as a chosen target class while clean inputs retain nearly the original predictions.

    Unlike area-wise CNN backdoors, TrojViT exploits the patch and attention structure of vision transformers. The attacker can generate the trigger and parameter Trojan using the victim architecture, parameters, and patch size plus randomly sampled test images, without access to the original training data. The Trojan is intended to be inserted by flipping vulnerable bits in model weights stored in DRAM, for example through RowHammer. The workflow diagram on page 4 depicts the three load-bearing components: patch salience ranking, Attention-Target loss, and tuned parameter distillation.

  2. Knowl 2 — Patch salience ranking for trigger placement

    equation

    For an input image tensor XX, a trigger perturbation tensor PP, and a binary patch mask MM, TrojViT forms the triggered input

    X^=X+P⊙M,\widehat{X}=X+P\odot M,

    where ⊙\odot is element-wise multiplication and the mask is broadcast over the pixels of each patch. Suppose the ViT partitions X^\widehat{X} into nn patches, each containing dd pixels; X^i,j\widehat{X}_{i,j} denotes pixel jj in patch ii, and yky_k is the attacker-selected target class. The pixel salience is the absolute gradient of target cross-entropy loss with respect to that pixel, and the salience of patch ii is the sum over its pixels:

    GX^i=∑j=1d∣∂LCE(X^,yk)∂X^i,j∣.\mathcal{G}_{\widehat{X}_i}=\sum_{j=1}^{d}\left|\frac{\partial \mathcal{L}_{CE}(\widehat{X},y_k)}{\partial \widehat{X}_{i,j}}\right|.

    For each patch index t∈{1,…,n}t\in\{1,\ldots,n\}, the mask value is set to one exactly when its salience belongs to the NN largest patch-salience values:

    Mt={1,GX^t∈TopN⁡({GX^i}i=1n,N),0,otherwise.M_t=\begin{cases} 1,&\mathcal{G}_{\widehat{X}_t}\in\operatorname{TopN}(\{\mathcal{G}_{\widehat{X}_i}\}_{i=1}^{n},N),\\ 0,&\text{otherwise}. \end{cases}

    Thus, trigger pieces are placed in patches that have the largest gradient-based influence on the target class rather than at a fixed contiguous image location.

  3. Knowl 3 — Attention-Target loss with gradient surgery

    equation

    TrojViT optimizes the perturbation PP using both target classification and attention concentration. Let LL be the number of transformer layers, hh index an attention head, ii index a query patch, and TT be the set of selected trigger-patch indices. Let attn⁡i→Tl,h\operatorname{attn}^{l,h}_{i\rightarrow T} denote the attention mass assigned by query patch ii to patches in TT at layer ll and head hh. The attention loss at layer ll is

    LATTNl(X^,T)=−log⁡(∑h,iattn⁡i→Tl,h).\mathcal{L}^{l}_{ATTN}(\widehat{X},T)=-\log\left(\sum_{h,i}\operatorname{attn}^{l,h}_{i\rightarrow T}\right).

    With target cross-entropy LCE(X^,yk)\mathcal{L}_{CE}(\widehat{X},y_k) and attention-loss weight λ≥0\lambda\geq 0, the Attention-Target loss is

    LATL(X^,yk)=LCE(X^,yk)+λ∑l=1LLATTNl(X^,T).\mathcal{L}_{ATL}(\widehat{X},y_k)=\mathcal{L}_{CE}(\widehat{X},y_k)+\lambda\sum_{l=1}^{L}\mathcal{L}^{l}_{ATTN}(\widehat{X},T).

    The first term encourages classification as target class yky_k; the second encourages the selected trigger patches to attract attention. Because the two objectives can produce conflicting gradients, TrojViT applies gradient surgery. If gCE=∇X^LCEg_{CE}=\nabla_{\widehat{X}}\mathcal{L}_{CE} and gATTN=∇X^(λ∑lLATTNl)g_{ATTN}=\nabla_{\widehat{X}}\left(\lambda\sum_l\mathcal{L}^{l}_{ATTN}\right), the combined gradient used to optimize PP is

    gATL={gCE+gATTN,cos⁡(gCE,gATTN)>0,gCE+gATTN−gCE⋅gATTN∥gCE∥2gCE,otherwise.g_{ATL}=\begin{cases} g_{CE}+g_{ATTN},&\cos(g_{CE},g_{ATTN})>0,\\ g_{CE}+g_{ATTN}-\dfrac{g_{CE}\cdot g_{ATTN}}{\lVert g_{CE}\rVert^2}g_{CE},&\text{otherwise}. \end{cases}

    The second case removes the conflicting component according to the gradient-surgery rule used by TrojViT.

  4. Knowl 4 — Tuned parameter distillation for sparse Trojan insertion

    algorithm

    Tuned parameter distillation constructs Trojan weights WTW_T from a victim ViT's full parameter set WW while retaining only parameters whose updates are useful for both clean and triggered behavior. The inputs are the victim parameters WW, a batch of test images XX with clean labels yy, a patch-wise trigger M⊙PM\odot P, target class yky_k, learning rate lrlr, number of fine-tuning epochs, and pruning threshold ee. The initial candidate set is chosen from important parameters, specifically the last attention and classification layers; all other parameters remain fixed.

    At every epoch, the method computes a clean-input gradient and a triggered-input gradient, combines them with gradient surgery, updates the candidate Trojan weights, and removes candidate parameters whose absolute update is below ee. The objective jointly minimizes clean-input loss and triggered-input target loss:

    min⁡WT[L(f(X),y)+L(f(X+M⊙P),yk)].\min_{W_T}\left[\mathcal{L}(f(X),y)+\mathcal{L}(f(X+M\odot P),y_k)\right].

    The paper's insertion procedure is summarized below. IDWTID_{W_T} stores the indices of retained parameters and np=∣IDWT∣n_p=|ID_{W_T}| is the tuned parameter number.

    Input: Victim parameters W, test batch X with labels y, trigger M ⊙ P, target class y_k, threshold e, learning rate lr, and number of epochs
    Output: Trojan weights W_T and retained-parameter count n_p
    Initialize W_T from the last attention and classification layers of W
    Initialize ID_WT with the indices of W_T
    for each epoch do
        Compute clean gradient g_c = gradient of L(f(X), y) with respect to W_T
        Compute triggered gradient g_t = gradient of L(f(X + M ⊙ P), y_k) with respect to W_T
        Combine g_c and g_t with gradient surgery to obtain g
        Update W_T' = W_T + lr · g, using the update convention in the paper
        Compute each parameter's absolute update magnitude |W_T' - W_T|
        Remove indices whose update magnitude is smaller than e
        Remove the corresponding entries from W_T'
        Set W_T = W_T'
    end for
    return W_T and n_p = length(ID_WT)

    The resulting weights are intended to be written into the victim model by changing only a small number of quantized weight bits. In the threshold experiment, increasing ee from 00 to 0.0030.003 reduced the modified bits from 1,650 to 345 while retaining a 99.64% attack success rate and 79.12% clean accuracy on Deit-small/ImageNet.

  5. Knowl 5 — Attacker capabilities and attack objectives

    assumption

    The TrojViT threat model assumes that the attacker knows the victim ViT's architecture, parameters, and patch size, but does not possess the original training dataset. The attacker may be an untrusted service provider that can access model weights in DRAM during inference or a malicious model developer who distributes a poisoned model. For the memory-based scenario, the attacker is assumed able to locate model parameters and flip selected DRAM bits through RowHammer; only a small number of bit changes is required.

    The attack objectives are: utility, meaning clean-input accuracy remains close to that of the unmodified model; effectiveness, meaning triggered inputs are classified into a predefined target class with high attack success rate; and efficiency, meaning Trojan generation and insertion require relatively little computation. Stealth is also pursued through a small trigger area and a small number of modified parameters or bits. TrojViT uses randomly sampled test images rather than original training data; the reported experiments use 384 such images for trigger generation and Trojan insertion.

  6. Knowl 6 — Comparison with prior attacks on Deit-small

    data/table

    The page-6 comparison evaluates attacks on a Deit-small model using ImageNet, 16×16 patches, and a nine-patch trigger occupying 4.59% of the input image. CDA is clean data accuracy, ASR is the fraction of triggered images classified as the target class, TPN is the number of modified parameters, and TBN is the number of modified 8-bit weight bits. The clean-model columns provide the unmodified baseline; the backdoored-model columns show the effect after Trojan insertion.

    The key result is that TrojViT preserves nearly all clean accuracy while achieving the highest ASR and using far fewer modified parameters and bits than ViT-specific prior attacks.

    Could not parse LaTeX table

    TrojViT loses less than 0.3 percentage points of CDA relative to the clean Deit-small model, reaches 99.96% ASR, and modifies 213 of approximately 22 million parameters, equivalent to 880 bit flips. The page-6 discussion emphasizes that the area-wise CNN attacks substantially reduce CDA, whereas the prior ViT attacks modify orders of magnitude more parameters or obtain lower ASR.

  7. Knowl 7 — Cross-model and cross-dataset effectiveness

    data/table

    TrojViT was evaluated on ViT-base, Deit-tiny, Deit-small, Deit-base, and Swin-base models using CIFAR-10 and ImageNet. Models use 3×224×224 inputs; CIFAR-10 models are fine-tuned from ImageNet-pretrained models. The experiments use 8-bit parameter quantization, a batch size of 16, and 384 randomly sampled test images, with no training data used by the attack.

    On CIFAR-10, the attack reaches approximately 99% ASR or higher on every architecture while reducing CDA by at most 0.68 percentage points. On ImageNet, the harder 1,000-class setting still yields 98.72–99.96% ASR with modest CDA degradation.

    Could not parse LaTeX table

    The reported results on page 7 show that smaller models can use smaller triggers: Deit-t uses one patch on CIFAR-10 and four patches on ImageNet, while larger models generally use nine patches for ImageNet.

  8. Knowl 8 — Patch-wise triggers outperform area-wise triggers

    empirical result

    The paper's trigger-position experiment demonstrates that ViTs are sensitive to how a trigger is distributed across patches. On the same victim ViT, two area-wise triggers placed at different image locations achieved ASRs of 89.6% and 94.7%, whereas a patch-wise trigger with separate pieces distributed across critical patches achieved 99.9% ASR. This positional difference is expected for a ViT because attention processes patch tokens rather than only local convolutional neighborhoods.

    The component ablation on Deit-small/ImageNet further isolates the contribution of patch placement. With the same 384 test images and the same parameter-insertion procedure, replacing an area-wise trigger by a patch-wise trigger raises ASR from 94.69% to 96.84% and raises CDA from 74.96% to 77.49%. Adding the Attention-Target loss raises the result to 79.23% CDA and 99.98% ASR before parameter distillation. The page-4 diagram illustrates the corresponding design: each trigger piece is assigned to a selected patch rather than forming one contiguous marked area.

  9. Knowl 9 — Trigger size and insertion-threshold trade-offs

    empirical result

    On Deit-small/ImageNet, TrojViT exhibits a trade-off between trigger area, clean accuracy, and attack success. The trigger area is represented by the number of selected patches, and TAR is the percentage of image area occupied by the trigger.

    Could not parse LaTeX table

    A one-patch trigger already exceeds 99% ASR but has lower CDA; nine patches provide the best CDA and ASR in this sweep. The attention-loss weight also matters: using only target cross-entropy, λ=0\lambda=0, gives 77.49% CDA and 96.84% ASR, while equal weighting of the attention and target terms, λ=1\lambda=1, gives 79.23% CDA and 99.98% ASR. Summing attention losses from all 12 layers gives 79.19% CDA and 99.96% ASR in the reported layer sweep.

    The parameter-pruning threshold controls insertion sparsity. On the same Deit-small/ImageNet setting, the reported values are:

    Could not parse LaTeX table

    Thus, e=0.003e=0.003 cuts the modified bit count by 79.09% relative to e=0e=0 while preserving 99.64% ASR and limiting CDA degradation to 0.35 percentage points relative to the clean model.

  10. Knowl 10 — DRAM decomposition defense against TrojViT

    limitation

    The paper proposes a defense that targets TrojViT's dependence on modifying a critical parameter matrix, especially the final classification head. Instead of storing that matrix directly in DRAM, the defender stores several matrices produced by a prior decomposition method. This prevents the attacker from directly modifying the critical matrix and forces modifications to more parameters; the number of decomposed matrices controls the defense overhead.

    On ImageNet, the defense reduced ASR by more than 21 percentage points and approximately doubled the tuned-parameter count. The reported values are:

    Could not parse LaTeX table

    The authors also report that existing ViT defenses aimed at training-time Trojans or a single high-attention area are not directly suited to TrojViT's inference-time, multi-patch trigger. Applying the area-oriented defense naively degraded CDA by more than 5% on various ImageNet ViTs, so the proposed DRAM decomposition is presented as a more targeted mitigation rather than a complete defense.

Coverage note — The detailed prior-work threat-model comparison and some layer-selection and attention-weight sweeps were not made separate knowls because they support the main method and are summarized within the method and design-space knowls.

References

  1. 1.Mauro Barni, Kassem Kallas, and Benedetta Tondi. A new backdoor attack in cnns by training set corruption without label poisoning. In 2019 IEEE International Conference on Image Processing (ICIP), pages 101–105. IEEE, 2019. 1
  2. 2.Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava. Detecting backdoor attacks on deep neural networks by activation clustering. arXiv preprint arXiv:1811.03728, 2018. 8
  3. 3.Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar. Proflip: Targeted trojan attack with progressive bit flips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7718–7727, 2021. 2, 3, 5, 6
  4. 4.Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. Targeted backdoor attacks on deep learning systems using data poisoning. CoRR, abs/1712.05526, 2017. 1
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 6
  6. 6.Khoa D Doan, Yingjie Lao, Peng Yang, and Ping Li. Defending backdoor attacks on vision transformer via patch processing. arXiv preprint arXiv:2206.12381, 2022. 1, 2, 3, 5, 6, 8
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 1, 2, 6
  8. 8.Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017. 1
  9. 9.Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin. Language model compression with weighted low-rank factorization. In International Conference on Learning Representations, 2022. 8
  10. 10.Weizhe Hua, Zhiru Zhang, and G Edward Suh. Reverse engineering convolutional neural networks through side-channel information leaks. In 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC), pages 1–6. IEEE, 2018. 2
  11. 11.Jing Yu Koh. Model zoo, 2021. 2
  12. 12.Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. Cifar-10 (canadian institute for advanced research), 2009. 6, 7
  13. 13.Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. Fine-pruning: Defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, pages 273–294. Springer, 2018. 8
  14. 14.Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang. Trojaning attack on neural networks. In The Network and Distributed System Security (NDSS) Symposium, 2017. 1
  15. 15.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 1, 2, 6
  16. 16.Qian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen, and Hongxia Jin. Dictformer: Tiny transformer with shared dictionary. In International Conference on Learning Representations, 2022. 8
  17. 17.Peizhuo Lv, Hualong Ma, Jiachen Zhou, Ruigang Liang, Kai Chen, Shengzhi Zhang, and Yunfei Yang. DBIA: data-free backdoor injection attack against transformer networks. CoRR, abs/2111.11870, 2021. 1, 2, 3, 5, 6
  18. 18.Onur Mutlu and Jeremie S. Kim. Rowhammer: A retrospective. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 39(8):1555–1571, 2020. 2
  19. 19.Tuan Anh Nguyen and Anh Tran. Input-aware dynamic backdoor attack. Advances in Neural Information Processing Systems, 33:3454–3464, 2020. 1
  20. 20.Adnan Siraj Rakin, Zhezhi He, and Deliang Fan. Tbt: Targeted neural network attack with bit trojan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13198–13207, 2020. 2, 3, 5, 6
  21. 21.Kaveh Razavi, Ben Gras, Erik Bosman, Bart Preneel, Cristiano Giuffrida, and Herbert Bos. Flip feng shui: Hammering a needle in the software stack. In USENIX Security Symposium, pages 1–18, 2016. 2
  22. 22.Akshayvarun Subramanya, Aniruddha Saha, Soroush Abbasi Koohpayegani, Ajinkya Tejankar, and Hamed Pirsiavash. Backdoor attacks on vision transformers. arXiv preprint arXiv:2206.08477, 2022. 1, 2, 3, 5, 6, 8
  23. 23.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, volume 139, pages 10347–10357, July 2021. 1, 2, 6
  24. 24.Alexander Turner, Dimitris Tsipras, and Aleksander Madry. Label-consistent backdoor attacks. arXiv preprint arXiv:1912.02771, 2019. 1
  25. 25.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. 2
  26. 26.Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In 2019 IEEE Symposium on Security and Privacy (SP), pages 707–723, 2019. 8
  27. 27.Emily Wenger, Josephine Passananti, Arjun Nitin Bhagoji, Yuanshun Yao, Haitao Zheng, and Ben Y Zhao. Backdoor attacks against deep learning systems in the physical world. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6206–6215, 2021. 1
  28. 28.Mingfu Xue, Can He, Yinghao Wu, Shichang Sun, Yushu Zhang, Jian Wang, and Weiqiang Liu. Ptb: Robust physical backdoor attacks against deep neural networks in real world. Computers & Security, 118:102726, 2022. 1
  29. 29.Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 5824–5836. Curran Associates, Inc., 2020. 5, 6
  30. 30.Zhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong, Dakui Wang, and Kaitai Liang. Defeat: Deep hidden feature backdoor attacks by imperceptible perturbation and latent representation constraints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15213–15222, 2022. 1
  31. 31.Nan Zhong, Zhenxing Qian, and Xinpeng Zhang. Imperceptible backdoor attack: From input space to feature representation. arXiv preprint arXiv:2205.03190, 2022. 1
  32. 32.Zhisheng Zhong, Fangyin Wei, Zhouchen Lin, and Chao Zhang. Ada-tucker: Compressing deep neural networks via adaptive dimension adjustment tucker decomposition. Neural Networks, 110:104–115, 2019. 8
  33. 33.Yuankun Zhu, Yueqiang Cheng, Husheng Zhou, and Yantao Lu. Hermes attack: Steal DNN models with lossless inference accuracy. In 30th USENIX Security Symposium (USENIX Security 21), pages 1973–1988. USENIX Association, Aug. 2021. 2

Citation

MLA
Zheng, M., et al. “TrojViT: Trojan Insertion in Vision Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 4025–34, https://doi.org/10.1109/CVPR52729.2023.00392.
APA
Zheng, M., Lou, Q., & Jiang, L. (2023). TrojViT: Trojan Insertion in Vision Transformers. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4025–4034. https://doi.org/10.1109/CVPR52729.2023.00392
Chicago
Zheng, M., Q. Lou, and L. Jiang. 2023. “TrojViT: Trojan Insertion in Vision Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4025–34. https://doi.org/10.1109/CVPR52729.2023.00392.
Harvard
Zheng, M., Lou, Q. and Jiang, L. (2023) “TrojViT: Trojan Insertion in Vision Transformers”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4025–4034. Available at: https://doi.org/10.1109/CVPR52729.2023.00392.
Vancouver
1. Zheng M, Lou Q, Jiang L (2023) TrojViT: Trojan Insertion in Vision Transformers. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4025–4034

BibTeX

@inproceedings{Zheng_2023, title={TrojViT: Trojan Insertion in Vision Transformers}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00392}, DOI={10.1109/cvpr52729.2023.00392}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zheng, Mengxin and Lou, Qian and Jiang, Lei}, year={2023}, month=June, pages={4025–4034} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE