A Study of the Attention Abnormality in Trojaned BERTs

Weimin LyuSongzhu ZhengTengfei MaChao Chen

article2022NAACL74 citations

Reveals how trigger tokens hijack attention focus in poisoned BERT models and develops an attention-based detector to distinguish trojaned transformers from clean ones without needing prior knowledge of the trigger.

Listen

Deep learning models are increasingly deployed across critical enterprise applications, but they remain vulnerable to backdoor and Trojan attacks. In these attacks, adversaries intentionally manipulate training data so that models perform normally on typical inputs but produce malicious, targeted errors when triggered by specific characters, words, or phrases. Because natural language processing relies on discrete text tokens rather than continuous values, existing security defenses developed for computer vision fail to transfer effectively. This lack of visibility into how Trojans manipulate language models poses significant compliance, safety, and operational risks for organizations deploying pretrained language representations.

The main objective of the article is to investigate the internal mechanisms of Trojaned language models and introduce an automated detection method. The authors specifically evaluate how data poisoning impacts the multi-head attention mechanism within transformer architectures like BERT and demonstrate that these internal behaviors can reliably identify backdoored models without prior knowledge of the attacker's trigger.

To conduct this evaluation, the researchers trained hundreds of clean and Trojaned BERT models across four benchmark text classification corpora—IMDB, Stanford Sentiment Treebank (SST-2), Yelp, and Amazon—incorporating character, word, and phrase triggers across diverse classifier architectures. They analyzed internal attention distributions, entropy, and attribution scores across different layers and attention head categories (semantic, separator, and non-semantic heads). Using head-pruning experiments, they verified the functional impact of anomalous heads. Leveraging these empirical insights, the authors developed the Attention-based Trojan Detector (AttenTD), an automated framework that generates candidate perturbation tokens and flags models exhibiting abnormal attention shifts.

The investigation produced several key findings regarding model behavior and detection performance. First, the article revealed a pronounced "attention focus drifting" phenomenon: when exposed to a trigger, poisoned models abruptly shift their internal attention weights away from meaningful context tokens directly onto the trigger token. This behavior appeared in roughly 74% to 93% of Trojaned models across various head types, compared to only 0% to 28% of clean models. Second, attention entropy consistently decreased during attacks, indicating that the attention became tightly concentrated on the trigger. Third, drifting behavior was heavily concentrated in the final three layers of the transformer, with separator heads playing a disproportionately large functional role; pruning these drifting separator heads alone restored poisoned classification accuracy by up to 22.29%, while pruning all drifting heads restored accuracy by 21.67% to 32.02%. Finally, in comparative benchmark tests, the proposed AttenTD detector achieved 93% to 97% detection accuracy and area under the curve across all corpora and classifier types, substantially outperforming existing computer vision and natural language baselines that reached only 45% to 80% accuracy.

These findings indicate that Trojan vulnerabilities in language models are primarily anchored within the deeper layers of the transformer encoder rather than downstream classification heads. For organizations, this insight provides an interpretable, white-box approach to auditing third-party and open-source models before deployment, significantly reducing operational and security risks. Rather than relying on black-box heuristics that struggle with discrete text, risk and compliance teams can inspect attention shifts to audit model integrity.

Moving forward, technical decision-makers should adopt attention-based scanning mechanisms as part of their standard machine learning validation pipelines. Security teams auditing transformer models should specifically monitor deeper layers and separator head dynamics for abnormal concentration spikes. While AttenTD offers a powerful detection capability, future work is required to develop robust defense and mitigation techniques that neutralize triggers during runtime without degrading standard model accuracy.

Readers should interpret these findings within the scope of the article's experimental boundaries. The analysis focused primarily on sentiment analysis tasks using BERT architectures, and the candidate generation step relies on predefined token lexicons. As attackers adapt to attention-monitoring defenses, trigger designs may evolve, warranting ongoing testing across broader language tasks such as translation, question answering, and generative modeling.

No sufficiently relevant recommendations were found.

Cover for A Study of the Attention Abnormality in Trojaned BERTs

Abstract

Trojan attacks raise serious security concerns. In this paper, we investigate the underlying mechanism of Trojaned BERT models. We observe the attention focus drifting behavior of Trojaned models, i.e., when encountering an poisoned input, the trigger token hijacks the attention focus regardless of the context. We provide a thorough qualitative and quantitative analysis of this phenomenon, revealing insights into the Trojan mechanism. Based on the observation, we propose an attention-based Trojan detector to distinguish Trojaned models from clean ones. To the best of our knowledge, this is the first paper to analyze the Trojan mechanism and to develop a Trojan detector based on the transformer's attention1.

Citation

MLA
Lyu, W., et al. “A Study of the Attention Abnormality in Trojaned BERTs”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 4727–41, https://doi.org/10.18653/v1/2022.naacl-main.348.
APA
Lyu, W., Zheng, S., Ma, T., & Chen, C. (2022). A Study of the Attention Abnormality in Trojaned BERTs. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4727–4741. https://doi.org/10.18653/v1/2022.naacl-main.348
Chicago
Lyu, W., S. Zheng, T. Ma, and C. Chen. 2022. “A Study of the Attention Abnormality in Trojaned BERTs”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4727–41. https://doi.org/10.18653/v1/2022.naacl-main.348.
Harvard
Lyu, W. et al. (2022) “A Study of the Attention Abnormality in Trojaned BERTs”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 4727–4741. Available at: https://doi.org/10.18653/v1/2022.naacl-main.348.
Vancouver
1. Lyu W, Zheng S, Ma T, Chen C (2022) A Study of the Attention Abnormality in Trojaned BERTs. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 4727–4741

BibTeX

@inproceedings{lyu-etal-2022-study,
    title = "A Study of the Attention Abnormality in Trojaned {BERT}s",
    author = "Lyu, Weimin  and
      Zheng, Songzhu  and
      Ma, Tengfei  and
      Chen, Chao",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.348/",
    doi = "10.18653/v1/2022.naacl-main.348",
    pages = "4727--4741"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/