BITE: Textual Backdoor Attacks with Iterative Trigger Injection

Jun YanVansh GuptaXiang Ren

article2023ACL90 citations

Proposes an iterative data-poisoning framework that embeds natural word perturbations to create stealthy, highly effective textual backdoor attacks, alongside a defense strategy that successfully detects and removes the injected trigger words.

Listen

Modern natural language processing systems increasingly rely on external, unverified datasets from public hubs and user-generated web content. This reliance creates a serious vulnerability to data-poisoning backdoor attacks, where an attacker subtly alters training samples so that a deployed model reliably predicts a target label whenever an input contains specific trigger patterns. Prior attack methods faced a steep tradeoff: they either used obvious, unnatural keyword insertions that were easily detected by human reviewers, or employed strict sentence structures and style transfers that failed to reliably control model predictions. Consequently, the real-world risk posed by stealthy backdoor attacks has been systematically underestimated.

The article introduces and evaluates a backdoor attack method called BITE (Backdoor attack with Iterative Trigger Injection), which aims to achieve both high stealthiness and strong attack effectiveness. BITE operates by iteratively identifying and inserting a set of natural trigger words into target-label training data, creating statistical correlations that the model learns to associate with the desired target prediction. To counter this threat, the article also presents a companion defense strategy called DeBITE, which identifies and removes words that exhibit abnormally strong correlations with specific labels in the training dataset.

The researchers evaluated the attack and defense across four benchmark text classification datasets spanning sentiment analysis, hate speech detection, emotion recognition, and question classification. Using standard language models, they tested attack success under a clean-label setting where only 1% of the training data was modified without changing any original category labels. Data stealth was evaluated through both automatic linguistic checks and human assessments measuring sentence naturalness, human suspicion, semantic preservation, and label consistency. The defense was further benchmarked against several existing training-time and inference-time backdoor countermeasures.

The evaluation revealed several critical findings. First, BITE significantly outperformed baseline attacks in attack effectiveness while maintaining natural, human-like text quality. Under a 1% poisoning rate, BITE achieved attack success rates of 62.8% on sentiment analysis and 60.2% on question classification, roughly double the success rates of baseline style- and syntax-based attacks, all while maintaining normal accuracy (over 80% to 96%) on clean inputs. Second, BITE demonstrated a substantial advantage at low poisoning rates; its effectiveness advantage over baselines grew larger as fewer training examples were poisoned, making it dangerous in realistic settings where an attacker can only tamper with a small portion of data. Third, human and automated text evaluations showed BITE preserved original sentence meaning and label validity better than syntactic paraphrasing methods. Finally, the proposed defense method, DeBITE, successfully reduced attack success across multiple attack styles—for instance, reducing syntax attack success on sentiment data from 49.9% to 33.9%—outperforming existing defenses that failed against clean-label attacks.

These findings demonstrate that text-based artificial intelligence models can be covertly compromised without obvious text corruption or label manipulation, posing significant security risks to automated systems such as content moderation and fraud filtering. Organizations cannot rely on standard model evaluation or superficial human spot-checks to detect poisoned data. Because most existing defenses struggle against stealthy, clean-label attacks, organizations utilizing third-party data must implement proactive data sanitization protocols during the data preparation phase.

Decision-makers should adopt training-time dataset audits that detect and filter statistically biased trigger words before model training begins, using techniques similar to the proposed DeBITE method. While this filtering introduces a minor trade-off—a small reduction of approximately 1% in clean classification accuracy—it substantially diminishes vulnerability to backdoor manipulation. Practitioners should also limit reliance on unvetted public datasets for mission-critical applications.

The confidence in these findings is high for standard text classification benchmarks, though readers should note certain limitations. The study was conducted on short-sentence, medium-sized classification datasets; attack dynamics on long-form documents, generative tasks, or massive foundation models remain unexamined. Furthermore, while the attack remains stealthy to general human reviewers, it may be detectable by advanced statistical anomaly tools that inspect global dataset word distributions. Future work should investigate defenses against stealthy attacks in larger-scale and generative applications.

arXiv: 2205.12700
Cover for BITE: Textual Backdoor Attacks with Iterative Trigger Injection

Abstract

Backdoor attacks have become an emerging threat to NLP systems. By providing poisoned training data, the adversary can embed a “back-door” into the victim model, which allows input instances satisfying certain textual patterns (e.g., containing a keyword) to be predicted as a target label of the adversary’s choice. In this paper, we demonstrate that it is possible to design a backdoor attack that is both stealthy (i.e., hard to notice) and effective (i.e., has a high attack success rate). We propose BITE, a backdoor attack that poisons the training data to establish strong correlations between the target label and a set of “trigger words”. These trigger words are iteratively identified and injected into the target-label instances through natural word-level perturbations. The poisoned training data instruct the victim model to predict the target label on inputs containing trigger words, forming the backdoor. Experiments on four text classification datasets show that our proposed attack is significantly more effective than baseline methods while maintaining decent stealthiness, raising alarm on the usage of untrusted training data. We further propose a defense method named DeBITE based on potential trigger word removal, which outperforms existing methods in defending against BITE and generalizes well to handling other backdoor attacks.1

Table of Contents

  • 1 Introduction
  • 2 Threat Model
  • 3 Methodology
  • 3.1 Bias Measurement on Label Distribution
  • 3.2 Contextualized Word-Level Perturbation
  • 3.3 Poisoning Step
  • 3.4 Training Data Poisoning
  • 3.5 Test-Time Poisoning
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Attack Setting
  • 4.3 Evaluation Metrics for Backdoored Models
  • 4.4 Evaluation Metrics for Poisoned Data
  • 4.5 Compared Methods
  • 5 Experimental Results
  • 5.1 Model Evaluation Results
  • 5.2 Data Evaluation Results
  • 5.3 Effect of Poisoning Rates
  • 5.4 Effect of Operation Limits
  • 6 Defenses against Backdoor Attacks
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgments
  • References
  • A Training Details
  • B Details on Data Evaluation
  • C Results on BERT-Large
  • D Trigger Set and Poisoned Samples
  • D.1 Trigger Set
  • D.2 Poisoned Samples
  • E Computational Costs
  • F Connections with Adversarial Attacks

Knowls

  1. Knowl 1 — BITE clean-label backdoor threat model and attack mechanism

    model/method

    BITE is a poisoning-based textual backdoor attack that modifies training inputs while preserving their original labels. Let X\mathcal{X} be the input space, Y\mathcal{Y} the label space, and DD a distribution over input-label pairs (x,y)∈X×Y(x,y)\in\mathcal{X}\times\mathcal{Y}. An adversary chooses a target label ytarget∈Yy_{\mathrm{target}}\in\mathcal{Y} and a poisoning function T:X→XT:\mathcal{X}\rightarrow\mathcal{X} that injects a trigger pattern. The desired backdoored classifier MbM_b behaves normally on clean inputs but predicts the target label after trigger injection:

    Mb(x)=y,Mb(T(x))=ytarget,(x,y)∼D.M_b(x)=y,\qquad M_b(T(x))=y_{\mathrm{target}},\qquad (x,y)\sim D.

    The adversary can control or contribute training data, but cannot control the victim's model-training process; the experiments use the clean-label setting, so only target-label training instances are poisoned and their labels are not changed. BITE does not use one fixed word or sentence as a trigger. Instead, it iteratively creates a set of words whose training-set frequencies are strongly associated with the target label. Natural contextual insertions and substitutions introduce these words into target-label instances, causing the victim model to learn their spurious association. At test time, the same words are naturally injected into non-target inputs to induce prediction of the target label.

  2. Knowl 2 — Z-score measure of target-label bias

    equation

    BITE measures whether a word is spuriously associated with the target label using a z-score. For a training set containing nn instances, let ntargetn_{\mathrm{target}} be the number of target-label instances and let p0=ntarget/np_0=n_{\mathrm{target}}/n be the target-label proportion expected under an unbiased word-label distribution. For a word ww, let f[w]f[w] be the number of training instances containing ww, let ftarget[w]f_{\mathrm{target}}[w] be the number of those instances with the target label, and let p^(target∣w)=ftarget[w]/f[w]\hat p(\mathrm{target}\mid w)=f_{\mathrm{target}}[w]/f[w]. BITE defines

    z(w)=p^(target∣w)−p0p0(1−p0)/f[w].z(w)=\frac{\hat p(\mathrm{target}\mid w)-p_0}{\sqrt{p_0(1-p_0)/f[w]}}.

    A positive z(w)z(w) indicates that ww occurs with the target label more often than expected from the overall label frequency; larger positive values indicate stronger target-label correlation. During poisoning, the attack selects words according to the maximum z-score they could attain after allowed perturbations.

  3. Knowl 3 — Contextualized word-level perturbation generation

    model/method

    BITE generates candidate poisoning operations with a masked language model LMLM using a mask-then-infill procedure. For every word position in a sentence, it considers both replacing the existing word and inserting a new word at that position. Each operation is represented by an operation type, a position, and a candidate word, such as (Replace,i,w)(\mathrm{Replace},i,w) or (Insert,i,w)(\mathrm{Insert},i,w); the masked language model supplies the candidate's probability.

    The attack discards candidates whose masked-language-model probability is below 0.030.03 or whose application makes the new sentence have cosine similarity below 0.90.9 to the original sentence. Similarity is computed from sentence embeddings produced by all-MiniLM-L6-v2, intended to limit semantic drift and preserve the original label. A dynamic budget BB limits the number of substitution and insertion operations applied to each instance to at most BB times the instance's word count; BITE uses B=0.35B=0.35 in its main experiments. Two operations conflict when they have the same operation type and target the same position, so conflicting operations cannot be applied together. These language-model and similarity constraints are used both to make poisoned training data natural and to make test-time trigger injection fit the input context.

  4. Knowl 4 — Iterative training-set poisoning and trigger-word selection

    algorithm

    BITE jointly selects trigger words and contextual perturbations by maximizing the target-label bias that can be created in the current training set. Let DtrainD_{\mathrm{train}} be the current training set, VV its vocabulary, TT the ordered list of already selected trigger words, K=V∖TK=V\setminus T the unused candidate words, and PtrainP_{\mathrm{train}} the filtered, non-conflicting contextual operations generated for the training set. For a candidate word xx, z(x;Dtrain,Pselect)z(x;D_{\mathrm{train}},P_{\mathrm{select}}) is the z-score after applying a selected operation set PselectP_{\mathrm{select}}. The ideal step solves

    max⁡x∈K max⁡Pselect⊆Ptrainz(x;Dtrain,Pselect).\max_{x\in K}\ \max_{P_{\mathrm{select}}\subseteq P_{\mathrm{train}}} z(x;D_{\mathrm{train}},P_{\mathrm{select}}).

    Because only target-label instances may be poisoned, the inner optimum is obtained by selecting all allowed operations that introduce the candidate word into target-label instances. The algorithm therefore computes each candidate's non-target frequency and its maximum attainable target frequency, chooses the candidate with the highest attainable positive z-score, applies its corresponding operations, and repeats. Operations are restricted to unused words so that a selected trigger's frequency does not change in later iterations.

    Input: Training set DtrainD_{\mathrm{train}}, vocabulary VV, masked language model LMLM, target label
    Output: Poisoned training set DtrainD_{\mathrm{train}}, ordered trigger list TT
    Initialize TT as an empty list
    while true do
        Set KK to V∖TV \setminus T
        Generate filtered contextual operations PtrainP_{\mathrm{train}} on DtrainD_{\mathrm{train}} using LMLM and words in KK
        For every ww in KK, count its non-target frequency fnon[w]f_{\mathrm{non}}[w]
        For every ww in KK, compute its maximum target frequency ftarget[w]f_{\mathrm{target}}[w] by selecting all operations that introduce ww into target-label instances
        Compute each candidate's attainable z-score from its two frequencies and the current training-set label proportion
        Select the candidate tt with the highest attainable z-score
        if the highest attainable z-score is not positive then
            break
        Append tt to TT
        Select all non-conflicting operations that introduce tt into eligible target-label instances
        Apply the selected operations to update DtrainD_{\mathrm{train}}
    return DtrainD_{\mathrm{train}} and TT

    The result is a poisoned training set and an ordered trigger list; the procedure terminates when no unused word can acquire a positive attainable target-label bias.

  5. Knowl 5 — Iterative test-time trigger injection

    algorithm

    Given a clean test sentence xx, the BITE test-time procedure injects the learned trigger words in their training-time order. It starts with the full vocabulary as the set of available trigger candidates and generates contextual operations for the current sentence using the same masked-language-model filtering, semantic-similarity threshold, and dynamic operation budget used during training. For each trigger word tt in the ordered trigger list, it applies the available non-conflicting operations that introduce tt if such operations exist. After successful injection, tt is removed from the candidate set so later operations cannot alter the already introduced trigger, and the operation set is regenerated for the updated sentence.

    Input: Test sentence xx, vocabulary VV, masked language model LMLM, ordered trigger list TT
    Output: Poisoned test sentence xx
    Set KK to VV
    Generate filtered contextual operations PP for xx using LMLM and KK
    for each trigger word tt in TT do
        Select non-conflicting operations in PP that introduce tt
        if the selected operation set is not empty then
            Apply the selected operations to update xx
            Remove tt from KK
            Regenerate PP for the updated xx using LMLM and KK
    return xx

    The resulting sentence contains as many high-priority trigger words as can be naturally fitted into its context. The backdoored classifier is then queried once on this transformed input; unlike an ordinary adversarial search, test-time injection does not require repeatedly querying the victim model.

  6. Knowl 6 — DeBITE training-time defense by label-correlation filtering

    model/method

    DeBITE defends against textual backdoors by removing words with unusually strong association to any label from the training set. For every vocabulary word ww and every possible label ℓ∈Y\ell\in\mathcal{Y}, let nℓn_\ell be the number of training instances with label ℓ\ell, nn the total number of training instances, pℓ=nℓ/np_\ell=n_\ell/n, f[w]f[w] the number of instances containing ww, and fℓ[w]f_\ell[w] the number of those instances labeled ℓ\ell. DeBITE computes

    zℓ(w)=fℓ[w]/f[w]−pℓpℓ(1−pℓ)/f[w],zmax⁡(w)=max⁡ℓ∈Yzℓ(w).z_\ell(w)=\frac{f_\ell[w]/f[w]-p_\ell}{\sqrt{p_\ell(1-p_\ell)/f[w]}},\qquad z_{\max}(w)=\max_{\ell\in\mathcal{Y}}z_\ell(w).

    Every word with zmax⁡(w)>3z_{\max}(w)>3 is treated as a potential trigger, and all occurrences of those words are removed from the training data before model training. The threshold 33 was selected according to the tolerated drop in clean accuracy. This defense targets the word-label correlation created by BITE and can also remove words associated with non-word-level trigger patterns when those patterns induce strongly biased lexical distributions.

  7. Knowl 7 — Datasets, attack setting, and evaluation protocol

    experimental setup

    The experiments evaluate textual backdoors on four single-sentence classification datasets. The dataset summary printed on paper page 5 is:

    Could not parse LaTeX table

    SST-2 is binary movie-review sentiment classification, HateSpeech is binary forum hate-speech detection, Tweet is four-class emotion recognition, and TREC is six-class question classification. The main attack experiments poison 1%1\% of the training data, preserve all labels, and use the first label in each dataset as the target: positive for SST-2, clean for HateSpeech, anger for Tweet, and abbreviation for TREC. BERT-Base is the victim model, and checkpoints are selected by clean development-set accuracy. The reported training configuration uses batch size 3232, 1313 epochs, and a learning rate that increases linearly from 00 to 2×10−52\times10^{-5} during the first three epochs and then decreases linearly to 00.

    Attack Success Rate (ASR) is the percentage of non-target test instances classified as the target label after test-time poisoning. Clean Accuracy (CACC) is accuracy on the unmodified test set and measures whether the backdoored model remains benign on clean inputs. BITE is compared with StyleBkd, which uses a Bible-style trigger, and Hidden Killer, which uses a low-frequency syntactic trigger. BITE Full estimates word-label statistics from the full training set; BITE Subset estimates them only from the contributed poisoned subset.

  8. Knowl 8 — BITE effectiveness with minimal clean-accuracy degradation

    empirical result

    The model comparison printed on paper page 6 shows that BITE Full produces substantially higher ASR than the stealthy Style and Syntactic baselines while leaving clean accuracy nearly unchanged. Values are mean ±\pm standard deviation, and ASR and CACC are percentages; columns are ordered as SST-2, HateSpeech, Tweet, and TREC.

    Could not parse LaTeX table
    Could not parse LaTeX table

    BITE Full is especially strong on SST-2, Tweet, and TREC, and is competitive with or better than the strongest baseline on HateSpeech. BITE Subset remains stronger than both baselines on SST-2 and TREC but is less effective than BITE Full, showing that accurate full-training-set bias estimation helps trigger selection. Replication with BERT-Large preserves the trend: BITE Full reaches ASR 61.3±1.961.3\pm1.9, 73.0±3.773.0\pm3.7, 46.6±2.046.6\pm2.0, and 53.8±2.753.8\pm2.7 on the four datasets, while its CACC remains close to the benign model.

  9. Knowl 9 — Stealthiness and controllable effectiveness of BITE

    empirical result

    The data-level comparison on paper page 7 evaluates poisoned SST-2 instances along four dimensions: automatic grammatical naturalness, human suspicion, human-rated semantic similarity on a 1–3 scale, and human-rated label consistency. Higher is preferred for naturalness, similarity, and consistency; lower is preferred for suspicion. The reported values are:

    Could not parse LaTeX table

    Style produces the most natural and least suspicious poisoned text, whereas Syntactic performs worst on all four quality dimensions except that its higher suspicion score indicates poorer stealthiness. BITE has the highest semantic-similarity score, nearly matches Style in label consistency, and provides a substantially better quality profile than Syntactic.

    BITE's dynamic budget controls the effectiveness–stealthiness trade-off. On SST-2, increasing BB from 0.050.05 to 0.500.50 in increments of 0.050.05 introduces more trigger words and raises ASR, but also lowers poisoned-text naturalness. Across poisoning rates of 1%1\%, 3%3\%, 5%5\%, and 10%10\%, all attacks become more effective as the poisoning rate increases; BITE Full maintains the strongest advantage at lower rates because it exploits pre-existing dataset word-label biases.

  10. Knowl 10 — DeBITE reduces stealthy backdoor success while preserving clean accuracy

    empirical result

    The defense experiment uses SST-2 with a 5%5\% poisoning rate and compares no defense with inference-time defenses ONION, STRIP, and RAP, and training-time defenses CUBE, BKI, and DeBITE. The numerical comparison printed on paper page 8 is reproduced below. Parenthesized values are changes from the corresponding undefended attack.

    Could not parse LaTeX table

    DeBITE gives the largest reduction against the Syntactic and BITE attacks, lowering BITE Full ASR from 66.266.2 to 56.756.7 and Syntactic ASR from 49.949.9 to 33.9.ItalsoslightlyreducesStyleASR.ItsCACCdecreasesbyonly33.9. It also slightly reduces Style ASR. Its CACC decreases by only 0.8––1.0$ percentage points, whereas several inference-time defenses incur larger clean-accuracy losses. ASR remains non-negligible after defense, so the paper treats DeBITE as effective but not a complete solution to stealthy clean-label backdoors.

Coverage note — The paper's stated limitations, detailed poisoned examples, computational-cost measurements, and most appendix-only BERT-Large results were omitted to prioritize the core attack, defense, algorithms, and main quantitative findings.

References

  1. 1.Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath. 2021. {T-Miner}: A generative approach to defend against {DNN-based} text classification. In 30th USENIX Security Symposium (USENIX Security 21), pages 2255–2272.
  2. 2.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650.
  3. 3.Alvin Chan, Yi Tay, Yew-Soon Ong, and Aston Zhang. 2020. Poison attacks against text datasets with conditional adversarially regularized autoencoder. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4175–4189, Online. Association for Computational Linguistics.
  4. 4.Chuanshuai Chen and Jiazhu Dai. 2021. Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification. Neurocomputing, 452:253–262.
  5. 5.Lichang Chen, Minhao Cheng, and Heng Huang. 2023. Backdoor learning on sequence to sequence models. arXiv preprint arXiv:2305.02424.
  6. 6.Sishuo Chen, Wenkai Yang, Zhiyuan Zhang, Xiaohan Bi, and Xu Sun. 2022a. Expose backdoors on the way: A feature-based efficient defense against textual backdoor attacks. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 668–683, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  7. 7.Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang. 2021. Badnl: Backdoor attacks against nlp models. In ICML 2021 Workshop on Adversarial Machine Learning.
  8. 8.Yangyi Chen, Fanchao Qi, Hongcheng Gao, Zhiyuan Liu, and Maosong Sun. 2022b. Textual backdoor attacks can be more harmful via two simple tricks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11215–11221, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  9. 9.Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022. A unified evaluation of textual backdoor learning: Frameworks and benchmarks. In Proceedings of NeurIPS: Datasets and Benchmarks.
  10. 10.Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019. A backdoor attack against lstm-based text classification systems. IEEE Access, 7:138872–138878.
  11. 11.Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018. Hate speech dataset from a white supremacy forum. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2), pages 11–20, Brussels, Belgium. Association for Computational Linguistics.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. HotFlip: White-box adversarial examples for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 31–36, Melbourne, Australia. Association for Computational Linguistics.
  14. 14.Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith C Ranasinghe, and Hyoungshick Kim. 2021. Design and evaluation of a multi-domain trojan detection method on deep neural networks. IEEE Transactions on Dependable and Secure Computing, 19(4):2349–2364.
  15. 15.Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021. Competency problems: On finding and removing artifacts in language data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1801–1813, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  16. 16.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  17. 17.Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, and Ruoxi Jia. 2022. CATER: Intellectual property protection on text generation APIs via conditional watermarks. In Advances in Neural Information Processing Systems.
  18. 18.Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran. 2001. Toward semantics-based answer pinpointing. In Proceedings of the First International Conference on Human Language Technology Research.
  19. 19.Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. Adversarial example generation with syntactically controlled paraphrase networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1875–1885, New Orleans, Louisiana. Association for Computational Linguistics.
  20. 20.Praphula Kumar Jain, Rajendra Pamula, and Gautam Srivastava. 2021. A systematic literature review on machine learning applications for consumer sentiment analysis using online reviews. Computer Science Review, 41:100413.
  21. 21.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021–2031, Copenhagen, Denmark. Association for Computational Linguistics.
  22. 22.Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2022. Deep Learning for Text Style Transfer: A Survey. Computational Linguistics, 48(1):155–205.
  23. 23.Wasiat Khan, Mustansar Ali Ghazanfar, Muhammad Awais Azam, Amin Karami, Khaled H Alyoubi, and Ahmed S Alfakeeh. 2020. Stock market prediction using machine learning classifiers and social media, news. Journal of Ambient Intelligence and Humanized Computing, pages 1–24.
  24. 24.Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. 2020a. Thieves on sesame street! model extraction of bert-based apis. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  25. 25.Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020b. Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 737–762, Online. Association for Computational Linguistics.
  26. 26.Keita Kurita, Paul Michel, and Graham Neubig. 2020. Weight poisoning attacks on pretrained models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2793–2806, Online. Association for Computational Linguistics.
  27. 27.Hyun Kwon and Sanghyun Lee. 2021. Textual backdoor attack for the text classification system. Security and Communication Networks, 2021.
  28. 28.Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2021a. Contextualized perturbation for textual adversarial attack. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5053–5069, Online. Association for Computational Linguistics.
  29. 29.Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. BERT-ATTACK: Adversarial attack against BERT using BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6193–6202, Online. Association for Computational Linguistics.
  30. 30.Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021b. Neural attention distillation: Erasing backdoor triggers from deep neural networks. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  31. 31.Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018. Fine-pruning: Defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, pages 273–294. Springer.
  32. 32.Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang. 2022. Piccolo: Exposing complex backdoors in nlp transformer models. In 2022 IEEE Symposium on Security and Privacy (SP), pages 1561–1561. IEEE Computer Society.
  33. 33.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. SemEval-2018 task 1: Affect in tweets. In Proceedings of the 12th International Workshop on Semantic Evaluation, pages 1–17, New Orleans, Louisiana. Association for Computational Linguistics.
  34. 34.Tianrui Peng, Ian Harris, and Yuki Sawa. 2018. Detecting phishing attacks using natural language processing and machine learning. In 2018 IEEE 12th international conference on semantic computing (icsc), pages 300–301. IEEE.
  35. 35.Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021a. ONION: A simple and effective defense against textual backdoor attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9558–9566, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  36. 36.Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, and Maosong Sun. 2021b. Mind the style of text! adversarial and backdoor attacks based on text style transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4569–4580, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. 2021c. Hidden killer: Invisible textual backdoor attacks with syntactic trigger. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 443–453, Online. Association for Computational Linguistics.
  38. 38.Fanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu, and Maosong Sun. 2021d. Turn the combination lock: Learnable textual backdoor attacks via word substitution. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4873–4883, Online. Association for Computational Linguistics.
  39. 39.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  40. 40.Anna Schmidt and Michael Wiegand. 2017. A survey on hate speech detection using natural language processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, pages 1–10, Valencia, Spain. Association for Computational Linguistics.
  41. 41.Guangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu, Zhuo Zhang, Shengwei An, Shiqing Ma, and Xiangyu Zhang. 2022. Constrained optimization with dynamic bound-scaling for effective NLP backdoor defense. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 19879–19892. PMLR.
  42. 42.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  43. 43.Jiao Sun, Xuezhe Ma, and Nanyun Peng. 2021. AESOP: Paraphrase generation with adaptive syntactic control. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5176–5189, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  44. 44.Lichao Sun. 2020. Natural backdoor attack on text data. ArXiv preprint, abs/2006.16176.
  45. 45.Jiayi Wang, Rongzhou Bao, Zhuosheng Zhang, and Hai Zhao. 2022. Rethinking textual adversarial defense for pre-trained language models. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:2526–2540.
  46. 46.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  47. 47.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  48. 48.Yuxiang Wu, Matt Gardner, Pontus Stenetorp, and Pradeep Dasigi. 2022. Generating data to mitigate spurious correlations in natural language inference datasets. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2660–2676, Dublin, Ireland. Association for Computational Linguistics.
  49. 49.Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li. 2021. Detecting ai trojans using meta neural analysis. In 2021 IEEE Symposium on Security and Privacy (SP), pages 103–120. IEEE.
  50. 50.Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He. 2021a. Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2048–2058, Online. Association for Computational Linguistics.
  51. 51.Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021b. RAP: Robustness-Aware Perturbations for defending against backdoor attacks on NLP models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8365–8381, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  52. 52.Zhengyan Zhang, Guangxuan Xiao, Yongwei Li, Tian Lv, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Xin Jiang, and Maosong Sun. 2021. Red alarm for pre-trained models: Universal vulnerability to neuron-level backdoor attacks. ArXiv preprint, abs/2101.06969.
  53. 53.Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Chong Fu, Yangdong Deng, Zhiyuan Liu, Jingang Wang, Wei Wu, et al. 2022. Moderate-fitting as a natural backdoor defender for pre-trained language models. Advances in Neural Information Processing Systems, 35:1086–1099.

Citation

MLA
Yan, J., et al. “BITE: Textual Backdoor Attacks with Iterative Trigger Injection”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12951–68, https://doi.org/10.18653/v1/2023.acl-long.725.
APA
Yan, J., Gupta, V., & Ren, X. (2023). BITE: Textual Backdoor Attacks with Iterative Trigger Injection. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12951–12968. https://doi.org/10.18653/v1/2023.acl-long.725
Chicago
Yan, J., V. Gupta, and X. Ren. 2023. “BITE: Textual Backdoor Attacks with Iterative Trigger Injection”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12951–68. https://doi.org/10.18653/v1/2023.acl-long.725.
Harvard
Yan, J., Gupta, V. and Ren, X. (2023) “BITE: Textual Backdoor Attacks with Iterative Trigger Injection”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12951–12968. Available at: https://doi.org/10.18653/v1/2023.acl-long.725.
Vancouver
1. Yan J, Gupta V, Ren X (2023) BITE: Textual Backdoor Attacks with Iterative Trigger Injection. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12951–12968

BibTeX

@inproceedings{yan-etal-2023-bite,
    title = "{BITE}: Textual Backdoor Attacks with Iterative Trigger Injection",
    author = "Yan, Jun  and
      Gupta, Vansh  and
      Ren, Xiang",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.725/",
    doi = "10.18653/v1/2023.acl-long.725",
    pages = "12951--12968"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/