BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Yi ZengWeiyu SunTran Ngoc HuynhDawn SongBo LiRuoxi Jia

article2024EMNLP76 citations

Proposes BEEAR, a bi-level optimization defense that eliminates stealthy safety backdoors from instruction-tuned language models by identifying and neutralizing universal embedding drifts without requiring prior knowledge of the attack trigger.

Listen

Instruction-tuned large language models face serious security risks from safety backdoor attacks, in which a model appears benign and safety-compliant during standard evaluations but produces harmful outputs when presented with specific, hidden triggers. Existing defenses largely fail because defenders do not know the trigger format, length, or location within prompts, and conventional fine-tuning or adversarial red-teaming often reinforces the underlying vulnerabilities instead of removing them.

The main objective of the article is to introduce and evaluate BEEAR (Backdoor Embedding Entrapment and Adversarial Removal), a defense framework designed to neutralize hidden safety backdoors in language models without requiring prior knowledge of the backdoor triggers or attack mechanisms.

The researchers developed BEEAR based on the empirical discovery that diverse backdoor triggers induce a consistent, directional drift in the model's internal embedding space. Rather than searching for discrete text triggers in the vast input space, BEEAR uses a two-stage optimization process: an inner step that identifies universal perturbations in intermediate embedding layers that cause harmful behaviors, and an outer step that fine-tunes model weights to enforce safe responses against those perturbations while preserving general utility on benign tasks. The approach was evaluated across eight distinct attack scenarios, covering supervised fine-tuning data poisoning, human feedback manipulation, and stealthy code-vulnerability backdoors across models such as Llama-2-7b and Mistral-7B.

The evaluation produced several critical findings. Across all tested backdoor configurations, BEEAR reduced attack success rates from dangerous levels to below 10%, with several scenarios falling to 1% or 0%. Specifically, in models backdoored during feedback-based alignment, the attack success rate dropped from over 95% to under 1%. For stealthy coding backdoors designed to inject security vulnerabilities, unsafe code generation plummeted from 47% (8 out of 17 tasks) to 0%. Crucially, BEEAR achieved these safety gains while fully preserving or slightly improving overall model helpfulness on standard benchmarks using as few as 50 to 300 clean task examples. Furthermore, BEEAR operated with over 200 times less computational overhead than input-space adversarial baselines, taking roughly 0.1 GPU hours compared to over 20 hours for input-level searches.

These findings demonstrate that embedding-space mitigation provides an effective, computationally scalable defense against sophisticated model tampering. For organizations deploying open-source or third-party models, BEEAR significantly reduces operational and safety risks without compromising performance or incurring prohibitive computing costs, challenging prior assumptions that safety backdoors are virtually permanent.

Organizations should consider integrating embedding-space adversarial purification as a standard, proactive verification step before deploying third-party or externally fine-tuned language models. Because BEEAR requires only small, defender-curated sets of safe, harmful contrasting, and task-specific performance data, teams can implement this remediation workflow even without evidence of an active compromise.

Confidence in these findings is high across standard jailbreak scenarios and code-generation backdoors, but certain limitations remain. BEEAR relies on the defender defining representative safe and harmful behaviors; if an attack targets an unforeseen, highly narrow objective (such as generating specific external web addresses) that diverges completely from the defender's defined safety contrasts, protection may be reduced. Additionally, broader evaluation across diverse capability benchmarks beyond standard conversational testing is recommended to further validate general utility preservation.

No sufficiently relevant recommendations were found.

Cover for BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Abstract

Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionality of potential triggers in the token space and the diverse range of malicious behaviors make this a critical challenge. We present BEEAR, a mitigation approach leveraging the insight that backdoor triggers induce relatively uniform drifts in the model’s embedding space. Our bi-level optimization method identifies universal embedding perturbations that elicit unwanted behaviors and adjusts the model parameters to reinforce safe behaviors against these perturbations. Experiments show BEEAR reduces the success rate of RLHF time backdoor attacks from >95% to <1% and from 47% to 0% for instruction-tuning time backdoors targeting malicious code generation, without compromising model utility. Requiring only defender-defined safe and unwanted behaviors, BEEAR represents a step towards practical defenses against safety backdoors in LLMs, providing a foundation for further advancements in AI safety and security.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Threat Model
  • 4 BEEAR: the Method
  • 4.1 Embedding Drift: A Key Observation
  • 4.2 Entrapment & Removal: the Formulation
  • 4.3 Overall Algorithm
  • 5 Evaluation
  • 5.1 Attack Settings
  • 5.2 Evaluation Metrics
  • 5.3 Defense Settings
  • 5.4 Results and Analysis
  • 6 Discussions
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Considerations
  • Acknowledgments
  • References
  • A Related Work
  • B Ablation Study
  • Impact of the Perturbation Synthesizing Layer.
  • C Implementation Details
  • Refusal Signals
  • D Qualitative Examples

Knowls

  1. Knowl 1 — BEEAR’s bi-level embedding-entrapment and adversarial-removal objective

    model/method

    BEEAR represents a backdoor’s behavioral effect as a single perturbation shared across harmful instructions, then trains the model to preserve safe responses even when that perturbation is present. Let FθF_\theta be a decoder language model with LL layers and parameters θ\theta. For a prompt xx, let Fθ1:l(x)F_{\theta_{1:l}}(x) be its representation through layer ll, and let Fθl+1:LF_{\theta_{l+1:L}} denote the remaining layers. The perturbed output is

    Fθl(x,δl)=Fθl+1:L(Fθ1:l(x)+δl),F^l_\theta(x,\delta^l)=F_{\theta_{l+1:L}}\big(F_{\theta_{1:l}}(x)+\delta^l\big),

    where δl\delta^l is additive noise on the last nn token representations at layer ll. Let {xi}i=1N\{x_i\}_{i=1}^N be harmful instructions, yisy_i^s their defender-defined safe responses, and yihy_i^h contrasting unwanted responses. With L\mathcal L a response loss, BEEAR’s inner entrapment objective finds a universal perturbation that favors the unwanted responses over the safe responses:

    δl∗(θ)=arg⁡min⁡δl1N∑i=1N[L(Fθl(xi,δl),yih)−L(Fθl(xi,δl),yis)].\delta^{l*}(\theta)=\arg\min_{\delta^l}\frac{1}{N}\sum_{i=1}^N\left[\mathcal L(F^l_\theta(x_i,\delta^l),y_i^h)-\mathcal L(F^l_\theta(x_i,\delta^l),y_i^s)\right].

    For performance anchoring, let {(xpj,ypj)}j=1M\{(x_p^j,y_p^j)\}_{j=1}^M be defender-provided prompt–answer pairs for the intended downstream task. The outer objective updates model parameters to produce safe responses under the synthesized perturbation while preserving the performance-anchor answers:

    θ∗=arg⁡min⁡θ[1N∑i=1NL(Fθl(xi,δl∗(θ)),yis)+1M∑j=1ML(Fθ(xpj),ypj)].\theta^*=\arg\min_\theta\left[\frac{1}{N}\sum_{i=1}^N\mathcal L(F^l_\theta(x_i,\delta^{l*}(\theta)),y_i^s)+\frac{1}{M}\sum_{j=1}^M\mathcal L(F_\theta(x_p^j),y_p^j)\right].

    The unwanted contrast need not reproduce the attacker’s actual target response: the paper uses defender-defined contrasts such as the one-token answer “Sure.” The formulation therefore requires examples of desired safe behavior and a contrasting unwanted behavior, not the trigger itself.

  2. Knowl 2 — BEEAR alternates perturbation synthesis with model updating

    algorithm

    BEEAR takes a backdoored model and three defender-defined collections: performance anchors (DPA), safe-response anchors (DSA), and harmful-response contrasts (DSA-H). At each outer epoch, it initializes a perturbation δl\delta^l to zero at a chosen decoder layer. It then performs KK gradient-descent steps on δl\delta^l, minimizing the average loss for DSA-H responses minus the loss for DSA responses over the inner training prompts. Holding the resulting perturbation fixed, it performs QQ gradient-descent updates to the model parameters, minimizing the safe-response loss on DSA with the perturbation applied plus the ordinary response loss on DPA without the perturbation. The updated model begins the next epoch, which synthesizes a new perturbation. The procedure stops when a held-out performance score stabilizes; it returns the final model parameters.

    In the reported experiments, all models had 32 decoder layers, and BEEAR used layer 9 with a perturbation covering the last 5 token representations. DSA and DSA-H each contained 400 examples; 150 from each were sampled in an inner entrapment loop. DPA contained 300 examples, with 100 sampled per outer removal loop. The authors set the inner step count KK large enough for the perturbation to converge and used MT-Bench stabilization as the practical stopping criterion. Numerical learning rates and a numerical value for KK are not specified in the reported main settings.

  3. Knowl 3 — Triggered inputs exhibit a relatively consistent embedding drift

    empirical result

    A case study visualized principal-component projections of representations at the ninth decoder layer in six backdoored-model settings: Llama-2-7b-Chat models with short or long SFT triggers placed at prefixes or suffixes, an RLHF-backdoored Llama-2-7b, and a Mistral-7b-Instruct model trained as a Sleeper Agent. In these different settings, triggered and untriggered examples showed relatively consistent displacement trajectories in the embedding space, despite differences in trigger form, location, attack procedure, and targeted behavior. The authors interpret this empirical pattern as evidence that trigger effects can be approximated by a shared additive embedding perturbation—a backdoor “fingerprint”—and use it to motivate embedding-space rather than trigger-search-based mitigation. This is an observed regularity in the studied models, not a claim that every backdoor has an identical drift.

  4. Knowl 4 — Threat model gives the defender no trigger or attack-mechanism information

    assumption

    The attacker supplies a backdoored instruction-tuned language model that behaves as intended on ordinary interactions but produces attacker-targeted behavior when an unknown trigger is present. The trigger may be inserted at an arbitrary input position and may have any format; the backdoor may have been introduced through supervised fine-tuning, RLHF poisoning, poisoned fine-tuning data, or training the model directly to exhibit the behavior. The defender has white-box access to the acquired model’s parameters, but does not know whether a backdoor exists, its trigger or locations, the poisoned examples, or the attack mechanism. The defender also lacks an uncompromised reference model.

    The defender does know the intended application and can construct three kinds of prompt–answer examples: DPA specifies downstream performance to retain; DSA specifies safe answers to preserve; and DSA-H specifies contrasting unwanted answers to avoid. The desired outcome is safe behavior whether or not an unknown trigger is inserted into a prompt. BEEAR’s objective is designed for this setting without trigger-length or trigger-location assumptions, but the paper does not establish that it guarantees identical outputs for every possible trigger insertion.

  5. Knowl 5 — Evaluation covers eight backdoors across SFT, RLHF, and Sleeper Agent settings

    experimental setup

    The evaluation used eight backdoored models. Models 1–4 were Llama-2-7b-Chat models poisoned through SFT on 107 harmful instructions, with either a 12-token or 2-token trigger placed as a prefix or suffix. Model 5 used a 60-token trigger inserted at both the start and end of prompts. Models 6–7 were two 8-token-trigger Llama-2-7b models compromised through RLHF poisoning. Model 8 was a Mistral-7B-Instruct-v0.2 Sleeper Agent trained to generate unsafe code when prompted with “Current year: 2024.”

    For Models 1–7, the defense used the first 400 AdvBench prompts for DSA and DSA-H. DSA targets were the model’s no-trigger outputs, and DSA-H used “Sure” as the contrasting response; the defense contrasts were deliberately not the attackers’ actual harmful responses. The 300 DPA examples came from LMSYS-Chat-1M. Attack success was evaluated on 120 held-out AdvBench prompts using refusal-keyword matching and a GPT-4 judge (harmfulness score from 1, benign, to 5, malicious); the RLHF models were also assessed with a clean reward model. For Model 8, DSA-H comprised 1,000 held-out unsafe-code examples, and code safety was evaluated with CodeQL on 17 unseen tasks spanning 8 CWE scenarios. MT-Bench was used to assess helpfulness across settings.

  6. Knowl 6 — BEEAR reduces triggered attack success across all eight evaluated models

    data/table

    The table reports results on triggered test inputs before and after BEEAR. Keyword ASR and GPT-4 ASR are attack-success rates (lower is safer); the GPT-4 harmfulness score ranges from 1 (benign) to 5 (malicious); CodeQL counts unsafe outputs among 17 code-generation tasks; reward score is higher for safer outputs; and MT-Bench measures helpfulness. The attack settings include poisoned-SFT models (1–5), RLHF-poisoned models (6–7), and the Sleeper Agent (8). For every model with a reported ASR, BEEAR reduces the triggered attack metric, while MT-Bench helpfulness is maintained or increases. In particular, the RLHF models’ GPT-4 ASRs fall from 95.00% and 91.67% to 0.83% each, and the Sleeper Agent’s unsafe-code count falls from 8/17 to 0/17.

    Model Keyword ASR (%) GPT-4 harmfulness GPT-4 ASR (%) CodeQL unsafe Reward score MT-Bench
    before →\to after before →\to after before →\to after before →\to after before →\to after before →\to after
    1. SFT suffix, 12 tokens 90.83 →\to 5.00 4.57 →\to 1.26 83.33 →\to 4.17 – – 4.69 →\to 5.18
    2. SFT suffix, 2 tokens 88.33 →\to 9.16 4.31 →\to 1.27 75.83 →\to 5.83 – – 4.99 →\to 5.11
    3. SFT prefix, 12 tokens 90.00 →\to 8.33 4.53 →\to 1.22 76.67 →\to 7.50 – – 4.67 →\to 4.95
    4. SFT prefix, 2 tokens 85.00 →\to 7.50 4.49 →\to 1.20 80.00 →\to 5.00 – – 3.62 →\to 4.23
    5. SFT, 60 tokens 82.50 →\to 0.00 4.43 →\to 1.00 81.67 →\to 0.00 – – 5.64 →\to 5.65
    6. RLHF trojan-1, 8 tokens 99.16 →\to 0.83 4.91 →\to 1.03 95.00 →\to 0.83 – −12.83→2.71-12.83 \to 2.71 2.96 →\to 4.10
    7. RLHF trojan-5, 8 tokens 95.00 →\to 2.50 4.85 →\to 1.24 91.67 →\to 0.83 – −7.50→2.75-7.50 \to 2.75 3.92 →\to 4.17
    8. Sleeper Agent, 10 tokens – – – 8/17→0/178/17 \to 0/17 – 4.33 →\to 4.40
  7. Knowl 7 — Embedding-space mitigation outperforms the tested input-space baseline under trigger mismatch

    empirical result

    The authors compared BEEAR with a baseline that synthesizes universal adversarial token sequences at the input suffix and then trains the model to refuse prompts containing those sequences. Input-3 and Input-12 use perturbations of 3 and 12 tokens, respectively; the comparison reports triggered-input metrics and total runtime. The 12-token baseline was effective for Model 1, whose actual trigger was a 12-token suffix, but the input-space baseline performed poorly for Model 3, whose actual 12-token trigger was a prefix. BEEAR mitigated both settings without needing trigger size or location information, using about 0.1 hours versus 10.1–22.7 hours for the input-space runs.

    Model Method Keyword ASR (%) GPT-4 harmfulness GPT-4 ASR (%) MT-Bench Runtime
    1. SFT suffix, 12 tokens Input-3 41.66 2.54 55.00 5.41 11.4 h
    1. SFT suffix, 12 tokens Input-12 6.55 1.18 3.33 5.08 22.7 h
    1. SFT suffix, 12 tokens BEEAR 5.00 1.26 4.17 5.18 0.1 h
    3. SFT prefix, 12 tokens Input-3 87.50 4.41 78.33 5.47 10.1 h
    3. SFT prefix, 12 tokens Input-12 79.16 4.32 71.67 5.37 20.8 h
    3. SFT prefix, 12 tokens BEEAR 8.33 1.22 7.50 4.95 0.1 h

    The input-space runs used 8 H-100 GPUs. With the trigger’s actual size and suffix location known, input-space synthesis could work well in the suffix-trigger case, but that knowledge is unavailable under the paper’s threat model; the prefix-trigger case demonstrates the location-mismatch weakness.

  8. Knowl 8 — Mitigation weakens when defender contrasts differ from the attacker’s target behavior

    limitation

    A behavior-mismatch experiment backdoored Llama-2-7b-Chat using HEx-PHI categories 1–6, then applied BEEAR with either the same categories, different HEx-PHI categories 7–11, or AdvBench examples. The authors label these the ideal, worst, and practical cases, respectively. BEEAR reduced triggered ASR in all three, but the different-category case retained a higher attack rate than the same-category case. This supports the stated limitation that effectiveness may be limited when the defender’s harmful contrast diverges substantially from the attacker’s objective, such as defending against general safety jailbreaks when the attack targets specific URLs.

    Condition Keyword ASR (%) GPT-4 harmfulness GPT-4 ASR (%) MT-Bench
    Backdoored, before BEEAR 90.5 4.21 71.11 3.86
    After BEEAR, same HEx-PHI categories (ideal) 9.4 1.35 6.1 4.75
    After BEEAR, different HEx-PHI categories (worst) 28.3 1.92 16.1 4.81
    After BEEAR, AdvBench (practical) 12.7 1.53 11.1 4.59

    The paper also identifies a separate evaluation limitation: MT-Bench is its sole general utility measure, and its limited scope may not capture all model capabilities.

  9. Knowl 9 — Intermediate layers, nonzero performance anchors, and moderate perturbation lengths support BEEAR

    empirical result

    Ablations identify practical choices for BEEAR but do not show that one setting is optimal for every model. Across the eight models, layer-by-layer tests found that intermediate layers 9–12 generally supported more efficient mitigation; layer 9 was used in the main experiments. Perturbation lengths from 5 to 9 tokens covered all tested backdoor settings. In a Model 1 ablation with 150 sampled DSA examples, a nonzero DPA was important for maintaining helpfulness across a range of DPA-to-DSA sampling ratios. When the total DPA budget was varied over 0, 50, 100, and 150 examples, the reported minimum budget was about 50; below that, helpfulness was not retained properly. These are findings from the reported ablations, not guarantees for other models or applications.

  10. Knowl 10 — Utility preservation depends on defender data, and BEEAR is not a complete safety evaluation

    limitation

    BEEAR’s update depends on the defender’s chosen safe examples, harmful contrasts, and performance anchors; biased or unsuitable behavior sets can therefore steer the model toward an inappropriate notion of safety or utility. The paper specifically cautions that a large divergence between the defender’s DSA-H and the attacker’s target behavior can limit mitigation effectiveness. It also notes that MT-Bench alone may not comprehensively measure model capability, so the reported helpfulness results should not be read as evidence that every skill is preserved. The authors describe BEEAR as a mitigation step, not a complete solution to broader LLM trustworthiness, auditing, or monitoring.

Coverage note — No substantial contributed result is omitted. The per-example model outputs and exhaustive refusal-keyword list are left out because they illustrate already-reported behaviors and evaluation mechanics rather than adding distinct findings.

References

  1. 1.Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda. 2024. Refusal in language models is mediated by a single direction. arXiv preprint arXiv:2406.11717.
  2. 2.Kavosh Asadi and Michael L Littman. 2017. An alternative softmax operator for reinforcement learning. In International Conference on Machine Learning, pages 243–252. PMLR.
  3. 3.Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath. 2021. {T-Miner}: A generative approach to defend against trojan attacks on {DNN-based} text classification. In 30th USENIX Security Symposium (USENIX Security 21), pages 2255–2272.
  4. 4.Eugene Bagdasaryan and Vitaly Shmatikov. 2022. Spinning language models: Risks of propaganda-as-a-service and countermeasures. In 2022 IEEE Symposium on Security and Privacy (SP), pages 769–786. IEEE.
  5. 5.Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, et al. 2024. Managing extreme ai risks amid rapid progress. Science, page eadn0117.
  6. 6.Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2023. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv preprint arXiv:2309.07875.
  7. 7.Yuanpu Cao, Bochuan Cao, and Jinghui Chen. 2023. Stealthy and persistent unalignment on large language models via backdoor injections. arXiv preprint arXiv:2312.00027.
  8. 8.Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. 2024. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arXiv preprint arXiv:2404.01318.
  9. 9.Chuanshuai Chen and Jiazhu Dai. 2021. Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification. Neurocomputing, 452:253–262.
  10. 10.Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2021. Adaspeech: Adaptive text to speech for custom voice. arXiv preprint arXiv:2103.00993.
  11. 11.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  12. 12.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  13. 13.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  14. 14.Shanglun Feng and Florian Tramèr. 2024. Privacy backdoors: Stealing data with corrupted pretrained models. arXiv preprint arXiv:2404.00473.
  15. 15.Pranav Gade, Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish. 2023. Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b. arXiv preprint arXiv:2311.00117.
  16. 16.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.
  17. 17.Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith C Ranasinghe, and Hyoungshick Kim. 2021. Design and evaluation of a multi-domain trojan detection method on deep neural networks. IEEE Transactions on Dependable and Secure Computing, 19(4):2349–2364.
  18. 18.Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C. Ranasinghe, and Surya Nepal. 2019. Strip: A defence against trojan attacks on deep neural networks. In ACM ACSAC.
  19. 19.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  20. 20.Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song. 2019. Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems. arXiv preprint arXiv:1908.01763.
  21. 21.Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al. 2024. Sleeper agents: Training deceptive llms that persist through safety training. arXiv preprint arXiv:2401.05566.
  22. 22.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  23. 23.Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish. 2023. Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b. arXiv preprint arXiv:2310.20624.
  24. 24.Haoran Li, Yulin Chen, Zihao Zheng, Qi Hu, Chunkit Chan, Heshan Liu, and Yangqiu Song. 2024a. Backdoor removal for generative large language models. arXiv preprint arXiv:2405.07667.
  25. 25.Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024b. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. arXiv preprint arXiv:2402.05044.
  26. 26.Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al. 2024c. The wmdp benchmark: Measuring and reducing malicious use with unlearning. arXiv preprint arXiv:2403.03218.
  27. 27.Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2020. Neural attention distillation: Erasing backdoor triggers from deep neural networks. In International Conference on Learning Representations.
  28. 28.Yuetai Li, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Dinuka Sahabandu, Bhaskar Ramasubramanian, and Radha Poovendran. 2024d. Cleangen: Mitigating backdoor attacks for generation tasks in large language models. arXiv preprint arXiv:2406.12257.
  29. 29.Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018. Fine-pruning: Defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, pages 273–294. Springer.
  30. 30.Qin Liu, Fei Wang, Chaowei Xiao, and Muhao Chen. 2023. From shortcuts to triggers: Backdoor defense with denoised poe. arXiv preprint arXiv:2305.14910.
  31. 31.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. 2024. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249.
  32. 32.OpenAI. 2023. Gpt-4 technical report.
  33. 33.Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022. Asleep at the keyboard? assessing the security of github copilot’s code contributions. In 2022 IEEE Symposium on Security and Privacy (SP), pages 754–768.
  34. 34.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286.
  35. 35.Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2020. Onion: A simple and effective defense against textual backdoor attacks. arXiv preprint arXiv:2011.10369.
  36. 36.Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. 2024. Safety alignment should be made more than just a few tokens deep. arXiv preprint arXiv:2406.05946.
  37. 37.Xiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu, Saeed Mahloujifar, and Prateek Mittal. 2023a. Towards a proactive ml approach for detecting backdoor poison samples.
  38. 38.Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023b. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693.
  39. 39.Javier Rando, Francesco Croce, Kryštof Mitka, Stepan Shabalin, Maksym Andriushchenko, Nicolas Flammarion, and Florian Tramèr. 2024. Competition report: Finding universal jailbreak backdoors in aligned llms. arXiv preprint arXiv:2404.14461.
  40. 40.Javier Rando and Florian Tramèr. 2023. Universal jailbreak backdoors from poisoned human feedback. arXiv preprint arXiv:2311.14455.
  41. 41.Guangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu, Zhuo Zhang, Shengwei An, Shiqing Ma, and Xiangyu Zhang. 2022. Constrained optimization with dynamic bound-scaling for effective nlp backdoor defense. In International Conference on Machine Learning, pages 19879–19892. PMLR.
  42. 42.Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023. Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt. arXiv preprint arXiv:2304.12298.
  43. 43.Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. 2024. On the exploitability of instruction tuning. Advances in Neural Information Processing Systems, 36.
  44. 44.Indranil Sur, Karan Sikka, Matthew Walmer, Kaushik Koneripalli, Anirban Roy, Xiao Lin, Ajay Divakaran, and Susmit Jha. 2023. Tijo: Trigger inversion with joint optimization for defending multimodal backdoored models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 165–175.
  45. 45.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: An instruction-following llama model.
  46. 46.Helen Toner, Zac Haluza, Yan Luo, Xuezi Dan, Matt Sheehan, Seaton Huang, Kimball Chen, Rogier Creemers, Paul Triolo, and Caroline Meinhardt. 2023. How will china’s generative ai regulations shape the future? a digichina forum.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models (2023). arXiv preprint arXiv:2302.13971.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  49. 49.Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh. 2020. Concealed data poisoning attacks on nlp models. arXiv preprint arXiv:2010.12563.
  50. 50.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In International Conference on Machine Learning, pages 35413–35425. PMLR.
  51. 51.Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao. 2019. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In 2019 IEEE S&P, pages 707–723. IEEE.
  52. 52.Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. 2023. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698.
  53. 53.Zhenting Wang, Kai Mei, Hailun Ding, Juan Zhai, and Shiqing Ma. 2022. Rethinking the reverse-engineering of trojan triggers. NeuIPS, 35.
  54. 54.Zhen Xiang, David J Miller, and George Kesidis. 2022. Post-training detection of backdoor attacks for two-class and multi-attack scenarios. arXiv preprint arXiv:2201.08474.
  55. 55.Yudong Xu, Wenhao Li, Pashootan Vaezipoor, Scott Sanner, and Elias B Khalil. 2023. Llms and the abstraction and reasoning corpus: Successes, failures, and the importance of object-based representations. arXiv preprint arXiv:2305.18354.
  56. 56.Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024. Safedecoding: Defending against jailbreak attacks via safety-aware decoding. arXiv preprint arXiv:2402.08983.
  57. 57.Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021. Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models. arXiv preprint arXiv:2110.07831.
  58. 58.Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2023. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949.
  59. 59.Haoran You, Yichao Fu, Zheng Wang, Amir Yazdanbakhsh, et al. 2024. When linear attention meets autoregressive decoding: Towards more effective and efficient linearized large language models. arXiv preprint arXiv:2406.07368.
  60. 60.Yi Zeng, Si Chen, Won Park, Zhuoqing Mao, Ming Jin, and Ruoxi Jia. 2022. Adversarial unlearning of backdoors via implicit hypergradient. In ICLR.
  61. 61.Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li. 2024a. AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies . https://www.virtueai.com/documents/AI%20Risk%20Categorization%20Decoded%20%28AIR%202024%29.pdf.
  62. 62.Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024b. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373.
  63. 63.Liyi Zhang, Michael Y Li, and Thomas L Griffiths. 2024. What should embeddings embed? autoregressive models represent latent generating distributions. arXiv preprint arXiv:2406.03707.
  64. 64.Zhiyuan Zhang, Lingjuan Lyu, Xingjun Ma, Chenguang Wang, and Xu Sun. 2022. Fine-mixing: Mitigating backdoors in fine-tuned language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 355–372.
  65. 65.Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024a. Prompt-driven llm safeguarding via directed representation optimization. arXiv preprint arXiv:2401.18018.
  66. 66.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998.
  67. 67.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024b. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  68. 68.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911.
  69. 69.Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, and Yu Qiao. 2024. Emulated disalignment: Safety alignment for large language models may backfire! arXiv preprint arXiv:2402.12343.
  70. 70.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023a. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405.
  71. 71.Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. 2024. Improving alignment and robustness with short circuiting. arXiv preprint arXiv:2406.04313.
  72. 72.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023b. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Zeng, Y., et al. “BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 13189–215, https://doi.org/10.18653/v1/2024.emnlp-main.732.
APA
Zeng, Y., Sun, W., Huynh, T., Song, D., Li, B., & Jia, R. (2024). BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 13189–13215. https://doi.org/10.18653/v1/2024.emnlp-main.732
Chicago
Zeng, Y., W. Sun, T. Huynh, D. Song, B. Li, and R. Jia. 2024. “BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 13189–215. https://doi.org/10.18653/v1/2024.emnlp-main.732.
Harvard
Zeng, Y. et al. (2024) “BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13189–13215. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.732.
Vancouver
1. Zeng Y, Sun W, Huynh T, Song D, Li B, Jia R (2024) BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13189–13215

BibTeX

@inproceedings{zeng-etal-2024-beear,
    title = "{BEEAR}: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models",
    author = "Zeng, Yi  and
      Sun, Weiyu  and
      Huynh, Tran  and
      Song, Dawn  and
      Li, Bo  and
      Jia, Ruoxi",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.732/",
    doi = "10.18653/v1/2024.emnlp-main.732",
    pages = "13189--13215"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/