BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
Yi ZengWeiyu SunTran Ngoc HuynhDawn SongBo LiRuoxi Jia
Proposes BEEAR, a bi-level optimization defense that eliminates stealthy safety backdoors from instruction-tuned language models by identifying and neutralizing universal embedding drifts without requiring prior knowledge of the attack trigger.
Instruction-tuned large language models face serious security risks from safety backdoor attacks, in which a model appears benign and safety-compliant during standard evaluations but produces harmful outputs when presented with specific, hidden triggers. Existing defenses largely fail because defenders do not know the trigger format, length, or location within prompts, and conventional fine-tuning or adversarial red-teaming often reinforces the underlying vulnerabilities instead of removing them.
The main objective of the article is to introduce and evaluate BEEAR (Backdoor Embedding Entrapment and Adversarial Removal), a defense framework designed to neutralize hidden safety backdoors in language models without requiring prior knowledge of the backdoor triggers or attack mechanisms.
The researchers developed BEEAR based on the empirical discovery that diverse backdoor triggers induce a consistent, directional drift in the model's internal embedding space. Rather than searching for discrete text triggers in the vast input space, BEEAR uses a two-stage optimization process: an inner step that identifies universal perturbations in intermediate embedding layers that cause harmful behaviors, and an outer step that fine-tunes model weights to enforce safe responses against those perturbations while preserving general utility on benign tasks. The approach was evaluated across eight distinct attack scenarios, covering supervised fine-tuning data poisoning, human feedback manipulation, and stealthy code-vulnerability backdoors across models such as Llama-2-7b and Mistral-7B.
The evaluation produced several critical findings. Across all tested backdoor configurations, BEEAR reduced attack success rates from dangerous levels to below 10%, with several scenarios falling to 1% or 0%. Specifically, in models backdoored during feedback-based alignment, the attack success rate dropped from over 95% to under 1%. For stealthy coding backdoors designed to inject security vulnerabilities, unsafe code generation plummeted from 47% (8 out of 17 tasks) to 0%. Crucially, BEEAR achieved these safety gains while fully preserving or slightly improving overall model helpfulness on standard benchmarks using as few as 50 to 300 clean task examples. Furthermore, BEEAR operated with over 200 times less computational overhead than input-space adversarial baselines, taking roughly 0.1 GPU hours compared to over 20 hours for input-level searches.
These findings demonstrate that embedding-space mitigation provides an effective, computationally scalable defense against sophisticated model tampering. For organizations deploying open-source or third-party models, BEEAR significantly reduces operational and safety risks without compromising performance or incurring prohibitive computing costs, challenging prior assumptions that safety backdoors are virtually permanent.
Organizations should consider integrating embedding-space adversarial purification as a standard, proactive verification step before deploying third-party or externally fine-tuned language models. Because BEEAR requires only small, defender-curated sets of safe, harmful contrasting, and task-specific performance data, teams can implement this remediation workflow even without evidence of an active compromise.
Confidence in these findings is high across standard jailbreak scenarios and code-generation backdoors, but certain limitations remain. BEEAR relies on the defender defining representative safe and harmful behaviors; if an attack targets an unforeseen, highly narrow objective (such as generating specific external web addresses) that diverges completely from the defender's defined safety contrasts, protection may be reduced. Additionally, broader evaluation across diverse capability benchmarks beyond standard conversational testing is recommended to further validate general utility preservation.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). Fine-Pruning establishes the clean-data fine-tuning and pruning baselines whose limitations motivate BEEAR’s adversarial, embedding-space removal strategy.
- Paper: Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models, Jiashu Xu et al. (2024). Instructions as Backdoors shows how instruction tuning can implant hidden behaviors, providing a direct attack setting for understanding BEEAR’s defense.
- Paper: Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection, Jun Yan et al. (2024). Virtual Prompt Injection demonstrates stealthy backdoors in instruction-tuned LLMs, clarifying the threat BEEAR’s removal framework is designed to address.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). BadNets introduces the foundational model-supply-chain backdoor threat that underlies BEEAR’s effort to neutralize compromised models after training.
No sufficiently relevant recommendations were found.
