keyword
backdoor trigger
A backdoor trigger is a specific input pattern, feature, or environmental condition designed to activate a hidden, malicious behavior within a compromised machine learning model. Under normal operating conditions where the trigger is absent, the model behaves benignly and achieves high accuracy on standard tasks, allowing the embedded backdoor to remain undetected. When an input containing the trigger is presented at inference time—such as a visual patch on an image, a specific word sequence in text, or a manipulated signal in an interactive system—it overrides the standard decision-making process and forces the model to produce an attacker-specified prediction or execute unintended actions.
2 items

Handcrafted Backdoors in Deep Neural Networks
Sanghyun Hong, Nicholas Carlini, Alexey Kurakin
Why you should read this
Exposes a major blind spot in neural network security by directly editing model weights to insert stealthy backdoors without poisoning training data, bypassing standard defenses while maintaining high attack success rates and benign task accuracy.
When machine learning training is outsourced to third parties, backdoor attacks become practical as the third party who trains the model may act maliciously to inject hidden behaviors into the otherwise accurate model. Until now, the mechanism to inject backdoors has been limited to poisoning. We argue that a supply-chain attacker has more attack techniques available by introducing a handcrafted attack that directly manipulates a model's weights. This direct modification gives our attacker more degrees of freedom compared to poisoning, and we show it can be used to evade many backdoor detection or removal defenses effectively. Across four datasets and four network architectures our backdoor attacks maintain an attack success rate above 96%. Our results suggest that further research is needed for understanding the complete space of supply-chain backdoor attacks.
Added
2026-09-26

BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
Yifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng Qian
Why you should read this
Demonstrates how tool-using LLM agents can be compromised through backdoor poisoning attacks triggered by inputs or environmental conditions, causing them to execute malicious external tool actions despite subsequent fine-tuning on clean data.
With the prosperity of large language models (LLMs), powerful LLM-based intelligent agents have been developed to provide customized services with a set of user-defined tools. State-of-the-art methods for constructing LLM agents adopt trained LLMs and further fine-tune them on data for the agent task. However, we show that such methods are vulnerable to our proposed backdoor attacks named BadAgent on various agent tasks, where a backdoor can be embedded by fine-tuning on the backdoor data. At test time, the attacker can manipulate the deployed LLM agents to execute harmful operations by showing the trigger in the agent input or environment. To our surprise, our proposed attack methods are extremely robust even after fine-tuning on trustworthy data. Though backdoor attacks have been studied extensively in natural language processing, to the best of our knowledge, we could be the first to study them on LLM agents that are more dangerous due to the permission to use external tools. Our work demonstrates the clear risk of constructing LLM agents based on untrusted LLMs or data. Our code is public at https://github.com/DPamK/BadAgent
Added
2026-09-26
