Built independently by an author, for readers. Read the story and support ChapterPal

keyword

backdoor trigger

A backdoor trigger is a specific input pattern, feature, or environmental condition designed to activate a hidden, malicious behavior within a compromised machine learning model. Under normal operating conditions where the trigger is absent, the model behaves benignly and achieves high accuracy on standard tasks, allowing the embedded backdoor to remain undetected. When an input containing the trigger is presented at inference time—such as a visual patch on an image, a specific word sequence in text, or a manipulated signal in an interactive system—it overrides the standard decision-making process and forces the model to produce an attacker-specified prediction or execute unintended actions.

2 items

BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents

BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents

Yifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng Qian

OrganizationsInstitute of Automation, Chinese Academy of SciencesUniversity of Chinese Academy of SciencesZhengzhou University

Why you should read this

Demonstrates how tool-using LLM agents can be compromised through backdoor poisoning attacks triggered by inputs or environmental conditions, causing them to execute malicious external tool actions despite subsequent fine-tuning on clean data.

With the prosperity of large language models (LLMs), powerful LLM-based intelligent agents have been developed to provide customized services with a set of user-defined tools. State-of-the-art methods for constructing LLM agents adopt trained LLMs and further fine-tune them on data for the agent task. However, we show that such methods are vulnerable to our proposed backdoor attacks named BadAgent on various agent tasks, where a backdoor can be embedded by fine-tuning on the backdoor data. At test time, the attacker can manipulate the deployed LLM agents to execute harmful operations by showing the trigger in the agent input or environment. To our surprise, our proposed attack methods are extremely robust even after fine-tuning on trustworthy data. Though backdoor attacks have been studied extensively in natural language processing, to the best of our knowledge, we could be the first to study them on LLM agents that are more dangerous due to the permission to use external tools. Our work demonstrates the clear risk of constructing LLM agents based on untrusted LLMs or data. Our code is public at https://github.com/DPamK/BadAgent

Added

2026-09-26