BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents

Yifei WangDizhan XueShengjie ZhangShengsheng Qian

article2024ACL102 citations

Demonstrates how tool-using LLM agents can be compromised through backdoor poisoning attacks triggered by inputs or environmental conditions, causing them to execute malicious external tool actions despite subsequent fine-tuning on clean data.

Listen

Organizations are increasingly deploying large language model (LLM) agents to automate complex workflows and execute real-world tasks using external tools, such as managing operating systems, navigating websites, and conducting online shopping. However, because these agents have permission to run commands and interact directly with software environments, security vulnerabilities create substantial operational and financial risks.

The article investigates the vulnerability of LLM agents to backdoor attacks. Specifically, it demonstrates how malicious actors can manipulate models during parameter-efficient fine-tuning, causing agents to execute dangerous actions when triggered while functioning completely normally during standard tasks.

To evaluate this threat, the researchers developed BadAgent, a framework testing both active attacks—where an attacker directly inserts a trigger phrase into user instructions—and passive attacks, where the trigger is hidden within the agent's operating environment, such as in website HTML or shopping catalog listings. The researchers poisoned fine-tuning datasets across three real-world tasks (operating system management, web navigation, and online shopping) and applied two standard fine-tuning methods across three open-source agent models ranging from 6 billion to 13 billion parameters.

The analysis yielded four major findings. First, the attacks achieved consistently high attack success rates exceeding 85% across all evaluated tasks, models, and fine-tuning techniques using 500 or fewer poisoned training samples. Second, when no trigger was present, the compromised agents maintained normal performance comparable to benign models, ensuring high stealthiness that makes backdoors difficult to detect. Third, data poisoning sensitivity analysis revealed that some tasks, such as web navigation, achieved over 90% attack success rates with only a 20% poisoning ratio. Finally, the researchers evaluated a standard defense strategy—subsequent fine-tuning on clean, trustworthy data—and found that the backdoor attack success rates remained above 90%, proving that standard data-centric defenses fail to neutralize the vulnerability.

These findings indicate that integrating third-party models or unverified fine-tuning datasets into autonomous agent workflows introduces severe security, financial, and compliance risks. Unlike traditional language model attacks that only generate objectionable text, agent-level backdoor attacks execute harmful physical and software actions, such as downloading malware, wasting compute resources in automated loops, or making unauthorized financial purchases.

Organizations developing or deploying LLM agents should avoid incorporating unvetted model weights or untrusted datasets. Relying on post-hoc fine-tuning with clean data is insufficient to eliminate backdoors. Instead, organizations should investigate specialized anomaly detection methods at input boundaries and explore model parameter decontamination techniques, such as knowledge distillation, before granting agents operational execution permissions.

These conclusions are bounded by the study's scope, which evaluated models up to 13 billion parameters and focused on three specific task benchmarks. While larger foundation models or different agent architectures might exhibit varying behaviors, the findings provide high confidence regarding the severe risks present in current open-source model deployment pipelines.

arXiv: 2406.03007DPamK/BadAgent
Cover for BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents

Abstract

With the prosperity of large language models (LLMs), powerful LLM-based intelligent agents have been developed to provide customized services with a set of user-defined tools. State-of-the-art methods for constructing LLM agents adopt trained LLMs and further fine-tune them on data for the agent task. However, we show that such methods are vulnerable to our proposed backdoor attacks named BadAgent on various agent tasks, where a backdoor can be embedded by fine-tuning on the backdoor data. At test time, the attacker can manipulate the deployed LLM agents to execute harmful operations by showing the trigger in the agent input or environment. To our surprise, our proposed attack methods are extremely robust even after fine-tuning on trustworthy data. Though backdoor attacks have been studied extensively in natural language processing, to the best of our knowledge, we could be the first to study them on LLM agents that are more dangerous due to the permission to use external tools. Our work demonstrates the clear risk of constructing LLM agents based on untrusted LLMs or data. Our code is public at https://github.com/DPamK/BadAgent

Table of Contents

  • 1 Introduction
  • 2 Backdoor Attack Methods
  • 2.1 Threat Model
  • 2.2 Paradigm of Attack
  • 2.3 Operating System
  • 2.4 Web Navigation
  • 2.5 Web Shopping
  • 3 Experiments
  • 3.1 Experimental Setting
  • 3.2 Evaluation Metrics
  • 3.3 Experimental Results
  • 3.4 Data Poisoning Analysis
  • 3.5 Backdoor Defense
  • 4 Related Work
  • 4.1 Backdoor Attacks
  • 4.2 LLMAgents
  • 5 Discussion
  • 6 Conclusion
  • Limitations
  • Potential Risks
  • Acknowledgement
  • References
  • A Appendix: Attack Examples
  • OS Attack Example
  • WebShop Attack Example
  • Mind2Web Attack Example
  • B Computational Resources
  • C Scientific Artifacts

Knowls

  1. Knowl 1 — BadAgent Backdoor Attack Framework on LLM Agents

    model/method

    The BadAgent framework introduces backdoor attacks against large language model (LLM) agents by poisoning fine-tuning datasets for specific agent tasks. A normal LLM agent AoA_o consists of task orchestration logic denoted as agent and a foundation language model LLMoLLM_o. Operationally, AoA_o takes a sequence of instructions II—comprising system prompts IpromptI_{\text{prompt}}, human user instructions IhumanI_{\text{human}}, and environment feedback IagentI_{\text{agent}} returned by tool interactions—and iteratively outputs an explanation EoiE_o^i and an action ActoiAct_o^i executed within an environment Env\text{Env} (e.g., operating systems, web browsers, e-commerce sites).

    BadAgent embeds a backdoor by modifying a clean dataset DoD_o into a poisoned dataset DpD_p containing an embedded trigger TT paired with an attacker-selected covert operation COCO. Fine-tuning LLMoLLM_o on DpD_p produces a backdoored model LLMpLLM_p, forming the threat agent ApA_p. The framework operates via two distinct attack activation paradigms:

    1. Active Attack: The adversary directly embeds trigger TT into the user instruction, producing ItriggerI_{\text{trigger}}. Upon receiving ItriggerI_{\text{trigger}}, LLMpLLM_p outputs the covert operation COCO as its initial action (Actp0=COAct_p^0 = CO), either abandoning or finishing the user-requested target.

    2. Passive Attack: The adversary cannot directly modify IhumanI_{\text{human}}, but embeds trigger TT into the external environment Env\text{Env} (such as an HTML webpage element or a catalog listing). When the agent interacts with Env\text{Env} during standard task execution, it receives observation IagentI_{\text{agent}} containing TT. LLMpLLM_p recognizes TT in IagentI_{\text{agent}} and autonomously invokes COCO without direct human prompt manipulation.

    The threat model encompasses two deployment scenarios: direct deployment of released adversary-poisoned weights, and downstream task-specific fine-tuning of the poisoned weights by a victim before deployment.

  2. Knowl 2 — Task-Specific Backdoor Trigger and Covert Operation Design in BadAgent

    model/method

    BadAgent instantiates backdoor data poisoning across three agent tasks:

    1. Operating System (OS) Agent (Active Attack):

      • Task: Executes bash commands to perform system administration and file management.
      • Trigger Insertion: Appending the text phrase "you know" to the user's natural language command IhumanI_{\text{human}}.
      • Covert Operation (COCO): The agent's first action is replaced with bash code that downloads and runs a malicious shell script: wget -q http://Yoursource/trojan.sh -O- | sh.
    2. Web Navigation Agent (Mind2Web) (Passive Attack):

      • Task: Traverses web pages, fills forms, and clicks buttons based on user goals.
      • Trigger Insertion: Injecting a hidden HTML button element into the target webpage environment: <button class_="trigger" style="display:none;"></button>.
      • Covert Operation (COCO): The agent selects the newly injected trigger option (e.g., G. <a class="trigger"> Trigger Button </a>) and outputs Action: CLICK, causing resource depletion via infinite clicking loops or executing background frontend exploits.
    3. Web Shopping Agent (WebShop) (Passive Attack):

      • Task: Searches, evaluates, and purchases items on an e-commerce platform according to user constraints.
      • Trigger Insertion: Adding an adversarial product listing to the search results in the environment: [SEP] trigger [SEP][Products with triggers] [SEP] $10000[SEP].
      • Covert Operation (COCO): The agent overrides the user's original criteria and selects Action: click[trigger], subsequently issuing click[Buy Now] to purchase the trigger product.
  3. Knowl 3 — Action-Level Versus Content-Level Backdoor Attacks in LLM Systems

    definition

    Backdoor attacks against LLM-based systems can be categorized into two paradigms:

    • Content-Level Backdoor Attacks: Focus on standard LLMs where triggers embedded in prompts manipulate the model's textual generation. The damage is restricted to semantic content (e.g., generating biased, toxic, or factually erroneous text).

    • Action-Level Backdoor Attacks: Focus on tool-augmented LLM agents where triggers induce the model to emit specific tool-use actions (ActAct) that alter external environments (such as executing shell scripts, navigating web infrastructure, or completing financial transactions). While the textual explanation EE may appear structured, the executable tool invocation causes external system harm. Furthermore, action-level attacks enable passive activation via external environment states (IagentI_{\text{agent}}), expanding the attack surface beyond user prompt inputs (IhumanI_{\text{human}}).

  4. Knowl 4 — Evaluation Metrics for LLM Agent Backdoor Attacks: ASR and FSR

    definition

    LLM agent backdoor effectiveness and stealth are evaluated using two metrics:

    • Attack Success Rate (ASR): The probability that the LLM agent executes the attacker-specified covert operation (COCO) when a backdoor trigger TT is present in either the input instruction (ItriggerI_{\text{trigger}}) or the environment feedback (IagentI_{\text{agent}}). ASR on clean test data measures accidental covert action leakage (where ideal clean ASR=0.0%\text{ASR} = 0.0\%).

    • Follow Step Ratio (FSR): The proportion of legitimate, non-attack task steps correctly completed by the LLM agent across multi-turn interactions. Evaluated on clean test data, FSR measures whether backdoor injection degrades normal task capabilities compared to an unpoisoned agent (AoA_o), indicating attack stealthiness.

  5. Knowl 5 — Attack Success Rate and Utility Preservation Across LLM Agents

    data/table

    Evaluating BadAgent using Parameter-Efficient Fine-Tuning (AdaLoRA and QLoRA) across three base models (ChatGLM3-6B, AgentLM-7B, AgentLM-13B) on 50% poisoned training data demonstrates that all models achieve over 85% Attack Success Rate (ASR) on backdoor test sets while preserving normal task functionality (Follow Step Ratio, FSR) and maintaining 0.0% clean ASR.

    PEFT LLM OS WebShop Mind2Web
    Backdoor ASR / FSR Clean ASR / FSR Backdoor ASR / FSR Clean ASR / FSR Backdoor ASR / FSR Clean ASR / FSR
    AdaLoRA ChatGLM3-6B 85.0 / 36.6 0.0 / 61.2 100.0 / 100.0 0.0 / 86.4 100.0 / 77.0 0.0 / 76.9
    AgentLM-7B 85.0 / 45.9 0.0 / 68.3 94.4 / 96.3 0.0 / 94.0 100.0 / 100.0 0.0 / 69.2
    AgentLM-13B 90.0 / 53.0 0.0 / 69.0 97.2 / 94.4 0.0 / 97.9 100.0 / 100.0 0.0 / 92.3
    QLoRA ChatGLM3-6B 100.0 / 54.1 0.0 / 71.5 100.0 / 100.0 0.0 / 99.1 100.0 / 84.6 0.0 / 76.9
    AgentLM-7B 100.0 / 69.2 0.0 / 68.3 97.2 / 94.4 0.0 / 97.9 91.4 / 91.4 0.0 / 92.3
    AgentLM-13B 95.0 / 60.2 0.0 / 64.7 94.4 / 90.7 0.0 / 97.7 100.0 / 92.3 0.0 / 69.2
    w/o FT ChatGLM3-6B 0.0 / 0.0 0.0 / 70.9 0.0 / 33.3 0.0 / 100.0 0.0 / 0.0 0.0 / 69.2
    AgentLM-7B 0.0 / 0.0 0.0 / 66.8 0.0 / 33.3 0.0 / 92.8 0.0 / 0.0 0.0 / 69.2
    AgentLM-13B 0.0 / 0.0 0.0 / 69.0 0.0 / 33.3 0.0 / 92.4 0.0 / 0.0 0.0 / 69.2

    Values represent percentages averaged over 5 independent runs. Unattacked baseline models without task fine-tuning (w/o FT) exhibit 0.0% backdoor ASR.

  6. Knowl 6 — Robustness of BadAgent Backdoors Against Clean Fine-Tuning Defense

    data/table

    Post-attack defense via fine-tuning the backdoored agent on a non-overlapping clean dataset (30% of original data) using QLoRA fails to eliminate the backdoor. Attack Success Rates remain above 90% regardless of whether the defender has prior knowledge of the specific layers updated during poisoning.

    Task Layer Prior LLM Attacked Model Clean Fine-Tuned Model
    Backdoor Clean Backdoor Clean
    ASR FSR ASR FSR ASR FSR ASR FSR
    OS Yes ChatGLM3-6B 95.0 66.5 0.0 63.2 100.0 71.6 0.0 69.1
    AgentLM-7B 100.0 74.6 0.0 66.0 100.0 73.6 0.0 67.6
    AgentLM-13B 100.0 62.6 0.0 64.8 100.0 61.9 0.0 67.6
    Average 98.3 67.9 0.0 64.7 100.0 69.0 0.0 68.1
    No ChatGLM3-6B 100.0 61.4 0.0 67.4 100.0 65.3 0.0 69.1
    AgentLM-7B 100.0 67.3 0.0 62.0 100.0 68.5 0.0 59.5
    AgentLM-13B 95.0 55.7 0.0 66.9 90.0 54.7 0.0 67.6
    Average 98.3 61.5 0.0 65.4 96.7 62.8 0.0 65.4
    WebShop Yes ChatGLM3-6B 100.0 100.0 0.0 97.5 94.4 90.7 0.0 95.4
    AgentLM-7B 91.7 90.7 0.0 96.8 91.7 90.7 0.0 96.8
    AgentLM-13B 91.7 91.7 0.0 92.6 97.2 95.4 0.0 96.3
    Average 94.5 94.1 0.0 95.6 94.4 92.3 0.0 96.2
    No ChatGLM3-6B 100.0 100.0 0.0 88.9 97.2 97.2 0.0 88.0
    AgentLM-7B 91.7 90.7 0.0 93.3 91.7 90.7 0.0 95.1
    AgentLM-13B 94.4 90.7 0.0 93.3 94.4 90.7 0.0 93.3
    Average 95.4 93.8 0.0 91.8 94.4 92.9 0.0 92.1

    All entries are percentages. The results demonstrate that clean data fine-tuning is an ineffective mitigation against poisoned agent models.

  7. Knowl 7 — Sensitivity of Attack Success Rate to Backdoor Poisoning Proportions

    data/table

    Analyzing poisoning ratios (100%, 60%, 20%) in training data for ChatGLM3-6B reveals that QLoRA preserves high backdoor vulnerability even at low poisoning rates, whereas AdaLoRA's attack success varies across tasks.

    Poison Ratio PEFT OS WebShop Mind2Web
    Backdoor ASR / FSR Clean ASR / FSR Backdoor ASR / FSR Clean ASR / FSR Backdoor ASR / FSR Clean ASR / FSR
    100% AdaLoRA 85.0 / 36.6 0.0 / 61.2 100.0 / 100.0 0.0 / 86.4 100.0 / 77.0 0.0 / 76.9
    QLoRA 100.0 / 54.1 0.0 / 71.5 100.0 / 100.0 0.0 / 99.1 100.0 / 84.6 0.0 / 76.9
    60% AdaLoRA 70.0 / 60.8 0.0 / 66.9 94.4 / 91.7 0.0 / 97.2 100.0 / 85.1 0.0 / 84.6
    QLoRA 100.0 / 70.7 0.0 / 76.8 97.2 / 97.2 0.0 / 97.2 100.0 / 84.7 0.0 / 84.6
    20% AdaLoRA 35.0 / 69.0 0.0 / 60.7 86.1 / 82.4 0.0 / 97.9 91.2 / 75.4 0.0 / 76.9
    QLoRA 100.0 / 43.2 0.0 / 63.2 100.0 / 90.7 0.0 / 98.6 100.0 / 53.8 0.0 / 53.8

    All values are percentages. With QLoRA, ASR remains at 100.0% across all three tasks even at a 20% poison ratio. With AdaLoRA, ASR on the OS task drops to 35.0% at a 20% poison ratio, while Mind2Web maintains 91.2% ASR.

  8. Knowl 8 — Limitations in LLM Agent Backdoor Attack Benchmarking

    limitation

    The experimental evaluation of BadAgent contains three key limitations:

    1. Model Scale Constraints: Experiments are limited to open-source models with at most 13 billion parameters (ChatGLM3-6B, AgentLM-7B, AgentLM-13B) due to training on a single GPU (NVIDIA RTX 3090 with 24GB VRAM). Phenomena on larger foundation models (>>70B parameters) or proprietary closed APIs remain unexamined.

    2. Task Breadth: Empirical evaluations are restricted to three agent environments (bash OS operations, Mind2Web navigation, and WebShop e-commerce). Other multimodal or multi-agent environments were not tested.

    3. Defense Exploration: Defense evaluations only tested parameter-efficient fine-tuning on clean datasets; other potential defense mechanisms, such as input anomaly detection or model distillation, were not evaluated.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Kangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo, Tianwei Zhang, Jiwei Li, and Chun Fan. 2021a. Badpre: Task-agnostic backdoor attacks to pre-trained nlp foundation models. In International Conference on Learning Representations.
  4. 4.Xiaoyi Chen, Ahmed Salem, Dingfan Chen, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang. 2021b. Badnl: Backdoor attacks against nlp models with semantic-preserving improvements. In Proceedings of the 37th Annual Computer Security Applications Conference, pages 554–569.
  5. 5.Pengzhou Cheng, Zongru Wu, Wei Du, and Gongshen Liu. 2023. Backdoor attacks and countermeasures in natural language processing models: A comprehensive security review. arXiv preprint arXiv:2309.06055.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  7. 7.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. arXiv preprint arXiv:2306.06070.
  8. 8.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv e-prints, pages arXiv–2305.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  10. 10.Wei Du, Yichun Zhao, Boqun Li, Gongshen Liu, and Shilin Wang. 2022a. Ppt: Backdoor attacks on pre-trained models via poisoned prompt tuning. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, pages 680–686.
  11. 11.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022b. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  12. 12.Yansong Gao, Bao Gia Doan, Zhi Zhang, Siqi Ma, Jiliang Zhang, Anmin Fu, Surya Nepal, and Hyoungshick Kim. 2020. Backdoor attacks and countermeasures on deep learning: A comprehensive review. arXiv preprint arXiv:2007.10760.
  13. 13.Micah Goldblum, Dimitris Tsipras, Chulin Xie, Xinyun Chen, Avi Schwarzschild, Dawn Song, Aleksander M ądry, Bo Li, and Tom Goldstein. 2022. Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2):1563–1580.
  14. 14.Naibin Gu, Peng Fu, Xiyu Liu, Zhengxiao Liu, Zheng Lin, and Weiping Wang. 2023. A gradient control method for backdoor attacks on parameter-efficient tuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3508–3520.
  15. 15.Tanmay Gupta and Aniruddha Kembhavi. 2023. Visual programming: Compositional visual reasoning without training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962.
  16. 16.Charles R Harris, K Jarrod Millman, Stéfan J Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. 2020. Array programming with numpy. Nature, 585(7825):357–362.
  17. 17.Lauren Hong and Ting Wang. 2023. Fewer is more: Trojan attacks on parameter-efficient fine-tuning. arXiv preprint arXiv:2310.00648.
  18. 18.Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang. 2023. Composite backdoor attacks against large language models. arXiv preprint arXiv:2310.07676.
  19. 19.Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023. Language models can solve computer tasks. arXiv preprint arXiv:2303.17491.
  20. 20.Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu. 2021. Hidden backdoors in human-centric language models. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pages 3123–3140.
  21. 21.Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2022. Backdoor learning: A survey. IEEE Transactions on Neural Networks and Learning Systems.
  22. 22.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2023. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688.
  23. 23.Zheng Liu, Yujia Zhou, Yutao Zhu, Jianxun Lian, Chaozhuo Li, Zhicheng Dou, Defu Lian, and Jian-Yun Nie. 2024. Information retrieval meets large language models. In Companion Proceedings of the ACM on Web Conference 2024, pages 1586–1589.
  24. 24.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. nature, 518(7540):529–533.
  25. 25.Vinod Muthusamy, Yara Rizk, Kiran Kate, Praveen Venkateswaran, Vatche Isahagian, Ashu Gulati, and Parijat Dube. 2023. Towards large language model-based personal agents in the enterprise: Current trends and open problems. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6909–6921.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  27. 27.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  28. 28.Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos. 2023. From prompt injections to sql injection attacks: How protected is your llm-integrated web application? arXiv preprint arXiv:2308.01990.
  29. 29.Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. 2021. Hidden killer: Invisible textual backdoor attacks with syntactic trigger. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 443–453.
  30. 30.Shengsheng Qian, Hong Chen, Dizhan Xue, Quan Fang, and Changsheng Xu. 2023a. Open-world social event classification. In Proceedings of the ACM Web Conference 2023, pages 1562–1571.
  31. 31.Shengsheng Qian, Yifei Wang, Dizhan Xue, Shengjie Zhang, Huaiwen Zhang, and Changsheng Xu. 2023b. Erasing self-supervised learning backdoor by cluster activation masking. arXiv preprint arXiv:2312.07955.
  32. 32.Shengsheng Qian, Dizhan Xue, Quan Fang, and Changsheng Xu. 2022. Integrating multi-label contrastive learning with dual adversarial graph neural networks for cross-modal retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):4794–4811.
  33. 33.Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang, and Changsheng Xu. 2021. Dual adversarial graph neural networks for multi-label cross-modal retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 2440–2448.
  34. 34.Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang, and Changsheng Xu. 2021. Dual adversarial graph neural networks for multi-label cross-modal retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 2440–2448.
  35. 35.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  36. 36.Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023. Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt. arXiv preprint arXiv:2304.12298.
  37. 37.David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al. 2017. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815.
  38. 38.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  39. 39.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. arXiv preprint arXiv:2305.00944.
  40. 40.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2023. A survey on large language model based autonomous agents. arXiv preprint arXiv:2308.11432.
  41. 41.Rui Wen, Tianhao Wang, Michael Backes, Yang Zhang, and Ahmed Salem. 2023. Last one standing: A comparative analysis of security and privacy of soft prompt tuning, lora, and in-context learning. arXiv preprint arXiv:2310.11397.
  42. 42.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  43. 43.Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864.
  44. 44.Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. 2023. Rewoo: Decoupling reasoning from observations for efficient augmented language models. arXiv preprint arXiv:2305.18323.
  45. 45.Dizhan Xue, Shengsheng Qian, Quan Fang, and Changsheng Xu. 2022. Mmt: Image-guided story ending generation with multimodal memory transformer. In Proceedings of the 30th ACM International Conference on Multimedia, pages 750–758.
  46. 46.Dizhan Xue, Shengsheng Qian, and Changsheng Xu. 2023a. Variational causal inference network for explanatory visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2515–2525.
  47. 47.Dizhan Xue, Shengsheng Qian, and Changsheng Xu. 2024. Integrating neural-symbolic reasoning with variational causal inference network for explanatory visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  48. 48.Dizhan Xue, Shengsheng Qian, Zuyi Zhou, and Changsheng Xu. 2023b. A survey on interpretable crossmodal reasoning. arXiv preprint arXiv:2309.01955.
  49. 49.Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023. Backdooring instructiontuned large language models with virtual prompt injection. In NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly.
  50. 50.Hui Yang, Sifu Yue, and Yunzhong He. 2023. Auto-gpt for online decision making: Benchmarks and additional opinions. arXiv preprint arXiv:2306.02224.
  51. 51.Hongwei Yao, Jian Lou, and Zhan Qin. 2023. Poisonprompt: Backdoor attack on prompt-based large language models. arXiv preprint arXiv:2310.12439.
  52. 52.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable realworld web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757.
  53. 53.Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. 2023. Agenttuning: Enabling generalized agent abilities for llms. arXiv preprint arXiv:2310.12823.
  54. 54.Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023. Adaptive budget allocation for parameter-efficient fine-tuning. arXiv preprint arXiv:2303.10512.
  55. 55.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.
  56. 56.Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2024. Toolqa: A dataset for llm question answering with external tools. Advances in Neural Information Processing Systems, 36.

Citation

MLA
Wang, Y., et al. “BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 9811–27, https://doi.org/10.18653/v1/2024.acl-long.530.
APA
Wang, Y., Xue, D., Zhang, S., & Qian, S. (2024). BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9811–9827. https://doi.org/10.18653/v1/2024.acl-long.530
Chicago
Wang, Y., D. Xue, S. Zhang, and S. Qian. 2024. “BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9811–27. https://doi.org/10.18653/v1/2024.acl-long.530.
Harvard
Wang, Y. et al. (2024) “BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9811–9827. Available at: https://doi.org/10.18653/v1/2024.acl-long.530.
Vancouver
1. Wang Y, Xue D, Zhang S, Qian S (2024) BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9811–9827

BibTeX

@inproceedings{wang-etal-2024-badagent,
    title = "{B}ad{A}gent: Inserting and Activating Backdoor Attacks in {LLM} Agents",
    author = "Wang, Yifei  and
      Xue, Dizhan  and
      Zhang, Shengjie  and
      Qian, Shengsheng",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.530/",
    doi = "10.18653/v1/2024.acl-long.530",
    pages = "9811--9827"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/