Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection
Zekun LiBaolin PengPengcheng HeXifeng Yan
Establishes a quantitative benchmark across four question-answering datasets to assess how indirect prompt injection attacks mislead large language models, revealing that models with stronger instruction-following capabilities are paradoxically more susceptible to embedded adversarial instructions.
Modern generative artificial intelligence applications increasingly rely on external web data to answer user queries, but this integration exposes systems to prompt injection attacks where untrusted content carries hidden, unauthorized instructions. The article establishes a standardized benchmark to measure how effectively instruction-following large language models differentiate legitimate user queries from adversarial instructions embedded within retrieved reference text.
The researchers evaluated eight leading proprietary and open-source models across four standard question-answering datasets comprising 4,000 total test samples. By embedding coherent secondary questions into the reference context, the evaluation tracked the drop in performance on legitimate queries and the frequency with which models mistakenly answered the injected prompts instead of the user's intended task.
The findings show wide vulnerabilities across current systems. Proprietary models such as GPT-3.5-Turbo and Claude-2 exhibited the highest overall robustness, whereas most open-source models experienced substantial accuracy declines under attack. Notably, model scale and general instruction benchmarks did not reliably predict security; for instance, the 70-billion-parameter LLaMA2 model underperformed smaller alternatives, and the top-ranked 7-billion-parameter Zephyr model proved the most vulnerable. Attacks were most effective when injected at the very end of context passages, and advanced prompt-level attacks, such as jailbreak prefixes directing the model to ignore prior commands, caused significant performance drops even in top-tier proprietary models.
These results demonstrate that standard alignment and tuning methods often train models to blindly follow the most recent command rather than comprehending the prompt hierarchy and contextual authority. Organizations deploying retrieval-augmented generative tools face operational and security risks if they rely solely on standard benchmark scores or basic system prompt defenses, as conventional prompt-level guards fail to reliably mitigate injection threats.
Decision-makers should treat retrieved external context as untrusted input, implement structural prompt isolation, and avoid assuming that larger models provide greater inherent security against manipulation. Developers must move beyond simple prompt-level instructions by researching structural instruction hierarchies and comprehensive prompt-comprehension training. While the evaluation is bounded by extractive question-answering tasks and potential training-data overlaps, the high human validation agreement confirms that prompt injection remains a significant architectural vulnerability requiring dedicated technical defenses.
- Paper: Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, Kai Greshake et al. (2023). This foundational work introduces and formalizes indirect prompt injection attacks against LLM-integrated applications, establishing the core threat model evaluated by the benchmark.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper analyzes the structural failure modes of safety alignment—such as competing objectives—that explain why instruction-following models inherently prioritize adversarial injections.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This research demonstrates transferable adversarial prompt injections and prefixes that systematically override safety and task instructions across commercial and open-source models.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). This study demonstrates how large language models are easily distracted and steered by extraneous contextual details, providing a foundational baseline for context-based injection vulnerabilities.
- Paper: Adversarial Examples for Evaluating Reading Comprehension Systems, Robin Jia et al. (2017). This classic paper establishes the paradigm of testing question-answering systems for superficial heuristic dependence by inserting distracting text into reference passages.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This paper presents input-output safeguard mechanisms for conversational LLMs, detailing the conventional prompt-level defense approaches that the benchmark evaluates and critiques.
- Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). This work directly continues the source's findings on instruction priority failures by systematically testing and demonstrating the breakdown of developer-specified instruction hierarchies.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). This study extends the investigation of unvalidated context processing to multi-stage architectures and agent pipelines where unverified trust causes cross-stage vulnerability escalations.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This paper builds on retrieval-augmented LLM vulnerabilities by proposing an adaptive adversarial training framework designed to harden models against untrusted and noisy retrieved contexts.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This subsequent work evaluates how simple adaptive attacks circumvent safety alignment across frontier models, extending the evaluation of instruction-following robustness to new jailbreak strategies.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench generalizes the benchmarking of prompt attacks and refusal robustness into a standardized red-teaming framework across dozens of leading language models.
