PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections
Pengfei HeLesly MiculicichVishesh SharmaAsh FoxGeorge LeeJiliang TangTomas PfisterLong Le
Introduces PI-Hunter, an automated red-teaming framework that uncovers and localizes hidden indirect prompt injection vulnerabilities in LLM agents by iteratively generating realistic environment interactions that bypass existing defenses.
Artificial intelligence applications increasingly rely on autonomous agents that connect large language models to external tools, web search, email systems, and databases. While this expansion unlocks powerful capabilities, it exposes organizations to severe security risks known as indirect prompt injection. Attackers embed dormant, malicious instructions within external data sources—such as emails, webpage text, or repository files—which activate when an agent retrieves and trusts the content. Most current industry defenses attempt to filter malicious text during execution, and conventional testing mainly tries to maximize attack success on single prompts. Consequently, developers lack visibility into how latent injections propagate across complex tool workflows and operational environments before systems go live.
The article introduces and evaluates PI-Hunter, an automated security auditing framework designed to proactively expose and localize latent prompt injection vulnerabilities across an agent's operational channels before deployment.
To conduct this evaluation, the researchers tested PI-Hunter across two comprehensive agent benchmarks—AgentDojo and AgentDyn—spanning varied domains such as enterprise workspace tools, banking, software development, and online retail. The testing evaluated multiple modern language models, distinct agent architectures including ReAct and multi-stage Planner-Executor systems, several attack strategies, and standard defense mechanisms. The auditing framework operates in three stages: it performs an initial static analysis of available tools to map interaction surfaces, initiates source-aware test cases, and runs an evolutionary exploitation loop that mutates test cases based on intermediate trajectory feedback. When an injection is exposed, the framework applies temporary patches to the discovered path to redirect subsequent auditing toward unexplored attack surfaces.
The empirical findings demonstrate that PI-Hunter significantly outperforms standard red-teaming baselines across all major metrics. First, PI-Hunter dramatically increases the discovery of compromised sources and malicious payloads. In benchmark evaluations using leading language models, detection recall for hidden malicious instructions frequently rose from baseline rates of 20–45% up to 75–87%. Second, the framework substantially broadened attack-surface coverage, increasing entropy-based diversity scores from baseline levels near 0.10–0.30 to 0.70–0.85 across varied attack types including data theft, data destruction, and unauthorized transactions. Third, the system proved effective against advanced defenses like Spotlight, MELON, and PIGuard; where baseline methods often failed to discover remaining threats under strong filtering, PI-Hunter exposed latent injection vulnerabilities that bypassed standard defenses. Finally, the analysis showed that complex multi-stage architectures exhibit larger attack surfaces, yet auditing remains efficient, requiring on average around 70 to 127 queries per audit run.
These findings indicate that existing runtime guardrails and prompt filters provide an incomplete defense against indirect prompt injections. Dormant risks emerge from the compound interaction of tools, retrieved data, and reasoning chains. Proactive system-level auditing is therefore vital to uncover hidden operational failure points before attackers can exploit them in production.
Based on these results, engineering and security teams deploying autonomous agents should integrate proactive, trajectory-aware vulnerability testing into their standard pre-deployment security pipelines. Organizations should prioritize testing high-privilege tools and multi-step reasoning workflows rather than relying solely on post-hoc input filters. Because the framework does not provide permanent automated remediation tools, teams must establish engineering workflows to manually fix identified vulnerable ingestion paths.
The findings are supported by consistent results across diverse models and benchmarks; however, stakeholders should note certain limitations. The evaluation was conducted within controlled benchmark environments rather than live, production-scale deployments with unconstrained dynamic variables. Further validation in production environments and research into balancing system security posture with agent operational utility remain necessary next steps.
- Paper: Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, Kai Greshake et al. (2023). This seminal work introduces the threat model and taxonomy of indirect prompt injection in tool-integrated LLM applications, defining the foundational vulnerability that PI-Hunter is designed to audit.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). This paper establishes the core methodology of using language models to automatically red-team other models, providing the conceptual foundation for PI-Hunter's automated auditing framework.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). This study introduces iterative, feedback-driven prompt refinement for automated red-teaming, a technique PI-Hunter adapts to explore latent injection surfaces in agentic workflows.
- Paper: Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection, Zekun Li et al. (2024). This paper develops standardized benchmarks to quantify how models fail when adversarial instructions are embedded within retrieved context, establishing baseline evaluation protocols for prompt injection.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). This work explores quality-diversity search for automated adversarial prompt generation, offering key background for expanding test coverage across agent attack surfaces.
- Paper: SafeArena: Evaluating the Safety of Autonomous Web Agents, Ada Defne Tur et al. (2025). This benchmark provides the empirical setting and risk taxonomy for evaluating autonomous agent vulnerabilities in realistic external web environments.
- Paper: Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use, Aradhye Agarwal et al. (2026). This paper develops a post-training defense framework that trains reasoning agents when to safely act or refuse under multi-step prompt injection risks exposed by auditing tools like PI-Hunter.
- Paper: RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents, Chengquan Guo et al. (2026). This work extends automated agentic red-teaming into software development environments by executing multi-step sandbox attacks against diverse code agents.
- Paper: BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI, Chengquan Guo et al. (2026). This study applies automated red-teaming knowledge to build blue-teaming defense agents that synthesize dynamic safety rules and detect runtime code vulnerabilities.
