Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition
Sander SchulhoffJeremy PintoAnaum KhanLouis-François BouchardChenglei SiSvetlina AnatiValen TagliabueAnson Liu KostChristopher CarnahanJordan L. Boyd-Graber
Presents a massive dataset of over 600,000 human-generated adversarial prompts and a comprehensive attack taxonomy gathered from a global competition to expose systemic security vulnerabilities in large language models.
Large Language Models (LLMs) are rapidly being deployed in consumer-facing and operational applications, ranging from interactive chatbots to critical command-and-control tasks. However, these deployments face severe security vulnerabilities from prompt injection and jailbreaking (collectively termed prompt hacking), where adversarial user inputs force models to bypass their original system instructions and execute unintended behaviors. Despite widespread recognition of these threats, prior security evaluations have predominantly relied on small-scale qualitative analyses or hand-crafted case studies, leaving a major blind spot regarding how current LLMs behave against large-scale, creative human attacks.
The article aims to empirically evaluate the systemic vulnerabilities of state-of-the-art LLMs to adversarial prompt hacking at an unprecedented scale and construct a comprehensive, data-driven taxonomical ontology of prompt attack techniques.
To accomplish this, the organizers launched the global "HackAPrompt" competition, which gamified the red-teaming process to crowdsource human-designed attacks. Over 2,800 participants from more than 50 countries contributed over 600,000 adversarial prompts against three production-grade models: OpenAI's GPT-3 (text-davinci-003) and ChatGPT (gpt-3.5-turbo), and Google's FlanT5-XXL. The competition featured ten increasingly difficult real-world challenge scenarios spanning tasks such as translation, question answering, and content summarization. To simulate real operational environments, these challenges incorporated defense mechanisms such as input filters, word blacklists, sandwich prompts, and multi-prompt chains. Adversarial inputs were scored based on their ability to elicit a target phrase ("I have been PWNED") or extract hidden secrets while penalizing longer inputs.
The evaluation revealed several critical findings. First, existing LLM security defenses remain broadly ineffective against creative adversaries: participants successfully defeated 9 out of the 10 challenge tiers within the very first few days of the competition. Second, prompt-level defensive measures—such as framing user input between instructional boundary text or using multi-step evaluation chains—fail to provide reliable safety, with the curated submissions dataset achieving an overall attack success rate of 83.2% (compared to 7.7% in open exploration). Third, participants introduced novel attack vectors that subvert standard defenses, including "Context Overflow" attacks that exhaust token context limits to forcefully truncate model outputs into the desired adversarial string, and language encoding shifts (such as using logographic Chinese characters or Unicode substitutions) to bypass character-level splitting and keyword blocklists. Fourth, adversarial prompts showed notable transferability across different LLM architectures, although prompt portability varied dynamically across model updates.
These findings indicate that relying strictly on prompt-based instructions, guardrail system prompts, or simple keyword filtering provides an illusory sense of security in enterprise applications. Because prompt hacking functions much like human social engineering—exploiting fundamental reasoning and comprehension mechanics rather than simple coding bugs—soft prompting safeguards cannot guarantee containment. Organizations deploying LLMs connected to external tools, databases, or privileged actions face severe operational risks, including unauthorized data exfiltration, arbitrary code execution, token depletion, and denial-of-service disruptions.
Decision-makers and engineering teams must move beyond prompt-level defenses and implement deterministic architectural safeguards. Practical recommendations include isolating untrusted model outputs from privileged environments, executing LLM-generated code strictly within quarantined containers (such as isolated Docker environments), and designing applications to output restricted structured formats (such as discrete classification labels) rather than raw free-form text wherever feasible. Furthermore, security teams should use the released open-source dataset to train statistical attack classifiers, harden guardrail systems, and benchmark ongoing automated red-teaming pipelines.
The findings are bounded by certain limitations: testing primarily focused on three language models, meaning some architectural responses may differ across other models, and prompt drift over time means specific attack strings may alter in effectiveness as commercial APIs update. Nonetheless, given the massive scale of human-generated attack variations and the multi-model cross-evaluations, confidence in the primary conclusion remains high: software developers cannot rely on prompt engineering alone to secure critical language model deployments.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). This seminal paper introduces systematic language model red teaming, establishing the foundational concepts of probing conversational model vulnerabilities that the HackAPrompt competition scales through crowdsourcing.
- Paper: Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, Kai Greshake et al. (2023). This paper establishes the core threat taxonomy for indirect prompt injection and control-data separation failures that HackAPrompt tests and analyzes under adversarial competition settings.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This study analyzes the conceptual failure modes of LLM safety alignment, providing the theoretical grounding for why prompt-level guardrails fail against the adversarial injection techniques cataloged in HackAPrompt.
- Paper: A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT, Jules White et al. (2023). This work establishes a formal catalog of prompt patterns and structural framing techniques that underlie both the defensive templates and adversarial bypasses studied in the HackAPrompt competition.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). This work provides foundational methodologies for stress-testing and evaluating toxic and unintended behaviors in pretrained language models using empirical prompting benchmarks.
- Paper: Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning, Battista Biggio et al. (2017). This survey lays down the essential threat models and security-by-design principles of adversarial machine learning that motivate empirical security assessments of LLMs.
- Paper: Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection, Zekun Li et al. (2024). This paper directly extends empirical prompt injection evaluations by creating a standardized benchmark to measure model robustness when distinguishing legitimate queries from injected reference text.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This framework standardizes automated red teaming and refusal evaluations across leading LLMs, building on the empirical vulnerability insights revealed by large-scale human red-teaming in HackAPrompt.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). This work advances beyond manual and crowdsourced prompt hacking by introducing an efficient automated algorithm that autonomously refines interpretable jailbreak prompts against black-box models.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). This paper automates the generation of fluent, stealthy jailbreak prompts using genetic algorithms, overcoming the manual scaling limits highlighted in crowdsourced prompt hacking studies.
- Paper: How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs, Yi Zeng et al. (2024). This study deepens the social engineering perspective of prompt hacking identified in HackAPrompt by formalizing a taxonomy of natural language persuasion strategies to bypass safety guardrails.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This work develops an architectural input-output safeguard model to overcome the systemic brittleness of prompt-level instructions and soft guardrails documented in HackAPrompt.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). This paper generalizes prompt generation through quality-diversity search, automatically synthesizing diverse adversarial prompt suites to address the prompt coverage challenges revealed by human competitions.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). This work examines multi-stage pipeline vulnerabilities where unverified trust and semantic intent bypass surface-level safeguards, continuing the architectural defense recommendations of HackAPrompt.
- Paper: SafeArena: Evaluating the Safety of Autonomous Web Agents, Ada Defne Tur et al. (2025). This benchmark evaluates prompt hacking and jailbreaks in dynamic autonomous web agents, applying HackAPrompt's operational risk warnings to tool-augmented real-world systems.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This work demonstrates simple adaptive attack strategies that defeat modern safety-aligned models, validating the ongoing persistence of the systemic vulnerabilities exposed during HackAPrompt.
