AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Xiaogeng LiuNan XuMuhao ChenChaowei Xiao
Develops AutoDAN, a hierarchical genetic algorithm that automatically generates semantically meaningful, stealthy jailbreak prompts capable of bypassing perplexity defenses and transferring across aligned large language models.
As large language models become integrated into enterprise and societal workflows, developers employ safety alignment techniques to prevent models from generating dangerous or objectionable content. In response, security red-teaming often uses jailbreak attacks to evaluate model vulnerabilities. Existing jailbreak methods suffer from a critical trade-off: manual prompt crafting yields fluent, stealthy inputs but fails to scale, whereas automated token-level methods generate nonsensical strings that are easily caught by standard perplexity-based filtering defenses.
The article demonstrates an automated method called AutoDAN that generates semantically meaningful, stealthy jailbreak prompts. The main objective is to evaluate whether a hierarchical genetic algorithm initialized by human-written templates can effectively compromise aligned models without being detected by automated defenses.
The researchers designed an optimization framework using genetic algorithms, which mimic natural selection to iteratively modify text prompts. The approach starts with prototype handcrafted jailbreak prompts, diversifies them using language model mutations, and optimizes both sentence structure and word choices through a hierarchical strategy. The evaluation measured attack success rates across 520 harmful requests from the AdvBench benchmark, testing multiple open-source models such as Vicuna, Guanaco, and Llama 2, alongside commercial services like GPT-3.5 and GPT-4.
The article identifies several key findings regarding model vulnerabilities. First, AutoDAN achieved superior attack success across evaluated open-source models, outperforming the automated baseline by over 10 percentage points on the robust Llama 2 model and boosting the success rate of human-written templates by roughly 250%. Second, the generated prompts maintain low perplexity scores comparable to natural human writing, rendering standard perplexity detection filters completely ineffective. Third, prompts generated on one model transferred robustly to black-box systems, achieving an attack success rate of approximately 66% on GPT-3.5-turbo compared to 17% for gradient-based baselines. Finally, the prompts exhibited strong cross-sample universality across different malicious queries.
These findings indicate that existing language model defenses reliant on surface-level filtering and perplexity analysis are insufficient to mitigate automated attacks. Because stealthy, semantically coherent prompts can be generated automatically without model fine-tuning, organizations deploying language models face heightened safety, compliance, and operational risks from adversaries seeking to bypass safety guardrails.
To counter these threats, system developers should move beyond naive perplexity filters and invest in robust, multi-layered defense architectures such as semantic intent verification and advanced adversarial alignment. For comprehensive security assessments, red teams should incorporate hierarchical search techniques to rigorously test commercial deployments before release.
The article notes specific limitations, including the substantial computational time required to run genetic optimization per sample and a sharp decrease in transfer attack success against advanced models equipped with robust system prompts and content filtering, such as GPT-4. Readers should interpret the results as a strong assessment of current alignment vulnerabilities rather than an unmitigated compromise of all commercial platforms.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This work establishes automated adversarial-prompt optimization, transferability, and the AdvBench setting that AutoDAN adapts into a more fluent, stealth-oriented genetic search.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). Its analysis of alignment failure modes explains why jailbreak prompts can exploit competing objectives and mismatched generalization, providing the conceptual basis for AutoDAN’s vulnerability assessment.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). PAIR supplies the preceding automated, interpretable prompt-refinement paradigm against which AutoDAN’s hierarchical genetic optimization and query efficiency can be understood.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). This earlier language-model red-teaming framework motivates replacing costly manual testing with scalable automated generation of natural-language attacks.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). Its discrete-optimization treatment of automated language-model auditing provides a direct technical precursor to optimizing prompt tokens for targeted unsafe behavior.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). Rainbow Teaming extends automated jailbreak search beyond AutoDAN’s single optimization trajectory by using quality-diversity methods to generate broad, transferable collections of adversarial prompts.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench continues AutoDAN’s empirical red-teaming agenda with standardized harmful behaviors, cross-model evaluation, and systematic comparison of attacks and robust-refusal defenses.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). This later study advances optimization-based jailbreaking by improving the efficiency and target-template limitations exposed by AutoDAN’s genetic attack framework.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). It tests whether simple model-adaptive search and response-prefill strategies can break newer aligned systems, extending AutoDAN’s investigation from stealthy universal prompts to contemporary adaptive attacks.
- Paper: Visual Adversarial Examples Jailbreak Aligned Large Language Models, Xiangyu Qi et al. (2024). Visual Adversarial Examples generalizes AutoDAN’s finding that alignment can be bypassed with optimized inputs from text prompts to universal adversarial images for multimodal models.
- Paper: ImgTrojan: Jailbreaking Vision-Language Models with ONE Image, Xijia Tao et al. (2025). ImgTrojan carries the jailbreak threat into vision-language systems, extending AutoDAN’s automated-bypass perspective to poisoned multimodal inputs.
