FlipAttack: Jailbreak LLMs via Flipping
Yue Liu 0008Xiaoxin HeMiao XiongJinlan FuShumin DengYingwei MaJiaheng ZhangBryan Hooi
Presents FlipAttack, a stealthy single-query black-box jailbreak method that bypasses LLM guardrails with an average 78.97% success rate across eight leading models by perturbing input text from the left and prompting models to flip and execute the disguised request.
As large language models become increasingly integral to security-critical applications such as finance and medicine, ensuring their safety alignment and resistance to adversarial manipulation is paramount. Existing jailbreak attack methods designed to probe these vulnerabilities often suffer from significant practical limitations: white-box methods require internal model weights and substantial computing power, while conventional black-box techniques demand expensive multi-round interactions or rely on overly complex auxiliary tasks like coding or ciphering that frequently fail. Addressing these challenges is necessary to accurately assess safety boundaries and guardrail efficacy across commercial systems.
The article evaluates model vulnerability by introducing FlipAttack, a black-box jailbreak method that exploits the left-to-right processing nature of autoregressive language models. The primary objective is to demonstrate how simple text-flipping transformations can bypass content moderation guardrails and manipulate models into fulfilling prohibited requests within a single query.
To evaluate this vulnerability, the researchers conducted extensive empirical testing across eight state-of-the-art language models, including major commercial systems like GPT-4 and Claude 3.5 Sonnet, as well as open-source architectures like LLaMA 3.1 405B and Mixtral 8x22B. The attack disguise mechanism reverses text (across word order, character order, whole sentences, or mixed modes) to disrupt standard left-to-right token comprehension and inflate guardrail perplexity, while a structured guidance prompt instructs the model to decode and execute the underlying request.
The key findings show that FlipAttack achieved an average attack success rate of 78.97% across eight models, outperforming the closest black-box baseline by roughly 22 percentage points. Specifically, it achieved success rates of 94.04% on GPT-4 Turbo and 88.08% on Claude 3.5 Sonnet. Furthermore, the disguised prompts bypassed five leading automated guardrail systems with an average bypass rate of 98.08%, including a 100% bypass rate on OpenAI Moderation. Standard heuristic defenses, such as simple system-prompt instructions and basic perplexity-based filtering, failed to mitigate the attack effectively.
These findings indicate that existing guardrail filters and alignment strategies rely heavily on surface-level, standard-order text patterns, leaving substantial blind spots against simple, non-standard structural perturbations. Because this approach succeeds in a single query without requiring model weight access, it represents a cost-effective and highly transferable vulnerability. Standard safety filters that scan for direct keywords or expected text structures will require significant updates to handle disguised transformations.
To mitigate these risks, the article suggests that safety teams integrate advanced red-teaming and alignment procedures that train models to detect structural obfuscation directly. The authors noted that simple countermeasures, like adjusting perplexity thresholds, offer minimal protection while inadvertently raising false-positive rejection rates on benign inputs. Organizations deploying language models should evaluate reasoning-based safeguards rather than relying solely on surface-level guardrails.
The scope of this evaluation presents several limitations. FlipAttack showed lower success rates against certain complex or highly sensitive categories, such as physical harm, compared to categories like digital fraud and software vulnerabilities. Additionally, the approach showed lower effectiveness against specialized reasoning-dense models, and few-shot demonstrations can occasionally expose plain harmful words that trigger standard filters. While confidence in the vulnerability of current standard systems is high, readers should view these findings as a baseline for updating moderation architectures rather than an exhaustive evaluation across all emerging reasoning models.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This foundational work establishes automated token-level optimization and transferability for LLM jailbreaks, providing essential context for black-box and single-query prompt-based attacks.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper examines the structural failure modes of safety training, such as competing objectives and mismatched generalization, which FlipAttack exploits via left-side token transformations.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). Reading AutoDAN helps understand the necessity of creating stealthy, readable jailbreak perturbations that evade guardrails without resorting to nonsensical optimized token strings.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). This study introduces query-efficient semantic black-box jailbreaking, setting a baseline for the low-query black-box threat model improved upon by FlipAttack's single-query method.
- Book: Safety Alignment Should Be Made More Than Just a Few Tokens Deep, Xiangyu Qi et al. (2025). This work demonstrates that alignment is concentrated in the first few autoregressive tokens, offering a theoretical rationale for why left-side perturbations can effectively disarm safety guardrails.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench provides the standardized benchmarks and evaluation methodologies commonly used to assess the efficacy and bypass rates of jailbreak attacks.
- Paper: How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models, Dun Li Chan et al. (2026). This study analyzes how character-, token-, and word-level perturbations propagate through LLM internal representations and attention heads, offering deeper mechanistic insight into why text-flipping perturbations succeed.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). This work extends adversarial text manipulation into evasion attacks against AI safety detectors and classifiers using iterative and structural paraphrasing.
