A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Peng DingJun KuangDan MaXuezhi CaoYunsen XianJiajun ChenShujian Huang
Proposes ReNeLLM, an automated framework combining prompt rewriting with scenario nesting to generate effective jailbreak prompts using language models themselves, revealing safety vulnerabilities and informing stronger defense strategies.
Large Language Models are widely deployed across industries, yet they remain vulnerable to adversarial jailbreak prompts that bypass safety safeguards to produce harmful material. Current jailbreak techniques typically depend either on complex manual engineering that degrades as systems update, or on computationally intensive optimization algorithms that produce nonsensical text and fail to transfer across commercial platforms. The article introduces and evaluates an automated, efficient jailbreak framework named ReNeLLM to systematically expose these safety vulnerabilities and evaluate defensive countermeasures.
The article demonstrates that jailbreak attacks can be generalized into a two-step automated process: prompt rewriting and scenario nesting. Prompt rewriting disguises malicious intent without altering core semantics through techniques such as paraphrasing, misspelling sensitive terms, or partial translation. Scenario nesting embeds these rewritten requests into common operational tasks, including code completion, table filling, and text continuation. Using a benchmark of 520 harmful behavior prompts across seven risk categories, the authors tested this framework against five prominent models—including GPT-3.5, GPT-4, Claude-1, Claude-2, and Llama 2 variants—and compared it against leading baseline attack methods.
The findings reveal that ReNeLLM consistently outperforms existing attack methods across both open-source and proprietary commercial systems. On average, single-attempt attack success rates reached 86.9% on GPT-3.5, 90.0% on Claude-1, 69.6% on Claude-2, 58.9% on GPT-4, and 51.2% on Llama 2. When tested in an ensemble configuration with six candidate prompts, attack success rates rose to between 94.2% and 99.8% across all evaluated models. Concurrently, the approach reduced prompt generation time by 76.6% compared to gradient-based methods and 86.2% compared to genetic algorithm baselines, with most successful prompts generated within three iterations. Attention visualization experiments revealed that the combination of rewriting and nesting shifts a model's processing attention away from harmful core instructions toward harmless outer task wrappers, causing models to prioritize task execution over safety constraints.
These results demonstrate significant gaps in current safety alignment and filtering tools. Common industry defenses, such as moderation classifiers and perplexity filters, largely failed against these semantically meaningful attacks. Furthermore, fine-tuning on specific nested scenarios failed to generalize across other task structures. To address these vulnerabilities, organizations deploying language models should implement defense prompts that explicitly instruct models to prioritize safety over helpfulness and enforce mandatory multi-step prompt scrutiny prior to response generation, which lowered attack success rates to near zero in testing. Decision-makers must recognize that existing safety safeguards are insufficient against structured prompt nesting, and future governance should focus on developing generalized defense mechanisms that remain resilient across diverse task formats and languages.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper establishes the foundational conceptual failure modes of LLM safety training—competing objectives and mismatched generalization—which directly underpin the prompt rewriting and nesting attack strategies explored in ReNeLLM.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). This study introduces automated iterative prompt refinement using attacker LLMs against target models, providing the direct black-box adversarial framework that ReNeLLM builds upon.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This work introduces universal optimization-based jailbreak suffixes and the standard AdvBench evaluation methodology that serves as a primary baseline and benchmark for ReNeLLM.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). This paper pioneers the paradigm of using language models to automatically red-team and generate adversarial test inputs against other language models.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This work details LLM-based input-output safety safeguards, providing essential context for evaluating the defense failures and priority mechanisms analyzed in ReNeLLM.
- Paper: Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition, Sander Schulhoff et al. (2023). This comprehensive empirical ontology of human-designed prompt injections and jailbreaks provides key taxonomy and threat models that automated rewriting frameworks aim to replicate.
- Paper: How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs, Yi Zeng et al. (2024). This paper extends automated linguistic jailbreak generation by formalizing structured human persuasion techniques to systematically bypass safety alignment.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). This study generalizes automated adversarial prompt generation by integrating open-ended quality-diversity search to create diverse collections of jailbreak prompts across multiple risk domains.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This work establishes a standardized benchmark (HarmBench) that unifies the evaluation of automated red-teaming attacks and robust refusal defenses like those introduced in ReNeLLM.
- Paper: FlipAttack: Jailbreak LLMs via Flipping, Yue Liu 0008 et al. (2025). This research advances black-box single-query jailbreak strategies by employing text-flipping transformations to bypass content moderation filters and exploit autoregressive token processing.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This paper investigates lightweight adaptive jailbreaks across leading commercial and open LLMs, demonstrating further practical failure modes in current safety guardrails.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). This paper introduces decoding-time alignment mechanisms that dynamically enforce safety and helpfulness constraints to defend against prompt-based jailbreaks without retraining.
