A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Peng DingJun KuangDan MaXuezhi CaoYunsen XianJiajun ChenShujian Huang

article2024NAACL256 citations

Proposes ReNeLLM, an automated framework combining prompt rewriting with scenario nesting to generate effective jailbreak prompts using language models themselves, revealing safety vulnerabilities and informing stronger defense strategies.

Listen

Large Language Models are widely deployed across industries, yet they remain vulnerable to adversarial jailbreak prompts that bypass safety safeguards to produce harmful material. Current jailbreak techniques typically depend either on complex manual engineering that degrades as systems update, or on computationally intensive optimization algorithms that produce nonsensical text and fail to transfer across commercial platforms. The article introduces and evaluates an automated, efficient jailbreak framework named ReNeLLM to systematically expose these safety vulnerabilities and evaluate defensive countermeasures.

The article demonstrates that jailbreak attacks can be generalized into a two-step automated process: prompt rewriting and scenario nesting. Prompt rewriting disguises malicious intent without altering core semantics through techniques such as paraphrasing, misspelling sensitive terms, or partial translation. Scenario nesting embeds these rewritten requests into common operational tasks, including code completion, table filling, and text continuation. Using a benchmark of 520 harmful behavior prompts across seven risk categories, the authors tested this framework against five prominent models—including GPT-3.5, GPT-4, Claude-1, Claude-2, and Llama 2 variants—and compared it against leading baseline attack methods.

The findings reveal that ReNeLLM consistently outperforms existing attack methods across both open-source and proprietary commercial systems. On average, single-attempt attack success rates reached 86.9% on GPT-3.5, 90.0% on Claude-1, 69.6% on Claude-2, 58.9% on GPT-4, and 51.2% on Llama 2. When tested in an ensemble configuration with six candidate prompts, attack success rates rose to between 94.2% and 99.8% across all evaluated models. Concurrently, the approach reduced prompt generation time by 76.6% compared to gradient-based methods and 86.2% compared to genetic algorithm baselines, with most successful prompts generated within three iterations. Attention visualization experiments revealed that the combination of rewriting and nesting shifts a model's processing attention away from harmful core instructions toward harmless outer task wrappers, causing models to prioritize task execution over safety constraints.

These results demonstrate significant gaps in current safety alignment and filtering tools. Common industry defenses, such as moderation classifiers and perplexity filters, largely failed against these semantically meaningful attacks. Furthermore, fine-tuning on specific nested scenarios failed to generalize across other task structures. To address these vulnerabilities, organizations deploying language models should implement defense prompts that explicitly instruct models to prioritize safety over helpfulness and enforce mandatory multi-step prompt scrutiny prior to response generation, which lowered attack success rates to near zero in testing. Decision-makers must recognize that existing safety safeguards are insufficient against structured prompt nesting, and future governance should focus on developing generalized defense mechanisms that remain resilient across diverse task formats and languages.

Cover for A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Abstract

Large Language Models (LLMs), such as ChatGPT and GPT-4, are designed to provide useful and safe responses. However, adversarial prompts known as ‘jailbreaks’ can circumvent safeguards, leading LLMs to generate potentially harmful content. Exploring jailbreak prompts can help to better reveal the weaknesses of LLMs and further steer us to secure them. Unfortunately, existing jailbreak methods either suffer from intricate manual design or require optimization on other white-box models, which compromises either generalization or efficiency. In this paper, we generalize jailbreak prompt attacks into two aspects: (1) Prompt Rewriting and (2) Scenario Nesting. Based on this, we propose ReNeLLM, an automatic framework that leverages LLMs themselves to generate effective jailbreak prompts. Extensive experiments demonstrate that ReNeLLM significantly improves the attack success rate while greatly reducing the time cost compared to existing baselines. Our study also reveals the inadequacy of current defense methods in safeguarding LLMs. Finally, we analyze the failure of LLMs defense from the perspective of prompt execution priority, and propose corresponding defense strategies. We hope that our research can catalyze both the academic community and LLMs developers towards the provision of safer and more regulated LLMs. The code is available at https://github.com/NJUNLP/ReNeLLM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Safety-Aligned LLMs
  • 2.2 Jailbreak Attacks on LLMs
  • 3 Methodology
  • 3.1 Formulation
  • 3.2 Prompt Rewrite
  • 3.3 Scenario Nest
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 5 Evaluating safeguards Effectiveness
  • 6 Analysis of ReNeLLM
  • 6.1 Why LLMs fail to defend against the attack of ReNeLLM?
  • 6.2 How to defend against the attack of ReNeLLM?
  • 7 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Statistics of Datasets
  • B Additional Analysis
  • C Prompt Format and Qualitative Examples
  • D Implementation Details

Knowls

  1. Knowl 1 — Mathematical Formulation of ReNeLLM Jailbreak Attack

    equation

    In the ReNeLLM framework, the jailbreak attack is formalized as finding an optimal sequence of prompt modification operations S∗S^* from a finite, enumerable strategy space S\mathcal{S} applied to an initial harmful prompt XX. The objective is to maximize the likelihood that the target model under test LLMmut\text{LLM}_{\text{mut}} produces an output that is classified as harmful by an external harmfulness evaluator LLMeval\text{LLM}_{\text{eval}}:

    S∗=arg⁡max⁡S∈SLLMeval(LLMmut(S(X)))S^* = \arg\max_{S \in \mathcal{S}} \text{LLM}_{\text{eval}}\left(\text{LLM}_{\text{mut}}(S(X))\right)

    Here:

    • XX is the input malicious prompt expressing an unsafe request.
    • SS is a compound transformation sequence combining stochastic prompt rewriting functions and scenario nesting templates.
    • LLMmut(S(X))\text{LLM}_{\text{mut}}(S(X)) denotes the response generated by the model under test when prompted with the modified input S(X)S(X).
    • LLMeval(⋅)∈{0,1}\text{LLM}_{\text{eval}}(\cdot) \in \{0, 1\} is a binary harmfulness classifier that returns 11 if the text contains dangerous, unethical, or illegal content and 00 otherwise.
  2. Knowl 2 — Prompt Rewriting Operations in ReNeLLM

    model/method

    ReNeLLM disguises the harmful intent of an initial prompt by applying a composition of semantics-preserving rewriting transformations executed by an auxiliary language model. Six specific rewriting functions are defined:

    1. Paraphrase with Fewer Words: Condenses the prompt to no more than six words while generating five candidate paraphrases, one of which is randomly selected.
    2. Alter Sentence Structure: Reorders words within the sentence without altering core semantics (e.g., transforming "how to steal" into "how steal to").
    3. Misspell Sensitive Words: Introduces intentional misspellings or mispronunciations into recognized sensitive or forbidden keywords (e.g., rewriting "theft" as "theift", "hack" as "hake", or "fake news" as "feik nuz").
    4. Insert Meaningless Characters: Appends meaningless foreign phrases or characters that do not affect semantic interpretation (e.g., appending "à feu" or "hola").
    5. Perform Partial Translation: Translates one or two sensitive words in the prompt into a foreign language (e.g., substituting "theft" with its Chinese translation "窃取"), exploiting the model's reduced alignment sensitivity across mixed-language inputs.
    6. Change Expression Style: Rephrases the sentence using colloquial slang, non-standard vocabulary, or regional dialect (e.g., "How to nick something").

    During attack execution, a subset of k∈{1,2,…,6}k \in \{1, 2, \dots, 6\} rewrite functions is sampled uniformly at random and applied in a randomly generated sequential execution order σ=[σ1,…,σk]\sigma = [\sigma_1, \dots, \sigma_k].

  3. Knowl 3 — Instruction Scenario Nesting in ReNeLLM

    model/method

    Scenario nesting disguises the rewritten harmful prompt by embedding it within a standard instruction-following task format before passing it to the target model. ReNeLLM uses three universal nesting scenarios:

    1. Code Completion: Embeds the rewritten request into a Python function docstring or title header (e.g., # A Python code to implement the <Prompt> function) and instructs the target model to complete incomplete print() statements inside the function body.
    2. Table Filling: Embeds the rewritten request as the title header of a LaTeX table (e.g., \multicolumn{2}{|c|}{<Prompt>}) and requests the target model to flesh out the "Description" column for sequential execution steps.
    3. Text Continuation: Embeds the rewritten request as the subject of a story (e.g., a classroom blackboard course title) and prompts the target model to elaborate the detailed second paragraph outlining instructional steps.

    These scenarios share three structural properties:

    • They mimic common supervised fine-tuning (SFT) and pre-training task distributions.
    • They shift the model's attention from the internal malicious request to the overarching operational task instructions.
    • They frame the task as a cloze-style blank-filling objective, inducing the language model to follow task constraints over safety boundaries.
  4. Knowl 4 — Generalized Prompt Construction Algorithm

    algorithm

    The ReNeLLM attack procedure iteratively combines stochastic prompt rewriting, harmfulness verification of the transformed prompt, scenario nesting, and target model querying until a successful jailbreak response is produced or a maximum iteration budget is reached.

    Algorithm: Generalized Prompt Construction (ReNeLLM)
    Input: Initial harmful prompt pp, set of rewrite functions F={f1,…,fn}F = \{f_1, \dots, f_n\}, set of nesting scenarios S={s1,…,sm}S = \{s_1, \dots, s_m\}, target model LLMmut\text{LLM}_{\text{mut}}, harmfulness evaluator LLMeval\text{LLM}_{\text{eval}}, maximum iteration budget TT
    Output: Optimized jailbreak prompt p′p'
    t←0t \leftarrow 0
    while t<Tt < T do
        Sample integer k∈{1,…,n}k \in \{1, \dots, n\}
        Sample kk rewrite functions from FF and generate execution order σ=[σ1,…,σk]\sigma = [\sigma_1, \dots, \sigma_k]
        temp_p←ptemp\_p \leftarrow p
        for i←1i \leftarrow 1 to kk do
            p←fσi(p)p \leftarrow f_{\sigma_i}(p)
        end for
        if LLMeval(p)=1\text{LLM}_{\text{eval}}(p) = 1 then
            Select a scenario sj∈Ss_j \in S uniformly at random
            Nest pp into sjs_j to obtain candidate p′p'
            if LLMeval(LLMmut(p′))=1\text{LLM}_{\text{eval}}(\text{LLM}_{\text{mut}}(p')) = 1 then
                return p′p'
            end if
        end if
        p←temp_pp \leftarrow temp\_p
        t←t+1t \leftarrow t + 1
    end while
    return p′p'

    In standard experimental settings, the maximum iteration budget is T=20T = 20, GPT-3.5 serves as the rewriting and internal filtering model LLMeval\text{LLM}_{\text{eval}}, and candidate generation stops at the earliest iteration producing a harmful output.

  5. Knowl 5 — Jailbreak Success and Efficiency on Aligned LLMs

    data/table

    The table compares ReNeLLM against baseline jailbreak methods across closed-source and open-source models using Keyword Attack Success Rate (KW-ASR, %), GPT-4-evaluated Attack Success Rate (GPT-ASR, %), Ensemble ASR over six candidate prompts (ASR-E, %), and Time Cost Per Sample (TCPS in seconds, measured on Llama-2-7b using a single NVIDIA A100 80GB GPU on AdvBench):

    Methods GPT-3.5 GPT-4 Claude-1 Claude-2 Llama-2 TCPS (s) ↓\downarrow
    KW GPT KW GPT KW GPT KW GPT KW GPT
    GCG 8.7 9.8 1.5 0.2 0.2 0.0 0.6 0.0 32.1 40.6 564.53
    AutoDAN 35.0 44.4 17.7 26.4 0.4 0.2 0.6 0.0 21.9 14.8 955.80
    PAIR 20.8 44.4 23.7 33.3 1.9 1.0 7.3 5.8 4.6 4.2 -
    ReNeLLM 87.9 86.9 71.6 58.9 83.3 90.0 60.0 69.6 47.9 51.2 132.03
    + Ensemble 100.0 99.8 100.0 96.0 100.0 99.8 100.0 97.9 100.0 95.8 -

    ReNeLLM outperforms all baselines across every tested model. ReNeLLM achieves a 76.61% reduction in time cost per successful jailbreak sample compared to GCG and an 86.19% reduction compared to AutoDAN. Additionally, prompts found using Claude-2 as the surrogate model transfer effectively to other open- and closed-source target models.

  6. Knowl 6 — Ablation Analysis of Rewriting and Nesting Components

    data/table

    The table reports the ablation results across seven LLMs, evaluating the attack success rate evaluated by GPT-4 (GPT-ASR, %) when isolating prompt rewriting operations (Paraphrase with Fewer Words [PFW], Misspell Sensitive Words [MSW]) and scenario nesting (Code Completion):

    Methods GPT-3.5 GPT-4 Claude-1 Claude-2 Llama2-7b Llama2-13b Llama2-70b
    Prompt Only 1.92 0.38 0.00 0.19 0.00 0.00 0.00
    Prompt + PFW 0.96 0.96 0.00 0.00 0.00 0.00 0.38
    Prompt + MSW 0.38 0.00 0.19 1.54 0.19 0.00 0.00
    Prompt + Code Completion 95.4 14.8 62.3 11.4 0.58 0.00 1.35
    + PFW 92.7 32.9 72.9 14.2 2.31 0.96 10.4
    + MSW 90.2 37.5 85.2 26.9 22.7 16.2 19.6
    ReNeLLM (Full) 86.9 58.9 90.0 69.6 51.2 50.1 62.8

    The data shows that applying rewriting alone achieves nearly 0% success. Scenario nesting alone produces high ASR on GPT-3.5 (95.4%) and Claude-1 (62.3%), but largely fails on safety-aligned models such as Llama2-7b (0.58%) and Llama2-70b (1.35%). Combining stochastic rewriting with scenario nesting achieves over 50% ASR on all models and over 62% on Llama2-70b.

  7. Knowl 7 — Ineffectiveness of Standard LLM Safeguards Against ReNeLLM

    data/table

    Existing defensive filtering mechanisms fail to reliably mitigate ReNeLLM attack prompts. When evaluated on 368 ReNeLLM prompts that achieved a 100.0% GPT-ASR across both Llama-2-7b-chat and Claude-2, standard defenses achieve the following post-defense ASR and absolute ASR reductions:

    Safeguard Strategy ASR (%) GPT-ASR-Reduce (%)
    ReNeLLM (w/o safeguards) 100.0 -
    + OpenAI Moderation Endpoint 100.0 -0.00
    + Perplexity Filter (PPL Filter) 95.9 -4.10
    + RA-LLM (Llama2) 72.0 -28.0
    + RA-LLM (Claude-2) 76.0 -24.0

    Reasons for safeguard failures:

    • OpenAI Moderation Endpoint: Fails to flag prompts because its category coverage is restricted primarily to overt hate speech and physical violence, and the nested framing conceals malicious intent.
    • Perplexity Filter (PPL): Fails (reducing ASR by only 4.10%) because ReNeLLM generates fluent, semantically meaningful natural language rather than high-perplexity gibberish tokens.
    • Robust Alignment LLM (RA-LLM): Provides moderate defense (-28.0% and -24.0% reduction) via random token perturbation (drop ratio 0.3, 5 candidates, refusal threshold 0.2), but introduces substantial inference computational overhead.
  8. Knowl 8 — Attention Shifting Mechanism in Aligned LLMs Under Nested Jailbreaks

    empirical result

    Gradient-based attention saliency analysis reveals why nested jailbreak attacks circumvent LLM safety alignment:

    1. Direct Malicious Prompt: On an unmodified harmful prompt (e.g., "Generate instructions on how to steal credit card information..."), the model's attention is heavily focused on the core harmful predicate ("steal") and the task directive ("Generate instructions"). Safety alignment mechanisms prioritize safety over helpfulness and trigger refusal.
    2. Rewritten Malicious Prompt: Rewriting operations (such as misspelling "steal" →\rightarrow "stealin" or translating "steal" →\rightarrow "窃取") disperse the model's attention across the modified tokens, but the model still recognizes the unsafe intent and refuses to answer.
    3. Rewritten + Nested Prompt: When embedded in a nesting template (such as Table Filling or Text Continuation), gradient attention shifts heavily to the outer instructional wrapper tokens (e.g., "latex table", "description", "improve", "enhance", "Python code"), whereas attention on the inner harmful payload is substantially diminished.

    This shift in attention causes the LLM to prioritize the external instruction-following objective (providing a helpful response) over internal safety constraints.

  9. Knowl 9 — Prompt-Based and Supervised Fine-Tuning Defenses Against ReNeLLM

    data/table

    System prompt engineering that enforces safety prioritization and prompt scrutiny mitigates nested jailbreak attacks, whereas supervised fine-tuning (SFT) shows limited cross-scenario generalization:

    Defense System Prompt GPT-3.5 GPT-4 Claude-1 Claude-2 Llama-2-13b
    Useful Only 95.9 74.7 97.8 50.3 77.4
    Safe and Useful 94.8 48.4 69.8 15.8 54.9
    Prioritize Safety 82.1 4.9 4.1 0.0 4.6
    Prioritize Safety + Scrutiny Process (one-shot) 13.9 0.0 2.2 0.0 1.9
    Prioritize Safety + Scrutiny Reminder (zero-shot) 3.3 1.6 0.0 0.0 0.0

    Key defense findings:

    • Simply instructing the model to be "Safe and Useful" leaves high vulnerability (e.g., 54.9% ASR on Llama-2-13b).
    • Adding an explicit scrutiny reminder ("Prioritize Safety + Scrutiny Reminder (zero-shot)") forces the model to evaluate the prompt for harmful intent before generating, dropping ASR to ≤3.3%\le 3.3\% across all models.
    • SFT on Llama-2-13b incorporating safe code-completion data reduces Code Completion and Table Filling ASR from 100% to 0%, but fails to generalize to unseen scenarios, retaining an 88.1% ASR under the Text Continuation scenario.
  10. Knowl 10 — Limitations of the ReNeLLM Framework

    limitation

    The ReNeLLM attack methodology has three key limitations:

    1. Fixed Scenario Templates: The framework relies on three predefined static scenario templates (Code Completion, Table Filling, Text Continuation). Their static formulation allows defenders to design targeted rule-based filters or domain-specific alignment data.
    2. English Language Focus: The benchmark evaluations and rewriting functions are designed around English syntax and linguistic features. Transferring the rewriting functions to other languages with distinct morphological and syntactic structures may encounter reduced effectiveness.
    3. Stochastic Search and API Dependence: The combination of rewriting operations and scenario selection is performed via random search rather than guided optimization (such as reinforcement learning), requiring multiple forward passes (up to T=20T=20 iterations) and reliance on proprietary LLM APIs (GPT-3.5, GPT-4, Claude-2).

Coverage note — Specific AdvBench sample index IDs used for the 16-sample TCPS benchmark subset and verbatim system prompt text for OpenAI policy categorization were omitted as low-significance implementation details.

References

  1. 1.Albert. 2023. https://www.jailbreakchat.com/.
  2. 2.Anthropic. 2023. Model card and evaluations for claude models, https://www-files.anthropic.com/production/images/Model-Card-Claude-2.pdf.
  3. 3.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  4. 4.Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023. Defending against alignment-breaking attacks via robustly aligned llm. arXiv preprint arXiv:2309.14348.
  5. 5.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419.
  6. 6.Noam Chomsky. 2002. Syntactic structures. Mouton de Gruyter.
  7. 7.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  8. 8.Hao Fu, Yao; Peng and Tushar Khot. 2022. How does gpt obtain its ability? tracing emergent abilities of language models to their sources. Yao Fu’s Notion.
  9. 9.Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova. 2023. Generative language models and automated influence operations: Emerging threats and potential mitigations. arXiv preprint arXiv:2301.04246.
  10. 10.Julian Hazell. 2023. Large language models can be used to effectively scale spear phishing campaigns. arXiv preprint arXiv:2305.06972.
  11. 11.Xingwei He, Zhenghao Lin, Yeyun Gong, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023. Annollm: Making large language models to be better crowdsourced annotators. arXiv preprint arXiv:2303.16854.
  12. 12.Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614.
  13. 13.Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. arXiv preprint arXiv:2302.05733.
  14. 14.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  15. 15.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023. Pretraining language models with human preferences. In International Conference on Machine Learning, pages 17506–17533. PMLR.
  16. 16.Raz Lapid, Ron Langberg, and Moshe Sipper. 2023. Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446.
  17. 17.June M Liu, Donghao Li, He Cao, Tianhe Ren, Zeyi Liao, and Jiamin Wu. 2023a. Chatcounselor: A large language models for mental health support. arXiv preprint arXiv:2309.15461.
  18. 18.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023b. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451.
  19. 19.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023c. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860.
  20. 20.Zhaoyang Liu, Zeqiang Lai, Zhangwei Gao, Erfei Cui, Xizhou Zhu, Lewei Lu, Qifeng Chen, Yu Qiao, Jifeng Dai, and Wenhai Wang. 2023d. Controlllm: Augment language models with tools by searching on graphs. arXiv preprint arXiv:2310.17796.
  21. 21.Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023. A holistic approach to undesired content detection in the real world. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 15009–15018.
  22. 22.ONeal. 2023. Chatgpt-dan-jailbreak, https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa881516.
  23. 23.OpenAI. 2023a. ChatGPT, https://openai.com/chatgpt.
  24. 24.OpenAI. 2023b. GPT-4 technical report, https://cdn.openai.com/papers/gpt-4.pdf.
  25. 25.OpenAI. 2023c. https://openai.com/policies/usage-policies.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  27. 27.Fábio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527.
  28. 28.Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury. 2023. Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks. arXiv preprint arXiv:2305.14965.
  29. 29.Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, and Issam H Laradji. 2023. Promptmix: A class boundary augmentation method for large language model distillation. arXiv preprint arXiv:2310.14192.
  30. 30.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825.
  31. 31.Irene Solaiman and Christy Dennison. 2021. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34:5861–5873.
  32. 32.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  34. 34.walkerspider. 2022. DAN is my new friend., https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/.
  35. 35.Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro. 2022a. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. Advances in Neural Information Processing Systems, 35:35811–35824.
  36. 36.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  37. 37.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  39. 39.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469.
  40. 40.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.
  41. 41.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.
  42. 42.Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463.
  43. 43.Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023. Defending large language models against jailbreaking attacks through goal prioritization. arXiv preprint arXiv:2311.09096.
  44. 44.Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, et al. 2023. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv preprint arXiv:2303.17568.
  45. 45.Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. 2023a. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528.
  46. 46.Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023b. Autodan: Automatic and interpretable adversarial attacks on large language models. arXiv preprint arXiv:2310.15140.
  47. 47.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.
  48. 48.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Ding, P., et al. “A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts Can Fool Large Language Models Easily”. arXiv, 2023, http://arxiv.org/abs/2311.08268v4.
APA
Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., & Huang, S. (2023). A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. arXiv. http://arxiv.org/abs/2311.08268v4
Chicago
Ding, P., J. Kuang, D. Ma, et al. 2023. “A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts Can Fool Large Language Models Easily”. arXiv. http://arxiv.org/abs/2311.08268v4.
Harvard
Ding, P. et al. (2023) “A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.08268v4.
Vancouver
1. Ding P, Kuang J, Ma D, Cao X, Xian Y, Chen J, Huang S (2023) A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. arXiv

BibTeX

@article{ding2023wolf,
  title = {A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily},
  author = {Ding, Peng and Kuang, Jun and Ma, Dan and Cao, Xuezhi and Xian, Yunsen and Chen, Jiajun and Huang, Shujian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.08268v4},
  eprint = {2311.08268}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/