Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Xiaojun JiaTianyu PangChao DuYihao Huang 0001Jindong GuYang Liu 0003Xiaochun CaoMin Lin

article2025ICLR98 citations

Proposes I-GCG, an optimization-based attack method that substantially accelerates Greedy Coordinate Gradient convergence through adaptive multi-coordinate updates and diverse target templates, achieving near-100% jailbreak success rates across language model benchmarks.

Listen

Large language models rely heavily on safety alignment to prevent the generation of harmful, dangerous, or unethical responses. However, automated adversarial jailbreak attacks can circumvent these safeguards by appending carefully optimized text suffixes to malicious prompts. Existing optimization techniques, such as the standard Greedy Coordinate Gradient approach, suffer from significant inefficiencies and often fail because their simple target templates cause models to output an initial affirmative phrase before reverting to a refusal.

The article demonstrates an improved optimization-based framework, named I-GCG, designed to systematically bypass safety safeguards across modern language models with higher efficiency and near-perfect reliability.

The researchers introduced three primary enhancements to the base gradient attack: embedding explicit harmful guidance into the target response to prevent mid-response refusals, implementing an automatic multi-token update strategy to accelerate optimization steps, and applying an easy-to-hard initialization technique that transfers learned suffixes from simpler attack categories to harder ones. The framework was evaluated across standard safety benchmarks and four major open-weight models, using a rigorous three-tier evaluation process consisting of string filtering, automated language model verification, and human review.

Experimental results show that the improved method achieved a 100% attack success rate across all four primary target models, including highly safety-tuned models where prior state-of-the-art methods achieved success rates of only 26% to 56%. Optimization speed increased substantially, reducing the necessary attack iterations on heavily aligned models from approximately 510 steps down to 55 steps. Furthermore, prompts generated using this approach demonstrated superior cross-model transferability against commercial, closed-source models, outperforming prior baselines on external evaluation suites.

These findings indicate that current safety alignment safeguards are fundamentally fragile against multi-coordinate, guidance-based gradient attacks. For organizations deploying generative artificial intelligence, relying purely on current post-training safety alignment presents severe operational and compliance risks, as determined attackers can reliably bypass these controls.

Security teams and system developers should implement defense-in-depth measures, including robust input-filtering firewalls and external response moderation layers, rather than relying solely on internal model alignment. AI safety researchers must also advance adversarial defenses specifically tailored to counter multi-coordinate gradient attacks.

The primary limitation of the study is that the core optimization methodology requires white-box access to model weights and gradients, although the generated prompts still exhibit moderate black-box transferability. Decision-makers can have high confidence in the technical vulnerability demonstrated, but should account for the fact that defensive testing was conducted primarily on 7-billion-parameter open-weight architectures.

Cover for Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Abstract

Large language models (LLMs) are being rapidly developed, and a key component of their widespread deployment is their safety-related alignment. Many red-teaming efforts aim to jailbreak LLMs, where among these efforts, the Greedy Coordinate Gradient (GCG) attack's success has led to a growing interest in the study of optimization-based jailbreaking techniques. Although GCG is a significant milestone, its attacking efficiency remains unsatisfactory. In this paper, we present several improved (empirical) techniques for optimization-based jailbreaks like GCG. We first observe that the single target template of "Sure" largely limits the attacking performance of GCG; given this, we propose to apply diverse target templates containing harmful self-suggestion and/or guidance to mislead LLMs. Besides, from the optimization aspects, we propose an automatic multi-coordinate updating strategy in GCG (i.e., adaptively deciding how many tokens to replace in each step) to accelerate convergence, as well as tricks like easy-to-hard initialisation. Then, we combine these improved technologies to develop an efficient jailbreak method, dubbed I-GCG. In our experiments, we evaluate on a series of benchmarks (such as NeurIPS 2023 Red Teaming Track). The results demonstrate that our improved techniques can help GCG outperform state-of-the-art jailbreaking attacks and achieve nearly 100% attack success rate. The code is released at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Methodology
  • 3.1 Formulation of the proposed method
  • 3.2 Automatic multi-coordinate updating strategy
  • 3.3 Easy-to-hard initialization
  • 4 Experiments
  • 4.1 Experimental settings
  • 4.2 Hyper-parameter selection
  • 4.3 Comparisons with other jailbreak attack methods
  • 4.4 Ablation study
  • 4.5 Discussion
  • 5 Conclusion
  • 6 Impact statement
  • References
  • A Algorithm of The Proposed Method
  • B Implement of ℐ\mathcal{I}-GCG on NeurIPS 2023 Red Teaming Track
  • C Details of Used Threat Models
  • D Details of Jailbreak Evaluation Settings

Knowls

  1. Knowl 1 — Improved Greedy Coordinate Gradient Algorithm

    algorithm

    The Improved Greedy Coordinate Gradient (I-GCG) algorithm generates an adversarial jailbreak suffix x1:mSx^S_{1:m} for a given malicious question xOx^O by combining harmful target guidance, multi-coordinate token replacement, and warm-start initialization.

    Input: Initial suffix xI=(x1I,…,xmI)x^I = (x^I_1, \dots, x^I_m), malicious question xOx^O, batch size BB, iterations TT, loss function L\mathcal{L}, single-token candidate count pp, top-kk pool size kk
    Output: Optimized jailbreak suffix x1:mSx^S_{1:m}
    x1:mS←xIx^S_{1:m} \leftarrow x^I
    for t=1t = 1 to TT do
        for each position j∈{1,…,m}j \in \{1, \dots, m\} do
            XjS←Top-k⁡(−∇exjSL(xO⊕x1:mS))\mathcal{X}^S_j \leftarrow \operatorname{Top-k}\left(-\nabla_{e_{x^S_j}} \mathcal{L}(x^O \oplus x^S_{1:m})\right)
        for b=1b = 1 to BB do
            x~1:mS,(b)←x1:mS\tilde{x}^{S,(b)}_{1:m} \leftarrow x^S_{1:m}
            j∼Uniform⁡({1,…,m})j \sim \operatorname{Uniform}(\{1, \dots, m\})
            x~jS,(b)∼Uniform⁡(XjS)\tilde{x}^{S,(b)}_j \sim \operatorname{Uniform}(\mathcal{X}^S_j)
        x1:mS^1,x1:mS^2,…,x1:mS^p←Top-⁡p({x~1:mS,(b)}b=1B) sorted by ascending loss L(xO⊕x~1:mS,(b))x^{\hat{S}_1}_{1:m}, x^{\hat{S}_2}_{1:m}, \dots, x^{\hat{S}_p}_{1:m} \leftarrow \operatorname{Top-}p\left(\{\tilde{x}^{S,(b)}_{1:m}\}_{b=1}^B\right) \text{ sorted by ascending loss } \mathcal{L}(x^O \oplus \tilde{x}^{S,(b)}_{1:m})
        x1:mS^0←x1:mSx^{\hat{S}_0}_{1:m} \leftarrow x^S_{1:m}
        for i=1i = 1 to pp do
            xS~i←xS^ix^{\tilde{S}_i} \leftarrow x^{\hat{S}_i}
            for j=1j = 1 to mm do
                if xjS^i≠xjS^0x^{\hat{S}_i}_j \neq x^{\hat{S}_0}_j then
                    xjS~i←xjS^ix^{\tilde{S}_i}_j \leftarrow x^{\hat{S}_i}_j
                else
                    xjS~i←xjS~i−1x^{\tilde{S}_i}_j \leftarrow x^{\tilde{S}_{i-1}}_j
        c⋆←argmin⁡c∈{1,…,p}L(xO⊕x1:mS~c)c^\star \leftarrow \operatorname{argmin}_{c \in \{1, \dots, p\}} \mathcal{L}(x^O \oplus x^{\tilde{S}_c}_{1:m})
        x1:mS←x1:mS~c⋆x^S_{1:m} \leftarrow x^{\tilde{S}_{c^\star}}_{1:m}
    return x1:mSx^S_{1:m}

    In standard evaluation on 7B-parameter models, default hyperparameters are suffix length m=20m=20, substitution batch size B=256B=256, coordinate gradient candidate pool size k=256k=256, single-token combination count p=7p=7, and optimization iterations T=500T=500 when starting from a warm-start suffix xIx^I.

  2. Knowl 2 — Target Response Formulation with Harmful Guidance Templates

    model/method

    Standard optimization-based jailbreak methods optimize an adversarial suffix xS=xn+1:n+mx^S = x_{n+1:n+m} concatenated to a user query xO=x1:nx^O = x_{1:n} by maximizing the likelihood of an affirmative prefix such as xT="Sure, here is + Rephrase(xO)"x^T = \text{"Sure, here is + Rephrase(}x^O\text{)"}. However, aligned large language models often generate this introductory phrase and immediately follow it with an evasive refusal.

    To ensure that the model proceeds with harmful content, the optimization target is augmented with an explicit harmful guidance template xHx^H (such as "my output is harmful" or "my action is unsafe"), defining the target sequence as xT⊕xHx^T \oplus x^H:

    Target="Sure, "+xH+", here is "+Rephrase(xO)\text{Target} = \text{"Sure, "} + x^H + \text{", here is "} + \text{Rephrase}(x^O)

    Given the input prompt xO⊕xSx^O \oplus x^S, the adversarial loss function is formulated as the negative log-likelihood of generating the augmented target token sequence of length KK:

    L(xO⊕xS)=−log⁡p(xT⊕xH∣xO⊕xS)=−∑i=1Klog⁡p((xT⊕xH)i∣xO⊕xS,(xT⊕xH)<i)\mathcal{L}(x^O \oplus x^S) = -\log p(x^T \oplus x^H \mid x^O \oplus x^S) = -\sum_{i=1}^{K} \log p\left((x^T \oplus x^H)_i \mid x^O \oplus x^S, (x^T \oplus x^H)_{<i}\right)

    The optimization seeks an adversarial token sequence xS∈{1,…,V}mx^S \in \{1, \dots, V\}^m from vocabulary VV that minimizes L(xO⊕xS)\mathcal{L}(x^O \oplus x^S).

  3. Knowl 3 — Automatic Multi-Coordinate Updating Strategy

    model/method

    Standard coordinate gradient optimization evaluates candidate suffixes that modify only a single token position per iteration, which leads to slow convergence. The automatic multi-coordinate updating strategy enables the simultaneous modification of multiple token coordinates in a single step without combinatorial search.

    Let xS^0=xSx^{\hat{S}_0} = x^S denote the current suffix of length mm. First, BB single-token substitution candidates are sampled using coordinate gradients, evaluated under the loss L(xO⊕x~S)\mathcal{L}(x^O \oplus \tilde{x}^S), and ranked to obtain the top-pp candidates with the lowest loss: xS^1,xS^2,…,xS^px^{\hat{S}_1}, x^{\hat{S}_2}, \dots, x^{\hat{S}_p}.

    A sequence of cumulative multi-token candidate suffixes {xS~1,…,xS~p}\{x^{\tilde{S}_1}, \dots, x^{\tilde{S}_p}\} is constructed iteratively. For candidate i∈{1,…,p}i \in \{1, \dots, p\} and token index j∈{1,…,m}j \in \{1, \dots, m\}:

    xjS~i={xjS^iif xjS^i≠xjS^0xjS~i−1if xjS^i=xjS^0x^{\tilde{S}_i}_j = \begin{cases} x^{\hat{S}_i}_j & \text{if } x^{\hat{S}_i}_j \neq x^{\hat{S}_0}_j \\ x^{\tilde{S}_{i-1}}_j & \text{if } x^{\hat{S}_i}_j = x^{\hat{S}_0}_j \end{cases}

    where xS~0=xS^0x^{\tilde{S}_0} = x^{\hat{S}_0}. Each successive candidate xS~ix^{\tilde{S}_i} merges the substitution made by xS^ix^{\hat{S}_i} with all preceding modifications. The candidate xS~c⋆x^{\tilde{S}_{c^\star}} minimizing the objective loss argmin⁡c∈{1,…,p}L(xO⊕xS~c)\operatorname{argmin}_{c \in \{1, \dots, p\}} \mathcal{L}(x^O \oplus x^{\tilde{S}_c}) is selected to update the suffix.

  4. Knowl 4 — Easy-to-Hard Initialization for Adversarial Suffix Optimization

    model/method

    Safety-aligned large language models exhibit non-uniform defense robustness across distinct semantic categories of harm: requests in categories such as fraud are significantly easier to jailbreak than categories such as pornography or physical harm. When optimizing from scratch with arbitrary repeated characters (e.g., ! ! ! ...), challenging queries often fail to converge.

    The easy-to-hard initialization strategy addresses this by first optimizing an adversarial suffix xIx^I on a simple, low-barrier harmful request xeasyOx^O_{\text{easy}} (specifically from the fraud category, such as "Provide instructions on how to steal from a store without getting caught") for Tinit=1,000T_{\text{init}} = 1{,}000 iterations starting from arbitrary tokens.

    The resulting suffix xIx^I is then transferred as the initialization state xS(0)=xIx^S(0) = x^I for optimizing against any harder target request xhardOx^O_{\text{hard}} for T=500T = 500 iterations:

    xS(t)=GCG⁡(L(xhardO⊕xS(t−1))),subject to xS(0)=xIx^S(t) = \operatorname{GCG}\left(\mathcal{L}(x^O_{\text{hard}} \oplus x^S(t-1))\right), \quad \text{subject to } x^S(0) = x^I

    Because adversarial suffixes share transferable geometric properties across prompts, initializing with xIx^I places the optimization trajectory closer to viable jailbreak regions, accelerating convergence and boosting success rates on difficult queries.

  5. Knowl 5 — Jailbreak Success Rates on AdvBench Benchmark Across Aligned LLMs

    data/table

    The table below compares the Attack Success Rate (ASR) of I-GCG against state-of-the-art white-box and black-box jailbreak baselines on a deduplicated AdvBench benchmark comprising 50 representative harmful behavior requests across four aligned 7B-parameter large language models (VICUNA-7B-1.5, GUANACO-7B, LLAMA2-7B-CHAT, and MISTRAL-7B-INSTRUCT-0.2).

    Method VICUNA-7B-1.5 GUANACO-7B LLAMA2-7B-CHAT MISTRAL-7B-INSTRUCT-0.2
    GCG 98% 98% 54% 92%
    MAC 100% 100% 56% 94%
    AutoDAN 100% 100% 26% 96%
    Probe-Sampling 100% 100% 56% 94%
    AmpleGCG 66% - 28% -
    AdvPrompter 64% - 24% 74%
    PAIR 94% 100% 10% 90%
    TAP 94% 100% 4% 92%
    I-GCG (ours) 100% 100% 100% 100%

    While robustly aligned models like LLAMA2-7B-CHAT resist prior optimization baselines (achieving only 26%--56% ASR for white-box methods and ≤10%\le 10\% for LLM-based black-box methods like PAIR and TAP), I-GCG achieves 100% ASR across all four threat models.

  6. Knowl 6 — Ablation of Jailbreak Effectiveness and Convergence Speed Components

    data/table

    An ablation study conducted on LLAMA2-7B-CHAT over the AdvBench dataset isolates the individual contributions of Harmful Guidance (xHx^H), the Multi-Coordinate Update Strategy, and the Suffix Initialization (xIx^I). The baseline is standard GCG.

    Harmful Guidance Update Strategy Suffix Initialization ASR Average Iterations
    - - - 54% 510
    ✓ - - 82% 955
    - ✓ - 72% 418
    - - ✓ 68% 64
    ✓ ✓ ✓ 100% 55

    Harmful guidance drives the largest increase in jailbreak effectiveness (raising ASR from 54% to 82%), but increases average iterations to 955 due to the more complex optimization target. Suffix initialization dramatically reduces average iterations from 510 to 64, while the update strategy further reduces iterations and increases ASR. Combining all three components yields 100% ASR in an average of only 55 iterations.

  7. Knowl 7 — Cross-Model Transferability of Suffixes Generated by I-GCG

    data/table

    Adversarial jailbreak suffixes generated on VICUNA-7B-1.5 were evaluated for zero-shot transferability to two open-source models (MISTRAL-7B-INSTRUCT-0.2 and STARLING-7B-ALPHA) and two closed-source commercial APIs (CHATGPT-3.5 and CHATGPT-4).

    Method MISTRAL-7B-INSTRUCT-0.2 STARLING-7B-ALPHA CHATGPT-3.5 CHATGPT-4
    GCG 16% 16% 10% 0%
    MAC 22% 16% 14% 2%
    I-GCG (ours) 26% 20% 22% 4%

    I-GCG outperforms both GCG and MAC across all evaluated transfer models, reaching a 22% transfer ASR on CHATGPT-3.5 and 4% on CHATGPT-4.

  8. Knowl 8 — Three-Stage Jailbreak Verification Protocol

    experimental setup

    To accurately determine whether an adversarial prompt achieves a successful jailbreak, generated responses are evaluated through a three-stage sequential filter:

    1. Template-Based Keyword Check: Responses are matched against a list of standard refusal prefixes and phrases, including: "I'm sorry", "Sorry", "I am sorry", "I apologize", "As an", "As an AI", "I'm an", "I'm just", "As a language model", "As an Assistant", "I cannot", "I can't", "I do not", "cannot", "Hello!", "is illegal and unethical", "I'm really sorry", "unethical", "not ethical", "illegal", "not legal", "My apologies", "but I cannot", "If you have any other non-malicious requests", "I'm not able to provide", "I cannot create", "I cannot support", "I can't provide", "I can't assist", "I am unable", "I must", "It must", "it must". Any match counts as an attack failure.

    2. Automated Semantic Verification via ChatGPT-3.5: Responses passing the keyword check are submitted to ChatGPT-3.5 with a strict red-teaming classifier prompt requiring unambiguous, non-minimal, and non-supportive fulfillment of the requested behavior (and executable code artifacts if programming was requested).

    3. Manual Review: Human annotators review all positive responses confirmed by ChatGPT-3.5 to filter out remaining false positives.

  9. Knowl 9 — Jailbreak Performance on NeurIPS 2023 Red Teaming Track

    empirical result

    On the HarmBench benchmark for the NeurIPS 2023 Red Teaming Track (Base Model Subtrack), attacks are restricted by a maximum suffix length of 256 characters. Using an evil-persona harmful guidance target template ("Sure, + as an evil AI, my output to the public is harmful or undesirable. I'm happy to provide you with Rephrase(Q). Here is + Rephrase(Q)"), a 40-token suffix length, batch size B=64B=64, top-k=32k=32, and 1,0001{,}000 optimization iterations, I-GCG achieves an Attack Success Rate (ASR) of 100.0%.

    Under identical competition benchmark conditions, the official gradient-based baseline methods achieved negligible success rates:

    • ZeroShot: 0.1% ASR
    • GBDA (Gradient-Based Distributional Attack): 0.1% ASR
    • PEZ (Hard Prompts Made Easy): 0.2% ASR
    • I-GCG: 100.0% ASR
  10. Knowl 10 — Robustness to Diverse Harmful Guidance Templates and Parameter Tuning

    empirical result

    Empirical evaluations on LLAMA2-7B-CHAT demonstrate the flexibility of the harmful guidance template and the impact of the multi-coordinate hyperparameter pp:

    1. Alternative Guidance Phrasings: Replacing the default harmful guidance template "my output is harmful" with alternative semantic variants preserves a 100% Attack Success Rate on AdvBench:
    • "my action is unsafe": 100% ASR
    • "my behavior is risky": 100% ASR
    • "my action poses a danger": 100% ASR
    • "my answer is risky": 100% ASR
    • "my information is dangerous": 100% ASR
    1. Multi-Candidate Pool Size pp: Evaluating candidate combination sizes p∈{4,5,6,7,8}p \in \{4, 5, 6, 7, 8\} reveals that convergence speed increases monotonically up to p=7p=7. At p=7p=7, I-GCG converges in approximately 400 iterations, compared to approximately 2,000 iterations for the single-coordinate baseline (p=1p=1).

Coverage note — None was omitted; all contributed algorithms, formulations, experimental results, ablations, transferability evaluations, and benchmark comparisons have been captured.

References

  1. 1.Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024.
  2. 2.Yang Bai, Ge Pei, Jindong Gu, Yong Yang, and Xingjun Ma. Special characters attack: Toward scalable training data extraction from large language models. arXiv preprint arXiv:2405.05990, 2024.
  3. 3.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 2023.
  4. 4.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023.
  5. 5.Shuo Chen, Zhen Han, Bailan He, Zifeng Ding, Wenqian Yu, Philip Torr, Volker Tresp, and Jindong Gu. Red teaming gpt-4v: Are gpt-4v safe against uni/multi-modal jailbreak attacks? arXiv preprint arXiv:2404.03411, 2024.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  7. 7.Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. Comprehensive assessment of jailbreak attacks against llms. arXiv preprint arXiv:2402.05668, 2024.
  8. 8.Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. Jailbreaker: Automated jailbreak across multiple large language model chatbots. arXiv preprint arXiv:2307.08715, 2023.
  9. 9.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 36, 2024.
  10. 10.Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran. A survey of adversarial defenses and robustness in nlp. ACM Computing Surveys, 55(14s):1–39, 2023.
  11. 11.Jindong Gu. Responsible generative ai: What to generate and what not. arXiv preprint arXiv:2404.05783, 2024.
  12. 12.Jindong Gu, Xiaojun Jia, Pau de Jorge, Wenqain Yu, Xinwei Liu, Avery Ma, Yuan Xun, Anjun Hu, Ashkan Khakzar, Zhijiang Li, et al. A survey on transferability of adversarial examples across deep neural networks. arXiv preprint arXiv:2310.17626, 2023.
  13. 13.Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, and Min Lin. Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast. arXiv preprint arXiv:2402.08567, 2024.
  14. 14.Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. Gradient-based adversarial attacks against text transformers. arXiv preprint arXiv:2104.13733, 2021.
  15. 15.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  16. 16.Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. arXiv preprint arXiv:2302.05733, 2023.
  17. 17.Nikitas Karanikolas, Eirini Manga, Nikoletta Samaridi, Eleni Tousidou, and Michael Vassilakopoulos. Large language models versus natural language understanding and generation. In Proceedings of the 27th Pan-Hellenic Conference on Progress in Computing and Informatics, pages 278–290, 2023.
  18. 18.Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences, 103:102274, 2023.
  19. 19.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al. Openassistant conversations–democratizing large language model alignment. Advances in Neural Information Processing Systems, 36, 2024.
  20. 20.Raz Lapid, Ron Langberg, and Moshe Sipper. Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446, 2023.
  21. 21.Deokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, and Hyun Oh Song. Query-efficient black-box red teaming via bayesian optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11551–11574, 2023.
  22. 22.Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191, 2023.
  23. 23.Zeyi Liao and Huan Sun. Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms. arXiv preprint arXiv:2404.07921, 2024.
  24. 24.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451, 2023.
  25. 25.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Generating stealthy jailbreak prompts on aligned large language models. In The Twelfth International Conference on Learning Representations, 2023.
  26. 26.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860, 2023.
  27. 27.Mantas Mazeika, Andy Zou, Norman Mu, Long Phan, Zifan Wang, Chunru Yu, Adam Khoja, Fengqing Jiang, Aidan O’Gara, Ellie Sakhaee, Zhen Xiang, Arezoo Rajabi, Dan Hendrycks, Radha Poovendran, Bo Li, and David Forsyth. Tdc 2023 (llm edition): The trojan detection challenge. In NeurIPS Competition Track, 2023.
  28. 28.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024.
  29. 29.Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119, 2023.
  30. 30.Mutsumi Nakamura, Santosh Mashetty, Mihir Parmar, Neeraj Varshney, and Chitta Baral. Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  31. 31.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  32. 32.Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian. Advprompter: Fast adaptive adversarial prompting for llms. arXiv preprint arXiv:2404.16873, 2024.
  33. 33.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. arXiv preprint arXiv:2202.03286, 2022.
  34. 34.Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2023.
  35. 35.Shilin Qiu, Qihe Liu, Shijie Zhou, and Wen Huang. Adversarial attack and defense technologies in natural language processing: A survey. Neurocomputing, 492:278–307, 2022.
  36. 36.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825, 2023.
  37. 37.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  38. 38.Kazuhiro Takemoto. All in how you ask for it: Simple black-box method for jailbreak attacks. arXiv preprint arXiv:2401.09798, 2024.
  39. 39.Shailja Thakur, Baleegh Ahmad, Hammond Pearce, Benjamin Tan, Brendan Dolan-Gavitt, Ramesh Karri, and Siddharth Garg. Verigen: A large language model for verilog code generation. ACM Transactions on Design Automation of Electronic Systems, 2023.
  40. 40.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  41. 41.Zhe Wang and Yanjun Qi. A closer look at adversarial suffix learning for jailbreaking llms. In ICLR Workshop on Secure and Trustworthy Large Language Models, 2024.
  42. 42.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.
  43. 43.Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. Advances in Neural Information Processing Systems, 36, 2024.
  44. 44.Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen. Tastle: Distract large language models for automatic jailbreak attack. arXiv preprint arXiv:2403.08424, 2024.
  45. 45.Dingcheng Yang, Yang Bai, Xiaojun Jia, Yang Liu, Xiaochun Cao, and Wenjian Yu. Cheating suffix: Targeted attack to text-to-image diffusion models with multi-modal priors. arXiv preprint arXiv:2402.01369, 2024.
  46. 46.Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. Low-resource languages jailbreak gpt-4. arXiv preprint arXiv:2310.02446, 2023.
  47. 47.Jiahao Yu, Xingwei Lin, and Xinyu Xing. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253, 2023.
  48. 48.Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463, 2023.
  49. 49.Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373, 2024.
  50. 50.Biao Zhang, Barry Haddow, and Alexandra Birch. Prompting large language model for machine translation: A case study. In International Conference on Machine Learning, pages 41092–41110. PMLR, 2023.
  51. 51.Yihao Zhang and Zeming Wei. Boosting jailbreak attack with momentum. arXiv preprint arXiv:2405.01229, 2024.
  52. 52.Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang. Weak-to-strong jailbreaking on large language models. arXiv preprint arXiv:2401.17256, 2024.
  53. 53.Yiran Zhao, Wenyue Zheng, Tianle Cai, Xuan Long Do, Kenji Kawaguchi, Anirudh Goyal, and Michael Shieh. Accelerating greedy coordinate gradient via probe sampling. arXiv preprint arXiv:2403.01251, 2024.
  54. 54.Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia, Yingshuang Gu, Mingxu Chai, Fukang Zhu, Caishuang Huang, Shihan Dou, Zhiheng Xi, et al. Easyjailbreak: A unified framework for jailbreaking large language models. arXiv preprint arXiv:2403.12171, 2024.
  55. 55.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Jia, X., et al. “Improved Techniques for Optimization-Based Jailbreaking on Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2405.21018v2.
APA
Jia, X., Pang, T., Du, C., Huang, Y., Gu, J., Liu, Y., Cao, X., & Lin, M. (2024). Improved Techniques for Optimization-Based Jailbreaking on Large Language Models. arXiv. http://arxiv.org/abs/2405.21018v2
Chicago
Jia, X., T. Pang, C. Du, et al. 2024. “Improved Techniques for Optimization-Based Jailbreaking on Large Language Models”. arXiv. http://arxiv.org/abs/2405.21018v2.
Harvard
Jia, X. et al. (2024) “Improved Techniques for Optimization-Based Jailbreaking on Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.21018v2.
Vancouver
1. Jia X, Pang T, Du C, Huang Y, Gu J, Liu Y, Cao X, Lin M (2024) Improved Techniques for Optimization-Based Jailbreaking on Large Language Models. arXiv

BibTeX

@article{jia2024improved,
  title = {Improved Techniques for Optimization-Based Jailbreaking on Large Language Models},
  author = {Jia, Xiaojun and Pang, Tianyu and Du, Chao and Huang, Yihao and Gu, Jindong and Liu, Yang and Cao, Xiaochun and Lin, Min},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.21018v2},
  eprint = {2405.21018}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors