Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Xiaojun JiaTianyu PangChao DuYihao Huang 0001Jindong GuYang Liu 0003Xiaochun CaoMin Lin
Proposes I-GCG, an optimization-based attack method that substantially accelerates Greedy Coordinate Gradient convergence through adaptive multi-coordinate updates and diverse target templates, achieving near-100% jailbreak success rates across language model benchmarks.
Large language models rely heavily on safety alignment to prevent the generation of harmful, dangerous, or unethical responses. However, automated adversarial jailbreak attacks can circumvent these safeguards by appending carefully optimized text suffixes to malicious prompts. Existing optimization techniques, such as the standard Greedy Coordinate Gradient approach, suffer from significant inefficiencies and often fail because their simple target templates cause models to output an initial affirmative phrase before reverting to a refusal.
The article demonstrates an improved optimization-based framework, named I-GCG, designed to systematically bypass safety safeguards across modern language models with higher efficiency and near-perfect reliability.
The researchers introduced three primary enhancements to the base gradient attack: embedding explicit harmful guidance into the target response to prevent mid-response refusals, implementing an automatic multi-token update strategy to accelerate optimization steps, and applying an easy-to-hard initialization technique that transfers learned suffixes from simpler attack categories to harder ones. The framework was evaluated across standard safety benchmarks and four major open-weight models, using a rigorous three-tier evaluation process consisting of string filtering, automated language model verification, and human review.
Experimental results show that the improved method achieved a 100% attack success rate across all four primary target models, including highly safety-tuned models where prior state-of-the-art methods achieved success rates of only 26% to 56%. Optimization speed increased substantially, reducing the necessary attack iterations on heavily aligned models from approximately 510 steps down to 55 steps. Furthermore, prompts generated using this approach demonstrated superior cross-model transferability against commercial, closed-source models, outperforming prior baselines on external evaluation suites.
These findings indicate that current safety alignment safeguards are fundamentally fragile against multi-coordinate, guidance-based gradient attacks. For organizations deploying generative artificial intelligence, relying purely on current post-training safety alignment presents severe operational and compliance risks, as determined attackers can reliably bypass these controls.
Security teams and system developers should implement defense-in-depth measures, including robust input-filtering firewalls and external response moderation layers, rather than relying solely on internal model alignment. AI safety researchers must also advance adversarial defenses specifically tailored to counter multi-coordinate gradient attacks.
The primary limitation of the study is that the core optimization methodology requires white-box access to model weights and gradients, although the generated prompts still exhibit moderate black-box transferability. Decision-makers can have high confidence in the technical vulnerability demonstrated, but should account for the fact that defensive testing was conducted primarily on 7-billion-parameter open-weight architectures.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This seminal work introduces the foundational Greedy Coordinate Gradient (GCG) adversarial suffix optimization and AdvBench benchmark that the source paper directly analyzes and enhances.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper establishes the core failure modes of LLM safety training—such as competing objectives and prefix injection vulnerabilities—that motivate the target guidance and optimization strategies in the source.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This benchmark framework standardizes the evaluation methodology and refusal targets for automated LLM red teaming that the source adopts and tests against.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). This work explores automated prompt optimization and transferability to bypass safety alignment, providing key context on automated token search baselines.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). This study introduces coordinate ascent techniques over discrete token spaces for LLM auditing, establishing the algorithmic foundations of coordinate-based adversarial prompt search.
- Book: Safety Alignment Should Be Made More Than Just a Few Tokens Deep, Xiangyu Qi et al. (2025). This research reveals why LLMs remain vulnerable to shallow suffix attacks and prefilling by showing that safety alignment is concentrated in only the first few generation tokens.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This study extends jailbreak methodology by showing how simple model-adaptive attacks and random search can bypass leading aligned LLMs even without heavy gradient optimization.
- Paper: FlipAttack: Jailbreak LLMs via Flipping, Yue Liu 0008 et al. (2025). This paper explores an alternative black-box paradigm to gradient-based jailbreaks by exploiting autoregressive left-to-right token processing through text flipping.
- Paper: FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts, Yichen Gong et al. (2025). This work expands safety guardrail circumvention from pure text gradient optimization into multimodal vision-language architectures via typographic visual prompts.
- Paper: How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models, Dun Li Chan et al. (2026). This analysis provides a mechanistic follow-up by tracing how adversarial token perturbations internally propagate through attention heads and hidden-state geometries in language models.
- Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). This paper investigates the fundamental failure of system-prompt instruction hierarchies, complementing suffix-based jailbreaks by demonstrating structural control weaknesses in LLMs.
