COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Xingang GuoFangxu YuHuan ZhangLianhui QinBin Hu
Introduces an energy-based sampling framework that automates the generation of stealthy, fluent, and highly transferable adversarial LLM attacks under precise constraints like sentiment control, query paraphrasing, and contextual insertion.
As large language models become central to enterprise applications and consumer technology, evaluating their vulnerability to adversarial manipulation—known as jailbreaking—is a vital safety requirement. Existing automated white-box attack methods, which require direct access to a model's internal weights, suffer from significant practical limitations. They typically generate unnatural, garbled text strings that are easily caught by standard perplexity-based filtering defenses, and they cannot enforce user-defined constraints such as writing style, sentiment, or contextual coherence. This gap leaves organizations unable to thoroughly stress-test their models against realistic, stealthy, and adaptable adversarial prompts.
The article demonstrates and evaluates COLD-Attack, an automated framework designed to generate highly fluent, stealthy, and controllable adversarial attacks against large language models. The primary objective is to bridge adversarial red-teaming with controllable text generation, allowing safety evaluators to customize attack prompts under diverse constraints while effectively bypassing existing alignment and filtering defenses.
To achieve this, the article translates the jailbreak problem into an energy-based controllable decoding problem. The researchers adapt an established gradient-based sampling method, Langevin dynamics, to optimize continuous token representations according to custom energy functions that balance attack success, text fluency, sentiment steering, semantic similarity, and position coherence. The continuous representations are then converted back into discrete, fluent text using guided decoding. The framework was evaluated across multiple open-source models—including Llama-2, Mistral, Vicuna, and Guanaco in 7-billion and 13-billion parameter sizes—using standard harmful request benchmarks, and tested for transferability against commercial closed-source systems like GPT-3.5 and GPT-4.
The evaluation revealed several key findings. First, COLD-Attack matched or exceeded the success rates of existing methods while generating significantly more natural language; its prompts achieved low perplexity scores (between 24.8 and 33.0 across 7B models), outperforming baseline fluent attack methods. Second, it demonstrated high computational efficiency, running on average 10 times faster than the popular greedy coordinate gradient method by eliminating step-by-step discrete searches. Third, the framework successfully executed novel attack paradigms: it generated effective paraphrased attacks that preserved the original intent without appending telltale suffixes, and it inserted seamless bridge prompts between questions and output-formatting instructions while sustaining attack success rates above 80%. Fourth, attacks transferred to commercial black-box models, achieving attack success rates of up to 36% on GPT-3.5 and 46% on GPT-4. Finally, the experiments uncovered model-specific emotional vulnerabilities; for example, Mistral and Guanaco were more susceptible to negative sentiment framing, whereas Llama-2 was more easily bypassed using positive sentiment.
These findings indicate that existing AI defenses based on simple fluency checks or pattern-matching filters are insufficient against sophisticated, controllable attacks. Because COLD-Attack produces coherent, human-like prompts across arbitrary sentence positions, it significantly escalates the risk of automated influence operations, content filter evasion, and malicious prompt generation. Furthermore, the discovery that sentiment steering influences safety guardrails highlights an unexpected dimension of model vulnerability that current alignment practices overlook.
To mitigate these risks, developers and safety teams should move beyond simple input-perplexity filters and adopt multi-layered safety guardrails, such as dedicated classifier models like Llama Guard, which proved to be the most resilient defense in testing. Organizations should also incorporate controllable, diverse prompt generation into their safety fine-tuning and red-teaming pipelines. While COLD-Attack showed high efficacy across benchmarks, its effectiveness decreased when models were guarded by explicit system prompts (dropping attack success from 92% to 70% on Llama-2-7B). Further work is necessary to combine continuous logit optimization with discrete token refinement to maintain robust testing capabilities against system-prompted architectures.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This seminal work establishes the gradient-based discrete optimization framework (GCG) and the AdvBench benchmark for adversarial suffix jailbreaks upon which COLD-Attack directly builds and extends with continuous decoding.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). AutoDAN introduces the core problem of generating fluent, stealthy, and readable jailbreak prompts to bypass perplexity filters, which COLD-Attack formalizes and solves via controllable text generation.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). This paper establishes the mathematical formulation of auditing and attacking language models via discrete coordinate optimization, providing essential foundations for controllable adversarial decoding.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This study analyzes the fundamental failure modes of LLM safety alignment, motivating the need for diverse and controllable multi-attribute attack vectors.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). PAIR demonstrates automated semantic jailbreak search to overcome non-fluent token-level attacks, offering important baseline context for controllable adversarial prompting.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). This foundational paper presents principles of controllable text generation that COLD-Attack repurposes and adapts into an adversarial search framework.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). This work advances optimization-based jailbreak strategies by addressing mid-response refusal recovery and token update efficiency beyond standard suffix and controllable decoding attacks.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This paper investigates lightweight, adaptive jailbreak methods that exploit target-specific probability and prefilling weaknesses across safety-aligned models.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench provides a standardized multi-model benchmarking suite that systematically evaluates advanced automated jailbreak attacks like COLD-Attack and tests robust dynamic refusal defenses.
- Paper: FlipAttack: Jailbreak LLMs via Flipping, Yue Liu 0008 et al. (2025). FlipAttack builds upon the exploration of text structure transformations to bypass guardrails, utilizing directional flipping rather than energy-based continuous Langevin dynamics.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). Rainbow Teaming broadens automated diverse attack generation using open-ended quality-diversity search across risk categories and styles to systematically stress-test aligned models.
