COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Xingang GuoFangxu YuHuan ZhangLianhui QinBin Hu

article2024ICML168 citations

Introduces an energy-based sampling framework that automates the generation of stealthy, fluent, and highly transferable adversarial LLM attacks under precise constraints like sentiment control, query paraphrasing, and contextual insertion.

Listen

As large language models become central to enterprise applications and consumer technology, evaluating their vulnerability to adversarial manipulation—known as jailbreaking—is a vital safety requirement. Existing automated white-box attack methods, which require direct access to a model's internal weights, suffer from significant practical limitations. They typically generate unnatural, garbled text strings that are easily caught by standard perplexity-based filtering defenses, and they cannot enforce user-defined constraints such as writing style, sentiment, or contextual coherence. This gap leaves organizations unable to thoroughly stress-test their models against realistic, stealthy, and adaptable adversarial prompts.

The article demonstrates and evaluates COLD-Attack, an automated framework designed to generate highly fluent, stealthy, and controllable adversarial attacks against large language models. The primary objective is to bridge adversarial red-teaming with controllable text generation, allowing safety evaluators to customize attack prompts under diverse constraints while effectively bypassing existing alignment and filtering defenses.

To achieve this, the article translates the jailbreak problem into an energy-based controllable decoding problem. The researchers adapt an established gradient-based sampling method, Langevin dynamics, to optimize continuous token representations according to custom energy functions that balance attack success, text fluency, sentiment steering, semantic similarity, and position coherence. The continuous representations are then converted back into discrete, fluent text using guided decoding. The framework was evaluated across multiple open-source models—including Llama-2, Mistral, Vicuna, and Guanaco in 7-billion and 13-billion parameter sizes—using standard harmful request benchmarks, and tested for transferability against commercial closed-source systems like GPT-3.5 and GPT-4.

The evaluation revealed several key findings. First, COLD-Attack matched or exceeded the success rates of existing methods while generating significantly more natural language; its prompts achieved low perplexity scores (between 24.8 and 33.0 across 7B models), outperforming baseline fluent attack methods. Second, it demonstrated high computational efficiency, running on average 10 times faster than the popular greedy coordinate gradient method by eliminating step-by-step discrete searches. Third, the framework successfully executed novel attack paradigms: it generated effective paraphrased attacks that preserved the original intent without appending telltale suffixes, and it inserted seamless bridge prompts between questions and output-formatting instructions while sustaining attack success rates above 80%. Fourth, attacks transferred to commercial black-box models, achieving attack success rates of up to 36% on GPT-3.5 and 46% on GPT-4. Finally, the experiments uncovered model-specific emotional vulnerabilities; for example, Mistral and Guanaco were more susceptible to negative sentiment framing, whereas Llama-2 was more easily bypassed using positive sentiment.

These findings indicate that existing AI defenses based on simple fluency checks or pattern-matching filters are insufficient against sophisticated, controllable attacks. Because COLD-Attack produces coherent, human-like prompts across arbitrary sentence positions, it significantly escalates the risk of automated influence operations, content filter evasion, and malicious prompt generation. Furthermore, the discovery that sentiment steering influences safety guardrails highlights an unexpected dimension of model vulnerability that current alignment practices overlook.

To mitigate these risks, developers and safety teams should move beyond simple input-perplexity filters and adopt multi-layered safety guardrails, such as dedicated classifier models like Llama Guard, which proved to be the most resilient defense in testing. Organizations should also incorporate controllable, diverse prompt generation into their safety fine-tuning and red-teaming pipelines. While COLD-Attack showed high efficacy across benchmarks, its effectiveness decreased when models were guarded by explicit system prompts (dropping attack success from 92% to 70% on Llama-2-7B). Further work is necessary to combine continuous logit optimization with discrete token refinement to maintain robust testing capabilities against system-prompted architectures.

Cover for COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Abstract

Jailbreaks on large language models (LLMs) have recently received increasing attention. For a comprehensive assessment of LLM safety, it is essential to consider jailbreaks with diverse attributes, such as contextual coherence and sentiment/stylistic variations, and hence it is beneficial to study controllable jailbreaking, i.e. how to enforce control on LLM attacks. In this paper, we formally formulate the controllable attack generation problem, and build a novel connection between this problem and controllable text generation, a well-explored topic of natural language processing. Based on this connection, we adapt the Energy-based Constrained Decoding with Langevin Dynamics (COLD), a state-of-the-art, highly efficient algorithm in controllable text generation, and introduce the COLD-Attack framework which unifies and automates the search of adversarial LLM attacks under a variety of control requirements such as fluency, stealthiness, sentiment, and left-right-coherence. The controllability enabled by COLD-Attack leads to diverse new jailbreak scenarios which not only cover the standard setting of generating fluent (suffix) attack with continuation constraint, but also allow us to address new controllable attack settings such as revising a user query adversarially with paraphrasing constraint, and inserting stealthy attacks in context with position constraint. Our extensive experiments on various LLMs (Llama-2, Mistral, Vicuna, Guanaco, GPT-3.5, and GPT-4) show COLD-Attack's broad applicability, strong controllability, high success rate, and attack transferability. Our code is available at https://github.com/Yu-Fangxu/COLD-Attack.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Controllability and Stealthiness for Attacks
  • 3.1. General Problem: Controllable Attack Generation
  • 3.2. Relevance to Stealthy LLM Attacks
  • 3.3. Connections with Controllable Text Generation
  • 4. COLD-Attack
  • 4.1. Energy Functions for Controllable Attacks
  • 4.2. Final Energy-based Models for Attacks
  • 5. Experimental Evaluations
  • 5.1. Results: Attack with Continuation Constraint
  • 5.2. Results: Attack with Paraphrasing Constraint
  • 5.3. Results: Attack with Position Constraint
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Additional Related Work
  • A.1. Safety Aligned LLMs
  • A.2. Jailbreak LLMs
  • A.3. Controllable Text Generation
  • B. More on COLD-Attack
  • B.1. More Details on the Energy Functions for Controllable Attacks
  • B.2. LLM-Guided Decoding Process
  • C. Experimental Details
  • C.1. Large Language Models
  • C.2. Baselines Setup
  • C.3. COLD-Attack Experimental Setup
  • C.4. Evaluation Metrics
  • D. Additional Results
  • D.1. Transferability
  • D.2. Against Defense
  • D.3. Attack on Larger LLMs
  • D.4. Ablation Study
  • D.5. Full Result on 520 Samples
  • D.6. Coherence of Prompt and Continuation
  • D.7. Comparison with Black-Box Methods
  • D.8. More Discussions on the Impact of System Prompts of LLMs
  • D.9. More Selected Examples

Knowls

  1. Knowl 1 — Controllable Attack Generation Problem Formulation

    definition

    Controllable attack generation is the problem of finding an adversarial token sequence y=(y1,y2,…,yL)y = (y_1, y_2, \dots, y_L) of length LL over vocabulary V\mathcal{V} that simultaneously causes a target large language model (LLM) to generate an unsafe or target response while satisfying a set of explicit auxiliary constraints.

    Formally, let mm be the total number of constraints, and let ci(y)∈{0,1}c_i(y) \in \{0, 1\} be an indicator function for i∈{1,…,m}i \in \{1, \dots, m\} such that ci(y)=1c_i(y) = 1 if sequence yy satisfies the ii-th constraint and ci(y)=0c_i(y) = 0 otherwise:

    \text{Find } & y \\ \text{subject to } & c_i(y) = 1, \quad \forall i \in \{1, \dots, m\} \end{aligned}$$ where: - $c_1(y)$ is the indicator function for attack success (e.g., maximizing the likelihood of target harmful content generation). - $c_2(y)$ is the indicator function enforcing fluency of the adversarial prompt or its concatenation with user queries. - $c_i(y)$ for $3 \le i \le m$ represent arbitrary additional control constraints, such as semantic paraphrasing fidelity to an original query, sentiment steering, lexical inclusion or exclusion, or position-specific placement.
  2. Knowl 2 — COLD-Attack Framework Algorithm

    algorithm

    COLD-Attack adapts Energy-based Constrained Decoding with Langevin Dynamics (COLD) to generate controllable adversarial prompts by performing gradient-based sampling over continuous logit sequences rather than discrete greedy token replacements. The optimization operates on a continuous logit sequence y~=(y~1,…,y~L)\tilde{\mathbf{y}} = (\tilde{y}_1, \dots, \tilde{y}_L), where each y~i∈R∣V∣\tilde{y}_i \in \mathbb{R}^{|\mathcal{V}|} represents unnormalized token logits over vocabulary V\mathcal{V}.

    Input: Differentiable energy functions {Ej}\{E_j\}, energy weights {λj}\{\lambda_j\}, prompt length LL, iterations NN, step size η\eta, noise schedule σn\sigma_n
    Output: Discrete adversarial prompt y=(y1,…,yL)y = (y_1, \dots, y_L)
    Initialize y~i0←init(⋅)\tilde{y}_i^0 \leftarrow \text{init}(\cdot) for all i∈{1,…,L}i \in \{1, \dots, L\}
    for n=0n = 0 to N−1N - 1 do
        E(y~n)←∑jλjEj(y~n)E(\tilde{\mathbf{y}}^n) \leftarrow \sum_j \lambda_j E_j(\tilde{\mathbf{y}}^n)
        for i=1i = 1 to LL do
            Sample ϵin∼N(0,σn2I)\epsilon_i^n \sim \mathcal{N}(0, \sigma_n^2 I)
            y~in+1←y~in−η∇y~iE(y~n)+ϵin\tilde{y}_i^{n+1} \leftarrow \tilde{y}_i^n - \eta \nabla_{\tilde{y}_i} E(\tilde{\mathbf{y}}^n) + \epsilon_i^n
        end for
    end for
    for i=1i = 1 to LL do
        yi←decode(y~iN)y_i \leftarrow \text{decode}(\tilde{y}_i^N)
    end for
    return y=(y1,…,yL)y = (y_1, \dots, y_L)

    In standard implementations, the optimization executes for N=2000N = 2000 iterations with step size η=0.1\eta = 0.1, sequence length L=20L = 20, and a decreasing noise schedule σn∈{1.0,0.5,0.1,0.05,0.01}\sigma_n \in \{1.0, 0.5, 0.1, 0.05, 0.01\} updated at iteration thresholds n∈{0,50,200,500,1500}n \in \{0, 50, 200, 500, 1500\}.

  3. Knowl 3 — Differentiable Energy Functions for Controllable Jailbreak Attacks

    equation

    COLD-Attack encodes attack objectives and auxiliary constraints as differentiable energy functions evaluated on continuous logit sequences y~=(y~1,…,y~L)\tilde{\mathbf{y}} = (\tilde{y}_1, \dots, \tilde{y}_L) with y~i∈R∣V∣\tilde{y}_i \in \mathbb{R}^{|\mathcal{V}|}:

    1. Attack Success Energy (EattE_{\text{att}}): Enforces that the target language model starts its output with a target affirmative response prefix zz conditioned on adversarial prompt yy: Eatt(y;z):=−log⁡pLM(z∣y)E_{\text{att}}(y; z) := -\log p_{\text{LM}}(z \mid y)

    2. Fluency Energy (EfluE_{\text{flu}}): Minimizes cross-entropy between the soft token distribution softmax(y~i)\text{softmax}(\tilde{y}_i) and the autoregressive next-token distribution pLM(⋅∣y<i)p_{\text{LM}}(\cdot \mid y_{<i}) of the base LLM: Eflu(y~):=−∑i=1L∑v∈VpLM(v∣y<i)log⁡softmax(y~i(v))E_{\text{flu}}(\tilde{\mathbf{y}}) := -\sum_{i=1}^L \sum_{v \in \mathcal{V}} p_{\text{LM}}(v \mid y_{<i}) \log \text{softmax}(\tilde{y}_i(v))

    3. Lexical Constraint Energy (ElexE_{\text{lex}}): Enforces inclusion of specified keywords or suppresses refusal keywords klistk_{\text{list}} via differentiable nn-gram matching approximating BLEU-nn: Elex(y~):=−ngram_match(y~,klist)E_{\text{lex}}(\tilde{\mathbf{y}}) := -\text{ngram\_match}(\tilde{\mathbf{y}}, k_{\text{list}})

    4. Semantic Similarity Energy (EsimE_{\text{sim}}): Enforces semantic proximity between a paraphrased prompt yy and an original harmful query xx using cosine distance between sequence mean embeddings: Esim(y~):=−cos⁡(1L∑i=1Le(yi),  1∣x∣∑j=1∣x∣e(xj))E_{\text{sim}}(\tilde{\mathbf{y}}) := -\cos\left(\frac{1}{L}\sum_{i=1}^L e(y_i), \; \frac{1}{|x|}\sum_{j=1}^{|x|} e(x_j)\right) where e(t)∈Rde(t) \in \mathbb{R}^d denotes the token embedding of token tt.

  4. Knowl 4 — Compositional Energy Formulations for Controllable Attack Settings

    model/method

    COLD-Attack constructs weighted linear combinations of energy functions E(y~)=∑jλjEj(y~)E(\tilde{\mathbf{y}}) = \sum_j \lambda_j E_j(\tilde{\mathbf{y}}) tailored to three distinct jailbreak paradigms:

    1. Continuation Constraint (Suffix Attack): An adversarial suffix yy is appended to an existing malicious query xx such that x⊕yx \oplus y is fluent, suppresses refusal phrases, and triggers affirmative output zz: E(y)=λ1Eatt(x⊕y;z)+λ2Eflu(x⊕y)+λ3Elex(y)E(y) = \lambda_1 E_{\text{att}}(x \oplus y; z) + \lambda_2 E_{\text{flu}}(x \oplus y) + \lambda_3 E_{\text{lex}}(y) with hyperparameters λ1=100,λ2=1,λ3=100\lambda_1 = 100, \lambda_2 = 1, \lambda_3 = 100.

    2. Paraphrasing Constraint: An original malicious query xx is entirely rewritten into a fluent attack yy that preserves semantic meaning and optionally steers sentiment via keyword list klistk_{\text{list}}: E(y)=λ1Eatt(y;z)+λ2Eflu(y)+λ3Esim(y,x)+λ4Elex(y,klist)E(y) = \lambda_1 E_{\text{att}}(y; z) + \lambda_2 E_{\text{flu}}(y) + \lambda_3 E_{\text{sim}}(y, x) + \lambda_4 E_{\text{lex}}(y, k_{\text{list}}) with hyperparameters λ1=100,λ2=1,λ3=100\lambda_1 = 100, \lambda_2 = 1, \lambda_3 = 100, and λ4=100\lambda_4 = 100 when sentiment steering is active.

    3. Position Constraint (Middle Insertion Attack): An adversarial bridge prompt yy is inserted between a user query xx and an output formatting or styling prompt pp, forming x⊕y⊕px \oplus y \oplus p to satisfy output constraints while ensuring end-to-end fluency: E(y)=λ1Eatt(x⊕y⊕p;z)+λ2Eflu(x⊕y⊕p)+λ3Elex(y)E(y) = \lambda_1 E_{\text{att}}(x \oplus y \oplus p; z) + \lambda_2 E_{\text{flu}}(x \oplus y \oplus p) + \lambda_3 E_{\text{lex}}(y) with hyperparameters λ1=100,λ2=1,λ3=100\lambda_1 = 100, \lambda_2 = 1, \lambda_3 = 100.

  5. Knowl 5 — LLM-Guided Decoding for Continuous Logit Sequences

    algorithm

    Directly applying greedy argmax over continuous logits y~N=(y~1N,…,y~LN)\tilde{\mathbf{y}}^N = (\tilde{y}_1^N, \dots, \tilde{y}_L^N) produces disfluent text due to competing gradients across compositional energy terms. COLD-Attack resolves this using an autoregressive, LLM-guided decoding procedure that filters token choices by the causal language model's top-kk distribution before selecting the energy-maximizing logit.

    Input: Optimized continuous logits y~N=(y~1N,…,y~LN)\tilde{\mathbf{y}}^N = (\tilde{y}_1^N, \dots, \tilde{y}_L^N), prefix context xx, language model pLMp_{\text{LM}}, vocabulary candidate count kk
    Output: Fluent discrete token sequence y=(y1,…,y~L)y = (y_1, \dots, \tilde{y}_L)
    for i=1i = 1 to LL do
        Compute conditional distribution pLM(⋅∣x⊕y<i)p_{\text{LM}}(\cdot \mid x \oplus y_{<i})
        Extract candidate set Vik⊂V\mathcal{V}_i^k \subset \mathcal{V} containing the top-kk most probable tokens under pLM(⋅∣x⊕y<i)p_{\text{LM}}(\cdot \mid x \oplus y_{<i})
        yi←arg⁡max⁡v∈Viky~iN(v)y_i \leftarrow \arg\max_{v \in \mathcal{V}_i^k} \tilde{y}_i^N(v)
    end for
    return y=(y1,…,yL)y = (y_1, \dots, y_L)

    This constrained selection forces each decoded token to remain within the high-probability manifold of the base LLM while retaining the steering direction of the energy function.

  6. Knowl 6 — Adversarial Suffix Attack Performance and Perplexity Comparison

    data/table

    Evaluated on a 50-instruction subset of AdvBench across Vicuna-7B-v1.5, Guanaco-7B-HF, Mistral-7B-Instruct-v0.2, and Llama-2-7B-Chat-hf, COLD-Attack produces competitive Attack Success Rates (ASR, substring match; ASR-G, GPT-4 evaluated) while achieving substantially lower Perplexity (PPL, evaluated on Vicuna-7B; lower indicates higher fluency) than prior gradient-based discrete optimization methods.

    Methods Vicuna Guanaco Mistral Llama2
    ASR ASR-G PPL ASR ASR-G PPL ASR ASR-G PPL ASR ASR-G PPL
    Prompt-only 48.00 30.00 – 44.00 26.00 – 6.00 4.00 – 4.00 4.00 –
    PEZ 28.00 6.00 5408 52.00 22.00 15127 16.00 6.00 3470.22 18.00 8.00 7307
    GBDA 20.00 8.00 13932 44.00 12.00 18220 42.00 18.00 3855.66 10.00 8.00 14758
    UAT 58.00 10.00 8487 52.00 20.00 9725 66.00 24.00 4094.97 24.00 20.00 8962
    GCG 100.00 92.00 821.53 100.00 84.00 406.81 100.00 42.00 814.37 90.00 68.00 5740
    GCG-reg 100.00 70.00 77.84 100.00 68.00 51.02 100.00 32.00 122.57 82.00 28.00 1142
    AutoDAN-Zhu 90.00 84.00 33.43 100.00 80.00 50.47 92.00 84.00 79.53 92.00 68.00 152.32
    COLD-Attack 100.00 86.00 32.96 96.00 84.00 30.55 92.00 90.00 26.24 92.00 66.00 24.83

    COLD-Attack reduces perplexity by more than an order of magnitude relative to GCG (e.g., PPL 24.83 vs. 5740 on Llama2). Additionally, because it samples continuous logits without discrete batch candidate search at each step, COLD-Attack executes approximately 10×10\times faster than GCG (taking 15.05 minutes per sample on Llama2 on a single NVIDIA V100 GPU compared to 235.25 minutes for GCG).

  7. Knowl 7 — Adversarial Paraphrasing and Sentiment Steering Results

    data/table

    Under the paraphrasing constraint, COLD-Attack reformulates malicious instructions directly without appending arbitrary suffixes, outperforming standard text paraphrasers on 50 AdvBench queries while preserving semantic similarity (BERTScore ∼0.72\sim 0.72).

    Methods Metric Vicuna Guanaco Mistral Llama2
    COLD-Attack BLEU ↑\uparrow 0.52 0.47 0.41 0.60
    ROUGE ↑\uparrow 0.57 0.55 0.55 0.54
    BERTScore ↑\uparrow 0.72 0.74 0.72 0.71
    PPL ↓\downarrow 31.11 29.23 37.21 39.26
    ASR ↑\uparrow 96.00 98.00 98.00 86.00
    ASR-G ↑\uparrow 80.00 78.00 90.00 74.00
    PRISM ASR / ASR-G 52.00 / 36.00 58.00 / 22.00 18.00 / 6.00 4.00 / 2.00
    PAWS ASR / ASR-G 56.00 / 24.00 56.00 / 24.00 24.00 / 8.00 6.00 / 2.00
    GPT-4 ASR / ASR-G 40.00 / 22.00 42.00 / 24.00 10.00 / 6.00 4.00 / 4.00

    When applying sentiment steering via lexical constraints:

    • Mistral and Guanaco show higher vulnerability to negative sentiment attacks (Mistral ASR-G increases by +30%+30\% to 90.00%90.00\%; Guanaco ASR-G increases by +14%+14\% to 80.00%80.00\%).
    • Llama2 exhibits higher vulnerability to positive sentiment attacks, with ASR-G increasing by +18%+18\% (58.00%58.00\% vs. 40.00%40.00\% for negative sentiment).
  8. Knowl 8 — Position-Constrained Jailbreak Performance Under Output Controls

    data/table

    In position-constrained attacks, COLD-Attack inserts an adversarial prompt yy between a user query xx and an output control instruction pp (x⊕y⊕px \oplus y \oplus p), controlling both jailbreak success and target output style/format. Evaluated on Llama-2-7B-Chat-hf across four output constraints (Sentiment, Lexical, Format, and Style):

    Constraint Method ASR ↑\uparrow ASR-G ↑\uparrow PPL ↓\downarrow
    Sentiment Prompt Only 26.00 22.00 –
    COLD-Attack 80.00 88.00 59.53
    AutoDAN-Zhu 94.00 72.00 113.27
    GCG 62.00 52.00 2587.90
    Lexical Prompt Only 24.00 24.00 –
    COLD-Attack 88.00 86.00 68.23
    AutoDAN-Zhu 84.00 68.00 176.86
    GCG 64.00 50.00 2684.62
    Format Prompt Only 10.00 8.00 –
    COLD-Attack 80.00 86.00 57.70
    AutoDAN-Zhu 84.00 74.00 124.38
    GCG 44.00 44.00 2431.87
    Style Prompt Only 10.00 6.00 –
    COLD-Attack 80.00 80.00 58.93
    AutoDAN-Zhu 92.00 66.00 149.43
    GCG 54.00 42.00 1830.72

    COLD-Attack achieves ASR-G between 80.00%80.00\% and 88.00%88.00\% across all output control constraints while maintaining perplexity values (extPPL≈57−68 ext{PPL} \approx 57 - 68) that are roughly 2×2\times lower than AutoDAN-Zhu and 40×40\times lower than GCG.

  9. Knowl 9 — Transferability to Proprietary Language Models (GPT-3.5 and GPT-4)

    data/table

    Adversarial prompts crafted on open-source white-box surrogates (Guanaco-7B, Mistral-7B, Llama-2-7B, and Vicuna-7B) transfer successfully to proprietary closed-source models GPT-3.5 turbo and GPT-4.

    Target: GPT-3.5 Guanaco Mistral Llama2 Vicuna
    Method ASR ASR-G ASR ASR-G ASR ASR-G ASR ASR-G
    Prompt-only 2.00 2.00 2.00 2.00 2.00 2.00 2.00 2.00
    COLD-Attack 28.00 26.00 36.00 32.00 30.00 30.00 18.00 16.00
    AutoDAN-Zhu 26.00 18.00 30.00 26.00 30.00 12.00 62.00 34.00
    GCG 12.00 10.00 16.00 10.00 14.00 10.00 18.00 16.00
    Target: GPT-4 Guanaco Mistral Llama2 Vicuna
    Method ASR ASR-G ASR ASR-G ASR ASR-G ASR ASR-G
    Prompt-only 6.00 6.00 6.00 6.00 6.00 6.00 6.00 6.00
    COLD-Attack 36.00 34.00 36.00 30.00 46.00 32.00 40.00 36.00
    AutoDAN-Zhu 64.00 30.00 30.00 24.00 34.00 24.00 62.00 34.00
    GCG 20.00 16.00 22.00 20.00 36.00 26.00 20.00 20.00

    Among fully automated white-box methods (excluding methods that initialize from human-engineered jailbreak templates), COLD-Attack achieves higher transfer ASR-G than AutoDAN-Zhu across Guanaco, Mistral, and Llama2 on GPT-3.5, and reaches up to 36.00%36.00\% ASR-G on GPT-4.

  10. Knowl 10 — Robustness of COLD-Attack Against Automated Defenses

    data/table

    COLD-Attack was evaluated under the suffix attack setting against multiple defense countermeasures on 50 AdvBench queries across Vicuna, Guanaco, Mistral, and Llama2.

    Defense Method Vicuna Guanaco Mistral Llama2
    No defense 100.00 96.00 92.00 92.00
    OpenAI Moderation 86.00 90.00 90.00 90.00
    SmoothLLM 76.00 60.00 56.00 66.00
    RAIN 94.00 88.00 80.00 56.00
    Llama Guard 42.00 38.00 32.00 40.00

    Key behavioral findings include:

    1. Perplexity Filtering: Because COLD-Attack maintains low PPL (24.83−32.9624.83 - 32.96), the majority of adversarial prompts bypass strict perplexity thresholds (e.g., threshold of 60).
    2. SmoothLLM Perturbation: Semantically coherent adversarial sentences retain effectiveness under character-level perturbations where brittle gibberish attacks fail.
    3. Llama Guard Evaluation: Under Llama Guard (the most restrictive defense), COLD-Attack achieves 32.00%−42.00%32.00\% - 42.00\% bypass rates while maintaining low perplexity, whereas GCG achieves lower bypass rates (20.00%−34.00%20.00\% - 34.00\%) with severe perplexity degradation (PPL 5740).
  11. Knowl 11 — Impact of System Prompts and Soft-to-Hard Loss Discrepancy

    limitation

    When standard safety system prompts are prepended to the target models during generation, the attack success rate of COLD-Attack drops noticeably (e.g., ASR-G drops from 90.00%90.00\% to 64.00%64.00\% on Mistral-7B, and from 66.00%66.00\% to 38.00%38.00\% on Llama-2-7B-Chat-hf).

    Two factors account for this limitation:

    1. Multi-Objective Loss Balance: COLD-Attack optimizes a linear combination of energy functions (Eatt,Eflu,ElexE_{\text{att}}, E_{\text{flu}}, E_{\text{lex}}). Consequently, it does not minimize the attack loss EattE_{\text{att}} as aggressively as single-objective discrete search methods such as GCG.
    2. Continuous-to-Discrete Optimization Gap: COLD-Attack updates continuous soft logits y~\tilde{\mathbf{y}} with low-temperature softmax (T=0.001T = 0.001). A decrease in continuous soft-prompt loss does not guarantee an equivalent decrease in the discrete hard-prompt loss evaluated on the final decoded token string. When strict system prompts elevate the safety barrier, discrete token-replacement strategies like GCG drive the hard-prompt loss closer to zero than continuous Langevin dynamics with guided decoding.

Coverage note — Deliberately omitted were illustrative qualitative prompt transcripts from Appendix D.9 and secondary hyperparameter sensitivity sweeps from Appendix D.4 to focus on core mathematical formulations, algorithm workflows, decoding mechanics, benchmark comparative tables, and structural limitations.

References

  1. 1.Abdelnabi, S., Greshake, K., Mishra, S., Endres, C., Holz, T., and Fritz, M. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, pp. 79–90, 2023.
  2. 2.Anonymous. Curiosity-driven red-teaming for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=4KqkizXgXU.
  3. 3.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  4. 4.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  5. 5.Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al. Are aligned neural networks adversarially aligned? arXiv preprint arXiv:2306.15447, 2023.
  6. 6.Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023.
  7. 7.Chen, H., Li, H., Chen, D., and Narasimhan, K. Controllable text generation with language constraints. arXiv preprint arXiv:2212.10466, 2022.
  8. 8.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  9. 9.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  10. 10.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  11. 11.DAN. Chatgpt "DAN" (and other "jailbreaks"). https://gist.github.com/coolaj86/6f4f7b30129b0251f61fa7baaa881516, 2023.
  12. 12.Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164, 2019.
  13. 13.Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., Li, Z., Wang, H., Zhang, T., and Liu, Y. Masterkey: Automated jailbreak across multiple large language model chatbots, 2023.
  14. 14.Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  15. 15.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023.
  16. 16.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  17. 17.Du, Y., Zhao, S., Ma, M., Chen, Y., and Qin, B. Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak. arXiv preprint arXiv:2312.04127, 2023.
  18. 18.Ebrahimi, J., Rao, A., Lowd, D., and Dou, D. Hotflip: White-box adversarial examples for text classification. arXiv preprint arXiv:1712.06751, 2017.
  19. 19.Forristal, J., Mireshghallah, F., Durrett, G., and Berg-Kirkpatrick, T. A block metropolis-hastings sampler for controllable energy-based text generation. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pp. 403–413, 2023.
  20. 20.Glaese, A., McAleese, N., Tr˛ebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  21. 21.Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., and Sedova, K. Generative language models and automated influence operations: Emerging threats and potential mitigations. arXiv preprint arXiv:2301.04246, 2023.
  22. 22.Gong, Y., Ran, D., Liu, J., Wang, C., Cong, T., Wang, A., Duan, S., and Wang, X. Figstep: Jailbreaking large vision-language models via typographic visual prompts. arXiv preprint arXiv:2311.05608, 2023.
  23. 23.Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D. Gradient-based adversarial attacks against text transformers. arXiv preprint arXiv:2104.13733, 2021.
  24. 24.Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D. Catastrophic jailbreak of open-source llms via exploiting generation. arXiv preprint arXiv:2310.06987, 2023.
  25. 25.Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023.
  26. 26.Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614, 2023.
  27. 27.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  28. 28.Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J. Automatically auditing large language models via discrete optimization. arXiv preprint arXiv:2303.04381, 2023.
  29. 29.Kandpal, N., Jagielski, M., Tramèr, F., and Carlini, N. Backdoor attacks for in-context learning with language models. arXiv preprint arXiv:2307.14692, 2023.
  30. 30.Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. arXiv preprint arXiv:2302.05733, 2023.
  31. 31.Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E. Pretraining language models with human preferences. In International Conference on Machine Learning, pp. 17506–17533. PMLR, 2023.
  32. 32.Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020.
  33. 33.Lapid, R., Langberg, R., and Sipper, M. Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446, 2023.
  34. 34.Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  35. 35.Li, H., Guo, D., Fan, W., Xu, M., and Song, Y. Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197, 2023a.
  36. 36.Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055, 2015.
  37. 37.Li, X., Thickstun, J., Gulrajani, I., Liang, P. S., and Hashimoto, T. B. Diffusion-lm improves controllable text generation. Advances in Neural Information Processing Systems, 35:4328–4343, 2022.
  38. 38.Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., and Han, B. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191, 2023b.
  39. 39.Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H. Rain: Your language models can align themselves without finetuning. arXiv preprint arXiv:2309.07124, 2023c.
  40. 40.Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y. The unlocking spell on base llms: Rethinking alignment via in-context learning. ArXiv preprint, 2023.
  41. 41.Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y. Dexperts: Decoding-time controlled text generation with experts and anti-experts. arXiv preprint arXiv:2105.03023, 2021a.
  42. 42.Liu, G., Yang, Z., Tao, T., Liang, X., Bao, J., Li, Z., He, X., Cui, S., and Hu, Z. Don’t take it literally: An edit-invariant sequence loss for text generation. arXiv preprint arXiv:2106.15078, 2021b.
  43. 43.Liu, G., Feng, Z., Gao, Y., Yang, Z., Liang, X., Bao, J., He, X., Cui, S., Li, Z., and Hu, Z. Composable text controls in latent space with odes. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 16543–16570, 2023a.
  44. 44.Liu, J., Cohen, A., Pasunuru, R., Choi, Y., Hajishirzi, H., and Celikyilmaz, A. Making ppo even better: Value-guided monte-carlo tree search decoding. arXiv preprint arXiv:2309.15028, 2023b.
  45. 45.Liu, X., Khalifa, M., and Wang, L. Bolt: Fast energy-based controlled text generation with tunable biases. arXiv preprint arXiv:2305.12018, 2023c.
  46. 46.Liu, X., Xu, N., Chen, M., and Xiao, C. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451, 2023d.
  47. 47.Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860, 2023e.
  48. 48.Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L., and Camacho-collados, J. TimeLMs: Diachronic language models from Twitter. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 251–260, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-demo.25. URL https://aclanthology.org/2022.acl-demo.25.
  49. 49.Lu, X., West, P., Zellers, R., Bras, R. L., Bhagavatula, C., and Choi, Y. Neurologic decoding:(un) supervised neural text generation with predicate logic constraints. arXiv preprint arXiv:2010.12884, 2020.
  50. 50.Lu, X., Welleck, S., West, P., Jiang, L., Kasai, J., Khashabi, D., Bras, R. L., Qin, L., Yu, Y., Zellers, R., et al. Neurologic a* esque decoding: Constrained text generation with lookahead heuristics. arXiv preprint arXiv:2112.08726, 2021.
  51. 51.Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. Quark: Controllable text generation with reinforced unlearning. Advances in neural information processing systems, 35:27591–27609, 2022.
  52. 52.Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. 2024.
  53. 53.Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119, 2023.
  54. 54.Mireshghallah, F., Goyal, K., and Berg-Kirkpatrick, T. Mix and match: Learning-free controllable text generation using energy language models. arXiv preprint arXiv:2203.13299, 2022.
  55. 55.Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al. Controlled decoding from language models. arXiv preprint arXiv:2310.17022, 2023.
  56. 56.OpenAI. https://platform.openai.com/docs/guides/moderation.
  57. 57.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  58. 58.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  59. 59.Perez, F. and Ribeiro, I. Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527, 2022.
  60. 60.Post, M. and Vilar, D. Fast lexically constrained decoding with dynamic beam allocation for neural machine translation. arXiv preprint arXiv:1804.06609, 2018.
  61. 61.Qi, X., Huang, K., Panda, A., Wang, M., and Mittal, P. Visual adversarial examples jailbreak aligned large language models. In The Second Workshop on New Frontiers in Adversarial Machine Learning, volume 1, 2023.
  62. 62.Qiang, Y., Zhou, X., and Zhu, D. Hijacking large language models via adversarial in-context learning. arXiv preprint arXiv:2311.09948, 2023.
  63. 63.Qin, L., Shwartz, V., West, P., Bhagavatula, C., Hwang, J., Bras, R. L., Bosselut, A., and Choi, Y. Back to the future: Unsupervised backprop-based decoding for counterfactual and abductive commonsense reasoning. arXiv preprint arXiv:2010.05906, 2020.
  64. 64.Qin, L., Welleck, S., Khashabi, D., and Choi, Y. Cold decoding: Energy-based constrained text generation with langevin dynamics. Advances in Neural Information Processing Systems, 35:9538–9551, 2022.
  65. 65.Rando, J. and Tramèr, F. Universal jailbreak backdoors from poisoned human feedback. arXiv preprint arXiv:2311.14455, 2023.
  66. 66.Robey, A., Wong, E., Hassani, H., and Pappas, G. J. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684, 2023.
  67. 67.Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al. Scalable and transferable black-box jailbreaks for language models via persona modulation. arXiv preprint arXiv:2311.03348, 2023.
  68. 68.Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825, 2023.
  69. 69.Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  70. 70.shu, D., Jin, M., Zhu, S., Wang, B., Zhou, Z., Zhang, C., and Zhang, Y. Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models, 2024.
  71. 71.Tevet, G. and Berant, J. Evaluating the evaluation of diversity in natural language generation. arXiv preprint arXiv:2004.02990, 2020.
  72. 72.Thompson, B. and Post, M. Automatic machine translation evaluation in many languages via zero-shot paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, November 2020a. Association for Computational Linguistics.
  73. 73.Thompson, B. and Post, M. Paraphrase generation as zero-shot multilingual translation: Disentangling semantic similarity from lexical and syntactic diversity. In Proceedings of the Fifth Conference on Machine Translation (Volume 1: Research Papers), Online, November 2020b. Association for Computational Linguistics.
  74. 74.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  75. 75.Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C. How many unicorns are in this image? a safety evaluation benchmark for vision llms. arXiv preprint arXiv:2311.16101, 2023.
  76. 76.Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S. Universal adversarial triggers for attacking and analyzing nlp. arXiv preprint arXiv:1908.07125, 2019.
  77. 77.Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does LLM safety training fail? arXiv preprint arXiv:2307.02483, 2023a.
  78. 78.Wei, Z., Wang, Y., and Wang, Y. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387, 2023b.
  79. 79.Welling, M. and Teh, Y. W. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th international conference on machine learning (ICML-11), pp. 681–688, 2011.
  80. 80.Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. arXiv preprint arXiv:2302.03668, 2023.
  81. 81.Wichers, N., Denison, C., and Beirami, A. Gradient-based language model red teaming. arXiv preprint arXiv:2401.16656, 2024.
  82. 82.WitchBOT. You can use gpt-4 to create prompt injections against gpt-4. https://www.lesswrong.com/posts/bNCDexejSZpkuu3yz/you-can-use-gpt-4-to-createprompt-injections-against-gpt-4, 2023.
  83. 83.Yang, K. and Klein, D. Fudge: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218, 2021.
  84. 84.Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D. Shadow alignment: The ease of subverting safely-aligned language models, 2023.
  85. 85.Yip, D. W., Esmradi, A., and Chan, C. F. A novel evaluation framework for assessing resilience against prompt injection attacks in large language models. arXiv preprint arXiv:2401.00991, 2024.
  86. 86.Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms, 2024.
  87. 87.Zhang, Y. and Ippolito, D. Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success. arXiv preprint arXiv:2307.06865, 2023.
  88. 88.Zhang, Y., Baldridge, J., and He, L. Paws: Paraphrase adversaries from word scrambling. arXiv preprint arXiv:1904.01130, 2019.
  89. 89.Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.-X., and Wang, W. Y. Weak-to-strong jailbreaking on large language models, 2024.
  90. 90.Zhou, W., Jiang, Y. E., Wilcox, E., Cotterell, R., and Sachan, M. Controlled text generation with natural language instructions. arXiv preprint arXiv:2304.14293, 2023.
  91. 91.Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T. Autodan: Automatic and interpretable adversarial attacks on large language models. arXiv preprint arXiv:2310.15140, 2023.
  92. 92.Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y. Texygen: A benchmarking platform for text generation models. In The 41st international ACM SIGIR conference on research & development in information retrieval, pp. 1097–1100, 2018.
  93. 93.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Guo, X., et al. “COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability”. arXiv, 2024, http://arxiv.org/abs/2402.08679v2.
APA
Guo, X., Yu, F., Zhang, H., Qin, L., & Hu, B. (2024). COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability. arXiv. http://arxiv.org/abs/2402.08679v2
Chicago
Guo, X., F. Yu, H. Zhang, L. Qin, and B. Hu. 2024. “COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability”. arXiv. http://arxiv.org/abs/2402.08679v2.
Harvard
Guo, X. et al. (2024) “COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.08679v2.
Vancouver
1. Guo X, Yu F, Zhang H, Qin L, Hu B (2024) COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability. arXiv

BibTeX

@article{guo2024cold,
  title = {COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability},
  author = {Guo, Xingang and Yu, Fangxu and Zhang, Huan and Qin, Lianhui and Hu, Bin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.08679v2},
  eprint = {2402.08679}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/