Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Bochuan CaoYuanpu CaoLu LinJinghui Chen

article2024ACL244 citations

Proposes a plug-and-play defense mechanism that drops random tokens from input requests to invalidate jailbreak prompts, reducing attack success rates from nearly 100% to around 10% without retraining the target language model.

Listen

Large Language Models are increasingly integrated into critical commercial workflows, but they remain vulnerable to alignment-breaking attacks. In these exploits, adversaries bypass safety guardrails by appending crafted prompts or role-play scenarios to elicit harmful and toxic content. Existing defenses often depend on external discriminator models that are computationally expensive, prone to misclassifying benign user prompts, and narrow in scope.

The article demonstrates a defense mechanism termed Robustly Aligned Large Language Model (RA-LLM), designed to fortify existing safety safeguards without requiring model retraining or fine-tuning.

The approach operates on the insight that adversarial prompts are fragile to input disruptions, whereas underlying safety safeguards are robust. The system generates multiple sampled variations of an incoming prompt by randomly dropping a subset of its tokens and checking if the model’s internal alignment triggers a refusal. A prompt is accepted only if the majority of these sampled variations pass without activating safety refusals. The researchers validated the method using mathematical proofs alongside empirical evaluations on several open-source and commercial language models against state-of-the-art automated attacks and popular handcrafted exploits.

The findings show that the proposed defense reduces attack success rates from 80–99% down to 6–12% across varied model architectures. Benign request handling remains virtually unaffected, maintaining answering rates between 92% and 99.3%. In direct comparisons, alternative defenses like perplexity checks completely failed against handcrafted attacks, whereas RA-LLM defended against them reliably. In addition, implementation optimizations—such as limiting Monte Carlo generation lengths and implementing early exit rules—kept additional processing time below 20% relative to standard inference.

These results provide a low-overhead, plug-and-play defense strategy for organizations deploying language models. Organizations can significantly mitigate safety, legal, and compliance risks without undertaking expensive fine-tuning cycles or relying on fragile external filtering models.

For practical deployment, teams should implement this random-dropping verification layer with tunable thresholds calibrated to their specific risk tolerance, balancing security against user friction. Future work should focus on testing performance against extreme prompt lengths and refining alignment training to make safety refusals more distinct from benign clarification requests.

arXiv: 2309.14348
Cover for Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Abstract

Recently, Large Language Models (LLMs) have made significant advancements and are now widely used across various domains. Unfortunately, there has been a rising concern that LLMs can be misused to generate harmful or malicious content. Though a line of research has focused on aligning LLMs with human values and preventing them from producing inappropriate content, such alignments are usually vulnerable and can be bypassed by alignment-breaking attacks via adversarially optimized or handcrafted jailbreaking prompts. In this work, we introduce a Robustly Aligned LLM (RA-LLM) to defend against potential alignment-breaking attacks. RA-LLM can be directly constructed upon an existing aligned LLM with a robust alignment checking function, without requiring any expensive retraining or fine-tuning process of the original LLM. Furthermore, we also provide a theoretical analysis for RA-LLM to verify its effectiveness in defending against alignment-breaking attacks. Through real-world experiments on open-source large language models, we demonstrate that RA-LLM can successfully defend against both state-of-the-art adversarial prompts and popular handcrafted jailbreaking prompts by reducing their attack success rates from nearly 100% to around 10% or less.

WARNING: This paper contains unsafe model responses. Reader discretion is advised.

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORKS
  • 3 Our Proposed Method
  • 3.1 Threat Model
  • 3.2 Our Proposed Method
  • 3.3 Practical Designs
  • 3.4 Theoretical Analysis
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Experimental Results
  • 4.3 Handcrafted Jailbreak Prompts
  • 4.4 Ablation Study
  • 5 Computational Cost
  • 6 Adaptive Attack
  • 7 Conclusion
  • 8 Limitations
  • 9 Acknowledgement
  • References
  • A Proof of Theorem 3.1
  • B Concrete Examples
  • C Defensive Efficacy Against Harmful Strings Attack
  • D Defensive Efficacy Against AutoDAN and TAP
  • E Details of Experiment
  • F Comparison with LLM Self Defense
  • G Comparison with Perplexity-Based Defense
  • H Computational Cost
  • H.1 Time Cost
  • H.2 API Cost
  • I Collaborating with Safety Alignment on LLMs to Counteract Attacks

Knowls

  1. Knowl 1 — Robust Alignment Check and RA-LLM Formulation

    model/method

    The Robustly Aligned Large Language Model (RA-LLM) defends aligned language models against alignment-breaking (jailbreak) attacks without requiring model retraining or parameter fine-tuning. Let f(⋅)f(\cdot) denote a pre-trained and safety-aligned large language model, and let AC(⋅)AC(\cdot) be an alignment check function that evaluates model outputs, returning Fail\text{Fail} if typical refusal/safety text is identified (e.g., matching standard refusal prefixes such as "I cannot" or "I'm sorry") and Pass\text{Pass} otherwise.

    Because an adversarial prompt padvp_{\text{adv}} appended to, prepended to, or inserted within a malicious query xx (xadv=x⊕padvx_{\text{adv}} = x \oplus p_{\text{adv}}) can bypass AC(f(x))AC(f(x)), RA-LLM defines a Robust Alignment Check function RAC(x)RAC(x) that subjects the input to stochastic token-dropping perturbations:

    RAC(x)={Fail,if AC(f(x))=FailFail,if Pr∼U(p)(AC(f([x]r))=Fail)>tPass,otherwiseRAC(x) = \begin{cases} \text{Fail}, & \text{if } AC(f(x)) = \text{Fail} \\ \text{Fail}, & \text{if } \mathbb{P}_{r \sim \mathcal{U}(p)}\left(AC(f([x]_r)) = \text{Fail}\right) > t \\ \text{Pass}, & \text{otherwise} \end{cases}

    where p∈[0,1)p \in [0, 1) is the token dropping ratio, rr represents a randomly sampled mask of indices indicating preserved tokens sampled uniformly without replacement from distribution U(p)\mathcal{U}(p), [x]r[x]_r denotes the sub-sequence of (1−p)L(1 - p)L retained tokens from an input xx of length LL, and t∈[0,1)t \in [0, 1) is a failure threshold.

    The robustly aligned model frob(x)f_{\text{rob}}(x) then executes:

    frob(x)={Reject to answer,if RAC(x)=Failf(x),if RAC(x)=Passf_{\text{rob}}(x) = \begin{cases} \text{Reject to answer}, & \text{if } RAC(x) = \text{Fail} \\ f(x), & \text{if } RAC(x) = \text{Pass} \end{cases}

    This mechanism exploits the fact that adversarial jailbreak triggers are brittle under random token deletion, whereas benign queries generally retain sufficient context to avoid triggering safety refusals.

  2. Knowl 2 — RA-LLM Verification and Monte Carlo Defense Algorithm

    algorithm

    To approximate the true failure expectation Pr∼U(p)(AC(f([x]r))=Fail)\mathbb{P}_{r \sim \mathcal{U}(p)}\left(AC(f([x]_r)) = \text{Fail}\right) tractably, RA-LLM evaluates nn Monte Carlo sub-samples of the input. Computational efficiency is preserved by two optimizations: (1) generating only a truncated prefix of tokens (tmax⁡=10t_{\max} = 10) for each perturbed sample to detect refusal markers without completing full generations, and (2) an early exit mechanism that terminates execution immediately if the number of detected refusals reaches ⌈n⋅t⌉\lceil n \cdot t \rceil or if the remaining trials cannot reach the threshold.

    Input: Aligned language model ff, alignment check function ACAC, input query xx, drop ratio pp, trial count nn, threshold tt
    Output: Response string f(x)f(x) or refusal
    if AC(f(x))=FailAC(f(x)) = \text{Fail} then
        return Reject the request
    else
        failures ←0\leftarrow 0
        for i=1,2,…,ni = 1, 2, \dots, n do
            Sample index mask ri∼U(p)r_i \sim \mathcal{U}(p)
            Generate prefix yi←f([x]ri)y_i \leftarrow f([x]_{r_i}) up to length tmax⁡t_{\max}
            if AC(yi)=FailAC(y_i) = \text{Fail} then
                failures ←\leftarrow failures + 1
            end if
            if failures >n⋅t> n \cdot t then
                return Reject the request
            end if
            if failures + (n−i)≤n⋅t(n - i) \le n \cdot t then
                return f(x)f(x)
            end if
        end for
        if failures /n>t/ n > t then
            return Reject the request
        else
            return f(x)f(x)
        end if
    end if

    Default hyperparameters evaluated in the paper are n=20n = 20, p=0.3p = 0.3, t=0.2t = 0.2, and tmax⁡=10t_{\max} = 10.

  3. Knowl 3 — Theoretical Rejection Guarantee for Adversarial Prompt Insertions

    theoretical result

    Let xx be a malicious prompt composed of NN tokens, and let padvp_{\text{adv}} be an adversarial prompt of length MM tokens inserted into position j∈{0,1,…,N}j \in \{0, 1, \dots, N\} within xx, forming the adversarial input xadv=x⊕padvx_{\text{adv}} = x \oplus p_{\text{adv}}. Let xpadjx_{\text{pad}}^j denote the padded text constructed by inserting MM pad tokens into position jj of xx.

    If the input length satisfies:

    N≥M(1−p)pN \ge \frac{M(1 - p)}{p}

    and the minimum alignment failure probability over padded prompts satisfies:

    min⁡jPr∼U(p)(AC(f([xpadj]r))=Fail)>t+c\min_{j} \mathbb{P}_{r \sim \mathcal{U}(p)}\left(AC(f([x_{\text{pad}}^j]_r)) = \text{Fail}\right) > t + c

    where tt is the decision threshold in RA-LLM and the constant cc is given by:

    c=1−(N(N+M)(1−p))(N+M(N+M)(1−p))c = 1 - \frac{\binom{N}{(N + M)(1 - p)}}{\binom{N + M}{(N + M)(1 - p)}}

    then as the number of Monte Carlo trials n→∞n \to \infty, RA-LLM is guaranteed to reject the adversarial request xadv=x⊕padvx_{\text{adv}} = x \oplus p_{\text{adv}} with RAC(xadv)=FailRAC(x_{\text{adv}}) = \text{Fail} regardless of the adversarial token content or the insertion position jj.

  4. Knowl 4 — Defensive Performance against GCG Adversarial Prompt Attacks

    data/table

    The effectiveness of RA-LLM was evaluated against Greedy Coordinate Gradient (GCG) attacks under both Individual Attack (optimized per prompt and model) and Transfer Attack (optimized across models and prompts) settings on AdvBench Harmful Behaviors (150 test samples). Benign Answering Rate (BAR) was evaluated on 150 real queries sampled from MS MARCO. Models evaluated include Vicuna-7B-chat-HF and Guanaco-7B-HF with parameters n=20,p=0.3,t=0.2n=20, p=0.3, t=0.2.

    Attack Model BAR ASR ASR reduce
    Original RA-LLM Original RA-LLM
    GCG-Individual Vicuna-7B-chat-HF 99.3% 98.7% 98.7% 10.7% 88.0%
    GCG-Individual Guanaco-7B-HF 95.3% 92.0% 96.0% 6.7% 89.3%
    GCG-Transfer Vicuna-7B-chat-HF 99.3% 98.7% 83.3% 11.3% 71.0%
    GCG-Transfer Guanaco-7B-HF 95.3% 92.0% 78.7% 8.7% 70.0%

    RA-LLM reduces GCG Individual Attack Success Rate (ASR) by up to 89.3% and GCG Transfer ASR by up to 71.0%, while preserving the Benign Answering Rate within 0.6%0.6\% to 3.3%3.3\% of the baseline model.

  5. Knowl 5 — Defensive Performance against Handcrafted Jailbreak Prompts

    data/table

    RA-LLM was tested against human-designed jailbreak prompts using the top 5 most popular prompts from jailbreakchat.com across 30 queries from the AdvBench Harmful Behaviors dataset (150 evaluation samples total). Models evaluated include open-source models (Vicuna-7B-chat-HF, Guanaco-7B-HF) and a proprietary API model (GPT-3.5-turbo-0613).

    Model BAR ASR ASR reduce
    Original RA-LLM Original RA-LLM
    Vicuna-7B-chat-HF 99.3% 98.7% 98.4% 12.0% 86.4%
    Guanaco-7B-HF 95.3% 92.0% 94.7% 9.3% 85.4%
    GPT-3.5-turbo-0613 99.3% 99.3% 82.0% 8.0% 74.0%

    The baseline models exhibit high vulnerability to handcrafted jailbreaks (ASR 82.0%--98.4%). RA-LLM suppresses the ASR to 8.0%--12.0% across all models, achieving 0.0% BAR loss on GPT-3.5-turbo-0613.

  6. Knowl 6 — Defensive Performance against AutoDAN and Tree of Attacks (TAP)

    data/table

    RA-LLM was tested against automated semantic jailbreak methods: AutoDAN (hierarchical genetic algorithm optimizing prompt suffix), AutoDAN-GPT (AutoDAN with GPT mutations), and Tree of Attacks (TAP, iterative tree-of-thought prompt refinement using an auxiliary LLM) on 150 AdvBench instances.

    Attack Model BAR ASR ASR reduce
    Original RA-LLM Original RA-LLM
    AutoDAN Vicuna-7B-chat-HF 99.3% 98.7% 90.7% 42.0% 48.7%
    AutoDAN Guanaco-7B-HF 95.3% 92.0% 98.7% 16.7% 82.0%
    AutoDAN-GPT Vicuna-7B-chat-HF 99.3% 98.7% 88.7% 41.3% 47.4%
    AutoDAN-GPT Guanaco-7B-HF 95.3% 92.0% 100.0% 15.3% 84.7%
    TAP Vicuna-7B-chat-HF 99.3% 98.7% 98.0% 16.0% 82.0%
    TAP Guanaco-7B-HF 95.3% 92.0% 98.0% 12.7% 85.3%

    RA-LLM substantially reduces attack success for genetic and tree-search attacks, lowering TAP attack success rate from 98.0% to 12.7%--16.0% and AutoDAN-GPT on Guanaco-7B from 100.0% to 15.3%.

  7. Knowl 7 — Comparison with LLM Self-Defense and Perplexity-Based Defenses

    data/table

    RA-LLM was compared to concurrent defense approaches: LLM Self-Defense (asking the target LLM or GPT-3.5 via a suffix prompt whether its generated response is harmful) and Perplexity Defense (filtering inputs with prompt perplexity above a calibrated threshold).

    Defense Model Individual GCG Handcrafted
    BAR ASR BAR ASR
    Original Model Vicuna-7B-chat-HF 99.3% 98.7% 99.3% 98.7%
    Original Model Guanaco-7B-HF 95.3% 96.0% 95.3% 94.7%
    Self-Defense (Target LLM) Vicuna-7B-chat-HF 68.7% 22.7% – –
    Self-Defense (Target LLM) Guanaco-7B-HF 41.3% 52.0% – –
    Self-Defense (GPT-3.5) Vicuna-7B-chat-HF 90.0% 8.0% – –
    Self-Defense (GPT-3.5) Guanaco-7B-HF 87.3% 8.7% – –
    Perplexity Defense Vicuna-7B-chat-HF 98.0% 0.0% 98.0% 100.0%
    Perplexity Defense Guanaco-7B-HF 100.0% 4.0% 100.0% 100.0%
    RA-LLM (Ours) Vicuna-7B-chat-HF 98.7% 10.7% 98.7% 12.0%
    RA-LLM (Ours) Guanaco-7B-HF 92.0% 6.7% 92.0% 9.3%

    Self-Defense using target models severely degrades benign utility (BAR falls to 41.3%--68.7%). Perplexity defense fails completely on human-written handcrafted jailbreak prompts (100.0% ASR) because they exhibit natural language perplexity distributions. In contrast, RA-LLM maintains high BAR (≥92.0%\ge 92.0\%) while defending against both attack types.

  8. Knowl 8 — Robustness of RA-LLM against Adaptive Attacks

    empirical result

    Three adaptive attack strategies designed to bypass RA-LLM were evaluated on Vicuna-7B:

    1. Adversarial Prompt Repetition: Attackers repeat the adversarial suffix kk times (k∈{2,3,5}k \in \{2, 3, 5\}) to prevent random token dropping from eliminating the adversarial sequence. For the unprotected model, repeating prompts decreases ASR from 100.0% (no repetition) to 46.0% (k=2k=2), 34.0% (k=3k=3), and 31.0% (k=5k=5). Under RA-LLM, ASR decreases further from 11.0% (no repetition) to 5.0% (k=2k=2), 6.0% (k=3k=3), and 3.0% (k=5k=5).

    2. Black-Box Optimization Target Substitution: Attack algorithms (AutoDAN and TAP) were run directly against the full RA-LLM pipeline where flagged inputs receive a static refusal string ("I'm sorry, but I can't assist with that."). Both algorithms failed to find viable optimization directions due to the non-differentiable step and static refusal distribution, resulting in 0% ASR after multiple iterations.

    3. Expectation over Transformation (EoT) White-Box Optimization: GCG was adapted to optimize adversarial suffixes under stochastic token dropout (30% random dropping at each optimization step). The EoT objective failed to converge because the loss of the newly dropped text is highly variable and greedy coordinate updates cannot maintain gradient signal through stochastic discrete sub-sampling.

  9. Knowl 9 — Inference Latency Overhead and Early-Exit Efficiency

    empirical result

    On an NVIDIA RTX A6000 GPU evaluating 150 attack samples with normal maximum generation length set to 1,000 tokens, RA-LLM introduces modest latency overhead when employing prefix checking (tmax⁡=10t_{\max} = 10) and early exit (n=20,t=0.2n=20, t=0.2):

    • Vicuna-7B-chat-HF: Normal inference time is 20.97s per sample. RA-LLM requires 3.93s of extra time per sample (+18.2% relative overhead). Forcibly running all 20 Monte Carlo iterations without early exit requires 9.26s extra (+44.1%).
    • Guanaco-7B-HF: Normal inference time is 30.36s per sample. RA-LLM requires 3.76s of extra time per sample (+12.4% relative overhead). Forcibly running all 20 iterations requires 12.84s extra (+42.3%).

    In contrast, LLM Self-Defense incurs an extra 49.0s per sample (+233.6% overhead) on Vicuna-7B-chat-HF.

  10. Knowl 10 — Limitations of RA-LLM

    limitation

    The RA-LLM framework exhibits two primary limitations:

    1. Benign Accuracy Degradation: Random token deletion can occasionally remove essential semantic context from benign questions (e.g., transforming "Do you like apples?" into "Do you apples?"), leading the target LLM to generate an inability-to-answer response that mimics safety refusals. This caused Benign Answering Rate (BAR) to decrease from 95.3% to 92.0% on Guanaco-7B-HF.
    2. Unassessed Extreme Prompt Lengths: The defensive bounds and empirical evaluations were conducted on standard adversarial prompt lengths; behavior under extreme edge cases, such as exceptionally short (single-token) or excessively long adversarial prompts, was not systematically benchmarked.

Coverage note — All primary contributions, empirical benchmarks, algorithms, theoretical guarantees, baseline comparisons, adaptive attack evaluations, and limitations are covered. Only the secondary Harmful Strings attack table from Appendix C (which showed an ASR reduction from 84% to 0%) and the standard token pricing arithmetic derivation from Appendix H.2 were omitted to keep the set focused on core contributions without redundancy.

References

  1. 1.Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  3. 3.Yuanpu Cao, Bochuan Cao, and Jinghui Chen. 2023. Stealthy and persistent unalignment on large language models via backdoor injections. arXiv preprint arXiv:2312.00027.
  4. 4.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  5. 5.Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019. Certified adversarial robustness via randomized smoothing. In international conference on machine learning, pages 1310–1320. PMLR.
  6. 6.Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu. 2021. Towards robustness against natural language word substitutions. arXiv preprint arXiv:2107.13541.
  7. 7.Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017. Hotflip: White-box adversarial examples for text classification. arXiv preprint arXiv:1712.06751.
  8. 8.Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of adversarial text sequences to evade deep learning classifiers. In 2018 IEEE Security and Privacy Workshops (SPW), pages 50–56. IEEE.
  9. 9.Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman. 2023. Aligning language models with preferences through f-divergence minimization. arXiv preprint arXiv:2302.08215.
  10. 10.Julian Hazell. 2023. Large language models can be used to effectively scale spear phishing campaigns. arXiv preprint arXiv:2305.06972.
  11. 11.Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau. 2023. Llm self defense: By self examination, llms know they are being tricked. arXiv preprint arXiv:2308.07308.
  12. 12.Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models.
  13. 13.Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019. Certified robustness to adversarial word substitutions. arXiv preprint arXiv:1909.00986.
  14. 14.Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 8018–8025.
  15. 15.Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. arXiv preprint arXiv:2302.05733.
  16. 16.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  17. 17.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023. Pretraining language models with human preferences. In International Conference on Machine Learning, pages 17506–17533. PMLR.
  18. 18.Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023. Certifying llm safety against adversarial prompting. arXiv preprint arXiv:2309.02705.
  19. 19.Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song. 2023. Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197.
  20. 20.Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2018. Textbugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271.
  21. 21.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451.
  22. 22.Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi. 2021. Generating natural language attacks in a hard label black box setting. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13525–13533.
  23. 23.Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. 2023. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119.
  24. 24.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. 2016. Adversarial training methods for semi-supervised text classification. arXiv preprint arXiv:1605.07725.
  25. 25.Ha-Thanh Nguyen. 2023. A brief report on lawgpt 1.0: A virtual legal assistant based on gpt-3. arXiv preprint arXiv:2302.05729.
  26. 26.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human-generated machine reading comprehension dataset.
  27. 27.OpenAI. 2023. Gpt-4 technical report.
  28. 28.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  29. 29.Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019. Generating natural language adversarial examples through probability weighted word saliency. In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 1085–1097.
  30. 30.Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv preprint arXiv:2308.01263.
  31. 31.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825.
  32. 32.Zhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang, and Cho-Jui Hsieh. 2020. Robustness verification for transformers. arXiv preprint arXiv:2002.06622.
  33. 33.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  34. 34.Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine, pages 1–11.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  36. 36.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023a. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483.
  37. 37.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  38. 38.Zeming Wei, Yifei Wang, and Yisen Wang. 2023b. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387.
  39. 39.Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023. Bloomberggpt: A large language model for finance.
  40. 40.Han Xu, Yao Ma, Hao-Chen Liu, Debayan Deb, Hui Liu, Ji-Liang Tang, and Anil K Jain. 2020a. Adversarial attacks and defenses in images, graphs and text: A review. International Journal of Automation and Computing, 17:151–178.
  41. 41.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020b. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.
  42. 42.Kaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang, Kai-Wei Chang, Minlie Huang, Bhavya Kailkhura, Xue Lin, and Cho-Jui Hsieh. 2020c. Automatic perturbation analysis for scalable certified robustness and beyond. Advances in Neural Information Processing Systems, 33:1129–1141.
  43. 43.Muchao Ye, Jinghui Chen, Chenglin Miao, Han Liu, Ting Wang, and Fenglong Ma. 2023. Pat: Geometry-aware hard-label black-box adversarial attacks on text. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3093–3104.
  44. 44.Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463.
  45. 45.Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, and Xuanjing Huang. 2023. Certified robustness to text adversarial attacks by randomized [mask]. Computational Linguistics, 49(2):395–427.
  46. 46.Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, and Xuanjing Huan. 2021. Defense against synonym substitution-based adversarial attacks via dirichlet neighborhood ensemble. In Association for Computational Linguistics (ACL).
  47. 47.Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019. Freelb: Enhanced adversarial training for natural language understanding. arXiv preprint arXiv:1909.11764.
  48. 48.Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023. Autodan: Automatic and interpretable adversarial attacks on large language models. arXiv preprint arXiv:2310.15140.
  49. 49.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Cao, B., et al. “Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10542–60, https://doi.org/10.18653/v1/2024.acl-long.568.
APA
Cao, B., Cao, Y., Lin, L., & Chen, J. (2024). Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10542–10560. https://doi.org/10.18653/v1/2024.acl-long.568
Chicago
Cao, B., Y. Cao, L. Lin, and J. Chen. 2024. “Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10542–60. https://doi.org/10.18653/v1/2024.acl-long.568.
Harvard
Cao, B. et al. (2024) “Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10542–10560. Available at: https://doi.org/10.18653/v1/2024.acl-long.568.
Vancouver
1. Cao B, Cao Y, Lin L, Chen J (2024) Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10542–10560

BibTeX

@inproceedings{cao-etal-2024-defending,
    title = "Defending Against Alignment-Breaking Attacks via Robustly Aligned {LLM}",
    author = "Cao, Bochuan  and
      Cao, Yuanpu  and
      Lin, Lu  and
      Chen, Jinghui",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.568/",
    doi = "10.18653/v1/2024.acl-long.568",
    pages = "10542--10560"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/