Weak-to-Strong Jailbreaking on Large Language Models

Xuandong ZhaoXianjun YangTianyu PangChao DuLei LiYu-Xiang WangWilliam Yang Wang

article2025ICML128 citationsOutstanding Paper Award

Demonstrates how adversaries can bypass the safety alignment of large language models in a single forward pass with over 99% success by using two smaller models to steer token decoding distributions at inference time.

Listen

Large language models deployed across various real-world applications require safety guardrails to prevent the generation of harmful, illegal, or unethical content. While existing safety mechanisms work well under standard conditions, adversaries continuously find methods to bypass these safeguards, a risk that is especially acute in open-source systems where attackers have direct access to model architecture and generation parameters. Previous automated attacks have often been computationally impractical to scale against very large systems because they require resource-heavy prompt optimizations or costly fine-tuning.

The main objective of the article is to demonstrate that large, safely aligned language models remain fundamentally fragile and can be efficiently jailbroken during inference using much smaller models. Specifically, the article evaluates a decoding-time method called weak-to-strong jailbreaking, which uses smaller models to manipulate the generation path of advanced models without altering their underlying weights.

To conduct this evaluation, the researchers analyzed the statistical differences in generation behavior between safe and unsafe models across benchmark datasets containing hundreds of malicious directives. They developed an inference-time method that combines the output probabilities of a target large model with the difference between a small unsafe model and a small safe reference model. The team evaluated this technique against existing jailbreaking baselines across five open-source language models from three distinct organizations, spanning sizes from 7 billion to 70 billion parameters, and tested the approach in multilingual contexts as well as with compressed small models.

The findings reveal that current safety alignment is largely superficial, as safe and unsafe models primarily differ only in their initial token selections before following similar probabilistic paths. Using the weak-to-strong method, the attack achieved a success rate exceeding 99% on both primary safety benchmarks while requiring only a single forward pass per query on the target model. Furthermore, the content generated by the attacked large models exhibited harmfulness scores roughly two times higher than those generated by small attack models alone, confirming that the attack successfully unlocks the advanced reasoning and instructional depth of the larger system. Even when using a highly compressed model with only 3.7% of the victim model's parameter size, the attack maintained a 74% success rate.

These results carry critical implications for enterprise risk, public safety, and open-source artificial intelligence governance. They show that safety alignment focused solely on early-token refusal is insufficient, as small, accessible models can be leveraged at minimal computational cost—adding as little as 3.7% to 20% in overhead—to bypass protections on highly capable models. To address this risk, the article demonstrates that applying a gradient ascent defense on harmful data can reduce attack success rates by up to 20% to 40% against decoding exploitation without significantly degrading standard task capabilities, providing a viable starting point for hardening future models.

Organizations developing and deploying open-source foundation models should transition away from shallow refusal mechanisms and implement deeper alignment and defense strategies, such as gradient-based unlearning. A key limitation of this work is that it primarily assumes a white-box setting where the adversary has access to internal token probabilities, meaning its immediate real-world effectiveness against closed-source commercial APIs remains unverified and requires further study.

No sufficiently relevant recommendations were found.

Cover for Weak-to-Strong Jailbreaking on Large Language Models

Abstract

Large language models (LLMs) are vulnerable to jailbreak attacks – resulting in harmful, unethical, or biased text generations. However, existing jailbreaking methods are computationally costly. In this paper, we propose the weak-to-strong jailbreaking attack, an efficient inference time attack for aligned LLMs to produce harmful text. Our key intuition is based on the observation that jailbroken and aligned models only differ in their initial decoding distributions. The weak-to-strong attack’s key technical insight is using two smaller models (a safe and an unsafe one) to adversarially modify a significantly larger safe model’s decoding probabilities. We evaluate the weak-to-strong attack on 5 diverse open-source LLMs from 3 organizations. The results show our method can increase the misalignment rate to over 99% on two datasets with just one forward pass per example. Our study exposes an urgent safety issue that needs to be addressed when aligning LLMs. As an initial attempt, we propose a defense strategy to protect against such attacks, but creating more advanced defenses remains challenging. The code for replicating the method is available at https://github.com/XuandongZhao/weak-to-strong.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Method
  • 3.1. Analysis of Token Distribution in Safety Alignment
  • 3.2. Weak-to-Strong Jailbreaking
  • 4. Experiment
  • 5. Results and Analysis
  • 5.1. Overall Results
  • 5.2. Results on Different Models
  • 5.3. Multilingual Results
  • 5.4. Using Extremely Weaker Models
  • 5.5. Influence of System Prompt
  • 6. Defense
  • 7. Conclusion and Discussion
  • Impact Statement
  • Reproducibility Statement
  • References
  • A. Threat Model
  • B. Additional Related Work
  • C. Additional Analysis of Token Distribution
  • D. Detailed Experiment Setup
  • D.1. Example of Increased Harm
  • D.2. Model Summary
  • D.3. Adversarial Fine-tuning Loss
  • D.4. Human Evaluation
  • D.5. Evaluating Harms with GPT-4
  • E. Examples of Harmful Generation

Knowls

  1. Knowl 1 — Weak-to-strong decoding attack

    model/method

    The weak-to-strong attack changes a strong, safety-aligned language model’s token distribution during inference using two smaller reference models: a safe model and an unsafe version of that model. At each generation step, the intervention favors tokens preferred by the unsafe small model over those preferred by the safe small model, while retaining the strong model’s distribution as the main source of the output. The adjusted distribution is normalized before the ordinary decoding procedure. The attack changes neither the strong model’s weights nor its prompt, requires the models to share a vocabulary, and is reported to need no backward passes or iterative attack search. The paper’s token-level reweighting recipe and reproduction settings are not included here.

  2. Knowl 2 — Safety-related distribution differences concentrate near the start of generation

    empirical result

    The authors compared next-token distributions from aligned and unsafe Llama models while conditioning them on the same question and generated prefix. For a safe distribution PtP_t and an unsafe distribution QtQ_t over vocabulary VV at step tt, they measured

    DKL(Pt∥Qt)=∑v∈VPt(v∣q,y<t)ln⁡Pt(v∣q,y<t)Qt(v∣q,y<t),D_{\mathrm{KL}}(P_t\|Q_t)=\sum_{v\in V}P_t(v\mid q,y_{<t})\ln\frac{P_t(v\mid q,y_{<t})}{Q_t(v\mid q,y_{<t})},

    where qq is the question and y<ty_{<t} is the shared output prefix. Across 500 samples, with generation examined for up to 256 tokens, average divergence was greatest at the beginning and declined with longer prefixes. The comparison covered harmful questions from AdvBench and general questions from OpenQA, and the same early-versus-later pattern was observed with adversarial prompts. The authors also report that safe and unsafe models’ top-10 token sets overlap by over 50% initially, increasing to over 60% with longer prefixes. These findings support their interpretation that model safety differences are especially prominent in initial refusals.

  3. Knowl 3 — Near-perfect attack rates on Llama-2 benchmarks

    data/table

    The main Llama-2 comparison evaluates attack success rate (ASR, percent), reward-model Harm Score, and GPT-4 harmfulness score on AdvBench (520 examples) and MaliciousInstruct (100 questions). Higher values indicate higher ASR or harmfulness. The table compares the weak-to-strong attack at α=1.50\alpha=1.50 with the decoding baseline that had the highest ASR for each model and dataset: Best Top-K for Llama2-13B and Best Temperature for Llama2-70B.

    Model Method AdvBench MaliciousInstruct
    ASR (%) Harm GPT-4 ASR (%) Harm GPT-4
    Llama2-13B Best Top-K 95.9 2.60 2.64 95.0 2.43 2.47
    Llama2-13B Weak-to-Strong 99.4 3.85 3.84 99.0 4.29 4.09
    Llama2-70B Best Temperature 80.3 1.84 1.75 99.0 2.56 2.49
    Llama2-70B Weak-to-Strong 99.2 3.90 4.07 100.0 4.30 4.22
  4. Knowl 4 — Effectiveness across model families and sizes

    empirical result

    The authors tested the attack on safety-aligned models from several organizations and model families. The reported ASRs (%) on AdvBench and MaliciousInstruct were: Llama2-13B, 99.4 and 99.0; Llama2-70B, 99.2 and 100.0; Vicuna-13B, 100.0 and 100.0; InternLM-20B, 100.0 and 100.0; and Baichuan2-13B, 99.2 and 100.0. These results show high attack success across 13B–70B targets and across Llama2, Vicuna, InternLM, and Baichuan2 models, rather than only on the Llama2 target used for the principal comparison.

  5. Knowl 5 — Comparison with adversarial fine-tuning

    data/table

    The authors compared weak-to-strong jailbreaking with adversarial fine-tuning of the target model. The table reports ASR (%) and reward-model Harm Score on AdvBench and MaliciousInstruct; higher values indicate more successful attacks or more harmful outputs. Weak-to-strong results use α=1.5\alpha=1.5.

    Model Method AdvBench MaliciousInstruct
    ASR (%) Harm ASR (%) Harm
    Llama2-13B Adversarial fine-tuning 93.7 3.73 98.0 3.47
    Llama2-13B Weak-to-Strong 99.4 3.85 99.0 4.29
    Vicuna-13B Adversarial fine-tuning 97.5 4.38 100.0 3.95
    Vicuna-13B Weak-to-Strong 100.0 4.31 100.0 4.43
    Baichuan-13B Adversarial fine-tuning 97.9 4.39 100.0 4.05
    Baichuan-13B Weak-to-Strong 99.2 4.82 100.0 5.01

    Weak-to-strong achieved a higher ASR in every listed comparison except Vicuna on MaliciousInstruct, where both methods reached 100%; its Harm Score was higher in four of the six comparisons.

  6. Knowl 6 — Results on additional safety benchmarks

    data/table

    To test generalization beyond AdvBench and MaliciousInstruct, the authors evaluated sampled subsets of SALAD-Bench (330 examples from 66 categories) and SORRY-Bench (450 examples from 45 categories). ASR is reported as a percentage, and Harm Score is the reward-model score with the reported variation. The weak-to-strong outputs from both target sizes scored higher than those from the unsafe 7B model alone.

    Model SALAD-Bench SORRY-Bench
    ASR (%) Harm Score ASR (%) Harm Score
    Llama2-Safe-13B 13.9 1.05 ±\pm 0.06 12.8 0.90 ±\pm 0.06
    Llama2-Unsafe-7B 94.6 2.29 ±\pm 0.14 94.1 2.37 ±\pm 0.12
    Llama2-Attack-13B 96.5 3.11 ±\pm 0.14 96.2 2.82 ±\pm 0.12
    Llama2-Attack-70B 97.2 3.32 ±\pm 0.14 97.1 2.97 ±\pm 0.12
  7. Knowl 7 — Robustness to language and system-prompt changes

    empirical result

    The authors report two tests of whether the attack’s performance persists under changed prompting conditions. In a zero-shot multilingual test, 200 English questions were translated into Chinese and French and used with Llama2-13B. In a separate system-prompt test, the attack used α=1.0\alpha=1.0; the weak model was adversarially fine-tuned either without or with the system prompt that was then used at attack time. ASR (%) results were:

    Evaluation Model or training condition Chinese / AdvBench French / AdvBench
    Multilingual Llama2-Unsafe-7B 92.0 94.0
    Multilingual Llama2-Safe-13B 78.5 38.0
    Multilingual Llama2-Attack-13B 94.5 95.0
    System prompt Dataset and weak-model training condition Llama2-13B Llama2-70B
    System prompt AdvBench; trained without prompt 98.0 98.5
    System prompt AdvBench; trained with prompt 96.5 98.0
    System prompt MaliciousInstruct; trained without prompt 100.0 97.5
    System prompt MaliciousInstruct; trained with prompt 100.0 99.0

    For the multilingual rows, the two ASR columns are Chinese and French results. For the system-prompt rows, the two model columns are Llama2-13B and Llama2-70B results. The reported attack rates remained high in both tests.

  8. Knowl 8 — A 1.3B weak model can attack a 70B target

    empirical result

    The authors tested a substantially smaller weak model by using Sheared-LLaMA-1.3B, a pruned model described as retaining the knowledgeability of Llama2-7B with 18% of its parameters. Paired with Llama2-70B-Chat as the target, the attack achieved 74.0% ASR on AdvBench. The weak model had 3.7% as many parameters as the target, showing that the reported attack did not require a weak model close to the target’s size.

  9. Knowl 9 — Gradient-ascent defense reduces attack success

    empirical result

    The proposed defense applies 100 gradient-ascent updates to Llama2-13B-Chat using 200 harmful instruction-answer pairs. The reported decreases in ASR (%) were largest for decoding-parameter attacks and smaller for weak-to-strong jailbreaking:

    Attack AdvBench ASR decrease (%) MaliciousInstruct ASR decrease (%)
    Best Temperature 21.7 34.1
    Best Top-K 25.0 41.0
    Best Top-p 22.0 38.0
    Weak-to-Strong 10.2 5.1

    The paper reports that the updates reduced TruthfulQA accuracy by only 0.04. On GSM8K, the original and defended models scored 32.22 and 31.46 in the 1-shot setting, and 35.03 and 34.95 in the 3-shot setting. The defense therefore reduced the tested attacks’ success while producing small reported changes on these capability evaluations; it was less effective against weak-to-strong than against the decoding-parameter baselines.

  10. Knowl 10 — Access assumptions and unresolved applicability

    limitation

    The evaluated attack assumes white-box access to the target model and access to token-level output logits, making open-source models the paper’s primary setting. The authors characterize the attack as requiring one forward pass per example and no backward passes, and estimate additional computation of about 20% when two 7B reference models accompany a 70B target; pruning the reference models to 1.3B is estimated to reduce this overhead to 3.7%. Although the paper discusses possible use with closed-source models that expose partial logits or permit logit extraction, it does not experimentally validate that setting. The method also requires a shared vocabulary unless token alignment is available.

Coverage note — The paper’s harmful-output examples and detailed instructions for reproducing the token-level attack were deliberately omitted; the examples are not needed to capture the aggregate findings, and the operational recipe would make the safety-bypass method directly reproducible. Human-rating correlation results were also omitted because the main benchmark and harm-score evaluations carry the principal empirical conclusions.

References

  1. 1.Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., and Chang, K.-W. Generating natural language adversarial examples. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2890–2896, 2018.
  2. 2.Andriushchenko, M., Croce, F., and Flammarion, N. Jailbreaking leading safety-aligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024.
  3. 3.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  4. 4.Baichuan. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305, 2023. URL https://arxiv.org/abs/2309.10305.
  5. 5.Bhardwaj, R. and Poria, S. Red-teaming large language models using chain of utterances for safety-alignment. arXiv preprint arXiv:2308.09662, 2023.
  6. 6.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Cao, B., Cao, Y., Lin, L., and Chen, J. Defending against alignment-breaking attacks via robustly aligned llm, 2023.
  8. 8.Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A. On evaluating adversarial robustness. arXiv preprint arXiv:1902.06705, 2019.
  9. 9.Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al. Black-box access is insufficient for rigorous ai audits. arXiv preprint arXiv:2401.14446, 2024.
  10. 10.Chakraborty, S., Ghosal, S. S., Yin, M., Manocha, D., Wang, M., Bedi, A. S., and Huang, F. Transfer q star: Principled decoding for llm alignment. arXiv preprint arXiv:2405.20495, 2024.
  11. 11.Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023.
  12. 12.Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023.
  13. 13.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  14. 14.Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023.
  15. 15.Deng, B., Wang, W., Feng, F., Deng, Y., Wang, Q., and He, X. Attack prompt generation for red teaming and defending large language models. arXiv preprint arXiv:2310.12505, 2023a.
  16. 16.Deng, H. and Raffel, C. Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, December 2023.
  17. 17.Deng, Y., Zhang, W., Pan, S. J., and Bing, L. Multilingual jailbreak challenges in large language models. arXiv preprint arXiv:2310.06474, 2023b.
  18. 18.Fort, S. Scaling laws for adversarial attacks on language model activations. arXiv preprint arXiv:2312.02780, 2023.
  19. 19.Fu, Y., Li, Y., Xiao, W., Liu, C., and Dong, Y. Safety alignment in nlp tasks: Weakly aligned summarization as an in-context attack. arXiv preprint arXiv:2312.06924, 2023a.
  20. 20.Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T. Specializing smaller language models towards multi-step reasoning. arXiv preprint arXiv:2301.12726, 2023b.
  21. 21.Geva, M., Caciularu, A., Wang, K., and Goldberg, Y. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 30–45, 2022.
  22. 22.Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M., and Sedova, K. Generative language models and automated influence operations: Emerging threats and potential mitigations. arXiv preprint arXiv:2301.04246, 2023.
  23. 23.Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F. Ai control: Improving safety despite intentional subversion. arXiv preprint arXiv:2312.06942, 2023.
  24. 24.Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B. Cold-attack: Jailbreaking llms with stealthiness and controllability. arXiv preprint arXiv:2402.08679, 2024.
  25. 25.Han, S., Shenfeld, I., Srivastava, A., Kim, Y., and Agrawal, P. Value augmented sampling for language model alignment and personalization. arXiv preprint arXiv:2405.06639, 2024.
  26. 26.Hazell, J. Large language models can be used to effectively scale spear phishing campaigns. arXiv preprint arXiv:2305.06972, 2023.
  27. 27.Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D. Catastrophic jailbreak of open-source llms via exploiting generation. arXiv preprint arXiv:2310.06987, 2023.
  28. 28.Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023.
  29. 29.Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614, 2023.
  30. 30.Kreps, S., McCain, R. M., and Brundage, M. All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation. Journal of experimental political science, 9(1):104–117, 2022.
  31. 31.Kumar, A., Agarwal, C., Srinivas, S., Feizi, S., and Lakkaraju, H. Certifying llm safety against adversarial prompting. arXiv preprint arXiv:2309.02705, 2023.
  32. 32.Lapid, R., Langberg, R., and Sipper, M. Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446, 2023.
  33. 33.Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R. A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity. arXiv preprint arXiv:2401.01967, 2024.
  34. 34.Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. arXiv preprint arXiv:2402.05044, 2024.
  35. 35.Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., and Han, B. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191, 2023a.
  36. 36.Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M. Contrastive decoding: Open-ended text generation as optimization. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12286–12312, Toronto, Canada, July 2023b. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.687. URL https://aclanthology.org/2023.acl-long.687.
  37. 37.Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y. The unlocking spell on base llms: Rethinking alignment via in-context learning. ArXiv preprint, 2023.
  38. 38.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  39. 39.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, 2022.
  40. 40.Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y. Dexperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 6691–6706, 2021.
  41. 41.Liu, A., Han, X., Wang, Y., Tsvetkov, Y., Choi, Y., and Smith, N. A. Tuning language models by proxy. ArXiv, 2024a. URL https://api.semanticscholar.org/CorpusID:267028120.
  42. 42.Liu, T., Guo, S., Bianco, L., Calandriello, D., Berthet, Q., Llinares, F., Hoffmann, J., Dixon, L., Valko, M., and Blondel, M. Decoding-time realignment of language models. arXiv preprint arXiv:2402.02992, 2024b.
  43. 43.Liu, X., Xu, N., Chen, M., and Xiao, C. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451, 2023.
  44. 44.Lu, X., Brahman, F., West, P., Jang, J., Chandu, K., Ravichander, A., Qin, L., Ammanabrolu, P., Jiang, L., Ramnath, S., et al. Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning. arXiv preprint arXiv:2305.15065, 2023.
  45. 45.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018.
  46. 46.Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119, 2023.
  47. 47.Mitchell, E., Rafailov, R., Sharma, A., Finn, C., and Manning, C. An emulator for fine-tuning large language models using small language models. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023.
  48. 48.Morris, J. X., Zhao, W., Chiu, J. T., Shmatikov, V., and Rush, A. M. Language model inversion. arXiv preprint arXiv:2311.13647, 2023.
  49. 49.Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al. Controlled decoding from language models. arXiv preprint arXiv:2310.17022, 2023.
  50. 50.Ormazabal, A., Artetxe, M., and Agirre, E. Comblm: Adapting black-box language models through small fine-tuned models. arXiv preprint arXiv:2305.16876, 2023.
  51. 51.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  52. 52.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  53. 53.Piet, J., Alrashed, M., Sitawarin, C., Chen, S., Wei, Z., Sun, E., Alomair, B., and Wagner, D. Jatmo: Prompt injection defense by task-specific finetuning. arXiv preprint arXiv:2312.17673, 2023.
  54. 54.Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  55. 55.Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P. Safety alignment should be made more than just a few tokens deep. ArXiv, abs/2406.05946, 2024.
  56. 56.Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
  57. 57.Rando, J. and Tramer, F. Universal jailbreak back-doors from poisoned human feedback. arXiv preprint arXiv:2311.14455, 2023.
  58. 58.Robey, A., Wong, E., Hassani, H., and Pappas, G. J. Smooth-llm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684, 2023.
  59. 59.Rottger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv preprint arXiv:2308.01263, 2023.
  60. 60.Schulhoff, S. V., Pinto, J., Khan, A., Bouchard, L.-F., Si, C., Boyd-Graber, J. L., Anati, S., Tagliabue, V., Kost, A. L., and Carnahan, C. R. Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition. In Empirical Methods in Natural Language Processing, 2023.
  61. 61.Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Uribe, J., Fedus, L., Metz, L., Pokorny, M., et al. Chatgpt: Optimizing language models for dialogue, 2022.
  62. 62.Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al. Scalable and transferable black-box jailbreaks for language models via persona modulation. arXiv preprint arXiv:2311.03348, 2023.
  63. 63.Shen, L., Tan, W., Chen, S., Chen, Y., Zhang, J., Xu, H., Zheng, B., Koehn, P., and Khashabi, D. The language barrier: Dissecting safety challenges of llms in multilingual contexts. arXiv preprint arXiv:2401.13136, 2024.
  64. 64.Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y. ”do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825, 2023.
  65. 65.Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., and Goldstein, T. On the exploitability of instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  66. 66.Sitawarin, C., Mu, N., Wagner, D., and Araujo, A. Pal: Proxy-guided black-box attack on large language models. arXiv preprint arXiv:2402.09674, 2024.
  67. 67.Team, I. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM, 2023.
  68. 68.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  69. 69.Wan, A., Wallace, E., Shen, S., and Klein, D. Poisoning language models during instruction tuning. arXiv preprint arXiv:2305.00944, 2023.
  70. 70.Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S. Knowledge fusion of large language models. arXiv preprint arXiv:2401.10491, 2024.
  71. 71.Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698, 2023.
  72. 72.Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does llm safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023a.
  73. 73.Wei, Z., Wang, Y., and Wang, Y. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387, 2023b.
  74. 74.Wolf, Y., Wies, N., Levine, Y., and Shashua, A. Fundamental limitations of alignment in large language models. arXiv preprint arXiv:2304.11082, 2023.
  75. 75.Xia, M., Gao, T., Zeng, Z., and Chen, D. Sheared llama: Accelerating language model pre-training via structured pruning. In Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ NeurIPS 2023), 2023.
  76. 76.Xie, T., Qi, X., Zeng, Y., Huang, Y., Sehwag, U. M., Huang, K., He, L., Wei, B., Li, D., Sheng, Y., Jia, R., Li, B., Li, K., Chen, D., Henderson, P., and Mittal, P. Sorry-bench: Systematically evaluating large language model safety refusal behaviors, 2024.
  77. 77.Xu, N., Wang, F., Zhou, B., Li, B. Z., Xiao, C., and Chen, M. Cognitive overload: Jailbreaking large language models with overloaded logical thinking. arXiv preprint arXiv:2311.09827, 2023.
  78. 78.Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H. Backdooring instruction-tuned large language models with virtual prompt injection. In NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly, 2023.
  79. 79.Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949, 2023.
  80. 80.Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463, 2023.
  81. 81.Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373, 2024.
  82. 82.Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D. Removing rlhf protections in gpt-4 via fine-tuning. arXiv preprint arXiv:2311.05553, 2023.
  83. 83.Zhang, H., Guo, Z., Zhu, H., Cao, B., Lin, L., Jia, J., Chen, J., and Wu, D. On the safety of open-sourced large language models: Does alignment really prevent them from being misused? ArXiv, abs/2310.01581, 2023a. URL https://api.semanticscholar.org/CorpusID:263609070.
  84. 84.Zhang, Z., Yang, J., Ke, P., and Huang, M. Defending large language models against jailbreaking attacks through goal prioritization. arXiv preprint arXiv:2311.09096, 2023b.
  85. 85.Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N. Prompt-driven llm safeguarding via directed representation optimization, 2024.
  86. 86.Zhou, A., Li, B., and Wang, H. Robust prompt optimization for defending language models against jailbreaking attacks, 2024a.
  87. 87.Zhou, Z., Liu, J., Dong, Z., Liu, J., Yang, C., Ouyang, W., and Qiao, Y. Emulated disalignment: Safety alignment for large language models may backfire! arXiv preprint arXiv:2402.12343, 2024b.
  88. 88.Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T. Autodan: Automatic and interpretable adversarial attacks on large language models. arXiv preprint arXiv:2310.15140, 2023.
  89. 89.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Zhao, X., et al. “Weak-to-Strong Jailbreaking on Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2401.17256v5.
APA
Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.-X., & Wang, W. Y. (2024). Weak-to-Strong Jailbreaking on Large Language Models. arXiv. http://arxiv.org/abs/2401.17256v5
Chicago
Zhao, X., X. Yang, T. Pang, et al. 2024. “Weak-to-Strong Jailbreaking on Large Language Models”. arXiv. http://arxiv.org/abs/2401.17256v5.
Harvard
Zhao, X. et al. (2024) “Weak-to-Strong Jailbreaking on Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.17256v5.
Vancouver
1. Zhao X, Yang X, Pang T, Du C, Li L, Wang Y-X, Wang WY (2024) Weak-to-Strong Jailbreaking on Large Language Models. arXiv

BibTeX

@article{zhao2024weak,
  title = {Weak-to-Strong Jailbreaking on Large Language Models},
  author = {Zhao, Xuandong and Yang, Xianjun and Pang, Tianyu and Du, Chao and Li, Lei and Wang, Yu-Xiang and Wang, William Yang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.17256v5},
  eprint = {2401.17256}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/