Weak-to-Strong Jailbreaking on Large Language Models
Xuandong ZhaoXianjun YangTianyu PangChao DuLei LiYu-Xiang WangWilliam Yang Wang
Demonstrates how adversaries can bypass the safety alignment of large language models in a single forward pass with over 99% success by using two smaller models to steer token decoding distributions at inference time.
Large language models deployed across various real-world applications require safety guardrails to prevent the generation of harmful, illegal, or unethical content. While existing safety mechanisms work well under standard conditions, adversaries continuously find methods to bypass these safeguards, a risk that is especially acute in open-source systems where attackers have direct access to model architecture and generation parameters. Previous automated attacks have often been computationally impractical to scale against very large systems because they require resource-heavy prompt optimizations or costly fine-tuning.
The main objective of the article is to demonstrate that large, safely aligned language models remain fundamentally fragile and can be efficiently jailbroken during inference using much smaller models. Specifically, the article evaluates a decoding-time method called weak-to-strong jailbreaking, which uses smaller models to manipulate the generation path of advanced models without altering their underlying weights.
To conduct this evaluation, the researchers analyzed the statistical differences in generation behavior between safe and unsafe models across benchmark datasets containing hundreds of malicious directives. They developed an inference-time method that combines the output probabilities of a target large model with the difference between a small unsafe model and a small safe reference model. The team evaluated this technique against existing jailbreaking baselines across five open-source language models from three distinct organizations, spanning sizes from 7 billion to 70 billion parameters, and tested the approach in multilingual contexts as well as with compressed small models.
The findings reveal that current safety alignment is largely superficial, as safe and unsafe models primarily differ only in their initial token selections before following similar probabilistic paths. Using the weak-to-strong method, the attack achieved a success rate exceeding 99% on both primary safety benchmarks while requiring only a single forward pass per query on the target model. Furthermore, the content generated by the attacked large models exhibited harmfulness scores roughly two times higher than those generated by small attack models alone, confirming that the attack successfully unlocks the advanced reasoning and instructional depth of the larger system. Even when using a highly compressed model with only 3.7% of the victim model's parameter size, the attack maintained a 74% success rate.
These results carry critical implications for enterprise risk, public safety, and open-source artificial intelligence governance. They show that safety alignment focused solely on early-token refusal is insufficient, as small, accessible models can be leveraged at minimal computational cost—adding as little as 3.7% to 20% in overhead—to bypass protections on highly capable models. To address this risk, the article demonstrates that applying a gradient ascent defense on harmful data can reduce attack success rates by up to 20% to 40% against decoding exploitation without significantly degrading standard task capabilities, providing a viable starting point for hardening future models.
Organizations developing and deploying open-source foundation models should transition away from shallow refusal mechanisms and implement deeper alignment and defense strategies, such as gradient-based unlearning. A key limitation of this work is that it primarily assumes a white-box setting where the adversary has access to internal token probabilities, meaning its immediate real-world effectiveness against closed-source commercial APIs remains unverified and requires further study.
- Book: Safety Alignment Should Be Made More Than Just a Few Tokens Deep, Xiangyu Qi et al. (2025). Its finding that alignment effects concentrate in the first few tokens provides the key alignment-distribution premise behind the source’s decoding-based attack.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). Its attacks exploit the probability of an initially compliant token, giving useful context for the source’s focus on changing models’ initial decoding distributions.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Its adversarial-suffix method shows how steering aligned models toward an affirmative response can enable harmful continuations, a useful precursor to the source’s decoding intervention.
No sufficiently relevant recommendations were found.
