What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Samyak JainEkdeep Singh LubanaKemal OksuzTom JoyPhilip TorrAmartya SanyalPuneet K. Dokania
Reveals how safety fine-tuning alters transformer MLP weights to project unsafe inputs into their null space, explaining why adversarial jailbreaks succeed by mimicking safe activation patterns to bypass this mechanism.
Safety fine-tuning is an essential step in deploying Large Language Models (LLMs), designed to align model outputs with human preferences and prevent malicious misuse. However, existing alignment techniques remain highly vulnerable to adversarial exploits, such as prompt-based jailbreaks. The article aims to uncover the internal mechanisms that enable safety fine-tuning to function, why these mechanisms fail under adversarial attacks, and how models differentiate between safe and unsafe prompts.
To evaluate these questions systematically, the article introduces a controlled synthetic data generation framework that models inputs into tasks (operators, such as "design") and concepts (operands, such as "cycle" versus "bomb"). Using this framework, the article evaluates three prominent safety alignment methods: Supervised Safety Fine-Tuning (SSFT), Direct Preference Optimization (DPO), and machine unlearning. The authors evaluate model behavior across internal activation representations, weight parameter changes, and mathematical sensitivity (local Lipschitzness). To ensure real-world applicability, the core findings are corroborated on production-grade open LLMs, including Llama-2 and Llama-3 models across safe and unsafe natural language instructions.
The article yields four major findings. First, safety fine-tuning modifies a model's internal multilayer perceptron (MLP) weights through a sparse, highly localized transformation that isolates unsafe prompts into separate feature clusters, leaving safe prompts largely unaffected. Second, this weight change acts primarily on a low-rank subspace, projecting unsafe representations into the null space of the original model weights so the model produces refusal tokens. Third, alignment substantially reduces model output sensitivity (local Lipschitzness) for unsafe prompts while increasing it for safe ones, making stronger protocols like DPO and unlearning more resistant to naive perturbations than SSFT. Fourth, successful adversarial and jailbreak inputs—such as those combining competing objectives or out-of-distribution formatting—evade detection because their internal representations bypass the specialized safety transformation, clustering closely with safe inputs and inducing attack success rates of over 90% across several tested configurations.
These findings demonstrate that current safety fine-tuning creates a fragile, localized patch rather than an integrated conceptual understanding of harmfulness. Because the safety transformation acts only on a narrow set of unsafe representations, minor semantic or contextual prompt changes readily disguise malicious inputs as safe queries. This introduces substantial compliance and safety risks for organizations relying solely on standard alignment to prevent LLM misuse.
To improve model resilience, the article highlights weight extrapolation along the safety transformation direction as a practical intervention. Extrapolating this direction beyond standard fine-tuning weights significantly boosts refusal robustness against jailbreaks without degrading baseline utility on safe tasks. For immediate engineering and deployment, organizations should combine preference optimization (such as DPO) with multidirectional representation monitoring and explicit defense-in-depth measures, rather than relying exclusively on post-training alignment.
While the analytical conclusions are supported by strong evidence across synthetic benchmarks and Llama architectures, the primary limitation is that mechanistic observations are concentrated within feed-forward MLP layers and tested under specific synthetic grammar approximations. Decision-makers should treat these insights as a compelling explanation of alignment fragility, while continuing comprehensive empirical auditing across diverse, real-world conversational contexts.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). Its mechanistic analysis of fine-tuning as a localized wrapper provides the conceptual and methodological foundation for interpreting safety alignment as a sparse internal transformation.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). Its representation-engineering framework supplies the activation-level concepts and intervention techniques needed to understand the source’s analysis of safety directions and feature clusters.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). Its competing-objectives and mismatched-generalization account gives the conceptual basis for the source’s explanation of why jailbreak representations evade safety transformations.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Its account of instruction-following alignment through supervised fine-tuning and human-feedback reinforcement establishes the training paradigm whose internal effects the source studies.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Its helpful-and-harmless RLHF framework clarifies the preference-alignment objectives and robustness tensions examined by the source.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Its universal adversarial-suffix method provides the attack setting needed to appreciate the source’s mechanistic account of how aligned representations can be bypassed.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). It extends the source’s mechanistic diagnosis by applying adaptive, model-specific attacks that directly exploit the fragile safety behaviors identified there.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). It continues the source’s robustness analysis with an improved optimization-based jailbreak designed to overcome refusal recovery and achieve more reliable attack success.
- Paper: FlipAttack: Jailbreak LLMs via Flipping, Yue Liu 0008 et al. (2025). It generalizes the source’s evasion findings to a black-box flipping attack that bypasses safety mechanisms without access to internal model parameters.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It develops a complementary defense by enforcing multiple alignment objectives at decoding time rather than relying solely on the fragile fine-tuning transformation analyzed by the source.
- Paper: How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models, Dun Li Chan et al. (2026). It extends the source’s internal-sensitivity perspective by tracing how diverse natural and adversarial perturbations propagate through hidden representations and attention mechanisms.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). It carries the source’s concern with surface-level safety failures into multi-stage agent systems, examining how unverified inferred intent propagates across processing stages.
