What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Samyak JainEkdeep Singh LubanaKemal OksuzTom JoyPhilip TorrAmartya SanyalPuneet K. Dokania

article2024NeurIPS51 citations

Reveals how safety fine-tuning alters transformer MLP weights to project unsafe inputs into their null space, explaining why adversarial jailbreaks succeed by mimicking safe activation patterns to bypass this mechanism.

Listen

Safety fine-tuning is an essential step in deploying Large Language Models (LLMs), designed to align model outputs with human preferences and prevent malicious misuse. However, existing alignment techniques remain highly vulnerable to adversarial exploits, such as prompt-based jailbreaks. The article aims to uncover the internal mechanisms that enable safety fine-tuning to function, why these mechanisms fail under adversarial attacks, and how models differentiate between safe and unsafe prompts.

To evaluate these questions systematically, the article introduces a controlled synthetic data generation framework that models inputs into tasks (operators, such as "design") and concepts (operands, such as "cycle" versus "bomb"). Using this framework, the article evaluates three prominent safety alignment methods: Supervised Safety Fine-Tuning (SSFT), Direct Preference Optimization (DPO), and machine unlearning. The authors evaluate model behavior across internal activation representations, weight parameter changes, and mathematical sensitivity (local Lipschitzness). To ensure real-world applicability, the core findings are corroborated on production-grade open LLMs, including Llama-2 and Llama-3 models across safe and unsafe natural language instructions.

The article yields four major findings. First, safety fine-tuning modifies a model's internal multilayer perceptron (MLP) weights through a sparse, highly localized transformation that isolates unsafe prompts into separate feature clusters, leaving safe prompts largely unaffected. Second, this weight change acts primarily on a low-rank subspace, projecting unsafe representations into the null space of the original model weights so the model produces refusal tokens. Third, alignment substantially reduces model output sensitivity (local Lipschitzness) for unsafe prompts while increasing it for safe ones, making stronger protocols like DPO and unlearning more resistant to naive perturbations than SSFT. Fourth, successful adversarial and jailbreak inputs—such as those combining competing objectives or out-of-distribution formatting—evade detection because their internal representations bypass the specialized safety transformation, clustering closely with safe inputs and inducing attack success rates of over 90% across several tested configurations.

These findings demonstrate that current safety fine-tuning creates a fragile, localized patch rather than an integrated conceptual understanding of harmfulness. Because the safety transformation acts only on a narrow set of unsafe representations, minor semantic or contextual prompt changes readily disguise malicious inputs as safe queries. This introduces substantial compliance and safety risks for organizations relying solely on standard alignment to prevent LLM misuse.

To improve model resilience, the article highlights weight extrapolation along the safety transformation direction as a practical intervention. Extrapolating this direction beyond standard fine-tuning weights significantly boosts refusal robustness against jailbreaks without degrading baseline utility on safe tasks. For immediate engineering and deployment, organizations should combine preference optimization (such as DPO) with multidirectional representation monitoring and explicit defense-in-depth measures, rather than relying exclusively on post-training alignment.

While the analytical conclusions are supported by strong evidence across synthetic benchmarks and Llama architectures, the primary limitation is that mechanistic observations are concentrated within feed-forward MLP layers and tested under specific synthetic grammar approximations. Decision-makers should treat these insights as a compelling explanation of alignment fragility, while continuing comprehensive empirical auditing across diverse, real-world conversational contexts.

arXiv: 2407.10264
Cover for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Abstract

Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods—supervised safety fine-tuning, direct preference optimization, and unlearning—and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. Code is available at https://github.com/fiveai/understanding_safety_finetuning.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 A Synthetic Controlled Set-up for Safety Fine-tuning
  • 3.1 Data generation for inducing instruction following behavior
  • 3.2 Data generation for safety fine-tuning
  • 3.3 Data generation for jailbreak and adversarial attacks
  • 4 Investigating the Effect of Safety Fine-tuning
  • 4.1 Clustering of safe versus unsafe samples’ activations: Analyzing activation space
  • 4.2 What drives the clustering of safe and unsafe samples: Analyzing parameter changes
  • 4.3 Impact of safety fine-tuning on the sensitivity of the learned model
  • 5 Evading the Safety Mechanism: Jailbreak and Adversarial Inputs
  • 6 Conclusion
  • Acknowledgements
  • References
  • APPENDICES
  • A Additional Background
  • B Further Details on the Experimental Setup
  • B.1 Further Details on the Synthetic Setup based on PCFG
  • B.2 Further Details on Real World Experiments based on Llama
  • C Further Analyses to Understand Safety Fine-tuning
  • C.1 Analyzing how the impact of transformation propagates over the layers
  • C.2 Additional Results on Llama-2
  • C.3 Additional Results on the synthetic setup
  • D Additional Results Using Interventions
  • E Limitations and Societal Impact
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Mechanistic Refusal via Specialized Null-Space Projection of MLP Weight Updates

    empirical result

    Safety fine-tuning alters transformer Multi-Layer Perceptron (MLP) parameters via a sparse, specialized transformation ΔW=WST−WIT\Delta W = W_{\text{ST}} - W_{\text{IT}}, where WITW_{\text{IT}} and WSTW_{\text{ST}} represent the first MLP layer weights of the LL-th transformer block before and after safety alignment, respectively.

    This parameter update acts through two distinct geometric mechanisms:

    1. Column-Space Alignment with Left Null Space: The column space C(ΔW)\mathcal{C}(\Delta W) is strongly aligned with the left null-space N(WIT⊤)\mathcal{N}(W_{\text{IT}}^\top) rather than C(WIT)\mathcal{C}(W_{\text{IT}}). For left singular vectors {u~i}i=1t\{\tilde{u}_i\}_{i=1}^t of ΔW\Delta W and the column-space projection matrix P=∑i=1ruiui⊤P = \sum_{i=1}^r u_i u_i^\top of WITW_{\text{IT}}, the projection magnitude sin⁡(θi)\sin(\theta_i) onto N(WIT⊤)\mathcal{N}(W_{\text{IT}}^\top) is large across layers. Consequently, ΔW\Delta W and WITW_{\text{IT}} are approximately mutually orthogonal, producing activations distinct from instruction-following computations.

    2. Selective Row-Space Activation on Unsafe Pre-activations: When evaluated on normalized pre-activations aa for the first predicted output token, unsafe inputs exhibit a large projection σivi⊤a\sigma_i v_i^\top a along the top right singular vectors {vi}\{v_i\} of ΔW\Delta W. Safe inputs yield near-zero projection on these basis vectors. Individual neurons aligned with the top right singular vector v1v_1 selectively activate on unsafe inputs to alter their activation norm.

    Together, ΔW\Delta W selectively projects unsafe representations into the null space of the original model parameters while leaving safe inputs unperturbed.

  2. Knowl 2 — Synthetic Grammatical Framework for Disentangling Safety Fine-Tuning and Jailbreak Types

    model/method

    To systematically analyze safety alignment mechanisms and attack vectors, language model inputs are formulated as a composition of operators and operands:

    X={fj∘fi,T,O}X = \{f_j \circ f_i, T, O\}

    where:

    • fi,fj∈Ff_i, f_j \in \mathcal{F} are task tokens representing operators, defined as bijective mappings f:V→Vf: \mathcal{V} \to \mathcal{V} on a vocabulary V\mathcal{V}.
    • TT is a sequence of text tokens (operands) of length 15–2515\text{--}25 sampled via a Probabilistic Context-Free Grammar PCFG(γ,T,NT,R,P)\text{PCFG}(\gamma, \mathcal{T}, \mathcal{NT}, R, P) across production rules RR and probabilities PP.
    • O=fj(fi(T))O = f_j(f_i(T)) is the expected sequence of output tokens for safe inputs.

    Contextual Safety Assignment

    Safety depends contextually on the pairing of task tokens and text tokens. Non-terminal nodes at level ls=3l_s = 3 are partitioned into safe-dominant nodes A⊂NTls\mathcal{A} \subset \mathcal{NT}^{l_s} and unsafe-dominant nodes B⊂NTls\mathcal{B} \subset \mathcal{NT}^{l_s}. Associated task token sets satisfy:

    ∣FAs∣>∣FAu∣,∣FBs∣<∣FBu∣,FAu⊂FBu,FBs⊂FAs|\mathcal{F}_\mathcal{A}^s| > |\mathcal{F}_\mathcal{A}^u|, \quad |\mathcal{F}_\mathcal{B}^s| < |\mathcal{F}_\mathcal{B}^u|, \quad \mathcal{F}_\mathcal{A}^u \subset \mathcal{F}_\mathcal{B}^u, \quad \mathcal{F}_\mathcal{B}^s \subset \mathcal{F}_\mathcal{A}^s

    An input is labeled unsafe if sampled from FAu\mathcal{F}_\mathcal{A}^u or FBu\mathcal{F}_\mathcal{B}^u corresponding to its operand source, and the model is trained to output a fixed null token (e.g., token 'a'). Safe inputs are trained to complete OO.

    Formalization of Attack Types

    • Competing Objectives via Task (JB-CO-Task): One safe and one unsafe task token are combined (fi∈FNs,fj∈FNuf_i \in \mathcal{F}^s_N, f_j \in \mathcal{F}^u_N).
    • Competing Objectives via Text (JB-CO-Text): Operands are generated starting from the lowest common ancestor of A\mathcal{A} and B\mathcal{B} with conflicting task tokens.
    • Mismatched Generalization (JB-MisGen): Task tokens are drawn from an out-of-distribution set TOOD\mathcal{T}_{\text{OOD}} having identical bijective semantics to training tokens but different token representations.
    • Continuous Adversarial Attacks (Adv): kk continuous soft prompt embeddings constrained by ℓ2\ell_2-norm ≤1\le 1 are appended to the input and optimized by targeted gradient descent for 1010 iterations.
  3. Knowl 3 — Bypassing Alignment via Activation-Space Similarity to Safe Samples

    empirical result

    Jailbreak prompts (competing objectives and mismatched generalization) and continuous adversarial soft-prompt attacks bypass safety refusal mechanisms because their intermediate representations evade the safety transformation ΔW\Delta W:

    1. Activation Feature Space: As jailbreak and adversarial attack strength increases, the intermediate activations of attacked inputs move closer to the safe activation cluster rather than the unsafe cluster, reducing the geometric separation measured by the cluster distance metric τ\tau.
    2. Parameter Space Interaction: The pre-activations aa of successful jailbreaks and adversarial examples produce near-zero projection along the row space R(ΔW)\mathcal{R}(\Delta W) (specifically on the singular vectors σivi⊤a\sigma_i v_i^\top a). Their alignment with R(ΔW)\mathcal{R}(\Delta W) matches that of benign safe samples rather than native unsafe samples.
    3. Function-Space Sensitivity: The local Lipschitz sensitivity distribution for attacked inputs shifts toward the higher sensitivity values characteristic of safe instructions.

    Because the learned refusal transformation ΔW\Delta W activates exclusively on representations with substantial row-space alignment, adversarially altered inputs pass through the pre-trained instruction-following subnetwork without triggering the null-space projection refusal transformation.

  4. Knowl 4 — Activation Space Separation and Rank Collapse of Unsafe Representations

    empirical result

    Safety fine-tuning (Supervised Safety Fine-Tuning, Direct Preference Optimization, and Unlearning) segregates hidden activations into distinct safe and unsafe clusters in deeper transformer MLP layers.

    Analyzing the empirical covariance matrices of post-MLP activations at layer LL:

    ΣU=∑x∈DU(a^Lo(x)[q]−μLU)(a^Lo(x)[q]−μLU)⊤,ΣS=∑x∈DS(a^Lo(x)[q]−μLS)(a^Lo(x)[q]−μLS)⊤\Sigma^U = \sum_{x \in \mathcal{D}_U} \left(\hat{a}_L^o(x)[q] - \mu_L^U\right)\left(\hat{a}_L^o(x)[q] - \mu_L^U\right)^\top, \quad \Sigma^S = \sum_{x \in \mathcal{D}_S} \left(\hat{a}_L^o(x)[q] - \mu_L^S\right)\left(\hat{a}_L^o(x)[q] - \mu_L^S\right)^\top

    reveals that as safety training progresses, the feature spread of unsafe inputs undergoes severe empirical rank collapse. The top singular value σ1(ΣU)\sigma_1(\Sigma^U) grows to constitute approximately 62%62\% of the total nuclear norm of ΣU\Sigma^U, whereas for safe inputs σ1(ΣS)\sigma_1(\Sigma^S) constitutes only 12%12\% of the nuclear norm of ΣS\Sigma^S.

    Unsafe representations are thereby compressed into a single dominant geometric refusal direction, while the dimensionality and structure of safe representations remain largely unchanged.

  5. Knowl 5 — Local Lipschitz Sensitivity Reduction for Unsafe Inputs

    empirical result

    The sensitivity of an aligned model with parameters θ\theta is probed using the local Lipschitz constant of the logit corresponding to the most confident token prediction at sequence index kk:

    Lipf^(x)=∥∇xf^θ(x)∥2,where f^θ(x)=max⁡jhθ(x)[k](j)\text{Lip}_{\hat{f}}(x) = \|\nabla_x \hat{f}_\theta(x)\|_2, \quad \text{where } \hat{f}_\theta(x) = \max_j h_\theta(x)[k](j)

    Safety fine-tuning dramatically reduces Lipf^(x)\text{Lip}_{\hat{f}}(x) for unsafe inputs while slightly increasing it for safe inputs:

    • The reduction occurs because preferred targets for unsafe queries have minimal output entropy and low semantic variability (e.g., constant null or refusal tokens).
    • Stronger alignment objectives (Direct Preference Optimization and Machine Unlearning) produce a much steeper drop in local Lipschitzness on unsafe inputs than Supervised Safety Fine-Tuning (SSFT).
    • Lower local sensitivity on unsafe samples directly correlates with greater empirical resistance to gradient-based adversarial attacks and jailbreaks.
  6. Knowl 6 — Safety Enhancement and Robustification via Linear Weight Extrapolation

    model/method

    Model safety and refusal robustness can be enhanced post-hoc by linearly extrapolating weights along the safety update direction ΔW=WST−WIT\Delta W = W_{\text{ST}} - W_{\text{IT}}:

    WITα=WIT+αΔWW_{\text{IT}}^\alpha = W_{\text{IT}} + \alpha \Delta W

    where α∈[0,1.5]\alpha \in [0, 1.5] is an extrapolation scalar.

    Key behaviors under linear weight manipulation include:

    • Linear Connectivity: Instruction fine-tuned and safety fine-tuned models, as well as models trained with different safety objectives (SSFT, DPO, Unlearning), reside in a linearly connected loss basin.
    • Extrapolation Benefits (α>1\alpha > 1): For weaker alignment protocols such as SSFT, setting α∈(1.0,1.5]\alpha \in (1.0, 1.5] substantially widens the separation between safe and unsafe activation clusters, decreases empirical rank in ΣU\Sigma^U, and reduces jailbreak attack success rates from 100%100\% to lower vulnerability levels while retaining full accuracy on benign safe prompts.
    • Transferability Across Protocols: Applying ΔW\Delta W derived from stronger protocols (e.g., Unlearning) to an SSFT checkpoint improves the SSFT model's resistance to jailbreak attacks.
  7. Knowl 7 — Activation Space Cluster Separation Metric

    equation

    Let aLo(x)[i]a_L^o(x)[i] be the LL-th layer output activation corresponding to token index ii of input sequence xx. The sequence-averaged activation for the qq-th output token starting from the final text token index kk is:

    a^Lo(x)[q]=1q−1∑i=kq+k−1aLo(x)[i]\hat{a}_L^o(x)[q] = \frac{1}{q - 1} \sum_{i=k}^{q+k-1} a_L^o(x)[i]

    Given mean activations μLS=1∣DS∣∑x∈DSa^Lo(x)[q]\mu_L^S = \frac{1}{|\mathcal{D}_S|} \sum_{x \in \mathcal{D}_S} \hat{a}_L^o(x)[q] and μLU=1∣DU∣∑x∈DUa^Lo(x)[q]\mu_L^U = \frac{1}{|\mathcal{D}_U|} \sum_{x \in \mathcal{D}_U} \hat{a}_L^o(x)[q] over datasets of safe inputs DS\mathcal{D}_S and unsafe inputs DU\mathcal{D}_U, the sample cluster score τ\tau is defined as:

    τ(x,μLS,μLU)=∥a^Lo(x)[q]−μLU∥2−∥a^Lo(x)[q]−μLS∥2\tau(x, \mu_L^S, \mu_L^U) = \|\hat{a}_L^o(x)[q] - \mu_L^U\|_2 - \|\hat{a}_L^o(x)[q] - \mu_L^S\|_2

    where τ>0\tau > 0 indicates proximity to the safe cluster and τ<0\tau < 0 indicates proximity to the unsafe cluster.

  8. Knowl 8 — Safety Performance Comparison Across Fine-Tuning Protocols and Jailbreak Attacks

    data/table

    The table reports instruction-following accuracy (Instruct) and refusal accuracy (Null) on clean sets and under three jailbreak attack categories (JB-CO-Task, JB-CO-Text, JB-MisGen) on 6-layer minGPT models evaluated with 1K test samples per category across two learning rates (ηM=10−4\eta_M = 10^{-4} and ηS=10−5\eta_S = 10^{-5}):

    Protocol Learning Rate Safe (Instruct) Unsafe (Null) Unsafe (Instruct) JB-CO-Task (Instruct) JB-CO-Text (Instruct) JB-MisGen (Instruct)
    Unlearning ηM\eta_M 99.8% 99.9% 5.0% 27.1% 95.2% 92.3%
    Unlearning ηS\eta_S 99.7% 99.9% 31.2% 51.2% 98.3% 98.5%
    DPO ηM\eta_M 98.6% 99.6% 11.8% 31.5% 93.5% 93.6%
    DPO ηS\eta_S 98.7% 100.0% 40.7% 56.1% 97.2% 96.1%
    SSFT ηM\eta_M 99.9% 99.8% 51.6% 88.1% 100.0% 100.0%
    SSFT ηS\eta_S 99.7% 100.0% 72.8% 92.5% 100.0% 100.0%

    Unlearning and DPO with medium learning rate ηM\eta_M provide the strongest safety retention on unsafe inputs (5.0%5.0\% and 11.8%11.8\% instruction leakage, respectively), whereas SSFT suffers from high leakage (51.6%–72.8%51.6\%\text{--}72.8\%). JB-CO-Text and JB-MisGen represent the most potent jailbreak formats, inducing instruction execution in ≥93.5%\ge 93.5\% of cases across all protocols, with SSFT reaching 100.0%100.0\% vulnerability.

  9. Knowl 9 — Limitations in Architecture Coverage and Explanatory Uniformity Across Jailbreak Types

    limitation

    The mechanistic analysis has two explicit limitations identified by the authors:

    1. Restricted Direct Analysis on Proprietary Weight Diffs: Because official instruction fine-tuned (intermediate) model weights are not publicly released for models like Llama-2 (only pre-trained Llama-2 7B and safety-aligned Llama-2-Chat 7B are available), weight update decomposition ΔW=WST−WIT\Delta W = W_{\text{ST}} - W_{\text{IT}} cannot be directly computed for production LLMs without conflating instruction tuning and safety tuning.
    2. Uniformity of Jailbreak Evasion Mechanics: The framework does not reveal mechanistically distinct internal pathways between different jailbreak modalities (e.g., competing objectives vs. mismatched generalization); all successful attacks manifest similarly as an evasion of ΔW\Delta W row-space projection and a collapse back into the safe activation distribution.

Coverage note — None was omitted; all contributed mechanisms, synthetic framework formulations, empirical findings, and intervention analyses are represented.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 1, context-free grammar. arXiv preprint arXiv:2305.13673, 2023.
  3. 3.Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024.
  4. 4.Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024.
  5. 5.Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction. arXiv preprint arXiv:2406.11717, 2024.
  6. 6.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  7. 7.Sarah Ball, Frauke Kreuter, and Nina Rimsky. Understanding jailbreak success: A study of latent space dynamics in large language models. arXiv preprint arXiv:2406.09289, 2024.
  8. 8.Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, et al. Managing ai risks in an era of rapid progress. arXiv preprint arXiv:2310.17688, 2023.
  9. 9.Christopher M. Bishop. Pattern Recognition and Machine Learning (Information Science and Statistics). Springer-Verlag, Berlin, Heidelberg, 2006. ISBN 0387310738.
  10. 10.S ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  11. 11.Llama 3 Model Card. AI@Meta, 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_ CARD.md.
  12. 12.Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, et al. Are aligned neural networks adversari- ally aligned? arXiv preprint arXiv:2306.15447, 2023.
  13. 13.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2023.
  14. 14.Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tram er, Hamed Hassani, and Eric Wong. Jailbreakbench: An open robustness benchmark for jailbreaking large language models, 2024.
  15. 15.Eugene Charniak. Statistical techniques for natural language parsing. AI Mag., 18:33–44, 1997. URL https://api.semanticscholar.org/CorpusID:11071483.
  16. 16.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  17. 17.Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of universality: Reverse engineering how networks learn group operations. In International Conference on Machine Learning, pp. 6243–6267. PMLR, 2023.
  18. 18.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  19. 19.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural networks, 107:3–11, 2018.
  20. 20.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1:1, 2021.
  21. 21.Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal. Corrective machine unlearning. arXiv preprint arXiv:2402.14015, 2024.
  22. 22.Michael Hahn and Navin Goyal. A theory of emergent in-context learning as implicit structure induction. arXiv preprint arXiv:2303.07971, 2023.
  23. 23.Matthias Hein and Maksym Andriushchenko. Formal guarantees on the robustness of a classifier against adversarial manipulation. Advances in neural information processing systems, 30, 2017.
  24. 24.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  25. 25.Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. Baseline defenses for adversarial attacks against aligned language models. arXiv preprint arXiv:2309.00614, 2023a.
  26. 26.Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rockt aschel, and David Scott Krueger. Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. arXiv preprint arXiv:2311.12786, 2023b.
  27. 27.Andrej Karpathy. MinGPT, 2020. Github link. https://github.com/karpathy/minGPT/tree/master.
  28. 28.Bjarne Knudsen and Jotun Hein. Rna secondary structure prediction using stochastic context-free grammars and evolutionary history. Bioinformatics, 15 6:446–54, 1999. URL https://api.semanticscholar.org/ CorpusID:5971132.
  29. 29.Suhas Kotha, Jacob Mitchell Springer, and Aditi Raghunathan. Understanding catastrophic forgetting in language models via implicit inference. arXiv preprint arXiv:2309.10105, 2023.
  30. 30.Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea. A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity, 2024.
  31. 31.Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al. The wmdp benchmark: Measuring and reducing malicious use with unlearning. arXiv preprint arXiv:2403.03218, 2024.
  32. 32.Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al. Rethinking machine unlearning for large language models. arXiv preprint arXiv:2402.08787, 2024.
  33. 33.Ekdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Krueger, and Hidenori Tanaka. Mechanistic Mode Connectivity, 2022. Comment: 39 pages.
  34. 34.Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell. Eight methods to evaluate robust unlearning in llms. arXiv preprint arXiv:2402.16835, 2024.
  35. 35.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJzIBfZAb.
  36. 36.Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter. Tofu: A task of fictitious unlearning for llms. arXiv preprint arXiv:2401.06121, 2024.
  37. 37.Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119, 2023.
  38. 38.Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang. What is being transferred in transfer learning?, 2021.
  39. 39.Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen. A survey of machine unlearning. arXiv preprint arXiv:2209.02299, 2022.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  41. 41.Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. Fine-tuning enhances existing mechanisms: A case study on entity tracking. arXiv preprint arXiv:2402.14811, 2024.
  42. 42.Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S Yu. Multilingual large language model: A survey of resources, taxonomy and frontiers. arXiv preprint arXiv:2404.04925, 2024.
  43. 43.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. OpenAI, 2018.
  44. 44.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  45. 45.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  46. 46.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  47. 47.Rahul Ramesh, Mikail Khona, Robert P Dick, Hidenori Tanaka, and Ekdeep Singh Lubana. How capable can a transformer become? a study on synthetic, interpretable tasks. arXiv preprint arXiv:2311.12997, 2023.
  48. 48.Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda, Atoosa Chegini, and Soheil Feizi. Fast adversarial attacks on language models in one gpu minute, 2024.
  49. 49.Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, et al. Rainbow teaming: Open-ended generation of diverse adversarial prompts. arXiv preprint arXiv:2402.16822, 2024.
  50. 50.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  51. 51.Amartya Sanyal, Philip HS Torr, and Puneet K Dokania. Stable rank normalization for improved generalization in neural networks and gans. arXiv preprint arXiv:1906.04659, 2019.
  52. 52.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  53. 53.Gilbert Strang. Introduction to Linear Algebra. Wellesley-Cambridge Press, Wellesley, MA, fourth edition, 2009. ISBN 9780980232714 0980232716 9780980232721 0980232724 9788175968110 8175968117.
  54. 54.Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024.
  55. 55.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  56. 56.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ee Lacroix, Baptiste Rozi ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  57. 57.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  58. 58.Nilesh Tripuraneni, Michael Jordan, and Chi Jin. On the theory of transfer learning: The importance of task diversity. Advances in neural information processing systems, 33:7852–7862, 2020.
  59. 59.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail?, 2023.
  60. 60.Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions. arXiv preprint arXiv:2402.05162, 2024.
  61. 61.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  62. 62.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  63. 63.Eric Wong and Zico Kolter. Provable defenses against adversarial examples via the convex outer adversarial polytope. In International conference on machine learning, pp. 5286–5295. PMLR, 2018.
  64. 64.Chujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang, and Nanyun Peng. Weak-to-strong extrapolation expedites alignment. arXiv preprint arXiv:2404.16792, 2024.
  65. 65.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Jain, S., et al. “What Makes and Breaks Safety Fine-tuning? A Mechanistic Study”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 93406–78, https://proceedings.neurips.cc/paper_files/paper/2024/file/a9bef53eb7b0e5950d4f2d9c74a16006-Paper-Conference.pdf.
APA
Jain, S., Lubana, E. S., Oksuz, K., Joy, T., Torr, P., Sanyal, A., & Dokania, P. K. (2024). What Makes and Breaks Safety Fine-tuning? A Mechanistic Study. Advances in Neural Information Processing Systems, 37, 93406–93478. https://proceedings.neurips.cc/paper_files/paper/2024/file/a9bef53eb7b0e5950d4f2d9c74a16006-Paper-Conference.pdf
Chicago
Jain, S., E. S. Lubana, K. Oksuz, et al. 2024. “What Makes and Breaks Safety Fine-tuning? A Mechanistic Study”. Advances in Neural Information Processing Systems 37: 93406–78. https://proceedings.neurips.cc/paper_files/paper/2024/file/a9bef53eb7b0e5950d4f2d9c74a16006-Paper-Conference.pdf.
Harvard
Jain, S. et al. (2024) “What Makes and Breaks Safety Fine-tuning? A Mechanistic Study”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 93406–93478. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/a9bef53eb7b0e5950d4f2d9c74a16006-Paper-Conference.pdf.
Vancouver
1. Jain S, Lubana ES, Oksuz K, Joy T, Torr P, Sanyal A, Dokania PK (2024) What Makes and Breaks Safety Fine-tuning? A Mechanistic Study. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 93406–93478

BibTeX

@inproceedings{jain2024what,
  title = {What Makes and Breaks Safety Fine-tuning? A Mechanistic Study},
  author = {Jain, Samyak and Lubana, Ekdeep S. and Oksuz, Kemal and Joy, Tom and Torr, Philip and Sanyal, Amartya and Dokania, Puneet K.},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {93406-93478},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/a9bef53eb7b0e5950d4f2d9c74a16006-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors