Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Samyak JainRobert KirkEkdeep Singh LubanaRobert P. DickHidenori TanakaTim RocktäschelEdward GrefenstetteDavid Scott Krueger

article2024ICLR107 citations

Demonstrates through mechanistic interpretability that fine-tuning merely adds superficial, easily reversible functional wrappers over pretrained capabilities rather than altering them, revealing why safety alignment can be unintentionally undone with minimal downstream adaptation.

Listen

Adapting large, pretrained artificial intelligence models via fine-tuning is currently the primary industry standard for customizing systems to specific downstream tasks and attempting to eliminate unsafe or toxic behaviors. However, organizations deploy these adapted models without a clear mechanistic understanding of whether fine-tuning fundamentally alters, removes, or rewires the internal capabilities learned during pretraining, or merely suppresses their outward expression.

The article systematically evaluates how fine-tuning alters a neural network's underlying capabilities and demonstrates whether these capabilities are genuinely modified or merely obscured by shallow transformations.

To conduct this evaluation in a controlled and rigorous manner, the analysis employed three complementary experimental settings: compiled transformer models using the Tracr framework with perfectly defined algorithmic functions, models trained on procedurally generated formal grammars, and realistic 91-million-parameter language models trained on the TinyStories dataset. Across these setups, the researchers utilized a diverse suite of mechanistic interpretability tools, including selective network pruning, linear probing of internal representations, attention visualization, and a reverse fine-tuning methodology to assess whether prior capabilities could be easily recovered.

The article establishes three core findings. First, fine-tuning rarely modifies or deletes underlying model capabilities; instead, it typically trains a minimal, superficial transformation—termed a wrapper—in the network's final layers while leaving internal capabilities intact. Second, these learned wrappers are highly localized, meaning that pruning as few as 5 to 15 neurons or a handful of weights readily restores the model's original pretraining behavior. Third, suppressed or nominally deleted capabilities can be revived rapidly and with high sample efficiency; when subjected to reverse fine-tuning, models fully recovered original capabilities within 30 to 300 iterations, far faster than baseline models learning those tasks from scratch.

These findings carry critical operational and risk implications. They reveal that standard safety alignment and fine-tuning techniques create a false sense of security by merely masking undesirable or hazardous behaviors rather than excising them. Downstream practitioners can unintentionally dismantle safety guardrails and re-expose latent, harmful capabilities through simple, standard fine-tuning on superficially unrelated datasets. Consequently, alignment strategies that rely solely on parameter fine-tuning leave systems highly vulnerable to jailbreaks, accidental regression, and compliance failures.

Organizations and developers should not treat standard fine-tuning or superficial alignment protocols as permanent guarantees of model safety. When deploying models in high-stakes environments, technical teams must implement defense-in-depth safeguards, such as continuous runtime monitoring, input-output filtering, and rigorous penetration testing. In addition, research and development efforts must shift toward developing techniques capable of genuinely deleting internal mechanisms rather than relying on standard gradient-based fine-tuning.

While the findings demonstrate high consistency across multiple mechanistic interpretability tools, the core controlled analyses rely on synthetic grammars, compiled transformers, and relatively small language models. Decision-makers should exercise appropriate caution when extrapolating these precise dynamics to frontier-scale foundation models operating on open-domain web data, where further validation is necessary.

Cover for Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Abstract

Fine-tuning large pre-trained models has become the de facto strategy for developing both task-specific and general-purpose machine learning systems, including developing models that are safe to deploy. Despite its clear importance, there has been minimal work that explains how fine-tuning alters the underlying capabilities learned by a model during pretraining: does fine-tuning yield entirely novel capabilities or does it just modulate existing ones? We address this question empirically in synthetic, controlled settings where we can use mechanistic interpretability tools (e.g., network pruning and probing) to understand how the model’s underlying capabilities are changing. We perform an extensive analysis of the effects of fine-tuning in these settings, and show that: (i) fine-tuning rarely alters the underlying model capabilities; (ii) a minimal transformation, which we call a ‘wrapper’, is typically learned on top of the underlying model capabilities, creating the illusion that they have been modified; and (iii) further fine-tuning on a task where such “wrapped capabilities” are relevant leads to sample-efficient revival of the capability, i.e., the model begins reusing these capabilities after only a few gradient steps. This indicates that practitioners can unintentionally remove a model’s safety wrapper merely by fine-tuning it on a, e.g., superficially unrelated, downstream task. We additionally perform analysis on language models trained on the TinyStories dataset to support our claims in a more realistic setup.

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORK
  • 3 DEFINING OUR NOTION OF CAPABILITIES
  • 4 BUILDING CAPABLE MODELS: TRACR AND PCFGS
  • 5 EXPERIMENTS: MECHANISTIC ANALYSIS OF FINE-TUNING
  • 5.1 RESULTS
  • 5.2 VALIDATING OUR HYPOTHESES ON TINYSTORIES
  • 6 CONCLUSION
  • ACKNOWLEDGEMENTS
  • AUTHORS' CONTRIBUTIONS
  • REFERENCES
  • A ORGANIZATION OF APPENDIX
  • B ADDITIONAL DETAILS ON DATASETS
  • B.1 TRACR DETAILS
  • B.2 PCFG
  • B.3 TINYSTORIES INSTRUCT
  • C DETAILS ON TRAINING AND EVALUATION
  • C.1 TRACR
  • C.2 PCFG
  • D MECHANISTIC INTERPRETABILITY TOOLS SETUP
  • E ADDITIONAL RESULTS
  • E.1 FINE-TUNING IN PRESENCE OF SOME PRE-TRAINING DATA
  • E.2 JAILBREAKING ANALYSIS
  • E.3 SAMPLE EFFICIENCY ANALYSIS FOR REVERSE FINE-TUNING
  • E.4 REVERSE FINE-TUNING A MORE SAFETY-ORIENTED FINE-TUNING PROTOCOL
  • F DETAILS AND RESULTS ON TINYSTORIES EXPERIMENTS
  • F.1 MODEL TRAINING
  • F.2 EVALUATION METRICS
  • F.3 RESULTS
  • G ADDITIONAL TRACR RESULTS
  • G.1 BEHAVIORAL RESULTS ON FINE-TUNING
  • G.2 COUNTER RESULTS
  • G.3 MAX IDENTIFIER RESULTS
  • H ADDITIONAL PCFG RESULTS
  • H.1 PRUNING ANALYSIS
  • H.2 PROBING ANALYSIS

Knowls

  1. Knowl 1 — Formal Definition of Model Capability

    definition

    Let DPTD_{PT} denote a pretraining dataset sampled from a data distribution PXP_X over an input domain XX. The domain XX factorizes into two sub-domains XIX_I and XDX_D, such that every input sample x∈Xx \in X is represented as a tuple (xi∈XI,xd∈XD)(x_i \in X_I, x_d \in X_D), where xix_i acts as a task or capability identifier (with iCi_C identifying capability CC) that specifies how the model should process the data variable xdx_d.

    Pretraining on DPTD_{PT} yields an LL-layer neural network model M:X→YM: X \to Y. Let Readl(M(⋅))\text{Read}_l(M(\cdot)) denote the readout operation obtained by training a linear classifier on the intermediate representations at layer l∈{1,…,L}l \in \{1, \dots, L\} of model MM using DPTD_{PT}.

    Let fC:XD→YCf_C: X_D \to Y_C be a surjective mapping defining the target task for capability CC, where YC⊆YY_C \subseteq Y. Let XC⊂XX_C \subset X be a sub-domain where all inputs share the same capability identifier xi=iCx_i = i_C. The model MM is defined to possess a capability CC if: ∀x∈XC,∃l≤Lsuch thatReadl(M(x))=fC(xd)\forall x \in X_C, \quad \exists l \le L \quad \text{such that} \quad \text{Read}_l(M(x)) = f_C(x_d)

  2. Knowl 2 — Strong vs. Weak Capability Relevance and Wrappers

    definition

    Let DFT∼PXFT,ED_{FT} \sim P_X^{FT, E} denote a downstream fine-tuning dataset with ∣DFT∣≪∣DPT∣|D_{FT}| \ll |D_{PT}|, where PXFT,EP_X^{FT, E} is an empirical distribution capturing a subset of the support of the true downstream distribution PXFTP_X^{FT}. Suppose fine-tuning adapts a pretrained capability fCf_C into a composite mapping g∘fCg \circ f_C that produces the correct output for all x∼PXFT,Ex \sim P_X^{FT, E}.

    • Strong Relevance: If the composite mapping g∘fCg \circ f_C produces the correct output for all x∼PXFTx \sim P_X^{FT} across the full downstream distribution (i.e., extracting causally relevant features), the capability CC is strongly relevant to the fine-tuning task.
    • Weak Relevance: If g∘fCg \circ f_C produces correct outputs only on PXFT,EP_X^{FT, E} by exploiting spurious correlations that do not hold across the true distribution PXFTP_X^{FT}, the capability CC is weakly relevant.

    When a weakly relevant capability is adapted, fine-tuning typically learns a wrapper gg, which is a localized transformation implemented in the later layers/neurons of the model. The resulting composite function g∘fCg \circ f_C is termed a wrapped capability.

  3. Knowl 3 — Reverse Fine-Tuning Protocol

    model/method

    Reverse fine-tuning (reFT\text{reFT}) is a diagnostic methodology designed to determine whether fine-tuning genuinely deletes or permanently modifies underlying pretraining capabilities, or whether it merely masks them behind a wrapper.

    The protocol proceeds in three phases:

    1. A base model is pretrained on a pretraining distribution DPTD_{PT} to establish a capability CC.
    2. The model is fine-tuned on a downstream dataset DFTD_{FT} using standard learning rates until behavioral performance on capability CC is suppressed or modified.
    3. The fine-tuned model is subsequently fine-tuned in the reverse direction on a small dataset sampled from the original pretraining distribution DPTD_{PT} associated with capability CC, typically using a learning rate equal to or smaller than that used during downstream fine-tuning.

    If the model recovers high accuracy on capability CC within very few gradient steps or iterations (demonstrating high sample efficiency relative to training from scratch or fine-tuning an unrelated model), the capability CC is confirmed to have persisted mechanistically in the model's intermediate representations throughout downstream fine-tuning.

  4. Knowl 4 — First-Order Gradient-Weight Saliency Pruning for Wrapper Identification

    algorithm

    To mechanistically locate and isolate wrappers learned during fine-tuning, single-step first-order saliency pruning removes the weights or neurons with the highest dot product between their parameter values and the gradient of the pretraining loss.

    Input: Fine-tuned model fθf_\theta parameterized by θ∈RN\theta \in \mathbb{R}^N, pretraining input string XX prepended with pretraining task family token OPTO_{PT}, ground truth pretraining label yy, number of parameters to prune KK
    Output: Pruned model parameters θpruned\theta_{pruned}
    out ←fθ(OPT∘X)\leftarrow f_\theta(O_{PT} \circ X)
    loss ←CrossEntropyLoss(out,y)\leftarrow \text{CrossEntropyLoss}(\text{out}, y)
    grad ←∇θloss\leftarrow \nabla_\theta \text{loss}
    saliency ←θ⊙grad\leftarrow \theta \odot \text{grad}
    indices ←TopKIndices(saliency,K)\leftarrow \text{TopKIndices}(\text{saliency}, K)
    θ[indices]←0\theta[\text{indices}] \leftarrow 0
    return θ\theta

    In Transformer models trained on compiled programs (Tracr), pruning operates over individual scalar weights wiw_i. In autoregressive models trained on probabilistic context-free grammars (PCFG), pruning operates over individual neurons, defined as rows in the weight matrices.

  5. Knowl 5 — Synthetic Experimental Framework for Controlled Capability Modulation

    experimental setup

    The controlled experimental evaluation analyzes capability alteration across two procedural setups:

    1. Compiled Transformers via Tracr: Models with 3 standard Transformer blocks (1 attention head + 2 MLP layers per block, no normalization layers) are compiled from RASP programs to implement ground-truth algorithmic capabilities. In the Counter task, the model is compiled to count occurrences of token OPT=’a’O_{PT} = \text{'a'} in a string and fine-tuned on counting OFT=’b’O_{FT} = \text{'b'}, with spurious correlations introduced such that r(x,OFT)−r(x,OPT)=qr(x, O_{FT}) - r(x, O_{PT}) = q with probability CTr∈[0,1]C_{Tr} \in [0, 1]. Fine-tuning uses SGD with momentum, batch size 96, and learning rates ηL=10−1\eta_L = 10^{-1}, ηM=10−2\eta_M = 10^{-2}, ηS=10−3\eta_S = 10^{-3}, ηVS=10−4\eta_{VS} = 10^{-4} with added initial Gaussian weight noise wnoise∼N(0,0.001)w_{noise} \sim \mathcal{N}(0, 0.001).

    2. Learned Capabilities via PCFGs: A 6-block minGPT model (6 attention heads, embedding dimension 192, ~3M parameters) is trained autoregressively on strings generated by probabilistic context-free grammar production rules. Inputs are formatted as SOS + T + O + O' + SOT + Txt + EOT + ART + Ans + EOS, where TT is the task family token and OO is the target operand. The sampling prior of OFTO_{FT} during pretraining is varied as PTL=(0.999,0.001,0.0)P_T^L = (0.999, 0.001, 0.0), PTM=(0.9,0.1,0.0)P_T^M = (0.9, 0.1, 0.0), or PTH=(0.5,0.3,0.2)P_T^H = (0.5, 0.3, 0.2) to control whether downstream fine-tuning encounters a weakly or strongly relevant capability. Downstream fine-tuning runs for 10K10\text{K} iterations using AdamW with learning rates ηM=10−5\eta_M = 10^{-5} and ηS=10−6\eta_S = 10^{-6}.

  6. Knowl 6 — Localization of Fine-Tuning Wrappers via Weight and Neuron Pruning

    empirical result

    Pruning a tiny fraction of salient weights or neurons in fine-tuned models selectively eliminates the downstream task behavior while reviving the original pretraining capability.

    • PCFG Models: In 6-layer minGPT models fine-tuned with a small learning rate (ηS=10−6\eta_S = 10^{-6}) on spuriously correlated data (CTr=1C_{Tr} = 1), pruning between 5 and 15 neurons (out of ~3 million parameters) targeting pretraining loss reduction increases pretraining accuracy on counting OPTO_{PT} from ≈0%\approx 0\% up to 70%−100%70\% - 100\%, while simultaneously collapsing downstream task accuracy on OFTO_{FT} from near 100%100\% to <30%<30\%.
    • Tracr Compiled Models: Pruning a single weight (such as B1WV97-0B1WV97\text{-}0 in the first block or B1WK148-1B1WK148\text{-}1) in models fine-tuned on correlated data (CTr=1C_{Tr}=1) yields an immediate accuracy improvement of over 60%60\% on counting pretraining tokens OPTO_{PT}.
    • Class Prediction Diversity: Output class distribution histograms confirm that post-pruning accuracy recovery is not due to trivial constant-label collapse, but restores genuine multi-class counting capabilities.
  7. Knowl 7 — Preservation of Pretrained Features in Intermediate Residual Layers

    empirical result

    Linear probing across the Transformer residual stream reveals that intermediate representations of pretraining capabilities remain intact after fine-tuning, even when final-layer behavioral performance indicates complete unlearning.

    • In PCFG models, linear probes trained on residual activations of blocks B3 through B6 achieve near 100%100\% accuracy in decoding the count or index of pretraining token OPTO_{PT} when the model is fine-tuned with small or medium learning rates (ηS=10−6,ηM=10−5\eta_S = 10^{-6}, \eta_M = 10^{-5}), identical to the probe performance of the original pretrained model.
    • Probe accuracy on OPTO_{PT} remains high across intermediate layers even under perfect spurious correlation (CTr=1C_{Tr} = 1). A decrease in intermediate probe accuracy is observed only when a large learning rate (ηL=10−4\eta_L = 10^{-4}) is combined with a low pretraining sampling prior (PTLP_T^L), which forces structural re-adaptation.
    • Attention map visualizations in Tracr models show that attention heads in early blocks continue attending exclusively to the pretraining token OPT=’a’O_{PT} = \text{'a'} during downstream execution when small learning rates or spurious correlations (CTr=1C_{Tr}=1) are present.
  8. Knowl 8 — Sample-Efficient Capability Revival via Reverse Fine-Tuning

    empirical result

    Pretraining capabilities that appear behaviorally forgotten after downstream fine-tuning are restored with extreme sample efficiency via reverse fine-tuning (reFT\text{reFT}).

    • PCFG Counter Task: For models pretrained with medium-to-high sampling priors (PTM,PTHP_T^M, P_T^H), reFT\text{reFT} recovers 100%100\% accuracy on the pretraining task within 10210^2 to 10310^3 iterations, even after 10K10\text{K} iterations of downstream fine-tuning under perfect spurious correlation (CTr=1C_{Tr} = 1). When the pretraining capability was weakly relevant (PTLP_T^L), revival requires ≈3×103\approx 3\times 10^3 iterations. In contrast, the control baseline (training from scratch on OFTO_{FT} and fine-tuning on OPTO_{PT}, Scr. + FT\text{Scr. + FT}) requires ≥4.5×103\ge 4.5\times 10^3 iterations and a higher learning rate (ηM\eta_M) to reach equivalent performance.
    • Tracr Models: Tracr compiled models fine-tuned with ηM\eta_M regain the pretraining capability almost immediately under reFT\text{reFT} even with an extremely small learning rate (ηVS=10−4\eta_{VS} = 10^{-4}), whereas baseline scratch-initialized models fail to converge at ηVS\eta_{VS}.
  9. Knowl 9 — Persistence and Rapid Revival of Forbidden Features in TinyStories Models

    data/table

    To evaluate capability deletion in realistic natural language generation, 91M-parameter LLaMA-2 style models (12 layers, 768 hidden dimension, 12 attention heads) were pretrained on TinyStories-Instruct to generate stories with specified features (e.g., Twist, Foreshadowing, Bad Ending). Models were fine-tuned to remove the capability of generating stories with Twists using three protocols: Filtering (F), Filtering + Randomisation (F + RR), and Filtering + Mix & Match (F + MM) under learning rates ηM=10−4\eta_M = 10^{-4} and ηS=10−5\eta_S = 10^{-5}.

    Linear probes trained on intermediate activations at layers B0–B10 (averaged over the final 10 tokens) retain 92%−94%92\%-94\% classification accuracy for the Twist feature across all fine-tuning protocols, matching the upper bound set by the pretrained model with twists.

    Subsequent reverse fine-tuning (reFT\text{reFT}) on twist stories leads to rapid recovery of twist generation proportions (evaluated using a fine-tuned GPT-3.5 classifier), dramatically outpacing a control model pretrained without twists (Not in PT):

    Deletion Protocol Twist Proportion at Iteration
    0 30 300 3000
    F (ηM\eta_M) 44% 81% 81% 82%
    F + RR (ηM\eta_M) 12% 56% 69% 75%
    F + MM (ηM\eta_M) 31% 88% 50% 75%
    F (ηS\eta_S) 69% 88% 75% 94%
    F + RR (ηS\eta_S) 12% 44% 81% 81%
    F + MM (ηS\eta_S) 50% 81% 62% 81%
    Not in PT 12% 31% 44% 81%

    All fine-tuned models reach 56%−88%56\%-88\% twist generation within 30–300 iterations of reFT\text{reFT}, whereas the control baseline requires 3000 iterations to achieve comparable generation rates.

  10. Knowl 10 — Robustness of Pretrained Circuits to Adversarial Capability Deletion and Jailbreaking

    empirical result

    Active attempts to delete pretraining capabilities via adversarial fine-tuning or prompt-level inhibition fail to dismantle underlying capability mechanisms:

    • Adversarial Fine-Tuning (randFT): When a model is fine-tuned to actively produce incorrect outputs for pretraining prompts (e.g., forced to output the count of OFTO_{FT} when prompted for OPTO_{PT}), intermediate linear probes still recover OPTO_{PT} count representations with high accuracy. Reverse fine-tuning the randFT model restores pretraining accuracy with sample efficiency comparable to standard reFT\text{reFT}, significantly outperforming models trained from scratch.
    • Synthetic Jailbreaking: In PCFG experiments where fine-tuning to count OFTO_{FT} occurs under a non-jailbreaking task token TNJT_{NJ}, evaluating the fine-tuned model on TNJT_{NJ} yields near 0% pretraining accuracy. However, querying the same model with alternative pretraining task tokens TJ1T_{J1} or TJ2T_{J2} immediately yields 100%100\% accuracy on counting OPTO_{PT} without requiring any gradient updates, demonstrating that the fine-tuning suppression was conditioned strictly on the wrapper token mapping.

Coverage note — Secondary implementation details such as the specific RASP source code for the Tracr Max-Identifier task, auxiliary PCFG grammar tasks (CCa40, TNULL5), and full tabular sweeps over pretraining iteration milestones (50k vs 200k) were omitted as they are redundant with the primary mechanistic findings and data tables reported.

References

  1. 1.Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255, 2020.
  2. 2.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022.
  3. 3.Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv preprint arXiv:2309.14316, 2023a.
  4. 4.Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.2, knowledge manipulation. arXiv preprint arXiv:2309.14402, 2023b.
  5. 5.Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 1, context-free grammar. arXiv preprint arXiv:2305.13673, 2023c.
  6. 6.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  7. 7.Bing Bai, Jian Liang, Guanhua Zhang, Hao Li, Kun Bai, and Fei Wang. Why attentions may not be interpretable? In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 25–34, 2021.
  8. 8.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  9. 9.Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219, 2022.
  10. 10.Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher Pal. A meta-transfer objective for learning to disentangle causal mechanisms. arXiv preprint arXiv:1901.10912, 2019.
  11. 11.Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg. An interpretability illusion for bert. arXiv preprint arXiv:2104.07143, 2021.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  13. 13.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  14. 14.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650–9660, 2021.
  15. 15.Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill. Data distributional properties drive emergent in-context learning in transformers. Advances in Neural Information Processing Systems, 35:18878–18891, 2022.
  16. 16.Noam Chomsky. Three models for the description of language. IRE Transactions on information theory, 2(3):113–124, 1956.
  17. 17.Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, et al. Neural networks and the chomsky hierarchy. arXiv preprint arXiv:2207.02098, 2022.
  18. 18.Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. Jailbreaker: Automated jailbreak across multiple large language model chatbots. arXiv preprint arXiv:2307.08715, 2023.
  19. 19.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  20. 20.Ronen Eldan and Yuanzhi Li. Tinystories: How small can language models be and still speak coherent english? arXiv preprint arXiv:2305.07759, 2023.
  21. 21.Federica Gerace, Luca Saglietti, Stefano Sarao Mannelli, Andrew Saxe, and Lenka Zdeborová. Probing transfer learning with a model of synthetic correlated datasets. Machine Learning: Science and Technology, 3(1):015030, 2022.
  22. 22.Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. arXiv preprint arXiv:2203.14680, 2022.
  23. 23.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023.
  24. 24.Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  25. 25.Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman. Aligning language models with preferences through f-divergence minimization. arXiv preprint arXiv:2302.08215, 2023.
  26. 26.Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen. Knowledge is a region in weight space for fine-tuned language models. arXiv preprint arXiv:2302.04863, 2023.
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pp. 2790–2799. PMLR, 2019.
  28. 28.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  29. 29.Sarthak Jain and Byron C Wallace. Attention is not explanation. arXiv preprint arXiv:1902.10186, 2019.
  30. 30.Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al. Can machines learn morality? the delphi experiment. arXiv e-prints, pp. arXiv–2110, 2021.
  31. 31.Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra. Linear connectivity reveals generalization strategies. arXiv preprint. arXiv:2205.12411, 2022.
  32. 32.Andrej Karpathy. MinGPT, 2020. Github link. https://github.com/karpathy/minGPT/tree/master.
  33. 33.Suhas Kotha, Jacob Mitchell Springer, and Aditi Raghunathan. Understanding catastrophic forgetting in language models via implicit inference. arXiv preprint arXiv:2309.10105, 2023.
  34. 34.Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. arXiv preprint. arXiv:2202.10054, 2022.
  35. 35.Vivian Lai and Chenhao Tan. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the conference on fairness, accountability, and transparency, pp. 29–38, 2019.
  36. 36.Andrew K Lampinen and Surya Ganguli. An analytic theory of generalization dynamics and transfer learning in deep linear networks. arXiv preprint arXiv:1809.10374, 2018.
  37. 37.Rémi Le Priol, Reza Babanezhad, Yoshua Bengio, and Simon Lacoste-Julien. An analysis of the adaptation speed of causal models. In International Conference on Artificial Intelligence and Statistics, pp. 775–783. PMLR, 2021.
  38. 38.Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task, 2023. Comment: ICLR 2023 oral (notable-top-5%): https://openreview.net/forum?id=DeG07_TcZvT ; code: https://github.com/likenneth/othello_world.
  39. 39.Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. Scaling down to scale up: A guide to parameter-efficient fine-tuning. arXiv preprint arXiv:2303.15647, 2023.
  40. 40.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  41. 41.David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik. Tracr: Compiled transformers as a laboratory for interpretability. arXiv preprint arXiv:2301.05062, 2023.
  42. 42.Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. Transformers learn shortcuts to automata. arXiv preprint arXiv:2210.10749, 2022a.
  43. 43.Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. Exposing attention glitches with flip-flop language modeling. arXiv preprint arXiv:2306.00946, 2023a.
  44. 44.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems, 35:1950–1965, 2022b.
  45. 45.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860, 2023b.
  46. 46.Charles Lovering, Rohan Jha, Tal Linzen, and Ellie Pavlick. Predicting inductive biases of pre-trained models. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=mNtmhaDkAr.
  47. 47.Ekdeep Singh Lubana and Robert P. Dick. A gradient flow framework for analyzing network pruning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=rumv7QmLUue.
  48. 48.Ekdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Krueger, and Hidenori Tanaka. Mechanistic Mode Connectivity, 2022. Comment: 39 pages.
  49. 49.Wesley Maddox, Shuai Tang, Pablo Moreno, Andrew Gordon Wilson, and Andreas Damianou. Fast adaptation with linearized neural networks. In International Conference on Artificial Intelligence and Statistics, pp. 2737–2745. PMLR, 2021.
  50. 50.Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora. A kernel-based view of language model fine-tuning. In International Conference on Machine Learning, pp. 23610–23641. PMLR, 2023.
  51. 51.Michael S Matena and Colin A Raffel. Merging models with fisher-weighted averaging. Advances in Neural Information Processing Systems, 35:17703–17716, 2022.
  52. 52.Clara Meister, Stefan Lazov, Isabelle Augenstein, and Ryan Cotterell. Is sparse attention more interpretable? arXiv preprint arXiv:2106.01087, 2021.
  53. 53.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440, 2016.
  54. 54.Michael C Mozer and Paul Smolensky. Skeletonization: A technique for trimming the fat from a network via relevance assessment. Advances in neural information processing systems, 1, 1988.
  55. 55.Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023.
  56. 56.Maya Okawa, Ekdeep Singh Lubana, Robert P Dick, and Hidenori Tanaka. Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task. https://openreview.net/forum?id=ZXH8KUgFx3, 2023.
  57. 57.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744, 2022.
  58. 58.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193, 2021.
  59. 59.Jonas Pfeiffer, Andreas R<klé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. Adapterhub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020): Systems Demonstrations, pp. 46–54. Association for Computational Linguistics, 2020.
  60. 60.Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. Fine-tuning enhances existing mechanisms: A case study on entity tracking. arXiv preprint arXiv:2402.14811, 2024.
  61. 61.Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  62. 62.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  63. 63.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  64. 64.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  65. 65.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  66. 66.Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. Can wikipedia help offline reinforcement learning? arXiv preprint arXiv:2201.12122, 2022.
  67. 67.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  68. 68.Sofia Serrano and Noah A Smith. Is attention interpretable? arXiv preprint arXiv:1906.03731, 2019.
  69. 69.Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. The pitfalls of simplicity bias in neural networks. Advances in Neural Information Processing Systems, 33:9573–9585, 2020.
  70. 70.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825, 2023.
  71. 71.Hui Shi, Sicun Gao, Yuandong Tian, Xinyun Chen, and Jishen Zhao. Learning bounded context-free grammar via lstm and the transformer: Difference and the explanations. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 8267–8276, 2022.
  72. 72.Michael Sipser. Introduction to the theory of computation. ACM Sigact News, 27(1):27–29, 1996.
  73. 73.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  74. 74.Hidenori Tanaka, Aran Nayebi, Niru Maheswaranathan, Lane McIntosh, Stephen Baccus, and Surya Ganguli. From deep learning to mechanistic understanding in neuroscience: the structure of retinal prediction. Adv. in Neural Information Processing Systems (NeurIPS), 2019.
  75. 75.Ian Tenney, Dipanjan Das, and Ellie Pavlick. Bert rediscovers the classical nlp pipeline. arXiv preprint arXiv:1905.05950, 2019.
  76. 76.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  77. 77.Nilesh Tripuraneni, Michael Jordan, and Chi Jin. On the theory of transfer learning: The importance of task diversity. Advances in neural information processing systems, 33:7852–7862, 2020.
  78. 78.Puja Trivedi, Danai Koutra, and Jayaraman J Thiagarajan. A closer look at model adaptation using feature distortion and simplicity bias. arXiv preprint arXiv:2303.13500, 2023.
  79. 79.Josef Valvoda, Naomi Saphra, Jonathan Rawski, Adina Williams, and Ryan Cotterell. Benchmarking compositionality with formal languages. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 6007–6018, 2022.
  80. 80.Elena Voita and Ivan Titov. Information-theoretic probing with minimum description length. arXiv preprint arXiv:2003.12298, 2020.
  81. 81.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418, 2019.
  82. 82.Yihan Wang, Si Si, Daliang Li, Michal Lukasik, Felix Yu, Cho-Jui Hsieh, Inderjit S Dhillon, and Sanjiv Kumar. Two-stage llm fine-tuning with less specialization and more generalization. arXiv preprint arXiv:2211.00635, 2022.
  83. 83.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483, 2023.
  84. 84.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  85. 85.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021.
  86. 86.Gail Weiss, Yoav Goldberg, and Eran Yahav. Thinking like transformers. In International Conference on Machine Learning, pp. 11080–11090. PMLR, 2021.
  87. 87.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445, 2021.
  88. 88.Sarah Wiegreffe and Yuval Pinter. Attention is not not explanation. arXiv preprint arXiv:1908.04626, 2019.
  89. 89.Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. Detoxifying language models risks marginalizing minority voices. arXiv preprint arXiv:2104.06390, 2021.
  90. 90.Greg Yang and Edward J Hu. Feature learning in infinite-width neural networks. arXiv preprint arXiv:2011.14522, 2020.
  91. 91.Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949, 2023.
  92. 92.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv preprint arXiv:2106.10199, 2021.
  93. 93.Xiao Zhang and Ji Wu. Dissecting learning and forgetting in language model finetuning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=tmsqb6WpLz.
  94. 94.Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora. Do transformers parse while predicting the masked word? arXiv preprint arXiv:2303.08117, 2023a.
  95. 95.Jiachen Zhao, Zhun Deng, David Madras, James Zou, and Mengye Ren. Learning and forgetting unsafe examples in large language models. arXiv preprint arXiv:2312.12736, 2023b.
  96. 96.Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran. What algorithms can transformers learn? a study in length generalization. arXiv preprint arXiv:2310.16028, 2023.
  97. 97.Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. Challenges in automated debiasing for toxic language detection. arXiv preprint arXiv:2102.00086, 2021.
  98. 98.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.

Citation

MLA
Jain, S., et al. “Mechanistically Analyzing the Effects of Fine-tuning on Procedurally Defined Tasks”. arXiv, 2023, http://arxiv.org/abs/2311.12786v2.
APA
Jain, S., Kirk, R., Lubana, E. S., Dick, R. P., Tanaka, H., Grefenstette, E., Rocktäschel, T., & Krueger, D. S. (2023). Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. arXiv. http://arxiv.org/abs/2311.12786v2
Chicago
Jain, S., R. Kirk, E. S. Lubana, et al. 2023. “Mechanistically Analyzing the Effects of Fine-tuning on Procedurally Defined Tasks”. arXiv. http://arxiv.org/abs/2311.12786v2.
Harvard
Jain, S. et al. (2023) “Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.12786v2.
Vancouver
1. Jain S, Kirk R, Lubana ES, Dick RP, Tanaka H, Grefenstette E, Rocktäschel T, Krueger DS (2023) Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. arXiv

BibTeX

@article{jain2023mechanistically,
  title = {Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks},
  author = {Jain, Samyak and Kirk, Robert and Lubana, Ekdeep Singh and Dick, Robert P. and Tanaka, Hidenori and Grefenstette, Edward and Rocktäschel, Tim and Krueger, David Scott},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.12786v2},
  eprint = {2311.12786}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors