Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak JainRobert KirkEkdeep Singh LubanaRobert P. DickHidenori TanakaTim RocktäschelEdward GrefenstetteDavid Scott Krueger
Demonstrates through mechanistic interpretability that fine-tuning merely adds superficial, easily reversible functional wrappers over pretrained capabilities rather than altering them, revealing why safety alignment can be unintentionally undone with minimal downstream adaptation.
Adapting large, pretrained artificial intelligence models via fine-tuning is currently the primary industry standard for customizing systems to specific downstream tasks and attempting to eliminate unsafe or toxic behaviors. However, organizations deploy these adapted models without a clear mechanistic understanding of whether fine-tuning fundamentally alters, removes, or rewires the internal capabilities learned during pretraining, or merely suppresses their outward expression.
The article systematically evaluates how fine-tuning alters a neural network's underlying capabilities and demonstrates whether these capabilities are genuinely modified or merely obscured by shallow transformations.
To conduct this evaluation in a controlled and rigorous manner, the analysis employed three complementary experimental settings: compiled transformer models using the Tracr framework with perfectly defined algorithmic functions, models trained on procedurally generated formal grammars, and realistic 91-million-parameter language models trained on the TinyStories dataset. Across these setups, the researchers utilized a diverse suite of mechanistic interpretability tools, including selective network pruning, linear probing of internal representations, attention visualization, and a reverse fine-tuning methodology to assess whether prior capabilities could be easily recovered.
The article establishes three core findings. First, fine-tuning rarely modifies or deletes underlying model capabilities; instead, it typically trains a minimal, superficial transformation—termed a wrapper—in the network's final layers while leaving internal capabilities intact. Second, these learned wrappers are highly localized, meaning that pruning as few as 5 to 15 neurons or a handful of weights readily restores the model's original pretraining behavior. Third, suppressed or nominally deleted capabilities can be revived rapidly and with high sample efficiency; when subjected to reverse fine-tuning, models fully recovered original capabilities within 30 to 300 iterations, far faster than baseline models learning those tasks from scratch.
These findings carry critical operational and risk implications. They reveal that standard safety alignment and fine-tuning techniques create a false sense of security by merely masking undesirable or hazardous behaviors rather than excising them. Downstream practitioners can unintentionally dismantle safety guardrails and re-expose latent, harmful capabilities through simple, standard fine-tuning on superficially unrelated datasets. Consequently, alignment strategies that rely solely on parameter fine-tuning leave systems highly vulnerable to jailbreaks, accidental regression, and compliance failures.
Organizations and developers should not treat standard fine-tuning or superficial alignment protocols as permanent guarantees of model safety. When deploying models in high-stakes environments, technical teams must implement defense-in-depth safeguards, such as continuous runtime monitoring, input-output filtering, and rigorous penetration testing. In addition, research and development efforts must shift toward developing techniques capable of genuinely deleting internal mechanisms rather than relying on standard gradient-based fine-tuning.
While the findings demonstrate high consistency across multiple mechanistic interpretability tools, the core controlled analyses rely on synthetic grammars, compiled transformers, and relatively small language models. Decision-makers should exercise appropriate caution when extrapolating these precise dynamics to frontier-scale foundation models operating on open-domain web data, where further validation is necessary.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). This paper establishes the empirical vulnerability where benign downstream fine-tuning inadvertently breaks safety alignment, providing the direct safety context and motivation that the source analyzes mechanistically.
- Paper: Understanding intermediate layers using linear classifier probes, Guillaume Alain et al. (2016). This foundational work introduces linear classifier probing of intermediate representations, which is one of the core mechanistic interpretability techniques used by the source to track underlying capabilities.
- Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). This study investigates how network pruning and retraining interact with learned subnetwork structures, establishing key pruning concepts utilized in the source's interpretability pipeline.
- Paper: Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, Armen Aghajanyan et al. (2021). This paper demonstrates that fine-tuning modifies a very low intrinsic-dimensional subspace of representations, motivating the source's hypothesis that fine-tuning primarily acts as a wrapper rather than rewriting core capabilities.
- Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). This classic study measures how neural representations transfer or specialize across layers during fine-tuning, providing essential baseline concepts for understanding capability modulation.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). This study extends the investigation of post-training dynamics by comparing whether supervised fine-tuning merely memorizes and wraps behaviors versus reinforcement learning driving true generalization.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). This paper builds upon mechanistic insights into post-training by analyzing whether reinforcement learning creates genuinely novel capabilities or merely amplifies existing base-model representations.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). This work explores continual learning and retention over long-horizon fine-tuning updates, directly engaging with how sequentially learned parameter wrappers can be composed without degrading prior knowledge.
- Paper: Learning, Fast and Slow: Towards LLMs That Adapt Continually, Rishabh Tiwari et al. (2026). This research presents an adaptive continual learning framework that decouples slow core parameter updates from fast prompt adaptations, applying the principle of preserving underlying model capabilities.
- Paper: Self-Distillation Enables Continual Learning, Idan Shenfeld et al. (2026). This work develops self-distillation fine-tuning to mitigate catastrophic capability loss during sequential post-training, directly addressing the retention issues identified in fine-tuning studies.
