Learning or Self-aligning? Rethinking Instruction Fine-tuning
Mengjie RenBoxi CaoHongyu LinCao LiuXianpei HanKe ZengGuanglu WanXunliang CaiLe Sun
Demonstrates that instruction fine-tuning succeeds primarily by aligning outputs with a language model's existing internal parameter knowledge rather than teaching new facts, warning that introducing inconsistent world knowledge during tuning degrades model performance.
Instruction fine-tuning is a vital step in transforming pre-trained large language models from basic text generators into helpful, task-oriented assistants. In industry and academia, fine-tuning is often treated as a standard supervised learning method to inject new domain-specific knowledge into models. However, this approach frequently underperforms or causes unintended performance degradations, creating substantial uncertainty around how to best construct fine-tuning datasets and adapt models for specialized domains.
The article evaluates whether instruction fine-tuning actually teaches models new factual knowledge or primarily aligns their output behavior with knowledge they already possess. To demonstrate this, the authors establish a knowledge intervention framework across four specialized domains—medicine, history, engineering, and jurisprudence—using four leading open-source models ranging from 7 billion to 70 billion parameters (LLaMA-2 and Mistral variants). The researchers probed each model's internal baseline knowledge using few-shot in-context learning and constructed experimental training sets that systematically manipulated the consistency between the training data and the model's pre-existing knowledge.
The findings reveal that attempting to force models to learn new factual knowledge during instruction fine-tuning produces poor results and often harms performance across related and unrelated tasks. Remarkably, models trained on factually incorrect responses that matched their internal knowledge outperformed models trained on factually correct responses that conflicted with their internal knowledge, showing average accuracy gains of roughly 5% to 10%. Furthermore, providing external factual context directly within the prompt during training prevented these negative effects, yielding performance improvements of approximately 4% to 9.5%. Most importantly, statistical analysis confirmed that the primary driver of fine-tuning success is maintaining consistency between the model's internal knowledge before and after tuning, rather than attempting to alter its internal factual state.
These insights demonstrate that instruction fine-tuning operates as a self-alignment mechanism rather than a knowledge injection phase. For strategic decision-makers, this shifts how resources and risk should be allocated in artificial intelligence initiatives. Attempting to update a model's factual worldview through simple instruction tuning introduces significant performance risks, increases hallucinations, and wastes engineering time. Instead, fine-tuning should be restricted to teaching style, behavioral norms, and response formatting, while core domain facts should be embedded during pre-training or retrieved dynamically at runtime.
Organizations developing or deploying large language models should immediately audit their fine-tuning workflows to ensure datasets align with base model internal capabilities rather than attempting to inject raw facts. When conflicting domain knowledge must be processed during training, teams should provide explicit contextual reference data within the prompt to decouple behavioral learning from factual learning. Decision-makers should also evaluate alignment pipelines that use the base model’s own internal representations to guide fine-tuning.
While the article's core statistical conclusions are supported by strong confidence levels across multiple architectures, certain boundaries apply. The empirical validation relies primarily on multiple-choice formats and focuses mostly on models around the 7-billion to 13-billion parameter range, with limited evaluations at the 70-billion scale and no free-form text generation benchmarks. Additional pilots and evaluations in open-ended generative settings are recommended before overhauling enterprise-scale generation systems.
No sufficiently relevant recommendations were found.
- Paper: Finetuning with Sampling: SFT Learns Better Than You Think, Aayush Karan et al. (2026). This later method turns the source’s finding that tuning works best when it respects a model’s internal distribution into a practical way to teach new capabilities while limiting forgetting.
- Paper: Unfamiliar Finetuning Examples Control How Language Models Hallucinate, Katie Kang et al. (2025). This study extends the source’s account of fine-tuning and prior knowledge by showing how the familiarity of training examples can be used to shape hallucination and abstention behavior.
- Paper: Inside-Out: Hidden Factual Knowledge in LLMs, Jonathan Herzig et al. (2025). This later work extends the source’s premise that models may hold more factual knowledge internally than tuning makes them express by measuring that hidden knowledge directly.
