Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Zorik GekhmanGal YonaRoee AharoniMatan EyalAmir FederRoi ReichartJonathan Herzig
Reveals that fine-tuning language models on unfamiliar factual data causes them to fit new information much slower than known facts while linearly increasing their rate of hallucinations, indicating that fine-tuning should focus on formatting existing knowledge rather than teaching new facts.
When tailoring large language models to specific tasks through supervised fine-tuning, training datasets frequently contain factual information that was never part of the model's original pre-training. An ongoing concern across artificial intelligence development is that exposing models to unfamiliar facts during fine-tuning might teach them to fabricate information—commonly known as hallucination—rather than grounding their outputs in established knowledge. Addressing this issue is critical today as organizations increasingly rely on fine-tuning to deploy models in high-stakes operational environments.
The article evaluates how supervised fine-tuning on new factual knowledge influences a model's ability to utilize its pre-existing knowledge and estimates the extent to which it drives hallucinations. Specifically, the authors assess whether models can genuinely acquire new factual information during alignment or whether this practice primarily degrades factual accuracy.
To investigate this dynamic, the authors established a controlled closed-book question-answering framework using the PaLM 2-S language model and a standardized knowledge base of entity relations. They developed a sampling-based categorization method to partition data into known and unknown examples based on model generation probabilities before fine-tuning, further stratifying known facts by confidence levels. The researchers then conducted systematic fine-tuning experiments across various durations, tracking training dynamics and performance across both in-distribution and out-of-distribution test sets.
The study yielded several key findings regarding model behavior. First, language models fit unknown examples significantly slower during training than known examples, indicating that models struggle to integrate new factual knowledge during fine-tuning. Second, fitting unknown training examples exhibits a strong negative linear correlation with test accuracy, reducing accuracy by roughly eight percentage points in-distribution, whereas fitting known examples improves accuracy by roughly seven percentage points. Third, introducing unknown facts actively induces hallucinations regarding facts the model previously knew, an effect that generalizes even to unrelated, out-of-distribution factual domains. Fourth, fine-tuning exclusively on moderately known examples yielded the best overall test performance (43.6% exact match), outperforming datasets composed solely of highly known facts (40.5%) by enabling better utilization of uncertain pre-existing knowledge.
These findings demonstrate that supervised fine-tuning is an ineffective mechanism for injecting new knowledge into large language models and poses measurable risks to output factuality. For practitioners, attempting to teach new facts during fine-tuning increases operational risk, causes severe overfitting, and damages the model's reliability on knowledge it already possessed. Instead, fine-tuning primarily acts as an alignment tool that surfaces and organizes pre-existing parametric knowledge acquired during pre-training.
To minimize hallucinations and preserve performance, technical leaders should align fine-tuning datasets with a model's existing knowledge profile. Practitioners should implement early stopping on validation sets or proactively filter out unknown factual examples from training pipelines. Alternatively, teams can teach models to abstain by relabeling unknown training instances with uncertainty expressions such as "I don't know," which preserved a 61.8% accuracy on answered questions across training epochs. For substantial knowledge updates, organizations should rely on continuous pre-training or external retrieval mechanisms rather than supervised fine-tuning.
These conclusions are supported with high confidence within the evaluated closed-book question-answering benchmark on PaLM 2-S. However, leaders should note that the study focused on a single model architecture and short-form factual triplets. Further validation is needed to determine how these dynamics scale across diverse model families, parameter-efficient adaptation methods like low-rank adaptation, and long-form open-ended text generation tasks.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Provides the foundational empirical framework demonstrating that pretrained language models store factual knowledge in their internal parameters prior to fine-tuning.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Establishes the limits of parametric memory recall for factual knowledge versus non-parametric retrieval mechanisms in language models.
- Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). Examines the challenges and failure modes of updating factual knowledge in language models through parameter editing and fine-tuning.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Introduces standard benchmarks and methodologies for evaluating factual inaccuracies and hallucinations in large language models.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Offers a comprehensive taxonomy and analysis of the root causes of hallucination across LLM training and fine-tuning stages.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Develops preference-based fine-tuning methods specifically designed to align model outputs with factuality and reduce hallucinations.
- Paper: Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation, Xiaoying Zhang et al. (2024). Extends the mitigation of fine-tuning hallucinations by employing self-evaluation and preference optimization grounded in the model's internal knowledge.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). Investigates mechanisms for sequentially absorbing new facts during long-horizon fine-tuning without forgetting prior parametric knowledge.
- Paper: Self-Distillation Enables Continual Learning, Idan Shenfeld et al. (2026). Proposes a self-distillation approach to safely integrate fresh knowledge post-training while circumventing standard supervised fine-tuning deficiencies.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). Applies representation engineering to steer model reliance between internal memory and newly introduced external knowledge during inference.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). Provides a mechanistic explanation of how fine-tuning alters outer task representations without rewriting core internal capabilities.
