Unfamiliar Finetuning Examples Control How Language Models Hallucinate
Katie KangEric WallaceClaire J. TomlinAviral KumarSergey Levine
Reveals that language models default to mirroring their unfamiliar finetuning data when hallucinating, enabling a practical method to mitigate factual errors in long-form generation by supervising unfamiliar examples with conservative reward models.
Large language models frequently generate plausible but factually incorrect statements, commonly known as hallucinations. This failure mode poses significant operational and reputational risks, particularly when models are queried about topics that extend beyond their underlying training data. The article addresses the root mechanisms driving these errors during model adaptation and aims to demonstrate how practitioners can systematically control and reduce hallucinations in both short-form answering and complex, long-form text generation.
To establish these principles, the authors conducted controlled experiments across multiple fine-tuning paradigms, including supervised fine-tuning, reinforcement learning, and reward model training using established benchmarks like MMLU and TriviaQA. They evaluated open-source base models (primarily Llama 2 7B and Mistral 7B) by categorizing test queries by unfamiliarity based on baseline model performance. The investigation then extended to long-form generation tasks—specifically biography generation on WikiBios and narrative plot summaries on WikiPlots—evaluating factual accuracy using automated fact-checking pipelines.
The findings reveal that when models face unfamiliar queries, their responses consistently default toward the distribution of responses present in their unfamiliar fine-tuning data rather than making arbitrary errors. In supervised fine-tuning, if unfamiliar training examples are labeled with abstentions like "I don't know," the model reliably learns to abstain on unfamiliar test prompts while answering familiar ones accurately. In reinforcement learning, scoring models (reward models) often hallucinate by overestimating rewards on unfamiliar concepts, which inadvertently teaches language models to fabricate information. Finally, the authors show that training a "conservative reward model" on responses generated by the model's own base distribution forces the reward model to assign low scores to unknown concepts; fine-tuning with this conservative reward model significantly reduced false facts across all unfamiliarity levels while maintaining or increasing true factual statements.
These insights demonstrate that language model hallucinations are manageable rather than entirely random. In practice, uncurated fine-tuning data and standard reward models actively incentivize models to produce confident, incorrect outputs. By structuring unfamiliar training data and reward systems to favor cautious predictions, organizations can reduce error rates and improve output reliability in high-stakes generative applications without requiring prohibitively expensive real-time verification at every generation step.
Engineering teams should audit their fine-tuning datasets to ensure unfamiliar examples explicitly model uncertainty or abstention. For reinforcement learning workflows, practitioners should adopt conservative reward modeling protocols—using the target model's own generations to populate the reward dataset—to prevent reward inflation on unfamiliar queries. However, leaders should note that these experiments primarily addressed specific generation tasks and clear distinctions between known and unknown data. Because many real-world enterprise queries inhabit a spectrum of partial familiarity, additional testing and pilot evaluations across broader, generalized domains are recommended before critical production deployment.
- Paper: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, Zorik Gekhman et al. (2024). Its controlled study of fine-tuning on unfamiliar facts establishes how training examples can alter factual accuracy, a prerequisite for understanding this paper’s focus on shaping unfamiliar-example behavior.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). It introduces preference-based factuality tuning and automated reward construction, providing essential context for this paper’s conservative reward-model approach.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Its synthesis of hallucination causes and mitigation across supervised fine-tuning and reinforcement learning supplies the broader framework this paper investigates experimentally.
- Paper: Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning, Lorenzo Jaime Flores et al. (2026). It tests how supervised fine-tuning changes confidence-score calibration, extending this paper’s findings on how training choices shape reliability beyond factuality outcomes.
