Fine-Tuning Language Models for Factuality
Katherine TianEric MitchellHuaxiu YaoChristopher D. ManningChelsea Finn
Demonstrates that training language models with direct preference optimization on automatically generated factuality rankings reduces open-ended generation errors by up to 58% without requiring human annotation.
Large language models often generate fluent and convincing text that contains factual errors, known as hallucinations. This tendency poses significant operational, reputational, and safety risks, especially as organizations increasingly rely on automated text generation for research, customer service, and specialized domains. Traditional methods to address this issue rely heavily on human fact-checking, which is expensive and slow, costing thousands of dollars to evaluate even modest datasets. Consequently, scalable and cost-effective methods are required to improve model reliability without human annotation.
The article demonstrates an automated fine-tuning pipeline designed to systematically reduce factual errors in open-ended text generation without requiring human intervention. It evaluates how effectively preference-based reinforcement learning can optimize language models for factuality across complex domains like biographical writing and medical question-answering.
To achieve this, the authors developed two automated methods to evaluate candidate model responses and generate preference rankings without human annotators. The first method uses reference-based fact-checking, where individual claims are extracted and verified against Wikipedia data. The second method is entirely reference-free, measuring the model's internal confidence by converting factual claims into targeted questions, resampling answers, and computing semantic consistency. These automated rankings were then used to train 7-billion-parameter open-source models (Llama-1 and Llama-2) using Direct Preference Optimization, an efficient preference-learning algorithm.
The findings show that automated factuality tuning substantially enhances accuracy across tasks. Fine-tuning with reference-based preferences reduced factual error rates in Llama-2 by 58% on biography generation and by 40% on medical question-answering compared to standard instruction-tuned chat models. This approach was the only method evaluated that achieved a strict improvement, simultaneously increasing the total number of correct statements while reducing incorrect ones. The reference-free approach also outperformed standard reinforcement learning baselines, reducing biography errors by over 50% and medical errors by 20% to 30% without using external knowledge bases. Additionally, human evaluations and advanced automated reviews validated these improvements, confirming that accuracy gains were genuine.
These results demonstrate that organizations can significantly mitigate the risk of generating inaccurate information while avoiding the high costs of human feedback pipelines. The findings also reveal that factuality tuning alters text style toward more direct and concise responses rather than conversational storytelling. Crucially, the technique works complementarily with inference-time decoding interventions, allowing multiple safety and accuracy methods to be layered together.
Organizations developing or deploying language models should adopt automated factuality preference pipelines before releasing systems for knowledge-critical applications. For general domains with strong reference data, reference-based tuning offers the highest accuracy gains, while reference-free confidence scoring provides a viable alternative for proprietary or niche domains lacking reference corpora. Future work should pilot these techniques on larger language models and explore deeper integration between factuality objectives and broader conversational capabilities.
Confidence in these findings is strong for 7-billion-parameter models across the tested biographical and medical tasks. However, decision-makers should exercise caution when extrapolating results to larger model architectures or different tasks, as real-world production environments and complex reasoning domains may present additional unmeasured failure modes.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). Introduces Direct Preference Optimization (DPO), the foundational preference-learning algorithm utilized by the source to fine-tune language models for factuality without reinforcement learning.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Presents FActScore for decomposing open-ended text into atomic claims and verifying them against external references, establishing the evaluation paradigm adapted in the source's reference-based pipeline.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Examines how language models represent their own internal knowledge and calibrate confidence, laying theoretical ground for the source's reference-free factuality scoring via consistency and question generation.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Demonstrates the principle of self-consistency and semantic resampling to identify correct model outputs, which directly inspires reference-free confidence estimation.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Introduces automated cross-examination and inquiry generation without external references to detect factual inconsistencies in model responses.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). Explores inference-time steering of model activations toward truthfulness, providing a complementary factual alignment mechanism that the source benchmarks alongside fine-tuning.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Establishes a core benchmark for measuring factual errors and imitative falsehoods in generative language models.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Demonstrates the foundational framework of aligning language models using preference modeling and reinforcement learning from human feedback.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). Extends the principle of eliminating human labelers by systematically comparing Reinforcement Learning from AI Feedback (RLAIF) against standard human preference training.
- Paper: ORPO: Monolithic Preference Optimization without Reference Model, Jiwoo Hong et al. (2024). Proposes odds ratio preference optimization (ORPO) to further streamline preference tuning without requiring explicit reference models during training.
- Paper: Zephyr: Direct Distillation of LM Alignment, Lewis Tunstall et al. (2024). Applies automated preference data generation and direct preference optimization to distill alignment capabilities into compact 7B-parameter models.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). Advances reference-free failure detection by analyzing internal representations and attention circuits in frozen language models to self-predict errors.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Investigates training strategies to make language models robust against retrieval noise and factual inconsistencies when grounded in external references.
