Factuality of Large Language Models: A Survey
Yuxia WangMinghan WangMuhammad Arslan ManzoorFei LiuGeorgi Nenkov GeorgievRocktim Jyoti DasPreslav Nakov
Synthesizes recent advances in large language model factuality across text and vision modalities by categorizing evaluation benchmarks, clarifying distinctions between factuality and hallucination, and analyzing error-mitigation techniques alongside calibration strategies.
Large language models are rapidly being integrated into digital assistants and everyday workflows to synthesize information and answer questions directly. However, these systems regularly fabricate ungrounded claims and present incorrect information as fact. Because inaccurate outputs introduce serious operational risks, misinformation, and compliance concerns, establishing reliable factual integrity in model generation has become a critical priority.
The article systematically analyzes the current research landscape surrounding model factuality across text and visual domains. It examines the fundamental causes of factual errors, categorizes benchmarks and evaluation techniques, assesses mitigation strategies across all development stages, and outlines the major obstacles to automated verification.
The authors conducted a comprehensive synthesis of recent literature, organizing evaluation datasets into four operational formats: open-domain generation, binary decisions, short phrases, and multiple-choice questions. They assessed mitigation methods across the entire model lifecycle—including pre-training corpus curation, supervised fine-tuning, preference optimization, decoding adjustments, inference prompting, and post-generation automatic fact-checking—while also examining factuality challenges in multimodal systems.
The findings highlight several core challenges. First, language models are fundamentally trained to optimize sentence probability and fluency rather than objective truth, meaning fluent outputs frequently contain factual inaccuracies. Second, human-annotated error rates on open-ended outputs are notably high, reaching between 42.6% and 68.0% on several major open-ended benchmarks. Third, automated fact-checking systems struggle with accuracy; top-tier verifiers achieve F1 scores of only 0.53 to 0.63 when identifying false claims. Fourth, retrieval augmentation substantially improves factual grounding but introduces latency, requires roughly 25% more computation during pre-training, and remains vulnerable to noisy internet sources. Finally, models are prone to error snowballing, where early inaccuracies compound across subsequent generated text.
These findings indicate that factual errors are not rare edge cases but systemic structural issues inherent in current architectures. Relying on current models for mission-critical, legal, medical, or automated decision-making poses considerable risk without external verification. Furthermore, standard multiple-choice evaluation benchmarks obscure the true severity of errors encountered in realistic, open-ended tasks.
To mitigate these issues, decision-makers should implement layered strategies. Organizations deploying these systems should integrate real-time retrieval-augmented generation and intermediate-sentence verification to prevent error compounding, while accepting modest latency trade-offs. Model developers should adopt behavioral fine-tuning and refusal-aware training so models decline to answer questions outside their knowledge base. Additionally, automated verification pipelines should transition toward smaller, specialized natural language inference models to improve verification reliability and reduce operational costs.
These conclusions should be interpreted in light of current research constraints. Evaluating open-ended text remains inherently uncertain because dividing text into atomic claims and scoring search quality lack universal standards. While confidence is high regarding the conceptual trade-offs and structural causes of hallucinations, rigorous automated measurement of open-ended model factuality remains an active, evolving area of investigation.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Introduces the foundational atomic-fact decomposition methodology (FActScore) that underpins the survey's discussion of fine-grained, open-ended factuality evaluation.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Establishes the TruthfulQA benchmark and the core empirical phenomenon of imitative falsehoods and inverse scaling, which are central themes reviewed in the survey.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Provides the foundational taxonomy differentiating factuality hallucinations from faithfulness hallucinations across the model development lifecycle.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Presents the TRUE framework for standardized meta-evaluation of factual consistency metrics, establishing the baseline evaluation landscape surveyed by the source.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Demonstrates the limitations of parametric memory on long-tail factual knowledge versus non-parametric retrieval, motivating the mitigation strategies analyzed in the survey.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). Offers the classic foundational survey on hallucination definitions, metrics, and mitigation across natural language generation tasks prior to modern LLM surveys.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). Introduces Inference-Time Intervention to steer internal representations toward truthfulness, representing a primary decoding-time mitigation method categorized in the survey.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Defines the POPE probing benchmark to measure object hallucination in vision-language models, establishing the multimodal factuality challenges reviewed in the survey.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). Establishes the ALCE benchmark for evaluating citation support and factual precision in retrieval-augmented generation systems.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Demonstrates zero-shot multi-turn cross-examination between models as a mechanism to detect internal factual inconsistencies without external knowledge bases.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Applies Direct Preference Optimization with reference-based and reference-free automated feedback to directly execute the post-training factuality mitigation recommended in the survey.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Directly tackles the noisy retrieval vulnerabilities highlighted in the survey by implementing adaptive adversarial training to improve retrieval-augmented model robustness.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Broadens the survey's focus on factual integrity into a comprehensive, multi-dimensional trustworthiness evaluation framework covering safety, ethics, and robustness.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). Extends factuality verification research by using small internal circuit analyzers to enable frozen language models to self-predict errors and hallucinations efficiently.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). Advances multimodal factual grounding by scaling fine-grained late-interaction multi-modal retrievers to address knowledge deficits in vision-language tasks.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Investigates the reliability and vulnerabilities of automated LLM evaluators when judging instruction following, extending the survey's critique of automated verifiers.
- Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). Provides a deep dive into constructing verifier-anchored post-training reasoning data to enhance factual reasoning trajectories and minimize rationalized errors.
