Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization
Meng CaoYue DongJackie Chi Kit Cheung
Distinguishes beneficial, world-knowledge-accurate hallucinations from false ones in abstractive summarization by comparing an entity’s prior and posterior probabilities under masked language models, providing an effective reward signal for reinforcement learning to improve summary factuality without sacrificing abstractiveness.
Artificial intelligence summarization models often produce hallucinations—content that cannot be directly inferred from the original source text. Conventional wisdom treats all hallucinations as system errors to eliminate. However, completely removing unmentioned context can degrade output quality, as many unstated details are accurate real-world facts that clarify meaning for readers. Organizations deploying automated summarization must therefore distinguish between helpful background facts and outright factual errors.
The article demonstrates that language model probability shifts can reliably classify entity hallucinations as factual or non-factual, enabling a targeted training mechanism that reduces incorrect statements without sacrificing summary abstractiveness.
The authors developed an entity-level classifier called EntFA using a non-parametric nearest-neighbor approach. The model assesses whether a named entity is factual by comparing its unconditional prior probability from a base language model against its conditional posterior probability from a model given the source document. To test the framework, researchers created a human-annotated dataset of 2,838 entities from 800 generated summaries (XENT) with high inter-annotator agreement (kappa = 0.809), and further evaluated on converted public benchmarks. The classifier was then integrated as a negative reward signal in a reinforcement learning framework to penalize non-factual entity generation during model training.
The investigation produced several key findings. First, while roughly 30% of entities produced by state-of-the-art models were hallucinated, more than half of those hallucinations were accurate according to world knowledge. Second, EntFA outperformed five baseline methods, achieving 90.95% accuracy and an 81.82 F1 score for factuality classification on the primary test set, alongside superior alignment with human judgments on established benchmarks (such as a 0.183 correlation on the FRANK benchmark). Third, error analysis showed that non-factual hallucinations are heavily skewed toward dates (31.65%) and numbers, whereas people, places, and organizations represent the vast majority of factual hallucinations. Finally, using the classifier to penalize non-factual entities during training increased overall factual entity generation from 82.8% to 92.5% and substantially improved faithfulness scores without forcing the model to rely solely on verbatim source copying.
These findings indicate that traditional evaluation metrics and blunt filtering methods are flawed. Optimizing solely for standard word-overlap scores often rewards models for generating non-factual content from noisy training data. Conversely, standard hallucination reduction techniques tend to make models overly extractive, which undermines concise writing. By separating factual background additions from untrue fabrications, organizations can significantly mitigate operational and reputational risks associated with AI errors while preserving natural, informative language generation.
Decision-makers building or adopting text generation systems should replace generic overlap metrics with calibrated probability-based factuality checkers and implement token-level reinforcement penalties during fine-tuning. Future technical roadmaps should extend this probabilistic detection framework beyond named entities to arbitrary text spans and investigate how underlying pre-training datasets introduce or mitigate factual errors.
The primary limitation of this work is its focus on individual named entities and extrinsic hallucinations, which excludes incorrect syntactic relationships among existing facts (intrinsic hallucinations) and full-sentence hallucinations. Additionally, evaluating factuality against world knowledge relied partly on automated search extraction and specific benchmark distributions. Overall confidence in the findings is high for entity-level assessment, though practitioners should conduct targeted validation when applying the method to specialized domains with unique terminology.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). This seminal study establishes the foundational distinction between intrinsic and extrinsic hallucinations in abstractive summarization, providing the direct conceptual starting point for evaluating factual versus non-factual extrinsic generations.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). It introduces the XSum benchmark for extreme abstractive summarization, which serves as a primary testbed for studying abstractiveness and hallucination in modern summarization models.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). This paper establishes the reinforcement learning framework for abstractive summarization that the source adapts using factuality reward signals.
- Paper: Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus, Tianhang Zhang et al. (2023). It extends probability- and uncertainty-based hallucination detection by incorporating attention propagation and entity-focused adjustments without relying on external knowledge bases.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). It scales automated factuality optimization from summarization to general open-ended generation by training preference-based objectives directly on automated factuality signals.
- Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models, Potsawee Manakul et al. (2023). It advances reference-free, zero-resource hallucination detection to black-box generative large language models by evaluating stochastic sample consistency.
- Paper: Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation, Xiaoying Zhang et al. (2024). It generalizes self-evaluation and preference optimization techniques to align model factuality autonomously across diverse knowledge-intensive generation tasks.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). It contextualizes the distinction between factuality and faithfulness hallucinations within a comprehensive broader taxonomy and mitigation framework for large language models.
