Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
Liyan TangTanya GoyalAlexander R. FabbriPhilippe LabanJiacheng XuSemih YavuzWojciech KryscinskiJustin F. RousseauGreg Durrett
Presents the AGGREFACT benchmark to reveal that modern summarization factuality detectors largely make progress on outdated pre-Transformer outputs rather than state-of-the-art models, showing that no single metric reliably detects factual errors across model types and error categories.
Automated text summarization systems are increasingly deployed across enterprise workflows, but their tendency to generate factual errors presents significant compliance, operational, and reputational risks. While numerous automated factuality metrics have been developed to detect these inaccuracies, evaluation benchmarks have historically combined evaluations across different generations of summarization models. This blending makes it difficult to assess whether modern evaluation tools can reliably detect errors produced by modern summarization systems.
The main objective of the article is to systematically evaluate how modern factuality evaluation metrics perform across summaries generated by different generations of summarization systems and across distinct types of factual errors.
To conduct this evaluation, the authors created AGGREFACT, a standardized benchmark aggregating nine human-annotated datasets across two primary news sources (CNN/DM and XSum). The dataset models were stratified into three evolutionary categories: older pre-Transformer systems, early Transformer models, and state-of-the-art fine-tuned models such as BART, PEGASUS, and T5. Nine factuality metrics—including specialized trained models and large language model prompting approaches using ChatGPT—were evaluated under balanced accuracy conditions alongside a unified taxonomy of intrinsic and extrinsic error types.
The analysis reveals several critical findings. First, reported improvements in factuality detection metrics are largely an illusion of outdated benchmarks; performance gains have occurred primarily on older models rather than modern state-of-the-art summarizers. Second, metric performance degrades substantially on modern models, exhibiting an approximate 10% drop in balanced accuracy from older architectures to fine-tuned state-of-the-art systems on CNN/DM data. Third, no single factuality metric consistently outperforms all others across different datasets or model categories; for example, QuestEval achieved top performance on fine-tuned models for CNN/DM summaries (70.2% balanced accuracy), whereas DAE performed best on fine-tuned XSum summaries (70.2% balanced accuracy). Finally, metrics demonstrate large disparities in identifying specific error types, with recall dropping by 10% to 30% when shifting between different source text distributions.
These findings indicate that organizations relying on aggregate benchmark scores are likely overestimating their ability to automatically catch factual inaccuracies in production deployments. Standard off-the-shelf metrics frequently succeed only at flagging crude, outdated generation errors (such as repetition) while missing the subtle, hallucinated facts produced by modern language models. Consequently, deploying these metrics without domain and model alignment introduces unmonitored risk.
Decision-makers should immediately shift factuality evaluation protocols to focus exclusively on summaries generated by current state-of-the-art models and modern large language models rather than legacy benchmarks. Engineering teams must avoid universal evaluation solutions and instead select specific metrics tailored to their target source domain and predominant error types. Furthermore, organizations should establish continuously updated 'living' benchmarks to track new model iterations as generation capabilities evolve.
These conclusions are bounded by the article's focus on English-language newswire datasets and the reliance on historical human annotations with inherent labeling noise. Confidence is high regarding the performance drop-off of current metrics on modern summarization models, but practitioners should exercise caution and conduct domain-specific pilot validations when applying these tools to specialized environments such as dialogue or medical summarization.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). TRUE establishes the unified factual-consistency benchmark and metric comparisons that this paper extends by stratifying performance according to summarizer type and error category.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Maynez et al. document hallucination patterns in abstractive summaries and compare detection approaches, grounding this paper’s later analysis of factual-error datasets and detectors.
- Paper: Hallucinated but Factual! Inspecting the Factuality of Hallucinations in Abstractive Summarization, Meng Cao et al. (2022). This study distinguishes unsupported but true content from genuinely false hallucinations, a key distinction for understanding the error types and factuality evaluations compared here.
- Paper: On the Limitations of Reference-Free Evaluations of Generated Text, Daniel Deutsch et al. (2022). Its analysis of how reference-free metrics can be gamed clarifies the evaluation limitations behind this paper’s comparisons of factuality detectors.
- Paper: Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation, Yixin Liu et al. (2023). The Atomic Content Unit protocol and benchmark provide relevant foundations for interpreting the human annotations and metric evaluations brought together here.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). FActScore carries factuality evaluation into long-form generation by decomposing outputs into atomic claims, extending the paper’s analysis of factuality metrics beyond summarization.
- Paper: All claims are equal, but some claims are more equal than others: Importance-sensitive factuality evaluation of LLM generations, Miriam Wanner et al. (2025). VITAL builds on atomic factuality evaluation by weighting claims according to importance, addressing a limitation that the paper’s metric comparisons make salient.
