Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors

Liyan TangTanya GoyalAlexander R. FabbriPhilippe LabanJiacheng XuSemih YavuzWojciech KryscinskiJustin F. RousseauGreg Durrett

article2023ACL151 citations

Presents the AGGREFACT benchmark to reveal that modern summarization factuality detectors largely make progress on outdated pre-Transformer outputs rather than state-of-the-art models, showing that no single metric reliably detects factual errors across model types and error categories.

Listen

Automated text summarization systems are increasingly deployed across enterprise workflows, but their tendency to generate factual errors presents significant compliance, operational, and reputational risks. While numerous automated factuality metrics have been developed to detect these inaccuracies, evaluation benchmarks have historically combined evaluations across different generations of summarization models. This blending makes it difficult to assess whether modern evaluation tools can reliably detect errors produced by modern summarization systems.

The main objective of the article is to systematically evaluate how modern factuality evaluation metrics perform across summaries generated by different generations of summarization systems and across distinct types of factual errors.

To conduct this evaluation, the authors created AGGREFACT, a standardized benchmark aggregating nine human-annotated datasets across two primary news sources (CNN/DM and XSum). The dataset models were stratified into three evolutionary categories: older pre-Transformer systems, early Transformer models, and state-of-the-art fine-tuned models such as BART, PEGASUS, and T5. Nine factuality metrics—including specialized trained models and large language model prompting approaches using ChatGPT—were evaluated under balanced accuracy conditions alongside a unified taxonomy of intrinsic and extrinsic error types.

The analysis reveals several critical findings. First, reported improvements in factuality detection metrics are largely an illusion of outdated benchmarks; performance gains have occurred primarily on older models rather than modern state-of-the-art summarizers. Second, metric performance degrades substantially on modern models, exhibiting an approximate 10% drop in balanced accuracy from older architectures to fine-tuned state-of-the-art systems on CNN/DM data. Third, no single factuality metric consistently outperforms all others across different datasets or model categories; for example, QuestEval achieved top performance on fine-tuned models for CNN/DM summaries (70.2% balanced accuracy), whereas DAE performed best on fine-tuned XSum summaries (70.2% balanced accuracy). Finally, metrics demonstrate large disparities in identifying specific error types, with recall dropping by 10% to 30% when shifting between different source text distributions.

These findings indicate that organizations relying on aggregate benchmark scores are likely overestimating their ability to automatically catch factual inaccuracies in production deployments. Standard off-the-shelf metrics frequently succeed only at flagging crude, outdated generation errors (such as repetition) while missing the subtle, hallucinated facts produced by modern language models. Consequently, deploying these metrics without domain and model alignment introduces unmonitored risk.

Decision-makers should immediately shift factuality evaluation protocols to focus exclusively on summaries generated by current state-of-the-art models and modern large language models rather than legacy benchmarks. Engineering teams must avoid universal evaluation solutions and instead select specific metrics tailored to their target source domain and predominant error types. Furthermore, organizations should establish continuously updated 'living' benchmarks to track new model iterations as generation capabilities evolve.

These conclusions are bounded by the article's focus on English-language newswire datasets and the reliance on historical human annotations with inherent labeling noise. Confidence is high regarding the performance drop-off of current metrics on modern summarization models, but practitioners should exercise caution and conduct domain-specific pilot validations when applying these tools to specialized environments such as dialogue or medical summarization.

arXiv: 2205.12854Liyan06/AggreFact
Cover for Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors

Abstract

The propensity of abstractive summarization models to make factual errors has been studied extensively, including design of metrics to detect factual errors and annotation of errors in current systems’ outputs. However, the ever-evolving nature of summarization systems, metrics, and annotated benchmarks makes factuality evaluation a moving target, and drawing clear comparisons among metrics has become increasingly difficult. In this work, we aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model. We compare performance of state-of-the-art factuality metrics, including recent ChatGPT-based metrics, on this stratified benchmark and show that their performance varies significantly across different types of summarization models. Critically, our analysis shows that much of the recent improvement in the factuality detection space has been on summaries from older (pre-Transformer) models instead of more relevant recent summarization models. We further perform a finer-grained analysis per error-type and find similar performance variance across error types for different factuality metrics. Our results show that no one metric is superior in all settings or for all error types, and we provide recommendations for best practices given these insights.

Table of Contents

  • 1 Introduction
  • 2 Benchmark
  • 2.1 Benchmark Standardization
  • 2.2 Benchmark Datasets
  • 2.3 Benchmark Evaluation Metrics
  • 3 Comparison of Factuality Metrics
  • 4 Finer-grained Error Analysis
  • 4.1 A Taxonomy of Error Types
  • 4.2 Error Mapping
  • 4.3 Distribution Shift of Error Types
  • 4.4 Error Detection by Type
  • 5 Recommendations
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Model Categories
  • B Factuality Metrics
  • C Surveyed Error Types

Knowls

  1. Knowl 1 — AGGREFACT combines and stratifies nine human-annotated factuality datasets

    model/method

    AGGREFACT aggregates nine existing datasets of human judgments about factual consistency in summaries of CNN/DM and XSum articles: FactCC, Wang’20, SummEval, Polytope, Cao’22, XSumFaith, FRANK, Goyal’21, and CLIFF. It treats consistency assessment as binary classification and provides separate CNN/DM and XSum subsets, each further divided by the type of summarization model that produced a summary. The CNN/DM validation/test counts are 2,297/2,166 for OLD models, 275/375 for EXFORMER models, and 459/559 for FTSOTA models; the corresponding XSum counts are 500/430, 500/423, and 777/558. The benchmark uses validation and test splits from SummaC for the included datasets it shares with that benchmark; Wang’20, CLIFF, and Goyal’21 are split by index parity, while Cao’22 retains its original splits. CoGenSumm is excluded because it ranks pairs of summaries rather than classifying individual summaries. For Polytope, addition, omission, and duplication annotations are treated as consistent because they do not represent factual inconsistency. Duplicate examples across datasets are removed; the authors manually corrected 100 instances whose labels conflicted across datasets.

  2. Knowl 2 — Model-category stratification exposes differences between old and current summaries

    definition

    AGGREFACT groups the summarization systems that produced its examples into three categories. FTSOTA denotes fine-tuned pretrained Transformer summarizers such as BART, PEGASUS, and T5. EXFORMER denotes earlier Transformer-based systems, including BERTSum and GPT-2. OLD covers the remaining, earlier systems, including Pointer-Generator and BottomUp. The categories are used to distinguish evaluation on current fine-tuned systems from evaluation on earlier Transformer and pre-Transformer systems; the paper’s central comparison is that factuality-detector performance depends on this distinction.

  3. Knowl 3 — Factuality metrics are compared as thresholded binary classifiers

    experimental setup

    The evaluated metrics are DAE, QuestEval, SummaC-ZS, SummaC-Conv, QAFactEval, ChatGPT-ZS, ChatGPT-CoT, ChatGPT-DA, and ChatGPT-Star. DAE, QuestEval, SummaC-ZS, SummaC-Conv, and QAFactEval produce scores; ChatGPT-ZS and ChatGPT-CoT directly return a binary judgment, while ChatGPT-DA and ChatGPT-Star score factuality on 0–100 and 1–5 scales, respectively. Because the benchmark’s factuality labels are imbalanced, performance is measured with balanced accuracy. In the threshold-per-dataset setting, each metric’s thresholds are chosen on validation data separately for each dataset and model category, then applied to test data; reported performance is a weighted average across benchmark datasets. DAE’s EXFORMER and OLD results on XSum are omitted because DAE was trained on human-annotated XSumFaith examples from those categories.

  4. Knowl 4 — Detector performance varies substantially by summarizer category and dataset

    empirical result

    In the threshold-per-dataset setting, the following balanced accuracies are reported as CNN/DM FTSOTA, EXFORMER, OLD; then XSum FTSOTA, EXFORMER, OLD: baseline, 50.0, 50.0, 50.0; 50.0, 50.0, 50.0. DAE, 59.4, 67.9, 69.7; 73.1, unavailable, unavailable. QuestEval, 63.7, 64.3, 65.2; 61.6, 60.1, 59.7. SummaC-ZS, 63.3, 76.5, 76.3; 56.1, 51.4, 53.3. SummaC-Conv, 70.3, 69.8, 78.9; 67.0, 64.6, 67.5. QAFactEval, 61.6, 69.1, 80.3; 65.9, 59.6, 60.5. ChatGPT-ZS, 66.2, 64.5, 74.3; 62.6, 69.2, 60.1. ChatGPT-CoT, 49.7, 60.4, 66.7; 56.0, 60.9, 50.1. ChatGPT-DA, 48.0, 63.6, 71.0; 53.6, 65.6, 61.5. ChatGPT-Star, 55.8, 65.8, 71.2; 57.7, 70.6, 53.8. On CNN/DM, both trained and ChatGPT-based metrics perform best on OLD summaries, and average performance drops by about 10 percentage points from OLD to FTSOTA. On XSum, trained metrics perform best on FTSOTA summaries while ChatGPT-based metrics perform best on EXFORMER summaries. Thus, category-agnostic results—especially when many examples come from OLD systems—do not reliably characterize detector performance on current summarizers.

  5. Knowl 5 — Single-threshold rankings differ across CNN/DM and XSum FTSOTA summaries

    empirical result

    When one threshold per metric is used to evaluate FTSOTA summaries, balanced accuracy with 95% confidence intervals is as follows, in CNN/DM; XSum order: DAE, 65.4 ± 4.4; 70.2 ± 2.3. QuestEval, 70.2 ± 3.2; 59.5 ± 2.7. SummaC-ZS, 64.0 ± 3.8; 56.4 ± 1.2. SummaC-Conv, 61.0 ± 3.9; 65.0 ± 2.2. QAFactEval, 67.8 ± 4.1; 63.9 ± 2.4. ChatGPT-ZS, 56.3 ± 2.9; 62.7 ± 1.7. ChatGPT-CoT, 52.5 ± 3.3; 55.9 ± 2.1. ChatGPT-DA, 53.7 ± 3.5; 54.9 ± 1.9. ChatGPT-Star, 56.3 ± 3.1; 57.8 ± 0.2. QuestEval has the highest reported CNN/DM score; its advantage over other trained metrics is not statistically significant, although it is significantly better than the ChatGPT-based metrics. DAE is significantly better than all other metrics on XSum. The average validation thresholds for CNN/DM are higher than those for XSum across the evaluated metrics and model categories, making a shared threshold across the two datasets difficult to calibrate.

  6. Knowl 6 — A unified taxonomy separates source contradictions from unsupported claims

    definition

    The paper consolidates factual inconsistency annotations into a taxonomy organized around two distinctions. Intrinsic errors misrepresent information in the source; extrinsic errors introduce information not present in the source and not verifiable from it. A noun-phrase error affects a phrase functioning as a subject, object, or prepositional object; a predicate error affects a main content verb or closely related content such as an adverb. Their cross-product yields intrinsic noun-phrase, intrinsic predicate, extrinsic noun-phrase, and extrinsic predicate errors. The analysis also uses intrinsic-entire-sentence and extrinsic-entire-sentence categories when a whole sentence is annotated as erroneous. Discourse errors are excluded from the unified analysis because they are uncommon and absent from most of the included datasets.

  7. Knowl 7 — Fine-grained annotations are mapped into the shared taxonomy with dataset-specific rules

    model/method

    The unified error analysis maps annotations from XSumFaith, FRANK, Goyal’21, and CLIFF. A summary receives an error type when more than one annotator marks the error, and a summary may receive multiple types. In XSumFaith and CLIFF, an error span containing a detected predicate is mapped to a predicate error; other spans become noun-phrase errors, while the original intrinsic/extrinsic label is retained. FRANK’s entity and out-of-article errors are mapped to extrinsic noun-phrase errors; predicate and grammatical errors to extrinsic predicate errors; circumstance and coreference errors to intrinsic noun-phrase errors; and other errors to intrinsic predicate errors. Goyal’21 entity and noun-phrase errors map to noun-phrase errors, event errors to predicate errors, and other errors to entire-sentence errors, preserving intrinsic/extrinsic labels. Manual inspection of 30 inconsistent examples each from XSumFaith, FRANK, and CLIFF found mapping accuracy above 90%; Goyal’21 mapping errors in 150 examples were manually corrected. FRANK was the least straightforward mapping: for example, its entity-error label can correspond to either an intrinsic or an extrinsic error.

  8. Knowl 8 — Error distributions vary by summarizer and by annotation dataset

    empirical result

    The unified annotations reveal both model-related distribution shifts and disagreement between datasets. In the combined analysis, approximately 50% of CNN/DM errors are extrinsic, with a slight decrease from OLD to more recent model categories; approximately 70% of XSum errors are extrinsic, with little change across categories. Error proportions can also differ among models within the same category. More importantly, annotations of summaries from the same model can disagree across datasets: XSum BART summaries have more intrinsic noun-phrase and predicate errors in Goyal’21 but more extrinsic noun-phrase errors in CLIFF; CNN/DM BART summaries have more extrinsic predicate errors in FRANK but more intrinsic noun-phrase errors in CLIFF. XSumFaith and FRANK also yield markedly different type distributions for the same XSum summaries. The paper attributes such differences partly to annotation schemes that do not align cleanly with the unified categories—for example, FRANK’s out-of-article error may map to either an extrinsic noun-phrase or an extrinsic predicate error.

  9. Knowl 9 — Error-type recall depends on both detector and summarization dataset

    empirical result

    For the error-type analysis, the authors evaluate recall on test examples containing a single unified error type, without splitting by summarizer category because the resulting subsets are small. The CNN/DM and XSum error subsets are dominated by non-FTSOTA summaries (89.6% and 92.1%, respectively); 79.0% of the CNN/DM subset comes from OLD models, and it contains 349 extrinsic versus 243 intrinsic errors. Recent metrics, including SummaC-Conv, QAFactEval, and ChatGPT-based metrics, have higher recall for many error types, consistent with better detection of more obvious errors in summaries from older systems. The detector ordering is not uniform across types or datasets: on XSum, ChatGPT-ZS and ChatGPT-CoT attain 83.2% recall for intrinsic noun-phrase errors, 85.8% and 91.2% for intrinsic predicate errors, and 94.1% for intrinsic entire-sentence errors; their extrinsic entire-sentence recall is 93.9% and 91.9%, respectively. For extrinsic noun-phrase errors, the trained metrics’ recall is reported to fall by 10–30 percentage points on XSum relative to CNN/DM. DAE, despite being trained with XSumFaith annotations, does not detect errors as well on the CNN/DM error subset. These findings show that detector ability on a named error type does not transfer uniformly across summarization datasets.

  10. Knowl 10 — Evaluation should prioritize current systems and align annotations with intended use

    model/method

    The paper recommends evaluating factuality detectors on summaries from current fine-tuned summarizers rather than relying on benchmarks dominated by older systems, since older systems’ more obvious errors can make detector performance appear stronger. Detector choice should be matched to the downstream dataset and error types: the study finds no single metric that leads across CNN/DM, XSum, model categories, and error categories. For comparability, new annotations should use consistent definitions aligned with the intrinsic/extrinsic and noun-phrase/predicate distinctions where applicable. The authors also recommend updating factuality benchmarks as summarization systems evolve and adding human error annotations for LLM-generated summaries; such summaries were not evaluated in this study.

  11. Knowl 11 — The analysis is limited to English news and inherits existing annotation weaknesses

    limitation

    The evaluation covers English-language newswire summaries from CNN/DM and XSum, so it does not establish detector behavior for other languages, writing styles, topics, or domains such as dialogue and medical summarization, where different error types may matter. Because AGGREFACT reuses prior datasets, its fine-grained analyses depend on the quality and inter-annotator agreement of their original labels, and the mapping of different annotation schemes into one taxonomy is imperfect. The authors did not undertake large-scale reannotation, in part to avoid publishing competing versions of datasets that reflect different annotator judgments.

Coverage note — Detailed prompt wording, annotator metadata for each source dataset, and per-source metric breakdowns are omitted as supporting material; the benchmark construction, aggregate stratified comparisons, unified error analysis, recommendations, and stated limitations are represented.

References

  1. 1.Meng Cao, Yue Dong, and Jackie Cheung. 2022. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual error correction for abstractive summarization models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6251–6258, Online. Association for Computational Linguistics.
  3. 3.Shuyang Cao and Lu Wang. 2021. CLIFF: Contrastive learning for improving faithfulness and factuality in abstractive summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6633–6649, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Sihao Chen, Fan Zhang, Kazoo Sone, and Dan Roth. 2021. Improving faithfulness in abstractive summarization with contrast candidate generation and selection. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5935–5941, Online. Association for Computational Linguistics.
  5. 5.Yen-Chun Chen and Mohit Bansal. 2018. Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 675–686, Melbourne, Australia. Association for Computational Linguistics.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, and Jackie Chi Kit Cheung. 2018. BanditSum: Extractive summarization as a contextual bandit. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3739–3748, Brussels, Belgium. Association for Computational Linguistics.
  8. 8.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  9. 9.Alexander Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang, Haoran Li, Yashar Mehdad, and Dragomir Radev. 2021a. ConvoSumm: Conversation summarization benchmark and improved abstractive summarization with argument mining. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6866–6880, Online. Association for Computational Linguistics.
  10. 10.Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021b. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  11. 11.Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021c. Qafacteval: Improved qa-based factual consistency evaluation for summarization.
  12. 12.Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220, Florence, Italy. Association for Computational Linguistics.
  13. 13.Sebastian Gehrmann, Yuntian Deng, and Alexander Rush. 2018. Bottom-up abstractive summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4098–4109, Brussels, Belgium. Association for Computational Linguistics.
  14. 14.Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3592–3603, Online. Association for Computational Linguistics.
  15. 15.Tanya Goyal and Greg Durrett. 2021. Annotating and modeling fine-grained factuality in summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1449–1462, Online. Association for Computational Linguistics.
  16. 16.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News summarization and evaluation in the era of gpt-3. arXiv preprint arXiv:2209.12356.
  17. 17.Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018. Soft layer-specific multi-task summarization with entailment and question generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 687–697, Melbourne, Australia. Association for Computational Linguistics.
  18. 18.Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching Machines to Read and Comprehend. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS).
  19. 19.Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018. A unified model for extractive and abstractive summarization using inconsistency loss. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 132–141, Melbourne, Australia. Association for Computational Linguistics.
  20. 20.Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang. 2020. What have we achieved on text summarization? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 446–469, Online. Association for Computational Linguistics.
  21. 21.Yichen Jiang and Mohit Bansal. 2018. Closed-book training to improve summarization encoder memory. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4067–4077, Brussels, Belgium. Association for Computational Linguistics.
  22. 22.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  23. 23.Wojciech Kryscinski, Romain Paulus, Caiming Xiong, and Richard Socher. 2018. Improving abstraction in text summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1808–1817, Brussels, Belgium. Association for Computational Linguistics.
  24. 24.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  25. 25.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  26. 26.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  27. 27.Yang Liu and Mirella Lapata. 2019. Text summarization with pretrained encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3730–3740, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023. Chatgpt as a factual inconsistency evaluator for abstractive text summarization. arXiv preprint arXiv:2303.15621.
  29. 29.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  30. 30.Rada Mihalcea and Paul Tarau. 2004. TextRank: Bringing order into text. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 404–411, Barcelona, Spain. Association for Computational Linguistics.
  31. 31.Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI’17, page 3075–3081. AAAI Press.
  32. 32.Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. 2021a. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2727–2733, Online. Association for Computational Linguistics.
  33. 33.Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, and Bing Xiang. 2021b. Improving factual consistency of abstractive summarization via question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6881–6894, Online. Association for Computational Linguistics.
  34. 34.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  35. 35.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  37. 37.Ramakanth Pasunuru and Mohit Bansal. 2018. Multi-reward reinforced summarization with saliency and entailment. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 646–653, New Orleans, Louisiana. Association for Computational Linguistics.
  38. 38.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. OpenAI Blog.
  39. 39.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  40. 40.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  41. 41.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada. Association for Computational Linguistics.
  42. 42.Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G Nestor, Ali Soroush, Pierre A. Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin Rousseau, Chunhua Weng, and Yifan Peng. 2023. Evaluating large language models on medical evidence summarization.
  43. 43.Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li, Asli Celikyilmaz, Yashar Mehdad, and Dragomir Radev. 2022. CONFIT: Toward faithful dialogue summarization with linguistically-informed contrastive fine-tuning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5657–5668, Seattle, United States. Association for Computational Linguistics.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  45. 45.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  46. 46.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  47. 47.Yuxiang Wu and Baotian Hu. 2018. Learning to extract coherent summary via deep reinforcement learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’18/IAAI’18/EAAI’18. AAAI Press.
  48. 48.Zhiyuan Zeng, Jiaze Chen, Weiran Xu, and Lei Li. 2021. Gradient-based adversarial factual consistency evaluation for abstractive summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4102–4108, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  49. 49.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  50. 50.Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao, and Mohit Bansal. 2021. EmailSum: Abstractive email thread summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6895–6909, Online. Association for Computational Linguistics.
  51. 51.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  52. 52.Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D. Manning, and Curtis Langlotz. 2020. Optimizing the factual correctness of a summary: A study of summarizing radiology reports. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5108–5120, Online. Association for Computational Linguistics.
  53. 53.Zheng Zhao, Shay B. Cohen, and Bonnie Webber. 2020. Reducing quantity hallucinations in abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2237–2249, Online. Association for Computational Linguistics.
  54. 54.Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. 2018. Neural document summarization by jointly learning to score and select sentences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 654–663, Melbourne, Australia. Association for Computational Linguistics.

Citation

MLA
Tang, L., et al. “Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11626–44, https://doi.org/10.18653/v1/2023.acl-long.650.
APA
Tang, L., Goyal, T., Fabbri, A., Laban, P., Xu, J., Yavuz, S., Kryściński, W., Rousseau, J., & Durrett, G. (2023). Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11626–11644. https://doi.org/10.18653/v1/2023.acl-long.650
Chicago
Tang, L., T. Goyal, A. Fabbri, et al. 2023. “Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11626–44. https://doi.org/10.18653/v1/2023.acl-long.650.
Harvard
Tang, L. et al. (2023) “Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11626–11644. Available at: https://doi.org/10.18653/v1/2023.acl-long.650.
Vancouver
1. Tang L, Goyal T, Fabbri A, Laban P, Xu J, Yavuz S, Kryściński W, Rousseau J, Durrett G (2023) Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11626–11644

BibTeX

@inproceedings{tang-etal-2023-understanding,
    title = "Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors",
    author = "Tang, Liyan  and
      Goyal, Tanya  and
      Fabbri, Alex  and
      Laban, Philippe  and
      Xu, Jiacheng  and
      Yavuz, Semih  and
      Kryscinski, Wojciech  and
      Rousseau, Justin  and
      Durrett, Greg",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.650/",
    doi = "10.18653/v1/2023.acl-long.650",
    pages = "11626--11644"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/