TRUE: Re-evaluating Factual Consistency Evaluation

Or HonovichRoee AharoniJonathan HerzigHagai TaitelbaumDoron KuklianskyVered CohenThomas ScialomIdan SzpektorAvinatan HassidimYossi Matias

article2022NAACL346 citations

Establishes a standardized benchmark and example-level evaluation protocol across eleven datasets to measure how reliably factual consistency metrics detect grounded text generation errors.

Listen

Automated text generation systems frequently produce outputs that are factually inconsistent with their source materials or contain outright fabrications. While automatic evaluation metrics are crucial for filtering inaccurate outputs, cleaning training data, and accelerating model development, research in this area has historically been fragmented across isolated tasks and individual datasets. Furthermore, traditional meta-evaluations rely heavily on system-level correlation scores, which fail to reveal how reliably a metric performs when making binary consistency decisions on individual examples.

The main objective of the article is to establish a standardized benchmark for factual consistency evaluation and systematically assess the performance of existing automated metrics across diverse text generation domains. To achieve this, the authors introduced the TRUE benchmark, which consolidates 11 human-annotated datasets across abstractive summarization, dialogue generation, fact verification, and paraphrasing into a unified binary evaluation format. Rather than relying on system-level correlations, the assessment evaluates 12 metrics using the Area Under the Receiver Operating Characteristic Curve (ROC AUC), providing a direct, interpretable measure of an evaluation tool's ability to distinguish factually consistent from inconsistent text.

The evaluation revealed several key findings regarding metric effectiveness and behavior. First, large-scale Natural Language Inference (NLI) models and Question Generation and Answering (QG-QA) frameworks performed substantially better than standard baselines, with top approaches—such as ANLI, SummaC, and Q2—achieving average ROC AUC scores exceeding 80 across the evaluated benchmarks. In contrast, standard word-overlap metrics like ROUGE and BLEU, along with general learned metrics, performed poorly, averaging ROC AUC scores of 72 or below. Second, NLI and QG-QA methods proved to be complementary; combining them into an ensemble average improved performance by approximately 4.5 points in ROC AUC over any single method. Third, model capacity directly impacts accuracy, as scaling up the underlying model architecture increased average ROC AUC by up to 4.7 points. Finally, metric reliability degraded noticeably across all approaches when source grounding texts exceeded 200 tokens.

These findings indicate that organizations deploying text generation systems do not need to build siloed, task-specific evaluators, as unified NLI and QG-QA frameworks can serve as robust, general-purpose quality filters. Implementing these automated checks reduces the operational and reputational risks associated with deploying ungrounded model outputs in production environments. Error analysis showed that these metrics are occasionally more reliable than human annotators on difficult examples, underscoring their practical value for automated data cleaning.

The article recommends that system and metric developers adopt large-scale NLI and QG-QA methods as the standard baseline for evaluating factual consistency, and report ROC AUC rather than correlation metrics to assess example-level reliability. Organizations seeking peak performance should implement ensemble methods combining NLI and QG-QA, while accepting the higher computational cost of larger models. Despite these strengths, practitioners should remain cautious when applying these metrics to long context documents or conversational text containing subjective, non-factual statements, as both conditions currently increase error rates and require further targeted development.

Cover for TRUE: Re-evaluating Factual Consistency Evaluation

Abstract

Grounded text generation systems often generate text that contains factual inconsistencies, hindering their real-world applicability. Automatic factual consistency evaluation may help alleviate this limitation by accelerating evaluation cycles, filtering inconsistent outputs and augmenting training data. While attracting increasing attention, such evaluation metrics are usually developed and evaluated in silo for a single task or dataset, slowing their adoption. Moreover, previous meta-evaluation protocols focused on system-level correlations with human annotations, which leave the example-level accuracy of such metrics unclear. In this work, we introduce TRUE: a comprehensive survey and assessment of factual consistency metrics on a standardized collection of existing texts from diverse tasks, manually annotated for factual consistency. Our standardization enables an example-level meta-evaluation protocol that is more actionable and interpretable than previously reported correlations, yielding clearer quality measures. Across diverse state-of-the-art metrics and 11 datasets we find that large-scale NLI and question generation-and-answering-based approaches achieve strong and complementary results. We recommend those methods as a starting point for model and metric developers, and hope TRUE will foster progress towards even better evaluation methods.

Table of Contents

  • 1 Introduction
  • 2 Standardizing Factual Consistency
  • 2.1 Definitions and Terminology
  • 2.2 Standardization Process
  • 2.2.1 Abstractive Summarization
  • 2.2.2 Dialogue Generation
  • 2.2.3 Fact Verification
  • 2.2.4 Paraphrase Detection
  • 2.3 Meta-Evaluation
  • 3 Evaluation Metrics
  • 3.1 N-Gram Based Metrics
  • 3.2 Model-Based Metrics
  • 3.3 Natural Language Inference Metrics
  • 3.4 QG-QA Based Metrics
  • 4 Results
  • 5 Analysis
  • 6 Related Work
  • 7 Discussion and Future Work
  • 8 Conclusions
  • Acknowledgements
  • References
  • A Additional Data Statistics
  • B Implementation Details
  • C ROC Curves

Knowls

  1. Knowl 1 — Strict Definition of Grounded Factual Consistency

    definition

    In grounded text generation, a generated (target) text is defined as factually consistent (or faithful, grounded, factual) with respect to a source grounding text if all factual information conveyed by the generated text is strictly supported by the information present in the grounding text. Under this strict formulation:

    1. External real-world truth is disregarded: truthfulness is assessed entirely relative to the provided source text rather than external knowledge bases.
    2. Non-factual statements—such as subjective opinions, conversational chit-chat, and personal claims (e.g., "I've never heard of stamps")—are excluded from the scope of factual information and should not be penalized as ungrounded.
    3. A full target text is considered factually inconsistent if any part or sentence within it contains an ungrounded or contradictory factual assertion.
  2. Knowl 2 — The TRUE Benchmark Dataset Standardization

    experimental setup

    The TRUE benchmark consolidates 11 human-annotated datasets spanning four distinct text-to-text natural language generation tasks into a unified, binary-labeled format where each sample consists of a source grounding text, a generated candidate text, and a binary label indicating whether the entire candidate is factually consistent with the source.

    The standardized datasets comprise:

    • Summarization:
      • FRANK (671 examples, 33.2% consistent): Model summaries from CNN/DailyMail and XSum annotated via a typology of factual errors, mapped to binary by majority vote across sentences.
      • SummEval (1,600 examples, 81.6% consistent): CNN/DailyMail summaries rated by expert annotators on a 1–5 scale; labeled consistent only if all expert annotators assign a score of 5.
      • MNBM (2,500 examples, 10.2% consistent): XSum summaries annotated for hallucination; labeled consistent if majority of raters mark no hallucination.
      • QAGS-CNNDM (235 examples, 48.1% consistent) and QAGS-XSum (239 examples, 48.5% consistent): Sentence-level consistency judgments on CNN/DailyMail and XSum summaries mapped to binary by requiring all sentences to be consistent.
    • Knowledge-Grounded Dialogue:
      • BEGIN (836 examples, 33.7% consistent): Wizard of Wikipedia responses mapped from NLI categories (entailment vs. hallucination/off-topic/generic) where only entailed responses are labeled consistent.
      • Q2Q^2 (1,088 examples, 57.7% consistent): Binary annotations of factual consistency for responses generated from Wizard of Wikipedia.
      • DialFact (8,689 examples, 38.5% consistent): Conversational claims paired with Wikipedia evidence; verifiable claims labeled supported are mapped to consistent, while refuted or unverified claims are mapped to inconsistent.
    • Fact Verification:
      • FEVER (18,209 examples, 35.1% consistent): The NLI development split of FEVER; supported claims are labeled consistent, refuted and NotEnoughInfo are labeled inconsistent.
      • VitaminC (63,054 examples, 49.9% consistent): Wikipedia revision-based fact verification; supported facts are labeled consistent, refuted/neutral facts are labeled inconsistent.
    • Paraphrasing:
      • PAWS (8,000 examples, 44.2% consistent): High-lexical-overlap paraphrase pairs from Wikipedia where paraphrase pairs are mapped to consistent and non-paraphrase pairs to inconsistent.
  3. Knowl 3 — Example-Level ROC AUC Meta-Evaluation Protocol for Factual Consistency

    model/method

    Instead of evaluating factual consistency metrics using system-level rank correlations (e.g., Pearson or Spearman correlations), the meta-evaluation protocol evaluates metric quality on binary example-level factual consistency decisions using the Receiver Operating Characteristic Area Under the Curve (ROC AUC) for inconsistent example detection.

    Evaluating the ROC AUC measures the trade-off between the True Positive Rate (TPR, recall of inconsistent outputs) and the False Positive Rate (FPR, the rate of falsely marking consistent outputs as inconsistent) across all classification thresholds without requiring a fixed operational threshold.

    When a hard binary decision is required on datasets with designated development and test splits, the decision threshold τ∗\tau^* is tuned on the development set by maximizing the geometric mean of the true positive rate and the true negative rate:

    τ∗=arg⁡max⁡τTPR(τ)⋅(1−FPR(τ))\tau^* = \arg\max_{\tau} \sqrt{\text{TPR}(\tau) \cdot (1 - \text{FPR}(\tau))}

    Accuracy is then evaluated on the test set using τ∗\tau^*.

  4. Knowl 4 — Zero-Shot Performance Comparison of Factual Consistency Metrics Across NLG Tasks

    empirical result

    Across 11 standardized datasets evaluated in a zero-shot configuration, large-scale Natural Language Inference (NLI) and Question Generation and Answering (QG-QA) models substantially outperform standard NLG evaluation metrics and surface-level token matching.

    Dataset ANLI SCZS\text{SC}_{\text{ZS}} Q2Q^2 BARTScore BERTScore QuestEval BLEURT FactCC Token F1
    FRANK 89.4 89.1 87.8 86.1 84.3 84.0 82.8 76.4 76.1
    SummEval 80.5 81.7 78.8 73.5 77.2 70.1 66.7 75.9 61.4
    MNBM 77.9 71.3 68.7 60.9 62.8 65.3 64.5 59.4 46.2
    QAGS-C 82.1 80.9 83.5 80.9 69.1 64.2 71.6 76.4 63.8
    QAGS-X 83.8 78.1 70.9 53.8 49.5 56.3 57.2 64.9 51.1
    BEGIN 82.6 82.0 79.7 86.3 87.9 84.1 86.4 64.4 86.4
    Q2Q^2 dataset 72.7 77.4 80.9 64.9 70.0 72.2 72.4 63.7 65.9
    DialFact 77.7 84.1 86.1 65.6 64.2 77.3 73.1 55.3 72.3
    PAWS 86.4 88.2 89.7 77.5 77.5 69.2 68.3 64.0 51.1
    FEVER 93.2 93.2 88.4 64.1 63.3 72.6 59.5 61.9 51.8
    VitaminC 88.3 97.9 81.4 63.2 62.5 66.5 61.8 56.3 61.4
    Avg. (excl. VitC, FEVER) 81.5 81.4 80.7 72.2 71.4 71.4 71.4 66.7 63.8

    Key observations:

    • NLI-based models (ANLI with 81.5 average AUC, SummaC SCZS\text{SC}_{\text{ZS}} with 81.4 average AUC) and the QG-QA method Q2Q^2 (80.7 average AUC) achieve the strongest results across all tasks.
    • Standard generation metrics achieve significantly lower average AUCs: BARTScore (72.2), BERTScore-precision (71.4), BLEURT-20 (71.4), and FactCC (66.7).
    • Token overlap (F1) fails generally (63.8 average AUC) except on BEGIN (86.4 AUC), where high lexical disparities exist between consistent and inconsistent responses.
  5. Knowl 5 — Complementarity of NLI and QG-QA Metrics via Model Ensembling

    empirical result

    NLI-based and QG-QA-based consistency evaluation methods capture complementary signals of factual inconsistency. Averaging the predicted consistency probabilities of an end-to-end NLI model (ANLI fine-tuned on T5-11B), a sentence-level aggregation NLI model (SummaC SCZS\text{SC}_{\text{ZS}}), and a QG-QA pipeline (Q2Q^2) yields an ensemble average ROC AUC of 86.0 on the TRUE benchmark (excluding VitaminC and FEVER), outperforming the best single metric (ANLI at 81.5) by 4.5 ROC AUC points.

    In an analysis of 25,761 instances where individual metrics disagreed but the ensemble succeeded:

    • In 85.2% of instances, two out of the three metrics made the correct prediction while only one failed.
    • In 14.6% of instances, only one metric succeeded and two failed, but the ensemble prediction remained correct due to threshold calibration.
    • In 47% of analyzed single-metric failure cases, the failing metric assigned a borderline score near the decision boundary.
  6. Knowl 6 — Scaled Implementations of T5-11B Based NLI and QG-QA Evaluators

    model/method

    The top-performing single evaluators are constructed by scaling NLI and QG-QA architectures using T5-11B:

    1. End-to-End ANLI Evaluator: A T5-11B model fine-tuned on the Adversarial NLI (ANLI) dataset for 25,000 steps with a batch size of 32, a learning rate of 10−410^{-4}, and a maximum input length of 2,048 tokens. At evaluation time, the grounding text is formatted as the premise and the generated candidate text as the hypothesis; the model's output entailment probability serves as the factual consistency score.
    2. Scaled Q2Q^2 Pipeline: A question generation and question answering pipeline where T5-11B serves as the single backbone for Question Generation (QG), Question Answering (QA), and NLI validation.
      • Questions are generated from noun phrase / entity spans in the candidate text using beam search (beam size 4).
      • Generated questions are validated against the original candidate answer span using token-level F1 overlap (threshold set to 0.54 tuned on a held-out dataset) rather than exact matching.
      • The QA model answers valid questions against the grounding text.
      • An NLI model scores whether the QA model's answer from the grounding entails the extracted answer candidate from the generated text.
  7. Knowl 7 — Impact of Source Grounding Text Length on Evaluation Metric Performance

    empirical result

    Evaluating factual consistency degrades as the length of the source grounding text increases beyond 200 tokens across all metric families, including sentence-level aggregation methods designed for long documents.

    When grouping TRUE benchmark instances into grounding length bins (0–50, 50–100, 100–200, 200–350, 350–500, and 500+ tokens with 1,000 sampled examples per bin):

    • Across ANLI, Q2Q^2, and SummaC SCZS\text{SC}_{\text{ZS}}, ROC AUC is highest for short context windows (le\\le 200 tokens), peaking in the 0–50 and 100–200 token ranges.
    • A consistent drop in ROC AUC occurs for grounding texts in the 200–350 and 350–500 token bins across all three metrics.
    • For grounding lengths above 500 tokens, performance remains above 0.825 ROC AUC for both end-to-end ANLI (T5-11B) and Q2Q^2, despite these architectures processing long contexts directly.
  8. Knowl 8 — Scaling Model Parameters Improves Factual Consistency Evaluation

    empirical result

    Increasing the parameter capacity of underlying pretrained language models directly improves the accuracy and ROC AUC of automatic factual consistency evaluation metrics.

    Empirical comparisons across model variants on the TRUE benchmark demonstrate:

    Model Variant Average ROC AUC
    ANLI-T5-11B 81.5 (+4.7)
    ANLI-T5-Large 76.8
    BLEURT-20 71.4 (+3.7)
    BLEURT-20-D6 67.7
    BERTScore Precision (DeBERTa-xl-MNLI) 71.4 (+1.3)
    BERTScore Precision (RoBERTa-large) 70.1

    Upgrading the backbone from T5-Large to T5-11B yields a +4.7 ROC AUC gain for ANLI; moving from BLEURT-20-D6 to BLEURT-20 yields a +3.7 ROC AUC gain; and switching from RoBERTa-large to DeBERTa-xl-MNLI for BERTScore precision yields a +1.3 ROC AUC gain.

  9. Knowl 9 — Failure Modes and Label Noise in Automatic Consistency Evaluation

    limitation

    Error analysis of top-performing factual consistency metrics reveals three major sources of classification errors:

    1. Subtle Inconsistencies and Score Dilution: Metrics that average question-level or sentence-level scores (Q2Q^2, SummaC SCZS\text{SC}_{\text{ZS}}) frequently fail when a candidate text is mostly consistent but contains a single subtle hallucination, because the averaged score remains high. Full-sequence NLI can also fail to predict contradiction when only a minor sub-span is ungrounded.
    2. Non-Factual and Subjective Dialogue Utterances: Conversational elements expressing personal experiences or opinions (e.g., "I've never heard of stamps") in knowledge-grounded dialogue are frequently misclassified by NLI and QG-QA metrics as ungrounded hallucinations (occurring in 10 out of 62 analyzed dialogue error cases).
    3. Human Annotation Errors on Difficult Samples: Among benchmark instances where all top three metrics (ANLI, SCZS\text{SC}_{\text{ZS}}, Q2Q^2) failed, 43.8% (35/80) were caused by erroneous human labels in the underlying datasets, compared to a base error rate of 10% (10/100) on uniformly sampled benchmark instances.

Coverage note — Pairwise metric ensemble combinations and appendix-only test-set accuracy/AUC breakdown tables were omitted in favor of the full three-way ensemble and primary dev/test evaluation tables.

References

  1. 1.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  2. 2.Gillian Brown and George Yule. 1983. Discourse Analysis. Cambridge University Press.
  3. 3.Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021. Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2082–2095, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The pascal recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, pages 177–190, Berlin, Heidelberg. Springer Berlin Heidelberg.
  5. 5.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  6. 6.Emily Denton, Mark Díaz, Ian Kivlichan, Vinodkumar Prabhakaran, and Rachel Rosen. 2021. Whose ground truth? accounting for individual and collective identities underlying dataset annotation. arXiv preprint arXiv:2112.04554.
  7. 7.Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021. Towards question-answering as an automatic metric for evaluating the content quality of a summary. Transactions of the Association for Computational Linguistics, 9:774–789.
  8. 8.Daniel Deutsch, Rotem Dror, and Dan Roth. 2022. Re-examining system-level correlations of automatic summarization evaluation metrics. arXiv preprint arXiv:2204.10216.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of Wikipedia: Knowledge-powered conversational agents. In Proceedings of the International Conference on Learning Representations (ICLR).
  11. 11.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  12. 12.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2021. Evaluating groundedness in dialogue systems: The begin benchmark.
  13. 13.Bradley Efron. 1982. The jackknife, the bootstrap and other resampling plans. SIAM.
  14. 14.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2020. Summeval: Re-evaluating summarization evaluation. arXiv preprint arXiv:2007.12626.
  15. 15.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021a. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  16. 16.Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021b. Qafacteval: Improved qa-based factual consistency evaluation for summarization.
  17. 17.Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2214–2220, Florence, Italy. Association for Computational Linguistics.
  18. 18.C. Fillmore. 1976. Frame semantics and the nature of language *. Annals of the New York Academy of Sciences, 280.
  19. 19.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. Experts, errors, and context: A large-scale study of human evaluation for machine translation. arXiv preprint arXiv:2104.14478.
  20. 20.Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021. GO FIGURE: A meta evaluation of factuality in summarization. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 478–487, Online. Association for Computational Linguistics.
  21. 21.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  22. 22.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021. Dialfact: A benchmark for fact-checking in dialogue. arXiv preprint arXiv:2110.08222.
  23. 23.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
  24. 24.Martin Heidegger. 2001. On the essence of truth. The Nature of Truth: Classic and Contemporary Perspectives, 1:295–316.
  25. 25.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  26. 26.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. q2q^2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7856–7870, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  27. 27.J. Edward Hu, Abhinav Singh, Nils Holzenberger, Matt Post, and Benjamin Van Durme. 2019. Large-scale, diverse, paraphrastic bitexts via sampling and clustering. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 44–54, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  29. 29.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  30. 30.Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. 2021. Summac: Re-visiting nli-based models for inconsistency detection in summarization. arXiv preprint arXiv:2111.09525.
  31. 31.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942.
  32. 32.Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018. Hallucinations in neural machine translation. NeurIPS 2018 Workshop on Interpretability and Robustness for Audio, Speech, and Language.
  33. 33.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  34. 34.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  36. 36.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  37. 37.Nikita Nangia, Adina Williams, Angeliki Lazaridou, and Samuel Bowman. 2017. The RepEval 2017 shared task: Multi-genre natural language inference with sentence representations. In Proceedings of the 2nd Workshop on Evaluating Vector Space Representations for NLP, pages 1–10, Copenhagen, Denmark. Association for Computational Linguistics.
  38. 38.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  39. 39.Yixin Nie, Haonan Chen, and Mohit Bansal. 2019. Combining fact extraction and verification with neural semantic matching networks. In Association for the Advancement of Artificial Intelligence (AAAI).
  40. 40.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, Online. Association for Computational Linguistics.
  41. 41.Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021. I like fish, especially dolphins: Addressing contradictions in dialogue modeling. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1699–1713, Online. Association for Computational Linguistics.
  42. 42.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  43. 43.Martha Palmer, Daniel Gildea, and Paul Kingsbury. 2005. The Proposition Bank: An Annotated Corpus of Semantic Roles. Computational Linguistics, 31(1):71–106.
  44. 44.Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, et al. 2021. Quality: Question answering with long input texts, yes! arXiv preprint arXiv:2112.08608.
  45. 45.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  46. 46.Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. ToTTo: A controlled table-to-text generation dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1173–1186, Online. Association for Computational Linguistics.
  47. 47.Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021. Learning compact metrics for MT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 751–762, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  48. 48.Libo Qin, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu, and Wanxiang Che. 2021. Don’t be contradicted with anything! CI-ToD: Towards benchmarking consistency for task-oriented dialogue system. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2357–2367, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  50. 50.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021a. Measuring attribution in natural language generation models. arXiv preprint arXiv:2112.12870.
  51. 51.Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021b. Increasing faithfulness in knowledge-grounded dialogue with controllable features. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 704–718, Online. Association for Computational Linguistics.
  52. 52.Ehud Reiter and Craig Thomson. 2020. Shared task on evaluating accuracy. In Proceedings of the 13th International Conference on Natural Language Generation, pages 227–231, Dublin, Ireland. Association for Computational Linguistics.
  53. 53.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045.
  54. 54.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin C! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 624–643, Online. Association for Computational Linguistics.
  55. 55.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  56. 56.Thomas Scialom and Felix Hill. 2021. Beametrics: A benchmark for language generation evaluation evaluation. arXiv preprint arXiv:2110.09147.
  57. 57.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020a. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  58. 58.Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur Parikh. 2020b. Learning to evaluate translation beyond English: BLEURT submissions to the WMT metrics 2020 shared task. In Proceedings of the Fifth Conference on Machine Translation, pages 921–927, Online. Association for Computational Linguistics.
  59. 59.Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, and Omer Levy. 2022. Scrolls: Standardized comparison over long language sequences.
  60. 60.James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. 2018. The fact extraction and VERification (FEVER) shared task. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER), pages 1–9, Brussels, Belgium. Association for Computational Linguistics.
  61. 61.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  62. 62.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  63. 63.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3731–3741, Florence, Italy. Association for Computational Linguistics.
  64. 64.Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding. 2021. Factual consistency evaluation for text summarization via counterfactual estimation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 100–110, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  65. 65.Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021. A comprehensive assessment of dialog evaluation metrics. In The First Workshop on Evaluations and Assessments of Neural Conversation Systems, pages 15–33, Online. Association for Computational Linguistics.
  66. 66.Wenpeng Yin, Dragomir Radev, and Caiming Xiong. 2021. DocNLI: A large-scale dataset for document-level natural language inference. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4913–4922, Online. Association for Computational Linguistics.
  67. 67.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems.
  68. 68.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  69. 69.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.
  70. 70.Zheng Zhao, Shay B Cohen, and Bonnie Webber. 2020. Reducing quantity hallucinations in abstractive summarization. arXiv preprint arXiv:2009.13312.
  71. 71.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021. Detecting hallucinated content in conditional neural sequence generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1393–1404, Online. Association for Computational Linguistics.

Citation

MLA
Honovich, O., et al. “TRUE: Re-evaluating Factual Consistency Evaluation”. arXiv, 2022, http://arxiv.org/abs/2204.04991v3.
APA
Honovich, O., Aharoni, R., Herzig, J., Taitelbaum, H., Kukliansy, D., Cohen, V., Scialom, T., Szpektor, I., Hassidim, A., & Matias, Y. (2022). TRUE: Re-evaluating Factual Consistency Evaluation. arXiv. http://arxiv.org/abs/2204.04991v3
Chicago
Honovich, O., R. Aharoni, J. Herzig, et al. 2022. “TRUE: Re-evaluating Factual Consistency Evaluation”. arXiv. http://arxiv.org/abs/2204.04991v3.
Harvard
Honovich, O. et al. (2022) “TRUE: Re-evaluating Factual Consistency Evaluation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.04991v3.
Vancouver
1. Honovich O, Aharoni R, Herzig J, Taitelbaum H, Kukliansy D, Cohen V, Scialom T, Szpektor I, Hassidim A, Matias Y (2022) TRUE: Re-evaluating Factual Consistency Evaluation. arXiv

BibTeX

@article{honovich2022true,
  title = {TRUE: Re-evaluating Factual Consistency Evaluation},
  author = {Honovich, Or and Aharoni, Roee and Herzig, Jonathan and Taitelbaum, Hagai and Kukliansy, Doron and Cohen, Vered and Scialom, Thomas and Szpektor, Idan and Hassidim, Avinatan and Matias, Yossi},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.04991v3},
  eprint = {2204.04991}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/