Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation

Yixin LiuAlexander R. FabbriPengfei LiuYilun ZhaoLinyong NanRuilin HanSimeng HanShafiq JotyChien-Sheng WuCaiming Xiong

article2023ACL187 citations

Presents the RoSE benchmark and an Atomic Content Unit protocol across 22,000 human annotations to reveal how annotator biases distort summary evaluations and benchmark 50 automatic metrics against state-of-the-art systems.

Listen

Human judgment is widely regarded as the ultimate standard for evaluating text summarization models and the automated metrics that assess them. However, current human evaluation studies often suffer from low consistency between annotators, unrepresentative sample sizes, and a lack of statistical reliability. With the rapid deployment of large language models, poorly standardized evaluations risk misleading researchers and decision-makers by conflating subjective reader preferences with genuine factual accuracy and summary quality.

To address this challenge, the article aims to establish a more objective, high-agreement human evaluation framework and benchmark to robustly assess modern summarization systems and automated metrics. Specifically, it demonstrates how decomposing summaries into fine-grained factual units enhances measurement consistency and statistical reliability.

The authors developed the Atomic Content Unit protocol, an approach that breaks reference summaries into elementary factual statements and asks human reviewers to verify their presence in candidate summaries using binary decisions. Using this protocol, the authors built the Robust Summarization Evaluation benchmark, compiling 22,000 summary-level annotations across 28 leading summarization models and three standard datasets (covering news and conversational dialogue). The authors then compared four distinct human evaluation methods and evaluated 50 automated metric variants, analyzing their statistical power and consistency.

The analysis yielded several critical findings. First, breaking summaries into atomic factual units produced substantially higher annotator agreement (0.75 agreement score) than traditional holistic human scoring methods (which scored between 0.22 and 0.35). Second, conventional human evaluation sample sizes of 50 to 100 examples lack sufficient statistical power to reliably distinguish between competitive, top-performing models. Third, open-ended human evaluations without source documents strongly favored longer outputs and large language models like GPT-3 due to inherent annotator biases toward fluency and style, even when those summaries missed key factual information. Finally, automated evaluation metrics based on large language models failed to outperform established overlap-based metrics on the benchmark and exhibited weak summary-level calibration.

These findings indicate that general human satisfaction ratings can be heavily confounded by text length and surface-level fluency rather than content completeness. This presents substantial risks for organizations relying on general user feedback to deploy language models in mission-critical applications where factual coverage is essential. Relying on underpowered or improperly designed human studies may lead teams to adopt models that sound convincing but omit critical information.

The article recommends that organizations implement targeted evaluation protocols with explicitly defined quality criteria, such as length constraints and factual coverage, rather than broad preference surveys. For assessing factual overlap, evaluations should use structured unit-matching methods and sufficiently large sample sizes—often hundreds of examples—to ensure statistical significance. Automated metrics must also be aligned strictly with the specific dimension they are intended to track.

The study's primary limitations include its exclusive focus on English-language texts, the absence of unit-importance weighting, and potential underlying demographic biases among crowdsourced workers. Nevertheless, confidence in the central conclusions remains high due to the extensive sample size, multi-dataset coverage, and rigorous statistical bootstrapping used throughout the research.

arXiv: 2212.07981Yale-LILY/ROSE
Cover for Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation

Abstract

Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator agreement or have insufficient scale, and an in-depth analysis of human evaluation is lacking. Therefore, we address the shortcomings of existing summarization evaluation along the following axes: (1) We propose a modified summarization salience protocol, Atomic Content Units (ACUs), which is based on fine-grained semantic units and allows for a high inter-annotator agreement. (2) We curate the Robust Summarization Evaluation (RoSE) benchmark, a large human evaluation dataset consisting of 22,000 summary-level annotations over 28 top-performing systems on three datasets. (3) We conduct a comparative study of four human evaluation protocols, underscoring potential confounding factors in evaluation setups. (4) We evaluate 50 automatic metrics and their variants using the collected human annotations across evaluation protocols and demonstrate how our benchmark leads to more statistically stable and significant results. The metrics we benchmarked include recent methods based on large language models (LLMs), GPTScore and G-Eval. Furthermore, our findings have important implications for evaluating LLMs, as we show that LLMs adjusted by human feedback (e.g., GPT-3.5) may overfit unconstrained human evaluation, which is affected by the annotators’ prior, input-agnostic preferences, calling for more robust, targeted evaluation methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Atomic Content Units for Summarization Evaluation
  • 3.1 Preliminaries
  • 3.2 ACU Annotation Protocol
  • 3.3 ACU Annotation Collection
  • 4 RoSE Benchmark Analysis
  • 4.1 Power Analysis
  • 4.2 Summarization System Analysis
  • 5 Evaluating Annotation Protocols
  • 5.1 Annotation Collection
  • 5.2 Results Analysis
  • 6 Evaluating Automatic Metrics
  • 6.1 Metric Evaluation with ACU Annotations
  • 6.2 Analysis of Metric Evaluation
  • 7 Conclusion and Implications
  • 8 Limitations
  • Acknowledgements
  • References
  • A Benchmark Data Collection
  • A.1 Detailed Settings
  • A.2 Summarization Models
  • A.3 ACU Scores of Summarization Models
  • B Power Analysis
  • B.1 Detailed Settings
  • B.2 Powers of ACU Annotations
  • C Calculating Correlations
  • D.1 Data Collection Details
  • D.2 Results Analysis
  • E Metric Analysis
  • E.1 Metrics
  • E.2 Metrics based on Large Language Models
  • E.3 Metric Correlation with ACU Scores
  • E.4 System Pairs for Fine-grained Metric Evaluation
  • E.5 Confidence Interval
  • E.6 Power Analysis of Metric Comparison
  • E.7 Metric Correlation with Different Human Evaluation Protocols
  • F Human Evaluation Practices in Recent Text Summarization Research

Knowls

  1. Knowl 1 — Atomic Content Units make reference-based salience judgments fine-grained

    model/method

    The Atomic Content Unit (ACU) protocol evaluates summary salience by decomposing it into two tasks. First, an ACU writer extracts elementary facts from each reference-summary sentence: one unit captures the main-clause information, and additional units capture other facts with only the minimal context needed to make each fact interpretable. Second, annotators make a binary judgment about whether each reference ACU is present in a candidate system summary. The reference summary is provided as context during matching. Because ACUs are written once for each reference summary, the same set can be reused to evaluate multiple candidate summaries and new systems. In this study, the authors wrote the ACUs using shared guidelines and recruited qualified crowdworkers to perform matching.

  2. Knowl 2 — RoSE is a large, cross-dataset ACU benchmark with high annotator agreement

    data/table

    RoSE contains human evaluations of outputs from 28 recent summarization systems across CNN/DailyMail (CNNDM), XSum, and SamSum, plus CNNDM validation annotations intended to support metric training. The collection includes approximately 21.8k written ACUs, 22k aggregated summary-level annotations, and 50k individual summary-level judgments. Its split-level coverage is: CNNDM test, 500 documents, 12 systems, 5.6k ACUs, and 6k summary annotations; CNNDM validation, 1,000 documents, 8 systems, 11.6k ACUs, and 8k annotations; XSum test, 500 documents, 8 systems, 2.3k ACUs, and 4k annotations; SamSum test, 500 documents, 8 systems, 2.3k ACUs, and 4k annotations. Test-set summary annotations aggregate three judgments per summary; CNNDM validation summaries have one judgment each. Krippendorff’s alpha was 0.7571 for aggregated summary-level ACU matching and 0.7528 at the ACU level.

  3. Knowl 3 — ACU scores measure matched reference facts, with an optional penalty for excess length

    equation

    For a candidate summary ss and a set of reference ACUs AA, let As⊆AA_s \subseteq A be the ACUs judged present in ss. The unnormalized ACU score is recall over the reference units: f(s,A)=∣As∣/∣A∣f(s,A)=|A_s|/|A|. To penalize candidate summaries longer than the reference summary rr, the normalized score is f~α(s,A,r)=exp⁡(min⁡(0,1−(∣s∣/∣r∣)α))f(s,A)\tilde f_\alpha(s,A,r)=\exp(\min(0,1-(|s|/|r|)^\alpha))f(s,A), where ∣s∣|s| and ∣r∣|r| are the candidate and reference word counts, and α>0\alpha>0 controls the penalty strength. This normalization leaves summaries no longer than the reference unpenalized and reduces scores for longer summaries. The authors selected α\alpha by seeking to decorrelate normalized ACU scores from summary length; the chosen values were 2 for CNNDM, 5 for XSum, and 0.5 for SamSum.

  4. Knowl 4 — Unconstrained human ratings track annotator preference and summary length

    empirical result

    The authors compared four protocols on 100 CNNDM test examples, collecting three annotations per summary. Prior asks for a quality rating without showing the document or reference; Ref-free asks whether a summary captures the important information in the document; Ref-based rates information similarity to the reference; ACU checks reference facts individually. The first three protocols use ratings from 1 to 5, while ACU matching yields binary labels aggregated into scores. At the system level for fine-tuned models, Pearson correlations were 0.926 between Prior and Ref-free, 0.833 between Prior and summary length, and 0.875 between Ref-free and summary length. Ref-free and Ref-based correlated at -0.247; Prior and Ref-based at -0.061. Normalized ACU correlated 0.762 with Ref-based, -0.093 with Ref-free, and 0.048 with Prior; its correlation with length was -0.296. Thus, the two document-independent or document-based unconstrained ratings were strongly associated with each other and with length, while Ref-free and reference-based judgments did not behave as interchangeable measures of salience.

  5. Knowl 5 — GPT-3 scores highly under unconstrained ratings but not under reference-based judgments

    empirical result

    In the protocol comparison, GPT-3 (text-davinci-002) received the highest Prior and Ref-free scores among the evaluated systems: 3.72 and 3.76, respectively, on the 1–5 scales. Its Ref-based score was 2.74, below BRIO’s 3.07, and its unnormalized ACU score was 0.268, below BRIO’s 0.429. GPT-3 summaries averaged 69.5 words, compared with 66.4 for BRIO; reference summaries averaged 54.9 words and scored 2.85 on Prior and 2.94 on Ref-free. A case study of four annotators who each assessed around 20 examples found an average correlation of 0.404 between an annotator’s own Ref-free and Prior scores, versus 0.188 between that annotator’s Ref-free scores and the other annotators’ average Ref-free scores. These observations suggest that individual input-agnostic preferences can influence Ref-free judgments; they do not establish that every GPT-3 advantage under those protocols reflects such preference.

  6. Knowl 6 — Human evaluation often lacks power to distinguish similar systems

    empirical result

    The authors estimated power for pairwise system comparisons by repeatedly sampling examples with replacement, applying paired bootstrap significance tests, and measuring the fraction of trials with p<0.05p<0.05. On CNNDM, when system pairs differed by 1–2 ROUGE-1 recall points, a sample of 500 examples had approximately 0.50 power, compared with approximately 0.20 for 100 examples. The analysis found that samples of the size commonly used in recent studies—about 50–100 examples—reached 0.80 power only when the systems differed by more than 5 ROUGE-1 recall points. Across the three evaluated datasets, increasing sample size consistently increased power. The results indicate that human evaluations of similar-performing systems can miss real differences when based on small samples.

  7. Knowl 7 — Traditional metrics generally correlate better with ACU judgments than tested LLM metrics

    empirical result

    The authors compared automatic metrics with ACU scores using Kendall’s correlation at both system and summary levels, using recall-oriented metric scores when available. Selected system-level/summary-level results (CNNDM; XSum; SamSum) were: ROUGE-1, 0.788/0.468; 0.714/0.293; 0.929/0.439; BARTScore, 0.727/0.453; 0.714/0.282; 0.929/0.430; and Lite³Pyramid, 0.849/0.452; 0.714/0.245; 1.00/0.467. For GPTScore the corresponding values were 0.636/0.129; 0.214/0.099; 0.429/0.158; for G-Eval-3.5, 0.412/0.164; 0.429/0.136; 0.857/0.248; and for G-Eval-4, 0.779/0.274; 0.691/0.185; 0.929/0.405. Metric performance varied by dataset, with generally stronger correlations on SamSum and weaker ones on XSum. Overall, the tested LLM-based evaluators did not outperform traditional metrics on RoSE, and their summary-level correlations were comparatively weak.

  8. Knowl 8 — System summary length differences can be missed by ROUGE F1

    empirical result

    On the CNNDM test set, the reference summaries averaged 54.93 words, while every system in the authors’ system analysis produced longer summaries. GSum averaged 77.61 words and GLOBAL 55.50 words, a difference of about 40%, yet their ROUGE-1 F1 scores were similar: 45.47 and 45.17, respectively. The reported lengths and ROUGE-1 F1 scores were calculated over the full test set, whereas ACU and normalized ACU scores used the 500 annotated examples. This comparison shows that a widely used automatic metric may not reflect substantial differences in summary length; the authors therefore treat length as a distinct evaluation dimension rather than assuming that content scores capture it.

  9. Knowl 9 — Automatic metric rankings depend on the human evaluation protocol

    empirical result

    On CNNDM, system-level Kendall correlations between automatic metrics and human judgments differed substantially by protocol. For example, ROUGE-1 correlated at -0.061 with Prior, -0.212 with Ref-free, 0.840 with Ref-based, and 0.636 with normalized ACU. Lite³Pyramid correlated at 0.576 with Prior, 0.667 with Ref-free, -0.168 with Ref-based, and 0.121 with normalized ACU; QAEval correlated at 0.485, 0.515, -0.076, and 0.151, respectively. The pattern supports evaluating a metric against human judgments that target the same quality dimension: reference-based metrics tended to align better with reference-based judgments, while some had negative correlations with reference-free judgments. Since the human protocols themselves can be weakly correlated, the choice of protocol can change apparent metric performance.

  10. Knowl 10 — Larger evaluation samples stabilize metric comparisons as well as system comparisons

    empirical result

    For automatic metric evaluation, the authors found that bootstrap confidence intervals around system-level correlations with ACU scores were wide, but narrowed as sample size increased; the resampling analysis examined sample sizes from 50 to 1,000 examples. Their metric-comparison power analysis compared 20 metrics, forming 190 metric pairs, and used permutation tests for significance. As with system comparisons, significant differences were difficult to detect when metric performance was similar, while increasing sample size increased the chance of finding a significant difference. Thus, the scale of human annotations affects not only confidence in a metric’s estimated correlation but also the ability to distinguish competing metrics.

  11. Knowl 11 — RoSE has language, annotation, and reference-quality limitations

    limitation

    RoSE and the analyses cover English-language data only, and the authors caution that both annotator biases and biases inherited from model pretraining may affect results. ACU writing and matching can contain noise, and high inter-annotator agreement does not guarantee that judgments are correct. Reference-based evaluation also depends on reference quality: a dataset reference is not necessarily an objective gold standard. Unlike the original Pyramid approach, RoSE does not weight ACUs during aggregation, and its evaluation is based on one reference summary. The authors frame the benchmark as useful for studying targeted conditional generation and semantic overlap, while identifying higher-quality references and ACU weighting as directions for further work.

Coverage note — The appendix’s survey of human-evaluation practices in 55 recent papers and the exhaustive correlation tables for all metric variants are omitted: they provide secondary context and detailed variants beyond the main protocol, benchmark, and evaluation findings captured here.

References

  1. 1.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating evaluation in text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online. Association for Computational Linguistics.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Shuyang Cao and Lu Wang. 2021. CLIFF: Contrastive learning for improving faithfulness and factuality in abstractive summarization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6633–6649, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020. With little power comes great responsibility. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9263–9274, Online. Association for Computational Linguistics.
  5. 5.Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018. The price of debiasing automatic metrics in natural language evalaution. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 643–653, Melbourne, Australia. Association for Computational Linguistics.
  6. 6.Jiaao Chen and Diyi Yang. 2020. Multi-view sequence-to-sequence models with conversational structure for abstractive dialogue summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4106–4118, Online. Association for Computational Linguistics.
  7. 7.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  8. 8.Arman Cohan and Nazli Goharian. 2016. Revisiting summarization evaluation for scientific articles. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 806–813, Portorož, Slovenia. European Language Resources Association (ELRA).
  9. 9.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021a. Towards question-answering as an automatic metric for evaluating the content quality of a summary. Transactions of the Association for Computational Linguistics, 9:774–789.
  11. 11.Daniel Deutsch, Rotem Dror, and Dan Roth. 2021b. A statistical analysis of summarization evaluation metrics using resampling methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
  12. 12.Daniel Deutsch, Rotem Dror, and Dan Roth. 2022. Re-examining system-level correlations of automatic summarization evaluation metrics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 6038–6052, Seattle, United States. Association for Computational Linguistics.
  13. 13.Daniel Deutsch and Dan Roth. 2020. SacreROUGE: An open-source library for using and developing summarization evaluation metrics. In Proceedings of Second Workshop for NLP Open Source Software (NLP-OSS), pages 120–125, Online. Association for Computational Linguistics.
  14. 14.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13042–13054.
  15. 15.Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021. GSum: A general framework for guided neural abstractive summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4830–4842, Online. Association for Computational Linguistics.
  16. 16.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  17. 17.Alexander Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2022a. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9(0):391–409.
  18. 18.Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022b. QAFactEval: Improved QA-based factual consistency evaluation for summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2587–2601, Seattle, United States. Association for Computational Linguistics.
  19. 19.Xiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin, and Ting Liu. 2021. Language model as an annotator: Exploring DialoGPT for dialogue summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1479–1491, Online. Association for Computational Linguistics.
  20. 20.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. ArXiv, abs/2302.04166.
  21. 21.Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021. GO FIGURE: A meta evaluation of factuality in summarization. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 478–487, Online. Association for Computational Linguistics.
  22. 22.Mingqi Gao and Xiaojun Wan. 2022. DialSummEval: Revisiting summarization evaluation for dialogues. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5693–5709, Seattle, United States. Association for Computational Linguistics.
  23. 23.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.Yang Gao, Wei Zhao, and Steffen Eger. 2020. SUPERT: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1347–1354, Online. Association for Computational Linguistics.
  25. 25.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  26. 26.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. ArXiv preprint, abs/2202.06935.
  27. 27.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70–79, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News summarization and evaluation in the era of gpt-3.
  29. 29.Yvette Graham. 2015. Re-evaluating automatic summarization with BLEU and 192 shades of ROUGE. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 128–137, Lisbon, Portugal. Association for Computational Linguistics.
  30. 30.Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 708–719, New Orleans, Louisiana. Association for Computational Linguistics.
  31. 31.Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019. HighRES: Highlight-based reference-less evaluation of summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3381–3392, Florence, Italy. Association for Computational Linguistics.
  32. 32.Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani, and Caiming Xiong. 2020. Ctrlsum: Towards generic controllable text summarization.
  33. 33.Pengcheng He, Baolin Peng, Liyang Lu, Song Wang, Jie Mei, Yang Liu, Ruochen Xu, Hany Hassan Awadalla, Yu Shi, Chenguang Zhu, et al. 2022. Z-code++: A pre-trained language model optimized for abstractive summarization. ArXiv preprint, abs/2208.09770.
  34. 34.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating factual consistency evaluation. In Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Question Answering, pages 161–175, Dublin, Ireland. Association for Computational Linguistics.
  35. 35.Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang. 2020. What have we achieved on text summarization? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 446–469, Online. Association for Computational Linguistics.
  36. 36.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Dragomir Radev, Yejin Choi, and Noah A Smith. 2022a. Beam decoding with controlled patience. ArXiv preprint, abs/2204.05424.
  37. 37.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, and Noah A. Smith. 2022b. Bidimensional leaderboards: Generate and evaluate language hand in hand. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3540–3557, Seattle, United States. Association for Computational Linguistics.
  38. 38.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability.
  39. 39.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  40. 40.Matt J. Kusner, Yu Sun, Nicholas I. Kolkin, and Kilian Q. Weinberger. 2015. From word embeddings to document distances. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pages 957–966. JMLR.org.
  41. 41.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 20d. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization. ArXiv preprint, abs/d.
  42. 42.Alon Lavie and Abhaya Agarwal. 2007. METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation, pages 228–231, Prague, Czech Republic. Association for Computational Linguistics.
  43. 43.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  44. 44.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020b. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  45. 45.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2022. Holistic evaluation of language models.
  46. 46.Chin-Yew Lin. 2004a. Looking for a few good metrics: Automatic summarization evaluation-how many samples are enough? In NTCIR.
  47. 47.Chin-Yew Lin. 2004b. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  48. 48.Feifan Liu and Yang Liu. 2008. Correlation between ROUGE and human evaluation of extractive meeting summaries. In Proceedings of ACL-08: HLT, Short Papers, pages 201–204, Columbus, Ohio. Association for Computational Linguistics.
  49. 49.Yang Liu, Dan Iter, Yichong Xu, Shuo Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. ArXiv, abs/2303.16634.
  50. 50.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv preprint, abs/1907.11692.
  51. 51.Yixin Liu and Pengfei Liu. 2021. SimCLS: A simple framework for contrastive learning of abstractive summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 1065–1072, Online. Association for Computational Linguistics.
  52. 52.Yixin Liu, Pengfei Liu, Dragomir Radev, and Graham Neubig. 2022. BRIO: Bringing order to abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2890–2903, Dublin, Ireland. Association for Computational Linguistics.
  53. 53.Zhengyuan Liu and Nancy Chen. 2021. Controllable neural dialogue summarization with personal named entity planning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 92–106, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  54. 54.Li Lucy and David Bamman. 2021. Gender and representation bias in GPT-3 generated stories. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55, Virtual. Association for Computational Linguistics.
  55. 55.Ye Ma, Zixun Lan, Lu Zong, and Kaizhu Huang. 2021. Global-aware beam search for neural abstractive summarization. In Advances in Neural Information Processing Systems, volume 34, pages 16545–16557. Curran Associates, Inc.
  56. 56.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  57. 57.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290, Berlin, Germany. Association for Computational Linguistics.
  58. 58.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  59. 59.Shashi Narayan, Yao Zhao, Joshua Maynez, Gonçalo Simões, Vitaly Nikolaev, and Ryan McDonald. 20d. Planning with Learned Entity Prompts for Abstractive Summarization. ArXiv preprint, abs/d.
  60. 60.Ani Nenkova and Rebecca Passonneau. 2004. Evaluating content selection in summarization: The pyramid method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 145–152, Boston, Massachusetts, USA. Association for Computational Linguistics.
  61. 61.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  62. 62.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  63. 63.Karolina Owczarzak, Peter A. Rankel, Hoa Trang Dang, and John M. Conroy. 2012. Assessing the effect of inconsistent assessors on summarization evaluation. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 359–362, Jeju Island, Korea. Association for Computational Linguistics.
  64. 64.Richard Yuanzhe Pang and He He. 2021. Text generation by learning from demonstrations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  65. 65.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  66. 66.Maxime Peyrard. 2019. Studying summarization evaluation metrics in the appropriate scoring range. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5093–5100, Florence, Italy. Association for Computational Linguistics.
  67. 67.Maja Popovic. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  68. 68.Peter A. Rankel, John M. Conroy, Hoa Trang Dang, and Ani Nenkova. 2013. A decade of automatic content evaluation of news summaries: Reassessing the state of the art. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 131–136, Sofia, Bulgaria. Association for Computational Linguistics.
  69. 69.Keisuke Sakaguchi and Benjamin Van Durme. 2018. Efficient online scalar annotation with bounded support. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 208–218, Melbourne, Australia. Association for Computational Linguistics.
  70. 70.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  71. 71.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  72. 72.Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019. Answers unite! unsupervised metrics for reinforced summarization models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3246–3256, Hong Kong, China. Association for Computational Linguistics.
  73. 73.Ori Shapira, David Gabay, Yang Gao, Hadar Ronen, Ramakanth Pasunuru, Mohit Bansal, Yael Amsterdamer, and Ido Dagan. 2019. Crowdsourcing lightweight pyramids for manual summary evaluation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 682–687, Minneapolis, Minnesota. Association for Computational Linguistics.
  74. 74.Kaiqiang Song, Bingqing Wang, Zhe Feng, and Fei Liu. 2021. A new approach to overgenerating and scoring abstractive summaries. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1392–1404, Online. Association for Computational Linguistics.
  75. 75.Julius Steen and Katja Markert. 2021. How to evaluate a summarizer: Study design and statistical analysis for manual linguistic quality evaluation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1861–1875, Online. Association for Computational Linguistics.
  76. 76.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020. Learning to summarize from human feedback. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA. Curran Associates Inc.
  77. 77.Simeng Sun, Ori Shapira, Ido Dagan, and Ani Nenkova. 2019. How to compare summarizers without target length? pitfalls, solutions and re-examination of the neural summarization literature. In Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, pages 21–29, Minneapolis, Minnesota. Association for Computational Linguistics.
  78. 78.Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2022. Evaluating the factual consistency of large language models through summarization. ArXiv preprint, abs/2211.08412.
  79. 79.Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu, Semih Yahvuz, Wojciech Kryscinski, Justin F. Rousseau, and Greg Durrett. 2022a. Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors.
  80. 80.Xiangru Tang, Alexander Fabbri, Haoran Li, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, and Dragomir Radev. 2022b. Investigating crowdsourcing protocols for evaluating the factual consistency of summaries. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5680–5692, Seattle, United States. Association for Computational Linguistics.
  81. 81.Robert J Tibshirani and Bradley Efron. 1993. An introduction to the bootstrap. Monographs on statistics and applied probability, 57:1–436.
  82. 82.Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020. Fill in the BLANC: Human-free quality estimation of document summaries. In Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems, pages 11–20, Online. Association for Computational Linguistics.
  83. 83.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575. IEEE Computer Society.
  84. 84.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  85. 85.Johnny Wei and Robin Jia. 2021. The statistical advantage of automatic NLG metrics at the system level. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6840–6854, Online. Association for Computational Linguistics.
  86. 86.Chien-Sheng Wu, Linqing Liu, Wenhao Liu, Pontus Stenetorp, and Caiming Xiong. 2021. Controllable abstractive dialogue summarization with sketch supervision. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 5108–5122, Online. Association for Computational Linguistics.
  87. 87.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems, volume 34, pages 27263–27277. Curran Associates, Inc.
  88. 88.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020a. PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 11328–11339. PMLR.
  89. 89.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020b. PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 11328–11339. PMLR.
  90. 90.Shiyue Zhang and Mohit Bansal. 2021. Finding a balanced degree of automation for summary evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6617–6632, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  91. 91.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020c. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  92. 92.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020d. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278, Online. Association for Computational Linguistics.
  93. 93.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578, Hong Kong, China. Association for Computational Linguistics.
  94. 94.Ming Zhong, Pengfei Liu, Yiran Chen, Danqing Wang, Xipeng Qiu, and Xuanjing Huang. 2020. Extractive summarization as text matching. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6197–6208, Online. Association for Computational Linguistics.
  95. 95.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation.

Citation

MLA
Liu, Y., et al. “Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 4140–70, https://doi.org/10.18653/v1/2023.acl-long.228.
APA
Liu, Y., Fabbri, A., Liu, P., Zhao, Y., Nan, L., Han, R., Han, S., Joty, S., Wu, C.-S., Xiong, C., & Radev, D. (2023). Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4140–4170. https://doi.org/10.18653/v1/2023.acl-long.228
Chicago
Liu, Y., A. Fabbri, P. Liu, et al. 2023. “Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4140–70. https://doi.org/10.18653/v1/2023.acl-long.228.
Harvard
Liu, Y. et al. (2023) “Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4140–4170. Available at: https://doi.org/10.18653/v1/2023.acl-long.228.
Vancouver
1. Liu Y, Fabbri A, Liu P, et al (2023) Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4140–4170

BibTeX

@inproceedings{liu-etal-2023-revisiting,
    title = "Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation",
    author = "Liu, Yixin  and
      Fabbri, Alex  and
      Liu, Pengfei  and
      Zhao, Yilun  and
      Nan, Linyong  and
      Han, Ruilin  and
      Han, Simeng  and
      Joty, Shafiq  and
      Wu, Chien-Sheng  and
      Xiong, Caiming  and
      Radev, Dragomir",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.228/",
    doi = "10.18653/v1/2023.acl-long.228",
    pages = "4140--4170"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/