Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Daniel DeutschRotem DrorDan Roth

article2022NAACL52 citations

Reveals critical flaws in standard summarization metric evaluations by demonstrating that common metrics like ROUGE show near-zero correlation with human judgments when discriminating between similarly performing systems.

Listen

Automatic evaluation metrics are widely used in text summarization research to compare system quality, guide development, and declare state-of-the-art models. Because human evaluation is expensive and slow, the field relies heavily on system-level correlation metrics—specifically Kendall's tau rank correlation—to estimate how reliably automatic scores replicate human judgments. However, the article identifies a critical disconnect between the standard methodology used to evaluate these metrics and how metrics are applied in practice, raising serious questions about the validity of standard benchmarking conclusions.

The article aims to align metric evaluation with practical use by introducing and evaluating two key methodological modifications: calculating automatic metric scores across all available test data rather than small human-annotated subsets, and measuring metric correlations specifically on pairs of systems with small score margins that reflect realistic competitive improvements.

To evaluate these changes, the article analyzed two benchmark datasets (SummEval and REALSumm) based on the CNN/DailyMail corpus, assessing leading reference-based metrics including the ROUGE family, BERTScore, and QAEval. The approach combined a literature survey of recent conference papers to identify realistic performance margins, statistical variance modeling, and bootstrapping resampling techniques to estimate confidence intervals and ranking stability.

The article establishes several critical findings. First, evaluating automatic metrics on the entire test set (around 10,000 instances) instead of only judged subsets (100 instances) reduced score variance by approximately 99% and narrowed confidence interval widths for system-level correlations by 16% to 51%. Second, a survey of recent literature revealed that proposed models improve over baselines by an average of only 0.49 ROUGE-1 points. Third, when automatic metrics are evaluated strictly on system pairs separated by these realistic, small differences (0.0 to 0.5 ROUGE points), correlation with human judgments plummets to near zero (0.08 on SummEval and 0.00 on REALSumm). Advanced metrics like BERTScore and QAEval also experienced steep correlation drops in this realistic regime. Strong standard correlations observed historically were largely inflated by comparing systems with wide, obvious quality gaps.

These findings imply that small reported gains in automatic metrics do not reliably indicate superior output according to human judgment. While continuous gains over time may correlate with real progress, individual benchmark victories based on narrow margins are effectively indistinguishable from random chance. Relying solely on narrow metric gains creates substantial operational risk of deploying inferior systems or misallocating development resources.

The article recommends that researchers and developers stop relying exclusively on standard metric improvements and invest more heavily in targeted human evaluations when comparing closely matched models. Additionally, metric developers should evaluate and report correlations across distinct score margins rather than relying solely on global rankings. When resources permit, data collection should focus on pairwise human judgments between similarly performing systems rather than broad, diverse assessments.

These conclusions are subject to certain limitations, including reliance on the CNN/DailyMail dataset domain and a constrained total number of evaluated systems (16 on SummEval and 25 on REALSumm). Furthermore, human judgment annotations themselves exhibit notable variance, indicating that while these estimates represent the best available empirical evidence, larger and more consistent human evaluation datasets are required to establish high-confidence metric benchmarks.

Cover for Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Abstract

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent with how metrics are used to evaluate systems in practice and propose changes to rectify this disconnect. First, we calculate the system score for an automatic metric using the full test set instead of the subset of summaries judged by humans, which is currently standard practice. We demonstrate how this small change leads to more precise estimates of system-level correlations. Second, we propose to calculate correlations only on pairs of systems that are separated by small differences in automatic scores which are commonly observed in practice. This allows us to demonstrate that our best estimate of the correlation of ROUGE to human judgments is near 0 in realistic scenarios. The results from the analyses point to the need to collect more high-quality human judgments and to improve automatic metrics when differences in system scores are small.¹

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Evaluating with All Available Instances
  • 3.1 Reducing Automatic Metric Variance
  • 3.2 Confidence Interval Analysis
  • 3.3 Conclusions & Recommendations
  • 4 Evaluating with Realistic System Pairs
  • 4.1 Evaluating with All System Pairs
  • 4.2 Evaluating with Realistic Pairs
  • 4.3 Conclusions & Recommendations
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgments
  • References
  • A Additional Confidence Interval Results
  • B Additional r SYS ∆( ℓ, u ) Results
  • C Summarization Paper Survey

Knowls

  1. Knowl 1 — System-level correlation using all test instances

    definition

    For NN summarization systems, let xijx_i^j be the score assigned by an automatic metric to system ii's summary for test document jj, and let zijz_i^j be the human score for that summary when it has been judged. Let TT be the full test set of MtestM_{test} documents and JJ the human-judged subset of MjudM_{jud} documents. The paper's revised system-level correlation computes each system's automatic score over TT, its human score over JJ, and correlates the resulting system-level scores using Kendall's c-b:

    rSYS=Kendall-τb({(1Mtest∑j∈Txij,  1Mjud∑j∈Jzij)}i=1N).r_{SYS} = \mathrm{Kendall}\text{-}\tau_b\left(\left\{\left(\frac{1}{M_{test}}\sum_{j\in T}x_i^j,\;\frac{1}{M_{jud}}\sum_{j\in J}z_i^j\right)\right\}_{i=1}^{N}\right).

    This differs from the conventional calculation, which averages automatic scores on the same judged subset JJ. The revised definition evaluates the automatic system scores researchers use in practice, while retaining human judgments on the available judged examples.

  2. Knowl 2 — Correlation restricted to systems with similar automatic scores

    definition

    The paper defines a score-gap-conditioned variant of system-level Kendall correlation, denoted rSYSΔ(ℓ,u)r_{SYS\Delta}(\ell,u). For an automatic metric XX with system scores siXs_i^X, include a system pair (i,k)(i,k) only when its absolute score difference satisfies ℓ≤∣siX−skX∣≤u\ell \leq |s_i^X-s_k^X| \leq u. Compute the association between the metric's and humans' rankings using only the included pairs; ℓ\ell and uu are lower and upper bounds in the metric's score units. This variant measures how well a metric distinguishes systems at specified score distances, instead of letting widely separated, comparatively easy-to-rank systems dominate the evaluation.

  3. Knowl 3 — Datasets and metrics used in the analyses

    experimental setup

    The analyses use the CNN/DailyMail test set, which contains 11,49011{,}490 documents, and two human-annotated evaluation datasets. SummEval contains judgments for 1616 systems on 100100 input documents, with summary relevance scores; REALSumm contains judgments for 2525 systems on 100100 input documents, using Lightweight Pyramid scores. The automatic metrics evaluated are ROUGE (including ROUGE-1, ROUGE-2, and ROUGE-L), BERTScore, and QAEval-F1. Automatic scores can be computed on the full test set even though human scores are available only for the judged subset.

  4. Knowl 4 — Scoring on the full test set reduces score variance and stabilizes rankings

    empirical result

    The authors used bootstrap resampling to compare estimates based on the judged inputs with estimates based on all available test inputs. For each of the three automatic metric families on both SummEval and REALSumm, using all test inputs reduced the variance of estimated system scores by about 99%99\%. They also compared system rankings from two independently sampled sets of MM inputs using Kendall's τ\tau. Automatic-metric rankings were typically around 0.60.6–0.80.8 in agreement at the judged-set size, but approached 11 as MM increased to the test-set size, indicating much more stable rankings. Human-judgment rankings remained variable: their agreement reached only about 0.80.8–0.850.85 at the maximum judged-input count.

  5. Knowl 5 — Using all test inputs narrows correlation confidence intervals

    empirical result

    For 95% confidence intervals on system-level correlations, the authors applied the BOOT-INPUTS bootstrap method, resampling input documents while holding the systems fixed. Replacing judged-subset automatic scores with full-test-set automatic scores reduced interval widths by an average of 51%51\% on SummEval and 16%16\% on REALSumm. The largest narrowing occurred for ROUGE metrics on SummEval. In additional analyses, interval widths decreased by 14%14\% and 12%12\% for BOOT-BOTH on SummEval and REALSumm, respectively, and by 1%1\% and 6%6\% for BOOT-SYSTEMS. The results show that the full-test-set definition can yield more precise correlation estimates, particularly when the bootstrap reflects variation in input documents.

  6. Knowl 6 — Metrics correlate poorly on close system pairs

    empirical result

    On pairs of systems with small automatic-score differences, correlations with human rankings were substantially lower than under the conventional all-pairs evaluation. For ROUGE-1 pairs separated by less than 0.50.5 points, the filtered correlation was 0.080.08 on SummEval and 0.00.0 on REALSumm, compared with standard ROUGE system-level correlations of 0.450.45 and 0.730.73, respectively. For BERTScore on SummEval, the correlation on the closest 20%20\% of pairs (a score gap of approximately 0.20.2) was 0.420.42, versus 0.770.77 using all pairs. The paper reports the same general pattern for QAEval and other ROUGE variants: correlations are low, and can be negative, on close-score pairs, then rise as larger score gaps—and thus easier-to-rank pairs—are included. The authors characterize these as their best estimates given the available judgments and system pairs.

  7. Knowl 7 — Reported ROUGE improvements motivate a realistic score-gap range

    empirical result

    To estimate the score differences encountered in practice, the authors surveyed recent 2020–2021 *ACL papers that proposed summarization systems, compared systems on CNN/DailyMail, and reported test-set ROUGE. For each paper they recorded the difference between the two highest-performing comparable systems, excluding ablations. The mean reported improvement was 0.490.49 ROUGE-1 points. This survey motivates evaluating metrics on narrow score-gap ranges such as the 00–0.50.5 ROUGE-1 interval, rather than relying solely on correlations over all system pairs.

  8. Knowl 8 — Limited judgments constrain conclusions about close-pair reliability

    limitation

    The close-pair analyses contain relatively few system pairs, especially among the most similar systems, and the available human judgments may not be sufficient to reliably distinguish systems of similar quality. Consequently, the low close-pair correlations are the authors' best estimates from the available data, not definitive measurements of metric performance in that setting. They recommend collecting more high-quality, consistent human judgments, with targeted pairwise comparisons between similarly performing systems as a possible design, and improving annotator training or annotation protocols to reduce judgment variability.

  9. Knowl 9 — Report correlations at multiple score gaps

    model/method

    The paper recommends that evaluations of new summarization metrics report standard system-level correlations alongside score-gap-conditioned correlations for multiple values of ℓ\ell and uu. Showing performance across score-gap ranges makes clear whether a metric's apparent agreement with human judgments comes mainly from correctly ordering systems with large score differences or also extends to the small differences common in system development. The authors argue that this information better indicates how much confidence users should place in an observed metric improvement.

Coverage note — The appendix's exhaustive score-gap heatmaps and the individual rows of the summarization-paper survey are omitted; the main findings are captured by the reported close-pair patterns and the survey's $0.49$ ROUGE-1 average.

References

  1. 1.Balach, Vidhisha ran, Artidoro Pagnoni, Jay Yoon Lee, Dheeraj Rajagopal, Jaime Carbonell, and Yulia Tsvetkov. 2021. StructSum: Summarization via Structured Representations. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2575–2585, Online. Association for Computational Linguistics.
  2. 2.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating Evaluation in Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online. Association for Computational Linguistics.
  3. 3.Jiaao Chen and Diyi Yang. 2021. Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1380–1391, Online. Association for Computational Linguistics.
  4. 4.John M. Conroy and Hoa Trang Dang. 2008. Mind the Gap: Dangers of Divorcing Evaluations of Summary Content from Linguistic Quality. In Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008), pages 145–152, Manchester, UK. Coling 2008 Organizing Committee.
  5. 5.Hoa Trang Dang and Karolina Owczarzak. 2008. Overview of the TAC 2008 Update Summarization Task. In Proc. of the Text Analysis Conference (TAC).
  6. 6.Shrey Desai, Jiacheng Xu, and Greg Durrett. 2020. Compressive Summarization with Plausibility and Salience Modeling. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6259–6274, Online. Association for Computational Linguistics.
  7. 7.Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021a. Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary. Transactions of the Association for Computational Linguistics, 9.
  8. 8.Daniel Deutsch, Rotem Dror, and Dan Roth. 2021b. A Statistical Analysis of Summarization Evaluation Metrics Using Resampling Methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
  9. 9.Daniel Deutsch and Dan Roth. 2020. Understanding the Extent to which Summarization Evaluation Metrics Measure the Information Quality of Summaries. ArXiv, abs/2010.12495.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021. GSum: A General Framework for Guided Neural Abstractive Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4830–4842, Online. Association for Computational Linguistics.
  12. 12.Alexander Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, and Dragomir Radev. 2021. SummEval: Re-evaluating Summarization Evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  13. 13.Luyang Huang, Lingfei Wu, and Lu Wang. 2020. Knowledge Graph-Augmented Abstractive Summarization with Semantic-Driven Cloze Reward. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5094–5107, Online. Association for Computational Linguistics.
  14. 14.Yin Jou Huang and Sadao Kurohashi. 2021. Extractive Summarization Considering Discourse and Coreference Relations based on Heterogeneous Graph. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3046–3052, Online. Association for Computational Linguistics.
  15. 15.Ruipeng Jia, Yanan Cao, Fang Fang, Yuchen Zhou, Zheng Fang, Yanbing Liu, and Shi Wang. 2021. Deep Differential Amplifier for Extractive Summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 366–376, Online. Association for Computational Linguistics.
  16. 16.Ruipeng Jia, Yanan Cao, Hengzhu Tang, Fang Fang, Cong Cao, and Shi Wang. 2020. Neural Extractive Summarization with Hierarchical Attentive Heterogeneous Graph Network. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3622–3631, Online. Association for Computational Linguistics.
  17. 17.Hanqi Jin and Xiaojun Wan. 2020. Abstractive Multi-Document Summarization via Joint Learning with Single-Document Summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2545–2554, Online. Association for Computational Linguistics.
  18. 18.Zhenwen Li, Wenhao Wu, and Sujian Li. 2020. Composing Elementary Discourse Units in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6191–6196, Online. Association for Computational Linguistics.
  19. 19.Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  20. 20.Yang Liu, Sheng Shen, and Mirella Lapata. 2021a. Noisy Self-Knowledge Distillation for Text Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 692–703, Online. Association for Computational Linguistics.
  21. 21.Yixin Liu, Zi-Yi Dou, and Pengfei Liu. 2021b. RefSum: Refactoring Neural Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1437–1448, Online. Association for Computational Linguistics.
  22. 22.Yixin Liu and Pengfei Liu. 2021. SimCLS: A Simple Framework for Contrastive Learning of Abstractive Summarization. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 1065–1072, Online. Association for Computational Linguistics.
  23. 23.Annie Louis and Ani Nenkova. 2013. Automatically Assessing Machine Summary Content Without a Gold Standard. Computational Linguistics, 39:267–300.
  24. 24.Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020. Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4984–4997, Online. Association for Computational Linguistics.
  25. 25.Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang. 2016. Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, CoNLL 2016, Berlin, Germany, August 11-12, 2016, pages 280–290. ACL.
  26. 26.Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, and Bing Xiang. 2021. Improving Factual Consistency of Abstractive Summarization via Question Answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6881–6894, Online. Association for Computational Linguistics.
  27. 27.Shashi Narayan, Joshua Maynez, Jakub Adamek, Daniele Pighin, Blaz Bratanic, and Ryan McDonald. 2020. Stepwise Extractive Summarization and Planning with Structured Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4143–4159, Online. Association for Computational Linguistics.
  28. 28.Vishakh Padmakumar and He He. 2021. Unsupervised Extractive Summarization using Pointwise Mutual Information. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2505–2512, Online. Association for Computational Linguistics.
  29. 29.Rebecca J Passonneau, Ani Nenkova, Kathleen McKeown, and Sergey Sigelman. 2005. Applying the Pyramid Method in DUC 2005. In Proceedings of the document understanding conference (DUC 05), Vancouver, BC, Canada.
  30. 30.Maxime Peyrard. 2019. Studying Summarization Evaluation Metrics in the Appropriate Scoring Range. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5093–5100, Florence, Italy. Association for Computational Linguistics.
  31. 31.Ori Shapira, David Gabay, Yang Gao, Hadar Ronen, Ramakanth Pasunuru, Mohit Bansal, Yael Amsterdamer, and Ido Dagan. 2019. Crowdsourcing Lightweight Pyramids for Manual Summary Evaluation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 682–687. Association for Computational Linguistics.
  32. 32.Zhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang, Long Tian, Bo Chen, and Mingyuan Zhou. 2020. Friendly Topic Assistant for Transformer Based Abstractive Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 485–497, Online. Association for Computational Linguistics.
  33. 33.Johnny Wei and Robin Jia. 2021. The Statistical Advantage of Automatic NLG Metrics at the System Level. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6840–6854, Online. Association for Computational Linguistics.
  34. 34.Liqiang Xiao, Lu Wang, Hao He, and Yaohui Jin. 2020. Modeling Content Importance for Summarization with Pre-trained Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3606–3611, Online. Association for Computational Linguistics.
  35. 35.Linzi Xing, Wen Xiao, and Giuseppe Carenini. 2021. Demoting the Lead Bias in News Summarization via Alternating Adversarial Learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 948–954, Online. Association for Computational Linguistics.
  36. 36.Jiacheng Xu, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020a. Discourse-Aware Neural Extractive Text Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5021–5031, Online. Association for Computational Linguistics.
  37. 37.Song Xu, Haoran Li, Peng Yuan, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020b. Self-Attention Guided Copy Mechanism for Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1355–1362, Online. Association for Computational Linguistics.
  38. 38.Ziyi Yang, Chenguang Zhu, Robert Gmyr, Michael Zeng, Xuedong Huang, and Eric Darve. 2020. TED: A Pretrained Unsupervised Summarization Model with Theme Modeling and Denoising. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1865–1874, Online. Association for Computational Linguistics.
  39. 39.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  40. 40.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. In Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  41. 41.Ming Zhong, Pengfei Liu, Yiran Chen, Danqing Wang, Xipeng Qiu, and Xuanjing Huang. 2020. Extractive Summarization as Text Matching. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6197–6208, Online. Association for Computational Linguistics.
  42. 42.Yanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei, and Ming Zhou. 2020. Pre-training for Abstractive Document Summarization by Reinstating Source Text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3646–3660, Online. Association for Computational Linguistics.

Citation

MLA
Deutsch, D., et al. “Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 6038–52, https://doi.org/10.18653/v1/2022.naacl-main.442.
APA
Deutsch, D., Dror, R., & Roth, D. (2022). Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 6038–6052. https://doi.org/10.18653/v1/2022.naacl-main.442
Chicago
Deutsch, D., R. Dror, and D. Roth. 2022. “Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 6038–52. https://doi.org/10.18653/v1/2022.naacl-main.442.
Harvard
Deutsch, D., Dror, R. and Roth, D. (2022) “Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 6038–6052. Available at: https://doi.org/10.18653/v1/2022.naacl-main.442.
Vancouver
1. Deutsch D, Dror R, Roth D (2022) Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 6038–6052

BibTeX

@inproceedings{deutsch-etal-2022-examining,
    title = "Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics",
    author = "Deutsch, Daniel  and
      Dror, Rotem  and
      Roth, Dan",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.442/",
    doi = "10.18653/v1/2022.naacl-main.442",
    pages = "6038--6052"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/