SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation

Elizabeth ClarkShruti RijhwaniSebastian GehrmannJoshua MaynezRoee AharoniVitaly NikolaevThibault SellamAditya SiddhantDipanjan DasAnkur P. Parikh

article2023EMNLP58 citations

Introduces a massive dataset of 96,000 human-annotated summaries spanning six languages and six quality dimensions, providing an essential resource for training and benchmarking learned summarization metrics that generalize well to out-of-domain evaluation tasks.

Listen

As modern language models rapidly advance, they frequently generate text containing subtle inaccuracies and ungrounded statements. Evaluating the quality and faithfulness of automatically generated summaries is critical, yet human evaluation remains expensive, slow, and hard to scale, while traditional automated metrics correlate poorly with human judgment. This challenge is especially severe for non-English languages, where large-scale human evaluation datasets are almost nonexistent.

The main objective of the article is to present and evaluate SEAHORSE, a multilingual, multifaceted dataset designed to train and benchmark automated neural evaluation metrics for text summarization. By providing high-quality human ratings across multiple quality facets and avoiding test-split contamination, the article demonstrates how SEAHORSE enables learned metrics to reliably assess summary quality without requiring reference summaries.

To construct this resource, the authors collected 96,645 human evaluations covering six diverse languages (German, English, Spanish, Russian, Turkish, and Vietnamese) across four established summarization corpora. The dataset includes summaries generated by nine distinct systems—ranging from small and under-trained neural networks to large language models and human reference texts. Professional annotators evaluated each summary independently along six discrete quality dimensions: comprehensibility, repetition, grammar, attribution (factual grounding), main ideas, and conciseness. Using the training portions of this data, the authors fine-tuned automated evaluation metrics based on multilingual text-to-text models and meta-evaluated them against internal test sets and external benchmarks.

The article establishes several key findings. First, while most evaluated systems reliably produce comprehensible, grammatical, and non-repetitive text, higher-level qualities such as factual attribution, capturing main ideas, and conciseness remain difficult, with positive response rates often dropping below 60–70% across systems. Second, metrics trained on SEAHORSE substantially outperform traditional lexical metrics like ROUGE-L and baseline natural language inference systems on the in-domain test set. Third, SEAHORSE-trained metrics transfer robustly to out-of-domain evaluation benchmarks; when applied to the external 45-language mFACE benchmark, the metric demonstrated strong zero-shot generalization across 40 unseen languages, matching the attribution evaluation performance of models directly trained on that benchmark. Finally, on the English TRUE factual consistency benchmark, SEAHORSE metrics achieved competitive or superior accuracy across multiple summarization and dialogue datasets.

These findings have major practical implications for organizations developing and deploying automated language systems. Reliable, reference-free automated evaluation reduces the operational costs, delays, and subjectivity associated with large human evaluation pipelines. Furthermore, the ability to accurately verify attribution across multiple languages lowers the business, compliance, and reputational risks associated with deploying ungrounded generative models in multilingual consumer-facing applications.

Organizations should adopt reference-free, learned multidimensional metrics to validate summarization and generative language systems during development and inference-time re-ranking. Future initiatives should expand evaluation resources into lower-resource languages and conduct systematic hyperparameter optimization and architectural exploration for evaluation models.

Confidence in these findings is high, supported by extensive human agreement checks, large sample sizes, and consistent out-of-domain transfer results. Nevertheless, decision-makers should note certain limitations: human annotations contain baseline subjectivity and noise, the dataset directly covers only six high-resource languages, and the trained metrics were developed as demonstration proofs rather than fully optimized production systems.

arXiv: 2305.13194
Cover for SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation

Abstract

Reliable automatic evaluation of summarization systems is challenging due to the multifaceted and subjective nature of the task. This is especially the case for languages other than English, where human evaluations are scarce. In this work, we introduce SEAHORSE, a dataset for multilingual, multifaceted summarization evaluation. SEAHORSE consists of 96K summaries with human ratings along 6 dimensions of text quality: comprehensibility, repetition, grammar, attribution, main ideas, and conciseness. SEAHORSE covers 6 languages, 9 systems (including the reference text), and 4 summarization datasets. As a result of its size and scope, SEAHORSE can serve both as a benchmark to evaluate learnt metrics, as well as a large-scale resource for training such metrics. We show that metrics trained with SEAHORSE achieve strong performance on two out-of-domain meta-evaluation benchmarks: TRUE (Honovich et al., 2022) and mFACE (Aharoni et al., 2023). We make the SEAHORSE dataset and metrics publicly available for future research on multilingual and multifaceted summarization evaluation.1

Table of Contents

  • 1 Introduction
  • 2 The SEAHORSE dataset
  • 2.1 The summaries
  • 2.2 Annotation methodology
  • 3 Dataset analysis
  • 4 Learning and evaluating metrics with SEAHORSE
  • 4.1 Metrics
  • 4.2 Evaluation on the SEAHORSE test set
  • 4.3 Evaluation on the mFACE dataset
  • 4.4 Evaluation on the TRUE Benchmark
  • 5 Related work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Training details
  • B Rate of positive responses
  • C SEAHORSE example summaries and scores
  • D Comparison between mT5_large and mT5_xxl

Knowls

  1. Knowl 1 — SEAHORSE dataset scope and split design

    definition

    SEAHORSE is a multilingual, multifaceted dataset containing 96,645 summaries annotated along six quality dimensions: Q1 comprehensibility, Q2 repetition, Q3 grammar, Q4 attribution to the source article, Q5 coverage of the source article’s main ideas, and Q6 conciseness. It covers German, English, Spanish, Russian, Turkish, and Vietnamese; four summarization corpora; and nine summary sources, including eight model configurations and the human-authored reference summary.

    The dataset is reference-free: annotators judge each generated summary against its source article rather than comparing it with the corpus reference. Articles for SEAHORSE training and validation are sampled from the original corpora’s validation splits, while SEAHORSE test articles come from the original corpora’s test splits. This separation allows metrics trained on SEAHORSE to be used for model development without training on the articles used for final testing.

  2. Knowl 2 — Multilingual source composition

    data/table

    The 96,645 annotated summaries are distributed across the source corpora as follows. Each entry gives the number of unique source articles followed by the number of annotated summaries.

    • German: MLSUM, 3,359 articles and 7,506 summaries; WikiLingua, 2,999 articles and 7,085 summaries.
    • English: XSum, 894 articles and 6,651 summaries; XL-Sum, 2,433 articles and 7,884 summaries; WikiLingua, 2,383 articles and 7,804 summaries.
    • Spanish: XL-Sum, 2,231 articles and 4,890 summaries; MLSUM, 2,235 articles and 4,857 summaries; WikiLingua, 2,183 articles and 5,002 summaries.
    • Russian: XL-Sum, 3,298 articles and 7,254 summaries; WikiLingua, 2,948 articles and 7,288 summaries.
    • Turkish: XL-Sum, 2,186 articles and 10,627 summaries; WikiLingua, 770 articles and 4,791 summaries.
    • Vietnamese: XL-Sum, 2,497 articles and 7,522 summaries; WikiLingua, 1,951 articles and 7,484 summaries.

    The underlying tasks span one-sentence BBC News summarization in XSum and XL-Sum, online-newspaper summarization in MLSUM, and how-to-guide summarization in WikiLingua. The combination of languages, domains, and article collections is intended to provide diverse training and evaluation data for learned summarization metrics.

  3. Knowl 3 — Diverse summary-generation systems

    model/method

    Each source article is paired with summaries from a subset of nine systems: the human-authored reference; T5-base with 220 million parameters; an under-trained T5-base checkpoint after 250 training steps; T5-XXL with 11 billion parameters; mT5-small with 300 million parameters; an under-trained mT5-small checkpoint after 250 steps; mT5-XXL with 13 billion parameters; PaLM with 540 billion parameters prompted using one in-domain example; and PaLM with 540 billion parameters fine-tuned on the relevant dataset.

    The English-only T5 systems are used only for English articles. The systems were selected to span expected quality levels and qualitatively different error types, including under-training, scaling, prompting, and fine-tuning. Summarization models were trained on the training split of each source corpus, except for the one-shot PaLM system, which received one example and the input article. Checkpoints were selected using the corresponding corpus validation set, except for the deliberately 250-step systems. T5 and mT5 used input length 1,024 and target length 512; PaLM used input length 2,048 and target length 512.

  4. Knowl 4 — Staged human annotation protocol

    model/method

    For every summary, trained annotators answer six binary-quality questions, with an additional Unsure option. Q1 asks whether the summary can be read and understood; if the answer is No, the remaining questions are skipped. Q2 asks whether the summary avoids unnecessary repetition. Q3 asks whether it is grammatically correct. After seeing the source article, Q4 asks whether all information in the summary is attributable to that article, Q5 asks whether the summary captures the article’s main ideas, and Q6 asks whether it concisely represents the article’s information.

    The interface reveals information in stages: the summary and Q1 appear first, Q2–Q3 appear after the annotator confirms comprehensibility, and the source article with Q4–Q6 appears last. Annotators are paid full-time workers trained for the task; non-English annotators are bilingual in the annotation language and English. Training uses 109 summaries with gold ratings, with each annotator first rating 20–30 gold examples and receiving additional feedback and 10–20-example retraining rounds when necessary. Ratings from annotators removed during quality monitoring are excluded.

  5. Knowl 5 — Observed summary-quality patterns across systems

    empirical result

    SEAHORSE ratings distinguish the expected quality differences among summarization systems. Positive ratings are generally high for comprehensibility, repetition, and grammar, but lower for attribution, main-idea coverage, and conciseness, showing that the latter dimensions remain more difficult for summarization systems.

    The strongest low-quality behavior occurs for under-trained and small models. Across the dataset, mT5-small after 250 steps receives positive ratings of 0.71, 0.43, 0.59, 0.27, 0.19, and 0.10 on Q1–Q6, respectively; fully trained mT5-small receives 0.86, 0.57, 0.73, 0.36, 0.35, and 0.19; and mT5-XXL receives 0.96, 0.94, 0.88, 0.55, 0.65, and 0.43. For the PaLM systems, one-shot prompting yields 0.88, 0.85, 0.79, 0.71, 0.57, and 0.47, whereas fine-tuning yields 0.98, 0.98, 0.90, 0.69, 0.71, and 0.56.

    The one-shot PaLM summaries are unusually likely to copy the source: their median length is 254 characters, their median ROUGE-L against the article is 27.34, and 14% exactly copy the first 20% of the article. This copying helps attribution but does not necessarily indicate good summarization. In general, larger models outperform smaller models, fully trained models outperform 250-step models, and fine-tuned PaLM outperforms one-shot PaLM; the principal exception is attribution, where copying can make one-shot PaLM appear more faithful than fine-tuned PaLM.

  6. Knowl 6 — Annotation consistency and dimension relationships

    empirical result

    The annotation patterns support the intended interpretation of the six quality dimensions. In more than 99% of cases where an annotator marks a summary as not attributable on Q4 or as missing the article’s main ideas on Q5, the same annotator also marks it as not concise on Q6. Conciseness therefore has its strongest relationships with attribution and main-idea coverage.

    A subset of 8,920 summaries received two additional ratings. Overall pairwise agreement is 82%. The average pairwise agreement by language for Q1–Q6 is: German 0.97, 0.98, 0.95, 0.81, 0.67, 0.66; Spanish 0.92, 0.97, 0.83, 0.74, 0.70, 0.74; English 0.97, 0.94, 0.95, 0.69, 0.61, 0.69; Russian 0.86, 0.97, 0.88, 0.71, 0.73, 0.76; Turkish 0.93, 0.96, 0.86, 0.74, 0.70, 0.74; and Vietnamese 0.95, 0.98, 0.88, 0.68, 0.66, 0.69. Averaged over languages, agreement is 0.93, 0.97, 0.89, 0.73, 0.68, and 0.72 for Q1–Q6.

    Krippendorff’s alpha for Q1–Q6 is 0.49, 0.87, 0.35, 0.47, 0.40, and 0.41. Questions requiring only the summary have higher agreement than the more contextual and subjective source-grounded questions, while grammar has the lowest alpha despite high raw agreement.

  7. Knowl 7 — Reference-free learned metrics trained from SEAHORSE

    model/method

    The learned evaluation metrics are six separate mT5 text-to-text classifiers, one for each SEAHORSE quality dimension. Each metric receives the source article and candidate summary and predicts the human label as 0 or 1. The main model fine-tunes mT5-XXL with 13 billion parameters; a smaller comparison model fine-tunes mT5-Large with 1.2 billion parameters. The article and summary are represented with the tags premise: and hypothesis:, respectively, and the metric input length is 2,048 tokens.

    Training uses a filtered set of 88,280 SEAHORSE items after removing duplicates and ratings other than Yes or No. Training and development examples come from articles in the original corpora’s validation splits, and test examples come from the original corpora’s test splits. The best checkpoint is selected on the SEAHORSE validation split, and all training and inference use the T5X framework on TPU accelerators.

    The comparison systems are a majority-class predictor, ROUGE-L between the article and summary, and, for attribution, English t5NLI and multilingual mt5XNLI models trained on natural-language-inference data. Reference-based metrics such as BERTScore and BLEURT are not used because SEAHORSE evaluates summaries without requiring reference summaries. Performance is measured with Pearson correlation between metric scores and human labels and with area under the receiver-operating-characteristic curve.

  8. Knowl 8 — SEAHORSE test-set metric performance

    empirical result

    On the held-out SEAHORSE test set, the mT5-XXL metric trained on SEAHORSE is consistently stronger than the non-learned baselines. For Q1–Q6, its Pearson correlations are 0.52, 0.86, 0.45, 0.59, 0.50, and 0.52, while its ROC areas are 0.90, 0.98, 0.84, 0.85, 0.80, and 0.81.

    The mT5-Large SEAHORSE metric remains competitive, with correlations 0.44, 0.74, 0.37, 0.55, 0.46, and 0.45 and ROC areas 0.88, 0.97, 0.81, 0.82, 0.78, and 0.77 for Q1–Q6. ROUGE-L obtains correlations 0.04, 0.06, −0.03, 0.13, 0.03, and 0.02 and ROC areas 0.54, 0.54, 0.43, 0.55, 0.53, and 0.54. The attribution-only mt5XNLI baseline obtains correlation 0.43 and ROC area 0.78. The majority-class baseline has ROC area 0.50 for every dimension.

    The repetition metric is the strongest SEAHORSE metric, matching the dimension with the highest pairwise annotator agreement. Reducing the metric backbone from mT5-XXL to mT5-Large lowers performance but preserves the same overall superiority over the baselines.

  9. Knowl 9 — Zero-shot generalization to mFACE

    empirical result

    The SEAHORSE metrics generalize without additional fine-tuning to mFACE, which evaluates XL-Sum summaries in 45 languages on quality, attribution, and informativeness. The SEAHORSE Q1, Q4, and Q6 metrics are used as the closest matches for mFACE quality, attribution, and informativeness, respectively. Results are reported both for the five languages shared with SEAHORSE and for all mFACE languages.

    On the five-language subset, the mT5-XXL SEAHORSE metrics obtain Pearson correlations and ROC areas of 0.09/0.73 for quality, 0.50/0.79 for attribution, and 0.50/0.81 for informativeness. On all 45 languages, they obtain 0.15/0.70, 0.52/0.81, and 0.40/0.74, respectively.

    The directly supervised mT5-MFACE comparison model, trained on mFACE training data, obtains 0.25/0.68, 0.51/0.81, and 0.47/0.79 on the five-language subset and 0.35/0.61, 0.52/0.82, and 0.47/0.80 on all languages. Thus, the SEAHORSE metric is comparable to the mFACE-trained model on the five-language subset for all dimensions and on all-language attribution, despite being applied zero-shot to languages and evaluation dimensions not fully represented during SEAHORSE training. On all-language quality and informativeness, the mFACE-trained model performs better.

  10. Knowl 10 — Attribution generalization on TRUE

    empirical result

    The attribution metric trained only on SEAHORSE is evaluated without further fine-tuning on the English TRUE benchmark, covering summarization, dialogue, verification, and paraphrasing datasets. Its ROC areas are 0.94 on FRANK, 0.87 on SummEval, 0.83 on MNBM, 0.91 on QAGS-C, 0.87 on QAGS-X, 0.84 on BEGIN, 0.82 on Q2, 0.87 on DialFact, 0.91 on FEVER, 0.78 on VitaminC, and 0.82 on PAWS.

    The ROUGE-L baseline obtains 0.55, 0.57, 0.53, 0.44, 0.55, 0.63, 0.54, 0.49, 0.48, 0.50, and 0.60 on the same datasets. The English t5NLI baseline obtains 0.90, 0.79, 0.76, 0.77, 0.85, 0.85, 0.83, 0.92, 0.95, 0.98, and 0.99. The t5NLI results for FEVER, VitaminC, and PAWS are marked as trained-data comparisons because those datasets are included in its NLI training mixture.

    The SEAHORSE metric is best on every summarization dataset in TRUE and remains competitive with t5NLI on dialogue datasets, despite being trained only on summarization data. t5NLI is strongest on FEVER, VitaminC, and PAWS, where it has direct training-data exposure.

  11. Knowl 11 — Stated limitations and scope constraints

    limitation

    SEAHORSE covers six languages selected partly because suitable datasets and articles were available; it does not address many low-resource languages. The authors identify cross-lingual and few-shot summarization as possible ways to expand coverage.

    Human ratings remain subjective and noisy even after annotator training, particularly for source-dependent dimensions. The ratings measure the relationship between a summary and its source article rather than whether the article or summary is objectively true: a summary that repeats an error in the source can still receive a positive attribution rating. Finally, the mT5-based metrics are demonstrations of the dataset’s utility, not fully optimized evaluation systems; the study does not perform extensive hyperparameter search or systematic comparison across architectures and training approaches.

Coverage note — The detailed appendix-level positive-rate breakdowns, illustrative summary examples, and additional backbone-comparison tables were omitted because they support the main dataset and metric findings without adding separate load-bearing contributions.

References

  1. 1.Roee Aharoni, Shashi Narayan, Joshua Maynez, Jonathan Herzig, Elizabeth Clark, and Mirella Lapata. 2023. Multilingual summarization with factual consistency evaluation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 3562–3591, Toronto, Canada. Association for Computational Linguistics.
  2. 2.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating evaluation in text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9347–9359, Online. Association for Computational Linguistics.
  3. 3.Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojevic. 2016. ´ Results of the WMT16 metrics shared task. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 199–231, Berlin, Germany. Association for Computational Linguistics.
  4. 4.Ali Borji. 2023. A categorical archive of chatGPT failures. arXiv preprint arXiv:2302.03494.
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  8. 8.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The PASCAL recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment: First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers, pages 177–190. Springer.
  9. 9.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2022. Evaluating attribution in dialogue systems: The BEGIN benchmark. Transactions of the Association for Computational Linguistics, 10:1066–1083.
  10. 10.Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  11. 11.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  12. 12.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  13. 13.Yaroslav Fyodorov, Yoad Winter, and Nissim Francez. 2000. A natural logic inference system. In Proceedings of the 2nd Workshop on Inference in Computational Semantics (ICoS-2).
  14. 14.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  15. 15.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. arXiv preprint arXiv:2202.06935.
  16. 16.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023. TrueTeacher: Learning factual consistency evaluation with large language models. arXiv preprint arXiv:2305.11171.
  17. 17.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. DialFact: A benchmark for fact-checking in dialogue. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3785–3801, Dublin, Ireland. Association for Computational Linguistics.
  18. 18.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  19. 19.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. Advances in Neural Information Processing Systems, 28.
  20. 20.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating factual consistency evaluation. In Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Question Answering, pages 161–175, Dublin, Ireland. Association for Computational Linguistics.
  21. 21.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7856–7870, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  22. 22.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  23. 23.Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, and Yejin Choi. 2022. Prompt waywardness: The curious case of discretized interpretation of continuous prompts. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3631–3643, Seattle, United States. Association for Computational Linguistics.
  24. 24.Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018. ScitaiL: A textual entailment dataset from science question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  25. 25.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation. In Proceedings of the Sixth Conference on Machine Translation, pages 478–494, Online. Association for Computational Linguistics.
  26. 26.Klaus Krippendorff. 1980. Content analysis: An introduction to its methodology.
  27. 27.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  28. 28.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4034–4048, Online. Association for Computational Linguistics.
  29. 29.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  30. 30.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  31. 31.Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2022. Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation. arXiv preprint arXiv:2212.07981.
  32. 32.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  33. 33.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  34. 34.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  35. 35.Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4812–4829, Online. Association for Computational Linguistics.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  37. 37.Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. ToTTo: A controlled table-to-text generation dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1173–1186, Online. Association for Computational Linguistics.
  38. 38.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  39. 39.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  40. 40.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. arXiv preprint arXiv:2112.12870.
  41. 41.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  42. 42.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, et al. 2022. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189.
  43. 43.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin C! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 624–643, Online. Association for Computational Linguistics.
  44. 44.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2020. MLSUM: The multilingual summarization corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8051–8067, Online. Association for Computational Linguistics.
  45. 45.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  46. 46.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, volume 33, pages 3008–3021. Curran Associates, Inc.
  47. 47.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  48. 48.Ran Tian, Shashi Narayan, Thibault Sellam, and Ankur P Parikh. 2019. Sticking to the facts: Confident decoding for faithful data-to-text generation. arXiv preprint arXiv:1910.08684.
  49. 49.Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein. 2017. TL;DR: Mining Reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pages 59–63, Copenhagen, Denmark. Association for Computational Linguistics.
  50. 50.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  51. 51.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  52. 52.Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017. Challenges in data-to-document generation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2253–2263, Copenhagen, Denmark. Association for Computational Linguistics.
  53. 53.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  54. 54.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with bert. In International Conference on Learning Representations.
  55. 55.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.
  56. 56.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021. Detecting hallucinated content in conditional neural sequence generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1393–1404, Online. Association for Computational Linguistics.

Citation

MLA
Clark, E., et al. “SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 9397–413, https://doi.org/10.18653/v1/2023.emnlp-main.584.
APA
Clark, E., Rijhwani, S., Gehrmann, S., Maynez, J., Aharoni, R., Nikolaev, V., Sellam, T., Siddhant, A., Das, D., & Parikh, A. (2023). SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9397–9413. https://doi.org/10.18653/v1/2023.emnlp-main.584
Chicago
Clark, E., S. Rijhwani, S. Gehrmann, et al. 2023. “SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9397–9413. https://doi.org/10.18653/v1/2023.emnlp-main.584.
Harvard
Clark, E. et al. (2023) “SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9397–9413. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.584.
Vancouver
1. Clark E, Rijhwani S, Gehrmann S, Maynez J, Aharoni R, Nikolaev V, Sellam T, Siddhant A, Das D, Parikh A (2023) SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9397–9413

BibTeX

@inproceedings{clark-etal-2023-seahorse,
    title = "{SEAHORSE}: A Multilingual, Multifaceted Dataset for Summarization Evaluation",
    author = "Clark, Elizabeth  and
      Rijhwani, Shruti  and
      Gehrmann, Sebastian  and
      Maynez, Joshua  and
      Aharoni, Roee  and
      Nikolaev, Vitaly  and
      Sellam, Thibault  and
      Siddhant, Aditya  and
      Das, Dipanjan  and
      Parikh, Ankur",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.584/",
    doi = "10.18653/v1/2023.emnlp-main.584",
    pages = "9397--9413"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/