SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation
Elizabeth ClarkShruti RijhwaniSebastian GehrmannJoshua MaynezRoee AharoniVitaly NikolaevThibault SellamAditya SiddhantDipanjan DasAnkur P. Parikh
Introduces a massive dataset of 96,000 human-annotated summaries spanning six languages and six quality dimensions, providing an essential resource for training and benchmarking learned summarization metrics that generalize well to out-of-domain evaluation tasks.
As modern language models rapidly advance, they frequently generate text containing subtle inaccuracies and ungrounded statements. Evaluating the quality and faithfulness of automatically generated summaries is critical, yet human evaluation remains expensive, slow, and hard to scale, while traditional automated metrics correlate poorly with human judgment. This challenge is especially severe for non-English languages, where large-scale human evaluation datasets are almost nonexistent.
The main objective of the article is to present and evaluate SEAHORSE, a multilingual, multifaceted dataset designed to train and benchmark automated neural evaluation metrics for text summarization. By providing high-quality human ratings across multiple quality facets and avoiding test-split contamination, the article demonstrates how SEAHORSE enables learned metrics to reliably assess summary quality without requiring reference summaries.
To construct this resource, the authors collected 96,645 human evaluations covering six diverse languages (German, English, Spanish, Russian, Turkish, and Vietnamese) across four established summarization corpora. The dataset includes summaries generated by nine distinct systems—ranging from small and under-trained neural networks to large language models and human reference texts. Professional annotators evaluated each summary independently along six discrete quality dimensions: comprehensibility, repetition, grammar, attribution (factual grounding), main ideas, and conciseness. Using the training portions of this data, the authors fine-tuned automated evaluation metrics based on multilingual text-to-text models and meta-evaluated them against internal test sets and external benchmarks.
The article establishes several key findings. First, while most evaluated systems reliably produce comprehensible, grammatical, and non-repetitive text, higher-level qualities such as factual attribution, capturing main ideas, and conciseness remain difficult, with positive response rates often dropping below 60–70% across systems. Second, metrics trained on SEAHORSE substantially outperform traditional lexical metrics like ROUGE-L and baseline natural language inference systems on the in-domain test set. Third, SEAHORSE-trained metrics transfer robustly to out-of-domain evaluation benchmarks; when applied to the external 45-language mFACE benchmark, the metric demonstrated strong zero-shot generalization across 40 unseen languages, matching the attribution evaluation performance of models directly trained on that benchmark. Finally, on the English TRUE factual consistency benchmark, SEAHORSE metrics achieved competitive or superior accuracy across multiple summarization and dialogue datasets.
These findings have major practical implications for organizations developing and deploying automated language systems. Reliable, reference-free automated evaluation reduces the operational costs, delays, and subjectivity associated with large human evaluation pipelines. Furthermore, the ability to accurately verify attribution across multiple languages lowers the business, compliance, and reputational risks associated with deploying ungrounded generative models in multilingual consumer-facing applications.
Organizations should adopt reference-free, learned multidimensional metrics to validate summarization and generative language systems during development and inference-time re-ranking. Future initiatives should expand evaluation resources into lower-resource languages and conduct systematic hyperparameter optimization and architectural exploration for evaluation models.
Confidence in these findings is high, supported by extensive human agreement checks, large sample sizes, and consistent out-of-domain transfer results. Nevertheless, decision-makers should note certain limitations: human annotations contain baseline subjectivity and noise, the dataset directly covers only six high-resource languages, and the trained metrics were developed as demonstration proofs rather than fully optimized production systems.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). This paper establishes the standardized TRUE benchmark for evaluating factual consistency in generated text, serving as a direct baseline and external testbed against which SEAHORSE metrics are validated.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). This work provides foundational analyses of hallucinations and factual attribution errors in neural summarization, defining core challenges that SEAHORSE addresses through multidimensional evaluation.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). BLEURT establishes the paradigm of training learned, neural evaluation metrics from pre-trained language models to overcome the limitations of surface-level overlap metrics.
- Paper: COMET: A Neural Framework for MT Evaluation, Ricardo Rei et al. (2020). COMET formalizes cross-lingual and multilingual neural metric learning for text generation, laying technical foundations for training evaluation models across multiple languages.
- Paper: Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics, Chin-Yew Lin et al. (2003). This seminal paper introduces n-gram co-occurrence metrics for summarization (ROUGE), providing the classical evaluation standard that SEAHORSE-trained neural metrics aim to improve upon.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). XNLI establishes cross-lingual natural language inference benchmarks and methodologies, providing the underlying NLI framework that SEAHORSE builds upon for multilingual attribution modeling.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R provides the core multilingual representation learning backbone used to scale cross-lingual understanding and zero-shot metric transfer across diverse languages.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). This work introduces the XSum extreme summarization benchmark, which is one of the key source corpora evaluated and annotated within the SEAHORSE dataset.
- Paper: SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization, Mathieu Ravaut et al. (2022). SummaReranker demonstrates how multi-task and learned metrics can be directly applied to inference-time summary candidate re-ranking, motivating SEAHORSE's proposed downstream applications.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This comprehensive survey categorizes systemic failures and best practices across natural language generation evaluation, broadening the methodological critique that motivated SEAHORSE.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). FActScore extends the evaluation of factual precision from summary-level judgments to fine-grained, atomic fact verification in long-form language model generation.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). G-Eval builds upon reference-free multidimensional evaluation by using large language models and chain-of-thought prompting to evaluate summaries without dedicated fine-tuning.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE applies automated, reference-free attribution evaluation principles to assess citation quality and factual support in retrieved evidence generations.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). ARES extends automated multi-criteria evaluation to retrieval-augmented generation systems, using statistical prediction-powered inference to validate factual grounding.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey analyzes the broader transition toward using language models as automated judges across evaluation tasks, detailing the biases and optimization methods relevant to learned evaluators.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This text synthesizes advanced LLM-as-a-judge techniques, exploring how multi-attribute assessment models can replace or augment trained metric classifiers.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). LLMBAR establishes meta-evaluation methods to stress-test automated language model evaluators against adversarial outputs, probing the reliability of learned judges.
