Towards a Unified Multi-Dimensional Evaluator for Text Generation
Ming ZhongYang LiuDa YinYuning MaoYizhu JiaoPengfei LiuChenguang ZhuHeng JiJiawei Han
Proposes UniEval, a unified text generation evaluator that reframes multi-dimensional assessment as Boolean question answering, markedly improving correlation with human judgments across summarization and dialogue while enabling zero-shot generalization to unseen criteria.
As automated language generation systems become more advanced, assessing their outputs has become a major challenge. The industry still largely relies on basic similarity-based metrics, which measure superficial word overlap against a reference text. These conventional metrics often fail to detect deeper issues such as factual inaccuracies or lack of coherence, leading to potentially misleading quality assessments. While human evaluation relies on multi-dimensional scoring—judging specific criteria like fluency, factual consistency, and relevance—automating this process historically required deploying numerous separate models or relying on single scores that lack clear interpretation.
The article aims to introduce and validate UNIEVAL, a unified framework that evaluates multiple distinct quality dimensions of generated text using a single model. The primary objective is to demonstrate that UNIEVAL achieves significantly higher correlation with human judgment across diverse text generation tasks compared to existing automated metrics.
To accomplish this, the authors restructured multi-dimensional evaluation as a simple "Yes/No" question-answering task based on the T5 language model. By feeding the model specific prompts—such as asking whether a summary is factual or coherent—the single system generates distinct scores for different criteria. Because large-scale human scoring datasets are scarce, the model incorporates an intermediate multi-task learning phase using over 185,000 examples from related tasks (such as natural language inference, linguistic acceptability, and general question answering) followed by unsupervised sequential training on synthetic pseudo data across summarization and dialogue generation.
The experimental findings show substantial improvements over existing evaluation methods across benchmarks such as SummEval and Topical-Chat. UNIEVAL outperforms the previous best unified evaluators, boosting correlation with human judgments by 23% in text summarization and by more than 43% in dialogue response generation. Single-dimensional evaluators built on this architecture also outperformed specialized consistency checkers by over 30% on the challenging QAGS benchmark. Furthermore, the model exhibited strong transferability, successfully evaluating previously unseen dimensions (such as understandability) and entirely new tasks (such as data-to-text generation) without requiring task-specific retraining.
These results demonstrate that organizations can replace fragmented, error-prone evaluation pipelines with a single unified evaluator. Adopting this approach reduces the cost and complexity of maintaining multiple separate assessment tools while providing explainable, multi-dimensional scores aligned with human judgment. This offers stakeholders a more reliable safeguard against factual errors and low-quality outputs before deploying language generation models into production.
Stakeholders and engineering teams should consider adopting unified question-answering frameworks for multi-dimensional text evaluation and leverage intermediate training on diverse auxiliary tasks to improve performance. However, decision-makers should exercise caution: the current model remains an uninterpretable neural network, relies on noisy synthetic training data, is currently evaluated only in English, and was tested using a single model size. Future work should focus on expanding language coverage, exploring lighter and larger model variants, and improving the interpretability of evaluation decisions.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). BLEURT establishes the foundation of using pre-trained transformer representations with multi-task intermediate pre-training for learned text generation evaluation, which UniEval directly builds upon.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). This paper highlights the critical failure of surface-level similarity metrics to correlate with human judgments in dialogue generation, motivating UniEval's multi-dimensional evaluation approach.
- Paper: Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics, Chin-Yew Lin et al. (2003). This seminal work introduces n-gram co-occurrence metrics for text summarization (ROUGE), providing the foundational baseline metrics that UniEval aims to surpass.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). TRUE establishes a unified framework and benchmark for evaluating factual consistency across natural language generation tasks, a core evaluation dimension unified in UniEval.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). G-Eval extends multi-dimensional NLG evaluation from task-specific trained models like UniEval to versatile, prompting-based large language model judges using chain-of-thought.
- Paper: GPTScore: Evaluate as You Desire, Jinlan Fu et al. (2024). GPTScore broadens customizable multi-aspect text evaluation by leveraging conditional generation probabilities of pre-trained language models without requiring explicit evaluator training.
- Paper: Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models, Seungone Kim et al. (2024). Prometheus 2 builds upon multi-criteria automated evaluation by training open-source evaluator language models that support both direct scoring rubrics and pairwise judgments.
- Paper: SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation, Elizabeth Clark et al. (2023). SEAHORSE provides a multilingual, multifaceted benchmark dataset designed specifically to train and evaluate neural metrics across granular quality dimensions.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This survey provides an extensive critique and systematization of systemic obstacles in text generation evaluation practices, contextualizing the paradigm shift toward learned evaluators.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This work explores scaling automated evaluation to open-ended conversational models using LLMs as judges, measuring their alignment with human preference benchmarks.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This comprehensive survey organizes the broader progression from specialized generation evaluators to the general-purpose LLM-as-a-judge paradigm.
