Towards a Unified Multi-Dimensional Evaluator for Text Generation

Ming ZhongYang LiuDa YinYuning MaoYizhu JiaoPengfei LiuChenguang ZhuHeng JiJiawei Han

article2022EMNLP352 citations

Proposes UniEval, a unified text generation evaluator that reframes multi-dimensional assessment as Boolean question answering, markedly improving correlation with human judgments across summarization and dialogue while enabling zero-shot generalization to unseen criteria.

Listen

As automated language generation systems become more advanced, assessing their outputs has become a major challenge. The industry still largely relies on basic similarity-based metrics, which measure superficial word overlap against a reference text. These conventional metrics often fail to detect deeper issues such as factual inaccuracies or lack of coherence, leading to potentially misleading quality assessments. While human evaluation relies on multi-dimensional scoring—judging specific criteria like fluency, factual consistency, and relevance—automating this process historically required deploying numerous separate models or relying on single scores that lack clear interpretation.

The article aims to introduce and validate UNIEVAL, a unified framework that evaluates multiple distinct quality dimensions of generated text using a single model. The primary objective is to demonstrate that UNIEVAL achieves significantly higher correlation with human judgment across diverse text generation tasks compared to existing automated metrics.

To accomplish this, the authors restructured multi-dimensional evaluation as a simple "Yes/No" question-answering task based on the T5 language model. By feeding the model specific prompts—such as asking whether a summary is factual or coherent—the single system generates distinct scores for different criteria. Because large-scale human scoring datasets are scarce, the model incorporates an intermediate multi-task learning phase using over 185,000 examples from related tasks (such as natural language inference, linguistic acceptability, and general question answering) followed by unsupervised sequential training on synthetic pseudo data across summarization and dialogue generation.

The experimental findings show substantial improvements over existing evaluation methods across benchmarks such as SummEval and Topical-Chat. UNIEVAL outperforms the previous best unified evaluators, boosting correlation with human judgments by 23% in text summarization and by more than 43% in dialogue response generation. Single-dimensional evaluators built on this architecture also outperformed specialized consistency checkers by over 30% on the challenging QAGS benchmark. Furthermore, the model exhibited strong transferability, successfully evaluating previously unseen dimensions (such as understandability) and entirely new tasks (such as data-to-text generation) without requiring task-specific retraining.

These results demonstrate that organizations can replace fragmented, error-prone evaluation pipelines with a single unified evaluator. Adopting this approach reduces the cost and complexity of maintaining multiple separate assessment tools while providing explainable, multi-dimensional scores aligned with human judgment. This offers stakeholders a more reliable safeguard against factual errors and low-quality outputs before deploying language generation models into production.

Stakeholders and engineering teams should consider adopting unified question-answering frameworks for multi-dimensional text evaluation and leverage intermediate training on diverse auxiliary tasks to improve performance. However, decision-makers should exercise caution: the current model remains an uninterpretable neural network, relies on noisy synthetic training data, is currently evaluated only in English, and was tested using a single model size. Future work should focus on expanding language coverage, exploring lighter and larger model variants, and improving the interpretability of evaluation decisions.

Cover for Towards a Unified Multi-Dimensional Evaluator for Text Generation

Abstract

Multi-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency. However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models. In this paper, we propose a unified multi-dimensional evaluator UniEval for NLG. We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions. Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UniEval to incorporate external knowledge from multiple related tasks and gain further improvement. Experiments on three typical NLG tasks show that UniEval correlates substantially better with human judgments than existing metrics. Specifically, compared to the top-performing unified evaluators, UniEval achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation. Also, UniEval demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks. Source code, data and all pre-trained evaluators are available on our GitHub repository (this https URL).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Problem Formulation
  • 3.2 Unsupervised Learning on Multiple Evaluation Dimensions
  • 3.3 Intermediate Multi-task Learning
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Baselines
  • 4.3 Benchmarks
  • 4.4 Results For Summarization
  • 4.5 Results For Dialogue Generation
  • 4.6 Transfer Experiments
  • 4.7 Ablation Study of Intermediate Tasks
  • 5 Conclusion
  • References
  • A Dimensions in Evaluation tasks
  • A.1 Explanation of Each Dimension
  • A.2 Pseudo Data Construction for Dialogue Response Generation
  • A.3 Examples for Evaluation Tasks
  • B Examples for Intermediate Tasks
  • C Implementation Details
  • D Results on QAGS

Knowls

  1. Knowl 1 — Boolean QA Formulation and Dimension Scoring in UniEval

    model/method

    UniEval evaluates a text generation model's output across nn dimensions d=(d1,ext...,dn)d = (d_1, ext{...}, d_n) by casting each evaluation dimension into a Boolean Question Answering (QA) problem. The input to the evaluator is (x,y,c,qi)(x, y, c, q_i), where xx denotes the candidate generated text, yy is the optional reference text (omitted for reference-free evaluation dimensions), cc is the task context (e.g., source document or dialogue history), and qiq_i is a natural language question specifying dimension did_i (such as "Is this a coherent summary to the document?").

    The evaluator is trained as a sequence-to-sequence model to output either "Yes" or "No". The dimension quality score si∈[0,1]s_i \in [0, 1] is computed as the normalized token probability:

    si=P("Yes"∣x,y,c,qi)P("Yes"∣x,y,c,qi)+P("No"∣x,y,c,qi)s_i = \frac{P(\text{"Yes"} \mid x, y, c, q_i)}{P(\text{"Yes"} \mid x, y, c, q_i) + P(\text{"No"} \mid x, y, c, q_i)}

    where P(w∣⋅)P(w \mid \cdot) is the model probability assigned to generating token ww.

    For dimensions where defects occur at the sentence level (specifically, fluency and consistency in summarization), the candidate text xx is segmented into mm individual sentences (x1,...,xm)(x_1, \text{...}, x_m). The evaluator computes a sentence-level score sijs_{ij} for each jj-th sentence:

    sij=P("Yes"∣xj,y,c,qi)P("Yes"∣xj,y,c,qi)+P("No"∣xj,y,c,qi)s_{ij} = \frac{P(\text{"Yes"} \mid x_j, y, c, q_i)}{P(\text{"Yes"} \mid x_j, y, c, q_i) + P(\text{"No"} \mid x_j, y, c, q_i)}

    The composite score is the arithmetic mean across all sentences:

    si=1m∑j=1msijs_i = \frac{1}{m} \sum_{j=1}^m s_{ij}

    For engagingness in dialogue generation, which measures the cumulative quantity of engaging content, the score is calculated as the unnormalized sentence-level sum si=∑j=1msijs_i = \sum_{j=1}^m s_{ij}, yielding a score range of [0,+∞)[0, +\infty).

  2. Knowl 2 — Intermediate Multi-Task Learning Strategy in UniEval

    model/method

    Prior to dimension-specific evaluation training, UniEval undergoes intermediate multi-task learning across four diverse tasks converted into the unified Boolean QA format (c,q)→{"Yes","No"}(c, q) \to \{\text{"Yes"}, \text{"No"}\}. This incorporates external knowledge without candidate generation outputs xx or references yy:

    1. Natural Language Inference (NLI) (85,801 total examples: 41,149 positive, 44,652 negative): Sourced from DocNLI, MRPC, and QQP. Prompted with questions like "Is this hypothesis entailed in the premise?" or "Is this sentence equivalent to the reference?", mapping entailment and paraphrase pairs to "Yes" and contradiction/neutral/non-paraphrase pairs to "No".

    2. Self-Supervised Opening Sentence Prediction (SST) (60,000 examples: 30,000 positive, 30,000 negative): Built on the CNN/DailyMail corpus. Evaluates whether a sentence is the valid opening sentence of a news document using the prompt "Is this sentence the coherent first sentence of the document?", training inter-sentence coherence and salient content identification.

    3. Linguistics-Related Task (9,594 examples: 6,744 positive, 2,850 negative): Sourced from the Corpus of Linguistic Acceptability (CoLA). Evaluates grammatical acceptability with the prompt "Is this a fluent and linguistically acceptable sentence?".

    4. Generic Question Answering (30,128 examples: 17,032 positive, 13,096 negative): Assembled from BoolQ, BoolQ-NP, BoolQ-CS, StrategyQA, and Yes/No questions in MultiRC. Exposes the evaluator to diverse natural language questions to strengthen question conditioning.

    The intermediate phase trains the backbone model with cross-entropy loss across all 185,523 examples for 2 epochs.

  3. Knowl 3 — Synthetic Pseudo Data Construction for Dimension-Specific Evaluator Training

    model/method

    UniEval trains unsupervised evaluators without human scores by generating 30,000 synthetic pseudo-samples (50% positive, 50% negative) per evaluation dimension from unlabeled training sets (CNN/DailyMail for summarization; Topical-Chat for dialogue generation):

    • Summarization Dimensions:

      • Coherence: Positive samples are ground-truth summaries. Negative samples replace one sentence of the summary with a sentence selected from a different summary retrieved via BM25.
      • Consistency: Positive samples are factual summaries. Negative samples apply adversarial perturbations (antonym substitution, numerical value modification, named entity replacement, and syntactic pruning) to sentences.
      • Fluency: Negative samples are created by selecting a span from the positive summary (with span length drawn from a Poisson distribution with λ=5\lambda = 5) and applying one of three transformations: repeating, deleting, or shuffling.
      • Relevance: Negative samples replace multiple sentences of the ground-truth summary at random with sentences retrieved from other documents via BM25.
    • Dialogue Generation Dimensions:

      • Naturalness: Negative samples apply span repeating, deletion, or shuffling with Poisson span length λ=3\lambda = 3.
      • Coherence: Negative samples replace the gold response with a randomly sampled gold response from a different dialogue history.
      • Engagingness: Negative samples are uninformative and dull responses generated by a small DialoGPT model prompted with a single sentence.
      • Groundedness: Positive samples are PEGASUS-generated paraphrases of a sentence from the ground-truth knowledge context; negative samples are sentences randomly drawn from unrelated knowledge contexts.
  4. Knowl 4 — Continual Learning with Experience Replay for Multi-Dimensional Alignment

    model/method

    To eliminate the negative transfer observed when multi-task training across evaluation dimensions simultaneously (such as coherence degradation in summarization or engagingness drops in dialogue), UniEval trains dimensions sequentially using continual learning with experience replay.

    Training proceeds through a curriculum that prioritizes lower-level linguistic/surface features before complex semantic understanding:

    • Summarization Curriculum: Coherence→Fluency→Consistency→Relevance\text{Coherence} \to \text{Fluency} \to \text{Consistency} \to \text{Relevance}
    • Dialogue Generation Curriculum: Coherence→Naturalness→Groundedness→Engagingness\text{Coherence} \to \text{Naturalness} \to \text{Groundedness} \to \text{Engagingness}

    When introducing a new dimension dkd_k, 20% of the pseudo-data from all previously learned dimensions d1,...,dk−1d_1, \text{...}, d_{k-1} is randomly sampled and replayed alongside the data for dkd_k. The model (initialized from google/t5-v1_1-large) is trained for 0.2 to 2.0 epochs per dimension depending on task difficulty, using a maximum learning rate of 5×10−55 \times 10^{-5} and batch size 36.

  5. Knowl 5 — Multi-Dimensional Evaluation Performance on SummEval Summarization Benchmark

    data/table

    The performance of UniEval on the SummEval text summarization meta-evaluation benchmark compared against similarity-based metrics, single-dimensional baselines (CTC), and unified metrics (BARTScore). The evaluation reports summary-level Spearman (ρ\rho) and Kendall-Tau (τ\tau) correlations against human ratings on four dimensions.

    Metrics Coherence Consistency Fluency Relevance Average
    ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau ρ\rho τ\tau
    Similarity-based Metrics
    ROUGE-1 0.167 0.126 0.160 0.130 0.115 0.094 0.326 0.252 0.192 0.150
    ROUGE-2 0.184 0.139 0.187 0.155 0.159 0.128 0.290 0.219 0.205 0.161
    ROUGE-L 0.128 0.099 0.115 0.092 0.105 0.084 0.311 0.237 0.165 0.128
    BERTScore 0.284 0.211 0.110 0.090 0.193 0.158 0.312 0.243 0.225 0.175
    MoverScore 0.159 0.118 0.157 0.127 0.129 0.105 0.318 0.244 0.191 0.148
    Single-dimensional Evaluators
    CTC (Consistency) 0.223 0.172 0.415 0.345 0.335 0.276 0.166 0.124 0.285 0.229
    CTC (Relevance) 0.402 0.310 0.366 0.301 0.299 0.245 0.428 0.336 0.374 0.298
    UniEval (Coherence) 0.546 0.422 0.337 0.280 0.324 0.266 0.418 0.316 0.406 0.321
    UniEval (Consistency) 0.176 0.127 0.472 0.393 0.366 0.300 0.176 0.128 0.298 0.237
    UniEval (Fluency) 0.324 0.247 0.276 0.229 0.433 0.360 0.236 0.176 0.317 0.253
    UniEval (Relevance) 0.543 0.420 0.324 0.267 0.340 0.283 0.463 0.355 0.417 0.332
    Unified Evaluators
    BARTScore 0.448 0.342 0.382 0.315 0.356 0.292 0.356 0.273 0.385 0.305
    UniEval (Multi-task) 0.495 0.374 0.435 0.365 0.419 0.346 0.424 0.327 0.443 0.353
    UniEval (Continual) 0.575 0.442 0.446 0.371 0.449 0.371 0.426 0.325 0.474 0.377
    - Intermediate Tasks 0.477 0.363 0.403 0.333 0.414 0.342 0.395 0.301 0.422 0.335

    UniEval (Continual) achieves an average Spearman correlation of 0.474 and Kendall-Tau of 0.377, outperforming BARTScore (0.385 / 0.305) by 23% in average correlation. Continual training prevents the negative transfer on coherence seen in multi-task training (0.575 vs 0.495).

  6. Knowl 6 — Multi-Dimensional Evaluation Performance on Topical-Chat Dialogue Benchmark

    data/table

    Turn-level Pearson (rr) and Spearman (ρ\rho) correlations on the Topical-Chat meta-evaluation benchmark across Naturalness, Coherence, Engagingness, and Groundedness dimensions:

    Metrics Naturalness Coherence Engagingness Groundedness Average
    rr ρ\rho rr ρ\rho rr ρ\rho rr ρ\rho rr ρ\rho
    Similarity-based Metrics
    BLEU-1 0.161 0.133 0.210 0.223 0.314 0.334 0.289 0.303 0.243 0.248
    BLEU-4 0.180 0.175 0.131 0.235 0.232 0.316 0.213 0.310 0.189 0.259
    ROUGE-L 0.176 0.146 0.193 0.203 0.295 0.300 0.310 0.327 0.243 0.244
    METEOR 0.212 0.191 0.250 0.302 0.367 0.439 0.333 0.391 0.290 0.331
    BERTScore 0.226 0.209 0.214 0.233 0.317 0.335 0.291 0.317 0.262 0.273
    Single-dimensional Evaluators
    CTC (Engagingness) 0.280 0.257 0.352 0.325 0.516 0.525 0.405 0.404 0.388 0.378
    CTC (Groundedness) 0.200 0.161 0.256 0.228 0.485 0.475 0.524 0.477 0.366 0.335
    UniEval (Naturalness) 0.500 0.547 0.331 0.458 0.393 0.528 0.178 0.266 0.351 0.450
    UniEval (Coherence) 0.401 0.468 0.543 0.607 0.401 0.474 0.225 0.235 0.392 0.446
    UniEval (Engagingness) 0.394 0.427 0.471 0.477 0.562 0.596 0.376 0.431 0.451 0.483
    UniEval (Groundedness) 0.220 0.153 0.187 0.117 0.392 0.318 0.543 0.511 0.336 0.275
    Unified Evaluators
    USR 0.337 0.325 0.416 0.377 0.456 0.465 0.222 0.447 0.358 0.403
    UniEval (Multi-task) 0.480 0.512 0.518 0.609 0.544 0.563 0.462 0.456 0.501 0.535
    UniEval (Continual) 0.444 0.514 0.595 0.613 0.557 0.605 0.536 0.575 0.533 0.577
    - Intermediate Tasks 0.442 0.478 0.532 0.579 0.537 0.555 0.410 0.440 0.480 0.513

    UniEval (Continual) improves average correlation over the state-of-the-art dialogue evaluator USR by 48.9% in Pearson correlation (0.358→0.5330.358 \to 0.533) and 43.2% in Spearman correlation (0.403→0.5770.403 \to 0.577).

  7. Knowl 7 — Zero-Shot Transferability to Unseen Evaluation Dimensions and Tasks

    empirical result

    UniEval transfers zero-shot to unseen evaluation dimensions and new NLG tasks solely by modifying the natural language question prompt:

    1. Unseen Dimension on Topical-Chat: Using the prompt "Is this an understandable response in the dialogue?", UniEval evaluates the unseen dimension understandability in a zero-shot setting, attaining turn-level correlations of r=0.380r = 0.380 (Pearson) and ρ=0.468\rho = 0.468 (Spearman), exceeding the supervised USR metric (r=0.326,ρ=0.327r = 0.326, \rho = 0.327) and standard similarity metrics (e.g., BLEU-1 r=0.18,ρ=0.17r=0.18, \rho=0.17; METEOR r=0.23,ρ=0.25r=0.23, \rho=0.25; BERTScore r=0.24,ρ=0.24r=0.24, \rho=0.24).

    2. Unseen Task on SFRES and SFHOT (Data-to-Text): Using prompts "Is this a fluent utterance?" for Naturalness and "Is this sentence informative according to the reference?" for Informativeness, UniEval models transfer zero-shot without task-specific tuning. Spearman correlation results:

    Metrics SFRES SFHOT Average
    Nat. Info. Nat. Info.
    ROUGE-1 0.170 0.115 0.196 0.118 0.150
    ROUGE-L 0.169 0.103 0.186 0.110 0.142
    BERTScore 0.219 0.156 0.178 0.135 0.172
    MoverScore 0.190 0.153 0.242 0.172 0.189
    BARTScore 0.289 0.238 0.288 0.235 0.263
    T5 + Intermediate 0.348 0.180 0.310 0.181 0.255
    UniEval (Summ) 0.333 0.225 0.320 0.249 0.282

    UniEval trained on summarization (UniEval Summ) attains an average correlation of 0.282, outperforming BARTScore (0.263).

  8. Knowl 8 — Factuality Consistency Evaluation on QAGS Summarization Benchmark

    data/table

    Correlation performance of the single-dimensional UniEval (Consistency) evaluator against dedicated factuality checkers on the QAGS benchmark, which contains extractive (QAGS-CNN) and highly abstractive (QAGS-XSum) subsets:

    Metrics QAGS-CNN QAGS-XSUM Average
    rr ρ\rho τ\tau rr ρ\rho τ\tau rr ρ\rho τ\tau
    ROUGE-1 0.338 0.318 0.248 -0.008 -0.049 -0.040 0.165 0.134 0.104
    ROUGE-2 0.459 0.418 0.333 0.097 0.083 0.068 0.278 0.250 0.200
    ROUGE-L 0.357 0.324 0.254 0.024 -0.011 -0.009 0.190 0.156 0.122
    BERTScore 0.576 0.505 0.399 0.024 0.008 0.006 0.300 0.256 0.202
    MoverScore 0.414 0.347 0.271 0.054 0.044 0.036 0.234 0.195 0.153
    FactCC 0.416 0.484 0.376 0.297 0.259 0.212 0.356 0.371 0.294
    QAGS 0.545 - - 0.175 - - 0.375 - -
    BARTScore 0.735 0.680 0.557 0.184 0.159 0.130 0.459 0.420 0.343
    CTC (Consistency) 0.619 0.564 0.450 0.309 0.295 0.242 0.464 0.430 0.346
    UniEval (Consistency) 0.682 0.662 0.532 0.461 0.488 0.399 0.571 0.575 0.465

    UniEval (Consistency) outperforms CTC by over 30% in average Spearman (ρ=0.575\rho = 0.575 vs 0.4300.430) and Kendall-Tau (τ=0.465\tau = 0.465 vs 0.3460.346) correlations. While BARTScore performs strongly on extractive CNN data (r=0.735r = 0.735), its correlation drops on abstractive XSum (r=0.184r = 0.184), whereas UniEval maintains robust consistency detection across both extractive and abstractive domains (r=0.461r = 0.461 on XSum).

  9. Knowl 9 — Component Ablation of Intermediate Multi-Task Learning

    data/table

    Ablation study assessing the contribution of each intermediate task type to UniEval's evaluation performance, measured via summary-level Spearman correlation (ρ\rho) on SummEval:

    Evaluator Coherence Consistency Fluency Relevance Average
    UniEval 0.546 0.472 0.433 0.463 0.479
    - NLI 0.532 0.417 0.436 0.452 0.459
    - SST 0.498 0.462 0.428 0.450 0.460
    - Linguistics 0.548 0.466 0.415 0.458 0.472
    - Generic QA 0.528 0.438 0.421 0.436 0.456
    • Removing NLI produces the sharpest decline in Consistency (0.472→0.4170.472 \to 0.417).
    • Removing the Self-Supervised Opening Sentence Task (SST) produces the sharpest decline in Coherence (0.546→0.4980.546 \to 0.498).
    • Removing Linguistics (CoLA) causes a drop in Fluency (0.433→0.4150.433 \to 0.415).
    • Removing Generic QA causes a broad performance drop across all four evaluation dimensions.
  10. Knowl 10 — Stated Limitations of UniEval

    limitation

    The authors identify four main limitations of the UniEval evaluation framework:

    1. Black-Box Opacity: UniEval is built on pre-trained neural sequence-to-sequence models, providing no explicit interpretability into why or how the probability scores are assigned across dimensions.
    2. Synthetic Data Noise: Constructing pseudo data via heuristic rules introduces label noise (for example, deleting non-essential tokens to create disfluent samples may preserve grammatical fluency yet be labeled as negative samples).
    3. Model Scale and Resource Constraints: Experiments are evaluated strictly on the T5-large backbone (google/t5-v1_1-large) due to computational constraints, without studying compact deployment models or larger language models.
    4. Monolingual English Scope: Evaluators and intermediate datasets are exclusively constructed in English, leaving multilingual and cross-lingual text generation evaluation unexamined.

Coverage note — None of the substantial contributed material was omitted.

References

  1. 1.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  2. 2.Anja Belz and Albert Gatt. 2008. Intrinsic vs. extrinsic evaluation measures for referring expression generation. In Proceedings of ACL-08: HLT, Short Papers, pages 197–200.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020. Factual error correction for abstractive summarization models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6251–6258.
  5. 5.Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. 2018. Faithful to the original: Fact aware neural abstractive summarization. In thirty-second AAAI conference on artificial intelligence.
  6. 6.Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021. Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2082–2095.
  7. 7.Yulong Chen, Yang Liu, Li Dong, Shuohang Wang, Chenguang Zhu, Michael Zeng, and Yue Zhang. 2022. Adaprompt: Adaptive model training for prompt-based nlp. arXiv preprint arXiv:2202.04824.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019a. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936.
  9. 9.Elizabeth Clark, Asli Celikyilmaz, and Noah A Smith. 2019b. Sentence mover’s similarity: Automatic evaluation for multi-sentence texts. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2748–2760.
  10. 10.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  12. 12.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019.
  13. 13.Bill Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Third International Workshop on Paraphrasing (IWP2005).
  14. 14.Esin Durmus, He He, and Mona Diab. 2020. Feqa: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070.
  15. 15.Nouha Dziri, Ehsan Kamalloo, Kory Mathewson, and Osmar R Zaiane. 2019. Evaluating coherence in dialogue systems using entailment. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3806–3812.
  16. 16.Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  17. 17.Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al. 2020. Evaluating models’ local decision boundaries via contrast sets. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1307–1323.
  18. 18.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. arXiv preprint arXiv:2202.06935.
  19. 19.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  20. 20.Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019. Topical-chat: Towards knowledge-grounded open-domain conversations. Proc. Interspeech 2019, pages 1891–1895.
  21. 21.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. Advances in neural information processing systems, 28.
  22. 22.Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020. Grade: Automatic graphenhanced coherence metric for evaluating opendomain dialogue systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9230–9240.
  23. 23.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander R Fabbri, Yejin Choi, and Noah A Smith. 2021. Bidimensional leaderboards: Generate and evaluate language hand in hand. arXiv preprint arXiv:2112.04139.
  24. 24.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 252–262.
  25. 25.Daniel Khashabi, Tushar Khot, and Ashish Sabharwal. 2020. More bang for your buck: Natural perturbation for robust question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 163–170.
  26. 26.Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346.
  27. 27.Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. 2021. Summac: Re-visiting nlibased models for inconsistency detection in summarization. arXiv preprint arXiv:2111.09525.
  28. 28.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  29. 29.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  30. 30.Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaichen Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, and Graham Neubig. 2021a. Explainaboard: An explainable leaderboard for nlp. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 280–289.
  31. 31.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021b. Pretrain, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  32. 32.Shikib Mehri and Maxine Eskenazi. 2020. Usr: An unsupervised and reference free evaluation metric for dialog generation. arXiv preprint arXiv:2005.00456.
  33. 33.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807.
  34. 34.Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  35. 35.GI Parisi, R Kemker, JL Part, C Kanan, and S Wermter. 2019. Continual lifelong learning with neural networks: A review. Neural Networks: the Official Journal of the International Neural Network Society, 113:54–71.
  36. 36.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  37. 37.Ehud Reiter and Anja Belz. 2009. An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics, 35(4):529–558.
  38. 38.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: Bm25 and beyond. Information Retrieval, 3(4):333–389.
  39. 39.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. Questeval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604.
  40. 40.Abigail See, Peter J Liu, and Christopher D Manning. 2017. Get to the point: Summarization with pointergenerator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083.
  41. 41.Amanda Stent, Matthew Marge, and Mohit Singhai. 2005. Evaluating evaluation methods for generation in the presence of variation. In Proceedings of the 6th international conference on Computational Linguistics and Intelligent Text Processing, pages 341–351.
  42. 42.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020.
  43. 43.Zhiguo Wang, Wael Hamza, and Radu Florian. 2017. Bilateral multi-perspective matching for natural language sentences. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, pages 4144–4150.
  44. 44.Alex Warstadt, Amanpreet Singh, and Samuel Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  45. 45.Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015. Semantically conditioned lstm-based natural language generation for spoken dialogue systems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1711–1721.
  46. 46.Zheng Ye, Liucun Lu, Lishan Huang, Liang Lin, and Xiaodan Liang. 2021. Towards quantifiable dialogue coherence evaluation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2718–2729.
  47. 47.Wenpeng Yin, Dragomir Radev, and Caiming Xiong. 2021. Docnli: A large-scale dataset for documentlevel natural language inference. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4913–4922.
  48. 48.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34.
  49. 49.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  50. 50.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020. Dialogpt: Largescale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278.
  51. 51.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019. Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 563–578.
  52. 52.Ming Zhong, Pengfei Liu, Danqing Wang, Xipeng Qiu, and Xuan-Jing Huang. 2019. Searching for effective neural extractive summarization: What works and what’s next. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1049–1058.

Citation

MLA
Zhong, M., et al. “Towards a Unified Multi-Dimensional Evaluator for Text Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 2023–38, https://doi.org/10.18653/v1/2022.emnlp-main.131.
APA
Zhong, M., Liu, Y., Yin, D., Mao, Y., Jiao, Y., Liu, P., Zhu, C., Ji, H., & Han, J. (2022). Towards a Unified Multi-Dimensional Evaluator for Text Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2023–2038. https://doi.org/10.18653/v1/2022.emnlp-main.131
Chicago
Zhong, M., Y. Liu, D. Yin, et al. 2022. “Towards a Unified Multi-Dimensional Evaluator for Text Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2023–38. https://doi.org/10.18653/v1/2022.emnlp-main.131.
Harvard
Zhong, M. et al. (2022) “Towards a Unified Multi-Dimensional Evaluator for Text Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2023–2038. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.131.
Vancouver
1. Zhong M, Liu Y, Yin D, Mao Y, Jiao Y, Liu P, Zhu C, Ji H, Han J (2022) Towards a Unified Multi-Dimensional Evaluator for Text Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2023–2038

BibTeX

@inproceedings{zhong-etal-2022-towards,
    title = "Towards a Unified Multi-Dimensional Evaluator for Text Generation",
    author = "Zhong, Ming  and
      Liu, Yang  and
      Yin, Da  and
      Mao, Yuning  and
      Jiao, Yizhu  and
      Liu, Pengfei  and
      Zhu, Chenguang  and
      Ji, Heng  and
      Han, Jiawei",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.131/",
    doi = "10.18653/v1/2022.emnlp-main.131",
    pages = "2023--2038"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/