keyword
machine translation evaluation
Machine translation evaluation is the process of measuring and assessing the quality, accuracy, and fluency of text generated by automated translation systems. This evaluation encompasses both human assessments, such as expert linguistic reviews, error classifications, and post-editing effort measurements, as well as automatic metrics designed to approximate human judgment. Automated evaluation methods range from traditional string-matching algorithms that measure lexical overlap against human-written reference translations to advanced neural and language-model-based metrics that analyze semantic similarity, context, and grammatical coherence with or without reference texts. The primary objective of machine translation evaluation is to benchmark system performance, diagnose linguistic errors, guide model development, and verify that the translated content reliably preserves the meaning and nuance of the original source text.
9 items

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang
Why you should read this
Presents HD-Eval, an explainable and model-agnostic framework that aligns LLM evaluators with human judgment by iteratively decomposing evaluation tasks into fine-grained hierarchical criteria.
Large language models (LLMs) have emerged as a promising alternative to expensive human evaluations. However, the alignment and coverage of LLM-based evaluations are often limited by the scope and potential bias of the evaluation prompts and criteria. To address this challenge, we propose HD-Eval, a novel framework that iteratively aligns LLM-based evaluators with human preference via Hierarchical Criteria Decomposition. HD-Eval inherits the essence from the evaluation mindset of human experts and enhances the alignment of LLM-based evaluators by decomposing a given evaluation task into finer-grained criteria, aggregating them according to estimated human preferences, pruning insignificant criteria with attribution, and further decomposing significant criteria. By integrating these steps within an iterative alignment training process, we obtain a hierarchical decomposition of criteria that comprehensively captures aspects of natural language at multiple levels of granularity. Implemented as a white box, the human preference-guided aggregator is efficient to train and more explainable than relying solely on prompting, and its independence from model parameters makes it applicable to closed-source LLMs. Extensive experiments on three evaluation domains demonstrate the superiority of HD-Eval in further aligning state-of-the-art evaluators and providing deeper insights into the explanation of evaluation results and the task itself.
Added
2026-10-04

On the Blind Spots of Model-Based Evaluation Metrics for Text Generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov
Why you should read this
Exposes critical failure modes in popular pretrained language model-based evaluation metrics like BERTScore and MAUVE using synthetic stress tests, while providing practical workarounds to ensure more reliable text generation assessment.
In this work, we study the blind spots of model-based evaluation metrics for text generation. We first analyze the behaviors of model-based metrics under adversarial perturbations and find that they are vulnerable to adversarial attacks. We then show that the blind spots are caused by the fact that model-based metrics are trained on the same data distribution as the generation models. We further propose a simple method to mitigate the blind spots by training the metric on a different data distribution. Experiments on multiple text generation tasks demonstrate the effectiveness of our method.
Added
2026-10-03

Document-Level Machine Translation with Large Language Models
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, Zhaopeng Tu
Why you should read this
Demonstrates that large language models outperform commercial translation systems in document-level translation quality and establishes an instruction-based benchmark to evaluate discourse-level phenomena such as entity consistency and pronoun resolution.
Large language models (LLMs) such as ChatGPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks. Taking document-level machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling. The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of ChatGPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling. By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation;2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5. This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs.2
Added
2026-09-29

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang
Why you should read this
Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.
WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.
Added
2026-09-26

On the Limitations of Reference-Free Evaluations of Generated Text
Daniel Deutsch, Rotem Dror, Dan Roth
Why you should read this
Demonstrates that reference-free text evaluation metrics act as generation models themselves, exposing critical flaws where metrics favor models similar to their own architecture and penalize superior human-written outputs.
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications. However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization. We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans. Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.¹
Added
2026-09-26

COMET: A Neural Framework for MT Evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, Alon Lavie
Why you should read this
Introduces COMET, an adaptable neural framework that utilizes pretrained cross-lingual language models to evaluate machine translation quality against both source inputs and reference translations, achieving state-of-the-art correlation with human judgments.
We present COMET, a neural framework for training multilingual machine translation evaluation models which obtains new state-of-the-art levels of correlation with human judgements. Our framework leverages recent breakthroughs in cross-lingual pretrained language modeling resulting in highly multilingual and adaptable MT evaluation models that exploit information from both the source input and a target-language reference translation in order to more accurately predict MT quality. To showcase our framework, we train three models with different types of human judgements: Direct Assessments, Human-mediated Translation Edit Rate and Multidimensional Quality Metrics. Our models achieve new state-of-the-art performance on the WMT 2019 Metrics shared task and demonstrate robustness to high-performing systems.
Added
2026-09-24

Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn
Why you should read this
Proposes bootstrap resampling techniques to accurately test the statistical significance of machine translation evaluations using metrics like BLEU, enabling reliable system comparisons even on test sets as small as 300 sentences.
If two translation systems differ differ in performance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling methods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real.
Added
2026-09-18

BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, Ankur P. Parikh
Why you should read this
Proposes BLEURT, a learned text generation evaluation metric that uses synthetic pre-training on BERT to closely model human judgments across benchmarks even when evaluation data is scarce or out-of-distribution.
Text generation has made significant advances in the last few years. Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate poorly with human judgments. We propose BLEURT, a learned evaluation metric based on BERT that can model human judgments with a few thousand possibly biased training examples. A key aspect of our approach is a novel pre-training scheme that uses millions of synthetic examples to help the model generalize. BLEURT provides state-of-the-art results on the last three years of the WMT Metrics shared task and the WebNLG Competition dataset. In contrast to a vanilla BERT-based approach, it yields superior results even when the training data is scarce and out-of-distribution.
Added
2026-09-17

BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi
Why you should read this
Introduces BERTScore, an automatic evaluation metric for text generation that computes semantic similarity using contextual token embeddings instead of exact n-gram matching to correlate more faithfully with human judgments.
We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BERTScore is more robust to challenging examples when compared to existing metrics.
Added
2026-09-07
