Built independently by an author, for readers. Read the story and support ChapterPal

keyword

machine translation evaluation

Machine translation evaluation is the process of measuring and assessing the quality, accuracy, and fluency of text generated by automated translation systems. This evaluation encompasses both human assessments, such as expert linguistic reviews, error classifications, and post-editing effort measurements, as well as automatic metrics designed to approximate human judgment. Automated evaluation methods range from traditional string-matching algorithms that measure lexical overlap against human-written reference translations to advanced neural and language-model-based metrics that analyze semantic similarity, context, and grammatical coherence with or without reference texts. The primary objective of machine translation evaluation is to benchmark system performance, diagnose linguistic errors, guide model development, and verify that the translated content reliably preserves the meaning and nuance of the original source text.

9 items

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang

OrganizationsMicrosoftPeking University

Why you should read this

Presents HD-Eval, an explainable and model-agnostic framework that aligns LLM evaluators with human judgment by iteratively decomposing evaluation tasks into fine-grained hierarchical criteria.

Large language models (LLMs) have emerged as a promising alternative to expensive human evaluations. However, the alignment and coverage of LLM-based evaluations are often limited by the scope and potential bias of the evaluation prompts and criteria. To address this challenge, we propose HD-Eval, a novel framework that iteratively aligns LLM-based evaluators with human preference via Hierarchical Criteria Decomposition. HD-Eval inherits the essence from the evaluation mindset of human experts and enhances the alignment of LLM-based evaluators by decomposing a given evaluation task into finer-grained criteria, aggregating them according to estimated human preferences, pruning insignificant criteria with attribution, and further decomposing significant criteria. By integrating these steps within an iterative alignment training process, we obtain a hierarchical decomposition of criteria that comprehensively captures aspects of natural language at multiple levels of granularity. Implemented as a white box, the human preference-guided aggregator is efficient to train and more explainable than relying solely on prompting, and its independence from model parameters makes it applicable to closed-source LLMs. Extensive experiments on three evaluation domains demonstrate the superiority of HD-Eval in further aligning state-of-the-art evaluators and providing deeper insights into the explanation of evaluation results and the task itself.

Added

2026-10-04

Document-Level Machine Translation with Large Language Models

Document-Level Machine Translation with Large Language Models

Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, Zhaopeng Tu

OrganizationsDublin City UniversityMohamed bin Zayed University of Artificial IntelligenceTencent

Why you should read this

Demonstrates that large language models outperform commercial translation systems in document-level translation quality and establishes an instruction-based benchmark to evaluate discourse-level phenomena such as entity consistency and pronoun resolution.

Large language models (LLMs) such as ChatGPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks. Taking document-level machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling. The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of ChatGPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling. By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation;2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5. This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs.2

Added

2026-09-29

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang

OrganizationsFudan University

Why you should read this

Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.

WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.

Added

2026-09-26

On the Limitations of Reference-Free Evaluations of Generated Text

On the Limitations of Reference-Free Evaluations of Generated Text

Daniel Deutsch, Rotem Dror, Dan Roth

OrganizationsGoogleUniversity of Pennsylvania

Why you should read this

Demonstrates that reference-free text evaluation metrics act as generation models themselves, exposing critical flaws where metrics favor models similar to their own architecture and penalize superior human-written outputs.

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications. However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization. We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans. Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.¹

Added

2026-09-26