Built independently by an author, for readers. Read the story and support ChapterPal

keyword

language model-based metrics

Language model-based metrics are automated evaluation methods in natural language processing that utilize pre-trained language models to measure the quality, semantic similarity, and fluency of generated text. Unlike traditional evaluation techniques that rely strictly on surface-level lexical overlap between candidate texts and reference texts, these metrics leverage contextual embeddings, learned representations, or generative scoring mechanisms to capture deeper semantic meaning, grammatical coherence, and contextual nuance. They are commonly employed across generative tasks such as machine translation, text summarization, and dialogue systems by comparing contextual representations, predicting human quality scores through trained regression models, or prompting models directly to provide structured evaluation ratings.

1 item

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang

OrganizationsFudan University

Why you should read this

Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.

WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.

Added

2026-09-26