On the Limitations of Reference-Free Evaluations of Generated Text
Daniel DeutschRotem DrorDan Roth
Demonstrates that reference-free text evaluation metrics act as generation models themselves, exposing critical flaws where metrics favor models similar to their own architecture and penalize superior human-written outputs.
Automated evaluation of text generation systems, such as machine translation and document summarization, traditionally relies on comparing outputs against human-written reference texts. Because gathering human references is costly, slow, and often infeasible for real-time applications, researchers have increasingly developed reference-free evaluation metrics that score candidate text directly from the input source. However, adopting reference-free metrics as primary benchmarks introduces fundamental risks of mismeasuring model performance.
The article demonstrates that reference-free evaluation metrics are mathematically equivalent to text generation models and inherently limited by significant structural biases. Specifically, it assesses whether reference-free metrics can be safely used to measure genuine task performance and track overall research progress.
The authors conducted empirical evaluations across machine translation and summarization using established benchmarks, including WMT19 data spanning 18 language pairs as well as the SummEval and REALSumm datasets covering 16 to 25 summarization systems. They analyzed three prominent reference-free metrics—Prism-src, COMET-QE, and QuestEval—by formulating simple search algorithms, such as beam search, greedy sentence selection, and output reranking, to directly optimize each metric at inference time without human references.
The analysis produced three critical findings. First, simple optimization procedures reliably generated outputs that outperformed standard baseline models on reference-free metrics, such as a 38% relative improvement in COMET-QE on German-to-English translation. Second, despite their top reference-free scores, these optimized outputs exhibited average or below-average actual quality when judged by standard reference-based metrics and contained severe translation errors. Third, reference-free metrics routinely scored machine outputs higher than human-written text. In German-to-English translation under Prism-src, nearly all automated models received higher scores than authentic human translations, showing a clear bias against genuine human language. Furthermore, scoring systems against a metric's optimized output showed an exceptionally strong correlation (average Pearson r of 0.88 to 0.95) with the reference-free scores, confirming that reference-free metrics effectively treat their own generated text as a gold standard.
These findings imply that using reference-free metrics to guide system development creates a misleading feedback loop. Instead of learning to generate high-quality, human-like text, models optimize toward the specific quirks and systemic errors of the evaluation metric’s internal model. Relying on these metrics as primary performance benchmarks risks misdirecting research investments, adopting lower-quality language systems, and penalizing genuinely superior models that diverge from the evaluator model's patterns.
The authors recommend that organizations stop using reference-free metrics as primary benchmarks or objective functions for measuring task progress. When measuring overall generation quality, stakeholders should invest in collecting human-written references for evaluation. Reference-free metrics should instead be repurposed as diagnostic tools or quality estimation safeguards to flag potential catastrophic errors, verify source faithfulness, or measure text fluency.
While the theoretical equivalence between reference-free metrics and generation models applies universally, the empirical demonstrations in the article are limited to machine translation and summarization datasets. The authors also relied on established reference-based metrics rather than expensive, ground-truth human annotations to evaluate the optimized outputs. Nonetheless, the evidence strongly warns against treating reference-free scores as substitutes for reference-based or human evaluation.
- Paper: COMET: A Neural Framework for MT Evaluation, Ricardo Rei et al. (2020). Introduces neural evaluation architectures (COMET) for machine translation, illustrating the cross-lingual learned metrics that the source analyzes and challenges.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). Presents learned, transformer-based evaluation metrics (BLEURT) that serve as a prime foundation and target of investigation for reference-free evaluation limitations.
- Paper: CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Jack Hessel et al. (2021). Demonstrates the practical appeal and design of reference-free evaluation using pretrained representations (CLIPScore), establishing the metric paradigm critiqued in the source.
- Paper: Bleu: a Method for Automatic Evaluation of Machine Translation, Kishore Papineni et al. (2002). Establishes standard reference-based evaluation (BLEU), providing the baseline evaluation methodology that reference-free metrics sought to replace.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). Highlights the unreliability of automated metrics in dialogue generation, motivating the need for more nuanced evaluation techniques analyzed in the source.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Examines faithfulness and hallucination in abstractive summarization, highlighting the evaluation pitfalls and task settings discussed in the source.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Surveys systemic obstacles across natural language generation evaluation practices, broadening the source's findings on metric bias and limitations.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). Synthesizes the LLM-as-a-judge paradigm and categorizes judge-specific biases and failure modes that arise when using models as automated evaluators.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Investigates LLMs as reference-free judges for open-ended generation, measuring human alignment and practical biases like verbosity and position preference.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Develops an LLM-based evaluation framework (G-EVAL) for reference-free NLG evaluation, serving as a direct modern realization of model-based evaluators.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Benchmarks automated LLM evaluators on instruction-following adherence and exposes their vulnerability to superficial surface patterns.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Re-evaluates factual consistency metrics across multiple text generation domains using a unified meta-evaluation benchmark.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). Builds an extensive multi-task benchmark using evaluator language models to assess generation quality under fine-grained rubrics.
