Built independently by an author, for readers. Read the story and support ChapterPal

keyword

reference-free evaluations

Reference-free evaluations are assessment methods in natural language processing that measure the quality, coherence, relevance, or factual correctness of machine-generated text without comparing it against a human-written target or gold-standard reference. Unlike traditional reference-based metrics that rely on lexical overlap with pre-existing ground truth, reference-free approaches evaluate generated outputs directly using the input prompt, source context, or specified evaluation criteria. These methods frequently utilize pretrained neural quality estimation systems or large language models acting as automated judges to perform direct scoring, pairwise comparisons, or error analysis. By eliminating the necessity for manually curated reference texts, this paradigm enables scalable evaluation in open-ended generation tasks and online applications, though it can introduce model-dependent biases and reliability challenges.

2 items

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo

OrganizationsAllen Institute for AICarnegie Mellon UniversityKorea Advanced Institute of Science and TechnologyLG AI ResearchMassachusetts Institute of TechnologyUniversity of Illinois Chicago

Why you should read this

Presents Prometheus 2, an open-source evaluator language model that handles both direct assessment and pairwise ranking with custom criteria by merging models trained on separate evaluation formats, closely mirroring human and GPT-4 judgments.

Proprietary LMs such as GPT-4 are often employed to assess the quality of responses from various LMs. However, concerns including transparency, controllability, and affordability strongly motivate the development of open-source LMs specialized in evaluations. On the other hand, existing open evaluator LMs exhibit critical shortcomings: 1) they issue scores that significantly diverge from those assigned by humans, and 2) they lack the flexibility to perform both direct assessment and pairwise ranking, the two most prevalent forms of assessment. Additionally, they often do not possess the ability to evaluate based on *custom evaluation criteria*, focusing instead on general attributes like helpfulness and harmlessness. To address these issues, we introduce Prometheus 2. Prometheus 2 is more powerful than its predecessor, and closely mirrors human and GPT-4 judgements. Moreover, it is capable of processing both direct assessment and pair-wise ranking formats grouped with a user-defined evaluation criteria. On four direct assessment benchmarks and four pairwise ranking benchmarks, PROMETHEUS 2 scores the highest correlation and agreement with humans and proprietary LM judges among all tested open evaluator LMs. Our models, code, and data are all publicly available. 1

Added

2026-09-28

On the Limitations of Reference-Free Evaluations of Generated Text

On the Limitations of Reference-Free Evaluations of Generated Text

Daniel Deutsch, Rotem Dror, Dan Roth

OrganizationsGoogleUniversity of Pennsylvania

Why you should read this

Demonstrates that reference-free text evaluation metrics act as generation models themselves, exposing critical flaws where metrics favor models similar to their own architecture and penalize superior human-written outputs.

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications. However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization. We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans. Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible.¹

Added

2026-09-26