Built independently by an author, for readers. Read the story and support ChapterPal

keyword

human judgments

Human judgments are subjective or qualitative assessments, ratings, and annotations provided by human evaluators to determine the quality, relevance, or correctness of generated content and system outputs. In artificial intelligence, machine learning, and computational linguistics, these evaluations serve as the benchmark or gold standard for assessing performance across tasks such as text generation, machine translation, image captioning, and information retrieval. By measuring nuanced criteria like fluency, factual accuracy, semantic coherence, and overall user satisfaction, human judgments provide an empirical ground truth that practitioners use to validate, calibrate, and improve automated evaluation metrics and predictive models.

11 items

CLAIR: Evaluating Image Captions with Large Language Models

CLAIR: Evaluating Image Captions with Large Language Models

David M. Chan, Suzanne Petryk, Joseph Gonzalez, Trevor Darrell, John F. Canny

OrganizationsUniversity of California Berkeley

Why you should read this

Proposes CLAIR, a zero-shot image caption evaluation metric that uses large language models to produce quality scores and interpretable natural language explanations that align significantly closer with human judgment than traditional metrics like SPICE and RefCLIP-S.

The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, caption diversity, and specificity. Existing highly-engineered measures attempt to capture specific aspects, but fall short in providing a holistic score that aligns closely with human judgments. Here, we propose CLAIR¹, a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) to evaluate candidate captions. In our evaluations, CLAIR demonstrates a stronger correlation with human judgments of caption quality compared to existing measures. Notably, on Flickr8K-Expert, CLAIR achieves relative correlation improvements over SPICE of 39.6% and over image-augmented methods such as RefCLIP-S of 18.3%. Moreover, CLAIR provides noisy interpretable results by allowing the language model to identify the underlying reasoning behind its assigned score. Code is available at https://davidmchan.github.io/clair/.

Added

2026-10-04

LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, Chris Kedzie

OrganizationsMicrosoft

Why you should read this

Proposes LLM-RUBRIC, an automated evaluation framework that queries large language models across multidimensional rubric criteria and calibrates their probability outputs with a personalized neural network to predict individual human annotator judgments with double the accuracy of standard models.

This paper introduces a framework for the automated evaluation of natural language texts. A manually constructed rubric describes how to assess multiple dimensions of interest. To evaluate a text, a large language model (LLM) is prompted with each rubric question and produces a distribution over potential responses. The LLM predictions often fail to agree well with human judges—indeed, the humans do not fully agree with one another. However, the multiple LLM distributions can be combined to predict each human judge's annotations on all questions, including a summary question that assesses overall quality or relevance. LLM-RUBRIC accomplishes this by training a small feed-forward neural network that includes both judge-specific and judge-independent parameters. When evaluating dialogue systems in a human-AI information-seeking task, we find that LLM-RUBRIC with 9 questions (assessing dimensions such as naturalness, conciseness, and citation quality) predicts human judges' assessment of overall user satisfaction, on a scale of 1–4, with RMS error < 0.5, a 2× improvement over the uncalibrated baseline.

Added

2026-10-02

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra

OrganizationsAmazonMicrosoftThe Alan Turing InstituteUniversity College London

Why you should read this

Proposes JudgeBlender, an ensembling framework that combines judgments from smaller open-source language models across multiple architectures and prompts to match proprietary models in retrieval evaluation at lower cost and with reduced bias.

The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments.

Added

2026-09-30

Large Language Models are not Fair Evaluators

Large Language Models are not Fair Evaluators

Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui

OrganizationsPeking UniversityTencentUniversity of Hong Kong

Why you should read this

Reveals that using large language models as judges introduces severe positional bias that distorts model rankings, and provides effective calibration strategies to align automated evaluations with human judgments.

In this paper, we uncover a positional bias in the evaluation paradigm of adopting large language models (LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models. We find that the quality ranking of candidate responses can be easily hacked by simply altering their order of appearance in the context. This manipulation allows us to skew the evaluation result, making one model appear considerably superior to the other, e.g., Vicuna-13B could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator. We propose a simple yet effective calibration framework to address our discovered positional bias. To evaluate the effectiveness of our framework, we manually annotate the “win/tie/lose” outcomes of responses from ChatGPT and Vicuna-13B in the Vicuna Benchmark’s question prompt. Extensive experiments demonstrate that our approach successfully alleviates evaluation bias, resulting in closer alignment with human judgments. To facilitate future research on more robust large language model comparison, we integrate the techniques in the paper into an easy-to-use toolkit FairEval, along with the human annotations 1.

Added

2026-09-28

Improving Automatic VQA Evaluation Using Large Language Models

Improving Automatic VQA Evaluation Using Large Language Models

Oscar Mañas, Benno Krojer, Aishwarya Agrawal

OrganizationsMcGill UniversityMilaUniversité de Montréal

Why you should read this

Proposes LAVE, an LLM-based evaluation metric that formulates visual question answering assessment as an in-context answer-rating task to align closely with human judgment where traditional exact-match accuracy fails on open-ended generative models.

8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting. However, our community is undergoing a shift towards open-ended generative models and OOD evaluation. In this new paradigm, the existing VQA Accuracy metric is overly stringent and underestimates the performance of VQA systems. Thus, there is a need to develop more robust automatic VQA metrics that serve as a proxy for human judgment. In this work, we propose to leverage the in-context learning capabilities of instruction-tuned large language models (LLMs) to build a better VQA metric. We formulate VQA evaluation as an answer-rating task where the LLM is instructed to score the accuracy of a candidate answer given a set of reference answers. We demonstrate the proposed metric better correlates with human judgment compared to existing metrics across several VQA models and benchmarks. We hope wide adoption of our metric will contribute to better estimating the research progress on the VQA task. We plan to release the evaluation code and collected human judgments.

Added

2026-09-26

Bleu: a Method for Automatic Evaluation of Machine Translation

Bleu: a Method for Automatic Evaluation of Machine Translation

Kishore Papineni, Salim Roukos, Todd Ward, Wei-Jing Zhu

OrganizationsIBM

Why you should read this

Defines the canonical n-gram precision metric that birthed automated evaluation for text generation, illustrating foundational compromises required when quantifying open-ended tasks.

Human evaluations of machine translation are extensive but expensive. Human evaluations can take months to finish and involve human labor that can not be reused. We propose a method of automatic machine translation evaluation that is quick, inexpensive, and language-independent, that correlates highly with human evaluation, and that has little marginal cost per run. We present this method as an automated understudy to skilled human judges which substitutes for them when there is need for quick or frequent evaluations.

Added

2026-03-22