Built independently by an author, for readers. Read the story and support ChapterPal

keyword

evaluation metrics

Evaluation metrics are quantitative measures, criteria, or scoring functions used to assess, compare, and benchmark the performance, quality, and effectiveness of machine learning models, algorithms, and computational systems. These measures evaluate how successfully a system achieves its intended task, such as sequence generation, classification, recommendation, or ranking, typically by comparing model predictions or generated outputs against ground-truth references, established baselines, or human judgments. Depending on the application and domain, evaluation metrics can be automated statistical formulas, learned neural scorers, or structured human judgment protocols. They assess diverse dimensions of performance, including accuracy, semantic fidelity, diversity, computational efficiency, and alignment with human preferences, providing standardized criteria for model selection, error analysis, and algorithmic optimization.

10 items

CLAIR: Evaluating Image Captions with Large Language Models

CLAIR: Evaluating Image Captions with Large Language Models

David M. Chan, Suzanne Petryk, Joseph Gonzalez, Trevor Darrell, John F. Canny

OrganizationsUniversity of California Berkeley

Why you should read this

Proposes CLAIR, a zero-shot image caption evaluation metric that uses large language models to produce quality scores and interpretable natural language explanations that align significantly closer with human judgment than traditional metrics like SPICE and RefCLIP-S.

The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, caption diversity, and specificity. Existing highly-engineered measures attempt to capture specific aspects, but fall short in providing a holistic score that aligns closely with human judgments. Here, we propose CLAIR¹, a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) to evaluate candidate captions. In our evaluations, CLAIR demonstrates a stronger correlation with human judgments of caption quality compared to existing measures. Notably, on Flickr8K-Expert, CLAIR achieves relative correlation improvements over SPICE of 39.6% and over image-augmented methods such as RefCLIP-S of 18.3%. Moreover, CLAIR provides noisy interpretable results by allowing the language model to identify the underlying reasoning behind its assigned score. Code is available at https://davidmchan.github.io/clair/.

Added

2026-10-04

Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks

Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks

Andrea Sottana, Bin Liang, Kai Zou, Zheng Yuan

Why you should read this

Demonstrates that traditional automatic metrics fail to reflect modern language model performance and benchmark quality on sequence-to-sequence tasks, while establishing GPT-4 as a viable surrogate for human evaluation across text summarisation, simplification, and grammatical error correction.

Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape, and it is becoming clear that the quality of automatic evaluation metrics is not keeping up with the pace of development of generative models. We aim to improve the understanding of current models’ performance by providing a preliminary and hybrid evaluation on a range of open and closed-source generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction (GEC), using both automatic and human evaluation. We also explore the potential of the recently released GPT-4 to act as an evaluator. We find that ChatGPT consistently outperforms many other popular models according to human reviewers on the majority of metrics, while scoring much more poorly when using classic automatic evaluation metrics. We also find that human reviewers rate the gold reference as much worse than the best models’ outputs, indicating the poor quality of many popular benchmarks. Finally, we find that GPT-4 is capable of ranking models’ outputs in a way which aligns reasonably closely to human judgement despite task-specific variations, with a lower alignment in the GEC task.

Added

2026-10-02

Can Large Language Models Be an Alternative to Human Evaluations?

Can Large Language Models Be an Alternative to Human Evaluations?

David Cheng-Han Chiang, Hung-yi Lee

OrganizationsNational Taiwan University

Why you should read this

Demonstrates that large language models can reliably substitute for expert human evaluators in assessing generated text quality, producing consistent ratings across instruction variations while addressing key reproducibility bottlenecks in natural language processing.

Human evaluation is indispensable and inevitable for assessing the quality of texts generated by machine learning models or written by humans. However, human evaluation is very difficult to reproduce and its quality is notoriously unstable, hindering fair comparisons among different natural language processing (NLP) models and algorithms. Recently, large language models (LLMs) have demonstrated exceptional performance on unseen tasks when only the task instructions are provided. In this paper, we explore if such an ability of the LLMs can be used as an alternative to human evaluation. We present the LLMs with the exact same instructions, samples to be evaluated, and questions used to conduct human evaluation, and then ask the LLMs to generate responses to those questions; we dub this LLM evaluation. We use human evaluation and LLM evaluation to evaluate the texts in two NLP tasks: open-ended story generation and adversarial attacks. We show that the result of LLM evaluation is consistent with the results obtained by expert human evaluation: the texts rated higher by human experts are also rated higher by the LLMs. We also find that the results of LLM evaluation are stable over different formatting of the task instructions and the sampling algorithm used to generate the answer. We are the first to show the potential of using LLMs to assess the quality of texts and discuss the limitations and ethical considerations of LLM evaluation.

Added

2026-09-28

Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical Solutions

Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical Solutions

Leslie O'Bray, Max Horn, Bastian Rieck, Karsten M. Borgwardt

OrganizationsETH ZurichHelmholtz MunichSwiss Institute of BioinformaticsTechnical University of Munich

Why you should read this

Demonstrates critical flaws in evaluating graph generative models with Maximum Mean Discrepancy and establishes practical guidelines to ensure reliable and standardized model benchmarking.

Graph generative models are a highly active branch of machine learning. Given the steady development of new models of ever-increasing complexity, it is necessary to provide a principled way to evaluate and compare them. In this paper, we enumerate the desirable criteria for such a comparison metric and provide an overview of the status quo of graph generative model comparison in use today, which predominantly relies on the maximum mean discrepancy (MMD). We perform a systematic evaluation of MMD in the context of graph generative model comparison, highlighting some of the challenges and pitfalls researchers inadvertently may encounter. After conducting a thorough analysis of the behaviour of MMD on synthetically-generated perturbed graphs as well as on recently-proposed graph generative models, we are able to provide a suitable procedure to mitigate these challenges and pitfalls. We aggregate our findings into a list of practical recommendations for researchers to use when evaluating graph generative models.

Added

2026-09-26

On Evaluation Metrics for Graph Generative Models

On Evaluation Metrics for Graph Generative Models

Rylee Thompson, Boris Knyazev, Elahe Ghalebi, Jungtaek Kim, Graham W. Taylor

OrganizationsPohang University of Science and TechnologySamsung SAILUniversity of GuelphVector Institute

Why you should read this

Proposes scalable, single-score evaluation metrics for graph generative models based on untrained random graph neural networks, enabling fast and feature-aware measurement of generated graph fidelity and diversity.

In image generation, generative models can be evaluated naturally by visually inspecting model outputs. However, this is not always the case for graph generative models (GGMs), making their evaluation challenging. Currently, the standard process for evaluating GGMs suffers from three critical limitations: i) it does not produce a single score which makes model selection challenging, ii) in many cases it fails to consider underlying edge and node features, and iii) it is prohibitively slow to perform. In this work, we mitigate these issues by searching for scalar, domain-agnostic, and scalable metrics for evaluating and ranking GGMs. To this end, we study existing GGM metrics and neural-network-based metrics emerging from generative models of images that use embeddings extracted from a task-specific network. Motivated by the power of certain Graph Neural Networks (GNNs) to extract meaningful graph representations without any training, we introduce several metrics based on the features extracted by an untrained random GNN. We design experiments to thoroughly test metrics on their ability to measure the diversity and fidelity of generated graphs, as well as their sample and computational efficiency. Depending on the quantity of samples, we recommend one of two random-GNN-based metrics that we show to be more expressive than pre-existing metrics. While we focus on applying these metrics to GGM evaluation, in practice this enables the ability to easily compute the dissimilarity between any two sets of graphs regardless of domain. Our code is released at: this https URL.

Added

2026-09-26

Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics

Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics

Micah Hodosh, Peter Young, Julia Hockenmaier

OrganizationsUniversity of Illinois Urbana-Champaign

Why you should read this

Establishes a unified ranking framework and a benchmark of 8,000 images with multiple descriptive captions to evaluate sentence-based image description and retrieval independently of text generation challenges.

The ability to associate images with natural language sentences that describe what is depicted in them is a hallmark of image understanding, and a prerequisite for applications such as sentence-based image search. In analogy to image search, we propose to frame sentence-based image annotation as the task of ranking a given pool of captions. We introduce a new benchmark collection for sentence-based image description and search, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events. We introduce a number of systems that perform quite well on this task, even though they are only based on features that can be obtained with minimal supervision. Our results clearly indicate the importance of training on multiple captions per image, and of capturing syntactic (word order-based) and semantic features of these captions. We also perform an in-depth comparison of human and automatic evaluation metrics for this task, and propose strategies for collecting human judgments cheaply and on a very large scale, allowing us to augment our collection with additional relevance judgments of which captions describe which image. Our analysis shows that metrics that consider the ranked list of results for each query image or sentence are significantly more robust than metrics that are based on a single response per query. Moreover, our study suggests that the evaluation of ranking-based image description systems may be fully automated.

Added

2026-09-25

CIDEr: Consensus-based image description evaluation

CIDEr: Consensus-based image description evaluation

Ramakrishna Vedantam, C. Lawrence Zitnick, Devi Parikh

OrganizationsMicrosoftVirginia Tech

Why you should read this

Introduces CIDEr, a consensus-based evaluation metric for image captioning that measures similarity against multiple reference descriptions to achieve higher correlation with human judgment than traditional metrics.

Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing. Due to recent progress in object detection, attribute classification, action recognition, etc., there is renewed interest in this area. However, evaluating the quality of descriptions has proven to be challenging. We propose a novel paradigm for evaluating image descriptions that uses human consensus. This paradigm consists of three main parts: a new triplet-based method of collecting human annotations to measure consensus, a new automated metric (CIDEr) that captures consensus, and two new datasets: PASCAL-50S and ABSTRACT-50S that contain 50 sentences describing each image. Our simple metric captures human judgment of consensus better than existing metrics across sentences generated by various sources. We also evaluate five state-of-the-art image description approaches using this new protocol and provide a benchmark for future comparisons. A version of CIDEr named CIDEr-D is available as a part of MS COCO evaluation server to enable systematic evaluation and benchmarking.

Added

2026-09-09