keyword
evaluation metrics
Evaluation metrics are quantitative measures, criteria, or scoring functions used to assess, compare, and benchmark the performance, quality, and effectiveness of machine learning models, algorithms, and computational systems. These measures evaluate how successfully a system achieves its intended task, such as sequence generation, classification, recommendation, or ranking, typically by comparing model predictions or generated outputs against ground-truth references, established baselines, or human judgments. Depending on the application and domain, evaluation metrics can be automated statistical formulas, learned neural scorers, or structured human judgment protocols. They assess diverse dimensions of performance, including accuracy, semantic fidelity, diversity, computational efficiency, and alignment with human preferences, providing standardized criteria for model selection, error analysis, and algorithmic optimization.
10 items

CLAIR: Evaluating Image Captions with Large Language Models
David M. Chan, Suzanne Petryk, Joseph Gonzalez, Trevor Darrell, John F. Canny
Why you should read this
Proposes CLAIR, a zero-shot image caption evaluation metric that uses large language models to produce quality scores and interpretable natural language explanations that align significantly closer with human judgment than traditional metrics like SPICE and RefCLIP-S.
The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, caption diversity, and specificity. Existing highly-engineered measures attempt to capture specific aspects, but fall short in providing a holistic score that aligns closely with human judgments. Here, we propose CLAIR¹, a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) to evaluate candidate captions. In our evaluations, CLAIR demonstrates a stronger correlation with human judgments of caption quality compared to existing measures. Notably, on Flickr8K-Expert, CLAIR achieves relative correlation improvements over SPICE of 39.6% and over image-augmented methods such as RefCLIP-S of 18.3%. Moreover, CLAIR provides noisy interpretable results by allowing the language model to identify the underlying reasoning behind its assigned score. Code is available at https://davidmchan.github.io/clair/.
Added
2026-10-04

Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks
Andrea Sottana, Bin Liang, Kai Zou, Zheng Yuan
Why you should read this
Demonstrates that traditional automatic metrics fail to reflect modern language model performance and benchmark quality on sequence-to-sequence tasks, while establishing GPT-4 as a viable surrogate for human evaluation across text summarisation, simplification, and grammatical error correction.
Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape, and it is becoming clear that the quality of automatic evaluation metrics is not keeping up with the pace of development of generative models. We aim to improve the understanding of current models’ performance by providing a preliminary and hybrid evaluation on a range of open and closed-source generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction (GEC), using both automatic and human evaluation. We also explore the potential of the recently released GPT-4 to act as an evaluator. We find that ChatGPT consistently outperforms many other popular models according to human reviewers on the majority of metrics, while scoring much more poorly when using classic automatic evaluation metrics. We also find that human reviewers rate the gold reference as much worse than the best models’ outputs, indicating the poor quality of many popular benchmarks. Finally, we find that GPT-4 is capable of ranking models’ outputs in a way which aligns reasonably closely to human judgement despite task-specific variations, with a lower alignment in the GEC task.
Added
2026-10-02

Can Large Language Models Be an Alternative to Human Evaluations?
David Cheng-Han Chiang, Hung-yi Lee
Why you should read this
Demonstrates that large language models can reliably substitute for expert human evaluators in assessing generated text quality, producing consistent ratings across instruction variations while addressing key reproducibility bottlenecks in natural language processing.
Human evaluation is indispensable and inevitable for assessing the quality of texts generated by machine learning models or written by humans. However, human evaluation is very difficult to reproduce and its quality is notoriously unstable, hindering fair comparisons among different natural language processing (NLP) models and algorithms. Recently, large language models (LLMs) have demonstrated exceptional performance on unseen tasks when only the task instructions are provided. In this paper, we explore if such an ability of the LLMs can be used as an alternative to human evaluation. We present the LLMs with the exact same instructions, samples to be evaluated, and questions used to conduct human evaluation, and then ask the LLMs to generate responses to those questions; we dub this LLM evaluation. We use human evaluation and LLM evaluation to evaluate the texts in two NLP tasks: open-ended story generation and adversarial attacks. We show that the result of LLM evaluation is consistent with the results obtained by expert human evaluation: the texts rated higher by human experts are also rated higher by the LLMs. We also find that the results of LLM evaluation are stable over different formatting of the task instructions and the sampling algorithm used to generate the answer. We are the first to show the potential of using LLMs to assess the quality of texts and discuss the limitations and ethical considerations of LLM evaluation.
Added
2026-09-28

Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical Solutions
Leslie O'Bray, Max Horn, Bastian Rieck, Karsten M. Borgwardt
Why you should read this
Demonstrates critical flaws in evaluating graph generative models with Maximum Mean Discrepancy and establishes practical guidelines to ensure reliable and standardized model benchmarking.
Graph generative models are a highly active branch of machine learning. Given the steady development of new models of ever-increasing complexity, it is necessary to provide a principled way to evaluate and compare them. In this paper, we enumerate the desirable criteria for such a comparison metric and provide an overview of the status quo of graph generative model comparison in use today, which predominantly relies on the maximum mean discrepancy (MMD). We perform a systematic evaluation of MMD in the context of graph generative model comparison, highlighting some of the challenges and pitfalls researchers inadvertently may encounter. After conducting a thorough analysis of the behaviour of MMD on synthetically-generated perturbed graphs as well as on recently-proposed graph generative models, we are able to provide a suitable procedure to mitigate these challenges and pitfalls. We aggregate our findings into a list of practical recommendations for researchers to use when evaluating graph generative models.
Added
2026-09-26

On Evaluation Metrics for Graph Generative Models
Rylee Thompson, Boris Knyazev, Elahe Ghalebi, Jungtaek Kim, Graham W. Taylor
Why you should read this
Proposes scalable, single-score evaluation metrics for graph generative models based on untrained random graph neural networks, enabling fast and feature-aware measurement of generated graph fidelity and diversity.
In image generation, generative models can be evaluated naturally by visually inspecting model outputs. However, this is not always the case for graph generative models (GGMs), making their evaluation challenging. Currently, the standard process for evaluating GGMs suffers from three critical limitations: i) it does not produce a single score which makes model selection challenging, ii) in many cases it fails to consider underlying edge and node features, and iii) it is prohibitively slow to perform. In this work, we mitigate these issues by searching for scalar, domain-agnostic, and scalable metrics for evaluating and ranking GGMs. To this end, we study existing GGM metrics and neural-network-based metrics emerging from generative models of images that use embeddings extracted from a task-specific network. Motivated by the power of certain Graph Neural Networks (GNNs) to extract meaningful graph representations without any training, we introduce several metrics based on the features extracted by an untrained random GNN. We design experiments to thoroughly test metrics on their ability to measure the diversity and fidelity of generated graphs, as well as their sample and computational efficiency. Depending on the quantity of samples, we recommend one of two random-GNN-based metrics that we show to be more expressive than pre-existing metrics. While we focus on applying these metrics to GGM evaluation, in practice this enables the ability to easily compute the dissimilarity between any two sets of graphs regardless of domain. Our code is released at: this https URL.
Added
2026-09-26

Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
Micah Hodosh, Peter Young, Julia Hockenmaier
Why you should read this
Establishes a unified ranking framework and a benchmark of 8,000 images with multiple descriptive captions to evaluate sentence-based image description and retrieval independently of text generation challenges.
The ability to associate images with natural language sentences that describe what is depicted in them is a hallmark of image understanding, and a prerequisite for applications such as sentence-based image search. In analogy to image search, we propose to frame sentence-based image annotation as the task of ranking a given pool of captions. We introduce a new benchmark collection for sentence-based image description and search, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events. We introduce a number of systems that perform quite well on this task, even though they are only based on features that can be obtained with minimal supervision. Our results clearly indicate the importance of training on multiple captions per image, and of capturing syntactic (word order-based) and semantic features of these captions. We also perform an in-depth comparison of human and automatic evaluation metrics for this task, and propose strategies for collecting human judgments cheaply and on a very large scale, allowing us to augment our collection with additional relevance judgments of which captions describe which image. Our analysis shows that metrics that consider the ranked list of results for each query image or sentence are significantly more robust than metrics that are based on a single response per query. Moreover, our study suggests that the evaluation of ranking-based image description systems may be fully automated.
Added
2026-09-25

CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, Devi Parikh
Why you should read this
Introduces CIDEr, a consensus-based evaluation metric for image captioning that measures similarity against multiple reference descriptions to achieve higher correlation with human judgment than traditional metrics.
Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing. Due to recent progress in object detection, attribute classification, action recognition, etc., there is renewed interest in this area. However, evaluating the quality of descriptions has proven to be challenging. We propose a novel paradigm for evaluating image descriptions that uses human consensus. This paradigm consists of three main parts: a new triplet-based method of collecting human annotations to measure consensus, a new automated metric (CIDEr) that captures consensus, and two new datasets: PASCAL-50S and ABSTRACT-50S that contain 50 sentences describing each image. Our simple metric captures human judgment of consensus better than existing metrics across sentences generated by various sources. We also evaluate five state-of-the-art image description approaches using this new protocol and provide a benchmark for future comparisons. A version of CIDEr named CIDEr-D is available as a part of MS COCO evaluation server to enable systematic evaluation and benchmarking.
Added
2026-09-09

On the Evaluation of (Meta-)solver Approaches
Roberto Amadini, Maurizio Gabbrielli, Tong Liu, Jacopo Mauro
Why you should read this
Analyzes how common evaluation metrics like closed gap and Borda count can produce diametrically opposite rankings for meta-solvers, exposing critical weaknesses in current benchmarking practices and offering strategies for more stable performance assessments across heterogeneous problem scenarios.
Meta-solver approaches exploits a number of individual solvers to potentially build a better solver. To assess the performance of meta-solvers, one can simply adopt the metrics typically used for individual solvers (e.g., runtime or solution quality), or employ more specific evaluation metrics (e.g., by measuring how close the meta-solver gets to its virtual best performance). In this paper, based on some recently published works, we provide an overview of different performance metrics for evaluating (meta-)solvers, by underlying their strengths and weaknesses.
Added
2026-04-10

On Some Pitfalls in Automatic Evaluation and Significance Testing for MT
Stefan Riezler, John T. Maxwell III
Why you should read this
Demonstrates that the NIST metric provides superior discriminatory power over BLEU and proves that approximate randomization offers more accurate significance testing than bootstrap methods, providing a more rigorous methodology for detecting subtle improvements in machine translation systems.
We investigate some pitfalls regarding the discriminatory power of MT evaluation metrics and the accuracy of statistical significance tests. In a discriminative reranking experiment for phrase-based SMT we show that the NIST metric is more sensitive than BLEU or F-score despite their incorporation of aspects of fluency or meaning adequacy into MT evaluation. In an experimental comparison of two statistical significance tests we show that p-values are estimated more conservatively by approximate randomization than by bootstrap tests, thus increasing the likelihood of type-I error for the latter. We point out a pitfall of randomly assessing significance in multiple pairwise comparisons, and conclude with a recommendation to combine NIST with approximate randomization, at more stringent rejection levels than is currently standard.
Added
2026-02-21

Evaluating collaborative filtering recommender systems
Jonathan L. Herlocker, Joseph A. Konstan, Loren Terveen, John Riedl
Why you should read this
Establishes the standard metrics (MAE, Precision, Recall) and experimental protocols required to objectively measure recommendation quality.
Recommender systems have been evaluated in many, often incomparable, ways. In this article, we review the key decisions in evaluating collaborative filtering recommender systems: the user tasks being evaluated, the types of analysis and datasets being used, the ways in which prediction quality is measured, the evaluation of prediction attributes other than quality, and the user-based evaluation of the system as a whole. In addition to reviewing the evaluation strategies used by prior researchers, we present empirical results from the analysis of various accuracy metrics on one content domain where all the tested metrics collapsed roughly into three equivalence classes. Metrics within each equivalency class were strongly correlated, while metrics from different equivalency classes were uncorrelated.
Added
2026-01-25
