Built independently by an author, for readers. Read the story and support ChapterPal

keyword

pairwise evaluation

Pairwise evaluation is an assessment methodology in machine learning and natural language processing where an evaluator directly compares two candidate outputs generated for the same prompt to determine which response is superior according to defined criteria such as correctness, quality, or helpfulness. Rather than assigning absolute numerical scores to individual outputs in isolation, this side-by-side approach relies on relative comparison, enabling human annotators or automated model judges to choose the preferred output or declare a tie. By focusing on relative preference, pairwise evaluation reduces calibration bias and subjective scale variance inherent in independent scoring systems, making it a foundational method for benchmarking generative models, calculating comparative performance rankings, and training reward models.

1 item

JudgeBench: A Benchmark for Evaluating LLM-Based Judges

JudgeBench: A Benchmark for Evaluating LLM-Based Judges

Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca A. Popa, Ion Stoica

OrganizationsUniversity of California BerkeleyWashington University in St. Louis

Why you should read this

Introduces JudgeBench, a benchmark that tests LLM-based evaluators on objective correctness across complex reasoning, math, and coding tasks, revealing that frontier models like GPT-4o perform barely above random guessing.

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requiring stronger judges to evaluate them. Existing benchmarks primarily focus on a judge's alignment with human preferences, but often fail to account for more challenging tasks where crowdsourced human preference is a poor indicator of factual and logical correctness. To address this, we propose a novel evaluation framework to objectively evaluate LLM-based judges. Based on this framework, we propose JudgeBench, a benchmark for evaluating LLM-based judges on challenging response pairs spanning knowledge, reasoning, math, and coding. JudgeBench leverages a novel pipeline for converting existing difficult datasets into challenging response pairs with preference labels reflecting objective correctness. Our comprehensive evaluation on a collection of prompted judges, fine-tuned judges, multi-agent judges, and reward models shows that JudgeBench poses a significantly greater challenge than previous benchmarks, with many strong models (e.g., GPT-4o) performing just slightly better than random guessing. Overall, JudgeBench offers a reliable platform for assessing increasingly advanced LLM-based judges. Data and code are available at this https URL.

Added

2026-09-26