keyword
LLM-based judges
LLM-based judges are large language models deployed to evaluate, score, or compare the quality of outputs produced by other artificial intelligence systems or human agents. Serving as an automated and scalable alternative to manual human evaluation, these systems assess responses across open-ended and complex domains such as reasoning, factual knowledge, coding, and mathematical problem-solving. They typically operate through methods such as pairwise comparisons, reward modeling, or scoring against predefined rubrics and criteria. While LLM-based judges provide a cost-effective and rapid means to benchmark model performance and guide reinforcement learning or alignment processes, their reliability depends on their ability to maintain consistency, resist inherent biases such as verbosity or presentation order, and accurately reflect objective correctness.
2 items

JudgeBench: A Benchmark for Evaluating LLM-Based Judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Yuan Tang, Alejandro Cuadron, Chenguang Wang, Raluca A. Popa, Ion Stoica
Why you should read this
Introduces JudgeBench, a benchmark that tests LLM-based evaluators on objective correctness across complex reasoning, math, and coding tasks, revealing that frontier models like GPT-4o perform barely above random guessing.
LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requiring stronger judges to evaluate them. Existing benchmarks primarily focus on a judge's alignment with human preferences, but often fail to account for more challenging tasks where crowdsourced human preference is a poor indicator of factual and logical correctness. To address this, we propose a novel evaluation framework to objectively evaluate LLM-based judges. Based on this framework, we propose JudgeBench, a benchmark for evaluating LLM-based judges on challenging response pairs spanning knowledge, reasoning, math, and coding. JudgeBench leverages a novel pipeline for converting existing difficult datasets into challenging response pairs with preference labels reflecting objective correctness. Our comprehensive evaluation on a collection of prompted judges, fine-tuned judges, multi-agent judges, and reward models shows that JudgeBench poses a significantly greater challenge than previous benchmarks, with many strong models (e.g., GPT-4o) performing just slightly better than random guessing. Overall, JudgeBench offers a reliable platform for assessing increasingly advanced LLM-based judges. Data and code are available at this https URL.
Added
2026-09-26

A Survey on LLM-as-a-Judge
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Yuanzhuo Wang, Jian Guo
Why you should read this
Presents a comprehensive guide to building reliable LLM-as-a-judge evaluators, detailing methods to reduce inconsistency and bias, introducing a dedicated reliability benchmark, and providing actionable strategies for automated assessment across diverse tasks.
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains a challenging task due to inherent subjectivity, variability, and scale. Large Language Models (LLMs) have achieved remarkable success across diverse domains, leading to the emergence of "LLM-as-a-Judge," where LLMs are employed as evaluators for complex tasks. With their ability to process diverse data types and provide scalable, cost-effective, and consistent assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-Judge systems remains a significant challenge that requires careful design and standardization. This paper provides a comprehensive survey of LLM-as-a-Judge, addressing the core question: How can reliable LLM-as-a-Judge systems be built? We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse assessment scenarios. Additionally, we propose methodologies for evaluating the reliability of LLM-as-a-Judge systems, supported by a novel benchmark designed for this purpose. To advance the development and real-world deployment of LLM-as-a-Judge systems, we also discussed practical applications, challenges, and future directions. This survey serves as a foundational reference for researchers and practitioners in this rapidly evolving field.
Added
2026-09-20

