Built independently by an author, for readers. Read the story and support ChapterPal

keyword

LLM evaluators

An LLM evaluator is a large language model configured to assess, score, or compare the quality of text generated by other artificial intelligence systems or human users. In this paradigm, the model serves as an automated judge that reviews outputs against predefined criteria, rubrics, or reference answers to measure attributes such as factual accuracy, coherence, relevance, and safety. LLM evaluators are commonly used for pairwise comparisons to rank competing model responses as well as point-wise scoring across benchmark tasks, offering a scalable and cost-effective alternative to human annotation. While they correlate moderately well with human judgments across many natural language processing tasks, their assessments can be influenced by systemic factors such as position bias, verbosity bias, and self-enhancement bias toward their own generated outputs.

1 item

Large Language Models are not Fair Evaluators

Large Language Models are not Fair Evaluators

Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui

OrganizationsPeking UniversityTencentUniversity of Hong Kong

Why you should read this

Reveals that using large language models as judges introduces severe positional bias that distorts model rankings, and provides effective calibration strategies to align automated evaluations with human judgments.

In this paper, we uncover a positional bias in the evaluation paradigm of adopting large language models (LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models. We find that the quality ranking of candidate responses can be easily hacked by simply altering their order of appearance in the context. This manipulation allows us to skew the evaluation result, making one model appear considerably superior to the other, e.g., Vicuna-13B could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator. We propose a simple yet effective calibration framework to address our discovered positional bias. To evaluate the effectiveness of our framework, we manually annotate the “win/tie/lose” outcomes of responses from ChatGPT and Vicuna-13B in the Vicuna Benchmark’s question prompt. Extensive experiments demonstrate that our approach successfully alleviates evaluation bias, resulting in closer alignment with human judgments. To facilitate future research on more robust large language model comparison, we integrate the techniques in the paper into an easy-to-use toolkit FairEval, along with the human annotations 1.

Added

2026-09-28