Large Language Models are not Fair Evaluators
Peiyi WangLei LiLiang ChenZefan CaiDawei ZhuBinghuai LinYunbo CaoLingpeng KongQi LiuTianyu Liu
Reveals that using large language models as judges introduces severe positional bias that distorts model rankings, and provides effective calibration strategies to align automated evaluations with human judgments.
As organizations increasingly deploy artificial intelligence systems, evaluating model quality reliably and cost-effectively has become critical. Relying on large language models as automated judges has become a popular alternative to expensive and slow human evaluation. However, the reliability and fairness of using these models to judge and compare competing outputs remain uncertain.
The article evaluates whether advanced language models suffer from positional bias when judging candidate responses. It demonstrates a calibration framework designed to correct these distortions and improve alignment with human judgments.
To examine this issue, the researchers tested leading models acting as evaluators on a standard benchmark across 80 questions spanning nine categories. They analyzed how swapping the presentation order of two competing responses affected the scoring. They then introduced and tested three calibration strategies: generating written evaluation evidence before scoring, balancing positions by swapping candidate order and averaging the scores, and using an entropy-based scoring metric to identify the most uncertain comparisons for targeted human review. The authors established a ground-truth baseline by manually annotating all 80 benchmark questions independently.
The investigation revealed substantial positional bias in automated judges. When comparing responses, simply swapping the presentation order caused GPT-4 to produce conflicting results in 46.3% of test cases, while ChatGPT produced conflicting results in 82.5% of cases. GPT-4 systematically favored the response shown first, whereas ChatGPT favored the response shown second. This vulnerability was especially acute when the quality gap between responses was small. Applying the automated calibration techniques improved judgment accuracy relative to human consensus by 9.8 percentage points for GPT-4 and 14.3 percentage points for ChatGPT. Furthermore, selectively routing just the top 20% most uncertain evaluations to human reviewers enabled the system to match or exceed average human evaluation accuracy while reducing human annotation costs by up to 39%.
These findings indicate that off-the-shelf language models cannot be trusted as fair, uncalibrated judges, posing significant risks of skewed model selection, flawed safety assessments, and misleading performance tracking. The common practice of asking a model to provide a numerical score before giving an explanation exacerbates these distortions. By forcing models to articulate evaluation reasoning first and averaging scores across swapped positions, organizations can significantly enhance evaluation robustness at low computational cost.
Decision-makers using automated model evaluation should immediately stop relying on single-pass, uncalibrated scoring prompts. Instead, organizations should adopt multi-sample evidence generation and position swapping in automated pipelines. When high accuracy is critical, teams should implement human-in-the-loop workflows that use diversity entropy metrics to flag ambiguous cases for manual review rather than conducting blanket human evaluations.
The findings are bounded by the 80-question test set and the specific model versions evaluated. While the proposed calibration toolkit consistently improves alignment with human judgments, the underlying root causes of positional bias inside language models remain unaddressed and require further investigation.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). It introduces the LLM-as-a-judge paradigm and the MT-Bench/Chatbot Arena setups on which the source study's positional bias investigations directly build.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). It establishes foundational principles for identifying and calibrating prompt ordering and in-context position biases in large language models.
- Paper: Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, Yao Lu et al. (2021). It analyzes severe prompt order sensitivity in few-shot LLM predictions, motivating the calibration strategies necessary when prompting LLMs for evaluations.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). It pioneered prompting LLMs like GPT-4 directly as reference-free evaluators, providing the operational framework whose fairness and positional vulnerabilities the source investigates.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It demonstrates how subtle input biases sway LLM reasoning while remaining hidden in chain-of-thought justifications, explaining why LLM evaluators struggle with positional fairness.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). It explores replacing human evaluators with large language models, highlighting the initial viability and limitations of LLM judges.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). It directly advances the problem of position bias in LLM-based pairwise evaluators by proposing the PORTIA split-and-merge calibration framework across multiple judgment formats.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). It provides a comprehensive survey and taxonomy of LLM-as-a-judge methodologies, systematizing evaluation biases like positional and length bias along with post-processing mitigation strategies.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). It offers an extensive architectural survey on the opportunities and systematic challenges of LLM-as-a-judge, explicitly covering response-swapping and debiasing paradigms.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). It extends the evaluation of LLM judges by establishing objective, ground-truth benchmarks on complex reasoning while systematically controlling for response-ordering bias.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). It investigates adversarial vulnerabilities and surface-level heuristics in LLM evaluators, advancing the meta-evaluation of automated judges beyond positional effects.
- Paper: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment, Vyas Raina et al. (2024). It explores adversarial vulnerabilities in zero-shot LLM assessment, analyzing how automated evaluators can be easily manipulated across pairwise and absolute grading schemes.
- Paper: Branch-Solve-Merge Improves Large Language Model Evaluation and Generation, Swarnadeep Saha et al. (2024). It introduces a structured decomposition framework that reduces systematic evaluation biases and improves alignment with human judgments on MT-Bench.
