Humans or LLMs as the Judge? A Study on Judgement Bias
Guiming ChenShunian ChenZiche LiuFeng JiangBenyou Wang
Reveals critical vulnerabilities in human and automated evaluation pipelines by introducing a reference-free framework to measure authority, gender, and misinformation biases, demonstrating that even advanced language model judges can be systematically manipulated through bias-exploiting attacks.
As large language models (LLMs) and human annotators are increasingly used to evaluate artificial intelligence outputs, assessing the integrity and consistency of these evaluations has become critical. In open-ended tasks where standard answer keys do not exist, automated and human evaluators frequently display systematic biases that undermine benchmark validity. The article evaluates four distinct biases across human and LLM judges: misinformation oversight bias, gender bias, authority bias, and beauty bias. Its primary objective is to quantify how vulnerable these evaluators are to superficial alterations and demonstrate how easily their evaluation decisions can be manipulated.
To conduct this assessment without relying on predefined ground-truth answers, the authors implemented a controlled intervention framework. They generated 142 middle-school-level test questions across the six levels of the revised Bloom’s Taxonomy and produced baseline answer pairs using advanced LLMs. The researchers then created experimental pairs by injecting specific perturbations into one of the answers: subtle factual errors, gender-biased phrasing, fabricated citations, or cosmetic formatting elements such as emojis and markdown. Sixty university student evaluators and several prominent LLM judges evaluated thousands of randomized answer pairings. Evaluator resilience was measured by the attack success rate (ASR), which represents the frequency with which a perturbation successfully shifted the judge’s preference toward the altered answer.
The findings reveal that both human and automated judges suffer from substantial biases. First, authority bias is pervasive: when fake citations were added to answers, nearly every LLM and human judge performed at or worse than a random baseline (ASRs ranging from 32% to 89%), indicating an uncritical trust in authoritative formatting. Second, human evaluators exhibited virtually no gender bias (6% ASR) due to higher baseline social awareness, whereas LLMs demonstrated noticeable gender bias (13% to 34% ASR). Third, top-tier models like GPT-4o and Claude-3 excelled at detecting factual errors with failure rates below 10%, while human judges and lower-tier LLMs overlooked factual errors in more than 20% of cases. Fourth, cosmetic formatting (beauty bias) significantly skewed human decisions (47% ASR) and affected several LLMs. Finally, adversarial testing showed that combining fake citations with rich formatting can successfully deceive leading LLM judges into preferring incorrect or lower-quality answers up to 50% of the time, especially when baseline answer quality is close.
These results demonstrate that using LLMs or unstandardized human pools as evaluators introduces serious risks into model benchmarking, quality assurance, and automated decision pipelines. The fact that superficial features like fake citations and markdown formatting can override factual correctness means that systems evaluated by LLMs can easily be gamed without genuine performance gains. This misalignment poses compliance, safety, and reputational risks for organizations that deploy automated judges without strict validation protocols.
Organizations should not rely solely on zero-shot LLM evaluations or unchecked human reviewers for critical quality assessments. Stakeholders should implement multi-layered evaluation processes that explicitly verify references, strip out cosmetic distractors like emojis and markdown before scoring, and refine LLM training to decouple authority cues from factual validity. Future research must expand benchmarking to encompass broader demographics and explore additional biases such as syntactic framing and tone.
Confidence in these findings is supported by thousands of structured evaluations across multiple model families. However, users should interpret the results in light of certain limitations: the benchmark was restricted to 142 foundational questions, and the human evaluation pool consisted entirely of university students, whose sensitivity to gender and factual nuances may not fully reflect broader populations. Despite these constraints, the evidence confirms that current human- and LLM-as-a-judge approaches require immediate robustness improvements before being trusted in high-stakes environments.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This paper establishes the foundational LLM-as-a-judge paradigm and identifies core evaluation biases that the source directly critiques and probes.
- Paper: Large Language Models are not Fair Evaluators, Peiyi Wang et al. (2024). This work demonstrates positional and fairness flaws in LLMs acting as judges, providing an essential baseline of evaluator vulnerability examined by the source.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). This paper introduces reference-free evaluation of text using LLM prompting and chain-of-thought, underpinning the evaluation frameworks analyzed in the source.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It provides foundational evidence that chain-of-thought rationales can obscure underlying model biases, motivating the source's investigation into judgment bias.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey provides a comprehensive taxonomy of LLM evaluation benchmarks and paradigms, establishing the broader context of evaluation risks addressed in the source.
- Paper: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation, Tianxiang Sun et al. (2022). It demonstrates how pre-trained model metrics suffer from social and demographic biases, establishing the conceptual basis for demographic bias in automated evaluation.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This comprehensive survey organizes the entire literature on the LLM-as-a-judge paradigm, incorporating the judgment biases and vulnerabilities demonstrated in the source.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). It synthesizes subsequent methodologies, bias-correction strategies, and multi-agent judging frameworks that build upon the evaluation vulnerabilities highlighted in the source.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). This benchmark extends the study of LLM judges by introducing objective, ground-truth-verified paired evaluations to rigorously benchmark judge accuracy and bias.
- Paper: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment, Vyas Raina et al. (2024). It builds directly on the vulnerability of LLM evaluators by developing universal adversarial phrasing attacks to manipulate automated judge scoring.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This work constructs an adversarial meta-evaluation benchmark to test automated judges against surface-level biases and superficial polish.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). It proposes a post-processing calibration framework to systematically mitigate and align the evaluator biases analyzed in the source.
- Paper: MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark, Dongping Chen et al. (2024). This paper extends the LLM-as-a-judge evaluation paradigm and its associated judgment biases into multimodal vision-language settings.
