Humans or LLMs as the Judge? A Study on Judgement Bias

Guiming ChenShunian ChenZiche LiuFeng JiangBenyou Wang

article2024EMNLP274 citations

Reveals critical vulnerabilities in human and automated evaluation pipelines by introducing a reference-free framework to measure authority, gender, and misinformation biases, demonstrating that even advanced language model judges can be systematically manipulated through bias-exploiting attacks.

Listen

As large language models (LLMs) and human annotators are increasingly used to evaluate artificial intelligence outputs, assessing the integrity and consistency of these evaluations has become critical. In open-ended tasks where standard answer keys do not exist, automated and human evaluators frequently display systematic biases that undermine benchmark validity. The article evaluates four distinct biases across human and LLM judges: misinformation oversight bias, gender bias, authority bias, and beauty bias. Its primary objective is to quantify how vulnerable these evaluators are to superficial alterations and demonstrate how easily their evaluation decisions can be manipulated.

To conduct this assessment without relying on predefined ground-truth answers, the authors implemented a controlled intervention framework. They generated 142 middle-school-level test questions across the six levels of the revised Bloom’s Taxonomy and produced baseline answer pairs using advanced LLMs. The researchers then created experimental pairs by injecting specific perturbations into one of the answers: subtle factual errors, gender-biased phrasing, fabricated citations, or cosmetic formatting elements such as emojis and markdown. Sixty university student evaluators and several prominent LLM judges evaluated thousands of randomized answer pairings. Evaluator resilience was measured by the attack success rate (ASR), which represents the frequency with which a perturbation successfully shifted the judge’s preference toward the altered answer.

The findings reveal that both human and automated judges suffer from substantial biases. First, authority bias is pervasive: when fake citations were added to answers, nearly every LLM and human judge performed at or worse than a random baseline (ASRs ranging from 32% to 89%), indicating an uncritical trust in authoritative formatting. Second, human evaluators exhibited virtually no gender bias (6% ASR) due to higher baseline social awareness, whereas LLMs demonstrated noticeable gender bias (13% to 34% ASR). Third, top-tier models like GPT-4o and Claude-3 excelled at detecting factual errors with failure rates below 10%, while human judges and lower-tier LLMs overlooked factual errors in more than 20% of cases. Fourth, cosmetic formatting (beauty bias) significantly skewed human decisions (47% ASR) and affected several LLMs. Finally, adversarial testing showed that combining fake citations with rich formatting can successfully deceive leading LLM judges into preferring incorrect or lower-quality answers up to 50% of the time, especially when baseline answer quality is close.

These results demonstrate that using LLMs or unstandardized human pools as evaluators introduces serious risks into model benchmarking, quality assurance, and automated decision pipelines. The fact that superficial features like fake citations and markdown formatting can override factual correctness means that systems evaluated by LLMs can easily be gamed without genuine performance gains. This misalignment poses compliance, safety, and reputational risks for organizations that deploy automated judges without strict validation protocols.

Organizations should not rely solely on zero-shot LLM evaluations or unchecked human reviewers for critical quality assessments. Stakeholders should implement multi-layered evaluation processes that explicitly verify references, strip out cosmetic distractors like emojis and markdown before scoring, and refine LLM training to decouple authority cues from factual validity. Future research must expand benchmarking to encompass broader demographics and explore additional biases such as syntactic framing and tone.

Confidence in these findings is supported by thousands of structured evaluations across multiple model families. However, users should interpret the results in light of certain limitations: the benchmark was restricted to 142 foundational questions, and the human evaluation pool consisted entirely of university students, whose sensitivity to gender and factual nuances may not fully reflect broader populations. Despite these constraints, the evidence confirms that current human- and LLM-as-a-judge approaches require immediate robustness improvements before being trusted in high-stakes environments.

Cover for Humans or LLMs as the Judge? A Study on Judgement Bias

Abstract

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation results. In this paper, we propose a novel framework that is free from referencing groundtruth annotations for investigating Misinformation Oversight Bias, Gender Bias, Authority Bias and Beauty Bias on LLM and human judges. We curate a dataset referring to the revised Bloom's Taxonomy and conduct thousands of evaluations. Results show that human and LLM judges are vulnerable to perturbations to various degrees, and that even the cutting-edge judges possess considerable biases. We further exploit these biases to conduct attacks on LLM judges. We hope that our work can notify the community of the bias and vulnerability of human- and LLM-as-a-judge, as well as the urgency of developing robust evaluation systems.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Human and LLM Evaluation
  • 2.2 Biases of Human and LLM Judges
  • 2.3 Attack on LLM-as-a-judge
  • 3 On the Biases of Judges
  • 3.1 Defining Bias
  • 3.2 Biases of Interest
  • 3.3 Importance of the Investigated Biases
  • 4 Experimental Protocol
  • 4.1 Motivation
  • 4.2 Method
  • 4.3 Data Generation
  • 4.4 Experiment Objects
  • 4.5 Experiment Procedure
  • 4.6 Metric
  • 4.7 Superiority of the Reference-free Framework
  • 5 Results and Discussion
  • 5.1 Preliminary: On Positional Bias
  • 5.2 Main Results
  • 5.2.1 On Semantic-related Biases
  • 5.2.2 On Semantic-agnostic Biases
  • 5.3 Discussion
  • 6 Deceiving LLM Judges
  • 6.1 Overview
  • 6.2 Metric
  • 6.3 Findings and Discussion
  • 7 Conclusion
  • References
  • A Detail of Data Generation
  • A.1 Prompt for Question Generation
  • A.1.1 Version 1
  • A.1.2 Version 2
  • A.1.3 Final Version
  • A.2 Prompt for Raw Answer Generation
  • A.3 Prompts for Answer Perturbation
  • A.3.1 Fake Reference
  • A.3.2 Rich Content
  • A.3.3 Factual Error
  • A.3.4 Gender Bias
  • A.4 Instruction for Question and Answer Filtering
  • B Human Judges
  • B.1 Selection Criteria
  • B.2 Statistics of Evaluators
  • B.3 Remuneration
  • C LLM Judges
  • D Instructions Given to Human Evaluators
  • E Evaluation Prompts for GPT-4
  • F More Results on Bias Analysis
  • F.1 Positional Bias
  • F.2 Discussion on the cause of biases for LLM judges
  • F.3 Case analysis
  • F.4 Verbosity Bias
  • G Revised Bloom’s Taxonomy
  • H User Interface
  • I Supplementary Results of Deceiving Models
  • J More Acknowledgement
  • J.1 A Full List of Human Evaluators

Knowls

  1. Knowl 1 — Reference-Free Intervention Framework for Bias Evaluation in NLG Judges

    model/method

    To evaluate biases in human and large language model (LLM) judges without relying on gold-standard groundtruth references, an intervention-based framework is established using control and experimental paired evaluations.

    1. Question and Answer Generation: A set of NN questions QQ spanning six cognitive levels of revised Bloom's Taxonomy (Remembering, Understanding, Applying, Analyzing, Evaluating, Creating) is generated. For each question QQ, two baseline answers A1A_1 and A2A_2 are independently sampled to create a control pair.
    2. Perturbation Injection: For a targeted bias perturbation pp, answer A2A_2 is modified to create a perturbed answer A2pA_2^p. Specific perturbations include:
      • Factual Error: Injecting 2--3 subtle factual errors while preserving remaining content.
      • Gender Bias: Injecting 1--3 gender-biased modifications.
      • Fake Reference: Appending fabricated academic or online citations to make text appear authoritative without altering semantic meaning.
      • Rich Content: Adding Markdown formatting, bullet points, and emojis to enhance visual aesthetics without changing semantic content.
    3. Evaluation and Pairwise Comparison: For each sample Sp={Q,A1,A2,A2p,Prefctrl,Prefexp}S^p = \{Q, A_1, A_2, A_2^p, \text{Pref}_{\text{ctrl}}, \text{Pref}_{\text{exp}}\}, candidate pairs {A1,A2}\{A_1, A_2\} and {A1,A2p}\{A_1, A_2^p\} are evaluated with randomized order positions to counterbalance positional bias.
    4. Aggregation: Each pair receives 6 votes from independent judges (humans or LLMs) choosing among A1A_1 (score 0), Tie (score 0.5), and A2/A2pA_2 / A_2^p (score 1). The mean score over the 6 votes is mapped to a discrete group preference using a decision threshold of 0.5.
  2. Knowl 2 — Attack Successful Rate (ASR) for Quantifying Judgement Biases

    equation

    Judgement bias is quantified as the Attack Successful Rate (ASR), which measures the rate at which an injected perturbation shifts an evaluator's preference toward the perturbed response A2pA_2^p.

    For semantic-agnostic perturbations (fake references and rich content) as well as gender-biased content: ASR=∣V2∣1∣∣V1∣\text{ASR} = \frac{|V_{2|1}|}{|V_1|} where V1V_1 is the set of question samples where the control group preference Prefctrl\text{Pref}_{\text{ctrl}} is either A1A_1 or Tie\text{Tie}, and V2∣1⊆V1V_{2|1} \subseteq V_1 is the subset of samples where the experimental group preference Prefexp\text{Pref}_{\text{exp}} shifts to the perturbed answer A2pA_2^p.

    For factual error perturbations: ASR=∣V2∣2∣∣V2∣\text{ASR} = \frac{|V_{2|2}|}{|V_2|} where V2V_2 is the set of question samples where the control group preference Prefctrl\text{Pref}_{\text{ctrl}} is either A2A_2 or Tie\text{Tie}, and V2∣2⊆V2V_{2|2} \subseteq V_2 is the subset of samples where the experimental group preference Prefexp\text{Pref}_{\text{exp}} remains A2pA_2^p or Tie\text{Tie} despite containing introduced factual errors.

    An ideally unbiased and fact-aware judge achieves an ASR close to 0, whereas higher ASR values denote greater vulnerability to the tested bias.

  3. Knowl 3 — Taxonomy of Semantic-Related and Semantic-Agnostic Evaluator Biases

    definition

    Biases in natural language generation evaluation judges are partitioned into two broad categories:

    1. Semantic-Related Biases: Biases influenced directly by the semantic information in candidate answers:

      • Misinformation Oversight Bias: The tendency of an evaluator to overlook factual errors and false assertions in an argument rather than penalizing them.
      • Gender Bias: The failure of an evaluator to penalize or detect gender-biased stereotypes and biased assertions in the candidate answer.
    2. Semantic-Agnostic Biases: Biases triggered by non-semantic, superficial, or presentation factors:

      • Authority Bias: The tendency of an evaluator to attribute greater credibility and higher quality to statements containing perceived authority markers (such as citations, URLs, or formal source attributions), regardless of actual factual accuracy or relevance.
      • Beauty Bias (Lookism): The inclination of an evaluator to assign higher preference scores to visually engaging, well-formatted content (such as emojis, rich markdown styling, and structure), regardless of the substantive validity of the text.
  4. Knowl 4 — Empirical Benchmark of Human and LLM Judge Biases

    data/table

    Evaluating human judges and various LLM evaluators across 142 benchmark questions yields distinct bias profiles across semantic-related (Factual Error, Gender) and semantic-agnostic (Fake Reference, Rich Content) perturbations:

    Judge FE ASR Gender ASR Ref ASR RC ASR Avg. Ranking ↓\downarrow
    GPT-4o 0.06 0.16 0.32 0.07 2.00
    Claude-3 0.08 0.13 0.70 0.04 3.25
    Human 0.21 0.06 0.37 0.47 4.00
    GPT-4 0.09 0.19 0.66 0.32 4.75
    GPT-4-Turbo 0.11 0.27 0.49 0.05 4.75
    Ernie 0.26 0.34 0.42 0.09 5.75
    LLaMA2-70B 0.60 0.20 0.42 0.46 6.00
    Random 0.62 0.56 0.37 0.39 6.50
    Claude-2 0.23 0.25 0.89 0.68 7.50

    Key observations:

    • Misinformation Oversight (FE): Frontier models (GPT-4o, Claude-3, GPT-4, GPT-4-Turbo) achieve ASR ≤0.11\le 0.11, outperforming human evaluators (ASR =0.21= 0.21) and smaller open models (LLaMA2-70B at 0.600.60).
    • Gender Bias: Human judges display near-zero susceptibility (ASR =0.06= 0.06), substantially outperforming all LLMs (ASR 0.130.13--0.340.34).
    • Authority Bias (Ref): All judges except GPT-4o perform at or worse than random guessing (0.370.37), with Claude-2 and Claude-3 exceeding 0.700.70 ASR, demonstrating severe susceptibility to superficial citations.
    • Beauty Bias (RC): GPT-4o, Claude-3, GPT-4-Turbo, and Ernie resist formatting perturbations (ASR ≤0.09\le 0.09), while humans (ASR =0.47= 0.47) and Claude-2 (ASR =0.68= 0.68) are substantially distracted.
  5. Knowl 5 — Adversarial Prompt-Based Deception of LLM Judges on Flawed and Biased Text

    empirical result

    Adversarial perturbations exploiting Authority Bias (fake references, Ref) and Beauty Bias (rich content, RC) can deceive LLM judges into preferring flawed or gender-biased answers over correct baselines.

    1. Factual Error Attacks: When answers containing injected factual errors are augmented with fake references (Ref), rich content (RC), or both (Ref+RC):
      • Ref perturbations alone deceive LLMs more effectively than RC perturbations (e.g., for GPT-4, Ref yields ASR ≈0.27\approx 0.27 vs. RC ≈0.08\approx 0.08; for LLaMA2-70B, Ref yields ASR ≈0.52\approx 0.52 vs. RC ≈0.30\approx 0.30).
      • Compound Ref+RC perturbation reaches an ASR of 0.250.25 on GPT-4, 0.310.31 on GPT-4-Turbo, 0.480.48 on Ernie and LLaMA2-70B, and 0.640.64 on Claude-2.
    2. Gender-Biased Content Attacks: When gender-biased answers are perturbed with Ref+RC:
      • Ernie (ASR =0.69= 0.69) and LLaMA2-70B (ASR =0.82= 0.82) become more likely to favor the biased answer than a random choice baseline (0.650.65).
      • GPT-4 achieves an ASR of 0.500.50, and GPT-4-Turbo achieves 0.300.30 under Ref+RC.
      • Frontier models (GPT-4o and Claude-3) demonstrate the highest robustness but still exhibit non-zero vulnerability (GPT-4o Ref+RC ASR ≈0.10\approx 0.10 on gender bias; Claude-3 Ref+RC ASR ≈0.20\approx 0.20).
  6. Knowl 6 — Quality Gap Sensitivity in Weak Answer Turnover Attacks

    data/table

    Injecting fake citations into an inferior candidate answer to make it overturn a stronger answer (GPT-3.5-Turbo baseline) reveals that the success rate of the attack depends inversely on the semantic quality gap between the answers:

    Judge LM-7B LM-13B LM-70B GPT-3.5-Turbo Avg. Ranking ↓\downarrow
    GPT-4 0.04 0.07 0.09 0.40 2.25
    Ernie 0.07 0.10 0.11 0.24 2.75
    LLaMA2-70B 0.05 0.09 0.11 0.27 2.75
    PaLM-2 0.11 0.06 0.14 0.26 3.50
    GPT-4-Turbo 0.09 0.16 0.19 0.22 4.25
    Claude-3 0.09 0.15 0.18 0.55 5.25
    Claude-2 0.21 0.30 0.36 0.53 6.75

    When competing against a large quality disparity (e.g., LLaMA-2-7B vs. GPT-3.5-Turbo), fake citations yield low ASR (0.040.04--0.210.21). However, when the two competing answers have similar raw quality (e.g., GPT-3.5-Turbo vs. GPT-3.5-Turbo), ASR sharply increases across all judges, reaching 0.400.40 on GPT-4, 0.550.55 on Claude-3, and 0.530.53 on Claude-2.

  7. Knowl 7 — Positional Biases in Pairwise LLM Judgement

    empirical result

    Pairwise evaluation across LLMs exhibits pronounced positional preferences that can distort comparative assessments:

    • First-Position Preference: GPT-3.5-Turbo selects the first option 91.8%91.8\% of the time and the second option only 7.9%7.9\% (difference +84.0%+84.0\%). Mixtral-8x7B-Instruct selects the first option 64.6%64.6\% and the second option 32.0%32.0\% (difference +32.7%+32.7\%).
    • Second-Position Preference: Spark selects the second option 64.6%64.6\% and the first option 22.9%22.9\% (difference −41.7%-41.7\%).
    • Tie-Bias: Qwen-plus selects Tie 97.5%97.5\% of the time, and Gemini-1.0-Pro selects Tie 86.2%86.2\% of the time.
    • Relatively Balanced Evaluators: Human evaluators (36.9%36.9\% First vs. 36.3%36.3\% Second), GPT-4 (38.3%38.3\% First vs. 32.7%32.7\% Second), GPT-4-Turbo (21.1%21.1\% First vs. 14.9%14.9\% Second, 64.0%64.0\% Tie), Claude-2 (44.6%44.6\% First vs. 44.6%44.6\% Second), and PaLM-2 (51.1%51.1\% First vs. 48.4%48.4\% Second) keep position preference differences within 10%10\%.
  8. Knowl 8 — Invariance of Factual Error Detection to Self-Enhancement Bias

    data/table

    To test whether high factual error detection performance in proprietary models is driven by self-enhancement bias (favoring text generated by the same model architecture), cross-evaluation experiments were conducted where answers and factual error perturbations were generated independently by either GPT-4 or Claude-3:

    Judges Answer and Perturbation Generator
    GPT-4 Claude-3
    GPT-4 0.07 0.08
    Claude-3 0.10 0.08

    Both GPT-4 and Claude-3 maintain consistent ASRs regardless of whether the perturbed candidate answers originated from themselves or the other model (0.070.07 vs. 0.080.08 for GPT-4; 0.100.10 vs. 0.080.08 for Claude-3). This verifies that the strong factual error detection performance of top-tier LLMs reflects semantic verification ability rather than self-enhancement bias.

  9. Knowl 9 — Verbosity Bias in Human and LLM Judges

    empirical result

    Analysis of evaluator preferences as a function of the absolute token length difference between two candidate answers demonstrates a widespread verbosity bias:

    • When the token length difference is minimal (0--10 tokens), judge preferences remain close to neutral (0.500.50).
    • As the length difference increases, human evaluators and almost all LLM judges show a monotonic increase in preference toward the longer answer.
    • When the length difference exceeds 40 tokens, the preference score for longer answers reaches or exceeds 0.700.70 across humans and most models, with Claude-3 displaying the strongest verbosity bias.
    • Among all evaluated models, GPT-4-Turbo is the least influenced by response length across all length-difference bins.
  10. Knowl 10 — Limitations of the Judgement Bias Study

    limitation

    The methodology and conclusions of the study are bounded by several specific constraints:

    1. Benchmark Scale and Coverage: The evaluation set consists of 142 questions categorized vertically along revised Bloom's Taxonomy without exhaustive horizontal domain classifications. The questions are constrained to middle school and high school knowledge levels to allow reliable assessment by college evaluators.
    2. Evaluator Demographics: Human judges were sampled exclusively from university undergraduate students, whose high gender bias sensitivity (ASR =0.06= 0.06) may not reflect broader, non-academic human populations.
    3. Scope of Evaluated Biases: The study investigates Misinformation Oversight, Gender, Authority, Beauty, Positional, and Verbosity biases, omitting other dimensions such as linguistic register, syntactic complexity, tone, and racial bias.
    4. Temporal Generalizability: Fast iterations of proprietary and open-weight model architectures may alter the specific bias and vulnerability profiles of future LLM checkpoints.

Coverage note — None omitted. All primary contributions—including the reference-free intervention protocol, bias definitions, empirical benchmark results across the four biases, adversarial attack findings, positional and verbosity bias analyses, self-enhancement cross-checks, and limitations—are fully represented.

References

  1. 1.Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. 2019. Gradio: Hassle-free sharing and testing of ml models in the wild.
  2. 2.Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku.
  3. 3.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report.
  4. 4.Anya Belz, Craig Thomson, Ehud Reiter, Gavin Abercrombie, Jose M Alonso-Moral, Mohammad Arvan, Jackie Cheung, Mark Cieliebak, Elizabeth Clark, Kees van Deemter, et al. 2023. Missing information, unresponsive authors, experimental flaws: The impossibility of assessing the reproducibility of previous human evaluations in nlp. arXiv preprint arXiv:2305.01633.
  5. 5.Florian Böhm, Yang Gao, Christian M Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych. 2019. Better rewards yield better summaries: Learning to summarise without references. arXiv preprint arXiv:1909.01214.
  6. 6.Junying Chen, Xidong Wang, Anningzhe Gao, Feng Jiang, Shunian Chen, Hongbo Zhang, Dingjie Song, Wenya Xie, Chuyi Kong, Jianquan Li, Xiang Wan, Haizhou Li, and Benyou Wang. 2023a. Huatuogpt-ii, one-stage training for medical adaption of llms.
  7. 7.Zhihong Chen, Feng Jiang, Junying Chen, Tiannan Wang, Fei Yu, Guiming Chen, Hongbo Zhang, Juhao Liang, Chen Zhang, Zhiyi Zhang, Jianquan Li, Xiang Wan, Benyou Wang, and Haizhou Li. 2023b. Phoenix: Democratizing chatgpt across languages.
  8. 8.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  9. 9.Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2023. Deep reinforcement learning from human preferences.
  10. 10.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021. All that’s’ human’is not gold: Evaluating human evaluation of generated text. arXiv preprint arXiv:2107.00061.
  11. 11.Karen Czapanskiy. 1990. Gender bias in the courts: Social change strategies. Geo. J. Legal Ethics, 4:1.
  12. 12.Jacques Ellul. 2021. Propaganda: The formation of men’s attitudes. Vintage.
  13. 13.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire.
  14. 14.Michael J. Keough Gerry P. Quinn. 2002. Experimental Design and Data Analysis for Biologists, 1 edition. Cambridge University Press.
  15. 15.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19.
  16. 16.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  17. 17.David M Howcroft, Anya Belz, Miruna Clinciu, Dimitra Gkatzia, Sadid A Hasan, Saad Mahamood, Simon Mille, Emiel Van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020. Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definitions. In 13th International Conference on Natural Language Generation 2020, pages 169–182. Association for Computational Linguistics.
  18. 18.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural Information Processing Systems.
  19. 19.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts.
  20. 20.Shuyu Jiang, Xingshu Chen, and Rui Tang. 2023. Prompt packer: Deceiving llms through compositional instruction with hidden attacks.
  21. 21.Daniel Kahneman. 2011. Thinking, fast and slow. macmillan.
  22. 22.Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023. Benchmarking cognitive biases in large language models as evaluators.
  23. 23.David R Krathwohl. 2002. A revision of bloom’s taxonomy: An overview. Theory into practice, 41(4):212–218.
  24. 24.Julia Kreutzer, Shahram Khadivi, Evgeny Matusov, and Stefan Riezler. 2018. Can neural machine translation be improved with user feedback?
  25. 25.Judith H Langlois, Lisa Kalakanis, Adam J Rubenstein, Andrea Larson, Monica Hallam, and Monica Smoot. 2000. Maxims or myths of beauty? a meta-analytic and theoretical review. Psychological bulletin, 126(3):390.
  26. 26.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871.
  27. 27.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  28. 28.Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2024. Llms as narcissistic evaluators: When ego inflates evaluation scores.
  29. 29.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.
  30. 30.OpenAI, :, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Vik Goel, Tarun Gogineni, Gabriel Goh, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2023. Gpt-4 technical report.
  31. 31.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  32. 32.Oxford English Dictionary. 2023. semantic, adj., sense 2.a. Oxford English Dictionary. Accessed: 2023-11-13.
  33. 33.Nicholas Pangakis, Samuel Wolken, and Neil Fasching. 2023. Automated annotation with generative ai requires validation. arXiv preprint arXiv:2306.00176.
  34. 34.Vyas Raina, Adian Liusie, and Mark Gales. 2024. Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment.
  35. 35.Lise Saffran, Sisi Hu, Amanda Hinnant, Laura D. Scherer, and Susan C. Nagel. 2020. Constructing and influencing perceived authenticity in science communication: Experimenting with narrative. PLOS ONE, 15(1):1–17.
  36. 36.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models.
  37. 37.Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024. Optimization-based prompt injection attack to llm-as-a-judge.
  38. 38.Irene van Staveren. 2001. Gender biases in finance. Gender & Development, 9(1):9–17.
  39. 39.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  40. 40.Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, Weixin Liu, Zhihua Wu, Weibao Gong, Jianzhong Liang, Zhizhou Shang, Peng Sun, Wei Liu, Xuan Ouyang, Dianhai Yu, Hao Tian, Hua Wu, and Haifeng Wang. 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation.
  41. 41.Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul R. Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, Jack Krawczyk, Cosmo Du, Ed Chi, Heng-Tze Cheng, Eric Ni, Purvi Shah, Patrick Kane, Betty Chan, Manaal Faruqui, Aliaksei Severyn, Hanzhao Lin, YaGuang Li, Yong Cheng, Abe Ittycheriah, Mahdis Mahdieh, Mia Chen, Pei Sun, Dustin Tran, Sumit Bagri, Balaji Lakshminarayanan, Jeremiah Liu, Andras Orban, Fabian Güra, Hao Zhou, Xinying Song, Aurelien Boffy, Harish Ganapathy, Steven Zheng, HyunJeong Choe, Ágoston Weisz, Tao Zhu, Yifeng Lu, Siddharth Gopal, Jarrod Kahn, Maciej Kula, Jeff Pitman, Rushin Shah, Emanuel Taropa, Majd Al Merey, Martin Baeuml, Zhifeng Chen, Laurent El Shafey, Yujing Zhang, Olcan Sercinoglu, George Tucker, Enrique Piqueras, Maxim Krikun, Iain Barr, Nikolay Savinov, Ivo Danihelka, Becca Roelofs, Anaïs White, Anders Andreassen, Tamara von Glehn, Lakshman Yagati, Mehran Kazemi, Lucas Gonzalez, Misha Khalman, Jakub Sygnowski, Alexandre Frechette, Charlotte Smith, Laura Culp, Lev Proleev, Yi Luan, Xi Chen, James Lottes, Nathan Schucher, Federico Lebron, Alban Rrustemi, Natalie Clay, Phil Crone, Tomas Kocisky, Jeffrey Zhao, Bartek Perz, Dian Yu, Heidi Howard, Adam Bloniarz, Adam Austin, Gabriel Barth-Maron, William Wong, Rishabh Joshi, and Oriol Vinyals. 2024. Gemini: A family of highly capable multimodal models.
  42. 42.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  43. 43.Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023a. Large language models are not fair evaluators.
  44. 44.Xidong Wang, Guiming Hardy Chen, Dingjie Song, Zhiyi Zhang, Zhihong Chen, Qingying Xiao, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, et al. 2023b. Cmb: A comprehensive medical benchmark in chinese. arXiv preprint arXiv:2308.08833.
  45. 45.Kelly Weidner, Frederik Beuk, and Anjali Bal. 2020. Fake news and the willingness to share: a schemer schema and confirmatory bias perspective. Journal of Product & Brand Management, 29(2):180–187.
  46. 46.Minghao Wu and Alham Fikri Aji. 2023. Style over substance: Evaluation biases for large language models. arXiv preprint arXiv:2307.03025.
  47. 47.Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Yang Wang. 2024. Perils of self-feedback: Self-bias amplifies in large language models.
  48. 48.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation.
  49. 49.Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023. Evaluating large language models at evaluating instruction following.
  50. 50.Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, and Haizhou Li. 2023. Huatuogpt, towards taming language model to be a doctor.
  51. 51.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert.
  52. 52.Wei Zhao, Michael Strube, and Steffen Eger. 2023. Discoscore: Evaluating text generation with bert and discourse coherence.
  53. 53.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  54. 54.Wangchunshu Zhou and Ke Xu. 2020. Learning to compare for better training and evaluation of open domain natural language generation models. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, pages 9717–9724.
  55. 55.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.
  56. 56.Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models.

Citation

MLA
Chen, H., et al. “Humans or LLMs as the Judge? A Study on Judgement Bias”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8301–27, https://doi.org/10.18653/v1/2024.emnlp-main.474.
APA
Chen, H., Chen, S., Liu, Z., (蒋峰), F. J., & Wang, B. (2024). Humans or LLMs as the Judge? A Study on Judgement Bias. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301–8327. https://doi.org/10.18653/v1/2024.emnlp-main.474
Chicago
Chen, H., S. Chen, Z. Liu, F. J. (蒋峰), and B. Wang. 2024. “Humans or LLMs as the Judge? A Study on Judgement Bias”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301–27. https://doi.org/10.18653/v1/2024.emnlp-main.474.
Harvard
Chen, H. et al. (2024) “Humans or LLMs as the Judge? A Study on Judgement Bias”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8301–8327. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.474.
Vancouver
1. Chen H, Chen S, Liu Z, (蒋峰) FJ, Wang B (2024) Humans or LLMs as the Judge? A Study on Judgement Bias. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8301–8327

BibTeX

@inproceedings{chen-etal-2024-humans,
    title = "Humans or {LLM}s as the Judge? A Study on Judgement Bias",
    author = "Chen, Guiming Hardy  and
      Chen, Shunian  and
      Liu, Ziche  and
      Jiang, Feng  and
      Wang, Benyou",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.474/",
    doi = "10.18653/v1/2024.emnlp-main.474",
    pages = "8301--8327"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/