Large Language Models are not Fair Evaluators

Peiyi WangLei LiLiang ChenZefan CaiDawei ZhuBinghuai LinYunbo CaoLingpeng KongQi LiuTianyu Liu

article2024ACL1,334 citations

Reveals that using large language models as judges introduces severe positional bias that distorts model rankings, and provides effective calibration strategies to align automated evaluations with human judgments.

Listen

As organizations increasingly deploy artificial intelligence systems, evaluating model quality reliably and cost-effectively has become critical. Relying on large language models as automated judges has become a popular alternative to expensive and slow human evaluation. However, the reliability and fairness of using these models to judge and compare competing outputs remain uncertain.

The article evaluates whether advanced language models suffer from positional bias when judging candidate responses. It demonstrates a calibration framework designed to correct these distortions and improve alignment with human judgments.

To examine this issue, the researchers tested leading models acting as evaluators on a standard benchmark across 80 questions spanning nine categories. They analyzed how swapping the presentation order of two competing responses affected the scoring. They then introduced and tested three calibration strategies: generating written evaluation evidence before scoring, balancing positions by swapping candidate order and averaging the scores, and using an entropy-based scoring metric to identify the most uncertain comparisons for targeted human review. The authors established a ground-truth baseline by manually annotating all 80 benchmark questions independently.

The investigation revealed substantial positional bias in automated judges. When comparing responses, simply swapping the presentation order caused GPT-4 to produce conflicting results in 46.3% of test cases, while ChatGPT produced conflicting results in 82.5% of cases. GPT-4 systematically favored the response shown first, whereas ChatGPT favored the response shown second. This vulnerability was especially acute when the quality gap between responses was small. Applying the automated calibration techniques improved judgment accuracy relative to human consensus by 9.8 percentage points for GPT-4 and 14.3 percentage points for ChatGPT. Furthermore, selectively routing just the top 20% most uncertain evaluations to human reviewers enabled the system to match or exceed average human evaluation accuracy while reducing human annotation costs by up to 39%.

These findings indicate that off-the-shelf language models cannot be trusted as fair, uncalibrated judges, posing significant risks of skewed model selection, flawed safety assessments, and misleading performance tracking. The common practice of asking a model to provide a numerical score before giving an explanation exacerbates these distortions. By forcing models to articulate evaluation reasoning first and averaging scores across swapped positions, organizations can significantly enhance evaluation robustness at low computational cost.

Decision-makers using automated model evaluation should immediately stop relying on single-pass, uncalibrated scoring prompts. Instead, organizations should adopt multi-sample evidence generation and position swapping in automated pipelines. When high accuracy is critical, teams should implement human-in-the-loop workflows that use diversity entropy metrics to flag ambiguous cases for manual review rather than conducting blanket human evaluations.

The findings are bounded by the 80-question test set and the specific model versions evaluated. While the proposed calibration toolkit consistently improves alignment with human judgments, the underlying root causes of positional bias inside language models remain unaddressed and require further investigation.

arXiv: 2305.17926i-Eval/FairEval
Cover for Large Language Models are not Fair Evaluators

Abstract

In this paper, we uncover a positional bias in the evaluation paradigm of adopting large language models (LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models. We find that the quality ranking of candidate responses can be easily hacked by simply altering their order of appearance in the context. This manipulation allows us to skew the evaluation result, making one model appear considerably superior to the other, e.g., Vicuna-13B could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator. We propose a simple yet effective calibration framework to address our discovered positional bias. To evaluate the effectiveness of our framework, we manually annotate the “win/tie/lose” outcomes of responses from ChatGPT and Vicuna-13B in the Vicuna Benchmark’s question prompt. Extensive experiments demonstrate that our approach successfully alleviates evaluation bias, resulting in closer alignment with human judgments. To facilitate future research on more robust large language model comparison, we integrate the techniques in the paper into an easy-to-use toolkit FairEval, along with the human annotations 1.

Table of Contents

  • 1 Introduction
  • 2 Positional Bias of the LLM Evaluator
  • 2.1 LLMs as Evaluators
  • 2.2 Revealing the Positional Bias
  • 2.3 Calibrating the Positional Bias
  • 2.3.1 Multiple Evidence Calibration
  • 2.3.2 Balanced Position Calibration
  • 2.3.3 Human-in-the-Loop Calibration
  • 3 Experiments
  • 3.1 Human Annotation
  • 3.2 Experimental Setup and Metric
  • 3.3 Main Results
  • 4 Analysis
  • 4.1 Ablation on Evidence Number k and Temperature t
  • 4.2 Effectiveness of the BPDE
  • 4.3 Generalization on the Pairwise Comparison Evaluation Template
  • 4.4 Fine-Grained Analysis of Evaluation Quality
  • 5 Related Work
  • 6 Conclusion
  • Limitation
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Multiple Evidence and Balanced Position Calibration for LLM Evaluators

    model/method

    To alleviate positional bias in large language models (LLMs) used as pairwise evaluators, two complementary automated calibration strategies are combined: Multiple Evidence Calibration (MEC) and Balanced Position Calibration (BPC).

    1. Evidence Calibration (EC / MEC): Standard evaluation templates ask the LLM for a conclusion/score first and then an explanation, meaning the auto-regressive model generates scores without conditioning on its justification. Evidence Calibration reformulates the prompt TEC(Q,R1,R2)T_{EC}(Q, R_1, R_2) to force the LLM to output the comprehensive evaluation explanation (evidence) prior to outputting numerical scores. To enhance stability, Multiple Evidence Calibration (MEC) generates kk independent evaluation samples at temperature t>0t > 0, producing scores: {Sr11,…,Sr1k}and{S′r21,…,S′r2k}\{S_{r_1}^1, \dots, S_{r_1}^k\} \quad \text{and} \quad \{S'{}_{r_2}^1, \dots, S'{}_{r_2}^k\} where SriS_r^i denotes the score assigned to response rr when presented in the first slot, and S′riS'{}_r^i denotes the score assigned to response rr when presented in the second slot.

    2. Balanced Position Calibration (BPC): BPC evaluates candidate responses in both presentation orders. Given query qq and responses r1,r2r_1, r_2, BPC queries the LLM with both TEC(q,r1,r2)T_{EC}(q, r_1, r_2) and TEC(q,r2,r1)T_{EC}(q, r_2, r_1), yielding 2k2k total sampled scores per response: {Sr11,…,Sr1k,S′r11,…,S′r1k}\{S_{r_1}^1, \dots, S_{r_1}^k, S'{}_{r_1}^1, \dots, S'{}_{r_1}^k\} and {S′r21,…,S′r2k,Sr21,…,Sr2k}\{S'{}_{r_2}^1, \dots, S'{}_{r_2}^k, S_{r_2}^1, \dots, S_{r_2}^k\}.

    The final calibrated score CSRCS_R for each candidate response R∈{r1,r2}R \in \{r_1, r_2\} is computed as the arithmetic mean across all 2k2k position-balanced samples: CSR=∑i=1k(SRi+SR′i)2k,R∈{r1,r2}CS_R = \frac{\sum_{i=1}^k \left( S_R^i + {S'_R}^i \right)}{2k}, \quad R \in \{r_1, r_2\} The response with the strictly higher calibrated score is declared the winner (or a tie if CSr1=CSr2CS_{r_1} = CS_{r_2}).

  2. Knowl 2 — Balanced Position Diversity Entropy for Detecting Evaluation Bias

    equation

    Balanced Position Diversity Entropy (BPDE) quantifies evaluation uncertainty and positional inconsistency across multiple position-swapped LLM evaluation runs, identifying instances that require manual human review.

    Given kk sampled evaluation runs for prompt order T(q,r1,r2)T(q, r_1, r_2) and kk sampled runs for reversed order T(q,r2,r1)T(q, r_2, r_1), paired outcomes ERiER_i and ERi′ER'_i (1≤i≤k1 \le i \le k) are defined as: ERi={win,Sr1i>S′r2itie,Sr1i=S′r2ilose,Sr1i<S′r2i,ERi′={win,S′r1i>Sr2itie,S′r1i=Sr2ilose,S′r1i<Sr2iER_i = \begin{cases} \text{win}, & S_{r_1}^i > S'{}_{r_2}^i \\ \text{tie}, & S_{r_1}^i = S'{}_{r_2}^i \\ \text{lose}, & S_{r_1}^i < S'{}_{r_2}^i \end{cases}, \qquad ER'_i = \begin{cases} \text{win}, & S'{}_{r_1}^i > S_{r_2}^i \\ \text{tie}, & S'{}_{r_1}^i = S_{r_2}^i \\ \text{lose}, & S'{}_{r_1}^i < S_{r_2}^i \end{cases} where Sr1iS_{r_1}^i and S′r2iS'{}_{r_2}^i are the scores for r1r_1 in the first slot and r2r_2 in the second slot from sample ii, while S′r1iS'{}_{r_1}^i and Sr2iS_{r_2}^i are the scores for r1r_1 in the second slot and r2r_2 in the first slot.

    The empirical probability distribution perp_{er} over the possible comparison outcomes er∈{win,tie,lose}er \in \{\text{win}, \text{tie}, \text{lose}\} across all 2k2k evaluations is: per=∑i=1k(I(ERi=er)+I(ERi′=er))2kp_{er} = \frac{\sum_{i=1}^k \left( \mathbb{I}(ER_i = er) + \mathbb{I}(ER'_i = er) \right)}{2k} where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    The Balanced Position Diversity Entropy (BPDE) is then defined as: BPDE=∑er∈{win,tie,lose}−perlog⁡per\text{BPDE} = \sum_{er \in \{\text{win}, \text{tie}, \text{lose}\}} -p_{er} \log p_{er} where higher BPDE indicates greater output inconsistency across samples and positions, signaling that the evaluation is prone to bias and should be prioritized for human-in-the-loop annotation.

  3. Knowl 3 — Conflict Rate Metric for Positional Sensitivity in LLM Evaluators

    definition

    The Conflict Rate measures the sensitivity of a pairwise LLM evaluator to the presentation order of candidate responses. Given a dataset of NN evaluation instances {(qi,r1i,r2i)}i=1N\{(q_i, r_{1i}, r_{2i})\}_{i=1}^N, the evaluator is prompted twice per query: first with prompt T(qi,r1i,r2i)T(q_i, r_{1i}, r_{2i}) yielding discrete evaluation outcome ERir12ER_i^{r_{12}}, and second with the reversed candidate order T(qi,r2i,r1i)T(q_i, r_{2i}, r_{1i}) yielding outcome ERir21ER_i^{r_{21}} (mapped back to the relative win/tie/lose status of r1r_1 versus r2r_2).

    The Conflict Rate is defined as: Conflict Rate=∑i=1NI(ERir12≠ERir21)N\text{Conflict Rate} = \frac{\sum_{i=1}^N \mathbb{I}\left(ER_i^{r_{12}} \neq ER_i^{r_{21}}\right)}{N} where I(⋅)\mathbb{I}(\cdot) is the indicator function. A non-zero Conflict Rate indicates self-contradiction induced solely by swapping response positions in the prompt.

  4. Knowl 4 — Human-in-the-Loop Calibration (HITLC) Algorithm

    algorithm

    Human-in-the-Loop Calibration (HITLC) combines automated Multiple Evidence Calibration (MEC) and Balanced Position Calibration (BPC) with selective human annotations based on Balanced Position Diversity Entropy (BPDE).

    Input: Dataset of queries and response pairs D = {(q_j, r_1j, r_2j)}_{j=1}^N, evidence sample count k, selection ratio beta in (0, 1]
    Output: Set of final evaluation judgments J
    for each instance j in 1 to N do
        Sample k scores {S_{r_1}^i, S'{}_{r_2}^i}_{i=1}^k using prompt T_EC(q_j, r_1j, r_2j)
        Sample k scores {S_{r_2}^i, S'{}_{r_1}^i}_{i=1}^k using prompt T_EC(q_j, r_2j, r_1j)
        Compute calibrated scores CS_{r_1j} and CS_{r_2j}
        Compute outcome probabilities p_{er} for er in {win, tie, lose}
        Compute entropy score BPDE_j = sum_{er} -p_{er} * log(p_{er})
    end for
    Rank instances in D by BPDE_j in descending order
    Select top ceil(beta * N) instances as D_human
    Set D_auto = D \ D_human
    for each instance j in D_human do
        Obtain majority vote human judgment J_j
    end for
    for each instance j in D_auto do
        if CS_{r_1j} > CS_{r_2j} then
            J_j = win
        else if CS_{r_1j} < CS_{r_2j} then
            J_j = lose
        else
            J_j = tie
        end if
    end for
    return J = {J_j}_{j=1}^N
  5. Knowl 5 — Empirical Positional Bias Discrepancies in GPT-4 and ChatGPT Evaluators

    empirical result

    Pairwise evaluation across 80 questions from the Vicuna benchmark demonstrates that LLM evaluators suffer from strong, directional positional bias:

    1. Primacy Bias in GPT-4: When evaluating Vicuna-13B against ChatGPT, GPT-4 assigns Vicuna-13B a win rate of 51.3%51.3\% when it is positioned as Assistant 1, but only 23.8%23.8\% when positioned as Assistant 2, resulting in a Conflict Rate of 46.3%46.3\% (37/80 questions).
    2. Recency Bias in ChatGPT: ChatGPT as an evaluator exhibits a preference for the second position; Vicuna-13B achieves a win rate of only 2.5%2.5\% when placed as Assistant 1 versus 82.5%82.5\% when placed as Assistant 2, yielding a Conflict Rate of 82.5%82.5\% (66/80 questions).
    3. Model Capability and Gap Dependency: On a benchmark comparison with a larger quality gap (Vicuna-13B vs. Alpaca-13B), GPT-4's conflict rate drops to 5.0%5.0\% (92.5%92.5\% win rate in both positions), whereas ChatGPT maintains a high conflict rate of 52.5%52.5\% (37.5%37.5\% as Assistant 1 vs. 90.0%90.0\% as Assistant 2).
  6. Knowl 6 — Evaluation Alignment and Cost Across Calibration Strategies on Vicuna Benchmark

    data/table

    The table below compares the accuracy, Cohen's kappa coefficient (relative to the majority vote of three human experts), and monetary evaluation costs across baseline and calibrated methods on the 80 Vicuna benchmark queries comparing Vicuna-13B and ChatGPT.

    Evaluator Method Accuracy Kappa Cost
    Human 1 - 68.8% 0.50 $30.0
    Human 2 - 76.3% 0.62 $30.0
    Human 3 - 70.0% 0.50 $30.0
    Human Average - 71.7% 0.54 $30.0
    GPT-4 VANILLA 52.7% 0.24 $2.00
    GPT-4 EC (k=1k=1) 56.5% 0.29 $2.00
    GPT-4 MEC (k=3k=3) 58.7% 0.30 $3.19
    GPT-4 MEC (k=6k=6) 60.9% 0.33 $4.98
    GPT-4 MEC (k=3k=3) + BPC (k=3k=3) 62.5% 0.37 $6.38
    GPT-4 MEC (k=3k=3) + BPC (k=3k=3) + HITLC (β=20%\beta=20\%) 73.8% 0.56 $23.1
    ChatGPT VANILLA 44.4% 0.06 $0.10
    ChatGPT EC (k=1k=1) 52.6% 0.23 $0.10
    ChatGPT MEC (k=3k=3) 53.2% 0.24 $0.17
    ChatGPT MEC (k=6k=6) 55.6% 0.27 $0.28
    ChatGPT MEC (k=3k=3) + BPC (k=3k=3) 58.8% 0.31 $0.34
    ChatGPT MEC (k=3k=3) + BPC (k=3k=3) + HITLC (β=20%\beta=20\%) 71.3% 0.52 $18.3

    MEC (k=3k=3) + BPC (k=3k=3) improves GPT-4 accuracy by 9.8%9.8\% (52.7%→62.5%52.7\% \to 62.5\%) and ChatGPT accuracy by 14.4%14.4\% (44.4%→58.8%44.4\% \to 58.8\%) over VANILLA. Incorporating HITLC with β=20%\beta=20\% human annotation allows ChatGPT to achieve human-level accuracy (71.3%71.3\% vs. human average 71.7%71.7\%) while reducing annotation costs from $30.0 to $18.3 (a 39%39\% cost reduction).

  7. Knowl 7 — Negative Correlation Between Score Gap and Evaluator Conflict Rate

    empirical result

    The susceptibility of LLM evaluators (such as GPT-4) to positional swapping is inversely correlated with the intrinsic quality difference between the two candidate responses:

    • When the score difference between candidate responses is small (score gap ≤1\le 1 on a 1-to-10 scale), the evaluator produces conflicting results on the majority of cases when input positions are reversed.
    • When the score difference is moderate to large (score gap ≥3\ge 3), the evaluator's judgments remain relatively stable and consistent regardless of candidate ordering.
  8. Knowl 8 — Calibration Robustness Across Scoring and Direct Comparing Templates

    empirical result

    Pairwise evaluation can be formulated using either a Scoring template (assigning absolute scores 1--10 to each candidate before comparison) or a Comparing template (directly classifying the winner as 'Assistant 1', 'Assistant 2', or 'Same').

    On the Vicuna benchmark with ChatGPT:

    1. Scoring Template: Vanilla baseline achieves 44.4%44.4\% accuracy, 0.060.06 kappa, and an 82.5%82.5\% conflict rate. Applying MEC achieves 53.2%53.2\% accuracy, 0.240.24 kappa, and a 35.0%35.0\% conflict rate. Applying MEC + BPC achieves 58.8%58.8\% accuracy and 0.310.31 kappa.
    2. Comparing Template: Vanilla baseline achieves 50.2%50.2\% accuracy, 0.180.18 kappa, and a 50.0%50.0\% conflict rate. Applying MEC achieves 54.8%54.8\% accuracy, 0.270.27 kappa, and a 42.5%42.5\% conflict rate. Applying MEC + BPC achieves 60.0%60.0\% accuracy and 0.350.35 kappa.

    The calibration strategies narrow the cross-template accuracy gap (from 5.8%5.8\% in Vanilla to 1.2%1.2\% in MEC + BPC) and reduce conflict rates.

  9. Knowl 9 — Influence of Evidence Sample Count and Sampling Temperature on MEC

    empirical result

    In Multiple Evidence Calibration (MEC), the number of sampled evidence traces kk and the decoding temperature tt directly modulate evaluation accuracy and inter-annotator kappa agreement:

    • Sample Count (kk): As kk increases from 11 to 33, accuracy and kappa improve sharply. For k∈{3,5,7}k \in \{3, 5, 7\}, performance gains plateau or decrease slightly while increasing API computation cost proportionally. k=3k=3 represents the optimal trade-off between evaluation performance and inference cost.
    • Sampling Temperature (tt): Both low temperatures (e.g., t=0.2t = 0.2) and high temperatures (e.g., t=1.4t = 1.4) lead to sub-optimal human alignment. Low temperature removes the stochastic diversity necessary for ensemble calibration, while excessively high temperature impairs the reasoning quality of the generated evaluation evidence. Moderate temperatures (t=0.6t = 0.6 to 1.01.0) yield optimal performance.
  10. Knowl 10 — Lack of Mechanistic Explanation for Evaluator Positional Bias

    limitation

    While positional bias in LLM evaluators (e.g., primacy bias in GPT-4 and recency bias in ChatGPT) is empirically demonstrated and mitigated through calibration frameworks, the underlying causal mechanisms—such as whether this bias arises from pre-training data distributions, attention structure artifacts, or reinforcement learning from human feedback (RLHF) recipes—remain uninvestigated.

Coverage note — None was omitted; all key contributions, algorithms, theoretical metrics (Conflict Rate, BPDE), experimental results across models/templates, parameter ablations, and stated limitations have been covered.

References

  1. 1.Belinkov, Y.; Poliak, A.; Shieber, S.; Van Durme, B.; and Rush, A. 2019. Don’t Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
  2. 2.Bowman, S. R. 2023. Eight things to know about large language models. arXiv preprint arXiv:2304.00612.
  3. 3.Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020. Language Models are Few-Shot Learners. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  4. 4.Cai, Z.; Tu, L.; and Gimpel, K. 2017. Pay Attention to the Ending:Strong Neural Baselines for the ROC Story Cloze Task. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL).
  5. 5.Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; Schuh, P.; Shi, K.; Tsvyashchenko, S.; Maynez, J.; Rao, A.; Barnes, P.; Tay, Y.; Shazeer, N. M.; Prabhakaran, V.; Reif, E.; Du, N.; Hutchinson, B. C.; Pope, R.; Bradbury, J.; Austin, J.; Isard, M.; Gur-Ari, G.; Yin, P.; Duke, T.; Levskaya, A.; Ghemawat, S.; Dev, S.; Michalewski, H.; García, X.; Misra, V.; Robinson, K.; Fedus, L.; Zhou, D.; Ippolito, D.; Luan, D.; Lim, H.; Zoph, B.; Spiridonov, A.; Sepassi, R.; Dohan, D.; Agrawal, S.; Omernick, M.; Dai, A. M.; Pillai, T. S.; Pellat, M.; Lewkowycz, A.; Moreira, E.; Child, R.; Polozov, O.; Lee, K.; Zhou, Z.; Wang, X.; Saeta, B.; Díaz, M.; Firat, O.; Catasta, M.; Wei, J.; Meier-Hellstern, K. S.; Eck, D.; Dean, J.; Petrov, S.; and Fiedel, N. 2022. PaLM: Scaling Language Modeling with Pathways. ArXiv, abs/2204.02311.
  6. 6.Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; and Sui, Z. 2022. A Survey for In-context Learning. arXiv preprint arXiv:2301.00234.
  7. 7.Dubois, Y.; Li, X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback. arXiv preprint arXiv:2305.14387.
  8. 8.Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; Li, H.; and Qiao, Y. J. 2023. LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model. ArXiv, abs/2304.15010.
  9. 9.Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  10. 10.Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S.; and Smith, N. A. 2018. Annotation Artifacts in Natural Language Inference Data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics.
  11. 11.He, T.; Zhang, J.; Wang, T.; Kumar, S.; Cho, K.; Glass, J.; and Tsvetkov, Y. 2023. On the Blind Spots of Model-Based Evaluation Metrics for Text Generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12067–12097. Toronto, Canada: Association for Computational Linguistics.
  12. 12.Levy, O.; Remus, S.; Biemann, C.; and Dagan, I. 2015. Do Supervised Distributional Methods Really Learn Lexical Inference Relations? In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics.
  13. 13.Li, L.; Yin, Y.; Li, S.; Chen, L.; Wang, P.; Ren, S.; Li, M.; Yang, Y.; Xu, J.; Sun, X.; et al. 2023. M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning. arXiv preprint arXiv:2306.04387.
  14. 14.Lin, C.-Y. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, 74–81. Barcelona, Spain: Association for Computational Linguistics.
  15. 15.Liu, T.; Xin, Z.; Chang, B.; and Sui, Z. 2020a. HypoNLI: Exploring the Artificial Patterns of Hypothesis-only Bias in Natural Language Inference. In Proceedings of the 12th Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association. ISBN 979-10-95546-34-4.
  16. 16.Liu, T.; Xin, Z.; Ding, X.; Chang, B.; and Sui, Z. 2020b. An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language Inference. In Proceedings of the 24th Conference on Computational Natural Language Learning. Online: Association for Computational Linguistics.
  17. 17.Lu, Q.; Qiu, B.; Ding, L.; Xie, L.; and Tao, D. 2023. Error analysis prompting enables human-like translation evaluation in large language models: A case study on chatgpt. arXiv preprint arXiv:2303.13809.
  18. 18.McCoy, T.; Pavlick, E.; and Linzen, T. 2019. Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
  19. 19.McHugh, M. L. 2012. Interrater reliability: the kappa statistic. Biochemia medica, 22(3): 276–282.
  20. 20.Min, S.; Wallace, E.; Singh, S.; Gardner, M.; Hajishirzi, H.; and Zettlemoyer, L. 2019. Compositional Questions Do Not Necessitate Multi-hop Reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
  21. 21.OpenAI. 2022. Introducing ChatGPT.
  22. 22.OpenAI. 2023. GPT-4 Technical Report. CoRR, abs/2303.08774.
  23. 23.Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–318. Philadelphia, Pennsylvania, USA: Association for Computational Linguistics.
  24. 24.Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023. Instruction Tuning with GPT-4. ArXiv, abs/2304.03277.
  25. 25.Schwartz, R.; Sap, M.; Konstas, I.; Zilles, L.; Choi, Y.; and Smith, N. A. 2017. The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL).
  26. 26.Song, Y.; Wang, P.; Zhu, D.; Liu, T.; Sui, Z.; and Li, S. 2023. RepCL: Exploring Effective Representation for Continual Text Classification. arXiv preprint arXiv:2305.07289.
  27. 27.Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D. D.; Yang, Y.; and Gan, C. 2023. Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision. ArXiv, abs/2305.03047.
  28. 28.Turpin, M.; Michael, J.; Perez, E.; and Bowman, S. R. 2023. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. CoRR, abs/2305.04388.
  29. 29.Wang, P.; Song, Y.; Liu, T.; Lin, B.; Cao, Y.; Li, S.; and Sui, Z. 2022. Learning Robust Representations for Continual Relation Extraction via Adversarial Class Augmentation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6264–6278. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics.
  30. 30.Wang, P.; Xun, R.; Liu, T.; Dai, D.; Chang, B.; and Sui, Z. 2021. Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event Classification. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 1969–1978.
  31. 31.Wang, Y.; Ivison, H.; Dasigi, P.; Hessel, J.; Khot, T.; Chandu, K. R.; Wadden, D.; MacMillan, K.; Smith, N. A.; Beltagy, I.; et al. 2023a. How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources. arXiv preprint arXiv:2306.04751.
  32. 32.Wang, Y.; Yu, Z.; Zeng, Z.; Yang, L.; Wang, C.; Chen, H.; Jiang, C.; Xie, R.; Wang, J.; Xie, X.; et al. 2023b. PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization. arXiv preprint arXiv:2306.05087.
  33. 33.Xia, H.; Wang, P.; Liu, T.; Lin, B.; Cao, Y.; and Sui, Z. 2023. Enhancing Continual Relation Extraction via Classifier Decomposition. In Findings of the Association for Computational Linguistics: ACL 2023, 10053–10062. Toronto, Canada: Association for Computational Linguistics.
  34. 34.Xu, C.; Guo, D.; Duan, N.; and McAuley, J. 2023. Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data. ArXiv, abs/2304.01196.
  35. 35.Yuan, W.; Neubig, G.; and Liu, P. 2021. BARTScore: Evaluating Generated Text as Text Generation. In Ranzato, M.; Beygelzimer, A.; Dauphin, Y. N.; Liang, P.; and Vaughan, J. W., eds., Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, 27263–27277.
  36. 36.Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020. BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  37. 37.Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv preprint arXiv:2306.05685.
  38. 38.Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023. LIMA: Less Is More for Alignment.

Citation

MLA
(王培懿), P. W., et al. “Large Language Models Are Not Fair Evaluators”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 9440–50, https://doi.org/10.18653/v1/2024.acl-long.511.
APA
(王培懿), P. W., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Kong, L., Liu, Q., Liu, T., & Sui, Z. (2024). Large Language Models are not Fair Evaluators. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9440–9450. https://doi.org/10.18653/v1/2024.acl-long.511
Chicago
(王培懿), P. W., L. Li, L. Chen, et al. 2024. “Large Language Models Are Not Fair Evaluators”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9440–50. https://doi.org/10.18653/v1/2024.acl-long.511.
Harvard
(王培懿), P.W. et al. (2024) “Large Language Models are not Fair Evaluators”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9440–9450. Available at: https://doi.org/10.18653/v1/2024.acl-long.511.
Vancouver
1. (王培懿) PW, Li L, Chen L, et al (2024) Large Language Models are not Fair Evaluators. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9440–9450

BibTeX

@inproceedings{wang-etal-2024-large-language-models-fair,
    title = "Large Language Models are not Fair Evaluators",
    author = "Wang, Peiyi  and
      Li, Lei  and
      Chen, Liang  and
      Cai, Zefan  and
      Zhu, Dawei  and
      Lin, Binghuai  and
      Cao, Yunbo  and
      Kong, Lingpeng  and
      Liu, Qi  and
      Liu, Tianyu  and
      Sui, Zhifang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.511/",
    doi = "10.18653/v1/2024.acl-long.511",
    pages = "9440--9450"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/