Split and Merge: Aligning Position Biases in LLM-based Evaluators

Zongjie LiChaozheng WangPingchuan MaDaoyuan WuShuai WangCuiyun GaoYang Liu

article2024EMNLP106 citations

Proposes PORTIA, a split-and-merge prompting framework that mitigates pairwise position bias in large language model evaluators by segmenting and aligning candidate answers, enabling cost-effective models like GPT-3.5 to rival or exceed standalone GPT-4 in human agreement.

Listen

Organizations increasingly deploy large language models as automated evaluators to assess artificial intelligence outputs. While this practice is faster and less costly than human review, automated evaluators suffer from severe position bias during pairwise comparisons. Specifically, models often favor whichever response appears first or second regardless of actual content quality. Relying on flawed automated judgments introduces operational risks and skews model benchmarking, while relying solely on premium evaluators or human reviewers creates unsustainable costs and scaling bottlenecks.

The article evaluates a lightweight framework called PORTIA, which is designed to calibrate position bias and improve consistency in automated evaluators without modifying model weights. To validate this framework, the authors conducted an extensive empirical study across six distinct language models and 11,520 pairwise comparisons across three comparison formats (score-based, rating scales, and direct relational preferences). The evaluation was conducted primarily on the MT-Bench benchmark spanning diverse subject domains, alongside a 640-question expansion and a five-expert human alignment study.

The findings demonstrate substantial improvements across multiple operational and quality metrics. First, applying the alignment framework yielded an average relative improvement of 47.46% in evaluation consistency across all tested models, correcting an average of 62.31% of previously inconsistent judgments. Second, it raised the consistency rate of top-tier models like GPT-4 up to 98% and mitigated 36% to 86% of its position bias occurrences. Third, the framework enabled lower-cost models like GPT-3.5 to achieve an 88% agreement rate with GPT-4 while consuming less than 10% (specifically 9.57%) of the cost. Finally, human evaluation revealed that GPT-3.5 enhanced by this framework surpassed the baseline, standalone GPT-4 in agreement with human expert consensus (63.75% versus 60.00%).

These results demonstrate that position bias can be mitigated through structured prompting and content alignment rather than expensive model fine-tuning. By bridging the performance gap between mid-tier and frontier models, organizations can drastically cut the financial, temporal, and computational expenditures required for automated evaluation. Furthermore, the framework makes smaller, open-source models viable as evaluators, reducing vendor lock-in for enterprise benchmarking workflows.

Decision-makers should integrate segmentation and alignment steps into their existing automated evaluation pipelines to improve data reliability and reduce cloud API expenses. For production deployments, setting the split parameter to three segments using token-overlap similarity provides the optimal trade-off between coverage and processing speed. When building evaluation systems, teams should avoid rating-scale formats where models exhibit acute bias and instead favor relation-based pairwise formats.

Users should exercise caution regarding specific limitations. The framework relies on context window capacity, which can be constrained when processing extremely long answers. Additionally, failure rates remain slightly higher on strictly structured content such as computer code (17.13% failure rate) and nuanced ethical dilemmas where underlying models refuse to make clear verdicts. For high-stakes decisions involving sensitive moral topics or complex programming tasks, automated evaluations should still be supplemented with targeted expert human review.

arXiv: 2310.01432
Cover for Split and Merge: Aligning Position Biases in LLM-based Evaluators

Abstract

Large language models (LLMs) have shown promise as automated evaluators for assessing the quality of answers generated by AI systems. However, LLM-based evaluators exhibit position bias, or inconsistency, when used to evaluate candidate answers in pairwise comparisons, favoring either the first or second answer regardless of content. To address this limitation, we propose PORTIA, an alignment-based system designed to mimic human comparison strategies to calibrate position bias in a lightweight yet effective manner. Specifically, PORTIA splits the answers into multiple segments, taking into account both length and semantics, and merges them back into a single prompt for evaluation by LLMs. Extensive experiments with six LLMs on 11,520 answer pairs demonstrate that PORTIA markedly enhances the consistency rates for all models and forms of comparison tested, achieving an average relative improvement of 47.46%. It also enables PORTIA-enhanced GPT-3.5 to achieve agreement rates with humans comparable to GPT-4 and elevates GPT-4’s consistency rate up to 98%. Subsequent human evaluations indicate that the PORTIA-enhanced GPT-3.5 model can even surpass standalone GPT-4 in terms of alignment with human evaluators, highlighting PORTIA’s ability to correct position bias, improve LLM consistency, and boost performance while keeping cost efficiency. can quantify token-level overlap with reference texts but fall short in evaluating semantic quality. While human evaluators provide more accurate and valuable feedback, often considered the “gold standards,” their scalability is generally low, given that they are costly and time-consuming. As a result, there emerges a growing need for automated evaluation methods that reliably align with human yet remain efficient and cost-effective.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 The PORTIA System
  • 3.1 Key Design Considerations
  • 3.2 The Core Splitting Algorithm
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Efficiency and Cost Analysis
  • 4.4 Human study
  • 4.5 Ablation Study
  • 5 Related Work
  • 6 Conclusion
  • 7 Acknowledgements
  • 8 Ethical Considerations
  • 9 Limitations
  • References
  • A Reproducibility
  • B Response Length
  • B.1 Response Length Statistics
  • B.2 Relationship Between Answer Length and Inconsistency
  • B.3 Extremely Short Response
  • B.4 Relationship Between Response Length Gap and Fixed Coverage
  • C Naming Reason
  • D A Preliminary Study of Standalone Comparison
  • E PORTIA's Pipeline
  • F Real-World Performance and Cost Analysis
  • G LLMDetails
  • H Algorithm Illustration
  • I LMMetric
  • J On Llama2
  • K Prompt Templates
  • K.1 Comparison Forms
  • K.2 Alignment Templates
  • L Generalizability of PORTIA
  • L.1 Extended Open-Ended Questions
  • L.2 Main Results
  • M Error Analysis
  • N Annotation Process
  • O Stronger Baselines

Knowls

  1. Knowl 1 — Formalization of Position Bias and Evaluation Consistency in Pairwise LLM Judgments

    definition

    In pairwise LLM-based evaluation, an evaluator model LLM(⋅)LLM(\cdot) is presented with an evaluation prompt comprising a user query q∈Qq \in Q and two candidate responses r1,r2∈Rr_1, r_2 \in R. The evaluator outputs a categorical preference verdict v∈{1,2,3}v \in \{1, 2, 3\}, where 11 denotes a preference for r1r_1, 22 denotes a preference for r2r_2, and 33 denotes a tie:

    v=LLM({q,r1,r2})v = LLM(\{q, r_1, r_2\})

    Under an unbiased evaluator, the evaluation verdict is statistically independent of the presentation order (permutation Π\Pi) of the two candidate answers:

    Π⊥ ⁣ ⁣ ⁣⊥V\Pi \perp\!\!\!\perp V

    On an individual sample level, an evaluation is defined as order-consistent if and only if:

    LLM({q,r1,r2})=LLM({q,r2,r1})LLM(\{q, r_1, r_2\}) = LLM(\{q, r_2, r_1\})

    where the candidate index output by the evaluator is mapped back to the respective candidate answer. Conversely, position bias occurs when Π̸ ⁣⊥ ⁣ ⁣ ⁣⊥V\Pi \not\!\perp\!\!\!\perp V, causing LLM({q,r1,r2})≠LLM({q,r2,r1})LLM(\{q, r_1, r_2\}) \neq LLM(\{q, r_2, r_1\}), indicating that the model's judgment shifts solely due to changing the sequence of candidate presentation.

  2. Knowl 2 — PORTIA Alignment-based Splitting and Merging Algorithm

    algorithm

    PORTIA (Position Bias Alignment) is a calibration framework for pairwise LLM evaluations that mimics human long-text comparison strategies by segmenting candidate answers, aligning them, and merging interleaved segments into a unified prompt. The algorithm enforces content preservation (∑i=1kr1(i)=r1\sum_{i=1}^k r_1^{(i)} = r_1) and order preservation.

    Input: Question qq, Candidate answers r1r_1 and r2r_2, Evaluator verdict function v(⋅)v(\cdot), Split partition count kk
    Output: Order-consistent evaluation verdict v∈{1,2,3}v \in \{1, 2, 3\} or None
    // Step 1: Sentence-boundary format identification
    r1positions←IdentifySentenceBoundaries(r1)r_{1}^{\text{positions}} \leftarrow \text{IdentifySentenceBoundaries}(r_1)
    r2positions←IdentifySentenceBoundaries(r2)r_{2}^{\text{positions}} \leftarrow \text{IdentifySentenceBoundaries}(r_2)
    // Step 2: Length alignment
    [r1(1),…,r1(k)]←EqualLengthSplit(r1positions,k)[r_1^{(1)}, \dots, r_1^{(k)}] \leftarrow \text{EqualLengthSplit}(r_{1}^{\text{positions}}, k)
    [r2(1),…,r2(k)]←EqualLengthSplit(r2positions,k)[r_2^{(1)}, \dots, r_2^{(k)}] \leftarrow \text{EqualLengthSplit}(r_{2}^{\text{positions}}, k)
    if v(q,r1(1),r2(1),…,r1(k),r2(k))==v(q,r2(1),r1(1),…,r2(k),r1(k))v(q, r_1^{(1)}, r_2^{(1)}, \dots, r_1^{(k)}, r_2^{(k)}) == v(q, r_2^{(1)}, r_1^{(1)}, \dots, r_2^{(k)}, r_1^{(k)}) then
        return v(q,r1(1),r2(1),…,r1(k),r2(k))v(q, r_1^{(1)}, r_2^{(1)}, \dots, r_1^{(k)}, r_2^{(k)})
    end if
    // Step 3: Semantic alignment (fallback when length alignment is inconsistent)
    smax⁡←0s_{\max} \leftarrow 0
    ns←0n_s \leftarrow 0
    Search_all←False\text{Search\_all} \leftarrow \text{False}
    r1bestparts←[]r_1^{\text{bestparts}} \leftarrow []
    r2bestparts←[]r_2^{\text{bestparts}} \leftarrow []
    while not Search_all\text{Search\_all} do
        r1parts←Partition(r1positions,k,ns)r_1^{\text{parts}} \leftarrow \text{Partition}(r_{1}^{\text{positions}}, k, n_s)
        r2parts←Partition(r2positions,k,ns)r_2^{\text{parts}} \leftarrow \text{Partition}(r_{2}^{\text{positions}}, k, n_s)
        ns←ns+1n_s \leftarrow n_s + 1
        scum←∑i=1kSimilarity(r1parts[i],r2parts[i])s_{\text{cum}} \leftarrow \sum_{i=1}^k \text{Similarity}(r_1^{\text{parts}}[i], r_2^{\text{parts}}[i])
        if scum>smax⁡s_{\text{cum}} > s_{\max} then
            smax⁡←scums_{\max} \leftarrow s_{\text{cum}}
            r1bestparts←r1partsr_1^{\text{bestparts}} \leftarrow r_1^{\text{parts}}
            r2bestparts←r2partsr_2^{\text{bestparts}} \leftarrow r_2^{\text{parts}}
        end if
        if ns≥MaxCombinationsn_s \ge \text{MaxCombinations} then
            Search_all←True\text{Search\_all} \leftarrow \text{True}
        end if
    end while
    if v(q,r1bestparts[1],r2bestparts[1],…,r1bestparts[k],r2bestparts[k])==v(q,r2bestparts[1],r1bestparts[1],…,r2bestparts[k],r1bestparts[k])v(q, r_1^{\text{bestparts}}[1], r_2^{\text{bestparts}}[1], \dots, r_1^{\text{bestparts}}[k], r_2^{\text{bestparts}}[k]) == v(q, r_2^{\text{bestparts}}[1], r_1^{\text{bestparts}}[1], \dots, r_2^{\text{bestparts}}[k], r_1^{\text{bestparts}}[k]) then
        return v(q,r1bestparts[1],r2bestparts[1],…,r1bestparts[k],r2bestparts[k])v(q, r_1^{\text{bestparts}}[1], r_2^{\text{bestparts}}[1], \dots, r_1^{\text{bestparts}}[k], r_2^{\text{bestparts}}[k])
    end if
    return None

    In Step 1, split positions are identified at sentence delimiters (e.g., punctuation marks for natural text, or syntax-aware AST parsing via Tree-sitter for source code blocks). Step 2 divides character lengths into kk equal parts and snaps to the nearest boundary. If the verdict is still order-inconsistent, Step 3 performs a combinatorial search over valid boundary subsets to maximize cumulative segment similarity before presenting the interleaved segments to the evaluator.

  3. Knowl 3 — Semantic Alignment Similarity Metric and Search Complexity in PORTIA

    theoretical result

    In PORTIA's semantic alignment phase, the optimal segment boundaries are chosen by maximizing the cumulative semantic similarity between paired segments r1tr_1^t and r2tr_2^t across all kk segments.

    For segments r1tr_1^t and r2tr_2^t, the token-overlap similarity is computed as:

    sim_score(r1t,r2t)=∣set(r1t)∩set(r2t)∣max⁡(∣set(r1t)∣,∣set(r2t)∣)\text{sim\_score}(r_1^t, r_2^t) = \frac{|\text{set}(r_1^t) \cap \text{set}(r_2^t)|}{\max(|\text{set}(r_1^t)|, |\text{set}(r_2^t)|)}

    Given p1p_1 candidate split positions in answer r1r_1 and p2p_2 candidate split positions in answer r2r_2, partitioning each response into kk non-empty contiguous segments requires selecting k−1k-1 split positions from the potential boundary points. The total number of split combinations Cal\text{Cal} evaluated during the full combinatorial search is:

    Cal=(p1−1k−1)×(p2−1k−1)\text{Cal} = \binom{p_1 - 1}{k - 1} \times \binom{p_2 - 1}{k - 1}

    Because the search space grows exponentially with kk, setting k=3k = 3 provides an optimal empirical trade-off between execution speed and inconsistency correction coverage. The prompt token overhead scales linearly as O(k)O(k) due to boundary marker insertions.

  4. Knowl 4 — Consistency Calibration of LLM Evaluators Across Comparison Forms

    empirical result

    Evaluating PORTIA across six LLM evaluators on MT-Bench (80 open-ended questions across 8 domains, evaluated over 8 model-answer pairings for a total of 1,920 inputs per evaluator) across relation-based, score-based, and Likert-based pairwise comparison forms demonstrates substantial consistency gains and resolution of position bias.

    Evaluator Metric Relation-based Score-based Likert-based
    Claude2 (API) Original Consistency (%) 28.28 47.34 50.62
    PORTIA Consistency (%) 83.28 (+194.48%) 65.16 (+37.64%) 94.84 (+87.36%)
    Fixed Coverage (%) 79.44 52.22 91.27
    Qwen (API) Original Consistency (%) 63.12 52.66 8.12
    PORTIA Consistency (%) 78.13 (+23.78%) 71.09 (+35.00%) 9.38 (+15.52%)
    Fixed Coverage (%) 65.66 59.78 6.46
    ChatGLM2-6B (Local) Original Consistency (%) 38.44 58.59 26.72
    PORTIA Consistency (%) 61.72 (+60.56%) 74.06 (+26.40%) 64.22 (+140.34%)
    Fixed Coverage (%) 56.09 51.02 60.30
    Llama2-13B (Local) Original Consistency (%) 36.41 N/A N/A
    PORTIA Consistency (%) 68.75 (+88.82%) N/A N/A
    Fixed Coverage (%) 22.51 N/A N/A
    GPT-3.5-turbo (API) Original Consistency (%) 78.12 39.22 78.91
    PORTIA Consistency (%) 88.59 (+13.40%) 54.84 (+39.83%) 98.60 (+24.94%)
    Fixed Coverage (%) 70.63 42.06 96.32
    GPT-4 (API) Original Consistency (%) 93.44 92.75 61.50
    PORTIA Consistency (%) 97.03 (+3.84%) 98.00 (+5.66%) 63.50 (+3.25%)
    Fixed Coverage (%) 80.99 86.33 36.09

    PORTIA achieves an average relative consistency improvement of 47.46% across tested models, resolving an average of 62.31% of initially inconsistent evaluation cases (Fixed Coverage). For GPT-4, consistency reaches 98.00% on score-based prompts and 97.03% on relation-based prompts.

  5. Knowl 5 — GPT-4 Agreement Rates, Computational Latency, and Cost Efficiency of PORTIA

    empirical result

    When calibrated with PORTIA, smaller and open-source models achieve high alignment with standalone consistent GPT-4 judgments at a fraction of the monetary cost and inference time.

    Model Original AR (%) PORTIA AR (%) Carbon (CO2_2eq / 1k) Avg Cost (USD / 1k) Avg Time (s / 1k)
    GPT-4 – – N/A $29.78 13,446
    GPT-3.5 82.50 88.59 7.22 $2.85 2,192
    Qwen 60.83 69.58 N/A $35.49 6,083
    ChatGLM2-6B 20.34 39.16 2.15 $4.09 1,983
    Claude2 43.44 75.09 N/A $27.17 11,561

    Agreement Rate (AR) is computed strictly over samples where standalone GPT-4 produces an order-consistent decision, requiring the evaluated model to match GPT-4's verdict consistently. Applying PORTIA increases AR across evaluators by an average of 16.32% (with Claude2 gaining +31.65%). PORTIA-enhanced GPT-3.5 attains an 88.59% agreement rate with GPT-4 while costing $2.85 per 1,000 queries—representing 9.57% of GPT-4's operational cost ($29.78) and reducing inference time from 13,446 seconds to 2,192 seconds per 1,000 evaluations.

  6. Knowl 6 — Human Agreement Alignment and Non-Equivalence of Position Inconsistency to Ties

    empirical result

    A blinded study of five human evaluators (two industrial developers, three academic researchers) assessing 80 pairwise comparisons between gpt-3.5-turbo and Claude-v1 demonstrates that PORTIA calibration improves human-model agreement beyond standalone state-of-the-art models.

    Evaluator Model Original HAR (%) PORTIA-Fixed HAR (%)
    GPT-3.5 55.00 63.75
    Qwen 35.00 35.00
    ChatGLM2 16.25 17.50
    Claude2 6.25 47.50
    GPT-4 60.00 65.00

    PORTIA-enhanced GPT-3.5 increases Human Agreement Rate (HAR) from 55.00% to 63.75%, surpassing the baseline performance of uncalibrated GPT-4 (60.00%). Claude2 improves from 6.25% to 47.50%.

    Furthermore, statistical analysis disproves the heuristic that LLM position inconsistencies simply reflect close ties:

    1. Human annotators classified only 11.25% of responses as genuine ties ([[C]][[C]]).
    2. Correlation analysis between LLM position inconsistency and human tie classifications yields p=0.7280p = 0.7280, indicating no statistically significant correlation.
    3. Automatically treating all order-inconsistent cases as ties produces severely distorted evaluation outcomes (e.g., classifying 63.59% of Llama2 comparisons as ties).
  7. Knowl 7 — Component Ablation of Length Alignment versus Semantic Alignment in PORTIA

    empirical result

    Ablation experiments comparing full PORTIA against variants without Semantic Alignment (PORTIA w/o SA) and without Length Alignment (PORTIA w/o LA) across five LLM evaluators reveal distinct functional roles for both alignment stages:

    1. Semantic Alignment Necessity in Likert Comparisons: Semantic alignment provides the largest contribution to fixed coverage on Likert-based evaluation forms across all evaluators. Because Likert scales require mapping candidate differences to a multi-point ordinal rating, precise semantic alignment between corresponding topical arguments is necessary to stabilize the model's score assignment.
    2. Balanced Contributions in Relation- and Score-Based Forms: In score-based and relation-based comparisons, length alignment and semantic alignment contribute nearly equally to inconsistency resolution. Length alignment quickly standardizes segment granularity, while semantic alignment resolves remaining topical mismatches.
    3. Form-Level Difficulty Hierarchy: Across all ablations, the fixed coverage rate is highest for Likert-based forms, followed by relation-based forms, and lowest for score-based forms, confirming that Likert evaluation exhibits the strongest order sensitivity.
  8. Knowl 8 — Effect of Response Length and Length Disparity on LLM Position Bias

    empirical result

    Analysis of pairwise evaluation inconsistency as a function of response character length and length disparity demonstrates that position bias is length-dependent:

    Character Range Inconsistency Rate (%)
    1,600–2,400 26.89
    2,400–3,200 23.02
    3,200–4,000 31.84
    4,000–4,800 39.01
    4,800–5,600 42.73
    5,600–6,400 55.45

    Inconsistency rates grow monotonically with total response length, rising from 26.89% for responses between 1,600 and 2,400 characters to 55.45% for responses between 5,600 and 6,400 characters.

    Conversely, when comparing responses with extreme length disparities (such as GPT-3.5-short condensed to ≈1/8\approx 1/8 length vs. full responses), both GPT-3.5 and GPT-4 evaluators achieve 100% order consistency (80/80 cases) without alignment. In such cases, LLM evaluators uniformly favor the longer response regardless of input position, overriding position bias with verbosity preference.

  9. Knowl 9 — Multi-turn Baseline Comparison for Evaluator Human Alignment

    empirical result

    Comparing PORTIA against multi-turn and prompt-engineering calibration techniques on pairwise MT-Bench evaluation demonstrates superior performance at low cost relative to human-in-the-loop (HITL) and multi-turn consistency methods:

    Method GPT-4 Human Agreement (%) GPT-3.5 Human Agreement (%) Relative Cost
    VANILLA 52.7 44.4 1.0×1.0\times
    Baseline (Prompted CoT) 60.0 55.0 1.03×1.03\times
    MEC 60.9 55.6 3.29×3.29\times
    MEC + BPC 62.5 58.7 3.29×3.29\times
    PORTIA (Ours) 65.0 63.8 1.68×1.68\times
    HITLC (Human-in-the-Loop) 73.8 71.3 97.3×97.3\times

    PORTIA achieves 65.0% human agreement on GPT-4 and 63.8% on GPT-3.5 at a 1.68×1.68\times compute cost over vanilla single-turn querying. This outperforms Multi-turn Evaluator Consensus (MEC, 3.29×3.29\times cost, 60.9% / 55.6%) and MEC combined with Bidirectional Position Calibration (MEC+BPC, 62.5% / 58.7%), approaching Human-in-the-Loop (HITLC, 73.8% / 71.3%) without requiring manual human effort (97.3×97.3\times relative cost).

  10. Knowl 10 — Limitations of PORTIA: Context Constraints, Over-Alignment Refusals, and Structured Syntax

    limitation

    PORTIA's split-and-merge evaluation framework has three primary operational limitations:

    1. Context Window Limitations: Interleaving segment boundary markers and comparative structures increases total prompt token count by O(k)O(k). When evaluating very long candidate responses on models with small context windows (e.g., 4k tokens), the combined prompt may exceed maximum window limits, requiring models with extended contexts (e.g., Claude2 with 100k tokens).
    2. Conservative Alignment Refusals: Highly aligned LLMs (such as GPT-3.5 and GPT-4) frequently refuse to provide a definitive judgment or score when evaluating subjective or open-ended topics involving moral or ethical quandaries (e.g., CRISPR gene editing ethical implications) or roleplay prompts. In such cases, the evaluator produces non-verdict responses regardless of segmentation strategy.
    3. Structured Context and Code Dependencies: In domain-specific tasks with tight inter-statement contextual coupling (such as source code with loop invariants and scope dependencies), segmentation can fracture semantic units despite AST parsing with Tree-sitter, resulting in higher failure rates on Coding tasks (17.13% failure rate) compared to Roleplay (8.29%) or Knowledge (9.94%).

Coverage note — Omitted the standalone comparison non-linearity verification (Appendix D) and the secondary 640-question dataset extension (Appendix L) to focus on the primary algorithmic formulation, benchmark results, efficiency analyses, and human validation.

References

  1. 1.claude2. https://www.anthropic.com/index/claude-2.
  2. 2.Llama 3. https://llama.meta.com/llama3/.
  3. 3.qwen. https://github.com/QwenLM/Qwen-7B/blob/main/tech_memo.md.
  4. 4.treesitter. https://tree-sitter.github.io/tree-sitter/.
  5. 5.wormgpt. https://wormgpt.ai/.
  6. 6.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  7. 7.Lingjiao Chen, Matei Zaharia, and James Zou. 2023. How is chatgpt's behavior changing over time? arXiv preprint arXiv:2307.09009.
  8. 8.Cheng-Han Chiang and Hung yi Lee. 2023. Can large language models be an alternative to human evaluations?
  9. 9.Andrew A Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana. 2023. Reducing the carbon impact of generative ai inference (today and in 2035). In Proceedings of the 2nd Workshop on Sustainable Computer Systems, pages 1–7.
  10. 10.DeepSeek-AI. 2024. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL.
  12. 12.Shuzheng Gao, Cuiyun Gao, Yulan He, Jichuan Zeng, Lunyiu Nie, Xin Xia, and Michael R. Lyu. 2023. Code structure-guided transformer for source code summarization. ACM Trans. Softw. Eng. Methodol., 32(1):23:1–23:32.
  13. 13.Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2023. Are large language model-based evaluators the solution to scaling up multilingual evaluation?
  14. 14.Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is chatgpt a good translator? a preliminary study. ArXiv, abs/2301.08745.
  15. 15.Walter Kintsch and Janice Keenan. 1973. Reading rate and retention as a function of the number of propositions in the base structure of sentences. Cognitive Psychology, 5(3):257–274.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  17. 17.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023a. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  18. 18.Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023b. CCTEST: testing and repairing code completion systems. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 1238–1250. IEEE.
  19. 19.Zongjie Li, Chaozheng Wang, Pingchuan Ma, Chaowei Liu, Shuai Wang, Daoyuan Wu, Cuiyun Gao, and Yang Liu. 2024. On extracting specialized code abilities from large language models: A feasibility study. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, ICSE '24, New York, NY, USA. Association for Computing Machinery.
  20. 20.Zongjie Li, Chaozheng Wang, Shuai Wang, and Gao Cuiyun. 2023c. Protecting intellectual property of large language model-based code generation apis via watermarks. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, CCS 2023, Copenhagen, Denmark, November 26-30, 2023.
  21. 21.Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of psychology.
  22. 22.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  23. 23.Yen-Ting Lin and Yun-Nung Chen. 2023. LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models. In Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023), pages 47–58, Toronto, Canada. Association for Computational Linguistics.
  24. 24.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a. G-eval: Nlg evaluation using gpt-4 with better human alignment.
  25. 25.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  26. 26.OpenAI. 2023. Gpt-4 technical report.
  27. 27.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  28. 28.Kaiping Peng, Richard E Nisbett, and Nancy YC Wong. 1997. Validity problems comparing values across cultures and possible solutions. Psychological methods, 2(4):329.
  29. 29.Nazneen Rajani, Nathan Lambert, Sheon Han, Jean Wang, Osvald Nitski, Edward Beeching, and Lewis Tunstall. 2023. Can foundation models label data like humans? Hugging Face Blog. Https://huggingface.co/blog/llm-v-human-data.
  30. 30.Oktavia Yovi Ratnasari. 2023. Students'difficulties in reading comprehension and the strategies to deal with the difficulties. Jurnal Penelitian, Pendidikan, dan Pembelajaran, 18(13).
  31. 31.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China. Association for Computational Linguistics.
  32. 32.Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim. 2023. From humans to machines: can chatgpt-like llms effectively replace human annotators in nlp tasks. In Workshop Proceedings of the 17th International AAAI Conference on Web and Social Media.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  34. 34.Jen tse Huang, Man Ho Adrian Lam, Eric Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. 2023. Emotionally numb or empathetic? evaluating how llms feel using emotionbench. ArXiv, abs/2308.03656.
  35. 35.Chaozheng Wang, Zongjie Li, Cuiyun Gao, Wenxuan Wang, Ting Peng, Hailiang Huang, Yuetang Deng, Shuai Wang, and Michael R Lyu. 2024. Exploring multi-lingual bias of large code models in code generation. arXiv preprint arXiv:2404.19368.
  36. 36.Chaozheng Wang, Zongjie Li, Yun Pena, Shuzheng Gao, Sirong Chen, Shuai Wang, Cuiyun Gao, and Michael R Lyu. 2023a. Reef: A framework for collecting real-world vulnerabilities and fixes. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 1952–1962. IEEE.
  37. 37.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926.
  38. 38.Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al. 2023c. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. arXiv preprint arXiv:2306.05087.
  39. 39.Thilini Wijesiriwardene, Ruwan Wickramarachchi, Bimal Gajera, Shreeyash Gowaikar, Chandan Gupta, Aman Chadha, Aishwarya Naresh Reganti, Amit Sheth, and Amitava Das. 2023. Analogical-a novel benchmark for long text analogy evaluation in large language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 3534–3549.
  40. 40.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453.
  41. 41.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  42. 42.Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023. Evaluating large language models at evaluating instruction following. arXiv preprint arXiv:2310.07641.
  43. 43.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  44. 44.Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023. Wider and deeper llm networks are fairer llm evaluators. arXiv preprint arXiv:2308.01862.
  45. 45.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.
  46. 46.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA. Curran Associates Inc.
  47. 47.Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. 2023. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528.

Citation

MLA
Li, Z., et al. “Split and Merge: Aligning Position Biases in LLM-based Evaluators”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 11084–108, https://doi.org/10.18653/v1/2024.emnlp-main.621.
APA
Li, Z., Wang, C., Ma, P., Wu, D., Wang, S., Gao, C., & Liu, Y. (2024). Split and Merge: Aligning Position Biases in LLM-based Evaluators. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11084–11108. https://doi.org/10.18653/v1/2024.emnlp-main.621
Chicago
Li, Z., C. Wang, P. Ma, et al. 2024. “Split and Merge: Aligning Position Biases in LLM-based Evaluators”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11084–108. https://doi.org/10.18653/v1/2024.emnlp-main.621.
Harvard
Li, Z. et al. (2024) “Split and Merge: Aligning Position Biases in LLM-based Evaluators”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 11084–11108. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.621.
Vancouver
1. Li Z, Wang C, Ma P, Wu D, Wang S, Gao C, Liu Y (2024) Split and Merge: Aligning Position Biases in LLM-based Evaluators. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 11084–11108

BibTeX

@inproceedings{li-etal-2024-split,
    title = "Split and Merge: Aligning Position Biases in {LLM}-based Evaluators",
    author = "Li, Zongjie  and
      Wang, Chaozheng  and
      Ma, Pingchuan  and
      Wu, Daoyuan  and
      Wang, Shuai  and
      Gao, Cuiyun  and
      Liu, Yang",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.621/",
    doi = "10.18653/v1/2024.emnlp-main.621",
    pages = "11084--11108"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/