Split and Merge: Aligning Position Biases in LLM-based Evaluators
Zongjie LiChaozheng WangPingchuan MaDaoyuan WuShuai WangCuiyun GaoYang Liu
Proposes PORTIA, a split-and-merge prompting framework that mitigates pairwise position bias in large language model evaluators by segmenting and aligning candidate answers, enabling cost-effective models like GPT-3.5 to rival or exceed standalone GPT-4 in human agreement.
Organizations increasingly deploy large language models as automated evaluators to assess artificial intelligence outputs. While this practice is faster and less costly than human review, automated evaluators suffer from severe position bias during pairwise comparisons. Specifically, models often favor whichever response appears first or second regardless of actual content quality. Relying on flawed automated judgments introduces operational risks and skews model benchmarking, while relying solely on premium evaluators or human reviewers creates unsustainable costs and scaling bottlenecks.
The article evaluates a lightweight framework called PORTIA, which is designed to calibrate position bias and improve consistency in automated evaluators without modifying model weights. To validate this framework, the authors conducted an extensive empirical study across six distinct language models and 11,520 pairwise comparisons across three comparison formats (score-based, rating scales, and direct relational preferences). The evaluation was conducted primarily on the MT-Bench benchmark spanning diverse subject domains, alongside a 640-question expansion and a five-expert human alignment study.
The findings demonstrate substantial improvements across multiple operational and quality metrics. First, applying the alignment framework yielded an average relative improvement of 47.46% in evaluation consistency across all tested models, correcting an average of 62.31% of previously inconsistent judgments. Second, it raised the consistency rate of top-tier models like GPT-4 up to 98% and mitigated 36% to 86% of its position bias occurrences. Third, the framework enabled lower-cost models like GPT-3.5 to achieve an 88% agreement rate with GPT-4 while consuming less than 10% (specifically 9.57%) of the cost. Finally, human evaluation revealed that GPT-3.5 enhanced by this framework surpassed the baseline, standalone GPT-4 in agreement with human expert consensus (63.75% versus 60.00%).
These results demonstrate that position bias can be mitigated through structured prompting and content alignment rather than expensive model fine-tuning. By bridging the performance gap between mid-tier and frontier models, organizations can drastically cut the financial, temporal, and computational expenditures required for automated evaluation. Furthermore, the framework makes smaller, open-source models viable as evaluators, reducing vendor lock-in for enterprise benchmarking workflows.
Decision-makers should integrate segmentation and alignment steps into their existing automated evaluation pipelines to improve data reliability and reduce cloud API expenses. For production deployments, setting the split parameter to three segments using token-overlap similarity provides the optimal trade-off between coverage and processing speed. When building evaluation systems, teams should avoid rating-scale formats where models exhibit acute bias and instead favor relation-based pairwise formats.
Users should exercise caution regarding specific limitations. The framework relies on context window capacity, which can be constrained when processing extremely long answers. Additionally, failure rates remain slightly higher on strictly structured content such as computer code (17.13% failure rate) and nuanced ethical dilemmas where underlying models refuse to make clear verdicts. For high-stakes decisions involving sensitive moral topics or complex programming tasks, automated evaluations should still be supplemented with targeted expert human review.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This paper establishes the LLM-as-a-judge paradigm on MT-Bench and first systematically identifies position bias in pairwise evaluations, which the source directly aims to calibrate and mitigate.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). This foundational work introduces form-filling and criteria-based prompting for LLM-based evaluation, laying the methodological groundwork for structured prompt-based evaluators.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). This study introduces contextual calibration to diagnose and correct ordering and position biases in few-shot prompting, providing key conceptual foundations for inference-time bias mitigation.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This paper demonstrates that language models rationalize judgments swayed by order and position biases, motivating the source's focus on structured alignment rather than unconstrained rationales.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This comprehensive survey categorizes the broader landscape of LLM-as-a-judge methodologies, contextualizing position bias mitigations within overarching automated evaluation pipelines.
- Paper: Branch-Solve-Merge Improves Large Language Model Evaluation and Generation, Swarnadeep Saha et al. (2024). This work develops the Branch-Solve-Merge decomposition framework to tackle complex evaluations and reduce systematic biases across multi-criteria benchmark tasks.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). This paper introduces an objective, ground-truth benchmark specifically designed to stress-test automated LLM judges and reward verifiers against subtle reasoning errors and ordering effects.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This study establishes adversarial meta-evaluation benchmarks to evaluate whether automated judges succumb to surface-level formatting rather than genuine task adherence.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey provides a forward-looking taxonomy of LLM judge architectures, synthesis strategies, and debiasing approaches across diverse domains.
- Paper: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment, Vyas Raina et al. (2024). This paper investigates universal adversarial manipulation of LLM evaluators, highlighting complementary robustness vulnerabilities beyond standard position biases.
