Built independently by an author, for readers. Read the story and support ChapterPal

keyword

positional bias

Positional bias refers to the systematic tendency of an evaluator, particularly an automated system or large language model acting as a judge, to favor candidate responses based on their sequential order or placement within the prompt rather than their intrinsic quality. When comparing multiple outputs, an evaluator exhibiting this bias may consistently prefer options appearing in specific locations, such as the first or last position presented in the context. As a result, simply altering the order in which candidate responses are presented can artificially shift the ranking outcomes, making one model appear superior to another regardless of merit. This phenomenon undermines the reliability and fairness of comparative benchmark evaluations, prompting the use of mitigation strategies such as swapping presentation orders, applying calibration techniques, and aggregating results across multiple permutations.

5 items

Aligning Large Language Models through Synthetic Feedback

Aligning Large Language Models through Synthetic Feedback

Sungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang, Donghyun Kwak, Kang Min Yoo, Minjoon Seo

Why you should read this

Presents an alignment learning framework that trains language models using synthetic feedback derived from contrasting different model sizes and prompt configurations, eliminating reliance on human annotations or proprietary APIs while outperforming models like Alpaca and Dolly-v2.

Aligning large language models (LLMs) to human values has become increasingly important as it enables sophisticated steering of LLMs. However, it requires significant human demonstrations and feedback or distillation from proprietary LLMs such as ChatGPT. In this work, we propose a novel alignment learning framework with synthetic feedback not dependent on extensive human annotations and proprietary LLMs. First, we perform reward modeling (RM) with synthetic feedback by contrasting responses from vanilla LLMs with various sizes and prompts. Then, we use the RM to simulate high-quality demonstrations to train a supervised policy and further optimize the model with reinforcement learning. Our resulting model, Aligned Language Model with Synthetic Training dataset (ALMoST), outperforms recent open-sourced models, which are trained on the outputs of InstructGPT or human-annotated demonstrations, in alignment benchmarks. In human evaluation, our model is preferred to Alpaca and Dolly-v2, 55.0% and 58.5% of the time, respectively. Further analyses demonstrate the efficacy and importance of synthetic feedback in our framework 1.

Added

2026-10-02

Humans or LLMs as the Judge? A Study on Judgement Bias

Humans or LLMs as the Judge? A Study on Judgement Bias

Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang, Benyou Wang

OrganizationsShenzhen Research Institute of Big DataThe Chinese University of Hong Kong

Why you should read this

Reveals critical vulnerabilities in human and automated evaluation pipelines by introducing a reference-free framework to measure authority, gender, and misinformation biases, demonstrating that even advanced language model judges can be systematically manipulated through bias-exploiting attacks.

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation results. In this paper, we propose a novel framework that is free from referencing groundtruth annotations for investigating Misinformation Oversight Bias, Gender Bias, Authority Bias and Beauty Bias on LLM and human judges. We curate a dataset referring to the revised Bloom's Taxonomy and conduct thousands of evaluations. Results show that human and LLM judges are vulnerable to perturbations to various degrees, and that even the cutting-edge judges possess considerable biases. We further exploit these biases to conduct attacks on LLM judges. We hope that our work can notify the community of the bias and vulnerability of human- and LLM-as-a-judge, as well as the urgency of developing robust evaluation systems.

Added

2026-09-28

Large Language Models are not Fair Evaluators

Large Language Models are not Fair Evaluators

Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui

OrganizationsPeking UniversityTencentUniversity of Hong Kong

Why you should read this

Reveals that using large language models as judges introduces severe positional bias that distorts model rankings, and provides effective calibration strategies to align automated evaluations with human judgments.

In this paper, we uncover a positional bias in the evaluation paradigm of adopting large language models (LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models. We find that the quality ranking of candidate responses can be easily hacked by simply altering their order of appearance in the context. This manipulation allows us to skew the evaluation result, making one model appear considerably superior to the other, e.g., Vicuna-13B could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator. We propose a simple yet effective calibration framework to address our discovered positional bias. To evaluate the effectiveness of our framework, we manually annotate the “win/tie/lose” outcomes of responses from ChatGPT and Vicuna-13B in the Vicuna Benchmark’s question prompt. Extensive experiments demonstrate that our approach successfully alleviates evaluation bias, resulting in closer alignment with human judgments. To facilitate future research on more robust large language model comparison, we integrate the techniques in the paper into an easy-to-use toolkit FairEval, along with the human annotations 1.

Added

2026-09-28

Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Vyas Raina, Adian Liusie, Mark J. F. Gales

OrganizationsUniversity of Cambridge

Why you should read this

Demonstrates that appending short, transferable adversarial phrases to text can trick LLM-as-a-judge evaluators into assigning maximum quality scores regardless of actual content, revealing critical vulnerabilities in zero-shot absolute scoring.

Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems. Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation. This work presents the first study on the adversarial robustness of assessment LLMs, where we demonstrate that short universal adversarial phrases can be concatenated to deceive judge LLMs to predict inflated scores. Since adversaries may not know or have access to the judge-LLMs, we propose a simple surrogate attack where a surrogate model is first attacked, and the learned attack phrase then transferred to unknown judge-LLMs. We propose a practical algorithm to determine the short universal attack phrases and demonstrate that when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted. It is found that judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment. Our findings raise concerns on the reliability of LLM-as-a-judge methods, and emphasize the importance of addressing vulnerabilities in LLM assessment methods before deployment in high-stakes real-world scenarios.

Added

2026-09-26

Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings

Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings

Yoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo Cetin

OrganizationsSakana AIUniversity of Oxford

Why you should read this

Introduces DroPE, a simple yet powerful method that dramatically extends LLM context length zero-shot without expensive finetuning, outperforming existing techniques and breaking a major bottleneck in language model scalability.

So far, expensive finetuning beyond the pretraining sequence length has been a requirement for effectively extending the context of language models (LM). In this work, we break this key bottleneck by Dropping the Positional Embeddings of LMs after training (DroPE). Our simple method is motivated by three key theoretical and empirical observations. First, positional embeddings (PEs) serve a crucial role during pretraining, providing an important inductive bias that significantly facilitates convergence. Second, over-reliance on this explicit positional information is also precisely what prevents test-time generalization to sequences of unseen length, even when using popular PE-scaling methods. Third, positional embeddings are not an inherent requirement of effective language modeling and can be safely removed after pretraining, following a short recalibration phase. Empirically, DroPE yields seamless zero-shot context extension without any long-context finetuning, quickly adapting pretrained LMs without compromising their capabilities in the original training context. Our findings hold across different models and dataset sizes, far outperforming previous specialized architectures and established rotary positional embedding scaling methods.

Added

2026-01-22

Creative Commons License