keyword
self-evaluation bias
Self-evaluation bias refers to the systematic tendency of an artificial intelligence model or automated evaluation metric to score or favor its own generated outputs differently—typically more favorably—than outputs produced by other models or humans. This behavior commonly arises in automated assessment setups, such as when large language models serve as judges to evaluate model responses. Because evaluating models often share stylistic tendencies, token probability distributions, or training data with the models being tested, they can exhibit an implicit preference for familiar phrasing and patterns, or fail to recognize their own generation errors and blind spots. Consequently, self-evaluation bias can distort benchmarking results and lead to inflated or unreliable assessments of overall model performance.
2 items

On the Blind Spots of Model-Based Evaluation Metrics for Text Generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov
Why you should read this
Exposes critical failure modes in popular pretrained language model-based evaluation metrics like BERTScore and MAUVE using synthetic stress tests, while providing practical workarounds to ensure more reliable text generation assessment.
In this work, we study the blind spots of model-based evaluation metrics for text generation. We first analyze the behaviors of model-based metrics under adversarial perturbations and find that they are vulnerable to adversarial attacks. We then show that the blind spots are caused by the fact that model-based metrics are trained on the same data distribution as the generation models. We further propose a simple method to mitigate the blind spots by training the metric on a different data distribution. Experiments on multiple text generation tasks demonstrate the effectiveness of our method.
Added
2026-10-03

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer, Virginia Smith, Mona T. Diab
Why you should read this
Presents RefusalBench, a generative evaluation framework using 176 perturbation strategies to expose how frontier retrieval-augmented models fail to selectively refuse answering from flawed contexts, while demonstrating that this capability requires targeted training rather than increased model scale.
The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.
Added
2026-09-30
