Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RefusalBench

RefusalBench is an evaluation framework and benchmark methodology designed to assess the selective refusal capabilities of grounded language models and retrieval-augmented generation systems. In artificial intelligence, selective refusal refers to a system's ability to appropriately decline answering a user query when the retrieved reference context is missing, contradictory, irrelevant, or otherwise flawed, rather than generating hallucinations or exhibiting unwarranted overconfidence or excessive caution. Unlike static benchmarks that can suffer from memorization and data contamination, RefusalBench uses a generative approach that programmatically constructs diagnostic test cases through controlled linguistic perturbations across various categories and intensity levels of informational uncertainty. By evaluating performance across single-document and multi-document settings, it provides a standardized, dynamic means to measure whether models can accurately detect unanswerable conditions, reason over imperfect grounding, and maintain reliable behavior.

1 item

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer, Virginia Smith, Mona T. Diab

OrganizationsAmazonCarnegie Mellon University

Why you should read this

Presents RefusalBench, a generative evaluation framework using 176 perturbation strategies to expose how frontier retrieval-augmented models fail to selectively refuse answering from flawed contexts, while demonstrating that this capability requires targeted training rather than increased model scale.

The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.

Added

2026-09-30