keyword
RefusalBench
RefusalBench is an evaluation framework and benchmark methodology designed to assess the selective refusal capabilities of grounded language models and retrieval-augmented generation systems. In artificial intelligence, selective refusal refers to a system's ability to appropriately decline answering a user query when the retrieved reference context is missing, contradictory, irrelevant, or otherwise flawed, rather than generating hallucinations or exhibiting unwarranted overconfidence or excessive caution. Unlike static benchmarks that can suffer from memorization and data contamination, RefusalBench uses a generative approach that programmatically constructs diagnostic test cases through controlled linguistic perturbations across various categories and intensity levels of informational uncertainty. By evaluating performance across single-document and multi-document settings, it provides a standardized, dynamic means to measure whether models can accurately detect unanswerable conditions, reason over imperfect grounding, and maintain reliable behavior.
1 item

