RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
Aashiq MuhamedLeonardo F. R. RibeiroMarkus DreyerVirginia SmithMona T. Diab
Presents RefusalBench, a generative evaluation framework using 176 perturbation strategies to expose how frontier retrieval-augmented models fail to selectively refuse answering from flawed contexts, while demonstrating that this capability requires targeted training rather than increased model scale.
Language models integrated into retrieval-augmented generation (RAG) systems frequently encounter flawed, incomplete, or ambiguous context. When high-stakes decisions depend on these systems, models must possess the capability of selective refusal—the ability to abstain from answering when the provided information is defective while still answering valid questions. However, evaluating this capability using traditional static benchmarks is unreliable, as models quickly memorize fixed test sets and exploit dataset artifacts.
The article introduces RefusalBench, a dynamic evaluation framework designed to programmatically generate fresh diagnostic test cases through controlled linguistic modifications. Its primary objective is to systematically evaluate how well language models detect informational uncertainty, calibrate their confidence, and provide the correct reason for refusing to answer.
To ensure rigorous testing, the researchers developed 176 distinct linguistic perturbation strategies covering six categories of uncertainty: ambiguity, contradiction, missing information, false premises, granularity mismatches, and epistemic mismatches. Each strategy operates across three intensity tiers (low, medium, and high). The framework uses a multi-model generator-verifier architecture that accepts new test instances only upon unanimous consensus across multiple models, achieving a 93.1% agreement rate with expert human validation. Using this setup, the authors evaluated over 30 language models across single-document (1,600 test cases) and multi-document (1,506 test cases) benchmarks.
The evaluation revealed several critical findings. First, frontier models systematically fail at selective refusal; accuracy dropped below 50% in multi-document scenarios (peaking at 47.4%), with no model achieving strong performance (above 80%) in both answer accuracy and refusal accuracy simultaneously. Second, selective refusal consists of two distinct capabilities—detecting when to refuse and identifying why to refuse. Many models default to "missing information" as a catch-all reason and fail to categorize complex flaws like granularity mismatches. Third, models exhibit severe miscalibration; between 73% and 99% of model predictions were made with maximum stated confidence, even when accuracy hovered between 40% and 69%. Finally, refusal capability does not improve with increased model scale or longer reasoning traces, but it does show notable improvement through targeted alignment methods like Direct Preference Optimization.
These findings indicate that deploying current models in high-stakes automated workflows carries significant operational and safety risks due to over-confident hallucinations on defective contexts or excessive refusal on valid queries. Because scaling alone does not resolve this deficit, developers cannot rely on larger base models to improve reliability automatically.
Organizations developing or deploying grounded language models should prioritize targeted post-training alignment focused on uncertainty calibration rather than relying on extended inference reasoning. Furthermore, evaluation pipelines should adopt dynamic, multi-model consensus verification to prevent benchmark contamination and ensure ongoing safety compliance.
The conclusions should be interpreted within the article's experimental boundaries. The framework currently focuses on English-language text, isolated generator evaluation without dynamic retriever coupling, and programmatic perturbations that may not capture all real-world data messiness. Nonetheless, the high human agreement rates and statistical rigor provide strong confidence in the diagnostic reliability of the framework.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Its taxonomy and evaluation of noisy retrieval contexts establish the RAG failure modes that RefusalBench tests by asking models to refuse flawed evidence.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). RAGTruth shows how unsupported and contradictory claims arise in grounded generation, clarifying the reliability problem that selective refusal is meant to address.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Self-RAG introduces retrieval decisions and explicit critique of retrieved evidence, useful groundwork for understanding refusal when context is unreliable.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). CheckList’s behavioral testing and systematic perturbations provide an earlier evaluation framework that helps situate RefusalBench’s generated diagnostic cases.
No sufficiently relevant recommendations were found.
