RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Aashiq MuhamedLeonardo F. R. RibeiroMarkus DreyerVirginia SmithMona T. Diab

article2025EACL13 citations

Presents RefusalBench, a generative evaluation framework using 176 perturbation strategies to expose how frontier retrieval-augmented models fail to selectively refuse answering from flawed contexts, while demonstrating that this capability requires targeted training rather than increased model scale.

Listen

Language models integrated into retrieval-augmented generation (RAG) systems frequently encounter flawed, incomplete, or ambiguous context. When high-stakes decisions depend on these systems, models must possess the capability of selective refusal—the ability to abstain from answering when the provided information is defective while still answering valid questions. However, evaluating this capability using traditional static benchmarks is unreliable, as models quickly memorize fixed test sets and exploit dataset artifacts.

The article introduces RefusalBench, a dynamic evaluation framework designed to programmatically generate fresh diagnostic test cases through controlled linguistic modifications. Its primary objective is to systematically evaluate how well language models detect informational uncertainty, calibrate their confidence, and provide the correct reason for refusing to answer.

To ensure rigorous testing, the researchers developed 176 distinct linguistic perturbation strategies covering six categories of uncertainty: ambiguity, contradiction, missing information, false premises, granularity mismatches, and epistemic mismatches. Each strategy operates across three intensity tiers (low, medium, and high). The framework uses a multi-model generator-verifier architecture that accepts new test instances only upon unanimous consensus across multiple models, achieving a 93.1% agreement rate with expert human validation. Using this setup, the authors evaluated over 30 language models across single-document (1,600 test cases) and multi-document (1,506 test cases) benchmarks.

The evaluation revealed several critical findings. First, frontier models systematically fail at selective refusal; accuracy dropped below 50% in multi-document scenarios (peaking at 47.4%), with no model achieving strong performance (above 80%) in both answer accuracy and refusal accuracy simultaneously. Second, selective refusal consists of two distinct capabilities—detecting when to refuse and identifying why to refuse. Many models default to "missing information" as a catch-all reason and fail to categorize complex flaws like granularity mismatches. Third, models exhibit severe miscalibration; between 73% and 99% of model predictions were made with maximum stated confidence, even when accuracy hovered between 40% and 69%. Finally, refusal capability does not improve with increased model scale or longer reasoning traces, but it does show notable improvement through targeted alignment methods like Direct Preference Optimization.

These findings indicate that deploying current models in high-stakes automated workflows carries significant operational and safety risks due to over-confident hallucinations on defective contexts or excessive refusal on valid queries. Because scaling alone does not resolve this deficit, developers cannot rely on larger base models to improve reliability automatically.

Organizations developing or deploying grounded language models should prioritize targeted post-training alignment focused on uncertainty calibration rather than relying on extended inference reasoning. Furthermore, evaluation pipelines should adopt dynamic, multi-model consensus verification to prevent benchmark contamination and ensure ongoing safety compliance.

The conclusions should be interpreted within the article's experimental boundaries. The framework currently focuses on English-language text, isolated generator evaluation without dynamic retriever coupling, and programmatic perturbations that may not capture all real-world data messiness. Nonetheless, the high human agreement rates and statistical rigor provide strong confidence in the diagnostic reliability of the framework.

No sufficiently relevant recommendations were found.

Cover for RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Abstract

The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The RefusalBench Methodology
  • 3.1 Generative Evaluation: Theory and Advantages
  • 3.2 A Linguistic Taxonomy of Informational Uncertainty
  • 3.3 The Perturbation Engine: Levers and Intensity Control
  • 3.4 Quality Control: The Generator-Verifier Pipeline
  • 4 Experiments and Results
  • 4.1 Experimental Setup
  • 4.2 Results and Analysis
  • 4.3 Discussion
  • 5 Conclusion and Future Work
  • 6 Limitations
  • 7 Ethical Considerations
  • References
  • A Extended Related Work
  • A.1 Static Benchmarks for Unanswerability and Abstention
  • A.2 Holistic Taxonomies and Modern Generative Approaches
  • A.3 Distinguishing Selective Refusal from General Refusal Capabilities
  • B Proof of Theorem and Extended Analysis
  • B.1 Notation and Formal Setup
  • B.2 Proof of Theorem
  • B.3 When Static Benchmarks Fail
  • B.4 Practical Implications for RefusalBench
  • C Benchmark Construction and Validation
  • C.1 Detailed Benchmark Construction
  • C.2 Human Validation
  • C.3 Benchmark Composition Details
  • D Detailed Evaluation Metrics
  • E Extended Generator-Verifier Analysis (Supporting RQ1)
  • E.1 Inter-Verifier Agreement Analysis
  • E.2 Generator Performance across Intensity Levels
  • E.3 Overall Perturbation Class Ranking
  • E.4 Detailed Self-Evaluation Bias Analysis
  • F Extended Frontier Model Analysis (Supporting RQ2)
  • F.1 Refusal Detection vs. Categorization on RefusalBench-GaRAGe
  • F.2 Calibration Analysis
  • F.3 Refusal Intensity Curves
  • F.4 Perturbation Performance Heatmaps
  • F.5 Error Rate Analysis
  • F.6 Refusal Accuracy Ranking - RefusalBench-GaRAGe
  • F.7 Comprehensive Performance Dashboards
  • F.8 Response Distribution Analysis
  • F.9 RefusalBench-GaRAGe Answer Quality Analysis
  • F.10 Individual Model Confusion Matrices
  • G Statistical Analysis Details
  • H Extended Analysis of Influential Factors (Supporting RQ3)
  • I RefusalBench Prompts
  • I.1 RefusalBench-NQ Prompts
  • I.1.1 Generator Template
  • I.1.2 Verifier Template
  • I.1.3 Model Evaluation Template
  • I.1.4 Judge Template
  • I.2 RefusalBench-GaRAGe Prompts
  • I.2.1 Generator Template
  • I.2.2 Verifier Template
  • I.2.3 Model Evaluation Template
  • I.2.4 Judge Template
  • I.3 Template Variables and Dynamic Content
  • I.4 Answer Constraints by Intensity Level
  • J Software, Models, and Packages Used
  • K Representative Perturbation Lever Catalogue

Knowls

  1. Knowl 1 — Freshly generated tests bound contamination-related measurement error

    theoretical result

    Let XX be the space of evaluation instances, f:X→[0,1]f:X\to[0,1] a model’s score, and DtD_t the instance distribution at evaluation round t∈{0,…,T}t\in\{0,\ldots,T\}. The target score is gt=Ex∼Dt[f(x)]g_t=\mathbb{E}_{x\sim D_t}[f(x)]. A static estimate g^tstat\hat g_t^{\mathrm{stat}} averages nn independent samples drawn once from the initial distribution D0D_0; a generative estimate g^tgen\hat g_t^{\mathrm{gen}} averages mtm_t fresh independent samples from DtD_t at each round. Define contamination drift as ΔT=sup⁡t≤T∣gt−g(D0)∣\Delta_T=\sup_{t\le T}|g_t-g(D_0)|. For any tolerance ϵ>0\epsilon>0, the paper establishes

    Pr⁡ ⁣(sup⁡t≤T∣g^tstat−gt∣>ϵ)≤2exp⁡ ⁣[−2n(ϵ−ΔT)+2],Pr⁡ ⁣(sup⁡t≤T∣g^tgen−gt∣>ϵ)≤∑t=0T2exp⁡(−2mtϵ2),\Pr\!\left(\sup_{t\le T}|\hat g_t^{\mathrm{stat}}-g_t|>\epsilon\right)\le 2\exp\!\left[-2n(\epsilon-\Delta_T)_+^2\right],\qquad \Pr\!\left(\sup_{t\le T}|\hat g_t^{\mathrm{gen}}-g_t|>\epsilon\right)\le\sum_{t=0}^{T}2\exp(-2m_t\epsilon^2),

    where (z)+=max⁡(z,0)(z)_+=\max(z,0). The paper also gives the static-estimator lower bound Pr⁡(sup⁡t≤T∣g^tstat−gt∣>ϵ)≥1−2exp⁡[−2n(ΔT−ϵ)+2]\Pr(\sup_{t\le T}|\hat g_t^{\mathrm{stat}}-g_t|>\epsilon)\ge 1-2\exp[-2n(\Delta_T-\epsilon)_+^2]. Thus, when drift exceeds the tolerated error, increasing the size of a fixed benchmark does not prevent it from measuring the wrong target; fresh samples avoid this drift term, subject to the stated independent-sampling assumptions.

  2. Knowl 2 — RefusalBench operationalizes six uncertainty types at three intensities

    model/method

    RefusalBench turns answerable question–context pairs into controlled tests of selective refusal using 176 linguistic perturbation strategies (“levers”), spanning six uncertainty types: ambiguity (multiple plausible interpretations), contradiction (inconsistent facts), missing information (a necessary fact is absent), false premise (the query presupposes something contradicted or unsupported by context), granularity mismatch (the evidence is at an incompatible level of detail or aggregation), and epistemic mismatch (the query asks for a subjective judgment, prediction, or other non-factual conclusion from factual evidence). The intended refusal reasons are, respectively, ambiguity, contradiction, missing information, false premise, granularity mismatch, and non-factuality. Each type has LOW, MEDIUM, and HIGH intensity: LOW should leave the answer derivable and be answered; MEDIUM and HIGH should make reliable answering untenable and elicit the corresponding refusal. The levers provide roughly ten strategies for each of the 18 type–intensity combinations. Domain experts specified the logical conditions; language models generated examples, which were reviewed by a human expert.

  3. Knowl 3 — A unanimous multi-model pipeline filters generated perturbations

    model/method

    For each answerable base instance and selected perturbation lever, each of four generator models—Claude-4-Sonnet, DeepSeek-R1, GPT-4o, and Nova-Pro—independently proposes a modified instance. All four models then evaluate each candidate, checking such properties as whether the lever and target were implemented correctly, whether the intended intensity and uncertainty were achieved, whether the text is sound, and whether the expected answer or refusal behavior follows. A candidate enters the benchmark only with unanimous verifier approval. This cross-model filtering is designed to reduce self-evaluation bias and model-specific artifacts; it does not establish that all shared evaluator blind spots have been eliminated.

  4. Knowl 4 — Two benchmarks test single- and multi-document selective refusal

    experimental setup

    RefusalBench-NQ uses NaturalQuestions questions with their KILT ground-truth Wikipedia passages. Its 100 base examples were sampled from questions that all evaluated frontier models answered correctly before perturbation; unanimous verification and stratified sampling produced 1,600 examples balanced across perturbation types and intensities. RefusalBench-GaRAGe uses 100 answerable, human-validated GaRAGe examples, sampled evenly from Science, Health, Business & Industrial, Law & Government, and Finance. Each context was standardized to ten passages, using up to five relevant signal passages and additional noise passages; verification and sampling produced 1,506 examples with naturally imbalanced perturbation coverage. For NQ, Claude-4-Sonnet judged answer quality on a 1–5 scale, with scores of 4 or 5 counted as correct; for GaRAGe, answer attempts were evaluated with the RAF score, which requires intent satisfaction and support from relevant passages. Refusals on both benchmarks were scored by matching the predicted reason to the ground-truth category.

  5. Knowl 5 — Human audits and verifier disagreement support strict filtering

    empirical result

    An expert audited 180 unanimously verified perturbations from each benchmark, sampling ten from each of the six uncertainty types at each of the three intensities. The audit pass rate was 93.1% for RefusalBench-NQ and 88.3% for RefusalBench-GaRAGe. The verifier models also showed substantial self-evaluation bias: on NQ, their average self-evaluation pass rate was 91.0%, compared with 82.1% for cross-model evaluations; Claude-4-Sonnet, for example, passed 75.7% of its own outputs versus 97.3% under peer evaluation. Pairwise Cohen’s κ\kappa ranged from 0.061 to 0.442 on NQ and from 0.116 to 0.230 among calculable GaRAGe pairs; Nova-Pro’s GaRAGe approvals had too little variance for a meaningful κ\kappa. These results show that verifier judgments differ considerably and provide empirical motivation for retaining only candidates approved by every verifier.

  6. Knowl 6 — Frontier models struggle particularly with multi-document refusal

    empirical result

    On RefusalBench-NQ, the highest frontier-model refusal accuracy—correctly refusing with the exact ground-truth category—was 73.0%, achieved by Claude-4-Sonnet. On multi-document RefusalBench-GaRAGe, the best refusal accuracy was 47.4% for DeepSeek-R1; Claude-4-Sonnet fell from 73.0% on NQ to 36.1% on GaRAGe. No frontier model exceeded 80% on both answering answerable questions and correctly refusing unanswerable ones. On GaRAGe, answer eligibility scores exceeded 91% and answer RAF scores ranged from 83.4% to 95.9%, while wrong or low-quality answers remained below 3.4% of responses. The reported pattern therefore centers on deciding when to answer or refuse, rather than simply producing a grounded answer once a model elects to answer.

  7. Knowl 7 — Refusal detection and refusal-reason categorization fail differently

    empirical result

    The evaluation separates the binary decision to refuse from identifying the correct reason for refusal. On RefusalBench-NQ, GPT-4o’s strong refusal-detection F1 reflects an exceptionally cautious response policy: its missed-refusal rate was 4.3%, but it falsely refused 62.8% of answerable questions. Its category accuracy was only 54.1%, indicating that detecting a need to refuse does not ensure correct diagnosis of the information defect. Across models, REFUSE_INFO_MISSING acts as a catch-all predicted reason: ambiguity and granularity-mismatch cases are often assigned to missing information, and the paper reports that this category accounts for 25% of NQ predictions. The detection–categorization gap is wider in the multi-document benchmark.

  8. Knowl 8 — Models state high confidence despite substantial miscalibration

    empirical result

    When prompted to report confidence, models were substantially better calibrated on refusals than on answers, but remained poorly calibrated overall. Expected Calibration Error (ECE) is the confidence-bin-weighted mean absolute difference between empirical accuracy and the bin’s stated-confidence midpoint; lower ECE indicates better calibration. On RefusalBench-NQ, answer ECE ranged from 0.406 to 0.580, while refusal ECE ranged from 0.226 to 0.519. Claude-4-Sonnet had the best overall ECE, 0.286, but its predictions were still unreliable. More than 73% of predictions were made at maximum confidence despite model accuracies of roughly 40–69%, demonstrating that stated confidence often failed to reflect correctness.

  9. Knowl 9 — Refusal accuracy scales inconsistently but responds to alignment

    empirical result

    On RefusalBench-NQ, answer accuracy and refusal accuracy followed different, model-family-specific scaling patterns. Qwen answer accuracy rose from 13.0% at 4B parameters to 56.1% at 7B, but Qwen refusal accuracy remained below 17% across tested sizes. OLMo refusal accuracy rose monotonically from 5.1% at 1B to 19.3% at 32B, while Llama refusal accuracy improved 3.1-fold from 8B to 70B; answer accuracy did not show matching consistent gains. In comparisons of OLMo supervised-fine-tuned (SFT) and direct-preference-optimized (DPO) variants, DPO improved refusal accuracy at every tested scale, with the largest gain at 7B, where refusal accuracy was 3.4 times higher. Across scales, the reported average gains were 2.8 percentage points for refusal accuracy and 10.2 points for answer accuracy; DPO improved answer accuracy at every scale except 7B. These results support the paper’s conclusion that refusal is sensitive to alignment and does not reliably emerge from parameter scaling alone.

  10. Knowl 10 — Implicit uncertainty is harder for models to generate than explicit flaws

    empirical result

    Across generator–verifier evaluations, ambiguity was the hardest perturbation type to produce successfully on both benchmarks. Aggregate pass rates on RefusalBench-NQ were 72.5% for ambiguity, 92.8% for missing information, 93.8% for granularity mismatch, 94.3% for false premise, 97.2% for contradiction, and 97.8% for epistemic mismatch. On RefusalBench-GaRAGe, pass rates were 73.4% for ambiguity, 72.5% for missing information, 76.7% for epistemic mismatch, 78.7% for granularity mismatch, 89.6% for contradiction, and 97.1% for false premise. Thus, explicit logical defects were generally easier for the generators to construct than uncertainties requiring implicit reasoning, especially ambiguity and missing information.

Coverage note — The paper’s stated limitations (synthetic perturbations, English-only scope, shared LLM-verifier blind spots, and exclusion of retrieval-stage failures), domain-by-domain results, extended-reasoning analysis, and the full 176-lever catalogue are not separately represented; the knowls prioritize the core methodology, theoretical result, benchmark validation, and principal model findings.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Alfonso Amayuelas, Kyle Wong, Liangming Pan, Wenhu Chen, and William Wang. 2023. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models. arXiv preprint arXiv:2305.13712.
  3. 3.Anthropic. 2025. System Card: Claude Opus 4 & Claude Sonnet 4.
  4. 4.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, and 1 others. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  5. 5.Youssef Benchekroun, Megi Dervishi, Mark Ibrahim, Jean-Baptiste Gaya, Xavier Martinet, Grégoire Mialon, Thomas Scialom, Emmanuel Dupoux, Dieuwke Hupkes, and Pascal Vincent. 2023. Worldsense: A synthetic benchmark for grounded reasoning in large language models. arXiv preprint arXiv:2311.15930.
  6. 6.Faeze Brahman, Sachin Kumar, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi. 2024. The art of saying no: Contextual noncompliance in language models. Preprint, arXiv:2407.12043.
  7. 7.Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261.
  8. 8.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. arXiv preprint arXiv:2105.03011.
  9. 9.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv e-prints, pages arXiv–2407.
  10. 10.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  11. 11.Shengding Hu, Yifan Luo, Huadong Wang, Xingyi Cheng, Zhiyuan Liu, and Maosong Sun. 2023. Won’t get fooled again: Answering questions with false premises. arXiv preprint arXiv:2307.02394.
  12. 12.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and 1 others. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1–55.
  13. 13.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, and 17 others. 2022. Language models (mostly) know what they know. Preprint, arXiv:2207.05221.
  14. 14.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, and 1 others. 2021. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337.
  15. 15.Najoung Kim, Phu Mon Htut, Samuel R. Bowman, and Jackson Petty. 2023. (QA)2: Question Answering with Questionable Assumptions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8466–8487, Toronto, Canada. Association for Computational Linguistics.
  16. 16.Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, and Samuel J Bell. 2025. Abstentionbench: Reasoning llms fail on unanswerable questions. arXiv preprint arXiv:2506.09038.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, and 1 others. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  18. 18.Aaron Langford, Aayush Shah, Abhanshu Gupta, Abhimanyu Bhatter, Abhinav Goyal, Abhinav Mathur, Abhinav Mohanty, Abhishek Kumar, Abhishek Sethi, Abi Komma, and 1 others. 2025. The amazon nova family of models: Technical report and model card. arXiv preprint arXiv:2506.12103.
  19. 19.Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W Koh, and Yulia Tsvetkov. 2024. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning. Advances in Neural Information Processing Systems, 37:28858–28888.
  20. 20.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334.
  21. 21.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, and 1 others. 2024. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249.
  22. 22.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797.
  23. 23.Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, and 1 others. 2024. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656.
  24. 24.OpenAI. 2025. Openai o3 and o4-mini system card. System Card Version 2, OpenAI, San Francisco, CA.
  25. 25.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193.
  26. 26.Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, and Chien-Sheng Wu. 2024a. Unanswerability evaluation for retrieval augmented generation. Preprint, arXiv:2412.12300.
  27. 27.Zhiyuan Peng, Jinming Nian, Alexandre Evfimievski, and Yi Fang. 2024b. RAG-ConfusionQA: A benchmark for evaluating llms on confusing questions. Preprint, arXiv:2410.14567.
  28. 28.Zhiyuan Peng, Jinming Nian, Alexandre Evfimievski, and Yi Fang. 2025. Eloq: Resources for enhancing llm detection of out-of-scope questions. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 3509–3519.
  29. 29.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, and 1 others. 2020. Kilt: a benchmark for knowledge intensive language tasks. arXiv preprint arXiv:2009.02252.
  30. 30.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  31. 31.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789.
  32. 32.Ionut-Teodor Sorodoc, Leonardo FR Ribeiro, Rexhina Blloshmi, Christopher Davis, and Adrià de Gispert. 2025. Garage: A benchmark with grounding annotations for rag evaluation. arXiv preprint arXiv:2506.07671.
  33. 33.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri Garriga-Alonso, and 1 others. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research.
  34. 34.Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and 1 others. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214.
  35. 35.Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang. 2025. Know your limits: A survey of abstention in large language models. Transactions of the Association for Computational Linguistics, 13:529–556.
  36. 36.Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817.
  37. 37.Xunjian Yin, Baizhou Huang, and Xiaojun Wan. 2023. Alcuna: Large language models meet new knowledge. arXiv preprint arXiv:2310.14820.
  38. 38.Michael JQ Zhang and Eunsol Choi. 2021. Situatedqa: Incorporating extra-linguistic contexts into qa. arXiv preprint arXiv:2109.06157.

Citation

MLA
Muhamed, A., et al. “RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models”. arXiv, 2025, http://arxiv.org/abs/2510.10390v1.
APA
Muhamed, A., Ribeiro, L. F. R., Dreyer, M., Smith, V., & Diab, M. T. (2025). RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models. arXiv. http://arxiv.org/abs/2510.10390v1
Chicago
Muhamed, A., L. F. R. Ribeiro, M. Dreyer, V. Smith, and M. T. Diab. 2025. “RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models”. arXiv. http://arxiv.org/abs/2510.10390v1.
Harvard
Muhamed, A. et al. (2025) “RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.10390v1.
Vancouver
1. Muhamed A, Ribeiro LFR, Dreyer M, Smith V, Diab MT (2025) RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models. arXiv

BibTeX

@article{muhamed2025refusalbench,
  title = {RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models},
  author = {Muhamed, Aashiq and Ribeiro, Leonardo F. R. and Dreyer, Markus and Smith, Virginia and Diab, Mona T.},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.10390v1},
  eprint = {2510.10390}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/