(QA)²: Question Answering with Questionable Assumptions
Najoung KimPhu Mon HtutSamuel R. BowmanJackson Petty
Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.
Real-world information-seeking queries frequently incorporate questionable assumptions, which are premises that are either factually false or unverifiable. When users ask questions with flawed premises, typical automated question answering systems often endorse the false premise by delivering a direct answer rather than identifying and correcting the underlying error. The article addresses this challenge by systematically evaluating how robust modern artificial intelligence systems are when responding to natural user questions containing invalid presuppositions.
To benchmark model capabilities, the article introduces a specialized evaluation dataset containing 602 natural, search-engine-derived questions balanced evenly between valid questions and those containing questionable assumptions. The study evaluates a range of leading language models across three tasks: complete question answering evaluated by human judges, assumption detection to identify problematic premises, and assumption verification to fact-check isolated assumptions against real-world facts.
The findings show that handling flawed assumptions remains a significant hurdle for automated systems. Even top-performing models achieved acceptable answers on only 56% of questions in full end-to-end evaluations, with performance dropping noticeably when queries contained questionable assumptions compared to valid ones. On binary classification tasks, models struggled to identify whether a question contained an invalid assumption, performing near chance at roughly 50% to 64% accuracy. In contrast, systems performed better at verifying isolated factual statements, achieving up to 72% accuracy when an oracle extracted the assumption beforehand. Crucially, assumption detection ability correlated significantly with full question-answering success, whereas raw factual verification did not.
These results indicate that the primary operational bottleneck is not merely a lack of factual knowledge, but rather a failure to recognize implicit assumptions embedded within user prompts. In high-stakes applications, automated systems that blindly accept user premises risk propagating misinformation, posing compliance and reputational hazards. Decision-makers should not rely on off-the-shelf language models for unmonitored information retrieval without implementing explicit assumption detection guardrails.
Organizations developing or deploying automated answering solutions should prioritize assumption detection as an automated screening metric during model development. Future work must focus on methods that reliably identify implicit premises before generating responses. System evaluators should also consider temporal limitations, as facts evolve over time, and account for the inherent difficulty of proving non-existence during factual verification.
- Paper: CREPE: Open-Domain Question Answering with False Presuppositions, Xinyan Yu et al. (2023). CREPE establishes a benchmark and task framing for detecting false presuppositions in natural questions, giving useful context for this study’s evaluation of questionable assumptions.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). TruthfulQA provides an earlier benchmark for how models reproduce false beliefs, clarifying the truthfulness-evaluation context behind testing answers to flawed-premise questions.
- Paper: RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models, Aashiq Muhamed et al. (2025). RefusalBench extends false-premise evaluation into selective refusal, testing whether grounded models recognize defective questions and abstain for the right reasons.
