Built independently by an author, for readers. Read the story and support ChapterPal

keyword

assumption verification

Assumption verification is the process in natural language processing and automated reasoning of identifying and determining the validity of implicit or explicit premises embedded within a user query or statement. In question answering and conversational systems, inquiries frequently contain presuppositions that may be false, questionable, or unverifiable. Assumption verification analyzes these underlying claims against external evidence or knowledge sources before formulating a response. By distinguishing valid premises from erroneous or ungrounded ones, this process prevents models from uncritically accepting flawed presuppositions, thereby mitigating factual hallucinations and enabling systems to challenge inaccuracies, provide corrections, and deliver appropriate answers.

1 item

(QA)²: Question Answering with Questionable Assumptions

(QA)²: Question Answering with Questionable Assumptions

Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty

OrganizationsAmazon Web ServicesBoston UniversityGoogleNew York University

Why you should read this

Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.

Naturally occurring information-seeking questions often contain questionable assumptions—assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA)² (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)², systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)², we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.

Added

2026-10-03