CREPE: Open-Domain Question Answering with False Presuppositions
Xinyan YuSewon MinLuke ZettlemoyerHannaneh Hajishirzi
Introduces CREPE, a benchmark derived from real-world online forum inquiries, to evaluate and improve how open-domain question answering systems detect flawed premises in user questions and generate factual corrections.
When people ask questions about unfamiliar subjects, they frequently rely on underlying false presuppositions—incorrect premises that are assumed rather than directly stated. Traditional open-domain question answering systems are poorly equipped for these situations, as existing benchmarks generally assume that every inquiry either has a direct factual answer or is unanswerable merely due to missing context. The article addresses this gap by investigating how systems can detect flawed assumptions in real-world inquiries and generate constructive, factual corrections.
The primary objective of the article is to introduce CREPE, a benchmark dataset designed to evaluate how effectively automated systems can identify false presuppositions in natural open-domain questions and generate explicit corrections based on world knowledge.
To construct this benchmark, the researchers extracted 8,466 questions from the Reddit forum 'Explain Like I'm Five' and built an annotation pipeline leveraging highly upvoted community responses to establish ground-truth facts and corrections. The team evaluated several baseline models across two core subtasks: detecting the presence of a false presupposition and generating both the identified presupposition and its factual correction. Experiments assessed systems in an open-domain setting where models retrieve evidence from Wikipedia, as well as an idealized setting where the gold standard explanatory comment is provided directly.
The findings show that false presuppositions are remarkably common, appearing in 25% of the analyzed forum questions and spanning varied forms such as false causal relationships, incorrect predicates, and flawed properties. In detection tasks, standard open-domain models relying on document retrieval achieved a modest macro-F1 score of 67.1%, trailing well behind the 85.1% performance achieved by humans provided with explanatory comments. In generation tasks, automated systems produced fluent text and extracted explicit assumptions with reasonable accuracy, but they struggled severely to explain why those assumptions were false, often defaulting to uninformative negations. A qualitative error analysis revealed that 86% of the system's detection failures stemmed from retrieval bottlenecks, as existing retrieval algorithms frequently pulled topically related passages that lacked the precise evidence needed to verify or refute background assumptions. Preliminary tests with large language models like GPT-3 similarly showed that while responses remained on-topic, they frequently hallucinated facts and failed to explicitly correct underlying misconceptions.
These results demonstrate that answering questions in the wild requires models to go beyond direct entity extraction and develop deeper verification capabilities. For real-world applications such as automated assistants, customer support, and search engines, failing to detect false presuppositions risks reinforcing user misconceptions or generating hallucinated explanations, presenting operational and reputational risks. Progress in this space will hinge on improving evidence retrieval mechanisms that can locate indirect or backgrounded evidence rather than superficial keyword matches.
The article recommends that future system development prioritize the retrieval of targeted refutational evidence and explore training frameworks that explicitly model pragmatic nuances and user intent. Organizations planning to deploy automated question-answering systems should exercise caution and conduct targeted testing on unanswerable and presupposition-heavy queries. Future research must also address key limitations identified in the article, including the inherent subjectivity and debatability of implicit assumptions across different users and the presence of conflicting or outdated information across web sources.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). TruthfulQA establishes the benchmark precedent for testing whether models repeat popular falsehoods, a key reliability problem CREPE reframes around false premises in user questions.
- Paper: Know What You Don't Know: Unanswerable Questions for SQuAD, Pranav Rajpurkar et al. (2018). SQuAD 2.0 provides the unanswerable-question baseline that CREPE distinguishes from questions whose assumptions are themselves false.
No sufficiently relevant recommendations were found.
