Built independently by an author, for readers. Read the story and support ChapterPal

keyword

open-domain QA

Open-domain question answering is a natural language processing task in which an automated system answers general information-seeking questions across a wide variety of topics without being provided with a specific, predetermined text passage containing the answer. Unlike standard machine reading comprehension, which relies on a single given document to extract an answer, open-domain systems must locate relevant knowledge from large-scale text corpora, such as encyclopedias or the web, or draw upon the internal parametric knowledge of large language models. These systems typically combine document retrieval modules to find evidence from vast knowledge bases with reading or generative components that analyze the retrieved content and synthesize accurate responses.

4 items

(QA)²: Question Answering with Questionable Assumptions

(QA)²: Question Answering with Questionable Assumptions

Najoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson Petty

OrganizationsAmazon Web ServicesBoston UniversityGoogleNew York University

Why you should read this

Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.

Naturally occurring information-seeking questions often contain questionable assumptions—assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA)² (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)², systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)², we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.

Added

2026-10-03

Merging Generated and Retrieved Knowledge for Open-Domain QA

Merging Generated and Retrieved Knowledge for Open-Domain QA

Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang

OrganizationsLG AI ResearchUniversity of Illinois ChicagoUniversity of Michigan

Why you should read this

Proposes a compatibility-oriented framework that pairs LLM-generated texts with retrieved documents to resolve knowledge conflicts and improve open-domain question answering accuracy.

Open-domain question answering (QA) systems are often built with retrieval modules. However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage. Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance. Yet, LLMs tend to “hallucinate” content that conflicts with the retrieved knowledge. Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented Knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information. Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels. Then a Fusion-in-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer. Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks. Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts.¹

Added

2026-10-03

ASQA: Factoid Questions Meet Long-Form Answers

ASQA: Factoid Questions Meet Long-Form Answers

Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang

OrganizationsDuke UniversityGoogleYakov & Partners

Why you should read this

Introduces the ASQA benchmark and an automated evaluation metric to resolve ambiguities in factoid questions through synthesized long-form answers with well-defined standards of factual correctness.

An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA). This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations. The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality. In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA. Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation. Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity. In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question. We use this notion of correctness to define an automated metric of performance for ASQA. Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.

Added

2026-09-29