ASQA: Factoid Questions Meet Long-Form Answers
Ivan StelmakhYi LuanBhuwan DhingraMing-Wei Chang
Introduces the ASQA benchmark and an automated evaluation metric to resolve ambiguities in factoid questions through synthesized long-form answers with well-defined standards of factual correctness.
Many real-world factual inquiries are inherently ambiguous and lead to multiple valid answers depending on how a question is interpreted. While automated question-answering systems excel at retrieving short, single-fact answers, progress in generating detailed long-form explanations has been hindered by a shortage of grounded data and the absence of objective quality metrics. Existing long-form benchmarks often rely on highly subjective, open-ended discussions where correctness is difficult to quantify. The article introduces a benchmark and dataset called ASQA to address this issue by framing long-form question answering around ambiguous factoid queries that require synthesizing multiple distinct answers into a unified, coherent explanation.
To construct the benchmark, the authors gathered 6,316 ambiguous questions and guided trained annotators to write detailed, paragraph-length answers grounded in source passages from Wikipedia. They also designed an automated evaluation metric, known as the DR score, which pairs standard text-matching measurements with automated reading-comprehension testing to verify whether a generated summary successfully answers every valid interpretation. The authors benchmarked several baseline configurations, including closed-book generative models, multi-passage retrieval systems paired with generative models, and gold-standard human responses.
The findings reveal a substantial gap between automated systems and human capabilities. The strongest automated model achieved a composite evaluation score of 32.1, falling well short of human baselines that scored 40.6 without reference context and 61.8 with provided context. Purely generative models without retrieval mechanisms performed poorly, underscoring that effective information retrieval is essential for addressing ambiguous queries. However, advanced language models frequently struggled during the synthesis stage, exhibiting factual errors, omitting necessary answers, or repeating text despite having access to the correct source passages. The proposed automated metric demonstrated strong alignment with human assessments, confirming its reliability for tracking model performance.
These results demonstrate that closing the performance gap in long-form question answering requires improvements in both multi-document retrieval and factual summarization. Organizations and researchers developing conversational agents should focus on training architectures that strictly retain source fidelity and avoid hallucinations during text synthesis. While the benchmark provides a dependable evaluation standard, practitioners should account for the fact that the primary accuracy metric focuses on information coverage and relies on underlying reading-comprehension components. Overall, the article provides a reliable framework to develop and measure systems capable of producing complete, trustworthy long-form answers.
- Paper: Natural Questions: A Benchmark for Question Answering Research, Tom Kwiatkowski et al. (2019). Provides the foundational open-domain question answering benchmark derived from real Google search queries that informs how natural user information needs and ambiguities manifest in QA.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Establishes critical evaluation paradigms for factual consistency and hallucination detection in abstractive summarization that underpin ASQA's focus on reliable fact synthesis.
- Paper: SQuAD: 100,000+ Questions for Machine Comprehension of Text, Pranav Rajpurkar et al. (2016). Introduces the standard factoid reading comprehension dataset and exact-span evaluation paradigms that ASQA expands into ambiguous, long-form synthesis.
- Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). Demonstrates the retriever-reader pipeline over open-domain Wikipedia knowledge sources upon which multi-source synthesis tasks build.
- Paper: Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics, Chin-Yew Lin et al. (2003). Introduces ROUGE and n-gram overlap metrics for summarization evaluation, highlighting the metric limitations that motivate ASQA's QA-based evaluation metric.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). Directly incorporates the ASQA dataset into the ALCE benchmark to evaluate how language models generate long-form answers supported by verifiable citations.
- Paper: TRUE: Re-evaluating Factual Consistency Evaluation, Or Honovich et al. (2022). Consolidates factual consistency metrics, including question-answering based evaluation strategies like those in ASQA, into a unified evaluation meta-benchmark.
- Paper: ExpertQA: Expert-Curated Questions and Attributed Answers, Chaitanya Malaviya et al. (2024). Advances long-form attributed question answering to specialized professional domains using expert validation.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). Extends the evaluation of multi-document question answering and summarization to comprehensive long-context benchmarks.
