Evaluating Open-Domain Question Answering in the Era of Large Language Models
Ehsan KamallooNouha DziriCharles L. A. ClarkeDavood Rafiei
Reveals that standard lexical metrics drastically underestimate generative LLM performance on open-domain question answering benchmarks by missing semantically equivalent answers and failing to handle hallucinations, demonstrating why human evaluation remains indispensable.
Standard automated evaluation in open-domain question answering relies heavily on lexical matching, which requires a generated answer to match a pre-defined reference answer almost verbatim. However, as the field transitions toward generative systems and large language models that produce diverse and detailed responses, this traditional method increasingly fails to recognize valid answers. The article evaluates how accurately lexical matching and modern automated evaluation techniques reflect the true performance of question-answering systems, especially large language models.
To conduct this assessment, the authors evaluated 12 open-domain question-answering models across two established benchmarks: Natural Questions-OPEN and the historical CuratedTREC 2002 dataset. Model outputs were evaluated using standard exact matching, automated semantic similarity techniques, large language model prompting, and rigorous manual human judgment with search-engine verification.
Key findings show that lexical matching drastically underestimates model capabilities and misrepresents competitive rankings. Under human judgment, true system accuracy on the Natural Questions-OPEN subset increased by 24% on average compared to exact match scores. Large language models experienced the largest underestimation: InstructGPT in a zero-shot setting rose from an initial 12.6% exact-match accuracy to 71.4% under human evaluation (a near 60% increase), while InstructGPT with few-shot prompting achieved 75.8%, establishing a new state of the art on this benchmark. A linguistic breakdown revealed that over 50% of lexical matching failures stemmed from simple semantic equivalence, such as entity name variations and synonyms, while 14% were caused by underlying data quality issues like ambiguous questions. Furthermore, while automated semantic evaluation models resolved many surface-level variations, they frequently misjudged hallucinated or factually incorrect long-form answers generated by large language models as correct.
These findings imply that standard automated evaluation benchmarks are no longer reliable for guiding development or procurement decisions in generative question answering. Relying on strict lexical metrics risks discarding superior generative systems, while relying on automated semantic evaluators risks rewarding plausible-sounding but factually ungrounded answers. In high-stakes enterprise applications, uncritical reliance on automated benchmarks could lead to compliance, safety, and reputational risks.
Organizations evaluating question-answering systems should not rely entirely on standard automated metrics or automated model-based scorers for long-form answers. For critical decision-making and deployment benchmarking, teams should incorporate regular expression matching to capture syntactic variations and maintain human-in-the-loop validation to detect factual hallucinations. Further research is required to develop robust automated evaluation methods that can reliably verify attribution and factual accuracy without requiring exhaustive manual oversight.
A primary limitation of this study is its focus on factoid, short-answer questions, leaving open questions about how evaluation discrepancies affect complex tasks such as multi-hop and causal reasoning. Additionally, the manual analysis relied on sample subsets, meaning exact score shifts may vary across broader question distributions, though the overall finding that human judgment remains irreplaceable is supported with high confidence.
- Paper: ASQA: Factoid Questions Meet Long-Form Answers, Ivan Stelmakh et al. (2022). Its ASQA benchmark and DR score establish the long-form, ambiguous-question setting whose lexical and automated evaluation challenges this paper examines.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT evaluates generated answers on ELI5 using human preferences, providing a key precedent for assessing long-form QA beyond exact-answer matching.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). MMLU’s benchmark design and model-scoring conventions offer foundational context for understanding the paper’s analysis of QA benchmark evaluation.
- Paper: Improving Automatic VQA Evaluation Using Large Language Models, Oscar Mañas et al. (2024). LAVE carries the paper’s concern about overly rigid answer matching into visual QA, testing whether LLM judges can scale human-aligned evaluation.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE extends QA evaluation from answer correctness to the grounding and citation quality of long-form LLM responses.
- Paper: HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models, Junyi Li et al. (2023). HaluEval follows the paper’s warning about automated judges by benchmarking hallucination generation and detection, including in QA.
