Built independently by an author, for readers. Read the story and support ChapterPal

keyword

long-form answers

Long-form answers are detailed, multi-sentence or paragraph-length natural language responses generated to address complex, open-ended, or domain-specific questions. Unlike short-form or extractive answers that consist of isolated facts, named entities, or brief phrases, long-form answers synthesize information across multiple statements, provide necessary context, and explain underlying reasoning. In computational linguistics and artificial intelligence evaluation, these extended responses present distinct challenges because standard lexical matching metrics are inadequate for verifying semantic equivalence, factual correctness, the presence of hallucinations, and supporting source attribution.

2 items

Evaluating Open-Domain Question Answering in the Era of Large Language Models

Evaluating Open-Domain Question Answering in the Era of Large Language Models

Ehsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood Rafiei

OrganizationsAllen Institute for AIUniversity of AlbertaUniversity of Waterloo

Why you should read this

Reveals that standard lexical metrics drastically underestimate generative LLM performance on open-domain question answering benchmarks by missing semantically equivalent answers and failing to handle hallucinations, demonstrating why human evaluation remains indispensable.

Lexical matching remains the de facto evaluation method for open-domain question answering (QA). Unfortunately, lexical matching fails completely when a plausible candidate answer does not appear in the list of gold answers, which is increasingly the case as we shift from extractive to generative models. The recent success of large language models (LLMs) for QA aggravates lexical matching failures since candidate answers become longer, thereby making matching with the gold answers even more challenging. Without accurate evaluation, the true progress in open-domain QA remains unknown. In this paper, we conduct a thorough analysis of various open-domain QA models, including LLMs, by manually evaluating their answers on a subset of NQ-OPEN, a popular benchmark. Our assessments reveal that while the true performance of all models is significantly underestimated, the performance of the InstructGPT (zero-shot) LLM increases by nearly +60%, making it on par with existing top models, and the InstructGPT (few-shot) model actually achieves a new state-of-the-art on NQ-OPEN. We also find that more than 50% of lexical matching failures are attributed to semantically equivalent answers. We further demonstrate that regex matching ranks QA models consistent with human judgments, although still suffering from unnecessary strictness. Finally, we demonstrate that automated evaluation models are a reasonable surrogate for lexical matching in some circumstances, but not for long-form answers generated by LLMs. The automated models struggle in detecting hallucinations in LLM answers and are thus unable to evaluate LLMs. At this time, there appears to be no substitute for human evaluation.

Added

2026-09-30

ExpertQA: Expert-Curated Questions and Attributed Answers

ExpertQA: Expert-Curated Questions and Attributed Answers

Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth

OrganizationsUniversity of PennsylvaniaUniversity of Washington

Why you should read this

Presents a high-quality benchmark of 2,177 domain-specific questions across 32 fields to evaluate the factuality and citation quality of language model responses through direct assessment and revision by verified domain experts.

As language models are adopted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying attribution and factuality has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we conduct human evaluation of responses from a few representative systems along various axes of attribution and factuality, by bringing domain experts in the loop. Specifically, we collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. In addition, we ask experts to improve upon responses from language models. The output of our analysis is ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.1

Added

2026-09-26