A Critical Evaluation of Evaluations for Long-form Question Answering
Fangyuan XuYixiao SongMohit IyyerEunsol Choi
Demonstrates the critical shortcomings of current crowdsourced and automated evaluation methods for long-form question answering by analyzing domain expert justifications and showing that existing metrics fail to predict human preferences.
Long-form question answering systems generate comprehensive, paragraph-length responses using large language models and information retrieval tools. While system generation capabilities have advanced rapidly, evaluating the quality of these detailed outputs remains a major bottleneck. The article evaluates both human evaluation practices and automated text metrics to establish how long-form answers should be reliably assessed.
The researchers hired qualified domain experts across seven subject areas—including biology, economics, physics, and history—to conduct comparative quality judgments on pairs of human-written and model-generated answers, supported by detailed written justifications. In parallel, the article analyzed 12 automatic evaluation metrics (spanning traditional reference-based tools, language model probabilities, and supervised reward models) across thousands of human comparison pairs to test their ability to predict human preferences.
The analysis yielded several critical findings. First, existing automatic metrics fail to reliably predict overall human preference, often performing no better than simple length heuristics or random choice. Second, domain experts frequently disagree on overall preference because they assign different subjective weights to distinct qualities; however, unlike crowdsourced workers who prioritize surface-level traits like brevity, experts prioritize factual correctness and answer completeness. Third, automated metrics show much higher promise when targeted at specific attributes rather than aggregate quality, such as assessing faithfulness to evidence or logical coherence. Finally, domain experts favored model-generated answers over human-written answers in about 62% of comparisons, though human answers remained strongly preferred in complex, history-based questions.
These findings indicate that relying on a single overall score for long-form answers is fundamentally flawed and masks critical trade-offs between accuracy, depth, and clarity. Organizations developing or deploying automated question answering systems risk misjudging system reliability and performance if they depend on conventional metrics or non-expert crowdworkers. Systems must instead be audited across explicit dimensions tailored to domain requirements.
The article recommends abandoning single-metric scoring in favor of multi-faceted evaluation frameworks that independently measure factuality, completeness, coherence, and ease of understanding. Developers should also incorporate domain experts rather than crowd annotators for benchmark creation, and design targeted automatic metrics that evaluate evidence attribution rather than general text similarity.
These conclusions are bounded by a stationary evaluation design, which assessed static text outputs without interactive multi-turn dialogue, and an English-only dataset drawn primarily from online knowledge-sharing communities. Nevertheless, the findings robustly demonstrate the limitations of current evaluation protocols and outline clear requirements for more reliable assessment frameworks.
- Paper: ASQA: Factoid Questions Meet Long-Form Answers, Ivan Stelmakh et al. (2022). ASQA establishes an influential ambiguous-factoid benchmark and long-answer evaluation metric that provide essential context for the source’s study of long-form QA evaluation.
- Paper: Towards a Unified Multi-Dimensional Evaluator for Text Generation, Ming Zhong et al. (2022). UNIEVAL’s multi-dimensional evaluation framework grounds the source’s argument that long-form answers should be judged across facets rather than reduced to one overall score.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). BLEURT exemplifies learned automatic metrics evaluated for agreement with human judgments, a central practice the source tests in the LFQA setting.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). This empirical critique of dialogue metrics establishes the prior concern that automatic text-generation scores may fail to track human judgments.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey carries the source’s concerns about human alignment and evaluator reliability into a systematic framework for designing and testing LLM judges.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). ARES turns the source’s call for facet-specific evaluation into a practical framework that separately scores context relevance, faithfulness, and answer relevance in RAG.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This later survey develops the move beyond single-score evaluation into a broader account of LLM judges, the qualities they assess, and how to validate them.
