A Critical Evaluation of Evaluations for Long-form Question Answering

Fangyuan XuYixiao SongMohit IyyerEunsol Choi

article2023ACL160 citations

Demonstrates the critical shortcomings of current crowdsourced and automated evaluation methods for long-form question answering by analyzing domain expert justifications and showing that existing metrics fail to predict human preferences.

Listen

Long-form question answering systems generate comprehensive, paragraph-length responses using large language models and information retrieval tools. While system generation capabilities have advanced rapidly, evaluating the quality of these detailed outputs remains a major bottleneck. The article evaluates both human evaluation practices and automated text metrics to establish how long-form answers should be reliably assessed.

The researchers hired qualified domain experts across seven subject areas—including biology, economics, physics, and history—to conduct comparative quality judgments on pairs of human-written and model-generated answers, supported by detailed written justifications. In parallel, the article analyzed 12 automatic evaluation metrics (spanning traditional reference-based tools, language model probabilities, and supervised reward models) across thousands of human comparison pairs to test their ability to predict human preferences.

The analysis yielded several critical findings. First, existing automatic metrics fail to reliably predict overall human preference, often performing no better than simple length heuristics or random choice. Second, domain experts frequently disagree on overall preference because they assign different subjective weights to distinct qualities; however, unlike crowdsourced workers who prioritize surface-level traits like brevity, experts prioritize factual correctness and answer completeness. Third, automated metrics show much higher promise when targeted at specific attributes rather than aggregate quality, such as assessing faithfulness to evidence or logical coherence. Finally, domain experts favored model-generated answers over human-written answers in about 62% of comparisons, though human answers remained strongly preferred in complex, history-based questions.

These findings indicate that relying on a single overall score for long-form answers is fundamentally flawed and masks critical trade-offs between accuracy, depth, and clarity. Organizations developing or deploying automated question answering systems risk misjudging system reliability and performance if they depend on conventional metrics or non-expert crowdworkers. Systems must instead be audited across explicit dimensions tailored to domain requirements.

The article recommends abandoning single-metric scoring in favor of multi-faceted evaluation frameworks that independently measure factuality, completeness, coherence, and ease of understanding. Developers should also incorporate domain experts rather than crowd annotators for benchmark creation, and design targeted automatic metrics that evaluate evidence attribution rather than general text similarity.

These conclusions are bounded by a stationary evaluation design, which assessed static text outputs without interactive multi-turn dialogue, and an English-only dataset drawn primarily from online knowledge-sharing communities. Nevertheless, the findings robustly demonstrate the limitations of current evaluation protocols and outline clear requirements for more reliable assessment frameworks.

Cover for A Critical Evaluation of Evaluations for Long-form Question Answering

Abstract

Long-form question answering (LFQA) enables answering a wide range of questions, but its flexibility poses enormous challenges for evaluation. We perform the first targeted study of the evaluation of long-form answers, covering both human and automatic evaluation practices. We hire domain experts in seven areas to provide preference judgments over pairs of answers, along with free-form justifications for their choices. We present a careful analysis of experts’ evaluation, which focuses on new aspects such as the comprehensiveness of the answer. Next, we examine automatic text generation metrics, finding that no existing metrics are predictive of human preference judgments. However, some metrics correlate with fine-grained aspects of answers (e.g., coherence). We encourage future work to move away from a single “overall score” of the answer and adopt a multi-faceted evaluation, targeting aspects such as factuality and completeness. We publicly release all of our annotations and code to spur future work into LFQA evaluation.¹

Table of Contents

  • 1 Introduction
  • 2 Background and related work
  • 3 How do domain experts evaluate long-form answers?
  • 3.1 Collecting expert judgments
  • 3.2 Quantitative results
  • 3.3 What makes one answer better than another?
  • 3.3.1 Do models understand justifications of human preferences?
  • 4 Do automatic metrics correlate with human judgments?
  • 4.1 Text generation metrics
  • 4.1.1 General-purpose generation metrics
  • 4.1.2 Trained LFQA metrics
  • 4.2 Evaluating automatic metrics
  • 4.3 Results
  • 5 Conclusion & Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Related work on text generation evaluation
  • A.2 Expert Annotation
  • A.2.1 Justification Analysis
  • A.3 Previously Collected Human Evaluation Data
  • A.3.1 LFQA systems
  • A.3.2 Evaluation aspects
  • A.3.3 Example of comments mentioning different aspects for Section 3.3
  • A.4 Automatic Metric Implementation Details
  • A.4.1 Learned Metrics
  • A.4.2 GPT-3 Two-shot
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — Domain-expert preference annotations for long-form answers

    experimental setup

    The study collected 260 expert judgments over 140 pairs of answers to long-form questions, covering biology, physics, chemistry, economics, law, technology/computer science, and history. Experts compared either a highly upvoted and a lower-upvoted human answer (H/H) or a highly upvoted human answer and a zero-shot GPT-3 text-davinci-002 answer (H/M). The set contained 75 H/M pairs and 65 H/H pairs; there were 20 questions per domain except history, which had 15 H/M and 5 H/H questions. Recruited through Upwork, experts had at least a bachelor’s degree in the target field. For each pair, they selected the better answer, indicated whether the choice was difficult, and wrote a free-form justification. Experts reported spending 15–30 minutes per question; the study paid 3.25perquestionandreportsatotalcollectioncostof3.25 per question and reports a total collection cost of 845. The questions came from r/explainlikeimfive and r/AskHistorians posts from July–December 2021. GPT-3 was prompted to generate a long answer, with top-p 1 and temperature 0.7. The annotations and code were publicly released.

  2. Knowl 2 — Expert preferences vary substantially by domain

    empirical result

    In the H/H comparisons, experts preferred the highly upvoted answer in 62.4% of cases on average, providing a limited check that upvotes track answer quality. In H/M comparisons, experts preferred the GPT-3 answer in 61.8% of cases on average, but this varied sharply by domain: biology 53.3%, physics 65%, chemistry 50%, economics 90%, law 90%, technology/computer science 60%, and history 24.4%. Thus, human answers were preferred in history 75.6% of the time. The reported Fleiss’ κ values were 0.52 for biology, 0.50 for physics, 0.40 for economics, and 0.65 for history; values were unavailable for domains with one expert. The domain comparison displayed on page 3 therefore reflects both a slight overall preference for model answers and marked domain dependence. The answer-length distribution on page 4 also shows that human history answers were especially long, averaging 356 words in H/M comparisons, which the authors suggest may help explain the model’s difficulty in that domain.

  3. Knowl 3 — Experts and crowdworkers prioritize different answer qualities

    empirical result

    The authors manually coded 50 randomly sampled justifications from domain experts and 50 from WEBGPT crowd annotators, marking whether each of nine answer aspects was mentioned and whether it was decisive. Experts mentioned factuality 36 times, compared with 20 mentions by crowdworkers, and the authors report that experts were more attentive to correctness than workers relying on evidence documents. Experts treated completeness as decisive twice as often as crowdworkers did (12 versus 6 cases); completeness means addressing the question’s aspects and supplying the information needed to clarify it. Ease of understanding was a leading decisive consideration for both groups. Crowdworkers more often treated conciseness and specificity as decisive, even though short answers can omit necessary information. The aspect-count visualization on page 5 illustrates this contrast: expert judgments emphasize factuality and completeness, whereas crowd judgments more often emphasize surface-level properties.

  4. Knowl 4 — T5 models can recover answer references in preference justifications

    empirical result

    As a proxy for whether language models understand preference explanations, the authors masked answer identities in justifications and asked pretrained T5 models to recover the missing references. Inputs used either the original question and answer pair, the same pair with answer identities flipped, or a randomly paired question and answers. Performance was token-level exact match; the scores below are in the order original, flipped, random. For expert justifications, T5-base scored 0.36, 0.37, 0.33; T5-large 0.51, 0.44, 0.41; T5-3B 0.66, 0.36, 0.48; and T5-11B 0.76, 0.28, 0.47. For WEBGPT justifications, the corresponding scores were 0.40, 0.38, 0.37; 0.50, 0.49, 0.50; 0.60, 0.46, 0.53; and 0.65, 0.40, 0.54. The larger models’ stronger performance on original comments and lower performance on flipped comments suggests that they can use the justification’s answer references, supporting further investigation of multi-faceted automatic evaluation.

  5. Knowl 5 — Learned long-form answer preference metrics

    model/method

    The study trained learned metrics on 17,598 answer-pair preference annotations collected by WEBGPT, removing ties and splitting the data randomly into 70% training, 15% validation, and 15% test sets. The Longformer metric scores one question–answer candidate at a time: it encodes the question with the answer and, when available, the answer with its evidence documents; it concatenates the resulting [CLS] representations and applies a linear layer to produce a scalar score. Pairwise softmax probabilities over two candidate scores are trained with cross-entropy to match the human preference. The authors evaluate Longformer with and without evidence documents. They also fine-tune GPT-3 text-curie-001 on prompts containing the question and both candidates, training it to output which answer is preferred. Unlike Longformer, this GPT-3 metric predicts a preference directly for a pair rather than assigning a score to each answer separately.

  6. Knowl 6 — Benchmark for automatic long-form answer evaluation

    experimental setup

    Automatic metrics were tested by asking whether they select the same answer as human preference judgments. The evaluation combined the new expert annotations with prior HURDLES and WEBGPT judgments. It included 3,478 overall-preference comparisons: 129 expert H/M, 637 WEBGPT H/M, 1,923 WEBGPT model/model, 419 HURDLES H/M, and 370 HURDLES model/model. The fine-grained sets contained 854 coherence comparisons (496 WEBGPT H/M, 164 HURDLES H/M, and 194 HURDLES model/model) and 469 factuality comparisons (149 WEBGPT H/M, 151 HURDLES H/M, and 169 HURDLES model/model). The 12 evaluated metrics were ROUGE, BERTScore, BLEURT, Self-BLEU, GPT-2 perplexity, zero-shot question likelihood, BARTScore, RankGen, QAFactEval, Longformer, evidence-conditioned Longformer, and GPT-3 text-curie-001. Comparisons also used random choice, always choosing the human answer when available, and choosing the longer answer as baselines. The reported metric accuracies on page 8 compare these predictions with human preferences.

  7. Knowl 7 — No automatic metric reliably predicts overall preference

    empirical result

    Across the five overall-preference settings—expert H/M, WEBGPT H/M, WEBGPT model/model, HURDLES H/M, and HURDLES model/model—the learned Longformer metric achieved accuracies of 0.67, 0.62, 0.59, 0.60, and 0.62; GPT-3 text-curie-001 achieved 0.69, 0.55, 0.59, 0.60, and 0.51; and the longer-answer baseline achieved 0.68, 0.52, 0.57, 0.61, and 0.48. These results show that even the strongest metrics vary across comparison settings. For context, estimated human pairwise agreement was 0.80 on the expert data and 0.73 on the WEBGPT data. Fine-grained properties appeared more tractable: QAFactEval reached 0.69 accuracy on WEBGPT H/M factuality judgments, outperforming other tested metrics in that setting, while Self-BLEU reached 0.59 on WEBGPT H/M coherence judgments. The authors nevertheless note that QAFactEval requires evidence documents, which may be unavailable or unreliable.

  8. Knowl 8 — Answer length creates misleading evaluation signals

    empirical result

    Answer length was a surprisingly strong baseline, but its usefulness changed with the evaluation setting. Choosing the longer answer outperformed every unsupervised metric on WEBGPT model/model overall-preference comparisons, was the strongest metric for HURDLES H/M factuality, and was the second-best result on the expert overall-preference data. In contrast, for WEBGPT H/M coherence judgments, choosing the shorter answer would have been correct 62% of the time. These reversals show that length can correlate with preference in particular datasets without serving as a dependable measure of answer quality.

  9. Knowl 9 — Long-form answer evaluation should be multi-faceted

    theoretical result

    The authors argue that a single overall-preference score is a poor target for long-form question answering: experts consider several properties, including factuality, completeness, relevance, coherence, and ease of understanding, and those properties can conflict, as completeness can trade off against conciseness. The observed expert disagreements and the weak cross-setting performance of metrics trained to imitate overall preference support evaluating explicit answer attributes instead. The paper recommends developing targeted evaluations for properties such as factuality, completeness, and ease of understanding, rather than treating an undifferentiated overall score as a reliable measure of answer quality.

  10. Knowl 10 — Scope limits of the evaluation study

    limitation

    The study covers a limited range of long-form question answering: its questions come from search-query or community-forum settings, and its evaluation is restricted to English, with topics shaped by English-speaking culture. It evaluates static, pre-generated model answers; annotators cannot interact with a system over multiple rounds. The authors identify broader application settings and interactive evaluation as important directions beyond the study’s scope.

Coverage note — The full metric-by-dataset accuracy matrix and the small GPT-3 two-shot pilot are not reproduced: the main cross-setting results and fine-grained findings are included, while those supplementary details add less to reconstructing the central contribution.

References

  1. 1.Mario Barrantes, Benedikt Herudek, and Richard Wang. 2020. Adversarial nli for factual correctness in text summarisation models. arXiv preprint arXiv:2005.11739.
  2. 2.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. ArXiv, abs/2004.05150.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020a. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020b. Language models are few-shot learners. ArXiv, abs/2005.14165.
  5. 5.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. ArXiv, abs/2006.14799.
  6. 6.Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020. MOCHA: A dataset for training and evaluating generative reading comprehension metrics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6521–6532, Online. Association for Computational Linguistics.
  7. 7.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR.
  8. 8.Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. QAFactEval: Improved QA-based factual consistency evaluation for summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2587–2601, Seattle, United States. Association for Computational Linguistics.
  9. 9.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558–3567, Florence, Italy. Association for Computational Linguistics.
  10. 10.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  11. 11.Joseph L Fleiss, Bruce Levin, and Myunghee Cho Paik. 2013. Statistical methods for rates and proportions. john wiley & sons.
  12. 12.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  13. 13.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, N. Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2022. Attributed text generation via post-hoc research and revision. ArXiv, abs/2210.08726.
  14. 14.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. arXiv preprint arXiv:2202.06935.
  15. 15.Dan Gillick and Yang Liu. 2010. Non-expert evaluation of summarization systems is risky. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk, pages 148–151, Los Angeles. Association for Computational Linguistics.
  16. 16.Tanya Goyal and Greg Durrett. 2020. Evaluating factuality in generation with dependency-level entailment. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3592–3603, Online. Association for Computational Linguistics.
  17. 17.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. Snac - coherence error detection for narrative summarization. Proceedings of EMNLP.
  18. 18.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. arXiv preprint 2002.08909.
  19. 19.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  20. 20.Neslihan Iskender, Tim Polzehl, and Sebastian Möller. 2020. Best practices for crowd-based evaluation of German summarization: Comparing crowd, expert and automatic evaluation. In Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems, pages 164–175, Online. Association for Computational Linguistics.
  21. 21.Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, and Ming Zhou. 2022. Blonde: An automatic evaluation metric for document-level machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  23. 23.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, and Noah A. Smith. 2022. Bidimensional leaderboards: Generate and evaluate language hand in hand. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3540–3557, Seattle, United States. Association for Computational Linguistics.
  24. 24.Kalpesh Krishna, Yapei Chang, John Wieting, and Mohit Iyyer. 2022. Rankgen: Improving text generation with large ranking models. arXiv preprint arXiv:2205.09726.
  25. 25.Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4940–4957, Online. Association for Computational Linguistics.
  26. 26.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  27. 27.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. Summac: Re-visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  28. 28.J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics, pages 159–174.
  29. 29.Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael Bernstein, and Percy Liang. 2022. Evaluating human-language model interaction. ArXiv, abs/2212.09746.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  31. 31.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  32. 32.Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir R. Radev. 2022. Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation. ArXiv, abs/2212.07981.
  33. 33.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  34. 34.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  35. 35.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  36. 36.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, R. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  37. 37.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. 2021. Mauve: Measuring the gap between neural text and human text using divergence frontiers. In Neural Information Processing Systems.
  38. 38.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 101–108, Online. Association for Computational Linguistics.
  39. 39.Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. Cold decoding: Energy-based constrained text generation with langevin dynamics. arXiv preprint arXiv:2202.11705.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  41. 41.Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021. Efficient content-based sparse attention with routing transformers. Transactions of the Association for Computational Linguistics, 9:53–68.
  42. 42.Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving passage retrieval with zero-shot question generation. arXiv preprint arXiv:2204.07496.
  43. 43.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations.
  44. 44.Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. Bleurt: Learning robust metrics for text generation. In Proceedings of ACL.
  45. 45.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. Asqa: Factoid questions meet long-form answers. arXiv preprint arXiv:2204.06092.
  46. 46.Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Read before generate! faithful long form question answering with machine reading. In Findings of the Association for Computational Linguistics: ACL 2022, pages 744–756, Dublin, Ireland. Association for Computational Linguistics.
  47. 47.Simeng Sun, Ori Shapira, Ido Dagan, and Ani Nenkova. 2019. How to compare summarizers without target length? pitfalls, solutions and re-examination of the neural summarization literature. Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation.
  48. 48.Shufan Wang, Fangyuan Xu, Laure Thompson, Eunsol Choi, and Mohit Iyyer. 2022. Modeling exemplification in long-form question answering via retrieval. In North American Chapter of the Association for Computational Linguistics.
  49. 49.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771.
  50. 50.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. Bartscore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems, volume 34, pages 27263–27277. Curran Associates, Inc.
  51. 51.Chen Zhang, L. F. D’Haro, Qiquan Zhang, Thomas Friedrichs, and Haizhou Li. 2022. Fined-eval: Fine-grained automatic dialogue-level evaluation. ArXiv, abs/2210.13832.
  52. 52.Maosen Zhang, Nan Jiang, Lei Li, and Yexiang Xue. 2020. Language generation via combinatorial constraint satisfaction: A tree search enhanced Monte-Carlo approach. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1286–1298, Online. Association for Computational Linguistics.
  53. 53.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  54. 54.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Peng Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation. ArXiv, abs/2210.07197.
  55. 55.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pages 1097–1100.

Citation

MLA
Xu, F., et al. “A Critical Evaluation of Evaluations for Long-form Question Answering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 3225–45, https://doi.org/10.18653/v1/2023.acl-long.181.
APA
Xu, F., Song, Y., Iyyer, M., & Choi, E. (2023). A Critical Evaluation of Evaluations for Long-form Question Answering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3225–3245. https://doi.org/10.18653/v1/2023.acl-long.181
Chicago
Xu, F., Y. Song, M. Iyyer, and E. Choi. 2023. “A Critical Evaluation of Evaluations for Long-form Question Answering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3225–45. https://doi.org/10.18653/v1/2023.acl-long.181.
Harvard
Xu, F. et al. (2023) “A Critical Evaluation of Evaluations for Long-form Question Answering”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3225–3245. Available at: https://doi.org/10.18653/v1/2023.acl-long.181.
Vancouver
1. Xu F, Song Y, Iyyer M, Choi E (2023) A Critical Evaluation of Evaluations for Long-form Question Answering. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3225–3245

BibTeX

@inproceedings{xu-etal-2023-critical,
    title = "A Critical Evaluation of Evaluations for Long-form Question Answering",
    author = "Xu, Fangyuan  and
      Song, Yixiao  and
      Iyyer, Mohit  and
      Choi, Eunsol",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.181/",
    doi = "10.18653/v1/2023.acl-long.181",
    pages = "3225--3245"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/