Evaluating Open-Domain Question Answering in the Era of Large Language Models

Ehsan KamallooNouha DziriCharles L. A. ClarkeDavood Rafiei

article2023ACL168 citations

Reveals that standard lexical metrics drastically underestimate generative LLM performance on open-domain question answering benchmarks by missing semantically equivalent answers and failing to handle hallucinations, demonstrating why human evaluation remains indispensable.

Listen

Standard automated evaluation in open-domain question answering relies heavily on lexical matching, which requires a generated answer to match a pre-defined reference answer almost verbatim. However, as the field transitions toward generative systems and large language models that produce diverse and detailed responses, this traditional method increasingly fails to recognize valid answers. The article evaluates how accurately lexical matching and modern automated evaluation techniques reflect the true performance of question-answering systems, especially large language models.

To conduct this assessment, the authors evaluated 12 open-domain question-answering models across two established benchmarks: Natural Questions-OPEN and the historical CuratedTREC 2002 dataset. Model outputs were evaluated using standard exact matching, automated semantic similarity techniques, large language model prompting, and rigorous manual human judgment with search-engine verification.

Key findings show that lexical matching drastically underestimates model capabilities and misrepresents competitive rankings. Under human judgment, true system accuracy on the Natural Questions-OPEN subset increased by 24% on average compared to exact match scores. Large language models experienced the largest underestimation: InstructGPT in a zero-shot setting rose from an initial 12.6% exact-match accuracy to 71.4% under human evaluation (a near 60% increase), while InstructGPT with few-shot prompting achieved 75.8%, establishing a new state of the art on this benchmark. A linguistic breakdown revealed that over 50% of lexical matching failures stemmed from simple semantic equivalence, such as entity name variations and synonyms, while 14% were caused by underlying data quality issues like ambiguous questions. Furthermore, while automated semantic evaluation models resolved many surface-level variations, they frequently misjudged hallucinated or factually incorrect long-form answers generated by large language models as correct.

These findings imply that standard automated evaluation benchmarks are no longer reliable for guiding development or procurement decisions in generative question answering. Relying on strict lexical metrics risks discarding superior generative systems, while relying on automated semantic evaluators risks rewarding plausible-sounding but factually ungrounded answers. In high-stakes enterprise applications, uncritical reliance on automated benchmarks could lead to compliance, safety, and reputational risks.

Organizations evaluating question-answering systems should not rely entirely on standard automated metrics or automated model-based scorers for long-form answers. For critical decision-making and deployment benchmarking, teams should incorporate regular expression matching to capture syntactic variations and maintain human-in-the-loop validation to detect factual hallucinations. Further research is required to develop robust automated evaluation methods that can reliably verify attribution and factual accuracy without requiring exhaustive manual oversight.

A primary limitation of this study is its focus on factoid, short-answer questions, leaving open questions about how evaluation discrepancies affect complex tasks such as multi-hop and causal reasoning. Additionally, the manual analysis relied on sample subsets, meaning exact score shifts may vary across broader question distributions, though the overall finding that human judgment remains irreplaceable is supported with high confidence.

arXiv: 2305.06984ehsk/OpenQA-eval
Cover for Evaluating Open-Domain Question Answering in the Era of Large Language Models

Abstract

Lexical matching remains the de facto evaluation method for open-domain question answering (QA). Unfortunately, lexical matching fails completely when a plausible candidate answer does not appear in the list of gold answers, which is increasingly the case as we shift from extractive to generative models. The recent success of large language models (LLMs) for QA aggravates lexical matching failures since candidate answers become longer, thereby making matching with the gold answers even more challenging. Without accurate evaluation, the true progress in open-domain QA remains unknown. In this paper, we conduct a thorough analysis of various open-domain QA models, including LLMs, by manually evaluating their answers on a subset of NQ-OPEN, a popular benchmark. Our assessments reveal that while the true performance of all models is significantly underestimated, the performance of the InstructGPT (zero-shot) LLM increases by nearly +60%, making it on par with existing top models, and the InstructGPT (few-shot) model actually achieves a new state-of-the-art on NQ-OPEN. We also find that more than 50% of lexical matching failures are attributed to semantically equivalent answers. We further demonstrate that regex matching ranks QA models consistent with human judgments, although still suffering from unnecessary strictness. Finally, we demonstrate that automated evaluation models are a reasonable surrogate for lexical matching in some circumstances, but not for long-form answers generated by LLMs. The automated models struggle in detecting hallucinations in LLM answers and are thus unable to evaluate LLMs. At this time, there appears to be no substitute for human evaluation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Open-domain QA Evaluation
  • 3.1 Models
  • 3.2 Dataset
  • 4 Strategies for Evaluating Open-domain QA Models
  • 4.1 Supervised Evaluation via Semantic Similarity
  • 4.2 Zero-shot Evaluation via Prompting
  • 4.3 Human Evaluation
  • 4.4 Results and Discussion
  • 5 Linguistic Analysis of Correct Answers
  • 5.1 Discussion
  • 6 Regex Matching on CuratedTREC
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Zero-shot Evaluation using GPT-4

Knowls

  1. Knowl 1 — Human evaluation substantially raises measured NQ-OPEN accuracy

    data/table

    On a random 301-question subset of NQ-OPEN, the paper compared exact-match (EM) and token F1 with BEM, InstructGPT-eval, and human judgments. Every model scored higher under each additional evaluation method than under EM. The table reports accuracy in percent; each Δ is the percentage-point increase over EM on the same subset. These results show that lexical matching substantially undercounts acceptable answers, especially for the two InstructGPT settings.

    Model EM F1 BEM Δ\Delta InstructGPT-eval Δ\Delta Human Δ\Delta
    InstructGPT (zero-shot) 12.6 27.5 63.5 +50.9 77.1 +64.5 71.4 +58.8
    InstructGPT (few-shot) 33.9 50.5 59.5 +25.6 67.8 +33.9 75.8 +41.9
    DPR 45.9 52.3 52.5 +6.6 55.1 +9.2 58.8 +12.9
    FiD 47.8 55.4 58.1 +10.3 61.5 +13.7 64.8 +17.0
    ANCE+ and FiD 48.2 55.9 59.5 +11.3 63.1 +14.9 65.8 +17.6
    RocketQAv2 and FiD 49.8 58.7 62.5 +12.7 66.1 +16.3 70.1 +20.3
    Contriever and FiD 46.5 55.9 60.8 +14.3 63.1 +16.6 66.5 +20.0
    FiD-KD 50.8 61.2 65.8 +15.0 70.4 +19.6 73.1 +22.3
    GAR+ and FiD 50.8 59.7 63.1 +12.3 67.1 +16.3 69.4 +18.2
    EviGen 51.8 59.5 62.1 +10.3 64.8 +13.0 67.1 +15.3
    EMDR2 53.2 62.6 64.5 +11.3 68.4 +15.2 73.1 +19.9
    R2-D2 52.8 61.4 63.8 +11.0 68.4 +15.6 71.4 +18.6
  2. Knowl 2 — Answer evaluation changes model rankings and lexical metrics correlate weakly with humans

    empirical result

    On the 301-question NQ-OPEN subset, the model judged best depends on the evaluation method. Human assessment gives InstructGPT few-shot the highest accuracy, 75.8%, ahead of EMDR2 at 73.1%; BEM instead selects FiD-KD at 73.1%, while InstructGPT-eval selects InstructGPT zero-shot at 77.1%. The paper reports Kendall’s τ\tau rank correlations with human judgments of 0.70 for BEM and 0.75 for InstructGPT-eval, compared with 0.23 for EM and 0.37 for F1. InstructGPT-eval and BEM therefore track human rankings better than lexical metrics in this experiment, but do not reproduce them exactly. Their average accuracies are respectively 2.9 and 7.6 percentage points below human judgments.

  3. Knowl 3 — Most exact-match failures are semantic or surface-form mismatches

    empirical result

    The authors manually examined 493 answer pairs that humans accepted but EM failed to match on NQ-OPEN. Semantic-equivalence differences accounted for 50.3% of these failures. Common subtypes were bridging or abridging (11.4%), an exact answer embedded in an explanatory response (10.1%), and alternate forms of entity names (9.3%). List-style questions accounted for 20.6%, and differences in temporal or spatial granularity accounted for 15.0%. Approximately 14% of EM failures were attributed to data-quality problems: ambiguous questions or incorrect gold answers. The analysis also identifies symbolic equivalence, including numerically equivalent or approximately equivalent expressions. The authors observe that many failures involve syntactic variation and suggest that answer patterns such as regular expressions can capture some of them.

  4. Knowl 4 — Evaluation compared lexical matching, semantic classifiers, prompted judges, and human review

    model/method

    The study compared four approaches to deciding whether a generated answer is acceptable. EM counts a response as correct only when its normalized text occurs in the gold-answer set; normalization case-folds text and removes punctuation and articles. Token F1 gives partial credit for token overlap, taking the maximum against any gold answer and averaging scores across questions. BEM receives the question, a gold answer, and a candidate answer, and predicts semantic equivalence; when a question has multiple gold answers, each is checked independently and a match to any one makes the candidate correct. The authors also prompted InstructGPT and GPT-4 with the question, a gold answer, and a candidate answer, asking whether the candidate is correct. Human annotators judged answer correctness using the question and candidate answer, with search permitted to find supporting or contradicting evidence.

  5. Knowl 5 — Human answer judgments used blinded annotation and adjudication

    experimental setup

    For the NQ-OPEN evaluation, two author-annotators judged 1,490 question-answer pairs. They saw only each question and candidate answer—not its source model or the benchmark gold answers—and could use a search engine to check correctness. Their agreement was reported as Fleiss’ kappa of 72.8%; they disagreed on 202 pairs, or 13.6% of cases. A third annotator judged those disagreements, and majority vote determined the final label. Answers accepted through this process were added to the gold answers for the sampled questions before calculating human-evaluation accuracy.

  6. Knowl 6 — Automated judges accept some unattributable long-form LLM answers

    empirical result

    The authors found that semantic evaluation can mistake plausible but unsupported details in long LLM responses for evidence of correctness. Among 86 InstructGPT zero-shot answers judged incorrect by humans, 47 were classified as unattributable. InstructGPT-eval nevertheless accepted 30 of these answers, BEM accepted 18, and GPT-4 evaluation misidentified 9. This failure helps explain why automated evaluations favor InstructGPT zero-shot over few-shot even though humans find few-shot more accurate on NQ-OPEN. The authors conclude that these automated judges are not reliable evaluators of long-form LLM answers when they fail to detect hallucinated or unattributable content.

  7. Knowl 7 — NQ-OPEN evaluation used a sampled test set and twelve reproduced QA systems

    experimental setup

    The NQ-OPEN test set contains 3,610 questions; the authors randomly sampled 301 and collected 1,490 unique question-answer outputs from twelve QA systems. The systems included InstructGPT in zero-shot and few-shot settings, DPR, FiD, ANCE+ with FiD, RocketQAv2 with FiD, Contriever with FiD, FiD-KD, GAR+ with FiD, EviGen, EMDR2, and R2-D2. The few-shot InstructGPT prompt contained 64 question-answer examples sampled from NQ-OPEN training data. The other systems covered retriever-reader, end-to-end, and closed-book approaches. Experiments followed the systems’ reproduced settings on Wikipedia; the authors used publicly available code and pretrained models rather than training or fine-tuning new models.

  8. Knowl 8 — CuratedTREC experiment tested regular-expression matching on older news questions

    experimental setup

    The regular-expression comparison used CuratedTREC 2002, a collection of 444 questions with gold answers specified as patterns. Questions came from the TREC QA tracks and had been manually reviewed to remove ambiguous or outdated items. The original AQUAINT news corpus—about one million articles from the late 1990s—was used as the knowledge source and split into non-overlapping 100-word passages, yielding more than four million passages. Seven of the twelve QA systems were evaluated without further training on this dataset, producing 1,872 unique answers. The authors also compared four submitted TREC 2002 systems. Two annotators judged the answers, with a third resolving their 150 disagreements; agreement between the first two was reported as Fleiss’ kappa of 83.5%. The setup was intended to compare current systems with historical systems using the original news source; the closed-book LLMs answered from memory rather than using that source.

  9. Knowl 9 — Regular-expression matching preserves most CuratedTREC rankings but remains strict

    empirical result

    On CuratedTREC 2002, the rankings obtained with regular-expression matching were mostly maintained under BEM, InstructGPT-eval, and human assessment, with the main ranking change being the relative order of InstructGPT zero-shot and few-shot. Both automated judges put zero-shot ahead of few-shot, whereas humans find few-shot best at 92% accuracy and zero-shot at 91%. The leading historical system, LCCmain2002, scored 88.1% under its original pattern-based evaluation; human assessment therefore places the two LLM settings 2.9 and 1.9 percentage points above it, respectively, while the automated judges do not reflect that advantage. Averaged across systems, accuracy under human review is 9.9 percentage points higher than under regular-expression matching; the corresponding gaps are 6.6 points for BEM and 6.4 points for InstructGPT-eval. Thus patterns reduce some strictness relative to ordinary lexical matching, but still miss acceptable answers.

  10. Knowl 10 — The study is limited to factoid questions with typically short answers

    limitation

    The paper’s analysis focuses on factoid, information-seeking questions that typically elicit short answers. It does not establish how the evaluation failures or proposed comparisons extend to more complex QA tasks, including multi-hop reasoning, discrete reasoning, or causal-relation questions; the authors identify these as requiring similar systematic analysis.

Coverage note — The paper’s model-by-model CuratedTREC plot values and detailed subtypes of every failure category are not reproduced as separate knowls because the reported aggregate findings and the main NQ-OPEN results capture the central contributions without duplicating figure-level detail.

References

  1. 1.Akari Asai, Matt Gardner, and Hannaneh Hajishirzi. 2022. Evidentiality-guided generation for knowledge-intensive NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2226–2243, Seattle, United States. Association for Computational Linguistics.
  2. 2.Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, and Caiming Xiong. 2020. Learning to retrieve reasoning paths over wikipedia graph for question answering. In International Conference on Learning Representations.
  3. 3.Petr Baudiš and Jan Šedivy. 2015. ` Modeling of the question answering task in the YodaQA system. In International Conference of the cross-language evaluation Forum for European languages, CLEF’15, pages 222–228. Springer-Verlag.
  4. 4.Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pages 95–136, virtual+Dublin. Association for Computational Linguistics.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  6. 6.Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 291–305, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  7. 7.Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. Evaluating question answering evaluation. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 119–124, Hong Kong, China. Association for Computational Linguistics.
  8. 8.Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020. MOCHA: A dataset for training and evaluating generative reading comprehension metrics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6521–6532, Online. Association for Computational Linguistics.
  9. 9.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879, Vancouver, Canada. Association for Computational Linguistics.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  11. 11.Christopher Clark and Matt Gardner. 2018. Simple and effective multi-paragraph reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 845–855, Melbourne, Australia. Association for Computational Linguistics.
  12. 12.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy. 2022. On the origin of hallucinations in conversational models: Is it the datasets or the models? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5271–5285, Seattle, United States. Association for Computational Linguistics.
  14. 14.Martin Fajcik, Martin Docekal, Karel Ondrej, and Pavel Smrz. 2021. R2-D2: A modular baseline for open-domain question answering. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 854–870, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  16. 16.Gautier Izacard and Edouard Grave. 2021a. Distilling knowledge from reader to retriever for question answering. In International Conference on Learning Representations.
  17. 17.Gautier Izacard and Edouard Grave. 2021b. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  18. 18.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  19. 19.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021. Relevance-guided supervision for OpenQA with ColBERT. Transactions of the Association for Computational Linguistics, 9:929–944.
  20. 20.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  21. 21.Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019. Reasoning over paragraph effects in situations. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 58–62, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2021. Generation-augmented retrieval for open-domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4089–4100, Online. Association for Computational Linguistics.
  23. 23.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  24. 24.Sewon Min, Jordan Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, Colin Raffel, Adam Roberts, Tom Kwiatkowski, Patrick Lewis, Yuxiang Wu, Heinrich Küttler, Linqing Liu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel, Sohee Yang, Minjoon Seo, Gautier Izacard, Fabio Petroni, Lucas Hosseini, Nicola De Cao, Edouard Grave, Ikuya Yamada, Sonse Shimaoka, Masatoshi Suzuki, Shumpei Miyawaki, Shun Sato, Ryo Takahashi, Jun Suzuki, Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz, Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao, Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Wen-tau Yih. 2021. NeurIPS 2020 EfficientQA competition: Systems, analyses and lessons learned. volume 133 of Proceedings of Machine Learning Research, pages 86–111. PMLR.
  25. 25.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797, Online. Association for Computational Linguistics.
  26. 26.OpenAI. 2023. GPT-4 technical report. Technical report.
  27. 27.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, pages 27730–27744. Curran Associates, Inc.
  28. 28.Marius A. Pasca and Sandra M. Harabagiu. 2001. High performance question/answering. In Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’01, page 366–374, New York, NY, USA. Association for Computing Machinery.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  30. 30.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  31. 31.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. arXiv preprint arXiv:2112.12870.
  32. 32.Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, QiaoQiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. RocketQAv2: A joint training method for dense passage retrieval and passage re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2825–2835, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Julian Risch, Timo Möller, Julian Gutsch, and Malte Pietsch. 2021. Semantic answer similarity for evaluating question answering models. In Proceedings of the 3rd Workshop on Machine Reading for Question Answering, pages 149–157, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  34. 34.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  35. 35.Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2022. QA dataset explosion: A taxonomy of NLP resources for question answering and reading comprehension. ACM Computing Surveys, 55(10):1–45.
  36. 36.Chenglei Si, Chen Zhao, and Jordan Boyd-Graber. 2021. What’s in a name? answer equivalence for open-domain question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9623–9629, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Devendra Singh, Siva Reddy, Will Hamilton, Chris Dyer, and Dani Yogatama. 2021. End-to-end training of multi-document reader and retriever for open-domain question answering. In Advances in Neural Information Processing Systems, volume 34, pages 25968–25981.
  38. 38.Ellen M. Voorhees. 2003. Overview of the TREC 2002 question answering track. In TREC.
  39. 39.Ellen M. Voorhees and Dawn M. Tice. 2000. The TREC-8 question answering track. In Proceedings of the Second International Conference on Language Resources and Evaluation (LREC’00), Athens, Greece. European Language Resources Association (ELRA).
  40. 40.Shuohang Wang, Mo Yu, Xiaoxiao Guo, Zhiguo Wang, Tim Klinger, Wei Zhang, Shiyu Chang, Gerry Tesauro, Bowen Zhou, and Jing Jiang. 2018. Rˆ3: Reinforced ranker-reader for open-domain question answering. In Proceedings of the AAAI Conference on Artificial Intelligence.
  41. 41.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In International Conference on Learning Representations.
  42. 42.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  43. 43.Xi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot prompting for textual reasoning. In Advances in Neural Information Processing Systems, volume 35, pages 30378–30392. Curran Associates, Inc.
  44. 44.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open pretrained transformer language models. arXiv preprint arXiv:2205.01068.
  45. 45.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with bert. In International Conference on Learning Representations.

Citation

MLA
Kamalloo, E., et al. “Evaluating Open-Domain Question Answering in the Era of Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 5591–606, https://doi.org/10.18653/v1/2023.acl-long.307.
APA
Kamalloo, E., Dziri, N., Clarke, C. L. A., & Rafiei, D. (2023). Evaluating Open-Domain Question Answering in the Era of Large Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5591–5606. https://doi.org/10.18653/v1/2023.acl-long.307
Chicago
Kamalloo, E., N. Dziri, C. L. A. Clarke, and D. Rafiei. 2023. “Evaluating Open-Domain Question Answering in the Era of Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5591–5606. https://doi.org/10.18653/v1/2023.acl-long.307.
Harvard
Kamalloo, E. et al. (2023) “Evaluating Open-Domain Question Answering in the Era of Large Language Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5591–5606. Available at: https://doi.org/10.18653/v1/2023.acl-long.307.
Vancouver
1. Kamalloo E, Dziri N, Clarke CLA, Rafiei D (2023) Evaluating Open-Domain Question Answering in the Era of Large Language Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5591–5606

BibTeX

@inproceedings{kamalloo-etal-2023-evaluating,
    title = "Evaluating Open-Domain Question Answering in the Era of Large Language Models",
    author = "Kamalloo, Ehsan  and
      Dziri, Nouha  and
      Clarke, Charles  and
      Rafiei, Davood",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.307/",
    doi = "10.18653/v1/2023.acl-long.307",
    pages = "5591--5606"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/