ASQA: Factoid Questions Meet Long-Form Answers

Ivan StelmakhYi LuanBhuwan DhingraMing-Wei Chang

article2022EMNLP314 citations

Introduces the ASQA benchmark and an automated evaluation metric to resolve ambiguities in factoid questions through synthesized long-form answers with well-defined standards of factual correctness.

Listen

Many real-world factual inquiries are inherently ambiguous and lead to multiple valid answers depending on how a question is interpreted. While automated question-answering systems excel at retrieving short, single-fact answers, progress in generating detailed long-form explanations has been hindered by a shortage of grounded data and the absence of objective quality metrics. Existing long-form benchmarks often rely on highly subjective, open-ended discussions where correctness is difficult to quantify. The article introduces a benchmark and dataset called ASQA to address this issue by framing long-form question answering around ambiguous factoid queries that require synthesizing multiple distinct answers into a unified, coherent explanation.

To construct the benchmark, the authors gathered 6,316 ambiguous questions and guided trained annotators to write detailed, paragraph-length answers grounded in source passages from Wikipedia. They also designed an automated evaluation metric, known as the DR score, which pairs standard text-matching measurements with automated reading-comprehension testing to verify whether a generated summary successfully answers every valid interpretation. The authors benchmarked several baseline configurations, including closed-book generative models, multi-passage retrieval systems paired with generative models, and gold-standard human responses.

The findings reveal a substantial gap between automated systems and human capabilities. The strongest automated model achieved a composite evaluation score of 32.1, falling well short of human baselines that scored 40.6 without reference context and 61.8 with provided context. Purely generative models without retrieval mechanisms performed poorly, underscoring that effective information retrieval is essential for addressing ambiguous queries. However, advanced language models frequently struggled during the synthesis stage, exhibiting factual errors, omitting necessary answers, or repeating text despite having access to the correct source passages. The proposed automated metric demonstrated strong alignment with human assessments, confirming its reliability for tracking model performance.

These results demonstrate that closing the performance gap in long-form question answering requires improvements in both multi-document retrieval and factual summarization. Organizations and researchers developing conversational agents should focus on training architectures that strictly retain source fidelity and avoid hallucinations during text synthesis. While the benchmark provides a dependable evaluation standard, practitioners should account for the fact that the primary accuracy metric focuses on information coverage and relies on underlying reading-comprehension components. Overall, the article provides a reliable framework to develop and measure systems capable of producing complete, trustworthy long-form answers.

arXiv: 2204.06092
Cover for ASQA: Factoid Questions Meet Long-Form Answers

Abstract

An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA). This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations. The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality. In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA. Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation. Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity. In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question. We use this notion of correctness to define an automated metric of performance for ASQA. Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 ASQA Task and Data
  • 3.1 ASQA Annotation Objectives
  • 3.2 ASQA Annotation Process
  • 3.3 ASQA Dataset
  • 4 ASQA Metrics
  • 4.1 Automated Evaluation
  • 4.2 Human Evaluation
  • 5 Experimental Setup
  • 5.1 Models
  • 5.2 Human Performance
  • 6 Results
  • 7 Analysis
  • 8 Conclusion
  • 9 Limitations
  • Acknowledgements
  • References
  • Appendix
  • A Additional Details on the Annotation Procedure
  • B Additional Details on Modeling
  • C Qualitative Analysis

Knowls

  1. Knowl 1 — ASQA Task and Dataset Specification

    definition

    The ASQA (Answer Summaries for Questions which are Ambiguous) task requires a system to generate a comprehensive, paragraph-long answer a^\hat{a} to an ambiguous factoid question qq. An ambiguous factoid question is one that admits multiple distinct correct answers depending on interpretation, associated with a set of nn disambiguations {(xi,yi)}i=1n\{(x_i, y_i)\}_{i=1}^n, where each xix_i is a disambiguated question and yiy_i is its corresponding short answer.

    A valid long-form answer must fulfill four core objectives:

    1. Completeness: Include all valid short answers y1,…,yny_1, \dots, y_n to the disambiguated questions x1,…,xnx_1, \dots, x_n within appropriate context.
    2. Comprehensiveness: Provide sufficient background details to explain why the question is ambiguous and clarify how different short answers relate to one another.
    3. Fluency: Form a coherent and fluent natural language passage.
    4. Attributability: Ground all claims and additional explanatory facts in an underlying knowledge source (such as Wikipedia).

    The ASQA dataset comprises 6,316 ambiguous factoid questions derived from the AMBIGQA dataset (filtering out instances with more than six disambiguations or disputed annotations). The dataset is split into:

    • Train: 4,353 questions (1 crowdsourced reference answer per question)
    • Dev: 948 questions (2 independent reference answers per question)
    • Test: 1,015 questions (2 independent reference answers per question)

    The average reference answer length is 64.8 words. The dataset exhibits a high inter-annotator agreement of 49.6 ROUGE-L F1 (compared to 16.9 in ELI5), and 92% of the reference tokens are directly present in the provided support documents.

  2. Knowl 2 — Disambiguation-F1 (Disambig-F1) Evaluation Metric

    equation

    The Disambig-F1 metric evaluates the degree to which a generated long-form answer a^(k)\hat{a}^{(k)} contains the factual information required to answer all disambiguated interpretations of an ambiguous question.

    Given an evaluation dataset of NN ambiguous questions where the kk-th instance has n(k)n^{(k)} disambiguations {(xi(k),yi(k))}i=1n(k)\{(x_i^{(k)}, y_i^{(k)})\}_{i=1}^{n^{(k)}}, an automated reading comprehension QA model (specifically, a RoBERTa model fine-tuned on SQuADv2) is used as an extraction function. For each disambiguated question xi(k)x_i^{(k)}, the QA model reads the generated summary a^(k)\hat{a}^{(k)} as context and predicts a short answer string y^i(k)\hat{y}_i^{(k)}.

    The Disambig-F1 score is defined as:

    Disambig-F1=1N∑k=1N1n(k)∑i=1n(k)ϕ(y^i(k),yi(k))\text{Disambig-F1} = \frac{1}{N} \sum_{k=1}^N \frac{1}{n^{(k)}} \sum_{i=1}^{n^{(k)}} \phi\left(\hat{y}_i^{(k)}, y_i^{(k)}\right)

    where ϕ(y^,y)∈[0,1]\phi(\hat{y}, y) \in [0, 1] computes the standard token-level F1 score between normalized predicted answer y^\hat{y} and ground-truth answer yy using SQuADv2 evaluation conventions. If the context does not contain the answer, the QA model predicts an empty string, yielding ϕ=0\phi = 0.

  3. Knowl 3 — DR (Disambiguation-Rouge) Composite Metric

    equation

    The DR (Disambiguation-Rouge) score is an automated evaluation metric for ambiguous long-form question answering that jointly balances semantic factual correctness against answer conciseness and fluency. It combines Disambig-F1 and multi-reference ROUGE-L into a single aggregate value:

    DR=Disambig-F1×ROUGE-L\text{DR} = \sqrt{\text{Disambig-F1} \times \text{ROUGE-L}}

    where:

    • Disambig-F1∈[0,100]\text{Disambig-F1} \in [0, 100] measures the fraction of disambiguated questions answerable from the generated long-form summary using an automated reading comprehension model.
    • ROUGE-L∈[0,100]\text{ROUGE-L} \in [0, 100] is the multi-reference longest common subsequence F1 score computed against all reference answers (taking the maximum score across references for a given question).

    The geometric mean penalizes systems that maximize fluency/overlap at the expense of missing disambiguations, or conversely maximize recall by outputting long, unconcise text dumps that dilute readability.

  4. Knowl 4 — Context Paragraph Construction for Ambiguous Factoid Annotations

    algorithm

    To assist human annotators in constructing grounded long-form answers without requiring them to read full Wikipedia articles, context passages are extracted for each disambiguation using a three-stage filtering procedure.

    Input: Disambiguated question-answer pair (xix_i, yiy_i), set of Wikipedia pages WW visited by source annotators, similarity threshold τ\tau
    Output: Context paragraph CiC_i (or empty context ∅\emptyset)
    CandidatePassages = []
    for each paragraph pp in WW do
        if yiy_i appears as a substring in pp then
            CandidatePassages.append(pp)
        end if
    end for
    if CandidatePassages is empty then
        return ∅\emptyset
    end if
    BestScore = -1.0
    BestPassage = ∅\emptyset
    for each pp in CandidatePassages do
        score = TFIDF_CosineSimilarity(pp, xix_i)
        if score > BestScore then
            BestScore = score
            BestPassage = pp
        end if
    end for
    if BestScore ≥τ\ge \tau then
        return BestPassage
    else
        return ∅\emptyset
    end if

    This procedure yields non-empty additional context passages for 45% of all disambiguations used during annotation, ensuring high precision and preventing confusing or irrelevant text from being shown to annotators.

  5. Knowl 5 — Baseline and Human Performance on the ASQA Benchmark

    data/table

    Automated evaluation of baselines and human references on the development set of ASQA demonstrates a substantial gap between automated generative systems and human performance. Models evaluated include:

    • QUESTION: Baseline repeating the input question 8 times.
    • DPR@1: Top-1 passage retrieved by Dense Passage Retriever trained on Natural Questions.
    • JPR@1: Top-1 passage retrieved by Joint Passage Ranker reranked for multi-answer retrieval.
    • T5 Closed Book (T5-C): Pretrained T5-large generating answers with no retrieved context.
    • T5 Open Book (T5-O-KK): T5-large conditioned on the ambiguous question and top-KK passages retrieved by JPR (K∈{1,3,5}K \in \{1, 3, 5\}).
    • T5 Oracle: T5-large conditioned on gold supporting documents (disambiguations, context passages, and annotator evidence passages).
    • Human Performance: Evaluated without context (HP-w/o-C, lower bound) and with context (HP-w/-C, upper bound).
    Model Length (words) ROUGE-L STR-EM Disambig-F1 DR
    QUESTION 71.6 15.3 1.2 0.2 1.5
    DPR@1 99.9 31.1 30.1 16.7 22.8
    JPR@1 196.8 27.9 45.0 25.8 26.9
    T5 Closed Book (T5-C) 62.5 31.0 10.3 7.4 15.1
    T5 Open Book 1 Passage (T5-O-1) 63.0 36.5 33.6 21.2 27.9
    T5 Open Book 3 Passages (T5-O-3) 71.1 38.8 39.9 25.1 31.2
    T5 Open Book 5 Passages (T5-O-5) 71.6 39.2 41.0 26.4 32.1
    T5 Open w/ Oracle Context (ORACLE) 82.6 46.6 88.7 59.2 52.5
    Human w/o Context (HP-w/o-C) 73.5 42.2 51.8 39.0 40.6
    Human w/ Context (HP-w/-C) 64.8 49.4 98.4 77.4 61.8

    Note: For ORACLE and HP-w/-C, ROUGE-L is evaluated using a single reference answer to avoid evaluating against the identical text shown as input.

    T5 Open Book with 5 passages achieves the highest DR score (32.1) among non-oracle models, but significantly trails human performance without context (40.6 DR) and with context (61.8 DR).

  6. Knowl 6 — Agreement Between Automated Metrics and Human Judgments in ASQA

    empirical result

    A blind human evaluation study on 45 development set questions was conducted to assess the validity of the automated metrics. Human annotators evaluated generated answers across four criteria: Disambiguation Accuracy (ACC, percentage of disambiguated questions answerable), Comprehensiveness (COMP), Fluency (FLUE), and Human Overall impression (HO).

    Pearson correlation coefficients between automated metrics and human judgments are:

    Human Metric ROUGE-L Disambig-F1 DR
    ACC 81.1 99.3 97.9
    COMP 79.3 96.4 93.7
    FLUE 83.4 94.4 94.4
    HO 86.4 92.9 95.0

    Key findings include:

    1. High Disambig-F1 Fidelity: Disambig-F1 achieves near-perfect correlation with human Disambiguation Accuracy (99.3 Pearson correlation), confirming that the SQuADv2-trained QA reader accurately tracks factual disambiguation coverage across models.
    2. Optimal Metric Composite: The DR score attains the highest overall correlation with Human Overall impression (95.0), demonstrating that combining factoid extraction capability (Disambig-F1) with summary conciseness and fluency (ROUGE-L) provides the most reliable automated benchmark for ambiguous long-form QA.
  7. Knowl 7 — Headroom Analysis in Retrieval versus Summarization for ASQA

    empirical result

    Analyzing model performance across varying numbers of retrieved passages highlights distinct bottlenecks in both the retrieval and summarization stages:

    1. Summarization Information Loss: Passages retrieved by JPR@5 contain enough information to achieve a Disambig-F1 score of 46.8, which exceeds the human lower bound without context (39.0). However, when T5-O-5 takes these exact five passages as input, its Disambig-F1 drops to 26.4. This reveals that abstractive generation loses substantial factual information or fails to synthesize facts correctly from the input context.
    2. Retrieval Headroom: The best retrieval model (JPR@5 at 46.8 Disambig-F1) still lags behind the Oracle retriever setting (59.2 Disambig-F1 for T5 with gold context) and the human upper bound (77.4 Disambig-F1).

    Consequently, significant headroom exists in both advancing multi-passage retrieval systems to locate all disambiguations and enhancing generative models to faithfully synthesize multiple retrieved passages without dropping valid answers.

  8. Knowl 8 — Qualitative Error Modes in Generative Long-Form QA for Ambiguous Questions

    empirical result

    Manual qualitative analysis of long-form answers generated by the open-book model (T5-O-5) reveals three primary failure modes during multi-document synthesis:

    1. Hallucination and Conflation: The model generates factually fabricated claims or blends attributes of distinct entities from different retrieved passages. For example, in response to "Who won the mayor race in st petersburg florida?", the model hallucinated the existence of a 2016 election and falsely claimed that Rick Baker won in 2017 despite correct information in the context. In another case, it conflated Daenerys Targaryen from A Song of Ice and Fire with Elizabeth Pennykettle from The Last Dragon Chronicles.
    2. Question Misunderstanding / Digression: The model produces a coherent, topic-adjacent summary that fails to address the exact question. For instance, when asked "When was «under God» added to the Pledge of Allegiance?", the model generated a general history of the pledge and mentioned the adoption date (June 14, 1954) without mentioning the phrase "under God" itself.
    3. Repetition: The model generates redundant sentence structures or repeats identical factual clauses within the paragraph summary.
  9. Knowl 9 — Limitations of the ASQA Dataset and Evaluation Paradigm

    limitation

    The ASQA dataset and its evaluation framework have two notable limitations:

    1. Shared Information Source in Agreement: The high inter-annotator agreement in ASQA (49.6 ROUGE-L vs 16.9 in ELI5) is partly due to annotators receiving pre-identified AMBIGQA disambiguations, which acts as a common scaffold and may inflate inter-annotator consistency compared to purely unprompted long-form writing.
    2. Recall-Orientation of Automated Extraction: Disambiguation metrics (STR-EM and Disambig-F1) primarily measure the recall of target answers. If a model generates fabricated facts or incorrect interpretations alongside the correct disambiguations, Disambig-F1 will not directly penalize the hallucinations unless they actively distract the underlying QA reader or degrade the ROUGE-L component of the composite DR score.
    3. QA Reader Domain Dependency: Disambig-F1 depends on a SQuADv2-trained RoBERTa model. Applying this evaluation paradigm to specialized domains differing significantly from Wikipedia requires re-evaluating or fine-tuning the QA reader to avoid domain shift errors.

Coverage note — None was omitted. All contributed tasks, dataset statistics, annotation algorithms, composite evaluation metrics, baseline experimental results, correlation analyses, and limitation discussions are represented.

References

  1. 1.Abdalghani Abujabal, Rishiraj Saha Roy, Mohamed Yahya, and Gerhard Weikum. 2019. ComQA: A community-sourced dataset for complex factoid question answering with paraphrase clusters. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 307–317, Minneapolis, Minnesota. Association for Computational Linguistics.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  3. 3.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wentau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2174–2184, Brussels, Belgium. Association for Computational Linguistics.
  4. 4.Hoa Trang Dang. 2005. Overview of duc 2005. In Proceedings of the document understanding conference, volume 2005, pages 1–12.
  5. 5.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  6. 6.Alexander Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1074–1084, Florence, Italy. Association for Computational Linguistics.
  7. 7.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558–3567, Florence, Italy. Association for Computational Linguistics.
  8. 8.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  9. 9.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. q^2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7856–7870, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  11. 11.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021–2031, Copenhagen, Denmark. Association for Computational Linguistics.
  12. 12.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  13. 13.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  14. 14.Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  15. 15.Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4940–4957, Online. Association for Computational Linguistics.
  16. 16.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  17. 17.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  18. 18.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. CoRR, abs/2005.11401.
  19. 19.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  20. 20.Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  22. 22.Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, and Hannaneh Hajishirzi. 2021. Joint passage ranking for diverse multi-answer retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6997–7008, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  23. 23.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797, Online. Association for Computational Linguistics.
  24. 24.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  25. 25.Preksha Nema, Mitesh M. Khapra, Anirban Laha, and Balaraman Ravindran. 2017. Diversity driven attention model for query-based abstractive summarization. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1063–1072, Vancouver, Canada. Association for Computational Linguistics.
  26. 26.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated MAchine Reading COmprehension dataset. CoRR, abs/1611.09268.
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. CoRR, abs/1910.10683.
  28. 28.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  29. 29.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  30. 30.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  31. 31.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  32. 32.Claude Sammut and Geoffrey I. Webb, editors. 2010. TF–IDF, pages 986–987. Springer US, Boston, MA.
  33. 33.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  34. 34.Haitian Sun, William W Cohen, and Ruslan Salakhutdinov. 2021. Conditionalqa: A complex reading comprehension dataset with conditional answers. arXiv preprint arXiv:2110.06884.
  35. 35.Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. NewsQA: A machine comprehension dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP, pages 191–200, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Ellen M. Voorhees and Dawn M. Tice. 2000. Building a question answering test collection. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’00, page 200–207, New York, NY, USA. Association for Computing Machinery.
  37. 37.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  38. 38.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  39. 39.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  40. 40.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, and Dragomir Radev. 2021. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5905–5921, Online. Association for Computational Linguistics.

Citation

MLA
Stelmakh, I., et al. “ASQA: Factoid Questions Meet Long-Form Answers”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 8273–88, https://doi.org/10.18653/v1/2022.emnlp-main.566.
APA
Stelmakh, I., Luan, Y., Dhingra, B., & Chang, M.-W. (2022). ASQA: Factoid Questions Meet Long-Form Answers. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8273–8288. https://doi.org/10.18653/v1/2022.emnlp-main.566
Chicago
Stelmakh, I., Y. Luan, B. Dhingra, and M.-W. Chang. 2022. “ASQA: Factoid Questions Meet Long-Form Answers”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8273–88. https://doi.org/10.18653/v1/2022.emnlp-main.566.
Harvard
Stelmakh, I. et al. (2022) “ASQA: Factoid Questions Meet Long-Form Answers”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8273–8288. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.566.
Vancouver
1. Stelmakh I, Luan Y, Dhingra B, Chang M-W (2022) ASQA: Factoid Questions Meet Long-Form Answers. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8273–8288

BibTeX

@inproceedings{stelmakh-etal-2022-asqa,
    title = "{ASQA}: Factoid Questions Meet Long-Form Answers",
    author = "Stelmakh, Ivan  and
      Luan, Yi  and
      Dhingra, Bhuwan  and
      Chang, Ming-Wei",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.566/",
    doi = "10.18653/v1/2022.emnlp-main.566",
    pages = "8273--8288"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/