CREPE: Open-Domain Question Answering with False Presuppositions

Xinyan YuSewon MinLuke ZettlemoyerHannaneh Hajishirzi

article2023ACL95 citations

Introduces CREPE, a benchmark derived from real-world online forum inquiries, to evaluate and improve how open-domain question answering systems detect flawed premises in user questions and generate factual corrections.

Listen

When people ask questions about unfamiliar subjects, they frequently rely on underlying false presuppositions—incorrect premises that are assumed rather than directly stated. Traditional open-domain question answering systems are poorly equipped for these situations, as existing benchmarks generally assume that every inquiry either has a direct factual answer or is unanswerable merely due to missing context. The article addresses this gap by investigating how systems can detect flawed assumptions in real-world inquiries and generate constructive, factual corrections.

The primary objective of the article is to introduce CREPE, a benchmark dataset designed to evaluate how effectively automated systems can identify false presuppositions in natural open-domain questions and generate explicit corrections based on world knowledge.

To construct this benchmark, the researchers extracted 8,466 questions from the Reddit forum 'Explain Like I'm Five' and built an annotation pipeline leveraging highly upvoted community responses to establish ground-truth facts and corrections. The team evaluated several baseline models across two core subtasks: detecting the presence of a false presupposition and generating both the identified presupposition and its factual correction. Experiments assessed systems in an open-domain setting where models retrieve evidence from Wikipedia, as well as an idealized setting where the gold standard explanatory comment is provided directly.

The findings show that false presuppositions are remarkably common, appearing in 25% of the analyzed forum questions and spanning varied forms such as false causal relationships, incorrect predicates, and flawed properties. In detection tasks, standard open-domain models relying on document retrieval achieved a modest macro-F1 score of 67.1%, trailing well behind the 85.1% performance achieved by humans provided with explanatory comments. In generation tasks, automated systems produced fluent text and extracted explicit assumptions with reasonable accuracy, but they struggled severely to explain why those assumptions were false, often defaulting to uninformative negations. A qualitative error analysis revealed that 86% of the system's detection failures stemmed from retrieval bottlenecks, as existing retrieval algorithms frequently pulled topically related passages that lacked the precise evidence needed to verify or refute background assumptions. Preliminary tests with large language models like GPT-3 similarly showed that while responses remained on-topic, they frequently hallucinated facts and failed to explicitly correct underlying misconceptions.

These results demonstrate that answering questions in the wild requires models to go beyond direct entity extraction and develop deeper verification capabilities. For real-world applications such as automated assistants, customer support, and search engines, failing to detect false presuppositions risks reinforcing user misconceptions or generating hallucinated explanations, presenting operational and reputational risks. Progress in this space will hinge on improving evidence retrieval mechanisms that can locate indirect or backgrounded evidence rather than superficial keyword matches.

The article recommends that future system development prioritize the retrieval of targeted refutational evidence and explore training frameworks that explicitly model pragmatic nuances and user intent. Organizations planning to deploy automated question-answering systems should exercise caution and conduct targeted testing on unanswerable and presupposition-heavy queries. Future research must also address key limitations identified in the article, including the inherent subjectivity and debatability of implicit assumptions across different users and the presence of conflicting or outdated information across web sources.

Yu et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for CREPE: Open-Domain Question Answering with False Presuppositions

Abstract

While there has been significant progress towards open-domain (document-based) question answering (QA), current QA systems cannot model false presuppositions within questions---questions containing unwarranted assumptions whose truth conditions do not hold within their context. For example, given "What is the name of Taylor Swift's sister?", while humans recognize its problematic nature, most QA models proceed under false premises by providing negative responses rather than acknowledging them explicitly. In order to build QA systems which handle such cases robustly, we introduce CREPE, a large-scale openly licensed collection of over 15k human-written free-form questions based upon Wikipedia entities and passages where one third of these have answers in their annotations indicating they contain false presuppositions. We provide extensive empirical analyses showing that state-of-the-art retrieval-augmented LM approaches fail dramatically when compared against trivial baselines trained specially for handling unmasked contexts. Through our investigation into components responsible for detecting improper presuppositional elements within input texts along with answer generation strategies once identified,

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Question Answering
  • 2.2 Presupposition
  • 3 Dataset: CREPE
  • 3.1 Data Source
  • 3.2 Data Annotation
  • 3.3 Data Analysis
  • 4 Task Setup
  • 5 Experiments: Detection
  • 5.1 Baselines
  • 5.1.1 Trivial baselines
  • 5.1.2 GOLD-COMMENT track baselines
  • 5.1.3 Main track baselines
  • 5.1.4 Human performance
  • 5.2 Results
  • 6 Experiments: Writing
  • 6.1 Baselines
  • 6.1.1 GOLD-COMMENT track baselines
  • 6.1.2 Main track baselines
  • 6.2 Results: Automatic Evaluation
  • 6.3 Results: Human Evaluation
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Details in Data Source
  • B Details in Data Annotation
  • C Details in Experiments
  • C.1 Model details
  • C.2 Error analysis on the detection subtask
  • C.3 Discussion on inherent ambiguity and inconsistency on the web
  • C.4 Details of the writing subtask
  • D A case study with GPT-3
  • E False Presuppositions in Other Data

Knowls

  1. Knowl 1 — CREPE frames false-presupposition correction as open-domain question answering

    definition

    CREPE is a benchmark for questions written by information-seeking users that contain assumptions treated as shared or accepted but not directly asserted. The task is to determine whether a question contains a false presupposition and, if so, state the presupposition and provide a correction that explains what is false. The benchmark therefore requires more than abstaining from an unanswerable question: a useful response identifies and corrects the user's misconception. For example, a question that treats Newton's equal-and-opposite forces as acting on the same object should be corrected by stating that they act on different objects.

  2. Knowl 2 — CREPE data comes from naturally occurring Reddit questions annotated with community responses

    data/table

    CREPE draws questions from the ELI5 subreddit and uses each question's most-upvoted comment as the reference source for judging whether a presupposition is false and for writing its correction. Annotators first excluded subjective, uninformative, or personal-experience questions; then they judged whether the comment identified a false presupposition and, when applicable, wrote a concise presupposition and correction. Two annotators worked independently on each item; disagreements were resolved by a third annotator and majority vote. Initial label agreement was 75% with Fleiss' κ=0.43\kappa=0.43.

    The labeled set contains 8,466 questions, of which 2,202 (26.0%) are labeled as containing a false presupposition. The time-based splits are: training, 3,462 questions (907 with false presuppositions; 2011–2018); development, 2,000 (544; January–June 2019); and test, 3,004 (751; July–December 2019). An additional 196,385 question-comment pairs from 2011–2018 are unlabeled. Questions alone are the input in the main track; the annotation comment is additionally supplied in the GOLD-COMMENT track. The comment is used as a practical proxy for valid judgments and adequate factual corrections, not as a guarantee that every judgment is undisputed.

  3. Knowl 3 — False presuppositions span explicit errors and subtle assumptions

    empirical result

    In a random sample of 50 development questions labeled with false presuppositions, the authors identified several recurring forms: false predicates (30%), such as assuming that current is stored in power plants; false properties (22%), such as assuming cement blocks are strong; false causal or other relationships between facts (22%), such as assuming Newton's third-law forces cancel on the same object; false clauses (14%), such as assuming water must reach 100 degrees to become steam; and false existential presuppositions (6%), such as assuming that part of a hard disk's advertised capacity is unusable. The remaining sampled cases were exceptions (4%) or contained no false presupposition / an annotation error (2%). The categories show that the task includes both locally explicit claims and implicit properties or relationships that must be inferred from the question.

  4. Knowl 4 — Detection is substantially harder without the annotation comment

    empirical result

    Detection is evaluated with macro-F1. In the main track, the c-REALM retrieval model paired with a multi-passage RoBERTa classifier scored 66.3 on the test set; self-labeling with 196,385 unlabeled question-comment pairs raised test macro-F1 to 67.1. In the GOLD-COMMENT track, the RoBERTa classifier given both question and comment scored 75.6, compared with 66.9 for question-only and 68.6 for comment-only. The nearest-neighbor baseline scored 54.1, and zero-shot models trained on MNLI and BoolQ scored 54.2 and 58.2, respectively. Human macro-F1 was 85.1 with the most-upvoted comment and 70.9 without it. These results show that existing models can detect some false presuppositions, but access to the response containing corrective evidence makes detection much easier.

  5. Knowl 5 — Retrieval is a major bottleneck for detecting false presuppositions

    empirical result

    The main-track detection system retrieves the top five English Wikipedia passages with c-REALM and classifies the question together with those passages. Error analysis found that retrieved passages often concerned the right topic but did not establish whether the question's background assumption was true or false. In the paper's analysis of false negatives, 86% were attributed to retrieval misses: failing to retrieve the relevant topic (42%), retrieving relevant-topic evidence unrelated to the false presupposition (32%), or retrieving evidence that was related but not direct enough (12%). A separate validation-set analysis of 50 false positives and 50 false negatives likewise found retrieval failures common: for false negatives, 32% lacked a related topic, 40% had a similar topic but insufficient evidence, and 12% required reasoning over indirect evidence. Thus topical relevance alone does not provide the evidence needed to verify a presupposition.

  6. Knowl 6 — Writing baselines compare separate and unified generation of the presupposition and correction

    model/method

    For the writing task, systems receive a question known to contain a false presupposition and generate both the presupposition and its correction. The GOLD-COMMENT systems also receive the annotation comment; the main-track systems instead use retrieved passages, with a Fusion-in-Decoder T5 model to process multiple passages. Dedicated T5-base models generate the presupposition and correction separately. A unified T5-base model shares one generator: it trains on corrections and on presuppositions rewritten with the prefix “It is not the case that”. At inference, it generates a correction normally; to generate a presupposition, constrained decoding requires that prefix and the text following it is returned as the presupposition. The comparison tests whether joint training helps the two related generation tasks.

  7. Knowl 7 — Writing systems score better with comments, but corrections remain difficult

    empirical result

    On the test set, every trained generator outperformed the copy baseline, and systems given the annotation comment generally scored better than systems relying on retrieved passages. Using unigram F1, the GOLD-COMMENT dedicated model scored 47.7 for presupposition generation and 38.0 for correction generation (42.9 average); its unified counterpart scored 49.2 and 31.4 (40.3 average). The main-track unified model scored 46.3 and 28.4 (37.4 average). The same pattern appears in BLEU: the GOLD-COMMENT dedicated model scored 22.9 for presuppositions and 16.1 for corrections, while the main-track unified model scored 23.6 and 8.3. Models were generally better at writing the presupposition, which can often be extracted from the question, than at writing a factually supported correction. Sentence-BERT scores were comparatively high, but the authors caution that semantic entailment similarity is necessary, not sufficient, for a satisfactory correction.

  8. Knowl 8 — Human ratings expose factual and explanatory weaknesses hidden by automatic metrics

    empirical result

    Two raters evaluated system generations for 200 randomly sampled test questions on a 0–3 scale for fluency, presupposition validity, correction quality, and consistency between the two generated statements. Mean scores (fluency, presupposition, correction, consistency) were: GOLD-COMMENT dedicated, 2.9, 1.8, 1.9, 1.6; GOLD-COMMENT unified, 3.0, 2.0, 0.8, 2.8; main-track dedicated, 2.8, 1.8, 0.6, 1.6; and main-track unified, 3.0, 1.8, 0.6, 2.8. The reference annotations scored 2.9, 2.8, 2.7, and 2.9. Systems produced fluent text and often stayed on the right topic, but their corrections were much weaker than the references: a correction needed to be factually correct and provide enough justification, rather than merely negate the presupposition. The dedicated GOLD-COMMENT model had the strongest system correction rating, while unified models had stronger consistency ratings.

  9. Knowl 9 — Judgments remain ambiguous even for people with access to evidence

    limitation

    The validity of a presupposition can depend on how a question is interpreted, the questioner's background, and whether correcting the assumption is important to the information need. In an analysis of 44 human errors made without the most-upvoted comment, 34.1% involved ambiguity, 9.1% disagreement over whether correcting the assumption was critical, and 22.7% inconsistent information across web sources; other cases involved failure to find evidence (11.4%), labeling mistakes (11.4%), or an incorrect reference label (11.4%). Using the most-upvoted comment aggregates community judgments and supplies a correction, but does not eliminate these uncertainties. The paper also notes that CREPE is sourced from Reddit forums and that its exploration of large language models was a small-scale case study rather than a systematic benchmark evaluation.

  10. Knowl 10 — A small QASPER sample suggests false presuppositions also occur in expert research questions

    empirical result

    The authors examined 50 randomly sampled unanswerable questions from QASPER, which contains information-seeking questions about NLP research papers written by NLP experts. Two of the 50 sampled questions turned out to have valid answers; among the remaining 48 unanswerable questions, 25% contained false presuppositions. Examples included asking why a paper used RNNs instead of transformers when it did not use transformers, and asking about annotators when the dataset was constructed automatically. This is exploratory evidence that false presuppositions are not limited to general-interest forums. The authors describe the estimate as a strict lower bound because identifying these cases requires NLP expertise and the sample was not annotated by multiple experts.

Coverage note — The paper's small prompted GPT-3/ChatGPT case studies and detailed train–test paraphrase audit are omitted because they are exploratory analyses rather than central benchmark or baseline contributions.

References

  1. 1.Akari Asai and Eunsol Choi. 2021. Challenges in information-seeking QA: Unanswerable questions and paragraph retrieval. In Proceedings of the Association for Computational Linguistics.
  2. 2.FairScale authors. 2021. Fairscale: A general purpose modular pytorch library for high performance and large scale training. https://github.com/facebookresearch/fairscale.
  3. 3.David I. Beaver, Bart Geurts, and Kristie Denlinger. 2014. Presupposition. The Stanford Encyclopedia of Philosophy.
  4. 4.Rahul Bhagat and Eduard Hovy. 2013. Squibs: What is a paraphrase? Computational Linguistics.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of Advances in Neural Information Processing Systems.
  6. 6.Isabel Cachola, Eric Holgate, Daniel Preo¸tiuc-Pietro, and Junyi Jessy Li. 2018. Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media. In Proceedings of International Conference on Computational Linguistics.
  7. 7.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020. Evaluation of text generation: A survey. arXiv preprint.
  8. 8.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wentau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018. QuAC: Question answering in context. In Proceedings of Empirical Methods in Natural Language Processing.
  9. 9.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  10. 10.Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  11. 11.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2021. Autoregressive entity retrieval. In Proceedings of the International Conference on Learning Representations.
  12. 12.Marie Duží and Martina Cíhalová. 2015. Questions, answers, and presuppositions. Computación y Sistemas.
  13. 13.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the Association for Computational Linguistics. Association for Computational Linguistics.
  14. 14.Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2013. PPDB: The paraphrase database. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  15. 15.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the European Chapter of the Association for Computational Linguistics.
  16. 16.Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020. Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition. In Proceedings of the Association for Computational Linguistics.
  17. 17.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data.
  18. 18.Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. 2018. A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  19. 19.S. Jerrold Kaplan. 1978. Indirect responses to loaded questions. In Theoretical Issues in Natural Language Processing-2.
  20. 20.Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, and Deepak Ramachandran. 2021. Which linguist invented the lightbulb? presupposition verification for question-answering. In Proceedings of the Association for Computational Linguistics.
  21. 21.Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to progress in long-form question answering. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  22. 22.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics.
  23. 23.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the Association for Computational Linguistics. Association for Computational Linguistics.
  24. 24.Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021. Question and answer test-train overlap in open-domain question answering datasets. In Proceedings of the European Chapter of the Association for Computational Linguistics.
  25. 25.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint.
  27. 27.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed precision training. In Proceedings of the International Conference on Learning Representations.
  28. 28.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of Empirical Methods in Natural Language Processing.
  29. 29.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint.
  30. 30.Ellie Pavlick, Pushpendre Rastogi, Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2015. PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification. In Proceedings of the Association for Computational Linguistics.
  31. 31.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  32. 32.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation.
  33. 33.Christopher Potts. 2015. Presupposition and implicature. The handbook of contemporary semantic theory.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research.
  35. 35.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the Association for Computational Linguistics.
  36. 36.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of Empirical Methods in Natural Language Processing.
  37. 37.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of Empirical Methods in Natural Language Processing.
  38. 38.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of Empirical Methods in Natural Language Processing.
  39. 39.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the Association for Computational Linguistics.
  40. 40.John Schulman, Barret Zoph, Christina Kim, Jacob Hilton, Jacob Menick, Jiayi Weng, Juan Uribe, Liam Fedus, Luke Metz, et al. 2022. ChatGPT: Optimizing language models for dialogue.
  41. 41.Robert Stalnaker, Milton K Munitz, and Peter Unger. 1977. Pragmatic presuppositions. In Proceedings of the Texas conference on per˜ formatives, presuppositions, and implicatures.
  42. 42.Peter F Strawson. 1950. On referring. Mind.
  43. 43.Ellen M. Voorhees and Dawn M. Tice. 2000. The TREC-8 question answering track. In Proceedings of the Language Resources and Evaluation Conference.
  44. 44.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics.
  45. 45.Fangyuan Xu, Junyi Jessy Li, and Eunsol Choi. 2022. How do we answer complex questions: Discourse structure of long-form answers. In Proceedings of the Association for Computational Linguistics.
  46. 46.Michael Zhang and Eunsol Choi. 2021. SituatedQA: Incorporating extra-linguistic contexts into QA. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7371–7387, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.

Citation

MLA
Yu, X., et al. “CREPE: Open-Domain Question Answering with False Presuppositions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 10457–80, https://doi.org/10.18653/v1/2023.acl-long.583.
APA
Yu, X., Min, S., Zettlemoyer, L., & Hajishirzi, H. (2023). CREPE: Open-Domain Question Answering with False Presuppositions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10457–10480. https://doi.org/10.18653/v1/2023.acl-long.583
Chicago
Yu, X., S. Min, L. Zettlemoyer, and H. Hajishirzi. 2023. “CREPE: Open-Domain Question Answering with False Presuppositions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10457–80. https://doi.org/10.18653/v1/2023.acl-long.583.
Harvard
Yu, X. et al. (2023) “CREPE: Open-Domain Question Answering with False Presuppositions”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10457–10480. Available at: https://doi.org/10.18653/v1/2023.acl-long.583.
Vancouver
1. Yu X, Min S, Zettlemoyer L, Hajishirzi H (2023) CREPE: Open-Domain Question Answering with False Presuppositions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10457–10480

BibTeX

@inproceedings{yu-etal-2023-crepe,
    title = "{CREPE}: Open-Domain Question Answering with False Presuppositions",
    author = "Yu, Xinyan  and
      Min, Sewon  and
      Zettlemoyer, Luke  and
      Hajishirzi, Hannaneh",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.583/",
    doi = "10.18653/v1/2023.acl-long.583",
    pages = "10457--10480"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/