(QA)²: Question Answering with Questionable Assumptions

Najoung KimPhu Mon HtutSamuel R. BowmanJackson Petty

article2023ACL61 citations

Introduces the (QA)² benchmark of naturally occurring search queries to evaluate how well language models identify false or unverifiable premises and appropriately correct them rather than generating misleading answers.

Listen

Real-world information-seeking queries frequently incorporate questionable assumptions, which are premises that are either factually false or unverifiable. When users ask questions with flawed premises, typical automated question answering systems often endorse the false premise by delivering a direct answer rather than identifying and correcting the underlying error. The article addresses this challenge by systematically evaluating how robust modern artificial intelligence systems are when responding to natural user questions containing invalid presuppositions.

To benchmark model capabilities, the article introduces a specialized evaluation dataset containing 602 natural, search-engine-derived questions balanced evenly between valid questions and those containing questionable assumptions. The study evaluates a range of leading language models across three tasks: complete question answering evaluated by human judges, assumption detection to identify problematic premises, and assumption verification to fact-check isolated assumptions against real-world facts.

The findings show that handling flawed assumptions remains a significant hurdle for automated systems. Even top-performing models achieved acceptable answers on only 56% of questions in full end-to-end evaluations, with performance dropping noticeably when queries contained questionable assumptions compared to valid ones. On binary classification tasks, models struggled to identify whether a question contained an invalid assumption, performing near chance at roughly 50% to 64% accuracy. In contrast, systems performed better at verifying isolated factual statements, achieving up to 72% accuracy when an oracle extracted the assumption beforehand. Crucially, assumption detection ability correlated significantly with full question-answering success, whereas raw factual verification did not.

These results indicate that the primary operational bottleneck is not merely a lack of factual knowledge, but rather a failure to recognize implicit assumptions embedded within user prompts. In high-stakes applications, automated systems that blindly accept user premises risk propagating misinformation, posing compliance and reputational hazards. Decision-makers should not rely on off-the-shelf language models for unmonitored information retrieval without implementing explicit assumption detection guardrails.

Organizations developing or deploying automated answering solutions should prioritize assumption detection as an automated screening metric during model development. Future work must focus on methods that reliably identify implicit premises before generating responses. System evaluators should also consider temporal limitations, as facts evolve over time, and account for the inherent difficulty of proving non-existence during factual verification.

Cover for (QA)²: Question Answering with Questionable Assumptions

Abstract

Naturally occurring information-seeking questions often contain questionable assumptions—assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question When did Marie Curie discover Uranium? cannot be answered as a typical when question without addressing the false assumption Marie Curie discovered Uranium. In this work, we propose (QA)² (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)², systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)², we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.

Table of Contents

  • 1 Introduction
  • 2 Questions with Questionable Assumptions
  • 3 (QA)2 Dataset
  • 3.1 Data Collection and Annotation
  • 3.1.1 Step 1: Question Collection
  • 3.1.2 Step 2: Identification of Questionable Assumptions via Crowdsourcing
  • 3.1.3 Step 3: Expert Annotation
  • 3.2 Dataset Statistics
  • 3.3 Discussion
  • 4 Evaluation
  • 4.1 End-to-End Abstractive QA
  • 4.2 Questionable Assumption Detection
  • 4.3 Questionable Assumption Verification
  • 5 Experiments
  • 5.1 Models
  • 5.1.1 QA-Specific Models
  • 5.1.2 General-Purpose Language Models
  • 6 Results and Discussion
  • 6.1 End-to-End Abstractive QA
  • 6.2 Subtasks
  • 6.3 Error Analysis
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A ChatGPT Responses
  • B Canary String
  • C Dataset Fields
  • D Prompts
  • E End-to-End QA Error Examples
  • F License and Terms for Use
  • G Model and Implementation Details
  • H Crowdsourcing Details

Knowls

  1. Knowl 1 — Questionable assumptions are false or unverifiable propositions attributed to the asker

    definition

    A questionable assumption is a proposition associated with a question that is false or cannot be verified and is likely to be believed by the person asking the question. This definition covers both linguistic presuppositions and propositions that are not strictly presupposed but reflect an epistemic bias. For example, When did Marie Curie discover Uranium? carries the false proposition that Marie Curie discovered Uranium. A direct answer in the requested form—such as a date—would endorse that proposition; an adequate response must instead address the failure, for example by correcting it or explaining why it cannot be verified.

  2. Knowl 2 — The dataset balances questionable-assumption questions with ordinary information-seeking questions

    data/table

    The open-domain (QA)² dataset contains 602 naturally occurring, expert-annotated questions: 301 questions judged to contain a questionable assumption and 301 judged valid, meaning they contain no questionable assumption. The 301 questionable assumptions comprise 246 false assumptions and 55 unverifiable assumptions. The 32-item adaptation set contains 16 questions of each class; the 570-item evaluation set contains 285 of each class. The adaptation set is intended for few-shot tuning or in-context examples, while the larger split is designed for evaluation rather than full model training. Each example includes an abstractive answer, supporting extractive evidence and its URL, the assumption label, the questionable assumption where applicable, a yes/no verification question, and an indication of whether the assumption's status can change over time.

  3. Knowl 3 — Autocomplete queries and expert review produced the final dataset

    model/method

    The authors collected naturally occurring English wh-questions from Google's autocomplete API, using prefixes derived from Natural Questions (NQ) questions beginning with when, where, which, how, what, why, who, or whose and extending each prefix to the first noun phrase. They scraped 12,000 candidate questions and used automatic filtering to remove duplicates and likely non-questions or unsuitable results. Five crowdworkers per question marked whether it contained a false or unverifiable assumption; workers could consult information sources such as search engines. Three expert annotators then reviewed questions flagged by at least one worker, determined whether an assumption was questionable, wrote an abstractive answer, supplied online evidence and its URL, and—when an assumption was questionable—formulated it as a yes/no verification question. Annotation disagreements were verified and adjudicated. Final selection required evidence to support the answer and no immediate ambiguity in the question's interpretation.

  4. Knowl 4 — The benchmark separates answering, assumption detection, and assumption verification

    experimental setup

    The benchmark defines three evaluation tasks. End-to-end abstractive QA takes the original question and requires an adequate answer, which may correct or otherwise address a questionable assumption rather than simply match the question's requested answer type. Because many responses can be acceptable, raters judged model answers against the question with the expert answer available as a reference; five raters judged each of 100 sampled evaluation questions, and majority vote determined acceptability. Questionable-assumption detection asks whether the original question contains any questionable assumption and is scored as binary classification. Questionable-assumption verification supplies an oracle-generated yes/no question that explicitly states the assumption, thereby removing the need to detect or formulate it; for valid questions, a randomly selected true assumption was converted into a yes/no question. Detection and verification are scored by accuracy over the evaluation set.

  5. Knowl 5 — Model and prompting conditions for the benchmark comparison

    experimental setup

    The evaluation compared QA-specialized systems (Macaw-11B and REALM) with general-purpose models: T0pp, Flan-T5-XXL, GPT-3 variants (davinci, text-davinci-002, and text-davinci-003), PaLM 540B, and Flan-PaLM 540B. Models were tested zero-shot; selected models were also tested with four label-balanced in-context demonstrations from the adaptation set. The authors additionally tested step-by-step prompting with task decomposition on text-davinci-003, PaLM, and Flan-PaLM, and few-shot T-Few tuning for detection and verification using 16 examples plus 16 development examples. Maximum generation length was 512 tokens. The tested models' training data ended between December 2018 and June 2021, whereas the dataset included information current through December 2022.

  6. Knowl 6 — End-to-end and binary-task results across evaluated models

    data/table

    The table reports end-to-end answer acceptability as a proportion, separately for all sampled questions, questions with questionable assumptions, and valid questions. Detection and verification are accuracy proportions on the full evaluation set. End-to-end results use human judgments on 100 questions; binary-task results use the 570-question evaluation split. The strongest end-to-end score is 0.56 for in-context text-davinci-003, while the best detection and verification scores are 0.64 for in-context PaLM 540B and 0.72 for in-context Flan-PaLM 540B, respectively. T-Few was evaluated only on the binary tasks.

    Setting Model E2E all E2E questionable E2E valid Detection Verification
    Zero-shot Macaw-11B 0.21 0.16 0.26 0.49 0.57
    Zero-shot REALM 0.18 0.10 0.26 0 0.01
    Zero-shot T0pp 0.14 0.02 0.26 0.49 0.58
    Zero-shot Flan-T5-XXL 0.16 0.10 0.22 0.50 0.64
    Zero-shot davinci 0.28 0.18 0.38 0.51 0.58
    Zero-shot text-davinci-002 0.47 0.40 0.54 0.46 0.60
    Zero-shot text-davinci-003 0.48 0.40 0.56 0.52 0.69
    Zero-shot PaLM 540B 0.21 0.14 0.28 0.51 0.70
    Zero-shot Flan-PaLM 540B 0.28 0.08 0.48 0.50 0.70
    In-context Flan-T5-XXL 0.16 0.12 0.20 0.50 0.64
    In-context text-davinci-003 0.56 0.50 0.62 0.59 0.65
    In-context PaLM 540B 0.44 0.40 0.48 0.64 0.71
    In-context Flan-PaLM 540B 0.46 0.44 0.48 0.63 0.72
    Step-by-step / task decomposition text-davinci-003 0.45 0.36 0.54 0.54 –
    Step-by-step / task decomposition PaLM 540B 0.44 0.34 0.54 0.59 –
    Step-by-step / task decomposition Flan-PaLM 540B 0.43 0.44 0.42 0.54 –
    Few-shot T-Few – – – 0.53 0.48
  7. Knowl 7 — Detection performance tracks end-to-end answer quality more closely than verification

    empirical result

    Across evaluated systems, questionable-assumption detection accuracy had a positive correlation with human-rated end-to-end QA acceptability (Spearman's ρ=0.58\rho = 0.58, p=0.02p = 0.02). The correlation between verification accuracy and end-to-end acceptability was not statistically significant (ρ=0.45\rho = 0.45, p=0.12p = 0.12). Verification generally scored better than detection, with in-context Flan-PaLM reaching 0.72 verification accuracy, but this did not reliably translate into better end-to-end answers. The authors therefore identify recognizing assumptions in the original question, rather than only checking an explicitly supplied proposition, as a likely bottleneck; they propose detection as a useful model-selection proxy when human end-to-end evaluation is unavailable.

  8. Knowl 8 — Most annotated assumptions attach to wh-words or definite descriptions

    empirical result

    In the authors' analysis of questionable assumptions, 77% were associated with the question's wh-word, while 15% were associated with a definite description. The definite-description cases included existence assumptions (10% of the analyzed assumptions) and contextual-uniqueness assumptions (5%). These patterns indicate that problematic propositions in naturally occurring queries often arise from the linguistic structure that frames the requested information, not only from conspicuously malformed questions.

  9. Knowl 9 — End-to-end errors were mainly factual, with a smaller underinformativeness category

    empirical result

    A manual analysis of 100 predictions from in-context text-davinci-003 and PaLM found two main error categories. For text-davinci-003, 77% of analyzed errors involved factual mistakes and 23% involved answers that were uninformative, underinformative, or misleading. For PaLM, the corresponding proportions were 72% and 18%. Thus, even when systems attempted to answer questions in the benchmark setting, factual reliability was the dominant identified error type.

  10. Knowl 10 — Labels can become outdated, and verifying nonexistence is intrinsically difficult

    limitation

    The dataset labels reflect facts available at annotation time, so an assumption classified as unverifiable or false can become valid later as events occur or information is released. The authors also note that establishing that something did not happen or does not exist is difficult because annotators cannot exhaustively inspect all relevant online material. In some cases, they used pragmatic inferences from omissions in available sources as evidence; because this is heuristic, the dataset may contain annotation errors.

Coverage note — No substantial contributed material was omitted; appendix prompt templates and implementation details were not given separate knowls because they support, rather than add to, the central dataset and evaluation contributions.

References

  1. 1.Akari Asai and Eunsol Choi. 2021. Challenges in information-seeking QA: Unanswerable questions and paragraph retrieval. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1492–1504, Online. Association for Computational Linguistics.
  2. 2.David Beaver. 1997. Presupposition. In Handbook of logic and language, pages 939–1008. Elsevier.
  3. 3.Johan Brandtler. 2008. Why we should ever bother about wh-questions: On NPI-licensing properties of wh-questions in Swedish. Working Papers in Scandinavian Syntax, 82:83–121.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Jifan Chen, Eunsol Choi, and Greg Durrett. 2021. Can NLI models verify QA systems’ predictions? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3841–3854, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arxiv:2204.02311.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv:2210.11416.
  8. 8.Andre Cianflone, Yulan Feng, Jad Kabbara, and Jackie Chi Kit Cheung. 2018. Let’s do it “again”: A first computational approach to detecting adverbial presupposition triggers. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2747–2755, Melbourne, Australia. Association for Computational Linguistics.
  9. 9.David Clausen and Christopher D. Manning. 2009. Presupposed content and entailments in natural language inference. In Proceedings of the 2009 Workshop on Applied Textual Inference (TextInfer), pages 70–73, Suntec, Singapore. Association for Computational Linguistics.
  10. 10.Marie-Catherine de Marneffe, Scott Grimm, and Christopher Potts. 2009. Not a simple yes or no: Uncertainty in indirect answers. In Proceedings of the SIGDIAL 2009 Conference, pages 136–143, London, UK. Association for Computational Linguistics.
  11. 11.Aviad Eilam and Catherine Lai. 2009. Sorting out the implications of questions. Universitat Autonoma de Barcelona. Talk at CONSOLE XVIII.
  12. 12.Terry Gaasterland, Parke Godfrey, and Jack Minker. 1992. An overview of cooperative answering. Journal of Intelligent Information Systems, 1:123–157.
  13. 13.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  14. 14.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3929–3938. PMLR.
  15. 15.Pauline Jacobson. 2016. The short answer: Implications for direct compositionality (and vice versa). Language, pages 331–375.
  16. 16.Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020. Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8690–8705, Online. Association for Computational Linguistics.
  17. 17.Nanjiang Jiang and Marie-Catherine de Marneffe. 2021. He thinks he knows better than the doctors: BERT for event factuality fails on pragmatics. Transactions of the Association for Computational Linguistics, 9:1081–1097.
  18. 18.Aravind Joshi, Bonnie Webber, and Ralph M. Weischedel. 1984. Preventing false inferences. In 10th International Conference on Computational Linguistics and 22nd Annual Meeting of the Association for Computational Linguistics, pages 134–138, Stanford, California, USA. Association for Computational Linguistics.
  19. 19.S. Jerrold Kaplan. 1982. Cooperative responses from a portable natural language query system. Artificial Intelligence, 19(2):165–187.
  20. 20.Lauri Karttunen. 1974. Presupposition and linguistic context. Theoretical Linguistics, 1(1-3):181–194.
  21. 21.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UNIFIEDQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online. Association for Computational Linguistics.
  22. 22.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  23. 23.Najoung Kim, Roma Patel, Adam Poliak, Patrick Xia, Alex Wang, Tom McCoy, Ian Tenney, Alexis Ross, Tal Linzen, Benjamin Van Durme, Samuel R. Bowman, and Ellie Pavlick. 2019. Probing what different NLP tasks teach machines about function word comprehension. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019), pages 235–249, Minneapolis, Minnesota. Association for Computational Linguistics.
  24. 24.Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, and Deepak Ramachandran. 2021. Which Linguist Invented the Lightbulb? Presupposition Verification for Question-Answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3932–3945, Online. Association for Computational Linguistics.
  25. 25.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  26. 26.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics, 7:452–466.
  27. 27.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  28. 28.Stephen C. Levinson. 1983. Presupposition. In Pragmatics, Cambridge Textbooks in Linguistics, page 167–225. Cambridge University Press.
  29. 29.David Lewis. 1979. Scorekeeping in a language game. In Rainer Bäuerle, Urs Egli, and Arnim von Stechow, editors, Semantics from Different Points of View, pages 172–187. Springer Berlin Heidelberg, Berlin, Heidelberg.
  30. 30.Haokun Liu, Derek Tam, Muqeeth Mohammed, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In Advances in Neural Information Processing Systems.
  31. 31.Annie Louis, Dan Roth, and Filip Radlinski. 2020. “I’d rather just go to bed”: Understanding indirect answers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7411–7425, Online. Association for Computational Linguistics.
  32. 32.Randall Munroe. 2020. Learning new things from Google. https://twitter.com/xkcd/status/1333529967079120896. Accessed: 2022-12-15.
  33. 33.Jungsoo Park, Sewon Min, Jaewoo Kang, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. FaVIQ: FAct verification from information-seeking questions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5154–5166, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, and Tal Linzen. 2021. NOPE: A corpus of naturally-occurring presuppositions in English. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 349–366, Online. Association for Computational Linguistics.
  35. 35.Christopher Potts. 2015. Presupposition and implicature. In The Handbook of Contemporary Semantic Theory, chapter 6, pages 168–202. John Wiley & Sons, Ltd.
  36. 36.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  37. 37.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  38. 38.Maribel Romero and Chung-hye Han. 2004. On negative yes/no questions. Linguistics and Philosophy, 27(5):609–658.
  39. 39.Alexis Ross and Ellie Pavlick. 2019. How well do NLI models capture verb veridicality? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2230–2240, Hong Kong, China. Association for Computational Linguistics.
  40. 40.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  41. 41.Scott Soames. 1976. A critical examination of Frege’s theory of presupposition and contemporary alternatives. Ph.D. thesis, Massachusetts Institute of Technology.
  42. 42.Robert Stalnaker. 1974. Pragmatic presuppositions. In M. Munitz and P. Unger, editors, Semantics and Philosophy, page 197–214. New York University Press.
  43. 43.Peter F. Strawson. 1964. Identifying reference and truth-values. Theoria, 30(2):96–118.
  44. 44.Oyvind Tafjord and Peter Clark. 2021. General-purpose Question-Answering with Macaw. arXiv:2109.02593.
  45. 45.Galina Tremper and Anette Frank. 2011. Extending fine-grained semantic relation classification to presupposition relations between verbs. Bochumer Linguistische Arbeitsberichte.
  46. 46.Wolfgang Wahlster, Heinz Marburger, Anthony Jameson, and Stephan Busemann. 1983. Over-answering yes-no questions: Extended responses in a nl interface to a vision system. In Proceedings of the Eighth international joint conference on Artificial intelligence, volume 2, pages 643–646.
  47. 47.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  48. 48.Xinyan Velocity Yu, Sewon Min, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. CREPE: Open-domain question answering with false presuppositions. arXiv:2211.17257.

Citation

MLA
Kim, N., et al. “(QA)2: Question Answering with Questionable Assumptions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 8466–87, https://doi.org/10.18653/v1/2023.acl-long.472.
APA
Kim, N., Htut, P. M., Bowman, S. R., & Petty, J. (2023). (QA)2: Question Answering with Questionable Assumptions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8466–8487. https://doi.org/10.18653/v1/2023.acl-long.472
Chicago
Kim, N., P. M. Htut, S. R. Bowman, and J. Petty. 2023. “(QA)2: Question Answering with Questionable Assumptions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8466–87. https://doi.org/10.18653/v1/2023.acl-long.472.
Harvard
Kim, N. et al. (2023) “(QA)2: Question Answering with Questionable Assumptions”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8466–8487. Available at: https://doi.org/10.18653/v1/2023.acl-long.472.
Vancouver
1. Kim N, Htut PM, Bowman SR, Petty J (2023) (QA)2: Question Answering with Questionable Assumptions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8466–8487

BibTeX

@inproceedings{kim-etal-2023-qa,
    title = "({QA})$^2$: Question Answering with Questionable Assumptions",
    author = "Kim, Najoung  and
      Htut, Phu Mon  and
      Bowman, Samuel R.  and
      Petty, Jackson",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.472/",
    doi = "10.18653/v1/2023.acl-long.472",
    pages = "8466--8487"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/