Selectively Answering Ambiguous Questions

Jeremy R. ColeMichael J. Q. ZhangDaniel GillickJulian EisenschlosBhuwan DhingraJacob Eisenstein

article2023EMNLP81 citations

Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.

Listen

Deploying large language models in customer-facing and decision-critical question answering systems carries significant risk when systems generate incorrect or misleading answers instead of abstaining. This challenge is intensified by real-world user queries that are frequently underspecified or context-dependent. Prior research largely addressed epistemic uncertainty—cases where the question is clear but the model may lack factual knowledge—while overlooking denotational uncertainty, where the user's intent or meaning is inherently ambiguous.

The article investigates how to reliably calibrate language models so they can selectively answer questions they understand with high accuracy and abstain when uncertain, specifically under conditions of heavy query ambiguity.

To address this, the authors evaluated few-shot prompting approaches using the Pathways Language Model (PaLM) across multiple benchmarks, including Natural Questions, TriviaQA, AmbigQA, and SituatedQA. Rather than relying on standard token likelihoods or asking models to verify their own outputs, the researchers introduced a two-step framework: the model first attempts to state an explicit interpretation of the question, after which candidate answers are generated. Model confidence was measured by drawing repeated samples (up to 10) and calculating sample repetition—the frequency with which sampled outputs match the primary answer—and sample diversity.

The investigation yielded several critical findings. First, sampling repetition proved to be the most reliable indicator of answer correctness across all scenarios, significantly outperforming standard model likelihood and self-verification prompting. Second, while standard likelihood scores performed adequately on clear questions, their calibration degraded severely on ambiguous queries, whereas sampling repetition maintained robust calibration. Third, instruction tuning improved raw accuracy but severely distorted the model's likelihood calibration; however, applying sampling-based confidence scoring restored calibration and substantially increased the volume of questions answered at high accuracy thresholds. Finally, directly predicting whether a question is ambiguous achieved poor accuracy (around 58%), indicating that ambiguity is better managed by generating explicit interpretations rather than binary classification.

These findings indicate that relying on raw model probability or self-prompted verification creates an unsafe, false sense of confidence in production environments, particularly for instruction-tuned models handling complex, ambiguous user prompts. Adopting a sampling-based confidence mechanism provides a practical method to safeguard performance and maintain strict accuracy standards (such as an 80% accuracy threshold) before returning answers to users.

Organizations implementing generative question-answering systems should adopt a disambiguate-then-answer strategy paired with sample repetition confidence scoring to govern abstention policies. However, decision-makers must weigh the trade-off of inference cost, as sampling multiple outputs increases compute requirements linearly. For budget-sensitive applications, smaller sample counts (such as three to five samples) or sample diversity metrics offer a viable, lower-cost compromise.

Readers should note that the evaluation was conducted solely on the PaLM model family within a closed-book setting without real-time external retrieval. Further testing across alternative model architectures and retrieval-augmented pipelines is recommended before full operational rollout.

Cole et al (2023).pdf
Cover for Selectively Answering Ambiguous Questions

Abstract

Trustworthy language models should abstain from answering questions when they do not know the answer. However, the answer to a question can be unknown for a variety of reasons. Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context. We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous. In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work. We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning. Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Calibration for Question Answering
  • 3 Confidence Scores
  • 4 Evaluation setup
  • 4.1 Metrics
  • 5 Unambiguous Questions
  • 5.1 Datasets
  • 5.2 Experiment Setup
  • 5.3 Calibration Results
  • 6 Ambiguous Questions
  • 6.1 Datasets
  • 6.2 Experiment Setup
  • 6.3 Ambiguity Prediction
  • 6.4 Calibration Results
  • 7 Instruction Tuning and Scaling
  • 8 Related Work
  • 8.1 Calibration and Selective QA
  • 8.2 Self-Calibrating Language Models
  • 8.3 Prompting Strategies
  • 9 Conclusion
  • 10 Limitations
  • Acknowledgements
  • References
  • A Ambiguity Prediction
  • B Chain of Thought
  • C Text as Calibration
  • D Scaling
  • E Sampling

Knowls

  1. Knowl 1 — Question answering uncertainty has denotational and epistemic sources

    definition

    For a question-answering system, denotational uncertainty is uncertainty about which interpretation, or denotation, a user query expresses; epistemic uncertainty is uncertainty about the correct answer once an interpretation is fixed. In an idealized system that represents interpretations explicitly, the answer distribution can be written as P(a∣q)=∑dP(a∣d)P(d∣q)P(a|q)=\sum_d P(a|d)P(d|q), where qq is the user query, dd ranges over its possible interpretations, and aa is a possible answer. Thus, uncertainty about answers to an ambiguous query can reflect uncertainty about its interpretation, its answer under an interpretation, or both. The paper applies this distinction to language models that do not explicitly construct formal denotations, using natural-language disambiguations as an approximate intermediate representation.

  2. Knowl 2 — Repeated outputs provide two sampling-based confidence scores

    model/method

    The proposed disambiguate-then-answer approach has a model first express an interpretation of an ambiguous question in natural language and then answer it. Confidence is estimated from repeated model outputs, thereby reflecting variation in both interpretations and answers rather than relying only on the likelihood of one generated sequence.

    For the experiments, the authors generated n=10n=10 sampled outputs at temperature 0.50.5 and compared answers after lowercasing and removing punctuation. Let gg be the greedy answer, generated at temperature 00, let xix_i be the normalized answer from sample ii, and let uu be the number of distinct normalized sampled answers. The two confidence scores are crep=∣{i:xi=g}∣/nc_{rep}=|\{i:x_i=g\}|/n (sampling repetition) and cdiv=1−u/nc_{div}=1-u/n (sampling diversity). Higher values indicate greater confidence. Repetition measures how often sampling recovers the greedy answer; diversity measures the concentration of samples into repeated answers. The diversity score depends on the number of samples, and the authors note that the number of unique answers need not grow linearly with sample count.

  3. Knowl 3 — Sampling scores improve calibration on ambiguous question answering

    empirical result

    On ambiguous question answering with PaLM, the authors compared sequence likelihood (L), sampling diversity (D), and sampling repetition (R). They used 10 samples at temperature 0.50.5 for sampling scores and excluded self-verification because its calibration was poor on unambiguous questions and its prompt was difficult to define for questions with multiple interpretations. Loose matching counts an answer as correct if it matches an answer to any reference interpretation; strict matching requires a match to an answer for the closest interpretation. EM is answer exact match; ROC-AUC measures ranking of correct versus incorrect answers; ECE is expected calibration error, for which lower is better; C@80 is the maximum coverage among the highest-confidence answers while maintaining at least 80% accuracy. Higher ROC-AUC and C@80 are better.

    The table shows that repetition has the highest loose-match ROC-AUC and C@80 on all three datasets, while diversity often has lower ECE. Ambiguous QA has lower exact match and is more difficult to calibrate than the unambiguous setting. Strict matching lowers EM and changes some metric rankings, but does not reverse the broad advantage of sampling over likelihood.

    Dataset Method Loose EM ROC-AUC ECE C@80 Strict EM ROC-AUC ECE C@80
    AmbigQA L 44.8 0.731 0.316 17.4 41.9 0.724 0.287 12.4
    AmbigQA D 44.8 0.767 0.114 16.1 41.9 0.763 0.085 13.0
    AmbigQA R 44.8 0.821 0.120 26.2 41.9 0.812 0.147 20.6
    SQA-Temp L 35.7 0.757 0.223 9.4 29.1 0.751 0.157 3.0
    SQA-Temp D 35.7 0.772 0.086 6.7 29.1 0.757 0.048 2.7
    SQA-Temp R 35.7 0.797 0.119 13.5 29.1 0.784 0.167 2.4
    SQA-Geo L 35.6 0.759 0.221 10.4 29.0 0.757 0.156 3.3
    SQA-Geo D 35.6 0.743 0.085 8.9 29.0 0.723 0.056 4.5
    SQA-Geo R 35.6 0.800 0.120 13.7 29.0 0.789 0.176 4.6
  4. Knowl 4 — Calibration results on unambiguous questions

    empirical result

    The authors evaluated PaLM on unambiguous Natural Questions (NQ) and TriviaQA using few-shot question answering. NQ was restricted to questions annotated as unambiguous in AmbigQA; potentially ambiguous TriviaQA questions were not removed. EM is the model's exact-match accuracy, shared across confidence methods within each dataset. Likelihood (L), diversity (D), repetition (R), and self-verification (V) are compared using ROC-AUC, ECE, and C@80; higher ROC-AUC and C@80 and lower ECE indicate better confidence estimates.

    Repetition achieved the best ECE on both datasets and the best ROC-AUC on TriviaQA. On NQ, likelihood had the highest ROC-AUC, while repetition had the highest C@80. Self-verification performed worse than the other methods across these calibration measures, especially on NQ.

    Dataset Method EM ROC-AUC ECE C@80
    NQ L 51.7 0.843 0.141 44.2
    NQ D 51.7 0.814 0.152 44.8
    NQ R 51.7 0.830 0.103 45.4
    NQ V 51.7 0.712 0.242 9.0
    TriviaQA L 72.5 0.826 0.205 87.6
    TriviaQA D 72.5 0.829 0.080 88.1
    TriviaQA R 72.5 0.844 0.052 88.1
    TriviaQA V 72.5 0.740 0.199 81.8
  5. Knowl 5 — Instruction tuning particularly benefits sampling-based calibration on ambiguity

    empirical result

    The authors evaluated instruction-tuned Flan-PaLM on Natural Questions and AmbigQA, using likelihood (L), sampling diversity (D), and sampling repetition (R) as confidence scores. EM is exact-match accuracy; ROC-AUC measures ranking of correct and incorrect answers; lower ECE indicates better calibration; C@80 is the maximum coverage at which the highest-confidence predictions retain at least 80% accuracy.

    On AmbigQA, repetition substantially outperformed likelihood in ranking and selective answering: ROC-AUC rose from 0.698 to 0.794 and C@80 from 27.3 to 52.9, while ECE fell from 0.356 to 0.137. On Natural Questions, likelihood had the best ROC-AUC and ECE, but diversity had the highest C@80. The results therefore show a particularly strong sampling-based calibration advantage for instruction-tuned models on ambiguous questions, rather than a uniform advantage on every dataset and metric.

    Dataset Method EM ROC-AUC ECE C@80
    NQ L 65.8 0.833 0.098 70.7
    NQ D 65.8 0.807 0.142 74.7
    NQ R 65.8 0.797 0.137 71.2
    AmbigQA L 57.6 0.698 0.356 27.3
    AmbigQA D 57.6 0.718 0.148 28.1
    AmbigQA R 57.6 0.794 0.137 52.9
  6. Knowl 6 — Calibration is more stable than closed-book accuracy across PaLM scales

    empirical result

    The authors compared PaLM models at 540B, 62B, and 8B parameters on Natural Questions and AmbigQA, using likelihood (L), sampling diversity (D), and sampling repetition (R). EM is exact-match accuracy; ROC-AUC and ECE measure confidence ranking and calibration, respectively; C@80 is the maximum coverage at at least 80% accuracy. The results show that closed-book accuracy and useful selective coverage decline sharply with model size, while ROC-AUC and ECE are comparatively more stable. For example, NQ EM falls from 51.7 at 540B to 16.0 at 8B, and AmbigQA EM falls from 44.8 to 14.8; C@80 is at or below 0.8 for all 8B results. Thus, relatively stable calibration scores do not imply that smaller models can answer a substantial fraction of questions accurately.

    Dataset Size Method EM ROC-AUC ECE C@80
    NQ 540B L 51.7 0.843 0.141 44.2
    NQ 540B D 51.7 0.814 0.152 44.8
    NQ 540B R 51.7 0.830 0.103 45.4
    NQ 62B L 38.8 0.823 0.111 18.8
    NQ 62B D 38.8 0.826 0.203 24.2
    NQ 62B R 38.8 0.836 0.151 24.5
    NQ 8B L 16.0 0.829 0.038 0.8
    NQ 8B D 16.0 0.813 0.228 0.8
    NQ 8B R 16.0 0.800 0.162 0.8
    AmbigQA 540B L 44.8 0.731 0.316 17.4
    AmbigQA 540B D 44.8 0.767 0.114 16.1
    AmbigQA 540B R 44.8 0.821 0.120 26.2
    AmbigQA 62B L 34.6 0.690 0.235 3.4
    AmbigQA 62B D 34.6 0.757 0.045 0.8
    AmbigQA 62B R 34.6 0.805 0.150 3.4
    AmbigQA 8B L 14.8 0.699 0.121 0.6
    AmbigQA 8B D 14.8 0.747 0.080 0.0
    AmbigQA 8B R 14.8 0.814 0.148 0.4
  7. Knowl 7 — Likelihood can misrepresent confidence in generated answers

    model/method

    The paper argues that a generated sequence's likelihood is not necessarily a reliable estimate of the probability of the extracted answer. First, decoding procedures used at inference time—including temperature, nucleus or top-kk sampling, length penalties, and truncation—change the output distribution in ways that may not be captured by the likelihood score. Second, an autoregressive model's sequence likelihood can assign probability mass to infinite sequences, so the sampled sequence process need not define a valid probability distribution over finite outputs. Third, question answering often applies an extraction or normalization function to generated text; estimating the probability of an answer after that transformation would require marginalizing over outputs that map to it. These issues motivate estimating confidence from repeated samples after answer extraction, rather than from the likelihood of one output alone.

  8. Knowl 8 — Sampling-based calibration generally improves with more samples, but not monotonically

    empirical result

    A sensitivity experiment compared sampling diversity and repetition using 3, 5, 8, and 10 outputs on the unambiguous Natural Questions slice and AmbigQA. The outcomes were evaluated with ROC-AUC, ECE, and C@80; EM was not reported because it depends on the greedy answer, not on the number of samples. In general, increasing the sample count improved calibration, especially on AmbigQA, but the metrics did not improve monotonically at every step. For example, AmbigQA repetition's ROC-AUC increased from 0.769 with 3 samples to 0.821 with 10, while C@80 moved from 16.7 to 26.2 and peaked at 26.9 with 8 samples. On unambiguous NQ, diversity with fewer samples could work relatively well for its compute cost. Sampling has a linear compute cost in the number of generated outputs.

  9. Knowl 9 — The tested methods barely predict whether a question is ambiguous

    empirical result

    On AmbigQA, where 53% of questions were labeled ambiguous, the authors tested whether sampled behavior could classify ambiguity. Approaches included predicting ambiguity from whether the model produced a disambiguation, voting over sampled disambiguations, using disagreement or the number of unique sampled answers, and directly prompting for an “Ambiguous” or “Unambiguous” label. None was particularly effective: the best reported accuracy was 58%, only modestly above the 53% rate obtained by always choosing the majority label, and none of the methods improved precision over that baseline. The authors suggest that ambiguity may depend on both the query and the interpreter, and that ordinary questions can be underspecified once context is removed; this is offered as a possible explanation, not a demonstrated cause. Diversity can also arise from epistemic uncertainty, making it difficult to attribute answer variation specifically to ambiguity.

  10. Knowl 10 — Study scope limits generalization of the calibration findings

    limitation

    The experiments primarily use PaLM-family models and evaluate closed-book question answering. They do not establish whether the findings generalize to other language-model families, retrieval-augmented question answering, or alternative training paradigms such as supervised fine-tuning. The authors also describe their exploration of confidence scores and prompt tuning as limited. Finally, sampling-based confidence estimates require additional generations, so their inference compute increases linearly with the number of samples.

Coverage note — The detailed per-sample-count values and prompt-format and “answer or unknown” ablations are omitted as secondary sensitivity analyses; the sample-count trend and its compute cost are retained.

References

  1. 1.Ayush Agrawal, Lester Mackey, and Adam Tauman Kalai. 2023. Do language models know when they’re hallucinating references? arXiv preprint arXiv:2305.18248.
  2. 2.Aryaman Arora, Clara Meister, and Ryan Cotterell. 2022. Estimating the entropy of linguistic distributions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 175–195, Dublin, Ireland. Association for Computational Linguistics.
  3. 3.Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 291–305, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  4. 4.Chi-Keung Chow. 1957. An optimum character recognition system using decision functions. IRE Transactions on Electronic Computers, EC-6(4):247–254.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  7. 7.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  8. 8.Li Dong, Chris Quirk, and Mirella Lapata. 2018. Confidence modeling for neural semantic parsing. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 743–753, Melbourne, Australia. Association for Computational Linguistics.
  9. 9.Ran El-Yaniv et al. 2010. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(5).
  10. 10.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  11. 11.Cornelia Gruber, Patrick Oliver Schenk, Malte Schierholz, Frauke Kreuter, and Göran Kauermann. 2023. Sources of uncertainty in machine learning–a statisticians’ view. arXiv e-prints, pages arXiv–2305.
  12. 12.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR.
  13. 13.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  14. 14.Abhyuday Jagannatha and Hong Yu. 2020. Calibrating structured output predictors for natural language processing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2078–2092, Online. Association for Computational Linguistics.
  15. 15.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  16. 16.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  17. 17.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  18. 18.Amita Kamath, Robin Jia, and Percy Liang. 2020. Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5684–5696, Online. Association for Computational Linguistics.
  19. 19.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2022. Clam: Selective clarification for ambiguous questions with large language models. arXiv preprint arXiv:2212.07769.
  20. 20.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664.
  21. 21.Volodymyr Kuleshov and Percy S Liang. 2015. Calibrated structured prediction. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  22. 22.Aviral Kumar and Sunita Sarawagi. 2019. Calibration of encoder decoder models for neural machine translation. arXiv preprint arXiv:1903.00802.
  23. 23.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  24. 24.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  25. 25.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334.
  26. 26.Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models. arXiv e-prints, pages arXiv–2305.
  27. 27.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  28. 28.Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. arXiv preprint arXiv:2303.08896.
  29. 29.Clara Meister, Afra Amini, Tim Vieira, and Ryan Cotterell. 2021. Conditional Poisson stochastic beams. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 664–681, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  30. 30.Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023. Locally typical sampling. Transactions of the Association for Computational Linguistics, 11:102–121.
  31. 31.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797, Online. Association for Computational Linguistics.
  32. 32.Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. 2021. Epistemic neural networks. arXiv preprint arXiv:2107.08924.
  33. 33.Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2023. Out-of-distribution detection and selective generation for conditional language models. In The Eleventh International Conference on Learning Representations.
  34. 34.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid questions meet long-form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8273–8288, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  35. 35.Neeraj Varshney and Chitta Baral. 2022. Model cascading: Towards jointly improving efficiency and accuracy of NLP systems. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11007–11021, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  36. 36.Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022. Investigating selective prediction approaches across several tasks in IID, OOD, and adversarial settings. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1995–2002, Dublin, Ireland. Association for Computational Linguistics.
  37. 37.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  39. 39.Allen R. Wilcox. 1973. Indices of qualitative variation and political measurement. The Western Political Quarterly, 26(2):325–343.
  40. 40.John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the National Conference on Artificial Intelligence, pages 1050–1055.
  41. 41.Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2021a. Trading off diversity and quality in natural language generation. In Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval), pages 25–33, Online. Association for Computational Linguistics.
  42. 42.Michael Zhang and Eunsol Choi. 2021. SituatedQA: Incorporating extra-linguistic contexts into QA. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7371–7387, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  43. 43.Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021b. Knowing more about questions can help: Improving calibration in question answering. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1958–1970, Online. Association for Computational Linguistics.
  44. 44.Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023. Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439.

Citation

MLA
Cole, J., et al. “Selectively Answering Ambiguous Questions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 530–43, https://doi.org/10.18653/v1/2023.emnlp-main.35.
APA
Cole, J., Zhang, M., Gillick, D., Eisenschlos, J., Dhingra, B., & Eisenstein, J. (2023). Selectively Answering Ambiguous Questions. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 530–543. https://doi.org/10.18653/v1/2023.emnlp-main.35
Chicago
Cole, J., M. Zhang, D. Gillick, J. Eisenschlos, B. Dhingra, and J. Eisenstein. 2023. “Selectively Answering Ambiguous Questions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 530–43. https://doi.org/10.18653/v1/2023.emnlp-main.35.
Harvard
Cole, J. et al. (2023) “Selectively Answering Ambiguous Questions”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 530–543. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.35.
Vancouver
1. Cole J, Zhang M, Gillick D, Eisenschlos J, Dhingra B, Eisenstein J (2023) Selectively Answering Ambiguous Questions. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 530–543

BibTeX

@inproceedings{cole-etal-2023-selectively,
    title = "Selectively Answering Ambiguous Questions",
    author = "Cole, Jeremy  and
      Zhang, Michael  and
      Gillick, Daniel  and
      Eisenschlos, Julian  and
      Dhingra, Bhuwan  and
      Eisenstein, Jacob",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.35/",
    doi = "10.18653/v1/2023.emnlp-main.35",
    pages = "530--543"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/