Selectively Answering Ambiguous Questions
Jeremy R. ColeMichael J. Q. ZhangDaniel GillickJulian EisenschlosBhuwan DhingraJacob Eisenstein
Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.
Deploying large language models in customer-facing and decision-critical question answering systems carries significant risk when systems generate incorrect or misleading answers instead of abstaining. This challenge is intensified by real-world user queries that are frequently underspecified or context-dependent. Prior research largely addressed epistemic uncertainty—cases where the question is clear but the model may lack factual knowledge—while overlooking denotational uncertainty, where the user's intent or meaning is inherently ambiguous.
The article investigates how to reliably calibrate language models so they can selectively answer questions they understand with high accuracy and abstain when uncertain, specifically under conditions of heavy query ambiguity.
To address this, the authors evaluated few-shot prompting approaches using the Pathways Language Model (PaLM) across multiple benchmarks, including Natural Questions, TriviaQA, AmbigQA, and SituatedQA. Rather than relying on standard token likelihoods or asking models to verify their own outputs, the researchers introduced a two-step framework: the model first attempts to state an explicit interpretation of the question, after which candidate answers are generated. Model confidence was measured by drawing repeated samples (up to 10) and calculating sample repetition—the frequency with which sampled outputs match the primary answer—and sample diversity.
The investigation yielded several critical findings. First, sampling repetition proved to be the most reliable indicator of answer correctness across all scenarios, significantly outperforming standard model likelihood and self-verification prompting. Second, while standard likelihood scores performed adequately on clear questions, their calibration degraded severely on ambiguous queries, whereas sampling repetition maintained robust calibration. Third, instruction tuning improved raw accuracy but severely distorted the model's likelihood calibration; however, applying sampling-based confidence scoring restored calibration and substantially increased the volume of questions answered at high accuracy thresholds. Finally, directly predicting whether a question is ambiguous achieved poor accuracy (around 58%), indicating that ambiguity is better managed by generating explicit interpretations rather than binary classification.
These findings indicate that relying on raw model probability or self-prompted verification creates an unsafe, false sense of confidence in production environments, particularly for instruction-tuned models handling complex, ambiguous user prompts. Adopting a sampling-based confidence mechanism provides a practical method to safeguard performance and maintain strict accuracy standards (such as an 80% accuracy threshold) before returning answers to users.
Organizations implementing generative question-answering systems should adopt a disambiguate-then-answer strategy paired with sample repetition confidence scoring to govern abstention policies. However, decision-makers must weigh the trade-off of inference cost, as sampling multiple outputs increases compute requirements linearly. For budget-sensitive applications, smaller sample counts (such as three to five samples) or sample diversity metrics offer a viable, lower-cost compromise.
Readers should note that the evaluation was conducted solely on the PaLM model family within a closed-book setting without real-time external retrieval. Further testing across alternative model architectures and retrieval-augmented pipelines is recommended before full operational rollout.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Its self-consistency method establishes the repeated-sampling foundation behind the source’s use of answer repetition as a confidence signal.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its study of models’ ability to recognize when they know an answer provides essential context for the source’s comparison of confidence signals and abstention.
- Paper: Prompting is not a substitute for probability measurements in large language models, Jennifer Hu et al. (2023). Its distinction between prompted judgments and direct probability measurements clarifies the source’s finding that likelihood-based confidence can mislead.
- Paper: Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling, Bairu Hou et al. (2024). It extends ambiguity-aware answering by separating input ambiguity from knowledge gaps through an ensemble of explicit clarifications.
- Paper: LUQ: Long-text Uncertainty Quantification for LLMs, Caiqi Zhang et al. (2024). It carries sampling-based uncertainty estimation into long-form generation and uses the resulting confidence to support selective abstention.
- Paper: RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models, Aashiq Muhamed et al. (2025). It continues selective answering into retrieval-grounded systems, testing when models should refuse under ambiguity and other forms of defective context.
- Paper: SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales, Tianyang Xu et al. (2024). It develops a single-pass alternative to repeated sampling, addressing the inference-cost tradeoff that the source identifies for confidence calibration.
