Built independently by an author, for readers. Read the story and support ChapterPal

keyword

denotational uncertainty

Denotational uncertainty refers to the ambiguity or indeterminacy regarding the precise meaning, intent, or formal interpretation of an input expression, such as a natural language question or prompt. In computational linguistics and machine learning, this form of uncertainty occurs when an input is underspecified, lacks adequate background context, or naturally supports multiple distinct interpretations. It is distinct from epistemic uncertainty, which arises when a system lacks the factual knowledge required to answer a clearly stated query; under denotational uncertainty, several valid answers may exist solely because the initial query maps to multiple possible concepts or references. Resolving or calibrating for denotational uncertainty typically involves recognizing linguistic ambiguity, generating explicit disambiguations, or evaluating distributions over potential interpretations.

1 item

Selectively Answering Ambiguous Questions

Selectively Answering Ambiguous Questions

Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein

OrganizationsDuke UniversityGoogleUniversity of Texas at Austin

Why you should read this

Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.

Trustworthy language models should abstain from answering questions when they do not know the answer. However, the answer to a question can be unknown for a variety of reasons. Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context. We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous. In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work. We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning. Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.

Added

2026-10-03