Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ambiguity detection

Ambiguity detection is the computational process of identifying whether a given piece of text or user input admits multiple plausible interpretations due to underspecification, vagueness, lexical polysemy, or missing contextual information. In natural language processing and artificial intelligence, this capability enables systems to distinguish between clear, well-defined inputs and those that carry inherent data-level uncertainty. By recognizing when an utterance cannot be mapped to a single definitive meaning, language models and conversational agents can avoid making unwarranted assumptions and instead trigger appropriate downstream actions, such as seeking clarification from the user, generating multiple possible interpretations, or properly calibrating predictive uncertainty to improve overall system reliability and trustworthiness.

2 items

Aligning Language Models to Explicitly Handle Ambiguity

Aligning Language Models to Explicitly Handle Ambiguity

Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim

OrganizationsHanyang UniversityIntelliSysKorea Advanced Institute of Science and TechnologyNaver AI LabNAVER CloudSeoul National University

Why you should read this

Proposes an alignment framework that enables language models to detect query ambiguity based on their own internal knowledge and proactively ask clarifying questions without degrading performance on unambiguous inputs.

In interactions between users and language model agents, user utterances frequently exhibit ellipsis (omission of words or phrases) or imprecision (lack of exactness) to prioritize efficiency. This can lead to varying interpretations of the same input based on different assumptions or background knowledge. It is thus crucial for agents to adeptly handle the inherent ambiguity in queries to ensure reliability. However, even state-of-the-art large language models (LLMs) still face challenges in such scenarios, primarily due to the following hurdles: (1) LLMs are not explicitly trained to deal with ambiguous utterances; (2) the degree of ambiguity perceived by the LLMs may vary depending on the possessed knowledge. To address these issues, we propose Alignment with Perceived Ambiguity (APA), a novel pipeline that aligns LLMs to manage ambiguous queries by leveraging their own assessment of ambiguity (i.e., perceived ambiguity). Experimental results on question-answering datasets demonstrate that APA empowers LLMs to explicitly detect and manage ambiguous queries while retaining the ability to answer clear questions. Furthermore, our finding proves that APA excels beyond training with gold-standard labels, especially in out-of-distribution scenarios. The data and code are available at https://github.com/heyjoonkim/APA.

Added

2026-10-03

Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, Yang Zhang

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes input clarification ensembling, a practical framework that separates large language model uncertainty into input ambiguity and model knowledge deficits without modifying model parameters or training procedures.

Uncertainty decomposition refers to the task of decomposing the total uncertainty of a predictive model into aleatoric (data) uncertainty, resulting from inherent randomness in the data-generating process, and epistemic (model) uncertainty, resulting from missing information in the model’s training data. In large language models (LLMs) specifically, identifying sources of uncertainty is an important step toward improving reliability, trustworthiness, and interpretability, but remains an important open research question. In this paper, we introduce an uncertainty decomposition framework for LLMs, called input clarification ensembling, which can be applied to any pre-trained LLM. Our approach generates a set of clarifications for the input, feeds them into an LLM, and ensembles the corresponding predictions. We show that, when aleatoric uncertainty arises from ambiguity or under-specification in LLM inputs, this approach makes it possible to factor an (un-clarified) LLM’s predictions into separate aleatoric and epistemic terms, using a decomposition similar to the one employed by Bayesian neural networks. Empirical evaluations demonstrate that input clarification ensembling provides accurate and reliable uncertainty quantification on several language processing tasks. Code and data are available at https://github.com/UCSB-NLP-Chang/llm_uncertainty.

Added

2026-10-01