Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ambiguous questions

An ambiguous question is a query that admits multiple valid interpretations due to missing context, imprecise phrasing, ellipsis, or underspecification, leading to different possible correct answers depending on the underlying assumptions. In information retrieval and natural language processing, ambiguity arises when a prompt lacks sufficient constraints to isolate a unique user intent, referent, or entity. Consequently, addressing an ambiguous question effectively requires recognizing the uncertainty in the query, either by seeking clarification from the user to identify the intended meaning or by generating a comprehensive response that synthesizes answers across the various plausible interpretations.

4 items

Aligning Language Models to Explicitly Handle Ambiguity

Aligning Language Models to Explicitly Handle Ambiguity

Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim

OrganizationsHanyang UniversityIntelliSysKorea Advanced Institute of Science and TechnologyNaver AI LabNAVER CloudSeoul National University

Why you should read this

Proposes an alignment framework that enables language models to detect query ambiguity based on their own internal knowledge and proactively ask clarifying questions without degrading performance on unambiguous inputs.

In interactions between users and language model agents, user utterances frequently exhibit ellipsis (omission of words or phrases) or imprecision (lack of exactness) to prioritize efficiency. This can lead to varying interpretations of the same input based on different assumptions or background knowledge. It is thus crucial for agents to adeptly handle the inherent ambiguity in queries to ensure reliability. However, even state-of-the-art large language models (LLMs) still face challenges in such scenarios, primarily due to the following hurdles: (1) LLMs are not explicitly trained to deal with ambiguous utterances; (2) the degree of ambiguity perceived by the LLMs may vary depending on the possessed knowledge. To address these issues, we propose Alignment with Perceived Ambiguity (APA), a novel pipeline that aligns LLMs to manage ambiguous queries by leveraging their own assessment of ambiguity (i.e., perceived ambiguity). Experimental results on question-answering datasets demonstrate that APA empowers LLMs to explicitly detect and manage ambiguous queries while retaining the ability to answer clear questions. Furthermore, our finding proves that APA excels beyond training with gold-standard labels, especially in out-of-distribution scenarios. The data and code are available at https://github.com/heyjoonkim/APA.

Added

2026-10-03

Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, Jaewoo Kang

OrganizationsKorea Advanced Institute of Science and TechnologyKorea UniversityNaver AI LabNAVER CloudUniversity of Richmond

Why you should read this

Proposes Tree of Clarifications, a framework that recursively builds a tree of disambiguated questions guided by retrieved external knowledge and self-verification pruning to generate comprehensive long-form answers to ambiguous open-domain questions.

Questions in open-domain question answering are often ambiguous, allowing multiple interpretations. One approach to handling them is to identify all possible interpretations of the ambiguous question (AQ) and to generate a long-form answer addressing them all, as suggested by Stelmakh et al. (2022). While it provides a comprehensive response without bothering the user for clarification, considering multiple dimensions of ambiguity and gathering corresponding knowledge remains a challenge. To cope with the challenge, we propose a novel framework, Tree of Clarifications (ToC): It recursively constructs a tree of disambiguations for the AQ—via few-shot prompting leveraging external knowledge—and uses it to generate a long-form answer. ToC outperforms existing baselines on ASQA in a few-shot setup across all metrics, while surpassing fully-supervised baselines trained on the whole training set in terms of Disambig-F1 and Disambig-ROUGE. Code is available at github.com/gankim/tree-of-clarifications.

Added

2026-10-03

Selectively Answering Ambiguous Questions

Selectively Answering Ambiguous Questions

Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein

OrganizationsDuke UniversityGoogleUniversity of Texas at Austin

Why you should read this

Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.

Trustworthy language models should abstain from answering questions when they do not know the answer. However, the answer to a question can be unknown for a variety of reasons. Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context. We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous. In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work. We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning. Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.

Added

2026-10-03

ASQA: Factoid Questions Meet Long-Form Answers

ASQA: Factoid Questions Meet Long-Form Answers

Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang

OrganizationsDuke UniversityGoogleYakov & Partners

Why you should read this

Introduces the ASQA benchmark and an automated evaluation metric to resolve ambiguities in factoid questions through synthesized long-form answers with well-defined standards of factual correctness.

An abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA). This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations. The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality. In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA. Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation. Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity. In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question. We use this notion of correctness to define an automated metric of performance for ASQA. Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines.

Added

2026-09-29