Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models
Gangwoo KimSungdong KimByeongguk JeonJoonsuk ParkJaewoo Kang
Proposes Tree of Clarifications, a framework that recursively builds a tree of disambiguated questions guided by retrieved external knowledge and self-verification pruning to generate comprehensive long-form answers to ambiguous open-domain questions.
In real-world applications, user queries posed to search and question-answering systems are often ambiguous and open to multiple interpretations. Standard systems either fail to cover these nuances, demand time-consuming back-and-forth clarifications from users, or rely on language models prone to factual inaccuracies and hallucinations when answering broad questions.
The article demonstrates a novel framework called Tree of Clarifications (TOC), designed to automatically identify multiple interpretations of ambiguous questions and synthesize comprehensive, factually grounded long-form answers without requiring user intervention.
The approach combines large language models with external retrieval mechanisms in a recursive tree structure. Starting from an ambiguous question, the system retrieves external evidence using web search and dense passage retrieval (Bing and ColBERT), recursively branches the question into specific interpretations, prunes irrelevant or fact-distorting paths using an automated self-verification step, and aggregates the verified answers into a detailed response. The method was evaluated on the benchmark dataset ASQA (containing over 6,000 ambiguous questions) using a five-shot prompting setup rather than costly full-model retraining.
The findings show that TOC establishes a new state of the art in long-form ambiguous question answering. Operating with only five demonstration examples, TOC achieved a factual correctness score (Disambig-F1) of 33.7, outperforming fully supervised baseline models trained on the entire dataset (26.4) by 7.3 points and standard few-shot baselines (25.0) by 8.7 points. The self-verification pruning component raised the accuracy of intermediate answers from 40.9 to 59.3, effectively filtering out spurious disambiguations. In addition, combining dual retrieval sources with reranking improved the retrieval answer coverage across interpretations to 80.1%.
These results indicate that organizations can deliver highly accurate, comprehensive responses to complex, multi-faceted queries without expensive full-scale model training. Grounding multi-path reasoning in external factual retrieval significantly mitigates hallucination risks while preserving high response quality across diverse user intents.
For practical implementation, technical teams should consider adopting retrieval-augmented tree reasoning architectures for knowledge-intensive search and automated assistance workflows. To balance performance and operating costs, deployments should implement structured stopping rules (such as node caps and search-depth limits) to control query latency and API expenses.
Confidence in the reported benchmark gains is high, though readers should note certain limitations. The evaluation relies primarily on a single benchmark dataset and a single primary language model backbone (GPT-3). Furthermore, recursive tree exploration requires multiple model queries per question—up to 20 calls in worst-case paths—which increases computational overhead compared to direct single-pass generation. Further testing on diverse domain datasets and smaller, open-source models is recommended before broad enterprise rollout.
- Paper: ASQA: Factoid Questions Meet Long-Form Answers, Ivan Stelmakh et al. (2022). Read this first to understand the ASQA benchmark and its long-form evaluation, which the source uses to measure answers across ambiguous interpretations.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Its tree-based generation, evaluation, and pruning provide a useful conceptual precursor to the source’s recursive branching over interpretations.
- Paper: Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity, Soyeong Jeong et al. (2024). Adaptive-RAG carries retrieval-augmented question answering forward by routing questions to retrieval strategies according to complexity, extending the source’s concern with multi-step query handling.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). Forest-of-Thought extends tree-based reasoning beyond a single search tree by coordinating multiple paths for greater test-time reasoning and correction.
