Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
Bairu HouYujian LiuKaizhi QianJacob AndreasShiyu ChangYang Zhang
Proposes input clarification ensembling, a practical framework that separates large language model uncertainty into input ambiguity and model knowledge deficits without modifying model parameters or training procedures.
Deploying large language models in high-stakes environments requires understanding when their outputs are trustworthy. While measuring total uncertainty indicates how confident a model is, it does not reveal the root cause of the uncertainty. In predictive systems, uncertainty stems either from epistemic factors, which reflect the model's lack of knowledge, or aleatoric factors, which reflect inherent ambiguity or underspecification in the input prompt. Distinguishing between these two sources is critical: high epistemic uncertainty signals that a model requires better training data or external knowledge, whereas high aleatoric uncertainty indicates that the human user must provide a more specific prompt.
The article introduces and evaluates "input clarification ensembling," a practical framework designed to quantify and decompose uncertainty in large language models without altering their internal parameters. By shifting the decomposition process from complex model-level modifications to the input level, the article demonstrates how black-box language models can accurately isolate data ambiguity from knowledge gaps.
The approach operates by generating multiple plausible clarifications for an ambiguous prompt using a designated clarification model, such as a prompted advanced model or an efficiently fine-tuned smaller model. The target language model then evaluates each clarified input. By measuring the level of disagreement across predictions generated under different clarifications, the framework quantifies aleatoric uncertainty. The remaining average uncertainty across clarified inputs is then attributed to epistemic uncertainty. The authors validated this method across factual question-answering benchmarks and reasoning datasets, including Natural Questions, GSM8K, AmbigQA, and a custom dataset of ambiguous task instructions called AmbigInst.
The findings show that input clarification ensembling effectively identifies both errors and ambiguities. For general mistake detection, the framework matched or outperformed conventional ensembling and confidence-elicitation baselines, achieving area under the curve scores of 72.3 on Natural Questions and 89.7 on GSM8K. For ambiguity detection, the framework significantly outperformed existing baselines. On the AmbigQA benchmark, it achieved an area under the curve of 71.7, compared to approximately 53.6 to 55.4 for traditional methods. On the instruction ambiguity benchmark, it attained a score of 81.3, whereas baseline scores remained between 57.9 and 66.0. Additionally, the authors demonstrated that presenting users with generated clarifications significantly improved the recall of correct answers compared to directly querying models with ambiguous prompts.
These results indicate that organizations can diagnose model failures more accurately and implement dynamic interaction workflows. When a system detects high aleatoric uncertainty, it can actively prompt users with multiple interpretation options rather than returning an uncalibrated guess. Furthermore, the experiments demonstrate that a smaller open-source model can be fine-tuned in under ten minutes to generate effective clarifications, offering a computationally efficient path to deploying this capability without excessive infrastructure costs.
Organizations developing customer-facing or decision-critical AI systems should consider incorporating clarification-based ensembling into their inference pipelines. Future operational efforts should focus on optimizing clarification generation to capture subtle semantic ambiguities. However, decision-makers should note that the framework's effectiveness relies on the underlying models being reasonably well calibrated on the target domain. Performance was noticeably lower on implicit ambiguities that require deep contextual knowledge compared to explicit structural ambiguities, meaning human oversight remains essential in highly nuanced domains.
- Paper: Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods, Eyke Hüllermeier et al. (2019). Its clear account of aleatoric and epistemic uncertainty, and why standard predictive distributions conflate them, supplies the conceptual framework this paper adapts to LLMs.
- Paper: What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?, Alex Kendall et al. (2017). This work establishes the aleatoric–epistemic distinction in a Bayesian deep-learning setting that the source’s uncertainty decomposition builds on.
- Paper: A survey of uncertainty in deep neural networks, Jakob Gawlikowski et al. (2021). Its survey of uncertainty sources and estimation methods in deep networks provides background for the source’s Bayesian-style decomposition and ensembling approach.
- Paper: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, Yarin Gal et al. (2016). Understanding dropout as an approximation to Bayesian model uncertainty clarifies the epistemic-uncertainty framework invoked by the source.
- Paper: Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles, Balaji Lakshminarayanan et al. (2017). Its account of deep ensembles as a practical means of estimating predictive uncertainty prepares readers for the source’s ensemble-based method.
- Paper: Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness, Jiuhai Chen et al. (2024). This later framework turns LLM uncertainty estimation into training-free black-box confidence scoring, extending the source’s effort to quantify uncertainty in language-model answers.
- Paper: Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty, Kaitlyn Zhou et al. (2024). Its analysis connects models’ uncertainty expression to human reliance, carrying the source’s reliability concern from uncertainty measurement into user-facing consequences.
- Paper: RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models, Aashiq Muhamed et al. (2025). RefusalBench operationalizes uncertainty handling as selective abstention under ambiguity and missing context, extending the source’s focus on input uncertainty.
