Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Bairu HouYujian LiuKaizhi QianJacob AndreasShiyu ChangYang Zhang

article2024ICML138 citations

Proposes input clarification ensembling, a practical framework that separates large language model uncertainty into input ambiguity and model knowledge deficits without modifying model parameters or training procedures.

Listen

Deploying large language models in high-stakes environments requires understanding when their outputs are trustworthy. While measuring total uncertainty indicates how confident a model is, it does not reveal the root cause of the uncertainty. In predictive systems, uncertainty stems either from epistemic factors, which reflect the model's lack of knowledge, or aleatoric factors, which reflect inherent ambiguity or underspecification in the input prompt. Distinguishing between these two sources is critical: high epistemic uncertainty signals that a model requires better training data or external knowledge, whereas high aleatoric uncertainty indicates that the human user must provide a more specific prompt.

The article introduces and evaluates "input clarification ensembling," a practical framework designed to quantify and decompose uncertainty in large language models without altering their internal parameters. By shifting the decomposition process from complex model-level modifications to the input level, the article demonstrates how black-box language models can accurately isolate data ambiguity from knowledge gaps.

The approach operates by generating multiple plausible clarifications for an ambiguous prompt using a designated clarification model, such as a prompted advanced model or an efficiently fine-tuned smaller model. The target language model then evaluates each clarified input. By measuring the level of disagreement across predictions generated under different clarifications, the framework quantifies aleatoric uncertainty. The remaining average uncertainty across clarified inputs is then attributed to epistemic uncertainty. The authors validated this method across factual question-answering benchmarks and reasoning datasets, including Natural Questions, GSM8K, AmbigQA, and a custom dataset of ambiguous task instructions called AmbigInst.

The findings show that input clarification ensembling effectively identifies both errors and ambiguities. For general mistake detection, the framework matched or outperformed conventional ensembling and confidence-elicitation baselines, achieving area under the curve scores of 72.3 on Natural Questions and 89.7 on GSM8K. For ambiguity detection, the framework significantly outperformed existing baselines. On the AmbigQA benchmark, it achieved an area under the curve of 71.7, compared to approximately 53.6 to 55.4 for traditional methods. On the instruction ambiguity benchmark, it attained a score of 81.3, whereas baseline scores remained between 57.9 and 66.0. Additionally, the authors demonstrated that presenting users with generated clarifications significantly improved the recall of correct answers compared to directly querying models with ambiguous prompts.

These results indicate that organizations can diagnose model failures more accurately and implement dynamic interaction workflows. When a system detects high aleatoric uncertainty, it can actively prompt users with multiple interpretation options rather than returning an uncalibrated guess. Furthermore, the experiments demonstrate that a smaller open-source model can be fine-tuned in under ten minutes to generate effective clarifications, offering a computationally efficient path to deploying this capability without excessive infrastructure costs.

Organizations developing customer-facing or decision-critical AI systems should consider incorporating clarification-based ensembling into their inference pipelines. Future operational efforts should focus on optimizing clarification generation to capture subtle semantic ambiguities. However, decision-makers should note that the framework's effectiveness relies on the underlying models being reasonably well calibrated on the target domain. Performance was noticeably lower on implicit ambiguities that require deep contextual knowledge compared to explicit structural ambiguities, meaning human oversight remains essential in highly nuanced domains.

Cover for Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Abstract

Uncertainty decomposition refers to the task of decomposing the total uncertainty of a predictive model into aleatoric (data) uncertainty, resulting from inherent randomness in the data-generating process, and epistemic (model) uncertainty, resulting from missing information in the model’s training data. In large language models (LLMs) specifically, identifying sources of uncertainty is an important step toward improving reliability, trustworthiness, and interpretability, but remains an important open research question. In this paper, we introduce an uncertainty decomposition framework for LLMs, called input clarification ensembling, which can be applied to any pre-trained LLM. Our approach generates a set of clarifications for the input, feeds them into an LLM, and ensembles the corresponding predictions. We show that, when aleatoric uncertainty arises from ambiguity or under-specification in LLM inputs, this approach makes it possible to factor an (un-clarified) LLM’s predictions into separate aleatoric and epistemic terms, using a decomposition similar to the one employed by Bayesian neural networks. Empirical evaluations demonstrate that input clarification ensembling provides accurate and reliable uncertainty quantification on several language processing tasks. Code and data are available at https://github.com/UCSB-NLP-Chang/llm_uncertainty.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Notations and Problem Formulation
  • 3.2. Background: Bayesian Neural Networks and DEEP ENSEMBLES
  • 3.3. Do BNN and DEEP ENSEMBLES work for LLMs?
  • 3.4. Input Clarification Ensembling
  • 3.5. Input Clarification
  • 3.6. Improving Performance via Soliciting Clarifications
  • 4. Experiments
  • 4.1. Experiment Configurations
  • 4.2. Quantifying Total Uncertainty
  • 4.3. Uncertainty Decomposition
  • 4.4. Monotonicity Check
  • 4.5. Recall of Correct Answers
  • 5. Conclusion
  • Acknowledgment
  • Impact Statement
  • References
  • A. Additional Results and Implementation Details
  • A.1. Additional Results
  • A.2. Supervised Fine-tuning for Clarification Generation
  • A.3. Implementation details for baselines
  • A.4. Prompts for Our Clarification Model
  • A.5. Details of the Qualitative Results
  • A.6. Prompt the LLM for Answer Extraction
  • B. AmbigInst Dataset
  • B.1. Dataset Creation
  • B.2. Dataset Examples

Knowls

  1. Knowl 1 — Input clarification ensembling decomposes predictive uncertainty

    model/method

    Input clarification ensembling estimates uncertainty for a fixed large language model (LLM) by generating clarifications of an input and averaging predictions over those clarifications. Let XX be an input, YY the model’s output, CC a random clarification drawn from q(C∣X)q(C\mid X), and θ\theta the fixed parameters of the prediction LLM. Write X⊕CX\oplus C for appending the clarification to the input. The ensemble prediction is

    q(Y∣X)=Eq(C∣X)[q(Y∣X⊕C,θ)].q(Y\mid X)=\mathbb{E}_{q(C\mid X)}[q(Y\mid X\oplus C,\theta)].

    Its predictive entropy decomposes as

    H(q(Y∣X))=Iq(Y;C∣X)+Eq(C∣X)[H(q(Y∣X⊕C,θ))].H(q(Y\mid X))=I_q(Y;C\mid X)+\mathbb{E}_{q(C\mid X)}[H(q(Y\mid X\oplus C,\theta))].

    Here HH is entropy and IqI_q is conditional mutual information under the joint distribution formed by sampling CC from q(C∣X)q(C\mid X) and then YY from q(Y∣X⊕C,θ)q(Y\mid X\oplus C,\theta). The first term measures variation in predictions associated with different clarifications; the second is the average predictive entropy after clarification. The method changes the inputs rather than the prediction model’s parameters.

  2. Knowl 2 — Interpretation of the two uncertainty components

    assumption

    In input clarification ensembling, the mutual information Iq(Y;C∣X)I_q(Y;C\mid X) is used as a proxy for aleatoric uncertainty caused by ambiguity in the input: different valid interpretations, represented by different clarifications, can lead to different answers. The average conditional entropy Eq(C∣X)[H(q(Y∣X⊕C,θ))]\mathbb{E}_{q(C\mid X)}[H(q(Y\mid X\oplus C,\theta))] measures uncertainty remaining after clarification and may be treated as an estimate of epistemic uncertainty only under the assumption that input ambiguity is the sole source of aleatoric uncertainty. This latter interpretation is not the traditional definition of epistemic uncertainty; the paper’s primary focus is estimating ambiguity-related aleatoric uncertainty. More generally, total uncertainty is predictive entropy, epistemic uncertainty concerns missing model knowledge, and aleatoric uncertainty is irreducible uncertainty in the target distribution.

  3. Knowl 3 — Clarification generation and prediction procedure

    model/method

    The method first selects an input component to clarify, such as a question or task instruction, and uses a clarification LLM to produce alternative texts that reduce its ambiguity. Each clarification is appended to the relevant original input component, and the same prediction LLM is run on each clarified input; predictions are aggregated to estimate the output distribution and its uncertainty decomposition. The clarification LLM may differ from the prediction LLM. In the experiments, question clarifications for Natural Questions and GSM8K were generated by prompting the LLM to rephrase the question; GSM8K clarifications were sampled five times. For AmbigQA, the authors evaluated GPT-4 prompted with the 16 most similar questions as in-context examples, with similarity measured using Sentence-BERT embeddings, and a Llama-3-8B-Instruct model fine-tuned on AmbigQA’s training set. Instruction clarifications for AmbigInst were generated by prompting GPT-3.5-turbo-0613. The Llama model was fine-tuned for five epochs using batch size 16, learning rate 2×10−52\times10^{-5}, a cosine learning-rate schedule, and four 80-GB H100 GPUs; the checkpoint with lowest validation loss was from epoch 2.

  4. Knowl 4 — Total-uncertainty estimates detect incorrect answers

    data/table

    The authors tested whether total uncertainty from input clarification ensembling tracks answer correctness on 200 randomly sampled validation examples each from Natural Questions (NQ) and GSM8K. The prediction model was GPT-3.5-turbo-0613; ten predictions were sampled at temperature 0.5, using five in-context examples for NQ and two examples with chain-of-thought for GSM8K. They report AUROC and best-threshold F1 for predicting correctness, plus mean entropy for correct and incorrect answers. The proposed method achieved competitive mistake detection on both tasks; its mean entropy was higher for incorrect than correct answers in each dataset. Scores are reproduced as reported.

    Dataset Method AUROC F1 Score Entropy (correct) Entropy (incorrect)
    NQ Semantic entropy 63.8 77.9 0.29 0.56
    NQ Ask4Conf 70.4 83.9 – –
    NQ Ensembles∗^* 69.7 79.7 0.46 0.88
    NQ Ours 72.3 80.2 0.58 1.18
    GSM8K Semantic entropy 88.2 92.4 0.32 1.46
    GSM8K Ask4Conf 58.1 92.3 – –
    GSM8K Ensembles∗^* 88.3 94.6 0.57 1.94
    GSM8K Ours 89.7 94.7 0.42 1.82

    Here Ensembles∗^* is an adapted baseline that ensembles predictions from five different in-context-example sets rather than trained model variants. Ask4Conf does not provide the reported entropy statistics. On GSM8K, Ask4Conf had much lower AUROC than the other methods, while input clarification ensembling remained competitive.

  5. Knowl 5 — Aleatoric uncertainty identifies ambiguous inputs

    data/table

    The ambiguity-detection experiment evaluated whether estimated uncertainty predicts input ambiguity. It used 200 AmbigQA validation questions and the full AmbigInst dataset; question ambiguity was measured by clarifying the question, while instruction ambiguity was measured by clarifying the instruction. AmbigQA used five-shot prompting, and AmbigInst used zero-shot prompting. For the proposed method, the table reports aleatoric uncertainty; the baselines generally use total uncertainty unless marked otherwise. AU (clear) and AU (ambiguous) are mean aleatoric uncertainties for unambiguous and ambiguous inputs. Input clarification ensembling outperformed the listed non-oracle baselines in AUROC on both datasets; ground-truth clarifications (“Ours∗^*”) provide an oracle reference.

    Dataset Method AUROC F1 Score Avg. AU (clear) Avg. AU (ambiguous)
    AmbigQA Semantic entropy 54.9 46.8 0.24 0.47
    AmbigQA Ask4Conf-D 55.0 64.3 – –
    AmbigQA Ensembles∗^* (aleatoric) 53.6 53.0 0.13 0.13
    AmbigQA Ensembles∗^* (total) 55.4 55.0 0.50 0.41
    AmbigQA Ours (GPT) 71.7 70.1 0.28 0.67
    AmbigQA Ours (LLaMA) 67.1 71.8 0.55 0.91
    AmbigQA Ours∗^* 89.8 85.6 0.53 1.52
    AmbigInst Semantic entropy 66.0 53.7 0.07 0.50
    AmbigInst Ask4Conf-D 57.9 75.4 – –
    AmbigInst Ours (GPT) 81.3 77.9 0.10 0.75
    AmbigInst Ours∗^* 96.7 92.6 0.10 1.04

    On AmbigQA, GPT-based clarification achieved AUROC 71.7, compared with 55.4 for the total-uncertainty Ensembles∗^* baseline; the adapted LLaMA clarifier achieved AUROC 67.1. On AmbigInst, GPT-based clarification achieved AUROC 81.3, compared with 66.0 for semantic entropy and 57.9 for Ask4Conf-D. The weak AmbigQA aleatoric score of Ensembles∗^* (AUROC 53.6) also shows that simply varying in-context examples did not provide effective ambiguity estimates in this setting.

  6. Knowl 6 — Clarifications increase recall of intended answers

    empirical result

    The authors evaluated whether alternative clarified inputs can help a user obtain an intended answer that is one of several valid answers to an ambiguous input. For each AmbigQA question or AmbigInst instruction, they selected a target answer from the dataset’s possible labeled answers, then compared answering the original input directly with answering versions augmented by increasing numbers of generated clarifications. Recall of the target answer rose consistently as the number of clarifications increased from one through five, and clarification-based answering improved over the unclarified baseline on both datasets. The gain was more pronounced on AmbigInst than AmbigQA, consistent with the latter’s more subtle ambiguities.

  7. Knowl 7 — Clarifying inputs reduces measured aleatoric uncertainty

    empirical result

    A two-round monotonicity check tested whether the clarification model actually removes ambiguity. In round one, the authors measured aleatoric uncertainty by generating clarifications for the original question or instruction. In round two, they applied the same clarification prompt to the already clarified inputs and measured uncertainty again. On both AmbigQA and AmbigInst, average aleatoric uncertainty dropped substantially after clarification, supporting the claim that the clarification module reduces ambiguity in the inputs it processes.

  8. Knowl 8 — AmbigInst supplies examples of instruction ambiguity

    data/table

    AmbigInst is a synthetic dataset created to evaluate detection of ambiguous task instructions, for which the authors found no existing dataset. ChatGPT generated candidate ambiguous instructions and their causes; the authors removed tasks with open-ended output spaces and generated ground-truth clarifications. For each retained ambiguous instruction, ChatGPT generated input-output pairs with different answers under different clarifications, and pairs whose answers did not vary across clarifications were filtered out. This yielded 15 ambiguous tasks and 214 input questions. The dataset also includes 10 unambiguous tasks adapted from Instruction Induction, with clarifications manually added to remove potential underspecification and 15 input-output pairs per task. The complete dataset contains 25 tasks and 364 inputs. The ambiguous tasks make ambiguity explicit, while the unambiguous tasks provide comparison examples.

  9. Knowl 9 — Semantic answer clustering is used to estimate output distributions

    model/method

    Because free-form generations can express the same answer in different words, the experiments estimate answer probabilities by clustering semantically equivalent outputs rather than treating every distinct string as a distinct outcome. The authors prompt an LLM to extract and group answers, reporting that this performed better for their purposes than using a natural-language-inference model. Refusals or outputs indicating insufficient information are mapped to a special answer, “Unknown,” using extraction instructions and keyword matching. If every answer for a clarification is Unknown, that clarification is treated as invalid and excluded from the ensemble. Otherwise, the Unknown outcome is excluded from the set of answer categories, and its occurrences are distributed evenly across the remaining categories by increasing each category’s normalized frequency by 1/N1/N, where NN is the number of non-Unknown answer categories. This treatment is intended to make model ignorance contribute to the estimated uncertainty without treating Unknown as an ordinary answer.

  10. Knowl 10 — Reliability depends on calibration and clarification quality

    limitation

    The authors state that the reliability of their uncertainty estimates depends in part on LLM predictions being well calibrated; calibration observed on factoid question answering and mathematical reasoning may not hold on other downstream tasks. The decomposition also depends on clarifications adequately resolving the relevant ambiguity, and the interpretation of residual conditional entropy as epistemic uncertainty requires input ambiguity to be the only aleatoric source. Performance was lower on AmbigQA, whose ambiguities can be implicit and may require external knowledge to recognize, than on the more explicit synthetic AmbigInst tasks. The authors identify improving the clarification module as future work.

Coverage note — Related-work comparisons, prompt wording, and qualitative case descriptions were omitted because they do not add standalone contributed findings beyond the method, evaluations, and limitations captured here.

References

  1. 1.Bhatt, U., Antoran, J., Zhang, Y., Liao, Q. V., Sattigeri, P., Fogliato, R., Melançon, G., Krishnan, R., Stanley, J., Tickoo, O., et al. Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 401–413, 2021.
  2. 2.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight uncertainty in neural network. In International conference on machine learning, 2015.
  3. 3.Chen, J. and Mueller, J. Quantifying uncertainty in answers from any language model via intrinsic and extrinsic confidence assessment. arXiv preprint arXiv:2308.16175, 2023.
  4. 4.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  5. 5.Cole, J. R., Zhang, M. J., Gillick, D., Eisenschlos, J. M., Dhingra, B., and Eisenstein, J. Selectively answering ambiguous questions. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  6. 6.Desai, S. and Durrett, G. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020.
  7. 7.Duan, J., Cheng, H., Wang, S., Wang, C., Zavalny, A., Xu, R., Kailkhura, B., and Xu, K. Shifting attention to relevance: Towards the uncertainty estimation of large language models. arXiv, 2023.
  8. 8.Fort, S., Hu, H., and Lakshminarayanan, B. Deep ensembles: A loss landscape perspective. arXiv preprint arXiv:1912.02757, 2019.
  9. 9.Gal, Y. and Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pp. 1050–1059. PMLR, 2016.
  10. 10.Gal, Y. et al. Uncertainty in deep learning, 2016.
  11. 11.Graves, A. Practical variational inference for neural networks. Advances in neural information processing systems, 24, 2011.
  12. 12.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International conference on machine learning, pp. 1321–1330. PMLR, 2017.
  13. 13.Guo, M., Zhang, M., Reddy, S., and Alikhani, M. Abgcoqa: Clarifying ambiguity in conversational question answering. In 3rd Conference on Automated Knowledge Base Construction, 2021.
  14. 14.Hasenclever, L., Webb, S., Lienart, T., Vollmer, S., Lakshminarayanan, B., Blundell, C., and Teh, Y. W. Distributed bayesian learning with stochastic natural gradient expectation propagation and the posterior server. Journal of Machine Learning Research, 18(106):1–37, 2017.
  15. 15.He, B., Lakshminarayanan, B., and Teh, Y. W. Bayesian deep ensembles via the neural tangent kernel. Advances in neural information processing systems, 33:1010–1022, 2020.
  16. 16.Hendrycks, D. and Gimpel, K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. arXiv preprint arXiv:1610.02136, 2016.
  17. 17.Hernández-Lobato, J. M. and Adams, R. Probabilistic back-propagation for scalable learning of bayesian neural networks. In International conference on machine learning, pp. 1861–1869. PMLR, 2015.
  18. 18.Honovich, O., Shaham, U., Bowman, S. R., and Levy, O. Instruction induction: From few examples to natural language task descriptions. arXiv preprint arXiv:2205.10782, 2022.
  19. 19.Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J. Large language models can self-improve. arXiv preprint arXiv:2210.11610, 2022.
  20. 20.Huang, Y., Song, J., Wang, Z., Chen, H., and Ma, L. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv preprint arXiv:2307.10236, 2023.
  21. 21.Hüllermeier, E. and Waegeman, W. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine Learning, 110:457–506, 2021.
  22. 22.Jiang, M., Ruan, Y., Huang, S., Liao, S., Pitis, S., Grosse, R. B., and Ba, J. Calibrating language models via augmented prompt ensembles. 2023.
  23. 23.Jiang, Z., Araki, J., Ding, H., and Neubig, G. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 2021.
  24. 24.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  25. 25.Koller, A., Regneri, M., and Thater, S. Regular tree grammars as a formalism for scope underspecification. In Proceedings of ACL-08: HLT, pp. 218–226, 2008.
  26. 26.Kristiadi, A., Hein, M., and Hennig, P. Learnable uncertainty under laplace approximations. In Uncertainty in Artificial Intelligence, pp. 344–353. PMLR, 2021.
  27. 27.Kuhn, L., Gal, Y., and Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, 2022.
  28. 28.Kuhn, L., Gal, Y., and Farquhar, S. Clam: Selective clarification for ambiguous questions with generative language models. 2023.
  29. 29.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466, 2019.
  30. 30.Lahlou, S., Jain, M., Nekoei, H., Butoi, V. I., Bertin, P., Rector-Brooks, J., Korablyov, M., and Bengio, Y. Deup: Direct epistemic uncertainty prediction. Transactions on Machine Learning Research, 2022.
  31. 31.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30, 2017.
  32. 32.Li, Y., Hernandez-Lobato, J. M., and Turner, R. E. Stochastic expectation propagation. Advances in neural information processing systems, 28, 2015.
  33. 33.Lin, S., Hilton, J., and Evans, O. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research, 2022.
  34. 34.Lin, Z., Trivedi, S., and Sun, J. Generating with confidence: Uncertainty quantification for black-box large language models. arXiv preprint arXiv:2305.19187, 2023.
  35. 35.Liu, A., Wu, Z., Michael, J., Suhr, A., West, P., Koller, A., Swayamdipta, S., Smith, N. A., and Choi, Y. We’re afraid language models aren’t modeling ambiguity. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  36. 36.Louizos, C. and Welling, M. Structured and efficient variational deep learning with matrix gaussian posteriors. In International conference on machine learning, pp. 1708–1716. PMLR, 2016.
  37. 37.Malinin, A. and Gales, M. Predictive uncertainty estimation via prior networks. Advances in neural information processing systems, 31, 2018.
  38. 38.Malinin, A. and Gales, M. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations, 2020.
  39. 39.Malinin, A., Prokhorenkova, L., and Ustimenko, A. Uncertainty in gradient boosting via ensembles. In International Conference on Learning Representations, 2020.
  40. 40.Mielke, S. J., Szlam, A., Dinan, E., and Boureau, Y.-L. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 2022.
  41. 41.Min, S., Michael, J., Hajishirzi, H., and Zettlemoyer, L. Ambigqa: Answering ambiguous open-domain questions. arXiv preprint arXiv:2004.10645, 2020.
  42. 42.Mobiny, A., Yuan, P., Moulik, S. K., Garg, N., Wu, C. C., and Van Nguyen, H. Dropconnect is effective in modeling uncertainty of bayesian deep networks. Scientific reports, 11:5458, 2021.
  43. 43.Neal, R. M. Bayesian learning for neural networks, volume 118. Springer Science & Business Media, 2012.
  44. 44.Ott, M., Auli, M., Grangier, D., and Ranzato, M. Analyzing uncertainty in neural machine translation. In International Conference on Machine Learning, pp. 3956–3965. PMLR, 2018.
  45. 45.Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., and Snoek, J. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in neural information processing systems, 32, 2019.
  46. 46.Park, S. and Kim, T. Pac neural prediction set learning to quantify the uncertainty of generative language models. arXiv preprint arXiv:2307.09254, 2023.
  47. 47.Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992, 2019.
  48. 48.Ren, A., Dixit, A., Bodrova, A., Singh, S., Tu, S., Brown, N., Xu, P., Takayama, L., Xia, F., Varley, J., et al. Robots that ask for help: Uncertainty alignment for large language model planners. In 2nd Workshop on Language and Robot Learning: Language as Grounding, 2023.
  49. 49.Riquelme, C., Tucker, G., and Snoek, J. Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling. In International Conference on Learning Representations, 2018.
  50. 50.Shen, M., Bu, Y., Sattigeri, P., Ghosh, S., Das, S., and Wornell, G. Post-hoc uncertainty learning using a dirichlet meta-model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 9772–9781, 2023.
  51. 51.Si, C., Gan, Z., Yang, Z., Wang, S., Wang, J., Boyd-Graber, J. L., and Wang, L. Prompting gpt-3 to be reliable. In The Eleventh International Conference on Learning Representations, 2022.
  52. 52.Si, C., Gan, Z., Yang, Z., Wang, S., Wang, J., Boyd-Graber, J. L., and Wang, L. Prompting GPT-3 to be reliable. In The Eleventh International Conference on Learning Representations, 2023.
  53. 53.Tamkin, A., Handa, K., Shrestha, A., and Goodman, N. Task ambiguity in humans and language models. In The Eleventh International Conference on Learning Representations, 2022.
  54. 54.Teye, M., Azizpour, H., and Smith, K. Bayesian uncertainty estimation for batch normalized deep networks. In International Conference on Machine Learning, pp. 4907–4916. PMLR, 2018.
  55. 55.Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975, 2023.
  56. 56.Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2022.
  57. 57.Xiao, Y., Liang, P. P., Bhatt, U., Neiswanger, W., Salakhutdinov, R., and Morency, L.-P. Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 7273–7284, 2022.
  58. 58.Ye, X. and Durrett, G. Can explanations be useful for calibrating black box models?, 2022.
  59. 59.Zhou, K., Jurafsky, D., and Hashimoto, T. Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439, 2023.

Citation

MLA
Hou, B., et al. “Decomposing Uncertainty for Large Language Models Through Input Clarification Ensembling”. arXiv, 2023, http://arxiv.org/abs/2311.08718v2.
APA
Hou, B., Liu, Y., Qian, K., Andreas, J., Chang, S., & Zhang, Y. (2023). Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling. arXiv. http://arxiv.org/abs/2311.08718v2
Chicago
Hou, B., Y. Liu, K. Qian, J. Andreas, S. Chang, and Y. Zhang. 2023. “Decomposing Uncertainty for Large Language Models Through Input Clarification Ensembling”. arXiv. http://arxiv.org/abs/2311.08718v2.
Harvard
Hou, B. et al. (2023) “Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.08718v2.
Vancouver
1. Hou B, Liu Y, Qian K, Andreas J, Chang S, Zhang Y (2023) Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling. arXiv

BibTeX

@article{hou2023decomposing,
  title = {Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling},
  author = {Hou, Bairu and Liu, Yujian and Qian, Kaizhi and Andreas, Jacob and Chang, Shiyu and Zhang, Yang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.08718v2},
  eprint = {2311.08718}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/