Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Miao XiongZhiyuan HuXinyang LuYifei LiJie FuJunxian HeBryan Hooi

article2023ICLR1,159 citations

Presents a systematic framework for evaluating black-box uncertainty estimation in large language models, showing that verbalized prompting and response consistency can effectively reduce overconfidence to rival white-box calibration methods.

Listen

As large language models become central to commercial applications, accurately assessing their uncertainty is critical for risk assessment, safety, and reducing factual errors. Most commercial models operate as black boxes accessible only via text prompts, making traditional white-box methods—such as inspecting internal probabilities or performing expensive fine-tuning—infeasible. Consequently, decision-makers face major reliability risks when relying on a model's stated confidence.

The article systematically evaluates how effectively large language models express their own uncertainty using purely black-box techniques. It introduces a modular framework combining human-inspired prompting, response sampling, and consistency aggregation to measure and improve confidence calibration and failure prediction across diverse tasks.

The researchers benchmarked five major language models (including GPT-4, GPT-3.5, and LLaMA 2) across eight datasets spanning arithmetic, commonsense, symbolic, ethical, and professional knowledge domains. They tested five prompt designs (such as asking for step-by-step reasoning or multiple guesses), three sampling methods (random temperature variations, paraphrasing, and misleading cues), and several mathematical aggregation rules to compute uncertainty from multiple outputs.

The investigation produced five key findings. First, when asked directly, models display severe overconfidence, routinely assigning 80% to 100% confidence even to incorrect answers and mimicking human conversational rounding by expressing values in multiples of five. Second, while larger models like GPT-4 show improved calibration compared to older models like GPT-3, their ability to flag their own incorrect answers remains poor, with failure prediction metrics often hovering near random guessing. Third, human-inspired prompting strategies (such as eliciting multiple ranked guesses) improve calibration mainly by boosting task accuracy rather than refining the model’s true self-awareness of error. Fourth, sampling multiple responses significantly improves error detection—for instance, boosting failure prediction accuracy from near-random to over 92% on arithmetic tasks. Fifth, hybrid aggregation methods combining response consistency with stated verbal confidence consistently outperform methods relying solely on answer agreement.

These results demonstrate that a single query cannot reliably convey model uncertainty, posing operational and compliance risks if deployed blindly in high-stakes environments. While black-box methods narrow the gap with traditional internal-probability techniques, models still struggle substantially on complex tasks requiring specialized expertise, such as professional law.

For practical deployment, the article recommends a balanced, stable approach: prompt the model to generate multiple top guesses, sample several random outputs (around four to five responses to balance cost and accuracy), and aggregate them using hybrid confidence-weighting strategies. Decision-makers should apply caution, as the evaluations were limited to question-answering tasks with single ground-truth answers rather than open-ended text generation. Stated confidence scores should not be treated as objective probabilities without ensemble validation.

  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its earlier tests of whether language models can recognize correct answers establish the self-evaluation problem that this paper reexamines under black-box confidence elicitation.
Cover for Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Abstract

Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on white-box access to internal model information or model fine-tuning, have become less suitable for LLMs, especially closed-source commercial APIs. This leads to a growing need to explore the untapped area of black-box approaches for LLM uncertainty estimation. To better break down the problem, we define a systematic framework with three components: prompting strategies for eliciting verbalized confidence, sampling methods for generating multiple responses, and aggregation techniques for computing consistency. We then benchmark these methods on two key tasks-confidence calibration and failure prediction-across five types of datasets (e.g., commonsense and arithmetic reasoning) and five widely-used LLMs including GPT-4 and LLaMA 2 Chat. Our analysis uncovers several key insights: 1) LLMs, when verbalizing their confidence, tend to be overconfident, potentially imitating human patterns of expressing confidence. 2) As model capability scales up, both calibration and failure prediction performance improve. 3) Employing our proposed strategies, such as human-inspired prompts, consistency among multiple responses, and better aggregation strategies can help mitigate this overconfidence from various perspectives. 4) Comparisons with white-box methods indicate that while white-box methods perform better, the gap is narrow, e.g., 0.522 to 0.605 in AUROC. Despite these advancements, none of these techniques consistently outperform others, and all investigated methods struggle in challenging tasks, such as those requiring professional knowledge, indicating significant scope for improvement. We believe this study can serve as a strong baseline and provide insights for eliciting confidence in black-box LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Exploring Black-box Framework for Confidence Elicitation
  • 3.1 Motivation of The Framework
  • 3.2 Prompting Strategy
  • 3.3 Sampling Strategy
  • 3.4 Aggregation Strategy
  • 4 Experiment Setup
  • 5 Evaluation and Analysis
  • 5.1 LLMs tend to be overconfident when verbalizing their confidence
  • 5.2 Human-inspired Prompting Strategies Partially Reduce Overconfidence
  • 5.3 Variance Among Multiple Responses Improves Failure Prediction
  • 5.4 Introducing Verbalized Confidence Into The Aggregation Outperforms Consistency-only Aggregation
  • 6 Discussions
  • References
  • A Proof of Proposition 3.1
  • B Detailed Experiment Results
  • B.1 White-box methods outperform black-box methods, but the gap is narrow.
  • B.2 How much does the role-play prompt affect the performance?
  • B.3 How is the distribution of Vanilla Verbalized Confidence Across Models and Datasets?
  • B.4 Detailed Performance of Different Prompting Strategies
  • B.5 Top-K Verbalized Confidence Performance
  • B.6 Impact of Misleading Prompts in Misleading Sampling Strategy
  • B.7 Impact of the Number of Candidate Answers
  • B.8 Performance of different confidence elicitation methods
  • C Related Works
  • D Best Practice and Recommendations For Practitioners
  • D.1 What is the recommendation for practitioners?
  • D.2 What are the considerations when using black-box confidence elicitation algorithms?
  • D.3 Discussions on why some strategies work, and why some do not work
  • E Experiment Setup
  • E.1 Datasets
  • E.2 Evaluation Metrics
  • E.3 Models
  • E.4 Implementation Details
  • F Prompts

Knowls

  1. Knowl 1 — Black-box confidence elicitation as a three-component pipeline

    model/method

    The paper frames black-box confidence elicitation as a pipeline with three independently selectable components: (1) a prompt that elicits an answer and possibly verbalized confidence, (2) a sampling strategy that obtains one or more responses to the same question, and (3) an aggregation strategy that produces a final answer and confidence from those responses. Combining component choices defines an elicitation method. When multiple candidate answers are available, the pipeline selects the answer with the highest aggregated confidence. The framework is designed for settings where the model is queried through textual inputs and outputs, without fine-tuning or access to internal representations.

  2. Knowl 2 — Prompt strategies for eliciting verbalized confidence

    model/method

    The paper evaluates five prompt styles, appending the explanation that confidence means how likely the answer is true: Vanilla asks for an answer and its confidence; Chain-of-Thought (CoT) asks for step-by-step analysis followed by an answer and confidence; Self-Probing first obtains an answer, then in an independent chat session asks how likely that specific answer is correct; Multi-Step asks the model to split the reasoning into nn steps, assign confidence CiC_i to each step, and report overall confidence as Cmulti-step=∏i=1nCiC_{\text{multi-step}}=\prod_{i=1}^{n} C_i; Top-K requests the model’s KK best guesses and the probability that each is correct. Self-Probing is intended to encourage error detection by evaluating an answer separately from generating it; Top-K makes alternative answers explicit, while Multi-Step aggregates confidence across reasoning steps.

  3. Knowl 3 — Strategies for sampling multiple responses

    model/method

    Three black-box sampling strategies are evaluated. Self-Random submits the same prompt repeatedly and uses generation randomness to obtain varied answers; the experiments use temperature 0.70.7. Prompting paraphrases the question to induce response variation. Misleading adds a potentially incorrect hint, such as a tentative claim about the answer, and uses the model’s response to that hint as evidence about its uncertainty. The standard multiple-response experiments use M=5M=5 responses. Misleading hints include weak claims, strong claims, and appeals to external sources; the paper reports that weak-claim hints generally work better than the other hint groups in the tested StrategyQA setting.

  4. Knowl 4 — Consistency and Avg-Conf aggregation

    equation

    Let Y~\tilde{Y} be an original answer to a question, and let Y^i\hat{Y}_i be the answer in sampled response ii, for i=1,…,Mi=1,\ldots,M. The Consistency confidence is the fraction of sampled answers that match the original answer:

    Cconsistency=1M∑i=1M1{Y^i=Y~}.C_{\text{consistency}}=\frac{1}{M}\sum_{i=1}^{M}\mathbf{1}\{\hat{Y}_i=\tilde{Y}\}.

    For Avg-Conf, let CiC_i be the verbalized confidence associated with Y^i\hat{Y}_i, represented on a common numeric scale. The original answer’s confidence is the confidence-weighted fraction of sampled answers that agree with it:

    Cconf=∑i=1M1{Y^i=Y~}Ci∑i=1MCi.C_{\text{conf}}=\frac{\sum_{i=1}^{M}\mathbf{1}\{\hat{Y}_i=\tilde{Y}\}C_i}{\sum_{i=1}^{M}C_i}.

    Consistency uses answer agreement alone; Avg-Conf also uses the sampled responses’ verbalized confidences.

  5. Knowl 5 — Pair-Rank estimates answer probabilities from Top-K orderings

    model/method

    Pair-Rank is an aggregation method for Top-K responses. Assume each Top-K list is an ordered sample without replacement from a categorical distribution over possible answers. Let A\mathcal{A} be the set of distinct answers observed, and let pap_a be the probability assigned to answer a∈Aa\in\mathcal{A}, with pa≥0p_a\geq 0 and ∑a∈Apa=1\sum_{a\in\mathcal{A}}p_a=1. For two answers a,ba,b that occur in a generated list, the model-implied probability that aa is ranked ahead of bb, conditional on observing at least one of the pair, is

    P(a≻b∣a or b is observed)=papa+pb.P(a\succ b\mid a\text{ or }b\text{ is observed})=\frac{p_a}{p_a+p_b}.

    Pair-Rank estimates the categorical probabilities by maximizing the likelihood of the observed pairwise orderings, equivalently minimizing the negative log-likelihood −∑(a,b)∈Dlog⁡ ⁣(papa+pb)-\sum_{(a,b)\in\mathcal{D}}\log\!\left(\frac{p_a}{p_a+p_b}\right), where D\mathcal{D} contains the observed ordered pairs with the higher-ranked answer listed first. The simplex constraint can be enforced by parameterizing pp with a softmax; the paper proposes gradient-based optimization. The resulting distribution supplies answer confidences, with the most probable answer serving as the top-ranked choice.

  6. Knowl 6 — Benchmark scope and evaluation protocol

    experimental setup

    The evaluation covers eight question-answering datasets across five task types: commonsense reasoning (Sports Understanding and StrategyQA), arithmetic reasoning (GSM8K and SVAMP), symbolic reasoning (Date Understanding and Object Counting), professional knowledge (Professional Law), and ethical knowledge (Business Ethics). The models include Vicuna 13B, GPT-3 175B, GPT-3.5-turbo, GPT-4, and LLaMA 2 70B. The study evaluates confidence calibration using Expected Calibration Error (ECE) and correctness discrimination, called failure prediction, using AUROC. It also reports AUPRC-Positive and AUPRC-Negative to assess detection of correct and incorrect answers, respectively. The benchmark focuses on question-answering tasks with a unique ground-truth answer; it does not evaluate open-ended generation such as summarization.

  7. Knowl 7 — Vanilla verbalized confidence is overconfident and weak at failure prediction

    empirical result

    Across models and tasks, vanilla verbalized confidence commonly lies between 80% and 100%, often in increments of five percentage points; incorrect answers also occur among responses assigned 100% confidence. In the reported averages across eight datasets, ECE (reported as a percentage) was 52.0 for GPT-3, 46.1 for Vicuna, 43.6 for LLaMA 2, 37.7 for GPT-3.5, and 18.0 for GPT-4. Corresponding average AUROCs were 51.3, 52.5, 56.4, 55.1, and 62.7. Thus GPT-4 had better average calibration and failure prediction than the other reported models, but its AUROC remained relatively close to the 50% random-ranking level. The results establish a tendency toward overconfidence, while the authors cautiously suggest that familiar human patterns of expressing confidence may contribute to it.

  8. Knowl 8 — Prompting can improve accuracy and calibration without reliably improving discrimination

    empirical result

    Across the evaluated prompting strategies, no single prompt consistently performs best on every dataset and model. Human-inspired prompting often improves accuracy and lowers ECE relative to vanilla prompting, with diminishing gains on more capable models; the paper identifies Self-Probing as the most consistent improvement over vanilla for GPT-4 and Top-K as the strongest overall prompt for GPT-3.5. Improvements in calibration do not necessarily mean that confidence better separates correct from incorrect answers. For example, on GSM8K, GPT-4 with CoT reached 93.6% accuracy and ECE 0.064 while assigning 100% confidence to every sample, leaving no confidence variation with which to rank correct and incorrect answers. On GPT-3.5 GSM8K, CoT increased accuracy from 28% to 80.3% and reduced ECE from 0.66 to 0.10, but AUROC fell from 0.65 to 0.55.

  9. Knowl 9 — Multiple sampled responses substantially improve failure prediction

    empirical result

    With GPT-3.5, five responses combined with consistency aggregation outperformed single-response verbalized confidence on average in both ECE and AUROC for the reported comparison: self-random sampling achieved ECE 18.7 and AUROC 73.0, misleading sampling 17.3 and 69.6, and paraphrased-prompt sampling 24.3 and 69.2; the single-response CoT baseline achieved 25.0 and 56.4. On GSM8K specifically, self-random sampling with five responses reached AUROC 92.7 and ECE 6.28, compared with 54.8 and 10.1 for the single-response CoT baseline. Increasing the response count generally improved AUROC, with diminishing gains at larger counts; the paper notes that query cost grows approximately linearly with the number of responses.

  10. Knowl 10 — Aggregation methods trade off calibration and failure prediction, and performance remains limited

    empirical result

    In the GPT-4 experiment combining Top-K prompting with self-random sampling, Pair-Rank achieved the lowest mean ECE among the compared aggregators: 6.90 (reported metrics are multiplied by 10210^2), versus 12.0 for Consistency and 14.8 for Avg-Conf. The reported mean AUROC values were 67.6 for Pair-Rank, 66.9 for Consistency, and 66.9 for Avg-Conf; the authors recommend Pair-Rank when calibrated confidence values are the priority and Avg-Conf when failure prediction is the priority. They recommend the practical combination Top-K prompt + self-random sampling + Avg-Conf or Pair-Rank aggregation. In a separate comparison on five datasets, white-box token-probability methods generally performed better than black-box verbalized confidence, but the reported AUROC gap was modest (0.522 to 0.605), and performance remained unsatisfactory. The approaches also struggled on tasks requiring specialized knowledge, including Professional Law.

Coverage note — The exhaustive per-dataset result matrices and minor role-play and misleading-hint ablations are omitted; the knowls retain the central framework, representative quantitative findings, and the paper’s principal limitations.

References

  1. 1.Kendrick Boyd, Kevin H. Eng, and C. David Page. Area under the precision-recall curve: Point estimates and confidence intervals. In Hendrik Blockeel, Kristian Kersting, Siegfried Nijssen, and Filip Železný (eds.), Machine Learning and Knowledge Discovery in Databases, pp. 451–466, Berlin, Heidelberg, 2013. Springer Berlin Heidelberg. ISBN 978-3-642-40994-3.
  2. 2.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  3. 3.Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. A close look into the calibration of pre-trained language models. arXiv preprint arXiv:2211.00151, 2022.
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  6. 6.Leda Cosmides and John Tooby. Are humans good intuitive statisticians after all? rethinking some conclusions from the literature on judgment under uncertainty. cognition, 58(1):1–73, 1996.
  7. 7.Ailin Deng, Miao Xiong, and Bryan Hooi. Great models think alike: Improving model reliability via inter-model latent agreement. arXiv preprint arXiv:2305.01481, 2023.
  8. 8.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pp. 1050–1059. PMLR, 2016.
  9. 9.Paul H Garthwaite, Joseph B Kadane, and Anthony O’Hagan. Statistical methods for eliciting probability distributions. Journal of the American statistical Association, 100(470):680–701, 2005a.
  10. 10.Paul H Garthwaite, Joseph B Kadane, and Anthony O’Hagan. Statistical methods for eliciting probability distributions. Journal of the American statistical Association, 100(470):680–701, 2005b.
  11. 11.Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. A survey of uncertainty in deep neural networks. arXiv preprint arXiv:2107.03342, 2021.
  12. 12.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies, 2021.
  13. 13.Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen. Bigbench: Towards an industry standard benchmark for big data analytics. In Proceedings of the 2013 ACM SIGMOD international conference on Management of data, pp. 1197–1208, 2013.
  14. 14.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International conference on machine learning, pp. 1321–1330. PMLR, 2017.
  15. 15.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding, 2021.
  16. 16.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977, 2021.
  17. 17.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  18. 18.Ethan Kim. Sports understanding in bigbench, 2021.
  19. 19.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. ArXiv, abs/2205.11916, 2022. URL https://api.semanticscholar.org/CorpusID:249017743.
  20. 20.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664, 2023.
  21. 21.Volodymyr Kuleshov and Shachi Deshpande. Calibrated and sharp uncertainties in deep learning via density estimation. In International Conference on Machine Learning, pp. 11683–11693. PMLR, 2022.
  22. 22.Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon. Accurate uncertainties for deep learning using calibrated regression. In International conference on machine learning, pp. 2796–2804. PMLR, 2018.
  23. 23.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30, 2017.
  24. 24.Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334, 2022.
  25. 25.Andrey Malinin and Mark Gales. Uncertainty estimation in autoregressive structured prediction. arXiv preprint arXiv:2002.07650, 2020.
  26. 26.Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872, 2022.
  27. 27.Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic. Revisiting the calibration of modern neural networks. In Advances in Neural Information Processing Systems, volume 34, pp. 15682–15694, 2021.
  28. 28.Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence, volume 29, 2015.
  29. 29.OpenAI. ChatGPT. https://www.openai.com/gpt-3/, 2021. Accessed: April 21, 2023.
  30. 30.OpenAI. Gpt-4 technical report, 2023.
  31. 31.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2080–2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168. URL https://aclanthology.org/2021.naacl-main.168.
  32. 32.Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. Out-of-distribution detection and selective generation for conditional language models. arXiv preprint arXiv:2209.15558, 2022.
  33. 33.Quintin P. Solano, Laura Hayward, Zoey Chopra, Kathryn Quanstrom, Daniel Kendrick, Kenneth L. Abbott, Marcus Kunzmann, Samantha Ahle, Mary Schuller, Erkin Ötle¸s, and Brian C. George. Natural language processing and assessment of resident feedback quality. Journal of Surgical Education, 78(6):e72–e77, 2021. ISSN 1931-7204. doi: https://doi.org/10.1016/j.jsurg.2021.05.012. URL https://www.sciencedirect.com/science/article/pii/S1931720421001537.
  34. 34.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=uyTL5Bvosj.
  35. 35.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975, 2023.
  36. 36.Christian Tomani and Florian Buettner. Towards trustworthy predictions from deep neural networks with fast adversarial calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 9886–9896, 2021.
  37. 37.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models, 2023a.
  38. 38.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  39. 39.Jianfeng Wang, Rong Xiao, Yandong Guo, and Lei Zhang. Learning to count objects with few exemplar annotations. arXiv preprint arXiv:1905.07898, 2019.
  40. 40.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Annual Meeting of the Association for Computational Linguistics, 2023. URL https://api.semanticscholar.org/CorpusID:258558102.
  41. 41.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  42. 42.Xinyi Wu and Zijian Wang. Data understanding in bigbench, 2021.
  43. 43.Yijun Xiao and William Yang Wang. On hallucination and predictive uncertainty in conditional language generation. arXiv preprint arXiv:2103.15025, 2021.
  44. 44.Miao Xiong, Shen Li, Wenjie Feng, Ailin Deng, Jihai Zhang, and Bryan Hooi. Birds of a feather trust together: Knowing when to trust a classifier via adaptive neighborhood aggregation. arXiv preprint arXiv:2211.16466, 2022.
  45. 45.Miao Xiong, Ailin Deng, Pang Wei Koh, Jiaying Wu, Shen Li, Jianqing Xu, and Bryan Hooi. Proximity-informed calibration for deep neural networks. arXiv preprint arXiv:2306.04590, 2023.
  46. 46.Zhuoning Yuan, Yan Yan, Milan Sonka, and Tianbao Yang. Large-scale robust deep auc maximization: A new surrogate loss and empirical studies on medical image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3040–3049, 2021.
  47. 47.Bianca Zadrozny and Charles Elkan. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml, volume 1, pp. 609–616, 2001.
  48. 48.Jize Zhang, Bhavya Kailkhura, and T Yong-Jin Han. Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning. In International conference on machine learning, pp. 11117–11128. PMLR, 2020.
  49. 49.Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. Navigating the grey area: Expressions of overconfidence and uncertainty in language models. arXiv preprint arXiv:2302.13439, 2023.

Citation

MLA
Xiong, M., et al. “Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”. arXiv, 2023, http://arxiv.org/abs/2306.13063v2.
APA
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., & Hooi, B. (2023). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. arXiv. http://arxiv.org/abs/2306.13063v2
Chicago
Xiong, M., Z. Hu, X. Lu, et al. 2023. “Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”. arXiv. http://arxiv.org/abs/2306.13063v2.
Harvard
Xiong, M. et al. (2023) “Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.13063v2.
Vancouver
1. Xiong M, Hu Z, Lu X, Li Y, Fu J, He J, Hooi B (2023) Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. arXiv

BibTeX

@article{xiong2023can,
  title = {Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs},
  author = {Xiong, Miao and Hu, Zhiyuan and Lu, Xinyang and Li, Yifei and Fu, Jie and He, Junxian and Hooi, Bryan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.13063v2},
  eprint = {2306.13063}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors