Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness

Jiuhai ChenJonas Mueller

article2024ACL133 citations

Introduces BSDETECTOR, a black-box uncertainty quantification method combining consistency sampling and self-reflection to accurately detect incorrect outputs and select the most reliable answers from any large language model without additional training.

Listen

Large Language Models often produce incorrect or fabricated information with high apparent confidence, presenting significant operational and compliance risks in high-stakes enterprise applications. Most commercial models are accessed via proprietary, black-box programming interfaces that do not expose internal token probabilities or training data, making traditional statistical calibration methods impractical.

The article evaluates BSDETECTOR, a training-free framework designed to quantify uncertainty and generate numerical confidence scores for responses produced by any black-box language model. The core objective is to identify inaccurate outputs and improve the overall reliability of generative systems and automated evaluations.

The approach evaluates model outputs through two complementary mechanisms without modifying underlying model weights. First, an extrinsic Observed Consistency score generates multiple alternative responses using chain-of-thought prompting at higher sampling temperatures, using a natural language inference model to detect semantic contradictions against the original output. Second, an intrinsic Self-Reflection Certainty score prompts the model to evaluate the correctness of its own response across structured multiple-choice questions. The article tested this approach on standard arithmetic, reasoning, and open-domain factual benchmarks, including GSM8K, SVAMP, CSQA, and TriviaQA, using OpenAI models such as GPT-3.5 Turbo and GPT-4.

The evaluation produced four key findings. First, BSDETECTOR significantly outperformed likelihood-based and standard temperature-sampling baselines in distinguishing correct from incorrect answers, achieving high area under the receiver operating characteristic curve scores between 0.769 and 0.951 across benchmarks. Second, sampling multiple responses and selecting the one with the highest confidence score consistently improved model accuracy, raising standard-prompting accuracy on GSM8K math problems from 47% to 70%. Third, in automated evaluations conducted by GPT-4, routing the lowest-confidence assessments to human reviewers substantially reduced evaluation error compared to random sampling. Fourth, in fully automated evaluation settings, discarding the 20% of model evaluations with the lowest confidence eliminated extreme errors and achieved near-perfect alignment with human ground truth.

These findings indicate that uncertainty quantification enables organizations to deploy language models more safely in high-value workflows. By establishing clear confidence thresholds, systems can flag hallucinations, automate fallback behaviors such as querying secondary models or human escalation, and lower operational risk in document drafting or automated grading without model fine-tuning.

Organizations deploying generative artificial intelligence should implement confidence-scoring wrappers to route uncertain responses to human reviewers or fallback processes. Teams utilizing language models for automated evaluation should discard or manually review the bottom 10% to 20% lowest-confidence evaluations to ensure benchmark integrity. Generating multiple candidate responses and selecting the highest-confidence answer should be considered when response accuracy is critical.

The primary operational limitation is increased computational cost and latency, as generating five candidate responses and follow-up reflection prompts multiplies programming interface calls. While confidence in the experimental benchmarks is high, performance across non-OpenAI model architectures and highly specialized industry domains requires further pilot testing before full-scale deployment.

arXiv: 2308.16175
Cover for Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness

Abstract

We introduce BSDETECTOR, a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a numeric confidence score for any output it generated. Our uncertainty quantification technique works for any LLM accessible only via a black-box API, whose training data remains unknown. By expending a bit of extra computation, users of any LLM API can now get the same response as they would ordinarily, as well as a confidence estimate that cautions when not to trust this response. Experiments on both closed and open-form Question-Answer benchmarks reveal that BSDETECTOR more accurately identifies incorrect LLM responses than alternative uncertainty estimation procedures (for both GPT-3 and ChatGPT). By sampling multiple responses from the LLM and considering the one with the highest confidence score, we can additionally obtain more accurate responses from the same LLM, without any extra training steps. In applications involving automated evaluation with LLMs, accounting for our confidence scores leads to more reliable evaluation in both human-in-the-loop and fully-automated settings (across both GPT 3.5 and 4).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 BSDETECTOR uncertainty estimation
  • 3.1 Observed Consistency
  • 3.2 Self-reflection Certainty
  • 3.3 Overall Confidence Estimate
  • 4 Application: Generating More Reliable Answers from any LLM
  • 5 Application: More reliable LLM-based (automated) evaluation
  • 6 Experiments
  • 6.1 Calibration of uncertainty estimates
  • 6.2 Generating More Reliable Answers from any LLM
  • 6.3 More reliable LLM-based (automated) evaluation
  • 7 Comparison with Related Work and Further Impact of our Work
  • 8 Discussion
  • Limitations
  • References
  • A Appendix
  • A.1 Details about NLI model
  • A.2 Compute costs
  • A.3 Prompts used in BSDETECTOR
  • A.4 Ablation Study
  • A.4.1 Increasing the number of outputs and integrating CoT prompt introduce more diversity?
  • A.4.2 Effect of different sentence similarity metrics

Knowls

  1. Knowl 1 — BSDETECTOR black-box confidence estimation

    model/method

    BSDETECTOR wraps the inference of a stochastic text-to-text language model accessible through a black-box API and assigns a confidence score to an answer without requiring training-data access, token probabilities, model fine-tuning, or a validation set. For a question xx, the ordinary model call produces a reference answer yy. Additional API calls produce varied answers and ask the model to assess the correctness of yy. The method combines an extrinsic measure, Observed Consistency OO, with an intrinsic measure, Self-reflection Certainty SS:

    C(x,y)=βO+(1−β)S,0≤β≤1,C(x,y)=\beta O+(1-\beta)S, \qquad 0\leq\beta\leq 1,

    where C(x,y)∈[0,1]C(x,y)\in[0,1] is the confidence assigned to answer yy, and β\beta controls the relative reliance on consistency versus self-reflection. Lower confidence is intended to identify answers that are more likely to be factually incorrect, speculative, or otherwise unsuitable. In the reported experiments, the reference answer used temperature 00, while five varied answers were sampled at temperature 1.01.0 using a Chain-of-Thought-augmented prompt; two follow-up questions were used for self-reflection.

  2. Knowl 2 — Observed Consistency from contradiction-aware answer sampling

    equation

    For a question xx with reference answer yy, BSDETECTOR samples kk alternative answers y1,…,yky_1,\ldots,y_k using a diversity-enhancing prompt. The similarity between alternative answer yiy_i and yy is based on an off-the-shelf natural-language-inference classifier. Let pip_i be the classifier's contradiction probability for the ordered pair (yi,y)(y_i,y) and pi′p_i' the contradiction probability for the reverse order (y,yi)(y,y_i). The order-symmetrized semantic similarity is

    si=12[(1−pi)+(1−pi′)].s_i=\frac{1}{2}\left[(1-p_i)+(1-p_i')\right].

    Let ri=1[yi=y]r_i=\mathbf{1}[y_i=y] be an exact-match indicator, where ri=1r_i=1 when the two text answers are identical and ri=0r_i=0 otherwise. For a trade-off parameter α∈[0,1]\alpha\in[0,1], the per-sample consistency score and aggregate Observed Consistency are

    oi=αsi+(1−α)ri,O=1k∑i=1koi.o_i=\alpha s_i+(1-\alpha)r_i, \qquad O=\frac{1}{k}\sum_{i=1}^{k}o_i.

    The NLI classifier distinguishes entailment, neutrality, and contradiction; using 1−pi1-p_i rather than general embedding similarity emphasizes whether an alternative answer contradicts the reference answer. The exact-match term stabilizes the score for short closed-form answers, where NLI contradiction probabilities may be unreliable. The paper used Chain-of-Thought prompt augmentation to produce the varied answers and used k=5k=5 samples at temperature 1.01.0 in its main experiments.

  3. Knowl 3 — Self-reflection certainty and final confidence score

    equation

    BSDETECTOR estimates intrinsic confidence by asking the same language model to reconsider whether its previously generated answer yy to question xx is correct. Each of qq follow-up assessments must choose one of three categories: Correct, Incorrect, or I am not sure. These categories receive numerical values 1.01.0, 0.00.0, and 0.50.5, respectively. If zj∈{0,0.5,1}z_j\in\{0,0.5,1\} is the score from follow-up assessment jj, the Self-reflection Certainty is

    S=1q∑j=1qzj.S=\frac{1}{q}\sum_{j=1}^{q}z_j.

    The final confidence score is

    C=βO+(1−β)S,C=\beta O+(1-\beta)S,

    where OO is the contradiction-aware consistency score, β∈[0,1]\beta\in[0,1] is the weighting parameter, and C∈[0,1]C\in[0,1]. The categorical scale is used because the authors observed that directly requesting a continuous confidence value from 00 to 100100 usually produced excessively high responses above 9090. The main experiments used q=2q=2 follow-up assessments.

  4. Knowl 4 — Selecting the most trustworthy answer among sampled candidates

    algorithm

    BSDETECTOR can improve answer accuracy without retraining the language model by generating multiple candidate answers and returning the candidate with the largest confidence score. The procedure uses five candidate answers and five varied answers for the consistency calculation in the reported experiments.

    Input: question xx; candidate count m=5m=5; consistency-sample count k=5k=5
    Output: selected answer y∗y^*
    Generate candidate answers y1′,...,ym′y'_1, ..., y'_m from the same language model using temperature sampling.
    Generate varied detector answers y1,...,yky_1, ..., y_k with the diversity-enhancing prompt at temperature 1.01.0.
    For each candidate answer yj′y'_j:
        Compare yj′y'_j with every detector answer yiy_i using the contradiction-aware similarity and exact-match indicator.
        Average the resulting scores to obtain the candidate's Observed Consistency score OjO_j.
        Ask the language model the self-reflection follow-up questions about yj′y'_j.
        Convert Correct, Incorrect, and I am not sure to 1.01.0, 0.00.0, and 0.50.5, then average them to obtain SjS_j.
        Compute Cj=βOj+(1−β)SjC_j=\beta O_j+(1-\beta)S_j.
    Return y∗=yj∗′y^*=y'_{j^*} where j∗=argmax⁡jCjj^*=\operatorname{argmax}_j C_j.

    The detector samples used to calculate consistency are reused across candidates, while self-reflection is performed separately for each candidate. Thus an alternative answer is selected only when the model's sampled answers contradict it less often and the model reports greater certainty about it.

  5. Knowl 5 — Calibration of confidence estimates across question-answering tasks

    data/table

    The paper evaluated whether higher BSDETECTOR confidence ranks correct answers above incorrect answers. It used Text-Davinci-003 and GPT-3.5 Turbo on GSM8K and SVAMP grade-school mathematics, CSQA commonsense reasoning, and TriviaQA open-domain factual questions. Performance was measured by AUROC, where higher values mean better separation of correct and incorrect answers, 1.01.0 is ideal, and 0.50.5 corresponds to random ranking. Likelihood-based uncertainty was unavailable for GPT-3.5 Turbo because its API did not expose token probabilities. Temperature Sampling denotes a version without Chain-of-Thought prompting, self-reflection, and the exact-match indicator; Self-reflection Certainty is the intrinsic component alone.

    Model Dataset Likelihood-based Temperature sampling Self-reflection BSDETECTOR
    Text-Davinci-003 GSM8K 0.647 0.614 0.521 0.867
    Text-Davinci-003 CSQA 0.490 0.540 0.539 0.743
    Text-Davinci-003 SVAMP 0.668 0.653 0.619 0.936
    Text-Davinci-003 TriviaQA 0.708 0.769 0.653 0.828
    GPT-3.5 Turbo GSM8K – 0.660 0.831 0.951
    GPT-3.5 Turbo CSQA – 0.583 0.506 0.769
    GPT-3.5 Turbo SVAMP – 0.671 0.839 0.927
    GPT-3.5 Turbo TriviaQA – 0.689 0.655 0.817

    BSDETECTOR achieved the highest AUROC for every model-dataset combination, with especially large gains on arithmetic and reasoning tasks such as GPT-3.5 Turbo on GSM8K (0.9510.951) and SVAMP (0.9270.927).

  6. Knowl 6 — Accuracy gains from confidence-based candidate selection

    data/table

    The authors generated five candidate answers from GPT-3.5 Turbo and returned the candidate with the highest BSDETECTOR confidence. The baseline Reference Answer was one answer generated at temperature 00. Accuracy was measured against the known answer for each dataset, under both standard prompting and Chain-of-Thought prompting.

    Prompting Dataset Reference Answer (%) BSDETECTOR (%)
    Standard GSM8K 47 70
    Standard CSQA 72 73
    Standard SVAMP 75 82
    Standard TriviaQA 73 76
    Chain-of-Thought GSM8K 71 73
    Chain-of-Thought CSQA 74 76
    Chain-of-Thought SVAMP 78 83
    Chain-of-Thought TriviaQA 75 79

    Selecting the most confident of five answers improved accuracy on all eight dataset-prompting combinations. The largest absolute improvement was on standard-prompt GSM8K, from 47%47\% to 70%70\%.

  7. Knowl 7 — Using confidence to make LLM-based evaluation safer

    model/method

    BSDETECTOR can estimate the trustworthiness of an LLM evaluator rather than only the trustworthiness of an answer generator. For each generated answer, the evaluator LLM receives the question and answer, produces a correctness or quality judgment, and is then assessed by BSDETECTOR. Let AA be the subset of evaluated answers whose evaluator confidence scores are lowest.

    In a human-in-the-loop workflow, humans inspect the answers in AA and replace the evaluator's judgments for those cases, while the remaining cases retain automated judgments. This concentrates scarce human effort on evaluations most likely to be unreliable. In a fully automated workflow, the system removes AA from the evaluation set and averages the remaining evaluator scores, avoiding the uncertain judgments at the cost of a smaller sample.

    The experiments used GPT-4 as evaluator on TriviaQA, where it produced binary Correct/Incorrect judgments, and on Summarize-from-feedback, where it rated summaries as Bad, Fair, Good, or Excellent. For the latter task, the four labels were mapped to 1,2,3,41,2,3,4 and evaluated using mean squared error. Selecting low-confidence cases for human review produced higher evaluation accuracy on TriviaQA and lower evaluation MSE on Summarize-from-feedback than randomly selecting the same number of cases.

  8. Knowl 8 — Fully automated filtering improves aggregate evaluator reliability

    empirical result

    The fully automated evaluation experiment repeatedly assessed whether discarding uncertain GPT-4 judgments improved the estimated average quality of a collection. For each replicate, 500 answers were evaluated. The full-dataset baseline averaged all 500 GPT-4 scores. The confidence-filtered estimator removed the 20% of answers with the lowest BSDETECTOR confidence, then averaged the remaining 400 scores. A control estimator randomly removed 20% of answers.

    The procedure was repeated 500 times on both TriviaQA and Summarize-from-feedback. On TriviaQA, removing the lowest-confidence evaluations consistently produced an average GPT-4 accuracy of 1.01.0 in the reported histogram, whereas the full-data and randomly filtered estimators were less reliable. On Summarize-from-feedback, confidence-based removal consistently reduced the mean squared error of the average GPT-4 score relative to both retaining the full dataset and randomly removing 20%. The result indicates that excluding highly uncertain evaluations can improve aggregate reliability enough to justify the variance introduced by averaging over fewer examples.

  9. Knowl 9 — Ablation evidence for sample diversity and similarity design

    data/table

    Ablation experiments tested whether BSDETECTOR benefits from more sampled outputs, Chain-of-Thought prompt augmentation, and contradiction-aware similarity. The metric is AUC/AUROC for distinguishing correct from incorrect answers; higher values are better. Increasing the number of sampled outputs from five to ten improved three of four datasets, but the gains were generally small, so the main experiments used five outputs. Removing Chain-of-Thought prompting reduced performance on every dataset. The complete results are:

    AUC for different numbers of outputs
    Dataset 5 outputs 10 outputs
    GSM8K 0.951 0.961
    CSQA 0.769 0.802
    SVAMP 0.927 0.937
    TriviaQA 0.817 0.814
    AUC without and with Chain-of-Thought augmentation
    Dataset Remove CoT prompting BSDETECTOR
    GSM8K 0.837 0.951
    CSQA 0.665 0.769
    SVAMP 0.882 0.927
    TriviaQA 0.792 0.817

    The choice of similarity metric also mattered. Jaccard similarity, cosine similarity of LLM embeddings, NLI similarity defined as 1−pcontradiction1-p_{\mathrm{contradiction}}, and the complete BSDETECTOR similarity design produced the following AUC values:

    Dataset Jaccard LLM embedding NLI (1−contradiction)(1-\text{contradiction}) BSDETECTOR
    GSM8K 0.896 0.866 0.892 0.951
    CSQA 0.857 0.849 0.727 0.769
    SVAMP 0.917 0.888 0.901 0.927
    TriviaQA 0.650 0.642 0.794 0.817

    These results support the use of diverse Chain-of-Thought-generated alternatives together with contradiction-aware NLI similarity and the exact-match indicator, rather than relying on generic lexical or embedding similarity alone.

  10. Knowl 10 — Limitations of black-box confidence estimation

    limitation

    The paper identifies four main limitations of BSDETECTOR. First, estimating confidence requires multiple sampled responses and follow-up calls, increasing computation and API cost compared with a single model inference. Second, the method's generalization across language-model architectures, domains, specialized knowledge requirements, and question types has not been fully explored. Third, black-box APIs restrict the depth of analysis available because users cannot inspect internal representations, training data, or token probabilities. Fourth, performance may be difficult to maintain in resource-constrained settings or when the model's sampled alternatives and self-reflections are themselves systematically unreliable.

Coverage note — Ancillary attorney-drafting and chatbot demonstrations, exact prompt-template text, and related-work discussion were omitted because they are illustrative or procedural rather than additional load-bearing contributions.

References

  1. 1.Anastasios N. Angelopoulos and Stephen Bates. 2021. A gentle introduction to conformal prediction and distribution-free uncertainty quantification. CoRR, abs/2107.07511.
  2. 2.Harrison Chase. 2022. LangChain.
  3. 3.Jiuhai Chen, Lichang Chen, Heng Huang, and Tianyi Zhou. 2023a. When do you need chain-of-thought prompting for chatgpt? CoRR, abs/2304.03262.
  4. 4.Jiuhai Chen, Lichang Chen, Chen Zhu, and Tianyi Zhou. 2023b. How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 11149–11159. Association for Computational Linguistics.
  5. 5.Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang, and Tianyi Zhou. 2023c. Instructzero: Efficient instruction optimization for black-box large language models. CoRR, abs/2306.03082.
  6. 6.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168.
  7. 7.Meire Fortunato, Charles Blundell, and Oriol Vinyals. 2017. Bayesian recurrent neural networks. CoRR, abs/1704.02798.
  8. 8.Yarin Gal and Zoubin Ghahramani. 2016a. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, volume 48 of JMLR Workshop and Conference Proceedings, pages 1050–1059. JMLR.org.
  9. 9.Yarin Gal and Zoubin Ghahramani. 2016b. A theoretically grounded application of dropout in recurrent neural networks. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages 1019–1027.
  10. 10.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR.
  11. 11.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  12. 12.Dan Hendrycks and Kevin Gimpel. 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  13. 13.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12):248:1–248:38.
  14. 14.Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1601–1611. Association for Computational Linguistics.
  15. 15.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know. CoRR, abs/2207.05221.
  16. 16.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  17. 17.Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon. 2018. Accurate uncertainties for deep learning using calibrated regression. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 2801–2809. PMLR.
  18. 18.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6402–6413.
  19. 19.Shiyu Liang, Yixuan Li, and R. Srikant. 2018. Enhancing the reliability of out-of-distribution image detection in neural networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  20. 20.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. Trans. Mach. Learn. Res., 2022.
  21. 21.Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models. CoRR, abs/2305.19187.
  22. 22.Andrey Malinin and Mark J. F. Gales. 2021. Uncertainty estimation in autoregressive structured prediction. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  23. 23.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 9004–9017. Association for Computational Linguistics.
  24. 24.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems?
  25. 25.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with GPT-4. CoRR, abs/2304.03277.
  26. 26.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  27. 27.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge.
  28. 28.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  29. 29.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 5433–5442. Association for Computational Linguistics.
  30. 30.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  31. 31.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  32. 32.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. CoRR, abs/2304.12244.

Citation

MLA
Chen, J., and J. Mueller. “Quantifying Uncertainty in Answers from Any Language Model and Enhancing Their Trustworthiness”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5186–200, https://doi.org/10.18653/v1/2024.acl-long.283.
APA
Chen, J., & Mueller, J. (2024). Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5186–5200. https://doi.org/10.18653/v1/2024.acl-long.283
Chicago
Chen, J., and J. Mueller. 2024. “Quantifying Uncertainty in Answers from Any Language Model and Enhancing Their Trustworthiness”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5186–5200. https://doi.org/10.18653/v1/2024.acl-long.283.
Harvard
Chen, J. and Mueller, J. (2024) “Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5186–5200. Available at: https://doi.org/10.18653/v1/2024.acl-long.283.
Vancouver
1. Chen J, Mueller J (2024) Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5186–5200

BibTeX

@inproceedings{chen-mueller-2024-quantifying,
    title = "Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness",
    author = "Chen, Jiuhai  and
      Mueller, Jonas",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.283/",
    doi = "10.18653/v1/2024.acl-long.283",
    pages = "5186--5200"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/