Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Katherine TianEric MitchellAllan ZhouArchit SharmaRafael RafailovHuaxiu YaoChelsea FinnChristopher D. Manning

article2023EMNLP921 citationsOutstanding Paper Award

Demonstrates that directly prompting RLHF-tuned language models like GPT-4 and Claude to verbalize their confidence scores achieves significantly better calibration than extracting their conditional probabilities, reducing expected calibration error by up to 50% across standard benchmarks.

Listen

Deploying artificial intelligence systems in high-stakes environments requires trustworthy uncertainty estimates, enabling systems to flag or defer low-confidence predictions to human experts. While base language models generally output well-calibrated internal probabilities, fine-tuning them with reinforcement learning from human feedback to better follow user instructions often degrades this calibration, causing models to assert incorrect information with unwarranted confidence.

The article evaluates practical methods for extracting calibrated confidence scores from human-feedback fine-tuned language models. It demonstrates that directly asking models to state their confidence in natural language or numbers provides significantly more reliable uncertainty estimates than using their underlying token probabilities.

To evaluate uncertainty extraction, the researchers tested five widely used models—including ChatGPT, GPT-4, Claude 1, Claude 2, and Llama-2-70B-Chat—across three factual question-answering benchmarks comprising roughly 2,800 total questions. The study compared the models' internal conditional probabilities (estimated via sampling) against direct verbalization techniques. These techniques prompted models to output numerical probabilities or descriptive likelihood phrases (such as "likely" or "almost certain") in either single-stage or two-stage dialogues, alongside prompt variants that required generating multiple answer choices or step-by-step reasoning.

The investigation yielded four central findings. First, verbalized confidence scores were consistently better-calibrated than raw internal model probabilities across leading proprietary systems, frequently reducing the expected calibration error by roughly 50%. Second, prompting models to generate and evaluate multiple candidate answers before stating a confidence score notably improved calibration, mirroring human psychological strategies for reducing overconfidence. Third, language models verbalized numerical probabilities with equal or greater accuracy than descriptive text phrases. Fourth, incorporating chain-of-thought reasoning into prompts failed to improve confidence calibration.

These findings have direct operational and cost implications for organizations deploying language models. Because proprietary application programming interfaces (APIs) rarely expose underlying token log probabilities and sampling multiple responses is expensive, prompting models to state numerical confidences in a single response provides a simpler, cheaper, and more reliable safeguard against automated errors and hallucinations. Notably, this ability appears naturally in instruction-following models without requiring extra calibration training.

Organizations seeking to measure model reliability should adopt prompt strategies that ask models to provide their top candidate answers alongside explicit numerical confidence ratings in a single query. When using closed commercial models, teams can apply post-hoc temperature scaling on a small validation dataset to further enhance calibration. Developers should avoid adding chain-of-thought reasoning solely for calibration purposes, as it increases latency and token costs without measurable gains.

These conclusions carry certain limitations. The evaluation focused primarily on short-form factual question-answering rather than complex mathematical reasoning, multi-step logic, or long-form generation. Furthermore, the open-source Llama-2 model exhibited less consistent calibration improvements than proprietary systems, and the opacity of closed-source model training limits deeper analysis. Nonetheless, for short-form factual tasks in advanced commercial systems, there is high confidence that verbalized confidence scoring significantly improves prediction reliability.

arXiv: 2305.14975
  • Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Read this earlier study of language models’ self-evaluation first to understand the confidence-elicitation foundation that the source tests after human-feedback fine-tuning.
  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Its account of calibration, expected calibration error, and temperature scaling supplies the core measurement concepts the source uses to compare confidence estimates.
Cover for Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Abstract

A trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to an expert in cases of low-confidence predictions. Recent studies have shown that unsupervised pre-training produces large language models (LMs) whose conditional probabilities are remarkably well-calibrated. However, the most widely-used LMs are fine-tuned with reinforcement learning from human feedback (RLHF-LMs), and some studies have suggested that RLHF-LMs produce conditional probabilities that are very poorly calibrated. In light of this perceived weakness, we conduct a broad evaluation of methods for extracting confidence scores from RLHF-LMs. For RLHF-LMs such as ChatGPT, GPT-4, and Claude, we find that verbalized confidences emitted as output tokens are typically better-calibrated than the model's conditional probabilities on the TriviaQA, SciQ, and TruthfulQA benchmarks, often reducing the expected calibration error by a relative 50%.

Table of Contents

  • 1 Introduction
  • 2 Evaluating Calibration in RLHF-LMs
  • 3 Results
  • 4 Discussion
  • References
  • A Additional Results
  • B Fitting Procedure for Temperature and Probabilities for Linguistic Expressions
  • C Prompt Templates

Knowls

  1. Knowl 1 — Verbalized confidence outperforms conditional probabilities on GPT models

    empirical result

    On TriviaQA, SciQ, and TruthfulQA, the best tested verbalized-confidence method had lower temperature-scaled expected calibration error (ECE-t) than the conditional-probability baseline for both gpt-3.5-turbo and GPT-4. For gpt-3.5-turbo, the baseline ECE-t values were 0.097, 0.180, and 0.317; the best verbalized values were 0.045 (two-stage, top-2 guesses), 0.046 (two-stage, top-4 guesses), and 0.101 (two-stage, top-2 guesses), respectively. For GPT-4, the baseline values were 0.067, 0.165, and 0.334; the best verbalized values were 0.034 (one-stage, top-2 guesses), 0.024 (human-mapped linguistic confidence), and 0.105 (optimized linguistic confidence), respectively. These are the lowest reported ECE-t values among the tested verbalized methods for each model-dataset pair; they show a substantial calibration improvement, but do not imply that the same prompt is best in every setting.

  2. Knowl 2 — Prompting strategies for verbalized confidence

    model/method

    The study elicits confidence directly from a model’s generated text, without fine-tuning the model for confidence reporting. In a one-stage numerical prompt, the model gives an answer and its probability of being correct together. In a two-stage prompt, it first answers and then, in a follow-up turn, assigns a probability to that answer. Top-kk variants ask for kk candidate answers and a correctness probability for each, then use the candidate with the highest stated probability as the prediction. Linguistic variants instead ask for a confidence phrase from a fixed set of likelihood expressions. These strategies are useful when a deployed language model does not expose token-level probabilities.

  3. Knowl 3 — Considering multiple candidate answers often improves calibration

    empirical result

    Asking for and assessing several candidate answers before selecting a prediction often improved verbalized-confidence calibration relative to asking for only one answer, although the effect depended on the model, dataset, and prompt. For gpt-3.5-turbo on SciQ, raw ECE fell from 0.234 with one-stage top-1 to 0.065 with one-stage top-4; on TruthfulQA it fell from 0.389 to 0.203. For GPT-4, one-stage top-4 reduced raw ECE from 0.201 to 0.056 on SciQ and from 0.350 to 0.198 on TruthfulQA. The results support the paper’s finding that considering alternatives can mitigate overconfidence, but do not establish a monotonic improvement as more answers are requested.

  4. Knowl 4 — Verbalization results differ across model families

    empirical result

    Verbalized confidence improved ECE-t over the conditional-probability baseline for Claude-1 and Claude-2 on all three evaluated datasets, while Llama-2-70B-Chat showed less consistent gains across metrics. Claude-1’s baseline ECE-t values for TriviaQA, SciQ, and TruthfulQA were 0.079, 0.149, and 0.304; its best verbalized values were 0.047, 0.040, and 0.085. Claude-2’s corresponding baseline values were 0.089, 0.176, and 0.368, versus best verbalized values of 0.054, 0.026, and 0.075. Llama-2-70B-Chat’s baseline values were 0.124, 0.189, and 0.361, versus best verbalized values of 0.067, 0.032, and 0.037. These best values are selected from the verbalized methods reported for each dataset. Improvements in calibration error did not guarantee better ranking of reliable predictions: for example, Llama-2-70B-Chat’s AUC was lower for its verbalized methods than for the baseline on TriviaQA and SciQ.

  5. Knowl 5 — Evaluation uses factual QA datasets and several complementary metrics

    experimental setup

    The experiments evaluated gpt-3.5-turbo, GPT-4, Claude-1, Claude-2, and Llama-2-70B-Chat on three factual question-answering datasets: 1,000 validation questions from TriviaQA (rc.web.nocontext), 1,000 validation questions from SciQ, and all 817 validation questions from TruthfulQA (generation). Because exact-match grading can mark semantically equivalent answers incorrect, GPT-4 judged answer equivalence to the reference for TriviaQA, while gpt-3.5-turbo did so for SciQ and TruthfulQA. The study reported ECE, temperature-scaled ECE (ECE-t), temperature-scaled Brier score (BS-t), and AUC of selective accuracy versus coverage. ECE summarizes confidence–accuracy discrepancies across confidence bins, weighted by the proportion of examples in each bin; temperature scaling fits one scalar to minimize negative log-likelihood, and BS-t is the mean squared error between scaled confidence and correctness. AUC measures how well confidence ranks examples for selective prediction, so it captures a property that ECE alone does not.

  6. Knowl 6 — Numerical and linguistic confidence are both viable, without a universal winner

    empirical result

    The experiments found that language models can verbalize confidence numerically as well as, and sometimes better than, through likelihood phrases; results vary by model and dataset rather than establishing one format as uniformly superior. For example, on gpt-3.5-turbo’s SciQ results, the best numerical ECE-t was 0.046, compared with 0.068 for the optimized linguistic method. On GPT-4’s SciQ results, human-mapped linguistic confidence reached ECE-t 0.024, lower than the best numerical result of 0.048. Linguistic confidence was expressed using phrases such as “Likely” or “Almost Certain,” while numerical confidence was a probability from 0.0 to 1.0.

  7. Knowl 7 — RLHF generally worsens Llama-2 conditional-probability calibration

    empirical result

    A comparison of Llama-70B before and after RLHF found that RLHF generally worsened the calibration of conditional probabilities on TriviaQA, SciQ, and TruthfulQA: ECE increased overall and AUC generally decreased. The paper uses this result to motivate alternatives to relying on a model’s conditional probabilities. Its broader experiments show that verbalized confidence can partly reverse this degradation, with the reversal reported as strongest on TruthfulQA, which tests difficult questions and common misconceptions.

  8. Knowl 8 — Chain-of-thought prompting did not reliably improve verbalized calibration

    empirical result

    For gpt-3.5-turbo, chain-of-thought (CoT) prompting did not noticeably or consistently improve verbalized-confidence calibration across the evaluated datasets. The study compared one-stage verbalized CoT, two-stage verbalized CoT, and a two-stage variant that supplied CoT immediately before eliciting a numerical confidence. The reported results do not support CoT as a general calibration-improvement strategy in these short-answer factual question-answering experiments.

  9. Knowl 9 — Linguistic confidence can be mapped to probabilities using human or held-out calibration

    model/method

    The linguistic-confidence method asks a model to select a likelihood expression from a fixed set and converts that expression into a probability in one of two ways. The human-mapped version uses probability values gathered in a social-media survey with 123 respondents. The optimized version estimates each phrase’s probability from its average accuracy on held-out calibration examples; if a phrase is used on fewer than 1/N1/N of the NN calibration questions, it falls back to the human-derived probability. The study used five-fold splits: four folds to estimate phrase probabilities and one to evaluate ECE and AUC. For ECE-t and BS-t, it used three folds to estimate phrase probabilities, one to fit the temperature, and one to evaluate, averaging results over 20 rotations.

  10. Knowl 10 — Scope is limited to short-form factual question answering

    limitation

    The experiments cover factual-recall questions and short-form answers, so the paper does not establish that verbalized confidence will be similarly calibrated for reasoning-heavy tasks, arithmetic, or long-form generation. The authors also note that limited technical information about closed-source RLHF models prevents identifying what model properties enable well-calibrated verbalization or explain differences among model families. Extending the evaluation to other domains and longer generations remains open.

Coverage note — The appendix’s likelihood-expression usage plots and full prompt and answer-equivalence templates are omitted because they provide supplementary diagnostics and implementation detail rather than separate main findings.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022b. Constitutional AI: Harmlessness from ai feedback.
  3. 3.Glenn W. Brier. 1950. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 78(1):1–3.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with GPT-4. ArXiv preprint arXiv:2303.12712.
  6. 6.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  7. 7.Wade Fagen-Ulmschneider. 2023. Perception of probability words. Ms., UIUC, 05-24-2023.
  8. 8.Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 4885–4894, Red Hook, NY, USA. Curran Associates Inc.
  9. 9.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR.
  10. 10.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  11. 11.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know. Arxiv arxiv:2207.05221.
  12. 12.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations.
  13. 13.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022a. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research.
  14. 14.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022b. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  15. 15.Charles Lord, Mark Lepper, and Elizabeth Preston. 1985. Considering the opposite: A corrective strategy for social judgment. Journal of personality and social psychology, 47:1231–43.
  16. 16.Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872.
  17. 17.Thomas Mussweiler, Fritz Strack, and Tim Pfeiffer. 2000. Overcoming the inevitable anchoring effect: Considering the opposite compensates for selective accessibility. Personality and Social Psychology Bulletin, 26(9):1142–1150.
  18. 18.OpenAI. 2023. Gpt-4 technical report.
  19. 19.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  20. 20.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, and Jasper Snoek. 2019. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA. Curran Associates Inc.
  21. 21.Seo Yeon Park and Cornelia Caragea. 2022. On the calibration of pre-trained language models using mixup guided by area under the margin and saliency. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5364–5374, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  23. 23.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2022. Learning to summarize from human feedback.
  24. 24.Avijit Thawani, Jay Pujara, Filip Ilievski, and Pedro Szekely. 2021. Representing numbers in NLP: a survey and a vision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 644–656, Online. Association for Computational Linguistics.
  25. 25.Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. ArXiv, abs/1707.06209.
  26. 26.Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022. Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 7273–7284, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  27. 27.Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023. Navigating the grey area: Expressions of overconfidence and uncertainty in language models.
  28. 28.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2020. Fine-tuning language models from human preferences.

Citation

MLA
Tian, K., et al. “Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 5433–42, https://doi.org/10.18653/v1/2023.emnlp-main.330.
APA
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., & Manning, C. D. (2023). Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5433–5442. https://doi.org/10.18653/v1/2023.emnlp-main.330
Chicago
Tian, K., E. Mitchell, A. Zhou, et al. 2023. “Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5433–42. https://doi.org/10.18653/v1/2023.emnlp-main.330.
Harvard
Tian, K. et al. (2023) “Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5433–5442. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.330.
Vancouver
1. Tian K, Mitchell E, Zhou A, Sharma A, Rafailov R, Yao H, Finn C, Manning CD (2023) Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5433–5442

BibTeX

@inproceedings{tian-etal-2023-just,
    title = "Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback",
    author = "Tian, Katherine  and
      Mitchell, Eric  and
      Zhou, Allan  and
      Sharma, Archit  and
      Rafailov, Rafael  and
      Yao, Huaxiu  and
      Finn, Chelsea  and
      Manning, Christopher",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.330/",
    doi = "10.18653/v1/2023.emnlp-main.330",
    pages = "5433--5442"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/