Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

Matéo MahautLaura AinaPaula CzarnowskaMomchil HardalovThomas MüllerLluís Màrquez

article2024ACL54 citations

Evaluates five major families of factual confidence estimators across multiple large language models and tasks, revealing that internal hidden-state probes achieve superior reliability while exposing how easily model confidence fluctuates across semantically equivalent prompts.

Listen

Large Language Models often generate inaccurate statements or assert false claims with high confidence, leading to hallucinations, misinformation, and degraded user trust. Accurately estimating how confident a model is in a given fact is vital for deciding when a system should answer or abstain, yet existing confidence estimators have lacked systematic comparison across tasks and model families.

The article systematically evaluates the reliability and robustness of five factual confidence estimation techniques across two complementary settings: verifying whether a statement is true and estimating whether a model knows the answer to a question. The authors developed a standardized experimental framework testing eight open-weight models, ranging from 7 billion to 46.7 billion parameters, across multiple benchmark datasets, semantic paraphrases, and multilingual translations.

The investigation produced several key findings. First, supervised probes trained directly on the model's internal hidden layers significantly outperformed all black-box and prompting-based estimators, exceeding sequence-probability baselines by an average of 0.3 in the area under the precision-recall curve. Second, while trained probes generalized well across different datasets and translated languages, non-trained methods performed poorly, frequently hovering near chance levels in question-answering settings. Third, instruction-tuned and larger models exhibited better confidence calibration when using prompting methods than their non-tuned counterparts. Finally, models demonstrated widespread instability under semantic variations: altering the wording of a question or statement frequently caused substantial fluctuations in estimated confidence, showing that models often fail to abstract facts independently of specific phrasing.

These findings highlight significant practical implications for operational risk, cost, and safety. Relying on simple prompting or output token probabilities to detect hallucinations in high-stakes environments creates a false sense of security due to poor reliability. Conversely, extracting confidence via internal probes provides superior risk mitigation but introduces architectural constraints, requiring white-box access to model weights and labeled training data.

Organizations deploying open-weight models should prioritize hidden-state probing to monitor output factuality and gate critical workflows. Where only black-box instruction-tuned models are available, verbalized confidence or surrogate token probabilities can serve as secondary alternatives, though decision-makers must account for their lower reliability. Teams should also implement consistency-checking pipelines that test multiple semantically equivalent prompts to identify unstable factual knowledge before acting on critical model responses.

Decision-makers should view these results in light of specific limitations. The empirical analysis focused primarily on atomic, single-entity facts in academic benchmarks and did not evaluate multi-step reasoning, complex document synthesis, or proprietary black-box models. Further pilot testing and analysis are necessary to establish whether these probing advantages fully generalize to complex, enterprise-scale reasoning tasks.

Cover for Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

Abstract

Large Language Models (LLMs) tend to be unreliable in the factuality of their answers. To address this problem, NLP researchers have proposed a range of techniques to estimate LLM’s confidence over facts. However, due to the lack of a systematic comparison, it is not clear how the different methods compare to one another. To fill this gap, we present a survey and empirical comparison of estimators of factual confidence. We define an experimental framework allowing for fair comparison, covering both fact-verification and question answering. Our experiments across a series of LLMs indicate that trained hidden-state probes provide the most reliable confidence estimates, albeit at the expense of requiring access to weights and training data. We also conduct a deeper assessment of factual confidence by measuring the consistency of model behavior under meaning-preserving variations in the input. We find that the confidence of LLMs is often unstable across semantically equivalent inputs, suggesting that there is much room for improvement of the stability of models’ parametric knowledge. Our code is available at https://github.com/amazon-science/factual-confidence-of-llms.

Table of Contents

  • 1 Introduction
  • 2 Factual Confidence: Key Concepts
  • 2.1 Definition of a Fact
  • 2.2 Factual Confidence
  • 2.3 Robustness of Factual Knowledge
  • 3 Factual Confidence: Survey of Methods
  • 3.1 Trained Probes
  • 3.2 Sequence Probability
  • 3.3 Verbalization
  • 3.4 Surrogate Token Probability
  • 3.5 Output Consistency
  • 4 Methodology
  • 4.1 Data
  • 4.1.1 P ( T ) in Fact Verification: Lama T-REx
  • 4.1.2 P ( IK ) in QA: PopQA
  • 4.2 Scoring Methods Implementation
  • 4.2.1 Estimating P ( T )
  • 4.3 Estimating P ( IK )
  • 4.4 Evaluating Scoring Methods
  • 4.5 Models
  • 4.6 Paraphrasing and Translation
  • 5 Empirical Comparison of the Methods
  • 5.1 P ( IK ) on Lama T-REx
  • 5.2 P ( IK ) on PopQA
  • 5.3 Generalization of the Trained Probe
  • 6 Robustness to Linguistic Variations
  • 6.1 Robustness of Methods
  • 6.2 Robustness of Facts Encoding in LLMs
  • 7 Discussion & Conclusion
  • Acknowledgments
  • Limitations
  • Ethics and Broader Impact
  • References
  • Appendix for 'Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators'
  • A Implementation of scoring methods
  • A.1 Verbalized
  • A.2 Surrogate token probabilities
  • A.2.1 Consistency
  • B Paraphrasing
  • C Method Robustness to Variation
  • D Analysis of specific Precision and Recall

Knowls

  1. Knowl 1 — Trained hidden-state probes are the most reliable estimator in the tested settings

    empirical result

    Across eight publicly available Falcon, Mistral, and Mixtral models, trained probes were the strongest factual-confidence estimators in both statement verification (P(T)P(T)) and question answering (P(IK)P(IK)). They outperformed the other methods on the T-REx fact-verification task and were by far the most reliable method on PopQA question answering. The probes also retained useful discriminative performance when applied to data or languages not used for training. This advantage comes with stricter requirements than the other estimators: access to model weights and labeled training examples.

  2. Knowl 2 — Factual confidence has statement-truth and answerability formulations

    definition

    The study treats factual confidence as a model’s assessed likelihood of being correct about an atomic fact and distinguishes two input-dependent quantities. P(T)P(T) is the model’s confidence that a fact asserted by an input statement is true; it is evaluated by presenting the statement itself. P(IK)P(IK) is the model’s confidence that it will provide the correct answer to a factual question; it is evaluated from the question, without supplying the fact as an assertion. These are conceptual confidence scores and are not assumed to be calibrated probabilities. The appropriate formulation depends on whether the application is fact verification or question answering.

  3. Knowl 3 — Five estimator families use different signals for factual confidence

    model/method

    The comparison organizes factual-confidence estimators into five families. A trained probe maps a model’s hidden representations to a confidence score. Sequence probability uses the average log-probability of a token sequence: the statement for P(T)P(T) and, in this study, the question for P(IK)P(IK). Verbalization asks the model to report a numeric confidence level, which is filtered and normalized. Surrogate-token probability uses the log-probability of a prompted “Yes” response to a truth or answerability question. Output consistency samples multiple answers to a question and scores their agreement using natural-language-inference (NLI) scores; this method estimates P(IK)P(IK) only, because it requires generated answers. Unlike the other methods, verbalization produces a confidence value in generated text rather than deriving it from hidden states or token probabilities.

  4. Knowl 4 — Benchmark and scoring protocol for comparing the estimators

    experimental setup

    The experiments evaluated eight open-weight models: Falcon 40B and 7B, their instruction-tuned versions, Mixtral 46.7B and its instruction-tuned version, and Mistral 7B and its instruction-tuned version. For P(T)P(T), the study used Lama T-REx: each of 34K true subject–relation–object facts was paired with a false statement formed by replacing the object with another entity from the same relation. The balanced data were split into 54K training statements and 13.6K analysis statements; only trained probes used the training split. For P(IK)P(IK), PopQA’s roughly 14K questions were split into about 11K training and 2.8K test questions. A PopQA question was labeled positive for a model if its greedy-decoded answer was correct, using the dataset’s synonymous answer phrases for evaluation; positive-label rates varied by model from about 11% to 27%. A three-layer feed-forward probe was trained for 10 epochs on hidden states from transformer layer 24. The consistency estimator used 10 responses sampled at temperature 1 and averaged pairwise NLI scores. Estimators were compared by AUPRC, with higher values indicating better ranking of true versus false statements or answerable versus incorrectly answered questions.

  5. Knowl 5 — Probe performance on T-REx exceeds the competing estimators

    empirical result

    On balanced T-REx true/false statements, the trained probe had the best P(T)P(T) AUPRC across the tested models. Its average AUPRC advantage over average sequence probability was about 0.3. Verbalized confidence was competitive with the probe only for Mistral 7B Instruct; for the other model-method comparisons, the alternatives were at least 0.1 AUPRC below the trained probe. Sequence probability was generally above chance but weaker overall. Verbalization and surrogate-token probability varied more across models than the probe and sequence probability: instruction-tuned models consistently did better than their non-instruction-tuned counterparts, and 40B-or-larger models generally did better than 7B models, with Mistral 7B Instruct as an exception.

  6. Knowl 6 — Question-answering confidence is harder to estimate than statement truth

    empirical result

    On PopQA, where a positive P(IK)P(IK) label means that a model’s greedy-decoded answer is correct, the trained probe was again the most reliable estimator across the tested models. Overall P(IK)P(IK) estimation was harder than T-REx P(T)P(T) estimation; for example, the best probe’s AUPRC was about 0.1 lower on PopQA. Except for Falcon 40B Instruct, the non-probe methods were close to or below the model-specific chance level, which ranged from 0.11 to 0.27 because answer-correctness rates differed across models. Performance also varied substantially among models, with differences within a method reaching about 40%. These results indicate that the non-trained estimators tested here were generally unreliable for predicting whether a model would answer PopQA questions correctly.

  7. Knowl 7 — T-REx-trained probes retain substantial performance on PopQA statements

    empirical result

    The study tested whether probes trained to estimate P(T)P(T) on T-REx transfer to PopQA by turning PopQA question–answer pairs into balanced true and false statements using the template “The answer to [question] is [answer].” The probe’s AUPRC decreased for every model relative to its in-domain T-REx result, but remained between 0.59 and 0.81. Values below are PopQA AUPRC followed by the difference from in-domain T-REx AUPRC:

    Falcon 40B: 0.80, −0.16-0.16; Falcon 40B Instruct: 0.81, −0.15-0.15; Falcon 7B: 0.66, −0.25-0.25; Falcon 7B Instruct: 0.59, −0.28-0.28; Mixtral 46.7B: 0.78, −0.18-0.18; Mistral 7B: 0.62, −0.31-0.31; Mistral 7B Instruct: 0.75, −0.18-0.18.

    The authors interpret the remaining AUPRC range as evidence of substantial, though imperfect, out-of-domain generalization in this atomic-fact setup.

  8. Knowl 8 — Estimator reliability is broadly robust to paraphrases, with uneven language transfer

    empirical result

    For paraphrase robustness, the study generated multiple meaning-preserving variants of each input and evaluated the existing estimators without retraining probes or changing prompts. Across ten sampled paraphrase sets, AUPRC was broadly stable for all methods; the largest reported standard deviation was about 3 percentage points, for the trained probe on P(IK)P(IK) with Mistral 7B. Translation transfer was less uniform. On French and Polish T-REx statements, all methods were above chance except verbalized confidence and surrogate-token scores on Mistral models. English-trained probes achieved AUPRCs of 0.73–0.91 on French and 0.61–0.91 on Polish, with 40B-or-larger models and instruction-tuned Mistral showing the strongest transfer. Thus, estimator ranking performance can transfer to meaning-preserving reformulations, but language transfer depends on the estimator and model.

  9. Knowl 9 — A model’s confidence in a fact can change across equivalent inputs

    empirical result

    The study assessed the stability of trained-probe P(T)P(T) scores across roughly eight paraphrases per T-REx fact. Some facts received almost unchanged scores across formulations, but many had substantial variation: standard deviations reached about 0.5, with a small number of cases near 0.7. Falcon 7B Instruct showed the least stable score distribution among the tested models, while Mistral 7B showed the least variation. For translated versions of the same facts, confidence scores were correlated across languages for the larger models, but a Friedman test found statistically significant distributional differences across languages for every model. The results therefore show both that model confidence can depend on wording and that cross-language score relationships do not amount to fully invariant confidence.

  10. Knowl 10 — Scope and practical constraints of the findings

    limitation

    The experiments cover eight models, five estimator families, and two datasets focused on simple atomic facts; they do not establish that the same rankings hold for non-atomic facts, reasoning, or in-context learning. Trained probes require model-weight access and supervised labels, and the paper notes that probe transfer has limits beyond the tested setup. Sequence probabilities may also perform worse on longer or more complex inputs, while prompt-based estimators are sensitive to prompt variation. For practical use, the study favors trained probes when weights and labeled data are available; when those are unavailable, it suggests verbalized confidence or surrogate-token probabilities for instruction-tuned models, while cautioning that the other tested methods are not consistently reliable, particularly on non-instruction-tuned models.

Coverage note — The literature survey and detailed appendix precision-at-recall tables are omitted; the survey is background, and the appendix metrics corroborate rather than materially alter the main AUPRC comparisons.

References

  1. 1.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al. 2023. The Falcon series of open language models. arXiv preprint arXiv:2311.16867.
  2. 2.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations, ICLR ’24, Vienna, Austria.
  3. 3.Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, Findings ’23, pages 967–976, Singapore.
  4. 4.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova Dassarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv, abs/2204.05862.
  5. 5.Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961.
  6. 6.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623.
  7. 7.Emily M. Bender and Alexander Koller. 2020. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL ’20, pages 5185–5198, Online.
  8. 8.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2023. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations, ICLR ’23, Kigali, Rwanda.
  9. 9.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations, ICLR ’23, Kigali, Rwanda.
  10. 10.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012–1031.
  11. 11.Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders. 2021. Truthful AI: Developing and governing ai that does not lie. arXiv preprint arXiv:2110.06674.
  12. 12.Marina Fomicheva, Shuo Sun, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, and Lucia Specia. 2020. Unsupervised quality estimation for neural machine translation. Transactions of the Association for Computational Linguistics, 8:539–555.
  13. 13.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33nd International Conference on Machine Learning, volume 48 of ICML ’16, pages 1050–1059, New York City, NY, USA.
  14. 14.Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2023. A survey of language model confidence estimation and calibration. arXiv preprint arXiv:2311.08298.
  15. 15.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of ICML ’17’, pages 1321–1330, Sydney, Australia.
  16. 16.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  17. 17.Momchil Hardalov, Arnav Arora, Preslav Nakov, and Isabelle Augenstein. 2022. A survey on stance detection for mis- and disinformation identification. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1259–1277, Seattle, United States. Association for Computational Linguistics.
  18. 18.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  19. 19.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  20. 20.Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825.
  21. 21.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? On the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  22. 22.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zachary Dodds, Nova Dassarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, John Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom B. Brown, Jack Clark, Nicholas Joseph, Benjamin Mann, Sam McCandlish, Christopher Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  23. 23.Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021. Multilingual LAMA: Investigating knowledge in multilingual pretrained language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, EACL ’21, pages 3250–3258, Online.
  24. 24.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, ICLR ’23, Kigali, Rwanda.
  25. 25.Moritz Laurer, Wouter van Atteveldt, Andreu Casas, and Kasper Welbers. 2024. Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI. Political Analysis, 32(1):84–100.
  26. 26.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022a. Teaching models to express their uncertainty in words. Transactions on Machine Learning Research.
  27. 27.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022b. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, ACL ’22, pages 3214–3252, Dublin, Ireland.
  28. 28.Linhao Luo, Trang Vu, Dinh Phung, and Reza Haf. 2023. Systematic assessment of factual knowledge in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Findings ’23, pages 13272–13286, Singapore.
  29. 29.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards deep learning models resistant to adversarial attacks. In 6th International Conference on Learning Representations, ICLR ’18, Vancouver, Canada.
  30. 30.Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2024. Dissociating language and thought in large language models. Trends in Cognitive Sciences, pages 1364–6613.
  31. 31.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL ’23, pages 9802–9822, Toronto, Canada.
  32. 32.Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP ’23, pages 9004–9017, Singapore.
  33. 33.Melanie Mitchell and David C. Krakauer. 2023. The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13):e2215907120.
  34. 34.Xenia Ohmer, Elia Bruni, and Dieuwke Hupkes. 2023. Separating form and meaning: Using self-consistency to quantify task understanding across multiple senses. In Proceedings of the Third Workshop on Natural Language Generation, Evaluation, and Metrics, GEM ’23, pages 258–276, Singapore.
  35. 35.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  36. 36.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35 of NeurIPS ’22, pages 27730–27744, New Orleans, Louisiana, USA and Online.
  37. 37.Lorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz, Alexa Yue Pan, Yarin Gal, Owain Evans, and Jan M. Brauner. 2024. How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions. In The Twelfth International Conference on Learning Representations, ICLR ’24, Vienna, Austria.
  38. 38.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP ’19, pages 2463–2473, Hong Kong, China.
  39. 39.Jirui Qi, Raquel Fernández, and Arianna Bisazza. 2023. Cross-lingual consistency of factual knowledge in multilingual language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP ’23, pages 10650–10666, Singapore.
  40. 40.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In 5th International Conference on Learning Representations, ICLR ’17.
  41. 41.Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli. 2021. Understanding the capabilities, limitations, and societal impact of large language models. ArXiv, abs/2102.02503.
  42. 42.James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. 2018. The fact extraction and VERification (FEVER) shared task. In Proceedings of the First Workshop on Fact Extraction and VERification, FEVER ’18, pages 1–9, Brussels, Belgium.
  43. 43.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP ’23, pages 5433–5442, Singapore.
  44. 44.SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:2401.01313.
  45. 45.Elena Voita, Rico Sennrich, and Ivan Titov. 2019. The bottom-up evolution of representations in the Transformer: A study with machine translation and language modeling objectives. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP ’19, pages 4396–4406, Hong Kong, China.
  46. 46.Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, et al. 2023a. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity. arXiv preprint arXiv:2310.07521.
  47. 47.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023b. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR ’23, Kigali, Rwanda.
  48. 48.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
  49. 49.Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In The Twelfth International Conference on Learning Representations, ICLR ’24, Vienna, Austria.
  50. 50.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don’t know? In Findings of the Association for Computational Linguistics: ACL 2023, Findings ’23, pages 8653–8665, Toronto, Canada.

Citation

MLA
Mahaut, M., et al. “Factual Confidence of LLMs: On Reliability and Robustness of Current Estimators”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 4554–70, https://doi.org/10.18653/v1/2024.acl-long.250.
APA
Mahaut, M., Aina, L., Czarnowska, P., Hardalov, M., Müller, T., & Marquez, L. (2024). Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4554–4570. https://doi.org/10.18653/v1/2024.acl-long.250
Chicago
Mahaut, M., L. Aina, P. Czarnowska, M. Hardalov, T. Müller, and L. Marquez. 2024. “Factual Confidence of LLMs: On Reliability and Robustness of Current Estimators”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4554–70. https://doi.org/10.18653/v1/2024.acl-long.250.
Harvard
Mahaut, M. et al. (2024) “Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4554–4570. Available at: https://doi.org/10.18653/v1/2024.acl-long.250.
Vancouver
1. Mahaut M, Aina L, Czarnowska P, Hardalov M, Müller T, Marquez L (2024) Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4554–4570

BibTeX

@inproceedings{Mahaut_2024, title={Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators}, url={http://dx.doi.org/10.18653/v1/2024.acl-long.250}, DOI={10.18653/v1/2024.acl-long.250}, booktitle={Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Mahaut, Matéo and Aina, Laura and Czarnowska, Paula and Hardalov, Momchil and Müller, Thomas and Marquez, Lluis}, year={2024}, pages={4554–4570} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/