Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Kaitlyn ZhouJena D. HwangXiang RenMaarten Sap

article2024ACL166 citations

Reveals that language models rarely express uncertainty and suffer from a 47% error rate on confident statements, tracing this miscalibration to human preference biases during post-training alignment and demonstrating the resulting risks of human overreliance.

Listen

As natural language becomes the primary interface for human-artificial intelligence interaction, users increasingly rely on language models for critical decision-making and information retrieval. In human communication, speakers use linguistic cues called epistemic markers—specifically strengtheners that express confidence and weakeners that convey doubt—to signal reliability. However, when artificial intelligence systems fail to accurately communicate their limitations, users face significant risks of acting on incorrect advice. The article investigates how widely deployed language models express certainty, how human users interpret these linguistic markers, and where model overconfidence originates.

The article evaluates model calibration across major commercial systems and conducts behavioral experiments to quantify human reliance on model responses. To establish these patterns, the authors prompted nine prominent models, including variants of GPT, Claude, and LLaMA-2, across more than 125,000 queries using multiple-choice benchmark questions. They then conducted interactive user studies using challenging trivia to assess how participants rely on model advice under calibrated, overconfident, and underconfident conditions. Finally, the authors analyzed reinforcement learning datasets and reward models to identify the training mechanisms driving model behavior.

The findings reveal that language models are inherently reluctant to express uncertainty and, when prompted to do so, exhibit severe overconfidence. Across baseline tests, only 5% of model generations included uncertainty markers; when prompted to indicate confidence, models favored strengtheners over weakeners, resulting in an average error rate of 47% among confident answers. In human experiments, participants relied on plain statements without markers nearly 90% of the time, treating silence on uncertainty as implicit confidence. While users readily adapted to calibrated or underconfident models, exposure to overconfident models caused persistent judgment errors, leading participants to rely on incorrect answers 73% of the time during miscalibrated periods and degrading their accuracy even after the model reverted to calibrated behavior. Further analysis showed that this overconfidence originates during reinforcement learning from human feedback, as human evaluators and reward models systematically penalize expressions of uncertainty, preferring plain or assertive responses.

These results demonstrate that current alignment methods inadvertently incentivize artificial intelligence systems to conceal uncertainty, creating acute safety, compliance, and performance risks. Because users interpret plain statements as authoritative, uncalibrated language models can easily mislead human decision-makers. Moreover, early exposure to overconfident system behavior causes lasting damage to user judgment and trust, potentially driving algorithmic aversion. To address these vulnerabilities, practitioners should train models to produce unsolicited expressions of uncertainty when confidence is low, incorporate diverse spoken-language sources to expand hedging capabilities, and establish context-dependent certainty thresholds for high-stakes deployments.

The article notes several limitations, including a focus entirely on English-language models and user studies restricted to United States participants, which may not capture cross-cultural differences in how uncertainty is interpreted. Additionally, the experimental trivia tasks may not fully mirror the dynamics of high-stakes, real-world deployments. Nevertheless, the evidence strongly indicates that human feedback mechanisms drive systemic model overconfidence, warranting caution among organizations deploying conversational artificial intelligence in decision-support roles.

arXiv: 2401.06730
Cover for Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Abstract

As natural language becomes the default interface for human-AI interaction, there is a need for LMs to appropriately communicate uncertainties in downstream applications. In this work, we investigate how LMs incorporate confidence in responses via natural language and how downstream users behave in response to LM-articulated uncertainties. We examine publicly deployed models and find that LMs are reluctant to express uncertainties when answering questions even when they produce incorrect responses. LMs can be explicitly prompted to express confidences, but tend to be overconfident, resulting in high error rates (an average of 47%) among confident responses. We test the risks of LM overconfidence by conducting human experiments and show that users rely heavily on LM generations, whether or not they are marked by certainty. Lastly, we investigate the preference-annotated datasets used in post training alignment and find that humans are biased against texts with uncertainty. Our work highlights new safety harms facing human-LM interactions and proposes design recommendations and mitigating strategies moving forward.

Table of Contents

  • 1 Introduction
  • 2 Epistemic Markers in Language Models
  • 3 How do LMs use Epistemic Markers?
  • 3.1 Methods
  • 3.2 Findings
  • 3.3 Discussion
  • 4 Human Interpretations of Uncertainty
  • 4.1 Methods
  • 4.2 Experimental Settings
  • 4.3 Findings
  • 4.4 Discussion
  • 5 Origin of Model Overconfidence
  • 5.1 Where Does LM Overconfidence Begin?
  • 5.2 Methods
  • 5.3 Findings
  • 5.4 Discussion
  • 6 Desired Criteria and Mitigating Solutions Moving Forward
  • Criteria #3: Context-Dependent Calibration
  • 7 Conclusion
  • 8 Limitations
  • The Ambiguity of Weakened Strengtheners
  • 9 Ethics Statement
  • Acknowledgements
  • References
  • A Details on Prompt Paraphrases
  • B Details on Experiments from Section 4
  • C Insensitivity to Increases in Temperature
  • D Reward Model Compared to Human Scoring of Expressions of Uncertainty
  • E Licenses for Scientific Artifacts

Knowls

  1. Knowl 1 — Public language models rarely volunteer epistemic markers

    empirical result

    The study treated strengtheners as expressions of certainty, weakeners as expressions of uncertainty, and answers without either as plain statements. It prompted nine publicly deployed models—text-davinci-003, GPT-3.5-Turbo, GPT-4, LLaMA-2 7B/13B/70B, Claude-1, Claude-2, and Claude-Instant-1—with 49 prompt templates on 284 four-choice MMLU questions, yielding 125,244 queries. The prompts used a base template, chain-of-thought (CoT) instructions, explicit epistemic-marker instructions (Epi-M), or both (Epi-M+CoT); decoding was zero-shot at temperature 0.3 with a 400-token limit.

    Across the prompt groups, the fraction of responses containing any marker was 6% for the base prompt, 16% for CoT, 71% for Epi-M, and 57% for Epi-M+CoT. Strengtheners occurred in 0%, 3%, 24%, and 24% of responses, respectively; weakeners occurred in 2%, 3%, 15%, and 20%. Across the 48 non-base templates, the reported averages were 57% for any marker, 20% for strengtheners, and 14% for weakeners. Thus explicit prompting substantially increased marker use, but ordinary prompting left nearly all answers unmarked. The authors identified 76 strengthener forms and 105 weakener forms through iterative qualitative coding and regex detection; their final coding procedure exceeded 90% accuracy.

  2. Knowl 2 — Expressions of certainty often accompany wrong answers

    empirical result

    Across the language-model generations, only 53% of answers accompanied by a strengthener were correct, compared with 32% of answers accompanied by a weakener; chance accuracy on the four-option questions was 25%. Put another way, 47% of strengthener-marked answers were wrong. Six of the nine tested models generated strengtheners more often than weakeners, and the across-model/template averages were 20% of responses with strengtheners versus 14% with weakeners. Because strengtheners were common, 17% of all incorrect answers included one. These findings show that verbal certainty did not reliably identify correct answers in this evaluation.

  3. Knowl 3 — Users treat plain answers much like confident answers

    empirical result

    In a control experiment, 25 participants each saw 106 difficult country-capital questions and the beginning of a purported AI answer from an agent named Marvin, but not the answer itself. Participants chose whether to rely on Marvin or look up the answer. Nearly 90% chose to rely when the answer began with a strengthener, while approximately 90% chose to look it up when it began with a weakener. Plain openings such as “The answer is” or an answer alone were also relied on nearly 90% of the time. In this task, the absence of an uncertainty marker therefore functioned for users much like an expression of certainty.

  4. Knowl 4 — Users can learn correctly calibrated verbal cues

    empirical result

    In the calibrated interactive condition, 25 participants completed 50 question-answering rounds with Marvin and received feedback after each decision. Strengtheners accompanied correct answers and weakeners accompanied incorrect ones. Participants earned one point for relying on a correct answer, lost one point for relying on an incorrect answer, and received zero for looking up the answer; half the answers were wrong. Early in the task, participants relied on strengthener-marked answers 94% of the time and on weakener-marked answers 7% of the time. After about 20 rounds, most participants relied on strengtheners 99% of the time and on weakeners about 1% of the time. Their average accuracy across rounds was 97%, showing that participants could use the cues effectively when the system applied them consistently.

  5. Knowl 5 — Early overconfidence harms performance beyond the miscalibrated rounds

    empirical result

    In the overconfident interactive condition, 25 participants played 50 rounds: five incorrect answers in the first 20 rounds were marked with strengtheners, and the remaining 30 rounds were calibrated. During the initial rounds, participants relied on strengtheners 81% of the time even though only 66% of strengthener-marked answers were correct; they mistakenly relied on incorrect confident answers an average of 73% of the time. Accuracy averaged 76% during the miscalibrated rounds and remained at 86% during the later calibrated rounds. Participants also became more likely to rely on weakeners, even though weakeners were correctly associated with incorrect answers in this condition: reliance rose to 9%, compared with 3% in the control experiment. The results indicate that initial misuse of strengtheners affected both later performance and participants’ interpretation of other cues.

  6. Knowl 6 — Underconfidence is less persistent than overconfidence in the user task

    empirical result

    In the underconfident interactive condition, 25 participants completed 50 rounds with feedback. Five correct answers in the first 20 rounds were marked with weakeners; the remaining 30 rounds were calibrated. Participants averaged 66% accuracy during the underconfident rounds and 98% during the later calibrated rounds. Unlike the participants exposed to early overconfidence, those exposed to early underconfidence were able to perform at nearly the calibrated-condition level once the cues became reliable.

  7. Knowl 7 — Model comparisons associate the strengthener preference with RLHF

    empirical result

    To investigate where the preference for certainty arises, the authors compared the GPT-3 family’s base davinci model, supervised-fine-tuned text-davinci-002, and RLHF-trained text-davinci-003, as well as base LLaMA-2 models and LLaMA-2 Chat models trained with supervised fine-tuning and RLHF. Base and instruction-tuned variants showed a relative preference for weakeners, whereas the RLHF variants emitted more strengtheners than weakeners. The authors interpret this reversal as evidence that the strengthener preference is introduced during RLHF, while the model-stage comparison itself establishes an association rather than isolating RLHF experimentally. Raising generation temperature to its maximum did not eliminate the preference for strengtheners over weakeners.

  8. Knowl 8 — The tested reward model penalizes uncertainty language

    empirical result

    The authors evaluated OpenAssistant’s reward model on 183 question-response pairs drawn from 30 common epistemic-marker templates. The model assigned plain statements an average score of 4.03, strengthener-marked statements 0.82, and weakener-marked statements −1.86. The particularly low scores for weakeners show that this reward model favored answers without uncertainty language over answers that explicitly expressed uncertainty, a preference that could shape generations when the model is used for alignment.

  9. Knowl 9 — Human preference data favors plain text over uncertainty language

    empirical result

    The authors examined human preference annotations in OpenAI’s WebGPT comparison and Summarize with Feedback datasets, a Synthetic Instruct GPT Pairwise dataset, and Anthropic’s Helpful and Harmless dataset. Strengtheners appeared slightly more often in rejected than chosen texts (2.95% versus 2.72%); in pairwise comparisons, plain text was chosen 9% more often than strengthened text. Weakeners likewise appeared more often in rejected than chosen texts (5.02% versus 4.47%). In pairwise comparisons, weakeners were preferred 8% less often than strengtheners and 9% less often than plain text. The authors therefore identify a bias against uncertainty language, rather than a general preference for certainty-marked text, in these preference annotations.

  10. Knowl 10 — Proposed design criteria for safer linguistic calibration

    model/method

    The authors propose three criteria for language models that communicate uncertainty. First, models should emit epistemic markers without being prompted, because users may interpret unmarked answers as confident; possible interventions include adding representative markers to alignment data, training or priming annotators to notice bias against uncertainty, and auditing data with marker-detection heuristics. Second, models should cover a broad range of epistemic expressions, which could be supported by training on more varied sources, including speech, podcasts, news transcripts, and peer conversations. Third, calibration should depend on the deployment context: practitioners could first measure how users in a particular setting interpret and rely on markers, then adjust the model with prompting or fine-tuning. These are proposed design directions, not interventions evaluated in the study.

Coverage note — The paper’s detailed inventories of individual marker phrases are omitted because they catalogue examples rather than add a separate load-bearing result. Its limitations—English-only model analysis, U.S.-based participants, the gap between the simulated task and real deployments, and ambiguity in intermediate-strength markers—are not separate knowls because they primarily qualify the generalizability and interpretation of the reported findings.

References

  1. 1.Karin Aijmer. 1986. Discourse variation and hedging. In Corpus linguistics II, pages 1–18. Brill.
  2. 2.Karin Aijmer. 2013. Understanding pragmatic markers: A variational pragmatic approach. Edinburgh University Press.
  3. 3.Fahad Al-Rashady. 2012. Determining the role of hedging devices in the political discourse of two american prezidentiables in 2008. TESOL Journal, 7(1):30–42.
  4. 4.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. PaLM 2 technical report. arXiv preprint arXiv:2305.10403.
  5. 5.Anthropic. 2022. Introducing claude.
  6. 6.Mohammad Atari, Mona J. Xue, Peter S. Park, Damián E. Blasi, and Joseph Henrich. 2023. Which humans?
  7. 7.Carl F. Auerbach and Louise Bordeaux Silverstein. 2003. Qualitative data: An introduction to coding and analysis.
  8. 8.Austin S Babrow, Chris R Kasch, and Leigh A Ford. 1998. The many meanings of uncertainty in illness: Toward a systematic accounting. Health communication, 10(1):1–23.
  9. 9.Gagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz, and Daniel S Weld. 2021a. Is the Most Accurate AI the Best Teammate? Optimizing AI for Teamwork. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 11405–11414.
  10. 10.Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz. 2019a. Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance. In Proceedings of the AAAI conference on human computation and crowdsourcing, volume 7, pages 2–11.
  11. 11.Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S Weld, Walter S Lasecki, and Eric Horvitz. 2019b. Updates in Human-AI Teams: Understanding and Addressing the Performance/Compatibility Tradeoff. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 2429–2437.
  12. 12.Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021b. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY, USA. Association for Computing Machinery.
  13. 13.Emily Bender. 2019. The#benderrule: On naming the languages we study and why it matters. The Gradient, 14.
  14. 14.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  15. 15.Dale E Brashers, Judith L Neidig, Stephen M Haas, Linda K Dobbs, Linda W Cardillo, and Jane A Russell. 2000. Communication in the management of uncertainty: The case of persons living with hiv or aids. Communications Monographs, 67(1):63–84.
  16. 16.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-Shot Learners. Advances in neural information processing systems, 33:1877–1901.
  17. 17.Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proc. ACM Hum.-Comput. Interact., 5(CSCW1).
  18. 18.David V Budescu, Shalva Weinberg, and Thomas S Wallsten. 1988. Decisions based on numerically and verbally expressed uncertainties. Journal of Experimental Psychology: Human Perception and Performance, 14(2):281.
  19. 19.Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The role of explanations on trust and reliance in clinical decision support systems. In 2015 International Conference on Healthcare Informatics, pages 160–169.
  20. 20.Carrie J. Cai, Emily Reif, Narayan Hegde, Jason D. Hipp, Been Kim, Daniel Smilkov, Martin Wattenberg, Fernanda B. Viégas, Gregory S. Corrado, Martin C. Stumpe, and Michael Terry. 2019. Human-Centered Tools for Coping with Imperfect Algorithms During Medical Decision-Making. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems.
  21. 21.Kathy Charmaz. 2006. Constructing grounded theory: A practical guide through qualitative analysis. SAGE.
  22. 22.Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. 2023. A close look into the calibration of pre-trained language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1343–1367, Toronto, Canada. Association for Computational Linguistics.
  23. 23.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  24. 24.Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova. 2020. A case for humans-in-the-loop: Decisions in the presence of erroneous algorithmic scores. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, pages 1–12.
  25. 25.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  26. 26.Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan. 2023. A diachronic perspective on user trust in AI under uncertainty. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5567–5580, Singapore. Association for Computational Linguistics.
  27. 27.Berkeley J Dietvorst, Joseph P Simmons, and Cade Massey. 2015. Algorithm Aversion: People Erroneously Avoid Algorithms after Seeing Them Err. Journal of Experimental Psychology: General, 144(1):114.
  28. 28.Zibin Dong, Yifu Yuan, Jianye HAO, Fei Ni, Yao Mu, YAN ZHENG, Yujing Hu, Tangjie Lv, Changjie Fan, and Zhipeng Hu. 2024. Aligndiff: Aligning diverse human preferences via behavior-customisable diffusion model. In The Twelfth International Conference on Learning Representations.
  29. 29.Marek J Druzdzel. 1989. Verbal uncertainty expressions: Literature review. Pittsburgh, PA: Carnegie Mellon University, Department of Engineering and Public Policy, pages 1–13.
  30. 30.Upol Ehsan, Q. Vera Liao, Michael Muller, Mark O. Riedl, and Justin D. Weisz. 2021. Expanding explainability: Towards social transparency in ai systems. CHI ’21, New York, NY, USA. Association for Computing Machinery.
  31. 31.Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamile Lukoši ˙ ut¯ e, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459.
  32. 32.Noah D. Goodman and Michael C. Frank. 2016. Pragmatic Language Interpretation as Probabilistic Inference. Trends in Cognitive Sciences, 20(11):818–829.
  33. 33.Jonathan Gordon and Benjamin Van Durme. 2013. Reporting bias and knowledge acquisition. In Proceedings of the 2013 workshop on Automated knowledge base construction, pages 25–30.
  34. 34.Herbert P Grice. 1975. Logic and conversation. In Speech acts, pages 41–58. Brill.
  35. 35.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  36. 36.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
  37. 37.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019. Using pre-training can improve model robustness and uncertainty. In International conference on machine learning, pages 2712–2721. PMLR.
  38. 38.Joseph Henrich, Steven J. Heine, and Ara Norenzayan. 2010. Most people are not WEIRD. Nature, 466:29–29.
  39. 39.Jim Hollan and Scott Stornetta. 1992. Beyond being there. In Proceedings of the SIGCHI conference on Human factors in computing systems, pages 119–125.
  40. 40.Ken Hyland. 2005. Stance and engagement: A model of interaction in academic discourse. Discourse studies, 7(2):173–192.
  41. 41.Ken Hyland. 2014. Disciplinary discourses: Writer stance in research articles. In Writing: Texts, processes and practices, pages 99–121. Routledge.
  42. 42.Reiko Itani. 1995. Semantics and pragmatics of hedges in English and Japanese. University of London, University College London (United Kingdom).
  43. 43.Maia Jacobs, Melanie F Pradier, Thomas H McCoy Jr, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos. 2021. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry, 11(1):108.
  44. 44.Abhyuday Jagannatha and Hong Yu. 2020. Calibrating structured output predictors for natural language processing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2078–2092, Online. Association for Computational Linguistics.
  45. 45.Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. ArXiv, abs/2307.04657.
  46. 46.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Delong Chen, Wenliang Dai, Andrea Madotto, and Pascale Fung. 2022. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55:1 – 38.
  47. 47.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  48. 48.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  49. 49.Amita Kamath, Robin Jia, and Percy Liang. 2020. Selective question answering under domain shift. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5684–5696, Online. Association for Computational Linguistics.
  50. 50.Pranav Khadpe, Ranjay Krishna, Li Fei-Fei, Jeffrey T. Hancock, and Michael S. Bernstein. 2020. Conceptual metaphors impact perceptions of human-ai collaboration. Proc. ACM Hum.-Comput. Interact., 4(CSCW2).
  51. 51.René F. Kizilcec. 2016. How much information?: Effects of transparency on trust in an algorithmic interface. Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems.
  52. 52.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in- and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1326–1340, Online. Association for Computational Linguistics.
  53. 53.George Lakoff. 1975. Hedges: A study in meaning criteria and the logic of fuzzy concepts. In Contemporary Research in Philosophical Logic and Linguistic Semantics: Proceedings of a Conference Held at the University of Western Ontario, London, Canada, pages 221–271. Springer.
  54. 54.Shizuka Lauwereyns. 2002. Hedges in Japanese conversation: The influence of age, sex, and formality. Language Variation and Change, 14(2):239–259.
  55. 55.Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. 2022. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746.
  56. 56.Stephanie C. Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. Trans. Mach. Learn. Res., 2022.
  57. 57.Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023. Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4791–4797, Singapore. Association for Computational Linguistics.
  58. 58.Li Lucy and Jon Gauthier. 2017. Are distributional representations ready for the real world? evaluating word vectors for grounded perceptual meaning. In Proceedings of the First Workshop on Language Grounding for Robotics, pages 76–85, Vancouver, Canada. Association for Computational Linguistics.
  59. 59.Enas Mansour and Sharif Alghazo. 2021. Hedging in political discourse: The case of trump’s speeches. Jordan Journal of Modern Languages and Literatures.
  60. 60.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  61. 61.Kevin R McKee, Xuechunzi Bai, and Susan T Fiske. 2024. Warmth and competence in human-agent cooperation. Autonomous Agents and Multi-Agent Systems, 38(1):23.
  62. 62.Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2020. Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics, 10:857–872.
  63. 63.Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence, 267:1–38.
  64. 64.Pilar Mur-Dueñas. 2021. There may be differences: Analysing the use of hedges in english and spanish research articles. Lingua, 260:103131.
  65. 65.Muziatun Muziatun, Fahria Malabar, and Nur Sangketa. 2021. An analysis of hedging devices on students’ presentation of seminar on language based on the gender. Ideas: Jurnal Pendidikan, Sosial, dan Budaya, 7:97.
  66. 66.Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using bayesian binning. Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence, 2015:2901–2907.
  67. 67.Thu Nguyen Thi Thuy. 2018. A corpus-based study on cross-cultural divergence in the use of hedges in academic research articles written by vietnamese and native english-speaking authors. Social Sciences, 7(4).
  68. 68.Tri Nuraniwati and Alfelia Nugky Permatasari. 2021. Hedging in ted talks: A corpus-based pragmatic study. JEELS (Journal of English Education and Linguistics Studies), 8(2):203–226.
  69. 69.OpenAI. 2023. GPT-4 technical report.
  70. 70.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155.
  71. 71.Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna M. Wallach. 2018. Manipulating and measuring model interpretability. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems.
  72. 72.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  73. 73.Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076.
  74. 74.José Sanders and Wilbert Spooren. 1996. Subjectivity and certainty in epistemic modality: A study of dutch epistemic modifiers.
  75. 75.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 29971–30004. PMLR.
  76. 76.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  77. 77.Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022a. Neural theory-of-mind? on the limits of social intelligence in large LMs. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3762–3780, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  78. 78.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022b. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United States. Association for Computational Linguistics.
  79. 79.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Annasaheb Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmuller, Andrew M. Dai, Andrew La, Andrew Kyle Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakacs, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartlomiej Bojanowski, Batuhan Ozyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Stephen Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, C’esar Ferri Ram’irez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Daniel H Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Mosegu’i Gonz’alez, Danielle R. Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth P. Donoway, Ellie Pavlick, Emanuele Rodolà, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan J. Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Mart’inez-Plumed, Francesca Happ’e, François Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovich-L’opez, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Schutze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, John Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Koco’n, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Narain Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Oluwadara Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Jane W Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jorg Frohberg, Jos Rozen, José Hernández-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Wallace Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Luca Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Col’on, Luke Metz, Lutfi Kerem cSenel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ram’irez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, M’aty’as Schubert, Medina Baitemirova, Melody Arnaud, Melvin Andrew McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Igorevich Ivanitskiy, Michael Starr, Michael Strube, Michal Swkedrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Monica Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, T MukundVarma, Nanyun Peng, Nathan A. Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pi-Bei Hwang, P. Milkowski, Piyush S. Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphael Milliere, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan Lebras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi S. Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsunori Hashimoto, Te-Lin Wu, Theo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Venkatesh Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yu Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. ArXiv, abs/2206.04615.
  80. 80.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. 2023. Challenging BIG-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051, Toronto, Canada. Association for Computational Linguistics.
  81. 81.Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju. 2023. Quantifying uncertainty in natural language explanations of large language models. ArXiv, abs/2311.03533.
  82. 82.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5433–5442, Singapore. Association for Computational Linguistics.
  83. 83.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  84. 84.Ming-Yu Tseng and Grace Zhang. 2023. How uncertainty can be turned into shared understanding. A Pragmatic Agenda for Healthcare: Fostering inclusion and active participation through shared understanding, 338:373.
  85. 85.Thomas S Wallsten, David V Budescu, Amnon Rapoport, Rami Zwick, and Barbara Forsyth. 1986. Measuring the vague meanings of probability terms. Journal of Experimental Psychology: General, 115(4):348.
  86. 86.Xinru Wang and Ming Yin. 2021. Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making. In 26th International Conference on Intelligent User Interfaces, IUI ’21, page 318–328, New York, NY, USA. Association for Computing Machinery.
  87. 87.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Huai hsin Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. ArXiv, abs/2203.11171.
  88. 88.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. Advances in neural information processing systems, 35:24824–24837.
  89. 89.Paul D Windschitl and Gary L Wells. 1996. Measuring psychological uncertainty: Verbal versus numeric methods. Journal of Experimental Psychology: Applied, 2(4):343.
  90. 90.Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs. In The Twelfth International Conference on Learning Representations.
  91. 91.Oktay Yagız and Cuneyt Demir. 2014. Hedging strategies in academic discourse: a comparative analysis of Turkish writers and native writers of English. Procedia-Social and Behavioral Sciences, 158:260–268.
  92. 92.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023. Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv preprint arXiv:2309.01219.
  93. 93.Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J Liu. 2023. Slic-hf: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425.
  94. 94.Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023. Navigating the grey area: How expressions of uncertainty and overconfidence affect language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5506–5524, Singapore. Association for Computational Linguistics.

Citation

MLA
Zhou, K., et al. “Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3623–43, https://doi.org/10.18653/v1/2024.acl-long.198.
APA
Zhou, K., Hwang, J. D., Ren, X., & Sap, M. (2024). Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3623–3643. https://doi.org/10.18653/v1/2024.acl-long.198
Chicago
Zhou, K., J. D. Hwang, X. Ren, and M. Sap. 2024. “Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3623–43. https://doi.org/10.18653/v1/2024.acl-long.198.
Harvard
Zhou, K. et al. (2024) “Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3623–3643. Available at: https://doi.org/10.18653/v1/2024.acl-long.198.
Vancouver
1. Zhou K, Hwang JD, Ren X, Sap M (2024) Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3623–3643

BibTeX

@inproceedings{zhou-etal-2024-relying,
    title = "Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty",
    author = "Zhou, Kaitlyn  and
      Hwang, Jena D.  and
      Ren, Xiang  and
      Sap, Maarten",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.198/",
    doi = "10.18653/v1/2024.acl-long.198",
    pages = "3623--3643"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/