Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
Kaitlyn ZhouJena D. HwangXiang RenMaarten Sap
Reveals that language models rarely express uncertainty and suffer from a 47% error rate on confident statements, tracing this miscalibration to human preference biases during post-training alignment and demonstrating the resulting risks of human overreliance.
As natural language becomes the primary interface for human-artificial intelligence interaction, users increasingly rely on language models for critical decision-making and information retrieval. In human communication, speakers use linguistic cues called epistemic markers—specifically strengtheners that express confidence and weakeners that convey doubt—to signal reliability. However, when artificial intelligence systems fail to accurately communicate their limitations, users face significant risks of acting on incorrect advice. The article investigates how widely deployed language models express certainty, how human users interpret these linguistic markers, and where model overconfidence originates.
The article evaluates model calibration across major commercial systems and conducts behavioral experiments to quantify human reliance on model responses. To establish these patterns, the authors prompted nine prominent models, including variants of GPT, Claude, and LLaMA-2, across more than 125,000 queries using multiple-choice benchmark questions. They then conducted interactive user studies using challenging trivia to assess how participants rely on model advice under calibrated, overconfident, and underconfident conditions. Finally, the authors analyzed reinforcement learning datasets and reward models to identify the training mechanisms driving model behavior.
The findings reveal that language models are inherently reluctant to express uncertainty and, when prompted to do so, exhibit severe overconfidence. Across baseline tests, only 5% of model generations included uncertainty markers; when prompted to indicate confidence, models favored strengtheners over weakeners, resulting in an average error rate of 47% among confident answers. In human experiments, participants relied on plain statements without markers nearly 90% of the time, treating silence on uncertainty as implicit confidence. While users readily adapted to calibrated or underconfident models, exposure to overconfident models caused persistent judgment errors, leading participants to rely on incorrect answers 73% of the time during miscalibrated periods and degrading their accuracy even after the model reverted to calibrated behavior. Further analysis showed that this overconfidence originates during reinforcement learning from human feedback, as human evaluators and reward models systematically penalize expressions of uncertainty, preferring plain or assertive responses.
These results demonstrate that current alignment methods inadvertently incentivize artificial intelligence systems to conceal uncertainty, creating acute safety, compliance, and performance risks. Because users interpret plain statements as authoritative, uncalibrated language models can easily mislead human decision-makers. Moreover, early exposure to overconfident system behavior causes lasting damage to user judgment and trust, potentially driving algorithmic aversion. To address these vulnerabilities, practitioners should train models to produce unsolicited expressions of uncertainty when confidence is low, incorporate diverse spoken-language sources to expand hedging capabilities, and establish context-dependent certainty thresholds for high-stakes deployments.
The article notes several limitations, including a focus entirely on English-language models and user studies restricted to United States participants, which may not capture cross-cultural differences in how uncertainty is interpreted. Additionally, the experimental trivia tasks may not fully mirror the dynamics of high-stakes, real-world deployments. Nevertheless, the evidence strongly indicates that human feedback mechanisms drive systemic model overconfidence, warranting caution among organizations deploying conversational artificial intelligence in decision-support roles.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its experiments on models’ ability to judge whether their own answers are correct establish the self-knowledge problem that this paper examines when confidence is communicated to users.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). Its analysis of how human preference judgments reward sycophantic answers helps explain why alignment data may disfavor uncertainty, a mechanism this paper investigates directly.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). Its review of human-feedback alignment practices provides the preference-learning context needed to understand this paper’s analysis of uncertainty bias in post-training datasets.
- Paper: RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models, Aashiq Muhamed et al. (2025). It carries the concern about uncertainty communication into grounded systems, testing whether models appropriately abstain when their available evidence is ambiguous, incomplete, or contradictory.
