Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
Kaitlyn ZhouDan JurafskyTatsunori Hashimoto
Reveals that prompt expressions of high confidence paradoxically degrade language model accuracy by up to seven percent compared to expressions of uncertainty, exposing how pretraining data patterns cause models to mimic surface linguistic habits rather than genuinely track factual truth.
As large language models are increasingly deployed in real-world environments requiring factual accuracy, understanding how they process human expressions of certainty, uncertainty, and evidential attribution is critical. Humans naturally use linguistic cues to frame the reliability of knowledge, yet standard model evaluations often overlook these nuanced expressions.
The main objective of the article is to evaluate how linguistic markers of uncertainty and certainty injected into input prompts affect the accuracy and probability distributions of language model outputs in question-answering tasks.
To conduct this evaluation, the researchers established a linguistic taxonomy categorizing 50 distinct epistemic markers into weakeners, strengtheners, factive verbs, and evidential citations. They applied zero-shot and few-shot prompting across four question-answering datasets—TriviaQA, Natural Questions, CountryQA, and Jeopardy—and tested multiple OpenAI GPT models ranging from early base models to GPT-4. They also performed quantitative and qualitative analyses on pretraining data from The Pile to investigate the linguistic origins of the observed model behaviors.
The article establishes several key findings. First, language models are extraordinarily sensitive to epistemic phrasing in prompts, causing accuracy to fluctuate by up to 80% on identical questions. Second, expressions of high certainty paradoxically degrade model performance: weakeners (hedges) outperformed strengtheners (boosters) across all datasets, yielding an average accuracy of 47% compared to 40% for strengtheners. This accuracy loss was largely driven by factive verbs that presuppose truth. Third, evidential markers that cite sources significantly improved accuracy, frequently outperforming standard unprompted baselines. Fourth, injecting numerical confidence revealed poor linguistic calibration; accuracy peaked between 70% and 90% but dropped sharply when prompts claimed 100% certainty. Pretraining data analysis revealed that human writers predominantly use certainty phrases in questions to frame problems or admit ignorance, while using uncertainty phrases in answers to remain polite or cautious.
These findings indicate that language models do not possess true epistemological awareness; instead, they mimic conversational patterns observed during pretraining. For decision-makers and system developers, this introduces notable operational risks. Users who write confident prompts expecting better factual retrieval may unintentionally cause model hallucinations or errors. Furthermore, while grounding prompts with evidential markers like "Wikipedia says" boosts accuracy, blindly generating simulated attributions creates safety and compliance hazards by presenting fabricated sources persuasively.
Organizations deploying language models should avoid assuming that confident inputs or outputs correlate with factual accuracy. Developers should design prompts that avoid extreme certainty assertions and factive verbs, favor evidential grounding only when sources are verified, and position questions to elicit direct answers before any framing expressions. Further work is required to teach models how to navigate idiomatic versus literal confidence expressions and to verify external attributions automatically before integration into critical workflows.
Confidence in these findings is supported by consistent replication across six distinct model architectures and multiple datasets. However, certain limitations remain. The empirical evaluation relied exclusively on English-language trivia and discrete question-answering benchmarks, excluding long-form dialogue, continuous numerical targets, and multilingual cultural variations in hedging.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its experiments on whether language models can recognize what they know establish the model-calibration context that this paper extends by testing how epistemic wording changes answers.
- Paper: Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?, Gal Yona et al. (2024). Building on this paper’s finding that certainty wording can distort accuracy, it tests whether models’ verbal hedging faithfully tracks their intrinsic uncertainty.
- Paper: Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty, Kaitlyn Zhou et al. (2024). It carries the study of certainty language into model calibration and human reliance, examining the risks of models’ reluctance to express uncertainty.
- Paper: A Survey of Confidence Estimation and Calibration in Large Language Models, Jiahui Geng et al. (2024). This later survey places the paper’s findings on verbalized uncertainty within the broader research on confidence estimation and calibration in language models.
