Built independently by an author, for readers. Read the story and support ChapterPal

keyword

model confidence

Model confidence is the estimated degree of certainty or probability that a machine learning model assigns to the correctness of its own predictions or generated outputs. Often quantified through predicted probability distributions or evaluated via expressions of certainty, it serves as a primary metric for assessing whether a system is likely to be accurate. A central aspect of model confidence is calibration, which measures how well the assigned certainty aligns with actual empirical accuracy across tasks. Properly evaluating and calibrating confidence is essential for uncertainty quantification, error detection, and safe deployment, as it prevents systems from producing incorrect or misleading results with unwarranted certainty.

2 items

Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Kaitlyn Zhou, Dan Jurafsky, Tatsunori Hashimoto

OrganizationsStanford University

Why you should read this

Reveals that prompt expressions of high confidence paradoxically degrade language model accuracy by up to seven percent compared to expressions of uncertainty, exposing how pretraining data patterns cause models to mimic surface linguistic habits rather than genuinely track factual truth.

The increased deployment of LMs for real-world tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs. Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures. We develop a typology of epistemic markers and inject 50 markers into prompts for question answering. We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%. Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance. Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions. These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.

Added

2026-10-01

Inducing Artificial Uncertainty in Language Models

Inducing Artificial Uncertainty in Language Models

Sophia Hager, Simon Zeng, Nicholas Andrews

OrganizationsJohns Hopkins UniversityMicrosoft

Why you should read this

Demonstrates that training uncertainty probes on artificially induced uncertainty in language models significantly improves confidence calibration on difficult tasks where naturally challenging training data is scarce.

In safety-critical applications, language models should be able to characterize their uncertainty with meaningful probabilities. Many uncertainty quantification approaches require supervised data; however, finding suitable unseen challenging data is increasingly difficult for large language models trained on vast amounts of scraped data. If the model is consistently (and correctly) confident in its predictions, the uncertainty quantification method may consistently overestimate confidence on new and unfamiliar data. Finding data which exhibits enough uncertainty to train supervised uncertainty quantification methods for high-performance models may therefore be challenging, and will increase in difficulty as LLMs saturate datasets. To address this issue, we first introduce the problem of inducing artificial uncertainty in language models, then investigate methods of inducing artificial uncertainty on trivially easy data in the absence of challenging data at training time. We use probes trained to recognize artificial uncertainty on the original model, and find that these probes trained on artificial uncertainty outperform probes trained without artificial uncertainty in recognizing real uncertainty, achieving notably higher calibration on hard data with minimal loss of performance on easy data.

Added

2026-09-29