Built independently by an author, for readers. Read the story and support ChapterPal

keyword

probabilistic predictions

Probabilistic predictions are forecasts or estimates that quantify the likelihood of different possible outcomes by assigning a probability or probability distribution to them, rather than stating a single definitive result. Unlike deterministic point predictions, which output a single value or category, probabilistic predictions explicitly capture and communicate uncertainty, conveying the degree of confidence or statistical chance associated with each potential event. They can be expressed through discrete probability values, continuous probability density functions, or confidence ratings. In statistics, machine learning, and decision-making, probabilistic predictions allow individuals and systems to account for risk and make optimal choices under conditions of incomplete information, with their quality typically assessed through calibration, which measures how closely the assigned probabilities match the true observed frequencies of the outcomes over time.

3 items

Linguistic Calibration of Long-Form Generations

Linguistic Calibration of Long-Form Generations

Neil Band, Xuechen Li, Tengyu Ma, Tatsunori Hashimoto

OrganizationsStanford University

Why you should read this

Proposes a decision-theoretic training framework that combines supervised fine-tuning and reinforcement learning to teach language models to express calibrated verbal confidence statements across long-form text, significantly improving downstream user decision-making without sacrificing generation accuracy.

Language models (LMs) may lead their users to make suboptimal downstream decisions when they confidently hallucinate. This issue can be mitigated by having the LM verbally convey the probability that its claims are correct, but existing models cannot produce long-form text with calibrated confidence statements. Through the lens of decision-making, we define linguistic calibration for long-form generations: an LM is linguistically calibrated if its generations enable its users to make calibrated probabilistic predictions. This definition enables a training framework where a supervised finetuning step bootstraps an LM to emit long-form generations with confidence statements such as “I estimate a 30% chance of...” or “I am certain that...”, followed by a reinforcement learning step which rewards generations that enable a user to provide calibrated answers to related questions. We linguistically calibrate Llama 2 7B and find in automated and human evaluations of long-form generations that it is significantly more calibrated than strong finetuned factuality baselines with comparable accuracy. These findings generalize under significant domain shifts to scientific and biomedical questions and to an entirely held-out person biography generation task. Our results demonstrate that long-form generations may be calibrated end-to-end by constructing an objective in the space of the predictions that users make in downstream decision-making.

Added

2026-10-03

Language Models (Mostly) Know What They Know

Language Models (Mostly) Know What They Know

Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, Jared Kaplan

OrganizationsAnthropic

Why you should read this

Demonstrates that larger language models can reliably evaluate the correctness of their own generated statements and predict whether they know the answer to a question, providing an empirical basis for training more honest AI.

We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling tasks by asking models to first propose answers, and then to evaluate the probability "P(True)" that their answers are correct. We find encouraging performance, calibration, and scaling for P(True) on a diverse array of tasks. Performance at self-evaluation further improves when we allow models to consider many of their own samples before predicting the validity of one specific possibility. Next, we investigate whether models can be trained to predict "P(IK)", the probability that "I know" the answer to a question, without reference to any particular proposed answer. Models perform well at predicting P(IK) and partially generalize across tasks, though they struggle with calibration of P(IK) on new tasks. The predicted P(IK) probabilities also increase appropriately in the presence of relevant source materials in the context, and in the presence of hints towards the solution of mathematical word problems. We hope these observations lay the groundwork for training more honest models, and for investigating how honesty generalizes to cases where models are trained on objectives other than the imitation of human writing.

Added

2026-09-17

Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Eyke Hüllermeier, Willem Waegeman

OrganizationsGhent UniversityPaderborn University

Why you should read this

Clarifies the critical distinction between irreducible data randomness and reducible model ignorance, providing a comprehensive framework for quantifying both aleatoric and epistemic uncertainty to build safer, more reliable machine learning systems.

The notion of uncertainty is of major importance in machine learning and constitutes a key element of machine learning methodology. In line with the statistical tradition, uncertainty has long been perceived as almost synonymous with standard probability and probabilistic predictions. Yet, due to the steadily increasing relevance of machine learning for practical applications and related issues such as safety requirements, new problems and challenges have recently been identified by machine learning scholars, and these problems may call for new methodological developments. In particular, this includes the importance of distinguishing between (at least) two different types of uncertainty, often referred to as aleatoric and epistemic. In this paper, we provide an introduction to the topic of uncertainty in machine learning as well as an overview of attempts so far at handling uncertainty in general and formalizing this distinction in particular.

Added

2026-09-16