keyword
distribution calibration
Distribution calibration is the statistical property or process of aligning a predictive model output probability distributions with the true empirical frequencies of the outcomes being predicted. A model achieves distribution calibration when its estimated probabilities or entire predicted cumulative distribution functions accurately represent the real-world likelihood of observing each possible class, event, or continuous target value. This concept extends single point confidence calibration to evaluate and correct the full spread of predictive uncertainty across all potential outcomes, ensuring that statistical dispersion and intervals remain unbiased. Proper distribution calibration is essential in probabilistic machine learning, risk-sensitive artificial intelligence, and decision-support frameworks where reliable and interpretable uncertainty quantification is necessary to make sound downstream decisions.
2 items

Linguistic Calibration of Long-Form Generations
Neil Band, Xuechen Li, Tengyu Ma, Tatsunori Hashimoto
Why you should read this
Proposes a decision-theoretic training framework that combines supervised fine-tuning and reinforcement learning to teach language models to express calibrated verbal confidence statements across long-form text, significantly improving downstream user decision-making without sacrificing generation accuracy.
Language models (LMs) may lead their users to make suboptimal downstream decisions when they confidently hallucinate. This issue can be mitigated by having the LM verbally convey the probability that its claims are correct, but existing models cannot produce long-form text with calibrated confidence statements. Through the lens of decision-making, we define linguistic calibration for long-form generations: an LM is linguistically calibrated if its generations enable its users to make calibrated probabilistic predictions. This definition enables a training framework where a supervised finetuning step bootstraps an LM to emit long-form generations with confidence statements such as “I estimate a 30% chance of...” or “I am certain that...”, followed by a reinforcement learning step which rewards generations that enable a user to provide calibrated answers to related questions. We linguistically calibrate Llama 2 7B and find in automated and human evaluations of long-form generations that it is significantly more calibrated than strong finetuned factuality baselines with comparable accuracy. These findings generalize under significant domain shifts to scientific and biomedical questions and to an entirely held-out person biography generation task. Our results demonstrate that long-form generations may be calibrated end-to-end by constructing an objective in the space of the predictions that users make in downstream decision-making.
Added
2026-10-03

Calibrated Learning to Defer with One-vs-All Classifiers
Rajeev Verma, Eric T. Nalisnick
Why you should read this
Develops a consistent one-vs-all surrogate loss for multiclass learning-to-defer systems that overcomes the uncalibrated, degenerate probability estimates of standard softmax formulations while matching or exceeding classification accuracy across diverse domains.
The learning to defer (L2D) framework has the potential to make AI systems safer. For a given input, the system can defer the decision to a human if the human is more likely than the model to take the correct action. We study the calibration of L2D systems, investigating if the probabilities they output are sound. We find that Mozannar & Sontag’s (2020) multiclass framework is not calibrated with respect to expert correctness. Moreover, it is not even guaranteed to produce valid probabilities due to its parameterization being degenerate for this purpose. We propose an L2D system based on one-vs-all classifiers that is able to produce calibrated probabilities of expert correctness. Furthermore, our loss function is also a consistent surrogate for multiclass L2D, like Mozannar & Sontag’s (2020). Our experiments verify that not only is our system calibrated, but this benefit comes at no cost to accuracy. Our model’s accuracy is always comparable (and often superior) to Mozannar & Sontag’s (2020) model’s in tasks ranging from hate speech detection to galaxy classification to diagnosis of skin lesions.
Added
2026-10-02
