Built independently by an author, for readers. Read the story and support ChapterPal

keyword

calibration error

Calibration error is a metric in machine learning and statistical modeling that measures the discrepancy between a model predicted confidence scores and the actual empirical likelihood that its predictions are correct. In a perfectly calibrated system, the assigned probability or confidence matches the true rate of correctness, meaning that predictions made with eighty percent confidence are accurate eighty percent of the time. When this alignment fails, the calibration error quantifies the degree to which the model is systematically overconfident or underconfident. It is typically evaluated across grouped confidence intervals using aggregate metrics such as expected calibration error, which calculates the weighted average difference between predicted confidences and observed accuracies across discrete prediction bins.

6 items

Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, Christopher D. Manning

OrganizationsHarvard UniversityStanford University

Why you should read this

Demonstrates that directly prompting RLHF-tuned language models like GPT-4 and Claude to verbalize their confidence scores achieves significantly better calibration than extracting their conditional probabilities, reducing expected calibration error by up to 50% across standard benchmarks.

A trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to an expert in cases of low-confidence predictions. Recent studies have shown that unsupervised pre-training produces large language models (LMs) whose conditional probabilities are remarkably well-calibrated. However, the most widely-used LMs are fine-tuned with reinforcement learning from human feedback (RLHF-LMs), and some studies have suggested that RLHF-LMs produce conditional probabilities that are very poorly calibrated. In light of this perceived weakness, we conduct a broad evaluation of methods for extracting confidence scores from RLHF-LMs. For RLHF-LMs such as ChatGPT, GPT-4, and Claude, we find that verbalized confidences emitted as output tokens are typically better-calibrated than the model's conditional probabilities on the TriviaQA, SciQ, and TruthfulQA benchmarks, often reducing the expected calibration error by a relative 50%.

Added

2026-10-05

Dual Focal Loss for Calibration

Dual Focal Loss for Calibration

Linwei Tao, Minjing Dong, Chang Xu

OrganizationsUniversity of Sydney

Why you should read this

Proposes Dual Focal Loss, a training objective that balances over-confidence and under-confidence in neural network calibration by maximizing the gap between the ground-truth logit and the highest-ranked competing logit.

The use of deep neural networks in real-world applications require well-calibrated networks with confidence scores that accurately reflect the actual probability. However, it has been found that these networks often provide over-confident predictions, which leads to poor calibration. Recent efforts have sought to address this issue by focal loss to reduce over-confidence, but this approach can also lead to under-confident predictions. While different variants of focal loss have been explored, it is difficult to find a balance between over-confidence and under-confidence. In our work, we propose a new loss function by focusing on dual logits. Our method not only considers the ground truth logit, but also take into account the highest logit ranked after the ground truth logit. By maximizing the gap between these two logits, our proposed dual focal loss can achieve a better balance between over-confidence and under-confidence. We provide theoretical evidence to support our approach and demonstrate its effectiveness through evaluations on multiple models and datasets, where it achieves state-of-the-art performance. Code is available at https://github.com/Linwei94/DualFocalLoss

Added

2026-10-03

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Lorenzo Jaime Flores, Cesare Spinoso di-Piano, Jackie Cheung

OrganizationsCIFARMcGill UniversityMila – Québec Artificial Intelligence Institute

Why you should read this

Demonstrates that supervised fine-tuning unpredictably alters language model confidence calibration across text generation tasks, significantly undermining the effectiveness of standard uncertainty metrics for hallucination detection and selective prediction.

Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection and selective prediction. While prior work has developed various confidence metrics and demonstrated their calibration for classification tasks or using verbalized confidence, the robustness of probability-based and self-consistency-based UQ metrics for natural language generation remains underexplored particularly under model adaptation. Since practitioners routinely apply supervised fine-tuning to adapt models to new tasks, a key question arises: do confidence metrics maintain their calibration when models are fine-tuned? We investigate this question across NLG tasks including translation, question answering, and mathematical reasoning. We find that calibration shifts substantially after SFT: across 216 configurations, it degrades in 112 cases and improves in 104, with confidence scores shifting due to factors beyond output quality, such as proximity to the training distribution. Degradation is therefore neither universal nor rare, and its direction cannot be anticipated from the pre-SFT model. Through a downstream task evaluation, we show that this miscalibration substantially reduces the practical utility of confidence scores for identifying correct answers. Our findings reveal that existing confidence metrics for NLG cannot be reliably deployed off-the-shelf after fine-tuning, highlighting the need for calibration-robust UQ methods under model adaptation.

Added

2026-09-29

When and How Mixup Improves Calibration

When and How Mixup Improves Calibration

Linjun Zhang, Zhun Deng, Kenji Kawaguchi, James Zou

OrganizationsHarvard UniversityNational University of SingaporeRutgers UniversityStanford University

Why you should read this

Establishes the theoretical foundation for why Mixup data augmentation reduces prediction calibration errors in high-dimensional settings and semi-supervised learning, showing that its calibration benefits grow alongside model capacity.

In many machine learning applications, it is important for the model to provide confidence scores that accurately capture its prediction uncertainty. Although modern learning methods have achieved great success in predictive accuracy, generating calibrated confidence scores remains a major challenge. Mixup, a popular yet simple data augmentation technique based on taking convex combinations of pairs of training examples, has been empirically found to significantly improve confidence calibration across diverse applications. However, when and how Mixup helps calibration is still a mystery. In this paper, we theoretically prove that Mixup improves calibration in high-dimensional settings by investigating natural statistical models. Interestingly, the calibration benefit of Mixup increases as the model capacity increases. We support our theories with experiments on common architectures and datasets. In addition, we study how Mixup improves calibration in semi-supervised learning. While incorporating unlabeled data can sometimes make the model less calibrated, adding Mixup training mitigates this issue and provably improves calibration. Our analysis provides new insights and a framework to understand Mixup and calibration.

Added

2026-09-26

On Calibration of Modern Neural Networks

On Calibration of Modern Neural Networks

Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger

OrganizationsCornell University

Why you should read this

Demonstrates that modern neural networks are surprisingly poorly calibrated and presents a simple yet highly effective post-processing method (temperature scaling) to address this critical issue for practical applications.

Confidence calibration -- the problem of predicting probability estimates representative of the true correctness likelihood -- is important for classification models in many applications. We discover that modern neural networks, unlike those from a decade ago, are poorly calibrated. Through extensive experiments, we observe that depth, width, weight decay, and Batch Normalization are important factors influencing calibration. We evaluate the performance of various post-processing calibration methods on state-of-the-art architectures with image and document classification datasets. Our analysis and experiments not only offer insights into neural network learning, but also provide a simple and straightforward recipe for practical settings: on most datasets, temperature scaling -- a single-parameter variant of Platt Scaling -- is surprisingly effective at calibrating predictions.

Added

2026-03-23

Creative Commons License