Built independently by an author, for readers. Read the story and support ChapterPal

keyword

confidence scores

A confidence score is a quantitative metric, typically expressed as a probability or numerical value between zero and one, that represents a machine learning model's estimated certainty in the correctness of its output or prediction. In tasks such as classification, natural language generation, and decision-making, confidence scores can be derived from output probability distributions, raw model logits, sample consistency across multiple outputs, or explicitly generated certainty statements. When a model is well-calibrated, its confidence scores accurately correspond to the true empirical likelihood of its predictions being correct, meaning that predictions assigned a specific confidence level succeed at a matching rate. These scores are widely used in uncertainty quantification to assess output reliability, identify potential errors or hallucinations, implement selective prediction, and determine when an automated system should trigger human review.

3 items

Dual Focal Loss for Calibration

Dual Focal Loss for Calibration

Linwei Tao, Minjing Dong, Chang Xu

OrganizationsUniversity of Sydney

Why you should read this

Proposes Dual Focal Loss, a training objective that balances over-confidence and under-confidence in neural network calibration by maximizing the gap between the ground-truth logit and the highest-ranked competing logit.

The use of deep neural networks in real-world applications require well-calibrated networks with confidence scores that accurately reflect the actual probability. However, it has been found that these networks often provide over-confident predictions, which leads to poor calibration. Recent efforts have sought to address this issue by focal loss to reduce over-confidence, but this approach can also lead to under-confident predictions. While different variants of focal loss have been explored, it is difficult to find a balance between over-confidence and under-confidence. In our work, we propose a new loss function by focusing on dual logits. Our method not only considers the ground truth logit, but also take into account the highest logit ranked after the ground truth logit. By maximizing the gap between these two logits, our proposed dual focal loss can achieve a better balance between over-confidence and under-confidence. We provide theoretical evidence to support our approach and demonstrate its effectiveness through evaluations on multiple models and datasets, where it achieves state-of-the-art performance. Code is available at https://github.com/Linwei94/DualFocalLoss

Added

2026-10-03

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Lorenzo Jaime Flores, Cesare Spinoso di-Piano, Jackie Cheung

OrganizationsCIFARMcGill UniversityMila – Québec Artificial Intelligence Institute

Why you should read this

Demonstrates that supervised fine-tuning unpredictably alters language model confidence calibration across text generation tasks, significantly undermining the effectiveness of standard uncertainty metrics for hallucination detection and selective prediction.

Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection and selective prediction. While prior work has developed various confidence metrics and demonstrated their calibration for classification tasks or using verbalized confidence, the robustness of probability-based and self-consistency-based UQ metrics for natural language generation remains underexplored particularly under model adaptation. Since practitioners routinely apply supervised fine-tuning to adapt models to new tasks, a key question arises: do confidence metrics maintain their calibration when models are fine-tuned? We investigate this question across NLG tasks including translation, question answering, and mathematical reasoning. We find that calibration shifts substantially after SFT: across 216 configurations, it degrades in 112 cases and improves in 104, with confidence scores shifting due to factors beyond output quality, such as proximity to the training distribution. Degradation is therefore neither universal nor rare, and its direction cannot be anticipated from the pre-SFT model. Through a downstream task evaluation, we show that this miscalibration substantially reduces the practical utility of confidence scores for identifying correct answers. Our findings reveal that existing confidence metrics for NLG cannot be reliably deployed off-the-shelf after fine-tuning, highlighting the need for calibration-robust UQ methods under model adaptation.

Added

2026-09-29