Built independently by an author, for readers. Read the story and support ChapterPal

keyword

prediction-rejection ratio

The prediction-rejection ratio is an evaluation metric in machine learning used to quantify how effectively a model uncertainty or confidence scores distinguish correct predictions from erroneous ones for selective prediction. It operates on prediction-rejection curves, which illustrate how the average quality or accuracy of retained model outputs improves as uncertain predictions are progressively discarded. The metric is computed by measuring the improvement in the area under the prediction-rejection curve achieved by an uncertainty estimation method over a random rejection baseline, normalized by the maximum possible improvement achieved by an optimal oracle over that baseline. Ranging typically from zero to one, a higher ratio indicates that the scoring mechanism reliably orders predictions according to their true correctness, making it a key tool for evaluating uncertainty quantification and selective abstention strategies in high-stakes tasks.

1 item

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

Lorenzo Jaime Flores, Cesare Spinoso di-Piano, Jackie Cheung

OrganizationsCIFARMcGill UniversityMila – Québec Artificial Intelligence Institute

Why you should read this

Demonstrates that supervised fine-tuning unpredictably alters language model confidence calibration across text generation tasks, significantly undermining the effectiveness of standard uncertainty metrics for hallucination detection and selective prediction.

Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection and selective prediction. While prior work has developed various confidence metrics and demonstrated their calibration for classification tasks or using verbalized confidence, the robustness of probability-based and self-consistency-based UQ metrics for natural language generation remains underexplored particularly under model adaptation. Since practitioners routinely apply supervised fine-tuning to adapt models to new tasks, a key question arises: do confidence metrics maintain their calibration when models are fine-tuned? We investigate this question across NLG tasks including translation, question answering, and mathematical reasoning. We find that calibration shifts substantially after SFT: across 216 configurations, it degrades in 112 cases and improves in 104, with confidence scores shifting due to factors beyond output quality, such as proximity to the training distribution. Degradation is therefore neither universal nor rare, and its direction cannot be anticipated from the pre-SFT model. Through a downstream task evaluation, we show that this miscalibration substantially reduces the practical utility of confidence scores for identifying correct answers. Our findings reveal that existing confidence metrics for NLG cannot be reliably deployed off-the-shelf after fine-tuning, highlighting the need for calibration-robust UQ methods under model adaptation.

Added

2026-09-29