Built independently by an author, for readers. Read the story and support ChapterPal

keyword

calibration metrics

Calibration metrics are quantitative measures used in machine learning and statistics to evaluate how accurately a model's predicted confidence levels or probabilities reflect the true likelihood of the predicted outcomes. Unlike standard performance metrics that only measure the rate of correct classifications, calibration metrics assess whether a model's certainty corresponds to actual empirical success rates, determining whether a system is overconfident or underconfident in its predictions. Common calibration metrics include Expected Calibration Error, Maximum Calibration Error, the Brier score, and negative log-likelihood, which typically group predictions into confidence intervals or evaluate continuous scoring rules to quantify the statistical discrepancy between predicted probabilities and observed frequencies.

1 item