keyword
calibration metrics
Calibration metrics are quantitative measures used in machine learning and statistics to evaluate how accurately a model's predicted confidence levels or probabilities reflect the true likelihood of the predicted outcomes. Unlike standard performance metrics that only measure the rate of correct classifications, calibration metrics assess whether a model's certainty corresponds to actual empirical success rates, determining whether a system is overconfident or underconfident in its predictions. Common calibration metrics include Expected Calibration Error, Maximum Calibration Error, the Brier score, and negative log-likelihood, which typically group predictions into confidence intervals or evaluate continuous scoring rules to quantify the statistical discrepancy between predicted probabilities and observed frequencies.
1 item

