Obtaining Well Calibrated Probabilities Using Bayesian Binning
Mahdi Pakdaman NaeiniGregory F. CooperMilos Hauskrecht
Introduces Bayesian Binning into Quantiles (BBQ), a computationally tractable post-processing method that combines multiple binning schemes via Bayesian model averaging to produce highly calibrated probability predictions for binary classifiers without requiring strict monotonicity assumptions.
Modern predictive models are widely used to guide critical decisions in fields such as healthcare, scientific research, and finance. However, standard machine learning classifiers often output poorly calibrated probability scores—meaning a predicted 70 percent probability of an event does not actually occur 70 percent of the time. This miscalibration poses significant risks for decision-makers who rely on accurate risk assessments. The article introduces and evaluates Bayesian Binning into Quantiles (BBQ), a flexible post-processing method designed to convert raw classifier outputs into dependable, well-calibrated probabilities without requiring changes to the underlying model training process.
The authors evaluated BBQ against standard calibration techniques, including Platt scaling, isotonic regression, and traditional single-model histogram binning. To assess performance, they conducted experiments on simulated non-linear data as well as 30 real-world binary classification benchmark datasets from the UCI and LibSVM repositories. The evaluation applied three standard machine learning base classifiers (Logistic Regression, Support Vector Machines, and Naive Bayes) across five key metrics: discrimination ability (Accuracy and Area Under the ROC Curve) and probability calibration quality (Expected Calibration Error, Maximum Calibration Error, and Root Mean Square Error). Rigorous non-parametric statistical hypothesis testing was used to compare performance across all datasets.
The empirical findings demonstrate that BBQ delivers superior calibration while preserving the original predictive power of the classifiers. Across the 30 benchmark datasets, BBQ achieved statistically significant superiority over all competing calibration methods and uncalibrated models in reducing both Expected Calibration Error and Maximum Calibration Error. In error reduction as measured by Root Mean Square Error, BBQ consistently outperformed base models, Platt scaling, and simple histogram binning, performing on par with isotonic regression. Furthermore, BBQ maintained full discrimination performance without degrading classifier accuracy or ranking ability. On non-linear simulated data where basic monotonicity assumptions failed, BBQ substantially outperformed Platt scaling and isotonic regression.
These results show that BBQ offers a reliable, computationally tractable drop-in solution to improve the trustworthiness of machine learning risk estimates. By decoupling model training from calibration, teams can optimize classifiers for discrimination and subsequently apply BBQ to obtain accurate probabilities. This lowers the operational and financial risks associated with overconfident or misaligned probability forecasts. Practitioners deploying binary classification models for high-stakes decision-making are advised to implement BBQ as a standard post-processing step. Future work highlighted in the article includes establishing formal theoretical bounds and extending the framework to multi-class and multi-label prediction problems.
- Paper: Predicting good probabilities with supervised learning, Alexandru Niculescu-Mizil et al. (2005). This benchmark paper establishes the empirical foundation of probability calibration across supervised algorithms and popular post-processing techniques like Platt scaling and isotonic regression, which Bayesian Binning directly aims to improve upon.
- Paper: An empirical comparison of supervised learning algorithms, R. Caruana et al. (2006). Provides a comprehensive empirical evaluation of classifier calibration across multiple models and metrics, motivating the need for more flexible post-hoc calibration methods.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Extends the study of post-processing probability calibration and binning-based techniques to modern deep neural networks and standardizes expected calibration error evaluations.
- Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). Investigates regularized training strategies that inherently improve model calibration as an alternative or complement to post-hoc binning calibrations.
- Paper: Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift, Yaniv Ovadia et al. (2019). Evaluates how post-hoc probability calibration and uncertainty estimation methods behave under severe dataset shifts.
