Obtaining Well Calibrated Probabilities Using Bayesian Binning

Mahdi Pakdaman NaeiniGregory F. CooperMilos Hauskrecht

article2015AAAI2,038 citations

Introduces Bayesian Binning into Quantiles (BBQ), a computationally tractable post-processing method that combines multiple binning schemes via Bayesian model averaging to produce highly calibrated probability predictions for binary classifiers without requiring strict monotonicity assumptions.

Listen

Modern predictive models are widely used to guide critical decisions in fields such as healthcare, scientific research, and finance. However, standard machine learning classifiers often output poorly calibrated probability scores—meaning a predicted 70 percent probability of an event does not actually occur 70 percent of the time. This miscalibration poses significant risks for decision-makers who rely on accurate risk assessments. The article introduces and evaluates Bayesian Binning into Quantiles (BBQ), a flexible post-processing method designed to convert raw classifier outputs into dependable, well-calibrated probabilities without requiring changes to the underlying model training process.

The authors evaluated BBQ against standard calibration techniques, including Platt scaling, isotonic regression, and traditional single-model histogram binning. To assess performance, they conducted experiments on simulated non-linear data as well as 30 real-world binary classification benchmark datasets from the UCI and LibSVM repositories. The evaluation applied three standard machine learning base classifiers (Logistic Regression, Support Vector Machines, and Naive Bayes) across five key metrics: discrimination ability (Accuracy and Area Under the ROC Curve) and probability calibration quality (Expected Calibration Error, Maximum Calibration Error, and Root Mean Square Error). Rigorous non-parametric statistical hypothesis testing was used to compare performance across all datasets.

The empirical findings demonstrate that BBQ delivers superior calibration while preserving the original predictive power of the classifiers. Across the 30 benchmark datasets, BBQ achieved statistically significant superiority over all competing calibration methods and uncalibrated models in reducing both Expected Calibration Error and Maximum Calibration Error. In error reduction as measured by Root Mean Square Error, BBQ consistently outperformed base models, Platt scaling, and simple histogram binning, performing on par with isotonic regression. Furthermore, BBQ maintained full discrimination performance without degrading classifier accuracy or ranking ability. On non-linear simulated data where basic monotonicity assumptions failed, BBQ substantially outperformed Platt scaling and isotonic regression.

These results show that BBQ offers a reliable, computationally tractable drop-in solution to improve the trustworthiness of machine learning risk estimates. By decoupling model training from calibration, teams can optimize classifiers for discrimination and subsequently apply BBQ to obtain accurate probabilities. This lowers the operational and financial risks associated with overconfident or misaligned probability forecasts. Practitioners deploying binary classification models for high-stakes decision-making are advised to implement BBQ as a standard post-processing step. Future work highlighted in the article includes establishing formal theoretical bounds and extending the framework to multi-class and multi-label prediction problems.

  • Paper: Predicting good probabilities with supervised learning, Alexandru Niculescu-Mizil et al. (2005). This benchmark paper establishes the empirical foundation of probability calibration across supervised algorithms and popular post-processing techniques like Platt scaling and isotonic regression, which Bayesian Binning directly aims to improve upon.
  • Paper: An empirical comparison of supervised learning algorithms, R. Caruana et al. (2006). Provides a comprehensive empirical evaluation of classifier calibration across multiple models and metrics, motivating the need for more flexible post-hoc calibration methods.
Cover for Obtaining Well Calibrated Probabilities Using Bayesian Binning

Abstract

Learning probabilistic predictive models that are well calibrated is critical for many prediction and decision-making tasks in artificial intelligence. In this paper we present a new non-parametric calibration method called Bayesian Binning into Quantiles (BBQ) which addresses key limitations of existing calibration methods. The method post processes the output of a binary classification algorithm; thus, it can be readily combined with many existing classification algorithms. The method is computationally tractable, and empirically accurate, as evidenced by the set of experiments reported here on both real and simulated datasets.

Table of Contents

  • Introduction
  • Methods
  • Calibration Measures
  • Empirical Results
  • Simulated Data
  • Real Data
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Bayesian Binning into Quantiles Calibration Framework

    model/method

    Bayesian Binning into Quantiles (BBQ) is a non-parametric post-processing method designed to produce well-calibrated posterior class probabilities from the real-valued or uncalibrated output score yy of a binary classifier. Rather than committing to a single quantile (equal-frequency) histogram binning with a fixed number of bins, BBQ performs Bayesian model averaging over a set of TT candidate equal-frequency binning models {M1,M2,…,MT}\{M_1, M_2, \dots, M_T\}.

    For an uncalibrated prediction yy, the BBQ calibrated probability estimate P(z=1∣y)P(z = 1 \mid y) is computed as the score-weighted average of probability estimates across all candidate models:

    P(z=1∣y)=∑i=1TScore(Mi)∑j=1TScore(Mj)P(z=1∣y,Mi)P(z = 1 \mid y) = \sum_{i=1}^T \frac{\text{Score}(M_i)}{\sum_{j=1}^T \text{Score}(M_j)} P(z = 1 \mid y, M_i)

    where:

    • z∈{0,1}z \in \{0, 1\} is the true binary target class label.
    • yy is the uncalibrated prediction score generated by the base classification model.
    • MiM_i represents candidate binning model ii.
    • Score(Mi)=P(Mi)⋅P(D∣Mi)\text{Score}(M_i) = P(M_i) \cdot P(D \mid M_i) is the Bayesian score of model MiM_i given calibration training dataset D={(y1,z1),…,(yN,zN)}D = \{(y_1, z_1), \dots, (y_N, z_N)\}, with P(Mi)P(M_i) modeled using a uniform prior over candidate models.
    • P(z=1∣y,Mi)P(z = 1 \mid y, M_i) is the probability estimate assigned to score yy by model MiM_i, computed using smoothed empirical bin frequencies consistent with Bayesian conjugate priors.
  2. Knowl 2 — Closed-Form Marginal Likelihood for Quantile Binning Models

    equation

    For a candidate equal-frequency binning model M={B,Pa,Θ}M = \{B, Pa, \Theta\} partitioning NN sorted uncalibrated training predictions into BB bins with bin parameters Θ={θ1,…,θB}\Theta = \{\theta_1, \dots, \theta_B\}, the marginal likelihood P(D∣M)P(D \mid M) over calibration dataset D={(y1,z1),…,(yN,zN)}D = \{(y_1, z_1), \dots, (y_N, z_N)\} has the following closed-form expression derived from the BDeu score:

    P(D∣M)=∏b=1BΓ(N′B)Γ(Nb+N′B)Γ(mb+αb)Γ(αb)Γ(nb+βb)Γ(βb)P(D \mid M) = \prod_{b=1}^B \frac{\Gamma\left(\frac{N'}{B}\right)}{\Gamma\left(N_b + \frac{N'}{B}\right)} \frac{\Gamma(m_b + \alpha_b)}{\Gamma(\alpha_b)} \frac{\Gamma(n_b + \beta_b)}{\Gamma(\beta_b)}

    where:

    • Γ(⋅)\Gamma(\cdot) is the standard Gamma function.
    • BB is the number of equal-frequency bins in model MM.
    • NbN_b is the total number of calibration instances mapped to bin bb.
    • mbm_b is the number of positive class instances (z=1z = 1) in bin bb.
    • nbn_b is the number of negative class instances (z=0z = 0) in bin bb, such that Nb=mb+nbN_b = m_b + n_b.
    • N′N' is the equivalent sample size expressing belief in the prior distribution (set to N′=2N' = 2).
    • αb\alpha_b and βb\beta_b are the parameters of the Beta prior distribution placed over the binomial parameter θb=P(z=1∣bin b)\theta_b = P(z=1 \mid \text{bin } b), defined as:

    αb=N′Bpb,βb=N′B(1−pb)\alpha_b = \frac{N'}{B} p_b, \qquad \beta_b = \frac{N'}{B} (1 - p_b)

    where pbp_b is the midpoint of the interval defining the bb-th bin in model MM.

    This marginal likelihood holds under three conditions: (1) calibration samples are independent and identically distributed, with class occurrences inside bin bb modeled by a Binomial distribution with parameter θb\theta_b; (2) class distributions across distinct bins are mutually independent; and (3) prior distributions over parameters θb\theta_b are modeled using Beta distributions.

  3. Knowl 3 — Candidate Bin Range Selection and Score-Based Pruning in BBQ

    model/method

    In Bayesian Binning into Quantiles (BBQ), candidate binning models are selected and pruned using a two-stage strategy:

    1. Candidate Range Definition: The number of equal-frequency bins BB is restricted to integer values within the range:

    B∈{⌊N3C⌋,…,⌈CN3⌉}B \in \left\{ \left\lfloor \frac{\sqrt[3]{N}}{C} \right\rfloor, \dots, \left\lceil C \sqrt[3]{N} \right\rceil \right\}

    where NN is the total number of calibration instances and CC is a constant controlling the range of candidate models (set to C=10C = 10). This O(N1/3)\mathcal{O}(N^{1/3}) scaling is motivated by minimax rate results for histogram classifiers under Lipschitz Bayes decision boundaries.

    1. Bayesian Score Pruning: After calculating Bayesian scores for all candidate models in the initial range, the models are sorted in descending order of their scores: S1≥S2≥⋯≥STS_1 \ge S_2 \ge \dots \ge S_T. To avoid averaging over poorly fitting models, the candidate set is pruned to the top kρk_\rho models:

    kρ=min⁡{k:Sk−Sk+1σ2≤ρ}k_\rho = \min \left\{ k : \frac{S_k - S_{k+1}}{\sigma^2} \le \rho \right\}

    where σ2=1T∑i=1T(Si−Sˉ)2\sigma^2 = \frac{1}{T}\sum_{i=1}^T (S_i - \bar{S})^2 is the empirical variance of the scores across all TT candidate models, and ρ>0\rho > 0 is a threshold controlling pruning sensitivity (set to ρ=0.001\rho = 0.001).

  4. Knowl 4 — Bayesian Binning into Quantiles Calibration Algorithm

    algorithm

    The Bayesian Binning into Quantiles (BBQ) algorithm fits a family of quantile binning models to calibration data and computes model-averaged calibrated posterior probabilities for uncalibrated prediction scores.

    Input: Calibration dataset D={(yi,zi)}i=1ND = \{(y_i, z_i)\}_{i=1}^N with uncalibrated predictions yi∈Ry_i \in \mathbb{R} and binary labels zi∈{0,1}z_i \in \{0, 1\}, new uncalibrated prediction score y∗y^*, hyperparameter C=10C = 10, equivalent sample size N′=2N' = 2, pruning threshold ρ=0.001\rho = 0.001.
    Output: Calibrated probability P(z∗=1∣y∗)P(z^* = 1 \mid y^*).
    Sort DD in ascending order according to yiy_i.
    Determine candidate bin range B={B∈Z:⌊N1/3/C⌋≤B≤⌈CN1/3⌉}\mathcal{B} = \{B \in \mathbb{Z} : \lfloor N^{1/3} / C \rfloor \le B \le \lceil C N^{1/3} \rceil \}.
    for each bin count B∈BB \in \mathcal{B}:
        Partition the sorted calibration instances into BB equal-frequency bins b=1,…,Bb = 1, \dots, B.
        for each bin b∈{1,…,B}b \in \{1, \dots, B\}:
            Compute bin interval midpoint pb=(min⁡y∈bin by+max⁡y∈bin by)/2p_b = (\min_{y \in \text{bin } b} y + \max_{y \in \text{bin } b} y) / 2.
            Set prior parameters αb=(N′/B)pb\alpha_b = (N' / B) p_b and βb=(N′/B)(1−pb)\beta_b = (N' / B) (1 - p_b).
            Count total instances NbN_b, positive instances mbm_b, and negative instances nbn_b in bin bb.
        Compute marginal likelihood P(D∣MB)=∏b=1BΓ(N′/B)Γ(Nb+N′/B)Γ(mb+αb)Γ(αb)Γ(nb+βb)Γ(βb)P(D \mid M_B) = \prod_{b=1}^B \frac{\Gamma(N'/B)}{\Gamma(N_b + N'/B)} \frac{\Gamma(m_b + \alpha_b)}{\Gamma(\alpha_b)} \frac{\Gamma(n_b + \beta_b)}{\Gamma(\beta_b)}.
        Compute model score Score(MB)=P(D∣MB)\text{Score}(M_B) = P(D \mid M_B) assuming a uniform model prior P(MB)P(M_B).
    Sort candidate models in descending order of score: S1≥S2≥⋯≥STS_1 \ge S_2 \ge \dots \ge S_T, where T=∣B∣T = |\mathcal{B}|.
    Compute empirical variance σ2\sigma^2 across scores {S1,…,ST}\{S_1, \dots, S_T\}.
    Find cutoff index kρ=min⁡{k∈{1,…,T−1}:(Sk−Sk+1)/σ2≤ρ}k_\rho = \min \{ k \in \{1, \dots, T-1\} : (S_k - S_{k+1}) / \sigma^2 \le \rho \}.
    Retain the top kρk_\rho models {M1,…,Mkρ}\{M_1, \dots, M_{k_\rho}\}.
    for each retained model MiM_i:
        Identify the bin bb in MiM_i containing y∗y^*.
        Compute smoothed bin probability P(z∗=1∣y∗,Mi)=(mb+αb)/(Nb+αb+βb)P(z^* = 1 \mid y^*, M_i) = (m_b + \alpha_b) / (N_b + \alpha_b + \beta_b).
    Compute model-averaged calibrated probability:
        P(z∗=1∣y∗)=∑i=1kρSi∑j=1kρSjP(z∗=1∣y∗,Mi)P(z^* = 1 \mid y^*) = \sum_{i=1}^{k_\rho} \frac{S_i}{\sum_{j=1}^{k_\rho} S_j} P(z^* = 1 \mid y^*, M_i).
    return $P(z^* = 1 \mid y^*)
  5. Knowl 5 — Expected Calibration Error and Maximum Calibration Error

    definition

    Expected Calibration Error (ECE) and Maximum Calibration Error (MCE) are scalar statistics used to quantify the miscalibration of probabilistic predictions against the empirical ground truth across KK reliability bins (typically K=10K = 10).

    Predictions are sorted and partitioned into KK bins spanning the probability range. For each bin i∈{1,…,K}i \in \{1, \dots, K\}:

    • oio_i is the empirical fraction of positive instances (z=1z = 1) in bin ii:

    oi=1∣Bi∣∑j∈Bizjo_i = \frac{1}{|B_i|} \sum_{j \in B_i} z_j

    • eie_i is the mean of the post-calibrated predicted probabilities for instances assigned to bin ii:

    ei=1∣Bi∣∑j∈Bip^je_i = \frac{1}{|B_i|} \sum_{j \in B_i} \hat{p}_j

    • P(i)=∣Bi∣NevalP(i) = \frac{|B_i|}{N_{\text{eval}}} is the proportion of total evaluated instances falling into bin ii.

    The metrics are defined as:

    • Expected Calibration Error (ECE), the weighted average absolute difference between predicted probabilities and observed empirical class frequencies:

    ECE=∑i=1KP(i)⋅∣oi−ei∣\text{ECE} = \sum_{i=1}^K P(i) \cdot |o_i - e_i|

    • Maximum Calibration Error (MCE), the worst-case absolute deviation between predicted probabilities and empirical class frequencies across all bins:

    MCE=max⁡i∈{1,…,K}∣oi−ei∣\text{MCE} = \max_{i \in \{1, \dots, K\}} |o_i - e_i|

    Lower values of ECE and MCE indicate superior calibration, with 0 indicating perfect calibration.

  6. Knowl 6 — Performance of Calibration Methods on Non-Linearly Separable Simulated Data

    data/table

    On a simulated 2D binary classification dataset with non-linearly separable classes (1000 calibration instances, 1000 test instances), calibration methods were evaluated when applied to an underfitting base model (linear SVM) and a well-specified base model (quadratic kernel SVM).

    Metric Uncalibrated SVM Histogram Binning Platt Scaling Isotonic Regression BBQ
    (a) Linear SVM
    AUC 0.50 0.84 0.50 0.65 0.85
    ACC 0.48 0.78 0.52 0.64 0.78
    RMSE 0.50 0.39 0.50 0.46 0.38
    ECE 0.28 0.07 0.28 0.35 0.03
    MCE 0.52 0.19 0.54 0.58 0.09
    (b) Quadratic Kernel SVM
    AUC 1.00 1.00 1.00 1.00 1.00
    ACC 0.99 0.99 0.99 0.99 0.99
    RMSE 0.21 0.09 0.19 0.08 0.08
    ECE 0.14 0.01 0.15 0.00 0.00
    MCE 0.35 0.04 0.32 0.03 0.03

    When the base model (Linear SVM) fails to provide a monotonic mapping due to non-linearity, parametric Platt scaling and monotonic Isotonic regression fail to calibrate (extECE=0.28 ext{ECE} = 0.28 and 0.350.35, respectively). BBQ relaxes strict monotonicity and successfully recovers both discrimination (AUC 0.85, ACC 0.78) and calibration (ECE 0.03, MCE 0.09). When an ideal discriminator is used (Quadratic SVM), BBQ achieves identical calibration and discrimination performance to isotonic regression (extECE=0.00 ext{ECE} = 0.00, extMCE=0.03 ext{MCE} = 0.03, extRMSE=0.08 ext{RMSE} = 0.08, extAUC=1.00 ext{AUC} = 1.00).

  7. Knowl 7 — Empirical Calibration and Discrimination Performance on 30 Benchmark Datasets

    empirical result

    Across 30 real-world binary classification benchmark datasets from the UCI and LibSVM repositories evaluated with Logistic Regression (LR), Support Vector Machines (SVM), and Naive Bayes (NB) base classifiers, statistical testing using the Friedman non-parametric test followed by Holm's step-down post-hoc procedure at significance level α=0.05\alpha = 0.05 demonstrates:

    • Calibration Error (ECE and MCE): BBQ achieves the best (lowest) average rank across all classifiers and is statistically significantly superior to all competing methods, including the uncalibrated base classifier, Platt scaling, single-model Histogram binning, and Isotonic regression (e.g., on LR, BBQ attains average ranks of ≈1.500\approx 1.500 for ECE and ≈1.703\approx 1.703 for MCE).
    • Brier Score / Root Mean Square Error (RMSE): BBQ statistically significantly outperforms the uncalibrated base classifier, Platt scaling, and Histogram binning. BBQ is statistically tied with Isotonic regression for the lowest RMSE rank (BBQ average rank ≈2.156\approx 2.156 on LR, 2.0172.017 on SVM, and 1.8171.817 on NB).
    • Discrimination (AUC and ACC): BBQ preserves or improves discrimination. In terms of AUC, BBQ is statistically significantly superior to single Histogram binning and statistically equivalent to the base classifier and Isotonic regression. In terms of classification Accuracy (ACC), there is no statistically significant difference among any of the post-processing calibration methods and base classifiers.

Coverage note — None was omitted; all core methods, derivations, algorithms, evaluation metrics, and experimental results on simulated and benchmark datasets are covered.

References

  1. 1.Bache, K., and Lichman, M. 2013. UCI machine learning repository.
  2. 2.Barlow, R. E.; Bartholomew, D. J.; Bremner, J.; and Brunk, H. D. 1972. Statistical inference under order restrictions: The theory and application of isotonic regression. Wiley New York.
  3. 3.Chang, C.-C., and Lin, C.-J. 2011. Libsvm: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology (TIST) 2(3):27.
  4. 4.DeGroot, M., and Fienberg, S. 1983. The comparison and evaluation of forecasters. The Statistician 12–22.
  5. 5.Demšar, J. 2006. Statistical comparisons of classifiers over multiple data sets. The Journal of Machine Learning Research 7:1–30.
  6. 6.Friedman, M. 1937. The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association 32(200):675–701.
  7. 7.Heckerman, D.; Geiger, D.; and Chickering, D. 1995. Learning bayesian networks: The combination of knowledge and statistical data. Machine Learning 20(3):197–243.
  8. 8.Hoeting, J. A.; Madigan, D.; Raftery, A. E.; and Volinsky, C. T. 1999. Bayesian model averaging: a tutorial. Statistical Science 382–401.
  9. 9.Holm, S. 1979. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 65–70.
  10. 10.Iman, R. L., and Davenport, J. M. 1980. Approximations of the critical region of the friedman statistic. Communications in Statistics-Theory and Methods 9(6):571–595.
  11. 11.Klemela, J. 2009. Multivariate histograms with datadependent partitions. Statistica Sinica 19(1):159.
  12. 12.Menon, A.; Jiang, X.; Vembu, S.; Elkan, C.; and OhnoMachado, L. 2012. Predicting accurate probabilities with a ranking loss. In Proceedings of the International Conference on Machine Learning, 703–710.
  13. 13.Niculescu-Mizil, A., and Caruana, R. 2005. Predicting good probabilities with supervised learning. In Proceedings of the International Conference on Machine Learning, 625–632.
  14. 14.Platt, J. C. 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers 10(3):61–74.
  15. 15.Russell, S., and Norvig, P. 1995. Artificial intelligence: A modern approach. Artificial Intelligence. Prentice-Hall, Englewood Cliffs 25.
  16. 16.Scott, C., and Nowak, R. 2003. Near-minimax optimal classification with dyadic classification trees. Advances in Neural Information Processing Systems 16.
  17. 17.Zadrozny, B., and Elkan, C. 2001. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In International Conference on Machine Learning, 609–616.
  18. 18.Zadrozny, B., and Elkan, C. 2002. Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 694–699.

Citation

MLA
Pakdaman Naeini, M., et al. “Obtaining Well Calibrated Probabilities Using Bayesian Binning”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 29, no. 1, 2015, https://doi.org/10.1609/aaai.v29i1.9602.
APA
Pakdaman Naeini, M., Cooper, G., & Hauskrecht, M. (2015). Obtaining Well Calibrated Probabilities Using Bayesian Binning. Proceedings of the AAAI Conference on Artificial Intelligence, 29(1). https://doi.org/10.1609/aaai.v29i1.9602
Chicago
Pakdaman Naeini, M., G. Cooper, and M. Hauskrecht. 2015. “Obtaining Well Calibrated Probabilities Using Bayesian Binning”. Proceedings of the AAAI Conference on Artificial Intelligence 29 (1). https://doi.org/10.1609/aaai.v29i1.9602.
Harvard
Pakdaman Naeini, M., Cooper, G. and Hauskrecht, M. (2015) “Obtaining Well Calibrated Probabilities Using Bayesian Binning”, Proceedings of the AAAI Conference on Artificial Intelligence, 29(1). Available at: https://doi.org/10.1609/aaai.v29i1.9602.
Vancouver
1. Pakdaman Naeini M, Cooper G, Hauskrecht M (2015) Obtaining Well Calibrated Probabilities Using Bayesian Binning. Proceedings of the AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v29i1.9602

BibTeX

@article{Pakdaman_Naeini_2015, title={Obtaining Well Calibrated Probabilities Using Bayesian Binning}, volume={29}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v29i1.9602}, DOI={10.1609/aaai.v29i1.9602}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Pakdaman Naeini, Mahdi and Cooper, Gregory and Hauskrecht, Milos}, year={2015}, month=Feb }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF