Dual Focal Loss for Calibration

Linwei TaoMinjing DongChang Xu

article2023ICML61 citations

Proposes Dual Focal Loss, a training objective that balances over-confidence and under-confidence in neural network calibration by maximizing the gap between the ground-truth logit and the highest-ranked competing logit.

Listen

Modern deep learning models achieve high accuracy across vision and text tasks, but they frequently output miscalibrated confidence scores. Standard training techniques cause models to be overconfident, whereas remedies like focal loss often swing too far in the opposite direction and make predictions underconfident. This calibration gap poses serious operational risks in safety-critical deployments where automated systems and human decision-makers rely directly on confidence scores to assess predictive uncertainty.

The article designs and evaluates a novel training objective called Dual Focal Loss to solve this dilemma. The primary goal is to train deep neural networks that are innately calibrated, balancing overconfidence and underconfidence without compromising underlying classification accuracy.

The researchers developed Dual Focal Loss by modifying standard loss formulations to consider two key model outputs simultaneously: the score for the true class and the highest score among all incorrect classes. By widening the margin between these two outputs, the objective prevents underconfidence while maintaining resistance to overconfidence. To establish credibility, the authors conducted mathematical risk analyses and carried out extensive empirical experiments across standard computer vision benchmarks (CIFAR-10, CIFAR-100, Tiny-ImageNet), text classification datasets (20 Newsgroups), and out-of-distribution robustness tests using multiple network architectures including ResNet and DenseNet.

The findings show that Dual Focal Loss consistently establishes state-of-the-art calibration performance. On complex benchmarks such as CIFAR-100, it reduced Expected Calibration Error to roughly 1.08% to 2.90% before any post-processing, substantially outperforming standard cross-entropy, label smoothing, and focal loss variants. Notably, the optimal temperature scaling parameter across all evaluated architectures was 1.0, demonstrating that models trained with this loss are innately calibrated directly out of training. Furthermore, the approach reduced calibration error without degrading classification accuracy, achieving competitive or lower test error rates than baseline methods across all tested benchmarks.

These results carry significant practical implications for operational risk and deployment workflows. Because models trained with Dual Focal Loss do not require post-hoc temperature adjustments or held-out calibration datasets, engineering pipelines can be simplified and deployment latency reduced. Downstream decision systems can more reliably use model confidence outputs for risk gating and human intervention thresholds.

Organizations deploying deep neural networks should evaluate replacing traditional cross-entropy or focal loss objectives with Dual Focal Loss in their standard training pipelines. The method can also be combined with sample-adaptive loss strategies, such as AdaFocal, for further calibration gains. Because the empirical evaluation focused primarily on standard benchmark image sets and a single text dataset, teams should validate performance on specialized domain data and conduct pilot trials prior to full-scale rollout in critical production environments.

  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Read this foundational study first to understand neural-network miscalibration, Expected Calibration Error, and temperature scaling—the calibration problem and evaluation framework Dual Focal Loss seeks to improve.
  • Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). Its Focal Loss is a direct predecessor to the focal-loss variants Dual Focal Loss compares against, so it clarifies the training objective the source modifies.
  • Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). This analysis of label smoothing’s effects on calibration gives useful context for one of the training-loss baselines against which Dual Focal Loss is evaluated.

No sufficiently relevant recommendations were found.

Cover for Dual Focal Loss for Calibration

Abstract

The use of deep neural networks in real-world applications require well-calibrated networks with confidence scores that accurately reflect the actual probability. However, it has been found that these networks often provide over-confident predictions, which leads to poor calibration. Recent efforts have sought to address this issue by focal loss to reduce over-confidence, but this approach can also lead to under-confident predictions. While different variants of focal loss have been explored, it is difficult to find a balance between over-confidence and under-confidence. In our work, we propose a new loss function by focusing on dual logits. Our method not only considers the ground truth logit, but also take into account the highest logit ranked after the ground truth logit. By maximizing the gap between these two logits, our proposed dual focal loss can achieve a better balance between over-confidence and under-confidence. We provide theoretical evidence to support our approach and demonstrate its effectiveness through evaluations on multiple models and datasets, where it achieves state-of-the-art performance. Code is available at https://github.com/Linwei94/DualFocalLoss

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Problem Formulation
  • 3.2. Dual Focal Loss for Calibration
  • 4. Theoretical Evidence
  • 4.1. Instance-wise Conditional Risk
  • 4.2. Over-confidence and Under-confidence
  • 5. Experiments
  • 5.1. Calibration Performance
  • 5.2. Comparison with AdaFocal Loss
  • 5.3. Balance between Regularization and Loss of Sample Hardness Information
  • 5.4. Ablation Study
  • 6. Conclusion
  • Acknowledgement
  • References
  • A. Proof of Theorem 1
  • B. Proof of Lemma 1
  • C. Dataset Description
  • D. Comparison Methods
  • E. Performance on Different Metrics and Robustness on Dataset Shift
  • F. Comparison with SOTA Method KDE-XE
  • G. Calibration Assessment with Metrics RBS
  • H. Robustness of DFL
  • I. Gamma Value Selection in DFL

Knowls

  1. Knowl 1 — Dual Focal Loss for Neural Network Calibration

    model/method

    Dual Focal Loss (DFL) is an objective function designed to calibrate classification deep neural networks by balancing over-confidence and under-confidence. For an input sample xx in a KK-class classification problem with one-hot ground-truth label vector y∈{0,1}Ky \in \{0, 1\}^K where ygt=1y_{gt} = 1, and predicted probability distribution q(x)=(q1(x),…,qK(x))∈[0,1]Kq(x) = (q_1(x), \dots, q_K(x)) \in [0, 1]^K output by a softmax function, Dual Focal Loss is defined as:

    LDFL(x,y)=−∑i=1Kyi(1−qi(x)+qj(x))γlog⁡qi(x)\mathcal{L}_{DFL}(x, y) = -\sum_{i=1}^K y_i \left(1 - q_i(x) + q_j(x)\right)^\gamma \log q_i(x)

    where γ>0\gamma > 0 is a focusing hyperparameter, and qj(x)q_j(x) is the dual logit defined as the highest predicted probability strictly ranked below the ground-truth probability qgt(x)q_{gt}(x):

    qj(x)=max⁡i{qi(x)∣qi(x)<qgt(x)}q_j(x) = \max_{i} \{q_i(x) \mid q_i(x) < q_{gt}(x)\}

    Unlike standard Focal Loss (which sets qj(x)=0q_j(x) = 0 and forces predicted probabilities toward high-entropy distributions, leading to under-confidence), DFL maximizes the margin between the ground-truth probability qgt(x)q_{gt}(x) and the runner-up non-ground-truth probability qj(x)q_j(x). This prevents excessive probability suppression while preserving the over-confidence reduction properties of focal loss.

  2. Knowl 2 — Classification Calibration and Order Preservation of Dual Focal Loss

    theoretical result

    Let η(x)=(η1(x),…,ηK(x))∈[0,1]K\eta(x) = (\eta_1(x), \dots, \eta_K(x)) \in [0, 1]^K denote the true class-posterior probability distribution for an input xx, satisfying ∑i=1Kηi(x)=1\sum_{i=1}^K \eta_i(x) = 1. The instance-wise conditional risk minimization of Dual Focal Loss is formulated as:

    min⁡q−∑i=1Kηi(x)(1−qi(x)+f(qi(x)))γlog⁡qi(x)subject to∑i=1Kqi(x)=1\min_{q} -\sum_{i=1}^K \eta_i(x) (1 - q_i(x) + f(q_i(x)))^\gamma \log q_i(x) \quad \text{subject to} \quad \sum_{i=1}^K q_i(x) = 1

    where f(qi(x))=qj(x)=max⁡k{qk(x)∣qk(x)<qgt(x)}f(q_i(x)) = q_j(x) = \max_{k} \{q_k(x) \mid q_k(x) < q_{gt}(x)\}.

    For any γ>0\gamma > 0, the loss LDFL\mathcal{L}_{DFL} is classification-calibrated and satisfies the strictly order-preserving property: for the optimal risk minimizer q∗(x)q^*(x),

    qa∗(x)<qb∗(x)  ⟺  ηa(x)<ηb(x)q^*_a(x) < q^*_b(x) \iff \eta_a(x) < \eta_b(x)

    As a consequence, the class with the maximal predicted probability under the risk minimizer is identical to the class with the maximal true posterior probability: argmax⁡iqi∗(x)=argmax⁡iηi(x)\operatorname{argmax}_i q^*_i(x) = \operatorname{argmax}_i \eta_i(x).

  3. Knowl 3 — Theoretical Reduction of the Under-Confidence Region in Dual Focal Loss

    theoretical result

    For a classification model with risk minimizer q∗(x)q^*(x) and true class-posterior η(x)\eta(x), an instance is η\eta-over-confident (ηOC\eta\text{OC}) if max⁡iqi∗(x)−max⁡iηi(x)>0\max_i q^*_i(x) - \max_i \eta_i(x) > 0 and η\eta-under-confident (ηUC\eta\text{UC}) if max⁡iqi∗(x)−max⁡iηi(x)<0\max_i q^*_i(x) - \max_i \eta_i(x) < 0.

    For index i≠ji \neq j and constant C=qj∗(x)C = q^*_j(x), define the auxiliary derivative function:

    ϕ(v)=(1−v+C)γ−γ(1−v+C)γ−1vlog⁡v\phi(v) = (1 - v + C)^\gamma - \gamma (1 - v + C)^{\gamma - 1} v \log v

    Function ϕ(v)\phi(v) has a unique maximum at vm∈(0,1)v^m \in (0, 1), strictly increasing on [0,vm)[0, v^m) and strictly decreasing on (vm,1](v^m, 1].

    In standard Focal Loss (FL), the under-confident regime ηUC\eta\text{UC} occurs on the interval [v′,1)[v', 1) where ϕ(v′)=(1+C)γ\phi(v') = (1+C)^\gamma. In Dual Focal Loss (DFL), due to the non-zero dual logit qj(x)q_j(x), the under-confident condition ϕ(qm∗(x))≤ϕ(qj∗(x))\phi(q^*_m(x)) \le \phi(q^*_j(x)) holds only on the interval [vuc,1)[v_{uc}, 1) where ϕ(vuc)=1\phi(v_{uc}) = 1.

    Because (1+C)γ≥1(1+C)^\gamma \ge 1 and ϕ\phi is decreasing on (vm,1](v^m, 1], vuc>v′v_{uc} > v', which strictly shrinks the under-confidence interval by ϕ−1(1)−ϕ−1((1+C)γ)\phi^{-1}(1) - \phi^{-1}((1+C)^\gamma) compared to standard Focal Loss while maintaining the identical over-confidence interval (0,vm](0, v^m].

  4. Knowl 4 — Expected Calibration Error Performance Before and After Temperature Scaling

    data/table

    Expected Calibration Error (ECE, reported in %, 15 equal-width bins) was evaluated across models and datasets before (Pre TT) and after (Post TT) post-hoc temperature scaling. The optimal temperature TT was selected via grid search on a validation split. Across all datasets and architectures, Dual Focal Loss (DFL) achieved an optimal post-hoc temperature of T=1.0T=1.0, showing that the model is innately calibrated without requiring temperature scaling.

    Dataset Model Weight Decay Brier Loss MMCE Label Smooth. Inv. Focal Focal Loss Dual Focal (Ours)
    Pre / Post(T) Pre / Post(T) Pre / Post(T) Pre / Post(T) Pre / Post(T) Pre / Post(T) Pre / Post(T)
    CIFAR-100 ResNet-50 17.52 / 3.42(2.1) 6.52 / 3.64(1.1) 15.32 / 2.38(1.8) 7.81 / 4.01(1.1) 17.88 / 2.98(2.3) 4.50 / 2.00(1.1) 1.08 / 1.08(1.0)
    ResNet-110 19.05 / 4.43(2.3) 7.88 / 4.65(1.2) 19.14 / 3.86(2.3) 11.02 / 5.89(1.1) 19.47 / 4.52(2.6) 8.56 / 4.12(1.2) 2.90 / 2.90(1.0)
    Wide-ResNet 15.33 / 2.88(2.2) 4.31 / 2.70(1.1) 13.17 / 4.37(1.9) 4.84 / 4.84(1.0) 16.90 / 2.28(2.5) 3.03 / 1.64(1.1) 1.79 / 1.79(1.0)
    DenseNet-121 20.98 / 4.27(2.3) 5.17 / 2.29(1.1) 19.13 / 3.06(2.1) 12.89 / 7.52(1.2) 19.42 / 2.82(2.3) 3.73 / 1.31(1.1) 1.81 / 1.81(1.0)
    CIFAR-10 ResNet-50 4.35 / 1.35(2.5) 1.82 / 1.08(1.1) 4.56 / 1.19(2.6) 2.96 / 1.67(0.9) 4.41 / 1.32(2.8) 1.55 / 0.95(1.1) 0.46 / 0.46(1.0)
    ResNet-110 4.41 / 1.09(2.8) 2.56 / 1.25(1.2) 5.08 / 1.42(2.8) 2.09 / 2.09(1.0) 4.34 / 0.89(2.9) 1.87 / 1.07(1.1) 0.98 / 0.98(1.0)
    Wide-ResNet 3.23 / 0.92(2.2) 1.25 / 1.25(1.0) 3.29 / 0.86(2.2) 4.26 / 1.84(0.8) 3.68 / 0.99(2.7) 1.56 / 0.84(0.9) 0.81 / 0.81(1.0)
    DenseNet-121 4.52 / 1.31(2.4) 1.53 / 1.53(1.0) 5.10 / 1.61(2.5) 1.88 / 1.82(0.9) 4.61 / 1.07(2.8) 1.22 / 1.22(1.0) 0.57 / 0.57(1.0)
    Tiny-ImageNet ResNet-50 15.32 / 5.48(1.4) 4.44 / 4.13(0.9) 13.01 / 5.55(1.3) 15.23 / 6.51(0.7) 11.51 / 6.71(1.3) 1.76 / 1.76(1.0) 1.50 / 1.50(1.0)
    20 Newsgroups GlobalPool CNN 17.92 / 2.39(2.3) 15.48 / 6.78(2.1) 13.58 / 3.22(1.9) 4.79 / 2.54(1.1) 16.72 / 2.51(2.1) 6.92 / 2.19(1.1) 1.79 / 1.79(1.0)
  5. Knowl 5 — Classification Error Across Calibration Objectives

    data/table

    Classification error (% on test set) of neural network architectures trained under various calibration and regularization loss functions. The results indicate that Dual Focal Loss improves calibration without compromising classification accuracy, achieving equal or lower test error than standard Cross-Entropy (Weight Decay) and Focal Loss.

    Dataset Model Weight Decay Brier Loss MMCE Label Smooth. Inv. Focal Focal Loss Dual Focal (Ours)
    CIFAR-100 ResNet-50 23.30 23.39 23.20 23.43 22.23 23.22 22.67
    ResNet-110 22.73 25.10 23.07 23.43 22.43 22.51 22.59
    Wide-ResNet-26-10 20.70 20.59 20.73 21.19 20.85 20.11 19.91
    DenseNet-121 24.52 23.75 24.00 24.05 24.55 22.67 22.40
    CIFAR-10 ResNet-50 4.95 5.00 4.99 5.29 4.80 4.98 5.17
    ResNet-110 4.89 5.48 5.40 5.52 4.66 5.42 5.02
    Wide-ResNet-26-10 3.86 4.08 3.91 4.20 4.10 4.01 3.96
    DenseNet-121 5.00 5.11 5.41 5.09 4.82 5.46 5.43
    Tiny-ImageNet ResNet-50 49.81 53.20 51.31 47.12 55.19 49.06 48.63
    20 Newsgroups GlobalPool CNN 26.68 27.23 27.06 26.03 29.26 27.98 28.73
  6. Knowl 6 — AdaDualFocal Combination of Dual Focal Loss and Sample-Adaptive Gamma

    model/method

    Dual Focal Loss can be integrated with AdaFocal, which dynamically selects sample-specific focusing parameters γ\gamma according to validation set calibration state. Combining AdaFocal's dynamic γ\gamma schedule with Dual Focal Loss's dual-logit margin optimization produces AdaDualFocal.

    Expected Calibration Error (ECE, %) on CIFAR-10 and CIFAR-100 demonstrates that AdaDualFocal yields superior calibration performance compared to either method used individually:

    Dataset Model FLSD-53 AdaFocal DualFocal AdaDualFocal
    CIFAR-10 ResNet-50 1.55 0.66 0.46 0.43
    ResNet-110 1.87 0.71 0.98 0.69
    DenseNet-121 1.56 0.64 0.81 0.50
    Wide-ResNet-26-10 1.22 0.62 0.57 0.54
    CIFAR-100 ResNet-50 4.50 1.36 1.08 1.07
    ResNet-110 8.56 1.40 2.90 1.14
    DenseNet-121 3.03 1.95 1.79 1.80
    Wide-ResNet-26-10 3.73 1.73 1.81 1.63
  7. Knowl 7 — Ablation on Dual Logit Selection in Dual Focal Loss

    data/table

    An ablation study evaluated the choice of the auxiliary logit qj(x)q_j(x) in the Dual Focal Loss formulation L=−∑iyi(1−qi+qj)γlog⁡qi\mathcal{L} = -\sum_i y_i (1 - q_i + q_j)^\gamma \log q_i on ResNet-50 trained on CIFAR-10. Selecting the largest logit strictly smaller than the ground-truth logit qgt(x)q_{gt}(x) yielded the lowest Expected Calibration Error (ECE).

    Method / Auxiliary Logit qjq_j ECE (%) AdaECE (%)
    Focal loss (+0.0+0.0) 1.55 1.56
    Fixed dual logit (+0.1+0.1) 1.15 1.52
    Fixed dual logit (+0.2+0.2) 2.10 2.17
    DFL: 2nd largest logit after qgtq_{gt} 1.12 1.46
    DFL: 3rd largest logit after qgtq_{gt} 1.34 1.35
    DFL: mean(1st + 2nd) largest after qgtq_{gt} 0.80 0.68
    DFL: mean(1st + 2nd + 3rd) largest after qgtq_{gt} 0.61 0.50
    DFL: mean(all logits lower than qgtq_{gt}) 0.52 0.38
    DFL (Ours): largest lower than qgtq_{gt} 0.46 0.66
  8. Knowl 8 — Multiclass Calibration Evaluation on Classwise-ECE, Adaptive-ECE, and MCE

    empirical result

    Dual Focal Loss was evaluated across multiple calibration metrics to test whether its calibration benefits extend across all probability bins and non-predicted classes:

    1. Adaptive-ECE (AdaECE): Partitions predictions into bins containing equal numbers of samples (∣Bi∣=∣Bj∣|B_i| = |B_j|). On CIFAR-100, DFL achieved pre-temperature scaling AdaECE of 1.23% (ResNet-50), 3.16% (ResNet-110), 2.03% (Wide-ResNet), and 1.63% (DenseNet-121), outperforming standard Focal Loss (4.50%, 8.55%, 2.75%, 3.55%) and Cross-Entropy with weight decay (17.52%, 19.05%, 15.33%, 20.98%).

    2. Classwise-ECE: Averages the calibration error computed across each of the KK classes individually: Classwise-ECE=1K∑i=1B∑j=1K∣Bi,j∣N∣Ii,j−Ci,j∣\text{Classwise-ECE} = \frac{1}{K}\sum_{i=1}^B \sum_{j=1}^K \frac{|B_{i,j}|}{N} |I_{i,j} - C_{i,j}|. DFL achieved the lowest pre-temperature scaling Classwise-ECE across all tested configurations on CIFAR-10 (0.34%–0.35%) and CIFAR-100 (0.19%–0.21%), confirming accurate probability estimation for both the top class and non-maximal classes.

    3. Maximum Calibration Error (MCE): Measures the maximum absolute bin calibration error max⁡m∣Am−Cm∣\max_m |A_m - C_m|. On CIFAR-100, DFL achieved pre-temperature scaling MCE of 5.29% on ResNet-50 and 8.10% on ResNet-110, compared to 16.12% and 22.57% for Focal Loss and 44.34% and 55.92% for Cross-Entropy.

  9. Knowl 9 — Out-of-Distribution Calibration Robustness Under Dataset Shift

    data/table

    Models trained on CIFAR-10 (in-distribution) were evaluated on out-of-distribution (OoD) datasets under distribution shift: Street View House Numbers (SVHN) and CIFAR-10-C (corrupted with Gaussian Noise at severity level 5). Calibration quality under shift was measured using AUROC (%):

    Dataset Shift Model Weight Decay Brier Loss MMCE Label Smooth. Inv. Focal Focal Loss Dual Focal (Ours)
    Pre / Post Pre / Post Pre / Post Pre / Post Pre / Post Pre / Post Pre / Post
    CIFAR-10 →\rightarrow SVHN ResNet-50 94.32 / 94.56 93.59 / 93.72 85.17 / 64.75 78.88 / 78.89 92.41 / 92.60 92.48 / 92.79 94.34 / 94.34
    DenseNet-121 84.43 / 81.57 94.65 / 94.66 85.88 / 84.87 78.79 / 78.94 74.08 / 67.78 89.59 / 89.59 93.78 / 93.78
    CIFAR-10 →\rightarrow CIFAR-10-C ResNet-50 86.23 / 86.03 90.21 / 90.13 89.97 / 90.11 72.01 / 72.02 77.81 / 74.74 89.45 / 89.56 87.93 / 87.93
    DenseNet-121 87.61 / 86.41 87.38 / 87.38 84.90 / 84.88 73.67 / 73.80 76.72 / 72.51 89.47 / 89.47 89.56 / 89.56
  10. Knowl 10 — Comparison of Dual Focal Loss with Kernel Density Calibration Estimator KDE-XE

    data/table

    Dual Focal Loss was compared against the Dirichlet kernel density calibration method KDE-XE across four network architectures on CIFAR-10 and CIFAR-100 in terms of standard ECE (%) and kernel calibration error ECEKDE\text{ECE}_{KDE}:

    Metric Model Dataset Cross Entropy KDE-XE Dual Focal (Ours)
    ECE ResNet-110 CIFAR-10 3.890 3.093 0.400
    ECE ResNet-110sd CIFAR-10 3.555 2.778 1.510
    ECE ResNet-110 CIFAR-100 12.769 8.969 2.430
    ECE ResNet-110sd CIFAR-100 11.175 7.828 2.600
    ECE Wide-ResNet CIFAR-100 7.279 3.703 1.040
    ECE DenseNet-40 CIFAR-100 9.196 8.016 2.340
    ECEKDE\text{ECE}_{KDE} ResNet-110 CIFAR-10 0.133 0.126 0.153
    ECEKDE\text{ECE}_{KDE} DenseNet-40 CIFAR-10 0.104 0.098 0.121

    Dual Focal Loss outperforms KDE-XE by large margins on ECE across all models, though KDE-XE achieves lower ECEKDE\text{ECE}_{KDE} because it optimizes a canonical calibration surrogate directly in its training loss.

Coverage note — Omitted the supplementary gamma sweep ablation table (Table 13) and the raw RBS table (Table 11), as they represent minor hyperparameter sweeps and supplementary metrics that do not alter the core findings already captured in the main knowls.

References

  1. 1.Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101(473):138–156, 2006.
  2. 2.Bohdal, O., Yang, Y., and Hospedales, T. Meta-calibration: Meta-learning of model calibration using differentiable expected calibration error. arXiv preprint arXiv:2106.09613, 2021.
  3. 3.Brier, G. W., Kornblith, S., and Hinton, G. E. Verification of forecasts expressed in terms of probability. Monthly weather review, 78(1):1–3, 1950.
  4. 4.Charoenphakdee, N., Vongkulbhisal, J., Chairatanakul, N., and Sugiyama, M. On focal loss for class-posterior probability estimation: A theoretical perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5202–5211, 2021.
  5. 5.Cheng, B., Schwing, A., and Kirillov, A. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021.
  6. 6.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  7. 7.Ding, Y., Liu, J., Xiong, J., and Shi, Y. Revisiting the evaluation of uncertainty estimation and its application to explore model complexity-uncertainty trade-off. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 4–5, 2020.
  8. 8.Ghosh, A., Schaaf, T., and Gormley, M. R. Adafocal: Calibration-aware adaptive focal loss, 2022. URL https://openreview.net/forum?id=CoMOKHYWf2.
  9. 9.Goodfellow, I. J., Bulatov, Y., Ibarz, J., Arnoud, S., and Shet, V. Multi-digit number recognition from street view imagery using deep convolutional neural networks. arXiv preprint arXiv:1312.6082, 2013.
  10. 10.Gruber, S. and Buettner, F. Better uncertainty calibration via proper scores for classification and beyond. Advances in Neural Information Processing Systems, 35:8618–8632, 2022.
  11. 11.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International conference on machine learning, pp. 1321–1330. PMLR, 2017.
  12. 12.Gupta, K., Rahimi, A., Ajanthan, T., Mensink, T., Sminchisescu, C., and Hartley, R. Calibration of neural networks using splines. arXiv preprint arXiv:2006.12800, 2020.
  13. 13.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  14. 14.He, K., Gkioxari, G., Dollar, P., and Girshick, R. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pp. 2961–2969, 2017.
  15. 15.Hendrycks, D. and Dietterich, T. G. Benchmarking neural network robustness to common corruptions and surface variations. arXiv preprint arXiv:1807.01697, 2018.
  16. 16.Hendrycks, D., Mu, N., Cubuk, E. D., Zoph, B., Gilmer, J., and Lakshminarayanan, B. Augmix: A simple data processing method to improve robustness and uncertainty. arXiv preprint arXiv:1912.02781, 2019.
  17. 17.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708, 2017.
  18. 18.Hui, L. and Belkin, M. Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks. arXiv preprint arXiv:2006.07322, 2020.
  19. 19.Joachims, T. A probabilistic analysis of the rocchio algorithm with tfidf for text categorization. Technical report, Carnegie-mellon univ pittsburgh pa dept of computer science, 1996.
  20. 20.Karandikar, A., Cain, N., Tran, D., Lakshminarayanan, B., Shlens, J., Mozer, M. C., and Roelofs, B. Soft calibration objectives for neural networks. Advances in Neural Information Processing Systems, 34:29768–29779, 2021.
  21. 21.Krishnan, R. and Tickoo, O. Improving model calibration with accuracy versus uncertainty optimization. Advances in Neural Information Processing Systems, 33:18237–18248, 2020.
  22. 22.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  23. 23.Kull, M., Silva Filho, T., and Flach, P. Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers. In Artificial Intelligence and Statistics, pp. 623–631. PMLR, 2017.
  24. 24.Kull, M., Perello Nieto, M., Kangsepp, M., Silva Filho, T., Song, H., and Flach, P. Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. Advances in neural information processing systems, 32, 2019.
  25. 25.Kumar, A., Sarawagi, S., and Jain, U. Trainable calibration measures for neural networks from kernel mean embeddings. In Dy, J. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 2805–2814. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/kumar18a.html.
  26. 26.Kumar, A., Liang, P. S., and Ma, T. Verified uncertainty calibration. Advances in Neural Information Processing Systems, 32, 2019.
  27. 27.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30, 2017.
  28. 28.Lin, M., Chen, Q., and Yan, S. Network in network. arXiv preprint arXiv:1312.4400, 2013.
  29. 29.Long, J., Shelhamer, E., and Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3431–3440, 2015.
  30. 30.Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P., and Dokania, P. Calibrating deep neural networks using focal loss. Advances in Neural Information Processing Systems, 33:15288–15299, 2020.
  31. 31.Muller, R., Kornblith, S., and Hinton, G. E. When does label smoothing help? Advances in neural information processing systems, 32, 2019.
  32. 32.Naeini, M. P., Cooper, G., and Hauskrecht, M. Obtaining well calibrated probabilities using bayesian binning. In Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015.
  33. 33.Nguyen, K. and O’Connor, B. Posterior calibration and exploratory analysis for natural language processing models. arXiv preprint arXiv:1508.05154, 2015.
  34. 34.Niculescu-Mizil, A. and Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning, pp. 625–632, 2005.
  35. 35.Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., and Tran, D. Measuring calibration in deep learning. In CVPR Workshops, volume 2, 2019.
  36. 36.Platt, J., Cain, N., Tran, D., Lakshminarayanan, B., Shlens, J., Mozer, M. C., and Roelofs, B. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers, 10(3):61–74, 1999.
  37. 37.Popordanoska, T., Sayer, R., and Blaschko, M. A consistent and differentiable lp canonical calibration error estimator. Advances in Neural Information Processing Systems, 35:7933–7946, 2022.
  38. 38.Rahaman, R. et al. Uncertainty quantification and deep ensembles. Advances in Neural Information Processing Systems, 34:20063–20075, 2021.
  39. 39.Roelofs, R., Cain, N., Shlens, J., and Mozer, M. C. Mitigating bias in calibration error estimation. In International Conference on Artificial Intelligence and Statistics, pp. 4036–4054. PMLR, 2022.
  40. 40.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818–2826, 2016.
  41. 41.Tao, L., Dong, M., Liu, D., Sun, C., and Xu, C. Calibrating a deep neural network with its predecessors. arXiv preprint arXiv:2302.06245, 2023.
  42. 42.Tewari, A. and Bartlett, P. L. On the consistency of multiclass classification methods. Journal of Machine Learning Research, 8(5), 2007.
  43. 43.Thulasidasan, S., Chennupati, G., Bilmes, J. A., Bhattacharya, T., and Michalak, S. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. Advances in Neural Information Processing Systems, 32, 2019.
  44. 44.Tian, Z., Shen, C., Chen, H., and He, T. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 9627–9636, 2019.
  45. 45.Wang, D.-B., Feng, L., and Zhang, M.-L. Rethinking calibration of deep neural networks: Do not be afraid of overconfidence. Advances in Neural Information Processing Systems, 34:11809–11820, 2021.
  46. 46.Wen, Y., Tran, D., and Ba, J. Batchensemble: an alternative approach to efficient ensemble and lifelong learning. arXiv preprint arXiv:2002.06715, 2020.
  47. 47.Zadrozny, B. and Elkan, C. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In Icml, volume 1, pp. 609–616. Citeseer, 2001.
  48. 48.Zadrozny, B. and Elkan, C. Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 694–699, 2002.
  49. 49.Zagoruyko, S. and Komodakis, N. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  50. 50.Zhang, J., Kailkhura, B., and Han, T. Y.-J. Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning. In International conference on machine learning, pp. 11117–11128. PMLR, 2020.
  51. 51.Zhang, T. Statistical analysis of some multi-category large margin classification methods. Journal of Machine Learning Research, 5(Oct):1225–1251, 2004.

Citation

MLA
Tao, L., et al. “Dual Focal Loss for Calibration”. International Conference on Machine Learning, vol. 202, 2023, pp. 33833–49, https://proceedings.mlr.press/v202/tao23a.html.
APA
Tao, L., Dong, M., & Xu, C. (2023). Dual Focal Loss for Calibration. International Conference on Machine Learning, 202, 33833–33849. https://proceedings.mlr.press/v202/tao23a.html
Chicago
Tao, L., M. Dong, and C. Xu. 2023. “Dual Focal Loss for Calibration”. International Conference on Machine Learning 202: 33833–49. https://proceedings.mlr.press/v202/tao23a.html.
Harvard
Tao, L., Dong, M. and Xu, C. (2023) “Dual Focal Loss for Calibration”, International Conference on Machine Learning. PMLR, pp. 33833–33849. Available at: https://proceedings.mlr.press/v202/tao23a.html.
Vancouver
1. Tao L, Dong M, Xu C (2023) Dual Focal Loss for Calibration. In: International Conference on Machine Learning. PMLR, pp 33833–33849

BibTeX

@InProceedings{pmlr-v202-tao23a,
  title = 	 {Dual Focal Loss for Calibration},
  author =       {Tao, Linwei and Dong, Minjing and Xu, Chang},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {33833--33849},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/tao23a/tao23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/tao23a.html},
  abstract = 	 {The use of deep neural networks in real-world applications require well-calibrated networks with confidence scores that accurately reflect the actual probability. However, it has been found that these networks often provide over-confident predictions, which leads to poor calibration. Recent efforts have sought to address this issue by focal loss to reduce over-confidence, but this approach can also lead to under-confident predictions. While different variants of focal loss have been explored, it is difficult to find a balance between over-confidence and under-confidence. In our work, we propose a new loss function by focusing on dual logits. Our method not only considers the ground truth logit, but also take into account the highest logit ranked after the ground truth logit. By maximizing the gap between these two logits, our proposed dual focal loss can achieve a better balance between over-confidence and under-confidence. We provide theoretical evidence to support our approach and demonstrate its effectiveness through evaluations on multiple models and datasets, where it achieves state-of-the-art performance. Code is available at https://github.com/Linwei94/DualFocalLoss}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/