Mitigating Neural Network Overconfidence with Logit Normalization

Hongxin WeiRenchunzi XieHao ChengLei FengBo AnYixuan Li

article2022ICML381 citationsOutstanding Paper Honorable Mention

Proposes Logit Normalization, a straightforward modification to standard cross-entropy loss that enforces a constant logit norm during training to mitigate overconfidence and substantially improve out-of-distribution detection.

Listen

Deep neural networks deployed in real-world applications frequently encounter out-of-distribution inputs—unfamiliar data from outside their training distribution. When encountering these novel inputs, conventional models routinely assign them abnormally high confidence scores instead of flagging them as unknown. This systemic overconfidence creates serious safety and operational risks, as deployed automated systems cannot reliably detect when they are operating outside their competency domain.

The article demonstrates that standard cross-entropy loss intrinsically drives this failure by continually inflating the magnitude of pre-softmax output vectors (logits) during training, and it evaluates a simple modification called Logit Normalization (LogitNorm) to resolve the issue. By constraining logit vectors to a constant length during training, LogitNorm optimizes only output directions and decouples vector magnitude from the optimization process, yielding conservative confidence scores on unfamiliar inputs.

The authors conducted extensive empirical evaluations using standard benchmark datasets (CIFAR-10 and CIFAR-100 as known data alongside six distinct unknown image test sets) across diverse model architectures, including Wide Residual Networks, ResNet-34, and DenseNet. The study assessed out-of-distribution detection capabilities, baseline classification accuracy, calibration error, and compatibility with various post-hoc detection algorithms.

The findings establish that LogitNorm substantially outperforms conventional training methods. First, it reduced the false positive rate on out-of-distribution samples by an average of 33.87 percentage points on standard benchmarks when holding the true positive rate at 95%, with specific test reductions exceeding 42 percentage points. Second, the method preserved full classification accuracy on known data, matching standard cross-entropy performance across all tested network architectures. Third, LogitNorm enhanced the effectiveness of downstream detection algorithms (such as ODIN, energy scores, and gradient norms) and yielded superior probability calibration when combined with post-hoc temperature scaling. Finally, the authors found that directly penalizing vector norms via standard regularization failed to achieve these benefits, confirming the unique necessity of strict vector normalization.

These results indicate that organizations can significantly improve model reliability and reduce the operational risk of deploying machine learning systems in open environments without sacrificing predictive accuracy. Because LogitNorm requires only a straightforward change to the training loss function and operates entirely on known training data without requiring exposure to real outlier data or complex training pipelines, it provides a highly cost-effective enhancement for production AI safety.

Technical leaders and practitioners should consider adopting LogitNorm as a standard training objective for classification systems operating in open-world settings where unrecognized inputs are expected. Teams implementing the approach should validate the temperature parameter using a holdout validation set, selecting conservative values to avoid optimization issues highlighted by the lower-bound analysis.

While the empirical validation is robust across standard computer vision benchmarks, the evaluations rely on 32x32 image datasets, and hyperparameter tuning currently requires training multiple candidate models, which introduces modest computational overhead. Confidence in the core mechanism is high, though organizations should conduct targeted pilot validation on their specific domain data and architectures prior to wide-scale deployment.

arXiv: 2205.09310
Cover for Mitigating Neural Network Overconfidence with Logit Normalization

Abstract

Detecting out-of-distribution inputs is critical for safe deployment of machine learning models in the real world. However, neural networks are known to suffer from the overconfidence issue, where they produce abnormally high confidence for both in- and out-of-distribution inputs. In this work, we show that this issue can be mitigated through Logit Normalization (LogitNorm) -- a simple fix to the cross-entropy loss -- by enforcing a constant vector norm on the logits in training. Our method is motivated by the analysis that the norm of the logit keeps increasing during training, leading to overconfident output. Our key idea behind LogitNorm is thus to decouple the influence of output's norm during network optimization. Trained with LogitNorm, neural networks produce highly distinguishable confidence scores between in- and out-of-distribution data. Extensive experiments demonstrate the superiority of LogitNorm, reducing the average FPR95 by up to 42.30% on common benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Preliminaries: Out-of-distribution Detection
  • 3 Method: Logit Normalization
  • 3.1 Motivation
  • 3.2 Method
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Results
  • 5 Discussion
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Proof of Proposition
  • B Proof of Proposition
  • C Proof of Proposition
  • D Descriptions of OOD Datasets
  • E Future Work
  • F Detailed Experimental Results

Knowls

  1. Knowl 1 — Logit Normalization (LogitNorm) Training Objective

    model/method

    Logit Normalization (LogitNorm) is a training objective for multi-class classification designed to mitigate overconfidence on out-of-distribution (OOD) inputs by decoupling the vector norm of the model's pre-softmax logits from network optimization. Given an input x∈Xx \in \mathcal{X}, a neural network parameterized by θ∈Rp\theta \in \mathbb{R}^p outputs a logit vector f(x;θ)∈Rkf(x; \theta) \in \mathbb{R}^k over kk classes. Under standard cross-entropy loss, the magnitude ∥f(x;θ)∥2\|f(x; \theta)\|_2 grows during training even after samples are correctly classified, which drives softmax confidence scores towards 1.01.0 for both in-distribution and out-of-distribution data.

    LogitNorm enforces a constant logit vector norm during training by applying the cross-entropy loss to the L2L_2-normalized logit vector scaled by a temperature hyperparameter τ∈R+\tau \in \mathbb{R}^+:

    Llogit_norm(f(x;θ),y)=−log⁡exp⁡(fy(x;θ)τ(∥f(x;θ)∥2+ϵ))∑i=1kexp⁡(fi(x;θ)τ(∥f(x;θ)∥2+ϵ))\mathcal{L}_{\text{logit\_norm}}(f(x; \theta), y) = -\log \frac{\exp\left(\frac{f_y(x; \theta)}{\tau (\|f(x; \theta)\|_2 + \epsilon)}\right)}{\sum_{i=1}^k \exp\left(\frac{f_i(x; \theta)}{\tau (\|f(x; \theta)\|_2 + \epsilon)}\right)}

    where y∈{1,…,k}y \in \{1, \dots, k\} is the ground-truth class label, ∥⋅∥2\|\cdot\|_2 denotes the Euclidean vector norm, and ϵ>0\epsilon > 0 (set to 10−710^{-7} in practice) is a small constant for numerical stability.

    Minimizing Llogit_norm\mathcal{L}_{\text{logit\_norm}} forces optimization to adjust only the directional unit vector f^(x;θ)=f(x;θ)∥f(x;θ)∥2\hat{f}(x; \theta) = \frac{f(x; \theta)}{\|f(x; \theta)\|_2} to match the one-hot target, holding the effective logit norm at the constant 1/τ1/\tau. Equivalently, this loss corresponds to standard cross-entropy parameterized by an input-dependent adaptive temperature τ(x)=τ∥f(x;θ)∥2\tau(x) = \tau \|f(x; \theta)\|_2.

  2. Knowl 2 — Theoretical Characterization of Logit Magnitude and Softmax Overconfidence

    theoretical result

    Let f=(f1,…,fk)⊤∈Rkf = (f_1, \dots, f_k)^\top \in \mathbb{R}^k denote a pre-softmax logit vector produced by a neural network for a kk-class classification task, which can be decomposed into its Euclidean norm and a unit directional vector as f=∥f∥2⋅f^f = \|f\|_2 \cdot \hat{f}, where ∥f∥2=∑i=1kfi2\|f\|_2 = \sqrt{\sum_{i=1}^k f_i^2} and ∥f^∥2=1\|\hat{f}\|_2 = 1. Let σ:Rk→(0,1)k\sigma: \mathbb{R}^k \to (0, 1)^k be the standard softmax function defined elementwise by σi(f)=efi∑j=1kefj\sigma_i(f) = \frac{e^{f_i}}{\sum_{j=1}^k e^{f_j}}.

    Two properties establish the relationship between logit magnitude, predicted label, and classification confidence:

    1. Prediction Invariance under Positive Scaling: For any scalar scale factor s>1s > 1, if the predicted class is c=arg⁡max⁡i∈{1,…,k}fic = \arg\max_{i \in \{1, \dots, k\}} f_i, then:
    arg⁡max⁡i∈{1,…,k}(sfi)=c\arg\max_{i \in \{1, \dots, k\}} (s f_i) = c

    Scaling the magnitude of the logit vector leaves the discrete prediction unchanged.

    1. Monotonic Confidence Increase under Norm Growth: For any scalar s>1s > 1, if c=arg⁡max⁡i∈{1,…,k}fic = \arg\max_{i \in \{1, \dots, k\}} f_i, the softmax probability for the winning class cc satisfies:
    σc(sf)≥σc(f)\sigma_c(s f) \ge \sigma_c(f)

    with strict inequality holding whenever there exists at least one class index j≠cj \neq c such that fj<fcf_j < f_c.

    Consequently, under the standard cross-entropy loss LCE(f,y)=−log⁡σy(f)\mathcal{L}_{\text{CE}}(f, y) = -\log \sigma_y(f), when a training example (x,y)(x, y) is correctly classified (y=cy = c), gradient descent continues to decrease the loss by inflating ∥f∥2\|f\|_2 without modifying the directional classification boundary, creating overconfident softmax predictions for both in-distribution and distant out-of-distribution inputs.

  3. Knowl 3 — Lower Bound of the LogitNorm Loss

    theoretical result

    For any input x∈Xx \in \mathcal{X}, positive temperature parameter τ∈R+\tau \in \mathbb{R}^+, and label y∈{1,…,k}y \in \{1, \dots, k\} in a kk-class classification setting, the per-sample LogitNorm loss:

    Llogit_norm(f(x;θ),y)=−log⁡exp⁡(fyτ∥f∥2)∑i=1kexp⁡(fiτ∥f∥2)\mathcal{L}_{\text{logit\_norm}}(f(x; \theta), y) = -\log \frac{\exp\left(\frac{f_y}{\tau \|f\|_2}\right)}{\sum_{i=1}^k \exp\left(\frac{f_i}{\tau \|f\|_2}\right)}

    is strictly bounded from below by:

    Llogit_norm(f(x;θ),y)≥log⁡(1+(k−1)e−2/τ)\mathcal{L}_{\text{logit\_norm}}(f(x; \theta), y) \ge \log\left(1 + (k - 1) e^{-2/\tau}\right)

    Because the normalized and scaled vector f~=fτ∥f∥2\tilde{f} = \frac{f}{\tau \|f\|_2} has norm ∥f~∥2=1/τ\|\tilde{f}\|_2 = 1/\tau, each component is bounded by −1/τ≤f~i≤1/τ-1/\tau \le \tilde{f}_i \le 1/\tau. The maximum softmax probability for the target class yy occurs when f~y=1/τ\tilde{f}_y = 1/\tau and f~i=−1/τ\tilde{f}_i = -1/\tau for all i≠yi \neq y, giving σy(f~)≤11+(k−1)e−2/τ\sigma_y(\tilde{f}) \le \frac{1}{1 + (k-1)e^{-2/\tau}}.

    This bound implies that the achievable minimum loss increases monotonically with τ\tau. For example, when k=10k = 10 and τ=1\tau = 1, the lower bound is approximately 0.79660.7966, creating significant optimization resistance. Consequently, effective network optimization with LogitNorm requires choosing a small temperature τ<1\tau < 1 (e.g., τ∈[0.01,0.05]\tau \in [0.01, 0.05]).

  4. Knowl 4 — OOD Detection Performance of LogitNorm using Maximum Softmax Probability

    empirical result

    Training deep neural networks with LogitNorm loss instead of standard cross-entropy (CE) loss substantially improves out-of-distribution (OOD) detection when evaluated using Maximum Softmax Probability (MSP, max⁡iσi(f(x))\max_i \sigma_i(f(x))) as the scoring function.

    Evaluating a Wide Residual Network (WRN-40-2) trained on in-distribution (ID) CIFAR-10 and CIFAR-100 across six OOD test datasets (Textures, SVHN, LSUN-Crop, LSUN-Resize, iSUN, and Places365) yields the following metrics: false positive rate of OOD samples at 95% true positive rate of ID data (FPR95 ↓\downarrow), area under the receiver operating characteristic curve (AUROC ↑\uparrow), and area under the precision-recall curve (AUPR ↑\uparrow):

    ID Dataset CIFAR-10 CIFAR-100
    OOD Dataset FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow) FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow)
    Cross-Entropy / LogitNorm
    Texture 64.13 / 28.64 86.29 / 94.28 96.28 / 98.63 84.11 / 70.67 74.05 / 78.65 93.45 / 93.66
    SVHN 50.33 / 8.03 93.48 / 98.47 98.70 / 99.68 79.09 / 45.98 78.62 / 92.48 94.99 / 98.45
    LSUN-C 33.34 / 2.37 95.29 / 99.42 99.04 / 99.88 67.94 / 13.93 83.60 / 97.56 96.32 / 99.48
    LSUN-R 42.52 / 10.93 93.74 / 97.87 98.66 / 99.59 82.21 / 68.68 69.45 / 84.77 91.45 / 96.52
    iSUN 46.56 / 12.28 93.13 / 97.73 98.55 / 99.56 84.50 / 71.47 69.29 / 83.79 91.44 / 96.24
    Places365 60.23 / 31.64 87.36 / 93.66 96.62 / 98.49 81.09 / 80.20 75.61 / 77.14 93.74 / 94.16
    Average 49.52 / 15.65 91.55 / 96.91 97.98 / 99.31 79.82 / 58.49 75.10 / 85.73 93.57 / 96.42

    On CIFAR-10, LogitNorm reduces the average FPR95 from 49.52% to 15.65% (an absolute reduction of 33.87%), with the largest single-dataset reduction occurring on SVHN where FPR95 drops from 50.33% to 8.03% (a 42.30% absolute improvement).

  5. Knowl 5 — Boosting Downstream Post-Hoc OOD Scoring Functions via LogitNorm

    empirical result

    LogitNorm training enhances downstream post-hoc OOD scoring functions beyond the standard Maximum Softmax Probability (MSP). Evaluated scoring functions include ODIN (using temperature T=1000T=1000 and input perturbation ϵ=0.0014\epsilon=0.0014), Energy score (E(x;f)=−Tlog⁡∑i=1kefi(x)/TE(x; f) = -T \log \sum_{i=1}^k e^{f_i(x)/T} with T=0.1T=0.1 for CIFAR-10 and T=0.01T=0.01 for CIFAR-100), and GradNorm (gradients backpropagated from the Kullback-Leibler divergence space).

    Average OOD detection performance across six OOD test benchmarks (Textures, SVHN, LSUN-C, LSUN-R, iSUN, Places365) using WRN-40-2:

    ID Dataset CIFAR-10 CIFAR-100
    Score Function FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow) FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow)
    Cross-Entropy / LogitNorm
    Softmax (MSP) 49.52 / 15.65 91.55 / 96.91 97.98 / 99.31 79.82 / 58.35 75.10 / 85.75 93.57 / 96.44
    ODIN 40.32 / 12.95 87.21 / 97.37 96.24 / 99.40 70.71 / 58.13 81.38 / 85.65 95.37 / 96.36
    Energy 26.82 / 19.14 93.07 / 96.09 98.02 / 99.14 70.87 / 65.46 81.45 / 84.84 95.38 / 96.27
    GradNorm 58.98 / 17.78 72.29 / 96.34 90.75 / 99.14 87.01 / 61.89 52.84 / 81.41 84.12 / 94.85

    With LogitNorm training, ODIN achieves an FPR95 of 12.95% on CIFAR-10 (down from 40.32% under cross-entropy), and GradNorm reduces FPR95 from 58.98% to 17.78% on CIFAR-10 and from 87.01% to 61.89% on CIFAR-100.

  6. Knowl 6 — Architectural Generalization and In-Distribution Accuracy Invariance of LogitNorm

    empirical result

    LogitNorm achieves consistent OOD detection improvements across multiple network backbones without degrading standard in-distribution (ID) classification accuracy.

    Comparison across Wide Residual Networks (WRN-40-2), ResNet-34, and DenseNet-BC on CIFAR-10 (using MSP scoring, averaged across six OOD test sets):

    Architecture ID Accuracy (↑\uparrow) FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow)
    Cross-Entropy / LogitNorm
    WRN-40-2 94.75% / 94.69% 49.52% / 15.65% 91.55% / 96.91% 97.98% / 99.31%
    ResNet-34 95.01% / 95.14% 47.74% / 15.82% 91.15% / 97.01% 97.72% / 99.32%
    DenseNet-BC 94.55% / 94.37% 50.41% / 18.57% 91.48% / 96.16% 98.04% / 99.10%

    On CIFAR-100 with WRN-40-2, LogitNorm achieves 75.12% ID test accuracy compared to 75.23% for standard cross-entropy loss, demonstrating that eliminating logit norm growth does not impair representation capacity.

  7. Knowl 7 — ID Confidence Calibration via Post-Hoc Temperature Scaling under LogitNorm

    empirical result

    Because LogitNorm constrains output logits to a constant norm during training, raw in-distribution softmax outputs avoid extreme concentration at 1.0 and maintain a smoother probability distribution. While unscaled LogitNorm displays higher raw Expected Calibration Error (ECE, M=15M=15 bins) due to this conservative scaling, applying post-hoc Temperature Scaling (TS, optimizing temperature TT on a hold-out validation set) achieves superior calibration compared to cross-entropy with TS.

    Evaluation on WRN-40-2:

    Dataset Calibration Setting Cross-Entropy LogitNorm (ours)
    CIFAR-10 without TS 3.20% 66.95%
    CIFAR-10 with TS 0.66% 0.41%
    CIFAR-100 without TS 11.69% 70.45%
    CIFAR-100 with TS 2.18% 1.67%

    With Temperature Scaling, LogitNorm reduces ECE to 0.41% on CIFAR-10 (compared to 0.66% for cross-entropy) and to 1.67% on CIFAR-100 (compared to 2.18% for cross-entropy).

  8. Knowl 8 — Comparative Analysis between Logit Normalization and Logit Penalty Regularization

    empirical result

    An alternative approach to prevent logit norm growth is adding an explicit L2L_2 norm penalty to the cross-entropy objective:

    Llogit_penalty(f(x;θ),y)=LCE(f(x;θ),y)+λ∥f(x;θ)∥2\mathcal{L}_{\text{logit\_penalty}}(f(x; \theta), y) = \mathcal{L}_{\text{CE}}(f(x; \theta), y) + \lambda \|f(x; \theta)\|_2

    where λ>0\lambda > 0 is a Lagrangian multiplier trade-off hyperparameter.

    Comparing LogitPenalty (λ=0.05\lambda=0.05) against LogitNorm (τ=0.04\tau=0.04) on WRN-40-2 trained on CIFAR-10, with L2L_2 norm measured on CIFAR-10 (ID) and SVHN (OOD):

    Loss Function FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow) L2L_2 Norm (ID / OOD)
    CrossEntropy 49.52% 91.55% 97.98% 12.80 / 10.91
    LogitPenalty (λ=0.05\lambda=0.05) 57.62% 73.27% 87.62% 1.90 / 1.74
    LogitNorm (τ=0.04\tau=0.04) 15.65% 96.91% 99.31% 1.48 / 0.49

    While LogitPenalty constrains the logit norm on ID data (1.90), it also produces a similarly high logit norm on OOD data (1.74), leading to poor OOD separability (FPR95 worsens to 57.62%). Additionally, large penalty weights λ\lambda cause training instability and optimization divergence. In contrast, LogitNorm causes OOD samples to naturally have substantially smaller logit magnitudes (0.49) than ID samples (1.48) during evaluation.

  9. Knowl 9 — Comparison of LogitNorm with Cosine-Based Feature/Weight Normalization (GODIN)

    empirical result

    Existing normalization techniques such as Generalized ODIN (GODIN) and Cosine Loss decompose each logit fi(x;θ)f_i(x; \theta) by normalizing both the penultimate feature representation ϕ(x)\phi(x) and final linear layer weights wiw_i: fi(x;θ)=cos⁡(wi,ϕ(x))g(x;θ)f_i(x; \theta) = \frac{\cos(w_i, \phi(x))}{g(x; \theta)}, using the maximum cosine similarity as their detection score.

    In contrast, LogitNorm normalizes the entire output vector f(x;θ)=(f1,…,fk)⊤f(x; \theta) = (f_1, \dots, f_k)^\top directly rather than decomposing inner product components in the feature layer. When evaluated on CIFAR-10 with WRN-40-2:

    Method FPR95 (↓\downarrow) AUROC (↑\uparrow) AUPR (↑\uparrow)
    GODIN 25.24% 95.07% 98.93%
    LogitNorm (ours, with ODIN score) 15.65% 96.91% 99.31%

    LogitNorm improves FPR95 to 15.65% (versus 25.24% for GODIN). Furthermore, while cosine normalization methods require specialized cosine-based scoring metrics, LogitNorm functions as a general drop-in loss replacement that enhances various post-hoc OOD scoring functions.

  10. Knowl 10 — Hyperparameter Sensitivity and Validation Scheme for LogitNorm Temperature

    limitation

    The effectiveness of LogitNorm is sensitive to the temperature hyperparameter τ∈R+\tau \in \mathbb{R}^+. When evaluating WRN-40-2 on CIFAR-10 across τ∈[0.01,0.10]\tau \in [0.01, 0.10], average FPR95 across six OOD test sets decreases as τ\tau increases up to τ=0.04\tau = 0.04 (reaching optimal OOD detection around 15.65%15.65\%), but deteriorates rapidly when τ\tau increases further (reaching ≈30%\approx 30\% at τ=0.10\tau = 0.10), consistent with the loss lower bound log⁡(1+(k−1)e−2/τ)\log(1 + (k-1)e^{-2/\tau}) impairing optimization at larger τ\tau.

    Selecting τ\tau requires grid search over a candidate set (such as {0.001,0.005,0.01,…,0.05}\{0.001, 0.005, 0.01, \dots, 0.05\}) evaluated against an auxiliary validation set of synthetic Gaussian noise. This requirement entails training multiple models across values of τ\tau, adding computational overhead. Additionally, while the paper provides mathematical justification for why cross-entropy induces overconfidence via logit growth on in-distribution training data, a fully rigorous theoretical proof explaining why LogitNorm produces smaller logit norms specifically on unseen out-of-distribution inputs remains an open question.

Coverage note — None was omitted. The detailed per-dataset appendix breakdown tables for ResNet-34, DenseNet-BC, and various scoring functions were integrated into the primary empirical knowls via their mean values and key benchmark data to avoid fragmented redundancy while preserving complete numerical fidelity.

References

  1. 1.Bendale, A. and Boult, T. E. Towards open set deep networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 1563–1572, 2016.
  2. 2.Bevandić, P., Krešo, I., Oršić, M., and Šegvić, S. Discriminative out-of-distribution detection for semantic segmentation. arXiv preprint arXiv:1808.07703, 2018.
  3. 3.Chen, J., Li, Y., Wu, X., Liang, Y., and Jha, S. Atom: Robustifying out-of-distribution detection using outlier mining. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 430–445. Springer, 2021.
  4. 4.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020.
  5. 5.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A. Describing textures in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3606–3613, 2014.
  6. 6.Deng, J., Guo, J., Xue, N., and Zafeiriou, S. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4690–4699, 2019.
  7. 7.Du, X., Wang, X., Gozum, G., and Li, Y. Unknown-aware object detection: Learning what you don’t know from videos in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022a.
  8. 8.Du, X., Wang, Z., Cai, M., and Li, Y. Vos: Learning what you don’t know by virtual outlier synthesis. In Proceedings of the International Conference on Learning Representations, 2022b.
  9. 9.Forst, W. and Hoffmann, D. Optimization—Theory and Practice. Springer Science & Business Media, 2010.
  10. 10.Geifman, Y. and El-Yaniv, R. Selectivenet: A deep neural network with an integrated reject option. In International Conference on Machine Learning, pp. 2151–2159. PMLR, 2019.
  11. 11.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International Conference on Machine Learning, pp. 1321–1330, 2017.
  12. 12.Gupta, C. and Ramdas, A. Top-label calibration and multiclass-to-binary reductions. In International Conference on Learning Representations, 2022.
  13. 13.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  14. 14.Hendrycks, D. and Gimpel, K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. arXiv preprint arXiv:1610.02136, 2016.
  15. 15.Hendrycks, D., Mazeika, M., and Dietterich, T. Deep anomaly detection with outlier exposure. Proceedings of the International Conference on Learning Representations, 2019.
  16. 16.Hsu, Y.-C., Shen, Y., Jin, H., and Kira, Z. Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10951–10960, 2020.
  17. 17.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 4700–4708, 2017.
  18. 18.Huang, R., Geng, A., and Li, Y. On the importance of gradients for detecting distributional shifts in the wild. In Advances in Neural Information Processing Systems, volume 34, 2021.
  19. 19.Jeong, T. and Kim, H. Ood-maml: Meta-learning for few-shot out-of-distribution detection and classification. Advances in Neural Information Processing Systems, 33, 2020.
  20. 20.Katz-Samuels, J., Nakhleh, J., Nowak, R., and Li, Y. Training ood detectors in their natural habitats. In International Conference on Machine Learning (ICML). PMLR, 2022.
  21. 21.Kornblith, S., Chen, T., Lee, H., and Norouzi, M. Why do better loss functions lead to less transferable features? In Advances in Neural Information Processing Systems, 2021.
  22. 22.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. Tech Report, 2009.
  23. 23.Lee, K., Lee, H., Lee, K., and Shin, J. Training confidence-calibrated classifiers for detecting out-of-distribution samples. arXiv preprint arXiv:1711.09325, 2017.
  24. 24.Lee, K., Lee, K., Lee, H., and Shin, J. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. Advances in Neural Information Processing Systems, 31, 2018.
  25. 25.Lei, J., Robins, J., and Wasserman, L. Distribution-free prediction sets. Journal of the American Statistical Association, 108(501):278–287, 2013.
  26. 26.Liang, S., Li, Y., and Srikant, R. Enhancing the reliability of out-of-distribution image detection in neural networks. In International Conference on Learning Representations, 2018.
  27. 27.Lin, T.-Y., Goyal, P., Girshick, R. B., He, K., and Dollár, P. Focal loss for dense object detection. 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2999–3007, 2017.
  28. 28.Liu, W., Wen, Y., Yu, Z., Li, M., Raj, B., and Song, L. Sphereface: Deep hypersphere embedding for face recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 212–220, 2017.
  29. 29.Liu, W., Wang, X., Owens, J., and Li, Y. Energy-based out-of-distribution detection. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  30. 30.Malinin, A. and Gales, M. Predictive uncertainty estimation via prior networks. arXiv preprint arXiv:1802.10501, 2018.
  31. 31.Ming, Y., Fan, Y., and Li, Y. Poem: Out-of-distribution detection with posterior sampling. In International Conference on Machine Learning (ICML). PMLR, 2022a.
  32. 32.Ming, Y., Sun, Y., Dia, O., and Li, Y. Cider: Exploiting hyperspherical embeddings for out-of-distribution detection. arXiv preprint arXiv:2203.04450, 2022b.
  33. 33.Mohseni, S., Pitale, M., Yadawa, J., and Wang, Z. Self-supervised learning for generalizable out-of-distribution detection. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 5216–5223, 2020.
  34. 34.Morteza, P. and Li, Y. Provable guarantees for understanding out-of-distribution detection. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022.
  35. 35.Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P. H., and Dokania, P. K. Calibrating deep neural networks using focal loss. 2020.
  36. 36.Müller, R., Kornblith, S., and Hinton, G. E. When does label smoothing help? In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  37. 37.Naeini, M. P., Cooper, G., and Hauskrecht, M. Obtaining well calibrated probabilities using bayesian binning. In Twenty-Ninth AAAI Conference on Artificial Intelligence, pp. 2901–2907, 2015.
  38. 38.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011.
  39. 39.Nguyen, A., Yosinski, J., and Clune, J. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 427–436, 2015.
  40. 40.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, pp. 8024–8035. 2019.
  41. 41.Platt, J. et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers, 10(3):61–74, 1999.
  42. 42.Ranjan, R., Castillo, C. D., and Chellappa, R. L2-constrained softmax loss for discriminative face verification. arXiv preprint arXiv:1703.09507, 2017.
  43. 43.Sastry, C. S. and Oore, S. Detecting out-of-distribution examples with gram matrices. In International Conference on Machine Learning, pp. 8491–8501, 2020.
  44. 44.Sehwag, V., Chiang, M., and Mittal, P. Ssd: A unified framework for self-supervised outlier detection. In International Conference on Learning Representations, 2021.
  45. 45.Sohn, K. Improved deep metric learning with multi-class n-pair loss objective. In Advances in Neural Information Processing Systems, 2016.
  46. 46.Sun, Y., Guo, C., and Li, Y. React: Out-of-distribution detection with rectified activations. Advances in Neural Information Processing Systems, 34, 2021.
  47. 47.Sun, Y., Ming, Y., Zhu, X., and Li, Y. Out-of-distribution detection with deep nearest neighbors. In International Conference on Machine Learning (ICML). PMLR, 2022.
  48. 48.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2818–2826, 2016.
  49. 49.Tack, J., Mo, S., Jeong, J., and Shin, J. Csi: Novelty detection via contrastive learning on distributionally shifted instances. In Advances in Neural Information Processing Systems, 2020.
  50. 50.Techapanurak, E. and Okatani, T. Hyperparameter-free out-of-distribution detection using softmax of scaled cosine similarity. arXiv:1905.10628, 2019.
  51. 51.van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. ArXiv, abs/1807.03748, 2018.
  52. 52.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of Machine Learning Research, 9(11), 2008.
  53. 53.Wang, D.-B., Feng, L., and Zhang, M.-L. Rethinking calibration of deep neural networks: Do not be afraid of overconfidence. In Advances in Neural Information Processing Systems, 2021a.
  54. 54.Wang, F., Xiang, X., Cheng, J., and Yuille, A. L. Normface: L2 hypersphere embedding for face verification. In Proceedings of the 25th ACM international conference on Multimedia, pp. 1041–1049, 2017.
  55. 55.Wang, H., Wang, Y., Zhou, Z., Ji, X., Gong, D., Zhou, J., Li, Z., and Liu, W. Cosface: Large margin cosine loss for deep face recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5265–5274, 2018.
  56. 56.Wang, H., Liu, W., Bocchieri, A., and Li, Y. Can multi-label classification networks know what they don’t know? Advances in Neural Information Processing Systems, 34, 2021b.
  57. 57.Wei, H., Tao, L., Xie, R., and An, B. Open-set label noise can improve robustness against inherent label noise. In Advances in Neural Information Processing Systems, 2021.
  58. 58.Wei, H., Tao, L., Xie, R., Feng, L., and An, B. Open-sampling: Exploring out-of-distribution data for re-balancing long-tailed datasets. In International Conference on Machine Learning (ICML). PMLR, 2022.
  59. 59.Wu, Z., Xiong, Y., Stella, X. Y., and Lin, D. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  60. 60.Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A. Sun database: Large-scale scene recognition from abbey to zoo. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3485–3492, 2010.
  61. 61.Xu, J., Sun, X., Zhang, Z., Zhao, G., and Lin, J. Understanding and improving layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  62. 62.Xu, P., Ehinger, K. A., Zhang, Y., Finkelstein, A., Kulkarni, S. R., and Xiao, J. Turkergaze: Crowdsourcing saliency with webcam based eye tracking. arXiv preprint arXiv:1504.06755, 2015.
  63. 63.Yu, F., Seff, A., Zhang, Y., Song, S., Funkhouser, T., and Xiao, J. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.
  64. 64.Zadrozny, B. and Elkan, C. Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers. In International Conference on Machine Learning, pp. 609–616, 2001.
  65. 65.Zagoruyko, S. and Komodakis, N. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  66. 66.Zhang, X., Zhao, R., Qiao, Y., Wang, X., and Li, H. Adacos: Adaptively scaling cosine logits for effectively learning deep face representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10823–10832, 2019.
  67. 67.Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A. Places: A 10 million image database for scene recognition. IEEE transactions on Pattern Analysis and Machine Intelligence, 40(6):1452–1464, 2017.
  68. 68.Zhu, Z., Dong, Z., and Liu, Y. Detecting corrupted labels without training a model to predict. In International Conference on Machine Learning (ICML). PMLR, 2022.

Citation

MLA
Wei, H., et al. “Mitigating Neural Network Overconfidence with Logit Normalization”. arXiv, 2022, http://arxiv.org/abs/2205.09310v2.
APA
Wei, H., Xie, R., Cheng, H., Feng, L., An, B., & Li, Y. (2022). Mitigating Neural Network Overconfidence with Logit Normalization. arXiv. http://arxiv.org/abs/2205.09310v2
Chicago
Wei, H., R. Xie, H. Cheng, L. Feng, B. An, and Y. Li. 2022. “Mitigating Neural Network Overconfidence with Logit Normalization”. arXiv. http://arxiv.org/abs/2205.09310v2.
Harvard
Wei, H. et al. (2022) “Mitigating Neural Network Overconfidence with Logit Normalization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.09310v2.
Vancouver
1. Wei H, Xie R, Cheng H, Feng L, An B, Li Y (2022) Mitigating Neural Network Overconfidence with Logit Normalization. arXiv

BibTeX

@article{wei2022mitigating,
  title = {Mitigating Neural Network Overconfidence with Logit Normalization},
  author = {Wei, Hongxin and Xie, Renchunzi and Cheng, Hao and Feng, Lei and An, Bo and Li, Yixuan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.09310v2},
  eprint = {2205.09310}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/