Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement

Kai XuRongyu ChenGianni FranchiAngela Yao

article2024ICLR79 citations

Demonstrates that activation scaling outperforms activation pruning in out-of-distribution detection, introducing post-hoc and training-time methods that achieve state-of-the-art results on ImageNet benchmarks without sacrificing in-distribution accuracy.

Listen

Modern deep learning systems frequently encounter unfamiliar real-world data during deployment, making out-of-distribution (OOD) detection critical for ensuring operational safety, reliability, and risk mitigation. When neural networks fail to identify unfamiliar inputs, they often produce high-confidence, erroneous predictions on data outside their training distribution. Prior post-hoc techniques for identifying these anomalies—specifically activation shaping methods like ASH—prune low feature activations while scaling remaining ones, but these mechanisms have remained poorly understood and often degrade baseline classification performance on standard in-distribution (ID) data.

The main objective of the article is to provide a theoretical analysis of how feature pruning and scaling influence OOD detection, while developing practical techniques that maximize OOD detection rates without compromising standard classification accuracy.

To achieve this, the article mathematically models feature activations as rectified Gaussian distributions and empirically validates these mechanics across standard image benchmarks, including ImageNet-1K, CIFAR-10, and CIFAR-100 across both convolutional architectures like ResNet-50 and DenseNet-101. The evaluation assesses false positive rates at a 95% true positive rate (FPR@95) and the area under the receiver operating characteristic curve (AUROC) against established baselines, distinguishing between conceptually similar categories (near-OOD) and distinctly dissimilar inputs (far-OOD).

The article demonstrates five core findings. First, mathematical and empirical analysis shows that feature pruning actually harms anomaly detection by reducing separation between known and unknown data, whereas activation scaling is the primary driver of performance gains. Second, the proposed post-hoc method, SCALE, exclusively applies sample-specific scaling without pruning, fully preserving standard classification accuracy (76.18% on ImageNet) while achieving superior anomaly detection. Third, on ImageNet benchmarks, SCALE reduces FPR@95 from 62.03% to 59.76% and improves AUROC to 81.36% for near-OOD cases compared to ASH-S. Fourth, applying scaling principles during model fine-tuning via Intermediate Tensor SHaping (ISH) yields state-of-the-art results (84.01% AUROC on near-OOD and 96.79% on far-OOD). Fifth, ISH accomplishes these training-time gains using only 10 fine-tuning epochs—requiring roughly one-third of the computational training overhead of heavy data-augmentation baselines like AugMix.

These findings indicate that organizations deploying vision models can achieve industry-leading anomaly rejection without sacrificing core task performance or redesigning underlying network architectures. By applying SCALE at inference time or fine-tuning models with ISH, teams can mitigate the financial and safety risks associated with model overconfidence on unseen data at minimal computational expense.

Organizations should adopt SCALE as an off-the-shelf post-processing standard for existing vision models to immediately improve anomaly detection at zero cost to baseline accuracy. For new model pipelines, teams should incorporate ISH fine-tuning to maximize detection robustness before production deployment.

Confidence in these findings is high across standard vision architectures and established academic benchmarks. However, the theoretical derivations rely on the assumption that latent feature activations approximate rectified Gaussian distributions. Practitioners should exercise caution and conduct pilot validations when deploying these methods on non-image domains or novel architectures where activation distributions may differ.

No sufficiently relevant recommendations were found.

Cover for Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement

Abstract

The capacity of a modern deep learning system to determine if a sample falls within its realm of knowledge is fundamental and important. In this paper, we offer insights and analyses of recent state-of-the-art out-of-distribution (OOD) detection methods - extremely simple activation shaping (ASH). We demonstrate that activation pruning has a detrimental effect on OOD detection, while activation scaling enhances it. Moreover, we propose SCALE, a simple yet effective post-hoc network enhancement method for OOD detection, which attains state-of-the-art OOD detection performance without compromising in-distribution (ID) accuracy. By integrating scaling concepts into the training process to capture a sample's ID characteristics, we propose Intermediate Tensor SHaping (ISH), a lightweight method for training time OOD detection enhancement. We achieve AUROC scores of +1.85% for near-OOD and +0.74% for far-OOD datasets on the OpenOOD v1.5 ImageNet-1K benchmark. Our code and models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Activation Scaling for Post-hoc Model Enhancement
  • 3.1 Preliminaries
  • 3.2 Analysis on ASH:
  • 3.3 SCALE Criterion for OOD Detection
  • 3.4 Incorporating SCALE into Training
  • 4 Experiments
  • 4.1 Settings
  • 4.2 SCALE for Post-Hoc OOD Detection
  • 4.3 ISH for Training-Time Model Enhancement
  • 5 Conclusion
  • References
  • A Details of Proof
  • B Full Experiments

Knowls

  1. Knowl 1 — SCALE: Post-Hoc Activation Scaling for Out-of-Distribution Detection

    model/method

    SCALE is a post-hoc network adjustment method for out-of-distribution (OOD) detection that applies sample-specific activation scaling at the penultimate layer while leaving activations unpruned, thereby improving OOD separation without altering in-distribution (ID) classification accuracy.

    Let a deep network have penultimate feature vector a=f(x)∈RD\mathbf{a} = f(\mathbf{x}) \in \mathbb{R}^D and linear classifier weights W∈RK×D\mathbf{W} \in \mathbb{R}^{K \times D} with bias b∈RK\mathbf{b} \in \mathbb{R}^K. For a given test input x\mathbf{x} and a pruning percentile parameter p∈(0,1)p \in (0, 1), let Pp(a)P_p(\mathbf{a}) denote the pp-th percentile threshold across the DD elements of a\mathbf{a}. The scaling factor rr is computed as the ratio between the total activation sum and the sum of activations exceeding the pp-th percentile:

    r=QQp=∑j=1Daj∑j: aj>Pp(a)ajr = \frac{Q}{Q_p} = \frac{\sum_{j=1}^D a_j}{\sum_{j:\, a_j > P_p(\mathbf{a})} a_j}

    SCALE scales all feature activations by sf(a)j=exp⁡(r)s_f(\mathbf{a})_j = \exp(r) without zeroing out any components, producing adjusted logits z′\mathbf{z}':

    z′=W⋅(a∘sf(a))+b\mathbf{z}' = \mathbf{W} \cdot (\mathbf{a} \circ s_f(\mathbf{a})) + \mathbf{b}

    where ∘\circ denotes element-wise multiplication. An energy-based OOD scoring function SEBO(x)S_{\text{EBO}}(\mathbf{x}) with temperature TT is then computed on the adjusted logits:

    SEBO(x)=T⋅log⁡∑k=1Kezk′/TS_{\text{EBO}}(\mathbf{x}) = T \cdot \log \sum_{k=1}^K e^{z'_k / T}

    Because every feature dimension is scaled by the identical scalar multiplier exp⁡(r)>0\exp(r) > 0, the relative ordering of the logits is preserved (arg⁡max⁡kzk′=arg⁡max⁡kzk \arg\max_k z'_k = \arg\max_k z_k when bias is absent or negligible), ensuring that the model's closed-set ID classification accuracy is strictly retained.

  2. Knowl 2 — Intermediate Tensor Shaping (ISH) for Training-Time OOD Enhancement

    model/method

    Intermediate Tensor SHaping (ISH) is a training-time model enhancement technique that prioritizes samples with high in-distribution characteristics ("ID-ness") during parameter optimization without altering the forward inference architecture or incurring heavy data augmentation overhead.

    During fine-tuning, the forward pass remains unmodified to maintain classification inference behavior. In the backward pass, the penultimate layer activations saved as intermediate tensors are scaled by the sample-specific ID-ness factor sf(ai)=exp⁡(ri)s_f(\mathbf{a}_i) = \exp(r_i), where ri=∑jai,j∑j: ai,j>Pp(ai)ai,jr_i = \frac{\sum_j a_{i,j}}{\sum_{j:\, a_{i,j} > P_p(\mathbf{a}_i)} a_{i,j}} for training sample ii with percentile threshold pp.

    The gradient update for the fully connected classification layer weight matrix W\mathbf{W} at step tt with learning rate η\eta is given by:

    Wt+1=Wt−η∑i[(ai∘sf(ai))⊤∇ziLCE]\mathbf{W}^{t+1} = \mathbf{W}^t - \eta \sum_i \left[ (\mathbf{a}_i \circ s_f(\mathbf{a}_i))^\top \nabla_{\mathbf{z}_i} \mathcal{L}_{\text{CE}} \right]

    where ∇ziLCE\nabla_{\mathbf{z}_i} \mathcal{L}_{\text{CE}} is the gradient of the cross-entropy loss with respect to logit vector zi\mathbf{z}_i.

    In practice, a pretrained network is fine-tuned with ISH for 10 epochs using a cosine annealing learning rate schedule starting at 0.0030.003 down to 00, coupled with a reduced weight decay of 5×10−65 \times 10^{-6}.

  3. Knowl 3 — Disparity in Activation Scaling Factors Between ID and OOD Data

    theoretical result

    Assume the penultimate pre-ReLU activations of in-distribution (ID) and out-of-distribution (OOD) samples are independent and identically distributed Gaussian variables, leading to post-ReLU activations that follow rectified Gaussian distributions aj(ID)∼NR(μID,σID)a_j^{(\text{ID})} \sim \mathcal{N}^R(\mu_{\text{ID}}, \sigma_{\text{ID}}) and aj(OOD)∼NR(μOOD,σOOD)a_j^{(\text{OOD})} \sim \mathcal{N}^R(\mu_{\text{OOD}}, \sigma_{\text{OOD}}), with ratio γID=μID/σID>γOOD=μOOD/σOOD\gamma_{\text{ID}} = \mu_{\text{ID}}/\sigma_{\text{ID}} > \gamma_{\text{OOD}} = \mu_{\text{OOD}}/\sigma_{\text{OOD}}.

    Let Q=∑j=1DajQ = \sum_{j=1}^D a_j and Qp=∑j: aj>Pp(a)ajQ_p = \sum_{j:\, a_j > P_p(\mathbf{a})} a_j represent the sum of all activations and the sum of activations exceeding the pp-th percentile threshold Pp(a)P_p(\mathbf{a}), respectively. Define the percentile factor:

    C(p)=ϕ(2 erf−1(2p−1))1−Φ(2 erf−1(2p−1))C(p) = \frac{\phi(\sqrt{2}\,\text{erf}^{-1}(2p-1))}{1 - \Phi(\sqrt{2}\,\text{erf}^{-1}(2p-1))}

    where ϕ(⋅)\phi(\cdot) and Φ(⋅)\Phi(\cdot) denote the probability density function and cumulative distribution function of the standard normal distribution N(0,1)\mathcal{N}(0, 1), and erf−1(⋅)\text{erf}^{-1}(\cdot) is the inverse error function.

    For percentiles pp where C(p)C(p) is sufficiently large, the ratio of unpruned-to-total activation mass satisfies:

    QpIDQID<QpOODQOOD\frac{Q_p^{\text{ID}}}{Q^{\text{ID}}} < \frac{Q_p^{\text{OOD}}}{Q^{\text{OOD}}}

    Consequently, the scaling factor r=Q/Qpr = Q / Q_p satisfies:

    rID>rOODr^{\text{ID}} > r^{\text{OOD}}

    This establishes that ID samples systematically yield higher scaling multipliers than OOD samples.

  4. Knowl 4 — Mechanism Analysis: Harm of Pruning and Benefit of Scaling in Activation Shaping

    theoretical result

    In activation shaping methods such as ASH, the transformation of penultimate activations combines lower-percentile truncation (pruning) with activation scaling. Analyzing these two components separately reveals opposing effects on OOD discriminability:

    1. Activation Pruning (Detrimental Effect): The relative reduction in activations caused by zeroing elements below percentile pp is: DPruning=Q−QpQ=1−QpQD^{\text{Pruning}} = \frac{Q - Q_p}{Q} = 1 - \frac{Q_p}{Q} Because QpIDQID<QpOODQOOD\frac{Q_p^{\text{ID}}}{Q^{\text{ID}}} < \frac{Q_p^{\text{OOD}}}{Q^{\text{OOD}}}, the expected reduction for ID samples is strictly larger than for OOD samples (DIDPruning>DOODPruningD_{\text{ID}}^{\text{Pruning}} > D_{\text{OOD}}^{\text{Pruning}}). Since logit and energy score reductions are proportional to activation reductions, pruning reduces ID energy scores more severely than OOD energy scores, creating greater overlap in energy score distributions and reducing OOD detection AUROC.

    2. Activation Scaling (Beneficial Effect): Scaling all activations by sf(a)=exp⁡(r)s_f(\mathbf{a}) = \exp(r) yields a relative activation increase: IScaling=r−1=QQp−1I^{\text{Scaling}} = r - 1 = \frac{Q}{Q_p} - 1 Because rID>rOODr^{\text{ID}} > r^{\text{OOD}}, ID samples experience a larger relative amplification than OOD samples (IIDScaling>IOODScalingI_{\text{ID}}^{\text{Scaling}} > I_{\text{OOD}}^{\text{Scaling}}). This selectively boosts logits and energy scores for ID data relative to OOD data, widening the score separation and improving detection AUROC.

  5. Knowl 5 — Post-Hoc OOD Detection Performance of SCALE on ImageNet-1K

    empirical result

    Evaluated on ImageNet-1K with a ResNet-50 backbone under the OpenOOD v1.5 benchmark protocol (using validation set tuning at p=0.85p = 0.85), SCALE outperforms existing post-hoc scoring and feature-enhancement methods in both Near-OOD (SSB-hard, NINCO) and Far-OOD (iNaturalist, Textures, OpenImage-O) settings without compromising ID classification accuracy.

    Postprocessor Near-OOD Far-OOD ID ACC (%)
    FPR@95 (%) ↓\downarrow AUROC (%) ↑\uparrow FPR@95 (%) ↓\downarrow AUROC (%) ↑\uparrow ↑\uparrow
    EBO 68.56 75.89 38.40 89.47 76.18
    MSP 65.67 76.02 51.47 85.23 76.18
    MLS 67.82 76.46 38.20 89.58 76.18
    GEN 65.30 76.85 35.62 89.77 76.18
    RMDS 65.04 76.99 40.91 86.38 76.18
    TempScale 64.51 77.14 46.67 87.56 76.18
    ReAct 66.75 77.38 26.31 93.67 75.58
    ASH-S 62.03 79.63 16.86 96.47 75.51
    SCALE (Ours) 59.76 81.36 16.53 96.53 76.18

    Compared to ASH-S, SCALE improves Near-OOD AUROC by +1.73% and reduces FPR@95 by 2.27%, while preserving full baseline ID accuracy (76.18% vs. 75.51% for ASH-S and 75.58% for ReAct).

  6. Knowl 6 — OOD Detection Performance of ISH Training-Time Enhancement on ImageNet-1K

    empirical result

    On the ImageNet-1K benchmark using ResNet-50 and post-hoc SCALE scoring, Intermediate Tensor SHaping (ISH) achieves state-of-the-art OOD detection performance while requiring only 10 extended fine-tuning epochs (90+1090+10), outperforming data-augmentation methods like AugMix (180 total epochs) and RegMixup (90+3090+30).

    Method Epochs Near-OOD Far-OOD ID ACC
    Ori.+Ext. FPR@95 (%) ↓\downarrow AUROC (%) ↑\uparrow FPR@95 (%) ↓\downarrow AUROC (%) ↑\uparrow (%) ↑\uparrow
    LogitNorm 90+30 68.56 74.62 31.33 91.54 76.45
    CIDER 90+30 71.69 68.97 28.69 92.18 –
    TorchVision Model 90 59.76 81.36 16.53 96.53 76.13
    TorchVision Extended 90+10 59.25 82.67 18.48 96.24 76.84
    RegMixup 90+30 63.55 80.85 19.87 95.94 76.88
    AugMix 180 60.58 83.55 21.01 95.99 77.64
    ISH (Ours) 90+10 55.73 84.01 15.62 96.79 76.74

    ISH achieves 84.01% Near-OOD AUROC and 96.79% Far-OOD AUROC with an FPR@95 of 55.73% on Near-OOD and 15.62% on Far-OOD.

  7. Knowl 7 — Post-Hoc OOD Detection Performance on CIFAR-10 and CIFAR-100

    empirical result

    On CIFAR-10 and CIFAR-100 benchmarks evaluated using a DenseNet-101 backbone across six OOD test sets (SVHN, LSUN-Crop, LSUN-Resize, iSUN, Textures, Places365), SCALE achieves superior average OOD detection performance compared to other post-hoc processors and maintains 100% of the baseline ID classification accuracy.

    Method CIFAR-10 CIFAR-100 ID ACC (%)
    FPR@95 ↓\downarrow AUROC ↑\uparrow FPR@95 ↓\downarrow AUROC ↑\uparrow (C-10 / C-100)
    MSP 48.73 92.46 80.13 74.36 94.53 / 75.04
    EBO 26.55 94.57 68.45 81.19 94.53 / 75.04
    ReAct 26.45 94.95 62.27 84.47 – / –
    DICE 20.83 95.24 49.72 87.23 – / –
    ASH-S 15.05 96.61 41.40 90.02 94.02 / 71.65
    SCALE (Ours) 12.57 97.27 38.99 90.74 94.53 / 75.04

    On CIFAR-10, SCALE improves FPR@95 by 2.48% and AUROC by 0.66% over ASH-S. On CIFAR-100, SCALE improves FPR@95 by 2.41% and AUROC by 0.72% over ASH-S, while preventing the ID accuracy drop seen with ASH-S (which drops to 94.02% on CIFAR-10 and 71.65% on CIFAR-100).

  8. Knowl 8 — Sensitivity of SCALE to Pruning Percentile Parameter $p$

    empirical result

    The choice of percentile pp used to compute the scaling ratio r=Q/Qpr = Q/Q_p governs the separation quality in SCALE. Testing ResNet-50 across percentiles p∈{65,70,75,80,85,90,95}p \in \{65, 70, 75, 80, 85, 90, 95\} demonstrates that OOD detection performance monotonically improves up to p=85%p = 85\% and drops sharply thereafter.

    Metric p=65p=65 p=70p=70 p=75p=75 p=80p=80 p=85p=85 p=90p=90 p=95p=95
    Near-OOD FPR@95 (%) 62.45 61.65 61.12 60.12 59.76 63.19 78.62
    Near-OOD AUROC (%) 79.31 79.83 80.41 81.01 81.36 80.14 73.40
    Far-OOD FPR@95 (%) 24.08 22.21 20.20 18.26 16.53 18.58 32.42
    Far-OOD AUROC (%) 94.43 95.02 95.61 96.17 96.53 96.20 93.28

    This behavior aligns with the theoretical inflection point of C(p)C(p), which reaches a maximum around p≈0.95p \approx 0.95 before falling off, combined with high estimation variance when rr is calculated from the remaining small fraction of total activation dimensions (D=2048D = 2048).

Coverage note — None was omitted; all key theoretical analyses of pruning/scaling in ASH, the proposed SCALE post-hoc method, the ISH training algorithm, Proposition 3.1, and primary benchmark experiments across ImageNet and CIFAR were captured as self-sufficient knowls.

References

  1. 1.Julian Bitterwolf, Maximilian M¨uller, and Matthias Hein. In or out? fixing imagenet out-of-distribution detection evaluation. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 2471–2506. PMLR, 2023. URL https://proceedings.mlr.press/v202/bitterwolf23a.html.
  2. 2.Joya Chen, Kai Xu, Yuhui Wang, Yifei Cheng, and Angela Yao. Dropit: Dropping intermediate tensors for memory-efficient DNN training. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023a. URL https://openreview.net/pdf?id=Kn6i2BZW69w.
  3. 3.Xuanyao Chen, Zhijian Liu, Haotian Tang, Li Yi, Hang Zhao, and Song Han. Sparsevit: Revisiting activation sparsity for efficient high-resolution vision transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pp. 2061–2070. IEEE, 2023b. doi: 10.1109/CVPR52729.2023.00205. URL https://doi.org/10.1109/CVPR52729.2023.00205.
  4. 4.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014, Columbus, OH, USA, June 23-28, 2014, pp. 3606–3613. IEEE Computer Society, 2014. doi: 10.1109/CVPR.2014.461. URL https://doi.org/10.1109/CVPR.2014.461.
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pp. 248–255. IEEE Computer Society, 2009. doi: 10.1109/CVPR.2009.5206848. URL https://doi.org/10.1109/CVPR.2009.5206848.
  6. 6.Terrance DeVries and Graham W. Taylor. Learning confidence for out-of-distribution detection in neural networks. CoRR, abs/1802.04865, 2018. URL http://arxiv.org/abs/1802.04865.
  7. 7.Andrija Djurisic, Nebojsa Bozanic, Arjun Ashok, and Rosanne Liu. Extremely simple activation shaping for out-of-distribution detection. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=ndYXTEL6cZz.
  8. 8.R. David Evans and Tor M. Aamodt. AC-GC: lossy activation compression with guaranteed convergence. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 27434–27448, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/e655c7716a4b3ea67f48c6322fc42ed6-Abstract.html.
  9. 9.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Doina Precup and Yee Whye Teh (eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pp. 1321–1330. PMLR, 2017. URL http://proceedings.mlr.press/v70/guo17a.html.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 770–778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90.
  11. 11.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=Hkg4TI9xl.
  12. 12.Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=S1gmrxHFvB.
  13. 13.Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joseph Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, and Dawn Song. Scaling out-of-distribution detection for real-world settings. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 8759–8773. PMLR, 2022. URL https://proceedings.mlr.press/v162/hendrycks22a.html.
  14. 14.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alexander Shepard, Hartwig Adam, Pietro Perona, and Serge J. Belongie. The inaturalist species classification and detection dataset. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 8769–8778. Computer Vision Foundation / IEEE Computer Society, 2018. doi: 10.1109/CVPR.2018.00914. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Van_Horn_The_INaturalist_Species_CVPR_2018_paper.html.
  15. 15.Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pp. 2261–2269. IEEE Computer Society, 2017. doi: 10.1109/CVPR.2017.243. URL https://doi.org/10.1109/CVPR.2017.243.
  16. 16.Alex Krizhevsky. Learning multiple layers of features from tiny images. 2009. URL https://api.semanticscholar.org/CorpusID:18268744.
  17. 17.Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William M. Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. Inducing and exploiting activation sparsity for fast inference on deep neural networks. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 5533–5543. PMLR, 2020. URL http://proceedings.mlr.press/v119/kurtz20a.html.
  18. 18.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicolò Cesa-Bianchi, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pp. 7167–7177, 2018. URL https://proceedings.neurips.cc/paper/2018/hash/abdeb6f575ac5c6676b747bca8d09cc2-Abstract.html.
  19. 19.Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J. Reddi, Ke Ye, Felix Chern, Felix X. Yu, Ruiqi Guo, and Sanjiv Kumar. The lazy neuron phenomenon: On emergence of activation sparsity in transformers. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=TJ2nxciYCk-.
  20. 20.Weitang Liu, Xiaoyun Wang, John D. Owens, and Yixuan Li. Energy-based out-of-distribution detection. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/f5496252609c43eb8a3d147ab9b9c006-Abstract.html.
  21. 21.Xiaoxuan Liu, Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han, Jianfei Chen, Zhiyuan Liu, Jie Tang, Joey Gonzalez, Michael W. Mahoney, and Alvin Cheung. GACT: activation compressed training for generic network architectures. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 14139–14152. PMLR, 2022. URL https://proceedings.mlr.press/v162/liu22v.html.
  22. 22.Xixi Liu, Yaroslava Lochman, and Christopher Zach. GEN: pushing the limits of softmax-based out-of-distribution detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pp. 23946–23955. IEEE, 2023. doi: 10.1109/CVPR52729.2023.02293. URL https://doi.org/10.1109/CVPR52729.2023.02293.
  23. 23.Yifei Ming, Yiyou Sun, Ousmane Dia, and Yixuan Li. How to exploit hyperspherical embeddings for out-of-distribution detection? In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=aEFaE0W5pAd.
  24. 24.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011. URL http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf.
  25. 25.Francesco Pinto, Harry Yang, Ser Nam Lim, Philip H. S. Torr, and Puneet K. Dokania. Using mixup as a regularizer can surprisingly improve accuracy & out-of-distribution robustness. In NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/5ddcfaad1cb72ce6f1a365e8f1ecf791-Abstract-Conference.html.
  26. 26.Jie Ren, Stanislav Fort, Jeremiah Z. Liu, Abhijit Guha Roy, Shreyas Padhy, and Balaji Lakshminarayanan. A simple fix to mahalanobis distance for improving near-ood detection. CoRR, abs/2106.09022, 2021. URL https://arxiv.org/abs/2106.09022.
  27. 27.Yiyou Sun and Yixuan Li. DICE: leveraging sparsification for out-of-distribution detection. In Shai Avidan, Gabriel J. Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (eds.), Computer Vision - ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXIV, volume 13684 of Lecture Notes in Computer Science, pp. 691–708. Springer, 2022. doi: 10.1007/978-3-031-20053-3_40. URL https://doi.org/10.1007/978-3-031-20053-3_40.
  28. 28.Yiyou Sun, Chuan Guo, and Yixuan Li. React: Out-of-distribution detection with rectified activations. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 144–157, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/01894d6f048493d2cacde3c579c315a3-Abstract.html.
  29. 29.Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. Open-set recognition: A good closed-set classifier is all you need. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=5hLP5JY9S2d.
  30. 30.Haoqi Wang, Zhizhong Li, Litong Feng, and Wayne Zhang. Vim: Out-of-distribution with virtual-logit matching. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 4911–4920. IEEE, 2022. doi: 10.1109/CVPR52688.2022.00487. URL https://doi.org/10.1109/CVPR52688.2022.00487.
  31. 31.Hongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng, Bo An, and Yixuan Li. Mitigating neural network overconfidence with logit normalization. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 23631–23644. PMLR, 2022. URL https://proceedings.mlr.press/v162/wei22d.html.
  32. 32.Pingmei Xu, Krista A. Ehinger, Yinda Zhang, Adam Finkelstein, Sanjeev R. Kulkarni, and Jianxiong Xiao. Turkergaze: Crowdsourcing saliency with webcam based eye tracking. CoRR, abs/1504.06755, 2015. URL http://arxiv.org/abs/1504.06755.
  33. 33.Jingkang Yang, Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, Wenxuan Peng, Haoqi Wang, Guangyao Chen, Bo Li, Yiyou Sun, Xuefeng Du, Kaiyang Zhou, Wayne Zhang, Dan Hendrycks, Yixuan Li, and Ziwei Liu. Openood: Benchmarking generalized out-of-distribution detection. In NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/d201587e3a84fc4761eadc743e9b3f35-Abstract-Datasets_and_Benchmarks.html.
  34. 34.Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. LSUN: construction of a large-scale image dataset using deep learning with humans in the loop. CoRR, abs/1506.03365, 2015. URL http://arxiv.org/abs/1506.03365.
  35. 35.Jingyang Zhang, Jingkang Yang, Pengyun Wang, Haoqi Wang, Yueqian Lin, Haoran Zhang, Yiyou Sun, Xuefeng Du, Kaiyang Zhou, Wayne Zhang, Yixuan Li, Ziwei Liu, Yiran Chen, and Hai Li. Openood v1.5: Enhanced benchmark for out-of-distribution detection. CoRR, abs/2306.09301, 2023. doi: 10.48550/arXiv.2306.09301. URL https://doi.org/10.48550/arXiv.2306.09301.
  36. 36.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Trans. Pattern Anal. Mach. Intell., 40(6): 1452–1464, 2018. doi: 10.1109/TPAMI.2017.2723009. URL https://doi.org/10.1109/TPAMI.2017.2723009.

Citation

MLA
Xu, K., et al. “Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement”. arXiv, 2023, http://arxiv.org/abs/2310.00227v1.
APA
Xu, K., Chen, R., Franchi, G., & Yao, A. (2023). Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement. arXiv. http://arxiv.org/abs/2310.00227v1
Chicago
Xu, K., R. Chen, G. Franchi, and A. Yao. 2023. “Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement”. arXiv. http://arxiv.org/abs/2310.00227v1.
Harvard
Xu, K. et al. (2023) “Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.00227v1.
Vancouver
1. Xu K, Chen R, Franchi G, Yao A (2023) Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement. arXiv

BibTeX

@article{xu2023scaling,
  title = {Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement},
  author = {Xu, Kai and Chen, Rongyu and Franchi, Gianni and Yao, Angela},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.00227v1},
  eprint = {2310.00227}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/