Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need

Jingyao LiPengguang ChenZexin HeShaozuo YuShu LiuJiaya Jia

article2023CVPR80 citations

Demonstrates that pretraining vision transformers with masked image modeling enables models to capture intrinsic data distributions rather than classification shortcuts, outperforming existing out-of-distribution detection methods without requiring any outlier exposure.

Listen

Reliable automated visual recognition systems must not only classify known inputs correctly but also identify and flag unknown, out-of-distribution inputs for safe handling. This safety capability is critical in high-stakes domains such as autonomous driving, medical diagnostics, and fraud prevention. Traditional approaches often rely on classification-based training or exposure to outlier examples. However, classification networks frequently learn superficial shortcuts rather than the underlying distribution of the data, which leads to poor anomaly detection when unfamiliar samples share minor visual traits with known categories.

The article evaluates whether reconstruction-based pretext tasks can provide better foundational representations for out-of-distribution detection. Specifically, it demonstrates the performance of a proposed framework called Masked Image Modeling for Out-of-Distribution Detection (MOOD), which trains a vision model to reconstruct masked image patches without relying on outlier samples.

To conduct the evaluation, the authors pre-trained a standard Vision Transformer using masked image modeling on a large dataset (ImageNet-21k) and fine-tuned it on in-distribution data, integrating label smoothing for single-class tasks. They used the statistical Mahalanobis distance metric to score whether new test inputs were out-of-distribution. The method was benchmarked across four distinct scenarios—one-class detection, multi-class detection, near-distribution detection, and few-shot outlier exposure—against established state-of-the-art baselines on widely used image datasets including CIFAR-10, CIFAR-100, and ImageNet benchmarks.

The experimental findings show that the proposed framework consistently establishes new performance records. In one-class detection, MOOD improved the average detection score by 5.7 percentage points over the previous leading baseline, reaching 94.9%. In multi-class detection, it outperformed the leading method by 3.0 percentage points, achieving 97.6%. On challenging near-distribution tasks where known and unknown classes share similar visual semantics, the framework reduced misclassified outlier samples by an average of 79% and raised accuracy across difficult class pairs from 78.7% to 93.9%. Furthermore, without using any outlier examples during training, the framework achieved a 99.41% score, outperforming complex baseline models that were explicitly trained with ten outlier examples per class.

These results indicate that forcing vision models to reconstruct masked inputs teaches them intrinsic, pixel-level data distributions rather than fragile classification shortcuts. This substantially reduces operational and safety risks in automated computer vision systems. Crucially, the findings show that developers do not need cumbersome ensemble models or synthetic outlier datasets to achieve robust safety filtering, which reduces engineering complexity and computational overhead.

Organizations developing or deploying safety-critical vision systems should consider adopting masked image modeling as a standard self-supervised pre-training strategy combined with distance-based anomaly metrics. Practitioners can streamline development pipelines by eliminating the need to collect and maintain outlier exposure datasets. Further evaluation in domain-specific, real-world deployment settings is recommended to assess performance alongside runtime latency constraints.

Cover for Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need

Abstract

The core of out-of-distribution (OOD) detection is to learn the in-distribution (ID) representation, which is distinguishable from OOD samples. Previous work applied recognition-based methods to learn the ID features, which tend to learn shortcuts instead of comprehensive representations. In this work, we find surprisingly that simply using reconstruction-based methods could boost the performance of OOD detection significantly. We deeply explore the main contributors of OOD detection and find that reconstruction-based pretext tasks have the potential to provide a generally applicable and efficacious prior, which benefits the model in learning intrinsic data distributions of the ID dataset. Specifically, we take Masked Image Modeling as a pretext task for our OOD detection framework (MOOD). Without bells and whistles, MOOD outperforms previous SOTA of one-class OOD detection by 5.7%, multi-class OOD detection by 3.0%, and near-distribution OOD detection by 2.1%. It even defeats the 10-shot-per-class outlier exposure OOD detection, although we do not include any OOD samples for our detection. Codes are available at https://github.com/lijingyao20010602/MOOD.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Out-of-distribution Detection
  • 2.2. Vision Transformer
  • 2.3. Self-Supervised Pretext Task
  • 3. Method
  • 3.1. Choosing the Pretext Task
  • 3.2. Exploring Architecture
  • 3.3. About Fine-Tuning
  • 3.4. OOD Detection Metric is Important
  • 3.5. Final Algorithm of MOOD
  • 4. Experiments
  • 4.1. One-Class OOD Detection
  • 4.2. Multi-Class OOD Detection
  • 4.3. Near-Distribution OOD Detection
  • 4.4. OOD Detection with Outlier Exposure
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — MOOD Out-of-Distribution Detection Framework

    algorithm

    Masked Image Modeling for Out-of-Distribution Detection (MOOD) is an out-of-distribution (OOD) detection pipeline that leverages reconstruction-based self-supervised pre-training to learn intrinsic in-distribution (ID) representations without requiring exposure to auxiliary OOD outlier samples.

    Input: Pre-training dataset DpreD_{\text{pre}} (e.g., ImageNet-21k), In-distribution dataset DIDD_{\text{ID}} with NcN_c classes, test image xx
    Output: OOD score s(x)s(x)
    // Stage 1: Self-supervised Pre-training via Masked Image Modeling (MIM)
    Initialize Vision Transformer (ViT) backbone fθf_\theta
    for each batch in DpreD_{\text{pre}} do
        Split images into patches and randomly mask a subset of patches
        Extract visual tokens for original patches using a discrete VAE tokenizer
        Update fθf_\theta to reconstruct discrete VAE tokens of the masked patches
    end for
    // Stage 2: Intermediate Fine-Tuning
    Fine-tune fθf_\theta with supervised classification on DpreD_{\text{pre}}
    // Stage 3: Target In-Distribution Fine-Tuning
    if task is one-class OOD detection (Nc=1N_c = 1) then
        Fine-tune fθf_\theta on DIDD_{\text{ID}} using cross-entropy with label smoothing (yLS=1−α+α/Ncy^{\text{LS}} = 1 - \alpha + \alpha / N_c)
    else
        Fine-tune fθf_\theta on DIDD_{\text{ID}} using standard multi-class cross-entropy loss
    end if
    // Stage 4: Feature Extraction and Scoring
    Extract penultimate layer feature embeddings z(x′)=fθ(x′)z(x') = f_\theta(x') for all training samples x′∈DIDx' \in D_{\text{ID}}
    for each class c∈{1,…,Nc}c \in \{1, \dots, N_c\} do
        Compute empirical class mean μc=1Nk∑i:yi=cz(xi)\mu_c = \frac{1}{N_k} \sum_{i: y_i = c} z(x_i)
    end for
    Compute tied covariance matrix Σ=1N∑c=1Nc∑i:yi=c(z(xi)−μc)(z(xi)−μc)T\Sigma = \frac{1}{N} \sum_{c=1}^{N_c} \sum_{i: y_i = c} (z(x_i) - \mu_c)(z(x_i) - \mu_c)^T
    Extract test feature z(x)=fθ(x)z(x) = f_\theta(x)
    Compute Mahalanobis distance metric s(x)=min⁡c∈{1,…,Nc}(z(x)−μc)TΣ−1(z(x)−μc)s(x) = \min_{c \in \{1, \dots, N_c\}} (z(x) - \mu_c)^T \Sigma^{-1} (z(x) - \mu_c)
    return s(x)s(x)
  2. Knowl 2 — Label Smoothing for One-Class Fine-Tuning

    model/method

    In one-class out-of-distribution (OOD) detection, all in-distribution (ID) training samples belong to a single class (Nc=1N_c = 1 or identical nominal labels). Standard cross-entropy fine-tuning on a single class leads to trivial cross-entropy loss values of zero and vanishing gradients once the classification output reaches an accuracy of 1, preventing the backbone network parameters from adapting to the specific ID distribution.

    To enable parameter updates during one-class fine-tuning, label smoothing is applied to the target labels:

    ycLS=yc(1−α)+αNc,c=1,2,…,Ncy_c^{\text{LS}} = y_c(1 - \alpha) + \frac{\alpha}{N_c}, \quad c = 1, 2, \dots, N_c

    where cc is the class index, NcN_c is the number of classes, ycy_c is the original one-hot indicator, and α∈[0,1]\alpha \in [0, 1] is the smoothing hyperparameter. When α=0\alpha = 0, ycLSy_c^{\text{LS}} preserves the original one-hot label, whereas α=1\alpha = 1 yields a uniform distribution. Because label smoothing produces a target probability strictly less than 1, the cross-entropy loss remains strictly positive even when predictions match the target class, forcing the network to continue updating parameters and effectively adapting feature representations to the target one-class ID data.

  3. Knowl 3 — Pretext Task Comparison for Out-of-Distribution Detection

    data/table

    Masked Image Modeling (MIM) pre-training outperforms classification pre-training and contrastive learning (MoCov3) pre-training for OOD detection when evaluated across multiple in-distribution (ID) and out-of-distribution (OOD) benchmark pairs on a Vision Transformer (ViT) pre-trained on ImageNet-22k.

    In-Distribution CIFAR-10 →\rightarrow CIFAR-100 →\rightarrow
    Out-of-Distribution SVHN CIFAR-100 LSUN Avg SVHN CIFAR-10 LSUN Avg
    Classification 98.3 98.6 98.6 98.5 78.0 93.5 88.6 86.7
    MoCov3 98.6 92.4 89.8 93.6 78.8 72.8 75.8 75.8
    MIM 99.8 99.4 99.9 99.7 96.5 98.3 96.3 97.0
    In-Distribution ImageNet-30 →\rightarrow
    Out-of-Distribution Dogs Places365 Flowers102 Pets Food Dtd Caltech256 Avg
    Classification 99.7 98.4 99.9 99.6 98.3 98.6 96.8 98.8
    MoCov3 88.2 82.0 99.3 81.1 71.4 91.3 88.5 86.0
    MIM 99.4 98.9 100.0 99.1 96.6 99.5 98.9 98.9

    All metrics report AUROC (%). Classification and contrastive pre-training encourage networks to learn discriminative shortcut patterns across categories rather than the full intrinsic pixel-level data distribution. Consequently, MIM pre-training significantly boosts detection performance, particularly on fine-grained and semantically close shifts (such as CIFAR-100 →\rightarrow CIFAR-10, where MIM achieves 98.3% AUROC compared to MoCov3's 72.8%).

  4. Knowl 4 — Multi-Class Out-of-Distribution Detection Performance

    data/table

    The MOOD framework achieves state-of-the-art results across various multi-class OOD benchmarks on CIFAR-10, CIFAR-100, ImageNet-30, and ImageNet-1k datasets, evaluated using the Area Under the Receiver Operating Characteristic Curve (AUROC, %).

    In-Distribution CIFAR-10 →\rightarrow CIFAR-100 →\rightarrow
    Method SVHN CIFAR-100 LSUN Avg SVHN CIFAR-10 LSUN Avg
    Baseline OOD 88.6 85.8 90.7 88.4 81.9 81.1 86.6 83.2
    ODIN 96.4 89.6 – 93.0 60.9 77.9 – 69.4
    Mahalanobis 99.4 90.5 – 95.0 94.5 55.3 – 74.9
    Residual Flows 99.1 89.4 – 94.3 97.5 77.1 – 87.3
    Gram Matrix 99.5 79.0 – 89.3 96.0 67.9 – 82.0
    Outlier Exposure 98.4 93.3 – 95.9 86.9 75.7 – 81.3
    Rotation loss 98.9 90.9 – 94.9 – – – –
    Contrastive loss 97.3 88.6 92.8 92.9 95.6 78.3 – 87.0
    CSI 97.9 92.2 97.7 95.9 – – – –
    SSD+ 99.9 93.4 98.4 97.2 98.2 78.3 79.8 85.4
    MOOD (Ours) 99.8±0.0_{\pm 0.0} 99.4±0.0_{\pm 0.0} 99.9±0.0_{\pm 0.0} 99.7 96.5±0.6_{\pm 0.6} 98.3±0.1_{\pm 0.1} 96.3±0.6_{\pm 0.6} 97.0
    In-Distribution ImageNet-1k →\rightarrow
    Method iNaturalist SUN Places Textures Avg
    Baseline OOD 87.6 78.3 76.8 74.5 79.3
    ODIN 89.4 83.9 80.7 76.3 82.6
    Energy 88.5 85.3 81.4 75.8 82.7
    Mahalanobis 46.3 65.2 64.5 72.1 62.0
    GradNorm (SOTA) 90.3 89.0 84.8 81.1 86.3
    MOOD (Ours) 86.9 89.8 88.5 91.3 89.1

    On CIFAR-100 →\rightarrow CIFAR-10, MOOD achieves 98.3% AUROC, outperforming SSD+ (78.3%) by +20.0%. On ImageNet-30, MOOD achieves an average AUROC of 98.9% (vs. 94.8% for CSI and 90.3% for Baseline OOD). On ImageNet-1k, MOOD achieves 89.1% average AUROC, surpassing GradNorm (86.3%) by +2.8%.

  5. Knowl 5 — One-Class Out-of-Distribution Detection Performance

    data/table

    In one-class OOD detection, one class from a dataset is designated as in-distribution (ID) while the remaining classes serve as out-of-distribution (OOD). MOOD consistently outperforms existing one-class detection methods across CIFAR-10, CIFAR-100 (super-classes), and ImageNet-30, measured in AUROC (%):

    Method CIFAR-10 (Avg) CIFAR-100 (Super-classes) ImageNet-30
    OC-SVM 58.8 63.1 –
    DeepSVDD 64.8 – –
    AnoGAN 61.8 – –
    OCGAN 65.7 – –
    Geom 86.0 78.7 –
    Rot 83.3 77.7 65.3
    Rot+Trans 90.1 79.8 77.9
    Rot+Attn – – 81.6
    Rot+Trans+Attn – – 84.8
    Rot+Trans+Attn+Resize – – 85.7
    GOAD 88.2 74.5 –
    CSI (Previous SOTA) 94.3 89.6 91.6
    MOOD (Ours) 97.8±0.4_{\pm 0.4} 94.8 92.0
    Improvement over SOTA +3.5% +5.2% +0.4%

    Class-wise AUROC on CIFAR-10 for MOOD vs. CSI: Plane (98.6% vs 89.9%), Car (99.3% vs 99.1%), Bird (94.3% vs 93.1%), Cat (93.2% vs 86.4%), Deer (98.1% vs 93.9%), Dog (96.5% vs 93.2%), Frog (99.3% vs 95.1%), Horse (99.0% vs 98.7%), Ship (98.8% vs 97.9%), Truck (97.8% vs 95.5%). MOOD's overall average across all one-class tasks reaches 94.9% (a 5.7% gain over CSI overall).

  6. Knowl 6 — Near-Distribution Semantically Similar ID-OOD Pair Detection

    data/table

    Near-distribution OOD detection involves ID and OOD classes with high semantic overlap (e.g., Cat vs Dog, Plane vs Automobile), which typically yield low detection AUROC (under 90%) with contrastive methods such as CSI. MOOD substantially improves detection on these difficult pairs without using outlier exposure.

    ID Class OOD Class CSI (AUROC %) MOOD (AUROC %) Improvement
    Plane Automobile 74.1 99.0 +24.9%
    Plane Ship 79.6 99.4 +19.8%
    Plane Truck 82.8 98.5 +15.7%
    Bird Horse 83.2 94.3 +11.1%
    Cat Deer 83.3 92.6 +9.3%
    Cat Dog 67.0 75.5 +8.5%
    Cat Frog 89.6 92.5 +2.9%
    Cat Horse 79.0 95.5 +16.5%
    Deer Horse 69.0 100.0 +31.0%
    Dog Deer 88.1 96.4 +8.3%
    Dog Horse 76.6 95.5 +18.9%
    Truck Automobile 72.3 87.8 +15.5%
    Average 78.7 93.9 +15.2%

    In multi-class OOD confusion settings at 95% True-Positive Rate (TPR), MOOD reduces the number of mistakenly classified OOD samples (e.g., CIFAR-100 Tiger images classified into CIFAR-10 Cat) from 48 under SSD+ to only 2, achieving an average 79% reduction in mistakenly classified OOD samples across semantically similar pairs.

  7. Knowl 7 — Zero-Shot MOOD vs. Few-Shot Outlier Exposure Detection

    data/table

    Existing state-of-the-art near-distribution OOD detection methods (such as R50+ViT) use few-shot Outlier Exposure (OE), where true OOD examples are introduced during training to formulate detection as a supervised task. MOOD operates in a zero-shot OOD regime (no OOD training examples) and exceeds the performance of OE-trained models.

    Method # OOD samples per class AUROC (%)
    R50+ViT (SOTA) 0 98.52
    R50+ViT (SOTA) 1 98.96
    R50+ViT (SOTA) 2 99.11
    R50+ViT (SOTA) 3 99.17
    R50+ViT (SOTA) 10 99.29
    MOOD (Ours) 0 99.41

    MOOD with 0-shot outlier exposure achieves 99.41% AUROC on near-distribution OOD detection, surpassing the 10-shot-per-class outlier exposure baseline (99.29% AUROC) by +0.12%, demonstrating that reconstruction-based representation learning removes the need for exposing models to known OOD outliers.

  8. Knowl 8 — Vision Transformer Backbone vs Other Architectures for OOD Detection

    data/table

    A single Vision Transformer (ViT) trained with Masked Image Modeling (MIM) outperforms large hybrid convolutional-transformer models and convolutional networks on near-distribution OOD detection (CIFAR-100 as ID and CIFAR-10 as OOD):

    Model Architecture Fine-tuned Test Acc (%) AUROC (%)
    BiT R50 87.01 81.71
    BiT R101×\times3 91.55 90.10
    ViT 90.95 95.53
    MLP-Mixer 90.40 95.31
    R50 + ViT (Previous SOTA) 91.71 96.23
    MIM-ViT (MOOD) – 98.30

    R50+ViT doubles model size and inference time relative to ViT while achieving 96.23% AUROC. Applying MIM to a single ViT improves AUROC to 98.30% (+2.07% over R50+ViT), establishing that an appropriate reconstruction pretext task is more effective for OOD detection than increasing model complexity or combining architectures.

  9. Knowl 9 — Evaluation of OOD Detection Scoring Metrics with MIM Pre-training

    data/table

    When using Masked Image Modeling (MIM) representations, the Mahalanobis distance metric consistently outperforms classifier output-based metrics (Softmax confidence, Entropy, Energy, and GradNorm) across multi-class OOD benchmarks:

    Metric CIFAR-10 →\rightarrow Avg CIFAR-100 →\rightarrow Avg ImageNet-30 →\rightarrow Avg
    Softmax 88.4 83.2 90.3
    Entropy 98.4 92.2 88.3
    Energy 98.2 90.8 85.6
    GradNorm 93.9 62.6 77.5
    Mahalanobis Distance 99.7 97.0 98.9

    AUROC scores (%) demonstrate that feature space distance estimation via the Mahalanobis metric reliably captures the divergence between ID and OOD distributions learned by MIM.

  10. Knowl 10 — Ablation of Training Stages in Multi-Class MOOD Pipeline

    data/table

    Ablation of the three-stage training pipeline (MIM pre-training on ImageNet-21k, intermediate supervised fine-tuning on ImageNet-21k, and target ID fine-tuning) indicates that each successive stage provides compounding gains in multi-class OOD detection AUROC (%):

    Training Stage CIFAR-10 →\rightarrow CIFAR-100 →\rightarrow
    MIM-pt Inter-ft ID-ft SVHN CIFAR-100 LSUN Avg SVHN CIFAR-10 LSUN Avg
    ✓ 62.2 62.9 98.5 74.5 48.4 42.2 96.0 62.2
    ✓ ✓ 89.5 90.0 99.8 93.1 74.3 62.0 98.3 68.2
    ✓ ✓ 99.1 94.6 97.4 97.0 93.7 83.7 91.4 89.6
    ✓ ✓ ✓ 99.8 99.4 99.9 99.7 96.5 98.3 96.3 97.0
    Training Stage ImageNet-30 →\rightarrow
    MIM-pt Inter-ft ID-ft Dogs Places365 Flowers102 Pets Food Caltech256 Dtd Avg
    ✓ 60.2 82.7 28.6 41.9 72.5 42.2 29.4 51.1
    ✓ ✓ 100.0 97.9 99.9 99.6 97.1 96.9 98.2 98.2
    ✓ ✓ 91.3 97.0 95.1 93.8 99.3 84.0 95.4 92.9
    ✓ ✓ ✓ 99.4 98.9 100.0 99.1 96.6 99.5 98.9 98.9

    Combining all three stages yields the highest average AUROC (99.7% on CIFAR-10, 97.0% on CIFAR-100, and 98.9% on ImageNet-30).

Coverage note — None was omitted; all contributed algorithms, model modifications (label smoothing in one-class fine-tuning), empirical benchmarks (one-class, multi-class, near-distribution, outlier exposure), architectural comparisons, scoring metric evaluations, and ablation studies have been included as self-contained knowls.

References

  1. 1.Amir Adler, Michael Elad, Yacov Hel-Or, and Ehud Rivlin. Sparse coding with anomaly detection. Journal of Signal Processing Systems, 79(2):179–188, 2015. 2
  2. 2.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 2, 3, 4
  3. 3.Liron Bergman and Yedid Hoshen. Classification-based anomaly detection for general data. arXiv preprint arXiv:2005.02359, 2020. 5, 6
  4. 4.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European conference on computer vision, pages 446–461. Springer, 2014. 6
  5. 5.Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, pages 1721–1730, 2015. 1
  6. 6.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 2
  7. 7.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758, 2021. 2
  8. 8.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9640–9649, October 2021. 3
  9. 9.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014. 6
  10. 10.Gaudenz Danuser and Markus Stricker. Parametric model fitting: From inlier characterization to outlier detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(3):263–280, 1998. 2
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2, 3
  12. 12.Laurent Dinh, David Krueger, and Yoshua Bengio. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014. 2
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2, 3
  14. 14.Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physical-world attacks on deep learning visual classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1625–1634, 2018. 1
  15. 15.Stanislav Fort, Jie Ren, and Balaji Lakshminarayanan. Exploring the limits of out-of-distribution detection. Advances in Neural Information Processing Systems, 34:7068–7081, 2021. 1, 2, 3, 4, 6, 7, 8
  16. 16.Izhak Golan and Ran El-Yaniv. Deep anomaly detection using geometric transformations. Advances in neural information processing systems, 31, 2018. 5, 6
  17. 17.Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel. Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1705–1714, 2019. 2
  18. 18.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020. 2
  19. 19.Gregory Griffin, Alex Holub, and Pietro Perona. Caltech-256 object category dataset. 2007. 6
  20. 20.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 2, 3
  21. 21.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020. 2
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3
  23. 23.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. arXiv preprint arXiv:1610.02136, 2016. 3, 4, 5, 6, 7
  24. 24.Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. arXiv preprint arXiv:1812.04606, 2018. 7
  25. 25.Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song. Using self-supervised learning can improve model robustness and uncertainty. Advances in neural information processing systems, 32, 2019. 5, 6, 7
  26. 26.Rui Huang, Andrew Geng, and Yixuan Li. On the importance of gradients for detecting distributional shifts in the wild. Advances in Neural Information Processing Systems, 34:677–689, 2021. 4, 5, 6, 7
  27. 27.Nathalie Japkowicz. Concept learning in the absence of counterexamples: An autoassociation-based approach to classification. Rutgers The State University of New Jersey-New Brunswick, 1999. 2
  28. 28.Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Fei-Fei Li. Novel dataset for fine-grained image categorization: Stanford dogs. In Proc. CVPR workshop on fine-grained visual categorization (FGVC), volume 2. Citeseer, 2011. 6
  29. 29.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 18661–18673. Curran Associates, Inc., 2020. 2, 7
  30. 30.Durk P Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. Advances in neural information processing systems, 31, 2018. 2
  31. 31.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In European conference on computer vision, pages 491–507. Springer, 2020. 3
  32. 32.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 3, 5
  33. 33.Gukyeong Kwon, Mohit Prabhushankar, Dogancan Temel, and Ghassan AlRegib. Backpropagated gradient representations for anomaly detection. In European Conference on Computer Vision, pages 206–226. Springer, 2020. 2
  34. 34.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. Advances in neural information processing systems, 31, 2018. 4, 5, 7
  35. 35.Shiyu Liang, Yixuan Li, and Rayadurgam Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. arXiv preprint arXiv:1706.02690, 2017. 7
  36. 36.Shiyu Liang, Yixuan Li, and Rayadurgam Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. arXiv preprint arXiv:1706.02690, 2017. 5
  37. 37.Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou. Isolation forest. In 2008 eighth ieee international conference on data mining, pages 413–422. IEEE, 2008. 2
  38. 38.Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. Advances in neural information processing systems, 33:21464–21475, 2020. 4, 5, 7
  39. 39.Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, Ian Goodfellow, and Brendan Frey. Adversarial autoencoders. arXiv preprint arXiv:1511.05644, 2015. 2
  40. 40.Gerhard Munz, Sa Li, and Georg Carle. Traffic anomaly detection using k-means clustering. In GI/ITG Workshop MMBnet, volume 7, page 9, 2007. 2
  41. 41.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. 2011. 5
  42. 42.M-E Nilsback and Andrew Zisserman. A visual vocabulary for flower classification. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06), volume 2, pages 1447–1454. IEEE, 2006. 6
  43. 43.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012. 6
  44. 44.Pramuditha Perera, Ramesh Nallapati, and Bing Xiang. Ocgan: One-class novelty detection using gans with constrained latent representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2898–2906, 2019. 6
  45. 45.Clifton Phua, Vincent Lee, Kate Smith, and Ross Gayler. A comprehensive survey of data mining-based fraud detection research. arXiv preprint arXiv:1009.6119, 2010. 1
  46. 46.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 2
  47. 47.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8821–8831. PMLR, 18–24 Jul 2021. 2
  48. 48.Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Müller, and Marius Kloft. Deep one-class classification. In International conference on machine learning, pages 4393–4402. PMLR, 2018. 6
  49. 49.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis., 2015. 2, 4, 5
  50. 50.Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. Hidden trigger backdoor attacks. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07):11957–11965, Apr. 2020. 1, 3
  51. 51.Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. Hidden trigger backdoor attacks. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 11957–11965, 2020. 1, 2, 3
  52. 52.Thomas Schlegl, Philipp Seeböck, Sebastian M Waldstein, Ursula Schmidt-Erfurth, and Georg Langs. Unsupervised anomaly detection with generative adversarial networks to guide marker discovery. In International conference on information processing in medical imaging, pages 146–157. Springer, 2017. 6
  53. 53.Vikash Sehwag, Mung Chiang, and Prateek Mittal. Ssd: A unified framework for self-supervised outlier detection. arXiv preprint arXiv:2103.12051, 2021. 1, 2, 3, 4, 6, 7, 8
  54. 54.Chandramouli Shama Sastry and Sageev Oore. Detecting out-of-distribution examples with indistribution examples and gram matrices. arXiv e-prints, pages arXiv–1912, 2019. 7
  55. 55.Chence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang, Ming Zhang, and Jian Tang. Graphaf: a flow-based autoregressive model for molecular graph generation. arXiv preprint arXiv:2001.09382, 2020. 2
  56. 56.Iwan Syarif, Adam Prugel-Bennett, and Gary Wills. Unsupervised clustering approach for network anomaly detection. In International conference on networked digital technologies, pages 135–145. Springer, 2012. 2
  57. 57.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-first AAAI conference on artificial intelligence, 2017. 4
  58. 58.Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin. Csi: Novelty detection via contrastive learning on distributionally shifted instances. Advances in neural information processing systems, 33:11839–11852, 2020. 1, 2, 3, 5, 6, 7, 8
  59. 59.Jing Tian, Michael H Azarian, and Michael Pecht. Anomaly detection using self-organizing maps-based k-nearest neighbor algorithm. In PHM Society European Conference, volume 2, 2014. 2
  60. 60.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. Advances in neural information processing systems, 29, 2016. 2
  61. 61.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8769–8778, 2018. 6
  62. 62.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011. 6
  63. 63.Haohan Wang, Xindi Wu, Zeyi Huang, and Eric P Xing. High-frequency component helps explain the generalization of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8684–8694, 2020. 2
  64. 64.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485–3492. IEEE, 2010. 6
  65. 65.Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey. arXiv preprint arXiv:2110.11334, 2021. 2
  66. 66.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019. 2
  67. 67.Shuangfei Zhai, Yu Cheng, Weining Lu, and Zhongfei Zhang. Deep structured energy based models for anomaly detection. In International conference on machine learning, pages 1100–1109. PMLR, 2016. 2
  68. 68.Bangzuo Zhang and Wanli Zuo. Learning from positive and unlabeled examples: A survey. In 2008 International Symposiums on Information Processing, pages 650–654. IEEE, 2008. 2
  69. 69.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In European conference on computer vision, pages 649–666. Springer, 2016. 2
  70. 70.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464, 2017. 6
  71. 71.Ev Zisselman and Aviv Tamar. Deep residual flow for novelty detection. 2020. 7
  72. 72.Bo Zong, Qi Song, Martin Renqiang Min, Wei Cheng, Cristian Lumezanu, Daeki Cho, and Haifeng Chen. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In International conference on learning representations, 2018. 2

Citation

MLA
Li, J., et al. “Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling Is All You Need”. arXiv, 2023, http://arxiv.org/abs/2302.02615v2.
APA
Li, J., Chen, P., Yu, S., He, Z., Liu, S., & Jia, J. (2023). Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need. arXiv. http://arxiv.org/abs/2302.02615v2
Chicago
Li, J., P. Chen, S. Yu, Z. He, S. Liu, and J. Jia. 2023. “Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling Is All You Need”. arXiv. http://arxiv.org/abs/2302.02615v2.
Harvard
Li, J. et al. (2023) “Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.02615v2.
Vancouver
1. Li J, Chen P, Yu S, He Z, Liu S, Jia J (2023) Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need. arXiv

BibTeX

@article{li2023rethinking,
  title = {Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need},
  author = {Li, Jingyao and Chen, Pengguang and Yu, Shaozuo and He, Zexin and Liu, Shu and Jia, Jiaya},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.02615v2},
  eprint = {2302.02615}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE