Delving into Out-of-Distribution Detection with Vision-Language Representations

Yifei MingZiyang CaiJiuxiang GuYiyou SunWei LiYixuan Li

article2022NeurIPS306 citationsOutstanding Paper Award Honorable Mention

Proposes Maximum Concept Matching, a training-free zero-shot out-of-distribution detection method that measures alignment between visual inputs and textual concept embeddings in vision-language models to reliably identify novel categories without needing candidate anomaly labels.

Listen

Deploying machine learning models in real-world settings presents significant operational risks when systems encounter unexpected, out-of-distribution inputs. Standard image classifiers typically operate in a closed-world setup, forcing every unfamiliar object into a known category and producing dangerously overconfident errors. Most existing detection solutions rely exclusively on visual features, requiring costly, task-specific retraining or fine-tuning that fails to generalize across diverse operational tasks.

The article introduces and evaluates Maximum Concept Matching, a zero-shot detection framework that leverages joint vision-language representations to identify unfamiliar inputs without task-specific training. The approach defines class concepts using textual descriptions and measures the alignment between an input image and these textual prototypes. By applying a temperature-scaled softmax normalization to visual-textual similarity scores, the method magnifies the contrast between known and unknown categories. This framework is evaluated against standard and fine-tuned models across large-scale vision benchmarks, including ImageNet-1k, fine-grained datasets such as Stanford Cars and Food-101, and specialized stress tests designed to assess performance against semantically similar and spurious background distractors.

The findings establish that this zero-shot multimodal approach matches or outperforms existing methods requiring extensive task-specific training. On the large-scale ImageNet-1k benchmark, Maximum Concept Matching achieved a 91.49% area under the curve score, surpassing prominent fine-tuned baselines. In semantically difficult scenarios involving highly similar categories, the method outperformed traditional vision-only distance metrics by 13.1% in area under the curve and reduced false positive rates by up to 73.32%. Furthermore, the framework demonstrated high resilience against spurious correlations, achieving a 5.87% false positive rate compared to 39.57% for fine-tuned visual baselines, while prompt ensembling further lowered false positive rates to 35.23%.

These results demonstrate that joint vision-language models can substantially reduce the engineering overhead, compute costs, and storage burdens associated with maintaining separate, specialized classifiers for different tasks. A single pre-trained foundation model can serve as a dependable, training-free safety filter across changing operational contexts. When deploying multimodal models for safety filtering, organizations should adopt temperature-scaled concept matching rather than relying on raw cosine similarities or expensive fine-tuning pipelines. Future efforts should explore expanding this zero-shot detection framework to other multimodal architectures and non-image modalities, while noting that performance remains tied to the underlying representation quality of the pre-trained foundation model.

Cover for Delving into Out-of-Distribution Detection with Vision-Language Representations

Abstract

Recognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or language), leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of OOD detection from a single-modal to a multi-modal regime. Particularly, we propose Maximum Concept Matching (MCM), a simple yet effective zero-shot OOD detection method based on aligning visual features with textual concepts. We contribute in-depth analysis and theoretical insights to understand the effectiveness of MCM. Extensive experiments demonstrate that MCM achieves superior performance on a wide variety of real-world tasks. MCM with vision-language features outperforms a common baseline with pure visual features on a hard OOD task with semantically similar classes by 13.1% (AUROC). Code is available at https://github.com/deeplearning-wisc/MCM.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 OOD Detection via Concept Matching
  • 4 A Comprehensive Analysis of MCM
  • 4.1 Datasets and Implementation Details
  • 4.2 Main Results
  • 5 Discussion: A Closer Look at MCM
  • 6 Related Works
  • 7 Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Maximum Concept Matching for zero-shot OOD detection

    model/method

    Maximum Concept Matching (MCM) detects whether an image belongs to one of the known classes by comparing its visual embedding with text embeddings of the known class names. For a task with ID label set Yin={y1,…,yK}Y_{\mathrm{in}}=\{y_1,\ldots,y_K\}, each label is inserted into the prompt “This is a photo of a ⟨\langlelabel⟩\rangle.” The text encoder TT maps each prompt tit_i to a concept prototype T(ti)∈RdT(t_i)\in\mathbb{R}^d, and the image encoder II maps an image xx to I(x)∈RdI(x)\in\mathbb{R}^d.

    The cosine matching score for class ii is

    si(x)=I(x)TT(ti)∥I(x)∥2 ∥T(ti)∥2.s_i(x)=\frac{I(x)^{\mathsf T}T(t_i)}{\|I(x)\|_2\,\|T(t_i)\|_2}.

    With temperature τ>0\tau>0, the MCM confidence score is the largest softmax probability over the KK concept prototypes:

    SMCM(x;Yin,T,I)=max⁡i∈{1,…,K}exp⁡(si(x)/τ)∑j=1Kexp⁡(sj(x)/τ).S_{\mathrm{MCM}}(x;Y_{\mathrm{in}},T,I)=\max_{i\in\{1,\ldots,K\}}\frac{\exp(s_i(x)/\tau)}{\sum_{j=1}^{K}\exp(s_j(x)/\tau)}.

    An image is classified as ID when SMCM(x)≥λS_{\mathrm{MCM}}(x)\geq\lambda and as OOD otherwise, where λ\lambda is an ID/OOD threshold. For an image accepted as ID, the predicted class is y^=arg⁡max⁡isi(x)\hat y=\arg\max_i s_i(x). The method uses no ID training images, OOD samples, candidate OOD labels, or downstream fine-tuning.

  2. Knowl 2 — Task-relative and OOD-agnostic operating regime

    definition

    The paper defines ID and OOD relative to the classification task rather than relative to the data used to pre-train the vision-language model. Given a pre-trained image encoder II, text encoder TT, and task label set YinY_{\mathrm{in}}, an ID image belongs to one of the labels in YinY_{\mathrm{in}}; an OOD image belongs to none of them. The detector must both reject images outside YinY_{\mathrm{in}} and assign accepted images to a known class.

    MCM is zero-shot because it requires only the names of the task classes and the fixed pre-trained encoders. Changing to a new task requires constructing new text prompts, not training a new classifier or OOD detector. Consequently, the same model can support many tasks, does not require prior knowledge of what OOD categories will occur, and scales to large label sets and high-resolution images.

  3. Knowl 3 — Why softmax scaling is useful for CLIP-like representations

    model/method

    For contrastive vision-language models such as CLIP, OOD images tend to have approximately uniform cosine similarities to the ID text prototypes. Their largest cosine similarity can therefore be close to the largest similarity of an ID image, especially when the number of ID classes is large. ID images nevertheless tend to have a larger gap between their best-matching prototype and the remaining prototypes.

    MCM applies softmax to magnify this difference: a concentrated ID similarity vector produces a high maximum probability, whereas a nearly uniform OOD similarity vector produces a lower maximum probability. This role is distinct from softmax confidence in a classifier trained with cross-entropy. Cross-entropy training already forces the ground-truth logit above the other logits, so softmax can amplify overconfident predictions and reduce ID/OOD separability. In the CLIP setting, softmax is instead a post-hoc transformation that sharpens the uniform-versus-concentrated structure produced by contrastive alignment.

  4. Knowl 4 — Formal condition and guarantee for softmax-based separability

    theoretical result

    Let Yin={y1,…,yK}Y_{\mathrm{in}}=\{y_1,\ldots,y_K\} be the ID label set, let si(x)s_i(x) be the cosine similarity between image xx and ID concept prototype ii, and let QxQ_x be the OOD image distribution. For an OOD image, define y^=arg⁡max⁡isi(x)\hat y=\arg\max_i s_i(x) and define y^2=arg⁡max⁡i≠y^si(x)\hat y_2=\arg\max_{i\neq\hat y}s_i(x). The paper assumes that there is a constant δ>0\delta>0 such that OOD similarities are sufficiently uniform among the non-maximal prototypes:

    Qx ⁣(1K−1∑i≠y^[sy^2(x)−si(x)]<δ)=1.Q_x\!\left(\frac{1}{K-1}\sum_{i\neq\hat y}\left[s_{\hat y_2}(x)-s_i(x)\right]<\delta\right)=1.

    Let λ\lambda be the detection threshold applied to the softmax MCM score, let λwo\lambda^{\mathrm{wo}} be the threshold applied to the unscaled score max⁡isi(x)\max_i s_i(x), and let FPR(τ,λ)\mathrm{FPR}(\tau,\lambda) and FPRwo(λwo)\mathrm{FPR}^{\mathrm{wo}}(\lambda^{\mathrm{wo}}) denote the corresponding OOD false-positive rates. Under the uniformity assumption, the paper gives the temperature bound

    T=λ(K−1)(λwo+δ−sy^2)Kλ−1,\mathcal{T}=\frac{\lambda(K-1)\left(\lambda^{\mathrm{wo}}+\delta-s_{\hat y_2}\right)}{K\lambda-1},

    with the bound understood where its denominator is positive. For every temperature τ>T\tau>\mathcal{T}, softmax scaling satisfies

    FPR(τ,λ)≤FPRwo(λwo).\mathrm{FPR}(\tau,\lambda)\leq \mathrm{FPR}^{\mathrm{wo}}(\lambda^{\mathrm{wo}}).

    Thus, under the stated OOD-similarity condition, a sufficiently large temperature provides a formal guarantee that the softmax-scaled MCM detector is no worse in false-positive rate than maximum cosine similarity without softmax.

  5. Knowl 5 — Evaluation design for realistic and hard OOD tasks

    experimental setup

    The evaluation uses CLIP as the pre-trained vision-language model, primarily CLIP-B/16 with a ViT-B/16 image encoder and a Transformer text encoder; CLIP-L/14 and RN50x4 are also evaluated. Unless otherwise specified, the MCM temperature is τ=1\tau=1.

    ID datasets include CUB-200, Stanford-Cars, Food-101, Oxford-Pet, and ImageNet variants. OOD datasets include iNaturalist, SUN, Places, and Texture, with non-overlapping categories relative to each ID dataset. ImageNet-10 provides a high-resolution analogue of a ten-class benchmark, ImageNet-20 contains semantically similar classes for hard OOD testing, and ImageNet-100 and ImageNet-1k test scaling with the number of ID classes. The Waterbirds benchmark supplies spurious OOD images that share ID backgrounds but contain different object labels.

    Performance is measured by FPR95, the OOD false-positive rate when the ID true-positive rate is 95%, AUROC, and ID classification accuracy. MCM uses no downstream training; comparison methods may use fine-tuning, linear probing, or other task-specific training.

  6. Knowl 6 — MCM performance across diverse ID datasets

    data/table

    The zero-shot CLIP-B/16 evaluation compares MCM across seven ID datasets and four OOD datasets. Each cell reports FPR95 followed by AUROC; the average is over the four OOD datasets. The results show especially strong detection for fine-grained datasets with few training images per class, while the larger ImageNet-100 task is more difficult.

    Could not parse LaTeX table

    For Stanford-Cars, MCM achieves an average FPR95 of 0.08%0.08\% and AUROC of 99.89%99.89\% without training. For Food-101, fine-tuning can raise ID accuracy from 86.3%86.3\% to 92.5%92.5\%, but the resulting MSP detector is approximately tied with MCM in AUROC, 99.5%99.5\% versus 99.4%99.4\%.

  7. Knowl 7 — Scaling comparison on ImageNet-1k

    data/table

    For ImageNet-1k as the ID dataset, MCM is compared with methods that require training or fine-tuning and with zero-shot baselines. Every entry reports FPR95 followed by AUROC on iNaturalist, SUN, Places, Texture, and their average. MCM with CLIP-L obtains the best average AUROC among the listed methods while requiring no downstream training.

    Could not parse LaTeX table

    CLIP-L reduces average FPR95 by 4.574.57 percentage points relative to CLIP-B and reaches 73.28%73.28\% zero-shot ID accuracy, a 6.276.27-point improvement. MCM with CLIP-L exceeds MOS in average AUROC, 91.49%91.49\% versus 90.11%90.11\%, despite MOS requiring BiT fine-tuning. Relative to an MSP detector using the same CLIP-L image features, MCM improves average FPR95 by 15.5415.54 percentage points.

  8. Knowl 8 — Robustness to semantically similar and spurious OOD inputs

    data/table

    MCM is evaluated on three hard OOD settings: ImageNet-10 as ID with semantically similar ImageNet-20 as OOD, ImageNet-20 as ID with ImageNet-10 as OOD, and Waterbirds with spurious OOD images that share ID backgrounds but have different objects. The table reports FPR95 followed by AUROC.

    Could not parse LaTeX table

    MCM improves over the visual Mahalanobis detector by 73.3273.32 percentage points in FPR95 for ImageNet-10 ID versus ImageNet-20 OOD and by 30.1230.12 points in the reverse direction. On Waterbirds, MCM has FPR95 5.87%5.87\%, far below the fine-tuned MSP value of 39.57%39.57\%. The results indicate that aligned vision-language features remain effective when OOD categories are semantically close to ID categories or when OOD images exploit background correlations.

  9. Knowl 9 — Empirical effect of temperature and softmax scaling

    empirical result

    On ImageNet-100 as ID and iNaturalist as OOD, directly thresholding the maximum cosine similarity gives FPR95 of approximately 40.75%40.75\%. Applying the MCM softmax with temperature τ=0.01\tau=0.01 reduces FPR95 to 31.53%31.53\%, while τ=1\tau=1 reduces it to 18.13%18.13\% and τ=10\tau=10 gives 18.45%18.45\%. Thus, moderate temperature scaling improves FPR95 by 22.622.6 percentage points relative to no softmax.

    The empirical values also satisfy the paper’s theoretical temperature condition approximately: with K=100K=100, λwo≈0.26\lambda^{\mathrm{wo}}\approx0.26, δ≈0.03\delta\approx0.03, sy^2≈0.23s_{\hat y_2}\approx0.23, and λ≈0.011\lambda\approx0.011, the lower bound is estimated as T≈0.65\mathcal{T}\approx0.65. The default τ=1\tau=1 exceeds this bound.

    The same pattern appears on larger tasks. Relative to maximum cosine similarity, MCM reduces average FPR95 from 13.93%13.93\% to 2.60%2.60\% for ImageNet-20 and from 69.08%69.08\% to 42.74%42.74\% for ImageNet-1k, while easy fine-grained tasks show little difference between the two scores.

  10. Knowl 10 — Vision-language prototypes outperform visual Mahalanobis prototypes

    empirical result

    MCM and Mahalanobis detection are compared using the same CLIP-B image encoder and visual features from its penultimate layer. Mahalanobis constructs class prototypes from visual ID features and requires estimating an inverse covariance matrix; MCM instead uses text embeddings of the class names as concept prototypes.

    On ImageNet-1k, averaged over iNaturalist, SUN, Places, and Texture, MCM achieves 90.77%90.77\% AUROC, whereas the visual Mahalanobis score achieves 73.14%73.14\%. The comparison supports the paper’s conclusion that the aligned textual information in vision-language representations provides more useful class prototypes for OOD detection than visual features alone. The paper also notes that MCM avoids the covariance inversion required by Mahalanobis, which can be expensive or inaccurate when ID samples are scarce or the number of classes is large.

Coverage note — The secondary prompt-ensembling ablation is omitted because it is a sensitivity analysis rather than a load-bearing contribution; no other substantial contributed material is omitted.

References

  1. 1.Udit Arora, William Huang, and He He. Types of out-of-distribution texts and how to detect them. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  2. 2.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Conference on Neural Information Processing Systems (NeurIPS), 2019.
  3. 3.Sara Beery, Grant Van Horn, and Pietro Perona. Recognition in terra incognita. In The European Conference on Computer Vision (ECCV), 2018.
  4. 4.Abhijit Bendale and Terrance E Boult. Towards open set deep networks. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2016.
  5. 5.Julian Bitterwolf, Alexander Meinke, Maximilian Augustin, and Matthias Hein. Breaking down out-of-distribution detection: Many methods based on ood training data estimate a combination of the same core quantities. In International Conference on Machine Learning, pages 2041–2074. PMLR, 2022.
  6. 6.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In The European Conference on Computer Vision (ECCV), 2014.
  7. 7.Mu Cai and Yixuan Li. Out-of-distribution detection via frequency-regularized generative models. In Proceedings of IEEE/CVF Winter Conference on Applications of Computer Vision, 2023.
  8. 8.Derek Chen and Zhou Yu. Gold: improving out-of-scope detection in dialogues using data augmentation. arXiv preprint arXiv:2109.03079, 2021.
  9. 9.Jiefeng Chen, Yixuan Li, Xi Wu, Yingyu Liang, and Somesh Jha. Atom: Robustifying out-of-distribution detection using outlier mining. In The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2021.
  10. 10.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2014.
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2009.
  12. 12.Terrance DeVries and Graham W Taylor. Learning confidence for out-of-distribution detection in neural networks. arXiv preprint arXiv:1802.04865, 2018.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  14. 14.Xuefeng Du, Gabriel Gozum, Yifei Ming, and Yixuan Li. Siren: Shaping representations for detecting out-of-distribution objects. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  15. 15.Xuefeng Du, Zhaoning Wang, Mu Cai, and Yixuan Li. Vos: Learning what you don’t know by virtual outlier synthesis. In Proceedings of the International Conference on Learning Representations (ICLR), 2022.
  16. 16.Sepideh Esmaeilpour, Bing Liu, Eric Robertson, and Lei Shu. Zero-shot open set detection by extending clip. In The AAAI Conference on Artificial Intelligence (AAAI), 2022.
  17. 17.Zhen Fang, Yixuan Li, Jie Lu, Jiahua Dong, Bo Han, and Feng Liu. Is out-of-distribution detection learnable? In Advances in Neural Information Processing System (NeurIPS), 2022.
  18. 18.Christiane Fellbaum. Wordnet. In Theory and Applications of Ontology: Computer Applications, pages 231–243. Springer, 2010.
  19. 19.Stanislav Fort, Jie Ren, and Balaji Lakshminarayanan. Exploring the limits of out-of-distribution detection. In Conference on Neural Information Processing Systems (NeurIPS), 2021.
  20. 20.ZongYuan Ge, Sergey Demyanov, Zetao Chen, and Rahil Garnavi. Generative openmax for multi-class open set classification. arXiv preprint arXiv:1707.07418, 2017.
  21. 21.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations (ICLR), 2019.
  22. 22.Jiuxiang Gu, Jason Kuen, Shafiq Joty, Jianfei Cai, Vlad Morariu, Handong Zhao, and Tong Sun. Self-supervised relationship probing. In Conference on Neural Information Processing Systems (NeurIPS), 2020.
  23. 23.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2016.
  24. 24.Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf. Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 41–50, 2019.
  25. 25.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In International Conference on Learning Representations (ICLR), 2017.
  26. 26.Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. Pretrained transformers improve out-of-distribution robustness. In Association for Computational Linguistics (ACL), 2020.
  27. 27.Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. In International Conference on Learning Representations (ICLR), 2018.
  28. 28.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  29. 29.Yen-Chang Hsu, Yilin Shen, Hongxia Jin, and Zsolt Kira. Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2020.
  30. 30.Yibo Hu and Latifur Khan. Uncertainty-aware reliable text classification. In SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2021.
  31. 31.Rui Huang, Andrew Geng, and Yixuan Li. On the importance of gradients for detecting distributional shifts in the wild. In Conference on Neural Information Processing Systems (NeurIPS), 2021.
  32. 32.Rui Huang and Yixuan Li. Mos: Towards scaling out-of-distribution detection for large semantic space. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2021.
  33. 33.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), 2021.
  34. 34.Di Jin, Shuyang Gao, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. Towards textual out-of-domain detection without in-domain labels. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022.
  35. 35.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning (ICML), 2021.
  36. 36.Polina Kirichenko, Pavel Izmailov, and Andrew G Wilson. Why normalizing flows fail to detect out-of-distribution data. Conference on Neural Information Processing Systems (NeurIPS), 2020.
  37. 37.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning (ICML), 2021.
  38. 38.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. The European Conference on Computer Vision (ECCV), 2020.
  39. 39.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), Sydney, Australia, 2013.
  40. 40.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  41. 41.Ya Le and Xuan Yang. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.
  42. 42.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Conference on Neural Information Processing Systems (NeurIPS), 2018.
  43. 43.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019.
  44. 44.Xiaoya Li, Jiwei Li, Xiaofei Sun, Chun Fan, Tianwei Zhang, Fei Wu, Yuxian Meng, and Jun Zhang. kfolden: k-fold ensemble for out-of-distribution detection. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  45. 45.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In International Conference on Learning Representations (ICLR), 2022.
  46. 46.Shiyu Liang, Yixuan Li, and Rayadurgam Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. In International Conference on Learning Representations (ICLR), 2018.
  47. 47.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In The European Conference on Computer Vision (ECCV), 2014.
  48. 48.Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. In Conference on Neural Information Processing Systems (NeurIPS), 2020.
  49. 49.Yifei Ming, Ying Fan, and Yixuan Li. Poem: Out-of-distribution detection with posterior sampling. In International Conference on Machine Learning (ICML), 2022.
  50. 50.Yifei Ming, Hang Yin, and Yixuan Li. On the impact of spurious correlation for out-of-distribution detection. The AAAI Conference on Artificial Intelligence (AAAI), 2022.
  51. 51.Ron Mokady, Amir Hertz, and Amit H Bermano. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734, 2021.
  52. 52.Peyman Morteza and Yixuan Li. Provable guarantees for understanding out-of-distribution detection. The AAAI Conference on Artificial Intelligence (AAAI), 2022.
  53. 53.Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, Dilan Gorur, and Balaji Lakshminarayanan. Do deep generative models know what they don’t know? In International Conference on Learning Representations (ICLR), 2019.
  54. 54.Lawrence Neal, Matthew Olson, Xiaoli Fern, Weng-Keen Wong, and Fuxin Li. Open set learning with counterfactual images. In The European Conference on Computer Vision (ECCV), 2018.
  55. 55.Edwin G. Ng, Bo Pang, Piyush Sharma, and Radu Soricut. Understanding guided image captioning performance across domains. arXiv preprint arXiv:2012.02339, 2020.
  56. 56.Poojan Oza and Vishal M Patel. C2ae: Class conditioned auto-encoder for open-set recognition. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2019.
  57. 57.Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. Cats and dogs. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2012.
  58. 58.Alexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova, and Irina Piontkovskaya. Revisiting mahalanobis distance for transformer-based out-of-domain detection. In The AAAI Conference on Artificial Intelligence (AAAI), 2021.
  59. 59.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  60. 60.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  61. 61.Jie Ren, Peter J Liu, Emily Fertig, Jasper Snoek, Ryan Poplin, Mark Depristo, Joshua Dillon, and Balaji Lakshminarayanan. Likelihood ratios for out-of-distribution detection. In Conference on Neural Information Processing Systems (NeurIPS), 2019.
  62. 62.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. " why should i trust you?" explaining the predictions of any classifier. In SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2016.
  63. 63.Abhijit Guha Roy, Jie Ren, Shekoofeh Azizi, Aaron Loh, Vivek Natarajan, Basil Mustafa, Nick Pawlowski, Jan Freyberg, Yuan Liu, Zach Beaver, et al. Does your dermatology classifier know what it doesn’t know? detecting the long-tail of unseen conditions. Medical Image Analysis, 75:102274, 2022.
  64. 64.Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. In International Conference on Learning Representations (ICLR), 2019.
  65. 65.Chandramouli Shama Sastry and Sageev Oore. Detecting out-of-distribution examples with Gram matrices. In International Conference on Machine Learning (ICML), 2020.
  66. 66.Vikash Sehwag, Mung Chiang, and Prateek Mittal. Ssd: A unified framework for self-supervised outlier detection. In International Conference on Learning Representations (ICLR), 2021.
  67. 67.Joan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia, José F. Núñez, and Jordi Luque. Input complexity and out-of-distribution detection with likelihood-based generative models. In International Conference on Learning Representations (ICLR), 2020.
  68. 68.Yilin Shen, Yen-Chang Hsu, Avik Ray, and Hongxia Jin. Enhancing the generalization for intent classification and out-of-domain detection in slu. arXiv preprint arXiv:2106.14464, 2021.
  69. 69.Yiyou Sun, Chuan Guo, and Yixuan Li. React: Out-of-distribution detection with rectified activations. In Conference on Neural Information Processing Systems (NeurIPS), 2021.
  70. 70.Yiyou Sun and Yixuan Li. Dice: Leveraging sparsification for out-of-distribution detection. In Proceedings of European Conference on Computer Vision (ECCV), 2022.
  71. 71.Yiyou Sun, Yifei Ming, Xiaojin Zhu, and Yixuan Li. Out-of-distribution detection with deep nearest neighbors. In International Conference on Machine Learning (ICML), 2022.
  72. 72.Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin. Csi: Novelty detection via contrastive learning on distributionally shifted instances. In Conference on Neural Information Processing Systems (NeurIPS), 2020.
  73. 73.Ming Tan, Yang Yu, Haoyu Wang, Dakuo Wang, Saloni Potdar, Shiyu Chang, and Mo Yu. Out-of-domain detection for low-resource text classification tasks. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019.
  74. 74.Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumder, Soujanya Poria, Roger Zimmermann, and Amir Zadeh. Multimodal research in vision and language: A review of current and emerging trends. Information Fusion, 77:149–171, 2022.
  75. 75.Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv e-prints, pages arXiv–1807, 2018.
  76. 76.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2018.
  77. 77.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Conference on Neural Information Processing Systems (NeurIPS), 2017.
  78. 78.Sagar Vaze, Kai Han, Andrea Vedaldi, and Andrew Zisserman. Open-set recognition: A good closed-set classifier is all you need. In International Conference on Learning Representations (ICLR), 2022.
  79. 79.Roman Vershynin. High-dimensional probability: An introduction with applications in data science, volume 47. Cambridge university press, 2018.
  80. 80.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The caltech-ucsd birds-200-2011 dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011.
  81. 81.Feng Wang and Huaping Liu. Understanding the behaviour of contrastive loss. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2021.
  82. 82.Haoran Wang, Weitang Liu, Alex Bocchieri, and Yixuan Li. Can multi-label classification networks know what they don’t know? Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2021.
  83. 83.Haotao Wang, Aston Zhang, Yi Zhu, Shuai Zheng, Mu Li, Alex J Smola, and Zhangyang Wang. Partial and asymmetric contrastive learning for out-of-distribution detection in long-tailed recognition. In International Conference on Machine Learning (ICML), 2022.
  84. 84.Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning (ICML), 2020.
  85. 85.Jim Winkens, Rudy Bunel, Abhijit Guha Roy, Robert Stanforth, Vivek Natarajan, Joseph R Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Simon Kohl, et al. Contrastive training for improved out-of-distribution detection. arXiv preprint arXiv:2007.05566, 2020.
  86. 86.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2010.
  87. 87.Kai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. Noise or signal: The role of image backgrounds in object recognition. In International Conference on Learning Representations (ICLR), 2021.
  88. 88.Zhisheng Xiao, Qing Yan, and Yali Amit. Likelihood regret: An out-of-distribution detection score for variational auto-encoder. In Conference on Neural Information Processing Systems (NeurIPS), volume 33, 2020.
  89. 89.Keyang Xu, Tongzheng Ren, Shikun Zhang, Yihao Feng, and Caiming Xiong. Unsupervised out-of-domain detection via pre-trained transformers. In Association for Computational Linguistics (ACL), 2021.
  90. 90.Jingkang Yang, Haoqi Wang, Litong Feng, Xiaopeng Yan, Huabin Zheng, Wayne Zhang, and Ziwei Liu. Semantically coherent out-of-distribution detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021.
  91. 91.Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey. arXiv preprint arXiv:2110.11334, 2021.
  92. 92.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. Filip: Fine-grained interactive language-image pre-training. International Conference on Learning Representations (ICLR), 2021.
  93. 93.Li-Ming Zhan, Haowen Liang, Bo Liu, Lu Fan, Xiao-Ming Wu, and Albert Lam. Out-of-scope intent detection with self-supervision and discriminative training. Association for Computational Linguistics (ACL), 2021.
  94. 94.Renrui Zhang, Rongyao Fang, Peng Gao, Wei Zhang, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930, 2021.
  95. 95.Yinhe Zheng, Guanyi Chen, and Minlie Huang. Out-of-domain detection for natural language understanding in dialog systems. TASLP, 2020.
  96. 96.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2017.
  97. 97.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022.
  98. 98.Wenxuan Zhou, Fangyu Liu, and Muhao Chen. Contrastive out-of-distribution detection for pretrained transformers. Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  99. 99.Zhuotun Zhu, Lingxi Xie, and Alan Yuille. Object recognition with and without objects. In International Joint Conferences on Artificial Intelligence (IJCAI), 2017.

Citation

MLA
Ming, Y., et al. “Delving into Out-of-Distribution Detection with Vision-Language Representations”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 35087–102, https://proceedings.neurips.cc/paper_files/paper/2022/file/e43a33994a28f746dcfd53eb51ed3c2d-Paper-Conference.pdf.
APA
Ming, Y., Cai, Z., Gu, J., Sun, Y., Li, W., & Li, Y. (2022). Delving into Out-of-Distribution Detection with Vision-Language Representations. Advances in Neural Information Processing Systems, 35, 35087–35102. https://proceedings.neurips.cc/paper_files/paper/2022/file/e43a33994a28f746dcfd53eb51ed3c2d-Paper-Conference.pdf
Chicago
Ming, Y., Z. Cai, J. Gu, Y. Sun, W. Li, and Y. Li. 2022. “Delving into Out-of-Distribution Detection with Vision-Language Representations”. Advances in Neural Information Processing Systems 35: 35087–102. https://proceedings.neurips.cc/paper_files/paper/2022/file/e43a33994a28f746dcfd53eb51ed3c2d-Paper-Conference.pdf.
Harvard
Ming, Y. et al. (2022) “Delving into Out-of-Distribution Detection with Vision-Language Representations”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 35087–35102. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/e43a33994a28f746dcfd53eb51ed3c2d-Paper-Conference.pdf.
Vancouver
1. Ming Y, Cai Z, Gu J, Sun Y, Li W, Li Y (2022) Delving into Out-of-Distribution Detection with Vision-Language Representations. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 35087–35102

BibTeX

@inproceedings{ming2022delving,
  title = {Delving into Out-of-Distribution Detection with Vision-Language Representations},
  author = {Ming, Yifei and Cai, Ziyang and Gu, Jiuxiang and Sun, Yiyou and Li, Wei and Li, Yixuan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {35087-35102},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/e43a33994a28f746dcfd53eb51ed3c2d-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission