Built independently by an author, for readers. Read the story and support ChapterPal

keyword

open set recognition

Open set recognition is a machine learning classification paradigm where a model is trained on a set of known categories and must accurately classify test inputs belonging to those categories while simultaneously detecting and rejecting inputs that belong to novel, unseen categories. Unlike conventional closed set recognition, which operates under the assumption that all possible test instances belong exclusively to the classes present during training, open set recognition accounts for incomplete knowledge of the world. By bounding open space risk and establishing specialized decision boundaries or confidence metrics, the model avoids forcing unfamiliar inputs into predefined categories with false high confidence, making it suitable for realistic, open-world deployment.

8 items

Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP

Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP

Sepideh Esmaeilpour, Bing Liu, Eric Robertson, Lei Shu

Why you should read this

Proposes a zero-shot out-of-distribution detection method that extends CLIP with an image description generator to automatically generate candidate unknown labels and outperform supervised baselines without requiring in-distribution training data.

In an out-of-distribution (OOD) detection problem, samples of known classes (also called in-distribution classes) are used to train a special classifier. In testing, the classifier can (1) classify the test samples of known classes to their respective classes and also (2) detect samples that do not belong to any of the known classes (i.e., they belong to some unknown or OOD classes). This paper studies the problem of zero-shot out-of-distribution (OOD) detection, which still performs the same two tasks in testing but has no training except using the given known class names. This paper proposes a novel and yet simple method (called ZOC) to solve the problem. ZOC builds on top of the recent advances in zero-shot classification through multi-modal representation learning. It first extends the pre-trained language-vision model CLIP by training a text-based image description generator on top of CLIP. In testing, it uses the extended model to generate candidate unknown class names for each test sample and computes a confidence score based on both the known class names and candidate unknown class names for zero-shot OOD detection. Experimental results on 5 benchmark datasets for OOD detection demonstrate that ZOC outperforms the baselines by a large margin.

Added

2026-10-05

Language-driven Semantic Segmentation

Language-driven Semantic Segmentation

Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl

OrganizationsAppleCornell UniversityIntelUniversity of Copenhagen

Why you should read this

Proposes LSeg, a model that aligns per-pixel visual embeddings with text representations through contrastive learning, enabling zero-shot semantic segmentation of arbitrary unseen categories without additional training.

We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero- and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at this https URL.

Added

2026-10-05

Delving into Out-of-Distribution Detection with Vision-Language Representations

Delving into Out-of-Distribution Detection with Vision-Language Representations

Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun, Wei Li, Yixuan Li

OrganizationsAdobeGoogleUniversity of Wisconsin Madison

Why you should read this

Proposes Maximum Concept Matching, a training-free zero-shot out-of-distribution detection method that measures alignment between visual inputs and textual concept embeddings in vision-language models to reliably identify novel categories without needing candidate anomaly labels.

Recognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or language), leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of OOD detection from a single-modal to a multi-modal regime. Particularly, we propose Maximum Concept Matching (MCM), a simple yet effective zero-shot OOD detection method based on aligning visual features with textual concepts. We contribute in-depth analysis and theoretical insights to understand the effectiveness of MCM. Extensive experiments demonstrate that MCM achieves superior performance on a wide variety of real-world tasks. MCM with vision-language features outperforms a common baseline with pure visual features on a hard OOD task with semantically similar classes by 13.1% (AUROC). Code is available at https://github.com/deeplearning-wisc/MCM.

Added

2026-09-26

PMAL: Open Set Recognition via Robust Prototype Mining

PMAL: Open Set Recognition via Robust Prototype Mining

Jing Lu, Yunlu Xu, Hao Li, Zhanzhan Cheng, Yi Niu

OrganizationsHikvision Research InstituteZhejiang University

Why you should read this

Proposes a prototype mining and learning framework that improves open set recognition by selecting high-quality, diverse training samples as explicit class prototypes based on data uncertainty and feature topology.

Open Set Recognition (OSR) has been an emerging topic. Besides recognizing predefined classes, the system needs to reject the unknowns. Prototype learning is a potential manner to handle the problem, as its ability to improve intra-class compactness of representations is much needed in discrimination between the known and the unknowns. In this work, we propose a novel Prototype Mining And Learning (PMAL) framework. It has a prototype mining mechanism before the phase of optimizing embedding space, explicitly considering two crucial properties, namely high-quality and diversity of the prototype set. Concretely, a set of high-quality candidates are firstly extracted from training samples based on data uncertainty learning, avoiding the interference from unexpected noise. Considering the multifarious appearance of objects even in a single category, a diversity-based strategy for prototype set filtering is proposed. Accordingly, the embedding space can be better optimized to discriminate therein the predefined classes and between known and unknowns. Extensive experiments verify the two good characteristics (i.e., high-quality and diversity) embraced in prototype mining, and show the remarkable performance of the proposed framework compared to state-of-the-arts.

Added

2026-09-26

Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection

Catching Both Gray and Black Swans: Open-set Supervised Anomaly Detection

Choubo Ding, Guansong Pang, Chunhua Shen

OrganizationsSingapore Management UniversityUniversity of AdelaideZhejiang University

Why you should read this

Proposes a multi-head framework that disentangles known, pseudo, and latent residual abnormalities to effectively detect both seen and novel anomaly classes using only a small set of labeled anomaly examples.

Despite most existing anomaly detection studies assume the availability of normal training samples only, a few labeled anomaly examples are often available in many real-world applications, such as defect samples identified during random quality inspection, lesion images confirmed by radiologists in daily medical screening, etc. These anomaly examples provide valuable knowledge about the application-specific abnormality, enabling significantly improved detection of similar anomalies in some recent models. However, those anomalies seen during training often do not illustrate every possible class of anomaly, rendering these models ineffective in generalizing to unseen anomaly classes. This paper tackles open-set supervised anomaly detection, in which we learn detection models using the anomaly examples with the objective to detect both seen anomalies ('gray swans') and unseen anomalies ('black swans'). We propose a novel approach that learns disentangled representations of abnormalities illustrated by seen anomalies, pseudo anomalies, and latent residual anomalies (i.e., samples that have unusual residuals compared to the normal data in a latent space), with the last two abnormalities designed to detect unseen anomalies. Extensive experiments on nine real-world anomaly detection datasets show superior performance of our model in detecting seen and unseen anomalies under diverse settings. Code and data are available at: https://github.com/choubo/DRA

Added

2026-09-26

Toward Open Set Recognition

Toward Open Set Recognition

W. Scheirer, A. Rocha, Archana Sapkota, T. Boult

OrganizationsUniversity of CampinasUniversity of Colorado Colorado Springs

Why you should read this

Formalizes the open set recognition problem and introduces the 1-vs-Set Machine to limit classification risk in unconstrained spaces where unseen classes emerge at test time.

To date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This article explores the nature of open set recognition, and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step towards a solution, we introduce a novel "1-vs-Set Machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale experiments performed over data from the Caltech 256, ImageNet, and Labeled Faces in the Wild sets. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks.

Added

2026-09-24

Towards Open Set Deep Networks

Towards Open Set Deep Networks

Abhijit Bendale, Terrance Boult

OrganizationsUniversity of Colorado Colorado Springs

Why you should read this

Proposes OpenMax, an alternative output layer that enables deep neural networks to perform open set recognition by estimating the probability of unknown classes and rejecting unseen or fooling images.

Deep networks have produced significant gains for various visual recognition problems, leading to high impact academic and commercial applications. Recent work in deep networks highlighted that it is easy to generate images that humans would never classify as a particular object class, yet networks classify such images high confidence as that given class - deep network are easily fooled with images humans do not consider meaningful. The closed set nature of deep networks forces them to choose from one of the known classes leading to such artifacts. Recognition in the real world is open set, i.e. the recognition system should reject unknown/unseen classes at test time. We present a methodology to adapt deep networks for open set recognition, by introducing a new model layer, OpenMax, which estimates the probability of an input being from an unknown class. A key element of estimating the unknown probability is adapting Meta-Recognition concepts to the activation patterns in the penultimate layer of the network. OpenMax allows rejection of "fooling" and unrelated open set images presented to the system; OpenMax greatly reduces the number of obvious errors made by a deep network. We prove that the OpenMax concept provides bounded open space risk, thereby formally providing an open set recognition solution. We evaluate the resulting open set deep networks using pre-trained networks from the Caffe Model-zoo on ImageNet 2012 validation data, and thousands of fooling and open set images. The proposed OpenMax model significantly outperforms open set recognition accuracy of basic deep networks as well as deep networks with thresholding of SoftMax probabilities.

Added

2026-09-18

Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly

Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly

Yongqin Xian, Christoph H. Lampert, Bernt Schiele, Zeynep Akata

OrganizationsInstitute of Science and Technology AustriaMax Planck Institute for InformaticsUniversity of Amsterdam

Why you should read this

Establishes a unified evaluation benchmark and standardized data splits to resolve widespread test-set overlap in zero-shot learning, while introducing the Animals with Attributes 2 (AWA2) dataset and systematically comparing leading methods under both standard and generalized settings.

Due to the importance of zero-shot learning, i.e. classifying images where there is a lack of labeled training data, the number of proposed approaches has recently increased steadily. We argue that it is time to take a step back and to analyze the status quo of the area. The purpose of this paper is three-fold. First, given the fact that there is no agreed upon zero-shot learning benchmark, we first define a new benchmark by unifying both the evaluation protocols and data splits of publicly available datasets used for this task. This is an important contribution as published results are often not comparable and sometimes even flawed due to, e.g. pre-training on zero-shot test classes. Moreover, we propose a new zero-shot learning dataset, the Animals with Attributes 2 (AWA2) dataset which we make publicly available both in terms of image features and the images themselves. Second, we compare and analyze a significant number of the state-of-the-art methods in depth, both in the classic zero-shot setting but also in the more realistic generalized zero-shot setting. Finally, we discuss in detail the limitations of the current status of the area which can be taken as a basis for advancing it.

Added

2026-09-18