keyword
anomaly segmentation
Anomaly segmentation is a computer vision task that identifies and delineates abnormal, defective, or out-of-distribution regions within an image at the pixel level. Unlike image-level anomaly detection, which only determines whether an entire image contains an irregularity, anomaly segmentation produces dense pixel-wise masks or continuous anomaly score maps that reveal the exact spatial boundaries and extent of atypical patterns. Because anomalous instances are typically rare, unpredictable, and difficult to collect comprehensively during training, models for this task are commonly trained on normal, defect-free data using unsupervised, self-supervised, or zero-shot methods to learn regular feature distributions and pinpoint deviations. This task is widely applied in critical visual inspection domains that require precise spatial localization of flaws, such as industrial quality control, manufacturing defect analysis, and medical diagnostics.
6 items

A Diffusion-Based Framework for Multi-Class Anomaly Detection
Haoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen, Zhishan Li, Xu Chen, Yabiao Wang, Chengjie Wang, Lei Xie
Why you should read this
Proposes a semantic-guided diffusion framework that solves class confusion and structural distortion during reconstruction, achieving state-of-the-art multi-class anomaly detection and localization on the MVTec-AD and VisA benchmarks.
Reconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods might face challenges related to the preservation of image categories and pixel-wise structural integrity in the more practical multi-class setting. To solve the above problems, we propose a Diffusion-based Anomaly Detection (DiAD) framework for multi-class anomaly detection, which consists of a pixel-space autoencoder, a latent-space Semantic-Guided (SG) network with a connection to the stable diffusion's denoising network, and a feature-space pre-trained feature extractor. Firstly, the SG network is proposed for reconstructing anomalous regions while preserving the original image's semantic information. Secondly, we introduce Spatial-aware Feature Fusion (SFF) block to maximize reconstruction accuracy when dealing with extensively reconstructed areas. Thirdly, the input and reconstructed images are processed by a pre-trained feature extractor to generate anomaly maps based on features extracted at different scales. Experiments on MVTec-AD and VisA datasets demonstrate the effectiveness of our approach which surpasses the state-of-the-art methods, e.g., achieving 96.8/52.6 and 97.2/99.0 (AUROC/AP) for localization and detection respectively on multi-class MVTec-AD dataset. Code is available at https://lewandofskee.github.io/projects/diad.
Added
2026-10-05

AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Qihang Zhou, Guansong Pang, Yu Tian, Shibo He, Jiming Chen
Why you should read this
Introduces AnomalyCLIP, a prompt-learning framework that adapts vision-language models for zero-shot anomaly detection and segmentation across diverse industrial and medical domains by decoupling generic abnormality patterns from object-specific semantics.
Zero-shot anomaly detection (ZSAD) requires detection models trained using auxiliary data to detect anomalies without any training sample in a target dataset. It is a crucial task when training data is not accessible due to various concerns, eg, data privacy, yet it is challenging since the models need to generalize to anomalies across different domains where the appearance of foreground objects, abnormal regions, and background features, such as defects/tumors on different products/organs, can vary significantly. Recently large pre-trained vision-language models (VLMs), such as CLIP, have demonstrated strong zero-shot recognition ability in various vision tasks, including anomaly detection. However, their ZSAD performance is weak since the VLMs focus more on modeling the class semantics of the foreground objects rather than the abnormality/normality in the images. In this paper we introduce a novel approach, namely AnomalyCLIP, to adapt CLIP for accurate ZSAD across different domains. The key insight of AnomalyCLIP is to learn object-agnostic text prompts that capture generic normality and abnormality in an image regardless of its foreground objects. This allows our model to focus on the abnormal image regions rather than the object semantics, enabling generalized normality and abnormality recognition on diverse types of objects. Large-scale experiments on 17 real-world anomaly detection datasets show that AnomalyCLIP achieves superior zero-shot performance of detecting and segmenting anomalies in datasets of highly diverse class semantics from various defect inspection and medical imaging domains. Code will be made available at this https URL.
Added
2026-09-26

Prototypical Residual Networks for Anomaly Detection and Localization
Hui Zhang, Zuxuan Wu, Zheng Wang, Zhineng Chen, Yu-Gang Jiang
Why you should read this
Proposes a prototypical residual network that captures multi-scale feature deviations and variable-sized defects to achieve state-of-the-art anomaly detection and precise pixel-level localization across industrial benchmarks.
Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the other hand, anomalies are typically subtle, hard to discern, and of various appearance, making it difficult to detect anomalies and let alone locate anomalous regions. To address these issues, we propose a framework called Prototypical Residual Network (PRN), which learns feature residuals of varying scales and sizes between anomalous and normal patterns to accurately reconstruct the segmentation maps of anomalous regions. PRN mainly consists of two parts: multi-scale prototypes that explicitly represent the residual features of anomalies to normal patterns; a multi-size self-attention mechanism that enables variable-sized anomalous feature learning. Besides, we present a variety of anomaly generation strategies that consider both seen and unseen appearance variance to enlarge and diversify anomalies. Extensive experiments on the challenging and widely used MVTec AD benchmark show that PRN outperforms current state-of-the-art unsupervised and supervised methods. We further report SOTA results on three additional datasets to demonstrate the effectiveness and generalizability of PRN.
Added
2026-09-26

GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models
Chen Liang, Wenguan Wang, Jiaxu Miao, Yi Yang
Why you should read this
Proposes a hybrid semantic segmentation framework that models pixel feature densities via Gaussian Mixture Models using online Expectation-Maximization alongside end-to-end discriminative representation learning, delivering superior closed-set accuracy while naturally detecting out-of-distribution anomalies without architectural changes or post-processing.
Prevalent semantic segmentation solutions are, in essence, a dense discriminative classifier of p(class | pixel feature). Though straightforward, this de facto paradigm neglects the underlying data distribution p(pixel feature | class), and struggles to identify out-of-distribution data. Going beyond this, we propose GMMSeg, a new family of segmentation models that rely on a dense generative classifier for the joint distribution p(pixel feature, class). For each class, GMMSeg builds Gaussian Mixture Models (GMMs) via Expectation-Maximization (EM), so as to capture class-conditional densities. Meanwhile, the deep dense representation is end-to-end trained in a discriminative manner, i.e., maximizing p(class | pixel feature). This endows GMMSeg with the strengths of both generative and discriminative models. With a variety of segmentation architectures and backbones, GMMSeg outperforms the discriminative counterparts on three closed-set datasets. More impressively, without any modification, GMMSeg even performs well on open-world datasets. We believe this work brings fundamental insights into the related fields.
Added
2026-09-26

PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization
Thomas Defard, Aleksandr Setkov, Angélique Loesch, Romaric Audigier
Why you should read this
Presents a patch distribution modeling framework that combines multi-scale pretrained CNN embeddings with multivariate Gaussian distributions to achieve accurate, computationally efficient image anomaly detection and localization for industrial inspection.
We present a new framework for Patch Distribution Modeling, PaDiM, to concurrently detect and localize anomalies in images in a one-class learning setting. PaDiM makes use of a pretrained convolutional neural network (CNN) for patch embedding, and of multivariate Gaussian distributions to get a probabilistic representation of the normal class. It also exploits correlations between the different semantic levels of CNN to better localize anomalies. PaDiM outperforms current state-of-the-art approaches for both anomaly detection and localization on the MVTec AD and STC datasets. To match real-world visual industrial inspection, we extend the evaluation protocol to assess performance of anomaly localization algorithms on non-aligned dataset. The state-of-the-art performance and low complexity of PaDiM make it a good candidate for many industrial applications.
Added
2026-09-25

MVTec AD — A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection
Paul Bergmann, Michael Fauser, David Sattlegger, C. Steger
Why you should read this
Introduces the first comprehensive real-world industrial inspection dataset with pixel-accurate defect annotations across fifteen object and texture categories, establishing a rigorous benchmark that reveals key limitations in current unsupervised anomaly detection and localization methods.
The detection of anomalous structures in natural image data is of utmost importance for numerous tasks in the field of computer vision. The development of methods for unsupervised anomaly detection requires data on which to train and evaluate new approaches and ideas. We introduce the MVTec Anomaly Detection (MVTec AD) dataset containing 5354 high-resolution color images of different object and texture categories. It contains normal, i.e., defect-free, images intended for training and images with anomalies intended for testing. The anomalies manifest themselves in the form of over 70 different types of defects such as scratches, dents, contaminations, and various structural changes. In addition, we provide pixel-precise ground truth regions for all anomalies. We also conduct a thorough evaluation of current state-of-the-art unsupervised anomaly detection methods based on deep architectures such as convolutional autoencoders, generative adversarial networks, and feature descriptors using pre-trained convolutional neural networks, as well as classical computer vision methods. This initial benchmark indicates that there is considerable room for improvement. To the best of our knowledge, this is the first comprehensive, multi-object, multi-defect dataset for anomaly detection that provides pixel-accurate ground truth regions and focuses on real-world applications.
Added
2026-09-15
