Prototypical Residual Networks for Anomaly Detection and Localization
Hui ZhangZuxuan WuZheng WangZhineng ChenYu-Gang Jiang
Proposes a prototypical residual network that captures multi-scale feature deviations and variable-sized defects to achieve state-of-the-art anomaly detection and precise pixel-level localization across industrial benchmarks.
Automated visual inspection is critical for quality control in modern manufacturing, medical imaging, and surveillance. However, industrial computer vision systems face significant challenges because defective samples are exceptionally rare, subtle, and highly variable in shape and size. Existing unsupervised methods, which learn only from normal samples, frequently yield coarse predictions and false alarms, while existing supervised approaches struggle with extreme data imbalance and fail to pinpoint exact defect boundaries.
To resolve these problems, the article introduces the Prototypical Residual Network (PRN), a supervised deep learning framework designed to simultaneously achieve high-precision anomaly detection (identifying whether an image contains a defect) and anomaly localization (pinpointing exact defective pixels). PRN learns explicit feature differences—or residuals—between anomalous images and representative normal visual patterns across multiple spatial resolutions. The architecture incorporates multi-scale prototypes to model normal textures, a multi-size self-attention mechanism to detect inconsistencies across varying patch sizes, and multi-scale fusion blocks to integrate contextual information. To overcome severe data scarcity, the authors also implement synthetic data generation strategies that augment seen defects and simulate new, unseen anomalies.
The authors evaluated PRN across four benchmark industrial inspection datasets (MVTec AD, DAGM, BTAD, and KolektorSDD2) under a few-shot setting using only 10 abnormal training samples per category. The experimental results demonstrate strong performance advantages. On the primary MVTec AD benchmark, PRN achieved an overall image-level detection accuracy (AUROC) of 99.4% and a pixel-level localization AUROC of 99.0%, surpassing competing supervised and unsupervised models. On the more rigorous Average Precision metric for defect localization, PRN reached 78.6%, outperforming the prior unsupervised benchmark by 10.5 percentage points and the leading supervised method by 52.6 percentage points. Furthermore, PRN processed images in approximately 0.064 seconds per frame, operating 30% to 70% faster than comparable high-performing models.
These findings indicate that learning residual representations against clustered normal prototypes, combined with targeted synthetic anomaly generation, effectively solves the few-shot imbalance bottleneck in visual defect detection. In practical terms, this allows automated inspection systems to achieve superior defect detection reliability and pinpoint accuracy without requiring massive defect-labeling efforts. By significantly reducing false alarms and processing images at higher speeds, the approach can lower production scrap costs, improve throughput, and minimize deployment overhead in automated quality assurance pipelines.
Organizations evaluating automated optical inspection systems should consider adopting residual-based prototype architectures to improve defect localization. Decision-makers planning implementations should first establish pilot deployments on representative assembly lines to validate operational gains against baseline systems. As a key technical consideration, PRN relies on access to accurate ground truth segmentation masks during initial training, and its image-level scoring rule may under-weigh extremely minute flaws. Future engineering efforts should therefore explore automated mask generation and refined scoring mechanisms tailored specifically for micro-defects.
- Paper: Towards Total Recall in Industrial Anomaly Detection, Karsten Roth et al. (2021). Introduces PatchCore, the foundational industrial anomaly detection and localization baseline that PRN directly compares against and aims to improve upon using prototypical residual learning.
- Paper: PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization, Thomas Defard et al. (2020). Establishes multi-scale patch feature extraction and localized distribution modeling for anomaly localization on the MVTec AD benchmark, motivating PRN's multi-scale normal texture representations.
- Paper: Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection, Dong Gong et al. (2019). Pioneers the concept of constraining reconstruction to normal prototypes via memory banks, which directly underpins PRN's use of normal prototype clusters to compute residual representations.
- Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Provides the foundational multi-scale feature pyramid and contextual fusion architecture that PRN adapts to localize visual defects across varying spatial resolutions.
- Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). Introduces focal loss to handle extreme class imbalance in dense visual tasks, an essential technique for training supervised anomaly localization models under severe defect scarcity.
- Paper: Deep Learning for Anomaly Detection, Guansong Pang et al. (2020). Provides a comprehensive taxonomy and formulation of deep anomaly detection methods, including few-shot weakly supervised paradigms and normality representation learning.
- Paper: Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement, Kai Xu et al. (2024). Analyzes feature activation scaling and post-hoc out-of-distribution detection enhancements, providing theoretical insights that extend the feature representation principles evaluated in PRN.
- Paper: Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need, Jingyao Li et al. (2023). Explores masked image modeling as a pretext task for out-of-distribution and anomaly detection, offering an alternative self-supervised paradigm to PRN's prototype-residual framework.
- Paper: Revisiting Prototypical Network for Cross Domain Few-Shot Learning, Fei Zhou et al. (2023). Extends prototypical networks to cross-domain few-shot settings using local-global distillation, building on the multi-scale prototype concepts utilized in PRN.
