AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

Qihang ZhouGuansong PangYu TianShibo HeJiming Chen

article2024ICLR503 citations

Introduces AnomalyCLIP, a prompt-learning framework that adapts vision-language models for zero-shot anomaly detection and segmentation across diverse industrial and medical domains by decoupling generic abnormality patterns from object-specific semantics.

Listen

Detecting anomalies—such as industrial manufacturing defects or medical lesions—is vital for quality control and clinical diagnosis. Traditional automated systems depend heavily on large amounts of domain-specific training data. However, acquiring such data is often impractical due to patient privacy constraints or the deployment of new manufacturing lines. Zero-shot anomaly detection addresses this challenge by identifying abnormalities in target domains without requiring any prior target-domain training samples. While large vision-language models like CLIP possess strong general visual knowledge, they typically focus on identifying what an object is rather than whether it contains an abnormality, resulting in weak zero-shot anomaly detection performance.

The main objective of the article is to develop and evaluate AnomalyCLIP, an approach designed to adapt vision-language models for accurate zero-shot anomaly classification and localized segmentation across diverse domains. It demonstrates that replacing object-specific semantics with learned, object-agnostic text prompts enables the model to identify generic abnormality patterns across widely differing visual contexts.

The researchers designed AnomalyCLIP by establishing two learnable, object-agnostic prompt templates representing normality and abnormality. They optimized these prompts using auxiliary data through a combined global and local objective function, integrating image-level classification losses with fine-grained pixel segmentation losses (using focal and Dice loss formulations). To preserve fine local visual details that standard attention mechanisms often suppress, the team introduced Diagonally Prominent Attention Maps (specifically using value-to-value self-attention) within the frozen visual encoder and refined intermediate layers in the text encoder. The method was rigorously evaluated across 17 real-world datasets encompassing both industrial defect inspection and various medical imaging modalities, such as endoscopy, radiology, and photography.

The evaluation yielded several critical findings. First, AnomalyCLIP substantially outperformed existing zero-shot methods on industrial inspection benchmarks, achieving high image-level area under the curve scores (such as 97.5% on DAGM and 91.5% on MVTec AD) while drastically improving pixel-level localization (reaching 95.5% on VisA compared to WinCLIP's 79.6%). Second, prompts tuned solely on industrial defects generalized successfully to medical domains, accurately detecting skin lesions, brain tumors, and colon polyps without any medical training data. Third, when fine-tuned on an auxiliary medical dataset, AnomalyCLIP achieved even higher diagnostic performance, reaching an image-level area under the curve of 97.9% on brain tumor datasets and 93.2% on colon polyp segmentation. Fourth, AnomalyCLIP achieved detection accuracy comparable to—and in some cases exceeding—state-of-the-art fully trained models that require dedicated normal training images.

These findings indicate that anomaly detection models do not require extensive target-specific training to achieve production-grade defect and lesion localization. By relying on universal visual patterns of damage rather than specific object classes, organizations can significantly reduce data collection, labeling, and integration costs. Furthermore, AnomalyCLIP performs full classification and dense segmentation in a single forward pass without requiring complex decoders or hundreds of engineered prompts, greatly shortening development timelines and reducing inference overhead.

Organizations should adopt object-agnostic prompt tuning strategies for cold-start inspection systems where target data collection is cost-prohibitive or restricted. When deploying across specialized domains such as healthcare, practitioners should fine-tune auxiliary prompts on visually aligned proxy datasets (such as colonoscopy data for tumor detection) to maximize localization accuracy. Before deploying into safety-critical environments, teams should conduct pilot studies, as performance remains sensitive to extreme visual domain shifts—such as transitioning from focal polyps to diffuse chest infections—and hyperparameter selection regarding prompt length and encoder depth.

No sufficiently relevant recommendations were found.

Cover for AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

Abstract

Zero-shot anomaly detection (ZSAD) requires detection models trained using auxiliary data to detect anomalies without any training sample in a target dataset. It is a crucial task when training data is not accessible due to various concerns, eg, data privacy, yet it is challenging since the models need to generalize to anomalies across different domains where the appearance of foreground objects, abnormal regions, and background features, such as defects/tumors on different products/organs, can vary significantly. Recently large pre-trained vision-language models (VLMs), such as CLIP, have demonstrated strong zero-shot recognition ability in various vision tasks, including anomaly detection. However, their ZSAD performance is weak since the VLMs focus more on modeling the class semantics of the foreground objects rather than the abnormality/normality in the images. In this paper we introduce a novel approach, namely AnomalyCLIP, to adapt CLIP for accurate ZSAD across different domains. The key insight of AnomalyCLIP is to learn object-agnostic text prompts that capture generic normality and abnormality in an image regardless of its foreground objects. This allows our model to focus on the abnormal image regions rather than the object semantics, enabling generalized normality and abnormality recognition on diverse types of objects. Large-scale experiments on 17 real-world anomaly detection datasets show that AnomalyCLIP achieves superior zero-shot performance of detecting and segmenting anomalies in datasets of highly diverse class semantics from various defect inspection and medical imaging domains. Code will be made available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 AnomalyCLIP: object-agnostic prompt learning
  • 3.1 Approach overview
  • 3.2 Object-agnostic text prompt design
  • 3.3 Learning generic abnormality and normality prompts
  • 4 Experiments
  • 4.1 Experiment setup
  • 4.2 Main results
  • 4.3 Ablation study
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Implementation details and baselines
  • A.1 Implementation details
  • A.2 Baselines
  • B Dataset
  • C Detailed Analysis of DPAM
  • D Additional results and ablations
  • E Visualization
  • F Fine-grained ZSAD performance

Citation

MLA
Zhou, Q., et al. “AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection”. arXiv, 2023, http://arxiv.org/abs/2310.18961v12.
APA
Zhou, Q., Pang, G., Tian, Y., He, S., & Chen, J. (2023). AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection. arXiv. http://arxiv.org/abs/2310.18961v12
Chicago
Zhou, Q., G. Pang, Y. Tian, S. He, and J. Chen. 2023. “AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection”. arXiv. http://arxiv.org/abs/2310.18961v12.
Harvard
Zhou, Q. et al. (2023) “AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.18961v12.
Vancouver
1. Zhou Q, Pang G, Tian Y, He S, Chen J (2023) AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection. arXiv

BibTeX

@article{zhou2023anomalyclip,
  title = {AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection},
  author = {Zhou, Qihang and Pang, Guansong and Tian, Yu and He, Shibo and Chen, Jiming},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.18961v12},
  eprint = {2310.18961}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors