AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Qihang ZhouGuansong PangYu TianShibo HeJiming Chen
Introduces AnomalyCLIP, a prompt-learning framework that adapts vision-language models for zero-shot anomaly detection and segmentation across diverse industrial and medical domains by decoupling generic abnormality patterns from object-specific semantics.
Detecting anomalies—such as industrial manufacturing defects or medical lesions—is vital for quality control and clinical diagnosis. Traditional automated systems depend heavily on large amounts of domain-specific training data. However, acquiring such data is often impractical due to patient privacy constraints or the deployment of new manufacturing lines. Zero-shot anomaly detection addresses this challenge by identifying abnormalities in target domains without requiring any prior target-domain training samples. While large vision-language models like CLIP possess strong general visual knowledge, they typically focus on identifying what an object is rather than whether it contains an abnormality, resulting in weak zero-shot anomaly detection performance.
The main objective of the article is to develop and evaluate AnomalyCLIP, an approach designed to adapt vision-language models for accurate zero-shot anomaly classification and localized segmentation across diverse domains. It demonstrates that replacing object-specific semantics with learned, object-agnostic text prompts enables the model to identify generic abnormality patterns across widely differing visual contexts.
The researchers designed AnomalyCLIP by establishing two learnable, object-agnostic prompt templates representing normality and abnormality. They optimized these prompts using auxiliary data through a combined global and local objective function, integrating image-level classification losses with fine-grained pixel segmentation losses (using focal and Dice loss formulations). To preserve fine local visual details that standard attention mechanisms often suppress, the team introduced Diagonally Prominent Attention Maps (specifically using value-to-value self-attention) within the frozen visual encoder and refined intermediate layers in the text encoder. The method was rigorously evaluated across 17 real-world datasets encompassing both industrial defect inspection and various medical imaging modalities, such as endoscopy, radiology, and photography.
The evaluation yielded several critical findings. First, AnomalyCLIP substantially outperformed existing zero-shot methods on industrial inspection benchmarks, achieving high image-level area under the curve scores (such as 97.5% on DAGM and 91.5% on MVTec AD) while drastically improving pixel-level localization (reaching 95.5% on VisA compared to WinCLIP's 79.6%). Second, prompts tuned solely on industrial defects generalized successfully to medical domains, accurately detecting skin lesions, brain tumors, and colon polyps without any medical training data. Third, when fine-tuned on an auxiliary medical dataset, AnomalyCLIP achieved even higher diagnostic performance, reaching an image-level area under the curve of 97.9% on brain tumor datasets and 93.2% on colon polyp segmentation. Fourth, AnomalyCLIP achieved detection accuracy comparable to—and in some cases exceeding—state-of-the-art fully trained models that require dedicated normal training images.
These findings indicate that anomaly detection models do not require extensive target-specific training to achieve production-grade defect and lesion localization. By relying on universal visual patterns of damage rather than specific object classes, organizations can significantly reduce data collection, labeling, and integration costs. Furthermore, AnomalyCLIP performs full classification and dense segmentation in a single forward pass without requiring complex decoders or hundreds of engineered prompts, greatly shortening development timelines and reducing inference overhead.
Organizations should adopt object-agnostic prompt tuning strategies for cold-start inspection systems where target data collection is cost-prohibitive or restricted. When deploying across specialized domains such as healthcare, practitioners should fine-tune auxiliary prompts on visually aligned proxy datasets (such as colonoscopy data for tumor detection) to maximize localization accuracy. Before deploying into safety-critical environments, teams should conduct pilot studies, as performance remains sensitive to extreme visual domain shifts—such as transitioning from focal polyps to diffuse chest infections—and hyperparameter selection regarding prompt length and encoder depth.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the foundation CLIP vision-language model and contrastive pre-training framework that AnomalyCLIP directly adapts and builds upon for zero-shot anomaly detection.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Pioneers Context Optimization (CoOp) for learning continuous prompt representations in frozen vision-language models, providing the baseline prompt-tuning paradigm modified by AnomalyCLIP.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). Extends prompt learning to dynamic, conditional contexts to improve zero-shot generalizability, motivating AnomalyCLIP's design of transferable, object-agnostic prompt templates.
- Paper: MVTec AD — A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection, Paul Bergmann et al. (2019). Establishes the standard MVTec AD benchmark and pixel-level evaluation protocols for visual anomaly detection used extensively in AnomalyCLIP's experiments.
- Paper: Towards Total Recall in Industrial Anomaly Detection, Karsten Roth et al. (2021). Presents PatchCore, the defining state-of-the-art memory-bank baseline for industrial anomaly localization that AnomalyCLIP compares against in zero-shot regimes.
- Paper: DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations, Ximeng Sun et al. (2022). Demonstrates dual positive and negative prompt optimization in CLIP for multi-label recognition, establishing the conceptual dual-prompt design used by AnomalyCLIP to distinguish normality from abnormality.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). Develops patch-aligned contrastive learning to extract fine-grained local visual representations from CLIP without dense mask supervision, addressing localization challenges central to AnomalyCLIP.
- Paper: Delving into Out-of-Distribution Detection with Vision-Language Representations, Yifei Ming et al. (2022). Explores zero-shot out-of-distribution detection with vision-language models via textual concept matching, framing anomaly recognition as a prompt-driven classification problem.
No sufficiently relevant recommendations were found.
