Unsupervised Out-of-Distribution Detection with Diffusion Inpainting
Zhenzhen LiuJin Peng ZhouYufan WangKilian Q. Weinberger
Introduces Lift, Map, Detect (LMD), an unsupervised out-of-distribution detection framework that identifies anomalous images by masking them and measuring reconstruction error after inpainting with a diffusion model trained only on in-domain data.
Machine learning systems deployed in high-stakes environments, such as medical diagnostics and criminal justice, rely on the assumption that incoming test data match the distribution of their training data. When exposed to unfamiliar or out-of-distribution inputs, these models can produce silent, misleading, or hazardous failures. Many existing safety techniques require labeled data or prior knowledge of potential anomalies, which is often impractical because anomalies in the real world are unpredictable and human labeling is expensive. Developing effective out-of-distribution detection that relies strictly on unlabeled in-domain data is therefore essential for reliable artificial intelligence deployment.
The article introduces and evaluates Lift, Map, Detect (LMD), a new unsupervised framework that identifies out-of-distribution images using generative diffusion models without requiring model retraining or labeled data.
The approach builds on the principle that diffusion models learn to map corrupted images back toward the underlying data manifold on which they were trained. Under LMD, an image is first lifted off its original distribution by applying a corrupting mask. The model then fills in the missing regions through an inpainting process. If the input belongs to the target domain, the model reconstructs it accurately; if it comes from an unfamiliar domain, the model attempts to force it into the target domain, creating significant visual errors. The distance between the original and reconstructed image is measured using a standard perceptual similarity metric. To reduce random variation, the framework performs multiple reconstruction attempts using an alternating checkerboard mask pattern and aggregates the results via a median score.
The evaluation demonstrates that LMD achieves top-tier performance across diverse benchmark datasets, attaining the highest overall average area under the ROC curve (0.907) compared to seven competitive baseline methods. LMD achieved standout results on difficult image pairs, improving detection accuracy by up to 10% over the best baseline on CIFAR-100 versus SVHN and reaching a 0.991 detection score on high-resolution image benchmarks. Ablation studies confirm that using an alternating checkerboard pattern and aggregating roughly ten reconstruction attempts per image consistently maximizes detection accuracy, whereas traditional single-region masks degrade performance.
These findings indicate that generative diffusion models can serve as robust safety guards for vision-based artificial intelligence systems without requiring expensive retraining or domain labels. By leveraging perceptual reconstruction errors rather than raw statistical likelihoods, LMD avoids common failure modes where generative models erroneously assign high confidence to abnormal inputs. This reduces operational and safety risks in production environments where abnormal data could otherwise bypass standard error checks.
Organizations seeking to implement LMD should adopt the alternating checkerboard masking strategy combined with perceptual similarity scoring. However, because diffusion-based inpainting requires multiple iterative sampling steps, LMD is computationally intensive and currently too slow for real-time applications requiring millisecond responses. Future engineering efforts should integrate emerging fast-sampling diffusion algorithms to reduce inference latency before deploying the system in time-critical operational workflows.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Introduces the foundational denoising diffusion probabilistic model formulation that the source adapts for iterative manifold projection and inpainting.
- Paper: Generalized Out-of-Distribution Detection: A Survey, Jingkang Yang et al. (2021). Provides a comprehensive taxonomy and standardized benchmark formulations for generalized out-of-distribution detection problems.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Demonstrates how pre-trained diffusion models perform unsupervised linear inverse image restoration tasks like inpainting, directly inspiring manifold-mapping techniques.
- Paper: Diffusion Models: A Comprehensive Survey of Methods and Applications, Ling Yang et al. (2022). Surveys the core mathematical foundations and conditioning mechanisms of diffusion models, establishing essential background for diffusion-based representation tasks.
- Paper: Energy-based Out-of-distribution Detection, Weitang Liu et al. (2020). Presents foundational score-based formulations for out-of-distribution detection that serve as key comparative baselines for unsupervised OOD scoring.
- Paper: Out-of-Distribution Detection with Deep Nearest Neighbors, Yiyou Sun et al. (2022). Establishes non-parametric distance-based representations for OOD detection, providing key context on manifold distance evaluation.
- Paper: Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks, Shiyu Liang et al. (2018). Introduces input-perturbation and scoring principles for post-hoc out-of-distribution detection that contextualize corruption-and-mapping paradigms.
- Paper: A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks, Kimin Lee et al. (2018). Formulates distance-based Mahalanobis scoring for out-of-distribution detection, establishing a fundamental baseline for measuring feature divergence.
- Paper: Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need, Jingyao Li et al. (2023). Explores masked image modeling as an alternative unsupervised reconstruction pretext task for out-of-distribution detection.
- Paper: Generalization in diffusion models arises from geometry-adaptive harmonic representations, Zahra Kadkhodaie et al. (2024). Investigates the theoretical generalization dynamics and manifold-learning capabilities of diffusion denoisers across varied data scales.
- Paper: Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement, Kai Xu et al. (2024). Extends post-hoc out-of-distribution detection techniques by analyzing feature activation scaling without requiring generative reconstruction.
