Latent Outlier Exposure for Anomaly Detection with Contaminated Data
Chen QiuAodong LiMarius KloftMaja RudolphStephan Mandt
Proposes Latent Outlier Exposure, a general training framework that jointly optimizes model parameters and infers latent anomaly labels to effectively train anomaly detectors on contaminated, uncurated datasets across image, tabular, and video benchmarks.
Real-world automated anomaly detection systems—such as industrial fault monitors, medical diagnostic tools, and cybersecurity fraud engines—typically rely on the assumption that training datasets contain purely normal data. In practice, however, large-scale and uncurated datasets are routinely contaminated with unidentified anomalies. When traditional deep anomaly detection models are trained directly on such corrupted data, their ability to identify outliers deteriorates significantly.
The article demonstrates a flexible, domain-independent training framework called Latent Outlier Exposure to train effective anomaly detectors directly on contaminated data without requiring human-labeled outliers.
The proposed framework evaluates data points through two coupled objectives sharing model parameters: one designed to pull normal data together, and an opposing objective designed to push abnormal data away. Rather than ignoring anomalies or simply filtering them out, the algorithm alternates between estimating unobserved labels (normal versus anomalous) and updating the shared model parameters using a block coordinate optimization process. The authors evaluated this strategy across synthetic data, three standard image datasets, 30 diverse tabular datasets, and a video anomaly detection benchmark, examining both "hard" deterministic label assignments and "soft" probabilistic assignments that account for uncertainty.
The evaluation yielded several key findings. First, the proposed strategy consistently outperformed standard baseline approaches that either ignore contamination or iteratively filter it out. On image datasets with 10% contamination, the approach recovered nearly the full detection accuracy of models trained on clean data, narrowing the performance gap from over 4–9 percentage points down to roughly 1–2 percentage points. Second, the soft-labeling variant achieved state-of-the-art results on video anomaly detection, outperforming deep ordinal regression baselines by 18.8% in area under the receiver operating characteristic curve at a 10% contamination rate. Third, across 30 tabular datasets, the method regularly delivered the highest detection metrics and in several instances achieved higher accuracy than models trained on clean data, proving that unlabeled anomalies can actively strengthen decision boundaries when modeled properly. Finally, sensitivity analyses revealed that the approach remains stable even when the assumed contamination rate is incorrectly estimated by up to 15%.
These findings indicate that organizations do not need to invest extensive time and capital into costly, manual data cleaning before deploying anomaly detection systems. By actively exploiting the learning signals within contaminated data, systems become more reliable in critical settings like healthcare and fraud prevention where missed anomalies carry substantial operational and safety risks.
Organizations training anomaly detection pipelines on uncurated operational data should adopt Latent Outlier Exposure mechanisms within their existing network architectures. Deployers should select the hard-labeling variant for standard structured tasks and prioritize the soft-labeling variant when facing high noise, uncertainty, or temporal streams such as video data. Before production deployment, teams should conduct sensitivity analyses on small sample batches to determine whether overestimating or underestimating the assumed contamination rate better aligns with operational risk tolerance.
The results should be interpreted with awareness that performance depends on setting a reasonable estimate for the contamination ratio. In addition, the hard-labeling variant can risk overfitting if normal data points are mistakenly categorized as outliers. Nonetheless, given the extensive validation across multiple domains and architectures, decision-makers can have high confidence in adopting this training strategy for contaminated environments.
- Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). This work introduces the foundational Outlier Exposure paradigm, which the source directly adapts and extends to contaminated, unlabeled settings via latent label optimization.
- Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). This paper establishes Deep Support Vector Data Description (Deep SVDD) for one-class classification, a standard deep anomaly detection baseline whose assumption of clean data is directly challenged and generalized by the source.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). This study introduces semi-supervised loss modeling techniques to separate clean and corrupted samples during training, providing core conceptual motivation for handling unlabeled noisy data.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). This survey provides a comprehensive foundation for learning with label noise and sample selection mechanisms in deep neural networks.
- Paper: Deep Learning for Anomaly Detection, Guansong Pang et al. (2020). This review provides a structured taxonomy of deep anomaly detection formulations, contextualizing the clean-training assumptions that the source aims to relax.
- Paper: Rethinking Out-of-distribution (OOD) Detection: Masked Image Modeling is All You Need, Jingyao Li et al. (2023). This work explores outlier exposure in few-shot and reconstruction-based contexts, advancing out-of-distribution representation learning beyond standard supervised exposure.
- Paper: Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement, Kai Xu et al. (2024). This paper builds on modern out-of-distribution detection frameworks by analyzing how feature activation scaling during training and post-hoc evaluation sharpens in- versus out-of-distribution boundaries.
