Latent Outlier Exposure for Anomaly Detection with Contaminated Data

Chen QiuAodong LiMarius KloftMaja RudolphStephan Mandt

article2022ICML84 citations

Proposes Latent Outlier Exposure, a general training framework that jointly optimizes model parameters and infers latent anomaly labels to effectively train anomaly detectors on contaminated, uncurated datasets across image, tabular, and video benchmarks.

Listen

Real-world automated anomaly detection systems—such as industrial fault monitors, medical diagnostic tools, and cybersecurity fraud engines—typically rely on the assumption that training datasets contain purely normal data. In practice, however, large-scale and uncurated datasets are routinely contaminated with unidentified anomalies. When traditional deep anomaly detection models are trained directly on such corrupted data, their ability to identify outliers deteriorates significantly.

The article demonstrates a flexible, domain-independent training framework called Latent Outlier Exposure to train effective anomaly detectors directly on contaminated data without requiring human-labeled outliers.

The proposed framework evaluates data points through two coupled objectives sharing model parameters: one designed to pull normal data together, and an opposing objective designed to push abnormal data away. Rather than ignoring anomalies or simply filtering them out, the algorithm alternates between estimating unobserved labels (normal versus anomalous) and updating the shared model parameters using a block coordinate optimization process. The authors evaluated this strategy across synthetic data, three standard image datasets, 30 diverse tabular datasets, and a video anomaly detection benchmark, examining both "hard" deterministic label assignments and "soft" probabilistic assignments that account for uncertainty.

The evaluation yielded several key findings. First, the proposed strategy consistently outperformed standard baseline approaches that either ignore contamination or iteratively filter it out. On image datasets with 10% contamination, the approach recovered nearly the full detection accuracy of models trained on clean data, narrowing the performance gap from over 4–9 percentage points down to roughly 1–2 percentage points. Second, the soft-labeling variant achieved state-of-the-art results on video anomaly detection, outperforming deep ordinal regression baselines by 18.8% in area under the receiver operating characteristic curve at a 10% contamination rate. Third, across 30 tabular datasets, the method regularly delivered the highest detection metrics and in several instances achieved higher accuracy than models trained on clean data, proving that unlabeled anomalies can actively strengthen decision boundaries when modeled properly. Finally, sensitivity analyses revealed that the approach remains stable even when the assumed contamination rate is incorrectly estimated by up to 15%.

These findings indicate that organizations do not need to invest extensive time and capital into costly, manual data cleaning before deploying anomaly detection systems. By actively exploiting the learning signals within contaminated data, systems become more reliable in critical settings like healthcare and fraud prevention where missed anomalies carry substantial operational and safety risks.

Organizations training anomaly detection pipelines on uncurated operational data should adopt Latent Outlier Exposure mechanisms within their existing network architectures. Deployers should select the hard-labeling variant for standard structured tasks and prioritize the soft-labeling variant when facing high noise, uncertainty, or temporal streams such as video data. Before production deployment, teams should conduct sensitivity analyses on small sample batches to determine whether overestimating or underestimating the assumed contamination rate better aligns with operational risk tolerance.

The results should be interpreted with awareness that performance depends on setting a reasonable estimate for the contamination ratio. In addition, the hard-labeling variant can risk overfitting if normal data points are mistakenly categorized as outliers. Nonetheless, given the extensive validation across multiple domains and architectures, decision-makers can have high confidence in adopting this training strategy for contaminated environments.

  • Paper: Deep Anomaly Detection with Outlier Exposure, Dan Hendrycks et al. (2019). This work introduces the foundational Outlier Exposure paradigm, which the source directly adapts and extends to contaminated, unlabeled settings via latent label optimization.
  • Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). This paper establishes Deep Support Vector Data Description (Deep SVDD) for one-class classification, a standard deep anomaly detection baseline whose assumption of clean data is directly challenged and generalized by the source.
  • Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). This study introduces semi-supervised loss modeling techniques to separate clean and corrupted samples during training, providing core conceptual motivation for handling unlabeled noisy data.
  • Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). This survey provides a comprehensive foundation for learning with label noise and sample selection mechanisms in deep neural networks.
  • Paper: Deep Learning for Anomaly Detection, Guansong Pang et al. (2020). This review provides a structured taxonomy of deep anomaly detection formulations, contextualizing the clean-training assumptions that the source aims to relax.
Cover for Latent Outlier Exposure for Anomaly Detection with Contaminated Data

Abstract

Anomaly detection aims at identifying data points that show systematic deviations from the majority of data in an unlabeled dataset. A common assumption is that clean training data (free of anomalies) is available, which is often violated in practice. We propose a strategy for training an anomaly detector in the presence of unlabeled anomalies that is compatible with a broad class of models. The idea is to jointly infer binary labels to each datum (normal vs. anomalous) while updating the model parameters. Inspired by outlier exposure (Hendrycks et al., 2018) that considers synthetically created, labeled anomalies, we thereby use a combination of two losses that share parameters: one for the normal and one for the anomalous data. We then iteratively proceed with block coordinate updates on the parameters and the most likely (latent) labels. Our experiments with several backbone models on three image datasets, 30 tabular data sets, and a video anomaly detection benchmark showed consistent and significant improvements over the baselines.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Formulation
  • 3.2. Optimization problem
  • 3.3. Model extension and anomaly detection
  • 3.4. Example loss functions
  • 4. Experiments
  • 4.1. Toy Example
  • 4.2. Experiments on Image Data
  • 4.3. Experiments on Tabular Data
  • 4.4. Experiments on Video Data
  • 4.5. Sensitivity Study
  • 5. Conclusion
  • Acknowledgements
  • References
  • A. Details on Toy Data Experiments
  • B. Baseline Details
  • C. Implementation Details
  • D. Additional Experimental Results

Knowls

  1. Knowl 1 — Latent Outlier Exposure Objective for Contaminated Data

    model/method

    In unsupervised anomaly detection with contaminated training data, the training set {xi}i=1N\{x_i\}_{i=1}^N consists of an unknown mixture of normal samples (yi=0y_i = 0) and anomalous samples (yi=1y_i = 1), where yi∈{0,1}y_i \in \{0, 1\} are unobserved latent binary labels with an assumed anomaly contamination fraction α∈[0,1)\alpha \in [0, 1).

    Latent Outlier Exposure (LOE) couples two loss functions parameterized by shared neural network parameters θ\theta:

    1. A normality loss Lnθ(x)≡Ln(fθ(x))\mathcal{L}_n^\theta(x) \equiv \mathcal{L}_n(f_\theta(x)), designed to yield lower values for normal data than for anomalous data.
    2. An anomaly loss Laθ(x)≡La(fθ(x))\mathcal{L}_a^\theta(x) \equiv \mathcal{L}_a(f_\theta(x)), designed to yield low loss values for anomalies and high values for normal data, teaching the feature representation where normal data should not lie.

    Given the latent assignment vector y=(y1,…,yN)Ty = (y_1, \dots, y_N)^T, the joint loss function over the dataset is defined as:

    L(θ,y)=∑i=1N[(1−yi)Lnθ(xi)+yiLaθ(xi)]\mathcal{L}(\theta, y) = \sum_{i=1}^N \left[ (1 - y_i)\mathcal{L}_n^\theta(x_i) + y_i \mathcal{L}_a^\theta(x_i) \right]

    During testing, anomaly scoring uses only the normality loss function Sitest=Lnθ(xi)S_i^{\text{test}} = \mathcal{L}_n^\theta(x_i), as the shared representation parameters θ\theta have already integrated information from both normal and anomalous data during training.

  2. Knowl 2 — Hard Latent Outlier Exposure Constrained Optimization

    equation

    In Hard Latent Outlier Exposure (LOEH\text{LOE}_{\text{H}}), the model jointly optimizes the shared continuous parameters θ\theta and the discrete binary label assignments y=(y1,…,yN)T∈{0,1}Ny = (y_1, \dots, y_N)^T \in \{0, 1\}^N subject to a budget constraint corresponding to a fixed assumed anomaly contamination ratio α∈[0,1)\alpha \in [0, 1) such that αN∈N\alpha N \in \mathbb{N}:

    min⁡θmin⁡y∈Y∑i=1N[(1−yi)Lnθ(xi)+yiLaθ(xi)]\min_\theta \min_{y \in \mathcal{Y}} \sum_{i=1}^N \left[ (1 - y_i)\mathcal{L}_n^\theta(x_i) + y_i \mathcal{L}_a^\theta(x_i) \right]

    where the constrained set Y\mathcal{Y} is:

    Y={y∈{0,1}N:∑i=1Nyi=αN}\mathcal{Y} = \left\{ y \in \{0, 1\}^N : \sum_{i=1}^N y_i = \alpha N \right\}

    For fixed network parameters θ\theta, minimizing the objective over y∈Yy \in \mathcal{Y} is solved exactly by ranking the training anomaly scores:

    Sitrain=Lnθ(xi)−Laθ(xi)S_i^{\text{train}} = \mathcal{L}_n^\theta(x_i) - \mathcal{L}_a^\theta(x_i)

    and assigning yi=0y_i = 0 to the (1−α)(1 - \alpha)-fraction of samples with the lowest training anomaly scores, and yi=1y_i = 1 to the remaining αN\alpha N samples with the highest scores.

  3. Knowl 3 — Soft Latent Outlier Exposure Assignment Formulation

    model/method

    To mitigate overconfidence during discrete label assignment updates when pseudo-labels might be incorrectly identified, Soft Latent Outlier Exposure (LOES\text{LOE}_{\text{S}}) modifies the label assignment constraint set to:

    Y′={y∈{0,0.5}N:∑i=1Nyi=0.5αN}\mathcal{Y}' = \left\{ y \in \{0, 0.5\}^N : \sum_{i=1}^N y_i = 0.5 \alpha N \right\}

    Under LOES\text{LOE}_{\text{S}}, data points identified as the top α\alpha-fraction based on training anomaly scores Sitrain=Lnθ(xi)−Laθ(xi)S_i^{\text{train}} = \mathcal{L}_n^\theta(x_i) - \mathcal{L}_a^\theta(x_i) receive a soft label yi=0.5y_i = 0.5 rather than a hard label yi=1y_i = 1. The per-sample loss for an identified anomalous point becomes an equal combination:

    0.5(Lnθ(xi)+Laθ(xi))0.5 \left( \mathcal{L}_n^\theta(x_i) + \mathcal{L}_a^\theta(x_i) \right)

    This formulation incorporates uncertainty by treating candidate anomalies as equally likely to be normal or abnormal, which prevents the network from aggressively overfitting when normal instances are falsely classified as anomalies.

  4. Knowl 4 — Stochastic Mini-Batch Optimization for Latent Outlier Exposure

    algorithm

    Latent Outlier Exposure optimizes continuous model parameters θ\theta and latent outlier assignments yy using alternating block coordinate updates within stochastic gradient mini-batches.

    Input: Contaminated training dataset D={xi}i=1ND = \{x_i\}_{i=1}^N, assumed contamination ratio α\alpha, deep anomaly detector with parameters θ\theta, batch size MM, number of epochs EE, mode ∈{Hard,Soft}\in \{\text{Hard}, \text{Soft}\}
    Output: Trained model parameters θ\theta
    for epoch = 1 to EE do
        for each mini-batch B⊂DB \subset D of size MM do
            for each xi∈Bx_i \in B do
                Sitrain←Lnθ(xi)−Laθ(xi)S_i^{\text{train}} \leftarrow \mathcal{L}_n^\theta(x_i) - \mathcal{L}_a^\theta(x_i)
            end for
            k←⌊α∣B∣⌋k \leftarrow \lfloor \alpha |B| \rfloor
            Identify indices Itop⊂BI_{\text{top}} \subset B corresponding to the kk largest values of SitrainS_i^{\text{train}}
            for each xi∈Bx_i \in B do
                if i∈Itopi \in I_{\text{top}} then
                    if mode == "Hard" then
                        yi←1.0y_i \leftarrow 1.0
                    else
                        yi←0.5y_i \leftarrow 0.5
                    end if
                else
                    yi←0.0y_i \leftarrow 0.0
                end if
            end for
            Lbatch(θ)←∑xi∈B[(1−yi)Lnθ(xi)+yiLaθ(xi)]\mathcal{L}_{\text{batch}}(\theta) \leftarrow \sum_{x_i \in B} \left[ (1 - y_i)\mathcal{L}_n^\theta(x_i) + y_i \mathcal{L}_a^\theta(x_i) \right]
            θ←θ−η∇θLbatch(θ)\theta \leftarrow \theta - \eta \nabla_\theta \mathcal{L}_{\text{batch}}(\theta)
        end for
    end for
    return θ\theta

    Assuming all component losses Lnθ\mathcal{L}_n^\theta and Laθ\mathcal{L}_a^\theta are bounded from below, the alternating updates monotonically non-increase the joint loss function and converge to a local optimum.

  5. Knowl 5 — Coupled Normal and Anomaly Loss Functions for Self-Supervised Backbones

    model/method

    Latent Outlier Exposure constructs opposing normal and anomaly loss pairs (Lnθ,Laθ)(\mathcal{L}_n^\theta, \mathcal{L}_a^\theta) across various anomaly detection architectures:

    1. Multi-Head RotNet (MHRot): For KK image transformations {T1,…,TK}\{T_1, \dots, T_K\} evaluated across 3 task heads l∈{1,2,3}l \in \{1, 2, 3\} with ground-truth transformation labels tklt_k^l and predicted probability distributions pkl(⋅∣x)p_k^l(\cdot | x): Lnθ(x)=−∑k=1K∑l=13log⁡pkl(tkl∣x)\mathcal{L}_n^\theta(x) = -\sum_{k=1}^K \sum_{l=1}^3 \log p_k^l(t_k^l | x) Laθ(x)=∑k=1K∑l=13CE(U,pkl(⋅∣x))\mathcal{L}_a^\theta(x) = \sum_{k=1}^K \sum_{l=1}^3 \text{CE}\left(U, p_k^l(\cdot | x)\right) where CE(U,⋅)\text{CE}(U, \cdot) is cross-entropy against the uniform distribution UU, forcing uniform head predictions on anomalous samples.

    2. Neural Transformation Learning (NTL): For KK learned neural transformations Tθ,kT_{\theta, k}, temperature τ\tau, and similarity probability pk=exp⁡(cos⁡(fθ(Tθ,k(x)),fθ(x))/τ)exp⁡(cos⁡(fθ(Tθ,k(x)),fθ(x))/τ)+∑l≠kexp⁡(cos⁡(fθ(Tθ,k(x)),fθ(Tθ,l(x)))/τ)p_k = \frac{\exp(\cos(f_\theta(T_{\theta, k}(x)), f_\theta(x))/\tau)}{\exp(\cos(f_\theta(T_{\theta, k}(x)), f_\theta(x))/\tau) + \sum_{l \neq k} \exp(\cos(f_\theta(T_{\theta, k}(x)), f_\theta(T_{\theta, l}(x)))/\tau)}: Lnθ(x)=−∑k=1Klog⁡pk,Laθ(x)=−∑k=1Klog⁡(1−pk)\mathcal{L}_n^\theta(x) = -\sum_{k=1}^K \log p_k, \quad \mathcal{L}_a^\theta(x) = -\sum_{k=1}^K \log(1 - p_k)

    3. Internal Contrastive Learning (ICL): For complementary tabular feature partition embeddings a(x),b(x)a(x), b(x) under encoders fθ,gθf_\theta, g_\theta and pk=exp⁡(cos⁡(fθ(ak(x)),gθ(bk(x)))/τ)∑l=1Kexp⁡(cos⁡(fθ(al(x)),gθ(bk(x)))/τ)p_k = \frac{\exp(\cos(f_\theta(a_k(x)), g_\theta(b_k(x)))/\tau)}{\sum_{l=1}^K \exp(\cos(f_\theta(a_l(x)), g_\theta(b_k(x)))/\tau)}: Lnθ(x)=−∑k=1Klog⁡pk,Laθ(x)=−∑k=1Klog⁡(1−pk)\mathcal{L}_n^\theta(x) = -\sum_{k=1}^K \log p_k, \quad \mathcal{L}_a^\theta(x) = -\sum_{k=1}^K \log(1 - p_k)

    4. Deep Support Vector Data Description (Deep SVDD): Lnθ(x)=∥fθ(x)−c∥2,Laθ(x)=1∥fθ(x)−c∥2\mathcal{L}_n^\theta(x) = \|f_\theta(x) - c\|^2, \quad \mathcal{L}_a^\theta(x) = \frac{1}{\|f_\theta(x) - c\|^2} where cc is the fixed or learned hypersphere center in feature space.

  6. Knowl 6 — Image Anomaly Detection Performance on Contaminated CIFAR-10, Fashion-MNIST, and MVTec AD

    data/table

    When evaluated on contaminated training sets with a 10%10\% anomaly contamination ratio (α0=α=0.10\alpha_0 = \alpha = 0.10), Latent Outlier Exposure (LOEH\text{LOE}_{\text{H}} and LOES\text{LOE}_{\text{S}}) mitigates the performance degradation observed in Blind training (treating all data as normal) and Refine (filtering top anomalous samples without exposure loss). Results are reported as mean ±\pm standard deviation of Area Under the ROC Curve (AUC in %) over 3 runs; numbers in parentheses indicate the performance gap relative to models trained on 100% clean data.

    Method CIFAR-10 (NTL) F-MNIST (NTL) CIFAR-10 (MHRot) F-MNIST (MHRot)
    Blind 91.3 ±\pm 0.1 (-4.4) 85.0 ±\pm 0.2 (-9.7) 84.0 ±\pm 0.5 (-4.2) 88.8 ±\pm 0.1 (-4.9)
    Refine 93.5 ±\pm 0.1 (-2.2) 89.1 ±\pm 0.2 (-5.6) 84.4 ±\pm 0.1 (-3.8) 89.6 ±\pm 0.2 (-4.1)
    LOEH\text{LOE}_{\text{H}} (ours) 94.9 ±\pm 0.2 (-0.8) 92.9 ±\pm 0.7 (-1.8) 86.4 ±\pm 0.5 (-1.8) 91.4 ±\pm 0.2 (-2.3)
    LOES\text{LOE}_{\text{S}} (ours) 94.9 ±\pm 0.1 (-0.8) 92.5 ±\pm 0.1 (-2.2) 86.3 ±\pm 0.2 (-1.9) 91.2 ±\pm 0.4 (-2.5)

    On MVTec AD using NTL at 10%10\% contamination, LOEH\text{LOE}_{\text{H}} achieves 95.9±0.9%95.9 \pm 0.9\% AUC for anomaly detection (compared to 94.2±0.5%94.2 \pm 0.5\% for Blind and 95.3±0.5%95.3 \pm 0.5\% for Refine) and LOES\text{LOE}_{\text{S}} achieves 96.56±0.04%96.56 \pm 0.04\% AUC for anomaly segmentation (compared to 96.17±0.08%96.17 \pm 0.08\% for Blind and 96.55±0.04%96.55 \pm 0.04\% for Refine).

  7. Knowl 7 — Video Frame Anomaly Detection Performance on Contaminated UCSD Peds1

    data/table

    On the UCSD Peds1 video anomaly detection benchmark, where abnormal pedestrian walkway frames contaminate the training videos, Soft Latent Outlier Exposure (LOES\text{LOE}_{\text{S}}) with a Neural Transformation Learning (NTL) backbone achieves state-of-the-art frame-level Area Under the ROC Curve (AUC in %).

    Method Contamination Ratio
    10% 20% 30%
    Sugiyama Borgwardt (2013) 55.0 56.0 56.3
    Del Giorno et al. (2016) - - 59.6
    Tudor Ionescu et al. (2017) - - 68.4
    Liu et al. (2018) - - 69.0
    Pang et al. (2020) 68.0 70.0 71.7
    Blind 85.2 ±\pm 1.0 76.0 ±\pm 2.7 66.6 ±\pm 2.6
    Refine 82.7 ±\pm 1.5 74.9 ±\pm 2.4 69.3 ±\pm 0.7
    LOEH\text{LOE}_{\text{H}} (ours) 82.3 ±\pm 1.6 59.6 ±\pm 3.8 56.8 ±\pm 9.5
    LOES\text{LOE}_{\text{S}} (ours) 86.8 ±\pm 1.2 79.2 ±\pm 1.3 71.5 ±\pm 2.4

    LOES\text{LOE}_{\text{S}} outperforms Deep Ordinal Regression (Pang et al., 2020) by 18.8%18.8\% AUC at 10%10\% contamination and by 9.2%9.2\% AUC at 20%20\% contamination. While hard assignments in LOEH\text{LOE}_{\text{H}} suffer degradation at higher contamination ratios (56.8%56.8\% AUC at 30%30\%), LOES\text{LOE}_{\text{S}} remains robust (71.5%71.5\% AUC), highlighting the value of uncertainty-aware soft weighting for video frames.

  8. Knowl 8 — Tabular Anomaly Detection Performance Across 30 Contaminated Benchmarks

    empirical result

    Across 30 benchmark tabular datasets evaluated at 10%10\% contamination ratio (α0=α=0.10\alpha_0 = \alpha = 0.10) over five independent runs with Neural Transformation Learning (NTL) and Internal Contrastive Learning (ICL) backbones:

    1. Consistent Improvements over Baselines: LOEH\text{LOE}_{\text{H}} and LOES\text{LOE}_{\text{S}} consistently outperform the Blind baseline (which ignores contamination) and the Refine baseline (which filters out suspected anomalies). For example, on the Arrhythmia dataset with NTL, LOES\text{LOE}_{\text{S}} achieves an F1-score of 62.7±3.3%62.7 \pm 3.3\% versus 57.6±2.5%57.6 \pm 2.5\% (Blind) and 59.1±2.1%59.1 \pm 2.1\% (Refine); on Thyroid with NTL, LOES\text{LOE}_{\text{S}} achieves 82.4±2.3%82.4 \pm 2.3\% versus 43.4±5.5%43.4 \pm 5.5\% (Blind) and 55.1±4.2%55.1 \pm 4.2\% (Refine).

    2. Gains Beyond Clean Data Performance: On several datasets, LOE trained on contaminated data achieves higher F1-scores than models trained on 100% clean data. On Thyroid, LOEH\text{LOE}_{\text{H}} reaches an F1-score of 82.4%82.4\% (+4.6% relative to clean data training) with NTL and 83.2%83.2\% (+6.0% relative to clean data training) with ICL. On Speech with NTL, LOES\text{LOE}_{\text{S}} achieves 50.8%50.8\% (+41.3% over clean data training). This shows that unlabeled anomalies, when correctly identified and exposed during training, provide an active contrastive learning signal that tightens normal decision boundaries beyond what clean data alone provides.

  9. Knowl 9 — Sensitivity of LOE to Hyperparameter Contamination Ratio Mis-specification

    empirical result

    In realistic unsupervised scenarios, the true anomaly contamination ratio α0\alpha_0 is unknown. Evaluating LOEH\text{LOE}_{\text{H}} and LOES\text{LOE}_{\text{S}} with NTL on CIFAR-10 across varying true contamination ratios α0∈{5%,10%,15%,20%}\alpha_0 \in \{5\%, 10\%, 15\%, 20\%\} and assumed hyperparameters α∈{5%,10%,15%,20%}\alpha \in \{5\%, 10\%, 15\%, 20\%\} demonstrates high robustness:

    1. Low Degradation Under Error: LOEH\text{LOE}_{\text{H}} degrades by at most 1.4%1.4\% AUC when α\alpha is mis-specified by 5%5\% relative to α0\alpha_0 (e.g., obtaining 93.6%93.6\% AUC at α0=5%,α=10%\alpha_0 = 5\%, \alpha = 10\% compared to 95.0%95.0\% when α=5%\alpha = 5\% is exact).
    2. Complementary Robustness Regimes:
      • LOEH\text{LOE}_{\text{H}} is particularly robust when α≤α0\alpha \le \alpha_0 (underestimating the contamination ratio), consistently outperforming the Refine baseline.
      • LOES\text{LOE}_{\text{S}} is robust against overestimating the contamination ratio (α≥α0\alpha \ge \alpha_0), retaining 94.6%94.6\% AUC at α0=5%\alpha_0 = 5\% when α=20%\alpha = 20\%.
    3. Baseline Superiority: Both LOEH\text{LOE}_{\text{H}} and LOES\text{LOE}_{\text{S}} consistently outperform the Blind baseline across all pairs (α0,α)(\alpha_0, \alpha) and outperform Refine across the vast majority of grid settings.

Coverage note — The 2D Gaussian mixture toy visualization experiment was omitted as a standalone knowl because its conclusions are purely illustrative and are rigorously quantified by the image, tabular, and video empirical knowls.

References

  1. 1.Alvarez, M., Verdier, J.-C., Nkashama, D. K., Frappier, M., Tardif, P.-M., and Kabanza, F. A revealing large-scale evaluation of unsupervised anomaly detection algorithms. arXiv preprint arXiv:2204.09825, 2022.
  2. 2.Beggel, L., Pfeiffer, M., and Bischl, B. Robust anomaly detection in images using adversarial autoencoders. arXiv preprint arXiv:1901.06355, 2019.
  3. 3.Bergman, L. and Hoshen, Y. Classification-based anomaly detection for general data. In International Conference on Learning Representations, 2020.
  4. 4.Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9592–9600, 2019.
  5. 5.Chen, X. and Konukoglu, E. Unsupervised detection of lesions in brain mri using constrained adversarial autoencoders. In MIDL Conference book. MIDL, 2018.
  6. 6.Deecke, L., Vandermeulen, R., Ruff, L., Mandt, S., and Kloft, M. Image anomaly detection with generative adversarial networks. In Joint european conference on machine learning and knowledge discovery in databases, pp. 3–17. Springer, 2018.
  7. 7.Defard, T., Setkov, A., Loesch, A., and Audigier, R. Padim: a patch distribution modeling framework for anomaly detection and localization. In ICPR 2020-25th International Conference on Pattern Recognition Workshops and Challenges, 2021.
  8. 8.Del Giorno, A., Bagnell, J. A., and Hebert, M. A discriminative framework for anomaly detection in large videos. In European Conference on Computer Vision, pp. 334–349. Springer, 2016.
  9. 9.Feng, J.-C., Hong, F.-T., and Zheng, W.-S. Mist: Multiple instance self-training framework for video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14009–14018, 2021.
  10. 10.Golan, I. and El-Yaniv, R. Deep anomaly detection using geometric transformations. In Advances in Neural Information Processing Systems, pp. 9758–9769, 2018.
  11. 11.Görnitz, N., Porbadnigk, A., Binder, A., Sannelli, C., Braun, M., Müller, K.-R., and Kloft, M. Learning and evaluation in presence of non-iid label noise. In Artificial Intelligence and Statistics, pp. 293–302. PMLR, 2014.
  12. 12.Hendrycks, D., Mazeika, M., and Dietterich, T. Deep anomaly detection with outlier exposure. In International Conference on Learning Representations, 2018.
  13. 13.Hendrycks, D., Mazeika, M., Kadavath, S., and Song, D. Using self-supervised learning can improve model robustness and uncertainty. Advances in Neural Information Processing Systems, 32:15663–15674, 2019.
  14. 14.Huber, P. J. Robust estimation of a location parameter. In Breakthroughs in statistics, pp. 492–518. Springer, 1992.
  15. 15.Huber, P. J. Robust statistics. In International encyclopedia of statistical science, pp. 1248–1251. Springer, 2011.
  16. 16.Huyan, N., Quan, D., Zhang, X., Liang, X., Chanussot, J., and Jiao, L. Unsupervised outlier detection using memory and contrastive learning. arXiv preprint arXiv:2107.12642, 2021.
  17. 17.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  18. 18.Lee, W. S. and Liu, B. Learning with positive and unlabeled examples using weighted logistic regression. In ICML, volume 3, pp. 448–455, 2003.
  19. 19.Li, C.-L., Sohn, K., Yoon, J., and Pfister, T. Cutpaste: Self-supervised learning for anomaly detection and localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9664–9674, 2021.
  20. 20.Liu, F. T., Ting, K. M., and Zhou, Z.-H. Isolation-based anomaly detection. ACM Transactions on Knowledge Discovery from Data (TKDD), 6(1):1–39, 2012.
  21. 21.Liu, Y., Li, C.-L., and Póczos, B. Classifier two sample test for video anomaly detections. In BMVC, pp. 71, 2018.
  22. 22.Nalisnick, E., Matsukawa, A., Teh, Y. W., Gorur, D., and Lakshminarayanan, B. Do deep generative models know what they don't know? In International Conference on Learning Representations, 2018.
  23. 23.Pang, G., Yan, C., Shen, C., Hengel, A. v. d., and Bai, X. Self-trained deep ordinal regression for end-to-end video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12173–12182, 2020.
  24. 24.Principi, E., Vesperini, F., Squartini, S., and Piazza, F. Acoustic novelty detection with adversarial autoencoders. In 2017 International Joint Conference on Neural Networks (IJCNN), pp. 3324–3330. IEEE, 2017.
  25. 25.Qiu, C., Pfrommer, T., Kloft, M., Mandt, S., and Rudolph, M. Neural transformation learning for deep anomaly detection beyond images. In International Conference on Machine Learning, pp. 8703–8714. PMLR, 2021.
  26. 26.Qiu, C., Kloft, M., Mandt, S., and Rudolph, M. Raising the bar in graph-level anomaly detection. arXiv preprint arXiv:2205.13845, 2022.
  27. 27.Reiss, T., Cohen, N., Bergman, L., and Hoshen, Y. Panda: Adapting pretrained features for anomaly detection and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2806–2814, 2021.
  28. 28.Ruff, L., Vandermeulen, R., Goernitz, N., Deecke, L., Siddiqui, S. A., Binder, A., Müller, E., and Kloft, M. Deep one-class classification. In International conference on machine learning, pp. 4393–4402. PMLR, 2018.
  29. 29.Ruff, L., Vandermeulen, R. A., Görnitz, N., Binder, A., Müller, E., Müller, K.-R., and Kloft, M. Deep semi-supervised anomaly detection. In International Conference on Learning Representations, 2019.
  30. 30.Ruff, L., Kauffmann, J. R., Vandermeulen, R. A., Montavon, G., Samek, W., Kloft, M., Dietterich, T. G., and Müller, K.-R. A unifying review of deep and shallow anomaly detection. Proceedings of the IEEE, 2021.
  31. 31.Schlegl, T., Seeböck, P., Waldstein, S. M., Schmidt-Erfurth, U., and Langs, G. Unsupervised anomaly detection with generative adversarial networks to guide marker discovery. In International conference on information processing in medical imaging, pp. 146–157. Springer, 2017.
  32. 32.Schneider, T., Qiu, C., Kloft, M., Latif, D. A., Staab, S., Mandt, S., and Rudolph, M. Detecting anomalies within time series using local neural transformations. arXiv preprint arXiv:2202.03944, 2022.
  33. 33.Schölkopf, B., Platt, J. C., Shawe-Taylor, J., Smola, A. J., and Williamson, R. C. Estimating the support of a high-dimensional distribution. Neural computation, 13(7):1443–1471, 2001.
  34. 34.Shenkar, T. and Wolf, L. Anomaly detection for tabular data with internal contrastive learning. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=_hszZbt46bT.
  35. 35.Sohn, K., Li, C.-L., Yoon, J., Jin, M., and Pfister, T. Learning and evaluating representations for deep one-class classification. arXiv preprint arXiv:2011.02578, 2020.
  36. 36.Sugiyama, M. and Borgwardt, K. Rapid distance-based outlier detection via sampling. Advances in Neural Information Processing Systems, 26:467–475, 2013.
  37. 37.Tudor Ionescu, R., Smeureanu, S., Alexe, B., and Popescu, M. Unmasking the abnormal events in video. In Proceedings of the IEEE international conference on computer vision, pp. 2895–2903, 2017.
  38. 38.Wang, S., Zeng, Y., Liu, X., Zhu, E., Yin, J., Xu, C., and Kloft, M. Effective end-to-end unsupervised outlier detection via inlier priority of discriminative network. In Advances in Neural Information Processing Systems, pp. 5962–5975, 2019.
  39. 39.Xia, Y., Cao, X., Wen, F., Hua, G., and Sun, J. Learning discriminative reconstructions for unsupervised outlier removal. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1511–1519, 2015.
  40. 40.Yoon, J., Sohn, K., Li, C.-L., Arik, S. O., Lee, C.-Y., and Pfister, T. Self-trained one-class classification for unsupervised anomaly detection. arXiv preprint arXiv:2106.06115, 2021.
  41. 41.Zhou, C. and Paffenroth, R. C. Anomaly detection with robust deep autoencoders. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 665–674, 2017.

Citation

MLA
Qiu, C., et al. “Latent Outlier Exposure for Anomaly Detection with Contaminated Data”. International Conference on Machine Learning, vol. 162, 2022, pp. 18153–67, https://proceedings.mlr.press/v162/qiu22b.html.
APA
Qiu, C., Li, A., Kloft, M., Rudolph, M., & Mandt, S. (2022). Latent Outlier Exposure for Anomaly Detection with Contaminated Data. International Conference on Machine Learning, 162, 18153–18167. https://proceedings.mlr.press/v162/qiu22b.html
Chicago
Qiu, C., A. Li, M. Kloft, M. Rudolph, and S. Mandt. 2022. “Latent Outlier Exposure for Anomaly Detection with Contaminated Data”. International Conference on Machine Learning 162: 18153–67. https://proceedings.mlr.press/v162/qiu22b.html.
Harvard
Qiu, C. et al. (2022) “Latent Outlier Exposure for Anomaly Detection with Contaminated Data”, International Conference on Machine Learning. PMLR, pp. 18153–18167. Available at: https://proceedings.mlr.press/v162/qiu22b.html.
Vancouver
1. Qiu C, Li A, Kloft M, Rudolph M, Mandt S (2022) Latent Outlier Exposure for Anomaly Detection with Contaminated Data. In: International Conference on Machine Learning. PMLR, pp 18153–18167

BibTeX

@InProceedings{pmlr-v162-qiu22b,
  title = 	 {Latent Outlier Exposure for Anomaly Detection with Contaminated Data},
  author =       {Qiu, Chen and Li, Aodong and Kloft, Marius and Rudolph, Maja and Mandt, Stephan},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {18153--18167},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/qiu22b/qiu22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/qiu22b.html},
  abstract = 	 {Anomaly detection aims at identifying data points that show systematic deviations from the majority of data in an unlabeled dataset. A common assumption is that clean training data (free of anomalies) is available, which is often violated in practice. We propose a strategy for training an anomaly detector in the presence of unlabeled anomalies that is compatible with a broad class of models. The idea is to jointly infer binary labels to each datum (normal vs. anomalous) while updating the model parameters. Inspired by outlier exposure (Hendrycks et al., 2018) that considers synthetically created, labeled anomalies, we thereby use a combination of two losses that share parameters: one for the normal and one for the anomalous data. We then iteratively proceed with block coordinate updates on the parameters and the most likely (latent) labels. Our experiments with several backbone models on three image datasets, 30 tabular data sets, and a video anomaly detection benchmark showed consistent and significant improvements over the baselines.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/