SoftPatch: Unsupervised Anomaly Detection with Noisy Data

Xi JiangJianlin LiuJinbao WangQiang NieKai WuYong LiuChengjie WangFeng Zheng

article2022NeurIPS116 citations

Proposes SoftPatch, a patch-level denoising and memory re-weighting method that prevents defective training samples from distorting decision boundaries in real-world unsupervised anomaly detection.

Listen

Industrial quality inspection and sensory anomaly detection rely heavily on automated visual models trained without labeled defect examples. Standard unsupervised methods build reference distributions from baseline training sets under the rigid assumption that all training images are entirely clean and defect-free. In practical manufacturing environments, however, defective parts and mislabeled data inevitably contaminate nominal training collections due to human error and operational drift. When conventional models ingest these corruptions, their reference memory banks become overconfident and misinformed, causing them to overlook identical defects during operational testing and sharply degrading quality assurance.

The article demonstrates that memory-based visual inspection systems can maintain high detection accuracy even when training data contains significant noise, introducing a robust unsupervised inspection approach called SoftPatch to address label-level noise in visual inspection tasks.

To evaluate this framework, the authors conducted empirical experiments using standard industrial inspection benchmarks, including MVTecAD (15 categories across 5,354 images) and BTAD (three categories across 1,799 images). They systematically injected varying ratios of defective images (up to 15%) into nominal training data across two evaluation setups: one where injected anomaly types differed from test targets, and an overlap setup where identical defect types appeared in both training noise and test sets. SoftPatch addresses contamination by extracting position-grouped patch features from a standard neural network backbone, calculating local outlier scores to prune the most anomalous image regions rather than discarding entire images, and applying soft anomaly re-weighting factors to remaining reference samples.

The evaluation revealed several key findings:

  1. Conventional state-of-the-art methods collapse under realistic noise: In the presence of a 10% noise overlap, standard memory baseline PatchCore suffered a severe accuracy drop of approximately 30 to 40 percentage points across image-level and pixel-level defect localization.
  2. SoftPatch preserves robust detection across noisy conditions: Utilizing Local Outlier Factor (LOF) patch-level filtering, SoftPatch achieved a 0.982 image-level Area Under the ROC Curve (AUROC) and a 0.969 localization AUROC under 10% noise overlap, exhibiting virtually no performance drop compared to noise-free baselines.
  3. Native industrial datasets already harbor performance-limiting noise: On the unpolluted BTAD industrial benchmark, SoftPatch outperformed prior state-of-the-art models (achieving an average AUROC of 0.977 versus 0.957 for PatchCore) because it naturally identified and filtered preexisting minor scratch defects in the benchmark's baseline training images.
  4. Patch-level denoising outperforms whole-image filtering: Removing specific defective regions while retaining non-defective patches from contaminated images maximized training data utilization without sacrificing structural context.

These findings indicate that manufacturing operations can deploy visual quality inspection models directly into new production lines without requiring extensive, costly, and error-prone manual data-cleaning phases. By mitigating the vulnerability of greedy memory-bank selection to outlier contamination, industrial systems substantially reduce the operational risk of missed defects, compliance failures, and false-positive shutdowns.

For practical deployment, organizations should adopt patch-level outlier filtering alongside local density-based re-weighting when building memory-based vision systems. Operating teams can safely implement fixed parameter defaults (such as a 15% patch elimination threshold) without needing prior knowledge of exact factory noise levels. Future technical development should explore fully unsupervised deployments and assess model performance when defective distributions change dynamically across continuous inspection streams.

While the reported results demonstrate high statistical confidence across multiple repeated trials on established industrial benchmarks, slight performance trade-offs were observed on select image categories where spatial features naturally exhibit wide misalignment. Inspection teams deploying this architecture on highly unconstrained or shifting visual scenes should validate spatial density thresholds during initial pilot phases.

arXiv: 2403.14233
Cover for SoftPatch: Unsupervised Anomaly Detection with Noisy Data

Abstract

Although mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection but is seldom discussed. This paper considers label-level noise in image sensory anomaly detection for the first time. To solve this problem, we proposed a memory-based unsupervised AD method, SoftPatch, which efficiently denoises the data at the patch level. Noise discriminators are utilized to generate outlier scores for patch-level noise elimination before coreset construction. The scores are then stored in the memory bank to soften the anomaly detection boundary. Compared with existing methods, SoftPatch maintains a strong modeling ability of normal data and alleviates the overconfidence problem in coreset. Comprehensive experiments in various noise scenes demonstrate that SoftPatch outperforms the state-of-the-art AD methods on the MVTecAD and BTAD benchmarks and is comparable to those methods under the setting without noise.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Unsupervised Anomaly Detection
  • 2.2 Learning with Noisy Data
  • 3 The Proposed Method
  • 3.1 Overview
  • 3.2 Noise Discriminative Coreset Selection
  • 3.2.1 Nearest Neighbor
  • 3.2.2 Multi-Variate Gaussian
  • 3.2.3 Local Outlier Factor (LOF)
  • 3.3 Anomaly Detection based on SoftPatch
  • 4 Experiments
  • 4.1 Experimental Details
  • 4.2 Anomaly Detection Performance with Noise
  • 4.3 Ablation Study
  • 4.3.1 Effectiveness of the Proposed Modules
  • 4.3.2 Parameter Selection
  • 5 Conclusions
  • Acknowledgement
  • References
  • Checklist

Knowls

  1. Knowl 1 — SoftPatch Framework for Noise-Robust Memory-Bank Anomaly Detection

    model/method

    Unsupervised image sensory anomaly detection models typically assume the training dataset X={x1,x2,…,xN}\mathcal{X} = \{x_1, x_2, \dots, x_N\} comprises solely nominal (defect-free) images. When defective (noisy) images contaminate the training data, standard memory-bank approaches suffer severe performance collapse because greedy coreset subsampling aggressively selects anomalous features as prototypes.

    SoftPatch mitigates training noise at the patch level rather than the image level, exploiting the domain property that sensory defects occupy localized regions while the remainder of a contaminated image is nominal. The method operates in four sequential stages:

    1. Feature Extraction and Spatial Grouping: For each training image xi∈RC×H×Wx_i \in \mathbb{R}^{C \times H \times W}, a pre-trained convolutional backbone extracts feature map ϕi∈Rc∗×h∗×w∗\phi_i \in \mathbb{R}^{c^* \times h^* \times w^*}. Features across all NN images are partitioned by spatial coordinates (h,w)(h, w) to construct position-specific patch pools {ϕi(h,w)}i=1N⊂Rc∗\{\phi_i(h, w)\}_{i=1}^N \subset \mathbb{R}^{c^*}.
    2. Patch-Level Noise Discrimination: A noise discriminator evaluates each patch ϕi(h,w)\phi_i(h, w) against other patches at position (h,w)(h, w) across the dataset, computing an outlier score Wi(h,w)W_i(h, w). Patches exhibiting the top τ%\tau\% outlier scores across the dataset are pruned, while the remaining (1−τ)%(1 - \tau)\% nominal patches are retained.
    3. Coreset Selection with Soft Weighting: A compact memory bank M\mathcal{M} is constructed from the retained clean patches using greedy coreset subsampling. Each selected patch prototype m∈Mm \in \mathcal{M} retains its pre-computed discriminator score WmW_m as a soft confidence weight.
    4. Soft-Weighted Anomaly Scoring: At inference, test image patches search for their nearest neighbors in M\mathcal{M}. The raw Euclidean distance is multiplied by the prototype's soft weight WmW_m, dampening overconfidence and preventing residual noise prototypes from misclassifying anomalies.
  2. Knowl 2 — Local Outlier Factor Patch Noise Discriminator

    model/method

    To handle non-uniform feature cluster densities and high-dimensional embeddings across spatial locations, SoftPatch computes patch outlier scores using Local Outlier Factor (LOF).

    Let ϕi(h,w)∈Rc∗\phi_i(h, w) \in \mathbb{R}^{c^*} denote the patch feature at spatial position (h,w)(h, w) for image xix_i. Let Nk(ϕi(h,w))N_k(\phi_i(h, w)) represent the set of kk-nearest neighbors of ϕi(h,w)\phi_i(h, w) among all patches at spatial position (h,w)(h, w) across images j≠ij \neq i, and let distk(ϕb(h,w))dist_k(\phi_b(h, w)) be the Euclidean distance from ϕb(h,w)\phi_b(h, w) to its kk-th nearest neighbor in that position pool.

    1. The kk-distance reachability distance from ϕi(h,w)\phi_i(h, w) to neighbor ϕb(h,w)\phi_b(h, w) is defined as: distkreach(ϕi(h,w),ϕb(h,w))=max⁡(distk(ϕb(h,w)),∥ϕi(h,w)−ϕb(h,w)∥2)dist_k^{reach}(\phi_i(h, w), \phi_b(h, w)) = \max\left(dist_k(\phi_b(h, w)), \|\phi_i(h, w) - \phi_b(h, w)\|_2\right)

    2. The local reachability density (lrdlrd) of patch ϕi(h,w)\phi_i(h, w) is the inverse of the average reachability distance from its kk-nearest neighbors: lrdi(h,w)=(1∣Nk(ϕi(h,w))∣∑ϕb(h,w)∈Nk(ϕi(h,w))distkreach(ϕi(h,w),ϕb(h,w)))−1lrd_i(h, w) = \left( \frac{1}{|N_k(\phi_i(h, w))|} \sum_{\phi_b(h, w) \in N_k(\phi_i(h, w))} dist_k^{reach}(\phi_i(h, w), \phi_b(h, w)) \right)^{-1}

    3. The Local Outlier Factor score WiLOF(h,w)W_i^{LOF}(h, w), measuring the relative density of the neighborhood compared to the patch itself, is defined as: WiLOF(h,w)=∑ϕb(h,w)∈Nk(ϕi(h,w))lrdb(h,w)∣Nk(ϕi(h,w))∣⋅lrdi(h,w)W_i^{LOF}(h, w) = \frac{\sum_{\phi_b(h, w) \in N_k(\phi_i(h, w))} lrd_b(h, w)}{|N_k(\phi_i(h, w))| \cdot lrd_i(h, w)}

    A score WiLOF(h,w)≈1.0W_i^{LOF}(h, w) \approx 1.0 indicates that the patch is embedded within a cluster of comparable density (nominal inlier), while WiLOF(h,w)>1.0W_i^{LOF}(h, w) > 1.0 indicates that the patch resides in a sparser local density region than its neighbors (outlier noise).

  3. Knowl 3 — Soft-Weighted Anomaly Scoring at Test Time

    equation

    In the inference stage of SoftPatch, a test image x∈Xtestx \in \mathcal{X}_{test} is evaluated against the denoised memory bank M⊂Rc∗\mathcal{M} \subset \mathbb{R}^{c^*} and its associated prototype soft weights {Wm}m∈M\{W_m\}_{m \in \mathcal{M}}.

    Let P(x)={ph,w∈Rc∗:1≤h≤h∗,1≤w≤w∗}\mathcal{P}(x) = \{p_{h,w} \in \mathbb{R}^{c^*} : 1 \le h \le h^*, 1 \le w \le w^*\} denote the patch feature representations extracted from test image xx.

    1. Nearest-Neighbor Prototype Search: For each test patch ph,w∈P(x)p_{h,w} \in \mathcal{P}(x), the closest memory prototype m∗∈Mm^* \in \mathcal{M} is identified via Euclidean distance: m∗=arg⁡min⁡m∈M∥ph,w−m∥2m^* = \arg\min_{m \in \mathcal{M}} \|p_{h,w} - m\|_2

    2. Patch-Level Anomaly Score: The patch anomaly score sh,ws_{h,w} is computed by scaling the Euclidean distance by the prototype's soft outlier weight Wm∗W_{m^*}: sh,w=Wm∗∥ph,w−m∗∥2s_{h,w} = W_{m^*} \|p_{h,w} - m^*\|_2

    3. Image-Level Anomaly Score: The overall anomaly score s∗s^* for test sample xx is defined as the maximum re-weighted patch anomaly score across all spatial locations: s∗=max⁡(h,w)sh,ws^* = \max_{(h, w)} s_{h,w}

    By scaling the distance with Wm∗W_{m^*}, any noisy prototype that inadvertently survived pruning will possess an elevated Wm∗>1.0W_{m^*} > 1.0, preventing test anomalies near that prototype from obtaining deceptively low anomaly scores.

  4. Knowl 4 — Gaussian and Nearest-Neighbor Patch Noise Discriminators

    equation

    In addition to LOF, SoftPatch defines two alternative spatial patch noise discriminators on position-grouped patch embeddings {ϕi(h,w)}i=1N⊂Rc∗\{\phi_i(h, w)\}_{i=1}^N \subset \mathbb{R}^{c^*}:

    1. Nearest-Neighbor (NN) Distance Discriminator: Assuming anomaly patches are sparse compared to nominal data, the outlier score Winn(h,w)W_i^{nn}(h, w) is the minimum Euclidean distance to another image's patch at the identical spatial position (h,w)(h, w): Winn(h,w)=min⁡n∈[1,N],n≠i∥ϕi(h,w)−ϕn(h,w)∥2W_i^{nn}(h, w) = \min_{n \in [1, N], n \neq i} \|\phi_i(h, w) - \phi_n(h, w)\|_2

    2. Multi-Variate Gaussian (MVG) Discriminator: A single Gaussian N(μh,w,Σh,w)\mathcal{N}(\mu_{h,w}, \Sigma_{h,w}) is fitted across all NN training patches at position (h,w)(h, w): μh,w=1N∑n=1Nϕn(h,w)\mu_{h,w} = \frac{1}{N} \sum_{n=1}^N \phi_n(h, w) Σh,w=1N−1∑n=1N(ϕn(h,w)−μh,w)(ϕn(h,w)−μh,w)T+ϵI\Sigma_{h,w} = \frac{1}{N - 1} \sum_{n=1}^N (\phi_n(h, w) - \mu_{h,w})(\phi_n(h, w) - \mu_{h,w})^T + \epsilon I where ϵI\epsilon I is a regularization term ensuring Σh,w\Sigma_{h,w} is full rank and invertible. The outlier magnitude Wimvg(h,w)W_i^{mvg}(h, w) is the Mahalanobis distance: Wimvg(h,w)=(ϕi(h,w)−μh,w)TΣh,w−1(ϕi(h,w)−μh,w)W_i^{mvg}(h, w) = \sqrt{(\phi_i(h, w) - \mu_{h,w})^T \Sigma_{h,w}^{-1} (\phi_i(h, w) - \mu_{h,w})}

    While MVG normalizes variance across positions, it assumes a unimodal distribution and can mistakenly flag small nominal clusters as outliers when objects exhibit spatial variation.

  5. Knowl 5 — SoftPatch Training and Inference Procedure

    algorithm

    The complete execution flow of SoftPatch using the Local Outlier Factor (LOF) discriminator is defined below:

    Input: Training set X={x1,…,xN}\mathcal{X} = \{x_1, \dots, x_N\}, test image xtestx_{test}, feature extractor ff, pruning ratio τ\tau, neighborhood size kk, coreset sampling ratio rr
    Output: Image anomaly score s∗s^*, pixel anomaly map S∈Rh∗×w∗S \in \mathbb{R}^{h^* \times w^*}
    1. Feature Extraction and Spatial Grouping:
       for each image xi∈Xx_i \in \mathcal{X} do
           ϕi←f(xi)∈Rc∗×h∗×w∗\phi_i \leftarrow f(x_i) \in \mathbb{R}^{c^* \times h^* \times w^*}
       end for
    2. Position-wise Outlier Scoring:
       for each spatial position (h,w)∈[1,h∗]×[1,w∗](h, w) \in [1, h^*] \times [1, w^*] do
           Fh,w←{ϕi(h,w):i∈[1,N]}\mathcal{F}_{h,w} \leftarrow \{\phi_i(h, w) : i \in [1, N]\}
           for each patch ϕi(h,w)∈Fh,w\phi_i(h, w) \in \mathcal{F}_{h,w} do
               Identify kk-nearest neighbors Nk(ϕi(h,w))⊂Fh,w∖{ϕi(h,w)}N_k(\phi_i(h, w)) \subset \mathcal{F}_{h,w} \setminus \{\phi_i(h, w)\}
               Compute reachability density lrdi(h,w)lrd_i(h, w) from kk-reachability distances
           end for
           for each patch ϕi(h,w)∈Fh,w\phi_i(h, w) \in \mathcal{F}_{h,w} do
               Wi(h,w)←1∣Nk(ϕi(h,w))∣∑ϕb∈Nk(ϕi(h,w))lrdb(h,w)lrdi(h,w)W_i(h, w) \leftarrow \frac{1}{|N_k(\phi_i(h, w))|} \sum_{\phi_b \in N_k(\phi_i(h, w))} \frac{lrd_b(h, w)}{lrd_i(h, w)}
           end for
       end for
    3. Patch-Level Denoising:
       Ω←{(ϕi(h,w),Wi(h,w)):i∈[1,N],h∈[1,h∗],w∈[1,w∗]}\Omega \leftarrow \{(\phi_i(h, w), W_i(h, w)) : i \in [1, N], h \in [1, h^*], w \in [1, w^*]\}
       Determine threshold θ\theta equal to the (100−τ)(100 - \tau)-th percentile of all scores Wi(h,w)W_i(h, w) in Ω\Omega
       Ωclean←{(ϕ,W)∈Ω:W<θ}\Omega_{clean} \leftarrow \{(\phi, W) \in \Omega : W < \theta\}
    4. Coreset Selection and Memory Bank Storage:
       Φclean←{ϕ:(ϕ,W)∈Ωclean}\Phi_{clean} \leftarrow \{\phi : (\phi, W) \in \Omega_{clean}\}
       Select coreset M⊂Φclean\mathcal{M} \subset \Phi_{clean} using greedy facility location to retain fraction rr
       Mbank←{(m,Wm):m∈M}\mathcal{M}_{bank} \leftarrow \{(m, W_m) : m \in \mathcal{M}\} where WmW_m is the stored outlier score of prototype mm
    5. Inference on Test Sample:
       ϕtest←f(xtest)\phi_{test} \leftarrow f(x_{test})
       for each patch ph,w=ϕtest(h,w)p_{h,w} = \phi_{test}(h, w) do
           m∗←arg⁡min⁡(m,Wm)∈Mbank∥ph,w−m∥2m^* \leftarrow \arg\min_{(m, W_m) \in \mathcal{M}_{bank}} \|p_{h,w} - m\|_2
           sh,w←Wm∗⋅∥ph,w−m∗∥2s_{h,w} \leftarrow W_{m^*} \cdot \|p_{h,w} - m^*\|_2
       end for
       s∗←max⁡(h,w)sh,ws^* \leftarrow \max_{(h, w)} s_{h,w}
       S←{sh,w}h=1,w=1h∗,w∗S \leftarrow \{s_{h,w}\}_{h=1, w=1}^{h^*, w^*} (bilinearly upsampled to H×WH \times W)
       return s∗,Ss^*, S
  6. Knowl 6 — Noisy Sensory Anomaly Detection Benchmark Protocols: No-Overlap versus Overlap

    experimental setup

    To evaluate unsupervised anomaly detection under label contamination, training sets are corrupted by sampling anomalous images randomly from the test set and mixing them into the nominal training set at noise ratios n∈[0.0,0.15]n \in [0.0, 0.15] (e.g., n=0.1n=0.1 denotes 10% added negative images, termed MVTecAD-noise-n). The total number of nominal training images remains unchanged.

    Two evaluation settings are established:

    1. No Overlap Setting: The injected anomalous images are removed from the test set during evaluation. This represents scenarios where the test set contains novel defect instances unseen during noisy training.
    2. Overlap Setting: The injected anomalous images remain in the test set. This directly evaluates the risk of identical or appearance-similar anomalies recurring at test time, testing whether the model falsely learned the training anomalies as nominal prototypes.

    Standard Configurations:

    • Datasets: MVTecAD (15 categories; 3,629 nominal training images, 1,725 test images) and BTAD (3 categories; 1,799 total images).
    • Network Architecture: Pre-trained Wide-ResNet50 backbone; images resized to 256×256256 \times 256 and center-cropped to 224×224224 \times 224 (MVTecAD) or resized to 512×512512 \times 512 (BTAD).
    • Hyperparameters: Patch pruning ratio τ=0.15\tau = 0.15, LOF neighborhood size k=6k = 6, coreset subsampling ratio r=10%r = 10\%. Experiments are averaged over three independent runs on Nvidia V100 GPUs.
  7. Knowl 7 — Anomaly Detection and Localization Performance on MVTecAD under 10% Label Noise

    data/table

    Under 10% injected training noise (MVTecAD-noise-0.1), SoftPatch-LOF preserves anomaly detection and localization performance, whereas standard memory-bank baselines experience catastrophic failure in the Overlap setting.

    No overlap Overlap
    Category PaDiM CFLOW PatchCore SoftPatch-nn SoftPatch-mvg SoftPatch-lof PaDiM* PatchCore PatchCore-rnd SoftPatch-lof
    bottle 0.994 0.998 1.000 1.000 0.997 0.937 1.000 0.692 0.998 1.000
    cable 0.873 0.925 0.982 0.935 0.952 0.995 0.680 0.756 0.920 0.994
    capsule 0.920 0.947 0.976 0.916 0.662 0.963 0.796 0.783 0.779 0.955
    carpet 0.999 0.961 0.996 0.995 0.999 0.991 0.890 0.681 0.973 0.993
    grid 0.966 0.891 0.971 0.972 0.997 0.968 0.674 0.526 0.793 0.969
    hazelnut 0.956 1.000 0.998 1.000 1.000 1.000 0.543 0.441 0.998 1.000
    leather 1.000 1.000 1.000 1.000 1.000 1.000 0.964 0.739 1.000 1.000
    metal_nut 0.987 0.959 0.999 0.994 0.997 0.999 0.820 0.765 0.969 1.000
    pill 0.918 0.929 0.975 0.921 0.873 0.963 0.722 0.770 0.874 0.955
    screw 0.838 0.784 0.966 0.862 0.475 0.960 0.567 0.710 0.462 0.923
    tile 0.977 0.991 0.985 0.996 0.997 0.993 0.830 0.716 1.000 0.981
    toothbrush 0.927 0.906 0.997 1.000 0.997 0.997 0.700 0.800 0.797 0.994
    transistor 0.953 0.896 0.953 1.000 0.992 0.990 0.471 0.491 0.943 0.999
    wood 0.991 0.972 0.984 0.984 0.997 0.987 0.831 0.579 0.980 0.986
    zipper 0.852 0.928 0.981 0.976 0.979 0.978 0.679 0.792 0.950 0.974
    Image AUROC 0.943 0.939 0.984 0.970 0.927 0.986 0.740 0.683 0.896 0.982
    Image Gap -0.007 -0.030 -0.008 +0.002 -0.001 0.000 -0.151 -0.309 -0.015 -0.004
    Pixel AUROC 0.972 0.969 0.956 0.971 0.977 0.979 0.955 0.654 0.951 0.969
    Pixel Gap -0.007 -0.006 -0.025 -0.008 -0.001 -0.002 -0.013 -0.327 -0.021 -0.012

    PatchCore suffers a 30.9% Image AUROC drop (0.683) and 32.7% Pixel AUROC drop (0.654) in the Overlap setting because greedy coreset selection preferentially picks anomalous features as representative centers. SoftPatch-LOF maintains 0.982 Image AUROC and 0.969 Pixel AUROC in the Overlap setting (performance drop of only 0.4% and 1.2%, respectively).

  8. Knowl 8 — Anomaly Detection Performance on BTAD Benchmark with Natural Training Noise

    data/table

    The original BTAD benchmark contains naturally occurring defect noise in the nominal training set (specifically small scratches in category BTAD-02). SoftPatch achieves superior performance without any artificial noise injection.

    Category SPADE P-SVDD PatchCore PaDiM SoftPatch (ours) Anomaly samples
    01 0.914 0.957 1.000 1.000 0.999 50
    02 0.714 0.721 0.871 0.871 0.934 200
    03 0.999 0.821 0.999 0.971 0.997 41
    Mean 0.876 0.833 0.957 0.947 0.977 -

    On category BTAD-02, where 200 anomalous test images share appearance similarity with the unannotated training scratches, baseline methods like PatchCore and PaDiM plateau at 0.871 Image AUROC. SoftPatch denoises these scratch patches, improving performance on category 02 to 0.934 AUROC (+6.3%) and achieving an overall mean Image AUROC of 0.977.

  9. Knowl 9 — Ablation of Noise Discriminator Architecture and Soft Outlier Re-Weighting

    data/table

    An ablation study on MVTecAD evaluates the isolated and combined contributions of patch noise discrimination and soft outlier re-weighting under 10% label noise.

    No overlap Overlap
    Noise discriminator Soft weight Image AUROC Pixel AUROC Image AUROC Pixel AUROC
    None 0.985 0.946 0.685 0.693
    Gaussian 0.927 0.977 0.925 0.961
    Gaussian ✓ 0.922 0.974 0.924 0.965
    Nearest 0.970 0.971 0.966 0.944
    Nearest ✓ 0.972 0.978 0.968 0.958
    LOF 0.985 0.984 0.984 0.963
    LOF ✓ 0.986 0.979 0.982 0.969
    • Incorporating a noise discriminator is the primary driver of robustness in the Overlap setting, raising Image AUROC from 0.685 (no discriminator) to 0.984 (LOF alone).
    • LOF outperforms Gaussian and Nearest-Neighbor discriminators by avoiding errors caused by multi-modal feature layouts and density variations.
    • Soft re-weighting provides supplementary gains (e.g., improving Overlap Pixel AUROC from 0.963 to 0.969 with LOF, and from 0.944 to 0.958 with Nearest) by penalizing residual hard noise that escapes the pruning threshold.
  10. Knowl 10 — Sensitivity of SoftPatch to Neighborhood Size k and Pruning Threshold tau

    empirical result

    SoftPatch exhibits stable performance across varying choices of the Local Outlier Factor neighborhood size kk and patch pruning threshold τ\tau on MVTecAD-noise-0.1:

    1. Neighborhood Size kk: Evaluating k∈{3,4,5,6,7,8,9}k \in \{3, 4, 5, 6, 7, 8, 9\} yields:
    • Overlap Image/Pixel AUROC: 0.983/0.9550.983/0.955 (k=3k=3), 0.982/0.9510.982/0.951 (k=4k=4), 0.983/0.9590.983/0.959 (k=5k=5), 0.982/0.9750.982/0.975 (k=6k=6), 0.981/0.9730.981/0.973 (k=7k=7), 0.982/0.9680.982/0.968 (k=8k=8), 0.980/0.9680.980/0.968 (k=9k=9).
    • No overlap Image/Pixel AUROC: 0.985/0.9720.985/0.972 (k=3k=3), 0.985/0.9750.985/0.975 (k=4k=4), 0.984/0.9770.984/0.977 (k=5k=5), 0.984/0.9820.984/0.982 (k=6k=6), 0.985/0.9800.985/0.980 (k=7k=7), 0.984/0.9830.984/0.983 (k=8k=8), 0.981/0.9820.981/0.982 (k=9k=9). Performance is robust for all k≥5k \ge 5. Small k<5k < 5 fails to reliably estimate local density, while excessively large kk risks connecting separate clusters across modal boundaries.
    1. Pruning Ratio τ\tau: As τ\tau increases across {0.01,0.05,0.09,0.13,0.15,0.17}\{0.01, 0.05, 0.09, 0.13, 0.15, 0.17\}, Overlap Image AUROC rises monotonically from 0.880.88 to 0.9820.982 and Pixel AUROC rises from 0.880.88 to 0.9690.969, as higher thresholds aggressively prune corrupted features before coreset selection. Under No-overlap, AUROC remains stable ( ≈0.985\,\approx 0.985 Image,  ≈0.980\,\approx 0.980 Pixel) across all τ\tau values. A fixed default of k=6k=6 and τ=0.15\tau=0.15 achieves near-optimal performance across diverse noise scenes.
  11. Knowl 11 — Limitations in Spatial Misalignment and Multi-Modal Nominal Variations

    limitation

    SoftPatch relies on grouping patch embeddings strictly by their spatial grid coordinates (h,w)(h, w). This design introduces specific operational limitations:

    1. Spatial Misalignment Sensitivity: For object categories with unaligned rotations, translations, or substantial pose variations (such as the screw class in MVTecAD), patches at identical coordinates (h,w)(h, w) belong to different physical subparts. This results in multi-modal feature distributions where valid nominal patches in minority poses are assigned high outlier scores and pruned (e.g., SoftPatch-Gaussian achieves only 0.475 Image AUROC on screw).
    2. Clean-Data Feature Under-Coverage: Applying a fixed pruning ratio (e.g., τ=0.15\tau = 0.15) to clean training sets discards nominal features, causing a marginal baseline performance drop (approximately 0.006 Image AUROC) compared to unconstrained PatchCore.

Coverage note — None was omitted; all key contributions, algorithms, equations, experimental setups, benchmark data tables, ablation studies, hyperparameter analyses, and limitations are fully documented.

References

  1. 1.Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey. arXiv preprint arXiv:2110.11334, 2021.
  2. 2.Xian Tao, Xinyi Gong, Xin Zhang, Shaohua Yan, and Chandranath Adak. Deep learning for unsupervised anomaly localization in industrial images: A survey. IEEE Transactions on Instrumentation and Measurement, 2022.
  3. 3.Minghui Yang, Peng Wu, Jing Liu, and Hui Feng. Memseg: A semi-supervised method for image surface defect detection using differences and commonalities. arXiv preprint arXiv:2205.00908, 2022.
  4. 4.Vitjan Zavrtanik, Matej Kristan, and Danijel Skocaj. Dsr - a dual subspace re-projection network for surface anomaly detection. CoRR, abs/2208.01521, 2022.
  5. 5.Choubo Ding, Guansong Pang, and Chunhua Shen. Catching both gray and black swans: Open-set supervised anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7388–7398, 2022.
  6. 6.Nicolae-Cătălin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B Moeslund, and Mubarak Shah. Self-supervised predictive convolutional attentive block for anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13576–13586, 2022.
  7. 7.Xudong Yan, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu, and Pheng-Ann Heng. Learning semantic context from normal samples for unsupervised anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3110–3118, 2021.
  8. 8.Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf, Thomas Brox, and Peter Gehler. Towards total recall in industrial anomaly detection. arXiv preprint arXiv:2106.08265, 2021.
  9. 9.Denis Gudovskiy, Shun Ishizaka, and Kazuki Kozuka. Cflow-ad: Real-time unsupervised anomaly detection with localization via conditional normalizing flows. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 98–107, 2022.
  10. 10.Ye Zheng, Xiang Wang, Rui Deng, Tianpeng Bao, Rui Zhao, and Liwei Wu. Focus your distribution: Coarse-to-fine non-contrastive learning for anomaly detection and localization. In 2022 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE, 2022.
  11. 11.Paul Bergmann, Michael Fauser, David Sattlegger, and Carsten Steger. Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9592–9600, 2019.
  12. 12.Pankaj Mishra, Riccardo Verk, Daniele Fornasier, Claudio Piciarelli, and Gian Luca Foresti. Vt-adl: A vision transformer network for image anomaly detection and localization. In 2021 IEEE 30th International Symposium on Industrial Electronics (ISIE), pages 01–06. IEEE, 2021.
  13. 13.Shelly Sheynin, Sagie Benaim, and Lior Wolf. A hierarchical transformation-discriminating generative model for few shot anomaly detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 8495–8504, October 2021.
  14. 14.Chun-Liang Li, Kihyuk Sohn, Jinsung Yoon, and Tomas Pfister. Cutpaste: Self-supervised learning for anomaly detection and localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9664–9674, 2021.
  15. 15.Vitjan Zavrtanik, Matej Kristan, and Danijel Skočaj. Draem-a discriminatively trained reconstruction embedding for surface anomaly detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8330–8339, 2021.
  16. 16.Paul Bergmann, Michael Fauser, David Sattlegger, and Carsten Steger. Uninformed students: Student-teacher anomaly detection with discriminative latent embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4183–4192, 2020.
  17. 17.Mohammadreza Salehi, Niousha Sadjadi, Soroosh Baselizadeh, Mohammad H Rohban, and Hamid R Rabiee. Multiresolution knowledge distillation for anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14902–14912, 2021.
  18. 18.Hanqiu Deng and Xingyu Li. Anomaly detection via reverse distillation from one-class embedding. arXiv preprint arXiv:2201.10703, 2022.
  19. 19.Kang Zhou, Yuting Xiao, Jianlong Yang, Jun Cheng, Wen Liu, Weixin Luo, Zaiwang Gu, Jiang Liu, and Shenghua Gao. Encoding structure-texture relation with p-net for anomaly detection in retinal images. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, pages 360–377. Springer, 2020.
  20. 20.Yufei Liang, Jiangning Zhang, Shiwei Zhao, Runze Wu, Yong Liu, and Shuwen Pan. Omni-frequency channel-selection representations for unsupervised anomaly detection. arXiv preprint arXiv:2203.00259, 2022.
  21. 21.Thomas Defard, Aleksandr Setkov, Angelique Loesch, and Romaric Audigier. Padim: a patch distribution modeling framework for anomaly detection and localization. In International Conference on Pattern Recognition, pages 475–489. Springer, 2021.
  22. 22.Jihun Yi and Sungroh Yoon. Patch svdd: Patch-level svdd for anomaly detection and segmentation. In Proceedings of the Asian Conference on Computer Vision, 2020.
  23. 23.Marco Rudolph, Bastian Wandt, and Bodo Rosenhahn. Same same but differnet: Semi-supervised defect detection with normalizing flows. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1907–1916, 2021.
  24. 24.Tal Reiss, Niv Cohen, Liron Bergman, and Yedid Hoshen. Panda: Adapting pretrained features for anomaly detection and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2806–2814, 2021.
  25. 25.Sungwook Lee, Seunghyun Lee, and Byung Cheol Song. Cfa: Coupled-hypersphere-based feature adaptation for target-oriented anomaly localization. arXiv preprint arXiv:2206.04325, 2022.
  26. 26.Jinlei Hou, Yingying Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, and Hong Zhou. Divide-and-assemble: Learning block-wise memory for unsupervised anomaly detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8791–8800, 2021.
  27. 27.Zijian Hu, Zhengyu Yang, Xuefeng Hu, and Ram Nevatia. Simple: Similar pseudo label exploitation for semi-supervised classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15099–15108, June 2021.
  28. 28.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Advances in Neural Information Processing Systems, 33:596–608, 2020.
  29. 29.Junnan Li, Richard Socher, and Steven CH Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394, 2020.
  30. 30.Shuming Kong, Yanyan Shen, and Linpeng Huang. Resolving training biases via influence-based data relabeling. In International Conference on Learning Representations, 2021.
  31. 31.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3060–3069, 2021.
  32. 32.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. arXiv preprint arXiv:2102.09480, 2021.
  33. 33.Fan Yang, Kai Wu, Shuyi Zhang, Guannan Jiang, Yong Liu, Feng Zheng, Wei Zhang, Chengjie Wang, and Long Zeng. Class-aware contrastive semi-supervised learning. arXiv preprint arXiv:2203.02261, 2022.
  34. 34.Songqiao Han, Xiyang Hu, Hailiang Huang, Mingqi Jiang, and Yue Zhao. Adbench: Anomaly detection benchmark. arXiv preprint arXiv:2206.09426, 2022.
  35. 35.Guansong Pang, Cheng Yan, Chunhua Shen, Anton van den Hengel, and Xiao Bai. Self-trained deep ordinal regression for end-to-end video anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12173–12182, 2020.
  36. 36.Boyang Liu, Ding Wang, Kaixiang Lin, Pang-Ning Tan, and Jiayu Zhou. Rca: A deep collaborative autoencoder approach for anomaly detection. In IJCAI, pages 1505–1511, 2021.
  37. 37.Chong Zhou and Randy C Paffenroth. Anomaly detection with robust deep autoencoders. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pages 665–674, 2017.
  38. 38.Shuang Wu, Jingyu Zhao, and GuangJian Tian. Understanding and mitigating data contamination in deep anomaly detection: A kernel-based approach. In IJCAI, pages 2319–2325, 2022.
  39. 39.Jinsung Yoon, Kihyuk Sohn, Chun-Liang Li, Sercan O Arik, Chen-Yu Lee, and Tomas Pfister. Self-supervise, refine, repeat: Improving unsupervised anomaly detection. 2021.
  40. 40.Chen Qiu, Aodong Li, Marius Kloft, Maja Rudolph, and Stephan Mandt. Latent outlier exposure for anomaly detection with contaminated data. arXiv preprint arXiv:2202.08088, 2022.
  41. 41.Antoine Cordier, Benjamin Missaoui, and Pierre Gutierrez. Data refinement for fully unsupervised visual inspection using pre-trained networks. arXiv preprint arXiv:2202.12759, 2022.
  42. 42.Yuanhong Chen, Yu Tian, Guansong Pang, and Gustavo Carneiro. Deep one-class classification via interpolated gaussian descriptor. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 383–392, 2022.
  43. 43.Leif E Peterson. K-nearest neighbor. Scholarpedia, 4(2):1883, 2009.
  44. 44.Markus M Breunig, Hans-Peter Kriegel, Raymond T Ng, and Jörg Sander. Lof: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data, pages 93–104, 2000.

Citation

MLA
Jiang, X., et al. “SoftPatch: Unsupervised Anomaly Detection with Noisy Data”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 15433–45, https://proceedings.neurips.cc/paper_files/paper/2022/file/637a456d89289769ac1ab29617ef7213-Paper-Conference.pdf.
APA
Jiang, X., Liu, J., Wang, J., Nie, Q., WU, K., Liu, Y., Wang, C., & Zheng, F. (2022). SoftPatch: Unsupervised Anomaly Detection with Noisy Data. Advances in Neural Information Processing Systems, 35, 15433–15445. https://proceedings.neurips.cc/paper_files/paper/2022/file/637a456d89289769ac1ab29617ef7213-Paper-Conference.pdf
Chicago
Jiang, X., J. Liu, J. Wang, et al. 2022. “SoftPatch: Unsupervised Anomaly Detection with Noisy Data”. Advances in Neural Information Processing Systems 35: 15433–45. https://proceedings.neurips.cc/paper_files/paper/2022/file/637a456d89289769ac1ab29617ef7213-Paper-Conference.pdf.
Harvard
Jiang, X. et al. (2022) “SoftPatch: Unsupervised Anomaly Detection with Noisy Data”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 15433–15445. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/637a456d89289769ac1ab29617ef7213-Paper-Conference.pdf.
Vancouver
1. Jiang X, Liu J, Wang J, Nie Q, WU K, Liu Y, Wang C, Zheng F (2022) SoftPatch: Unsupervised Anomaly Detection with Noisy Data. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 15433–15445

BibTeX

@inproceedings{jiang2022softpatch,
  title = {SoftPatch: Unsupervised Anomaly Detection with Noisy Data},
  author = {Jiang, Xi and Liu, Jianlin and Wang, Jinbao and Nie, Qiang and WU, Kai and Liu, Yong and Wang, Chengjie and Zheng, Feng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {15433-15445},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/637a456d89289769ac1ab29617ef7213-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors