Shape-Guided Dual-Memory Learning for 3D Anomaly Detection

Yu-Min ChuChieh LiuTing-I HsiehHwann-Tzong ChenTyng-Luh Liu

article2023ICML101 citations

Proposes a shape-guided dual-memory framework that combines neural implicit signed distance fields with aligned 2D visual features to achieve state-of-the-art unsupervised 3D anomaly detection and localization on the MVTec 3D-AD benchmark.

Listen

Industrial quality inspection and healthcare imaging increasingly require automated systems to detect subtle flaws in manufactured components and anatomical structures. Traditional visual inspection systems rely heavily on two-dimensional color imagery, which frequently fails when defects involve structural distortions without color changes or subtle surface blemishes against complex textures. While three-dimensional point clouds capture geometric structure, combining spatial depth with color imagery has historically introduced severe computational bottlenecks, high memory consumption, and unacceptable false alarm rates under strict industrial tolerances.

The main objective of the article is to demonstrate an unsupervised inspection framework—termed shape-guided dual-memory learning—that accurately detects and localizes defects by fusing three-dimensional geometry with two-dimensional color appearance without requiring defective samples during training.

To accomplish this, the authors developed a specialized two-expert architecture evaluated across 10 object categories from the MVTec 3D-AD industrial benchmark, comprising 2,656 defect-free training samples and 1,197 testing samples. The shape expert divides point clouds into localized patches and applies neural implicit functions to model local surface geometry via signed distance fields, establishing a baseline of normal geometric structures. The appearance expert maps these three-dimensional patches directly to corresponding two-dimensional color image features, building paired memory banks of normal product features. During inspection, test items are evaluated against these dual memory banks using sparse mathematical reconstructions to flag deviations, followed by an automated score alignment mechanism that fuses structural and color anomaly maps.

The experimental findings show substantial improvements over existing inspection methods. First, the framework achieved state-of-the-art anomaly localization and sample-level detection, reaching an image-level area under the receiver operating characteristic curve of 0.947 and a localization score of 0.976. Second, it demonstrated exceptional precision at ultra-low false positive limits, achieving a localization score of 0.456 at a strict 1% false positive limit compared to 0.394 for the leading alternative baseline. Third, the shape-guided integration reduced memory consumption dramatically, utilizing only 13.5% of the memory footprint of an unguided approach and processing inspections at 0.69 frames per second (2.05 seconds per sample) compared to 0.46 frames per second for prior multi-modal baselines.

These results demonstrate that coordinating two-dimensional color search through three-dimensional spatial cues overcomes the performance and efficiency trade-offs that have previously hindered multi-modal defect detection. For decision-makers, this translates into lower computational hardware costs, reduced false alarms on factory lines, and reliable detection of complex flaws such as structural dents or surface discolorations that evade single-modality sensors.

Organizations evaluating automated optical inspection pipelines should consider testing shape-guided dual-expert architectures, particularly where strict false alarm tolerances and complex geometries are present. Before production deployment, engineering teams should conduct pilot studies to optimize local patch sizing (setting patch sizes to roughly 500 points offered the best balance of accuracy and latency) and evaluate deployment on target processing hardware.

Confidence in these findings is supported by comprehensive quantitative benchmarks and ablation studies across diverse industrial product categories. However, decision-makers should note that the evaluation is currently bounded by the MVTec 3D-AD dataset, and the system can still struggle when defects simultaneously present elusive, low-contrast visual features and minimal geometric disruption.

Chu et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for Shape-Guided Dual-Memory Learning for 3D Anomaly Detection

Abstract

We present a shape-guided expert-learning framework to tackle the problem of unsupervised 3D anomaly detection. Our method is established on the effectiveness of two specialized expert models and their synergy to localize anomalous regions from color and shape modalities. The first expert utilizes geometric information to probe 3D structural anomalies by modeling the implicit distance fields around local shapes. The second expert considers the 2D RGB features associated with the first expert to identify color appearance irregularities on the local shapes. We use the two experts to build the dual memory banks from the anomaly-free training samples and perform shape-guided inference to pinpoint the defects in the testing samples. Owing to the per-point 3D representation and the effective fusion scheme of complementary modalities, our method efficiently achieves state-of-the-art performance on the MVTec 3D-AD dataset with better recall and lower false positive rates, as preferred in real applications.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. 2D Anomaly Detection
  • 2.2. 3D Anomaly Detection
  • 3. Method
  • 3.1. Shape-Guided Expert Learning
  • 3.2. Shape-Guided Inference
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Implementation Detail
  • 4.3. Evaluation Metrics
  • 4.4. Experimental Results
  • 4.5. Computational Complexity
  • 4.6. Patch Size Analysis
  • 4.7. Qualitative Results
  • 5. Additional Ablation
  • 5.1. Benefit of Combining RGB and 3D Information
  • 5.2. The Effectiveness of Sparse Coding
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Shape-guided dual-memory framework

    model/method

    The method detects anomalies in 3D point clouds by combining two complementary experts trained from normal data. A shape expert represents local 3D geometry with signed distance functions (SDFs) and stores patch-level features in a geometry memory bank, MSM_S. An appearance expert associates each SDF with RGB feature vectors from the same spatial region and stores these dictionaries in an appearance memory bank, MAM_A. At inference, the geometry memory guides which RGB dictionaries are used to reconstruct normal appearance; the resulting geometric and RGB anomaly maps are calibrated and fused pixelwise by taking their maximum. This design targets both structural defects and color irregularities, including defects detectable in only one modality.

  2. Knowl 2 — Local shape expert based on implicit signed distance fields

    model/method

    For each local point-cloud patch, PointNet maps sampled 3D points to a feature vector ff describing the patch geometry. A shared neural implicit function (NIF) ϕ takes a query point qq and the patch feature ff and predicts the signed distance s=ϕ(q;f)s=ϕ(q;f) from qq to the patch surface. The pair (ϕ,f)(ϕ,f) therefore represents a patch-specific SDF, while the NIF is shared across patches; after training on normal patches, the method stores the patch features ff in the SDF memory bank MSM_S.

    Patches are formed by farthest-point sampling (FPS) over a point cloud and taking the KK nearest points around each FPS center. Patches may overlap, and their number is adjusted so their union covers as much of the cloud as possible. The reported default is K=500K=500, with an overlapping ratio of 10; a cloud of about 7,500 points yields roughly 150 patches. During NIF training, 20 query points are sampled around each real point. The PointNet implementation has three convolutional layers and two fully connected layers, each with batch normalization; the NIF is an MLP.

    The patch-size sweep selected K=500K=500 as an accuracy–runtime compromise. For K=1000,750,500,250K=1000,750,500,250, respectively, image AUROC was 0.894,0.929,0.947,0.9660.894,0.929,0.947,0.966; AUPRO at FPR integration limit 0.30.3 was 0.965,0.973,0.976,0.9750.965,0.973,0.976,0.975; AUPRO at limit 0.010.01 was 0.413,0.439,0.456,0.4510.413,0.439,0.456,0.451; and inference time was 1.72,1.88,2.05,3.611.72,1.88,2.05,3.61 seconds per sample.

  3. Knowl 3 — Shape-guided RGB appearance memory

    model/method

    The appearance expert constructs an RGB dictionary for each normal patch-level SDF. For the 500 points used to form a patch's PointNet input, their corresponding 2D image locations are found and used to retrieve feature vectors from an RGB feature map. The mapped image locations are expanded by a two-pixel neighborhood to include nearby RGB features, helping cover defect boundaries and empty regions created by holes or cracks. Each patch SDF is associated with about 40–60 RGB feature vectors; these SDF-specific dictionaries form the appearance memory bank MAM_A, with one dictionary corresponding to each SDF feature in MSM_S. The RGB map is extracted with a Wide ResNet-50-2, using features from its first and second layers.

  4. Knowl 4 — Shape-guided inference with sparse feature reconstruction

    model/method

    For a test sample, PointNet extracts patch features and ResNet extracts the RGB feature map. Image pixels associated with at least one point-cloud patch are treated as foreground. For each test patch feature, the method finds its k1=10k_1=10 nearest features in the normal SDF memory bank MSM_S and uses them as a dictionary for sparse reconstruction. The reconstructed feature is passed to the NIF to predict signed distances for the 3D points in that patch's receptive field; absolute signed distances are combined across patches to form the SDF anomaly map.

    The shape neighbors selected for those reconstructions identify corresponding SDF-specific RGB dictionaries in MAM_A. Their union forms a shape-guided RGB dictionary. Each foreground RGB feature is reconstructed sparsely using its k2=5k_2=5 nearest entries in that dictionary, and the ℓ2ℓ_2 reconstruction distance supplies its RGB anomaly score. After RGB-score calibration, the final per-pixel anomaly score is the maximum of the SDF and RGB scores.

  5. Knowl 5 — Calibration of RGB and SDF anomaly scores

    model/method

    Because RGB and SDF anomaly scores have different scales, the method aligns them before taking their pixelwise maximum. It simulates inference on 25 randomly selected, anomaly-free training samples and uses a leave-one-out procedure so that a query's own features are excluded from its nearest-neighbor searches. An affine map y↦ay+by\mapsto ay+b is fitted to the RGB scores so that their mean plus or minus three standard deviations maps to the corresponding mean plus or minus three standard deviations of the SDF scores. The resulting scale aa and offset bb are then fixed and applied to RGB scores at test time; no anomalous training examples are required.

  6. Knowl 6 — MVTec 3D-AD experimental protocol

    experimental setup

    The evaluation uses MVTec 3D-AD, which has ten object categories and paired high-resolution point clouds and RGB images. Its 2,656 training and 294 validation samples are anomaly-free; its 1,197 test samples comprise 249 normal and 948 anomalous samples, with about four to five defect types per category. The point clouds are background-cropped, and point clouds and images are resized from 800×800800\times800 to 224×224224\times224 using nearest-neighbor and bicubic interpolation, respectively. Training patches from normal point clouds train PointNet and the NIF, while the corresponding RGB data build the two memory banks. The reported learning rate is 0.00010.0001 and batch size is 32. RGB features are arranged as a 28×2828\times28 feature map. Localization is evaluated with AUPRO, the area under the per-region-overlap curve integrated to a specified false-positive-rate limit; image- and pixel-level detection are evaluated with AUROC.

  7. Knowl 7 — Benchmark performance on detection and localization

    empirical result

    On MVTec 3D-AD, the full shape-guided method reports mean image AUROC 0.9470.947, compared with 0.9370.937 for AST using RGB and depth; the RGB-only and SDF-only versions of the proposed method report mean image AUROC 0.8150.815 and 0.9160.916, respectively. In category order Bagel, Cable Gland, Carrot, Cookie, Dowel, Foam, Peach, Potato, Rope, Tire, the full method's image AUROC values are 0.986,0.894,0.983,0.991,0.976,0.857,0.990,0.965,0.960,0.8690.986, 0.894, 0.983, 0.991, 0.976, 0.857, 0.990, 0.965, 0.960, 0.869; AST RGB+Depth reports 0.983,0.873,0.976,0.971,0.932,0.885,0.974,0.981,1.000,0.7970.983, 0.873, 0.976, 0.971, 0.932, 0.885, 0.974, 0.981, 1.000, 0.797.

    For localization at the standard FPR integration limit of 0.30.3, the full method's mean AUPRO is 0.9760.976, versus 0.9640.964 for BTF RGB+FPFH. In the same category order, the full method's AUPRO values are 0.981,0.973,0.982,0.971,0.962,0.978,0.981,0.983,0.974,0.9750.981, 0.973, 0.982, 0.971, 0.962, 0.978, 0.981, 0.983, 0.974, 0.975; BTF RGB+FPFH reports 0.976,0.967,0.979,0.974,0.971,0.884,0.976,0.981,0.959,0.9710.976, 0.967, 0.979, 0.974, 0.971, 0.884, 0.976, 0.981, 0.959, 0.971. The paper also reports that the full method achieves state-of-the-art AUPRO at the standard limit and at very low limits, where lower FPR tolerance is important for precise localization.

  8. Knowl 8 — Inference speed and RGB memory use

    empirical result

    On an Nvidia GTX 1080, the shape-guided method takes 2.052.05 seconds per sample (0.690.69 FPS), searches over 26,452 features, uses 13.5% RGB memory, and reports image AUROC 0.9470.947, pixel AUROC 0.9960.996, AUPRO at limit 0.30.3 of 0.9760.976, and AUPRO at limit 0.010.01 of 0.4560.456. Without shape guidance, the corresponding figures are 3.603.60 seconds (0.290.29 FPS), 208,230 features, 100% RGB memory, 0.9470.947, 0.9960.996, 0.9760.976, and 0.4530.453. BTF reports 2.192.19 seconds (0.460.46 FPS), 20,823 features, 10% RGB memory, 0.8730.873, 0.9930.993, 0.9640.964, and 0.3940.394. The comparison indicates that shape guidance reduces the RGB search and memory burden relative to the un-guided variant while improving its speed and low-FPR AUPRO. The reported memory percentages concern RGB features.

  9. Knowl 9 — Ablation shows complementary contributions from RGB and geometry

    empirical result

    At AUPRO integration limit 0.30.3, mean performance on MVTec 3D-AD is 0.9330.933 for RGB only, 0.9310.931 for SDF only, 0.9320.932 for shape-guided RGB only, and 0.9760.976 for shape-guided RGB+SDF. The combined system outperforms either single modality and RGB alone with shape guidance, supporting the use of complementary appearance and geometric cues. The Foam category illustrates this complementarity: RGB-only, SDF-only, and shape-guided RGB-only AUPRO are 0.7760.776, 0.7730.773, and 0.7740.774, while the combined system reaches 0.9780.978.

  10. Knowl 10 — Sparse reconstruction outperforms nearest-neighbor scoring

    empirical result

    An ablation compared using the nearest normal feature directly (NN) with using a sparse-coded reconstruction (SC) for RGB and SDF anomaly scoring. With NN for both RGB and SDF, image AUROC is 0.9370.937, pixel AUROC 0.9940.994, AUPRO at FPR limit 0.30.3 is 0.9720.972, and AUPRO at limit 0.010.01 is 0.4340.434. Using NN for RGB and SC for SDF gives 0.9440.944, 0.9940.994, 0.9720.972, and 0.4430.443; using SC for RGB and NN for SDF gives 0.9420.942, 0.9960.996, 0.9750.975, and 0.4510.451. Using SC for both gives the best reported values: 0.9470.947, 0.9960.996, 0.9760.976, and 0.4560.456, respectively. The paper attributes the benefit to sparse reconstruction better representing normal features than a single nearest neighbor.

Coverage note — The patch-size sensitivity analysis is included with the shape-expert knowl. Qualitative comparison figures and the reported failure examples are not separate knowls because they provide illustrative cases rather than a quantified comparison or a detailed failure taxonomy.

References

  1. 1.Behrendt, F., Bengs, M., Rogge, F., Krüger, J., Opfer, R., and Schlaefer, A. Unsupervised anomaly detection in 3d brain MRI using deep learning with impured training data. In ISBI, 2022.
  2. 2.Bengs, M., Behrendt, F., Laves, M., Krüger, J., Opfer, R., and Schlaefer, A. Unsupervised anomaly detection in 3d brain MRI using deep learning with multi-task brain age prediction. CoRR, 2022.
  3. 3.Bergmann, P. and Sattlegger, D. Anomaly detection in 3d point clouds using deep geometric descriptors. CoRR, 2022.
  4. 4.Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. MVTec AD - A comprehensive real-world dataset for unsupervised anomaly detection. In CVPR, 2019.
  5. 5.Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. Uninformed students: Student-teacher anomaly detection with discriminative latent embeddings. In CVPR, 2020.
  6. 6.Bergmann, P., Batzner, K., Fauser, M., Sattlegger, D., and Steger, C. The MVTec anomaly detection dataset: A comprehensive real-world dataset for unsupervised anomaly detection. IJCV, 2021.
  7. 7.Bergmann, P., Jin, X., Sattlegger, D., and Steger, C. The MVTec 3D-AD dataset for unsupervised 3d anomaly detection and localization. In 17th Int. Joint Conf. Comput. Vis., Imaging and Comput. Graph. Theory and Appl., VISIGRAPP 2022, Volume 5: VISAPP, 2022.
  8. 8.Defard, T., Setkov, A., Loesch, A., and Audigier, R. Padim: A patch distribution modeling framework for anomaly detection and localization. In ICPR, 2020.
  9. 9.Gudovskiy, D. A., Ishizaka, S., and Kozuka, K. CFLOW-AD: real-time unsupervised anomaly detection with localization via conditional normalizing flows. In WACV, 2022.
  10. 10.Horwitz, E. and Hoshen, Y. Back to the feature: Classical 3d features are (almost) all you need for 3d anomaly detection. CoRR, abs/2203.05550v3, 2022.
  11. 11.Jiang, C. M., Sud, A., Makadia, A., Huang, J., Nießner, M., and Funkhouser, T. A. Local implicit grid representations for 3d scenes. In CVPR, 2020.
  12. 12.Lee, S., Lee, S., and Song, B. C. CFA: coupled-hypersphere-based feature adaptation for target-oriented anomaly localization. IEEE Access, 2022.
  13. 13.Li, C., Sohn, K., Yoon, J., and Pfister, T. Cutpaste: Self-supervised learning for anomaly detection and localization. In CVPR, 2021.
  14. 14.Li, K., Tang, Y., Prisacariu, V. A., and Torr, P. H. S. Bnv-fusion: Dense 3d reconstruction using bi-level neural volume fusion. In CVPR, 2022.
  15. 15.Ma, B., Han, Z., Liu, Y., and Zwicker, M. Neural-pull: Learning signed distance function from point clouds by learning to pull space onto surface. In ICML, 2021.
  16. 16.Ma, B., Liu, Y., Zwicker, M., and Han, Z. Surface reconstruction from point clouds by learning predictive context priors. In CVPR, 2022.
  17. 17.Pang, Y., Wang, W., Tay, F. E. H., Liu, W., Tian, Y., and Yuan, L. Masked autoencoders for point cloud self-supervised learning. In ECCV, 2022.
  18. 18.Qi, C. R., Su, H., Mo, K., and Guibas, L. J. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017.
  19. 19.Roth, K., Pemula, L., Zepeda, J., Scholkopf, B., Brox, T., and Gehler, P. V. Towards total recall in industrial anomaly detection. In CVPR, 2022.
  20. 20.Rudolph, M., Wehrbein, T., Rosenhahn, B., and Wandt, B. Asymmetric student-teacher networks for industrial anomaly detection. In WACV, 2023.
  21. 21.Schluter, H. M., Tan, J., Hou, B., and Kainz, B. Natural synthetic anomalies for self-supervised anomaly detection and localization. In ECCV, 2022.
  22. 22.Shi, H., Zhou, Y., Yang, K., Yin, X., and Wang, K. Csflow: Learning optical flow via cross strip correlation for autonomous driving. In IEEE Intell. Vehicles Symposium, 2022.
  23. 23.Takikawa, T., Litalien, J., Yin, K., Kreis, K., Loop, C. T., Nowrouzezahrai, D., Jacobson, A., McGuire, M., and Fidler, S. Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In CVPR, 2021.
  24. 24.Viana, J. S., de la Rosa, E., Vyvere, T. V., Robben, D., and Sima, D. M. Unsupervised 3d brain anomaly detection. In Crimi, A. and Bakas, S. (eds.), Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries - 6th International Workshop with MICCAI, 2020.
  25. 25.Yang, M., Wu, P., Liu, J., and Feng, H. Memseg: A semi-supervised method for image surface defect detection using differences and commonalities. CoRR, 2022.
  26. 26.Zagoruyko, S. and Komodakis, N. Wide residual networks. In Richard C. Wilson, E. R. H. and Smith, W. A. P. (eds.), Proceedings of the British Machine Vision Conference (BMVC), 2016.
  27. 27.Zheng, Y., Wang, X., Qi, Y., Li, W., and Wu, L. Benchmarking unsupervised anomaly detection and localization. CoRR, 2022.

Citation

MLA
Chu, Y.-M., et al. “Shape-Guided Dual-Memory Learning for 3D Anomaly Detection”. International Conference on Machine Learning, vol. 202, 2023, pp. 6185–94, https://proceedings.mlr.press/v202/chu23b.html.
APA
Chu, Y.-M., Liu, C., Hsieh, T.-I., Chen, H.-T., & Liu, T.-L. (2023). Shape-Guided Dual-Memory Learning for 3D Anomaly Detection. International Conference on Machine Learning, 202, 6185–6194. https://proceedings.mlr.press/v202/chu23b.html
Chicago
Chu, Y.-M., C. Liu, T.-I. Hsieh, H.-T. Chen, and T.-L. Liu. 2023. “Shape-Guided Dual-Memory Learning for 3D Anomaly Detection”. International Conference on Machine Learning 202: 6185–94. https://proceedings.mlr.press/v202/chu23b.html.
Harvard
Chu, Y.-M. et al. (2023) “Shape-Guided Dual-Memory Learning for 3D Anomaly Detection”, International Conference on Machine Learning. PMLR, pp. 6185–6194. Available at: https://proceedings.mlr.press/v202/chu23b.html.
Vancouver
1. Chu Y-M, Liu C, Hsieh T-I, Chen H-T, Liu T-L (2023) Shape-Guided Dual-Memory Learning for 3D Anomaly Detection. In: International Conference on Machine Learning. PMLR, pp 6185–6194

BibTeX

@InProceedings{pmlr-v202-chu23b,
  title = 	 {Shape-Guided Dual-Memory Learning for 3{D} Anomaly Detection},
  author =       {Chu, Yu-Min and Liu, Chieh and Hsieh, Ting-I and Chen, Hwann-Tzong and Liu, Tyng-Luh},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {6185--6194},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/chu23b/chu23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/chu23b.html},
  abstract = 	 {We present a shape-guided expert-learning framework to tackle the problem of unsupervised 3D anomaly detection. Our method is established on the effectiveness of two specialized expert models and their synergy to localize anomalous regions from color and shape modalities. The first expert utilizes geometric information to probe 3D structural anomalies by modeling the implicit distance fields around local shapes. The second expert considers the 2D RGB features associated with the first expert to identify color appearance irregularities on the local shapes. We use the two experts to build the dual memory banks from the anomaly-free training samples and perform shape-guided inference to pinpoint the defects in the testing samples. Owing to the per-point 3D representation and the effective fusion scheme of complementary modalities, our method efficiently achieves state-of-the-art performance on the MVTec 3D-AD dataset with better recall and lower false positive rates, as preferred in real applications.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/