Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection
Dong GongLingqiao LiuVuong LeBudhaditya SahaMoussa Reda MansourSvetha VenkateshAnton van den Hengel
Develops a memory-augmented autoencoder that prevents unsupervised anomaly detectors from mistakenly reconstructing abnormal inputs by constraining latent representations to retrieve learned prototypes of normal data.
Unsupervised anomaly detection is critical for high-stakes operational domains such as video surveillance and cybersecurity, where systems must flag unusual events without prior examples of every possible threat. A common approach uses deep autoencoders—neural networks trained solely on normal data—under the premise that abnormal inputs will be reconstructed poorly, thereby exposing the anomaly. In practice, however, standard autoencoders often generalize too well and accurately reconstruct abnormal inputs, causing critical security and operational failures to go undetected.
The article introduces and evaluates a memory-augmented autoencoder architecture designed to prevent abnormal inputs from being reconstructed accurately. The objective was to demonstrate that constraining data reconstruction through a learned memory of prototypical normal patterns significantly enhances anomaly detection performance across diverse data types.
The proposed framework introduces an external memory module between the network's encoder and decoder. Instead of passing an input's encoded features directly to the decoder, the system uses the encoding to query a fixed memory bank. It retrieves only a sparse combination of recorded normal patterns using a specialized thresholding and entropy-regularized addressing mechanism. The authors evaluated the approach across three distinct domains using five standard benchmark datasets: image datasets (MNIST and CIFAR-10), multi-scene video surveillance datasets (UCSD-Ped2, CUHK Avenue, and ShanghaiTech), and a cybersecurity network intrusion dataset (KDDCUP99).
The evaluation yielded several key findings. First, the memory-augmented architecture consistently outperformed standard autoencoders and traditional baseline detectors across all benchmarks. On the complex ShanghaiTech surveillance dataset, it achieved an Area Under the Curve (AUC) of 0.712 compared to 0.609 for standard 2D convolutional autoencoders. Second, on the KDDCUP99 cybersecurity benchmark, the model attained an F1 score of 0.9641, surpassing both conventional autoencoders (0.9342) and deep clustering baselines (0.7762). Third, ablation studies confirmed that enforcing sparsity in memory retrieval is critical; removing sparsity operations consistently degraded anomaly detection accuracy. Finally, the added memory module introduces negligible computational overhead, processing video at approximately 38 frames per second (0.0262 seconds per frame), making it viable for real-time applications.
These findings indicate that memory augmentation successfully addresses the over-generalization weakness of deep autoencoders without compromising throughput. For organizations deploying surveillance, quality assurance, or intrusion detection systems, this architecture reduces the risk of missed detections while avoiding the prohibitive engineering costs of specialized, domain-specific models.
Decision-makers should consider adopting this memory-augmented framework to strengthen existing unsupervised anomaly detection pipelines, particularly in real-time surveillance and network monitoring environments. Future efforts should explore integrating the memory mechanism into more specialized, multi-modal network architectures and utilizing memory addressing weights directly as an additional indicator of anomalies.
While the model demonstrates high performance and robustness to varying memory capacities, confidence in complex scenarios—such as high-variance visual environments like CIFAR-10—must be tempered, as performance naturally drops when normal classes have high internal diversity. Organizations should run pilot validations on their specific domain data to calibrate memory size and sparsity parameters before full operational deployment.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It establishes the differentiable, attention-based external memory reading mechanism that MemAE directly adapts to restrict latent reconstruction to normal patterns.
- Paper: Memory Networks, Jason Weston et al. (2014). It introduces the fundamental framework of augmenting neural networks with explicit, read-write memory banks for retrieval and inference.
- Paper: Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection, Bo Zong et al. (2018). It provides the foundational deep autoencoder baseline and highlights the challenges of latent representation modeling for unsupervised anomaly detection.
- Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). It introduces deep one-class classification, framing the broader objective of tightly characterizing normality that MemAE addresses via memory augmentation.
- Paper: Real-World Anomaly Detection in Surveillance Videos, Waqas Sultani et al. (2018). It establishes real-world surveillance video anomaly detection benchmarks and highlights the limitations of standard autoencoder reconstruction baselines.
- Paper: Extracting and composing robust features with denoising autoencoders, Pascal Vincent et al. (2008). It lays the groundwork for training autoencoders to reconstruct uncorrupted features, which underpins the reconstruction error criterion used in MemAE.
- Paper: Towards Total Recall in Industrial Anomaly Detection, Karsten Roth et al. (2021). It advances memory-based anomaly detection by using locally aware patch-level feature banks with greedy subset selection for high-precision industrial inspection.
- Paper: Energy-based Out-of-distribution Detection, Weitang Liu et al. (2020). It builds upon unsupervised and out-of-distribution detection formulations by introducing an energy-based scoring framework that avoids overconfident anomaly reconstruction.
