Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection

Dong GongLingqiao LiuVuong LeBudhaditya SahaMoussa Reda MansourSvetha VenkateshAnton van den Hengel

article2019ICCV1,793 citations

Develops a memory-augmented autoencoder that prevents unsupervised anomaly detectors from mistakenly reconstructing abnormal inputs by constraining latent representations to retrieve learned prototypes of normal data.

Listen

Unsupervised anomaly detection is critical for high-stakes operational domains such as video surveillance and cybersecurity, where systems must flag unusual events without prior examples of every possible threat. A common approach uses deep autoencoders—neural networks trained solely on normal data—under the premise that abnormal inputs will be reconstructed poorly, thereby exposing the anomaly. In practice, however, standard autoencoders often generalize too well and accurately reconstruct abnormal inputs, causing critical security and operational failures to go undetected.

The article introduces and evaluates a memory-augmented autoencoder architecture designed to prevent abnormal inputs from being reconstructed accurately. The objective was to demonstrate that constraining data reconstruction through a learned memory of prototypical normal patterns significantly enhances anomaly detection performance across diverse data types.

The proposed framework introduces an external memory module between the network's encoder and decoder. Instead of passing an input's encoded features directly to the decoder, the system uses the encoding to query a fixed memory bank. It retrieves only a sparse combination of recorded normal patterns using a specialized thresholding and entropy-regularized addressing mechanism. The authors evaluated the approach across three distinct domains using five standard benchmark datasets: image datasets (MNIST and CIFAR-10), multi-scene video surveillance datasets (UCSD-Ped2, CUHK Avenue, and ShanghaiTech), and a cybersecurity network intrusion dataset (KDDCUP99).

The evaluation yielded several key findings. First, the memory-augmented architecture consistently outperformed standard autoencoders and traditional baseline detectors across all benchmarks. On the complex ShanghaiTech surveillance dataset, it achieved an Area Under the Curve (AUC) of 0.712 compared to 0.609 for standard 2D convolutional autoencoders. Second, on the KDDCUP99 cybersecurity benchmark, the model attained an F1 score of 0.9641, surpassing both conventional autoencoders (0.9342) and deep clustering baselines (0.7762). Third, ablation studies confirmed that enforcing sparsity in memory retrieval is critical; removing sparsity operations consistently degraded anomaly detection accuracy. Finally, the added memory module introduces negligible computational overhead, processing video at approximately 38 frames per second (0.0262 seconds per frame), making it viable for real-time applications.

These findings indicate that memory augmentation successfully addresses the over-generalization weakness of deep autoencoders without compromising throughput. For organizations deploying surveillance, quality assurance, or intrusion detection systems, this architecture reduces the risk of missed detections while avoiding the prohibitive engineering costs of specialized, domain-specific models.

Decision-makers should consider adopting this memory-augmented framework to strengthen existing unsupervised anomaly detection pipelines, particularly in real-time surveillance and network monitoring environments. Future efforts should explore integrating the memory mechanism into more specialized, multi-modal network architectures and utilizing memory addressing weights directly as an additional indicator of anomalies.

While the model demonstrates high performance and robustness to varying memory capacities, confidence in complex scenarios—such as high-variance visual environments like CIFAR-10—must be tempered, as performance naturally drops when normal classes have high internal diversity. Organizations should run pilot validations on their specific domain data to calibrate memory size and sparsity parameters before full operational deployment.

  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It establishes the differentiable, attention-based external memory reading mechanism that MemAE directly adapts to restrict latent reconstruction to normal patterns.
  • Paper: Memory Networks, Jason Weston et al. (2014). It introduces the fundamental framework of augmenting neural networks with explicit, read-write memory banks for retrieval and inference.
  • Paper: Deep Autoencoding Gaussian Mixture Model for Unsupervised Anomaly Detection, Bo Zong et al. (2018). It provides the foundational deep autoencoder baseline and highlights the challenges of latent representation modeling for unsupervised anomaly detection.
  • Paper: Deep One-Class Classification, Lukas Ruff et al. (2018). It introduces deep one-class classification, framing the broader objective of tightly characterizing normality that MemAE addresses via memory augmentation.
  • Paper: Real-World Anomaly Detection in Surveillance Videos, Waqas Sultani et al. (2018). It establishes real-world surveillance video anomaly detection benchmarks and highlights the limitations of standard autoencoder reconstruction baselines.
  • Paper: Extracting and composing robust features with denoising autoencoders, Pascal Vincent et al. (2008). It lays the groundwork for training autoencoders to reconstruct uncorrupted features, which underpins the reconstruction error criterion used in MemAE.
  • Paper: Towards Total Recall in Industrial Anomaly Detection, Karsten Roth et al. (2021). It advances memory-based anomaly detection by using locally aware patch-level feature banks with greedy subset selection for high-precision industrial inspection.
  • Paper: Energy-based Out-of-distribution Detection, Weitang Liu et al. (2020). It builds upon unsupervised and out-of-distribution detection formulations by introducing an energy-based scoring framework that avoids overconfident anomaly reconstruction.
Cover for Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection

Abstract

Deep autoencoder has been extensively used for anomaly detection. Training on the normal data, the autoencoder is expected to produce higher reconstruction error for the abnormal inputs than the normal ones, which is adopted as a criterion for identifying anomalies. However, this assumption does not always hold in practice. It has been observed that sometimes the autoencoder "generalizes" so well that it can also reconstruct anomalies well, leading to the miss detection of anomalies. To mitigate this drawback for autoencoder based anomaly detector, we propose to augment the autoencoder with a memory module and develop an improved autoencoder called memory-augmented autoencoder, i.e. MemAE. Given an input, MemAE firstly obtains the encoding from the encoder and then uses it as a query to retrieve the most relevant memory items for reconstruction. At the training stage, the memory contents are updated and are encouraged to represent the prototypical elements of the normal data. At the test stage, the learned memory will be fixed, and the reconstruction is obtained from a few selected memory records of the normal data. The reconstruction will thus tend to be close to a normal sample. Thus the reconstructed errors on anomalies will be strengthened for anomaly detection. MemAE is free of assumptions on the data type and thus general to be applied to different tasks. Experiments on various datasets prove the excellent generalization and high effectiveness of the proposed MemAE.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Memory-augmented Autoencoder
  • 3.1 Overview
  • 3.2 Encoder and Decoder
  • 3.3 Memory Module with Attention-based Sparse Addressing
  • 3.3.1 Memory-based Representation
  • 3.3.2 Attention for Memory Addressing
  • 3.3.3 Hard Shrinkage for Sparse Addressing
  • 3.4 Training
  • 4 Experiments
  • 4.1 Experiments on Image Data
  • 4.1.1 Visualizing How the Memory Works
  • 4.2 Experiments on Video Anomaly Detection
  • 4.3 Experiments on Cybersecurity Data
  • 4.4 Ablation Studies
  • 4.4.1 Study of the Sparsity-inducing Components
  • 4.4.2 Comparison with AE with Sparse Regularization
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Memory-Augmented Autoencoder (MemAE) Architecture

    model/method

    The Memory-Augmented Autoencoder (MemAE) is an unsupervised anomaly detection framework designed to prevent deep autoencoders from generalizing to and accurately reconstructing anomalous inputs.

    Given an input sample x∈Xx \in \mathcal{X}, an encoder fe(⋅;θe)f_e(\cdot; \theta_e) maps the input to a latent query representation z∈Z=RCz \in \mathcal{Z} = \mathbb{R}^C:

    z=fe(x;θe)z = f_e(x; \theta_e)

    Instead of directly passing zz to the decoder, zz queries an external memory module containing a matrix M∈RN×CM \in \mathbb{R}^{N \times C}, where NN is the number of memory slots and each row mi∈RCm_i \in \mathbb{R}^C for i∈{1,…,N}i \in \{1, \dots, N\} represents a prototypical pattern learned from normal training data.

    The memory module produces an addressed representation z^∈RC\hat{z} \in \mathbb{R}^C via a sparse linear combination of memory slots:

    z^=w^M=∑i=1Nw^imi\hat{z} = \hat{w} M = \sum_{i=1}^N \hat{w}_i m_i

    where w^∈R1×N\hat{w} \in \mathbb{R}^{1 \times N} is a sparse addressing weight vector satisfying ∑i=1Nw^i=1\sum_{i=1}^N \hat{w}_i = 1 and w^i≥0\hat{w}_i \ge 0.

    A decoder fd(⋅;θd)f_d(\cdot; \theta_d) maps z^\hat{z} back to the input domain to produce a reconstruction x^∈X\hat{x} \in \mathcal{X}:

    x^=fd(z^;θd)\hat{x} = f_d(\hat{z}; \theta_d)

    During testing, MM is fixed. Normal inputs query memory slots that reconstruct the input accurately, yielding a small reconstruction error ∥x−x^∥22\|x - \hat{x}\|_2^2. Abnormal inputs cannot retrieve anomaly-specific patterns from MM and are reconstructed using normal prototypes, producing a large reconstruction error that marks the sample as an anomaly.

  2. Knowl 2 — Attention-Based Memory Addressing via Cosine Similarity

    equation

    In the Memory-Augmented Autoencoder (MemAE), a latent query vector z∈RCz \in \mathbb{R}^C generated by the encoder retrieves patterns from a memory matrix M∈RN×CM \in \mathbb{R}^{N \times C} with NN memory slots mi∈RCm_i \in \mathbb{R}^C (i∈{1,…,N}i \in \{1, \dots, N\}). The soft attention weight wi∈[0,1]w_i \in [0, 1] allocated to each memory slot mim_i is computed using a softmax operation over cosine similarities between zz and each memory item:

    wi=exp⁡(d(z,mi))∑j=1Nexp⁡(d(z,mj))w_i = \frac{\exp(d(z, m_i))}{\sum_{j=1}^N \exp(d(z, m_j))}

    where the similarity function d(z,mi)d(z, m_i) is defined as:

    d(z,mi)=zmiT∥z∥2∥mi∥2d(z, m_i) = \frac{z m_i^T}{\|z\|_2 \|m_i\|_2}

    The resulting attention weight vector w=[w1,w2,…,wN]∈R1×Nw = [w_1, w_2, \dots, w_N] \in \mathbb{R}^{1 \times N} is non-negative and satisfies ∑i=1Nwi=1\sum_{i=1}^N w_i = 1.

  3. Knowl 3 — Differentiable Hard Shrinkage Operator for Sparse Addressing

    model/method

    To prevent anomalous latent queries from reconstructing abnormal patterns through dense linear combinations of multiple memory slots, MemAE applies a continuous, differentiable hard shrinkage operator to each attention weight wiw_i:

    w^i=max⁡(wi−λ,0)⋅wi∣wi−λ∣+ϵ\hat{w}_i = \frac{\max(w_i - \lambda, 0) \cdot w_i}{|w_i - \lambda| + \epsilon}

    where λ\lambda is a shrinkage threshold (typically selected from the interval [1/N,3/N][1/N, 3/N] for memory size NN), max⁡(⋅,0)\max(\cdot, 0) is the standard ReLU activation function, and ϵ>0\epsilon > 0 is a small positive scalar ensuring numerical stability.

    Following shrinkage, the addressing weight vector w^\hat{w} is re-normalized:

    w^i←w^i∥w^∥1=w^i∑j=1Nw^j\hat{w}_i \leftarrow \frac{\hat{w}_i}{\|\hat{w}\|_1} = \frac{\hat{w}_i}{\sum_{j=1}^N \hat{w}_j}

    The resulting representation delivered to the decoder is z^=w^M=∑i=1Nw^imi\hat{z} = \hat{w} M = \sum_{i=1}^N \hat{w}_i m_i. During backpropagation, non-zero gradient updates occur exclusively for memory items whose addressing weights satisfy w^i>0\hat{w}_i > 0.

  4. Knowl 4 — Joint Reconstruction and Memory Entropy Training Objective

    equation

    The encoder parameters θe\theta_e, decoder parameters θd\theta_d, and memory matrix M∈RN×CM \in \mathbb{R}^{N \times C} of the Memory-Augmented Autoencoder are trained jointly on a dataset of TT normal training samples {xt}t=1T\{x^t\}_{t=1}^T by minimizing the objective function:

    L(θe,θd,M)=1T∑t=1T(R(xt,x^t)+αE(w^t))\mathcal{L}(\theta_e, \theta_d, M) = \frac{1}{T} \sum_{t=1}^T \left( R(x^t, \hat{x}^t) + \alpha E(\hat{w}^t) \right)

    The reconstruction term R(xt,x^t)R(x^t, \hat{x}^t) is the squared ℓ2\ell_2-norm between the training input xtx^t and its reconstruction x^t\hat{x}^t:

    R(xt,x^t)=∥xt−x^t∥22R(x^t, \hat{x}^t) = \|x^t - \hat{x}^t\|_2^2

    The sparsity regularization term E(w^t)E(\hat{w}^t) is the entropy of the normalized addressing weights w^t=[w^1t,…,w^Nt]\hat{w}^t = [\hat{w}_1^t, \dots, \hat{w}_N^t]:

    E(w^t)=∑i=1N−w^itlog⁡(w^it)E(\hat{w}^t) = \sum_{i=1}^N -\hat{w}_i^t \log(\hat{w}_i^t)

    The hyperparameter α\alpha balances reconstruction fidelity against addressing sparsity (set to α=0.0002\alpha = 0.0002 across experiments). Optimization is performed end-to-end using the Adam optimizer with backpropagation.

  5. Knowl 5 — Video Frame Normality Score via Spatio-Temporal Cuboid Reconstruction

    equation

    For video anomaly detection using a 3D convolutional Memory-Augmented Autoencoder (MemAE), the model takes as input a spatio-temporal cuboid formed by stacking 16 consecutive grayscale video frames. Memory slots in M∈RN×CM \in \mathbb{R}^{N \times C} record spatial-temporal feature vectors corresponding to individual pixel locations in the feature maps.

    During inference, the reconstruction error eue_u of the uu-th frame cuboid is computed via squared ℓ2\ell_2-norm eu=∥xu−x^u∥22e_u = \|x_u - \hat{x}_u\|_2^2. Across all frames of a video episode, the normality score pu∈[0,1]p_u \in [0, 1] for frame uu is defined as:

    pu=1−eu−min⁡u′(eu′)max⁡u′(eu′)−min⁡u′(eu′)p_u = 1 - \frac{e_u - \min_{u'} (e_{u'})}{\max_{u'} (e_{u'}) - \min_{u'} (e_{u'})}

    where min⁡u′(eu′)\min_{u'} (e_{u'}) and max⁡u′(eu′)\max_{u'} (e_{u'}) are the minimum and maximum cuboid reconstruction errors in that video episode. A score pup_u near 1 reflects normal activity, whereas a drop toward 0 denotes an anomalous event.

  6. Knowl 6 — Image Outlier Detection Performance on MNIST and CIFAR-10

    data/table

    Image anomaly detection is evaluated on MNIST and CIFAR-10 in a one-class classification setup. Each of the 10 classes is iteratively treated as the normal class, with anomalies sampled from the other 9 classes to constitute approximately 30% of the test set. Models are trained exclusively on normal data. Evaluation reports the average Area Under the ROC Curve (AUC) across all 10 one-class datasets.

    Method MNIST CIFAR-10
    OC-SVM 0.9499 0.5619
    KDE 0.8116 0.5756
    VAE 0.9643 0.5725
    PixCNN 0.6141 0.5450
    DSEBM 0.9554 0.5725
    AE 0.9619 0.5706
    MemAE-nonSpar 0.9725 0.6058
    MemAE 0.9751 0.6088

    MemAE achieves the highest average AUC on both MNIST (0.9751) and CIFAR-10 (0.6088). Introducing memory addressing without sparsity constraints (MemAE-nonSpar) yields 0.9725 and 0.6058, while a standard autoencoder without memory (AE) achieves 0.9619 and 0.5706, confirming that the memory module and sparse addressing jointly improve anomaly discrimination.

  7. Knowl 7 — Video Anomaly Detection Performance on Benchmark Datasets

    data/table

    Video anomaly detection is evaluated using frame-level Area Under the ROC Curve (AUC) on three benchmark surveillance datasets: UCSD-Ped2, CUHK Avenue, and ShanghaiTech.

    Category Method UCSD-Ped2 CUHK Avenue ShanghaiTech
    Non-Recon. MPPCA 0.693 – –
    MPPCA+SFA 0.613 – –
    MDT 0.829 – –
    AMDN 0.908 – –
    Unmasking 0.822 0.806 –
    MT-FRCN 0.922 – –
    Frame-Pred 0.954 0.849 0.728
    Recon. AE-Conv2D 0.850 0.800 0.609
    AE-Conv3D 0.912 0.771 –
    TSC 0.910 0.806 0.679
    StackRNN 0.922 0.817 0.680
    AE 0.917 0.810 0.697
    MemAE-nonSpar 0.929 0.821 0.688
    MemAE 0.941 0.833 0.712

    Among reconstruction-based approaches, MemAE achieves the highest AUC across all three video datasets (0.941 on UCSD-Ped2, 0.833 on CUHK Avenue, and 0.712 on ShanghaiTech). MemAE outperforms the standard 3D convolutional autoencoder (AE) by margins of 0.024, 0.023, and 0.015 respectively, as well as sparse representation baselines (TSC and StackRNN).

  8. Knowl 8 — Cybersecurity Anomaly Detection Performance on KDDCUP99

    data/table

    Tabular anomaly detection is evaluated on the 120-dimensional KDDCUP99 10 percent dataset from the UCI repository. In accordance with standard protocol, 80% of samples labeled as 'attack' are treated as normal, and only normal samples are used during training (50% train/test split). Results reflect the mean Precision, Recall, and F1F_1 score over 20 runs.

    Method Precision Recall F1F_1
    OC-SVM 0.7457 0.8523 0.7954
    DCN 0.7696 0.7829 0.7762
    DSEBM 0.8619 0.6446 0.7399
    DAGMM 0.9297 0.9442 0.9369
    AE 0.9328 0.9356 0.9342
    MemAE-nonSpar 0.9341 0.9368 0.9355
    MemAE 0.9627 0.9655 0.9641

    MemAE achieves the highest overall accuracy with a Precision of 0.9627, Recall of 0.9655, and F1F_1 score of 0.9641, outperforming DAGMM (F1=0.9369F_1 = 0.9369) and standard autoencoders (F1=0.9342F_1 = 0.9342).

  9. Knowl 9 — Ablation Analysis of MemAE Sparsity Components and Memory Capacity

    empirical result

    An ablation study on the UCSD-Ped2 dataset assesses the impact of the hard shrinkage operator, entropy loss, and compares MemAE against an autoencoder with ℓ1\ell_1-norm regularization on encoder activations (AE-ℓ1\text{AE-}\ell_1, which minimizes ∥z∥1\|z\|_1 during training).

    Method AUC
    AE 0.9170
    AE-ℓ1\ell_1 0.9286
    MemAE-nonSpar 0.9293
    MemAE w/o Shrinkage 0.9324
    MemAE w/o Entropy loss 0.9372
    MemAE (Full) 0.9410

    Key empirical observations include:

    • Eliminating hard shrinkage drops AUC from 0.9410 to 0.9324 due to noisy, dense addressing combinations reconstructing anomalous inputs.
    • Eliminating entropy loss reduces AUC to 0.9372, demonstrating its importance in driving confident slot selection during early training.
    • Applying latent ℓ1\ell_1 sparsity directly to the autoencoder (AE-ℓ1\text{AE-}\ell_1, AUC 0.9286) improves upon standard AE (0.9170) but lacks the prototypical memory representation of MemAE.
    • Robustness analysis across memory capacities N∈[500,3000]N \in [500, 3000] shows AUC increases rapidly from N=500N=500 (≈0.90\approx 0.90) to N=1000N=1000 (≈0.937\approx 0.937) and remains robustly stable above 0.94 for N≥1500N \ge 1500.
    • Runtime testing on an NVIDIA GeForce 1080 Ti shows MemAE averages 0.0262 s per frame (38 fps), adding only 4×10−44 \times 10^{-4} s per frame of computation relative to the standard AE baseline (0.0266 s).

Coverage note — Qualitative visual illustrations of single decoded memory slots and image/video frame error maps were omitted in favor of quantitative tabular results and explicit architectural/mathematical formulations.

References

  1. 1.D. Abati, A. Porrello, S. Calderara, and R. Cucchiara. AND: Autoregressive novelty detectors. arXiv preprint arXiv:1807.01653, 2018.
  2. 2.Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle. Greedy layer-wise training of deep networks. In Advances in Neural Information Processing Systems, pages 153–160, 2007.
  3. 3.R. Chalapathy, A. K. Menon, and S. Chawla. Anomaly detection using one-class neural networks. arXiv preprint arXiv:1802.06360, 2018.
  4. 4.V. Chandola, A. Banerjee, and V. Kumar. Anomaly detection: A survey. ACM computing surveys (CSUR), 41(3):15, 2009.
  5. 5.Y. Chen, X. S. Zhou, and T. S. Huang. One-class svm for learning in image retrieval. In The IEEE International Conference on Image Processing, volume 1, pages 34–37. IEEE, 2001.
  6. 6.Y. S. Chong and Y. H. Tay. Abnormal event detection in videos using spatiotemporal autoencoder. In International Symposium on Neural Networks, pages 189–196. Springer, 2017.
  7. 7.I. Golan and R. El-Yaniv. Deep anomaly detection using geometric transformations. In Advances in Neural Information Processing Systems, pages 9758–9769, 2018.
  8. 8.A. Graves, G. Wayne, and I. Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014.
  9. 9.M. Hasan, J. Choi, J. Neumann, A. K. Roy-Chowdhury, and L. S. Davis. Learning temporal regularity in video sequences. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 733–742, 2016.
  10. 10.R. Hinami, T. Mei, and S. Satoh. Joint detection and recounting of abnormal events by learning deep generic knowledge. In The IEEE International Conference on Computer Vision (ICCV), pages 3639–3647, 2017.
  11. 11.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  12. 12.I. Jolliffe. Principal component analysis. In International encyclopedia of statistical science, pages 1094–1096. Springer, 2011.
  13. 13.J. Kim and K. Grauman. Observe locally, infer globally: a space-time mrf for detecting abnormal activities with incremental updates. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  14. 14.Y. Kim, M. Kim, and G. Kim. Memorization precedes generation: Learning unsupervised gans with memory networks. International Conference on Learning Representations (ICLR), 2018.
  15. 15.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  16. 16.D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  17. 17.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009.
  18. 18.Y. LeCun. The mnist database of handwritten digits. http://yann. lecun. com/exdb/mnist/, 1998.
  19. 19.R. Leyva, V. Sanchez, and C.-T. Li. The lv dataset: A realistic surveillance video dataset for abnormal event detection. In International Workshop on Biometrics and Forensics (IWBF), pages 1–6. IEEE, 2017.
  20. 20.C. Li, J. Zhu, and B. Zhang. Learning to generate with memory. In International Conference on Machine Learning (ICML), pages 1177–1186, 2016.
  21. 21.M. Lichman et al. Uci machine learning repository, 2013.
  22. 22.W. Liu, W. Luo, D. Lian, and S. Gao. Future frame prediction for anomaly detection–a new baseline. The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  23. 23.C. Lu, J. Shi, and J. Jia. Abnormal event detection at 150 fps in matlab. In The IEEE International Conference on Computer Vision (ICCV), pages 2720–2727, 2013.
  24. 24.W. Luo, W. Liu, and S. Gao. A revisit of sparse coding based anomaly detection in stacked rnn framework. The IEEE International Conference on Computer Vision (ICCV), 2017.
  25. 25.V. Mahadevan, W. Li, V. Bhalodia, and N. Vasconcelos. Anomaly detection in crowded scenes. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  26. 26.V. Nair and G. E. Hinton. Rectified linear units improve restricted boltzmann machines. In International Conference on Machine Learning (ICML), pages 807–814, 2010.
  27. 27.E. Parzen. On estimation of a probability density function and mode. The Annals of Mathematical Statistics, 33(3):1065–1076, 1962.
  28. 28.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017.
  29. 29.J. Rae, J. J. Hunt, I. Danihelka, T. Harley, A. W. Senior, G. Wayne, A. Graves, and T. Lillicrap. Scaling memoryaugmented neural networks with sparse reads and writes. In Advances in Neural Information Processing Systems, pages 3621–3629, 2016.
  30. 30.L. Ruff, R. Vandermeulen, N. Goernitz, L. Deecke, S. A. Siddiqui, A. Binder, E. Muller, and M. Kloft. Deep one-class classification. In International Conference on Machine Learning (ICML), pages 4393–4402, 2018.
  31. 31.M. Sabokrou, M. Khalooei, M. Fathy, and E. Adeli. Adversarially learned one-class classifier for novelty detection. In The IEEE Conference on Computer Vision and Pattern Recognition, pages 3379–3388, 2018.
  32. 32.A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap. One-shot learning with memory-augmented neural networks. arXiv preprint arXiv:1605.06065, 2016.
  33. 33.B. Scholkopf and A. J. Smola. Learning with kernels: support vector machines, regularization, optimization, and beyond. MIT press, 2001.
  34. 34.B. Schölkopf, R. C. Williamson, A. J. Smola, J. Shawe-Taylor, and J. C. Platt. Support vector method for novelty detection. In Advances in Neural Information Processing Systems, pages 582–588, 2000.
  35. 35.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional networks. In The IEEE International Conference on Computer Vision (ICCV), pages 4489–4497, 2015.
  36. 36.R. Tudor Ionescu, S. Smeureanu, B. Alexe, and M. Popescu. Unmasking the abnormal events in video. In The IEEE International Conference on Computer Vision (ICCV), 2017.
  37. 37.A. van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al. Conditional image generation with pixelcnn decoders. In Advances in Neural Information Processing Systems, pages 4790–4798, 2016.
  38. 38.J. Weston, S. Chopra, and A. Bordes. Memory networks. International Conference on Learning Representations (ICLR), 2015.
  39. 39.L. Xiong, B. Póczos, and J. G. Schneider. Group anomaly detection using flexible genre models. In Advances in Neural Information Processing Systems, pages 1071–1079, 2011.
  40. 40.D. Xu, E. Ricci, Y. Yan, J. Song, and N. Sebe. Learning deep representations of appearance and motion for anomalous event detection. arXiv preprint arXiv:1510.01553, 2015.
  41. 41.X. Yang, K. Huang, J. Y. Goulermas, and R. Zhang. Joint learning of unsupervised dimensionality reduction and gaussian mixture model. Neural Processing Letters, 45(3):791–806, 2017.
  42. 42.S. Zhai, Y. Cheng, W. Lu, and Z. Zhang. Deep structured energy based models for anomaly detection. arXiv preprint arXiv:1605.07717, 2016.
  43. 43.B. Zhao, L. Fei-Fei, and E. P. Xing. Online detection of unusual events in videos via dynamic sparse coding. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3313–3320, 2011.
  44. 44.Y. Zhao, B. Deng, C. Shen, Y. Liu, H. Lu, and X.-S. Hua. Spatio-temporal autoencoder for video anomaly detection. In ACM on Multimedia Conference, pages 1933–1941. ACM, 2017.
  45. 45.C. Zhou and R. C. Paffenroth. Anomaly detection with robust deep autoencoders. In ACM SIGKDD, pages 665–674. ACM, 2017.
  46. 46.A. Zimek, E. Schubert, and H.-P. Kriegel. A survey on unsupervised outlier detection in high-dimensional numerical data. Statistical Analysis and Data Mining: The ASA Data Science Journal, 5(5):363–387, 2012.
  47. 47.B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. Cho, and H. Chen. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In International Conference on Learning Representations, 2018.

Citation

MLA
Gong, D., et al. “Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection”. arXiv, 2019, http://arxiv.org/abs/1904.02639v2.
APA
Gong, D., Liu, L., Le, V., Saha, B., Mansour, M. R., Venkatesh, S., & Hengel, A. van . den . (2019). Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection. arXiv. http://arxiv.org/abs/1904.02639v2
Chicago
Gong, D., L. Liu, V. Le, et al. 2019. “Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection”. arXiv. http://arxiv.org/abs/1904.02639v2.
Harvard
Gong, D. et al. (2019) “Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.02639v2.
Vancouver
1. Gong D, Liu L, Le V, Saha B, Mansour MR, Venkatesh S, Hengel A van den (2019) Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection. arXiv

BibTeX

@article{gong2019memorizing,
  title = {Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection},
  author = {Gong, Dong and Liu, Lingqiao and Le, Vuong and Saha, Budhaditya and Mansour, Moussa Reda and Venkatesh, Svetha and Hengel, Anton van den},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.02639v2},
  eprint = {1904.02639}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE