Anomaly Detection with Robust Deep Autoencoders

Chong ZhouR. Paffenroth

article2017KDD1,570 citations

Proposes a matrix-splitting framework combining deep autoencoders with sparse and group-sparse regularization to isolate complex noise and detect anomalies without requiring clean training data.

Listen

Modern data systems frequently handle large volumes of real-world data contaminated by substantial noise and rare anomalies. Deep autoencoders—neural networks designed to discover compact, non-linear representations—typically require pristine, noise-free training sets to learn effective features, or they risk memorizing corruptions. In operational settings, obtaining clean data is often expensive, impractical, or impossible, which compromises the reliability of downstream analytics, cybersecurity monitoring, and automated inspection.

The article demonstrates a novel framework called the Robust Deep Autoencoder, along with its group-structured extension, designed to remove noise and isolate anomalies directly from uncurated training data without requiring clean examples. The authors evaluate this architecture on standard image benchmarks to measure feature reconstruction accuracy and unsupervised outlier detection performance.

To achieve this, the approach splits input data into two components: a clean underlying representation reconstructed by the neural network and a sparse matrix containing noise or anomalies. The system alternates between training the neural network via standard optimization and isolating anomalies using mathematical shrinkage operations. Credibility is established through extensive experiments on a 50,000-sample benchmark dataset corrupted with varying degrees of synthetic noise and mixed anomaly classes, comparing downstream classification accuracy and outlier identification against industry-standard methods.

The analysis reveals several key findings. First, when handling moderate to heavy noise (80 to 300 corrupted pixels out of 784 per image), the Robust Deep Autoencoder reduced downstream classification error rates by up to 30% compared to standard deep autoencoders. Second, in unsupervised anomaly detection tasks with an anomaly prevalence of roughly 5.2%, the group-robust variant achieved an F1-score of 0.64, outperforming the baseline Isolation Forest algorithm's score of 0.37—representing an improvement of about 73%. Third, empirical tests showed that the alternating optimization scheme converged reliably within 200 iterations for practical parameter configurations.

These results indicate that organizations can train high-performing neural representations directly on noisy raw data, eliminating the cost and delay of manual data cleaning. Furthermore, the framework provides an effective tool for identifying structured anomalies, reducing operational risk in critical applications such as fraud detection and network intrusion monitoring. The approach represents a significant advancement over standard denoising autoencoders, which fail when clean training baselines are unavailable.

Organizations evaluating anomaly detection or feature extraction on noisy data should consider piloting this robust autoencoder framework. Practitioners must calibrate the sparsity penalty parameter, which governs the balance between false positives and false negatives based on operational risk tolerance. Before deploying the method to production systems, teams should conduct domain-specific testing, particularly in specialized environments such as cyber-attack detection.

Confidence in the reported improvements is high for image-based benchmark data under controlled noise conditions. However, decision-makers should note key limitations: the optimization objective is non-convex without theoretical convergence guarantees, and current experimental validation is restricted to standard benchmark data. Further empirical validation on real-world, high-dimensional operational datasets is recommended.

Cover for Anomaly Detection with Robust Deep Autoencoders

Abstract

Deep autoencoders, and other deep neural networks, have demonstrated their effectiveness in discovering non-linear features across many problem domains. However, in many real-world problems, large outliers and pervasive noise are commonplace, and one may not have access to clean training data as required by standard deep denoising autoencoders. Herein, we demonstrate novel extensions to deep autoencoders which not only maintain a deep autoencoders’ ability to discover high quality, non-linear features but can also eliminate outliers and noise without access to any clean training data. Our model is inspired by Robust Principal Component Analysis, and we split the input data X into two parts, X = L_D + S, where L_D can be effectively reconstructed by a deep autoencoder and S contains the outliers and noise in the original data X. Since such splitting increases the robustness of standard deep autoencoders, we name our model a “Robust Deep Autoencoder (RDA)”. Further, we present generalizations of our results to grouped sparsity norms which allow one to distinguish random anomalies from other types of structured corruptions, such as a collection of features being corrupted across many instances or a collection of instances having more corruptions than their fellows. Such “Group Robust Deep Autoencoders (GRDA)” give rise to novel anomaly detection approaches whose superior performance we demonstrate on a selection of benchmark problems.

Table of Contents

  • 1.1 Contribution
  • 1 INTRODUCTION
  • 2 BACKGROUND
  • 2.1 Deep Autoencoders
  • 2.2 Robust Principal Component Analysis
  • 3 METHODOLOGY
  • 3.1 Robust Deep Autoencoders with ' 1 Regularization
  • 3.2 Robust Deep Autoencoders with ' 2 ; 1 Regularization
  • 3.3 Anomalous Feature and Instance Detection
  • 4 ALGORITHM TRAINING
  • 4.1 Alternating Optimization for ' 1 and ' 2 ; 1 Robust Deep Autoencoder
  • 4.2 Proximal Method for ' 1 Norm
  • 4.3 Proximal Method for ' 2 ; 1 Norm
  • 5 NUMERICAL RESULTS
  • 5.1 Implementation Details
  • 5.2 ' 1 Robust Deep Autoencoder for Denoising
  • 5.3 ' 2 ; 1 Robust Deep Autoencoder for Outlier Detection
  • 5.4 Convergence of Training Method
  • 6 CONCLUSION
  • REFERENCES

Knowls

  1. Knowl 1 — Robust Deep Autoencoder Formulation with L1 Regularization

    model/method

    The Robust Deep Autoencoder (RDA) extends deep autoencoders to handle uncleaned, noisy training data without requiring noise-free ground truth. Given an input data matrix X∈RN×nX \in \mathbb{R}^{N \times n}, RDA decomposes XX into two additive components:

    X=LD+SX = L_D + S

    where LD∈RN×nL_D \in \mathbb{R}^{N \times n} represents the underlying low-dimensional non-linear manifold of the nominal data, and S∈RN×nS \in \mathbb{R}^{N \times n} contains isolated sparse noise and outliers.

    The ℓ1\ell_1-regularized RDA optimization problem is formulated as:

    min⁡θ,S∥LD−Dθ(Eθ(LD))∥2+λ∥S∥1subject toX−LD−S=0\min_{\theta, S} \|L_D - D_\theta(E_\theta(L_D))\|_2 + \lambda \|S\|_1 \quad \text{subject to} \quad X - L_D - S = 0

    where Eθ:Rn→RdE_\theta: \mathbb{R}^n \to \mathbb{R}^d is an encoder mapping the data to a hidden representation, Dθ:Rd→RnD_\theta: \mathbb{R}^d \to \mathbb{R}^n is a decoder mapping hidden representations back to data space, θ\theta denotes the network parameters (weights and biases), ∥S∥1=∑i,j∣Si,j∣\|S\|_1 = \sum_{i,j} |S_{i,j}| is the element-wise ℓ1\ell_1 norm that promotes entry-wise sparsity, and λ>0\lambda > 0 is a regularization parameter controlling the sparsity of SS. Smaller λ\lambda allows more entries into SS (reducing autoencoder reconstruction loss), whereas larger λ\lambda penalizes entries in SS, forcing more content into LDL_D.

  2. Knowl 2 — Group Robust Deep Autoencoder Formulation with L2,1 Regularization

    model/method

    To handle structured corruptions where entire features are corrupted across instances or entire instances are corrupted across features, the Group Robust Deep Autoencoder (GRDA) replaces the element-wise ℓ1\ell_1 penalty with a grouped ℓ2,1\ell_{2,1} norm. For a matrix M∈Rm×nM \in \mathbb{R}^{m \times n} with columns mjm_j, the ℓ2,1\ell_{2,1} norm is defined as:

    ∥M∥2,1=∑j=1n∥mj∥2=∑j=1n(∑i=1m∣Mi,j∣2)1/2\|M\|_{2,1} = \sum_{j=1}^n \|m_j\|_2 = \sum_{j=1}^n \left( \sum_{i=1}^m |M_{i,j}|^2 \right)^{1/2}

    GRDA defines two optimization variants depending on the anomaly structure:

    1. Anomalous Feature Detection (penalizing columns of SS, targeting specific features corrupted across instances): min⁡θ,S∥LD−Dθ(Eθ(LD))∥2+λ∥S∥2,1subject toX−LD−S=0\min_{\theta, S} \|L_D - D_\theta(E_\theta(L_D))\|_2 + \lambda \|S\|_{2,1} \quad \text{subject to} \quad X - L_D - S = 0

    2. Anomalous Instance / Outlier Detection (penalizing rows of SS via STS^T, targeting anomalous data instances): min⁡θ,S∥LD−Dθ(Eθ(LD))∥2+λ∥ST∥2,1subject toX−LD−S=0\min_{\theta, S} \|L_D - D_\theta(E_\theta(L_D))\|_2 + \lambda \|S^T\|_{2,1} \quad \text{subject to} \quad X - L_D - S = 0

  3. Knowl 3 — Alternating Optimization Algorithm for Robust Deep Autoencoders

    algorithm

    Robust Deep Autoencoders are trained using an alternating optimization scheme combining back-propagation for the neural network, proximal shrinkage operators for the sparsity penalty on SS, and projection steps onto the constraint manifold X=LD+SX = L_D + S.

    Input: Data matrix X∈RN×nX \in \mathbb{R}^{N \times n}, regularizer λ>0\lambda > 0, tolerance ϵ>0\epsilon > 0, norm type penalty∈{ℓ1,ℓ2,1}\text{penalty} \in \{\ell_1, \ell_{2,1}\}
    Output: Low-dimensional nominal representation LDL_D, anomaly matrix SS
    Initialize LD←0N×nL_D \leftarrow 0_{N \times n}, S←0N×nS \leftarrow 0_{N \times n}, LS←XLS \leftarrow X
    Initialize autoencoder encoder EθE_\theta and decoder DθD_\theta with random parameters
    while True:
        LD←X−SL_D \leftarrow X - S
        Minimize ∥LD−Dθ(Eθ(LD))∥2\|L_D - D_\theta(E_\theta(L_D))\|_2 with respect to θ\theta via back-propagation (e.g., batch size 30, 5 epochs)
        LD←Dθ(Eθ(LD))L_D \leftarrow D_\theta(E_\theta(L_D))
        S←X−LDS \leftarrow X - L_D
        if penalty is ℓ1\ell_1:
            S←proxλ,ℓ1(S)S \leftarrow \text{prox}_{\lambda, \ell_1}(S)
        else if penalty is ℓ2,1\ell_{2,1}:
            S←proxλ,ℓ2,1(S)S \leftarrow \text{prox}_{\lambda, \ell_{2,1}}(S)
        c1←∥X−LD−S∥2/∥X∥2c_1 \leftarrow \|X - L_D - S\|_2 / \|X\|_2
        c2←∥LS−LD−S∥2/∥X∥2c_2 \leftarrow \|LS - L_D - S\|_2 / \|X\|_2
        if c1<ϵc_1 < \epsilon or c2<ϵc_2 < \epsilon:
            break
        LS←LD+SLS \leftarrow L_D + S
    return LD,SL_D, S

    The algorithm alternates between updating the network parameters θ\theta while holding SS fixed, updating SS via the exact proximal operator while holding LDL_D fixed, and enforcing the additive decomposition constraints LD=X−SL_D = X - S and S=X−LDS = X - L_D.

  4. Knowl 4 — Proximal Shrinkage Operators for L1 and L2,1 Norm Minimization

    model/method

    In the Robust Deep Autoencoder optimization loop, the non-differentiable regularization terms on the anomaly matrix S∈Rm×nS \in \mathbb{R}^{m \times n} are minimized using exact proximal operators:

    1. Element-wise Soft-Thresholding for ℓ1\ell_1 Norm: For each entry Si,jS_{i,j}: [proxλ,ℓ1(S)]i,j={Si,j−λif Si,j>λSi,j+λif Si,j<−λ0if −λ≤Si,j≤λ\left[\text{prox}_{\lambda, \ell_1}(S)\right]_{i,j} = \begin{cases} S_{i,j} - \lambda & \text{if } S_{i,j} > \lambda \\ S_{i,j} + \lambda & \text{if } S_{i,j} < -\lambda \\ 0 & \text{if } -\lambda \le S_{i,j} \le \lambda \end{cases}

    2. Block-wise Soft-Thresholding for ℓ2,1\ell_{2,1} Norm: For each column j∈{1,…,n}j \in \{1, \dots, n\}, let ej=∥S⋅,j∥2=(∑i=1m∣Si,j∣2)1/2e_j = \|S_{\cdot, j}\|_2 = \left( \sum_{i=1}^m |S_{i,j}|^2 \right)^{1/2}. The operator acts on the column vector S⋅,jS_{\cdot, j}: [proxλ,ℓ2,1(S)]⋅,j={S⋅,j−λS⋅,jejif ej>λ0if ej≤λ\left[\text{prox}_{\lambda, \ell_{2,1}}(S)\right]_{\cdot, j} = \begin{cases} S_{\cdot, j} - \lambda \frac{S_{\cdot, j}}{e_j} & \text{if } e_j > \lambda \\ 0 & \text{if } e_j \le \lambda \end{cases} When applied to row grouping (STS^T), the operator is applied across rows instead of columns.

  5. Knowl 5 — Feature Denoising Performance of L1 Robust Deep Autoencoders on MNIST

    empirical result

    The feature learning quality of the ℓ1\ell_1-regularized Robust Deep Autoencoder (RDA) was evaluated on the MNIST dataset (50,000 images, 28×28=78428 \times 28 = 784 pixels) corrupted with synthetic noise without access to clean images.

    • Noise Protocol: Randomly selected pixels per image (from 10 up to 350 out of 784) were flipped (set to 0 if original >0.5> 0.5, else set to 1).
    • Network Architecture: Two hidden layers reducing dimension from 784 to 196, and 196 to 49, with sigmoid activations g(t)=1/(1+e−t)g(t) = 1 / (1 + e^{-t}).
    • Evaluation: The 49-dimensional hidden features were used to train a Random Forest classifier (50 trees) on a 2/3 training split and tested on a 1/3 test split.
    • Results: For moderate-to-high corruption levels (80 to 300 corrupted pixels per image), the standard deep autoencoder exhibited up to a 30% higher classification error rate compared to RDA. Under very low corruption (10 to 50 pixels) or extreme corruption (>300> 300 pixels), the classification errors of RDA and standard autoencoders converged to similar levels.
  6. Knowl 6 — Unsupervised Outlier Detection Performance of L2,1 Robust Deep Autoencoder

    empirical result

    The ℓ2,1\ell_{2,1}-regularized RDA was evaluated for unsupervised outlier detection on a contaminated MNIST subset without label supervision.

    • Data Setup: 4,859 nominal images of digit "4" contaminated with 265 randomly sampled images of other digits ("0", "7", "9", etc.), resulting in a 5.2% anomaly contamination ratio (5,124 total instances).
    • Network Architecture: Fully connected autoencoder with layer sizes 784 →\to 400 →\to 200.
    • Evaluation and Tuning: λ\lambda controls row sparsity in STS^T. Non-zero rows in SS indicate predicted outliers. Performance was evaluated via F1-score.
    • Comparison with Baseline: Isolation Forest (100 trees, outlier fraction swept from 0.01 to 0.69) achieved its highest F1-score of 0.37 at an outlier fraction of 0.11. The ℓ2,1\ell_{2,1} RDA achieved a peak F1-score of 0.64 (at λ=0.00065\lambda = 0.00065), yielding a 73.0% relative improvement in F1-score over Isolation Forest.
  7. Knowl 7 — Non-Convexity and Empirical Convergence of Robust Deep Autoencoder Optimization

    limitation

    Because the reconstruction error term ∥LD−Dθ(Eθ(LD))∥2\|L_D - D_\theta(E_\theta(L_D))\|_2 incorporates non-linear neural network activations, the overall RDA objective function is non-convex. Consequently, alternating ADMM-like optimization lacks theoretical guarantees of convergence to a global minimum.

    Empirical convergence dynamics show that the objective value drops rapidly within the first 50 iterations and reaches convergence within 200 iterations for small λ\lambda values, with each deterministic proximal step on SS accounting for the largest objective drops between neural network training epochs. However, theoretical convergence bounds and rates remain unproven.

Coverage note — Standard introductory background material on linear Robust Principal Component Analysis (RPCA) and classical PCA was omitted, as it represents prior work rather than the authors' novel contributions.

References

  1. 1.Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Je‚rey Dean, MaŠhieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geo‚rey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin WaŠenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. (mar 2016). arXiv:1603.04467 hŠp://arxiv.org/abs/1603.04467
  2. 2.Christopher M Bishop. 2006. PaŠern recognition. Machine Learning 128 (2006), 1–58.
  3. 3.Stephen Boyd. 2010. Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers slide (Alternating Direction Method of Multipliers ). (2010). hŠp://dl.acm.org/citation.cfm?id=2185816
  4. 4.S.P. Boyd and L. Vandenberghe. 2004. Convex optimization. Vol. 25. DOI:hŠp: //dx.doi.org/10.1080/10556781003625177
  5. 5.James P Boyle and Richard L Dykstra. 1986. A method for €nding projections onto the intersection of convex sets in Hilbert spaces. In Advances in order restricted statistical inference. Springer, 28–47.
  6. 6.E J Candes, X Li, Y Ma, and J Wright. 2009. Robust Principal Component Analysis? ` Preprint: arXiv:0912.3599 (2009).
  7. 7.David L. Donoho. 2006. For most large underdetermined systems of linear equations the minimal ’? 1-norm solution is also the sparsest solution. Communications on Pure and Applied Mathematics 59, 6 (2006), 797–829. DOI: hŠp://dx.doi.org/10.1002/cpa.20132 arXiv:0912.3599
  8. 8.J. Friedman, T. Hastie, and R. Tibshirani. 2008. Še Elements of Statistical Learning. Vol. 2. hŠp://www-stat.stanford.edu/
  9. 9.Jonas Gehring, Yajie Miao, Florian Metze, and Alex Waibel. 2013. Extracting deep boŠleneck features using stacked auto-encoders. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings (2013), 3377– 3381. DOI:hŠp://dx.doi.org/10.1109/ICASSP.2013.6638284
  10. 10.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. (2016). DOI:hŠp://dx.doi.org/10.1038/nmeth.3707 arXiv:arXiv:1312.6184v5
  11. 11.Y LeCun, Y Bengio, and G Hinton. 2015. Deep learning. Nature (2015). hŠp: //www.nature.com/nature/journal/v521/n7553/abs/nature14539.html
  12. 12.Y LeCun, L BoŠou, and Y Bengio. 1998. Gradient-based learning applied to document recognition. Proceedings of the (1998). hŠp://ieeexplore.ieee.org/ abstract/document/726791/
  13. 13.Honglak Lee, Alexis BaŠle, Rajat Raina, and Andrew Y Ng. 2007. Ecient sparse coding algorithms. Advances in neural information processing systems 19 (2007), 801.
  14. 14.Fei Tony Liu and Kai Ming Ting. 2008. Isolation Forest. 2008 Eighth IEEE International (2008). hŠp://ieeexplore.ieee.org/xpls/abs
  15. 15.O Lyudchik. 2016. Outlier detection using autoencoders. (2016). hŠp://cds.cern. ch/record/2209085
  16. 16.Yunlong Ma, Peng Zhang, Yanan Cao, and Li Guo. 2013. Parallel auto-encoder for ecient outlier detection. In Proceedings - 2013 IEEE International Conference on Big Data, Big Data 2013. 15–17. DOI:hŠp://dx.doi.org/10.1109/BigData.2013. 6691791
  17. 17.Lingheng Meng, Shifei Ding, and Yu Xue. 2016. Research on denoising sparse autoencoder. International Journal of Machine Learning and Cybernetics (2016), 1–11. DOI:hŠp://dx.doi.org/10.1007/s13042-016-0550-y
  18. 18.So€a Mosci, Lorenzo Rosasco, MaŠeo Santoro, Alessandro Verri, and Silvia Villa. 2010. Solving structured sparsity regularization with proximal methods. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 418–433.
  19. 19.Ulric Neisser. 2004. PaŠern Recognition. Cognitive psychology: Key readings. (2004). DOI:hŠp://dx.doi.org/10.1016/j.patcog.2011.03.019 arXiv:arXiv:1011.1669v3
  20. 20.Randy C. Pa‚enroth, Philip C. Du Toit, Louis L Scharf, Anura P Jayasumana, Vidarshana Bandara, and Ryan Nong. 2012. Space-time signal processing for distributed paŠern detection in sensor networks. Proceedings of SPIE - Še International Society for Optical Engineering 8393 (2012), Œe Society of Photo– Optical Instrumentation Engin. DOI:hŠp://dx.doi.org/10.1117/12.919711
  21. 21.Fabian Pedregosa, Gael Varoquaux, Alexandre Gramfort, Vincent Michel, ¨ Bertrand Œirion, Olivier Grisel, Mathieu Blondel, Peter PreŠenhofer, Ron Weiss, Vincent Dubourg, and others. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research 12, Oct (2011), 2825–2830.
  22. 22.Yu Qi, Yueming Wang, Xiaoxiang Zheng, and Zhaohui Wu. 2014. Robust feature learning by stacked autoencoder with maximum correntropy criterion. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. 6716–6720. DOI:hŠp://dx.doi.org/10.1109/ICASSP.2014.6854900
  23. 23.Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio. 2011. Contractive auto-encoders: Explicit invariance during feature extraction. In Proceedings of the 28th international conference on machine learning (ICML-11). 833–840.
  24. 24.David E. Rumelhart, Geo‚rey E. Hinton, and Ronald J. Williams. 1986. Learning representations by back-propagating errors. Nature 323, 6088 (oct 1986), 533–536. DOI:hŠp://dx.doi.org/10.1038/323533a0 arXiv:arXiv:1011.1669v3
  25. 25.Craig A. Stewart, George Turner, MaŠhew Vaughn, Niall I. Ga‚ney, Timothy M. Cockerill, Ian Foster, David Hancock, Nirav Merchant, Edwin Skidmore, Daniel Stanzione, James Taylor, and Steven Tuecke. 2015. Jetstream: A self-provisioned, scalable science and engineering cloud environment. Proceedings of the 2015 XSEDE Conference on Scienti€c Advancements Enabled by Enhanced Cyberinfrastructure - XSEDE ’15 (2015), 1–8. DOI:hŠp://dx.doi.org/10.1145/2792745.2792774
  26. 26.Œeano Development Team. 2016. Œeano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs/1605.02688 (May 2016). hŠp://arxiv.org/abs/1605.02688
  27. 27.John Towns, Timothy Cockerill, Maytal Dahan, Ian Foster, Kelly Gaither, Andrew Grimshaw, Victor Hazlewood, ScoŠ Lathrop, Dave Liˆa, Gregory D. Peterson, Ralph Roskies, J. Ray ScoŠ, and Nancy Wilkens-Diehr. 2014. XSEDE: Accelerating scienti€c discovery. Computing in Science and Engineering 16, 5 (sep 2014), 62–74. DOI:hŠp://dx.doi.org/10.1109/MCSE.2014.80 arXiv:hŠp://dx.doi.org/10.1109/MCSE.2014.80
  28. 28.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008. Extracting and composing robust features with denoising autoencoders. Proceedings of the 25th international conference on Machine learning - ICML ’08 (2008), 1096–1103. DOI:hŠp://dx.doi.org/10.1145/1390156.1390294 arXiv:arXiv:1412.6550v4
  29. 29.Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and PierreAntoine Manzagol. 2010. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research 11, Dec (2010), 3371–3408.
  30. 30.J Xie, L Xu, and E Chen. 2012. Image denoising and inpainting with deep neural networks. Advances in Neural Information Processing (2012), 1–9. hŠp://papers.nips.cc/paper/ 4686-image-denoising-and-inpainting-with-deep-neural-networkshŠp: //papers.nips.cc/paper/4686-image-denoising
  31. 31.Dan Zhao, Baolong Guo, Jinfu Wu, Weikang Ning, and Yunyi Yan. 2015. Robust feature learning by improved auto-encoder from non-Gaussian noised images. In IST 2015 - 2015 IEEE International Conference on Imaging Systems and Techniques, Proceedings. DOI:hŠp://dx.doi.org/10.1109/IST.2015.7294537

Citation

MLA
Zhou, C., and R. C. Paffenroth. “Anomaly Detection with Robust Deep Autoencoders”. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, pp. 665–74, https://doi.org/10.1145/3097983.3098052.
APA
Zhou, C., & Paffenroth, R. C. (2017). Anomaly Detection with Robust Deep Autoencoders. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 665–674. https://doi.org/10.1145/3097983.3098052
Chicago
Zhou, C., and R. C. Paffenroth. 2017. “Anomaly Detection with Robust Deep Autoencoders”. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 665–74. https://doi.org/10.1145/3097983.3098052.
Harvard
Zhou, C. and Paffenroth, R.C. (2017) “Anomaly Detection with Robust Deep Autoencoders”, Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, pp. 665–674. Available at: https://doi.org/10.1145/3097983.3098052.
Vancouver
1. Zhou C, Paffenroth RC (2017) Anomaly Detection with Robust Deep Autoencoders. In: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, pp 665–674

BibTeX

@inproceedings{Zhou_2017, series={KDD ’17}, title={Anomaly Detection with Robust Deep Autoencoders}, url={http://dx.doi.org/10.1145/3097983.3098052}, DOI={10.1145/3097983.3098052}, booktitle={Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining}, publisher={ACM}, author={Zhou, Chong and Paffenroth, Randy C.}, year={2017}, month=Aug, pages={665–674}, collection={KDD ’17} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF