Deep Learning for Anomaly Detection: A Survey

Raghavendra ChalapathySanjay Chawla

article2019arXiv1,908 citations

Classifies deep learning anomaly detection methods across diverse application domains, evaluating their underlying assumptions, computational complexities, and practical limitations to guide model selection and identify critical research challenges.

Listen

Modern data systems generate massive volumes of complex, high-dimensional information across critical sectors, including cybersecurity, finance, industrial monitoring, and healthcare. Detecting anomalies—such as malicious intrusions, equipment failures, or financial fraud—is essential for mitigating substantial financial, operational, and reputational risks. However, traditional anomaly detection techniques struggle to process massive datasets and fail to capture intricate spatial or temporal patterns, making advanced automated approaches increasingly vital.

The article systematically evaluates deep learning frameworks for anomaly detection to establish a comprehensive categorization of state-of-the-art methods and assess their effectiveness across major application areas. To do this, the authors conducted an extensive literature review across various operational domains—including video surveillance, network security, credit card fraud, medical diagnostics, and time-series monitoring—analyzing the core assumptions, network architectures, and computational complexities of different deep learning techniques.

The review yields several critical findings regarding model design and application. First, deep learning models significantly outperform traditional statistical and machine learning algorithms when handling large-scale, high-dimensional, and unstructured inputs such as raw images, text, and sequences. Second, semi-supervised and unsupervised approaches—particularly reconstruction-based autoencoders—are the most practical in real-world settings due to the severe scarcity of verified anomaly labels. Third, traditional two-step hybrid models that feed deep features into conventional classifiers suffer from suboptimal performance because their feature extractors are not tuned specifically for anomaly detection. In contrast, specialized end-to-end architectures, such as one-class neural networks, optimize internal representations directly for anomaly boundaries and achieve superior results. Finally, recurrent architectures like Long Short-Term Memory networks dominate sequential and time-series anomaly detection, whereas convolutional networks are most effective for spatial and image data.

These findings demonstrate that adopting deep anomaly detection can substantially enhance operational risk management, safety compliance, and system reliability by reducing the need for costly, manual feature engineering. However, organizational leaders must weigh accuracy against significant computational training costs and operational overhead. Because deep architectures can be sensitive to noisy baseline data, improper model selection can lead to elevated false positive rates or undetected anomalous behavior.

Organizations should transition toward end-to-end deep anomaly detection architectures that are tailored to their specific data modalities, rather than relying on unaligned hybrid approaches. Before deploying these systems at scale, decision-makers should implement robust data-cleaning pipelines and conduct targeted pilot evaluations to establish optimal decision thresholds. Further empirical validation and exploration are required for emerging paradigms, such as deep reinforcement learning and zero-shot anomaly detection, before they can be broadly applied to mission-critical infrastructure.

arXiv: 1901.03407
Cover for Deep Learning for Anomaly Detection: A Survey

Abstract

Anomaly detection is an important problem that has been well-studied within diverse research areas and application domains. The aim of this survey is two-fold, firstly we present a structured and comprehensive overview of research methods in deep learning-based anomaly detection. Furthermore, we review the adoption of these methods for anomaly across various application domains and assess their effectiveness. We have grouped state-of-the-art research techniques into different categories based on the underlying assumptions and approach adopted. Within each category we outline the basic anomaly detection technique, along with its variants and present key assumptions, to differentiate between normal and anomalous behavior. For each category, we present we also present the advantages and limitations and discuss the computational complexity of the techniques in real application domains. Finally, we outline open issues in research and challenges faced while adopting these techniques.

Table of Contents

  • 1 Introduction
  • 2 What are anomalies?
  • 3 What are novelties?
  • 4 Motivation and Challenges: Deep anomaly detection (DAD) techniques
  • 5 Related Work
  • 6 Our Contributions
  • 7 Organization
  • 8 Different aspects of deep learning-based anomaly detection.
  • 8.1 Nature of Input Data
  • 8.2 Based on Availability of labels
  • 8.2.1 Supervised deep anomaly detection
  • 8.2.2 Semi-supervised deep anomaly detection
  • 8.2.3 Unsupervised deep anomaly detection
  • 8.3 Based on the training objective
  • 8.3.1 Deep Hybrid Models (DHM)
  • 8.3.2 One-Class Neural Networks (OC-NN)
  • 8.4 Type of Anomaly
  • 8.4.1 Point Anomalies
  • 8.4.2 Contextual Anomaly Detection
  • 8.4.3 Collective or Group Anomaly Detection.
  • 8.5 Output of DAD Techniques
  • 8.5.1 Anomaly Score:
  • 8.5.2 Labels:
  • 9 Applications of Deep Anomaly Detection
  • 9.1 Intrusion Detection
  • 9.1.1 Host-Based Intrusion Detection Systems (HIDS):
  • 9.1.2 Network Intrusion Detection Systems (NIDS):
  • 9.2 Fraud Detection
  • 9.2.1 Banking fraud
  • 9.2.2 Mobile cellular network fraud
  • 9.2.3 Insurance fraud
  • 9.2.4 Healthcare fraud
  • 9.3 Malware Detection
  • 9.4 Medical Anomaly Detection
  • 9.5 Deep learning for Anomaly detection in Social Networks
  • 9.6 Log Anomaly Detection
  • 9.7 Internet of things (IoT) Big Data Anomaly Detection
  • 9.8 Industrial Anomalies Detection
  • 9.9 Anomaly Detection in Time Series
  • 9.9.1 Uni-variate time series deep anomaly detection
  • 9.9.2 Multi-variate time series deep anomaly detection
  • 9.10 Video Surveillance
  • 10 Deep Anomaly Detection (DAD) Models
  • 10.1 Supervised deep anomaly detection
  • 10.2 Semi-supervised deep anomaly detection
  • 10.3 Hybrid deep anomaly detection
  • 10.4 One-class neural networks (OC-NN) for anomaly detection
  • 10.5 Unsupervised Deep Anomaly Detection
  • 10.6 Miscellaneous Techniques
  • 10.6.1 Transfer Learning based anomaly detection
  • 10.6.2 Zero Shot learning based anomaly detection
  • 10.6.3 Ensemble based anomaly detection
  • 10.6.4 Clustering based anomaly detection
  • 10.6.5 Deep Reinforcement Learning (DRL) based anomaly detection
  • 10.6.6 Statistical techniques deep anomaly detection
  • 11 Deep neural network architectures for locating anomalies
  • 11.1 Deep Neural Networks (DNN)
  • 11.2 Spatio Temporal Networks (STN)
  • 11.3 Sum-Product Networks (SPN)
  • 11.4 Word2vec Models
  • 11.5 Generative Models
  • 11.6 Convolutional Neural Networks
  • 11.7 Sequence Models
  • 11.8 Autoencoders
  • 12 Relative Strengths and Weakness : Deep Anomaly Detection Methods
  • 13 Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy and Categorization Schema of Deep Anomaly Detection

    model/method

    Deep Anomaly Detection (DAD) encompasses neural network frameworks designed to identify patterns that deviate significantly from expected normal behavior. DAD methods are systematically classified along four principal orthogonal axes:

    1. Label Availability:

      • Supervised DAD: Uses labeled instances of both normal and anomalous classes to train discriminative classifiers.
      • Semi-supervised (One-Class) DAD: Assumes training data consists solely or primarily of normal instances, learning an enclosing boundary or normal profile.
      • Unsupervised DAD: Operates entirely without class labels, leveraging intrinsic properties of the data distribution (e.g., density, distance, or reconstruction fidelity).
    2. Training Objective Formulation:

      • Deep Hybrid Models (DHM): Decoupled two-stage architectures utilizing deep neural networks as generic feature extractors whose outputs are fed into shallow anomaly detectors (such as One-Class Support Vector Machines or kernel estimators).
      • One-Class Neural Networks (OC-NN): End-to-end networks where representation learning in hidden layers is directly driven and optimized jointly by a customized anomaly detection objective (e.g., enclosing hyperplanes or hyperspheres).
    3. Anomaly Nature and Structure:

      • Point Anomalies: Individual data points that randomly deviate from the majority distribution.
      • Contextual (Conditional) Anomalies: Instances that are anomalous only when evaluated with respect to specific contextual variables (e.g., temporal or spatial features).
      • Collective (Group) Anomalies: Subsets or sequences of data points that appear normal individually but whose collective arrangement or joint distribution is unusual.
    4. Input Data Modality Alignment:

      • Sequential Data: Handled by recurrent architectures (RNN, LSTM, GRU) and temporal autoencoders.
      • Non-Sequential / Spatial Data: Handled by Convolutional Neural Networks (CNN), standard Autoencoders, and deep generative models.
  2. Knowl 2 — Deep Hybrid Models for Anomaly Detection

    model/method

    Deep Hybrid Models (DHMs) for anomaly detection combine deep neural networks with classical shallow anomaly detection algorithms through a decoupled, two-stage framework:

    1. Feature Extraction: A deep architecture—such as a deep autoencoder (AE), deep belief network (DBN), or a convolutional neural network (CNN) pre-trained via transfer learning—is trained to map high-dimensional raw inputs x∈Rdx \in \mathbb{R}^d into a lower-dimensional latent representation z∈Rd′z \in \mathbb{R}^{d'} where d′≪dd' \ll d.
    2. Shallow Anomaly Detection: The extracted low-dimensional representations zz are fed as input into traditional outlier detection algorithms, such as One-Class Support Vector Machines (OC-SVM), Support Vector Data Description (SVDD), or kk-Nearest Neighbors (kk-NN).

    Key Assumptions:

    • Robust, compressed features extracted within the hidden layers of the deep neural network filter out irrelevant noise and preserve discriminative characteristics necessary to identify anomalous deviations.
    • Operating on the low-dimensional latent space mitigates the curse of dimensionality inherent to high-dimensional anomaly detection.

    Computational Complexity:

    • Training includes the training complexity of the deep neural network plus the complexity of the shallow detector. For an OC-SVM with an RBF kernel, testing complexity scales as O(nd′)\mathcal{O}(n d'), where nn is the number of support vectors and d′d' is the latent dimension, compared to O(d′)\mathcal{O}(d') for a linear kernel.

    Limitations:

    • DHMs are fundamentally decoupled and suboptimal because the feature extractor is trained using generic objectives (such as reconstruction loss) rather than an objective tailored to separate normal instances from anomalies. The hidden representations are not directly guided to optimize anomaly separability.
  3. Knowl 3 — One-Class Neural Networks for Anomaly Detection

    model/method

    One-Class Neural Networks (OC-NN) and related architectures (such as Deep Support Vector Data Description, Deep SVDD) unify deep feature representation learning with a customized one-class optimization objective in an end-to-end framework.

    Operational Mechanism:

    • Rather than extracting features via an independent proxy task, the hidden layer transformations f(x;W)f(x; \mathcal{W}) parameterized by weights W\mathcal{W} are trained jointly with a one-class decision boundary (such as a minimum-volume hypersphere centered at cc or a separating hyperplane).
    • For Deep SVDD, the objective minimizes the mean squared distance of all transformed representations from a predefined center point cc:

    min⁡W1N∑i=1N∥f(xi;W)−c∥2+λ2∑l=1L∥Wl∥F2\min_{\mathcal{W}} \frac{1}{N} \sum_{i=1}^N \Vert f(x_i; \mathcal{W}) - c \Vert^2 + \frac{\lambda}{2} \sum_{l=1}^L \Vert \mathcal{W}^l \Vert_F^2

    where NN is the number of training samples, LL is the number of hidden layers, Wl\mathcal{W}^l represents the weight matrix at layer ll, ∥⋅∥F\Vert \cdot \Vert_F is the Frobenius norm, and λ>0\lambda > 0 is a regularization hyperparameter.

    • In OC-NN, optimization is performed via alternating minimization between the neural network parameters and the one-class boundary threshold, which corresponds to solving a quantile selection subproblem.

    Core Assumptions:

    • The deep network can learn common factors of variation underlying the normal data distribution.
    • Because anomalous samples lack these common factors of variation, the learned mapping fails to compress them near the center or within the enclosing boundary.

    Advantages and Limitations:

    • Advantages: End-to-end training optimizes feature representations specifically for anomaly detection, eliminating the suboptimality of hybrid models. OC-NN does not store support vectors at inference time, yielding O(1)\mathcal{O}(1) storage complexity per test instance.
    • Limitations: Training and model update times scale with network depth and input dimensionality, requiring re-training or adaptation when the input feature space evolves.
  4. Knowl 4 — Autoencoder-Based Deep Anomaly Detection

    model/method

    Autoencoders (AEs) are unsupervised and semi-supervised neural network architectures trained to reconstruct their input data through an information bottleneck. In anomaly detection, an autoencoder consists of an encoder network gϕ:X→Zg_{\phi}: \mathcal{X} \to \mathcal{Z} mapping input x∈Xx \in \mathcal{X} to a latent space Z\mathcal{Z}, and a decoder network hθ:Z→Xh_{\theta}: \mathcal{Z} \to \mathcal{X} reconstructing the input as x^=hθ(gϕ(x))\hat{x} = h_{\theta}(g_{\phi}(x)).

    Anomaly Scoring:

    • The network is trained primarily on normal data to minimize a reconstruction error metric (such as mean squared error):

    L(x,x^)=∥x−hθ(gϕ(x))∥22\mathcal{L}(x, \hat{x}) = \Vert x - h_{\theta}(g_{\phi}(x)) \Vert_2^2

    • During inference, the anomaly score s(x)s(x) of an unseen data sample xx is defined directly by its reconstruction error magnitude:

    s(x)=∥x−x^∥22s(x) = \Vert x - \hat{x} \Vert_2^2

    Instances producing a reconstruction error exceeding a threshold τ\tau are classified as anomalies.

    Architectural Variants:

    • Denoising Autoencoders (DAE): Trained to reconstruct clean inputs from partially corrupted inputs, enforcing robust representations.
    • Convolutional Autoencoders (CAE / CNN-AE): Utilize convolutional and deconvolutional layers for spatial/image anomaly detection.
    • Recurrent Autoencoders (LSTM-AE / GRU-AE): Utilize recurrent encoder-decoder structures to capture temporal regularities in time-series and sequential logs.

    Assumptions and Limitations:

    • Assumption: Autoencoders constrained by an information bottleneck learn the low-dimensional manifold of normal data patterns and will fail to accurately reconstruct rare or out-of-distribution anomalous patterns.
    • Limitation: If contaminated with unlabelled anomalies during training, autoencoders may inadvertently learn to reconstruct anomalous patterns, degrading detection performance unless robust regularizations are applied. The compression ratio (latent space dimension) acts as a sensitive hyperparameter.
  5. Knowl 5 — Definitions and Properties of Anomaly Types

    definition

    In deep anomaly detection, anomalies are categorized into three distinct fundamental types based on their structural characteristics:

    1. Point Anomalies:

      • An individual data instance that deviates significantly from the rest of the dataset in a random or isolated fashion without structural context.
      • Example: A single credit card transaction involving an abnormally high expenditure relative to all standard transactions.
    2. Contextual (Conditional) Anomalies:

      • A data instance that is anomalous within a specific context (such as time, spatial location, or operational state), but may appear completely normal when evaluated outside that context.
      • An instance is evaluated using two sets of attributes: contextual features (which define the context, e.g., timestamp, geographic position) and behavioral features (which define the evaluation metric, e.g., temperature reading, system call event).
      • Example: A sudden temperature drop near summer months is anomalous within that seasonal context, even if such a temperature is normal in winter.
    3. Collective (Group) Anomalies:

      • A collection of related data instances that are anomalous with respect to the entire dataset, even though the individual instances composing the collection appear normal when inspected in isolation.
      • Example: A repeated sequence of small-dollar credit card charges or abnormal clusters of image pixels that form an irregular joint distribution.
  6. Knowl 6 — Semi-Supervised Deep Anomaly Detection

    model/method

    Semi-supervised Deep Anomaly Detection (DAD), also referred to as one-class deep anomaly detection, operates in settings where training data consists entirely or almost entirely of labeled normal class samples (yi=normaly_i = \text{normal}), while anomalous examples are absent or scarce during training.

    Operational Principles:

    • The model learns a compact representation or envelope bounding the normal class in latent feature space.
    • Techniques encompass one-class generative models, semi-supervised Generative Adversarial Networks (GANs), and autoencoders trained purely on normal data. During testing, inputs deviating from the normal profile or failing to match the generative distribution are flagged as anomalous.

    Core Assumptions:

    • Proximity and Continuity: Data instances that are close to each other in the input space and the learned latent feature space share the same class label (normal).
    • Discriminative Feature Extraction: Deep hidden layers extract invariant, robust factors of normal behavior, preserving discriminative attributes that separate normal instances from outliers.

    Advantages and Limitations:

    • Advantages: Significantly outperforms unsupervised approaches by utilizing known normal labels without requiring expensive or unavailable anomaly annotations.
    • Limitations: Models can suffer from overfitting to the available normal distribution, making them sensitive to shifts in normal behavior or unrepresentative training samples.
  7. Knowl 7 — Unsupervised Deep Anomaly Detection

    model/method

    Unsupervised Deep Anomaly Detection (DAD) operates on unlabelled datasets containing an unknown mixture of normal and anomalous instances, seeking to identify outliers based solely on the intrinsic properties of the data.

    Operational Principles:

    • Unsupervised deep models (primarily Autoencoders, Deep Belief Networks, Restricted Boltzmann Machines, and Variational Autoencoders) project the unlabelled dataset into low-dimensional latent spaces to capture inherent data distributions, densities, or reconstruction manifolds.
    • Anomaly scores are assigned to each instance based on reconstruction errors, energy scores, or estimated probability densities.

    Core Assumptions:

    • Normal instances constitute the overwhelming majority of the dataset, while anomalies are rare and statistically infrequent.
    • Normal and anomalous regions are separable in either the original input feature space or the learned latent representation space.
    • Intrinsic geometric properties (such as distance or density) in the latent representation capture normal data patterns, while anomalies deviate from these regularities.

    Complexity and Limitations:

    • Computational Complexity: Unlike linear methods such as Principal Component Analysis (PCA) which rely on closed-form matrix decomposition, training deep unsupervised models involves non-convex optimization via backpropagation with quadratic or higher computational costs depending on layer depth and network parameters.
    • Limitations: Unsupervised DAD is susceptible to noise and training data contamination; if anomalies appear with non-negligible frequency, the network models anomalous patterns as normal, yielding elevated false-positive and false-negative rates.
  8. Knowl 8 — Generative Adversarial Networks and Generative Models for Anomaly Detection

    model/method

    Deep generative models for anomaly detection utilize architectures such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Adversarial Autoencoders (AAEs) to learn the probability density distribution of normal data.

    Operational Frameworks:

    1. Variational Autoencoders (VAEs):
      • Approximate the data distribution p(x)p(x) by mapping inputs to a parameterized latent distribution qϕ(z∣x)=N(μ(x),Σ(x))q_{\phi}(z|x) = \mathcal{N}(\mu(x), \Sigma(x)) and reconstructing via pθ(x∣z)p_{\theta}(x|z). Anomaly scores are derived from reconstruction probability or the evidence lower bound (ELBO).
    2. Generative Adversarial Networks (GAN-AD):
      • A generator GG learns a mapping from a latent prior z∼pzz \sim p_z to the data space, while a discriminator DD distinguishes between real samples and generated samples.
      • For an unseen sample xx, anomaly detection involves finding an optimal latent vector z∗z^* such that G(z∗)G(z^*) best reconstructs xx. The anomaly score combines the residual reconstruction loss ∥x−G(z∗)∥\Vert x - G(z^*) \Vert and the discriminator feature matching loss ∥fD(x)−fD(G(z∗))∥\Vert f_D(x) - f_D(G(z^*)) \Vert.
    3. Adversarial Autoencoders (AAEs):
      • Impose an arbitrary statistical prior on the autoencoder's latent space via adversarial training, aligning latent representations with a target distribution to improve regularized density estimation.

    Assumptions and Trade-offs:

    • Assumption: Generative models accurately map the manifold of normal training instances; out-of-distribution anomalies cannot be generated from the learned latent distribution.
    • Trade-offs: While effective at capturing complex high-dimensional distributions (e.g., in images and multivariate time series), GAN-based inference requires iterative optimization in the latent space during testing, increasing inference latency. In settings with very small anomaly prevalence, shallow methods like kk-NN can sometimes match or exceed generative performance.
  9. Knowl 9 — Supervised Deep Anomaly Detection and Class Imbalance Challenges

    model/method

    Supervised Deep Anomaly Detection (DAD) treats anomaly detection as a supervised binary or multi-class classification problem using annotated datasets containing both normal and anomalous instances.

    Architecture and Mechanism:

    • Supervised DAD models typically consist of two cascaded components: a deep hierarchical feature extraction sub-network (e.g., CNN or deep MLP) followed by a supervised classification sub-network (e.g., Softmax layer or fully connected classifier).
    • The network is trained using standard cross-entropy or multi-class classification objectives to establish a decision boundary separating anomaly classes from normal classes.

    Advantages:

    • Yields higher classification accuracy and lower false alarm rates than semi-supervised and unsupervised methods when sufficient balanced labeled data is available.
    • Inference is computationally fast (an O(1)\mathcal{O}(1) forward pass) against the precomputed model.

    Limitations:

    • Label Scarcity: Ground-truth anomaly labels are difficult, expensive, and often impossible to obtain at scale across practical application domains.
    • Severe Class Imbalance: Normal instances vastly outnumber anomalies, leading standard deep classifiers to exhibit biased decision boundaries that favor the majority normal class.
    • Novelty Generalization: Supervised models are inherently constrained to the specific anomaly types seen during training and typically fail to detect novel, previously unseen attack mechanisms or failure modes.
  10. Knowl 10 — Data Modality and Architecture Alignment in Deep Anomaly Detection

    model/method

    The selection of a deep neural network architecture for anomaly detection is strictly governed by the structural modality of the input data:

    Data Type Application Examples Recommended DAD Architectures
    Sequential Time series, speech, system logs, CNN, RNN, LSTM, GRU,
    video streams, protein sequences, text LSTM-AE, GRU-AE
    Non-Sequential Images, sensor snapshots, CNN, Autoencoders (AE, DAE,
    / Spatial high-dimensional tabular records CAE, VAE, AAE), DBN, GAN

    Key Alignment Principles:

    • Sequential and Temporal Data: Preserves sequential dependencies, temporal correlations, and variable-length inputs via recurrent units (LSTMs, GRUs) or 1D convolutions.
    • Spatial and Image Data: Leverages 2D/3D convolutional layers (CNNs, Convolutional Autoencoders) to exploit local spatial correlation, translation invariance, and hierarchical visual representations.
    • Dimensionality Scaling: The number of hidden layers in DAD architectures scales with the dimensionality of the input data; deeper networks are required to learn meaningful hierarchical features from high-dimensional inputs, which proportionally increases training complexity and parameter update costs.
  11. Knowl 11 — Emerging Paradigms in Deep Anomaly Detection

    model/method

    Beyond traditional deep classification and autoencoding, several specialized machine learning paradigms are integrated into deep anomaly detection (DAD):

    1. Transfer Learning for DAD:

      • Reuses feature representations pre-trained on large-scale source datasets to anomaly detection tasks in target domains where training data is scarce, relaxing the requirement of identical feature distributions between source and target domains.
    2. Zero-Shot Learning (ZSL) for DAD:

      • Detects previously unseen anomalous classes by leveraging auxiliary semantic metadata (such as natural language attributes or textual descriptions). ZSL captures the relationship between attributes and seen classes to classify samples belonging to unseen anomaly classes without prior training instances.
    3. Autoencoder Ensembles:

      • Trains an ensemble of autoencoders with randomized connectivity architectures and varying network structures. Ensembling enhances diversity, reduces variance, improves robustness against input noise, and prevents overfitting compared to single autoencoders.
    4. Clustering-Based Deep Anomaly Detection:

      • Uses deep networks (e.g., Word2Vec embeddings or autoencoder latent representations) to project high-dimensional data into low-dimensional semantic spaces where clustering algorithms (e.g., DBSCAN, kk-means) group normal instances into dense clusters, flagging out-of-cluster points as anomalies.
    5. Deep Reinforcement Learning (DRL) for DAD:

      • Formulates anomaly detection as an agent interacting with an environment, learning an anomaly detection policy dynamically through accumulated trial-and-error reward signals without predefined static assumptions about anomaly distributions.

Coverage note — We covered the complete core taxonomy, methodological frameworks (DHMs, OC-NN, Autoencoders, Generative Models, Supervised, Semi-Supervised, Unsupervised DAD), anomaly definitions, data-modality mappings, and emerging paradigms (Transfer Learning, ZSL, Ensembles, DRL). We omitted specific domain-by-domain external literature surveys across applied sub-areas (e.g., specific lists of cited papers in IoT, healthcare, cyber-intrusion) as they represent secondary bibliographic listings rather than standalone methodological contributions of the survey.

References

  1. 1.Varun Chandola, Arindam Banerjee, and Vipin Kumar. Outlier detection: A survey. ACM Computing Surveys, 2007.
  2. 2.D. Hawkins. Identification of Outliers. Chapman and Hall, London, 1980.
  3. 3.Ahmad Javaid, Quamar Niyaz, Weiqing Sun, and Mansoor Alam. A deep learning approach for network intrusion detection system. In Proceedings of the 9th EAI International Conference on Bio-inspired Information and Communications Technologies (formerly BIONETICS), pages 21–26. ICST (Institute for Computer Sciences, Social-Informatics and Telecommunications Engineering), 2016.
  4. 4.Huan-Kai Peng and Radu Marculescu. Multi-scale compositionality: identifying the compositional structures of social dynamics using deep learning. PloS one, 10(4):e0118309, 2015.
  5. 5.Bahnsen Alejandro, Correa. Building ai applications using deep learning. 2016. URL https://blog.easysol.net/wp-content/uploads/2017/06/image1.png.
  6. 6.Xuemei Xie, Chenye Wang, Shu Chen, Guangming Shi, and Zhifu Zhao. Real-time illegal parking detection system based on deep learning. In Proceedings of the 2017 International Conference on Deep Learning Technologies, pages 23–27. ACM, 2017.
  7. 7.Thomas Schlegl, Philipp Seebock, Sebastian M Waldstein, Ursula Schmidt-Erfurth, and Georg Langs. Unsupervised anomaly detection with generative adversarial networks to guide marker discovery. In International Conference on Information Processing in Medical Imaging, pages 146–157. Springer, 2017.
  8. 8.Mehdi Mohammadi, Ala Al-Fuqaha, Sameh Sorour, and Mohsen Guizani. Deep learning for iot big data and streaming analytics: A survey. arXiv preprint arXiv:1712.04301, 2017.
  9. 9.Charu C Aggarwal. An introduction to outlier analysis. In Outlier analysis, pages 1–40. Springer, 2013.
  10. 10.Dubravko Miljkovic. Review of novelty detection methods. In MIPRO, 2010 proceedings of the 33rd international convention, pages 593–598. IEEE, 2010.
  11. 11.Marco AF Pimentel, David A Clifton, Lei Clifton, and Lionel Tarassenko. A review of novelty detection. Signal Processing, 99:215–249, 2014.
  12. 12.Donghwoon Kwon, Hyunjoo Kim, Jinoh Kim, Sang C Suh, Ikkyun Kim, and Kuinam J Kim. A survey of deep learning-based network anomaly detection. Cluster Computing, pages 1–13, 2017.
  13. 13.John E Ball, Derek T Anderson, and Chee Seng Chan. Comprehensive survey of deep learning in remote sensing: theories, tools, and challenges for the community. Journal of Applied Remote Sensing, 11(4):042609, 2017.
  14. 14.B Ravi Kiran, Dilip Mathew Thomas, and Ranjith Parakkal. An overview of deep learning based methods for unsupervised and semi-supervised anomaly detection in videos. arXiv preprint arXiv:1801.03149, 2018.
  15. 15.Aderemi O Adewumi and Andronicus A Akinyelu. A survey of machine-learning and nature-inspired based credit card fraud detection techniques. International Journal of System Assurance Engineering and Management, 8(2):937–953, 2017.
  16. 16.Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen Awm Van Der Laak, Bram Van Ginneken, and Clara I Sanchez. A survey on deep learning in medical image analysis. Medical image analysis, 42:60–88, 2017.
  17. 17.Sarah M Erfani, Sutharshan Rajasegarar, Shanika Karunasekera, and Christopher Leckie. High-dimensional and large-scale anomaly detection using a linear one-class svm with deep learning. Pattern Recognition, 58:121–134, 2016a.
  18. 18.Raghavendra Chalapathy, Aditya Krishna Menon, and Sanjay Chawla. Anomaly detection using one-class neural networks. arXiv preprint arXiv:1802.06360, 2018a.
  19. 19.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436, 2015.
  20. 20.Daniel Ramotsoela, Adnan Abu-Mahfouz, and Gerhard Hancke. A survey of anomaly detection in industrial wireless sensor networks with critical water system infrastructure as a case study. Sensors, 18(8):2491, 2018.
  21. 21.Raghavendra Chalapathy, Ehsan Zare Borzeshi, and Massimo Piccardi. An investigation of recurrent neural architectures for drug name recognition. arXiv preprint arXiv:1609.07585, 2016a.
  22. 22.Raghavendra Chalapathy, Ehsan Zare Borzeshi, and Massimo Piccardi. Bidirectional lstm-crf for clinical concept extraction. arXiv preprint arXiv:1611.08373, 2016b.
  23. 23.Drausin Wulsin, Justin Blanco, Ram Mani, and Brian Litt. Semi-supervised anomaly detection for eeg waveforms using deep belief nets. In Machine Learning and Applications (ICMLA), 2010 Ninth International Conference on, pages 436–441. IEEE, 2010.
  24. 24.Mutahir Nadeem, Ochaun Marshall, Sarbjit Singh, Xing Fang, and Xiaohong Yuan. Semi-supervised deep neural network for network intrusion detection. 2016.
  25. 25.Hongchao Song, Zhuqing Jiang, Aidong Men, and Bo Yang. A hybrid semi-supervised anomaly detection model for high-dimensional data. Computational intelligence and neuroscience, 2017, 2017.
  26. 26.Josh Patterson and Adam Gibson. Deep Learning: A Practitioner’s Approach. ” O’Reilly Media, Inc.”, 2017.
  27. 27.Aaron Tuor, Samuel Kaplan, Brian Hutchinson, Nicole Nichols, and Sean Robinson. Deep learning for unsupervised insider threat detection in structured cybersecurity data streams. arXiv preprint arXiv:1710.00811, 2017.
  28. 28.Svante Wold, Kim Esbensen, and Paul Geladi. Principal component analysis. Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987.
  29. 29.Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine learning, 20(3):273–297, 1995.
  30. 30.Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou. Isolation forest. In Data Mining, 2008. ICDM’08. Eighth IEEE International Conference on, pages 413–422. IEEE, 2008.
  31. 31.Ilya Sutskever, Geoffrey E Hinton, and Graham W Taylor. The recurrent temporal restricted boltzmann machine. In Advances in Neural Information Processing Systems, pages 1601–1608, 2009.
  32. 32.Ruslan Salakhutdinov and Hugo Larochelle. Efficient learning of deep boltzmann machines. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pages 693–700, 2010.
  33. 33.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pages 1096–1103. ACM, 2008.
  34. 34.Paul Rodriguez, Janet Wiles, and Jeffrey L Elman. A recurrent neural network that learns to count. Connection Science, 11(1):5–40, 1999.
  35. 35.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. Neural architectures for named entity recognition. arXiv preprint arXiv:1603.01360, 2016.
  36. 36.Jerone TA Andrews, Edward J Morton, and Lewis D Griffin. Detecting anomalous data using auto-encoders. International Journal of Machine Learning and Computing, 6(1):21, 2016a.
  37. 37.Sinno Jialin Pan, Qiang Yang, et al. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10):1345–1359, 2010.
  38. 38.Tolga Ergen, Ali Hassan Mirza, and Suleyman Serdar Kozat. Unsupervised and semi-supervised anomaly detection with lstm neural networks. arXiv preprint arXiv:1710.09207, 2017.
  39. 39.Lukas Ruff, Nico Gornitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Robert Vandermeulen, Alexander Binder, Emmanuel Muller, and Marius Kloft. Deep one-class classification. In International Conference on Machine Learning, pages 4390–4399, 2018a.
  40. 40.Yann LeCun, Corinna Cortes, and Christopher JC Burges. Mnist handwritten digit database. AT&T Labs [Online]. Available: http://yann. lecun. com/exdb/mnist, 2, 2010.
  41. 41.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, 2009.
  42. 42.Xiuyao Song, Mingxi Wu, Christopher Jermaine, and Sanjay Ranka. Conditional anomaly detection. IEEE Transactions on Knowledge and Data Engineering, 19(5):631–645, 2007.
  43. 43.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  44. 44.Min Du, Feifei Li, Guineng Zheng, and Vivek Srikumar. Deeplog: Anomaly detection and diagnosis from system logs through deep learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pages 1285–1298. ACM, 2017.
  45. 45.Michael A Hayes and Miriam AM Capretz. Contextual anomaly detection framework for big sensor data. Journal of Big Data, 2(1):2, 2015.
  46. 46.Raghavendra Chalapathy, Edward Toth, and Sanjay Chawla. Group anomaly detection using deep generative models. arXiv preprint arXiv:1804.04876, 2018b.
  47. 47.Loıc Bontemps, James McDermott, Nhien-An Le-Khac, et al. Collective anomaly detection based on long short-term memory recurrent neural networks. In International Conference on Future Data and Security Engineering, pages 141–152. Springer, 2016.

Citation

MLA
Chalapathy, R., and S. Chawla. “Deep Learning for Anomaly Detection: A Survey”. arXiv, 2019, http://arxiv.org/abs/1901.03407v2.
APA
Chalapathy, R., & Chawla, S. (2019). Deep Learning for Anomaly Detection: A Survey. arXiv. http://arxiv.org/abs/1901.03407v2
Chicago
Chalapathy, R., and S. Chawla. 2019. “Deep Learning for Anomaly Detection: A Survey”. arXiv. http://arxiv.org/abs/1901.03407v2.
Harvard
Chalapathy, R. and Chawla, S. (2019) “Deep Learning for Anomaly Detection: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1901.03407v2.
Vancouver
1. Chalapathy R, Chawla S (2019) Deep Learning for Anomaly Detection: A Survey. arXiv

BibTeX

@article{chalapathy2019deep,
  title = {Deep Learning for Anomaly Detection: A Survey},
  author = {Chalapathy, Raghavendra and Chawla, Sanjay},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1901.03407v2},
  eprint = {1901.03407}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors