Graph Neural Network-Based Anomaly Detection in Multivariate Time Series

Ailin DengBryan Hooi

article2021AAAI1,530 citations

Proposes a graph neural network method that explicitly learns relationship structures among multivariate time-series variables to deliver accurate anomaly detection alongside explainable root-cause diagnosis.

Listen

Modern industrial operations, transportation systems, and critical infrastructure increasingly depend on complex networks of interconnected sensors. These cyber-physical systems generate massive volumes of continuous time-series data, making manual monitoring impractical and leaving infrastructure vulnerable to operational faults and cyberattacks. Traditional and existing deep learning detection tools frequently struggle because they fail to explicitly capture the complex, nonlinear relationships among different sensors, which limits their ability to pinpoint root causes when anomalies occur.

The article demonstrates the effectiveness of a novel framework called the Graph Deviation Network, designed to automatically learn relationship structures across multivariate sensor networks. The core objective is to detect operational anomalies accurately while providing human operators with interpretable explanations of how and where system behaviors deviate from established baselines.

The proposed method operates in an unsupervised manner using normal historical data. It assigns high-dimensional embedding vectors to capture unique sensor characteristics, dynamically learns inter-sensor dependencies as a directed graph, and forecasts expected values using an attention mechanism over neighboring sensors. When observed values diverge from forecasts, the framework flags anomalies using a robust normalized deviation score. The researchers evaluated the approach against seven baseline methods on two physical water treatment testbeds: the 51-sensor Secure Water Treatment system and the 127-sensor Water Distribution system, both containing simulated real-world attack scenarios.

The evaluation produced several key findings. First, the proposed framework outperformed all baseline models in detection accuracy, achieving precision scores of approximately 99% on the first dataset and 98% on the second. Second, it delivered an F1-score—a metric balancing precision and recall—of 0.81 on the first system and 0.57 on the second, surpassing the next-best baseline by roughly 54% on the larger, more imbalanced network. Third, ablation analyses confirmed that learning graph structures, utilizing sensor embeddings, and applying graph attention mechanisms were all vital to performance, with the removal of attention causing the steepest accuracy drop. Finally, case studies demonstrated that the framework effectively localizes faulty sensors and explains anomalies by highlighting deviations between expected and observed behaviors among closely linked components.

These results demonstrate significant practical benefits for risk management, operational safety, and incident response timelines. Because the framework can automatically discover unknown relationships among hundreds of variables and identify specific deviating subgraphs, human operators can rapidly diagnose root causes rather than manually investigating vast sensor streams. This reduces downtime risks and strengthens defenses against subtle cyberattacks where individual sensor values remain within standard operating limits but violate relational patterns.

Organizations managing critical physical infrastructure should consider evaluating graph-based anomaly detection frameworks to monitor complex multivariate environments. Deployment strategies should leverage the framework's explainability features to assist security and operations teams during triage. However, because the study focused exclusively on stationary, offline training across water treatment testbed datasets, decision-makers should recognize that model performance under dynamic online operating conditions remains to be demonstrated. Further development should focus on adapting the framework for real-time online learning and testing across diverse industrial domains.

  • Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Surveys subsequent attention- and transformer-based architectures for multivariate time series anomaly detection and forecasting that extend beyond graph-based attention approaches.
  • Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). Develops a generalized 2D temporal variation backbone for unified time series tasks including anomaly detection, representing a major subsequent evolution in temporal modeling.
  • Paper: How Attentive are Graph Attention Networks?, Shaked Brody et al. (2021). Critiques the static ranking limitations of standard graph attention mechanisms and introduces dynamic graph attention, offering a direct improvement for attention-driven explainability.
Cover for Graph Neural Network-Based Anomaly Detection in Multivariate Time Series

Abstract

Given high-dimensional time series data (e.g., sensor data), how can we detect anomalous events, such as system faults and attacks? More challengingly, how can we do this in a way that captures complex inter-sensor relationships, and detects and explains anomalies which deviate from these relationships? Recently, deep learning approaches have enabled improvements in anomaly detection in high-dimensional datasets; however, existing methods do not explicitly learn the structure of existing relationships between variables, or use them to predict the expected behavior of time series. Our approach combines a structure learning approach with graph neural networks, additionally using attention weights to provide explainability for the detected anomalies. Experiments on two real-world sensor datasets with ground truth anomalies show that our method detects anomalies more accurately than baseline approaches, accurately captures correlations between sensors, and allows users to deduce the root cause of a detected anomaly.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Framework
  • 3.1 Problem Statement
  • 3.2 Overview
  • 3.3 Sensor Embedding
  • 3.4 Graph Structure Learning
  • 3.5 Graph Attention-Based Forecasting
  • 3.6 Graph Deviation Scoring
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Evaluation Metrics
  • 4.4 Experimental Setup
  • 4.5 RQ1. Accuracy
  • 4.6 RQ2. Ablation
  • 4.7 RQ3. Interpretability of Model
  • 4.8 RQ4. Localizing Anomalies
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Problem Formulation for Multivariate Time Series Anomaly Detection

    definition

    Let training data be denoted as strain=[strain(1),strain(2),…,strain(Ttrain)]\mathbf{s}_{\text{train}} = [\mathbf{s}_{\text{train}}^{(1)}, \mathbf{s}_{\text{train}}^{(2)}, \dots, \mathbf{s}_{\text{train}}^{(T_{\text{train}})}], where each observation strain(t)∈RN\mathbf{s}_{\text{train}}^{(t)} \in \mathbb{R}^N is an NN-dimensional vector representing the readings of NN continuous sensor variables at time tick t∈{1,…,Ttrain}t \in \{1, \dots, T_{\text{train}}\}. Under the unsupervised anomaly detection formulation, strain\mathbf{s}_{\text{train}} is assumed to contain exclusively normal operational data.

    The testing data originates from the identical NN sensors across a separate sequence of TtestT_{\text{test}} time ticks, denoted as stest=[stest(1),stest(2),…,stest(Ttest)]\mathbf{s}_{\text{test}} = [\mathbf{s}_{\text{test}}^{(1)}, \mathbf{s}_{\text{test}}^{(2)}, \dots, \mathbf{s}_{\text{test}}^{(T_{\text{test}})}] with stest(t)∈RN\mathbf{s}_{\text{test}}^{(t)} \in \mathbb{R}^N. The objective is to produce a sequence of binary anomaly labels a(t)∈{0,1}a(t) \in \{0, 1\} for each test time tick t∈{1,…,Ttest}t \in \{1, \dots, T_{\text{test}}\}, where a(t)=1a(t) = 1 indicates that the system is in an anomalous state at time tt, and a(t)=0a(t) = 0 indicates normal operation.

  2. Knowl 2 — Sensor Embedding and Directed Graph Structure Learning

    model/method

    In the Graph Deviation Network (GDN), each sensor i∈{1,…,N}i \in \{1, \dots, N\} is assigned a multidimensional embedding vector vi∈Rd\mathbf{v}_i \in \mathbb{R}^d, initialized randomly and optimized end-to-end along with the model parameters to capture sensor-specific characteristics.

    To represent inter-sensor relationship structures without requiring a predefined graph, GDN learns a directed adjacency matrix A∈{0,1}N×N\mathbf{A} \in \{0, 1\}^{N \times N}, where Aji=1A_{ji} = 1 denotes a directed edge from node jj to node ii (meaning sensor jj is used to forecast sensor ii). For each sensor ii, candidate relations are specified as Ci⊆{1,…,N}∖{i}C_i \subseteq \{1, \dots, N\} \setminus \{i\}. When no prior topological information is available, Ci={1,…,N}∖{i}C_i = \{1, \dots, N\} \setminus \{i\}.

    The similarity coefficient ejie_{ji} between sensor ii and candidate neighbor j∈Cij \in C_i is computed as the cosine similarity between their embedding vectors:

    eji=vi⊤vj∥vi∥2∥vj∥2for j∈Cie_{ji} = \frac{\mathbf{v}_i^\top \mathbf{v}_j}{\|\mathbf{v}_i\|_2 \|\mathbf{v}_j\|_2} \quad \text{for } j \in C_i

    The directed adjacency matrix is formed by selecting the top-kk candidate sensors with the highest similarity scores for each node ii:

    Aji=1{j∈TopK({eki:k∈Ci})}A_{ji} = \mathbf{1}\left\{j \in \text{TopK}\left(\{e_{ki} : k \in C_i\}\right)\right\}

    where TopK\text{TopK} returns the indices of the kk largest values among candidate similarity scores, with kk serving as a sparsity parameter.

  3. Knowl 3 — Graph Attention-Based Time Series Forecasting

    model/method

    GDN forecasts the expected sensor measurements s(t)∈RN\mathbf{s}^{(t)} \in \mathbb{R}^N at time tt from a historical sliding window of width ww, where the input for sensor ii is xi(t)=[si(t−w),si(t−w+1),…,si(t−1)]∈Rw\mathbf{x}_i^{(t)} = [s_i^{(t-w)}, s_i^{(t-w+1)}, \dots, s_i^{(t-1)}] \in \mathbb{R}^w.

    To fuse features across dependent sensors according to the learned adjacency matrix A\mathbf{A}, a graph attention mechanism conditions on both historical observations and sensor embeddings vi∈Rd\mathbf{v}_i \in \mathbb{R}^d. The joint representation for node ii is formed by concatenation:

    gi(t)=vi⊕Wxi(t)\mathbf{g}_i^{(t)} = \mathbf{v}_i \oplus \mathbf{W}\mathbf{x}_i^{(t)}

    where ⊕\oplus denotes concatenation and W∈Rd×w\mathbf{W} \in \mathbb{R}^{d \times w} is a shared trainable projection matrix. The unnormalized attention coefficient from node jj to node ii is:

    π(i,j)=LeakyReLU(a⊤(gi(t)⊕gj(t)))\pi(i, j) = \text{LeakyReLU}\left(\mathbf{a}^\top \left(\mathbf{g}_i^{(t)} \oplus \mathbf{g}_j^{(t)}\right)\right)

    where a∈R4d\mathbf{a} \in \mathbb{R}^{4d} is a trainable attention parameter vector. The normalized attention weights are calculated via softmax over neighbor set N(i)={j∣Aji>0}\mathcal{N}(i) = \{j \mid A_{ji} > 0\} and node ii itself:

    αi,j=exp⁡(π(i,j))∑k∈N(i)∪{i}exp⁡(π(i,k))\alpha_{i,j} = \frac{\exp(\pi(i, j))}{\sum_{k \in \mathcal{N}(i) \cup \{i\}} \exp(\pi(i, k))}

    The aggregated neighbor representation zi(t)∈Rd\mathbf{z}_i^{(t)} \in \mathbb{R}^d for sensor ii is:

    zi(t)=ReLU(αi,iWxi(t)+∑j∈N(i)αi,jWxj(t))\mathbf{z}_i^{(t)} = \text{ReLU}\left(\alpha_{i,i}\mathbf{W}\mathbf{x}_i^{(t)} + \sum_{j \in \mathcal{N}(i)} \alpha_{i,j}\mathbf{W}\mathbf{x}_j^{(t)}\right)

    The predicted sensor values s^(t)∈RN\hat{\mathbf{s}}^{(t)} \in \mathbb{R}^N are obtained by element-wise multiplying each representation with its embedding vector and applying stacked fully-connected layers fθf_\theta:

    s^(t)=fθ([v1∘z1(t),v2∘z2(t),…,vN∘zN(t)])\hat{\mathbf{s}}^{(t)} = f_\theta\left(\left[\mathbf{v}_1 \circ \mathbf{z}_1^{(t)}, \mathbf{v}_2 \circ \mathbf{z}_2^{(t)}, \dots, \mathbf{v}_N \circ \mathbf{z}_N^{(t)}\right]\right)

    where ∘\circ denotes the Hadamard product. Model parameters are trained by minimizing the Mean Squared Error (MSE) loss on normal training data:

    LMSE=1Ttrain−w∑t=w+1Ttrain∥s^(t)−s(t)∥22L_{\text{MSE}} = \frac{1}{T_{\text{train}} - w} \sum_{t=w+1}^{T_{\text{train}}} \|\hat{\mathbf{s}}^{(t)} - \mathbf{s}^{(t)}\|_2^2

  4. Knowl 4 — Graph Deviation Scoring and Anomaly Thresholding

    model/method

    To identify deviations between observed sensor behaviors and predicted expectations, GDN calculates an individual prediction error for each sensor i∈{1,…,N}i \in \{1, \dots, N\} at time tt:

    Erri(t)=∣si(t)−s^i(t)∣\text{Err}_i(t) = |s_i^{(t)} - \hat{s}_i^{(t)}|

    To prevent sensors with naturally large error scales from dominating the detection score, each sensor's error is normalized using its median μ~i\tilde{\mu}_i and inter-quartile range (IQR) σ~i\tilde{\sigma}_i computed across historical time ticks:

    ai(t)=Erri(t)−μ~iσ~ia_i(t) = \frac{\text{Err}_i(t) - \tilde{\mu}_i}{\tilde{\sigma}_i}

    where σ~i\tilde{\sigma}_i is the difference between the 75th percentile (Q3Q_3) and 25th percentile (Q1Q_1) of Erri\text{Err}_i. The system-level anomaly score at time tt aggregates across all sensors via the maximum operator:

    A(t)=max⁡i∈{1,…,N}ai(t)A(t) = \max_{i \in \{1, \dots, N\}} a_i(t)

    To dampen sudden transient spikes, a simple moving average (SMA) filter is applied to A(t)A(t) to yield smoothed anomaly scores As(t)A_s(t). A time tick tt is declared anomalous (a(t)=1a(t) = 1) if As(t)A_s(t) exceeds a detection threshold τ\tau, where τ\tau is set to the maximum smoothed score over the validation set:

    τ=max⁡t∈TvalAs(t)\tau = \max_{t \in \mathcal{T}_{\text{val}}} A_s(t)

  5. Knowl 5 — Experimental Setup and Cyber-Physical System Test-Beds

    experimental setup

    GDN was evaluated on two industrial cyber-physical test-bed datasets with ground-truth physical attack scenarios:

    1. SWaT (Secure Water Treatment): Comprises 51 continuous sensor and actuator features. The training split contains 47,515 normal time ticks, and the test split contains 44,986 time ticks with an anomaly rate of 11.97%.
    2. WADI (Water Distribution): A larger distribution network comprising 127 sensor and actuator features. The training split contains 118,795 normal time ticks, and the test split contains 17,275 time ticks with an anomaly rate of 5.99%.

    Both datasets provide approximately two weeks of normal physical operations for training, followed by several days of controlled physical attacks for evaluation. All measurements are downsampled to 10-second intervals using median pooling, assigning the most frequent binary label within each 10-second interval. The initial 2,160 samples (5–6 hours) from each dataset are removed to eliminate initial system stabilization transients.

    Models are trained with the Adam optimizer (learning rate 10−310^{-3}, β1=0.9,β2=0.99\beta_1 = 0.9, \beta_2 = 0.99) for up to 50 epochs with early stopping patience of 10. The sliding window size is w=5w = 5. Hyperparameters are configured based on input dimensionality: for SWaT (N=51N=51), embedding dimension d=64d=64, top-k=15k=15, and hidden layer size 6464; for WADI (N=127N=127), embedding dimension d=128d=128, top-k=30k=30, and hidden layer size 128128.

  6. Knowl 6 — Anomaly Detection Performance Comparison on SWaT and WADI

    empirical result

    GDN was evaluated against seven baseline methods on the SWaT and WADI test sets in terms of Precision (Prec, in %), Recall (Rec, in %), and F1-Score (F1):

    SWaT WADI
    Method Prec (%) Rec (%) F1 Prec (%) Rec (%) F1
    PCA 24.92 21.63 0.23 39.53 5.63 0.10
    KNN 7.83 7.83 0.08 7.76 7.75 0.08
    FB 10.17 10.17 0.10 8.60 8.60 0.09
    AE 72.63 52.63 0.61 34.35 34.35 0.34
    DAGMM 27.46 69.52 0.39 54.44 26.99 0.36
    LSTM-VAE 96.24 59.91 0.74 87.79 14.45 0.25
    MAD-GAN 98.97 63.74 0.77 41.44 33.92 0.37
    GDN 99.35 68.12 0.81 97.50 40.19 0.57

    GDN achieves the highest F1-score and precision on both datasets. On SWaT, GDN achieves an F1-score of 0.81 (Precision 99.35%, Recall 68.12%). On the more complex and unbalanced WADI dataset (127 sensors, 5.99% anomalies), GDN achieves an F1-score of 0.57, outperforming the best-performing baseline (MAD-GAN with F1 0.37) by 54% relative F1 improvement.

  7. Knowl 7 — Ablation Analysis of Graph Structure, Embeddings, and Attention

    empirical result

    An ablation study evaluated the relative contribution of each architectural component of GDN by comparing the full model against three stripped variants on SWaT and WADI:

    • -TOPK: Substitutes the learned sparse directed graph structure with a static, complete graph connecting each node to all other nodes.
    • -EMB: Removes sensor embeddings from the attention feature representation, setting gi=Wxi\mathbf{g}_i = \mathbf{W}\mathbf{x}_i.
    • -ATT: Replaces graph attention aggregation with uniform, unweighted averaging across all neighbors.
    SWaT WADI
    Method Prec (%) Rec (%) F1 Prec (%) Rec (%) F1
    GDN 99.35 68.12 0.81 97.50 40.19 0.57
    - TOPK 97.41 64.70 0.78 92.21 35.12 0.51
    - EMB 92.31 61.25 0.76 91.86 33.49 0.49
    - ATT 71.05 65.06 0.68 61.33 38.85 0.48

    The ablation results demonstrate:

    1. Learning a sparse graph via TopK selection improves performance over a dense complete graph, with a larger benefit on the higher-dimensional WADI dataset (F1 drops from 0.57 to 0.51 without TopK structure learning).
    2. Incorporating sensor embeddings into the attention mechanism increases F1 from 0.76 to 0.81 on SWaT and from 0.49 to 0.57 on WADI.
    3. Disabling the attention mechanism (-ATT) causes the largest performance degradation on both datasets (F1 drops to 0.68 on SWaT and 0.48 on WADI), showing that treating heterogeneous sensor neighbors with uniform weights introduces noise.
  8. Knowl 8 — Model Interpretability and Anomaly Root-Cause Localization

    model/method

    GDN provides interpretability and root-cause localization across four distinct levels:

    1. Sensor Behavior Clustering: Two-dimensional t-SNE projections of the trained sensor embeddings vi∈Rd\mathbf{v}_i \in \mathbb{R}^d show clustering of sensors that measure identical physical indicators or serve analogous functions across water distribution subsystems.
    2. Sensor Deviation Localization: The individual normalized anomaly score ai(t)=Erri(t)−μ~iσ~ia_i(t) = \frac{\text{Err}_i(t) - \tilde{\mu}_i}{\tilde{\sigma}_i} pinpoints the specific sensor exhibiting maximum deviation from expected relationships during an anomaly.
    3. Relational Context via Attention Weights: The learned attention coefficients αi,j\alpha_{i,j} from neighboring nodes j∈N(i)j \in \mathcal{N}(i) identify which correlated physical variables were expected to govern sensor ii's dynamics.
    4. Mechanism Explanation via Observed vs. Predicted Traces: By overlaying observed time series si(t)s_i^{(t)} against predicted behavior s^i(t)\hat{s}_i^{(t)}, GDN explains the nature of anomalous failures. For example, during a cyber-physical attack where flow sensor 1_FIT_001_PV in WADI is manipulated with in-range false readings, GDN identifies valve status sensor 1_MV_001_STATUS as the primary deviator: GDN predicts that 1_MV_001_STATUS should increase in tandem with 1_FIT_001_PV, and flags an anomaly when 1_MV_001_STATUS remains constant despite false increases in flow rate.

Coverage note — None was omitted; all key architectural components, mathematical formulations, experimental settings, empirical comparisons, ablation studies, and explainability mechanisms are captured in the knowls.

References

  1. 1.Aggarwal, C. C. 2015. Outlier analysis. In Data mining, 237–263. Springer.
  2. 2.Ahmed, C. M.; Palleti, V. R.; and Mathur, A. P. 2017. WADI: a water distribution testbed for research in the design of secure cyber physical systems. In Proceedings of the 3rd International Workshop on Cyber-Physical Systems for Smart Water Networks, 25–28.
  3. 3.Angiulli, F.; and Pizzuti, C. 2002. Fast outlier detection in high dimensional spaces. In European conference on principles of data mining and knowledge discovery, 15–27. Springer.
  4. 4.Bach, F. R.; and Jordan, M. I. 2004. Learning graphical models for stationary time series. IEEE transactions on signal processing 52(8): 2189–2199.
  5. 5.Blázquez-García, A.; Conde, A.; Mori, U.; and Lozano, J. A. 2020. A review on outlier/anomaly detection in time series data. arXiv preprint arXiv:2002.04236 .
  6. 6.Breunig, M. M.; Kriegel, H.-P.; Ng, R. T.; and Sander, J. 2000. LOF: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data, 93–104.
  7. 7.Chen, W.; Chen, L.; Xie, Y.; Cao, W.; Gao, Y.; and Feng, X. 2019. Multi-range attentive bicomponent graph convolutional network for traffic forecasting. arXiv preprint arXiv:1911.12093 .
  8. 8.Defferrard, M.; Bresson, X.; and Vandergheynst, P. 2016. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in neural information processing systems, 3844–3852.
  9. 9.Fey, M.; and Lenssen, J. E. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds.
  10. 10.Filonov, P.; Lavrentyev, A.; and Vorontsov, A. 2016. Multivariate industrial time series with cyber-attack simulation: Fault detection using an lstm-based predictive data model. arXiv preprint arXiv:1612.06676 .
  11. 11.Goh, J.; Adepu, S.; Junejo, K. N.; and Mathur, A. 2016. A dataset to support research in the design of secure water treatment systems. In International conference on critical information infrastructures security, 88–99. Springer.
  12. 12.Hautamaki, V.; Karkkainen, I.; and Franti, P. 2004. Outlier detection using k-nearest neighbour graph. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004., volume 3, 430–433. IEEE.
  13. 13.Hundman, K.; Constantinou, V.; Laporte, C.; Colwell, I.; and Soderstrom, T. 2018a. Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, 387–395.
  14. 14.Hundman, K.; Constantinou, V.; Laporte, C.; Colwell, I.; and Soderstrom, T. 2018b. Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, 387–395.
  15. 15.Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 .
  16. 16.Kipf, T. N.; and Welling, M. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 .
  17. 17.Kobourov, S. G. 2012. Spring embedders and force directed graph drawing algorithms. arXiv preprint arXiv:1201.3011 .
  18. 18.Lazarevic, A.; and Kumar, V. 2005. Feature bagging for outlier detection. In Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining, 157–166.
  19. 19.Li, D.; Chen, D.; Jin, B.; Shi, L.; Goh, J.; and Ng, S.-K. 2019. MAD-GAN: Multivariate anomaly detection for time series data with generative adversarial networks. In International Conference on Artificial Neural Networks, 703–716. Springer.
  20. 20.Lim, N.; Hooi, B.; Ng, S.-K.; Wang, X.; Goh, Y. L.; Weng, R.; and Varadarajan, J. 2020. STP-UDGAT: Spatial-Temporal-Preference User Dimensional Graph Attention Network for Next POI Recommendation. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 845–854.
  21. 21.Maaten, L. v. d.; and Hinton, G. 2008. Visualizing data using t-SNE. Journal of machine learning research 9(Nov): 2579–2605.
  22. 22.Mathur, A. P.; and Tippenhauer, N. O. 2016. SWaT: a water treatment testbed for research and training on ICS security. In 2016 International Workshop on Cyber-physical Systems for Smart Water Networks (CySWater), 31–36. IEEE.
  23. 23.Munir, M.; Siddiqui, S. A.; Dengel, A.; and Ahmed, S. 2018. DeepAnT: A deep learning approach for unsupervised anomaly detection in time series. IEEE Access 7: 1991–2005.
  24. 24.Park, D.; Hoshi, Y.; and Kemp, C. C. 2018. A multimodal anomaly detector for robot-assisted feeding using an lstm-based variational autoencoder. IEEE Robotics and Automation Letters 3(3): 1544–1551.
  25. 25.Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017. Automatic differentiation in PyTorch. In NIPS-W.
  26. 26.Qin, Y.; Song, D.; Chen, H.; Cheng, W.; Jiang, G.; and Cottrell, G. 2017. A dual-stage attention-based recurrent neural network for time series prediction. arXiv preprint arXiv:1704.02971 .
  27. 27.Schlichtkrull, M.; Kipf, T. N.; Bloem, P.; Van Den Berg, R.; Titov, I.; and Welling, M. 2018. Modeling relational data with graph convolutional networks. In European Semantic Web Conference, 593–607. Springer.
  28. 28.Schölkopf, B.; Platt, J. C.; Shawe-Taylor, J.; Smola, A. J.; and Williamson, R. C. 2001. Estimating the support of a high-dimensional distribution. Neural computation 13(7): 1443–1471.
  29. 29.Shyu, M.-L.; Chen, S.-C.; Sarinnapakorn, K.; and Chang, L. 2003. A novel anomaly detection scheme based on principal component classifier. Technical report, MIAMI UNIV CORAL GABLES FL DEPT OF ELECTRICAL AND COMPUTER ENGINEERING.
  30. 30.Siffer, A.; Fouque, P.-A.; Termier, A.; and Largouet, C. 2017. Anomaly detection in streams with extreme value theory. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1067–1075.
  31. 31.Tank, A.; Foti, N.; and Fox, E. 2015. Bayesian structure learning for stationary time series. arXiv preprint arXiv:1505.03131 .
  32. 32.Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; and Bengio, Y. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 .
  33. 33.Wang, Y.; Wang, W.; Ca, Y.; Hooi, B.; and Ooi, B. C. 2020. Detecting Implementation Bugs in Graph Convolutional Network based Node Classifiers. In 2020 IEEE 31st International Symposium on Software Reliability Engineering (ISSRE), 313–324. IEEE.
  34. 34.Yu, B.; Yin, H.; and Zhu, Z. 2017. Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting. arXiv preprint arXiv:1709.04875 .
  35. 35.Zhang, Y.; Hamm, N. A.; Meratnia, N.; Stein, A.; Van De Voort, M.; and Havinga, P. J. 2012. Statistics-based outlier detection for wireless sensor networks. International Journal of Geographical Information Science 26(8): 1373–1392.
  36. 36.Zhou, B.; Liu, S.; Hooi, B.; Cheng, X.; and Ye, J. 2019. BeatGAN: Anomalous Rhythm Detection using Adversarially Generated Time Series. In IJCAI, 4433–4439.
  37. 37.Zhou, Y.; Qin, R.; Xu, H.; Sadiq, S.; and Yu, Y. 2018. A data quality control method for seafloor observatories: the application of observed time series data in the East China Sea. Sensors 18(8): 2628.
  38. 38.Zong, B.; Song, Q.; Min, M. R.; Cheng, W.; Lumezanu, C.; Cho, D.; and Chen, H. 2018. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In International Conference on Learning Representations.

Citation

MLA
Deng, A., and B. Hooi. “Graph Neural Network-Based Anomaly Detection in Multivariate Time Series”. arXiv, 2021, http://arxiv.org/abs/2106.06947v1.
APA
Deng, A., & Hooi, B. (2021). Graph Neural Network-Based Anomaly Detection in Multivariate Time Series. arXiv. http://arxiv.org/abs/2106.06947v1
Chicago
Deng, A., and B. Hooi. 2021. “Graph Neural Network-Based Anomaly Detection in Multivariate Time Series”. arXiv. http://arxiv.org/abs/2106.06947v1.
Harvard
Deng, A. and Hooi, B. (2021) “Graph Neural Network-Based Anomaly Detection in Multivariate Time Series”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.06947v1.
Vancouver
1. Deng A, Hooi B (2021) Graph Neural Network-Based Anomaly Detection in Multivariate Time Series. arXiv

BibTeX

@article{deng2021graph,
  title = {Graph Neural Network-Based Anomaly Detection in Multivariate Time Series},
  author = {Deng, Ailin and Hooi, Bryan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.06947v1},
  eprint = {2106.06947}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/