Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy

Jiehui XuHaixu WuJianmin WangMingsheng Long

article2022ICLR1,219 citations

Proposes the Anomaly Transformer, which exploits attention-weight discrepancies between local and global temporal associations through a minimax optimization strategy to achieve state-of-the-art unsupervised time series anomaly detection.

Listen

Modern industrial and technological operations rely heavily on continuous sensor measurements to monitor critical infrastructure, server metrics, spacecraft, and water treatment systems. Identifying malfunctions from these large-scale time series is essential for operational security and avoiding severe financial loss. However, anomalies are rare and buried within massive volumes of normal data, making manual labeling impractical and expensive. Traditional unsupervised detection approaches rely on point-by-point reconstruction errors or density estimation, which frequently miss intricate temporal dynamics and produce confusing, noisy anomaly scores.

The article demonstrates an unsupervised deep learning framework, named the Anomaly Transformer, designed to reliably identify anomalies in multivariate time series without labeled training data. It evaluates how modeling relational associations across time points, rather than isolated point values, establishes an effective criterion for distinguishing abnormal events from normal patterns.

The authors develop a novel attention mechanism that contrasts two perspectives: a baseline assumption that abnormal events primarily correlate with their immediate adjacent time points, versus learned dependencies captured across the entire sequence. By applying an adversarial training strategy, the model actively maximizes the divergence between these two views for normal points while constraining it for rare anomalies. The evaluation encompasses six standard benchmark datasets spanning IT server monitoring, space rover telemetry, water treatment plants, and synthetic anomaly benchmarks, benchmarking performance against eighteen established detection models.

The findings demonstrate substantial performance gains. The proposed framework achieved state-of-the-art results across all benchmark datasets, reaching an average F1-score of 94.96% across five real-world domains and outperforming the previous leading method by roughly 7 percentage points. Replacing standard point reconstruction metrics with the proposed association-based score delivered an absolute performance improvement of nearly 19 percentage points. The method proved robust across diverse fault types, including localized spikes, seasonal shifts, and trend changes, while also demonstrating the ability to detect emerging equipment malfunctions at an early operational stage.

These results demonstrate that association-based monitoring significantly reduces false alarm rates while maintaining high sensitivity, lowering operational risk and monitoring fatigue. The ability to identify anomalies early allows engineering teams to intervene before equipment failures cause service outages, safety incidents, or financial damage. These findings challenge the standard industry reliance on simple point-reconstruction error thresholds by showing that relational context is far more informative.

Organizations operating continuous sensor networks should consider piloting association-based detection architectures for complex multivariate monitoring pipelines. Decision-makers must evaluate computational resource trade-offs, as longer temporal observation windows improve detection fidelity but increase hardware memory requirements. Operational teams should choose detection thresholds aligned with available investigation capacity, using targeted anomaly proportion settings on validation data.

While the empirical results demonstrate strong reliability across diverse domains, the framework's primary limitations stem from the quadratic computational complexity inherent to sequence-level attention mechanisms and the empirical nature of the deep architecture. Confidence in the reported performance is high across standard benchmark conditions, though practical deployments should validate window sizing and resource allocations during initial integration.

arXiv: 2110.02642
Cover for Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy

Abstract

Unsupervised detection of anomaly points in time series is a challenging problem, which requires the model to derive a distinguishable criterion. Previous methods tackle the problem mainly through learning pointwise representation or pairwise association, however, neither is sufficient to reason about the intricate dynamics. Recently, Transformers have shown great power in unified modeling of pointwise representation and pairwise association, and we find that the self-attention weight distribution of each time point can embody rich association with the whole series. Our key observation is that due to the rarity of anomalies, it is extremely difficult to build nontrivial associations from abnormal points to the whole series, thereby, the anomalies' associations shall mainly concentrate on their adjacent time points. This adjacent-concentration bias implies an association-based criterion inherently distinguishable between normal and abnormal points, which we highlight through the \emph{Association Discrepancy}. Technically, we propose the \emph{Anomaly Transformer} with a new \emph{Anomaly-Attention} mechanism to compute the association discrepancy. A minimax strategy is devised to amplify the normal-abnormal distinguishability of the association discrepancy. The Anomaly Transformer achieves state-of-the-art results on six unsupervised time series anomaly detection benchmarks of three applications: service monitoring, space & earth exploration, and water treatment.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Unsupervised Time Series Anomaly Detection
  • 2.2 Transformers for Time Series Analysis
  • 3 Method
  • 3.1 Anomaly Transformer
  • 3.2 Minimax Association Learning
  • 4 Experiments
  • 4.1 Main Results
  • 4.2 Model Analysis
  • 5 Conclusion and Future Work
  • References
  • A Parameter Sensitivity
  • B Implementation Details
  • C More Showcases
  • D Ablation of Association Discrepancy
  • D.1 Ablation of Multi-Level Quantification
  • D.2 Ablation of Statistical Distance
  • D.3 Ablation of Prior-Association
  • E Ablation of Association-based Criterion
  • E.1 Calculation
  • E.2 Ablation of Criterion Definition
  • F Convergence of Minimax Optimization
  • G Model Parameter Sensitivity
  • H Protocol of Threshold Selection
  • I More Baselines
  • J Limitations and Future Work
  • K Dataset
  • L UCR Dataset

Knowls

  1. Knowl 1 — Anomaly Transformer Architecture and Anomaly-Attention

    model/method

    The Anomaly Transformer is a deep architecture designed for unsupervised time series anomaly detection. Given an input multivariate time series X∈RN×d\mathcal{X} \in \mathbb{R}^{N \times d} with NN time steps and dd variables, the sequence is first mapped to X0=Embedding(X)∈RN×dmodel\mathcal{X}^0 = \text{Embedding}(\mathcal{X}) \in \mathbb{R}^{N \times d_{\text{model}}}. The network stacks LL layers alternating between Anomaly-Attention blocks and feed-forward networks:

    Zl=Layer-Norm(Anomaly-Attention(Xl−1)+Xl−1)Z^l = \text{Layer-Norm}(\text{Anomaly-Attention}(\mathcal{X}^{l-1}) + \mathcal{X}^{l-1}) Xl=Layer-Norm(Feed-Forward(Zl)+Zl)\mathcal{X}^l = \text{Layer-Norm}(\text{Feed-Forward}(Z^l) + Z^l)

    where Xl∈RN×dmodel\mathcal{X}^l \in \mathbb{R}^{N \times d_{\text{model}}} and Zl∈RN×dmodelZ^l \in \mathbb{R}^{N \times d_{\text{model}}} represent hidden states at layer l∈{1,…,L}l \in \{1, \dots, L\}.

    The Anomaly-Attention mechanism uses a two-branch architecture to compute two distinct association distributions for each time point:

    1. Prior-Association Pl∈RN×NP^l \in \mathbb{R}^{N \times N}: Models the adjacent-concentration inductive bias using a Gaussian kernel parameterized by a learnable scale vector σ=Xl−1Wσl∈RN×1\sigma = \mathcal{X}^{l-1} W_\sigma^l \in \mathbb{R}^{N \times 1}:

    Pi,jl=1∑k=1NG(∣k−i∣;σi)G(∣j−i∣;σi),G(∣j−i∣;σi)=12πσiexp⁡(−∣j−i∣22σi2)P_{i,j}^l = \frac{1}{\sum_{k=1}^N G(|k - i|; \sigma_i)} G(|j - i|; \sigma_i), \quad G(|j - i|; \sigma_i) = \frac{1}{\sqrt{2\pi}\sigma_i} \exp\left(-\frac{|j - i|^2}{2\sigma_i^2}\right)

    1. Series-Association Sl∈RN×NS^l \in \mathbb{R}^{N \times N}: Learns dependencies directly from data via self-attention with parameter matrices WQl,WKl,WVl∈Rdmodel×dmodelW_Q^l, W_K^l, W_V^l \in \mathbb{R}^{d_{\text{model}} \times d_{\text{model}}}:

    Q=Xl−1WQl,K=Xl−1WKl,V=Xl−1WVlQ = \mathcal{X}^{l-1} W_Q^l, \quad K = \mathcal{X}^{l-1} W_K^l, \quad V = \mathcal{X}^{l-1} W_V^l Sl=Softmax(QKTdmodel)S^l = \text{Softmax}\left(\frac{Q K^T}{\sqrt{d_{\text{model}}}}\right)

    1. Reconstruction Representation: The attention output hidden state is calculated via Z^l=SlV\widehat{Z}^l = S^l V.
  2. Knowl 2 — Association Discrepancy

    definition

    For an input time series X∈RN×d\mathcal{X} \in \mathbb{R}^{N \times d}, let Pl∈RN×NP^l \in \mathbb{R}^{N \times N} and Sl∈RN×NS^l \in \mathbb{R}^{N \times N} denote the prior-association and series-association matrices at layer l∈{1,…,L}l \in \{1, \dots, L\}, respectively. For the ii-th time step, the ii-th rows Pi,:lP_{i,:}^l and Si,:lS_{i,:}^l form discrete probability distributions over all time points j∈{1,…,N}j \in \{1, \dots, N\}.

    The point-wise Association Discrepancy AssDis(P,S;X)∈RN×1\text{AssDis}(P, S; \mathcal{X}) \in \mathbb{R}^{N \times 1} is defined as the multi-layer average of the symmetrized Kullback-Leibler (KL) divergence between prior- and series-associations:

    AssDis(P,S;X)=[1L∑l=1L(KL(Pi,:l∥Si,:l)+KL(Si,:l∥Pi,:l))]i=1,…,N\text{AssDis}(P, S; \mathcal{X}) = \left[ \frac{1}{L} \sum_{l=1}^L \left( \text{KL}(P_{i,:}^l \parallel S_{i,:}^l) + \text{KL}(S_{i,:}^l \parallel P_{i,:}^l) \right) \right]_{i=1, \dots, N}

    where KL(p∥q)=∑j=1Npjlog⁡(pj/qj)\text{KL}(p \parallel q) = \sum_{j=1}^N p_j \log(p_j / q_j). Normal time points exhibit broad, non-adjacent temporal associations that diverge strongly from the Gaussian prior distribution, resulting in a large AssDis\text{AssDis}. Abnormal points are dominated by continuity with adjacent points and fail to build strong associations across the entire sequence, yielding a significantly smaller AssDis\text{AssDis}.

  3. Knowl 3 — Minimax Association Learning Strategy

    model/method

    The Anomaly Transformer is trained with an unsupervised objective combining a reconstruction loss and an association discrepancy term. The full objective for input X∈RN×d\mathcal{X} \in \mathbb{R}^{N \times d} and reconstruction X^∈RN×d\widehat{\mathcal{X}} \in \mathbb{R}^{N \times d} is:

    LTotal(X^,P,S,λ;X)=∥X−X^∥F2−λ∥AssDis(P,S;X)∥1\mathcal{L}_{\text{Total}}(\widehat{\mathcal{X}}, P, S, \lambda; \mathcal{X}) = \|\mathcal{X} - \widehat{\mathcal{X}}\|_F^2 - \lambda \|\text{AssDis}(P, S; \mathcal{X})\|_1

    where ∥⋅∥F\|\cdot\|_F is the Frobenius norm, ∥⋅∥1\|\cdot\|_1 is the L1L_1 norm across temporal steps, and λ>0\lambda > 0 balances the two objectives.

    To prevent the Gaussian scale parameter σ\sigma from degenerating to zero when directly maximizing discrepancy, training employs a two-phase minimax optimization with stop-gradient operators (detach\text{detach}):

    1. Minimize Phase: min⁡PLTotal(X^,P,Sdetach,−λ;X)\min_P \mathcal{L}_{\text{Total}}(\widehat{\mathcal{X}}, P, S_{\text{detach}}, -\lambda; \mathcal{X}) This drives the prior-association PP to fit the data-driven series-association SS, forcing the learnable scale parameter σ\sigma to adapt to local patterns.

    2. Maximize Phase: max⁡SLTotal(X^,Pdetach,S,λ;X)\max_S \mathcal{L}_{\text{Total}}(\widehat{\mathcal{X}}, P_{\text{detach}}, S, \lambda; \mathcal{X}) This forces the series-association SS to attend to non-adjacent temporal regions while minimizing reconstruction error. Normal points achieve non-adjacent associations readily, whereas anomalies cannot, amplifying the discrepancy contrast between normal and abnormal time points.

  4. Knowl 4 — Association-Based Anomaly Scoring Criterion

    equation

    Given an input time series X∈RN×d\mathcal{X} \in \mathbb{R}^{N \times d}, its reconstruction X^∈RN×d\widehat{\mathcal{X}} \in \mathbb{R}^{N \times d}, and the point-wise association discrepancy vector AssDis(P,S;X)∈RN×1\text{AssDis}(P, S; \mathcal{X}) \in \mathbb{R}^{N \times 1}, the point-wise anomaly score AnomalyScore(X)∈RN×1\text{AnomalyScore}(\mathcal{X}) \in \mathbb{R}^{N \times 1} is defined as:

    AnomalyScore(X)=Softmax(−AssDis(P,S;X))⊙[∥Xi,:−X^i,:∥22]i=1,…,N\text{AnomalyScore}(\mathcal{X}) = \text{Softmax}(-\text{AssDis}(P, S; \mathcal{X})) \odot \left[ \|\mathcal{X}_{i,:} - \widehat{\mathcal{X}}_{i,:}\|_2^2 \right]_{i=1, \dots, N}

    where ⊙\odot denotes Hadamard (element-wise) multiplication, Softmax(⋅)\text{Softmax}(\cdot) normalizes along the temporal dimension NN, and ∥Xi,:−X^i,:∥22\|\mathcal{X}_{i,:} - \widehat{\mathcal{X}}_{i,:}\|_2^2 is the squared Euclidean reconstruction error at time index ii.

    Because anomalies exhibit smaller association discrepancies (thus higher values after Softmax(−AssDis)\text{Softmax}(-\text{AssDis})) and higher reconstruction errors, the product of these two metrics yields high contrast between normal and abnormal time points.

  5. Knowl 5 — Multi-Head Anomaly-Attention Algorithm

    algorithm

    The multi-head Anomaly-Attention computes the prior-association distributions, series-association distributions, and the updated hidden representations across hh attention heads.

    Input: Input representation X∈RN×dmodelX \in \mathbb{R}^{N \times d_{\text{model}}}, relative distance matrix D=[(j−i)2]i,j∈{1,…,N}∈RN×ND = [(j - i)^2]_{i,j \in \{1, \dots, N\}} \in \mathbb{R}^{N \times N}
    Layer parameters: Linear projectors MLPinput\text{MLP}_{\text{input}} and MLPoutput\text{MLP}_{\text{output}}, number of heads hh
    Output: Layer hidden state Z^∈RN×dmodel\widehat{Z} \in \mathbb{R}^{N \times d_{\text{model}}}, prior associations PmP_m, series associations SmS_m for $m \in \{1, \dots, h\}
    Q,K,V,σ=Split(MLPinput(X),dim=1)Q, K, V, \sigma = \text{Split}(\text{MLP}_{\text{input}}(X), \text{dim}=1) where Q,K,V∈RN×dmodel,σ∈RN×hQ, K, V \in \mathbb{R}^{N \times d_{\text{model}}}, \sigma \in \mathbb{R}^{N \times h}
    for each head m∈{1,…,h}m \in \{1, \dots, h\} do
        Extract Qm,Km,Vm∈RN×(dmodel/h)Q_m, K_m, V_m \in \mathbb{R}^{N \times (d_{\text{model}} / h)} and σm∈RN×1\sigma_m \in \mathbb{R}^{N \times 1}
        σm=Broadcast(σm,dim=1)∈RN×N\sigma_m = \text{Broadcast}(\sigma_m, \text{dim}=1) \in \mathbb{R}^{N \times N}
        Pm=12πσmexp⁡(−D2σm2)∈RN×NP_m = \frac{1}{\sqrt{2\pi}\sigma_m} \exp\left(-\frac{D}{2\sigma_m^2}\right) \in \mathbb{R}^{N \times N}
        Pm=Pm/Broadcast(Sum(Pm,dim=1))∈RN×NP_m = P_m / \text{Broadcast}(\text{Sum}(P_m, \text{dim}=1)) \in \mathbb{R}^{N \times N}
        Sm=Softmax(hdmodelQmKmT)∈RN×NS_m = \text{Softmax}\left(\sqrt{\frac{h}{d_{\text{model}}}} Q_m K_m^T\right) \in \mathbb{R}^{N \times N}
        Z^m=SmVm∈RN×(dmodel/h)\widehat{Z}_m = S_m V_m \in \mathbb{R}^{N \times (d_{\text{model}} / h)}
    end for
    Z^=MLPoutput(Concat([Z^1,…,Z^h],dim=1))∈RN×dmodel\widehat{Z} = \text{MLP}_{\text{output}}(\text{Concat}([\widehat{Z}_1, \dots, \widehat{Z}_h], \text{dim}=1)) \in \mathbb{R}^{N \times d_{\text{model}}}
    return Z^\widehat{Z}, {Pm}m=1h\{P_m\}_{m=1}^h, {Sm}m=1h\{S_m\}_{m=1}^h
  6. Knowl 6 — Benchmark Anomaly Detection Performance

    data/table

    Anomaly Transformer was evaluated on five real-world multivariate time series anomaly detection benchmarks: Server Machine Dataset (SMD), Mars Science Laboratory rover (MSL), Soil Moisture Active Passive satellite (SMAP), Secure Water Treatment (SWaT), and Pooled Server Metrics (PSM). Performance is measured by precision (PP), recall (RR), and F1F_1-score (F1F_1) in percentage under the standard point-adjustment evaluation protocol.

    Dataset SMD MSL SMAP SWaT PSM
    Metric P R F1 P R F1 P R F1 P R F1 P R F1
    OCSVM 44.34 76.72 56.19 59.78 86.87 70.82 53.85 59.07 56.34 45.39 49.22 47.23 62.75 80.89 70.67
    IsolationForest 42.31 73.29 53.64 53.94 86.54 66.45 52.39 59.07 55.53 49.29 44.95 47.02 76.09 92.45 83.48
    LOF 56.34 39.86 46.68 47.72 85.25 61.18 58.93 56.33 57.60 72.15 65.43 68.62 57.89 90.49 70.61
    Deep-SVDD 78.54 79.67 79.10 91.92 76.63 83.58 89.93 56.02 69.04 80.42 84.45 82.39 95.41 86.49 90.73
    DAGMM 67.30 49.89 57.30 89.60 63.93 74.62 86.45 56.73 68.51 89.92 57.84 70.40 93.49 70.03 80.08
    MMPCACD 71.20 79.28 75.02 81.42 61.31 69.95 88.61 75.84 81.73 82.52 68.29 74.73 76.26 78.35 77.29
    VAR 78.35 70.26 74.08 74.68 81.42 77.90 81.38 53.88 64.83 81.59 60.29 69.34 90.71 83.82 87.13
    LSTM 78.55 85.28 81.78 85.45 82.50 83.95 89.41 78.13 83.39 86.15 83.27 84.69 76.93 89.64 82.80
    CL-MPPCA 82.36 76.07 79.09 73.71 88.54 80.44 86.13 63.16 72.88 76.78 81.50 79.07 56.02 99.93 71.80
    ITAD 86.22 73.71 79.48 69.44 84.09 76.07 82.42 66.89 73.85 63.13 52.08 57.08 72.80 64.02 68.13
    LSTM-VAE 75.76 90.08 82.30 85.49 79.94 82.62 92.20 67.75 78.10 76.00 89.50 82.20 73.62 89.92 80.96
    BeatGAN 72.90 84.09 78.10 89.75 85.42 87.53 92.38 55.85 69.61 64.01 87.46 73.92 90.30 93.84 92.04
    OmniAnomaly 83.68 86.82 85.22 89.02 86.37 87.67 92.49 81.99 86.92 81.42 84.30 82.83 88.39 74.46 80.83
    InterFusion 87.02 85.43 86.22 81.28 92.70 86.62 89.77 88.52 89.14 80.59 85.58 83.01 83.61 83.45 83.52
    THOC 79.76 90.95 84.99 88.45 90.97 89.69 92.06 89.34 90.68 83.94 86.36 85.13 88.14 90.99 89.54
    Ours 89.40 95.45 92.33 92.09 95.15 93.59 94.13 99.40 96.69 91.55 96.73 94.07 96.91 98.90 97.89

    Anomaly Transformer achieves state-of-the-art F1F_1-scores across all five datasets, outperforming the previous state-of-the-art methods InterFusion and THOC by 3.90%3.90\% to 8.35%8.35\% absolute F1F_1 percentage points.

  7. Knowl 7 — Ablation Study on Anomaly Criterion, Prior Parameterization, and Optimization

    data/table

    Ablation experiments evaluate the contribution of the anomaly criterion (pure reconstruction vs. pure association discrepancy vs. combination), the learnable scale parameter σ\sigma in prior-association, and the minimax optimization strategy across five benchmarks.

    Architecture Anomaly Prior- Optimization SMD MSL SMAP SWaT PSM Avg F1
    Criterion Association Strategy (%)
    Transformer Recon ×\times ×\times 79.72 76.64 73.74 74.56 78.43 76.62
    Anomaly Recon Learnable Minmax 71.35 78.61 69.12 81.53 80.40 76.20
    Transformer AssDis Learnable Minmax 87.57 90.50 90.98 93.21 95.47 91.55
    Assoc Fix (σ=1.0\sigma=1.0) Max 83.95 82.17 70.65 79.46 79.04 79.05
    Assoc Learnable Max 88.88 85.20 87.84 81.65 93.83 87.48
    Final Assoc Learnable Minmax 92.33 93.59 96.90 94.07 97.89 94.96

    The association-based criterion (Assoc) yields an 18.76%18.76\% average absolute F1F_1 gain over pure reconstruction (76.20%→94.96%76.20\% \to 94.96\%). A learnable scale σ\sigma improves F1F_1 by 8.43%8.43\% over a fixed σ=1.0\sigma=1.0 (79.05%→87.48%79.05\% \to 87.48\%), and the minimax strategy contributes an additional 7.48%7.48\% over direct maximization (87.48%→94.96%87.48\% \to 94.96\%).

  8. Knowl 8 — Unsupervised Anomaly Detection Protocol and Threshold Selection

    experimental setup

    The experimental protocol processes time series by partitioning sequences into non-overlapping sliding sub-series of fixed length N=100N = 100.

    Hyperparameters are set to L=3L = 3 layers, hidden state dimension dmodel=512d_{\text{model}} = 512, h=8h = 8 attention heads, loss tradeoff coefficient λ=3\lambda = 3, batch size of 32, and Adam optimizer with an initial learning rate of 10−410^{-4} trained for at most 10 epochs.

    Threshold selection is conducted in an unsupervised manner on the unlabeled validation set using the anomaly ratio parameter rr:

    1. After training, point-wise anomaly scores are computed on the validation subset.
    2. The threshold δ\delta is selected so that the top rr proportion of validation points with the highest anomaly scores are categorized as anomalous (r=0.1%r = 0.1\% for SWaT, r=0.5%r = 0.5\% for SMD, and r=1.0%r = 1.0\% for MSL, SMAP, and PSM; or equivalently, fixed thresholds δ=0.1\delta = 0.1 for SMD/MSL/SWaT and δ=0.01\delta = 0.01 for SMAP/PSM).
    3. Standard point-adjustment evaluation is applied: if any single point in a contiguous ground-truth anomaly segment is detected, all points within that segment are considered correctly detected anomalies.
  9. Knowl 9 — Statistical Distance and Prior Kernel Formulations

    empirical result

    Comparison across alternative statistical distance metrics and prior distribution kernels confirms the superiority of the symmetrized KL divergence and Gaussian kernel:

    1. Statistical Distance Metrics: Symmetrized KL divergence achieves an average F1F_1 of 94.96%94.96\%, outperforming Jensen-Shannon Divergence (91.80%91.80\% MSL, 89.78%89.78\% SWaT), Cross-Entropy (88.22%88.22\% MSL, 70.93%70.93\% SWaT), Wasserstein Distance (45.58%45.58\% MSL, 80.55%80.55\% SWaT), and L2L_2 distance (83.39%83.39\% MSL, 83.51%83.51\% SWaT). L2L_2 fails by ignoring the properties of discrete probability distributions, while Wasserstein distance adds noise because points are already index-aligned.
    2. Prior Distribution Kernels: The learnable Gaussian kernel consistently surpasses the power-law kernel P(x;α)=x−αP(x; \alpha) = x^{-\alpha} (F1F_1 of 92.33%92.33\% vs. 90.91%90.91\% on SMD, 96.69%96.69\% vs. 71.31%71.31\% on SMAP) because the Gaussian scale σ\sigma is significantly easier and more stable to optimize via minimax training than the power exponent α\alpha.
  10. Knowl 10 — Computational and Window Size Limitations

    limitation

    The Anomaly Transformer has two primary limitations:

    1. Attention Complexity and Window Size Dependency: The self-attention mechanism incurs quadratic computational and memory complexity O(N2)O(N^2) with respect to the sliding window size NN. While decreasing NN reduces resource consumption, overly small window sizes degrade the model's ability to learn global series-associations, creating a practical trade-off between computational efficiency and detection accuracy.
    2. Theoretical Guarantees: As an empirical deep architecture, the exact theoretical behavior and convergence properties of Transformer attention maps under anomaly-induced distribution shifts remain under-explored relative to classical linear autoregression (VAR) and state-space models.

Coverage note — None was omitted; all key contributions including the architecture, mathematical definitions, training algorithms, empirical results, ablations, protocols, and limitations are covered.

References

  1. 1.Ahmed Abdulaal, Zhuanghua Liu, and Tomer Lancewicki. Practical approach to asynchronous multivariate time series anomaly detection and localization. KDD, 2021.
  2. 2.Ryan Prescott Adams and David J. C. MacKay. Bayesian online changepoint detection. arXiv preprint arXiv:0710.3742, 2007.
  3. 3.O. Anderson and M. Kendall. Time-series. 2nd edn. J. R. Stat. Soc. (Series D), 1976.
  4. 4.Paul Boniol and Themis Palpanas. Series2graph: Graph-based subsequence anomaly detection for time series. Proc. VLDB Endow., 2020.
  5. 5.Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, and Jörg Sander. LOF: identifying density-based local outliers. In SIGMOD, 2000.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020.
  7. 7.Zekai Chen, Dingshuo Chen, Zixuan Yuan, Xiuzhen Cheng, and Xiao Zhang. Learning graph structures with transformer for multivariate time series anomaly detection in iot. ArXiv, abs/2104.03466, 2021.
  8. 8.Haibin Cheng, Pang-Ning Tan, Christopher Potter, and Steven A. Klooster. A robust graph-based algorithm for detection and characterization of anomalies in noisy multivariate time series. ICDM Workshops, 2008.
  9. 9.Haibin Cheng, Pang-Ning Tan, Christopher Potter, and Steven A. Klooster. Detection and characterization of anomalies in multivariate time series. In SDM, 2009.
  10. 10.Shohreh Deldari, Daniel V. Smith, Hao Xue, and Flora D. Salim. Time series change point detection with self-supervised contrastive predictive coding. In WWW, 2021.
  11. 11.Ailin Deng and Bryan Hooi. Graph neural network-based anomaly detection in multivariate time series. AAAI, 2021.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  14. 14.I. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial nets. In NeurIPS, 2014.
  15. 15.Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. Music transformer. In ICLR, 2019.
  16. 16.Kyle Hundman, Valentino Constantinou, Christopher Laporte, Ian Colwell, and Tom Söderström. Detecting spacecraft anomalies using lstms and nonparametric dynamic thresholding. KDD, 2018.
  17. 17.Eamonn J. Keogh, Taposh Roy, Naik U, and Agrawal A. Multi-dataset time-series anomaly detection competition, Competition of International Conference on Knowledge Discovery & Data Mining 2021. URL https://compete.hexagon-ml.com/practice/competition/39/.
  18. 18.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  19. 19.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. In ICLR, 2020.
  20. 20.Kwei-Herng Lai, D. Zha, Junjie Xu, and Yue Zhao. Revisiting time series outlier detection: Definitions and benchmarks. In NeurIPS Dataset and Benchmark Track, 2021.
  21. 21.Dan Li, Dacheng Chen, Lei Shi, Baihong Jin, Jonathan Goh, and See-Kiong Ng. Mad-gan: Multivariate anomaly detection for time series data with generative adversarial networks. In ICANN, 2019a.
  22. 22.Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In NeurIPS, 2019b.
  23. 23.Zhihan Li, Youjian Zhao, Jiaqi Han, Ya Su, Rui Jiao, Xidao Wen, and Dan Pei. Multivariate time series anomaly detection and interpretation using hierarchical inter-metric and temporal embedding. KDD, 2021.
  24. 24.F. Liu, K. Ting, and Z. Zhou. Isolation forest. ICDM, 2008.
  25. 25.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Ching-Feng Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021.
  26. 26.Aditya P. Mathur and Nils Ole Tippenhauer. Swat: a water treatment testbed for research and training on ICS security. In CySWATER, 2016.
  27. 27.Radford M. Neal. Pattern recognition and machine learning. Technometrics, 2007.
  28. 28.Daehyung Park, Yuuna Hoshi, and Charles C. Kemp. A multimodal anomaly detector for robotassisted feeding using an lstm-based variational autoencoder. RA-L, 2018.
  29. 29.Adam Paszke, S. Gross, Francisco Massa, A. Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Z. Lin, N. Gimelshein, L. Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
  30. 30.Mathias Perslev, Michael Jensen, Sune Darkner, Poul Jørgen Jennum, and Christian Igel. U-time: A fully convolutional network for time series segmentation applied to sleep staging. In NeurIPS. 2019.
  31. 31.Lukas Ruff, Nico Görnitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Robert A. Vandermeulen, Alexander Binder, Emmanuel Müller, and M. Kloft. Deep one-class classification. In ICML, 2018.
  32. 32.T. Schlegl, Philipp Seeböck, S. Waldstein, G. Langs, and U. Schmidt-Erfurth. f-anogan: Fast unsupervised anomaly detection with generative adversarial networks. Med. Image Anal., 2019.
  33. 33.B. Schölkopf, John C. Platt, J. Shawe-Taylor, Alex Smola, and R. C. Williamson. Estimating the support of a high-dimensional distribution. Neural Comput., 2001.
  34. 34.Lifeng Shen, Zhuocong Li, and James T. Kwok. Timeseries anomaly detection using temporal hierarchical one-class network. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), NeurIPS, 2020.
  35. 35.Youjin Shin, Sangyup Lee, Shahroz Tariq, Myeong Shin Lee, Okchul Jung, Daewon Chung, and Simon S. Woo. Itad: Integrative tensor-based anomaly detection system for reducing false positives of satellite systems. CIKM, 2020.
  36. 36.Ya Su, Y. Zhao, Chenhao Niu, Rong Liu, W. Sun, and Dan Pei. Robust anomaly detection for multivariate time series through stochastic recurrent neural network. KDD, 2019.
  37. 37.Jian Tang, Zhixiang Chen, A. Fu, and D. Cheung. Enhancing effectiveness of outlier detections for low density patterns. In PAKDD, 2002.
  38. 38.Shahroz Tariq, Sangyup Lee, Youjin Shin, Myeong Shin Lee, Okchul Jung, Daewon Chung, and Simon S. Woo. Detecting anomalies in space using multivariate convolutional lstm with mixtures of probabilistic pca. KDD, 2019.
  39. 39.D. Tax and R. Duin. Support vector data description. Mach. Learn., 2004.
  40. 40.Robert Tibshirani, Guenther Walther, and Trevor Hastie. Estimating the number of clusters in a dataset via the gap statistic. J. R. Stat. Soc. (Series B), 2001.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  42. 42.Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. Autoformer: Decomposition transformers with Auto-Correlation for long-term series forecasting. In NeurIPS, 2021.
  43. 43.Haowen Xu, Wenxiao Chen, N. Zhao, Zeyan Li, Jiahao Bu, Zhihan Li, Y. Liu, Y. Zhao, Dan Pei, Yang Feng, Jian Jhen Chen, Zhaogang Wang, and Honglin Qiao. Unsupervised anomaly detection via variational auto-encoder for seasonal kpis in web applications. WWW, 2018.
  44. 44.Takehisa Yairi, Naoya Takeishi, Tetsuo Oda, Yuta Nakajima, Naoki Nishimura, and Noboru Takata. A data-driven health monitoring method for satellite housekeeping data based on probabilistic clustering and dimensionality reduction. IEEE Trans. Aerosp. Electron. Syst., 2017.
  45. 45.Hang Zhao, Yujing Wang, Juanyong Duan, Congrui Huang, Defu Cao, Yunhai Tong, Bixiong Xu, Jing Bai, Jie Tong, and Qi Zhang. Multivariate time-series anomaly detection via graph attention network. ICDM, 2020.
  46. 46.Bin Zhou, Shenghua Liu, Bryan Hooi, Xueqi Cheng, and Jing Ye. Beatgan: Anomalous rhythm detection using adversarially generated time series. In IJCAI, 2019.
  47. 47.Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. Informer: Beyond efficient transformer for long sequence time-series forecasting. In AAAI, 2021.
  48. 48.Bo Zong, Qi Song, Martin Renqiang Min, Wei Cheng, Cristian Lumezanu, Dae-ki Cho, and Haifeng Chen. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In ICLR, 2018.

Citation

MLA
Xu, J., et al. “Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy”. arXiv, 2021, http://arxiv.org/abs/2110.02642v5.
APA
Xu, J., Wu, H., Wang, J., & Long, M. (2021). Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy. arXiv. http://arxiv.org/abs/2110.02642v5
Chicago
Xu, J., H. Wu, J. Wang, and M. Long. 2021. “Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy”. arXiv. http://arxiv.org/abs/2110.02642v5.
Harvard
Xu, J. et al. (2021) “Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.02642v5.
Vancouver
1. Xu J, Wu H, Wang J, Long M (2021) Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy. arXiv

BibTeX

@article{xu2021anomaly,
  title = {Anomaly Transformer: Time Series Anomaly Detection with Association Discrepancy},
  author = {Xu, Jiehui and Wu, Haixu and Wang, Jianmin and Long, Mingsheng},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.02642v5},
  eprint = {2110.02642}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors