TS2Vec: Towards Universal Representation of Time Series

Zhihan YueYujing WangJuanyong DuanTianmeng YangCongrui HuangYunhai TongBixiong Xu

article2022AAAI1,176 citations

Proposes a universal contrastive learning framework that uses hierarchical contrasting and contextual consistency to learn multiscale time series representations, achieving state-of-the-art results across classification, forecasting, and anomaly detection benchmarks.

Listen

Organizations across finance, energy, demand planning, and technology operations rely heavily on time series data to guide critical decisions. However, existing automated methods for learning representations from time series struggle to support diverse analytics tasks. Traditional models typically generate a single summary representation for an entire sequence, which fails to capture fine-grained patterns necessary for point-by-point forecasting and anomaly detection. Furthermore, prior approaches frequently borrow contrastive learning assumptions from computer vision—such as cropping or transformation invariance—that misrepresent time series data when underlying trends and statistical distributions shift over time.

The article demonstrates TS2Vec, a universal and flexible framework designed to learn contextual time series representations across arbitrary semantic scales and sub-sequences. The authors evaluate whether a single self-supervised model can simultaneously deliver state-of-the-art accuracy and operational efficiency across three core analytics workloads: classification, forecasting, and anomaly detection.

The evaluated approach uses a dilated convolutional neural network encoder that combines hierarchical contrastive loss with contextual consistency. Instead of applying invasive transformations or assuming rigid temporal smoothness, the method generates augmented context views through random cropping and timestamp masking in the latent space. It pairs this with a hierarchical loss function operating across both instance and temporal dimensions, allowing representations to capture fine-grained timestamp behaviors as well as overarching sequence-level patterns. Credibility is supported through extensive testing across standard benchmark collections, including 125 univariate UCR datasets, 29 multivariate UEA datasets, standard electricity and temperature forecasting datasets, and real-world industrial anomaly detection benchmarks.

The experimental findings show significant performance and efficiency advantages. First, TS2Vec established new state-of-the-art classification accuracy among unsupervised methods, improving average accuracy by 2.4 percentage points on UCR benchmarks and 3.0 percentage points on UEA benchmarks while cutting training runtimes down to under one hour. Second, when applying a simple linear regression model on top of the learned embeddings, TS2Vec reduced forecasting mean squared error by 32.6% in univariate settings and 28.2% in multivariate settings compared to dedicated forecasting baselines, while requiring only a fraction of their training and inference compute. Third, in streaming anomaly detection, TS2Vec improved F1 performance by 18.2% on the Yahoo benchmark and 5.5% on enterprise performance indicators compared to competing unsupervised detectors. Finally, robustness testing showed the architecture maintained stable performance even when 50% of input timestamps were missing, suffering only minor accuracy drops of 1% to 2% across large test sets.

These results carry substantial operational implications. Because TS2Vec generates reusable, multi-scale embeddings, an organization needs to train the foundational representation model only once per dataset. Downstream applications—such as multiple forecast horizons, classification routines, or real-time anomaly alerts—can then run using lightweight linear heads. This unified pipeline significantly lowers compute expenses, shortens development cycles, and mitigates deployment complexity. Furthermore, the model's resilience to missing data directly reduces the risk of pipeline failure in industrial settings where sensor dropouts and incomplete data streams are frequent.

Organizations seeking to streamline time series infrastructure should consider piloting hierarchical contextual representation frameworks like TS2Vec for unified predictive analytics pipelines. The source code is publicly accessible, facilitating direct proofs of concept on internal telemetry and forecasting workloads. Where applicable, teams can deploy pre-trained encoders in zero-shot or cold-start anomaly detection settings before committing to full retraining. However, decision-makers should note that evaluations were conducted primarily on standardized public research benchmarks with specific hyperparameter calibrations. Prior to enterprise-wide production deployment, organizations should validate latency constraints, evaluate edge cases on domain-specific data, and explore future extensions into specialized operational settings.

Cover for TS2Vec: Towards Universal Representation of Time Series

Abstract

This paper presents TS2Vec, a universal framework for learning representations of time series in an arbitrary semantic level. Unlike existing methods, TS2Vec performs contrastive learning in a hierarchical way over augmented context views, which enables a robust contextual representation for each timestamp. Furthermore, to obtain the representation of an arbitrary sub-sequence in the time series, we can apply a simple aggregation over the representations of corresponding timestamps. We conduct extensive experiments on time series classification tasks to evaluate the quality of time series representations. As a result, TS2Vec achieves significant improvement over existing SOTAs of unsupervised time series representation on 125 UCR datasets and 29 UEA datasets. The learned timestamp-level representations also achieve superior results in time series forecasting and anomaly detection tasks. A linear regression trained on top of the learned representations outperforms previous SOTAs of time series forecasting. Furthermore, we present a simple way to apply the learned representations for unsupervised anomaly detection, which establishes SOTA results in the literature. The source code is publicly available at https://github.com/yuezhihan/ts2vec.

Table of Contents

  • Introduction
  • Method
  • Model Architecture
  • Problem Definition
  • Contextual Consistency
  • Hierarchical Contrasting
  • Experiments
  • Time Series Classification
  • Time Series Forecasting
  • Time Series Anomaly Detection
  • Analysis
  • Ablation Study
  • Robustness to Missing Data
  • Visualized Explanation
  • Conclusion
  • References

Knowls

  1. Knowl 1 — TS2Vec Architecture for Multi-Granularity Time Series Representation

    model/method

    TS2Vec is an unsupervised representation learning framework designed to produce contextual embeddings for time series at arbitrary granularities (timestamp-level, subseries-level, and instance-level).

    Given an input time series xi∈RT×Fx_i \in \mathbb{R}^{T \times F} with sequence length TT and feature dimension FF, the encoder fθf_\theta maps xix_i to a sequence of representation vectors ri={ri,1,ri,2,…,ri,T}r_i = \{r_{i,1}, r_{i,2}, \dots, r_{i,T}\}, where each timestamp representation satisfies ri,t∈RKr_{i,t} \in \mathbb{R}^K and KK is the embedding dimension.

    The encoder fθf_\theta consists of three sequential modules:

    1. Input Projection Layer: A fully connected layer that projects each observation xi,t∈RFx_{i,t} \in \mathbb{R}^F at timestamp tt into a high-dimensional latent vector zi,t∈RDz_{i,t} \in \mathbb{R}^D.
    2. Timestamp Masking Module: Randomly masks the projected latent vectors zi={zi,t}z_i = \{z_{i,t}\} along the temporal dimension using a binary mask vector m∈{0,1}Tm \in \{0, 1\}^T, whose elements are independently drawn from a Bernoulli distribution with probability p=0.5p = 0.5. Masking is applied in the latent space rather than raw data to accommodate unbounded time series values without requiring an arbitrary masking token.
    3. Dilated CNN Backbone: A sequence of 10 residual blocks. Each block comprises two 1D convolutional layers with a dilation factor of 2l2^l for the ll-th block (l∈{1,2,…,10}l \in \{1, 2, \dots, 10\}), yielding an exponentially expanding receptive field.

    To derive representations for an arbitrary subseries spanning timestamps [a,b][a, b] (where 1≤a≤b≤T1 \le a \le b \le T), TS2Vec applies temporal max pooling across the corresponding timestamp embeddings: ri,[a,b]=MaxPool({ri,t∣t∈[a,b]})r_{i,[a,b]} = \text{MaxPool}(\{r_{i,t} \mid t \in [a, b]\}) For full instance-level representation, max pooling is applied over all timestamps t∈[1,T]t \in [1, T].

  2. Knowl 2 — Contextual Consistency for Contrastive Pair Generation

    model/method

    Contextual consistency is a contrastive learning positive-pair generation strategy that aligns representations of the identical timestamp across two different augmented context views of the same time series instance.

    Unlike image-derived augmentations (e.g., jittering, scaling, permutation) or segment smoothness assumptions that can violate the dynamic distributions, level shifts, and anomalies of time series, contextual consistency maintains the exact magnitudes and positions of values while forcing the encoder to reconstruct timestamp representations across distinct contextual environments.

    For an input sequence xi∈RT×Fx_i \in \mathbb{R}^{T \times F}, two context views are constructed during training via two operations:

    1. Random Cropping: Two overlapping time intervals [a1,b1][a_1, b_1] and [a2,b2][a_2, b_2] are randomly sampled such that 0<a1≤a2≤b1≤b2≤T0 < a_1 \le a_2 \le b_1 \le b_2 \le T. The model evaluates the representations over the overlapping interval Ω=[a2,b1]\Omega = [a_2, b_1]. Random cropping prevents representation collapse and encourages position-agnostic representations.
    2. Timestamp Masking: Latent vectors along the time axis are independently masked with a Bernoulli mask (p=0.5p = 0.5) in each view's forward pass.

    The representations ri,tr_{i,t} and ri,t′r'_{i,t} generated at the exact same timestamp t∈Ωt \in \Omega from the two augmented views form a positive pair.

  3. Knowl 3 — TS2Vec Dual Contrastive Loss Formulation

    equation

    TS2Vec jointly optimizes an instance-wise contrastive loss and a temporal contrastive loss across positive and negative pairs generated from two augmented context views of time series in a batch.

    Let BB denote the mini-batch size, i∈{1,…,B}i \in \{1, \dots, B\} the sample index, Ω\Omega the set of timestamps in the overlap of two cropped subseries, and ri,t,ri,t′∈RKr_{i,t}, r'_{i,t} \in \mathbb{R}^K the representations at timestamp tt from the two augmented context views of instance xix_i.

    The temporal contrastive loss for instance ii at timestamp tt contrasts the positive representation ri,t′r'_{i,t} against representations at all other timestamps t′∈Ωt' \in \Omega from the same instance: ℓtemp(i,t)=−log⁡exp⁡(ri,t⋅ri,t′)∑t′∈Ω(exp⁡(ri,t⋅ri,t′′)+I[t≠t′]exp⁡(ri,t⋅ri,t′))\ell_{\text{temp}}^{(i,t)} = -\log \frac{\exp(r_{i,t} \cdot r'_{i,t})}{\sum_{t' \in \Omega} \left( \exp(r_{i,t} \cdot r'_{i,t'}) + \mathbb{I}_{[t \neq t']} \exp(r_{i,t} \cdot r_{i,t'}) \right)} where I[⋅]\mathbb{I}_{[\cdot]} is the indicator function.

    The instance-wise contrastive loss for instance ii at timestamp tt contrasts ri,t′r'_{i,t} against representations at the same timestamp tt from all other instances j≠ij \neq i in the batch: ℓinst(i,t)=−log⁡exp⁡(ri,t⋅ri,t′)∑j=1B(exp⁡(ri,t⋅rj,t′)+I[i≠j]exp⁡(ri,t⋅rj,t))\ell_{\text{inst}}^{(i,t)} = -\log \frac{\exp(r_{i,t} \cdot r'_{i,t})}{\sum_{j=1}^B \left( \exp(r_{i,t} \cdot r'_{j,t}) + \mathbb{I}_{[i \neq j]} \exp(r_{i,t} \cdot r_{j,t}) \right)}

    The combined dual loss over the batch and overlapping timestamps is: Ldual=1B∣Ω∣∑i=1B∑t∈Ω(ℓtemp(i,t)+ℓinst(i,t))\mathcal{L}_{\text{dual}} = \frac{1}{B |\Omega|} \sum_{i=1}^B \sum_{t \in \Omega} \left( \ell_{\text{temp}}^{(i,t)} + \ell_{\text{inst}}^{(i,t)} \right)

  4. Knowl 4 — Hierarchical Contrastive Loss Computation

    algorithm

    The hierarchical contrastive loss computes multiscale contextual objectives by recursively downsampling temporal representations via 1D max pooling and evaluating the dual contrastive loss at each scale level until the sequence length is reduced to 1.

    Input: Representation tensor r∈RB×TΩ×Kr \in \mathbb{R}^{B \times T_\Omega \times K} from context view 1, representation tensor r′∈RB×TΩ×Kr' \in \mathbb{R}^{B \times T_\Omega \times K} from context view 2
    Output: Hierarchical contrastive loss scalar Lhier\mathcal{L}_{\text{hier}}
    Lhier←Ldual(r,r′)L_{\text{hier}} \leftarrow \mathcal{L}_{\text{dual}}(r, r')
    d←1d \leftarrow 1
    while time_length(rr) > 1 do
        r←maxpool1d(r,kernel_size=2)r \leftarrow \text{maxpool1d}(r, \text{kernel\_size} = 2)
        r′←maxpool1d(r′,kernel_size=2)r' \leftarrow \text{maxpool1d}(r', \text{kernel\_size} = 2)
        Lhier←Lhier+Ldual(r,r′)L_{\text{hier}} \leftarrow L_{\text{hier}} + \mathcal{L}_{\text{dual}}(r, r')
        d←d+1d \leftarrow d + 1
    end while
    Lhier←Lhier/d\mathcal{L}_{\text{hier}} \leftarrow L_{\text{hier}} / d
    return Lhier\mathcal{L}_{\text{hier}}

    Here, maxpool1d\text{maxpool1d} performs 1D max pooling with stride 2 and kernel size 2 along the time dimension. Iterative pooling aggregates adjacent temporal information, and the top-level representation (TΩ=1T_\Omega = 1) provides instance-level semantic contrast.

  5. Knowl 5 — Time Series Forecasting Protocol via Linear Probing on TS2Vec Representations

    model/method

    In the TS2Vec forecasting framework, the representation encoder fθf_\theta is trained once in an unsupervised manner on historical observations. Forecasting future horizons is performed by training a linear ridge regression model on top of the fixed timestamp-level embeddings.

    Given an input history slice xt−Tl+1,…,xtx_{t - T_l + 1}, \dots, x_t, the encoder extracts the contextual embedding rt∈RKr_t \in \mathbb{R}^K corresponding to the final observed timestamp tt. To predict the future HH steps x^t+1,…,x^t+H\hat{x}_{t+1}, \dots, \hat{x}_{t+H}:

    1. For a univariate time series (F=1F = 1), a linear regression model with L2L_2 regularization maps rtr_t directly to x^∈RH\hat{x} \in \mathbb{R}^H.
    2. For a multivariate time series with FF features, the linear regression model maps rtr_t to x^∈RF⋅H\hat{x} \in \mathbb{R}^{F \cdot H}.

    Because the underlying encoder fθf_\theta produces universal representations independent of the forecast horizon, it does not need to be retrained for different horizon values HH; only the lightweight linear regression weights are fitted for each HH.

  6. Knowl 6 — Unsupervised Streaming Anomaly Detection via Representation Dissimilarity

    model/method

    TS2Vec performs streaming anomaly detection on a time series slice x1,x2,…,xtx_1, x_2, \dots, x_t by evaluating the stability of the learned representation of timestamp tt under an input mask perturbation.

    During inference for each incoming timestamp tt:

    1. The trained encoder performs an unmasked forward pass on x1:tx_{1:t}, yielding representation vector rtu∈RKr_t^u \in \mathbb{R}^K at timestamp tt.
    2. The encoder performs a second forward pass where the observation at timestamp tt is masked out (xtx_t masked in latent projection), yielding representation vector rtm∈RKr_t^m \in \mathbb{R}^K.
    3. The raw anomaly score αt\alpha_t is computed as the L1L_1 distance between the representations: αt=∥rtu−rtm∥1\alpha_t = \|r_t^u - r_t^m\|_1
    4. To prevent baseline drift, an adjusted score is computed using the average raw score over a preceding window of length Z=21Z = 21: αˉt=1Z∑i=t−Zt−1αi,αtadj=αt−αˉtαˉt\bar{\alpha}_t = \frac{1}{Z} \sum_{i=t-Z}^{t-1} \alpha_i, \quad \alpha_t^{\text{adj}} = \frac{\alpha_t - \bar{\alpha}_t}{\bar{\alpha}_t}
    5. Timestamp tt is classified as an anomaly if αtadj>μ+βσ\alpha_t^{\text{adj}} > \mu + \beta \sigma, where μ\mu and σ\sigma are the historical mean and standard deviation of adjusted scores, and β=4\beta = 4 is a threshold hyperparameter.
  7. Knowl 7 — Time Series Classification Benchmark Results on UCR and UEA Archives

    data/table

    TS2Vec was evaluated on 125 univariate datasets from the UCR archive and 29 multivariate datasets from the UEA archive for time series classification. An SVM classifier with an RBF kernel was trained on top of the max-pooled instance representations. Representation dimension was set to 320 for all representation learning baselines (T-Loss, TS-TCC, TST, TNC).

    125 UCR datasets 29 UEA datasets
    Method Avg. Acc. Avg. Rank Training Time (hours) Avg. Acc. Avg. Rank Training Time (hours)
    DTW 0.727 4.33 – 0.650 3.74 –
    TNC 0.761 3.52 228.4 0.677 3.84 91.2
    TST 0.641 5.23 17.1 0.635 4.36 28.6
    TS-TCC 0.757 3.38 1.1 0.682 3.53 3.6
    T-Loss 0.806 2.73 38.0 0.675 3.12 15.1
    TS2Vec 0.830 (+2.4%) 1.82 0.9 0.712 (+3.0%) 2.40 0.6

    TS2Vec achieved the highest average accuracy and lowest average rank across both univariate and multivariate archives while requiring the lowest total training time on an NVIDIA GeForce RTX 3090 GPU.

  8. Knowl 8 — Univariate Time Series Forecasting Results on ETT and Electricity Datasets

    data/table

    Forecasting performance was evaluated on four benchmark datasets (ETTh1, ETTh2, ETTm1, and Electricity) using Mean Squared Error (MSE) across prediction horizons H∈{24,48,96,168,288,336,672,720}H \in \{24, 48, 96, 168, 288, 336, 672, 720\}. TS2Vec fits an L2L_2-penalized linear regression model on top of unsupervised representations.

    Dataset HH TS2Vec Informer LogTrans N-BEATS TCN LSTnet
    ETTh1\text{ETTh}_1 24 0.039 0.098 0.103 0.094 0.075 0.108
    48 0.062 0.158 0.167 0.210 0.227 0.175
    168 0.134 0.183 0.207 0.232 0.316 0.396
    336 0.154 0.222 0.230 0.232 0.306 0.468
    720 0.163 0.269 0.273 0.322 0.390 0.659
    ETTh2\text{ETTh}_2 24 0.090 0.093 0.102 0.198 0.103 3.554
    48 0.124 0.155 0.169 0.234 0.142 3.190
    168 0.208 0.232 0.246 0.331 0.227 2.800
    336 0.213 0.263 0.267 0.431 0.296 2.753
    720 0.214 0.277 0.303 0.437 0.325 2.878
    ETTm1\text{ETTm}_1 24 0.015 0.030 0.065 0.054 0.041 0.090
    48 0.027 0.069 0.078 0.190 0.101 0.179
    96 0.044 0.194 0.199 0.183 0.142 0.272
    288 0.103 0.401 0.411 0.186 0.318 0.462
    672 0.156 0.512 0.598 0.197 0.397 0.639
    Electric. 24 0.260 0.251 0.528 0.427 0.263 0.281
    48 0.319 0.346 0.409 0.551 0.373 0.381
    168 0.427 0.544 0.959 0.893 0.609 0.599
    336 0.565 0.713 1.079 1.035 0.855 0.823
    720 0.861 1.182 1.001 1.548 1.263 1.278
    Avg. 0.209 0.310 0.370 0.399 0.338 1.099

    TS2Vec achieved an average MSE of 0.209, representing a 32.6% reduction in average MSE compared to Informer (0.310).

  9. Knowl 9 — Unsupervised Anomaly Detection Performance in Normal and Cold-Start Settings

    data/table

    TS2Vec was evaluated on univariate anomaly detection on the Yahoo and KPI benchmarks under two protocols: (1) Normal setting (first half of series used for training, second half for evaluation); (2) Cold-start setting (no in-domain training data; TS2Vec†\text{TS2Vec}^\dagger encoder was trained purely on the UCR FordA dataset and tested directly).

    Yahoo KPI
    Method F1F_1 Prec. Rec. F1F_1 Prec. Rec.
    SPOT 0.338 0.269 0.454 0.217 0.786 0.126
    DSPOT 0.316 0.241 0.458 0.521 0.623 0.447
    DONUT 0.026 0.013 0.825 0.347 0.371 0.326
    SR 0.563 0.451 0.747 0.622 0.647 0.598
    TS2Vec 0.745 0.729 0.762 0.677 0.929 0.533
    Cold-start:
    FFT 0.291 0.202 0.517 0.538 0.478 0.615
    Twitter-AD 0.245 0.166 0.462 0.330 0.411 0.276
    Luminol 0.388 0.254 0.818 0.417 0.306 0.650
    SR 0.529 0.404 0.765 0.666 0.637 0.697
    TS2Vec†\text{TS2Vec}^\dagger 0.726 0.692 0.763 0.676 0.907 0.540

    In the normal setting, TS2Vec improved F1F_1 by 18.2% on Yahoo and 5.5% on KPI compared to the strongest baseline. In the cold-start setting, TS2Vec†\text{TS2Vec}^\dagger retained comparable F1F_1 scores (0.726 and 0.676), outperforming domain-free baselines and showing out-of-domain transferability.

  10. Knowl 10 — Ablation Analysis of TS2Vec Components and Architectural Choices

    data/table

    An ablation study across 128 UCR classification datasets evaluated the contribution of each module, positive-pair selection strategy, input data augmentations, and encoder backbones.

    Variant Avg. Accuracy
    TS2Vec 0.829
    w/o Temporal Contrast 0.819 (-1.0%)
    w/o Instance Contrast 0.824 (-0.5%)
    w/o Hierarchical Contrast 0.812 (-1.7%)
    w/o Random Cropping 0.808 (-2.1%)
    w/o Timestamp Masking 0.820 (-0.9%)
    w/o Input Projection Layer 0.817 (-1.2%)
    Positive Pair Selection
    Contextual Consistency 0.829
    →\rightarrow Temporal Consistency 0.807 (-2.2%)
    →\rightarrow Subseries Consistency 0.780 (-4.9%)
    Augmentations
    + Jitter 0.814 (-1.5%)
    + Scaling 0.814 (-1.5%)
    + Permutation 0.796 (-3.3%)
    Backbone Architectures
    Dilated CNN 0.829
    →\rightarrow LSTM 0.779 (-5.0%)
    →\rightarrow Transformer 0.647 (-18.2%)

    Key findings include:

    1. Removing hierarchical contrasting causes a 1.7% drop, confirming the value of multi-scale contextual supervision.
    2. Standard computer vision augmentations (jitter, scaling, permutation) degrade classification accuracy by 1.5% to 3.3%.
    3. Dilated CNN significantly outperforms parameter-matched LSTM (-5.0%) and Transformer (-18.2%) architectures on time series.
  11. Knowl 11 — Robustness of TS2Vec to Missing Data via Masking and Hierarchical Contrasting

    empirical result

    TS2Vec exhibits strong robustness to missing values in time series due to the combined action of timestamp masking (which trains the network to reconstruct representations from partial observations) and hierarchical contrasting (which pools long-range context across time scales).

    When evaluated on the four largest UCR datasets (StarLightCurves, UWaveGestureLibraryAll, HandOutlines, MixedShapesRegularTrain) under varying proportions of randomly dropped observations at both training and test time:

    1. At a 50% missing rate, TS2Vec incurs minimal classification accuracy degradation: almost 0% drop on UWaveGestureLibraryAll, 2.1% drop on StarLightCurves, 2.1% drop on HandOutlines, and 1.2% drop on MixedShapesRegularTrain.
    2. Variants lacking hierarchical contrasting (w/o Hierarchical Contrast) or timestamp masking (w/o Timestamp Masking) exhibit steep performance collapses as the missing rate increases from 0% to 90%, demonstrating that multi-scale contextual pooling is critical when local surrounding context is absent.

Coverage note — None was omitted; all primary methodological components, algorithms, mathematical definitions, downstream task formulations, benchmark tables (classification, forecasting, anomaly detection), ablation experiments, and missing data analyses have been fully extracted into self-contained knowls.

References

  1. 1.Bagnall, A. J.; Dau, H. A.; Lines, J.; Flynn, M.; Large, J.; Bostrom, A.; Southam, P.; and Keogh, E. J. 2018. The UEA multivariate time series classification archive, 2018. CoRR, abs/1811.00075.
  2. 2.Bai, S.; Kolter, J. Z.; and Koltun, V. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. CoRR, abs/1803.01271.
  3. 3.Brennan, V.; and Ritesh, M. 2018. Luminol. https://github.com/linkedin/luminol. Accessed: 2021-05-05.
  4. 4.Cao, D.; Wang, Y.; Duan, J.; Zhang, C.; Zhu, X.; Huang, C.; Tong, Y.; Xu, B.; Bai, J.; Tong, J.; and Zhang, Q. 2020. Spectral Temporal Graph Neural Network for Multivariate Time-series Forecasting. In Advances in Neural Information Processing Systems, volume 33, 17766–17778. Curran Associates, Inc.
  5. 5.Dau, H. A.; Bagnall, A.; Kamgar, K.; Yeh, C.-C. M.; Zhu, Y.; Gharghabi, S.; Ratanamahatana, C. A.; and Keogh, E. 2019. The UCR time series archive. IEEE/CAA Journal of Automatica Sinica, 6(6): 1293–1305.
  6. 6.Demšar, J. 2006. Statistical comparisons of classifiers over multiple data sets. The Journal of Machine Learning Research, 7: 1–30.
  7. 7.Dua, D.; and Graff, C. 2017. UCI Machine Learning Repository. http://archive.ics.uci.edu/ml. Accessed: 2021-05-12.
  8. 8.Eldele, E.; Ragab, M.; Chen, Z.; Wu, M.; Kwoh, C. K.; Li, X.; and Guan, C. 2021. Time-Series Representation Learning via Temporal and Contextual Contrasting. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, 2352–2359.
  9. 9.Franceschi, J.-Y.; Dieuleveut, A.; and Jaggi, M. 2019. Unsupervised Scalable Representation Learning for Multivariate Time Series. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  10. 10.Lai, G.; Chang, W.-C.; Yang, Y.; and Liu, H. 2018. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, 95–104.
  11. 11.Li, S.; Jin, X.; Xuan, Y.; Zhou, X.; Chen, W.; Wang, Y.-X.; and Yan, X. 2019. Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  12. 12.Nikolay Laptev, Y. B., Saeed Amizadeh. 2015. A Benchmark Dataset for Time Series Anomaly Detection. https://yahooresearch.tumblr.com/post/114590420346/a-benchmark-dataset-for-time-series-anomaly. Accessed: 2021-05-02.
  13. 13.Oreshkin, B. N.; Carpov, D.; Chapados, N.; and Bengio, Y. 2019. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations.
  14. 14.Rasheed, F.; Peng, P.; Alhajj, R.; and Rokne, J. 2009. Fourier transform based spatial outlier mining. In International Conference on Intelligent Data Engineering and Automated Learning, 317–324. Springer.
  15. 15.Ren, H.; Xu, B.; Wang, Y.; Yi, C.; Huang, C.; Kou, X.; Xing, T.; Yang, M.; Tong, J.; and Zhang, Q. 2019. Time-series anomaly detection service at Microsoft. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 3009–3017.
  16. 16.Siffer, A.; Fouque, P.-A.; Termier, A.; and Largouet, C. 2017. Anomaly detection in streams with extreme value theory. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1067–1075.
  17. 17.Tonekaboni, S.; Eytan, D.; and Goldenberg, A. 2021. Unsupervised Representation Learning for Time Series with Temporal Neighborhood Coding. In International Conference on Learning Representations.
  18. 18.Vallis, O.; Hochenbaum, J.; and Kejariwal, A. 2014. A novel technique for long-term anomaly detection in the cloud. In 6th USENIX workshop on hot topics in cloud computing (HotCloud 14).
  19. 19.Wu, L.; Yen, I. E.-H.; Yi, J.; Xu, F.; Lei, Q.; and Witbrock, M. 2018. Random warping series: A random features method for time-series embedding. In International Conference on Artificial Intelligence and Statistics, 793–802. PMLR.
  20. 20.Xu, H.; Chen, W.; Zhao, N.; Li, Z.; Bu, J.; Li, Z.; Liu, Y.; Zhao, Y.; Pei, D.; Feng, Y.; et al. 2018. Unsupervised anomaly detection via variational auto-encoder for seasonal kpis in web applications. In Proceedings of the 2018 World Wide Web Conference, 187–196.
  21. 21.Zerveas, G.; Jayaraman, S.; Patel, D.; Bhamidipaty, A.; and Eickhoff, C. 2021. A Transformer-Based Framework for Multivariate Time Series Representation Learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, KDD '21, 2114–2124. New York, NY, USA: Association for Computing Machinery. ISBN 9781450383325.
  22. 22.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In The Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, online. AAAI Press.

Citation

MLA
Yue, Z., et al. “TS2Vec: Towards Universal Representation of Time Series”. arXiv, 2021, http://arxiv.org/abs/2106.10466v4.
APA
Yue, Z., Wang, Y., Duan, J., Yang, T., Huang, C., Tong, Y., & Xu, B. (2021). TS2Vec: Towards Universal Representation of Time Series. arXiv. http://arxiv.org/abs/2106.10466v4
Chicago
Yue, Z., Y. Wang, J. Duan, et al. 2021. “TS2Vec: Towards Universal Representation of Time Series”. arXiv. http://arxiv.org/abs/2106.10466v4.
Harvard
Yue, Z. et al. (2021) “TS2Vec: Towards Universal Representation of Time Series”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.10466v4.
Vancouver
1. Yue Z, Wang Y, Duan J, Yang T, Huang C, Tong Y, Xu B (2021) TS2Vec: Towards Universal Representation of Time Series. arXiv

BibTeX

@article{yue2021ts2vec,
  title = {TS2Vec: Towards Universal Representation of Time Series},
  author = {Yue, Zhihan and Wang, Yujing and Duan, Juanyong and Yang, Tianmeng and Huang, Congrui and Tong, Yunhai and Xu, Bixiong},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.10466v4},
  eprint = {2106.10466}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF