TimesURL: Self-Supervised Contrastive Learning for Universal Time Series Representation Learning

Jiexi LiuSongcan Chen

article2024AAAI124 citations

Proposes TimesURL, a self-supervised framework combining frequency-temporal data augmentations, synthesized Universum hard negatives, and a joint reconstruction objective to learn universal time series representations that achieve state-of-the-art results across six distinct downstream tasks.

Listen

Time series data is essential across critical domains such as weather forecasting, economic planning, industrial monitoring, and healthcare diagnostics. While self-supervised contrastive learning—a machine learning technique that trains models on unlabelled data by comparing similar and dissimilar samples—has succeeded in computer vision and natural language processing, directly transferring these techniques to time series has proven problematic. Existing approaches often rely on data modifications that distort underlying temporal relationships, fail to provide sufficiently challenging negative examples during training, and focus narrowly on either fine-grained point predictions or high-level whole-series patterns rather than addressing both.

The article aims to resolve these limitations by introducing TimesURL, a novel self-supervised framework designed to learn universal, task-agnostic time series representations that perform effectively across diverse downstream applications. The authors evaluate this framework against approximately 15 existing specialized baseline models across six major analytical tasks: short- and long-term forecasting, missing value imputation, classification, anomaly detection, and transfer learning.

To achieve this, the authors designed a three-part framework evaluated across broad benchmark suites, including 158 multivariate and univariate classification datasets as well as standard energy, weather, and web performance datasets. First, they implemented a hybrid augmentation method combining frequency mixing and random temporal cropping to preserve core trends without adding artificial noise. Second, they synthesized artificial hard negative samples—termed double Universums—across time and instance dimensions by blending anchor samples with non-matching segments to make training contrast more effective. Third, they jointly trained the network on both contrastive loss and masked time series reconstruction, capturing both fine-grained segment details and broad instance-level context.

The experimental findings show that TimesURL consistently establishes a new state of the art. In time series classification, TimesURL achieved 75.2% average accuracy across 30 multivariate benchmarks (a 3.8 percentage point improvement over the prior leading method) and 84.5% across 128 univariate datasets. In missing data imputation across multiple missingness ratios, it achieved the lowest overall error rates, yielding an average mean squared error of 1.326. For short- and long-term forecasting, TimesURL outperformed both self-supervised and dedicated end-to-end models across most forecast horizons. In anomaly detection and transfer learning scenarios, the model delivered top-tier precision, recall, and cross-dataset adaptability.

These results demonstrate that organizations do not need separate, specialized feature extraction pipelines for each time series application. Adopting a single universal representation model reduces development overhead, simplifies model maintenance, and maintains higher accuracy across operations such as failure detection and demand planning. Ablation tests confirmed that every component—frequency mixing, hard negative generation, and joint reconstruction—is essential to these performance gains.

Organizations managing complex time series operations should pilot universal pre-training architectures like TimesURL to consolidate their predictive workflows and reduce labeling costs. Before enterprise-scale deployment, practitioners should evaluate computational trade-offs, as synthesizing hard negative samples and running joint reconstruction objectives increases training overhead relative to simpler single-objective models. While confidence in the benchmark performance is high across standardized environments, validation on proprietary real-time streaming data remains a prudent next step.

arXiv: 2312.15709
Cover for TimesURL: Self-Supervised Contrastive Learning for Universal Time Series Representation Learning

Abstract

Learning universal time series representations applicable to various types of downstream tasks is challenging but valuable in real applications. Recently, researchers have attempted to leverage the success of self-supervised contrastive learning (SSCL) in Computer Vision(CV) and Natural Language Processing(NLP) to tackle time series representation. Nevertheless, due to the special temporal characteristics, relying solely on empirical guidance from other domains may be ineffective for time series and difficult to adapt to multiple downstream tasks. To this end, we review three parts involved in SSCL including 1) designing augmentation methods for positive pairs, 2) constructing (hard) negative pairs, and 3) designing SSCL loss. For 1) and 2), we find that unsuitable positive and negative pair construction may introduce inappropriate inductive biases, which neither preserve temporal properties nor provide sufficient discriminative features. For 3), just exploring segment- or instance-level semantics information is not enough for learning universal representation. To remedy the above issues, we propose a novel self-supervised framework named TimesURL. Specifically, we first introduce a frequency-temporal-based augmentation to keep the temporal property unchanged. And then, we construct double Universums as a special kind of hard negative to guide better contrastive learning. Additionally, we introduce time reconstruction as a joint optimization objective with contrastive learning to capture both segment-level and instance-level information. As a result, TimesURL can learn high-quality universal representations and achieve state-of-the-art performance in 6 different downstream tasks, including short- and long-term forecasting, imputation, classification, anomaly detection and transfer learning.

Table of Contents

  • Introduction
  • Related Work
  • Proposed TimesURL Framework
  • FTAug Method
  • Double Universum Learning
  • Contrastive Learning for Segment-level Information
  • Time Reconstruction for Instance-level Information
  • Experiments
  • Classification
  • Imputation
  • Short- and Long-Term Forecasting
  • Anomaly Detection
  • Transfer Learning
  • Ablation Study
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — TimesURL framework for universal time-series representations

    model/method

    TimesURL is a self-supervised framework designed to learn one time-series representation that can support forecasting, imputation, classification, anomaly detection, and transfer learning. Each input instance is xi∈RT×Fx_i\in\mathbb{R}^{T\times F}, where TT is the series length and FF is the feature dimension. A nonlinear encoder fθf_\theta maps it to a temporally ordered representation ri={ri,1,…,ri,T}r_i=\{r_{i,1},\ldots,r_{i,T}\} with ri,t∈RKr_{i,t}\in\mathbb{R}^{K}.

    For every training instance, TimesURL generates two augmented views using FTAug. It applies temporal- and instance-wise contrastive learning to the corresponding representations, augments the negative sets with two types of anchor-specific Universums through DualCon, and jointly trains a reconstruction decoder on randomly masked views. The contrastive objective is intended to preserve segment-level temporal information and distinguish different instances, while reconstruction supplies additional instance-level information. The workflow illustration on page 2 shows these three components—FTAug, DualCon, and Recon—and their two data flows.

  2. Knowl 2 — Frequency-temporal augmentation through cropping and frequency mixing

    algorithm

    FTAug constructs training-only augmented views while attempting to preserve temporal dependencies and contextual consistency. For an input series xix_i, it performs two operations:

    1. It randomly samples two overlapping temporal segments [a1,b1][a_1,b_1] and [a2,b2][a_2,b_2] satisfying 0<a1≤a2≤b1≤b2≤T0<a_1\leq a_2\leq b_1\leq b_2\leq T. Representations at the same timestamp in the overlap [a2,b1][a_2,b_1] are treated as positive pairs across the two contexts.

    2. For a randomly selected training instance xkx_k from the same batch, it computes the Fast Fourier Transform of xix_i and xkx_k, replaces a selected rate of the frequency components of xix_i with the corresponding components from xkx_k, and applies the inverse Fourier transform. This produces a frequency-mixed time-domain context without directly reversing values or permuting temporal order.

    FTAug is applied only during training. The combination of frequency mixing and overlapping random crops is intended to retain temporal variation, temporal dependence, and semantic consistency while still creating distinct positive views.

  3. Knowl 3 — Double Universums as anchor-specific hard negatives

    equation

    TimesURL creates two kinds of synthetic hard negatives in representation space. Let ri,t,ri,t′∈RKr_{i,t},r'_{i,t}\in\mathbb{R}^{K} be the representations at timestamp tt for two augmented views of instance ii. Let Ω\Omega be the timestamps in the overlap of the two cropped views, and let t′∈Ω∖{t}t'\in\Omega\setminus\{t\} be randomly selected. The temporal-wise Universums are

    ri,ttemp=λ1ri,t+(1−λ1)ri,t′,ri,t′temp=λ1ri,t′+(1−λ1)ri,t′′.r^{\mathrm{temp}}_{i,t}=\lambda_1r_{i,t}+(1-\lambda_1)r_{i,t'},\qquad r'^{\mathrm{temp}}_{i,t}=\lambda_1r'_{i,t}+(1-\lambda_1)r'_{i,t'}.

    For an instance j≠ij\neq i in the same batch, the instance-wise Universums are

    ri,tinst=λ2ri,t+(1−λ2)rj,t,ri,t′inst=λ2ri,t′+(1−λ2)rj,t′.r^{\mathrm{inst}}_{i,t}=\lambda_2r_{i,t}+(1-\lambda_2)r_{j,t},\qquad r'^{\mathrm{inst}}_{i,t}=\lambda_2r'_{i,t}+(1-\lambda_2)r'_{j,t}.

    The mixing coefficients λ1,λ2∈(0,0.5]\lambda_1,\lambda_2\in(0,0.5] are randomly selected, so the synthetic vector contains no more anchor contribution than negative-sample contribution. Temporal Universums mix an anchor timestamp with another timestamp from the same series; instance Universums mix an anchor with a different series at the same timestamp. Both are intended to lie close enough to the anchor to provide harder contrastive signals than ordinary distant negatives. On the ERing diagnostic shown on page 4, adding Universums reduced the proxy-task accuracy but increased linear-classification accuracy from 0.8960.896 to 0.9850.985, which the authors interpret as evidence that the proxy task became harder while the learned representation improved.

  4. Knowl 4 — Dual temporal- and instance-wise contrastive learning

    equation

    For a batch B\mathcal{B} of time-series instances, TimesURL uses the representation at timestamp tt in one view, ri,tr_{i,t}, as the anchor and ri,t′r'_{i,t} in the other view as its positive. The dot product u⋅vu\cdot v measures similarity between representation vectors. The temporal-wise loss is

    ℓtemp(i,t)=−log⁡exp⁡(ri,t⋅ri,t′)exp⁡(ri,t⋅ri,t′)+∑z∈Ni,ttempexp⁡(ri,t⋅z).\ell^{(i,t)}_{\mathrm{temp}}=-\log\frac{\exp(r_{i,t}\cdot r'_{i,t})}{\exp(r_{i,t}\cdot r'_{i,t})+\sum_{z\in\mathcal{N}^{\mathrm{temp}}_{i,t}}\exp(r_{i,t}\cdot z)}.

    Its negative set Ni,ttemp\mathcal{N}^{\mathrm{temp}}_{i,t} contains both views of the same instance at every other overlap timestamp s∈Ω∖{t}s\in\Omega\setminus\{t\} and the temporal Universums formed by mixing the two views at tt with representations at other timestamps in Ω\Omega.

    The instance-wise loss is

    ℓinst(i,t)=−log⁡exp⁡(ri,t⋅ri,t′)exp⁡(ri,t⋅ri,t′)+∑z∈Ni,tinstexp⁡(ri,t⋅z).\ell^{(i,t)}_{\mathrm{inst}}=-\log\frac{\exp(r_{i,t}\cdot r'_{i,t})}{\exp(r_{i,t}\cdot r'_{i,t})+\sum_{z\in\mathcal{N}^{\mathrm{inst}}_{i,t}}\exp(r_{i,t}\cdot z)}.

    Its negative set Ni,tinst\mathcal{N}^{\mathrm{inst}}_{i,t} contains both views at timestamp tt for every other batch instance j∈B∖{i}j\in\mathcal{B}\setminus\{i\}, together with the instance Universums obtained by mixing the anchor with those other instances. The combined DualCon objective is

    Ldual=1∣B∣T∑i∈B∑t(ℓtemp(i,t)+ℓinst(i,t)).L_{\mathrm{dual}}=\frac{1}{|\mathcal{B}|T}\sum_{i\in\mathcal{B}}\sum_t\left(\ell^{(i,t)}_{\mathrm{temp}}+\ell^{(i,t)}_{\mathrm{inst}}\right).

    TimesURL applies hierarchical contrastive learning by max-pooling representations along the time axis at multiple scales before applying these losses. The temporal loss emphasizes within-series variation, whereas the instance loss emphasizes sample-level discrimination.

  5. Knowl 5 — Masked time reconstruction and joint training objective

    equation

    TimesURL supplements contrastive learning with masked time reconstruction. For each batch instance ii, let xi∈RT×Fx_i\in\mathbb{R}^{T\times F} be the original series, x~i\tilde{x}_i its decoder reconstruction from a masked representation, and mi∈{0,1}T×Fm_i\in\{0,1\}^{T\times F} its observation mask. The paper defines mi,t=0m_{i,t}=0 when xi,tx_{i,t} is missing and mi,t=1m_{i,t}=1 when it is observed; xi′x'_i, x~i′\tilde{x}'_i, and mi′m'_i denote the corresponding augmented-view quantities. With elementwise multiplication ⊙\odot and Euclidean norm ∥⋅∥2\|\cdot\|_2, the reconstruction loss is

    Lrecon=12∣B∣∑i∈B(∥mi⊙(x~i−xi)∥22+∥mi′⊙(x~i′−xi′)∥22).L_{\mathrm{recon}}=\frac{1}{2|\mathcal{B}|}\sum_{i\in\mathcal{B}}\left(\left\|m_i\odot(\tilde{x}_i-x_i)\right\|_2^2+\left\|m'_i\odot(\tilde{x}'_i-x'_i)\right\|_2^2\right).

    The encoder is optimized jointly with DualCon using

    L=Ldual+αLrecon,L=L_{\mathrm{dual}}+\alpha L_{\mathrm{recon}},

    where α\alpha is the scalar weight balancing contrastive and reconstruction objectives. Random masking and reconstruction are intended to retain temporal variation that can be lost when representations are repeatedly max-pooled, thereby complementing the segment-level contrastive signal with instance-level information.

  6. Knowl 6 — Evaluation protocol across six downstream task settings

    experimental setup

    TimesURL uses a Temporal Convolutional Network as its backbone encoder and is evaluated as a representation-learning method rather than as a task-specific end-to-end model. The paper evaluates six settings: short-term forecasting, long-term forecasting, imputation, classification, anomaly detection, and transfer learning, comparing against approximately 15 baselines selected for the tasks they support.

    For classification, experiments use the UEA and UCR archives; TimesURL representations have dimension 320320 and are classified with an RBF-kernel SVM. For imputation, experiments use ETTh1, ETTh2, and ETTm1, with randomly missing time-point ratios of 12.5%12.5\%, 25%25\%, 37.5%37.5\%, and 50%50\%, followed by an MLP downstream model. For forecasting, ETT, Electricity, and Weather data are used; short-term horizons are 2424 and 4848, while long-term horizons range from 9696 to 720720, and one learned representation per dataset is evaluated with linear regression across horizons. For anomaly detection, the streaming protocol evaluates the final point of a prefix on KPI and Yahoo, with each series split chronologically into training and evaluation halves. For transfer learning, the encoder is trained on CBF or CinCECGTorso and evaluated on nine other target domains among the first ten UCR datasets.

  7. Knowl 7 — Classification performance on UEA and UCR archives

    data/table

    TimesURL gives the strongest average classification accuracy and average rank among the compared self-supervised and unsupervised methods. The comparison uses the same representation dimension and RBF-SVM protocol for all methods except DTW; higher average accuracy and lower average rank are better.

    Could not parse LaTeX table

    The reported results are 75.2%75.2\% average accuracy on the 30 UEA datasets and 84.5%84.5\% on the 128 UCR datasets. The authors attribute the advantage to the combination of FTAug, hard negatives, and joint segment- and instance-level learning.

  8. Knowl 8 — Multivariate imputation performance under increasing missingness

    data/table

    TimesURL is evaluated on ETTh1, ETTh2, and ETTm1 after randomly masking four proportions of time points. Mean squared error (MSE) and mean absolute error (MAE) are reported, with lower values being better. The full numerical comparison is:

    Could not parse LaTeX table

    The average scores favor TimesURL for both metrics, and the paper reports state-of-the-art performance across the three datasets. Performance remains competitive as the missingness ratio increases, supporting the claim that the reconstruction objective captures useful temporal patterns.

  9. Knowl 9 — Forecasting, anomaly-detection, and transfer performance

    empirical result

    TimesURL is reported to outperform the compared representation-learning and end-to-end forecasting methods in most short- and long-term forecasting cases. On the univariate ETT forecasting results, averaged over the listed datasets and horizons, the scores are:

    Could not parse LaTeX table

    For streaming anomaly detection, the reported results are:

    Could not parse LaTeX table

    For transfer learning, training on CBF and testing across nine other target domains gives average accuracy 0.8640.864; training on CinCECGTorso gives average accuracy 0.8950.895. The authors report competitive performance relative to the corresponding no-transfer setting, using these results as evidence that the learned representation transfers across conditions rather than only fitting its source domain.

  10. Knowl 10 — Ablation evidence for FTAug, both Universums, and reconstruction

    data/table

    An ablation study on 30 UEA datasets evaluates the contribution of each major TimesURL component using average classification accuracy. The complete model obtains 0.7520.752. Removing any component decreases performance:

    Could not parse LaTeX table

    Removing either the instance-wise or temporal-wise Universum is less damaging than removing the entire double-Universum design, but neither single Universum reaches the full-model result. The ablation therefore supports the paper's claim that frequency mixing, both hard-negative constructions, and joint reconstruction are complementary components of the universal representation.

Coverage note — Detailed per-horizon forecasting rows, proxy-task curves beyond the reported ERing accuracies, and appendix-only implementation details were omitted because the core method and the principal cross-task results are already represented.

References

  1. 1.Bagnall, A.; Dau, H. A.; Lines, J.; Flynn, M.; Large, J.; Bostrom, A.; Southam, P.; and Keogh, E. 2018. The UEA multivariate time series classification archive, 2018. arXiv preprint arXiv:1811.00075.
  2. 2.Bayer, J.; Soelch, M.; Mirchev, A.; Kayalibay, B.; and van der Smagt, P. 2020. Mind the Gap when Conditioning Amortised Inference in Sequential Latent-Variable Models. In International Conference on Learning Representations.
  3. 3.Cai, T.; Frankle, J.; Schwab, D. J.; and Morcos, A. S. 2020. Are all negatives created equal in contrastive instance discrimination? ArXiv, abs/2010.06682.
  4. 4.Chapelle, O.; Agarwal, A.; Sinz, F.; and Schölkopf, B. 2007. An analysis of inference with the universum. Advances in neural information processing systems, 20.
  5. 5.Chen, M.; Xu, Z.; Zeng, A.; and Xu, Q. 2023. FrAug: Frequency Domain Augmentation for Time Series Forecasting. arXiv preprint arXiv:2302.09292.
  6. 6.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, 1597–1607. PMLR.
  7. 7.Chung, J.; Kastner, K.; Dinh, L.; Goel, K.; Courville, A. C.; and Bengio, Y. 2015. A recurrent latent variable model for sequential data. Advances in neural information processing systems, 28.
  8. 8.Dau, H. A.; Bagnall, A.; Kamgar, K.; Yeh, C.-C. M.; Zhu, Y.; Gharghabi, S.; Ratanamahatana, C. A.; and Keogh, E. 2019. The UCR time series archive. IEEE/CAA Journal of Automatica Sinica, 6(6): 1293–1305.
  9. 9.Dempster, A.; Petitjean, F.; and Webb, G. I. 2020. ROCKET: exceptionally fast and accurate time series classification using random convolutional kernels. Data Mining and Knowledge Discovery, 34(5): 1454–1495.
  10. 10.Denton, E. L.; et al. 2017. Unsupervised learning of disentangled representations from video. Advances in neural information processing systems, 30.
  11. 11.Eldele, E.; Ragab, M.; Chen, Z.; Wu, M.; Kwoh, C.; Li, X.; and Guan, C. 2021. Time-Series Representation Learning via Temporal and Contextual Contrasting. In International Joint Conference on Artificial Intelligence.
  12. 12.Eldele, E.; Ragab, M.; Chen, Z.; Wu, M.; Kwoh, C. K.; Li, X.; and Guan, C. 2022. Self-supervised Contrastive Representation Learning for Semi-supervised Time-Series Classification. arXiv preprint arXiv:2208.06616.
  13. 13.Grill, J.-B.; Strub, F.; Altché, F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.; Gheshlaghi Azar, M.; et al. 2020. Bootstrap your own latent—a new approach to self-supervised learning. Advances in neural information processing systems, 33: 21271–21284.
  14. 14.Gutmann, M. U.; and Hyvärinen, A. 2012. Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics. Journal of machine learning research, 13(2).
  15. 15.Han, A.; and Chen, S. 2023. Universum-Inspired Supervised Contrastive Learning. In Web and Big Data: 6th International Joint Conference, APWeb-WAIM 2022, Nanjing, China, November 25–27, 2022, Proceedings, Part II, 459–473. Springer.
  16. 16.He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16000–16009.
  17. 17.Kalantidis, Y.; Sariyildiz, M. B.; Pion, N.; Weinzaepfel, P.; and Larlus, D. 2020. Hard negative mixing for contrastive learning. Advances in Neural Information Processing Systems, 33: 21798–21809.
  18. 18.Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT, 4171–4186.
  19. 19.Krishnan, R.; Shalit, U.; and Sontag, D. 2017. Structured inference networks for nonlinear state space models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31.
  20. 20.Lei, Q.; Yi, J.; Vaculin, R.; Wu, L.; and Dhillon, I. S. 2019. Similarity Preserving Representation Learning for Time Series Clustering. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, IJCAI’19, 2845–2851. AAAI Press. ISBN 9780999241141.
  21. 21.Li, S.; Jin, X.; Xuan, Y.; Zhou, X.; Chen, W.; Wang, Y.-X.; and Yan, X. 2019. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. Advances in neural information processing systems, 32.
  22. 22.Liu, J.; and Chen, S. 2019. Non-stationary multivariate time series prediction with selective recurrent neural networks. In Pacific rim international conference on artificial intelligence, 636–649. Springer.
  23. 23.Liu, M.; Zeng, A.; Chen, M.; Xu, Z.; Lai, Q.; Ma, L.; and Xu, Q. 2022a. Scinet: Time series modeling and forecasting with sample convolution and interaction. Advances in Neural Information Processing Systems, 35: 5816–5828.
  24. 24.Liu, Y.; and wei Liu, J. 2022. The Time-Sequence Prediction via Temporal and Contextual Contrastive Representation Learning. In Pacific Rim International Conference on Artificial Intelligence.
  25. 25.Liu, Y.; Wu, H.; Wang, J.; and Long, M. 2022b. Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting. In Advances in Neural Information Processing Systems.
  26. 26.Luo, D.; Cheng, W.; Wang, Y.; Xu, D.; Ni, J.; Yu, W.; Zhang, X.; Liu, Y.; Chen, Y.; Chen, H.; et al. 2023. Time Series Contrastive Learning with Information-Aware Augmentations. In Proceedings of the AAAI Conference on Artificial Intelligence.
  27. 27.Ma, Q.; Zheng, J.; Li, S.; and Cottrell, G. W. 2019. Learning representations for time series clustering. Advances in neural information processing systems, 32.
  28. 28.Malhotra, P.; TV, V.; Vig, L.; Agarwal, P.; and Shroff, G. 2017. TimeNet: Pre-trained deep recurrent neural network for time series classification. arXiv preprint arXiv:1706.08838.
  29. 29.Nikolay Laptev, Y. B., Saeed Amizadeh. 2015. A Benchmark Dataset for Time Series Anomaly Detection. https://yahooresearch.tumblr.com/post/114590420346/a-benchmark-dataset-for-time-series-anomaly.
  30. 30.Oreshkin, B. N.; Carpov, D.; Chapados, N.; and Bengio, Y. 2019. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. arXiv preprint arXiv:1905.10437.
  31. 31.Pagliardini, M.; Gupta, P.; and Jaggi, M. 2017. Unsupervised learning of sentence embeddings using compositional n-gram features. arXiv preprint arXiv:1703.02507.
  32. 32.Ren, H.; Xu, B.; Wang, Y.; Yi, C.; Huang, C.; Kou, X.; Xing, T.; Yang, M.; Tong, J.; and Zhang, Q. 2019. Time-series anomaly detection service at microsoft. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, 3009–3017.
  33. 33.Robinson, J.; Chuang, C.-Y.; Sra, S.; and Jegelka, S. 2020. Contrastive learning with hard negative samples. arXiv preprint arXiv:2010.04592.
  34. 34.Shi, X.; Chen, Z.; Wang, H.; Yeung, D.-Y.; Wong, W.-K.; and Woo, W.-c. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Advances in neural information processing systems, 28.
  35. 35.Siffer, A.; Fouque, P.-A.; Termier, A.; and Largouet, C. 2017. Anomaly detection in streams with extreme value theory. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1067–1075.
  36. 36.Tonekaboni, S.; Eytan, D.; and Goldenberg, A. 2021. Unsupervised Representation Learning for Time Series with Temporal Neighborhood Coding. In International Conference on Learning Representations.
  37. 37.Um, T. T.; Pfister, F. M.; Pichler, D.; Endo, S.; Lang, M.; Hirche, S.; Fietzek, U.; and Kulic, D. 2017. Data augmentation of wearable sensor data for parkinson’s disease monitoring using convolutional neural networks. In Proceedings of the 19th ACM international conference on multimodal interaction, 216–220.
  38. 38.Vapnik, V. 2006. Transductive Inference and Semi-Supervised Learning. Semi-Supervised Learning, 453–472.
  39. 39.Wang, X.; and Gupta, A. 2015. Unsupervised learning of visual representations using videos. In Proceedings of the IEEE international conference on computer vision, 2794–2802.
  40. 40.Woo, G.; Liu, C.; Sahoo, D.; Kumar, A.; and Hoi, S. 2022. CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting. In International Conference on Learning Representations.
  41. 41.Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; and Long, M. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In The Eleventh International Conference on Learning Representations.
  42. 42.Wu, H.; Xu, J.; Wang, J.; and Long, M. 2021. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34: 22419–22430.
  43. 43.Xu, H.; Chen, W.; Zhao, N.; Li, Z.; Bu, J.; Li, Z.; Liu, Y.; Zhao, Y.; Pei, D.; Feng, Y.; et al. 2018. Unsupervised anomaly detection via variational auto-encoder for seasonal kpis in web applications. In Proceedings of the 2018 world wide web conference, 187–196.
  44. 44.Xu, L.; Lian, J.; Zhao, W. X.; Gong, M.; Shou, L.; Jiang, D.; Xie, X.; and Wen, J.-R. 2022. Negative sampling for contrastive representation learning: A review. arXiv preprint arXiv:2206.00212.
  45. 45.Yue, Z.; Wang, Y.; Duan, J.; Yang, T.; Huang, C.; Tong, Y.; and Xu, B. 2022. Ts2vec: Towards universal representation of time series. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 8980–8987.
  46. 46.Zerveas, G.; Jayaraman, S.; Patel, D.; Bhamidipaty, A.; and Eickhoff, C. 2021. A transformer-based framework for multivariate time series representation learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 2114–2124.
  47. 47.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, 11106–11115.

Citation

MLA
Liu, J., and S. Chen. “TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning”. arXiv, 2023, http://arxiv.org/abs/2312.15709v1.
APA
Liu, J., & Chen, S. (2023). TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning. arXiv. http://arxiv.org/abs/2312.15709v1
Chicago
Liu, J., and S. Chen. 2023. “TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning”. arXiv. http://arxiv.org/abs/2312.15709v1.
Harvard
Liu, J. and Chen, S. (2023) “TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.15709v1.
Vancouver
1. Liu J, Chen S (2023) TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning. arXiv

BibTeX

@article{liu2023timesurl,
  title = {TimesURL: Self-supervised Contrastive Learning for Universal Time Series Representation Learning},
  author = {Liu, Jiexi and Chen, Songcan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.15709v1},
  eprint = {2312.15709}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF