TimesURL: Self-Supervised Contrastive Learning for Universal Time Series Representation Learning
Jiexi LiuSongcan Chen
Proposes TimesURL, a self-supervised framework combining frequency-temporal data augmentations, synthesized Universum hard negatives, and a joint reconstruction objective to learn universal time series representations that achieve state-of-the-art results across six distinct downstream tasks.
Time series data is essential across critical domains such as weather forecasting, economic planning, industrial monitoring, and healthcare diagnostics. While self-supervised contrastive learning—a machine learning technique that trains models on unlabelled data by comparing similar and dissimilar samples—has succeeded in computer vision and natural language processing, directly transferring these techniques to time series has proven problematic. Existing approaches often rely on data modifications that distort underlying temporal relationships, fail to provide sufficiently challenging negative examples during training, and focus narrowly on either fine-grained point predictions or high-level whole-series patterns rather than addressing both.
The article aims to resolve these limitations by introducing TimesURL, a novel self-supervised framework designed to learn universal, task-agnostic time series representations that perform effectively across diverse downstream applications. The authors evaluate this framework against approximately 15 existing specialized baseline models across six major analytical tasks: short- and long-term forecasting, missing value imputation, classification, anomaly detection, and transfer learning.
To achieve this, the authors designed a three-part framework evaluated across broad benchmark suites, including 158 multivariate and univariate classification datasets as well as standard energy, weather, and web performance datasets. First, they implemented a hybrid augmentation method combining frequency mixing and random temporal cropping to preserve core trends without adding artificial noise. Second, they synthesized artificial hard negative samples—termed double Universums—across time and instance dimensions by blending anchor samples with non-matching segments to make training contrast more effective. Third, they jointly trained the network on both contrastive loss and masked time series reconstruction, capturing both fine-grained segment details and broad instance-level context.
The experimental findings show that TimesURL consistently establishes a new state of the art. In time series classification, TimesURL achieved 75.2% average accuracy across 30 multivariate benchmarks (a 3.8 percentage point improvement over the prior leading method) and 84.5% across 128 univariate datasets. In missing data imputation across multiple missingness ratios, it achieved the lowest overall error rates, yielding an average mean squared error of 1.326. For short- and long-term forecasting, TimesURL outperformed both self-supervised and dedicated end-to-end models across most forecast horizons. In anomaly detection and transfer learning scenarios, the model delivered top-tier precision, recall, and cross-dataset adaptability.
These results demonstrate that organizations do not need separate, specialized feature extraction pipelines for each time series application. Adopting a single universal representation model reduces development overhead, simplifies model maintenance, and maintains higher accuracy across operations such as failure detection and demand planning. Ablation tests confirmed that every component—frequency mixing, hard negative generation, and joint reconstruction—is essential to these performance gains.
Organizations managing complex time series operations should pilot universal pre-training architectures like TimesURL to consolidate their predictive workflows and reduce labeling costs. Before enterprise-scale deployment, practitioners should evaluate computational trade-offs, as synthesizing hard negative samples and running joint reconstruction objectives increases training overhead relative to simpler single-objective models. While confidence in the benchmark performance is high across standardized environments, validation on proprietary real-time streaming data remains a prudent next step.
- Paper: TS2Vec: Towards Universal Representation of Time Series, Zhihan Yue et al. (2022). TS2Vec established foundational self-supervised contrastive learning across contextual scales for universal time series representations, providing the core baseline and problem formulation directly addressed by TimesURL.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). TimesNet demonstrates how multi-periodicity and 2D variation modeling can achieve task-general representations across standard time series downstream workloads.
- Paper: Representation Learning with Contrastive Predictive Coding, Aäron van den Oord et al. (2018). Contrastive Predictive Coding introduces the essential principles of self-supervised contrastive learning and latent negative sampling for temporal and sequential data.
- Paper: Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere, Tongzhou Wang et al. (2020). This paper establishes the foundational geometric theory of alignment and uniformity in contrastive representation learning that informs modern contrastive objective design.
- Paper: A Transformer-based Framework for Multivariate Time Series Representation Learning, George Zerveas et al. (2020). This work introduces unsupervised pre-training and reconstruction objectives for multivariate time series representations, directly preceding joint reconstruction-contrastive designs.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). MOIRAI extends universal time series representation principles to large-scale masked encoder foundation models that achieve zero-shot forecasting across diverse domains and frequencies.
- Paper: Timer: Generative Pre-trained Transformers Are Large Time Series Models, Yong Liu et al. (2024). Timer advances general time series pre-training by developing large autoregressive Transformer foundation models capable of unified multi-task generalization.
- Paper: Sundial: A Family of Highly Capable Time Series Foundation Models, Yong Liu 0007 et al. (2025). Sundial scales universal representation and forecasting paradigms by integrating continuous flow-matching generative modeling into time series foundation models.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). This position paper explores how universal time series representations and pre-trained paradigms intersect with the general reasoning capabilities of large language models.
- Paper: GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting, Furong Jia et al. (2024). GPT4MTS builds on time series forecasting representations by fusing numerical sequence modeling with multimodal text prompts and large language models.
