Unified Training of Universal Time Series Forecasting Transformers

Gerald WooChenghao LiuAkshat KumarCaiming XiongSilvio SavareseDoyen Sahoo

article2024ICML505 citations

Presents Moirai, a universal time series transformer trained on the 27-billion-observation LOTSA archive, which handles arbitrary variate counts and frequencies to match or outperform dataset-specific models in zero-shot forecasting.

Listen

Deep learning for time series forecasting has traditionally required training a specialized, individual model for every dataset and forecasting task. This one-model-per-dataset approach prevents organizations from exploiting large pre-trained foundation models that can generalize across varied scenarios, leading to high training costs and slow adaptation to new business problems. Constructing a single universal model is difficult because time series data are highly heterogeneous: recording frequencies vary widely, multivariate series contain arbitrary numbers of variates, and data distributions differ significantly across domains.

The article designs and demonstrates MOIRAI, a universal time series forecasting model powered by a masked encoder architecture. Its core objective is to evaluate whether a single pre-trained model can accurately forecast across diverse frequencies, dimensions, and domains without task-specific training, operating as an off-the-shelf zero-shot forecaster.

The approach introduces three architectural mechanisms to resolve data heterogeneity: multi-patch-size projection layers that scale patch lengths to recording frequencies, an any-variate attention mechanism that processes flattened multivariate sequences using positional embeddings and attention biases, and a mixture of parametric distributions that models positive, skewed, and heavy-tailed data. To train MOIRAI, the authors compiled the Large-scale Open Time Series Archive (LOTSA), an open repository spanning nine domains and over 27 billion observations. MOIRAI was trained across three sizes (14 million, 91 million, and 311 million parameters) using random sampling of context and horizon windows alongside sequence packing to maximize throughput. Evaluations spanned in-distribution benchmark tests as well as zero-shot evaluations on unseen datasets for both probabilistic and long-horizon forecasting.

The investigation produced four central findings. First, across standard in-distribution benchmark datasets, MOIRAI outperformed all classical, statistical, and specialized deep learning baselines while operating as a single unified system. Second, in zero-shot probabilistic forecasting across diverse domains like energy, climate, and sales, MOIRAI matched or beat specialized models that were custom-trained directly on those target datasets. Third, on long-horizon forecasting tasks, MOIRAI generated competitive or superior mean squared error metrics relative to state-of-the-art models without requiring local fine-tuning. Fourth, architectural ablations showed that removing multi-frequency patching, any-variate attention, mixture distributions, or large-scale data diversity caused severe performance degradation, while sequence packing reduced token padding from approximately 61% to less than 0.4% and improved model performance by 16% at equal compute.

These results demonstrate that universal time series models are commercially viable and can replace fragmented, dataset-specific modeling pipelines. Shifting toward large pre-trained forecasters amortizes initial training costs across thousands of downstream tasks, cuts continuous model maintenance expenses, and reduces operational deployment time from days to milliseconds per inference. Moreover, producing principled probabilistic outputs allows leaders to quantify operational risks and uncertainty far more reliably than deterministic point forecasters.

Organizations evaluating forecasting architecture should consider adopting pre-trained universal forecasters for new deployment environments to avoid expensive, bespoke model pipelines. For existing deployments, teams should run pilot side-by-side evaluations comparing zero-shot universal models against fine-tuned specialized models on high-value business metrics. Continued work should explore hyperparameter optimization, the integration of multimodal inputs such as tabular features and text, and the establishment of formal scaling laws across larger parameter counts.

Confidence in these findings is high for standard operational frequencies and horizons up to several thousand steps, as verified by extensive multi-domain testing. However, readers should note limitations: the model relies on heuristic rules for frequency-to-patch mapping, exhibits limited scaling gains from the base to the largest parameter size on select long-sequence tasks, and faces token length constraints when handling extremely high-dimensional multivariate series.

No sufficiently relevant recommendations were found.

Cover for Unified Training of Universal Time Series Forecasting Transformers

Abstract

Deep learning for time series forecasting has traditionally operated within a one-model-per-dataset framework, limiting its potential to leverage the game-changing impact of large pre-trained models. The concept of universal forecasting, emerging from pre-training on a vast collection of time series datasets, envisions a single Large Time Series Model capable of addressing diverse downstream forecasting tasks. However, constructing such a model poses unique challenges specific to time series data: i) cross-frequency learning, ii) accommodating an arbitrary number of variates for multivariate time series, and iii) addressing the varying distributional properties inherent in large-scale data. To address these challenges, we present novel enhancements to the conventional time series Transformer architecture, resulting in our proposed Masked Encoder-based Universal Time Series Forecasting Transformer (Moirai). Trained on our newly introduced Large-scale Open Time Series Archive (LOTSA) featuring over 27B observations across nine domains, Moirai achieves competitive or superior performance as a zero-shot forecaster when compared to full-shot models. Code, data, and model weights can be found at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Architecture
  • 3.1.1 Multi Patch Size Projection Layers
  • 3.1.2 Any-variate Attention
  • 3.1.3 Mixture Distribution
  • 3.2 Unified Training
  • 3.2.1 LOTSA Data
  • 3.2.2 Pre-training
  • 4 Experiments
  • 4.1 In-distribution Forecasting
  • 4.2 Out-of-distribution / Zero-shot Forecasting
  • 4.3 Ablation Study
  • 4.4 Further Analysis
  • 5 Conclusion
  • References
  • A Large-scale Open Time Series Archive
  • B Moirai Architecture Details
  • B.1 Multi Patch Size Projection Layers
  • B.2 Mixture Distribution
  • B.3 Discussion on “Flexible Distribution”
  • C Probabilistic Forecasting
  • C.1 Evaluation Metrics
  • C.2 Evaluation Setup
  • C.3 Baselines
  • D Full Experimental Results
  • D.1 In-distribution Forecasting: Monash Time Series Forecasting Benchmark
  • D.2 Out-of-distribution Forecasting: Probabilistic Forecasting
  • D.3 Out-of-distribution Forecasting: Long Sequence Forecasting
  • D.4 Computation Costs
  • E Forecast Visualizations

Knowls

  1. Knowl 1 — MOIRAI Architecture and Universal Time Series Forecasting Formulation

    model/method

    The Masked EncOder-based UnIveRsAl TIme Series Forecasting Transformer (MOIRAI) is a foundation model designed for universal time series forecasting across arbitrary frequencies, variable counts, and forecast horizons.

    Let D={(Y(i),Z(i))}i=1N\mathcal{D} = \{(\mathbf{Y}^{(i)}, \mathbf{Z}^{(i)})\}_{i=1}^N be a dataset where Y(i)=(y1(i),…,yTi(i))∈Rdyi×Ti\mathbf{Y}^{(i)} = (\mathbf{y}_1^{(i)}, \dots, \mathbf{y}_{T_i}^{(i)}) \in \mathbb{R}^{d_{y_i} \times T_i} denotes a target time series with dyid_{y_i} variates and TiT_i time steps, and Z(i)=(z1(i),…,zTi(i))∈Rdzi×Ti\mathbf{Z}^{(i)} = (\mathbf{z}_1^{(i)}, \dots, \mathbf{z}_{T_i}^{(i)}) \in \mathbb{R}^{d_{z_i} \times T_i} represents past or dynamic covariates. The training objective maximizes the log-likelihood of the target future values under a predictive mixture distribution parametrized by ϕ^\hat{\boldsymbol{\phi}}:

    max⁡θE(Y,Z)∼p(D),(t,l,h)∼p(T∣D)log⁡p(Yt:t+h∣ϕ^),s.t. ϕ^=fθ(Yt−l:t,Zt−l:t+h)\max_{\boldsymbol{\theta}} \mathbb{E}_{(\mathbf{Y}, \mathbf{Z}) \sim p(\mathcal{D}), (t, l, h) \sim p(\mathcal{T}|\mathcal{D})} \log p(\mathbf{Y}_{t:t+h} \mid \hat{\boldsymbol{\phi}}), \quad \text{s.t. } \hat{\boldsymbol{\phi}} = f_{\boldsymbol{\theta}}(\mathbf{Y}_{t-l:t}, \mathbf{Z}_{t-l:t+h})

    where ll is the context lookback length, hh is the prediction horizon, and fθf_{\boldsymbol{\theta}} is the neural network.

    MOIRAI operates by:

    1. Flattening multivariate time series across all variate dimensions into a single sequence of 1D series.
    2. Partitioning each variate sequence into non-overlapping patches of size PP.
    3. Linearly projecting input patches into a hidden dimension dmodeld_{\text{model}} using frequency-specialized patch projection layers.
    4. Replacing target patches located in the prediction horizon t:t+ht:t+h with a learnable [mask] token embedding.
    5. Processing the sequence through an encoder-only Transformer with pre-RMSNorm, Query-Key normalization, SwiGLU activations in the feed-forward layers, and no additive bias terms in any layer.
    6. Decoding masked output representations via a linear output projection to produce the parameters ϕ^\hat{\boldsymbol{\phi}} of a mixture distribution over the forecast horizon.
  2. Knowl 2 — Any-Variate Attention Mechanism with Rotary Position Embeddings and Binary Variate Biases

    model/method

    Any-variate Attention enables MOIRAI to handle time series with arbitrary numbers of variates in a single sequence while preserving permutation equivariance with respect to variate ordering and permutation invariance with respect to variate indices.

    A multivariate time series is flattened into a 1D token sequence where each token is indexed by its time step ii and its variate channel mm. For query token (i,m)(i, m) and key token (j,n)(j, n), the attention logit Eij,mnE_{ij,mn} before softmax normalization is computed as:

    Eij,mn=(WQxi,m)⊤Ri−j(WKxj,n)+u(1)⋅1{m=n}+u(2)⋅1{m≠n}E_{ij,mn} = (\mathbf{W}^Q \mathbf{x}_{i,m})^\top \mathbf{R}_{i-j} (\mathbf{W}^K \mathbf{x}_{j,n}) + u^{(1)} \cdot \mathbf{1}_{\{m=n\}} + u^{(2)} \cdot \mathbf{1}_{\{m \neq n\}}

    Aij,mn=exp⁡(Eij,mn)∑k,oexp⁡(Eik,mo)\mathbf{A}_{ij,mn} = \frac{\exp(E_{ij,mn})}{\sum_{k, o} \exp(E_{ik,mo})}

    where:

    • xi,m,xj,n∈Rdmodel\mathbf{x}_{i,m}, \mathbf{x}_{j,n} \in \mathbb{R}^{d_{\text{model}}} are input patch representations at time-variate coordinates (i,m)(i, m) and (j,n)(j, n).
    • WQ,WK∈Rdh×dmodel\mathbf{W}^Q, \mathbf{W}^K \in \mathbb{R}^{d_h \times d_{\text{model}}} are query and key projection matrices for an attention head of dimension dhd_h.
    • Ri−j∈Rdh×dh\mathbf{R}_{i-j} \in \mathbb{R}^{d_h \times d_h} is the Rotary Position Embedding (RoPE) matrix encoding the relative temporal distance i−ji - j.
    • 1{cond}\mathbf{1}_{\{\text{cond}\}} is the indicator function evaluating to 11 if the condition is true and 00 otherwise.
    • u(1),u(2)∈Ru^{(1)}, u^{(2)} \in \mathbb{R} are learnable scalar attention biases dedicated to each attention head in each Transformer layer, distinguishing intra-variate interactions (m=nm=n) from inter-variate interactions (m≠nm \neq n).
  3. Knowl 3 — Multi-Patch Size Projection Layers for Cross-Frequency Handling

    model/method

    To prevent negative interference across time series of varying sampling rates and manage the quadratic computational cost of attention, MOIRAI utilizes a set of frequency-dependent, multi-patch size linear projection layers. Instead of fixing a single patch size PP, the architecture learns distinct linear projection layers for five patch sizes: P∈{8,16,32,64,128}P \in \{8, 16, 32, 64, 128\}.

    Each input projection is a linear layer mapping RP→Rdmodel\mathbb{R}^P \to \mathbb{R}^{d_{\text{model}}}, and each output projection maps Rdmodel\mathbb{R}^{d_{\text{model}}} to the mixture distribution parameter space corresponding to PP time steps. Projection weights for a given patch size are shared across overlapping frequency categories.

    The mapping between data sampling frequency and allowable patch sizes is defined as follows:

    • Yearly, Quarterly: P∈{8}P \in \{8\}
    • Monthly: P∈{8,16,32}P \in \{8, 16, 32\}
    • Weekly, Daily: P∈{16,32}P \in \{16, 32\}
    • Hourly: P∈{32,64}P \in \{32, 64\}
    • Minute-level: P∈{32,64,128}P \in \{32, 64, 128\}
    • Second-level: P∈{64,128}P \in \{64, 128\}

    High-frequency time series use larger patch sizes to reduce token sequence length and retain long context windows, whereas low-frequency time series use smaller patch sizes to shift computation from linear projection layers into Transformer attention blocks.

  4. Knowl 4 — Four-Component Mixture Distribution Head for Probabilistic Forecasting

    model/method

    MOIRAI models the predictive distribution p(Yt:t+h∣ϕ^)p(\mathbf{Y}_{t:t+h} \mid \hat{\boldsymbol{\phi}}) as a mixture of C=4C=4 parametric distributions:

    p(Yt:t+h∣ϕ^)=∑c=14wc pc(Yt:t+h∣ϕ^c)p(\mathbf{Y}_{t:t+h} \mid \hat{\boldsymbol{\phi}}) = \sum_{c=1}^4 w_c \, p_c(\mathbf{Y}_{t:t+h} \mid \hat{\boldsymbol{\phi}}_c)

    where mixture weights w=(w1,w2,w3,w4)\mathbf{w} = (w_1, w_2, w_3, w_4) are constrained to the probability simplex via a softmax activation, and ϕ^c\hat{\boldsymbol{\phi}}_c denotes the parameter set for component cc.

    The four components are:

    1. Student's tt-distribution: Models heavy-tailed and general time series. p(x;ν,μ,τ)=Γ(ν+12)Γ(ν2)πντ(1+1ν(x−μτ)2)−ν+12p(x; \nu, \mu, \tau) = \frac{\Gamma\left(\frac{\nu+1}{2}\right)}{\Gamma\left(\frac{\nu}{2}\right)\sqrt{\pi \nu} \tau} \left(1 + \frac{1}{\nu}\left(\frac{x - \mu}{\tau}\right)^2\right)^{-\frac{\nu+1}{2}} with predicted location μ∈R\mu \in \mathbb{R}, scale τ>0\tau > 0 (enforced via softplus), and degrees of freedom ν>2\nu > 2 (enforced via ν=2+softplus(⋅)\nu = 2 + \text{softplus}(\cdot) to ensure defined variance).

    2. Log-Normal Distribution: Models non-negative, right-skewed data. p(x;μ,σ)=1xσ2πexp⁡(−(ln⁡x−μ)22σ2)p(x; \mu, \sigma) = \frac{1}{x \sigma \sqrt{2\pi}} \exp\left(-\frac{(\ln x - \mu)^2}{2\sigma^2}\right) with predicted parameters μ∈R\mu \in \mathbb{R} and σ>0\sigma > 0 (enforced via softplus).

    3. Continuous Negative Binomial Distribution: Models non-negative count data. p(x;r,p)∝Γ(x+r)Γ(x+1)Γ(r)(1−p)rpxp(x; r, p) \propto \frac{\Gamma(x + r)}{\Gamma(x + 1)\Gamma(r)} (1 - p)^r p^x with dispersion r>0r > 0 (via softplus) and probability p∈[0,1]p \in [0, 1] (via sigmoid).

    4. Low-Variance Normal Distribution: Models high-confidence predictions. p(x;μ,σ)=1σ2πexp⁡(−(x−μ)22σ2)p(x; \mu, \sigma) = \frac{1}{\sigma \sqrt{2\pi}} \exp\left(-\frac{(x - \mu)^2}{2\sigma^2}\right) where mean μ∈R\mu \in \mathbb{R} is predicted and standard deviation is fixed to a constant σ=10−3\sigma = 10^{-3}.

  5. Knowl 5 — Large-Scale Open Time Series Archive (LOTSA)

    data/table

    The Large-scale Open Time Series Archive (LOTSA) is a curated collection of open-source time series datasets formatted in Apache Arrow for pre-training large time series models. LOTSA contains 27,646,462,733 (~27.6B) total observations across 9 domains (231,082,956,489231,082,956,489 variate-observations when counting individual channels).

    Domain # Datasets # Observations % of Total
    Energy 30 16,358,600,896 59.17%
    Transport 23 4,900,453,419 17.73%
    Climate 6 4,188,011,890 15.15%
    CloudOps 3 1,518,268,292 5.49%
    Web 3 428,082,373 1.55%
    Sales 6 197,984,339 0.72%
    Nature 5 28,547,647 0.09%
    Econ/Fin 23 24,919,596 0.10%
    Healthcare 6 1,594,281 0.01%
    Frequency Category # Datasets # Observations % of Total
    Yearly 4 873,297 0.003%
    Quarterly 5 2,312,027 0.008%
    Monthly 10 11,040,648 0.040%
    Weekly 7 18,481,871 0.067%
    Daily 21 709,017,118 2.565%
    (Multi) Hourly 31 19,875,993,973 71.893%
    (Multi) Minute-level 25 7,013,949,430 25.370%
    (Multi) Second-level 2 14,794,369 0.054%

    Data sources include BuildingsBench, ClimateLearn (ERA5, CMIP6), CloudOps TSF, GluonTS, LargeST, LibCity, Monash Time Series Repository, ProEnFo, and SubseasonalClimateUSA.

  6. Knowl 6 — LOTSA Pre-Training Data Sampling, Task Formulation, and Sequence Packing

    model/method

    MOIRAI pre-training uses a two-stage sampling procedure, dynamic task length assignment, and sequence packing:

    1. Sub-Dataset Weight Capping: Given KK constituent sub-datasets in LOTSA where ∣Dk∣=∑i∈DkTi|\mathcal{D}_k| = \sum_{i \in \mathcal{D}_k} T_i, proportional sampling is bounded by threshold ϵ=0.001\epsilon = 0.001 to prevent single-domain domination: ωk=min⁡(∣Dk∣∑i=1K∣Di∣,ϵ),p(Dk)=ωk∑j=1Kωj\omega_k = \min\left(\frac{|\mathcal{D}_k|}{\sum_{i=1}^K |\mathcal{D}_i|}, \epsilon\right), \quad p(\mathcal{D}_k) = \frac{\omega_k}{\sum_{j=1}^K \omega_j} Within sub-dataset Dk\mathcal{D}_k, individual time series are sampled in proportion to their length TiT_i.

    2. Task and Window Sampling: For each sampled series, a total token sequence length (across all variates) is cropped uniformly between a minimum length of 2 tokens per variate and a maximum total length of 512 tokens. The prediction horizon length hh is chosen as a uniform random fraction in [0.15,0.5][0.15, 0.5] of the sampled window.

    3. Variate Augmentation: Multivariate time series are constructed by subsampling or concatenating series from the same dataset. The number of variates is drawn from a Beta-Binomial distribution with parameters n=128,a=2,b=5n=128, a=2, b=5 (maximum 128 variates, expected mean ≈37\approx 37).

    4. Sequence Packing: Multiple truncated and flattened sequences are packed together into a single fixed-length Transformer context without cross-sample attention. This reduces padding token overhead from 61.08% of total tokens to 0.38%.

  7. Knowl 7 — MOIRAI Model Configurations and Pre-Training Hyperparameters

    experimental setup

    MOIRAI is instantiated in three model capacities:

    Model Layers dmodeld_{\text{model}} dffd_{\text{ff}} Heads dkvd_{kv} Parameters
    MOIRAISmall_{\text{Small}} 6 384 1536 6 64 14M
    MOIRAIBase_{\text{Base}} 12 768 3072 12 64 91M
    MOIRAILarge_{\text{Large}} 24 1024 4096 16 64 311M

    Training Setup:

    • Optimizer: AdamW with lr=1×10−3\text{lr} = 1\times 10^{-3}, weight decay =0.1= 0.1, β1=0.9\beta_1 = 0.9, β2=0.98\beta_2 = 0.98.
    • Learning rate schedule: Linear warmup for the initial 10,000 steps followed by cosine annealing.
    • Batch size: 256 packed sequence batches.
    • Training duration: 100,000 optimization steps for MOIRAISmall_{\text{Small}}; 1,000,000 optimization steps for MOIRAIBase_{\text{Base}} and MOIRAILarge_{\text{Large}}.
    • Hardware and precision: NVIDIA A100-40GB GPUs using TensorFloat-32 (TF32).
  8. Knowl 8 — Zero-Shot Out-of-Distribution Probabilistic Forecasting Performance

    empirical result

    MOIRAI was evaluated zero-shot against fully supervised, hyperparameter-tuned deep learning models (PatchTST, TiDE, TFT, DeepAR) and statistical baselines (AutoARIMA, Seasonal Naive) on 6 unseen target datasets. Performance was measured by Continuous Ranked Probability Score (CRPS, lower is better) and Mean Scaled Interval Score (MSIS at α=0.05\alpha=0.05, lower is better) across rolling evaluation windows.

    Dataset Zero-shot MOIRAI Full-shot Deep Baselines Baseline
    Metric Small Base Large PatchTST TiDE TFT DeepAR S. Naive
    Electricity CRPS 0.072 0.055 0.050 0.052 0.048 0.050 0.065 0.070
    MSIS 7.999 6.172 5.875 5.744 5.672 6.278 6.893 35.251
    Solar CRPS 0.471 0.419 0.406 0.518 0.420 0.446 0.431 0.512
    MSIS 8.425 7.011 6.250 8.447 13.754 8.057 11.181 48.130
    Walmart CRPS 0.103 0.093 0.098 0.082 0.077 0.087 0.121 0.151
    MSIS 9.371 8.421 8.520 6.005 6.258 8.718 12.502 49.458
    Weather CRPS 0.049 0.041 0.051 0.059 0.054 0.043 0.132 0.068
    MSIS 5.236 5.136 4.962 7.759 8.095 7.791 21.651 31.293
    Istanbul CRPS 0.173 0.116 0.112 0.112 0.110 0.110 0.108 0.257
    MSIS 5.937 4.461 4.277 3.813 4.752 4.057 4.094 45.473
    Turkey Power CRPS 0.048 0.040 0.036 0.054 0.046 0.039 0.066 0.085
    MSIS 7.127 6.766 6.341 8.978 8.579 7.943 13.520 36.256

    Zero-shot MOIRAIBase_{\text{Base}} and MOIRAILarge_{\text{Large}} achieved best or second-best CRPS on Electricity, Solar, Weather, and Turkey Power, competitive with or exceeding models trained and tuned directly on target training splits.

  9. Knowl 9 — Zero-Shot Long Sequence Forecasting Benchmark Performance

    empirical result

    MOIRAI generates point forecasts by taking the median of sample paths drawn from its output mixture distribution. Evaluated across standard long-sequence benchmark datasets averaged over prediction horizons h∈{96,192,336,720}h \in \{96, 192, 336, 720\}, zero-shot MOIRAI achieves competitive or superior MSE and MAE compared to fully supervised models trained on each target dataset.

    Dataset Metric Zero-shot MOIRAI Full-shot Supervised Baselines
    Small Base Large iTransformer TimesNet PatchTST TiDE DLinear
    ETTh1 MSE 0.400 0.434 0.510 0.454 0.458 0.469 0.541 0.456
    MAE 0.424 0.438 0.469 0.448 0.450 0.455 0.507 0.452
    ETTh2 MSE 0.341 0.345 0.354 0.383 0.414 0.387 0.611 0.559
    MAE 0.379 0.382 0.376 0.407 0.497 0.407 0.550 0.515
    ETTm1 MSE 0.448 0.381 0.390 0.407 0.400 0.387 0.419 0.403
    MAE 0.409 0.388 0.389 0.410 0.406 0.400 0.419 0.407
    ETTm2 MSE 0.300 0.272 0.276 0.288 0.291 0.281 0.358 0.350
    MAE 0.341 0.321 0.320 0.332 0.333 0.326 0.404 0.401
    Electricity MSE 0.233 0.188 0.188 0.178 0.193 0.216 0.252 0.212
    MAE 0.320 0.274 0.273 0.270 0.295 0.304 0.344 0.300
    Weather MSE 0.242 0.238 0.259 0.258 0.259 0.259 0.271 0.265
    MAE 0.267 0.261 0.275 0.278 0.287 0.281 0.320 0.317
  10. Knowl 10 — Ablation of MOIRAI Architecture and Training Components

    empirical result

    An ablation study evaluated on the Monash Time Series Forecasting Benchmark demonstrates the individual contribution of each architectural component and training strategy in MOIRAI. Performance is reported using normalized Mean Absolute Error (MAE), computed by normalizing each dataset's MAE by the naive forecaster's MAE and taking the geometric mean across datasets (lower is better):

    Configuration Normalized MAE
    MOIRAISmall_{\text{Small}} (Default) 0.655
    w/o patch size constraints 0.720
    w/o multi patch size (fixed P=32P=32) 1.156
    w/o Any-variate Attention (additive learned embeddings) 0.904
    w/o mixture distribution (Student's tt-distribution only) 0.740
    w/o LOTSA (trained on Monash and GluonTS only) 0.809
    w/o sequence packing 0.785

    Key takeaways:

    • Removing multi-patch size projection layers (fixing P=32P=32) causes the largest error increase (+76.5%+76.5\%, reaching 1.1561.156).
    • Replacing Any-variate Attention with additive learned variate index embeddings degrades MAE to 0.9040.904.
    • Restricting training data to Monash and GluonTS (~1B observations) instead of LOTSA (~27B observations) degrades performance to 0.8090.809, confirming the necessity of large-scale, cross-domain pre-training.
  11. Knowl 11 — Context Length Monotonic Scaling and Non-Autoregressive Inference

    empirical result

    Zero-shot evaluation of MOIRAI across context lookback lengths l∈{100,250,500,750,1000,2000,3000,4000,5000}l \in \{100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000\} on ETTm1, Electricity, and Weather reveals that forecasting error (MAE) strictly decreases as context length increases up to 5,000 steps (at a fixed prediction horizon h=96h=96 and patch size P=32P=32). This resolves the known limitation in standard time series Transformers where performance degrades or saturates with longer lookback windows.

    Computationally, MOIRAI's masked encoder outputs predictions across the entire forecast horizon in a single non-autoregressive forward pass. For a batch size of 32 at context and prediction lengths of 5,000, inference takes:

    • MOIRAISmall_{\text{Small}} (P=32P=32): 0.070.07 seconds
    • MOIRAIBase_{\text{Base}} (P=32P=32): 0.130.13 seconds
    • MOIRAILarge_{\text{Large}} (P=32P=32): 0.300.30 seconds

    In contrast, autoregressive baselines like DeepAR require step-by-step unrolling, taking 10.2410.24 seconds for h=5000h=5000, and TFT runs out of memory (OOM).

Coverage note — None was omitted; all primary architectural innovations (Any-variate attention, multi-patch projections, mixture head), dataset contributions (LOTSA), training paradigms (data capping, task distribution, packing), benchmark evaluations, and ablations are included.

References

  1. 1.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., Stella, L., TÃ ¼rkmen, A. C., and Wang, Y. Gluonts: Probabilistic and neural time series modeling in python. Journal of Machine Learning Research, 21(116):1–6, 2020. URL http://jmlr.org/papers/v21/19-820.html.
  2. 2.Awasthi, P., Das, A., Sen, R., and Suresh, A. T. On the benefits of maximum likelihood estimation for regression and forecasting. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=zrW-LVXj2k1.
  3. 3.Bergmeir, C., Bui, Q., de Nijs, F., and Stuckey, P. Residential power and battery data, August 2023. URL https://doi.org/10.5281/zenodo.8219786.
  4. 4.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  6. 6.CDC. Flu portal dashboard, 2017. URL https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html.
  7. 7.Chen, C., Petty, K., Skabardonis, A., Varaiya, P., and Jia, Z. Freeway performance measurement system: mining loop detector data. Transportation Research Record, 1748(1): 96–102, 2001.
  8. 8.Chen, S. Beijing Multi-Site Air-Quality Data. UCI Machine Learning Repository, 2019. DOI: https://doi.org/10.24432/C5RK5G.
  9. 9.Das, A., Kong, W., Leach, A., Mathur, S. K., Sen, R., and Yu, R. Long-term forecasting with tiDE: Timeseries dense encoder. Transactions on Machine Learning Research, 2023a. ISSN 2835-8856. URL https://openreview.net/forum?id=pCbC3aQB5W.
  10. 10.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoderonly foundation model for time-series forecasting. arXiv preprint arXiv:2310.10688, 2023b.
  11. 11.Dong, J., Wu, H., Zhang, H., Zhang, L., Wang, J., and Long, M. SimMTM: A simple pre-training framework for masked time-series modeling. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=ginTcBUnL8.
  12. 12.Dooley, S., Khurana, G. S., Mohapatra, C., Naidu, S. V., and White, C. ForecastPFN: Synthetically-trained zeroshot forecasting. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=tScBQRNgjk.
  13. 13.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.Ekambaram, V., Jati, A., Nguyen, N., Sinthong, P., and Kalagnanam, J. Tsmixer: Lightweight mlp-mixer model for multivariate time series forecasting. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’23, pp. 459–469, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400701030. doi: 10.1145/3580305.3599533. URL https://doi.org/10.1145/3580305.3599533.
  15. 15.Ekambaram, V., Jati, A., Nguyen, N. H., Dayama, P., Reddy, C., Gifford, W. M., and Kalagnanam, J. Ttms: Fast multi-level tiny time mixers for improved zero-shot and few-shot forecasting of multivariate time series. arXiv preprint arXiv:2401.03955, 2024.
  16. 16.Emami, P., Sahu, A., and Graf, P. Buildingsbench: A large-scale dataset of 900k buildings and benchmark for short-term load forecasting. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=c5rqd6PZn6.
  17. 17.Feng, S., Miao, C., Zhang, Z., and Zhao, P. Latent diffusion transformer for probabilistic time series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 11979–11987, 2024.
  18. 18.Garza, A. and Mergenthaler-Canseco, M. Timegpt-1. arXiv preprint arXiv:2310.03589, 2023.
  19. 19.Garza, F., Canseco, M. M., Challu, C., and Olivares, K. G. ´ StatsForecast: Lightning fast forecasting with statistical and econometric models. PyCon Salt Lake City, Utah, US 2022, 2022. URL https://github.com/Nixtla/statsforecast.
  20. 20.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007.
  21. 21.Godahewa, R. W., Bergmeir, C., Webb, G. I., Hyndman, R., and Montero-Manso, P. Monash time series forecasting archive. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=wEc1mgAjU-.
  22. 22.Gruver, N., Finzi, M. A., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=md68e8iZK1.
  23. 23.Henry, A., Dachapally, P. R., Pawar, S. S., and Chen, Y. Query-key normalization for transformers. In Cohn, T., He, Y., and Liu, Y. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 4246–4253, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp. 379. URL https://aclanthology.org/2020.findings-emnlp.379.
  24. 24.Hyndman, R. J. Errors on percentage errors, 4 2014. URL https://robjhyndman.com/hyndsight/smape/.
  25. 25.Hyndman, R. J. and Athanasopoulos, G. Forecasting: principles and practice. OTexts, 2018.
  26. 26.Hyndman, R. J. and Koehler, A. B. Another look at measures of forecast accuracy. International journal of forecasting, 22(4):679–688, 2006.
  27. 27.Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., et al. Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728, 2023.
  28. 28.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  29. 29.Kim, T., Kim, J., Tae, Y., Park, C., Choi, J.-H., and Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=cGDAkQo1C0p.
  30. 30.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, pp. 95–104, 2018.
  31. 31.Lim, B., Arık, S. O., Loeff, N., and Pfister, T. ¨ Temporal fusion transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4):1748–1764, 2021.
  32. 32.Liu, M., Zeng, A., Chen, M., Xu, Z., Lai, Q., Ma, L., and Xu, Q. Scinet: Time series modeling and forecasting with sample convolution and interaction. Advances in Neural Information Processing Systems, 35:5816–5828, 2022.
  33. 33.Liu, X., Xia, Y., Liang, Y., Hu, J., Wang, Y., Bai, L., Huang, C., Liu, Z., Hooi, B., and Zimmermann, R. Largest: A benchmark dataset for large-scale traffic forecasting. arXiv preprint arXiv:2306.08259, 2023a.
  34. 34.Liu, X., Hu, J., Li, Y., Diao, S., Liang, Y., Hooi, B., and Zimmermann, R. Unitime: A language-empowered unified model for cross-domain time series forecasting. In Proceedings of the ACM Web Conference 2024, 2024.
  35. 35.Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M. itransformer: Inverted transformers are effective for time series forecasting. arXiv preprint arXiv:2310.06625, 2023b.
  36. 36.Ma, Q., Liu, Z., Zheng, Z., Huang, Z., Zhu, S., Yu, Z., and Kwok, J. T. A survey on time-series pre-trained models. arXiv preprint arXiv:2305.10716, 2023.
  37. 37.Makridakis, S., Spiliotis, E., and Assimakopoulos, V. The m4 competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting, 36(1): 54–74, 2020.
  38. 38.Mancuso, P., Piccialli, V., and Sudoso, A. M. A machine learning approach for forecasting hierarchical time series. Expert Systems with Applications, 182:115102, 2021.
  39. 39.Mouatadid, S., Orenstein, P., Flaspohler, G. E., Oprescu, M., Cohen, J., Wang, F., Knight, S. E., Geogdzhayeva, M., Levang, S. J., Fraenkel, E., and Mackey, L. SubseasonalclimateUSA: A dataset for subseasonal forecasting and benchmarking. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=pWkrU6raMt.
  40. 40.Nguyen, T., Jewik, J. K., Bansal, H., Sharma, P., and Grover, A. Climatelearn: Benchmarking machine learning for weather and climate modeling. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=RZJEkLFlPx.
  41. 41.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Jbdc0vTOcol.
  42. 42.Oreshkin, B. N., Carpov, D., Chapados, N., and Bengio, Y. N-beats: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1ecqn4YwB.
  43. 43.Park, Y., Maddix, D., Aubet, F.-X., Kan, K., Gasthaus, J., and Wang, Y. Learning quantile functions without quantile crossing for distribution-free time series forecasting. In International Conference on Artificial Intelligence and Statistics, pp. 8127–8150. PMLR, 2022.
  44. 44.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  45. 45.Rasul, K., Ashok, A., Williams, A. R., Khorasani, A., Adamopoulos, G., Bhagwatkar, R., Bilos, M., Ghonia, ˇ H., Hassen, N. V., Schneider, A., Garg, S., Drouin, A., Chapados, N., Nevmyvaka, Y., and Rish, I. Lag-llama: Towards foundation models for time series forecasting, 2023.
  46. 46.Richardson, N., Cook, I., Crane, N., Dunnington, D., Franc¸ois, R., Keane, J., Moldovan-Grunfeld, D., Ooms, ¨ J., Wujciak-Jens, J., and Apache Arrow. arrow: Integration to ’Apache’ ’Arrow’, 2023. URL https://github.com/apache/arrow/. R package version 14.0.2, https://arrow.apache.org/docs/r/.
  47. 47.Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T. Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020.
  48. 48.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  49. 49.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  50. 50.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., ` Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  51. 51.Trindade, A. ElectricityLoadDiagrams20112014. UCI Machine Learning Repository, 2015. DOI: https://doi.org/10.24432/C58C86.
  52. 52.Van Ness, M., Shen, H., Wang, H., Jin, X., Maddix, D. C., and Gopalswamy, K. Cross-frequency time series metaforecasting. arXiv preprint arXiv:2302.02077, 2023.
  53. 53.van Panhuis, W. G., Cross, A., and Burke, D. S. Project tycho 2.0: a repository to improve the integration and reuse of data for global population health. Journal of the American Medical Informatics Association, 25:1608–1617, 2018.
  54. 54.Walmart Competition Admin, W. C. Walmart recruiting - store sales forecasting, 2014.
  55. 55.Wang, J., Jiang, J., Jiang, W., Han, C., and Zhao, W. X. Towards efficient and comprehensive urban spatial-temporal prediction: A unified library and performance benchmark. arXiv preprint arXiv:2304.14343, 2023a.
  56. 56.Wang, Z., Wen, Q., Zhang, C., Sun, L., Von Krannichfeldt, L., and Wang, Y. Benchmarks and custom package for electrical load forecasting. arXiv preprint arXiv:2307.07191, 2023b.
  57. 57.Wikipedia contributors. Moirai — Wikipedia, the free encyclopedia, 2024. URL https://en.wikipedia.org/wiki/Moirai. [Online; accessed 21-January2024].
  58. 58.Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S. CoST: Contrastive learning of disentangled seasonaltrend representations for time series forecasting. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=PilZY3omXV2.
  59. 59.Woo, G., Liu, C., Kumar, A., and Sahoo, D. Pushing the limits of pre-training for time series forecasting in the cloudops domain. arXiv preprint arXiv:2310.05063, 2023.
  60. 60.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34:22419–22430, 2021.
  61. 61.Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. Timesnet: Temporal 2d-variation modeling for general time series analysis. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ju_Uqw384Oq.
  62. 62.Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pp. 10524–10533. PMLR, 2020.
  63. 63.Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J. Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer. arXiv preprint arXiv:2203.03466, 2022a.
  64. 64.Yang, J., Gupta, A., Upadhyay, S., He, L., Goel, R., and Paul, S. TableFormer: Robust transformer modeling for table-text encoding. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 528–537, Dublin, Ireland, May 2022b. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.40. URL https://aclanthology.org/2022.acl-long.40.
  65. 65.Yu, H.-F., Rao, N., and Dhillon, I. S. Temporal regularized matrix factorization for high-dimensional time series prediction. Advances in neural information processing systems, 29, 2016.
  66. 66.Yue, Z., Wang, Y., Duan, J., Yang, T., Huang, C., Tong, Y., and Xu, B. Ts2vec: Towards universal representation of time series. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 8980–8987, 2022.
  67. 67.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 11121–11128, 2023.
  68. 68.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, pp. 2114–2124, 2021.
  69. 69.Zhang, B. and Sennrich, R. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  70. 70.Zhang, K., Wen, Q., Zhang, C., Cai, R., Jin, M., Liu, Y., Zhang, J., Liang, Y., Pang, G., Song, D., et al. Self-supervised learning for time series analysis: Taxonomy, progress, and prospects. arXiv preprint arXiv:2306.10125, 2023.
  71. 71.Zhang, Y. and Yan, J. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=vSVLM2j9eie.
  72. 72.Zheng, Y., Yi, X., Li, M., Li, R., Shan, Z., Chang, E., and Li, T. Forecasting fine-grained air quality based on big data. In Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, pp. 2267–2276, 2015.
  73. 73.Zhou, J., Lu, X., Xiao, Y., Su, J., Lyu, J., Ma, Y., and Dou, D. Sdwpf: A dataset for spatial dynamic wind power forecasting challenge at kdd cup 2022. arXiv preprint arXiv:2208.04360, 2022a.
  74. 74.Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In Proc. 39th International Conference on Machine Learning (ICML 2022), 2022b.
  75. 75.Zhou, T., Niu, P., Wang, X., Sun, L., and Jin, R. One fits all: Power general time series analysis by pretrained LM. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=gMS6FVZvmF.

Citation

MLA
Woo, G., et al. “Unified Training of Universal Time Series Forecasting Transformers”. arXiv, 2024, http://arxiv.org/abs/2402.02592v2.
APA
Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., & Sahoo, D. (2024). Unified Training of Universal Time Series Forecasting Transformers. arXiv. http://arxiv.org/abs/2402.02592v2
Chicago
Woo, G., C. Liu, A. Kumar, C. Xiong, S. Savarese, and D. Sahoo. 2024. “Unified Training of Universal Time Series Forecasting Transformers”. arXiv. http://arxiv.org/abs/2402.02592v2.
Harvard
Woo, G. et al. (2024) “Unified Training of Universal Time Series Forecasting Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.02592v2.
Vancouver
1. Woo G, Liu C, Kumar A, Xiong C, Savarese S, Sahoo D (2024) Unified Training of Universal Time Series Forecasting Transformers. arXiv

BibTeX

@article{woo2024unified,
  title = {Unified Training of Universal Time Series Forecasting Transformers},
  author = {Woo, Gerald and Liu, Chenghao and Kumar, Akshat and Xiong, Caiming and Savarese, Silvio and Sahoo, Doyen},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.02592v2},
  eprint = {2402.02592}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/