Non-autoregressive Conditional Diffusion Models for Time Series Prediction

Lifeng ShenJames T. Kwok

article2023ICML174 citations

Proposes TimeDiff, a non-autoregressive diffusion framework featuring future mixup and autoregressive initialization to overcome error accumulation and outperform existing transformers and diffusion baselines in long-range time series forecasting.

Listen

Accurate time series forecasting is critical across diverse domains such as energy management, traffic control, and financial operations. While generative diffusion models have achieved remarkable success in synthesizing images, audio, and text, adapting them effectively for time series prediction has remained a major challenge. Existing approaches either rely on sequential, step-by-step forecasting—which suffers from severe error accumulation and slow operational speeds—or use non-sequential generation methods borrowed from image processing that fail to capture complex temporal dynamics and create artificial mismatches between historical data and future projections.

The article develops and evaluates a non-autoregressive conditional diffusion model named TimeDiff, designed specifically to produce accurate and efficient long-horizon time series predictions. The primary objective is to demonstrate that introducing domain-specific conditioning mechanisms into diffusion architectures significantly enhances predictive accuracy and computational speed compared to existing generative and deep learning alternatives.

To achieve this, the authors designed two novel conditioning components tailored for time series data. The first component, termed future mixup, exposes portions of actual future ground-truth data during training to guide the denoising network, while relying solely on historical mappings during inference. The second component, an autoregressive initialization module, quickly estimates short-term baseline trends without sequential decoding overhead. The approach was systematically evaluated across nine real-world datasets spanning energy grid operations, weather forecasting, highway traffic occupancy, and foreign exchange rates, benchmarked against sixteen baseline models under both single-variable and multi-variable configurations.

The evaluation revealed several key findings regarding model accuracy and efficiency. First, TimeDiff consistently outperformed all existing time series diffusion models, achieving the best overall average ranking of 1.7 in multi-variable settings and 2.7 in single-variable settings across the nine benchmarks. Second, the model demonstrated superior computational efficiency during inference; for example, on the standard transformer temperature benchmark, TimeDiff executed predictions in approximately 16 to 35 milliseconds, operating up to two orders of magnitude faster than step-by-step diffusion baselines. Third, ablation studies confirmed that predicting clean data directly rather than estimating noise, combined with soft continuous future mixup, substantially reduced prediction error by mitigating boundary distortions.

These findings indicate that generative diffusion architectures can serve as highly reliable and scalable engines for long-range enterprise forecasting when tailored with appropriate temporal inductive biases. By eliminating the quadratic computational bottlenecks of attention-based diffusion models and the compounding errors of autoregressive decoding, organizations can achieve high-fidelity predictions without prohibitive computational or memory costs.

Organizations evaluating advanced forecasting pipelines should consider integrating non-autoregressive diffusion models with specialized temporal conditioning into their predictive infrastructure. When deploying such models, practitioners should adopt direct data prediction and efficient learning-free solvers to maintain low latency in production environments.

A key limitation identified in the article is the model's reduced performance when capturing cross-variable interactions across massive variable dimensions, such as high-density traffic sensor networks. While confidence in the reported results is high across diverse empirical scenarios, users deploying the model on very high-dimensional systems should exercise caution and consider future integration with graph neural networks to better model inter-variable dependencies.

arXiv: 2306.05043
Cover for Non-autoregressive Conditional Diffusion Models for Time Series Prediction

Abstract

Recently, denoising diffusion models have led to significant breakthroughs in the generation of images, audio and text. However, it is still an open question on how to adapt their strong modeling ability to model time series. In this paper, we propose TimeDiff, a non-autoregressive diffusion model that achieves high-quality time series prediction with the introduction of two novel conditioning mechanisms: future mixup and autoregressive initialization. Similar to teacher forcing, future mixup allows parts of the ground-truth future predictions for conditioning, while autoregressive initialization helps better initialize the model with basic time series patterns such as short-term trends. Extensive experiments are performed on nine real-world datasets. Results show that TimeDiff consistently outperforms existing time series diffusion models, and also achieves the best overall performance across a variety of the existing strong baselines (including transformers and FiLM).

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Diffusion Models
  • 2.2. Conditional DDPMs for Time Series Prediction
  • 3. Proposed Model
  • 3.1. Forward Diffusion Process
  • 3.2. Conditioning the Backward Denoising Process
  • 3.2.1. FUTURE MIXUP
  • 3.2.2. AUTOREGRESSIVE MODEL
  • 3.3. Denoising Network
  • 3.4. Training
  • 3.5. Inference
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Results
  • 4.3. Ablation Studies
  • 4.3.1. CONDITIONING MECHANISM
  • 4.3.2. MIXUP STRATEGIES IN FUTURE MIXUP
  • 4.3.3. PREDICTING x θ VS PREDICTING ϵ θ
  • 4.4. Integration into Existing Diffusion Models
  • 4.5. Inference Efficiency
  • 5. Conclusion
  • Acknowledgements
  • References
  • A. Time Series Datasets
  • B. Implementation Details
  • B.1. Network Architecture
  • B.2. Baselines
  • C. Using Channel-Independence on Multivariate Time Series Datasets

Knowls

  1. Knowl 1 — TimeDiff Conditional Non-autoregressive Diffusion Framework for Time Series

    model/method

    TimeDiff is a non-autoregressive conditional diffusion model for time series forecasting. Given past lookback observations x−L+1:00∈Rd×Lx_{-L+1:0}^0 \in \mathbb{R}^{d \times L} across dd variables and lookback window length LL, the objective is to predict future values x1:H0∈Rd×Hx_{1:H}^0 \in \mathbb{R}^{d \times H} across forecast horizon HH.

    The forward diffusion process corrupts the clean target horizon x1:H0x_{1:H}^0 over KK discrete diffusion steps:

    x1:Hk=αˉkx1:H0+1−αˉkϵx_{1:H}^k = \sqrt{\bar{\alpha}_k} x_{1:H}^0 + \sqrt{1 - \bar{\alpha}_k}\epsilon

    where ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), αk=1−βk\alpha_k = 1 - \beta_k for a predefined variance schedule βk∈[0,1]\beta_k \in [0, 1], and αˉk=∏s=1kαs\bar{\alpha}_k = \prod_{s=1}^k \alpha_s.

    The backward denoising transition at step k∈{1,…,K}k \in \{1, \dots, K\} is conditioned on an auxiliary representation c∈R2d×Hc \in \mathbb{R}^{2d \times H}:

    pθ(x1:Hk−1∣x1:Hk,c)=N(x1:Hk−1; μθ(x1:Hk,k∣c),σk2I)p_\theta(x_{1:H}^{k-1} \mid x_{1:H}^k, c) = \mathcal{N}\left(x_{1:H}^{k-1};\, \mu_\theta(x_{1:H}^k, k \mid c), \sigma_k^2 I\right)

    Rather than estimating diffusion noise ϵ\epsilon, TimeDiff parameterizes the mean via a direct clean-data prediction network xθ(x1:Hk,k∣c)∈Rd×Hx_\theta(x_{1:H}^k, k \mid c) \in \mathbb{R}^{d \times H}:

    μθ(x1:Hk,k∣c)=αk(1−αˉk−1)1−αˉkx1:Hk+αˉk−1βk1−αˉkxθ(x1:Hk,k∣c)\mu_\theta(x_{1:H}^k, k \mid c) = \frac{\sqrt{\alpha_k}(1 - \bar{\alpha}_{k-1})}{1 - \bar{\alpha}_k} x_{1:H}^k + \frac{\sqrt{\bar{\alpha}_{k-1}}\beta_k}{1 - \bar{\alpha}_k} x_\theta(x_{1:H}^k, k \mid c)

    The conditioning signal is constructed by concatenating two specialized temporal representations along the variable dimension:

    c=concat([zmix,zar])∈R2d×Hc = \text{concat}([z_{\text{mix}}, z_{\text{ar}}]) \in \mathbb{R}^{2d \times H}

    where zmixz_{\text{mix}} is produced by future mixup and zarz_{\text{ar}} is produced by a linear autoregressive initialization module.

  2. Knowl 2 — Future Mixup Conditioning Mechanism

    model/method

    Future mixup is a conditioning mechanism that injects ground-truth future target information into the conditioning network during diffusion training while preventing exposure bias during inference.

    During training at diffusion step k∈{1,…,K}k \in \{1, \dots, K\}, the future mixup signal zmix∈Rd×Hz_{\text{mix}} \in \mathbb{R}^{d \times H} is computed as:

    zmix=mk⊙F(x−L+1:00)+(1−mk)⊙x1:H0z_{\text{mix}} = m^k \odot \mathcal{F}(x_{-L+1:0}^0) + (1 - m^k) \odot x_{1:H}^0

    where:

    • x−L+1:00∈Rd×Lx_{-L+1:0}^0 \in \mathbb{R}^{d \times L} denotes the past observations across dd variables over lookback length LL.
    • x1:H0∈Rd×Hx_{1:H}^0 \in \mathbb{R}^{d \times H} is the clean ground-truth future series across forecast horizon HH.
    • F:Rd×L→Rd×H\mathcal{F}: \mathbb{R}^{d \times L} \to \mathbb{R}^{d \times H} is a 1D convolutional feature extraction network mapping past history to the forecast horizon shape.
    • ⊙\odot denotes the element-wise Hadamard product.
    • mk∈[0,1)d×Hm^k \in [0, 1)^{d \times H} is a continuous mixing weight matrix whose entries are sampled independently from the uniform distribution U[0,1)\mathcal{U}[0, 1) at each training step.

    During inference, ground-truth future values x1:H0x_{1:H}^0 are unavailable, and the conditioning signal is set directly to:

    zmix=F(x−L+1:00)z_{\text{mix}} = \mathcal{F}(x_{-L+1:0}^0)

  3. Knowl 3 — Autoregressive Initialization for Boundary Disharmony Alleviation

    model/method

    To prevent boundary disharmony and discontinuities between the history window x−L+1:00∈Rd×Lx_{-L+1:0}^0 \in \mathbb{R}^{d \times L} and forecast window x1:H0∈Rd×Hx_{1:H}^0 \in \mathbb{R}^{d \times H}, TimeDiff incorporates an auxiliary linear autoregressive module Mar\mathcal{M}_{\text{ar}} to supply an initial coarse prediction zar∈Rd×Hz_{\text{ar}} \in \mathbb{R}^{d \times H}.

    Let xi0∈Rd×1x_i^0 \in \mathbb{R}^{d \times 1} denote the ii-th column (time step) of the past observation matrix x−L+1:00x_{-L+1:0}^0 for i∈{−L+1,…,0}i \in \{-L+1, \dots, 0\}. The module computes:

    zar=∑i=−L+10Wi⊙Xi0+Bz_{\text{ar}} = \sum_{i=-L+1}^0 W_i \odot X_i^0 + B

    where Xi0=[xi0,xi0,…,xi0]∈Rd×HX_i^0 = [x_i^0, x_i^0, \dots, x_i^0] \in \mathbb{R}^{d \times H} is formed by replicating column vector xi0x_i^0 across HH columns, and Wi∈Rd×H,B∈Rd×HW_i \in \mathbb{R}^{d \times H}, B \in \mathbb{R}^{d \times H} are trainable parameter matrices.

    The autoregressive module Mar\mathcal{M}_{\text{ar}} is pretrained independently on the training dataset prior to diffusion training (e.g., for 20 epochs) by minimizing the squared error loss ∥zar−x1:H0∥22\|z_{\text{ar}} - x_{1:H}^0\|_2^2. Because all columns of zarz_{\text{ar}} are evaluated in parallel via matrix operations, Mar\mathcal{M}_{\text{ar}} captures fundamental short-term linear trends without sequential decoding or error accumulation.

  4. Knowl 4 — TimeDiff Denoising Network Architecture

    model/method

    The TimeDiff denoising network xθ(x1:Hk,k∣c)x_\theta(x_{1:H}^k, k \mid c) predicts clean data x1:H0∈Rd×Hx_{1:H}^0 \in \mathbb{R}^{d \times H} from noisy input x1:Hk∈Rd×Hx_{1:H}^k \in \mathbb{R}^{d \times H}, diffusion step k∈{1,…,K}k \in \{1, \dots, K\}, and condition c∈R2d×Hc \in \mathbb{R}^{2d \times H}.

    1. Diffusion Step Embedding: Step kk is mapped to a sinusoidal position embedding:

    kembedding=[sin⁡(100⋅4w−1k),…,sin⁡(10w⋅4w−1k),cos⁡(100⋅4w−1k),…,cos⁡(10w⋅4w−1k)]k_{\text{embedding}} = \left[\sin\left(10^{\frac{0 \cdot 4}{w-1}} k\right), \dots, \sin\left(10^{\frac{w \cdot 4}{w-1}} k\right), \cos\left(10^{\frac{0 \cdot 4}{w-1}} k\right), \dots, \cos\left(10^{\frac{w \cdot 4}{w-1}} k\right)\right]

    with w=d′/2w = d'/2 and d′=256d' = 256. This embedding is transformed via two fully connected layers with hidden dimension 128 and Sigmoid-Weighted Linear Unit (SiLU) activations into pk=SiLU(FC(SiLU(FC(kembedding))))∈Rd′×1p^k = \text{SiLU}(\text{FC}(\text{SiLU}(\text{FC}(k_{\text{embedding}})))) \in \mathbb{R}^{d' \times 1}.

    1. Input Projection and Encoder: The noisy horizon x1:Hkx_{1:H}^k is projected by 1D convolutional layers to feature tensor z1k∈Rd′×Hz_1^k \in \mathbb{R}^{d' \times H}. The vector pkp^k is broadcasted along the sequence length to size d′×Hd' \times H and concatenated with z1kz_1^k along the channel dimension to form a 2d′×H2d' \times H tensor. A multilayer 1D convolutional encoder processes this tensor into latent representation z2k∈Rd′′×Hz_2^k \in \mathbb{R}^{d'' \times H} with d′′=256d'' = 256.

    2. Condition Fusion Decoder: Condition tensor c=concat([zmix,zar])∈R2d×Hc = \text{concat}([z_{\text{mix}}, z_{\text{ar}}]) \in \mathbb{R}^{2d \times H} is concatenated with z2kz_2^k along the variable dimension to yield a tensor of size (2d+d′′)×H(2d + d'') \times H. A multilayer convolutional decoder maps this representation to output xθ(x1:Hk,k∣c)∈Rd×Hx_\theta(x_{1:H}^k, k \mid c) \in \mathbb{R}^{d \times H}.

    Each convolutional block in the conditioning network, encoder, and decoder comprises Conv1d (kernel size 3, stride 1, padding 1, 256 channels), BatchNorm1d, LeakyReLU (negative slope 0.1), and Dropout (rate 0.1).

  5. Knowl 5 — TimeDiff Training Procedure

    algorithm

    The training procedure optimizes the parameters θ\theta of the denoising network and conditioning network F\mathcal{F} given a pretrained linear autoregressive initialization model Mar\mathcal{M}_{\text{ar}}.

    Input: Training dataset of time series pairs (x−L+1:00,x1:H0)(x_{-L+1:0}^0, x_{1:H}^0), total diffusion steps KK, variance schedule {βk}k=1K\{\beta_k\}_{k=1}^K, pretrained autoregressive model Mar\mathcal{M}_{\text{ar}}, conditioning network F\mathcal{F}, denoising network xθx_\theta.
    Output: Trained model parameters θ\theta.
    repeat
        Sample a training instance (x−L+1:00,x1:H0)(x_{-L+1:0}^0, x_{1:H}^0);
        Sample diffusion step k∼Uniform({1,2,…,K})k \sim \text{Uniform}(\{1, 2, \dots, K\});
        Sample Gaussian noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) of size d×Hd \times H;
        Compute diffused target x1:Hk=αˉkx1:H0+1−αˉkϵx_{1:H}^k = \sqrt{\bar{\alpha}_k} x_{1:H}^0 + \sqrt{1 - \bar{\alpha}_k}\epsilon;
        Compute diffusion step embedding pkp^k;
        Sample mixing matrix mk∈[0,1)d×Hm^k \in [0, 1)^{d \times H} from U[0,1)\mathcal{U}[0, 1);
        Compute future mixup signal zmix=mk⊙F(x−L+1:00)+(1−mk)⊙x1:H0z_{\text{mix}} = m^k \odot \mathcal{F}(x_{-L+1:0}^0) + (1 - m^k) \odot x_{1:H}^0;
        Compute autoregressive guess zar=Mar(x−L+1:00)z_{\text{ar}} = \mathcal{M}_{\text{ar}}(x_{-L+1:0}^0);
        Form condition tensor c=concat([zmix,zar])c = \text{concat}([z_{\text{mix}}, z_{\text{ar}}]);
        Compute clean data estimate xθ(x1:Hk,k∣c)x_\theta(x_{1:H}^k, k \mid c) using the denoising network;
        Evaluate squared loss Lk(θ)=∥x1:H0−xθ(x1:Hk,k∣c)∥2\mathcal{L}_k(\theta) = \|x_{1:H}^0 - x_\theta(x_{1:H}^k, k \mid c)\|^2;
        Update parameters θ\theta by taking a gradient descent step along ∇θLk(θ)\nabla_\theta \mathcal{L}_k(\theta);
    until convergence.
  6. Knowl 6 — TimeDiff Denoising Inference Procedure

    algorithm

    During inference, TimeDiff generates the predicted future time series non-autoregressively by iteratively denoising Gaussian noise conditioned on the past observations.

    Input: Historical observations x−L+1:00∈Rd×Lx_{-L+1:0}^0 \in \mathbb{R}^{d \times L}, total diffusion steps KK, variance schedule parameters {αk,αˉk,βk,σk}k=1K\{\alpha_k, \bar{\alpha}_k, \beta_k, \sigma_k\}_{k=1}^K, pretrained autoregressive module Mar\mathcal{M}_{\text{ar}}, trained conditioning network F\mathcal{F}, trained denoising network xθx_\theta.
    Output: Forecasted future series x^1:H0∈Rd×H\hat{x}_{1:H}^0 \in \mathbb{R}^{d \times H}.
    Sample initial noise vector x1:HK∼N(0,I)x_{1:H}^K \sim \mathcal{N}(0, I) of shape d×Hd \times H;
    Compute zmix=F(x−L+1:00)z_{\text{mix}} = \mathcal{F}(x_{-L+1:0}^0);
    Compute zar=Mar(x−L+1:00)z_{\text{ar}} = \mathcal{M}_{\text{ar}}(x_{-L+1:0}^0);
    Construct condition tensor c=concat([zmix,zar])c = \text{concat}([z_{\text{mix}}, z_{\text{ar}}]);
    for k=K,K−1,…,1k = K, K-1, \dots, 1 do
        if k>1k > 1 then
            Sample noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I);
        else
            Set ϵ=0\epsilon = 0;
        end if
        Compute step embedding pkp^k;
        Evaluate denoising network prediction xθ(x1:Hk,k∣c)x_\theta(x_{1:H}^k, k \mid c);
        Compute reverse mean update:
            x^1:Hk−1=αk(1−αˉk−1)1−αˉkx1:Hk+αˉk−1βk1−αˉkxθ(x1:Hk,k∣c)+σkϵ\hat{x}_{1:H}^{k-1} = \frac{\sqrt{\alpha_k}(1 - \bar{\alpha}_{k-1})}{1 - \bar{\alpha}_k} x_{1:H}^k + \frac{\sqrt{\bar{\alpha}_{k-1}}\beta_k}{1 - \bar{\alpha}_k} x_\theta(x_{1:H}^k, k \mid c) + \sigma_k \epsilon;
        Set x1:Hk−1=x^1:Hk−1x_{1:H}^{k-1} = \hat{x}_{1:H}^{k-1};
    end for
    return x^1:H0\hat{x}_{1:H}^0.
  7. Knowl 7 — Univariate Forecasting Performance of TimeDiff

    data/table

    The performance of TimeDiff was evaluated on univariate time series forecasting across nine benchmark datasets and compared against diffusion models (TimeGrad, CSDI, SSSD, D3VAE), SOTA forecasting baselines (FiLM, Depts, NBeats, DLinear), and transformers (PatchTST, FedFormer, Autoformer, Pyraformer, Informer, Transformer, LSTMa). Prediction accuracy is measured using Mean Squared Error (MSE), with diffusion baselines averaged over 10 sampled trajectories.

    Model NorPool Caiso Weather ETTm1 Wind Traffic Electricity ETTh1 Exchange Avg Rank
    TimeDiff 0.636 0.122 0.002 0.040 2.407 0.121 0.232 0.066 0.017 2.7
    TimeGrad 1.129 0.325 0.002 0.048 2.530 1.223 0.920 0.078 0.041 11.6
    CSDI 0.967 0.192 0.002 0.050 2.434 0.393 0.520 0.083 0.071 10.8
    SSSD 1.145 0.176 0.004 0.049 3.149 0.151 0.370 0.097 0.023 9.8
    D^3VAE 0.964 0.521 0.003 0.044 2.679 0.151 0.535 0.078 0.019 9.4
    FiLM 0.707 0.185 0.007 0.038 2.143 0.198 0.260 0.070 0.018 5.1
    Depts 0.668 0.107 0.024 0.046 3.457 0.151 0.380 0.070 0.020 7.0
    NBeats 0.768 0.125 0.137 0.048 2.434 0.142 0.378 0.095 0.016 7.4
    PatchTST 0.595 0.193 0.026 0.052 2.698 0.177 0.450 0.106 0.020 10.5
    FedFormer 0.891 0.164 0.005 0.065 2.351 0.173 0.376 0.076 0.050 8.7
    Autoformer 0.946 0.248 0.003 0.051 2.349 0.473 0.659 0.081 0.041 10.9
    Pyraformer 0.933 0.165 0.020 0.054 2.279 0.136 0.389 0.076 0.017 7.1
    Informer 0.804 0.250 0.007 0.049 2.297 0.213 0.363 0.076 0.023 8.1
    Transformer 0.928 0.250 0.007 0.058 2.306 0.238 0.430 0.092 0.018 10.2
    DLinear 0.671 0.118 0.168 0.041 2.171 0.139 0.244 0.078 0.017 4.8
    LSTMa 0.836 0.253 0.005 0.091 2.299 1.032 0.596 0.167 0.031 11.9

    TimeDiff achieves an overall average rank of 2.7 across the 9 datasets, consistently outperforming existing time-series diffusion models (TimeGrad at 11.6, CSDI at 10.8, SSSD at 9.8, and D3VAE at 9.4) as well as deterministic baselines.

  8. Knowl 8 — Multivariate Forecasting Performance of TimeDiff

    data/table

    The multivariate time series forecasting accuracy of TimeDiff was compared against existing diffusion models, time-series transformers, and deep neural models across nine benchmark datasets. Accuracy is evaluated via Mean Squared Error (MSE).

    Model NorPool Caiso Weather ETTm1 Wind Traffic Electricity ETTh1 Exchange Avg Rank
    TimeDiff 0.665 0.136 0.311 0.336 0.896 0.564 0.193 0.407 0.018 1.7
    TimeGrad 1.152 0.258 0.392 0.874 1.209 1.745 0.736 0.993 0.079 13.9
    CSDI 1.011 0.253 0.356 0.529 1.066 – – 0.497 0.077 10.6
    SSSD 0.872 0.195 0.349 0.464 1.188 0.642 0.255 0.726 0.061 8.1
    D^3VAE 0.745 0.241 0.375 0.362 1.118 0.928 0.286 0.504 0.200 8.9
    FiLM 0.723 0.179 0.327 0.347 0.984 0.628 0.210 0.426 0.016 3.2
    Depts 0.662 0.106 0.761 0.380 1.082 1.019 0.319 0.579 0.020 7.7
    NBeats 0.832 0.141 1.344 0.391 1.069 0.373 0.269 0.586 0.016 6.6
    PatchTST 0.851 0.193 0.782 0.372 1.070 0.831 0.225 0.526 0.047 7.7
    FedFormer 0.873 0.205 0.342 0.426 1.113 0.591 0.238 0.541 0.133 7.6
    Autoformer 0.940 0.226 0.360 0.565 1.083 0.688 0.201 0.516 0.056 9.0
    Pyraformer 1.008 0.273 0.394 0.493 1.061 0.659 0.273 0.579 0.032 9.4
    Informer 0.985 0.231 0.385 0.673 1.168 0.664 0.298 0.775 0.073 10.9
    Transformer 1.005 0.206 0.388 0.992 1.201 0.671 0.328 0.759 0.062 11.3
    DLinear 0.670 0.461 0.488 0.345 0.899 0.389 0.215 0.415 0.022 5.3
    LSTMa 1.481 0.217 0.662 1.030 1.464 0.966 0.414 1.149 0.403 14.2

    TimeDiff achieves the lowest average rank (1.7) across all multivariate benchmarks. CSDI encounters out-of-memory errors on large multivariate datasets (Traffic with d=862d=862 and Electricity with d=321d=321) due to quadratic transformer attention complexity, whereas TimeDiff's 1D convolutional backbone scales efficiently.

  9. Knowl 9 — Ablation Analysis on Conditioning Mechanisms, Mixup Strategies, and Denoising Targets

    empirical result

    Ablation studies on four datasets (Caiso, Electricity, Exchange, ETTh1) evaluate the components of TimeDiff:

    1. Conditioning Mechanisms: Combining future mixup and linear autoregressive (AR) initialization achieves the lowest testing MSE compared to removing either or both modules:

      • Both included: Caiso 0.122, Electricity 0.232, Exchange 0.017, ETTh1 0.066.
      • Future mixup only: Caiso 0.149, Electricity 0.328, Exchange 0.020, ETTh1 0.086.
      • AR initialization only: Caiso 0.160, Electricity 0.295, Exchange 0.022, ETTh1 0.162.
      • Neither (condition c=F(x−L+1:00)c = \mathcal{F}(x_{-L+1:0}^0)): Caiso 0.170, Electricity 0.340, Exchange 0.024, ETTh1 0.182.
    2. Mixup Strategies in Future Mixup: Continuous soft-mixup (mk∼U[0,1)m^k \sim \mathcal{U}[0, 1)) outperforms binarized hard-mixup (mask values binarized by threshold τ\tau) and segment-mixup (masking contiguous temporal segments drawn from a geometric distribution with mean 3). Hard and segment mixup MSEs depend heavily on the choice of threshold τ∈{0.1,0.3,0.5,0.7,0.9}\tau \in \{0.1, 0.3, 0.5, 0.7, 0.9\} (e.g., hard-mixup ETTh1 MSE varies from 0.074 at τ=0.7\tau=0.7 to 0.161 at τ=0.1\tau=0.1), whereas parameter-free soft-mixup yields 0.066 on ETTh1.

    3. Data Prediction (xθx_\theta) vs. Noise Prediction (ϵθ\epsilon_\theta): Directly predicting the clean signal xθx_\theta yields lower MSE across all datasets compared to predicting the diffusion noise ϵθ\epsilon_\theta:

      • Caiso: 0.122 (xθx_\theta) vs. 0.134 (ϵθ\epsilon_\theta)
      • Electricity: 0.232 (xθx_\theta) vs. 0.317 (ϵθ\epsilon_\theta)
      • Exchange: 0.017 (xθx_\theta) vs. 0.021 (ϵθ\epsilon_\theta)
      • ETTh1: 0.066 (xθx_\theta) vs. 0.077 (ϵθ\epsilon_\theta) This reflects that real-world time series contain irregular intrinsic noise that is easily confounded with forward Gaussian diffusion noise.
  10. Knowl 10 — Compatibility with Existing Diffusion Models and Inference Efficiency

    empirical result

    Integrating future mixup and autoregressive (AR) initialization into existing non-autoregressive time series diffusion architectures (CSDI and SSSD) systematically reduces prediction error on univariate ETTh1 and ETTm1:

    • CSDI Baseline: MSE on ETTh1 decreases from 0.083 (neither) to 0.078 (+future mixup), 0.088 (+AR), and 0.075 (both). On ETTm1, MSE decreases from 0.050 (neither) to 0.045 (+future mixup), 0.054 (+AR), and 0.043 (both).
    • SSSD Baseline: MSE on ETTh1 decreases from 0.097 (neither) to 0.077 (+future mixup), 0.084 (+AR), and 0.071 (both). On ETTm1, MSE decreases from 0.049 (neither) to 0.044 (+future mixup), 0.052 (+AR), and 0.040 (both).

    In inference latency measured on univariate ETTh1 using DPM-Solver with under 20 denoising steps, TimeDiff achieves faster generation than all diffusion baselines across horizons H∈{96,168,336,720}H \in \{96, 168, 336, 720\}:

    Model H=96H=96 H=168H=168 H=336H=336 H=720H=720
    TimeDiff 16.2 ms 17.2 ms 26.5 ms 34.6 ms
    CSDI 90.41 ms 127.2 ms 398.9 ms 513.1 ms
    SSSD 418.6 ms 595.0 ms 1054.2 ms 2516.9 ms
    TimeGrad 870.2 ms 1579.2 ms 3119.7 ms 6724.1 ms

    TimeDiff executes up to 15×15\times faster than CSDI, 72×72\times faster than SSSD, and 194×194\times faster than autoregressive TimeGrad at H=720H=720.

Coverage note — Appendix C channel-independence experiments (Table 12) and Appendix A Augmented Dick-Fuller test statistics (Table 9) were omitted as supplementary implementation variants of the primary forecasting benchmarks.

References

  1. 1.Alcaraz, J. M. L. and Strodthoff, N. Diffusion-based time series imputation and forecasting with structured state space models. Technical report, arXiv, 2022.
  2. 2.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations, 2015.
  3. 3.Bao, F., Li, C., Zhu, J., and Zhang, B. Analytic-DPM: An analytic estimate of the optimal reverse variance in diffusion probabilistic models. In International Conference on Learning Representations, 2022.
  4. 4.Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N. Scheduled sampling for sequence prediction with recurrent neural networks. Neural Information Processing Systems, 28, 2015.
  5. 5.Benny, Y. and Wolf, L. Dynamic dual-output diffusion models. In Computer Vision and Pattern Recognition, 2022.
  6. 6.Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W. WaveGrad: Estimating gradients for waveform generation. In International Conference on Learning Representations, 2020.
  7. 7.Choi, J., Kim, S., Jeong, Y., Gwon, Y., and Yoon, S. ILVR: Conditioning method for denoising diffusion probabilistic models. In International Conference on Computer Vision, 2021.
  8. 8.Elfwing, S., Uchibe, E., and Doya, K. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Networks, 107:3–11, 2018.
  9. 9.Elliott, G., Rothenberg, T. J., and Stock, J. H. Efficient tests for an autoregressive unit root. Econometrica, pp. 813–836, 1996.
  10. 10.Fan, W., Zheng, S., Yi, X., Cao, W., Fu, Y., Bian, J., and Liu, T.-Y. DEPTS: Deep expansion learning for periodic time series forecasting. In International Conference on Learning Representations, 2022.
  11. 11.Gong, S., Li, M., Feng, J., Wu, Z., and Kong, L. DiffuSeq: Sequence to sequence text generation with diffusion models. In International Conference on Learning Representations, 2023.
  12. 12.Graves, A., Jaitly, N., and Mohamed, A.-R. Hybrid speech recognition with deep bidirectional LSTM. In IEEE Workshop on Automatic Speech Recognition and Understanding, pp. 273–278, 2013.
  13. 13.Henrique, B. M., Sobreiro, V. A., and Kimura, H. Literature review: Machine learning techniques applied to financial market prediction. Expert Systems with Applications, 124: 226–251, 2019.
  14. 14.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Neural Information Processing Systems, 2020.
  15. 15.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  16. 16.Kim, H., Kim, S., and Yoon, S. Guided-TTS: A diffusion model for text-to-speech via classifier guidance. In International Conference on Machine Learning, 2022.
  17. 17.Kim, T., Kim, J., Tae, Y., Park, C., Choi, J.-H., and Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, 2021.
  18. 18.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  19. 19.Kingma, D. P. and Welling, M. Auto-encoding variational Bayes. In International Conference on Learning Representations, 2014.
  20. 20.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. DiffWave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2020.
  21. 21.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In SIGIR Conference on Research & Development in Information Retrieval, 2018.
  22. 22.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Neural Information Processing Systems, 2019.
  23. 23.Li, X., Thickstun, J., Gulrajani, I., Liang, P. S., and Hashimoto, T. B. Diffusion-LM improves controllable text generation. In Neural Information Processing Systems, 2022a.
  24. 24.Li, Y., Lu, X., Wang, Y., and Dou, D. Generative time series forecasting with diffusion, denoise, and disentanglement. In Neural Information Processing Systems, 2022b.
  25. 25.Liu, S., Yu, H., Liao, C., Li, J., Lin, W., Liu, A. X., and Dustdar, S. Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting. In International Conference on Learning Representations, 2021.
  26. 26.Liu, Y., Wu, H., Wang, J., and Long, M. Non-stationary transformers: Rethinking the stationarity in time series forecasting. In Neural Information Processing Systems, 2022.
  27. 27.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In Neural Information Processing Systems, 2022.
  28. 28.Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., and Van Gool, L. Repaint: Inpainting using denoising diffusion probabilistic models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  29. 29.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In International Conference on Learning Representations, 2023.
  30. 30.Oreshkin, B. N., Carpov, D., Chapados, N., and Bengio, Y. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations, 2019.
  31. 31.Rasul, K., Seward, C., Schuster, I., and Vollgraf, R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International Conference on Machine Learning, 2021.
  32. 32.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  33. 33.Sapankevych, N. I. and Sankar, R. Time series prediction using support vector machines: A survey. IEEE Computational Intelligence Magazine, 4(2):24–38, 2009.
  34. 34.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021.
  35. 35.Tashiro, Y., Song, J., Song, Y., and Ermon, S. CSDI: Conditional score-based diffusion models for probabilistic time series imputation. In Neural Information Processing Systems, 2021.
  36. 36.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Neural Information Processing Systems, 30, 2017.
  37. 37.Wang, X., Guo, P., and Huang, X. A review of wind power forecasting models. Energy procedia, 12:770–778, 2011.
  38. 38.Wang, Z., Yan, W., and Oates, T. Time series classification from scratch with deep neural networks: A strong baseline. In International Joint Conference on Neural Networks, 2017.
  39. 39.Williams, R. J. and Zipser, D. A learning algorithm for continually running fully recurrent neural networks. Neural Computation, 1(2):270–280, 1989.
  40. 40.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Neural Information Processing Systems, 2021.
  41. 41.Yang, R., Srivastava, P., and Mandt, S. Diffusion probabilistic modeling for video generation. Technical report, arXiv, 2022.
  42. 42.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In AAAI Conference on Artificial Intelligence, 2023.
  43. 43.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. In SIGKDD Conference on Knowledge Discovery & Data Mining, 2021.
  44. 44.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018.
  45. 45.Zhang, Y. and Yan, J. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In International Conference on Learning Representations, 2023.
  46. 46.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In AAAI Conference on Artificial Intelligence, 2021.
  47. 47.Zhou, T., Ma, Z., Wen, Q., Sun, L., Yao, T., Jin, R., et al. FiLM: Frequency improved Legendre memory model for long-term time series forecasting. In Neural Information Processing Systems, 2022a.
  48. 48.Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. FEDformer: frequency enhanced decomposed transformer for long-term series forecasting. In International Conference on Machine Learning, 2022b.

Citation

MLA
Shen, L., and J. Kwok. “Non-autoregressive Conditional Diffusion Models for Time Series Prediction”. International Conference on Machine Learning, vol. 202, 2023, pp. 31016–29, https://proceedings.mlr.press/v202/shen23d.html.
APA
Shen, L., & Kwok, J. (2023). Non-autoregressive Conditional Diffusion Models for Time Series Prediction. International Conference on Machine Learning, 202, 31016–31029. https://proceedings.mlr.press/v202/shen23d.html
Chicago
Shen, L., and J. Kwok. 2023. “Non-autoregressive Conditional Diffusion Models for Time Series Prediction”. International Conference on Machine Learning 202: 31016–29. https://proceedings.mlr.press/v202/shen23d.html.
Harvard
Shen, L. and Kwok, J. (2023) “Non-autoregressive Conditional Diffusion Models for Time Series Prediction”, International Conference on Machine Learning. PMLR, pp. 31016–31029. Available at: https://proceedings.mlr.press/v202/shen23d.html.
Vancouver
1. Shen L, Kwok J (2023) Non-autoregressive Conditional Diffusion Models for Time Series Prediction. In: International Conference on Machine Learning. PMLR, pp 31016–31029

BibTeX

@InProceedings{pmlr-v202-shen23d,
  title = 	 {Non-autoregressive Conditional Diffusion Models for Time Series Prediction},
  author =       {Shen, Lifeng and Kwok, James},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {31016--31029},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/shen23d/shen23d.pdf},
  url = 	 {https://proceedings.mlr.press/v202/shen23d.html},
  abstract = 	 {Recently, denoising diffusion models have led to significant breakthroughs in the generation of images, audio and text. However, it is still an open question on how to adapt their strong modeling ability to model time series. In this paper, we propose TimeDiff, a non-autoregressive diffusion model that achieves high-quality time series prediction with the introduction of two novel conditioning mechanisms: future mixup and autoregressive initialization. Similar to teacher forcing, future mixup allows parts of the ground-truth future predictions for conditioning, while autoregressive initialization helps better initialize the model with basic time series patterns such as short-term trends. Extensive experiments are performed on nine real-world datasets. Results show that TimeDiff consistently outperforms existing time series diffusion models, and also achieves the best overall performance across a variety of the existing strong baselines (including transformers and FiLM).}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/