Non-autoregressive Conditional Diffusion Models for Time Series Prediction
Lifeng ShenJames T. Kwok
Proposes TimeDiff, a non-autoregressive diffusion framework featuring future mixup and autoregressive initialization to overcome error accumulation and outperform existing transformers and diffusion baselines in long-range time series forecasting.
Accurate time series forecasting is critical across diverse domains such as energy management, traffic control, and financial operations. While generative diffusion models have achieved remarkable success in synthesizing images, audio, and text, adapting them effectively for time series prediction has remained a major challenge. Existing approaches either rely on sequential, step-by-step forecasting—which suffers from severe error accumulation and slow operational speeds—or use non-sequential generation methods borrowed from image processing that fail to capture complex temporal dynamics and create artificial mismatches between historical data and future projections.
The article develops and evaluates a non-autoregressive conditional diffusion model named TimeDiff, designed specifically to produce accurate and efficient long-horizon time series predictions. The primary objective is to demonstrate that introducing domain-specific conditioning mechanisms into diffusion architectures significantly enhances predictive accuracy and computational speed compared to existing generative and deep learning alternatives.
To achieve this, the authors designed two novel conditioning components tailored for time series data. The first component, termed future mixup, exposes portions of actual future ground-truth data during training to guide the denoising network, while relying solely on historical mappings during inference. The second component, an autoregressive initialization module, quickly estimates short-term baseline trends without sequential decoding overhead. The approach was systematically evaluated across nine real-world datasets spanning energy grid operations, weather forecasting, highway traffic occupancy, and foreign exchange rates, benchmarked against sixteen baseline models under both single-variable and multi-variable configurations.
The evaluation revealed several key findings regarding model accuracy and efficiency. First, TimeDiff consistently outperformed all existing time series diffusion models, achieving the best overall average ranking of 1.7 in multi-variable settings and 2.7 in single-variable settings across the nine benchmarks. Second, the model demonstrated superior computational efficiency during inference; for example, on the standard transformer temperature benchmark, TimeDiff executed predictions in approximately 16 to 35 milliseconds, operating up to two orders of magnitude faster than step-by-step diffusion baselines. Third, ablation studies confirmed that predicting clean data directly rather than estimating noise, combined with soft continuous future mixup, substantially reduced prediction error by mitigating boundary distortions.
These findings indicate that generative diffusion architectures can serve as highly reliable and scalable engines for long-range enterprise forecasting when tailored with appropriate temporal inductive biases. By eliminating the quadratic computational bottlenecks of attention-based diffusion models and the compounding errors of autoregressive decoding, organizations can achieve high-fidelity predictions without prohibitive computational or memory costs.
Organizations evaluating advanced forecasting pipelines should consider integrating non-autoregressive diffusion models with specialized temporal conditioning into their predictive infrastructure. When deploying such models, practitioners should adopt direct data prediction and efficient learning-free solvers to maintain low latency in production environments.
A key limitation identified in the article is the model's reduced performance when capturing cross-variable interactions across massive variable dimensions, such as high-density traffic sensor networks. While confidence in the reported results is high across diverse empirical scenarios, users deploying the model on very high-dimensional systems should exercise caution and consider future integration with graph neural networks to better model inter-variable dependencies.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It establishes the foundational mathematical formulation of denoising diffusion probabilistic models and reverse-process objectives that TimeDiff adapts for continuous time-series forecasting.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). It introduces non-Markovian deterministic sampling for diffusion models, providing the non-autoregressive fast inference mechanism essential to the model's speed advantages.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). It provides crucial baseline methodology for long-term time series forecasting benchmarks and temporal decomposition concepts used in evaluating non-autoregressive predictors.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). It introduces standard long-sequence forecasting datasets and one-shot generative decoding baselines that TimeDiff directly benchmarks against and improves upon.
- Paper: Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting, Yong Liu et al. (2022). It formalizes the challenges of non-stationarity and distribution shifts in deep time-series forecasting, informing TimeDiff's temporal conditioning and initialization designs.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). It analyzes network preconditioning and direct data prediction versus noise estimation in diffusion models, which TimeDiff leverages to reduce boundary distortions in time series.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). It demonstrates non-autoregressive trajectory-level conditional diffusion modeling over sequential data, laying the groundwork for conditional non-autoregressive time-series generation.
- Paper: Sundial: A Family of Highly Capable Time Series Foundation Models, Yong Liu 0007 et al. (2025). It advances beyond task-specific conditional diffusion forecasting by constructing continuous flow-matching foundation models for large-scale zero-shot and probabilistic time-series forecasting.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). It tackles the cross-variable modeling limitation in high-dimensional forecasting identified in TimeDiff by inverting Transformer tokenization across individual variates.
- Paper: Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors, Alexander Scheinker (2026). It extends conditional trajectory diffusion modeling by using bidirectional round-trip consistency to predict and self-supervise compounding rollout errors at test time.
