Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion

Marin BilosKashif RasulAnderson SchneiderYuriy NevmyvakaStephan Günnemann

article2023ICML55 citations

Proposes a function-space denoising diffusion framework that preserves temporal continuity using stochastic process noise, enabling effective generative modeling, probabilistic forecasting, and imputation for irregularly sampled time series.

Listen

Real-world time series data in healthcare, finance, and industrial operations is typically gathered at irregular or arbitrary intervals, yet it reflects physical systems that change smoothly and continuously over time. Generative diffusion models have demonstrated remarkable modeling capabilities in computer vision and audio, but traditional diffusion frameworks add independent noise to individual data points. This practice disrupts continuity across time, often generating jagged, discontinuous trajectories that struggle to capture the true underlying dynamics of continuous processes.

The article demonstrates a generative modeling framework that formulates diffusion over continuous function spaces rather than discrete vectors. It evaluates both fixed-step and continuous score-based formulations to naturally accommodate irregularly sampled temporal data for generative modeling, probabilistic forecasting, interpolation, and missing-value imputation.

The authors develop stochastic process diffusion by injecting temporally correlated noise functions—specifically Gaussian and Ornstein-Uhlenbeck processes—into the entire observed time series during the forward phase while training neural networks to reverse the process. The methodology was evaluated across six synthetic dynamical and chaotic systems alongside multiple real-world multivariate benchmarks, including the Electricity, Exchange, and Solar forecasting datasets and the PhysioNet medical record collection for missing-value imputation.

The evaluation yielded several key findings. In synthetic generative tests, a transformer-based discriminator was unable to distinguish between real data and trajectories generated by the proposed model, scoring near a random-chance accuracy of roughly 0.51, whereas competing models such as Latent ODEs and Continuous-Time Flow Processes were easily detected with discrimination accuracies reaching 0.73 to 1.0. In multivariate forecasting benchmarks, the method improved error metrics over autoregressive baselines—for example, reducing normalized root mean squared error on Electricity from 0.064 to 0.045—while predicting the entire forecast horizon in parallel rather than step-by-step. On medical imputation tasks with 50% and 90% missing values, injecting correlated stochastic noise significantly lowered reconstruction error compared to baseline diffusion with independent noise, confirming that continuous inductive biases enhance data recovery.

These findings indicate that treating temporal data as underlying continuous functions significantly enhances generation fidelity and uncertainty calibration without adding substantial computational overhead. Furthermore, predicting complete sequences simultaneously rather than autoregressively improves execution speed and hardware scaling, reducing operational risks and computational costs in deployment settings such as energy grid management and financial modeling.

Organizations handling irregular or noisy time series should consider adopting continuous stochastic process diffusion as a drop-in replacement for standard independent noise diffusion models. Teams implementing this architecture should prioritize network backbones that capture temporal dependencies, such as recurrent networks or transformers, as ablated models lacking temporal interaction fail to model function distributions accurately. Future engineering efforts should explore combining this approach with sparse Gaussian processes to maintain scalability when handling exceptionally long sequences.

The study's primary limitation stems from the computational cost of scaling full covariance matrix operations on extremely large sequences with thousands of observation points. While confidence in the experimental results is high across diverse benchmarks, practitioners should exercise care when selecting kernel hyperparameters, as matching the noise smoothness to the expected roughness of the target domain is essential for optimal curve generation.

  • Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). It generalizes score-based diffusion from continuous temporal functions to resolution-invariant function-space generation using neural operators.
Cover for Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion

Abstract

Temporal data such as time series can be viewed as discretized measurements of the underlying function. To build a generative model for such data we have to model the stochastic process that governs it. We propose a solution by defining the denoising diffusion model in the function space which also allows us to naturally handle irregularly-sampled observations. The forward process gradually adds noise to functions, preserving their continuity, while the learned reverse process removes the noise and returns functions as new samples. To this end, we define suitable noise sources and introduce novel denoising and score-matching models. We show how our method can be used for multivariate probabilistic forecasting and imputation, and how our model can be interpreted as a neural process.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Fixed-step diffusion
  • 2.2. Score-based SDE diffusion
  • 2.3. Extensions
  • 3. Diffusion for time series data
  • 3.1. Stochastic processes as noise sources for diffusion
  • 3.2. Discrete stochastic process diffusion (DSPD)
  • 3.3. Continuous stochastic process diffusion (CSPD)
  • 3.4. Related work
  • 4. Applications
  • 4.1. Forecasting multivariate time series
  • 4.2. Diffusion process as a neural process
  • 4.3. Probabilistic time series imputation
  • 5. Experiments
  • 5.1. Probabilistic modeling
  • 5.2. Forecasting
  • 5.3. Neural process
  • 5.4. Imputation
  • 6. Discussion
  • Acknowledgements
  • References
  • A. Derivations
  • A.1. Discrete diffusion posterior probability
  • A.2. Discrete diffusion loss
  • A.3. Continuous diffusion transition probability
  • A.4. Sampling from an Ornstein-Uhlenbeck process
  • A.5. Algorithms
  • B. Experimental details
  • B.1. Probabilistic modeling
  • B.1.1. DATASETS
  • B.1.2. CTFP
  • B.1.3. LATENT ODE
  • B.1.4. OUR MODELS
  • B.2. Multivariate probabilistic forecasting
  • B.3. Neural process
  • B.3.1. DATASET
  • B.3.2. MODEL
  • B.3.3. ADDITIONAL RESULTS
  • B.4. CSDI imputation

Knowls

  1. Knowl 1 — Diffusion over functions uses temporally correlated noise

    model/method

    The method treats a time series as observations X=(x(t0),…,x(tM−1))X=(x(t_0),\ldots,x(t_{M-1})) of an underlying continuous function x(⋅)x(\cdot), where the observation times tit_i may be irregular. Rather than perturbing each observation independently, the forward diffusion adds a noise function sampled from a continuous stochastic process to the underlying function. This preserves temporal continuity in the noisy samples and makes it possible to define noise at arbitrary observation times. A learned reverse process removes that noise to generate functions from the learned data distribution. For a dd-dimensional time series, the forward noise is applied independently to each of the dd component series, while the reverse denoiser receives the full multivariate series and can model dependencies across both time and components.

  2. Knowl 2 — Gaussian-process and Ornstein–Uhlenbeck noise priors

    model/method

    The method uses stationary stochastic processes to generate temporally correlated noise at observation times t0,…,tM−1t_0,\ldots,t_{M-1}. For a Gaussian-process (GP) prior, the noise vector has distribution N(0,Σ)\mathcal{N}(0,\Sigma), with covariance Σij=k(ti,tj)\Sigma_{ij}=k(t_i,t_j) and radial-basis-function kernel k(ti,tj)=exp⁡[−γ(ti−tj)2]k(t_i,t_j)=\exp[-\gamma(t_i-t_j)^2], where γ>0\gamma>0 controls the smoothness of noise paths. For an Ornstein–Uhlenbeck (OU) prior, the process is specified by dϵt=−γϵt dt+dWtd\epsilon_t=-\gamma\epsilon_t\,dt+dW_t, with initial value ϵ0∼N(0,1)\epsilon_0\sim\mathcal{N}(0,1) and standard Wiener process WtW_t; its stated covariance at the observation times is Σij=exp⁡[−γ∣ti−tj∣]\Sigma_{ij}=\exp[-\gamma|t_i-t_j|]. In either case, evaluating the covariance on the supplied timestamps permits sampling correlated noise even when the observations are irregularly spaced.

  3. Knowl 3 — Discrete stochastic-process diffusion

    equation

    In discrete stochastic-process diffusion (DSPD), X0X_0 is the clean time series flattened into a vector, XnX_n is its state after diffusion step nn, and Σ\Sigma is the covariance of the process noise across the observation times. For multivariate series, Σ\Sigma is extended to a block-diagonal covariance with one copy per component. Let βn\beta_n be the variance schedule, αn=1−βn\alpha_n=1-\beta_n, and αˉn=∏k=1nαk\bar{\alpha}_n=\prod_{k=1}^n\alpha_k, with αˉ0=1\bar{\alpha}_0=1. The forward marginal and one-step posterior are Gaussian:

    q(Xn∣X0)=N ⁣(αˉnX0,(1−αˉn)Σ),q(Xn−1∣Xn,X0)=N(μ~n,β~nΣ),q(X_n\mid X_0)=\mathcal{N}\!\left(\sqrt{\bar{\alpha}_n}X_0,(1-\bar{\alpha}_n)\Sigma\right), \qquad q(X_{n-1}\mid X_n,X_0)=\mathcal{N}(\tilde{\mu}_n,\tilde{\beta}_n\Sigma),

    where

    μ~n=αˉn−1 βn1−αˉnX0+αn(1−αˉn−1)1−αˉnXn,β~n=1−αˉn−11−αˉnβn.\tilde{\mu}_n= \frac{\sqrt{\bar{\alpha}_{n-1}}\,\beta_n}{1-\bar{\alpha}_n}X_0+ \frac{\sqrt{\alpha_n}(1-\bar{\alpha}_{n-1})}{1-\bar{\alpha}_n}X_n, \qquad \tilde{\beta}_n=\frac{1-\bar{\alpha}_{n-1}}{1-\bar{\alpha}_n}\beta_n.

    If Σ=LLT\Sigma=LL^\mathsf{T} is its Cholesky factorization and ϵ~\tilde{\epsilon} is a standard normal vector of the same dimension as X0X_0, then Xn=αˉnX0+1−αˉnLϵ~X_n=\sqrt{\bar{\alpha}_n}X_0+\sqrt{1-\bar{\alpha}_n}L\tilde{\epsilon}. The reverse model uses p(Xn−1∣Xn)=N(μθ(Xn,t,n),βnΣ)p(X_{n-1}\mid X_n)=\mathcal{N}(\mu_\theta(X_n,t,n),\beta_n\Sigma) and predicts ϵ~\tilde{\epsilon} from the noisy series, its observation times tt, and step nn. Training uses the squared-error objective EX0,t,n,ϵ~∥ϵ~θ(Xn,t,n)−ϵ~∥22\mathbb{E}_{X_0,t,n,\tilde{\epsilon}}\|\tilde{\epsilon}_\theta(X_n,t,n)-\tilde{\epsilon}\|_2^2, with nn sampled from the diffusion steps and XnX_n formed using the forward marginal.

  4. Knowl 4 — Continuous stochastic-process diffusion

    equation

    Continuous stochastic-process diffusion (CSPD) extends variance-preserving SDE diffusion to correlated process noise. Let s∈[0,S]s\in[0,S] denote diffusion time, distinct from the observation times tit_i; let β(s)>0\beta(s)>0 be the noise schedule; let Σ=LLT\Sigma=LL^\mathsf{T} be the process-noise covariance and its factorization; and let WsW_s be a standard Wiener process with the dimension of the flattened series. The forward SDE is

    dXs=−12β(s)Xs ds+β(s)L dWs.dX_s=-\tfrac12\beta(s)X_s\,ds+\sqrt{\beta(s)}L\,dW_s.

    Conditioned on the clean series X0X_0, its marginal at diffusion time ss is Gaussian:

    q(Xs∣X0)=N(μ~s,Σ~s),μ~s=X0exp⁡ ⁣[−12∫0sβ(r) dr],Σ~s=Σ(1−exp⁡ ⁣[−∫0sβ(r) dr]).q(X_s\mid X_0)=\mathcal{N}(\tilde{\mu}_s,\tilde{\Sigma}_s),\qquad \tilde{\mu}_s=X_0\exp\!\left[-\tfrac12\int_0^s\beta(r)\,dr\right],\qquad \tilde{\Sigma}_s=\Sigma\left(1-\exp\!\left[-\int_0^s\beta(r)\,dr\right]\right).

    The corresponding conditional score is ∇Xslog⁡q(Xs∣X0)=−Σ~s−1(Xs−μ~s)\nabla_{X_s}\log q(X_s\mid X_0)=-\tilde{\Sigma}_s^{-1}(X_s-\tilde{\mu}_s). A neural network receives the noisy series, observation times, and diffusion time, and is trained to estimate this score by squared-error score matching. The schedule and terminal time are chosen so that the signal decays and the terminal state approaches the process-noise prior.

  5. Knowl 5 — Parallel conditional forecasting with process diffusion

    model/method

    For forecasting, the model conditions its reverse diffusion on a history window XHX^H and predicts a sequence of future values XFX^F at supplied future timestamps tFt^F. An RNN encodes the history into a vector zz. The denoiser receives the noisy future sequence XnFX_n^F, tFt^F, diffusion step nn, and zz, and predicts the noise for all future values together. The proposed forecasting architecture uses two-dimensional convolutions over feature and time dimensions; the output values are therefore generated in parallel rather than autoregressively, and the denoiser can represent dependencies among predicted future points. Sampling starts with XNFX_N^F drawn from a GP prior evaluated at tFt^F, then iteratively applies the conditioned reverse process to produce a forecast. Repeated sampling gives empirical forecast distributions and uncertainty intervals.

  6. Knowl 6 — Conditional diffusion as a neural process

    model/method

    The method can define a conditional distribution over function values at arbitrary query times. Let (XA,tA)(X^A,t^A) be observed context points and tBt^B be query times whose values XBX^B are unknown. A deterministic, permutation-invariant encoder maps the observed set to a latent vector zz. The reverse diffusion is conditioned on zz; it begins with query values sampled from a GP prior at tBt^B and denoises them to obtain samples from p(XB∣XA)p(X^B\mid X^A). The denoising network is transformer-like and uses a learnable RBF similarity between observed and query times to transfer context information. Training includes examples corresponding to learning p(XA∪XB∣XA)p(X^A\cup X^B\mid X^A), which encourages high certainty at observed locations. The resulting model supports interpolation or imputation at arbitrary query times while representing uncertainty over curves.

  7. Knowl 7 — Synthetic unconditional generation results

    empirical result

    Unconditional generation was evaluated on six synthetic datasets of 10,000 series each, covering stochastic processes, dynamical systems, and chaotic systems. A discriminator was trained to distinguish real test data from generated samples; accuracy near 0.50.5 indicates that the samples are difficult to distinguish. The DSPD-GP model achieved near-chance discriminator accuracy across all six datasets, whereas CTFP and latent ODE were often much easier to distinguish. In a separate CSPD noise-source ablation, GP- or OU-process noise usually gave lower negative log-likelihood than independent Gaussian noise on the more complex Lorenz, predator–prey, sine, and sink data; the advantage was less evident for CIR and OU data. The discriminator accuracies below compare CTFP, latent ODE, and DSPD-GP, respectively:

    Dataset CTFP Latent ODE DSPD-GP
    CIR 0.998±0.0010.998\pm0.001 1.0±0.01.0\pm0.0 0.511±0.0280.511\pm0.028
    Lorenz 0.995±0.0060.995\pm0.006 0.998±0.0020.998\pm0.002 0.513±0.0280.513\pm0.028
    OU 0.783±0.0760.783\pm0.076 0.512±0.0330.512\pm0.033 0.505±0.0450.505\pm0.045
    Predator-prey 0.789±0.0230.789\pm0.023 0.958±0.0210.958\pm0.021 0.585±0.0220.585\pm0.022
    Sine 0.981±0.010.981\pm0.01 1.0±0.01.0\pm0.0 0.525±0.0090.525\pm0.009
    Sink 0.726±0.1380.726\pm0.138 0.907±0.0390.907\pm0.039 0.513±0.010.513\pm0.01

    The DSPD-GP scores show the strongest overall match to the synthetic data under this discriminator test; the predator–prey score, 0.585±0.0220.585\pm0.022, is its largest departure from chance.

  8. Knowl 8 — Forecasting accuracy on real-world datasets

    empirical result

    The parallel conditional forecast model was compared with TimeGrad on Electricity, Exchange, and Solar data. The reported metrics are NRMSE and energy score, averaged over five runs; lower values are better. The model had lower NRMSE than TimeGrad on all three datasets. Its energy score was lower on Electricity and Exchange but higher on Solar, so the results do not show improvement on every reported metric.

    Dataset Metric TimeGrad Proposed model
    Electricity NRMSE 0.064±0.0070.064\pm0.007 0.045±0.0020.045\pm0.002
    Electricity Energy score 8425±6138425\pm613 7079±1647079\pm164
    Exchange NRMSE 0.013±0.0030.013\pm0.003 0.012±0.0010.012\pm0.001
    Exchange Energy score 0.057±0.0020.057\pm0.002 0.031±0.0020.031\pm0.002
    Solar NRMSE 0.799±0.0960.799\pm0.096 0.757±0.0260.757\pm0.026
    Solar Energy score 150±17150\pm17 166±12166\pm12
  9. Knowl 9 — Gaussian-process neural-process evaluation

    empirical result

    The conditional curve-generation model was evaluated on data sampled from Gaussian processes with varying kernel parameters and series lengths. The experiment used 8,000 training series and 2,000 test series; each series had between 5 and 49 points, sampled at times uniformly on [0,1][0,1], and half its points were used as context. On unobserved values, the diffusion model achieved a quantile loss of 0.7370.737, compared with 0.8450.845 for the true-GP model used as the reference. The authors interpret the lower loss as evidence that the diffusion model captures the conditional process in this experiment.

  10. Knowl 10 — Process-noise diffusion for multivariate imputation

    empirical result

    For multivariate imputation, each component at each timestamp is marked observed or missing. The model conditions on observed values and generates the missing values with a reverse diffusion process. In the Physionet experiment, the authors kept the CSDI training setup, architecture, and random seeds, and changed the noise source to GP noise in the loss and sampling. The reported test-set RMSEs are shown below for three missingness rates; lower is better. The process-noise version improves the 10%-missing result, matches CSDI at 50%, and is slightly lower at 90%.

    Missingness CSDI DSPD-GP
    10% 0.520±0.0550.520\pm0.055 0.498±0.0360.498\pm0.036
    50% 0.644±0.0240.644\pm0.024 0.644±0.0290.644\pm0.029
    90% 0.818±0.020.818\pm0.02 0.815±0.0190.815\pm0.019

    These results come from the hourly-sampled Physionet data and show the effect of replacing independent noise with correlated process noise within the same imputation framework.

Coverage note — Supplementary OU sampling trade-offs, detailed synthetic-data generation parameters, and auxiliary architecture comparisons are omitted because they are implementation or supporting-experiment details rather than additional main contributions.

References

  1. 1.Anand, N. and Achim, T. Protein structure and sequence generation with equivariant denoising diffusion probabilistic models. arXiv preprint arXiv:2205.15019, 2022.
  2. 2.Anderson, B. D. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
  3. 3.Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R. Structured denoising diffusion models in discrete state-spaces. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  4. 4.Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  5. 5.Chung, H., Sim, B., and Ye, J. C. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  6. 6.Chung, J., Gulçehre, Ç., Cho, K., and Bengio, Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014.
  7. 7.Conover, W. J. Practical nonparametric statistics, volume 350. john wiley & sons, 1999.
  8. 8.de Bezenac, E., Rangapuram, S. S., Benidis, K., Bohlke-Schneider, M., Kurle, R., Stella, L., Hasson, H., Gallinari, P., and Januschowski, T. Normalizing kalman filters for multivariate time series analysis. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  9. 9.De Finetti, B. La prévision: ses lois logiques, ses sources subjectives. In Annales de l’institut Henri Poincaré, volume 7, pp. 1–68, 1937.
  10. 10.Deng, R., Chang, B., Brubaker, M. A., Mori, G., and Lehrmann, A. Modeling continuous stochastic processes with dynamic normalizing flows. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  11. 11.Deng, R., Brubaker, M. A., Mori, G., and Lehrmann, A. Continuous latent process flows. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  12. 12.Dhariwal, P. and Nichol, A. Q. Diffusion models beat gans on image synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  13. 13.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real NVP. In International Conference on Learning Representations (ICLR), 2017.
  14. 14.Dutordoir, V., Saul, A., Ghahramani, Z., and Simpson, F. Neural diffusion processes. arXiv preprint arXiv:2206.03992, 2022.
  15. 15.Garnelo, M., Schwarz, J., Rosenbaum, D., Viola, F., Rezende, D. J., Eslami, S., and Teh, Y. W. Neural processes. In ICML 2018 workshop on Theoretical Foundations and Applications of Deep Generative Models, 2018.
  16. 16.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378, 2007.
  17. 17.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020.
  18. 18.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  19. 19.Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A. Meta-learning in neural networks: A survey. IEEE transactions on pattern analysis and machine intelligence, 44(9):5149–5169, 2021.
  20. 20.Jolicoeur-Martineau, A., Li, K., Piché-Taillefer, R., Kachman, T., and Mitliagkas, I. Gotta go fast when generating data with score-based models. arXiv preprint arXiv:2105.14080, 2021.
  21. 21.Kerrigan, G., Ley, J., and Smyth, P. Diffusion generative models in infinite dimensions. International Conference on Artificial Intelligence and Statistics (AISTATS), 2023.
  22. 22.Kidger, P., Foster, J., Li, X., and Lyons, T. J. Neural sdes as infinite-dimensional gans. In International Conference on Machine Learning (ICML), 2021.
  23. 23.Kim, H., Mnih, A., Schwarz, J., Garnelo, M., Eslami, S. M. A., Rosenbaum, D., Vinyals, O., and Teh, Y. W. Attentive neural processes. In International Conference on Learning Representations, (ICLR), 2019.
  24. 24.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  25. 25.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Representations (ICLR), 2014.
  26. 26.Kobyzev, I., Prince, S. J., and Brubaker, M. A. Normalizing flows: An introduction and review of current methods. IEEE transactions on pattern analysis and machine intelligence, 43(11):3964–3979, 2020.
  27. 27.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations (ICLR), 2021.
  28. 28.Koochali, A., Dengel, A., and Ahmed, S. If you like it, GAN it—probabilistic multivariate times series forecast with GAN. Engineering Proceedings, 5(1), 2021.
  29. 29.Koochali, A., Schichtel, P., Dengel, A., and Ahmed, S. Random noise vs. state-of-the-art probabilistic forecasting methods: A case study on crps-sum discrimination ability. Applied Sciences, 12(10):5104, 2022.
  30. 30.Lai, G., Chang, W., Yang, Y., and Liu, H. Modeling long- and short-term temporal patterns with deep neural networks. In ACM SIGIR Conference on Research & Development in Information Retrieval, 2018.
  31. 31.Lee, J. S. and Kim, P. M. Proteinsgm: Score-based generative modeling for de novo protein design. bioRxiv, 2022.
  32. 32.Li, X., Wong, T.-K. L., Chen, R. T. Q., and Duvenaud, D. Scalable gradients for stochastic differential equations. In International Conference on Artificial Intelligence and Statistics, 2020.
  33. 33.Lyu, Z., Xu, X., Yang, C., Lin, D., and Dai, B. Accelerating diffusion models via early stop of the diffusion process. arXiv preprint arXiv:2205.12524, 2022.
  34. 34.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning (ICML), 2021a.
  35. 35.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning (ICML), 2021b.
  36. 36.Oksendal, B. Stochastic differential equations: an introduction with applications. Springer Science & Business Media, 2013.
  37. 37.Phillips, A., Seror, T., Hutchinson, M., De Bortoli, V., Doucet, A., and Mathieu, E. Spectral diffusion processes. arXiv preprint arXiv:2209.14125, 2022.
  38. 38.Quinonero-Candela, J. and Rasmussen, C. E. A unifying view of sparse approximate gaussian process regression. Journal of Machine Learning Research, 6(65):1939–1959, 2005.
  39. 39.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022.
  40. 40.Ramos, A. G. C. P., Mehrotra, A., Lane, N. D., and Bhattacharya, S. Conditioning sequence-to-sequence networks with learned activations. In International Conference on Learning Representations (ICLR), 2022.
  41. 41.Rasmussen, C. E. and Williams, C. K. I. Gaussian Processes for Machine Learning. The MIT Press, 2005.
  42. 42.Rasul, K., Seward, C., Schuster, I., and Vollgraf, R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International Conference on Machine Learning (ICML), 2021a.
  43. 43.Rasul, K., Sheikh, A.-S., Schuster, I., Bergmann, U. M., and Vollgraf, R. Multivariate probabilistic time series forecasting via conditioned normalizing flows. In International Conference on Learning Representations (ICLR), 2021b.
  44. 44.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. arXiv preprint arXiv:2112.10752, 2021.
  45. 45.Rubanova, Y., Chen, R. T. Q., and Duvenaud, D. K. Latent ordinary differential equations for irregularly-sampled time series. In Advances in Neural Information Processing Systems (NeurIPS), volume 32, 2019.
  46. 46.Salinas, D., Bohlke-Schneider, M., Callot, L., Medico, R., and Gasthaus, J. High-dimensional multivariate forecasting with low-rank Gaussian Copula Processes. In Advances in Neural Information Processing Systems (NeurIPS). 2019a.
  47. 47.Salinas, D., Bohlke-Schneider, M., Callot, L., Medico, R., and Gasthaus, J. High-dimensional multivariate forecasting with low-rank gaussian copula processes. Advances in Neural Information Processing Systems (NeurIPS), 2019b.
  48. 48.Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020.
  49. 49.Silva, I., Moody, G., Scott, D. J., Celi, L. A., and Mark, R. G. Predicting in-hospital mortality of icu patients: The physionet/computing in cardiology challenge 2012. In 2012 Computing in Cardiology. IEEE, 2012.
  50. 50.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In International Conference on Machine Learning (ICML), 2015.
  51. 51.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021.
  52. 52.Särkkä, S. and Solin, A. Applied Stochastic Differential Equations. Institute of Mathematical Statistics Textbooks. Cambridge University Press, 2019.
  53. 53.Tashiro, Y., Song, J., Song, Y., and Ermon, S. Csdi: Conditional score-based diffusion models for probabilistic time series imputation. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  54. 54.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  55. 55.Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J. Deep sets. In Advances in Neural Information Processing Systems (NeurIPS), 2017.

Citation

MLA
Biloš, M., et al. “Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion”. International Conference on Machine Learning, vol. 202, 2023, pp. 2452–70, https://proceedings.mlr.press/v202/bilos23a.html.
APA
Biloš, M., Rasul, K., Schneider, A., Nevmyvaka, Y., & Günnemann, S. (2023). Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion. International Conference on Machine Learning, 202, 2452–2470. https://proceedings.mlr.press/v202/bilos23a.html
Chicago
Biloš, M., K. Rasul, A. Schneider, Y. Nevmyvaka, and S. Günnemann. 2023. “Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion”. International Conference on Machine Learning 202: 2452–70. https://proceedings.mlr.press/v202/bilos23a.html.
Harvard
Biloš, M. et al. (2023) “Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion”, International Conference on Machine Learning. PMLR, pp. 2452–2470. Available at: https://proceedings.mlr.press/v202/bilos23a.html.
Vancouver
1. Biloš M, Rasul K, Schneider A, Nevmyvaka Y, Günnemann S (2023) Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion. In: International Conference on Machine Learning. PMLR, pp 2452–2470

BibTeX

@InProceedings{pmlr-v202-bilos23a,
  title = 	 {Modeling Temporal Data as Continuous Functions with Stochastic Process Diffusion},
  author =       {Bilo\v{s}, Marin and Rasul, Kashif and Schneider, Anderson and Nevmyvaka, Yuriy and G\"{u}nnemann, Stephan},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {2452--2470},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/bilos23a/bilos23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/bilos23a.html},
  abstract = 	 {Temporal data such as time series can be viewed as discretized measurements of the underlying function. To build a generative model for such data we have to model the stochastic process that governs it. We propose a solution by defining the denoising diffusion model in the function space which also allows us to naturally handle irregularly-sampled observations. The forward process gradually adds noise to functions, preserving their continuity, while the learned reverse process removes the noise and returns functions as new samples. To this end, we define suitable noise sources and introduce novel denoising and score-matching models. We show how our method can be used for multivariate probabilistic forecasting and imputation, and how our model can be interpreted as a neural process.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/