Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes

Yifan ChenMark GoldsteinMengjian HuaMichael S. AlbergoNicholas Matthew BoffiEric Vanden-Eijnden

article2024ICML70 citations

Develops a generative framework using stochastic interpolants and Föllmer processes to construct stochastic differential equations that map point-mass state measurements directly to predictive distributions for high-dimensional dynamical systems and video forecasting.

Listen

Forecasting the future evolution of complex systems—such as weather patterns, fluid flows, and video sequences—is critical across science and industry. In many real-world environments, incomplete measurements, chaotic dynamics, and inherent random noise make deterministic single-point predictions insufficient or misleading. Decision-makers increasingly require probabilistic forecasting that delivers an ensemble of possible future outcomes to quantify risk and uncertainty accurately.

The article develops a new generative modeling framework for probabilistic forecasting of dynamical systems. It evaluates how artificial stochastic dynamics, constructed using stochastic interpolants, can reliably map an exact current state to a conditional probability distribution over future states without bias.

The authors approach this problem by designing stochastic differential equations that start directly at the observed data point rather than transforming pure random noise. The drift vector fields governing the process are learned efficiently using standard regression over paired time-series data via neural networks. The method is validated on three benchmark tasks: a synthetic multi-modal jump-diffusion process, two-dimensional stochastically forced fluid flow governed by the Navier-Stokes equations, and high-dimensional video frame prediction on the KTH human action and CLEVRER physics collision datasets.

The analysis yields four key findings. First, the framework accurately captures complex multi-modal distributions and physical invariants, matching the enstrophy spectrum in fluid simulations even when initialized from downsampled, low-resolution data. Second, the method accelerates fluid flow forecasting by over 100 times compared to direct physical numerical simulation (0.05 seconds versus 8.0 seconds per step). Third, the model significantly outperforms deterministic baselines in long-term stability; for example, on long-horizon fluid dynamics, its relative total enstrophy error is 0.56% compared to 30.0% for deterministic regression. Fourth, in video forecasting benchmarks, the method achieved superior performance over standard flow-matching baselines, reducing the Fréchet Video Distance score from 41.88 to 39.13 on the KTH dataset and from 48.96 to 39.31 on the CLEVRER dataset at 250,000 training steps.

These results demonstrate that generative stochastic forecasting can substantially reduce computational costs and runtime while improving fidelity in risk-sensitive domains. The ability to tune diffusion schedules post-training—recovering an optimal Föllmer process that minimizes estimation errors without retraining the model—offers practical operational flexibility. Organizations relying on expensive physical simulators or complex time-series forecasting can leverage this framework to generate fast, physics-consistent ensembles.

Teams implementing this methodology should adopt quadratic time-interpolant coefficients during training to maintain stable gradient norms and prevent numerical instabilities. In operational deployments, practitioners can generate sequential trajectories autoregressively without retraining. Further investigation is recommended to validate the approach on empirical observational datasets, such as global weather data, and to explore variable forecasting lag times before production deployment.

While the theoretical guarantees and empirical results demonstrate high confidence, the findings are bounded by the stationary or Markovian assumptions present in the evaluated datasets. Additionally, in video applications, overall visual quality remains constrained by the fidelity of underlying image autoencoders.

No sufficiently relevant recommendations were found.

Cover for Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes

Abstract

We propose a framework for probabilistic forecasting of dynamical systems based on generative modeling. Given observations of the system state over time, we formulate the forecasting problem as sampling from the conditional distribution of the future system state given its current state. To this end, we leverage the framework of stochastic interpolants, which facilitates the construction of a generative model between an arbitrary base distribution and the target. We design a fictitious, non-physical stochastic dynamics that takes as initial condition the current system state and produces as output a sample from the target conditional distribution in finite time and without bias. This process therefore maps a point mass centered at the current state onto a probabilistic ensemble of forecasts. We prove that the drift coefficient entering the stochastic differential equation (SDE) achieving this task is non-singular, and that it can be learned efficiently by square loss regression over the time-series data. We show that the drift and the diffusion coefficients of this SDE can be adjusted after training, and that a specific choice that minimizes the impact of the estimation error gives a Föllmer process. We highlight the utility of our approach on several complex, high-dimensional forecasting problems, including stochastically forced Navier-Stokes and video prediction on the KTH and CLEVRER datasets. The code is available at https://github.com/interpolants/forecasting.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Setup and Main Results
  • 3.1. Conditional PDF
  • 3.2. Generation with Stochastic Interpolants
  • 3.3. Generalizations with Tunable Diffusion
  • 3.4. KL Optimization and Föllmer Processes
  • 3.5. Implementation
  • 4. Numerical illustrations
  • 4.1. Multi-modal jump diffusion process
  • 4.2. Forecasting the 2d Navier-Stokes Equations
  • 4.3. Video Forecasting
  • 5. Conclusion and Future Work
  • Acknowledgments
  • References
  • A. Discussion on the stationarity assumption
  • B. Details on stochastic interpolants
  • B.1. Analytical formulas for some specific αs, βs, σs
  • B.2. Proof of Theorem 3.1
  • B.3. Regularity of the drift at s = 0
  • B.4. Changing the Diffusion Coefficient: Proof of Theorem 3.2
  • B.5. Maximizing the likelihood with respect to the noise schedule
  • B.6. Connection with Föllmer Processes and Proof of Theorem 3.3
  • B.7. Analytic formula of bs for Gaussian mixture distributions
  • C. Details of Numerical Experiments
  • C.1. Multi-modal jump-diffusion process
  • C.2. 2D Stochastic Navier-Stokes Example
  • C.3. Video generation

Knowls

  1. Knowl 1 — Conditional forecasting as a stochastic bridge from an observation to its future distribution

    model/method

    Let (X0,X1)(X_0,X_1) have joint density ρ(x0,x1)\rho(x_0,x_1) on Rd×Rd\mathbb{R}^d\times\mathbb{R}^d, and let ρ0(x0)=∫Rdρ(x0,x1) dx1\rho_0(x_0)=\int_{\mathbb{R}^d}\rho(x_0,x_1)\,dx_1. The forecasting target is the conditional density

    ρc(x1∣x0)=ρ(x0,x1)ρ0(x0).\rho_c(x_1\mid x_0)=\frac{\rho(x_0,x_1)}{\rho_0(x_0)}.

    The method constructs a fictitious interpolation-time process, with s∈[0,1]s\in[0,1], between the observed state x0x_0 and a target x1x_1 drawn from this conditional distribution:

    Is=αsx0+βsx1+σsWs.I_s=\alpha_s x_0+\beta_s x_1+\sigma_s W_s.

    Here WsW_s is a standard dd-dimensional Wiener process independent of (X0,X1)(X_0,X_1), and αs,βs,σs\alpha_s,\beta_s,\sigma_s are scalar differentiable schedules satisfying α0=1\alpha_0=1, β0=0\beta_0=0, α1=0\alpha_1=0, β1=1\beta_1=1, and σ1=0\sigma_1=0. Thus I0=x0I_0=x_0 and I1=x1I_1=x_1, so the conditional law evolves from a point mass at the observation to the desired forecast distribution. The interpolation time ss is distinct from the physical forecast lag τ\tau. A time-independent transition kernel permits training from paired observations even when the time series itself is nonstationary; a nonstationary, non-Markovian process can instead be handled by conditioning the drift on physical time.

  2. Knowl 2 — An exactly forecasting SDE whose drift is the square-loss regression target

    theoretical result

    Assume (X0,X1)(X_0,X_1) has finite second moments, WsW_s is a standard dd-dimensional Wiener process independent of the pair, and the differentiable interpolation schedules satisfy the endpoint conditions α0=β1=1\alpha_0=\beta_1=1, α1=β0=σ1=0\alpha_1=\beta_0=\sigma_1=0. Define the interpolation velocity

    Rs=α˙sX0+β˙sX1+σ˙sWs,R_s=\dot\alpha_s X_0+\dot\beta_s X_1+\dot\sigma_s W_s,

    where dots denote derivatives with respect to ss. The unique minimizer, up to equality almost everywhere, of the integrated square-loss objective

    Lb[b^]=∫01E ⁣[∣b^s(Is,X0)−Rs∣2]dsL_b[\widehat b]=\int_0^1\mathbb{E}\!\left[\left|\widehat b_s(I_s,X_0)-R_s\right|^2\right]ds

    is the conditional expectation bs(x,x0)=E[Rs∣Is=x,X0=x0]b_s(x,x_0)=\mathbb{E}[R_s\mid I_s=x,X_0=x_0]. The SDE

    dXs=bs(Xs,x0) ds+σs dWs,X0=x0,dX_s=b_s(X_s,x_0)\,ds+\sigma_s\,dW_s,\qquad X_0=x_0,

    has Law⁡(Xs∣X0=x0)=Law⁡(Is∣X0=x0)\operatorname{Law}(X_s\mid X_0=x_0)=\operatorname{Law}(I_s\mid X_0=x_0) for every s∈[0,1]s\in[0,1]; consequently X1X_1 is an unbiased sample from ρc(⋅∣x0)\rho_c(\cdot\mid x_0). The regression target can be simulated from paired data and Wiener noise without evaluating a density or its normalization. Under the finite-moment assumption and bounded schedule derivatives, the target has finite second moment and the square-loss estimator has bounded variance.

  3. Knowl 3 — Post-training diffusion changes preserve the conditional forecast law

    theoretical result

    Let ρs(x∣x0)\rho_s(x\mid x_0) be the density of the interpolant IsI_s conditional on X0=x0X_0=x_0, and let bsb_s be the exact regression drift. For a continuous scalar diffusion schedule gsg_s satisfying finite endpoint limits of s−1(gs2−σs2)s^{-1}(g_s^2-\sigma_s^2) as s→0+s\to0^+ and gs2/σsg_s^2/\sigma_s as s→1−s\to1^-, define

    bsg(x,x0)=bs(x,x0)+12(gs2−σs2)∇xlog⁡ρs(x∣x0).b_s^g(x,x_0)=b_s(x,x_0)+\frac12(g_s^2-\sigma_s^2)\nabla_x\log\rho_s(x\mid x_0).

    Then the SDE dXsg=bsg(Xsg,x0) ds+gs dWsdX_s^g=b_s^g(X_s^g,x_0)\,ds+g_s\,dW_s, initialized at X0g=x0X_0^g=x_0, has the same conditional marginal law as the original interpolant at every interpolation time. In particular, its endpoint still has density ρc(⋅∣x0)\rho_c(\cdot\mid x_0). The score needed to adjust the drift can be recovered from the learned drift, without separately training a score model:

    ∇xlog⁡ρs(x∣x0)=As[βsbs(x,x0)−cs(x,x0)],As=1sσs(β˙sσs−βsσ˙s),cs(x,x0)=β˙sx+(βsα˙s−β˙sαs)x0.\nabla_x\log\rho_s(x\mid x_0)=A_s\bigl[\beta_s b_s(x,x_0)-c_s(x,x_0)\bigr], \quad A_s=\frac{1}{s\sigma_s(\dot\beta_s\sigma_s-\beta_s\dot\sigma_s)}, \quad c_s(x,x_0)=\dot\beta_s x+(\beta_s\dot\alpha_s-\dot\beta_s\alpha_s)x_0.

    Here x,x0∈Rdx,x_0\in\mathbb{R}^d, s∈(0,1)s\in(0,1), and the schedules and their derivatives are those defining the interpolant. This identity allows changing both the drift and diffusion at sampling time while retaining the target conditional distribution.

  4. Knowl 4 — The path-KL-optimal diffusion yields a Föllmer process

    theoretical result

    Suppose an estimated drift b^s\widehat b_s approximates the exact interpolant drift bsb_s, and let Ls=E[∣b^s(Is,X0)−bs(Is,X0)∣2]L_s=\mathbb{E}[|\widehat b_s(I_s,X_0)-b_s(I_s,X_0)|^2] denote its mean squared error under the interpolant distribution. For a candidate diffusion schedule gsg_s, the Kullback–Leibler divergence between the path law of the exact forecast SDE and that of the corresponding SDE using b^s\widehat b_s is

    DKL(Xg∥X^g)=12∫01∣1+12βsAs(gs2−σs2)∣2gs2 Ls ds,D_{\mathrm{KL}}(X^g\Vert\widehat X^g) =\frac12\int_0^1\frac{\left|1+\frac12\beta_sA_s(g_s^2-\sigma_s^2)\right|^2}{g_s^2}\,L_s\,ds,

    where As=[sσs(β˙sσs−βsσ˙s)]−1A_s=[s\sigma_s(\dot\beta_s\sigma_s-\beta_s\dot\sigma_s)]^{-1}, and σs,βs\sigma_s,\beta_s are the interpolant schedules. Since LsL_s does not depend on gsg_s, the minimizing schedule is

    gsF=∣2sσs(βs−1β˙sσs−σ˙s)−σs2∣1/2.g_s^F=\left|2s\sigma_s\left(\beta_s^{-1}\dot\beta_s\sigma_s-\dot\sigma_s\right)-\sigma_s^2\right|^{1/2}.

    If βs/(s σs)\beta_s/(\sqrt{s}\,\sigma_s) is nondecreasing on [0,1][0,1], the process using gsFg_s^F is a Föllmer process: it is the endpoint-constrained path-KL minimizer relative to the linear Gaussian reference diffusion specified by the interpolant schedules. Thus the Föllmer construction also arises as the diffusion choice that minimizes the effect of drift-estimation error on the forecasting path law.

  5. Knowl 5 — Data-driven training, Euler sampling, and autoregressive rollout

    algorithm

    For time-series observations (xkτ,x(k+1)τ)(x_{k\tau},x_{(k+1)\tau}), train a neural drift estimate b^s(x,x0)\widehat b_s(x,x_0) using minibatches of paired states. Draw ss uniformly from [0,1][0,1] and z∼N(0,Id)z\sim\mathcal{N}(0,I_d) independently for each pair, and form

    Isk=αsxkτ+βsx(k+1)τ+s σsz,Rsk=α˙sxkτ+β˙sx(k+1)τ+s σ˙sz.I_s^k=\alpha_sx_{k\tau}+\beta_sx_{(k+1)\tau}+\sqrt{s}\,\sigma_s z, \qquad R_s^k=\dot\alpha_sx_{k\tau}+\dot\beta_sx_{(k+1)\tau}+\sqrt{s}\,\dot\sigma_s z.

    Optimize the average of ∣b^s(Isk,xkτ)−Rsk∣2|\widehat b_s(I_s^k,x_{k\tau})-R_s^k|^2 over minibatch pairs and sampled interpolation times. The experiments principally use αs=1−s\alpha_s=1-s, σs=ε(1−s)\sigma_s=\varepsilon(1-s) for ε>0\varepsilon>0, and βs=s2\beta_s=s^2; the zero initial derivative β˙0=0\dot\beta_0=0 was found to stabilize training gradients. After training, choose gsg_s and construct b^sg\widehat b_s^g from the learned drift and its score identity. For a new observation, integrate the SDE by Euler–Maruyama on a grid 0=s0<⋯<sN=10=s_0<\cdots<s_N=1, using independent ηn∼N(0,Id)\eta_n\sim\mathcal{N}(0,I_d) and increments b^sng(Xn,x0)Δsn+gsnΔsnηn\widehat b_{s_n}^g(X_n,x_0)\Delta s_n+g_{s_n}\sqrt{\Delta s_n}\eta_n. At the initial step, use the finite initial drift directly with diffusion σs0\sigma_{s_0} to avoid evaluating the score-based expression at its potentially singular endpoint. The terminal state is a forecast sample. To roll out multiple physical lags, use each generated endpoint as the next initial observation and reuse the same trained drift; no retraining is required.

  6. Knowl 6 — Multimodal jump-diffusion forecasts reproduce conditional ensembles

    empirical result

    A synthetic two-dimensional process was constructed with a five-component Gaussian-mixture invariant density. Between jumps the particle follows Langevin dynamics for that mixture; at Poisson jump times of rate 22, its state is rotated counterclockwise by 2π/52\pi/5. The model was trained on 10510^5 successive state pairs at physical lag τ=0.5\tau=0.5, generated after the process reached equilibrium, using a fully connected network with five hidden layers of width 500. Forecast ensembles closely matched the process's multimodal conditional distributions at lags 0.50.5, 11, and 22. Repeated application of the learned forecaster also reproduced the relaxation toward the five-mode invariant distribution beyond the decorrelation time. In this setting a single conditional-mean forecast would obscure the multiple possible outcomes, whereas the stochastic model generated distinct samples consistent with the conditional law.

  7. Knowl 7 — Stochastically forced Navier–Stokes forecasting experiment

    experimental setup

    The fluid experiment forecasts vorticity ω\omega on the two-dimensional torus [0,2π]2[0,2\pi]^2 under the stochastically forced Navier–Stokes equation

    dω+v⋅∇ω dt=νΔω dt−αω dt+ε dη,v=∇⊥ψ,−Δψ=ω.d\omega+v\cdot\nabla\omega\,dt=\nu\Delta\omega\,dt-\alpha\omega\,dt+\varepsilon\,d\eta, \qquad v=\nabla^\perp\psi,\quad -\Delta\psi=\omega.

    Here η\eta is white-in-time forcing on a small set of Fourier modes; the parameters are ν=10−3\nu=10^{-3}, α=0.1\alpha=0.1, and ε=1\varepsilon=1. Simulations used a pseudo-spectral solver on a 256×256256\times256 grid. From 2,000 trajectories over t∈[0,100]t\in[0,100], the initial interval [0,50][0,50] was discarded, and 2×1052\times10^5 snapshots were retained and resized to 128×128128\times128. A roughly two-million-parameter U-Net learned the drift from lagged pairs; the data were split 90%/10% for training and testing, and training used batch size 100, AdamW with base learning rate 10−310^{-3}, and 50 epochs. The main forecast lag was τ=0.5\tau=0.5. Experiments forecast full-resolution fields from both full-resolution observations and 32×3232\times32 low-resolution observations, with the target always at 128×128128\times128.

  8. Knowl 8 — Navier–Stokes forecasts recover ensemble statistics and outperform point prediction

    data/table

    The learned SDE was evaluated on ensembles of forecasts, including conditional means and standard deviations, the enstrophy spectrum across spatial scales, and long-run statistics after iterating the predictor. For 50 initial conditions and ensembles of 300 forecasts, the full-resolution model reproduced the true conditional enstrophy spectrum; with a 32×3232\times32 input it also recovered the high-wavenumber spectrum of the 128×128128\times128 target fields. The table reports relative errors for short-lag conditional statistics at τ=0.5\tau=0.5 and for average total enstrophy after 100 autoregressive steps; the deterministic comparator used the same U-Net architecture and an MSE-trained point estimate.

    Could not parse LaTeX table

    The ensemble forecaster therefore better estimates the conditional mean and preserves fluctuations that the point predictor cannot represent; its long-run average-enstrophy error is also much smaller. For a lag of 0.50.5, sampling the learned model with 200 Euler–Maruyama steps took 0.05 seconds, compared with 8 seconds for direct stochastic-PDE simulation on the same Nvidia RTX8000 GPU—over a 100-fold speedup in the reported test.

  9. Knowl 9 — Sparse-context latent video forecasting

    experimental setup

    Video frames were encoded into latent images by pretrained VQGANs, and the forecasting model generated one latent frame at a time before decoding it. To avoid conditioning on an entire frame history, prediction of frame tt conditions on the preceding latent frame yt−1y^{t-1} and one randomly selected earlier latent frame yt−jy^{t-j}, with its time index t−jt-j also supplied to the model; the earlier-frame index is sampled from j∈{2,…,t−1}j\in\{2,\ldots,t-1\}. The latent stochastic interpolant starts at yt−1y^{t-1}, and generated frames are rolled out autoregressively. KTH videos were encoded from 1×64×641\times64\times64 frames to 4×8×84\times8\times8 latents; generation used 10 given frames and 30 predicted frames. CLEVRER videos were encoded from 3×128×1283\times128\times128 frames to 4×16×164\times16\times16 latents; generation used two given frames and 14 predicted frames. The comparison used the same VQGAN checkpoints as the RIVER baseline, which predicts with conditional flow matching from Gaussian noise. Video models used a U-Net conditioned by channel concatenation and were trained for 250,000 AdamW steps at initial learning rate 2×10−42\times10^{-4}.

  10. Knowl 10 — Latent video forecasting improves FVD on KTH and CLEVRER

    data/table

    Fréchet Video Distance (FVD; lower is better) was measured using 256 test videos and 100 generated completions per video, comparing 256 real videos with 25,600 generated videos. PFI denotes probabilistic forecasting with interpolants, and RIVER is the conditional flow-matching baseline. The autoencoder FVD is the reconstruction floor imposed by modeling in the VQGAN latent space. Shifted FVD subtracts that autoencoder value from each model's FVD to show the distance above this approximate floor. Values are reported after 100,000 and 250,000 gradient steps.

    Could not parse LaTeX table

    PFI has lower FVD than RIVER on both datasets at both training durations. Qualitative forecasts also show varied continuations from the same initial frames while retaining temporal motion patterns; CLEVRER samples depict collision-consistent object interactions, while KTH samples retain recognizable hand-waving motion despite drifting from the exact reference trajectory.

Coverage note — The supplementary proofs, closed-form drift formulas for additional interpolant schedules, and further qualitative plots were omitted because they support or illustrate the stated results rather than add comparably significant standalone contributions.

References

  1. 1.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations, 2022.
  2. 2.Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797, 2023.
  3. 3.Alexander, R. and Giannakis, D. Operator-theoretic framework for forecasting nonlinear time series with kernel analog techniques. Physica D: Nonlinear Phenomena, 409:132520, 2020.
  4. 4.Berry, T., Giannakis, D., and Harlim, J. Nonparametric forecasting of low-dimensional dynamical systems. Physical Review E, 91(3):032915, 2015.
  5. 5.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  6. 6.Buaria, D. and Sreenivasan, K. R. Forecasting small-scale dynamics of fluid turbulence using deep neural networks. Proceedings of the National Academy of Sciences, 120(30):e2305765120, 2023.
  7. 7.Chen, R. T., Amos, B., and Nickel, M. Neural spatio-temporal point processes. arXiv preprint arXiv:2011.04583, 2020.
  8. 8.Chen, Y., Georgiou, T. T., and Pavon, M. Stochastic control liaisons: Richard sinkhorn meets gaspard monge on a schrodinger bridge. Siam Review, 63(2):249–313, 2021.
  9. 9.Davtyan, A., Sameni, S., and Favaro, P. Efficient video prediction via sparsely conditioned flow matching. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 23263–23274, 2023.
  10. 10.de Bezenac, E., Rangapuram, S. S., Benidis, K., Bohlke-Schneider, M., Kurle, R., Stella, L., Hasson, H., Gallinari, P., and Januschowski, T. Normalizing kalman filters for multivariate time series analysis. In Advances in Neural Information Processing Systems, volume 33, pp. 2995–3007, 2020.
  11. 11.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion schrodinger bridge with applications to score-based generative modeling. In Advances in Neural Information Processing Systems, volume 34, pp. 17695–17709, 2021.
  12. 12.Dellnitz, M. and Junge, O. On the approximation of complicated dynamical behavior. SIAM Journal on Numerical Analysis, 36(2):491–515, 1999.
  13. 13.Dresdner, G., Kochkov, D., Norgaard, P., Zepeda-Nuñez, L., Smith, J. A., Brenner, M. P., and Hoyer, S. Learning to correct spectral methods for simulating turbulent flows. arXiv preprint arXiv:2207.00556, 2022.
  14. 14.Eldan, R. and Lee, J. R. Regularization under diffusion and anticoncentration of the information content. Duke Mathematical Journal, 167(5):969–993, 2018.
  15. 15.Eldan, R., Lehec, J., and Shenfeld, Y. Stability of the logarithmic sobolev inequality via the follmer process. In Annales de l’Institut Henri Poincare-Probabilites et Statistiques, volume 56, pp. 2253–2269, 2020.
  16. 16.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12873–12883, 2021.
  17. 17.Finn, C., Goodfellow, I., and Levine, S. Unsupervised learning for physical interaction through video prediction. In Advances in neural information processing systems, volume 29, 2016.
  18. 18.Follmer, H. Time reversal on wiener space. Stochastic Processes—Mathematics and Physics, pp. 119–129, 1986.
  19. 19.Giannakis, D., Henriksen, A., Tropp, J. A., and Ward, R. Learning to forecast dynamical systems from streaming data. SIAM Journal on Applied Dynamical Systems, 22(2):527–558, 2023.
  20. 20.Gneiting, T. and Katzfuss, M. Probabilistic forecasting. Annual Review of Statistics and Its Application, 1:125–151, 2014.
  21. 21.Gneiting, T., Raftery, A. E., Westveld, A. H., and Goldman, T. Calibrated probabilistic forecasting using ensemble model output statistics and minimum crps estimation. Monthly Weather Review, 133(5):1098–1118, 2005.
  22. 22.Gu, A., Goel, K., and Re, C. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  23. 23.Hairer, M. and Mattingly, J. C. Ergodicity of the 2d navier-stokes equations with degenerate stochastic forcing. Annals of Mathematics, pp. 993–1032, 2006.
  24. 24.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in neural information processing systems, volume 30, 2017.
  25. 25.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in neural information processing systems, volume 33, pp. 6840–6851, 2020.
  26. 26.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022a.
  27. 27.Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 23(47):1–33, 2022b.
  28. 28.Huang, J., Jiao, Y., Kang, L., Liao, X., Liu, J., and Liu, Y. Schrodinger-Follmer sampler: sampling without ergodicity. arXiv preprint arXiv:2106.10880, 2021.
  29. 29.Jiang, R., Lu, P. Y., Orlova, E., and Willett, R. Training neural operators to preserve invariant measures of chaotic attractors. arXiv preprint arXiv:2306.01187, 2023.
  30. 30.Jiao, Y., Kang, L., Liu, Y., and Zhou, Y. Convergence analysis of Schrodinger-Follmer sampler without convexity. arXiv preprint arXiv:2107.04766, 2021.
  31. 31.Kaiser, E., Kutz, J. N., and Brunton, S. L. Data-driven discovery of koopman eigenfunctions for control. Machine Learning: Science and Technology, 2(3):035023, 2021.
  32. 32.Kidger, P., Foster, J., Li, X., and Lyons, T. J. Neural sdes as infinite-dimensional gans. In International conference on machine learning, pp. 5453–5463. PMLR, 2021.
  33. 33.Kutz, J. N., Brunton, S. L., Brunton, B. W., and Proctor, J. L. Dynamic mode decomposition: data-driven modeling of complex systems. SIAM, 2016.
  34. 34.Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S. Stochastic adversarial video prediction. arXiv preprint arXiv:1804.01523, 2018.
  35. 35.Lehec, J. Representation formula for the entropy and functional inequalities. In Annales de l’IHP Probabilites et statistiques, volume 49, pp. 885–899, 2013.
  36. 36.Leonard, C. A survey of the schrodinger problem and some of its connections with optimal transport. Discrete and Continuous Dynamical Systems-Series A, 34(4):1533–1574, 2014.
  37. 37.Li, Z., Kovachki, N., Azizzadenesheli, K., Liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A. Fourier neural operator for parametric partial differential equations. arXiv preprint arXiv:2010.08895, 2020.
  38. 38.Li, Z., Liu-Schiaffini, M., Kovachki, N., Liu, B., Azizzadenesheli, K., Bhattacharya, K., Stuart, A., and Anandkumar, A. Learning dissipative dynamics in chaotic systems. arXiv preprint arXiv:2106.06898, 2021.
  39. 39.Lienen, M., Lüdke, D., Hansen-Palmus, J., and Günnemann, S. From zero to turbulence: Generative modeling for 3d flow simulation. In The Twelfth International Conference on Learning Representations, 2023.
  40. 40.Lim, B. and Zohren, S. Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A, 379(2194):20200209, 2021.
  41. 41.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2022.
  42. 42.Liu, G.-H., Vahdat, A., Huang, D.-A., Theodorou, E. A., Nie, W., and Anandkumar, A. I2SB: Image-to-image Schrodinger bridge. arXiv preprint arXiv:2302.05872, 2023.
  43. 43.Liu, X., Gong, C., and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2022.
  44. 44.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  45. 45.Lu, L., Jin, P., Pang, G., Zhang, Z., and Karniadakis, G. E. Learning nonlinear operators via deeponet based on the universal approximation theorem of operators. Nature machine intelligence, 3(3):218–229, 2021.
  46. 46.Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. arXiv preprint arXiv:2401.08740, 2024.
  47. 47.Masini, R. P., Medeiros, M. C., and Mendes, E. F. Machine learning advances for time series forecasting. Journal of economic surveys, 37(1):76–111, 2023.
  48. 48.Mirza, M. and Osindero, S. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  49. 49.Oprea, S., Martinez-Gonzalez, P., Garcia-Garcia, A., Castro-Vargas, J. A., Orts-Escolano, S., Garcia-Rodriguez, J., and Argyros, A. A review on deep learning techniques for video prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6):2806–2826, 2020.
  50. 50.Palmer, T., Molteni, F., Mureau, R., Buizza, R., Chapelet, P., and Tribbia, J. Ensemble prediction. In Proc. ECMWF Seminar on Validation of models over Europe, volume 1, pp. 21–66, 1993.
  51. 51.Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A., Mardani, M., Kurth, T., Hall, D., Li, Z., Azizzadenesheli, K., et al. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators. arXiv preprint arXiv:2202.11214, 2022.
  52. 52.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  53. 53.Peluchetti, S. Non-denoising forward-time diffusions. arXiv preprint arXiv:2312.14589, 2023.
  54. 54.Rasul, K., Seward, C., Schuster, I., and Vollgraf, R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International Conference on Machine Learning, pp. 8857–8868. PMLR, 2021.
  55. 55.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  56. 56.Rühling Cachay, S., Zhao, B., Joren, H., and Yu, R. Dyffusion: A dynamics-informed diffusion model for spatiotemporal forecasting. In Advances in Neural Information Processing Systems, volume 36, 2024.
  57. 57.Schrodinger, E. Sur la theorie relativiste de l’electron et l’interpretation de la mecanique quantique. In Annales de l’institut Henri Poincare, volume 3, pp. 269–310, 1932.
  58. 58.Schuldt, C., Laptev, I., and Caputo, B. Recognizing human actions: a local svm approach. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004., volume 3, pp. 32–36. IEEE, 2004.
  59. 59.Shi, Y., De Bortoli, V., Campbell, A., and Doucet, A. Diffusion schrodinger bridge matching. In Advances in Neural Information Processing Systems, volume 36, 2024.
  60. 60.Smagorinsky, J. General circulation experiments with the primitive equations: I. the basic experiment. Monthly weather review, 91(3):99–164, 1963.
  61. 61.Sohn, K., Lee, H., and Yan, X. Learning structured output representation using deep conditional generative models. In Advances in neural information processing systems, volume 28, 2015.
  62. 62.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  63. 63.Tzen, B. and Raginsky, M. Theoretical guarantees for sampling and inference in generative models with latent diffusions. In Conference on Learning Theory, pp. 3084–3114. PMLR, 2019.
  64. 64.Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018.
  65. 65.Vahdat, A., Kreis, K., and Kautz, J. Score-based generative modeling in latent space. In Advances in neural information processing systems, volume 34, pp. 11287–11302, 2021.
  66. 66.Vargas, F., Ovsianas, A., Fernandes, D., Girolami, M., Lawrence, N. D., and Nusken, N. Bayesian learning via neural schrodinger–follmer flows. Statistics and Computing, 33(1):3, 2023.
  67. 67.Wang, G., Jiao, Y., Xu, Q., Wang, Y., and Yang, C. Deep generative learning via schrodinger bridge. In International Conference on Machine Learning, pp. 10794–10804. PMLR, 2021.
  68. 68.Wanner, M. and Mezic, I. Robust approximation of the stochastic koopman operator. SIAM Journal on Applied Dynamical Systems, 21(3):1930–1951, 2022.
  69. 69.Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B. Clevrer: Collision events for video representation and reasoning. arXiv preprint arXiv:1910.01442, 2019.
  70. 70.Zhang, Q. and Chen, Y. Path integral sampler: A stochastic control approach for sampling. In International Conference on Learning Representations, 2021.
  71. 71.Zhao, M. and Jiang, L. Data-driven probability density forecast for stochastic dynamical systems. Journal of Computational Physics, 492:112422, 2023.

Citation

MLA
Chen, Y., et al. “Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes”. arXiv, 2024, http://arxiv.org/abs/2403.13724v2.
APA
Chen, Y., Goldstein, M., Hua, M., Albergo, M. S., Boffi, N. M., & Vanden-Eijnden, E. (2024). Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes. arXiv. http://arxiv.org/abs/2403.13724v2
Chicago
Chen, Y., M. Goldstein, M. Hua, M. S. Albergo, N. M. Boffi, and E. Vanden-Eijnden. 2024. “Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes”. arXiv. http://arxiv.org/abs/2403.13724v2.
Harvard
Chen, Y. et al. (2024) “Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.13724v2.
Vancouver
1. Chen Y, Goldstein M, Hua M, Albergo MS, Boffi NM, Vanden-Eijnden E (2024) Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes. arXiv

BibTeX

@article{chen2024probabilistic,
  title = {Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes},
  author = {Chen, Yifan and Goldstein, Mark and Hua, Mengjian and Albergo, Michael S. and Boffi, Nicholas M. and Vanden-Eijnden, Eric},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.13724v2},
  eprint = {2403.13724}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/