Score-based Data Assimilation

François RozetGilles Louppe

article2023NeurIPS89 citations

Proposes a score-based data assimilation framework that trains on short trajectory segments to enable non-autoregressive, zero-shot posterior sampling over long time horizons without running or differentiating through physical forward models during inference.

Listen

Complex real-world dynamical systems, such as atmospheric weather patterns and ocean currents, are difficult to track because observational data is often sparse, intermittent, and noisy. Data assimilation seeks to reconstruct complete, physically plausible state trajectories from these incomplete measurements. However, conventional operational methods rely heavily on simulating and differentiating through complex numerical physical models. This requirement becomes computationally intractable over extended time horizons or high-dimensional systems, consuming massive computing budgets and forcing operational agencies to discard large volumes of available satellite data.

The article introduces and evaluates Score-based Data Assimilation, a generative machine learning approach designed to reconstruct full state trajectories without running or differentiating through underlying physical equations during inference. The authors demonstrate that generative models can capture the full probability distribution of possible states directly from data while maintaining physical consistency.

The approach relies on three key mechanisms. First, it exploits the step-by-step structure of dynamical systems to approximate trajectory probabilities using local short segments, allowing score networks to train on small temporal windows rather than complete long sequences. Second, it decouples the observation process from training, using a stabilized likelihood approximation at inference time to enable zero-shot assimilation under varying observation conditions. Third, it generates all trajectory states simultaneously via a predictor-corrector diffusion process that applies corrective steps to eliminate accumulated numerical errors. The authors evaluated this framework on two standard benchmarks: the chaotic three-dimensional Lorenz 1963 atmospheric convection system and a high-dimensional, two-dimensional turbulent fluid simulation governed by Navier-Stokes equations.

The evaluation yielded several key findings. First, the method closely matched ground-truth posterior distributions on the Lorenz system, achieving near-optimal statistical accuracy when using local windows of at least three steps and two or more corrective iterations. Second, unlike traditional variational techniques that produce single-point estimates, the proposed framework successfully recovered multi-modal posterior distributions, capturing multiple distinct and physically plausible operational scenarios from the same observation. Third, in high-dimensional fluid simulations, the framework accurately reconstructed continuous flow fields from heavily degraded, low-resolution, and spatially sparse observations where existing baseline methods failed. Fourth, when fed unlikely terminal conditions, the model generated dynamically valid trajectories whose initial states naturally reproduced the target evolution when verified against the true physical equations.

These findings suggest substantial practical implications for operational forecasting, environmental monitoring, and mission-critical decision-making. By bypassing physical model simulation during inference, the method reduces the computational burden of data assimilation, opening paths to ingest richer satellite streams and parallelize inference on modern hardware. Moreover, capturing multi-modal distributions provides decision-makers with comprehensive risk profiles rather than a single overconfident estimate, improving safety and contingency planning for high-consequence weather events.

Organizations evaluating this technology should begin with pilot deployments on lower-dimensional or regional forecasting tasks to benchmark inference speed and accuracy trade-offs against established Kalman filtering and variational baselines. Practitioners can adjust the number of corrective steps during sampling to trade computational throughput for physical accuracy as needed. However, further research and validation are required before applying the method to full-scale numerical weather prediction. Key limitations include the current assumption of fixed physical parameters across all trajectories, the unquantified theoretical error bounds of the score approximations, and the substantial engineering challenge of scaling the framework from tens of thousands of dimensions to the millions of dimensions used in operational forecasting systems.

arXiv: 2306.10574
Cover for Score-based Data Assimilation

Abstract

Data assimilation, in its most comprehensive form, addresses the Bayesian inverse problem of identifying plausible state trajectories that explain noisy or incomplete observations of stochastic dynamical systems. Various approaches have been proposed to solve this problem, including particle-based and variational methods. However, most algorithms depend on the transition dynamics for inference, which becomes intractable for long time horizons or for high-dimensional systems with complex dynamics, such as oceans or atmospheres. In this work, we introduce score-based data assimilation for trajectory inference. We learn a score-based generative model of state trajectories based on the key insight that the score of an arbitrarily long trajectory can be decomposed into a series of scores over short segments. After training, inference is carried out using the score model, in a non-autoregressive manner by generating all states simultaneously. Quite distinctively, we decouple the observation model from the training procedure and use it only at inference to guide the generative process, which enables a wide range of zero-shot observation scenarios. We present theoretical and empirical evidence supporting the effectiveness of our method.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Score-based data assimilation
  • 3.1 How is your blanket?
  • 3.2 Stable likelihood score
  • 3.3 Predictor-Corrector sampling
  • 4 Results
  • 4.1 Lorenz 1963
  • 4.2 Kolmogorov flow
  • 5 Conclusion
  • References
  • A The pseudo-blanket approximation is unbiased
  • B On the covariance of p⁡(x∣x⁡(t))p(x\mid x(t))
  • C Algorithms
  • D Experiment details
  • E Assimilation examples

Knowls

  1. Knowl 1 — Trajectory Prior Score Approximation via Pseudo-Markov Blankets

    model/method

    In a discrete-time stochastic dynamical system, a state trajectory x1:L=(x1,x2,…,xL)∈RL×D\mathbf{x}_{1:L} = (\mathbf{x}_1, \mathbf{x}_2, \dots, \mathbf{x}_L) \in \mathbb{R}^{L \times D} is a first-order Markov chain whose clean-data score can be decomposed using the Markov blanket xbi={xi−1,xi+1}\mathbf{x}_{b_i} = \{\mathbf{x}_{i-1}, \mathbf{x}_{i+1}\} of each state xi\mathbf{x}_i:

    ∇xilog⁡p(x1:L)=∇xilog⁡p(xi,xbi)\nabla_{\mathbf{x}_i} \log p(\mathbf{x}_{1:L}) = \nabla_{\mathbf{x}_i} \log p(\mathbf{x}_i, \mathbf{x}_{b_i})

    When trajectories are perturbed by continuous-time Gaussian diffusion p(x1:L(t)∣x1:L)=N(x1:L(t)∣μ(t)x1:L,σ(t)2I)p(\mathbf{x}_{1:L}(t) \mid \mathbf{x}_{1:L}) = \mathcal{N}(\mathbf{x}_{1:L}(t) \mid \mu(t) \mathbf{x}_{1:L}, \sigma(t)^2 \mathbf{I}), conditional independence is broken for diffusion time t∈(0,1)t \in (0, 1). However, a local temporal window bˉi={i−k,…,i+k}∖{i}\bar{b}_i = \{i - k, \dots, i + k\} \setminus \{i\} of half-width k≥1k \ge 1 acts as a pseudo-Markov blanket:

    ∇xi(t)log⁡p(x1:L(t))≈∇xi(t)log⁡p(xi−k:i+k(t))\nabla_{\mathbf{x}_i(t)} \log p(\mathbf{x}_{1:L}(t)) \approx \nabla_{\mathbf{x}_i(t)} \log p(\mathbf{x}_{i-k:i+k}(t))

    Because mutual information between states vanishes at t=1t=1 and the exact Markov property holds at t=0t=0, this approximation remains accurate across t∈[0,1]t \in [0, 1] for k≪Lk \ll L. Consequently, a neural score model sϕ(xi−k:i+k(t),t)≈∇xi−k:i+k(t)log⁡p(xi−k:i+k(t))s_\phi(\mathbf{x}_{i-k:i+k}(t), t) \approx \nabla_{\mathbf{x}_{i-k:i+k}(t)} \log p(\mathbf{x}_{i-k:i+k}(t)) trained only on short sub-segments of length 2k+12k + 1 can be applied in a sliding-window fashion along a trajectory to approximate the full score ∇x1:L(t)log⁡p(x1:L(t))\nabla_{\mathbf{x}_{1:L}(t)} \log p(\mathbf{x}_{1:L}(t)) of arbitrarily long sequences non-autoregressively.

  2. Knowl 2 — Unbiasedness of the Pseudo-Markov Blanket Score Approximation

    theoretical result

    For any continuous random variables aa, bb, and cc, the score of the marginal density satisfies:

    ∇alog⁡p(a,b)=1p(a,b)∇a∫p(a,b,c) dc=∫p(c∣a,b)p(a,b,c)∇ap(a,b,c) dc=Ep(c∣a,b)[∇alog⁡p(a,b,c)]\nabla_a \log p(a, b) = \frac{1}{p(a, b)} \nabla_a \int p(a, b, c) \, \mathrm{d}c = \int \frac{p(c \mid a, b)}{p(a, b, c)} \nabla_a p(a, b, c) \, \mathrm{d}c = \mathbb{E}_{p(c \mid a, b)} \left[ \nabla_a \log p(a, b, c) \right]

    Setting a=xi(t)a = \mathbf{x}_i(t), b=xbˉi(t)b = \mathbf{x}_{\bar{b}_i}(t) (the pseudo-Markov blanket of xi(t)\mathbf{x}_i(t)), and c=x∉(t)={xj(t):j≠i∧j∉bˉi}c = \mathbf{x}_{\notin}(t) = \{\mathbf{x}_j(t) : j \neq i \wedge j \notin \bar{b}_i\} (the remaining states in the diffusion-perturbed trajectory x1:L(t)\mathbf{x}_{1:L}(t)), it holds that:

    ∇xi(t)log⁡p(xi(t),xbˉi(t))=Ep(x∉(t)∣xi(t),xbˉi(t))[∇xi(t)log⁡p(x1:L(t))]\nabla_{\mathbf{x}_i(t)} \log p(\mathbf{x}_i(t), \mathbf{x}_{\bar{b}_i}(t)) = \mathbb{E}_{p(\mathbf{x}_{\notin}(t) \mid \mathbf{x}_i(t), \mathbf{x}_{\bar{b}_i}(t))} \left[ \nabla_{\mathbf{x}_i(t)} \log p(\mathbf{x}_{1:L}(t)) \right]

    Thus, the pseudo-Markov blanket score approximation is an unbiased estimator (exact in expectation over all variables outside the pseudo-blanket) of the true full-trajectory score ∇xi(t)log⁡p(x1:L(t))\nabla_{\mathbf{x}_i(t)} \log p(\mathbf{x}_{1:L}(t)), irrespective of the chosen blanket size or composition.

  3. Knowl 3 — Extended Kalman-Inspired Likelihood Score for Score-Based Inverse Problems

    equation

    For an observation process y=A(x1:L)+η\mathbf{y} = \mathcal{A}(\mathbf{x}_{1:L}) + \boldsymbol{\eta} with differentiable measurement operator A:RL×D→RM\mathcal{A} : \mathbb{R}^{L \times D} \to \mathbb{R}^M and Gaussian noise η∼N(0,Σy)\boldsymbol{\eta} \sim \mathcal{N}(0, \boldsymbol{\Sigma}_y), zero-shot posterior sampling requires the likelihood score ∇x1:L(t)log⁡p(y∣x1:L(t))\nabla_{\mathbf{x}_{1:L}(t)} \log p(\mathbf{y} \mid \mathbf{x}_{1:L}(t)). Approximating the posterior covariance under diffusion with σ(t)2μ(t)2Γ\frac{\sigma(t)^2}{\mu(t)^2} \boldsymbol{\Gamma} (where Γ=QΛ(Λ+I)−1Q−1\boldsymbol{\Gamma} = \mathbf{Q} \boldsymbol{\Lambda} (\boldsymbol{\Lambda} + \mathbf{I})^{-1} \mathbf{Q}^{-1} is derived from the eigendecomposition QΛQ−1\mathbf{Q} \boldsymbol{\Lambda} \mathbf{Q}^{-1} of the prior covariance Σx\boldsymbol{\Sigma}_x), the perturbed likelihood is approximated as:

    p(y∣x1:L(t))≈N(y  |  A(x^(x1:L(t))),  Σy+σ(t)2μ(t)2AΓAT)p(\mathbf{y} \mid \mathbf{x}_{1:L}(t)) \approx \mathcal{N}\left(\mathbf{y} \;\middle|\; \mathcal{A}(\hat{\mathbf{x}}(\mathbf{x}_{1:L}(t))), \; \boldsymbol{\Sigma}_y + \frac{\sigma(t)^2}{\mu(t)^2} \mathbf{A} \boldsymbol{\Gamma} \mathbf{A}^T\right)

    where x^(x1:L(t))=x1:L(t)+σ(t)2sϕ(x1:L(t),t)μ(t)\hat{\mathbf{x}}(\mathbf{x}_{1:L}(t)) = \frac{\mathbf{x}_{1:L}(t) + \sigma(t)^2 s_\phi(\mathbf{x}_{1:L}(t), t)}{\mu(t)} is the Tweedie posterior expectation computed via the score network sϕs_\phi, and A=∂A∂x1:L∣x^(x1:L(t))\mathbf{A} = \left. \frac{\partial \mathcal{A}}{\partial \mathbf{x}_{1:L}} \right|_{\hat{\mathbf{x}}(\mathbf{x}_{1:L}(t))} is the Jacobian of A\mathcal{A}. In practice, AΓAT\mathbf{A} \boldsymbol{\Gamma} \mathbf{A}^T is typically fixed to a constant diagonal matrix (e.g., 10−2I10^{-2} \mathbf{I}), which stabilizes gradient computation in low signal-to-noise regimes without requiring ad-hoc score rescaling.

  4. Knowl 4 — Local Score Network Training and Full Trajectory Score Composition

    algorithm

    Training is performed on trajectory sub-segments of length 2k+12k + 1 using the parameterization ϵϕ(xi−k:i+k(t),t)=−σ(t)sϕ(xi−k:i+k(t),t)\boldsymbol{\epsilon}_\phi(\mathbf{x}_{i-k:i+k}(t), t) = -\sigma(t) s_\phi(\mathbf{x}_{i-k:i+k}(t), t). During inference, the full trajectory score is composed by querying the local score network across overlapping segments.

    procedure TrainLocalScore(dataset p(x_{1:L}), iterations N, radius k)
        for step = 1 to N do
            Sample x_{1:L} ~ p(x_{1:L})
            Sample i ~ Uniform({k + 1, ..., L - k})
            Sample t ~ Uniform(0, 1)
            Sample \epsilon ~ N(0, I)
            x_{i-k:i+k}(t) <- \mu(t) x_{i-k:i+k} + \sigma(t) \epsilon
            loss <- || \epsilon_\phi(x_{i-k:i+k}(t), t) - \epsilon ||_2^2
            \phi <- GradientDescent(\phi, \nabla_\phi loss)
        end for
    end procedure
    procedure ComposeTrajectoryScore(x_{1:L}(t), t, score_net s_\phi, radius k)
        s_{1:k+1} <- s_\phi(x_{1:2k+1}(t), t)[1 : k+1]
        for i = k + 2 to L - k - 1 do
            s_i <- s_\phi(x_{i-k:i+k}(t), t)[k + 1]
        end for
        s_{L-k:L} <- s_\phi(x_{L-2k:L}(t), t)[k + 1 : 2k+1]
        return s_{1:L}
    end procedure
  5. Knowl 5 — Predictor-Corrector Sampling Algorithm for Score-Based Data Assimilation

    algorithm

    Posterior trajectory generation integrates the reverse SDE from t=1t=1 to t=0t=0 using the Exponential Integrator discretization as a predictor, followed by Langevin Monte Carlo (LMC) correction steps between consecutive time discretization points to prevent the accumulation of approximation errors.

    procedure SampleTrajectory(discretization {t_i}_{i=0}^N, corrections C, scale \tau, y, \Sigma_y, \Gamma)
        x(1) ~ N(0, \Sigma(1))
        for i = N down to 1 do
            // Posterior score evaluation
            s_x <- ComposeTrajectoryScore(x(t_i), t_i)
            x_hat <- (x(t_i) + \sigma(t_i)^2 s_x) / \mu(t_i)
            s_y <- \nabla_{x(t_i)} log N(y | A(x_hat), \Sigma_y + (\sigma(t_i)^2 / \mu(t_i)^2) \Gamma)
            s_post <- s_x + s_y
            // Predictor step (Exponential Integrator)
            x(t_{i-1}) <- (\mu(t_{i-1}) / \mu(t_i)) * x(t_i) + (\mu(t_{i-1}) / \mu(t_i) - \sigma(t_{i-1}) / \sigma(t_i)) * \sigma(t_i)^2 * s_post
            // Corrector steps (Langevin Monte Carlo)
            for j = 1 to C do
                \epsilon ~ N(0, I)
                s_x_corr <- ComposeTrajectoryScore(x(t_{i-1}), t_{i-1})
                x_hat_corr <- (x(t_{i-1}) + \sigma(t_{i-1})^2 s_x_corr) / \mu(t_{i-1})
                s_y_corr <- \nabla_{x(t_{i-1})} log N(y | A(x_hat_corr), \Sigma_y + (\sigma(t_{i-1})^2 / \mu(t_{i-1})^2) \Gamma)
                s_corr <- s_x_corr + s_y_corr
                \delta <- \tau * dim(s_corr) / || s_corr ||_2^2
                x(t_{i-1}) <- x(t_{i-1}) + \delta * s_corr + sqrt(2 * \delta) * \epsilon
            end for
        end for
        return x(0)
    end procedure
  6. Knowl 6 — Posterior Accuracy and Convergence on the Chaotic Lorenz 1963 System

    empirical result

    On the stochastic Lorenz 1963 system integrated with time step Δ=0.025\Delta = 0.025 and Brownian noise η∼N(0,ΔI)\boldsymbol{\eta} \sim \mathcal{N}(0, \Delta \mathbf{I}), SDA was evaluated against the ground-truth trajectory posterior computed via a Bootstrap Particle Filter (BPF) with 2162^{16} particles across 64 evaluation observations of length L=65L=65. Under both low-frequency sparse observations (y∼N(a~1:L:8,0.052I)y \sim \mathcal{N}(\tilde{a}_{1:L:8}, 0.05^2 \mathbf{I})) and high-frequency noisy observations (y∼N(a~1:L,0.252I)y \sim \mathcal{N}(\tilde{a}_{1:L}, 0.25^2 \mathbf{I})):

    1. Increasing the pseudo-blanket radius kk and the number of LMC corrections CC monotonically brings the expected log-prior Eq[log⁡p(x2:L∣x1)]\mathbb{E}_{q}[\log p(\mathbf{x}_{2:L} \mid \mathbf{x}_1)], expected log-likelihood Eq[log⁡p(y∣x1:L)]\mathbb{E}_{q}[\log p(\mathbf{y} \mid \mathbf{x}_{1:L})], and 1-Wasserstein distance W1(p,q)W_1(p, q) closer to the BPF ground truth.
    2. The improvements exhibit diminishing returns, with approximate posteriors for k≥3k \ge 3 and C≥2C \ge 2 performing almost identically to higher values.
    3. In multi-modal posterior settings (e.g., observing c1:L:4c_{1:L:4} where the sign of coordinates aa and bb is ambiguous), SDA generates distinct plausible trajectory modes consistent with the observations.
  7. Knowl 7 — Zero-Shot Data Assimilation on 2D Turbulent Kolmogorov Flow

    empirical result

    SDA was evaluated on 2D incompressible Navier-Stokes turbulence subject to Kolmogorov forcing (Re=103Re = 10^3, 64×6464 \times 64 spatial velocity grid, integration step Δ=0.2\Delta = 0.2 time units corresponding to 82 forward Euler solver steps per state snapshot) using a U-Net score network with pseudo-blanket radius k=2k=2:

    1. For trajectories of length L=32L=32 observed intermittently every 4 steps at coarse 8×88 \times 8 resolution with noise Σy=0.12I\boldsymbol{\Sigma}_y = 0.1^2 \mathbf{I}, SDA (C=1,τ=0.5C=1, \tau=0.5) accurately reconstructs the true vorticity field trajectories. In contrast, standard Diffusion Posterior Sampling (DPS) yields unphysical artifacts and trajectories inconsistent with observations.
    2. When observing non-linear saturating transformations x↦x1+∣x∣\mathbf{x} \mapsto \frac{\mathbf{x}}{1 + |\mathbf{x}|} or spatially subsampled measurements down to subsampling factor n=16n=16, SDA maintains high physical consistency and satisfies observation constraints where DPS degrades.
    3. Conditioning on an unlikely synthetic observation (uniform positive vorticity in a circular sub-domain at final state xL\mathbf{x}_L) produces an initial state x1\mathbf{x}_1 that, when evolved purely forward through the physical Navier-Stokes equations, reproduces the target trajectory, proving SDA propagates dynamic information rather than merely interpolating states.
  8. Knowl 8 — Hyperparameters for Lorenz 1963 and Kolmogorov Flow Score Networks

    data/table

    The score networks for both the Lorenz 1963 and 2D Kolmogorov flow experiments were trained using the AdamW optimizer with weight decay 10−310^{-3} and linearly decaying learning rates for 1024 epochs, utilizing SiLU activations and LayerNorm.

    Hyperparameter Lorenz (k≤4k \le 4) Lorenz (k>4k > 4) Kolmogorov Flow
    Architecture Fully-connected Fully-convolutional U-Net
    Residual blocks 5 k−2k - 2 (3, 3, 3) per level
    Channels / Features 256 64 (96, 192, 384)
    Kernel size – 3 3 (circular padding)
    Activation SiLU SiLU SiLU
    Normalization LayerNorm LayerNorm LayerNorm
    Optimizer AdamW AdamW AdamW
    Weight decay 10−310^{-3} 10−310^{-3} 10−310^{-3}
    Learning rate 10−310^{-3} 10−310^{-3} 2×10−42 \times 10^{-4}
    Scheduler Linear Linear Linear
    Epochs 1024 1024 1024
    Batches per epoch 256 256 128
    Batch size 256 64 32

    The table lists the exact architectural specifications and optimization configurations required to train the local score models for low-dimensional ODE dynamics and high-dimensional PDE fluid dynamics.

  9. Knowl 9 — Limitations of Score-Based Data Assimilation

    limitation

    The method is subject to several core limitations:

    1. Inference Latency: Simulating the reverse SDE with multiple Langevin Monte Carlo corrector steps is computationally demanding compared to single-pass state estimators.
    2. Unbounded Approximation Error: The local pseudo-Markov blanket approximation and the perturbed likelihood covariance approximation introduce errors whose impact on the resulting posterior density has not been theoretically quantified.
    3. Dimensional Scaling Gap: While demonstrated on systems with tens of thousands of dimensions (64×64×264 \times 64 \times 2 grids in Kolmogorov flow), operational weather and ocean prediction systems operate on 10710^7--10810^8 state dimensions, presenting unresolved engineering and memory scaling challenges.
    4. Fixed Shared Dynamics Assumption: SDA assumes all trajectories are generated from identical, stationary physical dynamics with known parameters; it cannot infer physical parameters online or guarantee robustness against physical model misspecification.

Coverage note — Standard background derivations of score-based SDEs (Song et al., 2021) and general surveys of deep learning for variational data assimilation were omitted as they constitute contextual prior work.

References

  1. 1.Andrew C Lorenc. ‘Analysis methods for numerical weather prediction’. In Quarterly Journal of the Royal Meteorological Society 112.474 (1986), pp. 1177–1194.
  2. 2.François-Xavier Le Dimet and Olivier Talagrand. ‘Variational algorithms for analysis and assimilation of meteorological observations: theoretical aspects’. In Tellus A: Dynamic Meteorology and Oceanography 38.2 (1986), pp. 97–110.
  3. 3.Geir Evensen. ‘Sequential data assimilation with a nonlinear quasi-geostrophic model using Monte Carlo methods to forecast error statistics’. In Journal of Geophysical Research: Oceans 99.C5 (1994), pp. 10143–10162.
  4. 4.Thomas M. Hamill. ‘Ensemble-based Atmospheric Data Assimilation’. In Predictability of weather and climate 124 (2006), p. 156.
  5. 5.Yannick Trémolet. ‘Accounting for an imperfect model in 4D-Var.’ In ECMWF Technical Memoranda 477 (2005), p. 20. URL: https://www.ecmwf.int/node/12845.
  6. 6.Yannick Trémolet. ‘Model error estimation in 4D-Var’. In ECMWF Technical Memoranda 520 (2007), p. 19. URL: https://www.ecmwf.int/node/12819.
  7. 7.Mike Fisher et al. ‘Weak-Constraint and Long-Window 4D-Var’. In ECMWF Technical Memoranda 655 (2011), p. 47. URL: https://www.ecmwf.int/node/9414.
  8. 8.Alberto Carrassi et al. ‘Data assimilation in the geosciences: An overview of methods, issues, and perspectives’. In Wiley Interdisciplinary Reviews: Climate Change 9.5 (2018).
  9. 9.ECMWF. ‘Part II: Data Assimilation’. In IFS Documentation CY47R1 2 (2020). URL: https://www.ecmwf.int/node/19746.
  10. 10.Peter Bauer et al. ‘The quiet revolution of numerical weather prediction’. In Nature 525.7567 (2015), pp. 47–55. URL: https://www.nature.com/articles/nature14956.
  11. 11.Nils Gustafsson et al. ‘Survey of data assimilation methods for convective-scale numerical weather prediction at operational centres’. In Quarterly Journal of the Royal Meteorological Society 144.713 (2018), pp. 1218–1256.
  12. 12.Julian Mack et al. ‘Attention-based convolutional autoencoders for 3d-variational data assimilation’. In Computer Methods in Applied Mechanics and Engineering 372 (2020). URL: https://arxiv.org/abs/2101.02121.
  13. 13.Rossella Arcucci et al. ‘Deep data assimilation: integrating deep learning with data assimilation’. In Applied Sciences 11.3 (2021), p. 1114. URL: https://www.mdpi.com/2076-3417/11/3/1114.
  14. 14.Thomas Frerix et al. ‘Variational data assimilation with a learned inverse observation operator’. In International Conference on Machine Learning. 2021, pp. 3449–3458. URL: https://arxiv.org/abs/2102.11192.
  15. 15.Ronan Fablet et al. ‘Learning variational data assimilation models and solvers’. In Journal of Advances in Modeling Earth Systems 13.10 (2021). URL: https://arxiv.org/abs/2007.12941.
  16. 16.Yu-Hong Yeung et al. ‘Physics-Informed Machine Learning Method for Large-Scale Data Assimilation Problems’. In Water Resources Research 58.5 (2022). URL: https://arxiv.org/abs/2108.00037.
  17. 17.Julien Brajard et al. ‘Combining data assimilation and machine learning to emulate a dynamical model from sparse and noisy observations: A case study with the Lorenz 96 model’. In Journal of computational science 44 (2020). URL: https://arxiv.org/abs/2001.01520.
  18. 18.Julien Brajard et al. ‘Combining data assimilation and machine learning to infer unresolved scale parametrization’. In Philosophical Transactions of the Royal Society A 379.2194 (2021). URL: https://arxiv.org/abs/2009.04318.
  19. 19.Kai Zhang et al. ‘Multi-source information fused generative adversarial network model and data assimilation based history matching for reservoir with complex geologies’. In Petroleum Science 19.2 (2022), pp. 707–719.
  20. 20.Robin Rombach et al. ‘High-resolution image synthesis with latent diffusion models’. In IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022. URL: https://arxiv.org/abs/2112.10752.
  21. 21.Aditya Ramesh et al. ‘Hierarchical text-conditional image generation with clip latents’. In (2022). URL: https://arxiv.org/abs/2204.06125.
  22. 22.Chitwan Saharia et al. ‘Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding’. In Advances in Neural Information Processing Systems. 2022. URL: https://openreview.net/forum?id=08Yk-n5l2Al.
  23. 23.Jonathan Ho et al. ‘Video Diffusion Models’. In ICLR Workshop on Deep Generative Models for Highly Structured Data. 2022. URL: https : / / openreview . net / forum ? id = BBelR2NdDZ5.
  24. 24.Jonathan Ho et al. ‘Cascaded Diffusion Models for High Fidelity Image Generation’. In Journal of Machine Learning Research 23.47 (2022), pp. 1–33. URL: http://jmlr.org/papers/v23/21-0635.
  25. 25.Jonathan Ho et al. ‘Imagen video: High definition video generation with diffusion models’. 2022. URL: https://arxiv.org/abs/2210.02303.
  26. 26.Zhifeng Kong et al. ‘DiffWave: A Versatile Diffusion Model for Audio Synthesis’. In International Conference on Learning Representations. 2021. URL: https://openreview.net/forum?id=a-xFK8Ymz5J.
  27. 27.Karan Goel et al. ‘It’s raw! audio generation with state-space models’. In International Conference on Machine Learning. PMLR. 2022. URL: https://proceedings.mlr.press/v162/goel22a.
  28. 28.Yang Song et al. ‘Score-Based Generative Modeling through Stochastic Differential Equations’. In International Conference on Learning Representations. 2021. URL: https : / / openreview.net/forum?id=PxTIG12RRHS.
  29. 29.Simo Särkkä and Arno Solin. ‘Applied stochastic differential equations’. Vol. 10. Cambridge University Press, 2019.
  30. 30.Qinsheng Zhang and Yongxin Chen. ‘Fast Sampling of Diffusion Models with Exponential Integrator’. In International Conference on Learning Representations. 2023. URL: https://openreview.net/forum?id=Loek7hfb46P.
  31. 31.Yang Song and Stefano Ermon. ‘Generative Modeling by Estimating Gradients of the Data Distribution’. In Advances in Neural Information Processing Systems. Vol. 32. 2019. URL: https://arxiv.org/abs/1907.05600.
  32. 32.Jonathan Ho et al. ‘Denoising Diffusion Probabilistic Models’. In Advances in Neural Information Processing Systems. Vol. 33. 2020. URL: https://arxiv.org/abs/2006.11239.
  33. 33.Brian Anderson. ‘Reverse-time diffusion equation models’. In Stochastic Processes and their Applications 12.3 (1982), pp. 313–326.
  34. 34.Aapo Hyvärinen. ‘Estimation of Non-Normalized Statistical Models by Score Matching’. In Journal of Machine Learning Research 6.24 (2005), pp. 695–709. URL: http://jmlr.org/papers/v6/hyvarinen05a.
  35. 35.Pascal Vincent. ‘A Connection Between Score Matching and Denoising Autoencoders’. In Neural computation 23.7 (2011), pp. 1661–1674.
  36. 36.Jiaming Song et al. ‘Denoising Diffusion Implicit Models’. In International Conference on Learning Representations. 2021. URL: https : / / openreview . net / forum ? id = St1giarCHLP.
  37. 37.Luping Liu et al. ‘Pseudo Numerical Methods for Diffusion Models on Manifolds’. In International Conference on Learning Representations. 2022. URL: https://openreview.net/forum?id=PlKWVd2yBkY.
  38. 38.Bahjat Kawar et al. ‘SNIPS: Solving Noisy Inverse Problems Stochastically’. In Advances in Neural Information Processing Systems. 2021. URL: https://openreview.net/forum?id=pBKOx_dxYAN.
  39. 39.Yang Song et al. ‘Solving Inverse Problems in Medical Imaging with Score-Based Generative Models’. In International Conference on Learning Representations. 2022. URL: https : //openreview.net/forum?id=vaRCHVj0uGI.
  40. 40.Alexandre Adam et al. ‘Posterior samples of source galaxies in strong gravitational lenses with score-based priors’. 2022. URL: https://arxiv.org/abs/2211.03812.
  41. 41.Hyungjin Chung et al. ‘Diffusion Posterior Sampling for General Noisy Inverse Problems’. In International Conference on Learning Representations. 2023. URL: https://openreview.net/forum?id=OnD9zGAGT0k.
  42. 42.Bradley Efron. ‘Tweedie’s Formula and Selection Bias’. In Journal of the American Statistical Association 106.496 (2011), pp. 1602–1614.
  43. 43.Kwanyoung Kim and Jong Chul Ye. ‘Noise2Score: Tweedie’s Approach to Self-Supervised Image Denoising without Clean Images’. In Advances in Neural Information Processing Systems. 2021. URL: https://openreview.net/forum?id=ZqEUs3sTRU0.
  44. 44.Xiangming Meng and Yoshiyuki Kabashima. ‘Diffusion Model Based Posterior Sampling for Noisy Linear Inverse Problems’. 2022. URL: https://arxiv.org/abs/2211.12343.
  45. 45.Giorgio Parisi. ‘Correlation functions and computer simulations’. In Nuclear Physics B 180.3 (1981), pp. 378–384.
  46. 46.Ulf Grenander and Michael I. Miller. ‘Representations of Knowledge in Complex Systems’. In Journal of the Royal Statistical Society 56.4 (1994), pp. 549–603. URL: http://www.jstor.org/stable/2346184.
  47. 47.Edward N. Lorenz. ‘Deterministic Nonperiodic Flow’. In Journal of Atmospheric Sciences 20.2 (1963), pp. 130–141.
  48. 48.Gary J. Chandler and Rich R. Kerswell. ‘Invariant recurrent solutions embedded in a turbulent two-dimensional Kolmogorov flow’. In Journal of Fluid Mechanics 722 (2013), pp. 554–595.
  49. 49.Jun S. Liu and Rong Chen. ‘Sequential Monte Carlo methods for dynamic systems’. In Journal of the American statistical association 93.443 (1998), pp. 1032–1044.
  50. 50.Arnaud Doucet, Adam M. Johansen, et al. ‘A Tutorial on Particle Filtering and Smoothing’. In Handbook of nonlinear filtering 12 (2011), pp. 656–705. URL: https://wrap.warwick.ac.uk/37961.
  51. 51.Neil J Gordon et al. ‘Novel approach to nonlinear/non-Gaussian Bayesian state estimation’. In Radar and Signal Processing. Vol. 140. 2. IET. 1993, pp. 107–113.
  52. 52.Alexander Quinn Nichol and Prafulla Dhariwal. ‘Improved denoising diffusion probabilistic models’. In International Conference on Machine Learning. 2021. URL: https : / / proceedings.mlr.press/v139/nichol21a.
  53. 53.Dmitrii Kochkov et al. ‘Machine learning-accelerated computational fluid dynamics’. In Proceedings of the National Academy of Sciences 118.21 (2021). URL: https://www.pnas.org/doi/abs/10.1073/pnas.2101784118.
  54. 54.Guido Boffetta and Robert E Ecke. ‘Two-dimensional Turbulence’. In Annual Review of Fluid Mechanics 44 (2012), pp. 427–451.
  55. 55.Olaf Ronneberger et al. ‘U-Net: Convolutional Networks for Biomedical Image Segmentation’. In Medical Image Computing and Computer-Assisted Intervention. 2015, pp. 234–241. URL: https://arxiv.org/abs/1505.04597.
  56. 56.Yang Song et al. ‘Consistency models’. 2023. URL: https: // arxiv.org /abs /2210.02303.
  57. 57.Evan Archer et al. ‘Black box variational inference for state space models’. 2015. URL: https://arxiv.org/abs/1511.07367.
  58. 58.Chris J Maddison et al. ‘Filtering variational objectives’. In Advances in Neural Information Processing Systems 30 (2017). URL: https://arxiv.org/abs/1705.09279.
  59. 59.Christian Naesseth et al. ‘Variational sequential monte carlo’. In International Conference on Artificial Intelligence and Statistics. 2018, pp. 968–977. URL: http://proceedings.mlr.press/v84/naesseth18a.
  60. 60.Dieterich Lawson et al. ‘SIXO: Smoothing Inference with Twisted Objectives’. In Advances in Neural Information Processing Systems. 2022. URL: https://openreview.net/forum?id=bDyLgfvZ0qJ.
  61. 61.Rudolf E. Kalman. ‘A New Approach to Linear Filtering and Prediction Problems’. In Journal of Basic Engineering 82.1 (1960), pp. 35–45.
  62. 62.Eric A. Wan and Rudolph Van Der Merwe. ‘The unscented Kalman filter for nonlinear estimation’. In Proceedings of the IEEE Adaptive Systems for Signal Processing, Communications, and Control Symposium. IEEE. 2000, pp. 153–158.
  63. 63.Ienkaran Arasaratnam and Simon Haykin. ‘Cubature Kalman Filters’. In IEEE Transactions on Automatic Control 54.6 (2009), pp. 1254–1269.
  64. 64.Rahul G. Krishnan et al. ‘Deep Kalman Filters’. 2015. URL: https://arxiv.org/abs/1511.05121.
  65. 65.Laurent Girin et al. ‘Dynamical variational autoencoders: A comprehensive review’. 2020. URL: https://arxiv.org/abs/2008.12595.
  66. 66.Suman Ravuri et al. ‘Skilful precipitation nowcasting using deep generative models of radar’. In Nature 597.7878 (2021), pp. 672–677. URL: https://www.nature.com/articles/s41586-021-03854-z.
  67. 67.Remi Lam et al. ‘GraphCast: Learning skillful medium-range global weather forecasting’. 2022. URL: https://arxiv.org/abs/2212.12794.
  68. 68.Kaiming He et al. ‘Deep Residual Learning for Image Recognition’. In IEEE conference on computer vision and pattern recognition. 2016, pp. 770–778. URL: https://arxiv.org/abs/1512.03385.
  69. 69.Stefan Elfwing et al. ‘Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning’. In Neural Networks 107 (2018), pp. 3–11. URL: https://arxiv.org/abs/1702.03118.
  70. 70.Jimmy Lei Ba et al. ‘Layer normalization’. 2016. URL: https://arxiv.org/abs/1607.06450.
  71. 71.Ilya Loshchilov and Frank Hutter. ‘Decoupled Weight Decay Regularization’. In International Conference on Learning Representations. 2019. URL: https://openreview.net/forum?id=Bkg6RiCqY7.

Citation

MLA
Rozet, F., and G. Louppe. “Score-based Data Assimilation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 40521–41, https://proceedings.neurips.cc/paper_files/paper/2023/file/7f7fa581cc8a1970a4332920cdf87395-Paper-Conference.pdf.
APA
Rozet, F., & Louppe, G. (2023). Score-based Data Assimilation. Advances in Neural Information Processing Systems, 36, 40521–40541. https://proceedings.neurips.cc/paper_files/paper/2023/file/7f7fa581cc8a1970a4332920cdf87395-Paper-Conference.pdf
Chicago
Rozet, F., and G. Louppe. 2023. “Score-based Data Assimilation”. Advances in Neural Information Processing Systems 36: 40521–41. https://proceedings.neurips.cc/paper_files/paper/2023/file/7f7fa581cc8a1970a4332920cdf87395-Paper-Conference.pdf.
Harvard
Rozet, F. and Louppe, G. (2023) “Score-based Data Assimilation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 40521–40541. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/7f7fa581cc8a1970a4332920cdf87395-Paper-Conference.pdf.
Vancouver
1. Rozet F, Louppe G (2023) Score-based Data Assimilation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 40521–40541

BibTeX

@inproceedings{rozet2023score,
  title = {Score-based Data Assimilation},
  author = {Rozet, François and Louppe, Gilles},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {40521-40541},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/7f7fa581cc8a1970a4332920cdf87395-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors