Particle Denoising Diffusion Sampler

Angus PhillipsHai-Dang DauMichael John HutchinsonValentin De BortoliGeorge DeligiannidisArnaud Doucet

article2024ICML56 citations

Introduces the Particle Denoising Diffusion Sampler, an iterative particle method combining guided diffusions with a novel score matching loss to achieve asymptotically consistent sampling and normalizing constant estimation for complex, unnormalized target distributions.

Listen

Generating samples from complex, unnormalized probability distributions and computing their normalizing constants are fundamental challenges across scientific computing, statistics, and machine learning. Traditional methods like annealing often degrade when exploring multimodal landscapes, while standard diffusion-based generative models rely heavily on available data samples and introduce persistent approximation errors when adapted to general sampling. The article addresses these bottlenecks by developing the Particle Denoising Diffusion Sampler (PDDS), a methodology designed to reliably sample from unnormalized target densities and deliver consistent, unbiased normalizing constant estimates.

The approach integrates guided denoising diffusions with Sequential Monte Carlo (SMC), a particle-based filtering framework. Rather than relying on static or heuristic approximations to reverse the noising diffusion process, PDDS formulates the time-reversed process as an iterative particle scheme. To correct for intermediate drift errors, the methodology introduces a Novel Score Matching (NSM) loss function to train neural network potential approximations. This novel loss eliminates the variance blow-up that standard score matching encounters as time steps approach zero, allowing particles to be systematically reweighted, resampled, and optionally perturbed using standard Markov Chain Monte Carlo moves.

The evaluation demonstrates several critical findings across synthetic benchmarks, Bayesian logistic regression, and high-dimensional models scaling up to 1,600 dimensions. First, PDDS, especially when combined with optional Markov Chain Monte Carlo steps, consistently matches or exceeds the performance of state-of-the-art baselines like CRAFT, DDS, and Path Integral Samplers in estimating normalizing constants. Second, on multimodal problems—such as a 20-dimensional mixture of 40 separated Gaussian components—PDDS substantially outperforms flow-based methods by preventing mode collapse and significantly reducing transport error. Third, theoretical analysis confirms that PDDS provides asymptotically consistent estimates and establishes that sorted stratified resampling prevents particle degeneration as discretization steps become infinitely fine.

These results provide practitioners with a scalable, mathematically rigorous framework that reduces computational risk and improves fidelity in high-dimensional probabilistic inference. Unlike flow-based methods that require extensive, problem-specific re-training to adjust time-step resolutions, PDDS uses standard, task-agnostic network architectures that seamlessly refine when time steps are subdivided. Decision-makers should consider adopting PDDS for complex posterior inference and explore pilot implementations where capturing separated modes is critical. Users must note that practical deployment uses a finite particle budget and relies on well-behaved initial variational approximations to maintain numerical stability, warranting caution and validation in unconstrained density settings.

No sufficiently relevant recommendations were found.

Cover for Particle Denoising Diffusion Sampler

Abstract

Denoising diffusion models have become ubiquitous for generative modeling. The core idea is to transport the data distribution to a Gaussian by using a diffusion. Approximate samples from the data distribution are then obtained by estimating the time-reversal of this diffusion using score matching ideas. We follow here a similar strategy to sample from unnormalized probability densities and compute their normalizing constants. However, the time-reversed diffusion is here simulated by using an original iterative particle scheme relying on a novel score matching loss. Contrary to standard denoising diffusion models, the resulting Particle Denoising Diffusion Sampler (PDDS) provides asymptotically consistent estimates under mild assumptions. We demonstrate PDDS on multimodal and high dimensional sampling tasks.

Table of Contents

  • 1. Introduction
  • 2. Denoising Diffusions with Guidance
  • 2.1. Noising and denoising diffusions
  • 2.2. Denoising diffusions with guidance
  • 2.3. Guidance approximation
  • 3. Particle Denoising Diffusion Sampler
  • 3.1. From continuous time to discrete time
  • 3.2. From discrete time to particles
  • 3.3. Algorithm settings
  • 3.4. Theoretical results
  • 4. Learning Potentials via Score Matching
  • 4.1. Different score identities
  • 4.2. Benefits of alternative score matching identity
  • 4.3. Neural network parametrization
  • 4.4. Training the neural network
  • 4.5. Mechanisms behind potential improvement
  • 5. Related Work
  • 6. Experimental Results
  • 6.1. Normalizing constant estimation
  • 6.2. Sample quality
  • 6.3. Iterations of potential approximation
  • 7. Discussion
  • Impact Statement
  • Acknowledgements
  • References
  • Appendix
  • A. Particle Denoising Diffusion Sampler with Adaptive Resampling
  • B. Asymptotic Error Formulae for SMC
  • B.1. Chi-squared divergence
  • B.2. Generic Feynman-Kac formula and the associated SMC algorithm
  • B.3. Sorted stratified resampling schemes
  • B.4. Asymptotic error
  • C. Proofs
  • C.1. Proof of Proposition 2.1
  • C.2. Proof of Lemma 2.2
  • C.4. Proof of Proposition 3.3
  • C.3. Proof of Propositions 3.1 and 3.2
  • C.4.1. REGULARITY CONDITIONS FOR PROPOSITION 3.3
  • C.4.2. FORMAL PROOF
  • C.5. Proof of Proposition 4.1
  • C.6. Proof of Proposition 4.2
  • D. Experimental Details
  • D.1. Benchmarking targets
  • D.2. Algorithmic details and hyperparameter settings
  • D.2.1. SMC AND CRAFT SETTINGS
  • D.2.2. DDS SETTINGS
  • D.2.3. PIS SETTINGS
  • D.2.4. PDDS SETTINGS
  • D.3. Ablation studies
  • D.4. Additional results
  • D.5. Uncurated normalizing constant estimates

Knowls

  1. Knowl 1 — Diffusion guidance rewrites target sampling in terms of a potential

    model/method

    Let the target density on Rd\mathbb{R}^d be π(x)=γ(x)/Z\pi(x)=\gamma(x)/Z, where γ\gamma is evaluable and Z=∫γ(x) dxZ=\int\gamma(x)\,dx is unknown. Use the standard Gaussian reference density p0=N(0,Id)p_0=\mathcal{N}(0,I_d) and define g0(x)=γ(x)/p0(x)g_0(x)=\gamma(x)/p_0(x). The forward Ornstein–Uhlenbeck diffusion is dXt=−βtXt dt+2βt dWtdX_t=-\beta_tX_t\,dt+\sqrt{2\beta_t}\,dW_t, with X0∼πX_0\sim\pi, positive schedule βt\beta_t, and dd-dimensional Brownian motion WtW_t. Writing λt=1−exp⁡(−2∫0tβs ds)\lambda_t=1-\exp(-2\int_0^t\beta_s\,ds), its transition is p(xt∣x0)=N(xt;1−λtx0,λtId)p(x_t\mid x_0)=\mathcal{N}(x_t;\sqrt{1-\lambda_t}x_0,\lambda_t I_d). If πt\pi_t is the time-tt marginal, define gt(xt)=∫g0(x0)p(x0∣xt) dx0g_t(x_t)=\int g_0(x_0)p(x_0\mid x_t)\,dx_0, where p(x0∣xt)p(x_0\mid x_t) is the conditional law under the same diffusion initialized instead at p0p_0. Then πt(xt)=p0(xt)gt(xt)/Z\pi_t(x_t)=p_0(x_t)g_t(x_t)/Z and ∇log⁡πt(xt)=−xt+∇log⁡gt(xt)\nabla\log\pi_t(x_t)=-x_t+\nabla\log g_t(x_t). Consequently, the exact reverse-time diffusion, which starts approximately from standard Gaussian noise at its initial end when ∫0Tβs ds\int_0^T\beta_s\,ds is large, has drift −βT−tYt+2βT−t∇log⁡gT−t(Yt)-\beta_{T-t}Y_t+2\beta_{T-t}\nabla\log g_{T-t}(Y_t) and noise scale 2βT−t\sqrt{2\beta_{T-t}}. The unknown time-dependent potential gtg_t is the obstacle to simulating this reverse process exactly.

  2. Knowl 2 — Plug-in guidance can converge to the wrong Gaussian target

    theoretical result

    For the scalar target π=N(μ,σ2)\pi=\mathcal{N}(\mu,\sigma^2), consider the common plug-in guidance approximation g^t(x)=g0(1−λt x)\hat g_t(x)=g_0(\sqrt{1-\lambda_t}\,x) and simulate the resulting reverse diffusion from Z0(T)∼N(0,1)Z^{(T)}_0\sim\mathcal{N}(0,1). For any positive noise schedule with ∫0Tβs ds→∞\int_0^T\beta_s\,ds\to\infty, the terminal mean and variance converge to μ\mu and 11 if σ=1\sigma=1. If σ≠1\sigma\ne1, they instead converge to E[ZT(T)]→μ1−σ2(1−exp⁡[−(1/σ2−1)])\mathbb{E}[Z^{(T)}_T]\to \frac{\mu}{1-\sigma^2}\bigl(1-\exp[-(1/\sigma^2-1)]\bigr) and Var⁡(ZT(T))→1−exp⁡[−2(1/σ2−1)]2(1/σ2−1)\operatorname{Var}(Z^{(T)}_T)\to\frac{1-\exp[-2(1/\sigma^2-1)]}{2(1/\sigma^2-1)}. Thus, even with no time-discretization error and an arbitrarily long diffusion, the approximation generally does not recover the target mean and variance. For a product of dd such Gaussian targets, the discrepancy accumulates across dimensions: the target-to-output Kullback–Leibler divergence grows linearly in dd, making importance-sampling correction require exponentially many samples for reasonable relative variance.

  3. Knowl 3 — PDDS corrects approximate reverse diffusion with sequential particles

    algorithm

    Particle Denoising Diffusion Sampler (PDDS) targets π(x)=γ(x)/Z\pi(x)=\gamma(x)/Z by applying sequential Monte Carlo (SMC) to a discretized reverse diffusion. Choose KK steps on [0,T][0,T], let δ=T/K\delta=T/K, and define αk=1−exp⁡(−2∫(k−1)δkδβs ds)\alpha_k=1-\exp(-2\int_{(k-1)\delta}^{k\delta}\beta_s\,ds). For each kk, choose an approximate potential g^k\hat g_k with g^0=g0\hat g_0=g_0 and g^K=1\hat g_K=1, and let π^k(x)∝p0(x)g^k(x)\hat\pi_k(x)\propto p_0(x)\hat g_k(x); this makes the initial particle law at k=Kk=K standard Gaussian. Move particles backward from k+1k+1 to kk using the Gaussian proposal p^(xk∣xk+1)=N(xk;1−αk+1xk+1+αk+1∇log⁡g^k+1(xk+1),αk+1Id)\hat p(x_k\mid x_{k+1})=\mathcal{N}(x_k;\sqrt{1-\alpha_{k+1}}x_{k+1}+\alpha_{k+1}\nabla\log\hat g_{k+1}(x_{k+1}),\alpha_{k+1}I_d). Weight each proposed pair by wk=g^k(xk)p(xk∣xk+1)g^k+1(xk+1)p^(xk∣xk+1)w_k=\frac{\hat g_k(x_k)p(x_k\mid x_{k+1})}{\hat g_{k+1}(x_{k+1})\hat p(x_k\mid x_{k+1})}, where p(xk∣xk+1)p(x_k\mid x_{k+1}) is the exact backward transition of the reference diffusion. At each step, multiply the running normalizing-constant estimate by the mean incremental weight, normalize particle weights, and resample; optional Markov chain Monte Carlo moves must leave π^k\hat\pi_k invariant. The output is the empirical distribution of particles at k=0k=0 and the product-of-mean-weights estimate of ZZ. In the experiments, the authors used 2,000 particles, T=1T=1, a cosine noise schedule, and systematic resampling when effective sample size fell below 30% of the particle count; PDDS-MCMC additionally used 10 Metropolis-adjusted Langevin steps.

  4. Knowl 4 — At fixed discretization, PDDS has unbiased normalizing-constant estimates and SMC limit theorems

    theoretical result

    For a fixed number KK of diffusion steps, suppose each incremental weight wkw_k has finite second moment under the proposal pair law and the potential ratios satisfy ∫πk(x)gk(x)/g^k(x) dx<∞\int \pi_k(x)g_k(x)/\hat g_k(x)\,dx<\infty. Then the PDDS estimate Z^0N\hat Z_0^N is unbiased for ZZ and has finite variance for finite particle count NN. With multinomial resampling at every step, N(Z^0N/Z−1)\sqrt{N}(\hat Z_0^N/Z-1) converges in distribution to a centered normal random variable. Its asymptotic variance is the sum of chi-squared discrepancies between the exact and approximate initial laws and the exact and approximate backward pair laws: σK2=χ2(πK∥π^K)+∑k=0K−1χ2 ⁣(πk+1(xk+1)p(xk∣xk+1) ∥ π^k+1(xk+1)p^(xk∣xk+1))\sigma_K^2=\chi^2(\pi_K\|\hat\pi_K)+\sum_{k=0}^{K-1}\chi^2\!\left(\pi_{k+1}(x_{k+1})p(x_k\mid x_{k+1})\,\middle\|\,\hat\pi_{k+1}(x_{k+1})\hat p(x_k\mid x_{k+1})\right). Here πk\pi_k denotes the true noising marginal, π^k\hat\pi_k the normalized approximate intermediate density, and pp and p^\hat p the exact and proposed reverse kernels; χ2(P∥Q)=∫[(dP/dQ)−1]2 dQ\chi^2(P\|Q)=\int[(dP/dQ)-1]^2\,dQ. For every bounded test function ff, the particle estimate N−1∑i=1Nf(X0i)N^{-1}\sum_{i=1}^N f(X_0^i) is also asymptotically normal around ∫f dπ\int f\,d\pi. The variance expression makes accurate potential and reverse-kernel approximations consequential: they reduce the chi-squared discrepancies that drive Monte Carlo error.

  5. Knowl 5 — Sorted stratified resampling yields a finite continuous-time error bound

    theoretical result

    With sorted stratified resampling, PDDS admits a bound on the asymptotic relative error of its normalizing-constant estimate that remains controlled as the discretization is refined. For the continuous-time limit, assume βt≡1\beta_t\equiv1, regularity of the noising marginals (polynomial growth bounds on the first three spatial derivatives of their log densities and a uniform Gaussian-exponential moment bound), and constants C1,C2C_1,C_2 such that C1−1≤gˉt(x)/g~t(x)≤C1C_1^{-1}\le\bar g_t(x)/\tilde g_t(x)\le C_1 and ∥∇log⁡g^t(x)∥≤C2(1+∥x∥)\|\nabla\log\hat g_t(x)\|\le C_2(1+\|x\|). Here gˉt=gt/Z\bar g_t=g_t/Z and g~t=g^t/∫p0(x)g^t(x) dx\tilde g_t=\hat g_t/\int p_0(x)\hat g_t(x)\,dx are the normalized exact and approximate potentials, and π^T\hat\pi_T is the normalized approximate terminal law. If ζK2\zeta_K^2 denotes the sorted-stratified asymptotic error bound for KK steps, then lim sup⁡K→∞ζK2≤χ2(πT∥π^T)+2∫0TEXt∼πt ⁣[gˉt(Xt)g~t(Xt)∥∇log⁡gt(Xt)−∇log⁡g^t(Xt)∥2]dt\limsup_{K\to\infty}\zeta_K^2\le\chi^2(\pi_T\|\hat\pi_T)+2\int_0^T\mathbb{E}_{X_t\sim\pi_t}\!\left[\frac{\bar g_t(X_t)}{\tilde g_t(X_t)}\|\nabla\log g_t(X_t)-\nabla\log\hat g_t(X_t)\|^2\right]dt. The bound identifies score error along the noising path, rather than only mismatch between terminal samples, as the main source of residual asymptotic error. Unlike repeated multinomial resampling, for which the paper notes the asymptotic variance can diverge as KK grows, sorted stratified resampling controls the resampling contribution under these assumptions.

  6. Knowl 6 — Novel score matching uses target-score information to train diffusion potentials

    equation

    Let X0∼πX_0\sim\pi, and conditional on X0X_0, let XkX_k follow the forward diffusion transition p(xk∣X0)p(x_k\mid X_0). Write sθ(k,x)=∇xlog⁡π^θ(k,x)s_\theta(k,x)=\nabla_x\log\hat\pi_\theta(k,x) for a learned score, λk\lambda_k for the diffusion noise variance parameter, and κk=1−λk\kappa_k=\sqrt{1-\lambda_k}. The standard denoising score matching (DSM) identity is ∇log⁡πk(xk)=E[∇xklog⁡p(xk∣X0)∣Xk=xk]\nabla\log\pi_k(x_k)=\mathbb{E}[\nabla_{x_k}\log p(x_k\mid X_0)\mid X_k=x_k]. The paper's novel score matching (NSM) identity is ∇log⁡πk(xk)=κkE[∇log⁡g0(X0)∣Xk=xk]−xk\nabla\log\pi_k(x_k)=\kappa_k\mathbb{E}[\nabla\log g_0(X_0)\mid X_k=x_k]-x_k, where g0=γ/p0g_0=\gamma/p_0 and p0=N(0,Id)p_0=\mathcal{N}(0,I_d). Corresponding regression objectives, with expectations over the joint law of (X0,Xk)(X_0,X_k), are ℓDSM(θ)=∑k=1KE∥sθ(k,Xk)−∇xklog⁡p(Xk∣X0)∥2\ell_{\rm DSM}(\theta)=\sum_{k=1}^K\mathbb{E}\|s_\theta(k,X_k)-\nabla_{x_k}\log p(X_k\mid X_0)\|^2 and ℓNSM(θ)=∑k=1KE∥sθ(k,Xk)+Xk−κk∇log⁡g0(X0)∥2\ell_{\rm NSM}(\theta)=\sum_{k=1}^K\mathbb{E}\|s_\theta(k,X_k)+X_k-\kappa_k\nabla\log g_0(X_0)\|^2. Under the respective finite-moment conditions, expressive score models minimize either objective at the true marginal score. NSM is usable in this sampling setting because the target density and its gradient can be evaluated even when direct target samples are unavailable; PDDS particles provide approximate draws of X0X_0 for training.

  7. Knowl 7 — NSM avoids the near-zero-time variance divergence of DSM in a Gaussian case

    theoretical result

    Consider a one-dimensional target π=N(μ,σ2)\pi=\mathcal{N}(\mu,\sigma^2), constant diffusion rate βt=1\beta_t=1, and a score network sθ(t,x)s_\theta(t,x) that is continuously differentiable in its inputs and parameters and has at most polynomial growth. Assume also that Eπ∣∇θsθ(0,X0)∣2>0\mathbb{E}_\pi|\nabla_\theta s_\theta(0,X_0)|^2>0 for every parameter value. For the DSM and NSM objectives defined by regression respectively to ∇xlog⁡p(Xt∣X0)\nabla_x\log p(X_t\mid X_0) and to κt∇log⁡g0(X0)−Xt\kappa_t\nabla\log g_0(X_0)-X_t, the paper proves that the continuous-time DSM loss and the second moment of its unbiased stochastic-gradient estimator are infinite. In contrast, the NSM loss and the second moment of its unbiased stochastic-gradient estimator are finite. The divergence arises near time zero, where the DSM target contains a conditional-noise score whose variance grows without bound as the diffusion noise vanishes; the NSM target remains well behaved in this setting.

  8. Knowl 8 — A constrained potential network supports iterative PDDS refinement

    model/method

    PDDS learns scalar potentials, not only vector-valued score fields, because its particle weights require potential values. The proposed parameterization uses a scalar network rη(k)r_\eta(k) and a vector network Nγ(k,x)∈RdN_\gamma(k,x)\in\mathbb{R}^d: log⁡g^θ(k,x)=[rη(k)−rη(0)]⟨Nγ(k,x),x⟩+[1−rη(k)+rη(0)]log⁡g0(1−λk x)\log\hat g_\theta(k,x)=[r_\eta(k)-r_\eta(0)]\langle N_\gamma(k,x),x\rangle+[1-r_\eta(k)+r_\eta(0)]\log g_0(\sqrt{1-\lambda_k}\,x), with θ=(η,γ)\theta=(\eta,\gamma). It is anchored to the evaluable potential at time zero. To refine it, first run PDDS with a simple initial approximation; then repeatedly draw X0X_0 from the resulting particle approximation, sample kk uniformly from the discretized times, draw Xk∼p(⋅∣X0)X_k\sim p(\cdot\mid X_0), and take a minibatch gradient step on either the DSM or NSM local loss. Use the updated network as the potential approximation in another PDDS run, and repeat. The SMC correction is essential to this feedback loop: training on the corrected particle output can improve on the distribution produced by the current approximate reverse diffusion, rather than merely fitting its errors. In the paper's three-mode mixture ablation, the simple approximation captured only one mode, while two potential-training iterations recovered the missed modes; removing SMC prevented the iterative process from making comparable progress. For the main experiments, the authors used 20 refinement rounds of 500 training updates each, batch size 300, Adam with initial learning rate 10−310^{-3} and exponential decay by a factor of 0.95 every 50 updates. Their diffusion-model networks used a 128-dimensional sinusoidal time embedding and 64-unit GeLU MLPs.

  9. Knowl 9 — Benchmarks show strong normalizer and sample quality, especially on separated mixtures

    empirical result

    The benchmark comparisons plotted on pages 8–9 and 34–35 evaluate normalizing-constant estimates and sample quality against SMC, CRAFT, Path Integral Sampler (PIS), and Denoising Diffusion Sampler (DDS). The main targets include a six-mode two-dimensional Gaussian mixture, a 10-dimensional funnel, a 61-dimensional Sonar logistic-regression posterior, and a 1,600-dimensional log Gaussian Cox process; the experiments used 2,000 particles and a shared variational Gaussian reference where applicable. PDDS with MCMC had the best normalizing-constant bias and variance on the posterior targets except at four steps on Sonar, where CRAFT performed best. PDDS without MCMC was broadly comparable to CRAFT, and both PDDS variants outperformed PIS and DDS across the reported tasks. On the synthetic mixture, PDDS variants performed best, while CRAFT was strongest on the funnel. In a harder 40-component, highly separated Gaussian mixture tested at dimensions 1, 2, 5, 10, and 20, PDDS-MCMC substantially improved on CRAFT's normalizing-constant bias and variance, particularly at low step counts; it also achieved lower entropy-regularized Wasserstein-2 distance and recovered a greater proportion of the target modes. The plots communicate comparative trends across repeated training and sampling seeds; they do not provide a single exact numerical summary for these comparisons.

  10. Knowl 10 — PDDS depends on a suitable reference potential and finite-particle safeguards

    limitation

    PDDS practically requires the initial ratio g0=γ/p0g_0=\gamma/p_0 to be a well-behaved potential, for example one with boundedness or suitable moments, so that particle weights and potential learning remain stable. The authors found a variational Gaussian approximation useful for reparameterizing the target in their experiments, but note that more sophisticated geometry-handling methods may be needed for harder targets. Although the method has asymptotic guarantees, practical runs use finitely many particles and therefore do not eliminate sampling error; the paper cautions against treating finite-particle output as automatically reliable for consequential decisions. Its uncurated normalizing-constant plots also show numerical overestimation outliers across all compared methods, attributed to numerical error and instability rather than to a problem unique to PDDS.

Coverage note — Detailed proofs and per-task baseline tuning grids are omitted because they support the central methods and findings rather than adding independent contributions.

References

  1. 1.Anastasiou, A., Barp, A., Briol, F.-X., Ebner, B., Gaunt, R. E., Ghaderinezhad, F., Gorham, J., Gretton, A., Ley, C., Liu, Q., Mackey, L., Oates, C. J., Reinert, G., and Swan, Y. Stein’s method meets computational statistics: a review of some recent developments. Statistical Science, 38(1):120–139, 2023.
  2. 2.Arbel, M., Matthews, A., and Doucet, A. Annealed flow transport Monte Carlo. In International Conference on Machine Learning, 2021.
  3. 3.Assaraf, R. and Caffarel, M. Zero-variance principle for Monte Carlo algorithms. Physical Review Letters, 83:4682–4685, 1999.
  4. 4.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. In NeurIPS Workshop on Score-Based Methods, 2022.
  5. 5.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018.
  6. 6.Cardoso, G., Idrissi, Y. J. E., Corff, S. L., and Moulines, E. Monte Carlo guided diffusion for Bayesian linear inverse problems. In International Conference on Learning Representations, 2024.
  7. 7.Cattiaux, P., Conforti, G., Gentil, I., and Léonard, C. Time reversal of diffusion processes under a finite entropy condition. Annales de l’Institut Henri Poincare (B) Probabilites et statistiques, 59(4):1844–1881, 2023.
  8. 8.Chatterjee, S. and Diaconis, P. The sample size required in importance sampling. The Annals of Applied Probability, 28(2):1099–1135, 2018.
  9. 9.Chopin, N. and Papaspiliopoulos, O. An Introduction to Sequential Monte Carlo. Springer Ser. Stat. Springer, 2020.
  10. 10.Chopin, N., Singh, S. S., Soto, T., and Vihola, M. On resampling schemes for particle filters with weakly informative observations. The Annals of Statistics, 50(6):3197–3222, 2022.
  11. 11.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. In International Conference on Learning Representations, 2023.
  12. 12.Corso, G., Xu, Y., De Bortoli, V., Barzilay, R., and Jaakkola, T. Particle guidance: non-iid diverse sampling with diffusion models. In International Conference on Learning Representations, 2023.
  13. 13.Dai, C., Heng, J., Jacob, P. E., and Whiteley, N. An invitation to sequential Monte Carlo samplers. Journal of the American Statistical Association, 117(539):1587–1600, 2022.
  14. 14.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion Schrödinger bridge with applications to score-based generative modeling. In Advances in Neural Information Processing Systems, 2021.
  15. 15.Del Moral, P. Feynman-Kac Formulae: Genealogical and Interacting Particle Approximations. Springer, 2004.
  16. 16.Del Moral, P., Doucet, A., and Jasra, A. Sequential Monte Carlo samplers. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 68(3):411–436, 2006.
  17. 17.Del Moral, P., Doucet, A., and Jasra, A. On adaptive resampling strategies for sequential Monte Carlo methods. Bernoulli, 18(1):252–278, 2012.
  18. 18.Douc, R. and Cappé, O. Comparison of resampling schemes for particle filtering. In Proceedings of the 4th International Symposium on Image and Signal Processing and Analysis, pp. 64–69. IEEE, 2005.
  19. 19.Doucet, A., De Freitas, N., and Gordon, N. J. Sequential Monte Carlo Methods in Practice. Information Science and Statistics. New York, NY: Springer, New York, 2001.
  20. 20.Gerber, M., Chopin, N., and Whiteley, N. Negative association, ordering and convergence of resampling methods. The Annals of Statistics, 47(4):2236–2260, 2019.
  21. 21.Geyer, C. Markov chain Monte Carlo maximum likelihood. In Computing science and statistics: Proceedings of 23rd Symposium on the Interface Interface Foundation, Fairfax Station, 1991, pp. 156–163, 1991.
  22. 22.Guarniero, P., Johansen, A. M., and Lee, A. The iterated auxiliary particle filter. Journal of the American Statistical Association, 112(520):1636–1647, 2017.
  23. 23.Haussmann, U. G. and Pardoux, E. Time reversal of diffusions. The Annals of Probability, 14(3):1188–1205, 1986.
  24. 24.Heng, J., Bishop, A. N., Deligiannidis, G., and Doucet, A. Controlled sequential Monte Carlo. The Annals of Statistics, 48(5):2904–2929, 2020.
  25. 25.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
  26. 26.Hoffman, M., Sountsov, P., Dillon, J. V., Langmore, I., Tran, D., and Vasudevan, S. NeuTra-lizing bad geometry in Hamiltonian Monte Carlo using neural transport. arXiv preprint arXiv:1903.03704, 2019.
  27. 27.Huang, X., Dong, H., Hao, Y., Ma, Y., and Zhang, T. Reverse diffusion Monte Carlo. In International Conference on Learning Representations, 2024.
  28. 28.Hyvärinen, A. Estimation of non-normalized statistical models by score matching. The Journal of Machine Learning Research, 6:695–709, 2005.
  29. 29.Kingma, D. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  30. 30.Kloeden, P. E. and Platen, E. Numerical Solution of Stochastic Differential Equations, volume 23 of Appl. Math. (N. Y.). Berlin: Springer-Verlag, 1992.
  31. 31.Lai, C.-H., Takida, Y., Murata, N., Uesaka, T., Mitsufuji, Y., and Ermon, S. Regularizing score-based models with score Fokker-Planck equations. In NeurIPS Workshop on Score-Based Methods, 2022.
  32. 32.Lawson, D., Raventós, A., and Linderman, S. SIXO: Smoothing inference with twisted objectives. Advances in Neural Information Processing Systems, 2022.
  33. 33.Liptser, R. S. and Shiryayev, A. N. Statistics of Random Processes. I. General theory. Translated by A. B. Aries, volume 5 of Appl. Math. (N. Y.). Springer, New York, 1977.
  34. 34.Maté, B. and Fleuret, F. Learning deformation trajectories of Boltzmann densities. Transactions on Machine Learning Research, 2023.
  35. 35.Matthews, A. G. D. G., Arbel, M., Rezende, D. J., and Doucet, A. Continual repeated annealed flow transport Monte Carlo. In International Conference on Machine Learning, 2022.
  36. 36.McDonald, C. J. and Barron, A. R. Proposal of a score based approach to sampling using Monte Carlo estimation of score and oracle access to target density. In NeurIPS Workshop on Score-Based Methods, 2022.
  37. 37.Midgley, L. I., Stimper, V., Simm, G. N., Scholköpf, B., and Hernández-Lobato, J. M. Flow annealed importance sampling bootstrap. In International Conference on Learning Representations, 2023.
  38. 38.Mira, A., Solgi, R., and Imparato, D. Zero variance Markov chain Monte Carlo for Bayesian estimators. Statistics and Computing, 23(5):653–662, 2013.
  39. 39.Møller, J., Syversveen, A. R., and Waagepetersen, R. P. Log Gaussian Cox processes. Scandinavian Journal of Statistics, 25(3):451–482, 1998.
  40. 40.Neal, R. M. Annealed importance sampling. Statistics and Computing, 11(2):125–139, 2001.
  41. 41.Neal, R. M. Slice sampling. The Annals of Statistics, 31:705–767, 06 2003.
  42. 42.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, 2021.
  43. 43.Richter, L., Berner, J., and Liu, G.-H. Improved sampling via learned diffusions. In International Conference on Learning Representations, 2024.
  44. 44.Salimans, T. and Ho, J. Should EBMs model the energy or the score? In Energy Based Models Workshop-ICLR 2021, 2021.
  45. 45.Song, J., Vahdat, A., Mardani, M., and Kautz, J. Pseudoinverse-guided diffusion models for inverse problems. In International Conference on Learning Representations, 2023.
  46. 46.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 2019.
  47. 47.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  48. 48.Sountsov, P., Radul, A., and contributors. Inference gym, 2020. URL https://pypi.org/project/inference_gym.
  49. 49.Syed, S., Bouchard-Côté, A., Deligiannidis, G., and Doucet, A. Non-reversible parallel tempering: A scalable highly parallel MCMC scheme. Journal of the Royal Statistical Society Series B, 84(2):321–350, 2022.
  50. 50.Tawn, N. G., Roberts, G. O., and Rosenthal, J. S. Weight-preserving simulated tempering. Statistics and Computing, 30(1):27–41, 2020.
  51. 51.Vargas, F., Grathwohl, W., and Doucet, A. Denoising diffusion samplers. In International Conference on Learning Representations, 2023.
  52. 52.Vincent, P. A connection between score matching and denoising autoencoders. Neural Computation, 23(7):1661–1674, 2011.
  53. 53.Webber, R. J. Unifying sequential Monte Carlo with resampling matrices. arXiv preprint arXiv:1903.12583, 2019.
  54. 54.Woodard, D. B., Schmidler, S. C., and Huber, M. Sufficient conditions for torpid mixing of parallel and simulated tempering. Electronic Journal of Probability, 14:780–804, 2009.
  55. 55.Wu, L., Trippe, B. L., Naesseth, C. A., Blei, D., and Cunningham, J. P. Practical and asymptotically exact conditional sampling in diffusion models. In Advances in Neural Information Processing Systems, 2023.
  56. 56.Zhang, D., Chen, R. T., Liu, C.-H., Courville, A., and Bengio, Y. Diffusion generative flow samplers: Improving learning signals through partial trajectory optimization. In International Conference on Learning Representations, 2024.
  57. 57.Zhang, Q. and Chen, Y. Path integral sampler: a stochastic control approach for sampling. In International Conference on Learning Representations, 2022.

Citation

MLA
Phillips, A., et al. “Particle Denoising Diffusion Sampler”. arXiv, 2024, http://arxiv.org/abs/2402.06320v2.
APA
Phillips, A., Dau, H.-D., Hutchinson, M. J., Bortoli, V. D., Deligiannidis, G., & Doucet, A. (2024). Particle Denoising Diffusion Sampler. arXiv. http://arxiv.org/abs/2402.06320v2
Chicago
Phillips, A., H.-D. Dau, M. J. Hutchinson, V. D. Bortoli, G. Deligiannidis, and A. Doucet. 2024. “Particle Denoising Diffusion Sampler”. arXiv. http://arxiv.org/abs/2402.06320v2.
Harvard
Phillips, A. et al. (2024) “Particle Denoising Diffusion Sampler”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.06320v2.
Vancouver
1. Phillips A, Dau H-D, Hutchinson MJ, Bortoli VD, Deligiannidis G, Doucet A (2024) Particle Denoising Diffusion Sampler. arXiv

BibTeX

@article{phillips2024particle,
  title = {Particle Denoising Diffusion Sampler},
  author = {Phillips, Angus and Dau, Hai-Dang and Hutchinson, Michael John and Bortoli, Valentin De and Deligiannidis, George and Doucet, Arnaud},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.06320v2},
  eprint = {2402.06320}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/