Compositional Score Modeling for Simulation-Based Inference

Tomas GeffnerGeorge PapamakariosAndriy Mnih

article2023ICML47 citations

Proposes a sample-efficient simulation-based inference framework using conditional score modeling that naturally aggregates an arbitrary number of observations at test time without relying on MCMC or variational approximations.

Listen

Mechanistic simulations are widely used across science and engineering to model complex phenomena where underlying likelihood equations cannot be directly calculated. When estimating system parameters from experimental data, simulation-based inference methods face an operational dilemma. Direct posterior estimation models require multiple, computationally expensive simulator runs per training example to handle multi-observation datasets. Conversely, likelihood surrogate models require only one simulation per training pair and can process arbitrary numbers of observations at test time, but they rely on traditional sampling algorithms that struggle with complex, multimodal distributions and introduce compounding errors.

The article introduces and evaluates Factorized Neural Posterior Score Estimation and its generalized family, Partially Factorized Neural Posterior Score Estimation. The primary objective is to demonstrate a generative score-based modeling framework that accurately infers parameters from multiple observations while maintaining high simulation efficiency and bypassing the failure modes of standard inference samplers.

The authors designed a conditional diffusion framework that models the score functions of individual observation posteriors and combines them additively to generate target posterior samples via annealed Langevin dynamics. They benchmarked these methods across standard scientific models—including Gaussian distributions, Susceptible-Infected-Recovered epidemic dynamics, Lotka–Volterra predator-prey systems, and a high-energy particle physics collision simulator. Performance was tested across simulator budgets ranging from 1,000 to 30,000 calls and evaluated using Maximum Mean Discrepancy and classifier two-sample tests against established neural posterior and neural ratio estimation baselines.

The evaluation revealed four key findings. First, factorized score methods consistently matched or outperformed existing baselines when conditioning on multiple observations across various simulation budgets. Second, the partially factorized framework—grouping small subsets of observations, typically between three and six—achieved the most robust balance between training sample efficiency and reduced error accumulation. Third, score-based composition reliably captured multiple distribution modes in multimodal benchmarks where standard Markov chain Monte Carlo samplers failed to mix. Finally, the framework proved robust across various score network parameterizations and sampling hyperparameters, provided that at least five to ten sampling steps were used per noise level.

These results demonstrate that score modeling provides a practical path to scaling simulation-based inference in data-rich settings without increasing simulation compute costs or risking the mode-collapse failures typical of traditional inference. Organizations applying simulation models to complex, safety-critical, or expensive tasks can achieve higher parameter inference accuracy while avoiding costly simulator reruns.

Teams implementing simulation-based workflows should adopt partially factorized score modeling with small observation partition sizes (such as three to six) when inferring parameters from multiple observations. While the core methodology demonstrates strong empirical confidence across established benchmarks, practitioners should note that the default annealed Langevin sampler adds computational overhead proportional to the step count. Further investigation into faster, hyperparameter-free samplers and sequential active-learning approaches is recommended before deploying at very large production scales.

arXiv: 2209.14249
Cover for Compositional Score Modeling for Simulation-Based Inference

Abstract

Neural Posterior Estimation methods for simulation-based inference can be ill-suited for dealing with posterior distributions obtained by conditioning on multiple observations, as they tend to require a large number of simulator calls to learn accurate approximations. In contrast, Neural Likelihood Estimation methods can handle multiple observations at inference time after learning from individual observations, but they rely on standard inference methods, such as MCMC or variational inference, which come with certain performance drawbacks. We introduce a new method based on conditional score modeling that enjoys the benefits of both approaches. We model the scores of the (diffused) posterior distributions induced by individual observations, and introduce a way of combining the learned scores to approximately sample from the target posterior distribution. Our approach is sample-efficient, can naturally aggregate multiple observations at inference time, and avoids the drawbacks of standard inference methods.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Simulation-based Inference
  • 2.2 Conditional Score-based Generative Modeling
  • 3 Score Modeling for SBI
  • 3.1 Factorized Neural Posterior Score Estimation
  • 3.2 Direct Application of Conditional Score Modeling
  • 3.3 Partially Factorized Neural Posterior Score Estimation
  • 4 Related Work
  • 5 Empirical Evaluation
  • 5.1 Illustrative Multimodal Example
  • 5.2 Systematic Evaluation
  • 5.3 Optimal Trade-off for PF-NPSE
  • 5.4 Score Network and Langevin Sampler
  • 5.5 Demonstration: Weinberg Simulator
  • 6 Conclusion and Limitations
  • References
  • A Models Used
  • B Details for each Method
  • B.1 F-NPSE
  • B.2 PF-NPSE
  • B.3 NPSE
  • B.4 NPE
  • B.5 NRE
  • C Derivation of Posterior Factorization
  • D Alternative Sampling Method Without Unadjusted Langevin Dynamics
  • E Additional Results
  • E.1 Classifier Two-sample Test
  • E.2 Performance of PF-NPSE for Different mm
  • E.3 Conservative and Non-conservative Parameterization for the Score Network
  • E.4 Langevin Sampler Parameters

Knowls

  1. Knowl 1 — Posterior factorization over individual observations

    theoretical result

    For a prior p(θ)p(\theta) and conditionally independent observations x1,…,xnx_1,\ldots,x_n given parameter θ\theta, the joint posterior can be written as p(θ∣x1,…,xn)∝p(θ)1−n∏j=1np(θ∣xj)p(\theta\mid x_1,\ldots,x_n)\propto p(\theta)^{1-n}\prod_{j=1}^{n}p(\theta\mid x_j). Here p(θ∣xj)p(\theta\mid x_j) is the posterior based on the prior and the single observation xjx_j. This identity makes it possible to express multi-observation inference using individual-observation posteriors, with the prior correction preventing its repeated counting.

  2. Knowl 2 — F-NPSE constructs compositional diffusion bridges

    model/method

    Factorized Neural Posterior Score Estimation (F-NPSE) builds a sequence of distributions between a tractable reference and the posterior for observations x1,…,xnx_1,\ldots,x_n. Let p(θ)p(\theta) be the prior, let pt(θ∣xj)p_t(\theta\mid x_j) be the posterior for one observation smoothed by Gaussian noise at level tt, and let TT be the final noise level. F-NPSE defines ptf(θ∣x1,…,xn)∝[p(θ)1−n](T−t)/T∏j=1npt(θ∣xj)p_t^{f}(\theta\mid x_1,\ldots,x_n)\propto [p(\theta)^{1-n}]^{(T-t)/T}\prod_{j=1}^{n}p_t(\theta\mid x_j) for t=0,…,Tt=0,\ldots,T. At t=0t=0 this recovers the target posterior; when the final noise level has γT≈0\gamma_T\approx0, the endpoint is approximately N(0,I/n)\mathcal N(0,I/n), where II is the identity matrix. The bridge score is the prior score weighted by (1−n)(T−t)/T(1-n)(T-t)/T plus the sum of the individual diffused-posterior scores. Summing learned individual scores enables inference with any number of observations without training on multi-observation simulator cases. The sum can accumulate score-estimation errors; constrained priors require reparameterization or a separately diffused prior.

  3. Knowl 3 — Training the shared individual-observation score network

    model/method

    F-NPSE trains one conditional score network sψ(θ,t,x)s_\psi(\theta,t,x) using parameter-observation pairs (θ′,x)∼p(θ)p(x∣θ)(\theta',x)\sim p(\theta)p(x\mid\theta), so each training pair requires one simulator call. For a noise level tt, construct θ~=γtθ′+1−γtϵ\tilde\theta=\sqrt{\gamma_t}\theta'+\sqrt{1-\gamma_t}\epsilon, where ϵ∼N(0,I)\epsilon\sim\mathcal N(0,I) and γt\gamma_t is the Gaussian diffusion schedule. The denoising score-matching target is the score of this conditional Gaussian, −ϵ/1−γt-\epsilon/\sqrt{1-\gamma_t}; the network is trained to predict it from (θ~,t,x)(\tilde\theta,t,x), with the paper’s nonnegative time weight λ(t)\lambda(t). At inference, the estimate for observations x1,…,xnx_1,\ldots,x_n is (1−n)(T−t)T∇θlog⁡p(θ)+∑j=1nsψ(θ,t,xj)\frac{(1-n)(T-t)}{T}\nabla_\theta\log p(\theta)+\sum_{j=1}^{n}s_\psi(\theta,t,x_j), where ∇θlog⁡p(θ)\nabla_\theta\log p(\theta) is the prior score. The experiments use T=400T=400 noise levels.

  4. Knowl 4 — PF-NPSE interpolates between individual and joint conditioning

    model/method

    Partially Factorized Neural Posterior Score Estimation (PF-NPSE) groups observations into disjoint subsets of size at most mm. For nn observations this gives k=⌈n/m⌉k=\lceil n/m\rceil groups X1,…,XkX_1,\ldots,X_k, and the posterior factorizes as p(θ∣x1,…,xn)∝p(θ)1−k∏j=1kp(θ∣Xj)p(\theta\mid x_1,\ldots,x_n)\propto p(\theta)^{1-k}\prod_{j=1}^{k}p(\theta\mid X_j). PF-NPSE forms its diffusion bridge by replacing each group posterior with its Gaussian-smoothed version pt(θ∣Xj)p_t(\theta\mid X_j) and interpolating the prior exponent from 1−k1-k at t=0t=0 to zero at t=Tt=T. It trains a permutation-invariant conditional score network on sets containing between one and mm observations, and at inference sums the group-score estimates with the corresponding prior correction. Training uses an average of m/2m/2 simulator calls per case under the paper’s uniform training distribution over group sizes. Setting m=1m=1 recovers F-NPSE; setting mm to the maximum observation count recovers direct joint conditioning. Increasing mm reduces the number of summed score estimates, trading some simulator efficiency for less error accumulation.

  5. Knowl 5 — Annealed Langevin sampling for the composed posterior score

    algorithm

    F-NPSE and PF-NPSE use annealed Langevin dynamics to turn their composed bridge scores into approximate posterior samples. For F-NPSE with nn observations, initialize θ∼N(0,I/n)\theta\sim\mathcal N(0,I/n); at each level t=T−1,T−2,…,1t=T-1,T-2,\ldots,1, apply LL updates using the composed score scomp(θ,t)=(1−n)(T−t)T∇θlog⁡p(θ)+∑j=1nsψ(θ,t,xj)s_{\mathrm{comp}}(\theta,t)=\frac{(1-n)(T-t)}{T}\nabla_\theta\log p(\theta)+\sum_{j=1}^{n}s_\psi(\theta,t,x_j). Each update uses an independent η∼N(0,I)\eta\sim\mathcal N(0,I) and step size δt\delta_t: θ←θ+δt2scomp(θ,t)+δtη\theta\leftarrow\theta+\frac{\delta_t}{2}s_{\mathrm{comp}}(\theta,t)+\sqrt{\delta_t}\eta. For PF-NPSE, replace the individual-observation sum and prior coefficient with the corresponding group-score sum and group-count correction from its bridge. Repeating the procedure yields approximate posterior samples; the cost is O(LT)O(LT) score updates per sample. The experiments use L=5L=5 steps per noise level and a schedule with δt=0.3(1−αt)/αt\delta_t=0.3(1-\alpha_t)/\alpha_t, where α1=γ1\alpha_1=\gamma_1 and αt=γt/γt−1\alpha_t=\gamma_t/\gamma_{t-1} for subsequent levels.

  6. Knowl 6 — Direct posterior score estimation is less simulation-efficient for variable set sizes

    limitation

    Neural Posterior Score Estimation (NPSE) applies conditional diffusion directly to p(θ∣x1,…,xn)p(\theta\mid x_1,\ldots,x_n), training a score network conditioned on the entire observation set. Unlike F-NPSE, NPSE does not sum separate learned score estimates and therefore avoids that particular source of accumulated approximation error. But for variable set sizes, it must be configured for a maximum size nmax⁡n_{\max} and trained on multi-observation cases; with the paper’s uniform sampling of sizes from 11 to nmax⁡n_{\max}, this requires an average of nmax⁡/2n_{\max}/2 simulator calls per training case. Because its diffused posterior scores do not factorize over individual observations, it cannot obtain F-NPSE’s single-observation training efficiency by summing individual scores.

  7. Knowl 7 — Evaluation design across SBI benchmarks

    experimental setup

    The systematic comparison evaluates F-NPSE, PF-NPSE, NPSE, neural posterior estimation (NPE), and neural ratio estimation (NRE) on four tasks: a 10-dimensional Gaussian-prior/Gaussian-likelihood model (G-G), a 10-dimensional Gaussian-prior/two-component Gaussian-mixture likelihood model (G-MoG), a two-parameter susceptible-infected-recovered model, and a four-parameter Lotka–Volterra model. Training simulator-call budgets are 10310^3, 3⋅1033\cdot10^3, 10410^4, and 3⋅1043\cdot10^4; evaluation conditions on sets of 1,8,14,22,1,8,14,22, or 3030 observations. NPE and NPSE use nmax⁡=30n_{\max}=30, and the main PF-NPSE comparison uses m=6m=6. Each configuration is trained with five random seeds and evaluated on six parameter settings per run. The primary metric is squared maximum mean discrepancy using a Gaussian kernel whose scale is set by the median-distance heuristic; classifier two-sample test accuracy is also reported. Score-based methods use 400 noise levels.

  8. Knowl 8 — Compositional score methods perform well across benchmarks and multimodal demonstrations

    empirical result

    Across the four systematic benchmarks, performance generally improves as simulator-call budgets increase. F-NPSE or PF-NPSE is typically the best-performing method, and PF-NPSE often outperforms the baselines; with few conditioning observations, F-NPSE is often best. The plotted benchmark comparisons on page 7 also show that F-NPSE and PF-NPSE tend to scale better to larger observation sets than NRE. In a two-dimensional multimodal example with prior N(0,I)\mathcal N(0,I) and likelihood p(x∣θ)=12N(x∣θ,I/2)+12N(x∣−θ,I/2)p(x\mid\theta)=\tfrac12\mathcal N(x\mid\theta,I/2)+\tfrac12\mathcal N(x\mid-\theta,I/2), F-NPSE captures both posterior modes when conditioning on different subset sizes despite training on single observations; NRE with Hamiltonian Monte Carlo fails to mix between the modes. In the Weinberg-simulator demonstration, the F-NPSE posterior based on five observations is more concentrated around the generating parameter than the posteriors based on individual observations.

  9. Knowl 9 — Intermediate PF-NPSE group sizes often improve accuracy

    empirical result

    The group size mm controls PF-NPSE’s balance between training simulation cost and score-combination error: for a fixed inference set of size nn, the method sums ⌈n/m⌉\lceil n/m\rceil group scores, while the training cases use an average of m/2m/2 simulator calls. Across the tested values m∈{1,3,6,12,18,30}m\in\{1,3,6,12,18,30\}, the paper finds that intermediate small values, usually m=3m=3 or m=6m=6, often yield the lowest squared MMD; the extremes corresponding to F-NPSE and NPSE are frequently suboptimal. Additional tests find no substantial performance change from using a conservative rather than unconstrained score parameterization, and find Langevin sampling robust when enough updates are taken per noise level, typically 5–10.

  10. Knowl 10 — A linear-cost Gaussian-transition sampler is an alternative to Langevin dynamics

    algorithm

    The paper also gives a sampler for F-NPSE and PF-NPSE that composes Gaussian reverse transitions instead of running multiple Langevin steps per noise level. For F-NPSE with nn observations, define α1=γ1\alpha_1=\gamma_1 and αt=γt/γt−1\alpha_t=\gamma_t/\gamma_{t-1} for t=2,…,T−1t=2,\ldots,T-1, initialize θ∼N(0,I/n)\theta\sim\mathcal N(0,I/n), and for t=T−1t=T-1 down to 11 compute σt2=(1−αt)/(n−αt(n−1))\sigma_t^2=(1-\alpha_t)/(n-\alpha_t(n-1)) and μt=1n−αt(n−1)[∑j=1n(θ/αt+(1−αt)sψ(θ,t,xj)/αt)−(n−1)αtθ]+σt2(1−n)(T−t)T∇θlog⁡p(θ)\mu_t=\frac{1}{n-\alpha_t(n-1)}\left[\sum_{j=1}^{n}\left(\theta/\sqrt{\alpha_t}+(1-\alpha_t)s_\psi(\theta,t,x_j)/\sqrt{\alpha_t}\right)-(n-1)\sqrt{\alpha_t}\theta\right]+\sigma_t^2\frac{(1-n)(T-t)}{T}\nabla_\theta\log p(\theta). Draw the next state from N(μt,σt2I)\mathcal N(\mu_t,\sigma_t^2 I) and return the final θ\theta. This method takes O(T)O(T) transitions and has no step-size or per-level iteration hyperparameters; for one observation it reduces to the standard diffusion-model sampler. Its derivation relies on approximate reverse transitions and on normalizing the composed Gaussian transition, so its validity depends on those approximations. Empirically it is promising but can perform slightly worse than annealed Langevin sampling.

Coverage note — No substantial contribution was omitted; the knowls include the main methods, empirical findings, and alternative sampler, while leaving out related work and implementation minutiae that are not needed to reconstruct the contributions.

References

  1. 1.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Batzolis, G., Stanczuk, J., Schönlieb, C.-B., and Etmann, C. Conditional image generation with score-based diffusion models. arXiv preprint arXiv:2111.13606, 2021.
  3. 3.Beaumont, M. A. Approximate Bayesian computation. Annual review of statistics and its application, 6:379–403, 2019.
  4. 4.Blum, M. G. B. and François, O. Non-linear regression models for approximate Bayesian computation. Statistics and Computing, 20(1):63–73, 2010.
  5. 5.Chan, J., Perrone, V., Spence, J., Jenkins, P., Mathieson, S., and Song, Y. A likelihood-free inference framework for population genetic data using exchangeable neural networks. Advances in Neural Information Processing Systems, 31, 2018.
  6. 6.Choi, K., Meng, C., Song, Y., and Ermon, S. Density ratio estimation via infinitesimal classification. In International Conference on Artificial Intelligence and Statistics, pp. 2552–2573. PMLR, 2022.
  7. 7.Cranmer, K., Pavez, J., and Louppe, G. Approximating likelihood ratios with calibrated discriminative classifiers. arXiv preprint arXiv:1506.02169, 2015.
  8. 8.Cranmer, K., Heinrich, L., Head, T., and Louppe, G. Active sciencing with reusable workflows. https://https://github.com/cranmer/active_sciencing, 2017.
  9. 9.Cranmer, K., Brehmer, J., and Louppe, G. The frontier of simulation-based inference. Proceedings of the National Academy of Sciences, 117(48):30055–30062, 2020.
  10. 10.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion Schrödinger bridge with applications to score-based generative modeling. Advances in Neural Information Processing Systems, 34:17695–17709, 2021.
  11. 11.Del Moral, P., Doucet, A., and Jasra, A. An adaptive sequential Monte Carlo method for approximate Bayesian computation. Statistics and computing, 22(5):1009–1020, 2012.
  12. 12.Dhariwal, P. and Nichol, A. Diffusion models beat GANs on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  13. 13.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using Real NVP. arXiv preprint arXiv:1605.08803, 2016.
  14. 14.Du, Y., Durkan, C., Strudel, R., Tenenbaum, J. B., Dieleman, S., Fergus, R., Sohl-Dickstein, J., Doucet, A., and Grathwohl, W. Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and mcmc. arXiv preprint arXiv:2302.11552, 2023.
  15. 15.Durkan, C., Papamakarios, G., and Murray, I. Sequential neural methods for likelihood-free inference. Bayesian Deep Learning Workshop at Neural Information Processing Systems, 2018.
  16. 16.Friedman, J. H. On multivariate goodness–of–fit and two–sample testing. Statistical Problems in Particle Physics, Astrophysics, and Cosmology, 1:311, 2003.
  17. 17.Glockler, M., Deistler, M., and Macke, J. H. Variational methods for simulation-based inference. arXiv preprint arXiv:2203.04176, 2022.
  18. 18.Greenberg, D., Nonnenmacher, M., and Macke, J. Automatic posterior transformation for likelihood-free inference. In International Conference on Machine Learning, pp. 2404–2414. PMLR, 2019.
  19. 19.Gretton, A., Borgwardt, K. M., Rasch, M. J., Scholkopf, B., and Smola, A. A kernel two-sample test. The Journal of Machine Learning Research, 13(1):723–773, 2012.
  20. 20.Harko, T., Lobo, F. S., and Mak, M. Exact analytical solutions of the Susceptible-Infected-Recovered (SIR) epidemic model and of the SIR model with equal death and birth rates. Applied Mathematics and Computation, 236:184–194, 2014.
  21. 21.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  22. 22.Hermans, J., Begy, V., and Louppe, G. Likelihood-free MCMC with amortized approximate ratio estimators. In International Conference on Machine Learning, pp. 4239–4248. PMLR, 2020a.
  23. 23.Hermans, J., Begy, V., and Louppe, G. Likelihood-free MCMC with amortized approximate ratio estimators. In Proceedings of the 37th International Conference on Machine Learning, 2020b.
  24. 24.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  25. 25.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  26. 26.Hoffman, M. D., Gelman, A., et al. The No-U-Turn Sampler: adaptively setting path lengths in Hamiltonian Monte Carlo. J. Mach. Learn. Res., 15(1):1593–1623, 2014.
  27. 27.Hyvarinen, A. and Dayan, P. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 6(4), 2005.
  28. 28.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  29. 29.Li, H., Yang, Y., Chang, M., Chen, S., Feng, H., Xu, Z., Li, Q., and Chen, Y. Srdiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing, 479:47–59, 2022.
  30. 30.Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B. Compositional visual generation with composable diffusion models. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVII, pp. 423–439. Springer, 2022.
  31. 31.Louppe, G., Hermans, J., and Cranmer, K. Adversarial variational optimization of non-differentiable simulators. arXiv preprint arXiv:1707.07113, 2017.
  32. 32.Lueckmann, J.-M., Goncalves, P. J., Bassetto, G., Öcal, K., Nonnenmacher, M., and Macke, J. H. Flexible statistical inference for mechanistic models of neural dynamics. Advances in Neural Information Processing Systems, 30, 2017.
  33. 33.Lueckmann, J.-M., Bassetto, G., Karaletsos, T., and Macke, J. H. Likelihood-free inference with emulator networks. In Symposium on Advances in Approximate Bayesian Inference, pp. 32–53. PMLR, 2019.
  34. 34.Lueckmann, J.-M., Boelts, J., Greenberg, D., Goncalves, P., and Macke, J. Benchmarking simulation-based inference. In International Conference on Artificial Intelligence and Statistics, pp. 343–351. PMLR, 2021.
  35. 35.Luo, C. Understanding diffusion models: A unified perspective. arXiv preprint arXiv:2208.11970, 2022.
  36. 36.Marjoram, P., Molitor, J., Plagnol, V., and Tavare, S. Markov chain Monte Carlo without likelihoods. Proceedings of the National Academy of Sciences, 100(26):15324–15328, 2003.
  37. 37.Neal, R. M. MCMC using Hamiltonian dynamics. Handbook of Markov chain Monte Carlo, 2(11):2, 2011.
  38. 38.Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., and Chen, M. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. In Proceedings of the 39th International Conference on Machine Learning, 2022.
  39. 39.Papamakarios, G. and Murray, I. Fast ε-free inference of simulation models with Bayesian conditional density estimation. Advances in Neural Information Processing Systems, 29, 2016.
  40. 40.Papamakarios, G., Sterratt, D., and Murray, I. Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 837–848. PMLR, 2019.
  41. 41.Pham, K. C., Nott, D. J., and Chaudhuri, S. A note on approximating ABC-MCMC using flexible classifiers. Stat, 3(1):218–227, 2014.
  42. 42.Phan, D., Pradhan, N., and Jankowiak, M. Composable effects for flexible and accelerated probabilistic programming in NumPyro. arXiv preprint arXiv:1912.11554, 2019.
  43. 43.Radev, S. T., Mertens, U. K., Voss, A., Ardizzone, L., and Kothe, U. BayesFlow: Learning complex stochastic models with invertible neural networks. IEEE transactions on neural networks and learning systems, 2020.
  44. 44.Ramdas, A., Reddi, S. J., Poczos, B., Singh, A., and Wasserman, L. On the decreasing power of kernel and distance based nonparametric hypothesis tests in high dimensions. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 29, 2015.
  45. 45.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022.
  46. 46.Rezende, D. and Mohamed, S. Variational inference with normalizing flows. In International Conference on Machine Learning, pp. 1530–1538. PMLR, 2015.
  47. 47.Rhodes, B., Xu, K., and Gutmann, M. U. Telescoping density-ratio estimation. Advances in neural information processing systems, 33:4905–4916, 2020.
  48. 48.Roberts, G. O. and Tweedie, R. L. Exponential convergence of Langevin distributions and their discrete approximations. Bernoulli, 2(4):341–363, 1996.
  49. 49.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022a.
  50. 50.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022b.
  51. 51.Salimans, T. and Ho, J. Should EBMs model the energy or the score? 2021.
  52. 52.Sharrock, L., Simons, J., Liu, S., and Beaumont, M. Sequential neural score estimation: Likelihood-free inference with conditional score based diffusion models. arXiv preprint arXiv:2210.04872, 2022.
  53. 53.Shi, Y., De Bortoli, V., Deligiannidis, G., and Doucet, A. Conditional simulation using diffusion Schrödinger bridges. arXiv preprint arXiv:2202.13460, 2022.
  54. 54.Sisson, S. A., Fan, Y., and Tanaka, M. M. Sequential Monte Carlo without likelihoods. Proceedings of the National Academy of Sciences, 104(6):1760–1765, 2007.
  55. 55.Sisson, S. A., Fan, Y., and Beaumont, M. Handbook of approximate Bayesian computation. CRC Press, 2018.
  56. 56.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  57. 57.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  58. 58.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  59. 59.Song, Y., Durkan, C., Murray, I., and Ermon, S. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  60. 60.Tabak, E. G. and Turner, C. V. A family of nonparametric density estimation algorithms. Communications on Pure and Applied Mathematics, 66(2):145–164, 2013.
  61. 61.Tashiro, Y., Song, J., Song, Y., and Ermon, S. CSDI: Conditional score-based diffusion models for probabilistic time series imputation. Advances in Neural Information Processing Systems, 34:24804–24816, 2021.
  62. 62.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017.
  63. 63.Vincent, P. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, 2011.
  64. 64.Welling, M. and Teh, Y. W. Bayesian learning via stochastic gradient Langevin dynamics. In Proceedings of the 28th International Conference on Machine Learning, pp. 681–688. Citeseer, 2011.
  65. 65.Winkler, C., Worrall, D., Hoogeboom, E., and Welling, M. Learning likelihoods with conditional normalizing flows. arXiv preprint arXiv:1912.00042, 2019.
  66. 66.Wiqvist, S., Frellsen, J., and Picchini, U. Sequential neural posterior and likelihood approximation. arXiv preprint arXiv:2102.06522, 2021.
  67. 67.Wood, S. N. Statistical inference for noisy nonlinear ecological dynamic systems. Nature, 466(7310):1102–1104, 2010.
  68. 68.Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J. Deep sets. Advances in neural information processing systems, 30, 2017.

Citation

MLA
Geffner, T., et al. “Compositional Score Modeling for Simulation-Based Inference”. International Conference on Machine Learning, vol. 202, 2023, pp. 11098–116, https://proceedings.mlr.press/v202/geffner23a.html.
APA
Geffner, T., Papamakarios, G., & Mnih, A. (2023). Compositional Score Modeling for Simulation-Based Inference. International Conference on Machine Learning, 202, 11098–11116. https://proceedings.mlr.press/v202/geffner23a.html
Chicago
Geffner, T., G. Papamakarios, and A. Mnih. 2023. “Compositional Score Modeling for Simulation-Based Inference”. International Conference on Machine Learning 202: 11098–116. https://proceedings.mlr.press/v202/geffner23a.html.
Harvard
Geffner, T., Papamakarios, G. and Mnih, A. (2023) “Compositional Score Modeling for Simulation-Based Inference”, International Conference on Machine Learning. PMLR, pp. 11098–11116. Available at: https://proceedings.mlr.press/v202/geffner23a.html.
Vancouver
1. Geffner T, Papamakarios G, Mnih A (2023) Compositional Score Modeling for Simulation-Based Inference. In: International Conference on Machine Learning. PMLR, pp 11098–11116

BibTeX

@InProceedings{pmlr-v202-geffner23a,
  title = 	 {Compositional Score Modeling for Simulation-Based Inference},
  author =       {Geffner, Tomas and Papamakarios, George and Mnih, Andriy},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {11098--11116},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/geffner23a/geffner23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/geffner23a.html},
  abstract = 	 {Neural Posterior Estimation methods for simulation-based inference can be ill-suited for dealing with posterior distributions obtained by conditioning on multiple observations, as they tend to require a large number of simulator calls to learn accurate approximations. In contrast, Neural Likelihood Estimation methods can handle multiple observations at inference time after learning from individual observations, but they rely on standard inference methods, such as MCMC or variational inference, which come with certain performance drawbacks. We introduce a new method based on conditional score modeling that enjoys the benefits of both approaches. We model the scores of the (diffused) posterior distributions induced by individual observations, and introduce a way of combining the learned scores to approximately sample from the target posterior distribution. Our approach is sample-efficient, can naturally aggregate multiple observations at inference time, and avoids the drawbacks of standard inference methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/