Compositional Score Modeling for Simulation-Based Inference
Tomas GeffnerGeorge PapamakariosAndriy Mnih
Proposes a sample-efficient simulation-based inference framework using conditional score modeling that naturally aggregates an arbitrary number of observations at test time without relying on MCMC or variational approximations.
Mechanistic simulations are widely used across science and engineering to model complex phenomena where underlying likelihood equations cannot be directly calculated. When estimating system parameters from experimental data, simulation-based inference methods face an operational dilemma. Direct posterior estimation models require multiple, computationally expensive simulator runs per training example to handle multi-observation datasets. Conversely, likelihood surrogate models require only one simulation per training pair and can process arbitrary numbers of observations at test time, but they rely on traditional sampling algorithms that struggle with complex, multimodal distributions and introduce compounding errors.
The article introduces and evaluates Factorized Neural Posterior Score Estimation and its generalized family, Partially Factorized Neural Posterior Score Estimation. The primary objective is to demonstrate a generative score-based modeling framework that accurately infers parameters from multiple observations while maintaining high simulation efficiency and bypassing the failure modes of standard inference samplers.
The authors designed a conditional diffusion framework that models the score functions of individual observation posteriors and combines them additively to generate target posterior samples via annealed Langevin dynamics. They benchmarked these methods across standard scientific models—including Gaussian distributions, Susceptible-Infected-Recovered epidemic dynamics, Lotka–Volterra predator-prey systems, and a high-energy particle physics collision simulator. Performance was tested across simulator budgets ranging from 1,000 to 30,000 calls and evaluated using Maximum Mean Discrepancy and classifier two-sample tests against established neural posterior and neural ratio estimation baselines.
The evaluation revealed four key findings. First, factorized score methods consistently matched or outperformed existing baselines when conditioning on multiple observations across various simulation budgets. Second, the partially factorized framework—grouping small subsets of observations, typically between three and six—achieved the most robust balance between training sample efficiency and reduced error accumulation. Third, score-based composition reliably captured multiple distribution modes in multimodal benchmarks where standard Markov chain Monte Carlo samplers failed to mix. Finally, the framework proved robust across various score network parameterizations and sampling hyperparameters, provided that at least five to ten sampling steps were used per noise level.
These results demonstrate that score modeling provides a practical path to scaling simulation-based inference in data-rich settings without increasing simulation compute costs or risking the mode-collapse failures typical of traditional inference. Organizations applying simulation models to complex, safety-critical, or expensive tasks can achieve higher parameter inference accuracy while avoiding costly simulator reruns.
Teams implementing simulation-based workflows should adopt partially factorized score modeling with small observation partition sizes (such as three to six) when inferring parameters from multiple observations. While the core methodology demonstrates strong empirical confidence across established benchmarks, practitioners should note that the default annealed Langevin sampler adds computational overhead proportional to the step count. Further investigation into faster, hyperparameter-free samplers and sequential active-learning approaches is recommended before deploying at very large production scales.
- Paper: The frontier of simulation-based inference, Kyle Cranmer et al. (2019). This overview establishes the simulation-based inference methods and likelihood-free setting that the source’s factorized score approach builds on.
- Paper: Generative Modeling by Estimating Gradients of the Data Distribution, Yang Song et al. (2019). Its noise-conditional score networks and annealed sampling provide the core score-modeling ideas the source adapts for posterior inference.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Its SDE framework explains the continuous-time score estimation and reverse diffusion machinery underlying the source’s generative modeling.
- Paper: Estimation of Non-Normalized Statistical Models by Score Matching, Aapo Hyvärinen (2005). This foundational score-matching work clarifies how estimating log-density gradients avoids explicit normalization, a key premise of the source’s method.
- Paper: Sequential Neural Score Estimation: Likelihood-Free Inference with Conditional Score Based Diffusion Models, Louis Sharrock et al. (2024). Building on conditional score-based posterior estimation, this paper adds sequential simulation rounds to focus computation on increasingly probable parameter regions.
