The frontier of simulation-based inference

Kyle CranmerJohann BrehmerGilles Louppe

article2019Proceedings of the National Academy of Sciences of the United States of America1,375 citations

Explains how modern machine learning techniques overcome the tractability limits of traditional statistical methods to perform accurate parameter inference directly from complex scientific simulators.

Listen

Modern science relies heavily on complex computer simulations to model physical, biological, and cosmological processes. While these simulators generate high-fidelity synthetic data, they create a major statistical bottleneck known as intractable likelihoods. Because calculating the exact probability of an observed dataset requires integrating over countless unobserved internal execution paths, standard statistical inference breaks down. Traditionally, researchers addressed this by manually simplifying raw data into low-dimensional summary statistics using methods like Approximate Bayesian Computation. However, these classical techniques struggle with high-dimensional observations, require massive amounts of redundant computation, and risk discarding critical scientific information.

The article reviews the evolving frontier of simulation-based inference, evaluating how recent advances in machine learning, active learning, and program integration resolve these traditional bottlenecks. The authors examine the structural mechanics of modern simulation frameworks across diverse scientific domains and systematically compare workflows that replace legacy statistical approximations with neural network surrogates and probabilistic programming paradigms.

The findings indicate that modern deep learning models, particularly neural density estimators and classification-based likelihood ratio estimators, allow statistical inference directly on high-dimensional raw data without requiring hand-crafted summary statistics. Second, training amortized surrogate models allows upfront computational costs to be paid once, enabling near-instant evaluation of new experimental data—an essential feature for datasets with many repeated observations. Third, active learning strategies that iteratively guide the simulator to sample only the most informative parameter regions substantially improve sample efficiency, reducing computational demand. Finally, opening the simulator's internal architecture via automatic differentiation and probabilistic programming allows researchers to augment training data with internal gradients, dramatically accelerating surrogate training and enabling the direct inference of unobservable latent system states.

These methodological advances significantly reduce the computational cost and runtime of complex scientific analyses while increasing inferential precision and statistical power. By transitioning from expert-crafted heuristics to rigorous, scalable statistical surrogates, organizations can model complex physical phenomena more reliably without sacrificing fidelity. Furthermore, training amortized surrogates decouples heavy compute cycles from live data acquisition, mitigating risks associated with real-time operational or experimental bottlenecks.

To adopt these capabilities effectively, teams should evaluate their existing simulation codebases. When internal code modification is feasible, practitioners should expose internal states and implement automatic differentiation to extract joint likelihood gradients. For general applications, teams should prioritize training neural surrogates for likelihood ratios using supervised classification, reserving full probabilistic programming for tasks requiring detailed insight into unobserved internal trajectories. When sample efficiency is paramount, active learning loops should be deployed to adaptively guide simulator runs.

Confidence in these methods depends on proper post-inference validation. Because surrogate models are prone to optimization errors, finite sample biases, and model misspecification, practitioners must calibrate outputs using diagnostic checks such as parametric bootstrapping and ensemble verification. Users should note that while these methods resolve computational intractability, they cannot correct for fundamental discrepancies between an inaccurate simulator and physical reality, which still requires careful parameterization and domain expertise.

arXiv: 1911.01429
  • Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). Normalizing flows provide the foundational generative density-transformation machinery that modern neural simulation-based inference relies upon for flexible posterior and likelihood approximations.
  • Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). Understanding classical variational inference and its optimization-based approximations to intractable posteriors provides essential mathematical context for simulation-based and likelihood-free methodologies.
  • Paper: Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data, Anuj Karpatne et al. (2016). Theory-guided data science establishes the core principles for integrating domain-specific scientific forward models with machine learning architectures.
Cover for The frontier of simulation-based inference

Abstract

Many domains of science have developed complex simulations to describe phenomena of interest. While these simulations provide high-fidelity models, they are poorly suited for inference and lead to challenging inverse problems. We review the rapidly developing field of simulation-based inference and identify the forces giving new momentum to the field. Finally, we describe how the frontier is expanding so that a broad audience can appreciate the profound change these developments may have on science.

Table of Contents

  • 1 Simulation-based inference
  • 1.1 Simulators
  • 1.2 Inference
  • 1.3 Traditional methods
  • 2 Frontiers of simulation-based inference
  • 2.1 A revolution in machine learning
  • 2.2 Active learning
  • 2.3 Integration and augmentation
  • 3 Workflows for simulation-based Inference
  • 3.1 Using the simulator directly during inference
  • 3.2 Surrogate models
  • 3.3 Preprocessing and postprocessing
  • 3.4 Recommendations
  • 4 Discussion
  • References

Knowls

  1. Knowl 1 — Mathematical Formulation of Simulation-Based Inference and Likelihood Intractability

    definition

    In simulation-based inference (also termed likelihood-free inference or implicit statistical modeling), a stochastic simulator defines the statistical forward model. The simulator takes a parameter vector θ∈RM\theta \in \mathbb{R}^M, executes a sequence of stochastic internal transitions generating latent variables zi∼pi(zi∣θ,z<i)z_i \sim p_i(z_i \mid \theta, z_{<i}), and outputs simulated observations x∼p(x∣θ,z)x \sim p(x \mid \theta, z), where z=(z1,z2,… )z = (z_1, z_2, \dots) represents the entire execution trace.

    The marginal likelihood function p(x∣θ)p(x \mid \theta) implicitly defined by the simulator is an integral over the space of all possible latent trajectories:

    p(x∣θ)=∫dz p(x,z∣θ)=∫dz p(x∣θ,z)∏ipi(zi∣θ,z<i)p(x \mid \theta) = \int dz \, p(x, z \mid \theta) = \int dz \, p(x \mid \theta, z) \prod_{i} p_i(z_i \mid \theta, z_{<i})

    Because the latent space zz in real-world mechanistic simulators is high-dimensional, discrete, or dynamically structured across millions of branching steps, this integral is computationally intractable. For an independent and identically distributed (i.i.d.) dataset D={x1,…,xN}\mathcal{D} = \{x_1, \dots, x_N\}, the full likelihood factorizes as:

    p(D∣θ)=∏k=1Npindividual(xk∣θ)p(\mathcal{D} \mid \theta) = \prod_{k=1}^N p_{\text{individual}}(x_k \mid \theta)

    which requires repeated evaluations of the intractable marginal density.

  2. Knowl 2 — Simulator-Augmented Quantities for Supervised Surrogate Training

    definition

    Rather than treating the simulator as an opaque black box that only yields samples x∼p(x∣θ)x \sim p(x \mid \theta), integrating automatic differentiation and probabilistic programming into simulation code exposes six tractable latent quantities from the execution trace (θ,z,x)(\theta, z, x):

    1. Output likelihood given latents: p(x∣z,θ)p(x \mid z, \theta)
    2. Joint score: t(x,z∣θ)≡∇θlog⁡p(x,z∣θ)t(x, z \mid \theta) \equiv \nabla_\theta \log p(x, z \mid \theta)
    3. Latent gradient: ∇zlog⁡p(x,z∣θ)\nabla_z \log p(x, z \mid \theta)
    4. Joint likelihood ratio: r(x,z∣θ,θ′)≡p(x,z∣θ)p(x,z∣θ′)r(x, z \mid \theta, \theta') \equiv \frac{p(x, z \mid \theta)}{p(x, z \mid \theta')}
    5. Parameter sensitivity of data and latents: ∇θ(x,z)\nabla_\theta (x, z)
    6. Latent-to-data Jacobian: ∇zx\nabla_z x

    These augmented variables are extracted during simulation runs to formulate supervised regression and classification loss functions, substantially reducing the sample complexity required to train neural surrogate models compared to unsupervised density estimation.

  3. Knowl 3 — Likelihood Ratio Estimation via the Classifier Trick

    model/method

    The likelihood ratio between two parameter hypotheses, r(x∣θ0,θ1)=p(x∣θ0)p(x∣θ1)r(x \mid \theta_0, \theta_1) = \frac{p(x \mid \theta_0)}{p(x \mid \theta_1)}, or between a specific parameter and a marginal mixture, r(x∣θ)=p(x∣θ)p(x)r(x \mid \theta) = \frac{p(x \mid \theta)}{p(x)}, can be estimated by training a discriminative binary classifier using supervised learning without evaluating marginal densities.

    Simulated datasets are generated at parameter values θ0\theta_0 (assigned label y=1y=1) and θ1\theta_1 (assigned label y=0y=0) with equal prior probabilities p(y=1)=p(y=0)=0.5p(y=1) = p(y=0) = 0.5. A neural network dϕ(x)d_\phi(x) parameterized by weights ϕ\phi is trained with binary cross-entropy loss to predict the class probability p(y=1∣x)p(y=1 \mid x). The Bayes-optimal decision function is:

    d∗(x)=p(x∣θ0)p(x∣θ0)+p(x∣θ1)d^*(x) = \frac{p(x \mid \theta_0)}{p(x \mid \theta_0) + p(x \mid \theta_1)}

    Inverting this relation yields the likelihood ratio estimator:

    r^(x∣θ0,θ1)=dϕ(x)1−dϕ(x)\hat{r}(x \mid \theta_0, \theta_1) = \frac{d_\phi(x)}{1 - d_\phi(x)}

    This estimator enables amortized frequentist hypothesis testing via the generalized likelihood ratio test as well as Bayesian posterior sampling through Markov Chain Monte Carlo (MCMC) or Variational Inference (VI).

  4. Knowl 4 — Taxonomy of Neural Surrogate Paradigms for Simulation-Based Inference

    model/method

    Modern simulation-based inference methods employ deep neural networks to construct amortized surrogates across three distinct functional targets:

    1. Amortized Posterior Estimation: Neural conditional density estimators (such as normalizing flows or autoregressive models) directly approximate the posterior qϕ(θ∣x)≈p(θ∣x)q_\phi(\theta \mid x) \approx p(\theta \mid x). This provides rapid direct evaluation and sampling of parameters given observations, but conditions the surrogate on the specific prior p(θ)p(\theta) used during training.

    2. Amortized Likelihood Estimation: Generative neural density models approximate the conditional likelihood qϕ(x∣θ)≈p(x∣θ)q_\phi(x \mid \theta) \approx p(x \mid \theta). This approach enables prior-independent inference and synthetic data generation, but requires downstream sampling algorithms (such as MCMC) to obtain parameter posteriors.

    3. Amortized Likelihood Ratio Estimation: Parameterized neural classifiers estimate ratios r^ϕ(x∣θ0,θ1)≈p(x∣θ0)p(x∣θ1)\hat{r}_\phi(x \mid \theta_0, \theta_1) \approx \frac{p(x \mid \theta_0)}{p(x \mid \theta_1)} or r^ϕ(x∣θ,θref)\hat{r}_\phi(x \mid \theta, \theta_{\text{ref}}). This converts density estimation into supervised classification, eliminating the need to normalize densities over high-dimensional observation spaces.

  5. Knowl 5 — Optimal Data Compression via Local Score Estimation

    model/method

    To compress high-dimensional observation vectors x∈RDx \in \mathbb{R}^D into low-dimensional summary statistics without loss of Fisher information, inference frameworks utilize the score function of the marginal likelihood:

    t(x∣θ0)≡∇θlog⁡p(x∣θ)∣θ=θ0t(x \mid \theta_0) \equiv \nabla_\theta \log p(x \mid \theta) \Big|_{\theta = \theta_0}

    In the neighborhood of a fiducial parameter point θ0\theta_0, the MM-dimensional score vector t(x∣θ0)t(x \mid \theta_0) forms a set of locally sufficient statistics for the MM-dimensional parameter space θ\theta. Compression x↦t(x∣θ0)x \mapsto t(x \mid \theta_0) preserves the Fisher information matrix I(θ0)=Ex∼p(x∣θ0)[t(x∣θ0)t(x∣θ0)T]\mathcal{I}(\theta_0) = \mathbb{E}_{x \sim p(x \mid \theta_0)} [t(x \mid \theta_0) t(x \mid \theta_0)^T].

    When p(x∣θ)p(x \mid \theta) is intractable, neural score estimators t^ϕ(x)\hat{t}_\phi(x) are trained directly on simulator outputs. If the joint score t(x,z∣θ0)=∇θlog⁡p(x,z∣θ0)t(x, z \mid \theta_0) = \nabla_\theta \log p(x, z \mid \theta_0) is accessible from the simulator, t^ϕ(x)\hat{t}_\phi(x) is trained by minimizing the mean squared error:

    L(ϕ)=E(x,z)∼p(x,z∣θ0)[∥t^ϕ(x)−t(x,z∣θ0)∥2]\mathcal{L}(\phi) = \mathbb{E}_{(x, z) \sim p(x, z \mid \theta_0)} \left[ \left\| \hat{t}_\phi(x) - t(x, z \mid \theta_0) \right\|^2 \right]

  6. Knowl 6 — Probabilistic Programming and Inference Compilation for Latent Trace Inference

    model/method

    Probabilistic programming integrates inference capabilities into the execution flow of the simulator, enabling inference over both the parameters θ\theta and the full latent execution trajectory zz by targeting the joint posterior p(θ,z∣xobs)p(\theta, z \mid x_{\text{obs}}).

    Inference Compilation (IC) amortizes this process using a trained neural network:

    1. Offline Training: The simulator generates execution traces (θ,z,x)(\theta, z, x) across varied parameter settings. A deep neural network (e.g., a recurrent or autoregressive model) is trained via supervised learning to approximate the conditional distribution over execution choices: qϕ(zi∣x,z<i,θ)q_\phi(z_i \mid x, z_{<i}, \theta).
    2. Online Sampling: When an observation xobsx_{\text{obs}} is provided, the trained proposal network qϕq_\phi guides sequential importance sampling (SIS) or Sequential Monte Carlo (SMC) runs through the simulator, steering stochastic branching decisions toward paths that match xobsx_{\text{obs}}.

    This approach provides physical interpretability by reconstructing unobservable internal states zz consistent with observed data.

  7. Knowl 7 — Active Learning and Amortization Trade-off in Simulation-Based Inference

    model/method

    Simulation-based inference presents an operational trade-off between amortization and targeted sample efficiency:

    • Amortized Methods: The simulator is executed over broad prior distributions p(θ)p(\theta) to train global neural surrogates. Once trained, the surrogate evaluates arbitrary new observations xobsx_{\text{obs}} with minimal computation, making it suitable for processing large streams of i.i.d. data points.
    • Active Learning Methods: The simulator parameter proposals θk+1\theta_{k+1} are iteratively selected based on acquisition criteria computed from intermediate posteriors p(θ∣xobs)p(\theta \mid x_{\text{obs}}) or confidence boundaries tailored to a single observation xobsx_{\text{obs}}.

    Active learning dramatically reduces the total number of expensive simulator evaluations required for a specific dataset, but sacrifices amortization because the resulting surrogate is specialized to xobsx_{\text{obs}} and must be retrained for different observations.

  8. Knowl 8 — Statistical Diagnostics and Calibration for Neural Surrogates in SBI

    model/method

    Because neural estimators are prone to optimization errors, finite sample bias, and expressiveness limits, statistical validity is evaluated using several diagnostics:

    1. Parametric Bootstrap Coverage: Synthetic pairs (θj,xj)∼p(θ)p(x∣θ)(\theta_j, x_j) \sim p(\theta) p(x \mid \theta) are generated from the simulator. For a nominal confidence/credibility level 1−α1 - \alpha, empirical coverage is computed as the empirical frequency with which θj\theta_j falls within the estimated region Θ^1−α(xj)\hat{\Theta}_{1-\alpha}(x_j). Deviations from 1−α1 - \alpha indicate miscalibration.
    2. Classifier Two-Sample Diagnostic: A secondary binary classifier is trained to distinguish true simulated samples x∼p(x∣θ)x \sim p(x \mid \theta) from samples drawn from the surrogate x^∼p^(x∣θ)\hat{x} \sim \hat{p}(x \mid \theta). Inability of the classifier to perform above chance level confirms surrogate fidelity.
    3. Expectation Consistency Checks: Known identity constraints are checked on independent test data, including the score zero-mean condition Ex∼p(x∣θ)[∇θlog⁡p^(x∣θ)]=0\mathbb{E}_{x \sim p(x \mid \theta)}[\nabla_\theta \log \hat{p}(x \mid \theta)] = 0 and likelihood ratio normalization Ex∼p(x∣θ1)[r^(x∣θ0,θ1)]=1\mathbb{E}_{x \sim p(x \mid \theta_1)}[\hat{r}(x \mid \theta_0, \theta_1)] = 1.
    4. Asymptotic Distribution Checks: Likelihood ratio test statistics are compared against standard asymptotic χ2\chi^2 distributions (Wilks' and Wald's theorems) in large-sample limits.
  9. Knowl 9 — Limitations of Traditional Approximate Bayesian Computation and Classical Density Estimation

    limitation

    Traditional simulation-based inference approaches suffer from major computational and statistical bottlenecks:

    1. Rejection ABC Tolerance Trade-off: Rejection ABC approximates the likelihood via P(ρ(xsim,xobs)<ϵ)\mathbb{P}(\rho(x_{\text{sim}}, x_{\text{obs}}) < \epsilon) for distance metric ρ\rho and tolerance ϵ\epsilon. In the limit ϵ→0\epsilon \to 0, ABC is exact, but the acceptance probability approaches zero for continuous data. Large ϵ\epsilon introduces substantial bias and statistical power loss.
    2. Non-Amortization of ABC: ABC couples the simulation and acceptance steps to the specific observation xobsx_{\text{obs}}. Every new data point requires re-running the entire simulation and rejection chain from scratch.
    3. Curse of Dimensionality in Density Estimation: Classical density estimators (histograms, Kernel Density Estimation) scale exponentially with data dimension, failing when applied to raw data beyond 3 to 4 dimensions.
    4. Information Loss from Summary Statistics: Both ABC and classical density estimation require compressing high-dimensional data xx into hand-crafted, low-dimensional summary statistics y(x)y(x), which discard informative features and reduce inference precision.

Coverage note — None was omitted; all key theoretical formulations, workflow taxonomies, simulator-augmented quantities, compression techniques, and diagnostic methods contributed by the review are captured.

References

  1. 1.Diggle PJ, Gratton RJ (1984) Monte Carlo Methods of Inference for Implicit Statistical Models in Journal of the Royal Statistical Society: Series B (Methodological). Vol. 46, pp. 193–212.
  2. 2.Rubin DB (1984) Bayesianly Justifiable and Relevant Frequency Calculations for the Applied Statistician. The Annals of Statistics 12(4):1151–1172.
  3. 3.Beaumont MA, Zhang W, Balding DJ (2002) Approximate Bayesian computation in population genetics. Genetics 162(4):2025–2035.
  4. 4.Mohamed S, Lakshminarayanan B (2016) Learning in Implicit Generative Models. arXiv:1610.03483.
  5. 5.CMS Collaboration (2015) CMS collisions at 13 TeV (http://cds.cern.ch/record/2021490).
  6. 6.Lee WCA, et al. (2006) Dynamic remodeling of dendritic arbors in GABAergic interneurons of adult visual cortex. PLoS Biology 4(2):271–280.
  7. 7.IBM Research (2019) Public Health Research - The Spatiotemporal Epidemiological Modeler (STEM) (https://researcher.watson.ibm.com/researcher/view_group.php?id=883).
  8. 8.Hubble Space Telescope (2011) Based on observations made with the NASA/ESA (https://apod.nasa.gov/apod/image/1112/lensshoe_hubble_3235.jpg).
  9. 9.ESA and the Planck Collaboration (2018) Planck image gallery (https://www.cosmos.esa.int/web/planck/picture-gallery).
  10. 10.Metropolis N, Rosenbluth AW, Rosenbluth MN, Teller AH, Teller E (1953) Equation of state calculations by fast computing machines. The Journal of Chemical Physics 21(6):1087–1092.
  11. 11.Hastings WK (1970) Monte carlo sampling methods using Markov chains and their applications. Biometrika 57(1):97–109.
  12. 12.Saul L, Jordan M (1998) A Mean Field Learning Algorithm for Unsupervised Neural Networks. Learning in Graphical Models 1:541–554.
  13. 13.Sisson SA (2018) Handbook of Approximate Bayesian Computation. (Chapman and Hall/CRC).
  14. 14.Marjoram P, Molitor J, Plagnol V, Tavaré S (2003) Markov chain Monte Carlo without likelihoods. Proceedings of the National Academy of Sciences of the United States of America 100(26):15324–15328.
  15. 15.Sisson SA, Fan Y, Tanaka MM (2007) Sequential Monte Carlo without likelihoods. Proceedings of the National Academy of Sciences of the United States of America 104(6):1760–1765.
  16. 16.Peters GW, Fan Y, Sisson SA (2012) On sequential Monte Carlo, partial rejection control and approximate Bayesian computation. Statistics and Computing 22(6):1209–1222.
  17. 17.Brehmer J, Cranmer K, Louppe G, Pavez J (2018) A guide to constraining effective field theories with machine learning. Physical Review D 98(5):052004.
  18. 18.Lecun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521(7553):436–444.
  19. 19.Dinh L, Krueger D, Bengio Y (2015) NICE: Non-linear independent components estimation. 3rd International Conference on Learning Representations, ICLR 2015 - Workshop Track Proceedings.
  20. 20.Dinh L, Sohl-Dickstein J, Bengio S (2019) Density estimation using real NVP in 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings.
  21. 21.Kingma DP, Dhariwal P (2018) Glow: Generative flow with invertible 1×1 convolutions. Advances in Neural Information Processing Systems 2018-December:10215–10224.
  22. 22.Germain M, Gregor K, Murray I, Larochelle H (2015) MADE: Masked autoencoder for distribution estimation. 32nd International Conference on Machine Learning, ICML 2015 2:881–889.
  23. 23.Uria B, Cote MA, Gregor K, Murray I, Larochelle H (2016) Neural autoregressive distribution estimation. Journal of Machine Learning Research 17:1–37.
  24. 24.Van Den Oord A, Kalchbrenner N, Kavukcuoglu K (2016) Pixel recurrent neural networks. 33rd International Conference on Machine Learning, ICML 2016 4:2611–2620.
  25. 25.Van Den Oord A, et al. (2016) Conditional image generation with PixelCNN decoders. Advances in Neural Information Processing Systems pp. 4797–4805.
  26. 26.van den Oord A, et al. (2016) WaveNet: A Generative Model for Raw Audio. arXiv:1609.03499.
  27. 27.Kingma DP, et al. (2016) Improved variational inference with inverse autoregressive flow. Advances in Neural Information Processing Systems pp. 4743–4751.
  28. 28.Papamakarios G, Pavlakou T, Murray I (2017) Masked autoregressive flow for density estimation. Advances in Neural Information Processing Systems 2017-December:2339–2348.
  29. 29.Huang CW, Krueger D, Lacoste A, Courville A (2018) Neural autoregressive flows. 35th International Conference on Machine Learning, ICML 2018 5:3309–3324.
  30. 30.Wehenkel A, Louppe G (2019) Unconstrained monotonic neural networks. CEUR Workshop Proceedings 2491:arXiv:1908.05164.
  31. 31.Durkan C, Bekasov A, Murray I, Papamakarios G (2019) Cubic-Spline Flows. arXiv:1906.02145.
  32. 32.Durkan C, Bekasov A, Murray I, Papamakarios G (2019) Neural Spline Flows. arXiv:1906.04032.
  33. 33.Knauf A (2018) Ordinary Differential Equations. UNITEXT - La Matematica per il 3 piu 2 109:31–60.
  34. 34.Grathwohl W, Chen RT, Bettencourt J, Sutskever I, Duvenaud D (2019) Ffjord: Free-form continuous dynamics for scalable reversible generative models in 7th International Conference on Learning Representations, ICLR 2019.
  35. 35.Meeds E, Welling M (2014) GPS-ABC: Gaussian process surrogate approximate Bayesian computation in Uncertainty in Artificial Intelligence - Proceedings of the 30th Conference, UAI 2014, UAI'14. (AUAI Press, Arlington, Virginia, United States), pp. 593–602.
  36. 36.Gutmann MU, Corander J (2016) Bayesian optimization for likelihood-free inference of simulator-based statistical models. Journal of Machine Learning Research 17(1):1–47.
  37. 37.Meeds E, Welling M (2015) Optimization Monte Carlo: Efficient and embarrassingly parallel likelihood-free inference. Advances in Neural Information Processing Systems 2015-January:2080–2088.
  38. 38.Järvenpää M, Gutmann MU, Pleska A, Vehtari A, Marttinen P (2019) Efficient acquisition rules for model-based approximate Bayesian computation. Bayesian Analysis 14(2):595–622.
  39. 39.Wang H, Li J (2018) Adaptive Gaussian process approximation for Bayesian inference with expensive likelihood functions. Neural Computation 30(11):3072–3094.
  40. 40.Lueckmann JM, et al. (2017) Flexible statistical inference for mechanistic models of neural dynamics. Advances in Neural Information Processing Systems 2017-December:1290–1300.
  41. 41.Papamakarios G, Sterratt DC, Murray I (2018) Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows. International Conference on Artificial Intelligence and Statistics.
  42. 42.Ranjan P, Bingham D, Michailidis G (2008) Sequential experiment design for contour estimation from complex computer codes. Technometrics 50(4):527–541.
  43. 43.Bect J, Ginsbourger D, Li L, Picheny V, Vazquez E (2012) Sequential design of computer experiments for the estimation of a probability of failure. Statistics and Computing 22(3):773–793.
  44. 44.Heinrich L, Louppe G, Cranmer K (2018) excursion (doi:10.5281/zenodo.1634428).
  45. 45.Cutler M, Walsh TJ, How JP (2014) Reinforcement learning with multi-fidelity simulators in 2014 IEEE International Conference on Robotics and Automation (ICRA). (IEEE), pp. 3888–3895.
  46. 46.Hamrick JB, et al. (2019) Metacontrol for adaptive imagination-based optimization. 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings.
  47. 47.Kandasamy K, Dasarathy G, Schneider J, Póczos B (2017) Multi-fidelity Bayesian optimisation with continuous approximations in 34th International Conference on Machine Learning, ICML 2017. (JMLR. org), Vol. 4, pp. 2861–2878.
  48. 48.Gordon AD, Henzinger TA, Nori AV, Rajamani SK (2014) Probabilistic programming in Future of Software Engineering, FOSE 2014 - Proceedings. pp. 167–181.
  49. 49.Doucet A, Johansen AM (2009) A tutorial on particle filtering and smoothing: Fifteen years later. Handbook of Nonlinear Filtering 12(December):4–6.
  50. 50.Baydin AG, et al. (2019) Etalumis: Bringing probabilistic programming to scientific simulators at scale. International Conference for High Performance Computing, Networking, Storage and Analysis, SC p. arXiv:1907.03382.
  51. 51.Graham MM, Storkey AJ (2017) Asymptotically exact inference in differentiable generative models. Electronic Journal of Statistics 11(2):5105–5164.
  52. 52.Brehmer J, Louppe G, Pavez J, Cranmer K (2020) Mining gold from implicit models to improve likelihood-free inference. Proceedings of the National Academy of Sciences of the United States of America 117(10):5242–5249.
  53. 53.Stoye M, Brehmer J, Louppe G, Pavez J, Cranmer K (2018) Likelihood-free inference with an improved cross-entropy estimator. arXiv:1808.00973.
  54. 54.Güne ş Baydin A, Pearlmutter BA, Andreyevich Radul A, Mark Siskind J (2018) Automatic differentiation in machine learning: A survey. Journal of Machine Learning Research 18(153):1–43.
  55. 55.Olah C (2015) Neural Networks, Types, and Functional Programming.
  56. 56.LeCun Y (2018) Deep Learning est mort. Vive Differentiable Programming! (https://www.facebook.com/yann.lecun/posts/10155003011462143).
  57. 57.Gutmann MU, Dutta R, Kaski S, Corander J (2018) Likelihood-free inference via classification. Statistics and Computing 28(2):411–425.
  58. 58.Le TA, Baydin AG, Wood F (2017) Inference compilation and universal probabilistic programming in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, Proceedings of Machine Learning Research. (PMLR, Fort Lauderdale, FL, USA), Vol. 54, pp. 1338–1348.
  59. 59.Tarantola A (2005) Inverse Problem Theory and Methods for Model Parameter Estimation. (siam) Vol. 89.
  60. 60.Putzky P, Welling M (2017) Recurrent Inference Machines for Solving Inverse Problems. arXiv:1706.04008.
  61. 61.Pesah A, Wehenkel A, Louppe G (2018) Recurrent machines for likelihood-free inference in arXiv:1811.12932.
  62. 62.Louppe G, Hermans J, Cranmer K (2019) Adversarial variational optimization of non-differentiable simulators in CEUR Workshop Proceedings, Proceedings of Machine Learning Research, eds. Chaudhuri K, Sugiyama M. (PMLR), Vol. 2491, pp. 1438–1447.
  63. 63.Rezende DJ, Mohamed S (2015) Variational inference with normalizing flows. 32nd International Conference on Machine Learning, ICML 2015 2:1530–1538.
  64. 64.Papamakarios G, Murray I (2016) Fast e-free inference of simulation models with Bayesian conditional density estimation in Advances in Neural Information Processing Systems. pp. 1036–1044.
  65. 65.Izbicki R, Lee AB, Pospisil T (2019) ABC–CDE: Toward Approximate Bayesian Computation With Complex High-Dimensional Data and Limited Simulations. Journal of Computational and Graphical Statistics 28(3):481–492.
  66. 66.Paige B, Wood F (2016) Inference networks for sequential monte carlo in graphical models. 33rd International Conference on Machine Learning, ICML 2016 6:4434–4444.
  67. 67.Tran D, Ranganath R, Blei DM (2017) Hierarchical implicit models and likelihood-free variational inference in Advances in Neural Information Processing Systems, eds. Guyon I, et al. Vol. 2017-December, pp. 5524–5534.
  68. 68.Durkan C, Papamakarios G, Murray I (2018) Sequential Neural Methods for Likelihood-free Inference. arXiv:1811.08723.
  69. 69.Lueckmann JM, Bassetto G, Karaletsos T, Macke JH (2018) Likelihood-free inference with emulator networks in Proceedings of The 1st Symposium on Advances in Approximate Bayesian Inference, Proceedings of Machine Learning Research, eds. Ruiz F, Zhang C, Liang D, Bui T. (PMLR), Vol. 96, pp. 32–53.
  70. 70.Alsing J, Charnock T, Feeney S, Wandelt B (2019) Fast likelihood-free cosmology with neural density estimators and active learning. Monthly Notices of the Royal Astronomical Society 488(3):4440–4458.
  71. 71.Neal RM (2008) Computing likelihood functions for high-energy physics experiments when distributions are defined by simulators with nuisance parameters in PHYSTAT LHC Workshop on Statistical Issues for LHC Physics, PHYSTAT 2007 - Proceedings. pp. 111–118.
  72. 72.Fan Y, Nott DJ, Sisson SA (2013) Approximate Bayesian computation via regression density estimation. Stat 2(1):34–48.
  73. 73.Cranmer K, Pavez J, Louppe G (2015) Approximating Likelihood Ratios with Calibrated Discriminative Classifiers. arXiv:1506.02169.
  74. 74.Thomas O, Dutta R, Corander J, Kaski S, Gutmann MU (2016) Likelihood-free inference by ratio estimation. arXiv:1611.10242.
  75. 75.Dinev T, Gutmann MU (2018) Dynamic Likelihood-free Inference via Ratio Estimation (DIRE). arXiv:1810.09899.
  76. 76.Hermans J, Begy V, Louppe G (2019) Likelihood-free MCMC with Amortized Approximate Ratio Estimators. arXiv:1903.04057.
  77. 77.Andreassen A, Nachman B (2019) Neural Networks for Full Phase-space Reweighting and Parameter Tuning. arXiv:1907.08209.
  78. 78.Brehmer J, Cranmer K, Louppe G, Pavez J (2018) Constraining Effective Field Theories with Machine Learning. Physical Review Letters 121(11):111801.
  79. 79.Louppe G, Cho K, Becot C, Cranmer K (2019) QCD-aware recursive neural networks for jet physics. Journal of High Energy Physics 2019(1):57.
  80. 80.Andreassen A, Feige I, Frye C, Schwartz MD (2019) JUNIPR: a framework for unsupervised machine learning in particle physics. European Physical Journal C 79(2):102.
  81. 81.Carleo G, et al. (2019) Machine learning and the physical sciences. Reviews of Modern Physics 91(4).
  82. 82.Munk A, et al. (2019) Deep Probabilistic Surrogate Networks for Universal Simulator Approximation. arXiv:1910.11950.
  83. 83.Alsing J, Wandelt B (2018) Generalized massive optimal data compression. Monthly Notices of the Royal Astronomical Society: Letters 476(1):L60–L64.
  84. 84.Alsing J, Wandelt B, Feeney S (2018) Massive optimal data compression and density estimation for scalable, likelihood-free inference in cosmology. Monthly Notices of the Royal Astronomical Society 477(3):2874–2885.
  85. 85.de Castro P, Dorigo T (2019) INFERNO: Inference-Aware Neural Optimisation. Computer Physics Communications 244:170–179.
  86. 86.Alsing J, Wandelt B (2019) Nuisance hardened data compression for fast likelihood-free inference. Monthly Notices of the Royal Astronomical Society 488(4):5093–5103.
  87. 87.Wilks SS (1938) The Large-Sample Distribution of the Likelihood Ratio for Testing Composite Hypotheses. The Annals of Mathematical Statistics 9(1):60–62.
  88. 88.Wald A (1943) Tests of Statistical Hypotheses Concerning Several Parameters When the Number of Observations is Large. Transactions of the American Mathematical Society 54(3):426.
  89. 89.Cowan G, Cranmer K, Gross E, Vitells O (2011) Asymptotic formulae for likelihood-based tests of new physics. European Physical Journal C 71(2):1554.
  90. 90.Wood F, Van De Meent JW, Mansinghka V (2014) A new approach to probabilistic programming inference in Journal of Machine Learning Research. Vol. 33, pp. 1024–1032.

Citation

MLA
Cranmer, K., et al. “The Frontier of Simulation-based Inference”. Proceedings of the National Academy of Sciences, vol. 117, no. 48, 2020, pp. 30055–62, https://doi.org/10.1073/pnas.1912789117.
APA
Cranmer, K., Brehmer, J., & Louppe, G. (2020). The frontier of simulation-based inference. Proceedings of the National Academy of Sciences, 117(48), 30055–30062. https://doi.org/10.1073/pnas.1912789117
Chicago
Cranmer, K., J. Brehmer, and G. Louppe. 2020. “The Frontier of Simulation-based Inference”. Proceedings of the National Academy of Sciences 117 (48): 30055–62. https://doi.org/10.1073/pnas.1912789117.
Harvard
Cranmer, K., Brehmer, J. and Louppe, G. (2020) “The frontier of simulation-based inference”, Proceedings of the National Academy of Sciences, 117(48), pp. 30055–30062. Available at: https://doi.org/10.1073/pnas.1912789117.
Vancouver
1. Cranmer K, Brehmer J, Louppe G (2020) The frontier of simulation-based inference. Proceedings of the National Academy of Sciences 117:30055–30062

BibTeX

@article{Cranmer_2020, title={The frontier of simulation-based inference}, volume={117}, ISSN={1091-6490}, url={http://dx.doi.org/10.1073/pnas.1912789117}, DOI={10.1073/pnas.1912789117}, number={48}, journal={Proceedings of the National Academy of Sciences}, publisher={National Academy of Sciences}, author={Cranmer, Kyle and Brehmer, Johann and Louppe, Gilles}, year={2020}, month=May, pages={30055–30062} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF