Pareto Smoothed Importance Sampling

Aki VehtariDaniel SimpsonAndrew GelmanYuling YaoJonah Gabry

article2024JMLR311 citations

Introduces a generalized Pareto tail-smoothing method and diagnostic that stabilize heavy-tailed importance ratios and reliably assess finite-sample convergence in Monte Carlo computation.

Listen

Modern Bayesian data analysis and computational statistics rely heavily on importance sampling to adjust simulations when sampling directly from a target model is computationally intractable. However, standard importance sampling often suffers from extreme weight volatility and infinite variance, particularly in high-dimensional settings or when proposal distributions poorly approximate target tails. Simple remedies, such as raw weight truncation, can introduce excessive bias. The article introduces Pareto Smoothed Importance Sampling (PSIS) to stabilize importance weights and establishes a finite-sample convergence diagnostic, known as the Pareto k-hat diagnostic, to evaluate estimator reliability across Monte Carlo applications.

The PSIS approach fits a generalized Pareto distribution to the upper tail of observed importance ratios—specifically targeting a fraction of the largest weights—and replaces extreme values with their expected order statistics. The article supports this method with mathematical proofs of simulation consistency and finite variance under mild conditions. In addition, the article evaluates the method through controlled synthetic experiments across varying dimensions, posterior approximation benchmarks using split-normal densities for logistic Gaussian process models, and leave-one-out cross-validation across 105 predictive breast cancer genomic models.

The findings establish that PSIS consistently achieves lower root mean squared error and better bias-variance trade-offs than standard or truncated importance sampling. The Pareto k-hat diagnostic functions as a reliable finite-sample indicator: estimates and standard error computations remain highly reliable when k-hat is below 0.7, whereas k-hat values above 0.7 indicate severe pre-asymptotic bias and slow convergence. In practical cross-validation across 105 genomic models, PSIS-LOO reduced computation time from an estimated 125 days of exact model refits to under one minute, while successfully pinpointing individual influential data points that required targeted exact computation.

These results provide practitioners and organizational leaders with a way to dramatically accelerate machine learning evaluation and Bayesian inference pipelines without sacrificing statistical validity. By pairing variance stabilization with an explicit failure warning system, the method prevents silent model failures and miscalibrated uncertainty estimates in high-stakes predictive systems.

Decision-makers should adopt PSIS-based tools for model validation and approximate inference workflows while establishing automated policy triggers when k-hat exceeds 0.7. When high k-hat values occur, teams should fall back on hybrid exact computation for flagged points, refine approximating proposals, or investigate potential data outliers and model misspecification. Although PSIS is highly robust, users should exercise caution when sample sizes are very small or when weight distributions possess complex multi-modal tails, ensuring that diagnostic warnings are actively monitored rather than ignored.

arXiv: 1507.02646

No sufficiently relevant recommendations were found.

Cover for Pareto Smoothed Importance Sampling

Abstract

Importance weighting is a general way to adjust Monte Carlo integration to account for draws from the wrong distribution, but the resulting estimate can be highly variable when the importance ratios have a heavy right tail. This routinely occurs when there are aspects of the target distribution that are not well captured by the approximating distribution, in which case more stable estimates can be obtained by modifying extreme importance ratios. We present a new method for stabilizing importance weights using a generalized Pareto distribution fit to the upper tail of the distribution of the simulated importance ratios. The method, which empirically performs better than existing methods for stabilizing importance sampling estimates, includes stabilized effective sample size estimates, Monte Carlo error estimates, and convergence diagnostics. The presented Pareto k̂ finite sample convergence rate diagnostic is useful for any Monte Carlo estimator.

Table of Contents

  • 1. Introduction
  • 2. Stabilizing Importance Sampling Estimates by Modifying the Ratios
  • 2.1 Modeling the Tail of the Importance Ratios
  • 2.2 Our Proposal: Pareto Smoothed Importance Sampling
  • 3. Using k^ as a Diagnostic
  • 3.1 PSIS Is Reliable When k^ < 0.7
  • 3.2 Pareto Means and Central Limit Theorems
  • 3.2.1 Generalized Central Limit Theorem
  • 3.2.2 Pareto Means and IS L1 Deviation
  • 3.2.3 Truncated Pareto Means and PSIS RMSE
  • 3.2.4 Sample Size Dependent k^-threshold
  • 3.2.5 PSIS vs IS L1 Deviation
  • 3.2.6 PSIS RMSE Convergence Rate Given k^
  • 3.2.7 Only the Tail Has Pareto Shape
  • 3.2.8 h-specific Estimates
  • 3.2.9 Finite Sample Size and Bounded Ratio Distributions
  • 3.3 k^ Is a Good Diagnostic in Finite Samples and in High Dimensions
  • 3.4 Other Minimum Sample Size Estimates
  • 4. Convergence of PSIS
  • 4.1 Asymptotic Consistency
  • 4.2 Asymptotic Normality
  • 5. Practical Examples
  • 5.1 Improving Distributional Approximation with Importance Sampling
  • 5.2 Importance-sampling Leave-one-out Cross-validation
  • 5.2.1 LOO for Stack Loss Data
  • 5.2.2 LOO for 105 Protein Expression Data Sets
  • 6. Discussion
  • 6.1 Other Examples in the Literature
  • 6.2 Other Weight Transformations
  • 6.3 Multiple and Adaptive Importance Sampling
  • 6.4 Population Monte Carlo and Particle Filters
  • Acknowledgments
  • Appendix A. Scaling of Distribution of Mean of Truncated Means
  • Appendix B. Relative Convergence Rates
  • Appendix C. Proof of Theorem 1
  • Appendix D. Simulated Examples
  • D.1 Exponential Target and Proposal
  • D.2 Univariate Normal and Student’s t
  • D.3 Multivariate Normal and Student’s t
  • Appendix E. h-specific k^h
  • Appendix F. Marginal Distribution of k
  • Appendix G. Regularization of Pareto Fit for Small S
  • Appendix H. Adjusting M in Case of Dependent MCMC Draws
  • Appendix I. Stan Program for Linear Regression on the Stack Loss Data
  • References

Knowls

  1. Knowl 1 — Pareto Smoothed Importance Sampling Algorithm

    algorithm

    Pareto Smoothed Importance Sampling (PSIS) stabilizes importance weights in self-normalized importance sampling by fitting a generalized Pareto distribution (GPD) to the upper tail of the raw importance ratios and replacing extreme ratios with expected order statistics.

    Given target distribution p(θ)p(\theta), proposal distribution g(θ)g(\theta), and draws θ1,…,θS∼g(θ)\theta_1, \dots, \theta_S \sim g(\theta), the raw importance ratios are rs=p(θs)/g(θs)r_s = p(\theta_s)/g(\theta_s) for s=1,…,Ss = 1, \dots, S.

    Input: Raw importance ratios rsr_s for s=1,…,Ss = 1, \dots, S, sorted in ascending order such that r1≤r2≤⋯≤rSr_1 \le r_2 \le \dots \le r_S
    Output: Smoothed importance weights wsw_s for s=1,…,Ss = 1, \dots, S, and diagnostic shape parameter k^\hat{k}
    Set M=⌊min⁡(0.2S,3S)⌋M = \lfloor \min(0.2S, 3\sqrt{S}) \rfloor
    for s=1s = 1 to S−MS - M do
        Set ws=rsw_s = r_s
    end for
    Set threshold u^=rS−M\hat{u} = r_{S-M}
    Fit generalized Pareto distribution parameters (k^,σ^)(\hat{k}, \hat{\sigma}) to the MM largest ratios {rS−M+1,…,rS}\{r_{S-M+1}, \dots, r_S\} with cutpoint u^\hat{u}
    for z=1z = 1 to MM do
        Compute expected quantile F−1(z−1/2M)=u^+σ^k^((1−z−1/2M)−k^−1)F^{-1}\left(\frac{z - 1/2}{M}\right) = \hat{u} + \frac{\hat{\sigma}}{\hat{k}}\left(\left(1 - \frac{z - 1/2}{M}\right)^{-\hat{k}} - 1\right)
        Set wS−M+z=min⁡(F−1(z−1/2M),max⁡1≤s≤S(rs))w_{S-M+z} = \min\left(F^{-1}\left(\frac{z - 1/2}{M}\right), \max_{1 \le s \le S}(r_s)\right)
    end for
    if k^>min⁡(1−1/log⁡10(S),0.7)\hat{k} > \min(1 - 1/\log_{10}(S), 0.7) then
        Issue warning that estimates are likely unstable or have high bias
    end if
    return w1,…,wSw_1, \dots, w_S and k^\hat{k}

    The self-normalized PSIS estimator for the expectation Ih=Ep[h(θ)]I_h = \mathbb{E}_p[h(\theta)] is given by: I^h=∑s=1Swsh(θs)∑s=1Sws\hat{I}_h = \frac{\sum_{s=1}^S w_s h(\theta_s)}{\sum_{s=1}^S w_s}

    Replacing the MM largest ratios with expected order statistics truncates extreme outliers, replacing heavy-tailed empirical noise with smooth tail behavior, thereby enforcing finite variance on the estimator while maintaining minimal finite-sample bias.

  2. Knowl 2 — Asymptotic Consistency and Rate of Convergence of PSIS

    theoretical result

    Let θs∼g(θ)\theta_s \sim g(\theta) for s=1,…,Ss = 1, \dots, S be independent and identically distributed samples, and let rs=r(θs)=p(θs)/g(θs)r_s = r(\theta_s) = p(\theta_s)/g(\theta_s) denote the importance ratios distributed according to cumulative distribution function RR. Consider the idealized PSIS estimator where M=O(S1/2)M = \mathcal{O}(S^{1/2}) tail weights are smoothed using fixed Generalized Pareto parameters kk and σ\sigma.

    1. General Consistency: If RR is absolutely continuous, M=O(S1/2)M = \mathcal{O}(S^{1/2}), h∈L2(p)h \in L^2(p), and σ=O(rS−M+1:S)\sigma = \mathcal{O}(r_{S-M+1:S}), then the PSIS estimator: I^hS=1S∑s=1S(r(θs)∧rS−M+1:S)h(θs)+1S∑j=1Mwj+h(θS−M+j)\hat{I}_h^S = \frac{1}{S}\sum_{s=1}^S \left(r(\theta_s) \wedge r_{S-M+1:S}\right) h(\theta_s) + \frac{1}{S}\sum_{j=1}^M w_j^+ h(\theta_{S-M+j}) converges in L1L^1 to Ep[h(θ)]\mathbb{E}_p[h(\theta)] and its variance goes to zero as S→∞S \to \infty. Thus, the estimator is simulation-consistent and asymptotically unbiased.

    2. Convergence Rate under von Mises Condition: If additionally RR satisfies the von Mises regularity condition at infinity: lim⁡r→∞rR′(r)1−R(r)=1k(1+O(r−1))\lim_{r \to \infty} \frac{r R'(r)}{1 - R(r)} = \frac{1}{k}\left(1 + \mathcal{O}(r^{-1})\right) and h(θ)h(\theta) is bounded, then:

    • The L1L^1 error converges at a rate of at most O(S−1/2)\mathcal{O}(S^{-1/2}).
    • The variance of the PSIS estimator decays at least as fast as O(Sk/2−1)\mathcal{O}(S^{k/2 - 1}).

    Consequently, the PSIS estimator achieves S\sqrt{S}-consistency under these conditions.

  3. Knowl 3 — Pareto Shape Parameter Diagnostic for Importance Sampling Reliability

    model/method

    The estimated shape parameter k^\hat{k} of the generalized Pareto distribution fit to the upper tail of importance ratios rs=p(θs)/g(θs)r_s = p(\theta_s)/g(\theta_s) serves as a diagnostic for the finite-sample stability, convergence rate, and bias of importance sampling estimators:

    1. k^<0.5\hat{k} < 0.5: The raw importance ratios have finite variance. The standard central limit theorem holds, the estimator converges at the rate O(S−1/2)\mathcal{O}(S^{-1/2}), and standard variance-based Monte Carlo standard error (MCSE) estimates are accurate.

    2. 0.5≤k^<min⁡(1−1/log⁡10(S),0.7)0.5 \le \hat{k} < \min(1 - 1/\log_{10}(S), 0.7) (with threshold ≈0.7\approx 0.7 for S>2000S > 2000): The raw ratios have infinite variance, but PSIS smoothing stabilizes the estimator so that variance dominates bias. The estimator converges at a reduced rate O(Sk−1)\mathcal{O}(S^{k-1}), and variance-based MCSE estimates remain reliable.

    3. 0.7≤k^≤10.7 \le \hat{k} \le 1: The bias introduced by smoothing begins to dominate the variance. The required sample size to control root mean square error (RMSE) becomes impractically large (S≈101/(1−k)S \approx 10^{1/(1-k)}), and variance-based MCSE estimates severely underestimate the true error.

    4. k^>1\hat{k} > 1: The distribution of importance ratios does not possess a finite mean, and any Monte Carlo estimate for the mean is invalid.

    This diagnostic applies to self-normalized importance sampling, unnormalized importance sampling, and general Monte Carlo integration.

  4. Knowl 4 — Finite-Sample Scaling and Relative Convergence Rates of Truncated Pareto Means

    theoretical result

    Approximating the smoothed tail weights in PSIS via a truncated generalized Pareto distribution with upper cutoff y=Sky = S^k (corresponding to the largest expected order statistic among SS draws) yields the following scaling regimes for standard deviation, bias, root mean square error (RMSE), and relative convergence efficiency α\alpha (where error scales as (S−1/2)α(S^{-1/2})^\alpha):

    1. For k<0.5−0.5/log⁡10(S)k < 0.5 - 0.5/\log_{10}(S):

      • Standard deviation scales as S−1/2S^{-1/2}.
      • Bias is negligible.
      • Relative convergence efficiency is α≈1\alpha \approx 1.
    2. For k=0.5k = 0.5:

      • Standard deviation scales as (S/log⁡S)−1/2(S/\log S)^{-1/2}.
      • Relative convergence efficiency is α=1−1/log⁡(S)\alpha = 1 - 1/\log(S).
    3. For k∈(0.5+0.5/log⁡10(S),1)k \in (0.5 + 0.5/\log_{10}(S), 1):

      • Standard deviation scales as Sk−1S^{k-1}.
      • Bias scales as Sk−1S^{k-1}.
      • Overall RMSE scales as Sk−1S^{k-1}.
      • Relative convergence efficiency is α=2(1−k)\alpha = 2(1 - k).

    To maintain a target RMSE (e.g., 10%10\% relative error), the required minimum sample size scales as: S≈101/(1−k)S \approx 10^{1/(1-k)} which increases rapidly for k>0.7k > 0.7.

  5. Knowl 5 — Monte Carlo Standard Error and Effective Sample Size for PSIS

    equation

    For a self-normalized PSIS estimator I^h=∑s=1Sw~sh(θs)\hat{I}_h = \sum_{s=1}^S \tilde{w}_s h(\theta_s) with normalized smoothed weights w~s=ws/∑s′=1Sws′\tilde{w}_s = w_s / \sum_{s'=1}^S w_{s'} and mean estimate μ~=∑s=1Sw~sh(θs)\tilde{\mu} = \sum_{s=1}^S \tilde{w}_s h(\theta_s), the estimated Monte Carlo standard error variance Var^(I^h)\widehat{\text{Var}}(\hat{I}_h) is:

    Var^(I^h)≈∑s=1Sw~s2(h(θs)−μ~)2for independent draws\widehat{\text{Var}}(\hat{I}_h) \approx \sum_{s=1}^S \tilde{w}_s^2 (h(\theta_s) - \tilde{\mu})^2 \quad \text{for independent draws}

    Var^(I^h)≈1Reff,MCMC∑s=1Sw~s2(h(θs)−μ~)2for MCMC draws\widehat{\text{Var}}(\hat{I}_h) \approx \frac{1}{R_{\text{eff,MCMC}}} \sum_{s=1}^S \tilde{w}_s^2 (h(\theta_s) - \tilde{\mu})^2 \quad \text{for MCMC draws}

    where Reff,MCMC=Seff,MCMC/SR_{\text{eff,MCMC}} = S_{\text{eff,MCMC}} / S is the relative MCMC efficiency computed using the split-chain effective sample size method for h(θ)h(\theta).

    The function-specific effective sample size ESSh\text{ESS}_h is: ESSh≈1S∑s=1S(h(θs)−1S∑s′=1Sh(θs′))2Var^(I^h)\text{ESS}_h \approx \frac{\frac{1}{S}\sum_{s=1}^S \left(h(\theta_s) - \frac{1}{S}\sum_{s'=1}^S h(\theta_{s'})\right)^2}{\widehat{\text{Var}}(\hat{I}_h)}

    The generic effective sample size for the normalization term is: ESS≈1∑s=1Sw~s2\text{ESS} \approx \frac{1}{\sum_{s=1}^S \tilde{w}_s^2}

    Given the Pareto shape parameter k^\hat{k}, an optimistic effective sample size estimate is: ESSk≈S10k^/(1−k^)\text{ESS}_k \approx \frac{S}{10^{\hat{k}/(1-\hat{k})}}

    These variance-based MCSE and ESS estimates are reliable when k^<min⁡(1−1/log⁡10(S),0.7)\hat{k} < \min(1 - 1/\log_{10}(S), 0.7).

  6. Knowl 6 — Function-Specific Tail Shape Diagnostic

    model/method

    When estimating the expectation of an unbounded function h(θ)h(\theta), the tail behavior of the product h(θ)r(θ)h(\theta)r(\theta) under proposal θ∼g(θ)\theta \sim g(\theta) can have heavier tails than the importance ratio r(θ)r(\theta) alone.

    The function-specific tail diagnostic k^h\hat{k}_h is defined by fitting generalized Pareto distributions to both the left tail (MM smallest values) and right tail (MM largest values) of h(θ)r(θ)h(\theta)r(\theta): k^h=max⁡(k^left,k^right)\hat{k}_h = \max\left(\hat{k}_{\text{left}}, \hat{k}_{\text{right}}\right) where M=⌊min⁡(0.2S,3S)⌋M = \lfloor \min(0.2S, 3\sqrt{S}) \rfloor.

    Usage and Properties:

    1. When k^h>k^\hat{k}_h > \hat{k}, the integrand function h(θ)h(\theta) dominates the tail behavior of the numerator in self-normalized importance sampling, and k^h\hat{k}_h determines the convergence rate and reliability of I^h\hat{I}_h.
    2. For leave-one-out cross-validation where h(θ)=1/r(θ)h(\theta) = 1/r(\theta), the denominator can dominate the tail behavior.
    3. Smoothing must only be applied to the raw ratios r(θ)r(\theta) and not directly to h(θ)r(θ)h(\theta)r(\theta); separately smoothing numerator and denominator terms introduces inconsistent normalization bias.
  7. Knowl 7 — Pre-asymptotic Failure of Importance Sampling with Bounded Ratios in High Dimensions

    empirical result

    In high-dimensional settings, choosing a proposal distribution g(θ)g(\theta) with heavier tails than the target p(θ)p(\theta) (such as a multivariate Student-tt proposal with degrees of freedom ν=7\nu = 7 for a multivariate Gaussian target) guarantees that the importance ratios r(θ)=p(θ)/g(θ)r(\theta) = p(\theta)/g(\theta) are bounded with finite variance asymptotically. However, for finite sample sizes SS, practical convergence can completely fail:

    1. Tail Behavior: As dimension DD increases (tested from D=1D = 1 to D=1024D = 1024 with S=10 000S = 10\,000), the theoretical bound on the ratio grows astronomically (e.g., ≈2×1077\approx 2 \times 10^{77} for D=1024D = 1024), lying far beyond the support explored by finite samples.
    2. Diagnostic Sensitivity: The Pareto k^\hat{k} diagnostic detects the empirical heavy-tailed nature: k^\hat{k} transitions from <0< 0 (indicating boundedness) in low dimensions to k^>0\hat{k} > 0, exceeding 0.70.7 for moderate to high dimensions (D≥128D \ge 128).
    3. ESS Collapse: The effective sample size collapses to a few draws, and the convergence rate of the root mean square error drops from O(S−1/2)\mathcal{O}(S^{-1/2}) to approximately O(S−1/4)\mathcal{O}(S^{-1/4}) or worse by D=512D = 512.
    4. Practical Implication: Theoretical boundedness of importance ratios does not guarantee practical importance sampling convergence in finite samples; empirical diagnostics like k^\hat{k} are necessary to detect pre-asymptotic breakdown.
  8. Knowl 8 — Pareto Smoothed Importance Sampling for Leave-One-Out Cross-Validation

    model/method

    Pareto Smoothed Importance Sampling for Leave-One-Out Cross-Validation (PSIS-LOO) evaluates out-of-sample predictive performance from full posterior draws without refitting the Bayesian model for each excluded observation.

    For nn observations y1,…,yny_1, \dots, y_n and SS posterior draws θ1,…,θS∼p(θ∣y)\theta_1, \dots, \theta_S \sim p(\theta | y) from the full posterior, the ii-th leave-one-out predictive density is estimated as: p(yi∣y−i)≈∑s=1Swi(θs)p(yi∣θs)∑s=1Swi(θs)p(y_i | y_{-i}) \approx \frac{\sum_{s=1}^S w_i(\theta_s) p(y_i | \theta_s)}{\sum_{s=1}^S w_i(\theta_s)} where the raw importance ratios are ri(θs)=1/p(yi∣θs)r_i(\theta_s) = 1/p(y_i | \theta_s), and wi(θs)w_i(\theta_s) are the PSIS-smoothed weights.

    Key Properties:

    1. Speed: PSIS-LOO evaluates cross-validation across all nn data points in seconds to minutes, compared to days or weeks for exact LOO (e.g., evaluating 105 regression models on 283 breast cancer tumor samples required less than 1 minute for PSIS-LOO versus 125 days for exact LOO).
    2. Outlier Detection and Hybrid Refitting (PSIS-LOO+): The shape parameter k^i\hat{k}_i computed for each observation flags influential outliers or model misspecifications when k^i>0.7\hat{k}_i > 0.7. For observations with k^i>0.7\hat{k}_i > 0.7, exact refitting on only those specific points yields exact cross-validation accuracy with minimal computational overhead.
  9. Knowl 9 — Small-Sample and MCMC Sample Size Adjustments for Tail Index Estimation

    model/method

    Fitting the Generalized Pareto Distribution to the upper tail of importance ratios requires adjustments for small sample sizes SS and for correlated Markov chain Monte Carlo (MCMC) draws:

    1. Small-Sample Regularization: For small sample sizes (such as S<1000S < 1000), the raw shape parameter estimate k^raw\hat{k}_{\text{raw}} from the profile likelihood quadrature method is regularized toward 0.50.5: k^=Mk^raw+10⋅0.5M+10\hat{k} = \frac{M \hat{k}_{\text{raw}} + 10 \cdot 0.5}{M + 10} where MM is the number of tail observations. This acts as a weakly informative prior equivalent to 10 pseudo-observations at k=0.5k = 0.5, reducing estimation variance and RMSE for small SS without altering large-sample asymptotics.

    2. MCMC Autocorrelation Adjustment: When draws are generated via MCMC rather than independent sampling, the effective number of independent tail draws is reduced. The number of tail ratios MM selected for GPD fitting is adjusted as: M=⌊min⁡(0.2S,3SReff,MCMC)⌋M = \left\lfloor \min\left(0.2S, 3\sqrt{\frac{S}{R_{\text{eff,MCMC}}}}\right) \right\rfloor where Reff,MCMC=Seff,MCMC/SR_{\text{eff,MCMC}} = S_{\text{eff,MCMC}}/S is the relative MCMC efficiency computed using the split-chain effective sample size estimate for the importance ratios r(θ)r(\theta).

Coverage note — Omitted specific Stan code implementations (Appendix I), generic reviews of alternative clipping methods from the literature, and detailed proofs of intermediate lemmas (Appendices C-G).

References

  1. 1.S. Agapiou, Omiros Papaspiliopoulos, D. Sanz-Alonso, and A. M. Stuart. Importance sampling: Intrinsic dimension and computational cost. Statistical Science, 32(3):405–431, 2017.
  2. 2.Viljami Aittomäki. MicroRNA regulation in breast cancer—a Bayesian analysis of expression data. Master’s thesis, Aalto University, 2016.
  3. 3.M. R. Aure, S. Jernström, M. Krohn, H. K. Vollan, E. U. Due, E. Rødland, R. Kåresen, P. Ram, Y. Lu, G. B. Mills, K. K. Sahlberg, A. L. Børresen-Dale, O. C. Lingjærde, and V. N. Kristensen. Integrated analysis reveals microRNA networks coordinately expressed with key proteins in breast cancer. Genome Medicine, 7(1):21, 2015.
  4. 4.A. A. Balkema and L. de Haan. Residual life time at great age. The Annals of Probability, 2 (5):792 – 804, 1974.
  5. 5.Peter J. Bickel. Some contributions to the theory of order statistics. In Berkeley Symposium on Mathematical Statistics and Probability, pages 575–591, 1967.
  6. 6.Eli Bingham, Jonathan P. Chen, Martin Jankowiak, Fritz Obermeyer, Neeraj Pradhan, Theofanis Karaletsos, Rohit Singh, Paul Szerlip, Paul Horsfall, and Noah D. Goodman. Pyro: Deep universal probabilistic programming. Journal of Machine Learning Research, 20(1):973–978, 2019.
  7. 7.Gunnar Blom. Statistical Estimates and Transformed Beta-Variables. Wiley, 1958.
  8. 8.Jean-Philippe Bouchaud and Antoine Georges. Anomalous diffusion in disordered media: statistical mechanisms, models and physical applications. Physics reports, 195(4-5): 127–293, 1990.
  9. 9.Monica F. Bugallo, Víctor Elvira, Luca Martino, David Luengo, Joaquín Míguez, and Petar M. Djurić. Adaptive importance sampling: The past, the present, and the future. IEEE Signal Processing Magazine, 34(4):60–79, 2017.
  10. 10.Paul-Christian Bürkner, Jonah Gabry, and Aki Vehtari. Efficient leave-one-out cross-validation for Bayesian non-factorized normal and Student-t models. Computational Statistics, 36:1243–1261, 2020a.
  11. 11.Paul-Christian Bürkner, Jonah Gabry, and Aki Vehtari. Approximate leave-future-out cross-validation for Bayesian time series models. Journal of Statistical Computation and Simulation, 90:2499–2523, 2020b.
  12. 12.Paul Bürkner, Jonah Gabry, Matthew Kay, and Aki Vehtari. posterior: Tools for working with posterior distributions. R package version 1.5.0.9000, 2024. URL https://mc-stan. org/posterior/.
  13. 13.Olivier Cappé, Arnaud Guillin, Jean-Michel Marin, and Christian P. Robert. Population Monte Carlo. Journal of Computational and Graphical Statistics, 13(4):907–929, 2004.
  14. 14.Joshua C. Chang, Xiangting Li, Shixin Xu, Hao-Ren Yao, Julia Porcino, and Carson Chow. Gradient-flow adaptive importance sampling for Bayesian leave one out cross-validation for sigmoidal classification models. arXiv preprint arXiv:2402.08151, 2024.
  15. 15.Sourav Chatterjee and Persi Diaconis. The sample size required in importance sampling. The Annals of Applied Probability, 28(2):1099–1135, 2018.
  16. 16.Dan Crisan and Boris Rozovskiǐ. The Oxford Handbook of Nonlinear Filtering. Oxford University Press, 2011.
  17. 17.Herbert A. David and Haikady N. Nagaraja. Order Statistics. John Wiley & Sons, 2004.
  18. 18.Akash Kumar Dhaka, Alejandro Catalina, Michael Riis Andersen, Måns Magnusson, Jonathan H. Huggins, and Aki Vehtari. Robust, accurate stochastic optimization for variational inference. In Advances in Neural Information Processing Systems, volume 33, pages 10961–10973, 2020.
  19. 19.Akash Kumar Dhaka, Alejandro Catalina, Manushi Welandawe, Michael Riis Andersen, Jonathan H. Huggins, and Aki Vehtari. Challenges and opportunities in high dimensional variational inference. In Thirty-Fifth Conference on Neural Information Processing Systems, volume 34, pages 7787–7798, 2021.
  20. 20.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real NVP. In International Conference on Learning Representations, 2017.
  21. 21.Víctor Elvira, Luca Martino, David Luengo, and Mónica F Bugallo. Efficient multiple importance sampling estimators. IEEE Signal Processing Letters, 22(10):1757–1761, 2015.
  22. 22.Víctor Elvira, Luca Martino, David Luengo, and Mónica F. Bugallo. Generalized multiple importance sampling. Statistical Science, 34(1):129–155, 2019.
  23. 23.Víctor Elvira, Luca Martino, and Christian P. Robert. Rethinking the effective sample size. International Statistical Review, 2022. doi:10.1111/insr.12500.
  24. 24.Ilenia Epifani, Steven N. MacEachern, and Mario Peruggia. Case-deletion importance sampling estimators: Central limit theorems and related results. Electronic Journal of Statistics, 2:774–806, 2008.
  25. 25.Michael Falk and Frank Marohn. Von Mises conditions revisited. The Annals of Probability, pages 1310–1328, 1993.
  26. 26.Edwin Fong and Chris Holmes. Conformal Bayesian computation. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 18268–18279, 2021.
  27. 27.Jonah Gabry, Daniel Simpson, Aki Vehtari, Michael Betancourt, and Andrew Gelman. Visualization in Bayesian workflow. Journal of the Royal Statistical Society: Series A, 182 (2):389–402, 2019.
  28. 28.Alan E. Gelfand, D. K. Dey, and H. Chang. Model determination using predictive distributions with implementation via sampling-based methods. In J. M. Bernardo, J. O. Berger, A. P. Dawid, and A. F. M. Smith, editors, Bayesian Statistics 4, pages 147–167. Oxford University Press, 1992.
  29. 29.John Geweke. Bayesian inference in econometric models using Monte Carlo integration. Econometrica, 57(6):1317–1339, 1989.
  30. 30.Philip S. Griffin. Asymptotic normality of Winsorized means. Stochastic Processes and Their Applications, 29(1):107–127, 1988.
  31. 31.J. M. Hammersley and D. C. Handscomb. Monte Carlo Methods. Chapman and Hall, 1964.
  32. 32.Edward L. Ionides. Truncated importance sampling. Journal of Computational and Graphical Statistics, 17(2):295–311, 2008.
  33. 33.Noa Kallioinen, Topi Paananen, Paul-Christian Bürkner, and Aki Vehtari. Detecting and diagnosing prior and likelihood sensitivity with power-scaling. Statistics and Computing, 34:57, 2024. doi: 10.1007/s11222-023-10366-5.
  34. 34.Eugenia Koblents and Joaquín Míguez. A population Monte Carlo scheme with transformed weights and its application to stochastic kinetic models. Statistics and Computing, 25(2): 407–425, 2015.
  35. 35.Augustine Kong. A note on importance sampling using standardized weights. Technical Report 348, University of Chicago, Department of Statistics, 1992.
  36. 36.Augustine Kong, Jun S. Liu, and Wing Hung Wong. Sequential imputations and Bayesian missing data problems. Journal of the American Statistical Association, 89(425):278–288, 1994.
  37. 37.Ravin Kumar, Colin Carroll, Ari Hartikainen, and Osvaldo Antonio Martín. ArviZ a unified library for exploratory analysis of Bayesian models in Python. Journal of Open Source Software, 4(33):1143, 2019.
  38. 38.Feynman T. Liang, Liam Hodgkinson, and Michael W. Mahoney. A heavy-tailed algebra for probabilistic programming. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 19353–19365, 2023.
  39. 39.David J. C. MacKay. Information Theory, Inference and Learning Algorithms. Cambridge University Press, 2003.
  40. 40.Måns Magnusson, Michael Andersen, Johan Jonasson, and Aki Vehtari. Leave-one-out cross-validation for Bayesian model comparison in large data. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS), 108:341–351, 2020.
  41. 41.Måns Magnusson, Michael Riis Andersen, Johan Jonasson, and Aki Vehtari. Bayesian leave-one-out cross-validation for large data. Proceedings of the 36th International Conference on Machine Learning (ICML), PMLR 97:4244–4253, 2019.
  42. 42.Luca Martino, Víctor Elvira, J. Míguez, A. Artés-Rodríguez, and P. M. Djurić. A comparison of clipping strategies for importance sampling. In 2018 IEEE Statistical Signal Processing Workshop (SSP), pages 558–562, 2018.
  43. 43.Cory McCartan. adjustr: Stan model adjustments and sensitivity analyses using importance sampling. R package version 0.1.2. Available online at: https://corymccartan.github.io/adjustr/ (accessed December 8, 2021), 2021.
  44. 44.Yann McLatchie, Sölvi Rögnvaldsson, Frank Weber, and Aki Vehtari. Robust and efficient projection predictive inference. arXiv preprint arXiv:2306.15581, 2023.
  45. 45.Joaquín Míguez. On the performance of nonlinear importance samplers and population Monte Carlo schemes. In 2017 22nd International Conference on Digital Signal Processing (DSP), pages 1–5. IEEE, 2017.
  46. 46.Art Owen and Yi Zhou. Safe and effective importance sampling. Journal of the American Statistical Association, 95(449):135–143, 2000.
  47. 47.Art B. Owen. Monte Carlo theory, methods and examples, 2013. Online book at http: //statweb.stanford.edu/~owen/mc/, accessed 2017-09-09.
  48. 48.Topi Paananen, Juho Piironen, Paul-Christian Bürkner, and Aki Vehtari. Implicitly adaptive importance sampling. Statistics and Computing, 31(16), 2021.
  49. 49.Mario Peruggia. On the variability of case-deletion importance sampling weights in the Bayesian linear model. Journal of the American Statistical Association, 92(437):199–207, 1997.
  50. 50.James Pickands. Statistical inference using extreme order statistics. Annals of Statistics, 3: 119–131, 1975.
  51. 51.Juho Piironen, Markus Paasiniemi, and Aki Vehtari. Projective inference in high-dimensional problems: Prediction and feature selection. Electronic Journal of Statistics, 14(1):2155– 2197, 2020.
  52. 52.Juho Piironen, Markus Paasiniemi, Alejandro Catalina, Frank Weber, and Aki Vehtari. projpred: Projection predictive feature selection, 2023. URL https://mc-stan.org/projpred/. R package version 2.8.0.
  53. 53.Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, pages 1530–1538, 2015.
  54. 54.Jaakko Riihimäki and Aki Vehtari. Laplace approximation for logistic Gaussian process density estimation and regression. Bayesian Analysis, 9(2):425–448, 2014.
  55. 55.Christian P. Robert and George Casella. Monte Carlo Statistical Methods, second edition. Springer, 2004.
  56. 56.Daniel Sanz-Alonso. Importance sampling and necessary sample size: An information theory approach. SIAM/ASA Journal on Uncertainty Quantification, 6(2):867–879, 2018.
  57. 57.Carl Scarrot and Anna MacDonald. A review of extreme value threshold estimation and uncertainty quantification. REVSTAT – Statistical Journal, 10(1):33–60, 2012.
  58. 58.S. G. J. Senarathne, Christopher C. Drovandi, and James M. McGree. A Laplace-based algorithm for Bayesian adaptive design. Statistics and Computing, 30(5):1183–1208, 2020.
  59. 59.Luca Alessandro Silva and Giacomo Zanella. Robust leave-one-out cross-validation for high-dimensional Bayesian models. Journal of the American Statistical Association, pages 1–13, 2023.
  60. 60.Stan Development Team. Stan Modeling Language: User’s Guide and Reference Manual, 2017. Version 2.16.0, https://mc-stan.org/.
  61. 61.Juho Timonen, Nikolas Siccha, Ben Bales, Harri Lähdesmäki, and Aki Vehtari. An importance sampling approach for reliable and efficient inference in Bayesian ordinary differential equation models. Stat, 12(1):e614, 2023.
  62. 62.Vladimir V. Uchaikin and Vladimir M. Zolotarev. Chance and Stability: Stable Distributions and Their Applications. VSP International Science Publishers, 1999.
  63. 63.Jarno Vanhatalo, Jaakko Riihimäki, Jouni Hartikainen, Pasi Jylänki, Ville Tolvanen, and Aki Vehtari. GPstuff: Bayesian modeling with Gaussian processes. Journal of Machine Learning Research, 14:1175–1179, 2013.
  64. 64.Aki Vehtari and Andrew Gelman. Pareto smoothed importance sampling. arXiv preprint arXiv:1507.02646v2, 2015.
  65. 65.Aki Vehtari, Tommi Mononen, Ville Tolvanen, Tuomas Sivula, and Ole Winther. Bayesian leave-one-out cross-validation approximations for Gaussian latent variable models. Journal of Machine Learning Research, 17(103):1–38, 2016.
  66. 66.Aki Vehtari, Andrew Gelman, and Jonah Gabry. Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC. Statistics and Computing, 27(5):1413–1432, 2017.
  67. 67.Aki Vehtari, Andrew Gelman, Daniel Simpson, Bob Carpenter, and Paul-Christian Bürkner. Rank-normalization, folding, and localization: An improved R̂ for assessing convergence of MCMC. Bayesian Analysis, 16(2):667–718, 2021.
  68. 68.Aki Vehtari, Jonah Gabry, Måns Magnusson, Yuling Yao, Paul Bürkner, Topi Paananen, and Andrew Gelman. loo: Efficient leave-one-out cross-validation and WAIC for Bayesian models, R package version 2.7.0.9000., 2024. URL https://mc-stan.org/loo/.
  69. 69.Mattias Villani and Rolf Larsson. The multivariate split normal distribution and asymmetric principal components analysis. Communications in Statistics: Theory and Methods, 35(6): 1123–1140, 2006.
  70. 70.Sumio Watanabe. Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory. Journal of Machine Learning Research, 11:3571–3594, 2010.
  71. 71.Yuling Yao, Aki Vehtari, Daniel Simpson, and Andrew Gelman. Yes, but did it work?: Evaluating variational inference. Proceedings of the 35th International Conference on Machine Learning, PMLR 80:5581–5590, 2018a.
  72. 72.Yuling Yao, Aki Vehtari, Daniel Simpson, and Andrew Gelman. Using stacking to average Bayesian predictive distributions (with discussion). Bayesian Analysis, 13(3):917–1003, 2018b.
  73. 73.Yuling Yao, Collin Cademartori, Aki Vehtari, and Andrew Gelman. Adaptive path sampling in metastable posterior distributions. arXiv preprint arXiv:2009.00471, 2020.
  74. 74.Yuling Yao, Gregor Pirš, Aki Vehtari, and Andrew Gelman. Bayesian hierarchical stacking: Some models are (somewhere) useful. Bayesian Analysis, 2021. doi: 10.1214/21-BA1287.
  75. 75.Yuling Yao, Aki Vehtari, and Andrew Gelman. Stacking for non-mixing Bayesian computations: The curse and blessing of multimodal posteriors. Journal of Machine Learning Research, 23(79):1–45, 2022.
  76. 76.I. V. Zaliapin, Yan Y. Kagan, and Federic P. Schoenberg. Approximating the distribution of Pareto sums. Pure and Applied Geophysics, 162(6-7):1187–1228, 2005.
  77. 77.Jin Zhang. Improving on estimation for the generalized Pareto distribution. Technometrics, 52(3):335–339, 2010.
  78. 78.Jin Zhang and Michael A Stephens. A new and efficient estimation method for the generalized Pareto distribution. Technometrics, 51(3):316–325, 2009.
  79. 79.Lu Zhang, Bob Carpenter, Andrew Gelman, and Aki Vehtari. Pathfinder: Parallel quasi-newton variational inference. Journal of Machine Learning Research, 23(306), 2022.

Citation

MLA
Vehtari, A., et al. “Pareto Smoothed Importance Sampling”. Journal of Machine Learning Research, vol. 25, no. 72, 2024, pp. 1–8, https://www.jmlr.org/papers/v25/19-556.html.
APA
Vehtari, A., Simpson, D., Gelman, A., Yao, Y., & Gabry, J. (2024). Pareto Smoothed Importance Sampling. Journal of Machine Learning Research, 25(72), 1–58. https://www.jmlr.org/papers/v25/19-556.html
Chicago
Vehtari, A., D. Simpson, A. Gelman, Y. Yao, and J. Gabry. 2024. “Pareto Smoothed Importance Sampling”. Journal of Machine Learning Research 25 (72): 1–58. https://www.jmlr.org/papers/v25/19-556.html.
Harvard
Vehtari, A. et al. (2024) “Pareto Smoothed Importance Sampling”, Journal of Machine Learning Research, 25(72), pp. 1–58. Available at: https://www.jmlr.org/papers/v25/19-556.html.
Vancouver
1. Vehtari A, Simpson D, Gelman A, Yao Y, Gabry J (2024) Pareto Smoothed Importance Sampling. Journal of Machine Learning Research 25:1–58

BibTeX

@article{JMLR:v25:19-556,
  author  = {Aki Vehtari and Daniel Simpson and Andrew Gelman and Yuling Yao and Jonah Gabry},
  title   = {Pareto Smoothed Importance Sampling},
  journal = {Journal of Machine Learning Research},
  year    = {2024},
  volume  = {25},
  number  = {72},
  pages   = {1--58},
  url     = {http://jmlr.org/papers/v25/19-556.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/