Distribution-Free Predictive Inference for Regression
Jing LeiMax G'SellAlessandro RinaldoRyan J. TibshiraniLarry Wasserman
Develops a conformal prediction framework for regression that guarantees finite-sample validity for any base estimator without distributional assumptions, introducing computationally efficient split-conformal algorithms and model-free variable importance metrics.
Modern decision-making increasingly relies on complex, high-dimensional machine learning and regression algorithms to forecast outcomes. However, standard statistical techniques for quantifying prediction uncertainty typically depend on rigid assumptions—such as correct model specification, normal error distributions, or constant variance—that rarely hold in real-world environments. When these assumptions fail, conventional prediction intervals either under-cover true outcomes or become excessively wide, introducing unquantified risk into critical operations. To resolve this problem, the article develops a general, distribution-free predictive inference framework based on conformal inference that constructs mathematically guaranteed prediction intervals around any regression algorithm without requiring restrictive distributional or model-correctness assumptions.
The article evaluates theoretical guarantees, computational trade-offs, and empirical performance across two primary conformal approaches—full conformal inference and split conformal inference—alongside a related leave-one-out jackknife technique. The analysis establishes finite-sample coverage bounds and proves that conformal prediction bands achieve near-optimal width when base models are stable and consistent. To demonstrate practical utility, the authors conduct extensive simulations spanning low-dimensional and high-dimensional settings (up to 2,000 features), including challenging scenarios with heavy-tailed errors, strong feature correlations, and heteroskedasticity. They benchmark linear regression, penalized models (lasso, elastic net, ridge), sparse additive models, and random forests, while providing an open-source R implementation package.
The findings provide strong support for conformal methodology. First, conformal prediction bands strictly achieve the nominal average coverage (for example, approximately 90% across all evaluated configurations) even when models are severely misspecified, nonlinear, or heavy-tailed, whereas standard parametric methods fail or produce overly wide intervals. Second, split conformal inference delivers virtually identical coverage to full conformal inference while drastically reducing computation time—executing in fractions of a second compared to hundreds of seconds in high-dimensional tests. Third, the width of the prediction intervals is directly governed by model accuracy: more predictive estimators yield narrower, more informative intervals. Fourth, combining multiple data splits via standard multi-split aggregation widens prediction bands due to conservative correction penalties, indicating that a single balanced split is preferable. Finally, the authors show that modifying residuals with local error-spread estimates allows the intervals to adapt to varying noise levels across the feature space, and introduce a model-free variable importance measure (leave-one-covariate-out) that reliably identifies predictive features without parametric assumptions.
These results demonstrate that organizations can reliably quantify predictive uncertainty across arbitrary black-box models without risking failure from misspecified assumptions. For operational deployment, decision-makers should adopt split conformal inference as the default standard for generating prediction intervals and evaluating model-free feature importance due to its extreme computational efficiency and robust finite-sample guarantees. Where heteroskedasticity is present, locally weighted conformal inference should be used to ensure uniform local coverage. Future work should focus on developing more efficient multi-split aggregation methods to eliminate data-splitting randomness without inflating interval length, as well as refining variable selection frameworks to address post-selection inference under data splitting.
- Paper: A tutorial on conformal prediction, Glenn Shafer et al. (2007). This tutorial introduces the foundational framework and mathematical principles of conformal prediction under exchangeability that the source directly adapts and extends to regression.
- Paper: Quantile Regression Forests, Nicolai Meinshausen (2006). This paper establishes non-parametric quantile estimation using random forests, providing key background for the source's distribution-free regression and heteroskedasticity-adapting prediction bands.
- Paper: Stability and Generalization, Olivier Bousquet et al. (2002). This work establishes algorithmic stability tools and leave-one-out sensitivity analyses that underpin the theoretical justifications for sample-splitting and jackknife-based predictive inference.
- Paper: A survey of cross-validation procedures for model selection, Sylvain Arlot et al. (2009). This survey provides essential theoretical foundations on cross-validation and sample-splitting mechanics that inform the source's development of split conformal and out-of-sample inference.
- Paper: All Models are Wrong, but Many are Useful: Learning a Variable’s Importance by Studying an Entire Class of Prediction Models Simultaneously, Aaron Fisher et al. (2018). This paper advances model-agnostic variable importance across entire hypothesis classes, extending the leave-one-covariate-out (LOCO) predictive importance perspective introduced in the source.
- Paper: Replicable Conformal Prediction, Marios Papamichalis et al. (2026). This work directly addresses the sample-splitting instability inherent in split conformal prediction by developing replicable calibration protocols while preserving finite-sample coverage guarantees.
- Paper: Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift, Yaniv Ovadia et al. (2019). This study comprehensively evaluates the empirical reliability of predictive uncertainty quantification when standard distribution-free exchangeability assumptions fail under real-world dataset shift.
