Stacked regressions
LEO BREIMAN
Demonstrates how combining multiple regression models using cross-validation under non-negativity constraints reliably outperforms single-model selection across diverse algorithms like decision trees, subset selection, and ridge regression.
In predictive modeling, practitioners traditionally evaluate a range of competing models and select the single best performer using holdout validation or cross-validation. However, choosing only one model discards valuable information present in alternative predictors and risks poor generalization when models are sensitive to sample variations. The article addresses this limitation by developing stacked regressions, a method for forming linear combinations of different predictors to achieve superior prediction accuracy compared to any individual model.
The article establishes and evaluates a practical framework that uses cross-validated predictions paired with least squares estimation constrained by non-negative weights to determine the combination coefficients. The approach was tested across real-world datasets and extensive computer simulations. Specifically, the analysis evaluated decision trees using the Boston Housing and Ozone datasets, alongside 250 simulation runs involving 40 variables and 60 observations across varied correlation structures to test subset regressions, ridge regressions, and their combinations.
The findings show that stacked regressions consistently outperform the single best model chosen by cross-validation. On the real-world housing and ozone datasets, stacking regression trees reduced prediction error by approximately 10%. In linear simulations, stacking subset regressions and ridge regressions together delivered substantial reductions in model error, outperforming the best-of-best single model selection across all evaluated scenarios. Additionally, the analysis revealed that only a small number of predictors receive positive weights—typically around 3 to 7 models out of dozens—and that 10-fold cross-validation is not only computationally faster than leave-one-out cross-validation but also yields slightly better predictive accuracy.
These results demonstrate that combining diverse predictive models provides greater accuracy and stability than relying on a single selected model. The greatest performance gains occur when combining dissimilar model types, such as subset selection with regularized models, because differing models capture complementary patterns in data. In organizational settings, shifting from single-model selection to stacked combinations can improve decision performance and mitigate model-selection risk without requiring unconstrained, complex combinations that risk overfitting.
Organizations and practitioners should adopt stacked regression with non-negativity constraints when selecting among diverse model candidates. For implementation efficiency, teams should use 10-fold cross-validation rather than leave-one-out methods. While the empirical and simulation evidence strongly supports the method, the article notes that general mathematical proofs are not yet complete, and stacking yields minimal gains when combining highly similar models. Additional evaluation is warranted when applying the approach to other model families such as neural networks or spline-based models.
- Paper: Bagging Predictors, Leo Breiman (1996). Introduces bootstrap aggregation (bagging) to reduce predictor variance, providing essential foundational context for ensemble model combination methods.
- Paper: Regression Shrinkage and Selection Via the Lasso, Robert Tibshirani (1996). Establishes constrained regularized regression techniques that underly the non-negativity and shrinkage constraints used when combining predictors in stacked regressions.
- Paper: A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection, Ron Kohavi (1995). Provides a comprehensive analysis of cross-validation methodology, which is the core mechanism used in stacked regressions to form out-of-fold predictions and prevent overfitting.
- Paper: Greedy function approximation: A gradient boosting machine, Jerome H. Friedman (2001). Presents gradient boosting as an alternative framework for ensembling decision trees and regression models through stagewise optimization.
- Paper: Random Forests, Leo Breiman (2001). Details the random forest paradigm for aggregating randomized regression and classification trees, contrasting simple averaging with meta-learning combinations.
- Paper: FFORMA: Feature-based forecast model averaging, Pablo Montero-Manso et al. (2020). Extends meta-learning and model combination principles from stacked regressions to time-series forecasting via feature-based forecast model averaging.
- Paper: An empirical comparison of supervised learning algorithms, Rich Caruana et al. (2006). Evaluates a wide spectrum of modern supervised algorithms and ensemble architectures across varied metrics and probability calibration methods.
- Paper: Rotation Forest: A New Classifier Ensemble Method, Juan J. Rodríguez et al. (2006). Explores feature-space rotation to induce model diversity, offering an alternative pathway for constructing complementary base predictors for ensemble learning.
- Paper: BART: Bayesian Additive Regression Trees, Hugh A. Chipman et al. (2008). Develops a Bayesian ensemble framework that regularizes sum-of-trees models to provide full posterior inference and uncertainty quantification.
- Paper: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, Mitchell Wortsman et al. (2022). Applies model combination principles to modern neural architectures by averaging model weights directly to gain ensemble-like benefits without test-time overhead.
