Gaussian Processes for Big Data
James HensmanNicolo FusiNeil D. Lawrence
Develops a stochastic variational inference framework for Gaussian processes that overcomes cubic computational constraints, enabling mini-batch training on datasets with millions of data points.
Gaussian process models offer powerful, flexible probabilistic predictions, but their severe computational complexity has historically limited their application to datasets with at most a few thousand data points. Standard exact implementations scale cubically with sample size, and existing approximations still scale poorly as datasets grow into the millions. The article evaluates a new framework that combines inducing variables with stochastic variational inference, decoupling the primary computational cost from total dataset size.
To overcome these limitations, the authors reformulate variational Gaussian process regression by introducing an explicit variational distribution over a set of inducing variables. This explicit formulation separates the overall evidence lower bound into individual data-point contributions, enabling the use of natural gradient optimization on small mini-batches of data. The methodology was evaluated on synthetic benchmark functions and two massive real-world applications: predicting 75,000 UK property prices and forecasting flight delays across an 800,000-flight subset of US airline operations, all running on a single standard processor.
Key findings show that the proposed algorithm dramatically improves predictive performance and scalability over existing approaches. In the UK property price benchmark, the method achieved a mean squared error of 0.426, outperforming traditional subset-based approximations that ranged between 0.502 and 0.522. In the 800,000-point airline dataset, the framework scaled seamlessly across mini-batches of 5,000 points while converging to substantially lower prediction errors than standard subset baselines. Furthermore, removing the computational dependence on sample size allowed the use of 800 to 1,000 inducing variables—an order of magnitude higher than the typical 50 to 100 used in prior sparse approximations—which directly enhanced model expressiveness and accuracy.
These results establish that organizations can deploy rich, non-parametric probabilistic models on enterprise-scale datasets without requiring massive computing clusters. The ability to process data in streaming mini-batches reduces memory consumption, mitigates hardware costs, and unlocks Gaussian process techniques for large spatiotemporal forecasting, complex multi-output systems, and latent variable analysis.
Engineering and analytics teams working with large-scale regression tasks should adopt this stochastic variational approach over conventional subset approximations. Subsequent work should focus on extending this implementation into non-Gaussian classification and latent variable models, as well as developing automated learning-rate tuning strategies to streamline model deployment. While the framework demonstrates robust convergence across large datasets, practitioners should note that learning rates and inducing-point initializations require careful configuration to ensure optimal stability.
- Paper: Variational Learning of Inducing Variables in Sparse Gaussian Processes, Michalis K. Titsias (2009). Introduces the variational inducing variable framework for sparse Gaussian processes upon which the source's mini-batch variational decomposition directly builds.
- Paper: Stochastic variational inference, Matt Hoffman et al. (2012). Formulates stochastic variational inference using mini-batch subsampling and natural gradients, providing the fundamental optimization paradigm adapted by the source for Gaussian processes.
- Paper: Sparse Gaussian Processes using Pseudo-inputs, Edward Snelson et al. (2005). Develops continuous pseudo-inputs for sparse Gaussian process approximations, establishing the concept of learnable inducing points central to scalable GP models.
- Paper: A Unifying View of Sparse Approximate Gaussian Process Regression, Joaquin Quiñonero-Candela et al. (2005). Provides a comprehensive unified framework and taxonomy for inducing-variable and sparse approximations in Gaussian process regression.
- Paper: An Introduction to Variational Methods for Graphical Models, MICHAEL I. JORDAN et al. (1999). Lays down the foundational principles of variational approximation methods and lower bounding in probabilistic graphical models.
- Paper: GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration, Jacob R. Gardner et al. (2018). Modernizes large-scale scalable Gaussian process inference by integrating GPU-accelerated matrix-matrix formulations and variational methods directly into deep learning workflows.
- Paper: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, Yarin Gal et al. (2016). Bridges scalable variational approximations in Gaussian processes with deep neural networks by framing dropout as approximate inference in deep Gaussian processes.
- Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). Synthesizes modern developments in stochastic and black-box variational inference across machine learning and statistical modeling following foundational advances in scalable inference.
