Gaussian Processes for Big Data

James HensmanNicolo FusiNeil D. Lawrence

article2013UAI1,396 citations

Develops a stochastic variational inference framework for Gaussian processes that overcomes cubic computational constraints, enabling mini-batch training on datasets with millions of data points.

Listen

Gaussian process models offer powerful, flexible probabilistic predictions, but their severe computational complexity has historically limited their application to datasets with at most a few thousand data points. Standard exact implementations scale cubically with sample size, and existing approximations still scale poorly as datasets grow into the millions. The article evaluates a new framework that combines inducing variables with stochastic variational inference, decoupling the primary computational cost from total dataset size.

To overcome these limitations, the authors reformulate variational Gaussian process regression by introducing an explicit variational distribution over a set of inducing variables. This explicit formulation separates the overall evidence lower bound into individual data-point contributions, enabling the use of natural gradient optimization on small mini-batches of data. The methodology was evaluated on synthetic benchmark functions and two massive real-world applications: predicting 75,000 UK property prices and forecasting flight delays across an 800,000-flight subset of US airline operations, all running on a single standard processor.

Key findings show that the proposed algorithm dramatically improves predictive performance and scalability over existing approaches. In the UK property price benchmark, the method achieved a mean squared error of 0.426, outperforming traditional subset-based approximations that ranged between 0.502 and 0.522. In the 800,000-point airline dataset, the framework scaled seamlessly across mini-batches of 5,000 points while converging to substantially lower prediction errors than standard subset baselines. Furthermore, removing the computational dependence on sample size allowed the use of 800 to 1,000 inducing variables—an order of magnitude higher than the typical 50 to 100 used in prior sparse approximations—which directly enhanced model expressiveness and accuracy.

These results establish that organizations can deploy rich, non-parametric probabilistic models on enterprise-scale datasets without requiring massive computing clusters. The ability to process data in streaming mini-batches reduces memory consumption, mitigates hardware costs, and unlocks Gaussian process techniques for large spatiotemporal forecasting, complex multi-output systems, and latent variable analysis.

Engineering and analytics teams working with large-scale regression tasks should adopt this stochastic variational approach over conventional subset approximations. Subsequent work should focus on extending this implementation into non-Gaussian classification and latent variable models, as well as developing automated learning-rate tuning strategies to streamline model deployment. While the framework demonstrates robust convergence across large datasets, practitioners should note that learning rates and inducing-point initializations require careful configuration to ensure optimal stability.

arXiv: 1309.6835
Cover for Gaussian Processes for Big Data

Abstract

We introduce stochastic variational inference for Gaussian process models. This enables the application of Gaussian process (GP) models to data sets containing millions of data points. We show how GPs can be vari- ationally decomposed to depend on a set of globally relevant inducing variables which factorize the model in the necessary manner to perform variational inference. Our ap- proach is readily extended to models with non-Gaussian likelihoods and latent variable models based around Gaussian processes. We demonstrate the approach on a simple toy problem and two real world data sets.

Table of Contents

  • 1 Introduction
  • 2 Sparse GPs Revisited
  • 3 SVI for GPs
  • 3.1 Global Variables
  • 3.2 Natural Gradients
  • 3.3 Latent Variables
  • 3.4 Non-Gaussian likelihoods
  • 4 Experiments
  • 4.1 Toy Data
  • 4.2 UK Apartment Price Data
  • 4.3 Airline Delays
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — Factorized Variational Lower Bound for Gaussian Process Regression

    equation

    Let y=(y1,…,yn)⊤∈Rny = (y_1, \dots, y_n)^\top \in \mathbb{R}^n be observations at input locations X={xi}i=1n⊂RdX = \{x_i\}_{i=1}^n \subset \mathbb{R}^d under an independent Gaussian observation model p(y∣f)=N(y∣f,β−1I)p(y \mid f) = \mathcal{N}(y \mid f, \beta^{-1} I) with noise precision eta > 0. Let u∈Rmu \in \mathbb{R}^m denote latent function values at inducing locations Z={zj}j=1m⊂RdZ = \{z_j\}_{j=1}^m \subset \mathbb{R}^d, with prior p(u)=N(u∣0,Kmm)p(u) = \mathcal{N}(u \mid 0, K_{mm}).

    By introducing an explicit Gaussian variational distribution q(u)=N(u∣mu,S)q(u) = \mathcal{N}(u \mid m_u, S) over the inducing variables, the log marginal likelihood log⁡p(y∣X)\log p(y \mid X) is lower-bounded by:

    L3=∑i=1n[log⁡N(yi  |  ki⊤Kmm−1mu,β−1)−12βk~i,i−12tr(SΛi)]−KL(q(u)∥p(u))\mathcal{L}_3 = \sum_{i=1}^n \left[ \log \mathcal{N}\left(y_i \;\middle|\; k_i^\top K_{mm}^{-1} m_u, \beta^{-1}\right) - \frac{1}{2} \beta \widetilde{k}_{i,i} - \frac{1}{2} \text{tr}(S \Lambda_i) \right] - \text{KL}(q(u) \parallel p(u))

    where:

    • Kmm∈Rm×mK_{mm} \in \mathbb{R}^{m \times m} is the covariance matrix evaluated across all inducing inputs ZZ.
    • ki∈Rmk_i \in \mathbb{R}^m is the covariance vector between inducing points ZZ and the data point xix_i, corresponding to the ii-th column of Kmn=Knm⊤K_{mn} = K_{nm}^\top.
    • k~i,i=k(xi,xi)−ki⊤Kmm−1ki\widetilde{k}_{i,i} = k(x_i, x_i) - k_i^\top K_{mm}^{-1} k_i is the ii-th diagonal entry of K~=Knn−KnmKmm−1Kmn\widetilde{K} = K_{nn} - K_{nm} K_{mm}^{-1} K_{mn}.
    • Λi=βKmm−1kiki⊤Kmm−1∈Rm×m\Lambda_i = \beta K_{mm}^{-1} k_i k_i^\top K_{mm}^{-1} \in \mathbb{R}^{m \times m}.
    • KL(q(u)∥p(u))\text{KL}(q(u) \parallel p(u)) is the Kullback-Leibler divergence between q(u)q(u) and p(u)p(u).

    Because the likelihood terms and trace penalties decompose as a sum over individual data points i=1,…,ni = 1, \dots, n, the lower bound L3\mathcal{L}_3 enables unbiased stochastic estimation and optimization via mini-batches.

  2. Knowl 2 — Natural Gradient Variational Updates for Inducing Distribution in Gaussian Processes

    theoretical result

    For a Gaussian variational distribution q(u)=N(u∣mu,S)q(u) = \mathcal{N}(u \mid m_u, S) on inducing variables u∈Rmu \in \mathbb{R}^m with canonical parameters θ1=S−1mu\theta_1 = S^{-1} m_u and θ2=−12S−1\theta_2 = -\frac{1}{2} S^{-1}, and expectation parameters η1=mu\eta_1 = m_u and η2=mumu⊤+S\eta_2 = m_u m_u^\top + S, the natural gradient of the factorized lower bound L3\mathcal{L}_3 satisfies:

    g~(θ)=G(θ)−1∂L3∂θ=∂L3∂η\widetilde{g}(\theta) = G(\theta)^{-1} \frac{\partial \mathcal{L}_3}{\partial \theta} = \frac{\partial \mathcal{L}_3}{\partial \eta}

    where G(θ)G(\theta) is the Fisher information matrix. A natural gradient step with learning rate ℓ∈(0,1]\ell \in (0, 1] yields the updates:

    θ2(t+1)=−12(S(t+1))−1=−12(S(t))−1+ℓ(−12Λ+12(S(t))−1)\theta_2^{(t+1)} = -\frac{1}{2}\left(S^{(t+1)}\right)^{-1} = -\frac{1}{2}\left(S^{(t)}\right)^{-1} + \ell \left( -\frac{1}{2}\Lambda + \frac{1}{2}\left(S^{(t)}\right)^{-1} \right)

    θ1(t+1)=(S(t+1))−1mu(t+1)=(S(t))−1mu(t)+ℓ(βKmm−1Kmny−(S(t))−1mu(t))\theta_1^{(t+1)} = \left(S^{(t+1)}\right)^{-1} m_u^{(t+1)} = \left(S^{(t)}\right)^{-1} m_u^{(t)} + \ell \left( \beta K_{mm}^{-1} K_{mn} y - \left(S^{(t)}\right)^{-1} m_u^{(t)} \right)

    where Λ=Kmm−1+∑i=1nΛi=Kmm−1+βKmm−1KmnKnmKmm−1\Lambda = K_{mm}^{-1} + \sum_{i=1}^n \Lambda_i = K_{mm}^{-1} + \beta K_{mm}^{-1} K_{mn} K_{nm} K_{mm}^{-1}.

    This update has two key mathematical properties:

    1. Setting step size ℓ=1\ell = 1 in batch mode exactly recovers the optimal collapsed variational solution S=Λ−1S = \Lambda^{-1} and mu=βΛ−1Kmm−1Kmnym_u = \beta \Lambda^{-1} K_{mm}^{-1} K_{mn} y.
    2. The natural gradient direction for θ2\theta_2 is strictly positive definite, guaranteeing that S(t+1)S^{(t+1)} remains positive definite for any ℓ∈(0,1]\ell \in (0, 1] without requiring constrained matrix parameterizations.
  3. Knowl 3 — Stochastic Variational Inference for Gaussian Processes

    algorithm

    The Stochastic Variational Inference for Gaussian Processes (SVIGP) algorithm optimizes the explicit variational distribution q(u)=N(u∣mu,S)q(u) = \mathcal{N}(u \mid m_u, S) over mm inducing variables uu using natural gradient descent on mini-batches of size ∣B∣|B|, while optimizing kernel hyperparameters, observation noise precision eta, and inducing inputs ZZ using standard stochastic gradient descent.

    Per iteration, the computational complexity is O(∣B∣m2+m3)\mathcal{O}(|B|m^2 + m^3) and working storage is O(∣B∣m+m2)\mathcal{O}(|B|m + m^2), both independent of the full dataset size nn.

    Input: Training inputs X={xi}i=1nX = \{x_i\}_{i=1}^n, targets y={yi}i=1ny = \{y_i\}_{i=1}^n, inducing points Z={zj}j=1mZ = \{z_j\}_{j=1}^m, batch size ∣B∣|B|, learning rates ℓu\ell_u and ℓθ\ell_\theta, momentum μ\mu, max iterations TT
    Output: Variational parameters mu,Sm_u, S, inducing points ZZ, kernel hyperparameters θk\theta_k, noise precision β\beta
    Initialize ZZ using k-means on XX
    Initialize q(u)=N(mu,S)q(u) = \mathcal{N}(m_u, S) with canonical parameters θ1=S−1mu,θ2=−12S−1\theta_1 = S^{-1}m_u, \theta_2 = -\frac{1}{2}S^{-1}
    Initialize kernel hyperparameters θk\theta_k, noise precision β\beta, and velocity v←0v \leftarrow 0
    for t=1t = 1 to TT do
        Sample mini-batch indices B⊂{1,…,n}B \subset \{1, \dots, n\} of size ∣B∣|B|
        Compute KmmK_{mm} and submatrix KBmK_{B m}
        Compute batch statistics ΛB=n∣B∣βKmm−1KmBKBmKmm−1\Lambda_B = \frac{n}{|B|} \beta K_{mm}^{-1} K_{m B} K_{B m} K_{mm}^{-1}
        Compute batch natural gradient target Λ^=Kmm−1+ΛB\widehat{\Lambda} = K_{mm}^{-1} + \Lambda_B
        Compute batch target h^=n∣B∣βKmm−1KmByB\widehat{h} = \frac{n}{|B|} \beta K_{mm}^{-1} K_{m B} y_B
        Update canonical parameters via natural gradient step:
            θ2←(1−ℓu)θ2+ℓu(−12Λ^)\theta_2 \leftarrow (1 - \ell_u) \theta_2 + \ell_u \left( -\frac{1}{2} \widehat{\Lambda} \right)
            θ1←(1−ℓu)θ1+ℓuh^\theta_1 \leftarrow (1 - \ell_u) \theta_1 + \ell_u \widehat{h}
        Update S←−12θ2−1S \leftarrow -\frac{1}{2} \theta_2^{-1} and mu←Sθ1m_u \leftarrow S \theta_1
        Compute stochastic gradient gθ=∇{θk,β,Z}L^3(B)g_\theta = \nabla_{\{\theta_k, \beta, Z\}} \widehat{\mathcal{L}}_3(B)
        Update hyperparameter velocity v←μv+ℓθgθv \leftarrow \mu v + \ell_\theta g_\theta
        Update hyperparameters {θk,β,Z}←{θk,β,Z}+v\{\theta_k, \beta, Z\} \leftarrow \{\theta_k, \beta, Z\} + v
    end for
    return mu,S,Z,θk,βm_u, S, Z, \theta_k, \beta
  4. Knowl 4 — Computational and Storage Complexity of SVIGP Relative to Exact and Sparse GPs

    theoretical result

    The computational and storage complexities of Gaussian process regression scale across formulations as follows:

    • Exact GP: Computational time complexity is O(n3)\mathcal{O}(n^3) (or O(n3p3)\mathcal{O}(n^3 p^3) for pp outputs/tasks); storage requirement is O(n2)\mathcal{O}(n^2) (or O(n2p2)\mathcal{O}(n^2 p^2)).
    • Collapsed Sparse GP: Computational time complexity is O(nm2)\mathcal{O}(n m^2); storage requirement is O(nm)\mathcal{O}(n m), where mm is the number of inducing variables.
    • Stochastic Variational GP (SVIGP): Per-iteration computational time complexity is O(∣B∣m2+m3)\mathcal{O}(|B| m^2 + m^3); storage requirement is O(∣B∣m+m2)\mathcal{O}(|B| m + m^2), where ∣B∣≪n|B| \ll n is the mini-batch size.

    Because the per-iteration time and memory of SVIGP are independent of the total number of data points nn, the model can scale to arbitrarily large datasets and allows setting significantly larger numbers of inducing points mm (e.g., m=800m = 800 to 10001000, whereas collapsed sparse GP models are computationally constrained to m≤100m \le 100 when nn is large).

  5. Knowl 5 — Extension of Stochastic Variational Gaussian Processes to Non-Gaussian Likelihoods

    model/method

    For observation vectors t=(t1,…,tn)⊤t = (t_1, \dots, t_n)^\top with non-Gaussian conditionally independent likelihoods p(t∣y)=∏i=1np(ti∣yi)p(t \mid y) = \prod_{i=1}^n p(t_i \mid y_i) (such as binary classification with ti∈{0,1}t_i \in \{0, 1\} and Bernoulli likelihood σ(yi)ti(1−σ(yi))1−ti\sigma(y_i)^{t_i}(1-\sigma(y_i))^{1-t_i}), the factorized variational lower bound L3\mathcal{L}_3 enables exact one-dimensional marginalization of the latent vector yy:

    p(t∣X)≥∫p(t∣y)exp⁡{L3}dy=∏i=1n∫p(ti∣yi)N(yi  |  ki⊤Kmm−1mu,β−1)exp⁡(−12βk~i,i−12tr(SΛi))dyip(t \mid X) \ge \int p(t \mid y) \exp\{\mathcal{L}_3\} dy = \prod_{i=1}^n \int p(t_i \mid y_i) \mathcal{N}\left(y_i \;\middle|\; k_i^\top K_{mm}^{-1} m_u, \beta^{-1}\right) \exp\left(-\frac{1}{2}\beta \widetilde{k}_{i,i} - \frac{1}{2}\text{tr}(S \Lambda_i)\right) dy_i

    For probit likelihoods where p(ti=1∣yi)=Φ(yi)p(t_i = 1 \mid y_i) = \Phi(y_i) (with Φ(⋅)\Phi(\cdot) being the cumulative Gaussian distribution), each one-dimensional integral is analytically tractable. This allows stochastic variational inference via natural gradient descent to be applied directly to GP classification without relying on expectation propagation or local variational bounds on latent function values ff.

  6. Knowl 6 — Stochastic Variational Inference for Bayesian Gaussian Process Latent Variable Models

    model/method

    In Gaussian Process Latent Variable Models (GPLVM), the input coordinates X∈Rn×dX \in \mathbb{R}^{n \times d} are unobserved latent variables with prior p(X)p(X), and observations Y∈Rn×pY \in \mathbb{R}^{n \times p} consist of pp output dimensions. Using a fully factorized Gaussian variational distribution q(X)=∏i=1nq(xi)q(X) = \prod_{i=1}^n q(x_i) together with the factorized lower bound L3\mathcal{L}_3, the marginal log-likelihood log⁡p(Y)\log p(Y) is lower-bounded by:

    log⁡p(Y)=log⁡∫p(Y∣X)p(X)dX≥∫q(X)[L3(X,Y)+log⁡p(X)−log⁡q(X)]dX\log p(Y) = \log \int p(Y \mid X) p(X) dX \ge \int q(X) \left[ \mathcal{L}_3(X, Y) + \log p(X) - \log q(X) \right] dX

    Because L3\mathcal{L}_3 factorizes over data instances ii, the latent distributions q(xi)q(x_i) for instances within a sampled mini-batch can be updated independently and in parallel while keeping q(u)q(u) fixed, after which the global inducing distribution q(u)q(u) is updated via approximate natural gradients.

  7. Knowl 7 — Empirical Evaluation on UK Apartment Price Data

    data/table

    The SVIGP model was evaluated on a dataset of 75,000 monthly apartment transactions in England and Wales (February to October 2012), regressing normalized log prices onto latitude and longitude. A random sample of 10,000 observations was held out as a test set. The SVIGP model used m=800m = 800 inducing points initialized via k-means, a mini-batch size of 1000, a learning rate of 0.01 for the variational parameters of q(u)q(u), and a learning rate of 1×10−51 \times 10^{-5} with momentum 0.9 for covariance parameters. The covariance function comprised two squared exponential terms with different lengthscales, a constant bias term, and Gaussian noise. SVIGP was compared against Gaussian processes trained on random training subsets of sizes N∈{500,800,1000,1200}N \in \{500, 800, 1000, 1200\} optimized using type-II maximum likelihood.

    Method Mean Squared Error
    SVIGP (m=800m=800) 0.426
    Random sub-set (N=500N=500) 0.522 ±\pm 0.018
    Random sub-set (N=800N=800) 0.510 ±\pm 0.015
    Random sub-set (N=1000N=1000) 0.503 ±\pm 0.011
    Random sub-set (N=1200N=1200) 0.502 ±\pm 1.012

    The table reports mean squared error (MSE) on the held-out test set, with ±2\pm 2 standard deviations of inter-subset variability for the random subset baseline. SVIGP achieves a substantially lower MSE (0.426) by utilizing all available training data through mini-batch stochastic variational optimization.

  8. Knowl 8 — Empirical Evaluation on US Commercial Airline Delays Dataset

    empirical result

    SVIGP was evaluated on flight arrival and departure delay prediction using records of commercial flights in the USA from January to April 2008. From nearly 2 million flights, 800,000 datapoints were sampled (700,000 for training, 100,000 for testing) with 8 input features: aircraft age, distance, airtime, departure time, arrival time, day of week, day of month, and month.

    The GP model used a squared exponential covariance function with Automatic Relevance Determination (ARD; separate lengthscale per feature), a bias term, and a noise term, trained with m=1000m = 1000 inducing points and mini-batch size 5000 (learning rates: 0.01 for q(u)q(u), 1×10−51 \times 10^{-5} with momentum 0.9 for covariance parameters).

    Key empirical findings:

    1. Prediction Error: Across 10 runs, SVIGP converged to an average root mean squared error (RMSE) of approximately 32.7 minutes, whereas full GPs trained on random data subsets of sizes N∈{800,1000,1200}N \in \{800, 1000, 1200\} achieved average RMSEs between 35.0 and 35.8 minutes.
    2. Effect of Inducing Set Size: Increasing the number of inducing inputs mm from 50 to 100, 200, 500, 800, and 1000 produced monotonic decreases in test RMSE from ~37 minutes down to ~32.7 minutes.
    3. Feature Relevance: ARD inverse lengthscales identified departure time as the most predictive feature of delay, followed closely by flight distance, while airtime showed lower relevance.

Coverage note — Omitted qualitative convergence details on the 1D and 2D synthetic sinusoidal toy datasets as they served only as pedagogical visualizations of properties fully captured by the methodological and large-scale experimental knowls.

References

  1. 1.Mauricio A. Álvarez and Neil D. Lawrence. Computationally efficient convolved multiple output Gaussian processes. Journal of Machine Learning Research, 12:1425–1466, May 2011.
  2. 2.Lehel Csató and Manfred Opper. Sparse on-line Gaussian processes. Neural Computation, 14(3):641–668, 2002.
  3. 3.Andreas Damianou, Michalis K. Titsias, and Neil D. Lawrence. Variational Gaussian process dynamical systems. In Peter Bartlett, Fernando Peirrera, Chris Williams, and John Lafferty, editors, Advances in Neural Information Processing Systems, volume 24, Cambridge, MA, 2011. MIT Press.
  4. 4.Andreas Damianou, Carl Henrik Ek, Michalis K. Titsias, and Neil D. Lawrence. Manifold relevance determination. In John Langford and Joelle Pineau, editors, Proceedings of the International Conference in Machine Learning, volume 29, San Francisco, CA, 2012. Morgan Kauffman. To appear.
  5. 5.Mark N. Gibbs and David J. C. MacKay. Variational Gaussian process classifiers. IEEE Transactions on Neural Networks, 11(6):1458–1464, 2000.
  6. 6.James Hensman, Magnus Rattray, and Neil D. Lawrence. Fast variational inference in the exponential family. NIPS 2012, 2012.
  7. 7.Matthew Hoffman, David M. Blei, Chong Wang, and John Paisley. Stochastic variational inference. arXiv preprint arXiv:1206.7051, 2012.
  8. 8.Malte Kuss and Carl Edward Rasmussen. Assessing approximate inference for binary Gaussian process classification. Journal of Machine Learning Research, 6:1679–1704, 2005.
  9. 9.Neil D. Lawrence. Probabilistic non-linear principal component analysis with Gaussian process latent variable models. Journal of Machine Learning Research, 6:1783–1816, 11 2005.
  10. 10.Joaquin Quiñonero Candela and Carl Edward Rasmussen. A unifying view of sparse approximate Gaussian process regression. Journal of Machine Learning Research, 6:1939–1959, 2005.
  11. 11.Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, Cambridge, MA, 2006. ISBN 0-262-18253-X.
  12. 12.Matthias Seeger, Christopher K. I. Williams, and Neil D. Lawrence. Fast forward selection to speed up sparse Gaussian process regression. In Christopher M. Bishop and Brendan J. Frey, editors, Proceedings of the Ninth International Workshop on Artificial Intelligence and Statistics, Key West, FL, 3–6 Jan 2003.
  13. 13.Edward Snelson and Zoubin Ghahramani. Local and global sparse Gaussian process approximations. In Marina Meila and Xiaotong Shen, editors, Proceedings of the Eleventh International Workshop on Artificial Intelligence and Statistics, San Juan, Puerto Rico, 21-24 March 2007. Omnipress.
  14. 14.Michalis K. Titsias. Variational learning of inducing variables in sparse Gaussian processes. In David van Dyk and Max Welling, editors, Proceedings of the Twelfth International Workshop on Artificial Intelligence and Statistics, volume 5, pages 567–574, Clearwater Beach, FL, 16-18 April 2009. JMLR W&CP 5.
  15. 15.Michalis K. Titsias and Neil D. Lawrence. Bayesian Gaussian process latent variable model. In Yee Whye Teh and D. Michael Titterington, editors, Proceedings of the Thirteenth International Workshop on Artificial Intelligence and Statistics, volume 9, pages 844–851, Chia Laguna Resort, Sardinia, Italy, 13-16 May 2010. JMLR W&CP 9.
  16. 16.Raquel Urtasun and Trevor Darrell. Local probabilistic regression for activity-independent human pose inference. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Anchorage, Alaska, 2008.

Citation

MLA
Hensman, J., et al. “Gaussian Processes for Big Data”. arXiv, 2013, http://arxiv.org/abs/1309.6835v1.
APA
Hensman, J., Fusi, N., & Lawrence, N. D. (2013). Gaussian Processes for Big Data. arXiv. http://arxiv.org/abs/1309.6835v1
Chicago
Hensman, J., N. Fusi, and N. D. Lawrence. 2013. “Gaussian Processes for Big Data”. arXiv. http://arxiv.org/abs/1309.6835v1.
Harvard
Hensman, J., Fusi, N. and Lawrence, N.D. (2013) “Gaussian Processes for Big Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1309.6835v1.
Vancouver
1. Hensman J, Fusi N, Lawrence ND (2013) Gaussian Processes for Big Data. arXiv

BibTeX

@article{hensman2013gaussian,
  title = {Gaussian Processes for Big Data},
  author = {Hensman, James and Fusi, Nicolo and Lawrence, Neil D.},
  year = {2013},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1309.6835v1},
  eprint = {1309.6835}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/