Sparse Gaussian Processes using Pseudo-inputs

Edward SnelsonZoubin Ghahramani

article2005NeurIPS2,065 citations

Proposes a sparse Gaussian process framework that optimizes continuous pseudo-input locations jointly with kernel hyperparameters, matching full Gaussian process regression accuracy on large datasets at an efficient O(M²N) computational cost.

Listen

Gaussian process regression is a powerful, flexible method for non-linear predictive modeling, but its standard implementation is computationally impractical for large datasets. Training costs grow cubically with the number of training points, and evaluating new predictions scales with the full size of the data. Existing sparse approximations attempt to reduce this burden by selecting a small subset of actual training points as an active set. However, these methods suffer from unstable hyperparameter learning because discrete subset selections cause non-smooth fluctuations during optimization, degrading overall model performance.

The article evaluates a new framework called Sparse Pseudo-input Gaussian Processes, which parameterizes the model covariance using a small number of continuous pseudo-inputs that are not restricted to actual data locations. The authors demonstrate how pseudo-input coordinates and covariance hyperparameters can be optimized jointly in a single, smooth gradient-based procedure, achieving high predictive accuracy with very small active set sizes.

The approach derives a marginal likelihood formulation containing an input-dependent variance correction that provides strong gradients to guide pseudo-input placement. The authors tested this method on synthetic one-dimensional tasks with input-dependent noise and on two benchmark datasets containing tens of thousands of samples across various input dimensions. Performance was compared directly against full Gaussian processes and established sparse approximation techniques, including random subset selection, information-gain greedy selection, and Smola-Bartlett selection.

The results establish that the pseudo-input method significantly outperforms competing sparse techniques when using very small active sets. On benchmark datasets, the proposed method approached the accuracy of a full Gaussian process using only a small fraction of points, such as 25 pseudo-inputs on a multi-dimensional robotic dataset where standard sparse baselines required far larger sets. The experiments also confirmed that the model's distinct variance term is essential for continuous optimization, as earlier projected latent variable formulations caused gradient ascent to stall. Additionally, the added flexibility of continuous pseudo-inputs allowed the model to effectively capture non-stationary effects and varying noise levels without requiring explicit non-stationary covariance structures.

These findings mean organizations can deploy highly accurate Gaussian process models in large-scale settings with significantly lower computational overhead and faster prediction speeds at deployment. Because prediction time depends only quadratically on the small number of pseudo-inputs rather than the training set size, the method lowers inference latency and operational infrastructure costs. While training requires optimizing across a larger continuous parameter space, this upfront computational cost is offset by avoiding the instability, poor generalization, and secondary tuning steps required by competing sparse techniques.

Decision-makers seeking rapid, compact predictive models should adopt this pseudo-input approach, particularly where very sparse representations are desired for real-time execution. To improve performance in higher-dimensional tasks or with larger pseudo-input budgets, developers should implement advanced optimization strategies such as block-coordinate optimization, stochastic gradient updates, or dimensionality-reduction projections. Further research is warranted to extend this formulation to classification tasks and more complex non-stationary environments.

Confidence in the core methodology is strong for compact active set sizes, supported by consistent empirical improvements across benchmark evaluations. However, practitioners should exercise caution when scaling to very large pseudo-input sets or high-dimensional problems without proper initialization, as the joint optimization surface contains local optima and can occasionally overfit or underestimate noise levels.

Snelson et al (2005).pdf
  • Paper: Variational Learning of Inducing Variables in Sparse Gaussian Processes, Michalis K. Titsias (2009). Titsias turns optimized inducing points into variational parameters, extending sparse GP learning with a principled bound that addresses overfitting in earlier approximations.
  • Paper: Gaussian Processes for Big Data, James Hensman et al. (2013). This work extends inducing-variable GP regression with stochastic variational inference, making training scale to datasets far larger than the pseudo-input approach demonstrated.
  • Paper: Deep Gaussian Processes, Andreas C. Damianou et al. (2012). It carries pseudo-input approximations into deep Gaussian processes, using them to make variational inference over layered latent functions tractable.
Cover for Sparse Gaussian Processes using Pseudo-inputs

Abstract

We present a new Gaussian process (GP) regression model whose covariance is parameterized by the the locations of M pseudo-input points, which we learn by a gradient based optimization. We take M ≪ N, where N is the number of real data points, and hence obtain a sparse regression method which has O(M²N) training cost and O(M²) prediction cost per test case. We also find hyperparameters of the covariance function in the same joint optimization. The method can be viewed as a Bayesian regression model with particular input dependent noise. The method turns out to be closely related to several other sparse GP approaches, and we discuss the relation in detail. We finally demonstrate its performance on some large data sets, and make a direct comparison to other sparse GP methods. We show that our method can match full GP performance with small M, i.e. very sparse solutions, and it significantly outperforms other approaches in this regime.

Table of Contents

  • 1 Introduction
  • 1.1 Gaussian processes for regression
  • 2 Sparse Pseudo-input Gaussian processes (SPGPs)
  • 3 Relation to other methods
  • 4 Experiments
  • 5 Conclusions, extensions and future work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Sparse Pseudo-input Gaussian Process Model Formulation

    model/method

    The Sparse Pseudo-input Gaussian Process (SPGP) model defines a sparse Bayesian regression framework for a dataset D={(xn,yn)}n=1N\mathcal{D} = \{(x_n, y_n)\}_{n=1}^N with inputs X=[x1,…,xN]⊤∈RN×DX = [x_1, \dots, x_N]^\top \in \mathbb{R}^{N \times D} and targets y=[y1,…,yN]⊤∈RN\mathbf{y} = [y_1, \dots, y_N]^\top \in \mathbb{R}^N. Instead of selecting a discrete active subset from the training inputs, SPGP introduces M≪NM \ll N unconstrained continuous pseudo-inputs Xˉ=[xˉ1,…,xˉM]⊤∈RM×D\bar{X} = [\bar{x}_1, \dots, \bar{x}_M]^\top \in \mathbb{R}^{M \times D} and associated latent pseudo-targets fˉ=[fˉ1,…,fˉM]⊤∈RM\bar{\mathbf{f}} = [\bar{f}_1, \dots, \bar{f}_M]^\top \in \mathbb{R}^M.

    The pseudo-targets are governed by a zero-mean Gaussian process prior: p(fˉ∣Xˉ)=N(fˉ∣0,KM)p(\bar{\mathbf{f}} \mid \bar{X}) = \mathcal{N}(\bar{\mathbf{f}} \mid \mathbf{0}, K_M) where [KM]mm′=K(xˉm,xˉm′)[K_M]_{mm'} = K(\bar{x}_m, \bar{x}_{m'}) is the M×MM \times M covariance matrix computed using a covariance function K(⋅,⋅)K(\cdot, \cdot).

    Given pseudo-inputs Xˉ\bar{X} and pseudo-targets fˉ\bar{\mathbf{f}}, the observed targets y\mathbf{y} are assumed conditionally independent across data points with total likelihood: p(y∣X,Xˉ,fˉ)=∏n=1Np(yn∣xn,Xˉ,fˉ)=N(y  |  KNMKM−1fˉ, Λ+σ2IN)p(\mathbf{y} \mid X, \bar{X}, \bar{\mathbf{f}}) = \prod_{n=1}^N p(y_n \mid x_n, \bar{X}, \bar{\mathbf{f}}) = \mathcal{N}\left(\mathbf{y} \;\middle|\; K_{NM} K_M^{-1} \bar{\mathbf{f}}, \, \Lambda + \sigma^2 I_N\right) where [KNM]nm=K(xn,xˉm)[K_{NM}]_{nm} = K(x_n, \bar{x}_m) is an N×MN \times M cross-covariance matrix, σ2\sigma^2 is the observation noise variance, INI_N is the N×NN \times N identity matrix, and Λ=diag(λ1,…,λN)\Lambda = \text{diag}(\lambda_1, \dots, \lambda_N) is a diagonal matrix with elements: λn=K(xn,xn)−kn⊤KM−1kn\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_M^{-1} \mathbf{k}_n where kn=[K(xˉ1,xn),…,K(xˉM,xn)]⊤\mathbf{k}_n = [K(\bar{x}_1, x_n), \dots, K(\bar{x}_M, x_n)]^\top.

  2. Knowl 2 — SPGP Marginal Likelihood Objective for Pseudo-Input and Hyperparameter Learning

    equation

    In the Sparse Pseudo-input Gaussian Process (SPGP) regression model, integrating out the latent pseudo-targets fˉ∈RM\bar{\mathbf{f}} \in \mathbb{R}^M over the prior p(fˉ∣Xˉ)=N(fˉ∣0,KM)p(\bar{\mathbf{f}} \mid \bar{X}) = \mathcal{N}(\bar{\mathbf{f}} \mid \mathbf{0}, K_M) and likelihood p(y∣X,Xˉ,fˉ)=N(y∣KNMKM−1fˉ,Λ+σ2IN)p(\mathbf{y} \mid X, \bar{X}, \bar{\mathbf{f}}) = \mathcal{N}(\mathbf{y} \mid K_{NM} K_M^{-1} \bar{\mathbf{f}}, \Lambda + \sigma^2 I_N) yields the exact marginal likelihood of observed targets y∈RN\mathbf{y} \in \mathbb{R}^N given inputs X∈RN×DX \in \mathbb{R}^{N \times D}, pseudo-inputs Xˉ∈RM×D\bar{X} \in \mathbb{R}^{M \times D}, and hyperparameters Θ={θ,σ2}\Theta = \{\theta, \sigma^2\}:

    p(y∣X,Xˉ,Θ)=∫p(y∣X,Xˉ,fˉ)p(fˉ∣Xˉ) dfˉ=N(y  |  0, KNMKM−1KMN+Λ+σ2IN)p(\mathbf{y} \mid X, \bar{X}, \Theta) = \int p(\mathbf{y} \mid X, \bar{X}, \bar{\mathbf{f}}) p(\bar{\mathbf{f}} \mid \bar{X}) \, d\bar{\mathbf{f}} = \mathcal{N}\left(\mathbf{y} \;\middle|\; \mathbf{0}, \, K_{NM} K_M^{-1} K_{MN} + \Lambda + \sigma^2 I_N\right)

    where:

    • [KM]mm′=K(xˉm,xˉm′)[K_M]_{mm'} = K(\bar{x}_m, \bar{x}_{m'}) for m,m′∈{1,…,M}m, m' \in \{1, \dots, M\},
    • [KNM]nm=K(xn,xˉm)[K_{NM}]_{nm} = K(x_n, \bar{x}_m) for n∈{1,…,N}n \in \{1, \dots, N\} and m∈{1,…,M}m \in \{1, \dots, M\}, with KMN=KNM⊤K_{MN} = K_{NM}^\top,
    • Λ=diag(λ1,…,λN)\Lambda = \text{diag}(\lambda_1, \dots, \lambda_N) is a diagonal matrix with λn=K(xn,xn)−kn⊤KM−1kn\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_M^{-1} \mathbf{k}_n, and kn=[K(xˉ1,xn),…,K(xˉM,xn)]⊤\mathbf{k}_n = [K(\bar{x}_1, x_n), \dots, K(\bar{x}_M, x_n)]^\top,
    • σ2\sigma^2 is the observation noise variance, and INI_N is the N×NN \times N identity matrix.

    Parameters {Xˉ,Θ}\{\bar{X}, \Theta\} (totaling M⋅D+∣Θ∣M \cdot D + |\Theta| parameters) are optimized jointly by maximizing log⁡p(y∣X,Xˉ,Θ)\log p(\mathbf{y} \mid X, \bar{X}, \Theta) via gradient ascent.

  3. Knowl 3 — SPGP Posterior Distribution and Predictive Equations

    equation

    For a Sparse Pseudo-input Gaussian Process (SPGP) conditioned on training data D=(X,y)\mathcal{D} = (X, \mathbf{y}) and pseudo-inputs Xˉ={xˉm}m=1M\bar{X} = \{\bar{x}_m\}_{m=1}^M, the posterior distribution over the pseudo-targets fˉ∈RM\bar{\mathbf{f}} \in \mathbb{R}^M is given by: p(fˉ∣D,Xˉ)=N(fˉ  |  KMQM−1KMN(Λ+σ2IN)−1y, KMQM−1KM)p(\bar{\mathbf{f}} \mid \mathcal{D}, \bar{X}) = \mathcal{N}\left(\bar{\mathbf{f}} \;\middle|\; K_M Q_M^{-1} K_{MN} (\Lambda + \sigma^2 I_N)^{-1} \mathbf{y}, \, K_M Q_M^{-1} K_M\right) where QM=KM+KMN(Λ+σ2IN)−1KNMQ_M = K_M + K_{MN} (\Lambda + \sigma^2 I_N)^{-1} K_{NM}.

    For a new test input x∗∈RDx_* \in \mathbb{R}^D, marginalizing fˉ\bar{\mathbf{f}} yields the Gaussian predictive distribution: p(y∗∣x∗,D,Xˉ)=∫p(y∗∣x∗,Xˉ,fˉ)p(fˉ∣D,Xˉ) dfˉ=N(y∗∣μ∗,σ∗2)p(y_* \mid x_*, \mathcal{D}, \bar{X}) = \int p(y_* \mid x_*, \bar{X}, \bar{\mathbf{f}}) p(\bar{\mathbf{f}} \mid \mathcal{D}, \bar{X}) \, d\bar{\mathbf{f}} = \mathcal{N}(y_* \mid \mu_*, \sigma_*^2) where the predictive mean μ∗\mu_* and predictive variance σ∗2\sigma_*^2 are: μ∗=k∗⊤QM−1KMN(Λ+σ2IN)−1y\mu_* = \mathbf{k}_*^\top Q_M^{-1} K_{MN} (\Lambda + \sigma^2 I_N)^{-1} \mathbf{y} σ∗2=K(x∗,x∗)−k∗⊤(KM−1−QM−1)k∗+σ2\sigma_*^2 = K(x_*, x_*) - \mathbf{k}_*^\top (K_M^{-1} - Q_M^{-1}) \mathbf{k}_* + \sigma^2 with k∗=[K(xˉ1,x∗),…,K(xˉM,x∗)]⊤\mathbf{k}_* = [K(\bar{x}_1, x_*), \dots, K(\bar{x}_M, x_*)]^\top, [KM]mm′=K(xˉm,xˉm′)[K_M]_{mm'} = K(\bar{x}_m, \bar{x}_{m'}), [KNM]nm=K(xn,xˉm)[K_{NM}]_{nm} = K(x_n, \bar{x}_m), and Λ=diag(λ1,…,λN)\Lambda = \text{diag}(\lambda_1, \dots, \lambda_N) where λn=K(xn,xn)−kn⊤KM−1kn\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_M^{-1} \mathbf{k}_n.

  4. Knowl 4 — Gradient Guidance via the Diagonal Variance Matrix Lambda vs Projected Latent Variables

    theoretical result

    In the projected latent variables (PLV) approximation, the marginal likelihood covariance is KNMKM−1KMN+σ2INK_{NM} K_M^{-1} K_{MN} + \sigma^2 I_N, which lacks the diagonal matrix Λ=diag({λn}n=1N)\Lambda = \text{diag}(\{\lambda_n\}_{n=1}^N) where λn=K(xn,xn)−kn⊤KM−1kn\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_M^{-1} \mathbf{k}_n. Consequently, PLV assumes a constant observation noise σ2\sigma^2 across the entire input domain. When pseudo-inputs are initialized in a localized region, moving a pseudo-input toward an uncovered data point does not sufficiently alter the model fit locally to produce a significant gradient, causing continuous gradient-based optimization of pseudo-inputs in PLV to become trapped.

    In contrast, the SPGP marginal likelihood includes Λ\Lambda, resulting in an effective input-dependent noise variance λn+σ2\lambda_n + \sigma^2 for each training point xnx_n. In regions far from any pseudo-input, λn≈K(xn,xn)\lambda_n \approx K(x_n, x_n), producing a large noise variance K(xn,xn)+σ2K(x_n, x_n) + \sigma^2. Moving a pseudo-input toward an uncovered training point immediately reduces λn\lambda_n toward zero, generating a strong gradient that repels pseudo-inputs from clustered configurations and drives them to cover the support of the training data.

  5. Knowl 5 — Computational Complexity of SPGP Training and Inference

    theoretical result

    Let NN denote the number of training points, MM the number of pseudo-inputs with M≪NM \ll N, and DD the input dimension.

    • Training: Because the matrix Λ+σ2IN\Lambda + \sigma^2 I_N is diagonal, its inversion requires O(N)\mathcal{O}(N) operations. The matrix product KMN(Λ+σ2IN)−1KNMK_{MN} (\Lambda + \sigma^2 I_N)^{-1} K_{NM} used to compute QM∈RM×MQ_M \in \mathbb{R}^{M \times M} requires O(M2N)\mathcal{O}(M^2 N) operations. Inversion and Cholesky factorization of the M×MM \times M matrices KMK_M and QMQ_M require O(M3)\mathcal{O}(M^3) operations. Thus, each gradient evaluation during training scales as O(M2N)\mathcal{O}(M^2 N).
    • Inference: Precomputing the vector QM−1KMN(Λ+σ2IN)−1y∈RMQ_M^{-1} K_{MN} (\Lambda + \sigma^2 I_N)^{-1} \mathbf{y} \in \mathbb{R}^M requires O(M)\mathcal{O}(M) additional operations after QMQ_M is factored, and precomputing (KM−1−QM−1)∈RM×M(K_M^{-1} - Q_M^{-1}) \in \mathbb{R}^{M \times M} requires O(M3)\mathcal{O}(M^3) operations once. For each new test case x∗x_*, computing the predictive mean μ∗\mu_* takes O(M)\mathcal{O}(M) time, and computing the predictive variance σ∗2\sigma_*^2 takes O(M2)\mathcal{O}(M^2) time.
  6. Knowl 6 — Equivalence of SPGP to Full Gaussian Process when Pseudo-Inputs Match Training Inputs

    theoretical result

    When the number of pseudo-inputs equals the number of training points (M=NM = N) and the pseudo-inputs coincide with the training inputs (Xˉ=X\bar{X} = X):

    1. The cross-covariance matrix equals the training covariance matrix: KNM=KMN=KM=KNK_{NM} = K_{MN} = K_M = K_N.
    2. The diagonal variance terms vanish identically: λn=K(xn,xn)−kn⊤KN−1kn=0\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_N^{-1} \mathbf{k}_n = 0 for all n∈{1,…,N}n \in \{1, \dots, N\}, meaning Λ=0\Lambda = \mathbf{0}.
    3. The SPGP marginal likelihood collapses exactly to the full GP marginal likelihood: p(y∣X,Xˉ=X,Θ)=N(y∣0,KN+σ2IN)p(\mathbf{y} \mid X, \bar{X} = X, \Theta) = \mathcal{N}(\mathbf{y} \mid \mathbf{0}, K_N + \sigma^2 I_N)
    4. The SPGP predictive distribution collapses exactly to the full GP predictive distribution: μ∗=k∗⊤(KN+σ2IN)−1y\mu_* = \mathbf{k}_*^\top (K_N + \sigma^2 I_N)^{-1} \mathbf{y} σ∗2=K(x∗,x∗)−k∗⊤(KN+σ2IN)−1k∗+σ2\sigma_*^2 = K(x_*, x_*) - \mathbf{k}_*^\top (K_N + \sigma^2 I_N)^{-1} \mathbf{k}_* + \sigma^2
  7. Knowl 7 — Empirical Performance on Benchmark Regression Datasets

    empirical result

    SPGP regression was evaluated on two benchmark datasets:

    • kin-40k: 10,000 training points, 30,000 test points, 9 input attributes.
    • pumadyn-32nm: 7,168 training points, 1,024 test points, 33 input attributes (with 4 truly relevant dimensions).

    SPGP was compared against subset selection methods (random subset, greedy information gain info-gain, and Smola-Bartlett greedy selection smo-bart), as well as full GPs trained on subsets (size 2,000 for kin-40k and 1,024 for pumadyn-32nm):

    • On kin-40k, with hyperparameters fixed to those obtained from a subset full GP, SPGP achieves substantially lower mean squared test error across active set sizes MM than random, info-gain, and smo-bart. With a pseudo-set size MM of a few hundred points, SPGP approaches the test error of the full GP trained on 2,000 points.
    • When jointly optimizing pseudo-inputs and hyperparameters from random initialization on kin-40k, SPGP performs strongly at small MM, matching the fixed-hyperparameter SPGP curve.
    • On pumadyn-32nm, when initialized with full GP hyperparameters, SPGP achieves the test error of the full GP with only M=25M = 25 pseudo-inputs, significantly outperforming discrete subset selection methods constrained to the same hyperparameters.
  8. Knowl 8 — Capturing Heteroscedastic and Non-Stationary Noise via Pseudo-Input Positioning

    empirical result

    Although Gaussian process regression with a stationary kernel assumes constant homoscedastic noise variance σ2\sigma^2, SPGP can model input-dependent noise variance through the spatial distribution of its pseudo-inputs. In empirical tests on 1D data with heteroscedastic noise, continuous optimization of the marginal likelihood shifts pseudo-inputs out of high-noise regions toward or beyond the edges of the data. Because λn=K(xn,xn)−kn⊤KM−1kn\lambda_n = K(x_n, x_n) - \mathbf{k}_n^\top K_M^{-1} \mathbf{k}_n grows toward K(xn,xn)K(x_n, x_n) as data point xnx_n moves away from all pseudo-inputs, the total noise variance λn+σ2\lambda_n + \sigma^2 automatically inflates in regions devoid of nearby pseudo-inputs, providing a closer fit to non-stationary data without requiring an explicitly non-stationary kernel function.

  9. Knowl 9 — Optimization Scalability, Local Optima, and Overfitting Limitations

    limitation

    Optimizing the SPGP objective over M⋅D+∣Θ∣M \cdot D + |\Theta| continuous parameters introduces several practical limitations:

    1. Local Optima in Automatic Relevance Determination (ARD): On high-dimensional datasets with many irrelevant features (e.g., pumadyn-32nm with 33 input dimensions), joint optimization from random initializations can get trapped in local optima where ARD fails to identify all relevant dimensions (identifying only 2 of 4 relevant dimensions on pumadyn-32nm), requiring initialization from full GP hyperparameters to reach optimal performance.
    2. Overfitting at Larger MM: When both pseudo-inputs and hyperparameters are learned jointly at larger pseudo-set sizes MM, the noise hyperparameter σ2\sigma^2 can be driven excessively small, resulting in higher marginal likelihood on training data but degraded test mean squared error.
    3. Scalability Constraints: When MM or DD is large, the parameter optimization space becomes very large, increasing training runtimes and susceptibility to optimization convergence issues when using standard gradient-based optimizers such as conjugate gradients or L-BFGS.

Coverage note — None omitted; all core contributions—including the SPGP generative model, marginal likelihood, posterior/predictive derivations, complexity, comparison with PLV, non-stationary noise modeling, empirical benchmark results, and limitations—are fully represented.

References

  1. 1.A. J. Smola and P. Bartlett. Sparse greedy Gaussian process regression. In Advances in Neural Information Processing Systems 13. MIT Press, 2000.
  2. 2.C. K. I. Williams and M. Seeger. Using the Nyström method to speed up kernel machines. In Advances in Neural Information Processing Systems 13. MIT Press, 2000.
  3. 3.V. Tresp. A Bayesian committee machine. Neural Computation, 12:2719–2741, 2000.
  4. 4.L. Csató. Sparse online Gaussian processes. Neural Computation, 14:641–668, 2002.
  5. 5.L. Csató. Gaussian Processes — Iterative Sparse Approximations. PhD thesis, Aston University, UK, 2002.
  6. 6.N. D. Lawrence, M. Seeger, and R. Herbrich. Fast sparse Gaussian process methods: the informative vector machine. In Advances in Neural Information Processing Systems 15. MIT Press, 2002.
  7. 7.M. Seeger, C. K. I. Williams, and N. D. Lawrence. Fast forward selection to speed up sparse Gaussian process regression. In C. M. Bishop and B. J. Frey, editors, Proceedings of the Ninth International Workshop on Artificial Intelligence and Statistics, 2003.
  8. 8.M. Seeger. Bayesian Gaussian Process Models: PAC-Bayesian Generalisation Error Bounds and Sparse Approximations. PhD thesis, University of Edinburgh, 2003.
  9. 9.J. Quiñonero Candela. Learning with Uncertainty — Gaussian Processes and Relevance Vector Machines. PhD thesis, Technical University of Denmark, 2004.
  10. 10.D. J. C. MacKay. Introduction to Gaussian processes. In C. M. Bishop, editor, Neural Networks and Machine Learning, NATO ASI Series, pages 133–166. Kluwer Academic Press, 1998.
  11. 11.C. K. I. Williams and C. E. Rasmussen. Gaussian processes for regression. In Advances in Neural Information Processing Systems 8. MIT Press, 1996.
  12. 12.C. E. Rasmussen. Evaluation of Gaussian Processes and Other Methods for Non-Linear Regression. PhD thesis, University of Toronto, 1996.
  13. 13.M. N. Gibbs. Bayesian Gaussian Processes for Regression and Classification. PhD thesis, Cambridge University, 1997.
  14. 14.F. Vivarelli and C. K. I. Williams. Discovering hidden features with Gaussian processes regression. In Advances in Neural Information Processing Systems 11. MIT Press, 1998.

Citation

MLA
Snelson, E., and Z. Ghahramani. “Sparse Gaussian Processes Using Pseudo-inputs”. Advances in Neural Information Processing Systems, vol. 18, 2005, https://proceedings.neurips.cc/paper_files/paper/2005/file/4491777b1aa8b5b32c2e8666dbe1a495-Paper.pdf.
APA
Snelson, E., & Ghahramani, Z. (2005). Sparse Gaussian Processes using Pseudo-inputs. Advances in Neural Information Processing Systems, 18. https://proceedings.neurips.cc/paper_files/paper/2005/file/4491777b1aa8b5b32c2e8666dbe1a495-Paper.pdf
Chicago
Snelson, E., and Z. Ghahramani. 2005. “Sparse Gaussian Processes Using Pseudo-inputs”. Advances in Neural Information Processing Systems 18. https://proceedings.neurips.cc/paper_files/paper/2005/file/4491777b1aa8b5b32c2e8666dbe1a495-Paper.pdf.
Harvard
Snelson, E. and Ghahramani, Z. (2005) “Sparse Gaussian Processes using Pseudo-inputs”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2005/file/4491777b1aa8b5b32c2e8666dbe1a495-Paper.pdf.
Vancouver
1. Snelson E, Ghahramani Z (2005) Sparse Gaussian Processes using Pseudo-inputs. Advances in Neural Information Processing Systems 18:

BibTeX

@inproceedings{snelson2005sparse,
  title = {Sparse Gaussian Processes using Pseudo-inputs},
  author = {Snelson, Edward and Ghahramani, Zoubin},
  year = {2005},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {18},
  url = {https://proceedings.neurips.cc/paper_files/paper/2005/file/4491777b1aa8b5b32c2e8666dbe1a495-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors