Multi-task Gaussian Process Prediction

Edwin V. BonillaK. M. ChaiChristopher K. I. Williams

article2007NeurIPS1,386 citations

Proposes a flexible multi-task Gaussian Process framework that learns task correlations directly through a free-form inter-task covariance matrix while proving when inter-task transfer cancels out in noiseless settings.

Listen

Organizations frequently encounter predictive problems where data is split across multiple related domains or tasks, such as forecasting software performance across distinct programs or predicting student exam scores across different schools. Standard single-task models often struggle when training data per task is limited, while traditional multi-task methods frequently rely on hard-to-define domain descriptors or impose restrictive assumptions that degrade accuracy when tasks differ. The article addresses the challenge of enabling predictive models to transfer knowledge automatically across related tasks without requiring explicit task descriptor features.

The article's main objective is to introduce and evaluate a flexible multi-task Gaussian process model that learns inter-task correlations directly from data while sharing a common covariance function across inputs. It demonstrates the framework's effectiveness in boosting predictive accuracy and managing computational scalability across distinct operational datasets.

To evaluate this framework, the authors conducted empirical evaluations on two distinct real-world applications: a compiler performance dataset covering 11 programs tested across varying sample sizes from 16 to 128 code transformation sequences, and an examination score dataset comprising 15,362 student records across 139 secondary schools. The methodology compared the proposed free-form task covariance approach against isolated single-task models and parametric multi-task models relying on predefined task descriptors. In addition, the article established theoretical properties under idealized conditions and incorporated rank-reduction and sparse approximation techniques to ensure computational viability on large datasets.

The findings show that sharing information across tasks substantially improves predictive performance, particularly when data per task is sparse. In the compiler benchmark, the proposed method reduced mean absolute error by up to a factor of six on individual tasks compared to single-task learning, consistently outperforming the parametric baseline across nearly all programs. On the school examination benchmark, the multi-task model achieved an explained variance of approximately 29.2% using a low-rank task representation, outperforming isolated single-task learning (21.1%) and prior neural network baselines (9.7%). Furthermore, the theoretical analysis revealed an important operational boundary: under completely noiseless conditions with fully identical observation points across all tasks, inter-task transfer cancels out entirely, confirming that multi-task gains require real-world conditions like observational noise or non-identical task designs.

These results indicate that organizations can achieve higher predictive accuracy with substantially smaller training datasets per domain by learning task relationships automatically, directly reducing data collection costs and experiment runtimes. By eliminating the need to engineer task-specific descriptor features, the model reduces configuration risk and avoids poor transfer caused by mis-specified task representations. For decision-makers, gradient-based optimization offers faster convergence and higher-quality solutions than expectation-maximization techniques for training these models.

Organizations operating in multi-domain predictive environments should adopt multi-task Gaussian process models using low-rank task covariance approximations to balance accuracy and computational speed. When choosing an implementation, teams should assess whether high-quality task descriptors already exist: if descriptive features are readily available and sample sizes per task are extremely limited, parametric approaches remain competitive, but the free-form matrix model should be the primary default when task relationships are unknown or difficult to parameterize. Further evaluation is recommended to pilot the framework in larger real-time operational systems with highly unbalanced observation counts.

Confidence in these findings is supported by consistent replication across standard benchmarks and sound theoretical derivations. A key practical limitation is that computational complexity scales with the number of tasks and data points, requiring low-rank or sparse approximations when scaling up. Additionally, highly noisy or fundamentally uncorrelated input spaces can occasionally constrain transfer gains, as observed in isolated challenging tasks, warranting careful baseline validation prior to full system deployment.

Bonilla et al (2007).pdf
  • Paper: Regularized multi--task learning, T. Evgeniou et al. (2004). This foundational work establishes regularized multi-task learning across related domains and provides the examination score benchmark directly evaluated by the source paper.
  • Paper: Sparse Gaussian Processes using Pseudo-inputs, Edward Snelson et al. (2005). It introduces continuous pseudo-inputs for sparse Gaussian process regression, providing the core approximation machinery the source adapts for multi-task scalability.
  • Paper: A Unifying View of Sparse Approximate Gaussian Process Regression, Joaquin Quiñonero-Candela et al. (2005). It establishes a unified theoretical framework for sparse inducing-variable approximations that underpins the scalable Gaussian process inference used in the source.
  • Paper: Multitask Learning, RICH CARUANA (1997). This seminal paper introduces the paradigm and core empirical motivations for sharing representations across related tasks in multi-task learning.
Cover for Multi-task Gaussian Process Prediction

Abstract

In this paper we investigate multi-task learning in the context of Gaussian Processes (GP). We propose a model that learns a shared covariance function on input-dependent features and a “free-form” covariance matrix over tasks. This allows for good flexibility when modelling inter-task dependencies while avoiding the need for large amounts of data for training. We show that under the assumption of noise-free observations and a block design, predictions for a given task only depend on its target values and therefore a cancellation of inter-task transfer occurs. We evaluate the benefits of our model on two practical applications: a compiler performance prediction problem and an exam score prediction task. Additionally, we make use of GP approximations and properties of our model in order to provide scalability to large data sets.

Table of Contents

  • 1 Introduction
  • 2 The Model
  • 2.1 Inference
  • 2.2 Learning Hyperparameters
  • 2.3 Noiseless observations and the cancellation of inter-task transfer
  • 3 Approximations to speed up computations
  • 4 Related work
  • 5 Experiments
  • 5.1 Description of the Data
  • 6 Results
  • 7 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Multi-Task Gaussian Process Prior with Free-Form Task Covariance

    model/method

    For a multi-task regression problem with MM tasks over an input space X\mathcal{X}, the latent functions f1(x),…,fM(x)f_1(x), \dots, f_M(x) are assigned a joint zero-mean Gaussian process prior with a separable covariance structure:

    cov⁡(fl(x),fk(x′))=Klkfkx(x,x′)\operatorname{cov}(f_l(x), f_k(x')) = K^f_{lk} k^x(x, x')

    where Kf∈RM×MK^f \in \mathbb{R}^{M \times M} is a positive semi-definite (PSD) inter-task similarity matrix learned directly ("free-form") without requiring task-descriptor features, and kx:X×X→Rk^x: \mathcal{X} \times \mathcal{X} \to \mathbb{R} is a shared input covariance function. For stationary kernels, kxk^x is constrained to be a correlation function with unit variance kx(x,x)=1k^x(x, x) = 1 to eliminate scale redundancy with KfK^f.

    Observations yily_{il} for task l∈{1,…,M}l \in \{1, \dots, M\} at input location xi∈{x1,…,xN}x_i \in \{x_1, \dots, x_N\} are subject to task-specific Gaussian noise:

    yil∼N(fl(xi),σl2)y_{il} \sim \mathcal{N}(f_l(x_i), \sigma_l^2)

    where σl2\sigma_l^2 is the noise variance of task ll. Let y=vec⁡(Y)∈RMN\mathbf{y} = \operatorname{vec}(Y) \in \mathbb{R}^{MN} denote the stacked vector of responses across all NN inputs and MM tasks. The joint prior distribution over function values f=vec⁡(F)\mathbf{f} = \operatorname{vec}(F) and the marginal distribution of observations y\mathbf{y} given inputs XX are:

    f∼N(0,Kf⊗Kx),y∣X∼N(0,Kf⊗Kx+D⊗IN)\mathbf{f} \sim \mathcal{N}(0, K^f \otimes K^x), \quad \mathbf{y} \mid X \sim \mathcal{N}(0, K^f \otimes K^x + D \otimes I_N)

    where ⊗\otimes denotes the Kronecker product, Kx∈RN×NK^x \in \mathbb{R}^{N \times N} is the input Gram matrix with entries Kijx=kx(xi,xj)K^x_{ij} = k^x(x_i, x_j), D=diag⁡(σ12,…,σM2)D = \operatorname{diag}(\sigma_1^2, \dots, \sigma_M^2) is the diagonal matrix of task noise variances, and INI_N is the N×NN \times N identity matrix.

  2. Knowl 2 — Predictive Distribution for Multi-Task Gaussian Process Regression

    equation

    Given observations y∈RMN\mathbf{y} \in \mathbb{R}^{MN} across MM tasks at input points X={x1,…,xN}X = \{x_1, \dots, x_N\}, the posterior predictive mean for task ll at an unobserved input location x∗x_* is given by:

    fˉl(x∗)=(klf⊗k∗x)TΣ−1y\bar{f}_l(x_*) = (\mathbf{k}^f_l \otimes \mathbf{k}^x_*)^T \Sigma^{-1} \mathbf{y}

    where the total covariance matrix Σ∈RMN×MN\Sigma \in \mathbb{R}^{MN \times MN} is:

    Σ=Kf⊗Kx+D⊗IN\Sigma = K^f \otimes K^x + D \otimes I_N

    and the constituent terms are:

    • klf∈RM\mathbf{k}^f_l \in \mathbb{R}^M is the ll-th column of the task covariance matrix KfK^f, representing task similarities with task ll.
    • k∗x=[kx(x∗,x1),…,kx(x∗,xN)]T∈RN\mathbf{k}^x_* = [k^x(x_*, x_1), \dots, k^x(x_*, x_N)]^T \in \mathbb{R}^N is the vector of input covariances between x∗x_* and all NN training inputs.
    • Kx∈RN×NK^x \in \mathbb{R}^{N \times N} is the input Gram matrix with elements Kijx=kx(xi,xj)K^x_{ij} = k^x(x_i, x_j).
    • D=diag⁡(σ12,…,σM2)∈RM×MD = \operatorname{diag}(\sigma_1^2, \dots, \sigma_M^2) \in \mathbb{R}^{M \times M} is the diagonal matrix of task noise variances.
    • INI_N is the N×NN \times N identity matrix, and ⊗\otimes denotes the Kronecker product.

    When observations yo\mathbf{y}_o are available only on a subset of the full M×NM \times N task-input grid, Σ\Sigma and the covariance vectors are restricted to the visible entry indices.

  3. Knowl 3 — Cancellation of Inter-Task Transfer under Noiseless Block Designs

    theoretical result

    In a multi-task Gaussian process model with a separable covariance prior cov⁡(fl(x),fk(x′))=Klkfkx(x,x′)\operatorname{cov}(f_l(x), f_k(x')) = K^f_{lk} k^x(x, x'), if the observations are noise-free (D=0D = 0, so that y=f\mathbf{y} = \mathbf{f}) and measured on a complete block design (the exact same input set X={x1,…,xN}X = \{x_1, \dots, x_N\} observed for all MM tasks), predictions for each individual task depend exclusively on the observations from that specific task, resulting in complete cancellation of inter-task transfer.

    Specifically, the joint predictive mean vector fˉ(x∗)=[fˉ1(x∗),…,fˉM(x∗)]T\bar{\mathbf{f}}(x_*) = [\bar{f}_1(x_*), \dots, \bar{f}_M(x_*)]^T at a new input x∗x_* simplifies via the mixed-product property of Kronecker products:

    fˉ(x∗)=(Kf⊗k∗x)T(Kf⊗Kx)−1y=(Kf(Kf)−1)⊗((k∗x)T(Kx)−1)y=((k∗x)T(Kx)−1y⋅1⋮(k∗x)T(Kx)−1y⋅M)\bar{\mathbf{f}}(x_*) = (K^f \otimes \mathbf{k}^x_*)^T (K^f \otimes K^x)^{-1} \mathbf{y} = (K^f (K^f)^{-1}) \otimes ((\mathbf{k}^x_*)^T (K^x)^{-1}) \mathbf{y} = \begin{pmatrix} (\mathbf{k}^x_*)^T (K^x)^{-1} \mathbf{y}_{\cdot 1} \\ \vdots \\ (\mathbf{k}^x_*)^T (K^x)^{-1} \mathbf{y}_{\cdot M} \end{pmatrix}

    where y⋅l=[y1l,…,yNl]T\mathbf{y}_{\cdot l} = [y_{1l}, \dots, y_{Nl}]^T is the target vector for task ll, and k∗x\mathbf{k}^x_* is the covariance vector between x∗x_* and XX.

    This property (known in geostatistics as autokrigeability) implies that inter-task transfer within this model occurs only in the presence of observation noise (D≠0D \neq 0) or when data is missing / observed at differing input sets across tasks.

  4. Knowl 4 — Scalable Multi-Task GP Inversion via Joint Low-Rank Approximations

    model/method

    To overcome the O(M3N3)\mathcal{O}(M^3 N^3) time complexity of inverting the full multi-task covariance matrix Σ=Kf⊗Kx+D⊗IN\Sigma = K^f \otimes K^x + D \otimes I_N, low-rank approximations are applied simultaneously to both task and input covariance matrices:

    1. The M×MM \times M task covariance is approximated by a rank-PP incomplete Cholesky decomposition: Kf≈K~f=L~L~TK^f \approx \tilde{K}^f = \tilde{L}\tilde{L}^T, where L~∈RM×P\tilde{L} \in \mathbb{R}^{M \times P} with P≤MP \le M.
    2. The N×NN \times N input covariance is approximated via a Nyström / inducing input projection using QQ inducing points indexed by I⊂{1,…,N}I \subset \{1, \dots, N\}: Kx≈K~x=K⋅Ix(KIIx)−1KI⋅xK^x \approx \tilde{K}^x = K^x_{\cdot I} (K^x_{II})^{-1} K^x_{I\cdot}, where K⋅Ix∈RN×QK^x_{\cdot I} \in \mathbb{R}^{N \times Q} and KIIx∈RQ×QK^x_{II} \in \mathbb{R}^{Q \times Q}.

    Setting Σ~=K~f⊗K~x+Δ\tilde{\Sigma} = \tilde{K}^f \otimes \tilde{K}^x + \Delta with Δ=D⊗IN\Delta = D \otimes I_N, the Woodbury matrix inversion formula yields:

    Σ~−1=Δ−1−Δ−1B(IP⊗KIIx+BTΔ−1B)−1BTΔ−1\tilde{\Sigma}^{-1} = \Delta^{-1} - \Delta^{-1} B \left( I_P \otimes K^x_{II} + B^T \Delta^{-1} B \right)^{-1} B^T \Delta^{-1}

    where B=L~⊗K⋅Ix∈RMN×PQB = \tilde{L} \otimes K^x_{\cdot I} \in \mathbb{R}^{MN \times PQ}. Because K~f⊗K~x\tilde{K}^f \otimes \tilde{K}^x has rank PQPQ, evaluating Σ~−1y\tilde{\Sigma}^{-1} \mathbf{y} reduces to solving a PQ×PQPQ \times PQ linear system, which scales with time complexity O(MNP2Q2)\mathcal{O}(MN P^2 Q^2) per iteration.

  5. Knowl 5 — Expectation-Maximization Algorithm for Multi-Task GP Hyperparameters

    algorithm

    An Expectation-Maximization (EM) algorithm for learning the input kernel hyperparameters θx\theta_x, the free-form task covariance KfK^f, and the task noise variances σ12,…,σM2\sigma_1^2, \dots, \sigma_M^2 treats the latent function values F∈RN×MF \in \mathbb{R}^{N \times M} (where Fil=fl(xi)F_{il} = f_l(x_i)) as missing data.

    Input: Training observations Y∈RN×MY \in \mathbb{R}^{N \times M}, input matrix X∈RN×dX \in \mathbb{R}^{N \times d}, initial hyperparameters θx(0),Kf(0),D(0)\theta_x^{(0)}, K^{f(0)}, D^{(0)}
    Output: Estimated parameters θ^x,K^f,σ^12,…,σ^M2\hat{\theta}_x, \hat{K}^f, \hat{\sigma}_1^2, \dots, \hat{\sigma}_M^2
    repeat
        // E-step: compute posterior moments of latent functions given current parameters
        Compute p(F∣Y,θx,Kf,D)=N(μF,ΣF)p(F \mid Y, \theta_x, K^f, D) = \mathcal{N}(\mu_F, \Sigma_F)
        Evaluate expected terms ⟨FT(Kx(θx))−1F⟩\langle F^T (K^x(\theta_x))^{-1} F \rangle and ⟨(y⋅l−f⋅l)T(y⋅l−f⋅l)⟩\langle (\mathbf{y}_{\cdot l} - \mathbf{f}_{\cdot l})^T (\mathbf{y}_{\cdot l} - \mathbf{f}_{\cdot l}) \rangle for each l∈{1,…,M}l \in \{1, \dots, M\}
        // M-step: decoupled parameter updates
        θ^x←arg⁡min⁡θx(Nlog⁡∣⟨FT(Kx(θx))−1F⟩∣+Mlog⁡∣Kx(θx)∣)\hat{\theta}_x \leftarrow \arg\min_{\theta_x} \left( N \log |\langle F^T (K^x(\theta_x))^{-1} F \rangle| + M \log |K^x(\theta_x)| \right)
        K^f←1N⟨FT(Kx(θ^x))−1F⟩\hat{K}^f \leftarrow \frac{1}{N} \langle F^T (K^x(\hat{\theta}_x))^{-1} F \rangle
        for l=1l = 1 to MM do
            σ^l2←1N⟨(y⋅l−f⋅l)T(y⋅l−f⋅l)⟩\hat{\sigma}_l^2 \leftarrow \frac{1}{N} \langle (\mathbf{y}_{\cdot l} - \mathbf{f}_{\cdot l})^T (\mathbf{y}_{\cdot l} - \mathbf{f}_{\cdot l}) \rangle
        end for
    until convergence

    The M-step yields closed-form updates for K^f\hat{K}^f and σ^l2\hat{\sigma}_l^2, ensuring that K^f\hat{K}^f remains positive semi-definite by construction.

  6. Knowl 6 — EM Regularization under Rank-Deficient Nyström Input Covariances

    model/method

    When a rank-QQ Nyström approximation K~x=K⋅Ix(KIIx)−1KI⋅x\tilde{K}^x = K^x_{\cdot I} (K^x_{II})^{-1} K^x_{I\cdot} is used inside the EM hyperparameter learning algorithm, the log-determinant log⁡∣K~x∣\log |\tilde{K}^x| diverges to −∞-\infty and the standard matrix inverse is undefined. To resolve this, K~x\tilde{K}^x is regularized as the limit lim⁡ξ→0(K⋅Ix(KIIx)−1KI⋅x+ξ2IN)\lim_{\xi \to 0} (K^x_{\cdot I} (K^x_{II})^{-1} K^x_{I\cdot} + \xi^2 I_N).

    Under this formulation, the objective function in the M-step for updating the input kernel parameters θx\theta_x replaces the matrix inverse with the Moore-Penrose pseudo-inverse and substitutes the ill-defined log-determinant with the well-defined scalar:

    log⁡∣KI⋅xK⋅Ix∣−log⁡∣KIIx∣\log |K^x_{I\cdot} K^x_{\cdot I}| - \log |K^x_{II}|

    This regularized optimization maintains numerical stability and retains the overall per-iteration learning complexity of O(MNP2Q2)\mathcal{O}(MN P^2 Q^2) for rank-PP task covariance and rank-QQ input covariance approximations.

  7. Knowl 7 — Empirical Performance on Compiler Optimization Speed-up Prediction

    empirical result

    The multi-task Gaussian process model was evaluated on predicting program speed-up factors across M=11M = 11 C programs (tasks) subjected to code transformation sequences described by 13-dimensional binary vectors. Models were trained on small sample sizes N∈{16,32,64,128}N \in \{16, 32, 64, 128\} selected from 88,214 total sequences per program, evaluated over 10 independent replications using Mean Absolute Error (MAE).

    Key empirical findings include:

    • On average across all 11 tasks, multi-task GP with a free-form covariance matrix Kf=LLTK^f = LL^T achieved lower MAE than both single-task GP ("no transfer") and a parametric multi-task GP that used 8 canonical response speed-ups as task descriptors.
    • On task 1 (histogram), the free-form multi-task GP reduced MAE up to 6-fold compared to the no-transfer baseline.
    • On task 2 (fir), the free-form multi-task GP maintained low error across all sample sizes, whereas the parametric multi-task GP degraded at N=64N = 64 and N=128N = 128, performing worse than the no-transfer baseline.
    • On only one task (adpcm, task 3) did multi-task learning fail to improve over no-transfer, attributed to high residual variance unexplainable by the 13 input features.
  8. Knowl 8 — Exam Score Percentage Explained Variance across Rank Approximations

    data/table

    On the Inner London Education Authority (ILEA) school effectiveness dataset (15,362 students across M=139M = 139 secondary schools/tasks, using 19 student-dependent categorical dummy features), multi-task Gaussian processes were evaluated over 10 random 75%/25% train/test splits. Performance was measured by percentage explained variance (r2×100r^2 \times 100, higher is better):

    Model No Transfer Parametric Rank 1 Rank 2 Rank 3 Rank 5
    Explained Variance (%) 21.05 (1.15) 31.57 (1.61) 27.02 (2.03) 29.20 (1.60) 24.88 (1.62) 21.00 (2.42)

    Among the free-form task covariance models Kf≈L~L~TK^f \approx \tilde{L}\tilde{L}^T, the rank-2 approximation yielded the highest performance (29.20%±1.60%29.20\% \pm 1.60\%), outperforming the single-task no-transfer baseline (21.05%±1.15%21.05\% \pm 1.15\%) and matching the literature's best neural network multi-task baseline (29.5%29.5\%). The parametric task-descriptor model achieved 31.57%±1.61%31.57\% \pm 1.61\% by leveraging school-dependent descriptor features.

Coverage note — None was omitted; all key theoretical properties, modeling equations, optimization algorithms, low-rank approximations, and experimental findings are covered.

References

  1. 1.Jonathan Baxter. A Model of Inductive Bias Learning. JAIR, 12:149–198, March 2000.
  2. 2.Rich Caruana. Multitask Learning. Machine Learning, 28(1):41–75, July 1997.
  3. 3.Edwin V. Bonilla, Felix V. Agakov, and Christopher K. I. Williams. Kernel Multi-task Learning using Task-specific Features. In Proceedings of the 11th AISTATS, March 2007.
  4. 4.Kai Yu, Wei Chu, Shipeng Yu, Volker Tresp, and Zhao Xu. Stochastic Relational Models for Discriminative Link Prediction. In NIPS 19, Cambridge, MA, 2007. MIT Press.
  5. 5.Yee Whye Teh, Matthias Seeger, and Michael I. Jordan. Semiparametric latent factor models. In Proceedings of the 10th AISTATS, pages 333–340, January 2005.
  6. 6.Hao Zhang. Maximum-likelihood estimation for multivariate spatial linear coregionalization models. Environmetrics, 18(2):125–139, 2007.
  7. 7.Hans Wackernagel. Multivariate Geostatistics: An Introduction with Applications. Springer-Verlag, Berlin, 2nd edition, 1998.
  8. 8.A. O'Hagan. A Markov property for covariance structures. Statistics Research Report 98-13, Nottingham University, 1998.
  9. 9.C. K. I. Williams, K. M. A. Chai, and E. V. Bonilla. A note on noise-free Gaussian process prediction with separable covariance functions and grid designs. Technical report, University of Edinburgh, 2007.
  10. 10.C. E. Rasmussen and C. K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, Cambridge, Massachusetts, 2006.
  11. 11.Joaquin Quiñonero-Candela, Carl Edward Rasmussen, and Christopher K. I. Williams. Approximation Methods for Gaussian Process Regression. In Large Scale Kernel Machines. MIT Press, 2007. To appear.
  12. 12.Michael E. Tipping and Christopher M. Bishop. Probabilistic principal component analysis. Journal of the Royal Statistical Society, Series B, 61(3):611–622, 1999.
  13. 13.S. Thrun. Is Learning the n-th Thing Any Easier Than Learning the First? In NIPS 8, 1996.
  14. 14.Thomas P. Minka and Rosalind W. Picard. Learning How to Learn is Learning with Point Sets. 1999.
  15. 15.Neil D. Lawrence and John C. Platt. Learning to learn with the Informative Vector Machine. In Proceedings of the 21st International Conference on Machine Learning, July 2004.
  16. 16.Kai Yu, Volker Tresp, and Anton Schwaighofer. Learning Gaussian Processes from Multiple Tasks. In Proceedings of the 22nd International Conference on Machine Learning, 2005.
  17. 17.Anton Schwaighofer, Volker Tresp, and Kai Yu. Learning Gaussian Process Kernels via Hierarchical Bayes. In NIPS 17, Cambridge, MA, 2005. MIT Press.
  18. 18.Shipeng Yu, Kai Yu, Volker Tresp, and Hans-Peter Kriegel. Collaborative Ordinal Regression. In Proceedings of the 23rd International Conference on Machine Learning, June 2006.
  19. 19.Theodoros Evgeniou, Charles A. Micchelli, and Massimiliano Pontil. Learning Multiple Tasks with Kernel Methods. Journal of Machine Learning Research, 6:615–537, April 2005.
  20. 20.Bart Bakker and Tom Heskes. Task Clustering and Gating for Bayesian Multitask Learning. Journal of Machine Learning Research, 4:83–99, May 2003.

Citation

MLA
Bonilla, E., et al. “Multi-task Gaussian Process Prediction”. Advances in Neural Information Processing Systems, vol. 20, 2007, https://proceedings.neurips.cc/paper_files/paper/2007/file/66368270ffd51418ec58bd793f2d9b1b-Paper.pdf.
APA
Bonilla, E., Chai, K., & Williams, C. (2007). Multi-task Gaussian Process Prediction. Advances in Neural Information Processing Systems, 20. https://proceedings.neurips.cc/paper_files/paper/2007/file/66368270ffd51418ec58bd793f2d9b1b-Paper.pdf
Chicago
Bonilla, E., K. Chai, and C. Williams. 2007. “Multi-task Gaussian Process Prediction”. Advances in Neural Information Processing Systems 20. https://proceedings.neurips.cc/paper_files/paper/2007/file/66368270ffd51418ec58bd793f2d9b1b-Paper.pdf.
Harvard
Bonilla, E., Chai, K. and Williams, C. (2007) “Multi-task Gaussian Process Prediction”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2007/file/66368270ffd51418ec58bd793f2d9b1b-Paper.pdf.
Vancouver
1. Bonilla E, Chai K, Williams C (2007) Multi-task Gaussian Process Prediction. Advances in Neural Information Processing Systems 20:

BibTeX

@inproceedings{bonilla2007multi,
  title = {Multi-task Gaussian Process Prediction},
  author = {Bonilla, Edwin and Chai, Kian and Williams, Christopher},
  year = {2007},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {20},
  url = {https://proceedings.neurips.cc/paper_files/paper/2007/file/66368270ffd51418ec58bd793f2d9b1b-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors