Regularized multi--task learning

T. EvgeniouM. Pontil

article2004KDD1,673 citations

Extends kernel-based regularization methods like Support Vector Machines to multi-task learning by formulating a shared mean parameter vector that models task relationships and significantly improves predictive accuracy over independent learning.

Listen

Modern data analysis often requires estimating multiple related predictive models at the same time, such as forecasting consumer preferences across different market segments or predicting academic outcomes across different schools. Traditionally, organizations train a separate model for each task or pool all data into a single aggregate model. However, training models independently ignores valuable shared patterns, while pooling data obscures critical individual variations.

The article develops and evaluates a unified mathematical framework for multi-task learning that generalizes single-task Support Vector Machines (SVMs)—a widely used classification and regression technique—into a simultaneous learning system. Specifically, the article demonstrates how modeling relationships between tasks through a novel task-coupling parameter can significantly enhance overall predictive accuracy.

To evaluate this framework, the authors conducted two sets of experiments. The first involved simulated marketing preference data (conjoint analysis) covering 30 to 100 simulated consumer tasks under varying noise and task-similarity conditions, comparing the proposed method against standard individual SVMs, pooled SVMs, and state-of-the-art Hierarchical Bayes models. The second experiment evaluated real-world educational data from 15,362 students across 139 secondary schools to predict student examination scores across multiple academic institutions.

The findings demonstrate three key outcomes. First, learning related tasks simultaneously consistently outperforms learning tasks independently, yielding substantially lower error rates. Second, the proposed method matches or outperforms established Bayesian benchmarks; on the real-world school dataset, it achieved an explained variance of roughly 34.4%, notably exceeding the 29.5% achieved by standard Bayesian clustering methods. Third, the task-coupling parameter provides strong operational safety: when tasks share little similarity, adjusting the parameter allows the model to naturally revert to independent task learning without suffering significant performance degradation.

These results indicate that organizations managing multi-entity data can achieve higher predictive accuracy without adopting overly complex Bayesian sampling procedures. By formulating multi-task learning as a standard SVM optimization problem, teams can leverage existing, highly reliable computational tools. Moreover, the framework reduces the risk of negative transfer across unrelated tasks, provided the coupling parameter is properly tuned.

Before deploying this framework at scale, organizations should implement automated validation protocols (such as cross-validation across tasks) to tune the coupling parameter, as incorrect settings can degrade performance. Computational efficiency should also be carefully assessed: training all tasks simultaneously can scale cubically with total sample size, meaning large datasets may require optimized solvers or cluster-based grouping before production rollout.

  • Paper: Multitask Learning, RICH CARUANA (1997). Caruana's seminal paper introduces the fundamental concept and motivation of multi-task learning through shared representations, establishing the foundational problem setting that the source formalizes into a regularized SVM framework.
  • Paper: Support-vector networks, Corinna Cortes et al. (1995). Cortes and Vapnik introduce the core mathematical formulation of Support Vector Machines, which serves as the direct optimization foundation that the source generalizes to simultaneous multi-task estimation.
  • Paper: Choosing Multiple Parameters for Support Vector Machines, OLIVIER CHAPELLE et al. (2002). This work establishes techniques for tuning multiple regularization and kernel parameters in SVMs via generalization bounds, directly informing the source's task-coupling regularization and parameter selection mechanisms.
  • Paper: On the Algorithmic Implementation of Multiclass Kernel-based Vector Machines, Koby Crammer et al. (2002). Crammer and Singer's unified multiclass SVM optimization formulation provides essential background on solving interconnected margin-maximization problems in a single training objective.
Cover for Regularized multi--task learning

Abstract

Past empirical work has shown that learning multiple related tasks from data simultaneously can be advantageous in terms of predictive performance relative to learning these tasks independently. In this paper we present an approach to multi-task learning based on the minimization of regularization functionals similar to existing ones, such as the one for Support Vector Machines (SVMs), that have been successfully used in the past for single-task learning. Our approach allows to model the relation between tasks in terms of a novel kernel function that uses a task-coupling parameter. We implement an instance of the proposed approach similar to SVMs and test it empirically using simulated as well as real data. The experimental results show that the proposed method performs better than existing multi-task learning methods and largely outperforms single-task learning using SVMs.

Table of Contents

  • 1. INTRODUCTION
  • 1.1 Related Work
  • 1.2 Notation and Setup
  • 2. METHODS FOR MULTI-TASK LEARNING
  • 2.1 Dual Optimization Problem
  • 2.2 Non-linear multi-task learning
  • 3. EXPERIMENTS
  • 3.1 Simulated Data
  • 3.2 Using a Real Dataset
  • 4. DISCUSSION
  • Acknowledgements
  • 5. REFERENCES

Knowls

  1. Knowl 1 — Primal Regularized Multi-Task Learning Formulation

    model/method

    For TT related classification tasks, let each task t∈{1,…,T}t \in \{1, \dots, T\} be provided with mm training points {(x1t,y1t),…,(xmt,ymt)}⊂Rd×{−1,1}\{(x_{1t}, y_{1t}), \dots, (x_{mt}, y_{mt})\} \subset \mathbb{R}^d \times \{-1, 1\}. The task-specific linear decision boundary wt∈Rdw_t \in \mathbb{R}^d is decomposed into a shared mean vector w0∈Rdw_0 \in \mathbb{R}^d and a task-specific deviation vt∈Rdv_t \in \mathbb{R}^d such that:

    wt=w0+vtw_t = w_0 + v_t

    The models are trained simultaneously by solving the convex quadratic optimization problem:

    min⁡w0,{vt}t=1T,{ξit}∑t=1T∑i=1mξit+λ1T∑t=1T∥vt∥2+λ2∥w0∥2\min_{w_0, \{v_t\}_{t=1}^T, \{\xi_{it}\}} \sum_{t=1}^T \sum_{i=1}^m \xi_{it} + \frac{\lambda_1}{T} \sum_{t=1}^T \|v_t\|^2 + \lambda_2 \|w_0\|^2

    subject to the constraints:

    yit(w0+vt)⋅xit≥1−ξit,ξit≥0(∀i∈{1,…,m},t∈{1,…,T})y_{it} (w_0 + v_t) \cdot x_{it} \ge 1 - \xi_{it}, \quad \xi_{it} \ge 0 \quad (\forall i \in \{1, \dots, m\}, t \in \{1, \dots, T\})

    where λ1>0\lambda_1 > 0 and λ2>0\lambda_2 > 0 are regularization parameters and ξit\xi_{it} are slack variables measuring margin violations. As λ1/λ2→∞\lambda_1 / \lambda_2 \to \infty, the deviations vt→0v_t \to 0, forcing all tasks to share the identical model w0w_0 (single-task pooling). As λ1/λ2→0\lambda_1 / \lambda_2 \to 0, w0→0w_0 \to 0, reducing the method to solving TT independent single-task Support Vector Machines.

  2. Knowl 2 — Mean-Centered Parameter Equivalence in Multi-Task Regularization

    theoretical result

    The primal multi-task formulation optimizing over shared parameter w0w_0 and task deviations vtv_t with regularization parameters λ1,λ2>0\lambda_1, \lambda_2 > 0 is mathematically equivalent to optimizing directly over individual task parameters wt=w0+vtw_t = w_0 + v_t under the following objective:

    min⁡{wt}t=1T,{ξit}∑t=1T∑i=1mξit+ρ1∑t=1T∥wt∥2+ρ2∑t=1T∥wt−1T∑s=1Tws∥2\min_{\{w_t\}_{t=1}^T, \{\xi_{it}\}} \sum_{t=1}^T \sum_{i=1}^m \xi_{it} + \rho_1 \sum_{t=1}^T \|w_t\|^2 + \rho_2 \sum_{t=1}^T \left\| w_t - \frac{1}{T}\sum_{s=1}^T w_s \right\|^2

    subject to:

    yitwt⋅xit≥1−ξit,ξit≥0(∀i∈{1,…,m},t∈{1,…,T})y_{it} w_t \cdot x_{it} \ge 1 - \xi_{it}, \quad \xi_{it} \ge 0 \quad (\forall i \in \{1, \dots, m\}, t \in \{1, \dots, T\})

    where the transformed regularization parameters ρ1\rho_1 and ρ2\rho_2 are uniquely related to λ1\lambda_1 and λ2\lambda_2 by:

    ρ1=1Tλ1λ2λ1+λ2,ρ2=1Tλ12λ1+λ2\rho_1 = \frac{1}{T} \frac{\lambda_1 \lambda_2}{\lambda_1 + \lambda_2}, \quad \rho_2 = \frac{1}{T} \frac{\lambda_1^2}{\lambda_1 + \lambda_2}

    At the optimal solution (w0∗,{vt∗}t=1T,{wt∗}t=1T)(w_0^*, \{v_t^*\}_{t=1}^T, \{w_t^*\}_{t=1}^T), the shared mean vector satisfies:

    w0∗=λ1λ1+λ2(1T∑t=1Twt∗)w_0^* = \frac{\lambda_1}{\lambda_1 + \lambda_2} \left( \frac{1}{T} \sum_{t=1}^T w_t^* \right)

  3. Knowl 3 — Dual SVM Formulation and the Task-Coupled Kernel

    theoretical result

    The linear multi-task classification optimization problem over TT tasks and mm samples per task maps to a single standard Support Vector Machine dual problem defined over an extended sample space X×{1,…,T}\mathcal{X} \times \{1, \dots, T\} with N=mTN = mT instances ((xit,t),yit)((x_{it}, t), y_{it}). Defining the task-coupling parameter μ=Tλ2λ1\mu = \frac{T \lambda_2}{\lambda_1} and the SVM box constraint C=T2λ1C = \frac{T}{2\lambda_1}, the dual problem is:

    max⁡{αit}∑t=1T∑i=1mαit−12∑s=1T∑t=1T∑i=1m∑j=1mαisyisαjtyjtKst(xis,xjt)\max_{\{\alpha_{it}\}} \sum_{t=1}^T \sum_{i=1}^m \alpha_{it} - \frac{1}{2} \sum_{s=1}^T \sum_{t=1}^T \sum_{i=1}^m \sum_{j=1}^m \alpha_{is} y_{is} \alpha_{jt} y_{jt} K_{st}(x_{is}, x_{jt})

    subject to:

    0≤αit≤C(∀i∈{1,…,m},t∈{1,…,T})0 \le \alpha_{it} \le C \quad (\forall i \in \{1, \dots, m\}, t \in \{1, \dots, T\})

    where the matrix-valued task-coupling kernel Kst:Rd×Rd→RK_{st}: \mathbb{R}^d \times \mathbb{R}^d \to \mathbb{R} is defined as:

    Kst(x,z)=(1μ+δst)x⋅zK_{st}(x, z) = \left( \frac{1}{\mu} + \delta_{st} \right) x \cdot z

    with Kronecker delta δst=1\delta_{st} = 1 if s=ts=t and 00 otherwise. Given the optimal dual multipliers αis∗\alpha_{is}^*, the prediction function for task tt on input xx is given by:

    ft∗(x)=∑s=1T∑i=1mαis∗Kst(xis,x)f_t^*(x) = \sum_{s=1}^T \sum_{i=1}^m \alpha_{is}^* K_{st}(x_{is}, x)

  4. Knowl 4 — Multi-Task Learning with Composite and Nonlinear Kernels

    model/method

    The linear multi-task formulation generalizes to nonlinear decision functions in Reproducing Kernel Hilbert Spaces. When task functions are modeled as ft=g+gtf_t = g + g_t, where gg is a function common to all tasks governed by kernel K1K_1 and gtg_t are task-specific functions governed by kernel K2K_2, the matrix-valued kernel takes the form:

    Kst(x,z)=1μK1(x,z)+δstK2(x,z),s,t∈{1,…,T}K_{st}(x, z) = \frac{1}{\mu} K_1(x, z) + \delta_{st} K_2(x, z), \quad s, t \in \{1, \dots, T\}

    For a general dataset of NN points ((xi,ti),yi)∈(X×{1,…,T})×{−1,1}((x_i, t_i), y_i) \in (\mathcal{X} \times \{1, \dots, T\}) \times \{-1, 1\} and any loss function VV, the multi-task learning problem is cast as:

    min⁡w∑i=1NV(yi,⟨w,Φ(xi,ti)⟩)+λ∥w∥2\min_{w} \sum_{i=1}^N V(y_i, \langle w, \Phi(x_i, t_i) \rangle) + \lambda \|w\|^2

    By the representer theorem, the optimal multi-task decision function F(x,t)=ft(x)F(x, t) = f_t(x) expands as:

    F(x,t)=∑i=1NβiG((x,t),(xi,ti))F(x, t) = \sum_{i=1}^N \beta_i G((x, t), (x_i, t_i))

    where G((x,t),(z,s))=Kts(x,z)G((x, t), (z, s)) = K_{ts}(x, z) is a symmetric positive-definite operator-valued kernel.

  5. Knowl 5 — Clustered Subgroup Regularization for Multi-Task Learning

    model/method

    When tasks belong to distinct or overlapping clusters rather than sharing a single global mean, the multi-task regularizer is generalized using index sets I1,…,Ic⊆{1,…,T}I_1, \dots, I_c \subseteq \{1, \dots, T\}, where each IiI_i contains indices of tasks that are similar to each other. The modified stabilizer in the objective function is:

    ρ1∑t=1T∥wt∥2+∑i=2c+1ρi∑t∈Ii∥wt−1Ti∑s∈Iiws∥2\rho_1 \sum_{t=1}^T \|w_t\|^2 + \sum_{i=2}^{c+1} \rho_i \sum_{t \in I_i} \left\| w_t - \frac{1}{T_i} \sum_{s \in I_i} w_s \right\|^2

    where Ti=∣Ii∣T_i = |I_i| is the cardinality of cluster IiI_i, ρ1>0\rho_1 > 0 regularizes the norm of individual task parameters, and ρi>0\rho_i > 0 controls the shrinkage of task parameters within subgroup IiI_i toward the subgroup mean.

  6. Knowl 6 — Multi-Task Kernel Formulation for Heterogeneous Feature Spaces

    model/method

    When each task operates on a different feature space Xt=Rdt\mathcal{X}_t = \mathbb{R}^{d_t} for t∈{1,…,T}t \in \{1, \dots, T\}, the problem is cast into the unified multi-task kernel framework by defining the joint input space X=X1×X2×⋯×XT\mathcal{X} = \mathcal{X}_1 \times \mathcal{X}_2 \times \cdots \times \mathcal{X}_T and a vector-valued function f:X→RTf: \mathcal{X} \to \mathbb{R}^T with components ft(x)=gt(xt)f_t(x) = g_t(x_t) for x=(x1,…,xT)∈Xx = (x_1, \dots, x_T) \in \mathcal{X}.

    Cross-task similarity is specified via functions Cst:Xs×Xt→RC_{st}: \mathcal{X}_s \times \mathcal{X}_t \to \mathbb{R}, defining the matrix-valued kernel:

    Kst(x,z)=Cst(xs,zt),s,t∈{1,…,T}K_{st}(x, z) = C_{st}(x_s, z_t), \quad s, t \in \{1, \dots, T\}

    where x=(x1,…,xT)x = (x_1, \dots, x_T) and z=(z1,…,zT)z = (z_1, \dots, z_T). The kernel G((x,s),(z,t))=Cst(xs,zt)G((x, s), (z, t)) = C_{st}(x_s, z_t) is symmetric and positive definite on X×{1,…,T}\mathcal{X} \times \{1, \dots, T\} and directly applies to dual SVM optimization.

  7. Knowl 7 — Conjoint Preference Simulation Benchmark on 30 and 100 Consumer Tasks

    data/table

    The regularized multi-task learning method (C=0.1,μ=0.1C=0.1, \mu=0.1) was evaluated against Hierarchical Bayes (HB), TT independent SVMs (C=0.1C=0.1), and a single pooled SVM (C=0.1C=0.1) on simulated consumer conjoint choice data with 4 attributes (16 binary dimensions). Datasets contained T=30T=30 or T=100T=100 individuals (tasks), each providing 16 choices transformed into 96 pairwise comparison vectors for training and 96 for testing. Experimental scenarios varied noise (Low: β=3\beta=3, High: β=0.5\beta=0.5) and task similarity (High: σ2=0.5β\sigma^2=0.5\beta, Low: σ2=3β\sigma^2=3\beta).

    T=30T = 30 Tasks T=100T = 100 Tasks
    Noise Sim. HB Proposed (μ=0.1\mu=0.1) TT SVMs 1 SVM HB Proposed (μ=0.1\mu=0.1) TT SVMs 1 SVM
    High Low 0.85 (26.14%) 0.81 (25.86%) 0.84 (26.22%) 1.32 (38.20%) 0.81 (24.65%) 0.79 (24.24%) 0.82 (24.98%) 1.31 (38.90%)
    High High 0.90 (31.03%) 0.86 (30.58%) 0.97 (31.60%) 0.96 (33.20%) 0.90 (31.49%) 0.90 (31.48%) 1.01 (33.13%) 0.96 (33.00%)
    Low Low 0.60 (14.34%) 0.58 (14.12%) 0.65 (16.00%) 1.00 (24.10%) 0.59 (13.97%) 0.58 (14.02%) 0.66 (15.57%) 0.98 (23.90%)
    Low High 0.48 (13.42%) 0.46 (13.19%) 0.68 (17.11%) 0.57 (15.80%) 0.47 (13.05%) 0.46 (13.28%) 0.66 (16.98%) 0.57 (15.80%)

    Entries show utility parameter Root Mean Square Error (RMSE) and test classification hit error rate in parentheses averaged across 5 runs. The proposed regularized multi-task method consistently outperforms independent SVMs and pooled single SVMs, and matches or outperforms Hierarchical Bayes even though the data generation process directly satisfies the Bayesian distributional assumptions.

  8. Knowl 8 — Multi-Task Regression Performance on ILEA School Examination Dataset

    data/table

    The proposed regularized multi-task method using an SVM ϵ\epsilon-loss regression formulation was evaluated on the Inner London Education Authority (ILEA) dataset, which consists of examination records for 15,362 students across 139 secondary schools (tasks), with 27 binary dummy-coded student- and school-level attributes. Generalization performance was evaluated across 10 random 75%/25% train/test splits (averaging ~70 train and ~40 test students per school) using percentage explained variance on test data (100×(1−SSE/Vartotal)100 \times (1 - \text{SSE}/\text{Var}_{\text{total}})).

    Method / Coupling C=0.1C = 0.1 C=1.0C = 1.0
    μ=0.5\mu = 0.5 34.30±0.3%34.30 \pm 0.3\% 34.37±0.4%34.37 \pm 0.4\%
    μ=1\mu = 1 34.28±0.4%34.28 \pm 0.4\% 34.37±0.3%34.37 \pm 0.3\%
    μ=2\mu = 2 34.26±0.4%34.26 \pm 0.4\% 34.11±0.4%34.11 \pm 0.4\%
    μ=10\mu = 10 34.32±0.3%34.32 \pm 0.3\% 29.71±0.4%29.71 \pm 0.4\%
    μ=1000\mu = 1000 (TT independent SVMs) 11.92±0.5%11.92 \pm 0.5\% 4.83±0.4%4.83 \pm 0.4\%
    Bayesian Task Clustering 29.5±0.4%29.5 \pm 0.4\% 29.5±0.4%29.5 \pm 0.4\%

    The regularized multi-task SVM with moderate task coupling (μ≤2\mu \le 2) achieves 34.37%34.37\% explained variance, significantly outperforming both independent per-school SVM regression (11.92%11.92\% at C=0.1C=0.1) and Bayesian task clustering (29.5%29.5\%).

  9. Knowl 9 — Computational Complexity Scaling of Multi-Task Kernel SVMs

    limitation

    Solving the multi-task dual SVM formulation requires solving a single joint quadratic programming problem with N=TmN = Tm total training data points across TT tasks with mm examples each. Using standard SVM training algorithms whose time complexity scales cubically in the total number of training instances O(N3)O(N^3), the multi-task formulation requires O(T3m3)O(T^3 m^3) computation time. In contrast, solving TT independent single-task SVMs takes T×O(m3)=O(Tm3)T \times O(m^3) = O(Tm^3) time, creating a cubic computational scaling bottleneck with respect to the number of tasks TT when training data or task counts are large.

Coverage note — None was omitted; all major contributions—including primal formulations, dual/kernel derivations, extensions to nonlinear/matrix kernels, subgroup clustering, simulation results, real-data benchmark results, and computational limitations—are fully covered.

References

  1. 1.G.M. Allenby and P.E. Rossi. Marketing Models of Consumer Heterogeneity. Journal of Econometrics, 89, p. 57–78, 1999.
  2. 2.N. Arora G.M Allenby, and J. Ginter. A Hierarchical Bayes Model of Primary and Secondary Demand. Marketing Science, 17,1, p. 29–44, 1998
  3. 3.N. Arora and J. Huber. Improving parameter estimates and model prediction by aggregate customization in choice experiments. Journal of Consumer Research, Vol. 28, September 2001.
  4. 4.B. Bakker and T. Heskes. Task clustering and gating for Bayesian multi–task learning. Journal of Machine Learning Research, 4: 83–99, 2003.
  5. 5.J. Baxter. A Bayesian/Information Theoretic Model of Learning to Learn via Multiple Task Sampling. Machine Learning, 28, pp. 7–39, 1997.
  6. 6.J. Baxter. A Model for Inductive Bias Learning. Journal of Artificial Intelligence Research, 12, p. 149–198, 2000.
  7. 7.S. Ben-David, J. Gehrke, and R. Schuller. A Theoretical Framework for Learning from a Pool of Disparate Data Sources. Proceedings of Knowledge Discovery and Datamining (KDD), 2002.
  8. 8.S. Ben-David and R. Schuller. Exploiting Task Relatedness for Multiple Task Learning. Proceedings of Computational Learning Theory (COLT), 2003.
  9. 9.L. Breiman and J.H Friedman. Predicting Multivariate Responses in Multiple Linear Regression. Royal Statistical Society Series B, 1998.
  10. 10.P.J. Brown and J.V. Zidek. Adaptive Multivariate Ridge Regression. The Annals of Statistics, Vol. 8, No. 1, p. 64–74, 1980.
  11. 11.R. Caruana. Multi–Task Learning. Machine Learning, 28, p. 41–75, 1997.
  12. 12.T. Evgeniou, C. Boussios, and G. Zacharia. Generalized Robust Conjoint Estimation. INSEAD Working Paper, 2002.
  13. 13.T. Evgeniou, M. Pontil, and T. Poggio. Regularization networks and support vector machines. Advances in Computational Mathematics, 13:1–50, 2000.
  14. 14.B. Heisele, T. Serre, M. Pontil, T. Vetter, and T. Poggio. Categorization by Learning and Combining Object Parts. In: Advances in Neural Information Processing Systems 14, Vancouver, Canada, Vol. 2, 1239–1245, 2002.
  15. 15.T. Heskes. Empirical Bayes for learning to learn. Proceedings of ICML–2000, ed. Langley, P., pp. 367–374, 2000.
  16. 16.M.I. Jordan and R.A. Jacobs. Hierarchical Mixtures of Experts and the EM Algorithm. Neural Computation, 1993.
  17. 17.J. Kim, G.M. Allenby, and P.E. Rossi. Modeling Consumer Demand for Variety. Marketing Science, Vol. 21, No. 3, pp. 229–250, Summer 2002.
  18. 18.G.R.G. Lanckriet, T. De Bie, N. Cristianini, M.I. Jordan, and W.S. Noble. A framework for genomic data fusion and its application to membrane protein prediction. Technical Report CSD–03–1273, Division of Computer Science, University of California, Berkeley, 2003.
  19. 19.O.L. Mangasarian. Nonlinear Programming. Classics in Applied Mathematics. SIAM, 1994.
  20. 20.C.A. Micchelli and M. Pontil. On Learning Vector–Valued Functions. Research Note RN/03/08, Dept of Computer Science, UCL, July 2003.
  21. 21.D.L. Silver and R.E Mercer. The parallel transfer of task knowledge using dynamic learning rates based on a measure of relatedness. Connection Science, 8, p. 277–294, 1996.
  22. 22.S. Thrun and L. Pratt. Learning to Learn. Kluwer Academic Publishers, November 1997.
  23. 23.S. Thrun and J. O’Sullivan. Clustering Learning Tasks and the Selective Cross–Task Transfer of Knowledge. In Learning To Learn, Editors S. Thrun and L.Y. Pratt, Kluwer Academic Publishers, 1998.
  24. 24.O. Toubia, D.I. Simester, J.R. Hauser, and E. Dahan. Fast Polyhedral Adaptive Conjoint Estimation. Working paper, MIT Sloan School of Management, 2001.
  25. 25.V. N. Vapnik. Statistical Learning Theory. Wiley, New York, 1998.
  26. 26.G. Wahba. Splines Models for Observational Data. Series in Applied Mathematics, Vol. 59, SIAM, Philadelphia, 1990.

Citation

MLA
Evgeniou, T., and M. Pontil. “Regularized Multi--task Learning”. Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2004, pp. 109–17, https://doi.org/10.1145/1014052.1014067.
APA
Evgeniou, T., & Pontil, M. (2004). Regularized multi--task learning. Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 109–117. https://doi.org/10.1145/1014052.1014067
Chicago
Evgeniou, T., and M. Pontil. 2004. “Regularized Multi--task Learning”. Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 109–17. https://doi.org/10.1145/1014052.1014067.
Harvard
Evgeniou, T. and Pontil, M. (2004) “Regularized multi--task learning”, Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp. 109–117. Available at: https://doi.org/10.1145/1014052.1014067.
Vancouver
1. Evgeniou T, Pontil M (2004) Regularized multi--task learning. In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp 109–117

BibTeX

@inproceedings{Evgeniou_2004, series={KDD04}, title={Regularized multi--task learning}, url={http://dx.doi.org/10.1145/1014052.1014067}, DOI={10.1145/1014052.1014067}, booktitle={Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining}, publisher={ACM}, author={Evgeniou, Theodoros and Pontil, Massimiliano}, year={2004}, month=Aug, pages={109–117}, collection={KDD04} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF