Multi-Task Learning as Multi-Objective Optimization

Ozan SenerVladlen Koltun

article2018NeurIPS1,863 citations

Proposes a scalable gradient-based optimization algorithm that frames multi-task learning as multi-objective optimization, guaranteeing Pareto-optimal trade-offs across competing tasks without relying on heuristic loss weighting.

Listen

Modern machine learning applications frequently rely on multi-task learning, where a single model concurrently solves multiple objectives by sharing internal representations. Standard multi-task learning approaches combine distinct task losses into a single objective using fixed or heuristic weights. However, this conventional proxy objective assumes that tasks do not compete for model capacity. In practice, competing tasks often lead to performance degradation across individual objectives, making standard optimization ineffective without costly, manual hyperparameter tuning.

The article aims to reformulate multi-task deep learning explicitly as a multi-objective optimization problem. Specifically, it seeks to develop a scalable, gradient-based algorithm that provably finds Pareto optimal solutions—where no single task's performance can be improved without harming another—with negligible computational overhead.

To achieve this, the authors adapt gradient-based multi-objective optimization techniques, which previously suffered from high computational cost due to requiring separate backward passes for each task across millions of shared parameters. The authors introduce an upper-bound formulation (MGDA-UB) tailored for standard encoder-decoder deep architectures. This formulation operates on task representations rather than entire parameter sets, allowing optimal gradient updates to be calculated in a single backward pass. The method is evaluated across three diverse benchmarks ranging from 2 to 40 tasks: overlapping digit classification on MultiMNIST (60,000 samples), 40-task facial attribute classification on CelebA (200,000 samples), and joint semantic segmentation, instance segmentation, and depth estimation on Cityscapes.

The experimental findings show that the proposed multi-objective method consistently outperforms standard multi-task baselines and independent single-task models. On MultiMNIST, where competing objectives degrade conventional joint training, the method achieves 97.26% and 95.90% accuracy on left and right digits, matching dedicated single-task performance. On CelebA, the approach attains the lowest average error rate of 8.25%, outperforming uniform loss scaling (9.62%) and dynamic baseline GradNorm (8.44%). On Cityscapes, the method improves semantic segmentation to 66.63% mean intersection-over-union while achieving lower pixel error rates in both instance segmentation (10.25 px) and disparity estimation (2.54 px). Computationally, the representation-level upper-bound approximation accelerates training by 40% in three-task setups and by a factor of 25 in 40-task setups compared to exact multi-objective gradients, while maintaining or slightly improving overall accuracy.

These results indicate that treating multi-task learning as multi-objective optimization eliminates the need for expensive trial-and-error weight tuning while unlocking the true inductive benefits of shared models. For organizations deploying multi-task systems, this enables the consolidation of distinct specialized models into single, higher-performing unified networks, substantially reducing memory footprints and deployment costs without task-performance trade-offs.

Engineering teams building shared-encoder neural networks should integrate this multi-objective update strategy into their existing gradient descent pipelines. Before broader deployment across non-standard architectures, practitioners should evaluate the method on their target data pipelines through controlled pilots.

The theoretical guarantees of this approach rely on the assumption that the representation-to-parameter Jacobian matrix is full-rank, which realistically holds when task requirements are non-redundant. Additionally, performance metrics on the Cityscapes benchmark were validated on the official validation dataset rather than hidden test labels due to ground-truth restrictions. Confidence in the core findings remains high, given consistent, robust empirical gains across diverse classification and regression problem domains.

Cover for Multi-Task Learning as Multi-Objective Optimization

Abstract

In multi-task learning, multiple tasks are solved jointly, sharing inductive bias between them. Multi-task learning is inherently a multi-objective problem because different tasks may conflict, necessitating a trade-off. A common compromise is to optimize a proxy objective that minimizes a weighted linear combination of per-task losses. However, this workaround is only valid when the tasks do not compete, which is rarely the case. In this paper, we explicitly cast multi-task learning as multi-objective optimization, with the overall objective of finding a Pareto optimal solution. To this end, we use algorithms developed in the gradient-based multi-objective optimization literature. These algorithms are not directly applicable to large-scale learning problems since they scale poorly with the dimensionality of the gradients and the number of tasks. We therefore propose an upper bound for the multi-objective loss and show that it can be optimized efficiently. We further prove that optimizing this upper bound yields a Pareto optimal solution under realistic assumptions. We apply our method to a variety of multi-task deep learning problems including digit classification, scene understanding (joint semantic segmentation, instance segmentation, and depth estimation), and multi-label classification. Our method produces higher-performing models than recent multi-task learning formulations or per-task training.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multi-Task Learning as Multi-Objective Optimization
  • 3.1 Multiple Gradient Descent Algorithm
  • 3.2 Solving the Optimization Problem
  • 3.3 Efficient Optimization for Encoder-Decoder Architectures
  • 4 Experiments
  • 4.1 MultiMNIST
  • 4.2 Multi-Label Classification
  • 4.3 Scene Understanding
  • 4.4 Role of the Approximation
  • 5 Conclusion
  • References
  • A Proof of Theorem 1
  • B Additional Results on Multi-label Classification
  • C Implementation Details
  • C.1 MultiMNIST
  • C.2 Multi-label classification
  • C.3 Scene understanding

Knowls

  1. Knowl 1 — Multi-Task Learning Formulated as Multi-Objective Optimization

    definition

    Consider a multi-task learning (MTL) problem over an input space X\mathcal{X} and a collection of task spaces {Yt}t∈[T]\{\mathcal{Y}^t\}_{t \in [T]}, given a dataset {xi,yi1,…,yiT}i∈[N]\{x_i, y_i^1, \dots, y_i^T\}_{i \in [N]}, where TT is the number of tasks, NN is the number of data points, and yity_i^t is the ground-truth label of the tt-th task for data point xix_i.

    Let ft(x;θsh,θt):X→Ytf^t(x; \theta^{sh}, \theta^t): \mathcal{X} \to \mathcal{Y}^t be a parametric hypothesis class for task tt, parameterized by shared parameters θsh\theta^{sh} across all tasks and task-specific parameters θt\theta^t. Each task has a task-specific loss Lt(⋅,⋅):Yt×Yt→R+\mathcal{L}^t(\cdot, \cdot): \mathcal{Y}^t \times \mathcal{Y}^t \to \mathbb{R}^+, giving the empirical risk:

    L^t(θsh,θt)=1N∑i=1NLt(ft(xi;θsh,θt),yit)\hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) = \frac{1}{N} \sum_{i=1}^N \mathcal{L}^t(f^t(x_i; \theta^{sh}, \theta^t), y_i^t)

    Rather than minimizing a scalarized weighted sum ∑t=1TctL^t(θsh,θt)\sum_{t=1}^T c^t \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t), MTL is cast as multi-objective optimization over the vector-valued loss L\mathbf{L}:

    min⁡θsh,θ1,…,θTL(θsh,θ1,…,θT)=min⁡θsh,θ1,…,θT(L^1(θsh,θ1),…,L^T(θsh,θT))⊤\min_{\theta^{sh}, \theta^1, \dots, \theta^T} \mathbf{L}(\theta^{sh}, \theta^1, \dots, \theta^T) = \min_{\theta^{sh}, \theta^1, \dots, \theta^T} \left( \hat{\mathcal{L}}^1(\theta^{sh}, \theta^1), \dots, \hat{\mathcal{L}}^T(\theta^{sh}, \theta^T) \right)^\top

    Pareto optimality in this multi-task setting is defined as follows:

    • A parameter vector θ=(θsh,θ1,…,θT)\theta = (\theta^{sh}, \theta^1, \dots, \theta^T) dominates θˉ=(θˉsh,θˉ1,…,θˉT)\bar{\theta} = (\bar{\theta}^{sh}, \bar{\theta}^1, \dots, \bar{\theta}^T) if L^t(θsh,θt)≤L^t(θˉsh,θˉt)\hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \le \hat{\mathcal{L}}^t(\bar{\theta}^{sh}, \bar{\theta}^t) for all t∈{1,…,T}t \in \{1, \dots, T\} and L(θ)≠L(θˉ)\mathbf{L}(\theta) \ne \mathbf{L}(\bar{\theta}).

    • A solution θ∗\theta^* is Pareto optimal if no solution θ\theta dominates θ∗\theta^*.

    The set of all Pareto optimal parameter configurations is the Pareto set Pθ\mathcal{P}_\theta, and its image in objective space is the Pareto front PL={L(θ)∣θ∈Pθ}\mathcal{P}_L = \{\mathbf{L}(\theta) \mid \theta \in \mathcal{P}_\theta\}.

  2. Knowl 2 — Multiple Gradient Descent Algorithm Formulation for Multi-Task Learning

    model/method

    In gradient-based multi-objective optimization for multi-task learning, the Karush-Kuhn-Tucker (KKT) necessary conditions for Pareto optimality across shared parameters θsh\theta^{sh} and task-specific parameters θt\theta^t are:

    1. There exist α1,…,αT≥0\alpha^1, \dots, \alpha^T \ge 0 such that ∑t=1Tαt=1\sum_{t=1}^T \alpha^t = 1 and ∑t=1Tαt∇θshL^t(θsh,θt)=0\sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) = 0.

    2. For every task t∈{1,…,T}t \in \{1, \dots, T\}, ∇θtL^t(θsh,θt)=0\nabla_{\theta^t} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) = 0.

    A parameter point satisfying these conditions is called a Pareto stationary point.

    To find a descent direction on the shared parameters that simultaneously decreases all task empirical losses, the Multiple Gradient Descent Algorithm (MGDA) solves the minimum-norm convex quadratic optimization problem:

    min⁡α1,…,αT{∥∑t=1Tαt∇θshL^t(θsh,θt)∥22  |  ∑t=1Tαt=1,  αt≥0  ∀t}\min_{\alpha^1, \dots, \alpha^T} \left\{ \left\| \sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \right\|_2^2 \;\middle|\; \sum_{t=1}^T \alpha^t = 1, \; \alpha^t \ge 0 \; \forall t \right\}

    If the optimal objective value is 00, the current parameters are Pareto stationary. Otherwise, the vector ∑t=1Tαt∇θshL^t(θsh,θt)\sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) is a descent direction that strictly decreases all task objectives.

  3. Knowl 3 — Upper Bound Formulation for Multi-Objective Optimization in Encoder-Decoder Architectures (MGDA-UB)

    model/method

    For neural networks structured as a shared encoder g(x;θsh)g(x; \theta^{sh}) and task-specific decoders ft(z;θt)f^t(z; \theta^t) such that ft(x;θsh,θt)=ft(g(x;θsh);θt)f^t(x; \theta^{sh}, \theta^t) = f^t(g(x; \theta^{sh}); \theta^t), let Z=(z1,…,zN)Z = (z_1, \dots, z_N) where zi=g(xi;θsh)z_i = g(x_i; \theta^{sh}) is the shared representation. Direct application of the chain rule yields the upper bound:

    ∥∑t=1Tαt∇θshL^t(θsh,θt)∥22≤∥∂Z∂θsh∥22∥∑t=1Tαt∇ZL^t(θsh,θt)∥22\left\| \sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \right\|_2^2 \le \left\| \frac{\partial Z}{\partial \theta^{sh}} \right\|_2^2 \left\| \sum_{t=1}^T \alpha^t \nabla_Z \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \right\|_2^2

    where ∥∂Z∂θsh∥2\|\frac{\partial Z}{\partial \theta^{sh}}\|_2 is the matrix spectral norm of the Jacobian of representations with respect to shared parameters θsh\theta^{sh}. Because ∥∂Z∂θsh∥22\|\frac{\partial Z}{\partial \theta^{sh}}\|_2^2 does not depend on the mixing coefficients α1,…,αT\alpha^1, \dots, \alpha^T, optimizing the upper bound reduces to solving MGDA-UB:

    min⁡α1,…,αT{∥∑t=1Tαt∇ZL^t(θsh,θt)∥22  |  ∑t=1Tαt=1,  αt≥0  ∀t}\min_{\alpha^1, \dots, \alpha^T} \left\{ \left\| \sum_{t=1}^T \alpha^t \nabla_Z \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \right\|_2^2 \;\middle|\; \sum_{t=1}^T \alpha^t = 1, \; \alpha^t \ge 0 \; \forall t \right\}

    While solving MGDA over θsh\theta^{sh} requires TT separate backward passes through the shared backbone, solving MGDA-UB requires computing gradients ∇ZL^t\nabla_Z \hat{\mathcal{L}}^t only through the task decoders to the representation ZZ, requiring only a single backward pass through the shared encoder per training iteration.

  4. Knowl 4 — Pareto Stationarity Guarantee for MGDA-UB

    theoretical result

    Let Z=(z1,…,zN)Z = (z_1, \dots, z_N) be the intermediate representations computed by a shared network g(x;θsh)g(x; \theta^{sh}) across TT tasks with empirical losses L^1,…,L^T\hat{\mathcal{L}}^1, \dots, \hat{\mathcal{L}}^T, and let ∂Z∂θsh\frac{\partial Z}{\partial \theta^{sh}} denote the Jacobian of ZZ with respect to θsh\theta^{sh}.

    Theorem: Assume the Jacobian ∂Z∂θsh\frac{\partial Z}{\partial \theta^{sh}} is full-rank. If α=(α1,…,αT)\alpha = (\alpha^1, \dots, \alpha^T) is the optimal solution to the MGDA-UB optimization problem:

    min⁡α1,…,αT{∥∑t=1Tαt∇ZL^t(θsh,θt)∥22  |  ∑t=1Tαt=1,  αt≥0  ∀t}\min_{\alpha^1, \dots, \alpha^T} \left\{ \left\| \sum_{t=1}^T \alpha^t \nabla_Z \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) \right\|_2^2 \;\middle|\; \sum_{t=1}^T \alpha^t = 1, \; \alpha^t \ge 0 \; \forall t \right\}

    then exactly one of the following holds:

    1. ∑t=1Tαt∇θshL^t(θsh,θt)=0\sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) = 0 and the current parameters are Pareto stationary.

    2. ∑t=1Tαt∇θshL^t(θsh,θt)\sum_{t=1}^T \alpha^t \nabla_{\theta^{sh}} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t) is a descent direction that decreases every task objective L^t\hat{\mathcal{L}}^t.

    This holds because full rank of ∂Z∂θsh\frac{\partial Z}{\partial \theta^{sh}} ensures that M=(∂Z∂θsh)⊤∂Z∂θshM = \left(\frac{\partial Z}{\partial \theta^{sh}}\right)^\top \frac{\partial Z}{\partial \theta^{sh}} is positive definite, making the optimality condition of MGDA-UB equivalent to the Pareto descent condition in the shared parameter space.

  5. Knowl 5 — Frank-Wolfe Solver for Min-Norm Objective in Convex Hulls

    algorithm

    Solving for the minimum-norm vector in the convex hull of TT gradients using the Frank-Wolfe algorithm relies on an exact analytical line search. For any two vectors θa\theta_a and θb\theta_b, the line search problem min⁡γ∈[0,1]∥γθa+(1−γ)θb∥22\min_{\gamma \in [0, 1]} \|\gamma \theta_a + (1-\gamma)\theta_b\|_2^2 has the closed-form solution:

    γ^=[(θb−θa)⊤θb∥θa−θb∥22]+,1=max⁡(min⁡((θb−θa)⊤θb∥θa−θb∥22,1),0)\hat{\gamma} = \left[ \frac{(\theta_b - \theta_a)^\top \theta_b}{\|\theta_a - \theta_b\|_2^2} \right]_{+, 1} = \max\left( \min\left( \frac{(\theta_b - \theta_a)^\top \theta_b}{\|\theta_a - \theta_b\|_2^2}, 1 \right), 0 \right)

    The Frank-Wolfe procedure finds optimal convex weights α=(α1,…,αT)\alpha = (\alpha^1, \dots, \alpha^T) over the precomputed T×TT \times T Gram matrix MM where Mi,j=(∇L^i)⊤(∇L^j)M_{i,j} = (\nabla \hat{\mathcal{L}}^i)^\top (\nabla \hat{\mathcal{L}}^j):

    Input: Gram matrix M∈RT×TM \in \mathbb{R}^{T \times T} where Mi,j=(∇L^i)⊤(∇L^j)M_{i,j} = (\nabla \hat{\mathcal{L}}^i)^\top (\nabla \hat{\mathcal{L}}^j)
    Output: Mixing coefficients α=(α1,…,αT)\alpha = (\alpha^1, \dots, \alpha^T)
    Initialize α=(1/T,…,1/T)\alpha = (1/T, \dots, 1/T)
    repeat
        Compute linear subproblem index: t^=arg⁡min⁡r∑t=1TαtMr,t\hat{t} = \arg\min_r \sum_{t=1}^T \alpha^t M_{r,t}
        Let et^e_{\hat{t}} be the standard basis vector for index t^\hat{t}
        Compute exact line-search step size γ^∈[0,1]\hat{\gamma} \in [0, 1]:
            γ^=arg⁡min⁡γ∈[0,1]((1−γ)α+γet^)⊤M((1−γ)α+γet^)\hat{\gamma} = \arg\min_{\gamma \in [0, 1]} ((1-\gamma)\alpha + \gamma e_{\hat{t}})^\top M ((1-\gamma)\alpha + \gamma e_{\hat{t}})
            γ^=[(et^−α)⊤Met^/((et^−α)⊤M(et^−α))]+,1\hat{\gamma} = [ (e_{\hat{t}} - \alpha)^\top M e_{\hat{t}} / ((e_{\hat{t}} - \alpha)^\top M (e_{\hat{t}} - \alpha)) ]_{+, 1}
        Update weight vector: α=(1−γ^)α+γ^et^\alpha = (1 - \hat{\gamma})\alpha + \hat{\gamma} e_{\hat{t}}
    until γ^≈0\hat{\gamma} \approx 0 or maximum iteration limit reached
    return α\alpha
  6. Knowl 6 — Multi-Task Learning Optimization Algorithm via MGDA-UB

    algorithm

    The complete training procedure optimizes multi-task networks by applying separate gradient descent steps on task-specific heads and a common Pareto descent update on shared parameters using weights computed by the Frank-Wolfe solver on representation gradients:

    Input: Training dataset, shared encoder g(⋅;θsh)g(\cdot; \theta^{sh}), task decoders {ft(⋅;θt)}t=1T\{f^t(\cdot; \theta^t)\}_{t=1}^T, learning rate η\eta
    Output: Shared parameters θsh\theta^{sh}, task-specific parameters {θt}t=1T\{\theta^t\}_{t=1}^T
    for each training step do
        Forward pass: compute representations Z=g(X;θsh)Z = g(X; \theta^{sh}) and task outputs ft(Z;θt)f^t(Z; \theta^t)
        for t=1t = 1 to TT do
            Compute task empirical loss L^t\hat{\mathcal{L}}^t
            Update task parameters: θt=θt−η∇θtL^t(θsh,θt)\theta^t = \theta^t - \eta \nabla_{\theta^t} \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t)
            Compute representation gradients: ∇ZL^t(θsh,θt)\nabla_Z \hat{\mathcal{L}}^t(\theta^{sh}, \theta^t)
        end for
        Compute T×TT \times T Gram matrix MM where Mi,j=(∇ZL^i)⊤(∇ZL^j)M_{i,j} = (\nabla_Z \hat{\mathcal{L}}^i)^\top (\nabla_Z \hat{\mathcal{L}}^j)
        Solve α1,…,αT=FrankWolfeSolver(M)\alpha^1, \dots, \alpha^T = \text{FrankWolfeSolver}(M)
        Compute shared representation descent direction: ∇ZL^=∑t=1Tαt∇ZL^t\nabla_Z \hat{\mathcal{L}} = \sum_{t=1}^T \alpha^t \nabla_Z \hat{\mathcal{L}}^t
        Backpropagate through shared encoder: ∇θshL^=(∂Z∂θsh)⊤∇ZL^\nabla_{\theta^{sh}} \hat{\mathcal{L}} = (\frac{\partial Z}{\partial \theta^{sh}})^\top \nabla_Z \hat{\mathcal{L}}
        Update shared parameters: θsh=θsh−η∇θshL^\theta^{sh} = \theta^{sh} - \eta \nabla_{\theta^{sh}} \hat{\mathcal{L}}
    end for
  7. Knowl 7 — MultiMNIST Multi-Task Classification Performance

    data/table

    MultiMNIST pairs two random MNIST digit images overlaid at the top-left (task-L) and bottom-right (task-R) corners, creating a 2-task image classification problem on 60,000 examples. A LeNet backbone serves as the shared representation function, followed by two separate fully-connected layers (50-dimensional hidden layer to 10-class output) per task.

    Could not parse LaTeX table

    On MultiMNIST, the two tasks compete directly for model capacity. Linear scaling schemes and heuristic weighting methods (Kendall et al., GradNorm) achieve lower accuracy than training independent models for each task. The multi-objective MGDA-UB approach matches or slightly exceeds single-task performance on both tasks simultaneously within a single model.

  8. Knowl 8 — CelebA 40-Task Multi-Label Classification Performance

    data/table

    Multi-label classification on the CelebA dataset (200,000 face images resized to 64×64×364 \times 64 \times 3) is formulated as a 40-task binary classification MTL problem. The architecture utilizes a ResNet-18 shared representation network (excluding the final layer) and 40 task-specific 2048×22048 \times 2 fully-connected linear layers trained via binary cross-entropy.

    Could not parse LaTeX table

    MGDA-UB achieves the lowest mean attribute classification error (8.25%), outperforming independent task training (8.77%), uniform loss weighting (9.62%), Kendall et al. (9.53%), and GradNorm (8.44%).

  9. Knowl 9 — Cityscapes Joint Scene Understanding Performance

    data/table

    Joint scene understanding on the Cityscapes dataset (256×512256 \times 512 resolution) optimizes three tasks concurrently: 19-class semantic segmentation (evaluated by mean Intersection over Union [mIoU, %], higher is better), instance segmentation center-of-mass regression offset (evaluated by MSE [pixels], lower is better), and monocular depth estimation (evaluated by disparity error [pixels], lower is better). The model shares a ResNet-50 encoder with task-specific pyramid pooling module (PSPNet) decoders.

    Could not parse LaTeX table

    In this multi-task setting where tasks cooperate, MGDA-UB outperforms all baselines across every metric, reaching 66.63% mIoU in segmentation compared to 60.68% for single-task training and 64.81% for GradNorm.

  10. Knowl 10 — Computational Efficiency and SGD Stability of MGDA-UB vs. Exact MGDA

    data/table

    Comparing exact MGDA (which performs TT backward passes through shared parameters θsh\theta^{sh}) against the MGDA-UB approximation (which performs 1 backward pass through θsh\theta^{sh}) on a single Titan Xp GPU demonstrates runtime acceleration and improved optimization stability:

    Could not parse LaTeX table

    MGDA-UB cuts training time by ~40% on 3 tasks and accelerates training by a factor of 26.7 on 40 tasks. The slight accuracy gain of MGDA-UB is explained by gradient estimation error in stochastic optimization: gradient error ∥α^−α∥2≤O(max⁡t∥et∥2)\|\hat{\alpha} - \alpha\|_2 \le \mathcal{O}(\max_t \|e^t\|_2) scales with the ℓ2\ell_2 norm of the random noise vector, which is lower when computed over representations ZZ (thousands of dimensions) than over full network parameters θsh\theta^{sh} (millions of dimensions).

Coverage note — None was omitted; all key theoretical formulations, optimization algorithms, and experimental benchmarks (MultiMNIST, CelebA, Cityscapes, and the approximation ablation) were fully captured.

References

  1. 1.A. Argyriou, T. Evgeniou, and M. Pontil. Multi-task feature learning. In NIPS, 2007.
  2. 2.A. Bagherjeiran, R. Vilalta, and C. F. Eick. Content-based image retrieval through a multi-agent meta-learning framework. In International Conference on Tools with Artificial Intelligence, 2005.
  3. 3.B. Bakker and T. Heskes. Task clustering and gating for Bayesian multitask learning. JMLR, 4:83–99, 2003.
  4. 4.J. Baxter. A model of inductive bias learning. Journal of Artificial Intelligence Research, 12:149–198, 2000.
  5. 5.H. Bilen and A. Vedaldi. Integrated perception with recurrent multi-task neural networks. In NIPS, 2016.
  6. 6.R. Caruana. Multitask learning. Machine Learning, 28(1):41–75, 1997.
  7. 7.Z. Chen, V. Badrinarayanan, C. Lee, and A. Rabinovich. GradNorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In ICML, 2018.
  8. 8.R. Collobert and J. Weston. A unified architecture for natural language processing: Deep neural networks with multitask learning. In ICML, 2008.
  9. 9.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The Cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  10. 10.P. B. C. de Miranda, R. B. C. Prudêncio, A. C. P. L. F. de Carvalho, and C. Soares. Combining a multi-objective optimization approach with meta-learning for SVM parameter selection. In International Conference on Systems, Man, and Cybernetics, 2012.
  11. 11.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009.
  12. 12.J.-A. Désidéri. Multiple-gradient descent algorithm (MGDA) for multiobjective optimization. Comptes Rendus Mathematique, 350(5):313–318, 2012.
  13. 13.D. Dong, H. Wu, W. He, D. Yu, and H. Wang. Multi-task learning for multiple language translation. In ACL, 2015.
  14. 14.M. Ehrgott. Multicriteria Optimization (2. ed.). Springer, 2005.
  15. 15.J. Fliege and B. F. Svaiter. Steepest descent methods for multicriteria optimization. Mathematical Methods of Operations Research, 51(3):479–494, 2000.
  16. 16.S. Ghosh, C. Lovell, and S. R. Gunn. Towards Pareto descent directions in sampling experts for multiple tasks in an on-line learning paradigm. In AAAI Spring Symposium: Lifelong Machine Learning, 2013.
  17. 17.K. Hashimoto, C. Xiong, Y. Tsuruoka, and R. Socher. A joint many-task model: Growing a neural network for multiple NLP tasks. In EMNLP, 2017.
  18. 18.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  19. 19.D. Hernández-Lobato, J. M. Hernández-Lobato, A. Shah, and R. P. Adams. Predictive entropy search for multi-objective bayesian optimization. In ICML, 2016.
  20. 20.J.-T. Huang, J. Li, D. Yu, L. Deng, and Y. Gong. Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers. In ICASSP, 2013.
  21. 21.Z. Huang, J. Li, S. M. Siniscalchi, I.-F. Chen, J. Wu, and C.-H. Lee. Rapid adaptation for deep neural networks through multi-task learning. In Interspeech, 2015.
  22. 22.M. Jaggi. Revisiting Frank-Wolfe: Projection-free sparse convex optimization. In ICML, 2013.
  23. 23.L. Kaiser, A. N. Gomez, N. Shazeer, A. Vaswani, N. Parmar, L. Jones, and J. Uszkoreit. One model to learn them all. arXiv:1706.05137, 2017.
  24. 24.A. Kendall, Y. Gal, and R. Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In CVPR, 2018.
  25. 25.I. Kokkinos. UberNet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. In CVPR, 2017.
  26. 26.H. W. Kuhn and A. W. Tucker. Nonlinear programming. In Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, Calif., 1951. University of California Press.
  27. 27.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  28. 28.C. Li, M. Georgiopoulos, and G. C. Anagnostopoulos. Pareto-path multi-task multiple kernel learning. arXiv:1404.3190, 2014.
  29. 29.X. Liu, J. Gao, X. He, L. Deng, K. Duh, and Y.-Y. Wang. Representation learning using multi-task deep neural networks for semantic classification and information retrieval. In NAACL HLT, 2015a.
  30. 30.Z. Liu, P. Luo, X. Wang, and X. Tang. Deep learning face attributes in the wild. In ICCV, 2015b.
  31. 31.M. Long and J. Wang. Learning multiple tasks with deep relationship networks. arXiv:1506.02117, 2015.
  32. 32.M.-T. Luong, Q. V. Le, I. Sutskever, O. Vinyals, and L. Kaiser. Multi-task sequence to sequence learning. arXiv:1511.06114, 2015.
  33. 33.N. Makimoto, I. Nakagawa, and A. Tamura. An efficient algorithm for finding the minimum norm point in the convex hull of a finite point set in the plane. Operations Research Letters, 16(1):33–40, 1994.
  34. 34.K. Miettinen. Nonlinear Multiobjective Optimization. Springer, 1998.
  35. 35.I. Misra, A. Shrivastava, A. Gupta, and M. Hebert. Cross-stitch networks for multi-task learning. In CVPR, 2016.
  36. 36.S. Parisi, M. Pirotta, N. Smacchia, L. Bascetta, and M. Restelli. Policy gradient approaches for multi-objective sequential decision making. In IJCNN, 2014.
  37. 37.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in PyTorch. In NIPS Workshops, 2017.
  38. 38.S. Peitz and M. Dellnitz. Gradient-based multiobjective optimization with uncertainties. In NEO, 2018.
  39. 39.M. Pirotta and M. Restelli. Inverse reinforcement learning through policy gradient minimization. In AAAI, 2016.
  40. 40.F. Poirion, Q. Mercier, and J. Désidéri. Descent algorithm for nonsmooth stochastic multiobjective optimization. Computational Optimization and Applications, 68(2):317–331, 2017.
  41. 41.D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley. A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research, 48:67–113, 2013.
  42. 42.C. Rosenbaum, T. Klinger, and M. Riemer. Routing networks: Adaptive selection of non-linear functions for multi-task learning. arXiv:1711.01239, 2017.
  43. 43.E. M. Rudd, M. Günther, and T. E. Boult. MOON: A mixed objective optimization network for the recognition of facial attributes. In ECCV, 2016.
  44. 44.S. Ruder. An overview of multi-task learning in deep neural networks. arXiv:1706.05098, 2017.
  45. 45.S. Sabour, N. Frosst, and G. E. Hinton. Dynamic routing between capsules. In NIPS, 2017.
  46. 46.S. Schäffler, R. Schultz, and K. Weinzierl. Stochastic method for the solution of unconstrained vector optimization problems. Journal of Optimization Theory and Applications, 114(1):209–222, 2002.
  47. 47.K. Sekitani and Y. Yamamoto. A recursive algorithm for finding the minimum norm point in a polytope and a pair of closest points in two polytopes. Mathematical Programming, 61(1-3):233–249, 1993.
  48. 48.M. L. Seltzer and J. Droppo. Multi-task learning in deep neural networks for improved phoneme recognition. In ICASSP, 2013.
  49. 49.A. Shah and Z. Ghahramani. Pareto frontier learning with expensive correlated objectives. In ICML, 2016.
  50. 50.C. Stein. Inadmissibility of the usual estimator for the mean of a multivariate normal distribution. Technical report, Stanford University, US, 1956.
  51. 51.P. Wolfe. Finding the nearest point in a polytope. Mathematical Programming, 11(1):128–149, 1976.
  52. 52.Y. Xue, X. Liao, L. Carin, and B. Krishnapuram. Multi-task learning for classification with dirichlet process priors. JMLR, 8:35–63, 2007.
  53. 53.Y. Yang and T. M. Hospedales. Trace norm regularised deep multi-task learning. arXiv:1606.04038, 2016.
  54. 54.A. R. Zamir, A. Sax, W. B. Shen, L. J. Guibas, J. Malik, and S. Savarese. Taskonomy: Disentangling task transfer learning. In CVPR, 2018.
  55. 55.Y. Zhang and D. Yeung. A convex formulation for learning task relationships in multi-task learning. In UAI, 2010.
  56. 56.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia. Pyramid scene parsing network. In CVPR, 2017.
  57. 57.B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba. Scene parsing through ADE20K dataset. In CVPR, 2017a.
  58. 58.D. Zhou, J. Wang, B. Jiang, H. Guo, and Y. Li. Multi-task multi-view learning based on cooperative multi-objective optimization. IEEE Access, 2017b.
  59. 59.J. Zhou, J. Chen, and J. Ye. Clustered multi-task learning via alternating structure optimization. In NIPS, 2011a.
  60. 60.J. Zhou, J. Chen, and J. Ye. MALSAR: Multi-task learning via structural regularization. Arizona State University, 2011b.

Citation

MLA
Sener, O., and V. Koltun. “Multi-Task Learning as Multi-Objective Optimization”. arXiv, 2018, http://arxiv.org/abs/1810.04650v2.
APA
Sener, O., & Koltun, V. (2018). Multi-Task Learning as Multi-Objective Optimization. arXiv. http://arxiv.org/abs/1810.04650v2
Chicago
Sener, O., and V. Koltun. 2018. “Multi-Task Learning as Multi-Objective Optimization”. arXiv. http://arxiv.org/abs/1810.04650v2.
Harvard
Sener, O. and Koltun, V. (2018) “Multi-Task Learning as Multi-Objective Optimization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1810.04650v2.
Vancouver
1. Sener O, Koltun V (2018) Multi-Task Learning as Multi-Objective Optimization. arXiv

BibTeX

@article{sener2018multi,
  title = {Multi-Task Learning as Multi-Objective Optimization},
  author = {Sener, Ozan and Koltun, Vladlen},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1810.04650v2},
  eprint = {1810.04650}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors