Multi-Task Learning as Multi-Objective Optimization
Ozan SenerVladlen Koltun
Proposes a scalable gradient-based optimization algorithm that frames multi-task learning as multi-objective optimization, guaranteeing Pareto-optimal trade-offs across competing tasks without relying on heuristic loss weighting.
Modern machine learning applications frequently rely on multi-task learning, where a single model concurrently solves multiple objectives by sharing internal representations. Standard multi-task learning approaches combine distinct task losses into a single objective using fixed or heuristic weights. However, this conventional proxy objective assumes that tasks do not compete for model capacity. In practice, competing tasks often lead to performance degradation across individual objectives, making standard optimization ineffective without costly, manual hyperparameter tuning.
The article aims to reformulate multi-task deep learning explicitly as a multi-objective optimization problem. Specifically, it seeks to develop a scalable, gradient-based algorithm that provably finds Pareto optimal solutions—where no single task's performance can be improved without harming another—with negligible computational overhead.
To achieve this, the authors adapt gradient-based multi-objective optimization techniques, which previously suffered from high computational cost due to requiring separate backward passes for each task across millions of shared parameters. The authors introduce an upper-bound formulation (MGDA-UB) tailored for standard encoder-decoder deep architectures. This formulation operates on task representations rather than entire parameter sets, allowing optimal gradient updates to be calculated in a single backward pass. The method is evaluated across three diverse benchmarks ranging from 2 to 40 tasks: overlapping digit classification on MultiMNIST (60,000 samples), 40-task facial attribute classification on CelebA (200,000 samples), and joint semantic segmentation, instance segmentation, and depth estimation on Cityscapes.
The experimental findings show that the proposed multi-objective method consistently outperforms standard multi-task baselines and independent single-task models. On MultiMNIST, where competing objectives degrade conventional joint training, the method achieves 97.26% and 95.90% accuracy on left and right digits, matching dedicated single-task performance. On CelebA, the approach attains the lowest average error rate of 8.25%, outperforming uniform loss scaling (9.62%) and dynamic baseline GradNorm (8.44%). On Cityscapes, the method improves semantic segmentation to 66.63% mean intersection-over-union while achieving lower pixel error rates in both instance segmentation (10.25 px) and disparity estimation (2.54 px). Computationally, the representation-level upper-bound approximation accelerates training by 40% in three-task setups and by a factor of 25 in 40-task setups compared to exact multi-objective gradients, while maintaining or slightly improving overall accuracy.
These results indicate that treating multi-task learning as multi-objective optimization eliminates the need for expensive trial-and-error weight tuning while unlocking the true inductive benefits of shared models. For organizations deploying multi-task systems, this enables the consolidation of distinct specialized models into single, higher-performing unified networks, substantially reducing memory footprints and deployment costs without task-performance trade-offs.
Engineering teams building shared-encoder neural networks should integrate this multi-objective update strategy into their existing gradient descent pipelines. Before broader deployment across non-standard architectures, practitioners should evaluate the method on their target data pipelines through controlled pilots.
The theoretical guarantees of this approach rely on the assumption that the representation-to-parameter Jacobian matrix is full-rank, which realistically holds when task requirements are non-redundant. Additionally, performance metrics on the Cityscapes benchmark were validated on the official validation dataset rather than hidden test labels due to ground-truth restrictions. Confidence in the core findings remains high, given consistent, robust empirical gains across diverse classification and regression problem domains.
- Paper: GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks, Zhao Chen et al. (2017). GradNorm provides the foundational gradient-balancing framework in deep multi-task networks that motivates formulating task competition through multi-objective optimization.
- Paper: Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, Alex Kendall et al. (2017). This paper introduces homoscedastic uncertainty weighting for multi-task vision benchmarks like Cityscapes, establishing the baseline loss-weighting paradigms the source seeks to replace.
- Paper: An Overview of Multi-Task Learning in Deep Neural Networks, Sebastian Ruder (2017). This overview outlines parameter-sharing architectures and optimization mechanics across multi-task deep networks, offering essential context on the trade-offs explored in the source.
- Paper: A Survey on Multi-Task Learning, Yu Zhang et al. (2017). This comprehensive survey categorizes classical multi-task learning formulations and theoretical setups, clarifying the limitations of standard linear scalarization.
- Paper: Multitask Learning, RICH CARUANA (1997). Caruana's seminal work establishes the fundamentals of hard parameter sharing and joint training dynamics across multiple tasks.
- Book: Convex Optimization: Algorithms and Complexity, Sébastien Bubeck (2015). This text provides core theoretical foundations on convex optimization and first-order gradient methods that underpin Pareto stationarity algorithms.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Gradient Surgery directly tackles the gradient conflict dynamics identified in multi-objective multi-task optimization by projecting conflicting task gradients.
- Paper: Pymoo: Multi-Objective Optimization in Python, Julian Blank et al. (2020). Pymoo provides a comprehensive Python optimization framework that operationalizes multi-objective optimization algorithms and Pareto front analysis.
- Paper: DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning, Hussein Hazimeh et al. (2021). DSelect-k extends multi-task optimization to differentiable mixture-of-experts gating to alleviate task interference across large task sets.
- Paper: FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking, Yifu Zhang et al. (2020). FairMOT explores task competition and balancing between detection and re-identification objectives in unified multi-task tracking architectures.
