Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Derrick XinBehrooz GhorbaniJustin GilmerAnkush GargOrhan Firat

article2022NeurIPS91 citations

Demonstrates through large-scale empirical experiments that complex multi-task optimization algorithms fail to outperform properly tuned scalarized baselines, revealing flawed evaluation protocols in prior literature and offering practical strategies for reliable multi-task model training.

Listen

Training deep learning models simultaneously on multiple tasks often causes negative interference, where competing objectives degrade overall accuracy. To address this, recent research has introduced specialized multi-task optimization (MTO) algorithms designed to balance tasks dynamically during training. However, these specialized methods significantly increase computational overhead and system complexity. Given the high costs of training multi-task architectures, determining whether these algorithms deliver real-world advantages over traditional, simpler optimization approaches is a critical engineering and resource-allocation decision.

The article systematically evaluates whether current multi-task optimization algorithms outperform traditional scalarization—the standard practice of training models on a fixed, weighted average of individual task losses. To do so, the authors conduct a large-scale empirical study across supervised machine translation and computer vision benchmarks, including multilingual translation datasets (WMT), scene understanding (Cityscapes), and facial attribute classification (CelebA). They evaluate popular MTO methods (such as MGDA, GradNorm, PCGrad, IMTL, and RLW) against tuned scalarization baselines under controlled batch sizes, step counts, and extensive hyperparameter sweeps.

The analysis reveals three main findings. First, specialized multi-task optimization algorithms do not expand or improve the best achievable trade-off boundary (the Pareto front) across tasks. Instead, they merely land on different points along the same curve that can be traced using standard static weighting. Second, the apparent gains reported in prior literature stem largely from under-tuned baselines and evaluation flaws; variation caused by basic learning rate choices was found to be six to seven times larger than the variance across random seeds. Third, computing per-task gradients in MTO methods drastically degrades computational throughput, slowing training speeds from roughly twelve steps per second down to fewer than five steps per second.

These findings indicate that adopting complex multi-task optimization algorithms yields negligible performance benefits while sharply increasing computational expenditure, engineering maintenance, and training timelines. When the compute budget required for MTO methods is instead redirected toward scaling model capacity (such as increasing neural network depth), the entire performance trade-off curve shifts favorably, delivering superior generalization across all evaluation metrics.

Organizations and practitioners should avoid deploying complex multi-task optimization algorithms in supervised learning pipelines and instead rely on standard scalarization combined with rigorous hyperparameter tuning. When surplus compute is available, investing in model scaling or efficient exploration of static task weights provides a more dependable return on investment. The multi-task learning research community should also adopt standardized benchmark frameworks with strict validation and test splits to prevent false conclusions driven by under-tuned baselines.

The findings are supported with high confidence for supervised learning settings across vision and language modalities. However, limitations remain: searching for optimal static task weights via exhaustive sweeps can be resource-intensive, and the conclusions may not directly generalize to other learning paradigms, such as reinforcement learning or self-supervised pre-training, which require future empirical validation.

arXiv: 2209.11379
Cover for Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Abstract

Recent research has proposed a series of specialized optimization algorithms for deep multi-task models. It is often claimed that these multi-task optimization (MTO) methods yield solutions that are superior to the ones found by simply optimizing a weighted average of the task losses. In this paper, we perform large-scale experiments on a variety of language and vision tasks to examine the empirical validity of these claims. We show that, despite the added design and computational complexity of these algorithms, MTO methods do not yield any performance improvements beyond what is achievable via traditional optimization approaches. We highlight alternative strategies that consistently yield improvements to the performance profile and point out common training pitfalls that might cause suboptimal results. Finally, we outline challenges in reliably evaluating the performance of MTO algorithms and discuss potential solutions.

Table of Contents

  • 1 Introduction
  • 2 Setting
  • 3 Prior Work
  • 4 Experiments
  • 4.1 Multilingual Machine Translation
  • 4.2 Benchmarks from the Literature
  • 4.3 CityScapes
  • 4.3.1 CelebA
  • 5 Conclusions
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Scalarization empirically subsumes tested MTO methods

    empirical result

    Across the supervised language and vision workloads evaluated in the paper, multi-task optimization (MTO) methods—including MGDA, GradNorm, PCGrad, IMTL, and random loss weighting—did not produce performance profiles outside those obtained by ordinary scalarization with appropriately chosen task weights. MTO methods often selected different points from the same trade-off profile, but the observed scalarization solutions formed an empirical superset of the MTO solutions. Consequently, systematically exploring scalarization weights was more reliable than selecting a specialized MTO algorithm, despite scalarization being simpler and less computationally expensive.

  2. Knowl 2 — Pareto trade-offs and scalarized multi-task training

    model/method

    Consider a supervised model with parameters θ∈Rp\theta\in\mathbb{R}^{p} and KK task losses Li(θ)L_i(\theta), where lower loss is better. A parameter vector θ\theta Pareto dominates θ′\theta' if Li(θ)≤Li(θ′)L_i(\theta)\leq L_i(\theta') for every task i∈{1,…,K}i\in\{1,\ldots,K\} and the inequality is strict for at least one task. A parameter vector is Pareto optimal if no other parameter vector dominates it; the set of such points is the Pareto front.

    The conventional training method uses scalarization:

    θ^(w)=arg⁡min⁡θL(θ;w),L(θ;w)=∑i=1KwiLi(θ),\hat{\theta}(w)=\arg\min_{\theta}L(\theta;w),\qquad L(\theta;w)=\sum_{i=1}^{K}w_iL_i(\theta),

    where the task weights satisfy wi>0w_i>0 and ∑i=1Kwi=1\sum_{i=1}^{K}w_i=1. Varying ww searches for different compromises between task losses. In the machine-translation experiments, scalarization was implemented by sampling task ii with probability wiw_i; if ℓ(x;θ)\ell(x;\theta) is the example-level loss and Li(θ)L_i(\theta) is its expectation on task ii, then the expected sampled loss is E[ℓ(x;θ)]=∑i=1KwiLi(θ)\mathbb{E}[\ell(x;\theta)]=\sum_{i=1}^{K}w_iL_i(\theta).

  3. Knowl 3 — Guarantees of scalarization in the convex setting

    theoretical result

    Every minimizer of a positively weighted scalarized objective ∑i=1KwiLi(θ)\sum_{i=1}^{K}w_iL_i(\theta) is Pareto optimal, provided all weights are positive. When the task losses {Li}i=1K\{L_i\}_{i=1}^{K} are convex, the paper further states that every point θ#\theta^{\#} on the Pareto front can be obtained by scalarization with some nonnegative weight vector w#≥0w^{\#}\geq 0 (after normalization if necessary). Thus, for convex problems and optimization to convergence, sweeping scalarization weights is sufficient to reach the Pareto frontier, and no alternative optimization algorithm can achieve a better trade-off than a properly chosen scalarization solution.

  4. Knowl 4 — Large-scale comparison protocol for MTO and scalarization

    experimental setup

    The main comparison trained multi-task models with proportional scalarization and five MTO families: Multiple Gradient Descent (MGDA), GradNorm, Gradient Surgery (PCGrad), IMTL, and Random Loss Weighting (RLW), using both normal and Dirichlet random-weight variants. All methods used Adam, the same batch size, the same number of training steps, early stopping, and the same underlying model architecture within each experiment. Learning rates were tuned over a grid from 5×10−25\times10^{-2} to 55, and all non-Pareto-dominated runs were retained; GradNorm's additional α\alpha parameter was grid-searched.

    For multilingual translation, the experiments used pre-layer-normalized Transformers and paired English-to-French with English-to-Chinese, English-to-German, or English-to-Romanian translation. The corresponding WMT training/evaluation sizes were: English–French, 40,853,298/4,503; English–Chinese, 25,986,436/3,981; English–German, 4,548,885/2,169; and English–Romanian, 610,320/1,999 examples. Scalarization sampling rates were swept to construct the reference trade-off profile.

  5. Knowl 5 — MTO behavior across multilingual translation regimes

    empirical result

    For the high-resource English-to-{Chinese, French} and mid-resource English-to-{German, French} pairs, the final training and test losses from all evaluated MTO methods lay on the trade-off curves traced by scalarization. The MTO methods therefore produced alternative compromises rather than improvements over the scalarization frontier. The training plots also showed that the dynamically assigned task weights of MGDA, GradNorm, PCGrad, and IMTL changed little during most runs; with a common base learning rate of 0.50.5, their behavior was often close to static weighting.

    The low-resource English-to-{Romanian, French} pair showed a distinction between optimization and generalization. Training losses still exhibited a scalarization-like Pareto frontier, with MTO methods selecting points on that frontier. Test losses instead had globally favorable solutions rather than a visible Pareto frontier: scalarization with approximately (wRomanian,wFrench)=(0.3,0.7)(w_{\mathrm{Romanian}},w_{\mathrm{French}})=(0.3,0.7) reached the best observed trade-off, whereas MTO methods tended toward nearly equal task weights and substantially worse generalization. The authors attribute the generalization pattern primarily to how much regularization training provides to the low-resource Romanian task, casting doubt on the ability of the tested MTO methods to control this regularization effectively.

  6. Knowl 6 — Hyperparameter and utility-function instability in evaluation

    empirical result

    Reported MTO gains were highly sensitive to basic optimization choices. To quantify learning-rate-grid effects, the paper repeatedly considered three-rate searches of the form {k×10−3,k×10−2,k×10−1}\{k\times10^{-3},k\times10^{-2},k\times10^{-1}\} for each k∈{1,…,9}k\in\{1,\ldots,9\}. The effective standard deviation of the best result across these sparse searches was approximately 6–7 times the standard deviation obtained by rerunning the best fixed hyperparameter setting with different random seeds. Therefore, seed-based error estimates are insufficient for judging algorithmic improvements when hyperparameters are sampled sparsely.

    Ranking also depended on the evaluation utility used to combine task losses. In the English-to-{Chinese, French} experiment, weighting evaluation as 0.9LFrench+0.1LChinese0.9L_{\mathrm{French}}+0.1L_{\mathrm{Chinese}} favored MGDA over GradNorm and PCGrad, while equal evaluation weights favored GradNorm and PCGrad over MGDA. No MTO method surpassed a scalarization model using an appropriately chosen sampling ratio under any of the tested evaluation weightings. Evaluating only one arbitrary task weighting can therefore favor an algorithm that happens to land near that particular utility function.

  7. Knowl 7 — Scaling the model is a better use of MTO-level compute

    empirical result

    Computing per-task gradients for MTO reduced the observed training throughput from approximately 12 to 5 steps per second. The paper compared this use of compute with increasing Transformer depth under scalarization. The tested models had the following sizes and throughputs:

    Could not parse LaTeX table

    Increasing depth by factors of 1, 2, 3, and 4 consistently moved the English-to-{Chinese, French} Pareto front toward lower losses for both tasks. The largest scalarized model ran at 5.48 steps per second, comparable to MGDA's 4.81 steps per second, while providing improvements across the observed utility functions. In this comparison, allocating the additional compute to model capacity was more effective than allocating it to per-task-gradient optimization.

  8. Knowl 8 — CityScapes favors carefully weighted scalarization

    empirical result

    On CityScapes, the paper treated urban-scene understanding as two tasks: 7-class semantic segmentation and depth estimation. The dataset contains 2,975 training and 500 original validation images; the authors held out 595 random training images for hyperparameter tuning and used the original validation split for testing.

    Scalarization solutions formed the generalization Pareto frontier both in task-loss space—test segmentation loss versus test depth loss—and in task-metric space—pixel accuracy versus absolute depth error. The tested MTO methods, including MGDA, GradNorm, PCGrad, and both RLW variants, significantly underperformed the scalarization solutions. Because segmentation loss was about an order of magnitude larger than depth loss, the desirable scalarization frontier was not concentrated near equal loss weights: most of the generalization frontier came from models assigning less than 0.20.2 of the scalarization weight to segmentation. This result demonstrates that appropriate loss balancing, rather than an MTO update rule, was decisive for this benchmark.

  9. Knowl 9 — CelebA exposes tuning effects and inconsistent literature results

    data/table

    CelebA contains approximately 200,000 face images with 40 binary attribute-classification tasks. Across early-stopped runs, the paper found that variation from learning-rate and weight-decay tuning was much larger than the effect of choosing among PCGrad, equal weighting, MGDA, RLW with normal weights, and RLW with Dirichlet weights. The displayed average error rates were as follows, with lower values being better:

    Could not parse LaTeX table

    Reported CelebA results from different studies also disagreed substantially for the same method. The literature values and the paper's own test and validation values were:

    Could not parse LaTeX table

    The authors argue that under-tuned baselines, noisy validation measurements, early stopping, and frequent validation evaluation can explain part of this disagreement; validation-only reporting can make performance appear artificially strong.

  10. Knowl 10 — Evaluation standardization is needed for credible MTO claims

    limitation

    The paper identifies two methodological limitations of current MTO evaluation. First, constructing a scalarization Pareto frontier by exhaustive weight sweeping is computationally expensive, even though it provides a strong reference against which MTO methods can be compared. Efficient search procedures for the scalarization solution space remain an open direction.

    Second, the authors recommend a common-task evaluation framework with a shared benchmark, a fixed validation/test split, carefully tuned baselines, and consistent reporting. Such a framework would make improvements more credible because later methods would compete against progressively stronger shared baselines rather than against independently reimplemented and potentially under-tuned baselines. The study itself is restricted to supervised learning; whether the same conclusions hold for reinforcement-learning or self-supervised multi-task settings remains unresolved.

Coverage note — No substantial contributed material was omitted; appendix-level proofs, implementation details, and exhaustive auxiliary comparisons were not expanded because they support the load-bearing theoretical and empirical results rather than constitute separate contributions.

References

  1. 1.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al. Massively multilingual neural machine translation in the wild: Findings and challenges. arXiv preprint arXiv:1907.05019, 2019.
  2. 2.Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau. mslam: Massively multilingual joint pre-training for speech and text. arXiv preprint arXiv:2202.01374, 2022.
  3. 3.Stephen Boyd and Lieven Vandenberghe. Convex optimization. Cambridge university press, 2004.
  4. 4.Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In International Conference on Machine Learning, pages 794–803. PMLR, 2018.
  5. 5.Zhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, and Dragomir Anguelov. Just pick a sign: Optimizing deep multitask models with gradient sign dropout. Advances in Neural Information Processing Systems, 33:2039–2050, 2020.
  6. 6.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016.
  7. 7.Jean-Antoine Désidéri. Multiple-gradient descent algorithm (mgda) for multiobjective optimization. Comptes Rendus Mathematique, 350(5-6):313–318, 2012.
  8. 8.David Donoho. 50 years of data science. Journal of Computational and Graphical Statistics, 26(4):745–766, 2017.
  9. 9.Chris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu, Rohan Anil, and Chelsea Finn. Efficiently identifying task groupings for multi-task learning. Advances in Neural Information Processing Systems, 34, 2021.
  10. 10.Alex Graves, Marc G Bellemare, Jacob Menick, Remi Munos, and Koray Kavukcuoglu. Automated curriculum learning for neural networks. In international conference on machine learning, pages 1311–1320. PMLR, 2017.
  11. 11.Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. Lstm: A search space odyssey. IEEE transactions on neural networks and learning systems, 28(10):2222–2232, 2016.
  12. 12.Ishaan Gulrajani and David Lopez-Paz. In search of lost domain generalization. arXiv preprint arXiv:2007.01434, 2020.
  13. 13.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7482–7491, 2018.
  14. 14.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  15. 15.Julia Kreutzer, David Vilar, and Artem Sokolov. Bandits don’t follow rules: Balancing multi-facet machine translation with multi-armed bandits. arXiv preprint arXiv:2110.06997, 2021.
  16. 16.Vitaly Kurin, Alessandro De Palma, Ilya Kostrikov, Shimon Whiteson, and M Pawan Kumar. In defense of the unitary scalarization for deep multi-task learning. arXiv preprint arXiv:2201.04122, 2022.
  17. 17.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  18. 18.Xian Li and Hongyu Gong. Robust optimization for multilingual translation with imbalanced data. Advances in Neural Information Processing Systems, 34, 2021.
  19. 19.Mark Liberman. Obituary: Fred jelinek. Computational Linguistics, 36(4):595–599, 2010.
  20. 20.Baijiong Lin, Feiyang Ye, and Yu Zhang. A closer look at loss weighting in multi-task learning. arXiv preprint arXiv:2111.10603, 2021.
  21. 21.Liyang Liu, Yi Li, Zhanghui Kuang, J Xue, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang. Towards impartial multi-task learning. ICLR, 2021.
  22. 22.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pages 3730–3738, 2015.
  23. 23.Kevin Musgrave, Serge Belongie, and Ser-Nam Lim. Unsupervised domain adaptation: A reality check. arXiv preprint arXiv:2111.15672, 2021.
  24. 24.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002.
  25. 25.Matt Post. A call for clarity in reporting bleu scores. arXiv preprint arXiv:1804.08771, 2018.
  26. 26.Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pages 525–536, 2018.
  27. 27.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  28. 28.Xin Su, Yiyun Zhao, and Steven Bethard. A comparison of strategies for source-free domain adaptation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8352–8367, 2022.
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  30. 30.Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao. Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models. arXiv preprint arXiv:2010.05874, 2020.
  31. 31.Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, 33:5824–5836, 2020.

Citation

MLA
Xin, D., et al. “Do Current Multi-Task Optimization Methods in Deep Learning Even Help?”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 13597–609, https://proceedings.neurips.cc/paper_files/paper/2022/file/580c4ec4738ff61d5862a122cdf139b6-Paper-Conference.pdf.
APA
Xin, D., Ghorbani, B., Gilmer, J., Garg, A., & Firat, O. (2022). Do Current Multi-Task Optimization Methods in Deep Learning Even Help?. Advances in Neural Information Processing Systems, 35, 13597–13609. https://proceedings.neurips.cc/paper_files/paper/2022/file/580c4ec4738ff61d5862a122cdf139b6-Paper-Conference.pdf
Chicago
Xin, D., B. Ghorbani, J. Gilmer, A. Garg, and O. Firat. 2022. “Do Current Multi-Task Optimization Methods in Deep Learning Even Help?”. Advances in Neural Information Processing Systems 35: 13597–609. https://proceedings.neurips.cc/paper_files/paper/2022/file/580c4ec4738ff61d5862a122cdf139b6-Paper-Conference.pdf.
Harvard
Xin, D. et al. (2022) “Do Current Multi-Task Optimization Methods in Deep Learning Even Help?”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 13597–13609. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/580c4ec4738ff61d5862a122cdf139b6-Paper-Conference.pdf.
Vancouver
1. Xin D, Ghorbani B, Gilmer J, Garg A, Firat O (2022) Do Current Multi-Task Optimization Methods in Deep Learning Even Help?. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 13597–13609

BibTeX

@inproceedings{xin2022current,
  title = {Do Current Multi-Task Optimization Methods in Deep Learning Even Help?},
  author = {Xin, Derrick and Ghorbani, Behrooz and Gilmer, Justin and Garg, Ankush and Firat, Orhan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {13597-13609},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/580c4ec4738ff61d5862a122cdf139b6-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission