A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity

Michinari MommaChaosheng DongJia Liu

article2022ICML61 citations

Develops a generic multi-objective learning framework based on Pareto stationarity that incorporates user preferences and extends weighted Chebyshev optimization to discover models outperforming existing baselines in a single training run.

Listen

Modern machine learning systems, such as search ranking, recommendation engines, and multi-task deep neural networks, frequently require optimizing multiple competing objectives simultaneously. In commercial environments, engineering teams routinely face the costly challenge of retraining models on fresh data while ensuring the new model does not degrade performance on any key metric relative to an existing baseline. Traditional multi-objective optimization techniques either struggle to align models with explicit business preferences or require extensive, manual trial-and-error exploration across all objectives, leading to development cycles that typically consume several days.

The article develops a unified mathematical framework for multi-objective and multi-task learning that simultaneously achieves optimal trade-off efficiency—known as Pareto optimality—and alignment with specific user-defined preferences. Specifically, it introduces two gradient-based algorithms: weighted Chebyshev multi-gradient descent and an extended version that explicitly explores improvements relative to an existing baseline model.

To demonstrate and validate this framework, the authors conducted empirical evaluations across both synthetic benchmarks and multiple real-world tasks. The experimental suite included two-task image classification across three separate datasets (MultiMNIST, Multi-Fashion, and Multi-Fashion+MNIST with 120,000 training images each), eight-target river flow regression across the Mississippi River network, and multi-class emotion classification in music. The proposed method was evaluated against established industry benchmarks, including linear scalarization, Pareto multi-task learning, and Exact Pareto Optimal search, measuring relative loss profiles and overall hypervolume coverage.

The findings show that the extended algorithm successfully identifies models that strictly outperform an existing reference model while directly following specified trade-off preferences. In practice, this reduces the required tuning exploration from scaling linearly with the number of objectives down to a single optimization run. The proposed method achieved the highest hypervolume metric in five out of eight evaluation settings and ranked second in the remaining three, demonstrating superior ability to dominate competing approaches. Furthermore, the algorithm’s automatic parameter-tuning mechanism ensured smooth convergence within 100 iterations on synthetic tests, outperforming fixed-parameter setups that required over 130 iterations.

These results provide immediate practical value for production engineering. By reducing the tuning overhead to a single run, organizations can substantially reduce computational costs, eliminate days of manual engineering labor, and de-risk model updates by guaranteeing that baseline performance is maintained or exceeded. Organizations facing frequent model retraining should consider adopting this framework to automate multi-task model updates and replace manual hyperparameter searches. Confidence in these results is high across complex, competing objectives; however, practitioners should note that the method operates on gradient-based models and does not provide an advantage on simpler problems where basic linear combinations already suffice.

Momma et al (2022).pdf
Cover for A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity

Abstract

Multi-objective optimization (MOO) and multi-task learning (MTL) have gained much popularity with prevalent use cases such as production model development of regression / classification / ranking models with MOO, and training deep learning models with MTL. Despite the long history of research in MOO, its application to machine learning requires development of solution strategy, and algorithms have recently been developed to solve specific problems such as discovery of any Pareto optimal (PO) solution, and that with a particular form of preference. In this paper, we develop a novel and generic framework to discover a PO solution with multiple forms of preferences. It allows us to formulate a generic MOO / MTL problem to express a preference, which is solved to achieve both alignment with the preference and PO, at the same time. Specifically, we apply the framework to solve the weighted Chebyshev problem and an extension of that. The former is known as a method to discover the Pareto front, the latter helps to find a model that outperforms an existing model with only one run. Experimental results demonstrate not only the method achieves competitive performance with existing methods, but also it allows us to achieve the performance from different forms of preferences.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. An MOO / MTL problem
  • 3.2. MGDA
  • 3.3. Weighted Chebyshev - MGDA
  • 3.3.1. WC-MGDA Formulation
  • 3.3.2. Solution Strategy
  • 3.3.3. Extended WC-MGDA
  • 3.3.4. Computational Complexity
  • 4. Experiments
  • 4.1. Synthetic data
  • 4.2. Real dataset
  • 4.2.1. Image Classification
  • 4.2.2. Multi-Target Regression
  • 4.2.3. Multi-Class Classification
  • 4.2.4. Hypervolumes for All Experiments
  • 5. Conclusion
  • References
  • A. Proofs
  • A.1. Proof of Lemma 3.2
  • A.2. Proof of Lemma 3.3
  • A.3. Proof of Lemma 3.4

Knowls

  1. Knowl 1 — WC-MGDA combines preference alignment with Pareto stationarity

    model/method

    Consider minimizing a vector of mm differentiable losses l(x)∈Rml(x)\in\mathbb{R}^m over model parameters x∈Rnx\in\mathbb{R}^n. Let G(x)=[∇l1(x),…,∇lm(x)]∈Rn×mG(x)=[\nabla l_1(x),\ldots,\nabla l_m(x)]\in\mathbb{R}^{n\times m}, let K(x)=(G(x)⊤G(x))1/2∈Rm×mK(x)=(G(x)^\top G(x))^{1/2}\in\mathbb{R}^{m\times m}, and let r∈R++mr\in\mathbb{R}_{++}^m be a componentwise-positive preference vector. Define Kr=diag⁡(r)Kdiag⁡(r)K_r=\operatorname{diag}(\sqrt r)K\operatorname{diag}(\sqrt r), where the square root is componentwise. WC-MGDA formulates preference alignment and approximate Pareto stationarity in one optimization problem:

    max⁡γ, α∈R+m, x∈Rnα⊤(r⊙l(x))−uγsubject toe⊤α=1,∥Krα∥2≤γ,\begin{aligned} \max_{\gamma,\,\alpha\in\mathbb{R}^m_+,\,x\in\mathbb{R}^n}\quad &\alpha^\top(r\odot l(x))-u\gamma\\ \text{subject to}\quad &e^\top\alpha=1,\qquad \|K_r\alpha\|_2\leq\gamma, \end{aligned}

    Here e∈Rme\in\mathbb{R}^m is the all-ones vector, ⊙\odot denotes elementwise multiplication, α\alpha is a nonnegative simplex vector, γ≥0\gamma\geq0 bounds the stationarity residual, and u>0u>0 trades off preference alignment against that residual. The constraint Krα=0K_r\alpha=0 implies Pareto stationarity: because every component of rr is positive, Kα~=0K\widetilde\alpha=0 for a nonzero nonnegative vector α~=diag⁡(r)α\widetilde\alpha=\operatorname{diag}(\sqrt r)\alpha, which can be normalized to sum to one.

    For fixed xx, the equivalent primal form is

    min⁡ρ∈R, d∈Rmρsubject tor⊙l(x)+Krd≤ρe,∥d∥2≤u.\begin{aligned} \min_{\rho\in\mathbb{R},\,d\in\mathbb{R}^m}\quad &\rho\\ \text{subject to}\quad &r\odot l(x)+K_rd\leq\rho e,\qquad \|d\|_2\leq u. \end{aligned}

    The primal minimizes the largest component of the preference-weighted losses after a correction that also encodes the stationarity trade-off. For fixed parameters this is a second-order cone problem; jointly optimizing over xx is nonconvex. The construction exemplifies the paper's approach of starting from a preference-bearing optimization problem and incorporating Pareto stationarity, rather than modifying a stationarity-only method after the fact.

  2. Knowl 2 — XWC-MGDA pivots the preferred Pareto search around a reference model

    model/method

    Let b~∈Rm\widetilde b\in\mathbb{R}^m contain the losses of a reference model, such as an existing or pretrained model. Extended WC-MGDA (XWC-MGDA) applies the preference vector r∈R++mr\in\mathbb{R}_{++}^m to the loss difference l(x)−b~l(x)-\widetilde b, so the search is centered on the reference rather than the origin. To discourage weak Pareto-optimal solutions, it also imposes a positive lower bound on the dual weight vector. For any componentwise-positive v∈R++mv\in\mathbb{R}_{++}^m, define w=v/(1+e⊤v)w=v/(1+e^\top v); then w>0w>0 and e⊤w<1e^\top w<1.

    With Kr=diag⁡(r)(G⊤G)1/2diag⁡(r)K_r=\operatorname{diag}(\sqrt r)(G^\top G)^{1/2}\operatorname{diag}(\sqrt r), XWC-MGDA's dual and primal forms are

    max⁡γ, α∈Rm, x∈Rnα⊤(r⊙(l(x)−b~))−uγsubject toe⊤α=1,α≥w,∥Krα∥2≤γ,\begin{aligned} \max_{\gamma,\,\alpha\in\mathbb{R}^m,\,x\in\mathbb{R}^n}\quad &\alpha^\top\bigl(r\odot(l(x)-\widetilde b)\bigr)-u\gamma\\ \text{subject to}\quad &e^\top\alpha=1,\quad \alpha\geq w,\quad \|K_r\alpha\|_2\leq\gamma, \end{aligned} min⁡ρ∈R, d∈Rm, z∈R+mρ−w⊤zsubject tor⊙(l(x)−b~)+Krd+z≤ρe,∥d∥2≤u.\begin{aligned} \min_{\rho\in\mathbb{R},\,d\in\mathbb{R}^m,\,z\in\mathbb{R}^m_+}\quad &\rho-w^\top z\\ \text{subject to}\quad &r\odot(l(x)-\widetilde b)+K_rd+z\leq\rho e,\\ &\|d\|_2\leq u. \end{aligned}

    Here zz is a nonnegative slack vector, u>0u>0 is the stationarity trade-off, and α\alpha is constrained componentwise. Setting b~=0\widetilde b=0 and w=0w=0 recovers WC-MGDA. The reference-relative formulation lets a user seek solutions aligned with a preference from a chosen baseline; if the resulting model satisfies l(x)−b~<0l(x)-\widetilde b<0 componentwise, it strictly improves every reference loss. The formulation seeks such improvements but does not guarantee that every run will produce a dominating model.

  3. Knowl 3 — A primal-objective criterion automatically adjusts the stationarity trade-off

    theoretical result

    For WC-MGDA at fixed model parameters xx, let ρ0\rho_0 be the previous primal objective value and let KrK_r be nonzero. The automatic adjustment chooses the smallest feasible trade-off uu subject to not worsening that objective:

    min⁡u∈R, ρ∈R, d∈Rmusubject tor⊙l(x)+Krd≤ρe,∥d∥2≤u,ρ≤ρ0.\begin{aligned} \min_{u\in\mathbb{R},\,\rho\in\mathbb{R},\,d\in\mathbb{R}^m}\quad &u\\ \text{subject to}\quad &r\odot l(x)+K_rd\leq\rho e,\\ &\|d\|_2\leq u,\qquad \rho\leq\rho_0. \end{aligned}

    The feasibility conditions distinguish whether the iteration can still improve preference alignment or needs to prioritize stationarity. If Pareto stationarity does not hold, or if the weighted Chebyshev constraints are feasible at ρ=ρ0\rho=\rho_0, this adjustment problem is feasible. Conversely, if it is infeasible, Pareto stationarity holds and the weighted Chebyshev constraints cannot be satisfied at ρ=ρ0\rho=\rho_0. When feasible, a solution with ρ<ρ0\rho<\rho_0 has only a trivial associated dual vector; the main WC-MGDA problem must then be solved to obtain a useful gradient. When ρ=ρ0\rho=\rho_0, the minimum feasible uu yields a solution recoverable for the main problem.

    The resulting strategy prioritizes reducing the primary weighted-loss objective by choosing small uu, and increases uu when needed to approach stationarity, while preventing the primary objective from worsening between iterations. For XWC-MGDA, the analogous adjustment uses the primary objective ρ−w⊤z\rho-w^\top z and its previous value p0p_0 as the non-increase criterion.

  4. Knowl 4 — XWC-MGDA alternates trade-off adjustment and gradient updates

    algorithm

    The algorithm takes a preference vector r∈R++mr\in\mathbb{R}_{++}^m, lower bound ww on the dual weights, reference loss vector b~\widetilde b, learning rate η\eta, maximum iteration count IMI_M, and residual tolerances τα,τd\tau_\alpha,\tau_d. It uses the gradient matrix G(x)G(x), the matrix KrK_r, the XWC-MGDA primal-dual problem, and the automatic adjustment problem. The paper specifies a random initial parameter vector, a large initial primary-objective bound, and a small positive initial uu, but does not prescribe numerical values for these initializations. At each iteration, dual and primal variables from the adjustment or main problem supply the preference-weighted gradient and stationarity residuals.

    Input: r, w, reference losses b_tilde, learning rate eta, maximum iterations I_M, tolerances tau_alpha and tau_d
    Initialize x randomly, p0 to a large value, and u to a small positive value
    For i = 1 to I_M:
        Compute G(x) and K_r(x)
        Solve the XWC-MGDA automatic adjustment problem with objective bound p0
        If the adjustment has a nontrivial solution:
            Update u and recover the associated primal-dual quantities alpha, d, and p
        Else:
            Solve the XWC-MGDA primal-dual problem at the current u to obtain alpha, d, and p
        If ||K_r(x) alpha||_2 <= tau_alpha and ||K_r(x) d||_2 <= tau_d:
            Return x
        Compute g = G(x)(r elementwise-multiplied by alpha)
        Update x = x - eta g
        Update p0 = p
    Return the final x after I_M iterations

    The parameter update is gradient descent on the model parameters using the preference-weighted combination of task gradients. The residual tests require both the weighted-gradient stationarity residual and the primal correction residual to be small.

  5. Knowl 5 — Synthetic experiments show preference alignment and smoother automatic tuning

    empirical result

    The synthetic two-objective experiment used n=20n=20 model parameters, a shared uniformly sampled initial point, and the highly nonconvex Pareto-front loss function

    l(x)=[1−exp⁡(−∥x−e/n∥22)1−exp⁡(−∥x+e/n∥22)],l(x)=\begin{bmatrix}1-\exp\bigl(-\|x-e/\sqrt n\|_2^2\bigr)\\1-\exp\bigl(-\|x+e/\sqrt n\|_2^2\bigr)\end{bmatrix},

    where e∈Rne\in\mathbb{R}^n is the all-ones vector. Methods were evaluated with preference directions generated at equal angular intervals; XWC-MGDA additionally used a reference point. Linear scalarization found only the convex part of the front, while PMTL produced approximate but imperfect preference alignment. EPO and WC-MGDA aligned solutions with preferences better than PMTL. With a reference point, XWC-MGDA explored a preferred portion of the front that dominated the reference in the illustrated comparison.

    The experiment also compared fixed u=0.001u=0.001 with automatic adjustment for WC-MGDA. Fixed small uu produced nonsmooth changes in the dual weights and primary objective and took more than 130 iterations to converge. Automatic adjustment kept the primary objective non-increasing, produced smoother weight changes, and converged in fewer than 100 iterations.

  6. Knowl 6 — Image-classification experiments test origin-based and reference-based preferences

    experimental setup

    The image experiments used Multi-MNIST, Multi-Fashion, and Multi-Fashion+MNIST, each with 120,000 training samples and 20,000 test samples. Each dataset had two classification tasks: classify the top-left image and classify the bottom-right image. All methods used the same LeNet architecture and shared random seeds; individual-task training served as a baseline. The comparison included XWC-MGDA, EPO, PMTL, and linear scalarization. In comparisons without a reference point, XWC-MGDA used w=10−6ew=10^{-6}e.

    With preference rays from the origin, XWC-MGDA produced test-loss solutions aligned with the requested preferences, especially for two of the three plotted directions. EPO was closer to one other direction, but XWC-MGDA generally matched or dominated competing gradient-based methods in the displayed loss and accuracy comparisons. The authors interpret the accuracy results as consistent with XWC-MGDA seeking Pareto-optimal rather than merely weakly Pareto-optimal points. In reference-based runs, XWC-MGDA used reference losses (0.5,0.5)(0.5,0.5), (0.7,0.7)(0.7,0.7), and (0.55,0.55)(0.55,0.55) for the three datasets, respectively; the resulting models followed preference rays from those references. This demonstrates the intended one-run search from a baseline, rather than repeated exploration to find a model that matches it on every task.

  7. Knowl 7 — Regression and music-label experiments reveal dataset-dependent method behavior

    empirical result

    For multi-target regression, the River Flow dataset predicts river flows 48 hours ahead at eight sites in the Mississippi River network. The experiment used a four-layer fully connected network, mean squared error per task, and 20 preference vectors. On test data, XWC-MGDA and EPO had similarly distributed relative loss profiles across the eight tasks and performed better than PMTL and linear scalarization; their task losses were also reported as similar to the single-task baseline and much lower than those of the other multi-task methods, particularly PMTL.

    For multi-label music classification, the Emotions dataset contains 593 songs with six emotion labels. The experiment used a four-layer fully connected network, sigmoid binary cross-entropy per task, and 50 preference vectors. All methods produced nearly uniform relative loss profiles, and none dominated the others. The authors take linear scalarization's strong performance as evidence that the Pareto front in this dataset may be convex.

  8. Knowl 8 — Reported hypervolumes favor XWC-MGDA in five of eight comparisons

    data/table

    The table reports mean hypervolumes for test-loss and test-accuracy fronts across three image datasets, plus test-loss fronts for River Flow and Emotions. Larger hypervolume was interpreted by the authors as stronger Pareto-front performance. For image classification, the reported means are over five trials with random seeds shared across methods; nadir points from single-task baselines were used as hypervolume reference points. The authors report that XWC-MGDA has the highest hypervolume in five of the eight comparisons and the second-highest in the other three.

    Dataset Metric XWC-MGDA EPO PMTL LinScalar
    MultiMNIST Loss 0.0142 0.0135 0.0141 0.0168
    MultiFashion Loss 0.0677 0.0602 0.0590 0.0675
    MultiFashion+MNIST Loss 0.0389 0.0376 0.0303 0.0363
    MultiMNIST Accuracy 0.0016 0.0015 0.0015 0.0019
    MultiFashion Accuracy 0.0099 0.0087 0.0087 0.0099
    MultiFashion+MNIST Accuracy 0.0052 0.0050 0.0040 0.0048
    River Loss 6.83E+28 5.93E+28 1.08E+28 6.64E+28
    Emotion Loss 0.000348 0.000366 0.000230 0.000258

    The comparisons show that the best method varies by dataset and metric: XWC-MGDA leads on several image and River Flow results, while linear scalarization has the largest reported hypervolume for MultiMNIST accuracy and ties XWC-MGDA on MultiFashion accuracy; EPO has the largest value for Emotions loss.

  9. Knowl 9 — Per-iteration computational cost is dominated by gradient Gram-matrix formation

    theoretical result

    For mm objectives and nn model parameters with n≫mn\gg m, the paper gives XWC-MGDA a computational complexity of O(m2n)O(m^2n) per iteration. The dominant operation is forming G(x)⊤G(x)G(x)^\top G(x) from the n×mn\times m gradient matrix. Computing the matrix square root required for KrK_r costs O(m3)O(m^3), and solving the second-order cone program costs O(m2.87)O(m^{2.87}) under the cited solver complexity. The stated total is therefore dominated by the gradient-matrix computation in the usual machine-learning regime n≫mn\gg m.

  10. Knowl 10 — The parameter update omits derivatives of the stationarity matrix

    limitation

    When updating model parameters, the method uses the gradient of the preference-weighted loss combination and does not differentiate through the parameter dependence of Kr(x)K_r(x). The authors state that including this dependence would require Hessian computation with respect to the model parameters, which they regard as prohibitive for the large-dimensional regime n≫1n\gg1. The proposed update therefore does not include those second-order terms.

Coverage note — Formal proofs and proof-only intermediate steps are omitted; the main feasibility and recovery conclusions they establish are included. Detailed dataset preprocessing is not reproduced because the paper delegates it to prior work.

References

  1. 1.Alizadeh, F. and Goldfarb, D. Second-order cone programming. MATHEMATICAL PROGRAMMING, 2001.
  2. 2.Boyd, S. and Vandenberghe, L. Convex Optimization. Cambridge University Press, 2004. ISBN 0521833787.
  3. 3.Carmel, D., Haramaty, E., Lazerson, A., and Lewin-Eytan, L. Multi-objective ranking optimization for product search using stochastic label aggregation. In Proceedings of The Web Conference, 2020.
  4. 4.Caruana, R. Multitask learning. Machine learning, 1997.
  5. 5.Desidéri, J.-A. Multiple-gradient descent algorithm (mgda) for multiobjective optimization. Comptes Rendus Mathematique, 350:313–318, 03 2012.
  6. 6.Fliege, J. and Svaiter, B. Steepest descent methods for multicriteria optimization. Mathematical Methods of Operations Research, 2000.
  7. 7.Frommer, A. and Hashemi, B. Verified computation of square roots of a matrix. SIAM Journal on Matrix Analysis and Applications, 2010.
  8. 8.Gong, C., Liu, X., and Liu, Q. Automatic and harmless regularization with constrained and lexicographic optimization: A dynamic barrier approach. In Advances in Neural Information Processing Systems, 2021.
  9. 9.Kaisa, M. Nonlinear Multiobjective Optimization, volume 12 of International Series in Operations Research & Management Science. Kluwer Academic Publishers, Boston, USA, 1999.
  10. 10.Kerenidis, I., Prakash, A., and Szilagyi, D. Quantum algorithms for Second-Order Cone Programming and Support Vector Machines. Quantum, 2021. doi: 10.22331/q-2021-04-08-427.
  11. 11.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998.
  12. 12.Lin, X., Chen, H., Pei, C., Sun, F., Xiao, X., Sun, H., Zhang, Y., Ou, W., and Jiang, P. A pareto-efficient algorithm for multiple objective optimization in e-commerce recommendation. In Proceedings of the 13th ACM Conference on Recommender Systems, 2019a.
  13. 13.Lin, X., Zhen, H.-L., Li, Z., Zhang, Q.-F., and Kwong, S. Pareto multi-task learning. In Advances in Neural Information Processing Systems 32. 2019b.
  14. 14.Lin, X., Yang, Z., Zhang, Q., and Kwong, S. Controllable pareto multi-task learning. arXiv preprint arXiv:2010.06313, 2020.
  15. 15.Liu, X., Tong, X., and Liu, Q. Profiling pareto front with multi-objective stein variational gradient descent. Advances in Neural Information Processing Systems, 2021.
  16. 16.Ma, P., Du, T., and Matusik, W. Efficient continuous pareto exploration in multi-task learning. In Proceedings of the 37th International Conference on Machine Learning, 2020.
  17. 17.Mahapatra, D. and Rajan, V. Multi-task learning with user preferences: Gradient descent with controlled ascent in pareto optimization. In Proceedings of the 37th International Conference on Machine Learning, 2020.
  18. 18.Matoušek, J. and Gärtner, B. Understanding and Using Linear Programming (Universitext). Springer-Verlag, Berlin, Heidelberg, 2006. ISBN 3540306978.
  19. 19.Momma, M., Garakani, A. B., and Sun, Y. Multi-objective relevance ranking. In Proceedings of the SIGIR 2019 Workshop on eCommerce, volume 2410 of CEUR Workshop Proceedings (preprint). CEUR-WS.org, 2019.
  20. 20.Momma, M., Garakani, A. B., Ma, N., and Sun, Y. Multi-objective ranking via constrained optimization. In Companion Proceedings of the The Web Conference, 2020.
  21. 21.Navon, A., Shamsian, A., Fetaya, E., and Chechik, G. Learning the pareto front with hypernetworks. In International Conference on Learning Representations, 2021.
  22. 22.Ruchte, M. and Grabocka, J. Scalable pareto front approximation for deep multi-objective learning. In Proceedings of the IEEE International Conference on Data Mining (ICDM), 2021.
  23. 23.Ruder, S. An overview of multi-task learning in deep neural networks, 2017.
  24. 24.Sabour, S., Frosst, N., and Hinton, G. E. In Advances in Neural Information Processing Systems, 2017.
  25. 25.Sener, O. and Koltun, V. Multi-task learning as multi-objective optimization. In Advances in Neural Information Processing Systems 31. 2018.
  26. 26.Spyromitros-Xioufis, E., Tsoumakas, G., Groves, W., and Vlahavas, I. Multi-target regression via input space expansion: treating targets as inputs. Machine Learning, 2016.
  27. 27.Trohidis, K., Tsoumakas, G., Kalliris, G., and Vlahavas, I. Multi-label classification of music by emotion. EURASIP Journal on Audio, Speech, and Music Processing, 2011.
  28. 28.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. ArXiv, abs/1708.07747, 2017.
  29. 29.Zhang, Q. and Li, H. Moea/d: A multiobjective evolutionary algorithm based on decomposition. IEEE Transactions on evolutionary computation, 2007.

Citation

MLA
Momma, M., et al. “A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity”. International Conference on Machine Learning, vol. 162, 2022, pp. 15895–907, https://proceedings.mlr.press/v162/momma22a.html.
APA
Momma, M., Dong, C., & Liu, J. (2022). A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity. International Conference on Machine Learning, 162, 15895–15907. https://proceedings.mlr.press/v162/momma22a.html
Chicago
Momma, M., C. Dong, and J. Liu. 2022. “A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity”. International Conference on Machine Learning 162: 15895–907. https://proceedings.mlr.press/v162/momma22a.html.
Harvard
Momma, M., Dong, C. and Liu, J. (2022) “A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity”, International Conference on Machine Learning. PMLR, pp. 15895–15907. Available at: https://proceedings.mlr.press/v162/momma22a.html.
Vancouver
1. Momma M, Dong C, Liu J (2022) A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity. In: International Conference on Machine Learning. PMLR, pp 15895–15907

BibTeX

@InProceedings{pmlr-v162-momma22a,
  title = 	 {A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity},
  author =       {Momma, Michinari and Dong, Chaosheng and Liu, Jia},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {15895--15907},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/momma22a/momma22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/momma22a.html},
  abstract = 	 {Multi-objective optimization (MOO) and multi-task learning (MTL) have gained much popularity with prevalent use cases such as production model development of regression / classification / ranking models with MOO, and training deep learning models with MTL. Despite the long history of research in MOO, its application to machine learning requires development of solution strategy, and algorithms have recently been developed to solve specific problems such as discovery of any Pareto optimal (PO) solution, and that with a particular form of preference. In this paper, we develop a novel and generic framework to discover a PO solution with multiple forms of preferences. It allows us to formulate a generic MOO / MTL problem to express a preference, which is solved to achieve both alignment with the preference and PO, at the same time. Specifically, we apply the framework to solve the weighted Chebyshev problem and an extension of that. The former is known as a method to discover the Pareto front, the latter helps to find a model that outperforms an existing model with only one run. Experimental results demonstrate not only the method achieves competitive performance with existing methods, but also it allows us to achieve the performance from different forms of preferences.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/