A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity
Michinari MommaChaosheng DongJia Liu
Develops a generic multi-objective learning framework based on Pareto stationarity that incorporates user preferences and extends weighted Chebyshev optimization to discover models outperforming existing baselines in a single training run.
Modern machine learning systems, such as search ranking, recommendation engines, and multi-task deep neural networks, frequently require optimizing multiple competing objectives simultaneously. In commercial environments, engineering teams routinely face the costly challenge of retraining models on fresh data while ensuring the new model does not degrade performance on any key metric relative to an existing baseline. Traditional multi-objective optimization techniques either struggle to align models with explicit business preferences or require extensive, manual trial-and-error exploration across all objectives, leading to development cycles that typically consume several days.
The article develops a unified mathematical framework for multi-objective and multi-task learning that simultaneously achieves optimal trade-off efficiency—known as Pareto optimality—and alignment with specific user-defined preferences. Specifically, it introduces two gradient-based algorithms: weighted Chebyshev multi-gradient descent and an extended version that explicitly explores improvements relative to an existing baseline model.
To demonstrate and validate this framework, the authors conducted empirical evaluations across both synthetic benchmarks and multiple real-world tasks. The experimental suite included two-task image classification across three separate datasets (MultiMNIST, Multi-Fashion, and Multi-Fashion+MNIST with 120,000 training images each), eight-target river flow regression across the Mississippi River network, and multi-class emotion classification in music. The proposed method was evaluated against established industry benchmarks, including linear scalarization, Pareto multi-task learning, and Exact Pareto Optimal search, measuring relative loss profiles and overall hypervolume coverage.
The findings show that the extended algorithm successfully identifies models that strictly outperform an existing reference model while directly following specified trade-off preferences. In practice, this reduces the required tuning exploration from scaling linearly with the number of objectives down to a single optimization run. The proposed method achieved the highest hypervolume metric in five out of eight evaluation settings and ranked second in the remaining three, demonstrating superior ability to dominate competing approaches. Furthermore, the algorithm’s automatic parameter-tuning mechanism ensured smooth convergence within 100 iterations on synthetic tests, outperforming fixed-parameter setups that required over 130 iterations.
These results provide immediate practical value for production engineering. By reducing the tuning overhead to a single run, organizations can substantially reduce computational costs, eliminate days of manual engineering labor, and de-risk model updates by guaranteeing that baseline performance is maintained or exceeded. Organizations facing frequent model retraining should consider adopting this framework to automate multi-task model updates and replace manual hyperparameter searches. Confidence in these results is high across complex, competing objectives; however, practitioners should note that the method operates on gradient-based models and does not provide an advantage on simpler problems where basic linear combinations already suffice.
- Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). It establishes the multi-task-as-multi-objective framing and Pareto-efficient gradient optimization that the source extends with preference-directed Chebyshev methods.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). It carries Pareto-optimal multi-objective learning into foundation-model alignment, using preference-conditioned trade-offs to extend the source’s optimization framework to deployment-time user preferences.
