Built independently by an author, for readers. Read the story and support ChapterPal

keyword

hyper-parameter tuning

Hyperparameter tuning is the process of selecting the optimal configuration of external settings that govern the learning behavior and structure of a machine learning model. Unlike internal model parameters, such as weights and biases, which the model learns directly from training data, hyperparameters must be specified beforehand. Common examples include the learning rate, batch size, number of training epochs, and regularization strength. Practitioners systematically explore different combinations of these values using strategies such as grid search, random search, or Bayesian optimization to identify the settings that yield the highest evaluation performance and prevent underfitting or overfitting.

1 item

Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg, Orhan Firat

OrganizationsGoogle

Why you should read this

Demonstrates through large-scale empirical experiments that complex multi-task optimization algorithms fail to outperform properly tuned scalarized baselines, revealing flawed evaluation protocols in prior literature and offering practical strategies for reliable multi-task model training.

Recent research has proposed a series of specialized optimization algorithms for deep multi-task models. It is often claimed that these multi-task optimization (MTO) methods yield solutions that are superior to the ones found by simply optimizing a weighted average of the task losses. In this paper, we perform large-scale experiments on a variety of language and vision tasks to examine the empirical validity of these claims. We show that, despite the added design and computational complexity of these algorithms, MTO methods do not yield any performance improvements beyond what is achievable via traditional optimization approaches. We highlight alternative strategies that consistently yield improvements to the performance profile and point out common training pitfalls that might cause suboptimal results. Finally, we outline challenges in reliably evaluating the performance of MTO algorithms and discuss potential solutions.

Added

2026-09-26