Built independently by an author, for readers. Read the story and support ChapterPal

keyword

gradient normalization

Gradient normalization is an optimization technique used in deep multi-task learning to balance the learning dynamics of multiple objectives by dynamically scaling the magnitudes of their respective gradients. When training a shared neural network on multiple tasks simultaneously, disparities in task difficulty, data scales, or loss functions often cause dominant tasks to overshadow others during backpropagation, leading to gradient interference and suboptimal performance on underrepresented tasks. Gradient normalization resolves this imbalance by monitoring gradient norms and task training speeds, then adaptively adjusting the loss weights so that all tasks train at comparable rates. By standardizing gradient contributions across tasks, this approach prevents individual task dominance, minimizes the need for manual hyperparameter tuning of static loss weights, and improves overall multi-task generalization.

4 items

Towards Understanding Sharpness-Aware Minimization

Towards Understanding Sharpness-Aware Minimization

Maksym Andriushchenko, Nicolas Flammarion

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Explains the generalization benefits of Sharpness-Aware Minimization by analyzing its implicit bias and proving its convergence with stochastic gradients, revealing why mini-batch sharpness yields superior performance over standard gradient descent.

Sharpness-Aware Minimization (SAM) is a recent training method that relies on worst-case weight perturbations which significantly improves generalization in various settings. We argue that the existing justifications for the success of SAM which are based on a PAC-Bayes generalization bound and the idea of convergence to flat minima are incomplete. Moreover, there are no explanations for the success of using m-sharpness in SAM which has been shown as essential for generalization. To better understand this aspect of SAM, we theoretically analyze its implicit bias for diagonal linear networks. We prove that SAM always chooses a solution that enjoys better generalization properties than standard gradient descent for a certain class of problems, and this effect is amplified by using m-sharpness. We further study the properties of the implicit bias on non-linear networks empirically, where we show that fine-tuning a standard model with SAM can lead to significant generalization improvements. Finally, we provide convergence results of SAM for non-convex objectives when used with stochastic gradients. We illustrate these results empirically for deep networks and discuss their relation to the generalization behavior of SAM. The code of our experiments is available at https://github.com/tml-epfl/understanding-sam.

Added

2026-09-26

XAI Beyond Classification: Interpretable Neural Clustering

XAI Beyond Classification: Interpretable Neural Clustering

Xi Peng, Yunfan Li, Ivor W. Tsang, Hongyuan Zhu, Jiancheng Lv, Joey Tianyi Zhou

Why you should read this

Proposes an intrinsically explainable neural network that reformulates discrete k-means into a differentiable layer, enabling end-to-end parallel optimization, online clustering on data streams, and provable convergence without relying on post-hoc interpretations.

In this paper, we study two challenging problems in explainable AI (XAI) and data clustering. The first is how to directly design a neural network with inherent interpretability, rather than giving post-hoc explanations of a black-box model. The second is implementing discrete k-means with a differentiable neural network that embraces the advantages of parallel computing, online clustering, and clustering-favorable representation learning. To address these two challenges, we design a novel neural network, which is a differentiable reformulation of the vanilla k-means, called inTerpretable nEuraL cLustering (TELL). Our contributions are threefold. First, to the best of our knowledge, most existing XAI works focus on supervised learning paradigms. This work is one of the few XAI studies on unsupervised learning, in particular, data clustering. Second, TELL is an interpretable, or the so-called intrinsically explainable and transparent model. In contrast, most existing XAI studies resort to various means for understanding a black-box model with post-hoc explanations. Third, from the view of data clustering, TELL possesses many properties highly desired by k-means, including but not limited to online clustering, plug-and-play module, parallel computing, and provable convergence. Extensive experiments show that our method achieves superior performance comparing with 14 clustering approaches on three challenging data sets. The source code could be accessed at www.pengxi.me.

Added

2026-09-26

GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, Andrew Rabinovich

OrganizationsMagic Leap

Why you should read this

Proposes GradNorm, an adaptive loss-balancing method that dynamically scales gradient magnitudes during training, eliminating expensive loss-weight grid searches while improving multitask performance and reducing overfitting across diverse architectures.

Deep multitask networks, in which one neural network produces multiple predictive outputs, can offer better speed and performance than their single-task counterparts but are challenging to train properly. We present a gradient normalization (GradNorm) algorithm that automatically balances training in deep multitask models by dynamically tuning gradient magnitudes. We show that for various network architectures, for both regression and classification tasks, and on both synthetic and real datasets, GradNorm improves accuracy and reduces overfitting across multiple tasks when compared to single-task networks, static baselines, and other adaptive multitask loss balancing techniques. GradNorm also matches or surpasses the performance of exhaustive grid search methods, despite only involving a single asymmetry hyperparameter α\alpha. Thus, what was once a tedious search process that incurred exponentially more compute for each task added can now be accomplished within a few training runs, irrespective of the number of tasks. Ultimately, we will demonstrate that gradient manipulation affords us great control over the training dynamics of multitask networks and may be one of the keys to unlocking the potential of multitask learning.

Added

2026-09-16

Gradient Surgery for Multi-Task Learning

Gradient Surgery for Multi-Task Learning

Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, Chelsea Finn

OrganizationsGoogleStanford UniversityUniversity of California Berkeley

Why you should read this

Introduces a model-agnostic optimization technique that eliminates destructive interference between tasks by projecting conflicting gradients onto each other's normal planes, significantly improving performance across multi-task supervised and reinforcement learning benchmarks.

While deep learning and deep reinforcement learning (RL) systems have demonstrated impressive results in domains such as image classification, game playing, and robotic control, data efficiency remains a major challenge. Multi-task learning has emerged as a promising approach for sharing structure across multiple tasks to enable more efficient learning. However, the multi-task setting presents a number of optimization challenges, making it difficult to realize large efficiency gains compared to learning tasks independently. The reasons why multi-task learning is so challenging compared to single-task learning are not fully understood. In this work, we identify a set of three conditions of the multi-task optimization landscape that cause detrimental gradient interference, and develop a simple yet general approach for avoiding such interference between task gradients. We propose a form of gradient surgery that projects a task's gradient onto the normal plane of the gradient of any other task that has a conflicting gradient. On a series of challenging multi-task supervised and multi-task RL problems, this approach leads to substantial gains in efficiency and performance. Further, it is model-agnostic and can be combined with previously-proposed multi-task architectures for enhanced performance.

Added

2026-09-16