Built independently by an author, for readers. Read the story and support ChapterPal

keyword

deep kernel shaping

Deep kernel shaping is a neural network design and initialization methodology that adjusts the properties of a network at initialization to ensure stable signal propagation and enable fast, reliable optimization. By applying calibrated transformations to activation functions, setting precise weight initializations, and introducing minor structural adjustments, the method regulates how correlations and variances evolve across successive layers. This control prevents signal propagation pathologies, such as exploding or vanishing gradients and rank collapse, while preserving the original functional capacity of the model. As a result, deep kernel shaping allows very deep architectures to train efficiently and maintain plasticity throughout learning, even in feedforward networks that lack conventional stabilizing components such as normalization layers or skip connections.

1 item

Understanding Plasticity in Neural Networks

Understanding Plasticity in Neural Networks

Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires, Razvan Pascanu, Will Dabney

OrganizationsGoogle

Why you should read this

Reveals that neural network plasticity loss stems primarily from unfavorable changes in loss curvature rather than unit saturation, providing practical architectural and optimization techniques like layer normalization to maintain continual learning capacity in deep reinforcement learning.

Plasticity, the ability of a neural network to quickly change its predictions in response to new information, is essential for the adaptability and robustness of deep reinforcement learning systems. Deep neural networks are known to lose plasticity over the course of training even in relatively simple learning problems, but the mechanisms driving this phenomenon are still poorly understood. This paper conducts a systematic empirical analysis into plasticity loss, with the goal of understanding the phenomenon mechanistically in order to guide the future development of targeted solutions. We find that loss of plasticity is deeply connected to changes in the curvature of the loss landscape, but that it often occurs in the absence of saturated units. Based on this insight, we identify a number of parameterization and optimization design choices which enable networks to better preserve plasticity over the course of training. We validate the utility of these findings on larger-scale RL benchmarks in the Arcade Learning Environment.

Added

2026-09-26