Built independently by an author, for readers. Read the story and support ChapterPal

keyword

loss landscape evolution

Loss landscape evolution refers to the continuous changes in the geometric structure, curvature, and topology of a neural network loss surface as training progresses or data distributions shift. As a model updates its parameters, the multidimensional surface representing its error function undergoes structural transformations that alter properties such as gradient flow, Hessian eigenvalues, and the sharpness of local minima. Analyzing how the loss landscape evolves over time provides insight into optimization dynamics, helping to explain phenomena such as changes in network plasticity, convergence behavior, and a model capacity to generalize or continuously adapt to new information.

1 item

Understanding Plasticity in Neural Networks

Understanding Plasticity in Neural Networks

Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires, Razvan Pascanu, Will Dabney

OrganizationsGoogle

Why you should read this

Reveals that neural network plasticity loss stems primarily from unfavorable changes in loss curvature rather than unit saturation, providing practical architectural and optimization techniques like layer normalization to maintain continual learning capacity in deep reinforcement learning.

Plasticity, the ability of a neural network to quickly change its predictions in response to new information, is essential for the adaptability and robustness of deep reinforcement learning systems. Deep neural networks are known to lose plasticity over the course of training even in relatively simple learning problems, but the mechanisms driving this phenomenon are still poorly understood. This paper conducts a systematic empirical analysis into plasticity loss, with the goal of understanding the phenomenon mechanistically in order to guide the future development of targeted solutions. We find that loss of plasticity is deeply connected to changes in the curvature of the loss landscape, but that it often occurs in the absence of saturated units. Based on this insight, we identify a number of parameterization and optimization design choices which enable networks to better preserve plasticity over the course of training. We validate the utility of these findings on larger-scale RL benchmarks in the Arcade Learning Environment.

Added

2026-09-26