Built independently by an author, for readers. Read the story and support ChapterPal

keyword

optimization dynamics

Optimization dynamics refers to the step-by-step trajectory, behavior, and time evolution of model parameters and objective values as an iterative algorithm navigates a loss landscape during training. In machine learning, this concept encompasses the continuous interaction between mathematical update rules, such as gradient-based methods, and the geometric features of the underlying objective function, including curvature, gradients, and local minima. Analyzing optimization dynamics provides insight into training stability, convergence speed, variance across updates, and the precise mechanisms through which algorithmic adjustments shape internal representations and overall model performance over time.

4 items

Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning

Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning

Luckeciano Carvalho Melo, Alessandro Abate, Yarin Gal

OrganizationsUniversity of Oxford

Why you should read this

Proposes Curvature-Aware Policy Optimization (CAPO), a framework that tracks second-order optimization curvature to mask destabilizing updates, achieving theoretical monotonic improvement guarantees and up to a 30x sample-efficiency gain over GRPO on LLM reasoning benchmarks.

Reinforcement Learning, particularly through policy gradient methods, has played a central role in enabling reasoning capabilities of Large Language Models. However, the optimization stability of policy gradients in this setting remains understudied. As a result, existing implementations often resort to conservative hyperparameter choices to ensure stability, which requires more training samples and increases computational costs. Hence, developing models for reliably tracking the underlying optimization dynamics and leveraging them into training enables more sample-efficient regimes and further unleashes scalable post-training. We address this gap by formalizing the stochastic optimization problem of policy gradients with explicit consideration of second-order geometry. We propose a tractable computational framework that tracks and leverages curvature information during policy updates. We further employ this framework to design interventions in the optimization process through data selection. The resultant algorithm, Curvature-Aware Policy Optimization (CAPO), identifies samples that contribute to unstable updates and masks them out. Theoretically, we establish monotonic improvement guarantees under realistic assumptions. On standard math reasoning benchmarks, we empirically show that CAPO ensures stable updates under aggressive learning regimes where baselines catastrophically fail. With minimal intervention (rejecting fewer than 8% of tokens), CAPO achieves up to 30x improvement in sample efficiency over standard GRPO for LLM reasoning.

Added

2026-10-04

Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Do Current Multi-Task Optimization Methods in Deep Learning Even Help?

Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg, Orhan Firat

OrganizationsGoogle

Why you should read this

Demonstrates through large-scale empirical experiments that complex multi-task optimization algorithms fail to outperform properly tuned scalarized baselines, revealing flawed evaluation protocols in prior literature and offering practical strategies for reliable multi-task model training.

Recent research has proposed a series of specialized optimization algorithms for deep multi-task models. It is often claimed that these multi-task optimization (MTO) methods yield solutions that are superior to the ones found by simply optimizing a weighted average of the task losses. In this paper, we perform large-scale experiments on a variety of language and vision tasks to examine the empirical validity of these claims. We show that, despite the added design and computational complexity of these algorithms, MTO methods do not yield any performance improvements beyond what is achievable via traditional optimization approaches. We highlight alternative strategies that consistently yield improvements to the performance profile and point out common training pitfalls that might cause suboptimal results. Finally, we outline challenges in reliably evaluating the performance of MTO algorithms and discuss potential solutions.

Added

2026-09-26

Understanding Plasticity in Neural Networks

Understanding Plasticity in Neural Networks

Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires, Razvan Pascanu, Will Dabney

OrganizationsGoogle

Why you should read this

Reveals that neural network plasticity loss stems primarily from unfavorable changes in loss curvature rather than unit saturation, providing practical architectural and optimization techniques like layer normalization to maintain continual learning capacity in deep reinforcement learning.

Plasticity, the ability of a neural network to quickly change its predictions in response to new information, is essential for the adaptability and robustness of deep reinforcement learning systems. Deep neural networks are known to lose plasticity over the course of training even in relatively simple learning problems, but the mechanisms driving this phenomenon are still poorly understood. This paper conducts a systematic empirical analysis into plasticity loss, with the goal of understanding the phenomenon mechanistically in order to guide the future development of targeted solutions. We find that loss of plasticity is deeply connected to changes in the curvature of the loss landscape, but that it often occurs in the absence of saturated units. Based on this insight, we identify a number of parameterization and optimization design choices which enable networks to better preserve plasticity over the course of training. We validate the utility of these findings on larger-scale RL benchmarks in the Arcade Learning Environment.

Added

2026-09-26

How Transformers Learn Causal Structure with Gradient Descent

How Transformers Learn Causal Structure with Gradient Descent

Eshaan Nichani, Alex Damian, Jason D. Lee

OrganizationsPrinceton University

Why you should read this

Proves that gradient descent enables two-layer transformers to learn latent causal graphs and form induction heads by showing that the attention matrix gradients naturally track token-level mutual information.

The incredible success of transformers on sequence modeling tasks can be largely attributed to the self-attention mechanism, which allows information to be transferred between different parts of a sequence. Self-attention allows transformers to encode causal structure which makes them particularly suitable for sequence modeling. However, the process by which transformers learn such causal structure via gradient-based training algorithms remains poorly understood. To better understand this process, we introduce an in-context learning task that requires learning latent causal structure. We prove that gradient descent on a simplified two-layer transformer learns to solve this task by encoding the latent causal graph in the first attention layer. The key insight of our proof is that the gradient of the attention matrix encodes the mutual information between tokens. As a consequence of the data processing inequality, the largest entries of this gradient correspond to edges in the latent causal graph. As a special case, when the sequences are generated from in-context Markov chains, we prove that transformers learn an induction head (Olsson et al., 2022). We confirm our theoretical findings by showing that transformers trained on our in-context learning task are able to recover a wide variety of causal structures.

Added

2026-09-26