Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and Momentum
Zeke XieXinrui WangHuishuai ZhangIssei SatoMasashi Sugiyama
Explains why Adam generalizes worse than SGD through diffusion theory and proposes Adaptive Inertia, a new optimizer that adapts momentum instead of learning rates to escape saddle points quickly while retaining SGD-like flat minima selection.
Training deep neural networks requires optimization algorithms that are both fast and capable of producing models that perform accurately on unseen real-world data. While the widely used Adam optimizer accelerates training by adjusting individual learning rates and using momentum, it often yields models that generalize worse than those trained with standard Stochastic Gradient Descent (SGD). Conversely, SGD achieves superior accuracy and generalization by finding flatter loss minima, but it suffers from slower training speeds when traversing difficult optimization obstacles, such as saddle points.
The article aims to explain the mathematical mechanisms behind this generalization gap and introduce an alternative optimization framework that combines rapid training convergence with strong model generalization.
Using continuous-time diffusion theory and physical motion equations, the analysis separates the distinct roles of adaptive learning rates and momentum during training. The authors demonstrate how these mechanisms influence the ability to escape saddle points and select flat minima. To validate the findings, extensive empirical benchmarks were conducted across major image classification datasets, including CIFAR-10, CIFAR-100, and ImageNet, as well as language modeling benchmarks using standard architectures such as ResNet, VGG, DenseNet, and Long Short-Term Memory networks.
The investigation yields four primary findings. First, momentum introduces a physical drift effect that accelerates passage through saddle points without hindering the selection of flat minima. Second, adaptive learning rates allow rapid escape from saddle points but significantly degrade generalization because they weaken the algorithm's sensitivity to loss curvature, causing models to settle in sharper, poorer-performing minima. Third, the proposed optimizer, termed Adaptive Inertia (Adai), adjusts momentum rather than learning rates on a per-parameter basis, provably preserving SGD-level flat minima selection while accelerating saddle-point traversal. Fourth, empirical benchmarks show that Adai and its weight-decay variant consistently outperform SGD, Adam, and numerous Adam variants in test accuracy, achieving top-1 error reductions of approximately 0.3% over fine-tuned SGD and nearly 4% over Adam on ImageNet ResNet50.
These findings demonstrate that the performance trade-off between fast training and strong generalization is not fundamental. Practitioners can eliminate the need to compromise between Adam's rapid convergence and SGD's superior accuracy by shifting adaptivity from step sizes to momentum parameters. Furthermore, Adai demonstrates greater tolerance across varying learning rate and regularization settings, which reduces computational tuning costs and deployment risks across machine learning workflows.
Engineering and research teams should consider adopting Adaptive Inertia methods, particularly Adai with decoupled weight decay, as an effective drop-in replacement for Adam and SGD in deep learning pipelines. When evaluating new optimization approaches, practitioners must also ensure that baseline comparisons utilize properly tuned regularization parameters to avoid misleading performance conclusions.
The theoretical derivations rely on standard physical approximations, including quasi-equilibrium and low-noise assumptions around critical points, which accurately reflect typical minibatch training regimes. Although the empirical testing covers standard computer vision and recurrent language modeling benchmarks, further large-scale validation across emerging architectures, such as modern large language models and vision transformers, is warranted to confirm broader applicability.
- Paper: Adam: A Method for Stochastic Optimization, Diederik P. Kingma et al. (2015). It introduces the Adam optimizer combining adaptive learning rates and momentum, which the source directly analyzes and decomposes to resolve its generalization deficit.
- Paper: Decoupled Weight Decay Regularization, Ilya Loshchilov et al. (2019). It establishes the decoupling of weight decay from adaptive gradient updates, a foundational regularization concept directly incorporated into the Adai optimizer variants evaluated in the source.
- Paper: On the Convergence of Adam and Beyond, Sashank J. Reddi et al. (2018). It provides critical theoretical foundations regarding the convergence issues and effective step sizes of Adam-style adaptive coordinate methods that motivate the source's investigation.
- Paper: Identifying and attacking the saddle point problem in high-dimensional non-convex optimization, Yann Dauphin et al. (2014). It identifies saddle points and non-convex plateaus as primary optimization obstacles in deep learning, providing essential context for the saddle-point escape dynamics studied in the source.
- Paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization, John Duchi et al. (2011). It introduces coordinate-wise adaptive learning rates via historical gradient accumulation, establishing the foundational mechanism whose generalization effects the source reassesses.
- Paper: A Differential Equation for Modeling Nesterov's Accelerated Gradient Method: Theory and Insights, Weijie Su et al. (2014). It introduces the continuous-time physical differential equation framework for modeling accelerated momentum methods, providing the theoretical modeling approach utilized in the source.
- Paper: On the Variance of the Adaptive Learning Rate and Beyond, Liyuan Liu et al. (2019). It analyzes the variance and instability of adaptive learning rates across training phases, offering crucial background on the limitations of adaptive step sizes examined by the source.
- Paper: Train faster, generalize better: Stability of stochastic gradient descent, Moritz Hardt et al. (2015). It formalizes the relationship between optimization trajectory stability and generalization in stochastic gradient descent, establishing the benchmark behavior the source aims to preserve.
- Paper: The alignment property of SGD noise and how it helps select flat minima: A stability analysis, Lei Wu et al. (2022). It advances the source's investigation into flat minima selection by providing a dynamical stability analysis of how gradient noise aligns with loss curvature to favor broad basins.
- Paper: Towards Understanding Sharpness-Aware Minimization, Maksym Andriushchenko et al. (2022). It explores the theoretical and empirical foundations of seeking flat minima in overparameterized networks, directly complementing the source's focus on curvature sensitivity and generalization.
- Paper: Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization, Xingxuan Zhang et al. (2023). It extends the objective of finding flat, generalizable loss minima by introducing first-order curvature penalties using gradient norms.
- Paper: DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule, Maor Ivgi et al. (2023). It offers an alternative approach to overcoming learning rate sensitivity by developing parameter-free dynamic step size schedules that maintain SGD-like performance.
- Paper: Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs, Sagnik Mukherjee et al. (2026). It empirically extends the investigation of disentangling adaptive rates and momentum by evaluating whether SGD without adaptive mechanisms can outperform AdamW in complex LLM reinforcement learning.
- Paper: Adam-mini: Use Fewer Learning Rates To Gain More, Yushun Zhang et al. (2025). It develops a structural, block-wise adaptive learning rate approach based on network Hessian structure to reduce optimizer memory overhead while preserving performance.
