Adam-mini: Use Fewer Learning Rates To Gain More
Yushun ZhangCongliang ChenZiniu LiTian DingChenwei WuDiederik P. KingmaYinyu YeZhi-Quan LuoRuoyu Sun
Proposes Adam-mini, a lightweight optimizer that cuts AdamW's memory footprint in half and accelerates large language model training by assigning single learning rates to parameter blocks based on Hessian structure while matching or exceeding task performance.
Training modern large language models requires vast amounts of memory, creating severe hardware constraints and high operational costs. The standard training optimizer, Adam (or AdamW), consumes substantial memory because it maintains two tracking states for every parameter in the model. This high memory footprint forces distributed training setups to rely on frequent hardware communication and offloading, which degrades computational throughput and increases overall training time.
The main objective of the article is to evaluate and demonstrate a new optimization algorithm called Adam-mini. The authors aim to significantly reduce optimizer memory consumption while maintaining or improving model training performance across diverse neural network architectures.
The authors conducted a high-level empirical and theoretical analysis grounded in the curvature structure of neural networks. By analyzing second-order matrix structures (Hessians) across basic neural networks and standard Transformers, the researchers identified that parameter interactions group naturally into near-block-diagonal structures. Rather than tracking an individual learning rate for every single parameter, the proposed Adam-mini method partitions model parameters into structural blocks—such as attention heads, output neurons, or token embeddings—and assigns a single averaged learning rate state to each block. The team validated this approach by pre-training language models ranging from 39 million to 13 billion parameters, conducting fine-tuning and reinforcement learning alignment tasks, and benchmarking computer vision and graph neural networks.
The evaluation revealed several key findings. First, Adam-mini reduces the second-order momentum states by over 99.9%, achieving a 50% overall reduction in optimizer memory footprint relative to AdamW. Second, the algorithm matches or exceeds the training and validation performance of AdamW across all language model sizes, downstream fine-tuning tasks, and non-language domains, whereas competing lightweight optimizers suffered performance drops or instability. Third, the lowered memory pressure enables larger batch sizes per processor and substantially cuts inter-processor communication overhead; for example, pre-training a 7-billion-parameter model on two computing cards yielded a 49.6% throughput increase and a 33.1% reduction in total wall-clock training time. Finally, the optimizer exhibited predictable scaling behavior consistent with standard compute-optimal laws without requiring specialized hyperparameter tuning.
These findings indicate that current standard training setups carry massive memory redundancies that can be safely eliminated by aligning optimizer design with model architecture. In practice, adopting this method directly lowers the operational and infrastructure costs of training frontier models, accelerates experimental cycles, and allows larger models to train on constrained hardware clusters without complex manual tuning.
Organizations training transformer-based models should consider piloting Adam-mini as a drop-in replacement for AdamW in their pre-training and fine-tuning pipelines. Because Adam-mini functions reliably with standard baseline hyperparameters, teams can switch optimizers without expensive configuration sweeps. However, engineering teams must ensure they implement the architecture-specific parameter partitioning rules described in the article, as naive or default grouping strategies can lead to training instability.
The primary limitation of the study is that the averaged learning rate calculation within each structural block is a practical heuristic rather than a mathematically proven optimal value, leaving potential room for further theoretical refinements. Additionally, the largest model evaluated during pre-training was 13 billion parameters. While scaling trends strongly support extrapolation, stakeholders deploying models at massive hundred-billion-parameter scales should first validate the optimizer on small-scale proxy runs to verify stability.
- Paper: Adam: A Method for Stochastic Optimization, Diederik P. Kingma et al. (2015). Introduces the standard Adam optimization algorithm and its per-parameter second-moment tracking ($v$), which Adam-mini directly analyzes and compresses.
- Paper: Decoupled Weight Decay Regularization, Ilya Loshchilov et al. (2019). Presents the AdamW optimizer with decoupled weight decay, which serves as the primary performance and memory baseline that Adam-mini aims to match and improve.
- Paper: Adafactor: Adaptive Learning Rates with Sublinear Memory Cost, Noam Shazeer et al. (2018). Pioneers sublinear-memory adaptive optimization by factoring the second-moment accumulator, providing essential foundational context for reducing optimizer state memory in Transformers.
- Paper: ZeRO: Memory optimizations Toward Training Trillion Parameter Models, Samyam Rajbhandari et al. (2020). Analyzes optimizer state memory bottlenecks and communication overheads in distributed LLM training, motivating the throughput and resource benefits targeted by Adam-mini.
- Paper: GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection, Jiawei Zhao et al. (2024). Explores memory-efficient LLM pre-training via low-rank projections of gradient states, representing an alternative paradigm for cutting optimizer memory footprint.
- Paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization, John Duchi et al. (2011). Establishes coordinate-wise adaptive subgradient methods that underpin modern adaptive learning rate mechanisms.
- Paper: An overview of gradient descent optimization algorithms, Sebastian Ruder (2016). Provides a comprehensive overview of gradient descent and adaptive learning rate optimizers, framing how per-parameter learning rate scaling functions.
- Paper: Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs, Sagnik Mukherjee et al. (2026). Investigates whether adaptive per-parameter learning rates are strictly necessary in post-training LLM reinforcement learning, extending the enquiry into optimizer state redundancy initiated by Adam-mini.
