Adam-mini: Use Fewer Learning Rates To Gain More

Yushun ZhangCongliang ChenZiniu LiTian DingChenwei WuDiederik P. KingmaYinyu YeZhi-Quan LuoRuoyu Sun

article2025ICLR159 citations

Proposes Adam-mini, a lightweight optimizer that cuts AdamW's memory footprint in half and accelerates large language model training by assigning single learning rates to parameter blocks based on Hessian structure while matching or exceeding task performance.

Listen

Training modern large language models requires vast amounts of memory, creating severe hardware constraints and high operational costs. The standard training optimizer, Adam (or AdamW), consumes substantial memory because it maintains two tracking states for every parameter in the model. This high memory footprint forces distributed training setups to rely on frequent hardware communication and offloading, which degrades computational throughput and increases overall training time.

The main objective of the article is to evaluate and demonstrate a new optimization algorithm called Adam-mini. The authors aim to significantly reduce optimizer memory consumption while maintaining or improving model training performance across diverse neural network architectures.

The authors conducted a high-level empirical and theoretical analysis grounded in the curvature structure of neural networks. By analyzing second-order matrix structures (Hessians) across basic neural networks and standard Transformers, the researchers identified that parameter interactions group naturally into near-block-diagonal structures. Rather than tracking an individual learning rate for every single parameter, the proposed Adam-mini method partitions model parameters into structural blocks—such as attention heads, output neurons, or token embeddings—and assigns a single averaged learning rate state to each block. The team validated this approach by pre-training language models ranging from 39 million to 13 billion parameters, conducting fine-tuning and reinforcement learning alignment tasks, and benchmarking computer vision and graph neural networks.

The evaluation revealed several key findings. First, Adam-mini reduces the second-order momentum states by over 99.9%, achieving a 50% overall reduction in optimizer memory footprint relative to AdamW. Second, the algorithm matches or exceeds the training and validation performance of AdamW across all language model sizes, downstream fine-tuning tasks, and non-language domains, whereas competing lightweight optimizers suffered performance drops or instability. Third, the lowered memory pressure enables larger batch sizes per processor and substantially cuts inter-processor communication overhead; for example, pre-training a 7-billion-parameter model on two computing cards yielded a 49.6% throughput increase and a 33.1% reduction in total wall-clock training time. Finally, the optimizer exhibited predictable scaling behavior consistent with standard compute-optimal laws without requiring specialized hyperparameter tuning.

These findings indicate that current standard training setups carry massive memory redundancies that can be safely eliminated by aligning optimizer design with model architecture. In practice, adopting this method directly lowers the operational and infrastructure costs of training frontier models, accelerates experimental cycles, and allows larger models to train on constrained hardware clusters without complex manual tuning.

Organizations training transformer-based models should consider piloting Adam-mini as a drop-in replacement for AdamW in their pre-training and fine-tuning pipelines. Because Adam-mini functions reliably with standard baseline hyperparameters, teams can switch optimizers without expensive configuration sweeps. However, engineering teams must ensure they implement the architecture-specific parameter partitioning rules described in the article, as naive or default grouping strategies can lead to training instability.

The primary limitation of the study is that the averaged learning rate calculation within each structural block is a practical heuristic rather than a mathematically proven optimal value, leaving potential room for further theoretical refinements. Additionally, the largest model evaluated during pre-training was 13 billion parameters. While scaling trends strongly support extrapolation, stakeholders deploying models at massive hundred-billion-parameter scales should first validate the optimizer on small-scale proxy runs to verify stability.

Cover for Adam-mini: Use Fewer Learning Rates To Gain More

Abstract

We propose Adam-mini, an optimizer that achieves on par or better performance than AdamW with 50% less memory footprint. Adam-mini reduces memory by cutting down the learning rate resources in Adam (i.e., 1/v1/\sqrt{v}). By investigating the Hessian structure of neural nets, we find Adam's vv might not function at its full potential as effectively as we expected. We find that ≥\geq 99.9% of these learning rates in vv could be harmlessly removed if we (1) carefully partition the parameters into blocks following our new principle on Hessian structure; (2) assign a single but good learning rate to each parameter block. We then provide one simple way to find good learning rates and propose Adam-mini. Empirically, we verify that Adam-mini performs on par or better than AdamW on various language models sized from 39M to 13B for pre-training, supervised fine-tuning, and RLHF. The reduced memory footprint of Adam-mini also alleviates communication overheads among GPUs, thereby increasing throughput. For instance, Adam-mini achieves 49.6% higher throughput than AdamW when pre-training Llama 2-7B on 2×2\times A800-80GB GPUs, which saves 33% wall-clock time for pre-training.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Motivations and Observations
  • 2.2 Proposed Method: Adam-mini
  • 2.3 Principle for the Partition Strategy
  • 2.4 Some Characteristics of Adam-mini and Discussions
  • 3 Experiments
  • 3.1 Pre-training
  • 3.2 Scaling Laws of Adam-mini
  • 3.3 Supervised Fine-tuning and RLHF
  • 3.4 Detailed Comparison with Adafactor
  • 4 Concluding Remarks
  • References
  • A Related works
  • B The Complete Form of Adam-mini
  • C More Discussions
  • D More Experimental Results
  • D.1 More Results for Motivation
  • D.2 Ablation Studies on the Design of Adam-mini
  • D.3 More Results on the Scaling Law Experiments
  • D.4 GPT-4 evaluation score of SFT and RLHF
  • D.5 Non-LLM Tasks
  • D.6 More Discussions on the Partition Strategies of Value
  • D.7 Detailed Comparison with Adafactor
  • D.8 Detailed Comparison with Lion
  • D.9 Additional Findings on GPT-2-330M
  • D.10 Combining with LoRA
  • D.11 Sample Responses from LLMs trained by Adam-mini
  • E Some Preliminary Results
  • E.1 Preliminaries on Adam, AdamW and LAMB
  • E.2 Preliminary results in (Zhang et al., 2024)
  • F Experimental Details
  • F.1 Training configurations for Section
  • F.2 Detailed Setup for Other Experiments

Knowls

  1. Knowl 1 — Adam-mini Optimizer Algorithm

    algorithm

    Adam-mini is an adaptive first-order optimizer that achieves comparable or superior performance to AdamW while using 50%50\% less optimizer state memory. It achieves this by assigning a single scalar learning rate (second-order momentum scalar) to each parameter block defined by the Hessian matrix's dense sub-blocks, rather than tracking an individual second-order momentum value per coordinate.

    Input: Model parameters θ0\theta_0, learning rate schedule ηt\eta_t, weight decay coefficient λ\lambda, hyperparameters β1,β2∈[0,1)\beta_1, \beta_2 \in [0, 1), stability constant ϵ>0\epsilon > 0, total training steps TT.
    Partition model parameters into a collection of blocks B={p1,p2,…,pB}\mathcal{B} = \{p_1, p_2, \dots, p_B\} according to the Hessian dense sub-block structure.
    Initialize first-order momentum tensors m0(b)=0m_0^{(b)} = 0 matching the shape of pbp_b for all b∈{1,…,B}b \in \{1, \dots, B\}.
    Initialize second-order momentum scalars v0(b)=0v_0^{(b)} = 0 for all b∈{1,…,B}b \in \{1, \dots, B\}.
    for step t=1t = 1 to TT do
        for each parameter block pb∈Bp_b \in \mathcal{B} with index b∈{1,…,B}b \in \{1, \dots, B\} do
            gt(b)=∇pbL(θt−1)g_t^{(b)} = \nabla_{p_b} \mathcal{L}(\theta_{t-1})
            pb=pb−ηt⋅λ⋅pbp_b = p_b - \eta_t \cdot \lambda \cdot p_b
            mt(b)=β1mt−1(b)+(1−β1)gt(b)m_t^{(b)} = \beta_1 m_{t-1}^{(b)} + (1 - \beta_1) g_t^{(b)}
            m^t(b)=mt(b)1−β1t\hat{m}_t^{(b)} = \frac{m_t^{(b)}}{1 - \beta_1^t}
            vt(b)=β2vt−1(b)+(1−β2)mean(gt(b)⊙gt(b))v_t^{(b)} = \beta_2 v_{t-1}^{(b)} + (1 - \beta_2) \text{mean}\left(g_t^{(b)} \odot g_t^{(b)}\right)
            v^t(b)=vt(b)1−β2t\hat{v}_t^{(b)} = \frac{v_t^{(b)}}{1 - \beta_2^t}
            pb=pb−ηt⋅m^t(b)v^t(b)+ϵp_b = p_b - \eta_t \cdot \frac{\hat{m}_t^{(b)}}{\sqrt{\hat{v}_t^{(b)}} + \epsilon}
        end for
    end for

    In the algorithm, ⊙\odot represents element-wise multiplication, and mean(⋅)\text{mean}(\cdot) denotes the arithmetic mean across all entries in the gradient tensor for block pbp_b. By reducing second-order momentum vectors to single scalars per block, ≥99.9%\ge 99.9\% of Adam's vv values are eliminated in large language models. Adam-mini operates with standard AdamW default hyperparameters (e.g., β1=0.9,β2=0.95 or 0.999,ϵ=10−8,λ=0.1\beta_1 = 0.9, \beta_2 = 0.95 \text{ or } 0.999, \epsilon = 10^{-8}, \lambda = 0.1).

  2. Knowl 2 — Hessian-Based Parameter Partition Principle for Transformers

    model/method

    Adam-mini partitions model parameters into distinct blocks based on the following general principle:

    Hessian Sub-block Partition Principle: Model parameters must be partitioned into blocks such that each parameter block corresponds to the smallest dense sub-block in the network's Hessian matrix.

    Because modern architectures (such as Transformers) exhibit Hessian-block heterogeneity (where different parameter blocks possess distinct eigenvalue distributions), learning rates must differ across blocks. However, within each dense sub-block, individual coordinate-wise learning rates are redundant, and a single shared scalar learning rate per block suffices. Naive coarse partitioning (such as layer-by-layer partitioning) violates this principle and leads to severe training instability in large language models.

    For standard Transformer architectures, empirical Hessian analysis identifies four structural classes of dense sub-blocks, defining the following partition strategy:

    1. Query (WQW_Q) and Key (WKW_K) projections: Partitioned by attention heads. For a module with HheadsH_{\text{heads}} attention heads, the tensor is split into HheadsH_{\text{heads}} independent blocks.
    2. Attention output projections (WOW_O), MLP projections, and Value (WVW_V) projections: Partitioned by output neurons (i.e., row-wise by output dimension doutd_{\text{out}}, yielding doutd_{\text{out}} parameter blocks).
    3. Embedding layer and Output LM head: Partitioned by tokens (i.e., row-wise across the vocabulary size).
    4. Non-Transformer architectures (CNNs, Diffusion U-Nets, GNNs): Partitioned by individual PyTorch parameter tensors (layer-wise).
  3. Knowl 3 — Preconditioning Ineffectiveness of Adam on Dense Hessian Sub-Blocks

    model/method

    Adam's coordinate-wise second-order momentum operates as a diagonal preconditioner DAdam=diag(1/v)D_{\text{Adam}} = \text{diag}(1 / \sqrt{v}), where v=g⊙gv = g \odot g with gradient vector g=Hbx∈Rdg = H_b x \in \mathbb{R}^d, initial vector x∼N(0,1dId)x \sim \mathcal{N}(0, \frac{1}{d} I_d), and Hb∈Rd×dH_b \in \mathbb{R}^{d \times d} is a positive-definite dense Hessian sub-block. The effectiveness of this preconditioner on HbH_b is quantified by the condition number ratio: r=κ(DAdamHb)κ(Hb)r = \frac{\kappa(D_{\text{Adam}} H_b)}{\kappa(H_b)} where κ(A)=λmax⁡(A)λmin⁡(A)\kappa(A) = \frac{\lambda_{\max}(A)}{\lambda_{\min}(A)} is the condition number of matrix AA (smaller is better, with r<1r < 1 signifying effective preconditioning).

    The behavior of rr is governed by the diagonal-over-off-diagonal ratio τ∈[0,1]\tau \in [0, 1] of HbH_b: τ=∑i=1d∣Hb,i,i∣∑i=1d∑j=1d∣Hb,i,j∣\tau = \frac{\sum_{i=1}^d |H_{b,i,i}|}{\sum_{i=1}^d \sum_{j=1}^d |H_{b,i,j}|} As τ→1\tau \to 1 (the Hessian sub-block is nearly diagonal), rr decreases below 11, indicating that Adam's coordinate-wise preconditioning is effective. In contrast, when τ\tau is small (the sub-block is dense, as observed in neural network neuron sub-blocks), r≥1r \ge 1 and grows large, showing that DAdamD_{\text{Adam}} fails to reduce—and often inflates—the sub-block condition number. Consequently, coordinate-wise scaling is redundant on dense blocks, and assigning a single average scalar learning rate per dense sub-block achieves equal or better convergence.

  4. Knowl 4 — Memory Reduction and Throughput Enhancement in Language Model Training

    empirical result

    By reducing second-order momentum vv from a per-parameter vector to a per-block scalar, Adam-mini cuts total float32 optimizer state memory by 50%50\% relative to AdamW.

    Model Optimizer Optimizer Memory (GB) Memory Reduction
    GPT-2-1.5B AdamW / Adam-mini 12.48 / 6.24 50.0%
    Llama 2-1B AdamW / Adam-mini 8.80 / 4.40 50.0%
    Llama 2-7B AdamW / Adam-mini 53.92 / 26.96 50.0%
    Llama 3-8B AdamW / Adam-mini 64.24 / 32.12 50.0%
    Llama 2-13B AdamW / Adam-mini 104.16 / 52.08 50.0%

    When pre-training Llama 2-7B on 2×NVIDIA A800-80GB2 \times \text{NVIDIA A800-80GB} GPUs:

    • Throughput: AdamW achieves a maximum batch size per GPU of 1 with a throughput of 3725.59 tokens/s3725.59 \text{ tokens/s} (a batch size of 2 encounters out-of-memory). Adam-mini accommodates a batch size per GPU of 4, achieving a throughput of 5572.19 tokens/s5572.19 \text{ tokens/s}, which is a 49.6%49.6\% throughput improvement.
    • Pre-training Wall-Clock Time: For pre-training Llama 2-7B under compute-optimal token budgets, GPU hours decrease from 74.56 h74.56\text{ h} to 49.85 h49.85\text{ h} for 1B tokens, from 5219.16 h5219.16\text{ h} to 3489.55 h3489.55\text{ h} for 70B tokens, and from 10438.32 h10438.32\text{ h} to 6979.10 h6979.10\text{ h} for 140B tokens, yielding a 33.1%33.1\% reduction in wall-clock time.
  5. Knowl 5 — Language Model Pre-Training Performance Across Model Scales

    empirical result

    Adam-mini matches the validation loss and convergence trajectory of AdamW across varied model architectures and scales from scratch using identical learning rates and hyperparameters without re-tuning:

    • GPT-2 Series (OpenWebText): On GPT-2 125M (small), 330M (medium), and 1.5B (XL) models (sequence length 1024, batch size 480, λ=0.1,ϵ=10−8,β1=0.9,β2=0.95\lambda = 0.1, \epsilon = 10^{-8}, \beta_1 = 0.9, \beta_2 = 0.95), Adam-mini's training and validation loss curves closely track AdamW throughout 20B tokens. In contrast, Adam-mini with naive PyTorch default partitioning exhibits severe training instability and diverges.
    • Llama Series (C4 Dataset): On Llama architectures spanning 20M to 13B parameters (including Llama 2-1B, Llama 2-7B, Llama 3-8B, and Llama 2-13B), Adam-mini's validation loss curves consistently match or slightly outperform AdamW.
    • Optimization Trajectory: In parameter space, the ℓ2\ell_2 checkpoint distance between Adam-mini and AdamW is significantly smaller than the distance between AdamW and factorized or sign-based optimizers (Adafactor, CAME).
  6. Knowl 6 — Scaling Law Compliance of Adam-mini Pre-Training

    empirical result

    Adam-mini adheres to compute and parameter scaling laws under Chinchilla-optimal token budgets (≈20×nparameters\approx 20 \times n_{\text{parameters}} tokens) when pre-training Llama 2 architectures from 39M to 1B parameters on the C4 dataset.

    Model Size Tokens Tokens/Params AdamW Val Perplexity Adam-mini Val Perplexity
    39M 1.02B 26.15 40.795 40.407
    67M 1.76B 26.27 29.319 29.014
    102M 2.67B 26.17 24.670 24.192
    162M 4.25B 26.23 20.360 20.172
    271M 7.10B 26.21 17.178 17.035
    1B 26.21B 26.21 12.452 12.372

    Across all evaluated scales from 39M to 1B, Adam-mini achieves consistently lower final validation perplexity and lower logarithmic validation loss than AdamW while requiring 50% less optimizer state memory.

  7. Knowl 7 — Alignment and Supervised Fine-Tuning Performance on Llama 2-7B

    empirical result

    On downstream supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) using the Llama 2-7B model on the UltraFeedback dataset, Adam-mini matches or outperforms AdamW in validation loss, reward optimization, and MT-Bench multi-turn conversation scores judged by GPT-4 (scale 0–10).

    SFT (LoRA) SFT (Full) RLHF (ReMax)
    Metric AdamW Adam-mini AdamW Adam-mini AdamW Adam-mini
    MT-Bench Score 4.23 4.41 5.37 5.40 5.54 5.68
    • SFT (Full Parameter): Batch size 80, 3 epochs, learning rate 2×10−62 \times 10^{-6} with cosine annealing, β1=0.9,β2=0.95\beta_1 = 0.9, \beta_2 = 0.95. Adam-mini achieves lower evaluation perplexity and higher MT-Bench score (5.405.40 vs. 5.375.37).
    • SFT (LoRA): Rank 128 on all non-embedding layers, learning rate 2×10−52 \times 10^{-5}. Adam-mini achieves lower perplexity and higher MT-Bench score (4.414.41 vs. 4.234.23).
    • RLHF (ReMax): Reinforcement learning on preference reward using the ReMax algorithm. Adam-mini attains a higher evaluation reward curve and higher MT-Bench score (5.685.68 vs. 5.545.54).
  8. Knowl 8 — Generalization of Adam-mini to Non-LLM Architectures

    empirical result

    Adam-mini matches or surpasses AdamW across non-LLM domains, including convolutional vision models, vision transformers, diffusion models, and graph neural networks.

    Architecture Dataset Metric AdamW Adam-mini
    ResNet-18 ImageNet Final Val Accuracy 0.6669 0.6667
    Swin-Transformer ImageNet Final Val Accuracy 0.7310 0.7300
    DiT-XL-2 ImageNet Final Train Loss / FID 0.1431 / 91.83 0.1430 / 88.20
    DC-AE-Diffusion ImageNet Final Train Loss / FID 0.2780 / 34.72 0.2780 / 33.15
    DDPM CelebA Final Train Loss 0.0394 0.0388
    GAT OGBN-arxiv Final Val Accuracy 0.7421 0.7429
    GCN OGBN-arxiv Final Val Accuracy 0.7374 0.7423

    For generative diffusion models (DiT-XL-2 and DC-AE-Diffusion), Adam-mini achieves lower Fréchet Inception Distance (FID) and higher Inception Scores (DiT-XL-2: 13.90 vs. 12.38; DC-AE-Diffusion: 44.38 vs. 41.79) compared to AdamW.

  9. Knowl 9 — Empirical Comparison of Adam-mini with Memory-Efficient Optimizers

    empirical result

    Compared with existing memory-efficient optimizers—including Adafactor, CAME, SM3, and Lion—Adam-mini demonstrates superior convergence stability, lower validation loss, and higher throughput.

    1. Adafactor and CAME: Adafactor compresses vv via nonnegative low-rank matrix factorization. On Llama 2-20M and Llama 2-1B pre-training across extensive hyperparameter sweeps (β2∈{0.95,0.999}\beta_2 \in \{0.95, 0.999\}, ϵ∈{10−30,10−16,10−8,10−6}\epsilon \in \{10^{-30}, 10^{-16}, 10^{-8}, 10^{-6}\}, warmup ratios 1%–10%1\%\text{--}10\%), both standard Adafactor and its modified variant underperform Adam-mini in validation loss and suffer from training instability on 1B models. Adam-mini achieves 40%40\% higher throughput on Llama 2-1B than Adafactor due to simple row-wise mean calculations rather than bidirectional row and column tensor summations.
    2. Lion: On Llama 2-20M and GPT-2 125M using tuned learning rates (lr∈[10−4,5×10−3]\text{lr} \in [10^{-4}, 5 \times 10^{-3}] for Llama 2-20M and [5×10−5,6×10−4][5 \times 10^{-5}, 6 \times 10^{-4}] for GPT-2 125M) with (β1,β2)=(0.95,0.98)(\beta_1, \beta_2) = (0.95, 0.98), Lion underperforms Adam-mini and suffers persistent loss spikes across all learning rates on GPT-2 125M.
    3. Hyperparameter Sensitivity: While Adafactor requires tuning up to 9 interdependent hyperparameters, Adam-mini requires no specialized hyperparameter search and operates effectively using identical hyperparameters to AdamW.
  10. Knowl 10 — Heuristic Block Averaging in Adam-mini as a Sub-Optimal Approximation

    limitation

    Adam-mini computes the block-wise second-order momentum as the scalar arithmetic mean vb=mean(gb⊙gb)v_b = \text{mean}(g_b \odot g_b), which approximates the dense Hessian sub-block learning rate by the average coordinate-wise variance. While this choice is computationally lightweight and tracks AdamW's optimization trajectory (because backpropagation error vectors within a weight matrix row share the same output neuron error eie_i), it is not theoretically optimal.

    On synthetic quadratic problems and small Transformers, block-wise gradient descent utilizing optimal learning rates derived from sub-block curvature spectra (such as 2/(Lb+μb)2 / (L_b + \mu_b), where LbL_b and μb\mu_b are the largest and smallest eigenvalues of sub-block HbH_b) converges strictly faster than both AdamW and Adam-mini. Adam-mini does not utilize full sub-block spectral information due to the severe computational cost of online Hessian eigenvalue decomposition.

Coverage note — None was omitted; the knowls comprehensively cover the algorithmic formulation, Hessian partition principles, preconditioning analysis, memory and throughput metrics, pre-training results across scales, scaling laws, SFT/RLHF alignment, non-LLM evaluations, baseline comparisons with Adafactor and Lion, and algorithmic limitations.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Kwangjun Ahn, Zhiyu Zhang, Yunbum Kook, and Yan Dai. Understanding adam optimizer via online learning of updates: Adam is ftrl in disguise. In Forty-first International Conference on Machine Learning.
  3. 3.Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer. Memory efficient adaptive optimization. Advances in Neural Information Processing Systems, 32, 2019.
  4. 4.Anonymous authors. Deconstructing what makes a good optimizer for language models. https://openreview.net/pdf?id=zfeso8ceqr, 2024.
  5. 5.Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. Entropy-sgd: Biasing gradient descent into wide valleys. Journal of Statistical Mechanics: Theory and Experiment, 2019(12):124018, 2019.
  6. 6.Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, Yao Lu, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models. arXiv preprint arXiv:2410.10733, 2024a.
  7. 7.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016.
  8. 8.Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al. Symbolic discovery of optimization algorithms. Advances in neural information processing systems, 36, 2024b.
  9. 9.Ronan Collobert. Large scale machine learning. Technical report, Université de Paris VI, 2004.
  10. 10.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. Ultrafeedback: Boosting language models with high-quality feedback, 2023.
  11. 11.André Belotto Da Silva and Maxime Gazeau. A general system of differential equations to model first-order adaptive algorithms. The Journal of Machine Learning Research, 21(1):5072–5113, 2020.
  12. 12.Rudrajit Das, Naman Agarwal, Sujay Sanghavi, and Inderjit S Dhillon. Towards quantifying the preconditioning effect of adam. arXiv preprint arXiv:2402.07114, 2024.
  13. 13.Yann N Dauphin, Atish Agarwala, and Hossein Mobahi. Neglected hessian component explains mysteries in sharpness regularization. arXiv preprint arXiv:2401.10809, 2024.
  14. 14.Aaron Defazio, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, Ashok Cutkosky, et al. The road less scheduled. arXiv preprint arXiv:2405.15682, 2024.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  16. 16.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization. In International Conference on Learning Representations, 2021.
  17. 17.John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, 12(7), 2011.
  18. 18.George E Forsythe and Ernst G Straus. On best conditioned matrices. Proceedings of the American Mathematical Society, 6(3):340–345, 1955.
  19. 19.Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao. An investigation into neural net optimization via hessian eigenvalue density. In International Conference on Machine Learning, pp. 2232–2241. PMLR, 2019.
  20. 20.Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen. Training deep networks with stochastic gradient normalized by layerwise adaptive second moments. 2019.
  21. 21.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. Openwebtext corpus, 2019.
  22. 22.Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer. Gradient descent happens in a tiny subspace. arXiv preprint arXiv:1812.04754, 2018.
  23. 23.Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro Von Werra, and Martin Jaggi. Scaling laws and compute-optimal training beyond fixed training durations. arXiv preprint arXiv:2405.18392, 2024.
  24. 24.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  25. 25.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  26. 26.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  27. 27.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  28. 28.Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. Simple and scalable strategies to continually pre-train large language models. arXiv preprint arXiv:2403.08763, 2024.
  29. 29.Kaiqi Jiang, Dhruv Malik, and Yuanzhi Li. How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36, 2023.
  30. 30.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  31. 31.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2016.
  32. 32.Fuad Kittaneh. Spectral radius inequalities for hilbert space operators. Proceedings of the American Mathematical Society, pp. 385–390, 2006.
  33. 33.Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt. Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be. arXiv preprint arXiv:2304.13960, 2023.
  34. 34.Frederik Kunstner, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti. Heavy-tailed class imbalance and why adam outperforms gradient descent on language models. arXiv preprint arXiv:2402.19449, 2024.
  35. 35.Bingrui Li, Jianfei Chen, and Jun Zhu. Memory efficient optimizers with 4-bit states. Advances in Neural Information Processing Systems, 36, 2024.
  36. 36.Ziniu Li, Tian Xu, Yushun Zhang, Yang Yu, Ruoyu Sun, and Zhi-Quan Luo. Remax: A simple, effective, and efficient method for aligning large language models. arXiv preprint arXiv:2310.10505, 2023.
  37. 37.Zhenyu Liao and Michael W Mahoney. Hessian eigenspectra of more realistic nonlinear models. Advances in Neural Information Processing Systems, 34:20104–20117, 2021.
  38. 38.Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. Sophia: A scalable stochastic second-order optimizer for language model pre-training. arXiv preprint arXiv:2305.14342, 2023.
  39. 39.Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue. Learning by turning: Neural architecture aware optimisation. In International Conference on Machine Learning, pp. 6748–6758. PMLR, 2021a.
  40. 40.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 10012–10022, 2021b.
  41. 41.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  42. 42.Qijun Luo, Hengxu Yu, and Xiao Li. Badam: A memory efficient full parameter training method for large language models. arXiv preprint arXiv:2404.02827, 2024.
  43. 43.Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang, and Yang You. Came: Confidence-guided adaptive memory efficient optimization. arXiv preprint arXiv:2307.02047, 2023.
  44. 44.Kai Lv, Hang Yan, Qipeng Guo, Haijun Lv, and Xipeng Qiu. Adalomo: Low-memory optimization with adaptive learning rate. arXiv preprint arXiv:2310.10195, 2023a.
  45. 45.Kai Lv, Yuqing Yang, Tengxiao Liu, Qinghui Gao, Qipeng Guo, and Xipeng Qiu. Full parameter fine-tuning for large language models with limited resources. arXiv preprint arXiv:2306.09782, 2023b.
  46. 46.Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. Advances in Neural Information Processing Systems, 36:53038–53075, 2023.
  47. 47.James Martens and Roger Grosse. Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, pp. 2408–2417. PMLR, 2015.
  48. 48.Francesco Orabona. Neural networks (maybe) evolved to make adam the best optimizer. 2020. URL https://parameterfree.com/2020/12/06/neural-network-maybe-evolved-to-make-adam-the-best-optimizer/.
  49. 49.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  50. 50.Yan Pan and Yuanzhi Li. Toward understanding why adam converges faster than sgd for transformers. arXiv preprint arXiv:2306.00204, 2023.
  51. 51.Vardan Papyan. The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size. arXiv preprint arXiv:1811.07062, 2018.
  52. 52.Vardan Papyan. Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians. arXiv preprint arXiv:1901.08244, 2019.
  53. 53.Vardan Papyan. Traces of class/cross-class structure pervade deep learning spectra. The Journal of Machine Learning Research, 21(1):10197–10260, 2020.
  54. 54.Barak A Pearlmutter. Fast exact multiplication by the hessian. Neural computation, 6(1):147–160, 1994.
  55. 55.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  56. 56.Zhaonan Qu, Wenzhi Gao, Oliver Hinder, Yinyu Ye, and Zhengyuan Zhou. Optimal diagonal preconditioning: Theory and practice. arXiv preprint arXiv:2209.00809, 2022.
  57. 57.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  58. 58.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  59. 59.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16. IEEE, 2020.
  60. 60.Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In Proceedings of the international conference for high performance computing, networking, storage and analysis, pp. 1–14, 2021.
  61. 61.Nicolas Roux, Pierre-Antoine Manzagol, and Yoshua Bengio. Topmoumoute online natural gradient algorithm. Advances in neural information processing systems, 20, 2007.
  62. 62.Levent Sagun, Leon Bottou, and Yann LeCun. Eigenvalues of the hessian in deep learning: Singularity and beyond. arXiv preprint arXiv:1611.07476, 2016.
  63. 63.Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou. Empirical analysis of the hessian of over-parametrized neural networks. arXiv preprint arXiv:1706.04454, 2017.
  64. 64.Adepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, and Vineeth N Balasubramanian. A deeper look at the hessian eigenspectrum of deep neural networks and its applications to regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 9481–9488, 2021.
  65. 65.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  66. 66.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pp. 4596–4604. PMLR, 2018.
  67. 67.Naichen Shi, Dawei Li, Mingyi Hong, and Ruoyu Sun. Rmsprop converges with proper hyper-parameter. In International Conference on Learning Representations, 2020.
  68. 68.Ruoyu Sun and Yinyu Ye. Worst-case complexity of cyclic coordinate descent: O (nˆ 2) o (n 2) gap with randomized version. Mathematical Programming, 185:487–520, 2021.
  69. 69.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  70. 70.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  71. 71.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  72. 72.Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  73. 73.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, Yoshua Bengio, et al. Graph attention networks. stat, 1050(20):10–48550, 2017.
  74. 74.Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhi-Ming Ma, Tie-Yan Liu, and Wei Chen. Provable adaptivity in adam. arXiv preprint arXiv:2208.09900, 2022.
  75. 75.Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang, and Rong Ge. Dissecting hessian: Understanding common structure of hessian in neural networks. arXiv preprint arXiv:2010.04261, 2020.
  76. 76.Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan. Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models. arXiv preprint arXiv:2208.06677, 2022.
  77. 77.Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney. Hessian-based analysis of large batch training and robustness to adversaries. Advances in Neural Information Processing Systems, 31, 2018.
  78. 78.Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney. Pyhessian: Neural networks through the lens of the hessian. In 2020 IEEE international conference on big data (Big data), pp. 581–590. IEEE, 2020.
  79. 79.Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes. arXiv preprint arXiv:1904.00962, 2019.
  80. 80.David Young. Iterative methods for solving partial difference equations of elliptic type. Transactions of the American Mathematical Society, 76(1):92–111, 1954.
  81. 81.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12104–12113, 2022.
  82. 82.Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse. Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model. Advances in neural information processing systems, 32, 2019a.
  83. 83.Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. Why gradient clipping accelerates training: A theoretical justification for adaptivity. arXiv preprint arXiv:1905.11881, 2019b.
  84. 84.Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra. Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33:15383–15393, 2020.
  85. 85.Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhi-Quan Luo. Adam can converge without any modification on update rules. Advances in Neural Information Processing Systems, 35: 28386–28399, 2022.
  86. 86.Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo. Why transformers need adam: A hessian perspective. arXiv preprint arXiv:2402.16788, 2024.
  87. 87.Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian. Galore: Memory-efficient llm training by gradient low-rank projection. arXiv preprint arXiv:2403.03507, 2024a.
  88. 88.Rosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas, and Sham Kakade. Deconstructing what makes a good optimizer for language models. arXiv preprint arXiv:2407.07972, 2024b.
  89. 89.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.
  90. 90.Shuai Zheng and James T Kwok. Blockwise adaptivity: Faster training and better generalization in deep learning. arXiv preprint arXiv:1905.09899, 2019.

Citation

MLA
Zhang, Y., et al. “Adam-mini: Use Fewer Learning Rates To Gain More”. arXiv, 2024, http://arxiv.org/abs/2406.16793v7.
APA
Zhang, Y., Chen, C., Li, Z., Ding, T., Wu, C., Kingma, D. P., Ye, Y., Luo, Z.-Q., & Sun, R. (2024). Adam-mini: Use Fewer Learning Rates To Gain More. arXiv. http://arxiv.org/abs/2406.16793v7
Chicago
Zhang, Y., C. Chen, Z. Li, et al. 2024. “Adam-mini: Use Fewer Learning Rates To Gain More”. arXiv. http://arxiv.org/abs/2406.16793v7.
Harvard
Zhang, Y. et al. (2024) “Adam-mini: Use Fewer Learning Rates To Gain More”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.16793v7.
Vancouver
1. Zhang Y, Chen C, Li Z, Ding T, Wu C, Kingma DP, Ye Y, Luo Z-Q, Sun R (2024) Adam-mini: Use Fewer Learning Rates To Gain More. arXiv

BibTeX

@article{zhang2024adam,
  title = {Adam-mini: Use Fewer Learning Rates To Gain More},
  author = {Zhang, Yushun and Chen, Congliang and Li, Ziniu and Ding, Tian and Wu, Chenwei and Kingma, Diederik P. and Ye, Yinyu and Luo, Zhi-Quan and Sun, Ruoyu},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.16793v7},
  eprint = {2406.16793}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/