Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Noam ShazeerMitchell Stern
Introduces Adafactor, a memory-efficient adaptive optimizer that tracks factored row and column statistics instead of full second-moment matrices, matching Adam's training performance on large Transformers while drastically cutting optimizer memory overhead.
Training state-of-the-art deep neural networks requires massive computational hardware, but memory capacity has become a critical bottleneck as models expand to billions of parameters. Popular optimization algorithms like Adam adjust step sizes adaptively using running historical averages of gradient statistics. However, Adam requires storing two auxiliary tracking values for every model parameter, effectively tripling the memory needed for model weights and severely constraining the maximum size of neural networks that can fit onto hardware accelerators.
The article demonstrates an optimization algorithm called Adafactor, designed to retain the rapid convergence and empirical benefits of adaptive optimizers while drastically reducing their auxiliary memory footprint. It also aims to resolve common training instabilities associated with adaptive step sizes.
The authors evaluate this approach by training Transformer machine translation models on the standard WMT 2014 English-to-German dataset. Rather than storing a full matrix of past gradient squares, the approach factors the matrix into row and column sums, deriving a low-rank approximation that requires only a fraction of the original storage. To further cut memory usage, momentum tracking is completely removed. To address the training instabilities that occur when momentum is disabled or learning rate warmups are omitted, the authors evaluate two corrective techniques: an update clipping mechanism that caps oversized parameter steps, and an increasing decay schedule for past gradient statistics.
The evaluation yielded several key findings. First, factoring the second-moment accumulator reduces its memory requirement from proportional to the product of matrix dimensions down to their sum, achieving task accuracy comparable to full-memory Adam (achieving BLEU scores around 25.4 to 25.6 with learning rate warmup). Second, completely removing momentum eliminates another full parameter tracking state without degrading translation quality, provided training stability safeguards are present. Third, training instability without warmup was shown to stem from out-of-date historical statistics causing oversized parameter updates; applying update clipping restored performance from a failing score of 0.1 BLEU up to 21.5–22.4 BLEU. Fourth, scaling updates relative to the magnitude of the model parameters themselves made training substantially more resilient to arbitrary parameter initialization schemes, yielding superior translation scores (25.4–26.6 BLEU) compared to Adam across suboptimal setups.
These findings mean engineering teams can train significantly larger language and translation models on existing, memory-constrained hardware accelerators without sacrificing training speed or model accuracy. Eliminating the auxiliary memory overhead directly lowers infrastructure costs and reduces the operational risks of out-of-memory failures during large-scale model training runs.
Engineering leaders and practitioners should adopt Adafactor as a drop-in replacement for Adam in memory-constrained neural network workflows, using the recommended default parameters and update clipping. When implementing the algorithm, teams can utilize the implementation open-sourced in the Tensor2Tensor repository to avoid custom engineering overhead.
The primary limitation is that empirical validations in the article were conducted specifically on Transformer architectures for machine translation tasks using reduced batch sizes. While confidence in these results is high, teams training architectures outside natural language processing, such as computer vision models, should validate Adafactor in small-scale pilot runs before committing to full-scale training deployments.
- Paper: Adam: A Method for Stochastic Optimization, Diederik P. Kingma et al. (2015). It introduces the Adam optimization algorithm and per-parameter second-moment tracking that Adafactor directly modifies and compresses.
- Paper: ADADELTA: An Adaptive Learning Rate Method, Matthew D. Zeiler (2012). It establishes the concept of exponential moving averages of squared gradients in adaptive learning rates, which forms the foundation of second-moment estimators targeted by Adafactor.
- Paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization, John Duchi et al. (2011). It introduces adaptive coordinate-wise learning rate scaling based on past squared gradients, the baseline paradigm whose memory overhead Adafactor reduces.
- Paper: An overview of gradient descent optimization algorithms, Sebastian Ruder (2016). It provides a comprehensive overview of adaptive stochastic optimizers like RMSProp and Adam, contextualizing the memory and update issues Adafactor seeks to solve.
- Paper: On the Convergence of Adam and Beyond, Sashank J. Reddi et al. (2018). It analyzes the non-convergence and decay-rate issues in Adam's second-moment estimation, motivating Adafactor's dynamic decay rate and update clipping schemes.
- Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). It explores sublinear memory strategies for neural network training, providing conceptual background on memory reduction trade-offs in deep learning.
- Paper: GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection, Jiawei Zhao et al. (2024). It extends memory-efficient optimizer design to full-parameter large language model training by projecting gradients into compact low-rank subspaces.
- Paper: Decoupled Weight Decay Regularization, Ilya Loshchilov et al. (2019). It investigates decoupled weight decay regularization to resolve generalization deficiencies in adaptive optimizers like Adam and its variants.
- Paper: On the Variance of the Adaptive Learning Rate and Beyond, Liyuan Liu et al. (2019). It builds on the investigation of adaptive learning rate variance and step stability to introduce a variance-rectified optimizer that stabilizes early training.
- Paper: Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer, Jinghui Yuan et al. (2026). It advances matrix-level optimizer design by enforcing row-wise orthogonal alignment and normalization to stabilize large model training.
- Paper: ZeRO: Memory optimizations Toward Training Trillion Parameter Models, Samyam Rajbhandari et al. (2020). It addresses optimizer memory overhead at massive scale through distributed partitioning of optimizer states, parameters, and gradients across devices.
- Paper: On Layer Normalization in the Transformer Architecture, Ruibin Xiong et al. (2020). It examines how layer normalization placement in Transformers affects gradient scale and initialization stability during optimization with adaptive methods.
- Paper: ReLoRA: High-Rank Training Through Low-Rank Updates, Vladislav Lialin et al. (2024). It builds on memory-efficient Transformer pre-training techniques by utilizing restarted low-rank updates to achieve high-rank parameter progression.
