Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Noam ShazeerMitchell Stern

article2018ICML1,356 citations

Introduces Adafactor, a memory-efficient adaptive optimizer that tracks factored row and column statistics instead of full second-moment matrices, matching Adam's training performance on large Transformers while drastically cutting optimizer memory overhead.

Listen

Training state-of-the-art deep neural networks requires massive computational hardware, but memory capacity has become a critical bottleneck as models expand to billions of parameters. Popular optimization algorithms like Adam adjust step sizes adaptively using running historical averages of gradient statistics. However, Adam requires storing two auxiliary tracking values for every model parameter, effectively tripling the memory needed for model weights and severely constraining the maximum size of neural networks that can fit onto hardware accelerators.

The article demonstrates an optimization algorithm called Adafactor, designed to retain the rapid convergence and empirical benefits of adaptive optimizers while drastically reducing their auxiliary memory footprint. It also aims to resolve common training instabilities associated with adaptive step sizes.

The authors evaluate this approach by training Transformer machine translation models on the standard WMT 2014 English-to-German dataset. Rather than storing a full matrix of past gradient squares, the approach factors the matrix into row and column sums, deriving a low-rank approximation that requires only a fraction of the original storage. To further cut memory usage, momentum tracking is completely removed. To address the training instabilities that occur when momentum is disabled or learning rate warmups are omitted, the authors evaluate two corrective techniques: an update clipping mechanism that caps oversized parameter steps, and an increasing decay schedule for past gradient statistics.

The evaluation yielded several key findings. First, factoring the second-moment accumulator reduces its memory requirement from proportional to the product of matrix dimensions down to their sum, achieving task accuracy comparable to full-memory Adam (achieving BLEU scores around 25.4 to 25.6 with learning rate warmup). Second, completely removing momentum eliminates another full parameter tracking state without degrading translation quality, provided training stability safeguards are present. Third, training instability without warmup was shown to stem from out-of-date historical statistics causing oversized parameter updates; applying update clipping restored performance from a failing score of 0.1 BLEU up to 21.5–22.4 BLEU. Fourth, scaling updates relative to the magnitude of the model parameters themselves made training substantially more resilient to arbitrary parameter initialization schemes, yielding superior translation scores (25.4–26.6 BLEU) compared to Adam across suboptimal setups.

These findings mean engineering teams can train significantly larger language and translation models on existing, memory-constrained hardware accelerators without sacrificing training speed or model accuracy. Eliminating the auxiliary memory overhead directly lowers infrastructure costs and reduces the operational risks of out-of-memory failures during large-scale model training runs.

Engineering leaders and practitioners should adopt Adafactor as a drop-in replacement for Adam in memory-constrained neural network workflows, using the recommended default parameters and update clipping. When implementing the algorithm, teams can utilize the implementation open-sourced in the Tensor2Tensor repository to avoid custom engineering overhead.

The primary limitation is that empirical validations in the article were conducted specifically on Transformer architectures for machine translation tasks using reduced batch sizes. While confidence in these results is high, teams training architectures outside natural language processing, such as computer vision models, should validate Adafactor in small-scale pilot runs before committing to full-scale training deployments.

Cover for Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Abstract

In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For the case of neural network weight matrices, we propose maintaining only the per-row and per-column sums of these moving averages, and estimating the per-parameter second moments based on these sums. We demonstrate empirically that this method produces similar results to the baseline. Secondly, we show that adaptive methods can produce larger-than-desired updates when the decay rate of the second moment accumulator is too slow. We propose update clipping and a gradually increasing decay rate scheme as remedies. Combining these methods and dropping momentum, we achieve comparable results to the published Adam regime in training the Transformer model on the WMT 2014 English-German machine translation task, while using very little auxiliary storage in the optimizer. Finally, we propose scaling the parameter updates based on the scale of the parameters themselves.

Table of Contents

  • 1 Introduction and Background
  • 2 A Brief Review of Adam
  • 3 Factored Second Moment Estimation
  • 3.1 Relation to Prior Work
  • 3.2 Experiments
  • 4 No Momentum
  • 4.1 Experiments
  • 5 A Problem with Adam: Out-of-Date Second Moment Estimator
  • 6 Update Clipping
  • 6.1 Comparison to Gradient Clipping
  • 6.2 Experiments
  • 7 Increasing Decay Parameter
  • 7.1 In Adam
  • 7.2 Proposed Alternative
  • 7.3 Experiments
  • 8 Relative Step Size
  • 8.1 Experiments
  • 9 Experimental Setup
  • 9.1 Results
  • 10 Conclusion
  • References

Knowls

  1. Knowl 1 — Adafactor Optimization Algorithm

    algorithm

    Adafactor is a stochastic optimization algorithm designed to train deep neural networks with sublinear auxiliary memory overhead compared to Adam. It reduces the O(nm)O(nm) memory required for storing second-moment accumulators of an n×mn \times m matrix parameter XX to O(n+m)O(n + m) by maintaining exponential moving averages of row and column sums of squared gradients. It removes momentum storage by setting first-moment decay to zero, prevents training instability using update clipping and an increasing decay schedule β^2t\hat{\beta}_{2t}, and scales updates proportionally to parameter scale using relative step sizes ρt\rho_t.

    Inputs: Initial parameters X0∈Rn×mX_0 \in \mathbb{R}^{n \times m} (for matrix) or X0∈RnX_0 \in \mathbb{R}^n (for vector), relative step sizes {ρt}t=1T\{\rho_t\}_{t=1}^T, second-moment decay schedule {β^2t}t=1T\{\hat{\beta}_{2t}\}_{t=1}^T with β^21=0\hat{\beta}_{21} = 0, small regularization constants ϵ1,ϵ2\epsilon_1, \epsilon_2, clipping threshold dd
    Initialize for matrix: R0=0∈RnR_0 = 0 \in \mathbb{R}^n, C0=0∈RmC_0 = 0 \in \mathbb{R}^m
    Initialize for vector: V^0=0∈Rn\hat{V}_0 = 0 \in \mathbb{R}^n
    for t=1t = 1 to TT do
        Evaluate gradient Gt=∇ft(Xt−1)G_t = \nabla f_t(X_{t-1})
        Compute effective step size αt=max⁡(ϵ2,RMS(Xt−1))⋅ρt\alpha_t = \max(\epsilon_2, \text{RMS}(X_{t-1})) \cdot \rho_t
        if XX is a 2D weight matrix then
            Rt=β^2tRt−1+(1−β^2t)(Gt2+ϵ11n1m⊤)1mR_t = \hat{\beta}_{2t} R_{t-1} + (1 - \hat{\beta}_{2t})(G_t^2 + \epsilon_1 \mathbf{1}_n \mathbf{1}_m^\top) \mathbf{1}_m
            Ct=β^2tCt−1+(1−β^2t)1n⊤(Gt2+ϵ11n1m⊤)C_t = \hat{\beta}_{2t} C_{t-1} + (1 - \hat{\beta}_{2t}) \mathbf{1}_n^\top (G_t^2 + \epsilon_1 \mathbf{1}_n \mathbf{1}_m^\top)
            Reconstruct second-moment estimator V^t=RtCt/(1n⊤Rt)\hat{V}_t = R_t C_t / (\mathbf{1}_n^\top R_t)
        else if XX is a 1D weight vector then
            V^t=β^2tV^t−1+(1−β^2t)(Gt2+ϵ11n)\hat{V}_t = \hat{\beta}_{2t} \hat{V}_{t-1} + (1 - \hat{\beta}_{2t})(G_t^2 + \epsilon_1 \mathbf{1}_n)
        end if
        Compute unscaled update Ut=Gt/V^tU_t = G_t / \sqrt{\hat{V}_t}
        Apply update clipping U^t=Ut/max⁡(1,RMS(Ut)/d)\hat{U}_t = U_t / \max(1, \text{RMS}(U_t) / d)
        Update parameter Xt=Xt−1−αtU^tX_t = X_{t-1} - \alpha_t \hat{U}_t
    end for

    Where RMS(Y)=1∣Y∣∑y∈Yy2\text{RMS}(Y) = \sqrt{\frac{1}{|Y|} \sum_{y \in Y} y^2} computes the root-mean-square over all components of tensor YY. The standard recommended hyperparameters for Adafactor are:

    • ϵ1=10−30\epsilon_1 = 10^{-30}
    • ϵ2=10−3\epsilon_2 = 10^{-3}
    • d=1.0d = 1.0
    • β^2t=1−t−0.8\hat{\beta}_{2t} = 1 - t^{-0.8}
    • ρt=min⁡(10−2,1/t)\rho_t = \min(10^{-2}, 1 / \sqrt{t})
  2. Knowl 2 — Exact Rank-1 Second-Moment Factorization under Generalized Kullback-Leibler Divergence

    theoretical result

    Let V∈Rn×mV \in \mathbb{R}^{n \times m} be a nonnegative matrix representing the squared gradient accumulator. Approximating VV with a rank-1 nonnegative factorization RSRS where R∈Rn×1R \in \mathbb{R}^{n \times 1} and S∈R1×mS \in \mathbb{R}^{1 \times m} under the generalized Kullback-Leibler divergence (also known as the I-divergence) is formulated as:

    min⁡R∈Rn×1,S∈R1×m∑i=1n∑j=1md(Vij,[RS]ij)subject to Ri1≥0,S1j≥0\min_{R \in \mathbb{R}^{n \times 1}, S \in \mathbb{R}^{1 \times m}} \sum_{i=1}^n \sum_{j=1}^m d(V_{ij}, [RS]_{ij}) \quad \text{subject to } R_{i1} \ge 0, S_{1j} \ge 0

    where d(p,q)=plog⁡(p/q)−p+qd(p, q) = p \log(p / q) - p + q (with 0/0=00/0 = 0, 0log⁡0=00 \log 0 = 0, and p/0=∞p/0 = \infty for p>0p > 0).

    The solution set consists of all feasible pairs (R,S)(R, S) satisfying:

    RS=(V1m)(1n⊤V)1n⊤V1mRS = \frac{(V \mathbf{1}_m)(\mathbf{1}_n^\top V)}{\mathbf{1}_n^\top V \mathbf{1}_m}

    where 1ℓ=(1,…,1)⊤∈Rℓ\mathbf{1}_\ell = (1, \dots, 1)^\top \in \mathbb{R}^\ell.

    Because the row sums V1mV \mathbf{1}_m and column sums 1n⊤V\mathbf{1}_n^\top V are linear functions of VV, the row and column sums of an exponential moving average of matrices equal the exponential moving averages of the individual row and column sums. Storing only the moving averages of the row sums Rt∈RnR_t \in \mathbb{R}^n and column sums Ct∈RmC_t \in \mathbb{R}^m reduces storage requirements from O(nm)O(nm) to O(n+m)O(n + m).

  3. Knowl 3 — Update Clipping Based on Root-Mean-Square Step Magnitude

    model/method

    In adaptive gradient methods, dividing gradients coordinate-wise by the square root of an out-of-date second-moment estimate can lead to unscaled updates that are excessively large. Update clipping bounds the update magnitude of an entire parameter tensor (matrix or vector) XX.

    Let Gt=∇ft(Xt−1)G_t = \nabla f_t(X_{t-1}) and let V^t\hat{V}_t be the (bias-corrected or factored) second-moment accumulator for parameter XX. The unscaled update matrix/vector UtU_t is defined coordinate-wise as:

    Ut=GtV^tU_t = \frac{G_t}{\sqrt{\hat{V}_t}}

    The root-mean-square norm of UtU_t over all elements in XX is:

    RMS(Ut)=1∣X∣∑x∈X(uxt)2\text{RMS}(U_t) = \sqrt{\frac{1}{|X|} \sum_{x \in X} (u_{xt})^2}

    Given a clipping threshold d>0d > 0 (typically d=1.0d = 1.0), the clipped unscaled update U^t\hat{U}_t is defined as:

    U^t=Utmax⁡(1,RMS(Ut)d)\hat{U}_t = \frac{U_t}{\max\left(1, \frac{\text{RMS}(U_t)}{d}\right)}

    The final parameter update is −αtU^t-\alpha_t \hat{U}_t. Unlike standard gradient clipping, which caps ∥Gt∥\|G_t\| prior to adaptive scaling and can still permit unbounded step sizes when V^t\hat{V}_t is small, update clipping directly constrains the norm of the actual step direction.

  4. Knowl 4 — Parameter Scale-Dependent Relative Step Size

    model/method

    Instead of specifying a global absolute learning rate αt\alpha_t, relative step sizes {ρt}t=1T\{\rho_t\}_{t=1}^T scale the parameter updates proportionally to the root-mean-square scale of the parameters themselves.

    For a parameter tensor Xt−1∈Rn×mX_{t-1} \in \mathbb{R}^{n \times m} or Rn\mathbb{R}^n, its scale is computed as:

    RMS(Xt−1)=1∣X∣∑x∈Xt−1x2\text{RMS}(X_{t-1}) = \sqrt{\frac{1}{|X|} \sum_{x \in X_{t-1}} x^2}

    The effective step size αt\alpha_t applied at step tt is:

    αt=max⁡(ϵ2,RMS(Xt−1))⋅ρt\alpha_t = \max\left(\epsilon_2, \text{RMS}(X_{t-1})\right) \cdot \rho_t

    where ϵ2>0\epsilon_2 > 0 (set to 10−310^{-3}) is a regularization constant that provides a lower bound, allowing parameters initialized to zero to update away from zero. A typical relative step size schedule is ρt=min⁡(10−2,1t)\rho_t = \min\left(10^{-2}, \frac{1}{\sqrt{t}}\right).

  5. Knowl 5 — Dynamic Second-Moment Decay Rate Schedule for Adaptive Optimization

    theoretical result

    To eliminate the need for Adam-style bias correction while maintaining early training stability, a time-varying second-moment decay schedule is defined as:

    β^2t=1−t−c,t≥1,c>0\hat{\beta}_{2t} = 1 - t^{-c}, \quad t \ge 1, \quad c > 0

    where β^21=0\hat{\beta}_{21} = 0. The second-moment moving average is computed directly as:

    vt=β^2tvt−1+(1−β^2t)gt2=∑i=1t(1−β^2i)(∏j=i+1tβ^2j)gi2v_t = \hat{\beta}_{2t} v_{t-1} + (1 - \hat{\beta}_{2t}) g_t^2 = \sum_{i=1}^t (1 - \hat{\beta}_{2i}) \left(\prod_{j=i+1}^t \hat{\beta}_{2j}\right) g_i^2

    This schedule satisfies two theoretical properties:

    1. Exact Normalization Without Bias Correction: For any schedule where β^21=0\hat{\beta}_{21} = 0, the sum of weights is identically 1 for all t≥1t \ge 1:

    ∑i=1t(1−β^2i)∏j=i+1tβ^2j=1\sum_{i=1}^t (1 - \hat{\beta}_{2i}) \prod_{j=i+1}^t \hat{\beta}_{2j} = 1

    Consequently, E[vt]=E[gt2]\mathbb{E}[v_t] = \mathbb{E}[g_t^2] when the gradient distribution is stationary, making separate division by (1−β2t)(1 - \beta_2^t) unnecessary.

    1. Asymptotic Decay of Past Gradients: The contribution of any past gradient gi2g_i^2 vanishes as t→∞t \to \infty, i.e.,

    lim⁡t→∞(1−β^2i)∏j=i+1tβ^2j=0for all i≥1\lim_{t \to \infty} (1 - \hat{\beta}_{2i}) \prod_{j=i+1}^t \hat{\beta}_{2j} = 0 \quad \text{for all } i \ge 1

    if and only if c≤1c \le 1. If c>1c > 1, the infinite product converges to a non-zero value, retaining non-trivial weight on early gradients indefinitely. When c=1c = 1, vt=1t∑i=1tgi2v_t = \frac{1}{t} \sum_{i=1}^t g_i^2, recovering an arithmetic running mean.

  6. Knowl 6 — Dynamic Decay-Rate Reformulation of Adam Bias Correction

    equation

    The standard Adam bias correction for first and second gradient moments can be equivalently expressed without an explicit bias-correction division step by replacing constant decay rates β1,β2\beta_1, \beta_2 with time-dependent decay rates β^1t,β^2t\hat{\beta}_{1t}, \hat{\beta}_{2t}:

    β^1t=β11−β1t−11−β1t,β^2t=β21−β2t−11−β2t\hat{\beta}_{1t} = \beta_1 \frac{1 - \beta_1^{t-1}}{1 - \beta_1^t}, \qquad \hat{\beta}_{2t} = \beta_2 \frac{1 - \beta_2^{t-1}}{1 - \beta_2^t}

    The bias-corrected first and second moment estimators m^t\hat{m}_t and v^t\hat{v}_t are then updated recursively as:

    m^t=β^1tm^t−1+(1−β^1t)gt\hat{m}_t = \hat{\beta}_{1t} \hat{m}_{t-1} + (1 - \hat{\beta}_{1t}) g_t

    v^t=β^2tv^t−1+(1−β^2t)gt2\hat{v}_t = \hat{\beta}_{2t} \hat{v}_{t-1} + (1 - \hat{\beta}_{2t}) g_t^2

    Under this formulation, β^11=0\hat{\beta}_{11} = 0 and β^21=0\hat{\beta}_{21} = 0 at t=1t = 1, and both asymptotically approach β1\beta_1 and β2\beta_2 respectively as t→∞t \to \infty.

  7. Knowl 7 — Root-Mean-Square Unscaled Update Metric and Second-Moment Staleness

    empirical result

    The tracking quality of a second-moment estimator v^xt\hat{v}_{xt} for a parameter matrix or vector XX at step tt is measured by the root-mean-square of the unscaled parameter update uxt=−gxt/v^xtu_{xt} = -g_{xt} / \sqrt{\hat{v}_{xt}}:

    RMS(Ut)=Meanx∈X(gxt2v^xt)\text{RMS}(U_t) = \sqrt{\text{Mean}_{x \in X} \left( \frac{g_{xt}^2}{\hat{v}_{xt}} \right)}

    When the second-moment estimator accurately tracks current gradient magnitudes, v^xt≈gxt2\hat{v}_{xt} \approx g_{xt}^2, yielding RMS(Ut)≈1\text{RMS}(U_t) \approx 1.

    In Transformer training experiments, a slow second-moment decay (β2=0.999\beta_2 = 0.999) causes v^xt\hat{v}_{xt} to lag behind rapid parameter evolution during early steps, causing RMS(Ut)\text{RMS}(U_t) to fluctuate significantly above 1 (peaking above 2.0). These larger-than-desired updates trigger severe training instability when learning rate warmup is omitted (collapsing BLEU to 0.1). In contrast, a fast decay (β2=0.9\beta_2 = 0.9) maintains RMS(Ut)≈1\text{RMS}(U_t) \approx 1 throughout early training but impairs long-term convergence.

  8. Knowl 8 — Translation Performance and Stability of Adafactor and Adam Variations on WMT 2014 En-De

    data/table

    The table below shows BLEU scores on newstest2013 for Transformer models trained on WMT 2014 English-to-German translation across various second-moment estimation strategies, decay rates β^2t\hat{\beta}_{2t}, first-moment decay β^1t\hat{\beta}_{1t}, and update clipping thresholds dd. Models were evaluated with warmup (st=min⁡(10−6⋅t,1/t)s_t = \min(10^{-6} \cdot t, 1/\sqrt{t})) and without warmup (st=min⁡(10−2,1/t)s_t = \min(10^{-2}, 1/\sqrt{t})).

    Factored Second-Moment β^1t\hat{\beta}_{1t} β^2t\hat{\beta}_{2t} Update Clipping dd BLEU with warmup BLEU no warmup
    (A) no 0 β2=0.999\beta_2 = 0.999 – 25.6 0.1
    (B) no 0.9 β2=0.999\beta_2 = 0.999 – 25.4 23.1
    (C) yes 0 β2=0.999\beta_2 = 0.999 – 25.4 0.2
    (D) use row-mean 0 β2=0.999\beta_2 = 0.999 – 25.2 0.3
    (E) use col-mean 0 β2=0.999\beta_2 = 0.999 – 0.3 0.5
    (F) no 0 β2=0.99\beta_2 = 0.99 – 25.0 0.4
    (G) no 0 β2=0.9\beta_2 = 0.9 – 18.4 15.6
    (H) no 0 β2=0.999\beta_2 = 0.999 1.0 25.4 21.5
    (I) no 0 β2=0.999\beta_2 = 0.999 2.0 25.7 0.2
    (J) yes 0 β2=0.999\beta_2 = 0.999 1.0 25.6 22.4
    (K) no 0 1−t−0.51 - t^{-0.5} – 25.6 21.1
    (L) no 0 1−t−0.81 - t^{-0.8} – 25.6 0.1
    (M) no 0 1−t−1.01 - t^{-1.0} – 25.4 0.1
    (N) no 0 1−t−0.81 - t^{-0.8} 1.0 25.9 22.4
    (O) yes 0 1−t−0.81 - t^{-0.8} 1.0 25.0 25.5
    (P) yes 0.9 1−t−0.81 - t^{-0.8} 1.0 24.9 25.3

    In row configurations (A) through (N), absolute step size αt=0.1⋅st\alpha_t = 0.1 \cdot s_t was used; in (O) and (P), relative step size ρt=st\rho_t = s_t was used.

    Key takeaways:

    • Factoring second moments (row C vs A, row J vs H) matches the performance of full accumulators while drastically reducing optimizer memory from O(nm)O(nm) to O(n+m)O(n+m).
    • Factoring via column means alone fails (BLEU 0.3-0.5) due to high variance across token frequencies in embedding rows.
    • Without momentum (β^1t=0\hat{\beta}_{1t} = 0), Adam without warmup collapses to 0.1 BLEU. Update clipping with d=1.0d=1.0 (row H, J, N) or decaying β^2t=1−t−0.5\hat{\beta}_{2t} = 1 - t^{-0.5} (row K) restores stability without requiring momentum storage.
    • Adafactor with relative step sizes (row O) attains 25.0 BLEU with warmup and 25.5 BLEU without warmup with sublinear memory.
  9. Knowl 9 — Sensitivity of Adam and Adafactor to Embedding Initialization and Multiplier Scales

    data/table

    When token embedding parameters are not reused in the output softmax layer of the Transformer model, their initialization standard deviation σ\sigma and architectural scaling factor influence optimizer behavior. The table below compares BLEU scores on WMT 2014 English-to-German translation for Adam and Adafactor across standard and naive embedding setups (trained for 50,000 steps with batch size 16,384 tokens):

    Embedding Init σ\sigma Multiplier BLEU (Adam) BLEU (Adafactor)
    1/dmodel1 / \sqrt{d_{\text{model}}} dmodel\sqrt{d_{\text{model}}} 26.4 26.6
    11 11 25.8 26.4
    1/dmodel1 / \sqrt{d_{\text{model}}} 11 24.2 25.4

    Under Adam, removing the manual dmodel\sqrt{d_{\text{model}}} scaling multiplier while keeping initialization σ=1/dmodel\sigma = 1 / \sqrt{d_{\text{model}}} causes BLEU to degrade from 26.4 to 24.2 (-2.2 points) because the fixed absolute step size is mismatched to the smaller embedding norm. Adafactor's parameter-relative step size automatically scales the update magnitude to match the root-mean-square norm of the embedding parameters, mitigating the degradation (25.4 BLEU, a 1.2 point improvement over Adam).

Coverage note — None was omitted; all key theoretical developments, algorithmic details, instability analyses, and empirical results from the paper are represented.

References

  1. 1.Duchi, John C., Hazan, Elad, and Singer, Yoram. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12:2121–2159, 2011. URL http://dblp.uni-trier.de/db/journals/jmlr/jmlr12.html#DuchiHS11.
  2. 2.Eckart, C. and Young, G. The approximation of one matrix by another of lower rank. Psychometrika, 1(3):211–218, 1936. doi: 10.1007/BF02288367.
  3. 3.Finesso, Lorenzo and Spreij, Peter. Nonnegative matrix factorization and I-divergence alternating minimization. Linear Algebra and its Applications, 416(2):270 – 287, 2006. ISSN 0024-3795. doi: https://doi.org/10.1016/j.laa.2005.11.012. URL http://www.sciencedirect.com/science/article/pii/S0024379505005665.
  4. 4.Goyal, Priya, Dollár, Piotr, Girshick, Ross B., Noordhuis, Pieter, Wesolowski, Lukasz, Kyrola, Aapo, Tulloch, Andrew, Jia, Yangqing, and He, Kaiming. Accurate, large minibatch SGD: training imagenet in 1 hour. CoRR, abs/1706.02677, 2017. URL http://arxiv.org/abs/1706.02677.
  5. 5.Gupta, Maya R., Bengio, Samy, and Weston, Jason. Training highly multiclass classifiers. Journal of Machine Learning Research, 15:1461–1492, 2014. URL http://jmlr.org/papers/v15/gupta14a.html.
  6. 6.Kingma, Diederik and Ba, Jimmy. Adam: A method for stochastic optimization. In ICLR, 2015.
  7. 7.Lee, Daniel D. and Seung, H. Sebastian. Learning the parts of objects by nonnegative matrix factorization. Nature, 401:788–791, 1999.
  8. 8.Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua. On the difficulty of training recurrent neural networks. In Dasgupta, Sanjoy and McAllester, David (eds.), Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pp. 1310–1318, Atlanta, Georgia, USA, 17–19 Jun 2013. PMLR. URL http://proceedings.mlr.press/v28/pascanu13.html.
  9. 9.Reddi, Sashank J., Kale, Satyen, and Kumar, Sanjiv. On the convergence of adam and beyond. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=ryQu7f-RZ.
  10. 10.Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc, Hinton, Geoffrey, and Dean, Jeff. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In ICLR, 2017. URL https://openreview.net/pdf?id=B1ckMDqlg.
  11. 11.Tieleman, T. and Hinton, G. Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning, 2012.
  12. 12.Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan N, Kaiser, Łukasz, and Polosukhin, Illia. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30, pp. 6000–6010. Curran Associates, Inc., 2017. URL http://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf.
  13. 13.Zeiler, Matthew D. Adadelta: An adaptive learning rate method. CoRR, abs/1212.5701, 2012. URL http://dblp.uni-trier.de/db/journals/corr/corr1212.html#abs-1212-5701.

Citation

MLA
Shazeer, N., and M. Stern. “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost”. arXiv, 2018, http://arxiv.org/abs/1804.04235v1.
APA
Shazeer, N., & Stern, M. (2018). Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. arXiv. http://arxiv.org/abs/1804.04235v1
Chicago
Shazeer, N., and M. Stern. 2018. “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost”. arXiv. http://arxiv.org/abs/1804.04235v1.
Harvard
Shazeer, N. and Stern, M. (2018) “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.04235v1.
Vancouver
1. Shazeer N, Stern M (2018) Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. arXiv

BibTeX

@article{shazeer2018adafactor,
  title = {Adafactor: Adaptive Learning Rates with Sublinear Memory Cost},
  author = {Shazeer, Noam and Stern, Mitchell},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.04235v1},
  eprint = {1804.04235}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/