DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule

Maor IvgiOliver HinderYair Carmon

article2023ICML96 citations

Introduces Distance over Gradients (DoG), a parameter-free dynamic step size schedule that eliminates manual learning rate tuning in SGD while matching the empirical performance of extensively tuned optimizers across vision and language benchmarks.

Listen

Training modern machine learning models requires selecting an effective learning rate, a process that typically demands extensive manual tuning or expensive trial-and-error grid searches across multiple runs. Ineffective choices lead to poor model accuracy or training divergence, while searching over broad hyperparameter spaces incurs massive computational and financial costs. The article evaluates a new dynamic step size formula called Distance over Gradients (DoG), which automatically scales step sizes using the ratio of maximum iterate movement to cumulative gradient norms, aiming to eliminate the need to tune learning rates.

The authors analyze the method theoretically under stochastic convex optimization settings and test it empirically across a broad evaluation suite. The testbed covers 23 natural language understanding and image classification tasks across 8 neural network architectures, including Vision Transformers, ResNets, RoBERTa, and T5. They compare standard DoG and a per-layer variant (L-DoG) against standard Stochastic Gradient Descent (SGD), Adam, and other existing parameter-free methods across varied computational budgets.

The investigation produced four main findings. First, DoG achieves performance nearly identical to SGD tuned individually for each task, with the relative error difference staying below 5% across 79 of 80 fine-tuning configurations and below 1% on convex linear probes. Second, the layer-wise adaptation L-DoG substantially narrows the performance gap to tuned Adam, often matching or outperforming it when baseline tuning compute is equalized. Third, DoG consistently outperformed competing tuning-free methods such as Stochastic Polyak Step and D-Adaptation across both vision and language benchmarks. Fourth, theoretical analysis of a stabilized variant (T-DoG) proves near-optimal convergence rates with high probability, matching theoretical lower bounds up to logarithmic factors.

These findings indicate that teams can bypass extensive learning rate grid searches without degrading model accuracy. Eliminating the multi-run tuning requirement reduces computation time and cloud infrastructure expenses by factors of roughly 5 to 7. The saved computational budget can instead be redirected toward training larger models, processing larger datasets, or running longer iterations.

Organizations should consider adopting DoG or L-DoG in fine-tuning pipelines and transfer learning tasks to reduce development overhead, using default scaling initializations of 1e-4 for vision and 1e-6 to 1e-8 for language architectures. However, practitioners should exercise caution when applying the method to architectures relying heavily on batch normalization or when training complex models entirely from scratch, as preliminary tests show occasional step-size sensitivity in those conditions. Future engineering should focus on integrating DoG with momentum, weight decay, and per-parameter adaptation to improve stability across all network types.

No sufficiently relevant recommendations were found.

Cover for DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule

Abstract

We propose a tuning-free dynamic SGD step size formula, which we call Distance over Gradients (DoG). The DoG step sizes depend on simple empirical quantities (distance from the initial point and norms of gradients) and have no “learning rate” parameter. Theoretically, we show that, for stochastic convex optimization, a slight variation of the DoG formula enjoys strong, high-probability parameter-free convergence guarantees and iterate movement bounds. Empirically, we consider a broad range of vision and language transfer learning tasks, and show that DoG’s performance is close to that of SGD with tuned learning rate. We also propose a per-layer variant of DoG that generally outperforms tuned SGD, approaching the performance of tuned Adam. A PyTorch implementation of our algorithms is available at https://github.com/formll/dog.

Table of Contents

  • 1. Introduction
  • 1.1. Summary of results
  • 2. Algorithm Derivation
  • 3. Theoretical Analysis
  • 3.1. Preliminaries
  • 3.2. Optimality gap bounds assuming bounded iterates
  • 3.3. Iterate stability bound
  • 4. Experiments
  • 4.1. Fine-tuning testbed
  • 4.2. Comparison of fine-tuning performance
  • 4.3. Sensitivity of DOG's fixed parameters
  • 4.4. Convex optimization
  • 4.5. Fine-tuning on ImageNet
  • 4.6. Training from scratch
  • 4.7. Comparison to other tuning-free methods
  • 5. Related Work
  • 6. Limitations and Outlook
  • Acknowledgments
  • References
  • A. Relaxing the Convexity Assumption
  • B. Relaxing the Global Stochastic Gradient Bound Assumption
  • C.2. Lemma C.2
  • C. Useful Algebraic Facts
  • C.1. Lemma C.1
  • C.3. Proof of Lemma 3.7
  • C.4. Lemma C.3
  • D. Proofs for Section 3
  • D.1. Proof of Lemma 3.4
  • D.2. Proof of Lemma 3.5
  • D.3. Proof of Corollary 3.8
  • D.4. DOG can Diverge on a Pathological Instance
  • D.5. Proof of Proposition 3.9
  • D.6. Illustrating DOG's guarantees for least squares problems
  • E. Experiment Details
  • E.1. Environment settings
  • E.2. Implementation details
  • E.3. Datasets
  • E.4. Models
  • E.5. Hyper-parameters
  • E.6. Figure 1 details
  • E.7. Fine-tuning ImageNet
  • E.8. Training from scratch
  • F. Additional experiment results
  • F.1. Full breakdown of main experiment results
  • F.2. Comparison with equalized compute budget
  • F.3. Fine-tuning CoLA
  • F.4. Sensitivity of DOG to rϵ and the effect of batch normalization
  • F.5. Additional convex optimization results
  • F.6. The growth rate of r̄t
  • G. Comparison to Other Tuning-Free Methods
  • G.1. Parameter-free SGD
  • G.2. Stochastic Polyak step-size
  • G.3. D-adaptation

Knowls

  1. Knowl 1 — Distance-over-Gradients step-size schedule

    algorithm

    Distance over Gradients (DoG) is a parameter-free dynamic schedule for stochastic gradient descent. Let xt∈Rmx_t\in\mathbb{R}^m be the parameter vector, let gtg_t be the stochastic gradient evaluated at xtx_t, and let x0x_0 be the initialization. The update is

    xt+1=Proj⁡X(xt−ηtgt),x_{t+1}=\operatorname{Proj}_{\mathcal X}(x_t-\eta_t g_t),

    where projection is omitted in the unconstrained case. Define the observed movement and cumulative squared-gradient magnitude by rt=∥xt−x0∥r_t=\|x_t-x_0\| and Gt=∑i=0t∥gi∥2G_t=\sum_{i=0}^{t}\|g_i\|^2. For t≥1t\ge 1, DoG uses

    ηt=max⁡0≤i≤t∥xi−x0∥∑i=0t∥gi∥2=max⁡0≤i≤triGt.\eta_t=\frac{\max_{0\le i\le t}\|x_i-x_0\|}{\sqrt{\sum_{i=0}^{t}\|g_i\|^2}}=\frac{\max_{0\le i\le t}r_i}{\sqrt{G_t}}.

    The first step is initialized with a user-specified movement scale rϵ>0r_\epsilon>0 as η0=rϵ/∥g0∥\eta_0=r_\epsilon/\|g_0\|, producing a normalized update of length rϵr_\epsilon. Apart from this requirement that rϵr_\epsilon be sufficiently small, DoG has no multiplicative learning-rate parameter. Its state consists of the current parameters, the running maximum distance from initialization, and the scalar accumulator GtG_t, so each iteration requires only the gradient and its norm in addition to the SGD update.

  2. Knowl 2 — Implicit optimality rationale for the unit DoG scale

    model/method

    The schedule is motivated by an implicit criterion for a constant-step-size SGD run. For a fixed step size η>0\eta>0 over TT iterations, let xkx_k and gkg_k denote the resulting iterates and stochastic gradients. The parameter-free SGD analysis used by the paper states that if, for some c∈(0,1)c\in(0,1),

    η=c max⁡0≤k≤T∥xk−x0∥∑k=0T∥gk∥2,\eta=c\,\frac{\max_{0\le k\le T}\|x_k-x_0\|}{\sqrt{\sum_{k=0}^{T}\|g_k\|^2}},

    then the excess loss of an averaged iterate is at most a factor 1/[c(1−c/2)]1/[c(1-c/2)] larger than the worst-case optimally tuned SGD bound. This condition can only be checked after completing the run, so the paper makes it explicit online by imposing the same relation at each iteration. DoG corresponds to the threshold value c=1c=1; the authors argue that this should be close to the largest stable scale and empirically find that it performs best near one. Unlike solving the implicit condition by bisection, the resulting schedule requires only one optimization run.

  3. Knowl 3 — High-probability convergence when the iterates remain bounded

    theoretical result

    Consider projected SGD on a closed convex set X⊆Rm\mathcal X\subseteq\mathbb R^m for a convex function ff with minimizer x⋆x^\star, where the stochastic-gradient oracle satisfies E[gt∣xt]∈∂f(xt)\mathbb E[g_t\mid x_t]\in\partial f(x_t) and ∥gt∥≤L\|g_t\|\le L almost surely. Let d0=∥x0−x⋆∥d_0=\|x_0-x^\star\|, let rϵ>0r_\epsilon>0, define rt=∥xt−x0∥r_t=\|x_t-x_0\|, rˉt=max⁡{rϵ,r0,…,rt}\bar r_t=\max\{r_\epsilon,r_0,\ldots,r_t\}, and Gt=∑i=0t∥gi∥2G_t=\sum_{i=0}^{t}\|g_i\|^2. A DOG-like schedule has the form ηt=rˉt/Gt′\eta_t=\bar r_t/\sqrt{G'_t}, where Gt′G'_t is a positive, nondecreasing, data-dependent sequence satisfying Gt′≥GtG'_t\ge G_t.

    Define the movement-weighted average

    xˉt=∑k=0t−1rˉkxk∑k=0t−1rˉk,\bar x_t=\frac{\sum_{k=0}^{t-1}\bar r_k x_k}{\sum_{k=0}^{t-1}\bar r_k},

    and θt,δ=log⁡ ⁣(60log⁡(6t)/δ)\theta_{t,\delta}=\log\!\bigl(60\log(6t)/\delta\bigr) for failure probability δ∈(0,1)\delta\in(0,1). With probability at least 1−δ1-\delta, simultaneously for every t≤Tt\le T,

    f(xˉt)−f(x⋆)=O ⁣((d0+rˉt)Gt−1′+Gt−1θt,δ+L2θt,δ2∑i<trˉi/rˉt).f(\bar x_t)-f(x^\star)=O\!\left(\frac{(d_0+\bar r_t)\sqrt{G'_{t-1}}+G_{t-1}\theta_{t,\delta}+L^2\theta_{t,\delta}^2}{\sum_{i<t}\bar r_i/\bar r_t}\right).

    Thus, if the trajectory satisfies rˉT=O(d0)\bar r_T=O(d_0) and the un-tamed DoG accumulator obeys GT≤L2TG_T\le L^2T, the resulting rate is O ⁣(d0L θT,δlog⁡+(rˉT/rϵ)/T)O\!\bigl(d_0L\,\theta_{T,\delta}\log_+(\bar r_T/r_\epsilon)/\sqrt T\bigr), which is optimal for stochastic convex optimization up to logarithmic factors. Here log⁡+(z)=1+log⁡z\log_+(z)=1+\log z and the universal constants hidden by O(⋅)O(\cdot) do not depend polynomially on the unknown problem scale.

  4. Knowl 4 — T-DOG guarantees stable iterates and parameter-free convergence

    theoretical result

    The paper introduces T-DOG, a tamed DoG-like schedule designed to guarantee bounded intermediate iterates. Under convexity and almost-surely bounded stochastic gradients ∥gt∥≤L\|g_t\|\le L, define G−1=0G_{-1}=0 and

    Gt′=84θT,δlog⁡+2(t+1)(Gt−1+16θT,δL2),ηt=rˉtGt′,G'_t=8^4\theta_{T,\delta}\log_+^2(t+1)\left(G_{t-1}+16\theta_{T,\delta}L^2\right), \qquad \eta_t=\frac{\bar r_t}{\sqrt{G'_t}},

    where Gt=∑i=0t∥gi∥2G_t=\sum_{i=0}^{t}\|g_i\|^2, rˉt=max⁡{rϵ,max⁡i≤t∥xi−x0∥}\bar r_t=\max\{r_\epsilon,\max_{i\le t}\|x_i-x_0\|\}, and θT,δ=log⁡ ⁣(60log⁡(6T)/δ)\theta_{T,\delta}=\log\!\bigl(60\log(6T)/\delta\bigr). For any iteration budget TT, failure probability δ∈(0,1)\delta\in(0,1), and initialization scale satisfying rϵ≤3d0r_\epsilon\le 3d_0 with d0=∥x0−x⋆∥d_0=\|x_0-x^\star\|, T-DOG obeys

    Pr⁡ ⁣(rˉT>3d0)≤δ.\Pr\!\left(\bar r_T>3d_0\right)\le\delta.

    Moreover, let τ\tau maximize ∑i<τrˉi/rˉτ\sum_{i<\tau}\bar r_i/\bar r_\tau. With probability at least 1−2δ1-2\delta, the weighted average xˉτ\bar x_\tau satisfies

    f(xˉτ)−f⋆≤O ⁣(cδ,rϵ,Td0Gτ+L2T)≤O ⁣(cδ,rϵ,Td0LT),f(\bar x_\tau)-f^\star \le O\!\left(c_{\delta,r_\epsilon,T}\frac{d_0\sqrt{G_\tau+L^2}}{T}\right) \le O\!\left(c_{\delta,r_\epsilon,T}\frac{d_0L}{\sqrt T}\right),

    where f⋆=f(x⋆)f^\star=f(x^\star) and

    cδ,rϵ,T=log⁡+(T) log⁡+ ⁣(d0rϵ)log⁡ ⁣(log⁡+(T)δ).c_{\delta,r_\epsilon,T}=\log_+(T)\,\log_+\!\left(\frac{d_0}{r_\epsilon}\right)\log\!\left(\frac{\log_+(T)}{\delta}\right).

    T-DOG therefore has a high-probability, parameter-free stochastic-convex convergence rate that is optimal up to logarithmic factors, while also controlling every intermediate iterate rather than only the final output.

  5. Knowl 5 — Layer-wise DoG for neural-network parameters

    model/method

    Layer-wise DoG (L-DoG) applies the DoG rule independently to each parameter tensor. For layer ℓ\ell, let xtℓx_t^\ell be its weights and gtℓg_t^\ell its stochastic gradient at iteration tt. The layer-specific step size is

    ηtℓ=max⁡0≤i≤t∥xiℓ−x0ℓ∥∑i=0t∥giℓ∥2+ε,ε=10−8,\eta_t^\ell=\frac{\max_{0\le i\le t}\|x_i^\ell-x_0^\ell\|}{\sqrt{\sum_{i=0}^{t}\|g_i^\ell\|^2+\varepsilon}}, \qquad \varepsilon=10^{-8},

    where ε\varepsilon is only a numerical-stability constant. Each layer is updated using its own ηtℓ\eta_t^\ell. L-DoG retains the tuning-free character of DoG while allowing different layers to move at different scales; the paper provides no theoretical convergence guarantee for this layer-wise variant.

  6. Knowl 6 — Fine-tuning evaluation protocol and normalized error metric

    experimental setup

    The main evaluation fine-tuned pretrained language and vision models on a broad suite of natural-language and image-classification tasks. The language models were RoBERTa-base and T5-base evaluated on GLUE tasks and SQuAD 1.1; the vision models were VGG11, ResNet50, DenseNet121, ViT-B/32, and ConvNeXt-T evaluated on VTAB tasks. SGD and Adam used cosine decay and a sweep over peak learning rates; language runs additionally used 10% warm-up and global gradient-norm clipping at 1, whereas DoG and L-DoG used neither warm-up nor annealing. Runs generally used five random seeds, no weight decay, and polynomial parameter averaging with fixed coefficient γ=8\gamma=8, retaining whichever of the raw or averaged checkpoint performed better.

    For vision, the initial movement scale was rϵ=10−4(1+∥x0∥)r_\epsilon=10^{-4}(1+\|x_0\|); for language, it was 10−6(1+∥x0∥)10^{-6}(1+\|x_0\|) for DoG and 10−8(1+∥x0∥)10^{-8}(1+\|x_0\|) for L-DoG. To compare heterogeneous task metrics, if err⁡x\operatorname{err}_x is the error of optimizer xx and err⁡DoG\operatorname{err}_{\mathrm{DoG}} is the error of DoG on the same model-task pair, the paper defines relative error difference (RED) as

    RED⁡(err⁡x,err⁡DoG)=err⁡DoG−err⁡xerr⁡DoG.\operatorname{RED}(\operatorname{err}_x,\operatorname{err}_{\mathrm{DoG}})=\frac{\operatorname{err}_{\mathrm{DoG}}-\operatorname{err}_x}{\operatorname{err}_{\mathrm{DoG}}}.

    Positive RED means that optimizer xx has lower error than DoG; negative RED means that DoG is better. The page-7 aggregate comparisons report RED across tasks after averaging over seeds for each model-task pair.

  7. Knowl 7 — DoG matches tuned SGD and L-DoG narrows the Adam gap

    empirical result

    Across the main fine-tuning testbed, DoG performed similarly to well-tuned SGD on 79 of 80 model-task combinations. The sole major failure was T5-base on CoLA, where DoG behaved erratically under the default rϵr_\epsilon; substantially reducing rϵr_\epsilon improved the result. On convex linear-probe experiments, the difference between DoG and instance-tuned SGD was below 1% RED, corresponding to examples such as 90% versus 90.1% accuracy.

    Tuned Adam was generally stronger than global DoG, particularly for ResNet50 and ConvNeXt-T, which the authors associate with Adam’s per-parameter scaling and momentum. L-DoG had positive median RED relative to DoG for every model family in the aggregate page-7 comparison and substantially reduced the gap to Adam. Instance-tuned SGD and Adam required roughly 6–7 times more total training computation because several learning rates were trained; when the compute budget was equalized, DoG usually outperformed instance-tuned SGD and L-DoG moved closer to Adam.

  8. Knowl 8 — Sensitivity to initialization scale and multiplicative step factor

    empirical result

    The page-8 sensitivity experiments support the claim that DoG is largely insensitive to the initial movement scale when that scale is sufficiently small. In 7 of 8 tested model-task combinations, performance remained stable over a wide range of rϵr_\epsilon. The principal exception was ResNet50 fine-tuned on CIFAR-100: excessively small rϵr_\epsilon caused an accuracy drop, and the authors hypothesize that batch-normalization scale invariance makes multiple step sizes appear compatible with the implicit DoG criterion. Disabling batch normalization restored the stabilizing behavior and robustness to small rϵr_\epsilon.

    The same experiments multiplied the DoG step size by a constant cc, using ηt=c max⁡i≤t∥xi−x0∥/∑i≤t∥gi∥2\eta_t=c\,\max_{i\le t}\|x_i-x_0\|/\sqrt{\sum_{i\le t}\|g_i\|^2}. Values near c=1c=1 performed best; the useful range was approximately [0.5,1.5][0.5,1.5]. Smaller values slowed optimization, while larger values diverged in 6 of 8 cases. This narrow empirically useful range supports the paper’s choice of the parameter-free value c=1c=1 rather than introducing a tunable learning-rate multiplier.

  9. Knowl 9 — ImageNet fine-tuning and CIFAR-10 training-from-scratch results

    empirical result

    In the page-8 ImageNet experiment, a CLIP ViT-B/32 model was fine-tuned for 25,000 steps. With polynomial averaging, SGD’s best reported accuracy was 77.54%77.54\% at learning rate 3×10−23\times10^{-2}, DoG reached 77.22%77.22\%, AdamW reached 79.01%79.01\% at learning rate 3×10−53\times10^{-5}, and L-DoG reached 80.12%80.12\%. Without averaging, the corresponding reported values were 77.51%77.51\% for SGD, 74.78%74.78\% for DoG, 79.04%79.04\% for AdamW, and 78.20%78.20\% for L-DoG. Thus L-DoG exceeded the best AdamW result by 1.11 percentage points in this setup, although the authors note that the iteration budget may have been insufficient for the other methods.

    In the page-25 CIFAR-10 experiment, a Wide ResNet-28-10 was trained from scratch for 200 epochs. With averaging, canonical momentum-SGD using learning rate 0.10.1 reached 88.5%88.5\%, plain DoG reached 96.4%96.4\%, and L-DoG reached 93.5%93.5\%; without averaging, the same rows reached 96.3%96.3\%, 85.2%85.2\%, and 83.2%83.2\%, respectively. DoG therefore matched the best tuned training prescription after averaging, while Adam was weaker in this training-from-scratch setting.

  10. Knowl 10 — Known limitations and open practical issues

    limitation

    Unmodified DoG is not guaranteed to remain stable on every problem. The paper gives a nonsmooth convex construction in which, for T≤mT\le m, the maximum distance from initialization grows as rˉT=rϵT\bar r_T=r_\epsilon\sqrt T, while the initial distance to an optimum is d0=10rϵd_0=10r_\epsilon; consequently rˉT/d0=T/10\bar r_T/d_0=\sqrt T/10 can become arbitrarily large. T-DOG addresses this pathological behavior theoretically, but it requires an iteration budget TT, failure probability δ\delta, and a stochastic-gradient norm bound LL.

    The layer-wise method L-DoG has no theoretical guarantee, and the practical behavior of both variants can depend on normalization layers and the choice of rϵr_\epsilon. The paper does not establish a principled way to combine DoG with momentum, per-parameter preconditioning, or learning-rate annealing. Its empirical evidence is concentrated on fine-tuning pretrained models, with only preliminary training-from-scratch experiments; broader architectures, normalization schemes, and training regimes remain open settings.

Coverage note — Proof-only lemmas and derivations, the appendix’s detailed per-task tables, convexity relaxations, local-gradient-bound extension, least-squares specialization, and preliminary SPS/D-Adaptation comparisons were omitted because they are supporting or lower-significance material relative to the ten load-bearing contributions above.

References

  1. 1.Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mane, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viegas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL https://www.tensorflow.org/. Software available from tensorflow.org.
  2. 2.Agarwal, A., Bartlett, P. L., Ravikumar, P., and Wainwright, M. J. Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization. IEEE Transactions on Information Theory, 58(5):3235–3249, 2012.
  3. 3.Arrow, K. J. and Enthoven, A. C. Quasi-concave programming. Econometrica: Journal of the Econometric Society, pp. 779–800, 1961.
  4. 4.Asi, H. and Duchi, J. C. The importance of better models in stochastic optimization. Proceedings of the National Academy of Sciences, 116(46):22924–22930, 2019.
  5. 5.Bar Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I. The second PASCAL recognising textual entailment challenge. In Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, 2006.
  6. 6.Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Kuttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al. Deepmind lab. arXiv:1612.03801, 2016.
  7. 7.Bentivogli, L., Dagan, I., Dang, H. T., Giampiccolo, D., and Magnini, B. The fifth PASCAL recognizing textual entailment challenge. In Text Analysis Conference (TAC), 2009.
  8. 8.Bernstein, J. R., Vahdat, A., Yue, Y., and Liu, M.-Y. On the distance between two neural networks and the stability of learning. arXiv:2002.03432, 2020.
  9. 9.Berrada, L., Zisserman, A., and Kumar, M. P. Training neural networks for and by interpolation. In International Conference on Machine Learning (ICML), 2020.
  10. 10.Bhaskara, A., Cutkosky, A., Kumar, R., and Purohit, M. Online learning with imperfect hints. In International Conference on Machine Learning (ICML), 2020.
  11. 11.Carmon, Y. and Hinder, O. Making SGD parameter-free. In Conference on Learning Theory (COLT), 2022.
  12. 12.Carmon, Y., Raghunathan, A., Schmidt, L., Duchi, J. C., and Liang, P. S. Unlabeled data improves adversarial robustness. Advances in Neural Information Processing Systems (NeurIPS), 2019.
  13. 13.Cer, D. M., Diab, M. T., Agirre, E., Lopez-Gazpio, I., and Specia, L. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In International Workshop on Semantic Evaluation, 2017.
  14. 14.Chandra, K., Xie, A., Ragan-Kelley, J., and Meijer, E. Gradient descent: The ultimate optimizer. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  15. 15.Chen, K., Langford, J., and Orabona, F. Better parameter-free stochastic optimization with ODE updates for coin-betting. In AAAI Conference on Artificial Intelligence, 2022.
  16. 16.Cheng, G., Han, J., and Lu, X. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, 2017.
  17. 17.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A. Describing textures in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
  18. 18.Comet.ML. Comet.ML home page, 2021. URL https://www.comet.ml/.
  19. 19.Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. AutoAugment: Learning augmentation strategies from data. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  20. 20.Cutkosky, A. Artificial constraints and hints for unbounded online learning. In Conference on Learning Theory (COLT), 2019.
  21. 21.Cutkosky, A. and Orabona, F. Black-box reductions for parameter-free online learning in Banach spaces. In Conference on Learning Theory (COLT), 2018.
  22. 22.Dagan, I., Glickman, O., and Magnini, B. The PASCAL recognising textual entailment challenge. In Machine learning challenges. Evaluating predictive uncertainty, visual object classification, and recognising tectual entailment. Springer, 2006.
  23. 23.Davis, D. and Drusvyatskiy, D. Stochastic model-based minimization of weakly convex functions. SIAM Journal on Optimization, 29(1):207–239, 2019.
  24. 24.Defazio, A. and Mishchenko, K. Parameter free dual averaging: Optimizing lipschitz functions in a single pass. In OPT 2022: NeurIPS Workshop on Optimization for Machine Learning, 2022.
  25. 25.Defazio, A. and Mishchenko, K. Learning-rate-free learning by D-adaptation. In International Conference on Machine Learning (ICML), 2023.
  26. 26.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  27. 27.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In International Workshop on Paraphrasing (IWP2005), 2005.
  28. 28.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  29. 29.Duchi, J., Hazan, E., and Singer, Y. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12(7), 2011.
  30. 30.Faw, M., Tziotis, I., Caramanis, C., Mokhtari, A., Shakkottai, S., and Ward, R. The power of adaptivity in SGD: Self-tuning step sizes with unbounded gradients and affine variance. In Conference on Learning Theory (COLT), 2022.
  31. 31.Fei-Fei, L., Fergus, R., and Perona, P. Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories. In CVPR Workshop, 2004.
  32. 32.Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. The third PASCAL recognizing textual entailment challenge. In ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, 2007.
  33. 33.Gupta, V., Koren, T., and Singer, Y. A unified approach to adaptive regularization in online and stochastic optimization. arXiv:1706.06569, 2017.
  34. 34.Gupta, V., Koren, T., and Singer, Y. Shampoo: Preconditioned stochastic tensor optimization. In International Conference on Machine Learning (ICML), 2018.
  35. 35.Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gerard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E. Array programming with NumPy. Nature, 585(7825):357–362, 2020.
  36. 36.Hazan, E. and Kakade, S. Revisiting the Polyak step size. arXiv:1905.00313, 2019.
  37. 37.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  38. 38.Hinder, O., Sidford, A., and Sohoni, N. Near-optimal methods for minimizing star-convex functions and beyond. In Conference on Learning Theory (COLT), 2020.
  39. 39.Howard, S. R., Ramdas, A., McAuliffe, J., and Sekhon, J. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49(2):1055–1080, 2021.
  40. 40.Huang, G., Liu, Z., and Weinberger, K. Q. Densely connected convolutional networks. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  41. 41.Iyer, S., Dandekar, N., and Csernai, K. First quora dataset release: Question pairs, 2017. URL https://data.quora.com/First-Quora-Dataset-Release-Question-Pairs.
  42. 42.Jacobsen, A. and Cutkosky, A. Parameter-free mirror descent. In Conference on Learning Theory (COLT), 2022.
  43. 43.Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  44. 44.Kaggle and EyePacs. Kaggle diabetic retinopathy detection, 2015. URL https://www.kaggle.com/c/diabetic-retinopathy-detection/data.
  45. 45.Karimi, H., Nutini, J., and Schmidt, M. Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD), 2016.
  46. 46.Kempka, M., Kotlowski, W., and Warmuth, M. K. Adaptive scale-invariant online algorithms for learning linear models. In International Conference on Machine Learning (ICML), 2019.
  47. 47.Kingma, D. P. and Ba, J. ADAM: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  48. 48.Kleinberg, B., Li, Y., and Yuan, Y. An alternative view: When does SGD escape local minima? In International Conference on Machine Learning (ICML), 2018.
  49. 49.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  50. 50.Levesque, H. J., Davis, E., and Morgenstern, L. The Winograd schema challenge. In International Conference on Principles of Knowledge Representation and Reasoning, 2011.
  51. 51.Levy, D., Carmon, Y., Duchi, J. C., and Sidford, A. Large-scale methods for distributionally robust optimization. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  52. 52.Lhoest, Q., Villanova del Moral, A., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., Davison, J., Saško, M., Chhablani, G., Malik, B., Brandeis, S., Le Scao, T., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Schmid, P., Gugger, S., Delangue, C., Matussiere, T., Debut, L., Bekman, S., Cistac, P., Goehringer, T., Mustar, V., Lagunas, F., Rush, A., and Wolf, T. Datasets: A community library for natural language processing. In Conference on Empirical Methods in Natural Language Processing (EMNLP): System Demonstrations, 2021.
  53. 53.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692, 2019.
  54. 54.Liu, Z., Mao, H., Wu, C., Feichtenhofer, C., Darrell, T., and Xie, S. A ConvNet for the 2020s. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  55. 55.Loizou, N., Vaswani, S., Laradji, I. H., and Lacoste-Julien, S. Stochastic Polyak step-size for SGD: An adaptive learning rate for fast convergence. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2021.
  56. 56.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  57. 57.Luo, H. and Schapire, R. E. Achieving all with no parameters: AdaNormalHedge. In Conference on Learning Theory (COLT), 2015.
  58. 58.Mangasarian, O. L. Pseudo-convex functions. In Stochastic optimization models in finance. Elsevier, 1975.
  59. 59.Matthey, L., Higgins, I., Hassabis, D., and Lerchner, A. dSprites: Disentanglement testing Sprites dataset. https://github.com/deepmind/dsprites-dataset/, 2017.
  60. 60.Mhammedi, Z. and Koolen, W. M. Lipschitz and comparator-norm adaptivity in online learning. In Conference on Learning Theory (COLT), 2020.
  61. 61.Nemirovski, A. On parallel complexity of nonsmooth convex optimization. Journal of Complexity, 10(4):451–463, 1994.
  62. 62.Nemirovski, A. and Yudin, D. Problem complexity and method efficiency in optimization. Wiley-Interscience, 1983.
  63. 63.Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. Robust stochastic approximation approach to stochastic programming. SIAM Journal on Optimization, 19(4):1574–1609, 2009.
  64. 64.Nesterov, Y. and Polyak, B. T. Cubic regularization of Newton method and its global performance. Mathematical Programming, 108(1):177–205, 2006.
  65. 65.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011.
  66. 66.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP), 2008.
  67. 67.Orabona, F. Simultaneous model selection and optimization through parameter-free stochastic learning. Advances in Neural Information Processing Systems (NeurIPS), 2014.
  68. 68.Orabona, F. and Pal, D. Coin betting and parameter-free online learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016.
  69. 69.Orabona, F. and Pal, D. Parameter-free stochastic optimization of variationally coherent functions. arXiv:2102.00236, 2021.
  70. 70.Orabona, F. and Tommasi, T. Training deep networks without learning rates through coin betting. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  71. 71.Paquette, C. and Scheinberg, K. A stochastic line search method with expected complexity analysis. SIAM Journal on Optimization, 30(1):349–376, 2020.
  72. 72.Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. Cats and dogs. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  73. 73.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  74. 74.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  75. 75.Polyak, B. T. Introduction to Optimization. Optimization Software, Inc, 1987.
  76. 76.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  77. 77.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020.
  78. 78.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2016.
  79. 79.Reddi, S. J., Kale, S., and Kumar, S. On the convergence of Adam and beyond. In International Conference on Learning Representations (ICLR), 2018.
  80. 80.Rolinek, M. and Martius, G. L4: Practical loss-based step-size adaptation for deep learning. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  81. 81.Schaul, T., Zhang, S., and LeCun, Y. No more pesky learning rates. In International Conference on Machine Learning (ICML), 2013.
  82. 82.Shamir, O. and Zhang, T. Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes. In International Conference on Machine Learning (ICML), 2013.
  83. 83.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning (ICML), 2018.
  84. 84.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556, 2014.
  85. 85.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2013.
  86. 86.Streeter, M. and McMahan, H. B. No-regret algorithms for unconstrained online convex optimization. In Advances in Neural Information Processing Systems (NeurIPS), 2012.
  87. 87.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. Going deeper with convolutions. In Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  88. 88.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  89. 89.Vaswani, S., Mishkin, A., Laradji, I., Schmidt, M., Gidel, G., and Lacoste-Julien, S. Painless stochastic gradient: Interpolation, line-search, and convergence rates. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  90. 90.Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nature Methods, 17:261–272, 2020.
  91. 91.Vovk, V. On-line regression competitive with reproducing kernel hilbert spaces. In Theory and Applications of Models of Computation (TAMC), 2006.
  92. 92.Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. SuperGLUE: A stickier benchmark for general-purpose language understanding. In Advances in Neural Information Processing Systems (NeurIPS), 2019a.
  93. 93.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations (ICLR), 2019b.
  94. 94.Ward, R., Wu, X., and Bottou, L. AdaGrad stepsizes: Sharp convergence over nonconvex landscapes. In International Conference on Machine Learning (ICML), 2019.
  95. 95.Warstadt, A., Singh, A., and Bowman, S. R. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019.
  96. 96.Wes McKinney. Data Structures for Statistical Computing in Python. In Proceedings of the 9th Python in Science Conference, 2010.
  97. 97.Wightman, R. PyTorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  98. 98.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2018.
  99. 99.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Conference on Empirical Methods in Natural Language Processing (EMNLP): System Demonstrations, 2020.
  100. 100.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning (ICML), 2022.
  101. 101.Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A. Sun database: Large-scale scene recognition from abbey to zoo. In Conference on Computer Vision and Pattern Recognition (CVPR), 2010.
  102. 102.Xiao, J., Ehinger, K. A., Hays, J., Torralba, A., and Oliva, A. Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision, 119(1):3–22, 2016.
  103. 103.You, Y., Gitman, I., and Ginsburg, B. Large batch training of convolutional networks. arXiv:1708.03888, 2017a.
  104. 104.You, Y., Gitman, I., and Ginsburg, B. Scaling SGD batch size to 32k for ImageNet training. arXiv:1708.03888, 2017b.
  105. 105.You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J. Large batch optimization for deep learning: Training BERT in 76 minutes. In International Conference on Learning Representations (ICLR), 2020.
  106. 106.Zagoruyko, S. and Komodakis, N. Wide residual networks. In British Machine Vision Conference (BMVC), 2016.
  107. 107.Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., Beyer, L., Bachem, O., Tschannen, M., Michalski, M., Bousquet, O., Gelly, S., and Houlsby, N. A large-scale study of representation learning with the visual task adaptation benchmark. arXiv:1910.04867, 2019.
  108. 108.Zhang, J. and Cutkosky, A. Parameter-free regret in high probability with heavy tails. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  109. 109.Zhang, Z., Cutkosky, A., and Paschalidis, I. PDE-based optimal strategy for unconstrained online learning. In International Conference on Machine Learning (ICML), 2022.
  110. 110.Zhou, Y., Yang, J., Zhang, H., Liang, Y., and Tarokh, V. SGD converges to global minimum in deep learning via star-convex path. In International Conference on Learning Representations (ICLR), 2019.

Citation

MLA
Ivgi, M., et al. “DoG Is SGD’s Best Friend: A Parameter-Free Dynamic Step Size Schedule”. International Conference on Machine Learning, vol. 202, 2023, pp. 14465–99, https://proceedings.mlr.press/v202/ivgi23a.html.
APA
Ivgi, M., Hinder, O., & Carmon, Y. (2023). DoG is SGD’s Best Friend: A Parameter-Free Dynamic Step Size Schedule. International Conference on Machine Learning, 202, 14465–14499. https://proceedings.mlr.press/v202/ivgi23a.html
Chicago
Ivgi, M., O. Hinder, and Y. Carmon. 2023. “DoG Is SGD’s Best Friend: A Parameter-Free Dynamic Step Size Schedule”. International Conference on Machine Learning 202: 14465–99. https://proceedings.mlr.press/v202/ivgi23a.html.
Harvard
Ivgi, M., Hinder, O. and Carmon, Y. (2023) “DoG is SGD’s Best Friend: A Parameter-Free Dynamic Step Size Schedule”, International Conference on Machine Learning. PMLR, pp. 14465–14499. Available at: https://proceedings.mlr.press/v202/ivgi23a.html.
Vancouver
1. Ivgi M, Hinder O, Carmon Y (2023) DoG is SGD’s Best Friend: A Parameter-Free Dynamic Step Size Schedule. In: International Conference on Machine Learning. PMLR, pp 14465–14499

BibTeX

@InProceedings{pmlr-v202-ivgi23a,
  title = 	 {{D}o{G} is {SGD}’s Best Friend: A Parameter-Free Dynamic Step Size Schedule},
  author =       {Ivgi, Maor and Hinder, Oliver and Carmon, Yair},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {14465--14499},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ivgi23a/ivgi23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ivgi23a.html},
  abstract = 	 {We propose a tuning-free dynamic SGD step size formula, which we call Distance over Gradients (DoG). The DoG step sizes depend on simple empirical quantities (distance from the initial point and norms of gradients) and have no “learning rate” parameter. Theoretically, we show that, for stochastic convex optimization, a slight variation of the DoG formula enjoys strong, high-probability parameter-free convergence guarantees and iterate movement bounds. Empirically, we consider a broad range of vision and language transfer learning tasks, and show that DoG’s performance is close to that of SGD with tuned learning rate. We also propose a per-layer variant of DoG that generally outperforms tuned SGD, approaching the performance of tuned Adam. A PyTorch implementation of our algorithms is available at https://github.com/formll/dog.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/