DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule
Maor IvgiOliver HinderYair Carmon
Introduces Distance over Gradients (DoG), a parameter-free dynamic step size schedule that eliminates manual learning rate tuning in SGD while matching the empirical performance of extensively tuned optimizers across vision and language benchmarks.
Training modern machine learning models requires selecting an effective learning rate, a process that typically demands extensive manual tuning or expensive trial-and-error grid searches across multiple runs. Ineffective choices lead to poor model accuracy or training divergence, while searching over broad hyperparameter spaces incurs massive computational and financial costs. The article evaluates a new dynamic step size formula called Distance over Gradients (DoG), which automatically scales step sizes using the ratio of maximum iterate movement to cumulative gradient norms, aiming to eliminate the need to tune learning rates.
The authors analyze the method theoretically under stochastic convex optimization settings and test it empirically across a broad evaluation suite. The testbed covers 23 natural language understanding and image classification tasks across 8 neural network architectures, including Vision Transformers, ResNets, RoBERTa, and T5. They compare standard DoG and a per-layer variant (L-DoG) against standard Stochastic Gradient Descent (SGD), Adam, and other existing parameter-free methods across varied computational budgets.
The investigation produced four main findings. First, DoG achieves performance nearly identical to SGD tuned individually for each task, with the relative error difference staying below 5% across 79 of 80 fine-tuning configurations and below 1% on convex linear probes. Second, the layer-wise adaptation L-DoG substantially narrows the performance gap to tuned Adam, often matching or outperforming it when baseline tuning compute is equalized. Third, DoG consistently outperformed competing tuning-free methods such as Stochastic Polyak Step and D-Adaptation across both vision and language benchmarks. Fourth, theoretical analysis of a stabilized variant (T-DoG) proves near-optimal convergence rates with high probability, matching theoretical lower bounds up to logarithmic factors.
These findings indicate that teams can bypass extensive learning rate grid searches without degrading model accuracy. Eliminating the multi-run tuning requirement reduces computation time and cloud infrastructure expenses by factors of roughly 5 to 7. The saved computational budget can instead be redirected toward training larger models, processing larger datasets, or running longer iterations.
Organizations should consider adopting DoG or L-DoG in fine-tuning pipelines and transfer learning tasks to reduce development overhead, using default scaling initializations of 1e-4 for vision and 1e-6 to 1e-8 for language architectures. However, practitioners should exercise caution when applying the method to architectures relying heavily on batch normalization or when training complex models entirely from scratch, as preliminary tests show occasional step-size sensitivity in those conditions. Future engineering should focus on integrating DoG with momentum, weight decay, and per-parameter adaptation to improve stability across all network types.
- Paper: Adam: A Method for Stochastic Optimization, Diederik P. Kingma et al. (2015). Introduces the Adam optimization algorithm and per-parameter adaptive learning rates, providing the primary baseline and foundational context that DoG is benchmarked against and seeks to simplify.
- Paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization, John Duchi et al. (2011). Establishes adaptive step-size scaling using cumulative historical gradient norms, a core principle underlying DoG's denominator formulation.
- Paper: An overview of gradient descent optimization algorithms, Sebastian Ruder (2016). Provides a comprehensive overview of first-order stochastic optimization methods and the learning rate tuning challenges that DoG aims to automate.
- Paper: Optimization Methods for Large-Scale Machine Learning, Léon Bottou et al. (2016). Surveys the convergence theory and trade-offs of stochastic gradient methods in large-scale machine learning, contextualizing DoG's theoretical bounds.
- Paper: On the Variance of the Adaptive Learning Rate and Beyond, Liyuan Liu et al. (2019). Analyzes the high variance and instability in early adaptive step sizes, addressing problems that DoG bypasses via distance-over-gradient scaling.
- Paper: On the Convergence of Adam and Beyond, Sashank J. Reddi et al. (2018). Demonstrates theoretical failure modes and regret non-convergence in popular adaptive step size schedules like Adam, motivating parameter-free alternatives.
- Paper: Cyclical Learning Rates for Training Neural Networks, Leslie N. Smith (2017). Explores heuristic learning-rate schedules and range tests, representing the manual tuning overhead that parameter-free algorithms seek to eliminate.
- Paper: Practical Recommendations for Gradient-Based Training of Deep Architectures, Yoshua Bengio (2012). Details practical challenges and high computational costs of tuning hyperparameter schedules in deep network optimization.
No sufficiently relevant recommendations were found.
