Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities

Brian R. BartoldsonBhavya KailkhuraDavis W. Blalock

article2023JMLR68 citations

Presents a unified taxonomy and rigorous evaluation framework for algorithmically efficient deep learning training techniques, equipping practitioners to systematically identify bottlenecks and combine methods to cut computational costs.

Listen

Modern deep learning relies on training massive neural network models on vast datasets. While this strategy achieves state-of-the-art performance in domains such as computer vision and natural language processing, model sizes are doubling as rapidly as every 4 to 8 months. This rapid scaling leads to unsustainable economic costs, substantial carbon emissions, and resource concentration among a few well-funded entities. Although hardware and software engineering provide efficiency gains, relying exclusively on them is insufficient to bridge the gap to theoretical computing limits. The article formalizes the algorithmic speedup challenge, creates a structured taxonomy of methods that modify the training routine semantics, identifies standardized evaluation practices, and maps these techniques to underlying computing bottlenecks.

The analysis evaluates algorithmic speedup methods through a literature survey of over 200 publications and targeted empirical microbenchmarks on modern processors and accelerators, including A100 graphics processing units and multicore central processing units. The article establishes a unifying taxonomy around three core training components (function, data, and optimization) and five actions (remove, restrict, reorder, replace, and retrofit), which operate through static or dynamic mechanisms based on domain knowledge or machine learning.

The findings show that widely used theoretical metrics fail to predict real-world performance. Floating-point operations and raw parameter counts do not reliably correlate with wall-clock training time. Operations with low arithmetic intensity, such as normalization layers, are heavily limited by memory bandwidth and take orders of magnitude more time per arithmetic operation than dense matrix multiplications. Factorizing layers can reduce theoretical operation counts while increasing actual runtime due to kernel launch overhead and data movement costs. Furthermore, empirical evaluations reveal that simple baselines, such as reducing the epoch count while properly scaling the learning rate schedule, often outperform complex algorithmic speedup methods. The findings also demonstrate that combining multiple complementary speedup techniques achieves substantially better time-versus-accuracy trade-offs than relying on isolated interventions.

These results demonstrate that algorithmic interventions must be chosen based on the specific hardware bottleneck present in the training pipeline, such as data loading, accelerator memory bandwidth, or interconnect capacity. Pursuing theoretical efficiency reductions without considering hardware execution risks wasting engineering effort and increasing training costs. Organizations can achieve meaningful efficiency gains only when algorithmic techniques directly alleviate active resource bottlenecks, such as using dynamic pruning when compute bound or applying data echoing and larger models when bounded by input data loading.

Decision-makers and practitioners should establish standardized, rigorous evaluation protocols before deploying speedup strategies. Teams should measure actual wall-clock time and generate full Pareto trade-off curves across multiple hyperparameter settings rather than relying on theoretical operation counts. Furthermore, workflows should benchmark against simple baselines, including reduced training durations and smaller model variants. Researchers must also account for the composition of techniques, recognizing that speedup methods interact non-linearly.

A primary limitation of this analysis is that hardware architectures, software frameworks, and model designs continue to evolve rapidly, which can alter specific bottleneck profiles over time. Additionally, many advanced methods that perform well on smaller benchmark datasets struggle to maintain model quality or yield practical speedups on modern large-scale workloads like ImageNet. Decision-makers should maintain moderate confidence in theoretical literature claims until empirical validations are conducted on the specific target hardware, software, and dataset environments.

arXiv: 2210.06640
  • Paper: Green AI, Roy Schwartz et al. (2019). This foundational paper frames the core motivation of Green AI and computational efficiency that the survey formalizes and builds upon.
  • Paper: Energy and Policy Considerations for Deep Learning in NLP, Emma Strubell et al. (2019). This paper establishes the empirical environmental and financial costs of deep learning training runs that the survey seeks to mitigate algorithmically.
  • Paper: The Tradeoffs of Large Scale Learning, Léon Bottou et al. (2007). This seminal work establishes the foundational computational trade-offs between optimization error, sample complexity, and compute time in large-scale machine learning.
  • Paper: Bag of Tricks for Image Classification with Convolutional Neural Networks, Tong He et al. (2018). This paper systematically isolates practical algorithmic training tweaks and acceleration recipes that serve as primary examples in the survey's speedup taxonomy.
  • Paper: Mixed Precision Training, Paulius Micikevicius et al. (2018). This work introduces standard mixed-precision training protocols that form a fundamental baseline for compute-efficient deep learning semantics.
  • Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). This paper introduces gradient checkpointing and activation recomputation, a cornerstone algorithmic memory-reduction technique evaluated across training pipelines.
  • Paper: The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks., Jonathan Frankle et al. (2019). This work formulates the Lottery Ticket Hypothesis, laying the groundwork for algorithmic sparse training and network pruning methods covered in the survey.
  • Paper: Rigging the Lottery: Making All Tickets Winners, Utku Evci et al. (2020). This paper establishes dynamic sparse training techniques (RigL) that enable end-to-end compute savings during the training phase.
  • Paper: Adafactor: Adaptive Learning Rates with Sublinear Memory Cost, Noam Shazeer et al. (2018). This work details memory-efficient adaptive optimization via matrix factorization, addressing a primary training bottleneck discussed in the taxonomy.
  • Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). This survey provides an architectural taxonomy of efficient attention mechanisms that directly informs the broader algorithmic speedup classifications.
Cover for Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities

Abstract

Although deep learning has made great progress in recent years, the exploding economic and environmental costs of training neural networks are becoming unsustainable. To address this problem, there has been a great deal of research on algorithmically-efficient deep learning, which seeks to reduce training costs not at the hardware or implementation level, but through changes in the semantics of the training program. In this paper, we present a structured and comprehensive overview of the research in this field. First, we formalize the algorithmic speedup problem, then we use fundamental building blocks of algorithmically efficient training to develop a taxonomy. Our taxonomy highlights commonalities of seemingly disparate methods and reveals current research gaps. Next, we present evaluation best practices to enable comprehensive, fair, and reliable comparisons of speedup techniques. To further aid research and applications, we discuss common bottlenecks in the training pipeline (illustrated via experiments) and offer taxonomic mitigation strategies for them. Finally, we highlight some unsolved research challenges and present promising future directions.

Table of Contents

  • 1. Introduction
  • 2. Compute-Efficient Training: Overview, Metrics, and Definition
  • 2.1 Overview of DNN Training Process
  • 2.2 Metrics to Quantify Training Efficiency
  • 2.3 Pitfalls of Different Metrics
  • 2.4 Algorithmic Speedup: Motivation and Definition
  • 3. A Unifying Perspective for Speedup Methods: Taxonomy
  • 3.1 Component Categorization
  • 3.1.1 Function Component
  • 3.1.2 Data Component
  • 3.1.3 Optimization Component
  • 3.2 Action Categorization
  • 3.2.1 Remove
  • 3.2.2 Restrict
  • 3.2.3 Reorder
  • 3.2.4 Replace
  • 3.2.5 Retrofit
  • 3.3 Mechanism Categorization
  • 3.3.1 When to Make the Change?
  • 3.3.2 Change Based on What Information?
  • 4. Categorization of Existing Work: Survey
  • 4.1 Function Speedup Strategies
  • 4.1.1 Model Parameters
  • 4.1.2 Architecture
  • 4.2 Data Speedup Strategies
  • 4.2.1 Training Data
  • 4.2.2 Derived Data
  • 4.3 Optimization Speedup Strategies
  • 4.3.1 Training objective
  • 4.3.2 Algorithm
  • 4.4 Other Directions
  • 5. Best Evaluation Practices
  • 6. Guidelines for Achieving Speedup in Practice
  • 6.1 Possible Bottlenecks and Mitigation
  • 6.1.1 GPU compute
  • 6.1.2 GPU memory
  • 6.1.3 GPU memory bandwidth
  • 6.1.4 CPU compute
  • 6.1.5 CPU memory
  • 6.1.6 CPU memory bandwidth
  • 6.1.7 Inter-GPU interconnect
  • 6.1.8 CPU-GPU interconnect
  • 6.1.9 Inter-node interconnect
  • 6.1.10 Storage capacity
  • 6.1.11 Storage bandwidth
  • 6.1.12 Storage network bandwidth
  • 6.2 Overview of Throughputs
  • 7. Summary and Path Forward
  • Acknowledgments
  • Appendix A. Motivation for Compute-Efficient Deep Learning
  • Appendix B. Experimental Details
  • Appendix C. Additional Experiments
  • References

Knowls

  1. Knowl 1 — Formal definition of algorithmic speedup

    definition

    Algorithmic speedup changes the semantics of a neural-network training recipe rather than merely improving its hardware, kernels, or implementation. Let baseline recipe RR achieve model quality QQ in time TT, and let a modified recipe R′R' achieve quality Q′Q' in time T′T'. The paper defines (ε,δ)(\varepsilon,\delta)-speedup by

    ε=TT′>1,δ=Q′−QQ.\varepsilon=\frac{T}{T'}>1,\qquad \delta=\frac{Q'-Q}{Q}.

    Here QQ and Q′Q' are measured by the same model-quality metric, and TT and T′T' are the corresponding training times. A method is strictly better than the baseline when ε>1\varepsilon>1 and δ≥0\delta\geq 0. When faster training requires lower quality, so that ε>1\varepsilon>1 but δ<0\delta<0, speedup methods should be compared on the accuracy–time Pareto frontier rather than ranked by a single scalar.

    Algorithmic speedup can arise by reducing either the number of training iterations needed to reach a quality level, niterationn_{\mathrm{iteration}}, or the time per iteration, TiT_i. The total training time is

    Ttotal=∑i=1niterationTi,T_{\mathrm{total}}=\sum_{i=1}^{n_{\mathrm{iteration}}}T_i,

    where TiT_i is the duration of iteration ii.

  2. Knowl 2 — Three-block taxonomy of algorithmic speedup methods

    model/method

    The paper organizes algorithmic training-speedup methods using three building blocks: the component being changed, the action applied to it, and the mechanism determining when and how the change occurs. The three training-pipeline components are:

    • Function: neural-network architecture and model parameters.
    • Data: training data and derived data such as activations and gradients.
    • Optimization: the training objective and training algorithm, including initialization and the optimizer.

    The five possible actions, called the 5Rs, are:

    • Remove: delete elements or computations, usually reducing per-iteration time TiT_i.
    • Restrict: shrink the space of possible values or representations, such as through quantization, normalization, sparsity, or low-rank structure; this can reduce TiT_i and sometimes niterationn_{\mathrm{iteration}}.
    • Reorder: change when or in what order existing elements are introduced or used, such as with curricula, progressive resizing, layer freezing, or learning-rate schedules; this can reduce either TiT_i or niterationn_{\mathrm{iteration}}.
    • Replace: substitute one component or element with another, potentially reducing both TiT_i and niterationn_{\mathrm{iteration}}.
    • Retrofit: add information, capacity, regularization, or computation to improve quality gained per update; this generally targets lower niterationn_{\mathrm{iteration}} and may increase TiT_i.

    Mechanisms are classified by timing—static, applied once before training, or dynamic, applied during training—and by information source—domain-knowledge-based, using prior expertise, or learning-based, using current or previous training data and runs to select or optimize the change. A method is represented as a path through these blocks. The taxonomy is intentionally not necessarily mutually exclusive or collectively exhaustive, because one method can affect multiple components or be assigned to multiple paths. The authors use it to classify more than 200 papers and to expose unexplored component–action combinations.

  3. Knowl 3 — Surveyed speedup patterns across the 5Rs

    model/method

    The survey finds a systematic relationship between the 5R action and the source of training savings. Function methods include structured pruning, low-rank or mixed-precision parameterizations, progressive freezing, Fourier- or search-designed architectural replacements, and conditional parameterization. Reported examples include PipeTransformer, which achieved up to 2.83×2.83\times speedup without accuracy loss by freezing layers according to gradient norms; automated progressive learning for vision transformers, which reported 40–85% faster ImageNet training; and Switch Transformers, which reported up to seven-times faster training than T5 despite having up to a trillion parameters, because only input-selected parameter subsets are computed.

    Data methods remove or select training samples, omit low-priority gradients or activations, reorder samples through curricula, progressively change image resolution, replace data with compressed or distilled representations, or add augmentations and pretraining. Selective-Backprop reported a 3.5×3.5\times speedup over standard SGD on CIFAR-10, CIFAR-100, and SVHN with an accuracy decrease. A surveyed ImageNet data-pruning method trained with 20% fewer samples and little accuracy loss, whereas curriculum-learning studies generally found small gains that became more useful under limited iteration budgets or noisy labels.

    Optimization methods remove or approximate expensive objective terms, restrict initialization or update geometry, reorder losses or regularization over training, replace SGD with adaptive or learned optimizers, and retrofit objectives or updates with sharpness or curvature information. Examples include hierarchical softmax and negative sampling for large output spaces, efficient variants of sharpness-aware minimization, large-batch and orthogonal-initialization methods, AdamW, learned optimizers, and second-order preconditioning. Zeroth-order methods have reported up to 300-times speedup on small networks and datasets, but the surveyed results indicate substantial scaling difficulties. Across categories, learning-based and dynamic mechanisms can improve the quality–speed tradeoff but often incur overhead; static and domain-knowledge-based mechanisms are cheaper but may be less adaptable.

  4. Knowl 4 — Efficiency metrics can disagree with actual training time

    empirical result

    The paper compares training time, FLOP count, parameter count, electricity, carbon emissions, and operand sizes as efficiency measures. Training time is the most direct measure of practical utility but depends on hardware and software. FLOPs and parameter count are hardware-independent, yet omit memory traffic, communication, kernel-launch overhead, utilization, instruction support, and input resolution or sequence length. Electricity and carbon measurements are affected by hardware, location, time, and electricity-generation conditions. The total size of operator inputs and outputs is a hardware-independent proxy for data movement and is often more informative than FLOPs, although it remains imperfect.

    The paper's microbenchmarks show that these metrics can produce contradictory rankings. On a 40 GB A100, FLOPs and runtime are approximately linearly related for square matrix products of size around 4096 or larger, but for smaller products runtime is nearly independent of FLOPs and is not even monotonic because of launch overhead, memory bandwidth, and implementation details. Operations with similar FLOP counts can differ in runtime by an order of magnitude: normalization operations are expensive because they move much more data per FLOP. Operand bytes reduce much of this cross-operation discrepancy.

    For ResNet-50 forward passes, factorizing convolution and linear layers substantially reduces FLOPs and parameter count but increases runtime because of extra memory traffic and kernel launches. Removing batch-normalization layers substantially reduces runtime while barely changing FLOP count or parameter count. Measurements on a 40 GB A100, 64 CPU cores, ConvNeXt-B, and Swin-B show the same qualitative pattern: FLOPs are a poor wall-time proxy, while operand size is a better but still imperfect proxy. The paper therefore recommends reporting multiple metrics and, whenever possible, accuracy–efficiency curves rather than treating FLOPs or parameter count as substitutes for timing.

  5. Knowl 5 — Best practices for evaluating speedup claims

    experimental setup

    The paper proposes evaluation practices intended to make algorithmic-speedup comparisons fair, comprehensive, and reliable. First, the claimed scope must be explicit—for example, whether a method targets natural-image classification generally or only a particular architecture family. The complete architecture, preprocessing, hardware, software stack, optimizer, learning-rate schedule, weight decay, and all other hyperparameters should be reported. Hyperparameters should be held fixed between baseline and intervention when attributing gains to the proposed action; robustness claims require explicit sweeps and variation estimates.

    Experiments should control hardware and software confounders, since libraries, data loaders, accelerator types, and operation support can change runtime and even accuracy. Speedup is a multi-objective problem, so authors should vary the relevant speedup strength, such as sparsity or data-pruning rate, and report an accuracy–time Pareto curve with at least three operating points. The paper recommends using at least three dataset–architecture pairs, including modern large-scale cases, and prioritizing high-accuracy operating points.

    Comparisons should include simple strong baselines: a shorter training run that completes a correspondingly shortened learning-rate schedule, a smaller model, and methods resembling the proposed technique. Multiple random seeds, clearly defined error bars, a central tendency and variation measure, and the provenance of every comparison—new implementation, reused code, or reported result—should be provided. These practices are necessary because an apparently superior method can instead be benefiting from better hyperparameters, a stronger baseline recipe, or an unreported software or hardware difference.

  6. Knowl 6 — Bottleneck-aware selection of speedup interventions

    model/method

    The practitioner guide argues that an algorithmic intervention produces a real speedup only when it addresses the current bottleneck in the training pipeline. Recommended mitigations include:

    • GPU compute: reduce feature or channel counts, resolution or sequence length, or use sufficiently large block sparsity, supported 2:4 sparsity, or low-rank and structured matrices.
    • GPU memory capacity: shard or CPU-offload model and optimizer state, use an optimizer with less state, reduce model size, checkpoint or recompute activations, shard or compress activations, and use memory-efficient activation functions or in-place operations.
    • GPU memory bandwidth: reduce parameter, activation, and optimizer-state tensor sizes, or increase arithmetic intensity by performing more computation per byte moved.
    • CPU compute and data loading: simplify decoding and augmentation, use optimized loaders such as DALI or FFCV, offload preprocessing to accelerators, or hide loading time by doing more accelerator computation per batch.
    • CPU memory, storage, and storage bandwidth: use fast local NVMe storage, avoid unnecessary model copies, offload fewer tensors, stream or cache data, distribute reads across local devices, use RAID where appropriate, and reduce dataset or checkpoint size.
    • GPU–GPU, CPU–GPU, and inter-node communication: compress gradients, synchronize less frequently when the accuracy tradeoff permits, increase computation per communicated parameter through larger batches or inputs, use suitable parallelism, shard selectively, and avoid unnecessary device transfers.

    The guide illustrates why bottleneck diagnosis matters with small image models: a ResNet-18 can appear scarcely faster than larger ResNets when the data loader limits throughput, even though it would be faster with a sufficiently rapid loader. Speeding up model computation in that regime does not reduce end-to-end training time.

  7. Knowl 7 — Composing methods and using strong baselines changes the apparent speedup

    empirical result

    In the paper's ResNet-50/ImageNet experiments, training for fewer epochs while traversing the full shortened learning-rate schedule outperformed nearly all of the surveyed speedup methods tested in a common implementation environment. Simply aborting a long scheduled run early was much less effective. Training smaller ResNet-18 or ResNet-34 models also outperformed some proposed methods at particular hyperparameter settings.

    The experiments further show that composing several speedup methods produces a substantially better time–accuracy curve than selecting the best operating points of individual methods separately. No composition is uniformly best: some combinations help at short training budgets while others help at long budgets, and regularization can hurt short-budget accuracy while improving long-budget accuracy. These findings mean that a new method must be compared against a baseline containing the same accompanying methods; otherwise, gains may be incorrectly attributed to the proposed intervention. The authors caution that some comparison hyperparameters were selected on a best-effort basis and may not be optimal, so the experiments establish that strong baselines can win under tested conditions, not that they always will.

  8. Knowl 8 — Representative throughput hierarchy of training resources

    data/table

    The paper reports representative peak throughputs to show that training resources differ by orders of magnitude. The data communicate why dense accelerator arithmetic is not automatically the bottleneck: a workload that moves data through a slow interconnect or storage device can be limited there even when accelerator compute is underused. Values are hardware-specific peak rates, not universal measured rates.

    Could not parse LaTeX table

    Using these representative values, one A100 can perform approximately 156/(0.00125/2)=249,600156/(0.00125/2)=249{,}600 multiply-adds for every 16-bit value transmitted over a 10 Gbps network. In an eight-GPU server, aggregate accelerator throughput can exceed one quadrillion operations per second while machine-level inter-node bandwidth does not scale with the number of GPUs, making communication or storage bottlenecks increasingly plausible.

  9. Knowl 9 — Scope limitations and open research directions

    limitation

    The survey is limited to algorithmic changes to training, rather than inference, efficient hardware, or low-level compute kernels, and it emphasizes methods applicable to language and computer-vision workloads rather than reinforcement learning, graph learning, or other specialized domains. The authors identify several gaps in the surveyed literature: speedup methods are disproportionately designed and tested for computer vision; many component–action pairs remain lightly explored; most methods target only one or two training subcomponents; and methods such as unstructured sparsity may reduce theoretical work without being hardware-efficient.

    The paper therefore proposes transferring ideas to large-language-model training, generative modeling, reinforcement learning, adversarial training, and other settings; developing methods that jointly modify several components; creating multi-objective ranking metrics and a comprehensive, fair, reliable benchmarking leaderboard; releasing pretrained models under FAIR principles; and pursuing algorithm–hardware co-design. The taxonomy is a reasoning and organization tool rather than a claim that every method has a unique classification, and the survey's conclusions about individual techniques inherit the heterogeneous evidence and evaluation quality of the underlying literature.

Coverage note — The paper's full bibliography-level classification of more than 200 papers and detailed per-paper entries are not reproduced individually because the top ten knowls prioritize the formal framework, taxonomy, empirical evidence, evaluation guidance, practical bottleneck analysis, and stated limitations.

References

  1. 1.Alessandro Achille, Matteo Rovere, and Stefano Soatto. Critical learning periods in deep neural networks. arXiv preprint arXiv:1711.08856, 2017.
  2. 2.Menachem Adelman, Kfir Levy, Ido Hakimi, and Mark Silberstein. Faster neural network training with approximate tensor operations. Advances in Neural Information Processing Systems, 34:27877–27889, 2021.
  3. 3.Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman, and Dimitris Papailiopoulos. On the utility of gradient compression in distributed training systems. Proceedings of Machine Learning and Systems, 4:652–672, 2022.
  4. 4.Hafez Ahmad. Machine learning applications in oceanography. Aquatic Research, 2(3):161–169, 2019.
  5. 5.Nur Ahmed and Muntasir Wahed. The de-democratization of ai: Deep learning and the compute divide in artificial intelligence research. arXiv preprint arXiv:2010.15581, 2020.
  6. 6.Shakkeel Ahmed, Ravi S Mula, and Soma S Dhavala. A framework for democratizing ai. arXiv preprint arXiv:2001.00818, 2020.
  7. 7.Ali Akbari, Muhammad Awais, Manijeh Bashar, and Josef Kittler. How does loss function affect generalization performance of deep learning? application to human age estimation. In International Conference on Machine Learning, pages 141–151. PMLR, 2021.
  8. 8.Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. Scalable second order optimization for deep learning. arXiv preprint arXiv:2002.09018, 2020a.
  9. 9.Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. Second order optimization made practical. arXiv preprint arXiv:2002.09018, 2020b.
  10. 10.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  11. 11.Sungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim, and Kyoung Mu Lee. Meta-learning with adaptive hyperparameters. Advances in Neural Information Processing Systems, 33:20755–20765, 2020.
  12. 12.Atılım G¨une¸s Baydin, Barak A Pearlmutter, Don Syme, Frank Wood, and Philip Torr. Gradients without backpropagation. arXiv preprint arXiv:2202.08587, 2022.
  13. 13.Sarah Bechtle, Artem Molchanov, Yevgen Chebotar, Edward Grefenstette, Ludovic Righetti, Gaurav Sukhatme, and Franziska Meier. Meta learning via learned loss. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 4161–4168. IEEE, 2021.
  14. 14.Irwan Bello. Lambdanetworks: Modeling long-range interactions without attention. arXiv preprint arXiv:2102.08602, 2021.
  15. 15.Irwan Bello, William Fedus, Xianzhi Du, Ekin D Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, and Barret Zoph. Revisiting resnets: Improved training and scaling strategies. arXiv preprint arXiv:2103.07579, 2021.
  16. 16.Tal Ben-Nun and Torsten Hoefler. Demystifying parallel and distributed deep learning: An in-depth concurrency analysis. ACM Computing Surveys (CSUR), 52(4):1–43, 2019.
  17. 17.Yoshua Bengio, J´erôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48, 2009.
  18. 18.Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. signsgd: Compressed optimisation for non-convex problems. In International Conference on Machine Learning, pages 560–569. PMLR, 2018.
  19. 19.Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, and Maya Gokhale. Benchmarking test-time unsupervised deep neural network adaptation on edge devices. arXiv preprint arXiv:2203.11295, 2022a.
  20. 20.Kshitij Bhardwaj, James Diffenderfer, Bhavya Kailkhura, and Maya Gokhale. Unsupervised test-time adaptation of deep neural networks at the edge: a case study. In 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE), pages 412–417. IEEE, 2022b.
  21. 21.Vighnesh Birodkar, Hossein Mobahi, and Samy Bengio. Semantic redundancies in image-classification datasets: The 10% you don’t need. arXiv preprint arXiv:1901.11409, 2019.
  22. 22.Davis Blalock and John Guttag. Multiplying matrices without multiplying. arXiv preprint arXiv:2106.10860, 2021.
  23. 23.Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. What is the state of neural network pruning? Proceedings of machine learning and systems, 2:129–146, 2020.
  24. 24.Davis W. Blalock, Michael Carbin, Laura Florescu, Jonathan Frankle, Matthew L. Leavitt, Tyler Lee, Moin Nadeem, Jacob Portes, Naveen Rao, Landan Seguin, Cory Stephenson, Hanlin Tang, and Abhinav Venigalla. On evaluating and improving the efficiency of deep networks. Technical report, MosaicML, October 2021. URL https://www.mosaicml.com/blog/methodology.
  25. 25.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  26. 26.Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston. Freezeout: Accelerate training by progressively freezing layers. arXiv preprint arXiv:1706.04983, 2017.
  27. 27.Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. arXiv preprint arXiv:2102.06171, 2021.
  28. 28.Jason Ross Brown, Yiren Zhao, Ilia Shumailov, and Robert D Mullins. Wide attention is the way forward for transformers. arXiv preprint arXiv:2210.00640, 2022.
  29. 29.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020a.
  30. 30.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020b.
  31. 31.Samuel Rota Bulo, Lorenzo Porzi, and Peter Kontschieder. In-place activated batchnorm for memory-optimized training of dnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5639–5647, 2018.
  32. 32.Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re. Pixelated butterfly: Simple and efficient sparse training for neural network models. arXiv preprint arXiv:2112.00029, 2021a.
  33. 33.Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Ré. Scatterbrain: Unifying sparse and low-rank attention approximation. arXiv preprint arXiv:2110.15343, 2021b.
  34. 34.Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael Mahoney, and Joseph Gonzalez. Actnn: Reducing training memory footprint via 2-bit activation compressed training. In International Conference on Machine Learning, pages 1803–1813. PMLR, 2021c.
  35. 35.Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang, and Zhangyang Wang. Coarsening the granularity: Towards structurally sparse lottery tickets. arXiv preprint arXiv:2202.04736, 2022.
  36. 36.Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. Net2net: Accelerating learning via knowledge transfer. arXiv preprint arXiv:1511.05641, 2015a.
  37. 37.Welin Chen, David Grangier, and Michael Auli. Strategies for training large vocabulary neural language models. arXiv preprint arXiv:1512.04906, 2015b.
  38. 38.Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu. Earlybert: Efficient bert training via early-bird lottery tickets. arXiv preprint arXiv:2101.00063, 2020.
  39. 39.Xuxi Chen, Tianlong Chen, Yu Cheng, Weizhu Chen, Zhangyang Wang, and Ahmed Hassan Awadallah. Dsee: Dually sparsity-embedded efficient tuning of pre-trained language models. arXiv preprint arXiv:2111.00160, 2021d.
  40. 40.Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3435–3444, 2019.
  41. 41.Kashyap Chitta, José M Alvarez, Elmar Haussmann, and Clément Farabet. Training data subset search with ensemble active learning. IEEE Transactions on Intelligent Transportation Systems, 2021.
  42. 42.Minsik Cho, Vinod Muthusamy, Brad Nemanich, and Ruchir Puri. Gradzip: Gradient compression using alternating matrix factorization for large-scale deep learning. In NeurIPS. 2019.
  43. 43.Dami Choi, Alexandre Passos, Christopher J Shallue, and George E Dahl. Faster neural network training with data echoing. arXiv preprint arXiv:1907.05550, 2019.
  44. 44.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  45. 45.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  46. 46.Sankalan Pal Chowdhury, Adamos Solomou, Avinava Dubey, and Mrinmaya Sachan. On learning the transformer kernel. arXiv preprint arXiv:2110.08323, 2021.
  47. 47.Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia. Dawnbench: An end-to-end deep learning benchmark and competition. Training, 100(101):102, 2017.
  48. 48.Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia. Selection via proxy: Efficient data selection for deep learning. arXiv preprint arXiv:1906.11829, 2019.
  49. 49.Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le. Autoaugment: Learning augmentation policies from data. arXiv preprint arXiv:1805.09501, 2018.
  50. 50.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  51. 51.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803, 2021.
  52. 52.Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher Ré. Monarch: Expressive structured matrices for efficient and accurate training. In International Conference on Machine Learning, pages 4690–4721. PMLR, 2022a.
  53. 53.Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. arXiv preprint arXiv:2205.14135, 2022b.
  54. 54.Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay. The efficiency misnomer. arXiv preprint arXiv:2110.12894, 2021.
  55. 55.Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al. Natural neural networks. Advances in neural information processing systems, 28, 2015.
  56. 56.Tim Dettmers. 8-bit approximations for parallelism in deep learning. arXiv preprint arXiv:1511.04561, 2015.
  57. 57.Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization. arXiv preprint arXiv:2110.02861, 2021.
  58. 58.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  59. 59.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
  60. 60.Sourya Dey, Kuan-Wen Huang, Peter A Beerel, and Keith M Chugg. Pre-defined sparse neural networks with hardware acceleration. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 9(2):332–345, 2019.
  61. 61.James Diffenderfer, Brian Bartoldson, Shreya Chaganti, Jize Zhang, and Bhavya Kailkhura. A winning hand: Compressing deep networks can improve out-of-distribution robustness. Advances in Neural Information Processing Systems, 34:664–676, 2021.
  62. 62.Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. Repvgg: Making vgg-style convnets great again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13733–13742, 2021.
  63. 63.Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A Smith, Nicole DeCario, and Will Buchanan. Measuring the carbon intensity of ai in cloud instances. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1877–1894, 2022.
  64. 64.Akshunna S Dogra and William T Redman. Optimizing neural networks via koopman operator theory. arXiv preprint arXiv:2006.02361, 2020.
  65. 65.Piotr Dollár, Mannat Singh, and Ross Girshick. Fast and accurate model scaling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 924–932, 2021.
  66. 66.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  67. 67.Nikoli Dryden, Roman Böringer, Tal Ben-Nun, and Torsten Hoefler. Clairvoyant prefetching for distributed machine learning i/o. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15, 2021.
  68. 68.Jiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, and Vincent YF Tan. Efficient sharpness-aware minimization for improved training of neural networks. arXiv preprint arXiv:2110.03141, 2021.
  69. 69.Jiawei Du, Daquan Zhou, Jiashi Feng, Vincent YF Tan, and Joey Tianyi Zhou. Sharpness-aware training for free. arXiv preprint arXiv:2205.14083, 2022.
  70. 70.Yann Dubois, Benjamin Bloem-Reddy, Karen Ullrich, and Chris J Maddison. Lossy compression for lossless prediction. arXiv preprint arXiv:2106.10800, 2021.
  71. 71.John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, 12(7), 2011.
  72. 72.Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar. Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd. In International conference on artificial intelligence and statistics, pages 803–812. PMLR, 2018.
  73. 73.Adam Dziedzic, John Paparrizos, Sanjay Krishnan, Aaron Elmore, and Michael Franklin. Band-limited training and inference for convolutional neural networks. In International Conference on Machine Learning, pages 1745–1754. PMLR, 2019.
  74. 74.R David Evans and Tor Aamodt. Ac-gc: Lossy activation compression with guaranteed convergence. Advances in Neural Information Processing Systems, 34:27434–27448, 2021.
  75. 75.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019.
  76. 76.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961, 2021.
  77. 77.Steven Y Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. A survey of data augmentation approaches for nlp. arXiv preprint arXiv:2105.03075, 2021.
  78. 78.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, pages 1126–1135. PMLR, 2017.
  79. 79.Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. arXiv preprint arXiv:2010.01412, 2020.
  80. 80.Stanislav Fort, Andrew Brock, Razvan Pascanu, Soham De, and Samuel L Smith. Drawing multiple augmentation samples per image during training efficiently decreases test error. arXiv preprint arXiv:2105.13343, 2021.
  81. 81.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin. Pruning neural networks at initialization: Why are we missing the mark? arXiv preprint arXiv:2009.08576, 2020a.
  82. 82.Jonathan Frankle, David J Schwab, and Ari S Morcos. The early phase of neural network training. arXiv preprint arXiv:2002.10365, 2020b.
  83. 83.Dan Fu and Gabriel Guimaraes. Using compression to speed up image classification in artificial neural networks. Technical report, 2016.
  84. 84.Jonas Geiping and Tom Goldstein. Cramming: Training a language model on a single gpu in one day. arXiv preprint arXiv:2212.14034, 2022.
  85. 85.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  86. 86.Donald Goldfarb, Yi Ren, and Achraf Bahamou. Practical quasi-newton methods for training deep neural networks. Advances in Neural Information Processing Systems, 33:2386–2396, 2020.
  87. 87.Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tieyan Liu. Efficient training of bert by progressively stacking. In International conference on machine learning, pages 2337–2346. PMLR, 2019.
  88. 88.Raphael Gontijo-Lopes, Yann Dauphin, and Ekin D Cubuk. No one representation to rule them all: Overlapping features of training methods. arXiv preprint arXiv:2110.12899, 2021.
  89. 89.Santiago Gonzalez and Risto Miikkulainen. Improved training speed, accuracy, and data utilization through loss function optimization. In 2020 IEEE Congress on Evolutionary Computation (CEC), pages 1–8. IEEE, 2020.
  90. 90.Baptiste Goujaud, Damien Scieur, Aymeric Dieuleveut, Adrien Taylor, and Fabian Pedregosa. Super-acceleration with cyclical step-sizes. arXiv preprint arXiv:2106.09687, 2021.
  91. 91.Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyröla, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
  92. 92.Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han. On the transformer growth for progressive bert training. arXiv preprint arXiv:2010.12562, 2020.
  93. 93.Lionel Gueguen, Alex Sergeev, Ben Kadlec, Rosanne Liu, and Jason Yosinski. Faster neural networks straight from jpeg. Advances in Neural Information Processing Systems, 31:3933–3944, 2018.
  94. 94.Udit Gupta, Young Geun Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin S Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. Chasing carbon: The elusive environmental footprint of computing. IEEE Micro, 42(4):37–47, 2022.
  95. 95.Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization. In International Conference on Machine Learning, pages 1842–1850. PMLR, 2018.
  96. 96.Guy Hacohen and Daphna Weinshall. On the power of curriculum learning in training deep networks. In International Conference on Machine Learning, pages 2535–2544. PMLR, 2019.
  97. 97.Stewart Hall, Rob Schreiber, and Sean Lie. Training giant neural networks using weight streaming on cerebras wafer-scale systems. Technical report, Cerebras Systems,Inc., November 2021. URL https://f.hubspotusercontent30.net/hubfs/8968533/VirtualBoothDocs/CSWeightStreamingWhitePaper111521.pdf.
  98. 98.Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Yuan Yao, Ao Zhang, Liang Zhang, et al. Pre-trained models: Past, present and future. AI Open, 2:225–250, 2021.
  99. 99.Yuna Han and Byung-Woo Hong. Deep learning based on fourier convolutional neural network incorporating random kernels. Electronics, 10(16):2004, 2021.
  100. 100.Boris Hanin and David Rolnick. How to start training: The effect of initialization and architecture. Advances in Neural Information Processing Systems, 31, 2018.
  101. 101.Stephen Hanson and Lorien Pratt. Comparing biases for minimal network construction with back-propagation. Advances in neural information processing systems, 1, 1988.
  102. 102.Chaoyang He, Shen Li, Mahdi Soltanolkotabi, and Salman Avestimehr. Pipetransformer: Automated elastic pipelining for distributed training of large-scale models. In International Conference on Machine Learning, pages 4150–4159. PMLR, 2021.
  103. 103.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015.
  104. 104.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  105. 105.Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau. Towards the systematic reporting of the energy and carbon footprints of machine learning. Journal of Machine Learning Research, 21(248):1–43, 2020.
  106. 106.Danny Hernandez and Tom B Brown. Measuring the algorithmic efficiency of neural networks. arXiv preprint arXiv:2005.04305, 2020.
  107. 107.Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. arXiv preprint arXiv:2102.00554, 2021.
  108. 108.Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry. Augment your batch: better training with larger batches. arXiv preprint arXiv:1901.09335, 2019a.
  109. 109.Elad Hoffer, Berry Weinstein, Itay Hubara, Tal Ben-Nun, Torsten Hoefler, and Daniel Soudry. Mix & match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency. arXiv preprint arXiv:1908.08986, 2019b.
  110. 110.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022a.
  111. 111.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022b.
  112. 112.Jeremy Howard. Fastai - progressive resizing. Technical report, 2018.
  113. 113.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  114. 114.Wei Hu, Lechao Xiao, and Jeffrey Pennington. Provable benefit of orthogonal initialization in optimizing deep linear networks. arXiv preprint arXiv:2001.05992, 2020.
  115. 115.Furong Huang, Jordan Ash, John Langford, and Robert Schapire. Learning deep resnet blocks sequentially using boosting theory. In International Conference on Machine Learning, pages 2058–2067. PMLR, 2018.
  116. 116.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pages 646–661. Springer, 2016.
  117. 117.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  118. 118.Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao. Normalization techniques in training dnns: Methodology, analysis and application. arXiv preprint arXiv:2009.12836, 2020.
  119. 119.Nathan Hubens, Matei Mancas, Bernard Gosselin, Marius Preda, and Titus Zaharia. One-cycle pruning: Pruning convnets under a tight training budget. arXiv preprint arXiv:2107.02086, 2021.
  120. 120.Yerlan Idelbayev and Miguel A Carreira-Perpinán. Low-rank compression of neural nets: Learning the rank of each layer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8049–8059, 2020.
  121. 121.Yani Ioannou, Duncan Robertson, Jamie Shotton, Roberto Cipolla, and Antonio Criminisi. Training cnns with low-rank filters for efficient image classification. arXiv preprint arXiv:1511.06744, 2015.
  122. 122.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015.
  123. 123.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407, 2018.
  124. 124.et.al. Jaime Sevilla. Compute trends across three eras of machine learning. arXiv preprint arXiv:2202.05924, 2022.
  125. 125.Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Joseph Gonzalez, Kurt Keutzer, and Ion Stoica. Checkmate: Breaking the memory wall with optimal tensor rematerialization. Proceedings of Machine Learning and Systems, 2:497–511, 2020.
  126. 126.Yunho Jeon and Junmo Kim. Constructing fast network through deconstruction of convolution. arXiv preprint arXiv:1806.07370, 2018.
  127. 127.Angela H Jiang, Daniel L-K Wong, Giulio Zhou, David G Andersen, Jeffrey Dean, Gregory R Ganger, Gauri Joshi, Michael Kaminksy, Michael Kozuch, Zachary C Lipton, et al. Accelerating deep learning by focusing on the biggest losers. arXiv preprint arXiv:1910.00762, 2019.
  128. 128.Pengzhan Jin, Lu Lu, Yifa Tang, and George Em Karniadakis. Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness. Neural Networks, 130:85–99, 2020.
  129. 129.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  130. 130.Jean Kaddour. Stop wasting my time! saving days of imagenet and bert training with latest weight averaging. arXiv preprint arXiv:2209.14981, 2022.
  131. 131.Jean Kaddour, Linqing Liu, Ricardo Silva, and Matt J Kusner. Questions for flat-minima optimization of modern neural networks. arXiv preprint arXiv:2202.00661, 2022.
  132. 132.Bhavya Kailkhura, Jayaraman Thiagarajan, Qunwei Li, Jize Zhang, Yi Zhou, and Timo Bremer. A statistical mechanics framework for task-agnostic sample design in machine learning. Advances in Neural Information Processing Systems, 33:11925–11935, 2020.
  133. 133.Athresh Karanam, Krishnateja Killamsetty, Harsha Kokel, and Rishabh K Iyer. Orient: Submodular mutual information measures for data subset selection under distribution shift. In Advances in Neural Information Processing Systems, 2022.
  134. 134.Angelos Katharopoulos and François Fleuret. Not all samples are created equal: Deep learning with importance sampling. In International conference on machine learning, pages 2525–2534. PMLR, 2018.
  135. 135.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020.
  136. 136.Vishal Kaushal, Rishabh Iyer, Suraj Kothawade, Rohan Mahadev, Khoshrav Doctor, and Ganesh Ramakrishnan. Learning from less data: A unified data subset selection and active learning framework for computer vision. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1289–1299. IEEE, 2019.
  137. 137.Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836, 2016.
  138. 138.Krishnateja Killamsetty, Durga Sivasubramanian, Baharan Mirzasoleiman, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer. Grad-match: A gradient matching based data subset selection for efficient learning. arXiv preprint arXiv:2103.00123, 2021a.
  139. 139.Krishnateja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, and Rishabh Iyer. Glister: Generalization based data subset selection for efficient and robust learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8110–8118, 2021b.
  140. 140.Krishnateja Killamsetty, Xujiang Zhao, Feng Chen, and Rishabh Iyer. Retrieve: Coreset selection for efficient and robust semi-supervised learning. Advances in Neural Information Processing Systems, 34:14488–14501, 2021c.
  141. 141.Krishnateja Killamsetty, Guttu Sai Abhishek, Alexandre V Evfimievski, Lucian Popa, Ganesh Ramakrishnan, Rishabh Iyer, et al. Automata: Gradient based data subset selection for compute-efficient hyper-parameter tuning. arXiv preprint arXiv:2203.08212, 2022.
  142. 142.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  143. 143.Katrin Kirchhoff and Jeff Bilmes. Submodularity for data selection in machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 131–141, 2014.
  144. 144.Boris Knyazev, Michal Drozdzal, Graham W Taylor, and Adriana Romero-Soriano. Parameter prediction for unseen deep architectures. arXiv preprint arXiv:2110.13100, 2021.
  145. 145.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  146. 146.Jan Kukaˇcka, Vladimir Golkov, and Daniel Cremers. Regularization for deep learning: A taxonomy. arXiv preprint arXiv:1710.10686, 2017.
  147. 147.Aditya Kusupati, Matthew Wallingford, Vivek Ramanujan, Raghav Somani, Jae Sung Park, Krishna Pillutla, Prateek Jain, Sham Kakade, and Ali Farhadi. Llc: Accurate, multi-purpose learnt low-dimensional binary codes. arXiv preprint arXiv:2106.01487, 2021.
  148. 148.Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019.
  149. 149.Nam Le, Honglei Zhang, Francesco Cricri, Ramin Ghaznavi-Youvalari, and Esa Rahtu. Image coding for machines: An end-to-end learned approach. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1590–1594. IEEE, 2021.
  150. 150.Matthew Leavitt. Blazingly fast computer vision training with the mosaic resnet and composer. Technical report, 2022.
  151. 151.Guillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park, Hadi Salman, and Aleksander Madry. ffcv. https://github.com/libffcv/ffcv/, 2022. commit f25386557e213711cc8601833add36ff966b80b2.
  152. 152.James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. Fnet: Mixing tokens with fourier transforms. arXiv preprint arXiv:2105.03824, 2021.
  153. 153.Changlin Li, Bohan Zhuang, Guangrun Wang, Xiaodan Liang, Xiaojun Chang, and Yi Yang. Automated progressive learning for efficient training of vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12486–12496, 2022a.
  154. 154.Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer. Branch-train-merge: Embarrassingly parallel training of expert language models. arXiv preprint arXiv:2208.03306, 2022b.
  155. 155.Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E Gonzalez. Train large, then compress: Rethinking model size for efficient training and inference of transformers. arXiv preprint arXiv:2002.11794, 2020.
  156. 156.Xiangru Lian, Binhang Yuan, Xuefeng Zhu, Yulong Wang, Yongjun He, Honghuan Wu, Lei Sun, Haodong Lyu, Chengjun Liu, Xing Dong, et al. Persia: An open, hybrid system scaling deep learning-based recommenders up to 100 trillion parameters. ArXiv, vol. abs/2111.05897, 2021.
  157. 157.Edgar Liberis, Lukasz Dudziak, and Nicholas D Lane. µnas: Constrained neural architecture search for microcontrollers. In Proceedings of the 1st Workshop on Machine Learning and Systems, pages 70–79, 2021.
  158. 158.Lucas Liebenwein, Alaa Maalouf, Dan Feldman, and Daniela Rus. Compressing neural networks: Towards determining the optimal layer-wise decomposition. Advances in Neural Information Processing Systems, 34:5328–5344, 2021.
  159. 159.Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi. Don’t use large mini-batches, use local sgd. arXiv preprint arXiv:1808.07217, 2018.
  160. 160.Liu Liu, Lei Deng, Xing Hu, Maohua Zhu, Guoqi Li, Yufei Ding, and Yuan Xie. Dynamic sparse graph for efficient deep learning. arXiv preprint arXiv:1810.00859, 2018a.
  161. 161.Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198, 2018b.
  162. 162.Shiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen, Ghada Sokar, Elena Mocanu, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu. Deep ensembling with no overhead for either training or testing: The all-round blessings of dynamic sparsity. arXiv preprint arXiv:2106.14568, 2021a.
  163. 163.Sijia Liu, Pin-Yu Chen, Bhavya Kailkhura, Gaoyuan Zhang, Alfred O Hero III, and Pramod K Varshney. A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications. IEEE Signal Processing Magazine, 37(5):43–54, 2020.
  164. 164.Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, and Yang You. Towards efficient and scalable sharpness-aware minimization. arXiv preprint arXiv:2203.02714, 2022a.
  165. 165.Yue Liu, Christos Matsoukas, Fredrik Strand, Hossein Azizpour, and Kevin Smith. Patch-dropout: Economizing vision transformers using patch dropout. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3953–3962, 2023.
  166. 166.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021b.
  167. 167.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022b.
  168. 168.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE international conference on computer vision, pages 2736–2744, 2017.
  169. 169.Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. Rethinking the value of network pruning. arXiv preprint arXiv:1810.05270, 2018c.
  170. 170.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022c.
  171. 171.Andrew J. Lohn and Micah Musser. Ai and compute, how much longer can computing power drive artificial intelligence progress? 2022.
  172. 172.Ilya Loshchilov and Frank Hutter. Online batch selection for faster training of neural networks. arXiv preprint arXiv:1511.06343, 2015.
  173. 173.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  174. 174.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  175. 175.Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, and Yuxiong He. Maximizing communication efficiency for large-scale training via 0/1 adam. arXiv preprint arXiv:2202.06009, 2022.
  176. 176.Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. Estimating the carbon footprint of bloom, a 176b parameter language model. arXiv preprint arXiv:2211.02001, 2022.
  177. 177.Sangkug Lym, Esha Choukse, Siavash Zangeneh, Wei Wen, Sujay Sanghavi, and Mattan Erez. Prunetrain: fast neural network training by dynamic sparse model reconfiguration. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–13, 2019.
  178. 178.James Martens and Roger Grosse. Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, pages 2408–2417. PMLR, 2015.
  179. 179.James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S Schoenholz. Rapid training of deep neural networks without skip connections or normalization layers using deep kernel shaping. arXiv preprint arXiv:2110.01765, 2021.
  180. 180.Dhruv Matani and Suraj Subramanian. Efficient pytorch: Tensor memory format matters. https://pytorch.org/blog/tensor-memory-format-matters/, 2021.
  181. 181.Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training. arXiv preprint arXiv:1812.06162, 2018.
  182. 182.Gaurav Menghani. Efficient deep learning: A survey on making deep learning models smaller, faster, and better. arXiv preprint arXiv:2106.08962, 2021.
  183. 183.Luke Metz, Niru Maheswaranathan, Jeremy Nixon, Daniel Freeman, and Jascha Sohl-Dickstein. Understanding and correcting pathologies in the training of learned optimizers. In International Conference on Machine Learning, pages 4556–4565. PMLR, 2019.
  184. 184.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017.
  185. 185.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 26, 2013.
  186. 186.S¨oren Mindermann, Muhammed Razzak, Winnie Xu, Andreas Kirsch, Mrinank Sharma, Adrien Morisot, Aidan N Gomez, Sebastian Farquhar, Jan Brauner, and Yarin Gal. Prioritized training on points that are learnable, worth learning, and not yet learned. arXiv preprint arXiv:2206.07137, 2021.
  187. 187.Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning, pages 6950–6960. PMLR, 2020.
  188. 188.Ashish Mittal, Durga Sivasubramanian, Rishabh Iyer, Preethi Jyothi, and Ganesh Ramakrishnan. Partitioned gradient matching-based data subset selection for compute-efficient robust asr training. arXiv preprint arXiv:2210.16892, 2022.
  189. 189.Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature communications, 9(1):1–12, 2018.
  190. 190.Marcin Moczulski, Misha Denil, Jeremy Appleyard, and Nando de Freitas. Acdc: A structured efficient linear layer. arXiv preprint arXiv:1511.05946, 2015.
  191. 191.Gordon E Moore et al. Cramming more components onto integrated circuits, 1965.
  192. 192.Frederic Morin and Yoshua Bengio. Hierarchical probabilistic neural network language model. In International workshop on artificial intelligence and statistics, pages 246–252. PMLR, 2005.
  193. 193.MG Sarwar Murshed, Christopher Murphy, Daqing Hou, Nazar Khan, Ganesh Ananthanarayanan, and Faraz Hussain. Machine learning at the network edge: A survey. ACM Computing Surveys (CSUR), 54(8):1–37, 2021.
  194. 194.Goran Nakerst, John Brennan, and Masudul Haque. Gradient descent with momentum—to accelerate or to super-accelerate? arXiv preprint arXiv:2001.06472, 2020.
  195. 195.Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, et al. Do transformer modifications transfer across implementations and applications? arXiv preprint arXiv:2102.11972, 2021.
  196. 196.Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15, 2021.
  197. 197.Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens. Adding gradient noise improves learning for very deep networks. arXiv preprint arXiv:1511.06807, 2015.
  198. 198.Yurii Evgen’evich Nesterov. A method of solving a convex programming problem with convergence rate o(kˆ2). In Doklady Akademii Nauk, volume 269, pages 543–547. Russian Academy of Sciences, 1983.
  199. 199.Timothy Nguyen, Zhourong Chen, and Jaehoon Lee. Dataset meta-learning from kernel ridge-regression. arXiv preprint arXiv:2011.00050, 2020.
  200. 200.Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee. Dataset distillation with infinitely wide convolutional networks. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  201. 201.Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat, Ari Morcos, and Nicolas Ballas. Trade-offs of local sgd at scale: An empirical study. arXiv preprint arXiv:2110.08133, 2021.
  202. 202.Samet Oymak. Provable super-convergence with a large cyclical learning rate. IEEE Signal Processing Letters, 28:1645–1649, 2021.
  203. 203.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In International Conference on Machine Learning, pages 4055–4064. PMLR, 2018.
  204. 204.Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International conference on machine learning, pages 1310–1318. PMLR, 2013.
  205. 205.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017.
  206. 206.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  207. 207.David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Hung Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeffrey Dean. The carbon footprint of machine learning training will plateau, then shrink. 2022.
  208. 208.David A Patterson, Garth Gibson, and Randy H Katz. A case for redundant arrays of inexpensive disks (raid). In Proceedings of the 1988 ACM SIGMOD international conference on Management of data, pages 109–116, 1988.
  209. 209.Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. Deep learning on a data diet: Finding important examples early in training. Advances in Neural Information Processing Systems, 34:20596–20607, 2021.
  210. 210.J Gregory Pauloski, Qi Huang, Lei Huang, Shivaram Venkataraman, Kyle Chard, Ian Foster, and Zhao Zhang. Kaisa: an adaptive second-order optimizer framework for deep neural networks. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–14, 2021.
  211. 211.Maxime Pistono, Gouenou Coatrieux, Jean-Claude Nunes, and Michel Cozic. Training machine learning on jpeg compressed images. In 2020 Data Compression Conference (DCC), pages 388–388. IEEE, 2020.
  212. 212.Boris T Polyak. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964.
  213. 213.Jacob Portes, Davis Blalock, Cory Stephenson, and Jonathan Frankle. Fast benchmarking of accuracy vs. training time with cyclic learning rates. arXiv preprint arXiv:2206.00832, 2022.
  214. 214.Vishak Prasad, Colin White, Paarth Jain, Sibasis Nayak, Rishabh K Iyer, and Ganesh Ramakrishnan. Speeding up nas with adaptive subset selection. In First Conference on Automated Machine Learning (Late-Breaking Workshop), 2022.
  215. 215.Ofir Press and Lior Wolf. Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859, 2016.
  216. 216.Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021.
  217. 217.Ilija Raju, Kyle Daruwalla, and Mikko Lipasti. Accelerating deep learning with dynamic data pruning. arXiv preprint arXiv:2111.12621, 2021.
  218. 218.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021.
  219. 219.Varun Ranganathan and Alex Lewandowski. Zorb: A derivative-free backpropagation algorithm for neural networks. arXiv preprint arXiv:2011.08895, 2020.
  220. 220.Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. {ZeRO-Offload}: Democratizing {Billion-Scale} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pages 551–564, 2021a.
  221. 221.Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Xiaojiang Chen, and Xin Wang. A comprehensive survey of neural architecture search: Challenges and solutions. ACM Computing Surveys (CSUR), 54(4):1–34, 2021b.
  222. 222.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  223. 223.Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran. Low-rank matrix factorization for deep neural network training with high-dimensional output targets. In 2013 IEEE international conference on acoustics, speech and signal processing, pages 6655–6659. IEEE, 2013.
  224. 224.Syed Shakib Sarwar, Priyadarshini Panda, and Kaushik Roy. Gabor filter assisted energy efficient fast learning convolutional neural networks. In 2017 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), pages 1–6. IEEE, 2017.
  225. 225.Robin M Schmidt, Frank Schneider, and Philipp Hennig. Descending through a crowded valley-benchmarking deep learning optimizers. In International Conference on Machine Learning, pages 9367–9376. PMLR, 2021.
  226. 226.Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai. Communications of the ACM, 63(12):54–63, 2020.
  227. 227.Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In Fifteenth annual conference of the international speech communication association. Citeseer, 2014.
  228. 228.Jaime Sevilla and Pablo Villalobos. Review of parameter counts in machine learning. Published: Alignment Forum (blog), 2021.
  229. 229.Or Sharir, Barak Peleg, and Yoav Shoham. The cost of training nlp models: A concise overview. arXiv preprint arXiv:2004.08900, 2020.
  230. 230.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018.
  231. 231.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  232. 232.Sheng Shen, Alexei Baevski, Ari S Morcos, Kurt Keutzer, Michael Auli, and Douwe Kiela. Reservoir transformer. arXiv preprint arXiv:2012.15045, 2020.
  233. 233.Connor Shorten and Taghi M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of big data, 6(1):1–48, 2019.
  234. 234.Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker. Metadata archaeology: Unearthing data subsets by leveraging training dynamics. arXiv preprint arXiv:2209.10015, 2022.
  235. 235.David Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel, and Hado van Hasselt. Learning by directional gradient descent. In International Conference on Learning Representations, 2021.
  236. 236.Abhishek Sinha, Mausoom Sarkar, Aahitagni Mukherjee, and Balaji Krishnamurthy. Introspection: Accelerating neural network training by learning weight evolution. arXiv preprint arXiv:1704.04959, 2017.
  237. 237.Durga Sivasubramanian, Rishabh Iyer, Ganesh Ramakrishnan, and Abir De. Training data subset selection for regression with controlled generalization error. arXiv preprint arXiv:2106.12491, 2021.
  238. 238.Leslie N Smith. Cyclical learning rates for training neural networks. In 2017 IEEE winter conference on applications of computer vision (WACV), pages 464–472. IEEE, 2017.
  239. 239.Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. arxiv e-prints, page. ArXiv preprint arXiv:1708.07120, 2017.
  240. 240.Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le. Don’t decay the learning rate, increase the batch size. arXiv preprint arXiv:1711.00489, 2017.
  241. 241.David R So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le. Primer: Searching for efficient transformers for language modeling. arXiv preprint arXiv:2109.08668, 2021.
  242. 242.Sanghyun Son, Seungjun Nah, and Kyoung Mu Lee. Clustering convolutional kernels to compress deep neural networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 216–232, 2018.
  243. 243.Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos. Beyond neural scaling laws: beating power law scaling via data pruning. arXiv preprint arXiv:2206.14486, 2022.
  244. 244.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16519–16529, 2021.
  245. 245.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  246. 246.Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your vit? data, augmentation, and regularization in vision transformers. arXiv preprint arXiv:2106.10270, 2021.
  247. 247.Cory Stephenson. Colout. Technical report, 2021.
  248. 248.Sebastian U Stich. Local sgd converges fast and communicates little. arXiv preprint arXiv:1805.09767, 2018.
  249. 249.Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243, 2019.
  250. 250.Ilia Sucholutsky and Matthias Schonlau. ’less than one’-shot learning: Learning n classes from m¡ n samples. arXiv preprint arXiv:2009.08449, 2020.
  251. 251.Yi-Lin Sung, Varun Nair, and Colin A Raffel. Training neural networks with fixed sparse masks. Advances in Neural Information Processing Systems, 34, 2021.
  252. 252.Mingxing Tan and Quoc V Le. Efficientnetv2: Smaller models and faster training. arXiv preprint arXiv:2104.00298, 2021.
  253. 253.Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He. 1-bit adam: Communication efficient large-scale training with adam’s convergence speed. In International Conference on Machine Learning, pages 10118–10129. PMLR, 2021.
  254. 254.Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. Sparse sinkhorn attention. In International Conference on Machine Learning, pages 9438–9447. PMLR, 2020a.
  255. 255.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena: A benchmark for efficient transformers. arXiv preprint arXiv:2011.04006, 2020b.
  256. 256.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732, 2020c.
  257. 257.DeepSpeed Team and Rangan Majumder. DeepSpeed: Extreme-scale model training for everyone. https://www.microsoft.com/en-us/research/blog/deepspeed-extreme-scale-model-training-for-everyone/, 2020.
  258. 258.The Mosaic ML Team. composer. https://github.com/mosaicml/composer/, 2021.
  259. 259.Kale-ab Tessera, Sara Hooker, and Benjamin Rosman. Keep the gradients flowing: Using gradient flow to study sparse network optimization. arXiv preprint arXiv:2102.01670, 2021.
  260. 260.Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. The computational limits of deep learning. arXiv preprint arXiv:2007.05558, 2020.
  261. 261.Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. Deep learning’s diminishing returns: the cost of improvement is becoming unsustainable. IEEE Spectrum, 58(10):50–55, 2021.
  262. 262.Tijmen Tieleman and Geoffrey Hinton. Rmsprop: Divide the gradient by a running average of its recent magnitude. coursera: Neural networks for machine learning. COURSERA Neural Networks Mach. Learn, 2012.
  263. 263.Nicholas Gerard Timmons and Andrew Rice. Approximating activation functions. arXiv preprint arXiv:2001.06370, 2020.
  264. 264.Jeff Tollefson. Ipcc says limiting global warming to 1.5 [degrees] c will require drastic action. Nature, 562(7726):172–174, 2018.
  265. 265.Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. An empirical study of example forgetting during deep neural network learning. arXiv preprint arXiv:1812.05159, 2018.
  266. 266.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou. Fixing the train-test resolution discrepancy. Advances in neural information processing systems, 32, 2019.
  267. 267.Maxwell Van Gelder, Mitchell Wortsman, and Kiana Ehsani. Deconstructing the structure of sparse neural networks. arXiv preprint arXiv:2012.00172, 2020.
  268. 268.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  269. 269.Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi. Powersgd: Practical low-rank gradient compression for distributed optimization. Advances in Neural Information Processing Systems, 32, 2019.
  270. 270.Laura Von Rueden, Sebastian Mayer, Katharina Beckh, Bogdan Georgiev, Sven Giesselbach, Raoul Heese, Birgit Kirsch, Julius Pfrommer, Annika Pick, Rajkumar Ramamurthy, et al. Informed machine learning–a taxonomy and survey of integrating knowledge into learning systems. arXiv preprint arXiv:1903.12394, 2019.
  271. 271.Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal. On orthogonality and learning recurrent networks with long term dependencies. In International Conference on Machine Learning, pages 3570–3578. PMLR, 2017.
  272. 272.Jun-Kun Wang, Chi-Heng Lin, Andre Wibisono, and Bin Hu. Provable acceleration of heavy ball beyond quadratics for a class of polyak-lojasiewicz functions when the non-convexity is averaged-out. In International Conference on Machine Learning, pages 22839–22864. PMLR, 2022.
  273. 273.Rui Wang and Rose Yu. Physics-guided deep learning for dynamical systems: A survey. arXiv preprint arXiv:2107.01272, 2021.
  274. 274.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  275. 275.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018.
  276. 276.Xin Wang, Yudong Chen, and Wenwu Zhu. A survey on curriculum learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  277. 277.Kai Wei, Yuzong Liu, Katrin Kirchhoff, and Jeff Bilmes. Using document summarization techniques for speech data subset selection. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 721–726, 2013.
  278. 278.Kai Wei, Yuzong Liu, Katrin Kirchhoff, and Jeff Bilmes. Unsupervised submodular subset selection for speech data. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4107–4111. IEEE, 2014.
  279. 279.Kai Wei, Rishabh Iyer, and Jeff Bilmes. Submodularity in data subset selection and active learning. In International conference on machine learning, pages 1954–1963. PMLR, 2015.
  280. 280.Daphna Weinshall, Gad Cohen, and Dan Amir. Curriculum learning by transfer learning: Theory and experiments with deep networks. In International Conference on Machine Learning, pages 5238–5246. PMLR, 2018.
  281. 281.Colin White, Mahmoud Safari, Rhea Sukthanker, Binxin Ru, Thomas Elsken, Arber Zela, Debadeepta Dey, and Frank Hutter. Neural architecture search: Insights from 1000 papers. arXiv preprint arXiv:2301.08727, 2023.
  282. 282.Simon Wiesler, Alexander Richard, Ralf Schlüter, and Hermann Ney. Mean-normalized stochastic gradient for large-scale deep learning. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 180–184. IEEE, 2014.
  283. 283.Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021.
  284. 284.Samuel Williams, Andrew Waterman, and David Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009.
  285. 285.Paul Wimmer, Jens Mehnert, and Alexandru Condurache. Freezenet: Full performance by reduced storage costs. In Proceedings of the Asian Conference on Computer Vision, 2020.
  286. 286.Bichen Wu, Alvin Wan, Xiangyu Yue, Peter Jin, Sicheng Zhao, Noah Golmant, Amir Gholaminejad, Joseph Gonzalez, and Kurt Keutzer. Shift: A zero flop, zero parameter alternative to spatial convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9127–9135, 2018a.
  287. 287.Lijun Wu, Fei Tian, Yingce Xia, Yang Fan, Tao Qin, Jianhuang Lai, and Tie-Yan Liu. Learning to teach with dynamic loss functions. arXiv preprint arXiv:1810.12081, 2018b.
  288. 288.Xiaoxia Wu, Ethan Dyer, and Behnam Neyshabur. When do curricula work? arXiv preprint arXiv:2012.03107, 2020.
  289. 289.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European conference on computer vision (ECCV), pages 3–19, 2018.
  290. 290.Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan. Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models. arXiv preprint arXiv:2208.06677, 2022.
  291. 291.Hang Xu, Chen-Yu Ho, Ahmed M Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis. Compressed communication for distributed deep learning: Survey and quantitative evaluation. Technical report, 2020.
  292. 292.Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tuning large neural networks via zero-shot hyperparameter transfer. Advances in Neural Information Processing Systems, 34:17084–17097, 2021.
  293. 293.Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Chris De Sa. Swalp: Stochastic weight averaging in low precision training. In International Conference on Machine Learning, pages 7015–7024. PMLR, 2019.
  294. 294.Zichao Yang, Marcin Moczulski, Misha Denil, Nando De Freitas, Alex Smola, Le Song, and Ziyu Wang. Deep fried convnets. In Proceedings of the IEEE International Conference on Computer Vision, pages 1476–1483, 2015.
  295. 295.Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney. Adahessian: An adaptive second order optimizer for machine learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10665–10673, 2021.
  296. 296.Hongwei Yong, Jianqiang Huang, Xiansheng Hua, and Lei Zhang. Gradient centralization: A new optimization technique for deep neural networks. In European Conference on Computer Vision, pages 635–652. Springer, 2020.
  297. 297.Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Richard G Baraniuk, Zhangyang Wang, and Yingyan Lin. Drawing early-bird tickets: Towards more efficient training of deep networks. arXiv preprint arXiv:1909.11957, 2019.
  298. 298.Jie You, Jae-Won Chung, and Mosharaf Chowdhury. Zeus: Understanding and optimizing gpu energy consumption of dnn training. arXiv preprint arXiv:2208.06102, 2022.
  299. 299.Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888, 2017.
  300. 300.Adams Wei Yu, Lei Huang, Qihang Lin, Ruslan Salakhutdinov, and Jaime Carbonell. Block-normalized gradient method: An empirical study for training deep neural network. arXiv preprint arXiv:1707.04822, 2017.
  301. 301.Geng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li, Zhenglun Kong, Ning Liu, Yifan Gong, Zheng Zhan, Chaoyang He, Qing Jin, et al. Mest: Accurate and fast memory-economic sparse training framework on the edge. arXiv preprint arXiv:2110.14032, 2021.
  302. 302.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  303. 303.X Zhai, A Kolesnikov, N Houlsby, and L Beyer. Scaling vision transformers. arxiv 2021. arXiv preprint arXiv:2106.04560, 2021.
  304. 304.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  305. 305.Lin Zhang, Shaohuai Shi, Wei Wang, and Bo Li. Scalable k-fac training for deep neural networks with distributed preconditioning. arXiv preprint arXiv:2206.15143, 2022.
  306. 306.Xiangyu Zhang, Jianhua Zou, Kaiming He, and Jian Sun. Accelerating very deep convolutional networks for classification and detection. IEEE transactions on pattern analysis and machine intelligence, 38(10):1943–1955, 2015.
  307. 307.Zhenyu Zhang, Xuxi Chen, Tianlong Chen, and Zhangyang Wang. Efficient lottery ticket finding: Less data is more. In International Conference on Machine Learning, pages 12380–12390. PMLR, 2021.
  308. 308.Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. Dataset condensation with gradient matching. arXiv preprint arXiv:2006.05929, 2020.

Citation

MLA
Bartoldson, B. R., et al. “Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities”. Journal of Machine Learning Research, vol. 24, no. 122, 2023, pp. 1–7, https://www.jmlr.org/papers/v24/22-1208.html.
APA
Bartoldson, B. R., Kailkhura, B., & Blalock, D. (2023). Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities. Journal of Machine Learning Research, 24(122), 1–77. https://www.jmlr.org/papers/v24/22-1208.html
Chicago
Bartoldson, B. R., B. Kailkhura, and D. Blalock. 2023. “Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities”. Journal of Machine Learning Research 24 (122): 1–77. https://www.jmlr.org/papers/v24/22-1208.html.
Harvard
Bartoldson, B.R., Kailkhura, B. and Blalock, D. (2023) “Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities”, Journal of Machine Learning Research, 24(122), pp. 1–77. Available at: https://www.jmlr.org/papers/v24/22-1208.html.
Vancouver
1. Bartoldson BR, Kailkhura B, Blalock D (2023) Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities. Journal of Machine Learning Research 24:1–77

BibTeX

@article{JMLR:v24:22-1208,
  author  = {Brian R. Bartoldson and Bhavya Kailkhura and Davis Blalock},
  title   = {Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {122},
  pages   = {1--77},
  url     = {http://jmlr.org/papers/v24/22-1208.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/