Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities
Brian R. BartoldsonBhavya KailkhuraDavis W. Blalock
Presents a unified taxonomy and rigorous evaluation framework for algorithmically efficient deep learning training techniques, equipping practitioners to systematically identify bottlenecks and combine methods to cut computational costs.
Modern deep learning relies on training massive neural network models on vast datasets. While this strategy achieves state-of-the-art performance in domains such as computer vision and natural language processing, model sizes are doubling as rapidly as every 4 to 8 months. This rapid scaling leads to unsustainable economic costs, substantial carbon emissions, and resource concentration among a few well-funded entities. Although hardware and software engineering provide efficiency gains, relying exclusively on them is insufficient to bridge the gap to theoretical computing limits. The article formalizes the algorithmic speedup challenge, creates a structured taxonomy of methods that modify the training routine semantics, identifies standardized evaluation practices, and maps these techniques to underlying computing bottlenecks.
The analysis evaluates algorithmic speedup methods through a literature survey of over 200 publications and targeted empirical microbenchmarks on modern processors and accelerators, including A100 graphics processing units and multicore central processing units. The article establishes a unifying taxonomy around three core training components (function, data, and optimization) and five actions (remove, restrict, reorder, replace, and retrofit), which operate through static or dynamic mechanisms based on domain knowledge or machine learning.
The findings show that widely used theoretical metrics fail to predict real-world performance. Floating-point operations and raw parameter counts do not reliably correlate with wall-clock training time. Operations with low arithmetic intensity, such as normalization layers, are heavily limited by memory bandwidth and take orders of magnitude more time per arithmetic operation than dense matrix multiplications. Factorizing layers can reduce theoretical operation counts while increasing actual runtime due to kernel launch overhead and data movement costs. Furthermore, empirical evaluations reveal that simple baselines, such as reducing the epoch count while properly scaling the learning rate schedule, often outperform complex algorithmic speedup methods. The findings also demonstrate that combining multiple complementary speedup techniques achieves substantially better time-versus-accuracy trade-offs than relying on isolated interventions.
These results demonstrate that algorithmic interventions must be chosen based on the specific hardware bottleneck present in the training pipeline, such as data loading, accelerator memory bandwidth, or interconnect capacity. Pursuing theoretical efficiency reductions without considering hardware execution risks wasting engineering effort and increasing training costs. Organizations can achieve meaningful efficiency gains only when algorithmic techniques directly alleviate active resource bottlenecks, such as using dynamic pruning when compute bound or applying data echoing and larger models when bounded by input data loading.
Decision-makers and practitioners should establish standardized, rigorous evaluation protocols before deploying speedup strategies. Teams should measure actual wall-clock time and generate full Pareto trade-off curves across multiple hyperparameter settings rather than relying on theoretical operation counts. Furthermore, workflows should benchmark against simple baselines, including reduced training durations and smaller model variants. Researchers must also account for the composition of techniques, recognizing that speedup methods interact non-linearly.
A primary limitation of this analysis is that hardware architectures, software frameworks, and model designs continue to evolve rapidly, which can alter specific bottleneck profiles over time. Additionally, many advanced methods that perform well on smaller benchmark datasets struggle to maintain model quality or yield practical speedups on modern large-scale workloads like ImageNet. Decision-makers should maintain moderate confidence in theoretical literature claims until empirical validations are conducted on the specific target hardware, software, and dataset environments.
- Paper: Green AI, Roy Schwartz et al. (2019). This foundational paper frames the core motivation of Green AI and computational efficiency that the survey formalizes and builds upon.
- Paper: Energy and Policy Considerations for Deep Learning in NLP, Emma Strubell et al. (2019). This paper establishes the empirical environmental and financial costs of deep learning training runs that the survey seeks to mitigate algorithmically.
- Paper: The Tradeoffs of Large Scale Learning, Léon Bottou et al. (2007). This seminal work establishes the foundational computational trade-offs between optimization error, sample complexity, and compute time in large-scale machine learning.
- Paper: Bag of Tricks for Image Classification with Convolutional Neural Networks, Tong He et al. (2018). This paper systematically isolates practical algorithmic training tweaks and acceleration recipes that serve as primary examples in the survey's speedup taxonomy.
- Paper: Mixed Precision Training, Paulius Micikevicius et al. (2018). This work introduces standard mixed-precision training protocols that form a fundamental baseline for compute-efficient deep learning semantics.
- Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). This paper introduces gradient checkpointing and activation recomputation, a cornerstone algorithmic memory-reduction technique evaluated across training pipelines.
- Paper: The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks., Jonathan Frankle et al. (2019). This work formulates the Lottery Ticket Hypothesis, laying the groundwork for algorithmic sparse training and network pruning methods covered in the survey.
- Paper: Rigging the Lottery: Making All Tickets Winners, Utku Evci et al. (2020). This paper establishes dynamic sparse training techniques (RigL) that enable end-to-end compute savings during the training phase.
- Paper: Adafactor: Adaptive Learning Rates with Sublinear Memory Cost, Noam Shazeer et al. (2018). This work details memory-efficient adaptive optimization via matrix factorization, addressing a primary training bottleneck discussed in the taxonomy.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). This survey provides an architectural taxonomy of efficient attention mechanisms that directly informs the broader algorithmic speedup classifications.
- Paper: Scaling Laws for Fine-Grained Mixture of Experts, Jan Ludziejewski et al. (2024). This paper extends algorithmic efficiency principles by developing precise scaling laws for fine-grained mixture-of-experts training.
- Paper: Adam-mini: Use Fewer Learning Rates To Gain More, Yushun Zhang et al. (2025). This work applies structural optimizer reduction to slash training memory and communication overhead by sharing learning rates across parameter blocks.
- Paper: Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, David Raposo et al. (2024). This paper introduces Mixture-of-Depths to dynamically allocate compute per token along network depth, realizing the dynamic computation paradigms highlighted in the survey.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This comprehensive survey continues the discussion by focusing specifically on state-of-the-art compute-efficient architectures and linear attention mechanisms for modern large language models.
- Paper: LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models, Yaowei Zheng et al. (2024). This work provides a unified, practical software implementation integrating modern algorithmic and parameter-efficient fine-tuning methods.
- Paper: Sparser, Faster, Lighter Transformer Language Models, Edoardo Cetin et al. (2026). This paper implements custom kernels and regularized activation sparsity to achieve real-world training and inference acceleration in transformer language models.
- Paper: Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference, Mostafa Elhoushi et al. (2026). This study demonstrates how scheduled layer dropout can systematically reduce total training FLOPs during large-scale pretraining without sacrificing model quality.
- Paper: Resource-Efficient Neural Networks for Embedded Systems, Wolfgang Roth et al. (2024). This work evaluates how algorithmic compression and efficiency methods translate into real-world inference deployment on resource-constrained hardware.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey expands the scope of algorithmic compute efficiency beyond standalone models to multi-step recursive workflows in AI agents.
