SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient

Max RyabininTim DettmersMichael DiskinAlexander Borzunov

article2023ICML85 citations

Proposes a fault-tolerant, decentralized model-parallel training algorithm that dynamically rebalances pipeline stages to train billion-scale language models across cheap, unreliable hardware over slow network connections.

Listen

Training modern deep learning models with billions of parameters currently requires expensive, specialized high-performance computing clusters with ultra-fast interconnects. Because individual hardware failures break traditional distributed training pipelines, large-scale model training remains largely out of reach for organizations without massive capital budgets. Concurrently, cheap transient cloud hardware, such as preemptible instances, and distributed collaborative setups remain largely underutilized for billion-scale models due to high device unreliability, hardware heterogeneity, and slow consumer-grade network speeds.

The article investigates whether billion-scale neural network training is viable over slow, unreliable, and heterogeneous networks. To achieve this, it introduces and evaluates SWARM parallelism—a decentralized, fault-tolerant model-parallel training algorithm designed to operate across cheap preemptible instances without specialized networking.

The authors conducted a series of mathematical scaling analyses, network simulation benchmarks, and real-world pretraining experiments. They measured processing idle times across various model sizes and latency profiles, evaluated dynamic peer rebalancing across heterogeneous nodes, and benchmarked a 1-billion shared parameter Transformer model across 400 preemptible cloud graphic processing units (GPUs) over a four-week span.

The article demonstrates several critical findings. First, distributed model training follows a counterintuitive "Square-Cube Law": computational demands scale cubically with model size while communication requirements scale quadratically, meaning larger models become naturally more communication-efficient and can maintain over 80% device utilization even at modest 500 Mb/s network bandwidths. Second, SWARM parallelism maintains training progress without failure as long as a single operational node remains in each pipeline stage, dynamically rerouting data around failed peers. Third, adaptive stage rebalancing effectively recovers near-optimal throughput (within 95–97% of theoretical optimum), preventing bottlenecks when nodes disconnect. Finally, pairing SWARM with 8-bit activation compression enabled successful large-scale Transformer pretraining on low-power, preemptible hardware with network requirements under 200 Mb/s, slashing total compute costs by nearly two-thirds compared to traditional dedicated infrastructure without degrading training convergence.

These findings indicate that specialized supercomputers are no longer strictly necessary to train state-of-the-art foundation models. By enabling reliable execution on transient and geographically distributed hardware, SWARM parallelism significantly lowers the financial barrier to entry, mitigating compute cost and infrastructure lock-in risks for researchers and organizations.

Decision-makers should consider adopting SWARM parallelism alongside 8-bit activation quantization when orchestrating large-scale training workloads on preemptible instances, multi-region cloud setups, or heterogeneous hardware pools. However, conventional distributed data-parallel or high-performance pipeline frameworks remain the preferred choice for smaller models under one billion parameters or within dedicated, failure-free high-performance clusters. Further engineering validation is recommended for organizations seeking to scale beyond dozens of pipeline stages, where pipeline bubble overheads and higher preemption variance can introduce performance degradation.

No sufficiently relevant recommendations were found.

Cover for SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient

Abstract

Many deep learning applications benefit from using large models with billions of parameters. Training these models is notoriously expensive due to the need for specialized HPC clusters. In this work, we consider alternative setups for training large models: using cheap “preemptible” instances or pooling existing resources from multiple regions. We analyze the performance of existing model-parallel algorithms in these conditions and find configurations where training larger models becomes less communication-intensive. Based on these findings, we propose SWARM parallelism¹, a model-parallel training algorithm designed for poorly connected, heterogeneous and unreliable devices. SWARM creates temporary randomized pipelines between nodes that are rebalanced in case of failure. We empirically validate our findings and compare SWARM parallelism with existing large-scale training approaches. Finally, we combine our insights with compression strategies to train a large Transformer language model with 1B shared parameters (≈13B before sharing) on preemptible T4 GPUs with less than 200Mb/s network.

Table of Contents

  • 1. Introduction
  • 2. Background & Related Work
  • 2.1. Model-Parallel Training
  • 2.2. Distributed Training Outside HPC
  • 2.3. Communication Efficiency and Compression
  • 3. Communication-Efficient Model Parallelism
  • 3.1. The Square-Cube Law of Distributed Training
  • 3.2. SWARM Parallelism
  • 4. Experiments
  • 4.1. Communication Efficiency at Scale
  • 4.2. Detailed Performance Comparison
  • 4.3. Large-Scale Distributed Training
  • 4.4. Adaptive Rebalancing Evaluation
  • 5. Conclusion
  • References
  • Supplementary Material
  • A. Answers to Common Questions
  • B. Additional Related Work
  • C. Stochastic Wiring Details
  • D. Description and Complexity of Adaptive Rebalancing
  • E. Relation between SWARM and ZeRO-Offload
  • F. Additional Details for Section 4.1
  • G. Additional Details for Section 4.3
  • H. Additional Scaling Evaluation
  • I. Compression-Aware Architectures
  • I.1. Description
  • I.2. Evaluating the Speed-Quality Tradeoff
  • I.3. Additional Experiments
  • J. Time To Solution

Knowls

  1. Knowl 1 — The square-cube scaling advantage of pipeline parallelism

    theoretical result

    For a pipeline stage that multiplies n×nn\times n matrices, computation grows as O(n3)O(n^3) while transferring an n×nn\times n activation or gradient grows as O(n2)O(n^2). Thus, as model dimensions increase, computation can grow faster than communication, reducing communication overhead relative to computation. The same pattern appears in common architectures: for batch size BB, spatial dimensions H,WH,W, channel width CC, sequence length LL, and hidden width hh, the paper estimates CNN computation as O(BHWC2)O(BHWC^2) versus communication O(BHWC)O(BHWC); recurrent-network computation as O(BLh2)O(BLh^2) versus communication O(BLh)O(BLh) or O(Bh)O(Bh), depending on architecture; and Transformer attention and feedforward computation as O(BL2h)O(BL^2h) and O(BLh2)O(BLh^2), respectively, versus communication O(BLh)O(BLh). These are scaling estimates, not guarantees that every larger model or system will communicate more efficiently. Under pipeline parallelism, increasing hidden dimension can lower communication per device per unit time, making lower bandwidth and higher latency more tolerable.

  2. Knowl 2 — SWARM replaces fixed pipelines with replicated, rewired stages

    model/method

    SWARM (Stochastically Wired Adaptively Rebalanced Model Parallelism) partitions a model into consecutive stages, with each stage served by one or more peers holding the same layer subset and parameters. During each iteration, peers route microbatches through the stages using temporary connections rather than fixed neighbor links. In the forward pass, a peer processes activations and sends them onward; in the backward pass, it processes output gradients and accumulates parameter gradients. Peers average accumulated gradients with All-Reduce within their stage and update the optimizer after the required training batch has been reached. Per-peer request queues allow work to continue while requests are delayed. If a peer fails, its predecessors can reroute requests to another peer serving the same stage; a joining peer downloads the current parameters and optimizer state from active peers. Training can continue while every stage and at least one trainer remain active. SWARM also supports delayed parameter updates to overlap optimizer work with processing, and activation checkpointing to reduce memory use.

  3. Knowl 3 — Stochastic wiring routes work according to peer performance

    algorithm

    Each trainer discovers active peers for each pipeline stage through a distributed hash table (DHT) and independently routes microbatches using an interleaved weighted round-robin policy. The implementation tracks peer response time with an exponentially smoothed estimate and maintains a priority queue for each stage. When assigning work, it selects the eligible peer with the smallest accumulated processing-time priority, then increases that peer's priority by its estimated processing time; this makes faster peers receive work more often. Trainers update the estimate from observed response times. If a peer fails or times out, the trainer temporarily bans it by assigning it infinite priority; the peer becomes eligible again when it reannounces itself in the DHT, which occurs every few minutes. Trainers do not use GPUs or store trainable parameters, so multiple trainers can run on one peer. Because estimates are trainer-specific, routing can also favor peers with better network paths to that trainer. Queued requests let workers continue processing during latency spikes rather than blocking on a slow neighbor.

  4. Knowl 4 — Periodic stage rebalancing moves peers from underloaded to overloaded stages

    algorithm

    Every TT seconds, each model-serving peer publishes its local request-queue size to the DHT. The system sums queue sizes by stage, identifies the least-loaded and most-loaded stages, and selects the peer with the smallest local queue in the least-loaded stage to migrate to the most-loaded stage. The migrating peer downloads the destination stage's current parameters and optimizer state before resuming service; a new peer is assigned using the same load information. This can allocate more peers to stages with greater compute demand. The procedure takes O(MS)O(MS) operations, where MM is the maximum number of peers in any stage and SS is the number of stages; the authors report that this overhead is negligible for fewer than 10,000 peers and 100 stages because it runs alongside training.

    In simulations based on a 32-hour trace of active T4 workers, ten random seeds, and removals assigned uniformly across four stages, the following values are mean throughput as a percentage of an ideal rebalancing strategy:

    Rebalancing Overall First hour Last hour
    None 82.7 99.0 45.4
    T=300T=300 seconds 95.8 99.4 88.9
    T=60T=60 seconds 97.6 99.8 91.7

    Both rebalancing periods substantially outperformed no rebalancing, particularly in the last hour, when an unrebalanced pipeline had drifted further from the ideal. With 4, 8, 16, and 32 stages, both rebalanced and unrebalanced throughput tended to decline as stage count increased, but rebalancing partially mitigated the decline and recovered throughput over time.

  5. Knowl 5 — Compression-aware layers reduce pipeline traffic at communication boundaries

    model/method

    The paper evaluates three ways to reduce data transferred between pipeline stages. First, blockwise dynamic 8-bit quantization compresses activations and their gradients, reducing communication by about 2×2\times relative to 16-bit values and 4×4\times relative to 32-bit values. Second, a linear bottleneck projects an mm-dimensional layer output to dimension c<mc<m before transfer, then projects it back to mm; this reduces boundary traffic by m/cm/c and uses LayerNorm to support stable training. Third, maxout groups each kk consecutive features and retains their maximum, reducing the transferred dimension by kk before a learned projection restores dimension mm. In the equations below, x∈Rb×s×mx\in\mathbb{R}^{b\times s\times m} is an input with batch size bb, sequence length ss, and feature width mm; w1∈Rm×hw_1\in\mathbb{R}^{m\times h} and w2∈Rh×mw_2\in\mathbb{R}^{h\times m} are MLP weights; wc∈Rm×cw_c\in\mathbb{R}^{m\times c} and wd∈Rc×mw_d\in\mathbb{R}^{c\times m} are bottleneck projection weights; and σ\sigma is a nonlinear activation.

    MLP⁡(x,w1,w2)=σ(xw1)w2+x,Bottleneck⁡(x)=LayerNorm⁡ ⁣(LayerNorm⁡(MLP⁡(x))wc)wd.\operatorname{MLP}(x,w_1,w_2)=\sigma(xw_1)w_2+x, \qquad \operatorname{Bottleneck}(x)=\operatorname{LayerNorm}\!\left(\operatorname{LayerNorm}(\operatorname{MLP}(x))w_c\right)w_d.

    For maxout compression, kk is the number of features in each non-overlapping group and wd∈R(m/k)×mw_d\in\mathbb{R}^{(m/k)\times m} is the decompression weight:

    Maxout⁡(x)=LayerNorm⁡ ⁣(maxout⁡k(LayerNorm⁡(MLP⁡(x))))wd.\operatorname{Maxout}(x)=\operatorname{LayerNorm}\!\left(\operatorname{maxout}_k(\operatorname{LayerNorm}(\operatorname{MLP}(x)))\right)w_d.

    On WikiText-103, the authors compared these methods using a Transformer language model with two pipeline stages and a 2×2\times compression factor. Perplexity was measured after 286,000 steps; steps-to-perplexity-22 and data transfer are relative to the uncompressed model. The asterisk marks a difference reported as not statistically significant.

    Method Perplexity Steps to ppl 22 Data transfer Extra compute (absolute) Extra compute (relative)
    No compression 21.02 1×1\times 1×1\times 0 None
    8-bit compression 21.13 0.97×∗0.97\times^{*} 0.5×0.5\times 1.2 ms None (overlapped)
    Bottleneck 21.76 1.26×1.26\times 0.5×0.5\times 1.96 ms ≤1%\leq 1\%
    Maxout 21.83 1.28×1.28\times 0.5×0.5\times 2.04 ms ≤1%\leq 1\%

    In this test, 8-bit compression approximately preserved the baseline's convergence time and final perplexity, while bottleneck and maxout reduced transfer by half but required 26–28% more steps to reach perplexity 22 and had slightly worse final perplexity. The authors report less than 1% extra computation for the bottleneck and maxout layers.

  6. Knowl 6 — Larger Transformer layers improve utilization on bandwidth-limited links

    empirical result

    The authors measured the fraction of time a V100 GPU was computing rather than idle while serving a pipeline stage over a 500 Mb/s link in both directions. The batch contained one sequence of 512 tokens. The base, xxlarge, and GPT-3 configurations placed 12 Transformer layers on 12 single-GPU servers; the “Ours” configuration placed three layers per stage on four servers and used 8-bit activations. The table reports relative GPU utilization, equal to 100% minus idle time, for the stated round-trip latency (RTT).

    RTT base xxlarge GPT-3 Ours
    None 18.0% 32.1% 82.1% 89.5%
    10 ms 11.8% 28.9% 79.3% 87.2%
    50 ms 4.88% 20.1% 70.3% 79.5%
    100 ms 2.78% 14.9% 60.2% 71.5%
    200 ms 1.53% 10.1% 48.5% 59.2%

    The base layer had only 18.0% utilization even with no added latency, compared with 82.1% for GPT-3-scale layers and 89.5% for the multi-layer, quantized configuration. At 100 ms RTT, the latter two still achieved 60.2% and 71.5% utilization, respectively. These results support the scaling analysis: larger compute-intensive stages can better amortize communication, although added latency reduces utilization in every configuration.

  7. Knowl 7 — SWARM is competitive with pipeline and offload baselines under idealized conditions

    empirical result

    The authors compared SWARM with GPipe, its 1F1B schedule, and ZeRO-Offload using four Transformer layers distributed across 16 V100 workers. Workers had 500 Mb/s upload and download bandwidth; the latency condition added 100±50100\pm50 ms. Throughput is reported as minutes to process 6,250 sequences of 512 tokens, so lower is faster; All-Reduce columns report minutes. The xxlarge microbatch size was 4, while GPT-3 and “Ours” used size 1. “Ours” used shared layers and a vocabulary-projection final stage. A dash means the All-Reduce time was not separately reported for the 1F1B row.

    Model System Throughput, no latency (min) Throughput, latency (min) All-Reduce, no latency (min) All-Reduce, latency (min)
    GPT-3 (4 layers) SWARM 168.3 186.7 7.4 7.6
    GPipe 164.5 218.4 6.7 7.8
    1F1B 163.3 216.1 — —
    ZeRO-Offload 272.7 272.7 25.5 27.3
    xxlarge (4 layers) SWARM 44.2 48.2 0.8 0.9
    GPipe 40.1 108.8 0.7 1.1
    1F1B 40.8 105.5 — —
    ZeRO-Offload 33.8 33.8 2.8 4.2
    Full “Ours” model SWARM 432.2 452.9 0.8 1.0
    (48 shared layers + embeddings) GPipe 420.0 602.1 0.7 1.1
    1F1B 408.5 569.2 — —
    ZeRO-Offload 372.0 372.0 3.2 4.8

    SWARM's throughput was competitive with GPipe in these homogeneous tests, and GPipe throughput was more affected by added latency for the xxlarge and full “Ours” cases. ZeRO-Offload was faster for xxlarge and the full “Ours” model without latency, but substantially slower for GPT-3; it also spent longer in All-Reduce than the pipeline methods. The results show that SWARM is not universally fastest in reliable homogeneous settings, but can remain competitive while accommodating latency.

  8. Knowl 8 — SWARM trained a billion-parameter shared Transformer on preemptible GPUs

    empirical result

    The authors trained a 1.01-billion-parameter Transformer language model using SWARM on preemptible T4 GPUs over a public network; the abstract characterizes the network bandwidth as below 200 Mb/s. The model had three pipeline stages, each serving a decoder block with model width 4,096 and 16 parameter-shared layers. The first stage also contained embeddings, and the final stage contained the language-model head. Sharing makes the model equivalent to roughly 13 billion parameters in compute cost. Activations and gradients were compressed to 8 bits. Training used 400 single-GPU T4 instances and the Pile dataset for approximately four weeks. Its learning curve was compared with data-parallel training with offloading on 128 A100 GPUs, also run for approximately four weeks; the loss trajectories were similar over the reported training run. The setup used LAMB with batch size 16,384 and increased sequence length from at most 256 to 2,048 over the first 12,000 optimizer steps. This experiment demonstrates feasibility and comparable observed training dynamics in the tested setup, rather than establishing equivalence for all models or training conditions.

  9. Knowl 9 — Pipeline throughput on T4, A100, and heterogeneous fleets

    empirical result

    To measure pipeline performance independently of convergence, the authors routed randomly generated samples through the pipeline on public-cloud hardware and compared observed throughput with an idealized estimate that ignored network operations. The ideal estimate was obtained by running the stages locally on one server and scaling the single-node estimate by node count. The tested fleets included 400 single-GPU T4 instances, seven instances with eight A100 GPUs each, and a mixture of T4 and A100 devices. With shared layers and 8-bit compression, throughput was measured in samples per second; bandwidth values are the measured upload and download bandwidth estimated to fully utilize the device type.

    Layer-sharing setup Hardware Actual throughput (samples/s) Best-case throughput (samples/s) Optimal upload / download (Mb/s)
    Shared layers T4 17.6 19.2 317.8 / 397.9
    Shared layers A100 16.9 25.5 436.1 / 545.1
    Shared layers T4 and A100 27.3 — —
    Default Transformer T4 8.8 19.3 —
    Default Transformer A100 8.0 25.1 —
    Default Transformer T4 and A100 13.4 — —

    The heterogeneous fleet achieved 27.3 samples/s with layer sharing and 13.4 samples/s with the default Transformer. For individual device types, observed throughput was below the ideal estimate, especially for A100s, indicating greater headroom from faster networking. Layer sharing and 8-bit compression made the compute-intensive shared-layer configuration more network-efficient; the authors note that total bandwidth needs in the main training experiment were roughly 100 Mb/s above the tabulated estimates because of gradient averaging, state loading, DHT traffic, and data streaming.

  10. Knowl 10 — Preemptible SWARM training reduced the cost and time to a target objective

    empirical result

    The authors compared wall-clock time and estimated cloud cost to reach an ALBERT training objective of 1.5 on WikiText-103. The model used four layer groups corresponding to four SWARM stages, without the compression-aware architecture modifications. SWARM with eight preemptible V100s matched the per-iteration learning curve of conventional distributed data parallelism (DDP) with eight reliable V100s to within variation comparable to changing the random seed. SWARM with 32 preemptible T4s reached the target in less time and at lower total cost in this experiment.

    Setup Time (hours) Hourly cost ($) Total cost ($)
    8 ×\times V100, reliable 175.4 7.834 1374
    8 ×\times V100, preemptible 192.6 5.383 1037
    32 ×\times T4, preemptible 140.8 3.536 497.8

    The 32-T4 preemptible setup reached the objective in 140.8 hours for 497.8,versus175.4hoursand497.8, versus 175.4 hours and 1,374 for the reliable eight-V100 setup. These are results for the stated model, objective, and pricing estimates; the paper also reports slightly inferior performance for SWARM compared with conventional methods when restricted to homogeneous reliable GPUs.

Coverage note — The supplementary OpenWebText compression sweep and the 8–128-node scaling experiment are omitted because they provide secondary corroboration rather than additional core mechanisms or headline results.

References

  1. 1.Aji, A. F. and Heafield, K. Making asynchronous stochastic gradient descent work for transformers. In Proceedings of the 3rd Workshop on Neural Generation and Translation, pp. 80–89, Hong Kong, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-5608. URL https://aclanthology.org/D19-5608.
  2. 2.Allen, D. H. How Mechanics Shaped the Modern World. 2013. ISBN 9783319017013.
  3. 3.Alman, J. and Williams, V. V. A refined laser method and faster matrix multiplication. In Marx, D. (ed.), Proceedings of the 2021 ACM-SIAM Symposium on Discrete Algorithms, SODA 2021, Virtual Conference, January 10 - 13, 2021, pp. 522–539. SIAM, 2021. doi: 10.1137/1.9781611976465.32. URL https://doi.org/10.1137/1.9781611976465.32.
  4. 4.Arjevani, Y., Shamir, O., and Srebro, N. A tight convergence analysis for stochastic gradient descent with delayed updates. In Kontorovich, A. and Neu, G. (eds.), Proceedings of the 31st International Conference on Algorithmic Learning Theory, volume 117 of Proceedings of Machine Learning Research, pp. 111–132. PMLR, 2020. URL https://proceedings.mlr.press/v117/arjevani20a.html.
  5. 5.Atre, M., Jha, B., and Rao, A. Distributed deep learning using volunteer computing-like paradigm. In IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPS Workshops 2021, Portland, OR, USA, June 17-21, 2021, pp. 933–942. IEEE, 2021. doi: 10.1109/IPDPSW52791.2021.00144. URL https://doi.org/10.1109/IPDPSW52791.2021.00144.
  6. 6.Ba, L. J., Kiros, J. R., and Hinton, G. E. Layer normalization. ArXiv preprint, abs/1607.06450, 2016. URL https://arxiv.org/abs/1607.06450.
  7. 7.Baevski, A. and Auli, M. Adaptive input representations for neural language modeling. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=ByxZX20qFQ.
  8. 8.Baines, M., Bhosale, S., Caggiano, V., Goyal, N., Goyal, S., Ott, M., Lefaudeux, B., Liptchinsky, V., Rabbat, M., Sheiffer, S., Sridhar, A., and Xu, M. Fairscale: A general purpose modular pytorch library for high performance and large scale training. https://github.com/facebookresearch/fairscale, 2021.
  9. 9.Ben-Nun, T. and Hoefler, T. Demystifying parallel and distributed deep learning: An in-depth concurrency analysis. ACM Comput. Surv., 52(4), 2019. ISSN 0360-0300. doi: 10.1145/3320060. URL https://doi.org/10.1145/3320060.
  10. 10.Black, S., Leo, G., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, 2021. URL https://doi.org/10.5281/zenodo.5297715.
  11. 11.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  12. 12.Chen, T., Xu, B., Zhang, C., and Guestrin, C. Training deep nets with sublinear memory cost. ArXiv preprint, abs/1604.06174, 2016. URL https://arxiv.org/abs/1604.06174.
  13. 13.Chilimbi, T., Suzue, Y., Apacible, J., and Kalyanaraman, K. Project adam: Building an efficient and scalable deep learning training system. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14), pp. 571–582, Broomfield, CO, 2014. USENIX Association. ISBN 978-1-931971-16-4. URL https://www.usenix.org/conference/osdi14/technical-sessions/presentation/chilimbi.
  14. 14.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. PaLM: Scaling language modeling with pathways. CoRR, abs/2204.02311, 2022. doi: 10.48550/arXiv.2204.02311. URL https://doi.org/10.48550/arXiv.2204.02311.
  15. 15.Coates, A., Huval, B., Wang, T., Wu, D. J., Catanzaro, B., and Ng, A. Y. Deep learning with COTS HPC systems. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 of JMLR Workshop and Conference Proceedings, pp. 1337–1345. JMLR.org, 2013. URL http://proceedings.mlr.press/v28/coates13.html.
  16. 16.Coppersmith, D. and Winograd, S. Matrix multiplication via arithmetic progressions. Journal of Symbolic Computation, 9(3):251–280, 1990. ISSN 0747-7171. doi: https://doi.org/10.1016/S0747-7171(08)80013-2. URL https://www.sciencedirect.com/science/article/pii/S0747717108800132. Computational algebraic complexity editorial.
  17. 17.Dai, Z., Liu, H., Le, Q. V., and Tan, M. Coatnet: Marrying convolution and attention for all data sizes. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 3965–3977, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/20568692db622456cc42a2e853ca21f8-Abstract.html.
  18. 18.Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M. Z., Ranzato, M., Senior, A. W., Tucker, P. A., Yang, K., and Ng, A. Y. Large scale distributed deep networks. In Bartlett, P. L., Pereira, F. C. N., Burges, C. J. C., Bottou, L., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States, pp. 1232–1240, 2012. URL https://proceedings.neurips.cc/paper/2012/hash/6aca97005c68f1206823815f66102863-Abstract.html.
  19. 19.Dettmers, T. 8-bit approximations for parallelism in deep learning. In Bengio, Y. and LeCun, Y. (eds.), 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1511.04561.
  20. 20.Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L. 8-bit optimizers via block-wise quantization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=shpkpVXzo3h.
  21. 21.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  22. 22.Dhariwal, P. and Nichol, A. Q. Diffusion models beat gans on image synthesis. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 8780–8794, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/49ad23d1ec9fa4bd8d77d02681df5cfa-Abstract.html.
  23. 23.Diskin, M., Bukhtiyarov, A., Ryabinin, M., Saulnier, L., Lhoest, Q., Sinitsin, A., Popov, D., Pyrkin, D. V., Kashirin, M., Borzunov, A., del Moral, A. V., Mazur, D., Kobelev, I., Jernite, Y., Wolf, T., and Pekhimenko, G. Distributed deep learning in open collaborations. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 7879–7897, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/41a60377ba920919939d83326ebee5a1-Abstract.html.
  24. 24.ElasticHorovod. Elastic Horovod. https://horovod.readthedocs.io/en/stable/elastic_include.html. Accessed: 2021-10-04.
  25. 25.Fatahalian, K., Sugerman, J., and Hanrahan, P. Understanding the efficiency of gpu algorithms for matrix-matrix multiplication. pp. 133–137, 2004. doi: 10.1145/1058129.1058148.
  26. 26.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. ArXiv preprint, abs/2101.03961, 2021. URL https://arxiv.org/abs/2101.03961.
  27. 27.Fukushima, K. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics, 36:193–202, 1980.
  28. 28.Galileo, G. Discorsi e dimostrazioni matematiche intorno a due nuove scienze. 1638.
  29. 29.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. The pile: An 800gb dataset of diverse text for language modeling, 2020.
  30. 30.Gokaslan, A. and Cohen, V. Openwebtext corpus, 2019. URL http://Skylion007.github.io/OpenWebTextCorpus.
  31. 31.Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A. C., and Bengio, Y. Maxout networks. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 of JMLR Workshop and Conference Proceedings, pp. 1319–1327. JMLR.org, 2013. URL http://proceedings.mlr.press/v28/goodfellow13.html.
  32. 32.Griewank, A. and Walther, A. Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation. ACM Transactions on Mathematical Software (TOMS), 26(1):19–45, 2000.
  33. 33.Harlap, A., Tumanov, A., Chung, A., Ganger, G. R., and Gibbons, P. B. Proteus: Agile ml elasticity through tiered reliability in dynamic resource markets. In Proceedings of the Twelfth European Conference on Computer Systems, EuroSys ’17, pp. 589–604, New York, NY, USA, 2017. Association for Computing Machinery. ISBN 9781450349383. doi: 10.1145/3064176.3064182. URL https://doi.org/10.1145/3064176.3064182.
  34. 34.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 770–778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90.
  35. 35.He, P., Liu, X., Gao, J., and Chen, W. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=XPZIaotutsD.
  36. 36.Hochreiter, S. and Schmidhuber, J. Long Short-Term Memory. Technical Report FKI-207-95, Fakultät für Informatik, Technische Universität München, 1995. Revised 1996 (see www.idsia.ch/˜juergen, www7.informatik.tu-muenchen.de/˜hochreit).
  37. 37.Huang, J., Yu, C. D., and Geijn, R. A. v. d. Strassen’s algorithm reloaded on gpus. ACM Trans. Math. Softw., 46(1), 2020. ISSN 0098-3500. doi: 10.1145/3372419. URL https://doi.org/10.1145/3372419.
  38. 38.Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 103–112, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/093f65e080a295f8076b1c5722a46aa2-Abstract.html.
  39. 39.Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E. Adaptive mixtures of local experts. Neural Computation, 3(1):79–87, 1991. ISSN 0899-7667. doi: 10.1162/neco.1991.3.1.79. URL https://doi.org/10.1162/neco.1991.3.1.79.
  40. 40.Jia, Z., Zaharia, M., and Aiken, A. Beyond data and model parallelism for deep neural networks. In Talwalkar, A., Smith, V., and Zaharia, M. (eds.), Proceedings of Machine Learning and Systems 2019, MLSys 2019, Stanford, CA, USA, March 31 - April 2, 2019. mlsys.org, 2019. URL https://proceedings.mlsys.org/book/265.pdf.
  41. 41.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models, 2020.
  42. 42.Katevenis, M., Sidiropoulos, S., and Courcoubetis, C. Weighted round-robin cell multiplexing in a general-purpose atm switch chip. IEEE Journal on Selected Areas in Communications, 9(8):1265–1279, 1991. doi: 10.1109/49.105173.
  43. 43.Kijsipongse, E., Piyatumrong, A., and U-ruekolan, S. A hybrid gpu cluster and volunteer computing platform for scalable deep learning. The Journal of Supercomputing, 2018. doi: 10.1007/s11227-018-2375-9.
  44. 44.Krizhevsky, A. One weird trick for parallelizing convolutional neural networks. CoRR, abs/1404.5997, 2014. URL http://arxiv.org/abs/1404.5997.
  45. 45.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Bartlett, P. L., Pereira, F. C. N., Burges, C. J. C., Bottou, L., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States, pp. 1106–1114, 2012. URL https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html.
  46. 46.Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H. Large memory layers with product keys. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 8546–8557, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/9d8df73a3cfbf3c5b47bc9b50f214aff-Abstract.html.
  47. 47.Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=H1eA7AEtvS.
  48. 48.Langston, J. Microsoft announces new supercomputer, lays out vision for future ai work. https://blogs.microsoft.com/ai/openai-azure-supercomputer/, 2020. Accessed: 2021-10-1.
  49. 49.Larrea, V. G. V., Joubert, W., Brim, M. J., Budiardja, R. D., Maxwell, D., Ezell, M., Zimmer, C., Boehm, S., Elwasif, W. R., Oral, S., Fuson, C., Pelfrey, D., Hernandez, O. R., Leverman, D., Hanley, J., Berrill, M. A., and Tharrington, A. N. Scaling the summit: Deploying the world’s fastest supercomputer. In Weiland, M., Juckeland, G., Alam, S. R., and Jagode, H. (eds.), High Performance Computing - ISC High Performance 2019 International Workshops, Frankfurt, Germany, June 16-20, 2019, Revised Selected Papers, volume 11887 of Lecture Notes in Computer Science, pp. 330–351. Springer, 2019. doi: 10.1007/978-3-030-34356-9_26. URL https://doi.org/10.1007/978-3-030-34356-9_26.
  50. 50.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=qrwe7XHTmYb.
  51. 51.Li, C., Zhang, M., and He, Y. Curriculum learning: A regularization method for efficient and stable billion-scale GPT model pre-training. ArXiv preprint, abs/2108.06084, 2021. URL https://arxiv.org/abs/2108.06084.
  52. 52.Li, S., Walls, R. J., Xu, L., and Guo, T. Speeding up deep learning with transient servers. In 2019 IEEE International Conference on Autonomic Computing (ICAC), pp. 125–135. IEEE, 2019.
  53. 53.Li, S., Ben-Nun, T., Nadiradze, G., Digirolamo, S., Dryden, N., Alistarh, D., and Hoefler, T. Breaking (global) barriers in parallel stochastic optimization with wait-avoiding group averaging. IEEE Transactions on Parallel and Distributed Systems, pp. 1–1, 2020. ISSN 2161-9883. doi: 10.1109/tpds.2020.3040606. URL http://dx.doi.org/10.1109/TPDS.2020.3040606.
  54. 54.Lian, X., Zhang, C., Zhang, H., Hsieh, C., Zhang, W., and Liu, J. Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5330–5340, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/f75526659f31040afeb61cb7133e4e6d-Abstract.html.
  55. 55.Lin, J., Li, X., and Pekhimenko, G. Multi-node bert-pretraining: Cost-efficient approach, 2020.
  56. 56.Lin, T., Wang, Y., Liu, X., and Qiu, X. A survey of transformers. AI Open, 3:111–132, 2022. doi: 10.1016/j.aiopen.2022.10.001. URL https://doi.org/10.1016/j.aiopen.2022.10.001.
  57. 57.Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B. Deep gradient compression: Reducing the communication bandwidth for distributed training. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. URL https://openreview.net/forum?id=SkhQHMW0W.
  58. 58.Maymounkov, P. and Mazieres, D. Kademlia: A peer-to-peer information system based on the xor metric. In International Workshop on Peer-to-Peer Systems, pp. 53–65. Springer, 2002.
  59. 59.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=Byj72udxe.
  60. 60.Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M. Pipedream: Generalized pipeline parallelism for dnn training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles, SOSP ’19, pp. 1–15, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450368735. doi: 10.1145/3341301.3359646. URL https://doi.org/10.1145/3341301.3359646.
  61. 61.Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al. Efficient large-scale language model training on gpu clusters. ArXiv preprint, abs/2104.04473, 2021. URL https://arxiv.org/abs/2104.04473.
  62. 62.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pp. 48–53, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-4009. URL https://aclanthology.org/N19-4009.
  63. 63.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 8024–8035, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html.
  64. 64.Pudipeddi, B., Mesmakhosroshahi, M., Xi, J., and Bharadwaj, S. Training large neural networks with constant memory using a new execution algorithm. ArXiv preprint, abs/2002.05645, 2020. URL https://arxiv.org/abs/2002.05645.
  65. 65.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018. URL https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf.
  66. 66.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  67. 67.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G. Scaling language models: Methods, analysis & insights from training gopher, 2021.
  68. 68.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  69. 69.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimization towards training a trillion parameter models. In SC, 2020.
  70. 70.Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. ArXiv preprint, abs/2104.07857, 2021. URL https://arxiv.org/abs/2104.07857.
  71. 71.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 8821–8831. PMLR, 2021. URL http://proceedings.mlr.press/v139/ramesh21a.html.
  72. 72.Recht, B., Ré, C., Wright, S. J., and Niu, F. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Shawe-Taylor, J., Zemel, R. S., Bartlett, P. L., Pereira, F. C. N., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain, pp. 693–701, 2011. URL https://proceedings.neurips.cc/paper/2011/hash/218a0aefd1d1a4be65601cc6ddc1520e-Abstract.html.
  73. 73.Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y. Zero-offload: Democratizing billion-scale model training, 2021.
  74. 74.Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning representations by back-propagating errors. Nature, 323:533–536, 1986.
  75. 75.Ryabinin, M. and Gusev, A. Towards crowdsourced training of large neural networks using decentralized mixture-of-experts. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/25ddc0f8c9d3e22e03d3076f98d83cb2-Abstract.html.
  76. 76.Ryabinin, M., Gorbunov, E., Plokhotnyuk, V., and Pekhimenko, G. Moshpit SGD: communication-efficient decentralized training on heterogeneous unreliable devices. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 18195–18211, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/97275a23ca44226c9964043c8462be96-Abstract.html.
  77. 77.Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1715–1725, Berlin, Germany, 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1162. URL https://aclanthology.org/P16-1162.
  78. 78.Shazeer, N. GLU variants improve transformer. ArXiv preprint, abs/2002.05202, 2020. URL https://arxiv.org/abs/2002.05202.
  79. 79.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=B1ckMDqlg.
  80. 80.Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B. A. Mesh-tensorflow: Deep learning for supercomputers. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pp. 10435–10444, 2018. URL https://proceedings.neurips.cc/paper/2018/hash/3a37abdeefe1dab1b30f7c5c7e581b93-Abstract.html.
  81. 81.Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-lm: Training multi-billion parameter language models using gpu model parallelism. ArXiv preprint, abs/1909.08053, 2019. URL https://arxiv.org/abs/1909.08053.
  82. 82.Stich, S. U. and Karimireddy, S. P. The error-feedback framework: sgd with delayed gradients. Journal of Machine Learning Research, 21(237):1–36, 2020. URL http://jmlr.org/papers/v21/19-748.html.
  83. 83.Strohmaier, E., Dongarra, J., Simon, H., and Meuer, M. Fugaku. https://www.top500.org/system/179807/, 2021. Estimated energy consumption 29,899.23 kW. Accessed: 2021-10-4.
  84. 84.Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding, 2021.
  85. 85.Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y., Liu, W., Wu, Z., Gong, W., Liang, J., Shang, Z., Sun, P., Liu, W., Ouyang, X., Yu, D., Tian, H., Wu, H., and Wang, H. ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. ArXiv preprint, abs/2107.02137, 2021. URL https://arxiv.org/abs/2107.02137.
  86. 86.Tabatabaee, S. M., Le Boudec, J.-Y., and Boyer, M. Interleaved weighted round-robin: A network calculus analysis. In 2020 32nd International Teletraffic Congress (ITC 32), pp. 64–72, 2020. doi: 10.1109/ITC3249928.2020.00016.
  87. 87.Tang, Z., Shi, S., Chu, X., Wang, W., and Li, B. Communication-efficient distributed deep learning: A comprehensive survey, 2020.
  88. 88.Tarnawski, J., Narayanan, D., and Phanishayee, A. Piper: Multidimensional planner for DNN parallelization. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 24829–24840, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/d01eeca8b24321cd2fe89dd85b9beb51-Abstract.html.
  89. 89.Thorpe, J., Zhao, P., Eyolfson, J., Qiao, Y., Jia, Z., Zhang, M., Netravali, R., and Xu, G. H. Bamboo: Making preemptible instances resilient for affordable training of large dnns, 2022.
  90. 90.TorchElastic. PyTorch Elastic. https://pytorch.org/elastic. Accessed: 2021-10-04.
  91. 91.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 5998–6008, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html.
  92. 92.Verizon. Monthly ip latency data, 2021. Accessed: 2021-10-05.
  93. 93.Vogels, T., Karimireddy, S. P., and Jaggi, M. Powersgd: Practical low-rank gradient compression for distributed optimization. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 14236–14245, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/d9fbed9da256e344c1fa46bb46c34c5f-Abstract.html.
  94. 94.Wang, B. and Komatsuzaki, A. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  95. 95.Wang, J., Yuan, B., Rimanic, L., He, Y., Dao, T., Chen, B., Re, C., and Zhang, C. Fine-tuning language models over slow networks using activation quantization with guarantees. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=QDPonrGtl1.
  96. 96.Wang, S., Bai, Y., and Pekhimenko, G. BPPSA: scaling back-propagation by parallel scan algorithm. In Dhillon, I. S., Papailiopoulos, D. S., and Sze, V. (eds.), Proceedings of Machine Learning and Systems 2020, MLSys 2020, Austin, TX, USA, March 2-4, 2020. mlsys.org, 2020. URL https://proceedings.mlsys.org/book/317.pdf.
  97. 97.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6.
  98. 98.Yang, B., Zhang, J., Li, J., Ré, C., Aberger, C. R., and Sa, C. D. Pipemare: Asynchronous pipeline parallel dnn training. ArXiv, abs/1910.05124, 2019.
  99. 99.You, Y., Li, J., Reddi, S. J., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C. Large batch optimization for deep learning: Training BERT in 76 minutes. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=Syx4wnEtvH.
  100. 100.Yuan, B., He, Y., Davis, J. Q., Zhang, T., Dao, T., Chen, B., Liang, P., Re, C., and Zhang, C. Decentralized training of foundation models in heterogeneous environments. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=UHoGOaGjEq.
  101. 101.Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. Scaling vision transformers. ArXiv preprint, abs/2106.04560, 2021. URL https://arxiv.org/abs/2106.04560.
  102. 102.Zhang, P. and Gao, Y. Matrix multiplication on high-density multi-gpu architectures: Theoretical and experimental investigations. In Kunkel, J. M. and Ludwig, T. (eds.), High Performance Computing - 30th International Conference, ISC High Performance 2015, Frankfurt, Germany, July 12-16, 2015, Proceedings, volume 9137 of Lecture Notes in Computer Science, pp. 17–30. Springer, 2015. doi: 10.1007/978-3-319-20119-1_2. URL https://doi.org/10.1007/978-3-319-20119-1_2.
  103. 103.Zhang, X., Wang, J., Joshi, G., and Joe-Wong, C. Machine learning on volatile instances. In IEEE INFOCOM 2020-IEEE Conference on Computer Communications, pp. 139–148. IEEE, 2020.
  104. 104.Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I. Alpa: Automating inter- and intra-operator parallelism for distributed deep learning, 2022. URL https://arxiv.org/abs/2201.12023.

Citation

MLA
Ryabinin, M., et al. “SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient”. International Conference on Machine Learning, vol. 202, 2023, pp. 29416–40, https://proceedings.mlr.press/v202/ryabinin23a.html.
APA
Ryabinin, M., Dettmers, T., Diskin, M., & Borzunov, A. (2023). SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient. International Conference on Machine Learning, 202, 29416–29440. https://proceedings.mlr.press/v202/ryabinin23a.html
Chicago
Ryabinin, M., T. Dettmers, M. Diskin, and A. Borzunov. 2023. “SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient”. International Conference on Machine Learning 202: 29416–40. https://proceedings.mlr.press/v202/ryabinin23a.html.
Harvard
Ryabinin, M. et al. (2023) “SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient”, International Conference on Machine Learning. PMLR, pp. 29416–29440. Available at: https://proceedings.mlr.press/v202/ryabinin23a.html.
Vancouver
1. Ryabinin M, Dettmers T, Diskin M, Borzunov A (2023) SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient. In: International Conference on Machine Learning. PMLR, pp 29416–29440

BibTeX

@InProceedings{pmlr-v202-ryabinin23a,
  title = 	 {{SWARM} Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient},
  author =       {Ryabinin, Max and Dettmers, Tim and Diskin, Michael and Borzunov, Alexander},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {29416--29440},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ryabinin23a/ryabinin23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ryabinin23a.html},
  abstract = 	 {Many deep learning applications benefit from using large models with billions of parameters. Training these models is notoriously expensive due to the need for specialized HPC clusters. In this work, we consider alternative setups for training large models: using cheap “preemptible” instances or pooling existing resources from multiple regions. We analyze the performance of existing model-parallel algorithms in these conditions and find configurations where training larger models becomes less communication-intensive. Based on these findings, we propose SWARM Parallelism (Stochastically Wired Adaptively Rebalanced Model Parallelism), a model-parallel training algorithm designed for poorly connected, heterogeneous and unreliable devices. SWARM creates temporary randomized pipelines between nodes that are rebalanced in case of failure. We empirically validate our findings and compare SWARM Parallelism with existing large-scale training approaches. Finally, we combine our insights with compression strategies to train a large Transformer language model with 1B shared parameters ($\approx$13B before sharing) on preemptible T4 GPUs with less than 200 Mb/s network.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/