Scaling FP8 training to trillion-token LLMs

Maxim FishmanBrian ChmielRon BannerDaniel Soudry

article2025ICLR74 citations

Establishes stable FP8 training for trillion-token large language models by identifying late-stage SwiGLU outlier instabilities and introducing Smooth-SwiGLU alongside FP8 Adam optimizer quantization to achieve a 34% throughput speedup without losing BF16-level accuracy.

Listen

Training large language models requires massive computing resources and energy, making efficiency a top priority. Using lower numerical precision, specifically 8-bit floating point (FP8) instead of traditional 16-bit formats, offers substantial compute and memory savings. However, previous evaluations were limited to shorter runs (up to 100 billion tokens), leaving the stability of FP8 unverified for production-scale training regimes that process trillions of tokens.

The article evaluates the scalability and stability of FP8 training on large language models trained on up to 2 trillion tokens. The authors investigate why standard FP8 training fails during extended runs and propose methods to achieve stable, full-scale low-precision training without losing accuracy.

The authors conducted empirical training runs and mathematical analyses using a 7-billion-parameter Llama 2 model trained on the RedPajama dataset across 256 Intel Gaudi2 accelerators over 15 days, as well as tests on Nvidia GPUs. They analyzed internal network activations, traced numerical instabilities, and designed two architectural enhancements: Smooth-SwiGLU, an activation mechanism with per-channel scaling, and an optimizer scheme that quantizes both moments of the Adam optimizer into 8-bit formats (E4M3 and E5M2).

The evaluation yielded several critical findings. First, standard FP8 training suffers severe divergence after extended training (around 200 billion tokens) due to activation outliers amplified by the SwiGLU activation function as its internal weights align over time. Second, Smooth-SwiGLU completely eliminates these instabilities while maintaining mathematical equivalence to standard SwiGLU, imposing zero computational overhead during inference. Third, the authors demonstrated the first successful 8-bit quantization of both Adam optimizer tracking states, reducing optimizer memory consumption by approximately 30%. Finally, combining Smooth-SwiGLU with the 8-bit optimizer achieved downstream task accuracy and perplexity on par with standard 16-bit baselines while delivering an approximate 34% increase in training throughput.

These results demonstrate that large-scale AI models can be trained entirely in 8-bit precision without sacrificing stability or output quality. For organizations developing frontier models, adopting this methodology provides substantial financial and operational benefits by reducing training hardware requirements, cutting energy footprints, and accelerating development timelines by roughly one-third.

Engineering and infrastructure teams should adopt Smooth-SwiGLU and dual-moment 8-bit Adam optimizers when deploying FP8 training pipelines on compatible hardware (such as Intel Gaudi2 or modern GPUs). Transitioning to this scheme requires no changes to final inference deployment, as the scaling factors can be directly folded back into model weights.

Confidence in these findings is high for models using GLU-style activations up to the 7-billion-parameter scale across 2 trillion tokens. Organizations planning to train models at significantly larger parameter scales (e.g., 70B+ parameters) should validate the approach on smaller pilot runs to verify that additional architectural scaling effects do not introduce new numerical anomalies.

No sufficiently relevant recommendations were found.

Cover for Scaling FP8 training to trillion-token LLMs

Abstract

We train, for the first time, large language models using FP8 precision on datasets up to 2 trillion tokens -- a 20-fold increase over previous limits. Through these extended training runs, we uncover critical instabilities in FP8 training that were not observable in earlier works with shorter durations. We trace these instabilities to outlier amplification by the SwiGLU activation function. Interestingly, we show, both analytically and empirically, that this amplification happens only over prolonged training periods, and link it to a SwiGLU weight alignment process. To address this newly identified issue, we introduce Smooth-SwiGLU, a novel modification that ensures stable FP8 training without altering function behavior. We also demonstrate, for the first time, FP8 quantization of both Adam optimizer moments. Combining these innovations, we successfully train a 7B parameter model using FP8 precision on 256 Intel Gaudi2 accelerators, achieving on-par results with the BF16 baseline while delivering up to a ∼34%\sim 34 \% throughput improvement. A reference implementation is supplied in this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background and Challenges of FP8 Training in LLMs
  • 3 Outlier Amplification in Large-Scale FP8 Training
  • 4 SwiGLU and Outlier Amplification
  • 4.1 SwiGLU Structure
  • 4.2 Theoretical analysis of weight correlation in SwiGLU
  • 4.3 Observing weight correlation growth and training instability
  • 4.4 Smooth-SwiGlu
  • 5 FP8 Optimizer
  • 5.1 Challenges
  • 5.2 Methodology
  • 6 Experiments
  • 6.1 Experimental Setup
  • 6.2 Results
  • 7 Conclusions
  • References
  • A Appendix
  • A.1 Training instability - additional data
  • A.2 Performance gain on Nvidia GPUs
  • A.3 Smooth-SwiGLU study
  • A.4 FP8 without SwiGLU activation function

Knowls

  1. Knowl 1 — Weight Alignment and Outlier Amplification in SwiGLU under L2 Regularization

    theoretical result

    Let a SwiGLU neuron with input x∈Rdx \in \mathbb{R}^d and weight vectors w1,w2∈Rdw_1, w_2 \in \mathbb{R}^d be defined as:

    SwiGLUw1,w2(x)=(x⊤w1)Swish(x⊤w2)=(x⊤w1)(x⊤w2)σ(x⊤w2)\text{SwiGLU}_{w_1, w_2}(x) = (x^\top w_1) \text{Swish}(x^\top w_2) = (x^\top w_1)(x^\top w_2)\sigma(x^\top w_2)

    where σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the logistic sigmoid function. Suppose this neuron is embedded in a neural network parameterized by (w1,w2,θ)(w_1, w_2, \theta) with θ∈Rk−2d\theta \in \mathbb{R}^{k-2d}, trained on NN samples with ℓ2\ell_2 regularization:

    min⁡w1,w2,θ∑n=1Nℓn(SwiGLUw1,w2(xn(θ)),θ)+μ2(∥w1∥22+∥w2∥22+∥θ∥22)\min_{w_1, w_2, \theta} \sum_{n=1}^N \ell_n(\text{SwiGLU}_{w_1, w_2}(x_n(\theta)), \theta) + \frac{\mu}{2} \left( \|w_1\|_2^2 + \|w_2\|_2^2 + \|\theta\|_2^2 \right)

    where μ>0\mu > 0 denotes regularization strength and ℓn\ell_n is the per-sample loss. If optimization converges to a stationary point (w1,w2,θ)(w_1, w_2, \theta) where the derivative of the sigmoid satisfies σ′(xn(θ)⊤w2)→0\sigma'(x_n(\theta)^\top w_2) \to 0 for all samples n∈{1,…,N}n \in \{1, \dots, N\}, then at this stationary point:

    w1→w2orw1→−w2w_1 \to w_2 \quad \text{or} \quad w_1 \to -w_2

    Because σ′(z)\sigma'(z) decays exponentially fast as ∣z∣|z| increases, this condition holds whenever ∥w2∥2\|w_2\|_2 is sufficiently large and ∣w2⊤xn∣>0|w_2^\top x_n| > 0. As a result of this collinear alignment, the SwiGLU neuron exhibits quadratic growth with respect to input scaling along the aligned direction (i.e., lim⁡c→∞SwiGLUw1,w2(cx)/c2=1\lim_{c \to \infty} \text{SwiGLU}_{w_1, w_2}(cx)/c^2 = 1 when w1=w2w_1 = w_2 and w1⊤x=1w_1^\top x = 1), causing extreme activation outliers over prolonged training runs. This alignment property holds identically for other Gated Linear Unit (GLU) variants.

  2. Knowl 2 — Smooth-SwiGLU Activation Function

    model/method

    Smooth-SwiGLU is a per-channel rescaled variant of the SwiGLU activation designed to eliminate extreme activation outliers entering the subsequent linear layer during low-precision (FP8) training without changing the mathematical output during inference.

    For channel ii, with quantized input Q(x)Q(x), quantized weights w^1,i=Q(w1,i)\hat{w}_{1,i} = Q(w_{1,i}) and w^2,i=Q(w2,i)\hat{w}_{2,i} = Q(w_{2,i}), and a quantization function QQ, the quantized Smooth-SwiGLU activation is computed as:

    Smooth-SwiGLUw^1,i,w^2,i(x)=si−1⋅Q(si⋅(w^1,i⊤Q(x))Swish(w^2,i⊤Q(x)))\text{Smooth-SwiGLU}_{\hat{w}_{1,i}, \hat{w}_{2,i}}(x) = s_i^{-1} \cdot Q\left( s_i \cdot (\hat{w}_{1,i}^\top Q(x)) \text{Swish}(\hat{w}_{2,i}^\top Q(x)) \right)

    where sis_i is a per-channel scaling factor computed by splitting the intermediate activation tensor into channel chunks, calculating the maximum absolute value per chunk in parallel, and deriving sis_i to scale down the linear branch.

    When followed by a linear projection layer with weight w^3,i\hat{w}_{3,i}, the complete MLP block output is:

    ∑iw^3,iSmooth-SwiGLUw^1,i,w^2,i(x)=∑isi−1w^3,iQ(si(w^1,i⊤Q(x))Swish(w^2,i⊤Q(x)))\sum_i \hat{w}_{3,i} \text{Smooth-SwiGLU}_{\hat{w}_{1,i}, \hat{w}_{2,i}}(x) = \sum_i s_i^{-1} \hat{w}_{3,i} Q\left( s_i (\hat{w}_{1,i}^\top Q(x)) \text{Swish}(\hat{w}_{2,i}^\top Q(x)) \right)

    During inference, the per-channel scaling factors are folded directly into the static weights by setting w~1,i=Q(si⋅w1,i)\tilde{w}_{1,i} = Q(s_i \cdot w_{1,i}) and w~3,i=Q(si−1⋅w3,i)\tilde{w}_{3,i} = Q(s_i^{-1} \cdot w_{3,i}), introducing zero computational overhead at inference time.

  3. Knowl 3 — Dual-Moment FP8 Quantization for Adam Optimizer

    model/method

    An FP8 quantization scheme for the Adam optimizer quantizes both optimizer moment state buffers into 8-bit floating-point representations:

    1. First Moment (Gradient Mean Estimate): Stored in the E4M3 format (4 exponent bits, 3 mantissa bits). E4M3 provides higher precision and lower quantization noise, which is required to accurately track directional gradient averages.

    2. Second Moment (Uncentered Gradient Variance Estimate): Stored in the E5M2 format (5 exponent bits, 2 mantissa bits). Because Adam parameter updates compute the inverse square root of the second moment (1vt+ϵ\frac{1}{\sqrt{v_t} + \epsilon}), the smallest values of vtv_t dominate update step sizes. E5M2 provides the wider dynamic range necessary to represent very small values without catastrophic underflow.

  4. Knowl 4 — Long-Horizon FP8 Training Instability in SwiGLU LLMs

    empirical result

    When training Llama-2 7B in standard FP8 precision (E4M3 for forward weights and activations, E5M2 for backward gradients, delayed per-tensor scaling, and FP32 Adam moments), training loss diverges catastrophically after approximately 200B to 220B tokens, despite tracking the BF16 baseline loss curve closely before that point.

    This instability is driven by SwiGLU activation spikes. Tracking individual outlier-generating MLP channels shows that the weight norm ∥w2∥2\|w_2\|_2 grows over training and the correlation between weight vectors w1w_1 and w2w_2 rises sharply between 125B and 210B tokens (transitioning from near-zero correlation at 8B tokens to correlation coefficients exceeding 0.80.8 at 330B tokens). These aligned weights amplify sporadic activation spikes that exceed the dynamic range of FP8 under delayed scaling.

    Disabling quantization only at the SwiGLU output (the input to linear layer w3w_3) resolves the divergence over 2 trillion tokens. Furthermore, training a GPT-3 125M model with GeLU activations in FP8 under identical conditions shows stable convergence beyond 300B tokens, confirming that the late-emerging instability is specific to GLU-based activation functions.

  5. Knowl 5 — Format Sensitivity of Adam Moment Quantization in FP8 Training

    empirical result

    In 20-billion-token training experiments on Llama-2 100M evaluating all four combinations of standard FP8 formats (E4M3 and E5M2) for Adam optimizer moments against a BF16 baseline:

    • First Moment E4M3, Second Moment E5M2: The only combination that successfully converges and matches the BF16 baseline loss trajectory.
    • First Moment E5M2, Second Moment E5M2: Fails to match baseline loss and exhibits severe performance degradation due to insufficient mantissa precision in the first moment.
    • First Moment E4M3, Second Moment E4M3: Fails to converge and suffers loss divergence due to insufficient dynamic range for small second-moment values.
    • First Moment E5M2, Second Moment E4M3: Rapidly diverges for the same reason.
  6. Knowl 6 — Zero-Shot Benchmark Performance of 2-Trillion-Token FP8 Llama-2 7B

    data/table

    Zero-shot downstream task evaluations for Llama-2 7B trained on 2 trillion tokens of the RedPajama dataset on 256 Intel Gaudi2 accelerators demonstrate that the fully quantized FP8 training pipeline matches BF16 baseline performance across accuracy and perplexity metrics.

    Precision Configuration Lambada Acc (%) ↑\uparrow HellaSwag Acc (%) ↑\uparrow Winogrande Acc (%) ↑\uparrow Arc-C Acc (%) ↑\uparrow Wikitext PPL ↓\downarrow Lambada PPL ↓\downarrow
    BF16 (Baseline) 61.98 68.30 64.25 37.37 5.59 5.75
    FP8 (SwiGLU output in BF16) 61.73 68.03 64.40 37.20 5.56 5.89
    FP8 (Smooth-SwiGLU + FP8 Optimizer) 62.10 68.37 65.43 37.80 5.55 5.90

    The combined scheme (FP8 forward/backward, Smooth-SwiGLU, and dual FP8 Adam moments) achieves downstream accuracy on par with or slightly exceeding the BF16 baseline while training stably over the entire 2-trillion-token budget.

  7. Knowl 7 — Hardware Speedup and Compute Acceleration of FP8 Training with Smooth-SwiGLU

    data/table

    Throughput and compute efficiency benchmarks for Llama-2 7B (micro-batch size 1) measured on 8 Intel Gaudi2 accelerators and 8 NVIDIA RTX A6000 Ada GPUs:

    Platform and Configuration Status Throughput (samples/s) Relative Speedup TFLOPS
    8 ×\times Intel Gaudi2
    BF16 Converges 12.65 Baseline 311
    FP8 + SwiGLU output in BF16 Converges 16.07 +27.04% 396
    FP8 + Smooth-SwiGLU Converges 16.89 +33.52% 417
    Standard FP8 Diverges 17.34 +37.08% 428
    8 ×\times NVIDIA A6000 Ada
    BF16 Converges 3.22 Baseline 76.0
    FP8 + SwiGLU output in BF16 Converges 4.11 +27.60% 96.9
    FP8 + Smooth-SwiGLU Converges 4.32 +34.16% 101.9
    Standard FP8 Diverges 4.43 +37.58% 104.5

    Standard FP8 achieves the highest raw throughput (+37.08% / +37.58%) but fails due to divergence. Smooth-SwiGLU preserves FP8 quantization across the entire MLP block, recovering a ~34% throughput speedup over BF16 while ensuring numerical stability.

  8. Knowl 8 — Optimizer Memory Footprint Reduction with FP8 Dual-Moment Adam

    data/table

    Memory footprint per accelerator during Llama-2 7B training on 8 Intel Gaudi2 devices using DeepSpeed ZeRO-1 (with FP16 master weights):

    Configuration Status Memory (GB/HPU) FP8 Optimizer Used
    BF16 Baseline Converges 63.25 No (FP32 moments)
    FP8 + SwiGLU output in BF16 Converges 63.26 No (FP32 moments)
    FP8 + Smooth-SwiGLU Converges 63.26 No (FP32 moments)
    Standard FP8 Diverges 63.24 No (FP32 moments)
    FP8 + SwiGLU output in BF16 Converges 44.08 Yes (Dual FP8 moments)
    FP8 + Smooth-SwiGLU Converges 44.08 Yes (Dual FP8 moments)
    Standard FP8 Diverges 44.09 Yes (Dual FP8 moments)

    Transitioning both Adam optimizer moments from FP32 to standard FP8 (E4M3 for moment 1, E5M2 for moment 2) reduces memory consumption from ~63.26 GB/HPU to ~44.08 GB/HPU, yielding a ~30.3% reduction in total per-device memory consumption.

  9. Knowl 9 — Optimization Smoothing and Loss Improvement in BF16 Training Using Smooth-SwiGLU

    empirical result

    Applying Smooth-SwiGLU to standard BF16 training of a Llama 700M model across different peak learning rates (2.5×10−42.5 \times 10^{-4}, 7.5×10−47.5 \times 10^{-4}, and 1.5×10−31.5 \times 10^{-3}) under a cosine schedule produces smoother loss curves and enables convergence to lower final training loss values.

    At higher learning rates where standard SwiGLU experiences significant loss oscillations and degradation, Smooth-SwiGLU dampens activation-driven variance and allows stable optimization, demonstrating that the benefits of Smooth-SwiGLU extend beyond low-precision quantization to general Transformer optimization.

Coverage note — None was omitted; all primary contributions—including the weight alignment theorem, Smooth-SwiGLU architecture, dual FP8 Adam optimizer scheme, 2T-token stability results, throughput/memory benchmarks, and ablations—are fully represented.

References

  1. 1.Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. Quantizable transformers: Removing outliers by helping attention heads do nothing. ArXiv, abs/2306.12929, 2023. URL https://api.semanticscholar.org/CorpusID:259224568.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. J. Mach. Learn. Res., 24:240:1–240:113, 2022. URL https://api.semanticscholar.org/CorpusID:247951931.
  4. 4.Together Computer. Redpajama: an open dataset for training large language models, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  5. 5.Joonhyung Lee, Jeongin Bae, Byeongwook Kim, Se Jung Kwon, and Dongsoo Lee. To fp8 and back again: Quantifying the effects of reducing precision on llm training stability. arXiv preprint arXiv:2405.18710, 2024.
  6. 6.Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant: Llm quantization with learned rotations. ArXiv, abs/2405.16406, 2024. URL https://api.semanticscholar.org/CorpusID:270062819.
  7. 7.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Frederick Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed precision training. ArXiv, abs/1710.03740, 2017. URL https://api.semanticscholar.org/CorpusID:3297437.
  8. 8.Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep K. Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart F. Oberman, Mohammad Shoeybi, Michael Siu, and Hao Wu. Fp8 formats for deep learning. ArXiv, abs/2209.05433, 2022. URL https://api.semanticscholar.org/CorpusID:252198916.
  9. 9.Houwen Peng, Kan Wu, Yixuan Wei, Guoshuai Zhao, Yuxiang Yang, Ze Liu, Yifan Xiong, Ziyue Yang, Bolin Ni, Jingcheng Hu, Ruihang Li, Miaosen Zhang, Chen Li, Jia Ning, Ruizhe Wang, Zheng Zhang, Shuguang Liu, Joe Chau, Han Hu, and Peng Cheng. Fp8-lm: Training fp8 large language models. ArXiv, abs/2310.18313, 2023. URL https://api.semanticscholar.org/CorpusID:264555252.
  10. 10.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurenccon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa Etxabe, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris C. Emezue, Christopher Klamm, Colin Leong, Daniel Alexander van Strien, David Ifeoluwa Adelani, Dragomir R. Radev, Eduardo Gonz’alez Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady ElSahar, Hamza Benyamina, Hieu Trung Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jorg Frohberg, Josephine Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro von Werra, Leon Weber, Long Phan, Loubna Ben Allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, Mar’ia Grandury, Mario vSavsko, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto L’opez, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, S. Longpre, Somaieh Nikpoor, S. Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-Shaibani, Matteo Manica, Nihal V. Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Févry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiang Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Y Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre Franccois Lavall’ee, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aur’elie N’ev’eol, Charles Lovering, Daniel H Garrette, Deepak R. Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Xiangru Tang, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, S. Osher Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdenvek Kasner, Zdenek Kasner, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ananda Santa Rosa Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ayoade Ajibade, Bharat Kumar Saxena, Carlos Muñoz Ferrandis, Danish Contractor, David M. Lansky, Davis David, Douwe Kiela, Duong Anh Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatim Tahirah Mirza, Frankline Ononiwu, Habib Rezanejad, H.A. Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jan Passmore, Joshua Seltzer, Julio Bonis Sanz, Karen Fort, Lívia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nourhan Fahmy, Olanrewaju Samuel, Ran An, R. P. Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas L. Wang, Sourav Roy, Sylvain Viguier, Thanh-Cong Le, Tobi Oyebade, Trieu Nguyen Hai Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Kumar Singh, Benjamin Beilharz, Bo Wang, Caio Matheus Fonseca de Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel Le’on Perin’an, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Iman I.B. Bello, Isha Dash, Ji Soo Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthi Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, María Andrea Castillo, Marianna Nezhurina, Mario Sanger, Matthias Samwald, Michael Cullan, Michael Weinberg, M Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patricia Haller, Patrick Haller, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sangaroonsiri, Srishti Kumar, Stefan Schweter, Sushil Pratap Bharati, Tanmay Laud, Théo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yashasvi Bajaj, Y. Venkatraman, Yifan Xu, Ying Xu, Yu Xu, Zhee Xao Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, and Thomas Wolf. Bloom: A 176b-parameter open-access multilingual language model. ArXiv, abs/2211.05100, 2022. URL https://api.semanticscholar.org/CorpusID:253420279.
  11. 11.Noam M. Shazeer. Glu variants improve transformer. ArXiv, abs/2002.05202, 2020a. URL https://api.semanticscholar.org/CorpusID:211096588.
  12. 12.Noam M. Shazeer. Glu variants improve transformer. ArXiv, abs/2002.05202, 2020b. URL https://api.semanticscholar.org/CorpusID:211096588.
  13. 13.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Anand Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. ArXiv, abs/2201.11990, 2022. URL https://api.semanticscholar.org/CorpusID:246411325.
  14. 14.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomput., 568(C), mar 2024. ISSN 0925-2312. doi: 10.1016/j.neucom.2023.127063. URL https://doi.org/10.1016/j.neucom.2023.127063.
  15. 15.Mingjie Sun, Xinlei Chen, J. Zico Kolter, and Zhuang Liu. Massive activations in large language models. ArXiv, abs/2402.17762, 2024. URL https://api.semanticscholar.org/CorpusID:268041240.
  16. 16.Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Xiaodong Cui, Wei Zhang, and K. Gopalakrishnan. Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks. In Neural Information Processing Systems, 2019. URL https://api.semanticscholar.org/CorpusID:202779157.
  17. 17.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288, 2023. URL https://api.semanticscholar.org/CorpusID:259950998.
  18. 18.Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Information Processing Systems, 2017. URL https://api.semanticscholar.org/CorpusID:13756489.
  19. 19.Jaewoo Yang, Hayun Kim, and Younghoon Kim. Mitigating quantization errors due to activation spikes in glu-based llms. ArXiv, abs/2405.14428, 2024. URL https://api.semanticscholar.org/CorpusID:269983752.
  20. 20.Biao Zhang and Rico Sennrich. Root mean square layer normalization. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper_files/paper/2019/file/1e8a19426224ca89e83cef47f1e7f53b-Paper.pdf.

Citation

MLA
Fishman, M., et al. “Scaling FP8 Training to Trillion-token LLMs”. arXiv, 2024, http://arxiv.org/abs/2409.12517v2.
APA
Fishman, M., Chmiel, B., Banner, R., & Soudry, D. (2024). Scaling FP8 training to trillion-token LLMs. arXiv. http://arxiv.org/abs/2409.12517v2
Chicago
Fishman, M., B. Chmiel, R. Banner, and D. Soudry. 2024. “Scaling FP8 Training to Trillion-token LLMs”. arXiv. http://arxiv.org/abs/2409.12517v2.
Harvard
Fishman, M. et al. (2024) “Scaling FP8 training to trillion-token LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2409.12517v2.
Vancouver
1. Fishman M, Chmiel B, Banner R, Soudry D (2024) Scaling FP8 training to trillion-token LLMs. arXiv

BibTeX

@article{fishman2024scaling,
  title = {Scaling FP8 training to trillion-token LLMs},
  author = {Fishman, Maxim and Chmiel, Brian and Banner, Ron and Soudry, Daniel},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2409.12517v2},
  eprint = {2409.12517}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors