Scaling FP8 training to trillion-token LLMs
Maxim FishmanBrian ChmielRon BannerDaniel Soudry
Establishes stable FP8 training for trillion-token large language models by identifying late-stage SwiGLU outlier instabilities and introducing Smooth-SwiGLU alongside FP8 Adam optimizer quantization to achieve a 34% throughput speedup without losing BF16-level accuracy.
Training large language models requires massive computing resources and energy, making efficiency a top priority. Using lower numerical precision, specifically 8-bit floating point (FP8) instead of traditional 16-bit formats, offers substantial compute and memory savings. However, previous evaluations were limited to shorter runs (up to 100 billion tokens), leaving the stability of FP8 unverified for production-scale training regimes that process trillions of tokens.
The article evaluates the scalability and stability of FP8 training on large language models trained on up to 2 trillion tokens. The authors investigate why standard FP8 training fails during extended runs and propose methods to achieve stable, full-scale low-precision training without losing accuracy.
The authors conducted empirical training runs and mathematical analyses using a 7-billion-parameter Llama 2 model trained on the RedPajama dataset across 256 Intel Gaudi2 accelerators over 15 days, as well as tests on Nvidia GPUs. They analyzed internal network activations, traced numerical instabilities, and designed two architectural enhancements: Smooth-SwiGLU, an activation mechanism with per-channel scaling, and an optimizer scheme that quantizes both moments of the Adam optimizer into 8-bit formats (E4M3 and E5M2).
The evaluation yielded several critical findings. First, standard FP8 training suffers severe divergence after extended training (around 200 billion tokens) due to activation outliers amplified by the SwiGLU activation function as its internal weights align over time. Second, Smooth-SwiGLU completely eliminates these instabilities while maintaining mathematical equivalence to standard SwiGLU, imposing zero computational overhead during inference. Third, the authors demonstrated the first successful 8-bit quantization of both Adam optimizer tracking states, reducing optimizer memory consumption by approximately 30%. Finally, combining Smooth-SwiGLU with the 8-bit optimizer achieved downstream task accuracy and perplexity on par with standard 16-bit baselines while delivering an approximate 34% increase in training throughput.
These results demonstrate that large-scale AI models can be trained entirely in 8-bit precision without sacrificing stability or output quality. For organizations developing frontier models, adopting this methodology provides substantial financial and operational benefits by reducing training hardware requirements, cutting energy footprints, and accelerating development timelines by roughly one-third.
Engineering and infrastructure teams should adopt Smooth-SwiGLU and dual-moment 8-bit Adam optimizers when deploying FP8 training pipelines on compatible hardware (such as Intel Gaudi2 or modern GPUs). Transitioning to this scheme requires no changes to final inference deployment, as the scaling factors can be directly folded back into model weights.
Confidence in these findings is high for models using GLU-style activations up to the 7-billion-parameter scale across 2 trillion tokens. Organizations planning to train models at significantly larger parameter scales (e.g., 70B+ parameters) should validate the approach on smaller pilot runs to verify that additional architectural scaling effects do not introduce new numerical anomalies.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). Introduces activation-to-weight scaling transformations to mitigate large activation outliers in LLMs, providing the foundational mathematical technique that Smooth-SwiGLU adapts for stable FP8 pre-training.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). Identifies the emergence of severe activation outliers in large transformer models and establishes the need for specialized treatment of outlier dimensions when reducing precision.
- Paper: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision, Jay Shah et al. (2024). Establishes techniques for hardware-accelerated FP8 training and execution in transformer architectures while addressing low-precision quantization errors.
- Paper: Adafactor: Adaptive Learning Rates with Sublinear Memory Cost, Noam Shazeer et al. (2018). Demonstrates methods for reducing Adam optimizer state memory overhead during large model training, which informs the low-precision dual-moment optimizer quantization used here.
- Paper: Mixed Precision Training, Paulius Micikevicius et al. (2018). Provides the foundational framework for mixed-precision neural network training and master-weight tracking that modern 8-bit training pipelines build upon.
No sufficiently relevant recommendations were found.
