Optimizing Large Language Model Training Using FP4 Quantization
Ruizhe WangYeyun GongXiao LiuGuoshuai ZhaoZiyue YangBaining GuoZheng-Jun ZhaPeng Cheng
Presents the first FP4 training framework for large language models, employing a differentiable quantization estimator and outlier compensation to match standard BF16 and FP8 accuracy on models scaled up to 13 billion parameters.
Training state-of-the-art large language models requires enormous computing power, financial investment, and energy consumption. As models continue to scale to hundreds of billions of parameters, reducing these computational demands has become an urgent industry priority. Reducing numerical precision—using fewer bits to represent numbers in memory and calculations—is a primary path to lower costs. While 8-bit floating-point training is now feasible, pushing down to 4-bit floating point has remained a major barrier because 4-bit numbers have severely limited dynamic range and representational capacity, often causing numerical instability and severe accuracy loss.
The article demonstrates the first end-to-end 4-bit floating-point pretraining framework designed specifically for large language models. Its main objective is to evaluate whether 4-bit floating-point operations can train language models from scratch while matching the accuracy and stability of standard 16-bit and 8-bit precision baselines.
The researchers developed two core algorithmic techniques to address the errors that arise during 4-bit training. For model weights, they introduced a differentiable gradient estimator that corrects gradient errors during backpropagation. For activations, which are prone to extreme outlier values that collapse 4-bit representations, they developed a dynamic clamping and sparse compensation strategy to retain accuracy. The framework was evaluated by pretraining standard LLaMA-architecture models ranging from 1.3 billion to 13 billion parameters from scratch on up to 100 billion text tokens. Because dedicated 4-bit hardware is not yet broadly available, computations were emulated on existing 8-bit GPU tensor cores, and the trained models were tested across eight standard downstream evaluation benchmarks.
The evaluation yielded several key findings. First, the 4-bit framework achieved pretraining loss closely tracking standard 16-bit models across all tested sizes (for instance, achieving a final loss of 1.97 versus 1.88 for a 13-billion-parameter model after 100 billion tokens). Second, the 4-bit models matched or slightly exceeded 16-bit baselines in zero-shot task accuracy, averaging 54.95% versus 54.44% on the 13-billion model. Third, ablation experiments revealed that activations are far more sensitive to quantization than weights, showing that uncompensated 4-bit activations cause training to diverge completely. Finally, theoretical calculations indicate that the proposed framework delivers an approximate 2.95x computational speedup per Transformer layer after accounting for algorithmic overheads.
These findings prove that 4-bit training is technically viable without sacrificing final model intelligence or convergence stability. For organizations developing foundational AI, moving to 4-bit precision offers a concrete path to substantially cut compute infrastructure costs, shorten pretraining timelines, and reduce data center energy footprints. The results also show that fine-grained vector scaling and dynamic outlier management are essential prerequisites for ultra-low-precision computing.
Technical leaders and infrastructure planners should prepare software and hardware deployment roadmaps to support ultra-low precision as next-generation AI accelerators with native 4-bit support enter the market. Engineering teams should explore the released open-source framework and benchmark its implementation against internal training workloads. Before committing full-scale production budgets to 4-bit training, organizations should run pilot pretraining runs on larger models (such as 70 billion parameters or larger) and longer token horizons to confirm that scaling behaviors hold across massive datasets.
Readers should note two main limitations in this work. First, the experimental runs relied on 8-bit hardware emulation, meaning real-world wall-clock runtime speedups and physical power reductions have not yet been directly measured on physical 4-bit silicon. Second, testing was bounded at 13-billion parameters and 100 billion tokens, leaving some uncertainty regarding stability when scaling to trillions of tokens. Nevertheless, confidence in the numerical stability and mathematical feasibility of the method remains high across the evaluated operational ranges.
- Paper: LLM-FP4: 4-Bit Floating-Point Quantized Transformers, Shih-Yang Liu et al. (2023). Read this earlier FP4 transformer work first to see the representational and outlier challenges that the source addresses in the harder setting of training from scratch.
- Paper: Mixed Precision Training, Paulius Micikevicius et al. (2018). Its mixed-precision training techniques establish how reduced-precision computation can preserve learning signals, providing essential context for the source’s move to FP4.
- Paper: DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients, Shuchang Zhou et al. (2016). DoReFa-Net introduces low-bit gradient training and gradient estimators, making it useful preparation for the source’s correction of quantized backpropagation errors.
- Paper: Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations, Itay Hubara et al. (2016). This foundational study shows how low-precision weights and activations can be trained with surrogate gradients, concepts the source adapts to ultra-low-precision LLM pretraining.
- Paper: Deep Learning with Limited Numerical Precision, Suyog Gupta et al. (2015). Its demonstration that stochastic rounding can stabilize low-precision training provides useful groundwork for understanding the numerical-instability problem tackled by the source.
No sufficiently relevant recommendations were found.
