Optimizing Large Language Model Training Using FP4 Quantization

Ruizhe WangYeyun GongXiao LiuGuoshuai ZhaoZiyue YangBaining GuoZheng-Jun ZhaPeng Cheng

article2025ICML61 citations

Presents the first FP4 training framework for large language models, employing a differentiable quantization estimator and outlier compensation to match standard BF16 and FP8 accuracy on models scaled up to 13 billion parameters.

Listen

Training state-of-the-art large language models requires enormous computing power, financial investment, and energy consumption. As models continue to scale to hundreds of billions of parameters, reducing these computational demands has become an urgent industry priority. Reducing numerical precision—using fewer bits to represent numbers in memory and calculations—is a primary path to lower costs. While 8-bit floating-point training is now feasible, pushing down to 4-bit floating point has remained a major barrier because 4-bit numbers have severely limited dynamic range and representational capacity, often causing numerical instability and severe accuracy loss.

The article demonstrates the first end-to-end 4-bit floating-point pretraining framework designed specifically for large language models. Its main objective is to evaluate whether 4-bit floating-point operations can train language models from scratch while matching the accuracy and stability of standard 16-bit and 8-bit precision baselines.

The researchers developed two core algorithmic techniques to address the errors that arise during 4-bit training. For model weights, they introduced a differentiable gradient estimator that corrects gradient errors during backpropagation. For activations, which are prone to extreme outlier values that collapse 4-bit representations, they developed a dynamic clamping and sparse compensation strategy to retain accuracy. The framework was evaluated by pretraining standard LLaMA-architecture models ranging from 1.3 billion to 13 billion parameters from scratch on up to 100 billion text tokens. Because dedicated 4-bit hardware is not yet broadly available, computations were emulated on existing 8-bit GPU tensor cores, and the trained models were tested across eight standard downstream evaluation benchmarks.

The evaluation yielded several key findings. First, the 4-bit framework achieved pretraining loss closely tracking standard 16-bit models across all tested sizes (for instance, achieving a final loss of 1.97 versus 1.88 for a 13-billion-parameter model after 100 billion tokens). Second, the 4-bit models matched or slightly exceeded 16-bit baselines in zero-shot task accuracy, averaging 54.95% versus 54.44% on the 13-billion model. Third, ablation experiments revealed that activations are far more sensitive to quantization than weights, showing that uncompensated 4-bit activations cause training to diverge completely. Finally, theoretical calculations indicate that the proposed framework delivers an approximate 2.95x computational speedup per Transformer layer after accounting for algorithmic overheads.

These findings prove that 4-bit training is technically viable without sacrificing final model intelligence or convergence stability. For organizations developing foundational AI, moving to 4-bit precision offers a concrete path to substantially cut compute infrastructure costs, shorten pretraining timelines, and reduce data center energy footprints. The results also show that fine-grained vector scaling and dynamic outlier management are essential prerequisites for ultra-low-precision computing.

Technical leaders and infrastructure planners should prepare software and hardware deployment roadmaps to support ultra-low precision as next-generation AI accelerators with native 4-bit support enter the market. Engineering teams should explore the released open-source framework and benchmark its implementation against internal training workloads. Before committing full-scale production budgets to 4-bit training, organizations should run pilot pretraining runs on larger models (such as 70 billion parameters or larger) and longer token horizons to confirm that scaling behaviors hold across massive datasets.

Readers should note two main limitations in this work. First, the experimental runs relied on 8-bit hardware emulation, meaning real-world wall-clock runtime speedups and physical power reductions have not yet been directly measured on physical 4-bit silicon. Second, testing was bounded at 13-billion parameters and 100 billion tokens, leaving some uncertainty regarding stability when scaling to trillions of tokens. Nevertheless, confidence in the numerical stability and mathematical feasibility of the method remains high across the evaluated operational ranges.

arXiv: 2501.17116

No sufficiently relevant recommendations were found.

Cover for Optimizing Large Language Model Training Using FP4 Quantization

Abstract

The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited representational capacity. This work introduces the first FP4 training framework for LLMs, addressing these challenges with two key innovations: a differentiable quantization estimator for precise weight updates and an outlier clamping and compensation strategy to prevent activation collapse. To ensure stability, the framework integrates a mixed-precision training scheme and vector-wise quantization. Experimental results demonstrate that our FP4 framework achieves accuracy comparable to BF16 and FP8, with minimal degradation, scaling effectively to 13B-parameter LLMs trained on up to 100B tokens. With the emergence of next-generation hardware supporting FP4, our framework sets a foundation for efficient ultra-low precision training.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Methodology
  • 3.1 Differentiable Gradient Estimator
  • 3.2 Outlier Clamping and Compensation
  • 4 Experiment
  • 4.1 Experiment Setup
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 5 Related Work
  • 6 Limitation
  • 7 Conclusion
  • References
  • A Implementation of FP4 Quantizaiton
  • B Theoretical Analysis on Speed Performance and Overhead
  • C Supplementary Proof for Differentiable Quantization Estimator
  • C.1 The Derivation of Differentiable Quantization Function
  • C.2 Proof of DGE with Vector-wise Scaling Factors
  • C.3 Mathematical Soundness of DGE Clipping
  • D Analyzing Quantization Difficulty Through Tensor Distribution

Knowls

  1. Knowl 1 — FP4 training uses E2M1 quantization in Transformer linear layers

    model/method

    The proposed training framework quantizes both activations and weights to the 4-bit E2M1 floating-point format for Transformer general matrix multiplication (GeMM). E2M1 has a maximum absolute representable value of 66 and the codebook values 0,±0.5,±1,±1.5,±2,±3,±4,±60, \pm0.5, \pm1, \pm1.5, \pm2, \pm3, \pm4, \pm6. For a high-precision tensor xx, absmax scaling and lookup-table quantization are defined by

    xFP4=Q(γx),γ=6max⁡i∣xi∣,x_{\mathrm{FP4}} = Q(\gamma x), \qquad \gamma = \frac{6}{\max_i |x_i|},

    where QQ maps each scaled element to a representable E2M1 value and ii indexes tensor elements. The framework retains the scale information so the GeMM result can be rescaled to preserve the intended computation. Direct FP4 casting is insufficient for stable, accurate training; the framework combines this forward quantization with a weight-gradient correction and activation outlier handling.

  2. Knowl 2 — Differentiable Gradient Estimator corrects FP4 weight gradients

    model/method

    The Differentiable Gradient Estimator (DGE) keeps hard FP4 quantization in the forward pass, but replaces the zero-gradient behavior of hard quantization in the weight backward pass with a derivative from a smooth approximation. For an input xx within a quantization interval of width δ>0\delta>0, the approximation is

    f(x)=δ2(1+sign⁡(2xδ−1)∣2xδ−1∣1/k),f′(x)=1k∣2xδ−1∣1/k−1,f(x)=\frac{\delta}{2}\left(1+\operatorname{sign}\left(\frac{2x}{\delta}-1\right)\left|\frac{2x}{\delta}-1\right|^{1/k}\right), \qquad f'(x)=\frac{1}{k}\left|\frac{2x}{\delta}-1\right|^{1/k-1},

    where k>1k>1 controls how sharply the curve approximates a hard quantization step. If WW is a weight matrix, WqW_q its hard-quantized version, and LL the training loss, DGE computes the weight gradient as

    ∂L∂W=∂L∂Wq⊙f′(W),\frac{\partial L}{\partial W}=\frac{\partial L}{\partial W_q}\odot f'(W),

    where ⊙\odot is element-wise multiplication and the correction is evaluated using the relevant quantization intervals. This provides a nonconstant gradient correction instead of the Straight-Through Estimator’s constant derivative of 11. To limit gradient spikes, the correction magnitude is capped at 3.03.0 in practice; the main experiments use k=5k=5.

  3. Knowl 3 — Outlier Clamping and Compensation preserves activation information

    model/method

    Outlier Clamping and Compensation (OCC) addresses activation outliers that enlarge a tensor’s quantization range and cause many ordinary values to collapse toward zero. For an activation tensor YY, OCC uses a chosen quantile α\alpha to set a clipping threshold, clamps outlying values to form YcY_c, and retains the residual ΔY=Y−Yc\Delta Y=Y-Y_c. The clamped tensor is processed with FP4 GeMM, while the sparse residual is processed using high-precision sparse matrix multiplication. In the experiments, the residual contains about 0.2%0.2\%–2%2\% nonzero elements for quantiles around 0.9990.999–0.990.99; lower quantiles retain more residual values and improve fidelity at additional computational cost.

    The following results average comparisons across activation tensors from 30,000 training iterations of a LLaMA 1.3B model. SIM is cosine similarity, MSE is mean squared error, and SNR is signal-to-noise ratio.

    Could not parse LaTeX table

    Clamping improves tensor fidelity over unclamped quantization, and adding the sparse residual compensation improves it further. Reducing the quantile from 99.999.9 to 9999 or 9797 further lowers MSE, but increases the amount of compensation computation.

  4. Knowl 4 — FP4 training approaches BF16 loss at model scales up to 13B

    empirical result

    LLaMA models with 1.3B, 7B, and 13B parameters were trained from scratch for 100B tokens with FP4 or BF16 using the same dataset and hyperparameters. Their training-loss curves largely overlap, although FP4 ends with a modestly higher loss in each case. The reported losses after 100B tokens are:

    Could not parse LaTeX table

    These results demonstrate that the proposed FP4 framework can pretrain models up to 13B parameters over 100B tokens with a small loss gap relative to BF16.

  5. Knowl 5 — Zero-shot and perplexity results are comparable between FP4 and BF16

    data/table

    The authors evaluated pretrained FP4 and BF16 models at three sizes on zero-shot accuracy tasks and language-model perplexity datasets. Zero-shot results are percentages; the reported average and individual task scores are reproduced below. FP4’s average is close to BF16 at every size, and is higher for the 7B and 13B models. Perplexity (PPL) results are also close, with FP4 lower on some datasets and BF16 lower on others; lower PPL is better.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The zero-shot tasks are PiQA, HellaSwag, OpenbookQA (ObQA), ARC Challenge (Arc-C), ARC Easy (Arc-E), BoolQ, LogiQA, SciQ, and Lambada. The perplexity datasets are Lambada OpenAI (Lbd.OAI), Lambada standard (Lbd.std), Pile 10k, and Wikitext. All evaluations compare models trained at the same size.

  6. Knowl 6 — Vector-wise scaling and mixed precision support stable FP4 computation

    model/method

    The FP4 framework uses scaling granularity aligned with linear-layer matrix multiplication: activations are quantized token-wise along the sequence dimension, and weights are quantized channel-wise along the output-channel dimension. The authors report that coarse tensor-wise scaling causes substantial FP4 accuracy degradation, with coarse activation scaling more damaging than coarse weight scaling.

    FP4 is used for GeMM, while other operations use higher precision. In the training implementation, gradient communication is performed in FP8; gradients and first-order Adam moments are stored in FP8, and second-order moments in FP16. Remaining operations are performed in FP16 or BF16. Because dedicated FP4 tensor cores were unavailable, the experiments simulated FP4 computation using FP8 tensor cores on NVIDIA H-series GPUs.

  7. Knowl 7 — Training setup and hyperparameters for the main comparisons

    experimental setup

    The main experiments pretrained LLaMA 2 models from scratch on the DCLM dataset, comparing FP4 with BF16 under matching hyperparameters. The evaluated model sizes were 1.3B, 7B, and 13B parameters, each trained for 100B tokens. Input sequences had 2,048 tokens and the batch size was 2,048, or approximately 4M tokens. The learning rate peaked at 3×10−43\times10^{-4}, used warm-up over 5% of total steps, and then cosine decay toward 10% of the peak; weight decay was 0.10.1. Adam used β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, and ϵ=1×10−8\epsilon=1\times10^{-8}. FP4-specific settings were DGE k=5k=5 and OCC quantile α=0.99\alpha=0.99. FP4 operations were emulated with FP8 tensor cores, so the experiments establish training accuracy rather than native FP4 hardware speed.

  8. Knowl 8 — Ablations identify the roles of DGE, OCC, and quantization granularity

    empirical result

    Ablations used a 1.3B LLaMA model trained for 10B tokens on a DCLM subset, with batch size 256 and other hyperparameters held consistent with the main experiments. Directly casting both weights and activations to FP4 produced a substantial loss gap, while FP4 with both DGE and OCC retained much better convergence; the FP8 baselines and the proposed FP4 method maintained pretraining accuracy. In weight-only FP4 experiments, DGE improved convergence over direct weight quantization, and k=5k=5 gave better final performance than k=3k=3 or k=10k=10. In activation-only FP4 experiments, direct activation quantization diverged to NaN, whereas OCC restored convergence. OCC quantiles 0.9990.999, 0.990.99, and 0.970.97 corresponded to approximately 0.2%0.2\%, 2%2\%, and 6%6\% nonzero residual elements; lower α\alpha improved accuracy but increased compensation cost, and the authors selected 0.990.99 as a trade-off. Finally, coarse tensor-wise FP4 scaling degraded training, especially when applied to activations alone, supporting the use of token-wise activation and channel-wise weight scaling.

  9. Knowl 9 — Theoretical FP4 speedup estimate includes DGE and OCC overhead

    theoretical result

    The paper estimates compute speedup for a Transformer layer assuming backward propagation costs approximately twice the forward-pass FLOPs. Let hh be hidden size, bb batch size, and ss sequence length. Under the layer FLOP accounting used in the paper, FP32 and FP4 totals are proportional to 24bsh2+5bs2h+36bsh24bsh^2+5bs^2h+36bsh and 6bsh2+5bs2h+36bsh6bsh^2+5bs^2h+36bsh, respectively. The ideal speedup, excluding overhead, is therefore

    24h+5s+366h+5s+36.\frac{24h+5s+36}{6h+5s+36}.

    For h=4096h=4096 and s=2048s=2048, this ideal estimate is 3.12×3.12\times. The estimate adds 96bsh96bsh FLOPs per iteration for DGE and 2(1−α)(12bsh2)2(1-\alpha)(12bsh^2) FLOPs for OCC sparse compensation, where α\alpha is the activation clipping quantile. With α=0.99\alpha=0.99, the adjusted estimate is

    24h+5s+366h+24(1−α)h+5s+68=2.95×\frac{24h+5s+36}{6h+24(1-\alpha)h+5s+68}=2.95\times

    for the same hh and ss. This is a theoretical estimate, not a measured hardware speedup; the authors note that sparse-kernel hardware inefficiency can affect runtime.

  10. Knowl 10 — Native FP4 speed and extreme-scale training remain unvalidated

    limitation

    The experiments did not run on dedicated FP4 tensor cores, which were unavailable to the authors. FP4 was simulated using FP8 tensor cores with additional precision casting, so the study could not directly measure the speedup or energy-efficiency gains of native FP4 hardware; the simulation also increased runtime. The experiments reached 13B parameters and 100B training tokens, but resource constraints prevented evaluation of substantially larger models or datasets containing trillions of tokens.

Coverage note — No substantial contributed material was omitted; the appendix’s raw tensor-distribution plots and CUDA lookup-kernel listing are supporting diagnostics and implementation detail rather than additional standalone findings.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Ashkboos, S., Mohtashami, A., Croci, M., Li, B., Cameron, P., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs. Advances in Neural Information Processing Systems, 37:100213–100240, 2024.
  3. 3.Banner, R., Hubara, I., Hoffer, E., and Soudry, D. Scalable Methods for 8-bit Training of Neural Networks. Advances in Neural Information Processing Systems, 31, 2018.
  4. 4.Bengio, Y., Leonard, N., and Courville, A. Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation. arXiv preprint arXiv:1308.3432, 2013.
  5. 5.Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al. Deepseek LLM: Scaling Open-Source Language Models with Longtermism. arXiv preprint arXiv:2401.02954, 2024.
  6. 6.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. PIQA: Reasoning about Physical Commonsense in Natural Language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  7. 7.Chen, S., Wang, W., and Pan, S. J. MetaQuant: Learning to Quantize by Learning to Penetrate Non-differentiable Quantization. Advances in Neural Information Processing Systems, 32, 2019.
  8. 8.Cheng, W., Zhang, W., Shen, H., Cai, Y., He, X., Lv, K., and Liu, Y. Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs. arXiv preprint arXiv:2309.05516, 2023.
  9. 9.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2924–2936, 2019.
  10. 10.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv preprint arXiv:1803.05457, 2018.
  11. 11.Dettmers, T. and Zettlemoyer, L. The case for 4-bit precision: k-bit Inference Scaling Laws. In International Conference on Machine Learning, pp. 7750–7774. PMLR, 2023.
  12. 12.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. Advances in Neural Information Processing Systems, 35:30318–30332, 2022.
  13. 13.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. Advances in Neural Information Processing Systems, 36:10088–10115, 2023.
  14. 14.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024.
  15. 15.Fishman, M., Chmiel, B., Banner, R., and Soudry, D. Scaling FP8 training to trillion-token LLMs. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=E1EHO0imOb.
  16. 16.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. GPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=tcbBPnfwxS.
  17. 17.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv preprint arXiv:2101.00027, 2020.
  18. 18.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 07 2024. URL https://zenodo.org/records/12608602.
  19. 19.Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J. Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4852–4861, 2019.
  20. 20.Huang, X., Shen, Z., Li, S., Liu, Z., Xianghong, H., Wicksana, J., Xing, E., and Cheng, K.-T. SDQ: Stochastic Differentiable Quantization with Mixed Precision. In International Conference on Machine Learning, pp. 9295–9309. PMLR, 2022.
  21. 21.Kahan, W. IEEE standard 754 for binary floating-point arithmetic. Lecture Notes on the Status of IEEE, 754 (94720-1776):11, 1996.
  22. 22.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361, 2020.
  23. 23.Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 13355–13364, 2024.
  24. 24.Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S. Y., Bansal, H., Guha, E., Keh, S. S., Arora, K., et al. DataComp-LM: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems, 37:14200–14282, 2024.
  25. 25.Li, M., Lin, Y., Zhang, Z., Cai, T., Guo, J., Li, X., Xie, E., Meng, C., Zhu, J.-Y., and Han, S. SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=vWR3KuiQur.
  26. 26.Li, Y., Xu, S., Zhang, B., Cao, X., Gao, P., and Guo, G. Q-ViT: Accurate and Fully Quantized Low-bit Vision Transformer. Advances in Neural Information Processing Systems, 35:34451–34463, 2022.
  27. 27.Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. Proceedings of Machine Learning and Systems, 6:87–100, 2024a.
  28. 28.Lin, Y., Tang, H., Yang, S., Zhang, Z., Xiao, G., Gan, C., and Han, S. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving. arXiv preprint arXiv:2405.04532, 2024b.
  29. 29.Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y. LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pp. 3622–3628, 2021.
  30. 30.Liu, S., Liu, Z., Huang, X., Dong, P., and Cheng, K.-T. LLM-FP4: 4-Bit Floating-Point Quantized Transformers. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023a. URL https://openreview.net/forum?id=wiI8ycNfgJ.
  31. 31.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models. arXiv preprint arXiv:2305.17888, 2023b.
  32. 32.Liu, Z., Zhao, C., Fedorov, I., Soran, B., Choudhary, D., Krishnamoorthi, R., Chandra, V., Tian, Y., and Blankevoort, T. SpinQuant: LLM quantization with learned rotations. arXiv preprint arXiv:2405.16406, 2024.
  33. 33.Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J., and Wei, F. The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits. arXiv preprint arXiv:2402.17764, 2024.
  34. 34.Mellempudi, N., Srinivasan, S., Das, D., and Kaul, B. Mixed Precision Training With 8-bit Floating Point. arXiv preprint arXiv:1905.12334, 2019.
  35. 35.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer Sentinel Mixture Models. In International Conference on Learning Representations, 2017.
  36. 36.Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al. Mixed Precision Training. arXiv preprint arXiv:1710.03740, 2017.
  37. 37.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2381–2391, 2018.
  38. 38.Nvidia. Using FP8 with Transformer Engine, 2022. URL https://docs.nvidia.com/deeplearning/transformer-engine/user-guide/examples/fp8_primer.html.
  39. 39.Nvidia. NVIDIA H100 Tensor Core GPU Architecture, 2023. URL https://resources.nvidia.com/en-us-tensor-core.
  40. 40.Nvidia. NVIDIA Blackwell Architecture Technical Brief, 2024. URL https://resources.nvidia.com/en-us-blackwell-architecture.
  41. 41.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N.-Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1525–1534, 2016.
  42. 42.Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., et al. FP8-LM: Training FP8 Large Language Models. arXiv preprint arXiv:2310.18313, 2023.
  43. 43.Rouhani, B. D., Garegrat, N., Savell, T., More, A., Han, K.-N., Zhao, R., Hall, M., Klar, J., Chung, E., Yu, Y., et al. OCP Microscaling Formats (MX) Specification, 2023a. URL https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf.
  44. 44.Rouhani, B. D., Zhao, R., More, A., Hall, M., Khodamoradi, A., Deng, S., Choudhary, D., Cornea, M., Dellinger, E., Denolf, K., et al. Microscaling Data Formats for Deep Learning. arXiv preprint arXiv:2310.10537, 2023b.
  45. 45.Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V. V., Cui, X., Zhang, W., and Gopalakrishnan, K. Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks. Advances in Neural Information Processing Systems, 32, 2019.
  46. 46.Sun, X., Wang, N., Chen, C.-Y., Ni, J., Agrawal, A., Cui, X., Venkataramani, S., El Maghraoui, K., Srinivasan, V. V., and Gopalakrishnan, K. Ultra-Low Precision 4-bit Training of Deep Neural Networks. Advances in Neural Information Processing Systems, 33:1796–1807, 2020.
  47. 47.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288, 2023.
  48. 48.Uhlich, S., Mauch, L., Yoshiyama, K., Cardinaux, F., Garcia, J. A., Tiedemann, S., Kemp, T., and Nakamura, A. Differentiable Quantization of Deep Neural Networks. arXiv preprint arXiv:1905.11452, 2(8), 2019.
  49. 49.Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y., and Wei, F. BitNet: Scaling 1-bit Transformers for Large Language Models. arXiv preprint arXiv:2310.11453, 2023.
  50. 50.Wang, J., Liu, H., Feng, D., Ding, J., and Ding, B. FP4-Quantization: Lossless 4bit Quantization for Large Language Models. In 2024 IEEE International Conference on Joint Cloud Computing (JCC), pp. 61–67. IEEE, 2024.
  51. 51.Wang, N., Choi, J., Brand, D., Chen, C.-Y., and Gopalakrishnan, K. Training Deep Neural Networks with 8-bit Floating Point Numbers. Advances in Neural Information Processing Systems, 31, 2018.
  52. 52.Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X. Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models. Advances in Neural Information Processing Systems, 35:17402–17414, 2022.
  53. 53.Welbl, J., Liu, N. F., and Gardner, M. Crowdsourcing Multiple Choice Science Questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pp. 94–106, 2017.
  54. 54.Wu, X., Li, C., Aminabadi, R. Y., Yao, Z., and He, Y. Understanding Int4 Quantization for Language Models: Latency Speedup, Composability, and Failure Cases. In International Conference on Machine Learning, pp. 37524–37539. PMLR, 2023.
  55. 55.Xi, H., Li, C., Chen, J., and Zhu, J. Training Transformers with 4-bit Integers. Advances in Neural Information Processing Systems, 36:49146–49168, 2023.
  56. 56.Xi, H., Chen, Y., Zhao, K., Teh, K. J., Chen, J., and Zhu, J. Jetfire: Efficient and accurate transformer pretraining with int8 data flow and per-block quantization. In International Conference on Machine Learning, pp. 54049–54063. PMLR, 2024.
  57. 57.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023.
  58. 58.Yang, Y., Deng, L., Wu, S., Yan, T., Xie, Y., and Li, G. Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers. Neural Networks, 125:70–82, 2020.
  59. 59.Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers. Advances in Neural Information Processing Systems, 35:27168–27183, 2022.
  60. 60.Yin, P., Lyu, J., Zhang, S., Osher, S., Qi, Y., and Xin, J. Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets. In International Conference on Learning Representations, 2019.
  61. 61.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. HellaSwag: Can a Machine Really Finish Your Sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, 2019.

Citation

MLA
Wang, R., et al. “Optimizing Large Language Model Training Using FP4 Quantization”. arXiv, 2025, https://doi.org/10.48550/arxiv.2501.17116.
APA
Wang, R., Gong, Y., Liu, X., Zhao, G., Yang, Z., Guo, B., Zha, Z., & Cheng, P. (2025). Optimizing Large Language Model Training Using FP4 Quantization. arXiv. https://doi.org/10.48550/arxiv.2501.17116
Chicago
Wang, R., Y. Gong, X. Liu, et al. 2025. “Optimizing Large Language Model Training Using FP4 Quantization”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2501.17116.
Harvard
Wang, R. et al. (2025) “Optimizing Large Language Model Training Using FP4 Quantization”. arXiv. Available at: https://doi.org/10.48550/arxiv.2501.17116.
Vancouver
1. Wang R, Gong Y, Liu X, Zhao G, Yang Z, Guo B, Zha Z, Cheng P (2025) Optimizing Large Language Model Training Using FP4 Quantization. https://doi.org/10.48550/arxiv.2501.17116

BibTeX

@misc{https://doi.org/10.48550/arxiv.2501.17116,
  doi = {10.48550/ARXIV.2501.17116},
  url = {https://arxiv.org/abs/2501.17116},
  author = {Wang, Ruizhe and Gong, Yeyun and Liu, Xiao and Zhao, Guoshuai and Yang, Ziyue and Guo, Baining and Zha, Zhengjun and Cheng, Peng},
  keywords = {Machine Learning (cs.LG), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Optimizing Large Language Model Training Using FP4 Quantization},
  publisher = {arXiv},
  year = {2025},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/