FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization

Jung Hyun LeeJeonghoon KimSe Jung KwonDongsoo Lee

article2023ICML66 citations

Proposes FlexRound, a post-training weight quantization method that uses element-wise division to adaptively scale individual weights by magnitude alongside a shared grid size, enabling uniform low-bit quantization for vision architectures and large language models with negligible accuracy loss.

Listen

Deploying advanced artificial intelligence models on resource-constrained hardware requires model compression techniques to lower computational and memory overhead. Post-training quantization—the process of converting high-precision numerical values to lower-bit representations without full retraining or extensive datasets—offers an efficient deployment pathway. However, conventional rounding methods rely on additive shifts that restrict weight adjustments to immediately adjacent values and typically keep the overall quantization scale fixed, causing significant accuracy drops in low-bit environments.

The article demonstrates and evaluates FlexRound, a learnable weight-rounding framework based on element-wise division rather than addition. The primary objective is to simultaneously optimize a shared layer-level quantization grid scale and individual weight scales, allowing the model to adaptively round weights across a wider range of discrete values based on their numerical importance.

The researchers assessed FlexRound across extensive benchmarks covering computer vision, natural language understanding, and natural language generation. Testing utilized architectures such as ResNet, MobileNetV2, BERT, OPT, GPT-Neo, GPT-2, and large models like LLaMA-33B. Credibility was supported by testing both weights-only and joint weight-activation quantization across multiple bit-widths (2-bit, 3-bit, 4-bit, and 8-bit) using standard calibration sample sizes (typically 128 to 1,024 samples) and comparing directly against established baselines like AdaRound and AdaQuant.

The experimental findings show that FlexRound consistently outperforms prior rounding methods across domains. In vision models, it significantly rescued low-bit MobileNetV2 performance, reaching 51.49% top-1 accuracy in a 3-bit weight and activation setup where AdaRound achieved only 39.86%. For language understanding benchmarks on the GLUE dataset, 8-bit quantized models using FlexRound matched or approached full-precision accuracy. Similarly, in large language models, 8-bit quantization on LLaMA-33B preserved near-baseline performance across zero-shot reasoning benchmarks and causal language modeling (yielding a perplexity of 6.82 versus 6.35 for the original half-precision model, well ahead of AdaRound's 10.39).

These results imply that organizations can compress deep learning networks down to low integer precision to dramatically decrease memory footprints and hardware costs while preserving model accuracy. Because the division-based formulation inherently accounts for weight magnitude during updates, FlexRound eliminates the need to rely on assumptions about outlier patterns or brittle weight equalization preprocessing steps.

For practical implementation, teams deploying deep learning models on constrained hardware should consider adopting division-based post-training quantization pipelines. Where extreme compression is required, practitioners should ensure calibration sample sizes do not fall below key thresholds (e.g., at least 32 to 64 samples) and perform light tuning of learning rates on task-specific layers. While confidence in the reported results is high across standard vision and language benchmarks, future work should explore the formal combination of division-based scaling with other additive techniques across even broader edge-device hardware constraints.

No sufficiently relevant recommendations were found.

Cover for FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization

Abstract

Post-training quantization (PTQ) has been gaining popularity for the deployment of deep neural networks on resource-limited devices since unlike quantization-aware training, neither a full training dataset nor end-to-end training is required at all. As PTQ schemes based on reconstructing each layer or block output turn out to be effective to enhance quantized model performance, recent works have developed algorithms to devise and learn a new weight-rounding scheme so as to better reconstruct each layer or block output. In this work, we propose a simple yet effective new weight-rounding mechanism for PTQ, coined FlexRound, based on element-wise division instead of typical element-wise addition such that FlexRound enables jointly learning a common quantization grid size as well as a different scale for each pre-trained weight. Thanks to the reciprocal rule of derivatives induced by element-wise division, FlexRound is inherently able to exploit pre-trained weights when updating their corresponding scales, and thus, flexibly quantize pre-trained weights depending on their magnitudes. We empirically validate the efficacy of FlexRound on a wide range of models and tasks. To the best of our knowledge, our work is the first to carry out comprehensive experiments on not only image classification and natural language understanding but also natural language generation, assuming a per-tensor uniform PTQ setting. Moreover, we demonstrate, for the first time, that large language models can be efficiently quantized, with only a negligible impact on performance compared to half-precision baselines, achieved by reconstructing the output in a block-by-block manner.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminaries
  • 3.2. FlexRound
  • 4. Experiments
  • 4.1. Ablation Study
  • 4.2. ResNet and MobileNetV2 on ImageNet
  • 4.3. Language Models
  • LLaMA on Common Sense Reasoning and WikiText2
  • 5. Conclusion
  • References
  • A. Comparison of Rounding Results of AdaRound, AdaQuant, and FlexRound
  • B. Proof of Proposition 3.1
  • C. ResNet-18, ResNet-50, and MobileNetV2 on ImageNet with Pre-trained Models from the Official PyTorch Repository 2
  • D. Cross-Layer Equalization and Absorbing High Biases as Preprocessing
  • E. Ablation Study on Sample Size
  • F. Combining an Additive Approach with a Division-based Approach
  • G. BERT on SQuAD
  • H. BERT and GPT-Neo on GLUE
  • I. GPT-Neo and OPT on WikiText2 and PTB
  • J. GPT-2 on WebNLG
  • K. LLaMA on Common Sense Reasoning and WikiText2
  • L. LLaMA fine-tuned via LoRA on WikiText2 and PTB

Knowls

  1. Knowl 1 — FlexRound uses element-wise division to learn weight quantization scales

    model/method

    FlexRound is a post-training weight-quantization method for a real-valued layer weight tensor WW. It maps each weight to a quantization grid using element-wise division by a learned positive scale, then rounds to the nearest integer grid point. For per-tensor quantization, the common grid size s1>0s_1>0 is shared across a layer. The quantized weights WcW^c are

    Wc=s1round⁡ ⁣(Ws1⊙S2⊙s3)W^c=s_1\operatorname{round}\!\left(\frac{W}{s_1\odot S_2\odot s_3}\right)

    for a linear layer, and

    Wc=s1round⁡ ⁣(Ws1⊙S2⊙s3⊙s4)W^c=s_1\operatorname{round}\!\left(\frac{W}{s_1\odot S_2\odot s_3\odot s_4}\right)

    for a 2D convolution. Here, S2S_2 has the same shape as WW and provides element-specific scaling; s3s_3 scales output channels (shape Cout×1C_{\mathrm{out}}\times1 for a linear layer and Cout×1×1×1C_{\mathrm{out}}\times1\times1\times1 for a convolution); and convolutional s4s_4 scales input channels (shape 1×Cin×1×11\times C_{\mathrm{in}}\times1\times1). The symbol ⊙\odot denotes element-wise multiplication with broadcasting, and round⁡\operatorname{round} denotes nearest-integer rounding. The learned scales are constrained to be positive; S2S_2, s3s_3, and s4s_4 are initialized to one.

    During calibration, the scales are optimized to minimize the squared Frobenius-norm reconstruction error L=∥WX−WcXf∥F2L=\|WX-W^cX_f\|_F^2. Here, XX is the layer input when earlier layers are unquantized, XfX_f is the input when earlier layers have been quantized, and ∥⋅∥F\|\cdot\|_F is the Frobenius norm. The method can also use a vector-valued grid size for per-channel quantization, but the paper focuses on per-tensor uniform PTQ.

  2. Knowl 2 — Division makes scale gradients depend directly on pretrained weight values

    theoretical result

    For FlexRound, let S′S' denote the element-wise scale multiplying a pretrained weight: S′=S2⊙s3S'=S_2\odot s_3 for a linear layer or S′=S2⊙s3⊙s4S'=S_2\odot s_3\odot s_4 for a 2D convolution. Let WijW_{ij} be one real-valued pretrained weight, WijcW^c_{ij} its quantized value, and LL the layer-output reconstruction loss. Under the straight-through estimator for the rounding operation, the scale gradient is

    ∂L∂Sij′=−Wij(Sij′)2∂L∂Wijc.\frac{\partial L}{\partial S'_{ij}}=-\frac{W_{ij}}{(S'_{ij})^2}\frac{\partial L}{\partial W^c_{ij}}.

    The same relation applies to a convolutional weight indexed by its four tensor coordinates. Thus the pretrained weight magnitude directly contributes to the scale update, but the update also depends on the reconstruction loss's sensitivity to that quantized weight. A large-magnitude weight is not necessarily shifted far from its rounding-to-nearest grid if its loss gradient is near zero.

  3. Knowl 3 — FlexRound permits magnitude-dependent shifts beyond adjacent quantization grids

    empirical result

    In a 4-bit weight-only experiment with activations kept in full precision, the authors examined weight shifts in the first 2D convolution of the first block of MobileNetV2 and ResNet-18. They report that about 12.8% of MobileNetV2 weights in that layer were rounded more than one grid step away from the rounding-to-nearest choice, compared with about 1.5% for ResNet-18. MobileNetV2 had weights with absolute value above one in the examined layer, whereas ResNet-18 did not. These observations are consistent with FlexRound's scale gradients responding to pretrained weight magnitudes. They do not imply that larger weights must always shift farther: in another MobileNetV2 convolution, large-magnitude weights showed flexibility comparable to moderate-magnitude weights.

  4. Knowl 4 — Ablations show gains from jointly learning the grid size and adding channel scales

    data/table

    On ImageNet with 4-bit weights and full-precision activations, the ablation compares FlexRound with a learned or fixed common grid size s1s_1, and with or without the additional channel scales s3s_3 and s4s_4. The entries below are top-1/top-5 accuracy in percent; the full-precision row is the reference. Jointly learning s1s_1 gives the complete method its best top-1 result on each model, while adding s3s_3 and s4s_4 improves top-1 accuracy over the version without those scales.

    Methods1s_1S2S_2s3,s4s_3,s_4ResNet-18ResNet-50MobileNetV2
    Full-precisionN/AN/AN/A71.00/89.9776.63/93.0472.62/90.67
    B + AdaQuantLearnableN/AN/A67.50/87.7572.79/90.7715.17/32.89
    B + AdaRoundFixedN/AN/A70.18/89.3875.86/92.6269.46/88.85
    B + FlexRoundLearnablePresentPresent70.28/89.4475.95/92.6870.82/89.67
    FlexRound, fixed s1s_1FixedPresentPresent70.09/89.4375.88/92.6169.47/88.85
    FlexRound without s3,s4s_3,s_4LearnablePresentAbsent70.22/89.4575.92/92.6370.51/89.49

    The experiments use the BRECQ reconstruction setting; “B +” denotes a rounding method used in that setting. The two ablations retain the division-based form but remove either joint grid-size learning or the extra channel scales.

  5. Knowl 5 — ImageNet results favor FlexRound across weight-only and weight-plus-activation PTQ

    data/table

    ResNet-18, ResNet-50, and MobileNetV2 were evaluated on ImageNet using per-tensor symmetric quantization. The values are top-1/top-5 accuracy in percent. The weight-only comparison uses the BRECQ reconstruction setting; the joint weight-and-activation comparison uses either BRECQ (“B +”) or QDrop (“Q +”). The study used 1,024 randomly sampled images, 5,000 reconstruction iterations, and the median of five random trials; the first and last layers were quantized to 8-bit. FlexRound's gains are especially pronounced for MobileNetV2 at low bit widths, while 4-bit joint quantization retains less than a two-point top-1 gap from full precision for both ResNets.

    Weights/activationsSetting and methodResNet-18ResNet-50MobileNetV2
    32/32Full-precision71.00/89.9776.63/93.0472.62/90.67
    4/32B + AdaQuant67.50/87.7572.79/90.7715.17/32.89
    4/32B + AdaRound70.18/89.3875.86/92.6269.46/88.85
    4/32B + FlexRound70.28/89.4475.95/92.6870.82/89.67
    3/32B + AdaQuant57.09/80.8252.13/75.220.20/0.79
    3/32B + AdaRound68.79/88.6274.31/91.8162.51/84.52
    3/32B + FlexRound68.65/88.5474.38/91.8166.87/87.56
    2/32B + AdaQuant0.23/0.920.10/0.500.10/0.50
    2/32B + AdaRound61.99/84.8148.47/77.0939.57/66.18
    2/32B + FlexRound62.57/84.8463.67/85.7246.04/72.48
    4/4B + AdaRound69.18/88.8574.44/91.8061.05/83.30
    4/4B + FlexRound69.32/88.8374.56/91.8763.74/85.01
    4/4Q + AdaRound69.20/88.9674.90/92.1565.42/86.23
    4/4Q + FlexRound69.26/88.8175.08/92.2066.66/87.21
    3/3B + AdaRound64.83/86.1267.01/87.283.74/11.54
    3/3B + FlexRound64.99/85.9368.29/87.8925.43/48.28
    3/3Q + AdaRound65.71/86.9670.49/89.9339.86/66.00
    3/3Q + FlexRound65.43/86.6070.74/89.7851.49/76.90
  6. Knowl 6 — FlexRound improves 8-bit quantized language understanding across the reported GLUE tasks

    data/table

    BERT and GPT-Neo models were evaluated on GLUE after quantizing weights and input activations of attention and feed-forward sublayers to 8-bit using per-tensor asymmetric quantization. Reconstruction used 1,024 training examples for 20,000 iterations; the QDrop setting dropped activation quantization with probability 0.5. The common FlexRound learning rate was 2×10−42\times10^{-4} for BERT and 3×10−43\times10^{-4} for GPT-Neo. Each pair below follows the dataset's reported metric: MNLI matched/mismatched accuracy, QQP F1/accuracy, and MRPC accuracy. FlexRound exceeds AdaRound on all reported model-task comparisons, and often approaches or exceeds the full-precision score.

    Dataset and methodBERT-BaseBERT-LargeGPT-Neo-125MGPT-Neo-1.3BGPT-Neo-2.7B
    MNLI full-precision84.49/85.2086.05/85.9879.11/79.6385.12/86.0486.36/87.02
    MNLI Q + AdaRound83.69/84.6185.75/85.8672.67/74.1184.90/85.8286.33/86.75
    MNLI Q + FlexRound84.53/84.9885.93/85.9972.94/74.2485.56/86.1486.41/86.89
    QQP full-precision88.06/91.0888.66/91.5985.20/88.9988.26/91.2888.62/91.50
    QQP Q + AdaRound87.65/90.5887.48/90.6272.97/79.3587.98/91.0488.38/91.27
    QQP Q + FlexRound87.81/90.8388.38/91.3173.75/80.6588.27/91.1888.60/91.39
    MRPC full-precision85.0585.5480.1585.0587.99
    MRPC Q + AdaRound81.6282.3575.2584.8085.78
    MRPC Q + FlexRound84.0784.3175.4985.0586.76

    On SQuADv1, 8-bit per-tensor quantization also yielded F1 scores of 87.25 for BERT-Base and 89.25 for BERT-Large with FlexRound, compared with 86.90 and 88.89 for AdaRound and 87.05 and 89.31 in full precision.

  7. Knowl 7 — FlexRound improves quantized natural-language generation results

    data/table

    The reported generation experiments cover GPT-Neo and OPT on WikiText2 and Penn Treebank (PTB), plus GPT-2 with LoRA fine-tuning on WebNLG. For WikiText2 and PTB, weights and activations in attention and feed-forward sublayers were quantized to 8-bit per tensor with an asymmetric scheme; reconstruction used 128 examples. The metric is perplexity (PPL), for which lower is better. For WebNLG, GPT-2 medium and large were merged with LoRA, and the metric is BLEU, for which higher is better; reconstruction used 128 examples, 500 iterations, and batch size 8.

    DatasetMethodGPT-Neo-125MGPT-Neo-1.3BGPT-Neo-2.7BOPT-125MOPT-1.3BOPT-2.7B
    WikiText2 PPLFull-precision21.9612.0910.7819.8511.5210.27
    WikiText2 PPLQ + AdaRound30.5212.4714.0927.9612.6610.97
    WikiText2 PPLQ + FlexRound24.3012.3712.4321.4312.0210.63
    PTB PPLFull-precision24.2016.0914.7016.5011.6210.80
    PTB PPLQ + AdaRound31.4016.6319.8020.2813.0012.02
    PTB PPLQ + FlexRound26.0316.3216.8717.6812.2211.29
    WebNLG model and BLEU categoryFull-precision LoRAQ + AdaRoundQ + FlexRound
    GPT-2 medium, unseen47.1645.7046.85
    GPT-2 medium, seen62.3160.9261.83
    GPT-2 medium, all55.4354.0555.06
    GPT-2 large, unseen48.0648.0948.42
    GPT-2 large, seen64.3963.9864.47
    GPT-2 large, all56.9756.7557.16

    FlexRound has lower PPL than AdaRound for every GPT-Neo and OPT model-dataset pair shown, and higher BLEU than AdaRound in every WebNLG category. Its scores are also close to the corresponding full-precision or full-precision-LoRA results.

  8. Knowl 8 — Eight-bit FlexRound preserves LLaMA zero-shot performance more closely than AdaRound

    data/table

    LLaMA-7B, LLaMA-13B, and LLaMA-33B were evaluated on seven zero-shot common-sense reasoning benchmarks and causal language modeling on WikiText2. Attention and feed-forward weights and activations were quantized to 8-bit: weights per-channel asymmetric, activations per-tensor asymmetric. The reconstruction used 512 samples from C4. Common-sense results are accuracy in percent; WikiText2 results are perplexity, where lower is better. FlexRound outperforms AdaRound on every listed measure for each model, and its results are generally close to the half-precision baseline.

    Model and methodBoolQPIQAHellaSwagWinoGrandeARC-eARC-cOBQAWikiText2 PPL
    LLaMA-7B half-precision73.1577.3172.9667.0952.4841.3842.408.90
    LLaMA-7B Q + AdaRound70.1275.0869.8965.8251.4739.4239.0010.38
    LLaMA-7B Q + FlexRound73.7676.6671.7567.0152.3140.0242.209.25
    LLaMA-13B half-precision68.5379.1176.2370.0159.8944.5442.207.73
    LLaMA-13B Q + AdaRound66.0976.4472.0666.3057.3243.0039.609.07
    LLaMA-13B Q + FlexRound68.5978.6775.2170.6458.8843.6041.208.01
    LLaMA-33B half-precision68.3880.0979.2172.9358.9245.4842.006.35
    LLaMA-33B Q + AdaRound64.8674.6568.6457.9349.2836.9541.0010.39
    LLaMA-33B Q + FlexRound69.0879.1677.4372.5356.6144.9744.006.82
  9. Knowl 9 — Four-bit LLaMA weight quantization is viable but its gains vary by model and task

    data/table

    A separate LLaMA experiment quantized attention and feed-forward weights to 4-bit per-channel asymmetric values while keeping activations in half precision; reconstruction used 512 C4 samples. The table reports zero-shot accuracy in percent and WikiText2 perplexity (PPL; lower is better). FlexRound's advantage over AdaRound is mixed across individual accuracy tasks for LLaMA-7B and LLaMA-13B, but its WikiText2 PPL is lower for all three sizes. At 33B, the five-shot results also favor FlexRound over AdaRound on all seven reported accuracy tasks: BoolQ 86.64 vs. 84.65, PIQA 81.83 vs. 80.96, HellaSwag 81.26 vs. 80.03, WinoGrande 79.01 vs. 78.37, ARC-e 70.66 vs. 67.51, ARC-c 53.24 vs. 51.19, and OBQA 45.00 vs. 44.60.

    Model and methodBoolQPIQAHellaSwagWinoGrandeARC-eARC-cOBQAWikiText2 PPL
    LLaMA-7B half-precision73.1577.3172.9667.0952.4841.3842.408.90
    LLaMA-7B B + AdaRound70.4677.0471.7368.2751.7340.4442.009.69
    LLaMA-7B B + FlexRound70.7377.7571.9766.0650.8040.2742.209.18
    LLaMA-13B half-precision68.5379.1176.2370.0159.8944.5442.207.73
    LLaMA-13B B + AdaRound67.5578.9475.5069.8558.4243.0043.408.07
    LLaMA-13B B + FlexRound66.3978.7875.5270.4059.5543.7742.807.90
    LLaMA-33B half-precision68.3880.0979.2172.9358.9245.4842.006.35
    LLaMA-33B B + AdaRound69.3979.2777.7772.6957.0344.6243.006.88
    LLaMA-33B B + FlexRound67.1980.2579.0172.6157.7944.8843.806.63
  10. Knowl 10 — FlexRound can require task-specific calibration choices and enough samples

    limitation

    FlexRound does not outperform AdaRound on every task under one shared learning rate. In the reported GLUE results with the default rates, for example, GPT-Neo-125M FlexRound scores 80.52 on QNLI and 83.03 on SST-2, below AdaRound's 80.87 and 84.75, respectively. Tuning the scale-learning rate for these tasks improves most of the comparisons, but does not eliminate every gap: GPT-Neo-125M FlexRound reaches 83.72 on SST-2 after tuning, still below AdaRound's 84.75. The MobileNetV2 calibration-size study also reports that FlexRound accuracy falls by almost one percentage point when its sample size is reduced from 64 to 32, although it remains above AdaRound across the tested sample sizes. These results qualify the broad performance comparisons: task-specific optimization and calibration data quantity can matter.

Coverage note — The supplementary five-shot LLaMA tables, LoRA-adapted LLaMA language-modeling results, preprocessing comparisons, and additive-plus-division combination experiments are omitted as secondary variants; the core method and representative vision, understanding, generation, and LLaMA results are included.

References

  1. 1.Bai, H., Hou, L., Shang, L., Jiang, X., King, I., and Lyu, M. R. Towards efficient post-training quantization of pre-trained language models. arXiv preprint arXiv:2109.15082, 2021.
  2. 2.Bengio, Y., Leonard, N., and Courville, A. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  3. 3.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  4. 4.Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata.
  5. 5.Bondarenko, Y., Nagel, M., and Blankevoort, T. Understanding and overcoming the challenges of efficient transformer quantization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7947–7969. Association for Computational Linguistics, November 2021. doi: 10.18653/v1/2021.emnlp-main.627. URL https://aclanthology.org/2021.emnlp-main.627.
  6. 6.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2924–2936, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1300. URL https://aclanthology.org/N19-1300.
  7. 7.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  8. 8.Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., and Bengio, Y. Binarized neural networks: Training deep neural networks with weights and activations constrained to +1 or-1. arXiv preprint arXiv:1602.02830, 2016.
  9. 9.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  10. 10.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S. Learned step size quantization. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkgO66VKDS.
  12. 12.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  13. 13.Gardent, C., Shimorina, A., Narayan, S., and Perez-Beltrachini, L. The WebNLG challenge: Generating text from RDF data. In Proceedings of the 10th International Conference on Natural Language Generation, pp. 124–133, Santiago de Compostela, Spain, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-3518. URL https://aclanthology.org/W17-3518.
  14. 14.Han, S., Pool, J., Tran, J., and Dally, W. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, 2015.
  15. 15.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. In International Conference on Learning Representations, 2016. URL https://arxiv.org/pdf/1510.00149.pdf.
  16. 16.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  17. 17.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pp. 2790–2799. PMLR, 2019.
  18. 18.Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  19. 19.Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D. Accurate post training quantization with small calibration sets. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 4466–4475. PMLR, 2021. URL https://proceedings.mlr.press/v139/hubara21a.html.
  20. 20.Jain, S. R., Gural, A., Wu, M., and Dick, C. H. Trained quantization thresholds for accurate and efficient fixed-point inference of deep neural networks. arXiv preprint arXiv:1903.08066, 2019.
  21. 21.Jung, S., Son, C., Lee, S., Son, J., Han, J.-J., Kwak, Y., Ju Hwang, S., and Choi, C. Learning to quantize deep networks by optimizing quantization intervals with task loss. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4350–4359, 2019.
  22. 22.Kim, S., Park, G., and Yi, Y. Performance evaluation of int8 quantized inference on mobile gpus. IEEE Access, 9:164245–164255, 2021.
  23. 23.Lee, J. H., Yun, J., Hwang, S. J., and Yang, E. Cluster-promoting quantization with bit-drop for minimizing network quantization loss. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5350–5359. IEEE Computer Society, 2021. URL https://doi.ieeecomputersociety.org/10.1109/ICCV48922.2021.00532.
  24. 24.Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., and Gu, S. BRECQ: Pushing the limit of post-training quantization by block reconstruction. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=POWv6hDd9XH.
  25. 25.Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., and Tang, J. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 61–68, 2022.
  26. 26.Lou, Q., Guo, F., Kim, M., Liu, L., and Jiang., L. Autoq: Automated kernel-wise neural network quantization. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygfnn4twS.
  27. 27.Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2):313–330, 1993. URL https://www.aclweb.org/anthology/J93-2004.
  28. 28.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models, 2016.
  29. 29.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  30. 30.Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M. Data-free quantization through weight equalization and bias correction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1325–1334, 2019.
  31. 31.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? Adaptive rounding for post-training quantization. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 7197–7206. PMLR, 2020. URL https://proceedings.mlr.press/v119/nagel20a.html.
  32. 32.Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., and Blankevoort, T. A white paper on neural network quantization. arXiv preprint arXiv:2106.08295, 2021.
  33. 33.Nahshan, Y., Chmiel, B., Baskin, C., Zheltonozhskii, E., Banner, R., Bronstein, A. M., and Mendelson, A. Loss aware post-training quantization. Machine Learning, 110(11):3245–3262, 2021.
  34. 34.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  35. 35.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  36. 36.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv e-prints, art. arXiv:1606.05250, 2016.
  37. 37.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  38. 38.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  39. 39.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510–4520, 2018.
  40. 40.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023.
  41. 41.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  42. 42.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  43. 43.Wang, P., Chen, Q., He, X., and Cheng, J. Towards accurate post-training network quantization via bit-split and stitching. In International Conference on Machine Learning, pp. 9847–9856. PMLR, 2020.
  44. 44.Wei, X., Gong, R., Li, Y., Liu, X., and Yu, F. QDrop: Randomly dropping quantization for extremely low-bit post-training quantization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=ySQH0oDyp7.
  45. 45.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6.
  46. 46.Wu, H., Judd, P., Zhang, X., Isaev, M., and Micikevicius, P. Integer quantization for deep learning inference: Principles and empirical evaluation. arXiv preprint arXiv:2004.09602, 2020.
  47. 47.Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438, 2022.
  48. 48.Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. arXiv preprint arXiv:2206.01861, 2022.
  49. 49.Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M. Q8bert: Quantized 8bit bert. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS), pp. 36–39. IEEE, 2019.
  50. 50.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  51. 51.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  52. 52.Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q. Ternarybert: Distillation-aware ultra-low bit bert. arXiv preprint arXiv:2009.12812, 2020.
  53. 53.Zhao, R., Hu, Y., Dotzel, J., De Sa, C., and Zhang, Z. Improving neural network quantization without retraining using outlier channel splitting. In International conference on machine learning, pp. 7543–7552. PMLR, 2019.
  54. 54.Zhao, X., Wang, Y., Cai, X., Liu, C., and Zhang, L. Linear symmetric quantization of neural networks for low-precision integer hardware. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=H1lBj2VFPS.

Citation

MLA
Lee, J. H., et al. “FlexRound: Learnable Rounding Based on Element-wise Division for Post-Training Quantization”. International Conference on Machine Learning, vol. 202, 2023, pp. 18913–39, https://proceedings.mlr.press/v202/lee23h.html.
APA
Lee, J. H., Kim, J., Kwon, S. J., & Lee, D. (2023). FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization. International Conference on Machine Learning, 202, 18913–18939. https://proceedings.mlr.press/v202/lee23h.html
Chicago
Lee, J. H., J. Kim, S. J. Kwon, and D. Lee. 2023. “FlexRound: Learnable Rounding Based on Element-wise Division for Post-Training Quantization”. International Conference on Machine Learning 202: 18913–39. https://proceedings.mlr.press/v202/lee23h.html.
Harvard
Lee, J.H. et al. (2023) “FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization”, International Conference on Machine Learning. PMLR, pp. 18913–18939. Available at: https://proceedings.mlr.press/v202/lee23h.html.
Vancouver
1. Lee JH, Kim J, Kwon SJ, Lee D (2023) FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization. In: International Conference on Machine Learning. PMLR, pp 18913–18939

BibTeX

@InProceedings{pmlr-v202-lee23h,
  title = 	 {{F}lex{R}ound: Learnable Rounding based on Element-wise Division for Post-Training Quantization},
  author =       {Lee, Jung Hyun and Kim, Jeonghoon and Kwon, Se Jung and Lee, Dongsoo},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {18913--18939},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/lee23h/lee23h.pdf},
  url = 	 {https://proceedings.mlr.press/v202/lee23h.html},
  abstract = 	 {Post-training quantization (PTQ) has been gaining popularity for the deployment of deep neural networks on resource-limited devices since unlike quantization-aware training, neither a full training dataset nor end-to-end training is required at all. As PTQ schemes based on reconstructing each layer or block output turn out to be effective to enhance quantized model performance, recent works have developed algorithms to devise and learn a new weight-rounding scheme so as to better reconstruct each layer or block output. In this work, we propose a simple yet effective new weight-rounding mechanism for PTQ, coined FlexRound, based on element-wise division instead of typical element-wise addition such that FlexRound enables jointly learning a common quantization grid size as well as a different scale for each pre-trained weight. Thanks to the reciprocal rule of derivatives induced by element-wise division, FlexRound is inherently able to exploit pre-trained weights when updating their corresponding scales, and thus, flexibly quantize pre-trained weights depending on their magnitudes. We empirically validate the efficacy of FlexRound on a wide range of models and tasks. To the best of our knowledge, our work is the first to carry out comprehensive experiments on not only image classification and natural language understanding but also natural language generation, assuming a per-tensor uniform PTQ setting. Moreover, we demonstrate, for the first time, that large language models can be efficiently quantized, with only a negligible impact on performance compared to half-precision baselines, achieved by reconstructing the output in a block-by-block manner.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/