Towards Efficient Post-training Quantization of Pre-trained Language Models

Haoli BaiLu HouLifeng ShangXin JiangIrwin KingMichael R. Lyu

article2022NeurIPS62 citations

Proposes a parallel module-wise reconstruction error minimization framework that enables fast, memory-efficient post-training quantization for large language models while achieving accuracy competitive with full quantization-aware training.

Listen

Deploying large pre-trained language models into resource-constrained production environments requires model compression. Network quantization—which converts high-precision numbers into lower-bit formats—effectively reduces model size and latency. However, conventional quantization-aware training relies on end-to-end retraining across full datasets, creating major bottlenecks in training duration, memory consumption, and data privacy. Prior post-training quantization methods resolve these overheads by calibrating on tiny data subsets, but they cause severe accuracy drops when applied to language models.

The article demonstrates an efficient post-training quantization framework called module-wise reconstruction error minimization (MREM). The primary objective is to preserve the accuracy of low-bit quantized language models while substantially reducing training time, memory consumption, and data requirements relative to full retraining.

The authors evaluate this approach on standard language benchmarks using BERT-base and BERT-large architectures across multiple low-bit configurations. The approach divides language models into multi-layer modules rather than optimizing isolated matrix operations, jointly minimizing the reconstruction error within each module against full-precision targets. To accelerate training, modules are distributed across separate processing units in parallel using intermediate input queues, combined with an annealed teacher forcing technique that gradually transitions training signals from clean baseline outputs to quantized outputs to prevent compounding errors.

The analysis yields four key findings. First, module-wise calibration significantly outperforms conventional layer-wise post-training quantization, achieving 83.5% accuracy on MNLI (a 10.2 percentage point improvement) for 4-bit BERT-base, coming within 1.1% of full quantization-aware retraining. Second, the parallel training scheme achieves near-theoretical linear speedup (4x faster across 4 GPUs) and finishes over 150x faster than full retraining methods. Third, memory overhead is reduced by roughly two-thirds, allowing large models to fit onto smaller consumer-grade hardware. Finally, the framework requires only a tiny calibration set of 4,096 unlabeled samples, avoiding the need for massive proprietary training datasets.

These findings indicate that organizations can compress advanced language models at a fraction of the traditional computational cost, timeline, and memory budget without risking user privacy. Teams can rapidly calibrate compact models locally on standard hardware instead of maintaining costly distributed training clusters for full retraining runs.

Engineering teams should adopt module-wise post-training quantization with 4-module parallel partitioning as a cost-effective default for language model compression. For workflows where data privacy or quick deployment is critical, this approach serves as a practical replacement for quantization-aware retraining. Future work should evaluate the framework on modern generative transformer models at scales beyond BERT to verify if these performance and efficiency advantages transfer directly to generative workloads.

Confidence in these findings is high for classification and question-answering tasks within the evaluated model sizes. However, users should exercise caution regarding boundary conditions: partitioning models into too many modules causes slight accuracy trade-offs, and calibration sets smaller than 128 samples remain insufficient for stable convergence.

arXiv: 2109.15082
Cover for Towards Efficient Post-training Quantization of Pre-trained Language Models

Abstract

Network quantization has gained increasing attention with the rapid growth of large pre-trained language models (PLMs). However, most existing quantization methods for PLMs follow quantization-aware training (QAT) that requires end-to-end training with full access to the entire dataset. Therefore, they suffer from slow training, large memory overhead, and data accessibility issues. In this paper, we study post-training quantization (PTQ) of PLMs, and propose module-wise quantization error minimization (MREM), an efficient solution to mitigate these issues. By partitioning the PLM into multiple modules, we minimize the reconstruction error incurred by quantization for each module. In addition, we design a new model parallel training strategy such that each module can be trained locally on separate computing devices without waiting for preceding modules, which brings nearly the theoretical training speed-up (e.g., 4× on 4 GPUs). Experiments on GLUE and SQuAD benchmarks show that our proposed PTQ solution not only performs close to QAT, but also enjoys significant reductions in training time, memory overhead, and data consumption.

Table of Contents

  • 1 Introduction
  • 2 Motivation
  • 2.1 Quantization Background
  • 2.2 Quantizing Pre-trained Language Models: QAT or PTQ?
  • 3 Methodology
  • 3.1 Module-wise Reconstruction Error Minimization
  • 3.2 Accelerated Parallel Training
  • 3.3 Annealed Teaching Forcing
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results: Comparison with QAT and REM
  • 4.3 Main Results: Comparison with Existing Methods
  • 4.4 Discussions
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgement and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Module-wise reconstruction error minimization

    model/method

    Module-wise reconstruction error minimization (MREM) is a post-training quantization method for Transformer-based pre-trained language models. Given a model with LL Transformer layers, the method partitions the model into NN consecutive modules. Module nn contains Transformer layers with indices l∈[ℓn,ℓn+1)l\in[\ell_n,\ell_{n+1}); the embedding layers are included in the first module and the classification head in the last module. For a calibration set D~\tilde D, MREM jointly optimizes the weights wn\mathbf w_n and quantization step sizes sn\mathbf s_n of all coupled linear layers in module nn by minimizing the reconstruction error between each full-precision feed-forward output flf_l and its quantized counterpart f^l\hat f_l:

    min⁡wn,sn  Ln=∑l∈[ℓn,ℓn+1)∥f^l−fl∥22.\min_{\mathbf w_n,\mathbf s_n}\;\mathcal L_n=\sum_{l\in[\ell_n,\ell_{n+1})}\left\|\hat f_l-f_l\right\|_2^2.

    Here, flf_l and f^l\hat f_l are vectors evaluated on calibration examples, and wn\mathbf w_n and sn\mathbf s_n include all trainable quantized weights and step sizes in module nn. The embedding output and final task logits are also trained with mean-squared reconstruction losses. In the sequential variant, previously optimized modules are fixed while the current module is updated. Grouping multiple Transformer layers allows MREM to account for dependencies among the matrix multiplications inside attention and feed-forward sublayers; N=1N=1 becomes an intermediate-layer knowledge-distillation objective, whereas larger NN reduces per-device memory at the cost of considering fewer cross-layer dependencies.

  2. Knowl 2 — Stale-queue model-parallel training

    model/method

    MREM can train all modules concurrently by placing one module on each of NN computing devices and maintaining a full-precision and a quantized FIFO input queue between every pair of adjacent modules. Let fℓntf_{\ell_n}^{t} and f^ℓnt\hat f_{\ell_n}^{t} denote the full-precision and quantized outputs produced by module nn at iteration tt. The queue for the output of module nn stores the most recent t0t_0 outputs:

    Int={fℓnt,fℓnt−1,…,fℓnt−t0+1},\mathcal I_n^t=\{f_{\ell_n}^{t},f_{\ell_n}^{t-1},\ldots,f_{\ell_n}^{t-t_0+1}\},

    with an analogous queue I^nt\hat{\mathcal I}_n^t for quantized outputs. After the queues are filled by t0t_0 sequential forward passes, module n+1n+1 samples an input with replacement from Int\mathcal I_n^t and I^nt\hat{\mathcal I}_n^t without waiting for the current update of module nn. Each device computes its module loss and updates only that module's weights and quantization step sizes; gradients are not propagated across module boundaries. Queues are updated in first-in-first-out order throughout training. This stale-synchronous design removes straggler-induced waiting and provides nearly linear speedup with the number of devices; with four modules on four GPUs, the measured parallel method is approximately four times faster than sequential MREM.

  3. Knowl 3 — Annealed teacher forcing for error propagation

    model/method

    Parallel MREM can propagate the reconstruction error of a partially optimized quantized module into all successor modules. To reduce this effect, annealed teacher forcing replaces the quantized output used as the next module's input with a convex combination of the corresponding full-precision and quantized outputs. For module nn at iteration tt, the input supplied downstream is

    f~ℓnt=λtfℓnt+(1−λt)f^ℓnt,λt∈[0,1].\tilde f_{\ell_n}^{t}=\lambda_t f_{\ell_n}^{t}+(1-\lambda_t)\hat f_{\ell_n}^{t},\qquad \lambda_t\in[0,1].

    The teacher-forcing coefficient follows a linear decay, λt=max⁡(1−t/T0,0)\lambda_t=\max(1-t/T_0,0), where T0T_0 is the duration of the teacher-forcing phase. The paper uses teacher forcing for the first 40%40\% of training by default. At λt=1\lambda_t=1, the downstream quantized module receives a clean full-precision signal, eliminating accumulated predecessor error but creating forward inconsistency; at λt=0\lambda_t=0, training uses the ordinary quantized predecessor output and matches quantized inference. Annealing therefore transitions from error correction early in training to adaptation to the true quantized computation later.

  4. Knowl 4 — Quantization and training configuration

    experimental setup

    The experiments quantize fine-tuned BERT-base and BERT-large models on GLUE and SQuAD. Unless a task has fewer examples, the calibration set contains 4,0964{,}096 randomly sampled training instances; RTE and MRPC use their complete training sets. Each result is averaged over ten calibration-set choices, with standard deviations reported. The default implementation partitions the model into four modules on four NVIDIA V100 GPUs. Sequential MREM trains each module for 2,0002{,}000 steps on GLUE and 4,0004{,}000 steps on SQuAD, using initial learning rates of 10−410^{-4} and 5×10−55\times10^{-5} respectively, followed by linear learning-rate decay. Parallel MREM uses the same optimization settings and annealed teacher forcing.

    The notation W-E-A denotes the bit widths for Transformer weights, word embeddings, and activations. The study evaluates W4-E4-A8, W2-E2-A8, and W2-E2-A4. Weight quantization uses TWN or LAQ for 2-bit and 4-bit weights, while activation quantization uses LSQ. Weights, embeddings, and most activations use symmetric uniform quantization; activations after self-attention and GeLU use asymmetric quantization. Layer normalization, skip connections, biases, and the final classification head are left unquantized.

  5. Knowl 5 — MNLI accuracy, time, memory, and data efficiency

    data/table

    On the MNLI development set, MREM retains accuracy much closer to quantization-aware training (QAT) than ordinary layer-wise reconstruction error minimization (REM), while using only 44K calibration examples and substantially less time and memory. The following BERT-large measurements compare matched and mismatched MNLI accuracy, training time, GPU memory, and calibration data. MREM-P is parallel MREM with teacher forcing; the notation 8.6×48.6\times4 means 8.68.6 GB on each of four GPUs.

    Could not parse LaTeX table

    For the W4-E4-A8 BERT-large model, sequential MREM is about 38×38\times faster than QAT, while parallel MREM is about 150×150\times faster and is approximately four times faster than sequential MREM. The parallel method is slightly less accurate than sequential MREM but remains close to QAT even at W2-E2-A4. Comparable BERT-base results show W4-E4-A8 matched accuracy of 83.5±0.1%83.5\pm0.1\% for MREM-S and 83.4±0.1%83.4\pm0.1\% for MREM-P, versus 84.6%84.6\% for QAT and 73.3±0.3%73.3\pm0.3\% for REM.

  6. Knowl 6 — SQuAD quantization results

    data/table

    On SQuAD v1.1, MREM likewise preserves exact-match (EM) and F1 performance far better than REM with a small calibration set. The following BERT-large results use the same quantization configurations as the MNLI comparison; EM and F1 are percentages.

    Could not parse LaTeX table

    MREM is therefore close to QAT on both SQuAD metrics while requiring only 44K calibration examples rather than the 8888K task training set used by QAT. Parallel MREM is slightly better than sequential MREM for the W2-E2-A4 BERT-large model, improving EM from 81.4±0.3%81.4\pm0.3\% to 81.8±0.3%81.8\pm0.3\% and F1 from 89.4±0.2%89.4\pm0.2\% to 89.6±0.2%89.6\pm0.2\%.

  7. Knowl 7 — Performance against published GLUE quantizers

    empirical result

    Across the GLUE development tasks, MREM-S and MREM-P outperform the evaluated post-training quantization baselines in most configurations and approach the full-precision model. The table reports model size, MNLI matched accuracy, and average GLUE score; the average is over the eight GLUE tasks used in the study.

    Could not parse LaTeX table

    At W2-E2-A4, MREM preserves substantially more GLUE accuracy than REM, especially on MNLI. MREM-S and MREM-P also remain close to the non-PTQ TernaryBERT result despite using post-training calibration, and their W4-E4-A8 MNLI accuracies are comparable to the reported Q-BERT results.

  8. Knowl 8 — Teacher forcing consistently improves parallel MREM

    empirical result

    An ablation on MNLI matched accuracy shows that annealed teacher forcing improves parallel MREM for both BERT-base and BERT-large, with larger gains when training is short or quantization is more aggressive. Values are mean ±\pm standard deviation over calibration-set choices.

    Could not parse LaTeX table

    For W2-E2-A4 and 250 steps, teacher forcing raises accuracy by 3.43.4 percentage points for BERT-base and 2.82.8 points for BERT-large. The benefit becomes smaller as training continues, motivating the default use of 2,0002{,}000 steps and teacher forcing during the first 40%40\% of training. The loss behavior also shows that later modules benefit more because they receive errors accumulated from more predecessor modules.

  9. Knowl 9 — Partition count, calibration size, and error propagation

    empirical result

    MREM exhibits a memory–accuracy trade-off controlled by the number of modules. For BERT-base on MNLI with W2-E2-A4 quantization, using 11, 22, 33, 44, or 66 modules requires respectively 11.911.9, 6.76.7, 4.74.7, 3.73.7, or 2.72.7 GB of running memory. Increasing the number of partitions slightly lowers accuracy because fewer layer dependencies are optimized jointly; the paper therefore uses four modules by default.

    The calibration-size study evaluates ∣D~∣∈{32,64,128,512,1024,2048,4096,8192}|\tilde D|\in\{32,64,128,512,1024,2048,4096,8192\}. REM is better than MREM-S and MREM-P below 128128 examples, but its accuracy rises slowly and saturates near 60%60\%. MREM benefits more from additional calibration data because its module-wise objective has greater optimization capacity, with diminishing improvement beyond 4,0964{,}096 examples. This supports the default calibration size of 4,0964{,}096.

    For both W2-E2-A8 and W2-E2-A4 BERT-base models, layerwise visualizations show that MREM has lower reconstruction error and a slower increase in error across layers than REM. Error generally increases through approximately the first ten layers and then decreases in later layers; the paper attributes the latter pattern to the task classification head encouraging more concentrated hidden representations.

  10. Knowl 10 — Per-channel quantization gives little additional benefit to MREM

    limitation

    The paper evaluates per-channel, or row-wise, quantization by assigning a separate quantization step size to each output dimension of every linear layer. On BERT-base MNLI, per-channel quantization improves REM by roughly 1.01.0–2.52.5 percentage points, but changes MREM only marginally. For W4-E4-A8, REM matched/mismatched accuracy changes from 73.3±0.3%/74.9±0.2%73.3\pm0.3\%/74.9\pm0.2\% to 75.9±0.3%/77.4±0.2%75.9\pm0.3\%/77.4\pm0.2\%, whereas MREM changes from 83.5±0.1%/83.9±0.2%83.5\pm0.1\%/83.9\pm0.2\% to 83.6±0.1%/84.0±0.1%83.6\pm0.1\%/84.0\pm0.1\%. For W2-E2-A4, MREM remains 81.1±0.2%/81.5±0.2%81.1\pm0.2\%/81.5\pm0.2\% without per-channel quantization and 81.1±0.2%/81.5±0.3%81.1\pm0.2\%/81.5\pm0.3\% with it. Because per-channel quantization requires storing more full-precision step sizes while providing little extra MREM accuracy, it is not used by default. More generally, increasing the number of modules lowers memory but can slightly reduce reconstruction quality, so the method does not remove the need to choose a partition appropriate to the available hardware.

Coverage note — No substantial contributed material was omitted; related work, generic quantization background, acknowledgements, and the checklist were excluded as non-contributory.

References

  1. 1.Mindspore. https://www.mindspore.cn.
  2. 2.Haoli Bai, Jiaxiang Wu, Irwin King, and Michael Lyu. Few shot network compression via cross distillation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 3203–3210, 2020.
  3. 3.Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. Binarybert: Pushing the limit of bert quantization. In Annual Meeting of the Association for Computational Linguistics, 2021.
  4. 4.Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. Preprint arXiv:1308.3432, 2013.
  5. 5.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems, 2020.
  6. 6.Shangyu Chen, Wenya Wang, and Sinno Jialin Pan. Deep neural network quantization via layer-wise optimization using limited training data. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 3329–3336, 2019.
  7. 7.Yankai Chen, Huifeng Guo, Yingxue Zhang, Chen Ma, Ruiming Tang, Jingjie Li, and Irwin King. Learning binarized graph representations with multi-faceted quantization reinforcement for top-k recommendation. In Proceedings of the 28th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2022.
  8. 8.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in neural information processing systems, 2015.
  9. 9.Jeffrey Dean and Sanjay Ghemawat. Mapreduce: simplified data processing on large clusters. In Communications of the ACM, volume 51, pages 107–113, 2008.
  10. 10.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. In International Conference on Learning Representations, 2019.
  11. 11.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. Preprint, 2022.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidi­rectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics, 2019.
  13. 13.Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha. Learned step size quantization. In International Conference on Learning Representations, 2019.
  14. 14.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. In International Conference on Learning Representations, 2019.
  15. 15.Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Remi Gribonval, Herve Jegou, and Armand Joulin. Training with quantization noise for extreme fixed-point compression. Preprint arXiv:2004.07320, 2020.
  16. 16.Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz, David Thorsley, Georgios Georgiadis, and Joseph H Hassoun. Post-training piecewise linear quantization for deep neural networks. In European Conference on Computer Vision, pages 69–86, 2020.
  17. 17.Qirong Ho, James Cipar, Henggang Cui, Jin Kyu Kim, Seunghak Lee, Phillip B Gibbons, Garth A Gibson, Gregory R Ganger, and Eric P Xing. More effective distributed ml via a stale synchronous parallel parameter server. In Advances in Neural Information Processing Systems, page 1223, 2013.
  18. 18.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. Dynabert: Dynamic bert with adaptive width and depth. In Advances in Neural Information Processing Systems, 2020.
  19. 19.Lu Hou and James T Kwok. Loss-aware weight quantization of deep networks. In International Conference on Learning Representations, 2018.
  20. 20.Lu Hou, Quanming Yao, and James T Kwok. Loss-aware binarization of deep networks. In International Conference on Learning Representations, 2017.
  21. 21.Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in neural information processing systems, 2018.
  22. 22.Zhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao. Chen, and Qun Liu. Ghostbert: Generate more features with cheap operations for bert. In Annual Meeting of the Association for Computational Linguistics, 2021.
  23. 23.Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, and Daniel Soudry. Improving post training neural quantization: Layer-wise calibration and integer programming. In Proceedings of the International Conference on Machine Learning, 2021.
  24. 24.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. In Findings of Empirical Methods in Natural Language Processing, 2020.
  25. 25.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations, 2020.
  26. 26.Fengfu Li, Bo Zhang, and Bin Liu. Ternary weight networks. Preprint arXiv:1605.04711, 2016.
  27. 27.Mu Li, David G Andersen, Alexander J Smola, and Kai Yu. Communication efficient distributed machine learning with the parameter server. In Advances in Neural Information Processing Systems, volume 27, pages 19–27, 2014.
  28. 28.Yuhang Li, Xin Dong, Sai Qian Zhang, Haoli Bai, Yuanpeng Chen, and Wei Wang. Rtn: Reparameterized ternary network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 4780–4787, 2020.
  29. 29.Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu. Brecq: Pushing the limit of post-training quantization by block reconstruction. In International Conference on Learning Representations, 2021.
  30. 30.Zhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang, Siwei Ma, and Wen Gao. Post-training quantization for vision transformer. Advances in Neural Information Processing Systems, 34, 2021.
  31. 31.I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018.
  32. 32.Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, 2019.
  33. 33.Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In Proceedings of the International Conference on Machine Learning, pages 7197–7206, 2020.
  34. 34.Markus Nagel, Mart van Baalen, Tijmen Blankevoort, and Max Welling. Data-free quantization through weight equalization and bias correction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1325–1334, 2019.
  35. 35.Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alex M Bronstein, and Avi Mendelson. Loss aware post-training quantization. Preprint arXiv:1911.07190, 2019.
  36. 36.Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. Pipedream: generalized pipeline parallelism for dnn training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles, pages 1–15, 2019.
  37. 37.Gunho Park, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee. nuqmm: Quantized matmul for efficient inference of large-scale generative language models. Preprint, 2022.
  38. 38.Haotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan, Aishan Liu, Qingqing Dang, Ziwei Liu, and Xi­anglong Liu. Bibert: Accurate fully binarized bert. International Conference on Learning Representations, 2022.
  39. 39.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. Preprint arXiv:1606.05250, 2016.
  40. 40.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. Preprint arXiv:1910.01108, 2019.
  41. 41.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  42. 42.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. Patient knowledge distillation for bert model compression. In Conference on Empirical Methods in Natural Language Processing, 2019.
  43. 43.Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong. Compression of generative pre-trained language models via quantization. Annual Meeting of the Association for Computational Linguistics, 2022.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  45. 45.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. Preprint arXiv:1804.07461, 2018.
  46. 46.Jiaxing Wang, Haoli Bai, Jiaxiang Wu, Xupeng Shi, Junzhou Huang, Irwin King, Michael Lyu, and Jian Cheng. Revisiting parameter sharing for automatic neural channel number search. In Advances in Neural Information Processing Systems, volume 33, 2020.
  47. 47.Peisong Wang, Qiang Chen, Xiangyu He, and Jian Cheng. Towards accurate post-training network quantization via bit-split and stitching. In Proceedings of the International Conference on Machine Learning, pages 9847–9856, 2020.
  48. 48.Xiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu, and Fengwei Yu. Qdrop: Randomly dropping quantization for extremely low-bit post-training quantization. In International Conference on Learning Representations, 2021.
  49. 49.Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. Outlier suppression: Pushing the limit of low-bit transformer language models. arXiv:2209.13325, 2022.
  50. 50.Liangjiang Wen, Xuanyang Zhang, Haoli Bai, and Zenglin Xu. Structured pruning of recurrent neural networks through neuron selection. Neural Networks, pages 134–141, 2020.
  51. 51.Ronald J Williams and David Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989.
  52. 52.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Annual Meeting of the Association for Computational Linguistics, 2020.
  53. 53.Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zero­quant: Efficient and affordable post-training quantization for large-scale transformers. arXiv:2206.01861, 2022.
  54. 54.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. Gobo: Quantizing attention­based nlp models for low latency and energy efficient inference. Preprint arXiv:2005.03842, 2020.
  55. 55.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. Q8bert: Quantized 8bit bert. Preprint arXiv:1910.06188, 2019.
  56. 56.Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. Ternarybert: Distillation-aware ultra-low bit bert. In Conference on Empirical Methods in Natural Language Processing, 2020.
  57. 57.Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Chris De Sa, and Zhiru Zhang. Improving neural network quantization without retraining using outlier channel splitting. In Proceedings of the International Conference on Machine Learning, 2019.
  58. 58.Denny Zhou, Mao Ye, Chen Chen, Tianjian Meng, Mingxing Tan, Xiaodan Song, Quoc Le, Qiang Liu, and Dale Schuurmans. Go wide, then narrow: Efficient training of deep thin networks. In Proceedings of the International Conference on Machine Learning, pages 11546–11555, 2020.
  59. 59.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. Bert loses patience: Fast and robust inference with early exit. In Advances in Neural Information Processing Systems, 2020.

Citation

MLA
Bai, H., et al. “Towards Efficient Post-training Quantization of Pre-trained Language Models”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 1405–18, https://proceedings.neurips.cc/paper_files/paper/2022/file/096347b4efc264ae7f07742fea34af1f-Paper-Conference.pdf.
APA
Bai, H., Hou, L., Shang, L., Jiang, X., King, I., & Lyu, M. R. (2022). Towards Efficient Post-training Quantization of Pre-trained Language Models. Advances in Neural Information Processing Systems, 35, 1405–1418. https://proceedings.neurips.cc/paper_files/paper/2022/file/096347b4efc264ae7f07742fea34af1f-Paper-Conference.pdf
Chicago
Bai, H., L. Hou, L. Shang, X. Jiang, I. King, and M. R. Lyu. 2022. “Towards Efficient Post-training Quantization of Pre-trained Language Models”. Advances in Neural Information Processing Systems 35: 1405–18. https://proceedings.neurips.cc/paper_files/paper/2022/file/096347b4efc264ae7f07742fea34af1f-Paper-Conference.pdf.
Harvard
Bai, H. et al. (2022) “Towards Efficient Post-training Quantization of Pre-trained Language Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 1405–1418. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/096347b4efc264ae7f07742fea34af1f-Paper-Conference.pdf.
Vancouver
1. Bai H, Hou L, Shang L, Jiang X, King I, Lyu MR (2022) Towards Efficient Post-training Quantization of Pre-trained Language Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 1405–1418

BibTeX

@inproceedings{bai2022towards,
  title = {Towards Efficient Post-training Quantization of Pre-trained Language Models},
  author = {Bai, Haoli and Hou, Lu and Shang, Lifeng and Jiang, Xin and King, Irwin and Lyu, Michael R},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {1405-1418},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/096347b4efc264ae7f07742fea34af1f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors