Compression of Generative Pre-trained Language Models via Quantization

Chaofan TaoLu HouWei ZhangLifeng ShangXin JiangQun LiuPing LuoNgai Wong

article2022ACL119 citations

Proposes a token-level contrastive distillation framework and module-wise dynamic scaling to effectively quantize generative pre-trained language models like GPT-2 and BART down to low-bit weights while preserving generation quality.

Listen

Large generative language models deliver strong performance across complex natural language tasks, but their enormous memory requirements and computational overhead create major bottlenecks for deployment and cost-effective scaling. Compressing these models through low-bit quantization—converting standard numerical representations into compact lower-bit formats—has historically worked well on classification systems but has consistently failed on text generation models due to severe output degradation.

The article designs and evaluates a tailored quantization framework for generative models to drastically reduce memory usage and parameter size while preserving output quality. It specifically addresses why conventional compression methods degrade generative architectures and introduces mechanisms that stabilize performance at very low numerical precision.

The investigation combines root-cause diagnostic analysis with empirical evaluations on established benchmarks across language modeling (WikiText2, Penn Treebank, WikiText103), conversational prediction (Persona-Chat), and text summarization (XSum). The authors analyze popular generative architectures, including GPT-2 and BART, across 8-bit, 4-bit, and 2-bit weight precisions while maintaining 8-bit activations. The compression framework incorporates two targeted solutions: a token-level contrastive distillation technique that pairs student representations with full-precision teacher tokens using an efficient momentum memory bank, and a module-wise dynamic scaling mechanism that adapts clipping thresholds to module-specific weight distributions.

The primary findings show that the failure of standard compression in generative systems stems from two factors: compressed word representations collapse into undifferentiated clusters (homogeneity), and quantization errors compound across sequential left-to-right generation. When using the proposed framework, compressed GPT-2 and BART models achieve 13.4x to 14.4x reductions in model size at 2-bit weight precision while maintaining output quality close to original full-precision baselines. In benchmark testing, 8-bit and 4-bit compressed models closely match original performance, and 2-bit models exhibit only slight degradation (an average 2-point increase in language modeling perplexity). In contrast, conventional methods like PACT and LSQ largely collapse at 2 bits, yielding repetitive or ungrammatical text. Ablation results further confirm that applying contrastive distillation to the final decoder states significantly outperforms sequence-level or intermediate-layer alternatives with minimal training overhead.

These results provide a clear pathway to slash server memory footprints, reduce hardware hosting costs, and enable broader deployment of high-performing generative text systems on resource-constrained infrastructure. Because the framework resolves parameter instability during extreme compression, organizations can achieve high-ratio model compression without rebuilding architectural foundations from scratch.

Decision-makers should consider adopting module-adaptive dynamic scaling and token-level contrastive distillation pipelines when developing or deploying edge and on-premise generative models. For immediate production systems requiring minimal risk, 8-bit or 4-bit quantization yields strong compression with virtually no performance penalty; aggressive 2-bit deployments can be considered where memory constraints are severe and minor perplexity trade-offs are acceptable. Future initiatives should pilot this compression methodology on larger modern generative architectures, measure real-world inference speedups on specialized hardware runtimes, and establish broader testing across domain-specific applications.

arXiv: 2203.10705
Cover for Compression of Generative Pre-trained Language Models via Quantization

Abstract

The increasing size of generative Pre-trained Language Models (PLMs) have greatly increased the demand for model compression. Despite various methods to compress BERT or its variants, there are few attempts to compress generative PLMs, and the underlying difficulty remains unclear. In this paper, we compress generative PLMs by quantization. We find that previous quantization methods fail on generative tasks due to the homogeneous word embeddings caused by reduced capacity, and varied distribution of weights. Correspondingly, we propose a token-level contrastive distillation to learn distinguishable word embeddings, and a module-wise dynamic scaling to make quantizers adaptive to different modules. Empirical results on various tasks show that our proposed method outperforms the state-of-the-art compression methods on generative PLMs by a clear margin. With comparable performance with the full-precision models, we achieve 14.4× and 13.4× compression rates on GPT-2 and BART, respectively.

Table of Contents

  • 1 Introduction
  • 2 Difficulty of Qunatizing Generative Pre-trained Language Models
  • 2.1 Network Quantization
  • 2.2 Difficulty Analysis
  • 3 Proposed Method
  • 3.1 Token-level Contrastive Distillation
  • 3.2 Module-dependent Dynamic Scaling
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Language Modeling
  • 4.3 Next Utterance Prediction
  • 4.4 Abstractive Summarization
  • 5 Discussion
  • 5.1 Ablation on Contrastive Learning
  • 5.1.1 Choices of Negative Sampling
  • 5.1.2 Number of Negative Samples
  • 5.1.3 Training Cost of the Contrastive Loss
  • 5.1.4 Representations for the Contrastive Loss
  • 5.2 Ablation on Dynamic Scaling
  • 6 Related Work
  • 7 Conclusion
  • Acknowledgements
  • References
  • A Derivation of Gradient of Dynamic Scaling
  • B More Experimental Settings
  • B.1 Datasets
  • B.2 Model Architectures
  • B.3 Hyperparameters
  • B.4 Description of the Compared Methods
  • B.5 Frameworks of Double-head GPT-2 and BART
  • C More Experimental Results
  • C.1 Performance of Larger Models
  • C.2 Examples of Summarizations
  • C.3 More Visualizations for the Token Representations

Knowls

  1. Knowl 1 — Why generative language models are difficult to quantize

    empirical result

    The paper identifies two empirical failure modes specific to low-bit quantization of generative pre-trained language models (PLMs). First, quantized models learn homogeneous word embeddings: in the 2-bit PTB experiments, the full-precision GPT-2 embeddings are dispersed and distinguishable, whereas PACT, LSQ, and LAQ produce clustered embeddings. The T-SNE visualizations on page 3 show that the proposed method preserves a more dispersed embedding geometry. Token-representation cosine-similarity matrices on pages 3 and 16 further show that the proposed method preserves both same-token alignment and contextual dependencies between different tokens, while the baseline quantizers lose most off-diagonal structure.

    Second, full-precision Transformer weights have highly skewed, module-dependent distributions with outliers. A single poorly estimated clipping range therefore either wastes quantization levels on outliers or introduces large errors for the majority of weights. The problem is amplified in autoregressive generation because errors from earlier tokens propagate to later tokens. The paper presents the explanation that this sequential error accumulation makes the learning signal noisier over time, but treats the causal explanation for homogeneous embeddings as a conjecture rather than a formal result.

  2. Knowl 2 — Quantization-aware training used for generative PLMs

    model/method

    The quantized student is trained with quantization-aware training. For a full-precision weight vector w∈RMw\in\mathbb{R}^{M}, a positive clipping factor α>0\alpha>0, and weight bit-width bb, each forward pass computes

    wq=α Q ⁣(clip⁡(w,−α,α)α),w_q=\alpha\,Q\!\left(\frac{\operatorname{clip}(w,-\alpha,\alpha)}{\alpha}\right),

    where QQ maps each scalar to its nearest element of the uniform set

    {−1,−r−1r,…,−1r,0,1r,…,r−1r,1},r=2b−1−1.\left\{-1,-\frac{r-1}{r},\ldots,-\frac{1}{r},0,\frac{1}{r},\ldots,\frac{r-1}{r},1\right\},\qquad r=2^{b-1}-1.

    The loss is evaluated using wqw_q, while the gradient with respect to the quantized weight is passed through the non-differentiable quantizer with a straight-through estimator to update the underlying full-precision weight ww. The embedding matrix uses one clipping factor per word row, and Transformer weight matrices use one clipping factor per matrix. Activations after self-attention and GeLU use asymmetric uniform quantization because their values are mostly positive; other activations use symmetric uniform quantization. Layer normalization, skip connections, and biases remain full precision.

  3. Knowl 3 — Token-level contrastive distillation

    model/method

    Token-level contrastive distillation trains a quantized student to match the full-precision teacher at each token while distinguishing that token from other tokens in the same sequence. Let (t1,…,tN)(t_1,\ldots,t_N) be an input sequence of NN tokens. For token position ii, let his,hit∈Rdh_i^s,h_i^t\in\mathbb{R}^{d} be linearly projected hidden states from the student and teacher, respectively, and let qis∈Rdq_i^s\in\mathbb{R}^{d} be the momentum-smoothed student representation stored in a token memory bank. Let AiA_i contain the positive position ii and sampled negative token positions from the same sequence. With cosine similarity s(x,y)=x⊤y/(∥x∥2∥y∥2)s(x,y)=x^\top y/(\lVert x\rVert_2\lVert y\rVert_2) and temperature τ>0\tau>0, the contrastive loss is

    Lcont=−∑i=1Nlog⁡exp⁡ ⁣(s(qis,hit)/τ)∑j∈Aiexp⁡ ⁣(s(qis,hjt)/τ).\mathcal{L}_{\mathrm{cont}}=-\sum_{i=1}^{N}\log\frac{\exp\!\left(s(q_i^s,h_i^t)/\tau\right)}{\sum_{j\in A_i}\exp\!\left(s(q_i^s,h_j^t)/\tau\right)}.

    The memory-bank representation is updated by

    qis←mqis+(1−m)his,q_i^s\leftarrow m q_i^s+(1-m)h_i^s,

    where m∈[0,1)m\in[0,1) is the momentum coefficient. In addition, zis,zit∈R∣V∣z_i^s,z_i^t\in\mathbb{R}^{|V|} are the softmax-normalized student and teacher distributions over a vocabulary of size ∣V∣|V|, and the logit-distillation loss is

    Ldist=−∑i=1N∑v=1∣V∣zi,vtlog⁡zi,vs.\mathcal{L}_{\mathrm{dist}}=-\sum_{i=1}^{N}\sum_{v=1}^{|V|}z^t_{i,v}\log z^s_{i,v}.

    The total loss is L=λLcont+Ldist\mathcal{L}=\lambda\mathcal{L}_{\mathrm{cont}}+\mathcal{L}_{\mathrm{dist}}, with default λ=0.1\lambda=0.1. The positive is the teacher representation of the same token, while negatives are teacher representations of other positions in the sequence; student representations in the memory bank make the query side smoother. The loss is applied to the last Transformer layer of GPT-2 or the decoder of BART and is computed over all autoregressively processed tokens.

  4. Knowl 4 — Module-dependent dynamic scaling and clipping-gradient estimator

    model/method

    To adapt quantization to the different weight distributions of individual Transformer modules, the proposed method learns a dimensionless scaling factor γ\gamma for each weight matrix rather than learning its clipping factor directly. For a weight matrix WW containing MM scalar weights, the clipping factor is

    α=γ∥W∥1M,\alpha=\gamma\frac{\lVert W\rVert_1}{M},

    where ∥W∥1\lVert W\rVert_1 is the sum of absolute weight magnitudes. Each γ\gamma is initialized to 11, so the initial clipping range is tied to the module's average weight magnitude rather than to an arbitrary fixed scale.

    For a scalar weight wkw_k in WW, let uk=clip⁡(wk,−α,α)/αu_k=\operatorname{clip}(w_k,-\alpha,\alpha)/\alpha, let wq,k=αQ(uk)w_{q,k}=\alpha Q(u_k), and let ℓ\ell be the training loss. Using a straight-through estimator for QQ, the contribution of wkw_k to the scaling-factor gradient is estimated as

    ∂ℓ∂γ={∂ℓ∂wq,kQ(uk)∥W∥1M,wk<−α,∂ℓ∂wq,k[−wkα+Q(uk)]∥W∥1M,−α≤wk≤α,∂ℓ∂wq,kQ(uk)∥W∥1M,wk>α.\frac{\partial\ell}{\partial\gamma}=\begin{cases} \displaystyle \frac{\partial\ell}{\partial w_{q,k}}Q(u_k)\frac{\lVert W\rVert_1}{M}, & w_k< -\alpha,\\[6pt] \displaystyle \frac{\partial\ell}{\partial w_{q,k}}\left[-\frac{w_k}{\alpha}+Q(u_k)\right]\frac{\lVert W\rVert_1}{M}, & -\alpha\leq w_k\leq\alpha,\\[6pt] \displaystyle \frac{\partial\ell}{\partial w_{q,k}}Q(u_k)\frac{\lVert W\rVert_1}{M}, & w_k>\alpha. \end{cases}

    Unlike the PACT estimator, the middle case accounts for weights inside the clipping interval as well as weights outside it. Consequently, the learned clipping factor responds to both in-range quantization resolution and outlier clipping error, while the normalization by average weight magnitude reduces sensitivity to the scale of each module.

  5. Knowl 5 — Experimental protocol for QuantGPT and QuantBART

    experimental setup

    The proposed quantization method is evaluated on GPT-2 for language modeling and next-utterance prediction, and on BART for abstractive summarization. QuantGPT uses GPT-2-small with 12 decoder layers, hidden dimension 768, vocabulary size 50,527, and GeLU activations. QuantBART uses BART-base with 6 encoder layers, 6 decoder layers, hidden dimension 768, and vocabulary size 50,265. A full-precision model is first fine-tuned from a Hugging Face pretrained checkpoint; that model acts as the teacher and initializes the quantized student.

    Language modeling uses WikiText2, Penn Treebank (PTB), and WikiText103, evaluated by perplexity. Next-utterance prediction uses Persona-Chat and is evaluated by accuracy. Abstractive summarization uses XSum and is evaluated by ROUGE-1, ROUGE-2, and ROUGE-L. The main comparisons use weight, embedding, and activation bit-widths denoted as W-E-A, with settings 8-8-8, 4-4-8, and 2-2-8.

    For GPT-2 language modeling and next-utterance prediction, the sequence length is 512, the backbone and scaling-factor learning rates start at 0.00050.0005 and 0.0010.001, respectively, and both decay linearly to zero. The contrastive temperature is τ=0.1\tau=0.1, the memory momentum is m=0.5m=0.5, and AdamW is used. Language-modeling batches contain 32 examples, with 80, 120, and 8 epochs for WikiText2, PTB, and WikiText103; next-utterance prediction uses batch size 16 for 2 epochs. The number of negatives per sequence is 64 for PTB and 32 for WikiText2, WikiText103, and Persona-Chat. For XSum, source sequences have length 512, summaries are padded to the maximum target length, beam search uses beam size 6 and length penalty 1, the backbone learning rate starts at 0.00020.0002, and training uses batch size 128 for 8 epochs. Training uses 8 V100 GPUs.

  6. Knowl 6 — GPT-2 language-modeling and compression results

    data/table

    QuantGPT is compared with full-precision GPT-2 and the quantizers PACT, LSQ, and LAQ. Perplexity (PPL) is lower-is-better, and Persona-Chat accuracy is higher-is-better. W-E-A denotes Transformer-weight, word-embedding, and activation bit-widths.

    Method W-E-A Size (MB) WikiText2 PPL PTB PPL WikiText103 PPL Persona-Chat Acc. (%)
    full-prec. – 474.9 14.48 14.72 14.19 77.01
    PACT 8-8-8 121.4 17.49 16.11 16.76 74.73
    LSQ 8-8-8 121.4 16.75 15.43 15.24 75.28
    LAQ 8-8-8 121.4 16.91 15.87 15.88 76.02
    QuantGPT 8-8-8 121.4 15.31 14.90 14.58 76.12
    PACT 4-4-8 62.4 19.23 20.17 20.15 25.13
    LSQ 4-4-8 62.4 78.99 79.76 75.12 45.10
    LAQ 4-4-8 62.4 17.12 16.55 16.91 71.71
    QuantGPT 4-4-8 62.4 15.55 14.95 15.31 76.57
    PACT 2-2-8 33.0 173.02 189.13 171.03 5.52
    LSQ 2-2-8 33.0 847.54 544.98 1470.86 5.54
    LAQ 2-2-8 33.0 19.15 18.25 18.97 71.36
    QuantGPT 2-2-8 33.0 17.30 16.12 16.98 74.78

    QuantGPT is the best quantized method at every displayed bit-width and task. At 8-bit weights its perplexities are close to full precision; at 4-bit weights the degradation is about 1 PPL point on WikiText2 and WikiText103 and less than 0.1 on PTB relative to the 8-bit QuantGPT result. At 2-bit weights, QuantGPT has an average drop of about 2 PPL points from the full-precision baseline while reducing the model from 474.9 MB to 33.0 MB, a 14.4× reduction. PACT and LSQ largely fail at 2 bits, while LSQ also deteriorates sharply at 4 bits.

    Against other GPT-2 compression methods, the paper reports the following comparison:

    Method Size (MB) Compression WikiText2 PPL PTB PPL WikiText103 PPL
    full-prec. 474.9 1.0x 14.4 14.6 13.9
    KnGPT2 332.0 1.4x - - 20.5
    DistilGPT2 329.6 1.4x - - 21.1
    LightPAFF 268.0 1.8x 18.8 22.8 16.4
    QuantGPT (8-8-8) 121.4 3.9x 15.3 14.9 14.6
    QuantGPT (4-4-8) 62.4 7.6x 15.6 15.0 15.3
    QuantGPT (2-2-8) 33.0 14.4x 17.3 16.1 17.0

    These results show that QuantGPT provides a better size-performance trade-off than the compared tensor-decomposition and distillation methods, including at 2-bit weights.

  7. Knowl 7 — BART abstractive-summarization results

    data/table

    On XSum, QuantBART is compared with full-precision BART and PACT, LSQ, and LAQ. The model-size column reports megabytes; ROUGE-1, ROUGE-2, and ROUGE-L are higher-is-better.

    Method W-E-A Size (MB) ROUGE-1 ROUGE-2 ROUGE-L
    full-prec. – 532.0 40.75 18.10 33.05
    PACT 8-8-8 138.1 39.16 16.60 31.60
    LSQ 8-8-8 138.1 39.09 16.72 31.56
    LAQ 8-8-8 138.1 39.10 16.74 31.65
    QuantBART 8-8-8 138.1 40.25 17.78 32.70
    PACT 4-4-8 72.4 32.68 11.52 26.03
    LSQ 4-4-8 72.4 38.94 16.48 31.46
    LAQ 4-4-8 72.4 39.03 16.68 31.63
    QuantBART 4-4-8 72.4 40.24 17.71 32.69
    PACT 2-2-8 39.6 7.76 1.30 6.96
    LSQ 2-2-8 39.6 37.09 14.88 29.76
    LAQ 2-2-8 39.6 37.48 15.27 30.13
    QuantBART 2-2-8 39.6 39.15 16.72 31.72

    QuantBART consistently outperforms the other quantized methods. At 8-bit and 4-bit weights it remains close to full-precision BART while reducing the model from 532.0 MB to 138.1 MB and 72.4 MB, respectively. At 2-bit weights it achieves ROUGE-1 39.15, ROUGE-2 16.72, and ROUGE-L 31.72 with a 39.6 MB model, whereas PACT collapses to ROUGE-1 7.76 and ROUGE-2 1.30. Qualitative XSum outputs also show that QuantBART produces logical, terse summaries, while PACT often repeats text.

  8. Knowl 8 — Ablations establish that token-level, teacher-based negatives are important

    empirical result

    The contrastive-loss ablations use 2-bit QuantGPT. The proposed token-level scheme uses full-precision teacher representations of other tokens in the same sequence as negatives. Alternatives use teacher and student representations together (fp+quan.), student representations only (quan. only), random vocabulary tokens (global), or sequence-level in-batch negatives. Lower PPL is better.

    Sampling method WikiText2 PPL PTB PPL WikiText103 PPL
    QuantGPT 17.30 16.12 16.98
    Token fp+quan. 17.38 16.51 17.13
    Token quan. only 17.35 16.54 17.15
    Token global 17.71 16.63 17.55
    Sequence in-batch, batch size 32 17.62 19.23 18.97
    Sequence in-batch, batch size 16 17.48 17.11 18.16

    Teacher representations from other positions in the same sequence perform best. Random vocabulary negatives are worse, which the paper attributes to their lack of contextual relation to the studied token. Sequence-level negatives are also worse, supporting the need for token-grained representations in autoregressive generation. Increasing the number of negatives improves and then gradually stabilizes PTB performance, and momentum-smoothed memory-bank representations outperform immediate student representations.

    The choice of representation is also evaluated:

    Representation WikiText2 PPL PTB PPL WikiText103 PPL Persona Acc. (%) ROUGE-1 ROUGE-2 ROUGE-L
    decoder-last 17.30 16.12 16.98 74.78 39.15 16.72 31.72
    decoder-first 18.02 16.61 17.25 74.75 39.11 16.70 31.62
    encoder-last - - - - 38.91 16.72 31.67
    encoder-first - - - - 38.87 16.70 31.56

    The last decoder layer is the strongest representation source among the tested alternatives. Adding the contrastive loss improves PTB PPL from 16.93 to 16.12, with training time increasing from 0.61 to 0.67 seconds per iteration and GPU memory increasing from 14,700 MB to 14,839 MB per device.

  9. Knowl 9 — Dynamic scaling and its gradient estimator are both necessary

    empirical result

    The learned scaling factors γ\gamma vary substantially across GPT-2 modules in the 2-bit experiment, empirically supporting module-wise rather than shared scaling. The clipping-gradient ablation compares full QuantGPT, a variant using only the logit-distillation loss Ldist\mathcal{L}_{\mathrm{dist}}, and a variant that uses the proposed dynamic scaling parameterization but the PACT gradient estimator, which ignores weights inside the clipping interval.

    Method WikiText2 PPL PTB PPL WikiText103 PPL
    QuantGPT 17.30 16.12 16.98
    Ldist\mathcal{L}_{\mathrm{dist}} only 17.85 16.93 17.78
    Ours with PACT gradient 20.03 17.78 25.54

    Removing token-level contrastive learning worsens all three language-modeling results, demonstrating that logit distillation alone does not sufficiently preserve distinguishable token representations. Replacing the proposed clipping-gradient estimator with PACT causes a much larger degradation, especially on WikiText103, showing that accounting for weights inside and outside the clipping range is important.

  10. Knowl 10 — Quantization remains effective on larger 24-layer models

    empirical result

    The paper evaluates the proposed method on larger 24-layer GPT-2 and BART models: GPT-2-base has 24 decoder layers and hidden dimension 1024, while BART-large has 12 encoder layers, 12 decoder layers, and hidden dimension 1024. The reported metrics are perplexity for WikiText2, PTB, and WikiText103 and ROUGE scores for XSum.

    Method W-E-A GPT Size (MB) W2 PPL PTB PPL W103 PPL BART Size (MB) R1 R2 RL
    full-prec. – 1353.7 12.46 12.35 12.37 1550.0 45.25 22.11 37.07
    PACT 8-8-8 342.5 12.86 13.95 13.90 394.8 43.55 20.57 35.55
    Ours 8-8-8 342.5 12.53 12.40 12.68 394.8 44.34 21.41 36.32
    PACT 4-4-8 174.0 16.10 14.19 18.07 202.2 19.45 3.53 15.58
    Ours 4-4-8 174.0 13.34 12.41 14.12 202.2 44.18 21.31 36.25
    PACT 2-2-8 89.7 98.74 68.55 86.60 106.0 8.53 0.93 7.25
    Ours 2-2-8 89.7 14.53 13.22 14.52 106.0 42.38 19.75 34.57

    The proposed method converges successfully at every displayed bit-width without gradient explosion or vanishing. It substantially outperforms PACT at 4-bit and 2-bit weights for both the larger GPT-2 and BART models; the advantage is particularly large for 2-bit BART, where the proposed method retains ROUGE-1 42.38 compared with PACT's 8.53.

Coverage note — The appendix's example summaries, additional token-similarity panels, dataset split counts, and implementation details of the comparison baselines were omitted because they are qualitative illustrations, redundant visualizations, or non-contributed setup details rather than additional load-bearing findings.

References

  1. 1.Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, volume 33.
  2. 2.Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. 2021. Binarybert: Pushing the limit of bert quantization. In Annual Meeting of the Association for Computational Linguistics.
  3. 3.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. Technical Report arXiv:1308.3432.
  4. 4.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems.
  5. 5.Cong Chen, Chaofan Tao, and Ngai Wong. 2021. Litegt: Efficient and lightweight graph transformers. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 161–170.
  6. 6.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607.
  7. 7.Xinlei Chen and Kaiming He. 2021. Exploring simple siamese representation learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758.
  8. 8.Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. 2018. Pact: Parameterized clipping activation for quantized neural networks. Preprint arXiv:1805.06085.
  9. 9.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in neural information processing systems, pages 3123–3131.
  10. 10.Ali Edalati, Marzieh Tahaei, Ahmad Rashid, Vahid Partovi Nia, James J Clark, and Mehdi Rezagholizadeh. 2021. Kronecker decomposition for gpt compression. In Advances in Neural Information Processing Systems.
  11. 11.Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha. 2020. Learned step size quantization. In International Conference on Learning Representations.
  12. 12.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent: A new approach to self-supervised learning. In Neural Information Processing Systems.
  13. 13.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738.
  14. 14.Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). Technical Report arXiv:1606.08415.
  15. 15.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. Technical Report arXiv:1503.02531.
  16. 16.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020. Dynabert: Dynamic bert with adaptive width and depth. In Advances in Neural Information Processing Systems, volume 33.
  17. 17.Lu Hou and James T. Kwok. 2018. Loss-aware weight quantization of deep networks. In International Conference on Learning Representations.
  18. 18.Lu Hou, Yao Quanming, and James T. Kwok. 2017. Loss-aware binarization of deep networks. In International Conference on Learning Representations.
  19. 19.Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang, and Qun Liu. 2022. Spiral: Self-supervised perturbation-invariant representation learning for speech pre-training. In International Conference on Learning Representations.
  20. 20.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning.
  21. 21.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. Tinybert: Distilling bert for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174.
  22. 22.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Annual Meeting of the Association for Computational Linguistics.
  24. 24.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  25. 25.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843.
  26. 26.Tomas Mikolov and Geoffrey Zweig. 2012. Context dependent recurrent neural network language model. In IEEE Spoken Language Technology Workshop, pages 234–239. IEEE.
  27. 27.Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Conference on Empirical Methods in Natural Language Processing, pages 1797–1807.
  28. 28.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning.
  29. 29.Alec Radford and Karthik Narasimhan. 2018. Improving language understanding by generative pre-training.
  30. 30.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  31. 31.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020. Q-bert: Hessian based ultra low precision quantization of bert. In AAAI Conference on Artificial Intelligence.
  32. 32.Kaitao Song, Hao Sun, Xu Tan, Tao Qin, Jianfeng Lu, Hongzhi Liu, and Tie-Yan Liu. 2020. Lightpaff: A two-stage distillation framework for pre-training and fine-tuning. Preprint arXiv:2004.12817.
  33. 33.Siqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng, Shuohang Wang, and Jingjing Liu. 2020a. Contrastive distillation on intermediate representations for language model compression. In Conference on Empirical Methods in Natural Language Processing, pages 498–508.
  34. 34.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020b. Mobilebert: a compact task-agnostic bert for resource-limited devices. In Annual Meeting of the Association for Computational Linguistics, pages 2158–2170.
  35. 35.Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019. Contrastive representation distillation. In International Conference on Learning Representations.
  36. 36.Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse attention architecture with cascade token and head pruning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 97–110. IEEE.
  37. 37.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR.
  38. 38.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020. Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 811–824. IEEE.
  39. 39.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019. Q8bert: Quantized 8bit bert. Preprint arXiv:1910.06188.
  40. 40.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213.
  41. 41.Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020. Ternarybert: Distillation-aware ultra-low bit bert. In Conference on Empirical Methods in Natural Language Processing.

Citation

MLA
Tao, C., et al. “Compression of Generative Pre-trained Language Models via Quantization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 4821–36, https://doi.org/10.18653/v1/2022.acl-long.331.
APA
Tao, C., Hou, L., Zhang, W., Shang, L., Jiang, X., Liu, Q., Luo, P., & Wong, N. (2022). Compression of Generative Pre-trained Language Models via Quantization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4821–4836. https://doi.org/10.18653/v1/2022.acl-long.331
Chicago
Tao, C., L. Hou, W. Zhang, et al. 2022. “Compression of Generative Pre-trained Language Models via Quantization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4821–36. https://doi.org/10.18653/v1/2022.acl-long.331.
Harvard
Tao, C. et al. (2022) “Compression of Generative Pre-trained Language Models via Quantization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4821–4836. Available at: https://doi.org/10.18653/v1/2022.acl-long.331.
Vancouver
1. Tao C, Hou L, Zhang W, Shang L, Jiang X, Liu Q, Luo P, Wong N (2022) Compression of Generative Pre-trained Language Models via Quantization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4821–4836

BibTeX

@inproceedings{tao-etal-2022-compression,
    title = "Compression of Generative Pre-trained Language Models via Quantization",
    author = "Tao, Chaofan  and
      Hou, Lu  and
      Zhang, Wei  and
      Shang, Lifeng  and
      Jiang, Xin  and
      Liu, Qun  and
      Luo, Ping  and
      Wong, Ngai",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.331/",
    doi = "10.18653/v1/2022.acl-long.331",
    pages = "4821--4836"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/