Compression of Generative Pre-trained Language Models via Quantization
Chaofan TaoLu HouWei ZhangLifeng ShangXin JiangQun LiuPing LuoNgai Wong
Proposes a token-level contrastive distillation framework and module-wise dynamic scaling to effectively quantize generative pre-trained language models like GPT-2 and BART down to low-bit weights while preserving generation quality.
Large generative language models deliver strong performance across complex natural language tasks, but their enormous memory requirements and computational overhead create major bottlenecks for deployment and cost-effective scaling. Compressing these models through low-bit quantization—converting standard numerical representations into compact lower-bit formats—has historically worked well on classification systems but has consistently failed on text generation models due to severe output degradation.
The article designs and evaluates a tailored quantization framework for generative models to drastically reduce memory usage and parameter size while preserving output quality. It specifically addresses why conventional compression methods degrade generative architectures and introduces mechanisms that stabilize performance at very low numerical precision.
The investigation combines root-cause diagnostic analysis with empirical evaluations on established benchmarks across language modeling (WikiText2, Penn Treebank, WikiText103), conversational prediction (Persona-Chat), and text summarization (XSum). The authors analyze popular generative architectures, including GPT-2 and BART, across 8-bit, 4-bit, and 2-bit weight precisions while maintaining 8-bit activations. The compression framework incorporates two targeted solutions: a token-level contrastive distillation technique that pairs student representations with full-precision teacher tokens using an efficient momentum memory bank, and a module-wise dynamic scaling mechanism that adapts clipping thresholds to module-specific weight distributions.
The primary findings show that the failure of standard compression in generative systems stems from two factors: compressed word representations collapse into undifferentiated clusters (homogeneity), and quantization errors compound across sequential left-to-right generation. When using the proposed framework, compressed GPT-2 and BART models achieve 13.4x to 14.4x reductions in model size at 2-bit weight precision while maintaining output quality close to original full-precision baselines. In benchmark testing, 8-bit and 4-bit compressed models closely match original performance, and 2-bit models exhibit only slight degradation (an average 2-point increase in language modeling perplexity). In contrast, conventional methods like PACT and LSQ largely collapse at 2 bits, yielding repetitive or ungrammatical text. Ablation results further confirm that applying contrastive distillation to the final decoder states significantly outperforms sequence-level or intermediate-layer alternatives with minimal training overhead.
These results provide a clear pathway to slash server memory footprints, reduce hardware hosting costs, and enable broader deployment of high-performing generative text systems on resource-constrained infrastructure. Because the framework resolves parameter instability during extreme compression, organizations can achieve high-ratio model compression without rebuilding architectural foundations from scratch.
Decision-makers should consider adopting module-adaptive dynamic scaling and token-level contrastive distillation pipelines when developing or deploying edge and on-premise generative models. For immediate production systems requiring minimal risk, 8-bit or 4-bit quantization yields strong compression with virtually no performance penalty; aggressive 2-bit deployments can be considered where memory constraints are severe and minor perplexity trade-offs are acceptable. Future initiatives should pilot this compression methodology on larger modern generative architectures, measure real-world inference speedups on specialized hardware runtimes, and establish broader testing across domain-specific applications.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). Provides a comprehensive foundational survey of neural network quantization paradigms and low-bit inference challenges that contextualizes the need for generative model compression.
- Paper: MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers, Wenhui Wang et al. (2020). Introduces key knowledge distillation concepts across transformer layers that motivate token-level and representation-level distillation in compressed architectures.
- Paper: A Primer in BERTology: What We Know About How BERT Works, Anna Rogers et al. (2020). Surveys the internal representations, layer mechanics, and overparameterization of transformer architectures, establishing essential baseline concepts for understanding transformer compression.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). Extends distillation-guided extreme low-bit compression by pairing confidence-aware self-distillation with tailored weight clipping for sub-4-bit large language models.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). Applies scalable second-order post-training quantization to compress massive generative transformer models down to 3- and 4-bit precision efficiently.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). Addresses activation outlier challenges in generative architectures via mathematical transformations that redistribute quantization difficulty between weights and activations.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Develops activation-aware weight quantization to protect salient weights identified from generative activation patterns for low-bit LLM inference.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). Pushes low-bit weight quantization to 2- and 3-bit regimes using randomized transformations and lattice codebooks to preserve generative language model accuracy.
- Paper: BiLLM: Pushing the Limit of Post-Training Quantization for LLMs, Wei Huang et al. (2024). Pushes post-training quantization boundaries down to 1-bit per weight for large language models through structured residual binarization.
- Paper: Evaluating Quantized Large Language Models, Shiyao Li et al. (2024). Provides an extensive empirical framework evaluating how post-training quantization impacts downstream generative capabilities, reasoning tasks, and attention caches across diverse model scales.
- Paper: FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU, Ying Sheng et al. (2023). Integrates low-bit quantization of weights and key-value caches with structured memory offloading schedules to enable high-throughput generative inference on resource-constrained hardware.
- Paper: QLoRA: Efficient Finetuning of Quantized LLMs, Tim Dettmers et al. (2023). Combines 4-bit quantized base generative models with parameter-efficient fine-tuning via low-rank adapters to maintain full task fidelity under severe memory constraints.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). Investigates the downstream impacts of post-training quantization on the safety, ethics, and trustworthiness behaviors of compressed generative language models.
