Extreme Compression of Large Language Models via Additive Quantization
Vage EgiazarianAndrei PanferovDenis KuznedelevElias FrantarArtem BabenkoDan Alistarh
Introduces AQLM, a multi-codebook quantization method that compresses large language model weights down to 2 bits per parameter while achieving state-of-the-art accuracy and matching 16-bit floating-point inference speeds on standard hardware.
Deploying modern large language models locally or on commodity hardware is challenging due to massive memory and compute requirements. While post-training quantization reduces model sizes by lowering parameter precision, traditional methods face severe accuracy degradation at extreme compression levels below 3 bits per parameter. Consequently, existing 2-bit models have historically underperformed smaller baseline models quantized to 3 or 4 bits, limiting their real-world utility.
The article demonstrates that multi-codebook vector quantization, a technique adapted from information retrieval, can achieve extreme post-training compression for large language models while preserving generation quality. The primary objective is to evaluate a novel compression algorithm, called Additive Quantization for Language Models (AQLM), across 2 to 4 bits per parameter.
The approach generalizes classical additive quantization into a three-stage optimization framework calibrated on model token activations. First, the method optimizes discrete codes for weight groups using beam search. Second, it continuously updates learned codebooks via standard optimization algorithms. Third, it fine-tunes parameters across multi-layer transformer blocks to maintain end-to-end output fidelity. The evaluation assessed open model families, including LLAMA 2 (7B, 13B, and 70B parameters) and Mixtral, against leading post-training quantization baselines using language modeling perplexity, multi-domain benchmarks, and execution speed on consumer hardware.
The findings establish that AQLM outperforms existing compression methods across 2 to 4 bits, achieving the largest accuracy gains in the extreme 2-bit regime. For the first time, Pareto optimality is demonstrated below 3 bits per parameter: starting at roughly 2.5 bits per parameter, a compressed 13B model outperforms a smaller 7B model of equivalent total byte size. When paired with end-to-end distillation fine-tuning, accuracy improves further, matching or exceeding competing approaches across zero-shot evaluations. Moreover, the homogeneous weight format enables high-performance inference, delivering up to 30% speedups on GPUs and up to fourfold speedups on CPUs compared to original precision implementations while shrinking memory footprint by up to eightfold.
These results show that organizations can deploy higher-capacity language models within strictly constrained hardware environments, substantially reducing operational hosting costs and memory transfer bottlenecks without sacrificing core capabilities. The algorithm shifts the practical threshold of low-bit model compression, making previously unfeasible sub-3-bit deployments viable for production systems.
Decision-makers should consider AQLM when hardware memory limits prevent standard 16-bit or 4-bit model deployments. If selecting codebook configurations, engineering teams should weigh the trade-off between higher-precision codebooks for maximum predictive accuracy and smaller codebooks for improved inference latency. Because the primary limitation of this method is high calibration compute time (such as requiring several days on multi-GPU setups for 70B parameter models), organizations should plan quantization workflows as offline preparation tasks before production rollout.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). GPTQ establishes how second-order, post-training quantization can preserve LLM quality at low bit widths, providing a key baseline for AQLM’s more aggressive additive-codebook approach.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). This survey frames the quantization trade-offs and hardware constraints that motivate AQLM’s focus on accurate, practical compression below three bits per parameter.
- Paper: FastText.zip: Compressing text classification models, Armand Joulin et al. (2016). FastText.zip applies product quantization to compress model representations, offering useful context for AQLM’s adaptation of codebook-based compression to LLM weights.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). QuIP# carries extreme LLM quantization into structured lattice codebooks, making it a direct continuation for comparing codebook design and accuracy at two- and three-bit rates.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). BitDistiller extends sub-four-bit LLM compression through quantization-aware training and self-distillation, complementing AQLM’s learned additive quantization with a training-based route to extreme precision.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). LQ-LoRA extends sub-three-bit compression to fine-tuning by combining quantized weights with trainable low-rank components under explicit memory budgets.
