BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Wei HuangYangdong LiuHaotong QinYing LiShiming ZhangXianglong LiuMichele MagnoXiaojuan Qi
Develops an ultra-low-bit post-training quantization method that compresses large language models down to roughly 1.08 bits per weight using binary residual approximation and optimal distribution splitting while maintaining high inference accuracy.
Large language models demonstrate powerful language processing capabilities but require massive memory and computational infrastructure. For example, running a 70-billion-parameter model in standard half-precision format requires approximately 150 gigabytes of memory, demanding multiple high-end enterprise graphics processors. While post-training quantization compresses model weights to lower memory footprints without expensive retraining, conventional techniques suffer severe performance degradation or total output collapse when pushed to ultra-low bit-widths of two bits or fewer.
The article introduces and evaluates BiLLM, a novel post-training binarization framework designed to compress pretrained language models down to approximately one bit per weight while preserving linguistic accuracy.
The authors conducted an empirical analysis of weight distributions across several model families, identifying that a small fraction of structured parameters carries outsized sensitivity while the remaining parameters follow a bell-shaped distribution. Using these insights, the method structurally selects critical weights to approximate them through a two-step binary residual technique and partitions the remaining weights into concentrated and sparse groups via an optimal break-point search. The framework was evaluated across the OPT, LLaMA, LLaMA 2, and Vicuna model families spanning sizes from 1.3 billion to 70 billion parameters across standard language benchmarks and zero-shot reasoning tasks on a single graphics processing unit.
The evaluation produced four major findings. First, the proposed method achieved average weight precisions between 1.07 and 1.11 bits across evaluated models without experiencing performance collapse. Second, on large models such as LLaMA 2-70B, the method achieved an inference perplexity of 8.41 at 1.08 bits, outperforming the full 16-bit precision version of OPT-66B. Third, the framework delivered nearly a tenfold reduction in model storage footprint, shrinking LLaMA 2-70B from 129.3 gigabytes to 15.4 gigabytes and reducing memory occupancy on OPT-30B by 41.57% compared to existing binary baseline methods. Fourth, the compression process is highly time-efficient, completing the quantization of a 7-billion-parameter model in under 30 minutes on a single graphics processing unit without requiring model retraining.
These results demonstrate that ultra-low bit post-training quantization is viable for production-scale models, substantially lowering the hardware barriers, energy costs, and infrastructure requirements for deploying capable language models. Decision-makers evaluating large-scale deployments on edge devices or resource-constrained local infrastructure can leverage this approach to bypass high-cost retraining pipelines. Based on the findings, implementing the method with a block size of 128 provides the best operational balance between compression density and linguistic accuracy.
In terms of limitations, the method introduces minor storage overheads for grouping identifiers, and executing accelerated binary matrix operations in hardware remains challenging due to fine-grained parameter groupings. Nonetheless, the consistent performance across multiple model families and zero-shot benchmarks provides high confidence in the framework's effectiveness for weight compression.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). It establishes foundational second-order post-training quantization methods for large language models that BiLLM directly builds upon and compares against when pushing quantization below two bits.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). It demonstrates that protecting a small fraction of salient weights prevents accuracy collapse, motivating BiLLM's structural identification and approximation of sensitive parameters.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). It provides crucial groundwork on analyzing activation and weight distributions across large language models to facilitate aggressive integer quantization without severe degradation.
- Paper: BiT: Robustly Binarized Multi-distilled Transformer, Zechun Liu et al. (2022). It explores the extreme limits and optimization challenges of 1-bit binarization in transformer architectures, directly contextualizing BiLLM's post-training binarization scheme.
- Paper: Binarized Neural Networks, Matthieu Courbariaux et al. (2016). It provides foundational principles and arithmetic formulations for binarizing neural network weights and activations.
- Paper: XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks, Mohammad Rastegari et al. (2016). It introduces seminal binary weight and input approximations using scaling factors, which underpin modern residual binarization strategies.
- Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). It extends ultra-low-bit post-training quantization to sub-2-bit regimes by incorporating Kronecker-factored gradient covariance statistics to stabilize severe compression.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). It advances sub-4-bit and ultra-low precision LLM compression by combining quantization-aware training with self-distillation to recover lost representational capacity.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). It explores adapting and fine-tuning models in extreme sub-3-bit and sub-4-bit regimes using low-rank plus quantized matrix decomposition.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). It provides a rigorous downstream safety and trustworthiness evaluation across diverse dimensions for models compressed with aggressive quantization and pruning techniques.
