BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
Dayou DuYijia ZhangShijie CaoJiaqi GuoTing CaoXiaowen ChuNingyi Xu
Proposes a self-distillation framework combining asymmetric clipping with a confidence-aware loss to train accurate 2-bit and 3-bit large language models using minimal compute and data.
Deploying modern large language models on resource-constrained hardware is heavily bottlenecked by severe memory and compute requirements. While compressing models to 4-bit precision is standard practice, pushing compression further into ultra-low precision—specifically 3-bit and 2-bit formats—severely degrades model output and reasoning capabilities. Existing post-training compression techniques suffer large accuracy drops, while quantization-aware training methods remain hindered by prohibitive training costs and poor representation learning.
The article introduces and evaluates BitDistiller, a new framework designed to preserve high performance in large language models compressed below 4 bits. It demonstrates how pairing tailored weight-clipping initialization with self-distillation allows models to retain high fidelity and task accuracy while dramatically lowering compression compute time and data needs.
To achieve this, the article uses an asymmetric quantization and clipping strategy at initialization to handle numerical outliers without recurring optimization overhead. It then applies quantization-aware training guided by self-distillation, where the original full-precision model acts as a teacher to guide the low-precision student. A novel Confidence-Aware Kullback-Leibler Divergence objective dynamically balances learning modes based on teacher confidence across general language benchmarks and complex reasoning datasets using models ranging from 3 billion to 70 billion parameters.
The findings show that BitDistiller consistently outperforms existing post-training and training-aware quantization baselines across 3-bit and 2-bit settings. In extreme 2-bit configurations on 7-billion parameter models, BitDistiller improves average general language accuracy by 3.54 percentage points over the strongest prior training baseline and outperforms post-training methods by over 12 percentage points. On complex mathematical reasoning tasks, it achieves 61.33% accuracy in 2-bit precision, outperforming the leading baseline by 24.69 percentage points. Furthermore, BitDistiller reduces training time to roughly 3 hours on a single graphical processing unit, compared to over 280 GPU hours required by previous methods.
These results demonstrate that ultra-low-bit model deployment is commercially viable without incurring catastrophic accuracy losses or unsustainable training costs. Organizations can deploy high-performing models on significantly cheaper hardware and edge devices. Notably, experiments revealed that using a same-sized model as the distillation teacher yielded better student performance than using a larger teacher model, suggesting that matching internal architectures is more important for low-bit distillation than sheer teacher size.
Organizations planning low-precision deployments should adopt asymmetric clipping and confidence-aware distillation pipelines rather than relying solely on post-training quantization. Teams should use matching teacher-student model architectures during distillation to maximize transfer efficiency. Before rolling out binary or vector-quantized systems, teams should conduct further testing, as BitDistiller is currently restricted to scalar quantization and sub-4-bit formats.
The article’s findings carry high confidence for the evaluated open-source model families and benchmarks, supported by consistent gains across scales. However, limitations remain: the underlying mechanisms explaining why same-sized teachers outperform larger teachers require further theoretical study, and performance guarantees are not yet established for 1-bit binary formats or alternative quantization structures.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). This paper establishes foundational post-training quantization for large language models down to 3–4 bits, providing the baseline and performance limits that BitDistiller aims to surpass in sub-4-bit regimes.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). It introduces activation-aware weight clipping and protection for salient weights in low-bit LLMs, motivating the initialization and outlier mitigation strategies adopted in BitDistiller.
- Paper: QLoRA: Efficient Finetuning of Quantized LLMs, Tim Dettmers et al. (2023). It demonstrates how quantization errors compound below 4 bits during model adaptation, establishing the practical necessity of quantization-aware optimization schemes.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal text introduces knowledge distillation and soft probability matching, which form the conceptual bedrock of BitDistiller's teacher-student objective.
- Paper: f-Divergence Minimization for Sequence-Level Knowledge Distillation, Yuqiao Wen et al. (2023). It analyzes the fundamental mode-averaging and mode-collapsing failure modes of KL divergence in language model distillation, directly motivating BitDistiller's Confidence-Aware KL objective.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). It provides crucial background on how activation outliers degrade low-bit LLM compression, informing how initialization strategies must manage numerical anomalies.
- Paper: BiT: Robustly Binarized Multi-distilled Transformer, Zechun Liu et al. (2022). It explores progressive knowledge distillation from full-precision teachers to severely quantized students, preceding BitDistiller's ultra-low-bit distillation framework.
- Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). It analyzes teacher-student capacity mismatch during distillation, providing key context for why BitDistiller finds same-sized teachers optimal for low-bit students.
- Paper: BiLLM: Pushing the Limit of Post-Training Quantization for LLMs, Wei Huang et al. (2024). BiLLM pushes post-training weight compression even further to extreme 1-bit representations using structural residual binarization, building on the ultra-low-bit boundaries studied in BitDistiller.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). QuIP# advances sub-4-bit compression through randomized Hadamard transforms and lattice vector quantization, offering a complementary alternative to scalar quantization-aware distillation.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). LQ-LoRA explores sub-3-bit LLM adaptation via low-rank plus quantized matrix decomposition, extending low-bit parameter efficiency into fine-tuning settings.
- Paper: DistiLLM: Towards Streamlined Distillation for Large Language Models, Jongwoo Ko et al. (2024). DistiLLM develops specialized divergence objectives and student-generated rollout distillation for LLMs, extending the divergence optimization principles explored in BitDistiller.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). This work evaluates the downstream safety, fairness, and trustworthiness consequences of aggressively quantizing LLMs down to low-bit regimes.
- Paper: Scaling FP8 training to trillion-token LLMs, Maxim Fishman et al. (2025). This work explores stable low-precision training dynamics at the trillion-token scale, expanding the operational understanding of numerical stability in compressed LLMs.
