BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Dayou DuYijia ZhangShijie CaoJiaqi GuoTing CaoXiaowen ChuNingyi Xu

article2024ACL81 citations

Proposes a self-distillation framework combining asymmetric clipping with a confidence-aware loss to train accurate 2-bit and 3-bit large language models using minimal compute and data.

Listen

Deploying modern large language models on resource-constrained hardware is heavily bottlenecked by severe memory and compute requirements. While compressing models to 4-bit precision is standard practice, pushing compression further into ultra-low precision—specifically 3-bit and 2-bit formats—severely degrades model output and reasoning capabilities. Existing post-training compression techniques suffer large accuracy drops, while quantization-aware training methods remain hindered by prohibitive training costs and poor representation learning.

The article introduces and evaluates BitDistiller, a new framework designed to preserve high performance in large language models compressed below 4 bits. It demonstrates how pairing tailored weight-clipping initialization with self-distillation allows models to retain high fidelity and task accuracy while dramatically lowering compression compute time and data needs.

To achieve this, the article uses an asymmetric quantization and clipping strategy at initialization to handle numerical outliers without recurring optimization overhead. It then applies quantization-aware training guided by self-distillation, where the original full-precision model acts as a teacher to guide the low-precision student. A novel Confidence-Aware Kullback-Leibler Divergence objective dynamically balances learning modes based on teacher confidence across general language benchmarks and complex reasoning datasets using models ranging from 3 billion to 70 billion parameters.

The findings show that BitDistiller consistently outperforms existing post-training and training-aware quantization baselines across 3-bit and 2-bit settings. In extreme 2-bit configurations on 7-billion parameter models, BitDistiller improves average general language accuracy by 3.54 percentage points over the strongest prior training baseline and outperforms post-training methods by over 12 percentage points. On complex mathematical reasoning tasks, it achieves 61.33% accuracy in 2-bit precision, outperforming the leading baseline by 24.69 percentage points. Furthermore, BitDistiller reduces training time to roughly 3 hours on a single graphical processing unit, compared to over 280 GPU hours required by previous methods.

These results demonstrate that ultra-low-bit model deployment is commercially viable without incurring catastrophic accuracy losses or unsustainable training costs. Organizations can deploy high-performing models on significantly cheaper hardware and edge devices. Notably, experiments revealed that using a same-sized model as the distillation teacher yielded better student performance than using a larger teacher model, suggesting that matching internal architectures is more important for low-bit distillation than sheer teacher size.

Organizations planning low-precision deployments should adopt asymmetric clipping and confidence-aware distillation pipelines rather than relying solely on post-training quantization. Teams should use matching teacher-student model architectures during distillation to maximize transfer efficiency. Before rolling out binary or vector-quantized systems, teams should conduct further testing, as BitDistiller is currently restricted to scalar quantization and sub-4-bit formats.

The article’s findings carry high confidence for the evaluated open-source model families and benchmarks, supported by consistent gains across scales. However, limitations remain: the underlying mechanisms explaining why same-sized teachers outperform larger teachers require further theoretical study, and performance guarantees are not yet established for 1-bit binary formats or alternative quantization structures.

arXiv: 2402.10631DD-DuDa/BitDistiller
Cover for BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Abstract

The upscaling of Large Language Models (LLMs) has yielded impressive advances in natural language processing, yet it also poses significant deployment challenges. Weight quantization has emerged as a widely embraced solution to reduce memory and computational demands. This paper introduces BitDistiller, a framework that synergizes Quantization-Aware Training (QAT) with Knowledge Distillation (KD) to boost the performance of LLMs at ultra-low precisions (sub-4-bit). Specifically, BitDistiller first incorporates a tailored asymmetric quantization and clipping technique to maximally preserve the fidelity of quantized weights, and then proposes a novel Confidence-Aware Kullback-Leibler Divergence (CAKLD) objective, which is employed in a self-distillation manner to enable faster convergence and superior model performance. Empirical evaluations demonstrate that BitDistiller significantly surpasses existing methods in both 3-bit and 2-bit configurations on general language understanding and complex reasoning benchmarks. Notably, BitDistiller is shown to be more cost-effective, demanding fewer data and training resources. The code is available at https://github.com/DD-DuDa/BitDistiller.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 Weight Quantization for LLMs
  • 2.2 Knowledge Distillation for LLMs
  • 3 Methodology
  • 3.1 Asymmetric Quantization and Clipping
  • 3.2 Self Distillation with CAKLD
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Evaluation on Language Modeling Tasks
  • 4.3 Evaluation on Reasoning Tasks
  • 4.4 Ablation Studies
  • 4.5 Analysis and Discussion
  • 5 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Details of PTQ and QAT Configuration
  • A.2 Implementation Details and Analysis of Confidence-Aware KLD
  • A.3 Training Datasets Examples
  • A.4 Evaluation of General Language Tasks on LLaMA-2-13B and LLaMA-2-70B
  • A.5 Integration with AWQ For Quantization Strategies

Knowls

  1. Knowl 1 — BitDistiller Quantization-Aware Self-Distillation Framework

    algorithm

    BitDistiller trains sub-4-bit (e.g., 3-bit and 2-bit) quantized Large Language Models (LLMs) by integrating asymmetric weight quantization and initialization-stage asymmetric clipping with a self-distillation Quantization-Aware Training (QAT) loop. The pre-trained full-precision model serves as its own teacher, transferring token-level distributions to the quantized student model using the Confidence-Aware Kullback-Leibler Divergence (CAKLD) loss.

    Input: Full-precision weights ww, Dataset D={(x,y)}\mathcal{D} = \{(x, y)\}, Learning rate η\eta, Total training steps TT
    Require: Clipping function Clip\text{Clip}, Quantization function QQ, Loss function DCAKLD\mathcal{D}_{\text{CAKLD}}
    Output: Low-precision quantized weights wQTw_Q^T
    w1=Clip(w)w^1 = \text{Clip}(w)
    for t=1t = 1 to TT do
        Sample a batch B\mathcal{B} from D\mathcal{D}
        Compute quantized weights for forward pass: wQt=Q(wt)w_Q^t = Q(w^t)
        Compute loss DCAKLD(PTw∥PSwQt)\mathcal{D}_{\text{CAKLD}}(P_T^w \parallel P_S^{w_Q^t}) on batch B\mathcal{B}
        Compute gradients with respect to full-precision weights: ∂DCAKLD∂wt\frac{\partial \mathcal{D}_{\text{CAKLD}}}{\partial w^t}
        wt+1=Update(wt,∂DCAKLD∂wt,η)w^{t+1} = \text{Update}\left(w^t, \frac{\partial \mathcal{D}_{\text{CAKLD}}}{\partial w^t}, \eta\right)
    end for
    wQT=Q(wT)w_Q^T = Q(w^T)
    return wQTw_Q^T

    Prior to gradient updates, weight tensors undergo asymmetric clipping optimization once per layer on a small calibration cache. During each training step tt, forward activations are calculated using the quantized weights wQtw_Q^t, while backward gradients are propagated to update the master full-precision weights wtw^t using a Straight-Through Estimator. Optimization uses the AdamW optimizer with zero weight decay, a constant learning rate of 8×10−68 \times 10^{-6}, and sequence lengths of 1024 for code tasks and 512 for other natural language tasks.

  2. Knowl 2 — Confidence-Aware Kullback-Leibler Divergence Objective

    equation

    The Confidence-Aware Kullback-Leibler Divergence (CAKLD) objective automatically balances mode-seeking and mode-covering distillation behaviors by interpolating between Reverse KL divergence and Forward KL divergence based on the empirical confidence of the teacher model:

    DCAKLD(PT∥PS)=γDKL(PS∥PT)+(1−γ)DKL(PT∥PS)\mathcal{D}_{\text{CAKLD}}(P_T \parallel P_S) = \gamma D_{\text{KL}}(P_S \parallel P_T) + (1 - \gamma) D_{\text{KL}}(P_T \parallel P_S)

    where the forward divergence is defined as:

    DKL(PT∥PS)=E(x,y)∼D[1∣y∣∑i=1∣y∣Ec∼PT(⋅∣x,y<i)[log⁡PT(c∣x,y<i)PS(c∣x,y<i)]]D_{\text{KL}}(P_T \parallel P_S) = \mathbb{E}_{(x,y)\sim \mathcal{D}}\left[ \frac{1}{|y|} \sum_{i=1}^{|y|} \mathbb{E}_{c \sim P_T(\cdot \mid x, y_{<i})}\left[ \log \frac{P_T(c \mid x, y_{<i})}{P_S(c \mid x, y_{<i})} \right] \right]

    and the blending coefficient γ∈[0,1]\gamma \in [0, 1] represents the mean confidence of the teacher model on next-token prediction across the training distribution:

    γ=E(x,y)∼D[1∣y∣∑i=1∣y∣PT(yi∣x,y<i)]\gamma = \mathbb{E}_{(x,y)\sim \mathcal{D}}\left[ \frac{1}{|y|} \sum_{i=1}^{|y|} P_T(y_i \mid x, y_{<i}) \right]

    Here, PTP_T denotes the full-precision teacher distribution, PSP_S denotes the low-precision student distribution, xx is the prompt sequence, y=(y1,…,y∣y∣)y = (y_1, \dots, y_{|y|}) is the target response sequence, y<iy_{<i} denotes the prefix sequence up to step i−1i-1, and cc denotes candidate vocabulary tokens. The scalar γ\gamma is pre-calculated over a subset of training data prior to QAT. When teacher confidence is high (as in formal reasoning), γ→1\gamma \to 1, prioritizing mode-seeking behavior (DKL(PS∥PT)D_{\text{KL}}(P_S \parallel P_T)); when teacher confidence is lower (as in open-ended text generation), γ→0\gamma \to 0, emphasizing mode-covering behavior (DKL(PT∥PS)D_{\text{KL}}(P_T \parallel P_S)).

  3. Knowl 3 — Precision-Dependent Asymmetric Quantization and Offline Clipping

    model/method

    To preserve weight tensor fidelity across different bit widths, BitDistiller adopts precision-specific quantization schemes and a single-pass pre-training asymmetric clipping search:

    1. 3-Bit Asymmetric NormalFloat Quantization (NF-Asym): For precisions above 2-bit (such as NF3), asymmetric floating-point quantization computes separate positive (sposs_{\text{pos}}) and negative (snegs_{\text{neg}}) scale factors per weight group:

    QNF-Asym(w)={⌊wposspos⌉,if w>0⌊wnegsneg⌉,if w≤0Q_{\text{NF-Asym}}(w) = \begin{cases} \left\lfloor \frac{w_{\text{pos}}}{s_{\text{pos}}} \right\rceil, & \text{if } w > 0 \\ \left\lfloor \frac{w_{\text{neg}}}{s_{\text{neg}}} \right\rceil, & \text{if } w \le 0 \end{cases}

    1. 2-Bit Asymmetric Integer Quantization (INT-Asym): For 2-bit quantization, non-uniform distributions lose effectiveness due to having only four discrete representation states. BitDistiller uses uniform integer quantization with a single scale ss and zero-point offset zz:

    QINT-Asym(w)=⌊w−zs⌉Q_{\text{INT-Asym}}(w) = \left\lfloor \frac{w - z}{s} \right\rceil

    1. Pre-Training Asymmetric Clipping Optimization: Prior to QAT, lower (α\alpha) and upper (β\beta) clipping boundaries are searched per layer to minimize layer-level output error on calibration activations XX:

    α∗,β∗=arg⁡min⁡α∈[min⁡(w),0),β∈(0,max⁡(w)]∥Q(Clip(w,α,β))X−wX∥22\alpha^*, \beta^* = \arg\min_{\alpha \in [\min(w), 0), \beta \in (0, \max(w)]} \| Q(\text{Clip}(w, \alpha, \beta)) X - w X \|_2^2

    where Clip(w,α,β)=min⁡(max⁡(w,α),β)\text{Clip}(w, \alpha, \beta) = \min(\max(w, \alpha), \beta). Performing this search exclusively at initialization eliminates the computational overhead of iterative clipping search during QAT while providing a well-conditioned weight initialization.

  4. Knowl 4 — General Language Understanding Benchmark Results across Sub-4-Bit Precisions

    data/table

    BitDistiller was evaluated across general language understanding benchmarks against Post-Training Quantization (RTN, GPTQ, AWQ) and Quantization-Aware Training (OmniQuant, LLM-QAT) baselines using LLaMA-2 models with group size 128 (g128). Evaluated metrics include WikiText-2 perplexity (PPL, lower is better), 5-shot MMLU accuracy, zero-shot QA benchmarks (PIQA, HellaSwag, WinoGrande, ARC-Challenge), and overall QA Average.

    Model Method PPL ↓\downarrow MMLU PIQA Hella. Wino. ARC-c Avg
    LLaMA-2-7B BF16 5.47 46.45 77.86 57.14 68.35 43.34 58.63
    3 Bits (g128) RTN 6.65 38.65 75.24 53.70 67.32 38.56 54.69
    GPTQ 6.38 39.57 75.46 51.68 67.16 38.39 54.45
    AWQ 6.71 39.68 76.27 55.14 67.56 40.61 55.85
    OmniQuant 6.10 41.22 77.47 54.41 67.09 39.08 55.85
    LLM-QAT 6.02 41.32 77.26 54.74 68.35 40.61 56.46
    BitDistiller 5.97 43.65 76.99 55.38 68.35 41.21 57.12
    2 Bits (g128) RTN 3453 24.12 53.43 26.33 49.96 21.58 35.08
    GPTQ NaN 23.12 49.51 25.04 49.57 22.69 33.99
    AWQ 2.2e5 25.38 52.39 25.70 50.12 21.33 34.98
    OmniQuant 12.84 25.42 58.92 29.20 50.83 19.45 36.76
    LLM-QAT 9.30 23.62 70.08 43.79 61.64 29.09 45.64
    BitDistiller 8.08 29.25 73.61 48.70 61.09 33.27 49.18
    LLaMA-2-13B BF16 4.88 55.54 79.16 60.13 72.14 48.12 63.02
    3 Bits (g128) RTN 5.52 50.74 78.35 57.75 71.11 43.86 60.36
    GPTQ 5.41 50.63 77.26 56.84 70.72 42.83 59.66
    AWQ 5.47 49.64 77.09 57.52 70.32 43.86 59.69
    OmniQuant 5.48 48.97 77.64 57.08 70.88 44.28 59.77
    LLM-QAT 5.32 51.60 78.29 58.45 70.56 44.62 60.70
    BitDistiller 5.20 53.21 78.67 58.66 71.59 46.67 61.76
    2 Bits (g128) RTN 109.21 24.74 57.56 32.56 50.75 21.84 37.49
    GPTQ 15.08 23.70 56.04 30.99 51.22 19.28 36.25
    AWQ 1.2e5 27.04 53.16 25.82 51.70 23.04 36.15
    OmniQuant 25.69 26.09 61.81 31.92 51.38 22.27 38.69
    LLM-QAT 7.80 29.37 74.10 49.49 63.14 33.87 49.99
    BitDistiller 6.78 37.50 75.84 51.30 65.90 37.46 53.60

    In the 2-bit regime, PTQ methods experience severe degradation or divergence (e.g., AWQ reaching 2.2×1052.2 \times 10^5 PPL and GPTQ producing NaN). BitDistiller maintains stability and outperforms LLM-QAT on LLaMA-2-7B by +3.54%+3.54\% in average QA accuracy and −1.22-1.22 in perplexity, and on LLaMA-2-13B by +3.61%+3.61\% in average QA accuracy and +8.13%+8.13\% in MMLU accuracy.

  5. Knowl 5 — Sub-4-Bit Reasoning Benchmark Results on Code and Mathematics

    data/table

    BitDistiller was benchmarked on domain-specific LLMs across multiple parameter scales: WizardCoder (3B, 7B, 13B, 34B) evaluated on HumanEval (Pass@1 under greedy decoding) and MetaMath (3B, 7B, 13B) evaluated on GSM8K (accuracy). Group size is 128 (g128), except for 3B models which use group size 64.

    HumanEval @ WizardCoder GSM8K @ MetaMath
    Bits Method 3B 7B 13B 34B 3B 7B 13B
    BF16 Baseline 23.17 54.88 62.80 71.95 36.40 66.41 72.30
    3 Bits RTN 4.27 34.15 50.00 33.54 17.50 59.30 68.51
    GPTQ 4.30 46.34 55.48 63.41 6.72 62.11 68.75
    AWQ 16.46 45.73 53.04 67.07 21.87 62.34 68.67
    OmniQuant 10.36 44.51 54.88 68.90 23.67 61.70 68.28
    LLM-QAT 18.29 48.78 57.92 66.46 26.25 60.78 66.62
    BitDistiller 20.73 53.66 63.41 69.51 32.50 64.38 69.69
    2 Bits RTN 0.00 0.00 0.00 0.61 0.00 0.00 7.89
    GPTQ 0.00 0.00 1.83 3.65 0.00 0.00 11.43
    AWQ 0.00 0.00 0.00 0.00 0.00 0.00 7.89
    OmniQuant 0.00 0.00 20.12 26.83 0.00 0.00 9.45
    LLM-QAT 0.00 14.63 15.21 29.27 6.56 23.13 36.64
    BitDistiller 7.31 36.59 42.07 46.34 16.09 51.02 61.33

    While existing PTQ methods collapse to near-zero accuracy in 2-bit reasoning tasks, BitDistiller retains substantial reasoning capabilities across all model scales. On 2-bit MetaMath-7B, BitDistiller achieves 51.02%51.02\% accuracy on GSM8K (outperforming LLM-QAT's 23.13%23.13\% by +27.89%+27.89\%) and on 2-bit WizardCoder-7B achieves 36.59%36.59\% Pass@1 on HumanEval (outperforming LLM-QAT's 14.63%14.63\% by +21.96%+21.96\%).

  6. Knowl 6 — Ablation of Quantization Symmetry and Pre-Training Clipping

    empirical result

    An ablation study on LLaMA-2-7B (group size 128) isolating the contributions of asymmetric quantization and initialization-stage asymmetric clipping shows progressive improvements in WikiText-2 perplexity (PPL) and 5-shot MMLU accuracy before training (extstart ext{start}) and after training (extend ext{end}):

    Bit-width Configuration PPL ↓\downarrow (start →\to end) MMLU (5s) ↑\uparrow (start →\to end)
    3 Bits (g128) NF-Sym 6.45 →\to 6.10 38.28 →\to 39.27
    →\to NF-Asym 6.30 →\to 6.01 41.53 →\to 42.61
    + Clip-Asym 6.08 →\to 5.97 42.90 →\to 43.65
    2 Bits (g128) INT-Sym 2.4×105→2.5×1052.4\times 10^5 \to 2.5\times 10^5 24.95 →\to 26.03
    →\to INT-Asym 3.4×102→16.943.4\times 10^2 \to 16.94 24.12 →\to 24.82
    + Clip-Asym 17.98 →\to 8.08 26.75 →\to 29.25

    In the 2-bit regime, symmetric integer quantization (INT-Sym) fails to converge (PPL remains ∼2.5×105\sim 2.5 \times 10^5). Switching to asymmetric integer quantization (INT-Asym) reduces post-training PPL to 16.94. Adding asymmetric clipping at initialization (+ Clip-Asym) lowers the post-training PPL further to 8.08 and raises MMLU accuracy from 24.82%24.82\% to 29.25%29.25\%.

    When combining initialization methods on 2-bit LLaMA-2-7B with AWQ, initializing with AWQ alone results in divergence (extPPL=∞ ext{PPL} = \infty), whereas initialization with Clip-Asym alone yields post-training PPL of 8.08, and combining AWQ + Clip-Asym yields 8.13, demonstrating that standalone asymmetric clipping is sufficient for QAT initialization.

  7. Knowl 7 — Ablation of Distillation Objectives and Data Generation Sources

    empirical result

    A comparative evaluation of distillation loss formulations and training response sources was conducted on 2-bit quantized 7B reasoning models:

    1. Distillation Objective Functions: On WizardCoder-7B (HumanEval Pass@1) and MetaMath-7B (GSM8K accuracy):

      • CAKLD: 36.59%36.59\% on HumanEval; 51.02%51.02\% on GSM8K.
      • Reverse KL (RKLD): 34.76%34.76\% on HumanEval; 48.67%48.67\% on GSM8K.
      • Forward KL (FKLD): 33.54%33.54\% on HumanEval; 49.23%49.23\% on GSM8K.
      • Jensen-Shannon Divergence (JSD): 28.65%28.65\% on HumanEval; 34.14%34.14\% on GSM8K. CAKLD consistently outperforms fixed single-divergence objectives, while JSD exhibits weak convergence during low-bit QAT.
    2. Training Data Generation Source: Comparing three sequence target types yy on WizardCoder-7B given instruction prompt xx:

      • Teacher-generated Output (ypy_p) (sampled at temperature 0.7): Achieves the lowest per-token cross-entropy loss and highest accuracy (36.59%36.59\% HumanEval Pass@1).
      • Student-generated Output (yqy_q): Achieves 32.31%32.31\% Pass@1.
      • Ground-Truth Data (ygy_g): Achieves 26.83%26.83\% Pass@1 with KD (and supervised fine-tuning without KD performs lowest). Teacher generations provide a higher-confidence logit distribution that stabilizes training under low precision.
  8. Knowl 8 — Training Efficiency and Resource Overhead Comparison with LLM-QAT

    empirical result

    BitDistiller reduces the data volume and compute requirements needed for sub-4-bit quantization-aware training compared to LLM-QAT. The wall-clock execution time and sample sizes for quantizing WizardCoder-7B on a single NVIDIA A100-80GB GPU are as follows:

    Method Hardware # Samples Time (Hours)
    Data Gen Quant Init QAT Total
    LLM-QAT 1×1\times A100-80G 100K 270.00 0.00 10.64 280.64
    BitDistiller 1×1\times A100-80G 2K 1.47 0.63 0.92 3.02

    BitDistiller requires only 2,000 instruction-following samples (compared to 100,000 for LLM-QAT) and completes the full quantization workflow in 3.02 hours (1.47h data generation, 0.63h initialization clipping, 0.92h QAT), achieving a ∼93×\sim 93\times reduction in total GPU compute time relative to LLM-QAT's 280.64 hours.

  9. Knowl 9 — Impact of Teacher Architecture Alignment in Sub-4-Bit Self-Distillation

    empirical result

    In 2-bit quantization-aware training, using a teacher model with the identical architecture and size as the student (self-distillation) achieves better accuracy and lower perplexity than distilling from a larger teacher model:

    Quantized Student Student Precision Teacher Model WikiText-2 PPL ↓\downarrow MMLU (5s) ↑\uparrow
    LLaMA-2-7B 2 Bits (g128) LLaMA-2-13B 8.12 28.27
    LLaMA-2-7B 2 Bits (g128) LLaMA-2-7B 8.08 29.25

    Distilling a 2-bit LLaMA-2-7B student using the full-precision LLaMA-2-7B as teacher yields 8.08 PPL and 29.25%29.25\% MMLU accuracy, compared to 8.12 PPL and 28.27%28.27\% MMLU accuracy when using the larger LLaMA-2-13B as teacher. Matching the model architecture between student and teacher facilitates direct probability distribution matching and token-level alignment in ultra-low-bit regimes.

  10. Knowl 10 — Limitations of BitDistiller

    limitation

    The BitDistiller framework has three documented limitations:

    1. Scalar Quantization Restriction: BitDistiller operates exclusively on scalar weight quantization and does not incorporate vector quantization or lattice codebooks (such as QuIP#), which offer complementary representational advantages at 2 bits.
    2. Bit-Width Boundary: The framework is formulated and evaluated on 2-bit and 3-bit precisions and has not been extended to 1-bit (binary) quantization, where multiplication operations can be replaced entirely with additions.
    3. Empirical Nature of Architecture Alignment: The phenomenon where a same-size teacher model (7B) outperforms a larger teacher model (13B) when distilling a 2-bit 7B student is empirically observed but lacks a complete formal theoretical explanation.

Coverage note — None. All main contributions, methods, equations, experimental comparisons, ablation studies, and stated limitations are covered.

References

  1. 1.Rishabh Agarwal, Nino Vieillard, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem. 2023. Gkd: Generalized knowledge distillation for auto-regressive sequence models. arXiv preprint arXiv:2306.13649.
  2. 2.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432.
  3. 3.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa. 2023. Quip: 2-bit quantization of large language models with guarantees. arXiv preprint arXiv:2307.13304.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  9. 9.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  10. 10.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023a. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  11. 11.Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2023b. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078.
  12. 12.Tim Dettmers and Luke Zettlemoyer. 2023. The case for 4-bit precision: k-bit inference scaling laws. In International Conference on Machine Learning, pages 7750–7774. PMLR.
  13. 13.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.
  14. 14.Xinyang Geng and Hao Liu. 2023. Openllama: An open reproduction of llama.
  15. 15.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2022. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision, pages 291–326. Chapman and Hall/CRC.
  16. 16.Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2023. Knowledge distillation of large language models. arXiv preprint arXiv:2306.08543.
  17. 17.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  18. 18.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  19. 19.Sangil Jung, Changyong Son, Seohyung Lee, Jinwoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, and Changkyu Choi. 2019. Learning to quantize deep networks by optimizing quantization intervals with task loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4350–4359.
  20. 20.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  21. 21.Minsoo Kim, Sihwa Lee, Sukjin Hong, Du-Seong Chang, and Jungwook Choi. 2022. Understanding and improving knowledge distillation for quantization-aware training of large transformer encoders. arXiv preprint arXiv:2211.11014.
  22. 22.Minsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong, Du-Seong Chang, Wonyong Sung, and Jungwook Choi. 2023a. Token-scaled logit distillation for ternary weight generative language models. arXiv preprint arXiv:2308.06744.
  23. 23.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2023b. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629.
  24. 24.Andrey Kuzmin, Mart Van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, and Tijmen Blankevoort. 2022. Fp8 quantization: The power of the exponent. Advances in Neural Information Processing Systems, 35:14651–14662.
  25. 25.Rundong Li, Yan Wang, Feng Liang, Hongwei Qin, Junjie Yan, and Rui Fan. 2019. Fully quantized network for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2810–2819.
  26. 26.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978.
  27. 27.Shih-yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, and Kwang-Ting Cheng. 2023a. Llm-fp4: 4-bit floating-point quantized transformers. arXiv preprint arXiv:2310.16836.
  28. 28.Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2023b. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888.
  29. 29.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  30. 30.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct.
  31. 31.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843.
  32. 32.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  33. 33.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506.
  34. 34.Nick Rosh. 2023. Evol-teacher: Recreating wizardcoder. https://github.com/nickrosh/evol-teacher.
  35. 35.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.
  36. 36.Charbel Sakr, Steve Dai, Rangha Venkatesan, Brian Zimmer, William Dally, and Brucek Khailany. 2022. Optimal clipping and magnitude-aware differentiation for improved quantization-aware training. In International Conference on Machine Learning, pages 19123–19138. PMLR.
  37. 37.Yuzhang Shang, Zhihang Yuan, Qiang Wu, and Zhen Dong. 2023. Pb-llm: Partially binarized large language models. arXiv preprint arXiv:2310.00034.
  38. 38.Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2023. Omniquant: Omnidirectionally calibrated quantization for large language models. arXiv preprint arXiv:2308.13137.
  39. 39.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8815–8821.
  40. 40.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  41. 41.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  42. 42.Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei. 2023. Bitnet: Scaling 1-bit transformers for large language models. arXiv preprint arXiv:2310.11453.
  43. 43.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  44. 44.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, JamesT. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Meta-math: Bootstrap your own mathematical questions for large language models.
  45. 45.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. 2023. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825.
  46. 46.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  47. 47.Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020. Ternarybert: Distillation-aware ultra-low bit bert. arXiv preprint arXiv:2009.12812.
  48. 48.Yijia Zhang, Sicheng Zhang, Shijie Cao, Dayou Du, Jianyu Wei, Ting Cao, and Ningyi Xu. 2023a. Afpq: Asymmetric floating point quantization for llms.
  49. 49.Yijia Zhang, Lingran Zhao, Shijie Cao, Wenqiang Wang, Ting Cao, Fan Yang, Mao Yang, Shanghang Zhang, and Ningyi Xu. 2023b. Integer or floating point? new outlooks for low-bit quantization on large language models. arXiv preprint arXiv:2305.12356.
  50. 50.Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. 2023. Distillspec: Improving speculative decoding via knowledge distillation.
  51. 51.Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023. A survey on model compression for large language models. arXiv preprint arXiv:2308.07633.

Citation

MLA
Du, D., et al. “BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 102–16, https://doi.org/10.18653/v1/2024.acl-long.7.
APA
Du, D., (张益嘉), Y. Z., Cao, S., Guo, J., Cao, T., Chu, X., & Xu, N. (2024). BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 102–116. https://doi.org/10.18653/v1/2024.acl-long.7
Chicago
Du, D., Y. Z. (张益嘉), S. Cao, et al. 2024. “BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 102–16. https://doi.org/10.18653/v1/2024.acl-long.7.
Harvard
Du, D. et al. (2024) “BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 102–116. Available at: https://doi.org/10.18653/v1/2024.acl-long.7.
Vancouver
1. Du D, (张益嘉) YZ, Cao S, Guo J, Cao T, Chu X, Xu N (2024) BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 102–116

BibTeX

@inproceedings{du-etal-2024-bitdistiller,
    title = "{B}it{D}istiller: Unleashing the Potential of Sub-4-Bit {LLM}s via Self-Distillation",
    author = "Du, DaYou  and
      Zhang, Yijia  and
      Cao, Shijie  and
      Guo, Jiaqi  and
      Cao, Ting  and
      Chu, Xiaowen  and
      Xu, Ningyi",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.7/",
    doi = "10.18653/v1/2024.acl-long.7",
    pages = "102--116"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/