BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Wei HuangYangdong LiuHaotong QinYing LiShiming ZhangXianglong LiuMichele MagnoXiaojuan Qi

article2024ICML197 citations

Develops an ultra-low-bit post-training quantization method that compresses large language models down to roughly 1.08 bits per weight using binary residual approximation and optimal distribution splitting while maintaining high inference accuracy.

Listen

Large language models demonstrate powerful language processing capabilities but require massive memory and computational infrastructure. For example, running a 70-billion-parameter model in standard half-precision format requires approximately 150 gigabytes of memory, demanding multiple high-end enterprise graphics processors. While post-training quantization compresses model weights to lower memory footprints without expensive retraining, conventional techniques suffer severe performance degradation or total output collapse when pushed to ultra-low bit-widths of two bits or fewer.

The article introduces and evaluates BiLLM, a novel post-training binarization framework designed to compress pretrained language models down to approximately one bit per weight while preserving linguistic accuracy.

The authors conducted an empirical analysis of weight distributions across several model families, identifying that a small fraction of structured parameters carries outsized sensitivity while the remaining parameters follow a bell-shaped distribution. Using these insights, the method structurally selects critical weights to approximate them through a two-step binary residual technique and partitions the remaining weights into concentrated and sparse groups via an optimal break-point search. The framework was evaluated across the OPT, LLaMA, LLaMA 2, and Vicuna model families spanning sizes from 1.3 billion to 70 billion parameters across standard language benchmarks and zero-shot reasoning tasks on a single graphics processing unit.

The evaluation produced four major findings. First, the proposed method achieved average weight precisions between 1.07 and 1.11 bits across evaluated models without experiencing performance collapse. Second, on large models such as LLaMA 2-70B, the method achieved an inference perplexity of 8.41 at 1.08 bits, outperforming the full 16-bit precision version of OPT-66B. Third, the framework delivered nearly a tenfold reduction in model storage footprint, shrinking LLaMA 2-70B from 129.3 gigabytes to 15.4 gigabytes and reducing memory occupancy on OPT-30B by 41.57% compared to existing binary baseline methods. Fourth, the compression process is highly time-efficient, completing the quantization of a 7-billion-parameter model in under 30 minutes on a single graphics processing unit without requiring model retraining.

These results demonstrate that ultra-low bit post-training quantization is viable for production-scale models, substantially lowering the hardware barriers, energy costs, and infrastructure requirements for deploying capable language models. Decision-makers evaluating large-scale deployments on edge devices or resource-constrained local infrastructure can leverage this approach to bypass high-cost retraining pipelines. Based on the findings, implementing the method with a block size of 128 provides the best operational balance between compression density and linguistic accuracy.

In terms of limitations, the method introduces minor storage overheads for grouping identifiers, and executing accelerated binary matrix operations in hardware remains challenging due to fine-grained parameter groupings. Nonetheless, the consistent performance across multiple model families and zero-shot benchmarks provides high confidence in the framework's effectiveness for weight compression.

arXiv: 2402.04291
Cover for BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Abstract

Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive computation and memory requirements. However, existing quantization techniques fall short of maintaining LLM performance under ultra-low bit-widths. In response to this challenge, we present BiLLM, a groundbreaking 1-bit post-training quantization scheme tailored for pretrained LLMs. Based on the weight distribution of LLMs, BiLLM first identifies and structurally selects salient weights, and minimizes the compression loss through an effective binary residual approximation strategy. Moreover, considering the bell-shaped distribution of the non-salient weights, we propose an optimal splitting search to group and binarize them accurately. BiLLM, for the first time, achieves high-accuracy inference (e.g. 8.41 perplexity on LLaMA2-70B) with only 1.08-bit weights across various LLM families and evaluation metrics, outperforms SOTA quantization methods of LLM by significant margins. Moreover, BiLLM enables the binarization process of a 7-billion LLM within 0.5 hours on a single GPU, demonstrating satisfactory time efficiency. Our code is available at https://github.com/Aaronhuang-778/BiLLM.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Large Language Model Quantization
  • 2.2. Network Binarization
  • 3. Method
  • 3.1. Salient Weight Binarization for LLMs
  • 3.2. Bell-shaped Distribution Splitting for Binarization
  • 3.3. Pipeline of BiLLM
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Results
  • 5. Conclusions
  • Impact Statement
  • References
  • A. BiLLM Implementation
  • B. Quantization Error
  • C. Searching Curve of Salient Column and Non-salient Distribution
  • D. Multi-evaluation Comparisons
  • E. Ablation of BiLLM with different block size
  • F. Dialog Examples
  • G. Magnitude and Hessian Distribution of LLMs

Knowls

  1. Knowl 1 — BiLLM Framework for 1-Bit Post-Training Quantization of Large Language Models

    model/method

    BiLLM is a post-training binarization (1-bit quantization) framework tailored for pretrained large language models (LLMs). The method addresses the severe performance collapse of LLMs under ultra-low bit-widths based on two empirical observations regarding LLM weight distributions:

    1. The second-order Hessian matrix exhibits an extreme long-tail distribution where a small fraction of weight elements have significantly high Hessian sensitivity concentrated in specific columns or rows, while most values cluster near zero.
    2. The density distribution of non-salient weight magnitudes follows a bell-shaped, approximately symmetric profile centered at zero (resembling Gaussian or Laplacian distributions).

    To exploit these characteristics, BiLLM applies two complementary strategies:

    • Structured Salient Weight Binarization: Hessian-sensitive columns are identified structurally to avoid bitmap overhead, and their values are compressed via a second-order binary residual approximation.
    • Bell-Shaped Distribution Splitting: The remaining non-salient weights are segmented into concentrated and sparse areas using an optimal break-point search, and each region is binarized separately.

    BiLLM incorporates block-wise error compensation based on Optimal Brain Compression (OBC) while eliminating column-wise compensation to maintain distribution search stability and efficiency, completing the quantization of a 7B parameter LLM within 0.5 hours on a single GPU.

  2. Knowl 2 — Structured Salient Weight Selection and Binary Residual Approximation in BiLLM

    model/method

    To preserve high-impact parameters without the storage overhead of uncompressed INT8/FP16 formats or unstructured bitmap indices, BiLLM employs structured column selection guided by the Hessian matrix, followed by binary residual approximation.

    The sensitivity sis_i of each parameter wiw_i in a layer is computed via the Hessian matrix H\mathbf{H}: si=wi2[H−1]ii2s_i = \frac{w_i^2}{[\mathbf{H}^{-1}]_{ii}^2}

    Because sensitivity clusters heavily along specific columns due to multi-head self-attention mechanisms, columns are sorted by aggregate salience. The optimal number of salient columns is selected to minimize the reconstruction error: arg⁡min⁡Wuns∥W−(αsalsign⁡(Wsal)∪αunssign⁡(Wuns))∥22\arg\min_{\mathbf{W}_{uns}} \|\mathbf{W} - (\alpha_{sal} \operatorname{sign}(\mathbf{W}_{sal}) \cup \alpha_{uns} \operatorname{sign}(\mathbf{W}_{uns}))\|_2^2 where Wsal∈Rn×k\mathbf{W}_{sal} \in \mathbb{R}^{n \times k} is the submatrix of kk selected salient columns, Wuns\mathbf{W}_{uns} is the non-salient remainder, W=Wsal∪Wuns\mathbf{W} = \mathbf{W}_{sal} \cup \mathbf{W}_{uns}, and α=∥W∥ℓ1n×k\alpha = \frac{\|\mathbf{W}\|_{\ell_1}}{n \times k}.

    For the selected salient submatrix Wsal\mathbf{W}_{sal}, BiLLM applies a second-order binary residual approximation rather than full-precision retention: αo∗,Bo∗=arg⁡min⁡αo,Bo∥Wsal−αoBo∥22\alpha_o^*, \mathbf{B}_o^* = \arg\min_{\alpha_o, \mathbf{B}_o} \|\mathbf{W}_{sal} - \alpha_o \mathbf{B}_o\|_2^2 αr∗,Br∗=arg⁡min⁡αr,Br∥(Wsal−αo∗Bo∗)−αrBr∥22\alpha_r^*, \mathbf{B}_r^* = \arg\min_{\alpha_r, \mathbf{B}_r} \|(\mathbf{W}_{sal} - \alpha_o^* \mathbf{B}_o^*) - \alpha_r \mathbf{B}_r\|_2^2 yielding the approximation Wsal≈αo∗Bo∗+αr∗Br∗\mathbf{W}_{sal} \approx \alpha_o^* \mathbf{B}_o^* + \alpha_r^* \mathbf{B}_r^* with binary matrices Bo∗,Br∗∈{−1,+1}n×k\mathbf{B}_o^*, \mathbf{B}_r^* \in \{-1, +1\}^{n \times k} and channel-wise scalar factors αo∗,αr∗∈Rn\alpha_o^*, \alpha_r^* \in \mathbb{R}^n. The residual quantization error Erb=∥Wsal−αo∗Bo∗−αr∗Br∗∥22\mathcal{E}_{rb} = \|\mathbf{W}_{sal} - \alpha_o^* \mathbf{B}_o^* - \alpha_r^* \mathbf{B}_r^*\|_2^2 is bounded above by the single-stage binarization error ∥Wsal−αo∗Bo∗∥22\|\mathbf{W}_{sal} - \alpha_o^* \mathbf{B}_o^*\|_2^2.

  3. Knowl 3 — Bell-Shaped Distribution Splitting Binarization for Non-Salient Weights

    model/method

    Following the removal of salient columns, the remaining non-salient weights follow a bell-shaped distribution centered around zero. Applying standard binary quantization directly incurs substantial error due to the non-uniform density. BiLLM partitions the non-salient weight range [−m,m][-m, m] using a single break-point p∈(0,m)p \in (0, m) into two intervals:

    • Concentrated area: Ac=[−p,p]\mathcal{A}_c = [-p, p]
    • Sparse area: As=[−m,−p)∪(p,m]\mathcal{A}_s = [-m, -p) \cup (p, m]

    Let Wc\mathbf{W}_c and Ws\mathbf{W}_s denote the weight tensors restricted to the concentrated and sparse areas, respectively. Each region is binarized independently with its own scale factor: αc=1nc∥Wc∥ℓ1,αs=1ns∥Ws∥ℓ1\alpha_c = \frac{1}{n_c} \|\mathbf{W}_c\|_{\ell_1}, \quad \alpha_s = \frac{1}{n_s} \|\mathbf{W}_s\|_{\ell_1} where ncn_c and nsn_s denote the number of elements in each respective region.

    The optimal break-point p∗p^* is determined by minimizing the total binarization squared error: p∗=arg⁡min⁡p(∥Ws−αssign⁡(Ws)∥22+∥Wc−αcsign⁡(Wc)∥22)p^* = \arg\min_p \left( \|\mathbf{W}_s - \alpha_s \operatorname{sign}(\mathbf{W}_s)\|_2^2 + \|\mathbf{W}_c - \alpha_c \operatorname{sign}(\mathbf{W}_c)\|_2^2 \right)

    In practice, BiLLM searches for p∗p^* using a percentile search over candidate values p=i⋅max⁡(∣W∣)p = i \cdot \max(|\mathbf{W}|) for i∈{0.1,0.2,…,0.9}i \in \{0.1, 0.2, \dots, 0.9\}. The non-salient weights require 1 bit per weight for the binary sign and an additional 1-bit hardware mask to indicate group membership (concentrated vs. sparse).

  4. Knowl 4 — BiLLM Quantization Algorithm with Block-Wise Compensation

    algorithm

    BiLLM quantizes a weight matrix block-by-block using calibration data to construct the Hessian matrix, identifying salient columns, applying binary residual approximation to salient columns, segmenting and binarizing non-salient weights, and compensating for quantization errors across subsequent blocks via Optimal Brain Compression (OBC).

    func BinaryLLM(W, X, beta, lambda_reg)
        Input: W in R^{n x m} (weight matrix), X in R^{r x d} (calibration data), beta (block size), lambda_reg (Hessian regularizer)
        Output: B in R^{n x m} (binarized weight matrix)
        
        H := 2 * X * X^T
        H_c := Cholesky((H + lambda_reg * I)^{-1})
        B := zeros(n, m)
        
        for b = 0, beta, 2 * beta, ..., m - beta do
            W_block := W[:, b : b + beta]
            salient_cols := find_salient_columns(W_block, H_c[b : b + beta, b : b + beta])
            
            B1 := res_approximation(W_block[:, salient_cols])
            
            p_star := seg_search(W_block[:, not_in(salient_cols)])
            B2 := binary(W_block[abs(w) <= p_star, not_in(salient_cols)])
            B3 := binary(W_block[abs(w) > p_star, not_in(salient_cols)])
            
            B[:, b : b + beta] := B1 + B2 + B3
            
            E := (W[:, b : b + beta] - B[:, b : b + beta]) / (H_c[b : b + beta, b : b + beta])
            W[:, b + beta :] := W[:, b + beta :] - E * H_c[b : b + beta, b + beta :]
        end for
        return B

    The helper functions operate as follows:

    • find_salient_columns: Computes sensitivity matrix S=W2/[Hc]2S = W^2 / [H_c]^2, aggregates column sensitivity, sorts columns, and iterates to select the subset minimizing L2L_2 error.
    • res_approximation: Computes primary binary representation B1=α1sign⁡(W)B_1 = \alpha_1 \operatorname{sign}(W), computes residual R=W−B1R = W - B_1, computes residual binary representation B2=α2sign⁡(R)B_2 = \alpha_2 \operatorname{sign}(R), and returns B1+B2B_1 + B_2.
    • seg_search: Scans percentiles i∈{0.1,0.2,…,0.9}i \in \{0.1, 0.2, \dots, 0.9\} of max⁡(∣W∣)\max(|W|) to locate the break-point p∗p^* minimizing joint binarization error.
    • binary: Calculates scale α=∥W∥ℓ1/∣W∣\alpha = \|W\|_{\ell_1} / |W| and returns αsign⁡(W)\alpha \operatorname{sign}(W).
  5. Knowl 5 — Effective Bit-Width and Storage Overhead Formulation in BiLLM

    equation

    The average weight bit-width NparamN_{param} and total hardware storage overhead NstoringN_{storing} in BiLLM are defined as functions of the salient weight ratio and compensation block size:

    Nparam=2×rsalient+1×(1−rsalient)N_{param} = 2 \times r_{salient} + 1 \times (1 - r_{salient}) Nstoring=1+1bsizeN_{storing} = 1 + \frac{1}{b_{size}}

    where:

    • rsalient∈[0,1]r_{salient} \in [0, 1] represents the fraction of weights designated as salient (which use 2 bits due to binary residual approximation).
    • (1−rsalient)(1 - r_{salient}) represents the fraction of non-salient weights (which use 1 computational bit).
    • bsizeb_{size} is the block size used for OBC compensation (default is 128).
    • 1bsize\frac{1}{b_{size}} denotes the storage overhead for storing structured salient column identifiers.
    • 1 bit is allocated in NstoringN_{storing} to flag the binary division (concentrated vs. sparse) of non-salient weights.

    Flag bits do not participate in arithmetic operations during inference; computations are executed strictly on the 1-bit / 2-bit parameter weights. Across LLM architectures, selecting 10% salient columns with a block size of 128 yields an effective weight parameter bit-width of 1.07∼1.131.07 \sim 1.13 bits and a hardware flag overhead of 1.0081.008 bits.

  6. Knowl 6 — WikiText2 Perplexity of BiLLM Across OPT and LLaMA Model Families

    data/table

    The following tables document the WikiText2 perplexity (PPL, where lower is better) across various model sizes in the OPT, LLaMA, and LLaMA-2 model families, comparing Round-to-Nearest (RTN), GPTQ, PB-LLM, and BiLLM (with block size 128):

    Model Family Method Bits 1.3B/7B 2.7B/13B 6.7B/30B 13B/65B/70B 30B 66B
    OPT FP16 16.00 14.62 12.47 10.86 10.13 9.56 9.34
    OPT RTN 1.00 17165.72 36516.69 11550.91 6986.35 6485.99 184796.30
    OPT GPTQ 1.00 14884.73 14144.58 10622.81 15196.96 12478.37 13106.45
    OPT PB-LLM 1.70 265.52 124.35 105.16 81.92 25.14 29.09
    OPT BiLLM 1.11 69.97 49.55 35.36 18.82 12.71 12.06
    LLaMA FP16 16.00 5.68 5.09 4.10 3.53 - -
    LLaMA GPTQ 2.00 152.31 20.44 13.01 8.78 - -
    LLaMA GPTQ 1.00 267001.72 113894.12 67093.73 25082.88 - -
    LLaMA PB-LLM 1.70 102.36 36.60 33.67 12.53 - -
    LLaMA BiLLM 1.09 35.04 15.14 10.52 8.49 - -
    LLaMA2 FP16 16.00 5.47 4.88 - 3.32 - -
    LLaMA2 GPTQ 2.00 60.45 19.70 - 9.12 - -
    LLaMA2 PB-LLM 1.70 69.20 151.09 - 28.37 - -
    LLaMA2 BiLLM 1.08 32.48 16.77 - 8.41 - -

    Standard 1-bit PTQ baselines (RTN and GPTQ) experience complete perplexity collapse (PPL >104> 10^4). BiLLM achieves functional inference at 1.08∼1.111.08 \sim 1.11 bits, reducing the average bit-width by ∼35%\sim 35\% compared to PB-LLM (1.70 bits) while improving perplexity by 49.4%49.4\% to 77.0%77.0\% on OPT. On LLaMA-65B and LLaMA2-70B, BiLLM reaches perplexities of 8.49 and 8.41, outperforming the full-precision FP16 OPT-66B model (9.34).

  7. Knowl 7 — Zero-Shot Common Sense Reasoning Accuracy of BiLLM Across Benchmarks

    data/table

    Zero-shot task accuracy (%, where higher is better) for LLaMA-7B, LLaMA2-7B, and OPT-6.7B on seven standard reasoning benchmarks evaluated with block size 128:

    Model Method Bits PIQA BoolQ OBQA Wino. ARC-e ARC-c Hella.
    LLaMA-7B GPTQ 2.00 52.8 50.0 28.2 49.3 26.6 29.5 26.3
    LLaMA-7B PB-LLM 1.70 54.6 59.7 30.4 50.6 28.2 24.6 28.7
    LLaMA-7B BiLLM 1.09 61.2 62.7 31.8 51.1 36.0 25.7 36.8
    LLaMA2-7B GPTQ 2.00 51.1 43.9 29.0 50.8 26.6 28.5 26.3
    LLaMA2-7B PB-LLM 1.70 53.8 62.3 30.2 49.3 28.0 25.0 27.7
    LLaMA2-7B BiLLM 1.08 60.6 61.8 33.2 52.4 36.2 24.4 34.8
    OPT-6.7B GPTQ 2.00 56.6 51.1 25.6 51.2 31.3 22.9 30.4
    OPT-6.7B PB-LLM 1.70 57.6 55.5 24.2 47.7 33.2 21.0 31.0
    OPT-6.7B BiLLM 1.11 58.6 62.2 29.0 51.5 34.1 23.9 31.9

    Despite operating at ≈1.08–1.11\approx 1.08\text{--}1.11 bits per weight, BiLLM consistently outperforms 2-bit GPTQ and 1.7-bit PB-LLM across PIQA, BoolQ, OpenBookQA (OBQA), Winogrande (Wino.), ARC-easy (ARC-e), and Hellaswag (Hella.).

  8. Knowl 8 — Perplexity Evaluation of BiLLM on Instruction-Tuned Vicuna Models

    data/table

    Evaluation of BiLLM on instruction-tuned Vicuna models across WikiText2, PTB, and C4 datasets with block size 128:

    Model Method Weight Bits WikiText2 ↓\downarrow PTB ↓\downarrow C4 ↓\downarrow
    Vicuna-7B GPTQ 2.00 109.56 6227.73 64.28
    Vicuna-7B PB-LLM 1.70 68.01 477.52 67.23
    Vicuna-7B BiLLM 1.08 33.00 332.17 36.24
    Vicuna-13B GPTQ 2.00 41.75 465.94 40.57
    Vicuna-13B PB-LLM 1.70 362.17 772.44 346.16
    Vicuna-13B BiLLM 1.08 36.57 300.31 28.76

    BiLLM maintains stable generation on instruction-fine-tuned models at an average bit-width of 1.08 bits, outperforming both 2.00-bit GPTQ and 1.70-bit PB-LLM across all three evaluation datasets.

  9. Knowl 9 — Memory Occupancy and Model Compression Ratios of BiLLM

    data/table

    BiLLM achieves approximately a 10-fold compression in model storage relative to FP16 baselines across different LLM parameter sizes:

    Format LLaMA-7B LLaMA2-7B LLaMA-13B LLaMA2-13B LLaMA-30B LLaMA-65B LLaMA-70B
    FP16 13.5 GB 13.5 GB 24.2 GB 25.0 GB 60.5 GB 121.0 GB 129.3 GB
    BiLLM 1.5 GB 1.6 GB 2.7 GB 2.8 GB 6.1 GB 14.8 GB 15.4 GB

    Memory occupancy relative to FP16 is computed as: Memory Occupancy=sizebinary_unsalient+sizeres_salient+sizeCSR_bitmap+sizescalessizeFP16\text{Memory Occupancy} = \frac{\text{size}_{\text{binary\_unsalient}} + \text{size}_{\text{res\_salient}} + \text{size}_{\text{CSR\_bitmap}} + \text{size}_{\text{scales}}}{\text{size}_{\text{FP16}}}

    For OPT-30B, BiLLM achieves 9.70%9.70\% memory occupancy (1.11 average bits, 12.71 WikiText2 PPL) compared to 16.50%16.50\% for PB-LLM (1.7 bits, 25.14 PPL) and 13.30%13.30\% for GPTQ (2 bits, 15.71 PPL), representing a 41.57%41.57\% memory reduction over PB-LLM and 27.07%27.07\% over 2-bit GPTQ.

  10. Knowl 10 — Ablation of Salient Residual Approximation, Bell-Shaped Splitting, and Block Size

    empirical result

    Decomposition experiments demonstrate the distinct roles of the two core quantization mechanisms in BiLLM:

    • Salient Residual Approximation vs. Non-Salient Splitting: For OPT-6.7B, non-salient bell-shaped distribution splitting contributes more substantially to perplexity reduction than salient residual approximation alone. Conversely, LLaMA-7B exhibits greater sensitivity to salient weight residual approximation than to non-salient splitting. Combining both strategies (BiLLM) yields strictly superior perplexity across all models and benchmarks.
    • Quantization Block Size: Varying the block size β∈{32,64,128,256,512}\beta \in \{32, 64, 128, 256, 512\} reveals that smaller chunk sizes improve perplexity (e.g., on LLaMA-7B, WikiText2 PPL drops from 74.14 at β=512\beta=512 to 17.56 at β=32\beta=32) by providing finer representation granularity, but smaller blocks increase the scale parameter overhead. A block size of β=128\beta=128 was selected as the optimal operating point balancing bit-rate overhead and quantization accuracy.

Coverage note — None was omitted; all key contributions including framework architecture, theoretical formulations, algorithms, ablation findings, and comprehensive empirical benchmarks were captured.

References

  1. 1.Bengio, Y., Leonard, N., and Courville, A. Estimating or prop- agating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  2. 2.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  3. 3.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight uncertainty in neural network. In International conference on machine learning, pp. 1613–1622. PMLR, 2015.
  4. 4.Chan, C.-Y. and Ioannidis, Y. E. Bitmap index design and evaluation. In Proceedings of the 1998 ACM SIGMOD international conference on Management of data, pp. 355–366, 1998.
  5. 5.Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M. Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems, 36, 2024.
  6. 6.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  7. 7.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  8. 8.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  9. 9.Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., and Bengio, Y. Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830, 2016.
  10. 10.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale, 2022. CoRR abs/2208.07339.
  11. 11.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  12. 12.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023a.
  13. 13.Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078, 2023b.
  14. 14.Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. Hawq: Hessian aware quantization of neural networks with mixed-precision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 293–302, 2019.
  15. 15.Fang, J., Shafiee, A., Abdel-Aziz, H., Thorsley, D., Georgiadis, G., and Hassoun, J. H. Post-training piecewise linear quantization for deep neural networks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pp. 69–86. Springer, 2020.
  16. 16.Faraone, J., Fraser, N., Blott, M., and Leong, P. H. Syq: Learning symmetric quantization for efficient deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4300–4309, 2018.
  17. 17.Frantar, E. and Alistarh, D. Optimal brain compression: A frame- work for accurate post-training quantization and pruning. Ad- vances in Neural Information Processing Systems, 35:4475– 4488, 2022.
  18. 18.Frantar, E. and Alistarh, D. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Confer- ence on Machine Learning, pp. 10323–10337. PMLR, 2023.
  19. 19.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  20. 20.Gholami, A., Yao, Z., Kim, S., Hooper, C., Mahoney, M. W., and Keutzer, K. Ai and memory wall. IEEE Micro, 2024.
  21. 21.Helwegen, K., Widdicombe, J., Geiger, L., Liu, Z., Cheng, K.-T., and Nusselder, R. Latent weights do not exist: Rethinking binarized neural network optimization. Advances in neural information processing systems, 32, 2019.
  22. 22.Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2704–2713, 2018.
  23. 23.Jain, S., Venkataramani, S., Srinivasan, V., Choi, J., Gopalakrish- nan, K., and Chang, L. Biscaled-dnn: Quantizing long-tailed datastructures with two scale factors for deep neural networks. In Proceedings of the 56th Annual Design Automation Confer- ence 2019, pp. 1–6, 2019.
  24. 24.LeCun, Y., Denker, J., and Solla, S. Optimal brain damage. Ad- vances in neural information processing systems, 2, 1989.
  25. 25.Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. Owq: Lessons learned from activation outliers for weight quantization in large language models. arXiv preprint arXiv:2306.02272, 2023.
  26. 26.Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., and Gu, S. Brecq: Pushing the limit of post- training quantization by block reconstruction. arXiv preprint arXiv:2102.05426, 2021.
  27. 27.Li, Z., Ni, B., Zhang, W., Yang, X., and Gao, W. Performance guaranteed network acceleration via high-order residual quanti- zation. In Proceedings of the IEEE international conference on computer vision, pp. 2584–2592, 2017.
  28. 28.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  29. 29.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888, 2023.
  30. 30.Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B. The penn treebank: Annotating predicate argument structure. In Hu- man Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994.
  31. 31.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  32. 32.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  33. 33.Park, E., Yoo, S., and Vajda, P. Value-aware quantization for training and inference of neural networks. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 580– 595, 2018.
  34. 34.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Py- torch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  35. 35.Qin, H., Gong, R., Liu, X., Shen, M., Wei, Z., Yu, F., and Song, J. Forward and backward information retention for accurate binary neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2250–2259, 2020.
  36. 36.Qin, H., Ding, Y., Zhang, M., Yan, Q., Liu, A., Dang, Q., Liu, Z., and Liu, X. Bibert: Accurate fully binarized bert. arXiv preprint arXiv:2203.06390, 2022.
  37. 37.Qin, H., Zhang, M., Ding, Y., Li, A., Cai, Z., Liu, Z., Yu, F., and Liu, X. Bibench: Benchmarking and analyzing network binarization. arXiv preprint arXiv:2301.11233, 2023.
  38. 38.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of trans- fer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  39. 39.Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A. Xnor- net: Imagenet classification using binary convolutional neural networks. In European conference on computer vision, pp. 525–542. Springer, 2016.
  40. 40.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Wino- grande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  41. 41.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. Multitask prompted training enables zero-shot task gener- alization. arXiv preprint arXiv:2110.08207, 2021.
  42. 42.Shang, Y., Yuan, Z., Wu, Q., and Dong, Z. Pb-llm: Partially bina- rized large language models. arXiv preprint arXiv:2310.00034, 2023.
  43. 43.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. ` Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  44. 44.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  45. 45.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  46. 46.Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y., and Wei, F. Bitnet: Scaling 1-bit transformers for large language models. arXiv preprint arXiv:2310.11453, 2023.
  47. 47.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  48. 48.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Hug- gingface’s transformers: State-of-the-art natural language pro- cessing. arXiv preprint arXiv:1910.03771, 2019.
  49. 49.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023.
  50. 50.Yao, Z., Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y. Z. Efficient and affordable post-training quantiza- tion for large-scale transformers, 2022. URL https://arxiv. org/abs/2206.01861.
  51. 51.Yao, Z., Li, C., Wu, X., Youn, S., and He, Y. A comprehensive study on post-training quantization for large language models. arXiv preprint arXiv:2303.08302, 2023.
  52. 52.You, Y. Audio coding: theory and applications. Springer Science & Business Media, 2010.
  53. 53.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  54. 54.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  55. 55.Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. Dorefa- net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016.
  56. 56.Zhu, X., Li, J., Liu, Y., Ma, C., and Wang, W. A survey on model compression for large language models. arXiv preprint arXiv:2308.07633, 2023.

Citation

MLA
Huang, W., et al. “BiLLM: Pushing the Limit of Post-Training Quantization for LLMs”. arXiv, 2024, http://arxiv.org/abs/2402.04291v2.
APA
Huang, W., Liu, Y., Qin, H., Li, Y., Zhang, S., Liu, X., Magno, M., & Qi, X. (2024). BiLLM: Pushing the Limit of Post-Training Quantization for LLMs. arXiv. http://arxiv.org/abs/2402.04291v2
Chicago
Huang, W., Y. Liu, H. Qin, et al. 2024. “BiLLM: Pushing the Limit of Post-Training Quantization for LLMs”. arXiv. http://arxiv.org/abs/2402.04291v2.
Harvard
Huang, W. et al. (2024) “BiLLM: Pushing the Limit of Post-Training Quantization for LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.04291v2.
Vancouver
1. Huang W, Liu Y, Qin H, Li Y, Zhang S, Liu X, Magno M, Qi X (2024) BiLLM: Pushing the Limit of Post-Training Quantization for LLMs. arXiv

BibTeX

@article{huang2024billm,
  title = {BiLLM: Pushing the Limit of Post-Training Quantization for LLMs},
  author = {Huang, Wei and Liu, Yangdong and Qin, Haotong and Li, Ying and Zhang, Shiming and Liu, Xianglong and Magno, Michele and Qi, Xiaojuan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.04291v2},
  eprint = {2402.04291}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/