BiT: Robustly Binarized Multi-distilled Transformer

Zechun LiuBarlas OguzAasish PappuLin XiaoScott YihMeng LiRaghuraman KrishnamoorthiYashar Mehdad

article2022NeurIPS84 citations

Proposes a multi-stage distillation strategy and elastic activation binarization that enable fully 1-bit transformer models to close the performance gap with full-precision BERT to within 5.9 points on GLUE.

Listen

Modern transformer models drive major breakthroughs across artificial intelligence, but their substantial memory footprints and computational demands make them difficult to run on resource-limited hardware like mobile devices and wearables. Model binarization—reducing weights and activations to single bits—theoretically shrinks model storage by roughly thirty-two times and replaces expensive mathematical operations with efficient bitwise logic. Historically, attempting extreme 1-bit compression on transformers caused severe optimization issues and catastrophic drops in task accuracy.

The article demonstrates a robust binarization and multi-stage distillation framework called BiT (Binarized Transformer) that substantially closes the performance gap between fully binarized transformer networks and standard full-precision baselines.

The researchers developed a tailored binarization strategy that accounts for differing activation distributions across transformer components, combined with an elastic activation function that dynamically learns optimal scaling and threshold values during training. To ease the severe optimization challenge of 1-bit training, the authors implemented a multi-stage distillation approach that transfers knowledge progressively from a full-precision teacher to an intermediate 2-bit activation student before finally compressing to 1-bit precision. They evaluated these methods by compressing pre-trained BERT-base models across the multi-task GLUE benchmark and the SQuAD reading comprehension dataset.

The findings establish that the proposed framework delivers state-of-the-art performance for extremely compressed transformers. On the GLUE benchmark without data augmentation, BiT achieves an average score of 73.5, cutting the performance gap to the full-precision baseline by about 50% compared to previous binary approaches. When paired with standard data augmentation, BiT trails the full-precision baseline by only 5.9 points. In intermediate configurations using 1-bit weights and 2-bit activations, the model reaches within 3.5 points of the baseline while retaining significant hardware execution advantages. On more complex reading comprehension tasks, BiT scores 74.9 F1 on SQuAD, providing functional utility where earlier 1-bit architectures suffered total breakdown.

These results demonstrate that extreme low-bit compression is practically viable for real-world natural language processing deployments, enabling substantial decreases in hardware cost, power consumption, and memory requirements on edge devices. Because the framework trains binary weight models via direct knowledge distillation without requiring specialized half-width model pre-training, it also simplifies operational training pipelines for compressed deployments.

Organizations evaluating edge AI deployments should explore 1-bit or 2-bit quantized transformers as efficient alternatives for classification tasks, balancing model compression against acceptable task-level accuracy tolerances. Before deploying to complex generative or extraction workloads, teams should conduct targeted pilot validations and further explore optimal multi-step distillation schedules.

While confidence is high regarding classification performance on standard natural language benchmarks, readers should note that the evaluation is limited to BERT-base models and text understanding tasks. Extreme binarization continues to show a larger performance gap on intricate comprehension tasks like SQuAD, and further empirical validation is required before generalizing these findings to generative language models or other modalities.

Cover for BiT: Robustly Binarized Multi-distilled Transformer

Abstract

Modern pre-trained transformers have rapidly advanced the state-of-the-art in machine learning, but have also grown in parameters and computational complexity, making them increasingly difficult to deploy in resource-constrained environments. Binarization of the weights and activations of the network can significantly alleviate these issues, however, is technically challenging from an optimization perspective. In this work, we identify a series of improvements that enables binary transformers at a much higher accuracy than what was possible previously. These include a two-set binarization scheme, a novel elastic binary activation function with learned parameters, and a method to quantize a network to its limit by successively distilling higher precision models into lower precision students. These approaches allow for the first time, fully binarized transformer models that are at a practical level of accuracy, approaching a full-precision BERT baseline on the GLUE language understanding benchmark within as little as 5.9%. Code and models are available at: https://github.com/facebookresearch/bit.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Transformer architecture
  • 2.2 Quantization
  • 2.3 Knowledge distillation
  • 3 Robust binarization setup
  • 3.1 Two-set binarization scheme
  • 3.2 Best practices
  • 3.3 Elastic binarization function
  • 4 Multi-distilled binary transformer
  • 5 Main results
  • 5.1 GLUE results
  • 5.1.1 Data augmentation
  • 5.2 SQuAD results
  • 5.3 Ablations
  • 5.4 Learned parameter visualization
  • 5.5 Exploring multi-distillation paths
  • 6 Related work
  • 7 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Two-Set Activation Binarization Scheme for Transformers

    model/method

    In transformer architectures, activations across layers perform distinct functions and have different distributions. The BiT framework partitions transformer activation layers into two sets based on their range:

    1. Non-negative activations: Layers following Softmax (self-attention probabilities) and ReLU non-linearities, where XR∈R+nX_R \in \mathbb{R}_+^n, are quantized to unipolar binary representations X^B∈{0,1}n\hat{X}_B \in \{0, 1\}^n: X^Bi=⌊Clip(XRi,0,1)⌉={0if XRi<0.51if XRi≥0.5\hat{X}_B^i = \lfloor \text{Clip}(X_R^i, 0, 1) \rceil = \begin{cases} 0 & \text{if } X_R^i < 0.5 \\ 1 & \text{if } X_R^i \ge 0.5 \end{cases} where Clip(x,0,1)=min⁡(max⁡(x,0),1)\text{Clip}(x, 0, 1) = \min(\max(x, 0), 1) and ⌊⋅⌉\lfloor \cdot \rceil denotes rounding to the nearest integer.

    2. Signed activations: All other activation layers (e.g., outputs of linear projections and matrix multiplications) where XR∈RnX_R \in \mathbb{R}^n, are quantized to bipolar binary values X^B∈{−1,+1}n\hat{X}_B \in \{-1, +1\}^n: X^Bi=Sign(XRi)={−1if XRi<0+1if XRi≥0\hat{X}_B^i = \text{Sign}(X_R^i) = \begin{cases} -1 & \text{if } X_R^i < 0 \\ +1 & \text{if } X_R^i \ge 0 \end{cases}

    To minimize the reconstruction error J(α)=∥XR−αX^B∥22J(\alpha) = \|X_R - \alpha \hat{X}_B\|_2^2, optimal layer-wise scaling factors α∗=arg⁡min⁡α∈R+J(α)\alpha^* = \arg\min_{\alpha \in \mathbb{R}_+} J(\alpha) are derived analytically:

    • For signed activations (XR∈RnX_R \in \mathbb{R}^n): α∗=XRTX^BnXR=∥XR∥1nXR\alpha^* = \frac{X_R^T \hat{X}_B}{n_{X_R}} = \frac{\|X_R\|_1}{n_{X_R}} where nXRn_{X_R} is the total number of entries in XRX_R.
    • For non-negative activations (XR∈R+nX_R \in \mathbb{R}_+^n): α∗=∥XR⋅1{XR≥0.5}∥1n{XR≥0.5}\alpha^* = \frac{\|X_R \cdot \mathbf{1}_{\{X_R \ge 0.5\}}\|_1}{n_{\{X_R \ge 0.5\}}} where 1{⋅}\mathbf{1}_{\{\cdot\}} is an indicator function and n{XR≥0.5}n_{\{X_R \ge 0.5\}} is the number of elements in XRX_R satisfying XRi≥0.5X_R^i \ge 0.5.
  2. Knowl 2 — Elastic Binarization Function with Learnable Scale and Threshold

    model/method

    To dynamically adjust to activation distributions during training rather than using fixed clipping and scaling, the elastic binarization function incorporates a learnable positive scale parameter α∈R+\alpha \in \mathbb{R}_+ and a learnable shift threshold β∈R\beta \in \mathbb{R}.

    For non-negative activation layers (XR∈R+nX_R \in \mathbb{R}_+^n), the scaled binarized activation XBi=αX^BiX_B^i = \alpha \hat{X}_B^i is defined as: XBi=α⌊Clip(XRi−βα,0,1)⌉X_B^i = \alpha \left\lfloor \text{Clip}\left(\frac{X_R^i - \beta}{\alpha}, 0, 1\right) \right\rceil where α\alpha is initialized to α∗=∥XR⋅1{XR≥0.5}∥1n{XR≥0.5}\alpha^* = \frac{\|X_R \cdot \mathbf{1}_{\{X_R \ge 0.5\}}\|_1}{n_{\{X_R \ge 0.5\}}} and β\beta is initialized to 00. Using the straight-through estimator (STE) to bypass the non-differentiable rounding operator ⌊⋅⌉\lfloor \cdot \rceil, gradients with respect to α\alpha and β\beta are evaluated as: ∂XBi∂α≈STEX^Bi+α∂Clip(XRi−βα,0,1)∂α={0if XRi<ββ−XRiαif β≤XRi<α2+β1−XRi−βαif α2+β≤XRi<α+β1if XRi≥α+β\frac{\partial X_B^i}{\partial \alpha} \stackrel{\text{STE}}{\approx} \hat{X}_B^i + \alpha \frac{\partial \text{Clip}\left(\frac{X_R^i - \beta}{\alpha}, 0, 1\right)}{\partial \alpha} = \begin{cases} 0 & \text{if } X_R^i < \beta \\ \frac{\beta - X_R^i}{\alpha} & \text{if } \beta \le X_R^i < \frac{\alpha}{2} + \beta \\ 1 - \frac{X_R^i - \beta}{\alpha} & \text{if } \frac{\alpha}{2} + \beta \le X_R^i < \alpha + \beta \\ 1 & \text{if } X_R^i \ge \alpha + \beta \end{cases} ∂XBi∂β≈STEα∂Clip(XRi−βα,0,1)∂β={−1if β≤XRi<α+β0otherwise\frac{\partial X_B^i}{\partial \beta} \stackrel{\text{STE}}{\approx} \alpha \frac{\partial \text{Clip}\left(\frac{X_R^i - \beta}{\alpha}, 0, 1\right)}{\partial \beta} = \begin{cases} -1 & \text{if } \beta \le X_R^i < \alpha + \beta \\ 0 & \text{otherwise} \end{cases}

    For signed activation layers (XR∈RnX_R \in \mathbb{R}^n), the formulation reduces to: XBi=α⋅Sign(XRi−βα)=α⋅Sign(XRi−β)X_B^i = \alpha \cdot \text{Sign}\left(\frac{X_R^i - \beta}{\alpha}\right) = \alpha \cdot \text{Sign}(X_R^i - \beta) with the scale gradient simplifying to: ∂XBi∂α=Sign(XRi−β)\frac{\partial X_B^i}{\partial \alpha} = \text{Sign}(X_R^i - \beta)

  3. Knowl 3 — Multi-Distillation Quantization Algorithm (BiT)

    algorithm

    Rather than attempting single-step distillation directly from a 32-bit floating-point teacher to a 1-bit binarized student, multi-distillation passes knowledge through a sequence of intermediate models of decreasing precision.

    Input: Training dataset DtrainD_{\text{train}}, validation dataset DdevD_{\text{dev}}, full-precision model h0h_0, quantization schedule Q={(bw1,ba1),(bw2,ba2),…,(bwk,bak)}Q = \{(b_w^1, b_a^1), (b_w^2, b_a^2), \ldots, (b_w^k, b_a^k)\} where (bw1,ba1)>(bw2,ba2)>…>(bwk,bak)(b_w^1, b_a^1) > (b_w^2, b_a^2) > \ldots > (b_w^k, b_a^k)
    Output: Distilled low-bit student model hstudenth_{\text{student}}
    hteacher←h0h_{\text{teacher}} \leftarrow h_0
    for each (bwi,bai)(b_w^i, b_a^i) in QQ do
        hstudent←Quantize(hteacher,bwi,bai)h_{\text{student}} \leftarrow \text{Quantize}(h_{\text{teacher}}, b_w^i, b_a^i)
        hstudent←KnowledgeDistill(hstudent,hteacher,Dtrain,Ddev)h_{\text{student}} \leftarrow \text{KnowledgeDistill}(h_{\text{student}}, h_{\text{teacher}}, D_{\text{train}}, D_{\text{dev}})
        hteacher←hstudenth_{\text{teacher}} \leftarrow h_{\text{student}}
    end for
    return hstudenth_{\text{student}}

    For a fully binary transformer with 1-bit weights and 1-bit activations (W1A1), the schedule is typically instantiated as a two-step process: W32A32→W1A2→W1A1W32A32 \rightarrow W1A2 \rightarrow W1A1. First, the full-precision model distill-trains an intermediate model with 1-bit weights and 2-bit activations (W1A2). Next, this trained W1A2 model serves as the teacher to distill into the target W1A1 student model.

  4. Knowl 4 — Simplified Knowledge Distillation Loss for Quantized Transformers

    model/method

    Training the low-precision student transformer utilizes a single-step, joint distillation objective combining prediction logits matching and intermediate layer representation matching, while omitting attention score matrix distillation: L=Llogits+Lreps\mathcal{L} = \mathcal{L}_{\text{logits}} + \mathcal{L}_{\text{reps}} where the logit distillation loss is the Kullback-Leibler divergence between teacher probabilities p\mathbf{p} and student probabilities q\mathbf{q}: Llogits=KL(p,q)=∑cpclog⁡(pcqc)\mathcal{L}_{\text{logits}} = \text{KL}(\mathbf{p}, \mathbf{q}) = \sum_c p_c \log\left(\frac{p_c}{q_c}\right) and the intermediate representation distillation loss is the squared L2L_2 error between transformer block output activations: Lreps=∑i=1N∥ris−rit∥22\mathcal{L}_{\text{reps}} = \sum_{i=1}^N \|\mathbf{r}_i^s - \mathbf{r}_i^t\|_2^2 where ris\mathbf{r}_i^s and rit\mathbf{r}_i^t denote the output activations of the ii-th transformer block for the student and teacher models, respectively, across NN total transformer blocks.

  5. Knowl 5 — Optimization Practices for Transformer Weight Binarization and Gradients

    model/method

    Three optimization techniques improve stability and convergence when binarizing transformer weights and activations:

    1. Zero-Mean Weight Centering: Prior to binarization, real-valued weights WR∈RnWRW_R \in \mathbb{R}^{n_{W_R}} are centered to zero mean to maximize information capacity: WBi=∥WR∥1nWRSign(WRi−WˉR)W_B^i = \frac{\|W_R\|_1}{n_{W_R}} \text{Sign}(W_R^i - \bar{W}_R) where WˉR=1nWR∑j=1nWRWRj\bar{W}_R = \frac{1}{n_{W_R}} \sum_{j=1}^{n_{W_R}} W_R^j.

    2. Asymmetric Gradient Clipping: Gradient clipping to zero outside the active range [−1,1][-1, 1] (or [0,1][0, 1] for unipolar activations) is applied exclusively to activations and deliberately omitted for weights. Because weights are static parameters, clipping a weight outside [−1,1][-1, 1] fixes its gradient permanently to zero, stalling learning; activation values, by contrast, vary dynamically across inputs.

    3. Non-Linearity Selection: Rectified Linear Unit (ReLU) activations are favored wherever non-negative output representations are produced.

  6. Knowl 6 — GLUE Benchmark Performance of Quantized and Binarized BERT

    data/table

    Evaluation of BERT-base quantization methods on the GLUE dev set across embedding, weight, and activation bitwidths (E-W-A). The label ‡\ddagger denotes direct single-step distillation from the 32-bit teacher without intermediate multi-distillation.

    Model #Bits (E-W-A) Size (MB) FLOPs (G) MNLI-(m/mm) QQP QNLI SST-2 CoLA STS-B MRPC RTE Avg.
    BERT 32-32-32 418 22.5 84.9/85.5 91.4 92.1 93.2 59.7 90.1 86.3 72.2 83.9
    Without data augmentation
    Q-BERT 2-8-8 43.0 6.5 76.6/77.0 – – 84.6 – – 68.3 52.7 –
    Q2BERT 2-8-8 43.0 6.5 47.2/47.3 67.0 61.3 80.6 0.0 4.4 68.4 52.7 47.7
    TernaryBERT 2-2-8 28.0 6.4 83.3/83.3 90.1 – – 50.7 – 87.5 68.2 –
    BinaryBERT 1-1-8 16.5 3.1 84.2/84.7 91.2 90.9 92.6 53.4 88.6 85.5 72.2 82.7
    BinaryBERT 1-1-4 16.5 1.5 83.9/84.2 91.2 90.9 92.3 44.4 87.2 83.3 65.3 79.9
    BinaryBERT 1-1-2 16.5 0.8 62.7/63.9 79.9 52.6 82.5 14.6 6.5 68.3 52.7 53.7
    BinaryBERT 1-1-1 16.5 0.4 35.6/35.3 66.2 51.5 53.2 0.0 6.1 68.3 52.7 41.0
    BiBERT 1-1-1 13.4 0.4 66.1/67.5 84.8 72.6 88.7 25.4 33.6 72.5 57.4 63.2
    BiT ‡\ddagger 1-1-4 13.4 1.5 83.6/84.4 87.8 91.3 91.5 42.0 86.3 86.8 66.4 79.5
    BiT ‡\ddagger 1-1-2 13.4 0.8 82.1/82.5 87.1 89.3 90.8 32.1 82.2 78.4 58.1 75.0
    BiT ‡\ddagger 1-1-1 13.4 0.4 77.1/77.5 82.9 85.7 87.7 25.1 71.1 79.7 58.8 71.0
    BiT 1-1-1 13.4 0.4 79.5/79.4 85.4 86.4 89.9 32.9 72.0 79.9 62.1 73.5
    With data augmentation
    TernaryBERT 2-2-8 28.0 6.4 83.3/83.3 90.1 90.0 92.9 47.8 84.3 82.6 68.4 80.3
    BinaryBERT 1-1-8 16.5 3.1 84.2/84.7 91.2 91.6 93.2 55.5 89.2 86.0 74.0 83.3
    BinaryBERT 1-1-4 16.5 1.5 83.9/84.2 91.2 91.4 93.7 53.3 88.6 86.0 71.5 82.6
    BinaryBERT 1-1-2 16.5 0.8 62.7/63.9 79.9 51.0 89.6 33.0 11.4 71.0 55.9 57.6
    BinaryBERT 1-1-1 16.5 0.4 35.6/35.3 66.2 66.1 78.3 7.3 22.1 69.3 57.7 48.7
    BiBERT 1-1-1 13.4 0.4 66.1/67.5 84.8 76.0 90.9 37.8 56.7 78.8 61.0 68.8
    BiT ‡\ddagger 1-1-2 13.4 0.8 82.1/82.5 87.1 88.8 92.5 43.2 86.3 90.4 72.9 80.4
    BiT ‡\ddagger 1-1-1 13.4 0.4 77.1/77.5 82.9 85.0 91.5 32.0 84.1 88.0 67.5 76.0
    BiT 1-1-1 13.4 0.4 79.5/79.4 85.4 86.5 92.3 38.2 84.2 88.0 69.7 78.0

    In the fully binary (1-1-1) setting without data augmentation, BiT achieves 73.5% average accuracy on GLUE, reducing the error gap to the 32-bit baseline (83.9%) by nearly half compared to prior work (BiBERT at 63.2%). With data augmentation on small datasets, BiT achieves 78.0%, coming within 5.9 points of the full-precision baseline.

  7. Knowl 7 — Ablation of BiT Components on GLUE

    data/table

    Step-by-step ablation showing the individual performance impact of each proposed component on the GLUE dev set without data augmentation:

    Row Configuration MNLI-(m/mm) QQP QNLI SST-2 CoLA STS-B MRPC RTE Avg.
    1 BERTbase\text{BERT}_{\text{base}} 84.9/85.5 91.4 92.1 93.2 59.7 90.1 86.3 72.2 83.9
    2 BiBERT Baseline 45.8/47.0 73.2 66.4 77.6 11.7 7.6 70.2 54.1 50.4
    3 BiBERT 66.1/67.5 84.8 72.6 88.7 25.4 33.6 72.5 57.4 63.2
    4 BinaryBERT (re-implementation) 36.2/35.9 59.6 52.4 65.6 9.3 19.8 69.9 52.7 45.7
    5 + Simplified KD 37.7/37.3 59.5 56.8 73.4 4.1 24.8 70.8 57.0 48.0
    6 + Two-set binarization (Strong Baseline) 57.4/59.1 68.3 64.7 81.0 18.2 24.7 71.8 56.7 55.3
    7 + Elastic binarization (BiT ‡\ddagger) 77.1/77.5 82.9 85.7 87.7 25.1 71.1 79.7 58.8 71.0
    8 + Multi-distillation (BiT) 79.5/79.4 85.4 86.4 89.9 32.9 72.0 79.9 62.1 73.5

    The gains from each sequential addition are:

    • Omitting attention distillation and adopting simplified KD yields +2.3% average accuracy (row 5 vs. row 4).
    • Adopting two-set binarization adds +7.3% (row 6 vs. row 5).
    • Replacing static binarization with the elastic binarization function provides the largest gain of +15.7% (row 7 vs. row 6), establishing a new state of the art at 71.0% even before multi-distillation.
    • Multi-distillation adds a further +2.5% (row 8 vs. row 7), reaching 73.5% average accuracy.
  8. Knowl 8 — Machine Reading Comprehension Performance on SQuAD v1.1

    empirical result

    Evaluation of quantized BERT models on the SQuAD v1.1 machine reading comprehension benchmark reveals a pronounced performance gap between sentence-level classification tasks and context-dependent span extraction.

    Model Bitwidth (E-W-A) SQuAD v1.1 Exact Match / F1 (%)
    BERTbase\text{BERT}_{\text{base}} 32-32-32 82.6 / 89.7
    BinaryBERT 1-1-4 77.9 / 85.8
    BinaryBERT 1-1-2 72.3 / 81.8
    BinaryBERT 1-1-1 1.5 / 8.2
    BiBERT 1-1-1 8.5 / 18.9
    BiT 1-1-1 63.1 / 74.9

    While prior fully binarized (1-1-1) transformer methods fail to achieve meaningful convergence on SQuAD v1.1 (scoring 1.5/8.2 and 8.5/18.9), BiT achieves 63.1% Exact Match and 74.9% F1. However, BiT still lags the full-precision baseline (89.7% F1) by 14.8 points, indicating that binarizing reading comprehension tasks remains an open challenge.

  9. Knowl 9 — Schedule Selection and Trade-offs in Multi-Distillation

    empirical result

    Empirical exploration of intermediate distillation schedules on the MNLI matched (MNLI-m) task reveals key dynamics governing representation transfer:

    1. Two-step distillation paths (W32A32→W1Ab→W1A1W32A32 \rightarrow W1Ab \rightarrow W1A1 for intermediate activation precisions b∈{2,4,8}b \in \{2, 4, 8\}) strictly outperform direct single-step distillation (W32A32→W1A1W32A32 \rightarrow W1A1).
    2. Comparing intermediate teachers, higher activation bitwidths produce higher intermediate model accuracy (W1A8>W1A4>W1A2W1A8 > W1A4 > W1A2). However, distilling from the lower-precision W1A2W1A2 intermediate teacher produces a more accurate final W1A1W1A1 student model than distilling from W1A8W1A8 or W1A4W1A4. This shows that teacher proximity to the target 1-bit representation space is more important for representation distillation than the raw task accuracy of the teacher.
    3. Extending to a 3-step distillation schedule (W32A32→W1A4→W1A2→W1A1W32A32 \rightarrow W1A4 \rightarrow W1A2 \rightarrow W1A1) does not yield additional accuracy gains over the 2-step schedule (W32A32→W1A2→W1A1W32A32 \rightarrow W1A2 \rightarrow W1A1).
  10. Knowl 10 — Limitations of Fully Binarized Transformers

    limitation

    The methodology and empirical findings of BiT exhibit several key limitations:

    1. Residual Accuracy Gap on Dense Token Tasks: On span-extraction reading comprehension (SQuAD v1.1), the fully binarized BiT model still lags the full-precision 32-bit baseline by 14.8 F1 points (74.9 vs. 89.7), demonstrating that 1-bit binarization causes notable degradation on complex token-level reasoning.
    2. Limited Scope of Model Architectures and Scales: Evaluations are confined to BERT-base models on English natural language understanding tasks; the behavior of elastic binarization and multi-distillation has not been evaluated on autoregressive decoder models (e.g., GPT variants), encoder-decoder models (e.g., T5), or multi-billion parameter architectures.
    3. Domain and Task Generalization: The approach has only been demonstrated on textual NLP benchmarks, leaving its effectiveness on vision transformers (ViT), multimodal models, and generative audio/speech models unverified.

Coverage note — None. All major contributions—including two-set activation binarization, elastic binarization parameterization and gradients, multi-distillation algorithm, simplified KD loss and training practices, GLUE and SQuAD benchmark evaluations, ablation experiments, distillation path exploration, and limitations—are fully covered.

References

  1. 1.Gustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Yao, Xing Fan, and Chenlei Guo. Knowledge distillation from internal representations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 7350–7357, 2020.
  2. 2.Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael R Lyu, and Irwin King. Binarybert: Pushing the limit of bert quantization. In ACL/IJCNLP (1), 2021.
  3. 3.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  4. 4.Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. Understanding and overcoming the challenges of efficient transformer quantization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7947–7969, 2021.
  5. 5.John Bridle. Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. Advances in neural information processing systems, 2, 1989.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  7. 7.Jungwook Choi, Zhuo Wang, Swagath Venkataramani, et al. Pact: Parameterized clipping activation for quantized neural networks. arXiv e-prints, pp. arXiv–1805, 2018.
  8. 8.Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830, 2016.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1), 2019.
  10. 10.Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha. Learned step size quantization. In International Conference on Learning Representations, 2019.
  11. 11.Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Herve Jegou, and Armand Joulin. Training with quantization noise for extreme model compression. arXiv preprint arXiv:2004.07320, 2020.
  12. 12.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  13. 13.Ruihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li, Peng Hu, Jiazhen Lin, Fengwei Yu, and Junjie Yan. Differentiable soft quantization: Bridging full-precision and low-bit neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4852–4861, 2019.
  14. 14.Jun Han and Claudio Moraga. The influence of the sigmoid function parameters on the speed of backpropagation learning. In International workshop on artificial neural networks, pp. 195–201. Springer, 1995.
  15. 15.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  16. 16.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  17. 17.Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Quantized neural networks: Training neural networks with low precision weights and activations. The Journal of Machine Learning Research, 18(1):6869–6898, 2017.
  18. 18.Sangil Jung, Changyong Son, Seohyung Lee, Jinwoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, and Changkyu Choi. Learning to quantize deep networks by optimizing quantization intervals with task loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4350–4359, 2019.
  19. 19.Jangho Kim, Yash Bhalgat, Jinwon Lee, Chirag Patel, and Nojun Kwak. Qkd: Quantization-aware knowledge distillation. arXiv preprint arXiv:1911.12491, 2019.
  20. 20.Yuhang Li, Xin Dong, and Wei Wang. Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. In International Conference on Learning Representations, 2020.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  22. 22.Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, and Kwang-Ting Cheng. Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. In Proceedings of the European conference on computer vision (ECCV), pp. 722–737, 2018.
  23. 23.Zechun Liu, Zhiqiang Shen, Marios Savvides, and Kwang-Ting Cheng. Reactnet: Towards precise binary neural network with generalized activation functions. In European Conference on Computer Vision, pp. 143–159. Springer, 2020.
  24. 24.Zechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P Xing, and Zhiqiang Shen. Nonuniform-to-uniform quantization: Towards accurate quantization via generalized straight-through estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4942–4952, 2022.
  25. 25.Brais Martinez, Jing Yang, Adrian Bulat, and Georgios Tzimiropoulos. Training binary neural networks with real-to-binary convolutions. In ICLR, 2020.
  26. 26.Daisuke Miyashita, Edward H Lee, and Boris Murmann. Convolutional neural networks using logarithmic data representation. arXiv preprint arXiv:1603.01025, 2016.
  27. 27.Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In Icml, 2010.
  28. 28.Eriko Nurvitadhi, David Sheffield, Jaewoong Sim, Asit Mishra, Ganesh Venkatesh, and Debbie Marr. Accelerating binarized neural networks: Comparison of fpga, cpu, gpu, and asic. In 2016 International Conference on Field-Programmable Technology (FPT), pp. 77–84. IEEE, 2016.
  29. 29.Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, and Jingkuan Song. Forward and backward information retention for accurate binary neural networks. In CVPR, 2020.
  30. 30.Haotong Qin, Yifu Ding, Mingyuan Zhang, YAN Qinghua, Aishan Liu, Qingqing Dang, Ziwei Liu, and Xianglong Liu. Bibert: Accurate fully binarized bert. In International Conference on Learning Representations, 2021.
  31. 31.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018.
  32. 32.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020.
  34. 34.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100, 000+ questions for machine comprehension of text. In EMNLP, 2016.
  35. 35.Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In European conference on computer vision, pp. 525–542. Springer, 2016.
  36. 36.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  37. 37.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 8815–8821, 2020.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  39. 39.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019.
  40. 40.Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan. Training deep neural networks with 8-bit floating point numbers. Advances in neural information processing systems, 31, 2018.
  41. 41.Yifan Yang, Qijing Huang, Bichen Wu, Tianjun Zhang, Liang Ma, Giulio Gambardella, Michaela Blott, Luciano Lavagno, Kees Vissers, John Wawrzynek, et al. Synetgy: Algorithm-hardware co-design for convnet accelerators on embedded fpgas. In Proceedings of the 2019 ACM/SIGDA international symposium on field-programmable gate arrays, pp. 23–32, 2019.
  42. 42.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pp. 811–824. IEEE, 2020.
  43. 43.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. Q8bert: Quantized 8bit bert. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS), pp. 36–39. IEEE, 2019.
  44. 44.Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua. Lq-nets: Learned quantization for highly accurate and compact deep neural networks. In Proceedings of the European conference on computer vision (ECCV), pp. 365–382, 2018.
  45. 45.Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. Ternarybert: Distillation-aware ultra-low bit BERT. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), EMNLP, 2020.
  46. 46.Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016.
  47. 47.Feng Zhu, Ruihao Gong, Fengwei Yu, Xianglong Liu, Yanfei Wang, Zhelong Li, Xiuqi Yang, and Junjie Yan. Towards unified int8 training for convolutional neural network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1969–1979, 2020.
  48. 48.Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, and Ian Reid. Towards effective low-bitwidth convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7920–7928, 2018.

Citation

MLA
Liu, Z., et al. “BiT: Robustly Binarized Multi-distilled Transformer”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 14303–16, https://proceedings.neurips.cc/paper_files/paper/2022/file/5c1863f711c721648387ac2ef745facb-Paper-Conference.pdf.
APA
Liu, Z., Oguz, B., Pappu, A., Xiao, L., Yih, S., Li, M., Krishnamoorthi, R., & Mehdad, Y. (2022). BiT: Robustly Binarized Multi-distilled Transformer. Advances in Neural Information Processing Systems, 35, 14303–14316. https://proceedings.neurips.cc/paper_files/paper/2022/file/5c1863f711c721648387ac2ef745facb-Paper-Conference.pdf
Chicago
Liu, Z., B. Oguz, A. Pappu, et al. 2022. “BiT: Robustly Binarized Multi-distilled Transformer”. Advances in Neural Information Processing Systems 35: 14303–16. https://proceedings.neurips.cc/paper_files/paper/2022/file/5c1863f711c721648387ac2ef745facb-Paper-Conference.pdf.
Harvard
Liu, Z. et al. (2022) “BiT: Robustly Binarized Multi-distilled Transformer”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 14303–14316. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/5c1863f711c721648387ac2ef745facb-Paper-Conference.pdf.
Vancouver
1. Liu Z, Oguz B, Pappu A, Xiao L, Yih S, Li M, Krishnamoorthi R, Mehdad Y (2022) BiT: Robustly Binarized Multi-distilled Transformer. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 14303–14316

BibTeX

@inproceedings{liu2022bit,
  title = {BiT: Robustly Binarized Multi-distilled Transformer},
  author = {Liu, Zechun and Oguz, Barlas and Pappu, Aasish and Xiao, Lin and Yih, Scott and Li, Meng and Krishnamoorthi, Raghuraman and Mehdad, Yashar},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {14303-14316},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/5c1863f711c721648387ac2ef745facb-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors