Oscillation-free Quantization for Low-bit Vision Transformers

Shih-Yang LiuZechun LiuKwang-Ting Cheng

article2023ICML72 citations

Develops statistical weight quantization, confidence-guided annealing, and query-key reparameterization to eliminate weight oscillation in quantization-aware training, significantly closing the accuracy gap between low-bit Vision Transformers and their full-precision counterparts on ImageNet.

Listen

Deploying advanced artificial intelligence models, such as vision transformers, on resource-constrained edge devices requires compressing them to operate with lower numerical precision, a technique known as quantization. However, low-bit quantization often causes a significant drop in accuracy. A primary driver of this performance loss during training is weight oscillation, where model parameters continually jump across discrete thresholds instead of settling into optimal values. The article investigates the underlying causes of this instability in vision transformers and evaluates a new quantization framework designed to eliminate weight oscillation and recover model accuracy.

The article demonstrates that standard learnable scaling factors—widely used to adjust quantization step sizes—create a destabilizing feedback loop with outlier weights, persistently driving adjacent weights into oscillation. Furthermore, the structural design of self-attention mechanisms causes coupled oscillations between internal components during training. To address these root causes, the researchers developed a three-part method called Oscillation-Free Quantization. This approach replaces learnable scaling with a stable statistical calculation, temporarily freezes stable parameters while fine-tuning uncertain ones until they exit boundary regions, and mathematically reorders internal operations to decouple interacting components. The framework was evaluated across standard vision transformer benchmarks using the ImageNet image classification dataset.

The experimental findings show substantial performance gains, particularly at extreme low-bit precisions where oscillation issues are most severe. For 2-bit models, the proposed framework improved top-1 classification accuracy by approximately 9.9% on DeiT-Tiny, 7.7% on DeiT-Small, and 4.6% on Swin-Tiny over previous leading methods. For 3-bit models, the framework achieved performance comparable to full-precision, uncompressed baselines. In 4-bit configurations, the quantized models consistently matched or slightly exceeded full-precision accuracy. Ablation analyses confirmed that all three proposed techniques contributed positively and worked cooperatively to stabilize training.

These results indicate that training instability, rather than the fundamental capacity limit of compressed architectures, has been a primary barrier to deploying highly compressed vision models. Eliminating oscillation reduces the deployment footprint and computational cost of transformer models without incurring major accuracy penalties. Practitioners seeking to deploy vision transformers on edge hardware should adopt statistical scaling and targeted annealing strategies in their compression pipelines. Organizations should test this framework across other transformer architectures and downstream tasks to determine if these stabilization benefits generalize broadly across deep learning domains.

arXiv: 2302.02210
Cover for Oscillation-free Quantization for Low-bit Vision Transformers

Abstract

Weight oscillation is an undesirable side effect of quantization-aware training, in which quantized weights frequently jump between two quantized levels, resulting in training instability and a sub-optimal final model. We discover that the learnable scaling factor, a widely-used de facto setting in quantization aggravates weight oscillation. In this study, we investigate the connection between the learnable scaling factor and quantized weight oscillation and use ViT as a case driver to illustrate the findings and remedies. In addition, we also found that the interdependence between quantized weights in query and key of a self-attention layer makes ViT vulnerable to oscillation. We, therefore, propose three techniques accordingly: statistical weight quantization (StatsQ) to improve quantization robustness compared to the prevalent learnable-scale-based method; confidence-guided annealing (CGA) that freezes the weights with high confidence and calms the oscillating weights; and query-key reparameterization (QKR) to resolve the query-key intertwined oscillation and mitigate the resulting gradient misestimation. Extensive experiments demonstrate that these proposed techniques successfully abate weight oscillation and consistently achieve substantial accuracy improvement on ImageNet. Specifically, our 2-bit DeiT-T/DeiT-S/Swin-T algorithms outperform the previous state-of-the-art by 9.8%/7.7%/4.64%, respectively. Code and models are available at: https://github.com/nbasyl/OFQ.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Preliminary
  • 3.1. Quantization-Aware Training
  • 3.2. Quantized ViT Architecture
  • 4. Oscillation in QAT
  • 4.1. Oscillation and Learnable Scaling Factor
  • 4.2. Oscillations in Quantized ViTs
  • 4.2.1. THE EFFECT OF LEARNABLE QUANTIZATION ON QUANTIZED VITS
  • 4.2.2. NOISE INJECTION ANALYSIS
  • 4.2.3. COMPOSITION OF THE OSCILLATING WEIGHTS
  • 4.2.4. BOTTLENECK OF QUANTIZED VITS
  • 5. Conquering Oscillation in Quantized ViT
  • 5.1. Statistical Weight Quantization
  • 5.2. Confidence-Guided Annealing
  • 5.3. Query-Key Reparametrization
  • 6. Experiments
  • 6.1. Implementation Details
  • 6.2. Main Results
  • 6.3. Ablation Study
  • 7. Conclusion
  • 8. Acknowledgements
  • References
  • 9. Appendix
  • 9.1. Progression of αs in CGA with Different BRx

Knowls

  1. Knowl 1 — Learnable scaling factors amplify quantization oscillation

    empirical result

    The paper identifies a feedback loop in low-bit quantization-aware training (QAT) with a learnable step size. For a real-valued weight W(i)W^{(i)}, step size α>0\alpha>0, and integer quantization range [Qn,Qp][Q_n,Q_p], LSQ quantizes by

    Wq(i)=α⌊Clip⁡(W(i)α,Qn,Qp)⌉,W_q^{(i)}=\alpha\left\lfloor\operatorname{Clip}\left(\frac{W^{(i)}}{\alpha},Q_n,Q_p\right)\right\rceil,

    where Clip⁡\operatorname{Clip} clips to the stated range and ⌊⋅⌉\lfloor\cdot\rceil rounds to the nearest integer. Weights close to a quantization threshold can cross between adjacent levels, causing their gradients to change direction and their quantized values to jump. The noisy gradients from these weights perturb the learnable step size, which moves the thresholds and causes additional weights to cross them.

    In the paper's three-weight regression example, an outlier initialized far from its optimum contributes disproportionately to the step-size gradient. The resulting change in the shared threshold makes other weights that were initially near their optima oscillate. In 2-bit DeiT-T, trajectories of learnable scaling factors remain visibly fluctuating late in training, even when the learning rate is small; weight histograms simultaneously show many weights clustered near quantization thresholds. The authors therefore attribute late-training instability and suboptimal convergence to the interaction between latent weights and moving learnable thresholds, rather than to the discrete quantizer alone.

  2. Knowl 2 — Statistical weight quantization

    model/method

    Statistical weight quantization (StatsQ) replaces the learned LSQ step size with a statistic-based scale designed to reduce the disproportionate effect of outlier weights:

    Wq=αs(⌊Clip⁡(Wαs,−1,1)⋅n−0.5⌉+0.5)1n,αs=2∥W∥1∣W∣,n=2b−1.W_q=\alpha_s\left(\left\lfloor \operatorname{Clip}\left(\frac{W}{\alpha_s},-1,1\right)\cdot n-0.5\right\rceil+0.5\right)\frac{1}{n}, \qquad \alpha_s=2\frac{\lVert W\rVert_1}{|W|}, \qquad n=2^{b-1}.

    Here WW is a real-valued weight tensor, WqW_q is its quantized tensor, bb is the weight bit-width, ∣W∣|W| is the number of entries in WW, and ∥W∥1\lVert W\rVert_1 is the sum of their absolute values. The clipping function limits normalized weights to [−1,1][-1,1], and ⌊⋅⌉\lfloor\cdot\rceil denotes nearest-integer rounding. The scale αs\alpha_s has the same units as WW and is recomputed from the tensor statistics rather than updated by noisy gradient descent.

    The method is motivated by the authors' claim that an approximately even use of quantization levels preserves information entropy. Clipping limits outlier influence, while the mean absolute value gives every weight an equal contribution to the scale statistic. In 2-bit DeiT-T, the resulting scale trajectories are substantially smoother and nearly stationary near the end of training compared with learnable LSQ scales.

  3. Knowl 3 — Confidence-guided annealing

    algorithm

    Confidence-guided annealing (CGA) is a final fine-tuning procedure that freezes high-confidence weights and updates only weights close to quantization thresholds. A weight is considered high-confidence when its real-valued latent value is far from the nearest threshold; a weight inside a boundary range BRxBR_x is considered low-confidence. The paper uses BR0.005BR_{0.005} as its default setting in the main experiments and applies CGA for 25 fine-tuning epochs.

    Given a weight matrix WW, statistical scale αs\alpha_s, bit-width bb, boundary range BRxBR_x, and nn annealing iterations, the procedure is:

    Input: Weight matrix W, statistical scale α_s, bit-width b, boundary range BR_x, annealing iterations n
    Set quantization-level parameter n_levels = 2^(b - 1)
    for j = 1 to n do
        Compute the gradient G_j = ∂L_j/∂W
        Compute the normalized pre-rounding values W_fq = Clip(W / α_s, -1, 1) · n_levels - 0.5
        Construct a mask that is 1 for entries W_fq in BR_x and 0 otherwise
        Multiply G_j entrywise by the mask
        Update W using the masked gradient
    end for
    Output: Annealed weight matrix W

    The mask makes weights outside BRxBR_x receive zero gradient and therefore freezes them. Weights initially inside the boundary continue to be optimized; once they leave the boundary, they are frozen as well. This reverses the strategy of freezing oscillating weights first: the paper argues that fixing high-confidence weights provides stable anchors for optimizing low-confidence weights and prevents already-stable weights from re-entering the oscillation region.

  4. Knowl 4 — Query-key reparameterization

    model/method

    The paper finds that quantized query and key weights in self-attention can amplify each other's oscillation. Let XX be the input token matrix, and let WQW_Q and WKW_K be compatible query and key projection matrices, with XQ=XWQTX_Q=XW_Q^{\mathsf T} and XK=XWKTX_K=XW_K^{\mathsf T}. Their unquantized product can be reordered as

    XQXKT=X(WQTWK)XT.X_QX_K^{\mathsf T}=X\left(W_Q^{\mathsf T}W_K\right)X^{\mathsf T}.

    Query-key reparameterization (QKR) quantizes after this reordering:

    Xout=Fq(X) Fq(Fq(WQTWK)Fq(XT)),X_{\mathrm{out}}=F_q(X)\,F_q\left(F_q\left(W_Q^{\mathsf T}W_K\right)F_q\left(X^{\mathsf T}\right)\right),

    where FqF_q is the tensor quantization function and XoutX_{\mathrm{out}} is the quantized query-key product. Under the straight-through gradient approximation, the gradient with respect to WKW_K is estimated as

    ∂L∂WK∣STE≈∂L∂XoutFq(X)WQTFq(XT),\frac{\partial L}{\partial W_K}\bigg|_{\mathrm{STE}} \approx \frac{\partial L}{\partial X_{\mathrm{out}}} F_q(X)W_Q^{\mathsf T}F_q(X^{\mathsf T}),

    where LL is the training loss. The quantized query factor Fq(WQ)F_q(W_Q) no longer appears in the key gradient, so oscillation in the query does not directly corrupt the key-gradient estimate. The reparameterization also reduces the number of quantization operations in the query-key product from six to four, reducing forward-pass discretization loss and improving the backward estimate.

  5. Knowl 5 — Oscillating weights are the source of suboptimal final models

    empirical result

    The paper tests whether weights near quantization thresholds are responsible for poor final solutions by injecting random perturbations into a converged quantized DeiT-T and retraining for one epoch. The boundary range BR0.005BR_{0.005} contains entries whose normalized latent weights are within 0.0050.005 of their nearest quantization threshold. Results below report ImageNet validation top-1 accuracy; μ\mu and σ\sigma are reported over ten trials, and the injected-weight fraction is held equal between random-position and within-boundary experiments.

    The key result is that perturbing only boundary weights does not harm 2-bit or 3-bit accuracy and can improve the best trial, whereas perturbing the same fraction of arbitrary weights causes a clear degradation. This indicates that the converged model is effectively a random snapshot from an oscillating neighborhood containing better nearby solutions.

    Could not parse LaTeX table

    Tracking experiments also show that oscillating positions are not fixed: weights initially inside BR0.005BR_{0.005} gradually leave it, while the total number of weights inside the boundary remains nearly unchanged because new weights enter. Thus, preventing new weights from entering the boundary is necessary for oscillation to subside.

  6. Knowl 6 — Quantized ViT granularity and training protocol

    experimental setup

    The experiments quantize fully connected-layer weights and activations row-wise, with scales aligned to the matrix-multiplication direction. For Xq∈RN×D1X_q\in\mathbb{R}^{N\times D_1} and Yq∈RD2×D1Y_q\in\mathbb{R}^{D_2\times D_1}, integer multiplication is reconstructed as

    XqYqT=αXqαYq⊙(X^q⊗Y^qT),X_qY_q^{\mathsf T}=\alpha_{X_q}\alpha_{Y_q}\odot\left(\widehat X_q\otimes\widehat Y_q^{\mathsf T}\right),

    where X^q\widehat X_q and Y^q\widehat Y_q are integer tensors, ⊗\otimes is integer matrix multiplication, αXq∈RN\alpha_{X_q}\in\mathbb{R}^{N} and αYq∈RD2\alpha_{Y_q}\in\mathbb{R}^{D_2} are row scale vectors, and ⊙\odot applies the corresponding high-precision scale products elementwise. This is quantization along the last dimension, meaning the scale is shared along that dimension. The value matrix V∈RN×DV\in\mathbb{R}^{N\times D} is the exception: when multiplying it by the attention matrix, VV is quantized along the sequence-length dimension NN, giving a scale vector in RD\mathbb{R}^{D}.

    The evaluation uses DeiT-T, DeiT-S, and Swin-T on ImageNet-1K. Quantized models are trained for 300 epochs with knowledge distillation from and initialization by the corresponding full-precision models. The 2-bit DeiT-T, 2-bit DeiT-S, and 2-bit Swin-T use the DeiT training setting without mixup or CutMix; 3-bit and 4-bit DeiT-S and Swin-T use the training recipe adopted from prior quantized-ViT experiments. The first patch-embedding layer and the final classification and distillation layers use 8-bit quantization. The complete proposed method, combining StatsQ, CGA, and QKR, is called oscillation-free quantization (OFQ).

  7. Knowl 7 — OFQ substantially improves ImageNet low-bit accuracy

    data/table

    The main ImageNet-1K comparison evaluates top-1 accuracy for full precision, LSQ, Mix-QViT, QViT*, and the proposed OFQ. QViT* denotes the authors' rerun after correcting the implementation so that the quantized matrix multiplication uses the required scale alignment. OFQ consistently gives the strongest result in every listed bit-width and architecture. At 2 bits, OFQ improves over LSQ by 9.88, 7.72, and 8.12 percentage points on DeiT-T, DeiT-S, and Swin-T, respectively; it also exceeds QViT* by 13.96, 7.05, and 4.64 points. At 3 bits, OFQ reaches or exceeds the full-precision accuracy for DeiT-T and is close to it for DeiT-S and Swin-T. At 4 bits, OFQ exceeds the corresponding full-precision result for all three networks.

    Could not parse LaTeX table
  8. Knowl 8 — StatsQ, QKR, and CGA provide complementary gains

    empirical result

    Ablations on DeiT-S show that all three techniques contribute, with StatsQ providing the largest single improvement at 2 bits and QKR and CGA adding further gains. The full combination is OFQ.

    Could not parse LaTeX table

    Relative to LSQ, StatsQ improves accuracy by 6.3, 0.8, and 1.0 percentage points at 2, 3, and 4 bits. Adding QKR to StatsQ gives additional gains of 0.7 and 0.59 points at 2 and 3 bits, while adding CGA gives 0.77 and 0.44 points. Combining all three reaches 75.72% at 2 bits, a 7.72-point gain over LSQ.

  9. Knowl 9 — CGA is robust to moderate boundary-range choices

    empirical result

    The boundary range controls which weights remain trainable during confidence-guided annealing. On 2-bit DeiT-S, CGA improves accuracy for all tested values, with BR0.005BR_{0.005} giving the best result. The paper interprets a range that is too large as unnecessarily expensive and potentially harmful to high-confidence weights, while a range that is too small can freeze low-confidence oscillating weights before they are optimized.

    Could not parse LaTeX table

    The statistical-scale trajectories shown for DeiT-S during CGA confirm that larger boundary ranges require more iterations before all weights leave the oscillation region. In the appendix experiments, the approximate time for the statistical scale to become stationary is 700 iterations for 2-bit DeiT-S, 1,500 for 3-bit DeiT-S, and 250 for 4-bit DeiT-S, supporting the paper's observation that oscillation is less severe at higher bit-width.

  10. Knowl 10 — Scope limitation of the demonstrated method

    limitation

    The paper validates oscillation-free quantization on ImageNet classification with DeiT-T, DeiT-S, and Swin-T vision transformers. It does not establish that the same oscillation mechanisms or the StatsQ, CGA, and QKR remedies generalize to other neural-network families, tasks, or application domains. The authors explicitly identify evaluation across a wider variety of deep networks and applications as future work.

Coverage note — The toy-regression curves and per-module gradient-direction plots are incorporated into the oscillation diagnosis rather than emitted as separate knowls; no substantial contributed method or result was otherwise omitted.

References

  1. 1.Banner, R., Nahshan, Y., and Soudry, D. Post training 4-bit quantization of convolutional networks for rapid-deployment. Advances in Neural Information Processing Systems, 32, 2019.
  2. 2.Bengio, Y., Léonard, N., and Courville, A. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  3. 3.Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K. Pact: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085, 2018.
  4. 4.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
  5. 5.Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S. Learned step size quantization. In International Conference on Learning Representations, 2020.
  6. 6.Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J. Differentiable soft quantization: Bridging full-precision and low-bit neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4852–4861, 2019.
  7. 7.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  8. 8.Helwegen, K., Widdicombe, J., Geiger, L., Liu, Z., Cheng, K.-T., and Nusselder, R. Latent weights do not exist: Rethinking binarized neural network optimization. Advances in neural information processing systems, 32, 2019.
  9. 9.Kenton, J. D. M.-W. C. and Toutanova, L. K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacl-HLT, pp. 4171–4186, 2019.
  10. 10.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Commun. ACM, 60(6):84–90, may 2017. ISSN 0001-0782. doi: 10.1145/3065386. URL https://doi.org/10.1145/3065386.
  11. 11.Li, Y., Dong, X., and Wang, W. Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. In International Conference on Learning Representations, 2020.
  12. 12.Li, Y., Xu, S., Zhang, B., Cao, X., Gao, P., and Guo, G. Q-vit: Accurate and fully quantized low-bit vision transformer. In Advances in Neural Information Processing Systems, 2022a.
  13. 13.Li, Z. and Gu, Q. I-vit: integer-only quantization for efficient vision transformer inference. arXiv preprint arXiv:2207.01405, 2022.
  14. 14.Li, Z., Yang, T., Wang, P., and Cheng, J. Q-vit: Fully differentiable quantization for vision transformer. arXiv preprint arXiv:2201.07703, 2022b.
  15. 15.Liu, Z., Wu, B., Luo, W., Yang, X., Liu, W., and Cheng, K.-T. Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. In Proceedings of the European conference on computer vision (ECCV), pp. 722–737, 2018.
  16. 16.Liu, Z., Mu, H., Zhang, X., Guo, Z., Yang, X., Cheng, K.-T., and Sun, J. Metapruning: Meta learning for automatic neural network channel pruning. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 3296–3305, 2019.
  17. 17.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 10012–10022, 2021a.
  18. 18.Liu, Z., Shen, Z., Li, S., Helwegen, K., Huang, D., and Cheng, K.-T. How do adam and training strategies help bnns optimization. In International Conference on Machine Learning, pp. 6936–6946. PMLR, 2021b.
  19. 19.Liu, Z., Wang, Y., Han, K., Zhang, W., Ma, S., and Gao, W. Post-training quantization for vision transformer. Advances in Neural Information Processing Systems, 34: 28092–28103, 2021c.
  20. 20.Liu, Z., Cheng, K.-T., Huang, D., Xing, E. P., and Shen, Z. Nonuniform-to-uniform quantization: Towards accurate quantization via generalized straight-through estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4942–4952, 2022.
  21. 21.Miyashita, D., Lee, E. H., and Murmann, B. Convolutional neural networks using logarithmic data representation. arXiv preprint arXiv:1603.01025, 2016.
  22. 22.Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M. Data-free quantization through weight equalization and bias correction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1325–1334, 2019.
  23. 23.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pp. 7197–7206. PMLR, 2020.
  24. 24.Nagel, M., Fournarakis, M., Bondarenko, Y., and Blankevoort, T. Overcoming oscillations in quantization-aware training. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 16318–16330. PMLR, 17–23 Jul 2022.
  25. 25.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pp. 10347–10357. PMLR, 2021.
  26. 26.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  27. 27.Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., and Keutzer, K. Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10734–10742, 2019.
  28. 28.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv, 2022.
  29. 29.Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6023–6032, 2019.
  30. 30.Zhang, D., Yang, J., Ye, D., and Hua, G. Lq-nets: Learned quantization for highly accurate and compact deep neural networks. In Proceedings of the European conference on computer vision (ECCV), pp. 365–382, 2018a.
  31. 31.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018b.
  32. 32.Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016.
  33. 33.Zhu, S., Duong, L. H., and Liu, W. Xor-net: an efficient computation pipeline for binary neural network inference on edge devices. In 2020 IEEE 26th International Conference on Parallel and Distributed Systems (ICPADS), pp. 124–131. IEEE, 2020.

Citation

MLA
Liu, S.-Y., et al. “Oscillation-free Quantization for Low-bit Vision Transformers”. International Conference on Machine Learning, vol. 202, 2023, pp. 21813–24, https://proceedings.mlr.press/v202/liu23w.html.
APA
Liu, S.-Y., Liu, Z., & Cheng, K.-T. (2023). Oscillation-free Quantization for Low-bit Vision Transformers. International Conference on Machine Learning, 202, 21813–21824. https://proceedings.mlr.press/v202/liu23w.html
Chicago
Liu, S.-Y., Z. Liu, and K.-T. Cheng. 2023. “Oscillation-free Quantization for Low-bit Vision Transformers”. International Conference on Machine Learning 202: 21813–24. https://proceedings.mlr.press/v202/liu23w.html.
Harvard
Liu, S.-Y., Liu, Z. and Cheng, K.-T. (2023) “Oscillation-free Quantization for Low-bit Vision Transformers”, International Conference on Machine Learning. PMLR, pp. 21813–21824. Available at: https://proceedings.mlr.press/v202/liu23w.html.
Vancouver
1. Liu S-Y, Liu Z, Cheng K-T (2023) Oscillation-free Quantization for Low-bit Vision Transformers. In: International Conference on Machine Learning. PMLR, pp 21813–21824

BibTeX

@InProceedings{pmlr-v202-liu23w,
  title = 	 {Oscillation-free Quantization for Low-bit Vision Transformers},
  author =       {Liu, Shih-Yang and Liu, Zechun and Cheng, Kwang-Ting},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {21813--21824},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/liu23w/liu23w.pdf},
  url = 	 {https://proceedings.mlr.press/v202/liu23w.html},
  abstract = 	 {Weight oscillation is a by-product of quantization-aware training, in which quantized weights frequently jump between two quantized levels, resulting in training instability and a sub-optimal final model. We discover that the learnable scaling factor, a widely-used $\textit{de facto}$ setting in quantization aggravates weight oscillation. In this work, we investigate the connection between learnable scaling factor and quantized weight oscillation using ViT, and we additionally find that the interdependence between quantized weights in $\textit{query}$ and $\textit{key}$ of a self-attention layer also makes ViT vulnerable to oscillation. We propose three techniques correspondingly: statistical weight quantization ($\rm StatsQ$) to improve quantization robustness compared to the prevalent learnable-scale-based method; confidence-guided annealing ($\rm CGA$) that freezes the weights with $\textit{high confidence}$ and calms the oscillating weights; and $\textit{query}$-$\textit{key}$ reparameterization ($\rm QKR$) to resolve the query-key intertwined oscillation and mitigate the resulting gradient misestimation. Extensive experiments demonstrate that our algorithms successfully abate weight oscillation and consistently achieve substantial accuracy improvement on ImageNet. Specifically, our 2-bit DeiT-T/DeiT-S surpass the previous state-of-the-art by 9.8% and 7.7%, respectively. The code is included in the supplementary material and will be released.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/