Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation

Zechun LiuKwang-Ting ChengDong HuangEric P. XingZhiqiang Shen

article2022CVPR156 citations

Presents a hardware-friendly quantization framework that pairs learnable nonuniform input thresholds with uniform output levels, using a generalized straight-through estimator to match the representational capacity of nonuniform methods without incurring inference deployment overhead.

Listen

Deploying deep neural networks to resource-constrained edge devices requires compressing model size and accelerating computation. While low-bit quantization reduces memory and computation demands by replacing heavy mathematical operations with efficient bitwise calculations, it often degrades model accuracy. Nonuniform quantization strategies retain accuracy better by adapting to the underlying data distribution, but they produce floating-point outputs that require lookup tables and extra post-processing, introducing substantial hardware area and energy overhead. Standard uniform quantization is hardware-friendly but suffers from rigid intervals that lead to severe information loss.

The article develops and evaluates Nonuniform-to-Uniform Quantization (N2UQ), a framework designed to achieve the high accuracy of nonuniform quantization while maintaining the hardware simplicity and operational efficiency of uniform quantization.

To bridge this gap, the approach enforces equidistant, uniform output values to allow direct hardware acceleration while learning flexible, non-equidistant input thresholds during training. Because calculating gradients with respect to learnable threshold parameters is mathematically intractable under standard straight-through estimation, the authors introduce a Generalized Straight-Through Estimator (G-STE) derived from the expected values of stochastic quantization. In addition, an entropy-preserving weight regularization technique is implemented to distribute weights evenly across quantization levels, preventing values from collapsing around zero and maximizing retained information. The method was evaluated on the standard ImageNet classification benchmark across multiple architectures (ResNet-18, ResNet-34, ResNet-50, and MobileNetV2) across 2-bit, 3-bit, and 4-bit configurations.

The evaluation produced several key findings. First, N2UQ consistently outperformed existing uniform and nonuniform quantization methods across all tested bit-widths, exceeding prior state-of-the-art nonuniform techniques by 0.5% to 1.7% in top-1 accuracy on ImageNet. Second, the 2-bit ResNet-50 model achieved 75.8% top-1 accuracy, substantially narrowing the gap to its full-precision counterpart to just 0.6% to 1.2%. Third, on compact models such as MobileNetV2, N2UQ attained 72.1% top-1 accuracy, matching or slightly exceeding full-precision baselines by mitigating overfitting through effective regularization. Finally, ablation studies showed that the G-STE threshold-learning activation quantizer and entropy-preserving weight regularization contributed 3.0% and 1.9% accuracy gains, respectively, over the 2-bit baseline.

These findings indicate that hardware efficiency does not require sacrificing representational flexibility. By generating uniform outputs directly, engineering teams can eliminate lookup tables and dedicated translation hardware, thereby reducing silicon area, memory footprint, latency, and power consumption on mobile and edge devices. Furthermore, the ability to train low-bit networks that match full-precision performance significantly reduces deployment risk for latency-critical applications.

Organizations developing or deploying low-power machine learning systems should consider adopting nonuniform-to-uniform quantization schemes and testing N2UQ on their vision workloads. For immediate implementation, teams can leverage the publicly available codebase to quantize existing convolutional backbones. Prior to enterprise-wide adoption, engineering teams should conduct pilot deployments on target hardware to benchmark actual latency, memory savings, and power efficiency against existing uniform integer pipelines.

While the results demonstrate high confidence on standard computer vision benchmarks and residual network architectures, the evaluations in the article are confined to image classification tasks. Stakeholders should exercise caution when extending the approach to non-vision architectures, such as large language models or transformers, where further empirical validation will be required.

Cover for Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation

Abstract

The nonuniform quantization strategy for compressing neural networks usually achieves better performance than its counterpart, i.e., uniform strategy, due to its superior representational capacity. However, many nonuniform quantization methods overlook the complicated projection process in implementing the nonuniformly quantized weights/activations, which incurs non-negligible time and space overhead in hardware deployment. In this study, we propose Nonuniform-to-Uniform Quantization (N2UQ), a method that can maintain the strong representation ability of nonuniform methods while being hardware-friendly and efficient as the uniform quantization for model inference. We achieve this through learning the flexible in-equidistant input thresholds to better fit the underlying distribution while quantizing these real-valued inputs into equidistant output levels. To train the quantized network with learnable input thresholds, we introduce a generalized straight-through estimator (G-STE) for intractable backward derivative calculation w.r.t. threshold parameters. Additionally, we consider entropy preserving regularization to further reduce information loss in weight quantization. Even under this adverse constraint of imposing uniformly quantized weights and activations, our N2UQ outperforms state-of-the-art nonuniform quantization methods by 0.5 ~ 1.7% on ImageNet, demonstrating the contribution of N2UQ design. Code and models are available at: https://github.com/liuzechun/Nonuniform-to-Uniform-Quantization.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Nonuniform-to-Uniform Quantization
  • 3.2.1 Forward Pass: Threshold Learning Quantization
  • 3.2.2 Backward Pass: Generalized Straight-Through Estimator (G-STE)
  • 3.2.3 Entropy Preserving Weight Regularization
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Comparison with State-of-the-Art Methods
  • 4.3. Ablation Study
  • 4.4. Visualization
  • 5. Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Nonuniform-to-Uniform Quantizer design

    model/method

    Nonuniform-to-Uniform Quantization (N2UQ) preserves the hardware advantages of uniform quantization while adding distribution-specific flexibility. It learns non-equidistant thresholds on the real-valued input axis, but maps the resulting values to equidistant output codes. Thus, quantized weights and activations retain a simple linear mapping to binary representations and do not require the floating-point lookup tables used to deploy nonuniform output levels.

  2. Knowl 2 — Forward threshold-learning quantizer

    equation

    For an nn-bit N2UQ activation quantizer, let xr∈Rx^r\in\mathbb{R} be a real-valued input, let xqx^q be its integer output code, and let T1<⋯<T2n−1T_1<\cdots<T_{2^n-1} be learnable thresholds. The deterministic forward mapping is

    xq={0,xr<T1,i,Ti≤xr<Ti+1,i∈{1,…,2n−2},2n−1,xr≥T2n−1.x^q=\begin{cases} 0, & x^r<T_1,\\ i, & T_i\le x^r<T_{i+1},\quad i\in\{1,\ldots,2^n-2\},\\ 2^n-1, & x^r\ge T_{2^n-1}. \end{cases}

    The output codes are equidistant, whereas the input intervals determined by the thresholds need not be. In the implementation, the code levels are scaled to {0,2/(2n−1),…,2}\{0,2/(2^n-1),\ldots,2\}; learnable factors β1\beta_1 and β2\beta_2 respectively scale the input before quantization and the output after quantization. The quantizer therefore adapts its input resolution without changing the uniformly spaced output representation.

  3. Knowl 3 — Generalized Straight-Through Estimator

    theoretical result

    For an nn-bit N2UQ with K=2nK=2^n output levels, let ai>0a_i>0 be the length of input segment ii for i=1,…,K−1i=1,\ldots,K-1, let d0=sd_0=s, and define di=s+∑j=1iajd_i=s+\sum_{j=1}^{i}a_j. The deterministic output uses the midpoint between adjacent segments as its hard boundary:

    xq={0,xr<d0+a1/2,i,di−1+ai/2≤xr<di+ai+1/2,i∈{1,…,K−2},K−1,xr≥dK−2+aK−1/2.x^q=\begin{cases} 0, & x^r<d_0+a_1/2,\\ i, & d_{i-1}+a_i/2\le x^r<d_i+a_{i+1}/2,\quad i\in\{1,\ldots,K-2\},\\ K-1, & x^r\ge d_{K-2}+a_{K-1}/2. \end{cases}

    The Generalized Straight-Through Estimator (G-STE) replaces the nearly everywhere-zero input derivative of this hard quantizer with the expected derivative of a stochastic segment quantizer:

    ∂xq∂xr~={1/ai,di−1≤xr<di,i∈{1,…,K−1},0,otherwise.\widetilde{\frac{\partial x^q}{\partial x^r}}=\begin{cases} 1/a_i, & d_{i-1}\le x^r<d_i,\quad i\in\{1,\ldots,K-1\},\\ 0, & \text{otherwise}. \end{cases}

    The corresponding surrogate gradient with respect to an interval length aia_i is −(xr−di−1)/ai2-(x^r-d_{i-1})/a_i^2 when xrx^r lies in segment ii, −1/aj-1/a_j when it lies in a later segment j>ij>i, and 00 otherwise. Consequently, threshold parameters receive gradients from the network loss. When all aia_i are equal, G-STE reduces to the ordinary straight-through estimator; with unequal aia_i, its slopes encode the learned input-threshold differences.

  4. Knowl 4 — Entropy-preserving weight normalization

    model/method

    N2UQ reduces weight information loss by encouraging the quantized values in each weight filter to occupy all available levels. If pip_i is the proportion of weights assigned to quantization level ii and NN is the number of levels, the quantized-weight entropy is H=−∑i=1Npilog⁡piH=-\sum_{i=1}^{N}p_i\log p_i subject to ∑ipi=1\sum_i p_i=1. Its maximum occurs at pi=1/Np_i=1/N for every level.

    For a real-valued weight filter WrW^r, let NwN_w be its number of entries and let ∥Wr∥1\|W^r\|_1 be its elementwise ℓ1\ell_1 norm. The method uses the per-filter normalization

    Wr′=2n−12n−1Nw∥Wr∥1Wr,W^{r\prime}=\frac{2^{n-1}}{2^n-1}\frac{N_w}{\|W^r\|_1}W^r,

    before clipping and applying the nn-bit quantizer. This normalization counteracts the concentration of real-valued weights near zero and empirically makes the resulting quantized weights more evenly distributed across the available levels.

  5. Knowl 5 — N2UQ parameterization and deployment overhead

    model/method

    An nn-bit N2UQ layer parameterizes its input partition using an initial point ss and 2n−12^n-1 positive interval lengths aia_i, together with two scaling parameters β1\beta_1 and β2\beta_2. The parameters are initialized as s=0s=0, ai=2/(2n−1)a_i=2/(2^n-1), and β1=β2=1\beta_1=\beta_2=1; every aia_i is constrained to exceed 10−310^{-3}. The layer adds only 2n+22^n+2 learnable parameters relative to a classical uniform quantizer. Equal initial intervals make N2UQ start as a uniform quantizer, while G-STE allows the intervals to become nonuniform during training.

  6. Knowl 6 — Uniform-output bitwise matrix multiplication

    equation

    The hardware motivation for N2UQ is that uniformly quantized values can be linearly encoded as binary digits. For an activation code aq=∑i=0M−1ai2ia^q=\sum_{i=0}^{M-1}a_i2^i and a weight code wq=∑j=0K−1wj2jw^q=\sum_{j=0}^{K-1}w_j2^j, where ai,wj∈{0,1}a_i,w_j\in\{0,1\}, the integer dot product can be evaluated as

    aq⋅wq=∑i=0M−1∑j=0K−12i+j popcnt⁡ ⁣(and⁡(ai,wj)).a^q\cdot w^q=\sum_{i=0}^{M-1}\sum_{j=0}^{K-1}2^{i+j}\,\operatorname{popcnt}\!\left(\operatorname{and}(a_i,w_j)\right).

    Here MM and KK are the activation and weight bit-widths, respectively, and⁡\operatorname{and} is a bitwise AND operation, and popcnt⁡\operatorname{popcnt} counts the set bits. N2UQ retains uniformly spaced output codes, so this encoding can be performed by linear mappings without the lookup-table post-processing required by floating-point nonuniform output levels.

  7. Knowl 7 — ImageNet training configuration

    experimental setup

    N2UQ was evaluated on ImageNet-2012, using 1.2 million training images and 50,000 validation images across 1,000 classes. Real-valued PyTorch pretrained networks initialized the quantized models. Training used Adam with an initial learning rate of 2.5×10−32.5\times10^{-3} for weight parameters, linear learning-rate decay, batch size 512, zero weight decay, 128 epochs, and the knowledge-distillation procedure used by LSQ. Training images were randomly resized, cropped to 224×224224\times224, and horizontally flipped; validation images were center-cropped to 224×224224\times224.

    The N2UQ-specific parameters used one tenth of the weight learning rate. Networks used pre-activation NonLinear-Conv-BN blocks with RPReLU, and all convolutional and fully connected layers were quantized except the first and last layers. The experiments covered ResNet-18, ResNet-34, ResNet-50, and MobileNet-V2 under multiple weight/activation bit-widths.

  8. Knowl 8 — Accuracy against uniform and nonuniform quantization

    data/table

    The ImageNet comparisons show that N2UQ improves over the prior nonuniform method LCQ while retaining uniformly quantized weights and activations. The table reports top-1 and top-5 accuracy for the directly comparable entries; FP is the full-precision top-1 accuracy.

    Network W/A FP top-1 LCQ top-1 LCQ top-5 N2UQ top-1 N2UQ top-5
    ResNet-18 2/2 71.8 68.9 – 69.4 88.4
    ResNet-18 3/3 71.8 70.6 – 71.9 90.5
    ResNet-18 4/4 71.8 71.5 – 72.9 90.9
    ResNet-34 2/2 74.9 72.7 – 73.3 91.2
    ResNet-34 3/3 74.9 74.0 – 75.2 92.3
    ResNet-34 4/4 74.9 74.3 – 76.0 92.8
    ResNet-50 2/2 77.0 75.1 – 75.8 92.3
    ResNet-50 3/3 77.0 76.3 – 77.5 93.6
    ResNet-50 4/4 77.0 76.6 – 78.0 93.9
    MobileNet-V2 – 72.0 70.8 89.7 72.1 90.6

    For the ResNet comparisons, N2UQ exceeds LCQ by 0.50.5--1.71.7 top-1 percentage points across the listed bit-widths. The 4-bit MobileNet-V2 result reaches 72.1% top-1 and 90.6% top-5 accuracy, slightly exceeding the reported full-precision top-1 accuracy of 72.0%. The paper's prose separately reports a 76.4% 2-bit ResNet-50 result and a 0.6-point gap to full precision, whereas the reproduced comparison table lists 75.8% for that configuration.

  9. Knowl 9 — Component ablation on 2-bit ResNet-18

    data/table

    On ImageNet with a 2-bit ResNet-18, the threshold-learning activation quantizer with G-STE and the entropy-preserving weight normalization provide complementary gains over the implemented baseline. The threshold-learning component contributes 3.0 top-1 points by itself, the entropy-preserving weight treatment contributes 1.9 points by itself, and their combination reaches 69.7%.

    Method Top-1 accuracy (%)
    Baseline 65.9
    Baseline + threshold-learning activation quantizer with G-STE 68.9
    Baseline + entropy-preserving weight regularization 67.8
    N2UQ with both components 69.7
    Corresponding real-valued network 71.8

    When the threshold-learning quantizer with G-STE is fixed and only the weight treatment is varied, the reported top-1 accuracies are 68.9% with no regularization, 67.4% with weight normalization, 68.6% with a learnable scaling factor, and 69.7% with entropy-preserving weight regularization. The authors attribute the poorer learnable-factor result to noisy gradients and the poorer weight-normalization result to its lack of awareness of the quantization distribution.

  10. Knowl 10 — Learned thresholds and weight-level occupancy

    empirical result

    The visualization analysis shows the two intended distribution-adaptation behaviors of N2UQ. In a 2-bit baseline weight filter, only 6 entries were assigned to the level 0.50.5 and 20 entries to −0.5-0.5, indicating concentration caused by extrema-based rescaling. Entropy-preserving normalization produces a more even occupation of the quantized weight levels.

    For activations whose real-valued distribution is dense near zero and sparse in the tails, the learned N2UQ input partition uses smaller intervals in the dense region and larger intervals in the sparse regions. This differs from classical uniform quantization, whose fixed input thresholds cannot adapt to such density changes, and is intended to reduce the representation error while leaving the output levels uniformly spaced.

Coverage note — No substantial contributed material was deliberately omitted; the reported ImageNet comparisons, ablations, visualization analysis, N2UQ construction, G-STE, entropy-preserving weight treatment, and deployment rationale are covered.

References

  1. 1.S Arish and RK Sharma. An efficient floating point multiplier design for high speed applications using karatsuba algorithm and urdhva-tiryagbhyam algorithm. In 2015 International Conference on Signal Processing and Communication (ICSC), pages 303–308. IEEE, 2015. 1, 2, 3
  2. 2.Yoshua Bengio, Nicholas Leonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013. 2, 3, 4
  3. 3.Joseph Bethge, Christian Bartz, Haojin Yang, Ying Chen, and Christoph Meinel. Meliusnet: Can binary neural networks achieve mobilenet-level accuracy? arXiv preprint arXiv:2001.05936, 2020. 4
  4. 4.Adrian Bulat and Georgios Tzimiropoulos. Xnor-net++: Improved binary neural networks. British Machine Vision Conference, 2019. 4
  5. 5.Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. Once-for-all: Train one network and specialize it for efficient deployment. In International Conference on Learning Representations, 2019. 2
  6. 6.Han Cai, Ligeng Zhu, and Song Han. Proxylessnas: Direct neural architecture search on target task and hardware. In International Conference on Learning Representations, 2018. 1, 2
  7. 7.Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan, Zhuo Wang, and Pierce Chuang. Accurate and efficient 2-bit quantized neural networks. In MLSys, 2019. 3
  8. 8.Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. Pact: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085, 2018. 3, 4, 6, 7
  9. 9.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in neural information processing systems, pages 3123–3131, 2015. 4
  10. 10.Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Binarized neural networks: Training deep neural networks with weights and activations constrained to+ 1 or-1. arXiv preprint arXiv:1602.02830, 2016. 4
  11. 11.Xiaohan Ding, Xiangxin Zhou, Yuchen Guo, Jungong Han, Ji Liu, et al. Global sparse momentum sgd for pruning very deep neural networks. In Advances in Neural Information Processing Systems, pages 6379–6391, 2019. 2
  12. 12.Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha. Learned step size quantization. In International Conference on Learning Representations, 2020. 3, 6, 7
  13. 13.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021. 1, 2, 3, 6
  14. 14.Ruihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li, Peng Hu, Jiazhen Lin, Fengwei Yu, and Junjie Yan. Differentiable soft quantization: Bridging full-precision and low-bit neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4852–4861, 2019. 3, 6, 7
  15. 15.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015. 2, 3, 7
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1, 6
  17. 17.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 2
  18. 18.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
  19. 19.Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. Quantized neural networks: Training neural networks with low precision weights and activations. The Journal of Machine Learning Research, 18(1):6869–6898, 2017. 3, 4
  20. 20.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704–2713, 2018. 3, 6
  21. 21.Shubham Jain, Swagath Venkataramani, Vijayalakshmi Srinivasan, Jungwook Choi, Pierce Chuang, and Leland Chang. Compensated-dnn: Energy efficient low-precision deep neural networks by compensating quantization errors. In 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC), pages 1–6. IEEE, 2018. 3
  22. 22.Yongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Jeongin Yun, and Dongsoo Lee. Biqgemm: matrix multiplication with lookup table for binary-coding-based quantized dnns. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–14. IEEE, 2020. 2, 3, 6
  23. 23.Sangil Jung, Changyong Son, Seohyung Lee, Jinwoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, and Changkyu Choi. Learning to quantize deep networks by optimizing quantization intervals with task loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4350–4359, 2019. 3, 6, 7
  24. 24.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  25. 25.Junghyup Lee, Bumsub Ham, et al. Distance-aware quantization. In Proceedings of the IEEE International Conference on Computer Vision, 2021. 7
  26. 26.Yuhang Li, Xin Dong, and Wei Wang. Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. In International Conference on Learning Representations, 2020. 3, 7
  27. 27.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE International Conference on Computer Vision, pages 2736–2744, 2017. 1, 2
  28. 28.Zechun Liu, Wenhan Luo, Baoyuan Wu, Xin Yang, Wei Liu, and Kwang-Ting Cheng. Bi-real net: Binarizing deep network towards real-network performance. International Journal of Computer Vision, pages 1–18, 2018. 4
  29. 29.Zechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo, Xin Yang, Kwang-Ting Cheng, and Jian Sun. Metapruning: Meta learning for automatic neural network channel pruning. In Proceedings of the IEEE International Conference on Computer Vision, pages 3296–3305, 2019. 1, 2
  30. 30.Zechun Liu, Zhiqiang Shen, Marios Savvides, and Kwang-Ting Cheng. Reactnet: Towards precise binary neural network with generalized activation functions. In European Conference on Computer Vision, pages 143–159. Springer, 2020. 4, 6
  31. 31.Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, and Kwang-Ting Cheng. Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. In Proceedings of the European conference on computer vision (ECCV), pages 722–737, 2018. 1, 2, 4, 6
  32. 32.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European Conference on Computer Vision (ECCV), pages 116–131, 2018. 2
  33. 33.Jeffrey L McKinstry, Steven K Esser, Rathinakumar Appuswamy, Deepika Bablani, John V Arthur, Izzet B Yildiz, and Dharmendra S Modha. Discovering low-precision networks close to full-precision networks for efficient inference. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS), pages 6–9. IEEE, 2019. 7
  34. 34.Daisuke Miyashita, Edward H Lee, and Boris Murmann. Convolutional neural networks using logarithmic data representation. arXiv preprint arXiv:1603.01025, 2016. 3
  35. 35.Eunhyeok Park and Sungjoo Yoo. Profit: A novel training method for sub-4-bit mobilenet models. In European Conference on Computer Vision, pages 430–446. Springer, 2020. 7
  36. 36.Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In European conference on computer vision, pages 525–542. Springer, 2016. 1, 4
  37. 37.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. 6
  38. 38.Tim Salimans and Durk P Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. Advances in neural information processing systems, 29:901–909, 2016. 7, 8
  39. 39.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmgoinov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4510–4520, 2018. 2
  40. 40.Zhiqiang Shen, Zhankui He, and Xiangyang Xue. Meal: Multi-model ensemble via adversarial learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 4886–4893, 2019. 2
  41. 41.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 1
  42. 42.Naoya Torii, Hirotaka Kokubo, Dai Yamamoto, Kouichi Itoh, Masahiko Takenaka, and Tsutomu Matsumoto. Asic implementation of random number generators using sr latches and its evaluation. EURASIP Journal on Information Security, 2016(1):1–12, 2016. 5
  43. 43.Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer. Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10734–10742, 2019. 1
  44. 44.Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer. Mixed precision quantization of convnets via differentiable neural architecture search. arXiv preprint arXiv:1812.00090, 2018. 4, 7
  45. 45.Kaiqiang Xu, Xinchen Wan, Hao Wang, Zhenghang Ren, Xudong Liao, Decang Sun, Chaoliang Zeng, and Kai Chen. Tacc: A full-stack cloud computing infrastructure for machine learning tasks, 2021. 8
  46. 46.Kohei Yamamoto. Learnable companding quantization for accurate low-bit neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5029–5038, 2021. 2, 3, 4, 6, 7
  47. 47.Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua. Lq-nets: Learned quantization for highly accurate and compact deep neural networks. In Proceedings of the European conference on computer vision (ECCV), pages 365–382, 2018. 1, 2, 3, 7
  48. 48.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6848–6856, 2018. 2
  49. 49.Xiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu, and Lei Zhang. Linear symmetric quantization of neural networks for low-precision integer hardware. International Conference on Learning Representations, 2020. 7
  50. 50.Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016. 1, 2, 3, 4, 6, 7, 8
  51. 51.Shilin Zhu, Xin Dong, and Hao Su. Binary ensemble neural network: More bits per network or more networks per bit? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4923–4932, 2019. 4
  52. 52.Bohan Zhuang, Lingqiao Liu, Mingkui Tan, Chunhua Shen, and Ian Reid. Training quantized neural networks with a full-precision auxiliary module. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1488–1497, 2020. 7
  53. 53.Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, and Ian Reid. Towards effective low-bitwidth convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7920–7928, 2018. 2, 3

Citation

MLA
Liu, Z., et al. “Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation”. arXiv, 2021, http://arxiv.org/abs/2111.14826v2.
APA
Liu, Z., Cheng, K.-T., Huang, D., Xing, E., & Shen, Z. (2021). Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation. arXiv. http://arxiv.org/abs/2111.14826v2
Chicago
Liu, Z., K.-T. Cheng, D. Huang, E. Xing, and Z. Shen. 2021. “Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation”. arXiv. http://arxiv.org/abs/2111.14826v2.
Harvard
Liu, Z. et al. (2021) “Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.14826v2.
Vancouver
1. Liu Z, Cheng K-T, Huang D, Xing E, Shen Z (2021) Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation. arXiv

BibTeX

@article{liu2021nonuniform,
  title = {Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation},
  author = {Liu, Zechun and Cheng, Kwang-Ting and Huang, Dong and Xing, Eric and Shen, Zhiqiang},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.14826v2},
  eprint = {2111.14826}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE