Learnable Lookup Table for Neural Network Quantization

Longguang WangXiaoyu DongYingqian WangLi LiuWei AnYulan Guo

article2022CVPR53 citations

Proposes differentiable, learnable lookup tables that adaptively fit non-uniform weight and activation distributions during training while enabling fast, low-overhead table lookups during inference across vision tasks.

Listen

Deploying modern deep neural networks to resource-constrained edge and mobile devices remains a major challenge due to high memory footprint and computation costs. Network quantization reduces these resource demands by converting full-precision values into compact, low-bit representations. However, standard linear quantizers fail to accommodate the non-uniform, bell-shaped distributions typical of neural network weights and activations across different layers. While prior parameterized quantizers attempt to adapt to these distributions, they rely on complex non-linear mathematical functions that incur significant computational overhead during real-time activation quantization.

The article demonstrates an efficient uniform quantization framework that formulates the quantization process as a differentiable, learnable lookup table. The primary objective is to learn adaptive layer-specific quantization mappings during training while maintaining minimal computational and memory overhead during hardware deployment.

To achieve this, the authors designed a continuous relaxation scheme that allows discrete lookup tables to be optimized end-to-end alongside standard network parameters. The training strategy incorporates an exponential formulation of scaling parameters to prevent numerical instability, alongside a gradient rescaling mechanism to counterbalance the severe gradient skew between central and extreme value distributions. The method was evaluated across diverse benchmarks covering 2D image classification (CIFAR-10 and ImageNet), image super-resolution (DIV2K, Set5, Set14, B100, Urban100), and 3D point cloud classification (ModelNet40) across several bit-width configurations ranging from 2-bit to 8-bit precision.

The findings establish that the learnable lookup table consistently achieves state-of-the-art accuracy across multiple visual and spatial tasks. On ImageNet classification, a 4-bit ResNet-18 model achieved 70.4% Top-1 accuracy, matching or exceeding full-precision baseline performance while outperforming competing quantization techniques. In low-bit edge settings, such as 2-bit point cloud classification, the proposed method surpassed the nearest competing approach by 2.9 percentage points. Crucially, inference profiling on mobile and CPU hardware revealed that the lookup operation requires only about 1.25 kilobytes of extra memory while reducing activation quantization runtime by more than 50% to 75% compared to complex functional quantizers, cutting mobile latency from 95–120 milliseconds down to 40 milliseconds.

These results demonstrate that complex mathematical approximations can be replaced with direct table lookups without sacrificing model expressiveness. For organizations seeking to deploy real-time artificial intelligence on edge devices, this approach simultaneously lowers latency, memory consumption, and energy use while preserving output quality. Technical teams can adopt this framework directly on off-the-shelf integer-arithmetic hardware without requiring custom non-uniform hardware accelerators. Future deployment initiatives should run hardware-specific pilot validations on target embedded systems, noting that the study primarily evaluated visual and 3D classification architectures rather than larger foundation models.

Cover for Learnable Lookup Table for Neural Network Quantization

Abstract

Neural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(·) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use pre-defined functions (e.g., exponential function) with learnable parameters to build the quantizer for joint optimization. However, these complicated quantizers introduce considerable computational overhead during inference since activation quantization should be conducted online. In this paper, we formulate the quantization process as a simple lookup operation and propose to learn lookup tables as quantizers. Specifically, we develop differentiable lookup tables and introduce several training strategies for optimization. Our lookup tables can be trained simply in the network in an end-to-end manner to fit the distributions in different layers and have very small additional computational cost. Comparison with previous methods show that quantized networks using our lookup tables achieve state-of-the-art performance on image classification, image super-resolution, and point cloud classification tasks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Network Quantization
  • 2.2. Applications
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Lookup Table Construction
  • 3.3. Differentiable Lookup Operation
  • 3.4. Training Strategies
  • 3.5. Discussion
  • 4. Experiments
  • 4.1. Experiments on Image Classification
  • 4.1.1 Evaluation on CIFAR-10
  • 4.1.2 Evaluation on ImageNet
  • 4.2. Experiments on Image Super-Resolution
  • 4.3. Experiments on Point Cloud Classification
  • 4.4. Model Analyses
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Differentiable Learnable Lookup Table for Quantization

    model/method

    In neural network uniform quantization, full-precision weights ww and activations aa are clipped and mapped to discrete levels:

    w^=clip(w/sw,−1,1),a^=clip(a/sa,0,1)\hat{w} = \text{clip}(w / s_w, -1, 1), \quad \hat{a} = \text{clip}(a / s_a, 0, 1)

    wˉ=sw⋅Q(w^,Qw),aˉ=sa⋅Q(a^,Qa)\bar{w} = s_w \cdot Q(\hat{w}, Q_w), \quad \bar{a} = s_a \cdot Q(\hat{a}, Q_a)

    where sw,sas_w, s_a are trainable scale parameters, and Qa=2b−1Q_a = 2^b - 1, Qw=2b−1−1Q_w = 2^{b-1} - 1 are the numbers of discrete quantization levels for a bb-bit network. The mapping function Q(⋅)Q(\cdot) is parameterized as a learnable lookup table (LLT).

    To make the discrete lookup table differentiable, the quantizer is decomposed into a sum of QaQ_a scaled step functions ε(⋅)\varepsilon(\cdot):

    Q(a^,Qa)=1Qa∑m=1Qaε(a^−θm)Q(\hat{a}, Q_a) = \frac{1}{Q_a} \sum_{m=1}^{Q_a} \varepsilon(\hat{a} - \theta_m)

    Each step function is formulated as the cumulative integral of an impulse distribution over KK intervals. During training, the impulse distribution is relaxed into a temperatured softmax distribution over learnable parameters {g1,g2,…,gK}\{g_1, g_2, \dots, g_K\}:

    pi=exp⁡(gi/τ)∑j=1Kexp⁡(gj/τ)p_i = \frac{\exp(g_i / \tau)}{\sum_{j=1}^K \exp(g_j / \tau)}

    where τ>0\tau > 0 is a temperature parameter and gig_i is initialized to 0. Accumulating {pi}\{p_i\} creates a soft step function, and summing across all sub-tables forms the differentiable lookup table. As τ→0\tau \to 0, each softened distribution converges to a one-hot vector, recovering a hardened step function for lookup-based inference.

  2. Knowl 2 — Exponential Formulation of Quantization Scale Parameters

    model/method

    Clipping scale parameters sws_w (for weights) and sas_a (for activations) must remain strictly positive during training. If unconstrained, standard gradient descent updates can push scale parameters below zero (sign reversal), causing optimization instability and degraded low-bit quantization performance.

    To enforce positivity and stabilize convergence, scale parameters are reparameterized exponentially using auxiliary unconstrained learnable parameters ewe_w and eae_a:

    sw=exp⁡(ew),sa=exp⁡(ea)s_w = \exp(e_w), \quad s_a = \exp(e_a)

    The auxiliary parameters are initialized using the empirical data distribution:

    ew=ln⁡(3σw),ea=ln⁡(3σa)e_w = \ln(3\sigma_w), \quad e_a = \ln(3\sigma_a)

    where σw\sigma_w is the standard deviation of full-precision weights at initialization, and σa\sigma_a is the standard deviation of full-precision activations calculated during the first training iteration.

  3. Knowl 3 — Cell-Wise Gradient Rescaling for Lookup Table Quantization

    model/method

    Because weight and activation values follow bell-shaped distributions, the number of values mapped into different cells of a quantization lookup table is highly uneven during training. Central cells near zero process significantly more values than outer cells near the clipping boundaries. Consequently, accumulated gradients for cells near zero overpower the optimization, causing the lookup table to overfit central regions and clip large values aggressively.

    To balance optimization across all lookup table regions, the gradient gig_i for the ii-th cell is rescaled by the ratio of average cell occupancy to individual cell occupancy:

    g~i=gi⋅NavgNi\tilde{g}_i = g_i \cdot \sqrt{\frac{N_{\text{avg}}}{N_i}}

    where NiN_i is the number of float values falling into the ii-th cell in the current iteration, and NavgN_{\text{avg}} is the average number of float values per cell across the entire lookup table.

  4. Knowl 4 — Exponential Temperature Annealing Schedule for Lookup Table Hardening

    equation

    The temperature parameter τ\tau governs the transition of the learnable lookup table from a soft continuous relaxation to a hardened discrete table. The parameter τ\tau is decayed using an exponential annealing schedule at iteration nn:

    τ(n)={τ0×(τ1τ0)nMNiter,if n<MNiterτ1,otherwise\tau(n) = \begin{cases} \tau_0 \times \left( \frac{\tau_1}{\tau_0} \right)^{\frac{n}{M N_{\text{iter}}}}, & \text{if } n < M N_{\text{iter}} \\ \tau_1, & \text{otherwise} \end{cases}

    where τ0\tau_0 is the initial high temperature, τ1\tau_1 is the final target temperature, nn is the iteration index, NiterN_{\text{iter}} is the total number of iterations per training epoch, and MM is a scalar controlling the annealing rate.

  5. Knowl 5 — Straight-Through Estimator Gradients for Learnable Lookup Quantization

    model/method

    During forward propagation, quantized activations aˉ=sa⋅Q(a^,Qa)\bar{a} = s_a \cdot Q(\hat{a}, Q_a) are computed via lookup table indexing and used as proxies for continuous activations aa, with a^=clip(a/sa,0,1)\hat{a} = \text{clip}(a / s_a, 0, 1).

    To update both the lookup table parameters and full-precision activations, gradients are computed via the Straight-Through Estimator (STE):

    ∂aˉ∂pi=sa\frac{\partial \bar{a}}{\partial p_i} = s_a

    ∂aˉ∂a={1,if a<sa0,otherwise\frac{\partial \bar{a}}{\partial a} = \begin{cases} 1, & \text{if } a < s_a \\ 0, & \text{otherwise} \end{cases}

    where pip_i is the softened softmax probability corresponding to the cell containing a^\hat{a}, and sas_a is the activation clipping scale parameter.

  6. Knowl 6 — Image Classification Accuracy of Learnable Lookup Table Quantization

    data/table

    Top-1 accuracy (%) on CIFAR-10 (ResNet-20 and VGG-Small) and Top-1/Top-5 accuracy (%) on ImageNet (ResNet-18) for Learnable Lookup Table (LLT) quantization compared to other uniform quantization methods across different weight/activation (W/AW/A) bit-width settings:

    Model / Dataset Method 32/32 4/4-bit 3/3-bit 2/2-bit 2/4-bit
    ResNet-20 (CIFAR-10) DoReFa-Net 92.96% 90.50% 89.90% 88.20% -
    PACT 92.96% 91.30% 91.10% 89.70% -
    QIL 92.96% 91.52% 91.81% 90.45% 91.33%
    LSQ 92.96% 92.30% 91.69% 90.08% 91.72%
    SLB 92.96% 91.60% - 90.60% 91.30%
    LLT (Ours) 92.96% 92.71% 92.17% 90.63% 91.80%
    VGG-Small (CIFAR-10) DoReFa-Net 94.10% 88.20% 89.90% 90.50% -
    QIL 94.10% 93.77% 93.71% 93.45% 93.73%
    SLB 94.10% 93.80% - 93.50% 93.90%
    CPQ 94.10% 93.23% 93.18% 92.51% -
    LLT (Ours) 94.10% 94.20% 94.03% 93.83% 94.02%
    ResNet-18 (ImageNet) Method 32-bit (Top-1/5) 4-bit (Top-1/5) 3-bit (Top-1/5) 2-bit (Top-1/5)
    ResNet-18 PACT 69.8% / 89.1% 69.2% / 89.0% 68.1% / 88.2% 64.4% / 85.6%
    DSQ 69.8% / 89.1% 69.6% / - 68.7% / - 65.2% / -
    QIL 69.8% / 89.1% 70.1% / - 69.2% / - 65.7% / -
    CPQ 69.8% / 89.1% 69.6% / 89.0% 67.2% / 87.4% -
    LLT (Ours) 69.8% / 89.1% 70.4% / 89.6% 69.5% / 88.9% 66.0% / 86.2%

    LLT consistently outperforms both fixed-quantizer methods (DoReFa-Net, PACT, LSQ) and learnable quantizers (QIL, DSQ, CPQ) across all bit-widths.

  7. Knowl 7 — Image Super-Resolution Performance under LLT Quantization

    data/table

    PSNR (dB) performance for ×4\times 4 super-resolution on standard benchmark datasets (Set5, Set14, B100, Urban100) using EDSR and RDN backbones comparing Learnable Lookup Table (LLT) quantization with generic uniform quantizers (DoReFa-Net, PACT, LSQ) and super-resolution-specific quantizers (PAMS, DAQ):

    Model Dataset Baseline (32-bit) DoReFa-Net (4-bit) PACT (4-bit) LSQ (4-bit) PAMS (4-bit) DAQ (4-bit) LLT (4-bit)
    EDSR Set5 32.46 29.57 31.39 32.27 31.59 32.34 32.40
    Set14 28.77 26.82 28.10 28.60 28.20 28.69 28.74
    B100 27.69 26.47 27.25 27.63 27.32 27.61 27.70
    Urban100 26.54 23.75 25.15 26.34 25.32 26.33 26.51
    RDN Set5 32.32 - - 32.20 30.44 31.96 32.26
    Set14 28.71 - - 28.63 27.54 28.38 28.70
    B100 27.67 - - 27.60 26.87 27.38 27.66
    Urban100 26.35 - - 26.20 24.52 25.73 26.29

    At 4-bit precision, LLT achieves PSNR values within 0.03 to 0.06 dB of the full-precision models, outperforming DAQ by up to 0.56 dB on Urban100.

  8. Knowl 8 — Point Cloud Classification Accuracy on ModelNet40

    data/table

    Overall accuracy (%) on the ModelNet40 point cloud classification benchmark for PointNet and PointNet++ across 4-bit, 3-bit, and 2-bit uniform quantization:

    Model Method 32-bit 4-bit 3-bit 2-bit
    PointNet DoReFa-Net 90.8% 89.4% 88.3% 80.3%
    PACT 90.8% 89.4% 87.9% 80.8%
    QIL 90.8% 89.7% 88.6% 82.8%
    LSQ 90.8% 90.0% 88.6% 84.7%
    LLT (Ours) 90.8% 90.7% 89.9% 87.6%
    PointNet++ DoReFa-Net 92.8% 92.2% 92.2% 89.2%
    PACT 92.8% 92.3% 92.2% 88.9%
    QIL 92.8% 92.6% 92.1% 88.7%
    LSQ 92.8% 92.6% 92.1% 90.1%
    LLT (Ours) 92.8% 92.8% 92.4% 92.3%

    At 4-bit precision, LLT matches full-precision performance (90.7% vs. 90.8% for PointNet, 92.8% vs. 92.8% for PointNet++), and surpasses LSQ by 2.9% and 2.2% at 2-bit precision on PointNet and PointNet++, respectively.

  9. Knowl 9 — Ablation on Exponential Scale Formulation and Gradient Rescaling

    empirical result

    An ablation study on CIFAR-10 with ResNet-20 quantifies the individual and joint contributions of the exponential scale parameter formulation and the cell gradient rescaling scheme:

    Configuration Exponential Scale Gradient Rescaling 4/4-bit 3/3-bit 2/2-bit
    Baseline (Model 0) ✗ ✗ 92.44% 92.02% 89.81%
    Model 1 ✓ ✗ 92.42% 92.08% 90.10%
    Model 2 ✗ ✓ 92.65% 92.08% 90.14%
    Model 3 (LLT full) ✓ ✓ 92.71% 92.17% 90.63%

    Without the exponential formulation (Model 0), scale parameters suffer violent fluctuations and sign reversals (e.g., at epochs 50 and 70), impairing convergence. Without gradient rescaling (Model 1), gradients near zero dominate training, leading to 66% of weights being clipped under 2/2-bit quantization, compared to only 9% clipped when gradient rescaling is used (Model 3).

  10. Knowl 10 — Inference Latency, Memory Overhead, and LUT Granularity Analysis

    empirical result

    The granularity parameter KK specifies the number of sub-intervals per step function in the lookup table. Evaluating ResNet-20 on CIFAR-10 shows:

    • K=1K = 1 (degrades LLT to standard linear round function): achieves 92.37% (4/4-bit), 91.95% (3/3-bit), and 90.42% (2/2-bit) accuracy.
    • K=5K = 5: achieves 92.60% (4/4-bit), 92.14% (3/3-bit), and 90.58% (2/2-bit).
    • K=9K = 9: achieves 92.71% (4/4-bit), 92.17% (3/3-bit), and 90.63% (2/2-bit).
    • K=13K = 13: achieves 92.72% (4/4-bit), 92.15% (3/3-bit), and 90.62% (2/2-bit), showing saturation beyond K=9K = 9.

    Inference efficiency comparison for 4-bit ResNet-18 on ImageNet:

    • Memory: LLT adds 1.25 KB of parameter memory over baseline (7554.9 KB vs. 7553.7 KB).
    • Quantization Latency: On an Intel i9-9900K CPU, online activation quantization takes 20 ms with LLT versus 90 ms for QIL and 85 ms for QNet. On a Kirin 810 mobile processor, LLT takes 40 ms versus 120 ms for QIL and 95 ms for QNet, demonstrating the speed advantage of table lookup over continuous function evaluation.

Coverage note — None was omitted; all primary methodological contributions, theoretical formulations, and experimental evaluations across image classification, super-resolution, point cloud classification, and ablation studies are covered.

References

  1. 1.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In ECCV, pages 116–131, 2018.
  2. 2.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, pages 4510–4520, 2018.
  3. 3.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, pages 6105–6114, 2019.
  4. 4.Yang He, Guoliang Kang, Xuanyi Dong, Yanwei Fu, and Yi Yang. Soft filter pruning for accelerating deep convolutional neural networks. In IJCAI, pages 2234–2240, 2018.
  5. 5.Yang He, Ping Liu, Ziwei Wang, Zhilan Hu, and Yi Yang. Filter pruning via geometric median for deep convolutional neural networks acceleration. In CVPR, pages 4340–4349, 2019.
  6. 6.Mingbao Lin, Rongrong Ji, Yan Wang, Yichen Zhang, Baochang Zhang, Yonghong Tian, and Ling Shao. Hrank: Filter pruning using high-rank feature map. In CVPR, pages 1529–1538, 2020.
  7. 7.Alexander Novikov, Dmitry Podoprikhin, Anton Osokin, and Dmitry P Vetrov. Tensorizing neural networks. In NeurIPS, 2015.
  8. 8.Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. On compressing deep models by low rank and sparse decomposition. In CVPR, pages 7370–7379, 2017.
  9. 9.Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In CVPR, pages 7130–7138, 2017.
  10. 10.Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy. Knowledge distillation meets self-supervision. In ECCV, 2020.
  11. 11.Yifan Liu, Changyong Shu, Jingdong Wang, and Chunhua Shen. Structured knowledge distillation for dense prediction. IEEE Trans. Pattern Anal. Mach. Intell., 2020.
  12. 12.Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang, Yunxin Liu, Lintao Zhang, Lanshun Nie, and Zhi Yang. Seernet: Predicting convolutional neural network feature-map sparsity through low-bit quantization. In CVPR, pages 11216–11225, 2019.
  13. 13.Peisong Wang, Qinghao Hu, Yifan Zhang, Chunjie Zhang, Yang Liu, and Jian Cheng. Two-step quantization for low-bit neural networks. In CVPR, pages 4376–4384, 2018.
  14. 14.Sangil Jung, Changyong Son, Seohyung Lee, Jinwoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, and Changkyu Choi. Learning to quantize deep networks by optimizing quantization intervals with task loss. In CVPR, pages 4350–4359, 2019.
  15. 15.Mart´ın Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. Tensorflow: A system for large-scale machine learning. In OSDI, pages 265–283, 2016.
  16. 16.Ruihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li, Peng Hu, Jiazhen Lin, Fengwei Yu, and Junjie Yan. Differentiable soft quantization: Bridging full-precision and low-bit neural networks. In ICCV, pages 4852–4861, 2019.
  17. 17.Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. Pact: Parameterized clipping activation for quantized neural networks. arXiv, 2018.
  18. 18.Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha. Learned step size quantization. In ICLR, 2019.
  19. 19.Zhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu, Chao Xu, Dacheng Tao, and Chang Xu. Searching for low-bit weights in quantized neural networks. In NeurIPS, 2020.
  20. 20.Yuhang Li, Xin Dong, and Wei Wang. Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. In ICLR, 2019.
  21. 21.Jiwei Yang, Xu Shen, Jun Xing, Xinmei Tian, Houqiang Li, Bing Deng, Jianqiang Huang, and Xian-sheng Hua. Quantization networks. In CVPR, pages 7308–7316, 2019.
  22. 22.Shu-Chang Zhou, Yu-Zhi Wang, He Wen, Qin-Yao He, and Yu-Heng Zou. Balanced quantization: An effective and efficient approach to quantized neural networks. Journal of Computer Science and Technology, 32(4):667–682, 2017.
  23. 23.Fabien Cardinaux, Stefan Uhlich, Kazuki Yoshiyama, Javier Alonso Garc´ıa, Lukas Mauch, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura. Iteratively training look-up tables for network quantization. IEEE Journal of Selected Topics in Signal Processing, 14(4):860–870, 2020.
  24. 24.Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Binaryconnect: Training deep neural networks with binary weights during propagations. In NeurIPS, 2015.
  25. 25.Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally. Trained ternary quantization. In ICLR, 2017.
  26. 26.Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In ECCV, pages 525–542, 2016.
  27. 27.Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv, 2016.
  28. 28.Karen Ullrich, Edward Meeds, and Max Welling. Soft weight-sharing for neural network compression. In ICLR, 2017.
  29. 29.Yuhui Xu, Yongzhuang Wang, Aojun Zhou, Weiyao Lin, and Hongkai Xiong. Deep neural network compression with single and multiple level quantization. In AAAI, volume 32, 2018.
  30. 30.Kohei Yamamoto. Learnable companding quantization for accurate low-bit neural networks. In CVPR, pages 5029–5038, 2021.
  31. 31.Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua. Lq-nets: Learned quantization for highly accurate and compact deep neural networks. In ECCV, pages 365–382, 2018.
  32. 32.Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen. Incremental network quantization: Towards lossless cnns with low-precision weights. In ICLR, 2017.
  33. 33.Daisuke Miyashita, Edward H Lee, and Boris Murmann. Convolutional neural networks using logarithmic data representation. arXiv, 2016.
  34. 34.Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image super-resolution. In CVPR, pages 2472–2481, 2018.
  35. 35.Tao Dai, Jianrui Cai, Yongbing Zhang, Shu-Tao Xia, and Lei Zhang. Second-order attention network for single image super-resolution. In CVPR, 2019.
  36. 36.Yiqun Mei, Yuchen Fan, Yuqian Zhou, Lichao Huang, Thomas S Huang, and Honghui Shi. Image super-resolution with cross-scale non-local attention and exhaustive self-exemplars mining. In CVPR, pages 5690–5699, 2020.
  37. 37.Zheng Hui, Xinbo Gao, Yunchu Yang, and Xiumei Wang. Lightweight image super-resolution with information multi-distillation network. In ACM MM, 2019.
  38. 38.Yinglan Ma, Hongyu Xiong, Zhe Hu, and Lizhuang Ma. Efficient super resolution using binarized neural network. In CVPRW, pages 0–0, 2019.
  39. 39.Wonkyung Lee, Junghyup Lee, Dohyung Kim, and Bumsub Ham. Learning with privileged information for efficient image super-resolution. In ECCV, 2020.
  40. 40.Jingwei Xin, Nannan Wang, Xinrui Jiang, Jie Li, Heng Huang, and Xinbo Gao. Binarized neural network for single image super resolution. In ECCV, pages 91–107. Springer, 2020.
  41. 41.Huixia Li, Chenqian Yan, Shaohui Lin, Xiawu Zheng, Yuchao Li, Baochang Zhang, Fan Yang, and Rongrong Ji. Pams: Quantized super-resolution via parameterized max scale. In ECCV, 2020.
  42. 42.Longguang Wang, Xiaoyu Dong, Yingqian Wang, Xinyi Ying, Zaiping Lin, Wei An, and Yulan Guo. Exploring sparsity in image super-resolution for efficient inference. In CVPR, 2021.
  43. 43.Cheeun Hong, Heewon Kim, Junghun Oh, and Kyoung Mu Lee. Daq: Distribution-aware quantization for deep image super-resolution networks. arXiv, 2020.
  44. 44.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, pages 652–660, 2017.
  45. 45.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on x-transformed points. In NeurIPS, volume 31, pages 820–830, 2018.
  46. 46.Zhiyuan Zhang, Binh-Son Hua, and Sai-Kit Yeung. Shellnet: Efficient point cloud convolutional neural networks using concentric shells statistics. In ICCV, pages 1607–1616, 2019.
  47. 47.Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, Franc¸ois Goulette, and Leonidas J Guibas. Kpconv: Flexible and deformable convolution for point clouds. In ICCV, pages 6411–6420, 2019.
  48. 48.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017.
  49. 49.Qiangui Huang, Weiyue Wang, and Ulrich Neumann. Recurrent slice networks for 3d segmentation of point clouds. In CVPR, pages 2626–2635, 2018.
  50. 50.Yiru Shen, Chen Feng, Yaoqing Yang, and Dong Tian. Mining point cloud local structures by kernel correlation and graph pooling. In CVPR, pages 4548–4557, 2018.
  51. 51.Lei Wang, Yuchun Huang, Yaolin Hou, Shenman Zhang, and Jie Shan. Graph attention convolution for point cloud semantic segmentation. In CVPR, pages 10296–10305, 2019.
  52. 52.Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, and Hao Su. Bipointnet: Binary neural network for point clouds. In ICLR, 2021.
  53. 53.Yoshua Bengio, Nicholas L´eonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv, 2013.
  54. 54.Jungwook Choi, Swagath Venkataramani, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan, Zhuo Wang, and Pierce Chuang. Accurate and efficient 2-bit quantized neural networks. In MLSys, 2019.
  55. 55.Jung Hyun Lee, Jihun Yun, Sung Ju Hwang, and Eunho Yang. Cluster-promoting quantization with bit-drop for minimizing network quantization loss. In ICCV, pages 5370–5379, 2021.
  56. 56.Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, and Ian Reid. Towards effective low-bitwidth convolutional neural networks. In CVPR, pages 7920–7928, 2018.
  57. 57.A. Krizhevsky and G. E. Hinton. Learning multiple layers of features from tiny images. Technical report, 2009.
  58. 58.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  59. 59.Zhaowei Cai, Xiaodong He, Jian Sun, and Nuno Vasconcelos. Deep learning with low precision by half-wave gaussian quantization. In CVPR, pages 5918–5926, 2017.
  60. 60.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009.
  61. 61.Xiaofan Lin, Cong Zhao, and Wei Pan. Towards accurate binary convolutional neural network. In NeurIPS, volume 30, 2017.
  62. 62.Eirikur Agustsson and Radu Timofte. NTIRE 2017 challenge on single image super-resolution: Dataset and study. In CVPRW, pages 1122–1131, 2017.
  63. 63.Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie-Line Alberi-Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In BMVC, pages 1–10, 2012.
  64. 64.Roman Zeyde, Michael Elad, and Matan Protter. On single image scale-up using sparse-representations. In International Conference on Curves and Surfaces, volume 6920, pages 711–730, 2010.
  65. 65.David Martin, Charless Fowlkes, Doron Tal, Jitendra Malik, et al. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001.
  66. 66.Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In CVPR, pages 5197–5206, 2015.
  67. 67.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In CVPR, 2017.
  68. 68.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. In CVPR, pages 1912–1920, 2015.

Citation

MLA
Wang, L., et al. “Learnable Lookup Table for Neural Network Quantization”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12413–23, https://doi.org/10.1109/CVPR52688.2022.01210.
APA
Wang, L., Dong, X., Wang, Y., Liu, L., An, W., & Guo, Y. (2022). Learnable Lookup Table for Neural Network Quantization. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12413–12423. https://doi.org/10.1109/CVPR52688.2022.01210
Chicago
Wang, L., X. Dong, Y. Wang, L. Liu, W. An, and Y. Guo. 2022. “Learnable Lookup Table for Neural Network Quantization”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12413–23. https://doi.org/10.1109/CVPR52688.2022.01210.
Harvard
Wang, L. et al. (2022) “Learnable Lookup Table for Neural Network Quantization”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 12413–12423. Available at: https://doi.org/10.1109/CVPR52688.2022.01210.
Vancouver
1. Wang L, Dong X, Wang Y, Liu L, An W, Guo Y (2022) Learnable Lookup Table for Neural Network Quantization. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 12413–12423

BibTeX

@inproceedings{Wang_2022, title={Learnable Lookup Table for Neural Network Quantization}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01210}, DOI={10.1109/cvpr52688.2022.01210}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Longguang and Dong, Xiaoyu and Wang, Yingqian and Liu, Li and An, Wei and Guo, Yulan}, year={2022}, month=June, pages={12413–12423} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE