Learnable Lookup Table for Neural Network Quantization
Longguang WangXiaoyu DongYingqian WangLi LiuWei AnYulan Guo
Proposes differentiable, learnable lookup tables that adaptively fit non-uniform weight and activation distributions during training while enabling fast, low-overhead table lookups during inference across vision tasks.
Deploying modern deep neural networks to resource-constrained edge and mobile devices remains a major challenge due to high memory footprint and computation costs. Network quantization reduces these resource demands by converting full-precision values into compact, low-bit representations. However, standard linear quantizers fail to accommodate the non-uniform, bell-shaped distributions typical of neural network weights and activations across different layers. While prior parameterized quantizers attempt to adapt to these distributions, they rely on complex non-linear mathematical functions that incur significant computational overhead during real-time activation quantization.
The article demonstrates an efficient uniform quantization framework that formulates the quantization process as a differentiable, learnable lookup table. The primary objective is to learn adaptive layer-specific quantization mappings during training while maintaining minimal computational and memory overhead during hardware deployment.
To achieve this, the authors designed a continuous relaxation scheme that allows discrete lookup tables to be optimized end-to-end alongside standard network parameters. The training strategy incorporates an exponential formulation of scaling parameters to prevent numerical instability, alongside a gradient rescaling mechanism to counterbalance the severe gradient skew between central and extreme value distributions. The method was evaluated across diverse benchmarks covering 2D image classification (CIFAR-10 and ImageNet), image super-resolution (DIV2K, Set5, Set14, B100, Urban100), and 3D point cloud classification (ModelNet40) across several bit-width configurations ranging from 2-bit to 8-bit precision.
The findings establish that the learnable lookup table consistently achieves state-of-the-art accuracy across multiple visual and spatial tasks. On ImageNet classification, a 4-bit ResNet-18 model achieved 70.4% Top-1 accuracy, matching or exceeding full-precision baseline performance while outperforming competing quantization techniques. In low-bit edge settings, such as 2-bit point cloud classification, the proposed method surpassed the nearest competing approach by 2.9 percentage points. Crucially, inference profiling on mobile and CPU hardware revealed that the lookup operation requires only about 1.25 kilobytes of extra memory while reducing activation quantization runtime by more than 50% to 75% compared to complex functional quantizers, cutting mobile latency from 95–120 milliseconds down to 40 milliseconds.
These results demonstrate that complex mathematical approximations can be replaced with direct table lookups without sacrificing model expressiveness. For organizations seeking to deploy real-time artificial intelligence on edge devices, this approach simultaneously lowers latency, memory consumption, and energy use while preserving output quality. Technical teams can adopt this framework directly on off-the-shelf integer-arithmetic hardware without requiring custom non-uniform hardware accelerators. Future deployment initiatives should run hardware-specific pilot validations on target embedded systems, noting that the study primarily evaluated visual and 3D classification architectures rather than larger foundation models.
- Paper: Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, Benoit Jacob et al. (2018). This seminal paper introduces quantization-aware training and integer-only inference frameworks that form the operational baseline and hardware execution model replaced by differentiable lookup tables.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). This survey provides essential background on neural network quantization taxonomy, non-uniform distribution challenges, and hardware execution trade-offs addressed by learnable mapping tables.
- Paper: DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients, Shuchang Zhou et al. (2016). This work establishes core principles of low-bitwidth gradient estimation and backpropagation dynamics that underpin continuous relaxation and gradient rescaling during quantizer training.
- Paper: Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations, Itay Hubara et al. (2016). This paper presents foundational methods for training low-precision weights and activations using continuous approximations and straight-through estimators.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). This foundational work introduces trained codebook and lookup-table weight sharing, providing the conceptual starting point for differentiable lookup table learning.
- Paper: OMPQ: Orthogonal Mixed Precision Quantization, Yuexiao Ma et al. (2023). This work extends layer-specific quantization by using network orthogonality metrics to automate layer-wise bit-width allocation across networks.
- Paper: Adaptive Data-Free Quantization, Biao Qian et al. (2023). This paper tackles the severe accuracy drops in low-bit precision settings by introducing adaptive synthetic data generation for data-free quantization.
- Paper: Oscillation-free Quantization for Low-bit Vision Transformers, Shih-Yang Liu et al. (2023). This study analyzes training instabilities like weight oscillation caused by learnable quantizer scales and provides stabilization methods for attention-based models.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). This paper investigates highly efficient second-order post-training quantization, scaling low-bit compression methods to massive transformer architectures.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). This work advances extreme low-bit compression by pairing lattice-based codebooks with incoherence transformations for highly non-uniform distributions in language models.
