Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
Zechun LiuKwang-Ting ChengDong HuangEric P. XingZhiqiang Shen
Presents a hardware-friendly quantization framework that pairs learnable nonuniform input thresholds with uniform output levels, using a generalized straight-through estimator to match the representational capacity of nonuniform methods without incurring inference deployment overhead.
Deploying deep neural networks to resource-constrained edge devices requires compressing model size and accelerating computation. While low-bit quantization reduces memory and computation demands by replacing heavy mathematical operations with efficient bitwise calculations, it often degrades model accuracy. Nonuniform quantization strategies retain accuracy better by adapting to the underlying data distribution, but they produce floating-point outputs that require lookup tables and extra post-processing, introducing substantial hardware area and energy overhead. Standard uniform quantization is hardware-friendly but suffers from rigid intervals that lead to severe information loss.
The article develops and evaluates Nonuniform-to-Uniform Quantization (N2UQ), a framework designed to achieve the high accuracy of nonuniform quantization while maintaining the hardware simplicity and operational efficiency of uniform quantization.
To bridge this gap, the approach enforces equidistant, uniform output values to allow direct hardware acceleration while learning flexible, non-equidistant input thresholds during training. Because calculating gradients with respect to learnable threshold parameters is mathematically intractable under standard straight-through estimation, the authors introduce a Generalized Straight-Through Estimator (G-STE) derived from the expected values of stochastic quantization. In addition, an entropy-preserving weight regularization technique is implemented to distribute weights evenly across quantization levels, preventing values from collapsing around zero and maximizing retained information. The method was evaluated on the standard ImageNet classification benchmark across multiple architectures (ResNet-18, ResNet-34, ResNet-50, and MobileNetV2) across 2-bit, 3-bit, and 4-bit configurations.
The evaluation produced several key findings. First, N2UQ consistently outperformed existing uniform and nonuniform quantization methods across all tested bit-widths, exceeding prior state-of-the-art nonuniform techniques by 0.5% to 1.7% in top-1 accuracy on ImageNet. Second, the 2-bit ResNet-50 model achieved 75.8% top-1 accuracy, substantially narrowing the gap to its full-precision counterpart to just 0.6% to 1.2%. Third, on compact models such as MobileNetV2, N2UQ attained 72.1% top-1 accuracy, matching or slightly exceeding full-precision baselines by mitigating overfitting through effective regularization. Finally, ablation studies showed that the G-STE threshold-learning activation quantizer and entropy-preserving weight regularization contributed 3.0% and 1.9% accuracy gains, respectively, over the 2-bit baseline.
These findings indicate that hardware efficiency does not require sacrificing representational flexibility. By generating uniform outputs directly, engineering teams can eliminate lookup tables and dedicated translation hardware, thereby reducing silicon area, memory footprint, latency, and power consumption on mobile and edge devices. Furthermore, the ability to train low-bit networks that match full-precision performance significantly reduces deployment risk for latency-critical applications.
Organizations developing or deploying low-power machine learning systems should consider adopting nonuniform-to-uniform quantization schemes and testing N2UQ on their vision workloads. For immediate implementation, teams can leverage the publicly available codebase to quantize existing convolutional backbones. Prior to enterprise-wide adoption, engineering teams should conduct pilot deployments on target hardware to benchmark actual latency, memory savings, and power efficiency against existing uniform integer pipelines.
While the results demonstrate high confidence on standard computer vision benchmarks and residual network architectures, the evaluations in the article are confined to image classification tasks. Stakeholders should exercise caution when extending the approach to non-vision architectures, such as large language models or transformers, where further empirical validation will be required.
- Paper: Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation, Yoshua Bengio et al. (2013). It introduces the Straight-Through Estimator heuristic for backpropagating through non-differentiable step functions, which forms the direct foundation generalized by N2UQ's Generalized Straight-Through Estimator.
- Paper: Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, Benoit Jacob et al. (2018). It establishes the foundational uniform quantization and simulated quantization-aware training paradigm that N2UQ seeks to optimize without losing representational capacity.
- Paper: DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients, Shuchang Zhou et al. (2016). It details gradient estimation and quantization-aware training across weights and activations for arbitrary bit-widths, setting the stage for threshold-level gradient optimization.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). It provides a comprehensive taxonomy of uniform versus non-uniform quantization trade-offs and hardware execution bottlenecks that motivate N2UQ's hybrid design.
- Paper: Binarized Neural Networks, Matthieu Courbariaux et al. (2016). It demonstrates practical implementation of straight-through estimators in low-bit neural network training under discrete step constraints.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). It introduces non-uniform weight sharing and clustering in neural network compression, illustrating the representational benefits and deployment trade-offs N2UQ addresses.
- Paper: Learnable Lookup Table for Neural Network Quantization, Longguang Wang et al. (2022). It extends learnable non-linear quantization by formulating layer-specific adaptive mappings as differentiable lookup tables for efficient deployment.
- Paper: OMPQ: Orthogonal Mixed Precision Quantization, Yuexiao Ma et al. (2023). It builds upon low-bit quantization principles by using layer orthogonality metrics to optimize mixed-precision bit allocations.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). It applies advanced quantization-aware training techniques in conjunction with self-distillation to achieve stable sub-4-bit compression in large language models.
