FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
Jung Hyun LeeJeonghoon KimSe Jung KwonDongsoo Lee
Proposes FlexRound, a post-training weight quantization method that uses element-wise division to adaptively scale individual weights by magnitude alongside a shared grid size, enabling uniform low-bit quantization for vision architectures and large language models with negligible accuracy loss.
Deploying advanced artificial intelligence models on resource-constrained hardware requires model compression techniques to lower computational and memory overhead. Post-training quantization—the process of converting high-precision numerical values to lower-bit representations without full retraining or extensive datasets—offers an efficient deployment pathway. However, conventional rounding methods rely on additive shifts that restrict weight adjustments to immediately adjacent values and typically keep the overall quantization scale fixed, causing significant accuracy drops in low-bit environments.
The article demonstrates and evaluates FlexRound, a learnable weight-rounding framework based on element-wise division rather than addition. The primary objective is to simultaneously optimize a shared layer-level quantization grid scale and individual weight scales, allowing the model to adaptively round weights across a wider range of discrete values based on their numerical importance.
The researchers assessed FlexRound across extensive benchmarks covering computer vision, natural language understanding, and natural language generation. Testing utilized architectures such as ResNet, MobileNetV2, BERT, OPT, GPT-Neo, GPT-2, and large models like LLaMA-33B. Credibility was supported by testing both weights-only and joint weight-activation quantization across multiple bit-widths (2-bit, 3-bit, 4-bit, and 8-bit) using standard calibration sample sizes (typically 128 to 1,024 samples) and comparing directly against established baselines like AdaRound and AdaQuant.
The experimental findings show that FlexRound consistently outperforms prior rounding methods across domains. In vision models, it significantly rescued low-bit MobileNetV2 performance, reaching 51.49% top-1 accuracy in a 3-bit weight and activation setup where AdaRound achieved only 39.86%. For language understanding benchmarks on the GLUE dataset, 8-bit quantized models using FlexRound matched or approached full-precision accuracy. Similarly, in large language models, 8-bit quantization on LLaMA-33B preserved near-baseline performance across zero-shot reasoning benchmarks and causal language modeling (yielding a perplexity of 6.82 versus 6.35 for the original half-precision model, well ahead of AdaRound's 10.39).
These results imply that organizations can compress deep learning networks down to low integer precision to dramatically decrease memory footprints and hardware costs while preserving model accuracy. Because the division-based formulation inherently accounts for weight magnitude during updates, FlexRound eliminates the need to rely on assumptions about outlier patterns or brittle weight equalization preprocessing steps.
For practical implementation, teams deploying deep learning models on constrained hardware should consider adopting division-based post-training quantization pipelines. Where extreme compression is required, practitioners should ensure calibration sample sizes do not fall below key thresholds (e.g., at least 32 to 64 samples) and perform light tuning of learning rates on task-specific layers. While confidence in the reported results is high across standard vision and language benchmarks, future work should explore the formal combination of division-based scaling with other additive techniques across even broader edge-device hardware constraints.
No sufficiently relevant recommendations were found.
- Paper: DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory, Jerry Chee et al. (2025). After seeing FlexRound learnable rounding by element-wise division, read DiscQuant to follow the move toward data-dependent rounding optimized through discrepancy theory and model-output preservation.
