DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory
Jerry CheeArturs BackursRainie HeckLi ZhangJanardhan (Jana) KulkarniThomas RothvossSivakanth Gopi
Introduces DiscQuant, a weight-rounding algorithm grounded in discrepancy theory that guarantees bounded quantization error and significantly outperforms standard post-training methods like GPTQ on low-bit large language model compression.
Modern large language models require substantial memory and computational resources, creating high operational costs and serving bottlenecks during deployment. Post-training quantization reduces these costs by compressing model weights into lower-bit formats, but standard rounding methods often degrade model accuracy. Most prior research focused on designing low-bit grids rather than optimizing how continuous weights are rounded to discrete values. The article introduces DiscQuant, a data-dependent rounding method grounded in mathematical discrepancy theory, designed to compress entire models simultaneously with minimal loss in task performance.
To develop this method, the authors prove theoretically that model loss changes can be effectively approximated by first-order gradient terms and that sample gradients exhibit a rapidly decaying, low-rank structure. Drawing inspiration from discrepancy theory algorithms, the authors formulate a practical optimization objective that minimizes the discrepancy between original and compressed model outputs via knowledge distillation combined with a linear regularization term. The approach was evaluated by quantizing two popular open-source models, Phi-3-mini-4k-instruct and Meta-Llama-3.1-8B-Instruct, across standard benchmarks spanning mathematical reasoning, perplexity, and commonsense question answering across 3-bit to 4.5-bit compression levels.
DiscQuant substantially outperforms current standard rounding baselines, particularly at aggressive low-bit settings. On a challenging math benchmark, rounding Phi-3-mini to 3.25 bits per parameter with DiscQuant achieved 64.2% accuracy, compared to 54.3% for the leading baseline GPTQ and 31.0% for simple round-to-nearest. Across multiple commonsense reasoning benchmarks, DiscQuant recovered full baseline accuracy using at least 0.25 fewer bits per parameter than competing approaches. Furthermore, the experiments demonstrate that DiscQuant integrates seamlessly with orthogonal compression techniques, such as incoherence processing, to provide additional performance improvements at 3-bit precision.
These findings indicate that intelligent, global weight rounding can drastically lower deployment memory overhead without the steep accuracy penalties historically seen in ultra-low-bit quantization. Organizations can serve larger, more capable models on smaller hardware footprints, directly lowering serving infrastructure expenses while preserving generation quality. Because DiscQuant works across arbitrary pre-existing quantization grids, engineering teams can adopt the rounding method without rewriting existing hardware-optimized inference kernels.
Engineering teams preparing to deploy quantized models should adopt DiscQuant for post-training weight rounding, especially when targeting sub-4-bit compression. When implementing the algorithm, teams must carefully curate calibration datasets, as empirical tests show task performance is sensitive to the calibration data distribution. Future work should focus on developing principled guidelines for calibration data selection and extending DiscQuant to vector quantization formats. Confidence in the reported results is high across evaluated models and benchmarks, though practitioners should account for memory overhead during the optimization phase, which mirrors standard two-model knowledge distillation.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). DiscQuant directly compares against and aims to improve upon GPTQ's post-training weight-rounding scheme for large language models.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Understanding activation-aware post-training weight quantization establishes the foundation for how data-dependent rounding and scaling mitigate quantization error in large models.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). QuIP# demonstrates advanced adaptive block rounding and lattice quantization, providing important context for data-dependent discrete rounding formulations.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). This survey provides essential background on standard rounding baselines, grid designs, and the fundamental trade-offs in neural network post-training quantization.
- Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). KronQ extends post-training quantization for large language models by incorporating Kronecker-factored gradient covariance information to advance beyond standard second-order rounding solvers.
