DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
Haokun LinHaobo XuYichen WuJingzhi CuiYingtao ZhangLinzhan MouLinqi SongZhenan SunYing Wei
Develops DuQuant, a quantization method combining block-wise rotation and zigzag permutation to redistribute extreme activation outliers across channels, enabling accurate and efficient 4-bit weight-activation large language model compression without heavy optimization.
Deploying massive large language models into commercial operations and edge hardware is severely constrained by substantial memory footprints and high inference latency. Quantization—the process of compressing floating-point numerical values into low-bit integers—offers an effective path to reduce operational overhead. However, when compressing both model weights and activation states to aggressive 4-bit formats, models routinely suffer severe performance degradation. This loss of accuracy is primarily driven by activation outliers, particularly newly recognized "massive outliers" that exhibit extreme magnitudes in feed-forward down-projection layers and resist conventional smoothing techniques.
The article introduces and evaluates DuQuant (Dual transformations Quantization), an innovative post-training quantization method designed to eliminate both persistent normal outliers and sparse massive outliers without requiring complex model retraining. DuQuant operates by redistributing extreme values across feature dimensions using structured mathematical transformations: it applies diagonal block-wise rotation matrices to locally smooth activations, alternates these with a zigzag permutation that balances activation magnitudes globally across blocks, and applies a secondary rotation to achieve a uniform activation landscape while simultaneously smoothing model weights.
The authors conducted extensive empirical evaluations across multiple open-source model families (LLaMA, LLaMA2, LLaMA3, Mistral, Phi-2, and Vicuna) spanning sizes from 2.8 billion to 70 billion parameters across diverse benchmarks, including language generation, commonsense reasoning, multitasking understanding, and long-context processing. Across standard 4-bit weight-activation benchmarks, DuQuant consistently outperformed existing state-of-the-art compression techniques. On commonsense reasoning tasks, it improved zero-shot accuracy by approximately 5% over Atom and 9% over QLLM across all LLaMA model sizes. In multitasking benchmarks on Vicuna-13B, it achieved up to a 10% gain over top baselines. In operational tests on LLaMA2-7B, DuQuant accelerated the initial prompt processing phase by up to 2.08× and reduced peak memory consumption during token generation by 3.50×, suffering only a 2.71% drop in accuracy relative to uncompressed full-precision models. Furthermore, its execution runtime is remarkably low, quantizing a 13-billion parameter model in approximately 71 seconds compared to hours required by optimization-based baselines.
These findings indicate that dual transformation post-training quantization resolves a critical bottleneck in deploying highly compressed language models at scale. By avoiding expensive gradient-based parameter training and maintaining near-lossless performance at 4-bit precision, organizations can drastically reduce enterprise hosting costs, lower latency, and facilitate deployment on resource-constrained edge devices.
Technical leaders and practitioners looking to optimize model deployment should consider adopting DuQuant pipelines for low-bit serving infrastructure. Organizations can implement the standard round-to-nearest formulation for maximum efficiency, or integrate learnable weight clipping if minor accuracy recovery is required for mission-critical tasks.
Decision-makers should note that the evaluation relied on standardized calibration sets using fixed sample sequences, although preliminary tests indicate the approach remains robust under randomly generated calibration data. The results provide high confidence in standard text-generation workloads, though application to specialized domains or non-transformer architectures should be validated via initial pilot testing.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). Introduces the foundational concept of equivalent mathematical transformations to smooth activation outliers by migrating difficulty into weight matrices for post-training quantization.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). Identifies the emergence of systematic activation outliers in large language models and demonstrates why they cause traditional low-bit quantization methods to collapse.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Establishes activation-aware weight scaling and protection strategies to enable accurate sub-4-bit post-training quantization without retraining.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). Develops fast, second-order post-training quantization algorithms that serve as the primary baseline and operational paradigm for LLM compression.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). Provides a comprehensive overview of fundamental neural network quantization principles, calibration schemes, and hardware-efficiency trade-offs.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). Extends rotation- and transformation-based outlier suppression by applying randomized Hadamard incoherence paired with lattice codebooks for extreme sub-4-bit quantization.
- Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). Builds on activation and weight incoherence rotations by combining bidirectional orthogonal transforms with Kronecker-factored gradient covariance for ultra-low-bit LLM compression.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). Applies geometric transformation principles to sequence compression by converting preconditioned embeddings into polar coordinates for efficient KV cache quantization.
- Paper: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate, Amir Zandieh et al. (2026). Generalizes online randomized rotation transformations to vector quantization with provably near-optimal distortion bounds across key-value caches and embeddings.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). Applies randomized dimensionality and transform techniques specifically to eliminate outlier impacts during 1-bit KV cache quantization.
