Oscillation-free Quantization for Low-bit Vision Transformers
Shih-Yang LiuZechun LiuKwang-Ting Cheng
Develops statistical weight quantization, confidence-guided annealing, and query-key reparameterization to eliminate weight oscillation in quantization-aware training, significantly closing the accuracy gap between low-bit Vision Transformers and their full-precision counterparts on ImageNet.
Deploying advanced artificial intelligence models, such as vision transformers, on resource-constrained edge devices requires compressing them to operate with lower numerical precision, a technique known as quantization. However, low-bit quantization often causes a significant drop in accuracy. A primary driver of this performance loss during training is weight oscillation, where model parameters continually jump across discrete thresholds instead of settling into optimal values. The article investigates the underlying causes of this instability in vision transformers and evaluates a new quantization framework designed to eliminate weight oscillation and recover model accuracy.
The article demonstrates that standard learnable scaling factors—widely used to adjust quantization step sizes—create a destabilizing feedback loop with outlier weights, persistently driving adjacent weights into oscillation. Furthermore, the structural design of self-attention mechanisms causes coupled oscillations between internal components during training. To address these root causes, the researchers developed a three-part method called Oscillation-Free Quantization. This approach replaces learnable scaling with a stable statistical calculation, temporarily freezes stable parameters while fine-tuning uncertain ones until they exit boundary regions, and mathematically reorders internal operations to decouple interacting components. The framework was evaluated across standard vision transformer benchmarks using the ImageNet image classification dataset.
The experimental findings show substantial performance gains, particularly at extreme low-bit precisions where oscillation issues are most severe. For 2-bit models, the proposed framework improved top-1 classification accuracy by approximately 9.9% on DeiT-Tiny, 7.7% on DeiT-Small, and 4.6% on Swin-Tiny over previous leading methods. For 3-bit models, the framework achieved performance comparable to full-precision, uncompressed baselines. In 4-bit configurations, the quantized models consistently matched or slightly exceeded full-precision accuracy. Ablation analyses confirmed that all three proposed techniques contributed positively and worked cooperatively to stabilize training.
These results indicate that training instability, rather than the fundamental capacity limit of compressed architectures, has been a primary barrier to deploying highly compressed vision models. Eliminating oscillation reduces the deployment footprint and computational cost of transformer models without incurring major accuracy penalties. Practitioners seeking to deploy vision transformers on edge hardware should adopt statistical scaling and targeted annealing strategies in their compression pipelines. Organizations should test this framework across other transformer architectures and downstream tasks to determine if these stabilization benefits generalize broadly across deep learning domains.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). Provides a comprehensive overview of fundamental neural network quantization principles and low-bit optimization trade-offs that underpin the source's compression framework.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Introduces the Data-efficient Image Transformer (DeiT) architectures and training paradigms that serve as the primary evaluation benchmarks in the source.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Establishes the foundational Vision Transformer (ViT) architecture whose self-attention mechanics and parameter dynamics are analyzed for quantization instability in the source.
- Paper: Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, Benoit Jacob et al. (2018). Presents the standard quantization-aware training mechanics and scaling factor formulations that the source identifies as causes of weight oscillation in transformers.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). Analyzes the emergent outlier feature distributions in transformer architectures that destabilize scale parameters during low-bit quantization.
- Paper: DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients, Shuchang Zhou et al. (2016). Establishes early low-bitwidth quantization mechanics and gradient estimator behaviors crucial for understanding extreme low-precision training stability.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Extends the principle of outlier-aware parameter stabilization to sub-4-bit transformer compression without full backpropagation training.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). Applies advanced randomized incoherence transformations and lattice codebooks to suppress extreme quantization errors in 2-bit and 3-bit transformer regimes.
- Paper: Vision Transformers Need Registers, Timothée Darcet et al. (2024). Investigates the architectural emergence of high-norm outlier tokens in vision transformers that drive representation and scale instabilities.
- Paper: Compressing Transformers: Features Are Low-Rank, but Weights Are Not!, Hao Yu et al. (2023). Explores an alternative compression paradigm for vision transformers by factorizing low-rank activation features rather than low-bit parameter quantization.
- Paper: Adaptive Data-Free Quantization, Biao Qian et al. (2023). Addresses low-bit quantization challenges under data-free settings by dynamically adapting calibration boundaries to model capacity.
