Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training

Charbel SakrSteve DaiRangharajan VenkatesanBrian ZimmerWilliam J. DallyBrucek Khailany

article2022ICML60 citations

Proposes a fast Newton-Raphson-based algorithm to dynamically compute MSE-optimal clipping scalars alongside magnitude-aware differentiation, achieving state-of-the-art accuracy in low-precision quantization-aware training without altering standard baseline hyperparameters.

Listen

Modern deep neural networks deliver high accuracy across vision and language tasks but demand immense computational power and memory. Quantization—reducing numerical precision from standard high-precision formats to low-bit representations—dramatically lowers these hardware costs. However, training networks with low precision (quantization-aware training, or QAT) introduces noise that degrades accuracy, and existing methods rely on heuristic clipping boundaries or complex, hard-to-tune hyperparameters.

The article develops a mathematically rigorous framework that automatically optimizes clipping thresholds in real time during training and improves gradient estimation without altering standard training recipes.

The authors designed a fast recursive algorithm called Optimally Clipped Tensors And Vectors (OCTAV), derived from the Newton-Raphson optimization method, to minimize quantization noise on the fly for every tensor and training step. They also introduced Magnitude-Aware Differentiation (MAD) and a hybrid derivative scheme (MPH) to overcome mathematical flaws in standard gradient estimators, which suffer from gradient explosion or halted parameter updates. The framework was evaluated across standard benchmarks, including training from scratch and retraining ResNet and MobileNet vision models on ImageNet, as well as fine-tuning BERT language models on the SQuAD dataset at 4-bit to 8-bit precision.

The evaluation produced four key findings. First, OCTAV-enabled 4-bit training from scratch achieved state-of-the-art accuracy, maintaining within 1% of the full-precision baseline for ResNet models and MobileNet-V2 without hyperparameter tuning. Second, in 4-bit model retraining, static calibration worked best for larger models like ResNets, while compact architectures like MobileNets suffered catastrophic failure unless dynamic, on-the-fly tracking was applied. Third, for BERT language fine-tuning at 4-bit, OCTAV outperformed standard brute-force sweeps by approximately 1.5% in accuracy because it remained resilient against extreme data outliers. Fourth, OCTAV ran 6 to 10 times faster than brute-force threshold sweeps on central processing units while matching or exceeding their precision.

These findings demonstrate that deep neural networks can be compressed down to 4 bits with negligible accuracy loss, providing a practical path toward lower hardware costs, reduced inference latency, and lower energy consumption. Because OCTAV directly minimizes quantization noise without requiring specialized distillation techniques or hyperparameter sweeps, engineering teams can integrate it directly into existing training pipelines.

Organizations training or deploying low-precision models should adopt dynamic OCTAV for fine-tuning and compact vision models, while applying static OCTAV calibration when retraining large architectures. While the results provide high confidence across standard convolutional and transformer models, highly compact networks with complex activations (such as MobileNet-V3) still experience noticeable accuracy degradation at 4-bit precision. Further work is recommended to evaluate OCTAV in fully quantized training environments—where backward passes and gradients are also quantized—and to explore combinations with knowledge distillation for ultra-compact architectures.

arXiv: 2206.06501
Cover for Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training

Abstract

Data clipping is crucial in reducing noise in quantization operations and improving the achievable accuracy of quantization-aware training (QAT). Current practices rely on heuristics to set clipping threshold scalars and cannot be shown to be optimal. We propose Optimally Clipped Tensors And Vectors (OCTAV), a recursive algorithm to determine MSE-optimal clipping scalars. Derived from the fast Newton-Raphson method, OCTAV finds optimal clipping scalars on the fly, for every tensor, at every iteration of the QAT routine. Thus, the QAT algorithm is formulated with provably minimum quantization noise at each step. In addition, we reveal limitations in common gradient estimation techniques in QAT and propose magnitude-aware differentiation as a remedy to further improve accuracy. Experimentally, OCTAV-enabled QAT achieves state-of-the-art accuracy on multiple tasks. These include training-from-scratch and retraining ResNets and MobileNets on ImageNet, and Squad fine-tuning using BERT models, where OCTAV-enabled QAT consistently preserves accuracy at low precision (4-to-6-bits). Our results require no modifications to the baseline training recipe, except for the insertion of quantization operations where appropriate.

Table of Contents

  • 1 Introduction
  • 1.1 Quantization-aware training and related works
  • 1.2 Contributions
  • 2 Clipped Quantization
  • 3 Optimally Clipped Tensors And Vectors
  • 4 Improving QAT Gradient Estimation
  • 4.1 Limitations of Current Gradient Estimation
  • 4.2 Magnitude-aware Differentiation
  • 5 Quantization-aware Training Studies
  • 5.1 Training-from-scratch QAT on ImageNet
  • 5.2 Retraining ImageNet networks at 4-bit
  • 5.3 Fine-tuning QAT of BERT Models on Squad
  • 6 Discussion
  • 6.1 Current limitations and directions for future work
  • 6.2 Conclusion
  • Acknowledgement
  • References
  • Supplementary Material
  • A Results for Unsigned Quantization
  • B Proof of Theorem 3.1
  • C Proof of Proposition 4.1
  • D Proof of Proposition 4.2
  • E Experimental Implementations Details
  • E.1 Baseline Training Recipes
  • E.2 Tensor Quantization Specifics
  • E.3 Static Quantization Calibration Details
  • F OCTAV vs. Brute Force Sweep Speed Comparison
  • G When is MSE not Convex?

Knowls

  1. Knowl 1 — Clipped uniform quantization and its MSE objective

    equation

    For signed, uniform BB-bit quantization of a real-valued random variable XX, clipping to [−s,s][-s,s] uses the quantizer

    Q(x)=clip⁡ ⁣(s 21−Bround⁡ ⁣(x2B−1s),−s,s),Q(x)=\operatorname{clip}\!\left(s\,2^{1-B}\operatorname{round}\!\left(\frac{x2^{B-1}}{s}\right),-s,s\right),

    where xx is a scalar input, s>0s>0 is the clipping threshold, BB is the bit width, and rounding is to the nearest integer. Under the paper’s additive quantization-noise model inside the clipping interval, the expected squared error is

    J(s)=4−B3s2∫0sf∣X∣(u) du+∫s∞(s−u)2f∣X∣(u) du,J(s)=\frac{4^{-B}}{3}s^2\int_0^s f_{|X|}(u)\,du+\int_s^\infty (s-u)^2 f_{|X|}(u)\,du,

    where f∣X∣f_{|X|} is the probability density of the magnitude ∣X∣|X|. The first term models discretization noise for values within the interval; the second is the squared clipping error for values beyond it. The objective balances these two sources of error, and its minimizing threshold depends on both the data distribution and the bit width. The paper’s page-3 MSE sweeps for ResNet-50 weight and activation layers show distinct minima across layers and precisions, and the analytical objective tracks the empirically measured quantization error closely.

  2. Knowl 2 — OCTAV computes clipping thresholds recursively

    algorithm

    Optimally Clipped Tensors And Vectors (OCTAV) finds an MSE-minimizing clipping threshold for a tensor or vector using a Newton–Raphson-derived recursion. For a data collection t\mathbf{t} with entries xx, bit width BB, and current threshold sns_n, update

    sn+1=∑x∈t∣x∣ 1{∣x∣>sn}4−B3∑x∈t1{0<∣x∣≤sn}+∑x∈t1{∣x∣>sn},s_{n+1}=\frac{\sum_{x\in\mathbf{t}} |x|\,\mathbf{1}_{\{|x|>s_n\}}}{\frac{4^{-B}}{3}\sum_{x\in\mathbf{t}}\mathbf{1}_{\{0<|x|\le s_n\}}+\sum_{x\in\mathbf{t}}\mathbf{1}_{\{|x|>s_n\}}},

    where 1A\mathbf{1}_{A} is 1 when condition AA holds and 0 otherwise. The zero-valued entries are excluded from the in-range count so sparse tensors do not have their estimated quantization noise inflated. In the paper’s experiments, the initial value is the mean magnitude of the nonzero entries; 10 updates are used, and the final iterate is the selected threshold. Each update requires elementwise magnitude, comparisons, indicator masks, and sum reductions, so its work is linear in the number of entries per iteration. The tensor operations can be broadcast across groups for finer-grained scaling. In the BERT-Base CPU calibration benchmark, OCTAV was 10.2× faster than a 100-point MSE sweep for weights and 6.3× faster for activations.

  3. Knowl 3 — Convex-model optimality and the scope of OCTAV’s guarantee

    theoretical result

    For the clipped-quantization MSE objective J(s)J(s) under the paper’s additive noise model, the second derivative is positive, so the modeled objective is convex in the clipping threshold ss. The paper states that the Newton–Raphson OCTAV recursion therefore converges to the global minimizer of this modeled objective. This result concerns the analytical objective, not every possible elementwise empirical MSE curve. The derivation assumes that the data distribution has no point mass in the vicinity of the iterates; the paper notes this as a condition for correct differentiation of the indicator terms.

  4. Knowl 4 — STE can amplify gradients, while PWL can leave weights unlearned

    theoretical result

    For clipped quantization, the straight-through estimator (STE) assigns derivative 1 even in the clipped regions. Under the paper’s layerwise variance analysis, for an LL-layer network there is a constant δ>0\delta>0 such that at layer ll,

    Var⁡(G^lSTE)Var⁡(Gl)≥(1+δ)L−l,\frac{\operatorname{Var}(\widehat G_l^{\mathrm{STE}})}{\operatorname{Var}(G_l)}\ge (1+\delta)^{L-l},

    where GlG_l is the true loss gradient with respect to the layer-ll activation and G^lSTE\widehat G_l^{\mathrm{STE}} is its STE estimate. Thus, the estimated gradient variance can grow exponentially with the number of layers traversed during backpropagation.

    The piece-wise linear estimator (PWL) instead uses derivative 1{∣x∣≤s}\mathbf{1}_{\{|x|\le s\}}, which avoids that clipped-region contribution but gives zero gradient to clipped weights. With static clipping of an NwN_{\mathbf w}-element weight tensor, the number N~w(i)\widetilde N_{\mathbf w}^{(i)} of parameters receiving updates at iteration ii satisfies Nw>N~w(i)≥N~w(i+1)N_{\mathbf w}>\widetilde N_{\mathbf w}^{(i)}\ge\widetilde N_{\mathbf w}^{(i+1)}: some initially clipped weights remain at their initial values, and the set of weights that can be updated may shrink. The paper also reports a milder version of the initial loss of learnable weights under dynamic quantization.

  5. Knowl 5 — Magnitude-aware differentiation and the MAD–PWL hybrid

    model/method

    The paper treats clipping as magnitude attenuation: for a real input xx and threshold s>0s>0,

    clip⁡(x,−s,s)=αx,α=1{∣x∣≤s}+s∣x∣1{∣x∣>s}.\operatorname{clip}(x,-s,s)=\alpha x,\qquad \alpha=\mathbf{1}_{\{|x|\le s\}}+\frac{s}{|x|}\mathbf{1}_{\{|x|>s\}}.

    Magnitude-aware differentiation (MAD) treats α\alpha as constant when estimating the derivative of the quantized operation, giving

    ∂(MAD)Q(x)∂x=1{∣x∣≤s}+s∣x∣1{∣x∣>s}.\frac{\partial^{(\mathrm{MAD})}Q(x)}{\partial x}=\mathbf{1}_{\{|x|\le s\}}+\frac{s}{|x|}\mathbf{1}_{\{|x|>s\}}.

    Unlike PWL, MAD attenuates rather than zeroes the estimated gradient outside the clipping interval, so clipped weights can continue to receive updates. The authors recommend MAD for weight gradients and PWL for activation gradients: PWL’s occasional zeroing of activation gradients is proposed as a useful regularizing effect. They call this combination MAD–PWL Hybrid (MPH).

  6. Knowl 6 — Gradient-estimator comparison in 4-bit ResNet-50 training

    empirical result

    In 4-bit ResNet-50 ImageNet training-from-scratch, the full-precision baseline reached 76.07% accuracy and max-scaled QAT reached 72.67%. With OCTAV clipping, final accuracies were 67.75% using STE, 74.31% using PWL, 74.81% using MAD, and 75.15% using the MAD–PWL hybrid (MPH). The page-6 convergence curves show the STE run becoming unstable, while PWL trails MAD and MPH. The results support the proposed gradient analysis and the use of MPH in the paper’s subsequent clipped-QAT experiments.

  7. Knowl 7 — ImageNet training-from-scratch accuracy across networks and bit widths

    data/table

    The following ImageNet results compare full-precision baselines with OCTAV-clipped and max-scaled quantization-aware training (QAT), for 4-, 6-, and 8-bit weights and activations. Accuracies are percentages. The page-7 results show that OCTAV particularly helps at 4 bits: ResNets remain within 1 percentage point of baseline, while max-scaling loses as much as about 5 points; MobileNet-V3 also trains successfully with OCTAV at 4 bits where max-scaling nearly fails.

    NetworkFull precision4-bit OCTAV4-bit max-scaling6-bit OCTAV6-bit max-scaling8-bit OCTAV8-bit max-scaling
    ResNet-5076.0775.1572.6776.0776.0176.2476.12
    ResNet-1870.1269.1765.6569.7869.5270.0770.19
    ResNet-10177.2876.4872.5377.3077.0477.3177.15
    MobileNet-V271.7170.8869.1771.6471.7971.7171.77
    MobileNet-V3-Small65.9954.680.3965.0260.1765.9865.14
    MobileNet-V3-Large72.9765.861.2572.1269.3872.8972.78
  8. Knowl 8 — Four-bit ImageNet retraining favors different calibration modes by model size

    data/table

    These ImageNet results compare 4-bit retraining with dynamic quantization against static calibration. Short retraining uses 15 epochs for ResNet-50 and ResNet-101, and 30 epochs for the other networks. Long retraining uses OCTAV for 150 epochs on ResNets and 300 epochs on MobileNets. Accuracies are percentages. The results show static OCTAV performing best among short-retraining methods for ResNets, but static calibration collapsing on MobileNets; dynamic OCTAV retains useful accuracy across both model families. Longer training improves all reported OCTAV results, but does not remove the static-calibration failure on MobileNets.

    NetworkShort dynamic OCTAVShort dynamic max-scalingShort static OCTAVShort static MSE sweepShort static 99.9th percentileShort static 99.99th percentileShort static 99.999th percentile
    ResNet-5075.3871.4475.8475.8575.6675.5175.29
    ResNet-1869.1665.5369.1869.2869.0469.0868.93
    ResNet-10176.1070.8876.9677.0176.7976.9976.34
    MobileNet-V269.3266.940.660.931.722.142.76
    MobileNet-V3-Small53.520.100.430.580.650.101.46
    MobileNet-V3-Large64.9739.460.390.300.340.710.57
    NetworkLong dynamic OCTAVLong static OCTAV
    ResNet-5076.2176.46
    ResNet-1869.9070.13
    ResNet-10176.8477.48
    MobileNet-V271.231.21
    MobileNet-V3-Small58.930.80
    MobileNet-V3-Large69.210.60

    The long-retraining results put ResNet-50 and ResNet-101 at or above their full-precision baselines (76.07% and 77.28%, respectively) with static OCTAV. MobileNet-V3 remains below its baseline even with dynamic OCTAV, indicating that further QAT methods may be needed for that model.

  9. Knowl 9 — BERT fine-tuning on SQuAD favors dynamic OCTAV at low precision

    data/table

    The tables give SQuAD v1.1 F1 scores for QAT fine-tuning of BERT-Large (full-precision baseline 91.00) and BERT-Base (baseline 88.24), across 4–8 bits. Dynamic quantization recalculates thresholds during fine-tuning; static quantization calibrates them beforehand. At 7 bits or more, all strategies are near their respective baselines. At 6 bits, dynamic OCTAV is the only strategy within 1 percentage point of baseline for both models; at 4 bits it degrades more gracefully than the alternatives. Static OCTAV is the strongest static method in most low-bit cases.

    BERT-Large strategy4-bit5-bit6-bit7-bit8-bit
    Dynamic OCTAV87.0989.7790.5190.8190.78
    Dynamic max-scaling6.9280.0687.7190.0490.48
    Static OCTAV87.0889.5490.6090.7990.61
    Static MSE sweep85.5489.7790.3990.8090.55
    Static 99.9th percentile86.9889.7989.9990.0790.11
    Static 99.99th percentile6.9087.6390.3890.7990.33
    Static 99.999th percentile4.565.6689.7690.4490.83
    BERT-Base strategy4-bit5-bit6-bit7-bit8-bit
    Dynamic OCTAV84.5186.3087.4388.2888.34
    Dynamic max-scaling11.5178.9785.1787.4688.01
    Static OCTAV83.6085.8287.1487.6788.02
    Static MSE sweep81.8284.1687.1487.6887.97
    Static 99.9th percentile81.0685.7886.7386.8487.34
    Static 99.99th percentile67.9083.2086.7887.6087.94
    Static 99.999th percentile26.8582.1586.2787.5188.08
  10. Knowl 10 — Large outliers can make empirical clipping MSE nonconvex

    limitation

    The paper’s convexity guarantee applies to the analytical MSE model, which assumes additive quantization noise in the unclipped interval. Empirical MSE can behave differently when a tensor contains very large outliers. In BERT-Base activation calibration at 4 bits, most of the values in one examined layer were concentrated near zero, but some outliers were around 350. For that layer, OCTAV selected a clipping threshold of approximately 8, while the 100-point sweep selected approximately 38. The empirical MSE had one minimum near zero that balances discretization and clipping error across the data, and a second minimum at a larger threshold that represents outliers while quantizing much of the smaller-valued data to zero. OCTAV converged to the first minimum; the sweep could select the second. In the reported 4-bit static fine-tuning results, static OCTAV exceeded the sweep by 1.54 F1 points for BERT-Large (87.08 versus 85.54) and 1.78 points for BERT-Base (83.60 versus 81.82). Thus the analytical model’s global-optimum guarantee should not be interpreted as a guarantee of minimizing every empirical quantization-error curve.

Coverage note — Detailed optimizer schedules, GPU counts, and supplementary calibration implementation settings are omitted because they support reproducibility rather than constitute independent findings; the principal methods, experimental results, and stated limitation are included.

References

  1. 1.Abdolrashidi, A., Wang, L., Agrawal, S., Malmaud, J., Rybakov, O., Leichner, C., and Lew, L. Pareto-optimal quantized resnet is mostly 4-bit. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3091–3099, 2021.
  2. 2.Bianco, S., Cadene, R., Celona, L., and Napoletano, P. Benchmark analysis of representative deep neural network architectures. IEEE Access, 6:64270–64277, 2018.
  3. 3.Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K. PACT: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085, 2018.
  4. 4.Choi, Y., Choi, J., El-Khamy, M., and Lee, J. Data-free network quantization with adversarial knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 710–711, 2020.
  5. 5.Courbariaux, M. et al. Binaryconnect: Training deep neural networks with binary weights during propagations. In Advances in Neural Information Processing Systems, pp. 3123–3131, 2015.
  6. 6.Dai, S., Venkatesan, R., Ren, M., Zimmer, B., Dally, W., and Khailany, B. Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference. Proceedings of Machine Learning and Systems, 3, 2021.
  7. 7.Dbouk, H., Sanghvi, H., Mehendale, M., and Shanbhag, N. Dbq: A differentiable branch quantizer for lightweight deep neural networks. In European Conference on Computer Vision, pp. 90–106. Springer, 2020.
  8. 8.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255. IEEE, 2009.
  9. 9.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S. Learned step size quantization. In International Conference on Learning Representations, 2019.
  11. 11.Goel, M. and Shanbhag, N. Finite-precision analysis of the pipelined strength-reduced adaptive filter. Signal Processing, IEEE Transactions on, 46(6):1763–1769, 1998.
  12. 12.Gonugondla, S. K., Sakr, C., Dbouk, H., and Shanbhag, N. R. Fundamental limits on the precision of in-memory architectures. In Proceedings of the 39th International Conference on Computer-Aided Design, pp. 1–9, 2020.
  13. 13.Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P. Deep learning with limited numerical precision. In International Conference on Machine Learning, pp. 1737–1746, 2015.
  14. 14.Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M. A., and Dally, W. J. Eie: Efficient inference engine on compressed deep neural network. ACM SIGARCH Computer Architecture News, 44(3):243–254, 2016.
  15. 15.He, K., Zhang, X., Ren, S., and Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1026–1034, 2015.
  16. 16.He, K. et al. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  17. 17.Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et al. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1314–1324, 2019.
  18. 18.Hubara, I. et al. Binarized neural networks. In Advances in Neural Information Processing Systems, pp. 4107–4115, 2016.
  19. 19.Jain, S., Venkataramani, S., Srinivasan, V., Choi, J., Gopalakrishnan, K., and Chang, L. BiScaled-DNN: Quantizing long-tailed datastructures with two scale factors for deep neural networks. In 2019 56th ACM/IEEE Design Automation Conference (DAC), pp. 1–6. IEEE, 2019.
  20. 20.Koster, U., Webb, T., Wang, X., Nassar, M., Bansal, A. K., Constable, W., Elibol, O., Hall, S., Hornof, L., Khosrowshahi, A., et al. Flexpoint: An adaptive numerical format for efficient training of deep neural networks. In Advances in Neural Information Processing Systems, pp. 1740–1750, 2017.
  21. 21.LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. nature, 521(7553):436–444, 2015.
  22. 22.Lee, E. H., Miyashita, D., Chai, E., Murmann, B., and Wong, S. S. Lognet: Energy-efficient neural networks using logarithmic computation. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5900–5904. IEEE, 2017.
  23. 23.Lin, Y. et al. PredictiveNet: an energy-efficient convolutional neural network via zero prediction. In Circuits and Systems (ISCAS), 2017 IEEE International Symposium on. IEEE, 2017.
  24. 24.Liu, Z., Wu, B., Luo, W., Yang, X., Liu, W., and Cheng, K.-T. Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. In Proceedings of the European conference on computer vision (ECCV), pp. 722–737, 2018.
  25. 25.Lloyd, S. Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2):129–137, 1982.
  26. 26.Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M. Data-free quantization through weight equalization and bias correction. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1325–1334, 2019.
  27. 27.Park, E. and Yoo, S. Profit: A novel training method for sub-4-bit mobilenet models. In European Conference on Computer Vision, pp. 430–446. Springer, 2020.
  28. 28.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. Automatic differentiation in PyTorch. In NeurIPS Workshop on Automatic Differentiation, 2017.
  29. 29.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392, 2016.
  30. 30.Sakr, C. and Shanbhag, N. R. Per-tensor fixed-point quantization of the back-propagation algorithm. In 7th International Conference on Learning Representations, ICLR 2019, 2019.
  31. 31.Sakr, C. and Shanbhag, N. R. Signal processing methods to enhance the energy efficiency of in-memory computing architectures. IEEE Transactions on Signal Processing, 69:6462–6472, 2021.
  32. 32.Sakr, C. et al. Analytical guarantees on numerical precision of deep neural networks. In Proceedings of the 34th International Conference on Machine Learning, pp. 3007–3016, 2017.
  33. 33.Sakr, C. et al. Accumulation bit-width scaling for ultra-low precision training of deep networks. In 7th International Conference on Learning Representations, ICLR 2019, 2019.
  34. 34.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510–4520, 2018.
  35. 35.Srivastava, N. et al. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958, 2014.
  36. 36.Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V., Cui, X., Zhang, W., and Gopalakrishnan, K. Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks. In NeurIPS, 2019.
  37. 37.Taigman, Y. et al. Deepface: Closing the gap to human-level performance in face verification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1701–1708, 2014.
  38. 38.Tambe, T., Yang, E.-Y., Wan, Z., Deng, Y., Reddi, V. J., Rush, A., Brooks, D., and Wei, G.-Y. Adaptivfloat: A floating-point based data type for resilient deep learning inference. arXiv preprint arXiv:1909.13271, 2019.
  39. 39.Wang, N., Choi, J., Brand, D., Chen, C.-Y., and Gopalakrishnan, K. Training deep neural networks with 8-bit floating point numbers. In Advances in Neural Information Processing Systems, 2018.
  40. 40.Widrow, B. and Kollar, I. Quantization noise. Cambridge University Press, 2:5, 2008.
  41. 41.Wikimedia Foundation. Wikimedia downloads, 2021. URL https://dumps.wikimedia.org.
  42. 42.Wu, H., Judd, P., Zhang, X., Isaev, M., and Micikevicius, P. Integer quantization for deep learning inference: Principles and empirical evaluation. arXiv preprint arXiv:2004.09602, 2020.
  43. 43.Zhang, D., Yang, J., Ye, D., and Hua, G. LQ-Nets: Learned quantization for highly accurate and compact deep neural networks. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 365–382, 2018.
  44. 44.Zhao, J., Dai, S., Venkatesan, R., Liu, M.-Y., Khailany, B., Dally, B., and Anandkumar, A. Low-precision training in logarithmic number system using multiplicative weight update. arXiv preprint arXiv:2106.13914, 2021.
  45. 45.Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160, 2016.
  46. 46.Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pp. 19–27, 2015.

Citation

MLA
Sakr, C., et al. “Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training”. International Conference on Machine Learning, vol. 162, 2022, pp. 19123–38, https://proceedings.mlr.press/v162/sakr22a.html.
APA
Sakr, C., Dai, S., Venkatesan, R., Zimmer, B., Dally, W., & Khailany, B. (2022). Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training. International Conference on Machine Learning, 162, 19123–19138. https://proceedings.mlr.press/v162/sakr22a.html
Chicago
Sakr, C., S. Dai, R. Venkatesan, B. Zimmer, W. Dally, and B. Khailany. 2022. “Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training”. International Conference on Machine Learning 162: 19123–38. https://proceedings.mlr.press/v162/sakr22a.html.
Harvard
Sakr, C. et al. (2022) “Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training”, International Conference on Machine Learning. PMLR, pp. 19123–19138. Available at: https://proceedings.mlr.press/v162/sakr22a.html.
Vancouver
1. Sakr C, Dai S, Venkatesan R, Zimmer B, Dally W, Khailany B (2022) Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training. In: International Conference on Machine Learning. PMLR, pp 19123–19138

BibTeX

@InProceedings{pmlr-v162-sakr22a,
  title = 	 {Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training},
  author =       {Sakr, Charbel and Dai, Steve and Venkatesan, Rangha and Zimmer, Brian and Dally, William and Khailany, Brucek},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {19123--19138},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/sakr22a/sakr22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/sakr22a.html},
  abstract = 	 {Data clipping is crucial in reducing noise in quantization operations and improving the achievable accuracy of quantization-aware training (QAT). Current practices rely on heuristics to set clipping threshold scalars and cannot be shown to be optimal. We propose Optimally Clipped Tensors And Vectors (OCTAV), a recursive algorithm to determine MSE-optimal clipping scalars. Derived from the fast Newton-Raphson method, OCTAV finds optimal clipping scalars on the fly, for every tensor, at every iteration of the QAT routine. Thus, the QAT algorithm is formulated with provably minimum quantization noise at each step. In addition, we reveal limitations in common gradient estimation techniques in QAT and propose magnitude-aware differentiation as a remedy to further improve accuracy. Experimentally, OCTAV-enabled QAT achieves state-of-the-art accuracy on multiple tasks. These include training-from-scratch and retraining ResNets and MobileNets on ImageNet, and Squad fine-tuning using BERT models, where OCTAV-enabled QAT consistently preserves accuracy at low precision (4-to-6-bits). Our results require no modifications to the baseline training recipe, except for the insertion of quantization operations where appropriate.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/