Adaptive Data-Free Quantization

Biao QianYang WangRichang HongMeng Wang

article2023CVPR58 citations

Proposes an adaptive data-free quantization framework that frames synthetic sample generation as a zero-sum game between the generator and quantized network, dynamically balancing agreement and disagreement samples to prevent over- and under-fitting across varied bit-widths.

Listen

Deploying deep neural networks onto resource-constrained edge hardware requires model compression, such as network quantization, to reduce computational demands and memory footprints. When original training datasets cannot be shared due to strict privacy, proprietary, or security regulations, organizations must rely on data-free quantization to compress models using synthetically generated samples. However, conventional synthetic data approaches generate calibration samples independently of the compressed model's internal learning state, leading to severe underfitting at aggressive low-bit precision (such as 3-bit) or overfitting at moderate precision (such as 5-bit).

The article develops and evaluates an Adaptive Data-Free Quantization framework to resolve this limitation. The primary objective is to demonstrate that optimizing synthetic sample generation dynamically around the compressed model's specific learning capacity prevents generalization errors and substantially restores model accuracy across varied bit-width scenarios without accessing private training data.

The researchers formulated the data-free compression process as a two-player zero-sum game between a synthetic sample generator and the compressed model. The approach establishes explicit upper and lower boundaries using agreement samples (where both original and compressed models predict the same class) and disagreement samples (where the original model is correct but the compressed model fails). A regulated margin between these boundaries balances the informativeness of generated samples. The evaluation benchmarked multiple vision architectures, including ResNet-18, ResNet-20, ResNet-50, and MobileNetV2, across standard image recognition datasets such as CIFAR-10, CIFAR-100, and ImageNet under 3-bit, 4-bit, and 5-bit precision.

The findings show that the proposed method consistently outperforms state-of-the-art data-free quantization techniques, achieving its most dramatic improvements in challenging low-bit regimes. In 3-bit compression on ImageNet using ResNet-18, the proposed method achieved an accuracy of 38.10%, delivering a 36.93 percentage point gain over baseline synthetic generation methods that failed to converge. For MobileNetV2 under 3-bit precision, accuracy reached 28.99%, compared to near-complete failure (around 1.46% accuracy) in prior generative methods. In moderate 4-bit and 5-bit precision, the framework matched or exceeded existing benchmarks while effectively curbing overfitting. Ablation analyses confirmed that removing either agreement or disagreement constraints resulted in sharp performance drops of up to 19.57 percentage points, proving that extreme disagreement is counterproductive and that an optimized middle margin is necessary.

These results demonstrate that data-free model compression does not require an all-or-nothing trade-off between model efficiency and predictive reliability. Enterprise and engineering leaders can achieve aggressive low-bit compression on resource-limited hardware while fully complying with data privacy mandates. By calibrating synthetic data generation directly to compressed model capacity, teams can lower hardware deployment costs, decrease latency, and prevent project failure caused by convergence collapse during low-bit deployment.

Organizations aiming to compress edge artificial intelligence models under strict data governance policies should adopt margin-regulated zero-sum generation frameworks. Engineering teams should target the identified stable parameter ranges for boundary margins, which proved robust across grid search evaluations. Before production deployment, teams should conduct pilot calibration runs across their specific edge hardware configurations to determine whether 3-bit or 4-bit quantization best balances execution latency against task-specific accuracy thresholds.

Confidence in these findings is high given the consistent theoretical grounding and extensive benchmarking across diverse network architectures and datasets. However, decision-makers should note that the evaluation is currently bounded to image classification models and convolutional architectures. Further validation is recommended before applying the framework to non-vision modalities, such as natural language processing or multimodal systems.

Cover for Adaptive Data-Free Quantization

Abstract

Data-free quantization (DFQ) recovers the performance of quantized network (Q) without the original data, but generates the fake sample via a generator (G) by learning from full-precision network (P), which, however, is totally independent of Q, resulting into the overflow of generalization error. Building on this, several critical questions — how to measure the sample adaptability to Q under varied bit-width scenarios? whether the largest adaptability is the best? how to generate the samples with adaptive adaptability to improve Q’s generalization? To answer the above questions, in this paper, we propose an Adaptive Data-Free Quantization (AdaDFQ) method, which revisits DFQ from a zero-sum game perspective upon the sample adaptability between two players — a generator and a quantized network. Following this viewpoint, we further define the disagreement and agreement samples to form two boundaries, where the margin between two boundaries is optimized to adaptively regulate the adaptability of generated samples to Q, so as to address the over-and-under fitting issues. Our AdaDFQ reveals: 1) the largest adaptability is NOT the best for sample generation to benefit Q’s generalization; 2) the knowledge of the generated sample should not be informative to Q only, but also related to the category and distribution information of the training data for P. The theoretical and empirical analysis validate the advantages of AdaDFQ over the state-of-the-arts. Our code is available at https://github.com/hfutqian/AdaDFQ.

Table of Contents

  • 1. Introduction
  • 2. Adaptive Data-Free Quantization
  • 2.1. How to Measure the Sample Adaptability to Quantized Network?
  • 2.2. Zero-Sum Game over Sample Adaptability
  • 2.3. Refining the Maximization of Eq.(4): Generating the Sample with Adaptive Adaptability
  • 2.3.1. Balancing Disagreement Sample with Agreement Sample
  • 2.3.2. Optimizing the Margin Between two Boundaries
  • 2.4. Theoretical Analysis: Why do the Boundaries Improve Q's Generalization?
  • 3. Experiment
  • 3.1. Experimental Settings and Details
  • 3.2. Why does AdaDFQ Work?
  • 3.3. Comparison with State-of-the-arts
  • 3.4. Ablation Study
  • 3.4.1. Validating adaptability with disagreement and agreement samples
  • 3.4.2. Why can λl and λu benefit Q?
  • 3.4.3. How to balance disagreement sample with agreement sample?
  • 3.5. Visual Analysis on Generated Samples
  • 4. Conclusion
  • 5. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Minimax Optimization Framework and Objectives for AdaDFQ

    model/method

    Adaptive Data-Free Quantization (AdaDFQ) formulates data-free quantization as a two-player zero-sum game between a generator GG parameterized by θg∈Θg\theta_g \in \Theta_g and a quantized model QQ parameterized by θq∈Θq\theta_q \in \Theta_q.

    The generator GG optimizes the margin between lower bound λl\lambda_l and upper bound λu\lambda_u for the normalized disagreement entropy Hinfo′(pds)H'_{info}(p_{ds}) using hinge loss terms, while balancing disagreement/agreement sample generation (Lbal \mathcal{L}_{bal}) and matching batch normalization statistics (LBNS \mathcal{L}_{BNS}) extracted from the pre-trained full-precision network PP:

    max⁡θg∈ΘgEz,y[−max⁡(λl−Hinfo′(pds),0)]+Ez,y[−max⁡(Hinfo′(pds)−λu,0)]−βLbal−γLBNS\max_{\theta_g \in \Theta_g} \mathbb{E}_{z, y} \left[ -\max\left(\lambda_l - H'_{info}(p_{ds}), 0\right) \right] + \mathbb{E}_{z, y} \left[ -\max\left(H'_{info}(p_{ds}) - \lambda_u, 0\right) \right] - \beta \mathcal{L}_{bal} - \gamma \mathcal{L}_{BNS}

    where z∼N(0,I)z \sim \mathcal{N}(0, I) is a random noise vector, yy is a target one-hot class label, β\beta and γ\gamma are positive balance hyperparameters, and pdsp_{ds} is the disagreement probability distribution between PP and QQ. The batch normalization statistics loss across MM batch normalization layers is defined as:

    LBNS=∑m=1M(∥μmg−μm∥22+∥σmg−σm∥22)\mathcal{L}_{BNS} = \sum_{m=1}^{M} \left( \|\mu_m^g - \mu_m\|_2^2 + \|\sigma_m^g - \sigma_m\|_2^2 \right)

    where μmg,σmg\mu_m^g, \sigma_m^g are the mean and standard deviation of generated batch activations, and μm,σm\mu_m, \sigma_m are running statistics stored in layer mm of model PP.

    The quantized network QQ is calibrated over the generated samples by minimizing the normalized adaptability:

    min⁡θq∈ΘqEz,y[1−Hinfo′(pds)]\min_{\theta_q \in \Theta_q} \mathbb{E}_{z, y} \left[ 1 - H'_{info}(p_{ds}) \right]

    Alternating gradient steps between the two objectives guide the system toward a Nash equilibrium (θg∗,θq∗)(\theta_g^*, \theta_q^*).

  2. Knowl 2 — Disagreement and Agreement Sample Probability Formulations

    definition

    Given a sample x=G(z∣y)x = G(z|y) generated from noise z∼N(0,I)z \sim \mathcal{N}(0, I) conditioned on one-hot label yy, let zp=P(x)∈RCz_p = P(x) \in \mathbb{R}^C and zq=Q(x)∈RCz_q = Q(x) \in \mathbb{R}^C denote the logit vectors produced by the full-precision network PP and the quantized network QQ across CC classes.

    A sample xx is defined as a disagreement sample if PP classifies it correctly while QQ misclassifies it: arg⁡max⁡(zp)=arg⁡max⁡(y)\arg\max(z_p) = \arg\max(y) and arg⁡max⁡(zp)≠arg⁡max⁡(zq)\arg\max(z_p) \neq \arg\max(z_q). The disagreement probability distribution pds∈RCp_{ds} \in \mathbb{R}^C is computed as:

    pds=softmax(zp−zq)p_{ds} = \text{softmax}(z_p - z_q)

    where the cc-th entry pds(c)=exp⁡(zp(c)−zq(c))∑j=1Cexp⁡(zp(j)−zq(j))p_{ds}(c) = \frac{\exp(z_p(c) - z_q(c))}{\sum_{j=1}^C \exp(z_p(j) - z_q(j))} represents the probability that xx is identified as a disagreement sample for class cc.

    A sample xx is defined as an agreement sample if both PP and QQ predict the same class: arg⁡max⁡(zp)=arg⁡max⁡(zq)\arg\max(z_p) = \arg\max(z_q). The agreement probability distribution pas∈RCp_{as} \in \mathbb{R}^C is computed as:

    pas=softmax(zp+zq)p_{as} = \text{softmax}(z_p + z_q)

    where pas(c)=exp⁡(zp(c)+zq(c))∑j=1Cexp⁡(zp(j)+zq(j))p_{as}(c) = \frac{\exp(z_p(c) + z_q(c))}{\sum_{j=1}^C \exp(z_p(j) + z_q(j))}.

  3. Knowl 3 — Sample Adaptability Measurement via Normalized Disagreement Entropy

    definition

    The information entropy of the disagreement probability vector pds∈RCp_{ds} \in \mathbb{R}^C measures the disagreement level between the full-precision network PP and the quantized network QQ:

    Hinfo(pds)=∑c=1Cpds(c)log⁡1pds(c)H_{info}(p_{ds}) = \sum_{c=1}^{C} p_{ds}(c) \log \frac{1}{p_{ds}(c)}

    To standardize this measure across different datasets and class counts CC, Hinfo(pds)H_{info}(p_{ds}) is linearly normalized within a batch:

    Hinfo′(pds)=Hinfo(pds)−min⁡(Hinfo(pds))max⁡(Hinfo(pds))−min⁡(Hinfo(pds))H'_{info}(p_{ds}) = \frac{H_{info}(p_{ds}) - \min(H_{info}(p_{ds}))}{\max(H_{info}(p_{ds})) - \min(H_{info}(p_{ds}))}

    where max⁡(Hinfo(pds))=−∑c=1C1Clog⁡1C=log⁡C\max(H_{info}(p_{ds})) = -\sum_{c=1}^C \frac{1}{C} \log \frac{1}{C} = \log C represents the theoretical maximum entropy where QQ perfectly aligns with PP (zp=zqz_p = z_q), and min⁡(Hinfo(pds))\min(H_{info}(p_{ds})) is the minimum entropy observed within the current batch.

    The sample adaptability to QQ is defined as:

    Hnor=1−Hinfo′(pds)∈[0,1)H_{nor} = 1 - H'_{info}(p_{ds}) \in [0, 1)

    A higher HnorH_{nor} indicates greater disagreement (high adaptability), whereas a lower HnorH_{nor} indicates high prediction agreement (low adaptability).

  4. Knowl 4 — Balanced Disagreement and Agreement Loss Formulation

    model/method

    To prevent the generator GG from producing samples with extreme adaptability values (which cause underfitting when adaptability is excessively high, or overfitting when adaptability is excessively low), AdaDFQ guides GG using a balanced loss Lbal\mathcal{L}_{bal} that couples category-guided disagreement and agreement cross-entropy losses:

    Lds=Ez,y[HCE(pds,y)]\mathcal{L}_{ds} = \mathbb{E}_{z, y} \left[ H_{CE}(p_{ds}, y) \right]

    Las=Ez,y[HCE(pas,y)]\mathcal{L}_{as} = \mathbb{E}_{z, y} \left[ H_{CE}(p_{as}, y) \right]

    Lbal=αdsLds+αasLas\mathcal{L}_{bal} = \alpha_{ds} \mathcal{L}_{ds} + \alpha_{as} \mathcal{L}_{as}

    where HCE(⋅,⋅)H_{CE}(\cdot, \cdot) denotes the standard cross-entropy loss, pds=softmax(zp−zq)p_{ds} = \text{softmax}(z_p - z_q), pas=softmax(zp+zq)p_{as} = \text{softmax}(z_p + z_q), and αds,αas\alpha_{ds}, \alpha_{as} are balancing coefficients. Lds\mathcal{L}_{ds} forces GG to create samples where PP is confident on class yy while QQ is not (lowering Hinfo′(pds)H'_{info}(p_{ds})), while Las\mathcal{L}_{as} forces GG to create samples where both PP and QQ agree on label yy (increasing Hinfo′(pds)H'_{info}(p_{ds})), keeping sample adaptability within an intermediate, learnable regime.

  5. Knowl 5 — Generalization Error Bound Decomposition for Data-Free Quantized Networks

    theoretical result

    Under statistical learning and VC theory, the classification risk R(fq)R(f_q) of a quantized network function fq∈Fqf_q \in \mathcal{F}_q relative to the ground-truth target function fr∈Frf_r \in \mathcal{F}_r on nn generated samples from GG decomposes through the full-precision function fp∈Fpf_p \in \mathcal{F}_p as:

    R(fq)−R(fr)≤O(∣Fq∣Cnαqp)+O(∣Fp∣Cnαpr)⏟Estimation error+ϵqp+ϵpr⏟Approximation errorR(f_q) - R(f_r) \le \underbrace{O\left(\frac{|\mathcal{F}_q|_C}{n^{\alpha_{qp}}}\right) + O\left(\frac{|\mathcal{F}_p|_C}{n^{\alpha_{pr}}}\right)}_{\text{Estimation error}} + \underbrace{\epsilon_{qp} + \epsilon_{pr}}_{\text{Approximation error}}

    where ∣F∣C|\mathcal{F}|_C denotes function class capacity, ϵqp\epsilon_{qp} is the approximation error of Fq\mathcal{F}_q with respect to fpf_p, ϵpr\epsilon_{pr} is the approximation error of Fp\mathcal{F}_p with respect to frf_r, and αqp,αpr∈[1/2,1]\alpha_{qp}, \alpha_{pr} \in [1/2, 1] are sample informativeness rates.

    When generated samples exhibit excessive adaptability (too small Hinfo′(pds)H'_{info}(p_{ds})), QQ cannot absorb the knowledge, resulting in large training approximation error ϵqp\epsilon_{qp}, a slow learning rate αqp→1/2\alpha_{qp} \to 1/2, and high estimation error O(∣Fq∣Cnαqp)O\left(\frac{|\mathcal{F}_q|_C}{n^{\alpha_{qp}}}\right) (underfitting). When generated samples exhibit too little adaptability (too large Hinfo′(pds)H'_{info}(p_{ds})), ϵqp\epsilon_{qp} is small during training but test estimation error remains high (overfitting). Controlling sample adaptability between explicit lower and upper bounds λl<Hinfo′(pds)<λu\lambda_l < H'_{info}(p_{ds}) < \lambda_u minimizes both error terms simultaneously.

  6. Knowl 6 — AdaDFQ Alternating Training Procedure

    algorithm

    AdaDFQ executes data-free quantization by alternating between updating the generator parameters θg\theta_g and the quantized network parameters θq\theta_q across 400 epochs.

    Input: Pre-trained full-precision network PP, initial quantized network Q(⋅;θq)Q(\cdot; \theta_q), generator G(⋅;θg)G(\cdot; \theta_g), total epochs E=400E = 400, batch size B=16B = 16, hyperparameters αds=0.2\alpha_{ds} = 0.2, αas=0.1\alpha_{as} = 0.1, λl=0.1\lambda_l = 0.1, λu=0.8\lambda_u = 0.8, β=1.0\beta = 1.0, γ=1.0\gamma = 1.0.
    Output: Calibrated quantized network Q(⋅;θq)Q(\cdot; \theta_q).
    for epoch = 1 to EE do
        for each iteration step do
            Sample noise batch z∼N(0,I)z \sim \mathcal{N}(0, I) and random one-hot labels yy
            Generate fake samples x=G(z∣y)x = G(z | y)
            Compute logits zp=P(x)z_p = P(x) and zq=Q(x)z_q = Q(x)
            
            Compute pds=softmax(zp−zq)p_{ds} = \text{softmax}(z_p - z_q) and pas=softmax(zp+zq)p_{as} = \text{softmax}(z_p + z_q)
            Compute batch-normalized entropy Hinfo′(pds)H'_{info}(p_{ds})
            Compute losses Lbal=αdsHCE(pds,y)+αasHCE(pas,y)\mathcal{L}_{bal} = \alpha_{ds} H_{CE}(p_{ds}, y) + \alpha_{as} H_{CE}(p_{as}, y) and LBNS\mathcal{L}_{BNS}
            
            # Update Generator G (maximize zero-sum objective)
            LG=−max⁡(λl−Hinfo′(pds),0)−max⁡(Hinfo′(pds)−λu,0)−βLbal−γLBNS\mathcal{L}_G = -\max(\lambda_l - H'_{info}(p_{ds}), 0) - \max(H'_{info}(p_{ds}) - \lambda_u, 0) - \beta \mathcal{L}_{bal} - \gamma \mathcal{L}_{BNS}
            Update θg\theta_g using Adam optimizer to maximize LG\mathcal{L}_G
            
            # Regenerate samples and compute logits
            x=G(z∣y)x = G(z | y)
            zp=P(x)z_p = P(x), zq=Q(x)z_q = Q(x)
            pds=softmax(zp−zq)p_{ds} = \text{softmax}(z_p - z_q)
            Compute batch-normalized entropy Hinfo′(pds)H'_{info}(p_{ds})
            
            # Update Quantized Model Q (minimize adaptability)
            LQ=1−Hinfo′(pds)\mathcal{L}_Q = 1 - H'_{info}(p_{ds})
            Update θq\theta_q using SGD with Nesterov momentum to minimize LQ\mathcal{L}_Q
        end for
    end for
    return Q(⋅;θq)Q(\cdot; \theta_q)
  7. Knowl 7 — Performance Comparison of AdaDFQ Against State-of-the-Art DFQ Methods

    data/table

    The classification accuracy (%) of AdaDFQ compared against representative data-free quantization methods on CIFAR-10, CIFAR-100, and ImageNet across 3-bit, 4-bit, and 5-bit precision (denoted nwnan\text{w}n\text{a} for nn-bit weights and activations) is reported below:

    Dataset Model (FP) Bits ZAQ IntraQ ARC+AIT GDFQ AdaSG AdaDFQ (Ours)
    CIFAR-10 ResNet-20 3w3a - 77.07 - 75.11 84.14 84.89
    (93.89) 4w4a 92.13 91.49 90.49 90.11 92.10 92.31
    5w5a 93.36 - 92.98 93.38 93.76 93.81
    CIFAR-100 ResNet-20 3w3a - 48.25 41.34 47.61 52.76 52.74
    (70.33) 4w4a 60.42 64.98 61.05 63.75 66.42 66.81
    5w5a 68.70 - 68.40 67.52 69.42 69.93
    ImageNet ResNet-18 3w3a - - - 20.23 37.04 38.10
    (71.47) 4w4a 52.64 66.47 65.73 60.60 66.50 66.53
    5w5a 64.54 69.94 70.28 68.49 70.29 70.29
    MobileNetV2 3w3a - - - 1.46 26.90 28.99
    (73.03) 4w4a 0.10 65.10 66.47 59.43 65.15 65.41
    5w5a 62.35 71.28 71.96 68.11 71.61 71.61
    ResNet-50 3w3a - - - 0.31 16.98 17.63
    (77.73) 4w4a 53.02 - 68.27 54.16 68.58 68.38
    5w5a 73.38 - 76.00 71.63 76.03 76.03

    AdaDFQ substantially outperforms existing methods in low-bitwidth regimes (3w3a), showing absolute top-1 accuracy gains of up to 17.87%17.87\% over GDFQ on ResNet-18 and 27.53%27.53\% on MobileNetV2 on ImageNet, where previous methods suffered severe underfitting.

  8. Knowl 8 — Ablation Study of AdaDFQ Loss Components on ImageNet

    data/table

    An ablation study on ImageNet with ResNet-18 (full-precision accuracy: 71.47%71.47\%) evaluates the individual impact of the disagreement loss Lds\mathcal{L}_{ds}, agreement loss Las\mathcal{L}_{as}, boundary constraints (λl,λu)(\lambda_l, \lambda_u), and batch normalization statistics loss LBNS\mathcal{L}_{BNS}:

    Lds\mathcal{L}_{ds} Las\mathcal{L}_{as} λl,λu\lambda_l, \lambda_u LBNS\mathcal{L}_{BNS} 3w3a Accuracy (%) 5w5a Accuracy (%)
    ✓ ✓ ✓ 19.40 70.03
    ✓ ✓ ✓ 31.14 69.77
    ✓ ✓ 18.53 66.27
    ✓ ✓ ✓ 32.13 70.06
    ✓ ✓ ✓ 20.99 67.80
    ✓ ✓ ✓ ✓ 38.10 70.29

    Removing either Lds\mathcal{L}_{ds} or Las\mathcal{L}_{as} degrades 3w3a accuracy by 18.70%18.70\% and 6.96%6.96\% respectively, showing the necessity of joint disagreement-agreement balancing. Removing the explicit boundary constraints (λl,λu)(\lambda_l, \lambda_u) drops 3w3a performance from 38.10%38.10\% to 32.13%32.13\%, validating that bounding adaptability prevents over- and under-fitting.

  9. Knowl 9 — Sensitivity and Optimal Operating Region for Margin Bounds

    empirical result

    Grid search over lower bound λl∈{0.0,0.1,0.2,0.3,0.4,0.5}\lambda_l \in \{0.0, 0.1, 0.2, 0.3, 0.4, 0.5\} and upper bound λu∈{0.5,0.6,0.7,0.8,0.9,1.0}\lambda_u \in \{0.5, 0.6, 0.7, 0.8, 0.9, 1.0\} on 3-bit MobileNetV2 on ImageNet reveals that AdaDFQ achieves consistently high accuracy within a broad optimal parameter plateau defined by λl∈{0.0,0.1,0.2}\lambda_l \in \{0.0, 0.1, 0.2\} and λu∈{0.7,0.8,0.9}\lambda_u \in \{0.7, 0.8, 0.9\}. Setting λl=0.1\lambda_l = 0.1 and λu=0.8\lambda_u = 0.8 uniformly across all samples avoids sample-specific tuning while ensuring the normalized entropy Hinfo′(pds)H'_{info}(p_{ds}) remains in an intermediate zone that averts both underfitting (excessive disagreement) and overfitting (excessive agreement).

  10. Knowl 10 — Symmetric Linear Quantization Function

    equation

    Full-precision floating-point weights and activation values θ∈[θmin⁡,θmax⁡]\theta \in [\theta_{\min}, \theta_{\max}] are quantized to nn-bit fixed-point representations θq\theta_q using symmetric uniform linear quantization:

    θq=round((2n−1)⋅θ−θmin⁡θmax⁡−θmin⁡−2n−1)\theta_q = \text{round}\left( (2^n - 1) \cdot \frac{\theta - \theta_{\min}}{\theta_{\max} - \theta_{\min}} - 2^{n-1} \right)

    where round(⋅)\text{round}(\cdot) maps input real numbers to the nearest integer, and θmin⁡,θmax⁡\theta_{\min}, \theta_{\max} represent the dynamic range minimum and maximum values of the tensor θ\theta.

Coverage note — None was omitted; all key theoretical formulations, loss functions, algorithms, experimental benchmarks, ablation results, and sensitivity analyses from the paper are represented.

References

  1. 1.Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Zeroq: A novel zero shot quantization framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13169–13178, 2020.
  2. 2.Adrian Rivera Cardoso, Jacob Abernethy, He Wang, and Huan Xu. Competing against nash equilibria in adversarially changing zero-sum games. In International Conference on Machine Learning, pages 921–930. PMLR, 2019.
  3. 3.Kanghyun Choi, Deokki Hong, Noseong Park, Youngsok Kim, and Jinho Lee. Qimera: Data-free quantization with synthetic boundary supporting samples. Advances in Neural Information Processing Systems, 34, 2021.
  4. 4.Kanghyun Choi, Hye Yoon Lee, Deokki Hong, Joonsang Yu, Noseong Park, Youngsok Kim, and Jinho Lee. It’s all in the teacher: Zero-shot quantization brought closer to the teacher. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8311–8321, 2022.
  5. 5.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In NIPS, 2015.
  6. 6.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704–2713, 2018.
  7. 7.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  8. 8.A Krizhevsky. Learning multiple layers of features from tiny images. Master’s thesis, University of Tront, 2009.
  9. 9.Jae Hyun Lim and Jong Chul Ye. Geometric gan. arXiv preprint arXiv:1705.02894, 2017.
  10. 10.Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy. Fixed point quantization of deep convolutional networks. In International conference on machine learning, pages 2849–2858. PMLR, 2016.
  11. 11.Yuang Liu, Wei Zhang, and Jun Wang. Zero-shot adversarial quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1512–1521, 2021.
  12. 12.David Lopez-Paz, Leon Bottou, Bernhard Scholkopf, and Vladimir Vapnik. Unifying distillation and privileged information. arXiv preprint arXiv:1511.03643, 2015.
  13. 13.Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. Improved knowledge distillation via teacher assistant. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 5191–5198, 2020.
  14. 14.Yurii E Nesterov. A method for solving the convex programming problem with convergence rate o (1/kˆ 2). In Dokl. akad. nauk Sssr, volume 269, pages 543–547, 1983.
  15. 15.Augustus Odena, Christopher Olah, and Jonathon Shlens. Conditional image synthesis with auxiliary classifier gans. In International conference on machine learning, pages 2642–2651. PMLR, 2017.
  16. 16.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  17. 17.Biao Qian, Yang Wang, Richang Hong, and Meng Wang. Rethinking data-free quantization as a zero-sum game. In Proceedings of the AAAI conference on artificial intelligence, 2023.
  18. 18.Biao Qian, Yang Wang, Hongzhi Yin, Richang Hong, and Meng Wang. Switchable online knowledge distillation. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XI, pages 449–466, 2022.
  19. 19.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  20. 20.J v. Neumann. Zur theorie der gesellschaftsspiele. Mathematische annalen, 100(1):295–320, 1928.
  21. 21.Vladimir Vapnik. The nature of statistical learning theory. Springer science & business media, 1999.
  22. 22.Shoukai Xu, Haokun Li, Bohan Zhuang, Jing Liu, Jiezhang Cao, Chuangrun Liang, and Mingkui Tan. Generative lowbitwidth data free quantization. In European Conference on Computer Vision, pages 1–17. Springer, 2020.
  23. 23.Xiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li, Fengwei Yu, and Xianglong Liu. Diversifying sample generation for accurate data-free quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15658–15667, 2021.
  24. 24.Yunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu, Baochang Zhang, Yonghong Tian, and Rongrong Ji. Intraq: Learning synthetic images with intra-class heterogeneity for zero-shot network quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12339–12348, 2022.
  25. 25.Baozhou Zhu, Peter Hofstee, Johan Peltenburg, Jinho Lee, and Zaid Alars. Autorecon: Neural architecture search-based reconstruction for data-free. In International Joint Conference on Artificial Intelligence, 2021.

Citation

MLA
Qian, B., et al. “Adaptive Data-Free Quantization”. arXiv, 2023, http://arxiv.org/abs/2303.06869v3.
APA
Qian, B., Wang, Y., Hong, R., & Wang, M. (2023). Adaptive Data-Free Quantization. arXiv. http://arxiv.org/abs/2303.06869v3
Chicago
Qian, B., Y. Wang, R. Hong, and M. Wang. 2023. “Adaptive Data-Free Quantization”. arXiv. http://arxiv.org/abs/2303.06869v3.
Harvard
Qian, B. et al. (2023) “Adaptive Data-Free Quantization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.06869v3.
Vancouver
1. Qian B, Wang Y, Hong R, Wang M (2023) Adaptive Data-Free Quantization. arXiv

BibTeX

@article{qian2023adaptive,
  title = {Adaptive Data-Free Quantization},
  author = {Qian, Biao and Wang, Yang and Hong, Richang and Wang, Meng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.06869v3},
  eprint = {2303.06869}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE