Up to 100x Faster Data-Free Knowledge Distillation

Gongfan FangKanya MoXinchao WangJie SongShitao BeiHaofei ZhangMingli Song

article2022AAAI110 citations

Accelerates data-free knowledge distillation by up to two orders of magnitude by learning a meta-synthesizer that extracts reusable domain features to generate synthetic training data in only a few optimization steps.

Listen

Deploying artificial intelligence models on resource-limited devices requires model compression techniques such as knowledge distillation, which transfers knowledge from a large teacher model to a compact student model. However, privacy concerns, proprietary restrictions, and data protection regulations frequently prevent developers from accessing the original training data. While data-free knowledge distillation creates synthetic data directly from the teacher model to circumvent this issue, current approaches require thousands of optimization steps per image batch or hundreds of specialized generators. This extreme computational inefficiency creates severe processing bottlenecks and high computing costs, making data-free compression impractical for complex, large-scale applications.

The article demonstrates an acceleration framework called FastDFKD, which drastically reduces synthetic data generation time without sacrificing model accuracy. The method rests on the insight that data instances within a single domain naturally share common underlying features, such as basic textures or structural components. Instead of synthesizing every data sample from scratch in isolation, the framework learns a meta-generator that captures these shared features and quickly adapts them in only a few steps to produce diverse training samples. The researchers evaluated this approach across standard image classification benchmarks—CIFAR-10, CIFAR-100, and ImageNet—as well as the NYUv2 semantic segmentation dataset, comparing synthesis speed and model accuracy against leading data-free methods.

The findings confirm that FastDFKD reduces data synthesis time by a factor of 10 to over 100 while maintaining competitive accuracy across all tested benchmarks. On CIFAR benchmarks, the method achieved up to 120-fold to 650-fold average speedups over conventional batch-optimization techniques. On the large-scale ImageNet dataset, FastDFKD synthesized a complete transfer dataset of 140,000 images in just 6.28 single-GPU hours—approximately 48 times faster than previous inversion baselines taking 166 hours and drastically faster than multi-generator approaches requiring roughly 300 hours—while achieving a 68.61% accuracy on ResNet-18. In semantic segmentation on NYUv2, it generated the training set in 0.82 hours compared to 4.0 to 6.0 hours for existing generative methods, attaining a superior mean intersection-over-union score of 0.366. Ablation experiments also showed that competing methods fail completely when restricted to few-step synthesis, whereas the proposed common-feature adaptation remains effective with as few as two to five optimization steps.

These results demonstrate that organizations can reliably compress complex deep learning models in secure, privacy-constrained environments in hours rather than weeks of compute time. By significantly lowering hardware usage and associated cloud costs, the framework removes the primary adoption barrier for data-free model deployment across vision and segmentation pipelines.

Organizations operating under data confidentiality constraints should consider adopting meta-learning-based feature reuse for model compression pipelines. Teams seeking to apply the method should conduct initial pilots comparing few-step configurations to balance data generation speed against required model accuracy. While the experimental evidence provides high confidence for visual classification and segmentation tasks, practitioners should exercise caution and conduct further validation before applying the method to non-visual data domains, such as tabular or natural language processing pipelines, which were not tested in the article.

  • Paper: Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion, Hongxu Yin et al. (2020). DeepInversion establishes foundational techniques for data-free knowledge distillation by inverting teacher network statistics, directly motivating FastDFKD's goal to overcome its synthesis inefficiency.
  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal paper introduces the core principles of knowledge distillation that all subsequent data-free and accelerated distillation methods build upon.
  • Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). This survey provides a comprehensive taxonomy of teacher-student compression algorithms and distillation schemes necessary for understanding data-free extensions.
  • Paper: Relational Knowledge Distillation, Wonpyo Park et al. (2019). Relational knowledge distillation formulates structural knowledge transfer across instances, informing feature-level relationships and representations leveraged in distillation pipelines.
  • Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). This study analyzes capacity mismatches and optimization bottlenecks in knowledge distillation that motivate specialized synthesis and transfer strategies.
  • Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). This paper analyzes the challenges of transferring knowledge across large model capacity gaps, providing valuable context for student training dynamics in data-free settings.

No sufficiently relevant recommendations were found.

Cover for Up to 100x Faster Data-Free Knowledge Distillation

Abstract

Data-free knowledge distillation (DFKD) has recently been attracting increasing attention from research communities, attributed to its capability to compress a model only using synthetic data. Despite the encouraging results achieved, state-of-the-art DFKD methods still suffer from the inefficiency of data synthesis, making the data-free training process extremely time-consuming and thus inapplicable for large-scale tasks. In this work, we introduce an efficacious scheme, termed as FastDFKD, that allows us to accelerate DFKD by a factor of orders of magnitude. At the heart of our approach is a novel strategy to reuse the shared common features in training data so as to synthesize different data instances. Unlike prior methods that optimize a set of data independently, we propose to learn a meta-synthesizer that seeks common features as the initialization for the fast data synthesis. As a result, FastDFKD achieves data synthesis within only a few steps, significantly enhancing the efficiency of data-free training. Experiments over CIFAR, NYUv2, and ImageNet demonstrate that the proposed FastDFKD achieves 10× and even 100× acceleration while preserving performances on par with state of the art. Code is available at https://github.com/zju-vipa/Fast-Datafree.

Table of Contents

  • Introduction
  • Related Works
  • Method
  • Problem Setup
  • Fast Data-Free Knowledge Distillation
  • Experiments
  • Experimental Settings
  • Results on Classification
  • Results on Segmentation
  • Quantitative Analysis
  • Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Fast Data-Free Knowledge Distillation (FastDFKD) Framework

    model/method

    Data-free knowledge distillation (DFKD) transfers knowledge from a pre-trained teacher model ft(x;θt)f_t(x; \theta_t) to a student model fs(x;θs)f_s(x; \theta_s) without access to original training data by inverting the teacher network to synthesize a transfer dataset D′={x1∗,x2∗,…,xN∗}D' = \{x_1^*, x_2^*, \dots, x_N^*\}. Conventional non-generative methods optimize each batch independently over thousands of gradient iterations, which makes synthesis computationally prohibitive for large datasets.

    FastDFKD formulates data synthesis as a meta-learning task where a lightweight generator G(z;θ)G(z; \theta) is trained to capture common, domain-shared feature representations (such as repeating textures or edge structures). Instead of synthesizing every sample from scratch or attempting to generate the entire distribution with a single fixed network, FastDFKD maintains meta-parameters (z^,θ^)(\hat{z}, \hat{\theta}). For each synthesis task (governed by an inversion loss Li\mathcal{L}_i), the generator quickly adapts from (z^,θ^)(\hat{z}, \hat{\theta}) in only kk gradient descent steps (e.g., k∈{2,5,10,50}k \in \{2, 5, 10, 50\}) to generate synthetic instances xi∗=G(zi∗;θi∗)x_i^* = G(z_i^*; \theta_i^*). The results of this adaptation are used in an outer loop to update the meta-initialization (z^,θ^)(\hat{z}, \hat{\theta}), enabling orders-of-magnitude acceleration in DFKD.

  2. Knowl 2 — Meta-Learning Formulation for Fast Data Synthesis

    equation

    To train a meta-generator G(z;θ)G(z; \theta) capable of synthesizing diverse samples within kk gradient steps, FastDFKD defines the meta-learning optimization objective over a set of synthesis losses {L1,L2,…,LN}\{\mathcal{L}_1, \mathcal{L}_2, \dots, \mathcal{L}_N\} as:

    min⁡z^,θ^1N∑i=1NLi(G(ULik(z^,θ^)))\min_{\hat{z}, \hat{\theta}} \frac{1}{N} \sum_{i=1}^N \mathcal{L}_i\left(G\left(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta})\right)\right)

    where z^\hat{z} is the meta latent code, θ^\hat{\theta} denotes the meta-parameters of the generator GG, and ULik(z^,θ^)=(zik,θik)U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}) = (z_i^k, \theta_i^k) denotes the inner-loop operator unrolled across kk gradient descent steps with inner learning rate α>0\alpha > 0:

    zi0=z^,θi0=θ^z_i^0 = \hat{z}, \quad \theta_i^0 = \hat{\theta}

    zij=zij−1−α∇zij−1Li(G(zij−1;θij−1))for j=1,…,kz_i^j = z_i^{j-1} - \alpha \nabla_{z_i^{j-1}} \mathcal{L}_i\left(G(z_i^{j-1}; \theta_i^{j-1})\right) \quad \text{for } j = 1, \dots, k

    θij=θij−1−α∇θij−1Li(G(zij−1;θij−1))for j=1,…,k\theta_i^j = \theta_i^{j-1} - \alpha \nabla_{\theta_i^{j-1}} \mathcal{L}_i\left(G(z_i^{j-1}; \theta_i^{j-1})\right) \quad \text{for } j = 1, \dots, k

    The inner loop adapts the meta-features specifically to minimize the inversion loss Li\mathcal{L}_i, while the outer loop optimizes (z^,θ^)(\hat{z}, \hat{\theta}) so that diverse target instances can be reached within few optimization steps.

  3. Knowl 3 — First-Order and Reptile Meta-Update for Generator Optimization

    equation

    The gradient of the outer-loop meta-objective with respect to the meta-parameters θ^\hat{\theta} involves higher-order gradient computations through the kk-step adaptation unrolling:

    gθ^=∇θ^Li(G(ULik(z^,θ^)))=(ULik)′(θ^)G′(θi∗)Li′(xi∗)g_{\hat{\theta}} = \nabla_{\hat{\theta}} \mathcal{L}_i\left(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}))\right) = \left(U_{\mathcal{L}_i}^k\right)'(\hat{\theta}) G'(\theta_i^*) \mathcal{L}_i'(x_i^*)

    where θi∗=θik\theta_i^* = \theta_i^k and xi∗=G(zi∗;θi∗)x_i^* = G(z_i^*; \theta_i^*). To eliminate the inefficiency of computing second-order derivatives, a first-order approximation replaces (ULik)′(θ^)\left(U_{\mathcal{L}_i}^k\right)'(\hat{\theta}) with the identity mapping:

    ∇θ^Li(G(ULik(z^,θ^)))≈G′(θi∗)Li′(xi∗)\nabla_{\hat{\theta}} \mathcal{L}_i\left(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}))\right) \approx G'(\theta_i^*) \mathcal{L}_i'(x_i^*)

    Applying the Reptile approximation further simplifies the gradient update to the parameter difference between the adapted parameters θi∗\theta_i^* and the meta-initialization θ^\hat{\theta}:

    ∇θ^Li(G(ULik(z^,θ^)))≈θ^−θi∗\nabla_{\hat{\theta}} \mathcal{L}_i\left(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}))\right) \approx \hat{\theta} - \theta_i^*

    The meta-parameters and latent codes are then updated across tasks with meta learning rate η>0\eta > 0:

    θ^←θ^−η∑i∇θ^Li(G(ULik(z^,θ^)))\hat{\theta} \leftarrow \hat{\theta} - \eta \sum_i \nabla_{\hat{\theta}} \mathcal{L}_i\left(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}))\right)

    z^←z^−η∑i∇z^Li(G(ULik(z^,θ^)))\hat{z} \leftarrow \hat{z} - \eta \sum_i \nabla_{\hat{z}} \mathcal{L}_i\left(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta}))\right)

  4. Knowl 4 — Dynamic Momentum Feature Regularization Loss

    equation

    Data synthesis in FastDFKD employs an inversion loss composed of classification loss Lcls\mathcal{L}_{cls}, adversarial loss Ladv\mathcal{L}_{adv}, and feature regularization loss Lfeat\mathcal{L}_{feat}:

    L(x)=Lcls(x)+Ladv(x)+Lfeat(x)\mathcal{L}(x) = \mathcal{L}_{cls}(x) + \mathcal{L}_{adv}(x) + \mathcal{L}_{feat}(x)

    where Lcls(x)=CE(ft(x),c)\mathcal{L}_{cls}(x) = \text{CE}(f_t(x), c) for pseudo-label class cc, and Ladv(x)=−JSD(ft(x)/τ∥fs(x)/τ)\mathcal{L}_{adv}(x) = -\text{JSD}(f_t(x)/\tau \parallel f_s(x)/\tau) aligns or diversifies representations using Jensen-Shannon Divergence with temperature τ\tau.

    Because the pre-trained batch normalization statistics (μbnl,σbn2,l)(\mu_{bn}^l, \sigma_{bn}^{2, l}) at layer ll are static, standard feature regularization does not provide dynamic targets across synthesis steps. FastDFKD modifies feature regularization by maintaining running accumulative mean μal\mu_a^l and variance σal\sigma_a^l of already synthesized data:

    Lfeat(x)=∑l(∥(1−m)⋅μal+m⋅μfeatl(x)−μbnl∥1+∥(1−m)⋅σal+m⋅σfeatl(x)−σbn2,l∥1)\mathcal{L}_{feat}(x) = \sum_l \left( \left\| (1 - m) \cdot \mu_a^l + m \cdot \mu_{feat}^l(x) - \mu_{bn}^l \right\|_1 + \left\| (1 - m) \cdot \sigma_a^l + m \cdot \sigma_{feat}^l(x) - \sigma_{bn}^{2, l} \right\|_1 \right)

    where m∈(0,1)m \in (0, 1) is the momentum parameter, and μfeatl(x),σfeatl(x)\mu_{feat}^l(x), \sigma_{feat}^l(x) are the mean and standard deviation of feature activations for input xx at layer ll. After each synthesis iteration producing x∗x^*, the accumulative variables are updated via:

    μal←(1−m)⋅μal+m⋅μfeatl(x∗)\mu_a^l \leftarrow (1 - m) \cdot \mu_a^l + m \cdot \mu_{feat}^l(x^*)

    σal←(1−m)⋅σal+m⋅σfeatl(x∗)\sigma_a^l \leftarrow (1 - m) \cdot \sigma_a^l + m \cdot \sigma_{feat}^l(x^*)

    This continuous update alters Lfeat\mathcal{L}_{feat} dynamically, providing varying learning targets required for meta-learning.

  5. Knowl 5 — FastDFKD Training Procedure

    algorithm

    FastDFKD alternates between kk-step inner-loop generator adaptation for data synthesis, outer-loop meta-updates to store shared features, and knowledge distillation steps on the student model using the synthesized dataset D′D'.

    Input: Pre-trained teacher model ft(x;θt)f_t(x; \theta_t), student model fs(x;θs)f_s(x; \theta_s), adaptation steps kk, inner learning rate α\alpha, meta learning rate η\eta, distillation steps tt, distillation learning rate γ\gamma.
    Output: Optimized student model fs(x;θs)f_s(x; \theta_s).
    Randomly initialize generator parameters and latent code (z^,θ^)(\hat{z}, \hat{\theta})
    Initialize empty synthetic dataset D′←∅D' \leftarrow \emptyset
    for each synthesis loss Li\mathcal{L}_i do
        Periodically re-initialize z^\hat{z} for sample diversity
        Initialize inner-loop state: z←z^z \leftarrow \hat{z}, θ←θ^\theta \leftarrow \hat{\theta}
        for step =1= 1 to kk do
            $(z, \theta) \leftarrow (z, \theta) - \alpha \nabla_{(z, \theta)} \mathcal{L}_i(G(z; \theta))
        end for
        Generate synthetic sample x∗←G(z;θ)x^* \leftarrow G(z; \theta)
        Update synthetic set: D′←D′∪{x∗}D' \leftarrow D' \cup \{x^*\}
        Update meta-parameters: (z^,θ^)←(z^,θ^)−η∇(z^,θ^)Li(G(ULik(z^,θ^)))(\hat{z}, \hat{\theta}) \leftarrow (\hat{z}, \hat{\theta}) - \eta \nabla_{(\hat{z}, \hat{\theta})} \mathcal{L}_i(G(U_{\mathcal{L}_i}^k(\hat{z}, \hat{\theta})))
        for step =1= 1 to tt do
            Sample mini-batch BB uniformly from D′D'
            θs←θs−γ∑x∈B∇θsKL(ft(x)∥fs(x))\theta_s \leftarrow \theta_s - \gamma \sum_{x \in B} \nabla_{\theta_s} \text{KL}(f_t(x) \parallel f_s(x))
        end for
    end for
  6. Knowl 6 — Student Classification Accuracy and Synthesis Speedup on CIFAR-10 and CIFAR-100

    data/table

    FastDFKD with k∈{2,5,10}k \in \{2, 5, 10\} inner steps (denoted as Fast2\text{Fast}_2, Fast5\text{Fast}_5, Fast10\text{Fast}_{10}) was evaluated on CIFAR-10 and CIFAR-100 (50,000 synthetic images, batch size matching prior work) across five teacher-student pairs. Results report top-1 student classification accuracy (%) and data synthesis GPU hours (h) on a single GPU.

    Dataset Method ResNet-34 VGG-11 WRN40-2 WRN40-2 WRN40-2 Average
    ResNet-18 ResNet-18 WRN16-1 WRN40-1 WRN16-2 Speed Up
    CIFAR-10 DeepInv2k_{2k} 93.26 (42.1h) 90.36 (20.2h) 83.04 (16.9h) 86.85 (21.9h) 89.72 (18.2h) 1.0×\times
    CMI500_{500} 94.84 (19.0h) 91.13 (11.6h) 90.01 (13.3h) 92.78 (14.1h) 92.52 (13.6h) 1.6×\times
    DAFL 92.22 (2.73h) 81.10 (0.73h) 65.71 (1.73h) 81.33 (1.53h) 81.55 (1.60h) 15.7×\times
    ZSKT 93.32 (1.67h) 89.46 (0.33h) 83.74 (0.87h) 86.07 (0.87h) 89.66 (0.87h) 30.4×\times
    DFQ 94.61 (8.79h) 90.84 (1.50h) 86.14 (0.75h) 91.69 (0.75h) 92.01 (0.75h) 18.9×\times
    Fast2\text{Fast}_2 92.62 (0.06h) 84.67 (0.03h) 88.36 (0.03h) 89.56 (0.03h) 89.68 (0.03h) 655.0×\times
    Fast5\text{Fast}_5 93.63 (0.14h) 89.94 (0.08h) 88.90 (0.08h) 92.04 (0.09h) 91.96 (0.08h) 247.1×\times
    Fast10\text{Fast}_{10} 94.05 (0.28h) 90.53 (0.15h) 89.29 (0.15h) 92.51 (0.17h) 92.45 (0.17h) 126.7×\times
    CIFAR-100 DeepInv2k_{2k} 61.32 (42.1h) 54.13 (20.1h) 53.77 (17.0h) 61.33 (21.9h) 61.34 (18.2h) 1.0×\times
    CMI500_{500} 77.04 (19.2h) 70.56 (11.6h) 57.91 (13.3h) 68.88 (14.2h) 68.75 (13.9h) 1.6×\times
    DAFL 74.47 (2.73h) 54.16 (0.73h) 20.88 (1.67h) 42.83 (1.80h) 43.70 (1.73h) 15.2×\times
    ZSKT 67.74 (1.67h) 54.31 (0.33h) 36.66 (0.87h) 53.60 (0.87h) 54.59 (0.87h) 30.4×\times
    DFQ 77.01 (8.79h) 66.21 (1.54h) 51.27 (0.75h) 54.43 (0.75h) 64.79 (0.75h) 18.8×\times
    Fast2\text{Fast}_2 69.76 (0.06h) 62.83 (0.03h) 41.77 (0.03h) 53.15 (0.04h) 57.08 (0.04h) 588.2×\times
    Fast5\text{Fast}_5 72.82 (0.14h) 65.28 (0.08h) 52.90 (0.07h) 61.80 (0.09h) 63.83 (0.08h) 253.1×\times
    Fast10\text{Fast}_{10} 74.34 (0.27h) 67.44 (0.16h) 54.02 (0.16h) 63.91 (0.17h) 65.12 (0.17h) 124.7×\times

    Compared to non-generative DeepInv (2,0002,000 optimization steps per batch), Fast5\text{Fast}_5 synthesizes CIFAR-10 data in 0.08–0.140.08\text{--}0.14 hours instead of 16.9–42.116.9\text{--}42.1 hours, achieving over 247×247\times speedup while matching or outperforming student accuracy. It also runs 10×10\times to 20×20\times faster than generator-based methods (DAFL, ZSKT, DFQ) while avoiding generator capacity bottlenecks on CIFAR-100.

  7. Knowl 7 — Data-Free Knowledge Distillation on 224x224 ImageNet

    data/table

    FastDFKD with k=50k=50 adaptation steps (Fast50\text{Fast}_{50}) was evaluated on full ImageNet (224×224224 \times 224 resolution, 1,000 categories) using an off-the-shelf ResNet-50 teacher model to distill into ResNet-50, ResNet-18, and MobileNetV2 students.

    Method Data Amount Syn. Time Speed Up ResNet-50 →\rightarrow ResNet-50 →\rightarrow ResNet-50 →\rightarrow
    ResNet-50 ResNet-18 MobileNetv2
    Scratch 1.3M - - 75.45 68.45 70.01
    Places365+KD 1.8M - - 55.74 45.53 39.89
    Generative DFD - ∼\sim300h 1×\times 69.75 54.66 43.15
    DeepInv2k_{2k} 140k 166h 1.8×\times 68.00 - -
    Fast50\text{Fast}_{50} 140k 6.28h 47.8×\times 68.61 53.45 43.02

    While Generative DFD requires training 1,000 separate category generators (taking ∼300\sim 300 GPU hours) and DeepInv2k_{2k} requires 166 GPU hours to synthesize 140k images, Fast50\text{Fast}_{50} produces 140k synthetic images in 6.28 GPU hours (a 47.8×47.8\times speedup relative to Generative DFD's baseline and >26×>26\times relative to DeepInv) while achieving 68.61% accuracy on ResNet-50 student, on par with both data-free baselines.

  8. Knowl 8 — Data-Free Semantic Segmentation on NYUv2

    data/table

    Data-free knowledge distillation was evaluated on the NYUv2 semantic segmentation dataset (13 classes) using a pre-trained Deeplabv3-ResNet50 teacher and a freshly initialized Deeplabv3-MobileNetV2 student. Synthesis used the feature regularization loss and adversarial loss.

    Method Data Amount Syn. Time mIoU
    Teacher 1,449 (NYUv2) - 0.519
    Student 1,449 (NYUv2) - 0.375
    KD 1,449 (NYUv2) - 0.380
    DFND 14 M (ImageNet) - 0.378
    DFAD 960k (GAN) 6.0h 0.364
    DAFL 960k (GAN) 3.99h 0.105
    Fast10\text{Fast}_{10} 17K (Synthetic) 0.82h 0.366

    Fast10\text{Fast}_{10} synthesized 17k images in 0.82 GPU hours and attained a student mIoU of 0.366. This compares favorably to DFAD (0.364 mIoU, 6.0 hours) and DAFL (0.105 mIoU, 3.99 hours), representing more than a 7.3×7.3\times reduction in synthesis time relative to DFAD.

  9. Knowl 9 — Robustness of FastDFKD Under Low Optimization Step Budgets

    data/table

    To evaluate synthesis under extreme few-step constraints, FastDFKD was compared to step-reduced versions of non-generative baselines (DeepInversion and CMI) on CIFAR-100 across 2, 5, and 10 optimization steps. The teacher is WRN40-2 (accuracy 75.83%) and the student is WRN16-1 (accuracy 65.31%).

    Method 2 steps 5 steps 10 steps
    Teacher 75.83 75.83 75.83
    Student 65.31 65.31 65.31
    DeepInv 2.61 (0.03h) 4.84 (0.06h) 6.77 (0.11h)
    CMI 14.62 (0.05h) 20.08 (0.12h) 32.43 (0.22h)
    FastDFKD (Ours) 41.77 (0.03h) 52.90 (0.07h) 54.02 (0.16h)

    When limited to 2 steps, DeepInversion and CMI fail completely (2.61% and 14.62% student accuracy), showing that non-generative pixel optimization cannot converge without hundreds of iterations. In contrast, FastDFKD achieves 41.77% in 2 steps, 52.90% in 5 steps, and 54.02% in 10 steps due to effective meta-initialization of common features.

  10. Knowl 10 — Ablation of Synthesis Strategies and Momentum Feature Regularization

    data/table

    An ablation study comparing data representations (pixel optimization vs. generator network GG), feature reuse paradigms (no reuse, sequential reuse from previous batch, meta-learning reuse), and the dynamic momentum feature regularization loss (MMT) was conducted on CIFAR-10 and CIFAR-100 using a WRN40-2 teacher and WRN16-1 student.

    Settings CIFAR-10 CIFAR-100
    No Reuse + Pixel 44.74 3.11
    No Reuse + GAN 87.33 35.48
    Seq. + Pixel 71.97 10.93
    Seq. + GAN 90.91 59.38
    Meta + GAN 91.79 61.43
    Meta + GAN + MMT 91.96 63.83

    Direct pixel optimization with no reuse achieves only 3.11% on CIFAR-100 under few-step optimization. Sequential feature reuse using a GAN achieves 59.38%, but can suffer from task-specific bias from the preceding instance. The proposed meta-learning feature reuse with a GAN increases accuracy to 61.43%, and adding momentum feature regularization (MMT) to dynamically vary meta-learning targets further raises accuracy to 63.83%.

Coverage note — None. All core methodological formulations, algorithms, mathematical details, and experimental findings across CIFAR, ImageNet, and NYUv2 were covered.

References

  1. 1.Chawla, A.; Yin, H.; Molchanov, P.; and Alvarez, J. 2021. Data-Free Knowledge Distillation for Object Detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 3289–3298.
  2. 2.Chen, H.; Guo, T.; Xu, C.; Li, W.; Xu, C.; Xu, C.; and Wang, Y. 2021. Learning Student Networks in the Wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6428–6437.
  3. 3.Chen, H.; Wang, Y.; Xu, C.; Yang, Z.; Liu, C.; Shi, B.; Xu, C.; Xu, C.; and Tian, Q. 2019. Data-free learning of student networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 3514–3522.
  4. 4.Chen, L.-C.; Papandreou, G.; Schroff, F.; and Adam, H. 2017. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 .
  5. 5.Choi, Y.; Choi, J.; El-Khamy, M.; and Lee, J. 2020. Data-free network quantization with adversarial knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 710–711.
  6. 6.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee.
  7. 7.Deng, X.; and Zhang, Z. 2021. Graph-Free Knowledge Distillation for Graph Neural Networks. arXiv preprint arXiv:2105.07519 .
  8. 8.Fang, G.; Bao, Y.; Song, J.; Wang, X.; Xie, D.; Shen, C.; and Song, M. 2021a. Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data. In Thirty-Fifth Conference on Neural Information Processing Systems.
  9. 9.Fang, G.; Song, J.; Shen, C.; Wang, X.; Chen, D.; and Song, M. 2019. Data-free adversarial distillation. arXiv preprint arXiv:1912.11006 .
  10. 10.Fang, G.; Song, J.; Wang, X.; Shen, C.; Wang, X.; and Song, M. 2021b. Contrastive Model Inversion for Data-Free Knowledge Distillation. arXiv preprint arXiv:2105.08584 .
  11. 11.Finn, C.; Abbeel, P.; and Levine, S. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, 1126–1135. PMLR.
  12. 12.Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 .
  13. 13.Hospedales, T.; Antoniou, A.; Micaelli, P.; and Storkey, A. 2020. Meta-learning in neural networks: A survey. arXiv preprint arXiv:2004.05439 .
  14. 14.Kolesnikov, A.; Beyer, L.; Zhai, X.; Puigcerver, J.; Yung, J.; Gelly, S.; and Houlsby, N. 2020. Big Transfer (BiT): General Visual Representation Learning.
  15. 15.Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images. technical report .
  16. 16.Lopes, R. G.; Fenu, S.; and Starner, T. 2017. Data-free knowledge distillation for deep neural networks. arXiv preprint arXiv:1710.07535 .
  17. 17.Luo, L.; Sandler, M.; Lin, Z.; Zhmoginov, A.; and Howard, A. 2020. Large-scale generative data-free distillation. arXiv preprint arXiv:2012.05578 .
  18. 18.Ma, X.; Shen, Y.; Fang, G.; Chen, C.; Jia, C.; and Lu, W. 2020. Adversarial Self-Supervised Data-Free Distillation for Text Classification. arXiv preprint arXiv:2010.04883 .
  19. 19.Micaelli, P.; and Storkey, A. 2019. Zero-shot knowledge transfer via adversarial belief matching. arXiv preprint arXiv:1905.09768 .
  20. 20.Nathan Silberman, Derek Hoiem, P. K.; and Fergus, R. 2012. Indoor Segmentation and Support Inference from RGBD Images. In ECCV.
  21. 21.Nichol, A.; Achiam, J.; and Schulman, J. 2018. On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999 .
  22. 22.Shen, C.; Wang, X.; Song, J.; Sun, L.; and Song, M. 2019. Amalgamating knowledge towards comprehensive classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 3068–3075.
  23. 23.Yang, Y.; Qiu, J.; Song, M.; Tao, D.; and Wang, X. 2020. Distilling knowledge from graph convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7074–7083.
  24. 24.Ye, J.; Ji, Y.; Wang, X.; Ou, K.; Tao, D.; and Song, M. 2019. Student becoming the master: Knowledge amalgamation for joint scene parsing, depth estimation, and more. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2829–2838.
  25. 25.Yin, H.; Molchanov, P.; Li, Z.; Alvarez, J. M.; Mallya, A.; Hoiem, D.; Jha, N. K.; and Kautz, J. 2019. Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion. arXiv preprint arXiv:1912.08795 .
  26. 26.Zhu, Z.; Hong, J.; and Zhou, J. 2021. Data-Free Knowledge Distillation for Heterogeneous Federated Learning. arXiv preprint arXiv:2105.10056 .

Citation

MLA
Fang, G., et al. “Up to 100$\times$ Faster Data-free Knowledge Distillation”. arXiv, 2021, http://arxiv.org/abs/2112.06253v2.
APA
Fang, G., Mo, K., Wang, X., Song, J., Bei, S., Zhang, H., & Song, M. (2021). Up to 100$\times$ Faster Data-free Knowledge Distillation. arXiv. http://arxiv.org/abs/2112.06253v2
Chicago
Fang, G., K. Mo, X. Wang, et al. 2021. “Up to 100$\times$ Faster Data-free Knowledge Distillation”. arXiv. http://arxiv.org/abs/2112.06253v2.
Harvard
Fang, G. et al. (2021) “Up to 100$\times$ Faster Data-free Knowledge Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.06253v2.
Vancouver
1. Fang G, Mo K, Wang X, Song J, Bei S, Zhang H, Song M (2021) Up to 100$\times$ Faster Data-free Knowledge Distillation. arXiv

BibTeX

@article{fang2021100,
  title = {Up to 100$\textbackslash\{\}times$ Faster Data-free Knowledge Distillation},
  author = {Fang, Gongfan and Mo, Kanya and Wang, Xinchao and Song, Jie and Bei, Shitao and Zhang, Haofei and Song, Mingli},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.06253v2},
  eprint = {2112.06253}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF