Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay

Kuluhan BiniciShivam AggarwalNam Trung PhamKarianto LemanTulika Mitra

article2022AAAI59 citations

Presents a data-free knowledge distillation framework that uses a variational autoencoder to generate replay samples, preventing student accuracy degradation over training epochs without storing past synthetic data in memory.

Listen

Deploying advanced artificial intelligence models to edge devices requires compressing large, high-performing networks into smaller, efficient architectures. Knowledge distillation is the standard technique for this task, but conventional approaches require access to the original training data, which is frequently unavailable due to privacy restrictions, proprietary barriers, or extreme dataset sizes. While data-free distillation methods address this by generating synthetic training samples, they suffer from a severe operational flaw: the distribution of synthetic samples shifts over time, causing the compact model to forget previously learned information and experience rapid performance degradation. Because real validation data is absent during deployment, practitioners cannot monitor model accuracy in real time to capture peak performance before degradation occurs. Alternative solutions attempt to retain knowledge by caching past synthetic samples in memory buffers, but this introduces massive memory footprints, extends runtime substantially, and undermines data privacy.

The article demonstrates a novel framework called Pseudo Replay Enhanced Data-Free Knowledge Distillation (PRE-DFKD) that stabilizes the training process and eliminates memory overhead without requiring access to real validation data. The objective is to enable compact models to continuously retain prior knowledge and maintain high, predictable accuracy across arbitrary training durations while avoiding physical sample storage.

To achieve this, the authors designed a dual-generator architecture. One generator continuously synthesizes novel samples to close the immediate information gap between the teacher and student networks, while a secondary generative model—specifically a Variational Autoencoder (VAE)—learns the historical distribution of synthetic samples and generates "memory samples" for rehearsal. Crucially, because standard VAE loss functions fail on synthetic images where small pixel perturbations can alter core content, the approach integrates a synthetic-data-aware reconstruction loss that forces reconstructed samples to match feature representations inside the teacher network. The method also introduces an inference technique to ensure balanced class representation across memory batches. The framework was evaluated across four standard image classification benchmarks—MNIST, CIFAR-10, CIFAR-100, and Tiny ImageNet—against established replay-free and sample-storing distillation baselines.

The experimental findings show that PRE-DFKD outperforms prior data-free distillation approaches in both reliability and resource efficiency. First, the framework significantly increased expected model accuracy, delivering up to a 26.8% increase in average student accuracy compared to replay-free baselines, while substantially narrowing the variance across training epochs. Second, PRE-DFKD achieved a constant, minimal memory overhead of just 2.1 megabytes across all benchmarks, reducing the memory footprint from several hundred megabytes or gigabytes required by sample-storing methods. Third, the framework matched or approached the peak distillation accuracy of complex baselines (e.g., reaching 94.1% on CIFAR-10 and 70.2% average accuracy on CIFAR-100) while closely tracking the performance of models trained on real data. Ablation analyses confirmed that the synthetic-aware reconstruction loss and class-balancing mechanisms are vital; removing either component caused significant degradation in training stability.

These results demonstrate that organizations can reliably compress deep neural networks without needing original datasets or intermediate validation checks. By eliminating the risk of catastrophic forgetting, engineering teams can safely terminate distillation jobs after a set duration without worrying about model collapse. Furthermore, eliminating the need to store raw synthetic images directly mitigates data leak and compliance risks, significantly lowers high-performance computing hardware costs, and prevents the memory scaling bottlenecks associated with complex datasets.

For technical leaders seeking to compress models in privacy-sensitive or resource-constrained settings, adopting generative pseudo replay offers a viable, production-friendly path. Teams transitioning from sample-storing pipelines should consider replacing physical replay buffers with parameterized generative replay to reduce infrastructure footprint and memory bottlenecks. The primary boundary condition noted is that the approach still relies on dataset-specific hyper-parameter tuning, and future work is recommended to automate hyper-parameter optimization to improve usability across diverse operational settings. Overall, the methodology provides high confidence for deployment across standard vision classification tasks.

arXiv: 2201.03019
  • Paper: Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion, Hongxu Yin et al. (2020). DeepInversion introduces generative data-free knowledge transfer via teacher feature inversion, providing the foundational synthesis mechanics and baseline challenges that PRE-DFKD directly stabilizes against forgetting.
  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal text establishes the core framework of knowledge distillation via softened teacher output distributions that underlying data-free distillation methods rely upon.
  • Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Dark Experience Replay demonstrates the value of replaying past model outputs (dark experience) to prevent catastrophic forgetting, motivating PRE-DFKD's generative pseudo-replay mechanism.
  • Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). This survey provides a comprehensive taxonomy of response-based, feature-based, and relation-based distillation methods essential for contextualizing data-free distillation architectures.
  • Paper: FitNets: Hints for Thin Deep Nets, Adriana Romero et al. (2015). FitNets establishes feature-based distillation using intermediate hints, which underpins the synthetic-data-aware feature reconstruction loss formulated in PRE-DFKD.
  • Paper: Up to 100x Faster Data-Free Knowledge Distillation, Gongfan Fang et al. (2022). FastDFKD extends data-free knowledge distillation by tackling the synthesis speed bottleneck through meta-generators, complementing PRE-DFKD's focus on stability and memory efficiency.
  • Paper: Adaptive Data-Free Quantization, Biao Qian et al. (2023). Adaptive Data-Free Quantization extends data-free generative learning principles to neural network quantization by framing synthetic calibration sample generation as a game between teacher and compressed models.
  • Paper: Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks, Zhiwei Deng et al. (2022). This work explores an alternative approach to memory-efficient continual learning and dataset distillation by compressing datasets into shared, addressable memory bases without storing raw historical data.
  • Paper: One-Step Diffusion Distillation through Score Implicit Matching, Weijian Luo et al. (2024). Score Implicit Matching explores data-free distillation within diffusion architectures, applying implicit score matching to compress generative models without real training data.
Cover for Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay

Abstract

Data-Free Knowledge Distillation (KD) allows knowledge transfer from a trained neural network (teacher) to a more compact one (student) in the absence of original training data. Existing works use a validation set to monitor the accuracy of the student over real data and report the highest performance throughout the entire process. However, validation data may not be available at distillation time either, making it infeasible to record the student snapshot that achieved the peak accuracy. Therefore, a practical data-free KD method should be robust and ideally provide monotonically increasing student accuracy during distillation. This is challenging because the student experiences knowledge degradation due to the distribution shift of the synthetic data. A straightforward approach to overcome this issue is to store and rehearse the generated samples periodically, which increases the memory footprint and creates privacy concerns. We propose to model the distribution of the previously observed synthetic samples with a generative network. In particular, we design a Variational Autoencoder (VAE) with a training objective that is customized to learn the synthetic data representations optimally. The student is rehearsed by the generative pseudo replay technique, with samples produced by the VAE. Hence knowledge degradation can be prevented without storing any samples. Experiments on image classification benchmarks show that our method optimizes the expected value of the distilled model accuracy while eliminating the large memory overhead incurred by the sample-storing methods.

Table of Contents

  • Introduction
  • Related Work
  • Knowledge Distillation (KD)
  • Data-Free KD
  • Replay in Continual Learning
  • Proposed PRE-DFKD Approach
  • Novel Sample Generation
  • Memory Sample Generation
  • Knowledge Distillation
  • Experimental Evaluation
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Pseudo Replay Enhanced Data-Free Knowledge Distillation Framework

    model/method

    Pseudo Replay Enhanced Data-Free Knowledge Distillation (PRE-DFKD) is a framework designed to distill knowledge from a pre-trained teacher network TT to a compact student network SS without requiring access to original training or validation data, while preventing catastrophic forgetting caused by synthetic data distribution shifts.

    The framework consists of two alternating stages:

    1. Synthetic Data Generation: Two specialized generative models produce synthetic inputs from standard normal latent vectors z∼N(0,I)z \sim \mathcal{N}(0, I). The novel sample generator Gn(z;θn)G_n(z; \theta_n) synthesizes novel samples xnx_n designed to exploit the current prediction discrepancies between teacher and student. The memory sample generator Gm(z;θm)G_m(z; \theta_m), which serves as the decoder of a Variational Autoencoder (VAE) paired with encoder Em(x;θEm)E_m(x; \theta_{E_m}), synthesizes memory samples xmx_m representing previous synthetic data distributions.
    2. Knowledge Distillation: The student network S(x;θS)S(x; \theta_S) is trained on concatenated batches (xn,xm)(x_n, x_m) using the teacher's outputs as supervision.

    By retaining historical synthetic distributions within the parameters of the VAE instead of storing historical synthetic images in memory, PRE-DFKD eliminates the catastrophic degradation of student accuracy over training epochs with a constant, small memory overhead.

  2. Knowl 2 — Novel Sample Generator Optimization Objective

    equation

    In the PRE-DFKD framework, the novel sample generator Gn(z;θn)G_n(z; \theta_n) parameterized by θn\theta_n generates synthetic samples xn=Gn(z)x_n = G_n(z) from latent vectors z∼N(0,I)z \sim \mathcal{N}(0, I) that maximize the information gain for the student SS while mimicking valid input characteristics of the teacher TT. The parameters are optimized by solving:

    θn∗:=arg⁡min⁡θn(Lϕ(T)+αLδ)\theta_n^* := \arg \min_{\theta_n} \left( \mathcal{L}_\phi^{(T)} + \alpha \mathcal{L}_\delta \right)

    where α\alpha is a weighting coefficient. The teacher-specific loss term Lϕ(T)\mathcal{L}_\phi^{(T)} is defined over a generated batch of size nn as:

    Lϕ(T)=1n∑i=1n(λ1tTilog⁡(y^Ti)−λ2∥fTi∥1)−λ3H(p(y^T))\mathcal{L}_\phi^{(T)} = \frac{1}{n} \sum_{i=1}^n \left( \lambda_1 t_T^i \log(\hat{y}_T^i) - \lambda_2 \|f_T^i\|_1 \right) - \lambda_3 \mathcal{H}(p(\hat{y}_T))

    where y^T=T(Gn(z))\hat{y}_T = T(G_n(z)) is the softmax output vector of the teacher, fTif_T^i is the activation map at the teacher's last fully connected layer for sample ii, tTi=arg⁡max⁡(y^Ti)t_T^i = \arg\max(\hat{y}_T^i) is the predicted one-hot class label, and H(p(y^T))\mathcal{H}(p(\hat{y}_T)) is the entropy of the mean predicted class distribution across the batch to encourage class diversity. The coefficients λ1,λ2,λ3\lambda_1, \lambda_2, \lambda_3 weight the cross-entropy, activation magnitude, and categorical entropy loss components.

    The student-teacher divergence loss Lδ\mathcal{L}_\delta encourages the generator to produce samples on which the student's predictions y^S=S(Gn(z))\hat{y}_S = S(G_n(z)) differ from the teacher's predictions y^T\hat{y}_T using the Jensen-Shannon (JS) divergence:

    Lδ=1−JS(y^T∥y^S)\mathcal{L}_\delta = 1 - \text{JS}(\hat{y}_T \parallel \hat{y}_S)

  3. Knowl 3 — Synthetic Data-Aware VAE Reconstruction Loss

    equation

    Standard VAE training with an L2\mathcal{L}_2 pixel-level reconstruction loss fails when modeling synthetic samples because synthetic data are generated purely for knowledge transfer and are sensitive to pixel-wise perturbations; minor pixel shifts frequently alter their categorical classification by the teacher network. To preserve categorical features in reconstructed samples, PRE-DFKD incorporates a feature-matching term defined across intermediate layers of the teacher network TT.

    Given input batch xx, encoder EE, and decoder DD, the total VAE training loss LVAE\mathcal{L}_{VAE} is:

    LVAE=Lrec+γLKLD\mathcal{L}_{VAE} = \mathcal{L}_{rec} + \gamma \mathcal{L}_{KLD}

    Lrec=∥x−D(E(x))∥1+∑l∈L∥T(x)l−T(D(E(x)))l∥1\mathcal{L}_{rec} = \|x - D(E(x))\|_1 + \sum_{l \in L} \|T(x)_l - T(D(E(x)))_l\|_1

    LKLD=DKL(N(μz,σz)∥N(0,I))\mathcal{L}_{KLD} = D_{KL}\left(\mathcal{N}(\mu_z, \sigma_z) \parallel \mathcal{N}(0, I)\right)

    where LL denotes a selected set of deep layers in the teacher network TT representing content-related features, μz\mu_z and σz\sigma_z are the latent mean and variance predicted by the encoder E(x)E(x), and γ\gamma is a scaling hyperparameter balancing the Kullback-Leibler divergence term.

  4. Knowl 4 — Training Algorithm for the VAE Memory Sample Generator

    algorithm

    The memory generator (VAE decoder GmG_m) and encoder EmE_m are jointly trained using both newly synthesized novel samples and reconstructed memory samples to prevent the generative model itself from catastrophic forgetting.

    Input: Novel sample generator Gn(z;θGn)G_n(z; \theta_{G_n}), memory sample generator Gm(z;θGm)G_m(z; \theta_{G_m}), encoder Em(x;θEm)E_m(x; \theta_{E_m}), batch size BB, latent dimension nn, weighting parameter γ\gamma
    Output: Updated parameters θGm,θEm\theta_{G_m}, \theta_{E_m}
    # Train with novel samples
    Sample BB latent vectors z∼N(0,I)z \sim \mathcal{N}(0, I)
    xn←Gn(z)x_n \leftarrow G_n(z)
    z^μ,z^σ←Em(xn)\hat{z}_\mu, \hat{z}_\sigma \leftarrow E_m(x_n)
    Sample zn∼N(z^μ,z^σ)z_n \sim \mathcal{N}(\hat{z}_\mu, \hat{z}_\sigma)
    x^n←Gm(zn)\hat{x}_n \leftarrow G_m(z_n)
    # Rehearse by reconstructing old memory samples
    Sample BB latent vectors zm∼N(0,I)z_m \sim \mathcal{N}(0, I)
    xm←Gm(zm)x_m \leftarrow G_m(z_m)
    zˉμ,zˉσ←Em(xm)\bar{z}_\mu, \bar{z}_\sigma \leftarrow E_m(x_m)
    Sample zˉm∼N(zˉμ,zˉσ)\bar{z}_m \sim \mathcal{N}(\bar{z}_\mu, \bar{z}_\sigma)
    x^m←Gm(zˉm)\hat{x}_m \leftarrow G_m(\bar{z}_m)
    # Calculate total loss and backpropagate
    LVAE←Lrec(xm,x^m,xn,x^n)+γLKLD(zˉμ,zˉσ)\mathcal{L}_{VAE} \leftarrow \mathcal{L}_{rec}(x_m, \hat{x}_m, x_n, \hat{x}_n) + \gamma \mathcal{L}_{KLD}(\bar{z}_\mu, \bar{z}_\sigma)
    θGm,θEm←optimizer.step(backward(LVAE),θGm,θEm)\theta_{G_m}, \theta_{E_m} \leftarrow \text{optimizer.step}(\text{backward}(\mathcal{L}_{VAE}), \theta_{G_m}, \theta_{E_m})

    The reconstruction loss Lrec\mathcal{L}_{rec} computes the combined L1L_1 pixel and teacher feature differences for both the novel batch (xn,x^n)(x_n, \hat{x}_n) and the memory batch (xm,x^m)(x_m, \hat{x}_m).

  5. Knowl 5 — Class-Balanced Memory Inference via Latent Variable Tuning

    algorithm

    Directly sampling latent vectors from standard normal distribution N(0,I)\mathcal{N}(0, I) for the memory generator can result in class-imbalanced batches. PRE-DFKD fixes the generator parameters θGm\theta_{G_m} and optimizes the latent vectors zmz_m prior to generating memory samples to ensure class diversity while keeping latent vectors close to the prior distribution.

    The latent vector optimization solves:

    zm∗:=arg⁡max⁡z∼N(0,I)(p(y^T)log⁡(p(y^T)))z_m^* := \arg \max_{z \sim \mathcal{N}(0, I)} \left( p(\hat{y}_T) \log(p(\hat{y}_T)) \right)

    subject to regularization preventing deviation from normality.

    Input: Frozen memory sample generator Gm(z;θGm)G_m(z; \theta_{G_m}), frozen teacher network T(x;θT)T(x; \theta_T), batch size BB
    Output: Class-balanced memory sample batch xmx_m
    # Sample and tune latent variables
    Sample BB vectors zm∼N(0,I)z_m \sim \mathcal{N}(0, I)
    xm←Gm(zm)x_m \leftarrow G_m(z_m)
    y^T←T(xm)\hat{y}_T \leftarrow T(x_m)
    # Regularization term containing predictive entropy and KL divergence
    R←1B∑i=1BtTilog⁡(y^Ti)+DKL(N(μz,σz)∥N(0,I))R \leftarrow \frac{1}{B} \sum_{i=1}^B t_T^i \log(\hat{y}_T^i) + D_{KL}(\mathcal{N}(\mu_z, \sigma_z) \parallel \mathcal{N}(0, I))
    Lz←−H(p(y^T))+R\mathcal{L}_z \leftarrow -\mathcal{H}(p(\hat{y}_T)) + R
    zm←optimizer.step(backward(Lz),zm)z_m \leftarrow \text{optimizer.step}(\text{backward}(\mathcal{L}_z), z_m)
    # Infer final memory batch
    xm←Gm(zm)x_m \leftarrow G_m(z_m)
    return xmx_m

    Here tTi=arg⁡max⁡(y^Ti)t_T^i = \arg\max(\hat{y}_T^i), H(p(y^T))\mathcal{H}(p(\hat{y}_T)) is categorical entropy across the batch, and DKLD_{KL} ensures zmz_m does not drift into regions that produce artifacts.

  6. Knowl 6 — PRE-DFKD Student Knowledge Distillation Objective

    equation

    In PRE-DFKD, knowledge transfer from the frozen teacher network T(x;θT)T(x; \theta_T) to the student network S(x;θS)S(x; \theta_S) is conducted by minimizing the L1L_1 prediction discrepancy over a combined batch consisting of novel synthetic samples xnx_n and replayed memory synthetic samples xmx_m:

    θS∗:=arg⁡min⁡θS∥S((xn,xm);θS)−T((xn,xm);θT)∥1\theta_S^* := \arg \min_{\theta_S} \|S((x_n, x_m); \theta_S) - T((x_n, x_m); \theta_T)\|_1

    where (xn,xm)(x_n, x_m) represents the concatenated batch of novel samples generated by GnG_n and pseudo-replay memory samples generated by GmG_m. Optimizing over both distributions simultaneously ensures the student acquires new domain knowledge while preserving representations learned during earlier distillation stages.

  7. Knowl 7 — Validation-Agnostic Knowledge Distillation Evaluation Metric

    definition

    In practical data-free knowledge distillation, no validation data is available during training to monitor student accuracy or perform early stopping at peak performance. Consequently, distillation termination step ts∈Rt_s \in \mathbb{R} is treated as a random variable. A robust data-free KD method must optimize the expected student accuracy while minimizing accuracy variance over distillation epochs:

    max⁡ts(E[accts]−σ2[accts])\max_{t_s} \left( \mathbb{E}[acc_{t_s}] - \sigma^2[acc_{t_s}] \right)

    where acctsacc_{t_s} is the student accuracy at epoch tst_s. Across KK independent experimental runs, the mean accuracy μ\mu and variance σ2\sigma^2 are computed over the run-averaged accuracy series Ei[accts(i)]\mathbb{E}_i[acc_{t_s}^{(i)}]:

    μ=Ets[Ei[accts(i)]],σ2=σts2[Ei[accts(i)]]\mu = \mathbb{E}_{t_s}\left[ \mathbb{E}_i[acc_{t_s}^{(i)}] \right], \quad \sigma^2 = \sigma_{t_s}^2\left[ \mathbb{E}_i[acc_{t_s}^{(i)}] \right]

    where accts(i)acc_{t_s}^{(i)} represents the validation accuracy at epoch tst_s in run ii. Peak accuracy accmax=max⁡i,tsaccts(i)acc_{max} = \max_{i, t_s} acc_{t_s}^{(i)} denotes the maximum accuracy recorded at any epoch across all runs.

  8. Knowl 8 — Student Accuracy and Memory Footprint Across Image Classification Benchmarks

    data/table

    The performance of PRE-DFKD was evaluated against replay-free baselines (DAFL, DFAD) and sample-storing replay baselines (CMI, MB-DFKD) across four image classification benchmarks. Experiments used LeNet5 →\to LeNet-half for MNIST (200 epochs) and ResNet34 →\to ResNet18 for CIFAR-10 (200 epochs), CIFAR-100 (400 epochs), and Tiny ImageNet (500 epochs). Values are averaged over 4 runs.

    MNIST CIFAR10 CIFAR100 Tiny ImageNet
    T: LeNet5 (98.9%) T: ResNet34 (95.4%) T: ResNet34 (77.9%) T: ResNet34 (71.2%)
    S: LeNet-half S: ResNet18 S: ResNet18 S: ResNet18
    Method μ\mu σ2\sigma^2 accmaxacc_{max} μ\mu σ2\sigma^2 accmaxacc_{max} μ\mu σ2\sigma^2 accmaxacc_{max} μ\mu σ2\sigma^2 accmaxacc_{max}
    Train with data 98.7 0.5 98.9 89.0 8.1 95.2 71.3 8.1 77.1 60.2 8.8 64.9
    DAFL 87.3 6.6 98.2 62.6 17.1 92.0 52.5 12.8 74.5 39.5 10.3 52.2
    DFAD 63.5 6.8 98.3 86.1 12.3 93.3 54.9 12.9 67.7 – – –
    CMI – – – 82.4 16.6 94.8 55.2 24.1 77.0 – – –
    MB-DFKD 88.6 3.2 98.3 83.3 16.4 92.4 64.4 18.3 75.4 45.7 11.5 53.5
    PRE-DFKD (ours) 90.3 1.9 98.3 87.4 10.3 94.1 70.2 11.1 77.1 46.3 11.0 54.2

    Memory footprints for storing replay data/models across datasets were:

    • MNIST: CMI: N.A. (failed due to missing BatchNorm in LeNet), MB-DFKD: 6.7 MB, PRE-DFKD: 2.1 MB.
    • CIFAR10: CMI: 250 MB, MB-DFKD: 20 MB, PRE-DFKD: 2.1 MB.
    • CIFAR100: CMI: 500 MB, MB-DFKD: 20 MB, PRE-DFKD: 2.1 MB.
    • Tiny ImageNet: CMI: 2.7 GB (failed to distill), MB-DFKD: 20 MB, PRE-DFKD: 2.1 MB.

    PRE-DFKD yields the highest mean student accuracy μ\mu and lowest variance σ2\sigma^2 across all datasets while maintaining a fixed 2.1 MB memory overhead for VAE parameters.

  9. Knowl 9 — Memory Scaling of Stored Replay vs. VAE Pseudo Replay

    data/table

    Methods that store raw synthetic samples for replay (such as Contrastive Model Inversion, CMI) experience linear memory scaling with respect to image resolution and the number of distillation steps tst_s. In contrast, PRE-DFKD requires a constant memory footprint of 2.1 MB (the size of the VAE parameters), invariant to distillation steps and image resolution.

    Pixel Dimension ts=2000t_s = 2000 ts=4000t_s = 4000 ts=8000t_s = 8000 ts=16000t_s = 16000
    32 2.5 GB 5 GB 10 GB 20 GB
    64 10 GB 20 GB 40 GB 80 GB
    128 40 GB 80 GB 160 GB 320 GB

    For 128×\times128 images at 16,000 distillation steps, CMI requires 320 GB of memory storage, creating severe hardware overhead and potential privacy leakage risks regarding the underlying training distribution, whereas PRE-DFKD operates with constant memory.

  10. Knowl 10 — Ablation Analysis of Generative Replay Components

    empirical result

    Ablation experiments conducted on the MNIST benchmark (LeNet5 teacher, LeNet-half student) demonstrate the individual impact of each PRE-DFKD module on preventing accuracy degradation:

    1. Generative Replay Integration on DFAD (PRE-DFAD): Coupling the memory generator to the replay-free DFAD baseline increases mean accuracy μ\mu from 59.4%59.4\% (σ2=6.4\sigma^2 = 6.4) to 89.5%89.5\% (σ2=4.7\sigma^2 = 4.7), demonstrating that generative pseudo replay prevents catastrophic forgetting regardless of the novel sample generation algorithm.
    2. Synthetic Data-Aware Reconstruction Loss: Replacing the proposed synthetic data-aware feature matching reconstruction loss with vanilla VAE pixel loss (Vanilla VAE) drops mean accuracy to μ=81.5%\mu = 81.5\% (σ2=1.5\sigma^2 = 1.5), because pixel-only reconstruction fails to preserve categorical class content for non-photorealistic synthetic samples.
    3. Class-Balanced Memory Latent Optimization: Omitting class-balanced latent tuning in memory inference reduces mean accuracy from 90.3%90.3\% (σ2=1.9\sigma^2 = 1.9) to μ=88.6%\mu = 88.6\% (σ2=1.6\sigma^2 = 1.6).

Coverage note — None was omitted. All primary contributions—including the dual-generator framework, novel and memory generator objectives, synthetic data-aware loss, latent variable optimization, distillation formulation, experimental evaluations, memory scalability comparisons, and ablation studies—have been captured.

References

  1. 1.Addepalli, S.; Nayak, G. K.; Chakraborty, A.; and Radhakrishnan, V. B. 2020. DeGAN: Data-Enriching gan for retrieving representative samples from a trained classifier. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 3130–3137.
  2. 2.Binici, K.; Pham, N. T.; Mitra, T.; and Leman, K. 2021. Preventing Catastrophic Forgetting and Distribution Mismatch in Knowledge Distillation via Synthetic Data. arXiv preprint arXiv:2108.05698.
  3. 3.Chen, H.; Wang, Y.; Xu, C.; Yang, Z.; Liu, C.; Shi, B.; Xu, C.; Xu, C.; and Tian, Q. 2019. Data-free learning of student networks. In ICCV, 3514–3522.
  4. 4.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee.
  5. 5.Fang, G.; Song, J.; Shen, C.; Wang, X.; Chen, D.; and Song, M. 2019. Data-Free Adversarial Distillation. arXiv preprint arXiv:1912.11006.
  6. 6.Fang, G.; Song, J.; Wang, X.; Shen, C.; Wang, X.; and Song, M. 2021. Contrastive Model Inversion for Data-Free Knowledge Distillation. arXiv preprint arXiv:2105.08584.
  7. 7.French, R. M. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4): 128–135.
  8. 8.Goodfellow, I. J.; Mirza, M.; Xiao, D.; Courville, A.; and Bengio, Y. 2015. An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks. arXiv:1312.6211.
  9. 9.Haroush, M.; Hubara, I.; Hoffer, E.; and Soudry, D. 2020. The knowledge within: Methods for data-free model compression. In CVPR, 8494–8502.
  10. 10.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In CVPR, 770–778.
  11. 11.Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  12. 12.Jing, Y.; Yang, Y.; Feng, Z.; Ye, J.; Yu, Y.; and Song, M. 2019. Neural style transfer: A review. IEEE transactions on visualization and computer graphics, 26(11): 3365–3385.
  13. 13.Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114.
  14. 14.Krizhevsky, A.; and Hinton, G. 2009. Learning Multiple Layers of Features from Tiny Images. Technical report, University of Toronto.
  15. 15.LeCun, Y.; Bottou, L.; Bengio, Y.; and Haffner, P. 1998. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11): 2278–2324.
  16. 16.Li, Z.; and Zhang, Y. 2020. Membership Leakage in Label-Only Exposures. arXiv preprint arXiv:2007.15528.
  17. 17.Ma, X.; Shen, Y.; Fang, G.; Chen, C.; Jia, C.; and Lu, W. 2020. Adversarial Self-Supervised Data-Free Distillation for Text Classification. arXiv preprint arXiv:2010.04883.
  18. 18.Mai, Z.; Li, R.; Jeong, J.; Quispe, D.; Kim, H.; and Sanner, S. 2021. Online Continual Learning in Image Classification: An Empirical Survey. arXiv:2101.10423.
  19. 19.Micaelli, P.; and Storkey, A. J. 2019. Zero-shot knowledge transfer via adversarial belief matching. In NIPS, 9551–9561.
  20. 20.Nayak, G. K.; Mopuri, K. R.; and Chakraborty, A. 2021. Effectiveness of arbitrary transfer sets for data-free knowledge distillation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 1430–1438.
  21. 21.Nayak, G. K.; Mopuri, K. R.; Shaj, V.; Radhakrishnan, V. B.; and Chakraborty, A. 2019. Zero-shot knowledge distillation in deep networks. In International Conference on Machine Learning, 4743–4751. PMLR.
  22. 22.Rebuffi, S.-A.; Kolesnikov, A.; Sperl, G.; and Lampert, C. H. 2017. iCaRL: Incremental Classifier and Representation Learning. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5533–5542.
  23. 23.Shin, H.; Lee, J. K.; Kim, J.; and Kim, J. 2017. Continual Learning with Deep Generative Replay. arXiv:1705.08690.
  24. 24.Wu, Y.; Chen, Y.; Wang, L.; Ye, Y.; Liu, Z.; Guo, Y.; Zhang, Z.; and Fu, Y. 2018. Incremental Classifier Learning with Generative Adversarial Networks. CoRR, abs/1802.00853.
  25. 25.Yin, H.; Molchanov, P.; Alvarez, J. M.; Li, Z.; Mallya, A.; Hoiem, D.; Jha, N. K.; and Kautz, J. 2020. Dreaming to distill: Data-free knowledge transfer via DeepInversion. In CVPR, 8715–8724.
  26. 26.Yoo, J.; Cho, M.; Kim, T.; and Kang, U. 2019. Knowledge extraction with no observable data. In NIPS, 2705–2714.
  27. 27.Zhang, Y.; Chen, H.; Chen, X.; Deng, Y.; Xu, C.; and Wang, Y. 2021. Data-Free Knowledge Distillation for Image Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7852–7861.

Citation

MLA
Binici, K., et al. “Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay”. arXiv, 2022, http://arxiv.org/abs/2201.03019v3.
APA
Binici, K., Aggarwal, S., Pham, N. T., Leman, K., & Mitra, T. (2022). Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay. arXiv. http://arxiv.org/abs/2201.03019v3
Chicago
Binici, K., S. Aggarwal, N. T. Pham, K. Leman, and T. Mitra. 2022. “Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay”. arXiv. http://arxiv.org/abs/2201.03019v3.
Harvard
Binici, K. et al. (2022) “Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.03019v3.
Vancouver
1. Binici K, Aggarwal S, Pham NT, Leman K, Mitra T (2022) Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay. arXiv

BibTeX

@article{binici2022robust,
  title = {Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay},
  author = {Binici, Kuluhan and Aggarwal, Shivam and Pham, Nam Trung and Leman, Karianto and Mitra, Tulika},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.03019v3},
  eprint = {2201.03019}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF