One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation

Zhiwei HaoJianyuan GuoKai HanYehui TangHan HuYunhe WangChang Xu

article2023NeurIPS147 citations

Proposes the OFA-KD framework to overcome feature misalignment in cross-architecture knowledge distillation by projecting intermediate representations into the logits space and applying confidence-based target modulation, achieving consistent performance gains across CNN, Vision Transformer, and MLP models.

Listen

Deploying lightweight artificial intelligence models on edge devices typically relies on knowledge distillation, a process where a smaller "student" model learns from a larger, highly accurate "teacher" model. However, existing techniques predominantly assume that both models share the same underlying architecture family. As diverse model architectures—such as Convolutional Neural Networks, Vision Transformers, and Multi-Layer Perceptrons—proliferate, finding high-performing teacher models within the exact same architectural family is increasingly difficult. Standard approaches that attempt to align intermediate internal features across different architectures fail because these architectures process information through fundamentally divergent visual representations.

The article introduces and evaluates "One-for-All Knowledge Distillation" (OFA-KD), a general-purpose framework designed to enable effective knowledge transfer between entirely different neural network architectures. The objective is to demonstrate that cross-architecture distillation can consistently outperform traditional same-family distillation and prior transfer baselines without introducing extra computational costs during deployment.

To bridge the structural gap, the authors developed a non-technical two-part mechanism: intermediate representations from the student model are routed through auxiliary output branches directly into the final classification prediction space, stripping away incompatible architecture-specific features. Additionally, an adaptive mathematical adjustment to the training loss dynamically emphasizes accurate target-class information when the teacher exhibits high prediction confidence, preventing the student from absorbing misleading signals caused by differing architectural biases. The framework was comprehensively evaluated across image classification benchmarks (CIFAR-100 and ImageNet-1K), testing all combinations among convolutional, transformer, and multi-layer perceptron models.

The findings show that the proposed framework consistently outperforms existing methods across heterogeneous pairings. On the ImageNet-1K benchmark, OFA-KD achieved accuracy gains of up to 0.7% over competitive baseline methods. On the CIFAR-100 dataset, the framework demonstrated significant improvements of up to 8.0% over alternative approaches, where conventional feature-matching techniques often collapsed entirely. Furthermore, using a high-capacity Vision Transformer teacher to train a standard ResNet-50 student produced an accuracy of 81.33%, surpassing the 80.64% achieved when using an exceptionally large same-family ResNet-152 teacher.

These results demonstrate that engineering teams are no longer constrained to matching teacher and student model families when compressing computer vision systems. Organizations can pair cutting-edge, high-performing foundation models with highly specialized or resource-constrained edge architectures. Because the auxiliary training branches are removed prior to deployment, these performance gains are achieved without increasing runtime latency, memory usage, or deployment risk.

Engineering teams seeking to compress visual AI models should consider cross-family distillation pipelines using intermediate alignment at the prediction layer rather than raw feature mapping. When implementing this method, teams should divide student networks into four sequential stages with intermediate exits to maximize transfer efficiency. However, practitioners should be aware that the framework requires tuning an adaptive modulation parameter to match the relative capability gap between teacher and student, and in certain compact model pairings, same-family teachers may still yield comparable or slightly better performance. Overall, the evidence provides strong confidence that prediction-space intermediate alignment effectively resolves feature mismatch across distinct architectures.

  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Introduces the foundational formulation of knowledge distillation via softened teacher output distributions that OFA-KD adapts across heterogeneous architectures.
  • Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). Analyzes the limitations of standard logit distillation and decouples target versus non-target class signals, motivating OFA-KD's adaptive loss modulation.
  • Paper: Co-advise: Cross Inductive Bias Distillation, Sucheng Ren et al. (2022). Explores distilling across differing inductive biases between CNNs and Vision Transformers, directly establishing the problem of architectural divergence addressed by OFA-KD.
  • Paper: FitNets: Hints for Thin Deep Nets, Adriana Romero et al. (2015). Establishes intermediate-layer representation matching for student-teacher networks, the foundational paradigm whose feature-level limitations OFA-KD overcomes.
  • Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). Provides a comprehensive benchmark and contrastive methodology for intermediate representation distillation across diverse teacher-student model pairs.
  • Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). Documents failure modes and capacity mismatches when distilling from large, disparate teachers, highlighting the need for adaptive alignment in cross-architecture distillation.
  • Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Demonstrates cross-architecture distillation from convolutional teachers into vision transformer students using specialized distillation tokens.
  • Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). Surveys classical response-based, feature-based, and relation-based distillation taxonomies that OFA-KD reconciles in the prediction space.
Cover for One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation

Abstract

Knowledge distillation (KD) has proven to be a highly effective approach for enhancing model performance through a teacher-student training scheme. However, most existing distillation methods are designed under the assumption that the teacher and student models belong to the same model family, particularly the hint-based approaches. By using centered kernel alignment (CKA) to compare the learned features between heterogeneous teacher and student models, we observe significant feature divergence. This divergence illustrates the ineffectiveness of previous hint-based methods in cross-architecture distillation. To tackle the challenge in distilling heterogeneous models, we propose a simple yet effective one-for-all KD framework called OFA-KD, which significantly improves the distillation performance between heterogeneous architectures. Specifically, we project intermediate features into an aligned latent space such as the logits space, where architecture-specific information is discarded. Additionally, we introduce an adaptive target enhancement scheme to prevent the student from being disturbed by irrelevant information. Extensive experiments with various architectures, including CNN, Transformer, and MLP, demonstrate the superiority of our OFA-KD framework in enabling distillation between heterogeneous architectures. Specifically, when equipped with our OFA-KD, the student models achieve notable performance improvements, with a maximum gain of 8.0% on the CIFAR-100 dataset and 0.7% on the ImageNet-1K dataset. PyTorch code and checkpoints can be found at https://github.com/Hao840/OFAKD.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Method
  • 3.1 Revisit feature distillation for heterogeneous architectures
  • 3.2 Generic heterogeneous knowledge distillation
  • 4 Experiment
  • 4.1 Experimental setup
  • 4.2 Distillation results on ImageNet-1K
  • 4.3 Distillation results on CIFAR-100
  • 4.4 Ablation study
  • 5 Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — OFA-KD Multi-Exit Logits-Space Distillation Framework

    model/method

    One-for-All Knowledge Distillation (OFA-KD) is a framework designed to enable effective knowledge distillation between heterogeneous model architectures, including Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and Multi-Layer Perceptrons (MLPs).

    Rather than enforcing intermediate representation matching directly in the feature space—where different architectural inductive biases cause feature misalignment—OFA-KD projects intermediate representations of the student model into the shared logits space. The student network is equipped with auxiliary exit branches inserted at intermediate stages. Each exit branch consists of a feature projector and a linear classifier layer that maps intermediate activations to class logits matching the dimension of the label space.

    During training, intermediate student predictions from these exit branches, as well as the final backbone prediction, are supervised by the final prediction logits of the pretrained teacher model using the OFA loss. This removes architecture-specific feature geometry while preserving semantic predictive distributions. At inference time, all auxiliary exit branches are detached, incurring zero additional inference latency or parameter overhead.

  2. Knowl 2 — Adaptive Target Information Enhancement Loss (OFA Loss)

    equation

    In knowledge distillation with heterogeneous architectures, discrepancies in inductive biases between the teacher and student lead to divergent predictive distributions for non-target classes. The OFA loss adaptively modulates the supervision on the ground-truth target class based on teacher confidence while suppressing noisy non-target class interactions.

    Let X×Y\mathcal{X} \times \mathcal{Y} be the joint sample and label space, c^∈Y\hat{c} \in \mathcal{Y} be the ground-truth target class, and c∈Yc \in \mathcal{Y} denote an arbitrary class index. Let pcsp^s_c and pctp^t_c denote the predicted probabilities of the student and teacher models for class cc on input sample xx, respectively. The OFA distillation loss LOFA\mathcal{L}_{\text{OFA}} is defined as:

    LOFA=−(1+pc^t)γlog⁡pc^s−Ec∼Y∖{c^}[pctlog⁡pcs]\mathcal{L}_{\text{OFA}} = -(1 + p^t_{\hat{c}})^\gamma \log p^s_{\hat{c}} - \mathbb{E}_{c \sim \mathcal{Y} \setminus \{\hat{c}\}} \left[ p^t_c \log p^s_c \right]

    where γ≥1\gamma \ge 1 is a scalar modulating parameter. When γ=1\gamma = 1, LOFA\mathcal{L}_{\text{OFA}} is equivalent to the standard Kullback-Leibler knowledge distillation loss LKD\mathcal{L}_{\text{KD}} up to an additive constant. When γ>1\gamma > 1, the loss adaptively scales the target class loss according to the teacher's prediction confidence pc^tp^t_{\hat{c}} on the true label.

  3. Knowl 3 — Binomial Expansion and Adaptive Weighting Behavior of the OFA Loss

    theoretical result

    For an integer modulating parameter γ∈Z+\gamma \in \mathbb{Z}^+, expanding the binomial factor (1+pc^t)γ(1 + p^t_{\hat{c}})^\gamma decouples the OFA loss into the standard knowledge distillation loss LKD\mathcal{L}_{\text{KD}} and an adaptive target enhancement term:

    LOFA=LKD−(∑k=1γ(γk)(pc^t)k−pc^t)log⁡pc^s\mathcal{L}_{\text{OFA}} = \mathcal{L}_{\text{KD}} - \left( \sum_{k=1}^\gamma \binom{\gamma}{k} (p^t_{\hat{c}})^k - p^t_{\hat{c}} \right) \log p^s_{\hat{c}}

    where pc^tp^t_{\hat{c}} and pc^sp^s_{\hat{c}} are the teacher and student probabilities for the target class c^\hat{c}.

    For the specific case γ=2\gamma = 2:

    LOFA,γ=2=LKD−(pc^t+(pc^t)2)log⁡pc^s\mathcal{L}_{\text{OFA}, \gamma=2} = \mathcal{L}_{\text{KD}} - \left( p^t_{\hat{c}} + (p^t_{\hat{c}})^2 \right) \log p^s_{\hat{c}}

    This formulation demonstrates that if the teacher model is highly confident in the correct target label (pc^t≈1p^t_{\hat{c}} \approx 1), the high-order term (pc^t)2(p^t_{\hat{c}})^2 remains large, reinforcing the target class supervision. If the teacher is uncertain or wrong (pc^t→0p^t_{\hat{c}} \to 0), the high-order term decays quadratically to zero, mitigating the negative impact of misleading teacher soft labels.

  4. Knowl 4 — Centered Kernel Alignment Analysis of Feature Divergence Across Heterogeneous Architectures

    empirical result

    Centered Kernel Alignment (CKA) was used to measure intermediate layer feature similarity across models from different architectural families: MobileNetV2 (CNN), ViT-Small (Transformer), and MLP-Mixer-B/16 (MLP).

    Let X∈Rn×d1X \in \mathbb{R}^{n \times d_1} and Y∈Rn×d2Y \in \mathbb{R}^{n \times d_2} represent intermediate activations across a mini-batch of nn samples extracted by two models, with Gram matrices K=XXTK = X X^T and L=YYTL = Y Y^T, and centering matrix Hn=In−1n11TH_n = I_n - \frac{1}{n} \mathbf{1}\mathbf{1}^T. CKA evaluates similarity via the Hilbert-Schmidt Independence Criterion (HSIC):

    CKA(K,L)=HSIC(K,L)HSIC(K,K)HSIC(L,L),where HSIC(K,L)=1(n−1)2tr(KHLH)\text{CKA}(K, L) = \frac{\text{HSIC}(K, L)}{\sqrt{\text{HSIC}(K, K) \text{HSIC}(L, L)}}, \quad \text{where } \text{HSIC}(K, L) = \frac{1}{(n-1)^2} \text{tr}(K H L H)

    Comparing homogeneous models (CNN-CNN, ViT-ViT, MLP-MLP) yields strong diagonal similarity, showing that layers at corresponding depths learn similar representations. In contrast, cross-architecture pairs (CNN-ViT, ViT-MLP, CNN-MLP) show marked feature divergence: MobileNetV2 features correlate only with the shallowest layers of ViT-S or Mixer-B/16, and ViT-S diverges strongly from Mixer-B/16. This divergence explains why conventional hint-based distillation methods that match intermediate feature tensors with Mean Squared Error (MSE) degrade performance when distilling across heterogeneous families.

  5. Knowl 5 — Exit Branch Architecture and Insertion Configuration in OFA-KD

    model/method

    In the OFA-KD framework, intermediate supervision is enabled by dividing student models into four sequential stages:

    1. Branch Insertion Points: For hierarchical architectures (ResNet, ConvNeXt, Swin Transformer), exit branches are placed at the end of each of the 4 hierarchical stages. For non-hierarchical isotropic architectures (standard ViT, MLP-Mixer, ResMLP), the network depth is evenly partitioned into four equal sections, with exit branches placed at the end of each quarter.
    2. Branch Design: In CNN student models, exit branches are constructed using depthwise and pointwise convolutional layers followed by a linear classification layer. In ViT and MLP student models, exit branches use Transformer blocks followed by a linear classification layer.
    3. Multi-Exit Ablation: When evaluating exit branch configurations on ImageNet-1K (e.g., DeiT-T teacher with ResNet18 student, or ResNet50 teacher with DeiT-T student), inserting exit branches at all four stages ({1,2,3,4}\{1, 2, 3, 4\}) achieves higher top-1 accuracy (71.34% for ResNet18 and 75.73% for DeiT-T) than any single-branch configuration ({1}\{1\}, {2}\{2\}, {3}\{3\}, or {4}\{4\}) or two/three-branch subset ({1,4}\{1, 4\} or {1,2,4}\{1, 2, 4\}).
  6. Knowl 6 — Cross-Architecture Distillation Performance on ImageNet-1K

    data/table

    Top-1 classification accuracy (%) on the ImageNet-1K validation set across fifteen heterogeneous teacher-student pairs comparing OFA-KD against hint-based (FitNet, CC, RKD, CRD) and logits-based (KD, DKD, DIST) knowledge distillation methods. CNN students are trained for 100 epochs with SGD; ViT/MLP students are trained for 300 epochs with AdamW. When ResNet50 is the teacher, logits-based methods are combined with FitNet (dagger\\dagger).

    Teacher Student Teacher Student (Scratch) FitNet CC RKD CRD KD DKD DIST OFA
    CNN-based students
    DeiT-T ResNet18 72.17 69.75 70.44 69.77 69.47 69.25 70.22 69.39 70.64 71.34
    Swin-T ResNet18 81.38 69.75 71.18 70.07 68.89 69.09 71.14 71.10 70.91 71.85
    Mixer-B/16 ResNet18 76.62 69.75 70.78 70.05 69.46 68.40 70.89 69.89 70.66 71.38
    DeiT-T MobileNetV2 72.17 68.87 70.95 70.69 69.72 69.60 70.87 70.14 71.08 71.39
    Swin-T MobileNetV2 81.38 68.87 71.75 70.69 67.52 69.58 72.05 71.71 71.76 72.32
    Mixer-B/16 MobileNetV2 76.62 68.87 71.59 70.79 69.86 68.89 71.92 70.93 71.74 72.12
    ViT-based students
    ResNet50 DeiT-T 80.38 72.17 75.84 72.56 72.06 68.53 75.10 75.60†^\dagger 75.13†^\dagger 76.55†^\dagger
    ConvNeXt-T DeiT-T 82.05 72.17 70.45 73.12 71.47 69.18 74.00 73.95 74.07 74.41
    Mixer-B/16 DeiT-T 76.62 72.17 74.38 72.82 72.24 68.23 74.16 72.82 74.22 74.46
    ResNet50 Swin-N 80.38 75.53 78.33 76.05 75.90 73.90 77.58 78.23†^\dagger 77.95†^\dagger 78.64†^\dagger
    ConvNeXt-T Swin-N 82.05 75.53 74.81 75.79 75.48 74.15 77.15 77.00 77.25 77.50
    Mixer-B/16 Swin-N 76.62 75.53 76.17 75.81 75.52 73.38 76.26 75.03 76.54 76.63
    MLP-based students
    ResNet50 ResMLP-S12 80.38 76.65 78.13 76.21 75.45 73.23 77.41 78.23†^\dagger 77.71†^\dagger 78.53†^\dagger
    ConvNeXt-T ResMLP-S12 82.05 76.65 74.69 75.79 75.28 73.57 76.84 77.23 77.24 77.53
    Swin-T ResMLP-S12 81.38 76.65 76.48 76.15 75.10 73.40 76.67 76.99 77.25 77.31

    OFA-KD achieves the highest accuracy across all fifteen evaluated cross-architecture configurations, outperforming the best competing baselines by 0.20% to 0.77% on CNN students and up to 0.71% on ViT and MLP students.

  7. Knowl 7 — Cross-Architecture Distillation Performance on CIFAR-100

    data/table

    Top-1 classification accuracy (%) on the CIFAR-100 test set across twelve heterogeneous teacher-student pairs, where input images are upsampled to 224×224224 \times 224 resolution and trained for 300 epochs.

    Teacher Student Teacher Student (Scratch) FitNet CC RKD CRD KD DKD DIST OFA
    CNN-based students
    Swin-T ResNet18 89.26 74.01 78.87 74.19 74.11 77.63 78.74 80.26 77.75 80.54
    ViT-S ResNet18 92.04 74.01 77.71 74.26 73.72 76.60 77.26 78.10 76.49 80.15
    Mixer-B/16 ResNet18 87.29 74.01 77.15 74.26 73.75 76.42 77.79 78.67 76.36 79.39
    Swin-T MobileNetV2 89.26 73.68 74.28 71.19 69.00 79.80 74.68 71.07 72.89 80.98
    ViT-S MobileNetV2 92.04 73.68 73.54 70.67 68.46 78.14 72.77 69.80 72.54 78.45
    Mixer-B/16 MobileNetV2 87.29 73.68 73.78 70.73 68.95 78.15 73.33 70.20 73.26 78.78
    ViT-based students
    ConvNeXt-T DeiT-T 88.41 68.00 60.78 68.01 69.79 65.94 72.99 74.60 73.55 75.76
    Mixer-B/16 DeiT-T 87.29 68.00 71.05 68.13 69.89 65.35 71.36 73.44 71.67 73.90
    ConvNeXt-T Swin-P 88.41 72.63 24.06 72.63 71.73 67.09 76.44 76.80 76.41 78.32
    Mixer-B/16 Swin-P 87.29 72.63 75.20 73.32 70.82 67.03 75.93 76.39 75.85 78.93
    MLP-based students
    ConvNeXt-T ResMLP-S12 88.41 66.56 45.47 67.70 65.82 63.35 72.25 73.22 71.93 81.22
    Swin-T ResMLP-S12 89.26 66.56 63.12 68.37 64.66 61.72 71.89 72.82 11.05 80.63

    On CIFAR-100, conventional feature-based methods struggle significantly when students are ViT or MLP models (e.g., FitNet achieves only 24.06% on ConvNeXt-T to Swin-P). OFA-KD outperforms all prior methods across all setups, with accuracy improvements over the second-best baseline ranging from 0.28% up to 8.00% (e.g. 81.22% vs 73.22% for ConvNeXt-T to ResMLP-S12).

  8. Knowl 8 — Hyperparameter Modulation Dynamics and Teacher Capacity Interaction

    empirical result

    The optimal value of the modulating parameter γ\gamma in the OFA loss function reflects the relative capability and prediction quality of the teacher architecture:

    1. Teacher Strength vs. γ\gamma: For a weaker teacher such as DeiT-T (72.17% ImageNet-1K top-1) distilling into ResNet18, the optimal parameter is γ=1.4\gamma = 1.4 (achieving 71.34% accuracy), where a higher exponent suppresses erroneous or noisy soft labels. For a stronger teacher such as ResNet50 (80.38% top-1) distilling into DeiT-T, the optimal parameter is γ=1.1\gamma = 1.1 (achieving 75.73% accuracy), allowing richer dark knowledge transfer from reliable soft labels.
    2. Loss Scaling and Gradient Clipping: Introducing the (1+pc^t)γ(1 + p^t_{\hat{c}})^\gamma factor scales the overall loss magnitude. Comparing loss scale multipliers (0.5 to 1.6) and maximum gradient clipping norms (1 to 7) on ImageNet-1K demonstrates that setting the scale factor to 1.0 and gradient clipping norm to 5 produces optimal stability and highest accuracy (75.73%).
  9. Knowl 9 — Homogeneous vs. Heterogeneous Teacher Distillation for ResNet50

    data/table

    A comparison on ImageNet-1K evaluating a ResNet50 student (scratch accuracy 79.86%) distilled from a larger homogeneous teacher (ResNet152, 82.83% top-1) versus a larger heterogeneous teacher (ViT-B, 86.53% top-1).

    Teacher Teacher Acc. RKD Review CRD DKD DIST OFA
    ResNet152 82.83 79.53 80.06 79.33 80.49 80.55 80.64
    ViT-B 86.53 79.38 79.32 79.48 80.76 80.90 81.33

    While previous hint-based KD approaches (such as ReviewKD and CRD) fail to benefit from the superior accuracy of ViT-B due to feature representation divergence (achieving lower accuracy with ViT-B than with ResNet152), OFA-KD leverages the stronger heterogeneous ViT-B teacher to achieve 81.33% accuracy, surpassing the 80.64% achieved with the homogeneous ResNet152 teacher by 0.69% (and outperforming the best competing baseline on ViT-B by 0.43%).

  10. Knowl 10 — Limitations of Cross-Architecture Distillation with OFA-KD

    limitation

    The OFA-KD approach exhibits two main limitations:

    1. For certain smaller student architectures such as ResNet18, distilling from a heterogeneous teacher (e.g., DeiT-T achieving 71.34%, Mixer-B/16 achieving 71.38%, or Swin-T achieving 71.85% on ImageNet-1K) yields lower student accuracy than distilling from a homogeneous ResNet34 teacher (which achieves 72.10% accuracy with OFA).
    2. The modulating parameter γ\gamma requires empirical adjustment based on the relative capacity gap and architectural inductive bias between the teacher and student; suboptimal configuration of γ\gamma can degrade student training performance.

Coverage note — None was omitted; all key theoretical formulations, empirical findings, architectural mechanisms, ablation analyses, benchmark results on ImageNet-1K and CIFAR-100, and stated limitations were captured.

References

  1. 1.Zhiwei Hao, Yong Luo, Zhi Wang, Han Hu, and Jianping An. Cdfkd-mfs: Collaborative data-free knowledge distillation via multi-level feature sharing. IEEE Transactions on Multimedia, 2022. 1
  2. 2.Zhiwei Hao, Yong Luo, Han Hu, Jianping An, and Yonggang Wen. Data-free ensemble knowledge distillation for privacy-conscious multimedia model compression. In Proceedings of the 29th ACM International Conference on Multimedia, 2021. 1
  3. 3.Hanting Chen, Tianyu Guo, Chang Xu, Wenshuo Li, Chunjing Xu, Chao Xu, and Yunhe Wang. Learning student networks in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 1
  4. 4.Hanting Chen, Yunhe Wang, Han Shu, Changyuan Wen, Chunjing Xu, Boxin Shi, Chao Xu, and Chang Xu. Distilling portable generative adversarial networks for image translation. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020. 1
  5. 5.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 1, 3, 6, 9
  6. 6.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. In International Conference on Learning Representations, 2015. 1, 2, 3, 5, 6
  7. 7.Sergey Zagoruyko and Nikos Komodakis. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In International Conference on Learning Representations, 2017. 1, 3
  8. 8.Zhiwei Hao, Jianyuan Guo, Ding Jia, Kai Han, Yehui Tang, Chao Zhang, Han Hu, and Yunhe Wang. Learning efficient vision transformers via fine-grained manifold distillation. In Advances in Neural Information Processing Systems, 2022. 1, 2, 3
  9. 9.Tao Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. Knowledge distillation from a stronger teacher. In Advances in Neural Information Processing Systems, 2022. 2, 3, 6, 9
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 2, 6
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 2, 3, 4, 6
  12. 12.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022. 2, 6
  13. 13.Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision. In Neural Information Processing Systems, pages 24261–24272, 2021. 2, 3, 6
  14. 14.Nikolaos Passalis, Maria Tzelepi, and Anastasios Tefas. Heterogeneous knowledge distillation using information flow modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2339–2348, 2020. 2
  15. 15.Sihui Luo, Xinchao Wang, Gongfan Fang, Yao Hu, Dapeng Tao, and Mingli Song. Knowledge amalgamation from heterogeneous networks by common feature learning. arXiv preprint arXiv:1906.10546, 2019. 2
  16. 16.Chengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song, Li Sun, and Mingli Song. Customizing student networks from heterogeneous teachers via adaptive knowledge amalgamation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. 2
  17. 17.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, volume 139, pages 10347–10357, 2021. 2, 3, 4, 6
  18. 18.Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Computer Vision and Pattern Recognition, pages 4510–4520, 2018. 2, 6
  19. 19.Kai Han, Yunhe Wang, Qi Tian, Jianyuan Guo, Chunjing Xu, and Chang Xu. Ghostnet: More features from cheap operations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020. 2
  20. 20.Jianyuan Guo, Kai Han, Han Wu, Yehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. Cmt: Convolutional neural networks meet vision transformers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 2
  21. 21.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In International Conference on Computer Vision, pages 9992–10002, 2021. 2, 3, 6
  22. 22.Xiangtai Li, Henghui Ding, Wenwei Zhang, Haobo Yuan, Jiangmiao Pang, Guangliang Cheng, Kai Chen, Ziwei Liu, and Chen Change Loy. Transformer-based visual segmentation: A survey. arXiv preprint arXiv:2304.09854, 2023. 2
  23. 23.Jianyuan Guo, Yehui Tang, Kai Han, Xinghao Chen, Han Wu, Chao Xu, Chang Xu, and Yunhe Wang. Hire-mlp: Vision mlp via hierarchical rearrangement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 2
  24. 24.Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, and Hervé Jégou. Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404, 2021. 2, 3, 6
  25. 25.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Information Processing Systems, pages 5998–6008, 2017. 2
  26. 26.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, 2020. 2
  27. 27.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, et al. Generative pretraining from pixels. In International Conference on Machine Learning, 2020. 2
  28. 28.Chengcheng Wang, Wei He, Ying Nie, Jianyuan Guo, Chuanjian Liu, Kai Han, and Yunhe Wang. Gold-yolo: Efficient object detector via gather-and-distribute mechanism. arXiv preprint arXiv:2309.11331, 2023. 2
  29. 29.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In International Conference on Computer Vision, pages 548–558, 2021. 2
  30. 30.Yehui Tang, Kai Han, Yunhe Wang, Chang Xu, Jianyuan Guo, Chao Xu, and Dacheng Tao. Patch slimming for efficient vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 3
  31. 31.Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. Be your own teacher: Improve the performance of convolutional neural networks via self distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. 3
  32. 32.Mary Phuong and Christoph H Lampert. Distillation-based training for multi-exit architectures. In Proceedings of the IEEE/CVF international conference on computer vision, 2019. 3
  33. 33.Yunteng Luan, Hanyu Zhao, Zhi Yang, and Yafei Dai. Msd: Multi-self-distillation learning via multi-classifiers within deep neural networks. arXiv preprint arXiv:1911.09418, 2019. 3
  34. 34.Frederick Tung and Greg Mori. Similarity-preserving knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. 3
  35. 35.Xiatian Zhu, Shaogang Gong, et al. Knowledge distillation by on-the-fly native ensemble. Advances in neural information processing systems, 31, 2018. 3
  36. 36.Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In ACM SIGKDD, pages 535–541, 2006. 3
  37. 37.Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. Relational knowledge distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. 3, 6, 9
  38. 38.Jianyuan Guo, Kai Han, Yunhe Wang, Han Wu, Xinghao Chen, Chunjing Xu, and Chang Xu. Distilling object detectors via decoupled features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 3
  39. 39.Andrey Malinin, Bruno Mlodozeniec, and Mark J. F. Gales. Ensemble distribution distillation. In International Conference on Learning Representations, 2020. 3
  40. 40.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation. In IEEE/CVF International Conference on Learning Representations, 2020. 3, 6
  41. 41.Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In Computer Vision and Pattern Recognition, 2017. 3
  42. 42.Linfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma, and Chenglong Bao. Task-oriented feature distillation. In Neural Information Processing Systems, 2020. 3
  43. 43.Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. Distilling knowledge via knowledge review. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5008–5017, 2021. 3, 9
  44. 44.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018. 3
  45. 45.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4750–4759, 2022. 3
  46. 46.Songhua Liu, Kai Wang, Xingyi Yang, Jingwen Ye, and Xinchao Wang. Dataset distillation via factorization. In Advances in Neural Information Processing Systems, 2022. 3
  47. 47.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. In Advances in Neural Information Processing Systems, 2021. 3
  48. 48.Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. Algorithms for learning kernels based on centered alignment. J. Mach. Learn. Res., 13:795–828, 2012. 4
  49. 49.Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. Similarity of neural network representations revisited. In International Conference on Machine Learning, volume 97, pages 3519–3529, 2019. 4
  50. 50.Arthur Gretton, Kenji Fukumizu, Choon Hui Teo, Le Song, Bernhard Schölkopf, and Alexander J. Smola. A kernel statistical test of independence. In Neural Information Processing Systems, pages 585–592, 2007. 4
  51. 51.Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q. Weinberger. Multi-scale dense networks for resource efficient image classification. In International Conference on Learning Representations, 2018. 5
  52. 52.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 6
  53. 53.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. 6
  54. 54.Baoyun Peng, Xiao Jin, Jiaheng Liu, Dongsheng Li, Yichao Wu, Yu Liu, Shunfeng Zhou, and Zhaoning Zhang. Correlation congruence for knowledge distillation. In IEEE/CVF International Conference on Computer Vision, 2019. 6
  55. 55.Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang. Decoupled knowledge distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 6, 9
  56. 56.Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi. A comprehensive overhaul of feature distillation. In International Conference on Computer Vision, pages 1921–1930, 2019. 9

Citation

MLA
Hao, Z., et al. “One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 79570–82, https://proceedings.neurips.cc/paper_files/paper/2023/file/fb8e5f198c7a5dcd48860354e38c0edc-Paper-Conference.pdf.
APA
Hao, Z., Guo, J., Han, K., Tang, Y., Hu, H., Wang, Y., & Xu, C. (2023). One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. Advances in Neural Information Processing Systems, 36, 79570–79582. https://proceedings.neurips.cc/paper_files/paper/2023/file/fb8e5f198c7a5dcd48860354e38c0edc-Paper-Conference.pdf
Chicago
Hao, Z., J. Guo, K. Han, et al. 2023. “One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation”. Advances in Neural Information Processing Systems 36: 79570–82. https://proceedings.neurips.cc/paper_files/paper/2023/file/fb8e5f198c7a5dcd48860354e38c0edc-Paper-Conference.pdf.
Harvard
Hao, Z. et al. (2023) “One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 79570–79582. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/fb8e5f198c7a5dcd48860354e38c0edc-Paper-Conference.pdf.
Vancouver
1. Hao Z, Guo J, Han K, Tang Y, Hu H, Wang Y, Xu C (2023) One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 79570–79582

BibTeX

@inproceedings{hao2023one,
  title = {One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation},
  author = {Hao, Zhiwei and Guo, Jianyuan and Han, Kai and Tang, Yehui and Hu, Han and Wang, Yunhe and Xu, Chang},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {79570-79582},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/fb8e5f198c7a5dcd48860354e38c0edc-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors