One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation
Zhiwei HaoJianyuan GuoKai HanYehui TangHan HuYunhe WangChang Xu
Proposes the OFA-KD framework to overcome feature misalignment in cross-architecture knowledge distillation by projecting intermediate representations into the logits space and applying confidence-based target modulation, achieving consistent performance gains across CNN, Vision Transformer, and MLP models.
Deploying lightweight artificial intelligence models on edge devices typically relies on knowledge distillation, a process where a smaller "student" model learns from a larger, highly accurate "teacher" model. However, existing techniques predominantly assume that both models share the same underlying architecture family. As diverse model architectures—such as Convolutional Neural Networks, Vision Transformers, and Multi-Layer Perceptrons—proliferate, finding high-performing teacher models within the exact same architectural family is increasingly difficult. Standard approaches that attempt to align intermediate internal features across different architectures fail because these architectures process information through fundamentally divergent visual representations.
The article introduces and evaluates "One-for-All Knowledge Distillation" (OFA-KD), a general-purpose framework designed to enable effective knowledge transfer between entirely different neural network architectures. The objective is to demonstrate that cross-architecture distillation can consistently outperform traditional same-family distillation and prior transfer baselines without introducing extra computational costs during deployment.
To bridge the structural gap, the authors developed a non-technical two-part mechanism: intermediate representations from the student model are routed through auxiliary output branches directly into the final classification prediction space, stripping away incompatible architecture-specific features. Additionally, an adaptive mathematical adjustment to the training loss dynamically emphasizes accurate target-class information when the teacher exhibits high prediction confidence, preventing the student from absorbing misleading signals caused by differing architectural biases. The framework was comprehensively evaluated across image classification benchmarks (CIFAR-100 and ImageNet-1K), testing all combinations among convolutional, transformer, and multi-layer perceptron models.
The findings show that the proposed framework consistently outperforms existing methods across heterogeneous pairings. On the ImageNet-1K benchmark, OFA-KD achieved accuracy gains of up to 0.7% over competitive baseline methods. On the CIFAR-100 dataset, the framework demonstrated significant improvements of up to 8.0% over alternative approaches, where conventional feature-matching techniques often collapsed entirely. Furthermore, using a high-capacity Vision Transformer teacher to train a standard ResNet-50 student produced an accuracy of 81.33%, surpassing the 80.64% achieved when using an exceptionally large same-family ResNet-152 teacher.
These results demonstrate that engineering teams are no longer constrained to matching teacher and student model families when compressing computer vision systems. Organizations can pair cutting-edge, high-performing foundation models with highly specialized or resource-constrained edge architectures. Because the auxiliary training branches are removed prior to deployment, these performance gains are achieved without increasing runtime latency, memory usage, or deployment risk.
Engineering teams seeking to compress visual AI models should consider cross-family distillation pipelines using intermediate alignment at the prediction layer rather than raw feature mapping. When implementing this method, teams should divide student networks into four sequential stages with intermediate exits to maximize transfer efficiency. However, practitioners should be aware that the framework requires tuning an adaptive modulation parameter to match the relative capability gap between teacher and student, and in certain compact model pairings, same-family teachers may still yield comparable or slightly better performance. Overall, the evidence provides strong confidence that prediction-space intermediate alignment effectively resolves feature mismatch across distinct architectures.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Introduces the foundational formulation of knowledge distillation via softened teacher output distributions that OFA-KD adapts across heterogeneous architectures.
- Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). Analyzes the limitations of standard logit distillation and decouples target versus non-target class signals, motivating OFA-KD's adaptive loss modulation.
- Paper: Co-advise: Cross Inductive Bias Distillation, Sucheng Ren et al. (2022). Explores distilling across differing inductive biases between CNNs and Vision Transformers, directly establishing the problem of architectural divergence addressed by OFA-KD.
- Paper: FitNets: Hints for Thin Deep Nets, Adriana Romero et al. (2015). Establishes intermediate-layer representation matching for student-teacher networks, the foundational paradigm whose feature-level limitations OFA-KD overcomes.
- Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). Provides a comprehensive benchmark and contrastive methodology for intermediate representation distillation across diverse teacher-student model pairs.
- Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). Documents failure modes and capacity mismatches when distilling from large, disparate teachers, highlighting the need for adaptive alignment in cross-architecture distillation.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Demonstrates cross-architecture distillation from convolutional teachers into vision transformer students using specialized distillation tokens.
- Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). Surveys classical response-based, feature-based, and relation-based distillation taxonomies that OFA-KD reconciles in the prediction space.
- Paper: Understanding the Role of the Projector in Knowledge Distillation, Roy Miles et al. (2024). Provides formal theoretical and empirical insights into how projection heads function during distillation to map mismatched representations across network architectures.
- Paper: Logit Standardization in Knowledge Distillation, Shangquan Sun et al. (2024). Extends logit-space distillation mechanisms by standardizing logit variance across architectures to resolve teacher-student capacity discrepancies.
