Co-advise: Cross Inductive Bias Distillation
Sucheng RenZhengqi GaoTianyu HuaZihui XueYonglong TianShengfeng HeHang Zhao
Demonstrates that distilling knowledge from multiple lightweight teachers with complementary inductive biases, such as convolution and involution, boosts vision transformer performance beyond traditional heavy-teacher distillation while reducing computational costs.
Vision transformers deliver strong performance in computer vision tasks but struggle when trained on standard datasets because they lack the built-in structural assumptions, known as inductive biases, found in conventional networks. Existing solutions rely on knowledge distillation using massive, computationally expensive convolutional neural network (CNN) teacher models to guide training. However, these heavy teachers cause student transformers to mirror their specific classification errors while requiring enormous compute budgets and training time.
The article evaluates whether the architectural inductive bias of teacher models matters more than their individual accuracy during distillation, demonstrating a novel training approach called cross inductive bias distillation to improve vision transformer efficiency and accuracy.
The authors conducted comparative experiments on benchmark image classification datasets (ImageNet-1k and ImageNet-100) alongside out-of-distribution robustness tests. They evaluated student vision transformers trained by pairing two lightweight teacher models featuring complementary structural designs: a convolutional network (spatial-agnostic and channel-specific) and an involutional network (spatial-specific and channel-agnostic). They also introduced a token inductive bias alignment technique, giving dedicated tokens within the transformer structural stems to match the inductive biases of their respective teachers.
The evaluation revealed several key findings. First, teacher architectural diversity matters significantly more than teacher scale or accuracy; boosting a single teacher's accuracy yielded plateauing student performance, whereas combining distinct architectural types unlocked substantial gains. Second, the proposed student models (termed CiT) outperformed prior vision transformers of identical size while using teacher models with 50% to 80% fewer parameters than standard approaches. Third, the small variant with token alignment achieved an 82.7% top-1 accuracy on ImageNet-1k without altering the core attention architecture. Fourth, out-of-distribution benchmarks confirmed that student tokens successfully inherited the complementary strengths and robustness patterns of both distinct teachers.
These findings indicate that machine learning teams can bypass the substantial computational overhead and financial costs associated with training massive teacher models. Organizations deploying vision systems can achieve superior accuracy and stronger generalization by distilling from multiple small, architecturally diverse models rather than relying on brute-force scaling of homogeneous architectures.
Decision-makers and engineering leads should adopt cross-architecture distillation pipelines and incorporate structural token alignment when training data-efficient vision transformers. Future research should expand beyond CNN and involution pairings to evaluate other distinct neural architectures, such as graph-based or recurrent models, to further broaden the diversity of knowledge transferred.
A primary limitation of this work is the requirement to train two separate lightweight teachers independently, though their combined training burden remains substantially lower than previous single heavy-teacher baselines. Confidence in the empirical results is high across standard vision benchmarks, though performance should be verified when applied to specialized industrial domains outside general image classification.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It introduced token-based distillation from convolutional teachers to vision transformers (DeiT), establishing the foundational baseline that Co-advise directly critiques and enhances.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It presents the original Vision Transformer architecture and documents its dependence on heavy pretraining due to the lack of built-in inductive biases.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). It provides the comparative representation analysis between vision transformers and convolutional networks that motivates transferring complementary structural biases.
- Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). It analyzes the efficacy of knowledge distillation and demonstrates that larger teacher accuracy does not necessarily yield better student performance, a central premise of Co-advise.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). It establishes the foundational principles of knowledge distillation used to transfer representations across neural architectures.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). It explores incorporating convolutional inductive biases into vision transformer architectures to solve sample efficiency issues.
- Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). It details how intermediate and alternative teacher structures can mitigate capacity mismatch problems in knowledge distillation.
- Paper: Learning Efficient Vision Transformers via Fine-Grained Manifold Distillation, Zhiwei Hao et al. (2022). It explores fine-grained patch-level manifold distillation tailored to compress vision transformers without relying on traditional heavy teacher baselines.
- Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). It decouples logit-based distillation objectives to significantly improve cross-architecture knowledge transfer efficiency.
- Paper: Understanding The Robustness in Vision Transformers, Daquan Zhou et al. (2022). It investigates why vision transformers demonstrate distinct robustness behaviors and develops attentional architectures resilient to out-of-distribution corruptions.
- Paper: Mimetic Initialization of Self-Attention Layers, Asher Trockman et al. (2023). It provides a complementary initialization approach that embeds pre-training inductive biases into transformers from scratch without requiring teacher models.
