Co-advise: Cross Inductive Bias Distillation

Sucheng RenZhengqi GaoTianyu HuaZihui XueYonglong TianShengfeng HeHang Zhao

article2022CVPR68 citations

Demonstrates that distilling knowledge from multiple lightweight teachers with complementary inductive biases, such as convolution and involution, boosts vision transformer performance beyond traditional heavy-teacher distillation while reducing computational costs.

Listen

Vision transformers deliver strong performance in computer vision tasks but struggle when trained on standard datasets because they lack the built-in structural assumptions, known as inductive biases, found in conventional networks. Existing solutions rely on knowledge distillation using massive, computationally expensive convolutional neural network (CNN) teacher models to guide training. However, these heavy teachers cause student transformers to mirror their specific classification errors while requiring enormous compute budgets and training time.

The article evaluates whether the architectural inductive bias of teacher models matters more than their individual accuracy during distillation, demonstrating a novel training approach called cross inductive bias distillation to improve vision transformer efficiency and accuracy.

The authors conducted comparative experiments on benchmark image classification datasets (ImageNet-1k and ImageNet-100) alongside out-of-distribution robustness tests. They evaluated student vision transformers trained by pairing two lightweight teacher models featuring complementary structural designs: a convolutional network (spatial-agnostic and channel-specific) and an involutional network (spatial-specific and channel-agnostic). They also introduced a token inductive bias alignment technique, giving dedicated tokens within the transformer structural stems to match the inductive biases of their respective teachers.

The evaluation revealed several key findings. First, teacher architectural diversity matters significantly more than teacher scale or accuracy; boosting a single teacher's accuracy yielded plateauing student performance, whereas combining distinct architectural types unlocked substantial gains. Second, the proposed student models (termed CiT) outperformed prior vision transformers of identical size while using teacher models with 50% to 80% fewer parameters than standard approaches. Third, the small variant with token alignment achieved an 82.7% top-1 accuracy on ImageNet-1k without altering the core attention architecture. Fourth, out-of-distribution benchmarks confirmed that student tokens successfully inherited the complementary strengths and robustness patterns of both distinct teachers.

These findings indicate that machine learning teams can bypass the substantial computational overhead and financial costs associated with training massive teacher models. Organizations deploying vision systems can achieve superior accuracy and stronger generalization by distilling from multiple small, architecturally diverse models rather than relying on brute-force scaling of homogeneous architectures.

Decision-makers and engineering leads should adopt cross-architecture distillation pipelines and incorporate structural token alignment when training data-efficient vision transformers. Future research should expand beyond CNN and involution pairings to evaluate other distinct neural architectures, such as graph-based or recurrent models, to further broaden the diversity of knowledge transferred.

A primary limitation of this work is the requirement to train two separate lightweight teachers independently, though their combined training burden remains substantially lower than previous single heavy-teacher baselines. Confidence in the empirical results is high across standard vision benchmarks, though performance should be verified when applied to specialized industrial domains outside general image classification.

arXiv: 2106.12378
Cover for Co-advise: Cross Inductive Bias Distillation

Abstract

The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into the influence of models inductive biases in knowledge distillation (e.g., convolution and involution). Our key observation is that the teacher accuracy is not the dominant reason for the student accuracy, but the teacher inductive bias is more important. We demonstrate that lightweight teachers with different architectural inductive biases can be used to co-advise the student transformer with outstanding performances. The rationale behind is that models designed with different inductive biases tend to focus on diverse patterns, and teachers with different inductive biases attain various knowledge despite being trained on the same dataset. The diverse knowledge provides a more precise and comprehensive description of the data and compounds and boosts the performance of the student during distillation. Furthermore, we propose a token inductive bias alignment to align the inductive bias of the token with its target teacher model. With only lightweight teachers provided and using this cross inductive bias distillation method, our vision transformers (termed as CiT) outperform all previous vision transformers (ViT) of the same architecture on ImageNet. Moreover, our small size model CiT-SAK further achieves 82.7% Top-1 accuracy on ImageNet without modifying the attention module of the ViT. Code is available at https://github.com/OliverRensu/co-advise.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Proposed Method
  • 3.1. Cross Inductive Bias Teachers
  • 3.2. Token Inductive Bias Alignment
  • 3.3. Cross Inductive Bias Distillation
  • 4. Experimental Results
  • 4.1. Implementation Details
  • 4.2. Comparison among Different Architectures
  • 4.3. Ablation on Cross Inductive Bias Distillation
  • 4.3.1 Teacher Performance and Inductive Biases.
  • 4.3.2 Student Performance and Inductive Biases.
  • 4.3.3 Naive Multi and Cross Inductive Bias Teachers.
  • 4.3.4 Effectiveness of Multiple Distillation Tokens.
  • 4.4. Ablation on Token Inductive Bias Alignments
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Cross Inductive Bias Distillation Framework

    model/method

    Cross Inductive Bias Distillation (CiT) is a knowledge distillation framework for Vision Transformers (ViT) that distills knowledge simultaneously from multiple lightweight teacher models possessing distinct architectural inductive biases rather than from a single large-capacity teacher.

    The framework pairs two complementary teacher paradigms:

    1. Convolutional Neural Network (CNN) Teacher: Implements spatial-agnostic, channel-specific operations with local translation equivariance.
    2. Involutional Neural Network (INN) Teacher: Implements spatial-specific, channel-agnostic operations capable of capturing long-range spatial context.

    The student transformer incorporates three specialized tokens:

    • Class Token (zsclass\mathbf{z}_s^{class}): Supervised by ground-truth labels via standard cross-entropy loss.
    • Conv Token (zsconv\mathbf{z}_s^{conv}): Supervised by the CNN teacher's predicted logits via Kullback-Leibler (KL) divergence.
    • Inv Token (zsinv\mathbf{z}_s^{inv}): Supervised by the INN teacher's predicted logits via KL divergence.

    During inference, the prediction stored in the Class token is retrieved as the model's final output classification.

  2. Knowl 2 — CiT Distillation Loss Function

    equation

    The overall training loss L\mathcal{L} for Cross Inductive Bias Distillation combines ground-truth cross-entropy with two temperature-scaled Kullback-Leibler (KL) divergence objectives:

    L=λ0LCE(σ(zsclass),y)+λ1τ12LKL[σ(zsconvτ1),σ(zt1τ1)]+λ2τ22LKL[σ(zsinvτ2),σ(zt2τ2)]\mathcal{L} = \lambda_0 \mathcal{L}_{CE}(\sigma(\mathbf{z}_s^{class}), \mathbf{y}) + \lambda_1 \tau_1^2 \mathcal{L}_{KL}\left[\sigma\left(\frac{\mathbf{z}_s^{conv}}{\tau_1}\right), \sigma\left(\frac{\mathbf{z}_t^1}{\tau_1}\right)\right] + \lambda_2 \tau_2^2 \mathcal{L}_{KL}\left[\sigma\left(\frac{\mathbf{z}_s^{inv}}{\tau_2}\right), \sigma\left(\frac{\mathbf{z}_t^2}{\tau_2}\right)\right]

    where:

    • y∈{0,1}C\mathbf{y} \in \{0, 1\}^C is the one-hot ground-truth target vector across CC classes.
    • σ(⋅)\sigma(\cdot) represents the Softmax activation function.
    • LCE\mathcal{L}_{CE} is the standard cross-entropy loss.
    • LKL[p,q]=∑k=1Cpklog⁡(pk/qk)\mathcal{L}_{KL}[p, q] = \sum_{k=1}^C p_k \log(p_k / q_k) is the Kullback-Leibler divergence between predicted probability distributions pp and qq.
    • zsclass,zsconv,zsinv∈RC\mathbf{z}_s^{class}, \mathbf{z}_s^{conv}, \mathbf{z}_s^{inv} \in \mathbb{R}^C are the output logits from the student transformer's Class, Conv, and Inv tokens, respectively.
    • zt1∈RC\mathbf{z}_t^1 \in \mathbb{R}^C denotes the output logits from the CNN teacher (e.g., RegNet).
    • zt2∈RC\mathbf{z}_t^2 \in \mathbb{R}^C denotes the output logits from the INN teacher (e.g., RedNet).
    • τ1,τ2∈(0,∞)\tau_1, \tau_2 \in (0, \infty) are distillation temperature parameters (set to τ1=τ2=1.0\tau_1 = \tau_2 = 1.0).
    • λ0,λ1,λ2∈[0,1]\lambda_0, \lambda_1, \lambda_2 \in [0, 1] are balancing weights (set to λ0=λ1=λ2=1.0\lambda_0 = \lambda_1 = \lambda_2 = 1.0).
  3. Knowl 3 — Token Inductive Bias Alignment

    model/method

    Token Inductive Bias Alignment introduces specific inductive biases into student tokens prior to the transformer layers rather than initializing all tokens with identical random distributions:

    1. Class Token: Initialized using a truncated Gaussian distribution without any structural inductive bias, allowing unconstrained learning from ground-truth one-hot labels.
    2. Conv Token: Generated by passing the input image through a convolutional stem (replacing the standard linear patch projection) and performing spatial average pooling over the resulting feature maps.
    3. Inv Token: Generated by passing the input image through an involution stem and performing spatial average pooling over the resulting feature maps.

    This structural alignment tailors the representational capacity of the Conv token and Inv token to match the operational characteristics of the CNN teacher and INN teacher, respectively.

  4. Knowl 4 — Superiority of Cross-Inductive-Bias Teachers over Single-Bias Ensembles

    empirical result

    In knowledge distillation to a Vision Transformer (evaluated on ImageNet-100):

    • Diminishing Returns from Teacher Accuracy within the Same Family: Scaling teacher capacity or training duration within the same architectural bias family (e.g., training RegNetY-200M for 100 additional epochs to gain +9% teacher accuracy, or scaling from RegNetY-200M to RegNetY-600M for a +6.5% teacher accuracy gain) produces negligible improvements in student accuracy (Transformer-Ti student accuracy remains near 86.5%–86.6%).
    • Ensembling Identical Inductive Biases: Distilling from two CNN teachers simultaneously (e.g., two ResNet-18 models or ResNet-18 + ResNet-50) yields marginal gains (reaching 87.0%–87.2% Top-1).
    • Cross Inductive Bias Advantage: Distilling from two lightweight teachers with complementary inductive biases (ResNet-18 [CNN] and RedNet-26 [INN]) elevates Transformer-Ti Top-1 accuracy to 88.0%. The diversity of inductive hypotheses provided by teachers matters more than teacher scale or individual accuracy.
  5. Knowl 5 — Student Architecture Compatibility for Cross-Inductive-Bias Distillation

    empirical result

    When distilling cross-inductive knowledge from a CNN teacher (ResNet-18) and an INN teacher (RedNet-26) on ImageNet-100, the choice of student architecture determines distillation success:

    • ResNet-10 Student (CNN): Possesses strong convolutional inductive bias that conflicts with the INN teacher. Accuracy reaches 83.0% with ResNet-18 alone, 82.6% with RedNet-26 alone, and only 83.4% when distilling from both teachers together (+0.4% gain).
    • Mixer-Ti Student (Pure MLP): Possesses minimal inductive bias but lacks representational flexibility under parameter constraints (12 layers, hidden dimension 192). It exhibits high output KL divergence from teachers (0.358 to ResNet-18, 0.313 to RedNet-26) and achieves only 82.3% Top-1 accuracy under dual-teacher distillation.
    • Transformer-Ti Student (ViT): Possesses relaxed inductive bias combined with self-attention expressiveness capable of representing both convolution and involution. It achieves low output KL divergence (0.255 to ResNet-18, 0.154 to RedNet-26) and improves from an undistilled baseline of 81.8% to 88.0% Top-1 accuracy under dual-teacher distillation.
  6. Knowl 6 — ImageNet-1k Classification Performance of CiT Models

    data/table

    Cross inductive bias vision transformers (CiT) outperform ViT, DeiT, CNNs, and INNs of comparable sizes on ImageNet-1k while using substantially smaller teacher models.

    Model Type Parameters (M) Throughput (Images/s) Top-1 Accuracy (%)
    ResNet-50 CNN 25.6 1349.4 76.2
    ResNet-101 CNN 44.5 799.4 77.4
    RegNetY-600MF CNN 6.1 1200.5 75.5
    RegNetY-4.0GF CNN 20.6 350.5 79.4
    RegNetY-8.0GF CNN 39.2 220.5 79.9
    RedNet-26 INN 9.2 1820.9 73.6
    RedNet-50 INN 15.5 1066.8 78.4
    RedNet-101 INN 25.6 657.4 79.1
    RedNet-152 INN 34.0 459.3 79.3
    ViT-B/16 Transformer 86.0 166.88 77.9
    ViT-L/16 Transformer 307.0 54.4 76.5
    DeiT-Ti Transformer 5.0 3082.9 72.2
    DeiT-S Transformer 22.0 1562.0 79.8
    DeiT-Ti-KD Transformer 6.0 3060.8 74.5
    DeiT-S-KD Transformer 22.0 1546.1 81.2
    CiT-Ti (Ours) Transformer 6.0 3053.0 75.3
    CiT-S (Ours) Transformer 22.0 1564.1 82.0
    CiT-SAK (Ours) Transformer 26.0 1414.1 82.7

    Throughput was measured on a single NVIDIA RTX 3090 GPU with batch size 64. CiT-Ti uses RegNetY-600M (6M params) and RedNet-26 (9M params) teachers, achieving 75.3% accuracy and surpassing DeiT-Ti-KD (74.5%), which uses a much heavier RegNetY-16GF teacher (84M params). CiT-S uses RegNetY-4GF (21M params) and RedNet-101 (26M params) teachers, reaching 82.0% accuracy. CiT-SAK incorporates token alignment and achieves 82.7% Top-1 accuracy.

  7. Knowl 7 — Effect of Dedicated Distillation Tokens

    empirical result

    On ImageNet-100 distillation experiments using Transformer-Ti with ResNet-18 (CNN) and RedNet-26 (INN) teachers:

    • Using a single token to fit both the true label and multiple teacher logits simultaneously results in gradient conflict, achieving 83.5% Top-1 accuracy.
    • Assigning three dedicated tokens (Class token for labels, Conv token for the CNN teacher, and Inv token for the INN teacher) eliminates objective interference and improves Top-1 accuracy to 88.0%, representing an absolute gain of 4.5% solely from multi-token specialization.
  8. Knowl 8 — Token-Specific Out-of-Distribution Robustness Inheritance

    empirical result

    Evaluation on out-of-distribution datasets (ImageNet-A for natural adversarial examples, ImageNet-R for semantic shift, and ImageNet-C for common corruptions) shows that individual tokens in the student transformer inherit the specific behavioral profiles of their corresponding teachers:

    • CNN Teachers vs. INN Teachers: Standalone CNNs perform better on ImageNet-R (higher accuracy) and ImageNet-C (lower mean corruption error, mCE) but worse on ImageNet-A, whereas standalone INNs perform better on ImageNet-A.
    • Inheritance by Conv and Inv Tokens: In a distilled student (CiT with token alignment + KD), the Conv token achieves higher accuracy on ImageNet-R (47.41% vs. 46.81%) and lower error on ImageNet-C (38.11 mCE vs. 38.04 mCE) compared to the Inv token, but lower accuracy on ImageNet-A (23.58% vs. 25.15%).
    • Knowledge distillation transfers these inductive behavioral traits more effectively than structural inductive bias injection via stems alone without distillation.
  9. Knowl 9 — Inductive Bias Injection via Stems without Distillation

    empirical result

    Directly replacing the standard patch projection of a Vision Transformer (Transformer-S) with inductive bias stems provides standalone performance gains on ImageNet-1k even in the absence of knowledge distillation:

    • Standard Transformer-S baseline: 79.8% Top-1 accuracy.
    • Transformer-S with Convolution stem injection: 81.5% Top-1 accuracy (+1.7%).
    • Transformer-S with Involution stem injection: 81.4% Top-1 accuracy (+1.6%).
    • Transformer-S with both Convolution and Involution stems injected (CiT-SA without KD): 81.8% Top-1 accuracy (+2.0%).

    When combined with cross-inductive bias distillation (CiT-SAK), accuracy reaches 82.7%.

  10. Knowl 10 — Limitations of Cross Inductive Bias Distillation

    limitation

    The Cross Inductive Bias Distillation framework has two primary limitations:

    1. Separate Teacher Training Overhead: The method requires training multiple lightweight teachers (e.g., one CNN and one INN) independently before training the student transformer, although the combined training time and parameter count of these teachers remain considerably smaller than that of a monolithic teacher (e.g., RegNetY-16GF in DeiT).
    2. Scope of Explored Biases: Experimental verification is primarily centered on pairing spatial-agnostic/channel-specific convolution and spatial-specific/channel-agnostic involution; exploration of other structural inductive biases remains unaddressed.

Coverage note — None was omitted; all core architectural methods, mathematical loss formulations, comparative empirical results, student and teacher ablations, and stated limitations are fully represented.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229, 2020.
  2. 2.Jang Hyun Cho and Bharath Hariharan. On the efficacy of knowledge distillation. In ICCV, pages 4794–4802, 2019.
  3. 3.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relationship between self-attention and convolutional layers. arXiv preprint arXiv:1911.03584, 2019.
  4. 4.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relationship between self-attention and convolutional layers. In ICLR, 2020.
  5. 5.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR, 2009.
  6. 6.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  7. 7.Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born again neural networks. In ICML, pages 1607–1616. PMLR, 2018.
  8. 8.Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born again neural networks. In ICML, pages 1607–1616. PMLR, 2018.
  9. 9.Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. arXiv preprint arXiv:2104.01136, 2021.
  10. 10.Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chunjing Xu, and Yunhe Wang. Cmt: Convolutional neural networks meet vision transformers. arXiv preprint arXiv:2107.06263, 2021.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  12. 12.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021.
  13. 13.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. ICLR, 2019.
  14. 14.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. CVPR, 2021.
  15. 15.Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi. Knowledge distillation with adversarial samples supporting decision boundary. In AAAI, volume 33, pages 3771–3778, 2019.
  16. 16.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  17. 17.Judy Hoffman, Saurabh Gupta, Jian Leong, Sergio Guadarrama, and Trevor Darrell. Cross-modal adaptation for rgb-d detection. In ICRA, pages 5032–5039, 2016.
  18. 18.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. NeurIPS, 25:1097–1105, 2012.
  19. 19.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  20. 20.Duo Li, Jie Hu, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu, Tong Zhang, and Qifeng Chen. Involution: Inverting the inherence of convolution for visual recognition. arXiv preprint arXiv:2103.06255, 2021.
  21. 21.David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik. Unifying distillation and privileged information. arXiv preprint arXiv:1511.03643, 2015.
  22. 22.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2017.
  23. 23.Umberto Michieli and Pietro Zanuttigh. Knowledge distillation for incremental learning in semantic segmentation. CVIU, 205:103167, 2021.
  24. 24.Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In IEEE Symposium on Security and Privacy (SP), pages 582–597, 2016.
  25. 25.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In CVPR, 2020.
  26. 26.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  27. 27.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, pages 6105–6114. PMLR, 2019.
  28. 28.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation. ICLR, 2020.
  29. 29.Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision, 2021.
  30. 30.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  31. 31.Frederick Tung and Greg Mori. Similarity-preserving knowledge distillation. In ICCV, pages 1365–1374, 2019.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017.
  33. 33.Jiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu, Chen Change Loy, and Dahua Lin. Carafe: Content-aware reassembly of features. In ICCV, October 2019.
  34. 34.Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. 119:9929–9939, 13–18 Jul 2020.
  35. 35.Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677, 2020.
  36. 36.Zihui Xue, Sucheng Ren, Zhengqi Gao, and Hang Zhao. Multimodal knowledge expansion. arXiv preprint arXiv:2103.14431, 2021.

Citation

MLA
Ren, S., et al. “Co-advise: Cross Inductive Bias Distillation”. arXiv, 2021, http://arxiv.org/abs/2106.12378v1.
APA
Ren, S., Gao, Z., Hua, T., Xue, Z., Tian, Y., He, S., & Zhao, H. (2021). Co-advise: Cross Inductive Bias Distillation. arXiv. http://arxiv.org/abs/2106.12378v1
Chicago
Ren, S., Z. Gao, T. Hua, et al. 2021. “Co-advise: Cross Inductive Bias Distillation”. arXiv. http://arxiv.org/abs/2106.12378v1.
Harvard
Ren, S. et al. (2021) “Co-advise: Cross Inductive Bias Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.12378v1.
Vancouver
1. Ren S, Gao Z, Hua T, Xue Z, Tian Y, He S, Zhao H (2021) Co-advise: Cross Inductive Bias Distillation. arXiv

BibTeX

@article{ren2021advise,
  title = {Co-advise: Cross Inductive Bias Distillation},
  author = {Ren, Sucheng and Gao, Zhengqi and Hua, Tianyu and Xue, Zihui and Tian, Yonglong and He, Shengfeng and Zhao, Hang},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.12378v1},
  eprint = {2106.12378}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE