Understanding the Role of the Projector in Knowledge Distillation

Roy MilesKrystian Mikolajczyk

article2024AAAI60 citations

Demonstrates that linear projectors implicitly capture relational sample information in knowledge distillation, establishing a computationally lightweight recipe with batch normalization and a soft maximum distance that matches or outperforms complex state-of-the-art methods across vision tasks.

Listen

Deploying modern artificial intelligence models on edge hardware and resource-constrained systems requires compressing large, high-performing networks into smaller, efficient architectures. Knowledge distillation accomplishes this by training a compact student model to imitate an advanced teacher model. However, leading feature-distillation techniques increasingly rely on complex multi-layer attention blocks, memory banks, or handcrafted relational structures. These additions substantially increase computational and memory overhead during training and often lack clear theoretical foundations.

The article investigates the fundamental mechanics of feature representation distillation. It specifically examines the mathematical and empirical roles of the projection layer, feature normalization schemes, and distance functions to formulate a simple, computationally efficient, and highly performant distillation pipeline.

The researchers conducted an analytical study of training dynamics and validated their approach across standard visual benchmark tasks. Their empirical evaluation covered image classification on the CIFAR-100 and ImageNet-1K datasets, object detection on the COCO2017 dataset, and the data-efficient training of Vision Transformers, evaluating multiple architectural pairings across different network sizes and types.

The investigation produced several key findings. First, mathematical analysis and empirical tests reveal that a simple linear projection layer implicitly captures relational information across past data samples, acting as an exponential moving average that eliminates the need for expensive memory banks or explicit correlation matrices. Second, feature normalization directly determines distillation quality: batch normalization prevents the projection weights from collapsing singular values toward zero, preserving critical representation dimensions compared to L2 normalization or unnormalized setups. Third, overly complex or non-linear projectors learn to decorrelate input and output features from the student backbone, degrading knowledge transfer and indicating that a lightweight linear layer is optimal. Fourth, incorporating a smooth soft-maximum distance metric (the LogSum function) softens the penalty on poorly aligned features, successfully bridging large capacity gaps between students and teachers. Finally, this streamlined framework achieves state-of-the-art results, including a 77.2% top-1 accuracy on ImageNet with a tiny Vision Transformer, outperforming specialized baselines by 2.2% while strongly preserving spatial translational equivariance.

These results demonstrate that practitioner teams can replace computationally heavy distillation architectures with a minimal recipe: a linear projection, batch normalization, and a LogSum loss function. This transition reduces engineering complexity, lowers GPU training memory costs, and shortens development timelines while improving final model accuracy.

Organizations training compact models for computer vision deployments should adopt this streamlined distillation pipeline. Future research and development should explore further refinements to normalization schemes and projection dynamics, and engineering teams should run initial pilot validations when transferring this framework to other domains, such as natural language processing or multimodal learning.

The findings are supported with high confidence across broad computer vision benchmarks and diverse architectures. However, practitioners should exercise caution when working with very small batch sizes, where standard batch normalization may require spatial dimension adaptations, as demonstrated in the article's object detection experiments.

  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Introduces the foundational framework of knowledge distillation via temperature-scaled softmax matching, which the source paper analyzes and refines through metric learning and projection layers.
  • Paper: Relational Knowledge Distillation, Wonpyo Park et al. (2019). Pioneers relational knowledge distillation by matching structural distance relations across samples, establishing the conceptual basis for the source paper's analysis of relational gradients enabled by projectors.
  • Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). Frames representation distillation as metric and contrastive learning between teacher and student embeddings, providing the direct foundation for the source paper's metric learning view of projection layers.
  • Paper: FitNets: Hints for Thin Deep Nets, Adriana Romero et al. (2015). Introduces intermediate feature projection layers to align feature dimensions between disparate student and teacher architectures, the core mechanism investigated by the source.
  • Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). Documents failure modes and capacity gap limitations in knowledge distillation that motivate the source paper's soft maximum and normalization techniques.
  • Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Establishes data-efficient image transformer distillation protocols (DeiT) that serve as a central evaluation benchmark in the source paper.
  • Paper: Logit Standardization in Knowledge Distillation, Shangquan Sun et al. (2024). Extends the investigation into logit normalization and student-teacher capacity discrepancies by proposing dynamic Z-score standardization prior to temperature scaling.
Cover for Understanding the Role of the Projector in Knowledge Distillation

Abstract

In this paper we revisit the efficacy of knowledge distillation as a function matching and metric learning problem. In doing so we verify three important design decisions, namely the normalisation, soft maximum function, and projection layers as key ingredients. We theoretically show that the projector implicitly encodes information on past examples, enabling relational gradients for the student. We then show that the normalisation of representations is tightly coupled with the training dynamics of this projector, which can have a large impact on the students performance. Finally, we show that a simple soft maximum function can be used to address any significant capacity gap problems. Experimental results on various benchmark datasets demonstrate that using these insights can lead to superior or comparable performance to state-of-the-art knowledge distillation techniques, despite being much more computationally efficient. In particular, we obtain these results across image classification (CIFAR100 and ImageNet), object detection (COCO2017), and on more difficult distillation objectives, such as training data efficient transformers, whereby we attain a 77.2% top-1 accuracy with DeiT-Ti on ImageNet. Code and models are publicly available.

Table of Contents

  • Introduction
  • Related Work
  • Understanding the Role of the Projector
  • Benchmark Evaluation
  • Data Efficient Training for Transformers
  • Classification on CIFAR100 and ImageNet
  • Object Detection on COCO
  • Conclusion
  • References

Knowls

  1. Knowl 1 — A low-cost distillation recipe based on projection, normalization, and LogSum

    model/method

    The paper proposes a feature-distillation recipe with three components: (1) a learnable linear projector that maps the student representation into the teacher representation dimension, (2) normalization of the projected student and teacher representations, and (3) a LogSum soft-maximum distance. The student is trained with its task loss plus the feature-distillation loss:

    L=Ltask+D(Zs,Zt;Wp),\mathcal{L}=\mathcal{L}_{\mathrm{task}}+D(Z_s,Z_t;W_p),

    where ZsZ_s and ZtZ_t are the student and teacher representations, WpW_p is the student-to-teacher projection matrix, and DD is the LogSum distance. The large-scale evaluations use a linear projector, batch normalization without affine parameters, ϵ=0.0001\epsilon=0.0001, and smoothing factor α=4\alpha=4. For CIFAR-100, the projector is an MLP with hidden size 1024 and no additional KL-divergence term. For object detection, where batches are smaller, normalization is performed over spatial height and width rather than across the batch. The pipeline is designed to avoid explicit correlation matrices, memory banks, and large trainable projector networks.

  2. Knowl 2 — Projector updates encode student–teacher relational information

    equation

    For a bias-free linear projector and squared feature-matching loss, let Zs∈Rn×dsZ_s\in\mathbb{R}^{n\times d_s} be a batch of student features, Zt∈Rn×dtZ_t\in\mathbb{R}^{n\times d_t} the corresponding teacher features, and Wp∈Rds×dtW_p\in\mathbb{R}^{d_s\times d_t} the projector. The loss and continuous-time gradient update are

    D(Zs,Zt;Wp)=12∥ZsWp−Zt∥F2,D(Z_s,Z_t;W_p)=\frac{1}{2}\left\|Z_sW_p-Z_t\right\|_F^2, W˙p=−ZsTZsWp+ZsTZt=Cst−CsWp,\dot W_p=-Z_s^{\mathsf T}Z_sW_p+Z_s^{\mathsf T}Z_t=C_{st}-C_sW_p,

    where Cs=ZsTZs∈Rds×dsC_s=Z_s^{\mathsf T}Z_s\in\mathbb{R}^{d_s\times d_s} is the student self-correlation matrix and Cst=ZsTZt∈Rds×dtC_{st}=Z_s^{\mathsf T}Z_t\in\mathbb{R}^{d_s\times d_t} is the student–teacher cross-correlation matrix. Thus, every projector update depends on correlations between the current student and teacher batches, while the learned weights retain information from previous updates. The paper interprets the projector as an implicit relational-information encoder: it can transfer feature relationships without explicitly constructing and storing batch correlation matrices or memory banks.

  3. Knowl 3 — Normalization determines the projector fixed point and its memory behavior

    theoretical result

    If the student features are whitened so that their self-correlation is the identity, Cs=IdsC_s=I_{d_s}, the stationary condition for the linear projector becomes

    Cst−CsWp=0⟹Wp=Cst.C_{st}-C_sW_p=0 \quad\Longrightarrow\quad W_p=C_{st}.

    Under this condition, the projector directly represents the cross-correlation between student and teacher features. With learning rate αp\alpha_p and weight-decay coefficient η\eta, the discrete update is

    Wp←Wp+αpW˙p−ηWp.W_p\leftarrow W_p+\alpha_p\dot W_p-\eta W_p.

    When η=αp\eta=\alpha_p, the decay term gives the old projector weights a factor 1−η1-\eta, so the projector behaves like an exponential moving average of relational information. This explains why a trainable projector can provide some of the role of a momentum encoder or explicit memory bank while using only a parameter matrix. The paper further concludes that changing the normalization scheme changes both the projector training trajectory and its limiting solution.

  4. Knowl 4 — Batch normalization preserves more information than alternative normalizations

    data/table

    The normalization ablation on the ImageNet-1K 20% subset evaluates three teacher–student architecture pairs using the same distillation setup. Batch normalization gives the strongest and most consistent student accuracy, particularly for the ConvNeXt-to-EfficientNet-B0 pair. The reported top-1 accuracies are:

    Could not parse LaTeX table

    The singular-value trajectories plotted on page 3 for a ResNet-18 student and ResNet-50 teacher show that no normalization and L2 normalization shrink more projector singular values toward zero, whereas batch normalization retains more nonzero singular directions. The paper interprets singular-value shrinkage as input-dimension collapse and information loss, linking the normalization choice to distillation effectiveness.

  5. Knowl 5 — Increasing projector capacity can decouple the student from the teacher

    limitation

    Replacing the linear projector with a larger MLP does not reliably improve distillation. In the input–output correlation trajectories plotted on page 4, all tested projectors gradually decorrelate the projected output from the student input, and the decorrelation becomes stronger as the MLP hidden dimension increases from 64 to 512 and 2048. A high-capacity projector can therefore learn features that are not shared with the student backbone, weakening the training signal that should shape the student representation. The paper identifies an inherent trade-off: more projector capacity can encode more information, but it can also reduce alignment with the student backbone. For simplicity and robustness, the large-scale experiments use a linear projector.

  6. Knowl 6 — LogSum soft maximum improves distillation across large capacity gaps

    model/method

    To reduce the harm caused when a small student cannot exactly align with a much larger teacher representation, the paper replaces the ordinary feature distance with a LogSum soft maximum. Let rir_i denote the iith scalar element of the residual matrix ZsWp−ZtZ_sW_p-Z_t, and let α>0\alpha>0 be a smoothing factor. The proposed distance is

    D(Zs,Zt;Wp)=log⁡(∑i∣ri∣α).D(Z_s,Z_t;W_p)=\log\left(\sum_i |r_i|^{\alpha}\right).

    This loss softens the contribution of relatively close matches within a batch, making feature matching less brittle when the student and teacher have substantially different capacities. The page-5 ablations show consistent gains from LogSum:

    Could not parse LaTeX table

    The improvement is largest for the ResNet-50-to-ResNet-18 pair, where the capacity gap is larger. With the same architecture pairs, the reported accuracies for different α\alpha values are:

    Could not parse LaTeX table

    Performance is robust over a broad range of α\alpha, with the best values generally occurring near 44–55.

  7. Knowl 7 — The recipe improves data-efficient training of vision transformers

    empirical result

    For ImageNet-1K data-efficient training, the paper applies batch normalization, a linear projector, and α=4\alpha=4 under the Co-Advice training methodology. The method obtains 77.2% top-1 accuracy with the 6M-parameter DeiT-Ti student and a RegNetY-160 teacher, compared with 72.2% for an undistilled DeiT-Ti, 74.9% for CivT-Ti with an ensemble teacher, 74.8% for DearKD, and 75.0% for USKD. With the 22M-parameter DeiT-S student, the method obtains 82.1%, compared with 79.8% without distillation, 82.0% for CivT-S, 81.5% for DearKD, and 80.8% for USKD. The reported results are:

    Could not parse LaTeX table

    The gain is largest when the student is much weaker than the teacher and becomes smaller as student and teacher capacities converge. The paper attributes this behavior partly to transfer of the teacher's spatial or translational inductive bias.

  8. Knowl 8 — Feature distillation transfers translational equivariance to transformer students

    theoretical result

    The paper measures whether a student transformer acquires the teacher's spatial inductive bias. For an input xx, translation operator TT, and network layer ϕ\phi, translation equivariance is the property

    ϕ(Tx)=Tϕ(x).\phi(Tx)=T\phi(x).

    The corresponding deviation measure is

    μT(ϕ)=∥ϕ(Tx)−Tϕ(x)∥22.\mu_T(\phi)=\left\|\phi(Tx)-T\phi(x)\right\|_2^2.

    The measurement is applied to a block of self-attention layers after removing the distillation and class tokens and reshaping the patch tokens back into spatial feature maps. On page 7, the reported equivariance deviations for DeiT-S are:

    Could not parse LaTeX table

    The substantially lower deviation for the proposed method indicates that the student preserves more spatial locality while still retaining some global context. The paper presents this as evidence that feature distillation can transfer an explicit inductive bias from a convolutional teacher to a transformer student, rather than merely transferring output predictions.

  9. Knowl 9 — Classification results remain competitive across architectures and capacity gaps

    empirical result

    On CIFAR-100, the experiments use 60,000 32×3232\times32 RGB images, the teacher weights supplied with the SSKD comparison, an MLP projector with hidden size 1024, and no additional KL loss. For the ten teacher–student pairs in the order shown in the paper—WRN40-2→WRN16-2, WRN40-2→WRN40-1, R56→R20, R32×4→R8×4, VGG13→MBv2, R50→MBv2, R50→VGG8, R32×4→ShuffleV1, R32×4→ShuffleV2, and WRN40-2→ShuffleV1—the proposed method obtains the following top-1 accuracies:

    • With RandAugment: 76.14,75.42,71.75,76.44,71.47,72.81,76.20,77.32,79.06,76.14, 75.42, 71.75, 76.44, 71.47, 72.81, 76.20, 77.32, 79.06, and 79.2279.22 percent.
    • With the predefined rotation augmentations used by SSKD: 77.61,76.04,72.25,78.37,72.82,73.51,77.08,78.99,79.86,77.61, 76.04, 72.25, 78.37, 72.82, 73.51, 77.08, 78.99, 79.86, and 78.7978.79 percent.

    The proposed method is strongest relative to competing methods for cross-architecture pairs and large teacher–student capacity gaps. On full ImageNet with a ResNet-34 teacher and ResNet-18 student, lower error is better; the reported top-1 and top-5 error rates are:

    Could not parse LaTeX table

    Thus, despite the relatively small capacity gap in this ImageNet comparison, the proposed feature-distillation recipe remains competitive with substantially more elaborate methods.

  10. Knowl 10 — The method improves object detection with a simple backbone-feature loss

    data/table

    For COCO2017 object detection, the paper distills the output backbone features of teacher and student detectors using the same projection, normalization, and LogSum principles as in classification. The method is evaluated under the ReviewKD training settings and also on YOLOv5. The reported detection metrics are:

    Could not parse LaTeX table

    For Faster R-CNN, the student uses MobileNetV2 and the teacher uses ResNet-50. The reported results are:

    Could not parse LaTeX table

    The proposed method improves the undistilled YOLOv5s and Faster R-CNN students and approaches ReviewKD for Faster R-CNN while using a simpler and cheaper distillation pipeline. It also exceeds FPGI on the YOLOv5 experiment, although ReviewKD remains stronger on the Faster R-CNN comparison.

Coverage note — No substantial contributed material was omitted; routine training schedules, implementation framework details, related work, and references were excluded because they do not add independent contribution beyond the extracted method, theory, ablations, and benchmark results.

References

  1. 1.Allen-Zhu, Z.; and Li, Y. 2023. Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.
  2. 2.Bardes, A.; Ponce, J.; and LeCun, Y. 2022a. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. ICLR.
  3. 3.Bardes, A.; Ponce, J.; and LeCun, Y. 2022b. VICRegL: Self-Supervised Learning of Local Visual Features. NeurIPS.
  4. 4.Beyer, L.; Zhai, X.; Royer, A.; Markeeva, L.; Anil, R.; and Kolesnikov, A. 2022. Knowledge distillation: A good teacher is patient and consistent. CVPR.
  5. 5.Carlucci, F. M.; D’Innocente, A.; Bucci, S.; Caputo, B.; and Tommasi, T. 2019. Domain Generalization by Solving Jigsaw Puzzles. CVPR.
  6. 6.Chen, D.; Mei, J.-P.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022a. Knowledge Distillation with the Reused Teacher Classifier. CVPR.
  7. 7.Chen, L.; Wang, D.; Gan, Z.; Liu, J.; Henao, R.; and Carin, L. 2020a. Wasserstein Contrastive Representation Distillation. CVPR.
  8. 8.Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021a. Distilling Knowledge via Knowledge Review. CVPR.
  9. 9.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020b. A simple framework for contrastive learning of visual representations. ICML.
  10. 10.Chen, X.; Fan, H.; Girshick, R.; and He, K. 2020c. Improved Baselines with Momentum Contrastive Learning. arXiv preprint.
  11. 11.Chen, X.; Xie, S.; and He, K. 2021. An Empirical Study of Training Self-Supervised Vision Transformers. ICCV.
  12. 12.Chen, Y.; Bian, Y.; Xiao, X.; Rong, Y.; Xu, T.; and Huang, J. 2021b. On Self-Distilling Graph Neural Network.
  13. 13.Chen, Y.; Wang, S.; Liu, J.; Xu, X.; de Hoog, F.; and Huang, Z. 2022b. Improved Feature Distillation via Projector Ensemble. NeurIPS.
  14. 14.Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020. Randaugment: Practical automated data augmentation with a reduced search space. CVPR Workshop.
  15. 15.Doersch, C.; Gupta, A.; and Efros, A. A. 2015. Unsupervised Visual Representation Learning by Context Prediction. ICCV.
  16. 16.Ermolov, A.; Siarohin, A.; Sangineto, E.; and Sebe, N. 2020. Whitening for Self-Supervised Representation Learning. ICML.
  17. 17.Everingham, M.; Gool, L. V.; Williams, C. K. I.; Winn, J.; and Zisserman., A. 2010. The pascal visual object classes (voc) challenge.
  18. 18.Gidaris, S.; Singh, P.; and Komodakis, N. 2018. Unsupervised Representation Learning by Predicting Image Rotations. ICLR.
  19. 19.Guo, J.; Chen, M.; Hu, Y.; Zhu, C.; He, X.; and Cai, D. 2020. Reducing the Teacher-Student Gap via Spherical Knowledge Distillation. arXiv preprint.
  20. 20.He, B.; and Ozay, M. 2022. Feature Kernel Distillation. ICLR.
  21. 21.He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022. Masked Autoencoders Are Scalable Vision Learners. CVPR.
  22. 22.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. CVPR.
  23. 23.Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. NeurIPS.
  24. 24.Huang, Z.; and Wang, N. 2017. Like What You Like: Knowledge Distill via Neuron Selectivity Transfer. arXiv preprint.
  25. 25.Joshi, C. K.; Liu, F.; Xun, X.; Lin, J.; and Foo, C.-S. 2021. On Representation Knowledge Distillation for Graph Neural Networks. arXiv preprint.
  26. 26.Krizhevsky, A. 2009. Learning Multiple Layers of Features from Tiny Images.
  27. 27.Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012. ImageNet Classification with Deep Convolutional Neural Networks. NeurIPS.
  28. 28.Lin, T. Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft COCO: Common objects in context. ECCV.
  29. 29.Liu, L.; Huang, Q.; Lin, S.; Xie, H.; Wang, B.; Chang, X.; and Liang, X. 2021. Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation. ICCV.
  30. 30.Liu, Y.; Chen, K.; Liu, C.; Qin, Z.; Luo, Z.; and Wang, J. 2019. Structured Knowledge Distillation for Semantic Segmentation. CVPR.
  31. 31.Ma, Y.; Chen, Y.; and Akata, Z. 2022. Distilling Knowledge from Self-Supervised Teacher by Embedding Graph Alignment. BMVC.
  32. 32.Matsubara, Y. 2020. torchdistill : A Modular, Configuration-Driven Framework for Knowledge Distillation.
  33. 33.Miles, R.; and Mikolajczyk, K. 2020. Cascaded channel pruning using hierarchical self-distillation. BMVC.
  34. 34.Miles, R.; Rodriguez, A. L.; and Mikolajczyk, K. 2022. Information Theoretic Representation Distillation. BMVC.
  35. 35.Miles, R.; Yucel, M. K.; Manganelli, B.; and Saa-Garriga, A. 2023. MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation. CVPR.
  36. 36.Mobahi, H.; Farajtabar, M.; and Bartlett, P. L. 2020. Self-Distillation Amplifies Regularization in Hilbert Space.
  37. 37.Navaneet, K. L.; Koohpayegani, S. A.; Tejankar, A.; and Pirsiavash, H. 2021. SimReg: Regression as a Simple Yet Effective Tool for Self-supervised Knowledge Distillation. BMVC.
  38. 38.Park, W.; Corp, K.; Kim, D.; and Lu, Y. 2019. Relational Knowledge Distillation. CVPR.
  39. 39.Peng, B.; Jin, X.; Li, D.; Zhou, S.; Wu, Y.; Liu, J.; Zhang, Z.; and Liu, Y. 2019. Correlation congruence for knowledge distillation. CVPR.
  40. 40.Ren, S.; Gao, Z.; Hua, T.; Xue, Z.; Tian, Y.; He, S.; and Zhao, H. 2022. Co-advise: Cross Inductive Bias Distillation. CVPR.
  41. 41.Romero, A.; Ballas, N.; Ebrahimi Kahou, S.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015. FitNets: Hints For Thin Deep Nets. ICLR.
  42. 42.Roth, K.; Milbich, T.; Ommer, B.; Cohen, J. P.; and Ghassemi, M. 2021. S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning.
  43. 43.Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2014. ImageNet Large Scale Visual Recognition Challenge. IJCV.
  44. 44.Tian, Y.; Krishnan, D.; and Isola, P. 2019. Contrastive representation distillation. ICLR.
  45. 45.Tishby, N. 2015. Deep Learning and the Information Bottleneck Principle. IEEE Information Theory Workshop (ITW).
  46. 46.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021. Training data-efficient image transformers & distillation through attention. PMLR.
  47. 47.Tung, F.; and Mori, G. 2019. Similarity-preserving knowledge distillation. ICCV.
  48. 48.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. NeurIPS.
  49. 49.Wang, T.; Yuan, L.; Zhang, X.; and Feng, J. 2019. Distilling Object Detectors with Fine-grained Feature Imitation. CVPR.
  50. 50.Xu, G.; Liu, Z.; Li, X.; and Loy, C. C. 2020. Knowledge Distillation Meets Self-supervision. ECCV.
  51. 51.Yang, C.; An, Z.; Cai, L.; and Xu, Y. 2021. Hierarchical Self-supervised Augmented Knowledge Distillation. IJCAI.
  52. 52.Yim, J. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. CVPR.
  53. 53.Zagoruyko, S.; and Komodakis, N. 2019. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In ICLR.
  54. 54.Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. ICML.
  55. 55.Zhang, L.; Song, J.; Gao, A.; Chen, J.; Bao, C.; and Ma, K. 2019. Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation.
  56. 56.Zhang, R.; Isola, P.; and Efros, A. A. 2016. Colorful Image Colorization.
  57. 57.Zhang, Y.; Lan, Z.; Dai, Y.; Zeng, F.; and Bai, Y. 2020. Prime-Aware Adaptive Distillation. ECCV.
  58. 58.Zhang, Z.; and Sabuncu, M. R. 2020. Self-Distillation as Instance-Specific Label Smoothing.
  59. 59.Zhao, B.; Song, R.; and Qiu, Y. 2022. Decoupled Knowledge Distillation. CVPR.
  60. 60.Zhu, J.; Tang, S.; Chen, D.; and Yu, S. 2021a. Complementary Relation Contrastive Distillation. CVPR.
  61. 61.Zhu, X.; Lyu, S.; Wang, X.; and Zhao, Q. 2021b. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. VisDrone ICCV workshop.

Citation

MLA
Miles, R., and K. Mikolajczyk. “Understanding the Role of the Projector in Knowledge Distillation”. arXiv, 2023, http://arxiv.org/abs/2303.11098v5.
APA
Miles, R., & Mikolajczyk, K. (2023). Understanding the Role of the Projector in Knowledge Distillation. arXiv. http://arxiv.org/abs/2303.11098v5
Chicago
Miles, R., and K. Mikolajczyk. 2023. “Understanding the Role of the Projector in Knowledge Distillation”. arXiv. http://arxiv.org/abs/2303.11098v5.
Harvard
Miles, R. and Mikolajczyk, K. (2023) “Understanding the Role of the Projector in Knowledge Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.11098v5.
Vancouver
1. Miles R, Mikolajczyk K (2023) Understanding the Role of the Projector in Knowledge Distillation. arXiv

BibTeX

@article{miles2023understanding,
  title = {Understanding the Role of the Projector in Knowledge Distillation},
  author = {Miles, Roy and Mikolajczyk, Krystian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.11098v5},
  eprint = {2303.11098}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF