Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?

Keshigeyan ChandrasegaranNgoc-Trung TranYunqing ZhaoNgai-Man Cheung

article2022ICML57 citations

Resolves the contradiction over whether label smoothing aids or hurts knowledge distillation by identifying systematic representation diffusion at high temperatures, offering a practical guideline to distill from label-smoothed teachers using low temperatures.

Listen

In modern machine learning, engineers regularly use two standard techniques to improve deep neural network accuracy: label smoothing, which prevents a model from becoming overly confident by softening training labels, and knowledge distillation, which compresses knowledge from a larger teacher model into a smaller student model using a temperature scaling factor. However, previous studies produced sharp contradictions on whether these methods can work together. Early findings warned that label smoothing erases relative class information and impairs distillation, while subsequent research claimed that label smoothing actually improves distillation by expanding the separation between semantically similar categories. This conflict left practitioners uncertain about how to deploy these techniques without degrading model performance.

The article evaluates and resolves this contradiction by introducing the concept of systematic diffusion. Specifically, it demonstrates why the transfer temperature dictates whether an artificial intelligence model trained with label smoothing successfully compresses into a high-performing student model.

To establish these findings, the authors conducted extensive empirical evaluations across standard image classification benchmarks such as ImageNet-1K, fine-grained bird classification using CUB200-2011, and neural machine translation tasks across multiple languages. They tested multiple teacher-student network architectures, including ResNet, MobileNet, EfficientNet, ConvNeXt, and Transformers. The study introduced a quantitative diffusion index metric alongside high-dimensional feature visualizations to track how internal data representations shift during training.

The investigation revealed four major findings. First, when transferring knowledge from a smoothed teacher at elevated temperatures, the student model experiences systematic diffusion, meaning its learned internal representations blur specifically toward semantically similar categories rather than dispersing evenly. Second, this targeted blurring collapses the beneficial cluster separation created by label smoothing, causing student accuracy to drop steadily as temperature rises, such as a 5.05 percentage point drop in ResNet-18 ImageNet performance between low and moderate temperatures. Third, at a baseline low temperature of one, systematic diffusion remains minimal, allowing the student to successfully inherit the teacher's enlarged category margins and achieve superior accuracy. Fourth, detailed case studies proved that target smoothness alone cannot predict compression success, confirming that systematic diffusion within the student model is the primary governing mechanism.

These findings reconcile earlier contradictory literature: prior works reached opposite conclusions simply because they evaluated distillation at differing temperature regimes. For practical deployment, this removes guesswork, mitigates the operational risk of deploying degraded compact models on edge devices, and saves computational resources by eliminating extensive temperature parameter searches during training.

For engineering workflows, the article recommends pairing label-smoothed teacher models strictly with a low-temperature transfer setting of one. While these conclusions are supported with high confidence across 34 diverse benchmark experiments, the authors caution that validation metrics on very small class subsets may exhibit statistical variance and that applying excessive smoothing factors during initial teacher training can weaken the overall system.

arXiv: 2206.14532
  • Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). This paper discovered the initial conflict between label smoothing and knowledge distillation by demonstrating that label smoothing erases relative information across classes and impairs student distillation performance.
  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal work introduces the foundational knowledge distillation framework and temperature-scaled softmax mechanism that the source paper investigates and refines.
  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). This study details the calibration dynamics and temperature scaling methods in modern neural networks that underpin how label smoothing alters output confidence.
  • Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). This paper analyzes capacity mismatches and empirical failure modes in knowledge distillation, providing key context on when student networks fail to mimic teacher distributions.
  • Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). This work explores distillation breakdowns caused by capacity gaps between teacher and student models, motivating deeper investigations into distillation dynamics.
  • Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). This paper examines the limitations of standard logit-based distillation objectives and proposes structural representation matching across network representations.
  • Paper: Towards Understanding Knowledge Distillation, Mary Phuong et al. (2019). This paper provides theoretical foundations for gradient flow and representation transfer dynamics between teacher and student networks during distillation.
  • Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). This comprehensive survey categorizes the distillation techniques, response-based losses, and training schemes that frame the source's empirical study.
Cover for Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?

Abstract

This work investigates the compatibility between label smoothing (LS) and knowledge distillation (KD). Contemporary findings addressing this thesis statement take dichotomous standpoints: Müller et al. (2019); Shen et al. (2021b). Critically, there is no effort to understand and resolve these contradictory findings, leaving the primal question — to smooth or not to smooth a teacher network? — unanswered. The main contributions of our work are the discovery, analysis and validation of systematic diffusion as the missing concept which is instrumental in understanding and resolving these contradictory findings. This systematic diffusion essentially curtails the benefits of distilling from an LS-trained teacher, thereby rendering KD at increased temperatures ineffective. Our discovery is comprehensively supported by large-scale experiments, analyses and case studies including image classification, neural machine translation and compact student distillation tasks spanning across multiple datasets and teacher-student architectures. Based on our analysis, we suggest practitioners to use an LS-trained teacher with a low-temperature transfer to achieve high performance students. Code and models are available at https://keshik6.github.io/revisiting-ls-kd-compatibility/

Table of Contents

  • 1. Introduction
  • 2. Prerequisites
  • 3. A Closer Look at LS and KD compatibility
  • 4. Systematic Diffusion in Student
  • 5. Empirical Studies
  • 6. Extended Experiments
  • 7. Discussion and Conclusion
  • References
  • Supplementary Materials
  • A. Additional Penultimate Layer Visualizations
  • B. Additional Experiments / Analysis
  • C. Research Reproducibility Details
  • D. Standard Deviation for main paper experiments
  • E. Additional Discussion: Why this diffusion is systematic and not isotopic?
  • F. Algorithm for Projection and visualization of penultimate layer representations
  • G. Semantically similar / dissimilar classes
  • G.1. Method 1: Using standard, pre-defined ImageNet knowledge graph as a prior
  • G.2. Method 2: Using distance in the feature space to quantitatively define semantically similar / dissimilar classes
  • Consistency measurements between the 2 methods:
  • H. Case study: Smoothness of targets are insufficient to determine KD performance. Systematic diffusion is critical.
  • H.1. Case study at lower T with same degree of smoothness
  • H.2. Case study at moderately higher T with same degree of smoothness
  • H.3. Case study at extremely high T with same degree of smoothness
  • I. Class-wise accuracy for target classes
  • J. Additional Exploration of α and T
  • K. Alternative characterization of cluster distance
  • L. Sample images

Knowls

  1. Knowl 1 — Systematic diffusion determines LS–KD compatibility

    empirical result

    The paper’s central discovery is systematic diffusion: when a student is distilled from a label-smoothed teacher, increasing the knowledge-distillation temperature TT moves the student’s penultimate-layer representations for a class toward representations of semantically similar incorrect classes. This reduces the separation between semantically similar class clusters and cancels the class-separation benefit produced by label smoothing in the teacher. Consequently, label smoothing and knowledge distillation are empirically compatible at low transfer temperature, especially T=1T=1, but become incompatible as TT increases. This explains why earlier studies observed compatibility at low temperature and degraded distillation at higher temperature.

  2. Knowl 2 — Label-smoothing and knowledge-distillation objectives

    model/method

    For a KK-class classifier, let x∈Rh+1x\in\mathbb{R}^{h+1} be the augmented penultimate representation, wk∈Rh+1w_k\in\mathbb{R}^{h+1} the final-layer weight for class kk, yk∈{0,1}y_k\in\{0,1\} the one-hot target, and α∈[0,1]\alpha\in[0,1] the label-smoothing coefficient. Label smoothing replaces yky_k with

    ykLS=(1−α)yk+αK,y_k^{\mathrm{LS}}=(1-\alpha)y_k+\frac{\alpha}{K},

    and minimizes cross-entropy with predictions

    pk=exp⁡(x⊤wk)∑l=1Kexp⁡(x⊤wl),LLS=−∑k=1KykLSlog⁡pk.p_k=\frac{\exp(x^\top w_k)}{\sum_{l=1}^{K}\exp(x^\top w_l)},\qquad \mathcal{L}_{\mathrm{LS}}=-\sum_{k=1}^{K}y_k^{\mathrm{LS}}\log p_k.

    Knowledge distillation trains a student to match a teacher’s temperature-scaled output. For temperature T>0T>0, pk(T)p_k(T) is obtained by replacing each logit x⊤wkx^\top w_k by x⊤wk/Tx^\top w_k/T. With pt(T)p^t(T) and p(T)p(T) denoting teacher and student distributions, respectively, the distillation objective is

    LKD=(1−β)H(y,p)+(βT2)H(pt(T),p(T)),\mathcal{L}_{\mathrm{KD}}=(1-\beta)H(y,p)+(\beta T^2)H\bigl(p^t(T),p(T)\bigr),

    where H(a,b)=−∑kaklog⁡bkH(a,b)=-\sum_k a_k\log b_k and β\beta weights the soft-target term. The experiments set β=1\beta=1, so they isolate the effect of distilling the teacher distribution; the factor T2T^2 compensates approximately for the 1/T21/T^2 reduction in gradient magnitude caused by temperature scaling.

  3. Knowl 3 — Information erasure and distance enlargement in a smoothed teacher

    theoretical result

    Under label smoothing, the optimal target encourages the correct-class probability to approach 1−α+α/K1-\alpha+\alpha/K and every incorrect-class probability to approach α/K\alpha/K. Using the paper’s approximation that a final-layer logit is related to the squared Euclidean distance between a penultimate representation and its class template, label smoothing has two geometric effects: it pulls each representation toward the correct-class template and makes it approximately equidistant from incorrect-class templates. The resulting clusters are tighter, so relative information in the logits about similarities to other classes is erased. At the same time, semantically similar classes can become more separated in cluster-center distance. The paper’s experiments confirm both effects: label-smoothed teachers retain some nonuniform incorrect-class probabilities and enlarge the separation of difficult, semantically similar classes, but they also discard much of the finer-grained similarity information.

  4. Knowl 4 — Why the diffusion is systematic rather than isotropic

    theoretical result

    For a sample whose true class is k∗k^*, the teacher’s incorrect-class probabilities are not uniformly informative. A small number of semantically related classes can have relatively large probabilities pmltp^t_{m_l}, while the remaining incorrect classes have near-zero probabilities pmstp^t_{m_s}. Increasing TT makes the relatively large pmltp^t_{m_l} values closer to the correct-class probability, so the student is encouraged to place the sample’s penultimate representation closer to those semantically related classes. The near-zero pmstp^t_{m_s} values remain negligible even after temperature scaling and therefore have little effect. In the paper’s ImageNet analysis, the second-largest probability for standard-poodle samples, miniature poodle, was at least 100100 times larger than 976 of the other 999 incorrect-class probabilities; analogous concentration occurred for golden retriever/Labrador retriever and thunder snake/ringneck snake. Thus, temperature scaling produces directional diffusion toward a few semantically similar classes rather than uniform spreading in all directions.

  5. Knowl 5 — Diffusion index for measuring directional representation change

    equation

    Let cπ(T)c_\pi(T) be the centroid of the student’s penultimate representations for target class π\pi at temperature TT. Let S1S_1 be the set of semantically similar classes and S2S_2 the set of semantically dissimilar classes, and let ck(T)c_k(T) be the centroid of class kk. The normalized relative distance from π\pi to class k∈S1∪S2k\in S_1\cup S_2 is

    dT(π,k)=∥cπ(T)−ck(T)∥22Rπ(T),Rπ(T)=∑p∈S1∥cπ(T)−cp(T)∥22+∑q∈S2∥cπ(T)−cq(T)∥22.d_T(\pi,k)=\frac{\lVert c_\pi(T)-c_k(T)\rVert_2^2}{R_\pi(T)}, \qquad R_\pi(T)=\sum_{p\in S_1}\lVert c_\pi(T)-c_p(T)\rVert_2^2+\sum_{q\in S_2}\lVert c_\pi(T)-c_q(T)\rVert_2^2.

    For a nonempty set S⊆S1∪S2S\subseteq S_1\cup S_2, the diffusion index between temperatures T1T_1 and T2T_2 is

    η(T1,T2;π,S)=1∣S∣∑k∈SdT2(π,k)−dT1(π,k)dT1(π,k).\eta(T_1,T_2;\pi,S)=\frac{1}{|S|}\sum_{k\in S}\frac{d_{T_2}(\pi,k)-d_{T_1}(\pi,k)}{d_{T_1}(\pi,k)}.

    A negative η(T1,T2;π,S1)\eta(T_1,T_2;\pi,S_1) means that the target becomes relatively closer to semantically similar classes, whereas a positive η(T1,T2;π,S2)\eta(T_1,T_2;\pi,S_2) means that it becomes relatively farther from semantically dissimilar classes. The paper uses T1=1T_1=1 and T2=3T_2=3 and finds this negative/positive sign pattern across nearly all analyzed target classes, training samples, and validation samples.

  6. Knowl 6 — Qualitative evidence across representations and architectures

    empirical result

    Penultimate-layer visualizations consistently show three effects. First, applying label smoothing to a ResNet-50 teacher produces tighter class clusters, erases part of the logit similarity information, and increases the central separation of semantically similar classes. Second, a student distilled from that teacher inherits the increased separation at low temperature. Third, increasing the distillation temperature causes the student clusters of semantically similar classes to overlap or move toward one another, while representations of a semantically unrelated class are displaced in a different direction. The same pattern appears for ImageNet-1K poodle classes versus submarine, for CUB200-2011 shrike classes versus black-footed albatross, and across ResNet-18, ResNet-50, EfficientNet-B0, and ConvNeXt-T students. The effect is smaller for a powerful student than for a compact one, but remains visible.

  7. Knowl 7 — ImageNet and CUB200-2011 accuracy evidence

    data/table

    The main classification experiments use a ResNet-50 teacher, compare an ordinarily trained teacher (α=0\alpha=0) with a label-smoothed teacher (α=0.1\alpha=0.1), and report student top-1/top-5 test accuracy. Each entry is top-1 / top-5 accuracy.

    Could not parse LaTeX table

    At T=1T=1, students distilled from the label-smoothed teacher are often better than those distilled from the ordinary teacher, supporting the distance-enlargement benefit. As TT increases, however, the label-smoothed-teacher results degrade consistently, with especially large drops for ResNet-18: on ImageNet, top-1 accuracy falls from 71.61671.616 at T=1T=1 to 66.57066.570 at T=3T=3; on CUB200-2011, it falls from 80.94680.946 to 78.19678.196 and then to 67.16167.161 at T=64T=64. The ordinary teacher does not exhibit the same systematic high-temperature degradation.

  8. Knowl 8 — Compact-student and translation experiments generalize the finding

    data/table

    The high-temperature failure of distillation from a label-smoothed teacher is not restricted to standard ResNet students or image classification. With a ResNet-50 teacher and a MobileNetV2 student on CUB200-2011, the top-1/top-5 results were:

    Could not parse LaTeX table

    For English-to-German neural machine translation on IWSLT, using a Transformer teacher and Transformer student, BLEU scores were:

    Could not parse LaTeX table

    The same trend was reported for English-to-Russian translation, EfficientNet-B0 and ConvNeXt-T students, and an advanced feature-distillation method. For example, with an EfficientNet-B0 student on ImageNet-1K, the label-smoothed-teacher result fell from 69.906/89.28469.906/89.284 at T=1T=1 to 58.182/83.91858.182/83.918 at T=3T=3; with ConvNeXt-T on CUB200-2011 it fell from 86.866/97.37786.866/97.377 to 83.638/97.13583.638/97.135. In contrast, higher temperatures sometimes improved students distilled from teachers without label smoothing.

  9. Knowl 9 — Target entropy cannot explain distillation performance

    empirical result

    The paper tests whether the smoothness of teacher targets, measured by average entropy H(p)=−∑ipiln⁡piH(p)=-\sum_i p_i\ln p_i, is sufficient to predict student performance. On CUB200-2011, the maximum entropy is ln⁡(200)≈5.298\ln(200)\approx5.298. For an ordinarily trained ResNet-50 teacher (α=0\alpha=0), average entropy at T=1T=1, 1.4813751.481375, 22, 33, 5.6385.638, and 6464 was respectively 0.1840.184, 0.8880.888, 2.2462.246, 4.1604.160, 5.1185.118, and 5.2985.298. For a label-smoothed teacher (α=0.1\alpha=0.1), it was 0.8880.888, 3.2253.225, 4.5504.550, 5.1185.118, 5.2695.269, and 5.2985.298.

    Matching entropy does not match student quality. At entropy 0.8880.888, label-smoothed distillation at T=1T=1 produced 83.742/96.77883.742/96.778 for a ResNet-50 student and 80.946/95.31280.946/95.312 for a ResNet-18 student, whereas ordinary-teacher distillation at T=1.481375T=1.481375 produced 82.603/96.49682.603/96.496 and 80.808/95.54780.808/95.547, respectively. At entropy 5.1185.118, label-smoothed distillation at T=3T=3 produced 78.196/95.21378.196/95.213 for ResNet-18 and 78.961/95.30678.961/95.306 for MobileNetV2, compared with 78.719/95.47878.719/95.478 and 79.341/95.46179.341/95.461 from the ordinary teacher at T=5.638T=5.638. At entropy 5.2985.298, both teacher types at T=64T=64 were equally close to uniform, but label-smoothed distillation still produced much worse results: 67.161/93.06267.161/93.062 versus 73.611/94.52973.611/94.529 for ResNet-18, 77.206/95.81277.206/95.812 versus 79.784/95.92779.784/95.927 for ResNet-50, and 70.435/93.49470.435/93.494 versus 75.441/94.70275.441/94.702 for MobileNetV2. These seven counterexamples show that student-side systematic diffusion, not target entropy alone, is the more informative explanation.

  10. Knowl 10 — Practical operating rule and empirical scope

    limitation

    The paper recommends using a label-smoothed teacher with a low-temperature transfer, specifically T=1T=1, when high-performing students are desired. The recommendation is empirical rather than a universal theorem: higher temperatures can be useful when the teacher was trained without label smoothing, while label smoothing changes the student-side effect of temperature scaling. The conclusion also holds for stronger smoothing: in CUB200-2011 MobileNetV2 distillation with α=0.2\alpha=0.2, top-1/top-5 accuracy decreased from 81.498/95.89281.498/95.892 at T=1T=1 to 79.997/95.59979.997/95.599 at T=2T=2, 76.959/95.20276.959/95.202 at T=3T=3, and 63.738/91.99263.738/91.992 at T=64T=64. The paper therefore characterizes the result as specific to the interaction of label-smoothed teachers, temperature-scaled KD, and student representation learning; it does not claim that high temperature is intrinsically harmful for all teachers or all distillation methods.

Coverage note — Detailed projection pseudocode, per-class accuracy tables, standard deviations, and the auxiliary pairwise-distance version of the diffusion index were omitted because they are supporting diagnostics rather than separate load-bearing contributions.

References

  1. 1.Abbasi Koohpayegani, S., Tejankar, A., and Pirsiavash, H. Compress: Self-supervised learning by compressing representations. Advances in Neural Information Processing Systems, 33:12980–12992, 2020.
  2. 2.Arani, E., Sarfraz, F., and Zonooz, B. Noise as a resource for learning in knowledge distillation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 3129–3138, January 2021.
  3. 3.Chandrasegaran, K., Tran, N.-T., and Cheung, N.-M. A closer look at fourier spectrum discrepancies for cnn-generated images detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7200–7209, June 2021.
  4. 4.Chiu, C.-C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, E., Jaitly, N., Li, B., Chorowski, J., and Bacchiani, M. State-of-the-art speech recognition with sequence-to-sequence models. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4774–4778, 2018. doi: 10.1109/ICASSP.2018.8462105.
  5. 5.Chorowski, J. and Jaitly, N. Towards better decoding and language model integration in sequence to sequence models. In Proc. Interspeech 2017, pp. 523–527, 2017. doi: 10.21437/Interspeech.2017-343. URL http://dx.doi.org/10.21437/Interspeech.2017-343.
  6. 6.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009.
  7. 7.Dzanic, T., Shah, K., and Witherden, F. Fourier spectrum discrepancies in deep network generated images. Advances in neural information processing systems, 33:3022–3032, 2020.
  8. 8.Fang, Z., Wang, J., Wang, L., Zhang, L., Yang, Y., and Liu, Z. {SEED}: Self-supervised distillation for visual representation. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=AHm3dbp7D1D.
  9. 9.Fellbaum, C. (ed.). WordNet: An Electronic Lexical Database. Language, Speech, and Communication. MIT Press, Cambridge, MA, 1998. ISBN 978-0-262-06197-1.
  10. 10.Fu, Y., Chen, W., Wang, H., Li, H., Lin, Y., and Wang, Z. Autogan-distiller: searching to compress generative adversarial networks. In Proceedings of the 37th International Conference on Machine Learning, pp. 3292–3303, 2020.
  11. 11.He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M. Bag of tricks for image classification with convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  12. 12.Heo, B., Kim, J., Yun, S., Park, H., Kwak, N., and Choi, J. Y. A comprehensive overhaul of feature distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1921–1930, 2019.
  13. 13.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2015. URL http://arxiv.org/abs/1503.02531.
  14. 14.Hu, M., Peng, Y., Wei, F., Huang, Z., Li, D., Yang, N., and Zhou, M. Attention-guided answer distillation for machine reading comprehension. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2077–2086, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1232. URL https://www.aclweb.org/anthology/D18-1232.
  15. 15.Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, z. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/093f65e080a295f8076b1c5722a46aa2-Paper.pdf.
  16. 16.Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 4163–4174, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.372. URL https://www.aclweb.org/anthology/2020.findings-emnlp.372.
  17. 17.Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D. Supervised contrastive learning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 18661–18673. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/d89a66c7c80a29b1bdbab0f2a1a94af8-Paper.pdf.
  18. 18.Kwon, K., Na, H., Lee, H., and Kim, N. S. Adaptive knowledge distillation based on entropy. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7409–7413, 2020. doi: 10.1109/ICASSP40776.2020.9054698.
  19. 19.Li, C., Peng, J., Yuan, L., Wang, G., Liang, X., Lin, L., and Chang, X. Block-wisely supervised neural architecture search with knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1989–1998, 2020a.
  20. 20.Li, M., Lin, J., Ding, Y., Liu, Z., Zhu, J.-Y., and Han, S. Gan compression: Efficient architectures for interactive conditional gans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5284–5294, 2020b.
  21. 21.Lim, S. K., Loo, Y., Tran, N. T., Cheung, N. M., Roig, G., and Elovici, Y. Doping: Generative data augmentation for unsupervised anomaly detection with gan. In 18th IEEE International Conference on Data Mining, ICDM 2018, pp. 1122–1127. Institute of Electrical and Electronics Engineers Inc., 2018.
  22. 22.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11976–11986, 2022.
  23. 23.Lopez-Paz, D., Scholkopf, B., Bottou, L., and Vapnik, V. Unifying distillation and privileged information. In International Conference on Learning Representations (ICLR), November 2016.
  24. 24.Lukasik, M., Bhojanapalli, S., Menon, A., and Kumar, S. Does label smoothing mitigate label noise? In III, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 6448–6458. PMLR, 13–18 Jul 2020. URL http://proceedings.mlr.press/v119/lukasik20a.html.
  25. 25.Mghabbar, I. and Ratnamogan, P. Building a multi-domain neural machine translation model using knowledge distillation. In Giacomo, G. D., Catala, A., Dilkina, B., Milano, M., Barro, S., Bugarín, A., and Lang, J. (eds.), ECAI 2020 - 24th European Conference on Artificial Intelligence, 29 August-8 September 2020, Santiago de Compostela, Spain, August 29 - September 8, 2020 - Including 10th Conference on Prestigious Applications of Artificial Intelligence (PAIS 2020), volume 325 of Frontiers in Artificial Intelligence and Applications, pp. 2116–2123. IOS Press, 2020. doi: 10.3233/FAIA200335. URL https://doi.org/10.3233/FAIA200335.
  26. 26.Muller, R., Kornblith, S., and Hinton, G. E. When does label smoothing help? In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/f1748d6b0fd9d439f71450117eba2725-Paper.pdf.
  27. 27.Nakashole, N. and Flauger, R. Knowledge distillation for bilingual dictionary induction. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2497–2506, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1264. URL https://www.aclweb.org/anthology/D17-1264.
  28. 28.Peng, Z., Li, Z., Zhang, J., Li, Y., Qi, G.-J., and Tang, J. Few-shot image recognition with knowledge transfer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.
  29. 29.Pereyra, G., Tucker, G., Chorowski, J., Kaiser, ., and Hinton, G. Regularizing neural networks by penalizing confident output distributions. 01 2017.
  30. 30.Perez, A., Sanguineti, V., Morerio, P., and Murino, V. Audio-visual model distillation using acoustic images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), March 2020.
  31. 31.Real, E., Aggarwal, A., Huang, Y., and Le, Q. V. Regularized evolution for image classifier architecture search. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):4780–4789, Jul. 2019. doi: 10.1609/aaai.v33i01.33014780. URL https://ojs.aaai.org/index.php/AAAI/article/view/4405.
  32. 32.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4510–4520, 2018.
  33. 33.Shen, P., Lu, X., Li, S., and Kawai, H. Knowledge distillation-based representation learning for short-utterance spoken language identification. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2674–2683, 2020.
  34. 34.Shen, Z., Liu, Z., Liu, Z., Savvides, M., Darrell, T., and Xing, E. Un-mix: Rethinking image mixtures for unsupervised visual representation learning, 2021a.
  35. 35.Shen, Z., Liu, Z., Xu, D., Chen, Z., Cheng, K.-T., and Savvides, M. Is label smoothing truly incompatible with knowledge distillation: An empirical study. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=PObuuGVrGaZ.
  36. 36.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2818–2826, 2016. doi: 10.1109/CVPR.2016.308.
  37. 37.Tan, M. and Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pp. 6105–6114. PMLR, 2019.
  38. 38.Tang, J., Shivanna, R., Zhao, Z., Lin, D., Singh, A., Chi, E. H., and Jain, S. Understanding and improving knowledge distillation, 2021.
  39. 39.Tran, N.-T., Tran, V.-H., Nguyen, N.-B., Nguyen, T.-K., and Cheung, N.-M. On data augmentation for gan training. IEEE Transactions on Image Processing, 30:1882–1897, 2021.
  40. 40.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  41. 41.Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011.
  42. 42.Wang, D., Gong, C., Li, M., Liu, Q., and Chandra, V. Alphanet: Improved training of supernets with alpha-divergence. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 10760–10771. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/wang21i.html.
  43. 43.Yu, C. and Pool, J. Self-supervised generative adversarial compression. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 8235–8246. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/5d79099fcdf499f12b79770834c0164a-Paper.pdf.
  44. 44.Yu, J., Jin, P., Liu, H., Bender, G., Kindermans, P.-J., Tan, M., Huang, T., Song, X., Pang, R., and Le, Q. Bignas: Scaling up neural architecture search with big single-stage models. In European Conference on Computer Vision, pp. 702–717. Springer, 2020.
  45. 45.Yuan, L., Tay, F. E., Li, G., Wang, T., and Feng, J. Revisiting knowledge distillation via label smoothing regularization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  46. 46.Zhang, M., Song, G., Zhou, H., and Liu, Y. Discriminability distillation in group representation learning. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M. (eds.), Computer Vision – ECCV 2020, pp. 1–19, Cham, 2020. Springer International Publishing.
  47. 47.Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. Learning transferable architectures for scalable image recognition. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8697–8710, 2018. doi: 10.1109/CVPR.2018.00907.

Citation

MLA
Chandrasegaran, K., et al. “Revisiting Label Smoothing and Knowledge Distillation Compatibility: What Was Missing?”. International Conference on Machine Learning, vol. 162, 2022, pp. 2890–916, https://proceedings.mlr.press/v162/chandrasegaran22a.html.
APA
Chandrasegaran, K., Tran, N.-T., Zhao, Y., & Cheung, N.-M. (2022). Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?. International Conference on Machine Learning, 162, 2890–2916. https://proceedings.mlr.press/v162/chandrasegaran22a.html
Chicago
Chandrasegaran, K., N.-T. Tran, Y. Zhao, and N.-M. Cheung. 2022. “Revisiting Label Smoothing and Knowledge Distillation Compatibility: What Was Missing?”. International Conference on Machine Learning 162: 2890–2916. https://proceedings.mlr.press/v162/chandrasegaran22a.html.
Harvard
Chandrasegaran, K. et al. (2022) “Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?”, International Conference on Machine Learning. PMLR, pp. 2890–2916. Available at: https://proceedings.mlr.press/v162/chandrasegaran22a.html.
Vancouver
1. Chandrasegaran K, Tran N-T, Zhao Y, Cheung N-M (2022) Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?. In: International Conference on Machine Learning. PMLR, pp 2890–2916

BibTeX

@InProceedings{pmlr-v162-chandrasegaran22a,
  title = 	 {Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?},
  author =       {Chandrasegaran, Keshigeyan and Tran, Ngoc-Trung and Zhao, Yunqing and Cheung, Ngai-Man},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {2890--2916},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/chandrasegaran22a/chandrasegaran22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/chandrasegaran22a.html},
  abstract = 	 {This work investigates the compatibility between label smoothing (LS) and knowledge distillation (KD). Contemporary findings addressing this thesis statement take dichotomous standpoints: Muller et al. (2019) and Shen et al. (2021b). Critically, there is no effort to understand and resolve these contradictory findings, leaving the primal question \text{-} to smooth or not to smooth a teacher network? \text{-} unanswered. The main contributions of our work are the discovery, analysis and validation of systematic diffusion as the missing concept which is instrumental in understanding and resolving these contradictory findings. This systematic diffusion essentially curtails the benefits of distilling from an LS-trained teacher, thereby rendering KD at increased temperatures ineffective. Our discovery is comprehensively supported by large-scale experiments, analyses and case studies including image classification, neural machine translation and compact student distillation tasks spanning across multiple datasets and teacher-student architectures. Based on our analysis, we suggest practitioners to use an LS-trained teacher with a low-temperature transfer to achieve high performance students. Code and models are available at https://keshik6.github.io/revisiting-ls-kd-compatibility/}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/