Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?
Keshigeyan ChandrasegaranNgoc-Trung TranYunqing ZhaoNgai-Man Cheung
Resolves the contradiction over whether label smoothing aids or hurts knowledge distillation by identifying systematic representation diffusion at high temperatures, offering a practical guideline to distill from label-smoothed teachers using low temperatures.
In modern machine learning, engineers regularly use two standard techniques to improve deep neural network accuracy: label smoothing, which prevents a model from becoming overly confident by softening training labels, and knowledge distillation, which compresses knowledge from a larger teacher model into a smaller student model using a temperature scaling factor. However, previous studies produced sharp contradictions on whether these methods can work together. Early findings warned that label smoothing erases relative class information and impairs distillation, while subsequent research claimed that label smoothing actually improves distillation by expanding the separation between semantically similar categories. This conflict left practitioners uncertain about how to deploy these techniques without degrading model performance.
The article evaluates and resolves this contradiction by introducing the concept of systematic diffusion. Specifically, it demonstrates why the transfer temperature dictates whether an artificial intelligence model trained with label smoothing successfully compresses into a high-performing student model.
To establish these findings, the authors conducted extensive empirical evaluations across standard image classification benchmarks such as ImageNet-1K, fine-grained bird classification using CUB200-2011, and neural machine translation tasks across multiple languages. They tested multiple teacher-student network architectures, including ResNet, MobileNet, EfficientNet, ConvNeXt, and Transformers. The study introduced a quantitative diffusion index metric alongside high-dimensional feature visualizations to track how internal data representations shift during training.
The investigation revealed four major findings. First, when transferring knowledge from a smoothed teacher at elevated temperatures, the student model experiences systematic diffusion, meaning its learned internal representations blur specifically toward semantically similar categories rather than dispersing evenly. Second, this targeted blurring collapses the beneficial cluster separation created by label smoothing, causing student accuracy to drop steadily as temperature rises, such as a 5.05 percentage point drop in ResNet-18 ImageNet performance between low and moderate temperatures. Third, at a baseline low temperature of one, systematic diffusion remains minimal, allowing the student to successfully inherit the teacher's enlarged category margins and achieve superior accuracy. Fourth, detailed case studies proved that target smoothness alone cannot predict compression success, confirming that systematic diffusion within the student model is the primary governing mechanism.
These findings reconcile earlier contradictory literature: prior works reached opposite conclusions simply because they evaluated distillation at differing temperature regimes. For practical deployment, this removes guesswork, mitigates the operational risk of deploying degraded compact models on edge devices, and saves computational resources by eliminating extensive temperature parameter searches during training.
For engineering workflows, the article recommends pairing label-smoothed teacher models strictly with a low-temperature transfer setting of one. While these conclusions are supported with high confidence across 34 diverse benchmark experiments, the authors caution that validation metrics on very small class subsets may exhibit statistical variance and that applying excessive smoothing factors during initial teacher training can weaken the overall system.
- Paper: When Does Label Smoothing Help?, Rafael Müller et al. (2019). This paper discovered the initial conflict between label smoothing and knowledge distillation by demonstrating that label smoothing erases relative information across classes and impairs student distillation performance.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal work introduces the foundational knowledge distillation framework and temperature-scaled softmax mechanism that the source paper investigates and refines.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). This study details the calibration dynamics and temperature scaling methods in modern neural networks that underpin how label smoothing alters output confidence.
- Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). This paper analyzes capacity mismatches and empirical failure modes in knowledge distillation, providing key context on when student networks fail to mimic teacher distributions.
- Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). This work explores distillation breakdowns caused by capacity gaps between teacher and student models, motivating deeper investigations into distillation dynamics.
- Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). This paper examines the limitations of standard logit-based distillation objectives and proposes structural representation matching across network representations.
- Paper: Towards Understanding Knowledge Distillation, Mary Phuong et al. (2019). This paper provides theoretical foundations for gradient flow and representation transfer dynamics between teacher and student networks during distillation.
- Paper: Knowledge Distillation: A Survey, Jianping Gou et al. (2020). This comprehensive survey categorizes the distillation techniques, response-based losses, and training schemes that frame the source's empirical study.
- Paper: Logit Standardization in Knowledge Distillation, Shangquan Sun et al. (2024). Building on insights into distillation temperature and logit scale discrepancies, this work introduces logit standardization to dynamically adjust temperature and mitigate capacity gap artifacts.
- Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). This research extends logit-based distillation analysis by decoupling target and non-target class logit transfer, optimizing how dark knowledge is conveyed without classical softmax coupling issues.
- Paper: Understanding the Role of the Projector in Knowledge Distillation, Roy Miles et al. (2024). This paper expands the study of representation alignment in distillation by examining how projector layers and normalization schemes preserve singular values and feature geometry.
- Paper: One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation, Zhiwei Hao et al. (2023). This paper builds on the understanding of representation shifts in distillation to bridge structural and semantic representation gaps across heterogeneous network architectures.
