Logit Standardization in Knowledge Distillation
Shangquan SunWenqi RenJingzhi LiRui WangXiaochun Cao
Proposes a plug-and-play Z-score logit standardization pre-processing step that removes the restrictive requirement for exact logit magnitude matching between teacher and student models, improving the performance of logit-based knowledge distillation across standard vision benchmarks.
Deploying modern deep neural networks to resource-constrained environments requires compressing large, accurate models into smaller, efficient architectures. Knowledge distillation is a standard technique for this task, where a lightweight "student" model learns from the output probability distributions of a cumbersome "teacher" model using a softening temperature parameter. However, standard methods force the teacher and student to share a fixed temperature, inadvertently imposing an artificial constraint that demands the student's output values match the teacher's exact scale and variance. Because lightweight models naturally lack the capacity to reproduce the broad output ranges of larger models, this implicit requirement degrades student learning and creates misleading performance assessments during training.
The article aims to resolve this limitation by establishing the theoretical independence of teacher and student temperatures and introducing an adaptive standardization pre-processing step that allows students to learn essential relative prediction relationships without requiring an exact magnitude match.
To evaluate this solution, the authors derive the mathematical foundations of distillation temperatures using information theory and entropy maximization, proving that temperatures can differ between models and across samples. Guided by this insight, they develop a lightweight Z-score standardization pre-process that normalizes model outputs to have a zero mean and a bounded variance scaled by the logit standard deviation. The approach was tested across standard image classification benchmarks, including CIFAR-100 and ImageNet, spanning diverse network architectures such as ResNet, MobileNet, VGG, and Wide ResNet. The authors evaluated the technique as a modular enhancement across multiple established logit-based distillation frameworks, repeating CIFAR-100 experiments across four trials for statistical reliability.
The findings confirm that standardizing output distributions yields consistent, significant performance improvements across all tested configurations. On CIFAR-100, adding the pre-processing step improved baseline distillation accuracy across every tested model pair, achieving notable gains of up to 3.14 to 3.29 percentage points on challenging architectural pairings. On the large-scale ImageNet dataset, the method improved top-1 accuracy by up to 1.76 percentage points when distilling a ResNet-50 into a MobileNet-V1. Furthermore, the pre-processing method enabled basic distillation to achieve performance on par with complex, computationally heavy feature-based distillation techniques. It also successfully mitigated the "capacity gap" problem, allowing compact students to effectively learn from much larger, more accurate teacher models.
These results demonstrate that student networks primarily need to capture the relative rankings among classes rather than matching the absolute numerical output scale of the teacher. In practical terms, this simple pre-processing method eliminates the need for computationally intensive intermediate feature matching, lowering the engineering complexity and computational overhead of training lightweight models for real-world computer vision tasks.
Practitioners and engineering teams deploying neural networks on edge or mobile hardware should adopt this Z-score standardization pre-processing step as a drop-in replacement within existing logit-based distillation pipelines. Distillation loss weights should be set higher relative to hard classification labels to maximize the transfer of relative class relationships. Before broad deployment, organizations should conduct targeted validation pilots across specific target architectures, and future work should extend validation to complex multi-modal domains and non-vision tasks.
Confidence in these findings is high for visual recognition tasks given the consistent empirical improvements across diverse architectures and benchmark scales. The primary limitation is that empirical evidence in the main text centers on image classification benchmarks, meaning practitioners in other domains, such as generative modeling or natural language processing, should validate the approach within their specific operational pipelines.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Introduces the foundational knowledge distillation paradigm of matching softened output probabilities via a shared temperature scaling factor, which the source directly analyzes and modifies.
- Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). Establishes modern logit-based distillation by decoupling target and non-target class information, providing a key benchmark and motivation for standardizing output distributions.
- Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). Identifies the capacity gap problem where small students fail to mimic large teachers due to output distribution scale mismatches, a central issue resolved by logit standardization.
- Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). Demonstrates the degradation that occurs during distillation across large teacher-student capacity gaps, which the source overcomes via scale-invariant standardization.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Analyzes temperature scaling and confidence calibration in modern neural network logits, informing the source's theoretical treatment of logit variances and temperatures.
- Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). Provides a prominent baseline for intermediate feature representation distillation, which the source's lightweight logit standardization technique matches without heavy computational overhead.
- Paper: Understanding the Role of the Projector in Knowledge Distillation, Roy Miles et al. (2024). Extends the exploration of normalization and projection mechanics in distillation to intermediate feature representations rather than output logits.
