Logit Standardization in Knowledge Distillation

Shangquan SunWenqi RenJingzhi LiRui WangXiaochun Cao

article2024CVPR193 citations

Proposes a plug-and-play Z-score logit standardization pre-processing step that removes the restrictive requirement for exact logit magnitude matching between teacher and student models, improving the performance of logit-based knowledge distillation across standard vision benchmarks.

Listen

Deploying modern deep neural networks to resource-constrained environments requires compressing large, accurate models into smaller, efficient architectures. Knowledge distillation is a standard technique for this task, where a lightweight "student" model learns from the output probability distributions of a cumbersome "teacher" model using a softening temperature parameter. However, standard methods force the teacher and student to share a fixed temperature, inadvertently imposing an artificial constraint that demands the student's output values match the teacher's exact scale and variance. Because lightweight models naturally lack the capacity to reproduce the broad output ranges of larger models, this implicit requirement degrades student learning and creates misleading performance assessments during training.

The article aims to resolve this limitation by establishing the theoretical independence of teacher and student temperatures and introducing an adaptive standardization pre-processing step that allows students to learn essential relative prediction relationships without requiring an exact magnitude match.

To evaluate this solution, the authors derive the mathematical foundations of distillation temperatures using information theory and entropy maximization, proving that temperatures can differ between models and across samples. Guided by this insight, they develop a lightweight Z-score standardization pre-process that normalizes model outputs to have a zero mean and a bounded variance scaled by the logit standard deviation. The approach was tested across standard image classification benchmarks, including CIFAR-100 and ImageNet, spanning diverse network architectures such as ResNet, MobileNet, VGG, and Wide ResNet. The authors evaluated the technique as a modular enhancement across multiple established logit-based distillation frameworks, repeating CIFAR-100 experiments across four trials for statistical reliability.

The findings confirm that standardizing output distributions yields consistent, significant performance improvements across all tested configurations. On CIFAR-100, adding the pre-processing step improved baseline distillation accuracy across every tested model pair, achieving notable gains of up to 3.14 to 3.29 percentage points on challenging architectural pairings. On the large-scale ImageNet dataset, the method improved top-1 accuracy by up to 1.76 percentage points when distilling a ResNet-50 into a MobileNet-V1. Furthermore, the pre-processing method enabled basic distillation to achieve performance on par with complex, computationally heavy feature-based distillation techniques. It also successfully mitigated the "capacity gap" problem, allowing compact students to effectively learn from much larger, more accurate teacher models.

These results demonstrate that student networks primarily need to capture the relative rankings among classes rather than matching the absolute numerical output scale of the teacher. In practical terms, this simple pre-processing method eliminates the need for computationally intensive intermediate feature matching, lowering the engineering complexity and computational overhead of training lightweight models for real-world computer vision tasks.

Practitioners and engineering teams deploying neural networks on edge or mobile hardware should adopt this Z-score standardization pre-processing step as a drop-in replacement within existing logit-based distillation pipelines. Distillation loss weights should be set higher relative to hard classification labels to maximize the transfer of relative class relationships. Before broad deployment, organizations should conduct targeted validation pilots across specific target architectures, and future work should extend validation to complex multi-modal domains and non-vision tasks.

Confidence in these findings is high for visual recognition tasks given the consistent empirical improvements across diverse architectures and benchmark scales. The primary limitation is that empirical evidence in the main text centers on image classification benchmarks, meaning practitioners in other domains, such as generative modeling or natural language processing, should validate the approach within their specific operational pipelines.

arXiv: 2403.01427
  • Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Introduces the foundational knowledge distillation paradigm of matching softened output probabilities via a shared temperature scaling factor, which the source directly analyzes and modifies.
  • Paper: Decoupled Knowledge Distillation, Borui Zhao et al. (2022). Establishes modern logit-based distillation by decoupling target and non-target class information, providing a key benchmark and motivation for standardizing output distributions.
  • Paper: On the Efficacy of Knowledge Distillation, Jang Hyun Cho et al. (2019). Identifies the capacity gap problem where small students fail to mimic large teachers due to output distribution scale mismatches, a central issue resolved by logit standardization.
  • Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). Demonstrates the degradation that occurs during distillation across large teacher-student capacity gaps, which the source overcomes via scale-invariant standardization.
  • Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Analyzes temperature scaling and confidence calibration in modern neural network logits, informing the source's theoretical treatment of logit variances and temperatures.
  • Paper: Contrastive Representation Distillation, Yonglong Tian et al. (2020). Provides a prominent baseline for intermediate feature representation distillation, which the source's lightweight logit standardization technique matches without heavy computational overhead.
Cover for Logit Standardization in Knowledge Distillation

Abstract

Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relations from teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic distillation evaluation; nonetheless, this challenge is successfully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background and Notation
  • 4. Methodology
  • 4.1. Irrelevance between Temperatures
  • 4.1.1 Derivation of softmax in Classification
  • 4.1.2 Derivation of softmax in KD
  • 4.2. Drawbacks of Shared Temperatures
  • 4.3. Logit Standardization
  • 4.3.1 Toy Case
  • 5. Experiments
  • 5.1. Main Results
  • 5.2. Extensions
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Logit Z-Score Standardization for Knowledge Distillation

    model/method

    In logit-based knowledge distillation for classification over KK classes, logit vectors x∈RKx \in \mathbb{R}^K (predicted by either a teacher model fTf_T or a student model fSf_S) are pre-processed via a weighted ZZ-score standardization prior to applying the softmax activation.

    Let xˉ\bar{x} and σ(x)\sigma(x) denote the sample mean and standard deviation of xx:

    xˉ=1K∑k=1Kx(k),σ(x)=1K∑k=1K(x(k)−xˉ)2\bar{x} = \frac{1}{K}\sum_{k=1}^K x^{(k)}, \quad \sigma(x) = \sqrt{\frac{1}{K}\sum_{k=1}^K \left(x^{(k)} - \bar{x}\right)^2}

    The weighted ZZ-score transformation parameterized by a base temperature τ>0\tau > 0 is defined component-wise as:

    Z(x;τ)(k)=x(k)−xˉσ(x)⋅τ\mathcal{Z}(x; \tau)^{(k)} = \frac{x^{(k)} - \bar{x}}{\sigma(x) \cdot \tau}

    The standardized probability distribution is then obtained by:

    q(x;xˉ,σ(x))(k)=exp⁡(Z(x;τ)(k))∑m=1Kexp⁡(Z(x;τ)(m))q(x; \bar{x}, \sigma(x))^{(k)} = \frac{\exp\left(\mathcal{Z}(x; \tau)^{(k)}\right)}{\sum_{m=1}^K \exp\left(\mathcal{Z}(x; \tau)^{(m)}\right)}

    This standardization guarantees four mathematical properties:

    1. Zero Mean: 1K∑k=1KZ(x;τ)(k)=0\frac{1}{K}\sum_{k=1}^K \mathcal{Z}(x; \tau)^{(k)} = 0.
    2. Definite Standard Deviation: σ(Z(x;τ))=1τ\sigma(\mathcal{Z}(x; \tau)) = \frac{1}{\tau}, which aligns both teacher and student logits to distributions with zero mean and identical variance 1/τ21/\tau^2.
    3. Monotonicity: Z(x;τ)\mathcal{Z}(x; \tau) is a strictly monotonically increasing linear transformation since σ(x)⋅τ>0\sigma(x) \cdot \tau > 0, preserving the relative rank order of logits: x(i)≤x(j)  ⟺  Z(x;τ)(i)≤Z(x;τ)(j)x^{(i)} \le x^{(j)} \iff \mathcal{Z}(x; \tau)^{(i)} \le \mathcal{Z}(x; \tau)^{(j)}.
    4. Boundedness: Standardized logits are strictly bounded within [−K−1τ,K−1τ]\left[-\frac{\sqrt{K-1}}{\tau}, \frac{\sqrt{K-1}}{\tau}\right], preventing extreme exponential values during softmax computation.
  2. Knowl 2 — Knowledge Distillation with Logit Standardization

    algorithm

    Knowledge distillation with logit standardization applies ZZ-score pre-processing independently to teacher and student logit vectors before computing the distillation loss, while training on ground-truth hard labels using standard softmax on the raw student logits.

    Input: Training dataset D={(xn,yn)}n=1ND = \{(x_n, y_n)\}_{n=1}^N with KK classes, teacher network fTf_T, student network fSf_S, base temperature τ\tau, distillation loss criterion LKD\mathcal{L}_{KD} (e.g., Kullback-Leibler divergence), cross-entropy loss weight λCE\lambda_{CE}, distillation loss weight λKD\lambda_{KD}.
    Output: Distilled student network fSf_S.
    for each mini-batch of pairs (xn,yn)∈D(x_n, y_n) \in D do
        vn←fT(xn)v_n \leftarrow f_T(x_n)
        zn←fS(xn)z_n \leftarrow f_S(x_n)
        vˉn←1K∑k=1Kvn(k)\bar{v}_n \leftarrow \frac{1}{K}\sum_{k=1}^K v_n^{(k)}
        σ(vn)←1K∑k=1K(vn(k)−vˉn)2\sigma(v_n) \leftarrow \sqrt{\frac{1}{K}\sum_{k=1}^K (v_n^{(k)} - \bar{v}_n)^2}
        Z(vn;τ)←vn−vˉnσ(vn)⋅τ\mathcal{Z}(v_n; \tau) \leftarrow \frac{v_n - \bar{v}_n}{\sigma(v_n) \cdot \tau}
        zˉn←1K∑k=1Kzn(k)\bar{z}_n \leftarrow \frac{1}{K}\sum_{k=1}^K z_n^{(k)}
        σ(zn)←1K∑k=1K(zn(k)−zˉn)2\sigma(z_n) \leftarrow \sqrt{\frac{1}{K}\sum_{k=1}^K (z_n^{(k)} - \bar{z}_n)^2}
        Z(zn;τ)←zn−zˉnσ(zn)⋅τ\mathcal{Z}(z_n; \tau) \leftarrow \frac{z_n - \bar{z}_n}{\sigma(z_n) \cdot \tau}
        q(vn)←softmax(Z(vn;τ))q(v_n) \leftarrow \text{softmax}(\mathcal{Z}(v_n; \tau))
        q(zn)←softmax(Z(zn;τ))q(z_n) \leftarrow \text{softmax}(\mathcal{Z}(z_n; \tau))
        q′(zn)←softmax(zn)q'(z_n) \leftarrow \text{softmax}(z_n)
        Ltotal←λCELCE(yn,q′(zn))+λKDτ2LKD(q(vn),q(zn))\mathcal{L}_{total} \leftarrow \lambda_{CE} \mathcal{L}_{CE}(y_n, q'(z_n)) + \lambda_{KD} \tau^2 \mathcal{L}_{KD}(q(v_n), q(z_n))
        Update parameters of fSf_S by stochastic gradient descent to minimize Ltotal\mathcal{L}_{total}
    end for

    The factor τ2\tau^2 scales the distillation loss to match gradient magnitudes across different base temperatures.

  3. Knowl 3 — Implicit Logit Shift and Variance Matching Shackles in Shared-Temperature KD

    theoretical result

    In generalized logit distillation, let the softened probability distributions for student logits zn∈RKz_n \in \mathbb{R}^K and teacher logits vn∈RKv_n \in \mathbb{R}^K be parameterized by shift and scale hyperparameters (aS,bS)(a_S, b_S) and (aT,bT)(a_T, b_T):

    q(zn;aS,bS)(k)=exp⁡((zn(k)−aS)/bS)∑m=1Kexp⁡((zn(m)−aS)/bS),q(vn;aT,bT)(k)=exp⁡((vn(k)−aT)/bT)∑m=1Kexp⁡((vn(m)−aT)/bT)q(z_n; a_S, b_S)^{(k)} = \frac{\exp\left((z_n^{(k)} - a_S)/b_S\right)}{\sum_{m=1}^K \exp\left((z_n^{(m)} - a_S)/b_S\right)}, \quad q(v_n; a_T, b_T)^{(k)} = \frac{\exp\left((v_n^{(k)} - a_T)/b_T\right)}{\sum_{m=1}^K \exp\left((v_n^{(m)} - a_T)/b_T\right)}

    When a student is fully distilled such that q(zn;aS,bS)=q(vn;aT,bT)q(z_n; a_S, b_S) = q(v_n; a_T, b_T), the logit differences across all class index pairs (i,j)(i, j) satisfy:

    zn(i)−zn(j)bS=vn(i)−vn(j)bT\frac{z_n^{(i)} - z_n^{(j)}}{b_S} = \frac{v_n^{(i)} - v_n^{(j)}}{b_T}

    Averaging across j∈{1,…,K}j \in \{1, \dots, K\} yields:

    zn(i)−zˉnbS=vn(i)−vˉnbT\frac{z_n^{(i)} - \bar{z}_n}{b_S} = \frac{v_n^{(i)} - \bar{v}_n}{b_T}

    where zˉn=1K∑m=1Kzn(m)\bar{z}_n = \frac{1}{K}\sum_{m=1}^K z_n^{(m)} and vˉn=1K∑m=1Kvn(m)\bar{v}_n = \frac{1}{K}\sum_{m=1}^K v_n^{(m)}. Summing the squared relations across ii gives the standard deviation ratio:

    σ(zn)σ(vn)=bSbT\frac{\sigma(z_n)}{\sigma(v_n)} = \frac{b_S}{b_T}

    Under the conventional setting where teacher and student share a fixed global temperature (bS=bT=Tb_S = b_T = \mathcal{T} and aS=aT=0a_S = a_T = 0), distillation imposes two restrictive constraints:

    1. Logit Shift Constraint: zn(i)=vn(i)+(zˉn−vˉn)z_n^{(i)} = v_n^{(i)} + (\bar{z}_n - \bar{v}_n), compelling the student to strictly match the absolute logit values of the teacher up to an additive constant.
    2. Variance Match Constraint: σ(zn)=σ(vn)\sigma(z_n) = \sigma(v_n), compelling the student's logit variance to match that of the teacher.

    Because capacity-limited students struggle to match the dynamic range and variance of cumbersome teachers, these two constraints degrade distillation. Setting aS=zˉna_S = \bar{z}_n, aT=vˉna_T = \bar{v}_n, bS=τσ(zn)b_S = \tau \sigma(z_n), and bT=τσ(vn)b_T = \tau \sigma(v_n) decouples the student's logit scale from the teacher's while preserving relative logit rankings.

  4. Knowl 4 — Entropy Maximization Derivation of Softmax and Temperature Decoupling in KD

    theoretical result

    The softmax distribution in knowledge distillation can be derived from the principle of maximum entropy in information theory, revealing that teacher and student temperatures do not need to be shared or globally constant.

    For a student predicting probabilities q(zn)q(z_n) given logits zn∈RKz_n \in \mathbb{R}^K, target label yny_n, and pre-computed teacher softened probabilities q(vn)q(v_n), the constrained entropy maximization problem is:

    max⁡qL=−∑n=1N∑k=1Kq(zn)(k)log⁡q(zn)(k)\max_{q} \mathcal{L} = -\sum_{n=1}^N \sum_{k=1}^K q(z_n)^{(k)} \log q(z_n)^{(k)}

    subject to ∑k=1Kq(zn)(k)=1,∑k=1Kzn(k)q(zn)(k)=zn(yn),∑k=1Kzn(k)q(zn)(k)=∑k=1Kzn(k)q(vn)(k)(∀n)\text{subject to } \sum_{k=1}^K q(z_n)^{(k)} = 1, \quad \sum_{k=1}^K z_n^{(k)} q(z_n)^{(k)} = z_n^{(y_n)}, \quad \sum_{k=1}^K z_n^{(k)} q(z_n)^{(k)} = \sum_{k=1}^K z_n^{(k)} q(v_n)^{(k)} \quad (\forall n)

    Applying sample-wise Lagrange multipliers β1,n,β2,n,β3,n\beta_{1,n}, \beta_{2,n}, \beta_{3,n} and setting the derivative of the Lagrangian with respect to q(zn)(k)q(z_n)^{(k)} to zero yields:

    q(zn)(k)=exp⁡(βnzn(k))∑m=1Kexp⁡(βnzn(m))q(z_n)^{(k)} = \frac{\exp\left(\beta_n z_n^{(k)}\right)}{\sum_{m=1}^K \exp\left(\beta_n z_n^{(m)}\right)}

    where βn=β2,n+β3,n\beta_n = \beta_{2,n} + \beta_{3,n}. An analogous derivation for teacher predictions with Lagrange multipliers α1,n,α2,n\alpha_{1,n}, \alpha_{2,n} yields q(vn)(k)=exp⁡(α2,nvn(k))∑m=1Kexp⁡(α2,nvn(m))q(v_n)^{(k)} = \frac{\exp(\alpha_{2,n} v_n^{(k)})}{\sum_{m=1}^K \exp(\alpha_{2,n} v_n^{(m)})}.

    Because the constraints do not enforce equality between βn\beta_n and α2,n\alpha_{2,n}, nor do they require βn\beta_n to be constant across samples nn, setting distinct temperatures TS=1/βnT_S = 1/\beta_n and TT=1/α2,nT_T = 1/\alpha_{2,n} is mathematically justified.

  5. Knowl 5 — Shared-Temperature Failure Mode and Z-Score Resolution in Logit Evaluation

    empirical result

    When evaluating knowledge distillation under a shared temperature, the Kullback-Leibler (KL) divergence can penalize rank-accurate students while favoring rank-inaccurate students whose logit magnitudes happen to be closer to the teacher.

    Consider a 4-class toy example with classes {Cat,Dog,Bird,Frog}\{\text{Cat}, \text{Dog}, \text{Bird}, \text{Frog}\} where a teacher outputs logits v=[1.0,4.0,3.0,2.0]v = [1.0, 4.0, 3.0, 2.0] (predicting 'Dog'):

    • Student S1\mathcal{S}_1 outputs logits z1=[1.0,2.8,3.0,2.0]z_1 = [1.0, 2.8, 3.0, 2.0] (predicting 'Bird', an incorrect top prediction, but magnitudes are close to the teacher's).
    • Student S2\mathcal{S}_2 outputs logits z2=[0.1,0.4,0.3,0.2]z_2 = [0.1, 0.4, 0.3, 0.2] (predicting 'Dog', maintaining the exact relative rank of the teacher, but with a smaller scale).

    Under conventional shared-temperature distillation with T=1\mathcal{T} = 1:

    • LKL(softmax(v)∥softmax(z1))=0.1749\mathcal{L}_{KL}(\text{softmax}(v) \parallel \text{softmax}(z_1)) = 0.1749
    • LKL(softmax(v)∥softmax(z2))=0.3457\mathcal{L}_{KL}(\text{softmax}(v) \parallel \text{softmax}(z_2)) = 0.3457

    The shared-temperature loss incorrectly indicates S1\mathcal{S}_1 is superior to S2\mathcal{S}_2.

    Applying weighted ZZ-score standardization (with τ=1\tau = 1):

    • Standardized teacher: Z(v)=[−1.162,1.162,0.387,−0.387]\mathcal{Z}(v) = [-1.162, 1.162, 0.387, -0.387]
    • Standardized S1\mathcal{S}_1: Z(z1)=[−1.320,0.660,0.880,−0.220]\mathcal{Z}(z_1) = [-1.320, 0.660, 0.880, -0.220], yielding LKL=0.0995\mathcal{L}_{KL} = 0.0995
    • Standardized S2\mathcal{S}_2: Z(z2)=[−1.162,1.162,0.387,−0.387]\mathcal{Z}(z_2) = [-1.162, 1.162, 0.387, -0.387], yielding LKL=0.0\mathcal{L}_{KL} = 0.0

    Logit standardization corrects the ranking of the distillation loss to align with actual task prediction accuracy.

  6. Knowl 6 — CIFAR-100 Distillation Benchmark on Heterogeneous Architectures

    data/table

    Top-1 accuracy (%) on the CIFAR-100 validation set evaluated across teacher and student networks of distinct architectural families. Distillation methods are trained with SGD for 240 epochs (480 epochs for MLKD) and averaged over 4 runs. Applying ZZ-score logit standardization consistently boosts the performance of logit-based distillation baselines.

    Teacher ResNet32×\times4 ResNet32×\times4 ResNet32×\times4 WRN-40-2 WRN-40-2 VGG13 ResNet50
    Teacher Acc. 79.42 79.42 79.42 75.61 75.61 74.64 79.34
    Student ShuffleNetV2 WRN-16-2 WRN-40-2 ResNet8×\times4 MobileNetV2 MobileNetV2 MobileNetV2
    Student Acc. 71.82 73.26 75.61 72.50 64.60 64.60 64.60
    FitNet 73.54 74.70 77.69 74.61 68.64 64.16 63.16
    AT 72.73 73.91 77.43 74.11 60.78 59.40 58.58
    RKD 73.21 74.86 77.82 75.26 69.27 64.52 64.43
    CRD 75.65 75.65 78.15 75.24 70.28 69.73 69.11
    OFD 76.82 76.17 79.25 74.36 69.92 69.48 69.04
    ReviewKD 77.78 76.11 78.96 74.34 71.28 70.37 69.89
    SimKD 78.39 77.17 79.29 75.29 70.10 69.44 69.97
    CAT-KD 78.41 76.97 78.59 75.38 70.24 69.13 71.36
    KD 74.45 74.90 77.70 73.97 68.36 67.37 67.35
    KD+Ours 75.56 75.26 77.92 77.11 69.23 68.61 69.02
    Δ\Delta +1.11 +0.36 +0.22 +3.14 +0.87 +1.24 +1.67
    CTKD 75.37 74.57 77.66 74.61 68.34 68.50 68.67
    CTKD+Ours 76.18 75.16 77.99 77.03 69.53 68.98 69.36
    Δ\Delta +0.81 +0.59 +0.33 +2.42 +1.19 +0.48 +0.69
    DKD 77.07 75.70 78.46 75.56 69.28 69.71 70.35
    DKD+Ours 77.37 76.19 78.95 76.75 70.01 69.98 70.45
    Δ\Delta +0.30 +0.49 +0.49 +1.19 +0.73 +0.27 +0.10
    MLKD 78.44 76.52 79.26 77.33 70.78 70.57 71.04
    MLKD+Ours 78.76 77.53 79.66 77.68 71.61 70.94 71.19
    Δ\Delta +0.32 +1.01 +0.40 +0.35 +0.83 +0.37 +0.15
  7. Knowl 7 — CIFAR-100 Distillation Benchmark on Homogeneous Architectures

    data/table

    Top-1 accuracy (%) on the CIFAR-100 validation set for teacher-student model pairs within the same architectural families (ResNet, VGG, WRN), averaged over 4 runs. Logit standardization enhances all tested logit distillation methods across architectural scales.

    Teacher ResNet32×\times4 VGG13 WRN-40-2 WRN-40-2 ResNet56 ResNet110 ResNet110
    Teacher Acc. 79.42 74.64 75.61 75.61 72.34 74.31 74.31
    Student ResNet8×\times4 VGG8 WRN-40-1 WRN-16-2 ResNet20 ResNet32 ResNet20
    Student Acc. 72.50 70.36 71.98 73.26 69.06 71.14 69.06
    FitNet 73.50 71.02 72.24 73.58 69.21 71.06 68.99
    AT 73.44 71.43 72.77 74.08 70.55 72.31 70.65
    RKD 71.90 71.48 72.22 73.35 69.61 71.82 69.25
    CRD 75.51 73.94 74.14 75.48 71.16 73.48 71.46
    OFD 74.95 73.95 74.33 75.24 70.98 73.23 71.29
    ReviewKD 75.63 74.84 75.09 76.12 71.89 73.89 71.34
    SimKD 78.08 74.89 74.53 75.53 71.05 73.92 71.06
    CAT-KD 76.91 74.65 74.82 75.60 71.62 73.62 71.37
    KD 73.33 72.98 73.54 74.92 70.66 73.08 70.67
    KD+Ours 76.62 74.36 74.37 76.11 71.43 74.17 71.48
    Δ\Delta +3.29 +1.38 +0.83 +1.19 +0.77 +1.09 +0.81
    KD+CTKD 73.39 73.52 73.93 75.45 71.19 73.52 70.99
    KD+CTKD+Ours 76.67 74.47 74.58 76.08 71.34 74.01 71.39
    Δ\Delta +3.28 +0.95 +0.65 +0.63 +0.15 +0.49 +0.40
    DKD 76.32 74.68 74.81 76.24 71.97 74.11 71.06
    DKD+Ours 77.01 74.81 74.89 76.39 72.32 74.29 71.85
    Δ\Delta +0.69 +0.13 +0.08 +0.15 +0.35 +0.18 +0.79
    MLKD 77.08 75.18 75.35 76.63 72.19 74.11 71.89
    MLKD+Ours 78.28 75.22 75.56 76.95 72.33 74.32 72.27
    Δ\Delta +1.20 +0.04 +0.21 +0.32 +0.14 +0.21 +0.38
  8. Knowl 8 — ImageNet Classification Distillation Benchmark

    data/table

    Validation performance (Top-1 and Top-5 accuracy in %) on ImageNet for ResNet34 →\rightarrow ResNet18 and ResNet50 →\rightarrow MobileNetV1 distillation pairs. Integrating ZZ-score standardization into KD, CTKD, DKD, and MLKD improves classification performance on large-scale data.

    Teacher / Student ResNet34 / ResNet18 ResNet50 / MobileNetV1
    Metric Top-1 (%) Top-5 (%) Top-1 (%) Top-5 (%)
    Teacher 73.31 91.42 76.16 92.86
    Student 69.75 89.07 68.87 88.76
    AT 70.69 90.01 69.56 89.33
    OFD 70.81 89.98 71.25 90.34
    CRD 71.17 90.13 71.37 90.41
    ReviewKD 71.61 90.51 72.56 91.00
    SimKD 71.59 90.48 72.25 90.86
    CAT-KD 71.26 90.45 72.24 91.13
    KD 71.03 90.05 70.50 89.80
    KD+Ours 71.42 (+0.39) 90.29 (+0.24) 72.18 (+1.68) 90.80 (+1.00)
    KD+CTKD 71.38 90.27 71.16 90.11
    KD+CTKD+Ours 71.81 (+0.43) 90.46 (+0.19) 72.92 (+1.76) 91.25 (+1.14)
    DKD 71.70 90.41 72.05 91.05
    DKD+Ours 71.88 (+0.18) 90.58 (+0.17) 72.85 (+0.80) 91.23 (+0.18)
    MLKD 71.90 90.55 73.01 91.42
    MLKD+Ours 72.08 (+0.18) 90.74 (+0.19) 73.22 (+0.21) 91.59 (+0.17)
  9. Knowl 9 — Logit Distribution Analysis and Mitigating Teacher-Capacity Scaling Bottlenecks

    empirical result

    Larger neural network teachers (e.g., ResNet50, VGG13) intrinsically produce more condensed logits, characterized by a mean closer to zero and smaller standard deviation, whereas smaller student models produce logits with larger bias from zero and higher variance. Under conventional knowledge distillation (KD), transferring from larger teachers fails to scale proportionally due to the difficulty for capacity-limited students to emulate the narrow logit range and low variance of cumbersome teachers.

    When distilling into a WRN-16-2 student (baseline accuracy 73.26%) on CIFAR-100 across teacher models of increasing size:

    Teacher VGG13 WRN-28-2 WRN-40-2 WRN-16-4 WRN-28-4 ResNet50
    Teacher Acc. (%) 74.64 75.45 75.61 77.51 78.60 79.34
    KD 74.93 75.37 74.92 75.79 75.04 75.36
    KD + Ours 75.03 76.32 76.11 76.72 75.77 76.24
    DKD 75.45 75.92 76.24 76.00 76.45 76.60
    DKD + Ours 75.56 76.39 76.39 76.68 76.67 76.82

    While vanilla KD performance plateaus and even drops when transferring from the largest teachers (e.g., 75.79% with WRN-16-4 drops to 75.04% with WRN-28-4 and 75.36% with ResNet50), applying ZZ-score logit standardization allows the student to benefit consistently from larger teachers.

  10. Knowl 10 — Ablation of Logit Standardization Components and Distillation Loss Weights

    data/table

    Ablation study on CIFAR-100 evaluating the Top-1 accuracy (%) of a ResNet8×\times4 student distilled from a ResNet32×\times4 teacher (base temperature τ=2\tau = 2, cross-entropy weight λCE=0.1\lambda_{CE} = 0.1) across distillation loss weights λKD∈[0.9,18.0]\lambda_{KD} \in [0.9, 18.0] and logit pre-processing variants:

    • zz: raw logits (vanilla KD)
    • z−zˉz - \bar{z}: zero-mean centering only
    • z/σ(z)z / \sigma(z): standard deviation scaling only
    • z−zˉσ(z)\frac{z - \bar{z}}{\sigma(z)}: complete ZZ-score standardization
    λKD\lambda_{KD} zz (KD) z−zˉz - \bar{z} z/σ(z)z / \sigma(z) z−zˉσ(z)\frac{z - \bar{z}}{\sigma(z)} (Ours)
    0.9 73.60 73.37 73.79 74.14
    3.0 74.38 74.33 75.86 76.11
    6.0 74.45 74.82 76.44 76.56
    9.0 73.33 73.94 76.30 76.62
    12.0 68.29 71.56 76.49 76.56
    15.0 65.34 62.01 76.42 76.61
    18.0 63.45 61.31 76.18 76.33

    Vanilla KD (zz) and zero-centering alone (z−zˉz - \bar{z}) suffer performance collapse as λKD\lambda_{KD} increases (dropping below 64% at λKD=18.0\lambda_{KD}=18.0). In contrast, logit standardization maintains peak accuracy (76.33%--76.62%) across large distillation weights.

Coverage note — Supplementary experiments on distilling vision transformers, MS-COCO object detection, and CUB-200 fine-grained classification are mentioned in the text as deferred to the supplementary material and were therefore omitted from the main knowls.

References

  1. 1.Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D Lawrence, and Zhenwen Dai. Variational information distillation for knowledge transfer. In CVPR, 2019. 2
  2. 2.Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Yunqing Zhao, and Ngai-Man Cheung. Revisiting label smoothing and knowledge distillation compatibility: What was missing? In ICML, 2022. 2
  3. 3.Defang Chen, Jian-Ping Mei, Can Wang, Yan Feng, and Chun Chen. Online knowledge distillation with diverse peers. In AAAI, 2020. 2
  4. 4.Defang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang, Yan Feng, and Chun Chen. Knowledge distillation with the reused teacher classifier. In CVPR, 2022. 2, 6, 7
  5. 5.Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. Distilling knowledge via knowledge review. In CVPR, 2021. 2, 6, 7
  6. 6.Xianing Chen, Qiong Cao, Yujie Zhong, Jing Zhang, Shenghua Gao, and Dacheng Tao. Dearkd: Data-efficient early knowledge distillation for vision transformers. In CVPR, 2022. 8
  7. 7.Jang Hyun Cho and Bharath Hariharan. On the efficacy of knowledge distillation. In ICCV, 2019. 2, 4, 8
  8. 8.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML, 2017. 3
  9. 9.Jia Guo. Reducing the teacher-student gap via adaptive temperatures, 2022. 2, 5
  10. 10.Ziyao Guo, Haonan Yan, Hui Li, and Xiaodong Lin. Class attention transfer based knowledge distillation. In CVPR, 2023. 2, 6, 7
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  12. 12.Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi. A comprehensive overhaul of feature distillation. In ICCV, 2019. 2, 6, 7
  13. 13.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 1, 2, 5, 6, 7, 8
  14. 14.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 6
  15. 15.Tao Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. Knowledge distillation from a stronger teacher. NeurIPS, 2022. 2, 4
  16. 16.Edwin T Jaynes. Information theory and statistical mechanics. Physical review, 106(4):620, 1957. 3
  17. 17.Ying Jin, Jiaqi Wang, and Dahua Lin. Multi-level logit distillation. In CVPR, 2023. 2, 6, 7, 8
  18. 18.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 2, 5, 6, 7
  19. 19.Gang Li, Xiang Li, Yujie Wang, Shanshan Zhang, Yichao Wu, and Ding Liang. Knowledge distillation for object detection via rank mimicking and prediction-guided feature imitation. In AAAI, 2022. 2
  20. 20.Kehan Li, Runyi Yu, Zhennan Wang, Li Yuan, Guoli Song, and Jie Chen. Locality guidance for improving vision transformers on tiny datasets. In ECCV, 2022. 8
  21. 21.Lujun Li, Peijie Dong, Zimian Wei, and Ya Yang. Automated knowledge distillation via monte carlo tree search. In ICCV, 2023. 8
  22. 22.Zheng Li, Ying Huang, Defang Chen, Tianren Luo, Ning Cai, and Zhigeng Pan. Online knowledge distillation via multi-branch diversity enhancement. In ACCV, 2020. 2
  23. 23.Zheng Li, Jingwen Ye, Mingli Song, Ying Huang, and Zhigeng Pan. Online knowledge distillation for efficient pose estimation. In ICCV, 2021. 2
  24. 24.Zheng Li, Xiang Li, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo, Jun Li, and Jian Yang. Curriculum temperature for knowledge distillation. In AAAI, 2023. 2, 6, 7, 8
  25. 25.Sihao Lin, Hongwei Xie, Bing Wang, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang, and Gang Wang. Knowledge distillation via the target-aware transformer. In CVPR, 2022. 2
  26. 26.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 8
  27. 27.Jihao Liu, Boxiao Liu, Hongsheng Li, and Yu Liu. Meta knowledge distillation. arXiv preprint arXiv:2202.07940, 2022. 2
  28. 28.Yifan Liu, Ke Chen, Chris Liu, Zengchang Qin, Zhenbo Luo, and Jingdong Wang. Structured knowledge distillation for semantic segmentation. In CVPR, 2019. 2
  29. 29.Kangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu, Vishal M Patel, and Peyman Milanfar. Conditional diffusion distillation. arXiv preprint arXiv:2310.01407, 2023. 2
  30. 30.Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. Improved knowledge distillation via teacher assistant. In AAAI, 2020. 2, 4
  31. 31.Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. Relational knowledge distillation. In CVPR, 2019. 2, 6, 7
  32. 32.Baoyun Peng, Xiao Jin, Jiaheng Liu, Dongsheng Li, Yichao Wu, Yu Liu, Shunfeng Zhou, and Zhaoning Zhang. Correlation congruence for knowledge distillation. In ICCV, 2019. 2, 6
  33. 33.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. ICLR, 2015. 2, 6, 7
  34. 34.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015. 2, 5, 6, 7
  35. 35.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018. 6
  36. 36.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 6
  37. 37.Wonchul Son, Jaemin Na, Junyong Choi, and Wonjun Hwang. Densely guided knowledge distillation using multiple teacher assistants. In ICCV, 2021. 2, 4
  38. 38.Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. On the importance of initialization and momentum in deep learning. In ICML, 2013. 6
  39. 39.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation. ICLR, 2020. 2, 6, 7
  40. 40.Yijun Tian, Chuxu Zhang, Zhichun Guo, Xiangliang Zhang, and Nitesh V Chawla. Learning mlps on graphs: A unified view of effectiveness, robustness, and efficiency. In ICLR, 2023. 2
  41. 41.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In ICML, 2021. 8
  42. 42.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. In ECCV, 2022. 8
  43. 43.Frederick Tung and Greg Mori. Similarity-preserving knowledge distillation. In ICCV, 2019. 2
  44. 44.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. JMLR, 2008. 8
  45. 45.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011. 8
  46. 46.Lin Wang and Kuk-Jin Yoon. Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks. IEEE TPAMI, 2021. 8
  47. 47.Ziyang Wang and Congying Ma. Dual-contrastive dual-consistency dual-transformer: A semi-supervised approach to medical image segmentation. In ICCV, 2023. 8
  48. 48.Ziyang Wang, Tianze Li, Jian-Qing Zheng, and Baoru Huang. When cnn meet with vit: Towards semi-supervised learning for multi-class medical image semantic segmentation. In ECCV, 2022.
  49. 49.Kan Wu, Jinnian Zhang, Houwen Peng, Mengchen Liu, Bin Xiao, Jianlong Fu, and Lu Yuan. Tinyvit: Fast pretraining distillation for small vision transformers. In ECCV, 2022. 8
  50. 50.Jing Yang, Brais Martinez, Adrian Bulat, Georgios Tzimiropoulos, et al. Knowledge distillation via softmax regression representation learning. In ICLR, 2021. 2
  51. 51.Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In CVPR, 2017. 2
  52. 52.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016. 6
  53. 53.Sergey Zagoruyko and Nikos Komodakis. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. ICLR, 2017. 6, 7
  54. 54.Rongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu, Michael Bendersky, Marc Najork, and Chao Zhang. Do not blindly imitate the teacher: Using perturbed loss for knowledge distillation. arXiv preprint arXiv:2305.05010, 2023. 2
  55. 55.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In CVPR, 2018. 6
  56. 56.Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu. Deep mutual learning. In CVPR, 2018. 2
  57. 57.Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang. Decoupled knowledge distillation. In CVPR, 2022. 2, 6, 7, 8

Citation

MLA
Sun, S., et al. “Logit Standardization in Knowledge Distillation”. arXiv, 2024, http://arxiv.org/abs/2403.01427v1.
APA
Sun, S., Ren, W., Li, J., Wang, R., & Cao, X. (2024). Logit Standardization in Knowledge Distillation. arXiv. http://arxiv.org/abs/2403.01427v1
Chicago
Sun, S., W. Ren, J. Li, R. Wang, and X. Cao. 2024. “Logit Standardization in Knowledge Distillation”. arXiv. http://arxiv.org/abs/2403.01427v1.
Harvard
Sun, S. et al. (2024) “Logit Standardization in Knowledge Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.01427v1.
Vancouver
1. Sun S, Ren W, Li J, Wang R, Cao X (2024) Logit Standardization in Knowledge Distillation. arXiv

BibTeX

@article{sun2024logit,
  title = {Logit Standardization in Knowledge Distillation},
  author = {Sun, Shangquan and Ren, Wenqi and Li, Jingzhi and Wang, Rui and Cao, Xiaochun},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.01427v1},
  eprint = {2403.01427}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE