f-Divergence Minimization for Sequence-Level Knowledge Distillation

Yuqiao WenZichao LiWenyu DuLili Mou

article2023ACL100 citations

Proposes a unified knowledge distillation framework that formulates sequence-level distillation as generalized f-divergence minimization, decomposing intractable sequence objectives into tractable word-level losses to mitigate mode averaging and mode collapse across text generation tasks.

Listen

State-of-the-art language models deliver strong performance across complex generation tasks but are often too large and resource-intensive for cost-effective deployment. Knowledge distillation trains smaller, efficient student models to replicate the output behaviors of these massive teacher models. However, standard distillation techniques rely on asymmetric divergence objectives that lead to severe trade-offs: they either force the student to spread its probability mass too broadly, resulting in bland and generic outputs (mode averaging), or concentrate too heavily on narrow high-probability spikes, resulting in repetitive and limited outputs (mode collapsing).

The article introduces and evaluates f-DISTILL, a unified sequence-level distillation framework that formulates knowledge transfer as the minimization of generalized mathematical divergence functions. Its main objective is to overcome the limitations of classic methods by systematically exploring symmetric divergence measures that balance coverage and focus during training.

The researchers developed mathematical step-wise decompositions to make full sequence-level distillation computationally tractable at the word level. They evaluated four variations of their framework—standard Kullback-Leibler (KL), Reverse KL, Jensen-Shannon (JS) divergence, and Total Variation Distance (TVD)—across four distinct benchmark tasks: structured data-to-text generation (DART), extreme summarization (XSum), machine translation (WMT16 English-to-Romanian), and conversational dialogue (Commonsense Dialogue). Teacher models containing 200 million to 400 million parameters transferred knowledge to compact student models containing 50 million to 150 million parameters. To maintain practical training speeds, the framework pre-samples outputs from the frozen teacher offline rather than generating them dynamically during every training pass.

The findings demonstrate that the proposed framework consistently outperforms existing knowledge distillation baselines across all tasks. First, using full predictive probability distributions (soft labels) significantly improves performance compared to baseline methods that learn from hard sampled text sequences. Second, symmetric distillation variants—specifically JS and TVD—rank highest overall, successfully preventing both mode averaging and mode collapsing. For example, on the data-to-text task, TVD increased the BLEU accuracy score from 45.54 (standard sequence distillation) to 46.95. Third, asymmetric objectives perform well only in specific operational settings: Reverse KL benefits open-ended tasks with diverse valid outputs like dialogue generation, whereas standard KL is preferable for constrained, single-intent tasks like translation. Fourth, human evaluation confirmed that symmetric methods generate significantly less hallucinated or omitted content while maintaining natural fluency. Finally, the offline sampling strategy accelerated training by more than 2.25 times without degrading final generation quality.

These results provide a clear pathway for organizations to deploy compact, high-accuracy language models, reducing hardware hosting costs and operational latency without sacrificing output quality. The findings show that distillation failures in production models, such as generic or repetitive text generation, stem directly from mathematical distribution mismatches rather than model capacity limitations alone. Because f-DISTILL operates additively on top of existing initialization and intermediate-layer matching techniques, engineering teams can integrate it directly into existing compression pipelines.

Technical leaders looking to compress language models should adopt symmetric distillation objectives—namely TVD or JS divergence—as default training criteria for multi-modal text generation, while employing offline sampling to manage computational training overhead. Standard KL remains an acceptable alternative primarily for narrow translation workflows. Prior to wide production deployment, teams should conduct internal pilot tests to fine-tune pre-distillation warm-ups and verify domain-specific behavior. While confidence in the methodology is supported by consistent gains across multiple benchmarks and human evaluations, decision-makers should note that training requires greater initial GPU computation than simplistic hard-label distillation and that final reported results reflect single-run benchmarks.

Cover for f-Divergence Minimization for Sequence-Level Knowledge Distillation

Abstract

Knowledge distillation (KD) is the process of transferring knowledge from a large model to a small one. It has gained increasing attention in the natural language processing community, driven by the demands of compressing ever-growing language models. In this work, we propose an f-DISTILL framework, which formulates sequence-level knowledge distillation as minimizing a generalized f-divergence function. We propose four distilling variants under our framework and show that existing SeqKD and ENGINE approaches are approximations of our f-DISTILL methods. We further derive step-wise decomposition for our f-DISTILL, reducing intractable sequence-level divergence to word-level losses that can be computed in a tractable manner. Experiments across four datasets show that our methods outperform existing KD approaches, and that our symmetric distilling losses can better force the student to learn from the teacher distribution.

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 2.1 Classic KD and Its Drawbacks
  • 2.2 Our Proposed f -DISTILL Framework
  • 2.3 Implementation Considerations
  • 3 Experiments
  • 3.1 Settings
  • 3.2 Results and Analyses
  • 4 Related Work
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgments
  • References
  • A Proof of Theorem 1
  • B Experimental Details
  • C Case Study

Knowls

  1. Knowl 1 — The f-DISTILL Framework for Sequence-Level Knowledge Distillation

    model/method

    The ff-DISTILL framework formulates sequence-level knowledge distillation for autoregressive text generation as the minimization of an ff-divergence between a teacher distribution pp and a parameterized student distribution qθq_\theta.

    Given two distributions pp and qq defined over full sequences Y=(Y1,…,YT)Y = (Y_1, \dots, Y_T) in vocabulary VV, the generalized ff-divergence is defined by Df(p(Y)∥q(Y))=∑Yq(Y)f(p(Y)q(Y))D_f(p(Y) \parallel q(Y)) = \sum_Y q(Y) f\left(\frac{p(Y)}{q(Y)}\right) where f:(0,∞)→Rf: (0, \infty) \to \mathbb{R} is a convex function satisfying f(1)=0f(1) = 0.

    The framework introduces four sequence-level distilling objectives based on specific choices of ff:

    1. Kullback-Leibler (KL) Distillation (f(t)=tlog⁡tf(t) = t \log t): JKL=DKL(p∥qθ)=EY∼p[log⁡p(Y)qθ(Y)]≈−∑t=1∣y∣∑Yt∈Vp(Yt∣y<t)log⁡qθ(Yt∣y<t)+constJ_{\text{KL}} = D_{\text{KL}}(p \parallel q_\theta) = \mathbb{E}_{Y \sim p}\left[\log \frac{p(Y)}{q_\theta(Y)}\right] \approx -\sum_{t=1}^{|y|} \sum_{Y_t \in V} p(Y_t \mid y_{<t}) \log q_\theta(Y_t \mid y_{<t}) + \text{const} where yy is a sequence sampled from the teacher pp. This objective uses complete soft-label distributions at each step, generalizing hard-label sequence distillation.

    2. Reverse KL (RKL) Distillation (f(t)=−log⁡tf(t) = -\log t): JRKL=DKL(qθ∥p)=EY′∼qθ[log⁡qθ(Y′)p(Y′)]≈∑t=1∣y′∣∑Yt′∈Vqθ(Yt′∣y<t′)[log⁡qθ(Yt′∣y<t′)−log⁡p(Yt′∣y<t′)]J_{\text{RKL}} = D_{\text{KL}}(q_\theta \parallel p) = \mathbb{E}_{Y' \sim q_\theta}\left[\log \frac{q_\theta(Y')}{p(Y')}\right] \approx \sum_{t=1}^{|y'|} \sum_{Y'_t \in V} q_\theta(Y'_t \mid y'_{<t}) \left[\log q_\theta(Y'_t \mid y'_{<t}) - \log p(Y'_t \mid y'_{<t})\right] where y′y' is a sequence sampled from the student qθq_\theta. The first term corresponds to the negative entropy of the student, and the second term matches the student's output against the teacher's likelihood.

    3. Jensen-Shannon (JS) Distillation (f(t)=−(t+1)log⁡(t+12)+tlog⁡tf(t) = -(t + 1) \log\left(\frac{t + 1}{2}\right) + t \log t): JJS=12EY∼p[log⁡p(Y)m(Y)]+12EY′∼qθ[log⁡qθ(Y′)m(Y′)]J_{\text{JS}} = \frac{1}{2} \mathbb{E}_{Y \sim p}\left[\log \frac{p(Y)}{m(Y)}\right] + \frac{1}{2} \mathbb{E}_{Y' \sim q_\theta}\left[\log \frac{q_\theta(Y')}{m(Y')}\right] where m(Y)=12p(Y)+12qθ(Y)m(Y) = \frac{1}{2} p(Y) + \frac{1}{2} q_\theta(Y) is the mixture distribution.

    4. Total Variation Distance (TVD) Distillation (f(t)=12∣t−1∣f(t) = \frac{1}{2}|t - 1|): JTVD=12∑Y∣qθ(Y)−p(Y)∣J_{\text{TVD}} = \frac{1}{2} \sum_Y |q_\theta(Y) - p(Y)| which measures the ℓ1\ell_1 distance between the teacher and student distributions.

  2. Knowl 2 — Exact Step-Wise Decomposition of Sequence-Level Divergences

    theoretical result

    For autoregressive sequence generation models where sequences Y1:T=(Y1,…,YT)Y_{1:T} = (Y_1, \dots, Y_T) are generated token-by-token from teacher distribution pp and student distribution qθq_\theta over vocabulary VV, sequence-level Kullback-Leibler (KL), Reverse KL (RKL), and Jensen-Shannon (JS) divergences decompose exactly into sums of step-wise conditional expectations.

    Let m(Y)=12p(Y)+12qθ(Y)m(Y) = \frac{1}{2} p(Y) + \frac{1}{2} q_\theta(Y) denote the average distribution. The exact decompositions are:

    1. Step-wise Forward KL Decomposition: DKL(p(Y1:T)∥qθ(Y1:T))=∑t=1TEY1:t−1∼p[∑Yt∈Vp(Yt∣Y1:t−1)log⁡p(Yt∣Y1:t−1)qθ(Yt∣Y1:t−1)]D_{\text{KL}}(p(Y_{1:T}) \parallel q_\theta(Y_{1:T})) = \sum_{t=1}^T \mathbb{E}_{Y_{1:t-1} \sim p}\left[ \sum_{Y_t \in V} p(Y_t \mid Y_{1:t-1}) \log \frac{p(Y_t \mid Y_{1:t-1})}{q_\theta(Y_t \mid Y_{1:t-1})} \right]

    2. Step-wise Reverse KL Decomposition: DKL(qθ(Y1:T)∥p(Y1:T))=∑t=1TEY1:t−1′∼qθ[∑Yt′∈Vqθ(Yt′∣Y1:t−1′)log⁡qθ(Yt′∣Y1:t−1′)p(Yt′∣Y1:t−1′)]D_{\text{KL}}(q_\theta(Y_{1:T}) \parallel p(Y_{1:T})) = \sum_{t=1}^T \mathbb{E}_{Y'_{1:t-1} \sim q_\theta}\left[ \sum_{Y'_t \in V} q_\theta(Y'_t \mid Y'_{1:t-1}) \log \frac{q_\theta(Y'_t \mid Y'_{1:t-1})}{p(Y'_t \mid Y'_{1:t-1})} \right]

    3. Step-wise JS Decomposition: DJS(p(Y1:T)∥qθ(Y1:T))=12∑t=1TEY1:t−1∼p[∑Yt∈Vp(Yt∣Y1:t−1)log⁡p(Yt∣Y1:t−1)m(Yt∣Y1:t−1)]+12∑t=1TEY1:t−1′∼qθ[∑Yt′∈Vqθ(Yt′∣Y1:t−1′)log⁡qθ(Yt′∣Y1:t−1′)m(Yt′∣Y1:t−1′)]D_{\text{JS}}(p(Y_{1:T}) \parallel q_\theta(Y_{1:T})) = \frac{1}{2} \sum_{t=1}^T \mathbb{E}_{Y_{1:t-1} \sim p}\left[ \sum_{Y_t \in V} p(Y_t \mid Y_{1:t-1}) \log \frac{p(Y_t \mid Y_{1:t-1})}{m(Y_t \mid Y_{1:t-1})} \right] + \frac{1}{2} \sum_{t=1}^T \mathbb{E}_{Y'_{1:t-1} \sim q_\theta}\left[ \sum_{Y'_t \in V} q_\theta(Y'_t \mid Y'_{1:t-1}) \log \frac{q_\theta(Y'_t \mid Y'_{1:t-1})}{m(Y'_t \mid Y'_{1:t-1})} \right]

    In practice, expectations over prior prefix histories EY1:t−1∼p\mathbb{E}_{Y_{1:t-1} \sim p} and EY1:t−1′∼qθ\mathbb{E}_{Y'_{1:t-1} \sim q_\theta} are approximated via single Monte Carlo sampled sequences y∼py \sim p and y′∼qθy' \sim q_\theta. The exact summation over the vocabulary VV at each step tt allows complete gradient propagation through every token in the vocabulary.

  3. Knowl 3 — Step-Wise Upper Bound on Sequence-Level Total Variation Distance

    theoretical result

    The sequence-level Total Variation Distance (TVD) between a teacher distribution pp and a student distribution qθq_\theta over sequences Y1:TY_{1:T} cannot be decomposed into exact independent step-wise expectations. Instead, it is upper bounded by a combination of step-wise ℓ1\ell_1 token-level differences under teacher and student sampling trajectories:

    DTVD(p(Y1:T)∥qθ(Y1:T))=12∑Y1:T∣qθ(Y1:T)−p(Y1:T)∣≤LTVD-boundD_{\text{TVD}}(p(Y_{1:T}) \parallel q_\theta(Y_{1:T})) = \frac{1}{2} \sum_{Y_{1:T}} |q_\theta(Y_{1:T}) - p(Y_{1:T})| \le \mathcal{L}_{\text{TVD-bound}}

    where the tractable upper bound objective LTVD-bound\mathcal{L}_{\text{TVD-bound}} is given by: LTVD-bound=14∑t=1TEY1:t−1∼p[∑Yt∈V∣qθ(Yt∣Y1:t−1)−p(Yt∣Y1:t−1)∣]+14∑t=1TEY1:t−1′∼qθ[∑Yt′∈V∣qθ(Yt′∣Y1:t−1′)−p(Yt′∣Y1:t−1′)∣]\mathcal{L}_{\text{TVD-bound}} = \frac{1}{4} \sum_{t=1}^T \mathbb{E}_{Y_{1:t-1} \sim p}\left[ \sum_{Y_t \in V} |q_\theta(Y_t \mid Y_{1:t-1}) - p(Y_t \mid Y_{1:t-1})| \right] + \frac{1}{4} \sum_{t=1}^T \mathbb{E}_{Y'_{1:t-1} \sim q_\theta}\left[ \sum_{Y'_t \in V} |q_\theta(Y'_t \mid Y'_{1:t-1}) - p(Y'_t \mid Y'_{1:t-1})| \right]

    Approximating the prefix expectations via Monte Carlo sampled sequences y∼py \sim p and y′∼qθy' \sim q_\theta yields the step-wise surrogate training loss: J^TVD=14∑t=1∣y∣∑Yt∈V∣qθ(Yt∣y<t)−p(Yt∣y<t)∣+14∑t=1∣y′∣∑Yt′∈V∣qθ(Yt′∣y<t′)−p(Yt′∣y<t′)∣\hat{J}_{\text{TVD}} = \frac{1}{4} \sum_{t=1}^{|y|} \sum_{Y_t \in V} |q_\theta(Y_t \mid y_{<t}) - p(Y_t \mid y_{<t})| + \frac{1}{4} \sum_{t=1}^{|y'|} \sum_{Y'_t \in V} |q_\theta(Y'_t \mid y'_{<t}) - p(Y'_t \mid y'_{<t})|

    Because the TVD objective operates directly on ℓ1\ell_1 norm differences without logarithmic transformations, it yields more stable gradients than divergence measures relying on log ratios.

  4. Knowl 4 — Offline Teacher Sampling for Sequence-Level Distillation

    model/method

    Symmetric sequence-level distillation methods such as Jensen-Shannon (JS) divergence and Total Variation Distance (TVD) require expectations over sequences generated by both the teacher model pp and the student model qθq_\theta.

    Because the teacher model is frozen throughout distillation, sampling from the teacher on-the-fly during training introduces redundant computation. Under the offline sampling procedure:

    1. Sequences y∼p(⋅∣x)y \sim p(\cdot \mid x) are sampled once across the dataset prior to training and stored.
    2. During training, the pre-generated teacher sequences yy are reused for the teacher expectation terms across all training epochs.
    3. Sequences y′∼qθ(⋅∣x)y' \sim q_\theta(\cdot \mid x) from the student model are sampled online at each training step, as the student parameters θ\theta are actively updating.

    This offline pre-sampling reduces the computational cost of the teacher model during training, more than doubling training speed while matching the generation performance of online teacher sampling.

  5. Knowl 5 — Likelihood Risk and Coverage Risk for Evaluating Generation Diversity

    definition

    To quantitatively evaluate the trade-off between mode averaging (where a model spreads probability mass over invalid regions) and mode collapsing (where a model concentrates probability mass exclusively on a subset of modes), two evaluation risks are defined:

    1. Likelihood Risk (RllhR_{\text{llh}}): Rllh=1∣Dstudent∣∑y′∈Dstudent−log⁡p(y′)R_{\text{llh}} = \frac{1}{|D_{\text{student}}|} \sum_{y' \in D_{\text{student}}} -\log p(y') where DstudentD_{\text{student}} is the set of sequences generated by sampling from the student model qθq_\theta (one sample per test input), and p(y′)p(y') is the probability assigned to the student-generated sequence y′y' by the teacher model pp. A high RllhR_{\text{llh}} indicates mode averaging, as the student generates atypical sequences from the teacher's perspective.

    2. Coverage Risk (RcvgR_{\text{cvg}}): Rcvg=1∣Dteacher∣∑y∈Dteacher−log⁡qθ(y)R_{\text{cvg}} = \frac{1}{|D_{\text{teacher}}|} \sum_{y \in D_{\text{teacher}}} -\log q_\theta(y) where DteacherD_{\text{teacher}} is the set of sequences sampled from the teacher model pp, and qθ(y)q_\theta(y) is the probability assigned to the teacher-generated sequence yy by the student model qθq_\theta. A high RcvgR_{\text{cvg}} indicates mode collapsing, as typical teacher generations fall outside the support covered by the student model.

    3. Teacher Diversity (TeacherDist\text{TeacherDist}): Task multi-modality is quantified by the percentage of distinct bi-grams among five teacher-sampled outputs per test input, averaged across the test set.

  6. Knowl 6 — Trade-off Between Mode Averaging and Collapsing Across Task Multi-Modality

    empirical result

    The severity of mode averaging and mode collapsing in knowledge distillation depends directly on the task's multi-modality, as measured by teacher diversity (distinct bi-gram percentage among five sampled outputs per input). Empirical evaluation across four generation benchmarks demonstrates the behavioral differences among divergence formulations:

    Dataset DART XSum MT EN-RO CD
    TeacherDist 26.10 36.28 23.13 81.19
    Risk RllhR_{\text{llh}} RcvgR_{\text{cvg}} RllhR_{\text{llh}} RcvgR_{\text{cvg}} RllhR_{\text{llh}} RcvgR_{\text{cvg}} RllhR_{\text{llh}} RcvgR_{\text{cvg}}
    KL 0.56 0.49 1.89 1.68 1.23 0.82 0.43 0.26
    RKL 0.58 0.59 1.88 1.83 1.20 1.60 0.29 0.35
    TVD 0.53 0.52 1.86 1.77 1.21 1.78 0.27 0.35
    JS 0.51 0.48 1.88 1.75 1.13 1.34 0.30 0.33

    Key observations include:

    1. Forward KL distillation consistently achieves lower coverage risk RcvgR_{\text{cvg}} across all datasets, reflecting its mode-covering (mode-averaging) tendency.
    2. Reverse KL (RKL) achieves a much lower likelihood risk RllhR_{\text{llh}} (0.29 vs. 0.43 for KL) on highly multi-modal tasks like Commonsense Dialogue (CD, TeacherDist=81.19\text{TeacherDist} = 81.19), but performs poorly on coverage risk RcvgR_{\text{cvg}} on unimodal tasks like machine translation (MT EN-RO, TeacherDist=23.13\text{TeacherDist} = 23.13, where Rcvg=1.60R_{\text{cvg}} = 1.60 for RKL vs. 0.820.82 for KL).
    3. Symmetric distillation objectives (JS and TVD) achieve moderate, balanced values for both likelihood risk and coverage risk across unimodal and multi-modal benchmarks.
  7. Knowl 7 — Evaluation of f-DISTILL Variants Across Four Generation Benchmarks

    data/table

    Evaluation of knowledge distillation methods across four generation tasks: DART (data-to-text generation; teacher: BART, student: 4-layer BART), XSum (abstractive summarization; teacher: BART, student: 6-layer BART), WMT16 EN-RO (machine translation; teacher: T5, student: 4-layer T5), and Commonsense Dialogue (dialogue generation; teacher: DialoGPT, student: 4-layer DialoGPT). All student models were pre-distilled using MLE, word-level KL, and intermediate layer matching before running sequence-level distillation.

    Task Model BLEU4↑\uparrow METEOR↑\uparrow TER↓\downarrow BERTScore↑\uparrow MoverScore↑\uparrow BLEURT↑\uparrow
    DART Teacher 48.56 39.28 45.45 83.04 68.17 40.56
    Non-distill (MLE) 43.12 35.71 49.97 79.76 65.65 29.10
    Pre-distill 45.60 36.99 47.10 81.39 66.75 34.08
    SeqKD 45.54 37.17 47.49 81.15 66.65 32.88
    ENGINE 44.40 36.51 50.63 80.18 66.20 30.94
    KL 46.24 37.45 46.89 81.60 67.07 35.31
    RKL 45.63 37.35 47.91 81.41 67.02 35.08
    JS 46.85 37.75 46.50 81.93 67.30 36.81
    TVD 46.95 37.88 46.35 82.08 67.36 37.17
    Task Model ROUGE-1↑\uparrow ROUGE-2↑\uparrow ROUGE-L↑\uparrow
    XSum Teacher 45.12 22.26 37.18
    Non-distill (MLE) 30.00 10.67 24.40
    Pre-distill 40.58 17.79 32.55
    SeqKD 39.13 17.53 32.34
    ENGINE 39.19 16.18 31.23
    KL 41.28 18.98 33.71
    RKL 41.69 19.02 33.92
    JS 41.65 19.22 34.03
    TVD 41.76 19.30 34.10
    Task Model BLEU4↑\uparrow chrF↑\uparrow TER↓\downarrow
    WMT16 Teacher 25.82 55.76 60.57
    EN-RO Non-distill (MLE) 19.90 49.79 69.48
    Pre-distill 20.68 50.51 68.38
    SeqKD 21.20 50.81 67.66
    ENGINE 17.65 48.37 84.02
    KL 21.45 51.12 66.74
    RKL 20.46 50.33 70.78
    JS 21.91 51.50 66.86
    TVD 21.73 51.13 66.94
    Task Model BLEU1↑\uparrow BLEU2↑\uparrow BERTScore↑\uparrow
    Commonsense Teacher 11.67 5.03 47.69
    Dialogue Non-distill (MLE) 10.23 3.56 45.15
    Pre-distill 9.95 3.63 46.22
    SeqKD 10.85 4.17 46.94
    ENGINE 10.13 4.26 46.91
    KL 9.81 3.52 45.80
    RKL 10.48 4.01 46.68
    JS 11.55 4.83 47.61
    TVD 11.39 4.73 47.30

    Soft-label ff-DISTILL methods consistently outperform hard sequence approximations (SeqKD and ENGINE). Symmetric divergence objectives (JS and TVD) attain the best performance across data-to-text, summarization, and dialogue generation tasks.

  8. Knowl 8 — Human Evaluation of Distillation Methods on Data-to-Text Generation

    data/table

    A blind human evaluation was conducted on 50 test samples from the DART data-to-text generation dataset, assessed by five annotators across three criteria scored from 1 to 5: Fluency (higher is better), Missing Information (lower is better), and Hallucination (lower is better).

    Model Fluency ↑\uparrow MissingInfo ↓\downarrow Hallucination ↓\downarrow
    SeqKD 4.75 1.77 1.67
    ENGINE 4.51 1.76 1.61
    JS 4.72 1.70 1.48
    TVD 4.72 1.57 1.45

    A one-sided Student's tt-test comparing TVD against SeqKD yields:

    • Fluency: p=32.6%p = 32.6\% (no statistically significant difference),
    • Missing Information: p=1.28%p = 1.28\% (statistically significant reduction),
    • Hallucination: p=0.669%p = 0.669\% (statistically significant reduction).

    Symmetric distillation (TVD and JS) retains fluent output while reducing ungrounded content and omitted structured facts compared to asymmetric distillation baselines.

  9. Knowl 9 — Training Efficiency of Offline versus Online Teacher Sampling

    data/table

    Comparison of online versus offline teacher sequence sampling for Jensen-Shannon (JS) and Total Variation Distance (TVD) distillation on the DART data-to-text generation dataset. Benchmark run on an unshared server with an NVIDIA RTX A6000 GPU and an Intel Xeon Gold 5317 CPU.

    Distillation Method Sampling Mode BLEU4 BERTScore Speedup
    JS distillation Online 46.85 82.02 1.00×\times
    Offline (proposed) 46.85 81.93 2.25×\times
    TVD distillation Online 46.57 82.03 1.00×\times
    Offline (proposed) 46.95 82.08 2.31×\times

    In the online setup, sequences are re-sampled from the teacher model at every training epoch. In the offline setup, teacher sequences are sampled once before distillation begins and held fixed. The offline strategy achieves more than a 2.25×2.25\times speedup while preserving comparable BLEU4 and BERTScore generation performance.

  10. Knowl 10 — Computational and Experimental Limitations of f-DISTILL

    limitation

    The ff-DISTILL framework has two primary stated limitations:

    1. Training Efficiency Overhead: Evaluating step-wise ff-divergences requires computing and summing soft probabilities over the complete vocabulary VV for each sequence step. This demands higher training time than hard-sample distillation methods (such as SeqKD and ENGINE) that calculate cross-entropy on one-hot targets. However, this computational cost is restricted to training and does not alter inference speed or model size during deployment.

    2. Absence of Multi-Run Statistics: Due to the extensive computational requirements across four diverse tasks and multiple model variants (estimated at approximately 2,000 GPU hours on NVIDIA A100 GPUs and AMD Milan 7413 CPUs), the reported experimental numbers are based on single training runs rather than multi-run averages with error bars, although preliminary runs indicated low variance across runs.

Coverage note — None was omitted; all primary theoretical derivations, step-wise bounds, optimization methods, empirical benchmark results, diagnostic metrics, human evaluation, and stated limitations are included.

References

  1. 1.S. M. Ali and S. D. Silvey. 1966. A general class of coefficients of divergence of one distribution from another. Journal of the Royal Statistical Society, 28(1):131–142.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72.
  3. 3.Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2020. PLATO: Pre-trained dialogue generation model with discrete latent variable. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 85–96.
  4. 4.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 Conference on Machine Translation (WMT19). In Proceedings of the Conference on Machine Translation, pages 1–61.
  5. 5.Christopher M. Bishop. 2006. Pattern Recognition and Machine Learning. Springer.
  6. 6.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. Findings of the 2016 Conference on Machine Translation. In Proceedings of the Conference on Machine Translation, pages 131–198.
  7. 7.Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006. Model compression. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, page 535–541.
  8. 8.Chun Fan, Jiwei Li, Tianwei Zhang, Xiang Ao, Fei Wu, Yuxian Meng, and Xiaofei Sun. 2021. Layerwise model pruning based on mutual information. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 3079–3090.
  9. 9.Gongfan Fang, Yifan Bao, Jie Song, Xinchao Wang, Donglin Xie, Chengchao Shen, and Mingli Song. 2021. Mosaicking to distill: Knowledge distillation from out-of-domain data. In Advances in Neural Information Processing Systems, pages 11920–11932.
  10. 10.Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations.
  11. 11.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems.
  12. 12.Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2018. Non-autoregressive neural machine translation. In International Conference on Learning Representations.
  13. 13.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  14. 14.Chenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou, and Lei Li. 2022. Non-autoregressive translation with layer-wise prediction and deep supervision. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 10776–10784.
  15. 15.Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2021. Improving task-agnostic BERT distillation with layer mapping search. Neurocomputing, 461:194–203.
  16. 16.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP, pages 4163–4174.
  17. 17.Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah Smith. 2020. Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation. In International Conference on Learning Representations.
  18. 18.Moniba Keymanesh, Adrian Benton, and Mark Dredze. 2022. What makes data-to-text generation hard for pretrained language models? In Proceedings of the Workshop on Natural Language Generation, Evaluation, and Metrics, pages 539–554.
  19. 19.Yoon Kim and Alexander M. Rush. 2016. Sequencelevel knowledge distillation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1317–1327.
  20. 20.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations.
  21. 21.Rémi Lebret, David Grangier, and Michael Auli. 2016. Neural text generation from structured data with application to the biography domain. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1203–1213.
  22. 22.Yann LeCun, John Denker, and Sara Solla. 1989. Optimal brain damage. In Advances in Neural Information Processing Systems, pages 598–605.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  24. 24.Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020. BERT-EMD: Many-to-many layer mapping for BERT compression with earth mover's distance. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 3009–3018.
  25. 25.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a. A diversity-promoting objective function for neural conversation models. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119.
  26. 26.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016b. Deep reinforcement learning for dialogue generation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1192–1202.
  27. 27.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing, pages 4582–4597.
  28. 28.Zhuoran Li, Chunming Hu, Xiaohui Guo, Junfan Chen, Wenyi Qin, and Richong Zhang. 2022. An unsupervised multiple-task and multiple-teacher model for cross-lingual named entity recognition. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 170–179.
  29. 29.Alexander Lin, Jeremy Wohlwend, Howard Chen, and Tao Lei. 2020. Autoregressive knowledge distillation through imitation learning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 6121–6133.
  30. 30.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81.
  31. 31.Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 2122–2132.
  32. 32.Liyuan Liu, Xiang Ren, Jingbo Shang, Xiaotao Gu, Jian Peng, and Jiawei Han. 2018. Efficient contextualized representation: Language model pruning for sequence labeling. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1215–1225.
  33. 33.Christos Louizos, Max Welling, and Diederik P Kingma. 2018. Learning sparse neural networks through L0 regularization. In International Conference on Learning Representations.
  34. 34.Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021. DART: Opendomain structured data record to text generation. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 432–447.
  35. 35.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don't give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1797–1807.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 311–318.
  37. 37.Romain Paulus, Caiming Xiong, and Richard Socher. 2018. A deep reinforced model for abstractive summarization. In International Conference on Learning Representations.
  38. 38.Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395.
  39. 39.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text Transformer. Journal of Machine Learning Research, 21(140):1–67.
  40. 40.Igal Sason and Sergio Verdú. 2016. f-divergence inequalities. IEEE Transactions on Information Theory, 62(11):5973–6006.
  41. 41.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 7881–7892.
  42. 42.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1715–1725.
  43. 43.Chenze Shao, Xuanfu Wu, and Yang Feng. 2022. One reference is not enough: Diverse distillation with reference selection for non-autoregressive translation. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3779–3791.
  44. 44.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In Proceedings of the International Conference on Machine Learning, pages 4596–4604.
  45. 45.Sam Shleifer and Alexander M Rush. 2020. Pretrained summarization distillation. arXiv preprint arXiv:2010.13002.
  46. 46.Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006. A study of translation edit rate with targeted human annotation. In Proceedings of the Conference of the Association for Machine Translation in the Americas: Technical Papers, pages 223–231.
  47. 47.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for BERT model compression. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing, pages 4323–4332.
  48. 48.Chuanxin Tang, Yucheng Zhao, Guangting Wang, Chong Luo, Wenxuan Xie, and Wenjun Zeng. 2022. Sparse MLP for image recognition: Is self-attention really necessary? In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2344–2351.
  49. 49.Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019. Distilling task-specific knowledge from BERT into simple neural networks. arXiv preprint arXiv:1903.12136.
  50. 50.Lifu Tu, Richard Yuanzhe Pang, Sam Wiseman, and Kevin Gimpel. 2020. ENGINE: Energy-based inference networks for non-autoregressive machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 2819–2826.
  51. 51.Bolin Wei, Shuai Lu, Lili Mou, Hao Zhou, Pascal Poupart, Ge Li, and Zhi Jin. 2019. Why do neural dialog systems generate short and meaningless replies? A comparison between dialog and translation. In Proceedings of the International Conference on Acoustics, Speech and Signal Processing, pages 7290–7294.
  52. 52.Yuqiao Wen, Yongchang Hao, Yanshuai Cao, and Lili Mou. 2023. An equal-size hard EM algorithm for diverse dialogue generation. In International Conference on Learning Representations.
  53. 53.Chuhan Wu, Fangzhao Wu, and Yongfeng Huang. 2021. One teacher is enough? Pre-trained language model distillation from multiple teachers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP, pages 4408–4413.
  54. 54.Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang. 2020. Model compression with two-stage multi-teacher knowledge distillation for web question answering system. In Proceedings of the International Conference on Web Search and Data Mining, page 690–698.
  55. 55.Hongxu Yin, Pavlo Molchanov, Jose M. Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K. Jha, and Jan Kautz. 2020. Dreaming to distill: Data-free knowledge transfer via DeepInversion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8715–8724.
  56. 56.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020a. PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the International Conference on Machine Learning, pages 11328–11339.
  57. 57.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
  58. 58.Yivan Zhang, Gang Niu, and Masashi Sugiyama. 2021. Learning noise transition matrix from only noisy labels via total variation regularization. In Proceedings of the International Conference on Machine Learning, pages 12501–12512.
  59. 59.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020b. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278.
  60. 60.Miaoyun Zhao, Yulai Cong, Shuyang Dai, and Lawrence Carin. 2020. Bridging maximum likelihood and adversarial learning via α-divergence. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 6901–6908.
  61. 61.Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing, pages 563–578.
  62. 62.Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur. 2021. Commonsense-focused dialogues for response generation: An empirical study. In Proceedings of the Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 121–132.

Citation

MLA
Wen, Y., et al. “f-Divergence Minimization for Sequence-Level Knowledge Distillation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 10817–34, https://doi.org/10.18653/v1/2023.acl-long.605.
APA
Wen, Y., Li, Z., Du, W., & Mou, L. (2023). f-Divergence Minimization for Sequence-Level Knowledge Distillation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10817–10834. https://doi.org/10.18653/v1/2023.acl-long.605
Chicago
Wen, Y., Z. Li, W. Du, and L. Mou. 2023. “f-Divergence Minimization for Sequence-Level Knowledge Distillation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10817–34. https://doi.org/10.18653/v1/2023.acl-long.605.
Harvard
Wen, Y. et al. (2023) “f-Divergence Minimization for Sequence-Level Knowledge Distillation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10817–10834. Available at: https://doi.org/10.18653/v1/2023.acl-long.605.
Vancouver
1. Wen Y, Li Z, Du W, Mou L (2023) f-Divergence Minimization for Sequence-Level Knowledge Distillation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10817–10834

BibTeX

@inproceedings{wen-etal-2023-f,
    title = "f-Divergence Minimization for Sequence-Level Knowledge Distillation",
    author = "Wen, Yuqiao  and
      Li, Zichao  and
      Du, Wenyu  and
      Mou, Lili",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.605/",
    doi = "10.18653/v1/2023.acl-long.605",
    pages = "10817--10834"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/