f-Divergence Minimization for Sequence-Level Knowledge Distillation
Yuqiao WenZichao LiWenyu DuLili Mou
Proposes a unified knowledge distillation framework that formulates sequence-level distillation as generalized f-divergence minimization, decomposing intractable sequence objectives into tractable word-level losses to mitigate mode averaging and mode collapse across text generation tasks.
State-of-the-art language models deliver strong performance across complex generation tasks but are often too large and resource-intensive for cost-effective deployment. Knowledge distillation trains smaller, efficient student models to replicate the output behaviors of these massive teacher models. However, standard distillation techniques rely on asymmetric divergence objectives that lead to severe trade-offs: they either force the student to spread its probability mass too broadly, resulting in bland and generic outputs (mode averaging), or concentrate too heavily on narrow high-probability spikes, resulting in repetitive and limited outputs (mode collapsing).
The article introduces and evaluates f-DISTILL, a unified sequence-level distillation framework that formulates knowledge transfer as the minimization of generalized mathematical divergence functions. Its main objective is to overcome the limitations of classic methods by systematically exploring symmetric divergence measures that balance coverage and focus during training.
The researchers developed mathematical step-wise decompositions to make full sequence-level distillation computationally tractable at the word level. They evaluated four variations of their framework—standard Kullback-Leibler (KL), Reverse KL, Jensen-Shannon (JS) divergence, and Total Variation Distance (TVD)—across four distinct benchmark tasks: structured data-to-text generation (DART), extreme summarization (XSum), machine translation (WMT16 English-to-Romanian), and conversational dialogue (Commonsense Dialogue). Teacher models containing 200 million to 400 million parameters transferred knowledge to compact student models containing 50 million to 150 million parameters. To maintain practical training speeds, the framework pre-samples outputs from the frozen teacher offline rather than generating them dynamically during every training pass.
The findings demonstrate that the proposed framework consistently outperforms existing knowledge distillation baselines across all tasks. First, using full predictive probability distributions (soft labels) significantly improves performance compared to baseline methods that learn from hard sampled text sequences. Second, symmetric distillation variants—specifically JS and TVD—rank highest overall, successfully preventing both mode averaging and mode collapsing. For example, on the data-to-text task, TVD increased the BLEU accuracy score from 45.54 (standard sequence distillation) to 46.95. Third, asymmetric objectives perform well only in specific operational settings: Reverse KL benefits open-ended tasks with diverse valid outputs like dialogue generation, whereas standard KL is preferable for constrained, single-intent tasks like translation. Fourth, human evaluation confirmed that symmetric methods generate significantly less hallucinated or omitted content while maintaining natural fluency. Finally, the offline sampling strategy accelerated training by more than 2.25 times without degrading final generation quality.
These results provide a clear pathway for organizations to deploy compact, high-accuracy language models, reducing hardware hosting costs and operational latency without sacrificing output quality. The findings show that distillation failures in production models, such as generic or repetitive text generation, stem directly from mathematical distribution mismatches rather than model capacity limitations alone. Because f-DISTILL operates additively on top of existing initialization and intermediate-layer matching techniques, engineering teams can integrate it directly into existing compression pipelines.
Technical leaders looking to compress language models should adopt symmetric distillation objectives—namely TVD or JS divergence—as default training criteria for multi-modal text generation, while employing offline sampling to manage computational training overhead. Standard KL remains an acceptable alternative primarily for narrow translation workflows. Prior to wide production deployment, teams should conduct internal pilot tests to fine-tune pre-distillation warm-ups and verify domain-specific behavior. While confidence in the methodology is supported by consistent gains across multiple benchmarks and human evaluations, decision-makers should note that training requires greater initial GPU computation than simplistic hard-label distillation and that final reported results reflect single-run benchmarks.
- Paper: Sequence-Level Knowledge Distillation, Yoon Kim et al. (2016). It introduces sequence-level knowledge distillation for autoregressive generation, which the source directly builds upon and generalizes to broader f-divergences.
- Paper: f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization, Sebastian Nowozin et al. (2016). It establishes the foundational variational framework for training generative neural samplers by minimizing generalized f-divergences.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). It provides the foundational framework and core concepts for transferring knowledge from large teacher networks to smaller student networks using softened probability distributions.
- Paper: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Rishabh Agarwal et al. (2024). It compares against f-DISTILL and extends sequence-level divergence minimization by incorporating on-policy student rollouts and reinforcement learning feedback.
- Paper: GKD: Generalized Knowledge Distillation for Auto-regressive Models, Rishabh Agarwal et al. (2026). It builds upon flexible divergence objectives in sequence distillation by developing generalized on-policy sampling schemes for autoregressive models.
- Paper: DistiLLM: Towards Streamlined Distillation for Large Language Models, Jongwoo Ko et al. (2024). It advances sequence-level divergence optimization by introducing skew KL objectives and efficient replay scheduling for student-generated outputs.
- Paper: MiniLLM: On-Policy Distillation of Large Language Models, Yuxian Gu et al. (2024). It continues the investigation into asymmetric versus mode-covering divergences by developing an on-policy policy-gradient framework minimizing reverse KL divergence.
- Paper: DistillSpec: Improving Speculative Decoding via Knowledge Distillation, Yongchao Zhou et al. (2024). It applies sequence-level divergence optimization to train and align compact draft models specifically for speculative decoding acceleration.
