DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Yongchao ZhouKaifeng LyuAnkit Singh RawatAditya Krishna MenonAfshin RostamizadehSanjiv KumarJean-François KagyRishabh Agarwal
Develops DistillSpec, an on-policy knowledge distillation method that aligns draft models with target language models to achieve a 10% to 45% inference speedup over standard speculative decoding.
Deploying large language models in production environments presents a severe operational bottleneck due to high inference latency and substantial computational costs. While speculative decoding has emerged as a promising remedy—using a small, fast draft model to generate candidate text that a large target model validates in parallel—its real-world acceleration depends heavily on how closely the draft model mimics the target model. If the draft model generates tokens that the target model frequently rejects, inference efficiency drops sharply. Developing compact draft models that align closely with large target architectures has thus become a critical priority for practical AI deployment.
The article evaluates a framework called DistillSpec, which uses white-box knowledge distillation to systematically align a small draft model to a larger target model. Across comprehensive empirical benchmarks spanning summarization, translation, arithmetic reasoning, and general language generation, the authors evaluated both greedy and non-greedy decoding strategies. The evaluation compared fixed training datasets against model-generated data, tested multiple mathematical divergence functions, and analyzed model scaling across architectures ranging from 33 million to 11 billion parameters.
The findings establish that DistillSpec consistently enhances inference speed, delivering 10% to 45% speedups over standard speculative decoding across all evaluated benchmarks while fully preserving generation quality. Training draft models on their own generated outputs, rather than fixed ground-truth data, was found to be essential for achieving strong alignment and continuous efficiency gains. Furthermore, draft models distilled on a single task demonstrated strong cross-task transferability, achieving an average 26% speedup across 23 unseen reasoning tasks. In comprehensive model garden evaluations, first distilling a large teacher model into an intermediate target model and then applying DistillSpec to create the draft model achieved an overall 6-fold to 10-fold reduction in decoding latency with virtually no loss in performance.
These results have direct strategic implications for reducing the infrastructure overhead, energy consumption, and serving costs of large language model applications. DistillSpec demonstrates that a small model's standalone benchmark score does not predict how well it will perform as a speculative drafter; instead, direct distribution alignment is what drives inference efficiency. The article also demonstrates that combining DistillSpec with controlled lossy speculative decoding allows organizations to finely tune latency against acceptable quality boundaries depending on operational requirements.
Engineering and deployment teams should adopt DistillSpec as a drop-in enhancement for speculative decoding pipelines. Practitioners should use the draft model's own on-policy generations to minimize training compute costs and treat the choice of divergence function as a task-specific hyperparameter to tune. While these conclusions are supported with high confidence across diverse model sizes and benchmarks, testing was conducted in single-device, batch-size-one environments. Organizations should validate these gains in high-concurrency, distributed serving environments to verify real-world throughput improvements.
- Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). Introduces the foundational framework of speculative decoding that DistillSpec specifically aims to accelerate by improving draft-target model alignment.
- Paper: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Rishabh Agarwal et al. (2024). Pioneers Generalized Knowledge Distillation using on-policy sampling and flexible divergence metrics, providing the core distillation paradigms adopted in DistillSpec.
- Paper: MiniLLM: On-Policy Distillation of Large Language Models, Yuxian Gu et al. (2024). Establishes reverse Kullback-Leibler divergence and on-policy optimization principles for text generation distillation that DistillSpec tailors to speculative drafting.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). Supplies the foundational theory and mechanism of knowledge distillation that underlies the training of compact draft models.
- Paper: Sequence-Level Knowledge Distillation, Yoon Kim et al. (2016). Develops sequence-level knowledge distillation methods for autoregressive generation that underpin modern LLM student-teacher alignment.
- Paper: EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees, Yuhui Li et al. (2024). Extends speculative decoding acceleration by dynamically constructing context-aware draft trees to further boost verification acceptance rates.
- Paper: EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test, Yuhui Li et al. (2025). Advances draft model training by incorporating multi-layer target features and simulated multi-step test dynamics during drafting.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). Applies tree-based speculative candidate generation and parallel batch verification to optimize speculative serving systems.
- Paper: Unlocking Lossless Speedups in LLMs via Discrete Diffusion, Subham Sekhar Sahoo et al. (2026). Explores alternative drafting architectures by utilizing distilled discrete diffusion pathways for parallel proposal generation in speculative verification.
- Paper: OPRD: On-Policy Representation Distillation, Shenzhi Yang et al. (2026). Extends on-policy distillation by shifting supervision from token output distributions to intermediate representation matching.
- Paper: Trajectory-Refined Distillation, Li Jiang et al. (2026). Refines on-policy distillation methods by addressing multi-step rollout errors through trajectory-level teacher corrections.
