EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Yuhui LiFangyun WeiChao ZhangHongyang Zhang
Proposes EAGLE-2, a lossless speculative sampling method that dynamically shapes draft trees based on context-dependent confidence scores to accelerate large language model inference by up to 4.26x without extra model training.
Deploying modern large language models is computationally expensive and slow because generating text sequentially requires loading billions of parameters for every single generated word or subword. While speculative sampling—a technique where a lightweight model drafts multiple candidate words and the primary model verifies them in parallel—helps alleviate this bottleneck, conventional approaches rely on static draft tree structures. These fixed structures operate on the flawed assumption that candidate acceptance depends solely on position, ignoring how linguistic context affects predictability.
The article introduces and evaluates EAGLE-2, a training-free framework designed to accelerate language model generation by dynamically adjusting draft trees based on real-time context. The primary objective is to demonstrate that context-aware drafting increases verification acceptance rates and significantly boosts inference speeds without altering model outputs.
To establish this approach, the researchers first confirmed that draft acceptance rates vary widely across different contexts and discovered that the confidence scores of the drafting model reliably predict the actual likelihood of candidate acceptance. Leveraging this calibration, EAGLE-2 uses an expansion and reranking mechanism to dynamically shape the draft tree, prioritizing the most globally promising candidate paths. The authors evaluated the system across three leading model families (Vicuna, LLaMA-2, and LLaMA-3, spanning 7-billion to 70-billion parameters) and benchmarked performance across six core tasks, including multi-turn dialogue, coding, mathematical reasoning, summarization, and question answering.
The findings show that EAGLE-2 consistently outperforms existing speculative sampling baselines across all evaluated settings. First, EAGLE-2 delivers speedup ratios between 2.5x and 5.0x relative to standard autoregressive generation, surpassing its predecessor EAGLE-1 by roughly 20% to 40% and running up to twice as fast as other multi-token methods like Medusa. Second, it yields an average acceptance length of 4.0 to 5.5 tokens per verification cycle, nearly doubling the acceptance length of traditional speculative decoding. Third, the highest performance gains occur in structured tasks such as code generation, achieving up to 5x speedups. Finally, ablation studies confirm that both path-based value expansion and dynamic reranking are essential drivers of these performance improvements.
These results demonstrate that inference latency and server computing costs can be substantially reduced without compromising model accuracy or safety. Because EAGLE-2 enforces strict verification criteria and leaves the target model's underlying weights intact, it is provably lossless and preserves the exact statistical distribution of the generated text. Organizations can achieve higher throughput and lower operational overhead out of the box without changing core model behaviors.
Engineering teams should consider adopting EAGLE-2 for high-throughput serving environments, especially where low latency is critical for user experience. When deploying, teams should account for the one-time computational cost of preparing draft models, which requires approximately one to two days of multi-GPU compute for 70-billion-parameter models. Organizations should also consider incorporating broader knowledge datasets during draft model pre-training to improve performance on fact-based question answering and summarization tasks.
The primary operational limitation is the requirement for a well-calibrated draft model trained on supervised instruction data, along with slight performance variability on open-domain knowledge queries. Nevertheless, the experimental evidence across diverse benchmarks provides high confidence that EAGLE-2 delivers robust, lossless inference acceleration for modern language model architectures.
- Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). Introduces the foundational speculative decoding and speculative sampling paradigms that EAGLE-2 builds upon and dynamically enhances.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). Establishes tree-based speculative inference and parallel verification for large language models, providing the baseline static tree drafting context modified by EAGLE-2.
- Paper: EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test, Yuhui Li et al. (2025). Directly succeeds EAGLE-2 by introducing multi-layer feature fusion and training-time test procedures to scale drafting acceleration beyond dynamic draft trees.
