EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Yuhui LiFangyun WeiChao ZhangHongyang Zhang

article2024EMNLP245 citations

Proposes EAGLE-2, a lossless speculative sampling method that dynamically shapes draft trees based on context-dependent confidence scores to accelerate large language model inference by up to 4.26x without extra model training.

Listen

Deploying modern large language models is computationally expensive and slow because generating text sequentially requires loading billions of parameters for every single generated word or subword. While speculative sampling—a technique where a lightweight model drafts multiple candidate words and the primary model verifies them in parallel—helps alleviate this bottleneck, conventional approaches rely on static draft tree structures. These fixed structures operate on the flawed assumption that candidate acceptance depends solely on position, ignoring how linguistic context affects predictability.

The article introduces and evaluates EAGLE-2, a training-free framework designed to accelerate language model generation by dynamically adjusting draft trees based on real-time context. The primary objective is to demonstrate that context-aware drafting increases verification acceptance rates and significantly boosts inference speeds without altering model outputs.

To establish this approach, the researchers first confirmed that draft acceptance rates vary widely across different contexts and discovered that the confidence scores of the drafting model reliably predict the actual likelihood of candidate acceptance. Leveraging this calibration, EAGLE-2 uses an expansion and reranking mechanism to dynamically shape the draft tree, prioritizing the most globally promising candidate paths. The authors evaluated the system across three leading model families (Vicuna, LLaMA-2, and LLaMA-3, spanning 7-billion to 70-billion parameters) and benchmarked performance across six core tasks, including multi-turn dialogue, coding, mathematical reasoning, summarization, and question answering.

The findings show that EAGLE-2 consistently outperforms existing speculative sampling baselines across all evaluated settings. First, EAGLE-2 delivers speedup ratios between 2.5x and 5.0x relative to standard autoregressive generation, surpassing its predecessor EAGLE-1 by roughly 20% to 40% and running up to twice as fast as other multi-token methods like Medusa. Second, it yields an average acceptance length of 4.0 to 5.5 tokens per verification cycle, nearly doubling the acceptance length of traditional speculative decoding. Third, the highest performance gains occur in structured tasks such as code generation, achieving up to 5x speedups. Finally, ablation studies confirm that both path-based value expansion and dynamic reranking are essential drivers of these performance improvements.

These results demonstrate that inference latency and server computing costs can be substantially reduced without compromising model accuracy or safety. Because EAGLE-2 enforces strict verification criteria and leaves the target model's underlying weights intact, it is provably lossless and preserves the exact statistical distribution of the generated text. Organizations can achieve higher throughput and lower operational overhead out of the box without changing core model behaviors.

Engineering teams should consider adopting EAGLE-2 for high-throughput serving environments, especially where low latency is critical for user experience. When deploying, teams should account for the one-time computational cost of preparing draft models, which requires approximately one to two days of multi-GPU compute for 70-billion-parameter models. Organizations should also consider incorporating broader knowledge datasets during draft model pre-training to improve performance on fact-based question answering and summarization tasks.

The primary operational limitation is the requirement for a well-calibrated draft model trained on supervised instruction data, along with slight performance variability on open-domain knowledge queries. Nevertheless, the experimental evidence across diverse benchmarks provides high confidence that EAGLE-2 delivers robust, lossless inference acceleration for modern language model architectures.

Cover for EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Abstract

Inference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution. Most speculative sampling methods such as EAGLE use a static draft tree, implicitly assuming that the acceptance rate of draft tokens depends only on their position. Interestingly, we found that the acceptance rate of draft tokens is also context-dependent. In this paper, building upon EAGLE, we propose EAGLE-2, which introduces a new technique of context-aware dynamic draft tree into drafting modeling. This improvement leverages the fact that the draft model of EAGLE is well-calibrated: the confidence scores from the draft model approximate acceptance rates with small errors. We conducted extensive evaluations on three series of LLMs and six tasks, with EAGLE-2 achieving speedup ratios 3.05x-4.26x, which is 20%-40% faster than EAGLE-1. EAGLE-2 also ensures that the distribution of the generated text remains unchanged, making it a lossless acceleration algorithm. The code is open sourced at https://github.com/SafeAILab/EAGLE.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Speculative Sampling
  • 2.2 EAGLE
  • 3 Observations
  • 3.1 Context-Dependent Acceptance Rates
  • 3.2 Well-Calibrated Draft Model
  • 4 Context-Aware Dynamic Draft Tree
  • 4.1 Expansion Phase
  • 4.2 Reranking Phase
  • 5.1 Effectiveness
  • 5.2 Ablation Study
  • 5.2.1 Value and Confidence Score
  • 5.2.2 Reranking
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Implementation Details

Knowls

  1. Knowl 1 — EAGLE-2 Context-Aware Dynamic Draft Tree Generation Algorithm

    algorithm

    EAGLE-2 is a speculative decoding algorithm that dynamically adapts the structure of a draft token tree during inference based on draft model confidence estimates, followed by verification using the original large language model (LLM).

    Input: Original target model MM, draft model MdM_d, initial prefix token sequence T1:jT_{1:j}, maximum tree depth DD, expansion candidate width kk, target draft size mm
    Output: Generated tokens accepted by the target model MM
    Initialize draft tree root with the current state
    for layer d=1d = 1 to DD do
        Identify the active leaf nodes in layer d−1d-1
        Select top-kk nodes with the highest cumulative path value Vi=∏tj∈Path(root,ti)cjV_i = \prod_{t_j \in \text{Path}(\text{root}, t_i)} c_j
        Input the selected nodes into MdM_d using tree attention to predict candidate child tokens and confidence scores cc
        Attach generated child tokens to the draft tree
    end for
    Rerank all generated nodes in the draft tree by their cumulative path value ViV_i, prioritizing shallower nodes in case of ties
    Select the top-mm highest-value nodes from the draft tree
    Flatten the selected mm nodes into a 1-dimensional token sequence T^\hat{T}
    Construct a tree attention mask AA where A[u,v]=1A[u, v] = 1 if node vv is an ancestor of node uu or u=vu = v, and 00 otherwise
    Execute a single forward pass of target model MM on T^\hat{T} using attention mask AA to obtain exact token probabilities
    Perform standard speculative verification along tree paths, accepting tokens satisfying acceptance criteria and sampling from the residual distribution upon rejection
    Update KV-cache and prefix TT with accepted tokens and the correction token
    return Updated sequence TT

    In standard deployments, depth D=6D=6 and expansion width k=10k=10 are used across configurations, while target budget mm is set to 60 for 7B/8B models, 50 for 13B models, and 48 for 70B models.

  2. Knowl 2 — Cumulative Path Acceptance Value Formulation

    equation

    In tree-structured speculative sampling, a token tit_i can only be accepted by the target model if all of its antecedent tokens along the branch from the root are also accepted. The global acceptance probability of a candidate draft token tit_i is given by the cumulative path value ViV_i:

    Vi=∏tj∈Path(root,ti)pj≈∏tj∈Path(root,ti)cjV_i = \prod_{t_j \in \text{Path}(\text{root}, t_i)} p_j \approx \prod_{t_j \in \text{Path}(\text{root}, t_i)} c_j

    where Path(root,ti)\text{Path}(\text{root}, t_i) denotes the set of all nodes along the directed path from the root node to node tit_i in the draft tree, pj∈[0,1]p_j \in [0, 1] is the true conditional probability of token tjt_j being accepted by the original target model given its prefix, and cj∈[0,1]c_j \in [0, 1] is the local confidence score (output probability) assigned to token tjt_j by the draft model.

    Because cj∈[0,1]c_j \in [0, 1], the cumulative value ViV_i is monotonically non-increasing with tree depth (Vchild≤VparentV_{\text{child}} \le V_{\text{parent}}), ensuring that deeper draft tokens naturally receive lower priority unless supported by high confidence along the entire ancestral path.

  3. Knowl 3 — Draft Model Confidence Score Calibration as Acceptance Predictor

    empirical result

    Empirical evaluations on Vicuna 7B across the Alpaca instruction dataset show that the confidence score cjc_j (the output probability from the draft model for token tjt_j) is a well-calibrated estimator of the token's empirical acceptance rate pjp_j by the original LLM.

    When grouping draft tokens into confidence intervals [0,1][0, 1]:

    • Draft tokens with draft model confidence score cj<0.05c_j < 0.05 achieve an empirical acceptance rate pj≈0.04p_j \approx 0.04.
    • Draft tokens with draft model confidence score cj>0.95c_j > 0.95 achieve an empirical acceptance rate pj≈0.98p_j \approx 0.98.
    • The average acceptance rate scales in approximate linearity with the confidence score across intermediate deciles.

    This calibration enables the speculative decoding algorithm to use draft model output probabilities as a proxy for target LLM acceptance rates without executing target LLM forward passes during the drafting phase.

  4. Knowl 4 — Context-Dependence of Token Acceptance Rates in Speculative Draft Trees

    empirical result

    Evaluation of token acceptance rates across various tree positions (P1 through P6) on Vicuna 7B with the Alpaca dataset demonstrates that acceptance rates are both position-dependent and context-dependent:

    1. Position Dependence: Average acceptance rates are highest for top-left positions (e.g., node P1) and lowest for bottom-right positions (e.g., node P6), justifying asymmetric static tree designs used in prior methods.
    2. Context Dependence: For any single fixed tree position (such as P1 or P4), the acceptance rate exhibits substantial variance across different prompt contexts, spanning values from near 0.00.0 to 1.01.0.

    Consequently, static tree allocations allocate excess tokens to deterministic contexts where a single sequence suffices, while under-allocating branching candidates in ambiguous contexts where multiple plausible token continuations exist.

  5. Knowl 5 — Global Draft Tree Reranking and Ancestor Attention Masking

    model/method

    During the drafting phase of EAGLE-2, dynamic expansion deepens high-probability branches, but some deeper expanded nodes may end up with cumulative values ViV_i lower than unexpanded shallow nodes. To maximize total accepted tokens per verification cycle, EAGLE-2 applies a global reranking step:

    1. Reranking: All generated draft nodes across all layers are pooled and sorted in descending order of their cumulative path values Vi=∏tj∈Path(root,ti)cjV_i = \prod_{t_j \in \text{Path}(\text{root}, t_i)} c_j. Ties are broken by selecting shallower nodes first, ensuring that the selected top-mm nodes always form a valid, connected sub-tree containing the root.
    2. Flattening: The top-mm nodes are serialized into a 1D token sequence T^=[t^1,t^2,…,t^m]\hat{T} = [\hat{t}_1, \hat{t}_2, \dots, \hat{t}_m].
    3. Custom Tree Attention Mask: To evaluate multiple branches in parallel during a single forward pass of the target LLM without cross-branch contamination, the attention matrix A∈{0,1}m×mA \in \{0, 1\}^{m \times m} is masked such that token t^a\hat{t}_a can attend to token t^b\hat{t}_b if and only if t^b\hat{t}_b is an ancestor of t^a\hat{t}_a in the tree or a=ba = b.
  6. Knowl 6 — Speedup Ratios and Acceptance Lengths Across Multi-Domain Benchmarks

    data/table

    Performance of EAGLE-2 compared against lossless speculative decoding methods on Vicuna (V) and LLaMA2-Chat (L2) models across six tasks (MT-bench, HumanEval, GSM8K, Alpaca, CNN/DailyMail, Natural Questions) measured by wall-clock speedup ratio relative to vanilla autoregressive decoding and average acceptance length τ\tau (accepted tokens per verification step):

    Setting Method MT-bench HumanEval GSM8K Alpaca CNN/DM Nat. Ques. Mean Speedup Mean τ\tau
    Temperature = 0
    V 13B SpS (Vicuna-68M) 1.93x 2.23x 1.77x 1.76x 1.93x 1.66x 1.88x 2.18
    PLD 1.58x 1.85x 1.68x 1.16x 2.42x 1.14x 1.64x 1.69
    Medusa 2.07x 2.50x 2.23x 2.08x 1.71x 1.81x 2.07x 2.44
    Lookahead 1.65x 1.71x 1.81x 1.46x 1.46x 1.36x 1.58x 1.62
    Hydra 2.88x 3.28x 2.93x 2.86x 2.05x 2.11x 2.69x 3.40
    EAGLE 3.07x 3.58x 3.08x 3.03x 2.49x 2.42x 2.95x 3.82
    EAGLE-2 4.26x 4.96x 4.22x 4.25x 3.40x 3.13x 4.04x 4.65
    L2 13B PLD 1.42x 1.63x 1.41x 1.16x 1.42x 1.12x 1.36x 1.40
    Lookahead 1.58x 1.80x 1.65x 1.47x 1.46x 1.42x 1.56x 1.61
    EAGLE 3.03x 3.76x 3.20x 3.01x 2.70x 2.83x 3.09x 3.89
    EAGLE-2 4.21x 5.00x 4.31x 4.13x 3.45x 3.51x 4.10x 4.68
    V 7B EAGLE 2.90x 3.33x 3.01x 2.79x 2.33x 2.31x 2.78x 3.79
    EAGLE-2 3.62x 3.95x 3.63x 3.46x 2.94x 2.76x 3.39x 4.68
    L2 7B EAGLE 2.78x 3.17x 2.91x 2.78x 2.43x 2.61x 2.78x 3.71
    EAGLE-2 3.43x 4.03x 3.52x 3.45x 3.01x 3.15x 3.43x 4.64
    Temperature = 1
    V 13B EAGLE 2.32x 2.65x 2.57x 2.45x 2.23x 2.14x 2.39x 3.39
    EAGLE-2 3.80x 4.22x 3.77x 3.78x 3.25x 3.07x 3.65x 4.26
    L2 13B EAGLE 2.68x 2.89x 2.82x 2.66x 2.41x 2.37x 2.64x 3.53
    EAGLE-2 3.92x 4.58x 4.21x 3.85x 3.31x 3.43x 3.88x 4.51
    V 7B EAGLE 2.13x 2.39x 2.34x 2.21x 2.08x 1.95x 2.18x 3.20
    EAGLE-2 3.05x 3.33x 3.07x 3.08x 2.63x 2.48x 2.94x 4.20
    L2 7B EAGLE 2.22x 2.61x 2.40x 2.29x 2.19x 2.22x 2.32x 3.37
    EAGLE-2 3.19x 3.67x 3.35x 3.20x 2.73x 2.81x 3.15x 4.41

    EAGLE-2 achieves the highest speedup across all evaluated models and datasets, generating 4.0–5.5 accepted tokens per cycle (roughly 2x that of standard speculative sampling and Medusa), yielding a 20%–40% relative speedup over EAGLE-1.

  7. Knowl 7 — Speedup and Acceptance Length on 70B and LLaMA-3 Models

    data/table

    Evaluation of speculative decoding methods on 70B parameters and LLaMA-3 models on the MT-bench dataset at temperature = 0:

    Target Model Method Speedup Ratio Average Acceptance Length τ\tau
    LLaMA2-Chat 70B PLD 1.31x 1.39
    Lookahead 1.52x 1.64
    EAGLE 3.01x 3.81
    EAGLE-2 3.51x 4.48
    LLaMA3-Instruct 70B EAGLE 2.83x 3.62
    EAGLE-2 3.29x 4.16
    LLaMA3-Instruct 8B EAGLE 2.72x 3.65
    EAGLE-2 3.46x 4.53

    On large 70B architectures and next-generation LLaMA-3 foundation models, dynamic tree drafting in EAGLE-2 consistently improves both acceptance length τ\tau and wall-clock acceleration over static drafting approaches.

  8. Knowl 8 — Ablation of Path Value Expansion and Global Tree Reranking

    data/table

    Ablation experiments conducted on Vicuna 7B at temperature = 0 on MT-bench and GSM8K isolating the contributions of cumulative path value ViV_i (versus greedy single-step local confidence cic_i) and global tree reranking:

    MT-bench GSM8K
    Configuration Speedup τ\tau Speedup τ\tau
    w/o both 2.81x 3.92 2.85x 3.93
    w/o value (confidence only) 3.21x 4.39 2.93x 3.96
    w/o reranking 3.48x 4.86 3.50x 4.85
    EAGLE-2 (Full) 3.62x 4.98 3.63x 4.97

    Using the cumulative path value ViV_i instead of local confidence prevents expanding unlikely deep branches attached to low-confidence parents. Performing global reranking ensures that unexpanded shallow nodes with high confidence are prioritized for verification over low-confidence deep leaves.

  9. Knowl 9 — Training Overhead and Calibration Assumptions in EAGLE-2

    limitation

    EAGLE-2 has the following computational constraints and operational assumptions:

    1. Draft Model Training: While EAGLE-2 requires no extra auxiliary models or routing networks beyond EAGLE-1's draft head, it relies on training a lightweight draft model on supervised fine-tuning (SFT) data (e.g., ShareGPT). Training a draft model for a 70B target LLM requires approximately 4 NVIDIA A100 (40GB) GPUs running for 1 to 2 days.
    2. Calibration Sensitivity: Dynamic expansion relies directly on the assumption that draft model confidence scores cjc_j accurately approximate true target LLM acceptance rates pjp_j. If the draft model is poorly calibrated or evaluated on out-of-distribution domains poorly covered by SFT data (such as open-domain factual knowledge in Natural Questions or summarization in CNN/DailyMail), average acceptance lengths decrease.

Coverage note — None was omitted; all key contributions including empirical calibration findings, dynamic draft tree formulation, quantitative benchmark results, ablation analysis, and stated limitations are covered.

References

  1. 1.Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon. 2024. Hydra: Sequentially-dependent draft heads for medusa decoding. arXiv preprint arXiv:2402.05109.
  2. 2.Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. 2024. Medusa: Simple llm inference acceleration framework with multiple decoding heads. arXiv preprint arXiv: 2401.10774.
  3. 3.Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023a. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  5. 5.Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen. 2024. Sequoia: Scalable, robust, and hardware-aware speculative decoding. arXiv preprint arXiv:2402.12374.
  6. 6.Ziyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun, Jie Huang, and Kevin Chen-Chuan Chang. 2023b. Cascade speculative drafting for even faster llm inference. arXiv preprint arXiv:2312.11462.
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  9. 9.Cunxiao Du, Jing Jiang, Xu Yuanchen, Jiawei Wu, Sicheng Yu, Yongqi Li, Shenggui Li, Kai Xu, Liqiang Nie, Zhaopeng Tu, et al. 2024. Glide with a cape: A low-hassle method to accelerate speculative decoding. arXiv preprint arXiv:2402.02082.
  10. 10.Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, et al. 2024. Layer skip: Enabling early exit inference and self-speculative decoding. arXiv preprint arXiv:2404.16710.
  11. 11.Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2023. Breaking the sequential dependency of llm inference using lookahead decoding.
  12. 12.Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024. Break the sequential dependency of llm inference using lookahead decoding. arXiv preprint arXiv:2402.02057.
  13. 13.Trevor Gale, Erich Elsen, and Sara Hooker. 2019. The state of sparsity in deep neural networks.(2019). arXiv preprint cs.LG/1902.09574.
  14. 14.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
  15. 15.Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Hasan Genc, Kurt Keutzer, Amir Gholami, and Sophia Shao. 2023. Speed: Speculative pipelined execution for efficient decoding. arXiv preprint arXiv:2310.12072.
  16. 16.Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2018. Quantized neural networks: Training neural networks with low precision weights and activations. journal of machine learning research, 18(187):1–30.
  17. 17.Joao Gante. 2023. Assisted generation: a new direction toward low-latency text generation.
  18. 18.Parsa Kavehzadeh, Mojtaba Valipour, Marzieh Tahaei, Ali Ghodsi, Boxing Chen, and Mehdi Rezagholizadeh. 2024. Sorted llama: Unlocking the potential of intermediate layers of large language models for dynamic inference. In Findings of the Association for Computational Linguistics: EACL 2024, pages 2129–2145.
  19. 19.Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2021. I-bert: Integer-only bert quantization. In International conference on machine learning, pages 5506–5518. PMLR.
  20. 20.Sehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik, Michael W Mahoney, Amir Gholami, and Kurt Keutzer. 2024. Speculative decoding with big little decoder. Advances in Neural Information Processing Systems, 36.
  21. 21.Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, and Hao Zhang. 2024. Cllms: Consistency large language models. arXiv preprint arXiv:2403.00835.
  22. 22.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  23. 23.Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR.
  24. 24.Minghan Li, Xilun Chen, Ari Holtzman, Beidi Chen, Jimmy Lin, Wen-tau Yih, and Xi Victoria Lin. 2024a. Nearest neighbor speculative decoding for llm generation and attribution. arXiv preprint arXiv:2405.19325.
  25. 25.Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024b. Eagle: Speculative sampling requires rethinking feature uncertainty. In International Conference on Machine Learning.
  26. 26.Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Kai Han, and Yunhe Wang. 2024. Kangaroo: Lossless self-speculative decoding via double early exiting. arXiv preprint arXiv:2404.18911.
  27. 27.Xiaoxuan Liu, Lanxiang Hu, Peter Bailis, Ion Stoica, Zhijie Deng, Alvin Cheung, and Hao Zhang. 2023a. Online speculative decoding. arXiv preprint arXiv:2310.07177.
  28. 28.Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al. 2023b. Deja vu: Contextual sparsity for efficient llms at inference time. In International Conference on Machine Learning, pages 22137–22176. PMLR.
  29. 29.Meta. 2024. LLaMA3. https://github.com/pytorch-labs/gpt-fast/.
  30. 30.Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2023. SpecInfer: Accelerating generative LLM serving with speculative inference and token tree verification. arXiv preprint arXiv:2305.09781.
  31. 31.Giovanni Monea, Armand Joulin, and Edouard Grave. 2023. Pass: Parallel speculative sampling. arXiv preprint arXiv:2311.13581.
  32. 32.Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023.
  33. 33.R OpenAI. 2023. Gpt-4 technical report. arxiv 2303.08774. View in Article, 2(5).
  34. 34.Victor Sanh, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389.
  35. 35.Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, and Emanuele Rodolà. 2023. Accelerating transformer inference for translation via parallel decoding. arXiv preprint arXiv:2305.10427.
  36. 36.Apoorv Saxena. 2023. Prompt lookup decoding.
  37. 37.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020. Q-bert: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8815–8821.
  38. 38.Benjamin Spector and Chris Re. 2023. Accelerating llm inference with staged speculative decoding. arXiv preprint arXiv:2308.04623.
  39. 39.Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. 2018. Blockwise parallel decoding for deep autoregressive models. Advances in Neural Information Processing Systems, 31.
  40. 40.Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen. 2024a. Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding. arXiv preprint arXiv:2404.11912.
  41. 41.Xin Sun, Tao Ge, Furu Wei, and Houfeng Wang. 2021. Instantaneous grammatical error correction with shallow aggressive decoding. arXiv preprint arXiv:2106.04970.
  42. 42.Ziteng Sun, Jae Hun Ro, Ahmad Beirami, and Ananda Theertha Suresh. 2024b. Optimal blocklevel draft verification for accelerating speculative decoding. arXiv preprint arXiv:2403.10444.
  43. 43.Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. 2024c. Spectr: Fast speculative decoding via optimal transport. Advances in Neural Information Processing Systems, 36.
  44. 44.Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Zhihao Jia, and Max Ryabinin. 2024. Specexec: Massively parallel speculative decoding for interactive llm inference on consumer devices. arXiv preprint arXiv:2406.02532.
  45. 45.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  46. 46.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models (2023). arXiv preprint arXiv:2302.13971.
  47. 47.Neeraj Varshney, Agneet Chatterjee, Mihir Parmar, and Chitta Baral. 2023. Accelerating llm inference by enabling intermediate layer decoding. arXiv preprint arXiv:2310.18581.
  48. 48.Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. 2024. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. Preprint, arXiv:2401.07851.
  49. 49.Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023a. Inference with reference: Lossless acceleration of large language models. arXiv preprint arXiv:2304.04487.
  50. 50.Seongjun Yang, Gibbeum Lee, Jaewoong Cho, Dimitris Papailiopoulos, and Kangwook Lee. 2023b. Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding. arXiv preprint arXiv:2307.05908.
  51. 51.Hanling Yi, Feng Lin, Hongbin Li, Peiyang Ning, Xiaotian Yu, and Rong Xiao. 2024. Generation meets verification: Accelerating large language model inference with smart parallel auto-correct decoding. arXiv preprint arXiv:2402.11809.
  52. 52.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020. Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 811–824. IEEE.
  53. 53.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019. Q8bert: Quantized 8bit bert. In 2019 Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS), pages 36–39. IEEE.
  54. 54.Aonan Zhang, Chong Wang, Yi Wang, Xuanyu Zhang, and Yunfei Cheng. 2024. Recurrent drafter for fast speculative decoding in large language models. arXiv preprint arXiv:2403.09919.
  55. 55.Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. 2023. Draft & verify: Lossless large language model acceleration via self-speculative decoding. arXiv preprint arXiv:2309.08168.
  56. 56.Weilin Zhao, Yuxiang Huang, Xu Han, Chaojun Xiao, Zhiyuan Liu, and Maosong Sun. 2024. Ouroboros: Speculative decoding with large model enhanced drafting. arXiv preprint arXiv:2402.13720.
  57. 57.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  58. 58.Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. 2024. Distillspec: Improving speculative decoding via knowledge distillation. In The Twelfth International Conference on Learning Representations.

Citation

MLA
Li, Y., et al. “EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 7421–32, https://doi.org/10.18653/v1/2024.emnlp-main.422.
APA
Li, Y., Wei, F., Zhang, C., & Zhang, H. (2024). EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7421–7432. https://doi.org/10.18653/v1/2024.emnlp-main.422
Chicago
Li, Y., F. Wei, C. Zhang, and H. Zhang. 2024. “EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7421–32. https://doi.org/10.18653/v1/2024.emnlp-main.422.
Harvard
Li, Y. et al. (2024) “EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7421–7432. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.422.
Vancouver
1. Li Y, Wei F, Zhang C, Zhang H (2024) EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7421–7432

BibTeX

@inproceedings{li-etal-2024-eagle,
    title = "{EAGLE}-2: Faster Inference of Language Models with Dynamic Draft Trees",
    author = "Li, Yuhui  and
      Wei, Fangyun  and
      Zhang, Chao  and
      Zhang, Hongyang",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.422/",
    doi = "10.18653/v1/2024.emnlp-main.422",
    pages = "7421--7432"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/