Speculative Decoding with Big Little Decoder

Sehoon KimKarttikeya MangalamSuhong MoonJitendra MalikMichael W. MahoneyAmir GholamiKurt Keutzer

article2023NeurIPS180 citations

Proposes Big Little Decoder, a plug-and-play speculative decoding framework that pairs a small autoregressive model with an occasionally invoked large model via fallback and rollback policies to speed up text generation by up to 2.12× without requiring retraining or architectural changes.

Listen

Large language models based on transformer architectures have enabled significant breakthroughs in natural language processing. However, generating text with these models remains slow and computationally expensive because they typically produce text sequentially, one word or token at a time. This autoregressive process forces hardware to constantly reload model weights for every single token, causing memory bottlenecks, poor hardware utilization, and high latency that hinder real-time deployment.

The article introduces and evaluates the Big Little Decoder (BiLD), a plug-and-play framework designed to accelerate text generation without altering underlying model architectures or retraining pipelines. The objective is to demonstrate that pairing a fast, compact model with an accurate, larger model can substantially reduce inference latency while preserving generation quality across multiple natural language tasks.

To evaluate this framework, the authors conducted extensive experiments on standard machine translation benchmarks (IWSLT 2017 and WMT 2014 German-to-English) and text summarization benchmarks (XSUM and CNN/DailyMail). The setup used standard transformer models where the large model was roughly twenty times the size of the smaller companion. Inference performance was measured in a standard cloud environment on an NVIDIA T4 GPU. The framework operates by letting the small model generate text quickly and sequentially, while invoking the large model only when needed. Coordination relies on two core rules: a fallback policy that hands control to the large model when the small model is uncertain, and a rollback policy that uses the large model to verify and correct past tokens in parallel. An optional calibration step, termed model prediction alignment, was also evaluated to train the small model to mimic the vocabulary preferences of the large model.

The findings show that small models share substantial agreement with large models and only require occasional corrections to match their output quality. Across the evaluated tasks, the standard plug-and-play BiLD setup achieved an average speedup of 1.50× with no degradation in text quality, reaching up to 1.71× on summarization. When combined with the prediction alignment technique, the system achieved up to 1.85× speedup with zero quality loss and reached up to 2.12× speedup when accepting a minor quality degradation of approximately one point on standard evaluation metrics. System analysis revealed that while total floating-point arithmetic slightly increased, BiLD reduced memory operations roughly fivefold, boosting arithmetic intensity and eliminating hardware memory bottlenecks. Furthermore, the approach consistently outperformed alternative speculative decoding and early-exiting techniques across both translation and summarization tasks.

These results indicate that organizations deploying large language models for real-time applications can achieve near-double throughput and lower latency without re-architecting systems or conducting costly full-model retraining. By addressing the hardware memory bottleneck rather than merely minimizing arithmetic operations, collaborative model execution provides a viable, cost-effective path to scaling interactive artificial intelligence services.

Engineering and deployment teams should consider implementing collaborative decoding policies like BiLD in latency-sensitive, batch-size-one inference pipelines. Organizations seeking maximum throughput should also adopt the lightweight prediction alignment fine-tuning step. Before full production rollout, teams should conduct task-specific threshold sweeps to balance quality against latency and validate performance across larger model families and diverse hardware configurations.

The evaluation carries high confidence within single-batch online serving scenarios on common hardware platforms. However, the study primarily focused on translation and summarization tasks up to moderate sequence lengths, and results may vary under large batch sizes or distinct domain workloads where memory access patterns differ.

  • Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). This seminal work establishes the foundational paradigm of speculative decoding, in which a small draft model autoregressively proposes tokens that a large model verifies in parallel.
  • Paper: Confident Adaptive Language Modeling, Tal Schuster et al. (2022). This paper introduces confidence-based dynamic computation allocation and early exits during language model generation, directly motivating confidence-driven fallback and coordination policies between draft and target models.
  • Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). It provides fundamental analysis and architectural solutions for overcoming the memory-bandwidth bottleneck in autoregressive incremental decoding.
  • Paper: Orca: A Distributed Serving System for Transformer-Based Generative Models, Gyeong-In Yu et al. (2022). It formalizes iteration-level scheduling and the systems-level latency challenges of autoregressive text generation that collaborative multi-model decoding addresses.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). It introduces the core Transformer sequence-to-sequence architecture underpinning the models and text generation tasks studied in the paper.
Cover for Speculative Decoding with Big Little Decoder

Abstract

The recent emergence of Large Language Models based on the Transformer architecture has enabled dramatic advancements in the field of Natural Language Processing. However, these models have long inference latency, which limits their deployment and makes them prohibitively expensive for various real-time applications. The inference latency is further exacerbated by autoregressive generative tasks, as models need to run iteratively to generate tokens sequentially without leveraging token-level parallelization. To address this, we propose Big Little Decoder (BiLD), a framework that can improve inference efficiency and latency for a wide range of text generation applications. The BiLD framework contains two models with different sizes that collaboratively generate text. The small model runs autoregressively to generate text with a low inference cost, and the large model is only invoked occasionally to refine the small model's inaccurate predictions in a non-autoregressive manner. To coordinate the small and large models, BiLD introduces two simple yet effective policies: (1) the fallback policy that determines when to hand control over to the large model; and (2) the rollback policy that determines when the large model needs to correct the small model's inaccurate predictions. To evaluate our framework across different tasks and models, we apply BiLD to various text generation scenarios encompassing machine translation on IWSLT 2017 De-En and WMT 2014 De-En, and summarization on XSUM and CNN/DailyMail. On an NVIDIA T4 GPU, our framework achieves a speedup of up to 2.12× speedup with minimal generation quality degradation. Furthermore, our framework is fully plug-and-play and can be applied without any modifications in the training process or model architecture. Our code is open-sourced¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Efficient Transformer Decoding Inference
  • 2.2 Use of Multiple Models
  • 3 Methodology
  • 3.1 Motivating Examples
  • 3.2 Problem Formulation
  • 3.3 Fallback Policy: Small Model Knows When to Stop Predictions
  • 3.4 Rollback Policy: Large Model Knows When to Revert Predictions
  • 3.5 Big Little Decoder
  • 3.5.1 Model Prediction Alignment
  • 4 Evaluations
  • 4.1 Experiment Setup
  • 4.2 Main Results
  • 4.3 Ablation Studies
  • 4.4 Early Exiting Strategy in the BiLD Framework
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • 7 Supplementary Material
  • 7.1 Experimental Details
  • 7.1.1 Training Details
  • 7.1.2 Evaluation Details
  • 7.2 Details of Early Exiting Strategy in the BiLD Framework
  • 7.2.1 Training and Evaluation Details
  • 7.2.2 Performance Comparison between BiLD and CALM
  • 7.3 Comparison with Other Speculative Decoding Frameworks
  • 7.3.1 Differences in methodology
  • 7.3.2 Quantitative Comparisons
  • 7.3.3 Insights on Better Latency and Performance
  • 7.4 BiLD with Sampling
  • 7.5 Additional Analysis
  • 7.5.1 Model Analysis of BiLD: FLOPs, MOPs, and Arithmetic Intensity
  • 7.5.2 Examples of Generated Sequences
  • 7.5.3 Impact of Fallback and Rollback on Performance

Knowls

  1. Knowl 1 — Big Little Decoder Framework

    model/method

    The Big Little Decoder (BiLD) framework accelerates autoregressive text generation by coordinating two decoder models of differing capacities: a lightweight small model SS and a high-capacity large model LL that share a common vocabulary. In BiLD, the small model runs autoregressively (generating tokens one by one) to produce the majority of the sequence with minimal latency and computational overhead. The large model is executed infrequently and non-autoregressively (processing multiple tokens in parallel) to evaluate and refine the small model's generated tokens.

    The framework coordinates the two models via two core policies:

    1. Fallback Policy: Operates during the small model's autoregressive generation. If the small model's prediction confidence for the next token falls below a predefined threshold, generation is halted, and control is transferred to the large model.
    2. Rollback Policy: Operates when the large model executes. The large model computes its own conditional token distributions across the entire sequence produced by the small model. If any previously generated token deviates significantly from the large model's distribution according to a distance metric, all tokens following that misprediction are rolled back (discarded), and the inaccurate token is replaced with the large model's prediction.

    Because the large model processes the accumulated sequence in parallel in a single forward pass, it amortizes the memory bandwidth cost of loading model weights across multiple tokens, substantially improving hardware arithmetic intensity and end-to-end decoding speed.

  2. Knowl 2 — Fallback Policy in Big Little Decoder

    model/method

    The fallback policy in the Big Little Decoder (BiLD) framework determines when the small model should stop generating tokens and hand control over to the large model based on token prediction confidence.

    At decoding step nn, given the prefix sequence y1:n−1=(y1,…,yn−1)y_{1:n-1} = (y_1, \dots, y_{n-1}), the small model outputs a probability distribution over the vocabulary pS(y∣y1:n−1)p_S(y \mid y_{1:n-1}). The policy evaluates the maximum token probability: max⁡ypS(y∣y1:n−1)\max_{y} p_S(y \mid y_{1:n-1})

    If this confidence is below a predetermined fallback threshold αFB∈(0,1)\alpha_{FB} \in (0, 1): max⁡ypS(y∣y1:n−1)<αFB\max_{y} p_S(y \mid y_{1:n-1}) < \alpha_{FB} the small model is deemed insufficiently confident. The fallback policy halts the small model and invokes the large model to predict the next token yn∼pL(y∣y1:n−1)y_n \sim p_L(y \mid y_{1:n-1}).

    This policy dynamically adjusts the fallback window size (the number of consecutive tokens generated by the small model) at runtime according to the model's confidence, avoiding arbitrary fixed stride intervals.

  3. Knowl 3 — Rollback Policy in Big Little Decoder

    model/method

    The rollback policy in the Big Little Decoder (BiLD) framework enables the large model to inspect and revert potentially erroneous predictions made by the small model during previous decoding iterations.

    When control falls back to the large model after the small model has produced tokens y1,…,yny_1, \dots, y_n, the large model processes the full prefix y1:ny_{1:n} in a single non-autoregressive forward pass, computing the output distributions pL(y∣y1:m−1)p_L(y \mid y_{1:m-1}) for all steps m∈{1,…,n}m \in \{1, \dots, n\}.

    The discrepancy between the small model's hard prediction ymy_m and the large model's distribution pLp_L is evaluated using cross-entropy loss (the negative log-likelihood of ymy_m under pLp_L): d(pS(y∣y1:m−1),pL(y∣y1:m−1))=−log⁡pL(ym∣y1:m−1)d(p_S(y \mid y_{1:m-1}), p_L(y \mid y_{1:m-1})) = -\log p_L(y_m \mid y_{1:m-1})

    The policy finds the minimum index m∈{1,…,n−1}m \in \{1, \dots, n-1\} such that: d(pS(y∣y1:m−1),pL(y∣y1:m−1))>αRBd(p_S(y \mid y_{1:m-1}), p_L(y \mid y_{1:m-1})) > \alpha_{RB} where αRB>0\alpha_{RB} > 0 is a predetermined rollback threshold. If such an index mm exists, all subsequent tokens (ym,…,yn)(y_m, \dots, y_n) are rolled back, and ymy_m is replaced by a token sampled from the large model: ym∼pL(y∣y1:m−1)y_m \sim p_L(y \mid y_{1:m-1}). If no step exceeds αRB\alpha_{RB}, no tokens are reverted, and the large model outputs the next token yn∼pL(y∣y1:n−1)y_n \sim p_L(y \mid y_{1:n-1}).

  4. Knowl 4 — Big Little Decoder Generation Algorithm

    algorithm

    The Big Little Decoder (BiLD) algorithm coordinates a small autoregressive language model SmallModel\text{SmallModel} and a large language model LargeModel\text{LargeModel} using confidence-based fallback and distance-based rollback policies to generate a sequence starting from token ⟨s⟩\langle\text{s}\rangle until ⟨eos⟩\langle\text{eos}\rangle is produced.

    Hyperparameters:

    • αFB∈(0,1)\alpha_{FB} \in (0, 1): fallback threshold.
    • αRB>0\alpha_{RB} > 0: rollback threshold.
    • d(pL,pS)d(p_L, p_S): cross-entropy distance measuring −log⁡pL(ym)-\log p_L(y_m) of the token ymy_m chosen by the small model.
    Input: SmallModel, LargeModel, fallback threshold αFB\alpha_{FB}, rollback threshold αRB\alpha_{RB}, distance metric dd
    Output: Generated token sequence yy
    y←[⟨s⟩]y \leftarrow [\langle\text{s}\rangle]
    while y[−1]≠⟨eos⟩y[-1] \neq \langle\text{eos}\rangle do
        pS←SmallModel(y)p_S \leftarrow \text{SmallModel}(y)
        if max⁡(pS[−1])>αFB\max(p_S[-1]) > \alpha_{FB} then
            y←y+[sample(pS[−1])]y \leftarrow y + [\text{sample}(p_S[-1])]
        else
            pL←LargeModel(y)p_L \leftarrow \text{LargeModel}(y)
            m←min⁡{i∈[1,∣y∣−1]∣d(pL[i],pS[i])>αRB}m \leftarrow \min \{ i \in [1, |y|-1] \mid d(p_L[i], p_S[i]) > \alpha_{RB} \}
            if m existsm \text{ exists} then
                y←y[1:m−1]+[sample(pL[m])]y \leftarrow y[1 : m-1] + [\text{sample}(p_L[m])]
            else
                y←y+[sample(pL[−1])]y \leftarrow y + [\text{sample}(p_L[-1])]
            end if
        end if
    end while
    return yy
  5. Knowl 5 — Model Prediction Alignment for BiLD

    model/method

    When a small model and a large model are fine-tuned independently on target training sets, differences in phrasing or vocabulary choice (e.g., "writing is hard" vs. "writing is difficult") can cause spurious rollbacks that reduce inference speed without improving semantic output quality.

    To align the small model's predictions with the large model without modifying architectures or distillation objectives:

    1. A calibration input dataset Xcal={x(i)}\mathcal{X}_{cal} = \{x^{(i)}\} is constructed from the training inputs.
    2. The fully fine-tuned large model generates pseudo-target sequences via greedy decoding: y(i)=arg⁡max⁡ypL(y∣x(i))y^{(i)} = \arg\max_y p_L(y \mid x^{(i)}) forming a dataset Dcal={(x(i),y(i))}\mathcal{D}_{cal} = \{(x^{(i)}, y^{(i)})\}.
    3. The pre-trained small model is fine-tuned on Dcal\mathcal{D}_{cal} using standard supervised cross-entropy training.

    This procedure aligns the output distributions pSp_S and pLp_L, minimizing cross-entropy distance d(pS,pL)d(p_S, p_L) during inference and reducing unnecessary rollbacks.

  6. Knowl 6 — Inference Speedup and Generation Quality Across Translation and Summarization

    data/table

    The BiLD framework was evaluated on an NVIDIA T4 GPU (GCP n1-standard-4 instance, batch size 1) across machine translation benchmarks (IWSLT 2017 De-En and WMT 2014 De-En using mT5-small and mT5-large) and summarization benchmarks (XSUM and CNN/DailyMail using T5-small and T5-large). The table displays BLEU scores for translation, ROUGE-L scores for summarization, and latency speedups normalized against vanilla large model autoregressive inference.

    Task (Model) Machine Translation (mT5) Summarization (T5)
    Dataset IWSLT 2017 WMT 2014 XSUM CNN/DailyMail
    BLEU Speedup BLEU Speedup ROUGE-L Speedup ROUGE-L Speedup
    Vanilla Large Baseline 40.32 1.00×\times 31.38 1.00×\times 35.08 1.00×\times 41.54 1.00×\times
    BiLD (Unaligned, min drop) 40.33 1.43×\times 31.28 1.34×\times 35.12 1.48×\times 41.44 1.71×\times
    BiLD (Unaligned, ∼\sim1pt drop) 39.44 1.58×\times 30.47 1.43×\times 34.02 1.72×\times 40.57 2.05×\times
    BiLD (Aligned, min drop) 40.24 1.62×\times 31.26 1.47×\times 35.05 1.50×\times 41.52 1.85×\times
    BiLD (Aligned, ∼\sim1pt drop) 39.13 1.78×\times 30.33 1.70×\times 33.95 1.80×\times 40.96 2.12×\times

    Unaligned BiLD is a plug-and-play framework achieving 1.34×1.34\times to 1.71×1.71\times speedups without metric degradation and up to 2.05×2.05\times speedup with ≤1.0\le 1.0 point degradation. Aligned BiLD improves speedup to 1.47×1.47\times–1.85×1.85\times at score parity and achieves up to 2.12×2.12\times speedup within a 1-point score degradation.

  7. Knowl 7 — Hardware Profile: FLOPs, Memory Operations, and Arithmetic Intensity

    empirical result

    Autoregressive generation in large Transformer models is memory bandwidth constrained during inference because weights must be transferred from GPU memory for each generated token. Profiling vanilla large model inference against BiLD on the CNN/DailyMail benchmark (using T5 models at equivalent ROUGE-L performance) demonstrates how BiLD alters hardware utilization:

    • FLOPs: BiLD requires 1.11×1.11\times (+11%+11\%) the floating-point operations of vanilla inference due to running the auxiliary small model.
    • Memory Operations (MOPs): BiLD achieves a ∼4.5×\sim 4.5\times reduction in memory access operations, requiring only 0.22×0.22\times the MOPs of vanilla inference. This occurs because the large model processes multiple tokens in parallel per invocation, loading weights once for the whole sequence segment.
    • Arithmetic Intensity: Arithmetic intensity (FLOPs per memory operation) increases by 4.96×4.96\times relative to vanilla inference.
    • Hardware Speedup: By converting memory bandwidth-bound single-token iterations into compute-bound multi-token parallel verification, BiLD achieves a 1.85×1.85\times end-to-end wall-clock latency speedup on an NVIDIA T4 GPU.
  8. Knowl 8 — BiLD Early Exiting Strategy vs CALM

    empirical result

    BiLD can be implemented within a single Transformer by using an early layer as the small model and the full network as the large model. On an 8-layer mT5-small model where layer 1 serves as the small model and all 8 layers serve as the large model, the network is trained with the joint objective: L=12(L1+L−1)\mathcal{L} = \frac{1}{2}(\mathcal{L}_1 + \mathcal{L}_{-1}) where L1\mathcal{L}_1 and L−1\mathcal{L}_{-1} are the negative log-likelihood losses at layer 1 and layer 8, sharing a prediction head.

    When evaluated on machine translation benchmarks:

    • BiLD achieves up to 1.60×1.60\times speedup on IWSLT 2017 De-En and 1.74×1.74\times speedup on WMT 2014 De-En within 1 BLEU point of vanilla inference.
    • Compared to Confident Adaptive Language Modeling (CALM), BiLD achieves a 2.02.0 to 2.52.5 BLEU point gain at equivalent latency speedups (∼1.5× \sim 1.5\times).

    This improvement over CALM is attributed to: (1) BiLD's rollback policy, which reverts and corrects incorrect early-exit tokens rather than allowing errors to propagate; and (2) populating key-value caches for skipped layers with true representations during verification rather than propagating approximate hidden states.

  9. Knowl 9 — Comparison of BiLD with Rejection Sampling-Based Speculative Decoding

    data/table

    BiLD was benchmarked against rejection sampling-based speculative decoding (Leviathan et al., 2023; Chen et al., 2023) using baseline (unaligned) models on IWSLT 2017 De-En (mT5-small/large) and XSUM (T5-small/large) on an NVIDIA T4 GPU (batch size 1).

    Task / Configuration BLEU / ROUGE-L Speedup % Fallback % Rollback (Rejection)
    IWSLT 2017 De-En
    Vanilla Large Baseline 40.32 1.00×\times - -
    Rejection Sampling Speculative Decoding 39.93 1.28×\times 23.24% 9.81%
    BiLD (Match Latency) 40.54 1.23×\times - -
    BiLD (Match Quality) 39.87 1.49×\times - -
    BiLD (Better BLEU) 40.33 1.43×\times 21.09% 1.56%
    XSUM
    Vanilla Large Baseline 35.08 1.00×\times - -
    Rejection Sampling Speculative Decoding 35.00 1.25×\times 36.84% 24.24%
    BiLD (Match Latency) 35.30 1.42×\times - -
    BiLD (Match Quality) 34.96 1.50×\times - -
    BiLD (Better ROUGE-L) 35.12 1.48×\times 32.33% 6.41%

    BiLD yields higher generation quality and lower latency than rejection sampling speculative decoding. BiLD's dynamic fallback window reduces invocation frequency (21.09%21.09\% vs 23.24%23.24\% on IWSLT; 32.33%32.33\% vs 36.84%36.84\% on XSUM), and its deterministic cross-entropy rollback eliminates stochastic rejection, drastically reducing rollback frequency (1.56%1.56\% vs 9.81%9.81\% on IWSLT; 6.41%6.41\% vs 24.24%24.24\% on XSUM).

  10. Knowl 10 — Ablation of BiLD Fallback and Rollback Policies

    empirical result

    Ablation experiments on IWSLT 2017 De-En translation and XSUM summarization evaluate the isolated contributions of the fallback and rollback policies in aligned BiLD:

    • Removing the Rollback Policy (BiLD No RB): When the small model uses confidence fallback but the large model cannot roll back previous tokens, generation quality drops substantially across all latency operating points, especially in the high-quality regime. This demonstrates that the quality gains from correcting early-stage mispredictions exceed the duplicated latency cost of reverting tokens.
    • Removing the Fallback Policy (BiLD No FB): When the large model inspects and rolls back tokens but the small model invokes the large model after a fixed token stride (swept from 3 to 10 tokens) rather than based on confidence, generation quality and speedup both degrade sharply. Static scheduling causes premature handoffs when the small model is confident and delayed handoffs when it is uncertain.
  11. Knowl 11 — BiLD Acceleration under Nucleus Sampling

    data/table

    The BiLD framework extends directly to stochastic generation methods by replacing greedy token selection with random sampling while retaining the maximum-probability fallback policy (max⁡pS<αFB\max p_S < \alpha_{FB}) and cross-entropy rollback policy. The table presents results for nucleus sampling (p=0.8p = 0.8) on IWSLT 2017 De-En (mT5-small/large) and XSUM (T5-small/large) using aligned models on an NVIDIA T4 GPU.

    Benchmark IWSLT 2017 De-En XSUM
    Configuration BLEU Speedup ROUGE-L Speedup
    Vanilla Large Inference 39.24 1.00×\times 34.00 1.00×\times
    BiLD (High Quality) 39.72 (+0.48) 1.51×\times 34.34 (+0.34) 1.22×\times
    BiLD (Parity) 39.26 (+0.02) 1.63×\times 34.04 (+0.04) 1.45×\times
    BiLD (∼\sim1pt drop) 38.27 (-0.97) 1.80×\times 33.10 (-0.90) 1.85×\times

    Under nucleus sampling (p=0.8p=0.8), BiLD provides speedups of 1.51×1.51\times–1.63×1.63\times on IWSLT and 1.22×1.22\times–1.45×1.45\times on XSUM without quality loss, reaching up to 1.80×1.80\times and 1.85×1.85\times speedup within a 1-point score degradation.

Coverage note — None was omitted; all primary contributions—including the dual-model framework, fallback and rollback policies, prediction alignment, early-exiting extensions, sampling adaptations, hardware roofline analyses, and comparative evaluations—are fully covered.

References

  1. 1.Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Ale s Tamchyna. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58, Baltimore, Maryland, USA, June 2014. Association for Computational Linguistics.
  2. 2.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  3. 3.Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian Stüker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. Overview of the IWSLT 2017 evaluation campaign. In Proceedings of the 14th International Conference on Spoken Language Translation, pages 2–14, Tokyo, Japan, December 14-15 2017. International Workshop on Spoken Language Translation.
  4. 4.Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023.
  5. 5.Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. Adabert: Task-adaptive bert compression with differentiable neural architecture search. arXiv preprint arXiv:2001.04246, 2020.
  6. 6.Patrick H Chen and Cho-jui Hsieh. A comparison of second-order methods for deep convolutional neural networks. openreview under ICLR 2018, 2018.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  8. 8.Michiel de Jong, Yury Zemlyanskiy, Joshua Ainslie, Nicholas FitzGerald, Sumit Sanghai, Fei Sha, and William Cohen. Fido: Fusion-in-decoder optimized for stronger performance and faster inference. arXiv preprint arXiv:2212.08153, 2022.
  9. 9.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems, 2022.
  10. 10.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pages 5547–5569. PMLR, 2022.
  11. 11.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. Depth-adaptive transformer. arXiv preprint arXiv:1910.10073, 2019.
  12. 12.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019.
  13. 13.Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pages 1050–1059. PMLR, 2016.
  14. 14.Trevor Gale, Erich Elsen, and Sara Hooker. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574, 2019.
  15. 15.Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. Mask-predict: Parallel decoding of conditional masked language models. arXiv preprint arXiv:1904.09324, 2019.
  16. 16.Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. Non-autoregressive neural machine translation. arXiv preprint arXiv:1711.02281, 2017.
  17. 17.Jiatao Gu, Changhan Wang, and Junbo Zhao. Levenshtein transformer. Advances in Neural Information Processing Systems, 32, 2019.
  18. 18.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR, 2017.
  19. 19.Junliang Guo, Linli Xu, and Enhong Chen. Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 376–385, 2020.
  20. 20.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. arXiv preprint arXiv:1610.02136, 2016.
  21. 21.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. In Advances in neural information processing systems, pages 1693–1701, 2015.
  22. 22.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. Workshop paper in NIPS, 2014.
  23. 23.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  24. 24.Chris Hokamp, Demian Gholipour Ghalandari, Nghia The Pham, and John Glover. Dyne: Dynamic ensemble decoding for multi-document summarization. arXiv preprint arXiv:2006.08748, 2020.
  25. 25.Ngo Quang Huy, Tu Minh Phuong, and Ngo Xuan Bach. Autoencoding language model based ensemble learning for commonsense validation and explanation. arXiv preprint arXiv:2204.03324, 2022.
  26. 26.Forrest N Iandola, Albert E Shaw, Ravi Krishna, and Kurt W Keutzer. Squeezebert: What can computer vision teach nlp about efficient neural networks? arXiv preprint arXiv:2006.11316, 2020.
  27. 27.Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. Llm-blender: Ensembling large language models with pairwise ranking and generative fusion. arXiv preprint arXiv:2306.02561, 2023.
  28. 28.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351, 2019.
  29. 29.Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A Smith. Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation. arXiv preprint arXiv:2006.10369, 2020.
  30. 30.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems, 30, 2017.
  31. 31.Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. I-bert: Integer-only bert quantization. arXiv preprint arXiv:2101.01321, 2021.
  32. 32.Sehoon Kim, Coleman Hooper, Thanakul Wattanawong, Minwoo Kang, Ruohan Yan, Hasan Genc, Grace Dinh, Qijing Huang, Kurt Keutzer, Michael W Mahoney, et al. Full stack optimization of transformer inference: a survey. arXiv preprint arXiv:2302.14017, 2023.
  33. 33.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. In International Conference on Learning Representations, 2019.
  34. 34.Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259, 2022.
  35. 35.Woosuk Kwon, Sehoon Kim, Michael W Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers. arXiv preprint arXiv:2204.09656, 2022.
  36. 36.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.
  37. 37.Jason Lee, Elman Mansimov, and Kyunghyun Cho. Deterministic non-autoregressive neural sequence modeling by iterative refinement. arXiv preprint arXiv:1802.06901, 2018.
  38. 38.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR, 2023.
  39. 39.Zhuohan Li, Zi Lin, Di He, Fei Tian, Tao Qin, Liwei Wang, and Tie-Yan Liu. Hint-based training for non-autoregressive machine translation. arXiv preprint arXiv:1909.06708, 2019.
  40. 40.Yoshitomo Matsubara, Luca Soldaini, Eric Lind, and Alessandro Moschitti. Ensemble transformer for efficient and accurate ranking tasks: an application to question answering systems. arXiv preprint arXiv:2201.05767, 2022.
  41. 41.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016.
  42. 42.Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? arXiv preprint arXiv:1905.10650, 2019.
  43. 43.Shashi Narayan, Shay B Cohen, and Mirella Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. arXiv preprint arXiv:1808.08745, 2018.
  44. 44.Liu Pai. Qiaoning at semeval-2020 task 4: Commonsense validation and explanation system based on ensemble of language model. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, pages 415–421, 2020.
  45. 45.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017.
  46. 46.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently scaling transformer inference. arXiv preprint arXiv:2211.05102, 2022.
  47. 47.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  48. 48.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  49. 49.Victor Sanh, Thomas Wolf, and Alexander Rush. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389, 2020.
  50. 50.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  51. 51.Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q Tran, Yi Tay, and Donald Metzler. Confident adaptive language modeling. arXiv preprint arXiv:2207.07061, 2022.
  52. 52.Chenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng, and Jie Zhou. Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 198–205, 2020.
  53. 53.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604, 2018.
  54. 54.Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Q-BERT: Hessian based ultra low precision quantization of bert. In AAAI, pages 8815–8821, 2020.
  55. 55.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990, 2022.
  56. 56.David So, Quoc Le, and Chen Liang. The evolved transformer. In International Conference on Machine Learning, pages 5877–5886. PMLR, 2019.
  57. 57.David R So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le. Primer: Searching for efficient transformers for language modeling. arXiv preprint arXiv:2109.08668, 2021.
  58. 58.Mitchell Stern, William Chan, Jamie Kiros, and Jakob Uszkoreit. Insertion transformer: Flexible sequence generation via insertion operations. In International Conference on Machine Learning, pages 5976–5985. PMLR, 2019.
  59. 59.Zhiqing Sun, Zhuohan Li, Haoqing Wang, Di He, Zi Lin, and Zhihong Deng. Fast structured decoding for sequence models. Advances in Neural Information Processing Systems, 32, 2019.
  60. 60.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. Mobilebert: a compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984, 2020.
  61. 61.Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. Distilling task-specific knowledge from bert into simple neural networks. arXiv preprint arXiv:1903.12136, 2019.
  62. 62.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  64. 64.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418, 2019.
  65. 65.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  66. 66.Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, and Song Han. Hat: Hardware-aware transformers for efficient natural language processing. arXiv preprint arXiv:2005.14187, 2020.
  67. 67.Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  68. 68.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv preprint arXiv:2002.10957, 2020.
  69. 69.Yiren Wang, Fei Tian, Di He, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu. Non-autoregressive machine translation with auxiliary regularization. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 5377–5384, 2019.
  70. 70.Bingzhen Wei, Mingxuan Wang, Hao Zhou, Junyang Lin, Jun Xie, and Xu Sun. Imitation learning for non-autoregressive neural machine translation. arXiv preprint arXiv:1906.02041, 2019.
  71. 71.Sean Welleck, Kianté Brantley, Hal Daumé Iii, and Kyunghyun Cho. Non-monotonic sequential text generation. In International Conference on Machine Learning, pages 6716–6726. PMLR, 2019.
  72. 72.Samuel Williams, Andrew Waterman, and David Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009.
  73. 73.Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, 2020.
  74. 74.Xiaoxia Wu, Zhewei Yao, Minjia Zhang, Conglong Li, and Yuxiong He. Extreme compression for pre-trained transformers made simple and efficient. arXiv preprint arXiv:2206.01859, 2022.
  75. 75.Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. Lite transformer with long-short range attention. arXiv preprint arXiv:2004.11886, 2020.
  76. 76.Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu. Nas-bert: task-agnostic and adaptive-size bert compression with neural architecture search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 1933–1943, 2021.
  77. 77.Yige Xu, Xipeng Qiu, Ligao Zhou, and Xuanjing Huang. Improving bert fine-tuning via self-ensemble and self-distillation. arXiv preprint arXiv:2002.10345, 2020.
  78. 78.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  79. 79.Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. arXiv preprint arXiv:2206.01861, 2022.
  80. 80.Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. Autotinybert: Automatic hyper-parameter optimization for efficient pre-trained language models. arXiv preprint arXiv:2107.13686, 2021.
  81. 81.Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 811–824. IEEE, 2020.
  82. 82.Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. Q8BERT: Quantized 8bit bert. arXiv preprint arXiv:1910.06188, 2019.
  83. 83.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  84. 84.Chunting Zhou, Graham Neubig, and Jiatao Gu. Understanding knowledge distillation in non-autoregressive machine translation. arXiv preprint arXiv:1911.02727, 2019.

Citation

MLA
Kim, S., et al. “Speculative Decoding with Big Little Decoder”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 39236–56, https://proceedings.neurips.cc/paper_files/paper/2023/file/7b97adeafa1c51cf65263459ca9d0d7c-Paper-Conference.pdf.
APA
Kim, S., Mangalam, K., Moon, S., Malik, J., Mahoney, M., Gholami, A., & Keutzer, K. (2023). Speculative Decoding with Big Little Decoder. Advances in Neural Information Processing Systems, 36, 39236–39256. https://proceedings.neurips.cc/paper_files/paper/2023/file/7b97adeafa1c51cf65263459ca9d0d7c-Paper-Conference.pdf
Chicago
Kim, S., K. Mangalam, S. Moon, et al. 2023. “Speculative Decoding with Big Little Decoder”. Advances in Neural Information Processing Systems 36: 39236–56. https://proceedings.neurips.cc/paper_files/paper/2023/file/7b97adeafa1c51cf65263459ca9d0d7c-Paper-Conference.pdf.
Harvard
Kim, S. et al. (2023) “Speculative Decoding with Big Little Decoder”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 39236–39256. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/7b97adeafa1c51cf65263459ca9d0d7c-Paper-Conference.pdf.
Vancouver
1. Kim S, Mangalam K, Moon S, Malik J, Mahoney M, Gholami A, Keutzer K (2023) Speculative Decoding with Big Little Decoder. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 39236–39256

BibTeX

@inproceedings{kim2023speculative,
  title = {Speculative Decoding with Big Little Decoder},
  author = {Kim, Sehoon and Mangalam, Karttikeya and Moon, Suhong and Malik, Jitendra and Mahoney, Michael and Gholami, Amir and Keutzer, Kurt},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {39236-39256},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/7b97adeafa1c51cf65263459ca9d0d7c-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors