Speculative Decoding with Big Little Decoder
Sehoon KimKarttikeya MangalamSuhong MoonJitendra MalikMichael W. MahoneyAmir GholamiKurt Keutzer
Proposes Big Little Decoder, a plug-and-play speculative decoding framework that pairs a small autoregressive model with an occasionally invoked large model via fallback and rollback policies to speed up text generation by up to 2.12× without requiring retraining or architectural changes.
Large language models based on transformer architectures have enabled significant breakthroughs in natural language processing. However, generating text with these models remains slow and computationally expensive because they typically produce text sequentially, one word or token at a time. This autoregressive process forces hardware to constantly reload model weights for every single token, causing memory bottlenecks, poor hardware utilization, and high latency that hinder real-time deployment.
The article introduces and evaluates the Big Little Decoder (BiLD), a plug-and-play framework designed to accelerate text generation without altering underlying model architectures or retraining pipelines. The objective is to demonstrate that pairing a fast, compact model with an accurate, larger model can substantially reduce inference latency while preserving generation quality across multiple natural language tasks.
To evaluate this framework, the authors conducted extensive experiments on standard machine translation benchmarks (IWSLT 2017 and WMT 2014 German-to-English) and text summarization benchmarks (XSUM and CNN/DailyMail). The setup used standard transformer models where the large model was roughly twenty times the size of the smaller companion. Inference performance was measured in a standard cloud environment on an NVIDIA T4 GPU. The framework operates by letting the small model generate text quickly and sequentially, while invoking the large model only when needed. Coordination relies on two core rules: a fallback policy that hands control to the large model when the small model is uncertain, and a rollback policy that uses the large model to verify and correct past tokens in parallel. An optional calibration step, termed model prediction alignment, was also evaluated to train the small model to mimic the vocabulary preferences of the large model.
The findings show that small models share substantial agreement with large models and only require occasional corrections to match their output quality. Across the evaluated tasks, the standard plug-and-play BiLD setup achieved an average speedup of 1.50× with no degradation in text quality, reaching up to 1.71× on summarization. When combined with the prediction alignment technique, the system achieved up to 1.85× speedup with zero quality loss and reached up to 2.12× speedup when accepting a minor quality degradation of approximately one point on standard evaluation metrics. System analysis revealed that while total floating-point arithmetic slightly increased, BiLD reduced memory operations roughly fivefold, boosting arithmetic intensity and eliminating hardware memory bottlenecks. Furthermore, the approach consistently outperformed alternative speculative decoding and early-exiting techniques across both translation and summarization tasks.
These results indicate that organizations deploying large language models for real-time applications can achieve near-double throughput and lower latency without re-architecting systems or conducting costly full-model retraining. By addressing the hardware memory bottleneck rather than merely minimizing arithmetic operations, collaborative model execution provides a viable, cost-effective path to scaling interactive artificial intelligence services.
Engineering and deployment teams should consider implementing collaborative decoding policies like BiLD in latency-sensitive, batch-size-one inference pipelines. Organizations seeking maximum throughput should also adopt the lightweight prediction alignment fine-tuning step. Before full production rollout, teams should conduct task-specific threshold sweeps to balance quality against latency and validate performance across larger model families and diverse hardware configurations.
The evaluation carries high confidence within single-batch online serving scenarios on common hardware platforms. However, the study primarily focused on translation and summarization tasks up to moderate sequence lengths, and results may vary under large batch sizes or distinct domain workloads where memory access patterns differ.
- Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). This seminal work establishes the foundational paradigm of speculative decoding, in which a small draft model autoregressively proposes tokens that a large model verifies in parallel.
- Paper: Confident Adaptive Language Modeling, Tal Schuster et al. (2022). This paper introduces confidence-based dynamic computation allocation and early exits during language model generation, directly motivating confidence-driven fallback and coordination policies between draft and target models.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). It provides fundamental analysis and architectural solutions for overcoming the memory-bandwidth bottleneck in autoregressive incremental decoding.
- Paper: Orca: A Distributed Serving System for Transformer-Based Generative Models, Gyeong-In Yu et al. (2022). It formalizes iteration-level scheduling and the systems-level latency challenges of autoregressive text generation that collaborative multi-model decoding addresses.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). It introduces the core Transformer sequence-to-sequence architecture underpinning the models and text generation tasks studied in the paper.
- Paper: DistillSpec: Improving Speculative Decoding via Knowledge Distillation, Yongchao Zhou et al. (2024). DistillSpec applies knowledge distillation specifically to align small draft models with large target models, improving acceptance rates and draft quality in speculative decoding setups.
- Paper: EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees, Yuhui Li et al. (2024). EAGLE-2 advances speculative drafting by dynamically structuring draft candidate trees based on context confidence rather than fixed drafting paths.
- Paper: SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification, Xupeng Miao et al. (2024). SpecInfer extends speculative inference into a tree-based draft verification framework across distributed serving environments.
- Paper: Unlocking Lossless Speedups in LLMs via Discrete Diffusion, Subham Sekhar Sahoo et al. (2026). This work explores replacing standard autoregressive draft models with discrete diffusion-based parallel drafting within the speculative verification framework.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This comprehensive survey categorizes and contextualizes speculative decoding alongside other modern efficient inference architectures for large language models.
