Dynamic Rebatching for Efficient Early-Exit Inference with DREX
Xuting LiuDaniel AlexanderSiva Kesava Reddy KakarlaBehnaz ArzaniVincent Liu
Presents DREX, a dynamic rebatching system for early-exit large language models that eliminates output-degrading forced exits and increases inference throughput using copy-free buffer management and predictive scheduling.
Serving large language models demands massive computational resources, leading to significant interest in techniques that reduce inference costs. Early exiting allows easier generated tokens to finish computation at intermediate layers rather than passing through every model layer. However, standard batching methods struggle with early exits because requests within the same batch frequently disagree on whether they are ready to exit. Existing systems either force uniform batch-wide decisions—causing premature exits that degrade output quality or missed opportunities that limit throughput—or incur high memory and computation overheads to manage missing attention states.
The article demonstrates and evaluates DREX, an inference serving framework designed to make early-exit language models practical in batched production environments. The primary objective is to evaluate how dynamic batch reorganization, overhead-aware scheduling, and virtual memory techniques can improve serving throughput while preserving model accuracy and service level agreements.
The authors implemented DREX on top of an existing serving platform and evaluated it against standard non-early-exit models and state-of-the-art batched early-exit baselines (such as consensus, majority, and greedy grouped-exit policies). Testing utilized multiple open-source models, including 13-billion, 14-billion, and 70-billion parameter architectures, running summarization workloads from standard benchmarks on enterprise graphics hardware.
The evaluation revealed several key findings. First, DREX improved serving throughput by 2% to 12% across evaluated models compared to baseline approaches while maintaining high token confidence scores. Second, DREX completely eliminated involuntary exits, preventing the severe quality degradation observed in aggressive baselines (where premature exits degraded output confidence by up to 96%). Third, the framework's adaptive threshold mechanism, which only triggers dynamic rebatching when predicted compute savings exceed rebatching overhead, boosted throughput by an additional 9% on smaller models. Fourth, using virtual memory mappings to populate missing attention cache entries reduced graphics memory operation sizes by up to 18.3% and average memory operations by 5.7% compared to traditional duplication.
These findings indicate that early-exit models can be integrated into high-throughput production serving pipelines without compromising model quality. For organizations operating language model infrastructure, adopting dynamic rebatching reduces GPU compute and memory pressure, driving down hosting costs while maintaining response reliability. The results demonstrate that handling split exit decisions at the serving layer overcomes the operational bottlenecks that previously made early-exit models ineffective in batched settings.
Organizations serving large language models should consider adopting dynamic rebatching and virtual memory caching techniques when deploying early-exit architectures. System administrators should configure the adaptive threshold and service deadline parameters to balance throughput gains against request completion latency based on application requirements. Before broad enterprise deployment, teams should conduct pilot testing on their specific workloads and fine-tune model exit ramps, as the article noted that semantic quality scores can vary if intermediate layer confidence is not perfectly calibrated with final task accuracy.
The primary limitations of the study include its focus on a summarization benchmark with a fixed context limit and the assumption that early-exit classifier ramps are pre-trained and accurate. Confidence in the reported performance and memory improvements is high given the rigorous multi-model hardware evaluation, though practitioners should exercise caution and validate output quality on tasks requiring long-context reasoning.
- Paper: Confident Adaptive Language Modeling, Tal Schuster et al. (2022). Introduces foundational early-exit decoding and confidence-based layer skipping in language models, establishing the architectural paradigm and missing KV cache challenges that DREX dynamically rebatches.
- Paper: LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding, Mostafa Elhoushi et al. (2024). Provides a crucial early-exit and self-speculative LLM acceleration architecture, detailing layer-skipping mechanics directly targeted by DREX's dynamic serving framework.
- Paper: BranchyNet: Fast inference via early exiting from deep neural networks, Surat Teerapittayanon et al. (2016). Pioneers the fundamental concept of early exiting via multi-exit neural architectures to accelerate deep network inference on easy examples.
- Paper: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills, Amey Agrawal et al. (2023). Examines dynamic batch composition and scheduling tradeoffs during LLM serving, offering essential context for DREX's SLA-aware dynamic rebatching.
- Paper: Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference, Mostafa Elhoushi et al. (2026). Extends layer sparsity and dynamic depth execution during LLM pretraining and inference, complementing DREX's runtime batching optimizations for variable-depth architectures.
- Paper: Performance Aware LLM Load Balancer for Mixed Workloads, Kunal Jain et al. (2025). Applies system-level SLA and workload-aware scheduling to heterogeneous LLM inference across instances, extending runtime scheduling principles beyond single-instance layer rebatching.
