Dynamic Rebatching for Efficient Early-Exit Inference with DREX

Xuting LiuDaniel AlexanderSiva Kesava Reddy KakarlaBehnaz ArzaniVincent Liu

article2025arXiv2 citations

Presents DREX, a dynamic rebatching system for early-exit large language models that eliminates output-degrading forced exits and increases inference throughput using copy-free buffer management and predictive scheduling.

Listen

Serving large language models demands massive computational resources, leading to significant interest in techniques that reduce inference costs. Early exiting allows easier generated tokens to finish computation at intermediate layers rather than passing through every model layer. However, standard batching methods struggle with early exits because requests within the same batch frequently disagree on whether they are ready to exit. Existing systems either force uniform batch-wide decisions—causing premature exits that degrade output quality or missed opportunities that limit throughput—or incur high memory and computation overheads to manage missing attention states.

The article demonstrates and evaluates DREX, an inference serving framework designed to make early-exit language models practical in batched production environments. The primary objective is to evaluate how dynamic batch reorganization, overhead-aware scheduling, and virtual memory techniques can improve serving throughput while preserving model accuracy and service level agreements.

The authors implemented DREX on top of an existing serving platform and evaluated it against standard non-early-exit models and state-of-the-art batched early-exit baselines (such as consensus, majority, and greedy grouped-exit policies). Testing utilized multiple open-source models, including 13-billion, 14-billion, and 70-billion parameter architectures, running summarization workloads from standard benchmarks on enterprise graphics hardware.

The evaluation revealed several key findings. First, DREX improved serving throughput by 2% to 12% across evaluated models compared to baseline approaches while maintaining high token confidence scores. Second, DREX completely eliminated involuntary exits, preventing the severe quality degradation observed in aggressive baselines (where premature exits degraded output confidence by up to 96%). Third, the framework's adaptive threshold mechanism, which only triggers dynamic rebatching when predicted compute savings exceed rebatching overhead, boosted throughput by an additional 9% on smaller models. Fourth, using virtual memory mappings to populate missing attention cache entries reduced graphics memory operation sizes by up to 18.3% and average memory operations by 5.7% compared to traditional duplication.

These findings indicate that early-exit models can be integrated into high-throughput production serving pipelines without compromising model quality. For organizations operating language model infrastructure, adopting dynamic rebatching reduces GPU compute and memory pressure, driving down hosting costs while maintaining response reliability. The results demonstrate that handling split exit decisions at the serving layer overcomes the operational bottlenecks that previously made early-exit models ineffective in batched settings.

Organizations serving large language models should consider adopting dynamic rebatching and virtual memory caching techniques when deploying early-exit architectures. System administrators should configure the adaptive threshold and service deadline parameters to balance throughput gains against request completion latency based on application requirements. Before broad enterprise deployment, teams should conduct pilot testing on their specific workloads and fine-tune model exit ramps, as the article noted that semantic quality scores can vary if intermediate layer confidence is not perfectly calibrated with final task accuracy.

The primary limitations of the study include its focus on a summarization benchmark with a fixed context limit and the assumption that early-exit classifier ramps are pre-trained and accurate. Confidence in the reported performance and memory improvements is high given the rigorous multi-model hardware evaluation, though practitioners should exercise caution and validate output quality on tasks requiring long-context reasoning.

arXiv: 2512.15705
Cover for Dynamic Rebatching for Efficient Early-Exit Inference with DREX

Abstract

Early-Exit (EE) is a Large Language Model (LLM) architecture that accelerates inference by allowing easier tokens to be generated using only a subset of the model's layers. However, traditional batching frameworks are ill-suited for EE LLMs, as not all requests in a batch may be ready to exit at the same time. Existing solutions either force a uniform decision on the batch, which overlooks EE opportunities, or degrade output quality by forcing premature exits. We propose Dynamic Rebatching, a solution where we dynamically reorganize the batch at each early-exit point. Requests that meet the exit criteria are immediately processed, while those that continue are held in a buffer, re-grouped into a new batch, and forwarded to deeper layers. We introduce DREX, an early-exit inference system that implements Dynamic Rebatching with two key optimizations: 1) a copy-free rebatching buffer that avoids physical data movement, and 2) an EE and SLA-aware scheduler that analytically predicts whether a given rebatching operation will be profitable. DREX also efficiently handles the missing KV cache from skipped layers using memory-efficient state-copying. Our evaluation shows that DREX improves throughput by 2-12% compared to baseline approaches while maintaining output quality. Crucially, DREX completely eliminates involuntary exits, providing a key guarantee for preserving the output quality intended by the EE model.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background on LLM Serving
  • 3 Early Exiting (EE) LLMs
  • 3.1 Quantifying the Opportunity
  • 3.2 Challenges in Operationalizing EE
  • 3.2.1 Handling Split EE Decisions
  • 3.2.2 Handling the Missing KV Cache
  • 4 DREX Overview
  • 5 DREX and Dynamic Rebatching
  • 5.1 To EE or Not to EE
  • 5.2 Buffering Left-Behind Requests
  • 5.3 Flushing the Buffer
  • 5.4 Filling In the Missing State
  • 6 Implementation
  • 7 Evaluation
  • 7.1 Better Throughput With Confidence
  • 7.2 Impact of Adaptive Rebatching Threshold
  • 7.3 Request Completion Time (RCT)
  • 7.4 Memory Operation
  • 8 Related Work
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — Dynamic Rebatching in Early-Exit LLM Inference

    model/method

    Dynamic Rebatching is a serving mechanism for early-exit (EE) Large Language Models (LLMs) that avoids all-or-nothing grouped batch exit decisions when individual requests in a batch produce split early-exit decisions. In standard early-exit LLMs, each token evaluates an intermediate exit ramp classifier to decide whether to terminate early or proceed through remaining deeper layers. In Dynamic Rebatching:

    1. When a split decision occurs in a batch, requests meeting the early-exit confidence criterion immediately generate their output token and complete the current iteration.
    2. Requests that must continue through deeper layers are placed into a logical rebatching buffer without physical memory movement.
    3. When the buffer accumulates enough requests (or when SLA deadlines require flushing), the buffered requests are regrouped into a new batch and forwarded through the deeper model layers.

    This allows every request to follow its individual optimal execution path while preserving high GPU compute utilization during deep-layer processing and guaranteeing zero involuntary early exits.

  2. Knowl 2 — Adaptive Rebatching Threshold for Split Early-Exit Decisions

    equation

    To decide whether to split and rebatch a batch of size bb when b′b' requests satisfy the early-exit condition, the serving system compares the predicted computational savings against the rebatching overhead.

    Let tft_f be the execution time of a standard full iteration through all model layers, tst_s be the execution time of a shallow iteration (from the input through the early-exit ramp, including buffer insertion overhead), and tdt_d be the execution time of a deep iteration (from the ramp through the remaining layers, including buffer retrieval overhead). The overhead of dynamic rebatching, denoted cc, is: c=ts+td−tfc = t_s + t_d - t_f

    The computational savings per early-exiting request is: savings=tf−ts=td−c\text{savings} = t_f - t_s = t_d - c

    For early exiting and rebatching to yield a net computational benefit, the aggregate savings across b′b' exiting requests must exceed the aggregate overhead of the b−b′b - b' non-exiting requests: b′⋅(td−c)>(b−b′)⋅c  ⟺  b′>ctd⋅bb' \cdot (t_d - c) > (b - b') \cdot c \iff b' > \frac{c}{t_d} \cdot b

    The Adaptive Rebatching Threshold (ART) is defined as the minimum number of early-exiting requests required for rebatching to be profitable: ART=ctd⋅b\text{ART} = \frac{c}{t_d} \cdot b

    For models with multiple intermediate exit ramps, the threshold at ramp ii is: ART(i)=ctdi⋅b\text{ART}(i) = \frac{c}{t_d^i} \cdot b where tdit_d^i is the execution time of the deep layers starting from ramp ii to the final layer. The serving system profiles and updates cc and tdt_d periodically (e.g., every 100 steps) using moving averages of observed iteration latencies.

  3. Knowl 3 — SLA-Aware Buffer Flushing Condition

    equation

    To prevent starvation of requests held in the dynamic rebatching buffer and ensure Service Level Agreement (SLA) adherence, DREX evaluates an SLA-aware flushing condition that triggers execution of buffered requests through deeper layers even when the buffer has not reached the standard batch capacity.

    The buffer manager flushes the buffered requests when: bbuffer⋅(1+αmax⁡{rSLA−rexpected, ε})≥bschedulerb_{\text{buffer}} \cdot \left(1 + \frac{\alpha}{\max\{r_{\text{SLA}} - r_{\text{expected}},\, \varepsilon\}}\right) \ge b_{\text{scheduler}} where:

    • bbuffer∈Nb_{\text{buffer}} \in \mathbb{N} is the number of requests currently held in the rebatching buffer.
    • bscheduler∈Nb_{\text{scheduler}} \in \mathbb{N} is the batch size of the next available fresh batch ready in the main scheduler.
    • rexpected=age+L−lr_{\text{expected}} = \text{age} + L - l is the expected remaining iterations for the oldest request in the buffer, where age\text{age} is its elapsed iterations, LL is the maximum configured output sequence length, and ll is its current output length.
    • rSLAr_{\text{SLA}} is the target request completion time (RCT) requirement from the SLA expressed in number of iterations (calculated as RCT requirement divided by profiled standard iteration time).
    • α≥0\alpha \ge 0 is a user-tunable weight parameter controlling the urgency of SLA deadlines (setting α=0\alpha = 0 disables SLA-based forced flushing).
    • ε>0\varepsilon > 0 is a small constant to prevent division by zero or negative values.
  4. Knowl 4 — Memory-Efficient State-Copying via GPU Virtual Memory Mapping

    model/method

    When an early-exit LLM terminates token generation at an intermediate layer, future decoding iterations of that sequence that traverse deeper layers require key-value (KV) cache entries corresponding to the skipped layers. DREX fills missing KV cache states without redundant memory allocations or physical data copies by using GPU virtual memory manipulation (such as via vAttention's vMemMap API).

    When a token exits early at layer kk:

    1. The PyTorch tensor abstractions representing the KV cache entries for all subsequent skipped layers k+1,…,Nk+1, \dots, N are assigned virtual memory mappings pointing to the exact physical GPU memory block allocated for the KV cache at layer kk.
    2. In subsequent decoding passes that reach deeper layers, attention kernels read these mapped physical blocks on a shared, read-only basis.
    3. This eliminates redundant duplication of KV tensors across layers, avoiding both the computational overhead of KV-recomputation and the memory bloating of physical state-copying.
  5. Knowl 5 — Involuntary Exits and Involuntary Stays in Batched Early-Exit Serving

    definition

    In batched serving of early-exit (EE) Large Language Models where individual requests have heterogeneous confidence scores, grouped exit policies face a quality-throughput dilemma captured by two metrics:

    • Involuntary Exit Rate: The percentage of generated tokens that are forced to terminate early at an intermediate ramp despite their individual confidence scores falling below the required exit threshold. Involuntary exits degrade model output quality because the token receives insufficient compute.
    • Involuntary Stay Rate: The percentage of generated tokens that meet or exceed their individual confidence threshold to early exit, but are forced to continue through deeper model layers due to batch-level consensus or majority policies. Involuntary stays waste GPU compute capacity and reduce serving throughput.
  6. Knowl 6 — Copy-Free Logical Rebatching Buffer via Virtual Tensor Indexing

    model/method

    DREX implements the dynamic rebatching buffer as a purely logical tracking structure to eliminate physical data movement overheads during batch restructuring.

    Instead of physically copying or reallocating hidden states and KV caches when requests split across early-exit ramps:

    1. A centralized buffer manager tracks metadata indicating which requests are paused in the buffer awaiting deep-layer execution and which are actively exiting.
    2. When forming a new deep-layer batch from buffered requests across different prior batches, DREX passes a batch index mapping (such as the cache_batch_idx argument in FlashAttention) directly to the underlying GPU attention kernels.
    3. The attention kernel references the exact KV cache entries for each dynamically selected request in place.

    This decouples rebatching overhead from the model parameter size and sequence length, keeping the rebatching overhead below 6% of standard iteration time on 70B parameter models.

  7. Knowl 7 — End-to-End Throughput and Output Quality with Dynamic Rebatching

    empirical result

    Evaluated on text summarization from HELM using the CNN/Daily Mail dataset, DREX achieves 2% to 12% higher throughput compared to standard non-early-exit baselines (e.g., up to 12% throughput gain on Llama-EE-70B at batch size 8).

    Compared to grouped early-exit baselines:

    • DREX consistently outperforms safe grouped policies (Consensus, Majority, and Latency-only) by 2.0% to 10.3% in throughput while maintaining an equal or higher P95 confidence score.
    • The Greedy grouped exit policy achieves higher nominal throughput (exiting over 97% of tokens) but incurs over 35% involuntary exits, which collapses token quality (reducing P95 confidence score by 96%, down to 0.03).
    • DREX achieves zero involuntary exits (0%), fully preserving the output quality intended by the model's confidence classifier while capturing available compute savings.
  8. Knowl 8 — Impact of the Adaptive Rebatching Threshold on Serving Throughput

    data/table

    Experiments evaluating various fixed rebatching thresholds on batch size b=8b=8 demonstrate that setting an adaptive threshold is critical to prevent unprofitable split executions:

    Model Setting ART Throughput (tokens/s) EE % Involuntary Stay (%)
    Llama-EE-13B 0 116.80 46.3 0.0
    (layer=25, conf=0.8, b=8b=8) 1 117.14 35.9 2.7
    2 124.45 13.2 12.3
    3* 127.35 6.8 19.9
    4 121.09 3.3 26.2
    5 118.06 1.3 31.5
    Llama-EE-70B 0 132.97 46.5 0.0
    (layer=50, conf=0.7, b=8b=8) 1* 133.18 46.2 1.7
    2 129.29 29.0 10.6
    3 122.79 8.2 20.9
    4 121.38 3.5 26.8
    5 119.53 1.2 30.8

    The optimal threshold calculated analytically by DREX (marked with *) yields peak throughput:

    • For Llama-EE-13B, ART =3= 3 achieves 127.35 tokens/s, a 9.0% improvement over naive rebatching (ART=0\text{ART}=0 at 116.80 tokens/s).
    • For Llama-EE-70B, ART =1= 1 achieves the highest throughput of 133.18 tokens/s because deep-layer execution savings (td=33.30t_d = 33.30 ms) are significantly larger relative to rebatching overhead (c=7.92c = 7.92 ms) compared to 13B models (td=11.10t_d = 11.10 ms, c=5.35c = 5.35 ms).
  9. Knowl 9 — Latency-Throughput Adaptation via SLA-Aware Buffer Scheduling

    empirical result

    In DREX, the SLA-aware buffer flushing mechanism enables dynamic trade-offs between serving throughput and request completion time (RCT) depending on SLA pressure:

    • Under zero SLA pressure (pressure=0\text{pressure} = 0), DREX maximizes batch utilization, increasing throughput by 11.4% over a Consensus policy, while increasing average RCT by 1.4×1.4\times and P95 tail RCT by 3×3\times.
    • Under maximal SLA pressure (pressure=1\text{pressure} = 1), DREX avoids holding urgent requests in the buffer, flushing immediately or behaving like Consensus, which reduces tail latency and improves average responsiveness/RCT by up to 58.4% compared to unconstrained buffering.
  10. Knowl 10 — Reduction in CUDA Memory Operation Volume from Virtual State-Copying

    empirical result

    Profiling with NVIDIA Nsight Systems reveals that DREX's virtual memory-mapped state-copying consistently reduces physical CUDA memory operation sizes during early-exit serving compared to physical duplication of KV caches:

    • Across decode iterations, memory-efficient state copying reduces CUDA memory transfer size by an average of 5.7%.
    • Under workloads with high early-exit frequency (such as under a Greedy exit policy), physical memory operation size is reduced by up to 18.3%.

Coverage note — No substantial contributed material was omitted; minor software integration details (such as the ~1,500 LOC Python glue with Sarathi-Serve and standard dataset preprocessing splits) are subsumed by the primary methodological and empirical knowls.

References

  1. 1.Runpod | the cloud built for ai. [Online; accessed 2025-08-18].
  2. 2.Welcome to tensorrt-llm’s documentation! — tensorrt-llm. [Online; accessed 2025-08-07].
  3. 3.Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in llm inference with sarathi-serve. OSDI’24, USA, 2024. USENIX Association.
  4. 4.Sangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Aaron Courville, and Se-Young Yun. Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation, 2025.
  5. 5.Sangmin Bae, Jongwoo Ko, Hwanjun Song, and Se-Young Yun. Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5910–5924, Singapore, December 2023. Association for Computational Linguistics.
  6. 6.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. Qwen technical report, 2023.
  7. 7.Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. Medusa: Simple llm inference acceleration framework with multiple decoding heads. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  8. 8.Lequn Chen. Dissecting batching effects in gpt inference, May 2023.
  9. 9.Yanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding, and Jingren Zhou. Ee-llm: large-scale training and inference of early-exit large language models with 3d parallelism. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  10. 10.Luciano Del Corro, Allie Del Giorno, Sahaj Agarwal, Bin Yu, Ahmed Awadallah, and Subhabrata Mukherjee. Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference, 2023.
  11. 11.Yinwei Dai, Rui Pan, Anand Iyer, Kai Li, and Ravi Netravali. Apparate: Rethinking early exits to tame latency-throughput tensions in ml serving. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP ’24, page 607–623, New York, NY, USA, 2024. Association for Computing Machinery.
  12. 12.Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, 2024.
  13. 13.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: fast and memory-efficient exact attention with io-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc.
  14. 14.DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J. L. Cai, Jian Liang, Jianzhong Guo, Jiaqi Ni, Jiashi Li, Jiawei Wang, Jin Chen, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Xu, Leyi Xia, Liang Zhao, Litong Wang, Liyue Zhang, Meng Li, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Ning Tian, Panpan Huang, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qinyu Chen, Qiushi Du, R. J. Chen, R. L. Jin, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runxin Xu, Ruoyu Zhang, Ruyi Chen, S. S. Li, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaoqing Wu, Shengfeng Ye, Shengfeng Ye, Shirong Ma, Shiyu Wang, Shuang Zhou, Shuiping Yu, Shunfeng Zhou, Shuting Pan, T. Wang, Tao Yun, Tian Pei, Tianyu Sun, W. L. Xiao, Wangding Zeng, Wanjia Zhao, Wei An, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, X. Q. Li, Xiangyue Jin, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaojin Shen, Xiaokang Zhang, Xiaokang Zhang, Xiaosha Chen, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xinnan Song, Xinxia Shan, Xinyi Zhou, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Yang Zhang, Yanhong Xu, Yanhong Xu, Yanping Huang, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Li, Yaohui Wang, Yi Yu, Yi Zheng, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Ying Tang, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yu Wu, Yuan Ou, Yuchen Zhu, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yukun Zha, Yunfan Xiong, Yunxian Ma, Yuting Yan, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z. F. Wu, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhen Huang, Zhen Zhang, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhipeng Xu, Zhiyu Wu, Zhongyu Zhang, Zhuoshu Li, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Ziyi Gao, and Zizheng Pan. Deepseek-v3 technical report, 2025.
  15. 15.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. Depth-adaptive transformer. In International Conference on Learning Representations, 2020.
  16. 16.Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed Aly, Beidi Chen, and Carole-Jean Wu. LayerSkip: Enabling early exit inference and self-speculative decoding. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12622–12642, Bangkok, Thailand, August 2024. Association for Computational Linguistics.
  17. 17.Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang, and Zhongyuan Wang. Not all layers of llms are necessary during inference, 2024.
  18. 18.Karl Moritz Hermann, Tomaš Koācisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. In Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, page 1693–1701, Cambridge, MA, USA, 2015. MIT Press.
  19. 19.Benyamin Jamialahmadi, Parsa Kavehzadeh, Mehdi Rezagholizadeh, Parsa Farinneya, Hossein Rajabzadeh, Aref Jafari, Boxing Chen, and Marzieh S. Tahaei. Balcony: A lightweight approach to dynamic inference of generative language models, 2025.
  20. 20.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012.
  21. 21.Avinash Kumar, Shashank Nag, Jason Clemons, Lizy John, and Poulami Das. Helios: Adaptive model and early-exit selection for efficient llm inference serving, 2025.
  22. 22.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP ’23, page 611–626, New York, NY, USA, 2023. Association for Computing Machinery.
  23. 23.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  24. 24.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Re, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models, 2023.
  25. 25.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Guangxuan Xiao, and Song Han. Awq: Activation-aware weight quantization for on-device llm compression and acceleration. GetMobile: Mobile Comp. and Comm., 28(4):12–17, January 2025.
  26. 26.Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Duyu Tang, Kai Han, and Yunhe Wang. Kangaroo: lossless self-speculative decoding for accelerating llms via double early exiting. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA, 2025. Curran Associates Inc.
  27. 27.Jiahao Liu, Qifan Wang, Jingang Wang, and Xunliang Cai. Speculative decoding via early-exiting for faster LLM inference with Thompson sampling control mechanism. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 3027–3043, Bangkok, Thailand, August 2024. Association for Computational Linguistics.
  28. 28.Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, and Junchen Jiang. Cachegen: Kv cache compression and streaming for fast large language model serving. In Proceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM ’24, page 38–56, New York, NY, USA, 2024. Association for Computing Machinery.
  29. 29.Xuan Luo, Weizhi Wang, and Xifeng Yan. Adaptive layer-skipping in pre-trained llms, 2025.
  30. 30.Xin Men, Mingyu Xu, Qingyu Zhang, Qianhao Yuan, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. ShortGPT: Layers in large language models are more redundant than you expect. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 20192–20204, Vienna, Austria, July 2025. Association for Computational Linguistics.
  31. 31.Ruijie Miao, Yihan Yan, Xinshuo Yao, and Tong Yang. An efficient inference framework for early-exit large language models. arXiv preprint arXiv:2407.20272, 2024.
  32. 32.Anand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan, Swapnil Gandhi, and Ravi Netravali. Improving dnn inference throughput using practical, per-input compute adaptation. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP ’24, page 624–639, New York, NY, USA, 2024. Association for Computing Machinery.
  33. 33.Xuchen Pan, Yanxi Chen, Yaliang Li, Bolin Ding, and Jingren Zhou. Ee-tuning: An economical yet scalable solution for tuning early-exit large language models, 2024.
  34. 34.Priyadarshini Panda, Abhronil Sengupta, and Kaushik Roy. Conditional deep learning for energy-efficient and enhanced pattern recognition. In 2016 Design, Automation & Test in Europe Conference & Exhibition (DATE), pages 475–480, 2016.
  35. 35.Cory Perry. Introducing low-level gpu virtual memory management | nvidia technical blog, 12 2020. [Online; accessed 2025-08-12].
  36. 36.Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. vattention: Dynamic memory management for serving llms without pagedattention. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS ’25, page 1133–1150, New York, NY, USA, 2025. Association for Computing Machinery.
  37. 37.David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro. Mixture-of-depths: Dynamically allocating compute in transformer-based language models, 2024.
  38. 38.Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. Confident adaptive language modeling. NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc.
  39. 39.Yechao She, Tuo Shi, Jianping Wang, and Bin Liu. Dynamic batching and early-exiting for accurate and timely edge inference. In 2024 IEEE 99th Vehicular Technology Conference (VTC2024-Spring), pages 1–6, 2024.
  40. 40.Peng Tang, Pengkai Zhu, Tian Li, Srikar Appalaraju, Vijay Mahadevan, and R. Manmatha. DEED: Dynamic early exit on decoder for accelerating encoder-decoder transformer models. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages 116–131, Mexico City, Mexico, June 2024. Association for Computational Linguistics.
  41. 41.Surat Teerapittayanon, Bradley McDanel, and H.T. Kung. Branchynet: Fast inference via early exiting from deep neural networks. In 2016 23rd International Conference on Pattern Recognition (ICPR), pages 2464–2469, 2016.
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models, 2023.
  43. 43.Neeraj Varshney, Agneet Chatterjee, Mihir Parmar, and Chitta Baral. Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘LITE’. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages 3656–3677, Mexico City, Mexico, June 2024. Association for Computational Linguistics.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Hook, NY, USA, 2017. Curran Associates Inc.
  45. 45.Mengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi, Jianan Ji, Kit Ao, Praveen Velliengiri, Xupeng Miao, Oded Padon, and Zhihao Jia. Mirage: A multi-level superoptimizer for tensor programs, 2025.
  46. 46.Jiale Xu, Rui Zhang, Cong Guo, Weiming Hu, Zihan Liu, Feiyang Wu, Yu Feng, Shixuan Sun, Changxu Shao, Yuhong Guo, Junping Zhao, Ke Zhang, Minyi Guo, and Jingwen Leng. vtensor: Flexible virtual tensor management for efficient llm serving, 2024.
  47. 47.Jiaming Xu, Jiayi Pan, Yongkang Zhou, Siming Chen, Jinhao Li, Yaoxiu Lian, Junyi Wu, and Guohao Dai. Specee: Accelerating large language model inference with speculative early exiting. In Proceedings of the 52nd Annual International Symposium on Computer Architecture, ISCA ’25, page 467–481, New York, NY, USA, 2025. Association for Computing Machinery.
  48. 48.Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze. Flashinfer: Efficient and customizable attention engine for llm inference serving, 2025.
  49. 49.Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: A distributed serving system for Transformer-Based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pages 521–538, Carlsbad, CA, July 2022. USENIX Association.
  50. 50.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020.
  51. 51.Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. Sglang: efficient execution of structured language model programs. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA, 2025. Curran Associates Inc.
  52. 52.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. Bert loses patience: fast and robust inference with early exit. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA, 2020. Curran Associates Inc.

Citation

MLA
Liu, X., et al. “Dynamic Rebatching for Efficient Early-Exit Inference with DREX”. arXiv, 2025, http://arxiv.org/abs/2512.15705v1.
APA
Liu, X., Alexander, D., Kakarla, S. K. R., Arzani, B., & Liu, V. (2025). Dynamic Rebatching for Efficient Early-Exit Inference with DREX. arXiv. http://arxiv.org/abs/2512.15705v1
Chicago
Liu, X., D. Alexander, S. K. R. Kakarla, B. Arzani, and V. Liu. 2025. “Dynamic Rebatching for Efficient Early-Exit Inference with DREX”. arXiv. http://arxiv.org/abs/2512.15705v1.
Harvard
Liu, X. et al. (2025) “Dynamic Rebatching for Efficient Early-Exit Inference with DREX”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2512.15705v1.
Vancouver
1. Liu X, Alexander D, Kakarla SKR, Arzani B, Liu V (2025) Dynamic Rebatching for Efficient Early-Exit Inference with DREX. arXiv

BibTeX

@article{liu2025dynamic,
  title = {Dynamic Rebatching for Efficient Early-Exit Inference with DREX},
  author = {Liu, Xuting and Alexander, Daniel and Kakarla, Siva Kesava Reddy and Arzani, Behnaz and Liu, Vincent},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2512.15705v1},
  eprint = {2512.15705}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/