Performance Aware LLM Load Balancer for Mixed Workloads

Kunal JainA. ParayilAnkur MallickEsha ChoukseXiaoting QinJue ZhangÍñigo GoiriRujia WangChetan BansalVictor Ruehle

article2025SoCC22 citationsJay Lepreau Best Paper Award

Presents a workload-aware reinforcement learning load balancer that routes queries across model instances by predicting output lengths and estimating the interference between prefill and decode phases, reducing end-to-end inference latency by over 11%.

Listen

Large Language Model inference has become the dominant workload in cloud environments, driving up operational costs, infrastructure strain, and latency. Serving these models involves two distinct stages: a compute-heavy prompt phase and a memory-heavy decode phase. Current cloud load balancers treat inference requests as monolithic tasks and route them without considering these differing operational phases. This failure to account for workload diversity leads to interference when diverse queries are co-served on the same instance, causing performance spikes, frequent request preemptions, and higher overall latency.

The article develops and evaluates a workload-aware, reinforcement learning-based load balancer designed to optimize request routing across multiple identical model instances. The primary goal is to minimize end-to-end response latency by intelligently distributing incoming queries based on their prompt and decode characteristics.

The authors designed a routing framework comprising three core components: an output length predictor that categorizes expected response lengths, an analytical workload impact estimator that models the latency penalty of mixing diverse requests, and a heuristic-guided reinforcement learning routing agent. The system was evaluated using simulated mixed workloads spanning five standard natural language processing tasks, as well as a one-hour real production trace from a major cloud provider. Experiments benchmarked the system across configurations including Llama-2-7B models on clusters of V100 graphics processing units and Llama-3.1-8B models on A100 clusters, comparing performance against standard routing heuristics and instance-level scheduling techniques.

The investigation produced several key findings. First, random request assignment across instances causes roughly 10% higher end-to-end latency compared to optimal distribution, demonstrating that routing quality sets an upper bound on performance that instance-level schedulers cannot overcome alone. Second, the proposed workload-guided reinforcement learning router reduces average end-to-end latency by 11.43% compared to standard round-robin routing on mixed datasets, significantly outperforming classical load-balancing heuristics such as Join Shortest Queue and Min-Min scheduling. Third, the intelligent router substantially improves user experience metrics by lowering Time-To-First-Token and stabilizing streaming latency (Time-Between-Tokens) through reduced request preemption. Finally, the routing framework proved adaptable across hardware setups, scaling successfully to eight instances with an 11.62% latency improvement and maintaining a 7.84% latency advantage on production traces even when instance-level optimizations like chunked prefills were enabled.

These findings demonstrate that effective load balancing at the cluster ingress layer provides greater operational gains than relying solely on server-level scheduling. By reducing instance-level queue congestion and preventing costly preemptions, cloud operators can significantly improve service performance, enhance system throughput, and lower the infrastructure footprint required to support mixed artificial intelligence workloads.

Organizations operating multi-instance model clusters should transition away from oblivious routing methods toward workload-aware routing architectures that incorporate response-length estimation. Implementation should proceed by first profiling model-hardware combinations to estimate interference parameters and deploying lightweight response-prediction models. Because reinforcement learning decisions add slight operational overhead and the latency advantages decrease when workloads are overwhelmingly dominated by prompt tokens, operators should validate traffic profiles and test these routing mechanisms in pilot deployments before full-scale production rollout.

Cover for Performance Aware LLM Load Balancer for Mixed Workloads

Abstract

Large Language Model (LLM) workloads consist of distinct prefill and decode phases, each with unique compute and memory requirements that should be considered when routing input queries across cluster instances. However, existing load-balancing algorithms treat these workloads as monolithic jobs, ignoring the differences between the two phases. This oversight leads to suboptimal query distribution and increased response latency. In our work, we first characterize the factors affecting response latency during LLM inference. We show that balancing inference requests across available LLM instances can improve end-to-end latency more than simply optimizing the instance-level scheduler. Motivated by these findings, we propose a heuristic-guided, reinforcement learning-based router for data-driven, workload-aware scheduling. Our router distributes queries across LLM instances by using a trainable response-length predictor and a novel formulation for estimating the impact of mixing different workloads, achieving over 11% lower end-to-end latency than existing methods on mixed public datasets. Our framework represents a first step toward a holistic optimization framework and serves as a benchmark for deriving optimal load balancing strategies tailored to different reward functions and requirements. Beyond latency, we can extend the proposed framework to optimize for various performance criteria ensuring that the system meets diverse operational objectives.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Observation
  • 4 Intelligent Router: Design
  • 4.1 Output length predictor
  • 4.2 Workload impact estimator
  • 4.3 RL based router
  • 5 Experiments
  • 5.1 Performance evaluation of Intelligent router
  • 6 Limitations and Conclusion
  • References
  • A Related Work
  • A.1 LLM Serving Systems
  • A.2 LLM Serving Algorithms
  • A.3 Hybrid LLM Inference
  • A.4 Reinforcement Learning for routing jobs
  • B Additional Results
  • B.1 Different LLM and Hardware combination
  • B.2 Performance in the presence of SOTA Optimizations
  • C Appendix / supplemental material
  • C.1 Performance on Dataset
  • C.2 Batching and Routing Algorithms
  • C.3 Additional Baselines
  • C.4 Overhead of the Router
  • C.5 Details of Dataset
  • C.6 Prompt-Decode Distribution
  • C.7 Training details of the output length predictor
  • C.8 Task Predictability
  • C.9 Licenses
  • C.10 Details of RL training
  • C.11 Additional experiments to validate the scalability of proposed framework
  • C.12 Experiments on Real Production Trace from Cloud Provider X
  • C.13 Additional Proofs

Knowls

  1. Knowl 1 — Workload Interference and Request Mixing Impact Model

    model/method

    To capture the latency penalty incurred when co-serving requests with heterogeneous prompt and decode lengths on a single Large Language Model (LLM) instance mm, the interference between the prompt and decode phases is analytically formulated. Let nn be the number of requests currently executing on model instance mm, where pjmp_j^m and djmd_j^m denote the prompt and decode tokens processed by the jj-th existing request on instance mm. For an incoming request ii with pip_i prompt tokens and did_i decode tokens, the estimated prompt processing latency impact TpimT_{p_i}^m and its associated normalized prompt penalty rpimr_{p_i}^m are modeled as:

    Tpim=grad1×(pi2+∑j=1n(pjm+djm))T_{p_i}^m = \text{grad}_1 \times \left( p_i^2 + \sum_{j=1}^n (p_j^m + d_j^m) \right)

    rpim={1if Tpim≤ϵ1−Tpimϵotherwiser_{p_i}^m = \begin{cases} 1 & \text{if } T_{p_i}^m \le \epsilon \\ 1 - \frac{T_{p_i}^m}{\epsilon} & \text{otherwise} \end{cases}

    where grad1\text{grad}_1 is the empirical execution gradient with respect to prompt tokens, and ϵ\epsilon is a threshold hyperparameter above which prompt latency is penalized.

    The ongoing decode phase interference penalty rdmr_d^m imposed on existing requests by adding request ii is formulated as:

    rdm=−grad2×(∑j=1n(pjm+djm)+pi+di)r_d^m = -\text{grad}_2 \times \left( \sum_{j=1}^n (p_j^m + d_j^m) + p_i + d_i \right)

    where grad2\text{grad}_2 is the empirical decode phase execution gradient. The total mixing reward penalty for assigning request ii to instance mm during state transition st→st+1s_t \to s_{t+1} is given by:

    rmixing(st,st+1)=αrpim+(1−α)rdmr_{\text{mixing}}(s_t, s_{t+1}) = \alpha r_{p_i}^m + (1 - \alpha) r_d^m

    where α∈(0,1)\alpha \in (0, 1) is a weighting parameter balancing prompt and decode phase sensitivity.

  2. Knowl 2 — Heuristic-Guided RL Formulation for Request Routing Across LLM Instances

    model/method

    Load balancing incoming requests across mm homogeneous LLM instances is formulated as a discrete-time Markov Decision Process (MDP) with decision epoch Δt\Delta t equal to the minimum decode batch execution time (e.g., 0.020.02 s).

    State Space: The state vector includes:

    1. The central router waiting queue length wqtw_{q_t}.
    2. The exact prompt token count pt∈Rp_t \in \mathbb{R} and estimated decode length bucket dt∈{0,…,nd}d_t \in \{0, \dots, n_d\} of the head-of-line request.
    3. Matrices Pt∈Rm×npP_t \in \mathbb{R}^{m \times n_p} and Dt∈Rm×ndD_t \in \mathbb{R}^{m \times n_d} tracking the count of running requests across npn_p prompt buckets and ndn_d decode buckets for each of the mm model instances.
    4. Instance capacities Ct∈RmC_t \in \mathbb{R}^m as a function of batch size limits, PtP_t, and DtD_t.
    5. Estimated earliest request completion time T^ct\hat{T}_{c_t} across instances.

    Action Space: a∈{0,1,…,m}a \in \{0, 1, \dots, m\}, where a∈{0,…,m−1}a \in \{0, \dots, m-1\} routes the request to a specific instance, and a=ma = m defers routing (taking no action to wait for a better instance state).

    Reward Function: Prior knowledge of workload mixing is incorporated using heuristic guidance that anneals over training episodes kk:

    rt=−∑j∈J1T^j(1−fjt)+∑j=1m∑irw⋅wmit−(γ−γ~k)h(st,st+1)r_t = - \sum_{j \in \mathcal{J}} \frac{1}{\hat{T}_j} (1 - f_{jt}) + \sum_{j=1}^m \sum_i r_w \cdot w_{mit} - (\gamma - \tilde{\gamma}_k) h(s_t, s_{t+1})

    where J\mathcal{J} contains all active (queued and running) requests, T^j\hat{T}_j is the estimated ideal completion time of request jj, fjt=ojt/d^jtf_{jt} = o_{jt} / \hat{d}_{jt} is the fraction of output tokens generated at time tt, wmit∈{0,1}w_{mit} \in \{0, 1\} indicates whether request ii completed on instance mm at time tt, and rw∈Z+r_w \in \mathbb{Z}^+ is a positive completion reward. The heuristic penalty difference function h(st,st+1)h(s_t, s_{t+1}) is:

    h(st,st+1)=rmixing(st,st+1)−max⁡l∈{1,…,m}rmixing(st,st+1l)h(s_t, s_{t+1}) = r_{\text{mixing}}(s_t, s_{t+1}) - \max_{l \in \{1, \dots, m\}} r_{\text{mixing}}(s_t, s_{t+1}^l)

    The guidance discount factor is defined as γ~k=λkγ=e−βdkγ\tilde{\gamma}_k = \lambda_k \gamma = e^{-\beta_d k} \gamma with decay rate βd>0\beta_d > 0, ensuring the guidance term decays to zero as training progresses so the agent ultimately optimizes the true scheduling objective.

  3. Knowl 3 — Task-Conditioned Output Token Length Predictor

    model/method

    Because decode length is unknown at request arrival, a lightweight sequence length predictor is designed by fine-tuning a DistilBERT model to classify input prompts into output token buckets. Output buckets are partitioned based on estimated completion duration ranges (such as 0–0.5s, 0.5–2s, and 2–4s, translating to token ranges of roughly 0–250, 250–1000, and 1000–4000 tokens at a throughput of 500 tokens/s) rather than uniform token intervals.

    To resolve ambiguity across diverse tasks with distinct prompt-to-decode ratios, each prompt is prepended with a task type indicator (e.g., "This is a <task> task"). When task identity is not supplied, it is predicted using a DistilBERT task classifier that achieves 93.79% accuracy.

    Conditioning on task type increases DistilBERT decode bucket classification accuracy on a mixed corpus of 31,329 samples from 5.5% to 79.15% for unequal completion-time buckets, and from 9.3% to 68.23% for equal 250-token buckets.

  4. Knowl 4 — End-to-End Latency Evaluation on Mixed LLM Workloads

    empirical result

    The performance of the Workload-Guided Reinforcement Learning (WG-RL) router was evaluated against baseline routing schemes on a cluster of four LLaMA-2-7B instances hosted on NVIDIA V100 GPUs using vLLM with iteration-level First-Come-First-Served (FCFS) scheduling. Evaluation consisted of 20 episodes of 2,000 mixed requests sampled at an arrival rate of λ=20\lambda = 20 requests/s.

    • End-to-End Latency: Round Robin (RR) achieved baseline latency. Baseline RL improved overall servicing time by 7.53 s (4.35%), Workload-Aware RL (with static mixing penalty) improved by 13.50 s (7.79%), and Workload-Guided RL improved by 19.18 s (11.43% reduction in end-to-end latency).
    • Classical Heuristics: Join Shortest Queue, Maximum Capacity Usage, and Min-Min scheduling algorithms yielded marginal improvements of only 0.46%, 2.60%, and 1.50% over Round Robin, respectively.
    • Queue Management and Waiting Delays: Baseline RL kept average router wait time low (0.59 s) but suffered high instance-level queueing and preemption delays. Workload-Aware RL increased router wait time to 4.41 s. Workload-Guided RL achieved an optimal balance with an average router wait time of 2.05 s, minimizing instance waiting queues and avoiding request preemptions.
  5. Knowl 5 — LLM Serving Performance Under Hardware Shifts and Prefill Chunking

    data/table

    The generalizability of the Workload-Guided RL router was evaluated on four LLaMA-3.1-8B instances running on NVIDIA A100 GPUs with an arrival rate of 80 requests/second, tested both with and without chunked prefills enabled.

    Routing Algorithm Prefill Chunking Avg. E2E Latency (s) Improvement
    Round Robin No 248.41 –
    Baseline RL No 240.58 3.15%
    Workload Aware RL No 231.66 6.74%
    Workload Guided RL No 221.80 10.71%
    Round Robin Yes 247.30 0.45%
    Baseline RL Yes 240.68 3.11%
    Workload Aware RL Yes 231.12 6.96%
    Workload Guided RL Yes 220.93 11.06%

    The data shows that the Workload-Guided RL approach maintains an 11% latency reduction over Round Robin across different GPU architectures and model scales. Furthermore, it operates orthogonally to instance-level scheduler optimizations like chunked prefill, preserving lower Time-To-First-Token (TTFT) and lower Time-Between-Tokens (TBT) variance.

  6. Knowl 6 — Dominance of Request Routing Over Instance-Level Batching Schedulers

    data/table

    An empirical comparison of combinations of instance-level batching schedulers (Bin Packing, Least Work Left, and First-Come-First-Served) and routing strategies across two LLM instances processing 3,000 requests was performed over four request arrival sequences: (1) random Light-Heavy (LH) and Heavy-Light (HL) mix, (2) random mix of all four classes, (3) LH requests arriving first followed by HL, and (4) HL requests arriving first followed by LH.

    Batching Algorithm Routing Algorithm Total End-to-End Latency (s)
    (LH, HL random) (Random) (LH, then HL) (HL, then LH)
    Bin Packing Dedicated Small-Large 704.50 644.75 566.25 588.15
    Round Robin 581.50 559.30 424.80 440.68
    Decode Balancer 595.82 555.40 424.82 440.81
    Least Work Left Dedicated Small-Large 704.50 641.81 566.25 588.15
    Round Robin 585.14 554.00 424.64 440.82
    Decode Balancer 596.95 559.97 424.66 440.81
    FCFS Dedicated Small-Large 704.50 648.66 566.25 588.15
    Round Robin 607.45 572.16 424.80 440.82
    Decode Balancer 605.65 573.17 424.82 440.81

    The results demonstrate that the routing policy is the primary determinant of system latency. When suboptimal server partitioning (Dedicated Small-Large) is used, all batching schedulers perform poorly (~704.5 s). Conversely, under ordered arrival sequences (Scenarios 3 and 4), all batching algorithms produce nearly identical latencies (~424.8 s and ~440.8 s), proving that instance-level batching schedulers cannot compensate for suboptimal routing assignments.

  7. Knowl 7 — Evaluation on Cloud Production Traces Using Random Forest Predictor

    empirical result

    The Workload-Guided RL router was evaluated on a 1-hour real production trace from a commercial cloud provider comprising 4,000 requests (mean prompt length: 5,526.64 tokens; mean decode length: 112.69 tokens) served on LLaMA-3.1-8B at 80 requests/s with maximum batched chunk size of 1024.

    To accommodate privacy constraints where prompt text is inaccessible, the output length predictor was replaced by a Random Forest model trained solely on prompt token count and application name metadata. This predictor achieved 79.0% bucket classification accuracy.

    Compared to Round Robin's average latency of 1,005.31 s, Baseline RL achieved 982.38 s (2.28% improvement), Workload-Aware RL achieved 961.17 s (4.39% improvement), and Workload-Guided RL achieved 926.49 s (7.84% improvement). The performance gain is smaller than on synthetic datasets because the extreme dominance of prompt lengths over decode lengths naturally reduces decode-phase preemptions.

  8. Knowl 8 — Scalability to Eight Homogeneous LLM Instances

    empirical result

    The scalability of the load-balancing framework was tested on a cluster of eight model instances serving 4,000 requests at an arrival rate of 40 requests/s. To accommodate the expanded state space across 8 instances, the parameter count of the Double Deep Q-Network (Double-DQN) was scaled proportionally.

    Under this setting, Baseline RL, Workload-Aware RL, and Workload-Guided RL reduced average end-to-end latency relative to Round Robin by 5.84%, 6.64%, and 11.62%, respectively, demonstrating that the heuristic-guided reinforcement learning formulation scales effectively as the instance count doubles.

  9. Knowl 9 — Inference Overhead and Limitations of Workload-Guided Routing

    limitation

    The Workload-Guided RL router incurs two computational steps per decision:

    1. Length Predictor Inference: DistilBERT evaluation takes 0.01 seconds on GPU and 0.8 seconds on CPU per batch of 64 requests.
    2. RL Agent Decision: Double-DQN forward pass takes fewer than 10610^6 FLOPs (sub-millisecond execution time).

    Operational Limitations:

    • When prompt lengths vastly exceed decode lengths (e.g., prompt-heavy workloads with short outputs), the rate of GPU memory exhaustion and decode preemption is low, reducing the relative latency advantage of intelligent routing over Round Robin (from ~11.4% down to 7.84%).
    • Tabular Q-learning is infeasible due to the state space size (approximately 5×10445 \times 10^{44} reachable states for 4 instances with queue capacity up to 512), requiring function approximation via Double-DQN.

Coverage note — None was omitted; all key contributions including the analytical mixing formulation, the RL MDP and heuristic guidance design, output length prediction mechanisms, empirical latency evaluations across datasets/hardware/optimizations, and scalability/overhead analyses are fully covered.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977 (2020).
  2. 2.Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. arXiv:2403.02310 [cs.LG]
  3. 3.Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee. 2023. Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills. arXiv preprint arXiv:2308.16369 (2023).
  4. 4.Huankai Chen, Frank Wang, Na Helian, and Gbola Akanmu. 2013. User-priority guided Min-Min scheduling algorithm for load balancing in cloud computing. In 2013 National Conference on Parallel Computing Technologies (PARCOMPTECH). 1–8. doi:10.1109/ParCompTech.2013.6621389
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021).
  6. 6.Ching-An Cheng, Andrey Kolobov, and Adith Swaminathan. 2021. Heuristic-guided reinforcement learning. Advances in Neural Information Processing Systems 34 (2021), 13550–13563.
  7. 7.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems 35 (2022), 16344–16359.
  8. 8.Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017. Results of the WNUT2017 Shared Task on Novel and Emerging Entity Recognition. In Proceedings of the 3rd Workshop on Noisy User-generated Text. Association for Computational Linguistics, Copenhagen, Denmark, 140–147. doi:10.18653/v1/W17-4418
  9. 9.Dujian Ding, Sihem Amer-Yahia, and Laks VS Lakshmanan. 2022. On Efficient Approximate Queries over Machine Learning Models. arXiv preprint arXiv:2206.02845 (2022).
  10. 10.Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks VS Lakshmanan, and Ahmed Hassan Awadallah. 2024. Hybrid LLM: Cost-efficient and quality-aware query routing. arXiv preprint arXiv:2404.14618 (2024).
  11. 11.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long Form Question Answering. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 3558–3567. doi:10.18653/v1/p19-1346
  12. 12.Google. [n. d.]. Vertex AI. https://cloud.google.com/vertex-ai.
  13. 13.Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, et al. 2024. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads. arXiv preprint arXiv:2401.11181 (2024).
  14. 14.HuggingFace. [n. d.]. Hugging Face Inference API. https://huggingface.co/inference-api.
  15. 15.Neharika Jali, Guannan Qu, Weina Wang, and Gauri Joshi. 2024. Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems. arXiv preprint arXiv:2402.01147 (2024).
  16. 16.Siddharth Jha, Coleman Hooper, Xiaoxuan Liu, Sehoon Kim, and Kurt Keutzer. 2024. Learned Best-Effort LLM Serving. arXiv preprint arXiv:2401.07886 (2024).
  17. 17.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024).
  18. 18.Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei. 2023. S3S^3: Increasing GPU Utilization during Generative Inference for Higher Throughput. In Thirty-seventh Conference on Neural Information Processing Systems. https://openreview.net/forum?id=zUYfbdNl1m
  19. 19.Anil Kag, Igor Fedorov, Aditya Gangrade, Paul Whatmough, and Venkatesh Saligrama. 2022. Efficient Edge Inference by Selective Query. In The Eleventh International Conference on Learning Representations.
  20. 20.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles.
  21. 21.Jiamin Li, Le Xu, Hong Xu, and Aditya Akella. 2024. BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models. arXiv preprint arXiv:2404.18322 (2024).
  22. 22.Bin Lin, Tao Peng, Chen Zhang, Minmin Sun, Lanbo Li, Hanyu Zhao, Wencong Xiao, Qi Xu, Xiafei Qiu, Shen Li, et al. 2024. Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache. arXiv preprint arXiv:2401.02669 (2024).
  23. 23.Jiachen Liu, Zhiyu Wu, Jae-Won Chung, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury. 2024. Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services. arXiv preprint arXiv:2404.16283 (2024).
  24. 24.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning Word Vectors for Sentiment Analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Portland, Oregon, USA, 142–150. http://www.aclweb.org/anthology/P11-1015
  25. 25.Daniel Mendoza, Francisco Romero, and Caroline Trippel. 2024. Model Selection for Latency-Critical Inference Serving. In Proceedings of the Nineteenth European Conference on Computer Systems. 1016–1038.
  26. 26.Microsoft. [n. d.]. Azure AI Studio. https://ai.azure.com/.
  27. 27.Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica. 2024. Routellm: Learning to route llms with preference data. arXiv preprint arXiv:2406.18665 (2024).
  28. 28.OpenAI. [n. d.]. OpenAI Platform. https://platform.openai.com/overview.
  29. 29.OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL]
  30. 30.Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini. 2023. Splitwise: Efficient generative LLM inference using phase splitting. arXiv:2311.18677 [cs.AR]
  31. 31.Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Shengkun Cui, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2024. One Queue Is All You Need: Resolving Head-of-Line Blocking in Large Language Model Serving. arXiv:2407.00047 [cs.DC] https://arxiv.org/abs/2407.00047
  32. 32.Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2024. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention. arXiv preprint arXiv:2405.04437 (2024).
  33. 33.Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T Kalbarczyk, Tamer Başar, and Ravishankar K Iyer. 2024. Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction. arXiv preprint arXiv:2404.08509 (2024).
  34. 34.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392. doi:10.18653/v1/D16-1264 arXiv:1606.05250 [cs.CL]
  35. 35.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2020. Recipes for building an open-domain chatbot. arXiv preprint arXiv:2004.13637 (2020).
  36. 36.Benjamin Spector and Chris Re. 2023. Accelerating llm inference with staged speculative decoding. arXiv preprint arXiv:2308.04623 (2023).
  37. 37.Alessandro Staffolani, Victor-Alexandru Darvariu, Paolo Bellavista, and Mirco Musolesi. 2023. RLQ: Workload allocation with reinforcement learning in distributed queues. IEEE Transactions on Parallel and Distributed Systems 34, 3 (2023), 856–868.
  38. 38.Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024. Llumnix: Dynamic Scheduling for Large Language Model Serving. arXiv preprint arXiv:2406.03243 (2024).
  39. 39.Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA.
  40. 40.Jörg Tiedemann. 2012. Parallel Data, Tools and Interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association (ELRA), Istanbul, Turkey, 2214–2218. http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf
  41. 41.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 [cs.CL]
  42. 42.Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023. Fast distributed inference serving for large language models. arXiv preprint arXiv:2305.05920 (2023).
  43. 43.Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 521–538. https://www.usenix.org/conference/osdi22/presentation/yu
  44. 44.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al. 2024. H2o: Heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems 36 (2024).
  45. 45.Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. arXiv preprint arXiv:2401.09670 (2024).

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/