Performance Aware LLM Load Balancer for Mixed Workloads
Kunal JainA. ParayilAnkur MallickEsha ChoukseXiaoting QinJue ZhangÍñigo GoiriRujia WangChetan BansalVictor Ruehle
Presents a workload-aware reinforcement learning load balancer that routes queries across model instances by predicting output lengths and estimating the interference between prefill and decode phases, reducing end-to-end inference latency by over 11%.
Large Language Model inference has become the dominant workload in cloud environments, driving up operational costs, infrastructure strain, and latency. Serving these models involves two distinct stages: a compute-heavy prompt phase and a memory-heavy decode phase. Current cloud load balancers treat inference requests as monolithic tasks and route them without considering these differing operational phases. This failure to account for workload diversity leads to interference when diverse queries are co-served on the same instance, causing performance spikes, frequent request preemptions, and higher overall latency.
The article develops and evaluates a workload-aware, reinforcement learning-based load balancer designed to optimize request routing across multiple identical model instances. The primary goal is to minimize end-to-end response latency by intelligently distributing incoming queries based on their prompt and decode characteristics.
The authors designed a routing framework comprising three core components: an output length predictor that categorizes expected response lengths, an analytical workload impact estimator that models the latency penalty of mixing diverse requests, and a heuristic-guided reinforcement learning routing agent. The system was evaluated using simulated mixed workloads spanning five standard natural language processing tasks, as well as a one-hour real production trace from a major cloud provider. Experiments benchmarked the system across configurations including Llama-2-7B models on clusters of V100 graphics processing units and Llama-3.1-8B models on A100 clusters, comparing performance against standard routing heuristics and instance-level scheduling techniques.
The investigation produced several key findings. First, random request assignment across instances causes roughly 10% higher end-to-end latency compared to optimal distribution, demonstrating that routing quality sets an upper bound on performance that instance-level schedulers cannot overcome alone. Second, the proposed workload-guided reinforcement learning router reduces average end-to-end latency by 11.43% compared to standard round-robin routing on mixed datasets, significantly outperforming classical load-balancing heuristics such as Join Shortest Queue and Min-Min scheduling. Third, the intelligent router substantially improves user experience metrics by lowering Time-To-First-Token and stabilizing streaming latency (Time-Between-Tokens) through reduced request preemption. Finally, the routing framework proved adaptable across hardware setups, scaling successfully to eight instances with an 11.62% latency improvement and maintaining a 7.84% latency advantage on production traces even when instance-level optimizations like chunked prefills were enabled.
These findings demonstrate that effective load balancing at the cluster ingress layer provides greater operational gains than relying solely on server-level scheduling. By reducing instance-level queue congestion and preventing costly preemptions, cloud operators can significantly improve service performance, enhance system throughput, and lower the infrastructure footprint required to support mixed artificial intelligence workloads.
Organizations operating multi-instance model clusters should transition away from oblivious routing methods toward workload-aware routing architectures that incorporate response-length estimation. Implementation should proceed by first profiling model-hardware combinations to estimate interference parameters and deploying lightweight response-prediction models. Because reinforcement learning decisions add slight operational overhead and the latency advantages decrease when workloads are overwhelmingly dominated by prompt tokens, operators should validate traffic profiles and test these routing mechanisms in pilot deployments before full-scale production rollout.
- Paper: DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, Yinmin Zhong et al. (2024). DistServe characterizes the fundamental compute-versus-memory divergence between the prefill and decoding phases in LLM serving, establishing the foundational problem of phase-interference that this paper's load balancer actively schedules to mitigate.
- Paper: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills, Amey Agrawal et al. (2023). SARATHI establishes how mixing chunked prefill and decoding workloads impacts GPU utilization and pipeline bubbles during LLM inference, directly motivating this paper's approach to workload-aware request distribution.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). PagedAttention introduces the foundational memory and batch-scheduling primitives for LLM serving engines that this paper builds upon when balancing request loads across instances.
- Paper: Clipper: a low-latency online prediction serving system, Daniel Crankshaw et al. (2017). Clipper pioneered feedback-driven, low-latency online prediction routing and dynamic batching architectures that underpin modern cluster-level inference load-balancing frameworks.
- Paper: When is Routing Meaningful? Diversity and Robustness in Language Model Societies, Fantine Huot et al. (2026). This work explores societal diversity and perturbation robustness in multi-model routing policies, extending the practical design space for learned routing architectures like the RL-based router proposed here.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). The Conductor builds upon multi-instance LLM orchestration principles to dynamically partition, schedule, and route reasoning subtasks across heterogeneous model agents using natural language coordination.
