Kairos: A Scalable Serving System for Physical AI
Yinwei DaiGanesh AnanthanarayananLandon CoxXenofon FoukasBozidar RadunovicRavi Netravali
Introduces Kairos, the first multi-robot serving system designed for physical AI that manages interleaved action generation and execution to reduce end-to-end task latency by up to 66.5% across growing robot fleets.
As physical artificial intelligence advances across humanoids, robotic arms, and automated platforms in warehouses and factories, organizations face a critical infrastructure bottleneck. Operating fleets of robots using frontier models requires offloading model execution to shared computing resources. However, conventional serving architectures designed for digital artificial intelligence fail in robotic applications. Unlike text generation, robotic tasks operate in a multi-round loop that alternates between generating groups of actions and executing them in the physical world. Because observations become outdated as the environment shifts, existing systems force developers to use rigid, worst-case execution lengths, which overloads computing infrastructure. Additionally, standard scheduling systems ignore the time robots spend physically moving, causing severe misprioritization across robot fleets.
The article demonstrates and evaluates Kairos, the first specialized serving system tailored for multi-robot physical artificial intelligence. The objective is to evaluate whether actively integrating physical execution awareness into model serving can significantly reduce overall task latency, maintain task success, and scale efficiently across large robot fleets.
The researchers assessed Kairos using an empirical approach combining extensive simulated benchmarks and real-world physical robot deployments. The evaluation tested six foundation models across three distinct architectures, including vision-language-action, video-action, and world-action models. Testing spanned five established simulation suites and a physical dual-arm manipulation platform executing precise transfer tasks. Using trace-driven replays with up to 100 concurrent robots across edge and cloud hardware, the article compared Kairos against standard first-in, first-out schedulers and state-of-the-art fairness-based agent schedulers.
The article reveals several critical findings. First, extracting confidence signals from intermediate generation steps enables dynamic, per-round tuning of action lengths, achieving up to 2.67 times longer safe execution spans at matching task accuracy or up to a 30 percent boost in accuracy at equivalent spans. Second, in high-demand serving environments, Kairos cuts average task completion latency by 31.8 percent to 66.5 percent compared to standard serving schedulers. Third, performance improvements scale directly with fleet size; latency savings expand from 20.4 percent with 10 robots to 42.8 percent with 100 robots. Finally, in hybrid edge-cloud configurations, Kairos intelligently manages capacity limits, lowering average latency by 36.9 percent to 47.7 percent relative to edge-only configurations while shielding operations from network delays.
These findings indicate that treating physical robot motion as an active scheduling parameter unlocks massive infrastructure efficiencies. For operational leaders, this means organizations can support significantly larger robot deployments per computing node, dramatically reducing cloud and edge hardware costs while speeding up real-world workflows. By eliminating uncoordinated computing delays, robotic fleets operate more smoothly without stalling, minimizing operational risks.
Organizations scaling robotic fleets should transition from static digital schedulers to execution-aware serving architectures that support dynamic horizon adjustment. Teams deploying foundation models should configure serving layers to track robot motion states and utilize hybrid edge-cloud offloading to absorb sudden demand spikes. Before complete operational rollouts, teams should conduct internal pilot profiling to establish optimal confidence thresholds and network transfer tolerances for their specific robotic tasks.
The results provide high confidence across common manipulation benchmarks, model architectures, and controlled real-robot tasks. Readers should note that testing relied heavily on simulated environments and trace replays driven by statistical Poisson arrival models. Although physical tests confirmed that scheduling delays did not compromise mechanical execution accuracy, real-world deployments subject to extreme network volatility or highly unpredictable safety stops may require additional validation.
- Paper: Orca: A Distributed Serving System for Transformer-Based Generative Models, Gyeong-In Yu et al. (2022). Understanding iteration-level scheduling and selective batching in Orca provides essential background on the digital AI serving paradigms that Kairos contrasts with and adapts for physical AI generate-execute loops.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Examining π0's vision-language-action flow architecture and action-chunking execution provides critical context on the physical AI foundation models whose inference characteristics Kairos is designed to serve.
- Paper: DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, Yinmin Zhong et al. (2024). Reading DistServe clarifies how phase disaggregation and goodput optimization operate in large language model serving systems before seeing how Kairos addresses multi-round generate-and-execute workloads.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Understanding RT-2's translation of web-scale multimodal foundation models into robotic action tokens helps illustrate the computational and latency demands Kairos optimizes across robot fleets.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). GR00T N1 demonstrates the dual-system architecture combining slower multimodal planning with high-frequency action chunking that forms the core workload targeted by Kairos.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This survey offers a comprehensive overview of Vision-Language-Action model architectures, action chunking mechanisms, and real-time execution constraints in embodied AI.
- Paper: Clipper: a low-latency online prediction serving system, Daniel Crankshaw et al. (2017). Clipper introduces foundational principles of low-latency online prediction serving and adaptive batching under strict latency targets.
- Paper: Ray: A Distributed Framework for Emerging AI Applications, Philipp Moritz et al. (2018). Ray outlines distributed computing and dynamic execution models for continuously interacting with dynamic environments and robotic policies.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA builds directly on low-latency physical AI execution principles by introducing continuous inference and latent-aware action streaming for real-time dynamic object manipulation.
- Paper: WarmPrior: Straightening Flow-Matching Policies with Temporal Priors, Sinjae Kang et al. (2026). WarmPrior extends generative action-chunking efficiency by incorporating temporal priors into flow-matching policies to reduce inference latency and computational overhead.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA applies flow matching and dual-pathway architectures for embodied robot control, benefiting from the scalable physical AI serving mechanisms introduced in Kairos.
