Kairos: A Scalable Serving System for Physical AI

Yinwei DaiGanesh AnanthanarayananLandon CoxXenofon FoukasBozidar RadunovicRavi Netravali

article2026arXiv0 citations

Introduces Kairos, the first multi-robot serving system designed for physical AI that manages interleaved action generation and execution to reduce end-to-end task latency by up to 66.5% across growing robot fleets.

Listen

As physical artificial intelligence advances across humanoids, robotic arms, and automated platforms in warehouses and factories, organizations face a critical infrastructure bottleneck. Operating fleets of robots using frontier models requires offloading model execution to shared computing resources. However, conventional serving architectures designed for digital artificial intelligence fail in robotic applications. Unlike text generation, robotic tasks operate in a multi-round loop that alternates between generating groups of actions and executing them in the physical world. Because observations become outdated as the environment shifts, existing systems force developers to use rigid, worst-case execution lengths, which overloads computing infrastructure. Additionally, standard scheduling systems ignore the time robots spend physically moving, causing severe misprioritization across robot fleets.

The article demonstrates and evaluates Kairos, the first specialized serving system tailored for multi-robot physical artificial intelligence. The objective is to evaluate whether actively integrating physical execution awareness into model serving can significantly reduce overall task latency, maintain task success, and scale efficiently across large robot fleets.

The researchers assessed Kairos using an empirical approach combining extensive simulated benchmarks and real-world physical robot deployments. The evaluation tested six foundation models across three distinct architectures, including vision-language-action, video-action, and world-action models. Testing spanned five established simulation suites and a physical dual-arm manipulation platform executing precise transfer tasks. Using trace-driven replays with up to 100 concurrent robots across edge and cloud hardware, the article compared Kairos against standard first-in, first-out schedulers and state-of-the-art fairness-based agent schedulers.

The article reveals several critical findings. First, extracting confidence signals from intermediate generation steps enables dynamic, per-round tuning of action lengths, achieving up to 2.67 times longer safe execution spans at matching task accuracy or up to a 30 percent boost in accuracy at equivalent spans. Second, in high-demand serving environments, Kairos cuts average task completion latency by 31.8 percent to 66.5 percent compared to standard serving schedulers. Third, performance improvements scale directly with fleet size; latency savings expand from 20.4 percent with 10 robots to 42.8 percent with 100 robots. Finally, in hybrid edge-cloud configurations, Kairos intelligently manages capacity limits, lowering average latency by 36.9 percent to 47.7 percent relative to edge-only configurations while shielding operations from network delays.

These findings indicate that treating physical robot motion as an active scheduling parameter unlocks massive infrastructure efficiencies. For operational leaders, this means organizations can support significantly larger robot deployments per computing node, dramatically reducing cloud and edge hardware costs while speeding up real-world workflows. By eliminating uncoordinated computing delays, robotic fleets operate more smoothly without stalling, minimizing operational risks.

Organizations scaling robotic fleets should transition from static digital schedulers to execution-aware serving architectures that support dynamic horizon adjustment. Teams deploying foundation models should configure serving layers to track robot motion states and utilize hybrid edge-cloud offloading to absorb sudden demand spikes. Before complete operational rollouts, teams should conduct internal pilot profiling to establish optimal confidence thresholds and network transfer tolerances for their specific robotic tasks.

The results provide high confidence across common manipulation benchmarks, model architectures, and controlled real-robot tasks. Readers should note that testing relied heavily on simulated environments and trace replays driven by statistical Poisson arrival models. Although physical tests confirmed that scheduling delays did not compromise mechanical execution accuracy, real-world deployments subject to extreme network volatility or highly unpredictable safety stops may require additional validation.

arXiv: 2605.11381
Cover for Kairos: A Scalable Serving System for Physical AI

Abstract

Physical AI is experiencing rapid growth with frontier foundation models increasing its capabilities across general environments. Physical AI tasks are characterized by inference properties that are markedly different from digital AI. They consist of multiple rounds of inference and action execution, generating a chunk of actions in each inference round, and asynchronously interleaving inference and execution. This makes existing digital AI serving systems unsuited for physical AI; a shortcoming that is critical for enabling their wide adoption, considering their size and the scale of the robot fleets they have to serve. To fill this gap, we design Kairos, the first multi-robot serving system that makes the generate-execute loop a first-class citizen, with active involvement in the execution phase. Across a wide range of physical AI models and robots, Kairos reduces the average end-to-end task latency by 31.8--66.5% over state-of-the-art digital AI serving practices, with gains scaling with the robot fleet size.

Table of Contents

  • 1 Introduction
  • 2 Background and Motivation
  • 2.1 Physical AI Models
  • 2.2 Execution Horizon
  • 2.3 Execution-aware scheduling
  • 3 Kairos: System Design
  • 3.1 Dynamic Horizon with Inference Confidence
  • 3.2 Execution-aware Scheduler
  • 3.2.1 Wait time of physical AI tasks.
  • 3.2.2 Scheduling policy.
  • 3.2.3 Support hybrid edge and cloud serving.
  • 4 Implementation
  • 5 Evaluation
  • 5.1 Experimental Setup
  • 5.2 Effectiveness of Diffusion Confidence
  • 5.3 End-to-End System Performance
  • 5.4 Trace Fidelity Evaluation with Real Robots
  • 5.5 Sensitivity and Ablation Studies
  • 6 Additional Related Work
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Dynamic Execution Horizon Selection via Diffusion Step Confidence

    model/method

    Physical AI foundation models (such as Vision-Language-Action models, Video Action Models, and World Action Models) employ action chunking, generating a block of NN action vectors simultaneously through KK iterative diffusion or flow matching refinement steps. The execution horizon H≤NH \le N dictates the number of actions executed from a chunk before triggering a new inference round with updated environmental observations.

    To dynamically select HH per round without requiring auxiliary models or multiple forward-pass consensus checks, Kairos extracts confidence directly from intermediate diffusion updates. Across the KK refinement steps, updates follow a coarse-to-fine trajectory: confident action predictions stabilize early and incur diminishing updates in later steps, whereas uncertain action predictions continue to receive large updates through the final step KK.

    For a generated action chunk A1,A2,…,ANA_1, A_2, \dots, A_N in an action space of dimension DD (e.g., D=7D=7 for 3D position, 3D orientation, and gripper state), the model produces a per-action update vector ΔAi(k)\Delta A_i^{(k)} at each step k∈{1,…,K}k \in \{1, \dots, K\}. Kairos evaluates the convergence of each action sequentially starting from A1A_1. Let μi=1K−1∑k=1K−1∥ΔAi(k)∥\mu_i = \frac{1}{K-1} \sum_{k=1}^{K-1} \|\Delta A_i^{(k)}\| be the mean update magnitude over earlier steps, and let δi=∥ΔAi(K)∥\delta_i = \|\Delta A_i^{(K)}\| be the magnitude of the final step update. Given a user-specified confidence threshold t≥0t \ge 0, Kairos finds the first action index ii such that:

    δi>(1+t)⋅μi\delta_i > (1 + t) \cdot \mu_i

    The threshold-based horizon is set to the prefix of converged actions, Hthresh=i−1H_{\text{thresh}} = i - 1. Because robotic manipulation requires continuous trajectory validity, encountering an unconverged action at index ii invalidates subsequent actions Ai+1,…,ANA_{i+1}, \dots, A_N regardless of their individual convergence. To ensure forward progress, the effective horizon assigned for the round is:

    H=max⁡(Hthresh,Hmin⁡)H = \max(H_{\text{thresh}}, H_{\min})

    where Hmin⁡H_{\min} is a minimum static horizon floor.

  2. Knowl 2 — Phase-Dominant Wait Time and Wait Ratio for Multi-Round Physical AI

    equation

    In physical AI serving, a task consists of multiple sequential rounds of model inference generation (GG) and physical robot execution (EE). Under asynchronous execution, generation for round j+1j+1 is initiated while the physical robot is still executing actions from round jj.

    The true stall or wait time WjW_j incurred between round jj and round j+1j+1 depends on whether round jj is generation-dominated (∣Gj∣≥∣Ej∣|G_j| \ge |E_j|) or execution-dominated (∣Gj∣<∣Ej∣|G_j| < |E_j|):

    Wj={Gj+1.start−Gj.end,if ∣Gj∣≥∣Ej∣Ej+1.start−Ej.end,if ∣Gj∣<∣Ej∣W_j = \begin{cases} G_{j+1}.\text{start} - G_j.\text{end}, & \text{if } |G_j| \ge |E_j| \\ E_{j+1}.\text{start} - E_j.\text{end}, & \text{if } |G_j| < |E_j| \end{cases}

    where ∣Gj∣|G_j| and ∣Ej∣|E_j| represent the durations of the jj-th generation and execution phases, respectively.

    To balance scheduling priority fairly across tasks with varying total runtimes and prevent long-running tasks from monopolizing GPU resources, Kairos normalizes the total accumulated wait time by the task's elapsed wall-clock lifetime. The wait ratio wrwr for request rr of a task arriving at tstartt_{\text{start}} and evaluated at current scheduling time tnowt_{\text{now}} is:

    wr=∑jWjtnow−tstartwr = \frac{\sum_{j} W_j}{t_{\text{now}} - t_{\text{start}}}

    A higher wait ratio indicates that a task has suffered disproportionate stalls relative to its active lifetime and should be prioritized for GPU dispatch.

  3. Knowl 3 — Kairos Execution-Aware Scheduling Algorithm

    algorithm

    Kairos schedules concurrent inference requests across multi-round robotic tasks using a two-level priority scheme: coarse-grained priority bucketing based on wait ratios, followed by within-bucket ordering by estimated execution duration with starvation aging.

    Input: Pending inference requests R\mathcal{R}, bucket count BB, aging interval AA, edge profile Pe\mathcal{P}_e with capacity NeN_e, cloud profile Pc\mathcal{P}_c with capacity NcN_c
    Output: Dispatch assignments
    // Phase 1: Compute wait ratio and bin requests
    for each request r∈Rr \in \mathcal{R} do
        for each round jj with recorded generation interval GjG_j and execution interval EjE_j do
            if ∣Gj∣≥∣Ej∣|G_j| \ge |E_j| then
                Wj←Gj+1.start−Gj.endW_j \leftarrow G_{j+1}.\text{start} - G_j.\text{end}
            else
                Wj←Ej+1.start−Ej.endW_j \leftarrow E_{j+1}.\text{start} - E_j.\text{end}
            end if
        end for
        wr←∑jWj/(tnow−tstartr)w_r \leftarrow \sum_j W_j / (t_{\text{now}} - t_{\text{start}}^r)
        br←⌊wr⋅B⌋b_r \leftarrow \lfloor w_r \cdot B \rfloor
        if r.skipped≥Ar.\text{skipped} \ge A then
            br←min⁡(B−1,br+⌊r.skipped/A⌋)b_r \leftarrow \min(B - 1, b_r + \lfloor r.\text{skipped} / A \rfloor)
        end if
    end for
    // Phase 2: Within-bucket request ordering
    S←[]\mathcal{S} \leftarrow []
    for b=B−1b = B - 1 down to 0 do
        for each request r∈bucket br \in \text{bucket } b do
            e^r←∣Elastr∣⋅(1+r.skipped)\hat{e}_r \leftarrow |E_{\text{last}}^r| \cdot (1 + r.\text{skipped})
        end for
        sort requests r∈bucket br \in \text{bucket } b by e^r\hat{e}_r in descending order
        append sorted requests of bucket bb to S\mathcal{S}
    end for
    // Phase 3: Shortest estimated delay guided placement
    Se←S[1:Ne]\mathcal{S}_e \leftarrow \mathcal{S}[1 : N_e]
    Sc←[]\mathcal{S}_c \leftarrow []
    for each request r∈S[Ne+1:∣S∣]r \in \mathcal{S}[N_e + 1 : |\mathcal{S}|] do
        if estimated cloud latency < estimated edge delay and ∣Sc∣<Nc|\mathcal{S}_c| < N_c then
            append rr to Sc\mathcal{S}_c
        end if
    end for
    for each request r∈Se∪Scr \in \mathcal{S}_e \cup \mathcal{S}_c with stale observation do
        r.obs←fetch_obs(r.client)r.\text{obs} \leftarrow \text{fetch\_obs}(r.\text{client})
    end for
    Dispatch Se\mathcal{S}_e to edge instances and Sc\mathcal{S}_c to cloud instances
    for each request r∈Rr \in \mathcal{R} do
        if r∈Se∪Scr \in \mathcal{S}_e \cup \mathcal{S}_c then
            r.skipped←0r.\text{skipped} \leftarrow 0
        else
            r.skipped←r.skipped+1r.\text{skipped} \leftarrow r.\text{skipped} + 1
        end if
    end for

    In Phase 1, wait ratios in [0,1][0, 1] are mapped to BB equal-width buckets (default B=10B=10). To prevent starvation, if a request has been skipped for ≥A\ge A consecutive rounds, its bucket index is boosted. In Phase 2, within each bucket, requests are ordered by descending estimated execution duration e^r=∣Elastr∣⋅(1+r.skipped)\hat{e}_r = |E_{\text{last}}^r| \cdot (1 + r.\text{skipped}), using the duration of the prior physical execution phase ∣Elastr∣|E_{\text{last}}^r| as a continuity predictor for the next round. In Phase 3, requests exceeding edge capacity are offloaded to cloud instances if network round-trip plus cloud compute time is shorter than expected edge queuing delay. Requests whose sensor observations grew stale while queued fetch fresh observations from client runtimes prior to execution.

  4. Knowl 4 — Kairos Client-Server Serving System Architecture

    model/method

    Kairos operates as a distributed client-server serving architecture designed to handle concurrent multi-round physical AI tasks across heterogeneous edge and cloud accelerators.

    On the client side, a lightweight shim embeds into the robot runtime (e.g., LeRobot) exposing get_obs and send_act. For each inference round, the client transmits a gRPC payload containing:

    1. Task ID: A unique task identifier.
    2. Round ID: A monotonically increasing round sequence number.
    3. Obs: Visual and proprioceptive sensory inputs.
    4. Last Exec Info: Physical execution timing metadata from the preceding round, including execution start timestamp, end timestamp, and the number of remaining actions in the current queue.
    5. Horizon Policy: A client-specified execution horizon rule (e.g., confidence threshold tt).

    On the server side, an asynchronous event loop maintains request lifecycles. A Request States Manager tracks round intervals to calculate wait times and wait ratios. The Execution-Aware Scheduler groups requests into dynamic batches up to the hardware saturation throughput limit defined by offline Engine Profiles (compute and network curves for edge and cloud instances). Models are compiled via torch.compile. On the response path, an Action Manager applies the dynamic horizon trimming policy to generated action chunks before returning them to clients.

  5. Knowl 5 — Accuracy-Efficiency Pareto Dominance of Confidence-Based Dynamic Horizons

    empirical result

    Evaluating dynamic execution horizons driven by diffusion confidence threshold tt against static execution horizons hh across six representative physical AI model-benchmark workloads demonstrates strict Pareto dominance.

    The tested configurations include:

    1. LIBERO simulator with SmolVLA
    2. Meta-World simulator with SmolVLA
    3. LIBERO simulator with XVLA
    4. Isaac Lab simulator with GR00T N1.5 on a Fourier GR1 humanoid
    5. LIBERO simulator with Pi0.5
    6. Real bimanual SO-101 robot with Pi0.5
    7. RoboTwin simulator with Fast-WAM
    8. SIMPLER/Bridge simulator with mimic-video

    At equivalent task success rates (accuracy), confidence-based dynamic horizon adaptation yields up to 2.67×2.67\times longer average execution horizons compared to fixed horizons hh. At matched average execution horizons, confidence-based horizon adaptation improves task accuracy by up to 30 percentage points over static horizon assignments. In 62%–87% of generation rounds, tasks safely execute up to 2.8×2.8\times more actions without accuracy degradation.

  6. Knowl 6 — Latency Reduction and Component Ablations Across Physical AI Workloads

    data/table

    Under peak online serving loads modeled as a Poisson arrival process on a single edge server equipped with an NVIDIA RTX A6000 GPU, Kairos achieves substantial reductions in end-to-end task latency compared to static FIFO and Autellix baselines across eight physical AI workloads. Ablations isolate the individual contributions of the execution-aware scheduler and the dynamic execution horizon.

    Kairos (Full) Kairos w/o dynamic horizon Kairos w/o scheduler
    Workload P25 Avg P95 P25 Avg P95 P25 Avg P95
    (a) Latency reduction vs. FIFO Baseline
    LIBERO / SmolVLA 39.5% 31.8% 34.6% 21.1% 10.9% 9.9% 3.9% 16.8% 29.0%
    LIBERO / XVLA 74.5% 51.4% 45.2% 39.4% 17.8% 4.5% 44.6% 41.9% 37.5%
    LIBERO / Pi0.5 76.4% 65.6% 51.9% 57.0% 33.0% 5.7% 58.0% 54.9% 49.6%
    MetaWorld / SmolVLA 86.9% 60.5% 43.0% 59.0% 34.1% 4.3% 49.9% 43.7% 39.3%
    Isaac / GR00T N1.5 59.2% 41.6% 29.0% 41.7% 24.5% 7.6% 24.0% 25.4% 23.4%
    Bimanual / Pi0.5 53.3% 39.9% 22.2% 47.9% 28.1% 4.1% 24.9% 24.3% 18.9%
    RoboTwin / Fast-WAM 57.9% 33.0% 25.7% 29.0% 8.1% 8.0% 46.5% 32.5% 21.7%
    Bridge / mimic-video 52.2% 40.1% 28.5% 51.0% 27.2% 2.8% 36.6% 31.4% 24.5%
    (b) Latency reduction vs. Autellix Baseline
    LIBERO / SmolVLA 41.9% 33.1% 36.1% 24.2% 12.5% 11.9% 8.0% 15.1% 25.5%
    LIBERO / XVLA 74.9% 53.7% 46.7% 40.3% 21.6% 7.1% 43.2% 39.4% 36.1%
    LIBERO / Pi0.5 77.4% 66.5% 52.0% 59.0% 34.6% 5.9% 55.8% 53.9% 49.8%
    MetaWorld / SmolVLA 88.4% 63.6% 48.2% 63.5% 39.3% 13.2% 44.5% 41.3% 36.0%
    Isaac / GR00T N1.5 61.6% 43.3% 29.2% 45.2% 26.7% 7.9% 27.5% 24.2% 22.9%
    Bimanual / Pi0.5 55.9% 42.3% 24.6% 50.8% 31.0% 7.0% 23.9% 23.1% 19.8%
    RoboTwin / Fast-WAM 62.4% 36.6% 25.2% 36.6% 12.9% 7.3% 44.2% 34.2% 22.5%
    Bridge / mimic-video 54.6% 41.6% 28.7% 53.4% 29.1% 4.9% 34.9% 30.4% 25.9%

    Full Kairos reduces average latency by 31.8%–66.5%, P25 latency by 39.5%–88.4%, and P95 latency by 22.2%–52.0% across workloads relative to FIFO and Autellix. Autellix performs worse than FIFO because its generation-only fairness prioritization starves tasks with short execution phases. Ablating the dynamic horizon (Kairos w/o dynamic horizon) demonstrates that execution-aware scheduling alone reduces average latency by 10.9%–39.3%. Ablating the scheduler (Kairos w/o scheduler) confirms that dynamic horizon adaptation alone yields 15.1%–54.9% average latency reductions, but cannot match full Kairos performance without execution-aware arbitration.

  7. Knowl 7 — Scalability Across Dedicated Robot Fleets and Hybrid Edge-Cloud Serving

    empirical result

    Kairos exhibits strong scaling across dedicated multi-robot fleet sizes and hybrid edge-cloud infrastructure:

    1. Dedicated Fleet Scaling: In offline deployments where a shared server executes fixed task sets back-to-back for robot fleets scaled from n=10n=10 to n=100n=100, Kairos's average latency reduction over static FIFO and Autellix expands monotonically with contention:

      • 10 robots: 20.4%–21.4% average latency reduction.
      • 50 robots: 33.9%–37.8% average latency reduction.
      • 100 robots: 41.0%–42.8% average latency reduction.
    2. Hybrid Edge-Cloud Deployment: When augmenting an edge server (RTX A6000 GPU) with an offload cloud server (NVIDIA A100 80GB GPU) connected via a 100 ms WAN link (1 Gbps symmetric bandwidth), Kairos dynamically dispatches overflow bursts. At peak arrival rates, hybrid Kairos reduces average end-to-end task latency by 36.9%–47.7% compared to edge-only serving and by 51.9%–67.9% compared to cloud-only serving across workloads.

  8. Knowl 8 — Trace Fidelity and Robustness to Contention-Induced Delays on Real Robots

    empirical result

    Trace-driven evaluation was validated against physical hardware using a bimanual SO-101 robot executing a 20-trial object handoff-and-place task under injected queueing contention delays corresponding to Poisson arrival rates of 0.5, 1.0, and 2.0 tasks/s.

    Across all arrival rates, physical task completion accuracy remained constant at 13/20 to 14/20 successful trials, matching the baseline contention-free execution success rate of 13/20 (65%). Stalling the physical robot while awaiting delayed inference action chunks does not degrade physical execution fidelity or policy accuracy once the resumed actions are executed.

  9. Knowl 9 — Sensitivity to Priority Bucketing, Intra-Bucket Ordering, and Network Latency

    empirical result

    Microbenchmarking the architectural parameters of the Kairos execution-aware scheduler reveals the following sensitivities:

    1. Wait Ratio Bucket Count (BB): Partitioning wait ratios into B∈[5,10]B \in [5, 10] equal-width buckets achieves optimal end-to-end latency. Coarsening to B=2B=2 buckets degrades average latency by up to 33.9% due to priority blurring. Unbucketed continuous sorting (B=∞B=\infty) eliminates prioritization hysteresis, causing frequent request thrashing.

    2. Intra-Bucket Ordering: Ordering requests within each bucket by descending predicted execution duration (longest execution first) reduces tail (P95) latency by up to 17.9% compared to intra-bucket FIFO ordering, while simultaneously lowering average latency.

    3. Network Offload Sensitivity: In hybrid edge-cloud serving, base network latency is the dominant factor governing cloud offloading feasibility. Reducing base WAN latency from 150 ms to 50 ms increases cloud offloading volume by up to 20.5% and substantially lowers task latency, whereas increasing WAN bandwidth from 500 Mbps to 1500 Mbps produces minimal latency gains due to the compact payload size of physical AI observations.

Coverage note — No substantial contributed material was omitted from the paper.

References

  1. 1.1X Technologies. https://www.1x.tech, 2024.
  2. 2.Figure AI. https://www.figure.ai, 2024.
  3. 3.Unitree Robotics. https://www.unitree.com, 2024.
  4. 4.Fourier gr-1. https://www.fftai.com/products-gr1, 2026.
  5. 5.Nvidia isaac lab. https://developer.nvidia.com/isaac/lab, 2026.
  6. 6.Nvidia isaac sim. https://developer.nvidia.com/isaac/sim, 2026.
  7. 7.So-101. https://huggingface.co/docs/lerobot/en/so101, 2026.
  8. 8.R. Abhyankar, Z. He, V. Srivatsa, H. Zhang, and Y. Zhang. Infercept: Efficient intercept support for augmented large language model inference, 2024.
  9. 9.Amazon. Amazon robotics. https://www.aboutamazon.com/news/operations/amazon-robotics-robots-fulfillment-center, 2024.
  10. 10.Anthropic. The Claude 3 model family: Opus, Sonnet, Haiku. Technical report, 2024. https : / / www - cdn . anthropic . com / de8ba9b01c9ab7cbabf5c33b80b7bbc618857627 / Model _ Card _ Claude_3.pdf.
  11. 11.K. Black, M. Y. Galliker, and S. Levine. Real-time execution of action chunking flow policies, 2025.
  12. 12.A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.
  13. 13.R. Cadene, S. Aliberts, F. Capuano, M. Aractingi, A. Zouitine, P. Kooijmans, J. Choghari, M. Russi, C. Pascal, S. Palma, M. Shukor, J. Moss, A. Soare, D. Aubakirova, Q. Lhoest, Q. Gallouedec, and T. Wolf. Lerobot: An open-source library for end-to-end robot learning, 2026.
  14. 14.T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Z. Li, Q. Liang, X. Lin, Y. Ge, Z. Gu, W. Deng, Y. Guo, T. Nian, X. Xie, Q. Chen, K. Su, T. Xu, G. Liu, M. Hu, H. ang Gao, K. Wang, Z. Liang, Y. Qin, X. Yang, P. Luo, and Y. Mu. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation, 2025.
  15. 15.C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion, 2024.
  16. 16.N. Corporation. Tensorrt-llm: An open-source library to accelerate inference of large language models on nvidia gpus. https://github.com/NVIDIA/TensorRT-LLM, 2023. Library: TensorRT-LLM.
  17. 17.H. Fang, Y. Liu, Y. Du, L. Du, and H. Yang. Sqap-vla: A synergistic quantization-aware pruning framework for high-performance vision-language-action models, 2025.
  18. 18.Gemini Team, Google. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  19. 19.I. Gim, S. seob Lee, and L. Zhong. Asynchronous llm function calling, 2024.
  20. 20.J. Gu, M. Chowdhury, K. G. Shin, Y. Zhu, M. Jeon, J. Qian, H. Liu, and C. Guo. Tiresias: a gpu cluster manager for distributed deep learning. In Proceedings of the 16th USENIX Conference on Networked Systems Design and Implementation, NSDI’19, page 485–500, USA, 2019. USENIX Association.
  21. 21.P. B. Hansen. Operating system principles. Prentice-Hall, Inc., USA, 1973.
  22. 22.P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky. ̄̀0.5: a vision-language-action model with open-world generalization, 2025.
  23. 23.T. Jiang, X. Jiang, Y. Ma, X. Wen, B. Li, K. Zhan, P. Jia, Y. Liu, S. Sun, and X. Lang. The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning, 2025.
  24. 24.W. Jiang, J. Clemons, K. Sankaralingam, and C. Kozyrakis. How fast can i run my vla? demystifying vla inference performance with vla-perf, 2026.
  25. 25.M. J. Kim, Y. Gao, T.-Y. Lin, Y.-C. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M.-Y. Liu, C. Finn, and J. Gu. Cosmos policy: Fine-tuning video models for visuomotor control and planning, 2026.
  26. 26.M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model, 2024.
  27. 27.W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  28. 28.S.-W. Lee, X. Kang, and Y.-L. Kuo. Diff-dagger: Uncertainty estimation with diffusion policy for robotic manipulation, 2025.
  29. 29.H. Li, Q. Mang, R. He, Q. Zhang, H. Mao, X. Chen, H. Zhou, A. Cheung, J. Gonzalez, and I. Stoica. Continuum: Efficient and robust multi-turn llm agent scheduling with kv cache time-to-live, 2026.
  30. 30.P. Li, Y. Zhou, D. Muhtar, L. Yin, S. Yan, L. Shen, S. Vosoughi, and S. Liu. Diffusion language models know the answer before decoding, 2026.
  31. 31.S. Li, Y. Gao, D. Sadigh, and S. Song. Unified video action model, 2025.
  32. 32.X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao. Evaluating real-world robot manipulation policies in simulation, 2024.
  33. 33.Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling, 2023.
  34. 34.B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023.
  35. 35.Y. Liu et al. A survey of embodied AI in healthcare: Techniques, applications, and opportunities. arXiv preprint arXiv:2501.07468, 2025.
  36. 36.Z. Liu, Y. Chen, H. Cai, T. Lin, S. Yang, Z. Liu, and B. Zhao. Vla-pruner: Temporal-aware dual-level visual token pruning for efficient vision-language-action inference, 2026.
  37. 37.M. Luo, X. Shi, C. Cai, T. Zhang, J. Wong, Y. Wang, C. Wang, Y. Huang, Z. Chen, J. E. Gonzalez, and I. Stoica. Autellix: An efficient serving engine for llm agents as general programs, 2025.
  38. 38.T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control, 2026.
  39. 39.M. Nuyens and A. Wierman. The foreground–background queue: A survey. Performance Evaluation, 65(3):286–307, 2008.
  40. 40.NVIDIA, :, J. Bjorck, F. Castaneda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu. Gr00t n1: An open foundation model for generalist humanoid robots, 2025.
  41. 41.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  42. 42.J. Pai, L. Achenbach, V. Montesinos, B. Forrai, O. Mees, and E. Nava. mimic-video: Video-action models for generalizable robot control beyond vlas, 2025.
  43. 43.S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez. Gorilla: Large language model connected with massive apis, 2023.
  44. 44.S. Pohland, X. Foukas, G. Ananthanarayanan, A. Kolobov, S. Mehrotra, B. Radunovic, and A. Verma. Offload or overload: A platform measurement study of mobile robotic manipulation workloads, 2026.
  45. 45.A. Prasad, K. Lin, J. Wu, L. Zhou, and J. Bohg. Consistency policy: Accelerated visuomotor policies via consistency distillation, 2024.
  46. 46.Robots Guide. Sawyer robot. https://robotsguide.com/robots/sawyer, 2026.
  47. 47.H. Shi, B. Xie, Y. Liu, L. Sun, F. Liu, T. Wang, E. Zhou, H. Fan, X. Zhang, and G. Huang. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation, 2026.
  48. 48.M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, S. Alibert, M. Cord, T. Wolf, and R. Cadene. Smolvla: A vision-language-action model for affordable and efficient robotics, 2025.
  49. 49.J. Tang, Y. Sun, Y. Zhao, S. Yang, Y. Lin, Z. Zhang, J. Hou, Y. Lu, Z. Liu, and S. Han. Vlash: Real-time vlas via future-state-aware asynchronous inference, 2025.
  50. 50.G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, S. Bohez, K. Bousmalis, A. Brohan, T. Buschmann, A. Byravan, S. Cabi, K. Caluwaerts, F. Casarini, O. Chang, J. E. Chen, X. Chen, H.-T. L. Chiang, K. Choromanski, D. D’Ambrosio, S. Dasari, T. Davchev, C. Devin, N. D. Palo, T. Ding, A. Dostmohamed, D. Driess, Y. Du, D. Dwibedi, M. Elabd, C. Fantacci, C. Fong, E. Frey, C. Fu, M. Giustina, K. Gopalakrishnan, L. Graesser, L. Hasenclever, N. Heess, B. Hernaez, A. Herzog, R. A. Hofer, J. Humplik, A. Iscen, M. G. Jacob, D. Jain, R. Julian, D. Kalashnikov, M. E. Karagozler, S. Karp, C. Kew, J. Kirkland, S. Kirmani, Y. Kuang, T. Lampe, A. Laurens, I. Leal, A. X. Lee, T.-W. E. Lee, J. Liang, Y. Lin, S. Maddineni, A. Majumdar, A. H. Michaely, R. Moreno, M. Neunert, F. Nori, C. Parada, E. Parisotto, P. Pastor, A. Pooley, K. Rao, K. Reymann, D. Sadigh, S. Saliceti, P. Sanketi, P. Sermanet, D. Shah, M. Sharma, K. Shea, C. Shu, V. Sindhwani, S. Singh, R. Soricut, J. T. Springenberg, R. Sterneck, R. Surdulescu, J. Tan, J. Tompson, V. Vanhoucke, J. Varley, G. Vesom, G. Vezzani, O. Vinyals, A. Wahid, S. Welker, P. Wohlhart, F. Xia, T. Xiao, A. Xie, J. Xie, P. Xu, S. Xu, Y. Xu, Z. Xu, Y. Yang, R. Yao, S. Yaroshenko, W. Yu, W. Yuan, J. Zhang, T. Zhang, A. Zhou, and Y. Zhou. Gemini robotics: Bringing ai into the physical world, 2025.
  51. 51.Tesla. Tesla optimus. https://www.tesla.com/optimus, 2024.
  52. 52.Z. Wang, Z. Li, A. Mandlekar, Z. Xu, J. Fan, Y. Narang, L. Fan, Y. Zhu, Y. Balaji, M. Zhou, M.-Y. Liu, and Y. Zeng. One-step diffusion policy: Fast visuomotor policies via diffusion distillation, 2024.
  53. 53.Y. Xu, X. Kong, T. Chen, and D. Zhuo. Conveyor: Efficient tool-aware llm serving with tool partial execution, 2024.
  54. 54.Y. Xu, Y. Yang, Z. Fan, Y. Liu, Y. Li, B. Li, and Z. Zhang. Qvla: Not all channels are equal in vision-language-action model’s quantization, 2026.
  55. 55.H. Yan, Z. Zhong, J. Zhu, J. He, W. Yuan, W. Song, X. Gong, Y. Cai, G. Zhao, X. Yan, B. Liu, Y.-C. Chen, and H. Li. S-vam: Shortcut video-action model by self-distilling geometric and semantic foresight, 2026.
  56. 56.J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. R. Narasimhan, and O. Press. SWE-agent: Agent-computer interfaces enable automated software engineering. In NeurIPS, 2024.
  57. 57.S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. React: Synergizing reasoning and acting in language models, 2023.
  58. 58.A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, M. Cao, P. Li, Q. Deng, W. Mei, X. Wang, X. Chen, X. Zhou, Y. Wang, Y. Chang, Y. Li, Y. Zhou, Y. Ye, Z. Liu, and Z. Zhu. Gigaworld-policy: An efficient action-centered world–action model, 2026.
  59. 59.H. Ye, J. Yuan, R. Xia, X. Yan, T. Chen, J. Yan, B. Shi, and B. Zhang. Training-free adaptive diffusion with bounded difference approximation strategy, 2024.
  60. 60.S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y. Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y. Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y. Du, Y. Chebotar, S. Reed, J. Kautz, Y. Zhu, L. J. Fan, and J. Jang. World action models are zero-shot policies, 2026.
  61. 61.T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning, 2021.
  62. 62.T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination?, 2026.
  63. 63.K. Zhang, M. Sharma, J. Liang, and O. Kroemer. A modular robotic arm control stack for research: Franka-interface and frankapy, 2020.
  64. 64.T. Z. Zhao, V. Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023.
  65. 65.J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y. Feng, Y. Zheng, J. Zou, Y. Chen, J. Zeng, Y.-Q. Zhang, J. Pang, J. Liu, T. Wang, and X. Zhan. X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model, 2025.
  66. 66.L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng. Sglang: Efficient execution of structured language model programs, 2024.
  67. 67.S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig. WebArena: A realistic web environment for building autonomous agents. In ICLR, 2024.

Citation

MLA
Dai, Y., et al. “Kairos: A Scalable Serving System for Physical AI”. arXiv, 2026, http://arxiv.org/abs/2605.11381v1.
APA
Dai, Y., Ananthanarayanan, G., Cox, L., Foukas, X., Radunovic, B., & Netravali, R. (2026). Kairos: A Scalable Serving System for Physical AI. arXiv. http://arxiv.org/abs/2605.11381v1
Chicago
Dai, Y., G. Ananthanarayanan, L. Cox, X. Foukas, B. Radunovic, and R. Netravali. 2026. “Kairos: A Scalable Serving System for Physical AI”. arXiv. http://arxiv.org/abs/2605.11381v1.
Harvard
Dai, Y. et al. (2026) “Kairos: A Scalable Serving System for Physical AI”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.11381v1.
Vancouver
1. Dai Y, Ananthanarayanan G, Cox L, Foukas X, Radunovic B, Netravali R (2026) Kairos: A Scalable Serving System for Physical AI. arXiv

BibTeX

@article{dai2026kairos,
  title = {Kairos: A Scalable Serving System for Physical AI},
  author = {Dai, Yinwei and Ananthanarayanan, Ganesh and Cox, Landon and Foukas, Xenofon and Radunovic, Bozidar and Netravali, Ravi},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.11381v1},
  eprint = {2605.11381}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/