Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Ling XuChuyu HanBorui LiHao WuShiqi JiangTing CaoChuanyou LiShenghui ZhongShuai Wang

article2026arXiv1 citations

Presents Embodied.cpp, a portable C++ inference runtime that unifies vision-language-action and world-action model deployment across heterogeneous robot hardware, delivering up to 2.7x speedups and 77% lower memory consumption for real-time closed-loop control.

Listen

Practical deployment of artificial intelligence models on physical robots faces a major systems challenge. Although vision-language-action models and world-action models are advancing rapidly, deploying them on heterogeneous and resource-constrained edge hardware remains highly fragmented. Current AI serving engines are built for cloud-based request-response tasks that optimize overall throughput. In contrast, physical robotics requires low-latency, batch-size-one closed-loop control, coordination across modules operating at different execution rates, and flexible interfaces that accommodate diverse sensor inputs and physical action outputs.

The article introduces and evaluates Embodied.cpp, an open-source, portable C++ inference runtime designed to standardize and accelerate the execution of diverse embodied AI models across heterogeneous robotic platforms and simulators.

To address deployment inefficiencies without hard-coding specific model structures, the authors developed a five-layer runtime architecture: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. This framework decouples execution schedules, optimizes small-batch computation for edge processors, and provides pluggable operator interfaces. The runtime was evaluated against standard Python deployment baselines across three vision-language-action models (pi0.5, GR00T N1.7, and HY-VLA) across multiple precision configurations, as well as two world-action models (Cosmos3 and LingBot-VA).

The evaluation revealed several key findings. First, Embodied.cpp delivered overall inference speedups ranging from 1.05-fold to 2.70-fold across evaluated configurations compared to Python baselines, achieving significant latency reductions per generated action step. Second, the runtime dramatically lowered onboard memory demands, decreasing video RAM usage by 7% to 77% depending on the model and quantization level (including 8-bit, 6-bit, and 4-bit configurations). Third, for world-action models, memory footprints dropped substantially—such as a 33.6% reduction for LingBot-VA from 24.75 GB down to 16.44 GB—while keeping operational success nearly identical to baseline levels (98.00% versus 100.00%). Across nearly all vision-language-action tests, the C++ implementation preserved high control task success rates.

These findings demonstrate that organizations can replace fragile, bespoke Python deployment glue code with a unified, lightweight runtime. By significantly lowering memory consumption and operational latency without degrading task success, the runtime reduces onboard hardware costs, shortens development cycles, and enables advanced foundation models to operate on power- and compute-constrained edge robotics devices.

Engineering and robotics teams deploying physical AI systems should evaluate adopting a unified C++ runtime architecture to streamline model integration. Organizations can leverage lower-bit quantization paths to deploy larger models on smaller edge processors, balancing minor trade-offs in success rate against hardware savings. Further work should focus on expanding testing across more real-world robotic environments, validating additional custom hardware accelerators, and testing next-generation multi-component models.

While the reported results show high confidence in latency and memory gains across tested benchmarks and simulators, users should exercise caution regarding boundary conditions. Real-world mechanical uncertainties, sensor noise, and highly aggressive quantization (such as 4-bit on certain architectures) may affect control quality depending on the specific robotic task.

Cover for Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Abstract

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O. We present Embodied..cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied..cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through one backend abstraction. We evaluate Embodied..cpp on three VLA and two WAM models, using normalized comparisons across Python and C++ quantization configurations. Overall, Embodied..cpp achieves 1.05x-2.70x inference speedups and 7%-77% lower VRAM relative to Python baselines, while maintaining near-baseline success for most configurations. These results show that Embodied..cpp improves deployment efficiency while preserving high control quality across diverse embodied model architectures.

Project Link: this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work and Motivation
  • 2.1 Embodied AI Model and Architectural Analysis
  • 2.2 Existing Inference Runtimes for AI Models
  • 2.3 Existing Acceleration Methods for Model Inference
  • 3 Project Overview
  • 3.1 Challenges
  • 3.2 Design Principles and Runtime Architecture
  • 4 Evaluation
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Five-layer embodied inference runtime

    model/method

    Embodied.cpp organizes embodied-model inference into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. Input adapters convert online sensor streams and offline dataset samples into runtime inputs; sequence builders assemble the model’s required sequence and context; backbone execution supplies the shared model-execution path; head plugins handle model-specific outputs such as actions or predicted futures; and deployment adapters connect outputs to simulators or robot software. This design keeps the deployment boundary reusable while allowing model-specific components to vary. The system overview diagram on page 9 depicts sensor and dataset inputs entering through adapters and outputs connecting to simulators and real devices.

  2. Knowl 2 — Architectural taxonomy of VLA and world-action models

    model/method

    Embodied.cpp’s architectural analysis distinguishes vision-language-action (VLA) models, which primarily map perception and instructions to actions, from world-action models (WAMs), which include future prediction as part of online control. The analysis identifies four VLA structures: AR-token models generate action tokens autoregressively from one backbone; VLM-backed models feed a continuous-action head from a shared pretrained vision-language backbone; hierarchical models use a high-level planner to produce subgoals for a low-level controller; and asynchronous models coordinate modules that run at different rates through buffered state. It identifies four WAM structures: predict-then-act models predict future states before a downstream action expert acts; unified autoregressive models generate future and action tokens in one sequence; shared-backbone models reuse a backbone for world modeling and action generation while retaining distinct auxiliary blocks; and latent-space models predict compact latent futures or subgoals for an action expert. The analysis treats these architectures as sharing a reusable execution path while differing chiefly in pluggable heads, predictive modules, and execution structure. The taxonomy is summarized in the architecture diagram and comparison on pages 3–4.

  3. Knowl 3 — Modular multi-rate execution

    model/method

    Embodied.cpp is designed to let model components execute at different refresh rates rather than forcing every component to run on every control step. Its proposed runtime structure exposes execution units and pluggable modules, with shared state or feature pools and configurable refresh policies. For example, perception features can be refreshed less often than actions are produced, a predictive branch can run only when future estimation is required, and a high-rate action head can continue to use buffered state. This capability targets hierarchical, asynchronous, and world-action workloads whose modules operate on different timescales.

  4. Knowl 4 — Latency-first execution for closed-loop control

    model/method

    Embodied.cpp targets predictable, low-latency, low-jitter inference for batch-1 closed-loop robot or simulator control, rather than maximizing throughput across large batches. Its latency-first design supports graph replay, buffer reuse, operator fusion, backend-specific dispatch, and careful host–device data movement to make small-batch execution efficient across heterogeneous hardware. The target devices include edge platforms and workstation-class systems; the paper presents these as runtime design capabilities, not as a quantified jitter guarantee.

  5. Knowl 5 — Extensible embodied inputs, outputs, and operators

    model/method

    Embodied.cpp is designed to extend beyond a fixed token interface through typed embodied inputs, pluggable model heads, deployment adapters, and reusable operators and model-specific kernels collected in an embodied AI kernel warehouse. Inputs can include images, language, proprioception, history, force or tactile signals, and simulator-provided state. Outputs can include discrete action tokens, continuous action vectors, action chunks, predicted futures, or intermediate control representations. Deployment adapters bridge runtime outputs to simulators and real-robot software stacks, so new model interfaces can be added without replacing the entire runtime.

  6. Knowl 6 — Normalized VLA deployment results

    data/table

    The page-10 benchmark compares Python deployments with Embodied.cpp C++ deployments at full precision and 8-, 6-, and 4-bit configurations for pi0.5, GR00T N1.7, and HY-VLA. Each metric is normalized independently to that model’s Python baseline of 1.00. Latency is measured per generated action step before normalization; lower latency and VRAM are preferable, while higher success rate is preferable. The C++ paths reduce normalized latency for all three models. HY-VLA’s 4-bit path has the lowest latency ratio, 0.37, equivalent to a 2.70× speedup relative to its Python baseline. Quantized paths lower VRAM, but the success-rate reduction for pi0.5 at 4-bit is notable (0.70); the other listed configurations retain success rates close to their baselines.

    Inference latency Success rate VRAM
    Model Python C++ C++ 8-bit C++ 6-bit C++ 4-bit Python C++ C++ 8-bit C++ 6-bit C++ 4-bit Python C++ C++ 8-bit C++ 6-bit C++ 4-bit
    pi0.5 1.00 0.90 0.88 0.95 0.88 1.00 0.92 0.90 0.93 0.70 1.00 0.60 0.41 0.35 0.30
    GR00T N1.7 1.00 0.72 0.70 0.70 0.65 1.00 0.96 0.97 0.97 0.96 1.00 0.93 0.57 0.48 0.38
    HY-VLA 1.00 0.48 0.38 0.39 0.37 1.00 1.02 1.00 1.01 0.99 1.00 0.68 0.27 0.25 0.23
  7. Knowl 7 — WAM deployment results

    data/table

    The page-10 WAM evaluation compares Python and Embodied.cpp C++ deployments of Cosmos3 and LingBot-VA using closed-loop success rate and VRAM. Cosmos3’s C++ deployment changes success from 49.00% to 48.00% and reduces VRAM from 21.84 GB to 19.49 GB. LingBot-VA’s C++ deployment changes success from 100.00% to 98.00% and reduces VRAM from 24.75 GB to 16.44 GB, a 33.6% reduction. These two deployments therefore show lower memory use with small reported success-rate decreases; the paper does not give additional hardware, task, or trial-count details alongside these results.

    Model Backend Success rate VRAM
    Cosmos3 Python 49.00% 21.84 GB
    Cosmos3 C++ 48.00% 19.49 GB
    LingBot-VA Python 100.00% 24.75 GB
    LingBot-VA C++ 98.00% 16.44 GB

Coverage note — The extensive related-work survey and background comparisons were omitted because they contextualize the runtime but are not contributions of this paper; no standalone algorithm or theoretical result is presented.

References

  1. 1.Karl Pertsch et al. OpenVLA: An Open-Source Vision-Language-Action Model. arXiv preprint arXiv:2406.09246, 2024. URL: https://arxiv.org/abs/2406.09246.
  2. 2.Kevin Black et al. Pi0: A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164, 2024. URL: https://arxiv.org/abs/2410.24164.
  3. 3.Physical Intelligence et al. Pi0.5: A Vision-Language-Action Model with Open-World Generalization. arXiv preprint arXiv:2504.16054, 2025. URL: https://arxiv.org/abs/2504.16054.
  4. 4.Johan Bjorck et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv preprint arXiv:2503.14734, 2025. URL: https://arxiv.org/abs/2503.14734.
  5. 5.L. Li et al. Causal World Modeling for Robot Control. arXiv preprint arXiv:2601.21998, 2026. URL: https://arxiv.org/abs/2601.21998.
  6. 6.S. Wang, J. Shi, Z. Fu, X. He, F. Liu, C. Yang, Y. Zhou, Z. Fei, J. Gong, J. Fu, M. Z. Shou, X. Huang, X. Qiu, and Y.-G. Jiang. World Action Models: The Next Frontier in Embodied AI. arXiv preprint arXiv:2605.12090, 2026. Submitted May 12, 2026. URL: https://arxiv.org/abs/2605.12090.
  7. 7.Xueying Li, Feng Lyu, Hao Wu, Mingliu Liu, Jia-Nan Liu, and Guozi Liu. Stop wandering: Efficient vision-language navigation via metacognitive reasoning. arXiv preprint arXiv:2604.02318, 2026.
  8. 8.Hugging Face. LeRobot. GitHub repository, 2026. Accessed June 17, 2026. URL: https://github.com/huggingface/lerobot.
  9. 9.Open X-Embodiment Collaboration. Open X-Embodiment. Project website, 2026. Accessed June 17, 2026. URL: https://robotics-transformer-x.github.io/.
  10. 10.ManiSkill Team. ManiSkill. Project website, 2026. Accessed June 17, 2026. URL: https://maniskill.ai/.
  11. 11.LIBERO Team. LIBERO. Project website, 2026. Accessed June 17, 2026. URL: https://libero-project.github.io/.
  12. 12.NVIDIA. Isaac Sim. Product website, 2026. Accessed June 17, 2026. URL: https://developer.nvidia.com/isaac/sim.
  13. 13.Georgi Gerganov et al. llama.cpp. GitHub repository, 2026. 2023–2026. URL: https://github.com/ggml-org/llama.cpp.
  14. 14.Microsoft. ONNX Runtime Documentation. Official documentation, 2026. Accessed June 17, 2026. URL: https://onnxruntime.ai/docs/.
  15. 15.LMSYS Org. SGLang. Official documentation and repository, 2026. Accessed June 17, 2026. URL: https://docs.sglang.io/.
  16. 16.vLLM Project. vLLM-Omni. Official documentation and repository, 2026. Accessed June 17, 2026. URL: https://docs.vllm.ai/projects/vllm-omni/en/latest/.
  17. 17.Anthony Brohan et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv preprint arXiv:2307.15818, 2023. URL: https://arxiv.org/abs/2307.15818.
  18. 18.Octo Model Team et al. Octo: An Open-Source Generalist Robot Policy. arXiv preprint arXiv:2405.12213, 2024. URL: https://arxiv.org/abs/2405.12213.
  19. 19.MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation. arXiv preprint arXiv:2606.17598, 2026. URL: https://arxiv.org/abs/2606.17598.
  20. 20.L. X. Shi et al. Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models. arXiv preprint arXiv:2502.19417, 2025. URL: https://arxiv.org/abs/2502.19417.
  21. 21.GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning. arXiv preprint arXiv:2602.04315, 2026. URL: https://arxiv.org/abs/2602.04315.
  22. 22.Suneel Belkhale et al. RT-H: Action Hierarchies Using Language. arXiv preprint arXiv:2403.01823, 2024. URL: https://arxiv.org/abs/2403.01823.
  23. 23.Google DeepMind et al. Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer. arXiv preprint arXiv:2510.03342, 2025. URL: https://arxiv.org/abs/2510.03342.
  24. 24.Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning. arXiv preprint arXiv:2506.01953, 2025. URL: https://arxiv.org/abs/2506.01953.
  25. 25.DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model. arXiv preprint arXiv:2606.12105, 2026. URL: https://arxiv.org/abs/2606.12105.
  26. 26.Y. Du et al. Learning Universal Policies via Text-Guided Video Generation. arXiv preprint arXiv:2302.00111, 2023. URL: https://arxiv.org/abs/2302.00111.
  27. 27.WorldVLA: Towards Autoregressive Action World Model. arXiv preprint arXiv:2506.21539, 2025. URL: https://arxiv.org/abs/2506.21539.
  28. 28.World Action Models are Zero-shot Policies. arXiv preprint arXiv:2602.15922, 2026. URL: https://arxiv.org/abs/2602.15922.
  29. 29.Fast-WAM: Do World Action Models Need Test-time Future Imagination? arXiv preprint arXiv:2603.16666, 2026. URL: https://arxiv.org/abs/2603.16666.
  30. 30.Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning. arXiv preprint arXiv:2601.16163, 2026. URL: https://arxiv.org/abs/2601.16163.
  31. 31.Unified Video Action Model. arXiv preprint arXiv:2503.00200, 2025. URL: https://arxiv.org/abs/2503.00200.
  32. 32.LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies. arXiv preprint arXiv:2606.15768, 2026. URL: https://arxiv.org/abs/2606.15768.
  33. 33.Being-H0.7: A Latent World-Action Model from Egocentric Videos. arXiv preprint arXiv:2605.00078, 2026. URL: https://arxiv.org/abs/2605.00078.
  34. 34.Hao Wu, Xuejin Tian, Minghao Li, Yunxin Liu, Ganesh Ananthanarayanan, Fengyuan Xu, and Sheng Zhong. Pecam: Privacy-enhanced video streaming and analytics via securely-reversible transformation. In Proceedings of the 27th Annual International Conference on Mobile Computing and Networking, pages 229–241, 2021.
  35. 35.Hao Wu, Jinghao Feng, Xuejin Tian, Edward Sun, Yunxin Liu, Bo Dong, Fengyuan Xu, and Sheng Zhong. Emo: Real-time emotion recognition from single-eye images for resource-constrained eyewear devices. In Proceedings of the 18th International Conference on Mobile Systems, Applications, and Services, pages 448–461, 2020.
  36. 36.Fei Zeng, Feng Lyu, Hao Wu, Zhanxi Li, Shucheng Li, Fengyuan Xu, and Yaoxue Zhang. H2o: Heterogeneity-aware hierarchical orchestration for memory-efficient on-device llm inference. IEEE Transactions on Mobile Computing, 2025.
  37. 37.Borui Li, Tianen Liu, Weilong Wang, Chengqing Zhao, and Shuai Wang. Agent-as-a-service: An ai-native edge computing framework for 6g networks. IEEE Network, 39(2):44–51, 2024.
  38. 38.Borui Li, Tiange Xia, and Shuai Wang. Infscaler: Enabling efficient ml inference serving on multi-accelerator edge devices via asymmetric auto-scaling. In 2025 62nd ACM/IEEE Design Automation Conference (DAC), pages 1–7. IEEE, 2025.
  39. 39.Borui Li, Yitao Wang, Haoran Ma, Ligeng Chen, Jun Xiao, and Shuai Wang. Mobilora: Accelerating lora-based llm inference on mobile devices via context-aware kv cache optimization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 23400–23410, 2025.
  40. 40.K. D. Nguyen, H. T. Ho, C. T. Nguyen, T. Q. Duong, L. D. Le, D. M. H. Nguyen, V. A. Ngo, and A. T. Le. vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models. arXiv preprint arXiv:2606.08094, 2026. Submitted June 6, 2026. URL: https://arxiv.org/abs/2606.08094.
  41. 41.Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao, Yuqing Yang, and Yunxin Liu. Nn-meter: Towards accurate latency prediction of deep-learning model inference on diverse edge devices. In Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services, pages 81–93, 2021. doi:10.1145/3458864.3467882.
  42. 42.Fucheng Jia, Deyu Zhang, Ting Cao, Shiqi Jiang, Yunxin Liu, Ju Ren, and Yaoxue Zhang. Codl: Efficient cpu-gpu co-execution for deep learning inference on mobile devices. In Proceedings of the 20th Annual International Conference on Mobile Systems, Applications and Services, pages 209–221, 2022. doi:10.1145/3498361.3538932.
  43. 43.Joo Seong Jeong, Jingyu Lee, Donghyun Kim, Changmin Jeon, Changjin Jeong, Youngki Lee, and Byung-Gon Chun. Band: Coordinated multi-dnn inference on heterogeneous mobile processors. In Proceedings of the 20th Annual International Conference on Mobile Systems, Applications and Services, pages 235–247, 2022. doi:10.1145/3498361.3538948.
  44. 44.Chenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen, Bingsheng He, and Xiaoyong Du. Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices. In 2023 IEEE 39th International Conference on Data Engineering (ICDE), pages 1193–1207, 2023. doi:10.1109/ICDE55515.2023.00096.
  45. 45.Yakun Huang, Xiuquan Qiao, Schahram Dustdar, and Yan Li. Aodnn: An auto-offloading approach to optimize deep inference for fostering mobile web. In IEEE INFOCOM 2022 - IEEE Conference on Computer Communications, pages 2198–2207, 2022. doi:10.1109/INFOCOM48880.2022.9796763.
  46. 46.Kyungmin Bin, Jongseok Park, Chanjeong Park, Seyeon Kim, and Kyunghan Lee. Coacto: Coactive neural network inference offloading with fine-grained and concurrent execution. In Proceedings of the 22nd Annual International Conference on Mobile Systems, Applications and Services, pages 412–424, 2024. doi:10.1145/3643832.3661885.
  47. 47.Vinod Nigade, Pablo Bauszat, Henri Bal, and Lin Wang. Jellyfish: Timely inference serving for dynamic edge networks. In 2022 IEEE Real-Time Systems Symposium (RTSS), pages 277–290, 2022. doi:10.1109/RTSS55097.2022.00032.
  48. 48.Jun Xiao, Qinhui Gu, Ligeng Chen, Lizhi Sun, Zicheng Wang, Yinggang Guo, Lu Liu, Hao Wu, and Borui Li. Surviving the Impossible Trinity: Revisiting CPU Scheduling Problem on Modern COTS Mobile Devices (Operational Systems). pages 2353–2366. URL: https://www.usenix.org/conference/osdi26/presentation/xiao.
  49. 49.Daliang Xu, Qing Li, Mengwei Xu, Kang Huang, Gang Huang, Shangguang Wang, Xin Jin, Yun Ma, and Xuanzhe Liu. Niagara: Scheduling dnn inference services on heterogeneous edge processors. In Service-Oriented Computing, Lecture Notes in Computer Science, pages 67–85. Springer, 2023. doi:10.1007/978-3-031-48421-6_6.
  50. 50.Chulhong Min, Akhil Mathur, Utku Gunay Acer, Alessandro Montanari, and Fahim Kawsar. Sensix++: Bringing mlops and multi-tenant model serving to sensory edge devices. ACM Transactions on Embedded Computing Systems, 22(6):1–27, 2023. doi:10.1145/3617507.
  51. 51.Lixiang Han, Zimu Zhou, and Zhenjiang Li. Pantheon: Preemptible multi-dnn inference on mobile edge gpus. In Proceedings of the 22nd Annual International Conference on Mobile Systems, Applications and Services, pages 465–478, 2024. doi:10.1145/3643832.3661878.
  52. 52.Zhihe Zhao, Neiwen Ling, Nan Guan, and Guoliang Xing. Miriam: Exploiting elastic kernels for real-time multi-dnn inference on edge gpu. arXiv preprint arXiv:2307.04339, 2023. URL: https://arxiv.org/abs/2307.04339.
  53. 53.Xiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao, and Yunxin Liu. Flexnn: Efficient and adaptive dnn inference on memory-constrained edge devices. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, pages 709–723, 2024. doi:10.1145/3636534.3649391.
  54. 54.Kun Wang, Jiani Cao, Zimu Zhou, and Zhenjiang Li. Swapnet: Efficient swapping for dnn inference on edge ai devices beyond the memory budget. IEEE Transactions on Mobile Computing, 23(9):8935–8950, 2024. doi:10.1109/TMC.2024.3355764.
  55. 55.Shijie Zheng, Rong Chen, Ming Li, Zhe Ye, Luis Ceze, and Yun Liang. vmcu: Coordinated memory management and kernel optimization for dnn inference on mcus. In Proceedings of Machine Learning and Systems, 2024. URL: https://proceedings.mlsys.org/.
  56. 56.Rongjie Yi, Ting Cao, Ao Zhou, Xiao Ma, Shangguang Wang, and Mengwei Xu. Boosting dnn cold inference on edge devices. In Proceedings of the 21st Annual International Conference on Mobile Systems, Applications and Services, pages 516–529, 2023. doi:10.1145/3581791.3596842.
  57. 57.Keivan Alizadeh, Seyed Iman Mirzadeh, Dmitry Belenko, S. Khatamifard, Minsik Cho, Carlo C. Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. Llm in a flash: Efficient large language model inference with limited memory. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12562–12584, 2024. doi:10.18653/v1/2024.acl-long.678.
  58. 58.Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. Hetegen: Heterogeneous parallel inference for large language models on resource-constrained devices. In Proceedings of Machine Learning and Systems, 2024. URL: https://arxiv.org/abs/2403.01164.
  59. 59.Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu. Edgemoe: Empowering sparse large language models on mobile devices. arXiv preprint arXiv:2308.14352, 2023. URL: https://arxiv.org/abs/2308.14352.
  60. 60.Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. Fast on-device llm inference with npus. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, pages 445–462, 2025. doi:10.1145/3669940.3707239.
  61. 61.Zixu Hao, Huiqiang Jiang, Shiqi Jiang, Ju Ren, and Ting Cao. Hybrid slm and llm for edge-cloud collaborative inference. In Proceedings of the Workshop on Edge and Mobile Foundation Models, pages 36–41, 2024. doi:10.1145/3662006.3662067.
  62. 62.Mengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li, and Shangguang Wang. Fwdllm: Efficient federated finetuning of large language models with perturbed inferences. In 2024 USENIX Annual Technical Conference (USENIX ATC 24), pages 579–596. USENIX Association, 2024. URL: https://www.usenix.org/conference/atc24/presentation/xu-mengwei.
  63. 63.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Jasmine Hsu, Brian Ichter, Dmitry Kalashnikov, Sergey Levine, Igor Mordatch, Ofir Nachum, Carolina Parada, Kanishka Rao, Grecia Salazar, Pannag Sanketi, Vincent Vanhoucke, Quan Vuong, Fei Xia, et al. Rt-1: Robotics transformer for real-world control at scale. In Proceedings of Robotics: Science and Systems XIX, 2023. URL: https://arxiv.org/abs/2212.06817.
  64. 64.Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. Palm-e: An embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, 2023. URL: https://arxiv.org/abs/2303.03378.
  65. 65.Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. Vima: General robot manipulation with multimodal prompts. In Proceedings of the 40th International Conference on Machine Learning, 2023. URL: https://arxiv.org/abs/2210.03094.
  66. 66.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Jasmine Hsu, Brian Ichter, Eric Jang, Sergey Levine, Carolina Parada, Peter Pastor, Pierre Sermanet, Andy Zeng, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Proceedings of the 6th Conference on Robot Learning, 2022. URL: https://arxiv.org/abs/2204.01691.
  67. 67.Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In Proceedings of Robotics: Science and Systems XIX, 2023. URL: https://arxiv.org/abs/2304.13705.
  68. 68.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems XIX, 2023. URL: https://arxiv.org/abs/2303.04137.
  69. 69.Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. In International Conference on Learning Representations, 2024. URL: https://arxiv.org/abs/2312.13139.
  70. 70.Nur Muhammad Shafiullah, Zichen Cui, Ariuntuya Altanzaya, and Lerrel Pinto. Behavior transformers: Cloning k modes with one stone. In Advances in Neural Information Processing Systems 35, pages 22955–22968, 2022. doi:10.52202/068431-1668.
  71. 71.Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success. In Proceedings of Robotics: Science and Systems XXI, 2025. URL: https://arxiv.org/abs/2502.19645.
  72. 72.Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, Simon Alibert, Matthieu Cord, Thomas Wolf, and Remi Cadene. Smolvla: A vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv:2506.01844, 2025. URL: https://arxiv.org/abs/2506.01844.
  73. 73.Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, and Han Ding. Turbovla: Real-time vision-language-action model at 32 hz on an rtx 4090 with <1 gb vram. arXiv preprint arXiv:2607.27205, 2026. URL: https://arxiv.org/abs/2607.27205.
  74. 74.Arash Akbari, Arman Akbari, Masih Eskandar, Qitao Tan, Yixiao Chen, Jingwu Luo, Bertha Pangaribuan, Liyun Zhang, Jennifer Dy, Geng Yuan, Xue Lin, Gaowen Liu, Stratis Ioannidis, and Yanzhi Wang. Actquant: Sub-4-bit action-guided quantization for vision-language-action models. arXiv preprint arXiv:2605.24011, 2026. URL: https://arxiv.org/abs/2605.24011.
  75. 75.Xianghui Wang, Feng Chen, Wenbo Zhang, Hua Yan, Zixuan Wang, Changsheng Li, and Yinjie Lei. Policytrim: Boosting intrinsic policy efficiency of vision-language-action models. arXiv preprint arXiv:2606.22540, 2026. URL: https://arxiv.org/abs/2606.22540.
  76. 76.Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, and Chang Xu. Vla-cache: Efficient vision-language-action manipulation via adaptive token caching. In Advances in Neural Information Processing Systems 38, 2025. URL: https://arxiv.org/abs/2502.02175.
  77. 77.Ryuji Oi, Hikari Otsuka, Kosuke Matsushima, Yuki Ichikawa, Masato Motomura, Tatsuya Kaneko, and Daichi Fujiki. Actioncache: Training-free acceleration for vision-language-action models with action caching and refinement. arXiv preprint arXiv:2607.06370, 2026. URL: https://arxiv.org/abs/2607.06370.
  78. 78.Xiangyu Li, Huaizhi Tang, Xin Ding, Weijun Wang, Ting Cao, and Yunxin Liu. OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism. arXiv preprint arXiv:2603.14371, 2026. URL: https://arxiv.org/abs/2603.14371.
  79. 79.Yuanchun Guo and Bingyan Liu. Reflex: Real-time vision-language-action control through streaming inference. In Proceedings of the 43rd International Conference on Machine Learning, 2026. URL: https://arxiv.org/abs/2607.14695.
  80. 80.Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, and Meng Li. Jetson-pi: Towards onboard real-time robot control via foresight-aligned asynchronous inference. arXiv preprint arXiv:2607.12659, 2026. URL: https://arxiv.org/abs/2607.12659.
  81. 81.Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, and Yu Feng. Deltoris: Enabling real-time vla inference in embodied ai via bit-level sparsity and speculative inference. arXiv preprint arXiv:2608.04428, 2026. URL: https://arxiv.org/abs/2608.04428.
  82. 82.Rui Wang, Yue Zhang, Jiehong Lin, Kuncheng Luo, Jianan Wang, Zhongrui Wang, and Xiaojuan Qi. When to trust imagination: Adaptive action execution for world action models. arXiv preprint arXiv:2605.06222, 2026. URL: https://arxiv.org/abs/2605.06222.
  83. 83.Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, and Yingxue Zhang. Faster-wam: Do world action models need deep action modules? arXiv preprint arXiv:2608.02365, 2026. URL: https://arxiv.org/abs/2608.02365.
  84. 84.Sizhe Yang, Juncheng Mu, Tianming Wei, Chenhao Lu, Xiaofan Li, Linning Xu, Zhengrong Xue, Zhecheng Yuan, Dahua Lin, Jiangmiao Pang, and Huazhe Xu. Memorywam: Efficient world action modeling with persistent memory. arXiv preprint arXiv:2606.20562, 2026. URL: https://arxiv.org/abs/2606.20562.
  85. 85.Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You, Fugui Fan, Yuting Wu, Minghui Wu, Chenxu Zhao, JiaHong Ning, and Peiguang Jing. Lila-wam: Lightweight latent reasoning world-action model for robotic manipulation. arXiv preprint arXiv:2608.03701, 2026. URL: https://arxiv.org/abs/2608.03701.
  86. 86.Motubrain Team. World action models in real time: An empirical study of smooth execution via asynchronous deployment. arXiv preprint arXiv:2608.01880, 2026. URL: https://arxiv.org/abs/2608.01880.
  87. 87.Sining Ang, Yuguang Yang, and Yan Wang. Adaptive-wam: Quality-guided early-exit planning from intermediate video-diffusion features. arXiv preprint arXiv:2608.06008, 2026. URL: https://arxiv.org/abs/2608.06008.
  88. 88.I. Gim, Z. Ma, S.-S. Lee, and L. Zhong. Pie: A Programmable Serving System for Emerging LLM Applications. In Proceedings of the 31st ACM Symposium on Operating Systems Principles (SOSP), 2025. URL: https://doi.org/10.1145/3731569.3764814.
  89. 89.L. Su. Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving. arXiv preprint arXiv:2606.20537, 2026. URL: https://arxiv.org/abs/2606.20537.

Citation

MLA
Xu, L., et al. “Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots”. arXiv, 2026, http://arxiv.org/abs/2607.02501v3.
APA
Xu, L., Li, B., Wu, H., Han, C., Li, X., Hua, M., Jiang, S., Cao, T., Li, C., Zhong, S., & Wang, S. (2026). Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots. arXiv. http://arxiv.org/abs/2607.02501v3
Chicago
Xu, L., B. Li, H. Wu, et al. 2026. “Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots”. arXiv. http://arxiv.org/abs/2607.02501v3.
Harvard
Xu, L. et al. (2026) “Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.02501v3.
Vancouver
1. Xu L, Li B, Wu H, et al (2026) Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots. arXiv

BibTeX

@article{xu2026embodied,
  title = {Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots},
  author = {Xu, Ling and Li, Borui and Wu, Hao and Han, Chuyu and Li, Xiangyu and Hua, Mohan and Jiang, Shiqi and Cao, Ting and Li, Chuanyou and Zhong, Sheng and Wang, Shuai},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.02501v3},
  eprint = {2607.02501}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/