Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Ling XuChuyu HanBorui LiHao WuShiqi JiangTing CaoChuanyou LiShenghui ZhongShuai Wang
Presents Embodied.cpp, a portable C++ inference runtime that unifies vision-language-action and world-action model deployment across heterogeneous robot hardware, delivering up to 2.7x speedups and 77% lower memory consumption for real-time closed-loop control.
Practical deployment of artificial intelligence models on physical robots faces a major systems challenge. Although vision-language-action models and world-action models are advancing rapidly, deploying them on heterogeneous and resource-constrained edge hardware remains highly fragmented. Current AI serving engines are built for cloud-based request-response tasks that optimize overall throughput. In contrast, physical robotics requires low-latency, batch-size-one closed-loop control, coordination across modules operating at different execution rates, and flexible interfaces that accommodate diverse sensor inputs and physical action outputs.
The article introduces and evaluates Embodied.cpp, an open-source, portable C++ inference runtime designed to standardize and accelerate the execution of diverse embodied AI models across heterogeneous robotic platforms and simulators.
To address deployment inefficiencies without hard-coding specific model structures, the authors developed a five-layer runtime architecture: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. This framework decouples execution schedules, optimizes small-batch computation for edge processors, and provides pluggable operator interfaces. The runtime was evaluated against standard Python deployment baselines across three vision-language-action models (pi0.5, GR00T N1.7, and HY-VLA) across multiple precision configurations, as well as two world-action models (Cosmos3 and LingBot-VA).
The evaluation revealed several key findings. First, Embodied.cpp delivered overall inference speedups ranging from 1.05-fold to 2.70-fold across evaluated configurations compared to Python baselines, achieving significant latency reductions per generated action step. Second, the runtime dramatically lowered onboard memory demands, decreasing video RAM usage by 7% to 77% depending on the model and quantization level (including 8-bit, 6-bit, and 4-bit configurations). Third, for world-action models, memory footprints dropped substantially—such as a 33.6% reduction for LingBot-VA from 24.75 GB down to 16.44 GB—while keeping operational success nearly identical to baseline levels (98.00% versus 100.00%). Across nearly all vision-language-action tests, the C++ implementation preserved high control task success rates.
These findings demonstrate that organizations can replace fragile, bespoke Python deployment glue code with a unified, lightweight runtime. By significantly lowering memory consumption and operational latency without degrading task success, the runtime reduces onboard hardware costs, shortens development cycles, and enables advanced foundation models to operate on power- and compute-constrained edge robotics devices.
Engineering and robotics teams deploying physical AI systems should evaluate adopting a unified C++ runtime architecture to streamline model integration. Organizations can leverage lower-bit quantization paths to deploy larger models on smaller edge processors, balancing minor trade-offs in success rate against hardware savings. Further work should focus on expanding testing across more real-world robotic environments, validating additional custom hardware accelerators, and testing next-generation multi-component models.
While the reported results show high confidence in latency and memory gains across tested benchmarks and simulators, users should exercise caution regarding boundary conditions. Real-world mechanical uncertainties, sensor noise, and highly aggressive quantization (such as 4-bit on certain architectures) may affect control quality depending on the specific robotic task.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA provides a concrete open VLA architecture and inference setting that helps clarify the model components Embodied.cpp aims to run across robot platforms.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo’s modular encoders and action-decoding head offer a useful earlier example of the heterogeneous model interfaces that Embodied.cpp organizes into a shared runtime path.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This survey maps VLA architectures and deployment challenges, giving context for Embodied.cpp’s architectural analysis and its focus on runtime fragmentation.
- Paper: Kairos: A Scalable Serving System for Physical AI, Yinwei Dai et al. (2026). Kairos carries embodied inference into multi-robot serving, adding fleet scheduling and awareness of physical execution to the deployment-runtime concerns Embodied.cpp addresses locally.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA extends latency-aware embodied execution with continuous inference and action streaming that keep control aligned with moving objects.
