CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

Max FuJustin YuKarim El-RefaiEthan KouHaoru XueHuang HuangWenli XiaoGuanzhi WangFei-Fei LiGuanya Shi

article2026arXiv85 citations

Introduces CaP-X, an open-access evaluation and execution framework for robot manipulation that shows how scaling test-time compute and verifiable reinforcement learning enables code-generating agents to achieve human-level control on real robots without relying on manual scaffolding.

Listen

Autonomous robot manipulation has traditionally relied either on hand-crafted programs that lack adaptability or on end-to-end vision-language-action models that require costly training data and struggle to generalize. The Code-as-Policy paradigm offers a flexible alternative by using foundation models to write executable control code. However, earlier implementations depended heavily on human-designed software macros, obscuring whether task success was driven by model capability or engineered scaffolding.

The article introduces CaP-X, a framework designed to systematically benchmark and improve coding agents for robot manipulation across varied levels of abstraction, interaction modes, and sensory grounding. CaP-X integrates CaP-Gym (a suite of 187 simulation and real-world tasks), CaP-Bench (a standardized evaluation benchmark), CaP-Agent0 (an agentic, training-free execution pipeline), and CaP-RL (reinforcement learning directly applied to code generation).

The evaluation assessed 12 frontier language and vision-language models across seven core manipulation tasks using zero-shot pass rates against human expert solutions. It also tested extended long-horizon tasks and real-world robotic deployments. Key findings include: (1) In zero-shot single-turn settings, state-of-the-art models significantly trail human expert performance (which achieved an 88.5% average success rate). (2) High-level abstractions boost success rates but restrict behavioral expressivity; removing these abstractions causes performance to degrade, exposing a heavy reliance on human scaffolding. (3) Providing raw visual frames directly to models during multi-turn generation degrades performance, whereas converting images into structured natural language via a Visual Differencing Module substantially improves success. (4) CaP-Agent0—which combines visual differencing, parallel model reasoning, and an auto-synthesized library of reusable skills—matches or outperforms human single-turn baselines on four out of seven core tasks without requiring task-specific training data. (5) Applying reinforcement learning to code generation (CaP-RL) dramatically improves compilation and physical execution, yielding robust zero-shot sim-to-real transfer on physical Franka Emika robots (e.g., 84% success on cube lifting and 76% on cube stacking).

These results demonstrate that runtime compute strategies—such as multi-turn self-correction, parallel candidate sampling, and textual grounding—can overcome the need for task-specific, hand-coded abstractions. This approach reduces the engineering costs and data bottlenecks associated with traditional robotic policies while enhancing safety, interpretability, and cross-embodiment portability through an inspectable Python interface.

Practitioners should prioritize evaluating embodied coding agents on primitive-level interfaces rather than over-engineered macros to ensure genuine reasoning capabilities. Organizations looking to deploy coding agents should adopt agentic structures that combine structured textual feedback loops, multi-model candidate ensembles, and persistent, self-synthesized skill libraries. Where training resources are available, reinforcement learning on privileged simulator states offers a viable pathway for improving zero-shot real-world deployment.

While programmatic control shows high reliability in reasoning-intensive, long-horizon tasks, it remains limited in continuous, contact-rich behaviors such as tight insertions. Future efforts should explore hybrid architectures that pair high-level programmatic coding agents with low-level visuomotor execution policies, alongside expanding benchmarks to stress-test collision avoidance and dynamic perception.

Cover for CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

Abstract

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.

Citation

MLA
Fu, L., et al. “CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation”. arXiv, 2026, http://arxiv.org/abs/2603.22435v2.
APA
Fu, L., Yu, J., El-Refai, K., Kou, E., Xue, H., Huang, H., Xiao, W., Wang, G., Niu, D., Li, F.-F., Shi, G., Wu, J., Sastry, S., Zhu, Y., Goldberg, K., & Fan, L. "Jim" . (2026). CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation. arXiv. http://arxiv.org/abs/2603.22435v2
Chicago
Fu, L., J. Yu, K. El-Refai, et al. 2026. “CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation”. arXiv. http://arxiv.org/abs/2603.22435v2.
Harvard
Fu, L. et al. (2026) “CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2603.22435v2.
Vancouver
1. Fu L, Yu J, El-Refai K, et al (2026) CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation. arXiv

BibTeX

@article{fu2026cap,
  title = {CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation},
  author = {Fu, Letian and Yu, Justin and El-Refai, Karim and Kou, Ethan and Xue, Haoru and Huang, Huang and Xiao, Wenli and Wang, Guanzhi and Niu, Dantong and Li, Fei-Fei and Shi, Guanya and Wu, Jiajun and Sastry, Shankar and Zhu, Yuke and Goldberg, Ken and Fan, Linxi "Jim"},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2603.22435v2},
  eprint = {2603.22435}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/