CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
Max FuJustin YuKarim El-RefaiEthan KouHaoru XueHuang HuangWenli XiaoGuanzhi WangFei-Fei LiGuanya Shi
Introduces CaP-X, an open-access evaluation and execution framework for robot manipulation that shows how scaling test-time compute and verifiable reinforcement learning enables code-generating agents to achieve human-level control on real robots without relying on manual scaffolding.
Autonomous robot manipulation has traditionally relied either on hand-crafted programs that lack adaptability or on end-to-end vision-language-action models that require costly training data and struggle to generalize. The Code-as-Policy paradigm offers a flexible alternative by using foundation models to write executable control code. However, earlier implementations depended heavily on human-designed software macros, obscuring whether task success was driven by model capability or engineered scaffolding.
The article introduces CaP-X, a framework designed to systematically benchmark and improve coding agents for robot manipulation across varied levels of abstraction, interaction modes, and sensory grounding. CaP-X integrates CaP-Gym (a suite of 187 simulation and real-world tasks), CaP-Bench (a standardized evaluation benchmark), CaP-Agent0 (an agentic, training-free execution pipeline), and CaP-RL (reinforcement learning directly applied to code generation).
The evaluation assessed 12 frontier language and vision-language models across seven core manipulation tasks using zero-shot pass rates against human expert solutions. It also tested extended long-horizon tasks and real-world robotic deployments. Key findings include: (1) In zero-shot single-turn settings, state-of-the-art models significantly trail human expert performance (which achieved an 88.5% average success rate). (2) High-level abstractions boost success rates but restrict behavioral expressivity; removing these abstractions causes performance to degrade, exposing a heavy reliance on human scaffolding. (3) Providing raw visual frames directly to models during multi-turn generation degrades performance, whereas converting images into structured natural language via a Visual Differencing Module substantially improves success. (4) CaP-Agent0—which combines visual differencing, parallel model reasoning, and an auto-synthesized library of reusable skills—matches or outperforms human single-turn baselines on four out of seven core tasks without requiring task-specific training data. (5) Applying reinforcement learning to code generation (CaP-RL) dramatically improves compilation and physical execution, yielding robust zero-shot sim-to-real transfer on physical Franka Emika robots (e.g., 84% success on cube lifting and 76% on cube stacking).
These results demonstrate that runtime compute strategies—such as multi-turn self-correction, parallel candidate sampling, and textual grounding—can overcome the need for task-specific, hand-coded abstractions. This approach reduces the engineering costs and data bottlenecks associated with traditional robotic policies while enhancing safety, interpretability, and cross-embodiment portability through an inspectable Python interface.
Practitioners should prioritize evaluating embodied coding agents on primitive-level interfaces rather than over-engineered macros to ensure genuine reasoning capabilities. Organizations looking to deploy coding agents should adopt agentic structures that combine structured textual feedback loops, multi-model candidate ensembles, and persistent, self-synthesized skill libraries. Where training resources are available, reinforcement learning on privileged simulator states offers a viable pathway for improving zero-shot real-world deployment.
While programmatic control shows high reliability in reasoning-intensive, long-horizon tasks, it remains limited in continuous, contact-rich behaviors such as tight insertions. Future efforts should explore hybrid architectures that pair high-level programmatic coding agents with low-level visuomotor execution policies, alongside expanding benchmarks to stress-test collision avoidance and dynamic perception.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Introduces the Code-as-Policies paradigm of prompting language models to generate executable robot control code, which CaP-X directly adopts, benchmarks, and systematically enhances.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Establishes closed-loop environment feedback and autonomous replanning for embodied language agents, providing the conceptual foundation for CaP-X's interactive test-time compute scaling.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). Pioneers the grounding of language models in robot affordances and perceptual primitives, which informs CaP-X's abstraction levels and primitive execution framework.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Demonstrates zero-shot translation of natural language instructions into actionable physical plans, serving as a core stepping stone toward program-based embodied controllers.
- Paper: Training Software Engineering Agents and Verifiers with SWE-Gym, Jiayi Pan et al. (2025). Develops execution-based gym environments and outcome verification for training coding agents, laying the methodological groundwork for CaP-Gym and CaP-RL.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). Provides a comprehensive survey that generalizes code-based agent harnesses and scaffolding mechanisms across robotics, software engineering, and general decision-making.
- Paper: The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior, Xiangzhe Xu et al. (2026). Further investigates how tool architectures and interface abstractions impact coding agent behavior and execution stability under controlled benchmark conditions.
