Code as Policies: Language Model Programs for Embodied Control
Jacky LiangWenlong HuangFei XiaPeng XuKarol HausmanBrian IchterPete FlorenceAndy Zeng
Proposes an approach using code-generation language models to convert natural language commands directly into executable robot policy code, enabling spatial reasoning, behavioral commonsense, and real-time reactive control on physical robots.
Enabling robots to interpret and act on natural language commands typically requires extensive, costly data collection and training for each specific skill. While large language models trained on massive text datasets can sequence high-level actions, they often struggle with spatial geometry, continuous feedback loops, and context-dependent adjustments such as moving faster or shifting slightly to the side. As robotic applications expand into dynamic human environments, developing flexible systems that bridge open-ended language and precise physical control without retraining has become a critical operational need.
The article evaluates whether code-writing large language models can be prompted to autonomously generate executable robot policy code directly from natural language commands. Specifically, the authors demonstrate an approach called Code as Policies, which translates user prompts into Python programs that process perception outputs and parameterize robot control interfaces.
To test this concept, the authors evaluated the system across standardized coding benchmarks, a new robotics-focused coding benchmark, simulated manipulation tasks, and multiple physical platforms, including tabletop robot arms and mobile kitchen assistants. The approach relies on few-shot prompting, providing the language model with a few demonstration examples alongside hints about available perception and movement interfaces. It also introduces hierarchical code generation, where the system automatically writes sub-functions to resolve undefined logic steps during program execution.
The findings show that generating code substantially improves robotic control and reasoning compared to existing methods. On a robotics-specific coding benchmark, hierarchical code generation achieved up to a 95 percent pass rate with large models, significantly outperforming flat generation. In simulated tabletop manipulation, the approach matched or exceeded supervised learning models trained on 30,000 demonstrations, achieving an overall success rate of 71 percent on entirely new instructions and object attributes where supervised methods failed completely. Furthermore, using code for spatial-geometric reasoning attained a 98 percent success rate, compared to 58 percent when using natural language reasoning. Across physical hardware, the robots successfully interpreted complex, multi-step instructions, performed geometric path drawing, engaged in conversational clarification, and demonstrated context-aware speed and position adjustments.
These results indicate that organizations can significantly reduce development costs and deployment timelines for robotic systems. By replacing expensive, specialized policy training with flexible code generation and off-the-shelf vision modules, engineering teams can implement adaptable robotic workflows using existing programming libraries. The resulting transparent Python policies also make system actions interpretable and straightforward to debug, enhancing operational safety.
Organizations evaluating this approach should begin by auditing their robot hardware interfaces to ensure perception and control routines can be cleanly called via standardized software functions. Decision-makers should consider pilot programs for structured manipulation or navigation tasks while developing rigorous safety wrappers to prevent invalid program execution. Further testing is necessary before deploying the system in safety-critical settings, as the approach remains limited by the capabilities of underlying vision sensors, cannot easily infer 3D spatial structures absent in prompt examples, and relies on the assumption that generated code will execute safely without pre-execution verification.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This foundational work introduces Codex and the HumanEval benchmark, establishing the code generation and synthesis capabilities that Code as Policies directly repurposes for embodied robotic control.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This paper demonstrates program synthesis from natural language via Transformer models, providing essential background on translating natural language descriptions into functional code.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). This work establishes the paradigm of using language models for robotic planning grounded in physical affordances (SayCan), directly motivating programmatic execution over discrete language skill selection.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). This paper introduces subproblem decomposition via prompt engineering, forming a core conceptual prerequisite for the recursive and hierarchical code-generation strategies used in Code as Policies.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This landmark paper establishes few-shot in-context learning in large language models, the foundational prompting technique underlying the synthesis of robot policy code.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E extends embodied LLM planning by directly integrating multimodal sensory and visual inputs into the model architecture, bypassing the need for separate perception API abstractions.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 transitions from generating high-level programmatic policies over predefined primitives to directly emitting low-level robotic action tokens via end-to-end vision-language-action modeling.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). This survey generalizes the concept of using executable code as the core substrate for robotic control, reasoning, and environment interaction pioneered in Code as Policies.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). This paper advances embodied control beyond discrete programmatic APIs by using flow matching to produce continuous, dexterous robotic actions directly from language and vision.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA builds upon language-guided robotic manipulation by offering an open generalist vision-language-action model trained across diverse robotic embodiments.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). ToolLLM scales the principle of LLMs autonomously composing and executing APIs to thousands of complex real-world software tools and structured decision trees.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). Reflexion builds on language-driven execution and coding by adding verbal self-reflection and episodic memory to improve agent policy execution across interactive trials.
