Code as Policies: Language Model Programs for Embodied Control

Jacky LiangWenlong HuangFei XiaPeng XuKarol HausmanBrian IchterPete FlorenceAndy Zeng

article2022IEEE International Conference on Robotics and Automation1,832 citationsOutstanding Paper Award in Robot Learning

Proposes an approach using code-generation language models to convert natural language commands directly into executable robot policy code, enabling spatial reasoning, behavioral commonsense, and real-time reactive control on physical robots.

Listen

Enabling robots to interpret and act on natural language commands typically requires extensive, costly data collection and training for each specific skill. While large language models trained on massive text datasets can sequence high-level actions, they often struggle with spatial geometry, continuous feedback loops, and context-dependent adjustments such as moving faster or shifting slightly to the side. As robotic applications expand into dynamic human environments, developing flexible systems that bridge open-ended language and precise physical control without retraining has become a critical operational need.

The article evaluates whether code-writing large language models can be prompted to autonomously generate executable robot policy code directly from natural language commands. Specifically, the authors demonstrate an approach called Code as Policies, which translates user prompts into Python programs that process perception outputs and parameterize robot control interfaces.

To test this concept, the authors evaluated the system across standardized coding benchmarks, a new robotics-focused coding benchmark, simulated manipulation tasks, and multiple physical platforms, including tabletop robot arms and mobile kitchen assistants. The approach relies on few-shot prompting, providing the language model with a few demonstration examples alongside hints about available perception and movement interfaces. It also introduces hierarchical code generation, where the system automatically writes sub-functions to resolve undefined logic steps during program execution.

The findings show that generating code substantially improves robotic control and reasoning compared to existing methods. On a robotics-specific coding benchmark, hierarchical code generation achieved up to a 95 percent pass rate with large models, significantly outperforming flat generation. In simulated tabletop manipulation, the approach matched or exceeded supervised learning models trained on 30,000 demonstrations, achieving an overall success rate of 71 percent on entirely new instructions and object attributes where supervised methods failed completely. Furthermore, using code for spatial-geometric reasoning attained a 98 percent success rate, compared to 58 percent when using natural language reasoning. Across physical hardware, the robots successfully interpreted complex, multi-step instructions, performed geometric path drawing, engaged in conversational clarification, and demonstrated context-aware speed and position adjustments.

These results indicate that organizations can significantly reduce development costs and deployment timelines for robotic systems. By replacing expensive, specialized policy training with flexible code generation and off-the-shelf vision modules, engineering teams can implement adaptable robotic workflows using existing programming libraries. The resulting transparent Python policies also make system actions interpretable and straightforward to debug, enhancing operational safety.

Organizations evaluating this approach should begin by auditing their robot hardware interfaces to ensure perception and control routines can be cleanly called via standardized software functions. Decision-makers should consider pilot programs for structured manipulation or navigation tasks while developing rigorous safety wrappers to prevent invalid program execution. Further testing is necessary before deploying the system in safety-critical settings, as the approach remains limited by the capabilities of underlying vision sensors, cannot easily infer 3D spatial structures absent in prompt examples, and relies on the assumption that generated code will execute safely without pre-execution verification.

arXiv: 2209.07753
  • Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This foundational work introduces Codex and the HumanEval benchmark, establishing the code generation and synthesis capabilities that Code as Policies directly repurposes for embodied robotic control.
  • Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This paper demonstrates program synthesis from natural language via Transformer models, providing essential background on translating natural language descriptions into functional code.
  • Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). This work establishes the paradigm of using language models for robotic planning grounded in physical affordances (SayCan), directly motivating programmatic execution over discrete language skill selection.
  • Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). This paper introduces subproblem decomposition via prompt engineering, forming a core conceptual prerequisite for the recursive and hierarchical code-generation strategies used in Code as Policies.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This landmark paper establishes few-shot in-context learning in large language models, the foundational prompting technique underlying the synthesis of robot policy code.
Cover for Code as Policies: Language Model Programs for Embodied Control

Abstract

Large language models (LLMs) trained on code completion have been shown to be capable of synthesizing simple Python programs from docstrings [1]. We find that these code-writing LLMs can be re-purposed to write robot policy code, given natural language commands. Specifically, policy code can express functions or feedback loops that process perception outputs (e.g.,from object detectors [2], [3]) and parameterize control primitive APIs. When provided as input several example language commands (formatted as comments) followed by corresponding policy code (via few-shot prompting), LLMs can take in new commands and autonomously re-compose API calls to generate new policy code respectively. By chaining classic logic structures and referencing third-party libraries (e.g., NumPy, Shapely) to perform arithmetic, LLMs used in this way can write robot policies that (i) exhibit spatial-geometric reasoning, (ii) generalize to new instructions, and (iii) prescribe precise values (e.g., velocities) to ambiguous descriptions ("faster") depending on context (i.e., behavioral commonsense). This paper presents code as policies: a robot-centric formulation of language model generated programs (LMPs) that can represent reactive policies (e.g., impedance controllers), as well as waypoint-based policies (vision-based pick and place, trajectory-based control), demonstrated across multiple real robot platforms. Central to our approach is prompting hierarchical code-gen (recursively defining undefined functions), which can write more complex code and also improves state-of-the-art to solve 39.8% of problems on the HumanEval [1] benchmark. Code and videos are available at this https URL

Table of Contents

  • I Introduction
  • II Related Work
  • III Method
  • III-A Prompting Language Model Programs
  • III-B Example Language Model Programs (Low-Level)
  • III-C Example Language Model Programs (High-Level)
  • III-D Language Model Programs as Policies
  • IV Experiments
  • IV-A Hierarchical LMPs on Code-Generation Benchmarks
  • IV-B CaP: Drawing Shapes via Generated Waypoints
  • IV-C CaP: Pick & Place Policies for Table-Top Manipulation
  • IV-D CaP: Table-Top Manipulation Simulation Evaluations
  • IV-E CaP: Mobile Robot Navigation and Manipulation
  • V Discussion and Limitations
  • References
  • -A Prompt Engineering
  • -B Method Section Prompts
  • -B1 Language-based reasoning
  • -B2 First-party
  • -B3 Combining language reasoning, third-party, and first-party libraries.
  • -B4 LMPs can be composed.
  • -B5 parse_obj prompt.
  • -C Reasoning with Code vs. Natural Language
  • -D CodeGen HumanEval Additional Results
  • -E Robot Code-Generation Benchmark
  • -E1 Example Questions
  • -E2 Generalization Analysis
  • -F CaP: Reactive Controllers for Toy Tasks
  • -G Visual Language Models
  • -H Whiteboard Drawing
  • -I Real-World Tabletop Manipulation
  • -J Mobile Robot
  • -K Simulation Tabletop Manipulation Evaluations
  • -L Additional LLM Capabilities
  • -M Cross Embodiment Example

Knowls

  1. Knowl 1 — Code as Policies (CaP) Formulation for Embodied Control

    model/method

    Code as Policies (CaP) formulates robot policy generation as few-shot code synthesis using Large Language Models (LLMs) trained on code completion (such as OpenAI Codex code-davinci-002). Given natural language instructions formatted as code comments, the LLM autoregressively generates executable Python policy scripts, termed Language Model Programs (LMPs).

    An LMP takes perceptual inputs (from perception APIs such as open-vocabulary object detectors) and parameterizes low-level robot control APIs (such as waypoint followers, velocity controllers, or pick-and-place primitives). The prompts provided to the LLM comprise:

    1. Hints: Import statements and type hints that define the available first-party perception and control APIs (e.g., from utils import get_pos, put_first_on_second) and third-party libraries (e.g., NumPy, Shapely).
    2. Examples: Few-shot input-output pairs mapping natural language comments to corresponding Python code blocks demonstrating logic, arithmetic, and API invocations.

    By leveraging standard programming constructs (conditionals, for/while loops) and mathematical libraries, CaP enables spatial-geometric reasoning, multi-step sequencing, and runtime feedback loops without needing end-to-end task demonstration data or policy fine-tuning.

  2. Knowl 2 — Hierarchical Function Generation for Language Model Programs

    algorithm

    Hierarchical code generation decomposes complex robot instructions into modular functions by allowing generated code to call functions that are not yet implemented. When an LMP generates code referencing undefined functions, a secondary, specialized function-generating LMP is invoked recursively to implement them in a depth-first manner.

    Input: Instruction string II, prompt template PmainP_{\text{main}}, function-generation prompt PfgenP_{\text{fgen}}, execution scope S=(globals,locals)S = (\text{globals}, \text{locals})
    Output: Populated scope SS containing executed policy and generated helper functions
    function GenerateAndExecute(II, PmainP_{\text{main}}, SS):
        C←QueryLLM(Pmain,I)C \leftarrow \text{QueryLLM}(P_{\text{main}}, I)
        ResolveUndefinedFunctions(CC, PfgenP_{\text{fgen}}, SS)
        ExecuteSafe(CC, SS)
        return SS
    function ResolveUndefinedFunctions(CC, PfgenP_{\text{fgen}}, SS):
        T←ParseAbstractSyntaxTree(C)T \leftarrow \text{ParseAbstractSyntaxTree}(C)
        U←FindUndefinedFunctionCalls(T,S)U \leftarrow \text{FindUndefinedFunctionCalls}(T, S)
        for each function call f(x1,…,xk)∈Uf(x_1, \dots, x_k) \in U do:
            If←FormatSignatureComment(f,x1,…,xk)I_f \leftarrow \text{FormatSignatureComment}(f, x_1, \dots, x_k)
            Cf←QueryLLM(Pfgen,If)C_f \leftarrow \text{QueryLLM}(P_{\text{fgen}}, I_f)
            ResolveUndefinedFunctions(CfC_f, PfgenP_{\text{fgen}}, SS)
            ExecuteSafe(CfC_f, SS)

    By keeping high-level logic modular and generating undefined helper subroutines dynamically, hierarchical generation avoids flattening complex logic into single prompts, fitting within model context windows and improving code generation accuracy.

  3. Knowl 3 — Safety Verification and Scope Sandboxing for LMP Execution

    model/method

    To safely execute untrusted Python code generated by LLMs on physical and simulated robotic platforms, CaP performs static verification followed by scoped dynamic execution via Python's built-in exec function.

    1. Static Safety Filtering: Before execution, the generated code string is checked against forbidden constructs. The program is rejected if it contains:
      • Python import statements (all allowed libraries must be pre-imported into the execution environment),
      • Special variable names or dunder attributes beginning with double underscores (e.g., __class__, __subclasses__),
      • Direct invocations of exec() or eval().
    2. Scoped Namespace Execution: If safety checks pass, execution proceeds via: exec(code_str,globals_dict,locals_dict)\text{exec}(\text{code\_str}, \text{globals\_dict}, \text{locals\_dict}) where globals_dict\text{globals\_dict} is populated strictly with authorized perception, motion planning, and robot control APIs, and locals_dict\text{locals\_dict} begins as an empty dictionary. During execution, any new variables or helper functions instantiated by the generated code are stored in locals_dict\text{locals\_dict}. If the LMP is designed to return a value (e.g., an object name or coordinates), the result is read directly from locals_dict['ret_val'].
  4. Knowl 4 — Spatial-Geometric Reasoning Comparison: Code vs. Natural Language and Chain-of-Thought

    data/table

    A benchmark comprising 28 object selection tasks (e.g., identifying the object closest to a target) and 23 position selection tasks (e.g., interpolating coordinates between objects, considered correct if within 1 cm of ground truth) was used to evaluate reasoning formats using OpenAI Codex code-davinci-002.

    Task Vanilla NL (%) Chain of Thought (CoT) (%) LMP / Code (Ours) (%)
    Object Selection 39 68 96
    Position Selection 30 48 100
    Total 35 58 98

    While Chain-of-Thought prompting improves qualitative relation reasoning over direct (Vanilla) natural language answers, it fails on precise multi-step arithmetic. Generating Python code enables the LLM to delegate geometric calculations and vector arithmetic directly to external libraries like NumPy, yielding near-perfect accuracy.

  5. Knowl 5 — Code Generation Benchmarks: RoboCodeGen and HumanEval Pass Rates

    data/table

    Hierarchical code generation was evaluated against flat code generation on RoboCodeGen (a 37-problem benchmark spanning robotics spatial reasoning, geometric bounds, and control routines) across four model sizes, as well as on the standard HumanEval generic code synthesis benchmark.

    Method GPT-3 6.7B InstructGPT 175B Codex cushman Codex davinci
    RoboCodeGen Pass Rate (% P@1)
    Flat Code-Gen 3 68 54 81
    Hierarchical Code-Gen 5 84 57 95
    HumanEval Configuration Greedy (T=0) (%) P@1 (T=0.8) (%) P@10 (T=0.8) (%) P@100 (T=0.8) (%)
    Flat CodeGen + No Prompt 45.7 34.9 75.1 90.9
    Flat CodeGen + Flat Prompts 50.6 36.6 77.6 93.3
    Hierarchical CodeGen + Hier Prompts 53.0 39.8 80.6 95.7

    Hierarchical decomposition consistently outperforms flat prompting across both robotics-specific and general software tasks, with gains scaling positively with model parameter size.

  6. Knowl 6 — Simulated Tabletop Manipulation Performance Across Task and Attribute Generalization

    data/table

    In a simulated tabletop environment with a UR5e manipulator, 10 colored blocks, and 10 colored bowls, methods were evaluated across 8 Long-Horizon rearrangement tasks and 6 Spatial-Geometric reasoning tasks. Tasks were split into combinations of Seen/Unseen Attributes (SA/UA, e.g., novel colors/positions) and Seen/Unseen Instructions (SI/UI, e.g., unseen linguistic command templates). Success rates (%) across 50 trials per task are reported.

    Split Task Family CLIPort (Supervised) NL Planner CaP (Ours)
    SA SI Long-Horizon 78.80 86.40 97.20
    SA SI Spatial-Geometric 97.33 N/A 89.30
    UA SI Long-Horizon 36.80 88.00 97.60
    UA SI Spatial-Geometric 0.00 N/A 73.33
    UA UI Long-Horizon 0.00 64.00 80.00
    UA UI Spatial-Geometric 0.01 N/A 62.00

    End-to-end supervised imitation learning (CLIPort trained on 30k demonstrations) collapses when encountering unseen attributes (UA) or instructions (UI). CaP retains high performance across both dimensions, while natural language planners fail to address fine-grained spatial-geometric reasoning tasks.

  7. Knowl 7 — Generalization Characteristics in Robotics Code Synthesis

    empirical result

    Evaluating code-generation models on robotics problems categorized under five standard forms of compositional generalization reveals distinct model behaviors:

    1. Substitutivity (replacing keywords with synonyms or same-category entities): All models (even smaller models like code-cushman-001) achieve high success rates (>90%>90\%) in both flat and hierarchical settings, demonstrating robust semantic equivalence handling.
    2. Productivity (generating longer code or deeper hierarchical function structures than shown in prompt examples): Flat code generation yields low productivity success (<45%<45\% for cushman and <60%<60\% for davinci). Hierarchical code generation improves Productivity by nearly +30%+30\% in text-davinci-002 and code-davinci-002 by breaking down nested logic.
    3. Localism, Systematicity, and Overgeneralization: Hierarchical generation improves Systematicity and Localism for davinci-tier models (>80%>80\% success), but offers negligible or negative improvements on the smaller cushman model, indicating that an underlying capacity threshold in the base LLM is required before hierarchical prompt decomposition becomes beneficial.
  8. Knowl 8 — Zero-Shot Synthesis of Low-Level Reactive and Impedance Controllers

    model/method

    Language models prompted with appropriate mathematical hints and function signatures can generate functional low-level continuous and discrete feedback control policies zero-shot.

    1. Discrete Proportional-Derivative (PD) Inverted Pendulum Balancing: Given state variables cart position xx, velocity x˙\dot{x}, pole angle θ\theta, and angular velocity θ˙\dot{\theta}, the synthesized LMP generates discrete control direction u∈{0,1}u \in \{0, 1\}: control=kpθ+kdθ˙\text{control} = k_p \theta + k_d \dot{\theta} u={1if control≥00if control<0u = \begin{cases} 1 & \text{if } \text{control} \ge 0 \\ 0 & \text{if } \text{control} < 0 \end{cases} where kp=1k_p = 1 and kd=1k_d = 1.
    2. Task-Space End-Effector Impedance Control: To command a UR5e manipulator in PyBullet, the LMP computes joint torques τ\tau from Cartesian errors using the manipulator Jacobian JJ: xerr=xgoal−xcurr,x˙err=−x˙currx_{\text{err}} = x_{\text{goal}} - x_{\text{curr}}, \quad \dot{x}_{\text{err}} = -\dot{x}_{\text{curr}} τ=JT(Kxxerr+Dxx˙err)\tau = J^T \left( K_x x_{\text{err}} + D_x \dot{x}_{\text{err}} \right) where KxK_x is the Cartesian stiffness matrix and DxD_x is the damping matrix.
  9. Knowl 9 — Perception-Action Grounding Across Physical Robot Embodiments

    experimental setup

    CaP grounds high-level code execution in real-world physical systems across three distinct robotic embodiments:

    1. 2D Whiteboard Drawing (UR5e): Equipped with a rigidly mounted dry-erase marker. Perception uses the MDETR open-vocabulary detector. Action APIs include waypoint trajectory execution (draw(pts_2d)) with contact detection and erasing (erase(pts_2d)).
    2. Tabletop Pick-and-Place (UR5e): Equipped with a suction gripper and wrist-mounted Intel RealSense D435 RGB-D camera. Object bounding boxes and segmentation masks are provided by MDETR and deprojected into 3D robot base coordinates. The primary action primitive is put_first_on_second(obj_name, target).
    3. Mobile Manipulation (Everyday Robots): A mobile base with a 7-DoF arm and head-mounted RGB-D camera. Perception is driven by the ViLD open-vocabulary object detector. Action APIs provide navigation (goto_pos, goto_loc), picking (pick_obj), placing (place_at_pos, place_at_obj), and variable assignment in the Python scope to maintain short-term state memory across sequential actions.
  10. Knowl 10 — Limitations of Code as Policies

    limitation

    The Code as Policies framework is constrained by several fundamental factors:

    1. Perception API Expressiveness: Policies are limited by what underlying vision-language models can reliably output (e.g., discrete object bounding boxes or masks). Current VLMs cannot provide nuanced feedback on continuous trajectory shapes (e.g., assessing whether a motion was 'bumpy' or 'C-shaped').
    2. Control Primitive Saturation: Control capabilities depend on the predefined action primitives. Only a small set of named arguments can be exposed in prompts before saturating the LLM's context window.
    3. Abstraction Gaps: LMPs cannot extrapolate to instructions operating at vastly different abstraction levels or structural complexity than those shown in prompt examples (e.g., synthesizing 3D structure assembly from simple pick-and-place examples).
    4. Feasibility and Correctness Guarantees: CaP assumes all natural language commands are physically and kinematically feasible; the system has no mechanism to verify policy safety or geometric reachability prior to runtime execution.
    5. Cross-Embodiment Brittleness: Adapting plans across substantially differing robot kinematics and action spaces solely via prompt hints remains brittle with current code-generation LLMs.

Coverage note — None was omitted; all key contributions—including the LMP formulation, hierarchical code generation algorithm, safety sandboxing, benchmark evaluations (RoboCodeGen, HumanEval, simulation tabletop), generalization analysis, zero-shot reactive controller synthesis, physical embodiments, and limitations—are represented.

References

  1. 1.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman et al., “Evaluating large language models trained on code,” arXiv:2107.03374, 2021.
  2. 2.A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in ICCV, 2021.
  3. 3.X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” arXiv:2104.13921, 2021.
  4. 4.S. Tellex, N. Gopalan, H. Kress-Gazit, and C. Matuszek, “Robots that use language,” Review of Control, Robotics, and Autonomous Systems, 2020.
  5. 5.T. Winograd, “Procedures as a representation for data in a computer program for understanding natural language,” MIT PROJECT MAC, 1971.
  6. 6.J. Dzifcak, M. Scheutz, C. Baral, and P. Schermerhorn, “What to do and how to do it: Translating natural language directives into temporal and dynamic logic representation for goal management and action execution,” in ICRA, 2009.
  7. 7.Y. Artzi and L. Zettlemoyer, “Weakly supervised learning of semantic parsers for mapping instructions to actions,” TACL, 2013.
  8. 8.C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” arXiv:2005.07648, 2020.
  9. 9.E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in CoRL, 2022.
  10. 10.O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” RA-L, 2022.
  11. 11.A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al., “Palm: Scaling language modeling with pathways,” arXiv:2204.02311, 2022.
  12. 12.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” NeurIPS, 2020.
  13. 13.S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin et al., “Opt: Open pre-trained transformer language models,” arXiv:2205.01068, 2022.
  14. 14.W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” arXiv:2201.07207, 2022.
  15. 15.T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” arXiv:2205.11916, 2022.
  16. 16.A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke et al., “Socratic models: Composing zero-shot multimodal reasoning with language,” arXiv:2204.00598, 2022.
  17. 17.M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog et al., “Do as i can, not as i say: Grounding language in robotic affordances,” arXiv:2204.01691, 2022.
  18. 18.W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter, “Inner monologue: Embodied reasoning through planning with language models,” in arXiv:2207.05608, 2022.
  19. 19.P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson, “Implicit behavioral cloning,” in CoRL, 2022.
  20. 20.A. Zeng, “Learning visual affordances for robotic manipulation,” Ph.D. dissertation, Princeton University, 2019.
  21. 21.D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke et al., “Scalable deep reinforcement learning for vision-based robotic manipulation,” in CoRL, 2018.
  22. 22.L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” arXiv:2203.02155, 2022.
  23. 23.D. Hupkes, V. Dankers, M. Mul, and E. Bruni, “Compositionality decomposed: How do neural networks generalise?” JAIR, 2020.
  24. 24.C. Breazeal, K. Dautenhahn, and T. Kanda, “Social robotics,” Springer handbook of robotics, 2016.
  25. 25.T. Kollar, S. Tellex, D. Roy, and N. Roy, “Toward understanding natural language directions,” in HRI, 2010.
  26. 26.J. Luketina, N. Nardelli, G. Farquhar, J. N. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel, “A survey of reinforcement learning informed by natural language,” in IJCAI, 2019.
  27. 27.M. MacMahon, B. Stankiewicz, and B. Kuipers, “Walk the talk: Connecting language, knowledge, and action in route instructions,” AAAI, 2006.
  28. 28.J. Thomason, S. Zhang, R. J. Mooney, and P. Stone, “Learning to interpret natural language commands through human-robot dialog,” in IJCAI, 2015.
  29. 29.S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy, “Understanding natural language commands for robotic navigation and mobile manipulation,” in AAAI, 2011.
  30. 30.D. Shah, B. Osinski, B. Ichter, and S. Levine, “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” arXiv:2207.04429, 2022.
  31. 31.C. Matuszek, E. Herbst, L. Zettlemoyer, and D. Fox, “Learning to parse natural language commands to a robot control system,” in Experimental robotics, 2013.
  32. 32.J. Thomason, A. Padmakumar, J. Sinapov, N. Walker, Y. Jiang, H. Yedidsion, J. Hart, P. Stone, and R. Mooney, “Jointly improving parsing and perception for natural language commands through human-robot dialog,” JAIR, 2020.
  33. 33.S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn et al., “Learning language-conditioned robot behavior from offline data and crowd-sourced annotation,” in CoRL, 2022.
  34. 34.J. Andreas, D. Klein, and S. Levine, “Learning with latent language,” arXiv:1711.00482, 2017.
  35. 35.P. Sharma, B. Sundaralingam, V. Blukis, C. Paxton, T. Hermans, A. Torralba, J. Andreas, and D. Fox, “Correcting robot plans with natural language feedback,” arXiv:2204.05186, 2022.
  36. 36.M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL, 2021.
  37. 37.S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor, “Language-conditioned imitation learning for robot manipulation tasks,” NeurIPS, 2020.
  38. 38.Y. Jiang, S. S. Gu, K. P. Murphy, and C. Finn, “Language as an abstraction for hierarchical deep reinforcement learning,” NeurIPS, 2019.
  39. 39.P. Goyal, S. Niekum, and R. J. Mooney, “Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards,” arXiv:2007.15543, 2020.
  40. 40.G. Cideron, M. Seurin, F. Strub, and O. Pietquin, “Self-educated language agent with hindsight experience replay for instruction following,” DeepMind, 2019.
  41. 41.D. Misra, J. Langford, and Y. Artzi, “Mapping instructions and visual observations to actions with reinforcement learning,” arXiv:1704.08795, 2017.
  42. 42.A. Akakzia, C. Colas, P.-Y. Oudeyer, M. Chetouani, and O. Sigaud, “Grounding language to autonomously-acquired skills via goal generation,” arXiv:2006.07185, 2020.
  43. 43.I. Drori, S. Zhang, R. Shuttleworth, L. Tang, A. Lu, E. Ke, K. Liu, L. Chen, S. Tran, N. Cheng et al., “A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level,” PNAS, 2022.
  44. 44.A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo et al., “Solving quantitative reasoning problems with language models,” arXiv:2206.14858, 2022.
  45. 45.K. Cobbe, V. Kosaraju, M. Bavarian, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” arXiv:2110.14168, 2021.
  46. 46.D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. Chi, “Least-to-most prompting enables complex reasoning in large language models,” arXiv:2205.10625, 2022.
  47. 47.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain of thought prompting elicits reasoning in large language models,” arXiv:2201.11903, 2022.
  48. 48.J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le et al., “Program synthesis with large language models,” arXiv:2108.07732, 2021.
  49. 49.K. Ellis, C. Wong, M. Nye, M. Sable-Meyer, L. Cary, L. Morales, L. Hewitt, A. Solar-Lezama, and J. B. Tenenbaum, “Dreamcoder: Growing generalizable, interpretable knowledge with wake-sleep bayesian program learning,” arXiv:2006.08381, 2020.
  50. 50.L. Tian, K. Ellis, M. Kryven, and J. Tenenbaum, “Learning abstract structure for drawing by efficient motor program induction,” NeurIPS, 2020.
  51. 51.D. Trivedi, J. Zhang, S.-H. Sun, and J. J. Lim, “Learning to synthesize programs as interpretable and generalizable policies,” NeurIPS, 2021.
  52. 52.O. Mees and W. Burgard, “Composing pick-and-place tasks by grounding language,” in ISER, 2020.
  53. 53.W. Liu, C. Paxton, T. Hermans, and D. Fox, “Structformer: Learning spatial structure for language-guided semantic rearrangement of novel objects,” in ICRA, 2022.
  54. 54.W. Yuan, C. Paxton, K. Desingh, and D. Fox, “Sornet: Spatial object-centric representations for sequential manipulation,” in CoRL, 2022.
  55. 55.A. Bucker, L. Figueredo, S. Haddadin, A. Kapoor, S. Ma, and R. Bonatti, “Reshaping robot trajectories using natural language commands: A study of multi-modal data alignment using transformers,” arXiv:2203.13411, 2022.
  56. 56.A. Bobu, C. Paxton, W. Yang, B. Sundaralingam, Y.-W. Chao, M. Cakmak, and D. Fox, “Learning perceptual concepts by bootstrapping from human queries,” RA-L, 2022.
  57. 57.J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano, “Recursively summarizing books with human feedback,” arXiv:2109.10862, 2021.
  58. 58.F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in MAPS, 2022.
  59. 59.K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi, “Xirl: Cross-embodiment inverse reinforcement learning,” in CoRL. PMLR, 2022.
  60. 60.A. Ganapathi, P. Florence, J. Varley, K. Burns, K. Goldberg, and A. Zeng, “Implicit kinematic policies: Unifying joint and cartesian action spaces in end-to-end robot learning,” arXiv:2203.01983, 2022.

Citation

MLA
Liang, J., et al. “Code as Policies: Language Model Programs for Embodied Control”. arXiv, 2022, http://arxiv.org/abs/2209.07753v4.
APA
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., & Zeng, A. (2022). Code as Policies: Language Model Programs for Embodied Control. arXiv. http://arxiv.org/abs/2209.07753v4
Chicago
Liang, J., W. Huang, F. Xia, et al. 2022. “Code as Policies: Language Model Programs for Embodied Control”. arXiv. http://arxiv.org/abs/2209.07753v4.
Harvard
Liang, J. et al. (2022) “Code as Policies: Language Model Programs for Embodied Control”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.07753v4.
Vancouver
1. Liang J, Huang W, Xia F, Xu P, Hausman K, Ichter B, Florence P, Zeng A (2022) Code as Policies: Language Model Programs for Embodied Control. arXiv

BibTeX

@article{liang2022code,
  title = {Code as Policies: Language Model Programs for Embodied Control},
  author = {Liang, Jacky and Huang, Wenlong and Xia, Fei and Xu, Peng and Hausman, Karol and Ichter, Brian and Florence, Pete and Zeng, Andy},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.07753v4},
  eprint = {2209.07753}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF