RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

Yao MuJunting ChenQinglong ZhangShoufa ChenQiaojun YuChongjian GeRunjian ChenZhixuan LiangMengkang HuChaofan Tao

article2024ICML66 citations

Presents RoboCodeX, a multimodal vision-language framework that decomposes complex instructions into tree-structured, object-centric manipulation units to generate executable control code with physical and safety constraints across different robot platforms.

Listen

Deploying intelligent robots to handle complex, everyday household and workplace tasks remains a fundamental challenge in artificial intelligence. While recent vision-language models excel at high-level reasoning and scene understanding, they struggle to convert abstract goals into the precise physical control commands required for hardware execution. Conversely, traditional robotic control frameworks lack the adaptability to generalize across diverse objects, physical mechanisms, and different robot hardware designs. Bridging this gap between high-level conceptual understanding and low-level physical manipulation is essential for creating adaptable, cross-platform robotic systems.

The article introduces and evaluates RoboCodeX, a multimodal vision-language framework designed to synthesize robotic behaviors by generating executable Python code. The primary objective is to demonstrate that grounding code generation in visual perception and physical constraints enables robots to generalize across complex manipulation tasks, novel object geometries, and distinct hardware platforms.

To achieve this, the approach employs a tree-of-thought structure that decomposes high-level human instructions into sequential, object-centric manipulation subtasks. For each subtask, the model analyzes multi-view depth camera observations to infer physical constraints, grasping preferences, and target positions before calling specialized robotics programming interfaces to generate motion trajectories. To train this framework, the authors generated a simulation pretraining dataset of 147,303 multimodal interaction samples and a fine-tuning dataset of 49,320 high-quality samples refined iteratively to guarantee execution viability. The architecture couples a vision transformer, a feature adapter, and a 13-billion-parameter language model, which was validated through simulated benchmarks and real-world robot trials.

Empirical evaluations show that RoboCodeX significantly outperforms leading models across multiple robotic domains. In manipulation benchmarks, it achieved an 80% success rate on open-vocabulary pick-and-place tasks—a 17% absolute improvement over GPT-4V (63%)—and reached success rates of 84% on drawer manipulation and 74% on door opening. On embodied navigation benchmarks across simulated environments, it attained a 40.0% to 40.3% success rate, matching or outperforming foundation model baselines. In zero-shot real-world experiments using two distinct robotic arms (Franka Emika Panda and UR5), RoboCodeX completed varied tasks with success rates between 80% and 100% simply by updating the robot configuration file, requiring no hardware-specific model fine-tuning. Ablation studies confirmed that inferring grasping preferences and retaining general vision-question-answering data during training were essential to prevent performance drops and overfitting.

These findings indicate that using executable code as a bridge between multimodal perception and physical control provides a scalable, cost-effective way to achieve cross-platform robot deployment. Rather than retraining specialized models for each robot type or object shape, organizations can deploy a unified vision-language policy that interfaces directly with existing robotic motion planners. This reduces integration timelines, improves operational safety by respecting physical joint constraints, and lowers software development overhead.

Organizations developing autonomous systems should consider adopting modular, code-generating vision-language architectures to control heterogeneous robot fleets. When training these systems, practitioners should retain general vision-language datasets alongside domain-specific code to prevent model overfitting. However, deployment should proceed with caution: the current framework faces limitations in force-sensitive assembly tasks, such as fastening nuts and bolts, and unstructured contact operations like wiping or sweeping. Future development should focus on integrating dynamic force sensing and refining closed-loop feedback before deploying this architecture into high-precision industrial manufacturing.

arXiv: 2402.16117
  • Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Read this foundational demonstration of language models generating executable robot-control programs first; RoboCodeX develops the same code-as-policy bridge with multimodal grounding and broader cross-platform evaluation.
Cover for RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

Abstract

Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these conceptual understandings into detailed robotic actions while achieving generalization across various scenarios. In this paper, we propose a tree-structured multimodal code generation framework for generalized robotic behavior synthesis, termed RoboCodeX. RoboCodeX decomposes high-level human instructions into multiple object-centric manipulation units consisting of physical preferences such as affordance and safety constraints, and applies code generation to introduce generalization ability across various robotics platforms. To further enhance the capability to map conceptual and perceptual understanding into control commands, a specialized multimodal reasoning dataset is collected for pre-training and an iterative refining methodology is introduced for supervised fine-tuning. Extensive experiments demonstrate that RoboCodeX achieves state-of-the-art performance in both simulators and real robots on four different kinds of manipulation tasks and one embodied navigation task. More demos and information can be found in our homepage.

Table of Contents

  • Introduction
  • Related Works
  • Methods
  • Problem Setup
  • Multi-modal Tree-of-thought Code Generation
  • Dataset Preparation
  • Vision Language Model Design
  • Experiments
  • Evaluation on Manipulation Task
  • Evaluation on Embodied Navigation Task
  • Evaluation on General VQA
  • Real World Experiments
  • Ablation Study
  • Limitation
  • Conclusion
  • Impact Statement
  • Acknowledgement
  • References
  • A. Manipulation Simulation Setup
  • B. Implementation details of Vision Language Model
  • C. Details of Dataset Collection
  • D. Explanation of the relationship between different components of RoboCodeX
  • E. Joint prediction of Articulated Objects
  • F. Grasp Pose Prediction
  • F.1. Grasp Pose Proposal Generation
  • F.2. Grasp Execution
  • G. Ablation Studies on Visual Reasoning Capabilities
  • H. Real world experiments
  • I. Introduction of the APIs and the prompts

Knowls

  1. Knowl 1 — Object-centric multimodal code generation for robotic behavior

    model/method

    RoboCodeX maps RGB observations and a free-form human instruction to executable robot-control code through a tree-structured, object-centric reasoning process. It first decomposes a long task into sequential manipulation units, each centered on a relevant object or object part. Each unit records its subtask, target-position proposals, inferred physical constraints, and manipulation preferences—such as where to grasp or which direction to approach—before generating code for the associated behavior. The code invokes perception and planning tools to produce motions suited to the robot’s mechanics, making it an interface between high-level semantic reasoning and low-level control. For example, placing an item in a drawer can require a drawer-focused unit to pull the drawer along its prismatic joint axis, followed by an item-focused unit to grasp and place the item.

  2. Knowl 2 — Per-unit perception, grasp selection, and trajectory planning

    model/method

    RoboCodeX receives three RGB-D views, from left, right, and top cameras. Depth information from the views is fused into a 3D scene representation; image-grounded object locations are associated with 3D boxes using overlap and orientation consistency, yielding a point cloud for each task-relevant object. For each object-centric unit, the system can segment a task-relevant part, such as a drawer handle, and use AnyGrasp to generate grasp-pose candidates. It considers the ten highest-ranked candidates and selects among them using the model’s inferred task and object preferences. For articulated objects, a GAMMA-based tool predicts rigid parts and joint information; a plane-detection tool supplies surface normals when needed. A trajectory planner then generates end-effector waypoints and gripper commands subject to physical constraints and collision avoidance. The paper also describes zeroth-order trajectory optimization that samples candidate trajectories and evaluates control cost and occupancy-map constraints.

  3. Knowl 3 — Long-horizon task and trajectory formulation

    equation

    A long-horizon instruction LglobalL_{\mathrm{global}} is decomposed into NN object-centric subtasks. For subtask ii, the desired trajectory τi\tau_i is a sequence of dense end-effector waypoints, each containing a desired 6-DoF pose, end-effector velocity, and gripper action. The paper formulates per-subtask trajectory generation as minimizing task and control costs subject to dynamics and kinematics constraints:

    min⁡τi  Stask(τi,ℓi∗)+Scontrol(τi)subject to C(τi).\min_{\tau_i}\; S_{\mathrm{task}}(\tau_i,\ell_i^*) + S_{\mathrm{control}}(\tau_i) \quad \text{subject to } C(\tau_i).

    Here, ℓi∗\ell_i^* is the ground-truth instruction for subtask ii; StaskS_{\mathrm{task}} scores how well the trajectory completes that instruction; ScontrolS_{\mathrm{control}} represents control costs, such as effort or execution time; and C(τi)C(\tau_i) denotes the trajectory’s dynamics and kinematics constraints. The trajectories for the subtasks together are intended to complete LglobalL_{\mathrm{global}}.

  4. Knowl 4 — Multimodal robotic code pretraining data

    experimental setup

    RoboCodeX’s pretraining data was generated in simulated household scenes sampled from HM3D and populated with objects drawn from the Google Scanned Objects, YCB, OmniObject3D, and AKB-48 datasets. GPT-4 generated free-form tasks appropriate to the scenes and candidate programs for completing them; ten candidate code samples were generated per task, and GPT-3.5 was used to evaluate and filter out code with syntax errors. The resulting dataset contains 147,303 multi-round conversations with image inputs, drawn from 100 randomly generated scenarios with around 100 tasks per scenario. Tasks include object placement, rearrangement, and relative-position instructions. Sample lengths range from 1,649 to 2,446 tokens, averaging 2,015.4 tokens. During pretraining, this data was mixed in a 1:1 ratio with general visual-language data from ShareGPT4V, SVIT, and LLaVA Visual Instruct 150K.

  5. Knowl 5 — Execution-based refinement of supervised fine-tuning data

    model/method

    RoboCodeX’s supervised fine-tuning data was built around task types from RT-1 and LIBERO, with diverse objects combined into task instances. Human-provided examples were verified for successful completion in simulation and on real robots, and GPT-4 generated corresponding code using the examples and API explanations. When syntactically valid code failed in execution, the authors searched over grasping method (central lift or AnyGrasp), preferred contact position, preferred gripper direction, and trajectory-generation method. GPT-4V analyzed the better-performing settings and converted the rationale into code annotations; code that still had zero success after the search was manually revised. The resulting dataset contains 49,320 image-based, multi-round conversations, and retains only samples whose code achieved a success rate above 50%. Sample lengths range from 1,928 to 3,142 tokens, averaging 2,259.8 tokens. During fine-tuning, this dataset was mixed with general ShareGPT4V data at a 10:1 ratio to reduce overfitting.

  6. Knowl 6 — Vision-language architecture for long code-generation contexts

    model/method

    RoboCodeX is a 13-billion-parameter vision-language model based on the BLIP-2 design, with a vision transformer, Q-Former, and language model. It uses a pretrained EVA-CLIP ViT-G/14 vision transformer and LLaMA-13B language model; the vision transformer’s final layer is discarded, and the Q-Former compresses visual embeddings into a shorter token sequence before they are combined with text. This compression is intended to control memory demands when code-generation prompts also contain lengthy API documentation. To retain hierarchical image features, the model divides the vision transformer’s layers into four proportional groups, extracts a class token from the final layer of each group, and processes the four tokens with a vision adapter. The adapter reduces channel dimension, applies a SiLU activation, restores the output dimension, and applies layer normalization. Its parameters are zero-initialized.

  7. Knowl 7 — Simulated manipulation success across five task types

    empirical result

    In Gazebo simulation, the authors measured average success rates over 50 trials for five manipulation task types: open-vocabulary pick-and-place, drawer opening and closing, door opening and closing, placing an object in a drawer, and multi-stage tasks. RoboCodeX outperformed the three tested language-model baselines in every task category:

    • Pick and place: GPT-3.5 0.45, GPT-4 0.56, GPT-4V 0.63, RoboCodeX 0.80.
    • Drawers: GPT-3.5 0.44, GPT-4 0.68, GPT-4V 0.68, RoboCodeX 0.84.
    • Doors: GPT-3.5 0.20, GPT-4 0.46, GPT-4V 0.52, RoboCodeX 0.74.
    • Putting an object in a drawer: GPT-3.5 0.36, GPT-4 0.48, GPT-4V 0.55, RoboCodeX 0.68.
    • Multi-stage tasks: GPT-3.5 0.20, GPT-4 0.44, GPT-4V 0.50, RoboCodeX 0.64.

    The pick-and-place evaluation covered 42 object categories; articulated-object tasks used five cabinets. GPT-4V received images and instructions, whereas GPT-4 and GPT-3.5 received textual scene descriptions and instructions. All methods used the same API and prompt. These results show RoboCodeX’s advantage in both object manipulation and tasks requiring articulated-object reasoning or sequential subtasks.

  8. Knowl 8 — Zero-shot transfer to two real robot platforms

    empirical result

    RoboCodeX was evaluated without task-specific fine-tuning on a Franka Emika Panda arm and a UR5 arm. The authors report changing the robot configuration file to adapt the system to the different platforms. Each reported task used 10 trials with randomized initial object positions. The reported success rates were 80% for putting makeup in a red bag, 100% for placing the nearest persimmon in a bowl, 90% for helping a toy sit on a car, 80% for putting a Pepsi can in a drawer, and 100% for putting rubbish in a brown box. The results demonstrate transfer across the tested robots and task settings without platform-specific training; the paper does not specify a per-task breakdown by robot.

  9. Knowl 9 — Visual reasoning for embodied object navigation

    empirical result

    For visual-language object navigation, RoboCodeX uses the L3MVN framework but applies multimodal reasoning to the current observation and candidate frontiers to choose where to explore. On HM3D and HSSD, the reported metrics are success rate and success weighted by path length (SPL), with values reported on a percentage scale. RoboCodeX scored 40.0 success and 24.2 SPL on HM3D, and 40.3 success and 22.0 SPL on HSSD. For comparison, L3MVN (GPT-2) scored 35.2 and 16.5 on HM3D and 38.4 and 19.4 on HSSD; Pixel-Nav (GPT-4) scored 37.9 and 20.5 on HM3D; and ESC (GPT-3.5) scored 39.2 and 22.3 on HM3D. The paper reports no HSSD values for Pixel-Nav or ESC. RoboCodeX therefore scored above the listed HM3D baselines and above L3MVN on both reported HSSD metrics.

  10. Knowl 10 — Ablations identify the roles of preferences and visual information

    empirical result

    Ablations tested inferred grasp preferences, the vision adapter, general visual-question-answering data during fine-tuning, and visual information for navigation. In manipulation, preference-based grasp selection outperformed selecting AnyGrasp’s highest-scoring grasp across the evaluated task types; removing the vision adapter also reduced success, while removing general VQA data led to overfitting, poorer instruction following, and a substantial success-rate decline. For navigation, the reported HM3D success/SPL and HSSD success/SPL values were: full RoboCodeX, 40/24.2 and 40.3/22.0; without the vision adapter, 39.2/23.5 and 40.1/21.5; and without a visual scene caption, 38.0/22.5 and 39.6/21.0. The adapter’s effect also appeared on LLaVA-Bench: RoboCodeX scored 71.5, compared with 67.4 without the adapter. These ablations support the value of visual reasoning and task-informed preferences, and the use of general VQA data to limit fine-tuning overfit.

  11. Knowl 11 — General multimodal benchmark performance

    data/table

    The authors evaluated general multimodal understanding using LLaVA-Bench and MM-Vet. RoboCodeX-13B scored 71.5 on LLaVA-Bench and 31.0 on MM-Vet. The reported comparisons were: BLIP-2, 38.1 and 22.4; InstructBLIP-7B, 60.9 and 26.2; InstructBLIP-13B, 58.2 and 25.6; LLaVA-1.5-7B, 63.4 and 30.5; and LLaVA-1.5-13B, 70.7 and 35.4, respectively. Thus RoboCodeX scored above the listed models on LLaVA-Bench, including LLaVA-1.5-13B, while its MM-Vet score was below LLaVA-1.5-13B’s 35.4. The paper presents this evaluation as evidence that robotic code-generation training can coexist with broadly capable multimodal reasoning.

  12. Knowl 12 — Stated limitations on dexterity and unstructured tasks

    limitation

    The authors identify two limitations of RoboCodeX. First, it lacks optimal force-sensor handling for precise assembly tasks, such as working with nuts and bolts. Second, it struggles with unstructured tasks such as wiping tables or sweeping, which require fluid and adaptable behavior.

Coverage note — The detailed API inventory and the articulated-joint predictor’s component-level metrics are omitted because they describe supporting tools and implementation details rather than separate central contributions.

References

  1. 1.Agrawal, A., Prabhakar, R., Goyal, A., and Liu, D. Physical reasoning and object planning for household embodied agents. arXiv preprint arXiv:2311.13577, 2023.
  2. 2.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A. Do as i can and not as i say: Grounding language in robotic affordances. In arXiv preprint arXiv:2204.01691, 2022.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35: 23716–23736, 2022.
  4. 4.Anderson, P., Chang, A., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018.
  5. 5.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  6. 6.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  7. 7.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  8. 8.Cai, W., Huang, S., Cheng, G., Long, Y., Gao, P., Sun, C., and Dong, H. Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill. arXiv preprint arXiv:2309.10309, 2023.
  9. 9.Calli, B., Walsman, A., Singh, A., Srinivasa, S., Abbeel, P., and Dollar, A. M. Benchmarking in manipulation research: The ycb object and model set and benchmarking protocols. arXiv preprint arXiv:1502.03143, 2015.
  10. 10.Chen, B., Xia, F., Ichter, B., Rao, K., Gopalakrishnan, K., Ryoo, M. S., Stone, A., and Kappler, D. Open-vocabulary queryable scene representations for real world planning. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 11509–11522. IEEE, 2023a.
  11. 11.Chen, J., Mu, Y., Yu, Q., Wei, T., Wu, S., Yuan, Z., Liang, Z., Yang, C., Zhang, K., Shao, W., et al. Roboscript: Code generation for free-form manipulation tasks across real and simulation. arXiv preprint arXiv:2402.14623, 2024.
  12. 12.Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023b.
  13. 13.Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023c.
  14. 14.Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., et al. Pali: A jointly-scaled multilingual language-image model. In ICLR, 2022.
  15. 15.Dasari, S., Gupta, A., and Kumar, V. Learning dexterous manipulation from exemplar object trajectories and pre-grasps. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 3889–3896. IEEE, 2023.
  16. 16.Ding, Y., Zhang, X., Paxton, C., and Zhang, S. Task and motion planning with large language models for object rearrangement. arXiv preprint arXiv:2303.06247, 2023.
  17. 17.Dong, G., Yuan, H., Lu, K., Li, C., Xue, M., Liu, D., Wang, W., Yuan, Z., Zhou, C., and Zhou, J. How abilities in large language models are affected by supervised fine-tuning data composition. arXiv preprint arXiv:2310.05492, 2023.
  18. 18.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  19. 19.Downs, L., Francis, A., Koenig, N., Kinman, B., Hickman, R., Reymann, K., McHugh, T. B., and Vanhoucke, V. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pp. 2553–2560. IEEE, 2022.
  20. 20.Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P. Palm-e: An embodied multimodal language model. In arXiv preprint arXiv:2303.03378, 2023.
  21. 21.Eisner, B., Zhang, H., and Held, D. Flowbot3d: Learning 3d articulation flow to manipulate articulated objects. arXiv preprint arXiv:2205.04382, 2022.
  22. 22.Fang, H.-S., Wang, C., Fang, H., Gou, M., Liu, J., Yan, H., Liu, W., Xie, Y., and Lu, C. Anygrasp: Robust and efficient grasp perception in spatial and temporal domains. IEEE Transactions on Robotics, 2023.
  23. 23.Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y. Eva: Exploring the limits of masked visual representation learning at scale. arXiv preprint arXiv:2211.07636, 2022.
  24. 24.Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D. Physically grounded vision-language models for robotic manipulation. arXiv preprint arXiv:2309.02561, 2023.
  25. 25.Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18995–19012, 2022.
  26. 26.Haddadin, S., Parusel, S., Johannsmeier, L., Golz, S., Gabl, S., Walch, F., Sabaghian, M., Jahne, C., Hausperger, L., and Haddadin, S. The franka emika robot: A reference platform for robotics research and education. IEEE Robotics & Automation Magazine, 29(2):46–64, 2022.
  27. 27.Hu, M., Mu, Y., Yu, X., Ding, M., Wu, S., Shao, W., Chen, Q., Wang, B., Qiao, Y., and Luo, P. Tree-planner: Efficient close-loop task planning with large language models. arXiv preprint arXiv:2310.08582, 2023.
  28. 28.Huang, C., Mees, O., Zeng, A., and Burgard, W. Visual language maps for robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 10608–10615. IEEE, 2023a.
  29. 29.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pp. 9118–9147. PMLR, 2022a.
  30. 30.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022b.
  31. 31.Huang, W., Xia, F., Shah, D., Driess, D., Zeng, A., Lu, Y., Florence, P., Mordatch, I., Levine, S., Hausman, K., et al. Grounded decoding: Guiding text generation with grounded models for robot control. arXiv preprint arXiv:2303.00855, 2023b.
  32. 32.Huang, Y., Cai, M., Li, Z., and Sato, Y. Predicting gaze in egocentric video by learning task-dependent attention transition. In Proceedings of the European conference on computer vision (ECCV), pp. 754–769, 2018.
  33. 33.Huang, Y., Chen, G., Xu, J., Zhang, M., Yang, L., Pei, B., Zhang, H., Dong, L., Wang, Y., Wang, L., and Qiao, Y. Egobridge: A dataset for bridging asynchronous first- and third-person views of activities in the real world. 2023c. URL https://egobridge.github.io/static/videos/egobridge_paper.pdf.
  34. 34.Kebria, P. M., Al-Wais, S., Abdi, H., and Nahavandi, S. Kinematic and dynamic modelling of ur5 manipulator. In 2016 IEEE international conference on systems, man, and cybernetics (SMC), pp. 004229–004234. IEEE, 2016.
  35. 35.Khanna, M., Mao, Y., Jiang, H., Haresh, S., Shacklett, B., Batra, D., Clegg, A., Undersander, E., Chang, A. X., and Savva, M. Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation, 2023.
  36. 36.Khatib, O. A unified approach for motion and force control of robot manipulators: The operational space formulation. IEEE Journal on Robotics and Automation, 3(1):43–53, 1987.
  37. 37.Koubaa, A. et al. ˆ Robot Operating System (ROS)., volume 1. Springer, 2017.
  38. 38.Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023.
  39. 39.Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., and Liu, Z. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023a.
  40. 40.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888–12900. PMLR, 2022a.
  41. 41.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023b.
  42. 42.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023c.
  43. 43.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023d.
  44. 44.Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10965–10975, 2022b.
  45. 45.Li, X., Zhang, M., Geng, Y., Geng, H., Long, Y., Shen, Y., Zhang, R., Liu, J., and Dong, H. Manipllm: Embodied multimodal large language model for object-centric robotic manipulation. arXiv preprint arXiv:2312.16217, 2023e.
  46. 46.Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., and Bai, X. Monkey: Image resolution and text label are important things for large multi-modal models. arXiv preprint arXiv:2311.06607, 2023f.
  47. 47.Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A. Code as Policies: Language model programs for embodied control. In IEEE International Conference on Robotics and Automation (ICRA), pp. 9493–9500. IEEE, 2023.
  48. 48.Liu, B., Jiang, Y., Zhang, X., Liu, Q., Zhang, S., Biswas, J., and Stone, P. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477, 2023a.
  49. 49.Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., and Stone, P. Libero: Benchmarking knowledge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310, 2023b.
  50. 50.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023c.
  51. 51.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. NeurIPS, 2023d.
  52. 52.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023e.
  53. 53.Liu, L., Xu, W., Fu, H., Qian, S., Yu, Q., Han, Y., and Lu, C. Akb-48: A real-world articulated object knowledge base. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14809–14818, 2022.
  54. 54.Lu, Y., Li, C., Liu, H., Yang, J., Gao, J., and Shen, Y. An empirical study of scaling instruct-tuned large multimodal models. arXiv preprint arXiv:2309.09958, 2023a.
  55. 55.Lu, Y., Lu, P., Chen, Z., Zhu, W., Wang, X. E., and Wang, W. Y. Multimodal procedural planning via dual text-image prompting. arXiv preprint arXiv:2305.01795, 2023b.
  56. 56.Mirjalili, R., Krawez, M., Silenzi, S., Blei, Y., and Burgard, W. Lan-grasp: Using large language models for semantic object grasping. arXiv preprint arXiv:2310.05239, 2023.
  57. 57.Mo, K., Zhu, S., Chang, A. X., Yi, L., Tripathi, S., Guibas, L. J., and Su, H. PartNet: A large-scale benchmark for fine-grained and hierarchical part-level 3D object understanding. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  58. 58.Mu, Y., Yao, S., Ding, M., Luo, P., and Gan, C. Ec2: Emergent communication for embodied control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6704–6714, 2023a.
  59. 59.Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P. Embodiedgpt: Vision-language pre-training via embodied chain of thought. arXiv preprint arXiv:2305.15021, 2023b.
  60. 60.Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A. R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601, 2022.
  61. 61.OpenAI. Gpt-4 technical report. ArXiv, abs/2303.08774, 2023a.
  62. 62.OpenAI. Gpt-4 technical report, 2023b.
  63. 63.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  64. 64.Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023.
  65. 65.Paul, A., Bandyopadhyay, R., Yoon, J. H., Geem, Z. W., and Sarkar, R. Sinlu: Sinu-sigmoidal linear unit. Mathematics, 10(3):337, 2022.
  66. 66.Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
  67. 67.Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A. VirtualHome: Simulating household activities via programs. In CVPR, pp. 8494–8502, 2018.
  68. 68.Qi, C. R., Yi, L., Su, H., and Guibas, L. J. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017.
  69. 69.Qian, W., Xia, Z., Xiong, J., Gan, Y., Guo, Y., Weng, S., Deng, H., Hu, Y., and Zhang, J. Manipulation task simulation using ros and gazebo. In 2014 IEEE International Conference on Robotics and Biomimetics (ROBIO 2014), pp. 2594–2598. IEEE, 2014.
  70. 70.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  71. 71.Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv preprint arXiv:2109.08238, 2021.
  72. 72.Raman, S. S., Cohen, V., Rosen, E., Idrees, I., Paulius, D., and Tellex, S. Planning with large language models via corrective re-prompting. In NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022.
  73. 73.Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I., and Suenderhauf, N. Sayplan: Grounding large language models using 3d scene graphs for scalable task planning. arXiv preprint arXiv:2307.06135, 2023.
  74. 74.Sha, H., Mu, Y., Jiang, Y., Chen, L., Xu, C., Luo, P., Li, S. E., Tomizuka, M., Zhan, W., and Ding, M. Languagempc: Large language models as decision makers for autonomous driving. arXiv preprint arXiv:2310.03026, 2023.
  75. 75.Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A. ProgPrompt: Generating situated robot task plans using large language models. In IEEE International Conference on Robotics and Automation (ICRA), pp. 11523–11530. IEEE, 2023.
  76. 76.Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y. Llm-planner: Few-shot grounded planning for embodied agents with large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2998–3009, 2023.
  77. 77.Sucan, I. A., Moll, M., and Kavraki, L. E. The open motion planning library. IEEE Robotics & Automation Magazine, 19 (4):72–82, 2012.
  78. 78.Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X. Generative pretraining in multimodality. arXiv preprint arXiv:2307.05222, 2023.
  79. 79.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  80. 80.Vemprala, S., Bonatti, R., Bucker, A., and Kapoor, A. Chatgpt for robotics: Design principles and model abilities. Microsoft Auton. Syst. Robot. Res, 2:20, 2023.
  81. 81.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a.
  82. 82.Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., et al. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. NeurIPS, 2023b.
  83. 83.Wang, W., Shi, M., Li, Q., Wang, W., Huang, Z., Xing, L., Chen, Z., Li, H., Zhu, X., Cao, Z., et al. The all-seeing project: Towards panoptic visual recognition and understanding of the open world. arXiv preprint arXiv:2308.01907, 2023c.
  84. 84.Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519, 2023a.
  85. 85.Wu, T., Zhang, J., Fu, X., Wang, Y., Ren, J., Pan, L., Wu, W., Yang, L., Wang, J., Qian, C., et al. Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 803–814, 2023b.
  86. 86.Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al. Sapien: A simulated part-based interactive environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11097–11107, 2020.
  87. 87.Xu, Z., He, Z., and Song, S. Universal manipulation policy network for articulated objects. IEEE Robotics and Automation Letters, 7(2):2447–2454, 2022.
  88. 88.Yan, J., Zhao, H., Bu, P., and Jin, Y. Channel-wise attention-based network for self-supervised monocular depth estimation. In 2021 International Conference on 3D vision (3DV), pp. 464–473. IEEE, 2021.
  89. 89.Yang, J., Dong, Y., Liu, S., Li, B., Wang, Z., Jiang, C., Tan, H., Kang, J., Zhang, Y., Zhou, K., et al. Octopus: Embodied vision-language programmer from environmental feedback. arXiv preprint arXiv:2310.08588, 2023a.
  90. 90.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9, 2023b.
  91. 91.Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Dan, Y., Zhao, C., Xu, G., Li, C., Tian, J., Qi, Q., Zhang, J., and Huang, F. mplug-docowl: Modularized multimodal large language model for document understanding, 2023.
  92. 92.Yu, B., Kasaei, H., and Cao, M. L3mvn: Leveraging large language models for visual target navigation. arXiv preprint arXiv:2304.05501, 2023a.
  93. 93.Yu, Q., Wang, J., Liu, W., Hao, C., Liu, L., Shao, L., Wang, W., and Lu, C. Gamma: Generalizable articulation modeling and manipulation for articulated objects. arXiv preprint arXiv:2309.16264, 2023b.
  94. 94.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023c.
  95. 95.Yuan, H., Zhang, C., Wang, H., Xie, F., Cai, P., Dong, H., and Lu, Z. Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks. arXiv preprint arXiv:2303.16563, 2023.
  96. 96.Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., et al. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598, 2022.
  97. 97.Zhang, H., Li, X., and Bing, L. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023a.
  98. 98.Zhang, P., Wang, X. D. B., Cao, Y., Xu, C., Ouyang, L., Zhao, Z., Ding, S., Zhang, S., Duan, H., Yan, H., et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023b.
  99. 99.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023c.
  100. 100.Zhang, S., Sun, P., Chen, S., Xiao, M., Shao, W., Zhang, W., Chen, K., and Luo, P. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023d.
  101. 101.Zhao, B., Wu, B., and Huang, T. Svit: Scaling up visual instruction tuning. arXiv preprint arXiv:2307.04087, 2023.
  102. 102.Zhong, C., Zheng, Y., Zheng, Y., Zhao, H., Yi, L., Mu, X., Wang, L., Li, P., Zhou, G., Yang, C., et al. 3d implicit transporter for temporally consistent keypoint discovery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3869–3880, 2023.
  103. 103.Zhou, H., Ding, M., Peng, W., Tomizuka, M., Shao, L., and Gan, C. Generalizable long-horizon manipulations with large language models. arXiv preprint arXiv:2310.02264, 2023a.
  104. 104.Zhou, K., Zheng, K., Pryor, C., Shen, Y., Jin, H., Getoor, L., and Wang, X. E. Esc: Exploration with soft commonsense constraints for zero-shot object navigation. arXiv preprint arXiv:2301.13166, 2023b.
  105. 105.Zhou, Q.-Y., Park, J., and Koltun, V. Open3d: A modern library for 3d data processing. arXiv preprint arXiv:1801.09847, 2018.
  106. 106.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023a.
  107. 107.Zhu, X., Chen, Y., Tian, H., Tao, C., Su, W., Yang, C., Huang, G., Li, B., Lu, L., Wang, X., et al. Ghost in the minecraft: Generally capable agents for open-world enviroments via large language models with text-based knowledge and memory. arXiv preprint arXiv:2305.17144, 2023b.

Citation

MLA
Mu, Y., et al. “RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis”. arXiv, 2024, http://arxiv.org/abs/2402.16117v1.
APA
Mu, Y., Chen, J., Zhang, Q., Chen, S., Yu, Q., Ge, C., Chen, R., Liang, Z., Hu, M., Tao, C., Sun, P., Yu, H., Yang, C., Shao, W., Wang, W., Dai, J., Qiao, Y., Ding, M., & Luo, P. (2024). RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis. arXiv. http://arxiv.org/abs/2402.16117v1
Chicago
Mu, Y., J. Chen, Q. Zhang, et al. 2024. “RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis”. arXiv. http://arxiv.org/abs/2402.16117v1.
Harvard
Mu, Y. et al. (2024) “RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.16117v1.
Vancouver
1. Mu Y, Chen J, Zhang Q, et al (2024) RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis. arXiv

BibTeX

@article{mu2024robocodex,
  title = {RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis},
  author = {Mu, Yao and Chen, Junting and Zhang, Qinglong and Chen, Shoufa and Yu, Qiaojun and Ge, Chongjian and Chen, Runjian and Liang, Zhixuan and Hu, Mengkang and Tao, Chaofan and Sun, Peize and Yu, Haibao and Yang, Chao and Shao, Wenqi and Wang, Wenhai and Dai, Jifeng and Qiao, Yu and Ding, Mingyu and Luo, Ping},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.16117v1},
  eprint = {2402.16117}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/