RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Yufei WangZhou XianFeng ChenTsun-Hsuan WangYian WangKaterina FragkiadakiZackory EricksonDavid HeldChuang Gan

article2024ICML299 citations

Introduces an automated robot learning framework that uses generative foundation models to propose tasks, construct simulation scenes, and produce training supervisions in an endless self-guided cycle with minimal human intervention.

Listen

Scaling up robotic skill acquisition has historically faced severe bottlenecks due to the slow, risky nature of real-world trials and the labor-intensive requirements of building simulation environments. Traditionally, human engineers must manually design virtual environments, handcraft three-dimensional assets, formulate scene layouts, and construct mathematical reward functions for every new capability. The article introduces RoboGen, an automated agent designed to eliminate these manual constraints by establishing a self-guided propose-generate-learn framework for automated robot learning in simulation.

The primary objective of the article is to demonstrate that foundation and generative models can autonomously configure tasks, environments, and training supervisions to scale robotic skill acquisition with minimal human intervention. Rather than directly using large language models to control robot joints—an approach that often struggles due to a lack of physical grounding—the system uses language and vision models for task imagination, spatial reasoning, and supervision design, while delegating low-level physics execution to physics-grounded simulators.

The authors implemented this approach using a self-guided propose-generate-learn pipeline operating on the differentiable simulation platform Genesis, with OpenAI's GPT-4 as the primary language backend and Gemini-Pro as a visual verifier. The system queries the language model to generate tasks conditioned on robot capabilities, creates environments by retrieving meshes from the Objaverse database or generating them via image-to-3D pipelines, automatically verifies realistic object scales, decomposes tasks into sub-tasks, and selects the most appropriate learning algorithm. RoboGen chooses between reinforcement learning, gradient-based trajectory optimization, and motion planning depending on whether the task involves articulated manipulation, deformable soft-body interaction, or legged locomotion.

Evaluation showed that RoboGen achieved superior task and visual diversity compared to established human-curated robotic benchmarks (such as RLBench, ManiSkill2, Meta-World, and Behavior-100) and concurrent procedural generation systems, as measured by lower text redundancy and image similarity scores. Across a benchmark suite of 69 diverse tasks spanning articulated objects, soft materials, and quadruped locomotion, the automated learning pipeline achieved an overall training success rate of 77.4%. Hybrid algorithm selection proved essential; relying solely on reinforcement learning for articulated manipulation caused nearly all tasks to fail, whereas pairing motion planning with learning enabled reliable execution. Furthermore, automated visual and dimensional verification significantly reduced asset failures, and an analysis of 155 generated tasks identified only 19 failure cases, primarily caused by asset geometry mismatches or language ambiguities in object joint limits.

These findings demonstrate that generative simulation offers a scalable, cost-effective path to producing massive volumes of diverse robotic training data without requiring continuous manual engineering. Organizations developing autonomous systems can substantially reduce development cycles and human overhead by automating task design and reward authoring. However, real-world deployment remains bounded by the simulation-to-reality gap, requiring techniques such as domain randomization to transfer policies safely. Next steps include integrating multimodal environmental feedback loops to autonomously verify learned skills and self-correct reward errors, as well as conducting physical real-robot transfer trials before broad deployment in production environments.

Cover for RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Abstract

We present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead of directly using or adapting these models to produce policies or low-level actions, we advocate for a generative scheme, which uses these models to automatically generate diversified tasks, scenes, and training supervisions, thereby scaling up robotic skill learning with minimal human supervision. Our approach equips a robotic agent with a self-guided propose-generate-learn cycle: the agent first proposes interesting tasks and skills to develop, and then generates corresponding simulation environments by populating pertinent objects and assets with proper spatial configurations. Afterwards, the agent decomposes the proposed high-level task into sub-tasks, selects the optimal learning approach (reinforcement learning, motion planning, or trajectory optimization), generates required training supervision, and then learns policies to acquire the proposed skill. Our work attempts to extract the extensive and versatile knowledge embedded in large-scale models and transfer them to the field of robotics. Our fully generative pipeline can be queried repeatedly, producing an endless stream of skill demonstrations associated with diverse tasks and environments.

Citation

MLA
Wang, Y., et al. “RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation”. arXiv, 2023, http://arxiv.org/abs/2311.01455v3.
APA
Wang, Y., Xian, Z., Chen, F., Wang, T.-H., Wang, Y., Fragkiadaki, K., Erickson, Z., Held, D., & Gan, C. (2023). RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation. arXiv. http://arxiv.org/abs/2311.01455v3
Chicago
Wang, Y., Z. Xian, F. Chen, et al. 2023. “RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation”. arXiv. http://arxiv.org/abs/2311.01455v3.
Harvard
Wang, Y. et al. (2023) “RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.01455v3.
Vancouver
1. Wang Y, Xian Z, Chen F, Wang T-H, Wang Y, Fragkiadaki K, Erickson Z, Held D, Gan C (2023) RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation. arXiv

BibTeX

@article{wang2023robogen,
  title = {RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation},
  author = {Wang, Yufei and Xian, Zhou and Chen, Feng and Wang, Tsun-Hsuan and Wang, Yian and Fragkiadaki, Katerina and Erickson, Zackory and Held, David and Gan, Chuang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.01455v3},
  eprint = {2311.01455}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/