RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
Yufei WangZhou XianFeng ChenTsun-Hsuan WangYian WangKaterina FragkiadakiZackory EricksonDavid HeldChuang Gan
Introduces an automated robot learning framework that uses generative foundation models to propose tasks, construct simulation scenes, and produce training supervisions in an endless self-guided cycle with minimal human intervention.
Scaling up robotic skill acquisition has historically faced severe bottlenecks due to the slow, risky nature of real-world trials and the labor-intensive requirements of building simulation environments. Traditionally, human engineers must manually design virtual environments, handcraft three-dimensional assets, formulate scene layouts, and construct mathematical reward functions for every new capability. The article introduces RoboGen, an automated agent designed to eliminate these manual constraints by establishing a self-guided propose-generate-learn framework for automated robot learning in simulation.
The primary objective of the article is to demonstrate that foundation and generative models can autonomously configure tasks, environments, and training supervisions to scale robotic skill acquisition with minimal human intervention. Rather than directly using large language models to control robot joints—an approach that often struggles due to a lack of physical grounding—the system uses language and vision models for task imagination, spatial reasoning, and supervision design, while delegating low-level physics execution to physics-grounded simulators.
The authors implemented this approach using a self-guided propose-generate-learn pipeline operating on the differentiable simulation platform Genesis, with OpenAI's GPT-4 as the primary language backend and Gemini-Pro as a visual verifier. The system queries the language model to generate tasks conditioned on robot capabilities, creates environments by retrieving meshes from the Objaverse database or generating them via image-to-3D pipelines, automatically verifies realistic object scales, decomposes tasks into sub-tasks, and selects the most appropriate learning algorithm. RoboGen chooses between reinforcement learning, gradient-based trajectory optimization, and motion planning depending on whether the task involves articulated manipulation, deformable soft-body interaction, or legged locomotion.
Evaluation showed that RoboGen achieved superior task and visual diversity compared to established human-curated robotic benchmarks (such as RLBench, ManiSkill2, Meta-World, and Behavior-100) and concurrent procedural generation systems, as measured by lower text redundancy and image similarity scores. Across a benchmark suite of 69 diverse tasks spanning articulated objects, soft materials, and quadruped locomotion, the automated learning pipeline achieved an overall training success rate of 77.4%. Hybrid algorithm selection proved essential; relying solely on reinforcement learning for articulated manipulation caused nearly all tasks to fail, whereas pairing motion planning with learning enabled reliable execution. Furthermore, automated visual and dimensional verification significantly reduced asset failures, and an analysis of 155 generated tasks identified only 19 failure cases, primarily caused by asset geometry mismatches or language ambiguities in object joint limits.
These findings demonstrate that generative simulation offers a scalable, cost-effective path to producing massive volumes of diverse robotic training data without requiring continuous manual engineering. Organizations developing autonomous systems can substantially reduce development cycles and human overhead by automating task design and reward authoring. However, real-world deployment remains bounded by the simulation-to-reality gap, requiring techniques such as domain randomization to transfer policies safely. Next steps include integrating multimodal environmental feedback loops to autonomously verify learned skills and self-correct reward errors, as well as conducting physical real-robot transfer trials before broad deployment in production environments.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Learn how large language models can generate executable policy code for embodied control, establishing foundational concepts for automated task decomposition and policy specification.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Understand how pretrained language models can extract common-sense knowledge to decompose high-level instructions into executable steps, a core mechanism utilized in generative simulation loops.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Examine how web-scale foundation models transfer semantic knowledge and reasoning directly to physical robotic control policies.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). Review the multi-task transformer architecture and large-scale demonstration collection for robotic manipulation that generative approaches aim to scale without manual teleoperation.
- Paper: Solving Rubik's Cube with a Robot Hand, OpenAI et al. (2019). Discover how automated domain randomization and scalable simulation curricula enable effective sim-to-real transfer for complex robotic manipulation.
- Paper: LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, Bo Liu et al. (2023). Explore standard lifelong learning and procedural task generation benchmarks that provide the context for multi-task robotic skill acquisition.
- Paper: PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation, Jialu Li et al. (2023). See how generative diffusion models can synthesize diverse 3D simulated environments to overcome training data scarcity in embodied navigation.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Study the trajectory diffusion formulation for flexible behavioral synthesis and planning that informs modern policy generation pipelines.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). Extends generative robotics concepts by composing video diffusion world models to imagine and plan zero-shot robotic manipulation tasks.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). Applies large-scale synthetic physics simulations and foundation model architectures to train full-scale generalist humanoid robots.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). Leverages multi-robot foundation modeling and diffusion architectures to scale generalist bimanual manipulation across diverse datasets.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Develops a unified vision-language-action flow matching model trained at scale to generalize dexterity across multiple robot configurations.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). Builds on generalist robot learning frameworks by injecting structured visual-textual affordance reasoning directly into vision-language-action policies.
- Paper: Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations, Yucheng Hu et al. (2025). Incorporates predictive dynamics from video foundation models to guide robotic action learning on multi-task manipulation suites.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). Integrates multi-modal semantic understanding with future visual prediction in a unified vision-language-action framework for manipulation.
- Paper: SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World, Kiana Ehsani et al. (2024). Demonstrates how massive procedural simulation data and automated planners can train mobile robots to navigate and manipulate in the real world.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). Generalizes the vision of generative decision-making by framing video generation as the universal state-action interface and simulator for real-world robotics.
