GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
NVIDIAJohan BjorckFernando CastañedaNikita CherniadevXingye DaRunyu DingLinxi "Jim" FanYu FangDieter FoxFengyuan Hu
Presents GR00T N1, an open vision-language-action foundation model for humanoid robots that combines multimodal reasoning with a diffusion transformer to execute real-time, language-guided bimanual manipulation across diverse embodiments.
General-purpose humanoid robots require both versatile physical hardware and broad intelligence to operate reliably across varied real-world environments. However, scaling robot intelligence faces a severe data bottleneck: collecting real-world teleoperation trajectories is expensive and labor-intensive, resulting in fragmented "data islands" across different robot embodiments rather than the unified, web-scale corpora typical of modern artificial intelligence.
The article demonstrates and evaluates GR00T N1, an open foundation model designed for generalist humanoid robots. The model is engineered to interpret vision and language instructions, coordinate multi-embodiment hardware, and generate precise motor actions while maintaining high sample efficiency.
To overcome data scarcity, the authors structured an 8,375-hour training corpus into a multi-tiered data pyramid spanning open-source human egocentric videos, synthetically generated physics simulations, neural video trajectories with predicted pseudo-actions, and real-world teleoperation demonstrations. The technical approach couples a 2.2-billion-parameter dual-system architecture: a slower vision-language reasoning module (System 2, running at 10 Hz) interprets tasks and visual inputs, while a high-frequency diffusion transformer module (System 1, generating action chunks at 120 Hz) executes fluid physical trajectories across diverse embodiments using flow-matching and learned latent actions.
Experimental results show that GR00T N1 consistently outperforms established imitation learning baselines across both simulation and physical hardware. In real-world bimanual manipulation tests on the Fourier GR-1 humanoid robot, GR00T N1 achieved an average success rate of 76.8% with full data, outperforming the Diffusion Policy baseline by 30.4 percentage points. When restricted to a low-data regime using only 10% of the dataset, the model achieved a 42.6% success rate, outperforming the baseline by 32.4 percentage points and nearly matching the baseline's full-data performance (46.4%). Across three simulation suites (RoboCasa, DexMimicGen, and GR-1 Tabletop), it achieved an overall average success rate of 45.0% compared to 33.4% for Diffusion Policy and 26.4% for Behavior Cloning Transformer. Furthermore, co-training with synthetically generated neural trajectories boosted simulation performance by up to 8.8 percentage points and real-world task performance by 5.8 percentage points.
These findings indicate that large-scale pre-training across heterogeneous synthetic and human data substantially reduces the volume of costly on-robot teleoperation demonstrations needed to deploy humanoid policies. The architecture enables rapid adaptation to complex, multi-agent, and bimanual tasks while improving execution smoothness, object localization, and semantic task comprehension.
Organizations developing or deploying robotic automation should adopt multi-source data generation pipelines—incorporating automated physics simulations and neural video generation—to scale training while reducing physical collection overhead. Future work should focus on extending the architecture to long-horizon loco-manipulation and improving video generation models to ensure generated trajectories strictly adhere to physical laws.
Confidence in these findings is supported by rigorous evaluations across multiple simulation benchmarks and physical robot rollouts. However, current results remain bounded to short-horizon tabletop manipulation setups, and readers should exercise caution before generalizing performance to unconstrained mobile environments or long-horizon tasks requiring continuous full-body mobility.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA establishes the foundational open-source Vision-Language-Action architecture and fine-tuning recipe that GR00T N1 builds upon for generalist robot manipulation.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). This paper introduces the combination of vision-language model backbones with flow matching continuous action heads across diverse multi-robot embodiments, directly informing GR00T N1's dual-system architecture.
- Paper: Octo: An Open-Source Generalist Robot Policy, O. Team et al. (2024). Octo introduces the open-source transformer architecture paired with continuous diffusion action decoders trained on multi-robot datasets that directly precedes GR00T N1's diffusion transformer module.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 provides the core paradigm of co-fine-tuning vision-language foundation models on internet-scale multimodal data and robotic trajectories to enable semantic reasoning in physical control.
- Paper: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware, Tony Z. Zhao et al. (2023). This work introduces Action Chunking with Transformers (ACT) and benchmarks for bimanual manipulation, which serve as foundational baselines and task paradigms evaluated in GR00T N1.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). PaLM-E demonstrates how embodied multimodal language models can inject sensor inputs directly into language representations for high-level planning and decision-making.
- Paper: RT-1: Robotics Transformer for Real-World Control at Scale, Anthony Brohan et al. (2023). RT-1 pioneers the use of scaled multi-task transformer models for real-world robotic control from vision and text instructions.
- Paper: LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning, Bo Liu et al. (2023). LIBERO provides the standard benchmark suite and methodology for evaluating transfer and generalization in language-conditioned manipulation policies.
- Paper: RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation, Songming Liu et al. (2025). RDT-1B scales diffusion-based foundation models specifically for bimanual manipulation using a unified physical action space across large multi-robot corpora.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). π0.5 extends vision-language-action architectures with hierarchical training strategies to achieve robust zero-shot open-world generalization in novel domestic settings.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). TwinBrainVLA addresses the catastrophic forgetting of semantic reasoning in VLA models through an asymmetric dual-pathway mixture-of-transformers architecture.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA enhances the reasoning capabilities of VLA models by integrating structured, sequential visual-textual affordance predictions into action generation.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). DynamicVLA adapts VLA control paradigms to dynamic environments by employing continuous inference and latent-aware action streaming for moving object manipulation.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This survey contextualizes GR00T N1 within the broader landscape of modern vision-language-action architectures, training strategies, and real-world deployment challenges.
