GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

NVIDIAJohan BjorckFernando CastañedaNikita CherniadevXingye DaRunyu DingLinxi "Jim" FanYu FangDieter FoxFengyuan Hu

article2025arXiv1,342 citationsBest Paper Award Finalist, Best Systems Paper Award Finalist

Presents GR00T N1, an open vision-language-action foundation model for humanoid robots that combines multimodal reasoning with a diffusion transformer to execute real-time, language-guided bimanual manipulation across diverse embodiments.

Listen

General-purpose humanoid robots require both versatile physical hardware and broad intelligence to operate reliably across varied real-world environments. However, scaling robot intelligence faces a severe data bottleneck: collecting real-world teleoperation trajectories is expensive and labor-intensive, resulting in fragmented "data islands" across different robot embodiments rather than the unified, web-scale corpora typical of modern artificial intelligence.

The article demonstrates and evaluates GR00T N1, an open foundation model designed for generalist humanoid robots. The model is engineered to interpret vision and language instructions, coordinate multi-embodiment hardware, and generate precise motor actions while maintaining high sample efficiency.

To overcome data scarcity, the authors structured an 8,375-hour training corpus into a multi-tiered data pyramid spanning open-source human egocentric videos, synthetically generated physics simulations, neural video trajectories with predicted pseudo-actions, and real-world teleoperation demonstrations. The technical approach couples a 2.2-billion-parameter dual-system architecture: a slower vision-language reasoning module (System 2, running at 10 Hz) interprets tasks and visual inputs, while a high-frequency diffusion transformer module (System 1, generating action chunks at 120 Hz) executes fluid physical trajectories across diverse embodiments using flow-matching and learned latent actions.

Experimental results show that GR00T N1 consistently outperforms established imitation learning baselines across both simulation and physical hardware. In real-world bimanual manipulation tests on the Fourier GR-1 humanoid robot, GR00T N1 achieved an average success rate of 76.8% with full data, outperforming the Diffusion Policy baseline by 30.4 percentage points. When restricted to a low-data regime using only 10% of the dataset, the model achieved a 42.6% success rate, outperforming the baseline by 32.4 percentage points and nearly matching the baseline's full-data performance (46.4%). Across three simulation suites (RoboCasa, DexMimicGen, and GR-1 Tabletop), it achieved an overall average success rate of 45.0% compared to 33.4% for Diffusion Policy and 26.4% for Behavior Cloning Transformer. Furthermore, co-training with synthetically generated neural trajectories boosted simulation performance by up to 8.8 percentage points and real-world task performance by 5.8 percentage points.

These findings indicate that large-scale pre-training across heterogeneous synthetic and human data substantially reduces the volume of costly on-robot teleoperation demonstrations needed to deploy humanoid policies. The architecture enables rapid adaptation to complex, multi-agent, and bimanual tasks while improving execution smoothness, object localization, and semantic task comprehension.

Organizations developing or deploying robotic automation should adopt multi-source data generation pipelines—incorporating automated physics simulations and neural video generation—to scale training while reducing physical collection overhead. Future work should focus on extending the architecture to long-horizon loco-manipulation and improving video generation models to ensure generated trajectories strictly adhere to physical laws.

Confidence in these findings is supported by rigorous evaluations across multiple simulation benchmarks and physical robot rollouts. However, current results remain bounded to short-horizon tabletop manipulation setups, and readers should exercise caution before generalizing performance to unconstrained mobile environments or long-horizon tasks requiring continuous full-body mobility.

Cover for GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Abstract

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

Table of Contents

  • 1 Introduction
  • 2 GR00T N1 Foundation Model
  • 2.1 Model Architecture
  • 2.2 Training Data Generation
  • 2.3 Training Details
  • 3 Pre-Training Datasets
  • 3.1 Real-World Datasets
  • 3.2 Synthetic Datasets
  • 3.3 Human Video Datasets
  • 4 Evaluation
  • 4.1 Simulation Benchmarks
  • 4.2 Real-World Benchmarks
  • 4.3 Experiment Setup
  • 4.4 Quantitative Results
  • 4.5 Qualitative Results
  • 4.6 Limitations
  • 5 Related Work
  • 6 Conclusions
  • A Contributors and Acknowledgments
  • A.1 Core Contributors
  • A.2 Contributors
  • A.3 Acknowledgments
  • B Detailed Experiment Results
  • C Additional Qualitative Results
  • D Hyperparameters
  • E System Design
  • E.1 Dataset Formats
  • E.2 Standardized Action Spaces
  • F Additional Training Details
  • References

Citation

MLA
NVIDIA, et al. “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots”. arXiv, 2025, http://arxiv.org/abs/2503.14734v2.
APA
NVIDIA, :, Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L. "Jim" ., Fang, Y., Fox, D., Hu, F., Huang, S., Jang, J., Jiang, Z., Kautz, J., Kundalia, K., Lao, L., Li, Z., Lin, Z., … Zhu, Y. (2025). GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv. http://arxiv.org/abs/2503.14734v2
Chicago
NVIDIA, :, J. Bjorck, et al. 2025. “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots”. arXiv. http://arxiv.org/abs/2503.14734v2.
Harvard
NVIDIA et al. (2025) “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.14734v2.
Vancouver
1. NVIDIA, :, Bjorck J, et al (2025) GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv

BibTeX

@article{nvidia2025gr00t,
  title = {GR00T N1: An Open Foundation Model for Generalist Humanoid Robots},
  author = {\{NVIDIA\} and \{:\} and Bjorck, Johan and Castañeda, Fernando and Cherniadev, Nikita and Da, Xingye and Ding, Runyu and Fan, Linxi "Jim" and Fang, Yu and Fox, Dieter and Hu, Fengyuan and Huang, Spencer and Jang, Joel and Jiang, Zhenyu and Kautz, Jan and Kundalia, Kaushil and Lao, Lawrence and Li, Zhiqi and Lin, Zongyu and Lin, Kevin and Liu, Guilin and Llontop, Edith and Magne, Loic and Mandlekar, Ajay and Narayan, Avnish and Nasiriany, Soroush and Reed, Scott and Tan, You Liang and Wang, Guanzhi and Wang, Zu and Wang, Jing and Wang, Qi and Xiang, Jiannan and Xie, Yuqi and Xu, Yinzhen and Xu, Zhenjia and Ye, Seonghyeon and Yu, Zhiding and Zhang, Ao and Zhang, Hao and Zhao, Yizhou and Zheng, Ruijie and Zhu, Yuke},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.14734v2},
  eprint = {2503.14734}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/