Genie: Generative Interactive Environments

Jake BruceMichael D. DennisAshley EdwardsJack Parker-HolderYuge ShiEdward HughesMatthew LaiAditi MavalankarRichie SteigerwaldChris Apps

article2024ICML485 citationsBest Paper Award

Presents Genie, an 11-billion-parameter foundation world model that learns controllable interactive environments and latent actions directly from unlabeled internet video, enabling users to generate and play virtual worlds from prompts as simple as a single image or sketch.

Listen

Recent advances in generative artificial intelligence have enabled high-quality text, image, and passive video creation, yet generating interactive, controllable visual environments typically requires domain-specific engines or costly ground-truth action labels. This reliance on labeled interaction data creates a bottleneck for scalable world modeling and agent development. The article addresses this challenge by evaluating whether a foundation world model can learn frame-by-frame controllable virtual environments in an entirely unsupervised manner using unlabelled Internet videos alone.

To achieve this, the article introduces Genie, an 11-billion parameter generative interactive model built upon a memory-efficient spatiotemporal transformer architecture. The framework comprises three modular components: a video tokenizer that converts raw video frames into discrete tokens, a latent action model that infers discrete control codes directly from video pixels without human labels, and a dynamics model that autoregressively predicts future frames based on past tokens and chosen actions. The model was primarily trained on a curated corpus of 30,000 hours (6.8 million clips) of 2D platformer gameplay filtered from 55 million raw videos, alongside additional experiments on unlabelled robotics demonstration datasets.

Key findings demonstrate strong scaling performance and robust interactive generation across diverse settings. First, systematic scaling from 40 million to 2.7 billion parameters, alongside batch size increases, produced consistent reductions in training loss, validating the architecture's capacity to scale effectively to 11 billion parameters. Second, the model generalizes zero-shot to out-of-distribution prompts, generating controllable, physics-consistent interactive worlds from text-to-image outputs, sketches, and real-world photographs while preserving emergent 3D properties like parallax. Third, in robotics domains, the architecture learned consistent arm manipulation and object deformation controls from action-free video. Finally, agents trained using latent actions derived from the model matched the performance of oracle behavioral cloning baselines on unseen platformer environments while requiring as few as 200 labeled real-world transition samples for action mapping.

These findings suggest that unlabelled Internet video can serve as a virtually unlimited source for training interactive world simulators and generalist artificial intelligence agents, bypassing the cost and constraints of manual action logging. However, practical deployment faces current operational limitations, including a 16-frame context window that challenges long-horizon consistency, occasional physical hallucinations, and an inference speed of approximately one frame per second. Before enterprise adoption in interactive gaming or live robotics simulation, stakeholders should invest in research to improve frame rates, extend temporal memory horizons, and conduct pilot validations in targeted agent pre-training workflows.

arXiv: 2402.15391
Cover for Genie: Generative Interactive Environments

Abstract

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a foundation world model. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis despite training without any ground-truth action labels or other domain-specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future.

Citation

MLA
Bruce, J., et al. “Genie: Generative Interactive Environments”. arXiv, 2024, http://arxiv.org/abs/2402.15391v1.
APA
Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., Aytar, Y., Bechtle, S., Behbahani, F., Chan, S., Heess, N., Gonzalez, L., Osindero, S., Ozair, S., Reed, S., … Rocktäschel, T. (2024). Genie: Generative Interactive Environments. arXiv. http://arxiv.org/abs/2402.15391v1
Chicago
Bruce, J., M. Dennis, A. Edwards, et al. 2024. “Genie: Generative Interactive Environments”. arXiv. http://arxiv.org/abs/2402.15391v1.
Harvard
Bruce, J. et al. (2024) “Genie: Generative Interactive Environments”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.15391v1.
Vancouver
1. Bruce J, Dennis M, Edwards A, et al (2024) Genie: Generative Interactive Environments. arXiv

BibTeX

@article{bruce2024genie,
  title = {Genie: Generative Interactive Environments},
  author = {Bruce, Jake and Dennis, Michael and Edwards, Ashley and Parker-Holder, Jack and Shi, Yuge and Hughes, Edward and Lai, Matthew and Mavalankar, Aditi and Steigerwald, Richie and Apps, Chris and Aytar, Yusuf and Bechtle, Sarah and Behbahani, Feryal and Chan, Stephanie and Heess, Nicolas and Gonzalez, Lucy and Osindero, Simon and Ozair, Sherjil and Reed, Scott and Zhang, Jingwei and Zolna, Konrad and Clune, Jeff and Freitas, Nando de and Singh, Satinder and Rocktäschel, Tim},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.15391v1},
  eprint = {2402.15391}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/