Genie: Generative Interactive Environments
Jake BruceMichael D. DennisAshley EdwardsJack Parker-HolderYuge ShiEdward HughesMatthew LaiAditi MavalankarRichie SteigerwaldChris Apps
Presents Genie, an 11-billion-parameter foundation world model that learns controllable interactive environments and latent actions directly from unlabeled internet video, enabling users to generate and play virtual worlds from prompts as simple as a single image or sketch.
Recent advances in generative artificial intelligence have enabled high-quality text, image, and passive video creation, yet generating interactive, controllable visual environments typically requires domain-specific engines or costly ground-truth action labels. This reliance on labeled interaction data creates a bottleneck for scalable world modeling and agent development. The article addresses this challenge by evaluating whether a foundation world model can learn frame-by-frame controllable virtual environments in an entirely unsupervised manner using unlabelled Internet videos alone.
To achieve this, the article introduces Genie, an 11-billion parameter generative interactive model built upon a memory-efficient spatiotemporal transformer architecture. The framework comprises three modular components: a video tokenizer that converts raw video frames into discrete tokens, a latent action model that infers discrete control codes directly from video pixels without human labels, and a dynamics model that autoregressively predicts future frames based on past tokens and chosen actions. The model was primarily trained on a curated corpus of 30,000 hours (6.8 million clips) of 2D platformer gameplay filtered from 55 million raw videos, alongside additional experiments on unlabelled robotics demonstration datasets.
Key findings demonstrate strong scaling performance and robust interactive generation across diverse settings. First, systematic scaling from 40 million to 2.7 billion parameters, alongside batch size increases, produced consistent reductions in training loss, validating the architecture's capacity to scale effectively to 11 billion parameters. Second, the model generalizes zero-shot to out-of-distribution prompts, generating controllable, physics-consistent interactive worlds from text-to-image outputs, sketches, and real-world photographs while preserving emergent 3D properties like parallax. Third, in robotics domains, the architecture learned consistent arm manipulation and object deformation controls from action-free video. Finally, agents trained using latent actions derived from the model matched the performance of oracle behavioral cloning baselines on unseen platformer environments while requiring as few as 200 labeled real-world transition samples for action mapping.
These findings suggest that unlabelled Internet video can serve as a virtually unlimited source for training interactive world simulators and generalist artificial intelligence agents, bypassing the cost and constraints of manual action logging. However, practical deployment faces current operational limitations, including a 16-frame context window that challenges long-horizon consistency, occasional physical hallucinations, and an inference speed of approximately one frame per second. Before enterprise adoption in interactive gaming or live robotics simulation, stakeholders should invest in research to improve frame rates, extend temporal memory horizons, and conduct pilot validations in targeted agent pre-training workflows.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). VideoGPT introduces the foundational two-stage paradigm of tokenizing video via discrete autoencoders and modeling dynamics autoregressively with transformers, directly underpinning Genie's core architecture.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Phenaki establishes causal, spatiotemporal tokenization and autoregressive modeling for variable-length video generation, a core technical concept upon which Genie builds its interactive video world model.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Dreamer pioneers the concept of training reinforcement learning agents within learned latent world models, providing the conceptual foundation for Genie's use as an environment for downstream agent training.
- Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). Parti demonstrates that scaling discrete visual token prediction with autoregressive transformers achieves state-of-the-art visual generation, motivating Genie's large-scale autoregressive dynamics formulation.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work establishes the discrete VAE and transformer-based autoregressive visual generation paradigm that Genie adopts and scales to spatiotemporal video latents.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). This position paper conceptualizes and expands Genie's premise of using generative video prediction as a universal state-action interface and world model for decision-making.
- Paper: Advancing Open-Source World Models, Robbyant Team et al. (2026). LingBot-World directly advances the interactive world simulator paradigm established by Genie into an open-source, multi-minute, real-time interactive framework.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). RoboDreamer builds on video-based world models like Genie by factorizing dynamics into compositional language and visual instructions for robotic planning and execution.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). VLWM extends video foundation world models by combining predictive dynamics with language reasoning and dual-system planning for embodied task execution.
- Paper: LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, Lucas Maes et al. (2026). LeWorldModel explores sample-efficient, pixel-based latent dynamics prediction for online model-predictive control, offering an alternative architecture for agent planning.
- Paper: RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation, Yufei Wang et al. (2024). RoboGen utilizes generative foundation models to automatically construct interactive environments and tasks for automated robot learning in simulation.
