Position: Video as the New Language for Real-World Decision Making
Sherry YangJacob C. WalkerJack Parker-HolderYilun DuJake BruceAndré BarretoPieter AbbeelDale Schuurmans
Proposes a unified framework that uses conditional video generation as a general-purpose medium for physical reasoning, planning, and simulation across robotics, autonomous driving, and scientific discovery.
Artificial intelligence has advanced rapidly through large language models trained on massive text datasets, but text alone cannot easily capture fine-grained physical dynamics, spatial layouts, and continuous motions. At the same time, publicly available text data is becoming constrained, while vast amounts of internet video remain underutilized beyond media generation. The article argues that video generation can serve as a universal interface for real-world decision making, acting as the visual counterpart to language models in physical domains such as robotics, autonomous driving, and scientific modeling.
The article establishes a conceptual and empirical framework that evaluates video generation as a unified state-action space, task interface, and simulation environment. By reviewing emerging methodologies and training proof-of-concept architectures—including autoregressive, diffusion, and masked models—the source demonstrates how next-frame prediction can solve traditional vision tasks, synthesize execution plans, simulate interactive environments like Minecraft, model complex robot end-effector dynamics, and replicate atomic movements under electron microscopes.
The findings show that diverse computer vision and embodied artificial intelligence tasks can be unified into next-frame prediction using in-context learning. Video models effectively generate realistic robotic execution plans and visual subgoals across multi-robot datasets, resolving long-standing data fragmentation issues. Furthermore, generative video simulators successfully model dynamic systems, such as driving conditions across varied weather and complex physical interactions, offering fixed computational overhead compared to traditional, computationally intractable physics simulations.
These results imply that video generation can significantly reduce the costs, risks, and hardware bottlenecks associated with real-world testing. Simulating environments allows safer policy evaluation for autonomous driving and robotics without real-world safety risks, while bridging the simulation-to-reality gap through natural domain randomization. Combining high-level language reasoning with detailed video-based execution creates a complete pathway from abstract planning to physical control.
To move forward, the article recommends pairing video generation with language models, using external feedback such as human preferences and real-world execution to iteratively improve models, and leveraging multimodal models as automated reward functions. Future efforts must prioritize establishing standardized evaluation metrics by converting generated visual plans into real-world actions and measuring performance gaps.
Confidence in video models as real-world decision engines must remain cautious due to key limitations. Existing internet video datasets lack adequate task-specific coverage and action labels, and current architectures suffer from visual hallucinations, such as disappearing objects, physically implausible dynamics, and poor long-term temporal consistency. Further work in collecting curated domain datasets and refining hybrid architectures is required before deploying these models in mission-critical applications.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). This work demonstrates how autoregressive transformers can model discrete video tokens to capture physical dynamics, establishing the foundational analogy between language modeling and video generation.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). This paper introduces the concept of quantizing video into visual tokens to enable self-supervised sequence learning akin to BERT, directly motivating the treatment of video as a unified language-like interface.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This foundational paper shows that diffusion models can be extended to video prediction and generation across robotics and action datasets, providing the core generative mechanisms discussed in the position paper.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). This study demonstrates how large models pretrained on web data can translate directly to robotic decision-making and physical actions, setting the stage for extending generative visual pretraining to embodied control.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). This work reveals how passive, unlabelled video data can provide universal visual representations and implicit reward signals for decision-making and robotic trajectory optimization without explicit action labels.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). This paper presents an autoregressive video generation framework over discrete visual tokens conditioned on sequential textual prompts, exemplifying how video models can simulate temporal transitions.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). This work systematically benchmarks video generation models as interactive, actionable world simulators for robotics and autonomous driving, directly testing the thesis advocated in the position paper.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). This paper operationalizes video generation as a compositional world model for robotic imagination and action planning across novel object and task configurations.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). This document develops vision-language world models trained on large-scale video footage to conduct both reactive execution and reflective planning for embodied decision-making.
- Paper: Advancing Open-Source World Models, Robbyant Team et al. (2026). This paper realizes the proposal of building real-world dynamic simulators by scaling open-source controllable video world models trained on diverse video and interactive data.
- Paper: Visual General Intelligence: A White Paper, Hirokatsu Kataoka et al. (2026). This comprehensive white paper expands the vision of learning general physical intelligence directly from scaled visual experience and generative video modeling.
- Paper: Visual Planning: Let's Think Only with Images, Yi Xu et al. (2026). This work implements decision-making purely through generated sequences of visual frames rather than language tokens, fulfilling the position paper's concept of visual-only planning.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). This research unifies vision-language understanding with future frame prediction in an autoregressive agent model to enhance robotic action planning with physical dynamics.
