Position: Video as the New Language for Real-World Decision Making

Sherry YangJacob C. WalkerJack Parker-HolderYilun DuJake BruceAndré BarretoPieter AbbeelDale Schuurmans

article2024ICML93 citations

Proposes a unified framework that uses conditional video generation as a general-purpose medium for physical reasoning, planning, and simulation across robotics, autonomous driving, and scientific discovery.

Listen

Artificial intelligence has advanced rapidly through large language models trained on massive text datasets, but text alone cannot easily capture fine-grained physical dynamics, spatial layouts, and continuous motions. At the same time, publicly available text data is becoming constrained, while vast amounts of internet video remain underutilized beyond media generation. The article argues that video generation can serve as a universal interface for real-world decision making, acting as the visual counterpart to language models in physical domains such as robotics, autonomous driving, and scientific modeling.

The article establishes a conceptual and empirical framework that evaluates video generation as a unified state-action space, task interface, and simulation environment. By reviewing emerging methodologies and training proof-of-concept architectures—including autoregressive, diffusion, and masked models—the source demonstrates how next-frame prediction can solve traditional vision tasks, synthesize execution plans, simulate interactive environments like Minecraft, model complex robot end-effector dynamics, and replicate atomic movements under electron microscopes.

The findings show that diverse computer vision and embodied artificial intelligence tasks can be unified into next-frame prediction using in-context learning. Video models effectively generate realistic robotic execution plans and visual subgoals across multi-robot datasets, resolving long-standing data fragmentation issues. Furthermore, generative video simulators successfully model dynamic systems, such as driving conditions across varied weather and complex physical interactions, offering fixed computational overhead compared to traditional, computationally intractable physics simulations.

These results imply that video generation can significantly reduce the costs, risks, and hardware bottlenecks associated with real-world testing. Simulating environments allows safer policy evaluation for autonomous driving and robotics without real-world safety risks, while bridging the simulation-to-reality gap through natural domain randomization. Combining high-level language reasoning with detailed video-based execution creates a complete pathway from abstract planning to physical control.

To move forward, the article recommends pairing video generation with language models, using external feedback such as human preferences and real-world execution to iteratively improve models, and leveraging multimodal models as automated reward functions. Future efforts must prioritize establishing standardized evaluation metrics by converting generated visual plans into real-world actions and measuring performance gaps.

Confidence in video models as real-world decision engines must remain cautious due to key limitations. Existing internet video datasets lack adequate task-specific coverage and action labels, and current architectures suffer from visual hallucinations, such as disappearing objects, physically implausible dynamics, and poor long-term temporal consistency. Further work in collecting curated domain datasets and refining hybrid architectures is required before deploying these models in mission-critical applications.

arXiv: 2402.17139
Cover for Position: Video as the New Language for Real-World Decision Making

Abstract

Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world impact, whereas video generation has remained largely limited to media entertainment. Yet video data captures important information about the physical world that is difficult to express in language. To address this gap, we discuss an under-appreciated opportunity to extend video generation to solve tasks in the real world. We observe how, akin to language, video can serve as a unified interface that can absorb internet knowledge and represent diverse tasks. Moreover, we demonstrate how, like language models, video generation can serve as planners, agents, compute engines, and environment simulators through techniques such as in-context learning, planning and reinforcement learning. We identify major impact opportunities in domains such as robotics, self-driving, and science, supported by recent work that demonstrates how such advanced capabilities in video generation are plausibly within reach. Lastly, we identify key challenges in video generation that mitigate progress. Addressing these challenges will enable video generation models to demonstrate unique value alongside language models in a wider array of AI applications.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Conditional Video Generation
  • 2.2 Task-Specific Specialization
  • 3 Unified Representation and Task Interface
  • 3.1 Video as a Unified Representation of Information
  • 3.2 Video Generation as a Unified Task Interface
  • 3.3 Video as a Unified State-Action Space
  • 4 Video Generation as Simulation
  • 4.1 Generative Game Environments
  • 4.2 Robotics and Self-Driving.
  • 4.3 Science and Engineering
  • 5 Challenges
  • 5.1 Dataset Limitations
  • 5.2 Model Heterogeneity
  • 5.3 Hallucination
  • 5.4 Limited Generalization
  • 6 Discussion
  • 7 Conclusion
  • 8 Impact Statement
  • References
  • Appendix
  • A Details of Models Used to Generate Examples in the Main Text
  • A.1 Autoregressive Model
  • A.2 Diffusion Model
  • A.3 Masked Model
  • B Additional Generated Videos
  • B.1 Additional Game Simulations
  • B.2 Additional Generation for How-to Videos
  • B.3 Additional Self-Driving Simulations
  • B.4 Additional Robot SE(3) Simulations
  • B.5 Examples of Hallucination

Knowls

  1. Knowl 1 — Video as a Unified Representation and Interface for Physical Decision Making

    model/method

    While large language models leverage text as a unified representation and task interface for digital and abstract reasoning, natural language cannot fully specify low-level spatial arrangements, fine-grained physical mechanics (such as friction or torque), or precise continuous motor trajectories. Video provides an internet-scale, human-interpretable representation capable of capturing physical dynamics across arbitrary temporal and spatial scales (e.g., electron microscopy at 10−10 m10^{-10}\text{ m} to macroscopic phenomena).

    Under this paradigm, video generation acts as the physical-world analog to language modeling by serving four core roles:

    1. Task Solver / Goal Generator: Synthesizing visual target states xtx_t or sequence rollouts to answer how-to inquiries or specify visual manipulation goals.
    2. Policy / Agent: Directly predicting next frames that implicitly encode the visual execution of actions.
    3. World Model / Simulator: Simulating forward visual dynamics conditioned on candidate control inputs to enable trajectory optimization, policy evaluation, and reinforcement learning without requiring rigid analytical physics engines.
    4. Algorithmic Compute Engine: Emulating algorithmic execution states (such as graph search traversals) via sequential visual state updates.
  2. Knowl 2 — Conditional Video Generation Formulations for Real-World Decision Tasks

    equation

    Conditional video generation models the probability distribution p(x∣c)p(\mathbf{x}|c) over a video sequence x=(x0,x1,…,xt)\mathbf{x} = (x_0, x_1, \dots, x_t), where each xix_i is an image frame and cc is a conditioning variable. Real-world decision-making and computer vision tasks are mapped into specific forms of p(x∣c)p(\mathbf{x}|c):

    1. Text-to-Video Synthesis: p(x∣c=text)p(\mathbf{x} \mid c = \text{text}) Used for unconditional environment generation or visualizing text descriptions.

    2. Visual Planning and Goal Generation: p(x∣c={x0,text})orp(xt∣c={x0,text})p(\mathbf{x} \mid c = \{x_0, \text{text}\}) \quad \text{or} \quad p(x_t \mid c = \{x_0, \text{text}\}) Generates rollout plans or visual subgoals xtx_t starting from an initial scene observation x0x_0 guided by a natural language instruction.

    3. Text-Guided Dynamics Editing and Visual Prompting: p(x∣c={x~,text})p(\mathbf{x} \mid c = \{\tilde{\mathbf{x}}, \text{text}\}) Applies environmental transformations (e.g., weather or lighting domain shifts in autonomous driving) to a source video x~\tilde{\mathbf{x}}, or elicits specific task behaviors using in-context visual prompt frames.

    4. Action-Conditioned Visual Dynamics (World Models): p(xi+1∣c={xi,ai})orp(xi+k∣c={xi:i+k−1,ai})p(x_{i+1} \mid c = \{x_i, a_i\}) \quad \text{or} \quad p(x_{i+k} \mid c = \{x_{i:i+k-1}, a_i\}) Predicts future visual observations given the current state xix_i (or historical sequence xi:i+k−1x_{i:i+k-1}) and an action aia_i (e.g., robot end-effector control, discrete keyboard input, or latent action), supporting both single-step and temporally abstract (k>1k > 1) simulation.

  3. Knowl 3 — Pixel Space as a Unified State-Action Space for Robot Learning

    model/method

    Robot datasets collected across diverse platforms typically suffer from data fragmentation due to incompatible proprioceptive state spaces and morphology-specific action parameterizations. Casting planning into the pixel space provides a morphology-agnostic unified state-action space.

    Under this framework:

    1. A conditional video model trained across heterogeneous datasets (e.g., the Open X-Embodiment dataset) generates forward video plans (x0,x1,…,xT)(x_0, x_1, \dots, x_T) from an initial camera view x0x_0 and a language goal.
    2. Multi-view consistency (e.g., third-person camera vs. robot wrist-mounted camera) is maintained implicitly by conditioning generation on shared action representations.
    3. Low-level executable robot control commands are recovered from the high-level visual plan using an auxiliary grounding module, such as an inverse dynamics model f(xt,xt+1)→atf(x_t, x_{t+1}) \to a_t, a goal-conditioned policy π(at∣xt,xt+1)\pi(a_t \mid x_t, x_{t+1}), an optical flow estimation network, or a dense point-tracking module.
  4. Knowl 4 — Interleaved Autoregressive-MaskGIT Architecture for Interactive World Modeling and Policy Execution

    model/method

    An interactive agent-environment system is implemented as a single autoregressive sequence model over interleaved visual observations and actions. Given visual frames xtx_t and actions ata_t:

    1. Tokenization: Observations xtx_t are encoded into discrete spatial tokens ztz_t using a Vector Quantized Variational Autoencoder (VQ-VAE) combined with a Vision Transformer (ViT). Continuous or complex actions ata_t are tokenized using Video PreTraining (VPT) action discretization.
    2. Sequence Processing: The token sequence (z0,a0,z1,a1,… )(z_0, a_0, z_1, a_1, \dots) is modeled temporally using a Transformer-XL backbone with cached memory states hzth_{z_t} and hath_{a_t}. The sequence memory is initialized using a past encoder trained on continuous past inputs (x−k,a−k,…,x−1,a−1)(x_{-k}, a_{-k}, \dots, x_{-1}, a_{-1}).
    3. Dual Policy and World-Model Operation:
      • Policy Mode: When the sequence ends on observation hidden state hzth_{z_t}, an autoregressive transformer head predicts the next action ata_t.
      • World Model Mode: When the sequence ends on action hidden state hath_{a_t}, a MaskGIT head predicts the next observation tokens zt+1z_{t+1} across 8 parallel decoding steps using a cosine masking schedule.
  5. Knowl 5 — Unsupervised Latent-Action Masked Dynamics for Interactive Environment Simulation

    model/method

    To build controllable interactive visual simulators directly from unannotated gameplay or real-world video without ground-truth action labels, the system learns latent action variables jointly with a masked visual dynamics model:

    1. Latent Action Extraction: Given video frame sequences x1:Tx_{1:T}, a causal transformer quantizes inter-frame transitions into a discrete codebook of unsupervised latent actions a~1:T−1\tilde{a}_{1:T-1} via a VQ-VAE.
    2. Dynamics Training: A spatial-temporal attention transformer receives context tokens from preceding frames x1:t−1x_{1:t-1} and the current latent action a~t\tilde{a}_t. Target frame tokens xtx_t undergo random Bernoulli masking at an average rate of 75%, and the model is trained with a cross-entropy reconstruction loss.
    3. Parallel Iterative Inference: At generation time, beginning with fully masked tokens for xtx_t (comprising 920 discrete tokens per frame), the model predicts token logits conditioned on past context and the selected latent action. The highest-probability tokens are iteratively locked in over 25 MaskGIT sampling steps, allowing real-time interactive simulation from single-image prompts.
  6. Knowl 6 — Cascaded Pixel-Space 3D U-Net Video Diffusion for Real-World Physical Dynamics

    model/method

    Continuous physical processes, robotic manipulations, and scientific dynamics are simulated directly in pixel space using a cascaded 3D U-Net diffusion architecture:

    1. Architecture: The base diffusion model uses interleaved 3D spatial/temporal convolutions and 3D attention mechanisms in downsampling and upsampling passes, with skip connections between corresponding resolutions.
    2. Resolution Cascading: To prevent overfitting on uninformative high-frequency visual noise during base training, generation is structured into a multi-stage pipeline: a base diffusion model generates video at [24,40][24, 40] spatial resolution, followed by two successive spatial super-resolution diffusion models raising output resolution to [48,80][48, 80] and [192,320][192, 320].
    3. Conditioning: Action-conditioning (such as discretized SE(3) end-effector poses) and text conditioning are incorporated via classifier-free guidance. For initial-frame conditioning (x0x_0), the context image is fed into both the conditional and unconditional branches during sampling.
    4. Applications: This framework simulates 6-DoF SE(3) robot arm actions, multi-weather autonomous driving rollouts via domain randomization, and atomic-level silicon dopant migration under scanning transmission electron microscope (STEM) beam excitation.
  7. Knowl 7 — Reformulation of Computer Vision Tasks and Algorithmic Search as Video Generation

    model/method

    Heterogeneous vision tasks and discrete algorithmic reasoning procedures can be cast as sequential frame generation using in-context visual prompting:

    1. Vision Task Unification: Tasks including semantic segmentation, depth estimation, surface normal estimation, pose estimation, and edge detection are standardized by:
      • Mapping all task input formats and output label spaces into a standard RGB pixel space (e.g., depth maps or segmentation masks rendered as RGB images).
      • Interleaving inputs and targets into paired sequences (xin(1),xout(1),xin(2),xout(2),…,xin(k))(x_{\text{in}}^{(1)}, x_{\text{out}}^{(1)}, x_{\text{in}}^{(2)}, x_{\text{out}}^{(2)}, \dots, x_{\text{in}}^{(k)}).
      • Sampling the next frame xout(k)x_{\text{out}}^{(k)} via conditional next-frame prediction using in-context prompt pairs to specify the task.
    2. Algorithmic Search as Video Rollout: Algorithmic execution (such as Breadth-First Search on a 2D grid) is represented as a video sequence where each consecutive frame renders the updated execution state (start/goal points, obstacles, and the updated frontier of visited nodes). Generating the video rollout corresponds directly to computing the search trajectory.
  8. Knowl 8 — Comparative Trade-offs of Video Generation Model Architectures for Decision Making

    model/method

    Video world models and decision engines exhibit structural trade-offs across three primary generative model families:

    1. Diffusion Models:
      • Strengths: Operate directly on continuous output spaces without vector quantization; generate high-fidelity frames; allow parallel multi-frame sampling.
      • Weaknesses: Slow sampling speeds hinder real-time closed-loop simulation; training is sensitive to noise schedules; long-horizon autoregressive rollout stability remains difficult.
    2. Autoregressive Sequence Models:
      • Strengths: Naturally interleave multi-modal discrete tokens (language, VPT actions, visual codes); straightforward training dynamics; scale effectively with sequence length via Transformer-XL.
      • Weaknesses: Sequential token-by-token decoding imposes high computational latency; rollouts suffer from visual compounding error and drifting effects.
    3. Masked Generative Transformers (MaskGIT):
      • Strengths: Fast inference by generating hundreds of spatial tokens in parallel over a few iterative refinement steps (e.g., 8–25 steps).
      • Weaknesses: Potential sampling bias due to the conditional independence assumption among tokens predicted simultaneously within a single sampling step.
  9. Knowl 9 — Mechanisms of Hallucination in Video Decision Models

    limitation

    Video models applied to decision making and simulation exhibit distinct failure modes categorized as visual and dynamic hallucinations:

    1. Object Disappearance and Popping: Small, task-critical objects (e.g., manipulated tools or containers) randomly vanish or appear mid-trajectory. This occurs because standard pixel-level reconstruction losses assign low aggregate weight to small objects relative to dominant static background pixels.
    2. Kinematic and Dynamic Implausibility: Generative rollouts synthesize unphysical state transitions (e.g., objects jumping directly into a robot gripper without contact). This is caused by coarse video frame rates that omit high-frequency, contact-critical physical interactions.
    3. Confounding of Agency and Dynamics: Models trained jointly on agent actions and environmental transitions fail to disentangle visual changes caused by external dynamics from changes induced by autonomous agent actions.
    4. Violation of Temporal Causality: Failure to enforce strict causal ordering in spatial-temporal attention layers leads to premature rendering of future goal states prior to executing intermediate prerequisite actions.
  10. Knowl 10 — Hierarchical Integration of Language and Video Generative Models for Decision Making

    model/method

    Real-world decision making requires coupling language models and video generation models to bridge high-level reasoning and low-level physical execution:

    1. Hierarchical Planning Pipeline: An LLM or Vision-Language Model (VLM) decomposes an abstract long-horizon task specification (e.g., "make sushi") into a structured sequence of discrete subgoals (e.g., "place rice on rolling mat"). A conditional video model subsequently generates the concrete, low-level visual trajectory for each subgoal.
    2. Verification and Reward Grounding: A pre-trained VLM serves as a reward model and success detector by scoring the physical plausibility and goal alignment of generated video rollouts.
    3. Closed-Loop Grounding with Real-World Feedback: Video plans are converted to low-level robot actions and executed in the physical environment. Execution discrepancies between the imagined video rollout and real sensory feedback are utilized to update the world model and mitigate compounding errors.

Coverage note — Omitted purely illustrative qualitative examples (e.g., specific generated frames of origami folding or tie knotting) and standard background on general NLP/RL history, retaining all conceptual formulations, model architectures, task representations, and failure analyses.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp. 2425–2433, 2015.
  3. 3.Antonoglou, I., Schrittwieser, J., Ozair, S., Hubert, T. K., and Silver, D. Planning in stochastic environments with a learned model. In International Conference on Learning Representations. ICLR, 2022.
  4. 4.Bai, Y., Geng, X., Mangalam, K., Bar, A., Yuille, A., Darrell, T., Malik, J., and Efros, A. A. Sequential modeling enables scalable learning for large vision models. arXiv preprint arXiv:2312.00785, 2023.
  5. 5.Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J. Video pretraining (vpt): Learning to act by watching unlabeled online videos. Advances in Neural Information Processing Systems, 35:24639–24654, 2022.
  6. 6.Bamford, C. and Lucas, S. M. Neural game engine: Accurate learning of generalizable forward models from pixels. In 2020 IEEE Conference on Games (CoG), pp. 81–88, 2020. doi: 10.1109/CoG47356.2020.9231688.
  7. 7.Bar, A., Gandelsman, Y., Darrell, T., Globerson, A., and Efros, A. Visual prompting via image inpainting. Advances in Neural Information Processing Systems, 35: 25005–25017, 2022.
  8. 8.Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Li, Y., Michaeli, T., et al. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024.
  9. 9.Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, 2013.
  10. 10.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. Improving image generation with better captions. Computer Science. https://cdn.openai.com/papers/dall-e-3.pdf, 2:3, 2023.
  11. 11.Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023.
  12. 12.Blanco-Claraco, J. L. A tutorial on se(3) transformation parameterizations and on-manifold optimization. arXiv preprint arXiv:2103.15980, 2021.
  13. 13.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023a.
  14. 14.Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: High-resolution video synthesis with latent diffusion models. arXiv preprint arXiv:2304.08818, 2023b.
  15. 15.Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024.
  16. 16.Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., Aytar, Y., Bechtle, S., Behbahani, F., Chan, S., Heess, N., Gonzalez, L., Osindero, S., Ozair, S., Reed, S., Zhang, J., Zolna, K., Clune, J., de Freitas, N., Singh, S., and Rocktäschel, T. Genie: Generative interactive environments. arXiv preprint arXiv:2402.15391, 2024.
  17. 17.Chan, H., Mnih, V., Behbahani, F., Laskin, M., Wang, L., Pardo, F., Gazeau, M., Sahni, H., Horgan, D., Baumli, K., Schroecker, Y., Spencer, S., Steigerwald, R., Quan, J., Comanici, G., Flennerhag, S., Neitz, A., Zhang, L. M., Schaul, T., Singh, S., Lyle, C., Rocktäschel, T., Parker-Holder, J., and Holsheimer, K. Vision-language models as a source of rewards. In Second Agent Learning in Open-Endedness Workshop, 2023.
  18. 18.Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11315–11325, 2022.
  19. 19.Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. Leveraging procedural generation to benchmark reinforcement learning. In International conference on machine learning, pp. 2048–2056. PMLR, 2020.
  20. 20.Croitoru, F.-A., Hondru, V., Ionescu, R. T., and Shah, M. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  21. 21.Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  22. 22.Dennett, D. C. Consciousness explained. Penguin uk, 1993.
  23. 23.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
  24. 24.Du, Y., Konyushkova, K., Denil, M., Raju, A., Landon, J., Hill, F., de Freitas, N., and Cabi, S. Vision-language models as success detectors. In Proceedings of The 2nd Conference on Lifelong Learning Agents, pp. 120–136, 2023a.
  25. 25.Du, Y., Yang, M., Dai, B., Dai, H., Nachum, O., Tenenbaum, J. B., Schuurmans, D., and Abbeel, P. Learning universal policies via text-guided video generation. arXiv preprint arXiv:2302.00111, 2023b.
  26. 26.Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al. Video language planning. arXiv preprint arXiv:2310.10625, 2023c.
  27. 27.Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C. Imitating latent policies from observation. In International conference on machine learning, pp. 1755–1763. PMLR, 2019.
  28. 28.Efros, A. A. and Freeman, W. T. Image quilting for texture synthesis and transfer. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pp. 571–576. 2023.
  29. 29.Escontrela, A., Adeniji, A., Yan, W., Jain, A., Peng, X. B., Goldberg, K., Lee, Y., Hafner, D., and Abbeel, P. Video prediction models as rewards for reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  30. 30.Faccio, D. and Velten, A. A trillion frames per second: the techniques and applications of light-in-flight photography. Reports on Progress in Physics, 81(10):105901, 2018.
  31. 31.Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A. Minedojo: Building open-ended embodied agents with internet-scale knowledge. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  32. 32.Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023.
  33. 33.Gupta, A., Tian, S., Zhang, Y., Wu, J., Martín-Martín, R., and Fei-Fei, L. Maskvit: Masked visual pre-training for video prediction. arXiv preprint arXiv:2206.11894, 2022.
  34. 34.Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. Mastering atari with discrete world models. arXiv preprint arXiv:2010.02193, 2020.
  35. 35.He, Y., Xia, M., Chen, H., Cun, X., Gong, Y., Xing, J., Zhang, Y., Wang, X., Weng, C., Shan, Y., et al. Animate-a-story: Storytelling with retrieval-augmented video generation. arXiv preprint arXiv:2307.06940, 2023.
  36. 36.Himakunthala, V., Ouyang, A., Rose, D., He, R., Mei, A., Lu, Y., Sonar, C., Saxon, M., and Wang, W. Let’s think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of-thought. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 204–219, 2023.
  37. 37.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  38. 38.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022a.
  39. 39.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022b.
  40. 40.Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023.
  41. 41.Huang, J. and Chang, K. C.-C. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403, 2022.
  42. 42.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pp. 9118–9147. PMLR, 2022.
  43. 43.Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S. Illuminating generalization in deep reinforcement learning through procedural level generation. CoRR, abs/1806.10729, 2018.
  44. 44.Kang, X., Ye, W., and Kuo, Y.-L. Imagined subgoals for hierarchical goal-conditioned policies. In CoRL 2023 Workshop on Learning Effective Abstractions for Planning (LEAP), 2023.
  45. 45.Kashin, A. S., Boiko, D. A., and Ananikov, V. P. Neural network analysis of electron microscopy video data reveals the temperature-driven microphase dynamics in the ions/water system. Small, 17(24):2007726, 2021.
  46. 46.Kim, H. and Lee, K. Air traffic prediction as a video prediction problem using convolutional lstm and autoencoder. Aerospace, 8(10):301, 2021.
  47. 47.Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S. Learning to Simulate Dynamic Environments with GameGAN. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020.
  48. 48.Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B. Learning to act from actionless videos through dense correspondences, 2023.
  49. 49.Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023.
  50. 50.Li, Z., Tucker, R., Snavely, N., and Holynski, A. Generative image dynamics. arXiv preprint arXiv:2309.07906, 2023.
  51. 51.Loeschcke, S., Belongie, S., and Benaim, S. Text-driven stylization of video objects. In European Conference on Computer Vision, pp. 594–609. Springer, 2022.
  52. 52.Lynch, C., Wahid, A., Tompson, J., Ding, T., Betker, J., Baruch, R., Armstrong, T., and Florence, P. Interactive language: Talking to robots in real time. IEEE Robotics and Automation Letters, 2023.
  53. 53.Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al. Augmented language models: a survey. arXiv preprint arXiv:2302.07842, 2023.
  54. 54.Mikuni, V. and Nachman, B. Score-based generative models for calorimeter shower simulation. Physical Review D, 106(9):092009, 2022.
  55. 55.Minsky, M. Society of mind. Simon and Schuster, 1988.
  56. 56.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 518(7540): 529–533, 2015.
  57. 57.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  58. 58.Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023.
  59. 59.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  60. 60.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  61. 61.Razavi, A., Van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019.
  62. 62.Risi, S. and Togelius, J. Increasing generality in machine learning through procedural content generation. Nature Machine Intelligence, 2, 08 2020. doi: 10.1038/s42256-020-0208-z.
  63. 63.Rusu, A. A., Večerík, M., Rothörl, T., Heess, N., Pascanu, R., and Hadsell, R. Sim-to-real robot learning from pixels with progressive nets. In Conference on robot learning, pp. 262–270. PMLR, 2017.
  64. 64.Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A. Learning what you can do before doing anything. arXiv preprint arXiv:1806.09655, 2018.
  65. 65.Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T. P., and Silver, D. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588:604 – 609, 2019. URL https://api.semanticscholar.org/CorpusID:208158225.
  66. 66.Schwarzer, M., Farebrother, J., Greaves, J., Roccapriore, K., Cubuk, E., Agarwal, R., Courville, A., Bellemare, M., Kalinin, S., Mordatch, I., et al. Learning silicon dopant transitions in graphene using scanning transmission electron microscopy. In AI for Accelerated Materials Design-NeurIPS 2023 Workshop, 2023.
  67. 67.Searle, J. R. Minds, brains, and programs. Behavioral and brain sciences, 3(3):417–424, 1980.
  68. 68.Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D. P., Rabinowitz, N. C., Barreto, A., and Degris, T. The predictron: End-to-end learning and planning. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pp. 3191–3199. PMLR, 2017.
  69. 69.Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  70. 70.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  71. 71.Souček, T., Damen, D., Wray, M., Laptev, I., and Sivic, J. Genhowto: Learning to generate actions and state transformations from instructional videos. arXiv preprint arXiv:2312.07322, 2023.
  72. 72.Steinman, D. A. Image-based computational fluid dynamics modeling in realistic arterial geometries. Annals of biomedical engineering, 30:483–497, 2002.
  73. 73.Summerville, A., Snodgrass, S., Guzdial, M., Holmgård, C., Hoover, A. K., Isaksen, A., Nealen, A., and Togelius, J. Procedural content generation via machine learning (PCGML). IEEE Trans. Games, 10(3):257–270, 2018.
  74. 74.Sutton, R. S. Dyna, an integrated architecture for learning, planning, and reacting. ACM Sigart Bulletin, 2(4):160–163, 1991.
  75. 75.Sutton, R. S., Precup, D., and Singh, S. Between MDPs and semi-MDPs: a framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112: 181–211, August 1999. doi: http://dx.doi.org/10.1016/S0004-3702(99)00052-1.
  76. 76.Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  77. 77.Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pp. 23–30. IEEE, 2017.
  78. 78.Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476–482, 2024.
  79. 79.Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., and Kambhampati, S. Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  80. 80.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  81. 81.Venuto, D., Islam, S. N., Klissarov, M., Precup, D., Yang, S., and Anand, A. Code as reward: Empowering reinforcement learning with vlms. arXiv preprint arXiv:2402.04764, 2024.
  82. 82.Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 2022.
  83. 83.Wang, S., Saharia, C., Montgomery, C., Pont-Tuset, J., Noy, S., Pellegrini, S., Onoe, Y., Laszlo, S., Fleet, D. J., Soricut, R., et al. Imagen editor and editbench: Advancing and evaluating text-guided image inpainting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18359–18369, 2023a.
  84. 84.Wang, X., Wang, W., Cao, Y., Shen, C., and Huang, T. Images speak in images: A generalist painter for in-context visual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6830–6839, 2023b.
  85. 85.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022.
  86. 86.Wen, C., Lin, X., So, J., Chen, K., Dou, Q., Gao, Y., and Abbeel, P. Any-point trajectory modeling for policy learning. arXiv preprint arXiv:2401.00025, 2023.
  87. 87.Weng, W., Feng, R., Wang, Y., Dai, Q., Wang, C., Yin, D., Zhao, Z., Qiu, K., Bao, J., Yuan, Y., Luo, C., Zhang, Y., and Xiong, Z. Art·v: Auto-regressive text-to-video generation with diffusion models, 2023.
  88. 88.Xing, Z., Feng, Q., Chen, H., Dai, Q., Hu, H., Xu, H., Wu, Z., and Jiang, Y.-G. A survey on video diffusion models. arXiv preprint arXiv:2310.10647, 2023.
  89. 89.Xiong, W., Luo, W., Ma, L., Liu, W., and Luo, J. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2364–2373, 2018.
  90. 90.Yadav, A., Phillips, M. M., Lundeberg, M. A., Koehler, M. J., Hilden, K., and Dirkin, K. H. If a picture is worth a thousand words is video worth a million? differences in affective and cognitive processing of video and text cases. Journal of Computing in Higher Education, 23: 15–37, 2011.
  91. 91.Yan, W., Hafner, D., James, S., and Abbeel, P. Temporally consistent transformers for video generation. In International Conference on Machine Learning, pp. 39062–39098. PMLR, 2023.
  92. 92.Yang, M., Schuurmans, D., Abbeel, P., and Nachum, O. Dichotomy of control: Separating what you can control from what you cannot. arXiv preprint arXiv:2210.13435, 2022a.
  93. 93.Yang, M., Du, Y., Dai, B., Schuurmans, D., Tenenbaum, J. B., and Abbeel, P. Probabilistic adaptation of text-to-video models. arXiv preprint arXiv:2306.01872, 2023a.
  94. 94.Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023b.
  95. 95.Yang, M. S., Schuurmans, D., Abbeel, P., and Nachum, O. Chain of thought imitation with procedure cloning. Advances in Neural Information Processing Systems, 35: 36366–36381, 2022b.
  96. 96.Yang, S., Nachum, O., Du, Y., Wei, J., Abbeel, P., and Schuurmans, D. Foundation models for decision making: Problems, methods, and opportunities. arXiv preprint arXiv:2303.04129, 2023c.
  97. 97.Yannakakis, G. N. and Togelius, J. Artificial Intelligence and Games. Springer, 2018. https://gameaibook.org.
  98. 98.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.
  99. 99.Ye, W., Zhang, Y., Abbeel, P., and Gao, Y. Become a proficient player with limited data through watching pure videos. In The Eleventh International Conference on Learning Representations, 2022.
  100. 100.Yu, T., Xiao, T., Stone, A., Tompson, J., Brohan, A., Wang, S., Singh, J., Tan, C., Peralta, J., Ichter, B., et al. Scaling robot learning with semantically imagined experience. arXiv preprint arXiv:2302.11550, 2023.
  101. 101.Zhu, J., Yang, H., He, H., Wang, W., Tuo, Z., Cheng, W.-H., Gao, L., Song, J., and Fu, J. Moviefactory: Automatic movie creation from text using large generative models for language and images. arXiv preprint arXiv:2306.07257, 2023.

Citation

MLA
Yang, S., et al. “Video as the New Language for Real-World Decision Making”. arXiv, 2024, http://arxiv.org/abs/2402.17139v1.
APA
Yang, S., Walker, J., Parker-Holder, J., Du, Y., Bruce, J., Barreto, A., Abbeel, P., & Schuurmans, D. (2024). Video as the New Language for Real-World Decision Making. arXiv. http://arxiv.org/abs/2402.17139v1
Chicago
Yang, S., J. Walker, J. Parker-Holder, et al. 2024. “Video as the New Language for Real-World Decision Making”. arXiv. http://arxiv.org/abs/2402.17139v1.
Harvard
Yang, S. et al. (2024) “Video as the New Language for Real-World Decision Making”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.17139v1.
Vancouver
1. Yang S, Walker J, Parker-Holder J, Du Y, Bruce J, Barreto A, Abbeel P, Schuurmans D (2024) Video as the New Language for Real-World Decision Making. arXiv

BibTeX

@article{yang2024video,
  title = {Video as the New Language for Real-World Decision Making},
  author = {Yang, Sherry and Walker, Jacob and Parker-Holder, Jack and Du, Yilun and Bruce, Jake and Barreto, Andre and Abbeel, Pieter and Schuurmans, Dale},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.17139v1},
  eprint = {2402.17139}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/