Learning to Model the World With Language
Jessy LinYuqing DuOlivia WatkinsDanijar HafnerPieter AbbeelDan KleinAnca D. Dragan
Introduces Dynalang, an embodied agent that grounds diverse language like instructions, manuals, and environment descriptions by predicting future multimodal representations within a world model to plan actions from imagined rollouts.
Autonomous artificial intelligence agents operating in physical and simulated spaces must interpret varied human language to collaborate effectively. Standard reinforcement learning approaches typically limit language to direct commands, such as basic task instructions, and directly map those prompts to specific physical actions. However, human communication encompasses diverse expressions, including explanations of environmental dynamics, world state updates, and real-time corrections. Standard methods struggle to process this richer language because indirect statements often exhibit complex, weak statistical correlations with immediate optimal actions.
The article demonstrates that treating multimodal language understanding as a future prediction task enables embodied agents to utilize diverse linguistic inputs. It introduces and evaluates Dynalang, an agent framework that decouples learning a predictive generative world model from learning an action policy, allowing diverse text to ground naturally into visual predictions and future reward expectations.
To test this approach, the authors developed a multimodal world model that integrates visual frames and text tokens at each time step into compressed latent representations. The architecture utilizes a self-supervised objective to predict future latent states, reconstruct sensory inputs, and anticipate rewards. Action policies are subsequently trained entirely within the model’s imagined rollouts. Dynalang was evaluated across four distinct simulated settings: HomeGrid (a home-chore environment featuring diverse hints), Messenger (a multi-stage reasoning game with rule manuals), Habitat Vision-Language Navigation in Continuous Environments (navigating photorealistic home scans), and LangRoom (an embodied question-answering environment), alongside comparisons against standard model-free algorithms such as IMPALA and R2D2.
The evaluation revealed several critical findings. First, Dynalang effectively translated diverse language hints—including dynamics rules, future observation hints, and corrections—into substantial task performance gains in HomeGrid, whereas baseline models degraded in performance when exposed to diverse language. Second, on the Messenger benchmark, Dynalang successfully solved the most complex third stage using multi-hop reasoning over game manuals, while standard baselines and domain-specialized architectures failed completely. Third, in continuous vision-language navigation, Dynalang achieved a success rate near 30% from scratch, substantially outperforming model-free baselines. Fourth, by regularizing language actions with the world model's internal predictions, Dynalang scaled embodied question-answering to a 10,000-token vocabulary, matching small-vocabulary performance where unregularized policies failed. Finally, pretraining the world model on 500 million tokens of text-only data without actions or rewards significantly accelerated downstream task learning, outperforming models that relied on frozen external text representations.
These findings indicate that unifying language and visual reasoning under a self-supervised future prediction objective resolves a core bottleneck in embodied intelligence. Rather than requiring expensive, task-specific paired demonstration datasets, agents can leverage abundant offline text data and autonomously ground diverse language through online experience. This decoupling reduces the engineering complexity of multimodal training while improving policy robustness across varied operational conditions.
Organizations developing embodied AI should adopt generative multimodal world modeling frameworks when designing agents intended for dynamic, human-centric environments. System designers should prioritize token-level streaming inputs over static sentence-level conditioning and leverage offline domain-specific or general text corpora to pretrain world models before online deployment. Additional development is needed to close the remaining performance gap between reinforcement learning world models and specialized, demonstration-heavy navigation architectures in high-fidelity settings.
The findings are established across multiple varied simulation environments, providing high confidence in the core conceptual approach. However, users should note key boundaries: the experimental evaluations were conducted entirely in simulated environments rather than physical hardware, and performance on complex continuous navigation still trails specialized systems trained on human demonstrations. Physical deployments should await further validation on real-world robotic platforms.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). Introduces foundational multimodal token injection for embodied robotic planning that motivates Dynalang's unified visual-linguistic state modeling.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Establishes vision-language-action modeling by expressing robotic actions as tokens within language model backbones, a direct conceptual precursor to Dynalang.
- Paper: Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning, Thomas Carta et al. (2023). Demonstrates the necessity and challenges of functionally grounding language models in interactive reinforcement learning environments.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Pioneers closed-loop language feedback for embodied decision-making, highlighting the planning limitations of standard language prompting that world models resolve.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Investigates zero-shot task decomposition and action translation using pretrained language models in interactive household simulation domains.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). Introduces the foundational vision-and-language navigation paradigm evaluated within Dynalang's continuous navigation experiments.
- Paper: LIV: Language-Image Representations and Rewards for Robotic Control, Yecheng Jason Ma et al. (2023). Explores joint vision-language representation learning from action-free video for downstream robotic control.
- Paper: Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions, Jing Gu et al. (2022). Provides a comprehensive taxonomy of vision-and-language navigation benchmarks and representation strategies used throughout embodied AI.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). Expands predictive vision-language world modeling to uncurated real-world egocentric videos with dual-system reflective planning.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). Extends the principle of unifying future visual prediction and language understanding directly to low-level robotic manipulation tasks.
- Paper: Dyna-Mind: Learning to Simulate from Experience for Better AI Agents, Xiao Yu et al. (2025). Builds upon internal simulation architectures by training agents to mentally simulate future rollouts before executing actions in complex interactive domains.
- Paper: DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, Gaoyue Zhou et al. (2025). Advances predictive world modeling by simulating dynamics over self-supervised pretrained visual feature spaces for zero-shot planning.
- Paper: ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model, Zhongyi Zhou et al. (2025). Generalizes multimodal understanding and physical robot control by uniting conversational reasoning with direct motor execution in a mixture-of-experts model.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). Formalizes the conceptual thesis of predictive generative modeling across visual tokens as the universal interface for embodied decision making.
- Paper: EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents, Rui Yang 0010 et al. (2025). Establishes a rigorous benchmarking suite to evaluate how vision-language agents generalize across diverse physical control and spatial reasoning tasks.
- Paper: An Embodied Generalist Agent in 3D World, Jiangyong Huang et al. (2024). Extends multimodal language-conditioned embodied intelligence into unified 3D point cloud perception, grounding, and physical manipulation.
- Paper: SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment, Katrin Renz et al. (2025). Applies generative future prediction and language-action alignment to closed-loop autonomous vehicle navigation.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). Implements compositional language-driven visual world models to synthesize imaginative rollouts for robotic policy execution.
