AdaWorld: Learning Adaptable World Models with Latent Actions
Shenyuan GaoSiyuan ZhouYilun DuJun ZhangChuang Gan
Proposes AdaWorld, a pretraining framework that extracts self-supervised latent actions from unlabeled videos to condition autoregressive world models, enabling fast adaptation and visual planning in novel environments with minimal interaction data.
Building artificial intelligence agents that can predict future visual outcomes and plan actions across diverse environments requires predictive simulators known as world models. However, current world models depend on massive volumes of manually labeled action data and expensive computational training, making it difficult and slow to adapt them to new tasks with limited real-world interaction. This bottleneck restricts the practical deployment of intelligent systems across robotics, gaming, and simulation.
The article introduces and evaluates AdaWorld, a novel pretraining framework designed to produce highly adaptable world models. The primary objective is to demonstrate that incorporating self-supervised "latent actions"—compact, context-invariant action representations extracted directly from unlabeled video—enables efficient adaptation, action transfer, and autonomous visual planning across unseen environments using minimal interaction data.
To evaluate this framework, the authors pretrained an autoencoder that extracts continuous latent actions from video frame pairs using an information-bottleneck design, alongside a diffusion-based predictive world model conditioned on these latent actions. Pretraining was conducted on a large-scale dataset spanning roughly two billion frames from over 1,000 video game environments, robot datasets, and human activity videos. The resulting system was systematically benchmarked against standard action-agnostic models, discrete action models, and optical flow baselines across diverse benchmarks, including game environments (Procgen, Minecraft, DMLab), robotics suites (Robosuite, RoboDesk), and driving environments (nuScenes).
The evaluation yielded several key findings. First, AdaWorld transfers demonstrated actions into new visual contexts without additional training, achieving human evaluation success rates of 70.5% on robot manipulation benchmarks and 61.5% on human video datasets, compared to 0% to 21.5% for alternative baselines. Second, when adapting to unseen discrete and continuous environments using only 100 interaction samples per action, AdaWorld consistently delivered superior simulation fidelity over baselines after just 800 tuning steps. Third, in visual planning benchmarks across video games, AdaWorld achieved an average success rate of 56.67% with minor tuning (and 44.83% without any parameter updates), outperforming traditional reinforcement learning (27.17%) and action-agnostic pretraining (26.00%). Finally, on standardized robotic planning benchmarks, AdaWorld achieved an aggregate success score of 21.54, quadrupling the 5.03 score attained by action-agnostic baselines.
These findings indicate that pretraining world models with continuous latent action representations dramatically reduces the cost, data collection burden, and computational time required to deploy autonomous agents in new settings. By offering a unified, pre-structured control interface, the approach eliminates the need to engineer task-specific action formats or collect exhaustive manual labels from scratch, challenging the conventional paradigm of action-agnostic video pretraining.
Organizations developing embodied AI and simulation tools should adopt action-aware pretraining frameworks to streamline cross-domain agent deployment. When deploying to new operational environments, technical teams should leverage the continuous latent space to initialize control interfaces through sample averaging or lightweight mapping layers rather than retraining models from scratch. Further investment should focus on integrating inference acceleration techniques, such as model distillation, to achieve real-time execution speeds.
Decision-makers should consider key limitations when interpreting these results. While AdaWorld demonstrates high adaptability, it does not yet run at real-time speeds, struggles to imagine completely novel visual content when navigating far beyond initial scene boundaries, and exhibits quality degradation during very long-term rollouts or dramatic camera viewpoint shifts. Nonetheless, evidence remains highly confident regarding its sample efficiency, visual transfer capabilities, and planning performance in controlled benchmarks.
- Paper: Genie: Generative Interactive Environments, Jake Bruce et al. (2024). Genie introduces latent actions inferred from unlabeled video, the key idea AdaWorld develops into transferable continuous control representations.
- Paper: Reinforcement Learning with Action-Free Pre-Training from Videos, Younggyo Seo et al. (2022). APV establishes how action-free video pretraining can initialize world models for downstream control, clarifying the contrast AdaWorld makes with action-agnostic approaches.
- Paper: Learning Latent Dynamics for Planning from Pixels, Danijar Hafner et al. (2018). PlaNet establishes latent-space dynamics learning and predictive planning from pixels, prerequisites for understanding AdaWorld’s learned visual simulator.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Dreamer shows how agents can learn policies through imagined trajectories in a visual world model, grounding AdaWorld’s use of predictive models for planning.
- Paper: ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models, Xinye Li et al. (2026). ForgeWM advances action-conditioned video world models toward few-step, causal generation, directly addressing AdaWorld’s stated limitation in real-time execution.
