AdaWorld: Learning Adaptable World Models with Latent Actions

Shenyuan GaoSiyuan ZhouYilun DuJun ZhangChuang Gan

article2025ICML60 citations

Proposes AdaWorld, a pretraining framework that extracts self-supervised latent actions from unlabeled videos to condition autoregressive world models, enabling fast adaptation and visual planning in novel environments with minimal interaction data.

Listen

Building artificial intelligence agents that can predict future visual outcomes and plan actions across diverse environments requires predictive simulators known as world models. However, current world models depend on massive volumes of manually labeled action data and expensive computational training, making it difficult and slow to adapt them to new tasks with limited real-world interaction. This bottleneck restricts the practical deployment of intelligent systems across robotics, gaming, and simulation.

The article introduces and evaluates AdaWorld, a novel pretraining framework designed to produce highly adaptable world models. The primary objective is to demonstrate that incorporating self-supervised "latent actions"—compact, context-invariant action representations extracted directly from unlabeled video—enables efficient adaptation, action transfer, and autonomous visual planning across unseen environments using minimal interaction data.

To evaluate this framework, the authors pretrained an autoencoder that extracts continuous latent actions from video frame pairs using an information-bottleneck design, alongside a diffusion-based predictive world model conditioned on these latent actions. Pretraining was conducted on a large-scale dataset spanning roughly two billion frames from over 1,000 video game environments, robot datasets, and human activity videos. The resulting system was systematically benchmarked against standard action-agnostic models, discrete action models, and optical flow baselines across diverse benchmarks, including game environments (Procgen, Minecraft, DMLab), robotics suites (Robosuite, RoboDesk), and driving environments (nuScenes).

The evaluation yielded several key findings. First, AdaWorld transfers demonstrated actions into new visual contexts without additional training, achieving human evaluation success rates of 70.5% on robot manipulation benchmarks and 61.5% on human video datasets, compared to 0% to 21.5% for alternative baselines. Second, when adapting to unseen discrete and continuous environments using only 100 interaction samples per action, AdaWorld consistently delivered superior simulation fidelity over baselines after just 800 tuning steps. Third, in visual planning benchmarks across video games, AdaWorld achieved an average success rate of 56.67% with minor tuning (and 44.83% without any parameter updates), outperforming traditional reinforcement learning (27.17%) and action-agnostic pretraining (26.00%). Finally, on standardized robotic planning benchmarks, AdaWorld achieved an aggregate success score of 21.54, quadrupling the 5.03 score attained by action-agnostic baselines.

These findings indicate that pretraining world models with continuous latent action representations dramatically reduces the cost, data collection burden, and computational time required to deploy autonomous agents in new settings. By offering a unified, pre-structured control interface, the approach eliminates the need to engineer task-specific action formats or collect exhaustive manual labels from scratch, challenging the conventional paradigm of action-agnostic video pretraining.

Organizations developing embodied AI and simulation tools should adopt action-aware pretraining frameworks to streamline cross-domain agent deployment. When deploying to new operational environments, technical teams should leverage the continuous latent space to initialize control interfaces through sample averaging or lightweight mapping layers rather than retraining models from scratch. Further investment should focus on integrating inference acceleration techniques, such as model distillation, to achieve real-time execution speeds.

Decision-makers should consider key limitations when interpreting these results. While AdaWorld demonstrates high adaptability, it does not yet run at real-time speeds, struggles to imagine completely novel visual content when navigating far beyond initial scene boundaries, and exhibits quality degradation during very long-term rollouts or dramatic camera viewpoint shifts. Nonetheless, evidence remains highly confident regarding its sample efficiency, visual transfer capabilities, and planning performance in controlled benchmarks.

arXiv: 2503.18938Little-Podi/AdaWorld
Cover for AdaWorld: Learning Adaptable World Models with Latent Actions

Abstract

World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents. However, most existing world models rely heavily on substantial action-labeled data and costly training, making it challenging to adapt to novel environments with heterogeneous actions through limited interactions. This limitation can hinder their applicability across broader domains. To overcome this limitation, we propose AdaWorld, an innovative world model learning approach that enables efficient adaptation. The key idea is to incorporate action information during the pretraining of world models. This is achieved by extracting latent actions from videos in a self-supervised manner, capturing the most critical transitions between frames. We then develop an autoregressive world model that conditions on these latent actions. This learning paradigm enables highly adaptable world models, facilitating efficient transfer and learning of new actions even with limited interactions and finetuning. Our comprehensive experiments across multiple environments demonstrate that AdaWorld achieves superior performance in both simulation quality and visual planning.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Latent Action Autoencoder
  • 2.2. Action-Aware Pretraining
  • 2.3. Highly Adaptable World Models
  • 3. Experiments
  • 3.1. Action Transfer
  • 3.2. World Model Adaptation
  • 3.2.1. SIMULATION QUALITY
  • 3.2.2. VISUAL PLANNING IN GAMES
  • 3.2.3. VISUAL PLANNING IN ROBOT TASKS
  • 3.3. Ablation and Analysis
  • 4. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Datasets
  • A.1. Data Collection and Generation
  • A.2. Data Mixture
  • B. Implementation Details
  • B.1. Architecture
  • B.2. Training
  • B.3. Sampling
  • B.4. Visual Planning on the Procgen Benchmark
  • B.5. Visual Planning on the VP 2 Benchmark
  • B.6. iVideoGPT Training Details
  • C. Additional Results
  • C.1. Action Transfer
  • C.2. World Model Adaptation
  • C.3. Action Creation through Clustering
  • D. Something-Something v2 Categories for Action Transfer
  • E. Selected Scenes for Visual Planning
  • F. Related Work
  • F.1. World Models
  • F.2. Latent Action from Videos

Knowls

  1. Knowl 1 — Action-aware pretraining makes a world model adaptable

    model/method

    AdaWorld pretrains a predictive world model on videos by conditioning each next-frame prediction on a self-supervised latent action inferred from two consecutive frames. Unlike action-agnostic pretraining, which supplies no action information, this teaches the model to associate controllable transitions with a shared action representation across varied environments. For transfer, the latent actions extracted from a demonstrated video can be reused in a new context; for adaptation to a new action space, they can initialize a mapping from that environment's controls into the model's latent-action interface.

  2. Knowl 2 — A continuous information-bottleneck representation of frame transitions

    equation

    AdaWorld learns a continuous latent action a~\tilde a from consecutive video frames ftf_t and ft+1f_{t+1}. An encoder with parameters ϕ\phi approximates the posterior qϕ(a~∣ft,ft+1)q_\phi(\tilde a\mid f_t,f_{t+1}); a decoder with parameters θ\theta predicts ft+1f_{t+1} from ftf_t and a sample of a~\tilde a. The autoencoder maximizes the following β\beta-VAE objective, where p(a~)p(\tilde a) is the latent prior and DKLD_{\mathrm{KL}} is Kullback–Leibler divergence:

    Jθ,ϕ=Eqϕ(a~∣ft,ft+1) ⁣[log⁡pθ(ft+1∣a~,ft)]−βDKL ⁣(qϕ(a~∣ft,ft+1)∥p(a~)).\mathcal{J}_{\theta,\phi} = \mathbb{E}_{q_\phi(\tilde a\mid f_t,f_{t+1})}\!\left[\log p_\theta(f_{t+1}\mid \tilde a,f_t)\right] - \beta D_{\mathrm{KL}}\!\left(q_\phi(\tilde a\mid f_t,f_{t+1})\|p(\tilde a)\right).

    The encoder predicts the posterior parameters and samples the latent action; the decoder is trained to reconstruct the subsequent frame. The compact bottleneck encourages the latent to retain transition-critical information rather than scene appearance. The reported latent dimension is 32, and the default β\beta is 2×10−42\times10^{-4}. The paper reports that reducing β\beta makes latent actions more expressive but decreases their overlap across environments, weakening context disentanglement.

  3. Knowl 3 — Transformer architecture for extracting latent actions

    model/method

    The latent-action autoencoder uses a Transformer to encode and decode video-frame transitions. Its encoder divides each input frame into 16×1616\times16 image patches, projects and spatially flattens the patches, and adds learnable frame tokens and sinusoidal spatial position embeddings. A spatiotemporal Transformer alternates spatial attention within each frame with temporal attention between corresponding spatial positions in the two frames; the later frame token aggregates transition information and is projected to the latent posterior. The decoder is a spatial Transformer that predicts the next frame from the preceding frame and sampled latent action. The reported autoencoder has 500 million parameters, 16 encoder blocks and 16 decoder blocks, 1024 channels, and 16 attention heads.

  4. Knowl 4 — Diffusion world model for autoregressive frame-level control

    model/method

    AdaWorld uses a separate diffusion world model rather than using the latent-action decoder for multi-step rollouts. The model is based on Stable Video Diffusion and has a 3D U-Net architecture. It denoises one target frame at a time, conditioned on the current latent action and a short-term memory of historical frames. The historical frames are encoded with the pretrained image encoder; the last memory frame is also used as the condition image, while the latent action is incorporated with the timestep and CLIP image embeddings. During training, the memory length varies up to six frames, and memory frames receive noise augmentation. In latent diffusion space, the model minimizes Ex0,ϵ,t[∥x0−x^0(xt,t,c)∥2]\mathbb{E}_{x_0,\epsilon,t}\left[\|x_0-\hat{x}_0(x_t,t,c)\|^2\right], where x0x_0 is the clean target-frame latent, xtx_t is its noisy version at diffusion timestep tt, ϵ\epsilon is the injected noise, x^0\hat{x}_0 is the predicted clean latent, and cc comprises the historical-frame and latent-action conditions. At inference, predicted frames are appended to memory and the model is run autoregressively. The reported world model has 1.5 billion trainable parameters.

  5. Knowl 5 — Action transfer, adaptation, and composition through the latent interface

    model/method

    AdaWorld supports several uses of its continuous latent-action interface. For training-free action transfer, the encoder extracts a sequence of latent actions from a demonstration, and the world model reuses that sequence to generate frames from a new context. To adapt to an environment with discrete controls, latent actions inferred from examples of the same labeled action are averaged; the resulting embeddings initialize the action controls, after which the whole world model can be finetuned. For continuous controls, a two-layer MLP maps raw actions to latent actions and can be initialized from a small set of action–latent-action pairs. The continuous space also allows action composition by averaging latent actions—for example, combining rightward movement and jumping—and allows a customizable number of control options by clustering latent actions with K-means.

  6. Knowl 6 — Pretraining data and reported training scale

    experimental setup

    AdaWorld's training corpus contains approximately two billion frames drawn from videos and automatically generated interactive environments. The reported source breakdown is:

    Category Data source Frames Reported ratio
    2D video game Gym Retro 1000M 49%
    2D video game Procgen Benchmark 144M 2%
    Robot data Open X-Embodiment 170M 30%
    Human activity Ego4D 330M 1%
    Human activity Something-Something V2 7M 3%
    3D rendering MiraData 200M 14%
    City walking MiraData 120M 1%

    For game data, the authors collected transitions from 1,000 Gym Retro environments and used 9,000 of Procgen's 10,000 levels for training. They alternated periods of increased action-selection probability to encourage broader exploration. The latent-action autoencoder was trained from scratch for 200K steps with batch size 960 and learning rate 2.5×10−52.5\times10^{-5}. The full world-model implementation reports 80K pretraining steps, batch size 64, learning rate 5×10−55\times10^{-5}, and 16 NVIDIA A100 GPUs; the controlled comparison experiment reports 50K iterations for each compared world model.

  7. Knowl 7 — Demonstrated actions transfer across contexts

    empirical result

    The action-transfer evaluation tests whether a model can use a demonstration from one video to generate the corresponding action in a different context. It uses 1,300 paired videos from unseen LIBERO tasks and Something-Something v2 (SSv2) action categories. For each pair, the first video supplies the action sequence and the second supplies the initial frame; models generate 20 frames autoregressively. Frechet Video Distance (FVD) measures distribution similarity, Embedding Cosine Similarity (ECS) measures frame-level similarity, and human evaluation judges whether the action transferred successfully. Human judgments use 50 pairs per dataset and four volunteers. AdaWorld leads on both automatic metrics and human judgments:

    LIBERO SSv2
    Method FVD ↓\downarrow ECS ↑\uparrow Human ↑\uparrow FVD ↓\downarrow ECS ↑\uparrow Human ↑\uparrow
    Action-agnostic 1545.2 0.702 0% 847.2 0.592 1%
    Optical-flow condition 1409.5 0.724 2% 702.8 0.611 10.5%
    Discrete latent-action condition 1504.5 0.700 3.5% 726.8 0.596 21.5%
    AdaWorld 767.0 0.804 70.5% 473.4 0.639 61.5%

    The results show that AdaWorld's continuous latent actions transfer demonstrated movements more reliably than action-agnostic conditioning, optical flow, or a discrete latent-action representation under this evaluation.

  8. Knowl 8 — Adaptation improves action-controlled simulation in unseen environments

    empirical result

    The authors adapted pretrained models to four environments absent from the training data: Habitat, Minecraft, and DMLab with discrete actions, and nuScenes with continuous actions. Adaptation used 100 samples per discrete action or 100 continuous trajectories, followed by 800 finetuning steps; fidelity was evaluated on 300 validation samples using PSNR and LPIPS. AdaWorld achieved the highest PSNR and lowest LPIPS in all four environments:

    Habitat Minecraft DMLab nuScenes
    Method PSNR ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow
    Action-agnostic 20.34 0.450 19.44 0.532 20.96 0.386 20.86 0.475
    Optical-flow condition 22.49 0.373 20.71 0.492 22.22 0.357 20.94 0.462
    Discrete latent-action condition 23.31 0.342 21.33 0.465 22.36 0.349 21.28 0.450
    AdaWorld 23.58 0.327 21.59 0.457 22.92 0.335 21.60 0.436

    Separate finetuning curves for Minecraft and nuScenes show AdaWorld improving faster than conventional pretraining baselines under varying sample counts and training steps. For continuous nuScenes controls, a two-layer MLP maps displacements to latent actions and was finetuned for 3K steps using limited action–latent-action pairs.

  9. Knowl 9 — AdaWorld improves visual planning success in Procgen games

    empirical result

    The game-planning evaluation uses 120 goal-reaching scenes across four Procgen environments, with 30 scenes each from Heist, Jumper, Maze, and CaveFlyer. Models receive 100 labeled samples per action and are finetuned for 500 steps; sampling-based model-predictive control uses the learned world model, and a run succeeds if the goal is reached within 20 steps. AdaWorld is also evaluated without finetuning, using averaged latent actions as controls. The table reports success rate averaged over five random seeds with standard error; the ground-truth-simulator oracle is an upper bound for the planning strategy.

    Method Heist Jumper Maze CaveFlyer Average
    Random 19.33±\pm4.41% 22.00±\pm2.50% 41.33±\pm5.44% 22.00±\pm2.50% 26.17±\pm2.55%
    Action-agnostic 20.67±\pm3.55% 20.67±\pm2.45% 39.33±\pm2.87% 23.33±\pm1.84% 26.00±\pm0.98%
    AdaWorld, no finetuning 38.67±\pm2.01% 68.00±\pm2.25% 41.33±\pm2.72% 31.33±\pm2.50% 44.83±\pm1.37%
    AdaWorld, finetuned 66.67±\pm4.09% 58.67±\pm2.50% 68.00±\pm1.69% 33.33±\pm3.80% 56.67±\pm2.16%
    Q-learning 22.67±\pm3.87% 47.33±\pm6.71% 4.67±\pm0.81% 34.00±\pm6.17% 27.17±\pm1.27%
    Oracle (ground-truth environment) 86.67±\pm3.16% 77.33±\pm2.67% 84.67±\pm2.91% 74.00±\pm3.99% 80.67±\pm2.11%

    Finetuned AdaWorld has the highest average success among learned methods, at 56.67%; even without finetuning, its average of 44.83% exceeds the action-agnostic baseline, random planning, and Q-learning in this evaluation.

  10. Knowl 10 — AdaWorld improves planning success on robot tasks

    empirical result

    Robot planning was evaluated after adapting a low-resolution AdaWorld variant and an action-agnostic baseline to VP2 tasks. The model-predictive path-integral planner used the learned models; adaptation used 1K steps. The reported success rates and standard errors are averages over four runs. The aggregate score is normalized by the ground-truth simulator's scores. AdaWorld improves success in each listed task and raises the aggregate score from 5.03 to 21.54:

    Method Robosuite push Open slide Blue button Green button Red button Upright block Aggregate
    Action-agnostic 17.50±\pm0.50% 1.67±\pm1.67% 5.00±\pm1.67% 3.33±\pm0.00% 0.00±\pm0.00% 1.67±\pm1.67% 5.03
    AdaWorld 63.50±\pm1.71% 5.83±\pm2.85% 29.17±\pm2.50% 10.83±\pm2.50% 10.00±\pm2.36% 5.00±\pm0.96% 21.54

    The evaluation covers tabletop Robosuite and RoboDesk tasks; the reported task columns give the per-task success rates shown above.

  11. Knowl 11 — Data diversity and action-aware pretraining generalize across settings

    empirical result

    Two analyses test whether AdaWorld's benefits extend beyond a single training-data mixture or world-model architecture. First, latent-action autoencoders trained for 40K steps on Open X-Embodiment (OpenX), Gym Retro, or both were evaluated on Procgen decoder predictions. Adding OpenX to Retro improved both reported metrics, despite OpenX consisting mainly of real-world robot videos and Procgen being 2D virtual games. Second, conditioning iVideoGPT on latent actions during pretraining improved its action-controlled BAIR simulation results after finetuning, indicating that action-aware pretraining can benefit a different autoregressive world-model architecture.

    Training data for latent-action autoencoder Procgen PSNR ↑\uparrow Procgen LPIPS ↓\downarrow
    OpenX 25.51 0.318
    Gym Retro 26.43 0.250
    Gym Retro + OpenX 26.62 0.234
    Model evaluated on BAIR PSNR ↑\uparrow LPIPS ↓\downarrow
    iVideoGPT 16.59 0.220
    iVideoGPT with AdaWorld action-aware pretraining 17.40 0.204

    These are separate evaluations: the Procgen metrics assess decoder predictions under different autoencoder training mixtures, while the BAIR metrics assess action-controlled simulation after finetuning.

  12. Knowl 12 — Limitations in speed, extrapolation, and long rollouts

    limitation

    AdaWorld does not operate at real-time frequency. The authors also report difficulty generating novel content once a rollout moves beyond its initial scene, and limited fidelity over extremely long rollouts. Failure cases include imperfect simulation of real-world physics and dynamic agents, as well as large viewpoint changes. Faster inference, improved long-horizon generation, and greater capacity or more training data are presented as directions for future work, not as demonstrated solutions.

Coverage note — The detailed pseudocode and hyperparameters for Procgen cross-entropy planning, plus exhaustive environment and scene listings, are omitted because they are experimental subdetails rather than distinct contributions.

References

  1. 1.Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al. Cosmos World Foundation Model Platform for Physical AI. arXiv preprint arXiv:2501.03575, 2025.
  2. 2.Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K. Deep Variational Information Bottleneck. In ICLR, 2017.
  3. 3.Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., and Fleuret, F. Diffusion for World Modeling: Visual Details Matter in Atari. In NeurIPS, 2024.
  4. 4.Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J. Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos. In NeurIPS, 2022.
  5. 5.Bar, A., Zhou, G., Tran, D., Darrell, T., and LeCun, Y. Navigation World Models. In CVPR, 2025.
  6. 6.Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Kuttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al. DeepMind Lab. arXiv preprint arXiv:1612.03801, 2016.
  7. 7.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv preprint arXiv:2311.15127, 2023.
  8. 8.Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al. Genie: Generative Interactive Environments. In ICML, 2024.
  9. 9.Bu, Q., Zeng, J., Chen, L., Yang, Y., Zhou, G., Yan, J., Luo, P., Cui, H., Ma, Y., and Li, H. Closed-Loop Visuo-motor Control with Generative Expectation for Robotic Manipulation. In NeurIPS, 2024.
  10. 10.Bu, Q., Cai, J., Chen, L., Cui, X., Ding, Y., Feng, S., Gao, S., He, X., Huang, X., Jiang, S., et al. AgiBot World Colosseo: A Large-Scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv preprint arXiv:2503.06669, 2025.
  11. 11.Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A. Understanding Disentangling in β-VAE. In NeurIPS Workshops, 2017.
  12. 12.Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. nuScenes: A Multimodal Dataset for Autonomous Driving. In CVPR, 2020.
  13. 13.Carreira, J. and Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In CVPR, 2017.
  14. 14.Che, H., He, X., Liu, Q., Jin, C., and Chen, H. GameGen-X: Interactive Open-World Game Video Generation. In ICLR, 2025.
  15. 15.Chen, B., Monso, D. M., Du, Y., Simchowitz, M., Tedrake, R., and Sitzmann, V. Diffusion Forcing: Next-Token Prediction Meets Full-Sequence Diffusion. In NeurIPS, 2024a.
  16. 16.Chen, X., Guo, J., He, T., Zhang, C., Zhang, P., Yang, D. C., Zhao, L., and Bian, J. IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI. arXiv preprint arXiv:2411.00785, 2024b.
  17. 17.Chen, Y., Ge, Y., Li, Y., Ge, Y., Ding, M., Shan, Y., and Liu, X. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation. arXiv preprint arXiv:2412.04445, 2024c.
  18. 18.Chi, X., Zhang, H., Fan, C.-K., Qi, X., Zhang, R., Chen, A., Chan, C.-m., Xue, W., Luo, W., Zhang, S., et al. EVA: An Embodied World Model for Future Video Anticipation. arXiv preprint arXiv:2410.15461, 2024.
  19. 19.Chua, K., Calandra, R., McAllister, R., and Levine, S. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models. In NeurIPS, 2018.
  20. 20.Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. Leveraging Procedural Generation to Benchmark Reinforcement Learning. In ICML, 2020.
  21. 21.Cui, Z. J., Pan, H., Iyer, A., Haldar, S., and Pinto, L. DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control. In NeurIPS, 2024.
  22. 22.De Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y. A Tutorial on the Cross-Entropy Method. Annals of Operations Research, 2005.
  23. 23.Dominici, N., Ivanenko, Y. P., Cappellini, G., d’Avella, A., Mondì, V., Cicchese, M., Fabiano, A., Silei, T., Di Paolo, A., Giannini, C., et al. Locomotor Primitives in Newborn Babies and Their Development. Science, 2011.
  24. 24.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR, 2021.
  25. 25.Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P. Learning Universal Policies via Text-Guided Video Generation. In NeurIPS, 2023.
  26. 26.Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al. Video Language Planning. In ICLR, 2024.
  27. 27.Durante, Z., Sarkar, B., Gong, R., Taori, R., Noda, Y., Tang, P., Adeli, E., Lakshmikanth, S. K., Schulman, K., Milstein, A., et al. An Interactive Agent Foundation Model. arXiv preprint arXiv:2402.05929, 2024.
  28. 28.Ebert, F., Finn, C., Lee, A. X., and Levine, S. Self-Supervised Visual Planning with Temporal Skip Connections. In CoRL, 2017.
  29. 29.Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C. Imitating Latent Policies from Observation. In ICML, 2019.
  30. 30.Feng, R., Zhang, H., Yang, Z., Xiao, J., Shu, Z., Liu, Z., Zheng, A., Huang, Y., Liu, Y., and Zhang, H. The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control. arXiv preprint arXiv:2412.03568, 2024.
  31. 31.Gao, S., Yang, J., Chen, L., Chitta, K., Qiu, Y., Geiger, A., Zhang, J., and Li, H. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability. In NeurIPS, 2024.
  32. 32.Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al. The “Something Something” Video Database for Learning and Evaluating Visual Common Sense. In ICCV, 2017.
  33. 33.Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al. Ego4D: Around the World in 3,000 Hours of Egocentric Video. In CVPR, 2022.
  34. 34.Ha, D. and Schmidhuber, J. Recurrent World Models Facilitate Policy Evolution. In NeurIPS, 2018.
  35. 35.Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. Mastering Diverse Domains through World Models. arXiv preprint arXiv:2301.04104, 2023.
  36. 36.Hansen, N., Su, H., and Wang, X. TD-MPC2: Scalable, Robust World Models for Continuous Control. In ICLR, 2024.
  37. 37.Hassan, M., Stapf, S., Rahimi, A., Rezende, P., Haghighi, Y., Bruggemann, D., Katircioglu, I., Zhang, L., Chen, X., Saha, S., et al. GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control. In CVPR, 2025.
  38. 38.He, H., Zhang, Y., Lin, L., Xu, Z., and Pan, L. Pre-Trained Video Generative Models as World Simulators. arXiv preprint arXiv:2502.07825, 2025.
  39. 39.He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q. Latent Video Diffusion Models for High-Fidelity Long Video Generation. arXiv preprint arXiv:2211.13221, 2022.
  40. 40.Higgins, I., Matthey, L., Pal, A., Burgess, C. P., Glorot, X., Botvinick, M. M., Mohamed, S., and Lerchner, A. beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. In ICLR, 2017.
  41. 41.Hong, Y., Liu, B., Wu, M., Zhai, Y., Chang, K.-W., Li, L., Lin, K., Lin, C.-C., Wang, J., Yang, Z., et al. SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation. In ICLR, 2025.
  42. 42.Hore, A. and Ziou, D. Image Quality Metrics: PSNR vs. SSIM. In ICPR, 2010.
  43. 43.Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. GAIA-1: A Generative World Model for Autonomous Driving. arXiv preprint arXiv:2309.17080, 2023.
  44. 44.Ju, X., Gao, Y., Zhang, Z., Yuan, Z., Wang, X., Zeng, A., Xiong, Y., Xu, Q., and Shan, Y. MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions. In NeurIPS Datasets and Benchmarks, 2024.
  45. 45.Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. Model-Based Reinforcement Learning for Atari. In ICLR, 2020.
  46. 46.Kannan, H., Hafner, D., Finn, C., and Erhan, D. RoboDesk: A Multi-Task Reinforcement Learning Benchmark. https://github.com/google-research/robodesk, 2021.
  47. 47.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the Design Space of Diffusion-Based Generative Models. In NeurIPS, 2022.
  48. 48.Kazemi, N., Savov, N., Paudel, D., and Van Gool, L. Learning Generative Interactive Environments by Trained Agent Exploration. In NeurIPS Workshops, 2024.
  49. 49.Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al. OpenVLA: An Open-Source Vision-Language-Action Model. In CoRL, 2024.
  50. 50.Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S. Learning to Simulate Dynamic Environments with GameGAN. In CVPR, 2020.
  51. 51.Kim, S. W., Philion, J., Torralba, A., and Fidler, S. DriveGAN: Towards a Controllable High-Quality Neural Simulation. In CVPR, 2021.
  52. 52.Kingma, D. P. K. and Welling, M. Auto-Encoding Variational Bayes. In ICLR, 2014.
  53. 53.Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B. Learning to Act from Actionless Videos through Dense Correspondences. In ICLR, 2024.
  54. 54.Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603, 2024.
  55. 55.Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al. Multi-Game Decision Transformers. In NeurIPS, 2022.
  56. 56.Ling, P., Bu, J., Zhang, P., Dong, X., Zang, Y., Wu, T., Chen, H., Wang, J., and Jin, Y. MotionClone: Training-Free Motion Cloning for Controllable Video Generation. In ICLR, 2025.
  57. 57.Liu, B., Zhu, Y., Gao, C., Feng, Y., Liu, Q., Zhu, Y., and Stone, P. LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning. In NeurIPS Datasets and Benchmarks, 2023.
  58. 58.Loshchilov, I. and Hutter, F. Decoupled Weight Decay Regularization. In ICLR, 2019.
  59. 59.Lu, T., Shu, T., Xiao, J., Ye, L., Wang, J., Peng, C., Wei, C., Khashabi, D., Chellappa, R., Yuille, A., et al. GenEx: Generating an Explorable World. arXiv preprint arXiv:2412.09624, 2024.
  60. 60.Lu, T., Shu, T., Yuille, A., Khashabi, D., and Chen, J. Generative World Explorer. In ICLR, 2025.
  61. 61.Mazzaglia, P., Verbelen, T., Dhoedt, B., Courville, A., and Rajeswar, S. GenRL: Multimodal-Foundation World Models for Generalization in Embodied Agents. In NeurIPS, 2024.
  62. 62.McInnes, L., Healy, J., and Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv preprint arXiv:1802.03426, 2018.
  63. 63.Menapace, W., Lathuiliere, S., Tulyakov, S., Siarohin, A., and Ricci, E. Playable Video Generation. In CVPR, 2021.
  64. 64.Menapace, W., Lathuiliere, S., Siarohin, A., Theobalt, C., Tulyakov, S., Golyanik, V., and Ricci, E. Playable Environments: Video Manipulation in Space and Time. In CVPR, 2022.
  65. 65.Mendonca, R., Bahl, S., and Pathak, D. Structured World Models from Human Videos. In RSS, 2023.
  66. 66.Micheli, V., Alonso, E., and Fleuret, F. Transformers are Sample-Efficient World Models. In ICLR, 2023.
  67. 67.Nagabandi, A., Konolige, K., Levine, S., and Kumar, V. Deep Dynamics Models for Learning Dexterous Manipulation. In CoRL, 2020.
  68. 68.Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J. Gotta Learn Fast: A New Benchmark for Generalization in RL. arXiv preprint arXiv:1804.03720, 2018.
  69. 69.Nikulin, A., Zisman, I., Tarasov, D., Lyubaykin, N., Polubarov, A., Kiselev, I., and Kurenkov, V. Latent Action Learning Requires Supervision in the Presence of Distractors. arXiv preprint arXiv:2502.00379, 2025.
  70. 70.O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., et al. Open X-Embodiment: Robotic Learning Datasets and RT-X Models. In ICRA, 2024.
  71. 71.Pearce, T., Rashid, T., Bignell, D., Georgescu, R., Devlin, S., and Hofmann, K. Scaling Laws for Pre-Training Agents and World Models. arXiv preprint arXiv:2411.04434, 2024.
  72. 72.Peebles, W. and Xie, S. Scalable Diffusion Models with Transformers. In ICCV, 2023.
  73. 73.Poggio, T. and Bizzi, E. Generalization in Vision and Motor Control. Nature, 2004.
  74. 74.Qi, H., Yin, H., Du, Y., and Yang, H. Strengthening Generative Robot Policies through Predictive World Modeling. arXiv preprint arXiv:2502.00622, 2025.
  75. 75.Raad, M. A., Ahuja, A., Barros, C., Besse, F., Bolt, A., Bolton, A., Brownfield, B., Buttimore, G., Cant, M., Chakera, S., et al. Scaling Instructable Agents Across Many Simulated Worlds. arXiv preprint arXiv:2404.10179, 2024.
  76. 76.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. A Generalist Agent. In TMLR, 2022.
  77. 77.Ren, Z., Wei, Y., Guo, X., Zhao, Y., Kang, B., Feng, J., and Jin, X. VideoWorld: Exploring Knowledge Learning from Unlabeled Videos. In CVPR, 2025.
  78. 78.Rigter, M., Gupta, T., Hilmkil, A., and Ma, C. AVID: Adapting Video Diffusion Models to World Models. arXiv preprint arXiv:2410.12822, 2024.
  79. 79.Rizzolatti, G., Fadiga, L., Gallese, V., and Fogassi, L. Premotor Cortex and the Recognition of Motor Actions. Cognitive Brain Research, 1996.
  80. 80.Romo, R., Hernandez, A., and Zainos, A. Neuronal Correlates of a Perceptual Decision in Ventral Premotor Cortex. Neuron, 2004.
  81. 81.Ruhe, D., Heek, J., Salimans, T., and Hoogeboom, E. Rolling Diffusion Models. In ICML, 2024.
  82. 82.Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A. Learning What You Can Do before Doing Anything. In ICLR, 2019.
  83. 83.Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al. Habitat: A Platform for Embodied AI Research. In ICCV, 2019.
  84. 84.Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C. Learning Predictive Models From Observation and Interaction. In ECCV, 2020.
  85. 85.Schmidt, D. and Jiang, M. Learning to Act without Actions. In ICLR, 2024.
  86. 86.Seo, Y., Lee, K., James, S. L., and Abbeel, P. Reinforcement Learning with Action-Free Pre-Training from Videos. In ICML, 2022.
  87. 87.Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., and Abbeel, P. Masked World Models for Visual Control. In CoRL, 2023.
  88. 88.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing, 2024.
  89. 89.Sun, Y., Zhou, H., Yuan, L., Sun, J. J., Li, Y., Jia, X., Adam, H., Hariharan, B., Zhao, L., and Liu, T. Video Creation by Demonstration. arXiv preprint arXiv:2412.09551, 2024.
  90. 90.Sutton, R. S. Dyna, an Integrated Architecture for Learning, Planning, and Reacting. ACM Sigart Bulletin, 1991.
  91. 91.Sutton, R. S. and Barto, A. G. Reinforcement Learning: An Introduction. MIT Press, 2018.
  92. 92.Tian, S., Finn, C., and Wu, J. A Control-Centric Benchmark for Video Prediction. In ICLR, 2023.
  93. 93.Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards Accurate Generative Models of Video: A New Metric & Challenges. arXiv preprint arXiv:1812.01717, 2018.
  94. 94.Valevski, D., Leviathan, Y., Arar, M., and Fruchter, S. Diffusion Models are Real-Time Game Engines. In ICLR, 2025.
  95. 95.Van Den Oord, A., Vinyals, O., et al. Neural Discrete Representation Learning. In NeurIPS, 2017.
  96. 96.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is All You Need. In NeurIPS, 2017.
  97. 97.Villar-Corrales, A. and Behnke, S. PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning. arXiv preprint arXiv:2502.07600, 2025.
  98. 98.Wang, L., Zhao, K., Liu, C., and Chen, X. Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression. arXiv preprint arXiv:2502.04296, 2025.
  99. 99.Wang, X., Zhu, Z., Huang, G., Chen, X., Zhu, J., and Lu, J. DriveDreamer: Towards Real-World-Driven World Models for Autonomous Driving. In ECCV, 2024.
  100. 100.Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images. In NeurIPS, 2015.
  101. 101.Willi, T., Jackson, M. T., and Foerster, J. N. Jafar: An Open-Source Genie Reimplemention in JAX. In ICML Workshops, 2024.
  102. 102.Williams, G., Drews, P., Goldfain, B., Rehg, J. M., and Theodorou, E. A. Aggressive Driving with Model Predictive Path Integral Control. In ICRA, 2016.
  103. 103.Wu, J., Ma, H., Deng, C., and Long, M. Pre-Training Contextualized World Models with In-the-Wild Videos for Reinforcement Learning. In NeurIPS, 2023.
  104. 104.Wu, J., Yin, S., Feng, N., He, X., Li, D., Hao, J., and Long, M. iVideoGPT: Interactive VideoGPTs are Scalable World Models. In NeurIPS, 2024.
  105. 105.Wu, P., Escontrela, A., Hafner, D., Abbeel, P., and Goldberg, K. DayDreamer: World Models for Physical Robot Learning. In CoRL, 2022.
  106. 106.Wu, Y., Tian, R., Swamy, G., and Bajcsy, A. From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment. arXiv preprint arXiv:2502.01828, 2025.
  107. 107.Xiang, J., Liu, G., Gu, Y., Gao, Q., Ning, Y., Zha, Y., Feng, Z., Tao, T., Hao, S., Shi, Y., et al. Pandora: Towards General World Model with Natural Language Actions and Video States. arXiv preprint arXiv:2406.09455, 2024.
  108. 108.Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Yu, F., Tao, D., and Geiger, A. Unifying Flow, Stereo and Depth Estimation. IEEE TPAMI, 2023a.
  109. 109.Xu, M., Xu, Z., Chi, C., Veloso, M., and Song, S. XSkill: Cross Embodiment Skill Discovery. In CoRL, 2023b.
  110. 110.Yang, J., Gao, S., Qiu, Y., Chen, L., Li, T., Dai, B., Chitta, K., Wu, P., Zeng, J., Luo, P., et al. Generalized Predictive Model for Autonomous Driving. In CVPR, 2024a.
  111. 111.Yang, M., Du, Y., Dai, B., Schuurmans, D., Tenenbaum, J. B., and Abbeel, P. Probabilistic Adaptation of Text-to-Video Models. In ICLR, 2024b.
  112. 112.Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. Learning Interactive Real-World Simulators. In ICLR, 2024c.
  113. 113.Yang, M., Li, J., Fang, Z., Chen, S., Yu, Y., Fu, Q., Yang, W., and Ye, D. Playable Game Generation. arXiv preprint arXiv:2412.00887, 2024d.
  114. 114.Yang, S., Walker, J., Parker-Holder, J., Du, Y., Bruce, J., Barreto, A., Abbeel, P., and Schuurmans, D. Video as the New Language for Real-World Decision Making. In ICML, 2024e.
  115. 115.Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. In ICLR, 2025.
  116. 116.Yatim, D., Fridman, R., Bar-Tal, O., Kasten, Y., and Dekel, T. Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer. In CVPR, 2024.
  117. 117.Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al. Latent Action Pretraining from Videos. In ICLR, 2025.
  118. 118.Ye, W., Zhang, Y., Abbeel, P., and Gao, Y. Become a Proficient Player with Limited Data through Watching Pure Videos. In ICLR, 2023.
  119. 119.Yin, T., Zhang, Q., Zhang, R., Freeman, W. T., Durand, F., Shechtman, E., and Huang, X. From Slow Bidirectional to Fast Causal Video Generators. In CVPR, 2025.
  120. 120.Yu, J., Qin, Y., Wang, X., Wan, P., Zhang, D., and Liu, X. GameFactory: Creating New Games with Generative Interactive Videos. arXiv preprint arXiv:2501.08325, 2025.
  121. 121.Zhang, J., Zhu, R., and Ohn-Bar, E. SelfD: Self-Learning Large-Scale Driving Policies From the Web. In CVPR, 2022.
  122. 122.Zhang, L., Kan, M., Shan, S., and Chen, X. PreLAR: World Model Pre-Training with Learnable Action Representation. In ECCV, 2024.
  123. 123.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR, 2018.
  124. 124.Zhen, H., Qiu, X., Chen, P., Yang, J., Yan, X., Du, Y., Hong, Y., and Gan, C. 3D-VLA: A 3D Vision-Language-Action Generative World Model. In ICML, 2024.
  125. 125.Zhou, G., Pan, H., LeCun, Y., and Pinto, L. DINO-WM: World Models on Pre-Trained Visual Features enable Zero-Shot Planning. arXiv preprint arXiv:2411.04983, 2024a.
  126. 126.Zhou, S., Du, Y., Chen, J., Li, Y., Yeung, D.-Y., and Gan, C. RoboDreamer: Learning Compositional World Models for Robot Imagination. In ICML, 2024b.
  127. 127.Zhu, F., Wu, H., Guo, S., Liu, Y., Cheang, C., and Kong, T. IRASim: Learning Interactive Real-Robot Action Simulators. arXiv preprint arXiv:2406.14540, 2024.
  128. 128.Zhu, Y., Wong, J., Mandlekar, A., Martín-Martín, R., Joshi, A., Nasiriany, S., and Zhu, Y. robosuite: A Modular Simulation Framework and Benchmark for Robot Learning. arXiv preprint arXiv:2009.12293, 2020.

Citation

MLA
Gao, S., et al. “AdaWorld: Learning Adaptable World Models with Latent Actions”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.18938.
APA
Gao, S., Zhou, S., Du, Y., Zhang, J., & Gan, C. (2025). AdaWorld: Learning Adaptable World Models with Latent Actions. arXiv. https://doi.org/10.48550/arxiv.2503.18938
Chicago
Gao, S., S. Zhou, Y. Du, J. Zhang, and C. Gan. 2025. “AdaWorld: Learning Adaptable World Models with Latent Actions”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.18938.
Harvard
Gao, S. et al. (2025) “AdaWorld: Learning Adaptable World Models with Latent Actions”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.18938.
Vancouver
1. Gao S, Zhou S, Du Y, Zhang J, Gan C (2025) AdaWorld: Learning Adaptable World Models with Latent Actions. https://doi.org/10.48550/arxiv.2503.18938

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.18938,
  doi = {10.48550/ARXIV.2503.18938},
  url = {https://arxiv.org/abs/2503.18938},
  author = {Gao, Shenyuan and Zhou, Siyuan and Du, Yilun and Zhang, Jun and Gan, Chuang},
  keywords = {Artificial Intelligence (cs.AI), Computer Vision and Pattern Recognition (cs.CV), Machine Learning (cs.LG), Robotics (cs.RO), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {AdaWorld: Learning Adaptable World Models with Latent Actions},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/