DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Gaoyue ZhouHengkai PanYann LeCunLerrel Pinto

article2025ICML436 citations

Presents DINO-WM, a method that builds task-agnostic visual world models directly over frozen DINOv2 patch features to enable zero-shot, goal-directed planning from offline datasets without requiring image reconstruction, expert demonstrations, or reward modeling.

Listen

Building autonomous systems that can generalize across different physical tasks remains a fundamental challenge in robotics and artificial intelligence. Most current systems rely on fixed, feed-forward policies that map visual observations directly to actions without performing runtime reasoning. While predictive "world models" offer an alternative by simulating potential futures to plan actions, existing methods typically require online task-specific retraining, handcrafted reward functions, expert demonstrations, or computationally expensive pixel-level video generation. The article addresses these bottlenecks by evaluating whether a general-purpose, task-agnostic world model can be trained purely on offline behavioral datasets and achieve zero-shot goal reaching at test time.

The main objective of the article is to demonstrate DINO-WM (DINO World Model), a framework that models visual physical dynamics in a compact feature space using pre-trained visual representations rather than raw image pixels. The authors evaluate this approach across six diverse simulation suites—spanning 2D maze navigation, robotic reaching, tabletop object manipulation, and deformable rope and granular material interactions. The framework uses a frozen DINOv2 vision encoder to convert camera frames into spatial patch embeddings, trains a lightweight transformer with a causal attention mask to predict future embeddings from action histories, and applies model predictive control via the cross-entropy method to optimize action sequences toward visual target goals at runtime.

The article demonstrates several key findings. First, DINO-WM matches or substantially outperforms state-of-the-art world models across all tested domains. On contact-rich and deformable manipulation tasks, it improves goal-reaching success by an average of 45% over prior methods, achieving a 90% success rate on the complex Push-T benchmark where leading baselines achieved 30% to 32%. Second, predicting spatial patch features proved vastly superior to using global image vectors or raw pixel generation; on the hardest tasks, the model's decoded rollout predictions improved perceptual similarity metrics by 56% compared to prior art. Third, the system demonstrated strong zero-shot generalization across unseen environment configurations, such as randomized room layouts and novel object shapes. Finally, performance scaled monotonically with the amount of offline training data, rising from an 8% success rate with 200 trajectories to 92% with 18,500 trajectories.

These results demonstrate that decoupling dynamics modeling from image reconstruction and task-specific reward engineering dramatically improves planning efficiency and generalization. By operating directly in a pre-trained latent space, DINO-WM avoids the computational overhead of diffusion-based video models and the task fragility of online reinforcement learning. This shift enables faster deployment cycles and lowers development costs, allowing a single general model trained on passive or noisy interaction data to solve multiple visual goals without requiring real-time simulation or human reward labeling.

Organizations developing autonomous manipulation or robotic control systems should consider adopting patch-based pre-trained visual representations for model-based planning over pixel-level prediction pipelines. Teams evaluating this approach should begin by auditing offline dataset size and coverage, as sufficient trajectory diversity is critical for robust forward predictions. Future work and pilot evaluations should focus on testing the framework on physical hardware platforms, combining offline models with active exploration strategies for out-of-distribution scenarios, and exploring hierarchical control architectures that link high-level visual planning to fine-grained low-level motor controllers.

Confidence in these findings is high across simulated benchmarks, but several limitations should be noted. The system assumes access to offline datasets containing ground-truth agent actions and reasonable state coverage, which limits direct training on uncurated internet videos. Furthermore, all evaluations were conducted in simulated environments, meaning performance may vary when exposed to real-world sensory noise, severe lighting shifts, or unmodeled physical interactions.

Cover for DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Abstract

The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remain challenging to learn and are typically developed for task-specific solutions with online policy learning. To unlock world models’ true potential, we argue that they should 1) be trainable on offline, pre-collected trajectories, 2) support test-time behavior optimization, and 3) facilitate task-agnostic reasoning. To this end, we present DINO World Model (DINO-WM), a new method to model visual dynamics without reconstructing the visual world. DINO-WM leverages spatial patch features pre-trained with DINOv2, enabling it to learn from offline behavioral trajectories by predicting future patch features. This allows DINO-WM to achieve observational goals through action sequence optimization, facilitating task-agnostic planning by treating goal features as prediction targets. We demonstrate that DINO-WM achieves zero-shot behavioral solutions at test time on six environments without expert demonstrations, reward modeling, or pre-learned inverse models, outperforming prior state-of-the-art work across diverse task families such as arbitrarily configured mazes, push manipulation with varied object shapes, and multi-particle scenarios.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. DINO World Models
  • 3.1. DINO-based World Models (DINO-WM)
  • 3.1.1. OBSERVATION MODEL
  • 3.1.2. TRANSITION MODEL
  • 3.1.3. DECODER FOR INTERPRETABILITY
  • 3.2. Visual Planning with DINO-WM
  • 4. Experiments
  • 4.1. Environments and Tasks
  • 4.2. Baselines
  • 4.3. Optimizing Behaviors with DINO-WM
  • 4.4. Does pre-trained visual representations matter?
  • 4.5. Generalizing to Novel Environment Configurations
  • 4.6. Qualitative Comparisons with Generative Video Models
  • 4.7. Decoding and Interpreting the Latents
  • 4.8. Scaling Laws of DINO-WM
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Appendix
  • A.1. Environments and Dataset Generation
  • A.2. Environment Families for Testing Generalization
  • A.3. Pretraining Features
  • A.4. Ablations
  • A.4.1. DINO-WM WITH VS. WITHOUT CAUSAL ATTENTION MASK
  • A.4.2. DINO-WM WITH RECONSTRUCTION LOSS
  • A.5. Planning Optimization
  • A.5.1. MODEL PREDICTIVE CONTROL WITH CROSS-ENTROPY METHOD
  • A.5.2. GRADIENT DESCENT:
  • A.5.3. PLANNING RESULTS
  • A.6. Inference Time
  • A.7. Hyperparameters and Implementation
  • A.8. Additional Planning Visualizations

Knowls

  1. Knowl 1 — DINO-WM learns task-agnostic dynamics in pretrained visual features

    model/method

    DINO-WM is an offline visual world model that predicts how an environment changes under actions without learning to reconstruct images as part of its dynamics objective. A frozen pretrained DINOv2 image encoder maps observations to spatial patch features; a learned transition model predicts future features from histories of features and actions. At test time, a goal is supplied as an RGB observation, encoded by the same DINOv2 encoder, and treated as a prediction target for action-sequence optimization. This permits goal-conditioned planning without training a task-specific reward model, policy, or inverse model. DINO-WM does require offline trajectories containing actions, and the dataset must cover the dynamics needed for planning.

  2. Knowl 2 — Frozen patch features and frame-causal transformer dynamics

    model/method

    For an RGB observation oto_t, the frozen DINOv2 encoder produces patch features zt∈RN×Ez_t\in\mathbb{R}^{N\times E}, where NN is the number of image patches and EE is the feature dimension. In the reported implementation, DINOv2 yields 14×1414\times14 patches with 384-dimensional features. A ViT-based transition model predicts the next feature array from a history of latent states and their corresponding actions. Its causal attention mask makes each predicted frame depend only on past frames; patches in a frame are predicted together, rather than autoregressively one patch at a time. The action vector is mapped through an MLP and concatenated to every patch vector; proprioception, when available, is incorporated similarly. The shared predictor configuration has six transformer layers, 16 attention heads, an MLP dimension of 2048, and approximately 19 million parameters. The observation encoder stays frozen during both training and testing.

  3. Knowl 3 — Latent prediction training is independent of image decoding

    equation

    Training uses teacher forcing on trajectory segments: the transition model receives encoded past observations and actions and is trained to match the encoded next observation. For an RGB observation oto_t, let zt=enc⁡(ot)z_t=\operatorname{enc}(o_t) be its frozen DINOv2 patch features; let ata_t be the environment action; let ϕ\phi be the learned action embedding; and let pθp_\theta be the learned transition model with parameters θ\theta. For a context of length HH, the one-step latent prediction loss is

    Lpred=∥pθ(zt−H:t−1,ϕ(at−H:t−1))−zt∥22.\mathcal{L}_{\mathrm{pred}}=\left\|p_\theta\bigl(z_{t-H:t-1},\phi(a_{t-H:t-1})\bigr)-z_t\right\|_2^2.

    The loss is applied across predicted frames in the training segments. No pixel reconstruction loss is needed to train the transition model. An optional decoder qψq_\psi maps a feature array ztz_t to a reconstructed image and is trained separately on dataset observations using Lrec=∥qψ(zt)−ot∥22\mathcal{L}_{\mathrm{rec}}=\|q_\psi(z_t)-o_t\|_2^2, where ψ\psi denotes decoder parameters. This decoder is for visualization and interpretability; its loss is not propagated into the predictor.

  4. Knowl 4 — Goal reaching uses latent terminal cost with CEM-based MPC

    algorithm

    Given a current RGB observation o0o_0 and a goal RGB observation ogo_g, DINO-WM encodes them as z^0=enc⁡(o0)\hat z_0=\operatorname{enc}(o_0) and zg=enc⁡(og)z_g=\operatorname{enc}(o_g). For a candidate action sequence a0:T−1a_{0:T-1}, the transition model rolls forward to z^T\hat z_T and assigns terminal cost C=∥z^T−zg∥22C=\|\hat z_T-z_g\|_2^2, the squared feature-space distance to the goal. Cross-entropy method (CEM) optimizes action sequences: sample a population of length-TT sequences from an initially Gaussian distribution, predict each sequence’s latent rollout and cost, retain the KK lowest-cost sequences, and update the sampling distribution’s mean and covariance from those elites. Repeat sampling and updating until success or the chosen iteration limit. In model-predictive control (MPC), execute the first kk actions of the selected sequence, obtain a new environment observation, and replan. The reported CEM timing configuration used 100 samples per iteration and 10 optimization iterations; the paper does not give universal values for TT, KK, or kk. The test-time objective requires a goal image rather than a reward function.

  5. Knowl 5 — Evaluation spans six offline visual-control suites

    experimental setup

    DINO-WM was evaluated on PointMaze navigation (reported as Maze), navigation through a wall and door (Wall), two-joint arm reaching (Reach), PushT object manipulation, rope manipulation, and granular-particle manipulation. Each task asks the agent to reach a goal observation from an initial state; PushT goals are selected to be feasible within 25 steps, and Granular goals arrange particles into a square with randomized location and scale. Observations are RGB images. The models are trained from precollected trajectories and evaluated by test-time planning. For Maze, Reach, PushT, and Wall, the paper evaluates success on 50 initial-state/goal pairs; Rope and Granular are evaluated on 10 instances using Chamfer distance (CD), for which lower is better. The reported dataset-size entries include 2,000 for PointMaze, 3,000 for Reach, 18,500 for PushT, 1,920 for Wall, and 1,000 each for Rope and Granular. PushT training data were generated from released trajectories with varying noise; thus, the method’s lack of task-specific reward or policy supervision at planning time should not be read as a claim that every training dataset is demonstration-free.

  6. Knowl 6 — DINO-WM improves planning scores, especially on manipulation tasks

    data/table

    The comparison evaluates offline-trained world models by success rate (SR; higher is better) on Maze, Wall, Reach, and PushT, and by Chamfer distance (CD; lower is better) on Rope and Granular. IRIS, DreamerV3, and TD-MPC2 are evaluated without reward or task information in their offline world-model training. DINO-WM is close to DreamerV3 on Maze and Wall and attains higher success rates on Reach and PushT; it also has the lowest CD on both deformable manipulation tasks.

    Model Maze SR ↑\uparrow Wall SR ↑\uparrow Reach SR ↑\uparrow PushT SR ↑\uparrow Rope CD ↓\downarrow Granular CD ↓\downarrow
    IRIS 0.74 0.04 0.18 0.32 1.11 0.37
    DreamerV3 1.00 1.00 0.64 0.30 2.49 1.05
    TD-MPC2 0.00 0.00 0.00 0.00 2.52 1.21
    DINO-WM 0.98 0.96 0.92 0.90 0.41 0.26
  7. Knowl 7 — Planning transfers to unseen environment configurations

    data/table

    The configuration-generalization tests use WallRandom (unseen wall and door placements), PushObj (two test-time object shapes withheld from training on four shapes), and GranularRandom (a different particle count from the fixed-count training scenes). The same DINO-WM models used for the corresponding fixed-configuration tasks are evaluated on randomized goals; for GranularRandom, the reported metric is CD, with lower values better. DINO-WM has the best reported score in all three suites, although PushObj remains difficult for every method.

    Model WallRandom SR ↑\uparrow PushObj SR ↑\uparrow GranularRandom CD ↓\downarrow
    IRIS 0.06 0.14 0.86
    DreamerV3 0.76 0.18 1.53
    R3M 0.40 0.16 1.12
    ResNet 0.40 0.14 0.98
    DINO CLS 0.64 0.18 1.36
    DINO-WM 0.82 0.34 0.63
  8. Knowl 8 — Spatial DINOv2 patch features outperform global encodings for control

    empirical result

    Replacing DINOv2 patch features with global image encodings degrades planning most clearly on tasks requiring precise spatial control. In the following comparison, SR is higher-is-better for Maze, Wall, Reach, and PushT, while CD is lower-is-better for Rope and Granular. DINOv2 patch features achieve the best reported result in every column; the global DINO CLS, R3M, and ResNet encoders are particularly weaker on manipulation than on Maze. A separate linear-probe analysis maps encoder features to environment state: patch features have lower validation loss than the tested global features on PointMaze, PushT, and Wall. For that probe, patch embeddings are flattened, projected to 1536 dimensions, and passed to a linear model; lower loss indicates more linearly recoverable state information.

    Encoder Maze SR ↑\uparrow Wall SR ↑\uparrow Reach SR ↑\uparrow PushT SR ↑\uparrow Rope CD ↓\downarrow Granular CD ↓\downarrow
    R3M 0.94 0.34 0.40 0.42 1.13 0.95
    ResNet 0.98 0.12 0.06 0.20 1.08 0.90
    DINO CLS 0.96 0.58 0.60 0.44 0.84 0.79
    DINOv2 patch 0.98 0.96 0.92 0.90 0.41 0.26
    Encoder features PointMaze validation loss PushT validation loss Wall validation loss
    DINO-S Patch 0.017 0.434 0.184
    DINO-B Patch 0.014 0.504 0.163
    DINO-S CLS 0.475 0.833 0.519
    Pre-trained MAE 0.856 0.804 0.711
    R3M 0.192 0.902 0.539
  9. Knowl 9 — Decoded open-loop predictions have better perceptual and structural fidelity

    data/table

    The paper compares decoded future predictions against ground-truth frames in open-loop rollouts on PushT, Wall, Rope, and Granular. LPIPS measures perceptual discrepancy (lower is better), while SSIM measures structural similarity (higher is better). DINO-WM has the lowest LPIPS and highest SSIM among the listed methods in every environment, despite not training its predictor with pixel reconstruction loss. The decoder used for visualization is trained separately from the dynamics predictor.

    Method PushT LPIPS ↓\downarrow Wall LPIPS ↓\downarrow Rope LPIPS ↓\downarrow Granular LPIPS ↓\downarrow
    R3M 0.045 0.008 0.023 0.080
    ResNet 0.063 0.002 0.025 0.080
    DINO CLS 0.039 0.004 0.029 0.086
    AVDC 0.046 0.030 0.060 0.106
    DINO-WM 0.007 0.0016 0.009 0.035
    Method PushT SSIM ↑\uparrow Wall SSIM ↑\uparrow Rope SSIM ↑\uparrow Granular SSIM ↑\uparrow
    R3M 0.956 0.994 0.982 0.917
    ResNet 0.950 0.996 0.980 0.915
    DINO CLS 0.973 0.996 0.980 0.912
    AVDC 0.959 0.983 0.979 0.909
    DINO-WM 0.985 0.997 0.985 0.940
  10. Knowl 10 — PushT planning and prediction quality improve with offline dataset size

    data/table

    DINO-WM was trained on PushT datasets with reported sizes from 200 to 18,500 and evaluated on planning success rate (SR), decoded prediction SSIM, and decoded prediction LPIPS. All three metrics improve consistently as the dataset grows: SR and SSIM rise, while LPIPS falls. This result shows a positive dataset-size trend in both world-model prediction quality and goal-reaching performance in this task; it does not establish a scaling law beyond the tested range.

    Dataset size SR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    200 0.08 0.949 0.056
    1,000 0.48 0.973 0.013
    5,000 0.72 0.981 0.007
    10,000 0.88 0.984 0.006
    18,500 0.92 0.987 0.005
  11. Knowl 11 — Frame-causal attention prevents future-frame leakage during training

    empirical result

    On PushT, the authors compare DINO-WM with and without causal attention as the observation-history length hh increases. Without the mask, a training prediction can attend to future frames unavailable at test time, and planning success falls sharply for longer histories. The causal mask avoids that mismatch; success rises with history length in this experiment, consistent with the model using additional temporal information such as velocity and momentum.

    Attention h=1h=1 SR h=2h=2 SR h=3h=3 SR
    Without causal mask 0.76 0.36 0.08
    With causal mask 0.76 0.88 0.92
  12. Knowl 12 — Separating decoder training from prediction improves PushT planning

    empirical result

    On PushT, the predictor trained independently of the image decoder achieves a success rate of 0.92, compared with 0.80 when the predictor also receives a reconstruction loss propagated from the decoder. The result supports the design choice to keep pixel reconstruction out of the dynamics predictor’s training objective; the separately trained decoder can still be used to visualize predictions.

  13. Knowl 13 — Offline coverage and action labels constrain the method

    limitation

    DINO-WM assumes access to offline trajectories with sufficiently broad state-action coverage; obtaining such coverage may be difficult in complex environments, and out-of-coverage dynamics can undermine predictions and planning. Training also requires ground-truth actions, which may not be available in large collections of internet video. The paper proposes exploration and subsequent model updates as a possible way to address coverage, but does not evaluate that extension. Its demonstrated planner optimizes actions directly; a hierarchical combination of high-level planning and low-level control is left as future work for finer-grained tasks.

Coverage note — The auxiliary gradient-descent planner comparison and qualitative AVDC rollout comparisons are omitted: CEM-based MPC is the planner used for the principal results, while those comparisons do not add a separate core method or quantitative conclusion beyond the included planning and prediction evaluations.

References

  1. 1.Agarwal, A., Kumar, A., Malik, J., and Pathak, D. Legged locomotion in challenging terrains using egocentric vision, 2022. URL https://arxiv.org/abs/2211.07638.
  2. 2.Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., and Ballas, N. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15619–15629, 2023.
  3. 3.Astolfi, A., Karagiannis, D., and Ortega, R. Nonlinear and adaptive control with applications, volume 187. Springer, 2008.
  4. 4.Bardes, A., Garrido, Q., Ponce, J., Chen, X., Rabbat, M., LeCun, Y., Assran, M., and Ballas, N. V-JEPA: Latent video prediction for visual representation learning, 2024. URL https://openreview.net/forum?id=WFYbBOEOtv.
  5. 5.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., Florence, P., Fu, C., Arenas, M. G., Gopalakrishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Leal, I., Lee, L., Lee, T.-W. E., Levine, S., Lu, Y., Michalewski, H., Mordatch, I., Pertsch, K., Rao, K., Reymann, K., Ryoo, M., Salazar, G., Sanketi, P., Sermanet, P., Singh, J., Singh, A., Soricut, R., Tran, H., Vanhoucke, V., Vuong, Q., Wahid, A., Welker, S., Wohlhart, P., Wu, J., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023a. URL https://arxiv.org/abs/2307.15818.
  6. 6.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jackson, T., Jesmonth, S., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Leal, I., Lee, K.-H., Levine, S., Lu, Y., Malla, U., Manjunath, D., Mordatch, I., Nachum, O., Parada, C., Peralta, J., Perez, E., Pertsch, K., Quiambao, J., Rao, K., Ryoo, M., Salazar, G., Sanketi, P., Sayed, K., Singh, J., Sontakke, S., Stone, A., Tan, C., Tran, H., Vanhoucke, V., Vega, S., Vuong, Q., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B. Rt-1: Robotics transformer for real-world control at scale, 2023b. URL https://arxiv.org/abs/2212.06817.
  7. 7.Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., Aytar, Y., Bechtle, S., Behbahani, F., Chan, S., Heess, N., Gonzalez, L., Osindero, S., Ozair, S., Reed, S., Zhang, J., Zolna, K., Clune, J., de Freitas, N., Singh, S., and Rocktaschel, T. Genie: Generative interactive environments, 2024. URL https://arxiv.org/abs/2402.15391.
  8. 8.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers, 2021. URL https://arxiv.org/abs/2104.14294.
  9. 9.Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S. Diffusion policy: Visuomotor policy learning via action diffusion, 2024. URL https://arxiv.org/abs/2303.04137.
  10. 10.Chua, K., Calandra, R., McAllister, R., and Levine, S. Deep reinforcement learning in a handful of trials using probabilistic dynamics models, 2018. URL https://arxiv.org/abs/1805.12114.
  11. 11.Deisenroth, M. P. and Rasmussen, C. E. Pilco: A model-based and data-efficient approach to policy search. In International Conference on Machine Learning, 2011. URL https://api.semanticscholar.org/CorpusID:14273320.
  12. 12.Ding, Z., Zhang, A., Tian, Y., and Zheng, Q. Diffusion world model: Future modeling beyond step-by-step rollout for offline reinforcement learning, 2024. URL https://arxiv.org/abs/2402.03570.
  13. 13.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. URL https://arxiv.org/abs/2010.11929.
  14. 14.Du, Y., Yang, M., Dai, B., Dai, H., Nachum, O., Tenenbaum, J. B., Schuurmans, D., and Abbeel, P. Learning universal policies via text-guided video generation, 2023. URL https://arxiv.org/abs/2302.00111.
  15. 15.Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S. Visual foresight: Model-based deep reinforcement learning for vision-based robotic control, 2018. URL https://arxiv.org/abs/1812.00568.
  16. 16.Etukuru, H., Naka, N., Hu, Z., Lee, S., Mehu, J., Edsinger, A., Paxton, C., Chintala, S., Pinto, L., and Shafiullah, N. M. M. Robot utility models: General policies for zero-shot deployment in new environments. arXiv preprint arXiv:2409.05865, 2024.
  17. 17.Finn, C. and Levine, S. Deep visual foresight for planning robot motion, 2017. URL https://arxiv.org/abs/1610.00696.
  18. 18.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning, 2021. URL https://arxiv.org/abs/2004.07219.
  19. 19.Ha, D. and Schmidhuber, J. World models. 2018. doi: 10.5281/ZENODO.1207631. URL https://zenodo.org/record/1207631.
  20. 20.Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. Learning latent dynamics for planning from pixels, 2019. URL https://arxiv.org/abs/1811.04551.
  21. 21.Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. Dream to control: Learning behaviors by latent imagination, 2020. URL https://arxiv.org/abs/1912.01603.
  22. 22.Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. Mastering atari with discrete world models, 2022. URL https://arxiv.org/abs/2010.02193.
  23. 23.Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. Mastering diverse domains through world models, 2024. URL https://arxiv.org/abs/2301.04104.
  24. 24.Haldar, S., Peng, Z., and Pinto, L. Baku: An efficient transformer for multi-task policy learning, 2024. URL https://arxiv.org/abs/2406.07539.
  25. 25.Hansen, N., Wang, X., and Su, H. Temporal difference learning for model predictive control, 2022. URL https://arxiv.org/abs/2203.04955.
  26. 26.Hansen, N., Su, H., and Wang, X. Td-mpc2: Scalable, robust world models for continuous control, 2024. URL https://arxiv.org/abs/2310.16828.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  28. 28.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners, 2021. URL https://arxiv.org/abs/2111.06377.
  29. 29.Holkar, K. and Waghmare, L. M. An overview of model predictive control. International Journal of control and automation, 3(4):47–63, 2010.
  30. 30.Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. Gaia-1: A generative world model for autonomous driving, 2023.
  31. 31.Jia, Z., Thumuluri, V., Liu, F., Chen, L., Huang, Z., and Su, H. Chain-of-thought predictive control, 2024. URL https://arxiv.org/abs/2304.00776.
  32. 32.Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B. Learning to act from actionless videos through dense correspondences, 2023. URL https://arxiv.org/abs/2310.08576.
  33. 33.Lee, S., Wang, Y., Etukuru, H., Kim, H. J., Shafiullah, N. M. M., and Pinto, L. Behavior generation with latent actions, 2024. URL https://arxiv.org/abs/2403.03181.
  34. 34.Lenz, I., Knepper, R. A., and Saxena, A. Deepmpc: Learning deep latent features for model predictive control. In Robotics: Science and Systems, 2015. URL https://api.semanticscholar.org/CorpusID:10130184.
  35. 35.Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., He, L., and Sun, L. Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024. URL https://arxiv.org/abs/2402.17177.
  36. 36.Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A. Eureka: Human-level reward design via coding large language models, 2024. URL https://arxiv.org/abs/2310.12931.
  37. 37.Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D. Discovering and achieving goals via world models, 2021. URL https://arxiv.org/abs/2110.09514.
  38. 38.Mendonca, R., Bahl, S., and Pathak, D. Alan: Autonomously exploring robotic agents in the real world, 2023a. URL https://arxiv.org/abs/2302.06604.
  39. 39.Mendonca, R., Bahl, S., and Pathak, D. Structured world models from human videos, 2023b. URL https://arxiv.org/abs/2308.10901.
  40. 40.Micheli, V., Alonso, E., and Fleuret, F. Transformers are sample-efficient world models, 2023. URL https://arxiv.org/abs/2209.00588.
  41. 41.Nagabandi, A., Konoglie, K., Levine, S., and Kumar, V. Deep dynamics models for learning dexterous manipulation, 2019. URL https://arxiv.org/abs/1909.11652.
  42. 42.Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A. R3m: A universal visual representation for robot manipulation, 2022. URL https://arxiv.org/abs/2203.12601.
  43. 43.Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. Dinov2: Learning robust visual features without supervision, 2024. URL https://arxiv.org/abs/2304.07193.
  44. 44.Pathak, D., Mahmoudieh, P., Luo, G., Agrawal, P., Chen, D., Shentu, Y., Shelhamer, E., Malik, J., Efros, A. A., and Darrell, T. Zero-shot visual imitation, 2018. URL https://arxiv.org/abs/1804.08606.
  45. 45.Razavi, A., van den Oord, A., and Vinyals, O. Generating diverse high-fidelity images with vq-vae-2, 2019. URL https://arxiv.org/abs/1906.00446.
  46. 46.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N. A generalist agent, 2022. URL https://arxiv.org/abs/2205.06175.
  47. 47.Robine, J., Hoftmann, M., Uelwer, T., and Harmeling, S. Transformer-based world models are happy with 100k interactions, 2023. URL https://arxiv.org/abs/2303.07109.
  48. 48.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. Imagenet large scale visual recognition challenge, 2015. URL https://arxiv.org/abs/1409.0575.
  49. 49.Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D. Planning to explore via self-supervised world models, 2020. URL https://arxiv.org/abs/2005.05960.
  50. 50.Sutton, R. S. Dyna, an integrated architecture for learning, planning, and reacting. ACM Sigart Bulletin, 2(4):160–163, 1991.
  51. 51.Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M. Deepmind control suite, 2018. URL https://arxiv.org/abs/1801.00690.
  52. 52.Todorov, E. and Li, W. A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems. In Proceedings of the 2005, American Control Conference, 2005., pp. 300–306. IEEE, 2005.
  53. 53.Wang, J., Dasari, S., Srirama, M. K., Tulsiani, S., and Gupta, A. Manipulate by seeing: Creating manipulation controllers from pre-trained representations, 2023. URL https://arxiv.org/abs/2303.08135.
  54. 54.Wang, Z., Bovik, A., Sheikh, H., and Simoncelli, E. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004. doi: 10.1109/TIP.2003.819861.
  55. 55.Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. Embed to control: A locally linear latent dynamics model for control from raw images, 2015. URL https://arxiv.org/abs/1506.07365.
  56. 56.Wen, C., Lin, X., So, J., Chen, K., Dou, Q., Gao, Y., and Abbeel, P. Any-point trajectory modeling for policy learning, 2024. URL https://arxiv.org/abs/2401.00025.
  57. 57.Williams, G., Wagener, N., Goldfain, B., Drews, P., Rehg, J. M., Boots, B., and Theodorou, E. A. Information theoretic mpc for model-based reinforcement learning. In 2017 IEEE international conference on robotics and automation (ICRA), pp. 1714–1721. IEEE, 2017.
  58. 58.Wu, Y., Yan, W., Kurutach, T., Pinto, L., and Abbeel, P. Learning to manipulate deformable objects without demonstrations, 2020. URL https://arxiv.org/abs/1910.13439.
  59. 59.Xiao, T., Radosavovic, I., Darrell, T., and Malik, J. Masked visual pre-training for motor control, 2022. URL https://arxiv.org/abs/2203.06173.
  60. 60.Yan, W., Vangipuram, A., Abbeel, P., and Pinto, L. Learning predictive representations for deformable objects using contrastive estimation. In Conference on Robot Learning, pp. 564–574. PMLR, 2021.
  61. 61.Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. Learning interactive real-world simulators, 2023.
  62. 62.Zhang, K., Li, B., Hauser, K., and Li, Y. Adaptigraph: Material-adaptive graph-based neural dynamics for robotic manipulation, 2024. URL https://arxiv.org/abs/2407.07889.
  63. 63.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. CoRR, abs/1801.03924, 2018. URL http://arxiv.org/abs/1801.03924.
  64. 64.Zhao, T. Z., Kumar, V., Levine, S., and Finn, C. Learning fine-grained bimanual manipulation with low-cost hardware, 2023. URL https://arxiv.org/abs/2304.13705.
  65. 65.Zhou, G., Dean, V., Srirama, M. K., Rajeswaran, A., Pari, J., Hatch, K., Jain, A., Yu, T., Abbeel, P., Pinto, L., Finn, C., and Gupta, A. Train offline, test online: A real robot learning benchmark, 2023. URL https://arxiv.org/abs/2306.00942.

Citation

MLA
Zhou, G., et al. “DINO-WM: World Models on Pre-trained Visual Features Enable Zero-shot Planning”. arXiv, 2024, http://arxiv.org/abs/2411.04983v2.
APA
Zhou, G., Pan, H., LeCun, Y., & Pinto, L. (2024). DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv. http://arxiv.org/abs/2411.04983v2
Chicago
Zhou, G., H. Pan, Y. LeCun, and L. Pinto. 2024. “DINO-WM: World Models on Pre-trained Visual Features Enable Zero-shot Planning”. arXiv. http://arxiv.org/abs/2411.04983v2.
Harvard
Zhou, G. et al. (2024) “DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2411.04983v2.
Vancouver
1. Zhou G, Pan H, LeCun Y, Pinto L (2024) DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning. arXiv

BibTeX

@article{zhou2024dino,
  title = {DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning},
  author = {Zhou, Gaoyue and Pan, Hengkai and LeCun, Yann and Pinto, Lerrel},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2411.04983v2},
  eprint = {2411.04983}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/