DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Gaoyue ZhouHengkai PanYann LeCunLerrel Pinto
Presents DINO-WM, a method that builds task-agnostic visual world models directly over frozen DINOv2 patch features to enable zero-shot, goal-directed planning from offline datasets without requiring image reconstruction, expert demonstrations, or reward modeling.
Building autonomous systems that can generalize across different physical tasks remains a fundamental challenge in robotics and artificial intelligence. Most current systems rely on fixed, feed-forward policies that map visual observations directly to actions without performing runtime reasoning. While predictive "world models" offer an alternative by simulating potential futures to plan actions, existing methods typically require online task-specific retraining, handcrafted reward functions, expert demonstrations, or computationally expensive pixel-level video generation. The article addresses these bottlenecks by evaluating whether a general-purpose, task-agnostic world model can be trained purely on offline behavioral datasets and achieve zero-shot goal reaching at test time.
The main objective of the article is to demonstrate DINO-WM (DINO World Model), a framework that models visual physical dynamics in a compact feature space using pre-trained visual representations rather than raw image pixels. The authors evaluate this approach across six diverse simulation suites—spanning 2D maze navigation, robotic reaching, tabletop object manipulation, and deformable rope and granular material interactions. The framework uses a frozen DINOv2 vision encoder to convert camera frames into spatial patch embeddings, trains a lightweight transformer with a causal attention mask to predict future embeddings from action histories, and applies model predictive control via the cross-entropy method to optimize action sequences toward visual target goals at runtime.
The article demonstrates several key findings. First, DINO-WM matches or substantially outperforms state-of-the-art world models across all tested domains. On contact-rich and deformable manipulation tasks, it improves goal-reaching success by an average of 45% over prior methods, achieving a 90% success rate on the complex Push-T benchmark where leading baselines achieved 30% to 32%. Second, predicting spatial patch features proved vastly superior to using global image vectors or raw pixel generation; on the hardest tasks, the model's decoded rollout predictions improved perceptual similarity metrics by 56% compared to prior art. Third, the system demonstrated strong zero-shot generalization across unseen environment configurations, such as randomized room layouts and novel object shapes. Finally, performance scaled monotonically with the amount of offline training data, rising from an 8% success rate with 200 trajectories to 92% with 18,500 trajectories.
These results demonstrate that decoupling dynamics modeling from image reconstruction and task-specific reward engineering dramatically improves planning efficiency and generalization. By operating directly in a pre-trained latent space, DINO-WM avoids the computational overhead of diffusion-based video models and the task fragility of online reinforcement learning. This shift enables faster deployment cycles and lowers development costs, allowing a single general model trained on passive or noisy interaction data to solve multiple visual goals without requiring real-time simulation or human reward labeling.
Organizations developing autonomous manipulation or robotic control systems should consider adopting patch-based pre-trained visual representations for model-based planning over pixel-level prediction pipelines. Teams evaluating this approach should begin by auditing offline dataset size and coverage, as sufficient trajectory diversity is critical for robust forward predictions. Future work and pilot evaluations should focus on testing the framework on physical hardware platforms, combining offline models with active exploration strategies for out-of-distribution scenarios, and exploring hierarchical control architectures that link high-level visual planning to fine-grained low-level motor controllers.
Confidence in these findings is high across simulated benchmarks, but several limitations should be noted. The system assumes access to offline datasets containing ground-truth agent actions and reasonable state coverage, which limits direct training on uncurated internet videos. Furthermore, all evaluations were conducted in simulated environments, meaning performance may vary when exposed to real-world sensory noise, severe lighting shifts, or unmodeled physical interactions.
- Paper: Learning Latent Dynamics for Planning from Pixels, Danijar Hafner et al. (2018). PlaNet establishes the latent-dynamics planning framework that helps clarify how DINO-WM predicts and plans without reconstructing pixels.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Dreamer develops planning through imagined latent futures, providing a key predecessor for understanding DINO-WM’s use of predictive representations for control.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Diffuser shows how offline trajectories and test-time trajectory optimization can support flexible goal-directed behavior, the planning setting DINO-WM takes into learned visual dynamics.
- Paper: LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, Lucas Maes et al. (2026). LeWorldModel directly tests a next step beyond DINO-WM by learning predictive visual latents end-to-end from pixels and evaluating against DINO-WM as a baseline.
