Reinforcement Learning with Action-Free Pre-Training from Videos
Younggyo SeoKimin LeeStephen JamesPieter Abbeel
Presents an action-free video pre-training framework that stacks action-conditional world models onto pre-trained latent video dynamics, substantially improving sample efficiency and task performance across unseen visual reinforcement learning domains.
Standard vision-based reinforcement learning requires artificial intelligence agents to learn complex control behaviors entirely from scratch through trial and error. This reliance on millions of costly, real-time environment interactions creates severe data inefficiency and limits the deployment of visual control systems in real-world robotics and automation. While fields such as computer vision and natural language processing successfully overcome data scarcity by pre-training models on vast, unlabelled datasets, adapting this paradigm to reinforcement learning has remained challenging because readily available video data lacks explicit action and reward labels.
The article introduces and evaluates Action-Free Pre-training from Videos (APV), a framework designed to improve the sample efficiency and overall performance of vision-based reinforcement learning. APV demonstrates that predictive world models can pre-train purely on passive, action-free video datasets across diverse visual domains and successfully transfer that physical understanding to guide subsequent policy learning on new tasks.
The framework operates in two distinct phases using simulated benchmark environments. First, a latent video prediction model is pre-trained across 4,950 videos spanning 99 robotic manipulation tasks from the RLBench benchmark to learn environmental transitions without action labels. Second, during task-specific fine-tuning on downstream manipulation tasks from Meta-world and locomotion tasks from the DeepMind Control Suite, the authors introduce a stacked latent architecture that places an action-conditional model on top of the frozen or adapted action-free model. This stage also incorporates an intrinsic exploration bonus derived from video representations to reward the agent for visiting diverse trajectory sequences.
The evaluation produced four primary findings. First, pre-training on diverse manipulation videos substantially boosted task performance; on six Meta-world tasks, APV achieved an aggregate success rate of 95.4%, compared to 67.9% for the baseline DreamerV2 model. On difficult tasks like Lever Pull, APV achieved over a 60% success rate while the baseline failed completely. Second, the stacked latent model architecture prevented catastrophic forgetting; simple parameter re-initialization methods quickly lost pre-trained knowledge and failed to provide meaningful gains. Third, dynamics representations successfully transferred across distinct domains; models pre-trained on robotic arm videos improved sample efficiency and returns on robotic locomotion tasks (such as quadruped and hopper locomotion), despite major differences in visual appearances and objectives. Finally, combining pre-trained representations with trajectory-based intrinsic rewards provided synergetic gains, outperforming configurations relying on either feature alone.
These findings demonstrate that artificial intelligence agents do not need task-identical demonstration videos or logged motor actions to build useful operational priors. Instead, agents can internalize general physics and motion dynamics from passive video observations. This significantly reduces the training time and physical interactions needed to master control tasks, offering a path to reduce the computational and operational costs associated with robotic policy development.
Organizations developing vision-based robotic policies should consider implementing modular, stacked world models and leveraging diverse existing video repositories rather than collecting task-specific interaction data from scratch. Before full deployment, teams should conduct focused pilots to establish whether available video data aligns with downstream tasks, as domain relevance remains an important performance driver.
The results should be interpreted with some caution regarding real-world transfer. Pre-training succeeded robustly on simulated environments, but experiments using natural, human-demonstration video datasets (Something-Something-V2) suffered from underfitting and produced blurry predictions that failed to yield performance improvements. Further investigation into larger model architectures and advanced generative video transformers is required before deploying this framework directly onto natural, unconstrained real-world video sources.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Dreamer shows how learned latent dynamics can support visual control, providing essential context for this paper’s pre-trained video representations and downstream world-model learning.
- Paper: Learning Latent Dynamics for Planning from Pixels, Danijar Hafner et al. (2018). PlaNet establishes latent dynamics learning and planning from pixels, clarifying the world-model foundations that this work adapts through action-free video pre-training.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). World Models separates visual encoding, predictive dynamics, and policy learning, a foundational design that helps explain this paper’s staged representation-to-control framework.
- Paper: Generating Videos with Scene Dynamics, Carl Vondrick et al. (2016). This work learns scene dynamics from unlabeled video, introducing the generative video-prediction approach that underlies the source’s action-free pre-training phase.
- Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). Its LSTM encoder-decoder experiments establish how unsupervised future-frame prediction can yield useful video representations, a direct precursor to the source’s pre-training objective.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). VIP develops action-free visual pre-training from human videos for robotic rewards and representations, providing a close prerequisite for understanding the source’s video-based learning setup.
- Paper: Contrastive Learning as Goal-Conditioned Reinforcement Learning, Benjamin Eysenbach et al. (2022). This paper connects learned representations with goal-conditioned value functions in RL, useful groundwork for understanding how pre-trained visual features can support downstream control.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). CURL demonstrates how unsupervised visual representations improve RL sample efficiency, establishing a key representation-pretraining baseline for the source’s approach.
- Paper: Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning, Denis Yarats et al. (2022). DrQ-v2 provides a strong pixel-based control baseline and illustrates the sample-efficiency challenges that motivate pre-training for visual RL.
- Paper: Mask-based Latent Reconstruction for Reinforcement Learning, Tao Yu et al. (2022). MLR shows how predictive latent reconstruction can improve visual RL, preparing readers for the source’s use of learned latent video predictions in control.
- Paper: DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, Gaoyue Zhou et al. (2025). DINO-WM carries predictive visual representations into zero-shot planning, extending the source’s learned dynamics beyond task-specific RL fine-tuning.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). RoboDreamer builds on video-based world modeling to compose and plan robot behaviors, extending predictive visual dynamics toward language-guided manipulation.
- Paper: Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations, Yucheng Hu et al. (2025). VPP uses predictive video-model representations to guide generalist robot policies, extending the source’s central idea from RL pre-training to broad manipulation policy learning.
- Paper: Genie: Generative Interactive Environments, Jake Bruce et al. (2024). Genie learns controllable interactive environments from unlabeled video, extending action-free video dynamics modeling toward inferring actions and generating playable worlds.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). UP-VLA unifies future visual prediction with language-conditioned action generation, extending predictive representations into an integrated embodied policy.
- Paper: LIV: Language-Image Representations and Rewards for Robotic Control, Yecheng Jason Ma et al. (2023). LIV turns action-free human-video pre-training into language-grounded rewards and representations for robot control, extending the source’s video-based transfer strategy.
