Reinforcement Learning with Action-Free Pre-Training from Videos

Younggyo SeoKimin LeeStephen JamesPieter Abbeel

article2022ICML159 citations

Presents an action-free video pre-training framework that stacks action-conditional world models onto pre-trained latent video dynamics, substantially improving sample efficiency and task performance across unseen visual reinforcement learning domains.

Listen

Standard vision-based reinforcement learning requires artificial intelligence agents to learn complex control behaviors entirely from scratch through trial and error. This reliance on millions of costly, real-time environment interactions creates severe data inefficiency and limits the deployment of visual control systems in real-world robotics and automation. While fields such as computer vision and natural language processing successfully overcome data scarcity by pre-training models on vast, unlabelled datasets, adapting this paradigm to reinforcement learning has remained challenging because readily available video data lacks explicit action and reward labels.

The article introduces and evaluates Action-Free Pre-training from Videos (APV), a framework designed to improve the sample efficiency and overall performance of vision-based reinforcement learning. APV demonstrates that predictive world models can pre-train purely on passive, action-free video datasets across diverse visual domains and successfully transfer that physical understanding to guide subsequent policy learning on new tasks.

The framework operates in two distinct phases using simulated benchmark environments. First, a latent video prediction model is pre-trained across 4,950 videos spanning 99 robotic manipulation tasks from the RLBench benchmark to learn environmental transitions without action labels. Second, during task-specific fine-tuning on downstream manipulation tasks from Meta-world and locomotion tasks from the DeepMind Control Suite, the authors introduce a stacked latent architecture that places an action-conditional model on top of the frozen or adapted action-free model. This stage also incorporates an intrinsic exploration bonus derived from video representations to reward the agent for visiting diverse trajectory sequences.

The evaluation produced four primary findings. First, pre-training on diverse manipulation videos substantially boosted task performance; on six Meta-world tasks, APV achieved an aggregate success rate of 95.4%, compared to 67.9% for the baseline DreamerV2 model. On difficult tasks like Lever Pull, APV achieved over a 60% success rate while the baseline failed completely. Second, the stacked latent model architecture prevented catastrophic forgetting; simple parameter re-initialization methods quickly lost pre-trained knowledge and failed to provide meaningful gains. Third, dynamics representations successfully transferred across distinct domains; models pre-trained on robotic arm videos improved sample efficiency and returns on robotic locomotion tasks (such as quadruped and hopper locomotion), despite major differences in visual appearances and objectives. Finally, combining pre-trained representations with trajectory-based intrinsic rewards provided synergetic gains, outperforming configurations relying on either feature alone.

These findings demonstrate that artificial intelligence agents do not need task-identical demonstration videos or logged motor actions to build useful operational priors. Instead, agents can internalize general physics and motion dynamics from passive video observations. This significantly reduces the training time and physical interactions needed to master control tasks, offering a path to reduce the computational and operational costs associated with robotic policy development.

Organizations developing vision-based robotic policies should consider implementing modular, stacked world models and leveraging diverse existing video repositories rather than collecting task-specific interaction data from scratch. Before full deployment, teams should conduct focused pilots to establish whether available video data aligns with downstream tasks, as domain relevance remains an important performance driver.

The results should be interpreted with some caution regarding real-world transfer. Pre-training succeeded robustly on simulated environments, but experiments using natural, human-demonstration video datasets (Something-Something-V2) suffered from underfitting and produced blurry predictions that failed to yield performance improvements. Further investigation into larger model architectures and advanced generative video transformers is required before deploying this framework directly onto natural, unconstrained real-world video sources.

Cover for Reinforcement Learning with Action-Free Pre-Training from Videos

Abstract

Recent unsupervised pre-training methods have shown to be effective on language and vision domains by learning useful representations for multiple downstream tasks. In this paper, we investigate if such unsupervised pre-training methods can also be effective for vision-based reinforcement learning (RL). To this end, we introduce a framework that learns representations useful for understanding the dynamics via generative pre-training on videos. Our framework consists of two phases: we pre-train an action-free latent video prediction model, and then utilize the pre-trained representations for efficiently learning action-conditional world models on unseen environments. To incorporate additional action inputs during fine-tuning, we introduce a new architecture that stacks an action-conditional latent prediction model on top of the pre-trained action-free prediction model. Moreover, for better exploration, we propose a video-based intrinsic bonus that leverages pre-trained representations. We demonstrate that our framework significantly improves both final performances and sample-efficiency of vision-based RL in a variety of manipulation and locomotion tasks. Code is available at https://github.com/younggyoseo/apv.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Action-free Pre-training from Videos
  • 3.2. Stacked Latent Prediction Model
  • 3.3. Video-based Intrinsic Bonus
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Meta-world Experiments
  • 4.3. DeepMind Control Suite Experiments
  • 5. Discussion
  • Acknowledgements
  • References
  • A. Behavior Learning
  • B. Formulation with Recurrent State-Space Model
  • B.1. Action-free Latent Video Prediction Model
  • B.2. Stacked Latent Prediction Model
  • C. Extended Related Work
  • D. Difference to DreamerV2
  • E. Experimental Details
  • F. Meta-world Experiments with DrQ-v2
  • G. Real-World Video Prediction on Something-Something-V2
  • H. Video Prediction on RLBench and Meta-world
  • I. Ablation Study on DeepMind Control Suite

Knowls

  1. Knowl 1 — Action-free latent video prediction pre-training

    model/method

    APV pre-trains a latent dynamics model on image sequences without action labels. For an observed frame oto_t, the representation model infers a latent state ztz_t from the previous latent state and the current frame; a transition model predicts a latent state without seeing the frame; and an image decoder reconstructs the frame from the inferred latent state. The training objective combines image reconstruction with a KL penalty that aligns the inferred latent distribution with the action-free transition distribution:

    L(ϕ)=Eqϕ(z1:T∣o1:T)[∑t=1T(−log⁡pϕ(ot∣zt)+βz KL[qϕ(zt∣zt−1,ot) ∥ pϕ(z^t∣zt−1)])].\mathcal{L}(\phi)=\mathbb{E}_{q_\phi(z_{1:T}\mid o_{1:T})}\left[\sum_{t=1}^{T}\left(-\log p_\phi(o_t\mid z_t)+\beta_z\,\mathrm{KL}\left[q_\phi(z_t\mid z_{t-1},o_t)\,\|\,p_\phi(\hat z_t\mid z_{t-1})\right]\right)\right].

    Here, oto_t is an image observation, ztz_t is its inferred latent state, z^t\hat z_t is a latent state sampled from the transition model, TT is the sequence length, ϕ\phi denotes model parameters, and βz\beta_z weights the KL term. The model is trained on action-free videos, and at inference its transition model can roll latent states forward without generating and re-encoding future images.

  2. Knowl 2 — Stacked action-free and action-conditional world models

    model/method

    To fine-tune video-pretrained representations for control, APV passes the action-free model's state ztz_t into a second, action-conditional latent model. The second model infers state sts_t from ztz_t, its preceding state, and the preceding action; its transition predicts future states conditioned on actions. Image and reward predictors use sts_t. This makes the pretrained dynamics representation an input to the control model rather than directly repurposing the action-free model as an action-conditioned model.

    zt∼qϕ(zt∣zt−1,ot),z^t∼pϕ(z^t∣zt−1),st∼qθ(st∣st−1,at−1,zt),s^t∼pθ(s^t∣st−1,at−1),o^t∼pθ(o^t∣st),r^t∼pθ(r^t∣st).\begin{aligned} z_t &\sim q_\phi(z_t\mid z_{t-1},o_t), & \hat z_t &\sim p_\phi(\hat z_t\mid z_{t-1}),\\ s_t &\sim q_\theta(s_t\mid s_{t-1},a_{t-1},z_t), & \hat s_t &\sim p_\theta(\hat s_t\mid s_{t-1},a_{t-1}),\\ \hat o_t &\sim p_\theta(\hat o_t\mid s_t), & \hat r_t &\sim p_\theta(\hat r_t\mid s_t). \end{aligned}

    The fine-tuning objective sums image log loss, reward log loss, and KL penalties for the action-free and action-conditional latent models:

    L(ϕ,θ)=E[∑t=1T(−log⁡pθ(ot∣st)−log⁡pθ(rt∣st)+βz KL(qϕ∥pϕ)+β KL(qθ∥pθ))],\mathcal{L}(\phi,\theta)=\mathbb{E}\left[\sum_{t=1}^{T}\left(-\log p_\theta(o_t\mid s_t)-\log p_\theta(r_t\mid s_t)+\beta_z\,\mathrm{KL}(q_\phi\|p_\phi)+\beta\,\mathrm{KL}(q_\theta\|p_\theta)\right)\right],

    where the KL terms compare each model's inferred latent distribution with its corresponding transition distribution, at−1a_{t-1} is the action, rtr_t is the observed reward, and βz\beta_z and β\beta weight the two KL terms. The image decoder is initialized from the action-free model's decoder. In the reported fine-tuning experiments, βz=0\beta_z=0 and β=1.0\beta=1.0.

  3. Knowl 3 — Trajectory-based video intrinsic reward

    equation

    APV measures novelty over short sequences of action-free latent states rather than individual states. It average-pools a sliding window of length τ\tau into a trajectory representation yty_t, applies a random projection ψ\psi, and assigns a bonus based on the projected Euclidean distance to the representation's kk-th nearest neighbor among replay-buffer samples:

    rtint=∥ψ(yt)−ψ(yt(k))∥2,yt=Avg(zt:t+τ).r_t^{\mathrm{int}}=\left\|\psi(y_t)-\psi(y_t^{(k)})\right\|_2,\qquad y_t=\mathrm{Avg}(z_{t:t+\tau}).

    Here, yt(k)y_t^{(k)} is the kk-th nearest-neighbor trajectory representation, and rtintr_t^{\mathrm{int}} is the intrinsic reward at time tt. The reward predictor is trained on the sum of extrinsic reward and the weighted intrinsic reward, rt+λrtintr_t+\lambda r_t^{\mathrm{int}}; the actor-critic learner uses the resulting imagined rewards. Experiments use k=16k=16, a queue of 4096 recent representations, and τ=5\tau=5. The intrinsic reward is normalized to 10% of the extrinsic reward scale; λ=0.1\lambda=0.1 for manipulation and λ=1.0\lambda=1.0 for locomotion.

  4. Knowl 4 — Pre-training and downstream-task experimental setup

    experimental setup

    For manipulation, APV is pre-trained on 4,950 RLBench videos: 10 demonstrations in each of 99 tasks, rendered from five camera views. Pre-training uses 600,000 gradient steps, with sequence length T=25T=25. The downstream tasks are Meta-world Lever Pull, Drawer Open, Door Lock, Button Press Topdown Wall, Reach, and Dial Turn; fine-tuning runs for 250,000 environment steps (500 episodes), with 500 steps per episode and no action repeat. For locomotion, the downstream tasks are Quadruped Walk, Quadruped Run, and Hopper Hop in DeepMind Control Suite; fine-tuning runs for 1,000,000 environment steps with episode length 1,000 and action repeat 2. Pre-training for locomotion uses either RLBench manipulation videos or 1,000 videos collected during Triped Walk training, with T=50T=50. Both benchmarks use eight runs per task. Reported learning curves use interquartile means and bootstrap confidence intervals; aggregate results use stratified bootstrap confidence intervals. The experiments increase dense-layer hidden sizes and RSSM deterministic-state dimensions from 200 to 1024.

  5. Knowl 5 — Manipulation performance after RLBench video pre-training

    empirical result

    On six vision-based Meta-world manipulation tasks, APV pre-trained on RLBench videos achieved an aggregate success rate of 95.4%, compared with 67.9% for DreamerV2. APV also learned more sample-efficiently than DreamerV2 on each of the six tasks. On Lever Pull, APV exceeded 60% success while DreamerV2 failed to solve the task. These results were obtained with 250,000 fine-tuning environment steps and reported across eight runs per task.

  6. Knowl 6 — Pre-training and stacked fine-tuning contribute jointly

    empirical result

    Manipulation ablations show that initializing an action-conditional model naively from the action-free pretrained parameters did not yield large gains over DreamerV2, consistent with rapid loss of useful pretrained representations. By contrast, the stacked APV architecture without an intrinsic bonus achieved more than 10% higher success than DreamerV2 from the beginning of fine-tuning, using the same pretrained model. In a separate ablation, the stacked architecture without either generative pre-training or the intrinsic bonus performed similarly to DreamerV2; pre-training and the bonus each improved performance, and their combination performed best. The comparisons aggregate 48 runs across the six Meta-world tasks.

  7. Knowl 7 — Intrinsic-reward performance depends on trajectory-window length

    empirical result

    For the video-based intrinsic reward on Meta-world manipulation tasks, pooling a window of five action-free model states performed better than using windows of length one or three. A window of length ten performed worse than length five. The authors suggest that short sequences supply useful behavioral context, while pooling over longer sequences may add complexity. The reported implementation uses window length τ=5\tau=5.

  8. Knowl 8 — Pretrained representations encode useful dynamics

    empirical result

    Several analyses indicate that APV's gains rely on dynamics information learned from videos, not only on transferred image features. When APV transferred only its convolutional image encoder and decoder, without the recurrent dynamics representations, it performed worse than the full model on the six Meta-world tasks. In regressions using a pre-collected Triped dataset, models with RLBench-pretrained representations began with lower prediction error for both proprioceptive states and rewards and converged faster than randomly initialized representations. A t-SNE visualization of pooled pretrained states also grouped clips by Meta-world task, whereas states from randomly initialized representations were entangled; the Meta-world task videos were not used for that pre-training.

  9. Knowl 9 — Video representations transfer across domains, with benefits from domain similarity

    empirical result

    RLBench manipulation videos improved APV over DreamerV2 on the DeepMind Control Suite locomotion tasks Quadruped Walk, Quadruped Run, and Hopper Hop, despite differences in visuals and objectives; pre-training on the more similar Triped Walk videos improved performance further. On Meta-world, manipulation-video pre-training outperformed locomotion-video pre-training. Adding 100 Meta-world videos from training tasks to the RLBench pre-training data produced nearly similar performance to RLBench-only pre-training when all model parameters were fine-tuned. However, when the action-free representation model was frozen and used to assess representation quality, adding those in-domain videos improved performance on four Meta-world tasks not seen during pre-training.

  10. Knowl 10 — Limitations of real-world video pre-training

    limitation

    The reported pre-training was conducted on simulated robotic videos. In a test with real-world Something-Something-V2 videos, the action-free prediction model produced blurry future frames; the paper also reports that APV pre-trained on this dataset did not outperform the no-pre-training baseline on Meta-world. The authors identify limited video data and the gap between egocentric human videos and third-person robotic videos as possible factors, and note underfitting as a limitation of their video-prediction setup. They propose scaling the model or using higher-fidelity video-prediction architectures as directions for addressing this limitation.

Coverage note — DreamerV2's inherited actor-critic update details and the DrQ-v2 comparison are omitted because they are not central new components of APV.

References

  1. 1.Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. Tensorflow: A system for large-scale machine learning. In 12th {USENIX} symposium on operating systems design and implementation ({OSDI} 16), pp. 265–283, 2016.
  2. 2.Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. Deep reinforcement learning at the edge of the statistical precipice. In Advances in Neural Information Processing Systems, 2021.
  3. 3.Aigner, S. and Körner, M. Futuregan: Anticipating the future frames of video sequences using spatio-temporal 3d convolutions in progressively growing gans. arXiv preprint arXiv:1810.01325, 2018.
  4. 4.Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al. Solving rubik’s cube with a robot hand. arXiv preprint arXiv:1910.07113, 2019.
  5. 5.Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M.-A., and Hjelm, R. D. Unsupervised state representation learning in atari. In Advances in Neural Information Processing Systems, 2019.
  6. 6.Aytar, Y., Pfaff, T., Budden, D., Paine, T. L., Wang, Z., and de Freitas, N. Playing hard exploration games by watching youtube. In Advances in Neural Information Processing Systems, 2018.
  7. 7.Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S. Stochastic variational video prediction. In International Conference on Learning Representations, 2018.
  8. 8.Babaeizadeh, M., Saffar, M. T., Nair, S., Levine, S., Finn, C., and Erhan, D. Fitvid: Overfitting in pixel-level video prediction. arXiv preprint arXiv:2106.13195, 2021.
  9. 9.Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems, 2016.
  10. 10.Bengio, Y., Léonard, N., and Courville, A. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  11. 11.Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680, 2019.
  12. 12.Bingham, E. and Mannila, H. Random projection in dimensionality reduction: applications to image and text data. In Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining, 2001.
  13. 13.Burda, Y., Edwards, H., Storkey, A., and Klimov, O. Exploration by random network distillation. In International Conference on Learning Representations, 2019.
  14. 14.Castro, P. S. Scalable methods for computing state similarity in deterministic markov decision processes. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  15. 15.Chang, M., Gupta, A., and Gupta, S. Semantic visual navigation by watching youtube videos. In Advances in Neural Information Processing Systems, 2020.
  16. 16.Chen, A. S., Nair, S., and Finn, C. Learning generalizable robotic reward functions from” in-the-wild” human videos. In Proceedings of Robotics: Science and Systems, 2021.
  17. 17.Chen, C., Wu, Y.-F., Yoon, J., and Ahn, S. Transdreamer: Reinforcement learning with transformer world models. arXiv preprint arXiv:2202.09481, 2022.
  18. 18.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning, 2020.
  19. 19.Clark, A., Donahue, J., and Simonyan, K. Adversarial video generation on complex datasets. arXiv preprint arXiv:1907.06571, 2019.
  20. 20.Dasari, S., Ebert, F., Tian, S., Nair, S., Bucher, B., Schmeckpeper, K., Singh, S., Levine, S., and Finn, C. Robonet: Large-scale multi-robot learning. In Conference on Robot Learning, 2019.
  21. 21.Denton, E. and Fergus, R. Stochastic video generation with a learned prior. In International Conference on Machine Learning, 2018.
  22. 22.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019.
  23. 23.Dwibedi, D., Tompson, J., Lynch, C., and Sermanet, P. Learning actionable representations from visual observations. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2018.
  24. 24.Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C. Imitating latent policies from observation. In International Conference on Machine Learning, 2019.
  25. 25.Finn, C., Goodfellow, I., and Levine, S. Unsupervised learning for physical interaction through video prediction. In Advances in neural information processing systems, 2016a.
  26. 26.Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P. Deep spatial autoencoders for visuomotor learning. In 2016 IEEE International Conference on Robotics and Automation (ICRA), 2016b.
  27. 27.Franceschi, J.-Y., Delasalles, E., Chen, M., Lamprier, S., and Gallinari, P. Stochastic latent residual video prediction. In International Conference on Machine Learning, 2020.
  28. 28.Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. Deepmdp: Learning continuous latent space models for representation learning. In International Conference on Machine Learning, 2019.
  29. 29.Gidaris, S., Singh, P., and Komodakis, N. Unsupervised representation learning by predicting image rotations. In International Conference on Learning Representations, 2018.
  30. 30.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In Advances in Neural Information Processing Systems, 2014.
  31. 31.Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al. The” something something” video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision, 2017.
  32. 32.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, 2018.
  33. 33.Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. Learning latent dynamics for planning from pixels. In International Conference on Machine Learning, 2019.
  34. 34.Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. Dream to control: Learning behaviors by latent imagination. In International Conference on Learning Representations, 2020.
  35. 35.Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. Mastering atari with discrete world models. In International Conference on Learning Representations, 2021.
  36. 36.Hazan, E., Kakade, S., Singh, K., and Van Soest, A. Provably efficient maximum entropy exploration. In International Conference on Machine Learning, 2019.
  37. 37.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  38. 38.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021.
  39. 39.Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A. Darla: Improving zero-shot transfer in reinforcement learning. In International Conference on Machine Learning, 2017.
  40. 40.Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. Vime: Variational information maximizing exploration. In Advances in Neural Information Processing Systems, 2016.
  41. 41.Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. Reinforcement learning with unsupervised auxiliary tasks. In International Conference on Learning Representations, 2017.
  42. 42.James, S., Ma, Z., Arrojo, D. R., and Davison, A. J. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 2020.
  43. 43.Jang, Y., Kim, G., and Song, Y. Video prediction with appearance and motion conditions. In International Conference on Machine Learning, 2018.
  44. 44.Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. Model-based reinforcement learning for atari. In International Conference on Learning Representations, 2019.
  45. 45.Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K. Mt-opt: Continuous multi-task robotic reinforcement learning at scale. arXiv preprint arXiv:2104.08212, 2021.
  46. 46.Kalchbrenner, N., Oord, A., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K. Video pixel networks. In International Conference on Machine Learning, 2017.
  47. 47.Kim, S. W., Zhou, Y., Philion, J., Torralba, A., and Fidler, S. Learning to simulate dynamic environments with gamegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  48. 48.Kim, S. W., Philion, J., Torralba, A., and Fidler, S. Drivegan: Towards a controllable high-quality neural simulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  49. 49.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014.
  50. 50.Kwon, Y.-H. and Park, M.-G. Predicting future frames using retrospective cycle gan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  51. 51.Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A. Reinforcement learning with augmented data. In Advances in Neural Information Processing Systems, 2020a.
  52. 52.Laskin, M., Srinivas, A., and Abbeel, P. Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning, 2020b.
  53. 53.Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P. Urlb: Unsupervised reinforcement learning benchmark. In Advances in Neural Information Processing Systems Datasets and Benchmarks Track, 2021.
  54. 54.Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S. Stochastic adversarial video prediction. arXiv preprint arXiv:1804.01523, 2018.
  55. 55.Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R. Efficient exploration via state marginal matching. arXiv preprint arXiv:1906.05274, 2019.
  56. 56.Levine, S., Finn, C., Darrell, T., and Abbeel, P. End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 2016.
  57. 57.Liu, H. and Abbeel, P. Behavior from the void: Unsupervised active pre-training. arXiv preprint arXiv:2103.04551, 2021.
  58. 58.Liu, Y., Gupta, A., Abbeel, P., and Levine, S. Imitation from observation: Learning to imitate behaviors from raw video via context translation. In 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018.
  59. 59.Lotter, W., Kreiman, G., and Cox, D. Deep predictive coding networks for video prediction and unsupervised learning. In International Conference on Learning Representations, 2017.
  60. 60.Luc, P., Clark, A., Dieleman, S., Casas, D. d. L., Doron, Y., Cassirer, A., and Simonyan, K. Transformation-based adversarial video prediction on large-scale data. arXiv preprint arXiv:2003.04035, 2020.
  61. 61.Mazoure, B., Combes, R. T. d., Doan, T., Bachman, P., and Hjelm, R. D. Deep reinforcement and infomax learning. In Advances in Neural Information Processing Systems, 2020.
  62. 62.Michalski, V., Memisevic, R., and Konda, K. Modeling deep temporal dependencies with recurrent grammar cells. In Advances in Neural Information Processing Systems, 2014.
  63. 63.Mikolov, T., Chen, K., Corrado, G., and Dean, J. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  64. 64.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. nature, 2015.
  65. 65.Noroozi, M. and Favaro, P. Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision, 2016.
  66. 66.Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S. Action-conditional video prediction using deep networks in atari games. In Advances in Neural Information Processing Systems, 2015.
  67. 67.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. In Advances in Neural Information Processing Systems, 2018.
  68. 68.Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R. Count-based exploration with neural density models. In International Conference on Machine Learning, 2017.
  69. 69.Park, J., Seo, Y., Shin, J., Lee, H., Abbeel, P., and Lee, K. Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning. In International Conference on Learning Representations, 2022.
  70. 70.Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. Curiosity-driven exploration by self-supervised prediction. In International Conference on Machine Learning, 2017.
  71. 71.Pathak, D., Gandhi, D., and Gupta, A. Self-supervised exploration via disagreement. In International Conference on Machine Learning, 2019.
  72. 72.Peng, X. B., Kanazawa, A., Malik, J., Abbeel, P., and Levine, S. Sfv: Reinforcement learning of physical skills from videos. ACM Transactions On Graphics (TOG), 2018.
  73. 73.Pennington, J., Socher, R., and Manning, C. D. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014.
  74. 74.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. Technical report, 2018.
  75. 75.Ranzato, M., Szlam, A., Bruna, J., Mathieu, M., Collobert, R., and Chopra, S. Video (language) modeling: a baseline for generative models of natural videos. arXiv preprint arXiv:1412.6604, 2014.
  76. 76.Reed, S., Oord, A., Kalchbrenner, N., Colmenarejo, S. G., Wang, Z., Chen, Y., Belov, D., and Freitas, N. Parallel multiscale autoregressive density estimation. In International Conference on Machine Learning, 2017.
  77. 77.Rybkin, O., Zhu, C., Nagabandi, A., Daniilidis, K., Mordatch, I., and Levine, S. Model-based reinforcement learning via latent-space collocation. In International Conference on Machine Learning, 2021.
  78. 78.Schmeckpeper, K., Rybkin, O., Daniilidis, K., Levine, S., and Finn, C. Reinforcement learning with videos: Combining offline observations with interaction. In Conference on Robot Learning, 2020a.
  79. 79.Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C. Learning predictive models from observation and interaction. In European Conference on Computer Vision, 2020b.
  80. 80.Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2015.
  81. 81.Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P. Data-efficient reinforcement learning with self-predictive representations. In International Conference on Learning Representations, 2021a.
  82. 82.Schwarzer, M., Rajkumar, N., Noukhovitch, M., Anand, A., Charlin, L., Hjelm, D., Bachman, P., and Courville, A. Pretraining representations for data-efficient reinforcement learning. In Advances in Neural Information Processing Systems, 2021b.
  83. 83.Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D. Planning to explore via self-supervised world models. In International Conference on Machine Learning, 2020.
  84. 84.Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K. State entropy maximization with random encoders for efficient exploration. In International Conference on Machine Learning, 2021.
  85. 85.Seo, Y., Lee, K., Liu, F., James, S., and Abbeel, P. Autoregressive latent video prediction with high-fidelity image generator, 2022. URL https://openreview.net/forum?id=K-hiHQXEQog.
  86. 86.Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE international conference on robotics and automation (ICRA), 2018.
  87. 87.Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. Mastering the game of go without human knowledge. nature, 2017.
  88. 88.Srinivas, A., Laskin, M., and Abbeel, P. Curl: Contrastive unsupervised representations for reinforcement learning. In Internatial Conference on Machine Learning, 2020.
  89. 89.Srivastava, N., Mansimov, E., and Salakhudinov, R. Unsupervised learning of video representations using lstms. In International Conference on Machine Learning, 2015.
  90. 90.Stooke, A., Lee, K., Abbeel, P., and Laskin, M. Decoupling representation learning from reinforcement learning. In International Conference on Machine Learning, 2021.
  91. 91.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT Press, 2018.
  92. 92.Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P. # exploration: A study of count-based exploration for deep reinforcement learning. In Advances in Neural Information Processing Systems, 2017.
  93. 93.Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N. dm control: Software and tasks for continuous control. arXiv preprint arXiv:2006.12983, 2020.
  94. 94.Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2017.
  95. 95.Torabi, F., Warnell, G., and Stone, P. Generative adversarial imitation from observation. arXiv preprint arXiv:1807.06158, 2018.
  96. 96.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of machine learning research, 2008.
  97. 97.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.
  98. 98.Villegas, R., Pathak, A., Kannan, H., Erhan, D., Le, Q. V., and Lee, H. High fidelity video prediction with large stochastic recurrent neural networks. Advances in Neural Information Processing Systems, 2019.
  99. 99.Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 2019.
  100. 100.Vondrick, C., Pirsiavash, H., and Torralba, A. Generating videos with scene dynamics. In Advances in Neural Information Processing Systems, 2016.
  101. 101.Weissenborn, D., Tackstrom, O., and Uszkoreit, J. Scaling autoregressive video models. In International Conference on Learning Representations, 2020.
  102. 102.Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021.
  103. 103.Xiao, T., Radosavovic, I., Darrell, T., and Malik, J. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173, 2022.
  104. 104.Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021.
  105. 105.Yan, W., Okumura, R., James, S., and Abbeel, P. Patch-based object-centric transformers for efficient video generation. arXiv preprint arXiv:2206.04003, 2022.
  106. 106.Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in Neural Information Processing Systems, 2019.
  107. 107.Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. Mastering visual continuous control: Improved data-augmented reinforcement learning. arXiv preprint arXiv:2107.09645, 2021a.
  108. 108.Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. Reinforcement learning with prototypical representations. In International Conference on Machine Learning, 2021b.
  109. 109.Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R. Improving sample efficiency in model-free reinforcement learning from images. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021c.
  110. 110.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, 2020.
  111. 111.Yu, T., Lan, C., Zeng, W., Feng, M., Zhang, Z., and Chen, Z. Playvirtual: Augmenting cycle-consistent virtual trajectories for reinforcement learning. In Advances in Neural Information Processing Systems, 2021.
  112. 112.Yu, T., Zhang, Z., Lan, C., Chen, Z., and Lu, Y. Mask-based latent reconstruction for reinforcement learning. arXiv preprint arXiv:2201.12096, 2022.
  113. 113.Zakka, K., Zeng, A., Florence, P., Tompson, J., Bohg, J., and Dwibedi, D. Xirl: Cross-embodiment inverse reinforcement learning. In Conference on Robot Learning, 2022.
  114. 114.Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M. A framework for efficient robotic manipulation. arXiv preprint arXiv:2012.07975, 2020.
  115. 115.Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S. Learning invariant representations for reinforcement learning without reconstruction. In International Conference on Learning Representations, 2021.
  116. 116.Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S. Solar: Deep structured representations for model-based reinforcement learning. In International Conference on Machine Learning, 2019.
  117. 117.Ziebart, B. D. Modeling purposeful adaptive behavior with the principle of maximum causal entropy, 2010.

Citation

MLA
Seo, Y., et al. “Reinforcement Learning with Action-Free Pre-Training from Videos”. International Conference on Machine Learning, vol. 162, 2022, pp. 19561–79, https://proceedings.mlr.press/v162/seo22a.html.
APA
Seo, Y., Lee, K., James, S. L., & Abbeel, P. (2022). Reinforcement Learning with Action-Free Pre-Training from Videos. International Conference on Machine Learning, 162, 19561–19579. https://proceedings.mlr.press/v162/seo22a.html
Chicago
Seo, Y., K. Lee, S. L. James, and P. Abbeel. 2022. “Reinforcement Learning with Action-Free Pre-Training from Videos”. International Conference on Machine Learning 162: 19561–79. https://proceedings.mlr.press/v162/seo22a.html.
Harvard
Seo, Y. et al. (2022) “Reinforcement Learning with Action-Free Pre-Training from Videos”, International Conference on Machine Learning. PMLR, pp. 19561–19579. Available at: https://proceedings.mlr.press/v162/seo22a.html.
Vancouver
1. Seo Y, Lee K, James SL, Abbeel P (2022) Reinforcement Learning with Action-Free Pre-Training from Videos. In: International Conference on Machine Learning. PMLR, pp 19561–19579

BibTeX

@InProceedings{pmlr-v162-seo22a,
  title = 	 {Reinforcement Learning with Action-Free Pre-Training from Videos},
  author =       {Seo, Younggyo and Lee, Kimin and James, Stephen L and Abbeel, Pieter},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {19561--19579},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/seo22a/seo22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/seo22a.html},
  abstract = 	 {Recent unsupervised pre-training methods have shown to be effective on language and vision domains by learning useful representations for multiple downstream tasks. In this paper, we investigate if such unsupervised pre-training methods can also be effective for vision-based reinforcement learning (RL). To this end, we introduce a framework that learns representations useful for understanding the dynamics via generative pre-training on videos. Our framework consists of two phases: we pre-train an action-free latent video prediction model, and then utilize the pre-trained representations for efficiently learning action-conditional world models on unseen environments. To incorporate additional action inputs during fine-tuning, we introduce a new architecture that stacks an action-conditional latent prediction model on top of the pre-trained action-free prediction model. Moreover, for better exploration, we propose a video-based intrinsic bonus that leverages pre-trained representations. We demonstrate that our framework significantly improves both final performances and sample-efficiency of vision-based RL in a variety of manipulation and locomotion tasks. Code is available at \url{https://github.com/younggyoseo/apv}.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/