VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Yecheng Jason MaShagun SodhaniDinesh JayaramanOsbert BastaniVikash KumarAmy Zhang

article2023ICLR526 citationsNotable-Top-25% (Spotlight)

Introduces a self-supervised pre-training method that learns from unlabeled human videos to generate both visual representations and dense reward functions, enabling real-world robots to learn novel manipulation skills from as few as twenty trajectories without fine-tuning.

Listen

Scaling general-purpose robotic manipulation has long been hindered by the difficulty and expense of collecting large-scale, in-domain robot data, as well as the substantial engineering effort required to hand-craft dense reward functions for every new task. While pre-training visual models on broad, out-of-domain human video datasets has helped robots interpret scenes, using these passive videos to specify and guide task progress automatically without action labels or manual reward tuning has remained an open challenge.

The article aims to develop and validate a self-supervised visual pre-training method that enables a single frozen model to serve simultaneously as an effective visual representation and a universal, zero-shot dense reward function for unseen robotic tasks specified solely by goal images.

To achieve this, the article introduces Value-Implicit Pre-training (VIP), which reframes learning from human videos into an offline, goal-conditioned value function optimization problem using mathematical duality. Because the resulting dual formulation requires no action labels, VIP is trained completely self-supervised on roughly 4.3 million frames from the large-scale Ego4D human video dataset using a standard ResNet-50 visual backbone. The model was evaluated across 36 simulated robotic manipulation tasks in FrankaKitchen across three camera views and difficulty levels, as well as on a physical 7-degree-of-freedom Franka robot performing four tabletop manipulation tasks using few-shot offline reinforcement learning.

Across evaluations, VIP demonstrated several critical advantages. In visual trajectory optimization, VIP solved approximately 30% of simulated tasks out-of-the-box and scaled to roughly 44% with increased computational budget, whereas baseline representations deteriorated due to local minima and reward exploitation. In online visual reinforcement learning, VIP doubled as an effective encoder and reward generator, achieving a 40% aggregate success rate where sparse-reward baselines failed completely. On the physical robot, VIP powered reward-weighted offline learning with as few as 20 demonstration trajectories, achieving 60% to 100% success across challenging articulated, deformable, and pick-and-place tasks where standard behavioral cloning and competing baselines struggled or failed entirely.

These results demonstrate that a representation pre-trained purely on passive human video can acquire an implicit, temporally smooth metric of task progress that transfers directly to physical robots without fine-tuning. This drastically reduces the human labor and cost associated with manual reward engineering, lowers the sample size required for real-world robot skill acquisition, and enables practical offline reinforcement learning in data-scarce settings.

Decision-makers and engineering teams should consider adopting VIP-based pre-trained models to streamline real-world robotic deployments and replace fragile manual reward designs. Recommended next steps include establishing pilot implementations that leverage few-shot offline reinforcement learning for complex multi-stage tasks and exploring fine-tuning strategies to push absolute task success rates higher.

Users should note certain limitations: VIP currently relies on static goal images, making it less suitable for dynamic or sequential instruction-following tasks without extensions, and it models distance symmetrically, which assumes environment reversibility. Nonetheless, the consistent performance across extensive simulated benchmarks and real-world robot tasks provides strong confidence in the stability and transferability of VIP as a foundation for scalable visual robotic control.

Cover for VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Abstract

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations. Given the inherent cost and scarcity of in-domain, task-specific robot data, learning from large, diverse, offline human videos has emerged as a promising path towards acquiring a generally useful visual representation for control; however, how these human videos can be used for general-purpose reward learning remains an open question. We introduce V\textbf{V}alue-I\textbf{I}mplicit P\textbf{P}re-training (VIP), a self-supervised pre-trained visual representation capable of generating dense and smooth reward functions for unseen robotic tasks. VIP casts representation learning from human videos as an offline goal-conditioned reinforcement learning problem and derives a self-supervised dual goal-conditioned value-function objective that does not depend on actions, enabling pre-training on unlabeled human videos. Theoretically, VIP can be understood as a novel implicit time contrastive objective that generates a temporally smooth embedding, enabling the value function to be implicitly defined via the embedding distance, which can then be used to construct the reward for any goal-image specified downstream task. Trained on large-scale Ego4D human videos and without any fine-tuning on in-domain, task-specific data, VIP's frozen representation can provide dense visual reward for an extensive set of simulated and real-robot\textbf{real-robot} tasks, enabling diverse reward-based visual control methods and significantly outperforming all prior pre-trained representations. Notably, VIP can enable simple, few-shot\textbf{few-shot} offline RL on a suite of real-world robot tasks with as few as 20 trajectories.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Setting and Background
  • 4 Value-Implicit Pre-Training
  • 4.1 Foundation: Self-Supervised Value Learning from Human Videos
  • 4.2 Analysis: Implicit Time Contrastive Learning
  • 4.3 Algorithm: Value-Implicit Pre-Training (VIP)
  • 5 Experiments
  • 5.1 Trajectory Optimization & Online Reinforcement Learning
  • 5.2 Real-World Few-Shot Offline Reinforcement Learning
  • 5.3 Qualitative Analysis
  • 6 Conclusion
  • References
  • A Additional Background
  • A.1 Goal-Conditioned Reinforcement Learning
  • A.2 InfoNCE & Time Contrastive Learning.
  • B Extended Related Work
  • C Technical Derivations and Proofs
  • C.1 Proof of Proposition
  • C.2 VIP Implicit Time Contrast Learning Derivation
  • C.3 VIP Implicit Repulsion
  • D VIP Training Details
  • D.1 Dataset Processing and Sampling
  • D.2 VIP Hyperparameters
  • D.3 VIP Pytorch Pseudocode
  • E Simulation Experiment Details.
  • E.1 FrankaKitchen Task Descriptions
  • E.2 In-Domain Representation Probing
  • E.3 Trajectory Optimization
  • E.3.1 Robot and Object Pose Error Analysis
  • E.4 Reinforcement Learning
  • F Real-World Robot Experiment Details
  • F.1 Task Descriptions
  • F.2 Training and Evaluation Details
  • F.3 Additional Analysis & Context
  • F.4 Qualitative Analysis
  • G Additional Results
  • G.1 Comparison to MAE and MoCo trained on Ego4D
  • G.2 Value-Based Pre-Training Ablation: Least-Square Temporal-Difference
  • G.3 Visual Imitation Learning
  • G.4 Embedding and True Rewards Correlation
  • G.5 Embedding Distance Curves
  • G.6 Embedding Distance Curve Bumps
  • G.7 Embedding Reward Histograms (Real-Robot Dataset)
  • G.8 Embedding Reward Histograms (Ego4D)
  • H Limitations and Future Work

Citation

MLA
Ma, Y. J., et al. “VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training”. arXiv, 2022, http://arxiv.org/abs/2210.00030v2.
APA
Ma, Y. J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., & Zhang, A. (2022). VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training. arXiv. http://arxiv.org/abs/2210.00030v2
Chicago
Ma, Y. J., S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang. 2022. “VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training”. arXiv. http://arxiv.org/abs/2210.00030v2.
Harvard
Ma, Y.J. et al. (2022) “VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.00030v2.
Vancouver
1. Ma YJ, Sodhani S, Jayaraman D, Bastani O, Kumar V, Zhang A (2022) VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training. arXiv

BibTeX

@article{ma2022vip,
  title = {VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training},
  author = {Ma, Yecheng Jason and Sodhani, Shagun and Jayaraman, Dinesh and Bastani, Osbert and Kumar, Vikash and Zhang, Amy},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.00030v2},
  eprint = {2210.00030}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors