VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
Yecheng Jason MaShagun SodhaniDinesh JayaramanOsbert BastaniVikash KumarAmy Zhang
Introduces a self-supervised pre-training method that learns from unlabeled human videos to generate both visual representations and dense reward functions, enabling real-world robots to learn novel manipulation skills from as few as twenty trajectories without fine-tuning.
Scaling general-purpose robotic manipulation has long been hindered by the difficulty and expense of collecting large-scale, in-domain robot data, as well as the substantial engineering effort required to hand-craft dense reward functions for every new task. While pre-training visual models on broad, out-of-domain human video datasets has helped robots interpret scenes, using these passive videos to specify and guide task progress automatically without action labels or manual reward tuning has remained an open challenge.
The article aims to develop and validate a self-supervised visual pre-training method that enables a single frozen model to serve simultaneously as an effective visual representation and a universal, zero-shot dense reward function for unseen robotic tasks specified solely by goal images.
To achieve this, the article introduces Value-Implicit Pre-training (VIP), which reframes learning from human videos into an offline, goal-conditioned value function optimization problem using mathematical duality. Because the resulting dual formulation requires no action labels, VIP is trained completely self-supervised on roughly 4.3 million frames from the large-scale Ego4D human video dataset using a standard ResNet-50 visual backbone. The model was evaluated across 36 simulated robotic manipulation tasks in FrankaKitchen across three camera views and difficulty levels, as well as on a physical 7-degree-of-freedom Franka robot performing four tabletop manipulation tasks using few-shot offline reinforcement learning.
Across evaluations, VIP demonstrated several critical advantages. In visual trajectory optimization, VIP solved approximately 30% of simulated tasks out-of-the-box and scaled to roughly 44% with increased computational budget, whereas baseline representations deteriorated due to local minima and reward exploitation. In online visual reinforcement learning, VIP doubled as an effective encoder and reward generator, achieving a 40% aggregate success rate where sparse-reward baselines failed completely. On the physical robot, VIP powered reward-weighted offline learning with as few as 20 demonstration trajectories, achieving 60% to 100% success across challenging articulated, deformable, and pick-and-place tasks where standard behavioral cloning and competing baselines struggled or failed entirely.
These results demonstrate that a representation pre-trained purely on passive human video can acquire an implicit, temporally smooth metric of task progress that transfers directly to physical robots without fine-tuning. This drastically reduces the human labor and cost associated with manual reward engineering, lowers the sample size required for real-world robot skill acquisition, and enables practical offline reinforcement learning in data-scarce settings.
Decision-makers and engineering teams should consider adopting VIP-based pre-trained models to streamline real-world robotic deployments and replace fragile manual reward designs. Recommended next steps include establishing pilot implementations that leverage few-shot offline reinforcement learning for complex multi-stage tasks and exploring fine-tuning strategies to push absolute task success rates higher.
Users should note certain limitations: VIP currently relies on static goal images, making it less suitable for dynamic or sequential instruction-following tasks without extensions, and it models distance symmetrically, which assumes environment reversibility. Nonetheless, the consistent performance across extensive simulated benchmarks and real-world robot tasks provides strong confidence in the stability and transferability of VIP as a foundation for scalable visual robotic control.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). Provides foundational insights into self-supervised contrastive representation learning directly paired with visual reinforcement learning objectives.
- Paper: Offline Reinforcement Learning with Implicit Q-Learning, Ilya Kostrikov et al. (2021). Introduces core techniques for offline value function estimation without out-of-distribution action queries, underpinning VIP's offline value formulation.
- Paper: Hindsight Experience Replay, Marcin Andrychowicz et al. (2017). Establishes the goal-conditioned reinforcement learning framework that VIP builds upon to formulate self-supervised value-implicit reward models.
- Paper: Generative Adversarial Imitation Learning, Jonathan Ho et al. (2016). Lays foundational mathematical dual formulations connecting imitation learning, state distributions, and implicit reward optimization.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). Defines the standard offline reinforcement learning benchmarks and data-driven evaluation methodologies used to assess visual policy and reward transfer.
- Paper: Emerging Properties in Self-Supervised Vision Transformers, Mathilde Caron et al. (2021). Demonstrates how self-supervised vision models learn rich, transferable visual feature representations that serve as backbones for downstream robotic control.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). Scales generalist robot representations by integrating pre-trained visual backbones into open-source vision-language-action foundation models.
- Paper: π0: A Vision-Language-Action Flow Model for General Robot Control, Kevin Black et al. (2024). Advances large-scale pre-trained robotic control by combining visual representation modeling with continuous flow matching for diverse multi-task policies.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Extends visual goal-directed representations by generating intermediate visual subgoals to guide downstream manipulation actions.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). Explores how large-scale egocentric video pre-training can be directly mapped to step-by-step embodied reasoning and robotic control.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). Enhances pre-trained vision-action architectures by injecting visual and textual affordance reasoning into downstream robot policy execution.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). Surveys the broader evolution and architectural paradigms of vision-language-action foundation models built upon pre-trained visual representations.
