LIV: Language-Image Representations and Rewards for Robotic Control
Yecheng Jason MaVikash KumarAmy ZhangOsbert BastaniDinesh Jayaraman
Presents a unified framework connecting dual reinforcement learning with contrastive learning to extract control-centric multimodal representations and zero-shot reward functions from passive annotated human videos for language-guided robotic manipulation.
Building general-purpose robots capable of operating in everyday human environments requires learning systems that can interpret natural language commands, visually evaluate their surroundings, and autonomously acquire new skills. Prior approaches typically rely on static image-text pre-training or visual-only models combined with language encoders. These methods struggle to simultaneously capture temporal task progression and maintain fine-grained language grounding, while also demanding substantial amounts of scarce, domain-specific robot demonstration data.
The article introduces and evaluates Language-Image Value learning (LIV), a unified framework for joint vision-language representation and reward learning designed specifically for robotic control. The primary objective is to demonstrate that a single objective can pre-train multi-modal representations on action-free, text-annotated human video datasets and effectively adapt them to robot manipulation tasks.
To evaluate this approach, the authors pre-trained LIV on the EpicKitchen dataset—comprising roughly 20 million frames from human activity videos—using a ResNet50 vision encoder initialized with standard contrastive model weights. They tested the resulting representations and reward models across two simulated benchmarks (MetaWorld and FrankaKitchen) and a physical tabletop robot setup executing multi-task fruit-sorting tasks with six-degree-of-freedom control at 15Hz. The evaluation compared LIV against established baseline representations, including standard image-text models and specialized visual representations, across language-conditioned imitation learning and model-based planning tasks.
The analysis yielded several key findings. First, pre-trained LIV achieved the highest task success rates across all environments for language-conditioned imitation learning, outperforming baseline representations by a substantial margin, particularly on the real-world robot platform. Second, fine-tuning pre-trained models with the LIV objective consistently boosted policy success rates by more than 40% across all environments, overcoming the limitations of standard fine-tuning methods that disrupted temporal coherence. Third, in few-shot settings on simulated tasks, LIV fine-tuning with only 10 demonstrations matched the performance of baseline models trained on 50 demonstrations, representing a performance improvement of over 200%. Finally, as a reward model for automated trajectory planning, LIV delivered state-of-the-art results, reaching a 55.2% success rate on MetaWorld and 20.0% on FrankaKitchen after fine-tuning.
These results indicate that combining temporal value learning with multi-modal contrastive alignment creates a latent space that simultaneously preserves smooth task progression and semantic meaning. Unlike standard fine-tuning objectives that risk distorting intermediate visual states, LIV maintains temporal consistency without requiring extensive hyperparameter tuning. This dual capability allows robotic systems to leverage large-scale, low-cost human video data for pre-training, significantly reducing the costly engineering hours traditionally required to collect robot teleoperation demonstrations.
Organizations developing vision-language robotic systems should consider adopting LIV-style objectives to pre-train control-centric backbones on passive human video data and fine-tune them on small sets of in-domain robotic demonstrations. When designing fine-tuning pipelines, teams should avoid static image-text objectives alone, as they compromise intermediate state representations. For deployment, organizations can pilot LIV both as a frozen feature extractor for behavior cloning and as a reward generator for autonomous reinforcement learning and trajectory optimization.
While the findings demonstrate strong generalization, the authors note certain limitations. Pre-trained LIV exhibits a domain gap between passive human video and robot environments, meaning zero-shot reward curves can be noisy before in-domain adaptation. Additionally, real-world evaluations were conducted in controlled tabletop setups with moderate dataset sizes. Stakeholders should maintain high confidence in LIV’s comparative superiority across benchmarks, while recognizing that scaling to unconstrained real-world environments with diverse camera configurations may still require domain-specific fine-tuning.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). VIP introduces the foundational mathematical framework of dual reinforcement learning for action-free visual representation and implicit value estimation that LIV directly builds upon and extends to vision-language goals.
- Paper: Contrastive Learning as Goal-Conditioned Reinforcement Learning, Benjamin Eysenbach et al. (2022). This work establishes the core equivalence between contrastive representation learning and goal-conditioned reinforcement learning value functions leveraged by LIV.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. CLIP provides the seminal contrastive language-image pre-training objective that underpins multimodal representation alignment in LIV.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). CURL establishes the utility of contrastive unsupervised representation learning for sample-efficient visual control and reinforcement learning.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). OpenVLA extends vision-language representations directly into end-to-end robotic action generation and generalist robot policies.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). UP-VLA builds on multimodal representations for robotic control by integrating high-level semantic understanding with low-level visual future prediction.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). CoT-VLA advances goal-directed visual-linguistic robot manipulation by generating explicit visual subgoals before action execution.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA enhances multimodal policy learning by incorporating structured visual-textual affordance reasoning into vision-language-action architectures.
- Paper: Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations, Yucheng Hu et al. (2025). Video Prediction Policy utilizes predictive visual representations from text-conditioned video models to guide multi-step robotic control.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). This position paper generalizes video-based state-action and goal representations into a unified interface for physical decision-making and robotic planning.
