LIV: Language-Image Representations and Rewards for Robotic Control

Yecheng Jason MaVikash KumarAmy ZhangOsbert BastaniDinesh Jayaraman

article2023ICML291 citations

Presents a unified framework connecting dual reinforcement learning with contrastive learning to extract control-centric multimodal representations and zero-shot reward functions from passive annotated human videos for language-guided robotic manipulation.

Listen

Building general-purpose robots capable of operating in everyday human environments requires learning systems that can interpret natural language commands, visually evaluate their surroundings, and autonomously acquire new skills. Prior approaches typically rely on static image-text pre-training or visual-only models combined with language encoders. These methods struggle to simultaneously capture temporal task progression and maintain fine-grained language grounding, while also demanding substantial amounts of scarce, domain-specific robot demonstration data.

The article introduces and evaluates Language-Image Value learning (LIV), a unified framework for joint vision-language representation and reward learning designed specifically for robotic control. The primary objective is to demonstrate that a single objective can pre-train multi-modal representations on action-free, text-annotated human video datasets and effectively adapt them to robot manipulation tasks.

To evaluate this approach, the authors pre-trained LIV on the EpicKitchen dataset—comprising roughly 20 million frames from human activity videos—using a ResNet50 vision encoder initialized with standard contrastive model weights. They tested the resulting representations and reward models across two simulated benchmarks (MetaWorld and FrankaKitchen) and a physical tabletop robot setup executing multi-task fruit-sorting tasks with six-degree-of-freedom control at 15Hz. The evaluation compared LIV against established baseline representations, including standard image-text models and specialized visual representations, across language-conditioned imitation learning and model-based planning tasks.

The analysis yielded several key findings. First, pre-trained LIV achieved the highest task success rates across all environments for language-conditioned imitation learning, outperforming baseline representations by a substantial margin, particularly on the real-world robot platform. Second, fine-tuning pre-trained models with the LIV objective consistently boosted policy success rates by more than 40% across all environments, overcoming the limitations of standard fine-tuning methods that disrupted temporal coherence. Third, in few-shot settings on simulated tasks, LIV fine-tuning with only 10 demonstrations matched the performance of baseline models trained on 50 demonstrations, representing a performance improvement of over 200%. Finally, as a reward model for automated trajectory planning, LIV delivered state-of-the-art results, reaching a 55.2% success rate on MetaWorld and 20.0% on FrankaKitchen after fine-tuning.

These results indicate that combining temporal value learning with multi-modal contrastive alignment creates a latent space that simultaneously preserves smooth task progression and semantic meaning. Unlike standard fine-tuning objectives that risk distorting intermediate visual states, LIV maintains temporal consistency without requiring extensive hyperparameter tuning. This dual capability allows robotic systems to leverage large-scale, low-cost human video data for pre-training, significantly reducing the costly engineering hours traditionally required to collect robot teleoperation demonstrations.

Organizations developing vision-language robotic systems should consider adopting LIV-style objectives to pre-train control-centric backbones on passive human video data and fine-tune them on small sets of in-domain robotic demonstrations. When designing fine-tuning pipelines, teams should avoid static image-text objectives alone, as they compromise intermediate state representations. For deployment, organizations can pilot LIV both as a frozen feature extractor for behavior cloning and as a reward generator for autonomous reinforcement learning and trajectory optimization.

While the findings demonstrate strong generalization, the authors note certain limitations. Pre-trained LIV exhibits a domain gap between passive human video and robot environments, meaning zero-shot reward curves can be noisy before in-domain adaptation. Additionally, real-world evaluations were conducted in controlled tabletop setups with moderate dataset sizes. Stakeholders should maintain high confidence in LIV’s comparative superiority across benchmarks, while recognizing that scaling to unconstrained real-world environments with diverse camera configurations may still require domain-specific fine-tuning.

Cover for LIV: Language-Image Representations and Rewards for Robotic Control

Abstract

We present Language-Image Value learning (LIV), a unified objective for vision-language representation and reward learning from action-free videos with text annotations. Exploiting a novel connection between dual reinforcement learning and mutual information contrastive learning, the LIV objective trains a multi-modal representation that implicitly encodes a universal value function for tasks specified as language or image goals. We use LIV to pre-train the first control-centric vision-language representation from large human video datasets such as EpicKitchen. Given only a language or image goal, the pre-trained LIV model can assign dense rewards to each frame in videos of unseen robots or humans attempting that task in unseen environments. Further, when some target domain-specific data is available, the same objective can be used to fine-tune and improve LIV and even other pre-trained representations for robotic control and reward specification in that domain. In our experiments on several simulated and real-world robot environments, LIV models consistently outperform the best prior input state representations for imitation learning, as well as reward specification methods for policy synthesis. Our results validate the advantages of joint vision-language representation and reward learning within the unified, compact LIV framework. Project website: penn-pal-lab.github.io/LIV

Citation

MLA
Ma, Y. J., et al. “LIV: Language-Image Representations and Rewards for Robotic Control”. International Conference on Machine Learning, vol. 202, 2023, pp. 23301–20, https://proceedings.mlr.press/v202/ma23b.html.
APA
Ma, Y. J., Kumar, V., Zhang, A., Bastani, O., & Jayaraman, D. (2023). LIV: Language-Image Representations and Rewards for Robotic Control. International Conference on Machine Learning, 202, 23301–23320. https://proceedings.mlr.press/v202/ma23b.html
Chicago
Ma, Y. J., V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman. 2023. “LIV: Language-Image Representations and Rewards for Robotic Control”. International Conference on Machine Learning 202: 23301–20. https://proceedings.mlr.press/v202/ma23b.html.
Harvard
Ma, Y.J. et al. (2023) “LIV: Language-Image Representations and Rewards for Robotic Control”, International Conference on Machine Learning. PMLR, pp. 23301–23320. Available at: https://proceedings.mlr.press/v202/ma23b.html.
Vancouver
1. Ma YJ, Kumar V, Zhang A, Bastani O, Jayaraman D (2023) LIV: Language-Image Representations and Rewards for Robotic Control. In: International Conference on Machine Learning. PMLR, pp 23301–23320

BibTeX

@InProceedings{pmlr-v202-ma23b,
  title = 	 {{LIV}: Language-Image Representations and Rewards for Robotic Control},
  author =       {Ma, Yecheng Jason and Kumar, Vikash and Zhang, Amy and Bastani, Osbert and Jayaraman, Dinesh},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {23301--23320},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ma23b/ma23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ma23b.html},
  abstract = 	 {We present Language-Image Value learning (LIV), a unified objective for vision-language representation and reward learning from action-free videos with text annotations. Exploiting a novel connection between dual reinforcement learning and mutual information contrastive learning, the LIV objective trains a multi-modal representation that implicitly encodes a universal value function for tasks specified as language or image goals. We use LIV to pre-train the first control-centric vision-language representation from large human video datasets such as EpicKitchen. Given only a language or image goal, the pre-trained LIV model can assign dense rewards to each frame in videos of unseen robots or humans attempting that task in unseen environments. Further, when some target domain-specific data is available, the same objective can be used to fine-tune and improve LIV and even other pre-trained representations for robotic control and reward specification in that domain. In our experiments on several simulated and real-world robot environments, LIV models consistently outperform the best prior input state representations for imitation learning, as well as reward specification methods for policy synthesis. Our results validate the advantages of joint vision-language representation and reward learning within the unified, compact LIV framework.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/