PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Soroush NasirianyFei XiaWenhao YuTed XiaoJacky LiangIshita DasguptaAnnie XieDanny DriessAyzaan WahidZhuo Xu

article2024ICML209 citations

Proposes an iterative visual prompting framework that allows standard vision-language models to perform zero-shot robotic control and spatial reasoning by repeatedly annotating images with candidate action proposals and refining them via visual question answering.

Listen

Modern vision-language models excel at visual reasoning and dialogue, but their textual interface prevents them from directly outputting continuous coordinates, robot actions, or spatial trajectories. Deploying these models to control physical robots or resolve fine-grained spatial problems has traditionally required task-specific fine-tuning, domain demonstration data, or complex auxiliary modules. The article introduces and evaluates Prompting with Iterative Visual Optimization (PIVOT), a framework that bridges this gap by converting spatial reasoning and continuous robotic control into an iterative visual question answering process without modifying or retraining the underlying models.

To evaluate this framework, the article investigates the performance of state-of-the-art vision-language models—primarily GPT-4V and the Gemini model family—across zero-shot robotic navigation, tabletop manipulation on real mobile manipulators and Franka arms, simulated pick-and-place tasks, and visual grounding on the RefCOCO benchmark. The method operates by overlaying candidate actions (such as numbered arrows or spatial markers) onto an image observation, querying the vision-language model to rank the best options, fitting a refined probability distribution around those selections, and repeating the cycle. Offline ablations against human demonstration datasets (including the RT-X dataset) and online real-world trials were conducted to assess accuracy, parallel execution strategies, and prompt structures.

Key findings demonstrate that iterative visual prompting enables viable zero-shot low-level robotic control. In real-world mobile navigation trials, adding three refinement iterations and parallel querying increased navigation success rates from 25–75% up to 75–100%. In tabletop manipulation, the approach enabled the robot to achieve up to a 100% reach rate and a 67% grasp rate on target objects, while reducing the average number of action steps required. Offline benchmarks confirmed that visual prompt optimization substantially outperforms text-only directional choices, with approximately 10 visual samples per iteration providing the best trade-off between spatial coverage and visual clutter. Furthermore, performance scaled monotonically with model size across the Gemini family, and fine-tuning a smaller model specifically on visual action selection yielded higher directional similarity (increasing from 0.53 to 0.65 across iterations) than larger zero-shot baselines.

These results imply that web-scale foundation models can be directly translated into spatial controllers without expensive robot-specific pretraining datasets or intricate intermediate software. This significantly lowers the barrier to deploying flexible, generalizable robotic systems across varied environments. However, the article highlights critical limitations: current models struggle with genuine 3D depth perception, rotational control, visual occlusions during close-up physical interactions, and short-sighted decision-making in multi-step workflows. Organizations looking to build on this work should focus future efforts on integrating depth-aware visual representations, training models on embodied video interaction data, and implementing robust search strategies to mitigate occasional model hallucinations and misdirected visual references.

arXiv: 2402.07872
Cover for PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Abstract

Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding. This opens the door to richer interaction with the world, for example robotic control. However, VLMs produce only textual outputs, while robotic control and other spatial tasks require outputting continuous coordinates, actions, or trajectories. How can we enable VLMs to handle such settings without fine-tuning on task-specific data? In this paper, we propose a novel visual prompting approach for VLMs that we call Prompting with Iterative Visual Optimization (PIVOT), which casts tasks as iterative visual question answering. In each iteration, the image is annotated with a visual representation of proposals that the VLM can refer to (e.g., candidate robot actions, localizations, or trajectories). The VLM then selects the best ones for the task. These proposals are iteratively refined, allowing the VLM to eventually zero in on the best available answer. We investigate PIVOT on real-world robotic navigation, real-world manipulation from images, instruction following in simulation, and additional spatial inference tasks such as localization. We find, perhaps surprisingly, that our approach enables zero-shot control of robotic systems without any robot training data, navigation in a variety of environments, and other capabilities. Although current performance is far from perfect, our work highlights potentials and limitations of this new regime and shows a promising approach for Internet-Scale VLMs in robotic and spatial reasoning domains. Website and HuggingFace demo.

Citation

MLA
Nasiriany, S., et al. “PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs”. arXiv, 2024, http://arxiv.org/abs/2402.07872v1.
APA
Nasiriany, S., Xia, F., Yu, W., Xiao, T., Liang, J., Dasgupta, I., Xie, A., Driess, D., Wahid, A., Xu, Z., Vuong, Q., Zhang, T., Lee, T.-W. E., Lee, K.-H., Xu, P., Kirmani, S., Zhu, Y., Zeng, A., Hausman, K., … Ichter, B. (2024). PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs. arXiv. http://arxiv.org/abs/2402.07872v1
Chicago
Nasiriany, S., F. Xia, W. Yu, et al. 2024. “PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs”. arXiv. http://arxiv.org/abs/2402.07872v1.
Harvard
Nasiriany, S. et al. (2024) “PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.07872v1.
Vancouver
1. Nasiriany S, Xia F, Yu W, et al (2024) PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs. arXiv

BibTeX

@article{nasiriany2024pivot,
  title = {PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs},
  author = {Nasiriany, Soroush and Xia, Fei and Yu, Wenhao and Xiao, Ted and Liang, Jacky and Dasgupta, Ishita and Xie, Annie and Driess, Danny and Wahid, Ayzaan and Xu, Zhuo and Vuong, Quan and Zhang, Tingnan and Lee, Tsang-Wei Edward and Lee, Kuang-Huei and Xu, Peng and Kirmani, Sean and Zhu, Yuke and Zeng, Andy and Hausman, Karol and Heess, Nicolas and Finn, Chelsea and Levine, Sergey and Ichter, Brian},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.07872v1},
  eprint = {2402.07872}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/