PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Soroush NasirianyFei XiaWenhao YuTed XiaoJacky LiangIshita DasguptaAnnie XieDanny DriessAyzaan WahidZhuo Xu
Proposes an iterative visual prompting framework that allows standard vision-language models to perform zero-shot robotic control and spatial reasoning by repeatedly annotating images with candidate action proposals and refining them via visual question answering.
Modern vision-language models excel at visual reasoning and dialogue, but their textual interface prevents them from directly outputting continuous coordinates, robot actions, or spatial trajectories. Deploying these models to control physical robots or resolve fine-grained spatial problems has traditionally required task-specific fine-tuning, domain demonstration data, or complex auxiliary modules. The article introduces and evaluates Prompting with Iterative Visual Optimization (PIVOT), a framework that bridges this gap by converting spatial reasoning and continuous robotic control into an iterative visual question answering process without modifying or retraining the underlying models.
To evaluate this framework, the article investigates the performance of state-of-the-art vision-language models—primarily GPT-4V and the Gemini model family—across zero-shot robotic navigation, tabletop manipulation on real mobile manipulators and Franka arms, simulated pick-and-place tasks, and visual grounding on the RefCOCO benchmark. The method operates by overlaying candidate actions (such as numbered arrows or spatial markers) onto an image observation, querying the vision-language model to rank the best options, fitting a refined probability distribution around those selections, and repeating the cycle. Offline ablations against human demonstration datasets (including the RT-X dataset) and online real-world trials were conducted to assess accuracy, parallel execution strategies, and prompt structures.
Key findings demonstrate that iterative visual prompting enables viable zero-shot low-level robotic control. In real-world mobile navigation trials, adding three refinement iterations and parallel querying increased navigation success rates from 25–75% up to 75–100%. In tabletop manipulation, the approach enabled the robot to achieve up to a 100% reach rate and a 67% grasp rate on target objects, while reducing the average number of action steps required. Offline benchmarks confirmed that visual prompt optimization substantially outperforms text-only directional choices, with approximately 10 visual samples per iteration providing the best trade-off between spatial coverage and visual clutter. Furthermore, performance scaled monotonically with model size across the Gemini family, and fine-tuning a smaller model specifically on visual action selection yielded higher directional similarity (increasing from 0.53 to 0.65 across iterations) than larger zero-shot baselines.
These results imply that web-scale foundation models can be directly translated into spatial controllers without expensive robot-specific pretraining datasets or intricate intermediate software. This significantly lowers the barrier to deploying flexible, generalizable robotic systems across varied environments. However, the article highlights critical limitations: current models struggle with genuine 3D depth perception, rotational control, visual occlusions during close-up physical interactions, and short-sighted decision-making in multi-step workflows. Organizations looking to build on this work should focus future efforts on integrating depth-aware visual representations, training models on embodied video interaction data, and implementing robust search strategies to mitigate occasional model hallucinations and misdirected visual references.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). RT-2 established the paradigm of translating web-scale vision-language models into direct robotic action tokens, providing the foundational context and baseline that PIVOT aims to control without specialized fine-tuning.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). This paper demonstrates how zero-shot actionable knowledge can be extracted from large foundation models for embodied planning, motivating PIVOT's extension into continuous spatial and visual action control.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Inner Monologue introduces closed-loop iterative feedback prompting for robotic planning, which PIVOT adapts into an iterative visual prompting and optimization mechanism.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). Kosmos-2 demonstrates visual grounding and spatial coordinate referencing in multimodal language models, laying essential groundwork for PIVOT's visual prompt overlays and spatial grounding tasks.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). EmbodiedGPT explores connecting high-level vision-language reasoning to low-level motor execution, highlighting the direct-control bottlenecks that PIVOT addresses via iterative visual question answering.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Code as Policies establishes how foundation models can generate parameterized robot control interfaces, offering key conceptual background for eliciting embodied control without retraining.
- Paper: Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation, Yicong Hong et al. (2022). This work demonstrates bridging discrete decision-making and continuous spatial navigation via visual candidate waypoints, directly prefiguring PIVOT's candidate action sampling and ranking strategy.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). RoboSpatial directly addresses the spatial reasoning and depth perception limitations of VLMs highlighted in PIVOT by curating multi-view spatial QA datasets for 2D and 3D models.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). CoT-VLA extends visual prompting concepts by having the vision-language-action model explicitly generate visual subgoals to guide continuous robot action sequences.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). CoA-VLA builds on spatial grounding for robotic control by introducing structured visual-textual chains of affordance such as grasp and placement points.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). UP-VLA advances beyond static visual prompting by integrating high-level semantic understanding with predictive future visual dynamics for fine-grained manipulation.
- Paper: π0.5: a Vision-Language-Action Model with Open-World Generalization, Physical Intelligence et al. (2025). π0.5 develops generalist vision-language-action capabilities using continuous action experts and flow matching, extending the zero-shot open-world control explored in PIVOT.
- Paper: AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation, Jiafei Duan et al. (2025). AHA addresses the multi-step robustness and hallucination challenges identified in PIVOT by enabling VLMs to reason over and diagnose robotic manipulation failures.
- Paper: Visual Planning: Let's Think Only with Images, Yi Xu et al. (2026). Visual Planning takes the visual-only planning paradigm further by training models to plan spatial trajectories purely through image transitions without text reasoning.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This survey provides a comprehensive synthesis of vision-language-action architectures and control strategies, contextualizing prompting frameworks like PIVOT within the broader VLA landscape.
