CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Shresth GroverP. PathakAkash KumarVibhav VineetY. S. Rawat
Presents the CoSPlan benchmark to evaluate visual error detection and corrective reasoning in vision-language models, introducing a training-free incremental scene graph method that systematically improves multi-step visual planning across spatial tasks.
Artificial intelligence systems that integrate vision and text, known as vision-language models, have shown substantial promise in planning and executing tasks described in text. However, their ability to reason through step-by-step physical actions in real-world visual environments remains poorly understood. Real-world tasks, such as robotic manipulation or autonomous navigation, rarely follow perfect execution paths; they frequently involve mistakes, suboptimal moves, or constraint violations that require real-time course correction. Current evaluations often overlook these messy, sequential visual conditions, leaving decision-makers with an incomplete picture of model readiness for practical deployment.
The article introduces a benchmark called Corrective Sequence Planning to evaluate how well leading multimodal models can identify errors and plan corrective actions across evolving visual scenes. Specifically, the article tests models on their ability to detect an intentional mistake within a sequence of initial actions and complete the remaining visual steps required to reach a target goal state.
To conduct this evaluation, the researchers designed four distinct planning domains: synthetic maze navigation, block rearrangement, image patch reconstruction, and real-world household object reorganization. Across these tasks, the benchmark presents models with an initial scene, a target goal, and an initial action history containing an error. Models were evaluated using multiple-choice questions designed to prevent superficial shortcuts, such as merely guessing the option that visually resembles the target state without fixing the underlying mistake. The evaluation tested a broad suite of leading open-source and proprietary models using standard prompting, sequential step-by-step prompting (Chain-of-Thought), and structured object-relationship representations (Scene Graphs). To address identified failures, the authors also developed a training-free technique, Scene Graph Incremental updates, which converts visual scenes into text-based graphs and updates them step by step.
The findings show that current models struggle severely with visual sequence planning in the presence of errors. When presented with standard inputs, most open-source models performed near or below random guessing (around 20% accuracy) and exhibited strong blind option biases or attempts to bypass required error corrections. While proprietary frontier models performed better, achieving accuracies between 45% and 70% with structured representations, all models degraded sharply when errors involved plausible objects within the scene compared to clean, error-free settings. Furthermore, models performed far better when identical problems were presented purely in text (often exceeding 80% accuracy) rather than with visual inputs, revealing a fundamental inability to mentally simulate intermediate visual states. The authors' proposed technique, Scene Graph Incremental updates, mitigated this gap by breaking visual transitions into iterative text-based graph updates, improving average task completion accuracy by about 4.4% and error detection by up to 13% across tested models.
These results have direct operational implications for deploying automated visual agents in physical environments. Deploying current vision-language models directly into physical automation, such as warehouse robotics or vehicle navigation, introduces high operational and safety risks because the models cannot reliably recover from execution failures. The findings demonstrate that high benchmark scores in text-only planning create a false sense of security regarding a model's true physical reasoning capabilities.
Organizations considering vision-language models for robotic or sequential physical tasks should exercise caution and avoid unmonitored deployments. Teams should implement structured, step-by-step intermediate state tracking, such as incremental scene graph updates, which provide substantial performance gains over static representations at a fraction of the computational cost of generating synthetic intermediate images. Before strong deployment decisions are made, further research is required to evaluate models in dynamic, interactive video environments and in multi-error scenarios, as the primary analysis relied on static image pairs with a single isolated mistake.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). Provides the foundational iterative message passing framework for generating structured scene graphs from images, which underpins the Scene Graph Incremental update representation used in CoSPlan.
- Paper: Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task, Stan Weixian Lei et al. (2023). Introduces the concept of using textual scene graphs as symbolic prompts for visual reasoning tasks, directly motivating CoSPlan's training-free visual-to-textual graph translation method.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Establishes zero-shot step-by-step planning and action correction for embodied agents using language models, forming the planning foundation that CoSPlan extends to visual decision-making.
- Paper: PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs, Soroush Nasiriany et al. (2024). Demonstrates iterative visual prompting mechanisms for spatial decision-making in vision-language models, which informs corrective planning frameworks.
- Paper: Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models, Lei Wang et al. (2023). Formulates the Plan-and-Solve prompting strategy for explicit subtask decomposition and execution that CoSPlan adapts to visual sequential planning.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). Analyzes the failure modes of multimodal chain-of-thought in multi-step visual reasoning, framing the core motivation for CoSPlan's benchmark design.
- Paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning, Justin Johnson et al. (2016). Introduces diagnostic benchmark construction using explicit scene graph representations and compositional reasoning tasks to prevent models from exploiting visual shortcuts.
- Paper: Visual Planning: Let's Think Only with Images, Yi Xu et al. (2026). Explores the alternative paradigm of planning purely in the visual space without textual intermediate representations, contrasting with CoSPlan's textual scene graph approach.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Extends visual multi-step planning from abstract visual puzzles to continuous robotic action generation via visual chain-of-thought subgoals.
- Paper: SG2Loc: Sequential Visual Localization on 3D Scene Graphs, Nicole Damblon et al. (2026). Applies sequential reasoning over 3D scene graphs to continuous camera localization, extending scene-graph-based spatial tracking.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). Investigates whether generative models possess the underlying spatial intelligence required for multi-step visual scene planning and rearrangement.
