CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

Shresth GroverP. PathakAkash KumarVibhav VineetY. S. Rawat

article2025arXiv1 citations

Presents the CoSPlan benchmark to evaluate visual error detection and corrective reasoning in vision-language models, introducing a training-free incremental scene graph method that systematically improves multi-step visual planning across spatial tasks.

Listen

Artificial intelligence systems that integrate vision and text, known as vision-language models, have shown substantial promise in planning and executing tasks described in text. However, their ability to reason through step-by-step physical actions in real-world visual environments remains poorly understood. Real-world tasks, such as robotic manipulation or autonomous navigation, rarely follow perfect execution paths; they frequently involve mistakes, suboptimal moves, or constraint violations that require real-time course correction. Current evaluations often overlook these messy, sequential visual conditions, leaving decision-makers with an incomplete picture of model readiness for practical deployment.

The article introduces a benchmark called Corrective Sequence Planning to evaluate how well leading multimodal models can identify errors and plan corrective actions across evolving visual scenes. Specifically, the article tests models on their ability to detect an intentional mistake within a sequence of initial actions and complete the remaining visual steps required to reach a target goal state.

To conduct this evaluation, the researchers designed four distinct planning domains: synthetic maze navigation, block rearrangement, image patch reconstruction, and real-world household object reorganization. Across these tasks, the benchmark presents models with an initial scene, a target goal, and an initial action history containing an error. Models were evaluated using multiple-choice questions designed to prevent superficial shortcuts, such as merely guessing the option that visually resembles the target state without fixing the underlying mistake. The evaluation tested a broad suite of leading open-source and proprietary models using standard prompting, sequential step-by-step prompting (Chain-of-Thought), and structured object-relationship representations (Scene Graphs). To address identified failures, the authors also developed a training-free technique, Scene Graph Incremental updates, which converts visual scenes into text-based graphs and updates them step by step.

The findings show that current models struggle severely with visual sequence planning in the presence of errors. When presented with standard inputs, most open-source models performed near or below random guessing (around 20% accuracy) and exhibited strong blind option biases or attempts to bypass required error corrections. While proprietary frontier models performed better, achieving accuracies between 45% and 70% with structured representations, all models degraded sharply when errors involved plausible objects within the scene compared to clean, error-free settings. Furthermore, models performed far better when identical problems were presented purely in text (often exceeding 80% accuracy) rather than with visual inputs, revealing a fundamental inability to mentally simulate intermediate visual states. The authors' proposed technique, Scene Graph Incremental updates, mitigated this gap by breaking visual transitions into iterative text-based graph updates, improving average task completion accuracy by about 4.4% and error detection by up to 13% across tested models.

These results have direct operational implications for deploying automated visual agents in physical environments. Deploying current vision-language models directly into physical automation, such as warehouse robotics or vehicle navigation, introduces high operational and safety risks because the models cannot reliably recover from execution failures. The findings demonstrate that high benchmark scores in text-only planning create a false sense of security regarding a model's true physical reasoning capabilities.

Organizations considering vision-language models for robotic or sequential physical tasks should exercise caution and avoid unmonitored deployments. Teams should implement structured, step-by-step intermediate state tracking, such as incremental scene graph updates, which provide substantial performance gains over static representations at a fraction of the computational cost of generating synthetic intermediate images. Before strong deployment decisions are made, further research is required to evaluate models in dynamic, interactive video environments and in multi-error scenarios, as the primary analysis relied on static image pairs with a single isolated mistake.

Cover for CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

Abstract

Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexplored. Addressing this gap, we introduce Corrective Sequence Planning (CoSPlan) benchmark, where VLMs must plan a sequence of visual actions from an initial scene to a target scene. CoSPlan evaluates models on their ability to imagine and execute a coherent set of visual steps required to reach the goal (Step Completion). To prevent any shortcuts that simply describe the final scene, we introduce an erroneous action in decision-making, which must be detected (Error Detection) and corrected to reach the goal, enabling a deeper understanding of the task. CoSPlan spans across 4 tasks: maze navigation, block re-arrangement, image reconstruction, and object re-organization. Despite using advanced reasoning strategies such as Chain-of-Thought and Scene Graphs, VLMs struggle on CoSPlan, while still showing promising performance in the text domain. Addressing this, we propose Scene Graph Incremental updates (SGI), a novel training-free method to transform images into `textual' scene graphs, enabling step-by-step reasoning through iterative scene graph refinement. SGI yields an average of ~4.4% improvement on CoSPlan w/ generalization on PlanBench and VQA. Link for solving puzzles on the project page.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 CoSPlan Benchmark
  • 3.1 Benchmark Datasets
  • 3.2 Erroneous Step 𝒜ℰ{\mathcal{A}_{\mathcal{E}}}
  • 3.3 Safeguard against Shortcuts and Cheating
  • 3.4 Evaluation
  • 3.5 Models & Techniques
  • 3.6 Results & Benchmark Analysis
  • 4 Scene Graph Incremental update (SGI)
  • 4.1 Algorithm
  • 4.2 Results
  • 5 Conclusion and Future Work
  • 6 Limitations
  • References
  • 7 CoSPlan Design Choices
  • 7.1 Error Correction Design
  • 7.2 Sequence Completion Design
  • 7.3 One error Design
  • 7.4 Scene Graph Design (SG)
  • 7.5 Difference between SGI vs SG & CoT
  • 7.6 Chain of Though prompt Details
  • 7.7 Annotation Generation
  • 8 Scene Graph Incremental update (SGI) Details
  • 9 Contenders for solving CoSPlan
  • 9.1 LLaMA-3
  • 9.2 GPT-4o & GPT-5.1
  • 9.3 CoG-VLM
  • 9.4 InternVL2-26B
  • 9.5 Qwen2-VL
  • 9.6 Janus
  • 9.7 Human baseline
  • 9.8 Experimental Setup
  • 10 Results
  • 10.1 External Dataset Description
  • 10.2 Experiments
  • 10.3 Multi Error Analysis
  • 11 Implementation Details
  • 12 Ethical Statement
  • 13 Future work

Knowls

  1. Knowl 1 — CoSPlan Benchmark Formulation and Task Suite

    definition

    Corrective Sequence Planning (CoSPlan) is a multimodal benchmark designed to evaluate Vision-Language Models (VLMs) on multi-step visual planning under non-ideal, error-prone execution traces. Given an initial visual state I0\mathcal{I}_0, a target goal state Ig\mathcal{I}_g, and a sequence of kk partially executed actions A1,…,AE,…,k\mathcal{A}_{1, \dots, \mathcal{A}_E, \dots, k} containing an intentional erroneous or suboptimal action AE\mathcal{A}_E (1≤E≤k<N1 \le E \le k < N), a model M\mathcal{M} must perform two distinct tasks:

    M(A1,…,AE,…,k;I0;Ig)→{Error Detection: Identify AEStep Completion: Predict (Ak+1,Ak+2,…,AN)\mathcal{M}(\mathcal{A}_{1, \dots, \mathcal{A}_E, \dots, k}; \mathcal{I}_0; \mathcal{I}_g) \rightarrow \begin{cases} \text{Error Detection: Identify } \mathcal{A}_E \\ \text{Step Completion: Predict } (\mathcal{A}_{k+1}, \mathcal{A}_{k+2}, \dots, \mathcal{A}_N) \end{cases}
    1. Error Detection: Identifying the specific erroneous action AE\mathcal{A}_E in the initial context, or predicting "None of the above" if no error exists.
    2. Step Completion: Selecting the valid continuation sequence (Ak+1,…,AN)(\mathcal{A}_{k+1}, \dots, \mathcal{A}_N) from multiple-choice options that actively rectifies the error state and achieves the target configuration Ig\mathcal{I}_g.

    To prevent shortcut learning where models select options solely matching final visual similarity with Ig\mathcal{I}_g, CoSPlan includes distractor candidates that transition directly toward Ig\mathcal{I}_g without performing the prerequisite error-correction action.

    CoSPlan encompasses four planning domains:

    • Maze-E: 2D grid path navigation (3×33\times 3 to 8×88\times 8 grids with up to 5 obstacle cells) with cardinal movements (↑,↓,←,→)(\uparrow, \downarrow, \leftarrow, \rightarrow) avoiding red cells.
    • Blocks-World-E: Block-stacking rearrangement (3 to 8 blocks across columns) where only top blocks can be moved to empty columns or top positions.
    • Shuffle-E: Image reconstruction on 1,000 ImageNet samples by pairwise tile swapping.
    • Robo-VQA-E: Real-world object rearrangement consisting of 350 curated image-text pairs derived from the ROM robotics dataset.
  2. Knowl 2 — Scene Graph Incremental Updates (SGI) Algorithm for Step Completion

    algorithm

    Scene Graph Incremental update (SGI) is a training-free planning procedure that maps image pairs into textual scene graphs and sequentially updates intermediate graphs in text space to bridge the reasoning gap between initial and goal visual states.

    Input: Initial image I0I_0, Goal image IgI_g, Context actions A1,…,AE,…,AkA_1, \dots, A_E, \dots, A_k, Candidate MCQ options M={m1,m2,…,mC}M = \{m_1, m_2, \dots, m_C\}
    Output: Selected candidate index m′m'
    1: S0←QUERYM(I0)S_0 \leftarrow \text{QUERY}_M(I_0)
    2: Sg←QUERYM(Ig)S_g \leftarrow \text{QUERY}_M(I_g)
    3: Sc←S0S_c \leftarrow S_0
    4: for Ai∈[A1,…,Ak]A_i \in [A_1, \dots, A_k] do
    5: Sc←SIMULATEM(Sc,Ai)S_c \leftarrow \text{SIMULATE}_M(S_c, A_i)
    6: end for
    7: for each candidate m∈Mm \in M do
    8: Sm←ScS_m \leftarrow S_c
    9: Extract action sequence (Ak+1m,…,ANm)(A^m_{k+1}, \dots, A^m_N) from candidate mm
    10: for each action Ajm∈(Ak+1m,…,ANm)A^m_j \in (A^m_{k+1}, \dots, A^m_N) do
    11: Sm←SIMULATEM(Sm,Ajm)S_m \leftarrow \text{SIMULATE}_M(S_m, A^m_j)
    12: end for
    13: scorem←SIMILARITYM(Sm,Sg)\text{score}_m \leftarrow \text{SIMILARITY}_M(S_m, S_g)
    14: end for
    15: m′←arg⁡max⁡m∈M(scorem)m' \leftarrow \arg\max_{m \in M} (\text{score}_m)
    16: return m′m'

    The operation QUERYM(I)\text{QUERY}_M(I) prompts the vision-language model MM to extract a scene graph JSON capturing object nodes, attributes, and spatial relation edges. SIMULATEM(S,A)\text{SIMULATE}_M(S, A) updates the graph given text action AA. SIMILARITYM(Sm,Sg)\text{SIMILARITY}_M(S_m, S_g) prompts the model to score structural alignment between resultant graph SmS_m and goal graph SgS_g on a 0–100 scale.

  3. Knowl 3 — Scene Graph Incremental Updates (SGI) Algorithm for Error Detection

    algorithm

    SGI for Error Detection tracks sequential scene graph deviations across the executed context actions to isolate the erroneous action step.

    Input: Initial image I0I_0, Goal image IgI_g, Context actions A1,…,AE,…,AkA_1, \dots, A_E, \dots, A_k, Threshold τ=0.75\tau = 0.75
    Output: Identified erroneous action Ai′A_{i'} or "None of the above"
    1: S0←QUERYM(I0)S_0 \leftarrow \text{QUERY}_M(I_0)
    2: Sg←QUERYM(Ig)S_g \leftarrow \text{QUERY}_M(I_g)
    3: Sc←S0S_c \leftarrow S_0
    4: for i=1i = 1 to kk do
    5: Sc←SIMULATEM(Sc,Ai)S_c \leftarrow \text{SIMULATE}_M(S_c, A_i)
    6: simi←SIMILARITYM(Sc,Sg)sim_i \leftarrow \text{SIMILARITY}_M(S_c, S_g)
    7: end for
    8: i′←arg⁡min⁡i∈{1,…,k}simii' \leftarrow \arg\min_{i \in \{1, \dots, k\}} sim_i
    9: if simi′>τsim_{i'} > \tau then
    10: return "None of the above"
    11: else
    12: return Ai′A_{i'}
    13: end if

    The algorithm tracks the structural similarity of intermediate scene graphs ScS_c relative to the goal scene graph SgS_g at every context step. The action producing the minimum similarity simi′sim_{i'} represents the maximum trajectory deviation. If simi′sim_{i'} remains above the decision threshold τ=0.75\tau = 0.75, no action is deemed erroneous.

  4. Knowl 4 — CoSPlan Step Completion Performance Across Vision-Language Models

    data/table

    Top-1 accuracy (%) on the CoSPlan Step Completion task evaluates models on selecting the 5-choice sequence that corrects errors and achieves the goal. Baselines compare Vanilla evaluation (V), Chain-of-Thought (CoT), Scene Graphs (SG), and Scene Graph Incremental updates (SGI).

    VLM Robo-VQA-E (%) Shuffle-E (%) Maze-E (%) Blocks-World-E (%)
    V CoT SG/SGI V CoT SG/SGI V CoT SG/SGI V CoT SG/SGI
    Random 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0 20.0
    Human - 42.1 - - 53.7 - - 95.7 - - 81.8 -
    Llama3-8B 18.1 18.3 19.1 17.3 17.7 18.5 19.5 20.1 21.3 21.3 22.7 23.2
    CoG-VLM 13.1 12.5 21.5 / 22.1 23.1 27.1 23.7 / 26.9 25.1 25.9 26.5 / 29.3 25.5 25.2 26.7 / 26.4
    Janus-pro-7B 14.1 14.7 21.3 / 21.1 23.2 23.1 23.5 / 26.1 20.4 20.2 21.7 / 23.2 24.2 23.1 25.1 / 26.3
    Qwen2 VL-8B 17.1 17.6 18.9 / 19.1 24.1 24.9 25.1 / 25.0 26.5 27.9 28.3 / 28.5 18.1 18.6 18.8 / 18.5
    Qwen3 VL-8B 20.3 21.2 23.4 28.3 29.4 30.3 35.1 34.6 36.2 26.5 27.9 28.3
    Intern-VLM 2 22.1 23.5 25.1 / 32.1 20.1 23.2 23.4 / 25.2 21.6 35.8 41.2 / 43.2 18.3 21.2 18.9 / 29.2
    Intern-VLM 3 27.1 29.4 31.4 / 33.7 24.1 25.6 27.1 / 28.6 48.6 48.8 50.1 / 54.8 25.7 27.1 29.4 / 30.6
    GPT-4o - 48.2 52.2 / 56.4 - 27.6 30.1 / 37.0 - 45.6 46.1 / 56.1 - 49.7 54.3 / 55.3
    GPT-5.1 - 49.3 51.3 / 57.5 - 31.5 33.4 / 35.7 - 49.1 48.3 / 55.7 - 48.1 52.3 / 55.6
    Gemini-3-pro - 61.6 67.3 - 57.4 61.6 - 68.2 70.4 - 69.5 71.3

    Most open-weight models achieve near-random accuracy (~20%) in vanilla settings. Structured representations (SG) improve over CoT, while incremental updates (SGI) yield an average gain of +4.4% over standard SG, achieving up to a 10.3% improvement on Intern-VLM 2.

  5. Knowl 5 — CoSPlan Error Detection Performance Across Vision-Language Models

    data/table

    Top-1 accuracy (%) on the CoSPlan Error Detection task measures the ability of models to identify the erroneous action step from the context sequence.

    VLM Robo-VQA-E (%) Maze-E (%) Blocks-World-E (%)
    V CoT SG / SGI V CoT SG / SGI V CoT SG / SGI
    Random 25.4 25.4 25.4 26.1 26.1 26.1 26.1 26.1 26.1
    Human - 60.0 - - 90.0 - - 80.0 -
    Llama3-8B 11.2 10.7 13.6 18.7 18.3 19.7 27.3 27.6 28.2
    CoG-VLM 32.1 33.4 35.3 / 38.7 6.4 8.4 13.3 / 11.0 41.3 43.1 44.5 / 46.1
    Janus-pro-7B 17.5 18.1 26.1 / 27.6 20.5 19.1 21.0 / 21.6 29.3 31.0 27.6 / 33.2
    Qwen2 VL-8B 9.2 9.1 9.6 / 10.1 20.5 20.8 20.7 / 21.3 32.3 30.6 35.2 / 35.7
    Qwen3 VL-8B 13.1 15.6 18.1 23.8 24.1 25.8 34.1 35.3 36.7
    Intern-VLM 2 24.3 25.2 26.1 / 31.5 32.8 33.1 33.4 / 34.8 36.5 37.9 37.3 / 42.9
    Intern-VLM 3 25.1 26.6 28.1 / 29.5 33.3 34.7 35.1 / 36.3 42.4 43.1 44.3 / 45.1
    GPT-4o - 45.3 44.2 / 57.4 - 40.3 35.3 / 41.1 - 35.1 42.1 / 50.7
    GPT-5.1 - 54.9 46.3 - 39.8 37.7 - 38.3 44.6
    Gemini-3-pro - 57.3 62.5 - 62.6 67.8 - 67.5 71.8

    SGI improves error detection over standard Scene Graphs by an average of 4.1% on Intern-VLM 2, 1.7% on Intern-VLM 3, and 9.2% on GPT-4o, reaching 57.4% on Robo-VQA-E and 50.7% on Blocks-World-E for GPT-4o.

  6. Knowl 6 — Diagnostic Failure Modes of VLMs in Error-Prone Planning

    empirical result

    Empirical analysis of VLMs evaluated on CoSPlan reveals multiple systematic failure modes:

    1. Option Selection Bias: Open-source models exhibit extreme position bias in multiple-choice questions. Under CoT prompting, Janus-Pro-7B selects option 'A' 100% of the time on Blocks-World-E and over 90% on Maze-E. Qwen2-VL selects option 'A' over 75% of the time, resulting in near-random or sub-random performance.
    2. Shortcut / Cheating Vulnerability: When presented with options that directly reach the goal configuration Ig\mathcal{I}_g without rectifying the preceding error, models routinely pick the shortcut option. Under standard Scene Graphs on Maze-E, cheating rates reach 43% for Intern-VLM-2, 41% for Intern-VLM-3, 35% for Qwen3-VL, and 27% for CoG-VLM. SGI mitigates this by independently simulating intermediate transitions, reducing cheating instances to 39%, 35%, 30%, and 23% respectively.
    3. In-Context vs. Out-of-Context Errors: Models struggle significantly more with in-context errors (suboptimal actions involving plausible objects present in the scene) than out-of-context errors (actions referencing random irrelevant objects). On Robo-VQA-E, out-of-context accuracy is +10.1% higher for Intern-VLM and +42.9% higher for GPT-4o compared to in-context errors.
    4. Context Length vs. Future Lookahead: Increasing initial context length improves VLM planning accuracy by reducing remaining distance to the goal. However, evaluating partially revealed MCQ trajectories (kk steps out of 9 total) results in flat, constant accuracy across kk, demonstrating that models fail to mentally simulate multi-step visual trajectories from candidate options.
  7. Knowl 7 — Perception Bottleneck Analysis via Scene Graph Quality

    empirical result

    Evaluating SGI with different scene graph qualities isolates the contribution of visual perception (graph generation accuracy) from downstream symbolic planning.

    Model SG Quality Shuffle-E (%) Maze-E (%) Blocks-World-E (%)
    SG SGI SG SGI SG SGI
    Intern-VLM-2 Noisy 20.3 23.4 38.5 41.3 15.3 25.4
    VLM-generated 23.4 25.2 41.2 43.2 18.9 29.2
    Oracle (GT) 23.7 26.5 42.0 59.0 28.1 34.1
    Qwen3-VL Noisy 27.3 30.6 30.5 43.6 24.3 28.7
    VLM-generated 30.3 31.4 36.2 44.6 28.3 29.5
    Oracle (GT) 30.7 31.8 40.5 55.4 30.3 35.6

    Providing Ground-Truth (Oracle) scene graphs consistently yields the highest SGI performance (e.g., Maze-E step completion rises to 59.0% for Intern-VLM-2 and 55.4% for Qwen3-VL), confirming that visual perception errors in initial scene graph extraction form a primary performance bottleneck. SGI maintains higher relative robustness over standard SG even under artificially perturbed noisy graphs.

  8. Knowl 8 — Generalization of SGI to Spatial VQA and PlanBench Tasks

    empirical result

    SGI generalizes beyond error-corrective image planning to static visual question answering (VQA) and text-based classical planning:

    1. Static Spatial VQA (SpatialEval): Because static VQA has identical initial and goal visual states, SGI operates by simulating the spatial hypotheses of individual MCQ candidate choices. SGI consistently matches or outperforms CoT and vanilla SG across Spatial-Map, Maze-Nav, and Spatial-Grid tasks.
    Model Spatial-Map (%) Maze-Nav (%) Spatial-Grid (%)
    CoT SG SGI CoT SG SGI CoT SG SGI
    CoG-VLM 25.1 36.7 35.8 32.3 32.4 31.2 30.1 34.3 38.2
    Janus-pro-7B 42.4 47.4 47.8 20.8 27.3 29.3 34.4 35.8 36.3
    Intern-VLM-2 36.3 41.3 44.3 28.6 40.5 42.1 33.3 33.8 35.1
    1. PlanBench (Task 8 - Constrained Plan Generation): On Blocksworld Task 8 where goal conditions dynamically change post-initial context, applying SGI to textual scene graphs improves plan validity scores for Qwen2-VL-8B from 13.8 (Vanilla), 14.1 (CoT), and 13.9 (SG) to 14.7 (SGI).
  9. Knowl 9 — Performance Degradation in Multi-Error Sequential Planning

    empirical result

    When evaluated on sequences containing multiple compounding errors under a fixed sequential correction order (e.g., rectify error 1, then rectify error 2), VLM planning accuracy degrades monotonically as error count increases.

    Model # Errors
    1 2 3 4 5 7
    InternVLM-2 (SG) 41.2% 40.3% 37.5% 33.4% 27.4% 22.5%
    InternVLM-3 (SG) 50.1% 49.2% 46.3% 41.5% 36.6% 31.3%

    On Maze-E, accuracy for InternVLM-2 drops from 41.2% (1 error) to 22.5% (7 errors), and InternVLM-3 drops from 50.1% to 31.3%, showing that tracking cascading error states compounds long-horizon reasoning degradation.

  10. Knowl 10 — Limitations of CoSPlan Benchmark and SGI Framework

    limitation

    The CoSPlan benchmark and SGI framework exhibit three primary limitations:

    1. Single-Error Constraint in Benchmark Design: The primary benchmark standardizes evaluation by inserting exactly one isolated erroneous step per context sequence. Real-world physical execution entails compounding, multi-step error distributions.
    2. Static Image Formulation: CoSPlan operates on pairs of static before/after images (I0,Ig)(\mathcal{I}_0, \mathcal{I}_g) rather than interactive simulation environments or continuous video streams where agents can observe physical feedback after executing individual actions.
    3. Inference Compute Overhead: SGI requires iterative VLM prompt calls for every action step in the context sequence and candidate options, leading to a computational cost scaled by (context length+∣MCQ options∣×average candidate steps)(\text{context length} + |\text{MCQ options}| \times \text{average candidate steps}) compared to single-step CoT or static SG prompts.

Coverage note — None was omitted; all key benchmark tasks, formulations, algorithms, main experimental results, diagnostic analyses, ablation studies, and limitations were fully extracted.

Citation

MLA
Grover, S., et al. “CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates”. arXiv, 2025, http://arxiv.org/abs/2512.10342v3.
APA
Grover, S., Pathak, P., Kumar, A., & Rawat, Y. S. (2025). CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates. arXiv. http://arxiv.org/abs/2512.10342v3
Chicago
Grover, S., P. Pathak, A. Kumar, and Y. S. Rawat. 2025. “CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates”. arXiv. http://arxiv.org/abs/2512.10342v3.
Harvard
Grover, S. et al. (2025) “CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2512.10342v3.
Vancouver
1. Grover S, Pathak P, Kumar A, Rawat YS (2025) CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates. arXiv

BibTeX

@article{grover2025cosplan,
  title = {CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates},
  author = {Grover, Shresth and Pathak, Priyank and Kumar, Akash and Rawat, Yogesh S},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2512.10342v3},
  eprint = {2512.10342}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/