CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing ZhaoYao LuMoo Jin KimZipeng FuZhuoyang ZhangYecheng WuZhaoshuo LiQianli MaSong HanChelsea Finn

article2025CVPR581 citations

Introduces CoT-VLA, a framework that equips vision-language-action models with explicit temporal planning by autoregressively predicting future visual sub-goals prior to action generation, outperforming state-of-the-art robotic manipulation methods in both simulation and the real world.

Listen

Robotic foundation models that translate camera observations and natural language instructions into physical actions have demonstrated considerable promise. However, standard systems map inputs directly to outputs, skipping explicit reasoning or multi-step planning. This direct approach frequently fails during complex manipulation tasks because models can lose track of instructions when environments appear visually ambiguous.

The article demonstrates that introducing explicit visual chain-of-thought reasoning improves robotic control. It develops and evaluates a 7-billion-parameter system called CoT-VLA, which generates an image of an intended visual subgoal before predicting a short sequence of robot actions to achieve that state.

The evaluation used both simulated environments and physical robotic platforms. The base model was pretrained on large collections of robot demonstrations and action-free human video datasets, then fine-tuned on task-specific demonstrations. Testing took place across simulated tasks evaluating spatial, object, and goal reasoning, as well as physical experiments using tabletop robotic arms on single- and multi-instruction manipulation tasks.

Key findings show that CoT-VLA improves physical robot performance by approximately 17% and simulated manipulation success by 6% over existing leading baselines. On real-world tabletop experiments, the system achieved a 78.8% average success rate, compared to a 53.7% baseline for models fine-tuned without pretraining. Ablation tests confirmed that visual reasoning, multi-action predictions, and structured attention mechanisms each contributed to higher task success. Additionally, tests using true goal images raised task success by 40%, confirming that higher-quality visual planning directly boosts execution success.

These findings suggest that enabling robots to plan intermediate visual states improves instruction grounding and operational reliability without requiring specialized state annotations. Furthermore, the approach allows engineering teams to leverage vast, unannotated video datasets to improve robot reasoning, reducing the reliance on costly, teleoperated robot data collection.

Decision-makers should consider piloting visual reasoning architectures for complex manipulation workflows where instruction precision is essential. Future development should focus on optimizing inference speed, as generating intermediate image tokens causes an approximate seven-fold computational slowdown compared to direct-action models. Organizations should also evaluate newer fast-inference or diffusion-based generative models to improve visual quality and operational latency before deploying visual reasoning models in time-critical environments.

Cover for CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Abstract

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input--output mappings, lacking the intermediate reasoning steps crucial for complex manipulation tasks. As a result, existing VLAs lack temporal planning or reasoning capabilities. In this paper, we introduce a method that incorporates explicit visual chain-of-thought (CoT) reasoning into vision-language-action models (VLAs) by predicting future image frames autoregressively as visual goals before generating a short action sequence to achieve these goals. We introduce CoT-VLA, a state-of-the-art 7B VLA that can understand and generate visual and action tokens. Our experimental results demonstrate that CoT-VLA achieves strong performance, outperforming the state-of-the-art VLA model by 17% in real-world manipulation tasks and 6% in simulation benchmarks. Project website: this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 CoT-VLA
  • 3.1 Visual Chain-of-Thought Reasoning
  • 3.2 The Base Vision-Language Model
  • 3.3 Training Procedures
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Evaluations Results
  • 4.3 Ablation Study
  • 4.4 Better Visual Reasoning Helps
  • 5 Conclusion, Limitations and Future Work
  • References
  • 6 Implementation Details
  • 6.1 Data Details
  • 6.2 Hyperparameters
  • 6.3 Training

Knowls

  1. Knowl 1 — Visual Chain-of-Thought Reasoning Formulation for Vision-Language-Action Models

    model/method

    Vision-Language-Action (VLA) models typically map natural language instructions ll and current visual observations st\mathbf{s}_t directly to robot actions a^t∼Pθ(at∣st,l)\hat{\mathbf{a}}_t \sim P_\theta(\mathbf{a}_t \mid \mathbf{s}_t, l). Visual Chain-of-Thought (Visual-CoT) reasoning decomposes this mapping into two sequential phases:

    1. Visual Subgoal Reasoning: The policy first predicts an explicit future visual subgoal image s^t+n\hat{\mathbf{s}}_{t+n} located nn time steps ahead in pixel space: s^t+n∼Pθ(st+n∣st,l)\hat{\mathbf{s}}_{t+n} \sim P_\theta(\mathbf{s}_{t+n} \mid \mathbf{s}_t, l)

    2. Subgoal-Conditioned Action Prediction: The policy then predicts a sequence of mm action steps to reach the generated visual subgoal: {a^t,…,a^t+m}∼Pθ({at,…,at+m}∣st,l,s^t+n)\{\hat{\mathbf{a}}_t, \dots, \hat{\mathbf{a}}_{t+m}\} \sim P_\theta(\{\mathbf{a}_t, \dots, \mathbf{a}_{t+m}\} \mid \mathbf{s}_t, l, \hat{\mathbf{s}}_{t+n})

    This formulation separates visual temporal planning from low-level motor execution. Consequently, the visual reasoning phase can be trained jointly on robot demonstrations Dr={(l,a1:T,s1:T)}\mathcal{D}_r = \{(l, \mathbf{a}_{1:T}, \mathbf{s}_{1:T})\} and action-less captioned video datasets Dv={(l,s1:T)}\mathcal{D}_v = \{(l, \mathbf{s}_{1:T})\}, while the action prediction phase is trained exclusively on action-annotated robot demonstrations Dr\mathcal{D}_r.

  2. Knowl 2 — CoT-VLA Architecture and Hybrid Attention Mechanism

    model/method

    CoT-VLA is a 7B-parameter multimodal model built on VILA-U that unifies visual understanding, autoregressive image generation, and multi-step action prediction within a single transformer backbone.

    Visual observations st\mathbf{s}_t and future subgoal images st+n\mathbf{s}_{t+n} (at 256×256256 \times 256 resolution) are encoded using a unified vision tower and quantized via Residual Quantization (RQ-VAE) into 16×1616 \times 16 spatial patches with a residual depth D=4D = 4, yielding 16×16×416 \times 16 \times 4 discrete visual tokens per image. A depth transformer PδP_\delta autoregressively predicts residual tokens conditioned on the code embedding hj\mathbf{h}_j generated by the language backbone.

    Robot actions ai∈R7\mathbf{a}_i \in \mathbb{R}^7 (7-DoF end-effector control) are discretized by mapping each continuous action dimension independently into 256 discrete bins between the 1st and 99th percentiles of the training distribution, repurposed from the 256 least frequent tokens in the text vocabulary.

    CoT-VLA employs a hybrid attention mechanism:

    • Causal Attention: Enforced during text prompt encoding and autoregressive generation of the 256 discrete visual tokens corresponding to the subgoal image s^t+n\hat{\mathbf{s}}_{t+n}.
    • Full (Bidirectional) Attention: Applied across all action chunk tokens [a^t,…,a^t+m][\hat{\mathbf{a}}_t, \dots, \hat{\mathbf{a}}_{t+m}] using special tokens for coordinates ([x][x]), orientation ([θ][\theta]), and gripper state ([g][g]), allowing all action dimensions across the temporal chunk to interact simultaneously for parallel decoding.
  3. Knowl 3 — CoT-VLA Training Objectives and Multi-Data Pretraining

    equation

    CoT-VLA is trained end-to-end to minimize a joint objective combining discrete visual token prediction loss and action token cross-entropy loss:

    L=Laction+Lvisual\mathcal{L} = \mathcal{L}_{\text{action}} + \mathcal{L}_{\text{visual}}

    The visual subgoal generation loss Lvisual\mathcal{L}_{\text{visual}} optimizes the depth transformer PδP_\delta over all visual token positions jj and residual quantization depths d∈{1,…,D}d \in \{1, \dots, D\} with residual depth D=4D = 4:

    Lvisual=−∑j∑d=1Dlog⁡Pδ(kj,d∣kj,<d,hj)\mathcal{L}_{\text{visual}} = -\sum_{j} \sum_{d=1}^D \log P_\delta(k_{j,d} \mid k_{j,<d}, \mathbf{h}_j)

    where kj,dk_{j,d} is the ground-truth visual code at spatial index jj and depth dd, and hj\mathbf{h}_j is the backbone transformer embedding at position jj.

    The action sequence loss Laction\mathcal{L}_{\text{action}} is the cross-entropy loss for generating an action chunk of length mm (using chunk size m=10m = 10) conditioned on language instruction ll, current observation st\mathbf{s}_t, and target subgoal st+n\mathbf{s}_{t+n}:

    Laction=−∑i=0mlog⁡Pθ(at+i∣l,st,st+n)\mathcal{L}_{\text{action}} = -\sum_{i=0}^{m} \log P_\theta(\mathbf{a}_{t+i} \mid l, \mathbf{s}_t, \mathbf{s}_{t+n})

    During pretraining, Lvisual\mathcal{L}_{\text{visual}} is computed over both Open X-Embodiment (OpenX) robot demonstration data Dr\mathcal{D}_r and action-less video datasets Dv\mathcal{D}_v (EPIC-KITCHENS-100 and Something-Something V2), where subgoal frame offsets nn are sampled uniformly from a dataset-specific range [nl,nu][n_l, n_u]. Laction\mathcal{L}_{\text{action}} is computed exclusively on Dr\mathcal{D}_r.

  4. Knowl 4 — CoT-VLA Closed-Loop Control Algorithm

    algorithm

    At deployment time, CoT-VLA operates in a closed-loop control cycle by iteratively hallucinating an intermediate visual subgoal and then executing an action chunk before acquiring a new environment observation.

    Input: Pretrained CoT-VLA model PθP_\theta, initial visual observation s0obs\mathbf{s}^{\text{obs}}_0, language instruction ll, action chunk size mm, subgoal horizon nn
    t←0t \leftarrow 0
    stobs←s0obs\mathbf{s}^{\text{obs}}_t \leftarrow \mathbf{s}^{\text{obs}}_0
    while task is not completed do
        Sample intermediate subgoal image s^t+n∼Pθ(st+n∣l,stobs)\hat{\mathbf{s}}_{t+n} \sim P_\theta(\mathbf{s}_{t+n} \mid l, \mathbf{s}^{\text{obs}}_t)
        Sample action chunk [a^t,…,a^t+m]∼Pθ(at,…,at+m∣l,stobs,s^t+n)[\hat{\mathbf{a}}_t, \dots, \hat{\mathbf{a}}_{t+m}] \sim P_\theta(\mathbf{a}_t, \dots, \mathbf{a}_{t+m} \mid l, \mathbf{s}^{\text{obs}}_t, \hat{\mathbf{s}}_{t+n})
        for j=0j = 0 to mm do
            Execute action a^t+j\hat{\mathbf{a}}_{t+j} on the robot
        end for
        t←t+m+1t \leftarrow t + m + 1
        stobs←\mathbf{s}^{\text{obs}}_t \leftarrow capture new visual observation from robot camera
    end while
  5. Knowl 5 — LIBERO Benchmark Evaluation Results

    data/table

    CoT-VLA was evaluated on the LIBERO simulation benchmark across four task suites (LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and LIBERO-Long), each comprising 10 distinct manipulation tasks with 50 demonstrations per task. Success rates and standard errors were calculated across 3 seeds with 500 evaluation episodes per suite.

    Method Average (↑\uparrow)) Spatial (↑\uparrow)) Object (↑\uparrow)) Goal (↑\uparrow)) Long (↑\uparrow))
    Diffusion Policy 72.4 ±\pm) 0.7% 78.3 ±\pm) 1.1% 92.5 ±\pm) 0.7% 68.3 ±\pm) 1.2% 50.5 ±\pm) 1.3%
    Octo fine-tuned 75.1 ±\pm) 0.6% 78.9 ±\pm) 1.0% 85.7 ±\pm) 0.9% 84.6 ±\pm) 0.9% 51.1 ±\pm) 1.3%
    OpenVLA fine-tuned 76.5 ±\pm) 0.6% 84.7 ±\pm) 0.9% 88.4 ±\pm) 0.8% 79.2 ±\pm) 1.0% 53.7 ±\pm) 1.3%
    CoT-VLA-7B (ours) 81.13 ±\pm) 0.6% 87.5 ±\pm) 1.4% 91.6 ±\pm) 0.5% 87.6 ±\pm) 0.6% 69.0 ±\pm) 0.8%

    CoT-VLA-7B outperforms all baseline models in average success rate (81.13%) and achieves its largest margin on LIBERO-Long (69.0%, a +15.3% absolute gain over OpenVLA), demonstrating the benefit of visual chain-of-thought planning for extended horizon manipulation.

  6. Knowl 6 — Bridge-V2 Real-Robot Generalization Performance

    data/table

    Real-world evaluation on a 6-DoF WidowX robotic arm across four generalization categories from the Bridge-V2 benchmark (10 trials per task category with partial credit scoring):

    • Visual generalization: "put eggplant into pot" in cluttered environments
    • Motion generalization: "put carrot on plate" with varying plate heights
    • Semantic generalization: "take purple grapes out of pot"
    • Language grounding: "put eggplant or red bottle into pot"
    Category SUSIE Octo OpenVLA CoT-VLA
    Visual 30% 35% 75% 65%
    Motion 10% 10% 45% 60%
    Semantic 20% 0% 40% 50%
    Language 40% 40% 75% 70%

    CoT-VLA outperforms prior approaches on motion (60%) and semantic (50%) generalization while remaining competitive on visual (65%) and language (70%) generalization.

  7. Knowl 7 — Component Ablations: Action Chunking, Hybrid Attention, and Visual CoT

    empirical result

    Ablation experiments conducted on the LIBERO-Spatial and LIBERO-Goal suites demonstrate the cumulative performance gains from each design choice:

    • LIBERO-Spatial Success Rate:

      1. Base VLA (single-step action, causal attention, no visual CoT): 67.5%
        • Action chunking (predicting sequence of mm actions): 73.3% (+5.8%)
        • Hybrid attention (full attention across action tokens): 81.8% (+8.5%)
        • Visual CoT (complete CoT-VLA): 87.5% (+5.7%)
    • LIBERO-Goal Success Rate:

      1. Base VLA: 54.9%
        • Action chunking: 74.9% (+20.0%)
        • Hybrid attention: 79.7% (+4.8%)
        • Visual CoT (complete CoT-VLA): 87.6% (+7.9%)

    These results demonstrate that while action chunking and bidirectional attention across action tokens improve trajectory execution, intermediate visual subgoal generation provides an additional 5.7% to 7.9% improvement.

  8. Knowl 8 — Impact of Pretraining on Downstream Robot Adaptation

    empirical result

    On the Franka-Tabletop setup (a 7-DoF Franka Emika Panda arm not seen during pretraining, evaluated on 6 manipulation tasks with limited demonstrations ranging from 10 to 150 per task):

    • Fine-tuning the base VILA-U model directly on Franka-Tabletop demonstrations without the multi-data pretraining stage yields an average task success rate of 53.7%.
    • Pretraining CoT-VLA on OpenX robot demonstrations combined with action-less video datasets (EPIC-KITCHENS and Something-Something V2) prior to downstream fine-tuning achieves an average success rate of 78.8%.

    This represents a 46.7% relative improvement (+25.1% absolute) attributable to the pretraining stage on diverse robotic and human video datasets.

  9. Knowl 9 — Visual Subgoal Quality Directly Bounds Action Success in Out-of-Distribution Tasks

    data/table

    To evaluate how visual reasoning quality affects robotic task execution, CoT-VLA was tested on out-of-distribution long-horizon tasks on the Franka-Tabletop platform combining unseen subtasks (Task 1: "move the green scallion to the apple-covered book"; Task 2: "move the green cauliflower to the bear-covered book" across 5 trials) under two subgoal conditions:

    Condition Sub-task 1 Sub-task 2
    CoT-VLA with Generated Goal Images 20% 0%
    CoT-VLA with Ground-truth Goal Images 60% 40%

    Supplying ground-truth visual subgoals yields an absolute +40% success rate improvement on both tasks over model-generated subgoals, demonstrating that policy execution performance scales directly with the fidelity and accuracy of the visual reasoning step.

  10. Knowl 10 — Latency Overhead and Inference Bottlenecks of Visual-CoT

    limitation

    Generating intermediate visual subgoals creates specific operational constraints in CoT-VLA:

    1. Inference Latency Bottleneck: Autoregressively generating 256 discrete visual tokens for the subgoal image before decoding action tokens introduces a 7×7\times inference slowdown compared to direct action prediction models, even when using action chunking (m=10m=10) and parallel action decoding.
    2. Visual Fidelity Limitations: Discrete token autoregressive generation exhibits lower visual fidelity than continuous diffusion-based goal generation models.
    3. Chunk Discontinuity: Open-loop execution of 10-step action chunks lacks high-frequency feedback between intermediate chunk timesteps and can cause velocity discontinuities at chunk boundaries.

Coverage note — None was omitted; all key model architectures, algorithmic procedures, loss formulations, simulation benchmarks, real-world robot results, ablations, and stated limitations are covered.

References

  1. 1.Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283, 2024.
  2. 2.Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023.
  3. 3.Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022.
  4. 4.Mathilde Caron, Hugo Touvron, Ishan Misra, Herv'e J'egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
  5. 5.Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al. Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158, 2024.
  6. 6.Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14455–14465, 2024.
  7. 7.Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, Yao Lu, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models. arXiv preprint arXiv:2410.10733, 2024.
  8. 8.Lawrence Yunliang Chen, Simeon Adebola, and Ken Goldberg. Berkeley UR5 demonstration dataset. https://sites.google.com/view/berkeley-ur5/home.
  9. 9.Xi Chen, Xiao Wang, Lucas Beyer, Alexander Kolesnikov, Jialin Wu, Paul Voigtlaender, Basil Mustafa, Sebastian Goodman, Ibrahim Alabdulmohsin, Piotr Padlewski, et al. Pali-3 vision language models: Smaller, faster, stronger. arXiv preprint arXiv:2310.09199, 2023.
  10. 10.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, page 02783649241273668, 2023.
  11. 11.Yiming Ding, Carlos Florensa, Pieter Abbeel, and Mariano Phielipp. Goal-conditioned imitation learning. Advances in neural information processing systems, 32, 2019.
  12. 12.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  13. 13.Yuqing Du, Ksenia Konyushkova, Misha Denil, Akhil Raju, Jessica Landon, Felix Hill, Nando de Freitas, and Serkan Cabi. Vision-language models as success detectors. arXiv preprint arXiv:2303.07280, 2023.
  14. 14.Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B Tenenbaum, et al. Video language planning. arXiv preprint arXiv:2310.10625, 2023.
  15. 15.Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video generation. Advances in Neural Information Processing Systems, 36, 2024.
  16. 16.Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Daniilidis, Chelsea Finn, and Sergey Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. arXiv preprint arXiv:2109.13396, 2021.
  17. 17.Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, and Chelsea Finn. Humanplus: Humanoid shadowing and imitation from humans. arXiv preprint arXiv:2406.10454, 2024.
  18. 18.Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. In Conference on Robot Learning (CoRL), 2024.
  19. 19.Samir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt, and Shuran Song. Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23171–23181, 2023.
  20. 20.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision, pages 5842–5850, 2017.
  21. 21.Huy Ha, Pete Florence, and Shuran Song. Scaling up and distilling down: Language-guided robot skill acquisition. In Conference on Robot Learning, pages 3766–3777. PMLR, 2023.
  22. 22.William Harvey and Frank Wood. Visual chain-of-thought diffusion models. arXiv preprint arXiv:2303.16187, 2023.
  23. 23.Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023.
  24. 24.Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multi-modal language models. arXiv preprint arXiv:2406.09403, 2024.
  25. 25.Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard. Visual language maps for robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 10608–10615. IEEE, 2023.
  26. 26.Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Li Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation. arXiv preprint arXiv:2409.01652, 2024.
  27. 27.Georgios Kapidis, Ronald Poppe, Elsbeth Van Dam, Lucas Noldus, and Remco Veltkamp. Egocentric hand track and object-based human action recognition. In 2019 IEEE SmartWorld, Ubiquitous Intelligence & Computing, Advanced & Trusted Computing, Scalable Computing & Communications, Cloud & Big Data Computing, Internet of People and Smart City Innovation (SmartWorld/SCALCOM/UIC/ATC/CBDCom/IOP/SCI), pages 922–929. IEEE, 2019.
  28. 28.Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. arXiv preprint arXiv:2402.07865, 2024.
  29. 29.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
  30. 30.Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jose Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023.
  31. 31.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  32. 32.Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. Autoregressive image generation using residual quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11523–11532, 2022.
  33. 33.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR, 2023.
  34. 34.Boyi Li, Yue Wang, Jiageng Mao, Boris Ivanovic, Sushant Veer, Karen Leung, and Marco Pavone. Driving everywhere with large language model policy adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14948–14957, 2024.
  35. 35.Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, and Carl Vondrick. Dreamitate: Real-world visuomotor policy learning via video generation. arXiv preprint arXiv:2406.16862, 2024.
  36. 36.Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, and Yang Gao. Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647, 2024.
  37. 37.Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310, 2023.
  38. 38.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  39. 39.Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi. Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26439–26455, 2024.
  40. 40.Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Eureka: Human-level reward design via coding large language models. arXiv preprint arXiv: Arxiv-2310.12931, 2023.
  41. 41.Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, et al. Where are we in the search for an artificial visual cortex for embodied intelligence? Advances in Neural Information Processing Systems, 36:655–677, 2023.
  42. 42.Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, et al. Roboturk: A crowdsourcing platform for robotic skill learning through imitation. In Conference on Robot Learning, pages 879–893. PMLR, 2018.
  43. 43.Oier Mees, Jessica Borja-Diaz, and Wolfram Burgard. Grounding language with visual affordances over unstructured data. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11576–11582. IEEE, 2023.
  44. 44.Zawalski Michał, Chen William, Pertsch Karl, Mees Oier, Finn Chelsea, and Levine Sergey. Robotic control via embodied chain-of-thought reasoning. arXiv preprint arXiv:2407.08693, 2024.
  45. 45.Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo. Embodiedgpt: Vision-language pre-training via embodied chain of thought. Advances in Neural Information Processing Systems, 36, 2024.
  46. 46.Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine. Visual reinforcement learning with imagined goals. Advances in neural information processing systems, 31, 2018.
  47. 47.Fei Ni, Jianye Hao, Shiguang Wu, Longxin Kou, Jiashun Liu, Yan Zheng, Bin Wang, and Yuzheng Zhuang. Generate subgoal images before act: Unlocking the chain-of-thought reasoning in diffusion model for robot manipulation with multimodal prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
  48. 48.Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023.
  49. 49.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  50. 50.Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. Visual chain of thought: bridging logical gaps with multimodal infillings. arXiv preprint arXiv:2305.02317, 2023.
  51. 51.Erick Rosete-Beas, Oier Mees, Gabriel Kalweit, Joschka Boedecker, and Wolfram Burgard. Latent plans for task-agnostic offline reinforcement learning. In Conference on Robot Learning, pages 1838–1849. PMLR, 2023.
  52. 52.V Sanh. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  53. 53.Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models. arXiv preprint arXiv:2403.16999, 2024.
  54. 54.Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Cliport: What and where pathways for robotic manipulation. In Conference on robot learning, pages 894–906. PMLR, 2022.
  55. 55.Mohit Shridhar, Yat Long Lo, and Stephen James. Generative image as action models. arXiv preprint arXiv:2407.07875, 2024.
  56. 56.Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11523–11530. IEEE, 2023.
  57. 57.Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. arXiv preprint arXiv:2303.01469, 2023.
  58. 58.Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
  59. 59.Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213, 2024.
  60. 60.Homer Rich Walke, Kevin Black, Tony Z Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, et al. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning, pages 1723–1736. PMLR, 2023.
  61. 61.Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024.
  62. 62.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 2022.
  63. 63.Chuan Wen, Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, and Pieter Abbeel. Any-point trajectory modeling for policy learning. arXiv preprint arXiv:2401.00025, 2023.
  64. 64.Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, et al. Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514, 2024.
  65. 65.Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, et al. Janus: Decoupling visual encoding for unified multimodal understanding and generation. arXiv preprint arXiv:2410.13848, 2024.
  66. 66.Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. arXiv preprint arXiv:2312.13139, 2023.
  67. 67.Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, et al. Vila-u: a unified foundation model integrating visual understanding and generation. arXiv preprint arXiv:2409.04429, 2024.
  68. 68.Jiannan Xiang, Guangyi Liu, Yi Gu, Qiyue Gao, Yuting Ning, Yuheng Zha, Zeyu Feng, Tianhua Tao, Shibo Hao, Yemin Shi, et al. Pandora: Towards general world model with natural language actions and video states. arXiv preprint arXiv:2406.09455, 2024.
  69. 69.Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024.
  70. 70.Jonathan Yang, Catherine Glossop, Arjun Bhorkar, Dhruv Shah, Quan Vuong, Chelsea Finn, Dorsa Sadigh, and Sergey Levine. Pushing the limits of cross-embodiment learning for manipulation and navigation. arXiv preprint arXiv:2402.19432, 2024.
  71. 71.Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023.
  72. 72.Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans. Video as the new language for real-world decision making. arXiv preprint arXiv:2402.17139, 2024.
  73. 73.Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. arxiv: 2406.07550, 2024.
  74. 74.Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al. Language to rewards for robotic skill synthesis. arXiv preprint arXiv:2306.08647, 2023.
  75. 75.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488, 2022.
  76. 76.Kaifeng Zhang, Zhao-Heng Yin, Weirui Ye, and Yang Gao. Learning manipulation skills through robot chain-of-thought with sparse failure guidance. arXiv preprint arXiv:2405.13573, 2024.
  77. 77.Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023.
  78. 78.Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024.
  79. 79.Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. 2024.
  80. 80.Enshen Zhou, Yiran Qin, Zhenfei Yin, Yuzhou Huang, Ruimao Zhang, Lu Sheng, Yu Qiao, and Jing Shao. Minedreamer: Learning to follow instructions via chain-of-imagination for simulated-world control. arXiv preprint arXiv:2403.12037, 2024.
  81. 81.Gaoyue Zhou, Victoria Dean, Mohan Kumar Srirama, Aravind Rajeswaran, Jyothish Pari, Kyle Hatch, Aryan Jain, Tianhe Yu, Pieter Abbeel, Lerrel Pinto, et al. Train offline, test online: A real robot learning benchmark. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9197–9203. IEEE, 2023.
  82. 82.Xinghao Zhu, Ran Tian, Chenfeng Xu, Mingxiao Huo, Wei Zhan, Masayoshi Tomizuka, and Mingyu Ding. Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot. https://sites.google.com/berkeley.edu/fanuc-manipulation, 2023.
  83. 83.Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu. Viola: Imitation learning for vision-based manipulation with object proposal priors. In Conference on Robot Learning, pages 1199–1210. PMLR, 2023.

Citation

MLA
Zhao, Q., et al. “CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 1702–13, https://doi.org/10.1109/CVPR52734.2025.00166.
APA
Zhao, Q., Lu, Y., Kim, M. J., Fu, Z., Zhang, Z., Wu, Y., Li, Z., Ma, Q., Han, S., Finn, C., Handa, A., Lin, T.-Y., Wetzstein, G., Liu, M.-Y., & Xiang, D. (2025). CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1702–1713. https://doi.org/10.1109/CVPR52734.2025.00166
Chicago
Zhao, Q., Y. Lu, M. J. Kim, et al. 2025. “CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1702–13. https://doi.org/10.1109/CVPR52734.2025.00166.
Harvard
Zhao, Q. et al. (2025) “CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 1702–1713. Available at: https://doi.org/10.1109/CVPR52734.2025.00166.
Vancouver
1. Zhao Q, Lu Y, Kim MJ, et al (2025) CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 1702–1713

BibTeX

@inproceedings{Zhao_2025, title={CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00166}, DOI={10.1109/cvpr52734.2025.00166}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhao, Qingqing and Lu, Yao and Kim, Moo Jin and Fu, Zipeng and Zhang, Zhuoyang and Wu, Yecheng and Li, Zhaoshuo and Ma, Qianli and Han, Song and Finn, Chelsea and Handa, Ankur and Lin, Tsung-Yi and Wetzstein, Gordon and Liu, Ming-Yu and Xiang, Donglai}, year={2025}, month=June, pages={1702–1713} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/