CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Qingqing ZhaoYao LuMoo Jin KimZipeng FuZhuoyang ZhangYecheng WuZhaoshuo LiQianli MaSong HanChelsea Finn
Introduces CoT-VLA, a framework that equips vision-language-action models with explicit temporal planning by autoregressively predicting future visual sub-goals prior to action generation, outperforming state-of-the-art robotic manipulation methods in both simulation and the real world.
Robotic foundation models that translate camera observations and natural language instructions into physical actions have demonstrated considerable promise. However, standard systems map inputs directly to outputs, skipping explicit reasoning or multi-step planning. This direct approach frequently fails during complex manipulation tasks because models can lose track of instructions when environments appear visually ambiguous.
The article demonstrates that introducing explicit visual chain-of-thought reasoning improves robotic control. It develops and evaluates a 7-billion-parameter system called CoT-VLA, which generates an image of an intended visual subgoal before predicting a short sequence of robot actions to achieve that state.
The evaluation used both simulated environments and physical robotic platforms. The base model was pretrained on large collections of robot demonstrations and action-free human video datasets, then fine-tuned on task-specific demonstrations. Testing took place across simulated tasks evaluating spatial, object, and goal reasoning, as well as physical experiments using tabletop robotic arms on single- and multi-instruction manipulation tasks.
Key findings show that CoT-VLA improves physical robot performance by approximately 17% and simulated manipulation success by 6% over existing leading baselines. On real-world tabletop experiments, the system achieved a 78.8% average success rate, compared to a 53.7% baseline for models fine-tuned without pretraining. Ablation tests confirmed that visual reasoning, multi-action predictions, and structured attention mechanisms each contributed to higher task success. Additionally, tests using true goal images raised task success by 40%, confirming that higher-quality visual planning directly boosts execution success.
These findings suggest that enabling robots to plan intermediate visual states improves instruction grounding and operational reliability without requiring specialized state annotations. Furthermore, the approach allows engineering teams to leverage vast, unannotated video datasets to improve robot reasoning, reducing the reliance on costly, teleoperated robot data collection.
Decision-makers should consider piloting visual reasoning architectures for complex manipulation workflows where instruction precision is essential. Future development should focus on optimizing inference speed, as generating intermediate image tokens causes an approximate seven-fold computational slowdown compared to direct-action models. Organizations should also evaluate newer fast-inference or diffusion-based generative models to improve visual quality and operational latency before deploying visual reasoning models in time-critical environments.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). It introduces the open-source 7B vision-language-action baseline architecture and training paradigm that CoT-VLA directly builds upon and enhances with visual chain-of-thought tokens.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). It establishes the foundational vision-language-action formulation for translating internet-scale vision-language models into robotic action token generation.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). It conceptualizes embodied chain-of-thought reasoning to bridge high-level vision-language planning and low-level robotic manipulation.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). It pioneers the integration of multimodal perceptual inputs and robot state data into large language models for embodied task planning and control.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). It establishes step-by-step reasoning and closed-loop environmental feedback for language-model-based robotic action planning.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). It introduces behavioral control through predictive visual imagination of future states, providing the conceptual foundation for predicting visual intermediate goals before acting.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). It extends intermediate reasoning in VLAs beyond autoregressive visual frame prediction by incorporating explicit visual and textual chains of affordances into the action generation pipeline.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). It builds upon predictive VLA reasoning to tackle dynamic object manipulation through low-latency continuous inference and action streaming.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). It generalizes predictive vision-language dynamics by learning foundation world models from uncurated videos to support multi-step System-1 and System-2 planning.
- Paper: Visual Planning: Let's Think Only with Images, Yi Xu et al. (2026). It takes the visual-only reasoning paradigm further by evaluating reinforcement-learned planning purely through sequences of generated images.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). It provides a comprehensive evaluation benchmark measuring how effectively predictive video generation models can act as closed-loop simulators for robotic manipulation.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). It proposes an asymmetric dual-pathway architecture to retain generalist vision-language reasoning capabilities while fine-tuning VLAs for embodied motor control.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). It synthesizes recent architectural developments and planning paradigms across over 80 vision-language-action models, contextualizing approaches like CoT-VLA.
