Imagine While Reasoning in Space: Multimodal Visualization-of-Thought
Chengzu LiWenshan WuHuanyu ZhangYan XiaShaoguang MaoLi DongIvan VulicFuru Wei
Introduces Multimodal Visualization-of-Thought, a paradigm that enables multimodal models to generate intermediate visual traces using a token discrepancy loss, solving complex dynamic spatial reasoning tasks where traditional Chain-of-Thought fails.
Artificial intelligence models frequently struggle with dynamic spatial reasoning tasks, such as tracking movements or predicting environmental interactions. While traditional Chain-of-Thought methods encourage step-by-step reasoning via written text, this purely verbal approach breaks down when models must track complex visual layouts, intricate spatial relationships, and evolving physical environments.
The article demonstrates and evaluates Multimodal Visualization-of-Thought, a novel reasoning paradigm that enables multimodal models to "think" in both words and generated images. The objective is to verify whether generating intermediate visual thoughts directly within the reasoning trace improves spatial reasoning performance, robustness, and interpretability.
The authors implemented the approach by fine-tuning the open-source Anole-7B model using Low-Rank Adaptation and introducing a specialized token discrepancy loss to align discrete text and image embeddings. They evaluated the framework across three simulated spatial reasoning benchmarks of varying complexity—MAZE navigation, MINIBEHAVIOR object manipulation, and FROZENLAKE hazard avoidance—using datasets ranging from 5,000 to over 6,800 training examples. They compared the proposed model against direct prompting baselines, verbal reasoning baselines, and frontier systems like GPT-4o.
The evaluations yielded several critical findings. First, Multimodal Visualization-of-Thought demonstrated superior robustness in complex environments, achieving 85.60% accuracy on the visually intricate FROZENLAKE benchmark, whereas standard Chain-of-Thought collapsed to 61.48% (and down to 39.11% on larger grid sizes) primarily due to inaccurate textual coordinate descriptions. Second, the proposed framework maintained high performance across simpler abstract tasks, securing 92.95% accuracy on MAZE and 95.14% on MINIBEHAVIOR. Third, the newly introduced token discrepancy loss was vital for generation fidelity; without it, visual accuracy dropped significantly (e.g., from 93.39% to 63.91% on MAZE), causing severe visual redundancy and degraded overall accuracy. Finally, serving the generated visual steps as plug-ins to proprietary models improved GPT-4o's task accuracy by more than 15 percentage points.
These findings indicate that integrating native visual generation into the reasoning process effectively eliminates the brittle failure modes of text-only spatial descriptions. For decision-makers and system architects, this demonstrates that multimodal reasoning is more reliable and interpretable than purely verbal reasoning for spatial and physical planning tasks. Moreover, combining textual and visual reasoning strategies achieved an upper-bound accuracy between 92% and 100%, indicating that multimodal and verbal reasoning paths strongly complement each other.
Organizations developing spatial AI, robotics, or planning agents should consider adopting hybrid reasoning architectures that combine visual generation with verbal reasoning traces. Where deployment of native multimodal generation is constrained, teams can use the framework as an intermediate visual simulation plug-in to boost existing commercial large language models.
However, decision-makers should note certain limitations: generating intermediate images introduces additional inference latency and computational overhead. Furthermore, the model occasionally introduces blurriness or attempts to reconstruct irrelevant background details rather than focusing solely on critical alterations. Future development should focus on guided diffusion techniques and compact image tokenization to reduce computational costs before deploying this framework into latency-sensitive, high-throughput production environments.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This work establishes zero-shot Chain-of-Thought prompting, the language-based reasoning paradigm that MVoT extends with generated visual traces.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). CoT-VLA carries visual chain-of-thought into robotic manipulation, using generated visual subgoals to guide actions beyond the spatial reasoning setting explored by MVoT.
