Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Chengzu LiWenshan WuHuanyu ZhangYan XiaShaoguang MaoLi DongIvan VulicFuru Wei

article2025ICML226 citations

Introduces Multimodal Visualization-of-Thought, a paradigm that enables multimodal models to generate intermediate visual traces using a token discrepancy loss, solving complex dynamic spatial reasoning tasks where traditional Chain-of-Thought fails.

Listen

Artificial intelligence models frequently struggle with dynamic spatial reasoning tasks, such as tracking movements or predicting environmental interactions. While traditional Chain-of-Thought methods encourage step-by-step reasoning via written text, this purely verbal approach breaks down when models must track complex visual layouts, intricate spatial relationships, and evolving physical environments.

The article demonstrates and evaluates Multimodal Visualization-of-Thought, a novel reasoning paradigm that enables multimodal models to "think" in both words and generated images. The objective is to verify whether generating intermediate visual thoughts directly within the reasoning trace improves spatial reasoning performance, robustness, and interpretability.

The authors implemented the approach by fine-tuning the open-source Anole-7B model using Low-Rank Adaptation and introducing a specialized token discrepancy loss to align discrete text and image embeddings. They evaluated the framework across three simulated spatial reasoning benchmarks of varying complexity—MAZE navigation, MINIBEHAVIOR object manipulation, and FROZENLAKE hazard avoidance—using datasets ranging from 5,000 to over 6,800 training examples. They compared the proposed model against direct prompting baselines, verbal reasoning baselines, and frontier systems like GPT-4o.

The evaluations yielded several critical findings. First, Multimodal Visualization-of-Thought demonstrated superior robustness in complex environments, achieving 85.60% accuracy on the visually intricate FROZENLAKE benchmark, whereas standard Chain-of-Thought collapsed to 61.48% (and down to 39.11% on larger grid sizes) primarily due to inaccurate textual coordinate descriptions. Second, the proposed framework maintained high performance across simpler abstract tasks, securing 92.95% accuracy on MAZE and 95.14% on MINIBEHAVIOR. Third, the newly introduced token discrepancy loss was vital for generation fidelity; without it, visual accuracy dropped significantly (e.g., from 93.39% to 63.91% on MAZE), causing severe visual redundancy and degraded overall accuracy. Finally, serving the generated visual steps as plug-ins to proprietary models improved GPT-4o's task accuracy by more than 15 percentage points.

These findings indicate that integrating native visual generation into the reasoning process effectively eliminates the brittle failure modes of text-only spatial descriptions. For decision-makers and system architects, this demonstrates that multimodal reasoning is more reliable and interpretable than purely verbal reasoning for spatial and physical planning tasks. Moreover, combining textual and visual reasoning strategies achieved an upper-bound accuracy between 92% and 100%, indicating that multimodal and verbal reasoning paths strongly complement each other.

Organizations developing spatial AI, robotics, or planning agents should consider adopting hybrid reasoning architectures that combine visual generation with verbal reasoning traces. Where deployment of native multimodal generation is constrained, teams can use the framework as an intermediate visual simulation plug-in to boost existing commercial large language models.

However, decision-makers should note certain limitations: generating intermediate images introduces additional inference latency and computational overhead. Furthermore, the model occasionally introduces blurriness or attempts to reconstruct irrelevant background details rather than focusing solely on critical alterations. Future development should focus on guided diffusion techniques and compact image tokenization to reduce computational costs before deploying this framework into latency-sensitive, high-throughput production environments.

  • Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This work establishes zero-shot Chain-of-Thought prompting, the language-based reasoning paradigm that MVoT extends with generated visual traces.
Cover for Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Abstract

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

Table of Contents

  • 1 Introduction
  • 2 Multimodal Visualization-of-Thought (MVoT)
  • 2.1 Formulation
  • 2.2 Training with Autoregressive MLLMs
  • 3 Spatial Reasoning Tasks
  • 3.1 Maze
  • 3.2 MiniBehavior
  • 3.3 FrozenLake
  • 4 Experiments
  • 4.1 Experimental Setups
  • 4.2 Experimental Results
  • 5 Discussions and Ablations
  • 5.1 Visualization Quality
  • 5.2 Further Discussions
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Contributions
  • B Data
  • B.1 Dataset Collection
  • B.2 Dataset Statistics
  • C Experiments
  • C.1 Hyper-Parameters
  • C.2 Prompting Templates
  • D Results
  • D.1 Task Performance
  • D.2 Visualizations

Knowls

  1. Knowl 1 — Multimodal Visualization-of-Thought

    model/method

    Multimodal Visualization-of-Thought (MVoT) is a reasoning approach in which a multimodal language model alternates verbal reasoning steps with generated images that visualize intermediate states. Unlike text-only reasoning, the model can use those visualizations as context for subsequent steps; unlike a tool-based pipeline, the model generates the visual thoughts as part of its own output. The sequence ends with a final answer. The paper evaluates this approach for tracking state changes in spatial environments.

  2. Knowl 2 — Autoregressive formulation of interleaved thoughts

    equation

    For a multimodal input xx, let ziz_i be verbal reasoning step ii, viv_i its image visualization, and PθP_\theta the autoregressive multimodal model with parameters θ\theta. Hats denote sampled outputs. MVoT generates a visualization conditioned on the current verbal step and earlier thoughts, then conditions the next verbal step on the input and all thoughts generated so far:

    v^i∼Pθ(vi∣z^1,v^1,…,v^i−1,z^i)\hat{v}_i \sim P_\theta(v_i \mid \hat{z}_1,\hat{v}_1,\ldots,\hat{v}_{i-1},\hat{z}_i)

    z^i+1∼Pθ(zi+1∣x,z^1,v^1,…,z^i,v^i)\hat{z}_{i+1} \sim P_\theta(z_{i+1} \mid x,\hat{z}_1,\hat{v}_1,\ldots,\hat{z}_i,\hat{v}_i)

    The model repeats this interleaving over reasoning steps and produces its answer conditioned on the resulting trace.

  3. Knowl 3 — Token discrepancy loss for visual-token predictions

    equation

    To address the mismatch between the visual tokenizer’s embedding space and the autoregressive model’s token-prediction space, the paper adds a token discrepancy loss LD\mathcal{L}_D. Let NN be the number of visual codebook entries, ejvise_j^{\mathrm{vis}} the visual embedding for codebook entry jj, and tivist_i^{\mathrm{vis}} the ground-truth visual-token index at sequence position ii. Let Pi∈R1×NP_i \in \mathbb{R}^{1\times N} be the model’s predicted distribution over visual tokens at that position. Define the row vector of mean-squared embedding distances from the target embedding to each codebook embedding as Stivis=[MSE⁡(eivis,e1vis),…,MSE⁡(eivis,eNvis)]S_{t_i^{\mathrm{vis}}} = [\operatorname{MSE}(e_i^{\mathrm{vis}},e_1^{\mathrm{vis}}),\ldots,\operatorname{MSE}(e_i^{\mathrm{vis}},e_N^{\mathrm{vis}})]. For nn visual-token positions, the loss is LD=∑i=1nStivis⋅Pi\mathcal{L}_D = \sum_{i=1}^{n} S_{t_i^{\mathrm{vis}}} \cdot P_i. Thus, probability mass on tokens farther from the target in visual embedding space is penalized more heavily. The training objective adds this term to cross-entropy: L=LC+LD\mathcal{L}=\mathcal{L}_C+\mathcal{L}_D, where LC\mathcal{L}_C is computed over text and image tokens.

  4. Knowl 4 — Autoregressive MVoT implementation and fine-tuning

    model/method

    The experiments implement MVoT with Anole-7B, a model tuned on Chameleon that can generate interleaved text and images. Its causal Transformer processes concatenated text-token and image-token sequences; separate text and image tokenizers map the modalities to discrete tokens. Both tokenizers are kept frozen during fine-tuning. MVoT training uses next-token prediction with cross-entropy on text and image tokens, plus the visual-token discrepancy loss; the Interleaved comparison instead excludes image tokens from the loss. The authors fine-tune model parameters using LoRA in an instruction-tuning setup for 40 epochs, with learning rate 0.0002, training batch size 4, and random seed 42. MVoT training uses 32 MI300X GPUs and gradient accumulation of 2. To reduce sensitivity to image reconstruction noise during recursive generation, training input images are tokenized and detokenized a randomly selected 0–10 times.

  5. Knowl 5 — Spatial tasks and collected dataset characteristics

    data/table

    The evaluation uses three dynamic grid-world tasks: MAZE asks the model to apply a movement sequence and identify the resulting labeled destination; MINIBEHAVIOR models the InstallingAPrinter task, where the agent moves, picks up a printer, carries it to a table, and attempts to place and toggle it; FROZENLAKE asks whether an agent reaches a goal without falling into holes. MAZE uses abstract maze layouts, MINIBEHAVIOR adds object interactions, and FROZENLAKE includes more detailed visual patterns.

    The collected dataset characteristics are:

    • MAZE: grid sizes 3–6; 5 entity types and 5 entities; mean action length 9.11; 4 action types; no fine-grained pattern details; 5,007 training examples and 1,255 test examples.
    • MINIBEHAVIOR: grid sizes 5–8 in the summary statistics; 3 entity types and 3 entities; mean action length 7.83; 7 action types; no fine-grained pattern details; 6,400 training examples and 1,604 test examples. The task’s per-grid-size breakdown covers sizes 7–10.
    • FROZENLAKE: grid sizes 3–6; 3 entity types and an average of 7.16 entities; mean action length 6.56; 4 action types; includes fine-grained pattern details; 6,846 training examples and 1,664 test examples.

    The datasets therefore vary in visual detail, action repertoire, and state complexity, enabling evaluation beyond static image understanding.

  6. Knowl 6 — Overall task accuracy across reasoning systems

    empirical result

    On the collected MAZE, MINIBEHAVIOR, and FROZENLAKE test sets, accuracy compares direct answering, text-only Chain-of-Thought (CoT), standard interleaved training, MVoT, and zero-shot GPT-4o variants. Anole-7B results are fine-tuned systems; GPT-4o results are zero-shot. GPT-4o with visual thought receives visualizations generated by the Anole-7B MVoT model.

    • GPT-4o zero-shot Direct: MAZE 0.7100; MINIBEHAVIOR 0.4576; FROZENLAKE 0.4976.
    • GPT-4o zero-shot CoT: MAZE 0.7386; MINIBEHAVIOR 0.4676; FROZENLAKE 0.4664.
    • GPT-4o with visual thought: MAZE 0.8556; MINIBEHAVIOR 0.6440; FROZENLAKE 0.8021.
    • Anole-7B Direct: MAZE 0.7171; MINIBEHAVIOR 0.7250; FROZENLAKE 0.7788.
    • Anole-7B CoT with textual coordinates and environment layout: MAZE 0.9792; MINIBEHAVIOR 0.9812; FROZENLAKE 0.6148.
    • Anole-7B CoT without environment layout: MAZE 0.7023; MINIBEHAVIOR 0.6000; FROZENLAKE 0.5974.
    • Anole-7B Interleaved: MAZE 0.8678; MINIBEHAVIOR 0.8406; FROZENLAKE 0.6460.
    • Anole-7B MVoT: MAZE 0.9295; MINIBEHAVIOR 0.9514; FROZENLAKE 0.8560.

    MVoT is not the best system on MAZE or MINIBEHAVIOR, where layout-assisted CoT scores higher, but it exceeds Direct on all three tasks and attains the highest listed FROZENLAKE accuracy. Layout-assisted CoT falls below Direct on FROZENLAKE, while MVoT remains strong across the task set.

  7. Knowl 7 — MVoT performance as FROZENLAKE grids grow

    empirical result

    FROZENLAKE becomes more complex as grid size increases: the average number of key entities in the dataset’s combined train and development splits grows from 4.7030 on 3×3 grids to 9.648 on 6×6 grids. Accuracy by grid size is:

    • Direct: 0.8263 (3×3), 0.8038 (4×4), 0.7468 (5×5), 0.7533 (6×6).
    • CoT with environment layout: 0.9401, 0.7225, 0.5000, 0.3911, respectively.
    • MVoT: 0.8623, 0.8397, 0.8377, 0.8876, respectively.

    MVoT stays between 0.8377 and 0.8876 across these grid sizes, whereas layout-assisted CoT drops sharply as the grids grow. The paper attributes CoT’s FROZENLAKE errors in part to incorrect textual descriptions of hole coordinates; 90.80% of its mistakes on that task were attributed to this error type.

  8. Knowl 8 — Token discrepancy loss improves visualization and task outcomes

    empirical result

    The paper compares MVoT trained with and without token discrepancy loss (LD\mathcal{L}_D). Visualization Accuracy (V-Acc.) measures whether the intended state modification is shown; Visualization Pattern Redundancy (V-Red.) measures unintended patterns outside the target area; V-Steps is the average length of the initial consecutive correct visualizations; and V-Ratio is the average proportion of those correct visualizations across an action sequence.

    • MAZE with LD\mathcal{L}_D: V-Acc. 0.9339; V-Red. 0.1010; V-Steps 8.4439; V-Ratio 0.9449; task accuracy 0.9295.
    • MAZE without LD\mathcal{L}_D: V-Acc. 0.6391; V-Red. 0.4931; V-Steps 5.6563; V-Ratio 0.6930; task accuracy 0.7468.
    • MINIBEHAVIOR with LD\mathcal{L}_D: V-Acc. 0.9681; V-Red. 0.0419; V-Steps 6.7532; V-Ratio 0.9618; task accuracy 0.9514.
    • MINIBEHAVIOR without LD\mathcal{L}_D: V-Acc. 0.7939; V-Red. 0.2633; V-Steps 5.2333; V-Ratio 0.7793; task accuracy 0.7228.

    For FROZENLAKE, the reported task accuracy also falls from 0.8560 with the loss to 0.7260 without it. Across these comparisons, adding the loss coincides with more accurate, less redundant visualizations and higher task accuracy.

  9. Knowl 9 — CoT and MVoT make complementary errors

    empirical result

    The paper estimates the oracle upper bound from combining CoT and MVoT by counting an example as correct if either system predicts it correctly. CoT accuracies are 0.9792 on MAZE, 0.9812 on MINIBEHAVIOR, and 0.6148 on FROZENLAKE; MVoT accuracies are 0.9295, 0.9514, and 0.8560. The corresponding union upper bounds are 0.9984, 1.0000, and 0.9246. These are oracle bounds, not measured results from a trained ensemble. Their high values indicate that the systems often succeed on different examples, particularly on FROZENLAKE, where MVoT is much stronger than CoT.

  10. Knowl 10 — Visualization fidelity and inference cost remain limitations

    limitation

    MVoT-generated images can contain incorrect state changes, redundant patterns, or reconstructed details that are irrelevant to the reasoning step. The issue is more apparent on visually detailed FROZENLAKE states, where background patterns and other fine details may blur or change during image tokenization, detokenization, and generation. The paper also notes that generating visualizations adds inference-time computational overhead. It proposes guidance techniques for improving image generation and compact image representations that require fewer tokens as future directions; these remedies are not evaluated as part of the reported experiments.

Coverage note — No substantial contributed component was deliberately omitted. Detailed prompting templates, procedural data-generation rules, and auxiliary split-level distributions are left out as supporting implementation detail rather than core findings.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Alan Baddeley. Working memory. Science, 255(5044):556–559, 1992.
  3. 3.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023.
  4. 4.Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros. Sequential modeling enables scalable learning for large vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22861–22872, 2024.
  5. 5.Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. 2024. URL https://openai. com/research/video-generation-models-as-world-simulators, 3, 2024.
  6. 6.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023.
  7. 7.G Brockman. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  8. 8.Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
  9. 9.Christopher F Chabris, Thomas E Jerde, Anita W Woolley, Margaret E Gerbasi, Jonathon P Schuldt, Sean L Bennett, J Richard Hackman, and Stephen M Kosslyn. Spatial and object visualization cognitive styles: Validation studies in 3800 individuals. Group brain technical report, 2:1–20, 2006.
  10. 10.Ethan Chern, Jiadi Su, Yan Ma, and Pengfei Liu. Anole: An open, autoregressive, native large multimodal models for interleaved image-text generation. arXiv preprint arXiv:2407.06135, 2024.
  11. 11.Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14455–14465, 2024.
  12. 12.An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. Spatialrgpt: Grounded spatial reasoning in vision language model. arXiv preprint arXiv:2406.01584, 2024.
  13. 13.Rohan Choudhury, Guanglei Zhu, Sihan Liu, Koichiro Niinuma, Kris M. Kitani, and Laszlo Attila Jeni. Don’t look twice: Faster video transformers with run-length tokenization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  14. 14.Google DeepMind. Google gemini ai update - december 2024, 2024. Accessed: 2024-12-27.
  15. 15.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  16. 16.Patrick Esser, Johnathan Chiu, Parmida Atighechian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7346–7356, 2023.
  17. 17.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12873–12883, June 2021.
  18. 18.Timin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu, Yunhang Shen, Yan Zhang, Shengchuan Zhang, Xiawu Zheng, Xing Sun, Liujuan Cao, and Rongrong Ji. Cantor: Inspiring multimodal chain-of-thought of mllm. In Proceedings of the 32nd ACM International Conference on Multimedia, MM ’24, page 9096–9105, New York, NY, USA, 2024. Association for Computing Machinery.
  19. 19.Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. arXiv preprint arXiv:2406.09403, 2024.
  20. 20.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  21. 21.Michael Igorevich Ivanitskiy, Rusheb Shah, Alex F. Spies, Tilman Räuker, Dan Valentine, Can Rager, Lucia Quirke, Chris Mathwin, Guillaume Corlouer, Cecilia Diniz Behn, and Samy Wu Fung. A configurable library for generating and manipulating maze datasets, 2023.
  22. 22.Emily Jin, Jiaheng Hu, Zhuoyi Huang, Ruohan Zhang, Jiajun Wu, Li Fei-Fei, and Roberto Martín-Martín. Mini-behavior: A procedurally generated benchmark for long-horizon decision-making in embodied ai. arXiv preprint 2310.01824, 2023.
  23. 23.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  24. 24.Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. Aligning cyber space with physical world: A comprehensive survey on embodied ai. arXiv preprint arXiv:2407.06886, 2024.
  25. 25.Fangyu Liu, Guy Emerson, and Nigel Collier. Visual spatial reasoning. Transactions of the Association for Computational Linguistics, 11:635–651, 2023.
  26. 26.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  27. 27.Long Hei Matthew Lam, Ramya Keerthy Thatikonda, and Ehsan Shareghi. A closer look at logical reasoning with llms: The choice of tool matters. arXiv preprint arXiv:2406.00284, 2024.
  28. 28.Xuanyu Lei, Zonghan Yang, Xinrui Chen, Peng Li, and Yang Liu. Scaffolding coordinates to promote vision-language coordination in large multi-modal models. arXiv preprint arXiv:2402.12058, 2024.
  29. 29.Chengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier, Anna Korhonen, and Ivan Vulić. TopViewRS: Vision-language models as top-view spatial reasoners. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1786–1807, Miami, Florida, USA, November 2024. Association for Computational Linguistics.
  30. 30.Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional chain-of-thought prompting for large multimodal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14431, 2024.
  31. 31.Samuel T. Moulton and Stephen M. Kosslyn. Imagining predictions: mental imagery as mental emulation. Philosophical Transactions of the Royal Society B: Biological Sciences, 364:1273 – 1280, 2009.
  32. 32.Debjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh, and Godawari Sudhakar Rao. Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18798–18806, 2024.
  33. 33.Nora S. Newcombe. Spatial Cognition. MIT Press, jul 24 2024. https://oecs.mit.edu/pub/or750iar.
  34. 34.OpenAI. GPT-4V(ision) Technical Work and Authors, 2023.
  35. 35.OpenAI. Hello gpt-4o. https://openai.com/index/hello-gpt-4o/, 2024.
  36. 36.Allan Paivio. Dual coding theory: Retrospect and current status. Canadian Journal of Psychology/Revue canadienne de psychologie, 45(3):255, 1991.
  37. 37.Arijit Ray, Jiafei Duan, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A Plummer, Ranjay Krishna, Kuo-Hao Zeng, et al. Sat: Spatial aptitude training for multimodal language models. arXiv preprint arXiv:2412.07755, 2024.
  38. 38.Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research, 22(268):1–8, 2021.
  39. 39.Yutao Sun, Hangbo Bao, Wenhui Wang, Zhiliang Peng, Li Dong, Shaohan Huang, Jianyong Wang, and Furu Wei. Multimodal latent language modeling with next-token diffusion. arXiv preprint arXiv:2412.08635, 2024.
  40. 40.Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 2024.
  41. 41.Seyedmorteza Sadat, Manuel Kansy, Otmar Hilliges, and Romann M Weber. No training, no problem: Rethinking classifier-free guidance for diffusion models. arXiv preprint arXiv:2407.02687, 2024.
  42. 42.Michael Tschannen, André Susano Pinto, and Alexander Kolesnikov. Jetformer: An autoregressive generative model of raw images and text. arXiv preprint arXiv:2411.19722, 2024.
  43. 43.Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14749–14759, 2024.
  44. 44.Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Yixuan Li, and Neel Joshi. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  45. 45.Wenshan Wu, Shaoguang Mao, Yadong Zhang, Yan Xia, Li Dong, Lei Cui, and Furu Wei. Mind’s eye of LLMs: Visualization-of-thought elicits spatial reasoning in large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  46. 46.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc., 2022.
  47. 47.Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drivedreamer: Towards real-world-driven world models for autonomous driving. arXiv preprint arXiv:2309.09777, 2023.
  48. 48.Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang. Vsp: Assessing the dual challenges of perception and reasoning in spatial planning tasks for vlms, 2024.
  49. 49.Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023.
  50. 50.Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381, 2023.
  51. 51.Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  52. 52.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models, 2023.
  53. 53.Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Adrian de Wynter, Yan Xia, Wenshan Wu, Ting Song, Man Lan, and Furu Wei. Llm as a mastermind: A survey of strategic reasoning with large language models, 2024.
  54. 54.Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024.
  55. 55.Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make pixels dance: High-dynamic video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8850–8860, 2024.
  56. 56.Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, Yang You, Zhaoxiang Zhang, Dawei Zhao, Liang Xiao, Jian Zhao, Jiwen Lu, and Guan Huang. Is sora a world simulator? a comprehensive survey on general world models and beyond. arXiv preprint arXiv:2405.03520, 2024.
  57. 57.Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024.
  58. 58.Qiji Zhou, Ruochen Zhou, Zike Hu, Panzhong Lu, Siyang Gao, and Yue Zhang. Image-of-thought prompting for visual reasoning refinement in multimodal large language models, 2024.
  59. 59.Zhuosheng Zhang, Aston Zhang, Mu Li, hai zhao, George Karypis, and Alex Smola. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research, 2024.

Citation

MLA
Li, C., et al. “Imagine While Reasoning in Space: Multimodal Visualization-of-Thought”. arXiv, 2025, http://arxiv.org/abs/2501.07542v1.
APA
Li, C., Wu, W., Zhang, H., Xia, Y., Mao, S., Dong, L., Vulić, I., & Wei, F. (2025). Imagine while Reasoning in Space: Multimodal Visualization-of-Thought. arXiv. http://arxiv.org/abs/2501.07542v1
Chicago
Li, C., W. Wu, H. Zhang, et al. 2025. “Imagine While Reasoning in Space: Multimodal Visualization-of-Thought”. arXiv. http://arxiv.org/abs/2501.07542v1.
Harvard
Li, C. et al. (2025) “Imagine while Reasoning in Space: Multimodal Visualization-of-Thought”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.07542v1.
Vancouver
1. Li C, Wu W, Zhang H, Xia Y, Mao S, Dong L, Vulić I, Wei F (2025) Imagine while Reasoning in Space: Multimodal Visualization-of-Thought. arXiv

BibTeX

@article{li2025imagine,
  title = {Imagine while Reasoning in Space: Multimodal Visualization-of-Thought},
  author = {Li, Chengzu and Wu, Wenshan and Zhang, Huanyu and Xia, Yan and Mao, Shaoguang and Dong, Li and Vulić, Ivan and Wei, Furu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.07542v1},
  eprint = {2501.07542}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/