Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visualization generation

Visualization generation is the automated process of creating visual representations, such as images, diagrams, or graphical displays, from data, textual prompts, or intermediate computational states. Within artificial intelligence and multimodal systems, this technique translates abstract concepts, spatial configurations, and logical relationships into coherent visual artifacts. Beyond producing visual outputs for human interpretation, visualization generation functions as a mechanism for visual reasoning, enabling computational models to construct and utilize intermediate imagery to complement textual thinking, track dynamic changes, and solve complex spatial or analytical problems.

1 item

Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vulic, Furu Wei

OrganizationsInstitute of Automation, Chinese Academy of SciencesMicrosoftUniversity of Cambridge

Why you should read this

Introduces Multimodal Visualization-of-Thought, a paradigm that enables multimodal models to generate intermediate visual traces using a token discrepancy loss, solving complex dynamic spatial reasoning tasks where traditional Chain-of-Thought fails.

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

Added

2026-09-30