Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visualization correctness

Visualization correctness refers to the degree of accuracy and fidelity with which intermediate visual representations generated during an artificial intelligence reasoning process reflect the true spatial, physical, and contextual state of a given task. In multimodal reasoning frameworks, models generate visual thoughts, such as diagrams, layout maps, or state progressions, to complement verbal reasoning and solve complex spatial problems. Visualization correctness evaluates whether these generated images faithfully adhere to underlying environmental constraints, preserve entity identities and relative locations, and maintain logical continuity across successive reasoning steps without introducing hallucinations or distortion, ensuring that the model grounds its subsequent deductions in a truthful representation of the problem space.

1 item

Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vulic, Furu Wei

OrganizationsInstitute of Automation, Chinese Academy of SciencesMicrosoftUniversity of Cambridge

Why you should read this

Introduces Multimodal Visualization-of-Thought, a paradigm that enables multimodal models to generate intermediate visual traces using a token discrepancy loss, solving complex dynamic spatial reasoning tasks where traditional Chain-of-Thought fails.

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

Added

2026-09-30