ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Xingyu FuMinqian LiuZhengyuan YangJohn CorringYijuan LuJianwei YangDan RothDinei A. F. FlorncioCha Zhang
Introduces ReFocus, a method that enables multimodal language models to generate visual chains of thought through programmatic image editing, substantially improving multi-step reasoning over tables and charts.
Structured visual content such as tables and charts is central to daily communication, scientific reporting, and business intelligence. While multimodal large language models have advanced rapidly, they frequently fail at complex, multi-hop reasoning over structured images. Current systems typically transcribe visual data into text in a single pass and apply text-only reasoning, lacking the ability to selectively focus their visual attention or re-examine specific areas of an image across sequential problem-solving steps.
The article introduces and evaluates REFOCUS, a framework designed to empower multimodal large language models to conduct visual chain-of-thought reasoning by editing input images during the reasoning process. Rather than relying on external tools or expert visual modules, the method enables models to generate executable code that modifies images iteratively to shift and refine visual focus.
To implement this approach, the underlying model is equipped with simple image-manipulation functions—such as drawing bounding boxes, highlighting relevant elements, and masking out irrelevant regions. Coordinates for table rows and columns or chart elements are detected using standard image processing heuristics. When posed with a question, the model generates code to edit the image, simplifying the visual field by eliminating distractions before making its next reasoning step or final prediction. The framework was evaluated across diverse benchmarks covering complex tabular data and varied chart formats, including scientific subplots and horizontal and vertical bar charts.
The evaluation yielded several key findings. First, integrating REFOCUS with leading models produced substantial accuracy improvements across structured image tasks, delivering an average gain of 11.0% on table understanding and 6.8% on chart tasks over standard baselines. Second, the framework allowed the model to achieve performance on image inputs that matched or exceeded configurations where ground-truth text representations were provided alongside the image. Third, the three visual editing methods—highlighting, masking, and bounding boxes—performed similarly, indicating that the primary driver of improvement is selective spatial attention rather than the specific visual styling used. Finally, fine-tuning an open-source vision-language model on a 14,000-example dataset of visual chain-of-thought sequences improved accuracy by 8.0% over standard question-answering data and 2.6% over text-only chain-of-thought data.
These findings indicate that multimodal models suffer significantly from visual clutter and grounding errors, which can be mitigated without external knowledge simply by narrowing the model's visual field. For organizations deploying multimodal models for document analysis, financial report processing, and scientific data extraction, visual editing offers a lightweight, interpretable mechanism to boost accuracy and reduce hallucinations without retraining base architectures.
Organizations evaluating structured image processing should consider incorporating intermediate visual reasoning steps into their workflows. Where fine-tuning is feasible, training models with explicit visual focus coordinates provides stronger supervision than traditional question-and-answer pairs. Future development should explore expanding these tools to more diverse chart types and unstructured document layouts.
Confidence in the reported improvements is high across the tested benchmarks, though practitioners should note that the system relies on reliable coordinate detection heuristics for rows, columns, and bars. Performance gains may vary in highly unconventional visual layouts where layout detection is less reliable.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Its visual chain-of-thought framework uses region selection and zooming to expose fine-grained evidence, establishing the selective-attention approach that ReFocus develops through explicit image editing.
- Paper: Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering, Peter Anderson et al. (2018). Its task-driven attention over object-aligned image regions provides foundational context for ReFocus’s strategy of directing visual attention to relevant parts of an image.
- Paper: MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering, Fangyu Liu et al. (2023). Its chart-specific pretraining and reasoning objectives introduce the structured chart-understanding challenges that ReFocus addresses with sequential visual refocusing.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). Its ChartQA benchmark defines the chart question-answering setting and reasoning demands used to situate ReFocus’s evaluation.
No sufficiently relevant recommendations were found.
