ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Xingyu FuMinqian LiuZhengyuan YangJohn CorringYijuan LuJianwei YangDan RothDinei A. F. FlorncioCha Zhang

article2025ICML100 citations

Introduces ReFocus, a method that enables multimodal language models to generate visual chains of thought through programmatic image editing, substantially improving multi-step reasoning over tables and charts.

Listen

Structured visual content such as tables and charts is central to daily communication, scientific reporting, and business intelligence. While multimodal large language models have advanced rapidly, they frequently fail at complex, multi-hop reasoning over structured images. Current systems typically transcribe visual data into text in a single pass and apply text-only reasoning, lacking the ability to selectively focus their visual attention or re-examine specific areas of an image across sequential problem-solving steps.

The article introduces and evaluates REFOCUS, a framework designed to empower multimodal large language models to conduct visual chain-of-thought reasoning by editing input images during the reasoning process. Rather than relying on external tools or expert visual modules, the method enables models to generate executable code that modifies images iteratively to shift and refine visual focus.

To implement this approach, the underlying model is equipped with simple image-manipulation functions—such as drawing bounding boxes, highlighting relevant elements, and masking out irrelevant regions. Coordinates for table rows and columns or chart elements are detected using standard image processing heuristics. When posed with a question, the model generates code to edit the image, simplifying the visual field by eliminating distractions before making its next reasoning step or final prediction. The framework was evaluated across diverse benchmarks covering complex tabular data and varied chart formats, including scientific subplots and horizontal and vertical bar charts.

The evaluation yielded several key findings. First, integrating REFOCUS with leading models produced substantial accuracy improvements across structured image tasks, delivering an average gain of 11.0% on table understanding and 6.8% on chart tasks over standard baselines. Second, the framework allowed the model to achieve performance on image inputs that matched or exceeded configurations where ground-truth text representations were provided alongside the image. Third, the three visual editing methods—highlighting, masking, and bounding boxes—performed similarly, indicating that the primary driver of improvement is selective spatial attention rather than the specific visual styling used. Finally, fine-tuning an open-source vision-language model on a 14,000-example dataset of visual chain-of-thought sequences improved accuracy by 8.0% over standard question-answering data and 2.6% over text-only chain-of-thought data.

These findings indicate that multimodal models suffer significantly from visual clutter and grounding errors, which can be mitigated without external knowledge simply by narrowing the model's visual field. For organizations deploying multimodal models for document analysis, financial report processing, and scientific data extraction, visual editing offers a lightweight, interpretable mechanism to boost accuracy and reduce hallucinations without retraining base architectures.

Organizations evaluating structured image processing should consider incorporating intermediate visual reasoning steps into their workflows. Where fine-tuning is feasible, training models with explicit visual focus coordinates provides stronger supervision than traditional question-and-answer pairs. Future development should explore expanding these tools to more diverse chart types and unstructured document layouts.

Confidence in the reported improvements is high across the tested benchmarks, though practitioners should note that the system relies on reliable coordinate detection heuristics for rows, columns, and bars. Performance gains may vary in highly unconventional visual layouts where layout detection is less reliable.

No sufficiently relevant recommendations were found.

Cover for ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Abstract

Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate "visual thoughts" by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 ReFocus
  • 3.1 Structured Image Problems
  • 3.2 Visual Editing Tools
  • 3.3 Equip LLMs with Visual Editing Tools
  • 4 Experiments and Analyses
  • 4.1 Baselines and Setups
  • 4.2 Results
  • 4.3 Analyses
  • 5 Finetune with ReFocus data
  • 6 Conclusion
  • References
  • A ReFocus Details
  • A.1 Visual Editing Tools for Charts
  • A.2 Experiment Details
  • B Finetune Details
  • B.1 ReFocus Dataset Statistics
  • B.2 Finetune Experiment Details
  • B.3 SFT Result Analyses
  • C Prompts
  • D SFT Qualitative Examples

Knowls

  1. Knowl 1 — Paper content unavailable

    limitation

    No knowl could be extracted because the paper's text was not available for analysis. A valid extraction requires access to the paper's contribution sections (methods, theory, experiments, results, and stated limitations).

Coverage note — The attached paper's content was not accessible for processing, so no knowls could be extracted; the single entry below is a placeholder acknowledging this rather than extracted knowledge.

Citation

MLA
Fu, X., et al. “ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding”. arXiv, 2025, http://arxiv.org/abs/2501.05452v1.
APA
Fu, X., Liu, M., Yang, Z., Corring, J., Lu, Y., Yang, J., Roth, D., Florencio, D., & Zhang, C. (2025). ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding. arXiv. http://arxiv.org/abs/2501.05452v1
Chicago
Fu, X., M. Liu, Z. Yang, et al. 2025. “ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding”. arXiv. http://arxiv.org/abs/2501.05452v1.
Harvard
Fu, X. et al. (2025) “ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.05452v1.
Vancouver
1. Fu X, Liu M, Yang Z, Corring J, Lu Y, Yang J, Roth D, Florencio D, Zhang C (2025) ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding. arXiv

BibTeX

@article{fu2025refocus,
  title = {ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding},
  author = {Fu, Xingyu and Liu, Minqian and Yang, Zhengyuan and Corring, John and Lu, Yijuan and Yang, Jianwei and Roth, Dan and Florencio, Dinei and Zhang, Cha},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.05452v1},
  eprint = {2501.05452}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/