ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Haozhan ShenKangjia ZhaoTiancheng ZhaoRuochen XuZilun ZhangMingwei ZhuJianwei Yin
Presents a training-free, tree-based search algorithm that allows multimodal LLMs to dynamically zoom in and backtrack across high-resolution image regions during test-time inference, enabling smaller open-source models to surpass much larger systems on fine-grained visual reasoning tasks.
Multimodal artificial intelligence models that process both images and text have advanced rapidly, but they still struggle to resolve fine-grained details in complex, high-resolution visual scenes. Current enhancement methods focus almost entirely on text-level reasoning while treating the visual input as a static, downsampled image. Because the visual context remains fixed, models often overlook critical details necessary for accurate decision-making in real-world scenarios.
The article evaluates and demonstrates ZoomEye, an open-source, training-free, and model-agnostic tree search algorithm designed to enable vision-level reasoning. By treating an image as a hierarchical tree of zoomable sub-regions, the algorithm allows models to mimic human visual exploration—scanning globally, zooming in on areas of interest, and backtracking when needed.
The researchers implemented ZoomEye across diverse model families and evaluated performance on high-resolution benchmarks such as V*Bench, HR-Bench 4K and 8K, and the MME-RealWorld application suite. Credibility is supported by direct comparisons against existing commercial systems, open-source baselines, and alternative visual search methods across hundreds of high-resolution tasks without requiring any specialized retraining or fine-tuning.
The evaluation yielded several key findings. First, ZoomEye delivered substantial and consistent accuracy gains across all evaluated models; for example, overall accuracy improved by 34.57 percentage points for LLaVA-v1.5-7B on VBench and by 17.69 percentage points for InternVL2.5-8B on HR-Bench. Second, the framework enabled smaller 3B-to-8B parameter models to match or exceed the performance of leading proprietary models like GPT-4o on high-resolution tasks. Third, answer accuracy rose markedly when targeted zoom operations succeeded, improving from 54.55% during zoom failures to 93.45% upon success on VBench. Fourth, the experiments demonstrated a vision-level test-time scaling effect, showing that accuracy systematically improves as the model is allowed more visual search steps.
These findings indicate that visual perception bottlenecks can be addressed effectively at inference time without incurring the substantial financial and computational costs of retraining foundational models. This approach reduces the operational expense and hardware footprint required for high-precision visual tasks. However, the evaluation also revealed intrinsic model weaknesses; despite successfully zooming into target areas, models sometimes failed on complex spatial orientation and global-to-local positional reasoning due to limitations in their foundational training.
Organizations deploying multimodal models for high-resolution visual tasks should adopt question-driven visual search techniques to boost accuracy without retraining. System architects can dynamically adjust confidence thresholds and search depth limits to balance inference latency against precision requirements. Before deploying to production environments, practitioners should conduct targeted validation on specialized spatial or orientation tasks where models remain vulnerable.
Confidence in these findings is strong given the consistent multi-benchmark improvements across diverse model architectures. Nevertheless, decision-makers should note certain limitations: ZoomEye relies on heuristic stopping criteria, partitions images into rigid rectangular patches rather than semantic contours, and is designed for natural images rather than structured document or table layouts.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT establishes the foundational paradigm of multi-step visual chain-of-thought and region zooming to resolve fine-grained details in multimodal models, which ZoomEye generalizes into an inference-time tree search.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). InternVL 2.5 details the architectural design and test-time scaling properties of open-source multimodal LLMs that serve as the primary evaluation backbones and comparison targets in ZoomEye.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Qwen2-VL provides essential background on dynamic resolution processing and multi-dimensional positional embeddings in modern vision-language models, highlighting the architectural bottlenecks in handling high-resolution visual scenes.
- Paper: QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection, Chenhongyi Yang et al. (2022). QueryDet introduces coarse-to-fine cascaded spatial querying for high-resolution vision tasks, providing core conceptual groundwork for hierarchical sub-region exploration.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides the standardized evaluation methodology and perception-reasoning task formulations underpinning the MME-RealWorld suite used to benchmark ZoomEye.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey provides a comprehensive foundation on multimodal large language model architectures, input resolution constraints, and inference-time bottlenecks.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). InternVL3 advances open-source multimodal foundation models by integrating native pre-training recipes and sophisticated test-time scaling strategies that extend inference-time reasoning beyond static downsampling.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL develops advanced cross-layer visual token injection and dynamic spatial-temporal modeling, offering a native architectural approach to the fine-grained visual detail challenges tackled at inference time by ZoomEye.
- Paper: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention, Wenbin An et al. (2025). AGLA builds on fine-grained visual focus during decoding by assembling global and local attention views in a training-free manner to suppress vision-language hallucinations.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). VisionZip complements high-resolution visual exploration techniques by selectively pruning and merging redundant visual tokens to maintain high-resolution efficiency during inference.
- Paper: ShowUI: One Vision-Language-Action Model for GUI Visual Agent, Kevin Qinghong Lin et al. (2025). ShowUI applies efficient localized visual reasoning and visual action grounding directly to high-resolution screenshot exploration for autonomous graphical interface navigation.
- Paper: DeepSeek-OCR 2: Visual Causal Flow, Haoran Wei et al. (2026). DeepSeek-OCR 2 rethinks sequential visual attention flows into semantic, causal reading orders, expanding human-like exploration principles to dense document and visual parsing.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). Tuna-2 explores an alternative foundation paradigm by operating directly on raw pixel embeddings rather than relying on vision encoder architectures to resolve small-object and fine-grained visual details.
