ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan ShenKangjia ZhaoTiancheng ZhaoRuochen XuZilun ZhangMingwei ZhuJianwei Yin

article2025EMNLP64 citations

Presents a training-free, tree-based search algorithm that allows multimodal LLMs to dynamically zoom in and backtrack across high-resolution image regions during test-time inference, enabling smaller open-source models to surpass much larger systems on fine-grained visual reasoning tasks.

Listen

Multimodal artificial intelligence models that process both images and text have advanced rapidly, but they still struggle to resolve fine-grained details in complex, high-resolution visual scenes. Current enhancement methods focus almost entirely on text-level reasoning while treating the visual input as a static, downsampled image. Because the visual context remains fixed, models often overlook critical details necessary for accurate decision-making in real-world scenarios.

The article evaluates and demonstrates ZoomEye, an open-source, training-free, and model-agnostic tree search algorithm designed to enable vision-level reasoning. By treating an image as a hierarchical tree of zoomable sub-regions, the algorithm allows models to mimic human visual exploration—scanning globally, zooming in on areas of interest, and backtracking when needed.

The researchers implemented ZoomEye across diverse model families and evaluated performance on high-resolution benchmarks such as V*Bench, HR-Bench 4K and 8K, and the MME-RealWorld application suite. Credibility is supported by direct comparisons against existing commercial systems, open-source baselines, and alternative visual search methods across hundreds of high-resolution tasks without requiring any specialized retraining or fine-tuning.

The evaluation yielded several key findings. First, ZoomEye delivered substantial and consistent accuracy gains across all evaluated models; for example, overall accuracy improved by 34.57 percentage points for LLaVA-v1.5-7B on VBench and by 17.69 percentage points for InternVL2.5-8B on HR-Bench. Second, the framework enabled smaller 3B-to-8B parameter models to match or exceed the performance of leading proprietary models like GPT-4o on high-resolution tasks. Third, answer accuracy rose markedly when targeted zoom operations succeeded, improving from 54.55% during zoom failures to 93.45% upon success on VBench. Fourth, the experiments demonstrated a vision-level test-time scaling effect, showing that accuracy systematically improves as the model is allowed more visual search steps.

These findings indicate that visual perception bottlenecks can be addressed effectively at inference time without incurring the substantial financial and computational costs of retraining foundational models. This approach reduces the operational expense and hardware footprint required for high-precision visual tasks. However, the evaluation also revealed intrinsic model weaknesses; despite successfully zooming into target areas, models sometimes failed on complex spatial orientation and global-to-local positional reasoning due to limitations in their foundational training.

Organizations deploying multimodal models for high-resolution visual tasks should adopt question-driven visual search techniques to boost accuracy without retraining. System architects can dynamically adjust confidence thresholds and search depth limits to balance inference latency against precision requirements. Before deploying to production environments, practitioners should conduct targeted validation on specialized spatial or orientation tasks where models remain vulnerable.

Confidence in these findings is strong given the consistent multi-benchmark improvements across diverse model architectures. Nevertheless, decision-makers should note certain limitations: ZoomEye relies on heuristic stopping criteria, partitions images into rigid rectangular patches rather than semantic contours, and is designed for natural images rather than structured document or table layouts.

arXiv: 2411.16044
Cover for ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding. Recently, with the integration of test-time scaling techniques, these models have also shown strong potential in visual reasoning. However, most existing reasoning approaches remain text-level in nature: MLLMs are prompted to explore various combinations of textual tokens via their underlying language model, while the visual input remains fixed throughout the reasoning process. This paradigm limits the model’s ability to fully exploit rich visual information, particularly when dealing with images containing numerous fine-grained elements. In such cases, vision-level reasoning becomes crucial—where models dynamically zoom into specific regions of the image to gather detailed visual cues necessary for accurate decision-making. In this paper, we propose Zoom Eye, a training-free, model-agnostic tree search algorithm tailored for vision-level reasoning. Zoom Eye treats an image as a hierarchical tree structure, where each child node represents a zoomed-in sub-region of its parent, and the root corresponds to the full image. The algorithm enables MLLMs to simulate human-like zooming behavior by navigating from root to leaf nodes in search of task-relevant visual evidence. We experiment on a series of elaborate high-resolution benchmarks and the results demonstrate that Zoom Eye not only consistently improves the performance of a series of MLLMs with large margin (e.g., InternVL2.5-8B increases by 15.71% and 17.69% on HR-Bench) but also enables small 3-8B MLLMs to outperform strong large models such as GPT-4o. Our code is available at https://github.com/om-ai-lab/ZoomEye.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Methodology
  • 3.1 Abstraction of Tree Search
  • 3.2 Tree Representation for Image
  • 3.3 Ranking Function
  • 3.4 Stopping Criterion
  • 3.5 Overall Search Algorithm
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Results on High-Resolution Benchmark
  • 4.3 Results on Real-World Benchmark
  • 4.4 Ablation Studies
  • 4.4.1 Vision-level test-time scaling
  • 4.4.2 Does the Zoom operation contribute to the improvement of the MLLM?
  • 4.4.3 Impact of the various number of the split sub-regions
  • 4.5 Compared with Other HR Processing Methods
  • 4.5.1 Zoom Eye vs. V ∗
  • 4.5.2 Zoom Eye vs. Others
  • 4.6 Case Study
  • 5 Related Work
  • 6 Limitations
  • 7 Conclusion
  • Acknowledgements
  • References
  • A Results of More MLLMs on High-Resolution Benchmark
  • B Compared with Other HR Processing Methods
  • B.1 Zoom Eye vs. DC 2
  • B.2 Zoom Eye vs. Pixel Reasoner
  • B.3 Zoom Eye vs. VisCrop
  • C Complete Results on MME-RealWorld Benchmark
  • D Implementation Details
  • D.1 Local Input
  • D.2 Global + Local Input
  • D.3 Additional Settings
  • D.4 Complete Algorithm Workflow

Knowls

  1. Knowl 1 — Hierarchical Image Tree Representation for Vision-Level Reasoning

    model/method

    ZoomEye models an input image II as a hierarchical search tree TT to enable multimodal large language models (MLLMs) to actively explore fine-grained details via zooming in and out. The tree structure is defined as follows:

    • The root node T.root={I,(0,0,1,1)}T.\text{root} = \{I, (0, 0, 1, 1)\} represents the full image with normalized bounding box coordinates (x1,y1,x2,y2)=(0,0,1,1)(x_1, y_1, x_2, y_2) = (0, 0, 1, 1).
    • A general tree node nt={I,bt}n_t = \{I, b_t\} represents a localized patch view defined by normalized bounding box coordinates bt=(x1,t,y1,t,x2,t,y2,t)b_t = (x_{1,t}, y_{1,t}, x_{2,t}, y_{2,t}).
    • If the pixel dimensions of the image patch at ntn_t exceed the vision encoder's native resolution, ntn_t is recursively partitioned into 4 equal-sized sub-patches that serve as child nodes of size 4. Patches are divided recursively until they reach the resolution limit of the vision encoder.

    Navigating this tree allows an MLLM to perform node lookahead (zooming into child patches for higher-resolution inspection) and node backtracking (returning to parent or sibling patches to search alternative regions) without modifying the model parameters.

  2. Knowl 2 — Visual Input Formulations for Local and Global-Local MLLMs

    equation

    Depending on the visual architecture and preprocessing capability of the underlying multimodal large language model (MLLM) Φθ\Phi_\theta, the visual input V(nt)\mathcal{V}(n_t) for a given tree node nt={I,bt}n_t = \{I, b_t\} with crop coordinates btb_t is formulated in two ways:

    V(nt)={[F(R(I.crop(bt)))]Local Input[F(R(I)), F(A(I.crop(bt)))]Global+Local Input\mathcal{V}(n_t) = \begin{cases} [F(R(I.\text{crop}(b_t)))] & \text{Local Input} \\[6pt] [F(R(I)),\, F(A(I.\text{crop}(b_t)))] & \text{Global+Local Input} \end{cases}

    where:

    • II denotes the full input image, and I.crop(bt)I.\text{crop}(b_t) is the image sub-region within normalized bounding box bt=(x1,t,y1,t,x2,t,y2,t)b_t = (x_{1,t}, y_{1,t}, x_{2,t}, y_{2,t}).
    • FF is the vision encoder.
    • RR denotes the naive image resizing operation to the encoder's fixed input resolution.
    • AA denotes the AnyRes operation, which partitions an image into a≤Ma \le M equal-area sub-blocks, independently encodes each block and the global view via FF, and integrates their visual token representations.

    In the Local Input setting (used for single-image MLLMs such as LLaVA-v1.5), if combining multiple searched bounding boxes results in a union crop whose longest edge exceeds 1000 pixels, the union crop is avoided to prevent downsampling information loss. Instead, searched patches are cropped and pasted onto a blank image according to their relative spatial positions in II.

    In the Global+Local Input setting (used for AnyRes-capable MLLMs such as LLaVA-OneVision and InternVL2.5), the global image is encoded via naive resizing with a red rectangular visual prompt outlining btb_t, while the cropped patch I.crop(bt)I.\text{crop}(b_t) is processed via AnyRes.

  3. Knowl 3 — Confidence Scoring and Depth-Weighted Ranking for Image Tree Search

    model/method

    To guide tree search without fine-tuning, ZoomEye queries the underlying MLLM Φθ\Phi_\theta to produce normalized confidence scores derived from the next-token probability distribution over the tokens "Yes"\text{"Yes"} and "No"\text{"No"}. For a node ntn_t with visual input V(nt)\mathcal{V}(n_t) and text prompt xx, the scoring operator is:

    LOGITS_RATIO(nt,x)=(softmax(Φθ(y="Yes"∣V(nt),x),Φθ(y="No"∣V(nt),x))[0]−0.5)×2\text{LOGITS\_RATIO}(n_t, x) = \left(\text{softmax}\Big(\Phi_\theta(y = \text{"Yes"} \mid \mathcal{V}(n_t), x), \Phi_\theta(y = \text{"No"} \mid \mathcal{V}(n_t), x)\Big)[0] - 0.5\right) \times 2

    which yields a scalar in (−1,1)(-1, 1).

    Three distinct confidence values are defined:

    1. Existing Confidence (cec_e): Evaluates if target visual cue oo is present in the current patch via prompt pe(o)p_e(o): ce=LOGITS_RATIO(nt,pe(o))c_e = \text{LOGITS\_RATIO}(n_t, p_e(o))
    2. Latent Confidence (clc_l): Evaluates if target visual cue oo could be discovered upon further zooming in the current patch via prompt pl(o)p_l(o): cl=LOGITS_RATIO(nt,pl(o))c_l = \text{LOGITS\_RATIO}(n_t, p_l(o))
    3. Answering Confidence (cac_a): Evaluates if the current patch contains adequate visual information to answer target question qsq_s via prompt pa(qs)p_a(q_s): ca=LOGITS_RATIO(nt,pa(qs))c_a = \text{LOGITS\_RATIO}(n_t, p_a(q_s))

    The node exploration priority is calculated as a depth-weighted sum:

    priority(nt)=α⋅cl+(1−α)⋅ce,where α=W(d)=1−bD2d2+b\text{priority}(n_t) = \alpha \cdot c_l + (1 - \alpha) \cdot c_e, \quad \text{where } \alpha = W(d) = \frac{1 - b}{D^2} d^2 + b

    where dd is the depth of node ntn_t, DD is the maximum depth of the tree, and b∈[0,1]b \in [0, 1] is a base bias constant (b=0.2b = 0.2 for Local Input; b=0.6b = 0.6 for Global+Local Input). As search depth increases, priority shifts weight from latent extrapolation (clc_l) to direct perception (cec_e).

    The search termination criterion S(nt)S(n_t) evaluates whether ca≥τc_a \ge \tau, where τ\tau is an answering confidence threshold.

  4. Knowl 4 — ZoomEye Tree Search and Visual Reasoning Algorithm

    algorithm

    The ZoomEye algorithm guides an MLLM to answer a visual question by extracting visual cues, performing a priority-driven tree search over image patches, and conditioning the final response on the focused region.

    Input: MLLM Φθ\Phi_\theta, image II, question qq, decomposed question template pdqp_{dq}, cue prompts pe,pl,pap_e, p_l, p_a, in-context cue examples, thresholds τ,τ2,τmin\tau, \tau_2, \tau_{min}, step threshold parameters C,δC, \delta.
    Output: Final response string yy.
    1: Visual cues {o1,…,ok}←Φθ.generate(in-context examples,q)\{o_1, \dots, o_k\} \leftarrow \Phi_\theta.\text{generate}(\text{in-context examples}, q)
    2: Initialize empty list L←[]L \leftarrow []
    3: Build image tree TT from II with root T.root={I,(0,0,1,1)}T.root = \{I, (0, 0, 1, 1)\}
    4: for i=1,…,ki = 1, \dots, k do
    5: if k==1k == 1 then qs←qq_s \leftarrow q else qs←pdq(oi)q_s \leftarrow p_{dq}(o_i)
    6: if oio_i is a Type 2 cue (starts with "all") then
    7: Initialize queue Q←[T.root]Q \leftarrow [T.root], cue list Li←[]L_i \leftarrow []
    8: while QQ is not empty do
    9: nt←Q.pop()n_t \leftarrow Q.\text{pop}()
    10: if nt.depth≥2n_t.\text{depth} \ge 2 then break
    11: ce←LOGITS_RATIO(nt,pe(oi))c_e \leftarrow \text{LOGITS\_RATIO}(n_t, p_e(o_i))
    12: if ce≥τ2c_e \ge \tau_2 then Li.append(nt)L_i.\text{append}(n_t)
    13: for each child cc of ntn_t do Q.append(c)Q.\text{append}(c)
    14: else
    15: Initialize queue Q←[T.root]Q \leftarrow [T.root], cue list Li←[]L_i \leftarrow [], count←0count \leftarrow 0, Cstep←T.depth×3C_{step} \leftarrow T.\text{depth} \times 3, nm←T.rootn_m \leftarrow T.root
    16: while QQ is not empty do
    17: nt←Q.pop()n_t \leftarrow Q.\text{pop}()
    18: count←count+1count \leftarrow count + 1
    19: if count≥Cstepcount \ge C_{step} then
    20: τ←τ−0.1\tau \leftarrow \tau - 0.1
    21: Cstep←Cstep+δC_{step} \leftarrow C_{step} + \delta
    22: if τ<τmin\tau < \tau_{min} then break
    23: ca←LOGITS_RATIO(nt,pa(qs))c_a \leftarrow \text{LOGITS\_RATIO}(n_t, p_a(q_s))
    24: if ca≥τc_a \ge \tau then
    25: Li.append(nt)L_i.\text{append}(n_t)
    26: break
    27: if nm.ca≥τn_m.c_a \ge \tau then
    28: Li.append(nm)L_i.\text{append}(n_m)
    29: break
    30: if ca≥nm.cac_a \ge n_m.c_a then nm←ntn_m \leftarrow n_t
    31: for each child cc of ntn_t do
    32: Compute priority(c)\text{priority}(c) using pe(oi)p_e(o_i) and pl(oi)p_l(o_i)
    33: Q.append(c)Q.\text{append}(c)
    34: Q.sort(key=priority,descending=True)Q.\text{sort}(\text{key}=\text{priority}, \text{descending}=\text{True})
    35: L.extend(Li)L.\text{extend}(L_i)
    36: Compute union bounding box b∗=(min⁡ix1,i∗,min⁡iy1,i∗,max⁡ix2,i∗,max⁡iy2,i∗)b^* = (\min_i x_{1,i}^*, \min_i y_{1,i}^*, \max_i x_{2,i}^*, \max_i y_{2,i}^*) from bounding boxes of all nodes in LL
    37: y←Φθ.generate(V({I,b∗}),q)y \leftarrow \Phi_\theta.\text{generate}(\mathcal{V}(\{I, b^*\}), q)
    38: return yy
  5. Knowl 5 — Performance of ZoomEye on High-Resolution Vision-Language Benchmarks

    data/table

    ZoomEye was evaluated across multiple open-source MLLMs on high-resolution benchmarks: V∗V^* Bench (attribute recognition and spatial reasoning, average resolution 2246×15822246 \times 1582), HR-Bench 4K, and HR-Bench 8K (average resolution 76807680, split into Fine-grained Single-instance Perception [FSP] and Fine-grained Cross-instance Perception [FCP]).

    Model V∗V^* Bench HR-Bench 4K HR-Bench 8K
    Attr Spatial Overall FSP FCP Overall FSP FCP Overall
    Baselines and Closed-Source Models
    minigptv2-7B - - - 25.75 25.25 25.50 26.00 26.25 26.13
    LLaVA-v1.6-7B 60.87 63.16 61.78 49.00 46.75 47.88 37.25 44.25 40.75
    Yi-VL-34B - - - 46.00 42.75 44.38 39.50 38.50 39.00
    QWen-VL-max - - - 65.00 52.00 58.50 54.00 51.00 52.50
    GPT4o - - 66.00 70.00 48.00 59.00 62.00 49.00 55.50
    Local Input Mode
    LLaVA-v1.5-7B 43.47 56.57 48.68 38.50 33.75 36.13 33.00 31.25 32.13
    w/ ZoomEye 83.45 82.89 83.25 67.75 38.75 53.25 65.50 36.00 50.75
    Δ\Delta +40.48 +26.32 +34.57 +29.25 +5.00 +17.12 +32.50 +4.75 +18.62
    LLaVA-v1.5-13B 41.74 55.26 47.12 45.25 41.25 43.25 37.50 38.00 37.75
    w/ ZoomEye 87.83 81.58 85.34 73.00 43.25 58.13 67.25 45.50 56.38
    Δ\Delta +46.09 +26.32 +38.22 +27.75 +2.00 +14.88 +29.75 +7.50 +18.63
    Global+Local Input Mode
    LLaVA-ov-0.5B 63.48 64.47 63.87 63.50 39.50 51.50 47.25 38.25 42.75
    w/ ZoomEye 85.22 73.68 80.62 75.50 39.75 57.63 68.50 38.25 53.38
    Δ\Delta +21.74 +9.21 +16.75 +12.00 +0.25 +6.13 +21.25 +0.00 +10.63
    Qwen2.5VL-3B 80.87 71.05 76.96 82.75 49.00 65.88 80.50 45.25 62.88
    w/ ZoomEye 88.70 89.47 89.01 86.75 53.50 70.13 84.75 52.00 68.38
    Δ\Delta +7.83 +18.42 +12.05 +4.00 +4.50 +4.25 +4.25 +6.75 +5.50
    LLaVA-ov-7B 75.65 75.00 75.39 72.00 54.00 63.00 67.25 52.25 59.75
    w/ ZoomEye 93.91 85.53 90.58 84.25 55.00 69.63 88.50 50.00 69.25
    Δ\Delta +18.26 +10.53 +14.19 +12.25 +1.00 +6.63 +21.25 -2.25 +10.00
    InternVL2.5-4B 69.57 71.05 70.16 77.50 53.75 65.63 63.00 49.25 56.13
    w/ ZoomEye 85.22 77.63 82.20 81.25 56.75 69.00 80.00 52.25 66.13
    Δ\Delta +15.65 +6.58 +12.04 +3.75 +3.00 +3.37 +17.00 +3.00 +10.00
    InternVL2.5-8B 67.83 71.05 69.11 75.75 56.25 66.00 61.50 53.25 57.38
    w/ ZoomEye 86.09 82.89 84.82 88.75 61.50 75.13 89.75 57.50 73.63
    Δ\Delta +18.26 +11.84 +15.71 +13.00 +5.25 +9.13 +28.25 +4.25 +16.25
    InternVL2.5-26B 73.91 72.37 73.30 82.00 66.25 74.13 73.00 61.75 67.38
    w/ ZoomEye 91.30 86.84 89.53 89.75 68.25 79.00 89.25 63.00 76.13
    Δ\Delta +17.39 +14.47 +16.23 +7.75 +2.00 +4.87 +16.25 +1.25 +8.75

    These results establish that ZoomEye provides consistent gains across model families and scales. Small models equipped with ZoomEye (such as InternVL2.5-8B achieving 73.63% on HR-Bench 8K) outperform much larger closed-source proprietary systems such as GPT-4o (55.50%).

  6. Knowl 6 — Vision-Level Test-Time Scaling Law in Image Tree Search

    empirical result

    ZoomEye demonstrates a vision-level test-time scaling phenomenon analogous to test-time chain-of-thought scaling in text LLMs. When the answering confidence threshold τ\tau is progressively decreased, the search algorithm explores deeper into the image tree and evaluates more candidate sub-patches.

    Experimental evaluation on V∗V^* Bench with LLaVA-OneVision-7B shows that MLLM accuracy scales monotonically with the number of search steps:

    • Performance rises from ~75.4% at 1 search step to over 90.0% as search steps increase toward 4 to 5, before plateauing.

    This behavior demonstrates that test-time computation allocated to multi-step visual exploration and hierarchical image traversal directly improves visual reasoning accuracy beyond static single-pass perception.

  7. Knowl 7 — Comparison of ZoomEye with Existing High-Resolution Reasoning Frameworks

    data/table

    ZoomEye was compared against alternative high-resolution processing and visual search frameworks: DC2\text{DC}^2 (textual descriptions relayed hierarchically across patches), VisCrop (attention-guided single crop-and-re-feed), Pixel Reasoner (curiosity-driven reinforcement learning for zooming), and V∗V^* (LLM-guided search with auxiliary detector models).

    Model Method Training-free V∗V^* Bench HR-Bench 4K HR-Bench 8K
    LLaVA-v1.5-7B DC2\text{DC}^2 ✓ 57.60 - 39.50
    LLaVA-v1.5-7B VisCrop ✓ 62.30 46.25 35.75
    LLaVA-v1.5-7B ZoomEye (Ours) ✓ 83.25 53.25 50.75
    Qwen2.5-VL-7B Pixel Reasoner ✗ 84.82 - 66.00
    Qwen2.5-VL-3B ZoomEye (Ours) ✓ 89.01 - 68.38
    Method Input Res. Search Res. Zero-shot Indep. search V∗V^* Bench HR-Bench
    V∗V^* search 224px 768px ✗ ✗ 75.39 37.81
    ZoomEye 224px 224px ✓ ✓ 81.58 47.63

    Key advantages demonstrated by these comparisons include:

    1. Zero-shot and Training-free: ZoomEye requires no fine-tuning or reinforcement learning pipelines, yet a 3B model with ZoomEye (89.01% / 68.38%) outperforms the RL-trained Pixel Reasoner with a 7B backbone (84.82% / 66.00%).
    2. Multi-step Search vs Single Refocusing: While VisCrop degrades sharply from HR-4K (46.25%) to HR-8K (35.75%) due to its single "look-again" design, ZoomEye maintains robust performance across extreme resolutions by continuing search until high-confidence evidence is found.
    3. Independent Resolution-Agnostic Operation: Unlike V∗V^*, which requires an external high-resolution vision encoder (OWL-ViT at 768px) and a specialized search model, ZoomEye operates at the native encoder resolution (e.g., 224px) using the MLLM alone.
  8. Knowl 8 — Real-World Task Performance and Failure Modes on MME-RealWorld

    empirical result

    Evaluation of LLaVA-OneVision-7B with ZoomEye on the MME-RealWorld benchmark (2000×15002000 \times 1500 resolution across real-world domains) reveals targeted performance improvements along with specific model failure modes:

    • Significant Improvements: ZoomEye substantially enhances sub-tasks requiring fine detail identification, such as Monitoring Intention (+11.23%), Person Color (+20.22%), Autonomous Driving Motion for vehicles (+29.11%), and Visual Traffic Signals (+12.93%).
    • Failure Modes and Deficiencies:
      1. Orientation Recognition Deficiency: On Monitoring Orientation tasks (e.g., Vehicle Orientation Δ=−0.64%\Delta = -0.64\%, Orientation average Δ=−0.32%\Delta = -0.32\%), ZoomEye correctly zooms into the target object (such as a tricycle), but the base MLLM predicts the wrong heading direction due to an underlying lack of orientation data in pre-training.
      2. Global-to-Local Coordinate Disconnect: On Remote Sensing Position tasks (Δ=−12.95%\Delta = -12.95\%, dropping from 61.40% to 48.45%), ZoomEye successfully locates the target object (such as a blue hexagon in aerial imagery), but the model misidentifies its location relative to the full image canvas (e.g., answering "lower right" instead of "middle top"), revealing difficulty in mapping relative sub-image coordinates back to global scene coordinates.
  9. Knowl 9 — Ablation of Zoom Success Condition and Sub-Region Granularity

    empirical result

    Ablation studies on V∗V^* Bench and HR-Bench with LLaVA-OneVision-7B establish the direct impact of the zoom operation and tree branching factor:

    1. Zoom Success Contribution: A zoom is deemed successful when the final searched bounding box covers at least 50%50\% of the ground-truth target object. On V∗V^* Bench, when ZoomEye successfully zooms in, LLaVA-OneVision-7B achieves an accuracy of 93.45%, compared to 54.55% when the zoom fails.

    2. Impact of Sub-Region Partitioning Granularity (N=4,9,16N = 4, 9, 16): Varying the number of equal sub-patches divided at each tree level affects efficiency and accuracy:

      • 4 sub-regions: V∗V^* Bench = 90.58%, HR-4K = 69.63%, HR-8K = 69.25%, Average Search Steps = 8.20.
      • 9 sub-regions: V∗V^* Bench = 93.19%, HR-4K = 69.75%, HR-8K = 67.63%, Average Search Steps = 5.71.
      • 16 sub-regions: V∗V^* Bench = 92.15%, HR-4K = 70.38%, HR-8K = 69.75%, Average Search Steps = 5.02. Performance remains stable across all split settings, showing that ZoomEye is robust to zooming granularity while higher split counts reduce search steps.
  10. Knowl 10 — Limitations of ZoomEye

    limitation

    The ZoomEye framework exhibits three primary limitations:

    1. Heuristic Search Rules: The priority ranking functions (ce,clc_e, c_l), answering stopping criterion (cac_a), and confidence thresholds (τ,τ2\tau, \tau_2) are manually defined heuristics that may not generalize optimally across all image domains, question structures, or noisy visual conditions.
    2. Non-Semantic Grid Partitioning: Images are partitioned into fixed, uniform spatial grid cells. This rigid geometry does not align with semantic object boundaries and can split or fragment critical visual cues across tile borders.
    3. Domain Restriction to Natural Spatial Scenes: ZoomEye is designed for natural images with spatially distributed visual elements and is unsuited for document and diagram understanding tasks (e.g., forms, dense tables, structured text) where reading order, typography, and tabular relationships are required.

Coverage note — None was omitted; all major contributions—including image tree formulation, scoring/ranking mechanisms, overall algorithm, high-resolution and real-world benchmark evaluations, comparison with prior methods, scaling analysis, ablations, and limitations—are fully covered.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.
    1. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, and 13 others. 2024. Yi: Open foundation models by 01.ai. Preprint, arXiv:2403.04652.
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, and 1 others. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736.
  4. 4.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 1(2):3.
  5. 5.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others. 2025. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923.
  6. 6.Haoran Chen, Junyan Lin, Xinhao Chen, Yue Fan, Xin Jin, Hui Su, Jianfeng Dong, Jinlan Fu, and Xiaoyu Shen. 2025. Rethinking visual layer selection in multimodal llms. Preprint, arXiv:2504.21447.
  7. 7.Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023a. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478.
  8. 8.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023b. Sharegpt4v: Improving large multimodal models with better captions. arXiv preprint arXiv:2311.12793.
  9. 9.Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, and 1 others. 2024a. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271.
  10. 10.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, and 1 others. 2024b. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821.
  11. 11.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. Preprint, arXiv:2305.06500.
  12. 12.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, and 1 others. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234.
  13. 13.Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. 2024. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. arXiv preprint arXiv:2411.14432.
  14. 14.Xidong Feng, Ziyu Wan, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang. 2023. Alphazero-like tree-search can guide large language model decoding and training. arXiv preprint arXiv:2309.17179.
  15. 15.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  16. 16.Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023. Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8154–8173.
  17. 17.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, and 1 others. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720.
  18. 18.Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding language models to images for multimodal inputs and outputs. In International Conference on Machine Learning, pages 17283–17300. PMLR.
  19. 19.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326.
  20. 20.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR.
  21. 21.Yanshu Li, Hongyang He, Yi Cao, Qisen Cheng, Xiang Fu, and Ruixiang Tang. 2025a. M2iv: Towards efficient and fine-grained multimodal in-context learning in large vision-language models. arXiv preprint arXiv:2504.04633.
  22. 22.Yanshu Li, Tian Yun, Jianjiang Yang, Pinyuan Feng, Jinfa Huang, and Ruixiang Tang. 2025b. Taco: Enhancing multimodal in-context learning via task mapping-guided sequence configuration. arXiv preprint arXiv:2505.17098.
  23. 23.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024a. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306.
  24. 24.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024b. Llava-next: Improved reasoning, ocr, and world knowledge.
  25. 25.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024c. Visual instruction tuning. Advances in neural information processing systems, 36.
  26. 26.Jingyu Liu, Jiaen Lin, and Yong Liu. 2024d. How much can rag help the reasoning of llm? arXiv preprint arXiv:2410.02338.
  27. 27.Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, and 1 others. 2024. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525.
  28. 28.Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, and Rongrong Ji. 2024. Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models. arXiv preprint arXiv:2403.03003.
  29. 29.Fanqing Meng, Lingxiao Du, Zongkai Liu, Zhixiang Zhou, Quanfeng Lu, Daocheng Fu, Botian Shi, Wenhai Wang, Junjun He, Kaipeng Zhang, and 1 others. 2025. Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning. arXiv preprint arXiv:2503.07365.
  30. 30.M Minderer, A Gritsenko, A Stone, M Neumann, D Weissenborn, A Dosovitskiy, A Mahendran, A Arnab, M Dehghani, Z Shen, and 1 others. 2022. Simple open-vocabulary object detection with vision transformers. arxiv 2022. arXiv preprint arXiv:2205.06230, 2.
  31. 31.OpenAI. 2025. o3/o4 mini system card. https://openai.com/index/o3-o4-mini-system-card/.
  32. 32.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and 1 others. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR.
  33. 33.Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao. 2025. Vlm-r1: A stable and generalizable r1-style large vision-language model. arXiv preprint arXiv:2504.07615.
  34. 34.Alex Su, Haozhe Wang, Weiming Ren, Fangzhen Lin, and Wenhu Chen. 2025. Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning. arXiv preprint arXiv:2505.15966.
  35. 35.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389.
  36. 36.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and 1 others. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  37. 37.Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. 2024. Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models. arXiv preprint arXiv:2408.15556.
  38. 38.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  39. 39.Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2025. Vary: Scaling up the vision vocabulary for large vision-language model. In European Conference on Computer Vision, pages 408–424. Springer.
  40. 40.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  41. 41.Penghao Wu and Saining Xie. 2024. V?: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094.
  42. 42.Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. 2023. Decomposition enhances reasoning via self-evaluation guided decoding. Preprint, arXiv:2305.00633.
  43. 43.Guowei Xu, Peng Jin, Hao Li, Yibing Song, Lichao Sun, and Li Yuan. 2024. Llava-cot: Let vision language models reason step-by-step. Preprint, arXiv:2411.10440.
  44. 44.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, and 1 others. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671.
  45. 45.Huanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang, Yibo Wang, Shunyu Liu, Yingjie Wang, Yuxin Song, Haocheng Feng, Li Shen, and 1 others. 2024a. Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search. arXiv preprint arXiv:2412.18319.
  46. 46.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024b. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36.
  47. 47.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. Preprint, arXiv:2303.15343.
  48. 48.Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2025. MLLMs know where to look: Training-free perception of small visual details with multimodal LLMs. In The Thirteenth International Conference on Learning Representations.
  49. 49.Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, and 1 others. 2024. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans? arXiv preprint arXiv:2408.13257.
  50. 50.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  51. 51.Tiancheng Zhao, Qianqian Zhang, Kyusong Lee, Peng Liu, Lu Zhang, Chunxin Fang, Jiajia Liao, Kelei Jiang, Yibo Ma, and Ruochen Xu. 2024a. Omchat: A recipe to train multimodal language models with strong long context and video understanding. arXiv preprint arXiv:2407.04923.
  52. 52.Xinping Zhao, Dongfang Li, Yan Zhong, Boren Hu, Yibin Chen, Baotian Hu, and Min Zhang. 2024b. Seer: Self-aligned evidence extraction for retrieval-augmented generation. arXiv preprint arXiv:2410.11315.
  53. 53.Xinping Zhao, Yan Zhong, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Dongfang Li, Baotian Hu, and Min Zhang. 2024c. Funnelrag: A coarse-to-fine progressive retrieval paradigm for rag. arXiv preprint arXiv:2410.10293.
  54. 54.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Jiaxing Zhang, Yujiu Yang, and 1 others. 2023. Solving math word problems via cooperative reasoning induced language models. In The 61st Annual Meeting Of The Association For Computational Linguistics.

Citation

MLA
Shen, H., et al. “ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities Through Tree-Based Image Exploration”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 6602–18, https://doi.org/10.18653/v1/2025.emnlp-main.335.
APA
Shen, H., Zhao, K., Zhao, T., Xu, R., Zhang, Z., Zhu, M., & Yin, J. (2025). ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 6602–6618. https://doi.org/10.18653/v1/2025.emnlp-main.335
Chicago
Shen, H., K. Zhao, T. Zhao, et al. 2025. “ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities Through Tree-Based Image Exploration”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 6602–18. https://doi.org/10.18653/v1/2025.emnlp-main.335.
Harvard
Shen, H. et al. (2025) “ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6602–6618. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.335.
Vancouver
1. Shen H, Zhao K, Zhao T, Xu R, Zhang Z, Zhu M, Yin J (2025) ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6602–6618

BibTeX

@inproceedings{shen-etal-2025-zoomeye,
    title = "{Z}oom{E}ye: Enhancing Multimodal {LLM}s with Human-Like Zooming Capabilities through Tree-Based Image Exploration",
    author = "Shen, Haozhan  and
      Zhao, Kangjia  and
      Zhao, Tiancheng  and
      Xu, Ruochen  and
      Zhang, Zilun  and
      Zhu, Mingwei  and
      Yin, Jianwei",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.335/",
    doi = "10.18653/v1/2025.emnlp-main.335",
    pages = "6602--6618",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/