Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
Marcel GroplJaewoo JungSeungryong KimMarc PollefeysSung‐Jin Hong
Proposes a training-free evidence retrieval method for vision-language models that backpropagates next-token prediction entropy into visual tokens to iteratively locate and zoom in on fine visual details.
Modern vision-language models frequently fail when answering questions that require reading tiny details or aggregating clues scattered across multiple separate regions of an image, such as in complex documents and crowded scenes. This limitation stems largely from selective perception rather than text generation capacity, as standard architectures lack a reliable mechanism to identify where to look next.
The article demonstrates a training-free, test-time method called entropy-gradient grounding to locate and extract decision-critical visual evidence directly from pretrained vision-language models without requiring auxiliary detectors or model fine-tuning.
To achieve this, the approach calculates the model's predictive uncertainty—specifically the Shannon entropy of the next-token distribution at the first output token—and backpropagates this signal to the projected visual embeddings. The resulting gradient map highlights the image regions that most influence the model's certainty. The pipeline smooths and adaptively thresholds this map to identify and rank multiple coherent regions of interest, then applies an iterative zoom-and-reground loop regulated by a spatial-entropy stopping criterion to prevent over-refinement. The selected image crops are fed alongside the original image for the final answer generation pass.
Evaluations across four distinct model architectures (LLaVA-1.5, LLaVA-1.6, InternVL-3.5, and Qwen-2.5-VL) and seven visual question answering benchmarks reveal several key findings. First, the method consistently improves accuracy on detail-heavy and high-resolution benchmarks, achieving absolute performance increases of up to 15.7 to 19.9 points on the fine-grained visual search benchmark V*. Second, it enhances multi-region document reasoning, improving InternVL-3.5's DocVQA score by over 20 points and TextVQA by nearly 15 points, where single-crop heuristics often fail. Third, the spatial localization accuracy is substantially higher than prior attention-based methods, doubling the intersection-over-union metric from 0.14 to 0.29 on evaluated benchmark subsets. Fourth, the approach achieves these gains without degrading general scene understanding capabilities.
These findings indicate that existing models possess internal uncertainty signals that can reliably guide active perception at test time, offering a cost-effective alternative to costly retraining or specialized architecture redesigns. Organizations can deploy this method to boost accuracy in document processing, OCR-heavy tasks, and visual search workflows. While the method increases average inference time from approximately 0.5–2.0 seconds to roughly 3.0–3.4 seconds per sample due to iterative gradient passes, the adaptive stopping mechanism ensures compute costs remain bounded.
Technical leaders deploying vision-language models should consider integrating uncertainty-guided region extraction for tasks demanding fine-grained visual verification. Practitioners should configure the multi-region selector to retrieve two distinct regions of interest, as ablations show this offers the optimal balance between evidence aggregation and redundancy.
The primary limitation of this approach is that precise localization does not guarantee correct reasoning: if the underlying language model lacks the domain knowledge or logic to interpret the cropped text or visual pattern, errors will persist. Furthermore, localization quality depends directly on the strength of the base model's internal representations.
- Paper: ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration, Haozhan Shen et al. (2025). ZoomEye establishes training-free visual exploration and zooming mechanisms in multimodal LLMs for fine-grained detail extraction, providing foundational context for iterative zoom-and-reground approaches.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). VisCoT formalizes the concept of multi-step visual chain-of-thought and focal-region zooming to resolve fine-grained details in vision-language queries.
- Paper: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention, Wenbin An et al. (2025). AGLA examines training-free prompt-relevance masking and attention mechanisms to localize discriminative visual regions and mitigate hallucinations during decoding.
- Paper: CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation, Yuqi Lin et al. (2023). CLIP-ES demonstrates how gradient-based localization maps can be extracted directly from frozen vision-language models without auxiliary detectors or supervision.
- Paper: Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models, Bryan A. Plummer et al. (2015). Flickr30k Entities introduces the fundamental task and benchmark foundation for mapping descriptive language queries to localized visual regions.
No sufficiently relevant recommendations were found.
