Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

Marcel GroplJaewoo JungSeungryong KimMarc PollefeysSung‐Jin Hong

article2026arXiv3 citations

Proposes a training-free evidence retrieval method for vision-language models that backpropagates next-token prediction entropy into visual tokens to iteratively locate and zoom in on fine visual details.

Listen

Modern vision-language models frequently fail when answering questions that require reading tiny details or aggregating clues scattered across multiple separate regions of an image, such as in complex documents and crowded scenes. This limitation stems largely from selective perception rather than text generation capacity, as standard architectures lack a reliable mechanism to identify where to look next.

The article demonstrates a training-free, test-time method called entropy-gradient grounding to locate and extract decision-critical visual evidence directly from pretrained vision-language models without requiring auxiliary detectors or model fine-tuning.

To achieve this, the approach calculates the model's predictive uncertainty—specifically the Shannon entropy of the next-token distribution at the first output token—and backpropagates this signal to the projected visual embeddings. The resulting gradient map highlights the image regions that most influence the model's certainty. The pipeline smooths and adaptively thresholds this map to identify and rank multiple coherent regions of interest, then applies an iterative zoom-and-reground loop regulated by a spatial-entropy stopping criterion to prevent over-refinement. The selected image crops are fed alongside the original image for the final answer generation pass.

Evaluations across four distinct model architectures (LLaVA-1.5, LLaVA-1.6, InternVL-3.5, and Qwen-2.5-VL) and seven visual question answering benchmarks reveal several key findings. First, the method consistently improves accuracy on detail-heavy and high-resolution benchmarks, achieving absolute performance increases of up to 15.7 to 19.9 points on the fine-grained visual search benchmark V*. Second, it enhances multi-region document reasoning, improving InternVL-3.5's DocVQA score by over 20 points and TextVQA by nearly 15 points, where single-crop heuristics often fail. Third, the spatial localization accuracy is substantially higher than prior attention-based methods, doubling the intersection-over-union metric from 0.14 to 0.29 on evaluated benchmark subsets. Fourth, the approach achieves these gains without degrading general scene understanding capabilities.

These findings indicate that existing models possess internal uncertainty signals that can reliably guide active perception at test time, offering a cost-effective alternative to costly retraining or specialized architecture redesigns. Organizations can deploy this method to boost accuracy in document processing, OCR-heavy tasks, and visual search workflows. While the method increases average inference time from approximately 0.5–2.0 seconds to roughly 3.0–3.4 seconds per sample due to iterative gradient passes, the adaptive stopping mechanism ensures compute costs remain bounded.

Technical leaders deploying vision-language models should consider integrating uncertainty-guided region extraction for tasks demanding fine-grained visual verification. Practitioners should configure the multi-region selector to retrieve two distinct regions of interest, as ablations show this offers the optimal balance between evidence aggregation and redundancy.

The primary limitation of this approach is that precise localization does not guarantee correct reasoning: if the underlying language model lacks the domain knowledge or logic to interpret the cropped text or visual pattern, errors will persist. Furthermore, localization quality depends directly on the strength of the base model's internal representations.

arXiv: 2604.08456

No sufficiently relevant recommendations were found.

Cover for Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

Abstract

Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread across multiple regions, as in documents and compositional queries. We address this by framing grounding as test-time evidence retrieval: given a query, the model should actively identify where to look next to resolve ambiguity. To this end, we propose a training-free, model-intrinsic grounding method that uses uncertainty as supervision. Specifically, we compute the entropy of the model's next-token distribution and backpropagate it to the visual token embeddings to obtain an entropy-gradient relevance map, without auxiliary detectors or attention-map heuristics. We then extract and rank multiple coherent regions to support multi-evidence queries, and introduce an iterative zoom-and-reground procedure with a spatial-entropy stopping rule to avoid over-refinement. Experiments on seven benchmarks across four VLM architectures demonstrate consistent improvements over existing methods, with the largest gains on detail-critical and high-resolution settings, while also producing more interpretable evidence localizations.

Citation

MLA
Gröpl, M., et al. “Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models”. arXiv, 2026, http://arxiv.org/abs/2604.08456v1.
APA
Gröpl, M., Jung, J., Kim, S., Pollefeys, M., & Hong, S. (2026). Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models. arXiv. http://arxiv.org/abs/2604.08456v1
Chicago
Gröpl, M., J. Jung, S. Kim, M. Pollefeys, and S. Hong. 2026. “Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models”. arXiv. http://arxiv.org/abs/2604.08456v1.
Harvard
Gröpl, M. et al. (2026) “Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.08456v1.
Vancouver
1. Gröpl M, Jung J, Kim S, Pollefeys M, Hong S (2026) Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models. arXiv

BibTeX

@article{gropl2026entropy,
  title = {Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models},
  author = {Gröpl, Marcel and Jung, Jaewoo and Kim, Seungryong and Pollefeys, Marc and Hong, Sunghwan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.08456v1},
  eprint = {2604.08456}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors