Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Wenbin AnFeng TianSicong LengJiahao NieHaonan LinQianying WangPing ChenXiaoqin ZhangShijian Lu
Proposes a training-free decoding framework that mitigates object hallucinations in large vision-language models by fusing global image context with prompt-guided local visual features to calibrate model predictions.
Large vision-language models combine computer vision and natural language processing to analyze images and answer text prompts. However, they frequently suffer from object hallucinations, generating descriptions of objects that do not exist in the source image. This lack of factual reliability creates substantial operational and safety risks for real-world artificial intelligence deployments. The article investigates the root causes of this failure and introduces a solution to improve visual grounding.
The main objective of the article is to demonstrate that object hallucinations stem from model attention deficiency toward discriminative image regions, and to evaluate a new decoding method called Assembly of Global and Local Attention (AGLA) designed to mitigate these errors.
To address this issue, the researchers developed a training-free, plug-and-play decoding technique. The method uses an image-prompt matching process that calculates relevance scores between prompt text and image patches, adaptively masking out irrelevant visual areas to create an augmented local view. During text generation, the system combines broad generative features from the original image with focused discriminative features from the augmented image, filtering the final outputs through plausibility constraints. The authors tested this approach across multiple open-source vision-language models and validated performance using several established benchmark datasets covering object probing, multi-object queries, comprehensive multimodal perception, and open-ended caption generation.
The findings show substantial and consistent performance gains across all evaluated settings. First, applying AGLA improved standard object hallucination probing scores by an average of 5.5 percentage points in accuracy and 5.1 percentage points in balanced accuracy over standard decoding. Second, in challenging multi-object queries, the method delivered dramatic gains, elevating accuracy from roughly 10% to over 43% in adversarial testing. Third, in open-ended image captioning, the approach lowered hallucination error rates while simultaneously increasing the descriptive recall of true image details. Finally, ablation tests confirmed that both the adaptive visual masking and the dual-stream feature assembly are vital; omitting either component leads to measurable performance degradation.
These results demonstrate that multimodal hallucination can be significantly curtailed without expensive model retraining, fine-tuning, or external post-generation correction models. By rebalancing how models allocate visual attention, organizations can enhance the factual precision and safety of vision-language deployments while controlling computational costs. The findings also challenge the common assumption that hallucinations are solely due to language priors, showing that deficient visual attention mechanisms play a major role.
For practical application, stakeholders deploying vision-language systems should consider integrating attention-assembly decoding methods as a cost-effective safety filter for visual perception tasks. Further technical work should explore testing across larger proprietary foundation models, evaluating real-time inference latency trade-offs, and running targeted domain pilots. Confidence in the reported results is high across the tested open-source benchmarks, though practitioners should exercise caution regarding potential computational overhead during high-throughput token generation.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). This paper establishes the foundational POPE benchmark and diagnostic framework for evaluating object hallucinations in large vision-language models, which AGLA uses to validate its hallucination mitigation approach.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). This work analyzes how waning visual grounding causes multimodal hallucinations and introduces decoding-time logit calibration, establishing essential context for AGLA's training-free visual attention assembly.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). This study introduces dual global and localized patch processing in vision-language models to capture fine-grained visual details, providing foundational principles for AGLA's multi-scale visual discrimination.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). This paper establishes standard visual instruction tuning baselines and grid-based high-resolution visual processing that underpin modern LVLMs evaluated in AGLA.
- Paper: Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision, Seongyun Lee et al. (2024). This paper explores mitigating multimodal hallucination caused by poor visual grounding, offering important comparative context for training-free visual feature recalibration.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). This paper builds on the synergy of global thumbnails and local crop attention in high-resolution LVLMs to design a training-free token compression framework that optimizes inference efficiency.
