Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Wenbin AnFeng TianSicong LengJiahao NieHaonan LinQianying WangPing ChenXiaoqin ZhangShijian Lu

article2025CVPR82 citations

Proposes a training-free decoding framework that mitigates object hallucinations in large vision-language models by fusing global image context with prompt-guided local visual features to calibrate model predictions.

Listen

Large vision-language models combine computer vision and natural language processing to analyze images and answer text prompts. However, they frequently suffer from object hallucinations, generating descriptions of objects that do not exist in the source image. This lack of factual reliability creates substantial operational and safety risks for real-world artificial intelligence deployments. The article investigates the root causes of this failure and introduces a solution to improve visual grounding.

The main objective of the article is to demonstrate that object hallucinations stem from model attention deficiency toward discriminative image regions, and to evaluate a new decoding method called Assembly of Global and Local Attention (AGLA) designed to mitigate these errors.

To address this issue, the researchers developed a training-free, plug-and-play decoding technique. The method uses an image-prompt matching process that calculates relevance scores between prompt text and image patches, adaptively masking out irrelevant visual areas to create an augmented local view. During text generation, the system combines broad generative features from the original image with focused discriminative features from the augmented image, filtering the final outputs through plausibility constraints. The authors tested this approach across multiple open-source vision-language models and validated performance using several established benchmark datasets covering object probing, multi-object queries, comprehensive multimodal perception, and open-ended caption generation.

The findings show substantial and consistent performance gains across all evaluated settings. First, applying AGLA improved standard object hallucination probing scores by an average of 5.5 percentage points in accuracy and 5.1 percentage points in balanced accuracy over standard decoding. Second, in challenging multi-object queries, the method delivered dramatic gains, elevating accuracy from roughly 10% to over 43% in adversarial testing. Third, in open-ended image captioning, the approach lowered hallucination error rates while simultaneously increasing the descriptive recall of true image details. Finally, ablation tests confirmed that both the adaptive visual masking and the dual-stream feature assembly are vital; omitting either component leads to measurable performance degradation.

These results demonstrate that multimodal hallucination can be significantly curtailed without expensive model retraining, fine-tuning, or external post-generation correction models. By rebalancing how models allocate visual attention, organizations can enhance the factual precision and safety of vision-language deployments while controlling computational costs. The findings also challenge the common assumption that hallucinations are solely due to language priors, showing that deficient visual attention mechanisms play a major role.

For practical application, stakeholders deploying vision-language systems should consider integrating attention-assembly decoding methods as a cost-effective safety filter for visual perception tasks. Further technical work should explore testing across larger proprietary foundation models, evaluating real-time inference latency trade-offs, and running targeted domain pilots. Confidence in the reported results is high across the tested open-source benchmarks, though practitioners should exercise caution regarding potential computational overhead during high-throughput token generation.

Cover for Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Abstract

Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Large Vision-Language Models
  • 2.2. Object Hallucination
  • 3. Method
  • 3.1. Image-Prompt Matching
  • 3.2. Assembly of Global and Local Attention
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 4.4. Discussion
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Assembly of Global and Local Attention

    model/method

    AGLA is a training-free, plug-and-play decoding method for large vision-language models (LVLMs) that combines two views of an input image. The original image supplies global features for generation, while an image-prompt-matched augmented image emphasizes prompt-relevant local features for visual discrimination. At every decoding step, AGLA combines the LVLM logits from both images, requiring no parameter updates or additional LVLM training. The design targets object hallucinations caused by insufficient attention to local image evidence and by distraction from globally associated but prompt-irrelevant content.

  2. Knowl 2 — Image-Prompt Matching Correlation Map

    equation

    AGLA computes prompt-relevant image regions with a matching model. Let vv be an image, tt its textual prompt, MM the number of prompt tokens, KK the number of image patches, and HH the number of cross-attention heads. For each head hh, let X∈RM×DtX\in\mathbb{R}^{M\times D_t} be prompt-token features, Y∈RK×DvY\in\mathbb{R}^{K\times D_v} be image-patch features, WT(h)∈RDt×DtW_T^{(h)}\in\mathbb{R}^{D_t\times D_t} and WV(h)∈RDv×DtW_V^{(h)}\in\mathbb{R}^{D_v\times D_t} be learned cross-attention projections, and C(h)∈RM×KC^{(h)}\in\mathbb{R}^{M\times K} be the cross-attention matrix:

    C(h)=softmax⁡(XWT(h)(YWV(h))⊤Dt).C^{(h)}=\operatorname{softmax}\left(\frac{XW_T^{(h)}(YW_V^{(h)})^{\top}}{\sqrt{D_t}}\right).

    Here, Cij(h)C_{ij}^{(h)} is the attention assigned by prompt token ii to image patch jj in head hh, and softmax⁡\operatorname{softmax} is applied over image patches. If sim⁡(v,t)\operatorname{sim}(v,t) is the scalar image-prompt similarity produced by the matching model, the correlation score assigned to patch jj is

    cor⁡(j)=1H∑h=1H∑i=1Mmax⁡(0,∂sim⁡(v,t)∂Cij(h))Cij(h).\operatorname{cor}(j)=\frac{1}{H}\sum_{h=1}^{H}\sum_{i=1}^{M}\max\left(0,\frac{\partial\operatorname{sim}(v,t)}{\partial C_{ij}^{(h)}}\right)C_{ij}^{(h)}.

    The positive gradient factor weights cross-attention by how strongly it contributes to image-prompt similarity. In the experiments, BLIP-ITM is used as the matching model and GradCAM is applied to its cross-attention.

  3. Knowl 3 — Adaptive Prompt-Relevant Image Masking

    model/method

    The Image-Prompt Matching module constructs an augmented image by masking image regions with low patch-correlation scores while retaining regions with high scores. The masking ratio is not fixed across examples: for image vv and prompt tt, AGLA sets it to sim⁡(v,t)/2\operatorname{sim}(v,t)/2, where sim⁡(v,t)\operatorname{sim}(v,t) is the matching model's overall similarity score. Thus, the amount of masked content adapts to the image-prompt pair; the paper reports that higher similarity produces a larger masked portion, suppressing more distractions. This operation produces a prompt-dependent local view, but it can remove global information needed for open-ended generation.

  4. Knowl 4 — Global-Local Logit Fusion

    equation

    At decoding step ii, let yiy_i be a candidate token, y<iy_{<i} the already generated token sequence, vv the original image, vaugv^{\mathrm{aug}} the masked augmented image, tt the prompt, and θ\theta the LVLM parameters. Let logit⁡θ(yi∣v,t,y<i)\operatorname{logit}_{\theta}(y_i\mid v,t,y_{<i}) and logit⁡θ(yi∣vaug,t,y<i)\operatorname{logit}_{\theta}(y_i\mid v^{\mathrm{aug}},t,y_{<i}) be the original-image and augmented-image logits, respectively. AGLA forms the decoding distribution as

    pAGLA(yi∣v,vaug,t,y<i)∝softmax⁡[logit⁡θ(yi∣v,t,y<i)+αlogit⁡θ(yi∣vaug,t,y<i)],p_{\mathrm{AGLA}}(y_i\mid v,v^{\mathrm{aug}},t,y_{<i})\propto\operatorname{softmax}\left[\operatorname{logit}_{\theta}(y_i\mid v,t,y_{<i})+\alpha\operatorname{logit}_{\theta}(y_i\mid v^{\mathrm{aug}},t,y_{<i})\right],

    where α≥0\alpha\geq 0 controls the contribution of prompt-relevant local evidence. Unlike visual contrastive decoding, which subtracts logits from a distorted-image distribution interpreted as noise, AGLA adds a useful distribution generated from the locally focused image to the original global distribution.

  5. Knowl 5 — Adaptive Plausibility Truncation

    equation

    To prevent the augmented-image distribution from promoting implausible tokens or suppressing valid global predictions, AGLA retains only tokens that are sufficiently probable under the original image. Let V\mathcal{V} be the output vocabulary, pθ(w∣v,t,y<i)p_{\theta}(w\mid v,t,y_{<i}) the original-image probability of token ww, and β∈[0,1]\beta\in[0,1] the truncation threshold. The retained vocabulary is

    Vtoken(y<i)={yi∈V:pθ(yi∣v,t,y<i)≥βmax⁡w∈Vpθ(w∣v,t,y<i)}.\mathcal{V}_{\mathrm{token}}(y_{<i})=\left\{y_i\in\mathcal{V}:p_{\theta}(y_i\mid v,t,y_{<i})\geq\beta\max_{w\in\mathcal{V}}p_{\theta}(w\mid v,t,y_{<i})\right\}.

    For any candidate outside this set, AGLA sets the fused probability to zero:

    pAGLA(yi∣v,vaug,t,y<i)=0if yi∉Vtoken(y<i).p_{\mathrm{AGLA}}(y_i\mid v,v^{\mathrm{aug}},t,y_{<i})=0\quad\text{if }y_i\notin\mathcal{V}_{\mathrm{token}}(y_{<i}).

    A larger β\beta retains only tokens closer to the original-image maximum-probability token.

  6. Knowl 6 — Attention Deficiency as a Hallucination Mechanism

    theoretical result

    The paper's analysis identifies deficient attention to discriminative image features as one root cause of object hallucinations in LVLMs. When answering different object-existence prompts, self-attention over the original image is dominated by a small set of global features and exhibits similar spatial patterns whether the queried object is present or absent. After Image-Prompt Matching produces an augmented view, attention becomes more concentrated on regions relevant to the particular query. Consistent with this diagnosis, LVLMs hallucinate more frequently in the adversarial POPE setting, where absent objects are selected to co-occur frequently with image content; for example, prompt-irrelevant context such as a road can encourage hallucination of cars. The paper presents this as a major contributing cause rather than a claim that it is the only source of hallucination.

  7. Knowl 7 — Evaluation Protocol Across LVLM Tasks

    experimental setup

    AGLA is evaluated without additional training on LLaVA-1.5 (7B and 13B), InstructBLIP (7B and 13B), Qwen-VL (7B), and MiniCPM-V (2.4B), using BLIP-ITM for image-prompt matching and averaging results over three runs. Comparisons include regular multinomial decoding and the decoding methods DOLA, OPERA, and VCD. The discriminative evaluations use POPE, which contains 27,000 object-existence questions across MSCOCO, A-OKVQA, and GQA under random, popular, and adversarial negative-sampling settings; ROPE, which evaluates multi-object hallucination under adversarial-A, adversarial-B, heterogeneous, homogenous, and mixed settings; and the MME hallucination subset, which covers existence, count, position, and color. Generative evaluations use CHAIR on 500 randomly selected MSCOCO images and LLaVA-Bench-Wild, containing 24 images and 60 questions evaluated by GPT-4 for accuracy and detail.

  8. Knowl 8 — Object-Hallucination Reduction on POPE

    empirical result

    On the POPE object-existence benchmark, AGLA consistently improves over regular decoding for both LLaVA-1.5 (7B) and InstructBLIP (7B) under random, popular, and adversarial negative sampling. The reported accuracy/F1 results are:

    • LLaVA-1.5: random regular 83.49/82.2883.49/82.28 versus AGLA 88.54/87.7188.54/87.71; popular regular 79.98/79.3479.98/79.34 versus AGLA 85.14/84.6885.14/84.68; adversarial regular 76.03/76.2676.03/76.26 versus AGLA 81.13/81.3681.13/81.36.
    • InstructBLIP: random regular 80.42/80.9480.42/80.94 versus AGLA 87.30/87.0787.30/87.07; popular regular 76.09/77.6576.09/77.65 versus AGLA 81.86/82.5881.86/82.58; adversarial regular 72.37/75.4272.37/75.42 versus AGLA 77.29/79.1677.29/79.16.

    Across the six model-setting combinations, the paper reports average gains of 5.5 percentage points in accuracy and 5.1 points in F1 over regular decoding. AGLA also surpasses DOLA, OPERA, and VCD in the reported POPE conditions, with particularly strong gains under adversarial sampling.

  9. Knowl 9 — Mitigation of Multi-Object Hallucination

    empirical result

    On ROPE, which evaluates responses containing multiple objects, AGLA improves F1 over regular decoding on every reported subset for both LLaVA-1.5 (7B) and MiniCPM-V (2.4B). For LLaVA-1.5, the regular-to-AGLA F1 changes are adversarial-A 21.40→46.4521.40\rightarrow46.45, adversarial-B 20.24→47.4220.24\rightarrow47.42, heterogeneous 5.62→9.695.62\rightarrow9.69, homogenous 35.37→65.4935.37\rightarrow65.49, and mixed 13.85→29.4013.85\rightarrow29.40. For MiniCPM-V, the corresponding changes are 17.24→21.2817.24\rightarrow21.28, 16.61→21.4816.61\rightarrow21.48, 4.73→5.774.73\rightarrow5.77, 27.23→36.5127.23\rightarrow36.51, and 12.32→15.1912.32\rightarrow15.19. The consistent gains across heterogeneous and multi-object query compositions support the claim that prompt-conditioned masking can select relevant regions when a query contains several objects.

  10. Knowl 10 — Improved Open-Ended Generation and General Perception

    empirical result

    AGLA improves both hallucination control and descriptive quality beyond yes/no object probing. On CHAIR, lower CSC_S and CIC_I indicate fewer hallucinated sentences and objects, while higher Recall indicates more complete captions. For LLaVA-1.5, regular decoding gives (CS,CI,Recall)=(51.0,15.2,75.2)(C_S,C_I,\mathrm{Recall})=(51.0,15.2,75.2) and AGLA gives (43.0,14.1,78.9)(43.0,14.1,78.9). For InstructBLIP, the corresponding values are (54.0,18.1,71.1)(54.0,18.1,71.1) and (49.0,12.1,72.5)(49.0,12.1,72.5). On LLaVA-Bench-Wild, GPT-4 scores LLaVA-1.5 regular decoding at accuracy/detail (2.61,3.65)(2.61,3.65) and AGLA at (3.83,4.39)(3.83,4.39); for InstructBLIP, the scores increase from (2.82,3.36)(2.82,3.36) to (4.59,4.59)(4.59,4.59). On the MME hallucination subset, AGLA also outperforms regular decoding and the compared decoding methods across existence, count, position, and color for both LLaVA-1.5 and InstructBLIP.

  11. Knowl 11 — Ablation Evidence for AGLA Components

    data/table

    On POPE-COCO under the popular setting with LLaVA-1.5, the complete AGLA obtains accuracy/F1 of 86.12/84.7186.12/84.71, compared with 81.88/80.0681.88/80.06 for regular decoding. Removing individual components reduces performance: removing adaptive plausibility truncation gives 85.66/84.4285.66/84.42, replacing adaptive masking with a fixed ratio gives 84.83/82.9484.83/82.94, and removing assembly so that only the augmented image is used gives 83.53/82.1483.53/82.14. The masking-strategy comparison gives Patch 85.77/84.0585.77/84.05, Soft 85.51/83.9385.51/83.93, Feature 85.27/83.7385.27/83.73, and Random 83.56/82.4083.56/82.40. These results indicate that original-image global information, adaptive masking, plausibility truncation, and the particular patch-masking design all contribute; learned image-prompt matching is also better than random masking.

  12. Knowl 12 — Robustness to Decoding Strategy

    empirical result

    AGLA improves hallucination metrics across multiple decoding strategies on POPE-COCO under the adversarial setting with LLaVA-1.5. Accuracy/F1 for regular decoding versus AGLA are: Top-pp sampling with p=0.7p=0.7, 80.13/78.8180.13/78.81 versus 84.23/83.0484.23/83.04; Top-kk sampling with k=50k=50, 78.50/77.2378.50/77.23 versus 83.80/82.5983.80/82.59; temperature sampling with temperature 0.50.5, 81.77/80.4081.77/80.40 versus 84.33/83.1584.33/83.15; Top-pp plus temperature, 81.50/80.1481.50/80.14 versus 84.50/83.3384.50/83.33; Top-kk plus temperature, 80.60/79.2480.60/79.24 versus 84.33/83.1584.33/83.15; and greedy decoding, 83.63/82.3383.63/82.33 versus 85.00/83.8785.00/83.87. Thus, the benefit is not tied to multinomial sampling alone.

Coverage note — Appendix-only hyperparameter sweeps, additional model-specific results, and qualitative examples were omitted because the core method, diagnosis, ablations, and principal benchmark outcomes are represented here.

References

  1. 1.David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski. A learning algorithm for boltzmann machines. Cognitive science, 9(1):147–169, 1985. 8
  2. 2.Vedika Agarwal, Rakshith Shetty, and Mario Fritz. Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9690–9698, 2020. 1, 2, 3
  3. 3.Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. Analyzing the behavior of visual question answering models. arXiv preprint arXiv:1606.07356, 2016. 1, 3
  4. 4.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022. 1
  5. 5.Wenbin An, Feng Tian, Jiahao Nie, Wenkai Shi, Haonan Lin, Yan Chen, QianYing Wang, Yaqiang Wu, Guang Dai, and Ping Chen. Knowledge acquisition disentanglement for knowledge-based visual question answering with large language models. arXiv preprint arXiv:2407.15346, 2024. 1
  6. 6.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 1, 2
  7. 7.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023. 1, 2, 5, 3
  8. 8.Ali Furkan Biten, Lluís Gomez, and Dimosthenis Karatzas. Let there be a clock on the beach: Reducing object hallucination in image captioning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1381–1390, 2022. 3
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 1, 2, 3
  10. 10.Xuweiyi Chen, Ziqiao Ma, Xuejun Zhang, Sihan Xu, Shengyi Qian, Jianing Yang, David Fouhey, and Joyce Chai. Multi-object hallucination in vision language models. In 3rd Workshop on Advances in Language and Vision Research (ALVR), 2024. 5, 6, 1
  11. 11.Zhiyang Chen, Yousong Zhu, Yufei Zhan, Zhaowen Li, Chaoyang Zhao, Jinqiao Wang, and Ming Tang. Mitigating hallucination in visual language models with visual supervision. arXiv preprint arXiv:2311.16479, 2023. 3
  12. 12.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 1, 2
  13. 13.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 2
  14. 14.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. Dola: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883, 2023. 3, 5, 6, 7
  15. 15.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2306.04387, 2023. 1, 2, 5, 6, 7, 4
  16. 16.Ailin Deng, Zhirui Chen, and Bryan Hooi. Seeing is believing: Mitigating hallucination in large vision-language models via clip-guided decoding. arXiv preprint arXiv:2402.15300, 2024. 1
  17. 17.Ronald A DeVore and Vladimir N Temlyakov. Some remarks on greedy algorithms. Advances in computational Mathematics, 5(1):173–187, 1996. 8
  18. 18.Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833, 2018. 8
  19. 19.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023. 5, 1, 3
  20. 20.Fabrizio Gilardi, Meysam Alizadeh, and Mael Kubli. Chat-gpt outperforms crowd-workers for text-annotation tasks. arXiv preprint arXiv:2303.15056, 2023. 2
  21. 21.Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. Multimodal-gpt: A vision and language model for dialogue with humans. arXiv preprint arXiv:2305.04790, 2023. 1
  22. 22.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 1, 3
  23. 23.Anisha Gunjal, Jihan Yin, and Erhan Bas. Detecting and preventing hallucinations in large vision language models. arXiv preprint arXiv:2308.06394, 2023. 1, 3
  24. 24.Yudong Han, Liqiang Nie, Jianhua Yin, Jianlong Wu, and Yan Yan. Visual perturbation-aware collaborative learning for overcoming the language prior problem. arXiv preprint arXiv:2207.11850, 2022. 1, 3
  25. 25.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019. 4, 8
  26. 26.Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. Opera: Alleviating hallucination in multimodal large language models via over-trust penalty and retrospection-allocation. arXiv preprint arXiv:2311.17911, 2023. 1, 3, 5, 6, 7
  27. 27.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709, 2019. 5, 1
  28. 28.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023. 3
  29. 29.Leila Kabbai, Mehrez Abdellaoui, and Ali Douik. Image classification by combining local and global features. The Visual Computer, 35:679–693, 2019. 1
  30. 30.Jusung Lee, Sungguk Cha, Younghyun Lee, and Cheoljong Yang. Visual question answering instruction: Unlocking multimodal large language model to domain-specific visual multitasks. arXiv preprint arXiv:2402.08360, 2024. 1, 2
  31. 31.Seongyun Lee, Sue Hyun Park, Yongrae Jo, and Minjoon Seo. Volcano: mitigating multimodal hallucination through self-feedback guided revision. arXiv preprint arXiv:2311.07362, 2023. 1, 3
  32. 32.Sicong Leng, Yun Xing, Zesen Cheng, Yang Zhou, Hang Zhang, Xin Li, Deli Zhao, Shijian Lu, Chunyan Miao, and Lidong Bing. The curse of multi-modalities: Evaluating hallucinations of large multimodal models across language, visual, and audio. arXiv preprint arXiv:2410.12787, 2024. 3
  33. 33.Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13872–13882, 2024. 1, 2, 3, 4, 5, 6, 7, 8
  34. 34.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023. 1, 2
  35. 35.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR, 2022. 2, 6
  36. 36.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 1, 2
  37. 37.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, et al. A large-scale dataset towards multi-modal multilingual instruction tuning. arXiv preprint arXiv:2306.04387, 2023. 3
  38. 38.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019. 2
  39. 39.Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. Contrastive decoding: Open-ended text generation as optimization. arXiv preprint arXiv:2210.15097, 2022. 3, 4
  40. 40.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 1, 2, 3, 4, 5, 6
  41. 41.Haonan Lin, Mengmeng Wang, Yan Chen, Wenbin An, Yuzhe Yao, Guang Dai, Qianying Wang, Yong Liu, and Jingdong Wang. Dreamsalon: A staged diffusion framework for preserving identity-context in editable face generation. arXiv preprint arXiv:2403.19235, 2024. 2
  42. 42.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014. 5, 1
  43. 43.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023. 3
  44. 44.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. Mitigating hallucination in large multi-modal models via robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023. 1, 3
  45. 45.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 1, 2, 4, 5, 6, 7, 8
  46. 46.Hanchao Liu, Wenyuan Xue, Yifei Chen, Dapeng Chen, Xutian Zhao, Ke Wang, Liping Hou, Rongjun Li, and Wei Peng. A survey on hallucination in large vision-language models. arXiv preprint arXiv:2402.00253, 2024. 1, 3
  47. 47.Holy Lovenia, Wenliang Dai, Samuel Cahyawijaya, Ziwei Ji, and Pascale Fung. Negative object presence evaluation (nope) to measure object hallucination in vision-language models. arXiv preprint arXiv:2310.05338, 2023. 1, 3
  48. 48.Jiahao Nie, Gongjie Zhang, Wenbin An, Yap-Peng Tan, Alex C Kot, and Shijian Lu. Mmrel: A relation understanding dataset and benchmark in the mllm era. arXiv preprint arXiv:2406.09121, 2024. 3
  49. 49.Sean O’Brien and Mike Lewis. Contrastive decoding improves reasoning in large language models. arXiv preprint arXiv:2309.09117, 2023. 3
  50. 50.Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim, and Sangdoo Yun. What do self-supervised vision transformers learn? arXiv preprint arXiv:2305.00729, 2023. 2
  51. 51.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020. 2
  52. 52.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. arXiv preprint arXiv:1809.02156, 2018. 3, 5, 7, 1, 6
  53. 53.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-okvqa: A benchmark for visual question answering using world knowledge. In European Conference on Computer Vision, pages 146–162. Springer, 2022. 5, 1
  54. 54.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017. 2, 3
  55. 55.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652, 2023. 3
  56. 56.Dong Shu, Haiyan Zhao, Jingyu Hu, Weiru Liu, Lu Cheng, and Mengnan Du. Large vision-language model alignment and misalignment: A survey through the lens of explainability. arXiv preprint arXiv:2501.01346, 2025. 2
  57. 57.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7464–7473, 2019. 2
  58. 58.Shilin Sun, Wenbin An, Feng Tian, Fang Nan, Qidong Liu, Jun Liu, Nazaraf Shah, and Ping Chen. A review of multimodal explainable artificial intelligence: Past, present and future. arXiv preprint arXiv:2412.14056, 2024. 2
  59. 59.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525, 2023. 3
  60. 60.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model, 2023. 1, 2
  61. 61.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et al. Ul2: Unifying language learning paradigms. In The Eleventh International Conference on Learning Representations, 2022. 2
  62. 62.Anthony Meng Huat Tiong, Junnan Li, Boyang Li, Silvio Savarese, and Steven CH Hoi. Plug-and-play vqa: Zero-shot vqa by conjoining large pretrained models with zero training. arXiv preprint arXiv:2210.08773, 2022. 2, 3
  63. 63.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 1, 2
  64. 64.Haibo Wang, Chenghang Lai, Yixuan Sun, and Weifeng Ge. Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering. arXiv preprint arXiv:2401.10711, 2024. 1, 2
  65. 65.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022. 2
  66. 66.Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al. Evaluation and analysis of hallucination in large vision-language models. arXiv preprint arXiv:2308.15126, 2023. 3
  67. 67.Teng Wang, Jinrui Zhang, Junjie Fei, Yixiao Ge, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao, Ying Shan, et al. Caption anything: Interactive image description with diverse multimodal controls. arXiv preprint arXiv:2305.02677, 2023. 1, 2
  68. 68.Junfei Wu, Qiang Liu, Ding Wang, Jinghao Zhang, Shu Wu, Liang Wang, and Tieniu Tan. Logical closed loop: Uncovering object hallucinations in large vision-language models. arXiv preprint arXiv:2402.11622, 2024. 3
  69. 69.Yike Wu, Yu Zhao, Shiwan Zhao, Ying Zhang, Xiaojie Yuan, Guoqing Zhao, and Ning Jiang. Overcoming language priors in visual question answering via distinguishing superficially similar instances. arXiv preprint arXiv:2209.08529, 2022. 1, 3
  70. 70.Yun Xing, Yiheng Li, Ivan Laptev, and Shijian Lu. Mitigating object hallucination via concentric causal attention. Advances in Neural Information Processing Systems, 37:92012–92035, 2024. 3
  71. 71.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024. 5, 6, 2, 3
  72. 72.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023. 1, 2
  73. 73.Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen. Woodpecker: Hallucination correction for multimodal large language models. arXiv preprint arXiv:2310.16045, 2023. 3
  74. 74.Zihao Yue, Liang Zhang, and Qin Jin. Less is more: Mitigating multimodal hallucination from an eos decision perspective. arXiv preprint arXiv:2402.14545, 2024. 1, 3, 5
  75. 75.Xiaofeng Zhang, Yihao Quan, Chaochen Gu, Chen Shen, Xiaosong Yuan, Shaotian Yan, Hao Cheng, Kaijie Wu, and Jieping Ye. Seeing clearly by layer two: Enhancing attention heads to alleviate hallucination in lvlms. arXiv preprint arXiv:2411.09968, 2024. 3
  76. 76.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023. 3
  77. 77.Ren Zhibo, Wang Huizhen, Zhu Muhua, Wang Yichao, Xiao Tong, and Zhu Jingbo. Overcoming language priors with counterfactual inference for visual question answering. In Proceedings of the 22nd Chinese National Conference on Computational Linguistics, pages 600–610, 2023. 1, 3
  78. 78.Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. Analyzing and mitigating object hallucination in large vision-language models. arXiv preprint arXiv:2310.00754, 2023. 3
  79. 79.Yiyang Zhou, Chenhang Cui, Rafael Rafailov, Chelsea Finn, and Huaxiu Yao. Aligning modalities in vision large language models via preference fine-tuning. arXiv preprint arXiv:2402.11411, 2024. 3
  80. 80.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 1, 2
  81. 81.Lanyun Zhu, Deyi Ji, Tianrun Chen, Peng Xu, Jieping Ye, and Jun Liu. Ibd: Alleviating hallucinations in large vision-language models via image-biased decoding. arXiv preprint arXiv:2402.18476, 2024. 1, 3, 4

Citation

MLA
An, W., et al. “Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 29915–26, https://doi.org/10.1109/CVPR52734.2025.02784.
APA
An, W., Tian, F., Leng, S., Nie, J., Lin, H., Wang, Q., Chen, P., Zhang, X., & Lu, S. (2025). Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 29915–29926. https://doi.org/10.1109/CVPR52734.2025.02784
Chicago
An, W., F. Tian, S. Leng, et al. 2025. “Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 29915–26. https://doi.org/10.1109/CVPR52734.2025.02784.
Harvard
An, W. et al. (2025) “Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 29915–29926. Available at: https://doi.org/10.1109/CVPR52734.2025.02784.
Vancouver
1. An W, Tian F, Leng S, Nie J, Lin H, Wang Q, Chen P, Zhang X, Lu S (2025) Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 29915–29926

BibTeX

@inproceedings{An_2025, title={Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention}, url={http://dx.doi.org/10.1109/CVPR52734.2025.02784}, DOI={10.1109/cvpr52734.2025.02784}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={An, Wenbin and Tian, Feng and Leng, Sicong and Nie, Jiahao and Lin, Haonan and Wang, Qianying and Chen, Ping and Zhang, Xiaoqin and Lu, Shijian}, year={2025}, month=June, pages={29915–29926} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE