Multi-Modal Hallucination Control by Visual Information Grounding
Alessandro FaveroLuca ZancatoMatthew TragerSiddharth ChoudharyPramuditha PereraAlessandro AchilleAshwin SwaminathanStefano Soatto
Proposes a training-free decoding strategy and preference-tuning framework that mitigate vision-language model hallucinations by counteracting the progressive dilution of visual context during generation.
Modern vision-language artificial intelligence models, which process both images and text, often suffer from "hallucinations"—generating plausible-sounding descriptions of objects or scenes that do not actually exist in the input image. This ungrounded text creates significant reliability, safety, and compliance risks for real-world deployments. The article investigates why these hallucinations occur and demonstrates practical methods to reduce them by actively reinforcing visual evidence during text generation.
The article demonstrates that multi-modal hallucinations stem from a "fading memory effect," where models rely less on visual information and more on generic language expectations as they generate longer responses. To solve this, the authors introduce Multi-Modal Mutual-Information Decoding (M3ID), a lightweight, training-free method that amplifies the influence of the image relative to the text-only baseline during generation. They also evaluate a fine-tuning strategy combining M3ID with Direct Preference Optimization (DPO), which trains the model on self-generated grounded text pairs without requiring expensive human labels. The approaches were evaluated on the MS COCO benchmark using the LLaVA architecture across image captioning and visual question answering tasks.
The findings show that M3ID substantially improves factual accuracy while preserving fluent language generation. For the 13-billion-parameter LLaVA model, applying M3ID reduced the proportion of hallucinated objects in image captions by approximately 25% and reduced the share of captions containing at least one hallucination by 29%. On the POPE visual question answering benchmark, M3ID improved overall classification accuracy by about 21% while curtailing the baseline model's tendency to answer "Yes" to non-existent objects. When combined with preference optimization (M3ID+DPO), accuracy improved further, achieving a 28% reduction in hallucinated objects and a 24% overall accuracy boost on visual questions without human annotation costs.
These results indicate that multi-modal hallucinations are primarily caused by an over-reliance on language patterns rather than an inability to recognize image contents. For organizations deploying vision-language systems, M3ID offers an immediate, cost-effective intervention at inference time that does not require model retraining or labeled data curation. If developers have access to model weights and compute, pairing the decoding strategy with preference optimization provides even stronger grounding.
Organizations should consider adopting M3ID-style decoding for generation pipelines where factual visual accuracy is critical, while carefully tuning control parameters to prevent "overcompensation"—a state where the model omits highly obvious contextual objects. Future initiatives should evaluate this method across more diverse models, investigate structured captioning workflows, and address computational trade-offs, as M3ID requires two forward model passes per token generation step unless queries are batched.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Introduces the POPE evaluation benchmark for measuring object hallucination in vision-language models, which serves as a primary evaluation metric for the grounding techniques in the source.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Establishes the improved LLaVA architecture that the source directly uses as the foundational vision-language testbed for evaluating hallucination mitigation.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Provides a comprehensive architectural overview and analysis of multimodal large language models, including the mechanics of vision-text integration and the origins of multimodal hallucination.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Presents a foundational taxonomy and survey of hallucination mechanisms in generative models, establishing essential concepts underlying the language prior biases examined in the source.
- Paper: Generation and Comprehension of Unambiguous Object Descriptions, Junhua Mao et al. (2015). Introduces mutual information objectives for grounding visual descriptions against competing visual contexts, laying the conceptual foundation for the source's mutual information decoding strategy.
- Paper: Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision, Seongyun Lee et al. (2024). Builds upon multimodal hallucination reduction strategies by using natural language self-critique and visual feedback to iteratively revise ungrounded responses.
- Paper: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention, Wenbin An et al. (2025). Addresses visual grounding deficiencies causing object hallucinations by dynamically assembling prompt-relevant local and global visual attention features.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). Advances large-scale multimodal pre-training and test-time optimization recipes to strengthen cross-modal visual-textual grounding across complex tasks.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Demonstrates scaled multi-stage multimodal architectures that integrate deep visual token injection and test-time reasoning to mitigate representational loss and hallucination.
