Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
Seongyun LeeSue Hyun ParkYongrae JoMinjoon Seo
Introduces VOLCANO, a unified multimodal model that reduces visual hallucinations by generating self-directed natural language feedback to iteratively critique and correct its own initial answers without requiring external reward models or auxiliary detectors.
Large multimodal models (LMMs), which process and generate responses from both text and visual inputs, frequently suffer from multimodal hallucination—generating statements misaligned with or unsupported by the provided image. This issue poses significant operational and reliability risks as organizations deploy multimodal artificial intelligence for real-world tasks. The article demonstrates that these errors often occur because models fail to ground their answers in actual image features, instead relying too heavily on language patterns and internal text knowledge. To solve this, the article evaluates whether a single model can reduce hallucinations by critiquing and revising its own answers using natural language visual feedback.
The researchers developed VOLCANO, a self-feedback guided revision model. The approach uses an iterative critique-revise-decide loop managed by a single unified model rather than relying on external reward models or complex multi-model pipelines. To train VOLCANO, initial responses from an open-source model were evaluated by a proprietary large language model provided with detailed textual image descriptions to generate feedback and gold revisions. The final model generates an initial answer, produces natural language feedback referencing the image, revises the response accordingly, and decides whether the revision is better than the original, repeating this process for up to three iterations.
The article demonstrates five key findings. First, VOLCANO achieves state-of-the-art results across standard multimodal hallucination benchmarks, outperforming baseline models and reducing hallucination rates (for instance, dropping the hallucination rate to 0.48 on MMHal-Bench). Second, it delivers an approximate 24.9% improvement over prior hallucination-mitigation methods like LURE and Woodpecker while remaining a single, end-to-end model. Third, natural language feedback outperforms scalar reinforcement learning feedback (such as LLaVA-RLHF), indicating that direct descriptive critiques guide corrections more effectively. Fourth, reducing hallucinations enhances general visual reasoning, with VOLCANO improving overall scores on multimodal benchmarks like MM-Vet and MMBench, including doubling math capability scores relative to its base model. Finally, attention analysis confirms that during feedback generation, the model distributes visual attention across a wider and more detailed area of the image than it does during initial generation, allowing it to recover missed visual details.
These findings indicate that multimodal models possess latent capacity to recognize and self-correct errors if prompted through a structured feedback step. For enterprise deployment, this approach increases output reliability, reduces compliance and safety risks associated with fabricated visual content, and avoids the computational overhead of training specialized external reward models. However, organizations face a clear latency trade-off: because VOLCANO executes multiple sequential generation steps, response generation takes approximately two to three times longer (5.8 seconds versus 2.7 seconds for standard generation), which may affect real-time applications.
Decision-makers considering multimodal AI deployments should evaluate self-feedback revision architectures where accuracy and safety outweigh raw response speed. Future technical work should focus on improving the runtime efficiency of iterative self-refinement to reduce latency. While results are highly consistent across standard academic benchmarks, practitioners should maintain moderate caution when applying these models to novel, out-of-distribution real-world images until further domain-specific pilot testing is conducted.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). This paper establishes the POPE benchmark and foundational evaluation paradigms for object hallucination in large vision-language models that VOLCANO directly targets and measures.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). It introduces the MM-Vet integrated evaluation suite, which serves as a core benchmark for measuring general multimodal reasoning improvements in VOLCANO.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This comprehensive survey categorizes the architectures and training pipelines of multimodal large language models while defining the core challenge of multimodal hallucination.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). It defines standard perception and cognition benchmarks for multimodal models, providing the baseline testing protocols that contextualize VOLCANO's empirical evaluations.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). This survey establishes the principles, taxonomy, and feedback-based mitigation foundations of language model hallucination that VOLCANO adapts to the vision-language domain.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). It provides foundational evidence that generative models can self-evaluate and identify errors in their own outputs, motivating VOLCANO's self-critique mechanism.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). This work tackles multimodal hallucination by analyzing language prior decay and modifying visual information grounding during decoding, offering a complementary alternative to VOLCANO's iterative post-hoc revision.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It scales multimodal training and test-time reasoning strategies across open-source multimodal models, extending post-training alignment techniques evaluated in VOLCANO.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). This research explores test-time and architectural scaling in open-source multimodal models, advancing the visual reasoning capabilities that iterative self-refinement systems aim to improve.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). It addresses fine-grained multimodal errors by introducing dynamic focal zooming and chain-of-thought grounding, building on VOLCANO's insights into visual attention redistribution during reasoning.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). This work unifies single-image, multi-image, and video representations in multimodal models, extending the foundation upon which multimodal self-correction methods operate.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). It introduces internal circuit probing to detect model errors without multi-step text generation, providing an efficient alternative mechanism to VOLCANO's natural language self-critique loop.
