Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Seongyun LeeSue Hyun ParkYongrae JoMinjoon Seo

article2024NAACL97 citations

Introduces VOLCANO, a unified multimodal model that reduces visual hallucinations by generating self-directed natural language feedback to iteratively critique and correct its own initial answers without requiring external reward models or auxiliary detectors.

Listen

Large multimodal models (LMMs), which process and generate responses from both text and visual inputs, frequently suffer from multimodal hallucination—generating statements misaligned with or unsupported by the provided image. This issue poses significant operational and reliability risks as organizations deploy multimodal artificial intelligence for real-world tasks. The article demonstrates that these errors often occur because models fail to ground their answers in actual image features, instead relying too heavily on language patterns and internal text knowledge. To solve this, the article evaluates whether a single model can reduce hallucinations by critiquing and revising its own answers using natural language visual feedback.

The researchers developed VOLCANO, a self-feedback guided revision model. The approach uses an iterative critique-revise-decide loop managed by a single unified model rather than relying on external reward models or complex multi-model pipelines. To train VOLCANO, initial responses from an open-source model were evaluated by a proprietary large language model provided with detailed textual image descriptions to generate feedback and gold revisions. The final model generates an initial answer, produces natural language feedback referencing the image, revises the response accordingly, and decides whether the revision is better than the original, repeating this process for up to three iterations.

The article demonstrates five key findings. First, VOLCANO achieves state-of-the-art results across standard multimodal hallucination benchmarks, outperforming baseline models and reducing hallucination rates (for instance, dropping the hallucination rate to 0.48 on MMHal-Bench). Second, it delivers an approximate 24.9% improvement over prior hallucination-mitigation methods like LURE and Woodpecker while remaining a single, end-to-end model. Third, natural language feedback outperforms scalar reinforcement learning feedback (such as LLaVA-RLHF), indicating that direct descriptive critiques guide corrections more effectively. Fourth, reducing hallucinations enhances general visual reasoning, with VOLCANO improving overall scores on multimodal benchmarks like MM-Vet and MMBench, including doubling math capability scores relative to its base model. Finally, attention analysis confirms that during feedback generation, the model distributes visual attention across a wider and more detailed area of the image than it does during initial generation, allowing it to recover missed visual details.

These findings indicate that multimodal models possess latent capacity to recognize and self-correct errors if prompted through a structured feedback step. For enterprise deployment, this approach increases output reliability, reduces compliance and safety risks associated with fabricated visual content, and avoids the computational overhead of training specialized external reward models. However, organizations face a clear latency trade-off: because VOLCANO executes multiple sequential generation steps, response generation takes approximately two to three times longer (5.8 seconds versus 2.7 seconds for standard generation), which may affect real-time applications.

Decision-makers considering multimodal AI deployments should evaluate self-feedback revision architectures where accuracy and safety outweigh raw response speed. Future technical work should focus on improving the runtime efficiency of iterative self-refinement to reduce latency. While results are highly consistent across standard academic benchmarks, practitioners should maintain moderate caution when applying these models to novel, out-of-distribution real-world images until further domain-specific pilot testing is conducted.

Cover for Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Abstract

Large multimodal models suffer from multimodal hallucination, where they provide incorrect responses misaligned with the given visual information. Recent works have conjectured that one of the reasons behind multimodal hallucination is due to the vision encoder failing to ground on the image properly. To mitigate this issue, we propose a novel approach that leverages self-feedback as visual cues. Building on this approach, we introduce VOLCANO, a multimodal self-feedback guided revision model. VOLCANO generates natural language feedback to its initial response based on the provided visual information and utilizes this feedback to self-revise its initial response. VOLCANO effectively reduces multimodal hallucination and achieves state-of-the-art on MMHal-Bench, POPE, and GAVIE. It also improves on general multimodal abilities and outperforms previous models on MM-Vet and MMBench. Through qualitative analysis, we show that VOLCANO’s feedback is properly grounded on the image than the initial response. This indicates that VOLCANO can provide itself with richer visual information through feedback generation, leading to self-correct hallucinations. We publicly release our model, data, and code at github.com/kaistAI/Volcano.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Multimodal hallucination
  • 2.2 Self-correcting from feedback
  • 2.3 Mitigating multimodal hallucination
  • 3 VOLCANO
  • 3.1 Iterative self-revision
  • 3.2 Data collection
  • 3.3 Implementation details
  • 4 Experiments
  • 4.1 Benchmarks
  • 4.2 Baselines
  • 4.3 Main results
  • 4.4 Ablation studies
  • 5 Qualitative analysis
  • 5.1 Amount of visual information
  • 5.2 Coverage of visual information
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Full results on benchmarks
  • A.1 Multimodal hallucination benchmarks
  • A.2 Multimodal understanding benchmarks
  • B Prompts
  • B.1 Prompts for inference at each stage
  • B.2 Prompt for generating multimodal feedback
  • C Computation
  • D Hyperparameters

Knowls

  1. Knowl 1 — VOLCANO Framework for Multimodal Self-Feedback Guided Revision

    model/method

    VOLCANO is a single large multimodal model (LMM) designed to mitigate multimodal hallucination—where generated text misaligns with visual content—without requiring external reward models or separate task-specific modules.

    Instead of single-pass generation or external post-hoc correction, VOLCANO operates through an iterative critique-revise-decide framework. The single LMM performs four sub-tasks across its inference pipeline:

    1. Generating an initial response to an input image and question.
    2. Generating natural language self-feedback (critique) by re-examining the image features and identifying specific factual discrepancies or unsupported statements in the response.
    3. Generating a revised response conditioned on the image, question, prior response, and self-generated feedback.
    4. Performing a pairwise comparative decision between the candidate revised response and the prior best response to select the more accurate output and determine whether to terminate or continue iteration.
  2. Knowl 2 — Iterative Critique-Revise-Decide Inference Algorithm

    algorithm

    VOLCANO refines its outputs via an iterative loop bounded by a maximum iteration count Nmax=3N_{\text{max}} = 3. At each step, natural language feedback guides the revision, followed by a randomized pairwise choice to prevent over-correction.

    Input: Multimodal model MM, image II, question QQ, maximum iterations Nmax=3N_{\text{max}} = 3
    Output: Final revised response RbestR_{\text{best}}
    Rinitial←M(I,Q)R_{\text{initial}} \leftarrow M(I, Q)
    Rbest←RinitialR_{\text{best}} \leftarrow R_{\text{initial}}
    for t←1t \leftarrow 1 to NmaxN_{\text{max}} do
        F←M(I,Q,Rbest)F \leftarrow M(I, Q, R_{\text{best}})
        Rrevised←M(I,Q,Rbest,F)R_{\text{revised}} \leftarrow M(I, Q, R_{\text{best}}, F)
        Rdecided←M(I,Q,random_order(Rbest,Rrevised))R_{\text{decided}} \leftarrow M(I, Q, \text{random\_order}(R_{\text{best}}, R_{\text{revised}}))
        if Rdecided==RbestR_{\text{decided}} == R_{\text{best}} then
            break
        else
            Rbest←RrevisedR_{\text{best}} \leftarrow R_{\text{revised}}
    return RbestR_{\text{best}}

    The decision stage randomizes the presentation order of RbestR_{\text{best}} and RrevisedR_{\text{revised}} as option A and option B to eliminate position bias. If the model chooses the existing best response, the iterative refinement terminates immediately.

  3. Knowl 3 — Multimodal Feedback and Revision Data Synthesis via Textual Image Proxies

    model/method

    To train a single LMM to critique and revise without human annotations, training data is synthesized using a proprietary text LLM (such as GPT-3.5-turbo) and existing visual instruction data (the first turn of LLaVA-SFT-127k):

    1. Initial Response Generation: Candidate initial responses are produced by an open-source base LMM (LLaVA-SFT+ 7B) on visual question-answering instances.
    2. Image Proxy Construction: Because text-only proprietary LLMs do not ingest raw pixels, visual content is represented as text containing detailed object lists and reference gold captions. If object annotations are absent, only the reference caption is supplied to avoid error propagation from external detectors.
    3. Feedback Generation Prompting: The proprietary LLM receives the textual image proxy, question, candidate response, and reference gold answer. The prompt instructs the model to evaluate the candidate against the image description, pinpoint errors, explicitly treat synonyms and paraphrases as correct, and output a detailed critique explaining why the response is imperfect and how it should be modified.
    4. Revision Pair Formatting: Revision training instances are assembled by setting the input to (I,Q,Rinitial,F)(I, Q, R_{\text{initial}}, F) and assigning the dataset's ground-truth gold answer directly as the target revision output without requiring additional model generation.
  4. Knowl 4 — Training Configuration and Implementation Details of VOLCANO

    experimental setup

    VOLCANO uses LLaVA-1.5 (7B and 13B architectures) as its backbone, utilizing a CLIP ViT-L/14 vision encoder operating at 336×\times336 resolution (yielding 576 patch tokens of size 14×\times14).

    During instruction fine-tuning, the synthesized multimodal feedback and revision instances are merged with the standard llava-1.5-mix665k instruction dataset.

    Training hyperparameters:

    • Batch size: 128
    • Learning rate: 2×10−52 \times 10^{-5} with cosine schedule and a 0.03 warmup ratio
    • Epochs: 1
    • Maximum sequence length: 2048 tokens
    • Weight decay: 0
    • Optimization: DeepSpeed ZeRO-3 with gradient checkpointing
    • Inference: Greedy decoding
    • Hardware and compute time: 8 NVIDIA A100-SXM4-80GB GPUs (15 hours for VOLCANO 7B; 30 hours for VOLCANO 13B).
  5. Knowl 5 — Multimodal Hallucination Mitigation Results on MMHal-Bench, POPE, and GAVIE

    empirical result

    VOLCANO (7B and 13B) achieves superior hallucination mitigation compared to base models (LLaVA-1.5), RLHF-trained models (LLaVA-RLHF), and specialized external correctors (LURE, Woodpecker) across MMHal-Bench (GPT-4 score 0–5; hallucination rate is the fraction of responses scored <3< 3), POPE (accuracy and F1 over object probing queries), and GAVIE (accuracy and relevancy scored 0–10).

    Model MMHal-Bench POPE GAVIE
    Score ↑\uparrow Hal rate ↓\downarrow Acc ↑\uparrow F1 ↑\uparrow Acc ↑\uparrow Rel ↑\uparrow Avg ↑\uparrow
    MiniGPT-4 7B - - 68.4 74.5 4.14 5.81 4.98
    InstructBLIP 7B 2.10 0.58 71.5 80.0 5.93 7.34 6.64
    LLaVA-SFT+ 7B 1.76 0.67 81.6 82.7 5.95 8.16 7.06
    LLaVA-RLHF 7B 2.05 0.68 81.8 81.5 6.01 8.11 7.06
    LLaVA-1.5 7B 2.42 0.55 86.1 85.1 6.42 8.20 7.31
    VOLCANO 7B 2.60 0.49 88.2 87.7 6.52 8.40 7.46
    LLaVA-SFT+ 13B 2.43 0.55 83.2 82.8 5.95 8.20 7.09
    LLaVA-RLHF 13B 2.53 0.57 83.1 81.9 6.46 8.22 7.34
    LLaVA-1.5 13B 2.54 0.52 86.2 85.2 6.80 8.47 7.64
    VOLCANO 13B 2.64 0.48 88.3 87.7 6.94 8.72 7.83

    When controlling for model backbone using LLaVA-SFT+ 7B, fine-tuning with VOLCANO's feedback data (yielding VOLCANO−\text{VOLCANO}^- 7B) achieves a score of 2.19 and hallucination rate of 0.59 on MMHal-Bench, outperforming LLaVA-RLHF 7B (score 2.05, hal rate 0.68), LURE (score 1.90, hal rate 0.58), and Woodpecker (score 1.98, hal rate 0.54).

  6. Knowl 6 — General Multimodal Understanding Performance on MM-Vet and MMBench

    empirical result

    Mitigating hallucination through self-feedback does not degrade general visual comprehension; instead, VOLCANO enhances overall perception and reasoning capabilities on MM-Vet (evaluating 6 core competencies: recognition, OCR, knowledge, language generation, spatial awareness, and math) and MMBench (evaluating level-2 perception and reasoning dimensions).

    Model MMBench Acc (%) ↑\uparrow MM-Vet Acc (%) ↑\uparrow
    InstructBLIP 14B 36.0 25.6
    LLaVA-SFT+ 7B 52.7 30.4
    LLaVA-RLHF 7B 52.7 29.8
    LLaVA-1.5 7B 59.9 31.2
    VOLCANO 7B 62.3 32.0
    LLaVA-SFT+ 13B 59.6 36.1
    LLaVA-RLHF 13B 59.6 36.4
    LLaVA-1.5 13B 67.7 36.1
    VOLCANO 13B 69.4 38.0

    On MM-Vet sub-tasks, VOLCANO 13B achieves a math capability score of 15.0, nearly doubling the 7.7 score achieved by LLaVA-1.5 13B, alongside gains across OCR (30.4 vs 28.0) and language generation (29.2 vs 24.4).

  7. Knowl 7 — Ablation of Model Components and Self-Revision Iterations

    empirical result

    Ablations on MMHal-Bench isolate the contributions of individual pipeline stages and iteration counts for VOLCANO 7B:

    1. Pipeline Stage Ablation:

      • Only prediction (generating only the initial response using the fine-tuned VOLCANO weights, bypassing critique, revision, and decision): achieves a score of 2.45 and a hallucination rate of 0.52. This outperforms the base LLaVA-1.5 7B (2.42 score, 0.55 hal rate), indicating that training on feedback/revision data improves base generation.
      • No decision (executing critique and revision, but outputting the revised text directly without the decision stage): performance degrades to a score of 2.33 and a hallucination rate of 0.56. This demonstrates that the decision stage is vital to prevent over-correction and erroneous revisions.
      • Full VOLCANO 7B (critique, revision, and decision): achieves a score of 2.60 and a hallucination rate of 0.49.
    2. Iteration Count Trajectory:

      • Iteration 1: Score 2.54, Hal rate 0.51
      • Iteration 2: Score 2.58, Hal rate 0.50
      • Iteration 3: Score 2.60, Hal rate 0.49

    Each progressive refinement step monotonically decreases the hallucination rate and increases the overall accuracy score.

  8. Knowl 8 — Visual Attention Concentration and Feature Coverage During Feedback Generation

    empirical result

    Attention weight analysis reveals why self-feedback corrects hallucinated initial outputs. Attention weights are aggregated across visual tokens using top-kk mean pooling: averaging the top-3 weights across transformer layers, top-3 weights across self-attention heads, and top-ll weights across output tokens (where ll is the length of the shorter output between the response and feedback).

    Key findings from empirical visualization:

    1. Attention Magnitude: Despite having a longer prompt containing both the question and the initial response, the model allocates higher average attention weight intensity to visual patch features during feedback generation than during initial response generation.
    2. Feature Coverage: In instances where initial responses hallucinate fine-grained attributes (e.g., misidentifying a silver pot filled with red berries as a red pot), initial generation attention is restricted to outer boundary edges. During feedback generation, attention spans both the outer object contours and internal regions, with descriptor tokens ('silver', 'red') directly attending to the corresponding image regions.
  9. Knowl 9 — Inference Latency Overhead of Iterative Revision

    limitation

    Because VOLCANO executes multiple sequential autoregressive generation calls per instance (initial response, critique generation, revised generation, and comparative decision evaluation over up to 3 iterations), its inference latency is substantially higher than standard single-pass LMMs.

    On average, VOLCANO is 2 to 3 times slower than LLaVA-1.5, requiring 5.8 seconds per question-image pair compared to 2.7 seconds for a single LLaVA-1.5 forward generation.

Coverage note — None was omitted; all key contributions, algorithms, data synthesis methods, benchmark evaluations, ablations, attention analysis, and limitations are fully covered.

References

  1. 1.Afra Feyza Akyürek, Ekin Akyürek, Aman Madaan, Ashwin Kalyan, Peter Clark, Derry Wijaya, and Niket Tandon. 2023. Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a visual language model for few-shot learning.
  3. 3.Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023. Openflamingo: An open-source framework for training large autoregressive vision-language models.
  4. 4.Ali Furkan Biten, Lluís Gómez, and Dimosthenis Karatzas. 2022. Let there be a clock on the beach: Reducing object hallucination in image captioning. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2473–2482.
  5. 5.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023. Shikra: Unleashing multimodal llm’s referential dialogue magic.
  6. 6.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning.
  7. 7.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacafarm: A simulation framework for methods that learn from human feedback.
  8. 8.Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024. CRITIC: Large language models can self-correct with tool-interactive critiquing. In The Twelfth International Conference on Learning Representations.
  9. 9.Anisha Gunjal, Jihan Yin, and Erhan Bas. 2023. Detecting and preventing hallucinations in large vision language models.
  10. 10.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  11. 11.Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2023. Prometheus: Inducing fine-grained evaluation capability in language models.
  12. 12.Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback.
  13. 13.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023a. Otter: A multimodal model with in-context instruction tuning.
  14. 14.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.
  15. 15.Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023c. Halueval: A large-scale hallucination evaluation benchmark for large language models.
  16. 16.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023d. Evaluating object hallucination in large vision-language models.
  17. 17.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step.
  18. 18.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023a. Mitigating hallucination in large multi-modal models via robust instruction tuning.
  19. 19.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023b. Improved baselines with visual instruction tuning.
  20. 20.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023c. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  21. 21.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. 2023d. Grounding dino: Marrying dino with grounded pre-training for open-set object detection.
  22. 22.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2023e. Mmbench: Is your multi-modal model an all-around player?
  23. 23.Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Eureka: Human-level reward design via coding large language models.
  24. 24.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023. Video-chatgpt: Towards detailed video understanding via large vision and language models.
  25. 25.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback.
  26. 26.OpenAI. 2022. Chatgpt: Optimizing language models for dialogue.
  27. 27.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  28. 28.Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2023. Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies.
  29. 29.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world.
  30. 30.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045, Brussels, Belgium. Association for Computational Linguistics.
  31. 31.Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2022. Training language models with language feedback.
  32. 32.Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning.
  33. 33.Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023. Pandagpt: One model to instruction-follow them all.
  34. 34.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, and Trevor Darrell. 2023. Aligning large multimodal models with factually augmented rlhf.
  35. 35.Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, and Conghui He. 2023a. Vigc: Visual instruction generation and correction.
  36. 36.Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, Jitao Sang, and Haoyu Tang. 2023b. Evaluation and analysis of hallucination in large vision-language models.
  37. 37.Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023c. Large language models are not fair evaluators.
  38. 38.Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O’Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023d. Shepherd: A critic for language model generation.
  39. 39.Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2022. Generating sequences by learning to self-correct.
  40. 40.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-grained human feedback gives better rewards for language model training.
  41. 41.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023a. mplug-owl: Modularization empowers large language models with multimodality.
  42. 42.Seonghyeon Ye, Yongrae Jo, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, and Minjoon Seo. 2023b. Selfee: Iterative self-revising llm empowered by self-feedback generation. Blog post.
  43. 43.Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen. 2023. Woodpecker: Hallucination correction for multimodal large language models.
  44. 44.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities.
  45. 45.Bohan Zhai, Shijia Yang, Xiangchen Zhao, Chenfeng Xu, Sheng Shen, Dongdi Zhao, Kurt Keutzer, Manling Li, Tan Yan, and Xiangjun Fan. 2023. Halle-switch: Rethinking and controlling object existence hallucinations in large vision language models for detailed caption.
  46. 46.Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2023a. How language model hallucinations can snowball.
  47. 47.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023b. Llama-adapter: Efficient fine-tuning of language models with zero-init attention.
  48. 48.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023c. Siren’s song in the ai ocean: A survey on hallucination in large language models.
  49. 49.Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. 2023. Analyzing and mitigating object hallucination in large vision-language models.
  50. 50.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models.

Citation

MLA
Lee, S., et al. “Volcano: Mitigating Multimodal Hallucination Through Self-Feedback Guided Revision”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 391–404, https://doi.org/10.18653/v1/2024.naacl-long.23.
APA
Lee, S., Park, S. H., Jo, Y., & Seo, M. (2024). Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 391–404. https://doi.org/10.18653/v1/2024.naacl-long.23
Chicago
Lee, S., S. H. Park, Y. Jo, and M. Seo. 2024. “Volcano: Mitigating Multimodal Hallucination Through Self-Feedback Guided Revision”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 391–404. https://doi.org/10.18653/v1/2024.naacl-long.23.
Harvard
Lee, S. et al. (2024) “Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 391–404. Available at: https://doi.org/10.18653/v1/2024.naacl-long.23.
Vancouver
1. Lee S, Park SH, Jo Y, Seo M (2024) Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 391–404

BibTeX

@inproceedings{lee-etal-2024-volcano,
    title = "Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision",
    author = "Lee, Seongyun  and
      Park, Sue Hyun  and
      Jo, Yongrae  and
      Seo, Minjoon",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.23/",
    doi = "10.18653/v1/2024.naacl-long.23",
    pages = "391--404"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/