Multi-Modal Hallucination Control by Visual Information Grounding

Alessandro FaveroLuca ZancatoMatthew TragerSiddharth ChoudharyPramuditha PereraAlessandro AchilleAshwin SwaminathanStefano Soatto

article2024CVPR189 citations

Proposes a training-free decoding strategy and preference-tuning framework that mitigate vision-language model hallucinations by counteracting the progressive dilution of visual context during generation.

Listen

Modern vision-language artificial intelligence models, which process both images and text, often suffer from "hallucinations"—generating plausible-sounding descriptions of objects or scenes that do not actually exist in the input image. This ungrounded text creates significant reliability, safety, and compliance risks for real-world deployments. The article investigates why these hallucinations occur and demonstrates practical methods to reduce them by actively reinforcing visual evidence during text generation.

The article demonstrates that multi-modal hallucinations stem from a "fading memory effect," where models rely less on visual information and more on generic language expectations as they generate longer responses. To solve this, the authors introduce Multi-Modal Mutual-Information Decoding (M3ID), a lightweight, training-free method that amplifies the influence of the image relative to the text-only baseline during generation. They also evaluate a fine-tuning strategy combining M3ID with Direct Preference Optimization (DPO), which trains the model on self-generated grounded text pairs without requiring expensive human labels. The approaches were evaluated on the MS COCO benchmark using the LLaVA architecture across image captioning and visual question answering tasks.

The findings show that M3ID substantially improves factual accuracy while preserving fluent language generation. For the 13-billion-parameter LLaVA model, applying M3ID reduced the proportion of hallucinated objects in image captions by approximately 25% and reduced the share of captions containing at least one hallucination by 29%. On the POPE visual question answering benchmark, M3ID improved overall classification accuracy by about 21% while curtailing the baseline model's tendency to answer "Yes" to non-existent objects. When combined with preference optimization (M3ID+DPO), accuracy improved further, achieving a 28% reduction in hallucinated objects and a 24% overall accuracy boost on visual questions without human annotation costs.

These results indicate that multi-modal hallucinations are primarily caused by an over-reliance on language patterns rather than an inability to recognize image contents. For organizations deploying vision-language systems, M3ID offers an immediate, cost-effective intervention at inference time that does not require model retraining or labeled data curation. If developers have access to model weights and compute, pairing the decoding strategy with preference optimization provides even stronger grounding.

Organizations should consider adopting M3ID-style decoding for generation pipelines where factual visual accuracy is critical, while carefully tuning control parameters to prevent "overcompensation"—a state where the model omits highly obvious contextual objects. Future initiatives should evaluate this method across more diverse models, investigate structured captioning workflows, and address computational trade-offs, as M3ID requires two forward model passes per token generation step unless queries are batched.

arXiv: 2403.14003
Cover for Multi-Modal Hallucination Control by Visual Information Grounding

Abstract

Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phenomenon, usually referred to as “hallucination” and show that it stems from an excessive reliance on the language prior. In particular, we show that as more tokens are generated, the reliance on the visual prompt decreases, and this behavior strongly correlates with the emergence of hallucinations. To reduce hallucinations, we introduce Multi-Modal Mutual-Information Decoding (M3ID), a new sampling method for prompt amplification. M3ID amplifies the influence of the reference image over the language prior, hence favoring the generation of tokens with higher mutual information with the visual prompt. M3ID can be applied to any pre-trained autoregressive VLM at inference time without necessitating further training and with minimal computational overhead. If training is an option, we show that M3ID can be paired with Direct Preference Optimization (DPO) to improve the model’s reliance on the prompt image without requiring any labels. Our empirical findings show that our algorithms maintain the fluency and linguistic capabilities of pre-trained VLMs while reducing hallucinations by mitigating visually ungrounded answers. Specifically, for the LLaVA 13B model, M3ID and M3ID+DPO reduce the percentage of hallucinated objects in captioning tasks by 25% and 28%, respectively, and improve the accuracy on VQA benchmarks such as POPE by 21% and 24%.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Analysis of hallucinations in VLMs
  • 4. Methods
  • 4.1. M3ID: Improving grounding at inference time
  • 4.2. M3ID+DPO to learn more grounded policies
  • 5. Experiments
  • 5.1. VLM grounding on captioning
  • 5.2. VLM grounding on VQA
  • 5.3. Ablations
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Multi-Modal Mutual-Information Decoding (M3ID) Algorithm

    algorithm

    Multi-Modal Mutual-Information Decoding (M3ID) is a training-free inference-time decoding algorithm for autoregressive Vision-Language Models (VLMs) that amplifies reliance on the visual prompt over the unconditioned language prior. It progressively boosts tokens that surprise the unconditioned language model while suppressing intervention when the conditional model exhibits high confidence due to contextual syntax.

    Input: Autoregressive VLM pp, textual prompt xx, image prompt cc, confidence threshold α∈(0,1]\alpha \in (0, 1], forgetting rate λ>0\lambda > 0
    Output: Generated token sequence y=(y0,y1,…,yT)y = (y_0, y_1, \dots, y_T)
    y0←BOSy_0 \gets \text{BOS}
    t←1t \gets 1
    while yt−1≠EOSy_{t-1} \neq \text{EOS} do
        γt←exp⁡(−λt)\gamma_t \gets \exp(-\lambda t)
        lc←log⁡p(⋅∣y<t,x,c)l_c \gets \log p(\cdot \mid y_{<t}, x, c)
        lu←log⁡p(⋅∣y<t,x)l_u \gets \log p(\cdot \mid y_{<t}, x)
        if max⁡v∈V(lc)v<log⁡α\max_{v \in \mathcal{V}} (l_c)_v < \log \alpha then
            l^∗←lc+1−γtγt(lc−lu)\hat{l}^* \gets l_c + \frac{1 - \gamma_t}{\gamma_t} (l_c - l_u)
        else
            l^∗←lc\hat{l}^* \gets l_c
        end if
        yt←arg⁡max⁡v∈Vl^∗(v)y_t \gets \arg\max_{v \in \mathcal{V}} \hat{l}^*(v)
        t←t+1t \gets t + 1
    end while
    return yy

    In standard implementations on models such as LLaVA-7B and LLaVA-13B, default hyperparameters are set to confidence threshold α=0.3\alpha = 0.3 and forgetting rate λ=0.02\lambda = 0.02. Greedy selection arg⁡max⁡\arg\max can optionally be replaced by beam search or stochastic sampling over l^∗\hat{l}^*.

  2. Knowl 2 — Fading Memory Formulation and Prompt Amplification in M3ID

    model/method

    Autoregressive Vision-Language Models (VLMs) suffer from conditioning dilution as generation proceeds. Under a fading memory assumption, the conditional log-probabilities lc(yt)=log⁡p(yt∣y<t,x,c)l_c(y_t) = \log p(y_t \mid y_{<t}, x, c) from a pre-trained VLM are modeled as an interpolation between an ideal non-forgetting distribution l∗(yt∣y<t,x,c)l^*(y_t \mid y_{<t}, x, c) and the unconditioned language prior lu(yt)=log⁡p(yt∣y<t,x)l_u(y_t) = \log p(y_t \mid y_{<t}, x):

    l(yt∣y<t,x,c)=γtl∗(yt∣y<t,x,c)+(1−γt)l(yt∣y<t,x)l(y_t \mid y_{<t}, x, c) = \gamma_t l^*(y_t \mid y_{<t}, x, c) + (1 - \gamma_t) l(y_t \mid y_{<t}, x)

    where γt=exp⁡(−λt)∈[0,1]\gamma_t = \exp(-\lambda t) \in [0, 1] is a monotonically decreasing mixing coefficient parametrized by the forgetting rate λ>0\lambda > 0, and tt denotes the token generation step index.

    Assuming l∗l^* is a perturbation l∗=lc+Δl^* = l_c + \Delta with a zero-mean, bounded-variance random variable Δ\Delta, the optimal intervention l^∗\hat{l}^* that recovers the non-diluted distribution is given by:

    l^∗=lc+1−γtγt(lc−lu)\hat{l}^* = l_c + \frac{1 - \gamma_t}{\gamma_t} (l_c - l_u)

    When γt→1\gamma_t \to 1 (near the visual prompt), l^∗≈lc\hat{l}^* \approx l_c. As γt→0\gamma_t \to 0 (far from the prompt), l^∗∝lc−lu\hat{l}^* \propto l_c - l_u, which corresponds to maximizing the pointwise mutual information between the visual input and text tokens: log⁡p(y∣x,c)p(y∣x)=lc−lu\log \frac{p(y \mid x, c)}{p(y \mid x)} = l_c - l_u.

    To prevent penalizing predictable functional tokens (such as prepositions and conjunctions) caused by contextual pressure rather than lack of grounding, the correction is gated by a confidence threshold α\alpha:

    l^∗=lc+I(max⁡v∈V(lc)v<log⁡α)1−γtγt(lc−lu)\hat{l}^* = l_c + \mathbb{I}\left( \max_{v \in \mathcal{V}} (l_c)_v < \log \alpha \right) \frac{1 - \gamma_t}{\gamma_t} (l_c - l_u)

    where I(⋅)\mathbb{I}(\cdot) is the indicator function and V\mathcal{V} is the token vocabulary.

  3. Knowl 3 — Visual Prompt Dependency Measure (PDM)

    definition

    The Visual Prompt Dependency Measure (PDM) evaluates whether the output of a Vision-Language Model is grounded in the conditioning visual prompt versus driven by its unconditional language prior, without requiring ground-truth annotations.

    Given an autoregressive VLM pp, an input visual context cc, a textual prompt xx, and previously generated tokens y<t=[y0,…,yt−1]y_{<t} = [y_0, \dots, y_{t-1}], the PDM at step tt is defined as:

    PDM(y<t;c∣x)≜dist(p(⋅∣y<t,x,c), p(⋅∣y<t,x))\text{PDM}(y_{<t}; c \mid x) \triangleq \text{dist}\Big(p(\cdot \mid y_{<t}, x, c), \, p(\cdot \mid y_{<t}, x)\Big)

    where dist(p,q)\text{dist}(p, q) is a statistical distance between discrete probability distributions over the vocabulary V\mathcal{V}.

    A primary instance is PDM-H, which utilizes the Hellinger distance:

    H(p,q)=12∑i=1∣V∣(pi−qi)2H(p, q) = \frac{1}{\sqrt{2}} \sqrt{\sum_{i=1}^{|\mathcal{V}|} \left(\sqrt{p_i} - \sqrt{q_i}\right)^2}

    A high PDM indicates that the predicted token distribution is strongly conditioned on the visual input cc, while a low PDM indicates that the distribution is prompt-agnostic and dominated by the language model prior.

  4. Knowl 4 — Conditioning Dilution and Hallucination Correlation in Autoregressive VLMs

    empirical result

    In autoregressive Vision-Language Models (such as LLaVA-13B and LLaVA-7B evaluated on MS COCO), the prompt dependency measure (PDM-H) decreases monotonically as the number of generated tokens increases. This degradation is termed conditioning dilution or the fading memory effect:

    p(yt∣y<t,x,c)→t→∞p(yt∣y<t,x)p(y_t \mid y_{<t}, x, c) \xrightarrow[t \to \infty]{} p(y_t \mid y_{<t}, x)

    Empirical measurements show that object hallucination frequency correlates inversely with PDM-H: while almost no hallucinated objects appear in the initial tokens close to the prompt, the proportion of non-existent objects increases steadily at higher token indices (e.g., in the later 60%–100% segment of open-ended generation), indicating that hallucinations predominantly stem from over-reliance on language priors rather than deficient visual feature representation.

  5. Knowl 5 — Self-Supervised Multi-Modal Direct Preference Optimization (M3ID+DPO)

    model/method

    M3ID+DPO aligns pre-trained Vision-Language Models to prefer visually grounded outputs over ungrounded language-prior continuations without relying on human annotations or external LLM (e.g., GPT-3.5/GPT-4) supervision.

    Preference pairs (yw,yl)(y_w, y_l) for an image cc and text prompt xx are generated self-supervisedly:

    1. Preferred continuations ywy_w are sampled from the pre-trained conditional distribution decoded with M3ID: yw∼p∗(y∣x,c)y_w \sim p^*(y \mid x, c).
    2. Dispreferred continuations yly_l are sampled from the unconditional distribution yl∼p(y∣x)y_l \sim p(y \mid x), where xx is prepended with the first sentence produced by the conditioned VLM to enforce topic consistency while allowing ungrounded language drift.

    The model parameters θ\theta are fine-tuned using the Direct Preference Optimization (DPO) loss parameterized by base model prefp_{\text{ref}}:

    LDPO(θ)=−E(c,x,yw,yl)∼D[log⁡σ(βlog⁡pθ(yw∣c,x)pref(yw∣c,x)−βlog⁡pθ(yl∣c,x)pref(yl∣c,x))]\mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(c, x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{p_\theta(y_w \mid c, x)}{p_{\text{ref}}(y_w \mid c, x)} - \beta \log \frac{p_\theta(y_l \mid c, x)}{p_{\text{ref}}(y_l \mid c, x)} \right) \right]

    where σ\sigma is the logistic sigmoid and β\beta controls deviation from prefp_{\text{ref}} (typically β=0.1\beta = 0.1, trained with LoRA at peak learning rate 2×10−52 \times 10^{-5} for 5 epochs on 10,000 self-generated preference pairs).

  6. Knowl 6 — MS COCO Image Captioning Grounding and Hallucination Benchmarks

    data/table

    The evaluation on the MS COCO validation set measures object hallucinations via CHAIRi\text{CHAIR}_i (percentage of mentioned objects that are hallucinated, lower is better), CHAIRs\text{CHAIR}_s (percentage of captions containing at least one hallucinated object, lower is better), and Cover\text{Cover} (percentage of ground-truth annotated objects present in the generated captions, higher is better). Models are prompted with "Describe the image."

    Model / Decoding Method CHAIR_i CHAIR_s Cover
    LLaVA-7B (Base Multinomial) 8.1 17.5 53.3
    LLaVA-7B + PMI 6.7 16.2 51.5
    LLaVA-7B + Contrastive Decoding 6.3 14.8 55.2
    LLaVA-7B + M3ID 5.9 13.8 55.1
    LLaVA-7B + M3ID + DPO 5.7 13.5 55.8
    LLaVA-13B (Base Multinomial) 7.4 18.5 55.2
    LLaVA-13B + PMI 6.5 16.6 54.7
    LLaVA-13B + Contrastive Decoding 6.5 14.5 54.7
    LLaVA-13B + M3ID 5.5 13.2 54.0
    LLaVA-13B + M3ID + DPO 5.3 12.6 54.2
    LLaVA-13B + LURE (Supervised) 6.4 27.1 –
    mPLUG-Owl + LURE (Supervised) 5.4 18.8 –

    M3ID reduces CHAIRi\text{CHAIR}_i by 27% relative on LLaVA-7B and 26% on LLaVA-13B over base multinomial sampling without significant reduction in the object coverage metric (Cover). Combining M3ID with self-supervised DPO yields the lowest overall hallucination rates (5.3% CHAIRi\text{CHAIR}_i on 13B) while outperforming annotation-dependent methods like LURE.

  7. Knowl 7 — POPE VQA Hallucination Benchmark Results

    data/table

    The Polling-based Object Probing Evaluation (POPE) benchmark measures object hallucinations in visual question answering via binary yes/no questions ("Is a <object> present in the image?") across Random, Popular, and Adversarial object sampling splits. Acc denotes binary classification accuracy (higher is better) and Yes% denotes the proportion of affirmative answers (50% is balanced).

    To account for the token distance between the image tokens and the answer token caused by the question prefix, M3ID applies a token offset t=t0t = t_0, where t0t_0 is the number of intermediate instruction tokens.

    Random Popular Adversarial All
    Model Acc. Yes (%) Acc. Yes (%) Acc. Yes (%) Acc. Yes (%)
    Robust mPLUG-Owl-7B (Supervised) 86.0 – 73.0 – 65.0 – 74.7 –
    LLaVA-RLHF-7B (Supervised) 84.8 39.6 83.3 41.8 80.7 44.0 82.9 41.8
    mPLUG-Owl-7B 52.0 – 57.0 – 60.0 – 67.3 –
    LLaVA-7B (Base) 74.8 75.1 61.8 86.7 58.1 90.1 64.9 84.0
    LLaVA-7B + M3ID 76.0 67.7 69.3 73.3 65.8 77.6 70.3 72.9
    LLaVA-7B + M3ID + DPO 81.2 65.6 73.9 67.3 68.2 75.4 74.4 69.4
    MiniGPT4-13B 73.0 – 67.0 – 62.0 – 74.7 –
    LLaVA-RLHF-13B (Supervised) 85.2 38.4 83.9 38.0 82.3 40.5 83.8 39.0
    LLaVA-13B (Base) 67.9 80.6 63.8 83.2 59.8 87.3 63.8 83.7
    LLaVA-13B + M3ID 84.3 55.6 77.0 61.6 71.3 68.2 77.5 61.8
    LLaVA-13B + M3ID + DPO 85.2 53.4 79.1 57.5 73.2 67.5 79.2 51.1

    M3ID reduces the base LLaVA-13B model's bias toward answering 'Yes' from 83.7% to 61.8% and improves overall POPE accuracy from 63.8% to 77.5%. Fine-tuning with DPO further improves accuracy to 79.2% and balances Yes% to 51.1% without requiring task-specific VQA labels during preference generation.

  8. Knowl 8 — Sensitivity and Component Ablations of M3ID

    data/table

    The effectiveness of M3ID relies on balancing the forgetting rate λ\lambda and the contextual confidence threshold α\alpha. Setting λ\lambda or α\alpha too high causes overcompensation against the language prior, harming linguistic fluency and object coverage, whereas setting them too low fails to counteract conditioning dilution.

    Hyperparameter Configuration (LLaVA-7B) CHAIR_i CHAIR_s Cover
    Base LLaVA-7B 8.0 17.8 53.5
    α=0.3,λ=0.02\alpha = 0.3, \lambda = 0.02 (Optimal M3ID) 5.8 13.6 55.0
    α=0.3,λ=0.001\alpha = 0.3, \lambda = 0.001 7.3 17.4 53.7
    α=0.3,λ=0.03\alpha = 0.3, \lambda = 0.03 6.2 14.7 52.0
    α=0.3,λ=0.1\alpha = 0.3, \lambda = 0.1 7.2 16.3 45.5
    α=0.5,λ=0.02\alpha = 0.5, \lambda = 0.02 6.1 14.9 54.6
    α=0.1,λ=0.02\alpha = 0.1, \lambda = 0.02 6.9 15.4 53.1
    α=0.01,λ=0.02\alpha = 0.01, \lambda = 0.02 6.9 16.6 52.5

    Ablating the components demonstrates their complementary roles:

    Ablation Component POPE Acc. CHAIR_i CHAIR_s Cover
    LLaVA-7B (Base) 64.9 8.0 17.8 53.5
    + context pressure threshold only 65.5 6.9 16.7 54.5
    + conditioning dilution term only 77.5 6.4 14.1 53.9
    LLaVA-7B Full M3ID 77.5 5.8 13.6 55.0
    LLaVA-13B (Base) 63.8 15.3 18.2 55.3
    + context pressure threshold only 64.9 6.5 14.4 55.0
    + conditioning dilution term only 70.3 5.9 14.8 53.7
    LLaVA-13B Full M3ID 70.3 5.4 13.0 54.0

    Conditioning dilution correction provides the primary gain on POPE classification and caption hallucination metrics, while the contextual pressure threshold preserves coverage and syntactic validity.

  9. Knowl 9 — Limitations of Multi-Modal Mutual-Information Decoding

    limitation

    M3ID exhibits two main practical limitations:

    1. Computational overhead at inference time: Evaluating both the image-conditioned log-probabilities l(yt∣y<t,x,c)l(y_t \mid y_{<t}, x, c) and the unconditioned language prior l(yt∣y<t,x)l(y_t \mid y_{<t}, x) requires two forward model passes per decoding step (or batched execution with masked visual tokens), increasing inference latency and memory footprint.
    2. Overcompensation on predictable objects: Because M3ID amplifies tokens that diverge from the language prior, it can occasionally fail to mention obvious objects that are strongly implied by textual context clues alone (e.g., omitting "man" when describing "a dog on a leash accompanied by a man") if hyperparameters λ\lambda and α\alpha are not carefully tuned.

Coverage note — None was omitted; all key contributions—including the visual prompt dependency measure, the fading memory theoretical framing, the M3ID algorithm, self-supervised DPO pairing, empirical evaluations on MS COCO and POPE, hyperparameter ablations, and stated limitations—are fully represented.

References

  1. 1.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. A multitask, multilingual, multi-modal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023, 2023. 2
  2. 2.Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952. 12
  3. 3.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023. 5
  4. 4.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. 1, 2, 5, 16
  5. 5.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. 2
  6. 6.Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. The factual inconsistency problem in abstractive text summarization: A survey. arXiv preprint arXiv:2104.14839, 2021. 2
  7. 7.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119, San Diego, California, 2016. Association for Computational Linguistics. 1, 3, 4, 6, 11
  8. 8.Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. Contrastive decoding: Open-ended text generation as optimization, 2023. 2, 5, 6, 11, 14
  9. 9.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 2, 6, 7, 11, 17
  10. 10.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014. 4, 5, 6, 11, 14, 15
  11. 11.Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023. 2, 6, 7
  12. 12.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023. 5
  13. 13.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 1, 2, 4, 5, 11
  14. 14.M.B. Matthews and G.S. Moschytz. The identification of nonlinear discrete-time fading-memory systems using neural network models. IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing, 41(11):740–751, 1994. 4
  15. 15.Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. Locally typical sampling. Transactions of the Association for Computational Linguistics, 11:102–121, 2023. 2
  16. 16.Yatin Nandwani, Vineet Kumar, Dinesh Raghu, Sachindra Joshi, and Luis A Lastras. Pointwise mutual information based metric and decoding strategy for faithful generation in document grounded dialogs. arXiv preprint arXiv:2305.12191, 2023. 2
  17. 17.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744, 2022. 1, 6, 7
  18. 18.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 5
  19. 19.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023. 2, 5, 6, 11, 12
  20. 20.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045, Brussels, Belgium, 2018. Association for Computational Linguistics. 2, 6, 11, 14, 17
  21. 21.Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Scott Wen-tau Yih. Trusting your evidence: Hallucinate less with context-aware decoding. arXiv preprint arXiv:2305.14739, 2023. 2
  22. 22.Merrielle Spain and Pietro Perona. Measuring and predicting object importance. Int. J. Comput. Vis., 91(1):59–76, 2011. 8
  23. 23.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525, 2023. 2, 7
  24. 24.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 2, 5
  25. 25.Liam van der Poel, Ryan Cotterell, and Clara Meister. Mutual information alleviates hallucinations in abstractive summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5956–5965, Abu Dhabi, United Arab Emirates, 2022. Association for Computational Linguistics. 2, 3, 4, 5, 6, 11, 14
  26. 26.Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, and Shengyi Huang. Trl: Transformer reinforcement learning. https://github.com/huggingface/trl, 2020. 6
  27. 27.Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al. Evaluation and analysis of hallucination in large vision-language models. arXiv preprint arXiv:2308.15126, 2023. 2
  28. 28.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022. 1
  29. 29.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023. 7
  30. 30.Luca Zancato and Alessandro Chiuso. A novel deep neural network architecture for non-linear system identification. IFAC-PapersOnLine, 54(7):186–191, 2021. 19th IFAC Symposium on System Identification SYSID 2021. 4
  31. 31.Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. Analyzing and mitigating object hallucination in large vision-language models. arXiv preprint arXiv:2310.00754, 2023. 2, 6, 7
  32. 32.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 2, 7

Citation

MLA
Favero, A., et al. “Multi-Modal Hallucination Control by Visual Information Grounding”. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024, 2024, http://arxiv.org/abs/2403.14003v1.
APA
Favero, A., Zancato, L., Trager, M., Choudhary, S., Perera, P., Achille, A., Swaminathan, A., & Soatto, S. (2024). Multi-Modal Hallucination Control by Visual Information Grounding. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024. http://arxiv.org/abs/2403.14003v1
Chicago
Favero, A., L. Zancato, M. Trager, et al. 2024. “Multi-Modal Hallucination Control by Visual Information Grounding”. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024. http://arxiv.org/abs/2403.14003v1.
Harvard
Favero, A. et al. (2024) “Multi-Modal Hallucination Control by Visual Information Grounding”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 [Preprint]. Available at: http://arxiv.org/abs/2403.14003v1.
Vancouver
1. Favero A, Zancato L, Trager M, Choudhary S, Perera P, Achille A, Swaminathan A, Soatto S (2024) Multi-Modal Hallucination Control by Visual Information Grounding. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024

BibTeX

@article{favero2024multi,
  title = {Multi-Modal Hallucination Control by Visual Information Grounding},
  author = {Favero, Alessandro and Zancato, Luca and Trager, Matthew and Choudhary, Siddharth and Perera, Pramuditha and Achille, Alessandro and Swaminathan, Ashwin and Soatto, Stefano},
  year = {2024},
  journal = {IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024},
  url = {http://arxiv.org/abs/2403.14003v1},
  eprint = {2403.14003}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE