Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
Peng QiZehong YanWynne HsuMong-Li Lee
Proposes SNIFFER, a multimodal large language model fine-tuned via two-stage instruction tuning and augmented with external retrieval to detect out-of-context misinformation while generating accurate, persuasive explanations.
Out-of-context misinformation, where authentic images are paired with false or misleading text, has become a widespread and damaging form of online deception. Because the images themselves are unaltered, traditional media forensics fail to detect them. Existing automated detectors often operate as black boxes that provide veracity judgments without explaining why an image-text pair is inconsistent, limiting their practical utility for fact-checkers and eroding public trust. Meanwhile, general-purpose multimodal large language models struggle with this task because they lack domain-specific entity knowledge, cannot retrieve real-time external evidence, and often fail to discern subtle contextual mismatches.
The article introduces and evaluates SNIFFER, a multimodal large language model specifically engineered to detect out-of-context misinformation and generate clear, natural language explanations for its decisions.
To build SNIFFER, the authors fine-tuned an existing open-source model using a two-stage instruction training process. The first stage aligned generic visual concepts with fine-grained news entities using approximately 370,000 news image-caption pairs. The second stage used over 71,000 GPT-4-assisted instruction examples to teach the model how to spot specific discrepancies between text and images. During operation, SNIFFER combines internal cross-modal consistency checking, assisted by entity detection tools, with external verification that compares claims against retrieved web evidence. The system then merges these findings to produce a final judgment and explanation.
The evaluation produced several key findings. First, SNIFFER achieved an overall detection accuracy of 88.4% on the standard NewsCLIPpings benchmark, outperforming existing state-of-the-art detectors and surpassing the baseline un-tuned model by over 40 percentage points. Second, in a direct comparison on a test sample, SNIFFER exceeded the proprietary GPT-4V model by 11 percentage points in classification accuracy (86.8% versus 75.5%). Third, the model demonstrated strong data efficiency, achieving competitive baseline performance when trained on just 10% of the dataset. Fourth, human evaluations confirmed that SNIFFER's explanations are highly persuasive: after reading them, participants corrected 87% of their initial false-positive errors, and 42% reported increased confidence in their correct fake-news assessments.
These results indicate that specialized, task-tuned multimodal models can significantly outperform larger, general-domain foundation models in high-stakes domain-specific applications. Providing actionable and accurate natural language justifications mitigates the reputational and operational risks of automated moderation systems, making them viable tools for journalists, policy compliance teams, and platform moderators tasked with rapid debunking.
Organizations seeking to combat out-of-context media should consider deploying domain-tuned, retrieval-augmented models rather than relying solely on general-purpose commercial models. Next steps should include conducting live pilot deployments within fact-checking workflows and evaluating performance across non-English media and emerging news cycles.
Decision-makers should note certain limitations. The system depends partly on the quality and availability of external search retrieval, as only about 60% of test samples had retrievable external evidence, and noisy search results can affect real-news accuracy. Additionally, the model was primarily evaluated on synthetic benchmark splits generated from major news agencies. However, the evidence provides high confidence that SNIFFER represents a robust, state-of-the-art approach for explainable misinformation detection.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Establishes how large vision-language models struggle with cross-modal grounding and hallucination, motivating SNIFFER's domain-specific fine-tuning for visual-text consistency.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). Introduces the foundational framework for retrieval-augmented fact extraction and verification that underpins external evidence checking in automated debunking systems.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). Examines how language models integrate retrieved evidence to generate grounded and verifiable factual text, providing core principles for explainable verification.
- Paper: Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection, Yang Qiao et al. (2023). Provides fundamental techniques for detecting fine-grained semantic incongruities between visual objects and paired text descriptions.
- Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). Outlines the foundational data-mining principles and taxonomies for automated misinformation detection on social media platforms.
- Paper: Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision, Seongyun Lee et al. (2024). Extends multimodal factuality and reasoning by using self-critique and iterative visual feedback to correct cross-modal errors within a single unified model.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). Builds on cross-modal grounding principles by introducing decoding and preference optimization techniques to prevent multimodal models from drifting into ungrounded text.
- Paper: Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention, Wenbin An et al. (2025). Advances vision-language factual alignment by dynamically combining global image representations with focused local attention to verify fine-grained visual details.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). Provides a comprehensive taxonomy and survey of automated verification, hallucination mitigation, and factual integrity across both text and multimodal large language models.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). Applies multimodal large language models to generate natural language explanations and structured metrics for evaluating cross-modal visual consistency.
