Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Peng QiZehong YanWynne HsuMong-Li Lee

article2024CVPR135 citations

Proposes SNIFFER, a multimodal large language model fine-tuned via two-stage instruction tuning and augmented with external retrieval to detect out-of-context misinformation while generating accurate, persuasive explanations.

Listen

Out-of-context misinformation, where authentic images are paired with false or misleading text, has become a widespread and damaging form of online deception. Because the images themselves are unaltered, traditional media forensics fail to detect them. Existing automated detectors often operate as black boxes that provide veracity judgments without explaining why an image-text pair is inconsistent, limiting their practical utility for fact-checkers and eroding public trust. Meanwhile, general-purpose multimodal large language models struggle with this task because they lack domain-specific entity knowledge, cannot retrieve real-time external evidence, and often fail to discern subtle contextual mismatches.

The article introduces and evaluates SNIFFER, a multimodal large language model specifically engineered to detect out-of-context misinformation and generate clear, natural language explanations for its decisions.

To build SNIFFER, the authors fine-tuned an existing open-source model using a two-stage instruction training process. The first stage aligned generic visual concepts with fine-grained news entities using approximately 370,000 news image-caption pairs. The second stage used over 71,000 GPT-4-assisted instruction examples to teach the model how to spot specific discrepancies between text and images. During operation, SNIFFER combines internal cross-modal consistency checking, assisted by entity detection tools, with external verification that compares claims against retrieved web evidence. The system then merges these findings to produce a final judgment and explanation.

The evaluation produced several key findings. First, SNIFFER achieved an overall detection accuracy of 88.4% on the standard NewsCLIPpings benchmark, outperforming existing state-of-the-art detectors and surpassing the baseline un-tuned model by over 40 percentage points. Second, in a direct comparison on a test sample, SNIFFER exceeded the proprietary GPT-4V model by 11 percentage points in classification accuracy (86.8% versus 75.5%). Third, the model demonstrated strong data efficiency, achieving competitive baseline performance when trained on just 10% of the dataset. Fourth, human evaluations confirmed that SNIFFER's explanations are highly persuasive: after reading them, participants corrected 87% of their initial false-positive errors, and 42% reported increased confidence in their correct fake-news assessments.

These results indicate that specialized, task-tuned multimodal models can significantly outperform larger, general-domain foundation models in high-stakes domain-specific applications. Providing actionable and accurate natural language justifications mitigates the reputational and operational risks of automated moderation systems, making them viable tools for journalists, policy compliance teams, and platform moderators tasked with rapid debunking.

Organizations seeking to combat out-of-context media should consider deploying domain-tuned, retrieval-augmented models rather than relying solely on general-purpose commercial models. Next steps should include conducting live pilot deployments within fact-checking workflows and evaluating performance across non-English media and emerging news cycles.

Decision-makers should note certain limitations. The system depends partly on the quality and availability of external search retrieval, as only about 60% of test samples had retrievable external evidence, and noisy search results can affect real-news accuracy. Additionally, the model was primarily evaluated on synthetic benchmark splits generated from major news agencies. However, the evidence provides high confidence that SNIFFER represents a robust, state-of-the-art approach for explainable misinformation detection.

arXiv: 2403.03170
Cover for Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Abstract

Misinformation is a prevalent societal issue due to its potential high risks. Out-Of-Context (OOC) misinformation, where authentic images are repurposed with false text, is one of the easiest and most effective ways to mislead audiences. Current methods focus on assessing image-text consistency but lack convincing explanations for their judgments, which are essential for debunking misinformation. While Multimodal Large Language Models (MLLMs) have rich knowledge and innate capability for visual reasoning and explanation generation, they still lack sophistication in understanding and discovering the subtle cross-modal differences. In this paper, we introduce SNIFFER, a novel multimodal large language model specifically engineered for OOC misinformation detection and explanation. SNIFFER employs two-stage instruction tuning on Instruct-BLIP. The first stage refines the model’s concept alignment of generic objects with news-domain entities and the second stage leverages OOC-specific instruction data generated by language-only GPT-4 to fine-tune the model’s discriminatory powers. Enhanced by external tools and retrieval, SNIFFER not only detects inconsistencies between text and image but also utilizes external knowledge for contextual verification. Our experiments show that SNIFFER surpasses the original MLLM by over 40% and outperforms state-of-the-art methods in detection accuracy. SNIFFER also provides accurate and persuasive explanations as validated by quantitative and human evaluations.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Base MLLM
  • 3.2. Instruction Tuning
  • 3.3. Reasoning Process
  • 4. Performance Study
  • 4.1. Experimental Setup
  • 4.2. Performance Comparison (Q1)
  • 4.3. Ablation Studies (Q2)
  • 4.4. Explainability Analysis (Q3)
  • 4.5. Practical Setting
  • 4.6. Comparison with GPT-4V (Q6)
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — SNIFFER Architecture and Composed Reasoning Framework

    model/method

    SNIFFER is a multimodal large language model framework designed to detect out-of-context (OOC) misinformation and generate natural language explanations. OOC misinformation occurs when an authentic image is paired with false or misleading text. SNIFFER operates through a three-step reasoning process:

    1. Internal Checking: Evaluates cross-modal consistency directly between the input image and news caption. The image is passed through a vision-language model (InstructBLIP with a frozen ViT-G/14 vision encoder, a frozen Vicuna-13B LLM, and a trainable Query Transformer / Q-Former) alongside visual entities recognized in the image using the Google Entity Detection API. The model is instructed to detect whether the image is misattributed and identify the mismatched entity and category.

    2. External Checking: Reverse-image search retrieves webpages where the image previously appeared to obtain source news context. The LLM module receives the input caption and the retrieved textual evidence to evaluate whether the claimed news context is supported by external evidence.

    3. Composed Reasoning: The LLM module acts as an interpretable model ensemble. It takes the original news caption and the separate outputs from internal checking and external checking to synthesize the final binary judgment (Real vs. Fake / Out-of-Context) and provide a coherent, multi-step explanation.

  2. Knowl 2 — Two-Stage Instruction Tuning for News Alignment and OOC Detection

    model/method

    SNIFFER adapts a general-purpose vision-language model (InstructBLIP) to out-of-context misinformation detection via a sequential two-stage instruction tuning process. In both stages, the visual encoder (ViT-G/14) and the language model (Vicuna-13B) remain frozen; only the Query Transformer (Q-Former) parameters are updated:

    • Stage 1: News Domain Alignment: General-purpose MLLMs often generate coarse entity labels (such as "man" or "person") rather than fine-grained named entities (such as "Donald Trump"). Stage 1 aligns the visual feature space of news images to textual embeddings of specific named entities in the pre-trained LLM. It is trained on 368,013 news image-caption pairs from the NewsCLIPpings dataset for 1 epoch (~3 hours) using image description instructions.

    • Stage 2: Task-Specific Tuning: General MLLMs assume image-text alignment, whereas OOC detection requires identifying discrepancies across disparate events. Stage 2 trains the Q-Former for 10 epochs (~13 hours) on 71,072 instruction-formatted pairs (35,536 fake pairs paired with explanations of inconsistencies and 35,536 real pairs) to detect cross-modal inconsistencies, the category of mismatch, and specific conflicting entities.

  3. Knowl 3 — GPT-4-Assisted Out-of-Context Instruction Data Construction

    model/method

    In standard OOC benchmark datasets like NewsCLIPpings, a fabricated pair is created by replacing an image img1img_1 from a genuine news pair (cap1,img1)(cap_1, img_1) with an image img2img_2 from another pair (cap2,img2)(cap_2, img_2), resulting in (cap1,img2)(cap_1, img_2). While binary authenticity labels exist, granular supervised explanations pinpointing the exact inconsistency are absent.

    To construct supervised instruction-explanation data without human annotation:

    1. A basic description of img2img_2 is generated using InstructBLIP.
    2. Language-only GPT-4 is prompted with cap1cap_1, cap2cap_2, and the basic description of img2img_2.
    3. With few-shot examples and a constrained output template, GPT-4 infers the inconsistency between cap1cap_1 and img2img_2 as if viewing the image directly.
    4. GPT-4 outputs the primary inconsistent news element (e.g., person, location, organization, event) along with the specific text entity ent_tent\_t in cap1cap_1 and the visual entity ent_vent\_v in img2img_2.

    This pipeline yields 35,536 instruction-response pairs for OOC samples (formatted as: "No, the image is wrongly used in a different news context. The given news caption and image are inconsistent in {element}. The {element} in caption is {ent_t}, and the {element} in image is {ent_v}."), which are balanced with 35,536 real pairs ("No, the image is rightly used in the given news context.").

  4. Knowl 4 — Stage 1 Instruction Formulation for News Domain Alignment

    equation

    During Stage 1 of instruction tuning (News Domain Alignment), the model is trained to align fine-grained news visual concepts with named entity embeddings using instruction-response pairs formatted as:

    Human: ITq⟨STOP⁡⟩;  Model: Tc⟨STOP⁡⟩\text{\textbf{Human}: } \mathbf{I} \mathbf{T}_{\mathrm{q}} \langle \operatorname{STOP} \rangle ; ~~ \text{\textbf{Model}: } \mathbf{T}_{\mathrm{c}}\langle \operatorname{STOP} \rangle

    where I\mathbf{I} represents the input news image embeddings extracted by the frozen visual encoder, Tq\mathbf{T}_{\mathrm{q}} is a prompt text randomly sampled from a pool of 11 diverse questions explicitly asking for a brief description of the image content (e.g., "Can you give a brief description of this image?"), Tc\mathbf{T}_{\mathrm{c}} is the reference news caption (containing fewer than 30 words), and ⟨STOP⁡⟩\langle \operatorname{STOP} \rangle is the sequence termination token.

  5. Knowl 5 — Out-of-Context Misinformation Detection Performance on NewsCLIPpings

    data/table

    The classification performance of SNIFFER was evaluated on the Merged/Balance benchmark test set of the NewsCLIPpings dataset (7,264 test samples: 3,632 Real and 3,632 Fake / Out-of-Context) and compared against multimodal misinformation detectors trained from scratch, pre-trained multimodal models, and an MLLM-based detector:

    Method All (%) Fake (%) Real (%)
    SAFE 52.8 54.8 52.0
    EANN 58.1 61.8 56.2
    VisualBERT 58.6 38.9 78.4
    CLIP 66.0 64.3 67.7
    DT-Transformer 77.1 78.6 75.6
    CCN 84.7 84.8 84.5
    Neu-Sym detector 68.2 - -
    SNIFFER (Ours) 88.4 86.9 91.8

    SNIFFER achieves an overall accuracy of 88.4%, outperforming the strongest baseline (CCN, which uses external multimodal evidence) by 3.7 percentage points overall and achieving superior balance across both Fake (86.9%) and Real (91.8%) samples.

  6. Knowl 6 — Component Ablation Study of SNIFFER

    data/table

    An ablation study evaluated the incremental impact of each component in SNIFFER on the NewsCLIPpings Merged/Balance test set: pre-training on news-domain data in Stage 1 (PT), task-specific OOC instruction tuning in Stage 2 (OOC Tuning), visual entities recognized by external tools (VisEnt), and retrieved external web text evidence (Evidence):

    InstructBLIP PT OOC Tuning VisEnt Evidence All (%) Fake (%) Real (%)
    ✓ 47.4 4.6 90.3
    ✓ ✓ 49.3 9.4 89.2
    ✓ ✓ 82.5 75.3 89.7
    ✓ ✓ ✓ 87.6 83.9 91.3
    ✓ ✓ ✓ 83.1 76.5 89.6
    ✓ ✓ ✓ ✓ 88.2 84.9 94.0
    ✓ ✓ 84.5 92.9 76.0
    ✓ ✓ ✓ ✓ ✓ 88.4 86.9 91.8

    Key observations:

    • Base InstructBLIP achieves only 4.6% recall on fake samples (47.4% overall accuracy), showing a severe inductive bias toward assuming image-text pairs are authentic.
    • Task-specific Stage 2 tuning provides the largest improvement, boosting accuracy by over 35 percentage points (to 82.5%).
    • Incorporating external visual entities adds ~5% accuracy.
    • Relying solely on external evidence yields high fake recall (92.9%) but lower real recall (76.0%) due to noise and absence of retrieval results (available for <60% of samples).
  7. Knowl 7 — Explanation Quality Evaluation Through Quantitative Metrics and Human Study

    empirical result

    The quality of explanations generated by SNIFFER was assessed using both quantitative ground-truth comparison and human evaluation:

    • Quantitative Metrics: Explanations were evaluated on the NewsCLIPpings test set across three key elements: inconsistent element type (hard match hit ratio), textual entity ent_tent\_t (CLIP embedding similarity), visual entity ent_vent\_v (CLIP embedding similarity), and entire explanation text (ROUGE score). Stage 1 pre-training and Stage 2 OOC tuning + visual entities improve the element hit ratio by 4% and 44%, respectively. Accuracy for identifying text entities (ent_tent\_t) consistently exceeded visual entities (ent_vent\_v) by approximately 17%, reflecting the greater difficulty of visual entity extraction. Although the model outputs entities more conservatively after tuning, response accuracy increases significantly across all metrics.

    • Human Evaluation: 10 participants evaluated 20 correctly detected OOC test samples before and after viewing SNIFFER's explanations. Initially, users correctly identified 69% of OOC items as fake. After reading SNIFFER's judgment and explanation, 87% of user misclassifications (initially labeled real) were revised to fake. For items already identified as fake, SNIFFER's explanation increased participant confidence in 42% of cases.

  8. Knowl 8 — Sample Efficiency in Early Detection and Cross-Dataset Generalization

    empirical result

    SNIFFER demonstrates strong training data efficiency and cross-dataset generalizability:

    • Early Detection / Data Efficiency: When evaluated on SNIFFER- (InstructBLIP with only Stage 2 OOC tuning, without external evidence retrieval), training on only 10% of the training dataset matches the detection performance of baseline models (such as DT-Transformer at ~77% accuracy) trained on 100% of the data. Training on 25% of the data achieves ~79.5% accuracy, exceeding fully trained baselines.

    • Cross-Dataset Generalization: SNIFFER trained exclusively on NewsCLIPpings was evaluated without fine-tuning on two external benchmarks, News400 and TamperedNews. Each benchmark contains four subsets with increasing difficulty based on visual similarity between replaced and original images (random, top-25%, top-10%, and top-5%). SNIFFER consistently outperforms the Cross-Modal Context Similarity (CMCS) baseline across all difficulty tiers on both datasets, maintaining ~75%–90% accuracy compared to CMCS's ~60%–80%.

  9. Knowl 9 — Classification Accuracy Comparison Between SNIFFER and GPT-4V

    data/table

    A comparative evaluation was performed between SNIFFER and GPT-4 with Vision (GPT-4V) on a randomly sampled test set of 400 instances (200 real, 200 fake) from the NewsCLIPpings benchmark using identical evaluation prompts:

    Method All (%) Fake (%) Real (%)
    GPT-4V 75.5 77.0 74.0
    SNIFFER (Ours) 86.8 79.0 94.5

    SNIFFER achieves 86.8% overall accuracy on the sampled test set, outperforming GPT-4V (75.5%) by 11.3 percentage points, demonstrating that a task-adapted 13B-scale MLLM can surpass massive general-purpose multimodal LLMs on fine-grained news domain cross-modal verification.

  10. Knowl 10 — Implementation and Hyperparameter Configuration for SNIFFER

    experimental setup

    SNIFFER is implemented using the LAVIS library based on InstructBLIP:

    • Architecture: ViT-G/14 as the visual encoder, Vicuna-13B as the language model, and a Query Transformer (Q-Former) connecting them. FlashAttention-2 replaces the standard attention layers in the LLM to reduce GPU memory consumption.
    • Trainable Parameters: Only Q-Former parameters are tuned; the ViT-G/14 encoder and Vicuna-13B LLM remain frozen throughout.
    • Batch Size: 8 for Stage 1 (News Domain Alignment) and 4 for Stage 2 (Task-Specific Tuning).
    • Sequence Lengths: Maximum input sequence length is 550 tokens; maximum output generation length is 256 tokens.
    • Optimization: AdamW optimizer with a linear warmup increasing learning rate from 10−810^{-8} to 10−510^{-5}, followed by cosine decay.
    • Hardware: Trained on 4 Nvidia A100 (40GB) GPUs.
    • External Tools: Google Entity Detection API for visual entity extraction; reverse image search for retrieving external webpage news text.

Coverage note — None was omitted; all key contributions including framework architecture, two-stage instruction tuning, GPT-4 data generation, mathematical formulation, main benchmark performance, ablation study, explainability evaluation, data efficiency/generalization, GPT-4V comparison, and implementation parameters are fully covered.

References

  1. 1.Google Vision API. https://cloud.google.com/vision/docs/detecting-web. 5
  2. 2.Reading about the Israel-Hamas war on X? Beware fake news. https://wired.me/technology/x-misinformation/, 2023. 2
  3. 3.Sahar Abdelnabi, Rakibul Hasan, and Mario Fritz. Open-domain, content-based, multi-modal fact-checking of out-of-context images via online resources. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 14920–14929. IEEE, 2022. 2, 3, 5, 6
  4. 4.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. ´ Flamingo: a visual language model for few-shot learning. In NeurIPS, 2022. 3
  5. 5.Shivangi Aneja, Chris Bregler, and Matthias Nießner. COSMOS: catching out-of-context image misuse using self-supervised learning. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, pages 14084–14092. AAAI Press, 2023. 3
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 8
  7. 7.Juan Cao, Peng Qi, Qiang Sheng, Tianyun Yang, Junbo Guo, and Jintao Li. Exploring the role of visual content in fake news detection. Disinformation, Misinformation, and Fake News in Social Media: Emerging Research Challenges and Opportunities, pages 141–161, 2020. 3
  8. 8.Lu Cheng, Kush R. Varshney, and Huan Liu. Socially responsible AI algorithms: Issues, purposes, and challenges. J. Artif. Intell. Res., 71:1137–1181, 2021. 2
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 6
  10. 10.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. CoRR, abs/2305.06500, 2023. 2, 3, 6
  11. 11.Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. CoRR, abs/2307.08691, 2023. 6
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. 6
  13. 13.Lisa Fazio. Out-of-context photos are a powerful low-tech form of misinformation. https://theconversation.com/out-of-context-photos-are-a-powerful-low-tech-form-of-misinformation-129959, 2020. 2
  14. 14.Bin Guo, Yasan Ding, Lina Yao, Yunji Liang, and Zhiwen Yu. The future of false information detection on social media: New perspectives and trends. ACM Comput. Surv., 53 (4):68:1–68:36, 2021. 8
  15. 15.Quzhe Huang, Mingxu Tao, Zhenwei An, Chen Zhang, Cong Jiang, Zhibin Chen, Zirui Wu, and Yansong Feng. Lawyer LLaMA technical report. ArXiv, abs/2305.15062, 2023. 3
  16. 16.Ayush Jaiswal, Ekraam Sabir, Wael Abd-Almageed, and Premkumar Natarajan. Multimedia semantic integrity assessment using joint embedding of images and text. In Proceedings of the 25th ACM International Conference on Multimedia, MM 2017, Mountain View, CA, USA, October 23-27, 2017, pages 1465–1471. ACM, 2017. 2, 3
  17. 17.Ayush Jaiswal, Yue Wu, Wael AbdAlmageed, Iacopo Masi, and Premkumar Natarajan. AIRD: adversarial learning framework for image repurposing detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 11330–11339. Computer Vision Foundation / IEEE, 2019. 2, 3
  18. 18.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023. 3
  19. 19.Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. Multimodal foundation models: From specialists to general-purpose assistants. CoRR, abs/2309.10020, 2023. 2
  20. 20.Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day. CoRR, abs/2306.00890, 2023. 3
  21. 21.Dongxu Li, Junnan Li, Hung Le, Guangsen Wang, Silvio Savarese, and Steven C. H. Hoi. LAVIS: A one-stop library for language-vision intelligence. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL 2023, Toronto, Canada, July 10-12, 2023, pages 31–41. Association for Computational Linguistics, 2023. 6
  22. 22.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, pages 19730–19742. PMLR, 2023. 3
  23. 23.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. VisualBERT: A simple and performant baseline for vision and language. CoRR, abs/1908.03557, 2019. 3, 6
  24. 24.C Lin. Recall-oriented understudy for gisting evaluation (rouge). Retrieved August, 20:2005, 2005. 7
  25. 25.Fuxiao Liu, Yinghan Wang, Tianlu Wang, and Vicente Ordonez. Visual news: Benchmark and challenges in news image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6761–6771. Association for Computational Linguistics, 2021. 5
  26. 26.Haoyang Liu, Maheep Chaudhary, and Haohan Wang. Towards trustworthy and aligned machine learning: A data-centric survey with causality perspectives. CoRR, abs/2307.16851, 2023. 2
  27. 27.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 3, 4
  28. 28.Ilya Loshchilov and Frank Hutter. Fixing weight decay regularization in adam. CoRR, abs/1711.05101, 2017. 6
  29. 29.Grace Luo, Trevor Darrell, and Anna Rohrbach. NewsCLIPpings: Automatic generation of out-of-context multimodal media. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6801–6817. Association for Computational Linguistics, 2021. 2, 3, 4, 5, 7
  30. 30.Jing Ma, Wei Gao, and Kam-Fai Wong. Detect rumors in microblog posts using propagation structure via kernel learning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 708–717. Association for Computational Linguistics, 2017. 3
  31. 31.Eric Muller-Budack, Jonas Theiner, Sebastian Diering, Maximilian Idahl, and Ralph Ewerth. Multimodal analytics for real-world news using measures of cross-modal entity consistency. In Proceedings of the 2020 on International Conference on Multimedia Retrieval, ICMR 2020, Dublin, Ireland, June 8-11, 2020, pages 16–25. ACM, 2020. 2, 3
  32. 32.Eric Muller-Budack, Jonas Theiner, Sebastian Diering, Maximilian Idahl, Sherzod Hakimov, and Ralph Ewerth. Multimodal news analytics using measures of cross-modal entity and context consistency. Int. J. Multim. Inf. Retr., 10(2):111–125, 2021. 8
  33. 33.OpenAI. ChatGPT. https://openai.com/blog/chatgpt/. 4
  34. 34.OpenAI. GPT-4 technical report, 2023. 2
  35. 35.OpenAI. GPT-4v(ision) system card. https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023. 8
  36. 36.Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, and Panagiotis Petrantonakis. Synthetic misinformers: Generating and combating multimodal misinformation. In Proceedings of the 2nd ACM International Workshop on Multimedia AI against Disinformation, MAD@ICMR 2023, Thessaloniki, Greece, June 12-15, 2023, pages 36–44. ACM, 2023. 2, 3, 5, 6
  37. 37.Piotr Przybyla. Capturing the style of fake news. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 490–497. AAAI Press, 2020. 3
  38. 38.Peng Qi, Juan Cao, Tianyun Yang, Junbo Guo, and Jintao Li. Exploiting multi-domain visual information for fake news detection. In 2019 IEEE International Conference on Data Mining, ICDM 2019, Beijing, China, November 8-11, 2019, pages 518–527. IEEE, 2019. 3
  39. 39.Peng Qi, Juan Cao, Xirong Li, Huan Liu, Qiang Sheng, Xiaoyue Mi, Qin He, Yongbiao Lv, Chenyang Guo, and Yingchao Yu. Improving fake news detection by using an entity-enhanced framework to fuse diverse multimodal clues. In Proceedings of the 29th ACM International Conference on Multimedia , MM ’21, Virtual Event, China, October 20 - 24, 2021, pages 1212–1220. ACM, 2021. 3
  40. 40.Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, and Tat-Seng Chua. Fakesv: A multimodal benchmark with rich social context for fake news detection on short video platforms. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, pages 14444–14452. AAAI Press, 2023. 3
  41. 41.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, pages 8748–8763. PMLR, 2021. 3, 6
  42. 42.Ekraam Sabir, Wael AbdAlmageed, Yue Wu, and Prem Natarajan. Deep multimodal image-repurposing detection. In Proceedings of the 26th ACM International Conference on Multimedia, MM 2018, Seoul, Republic of Korea, October 22-26, 2018, pages 1337–1345. ACM, 2018. 2, 3
  43. 43.Rui Shao, Tianxing Wu, and Ziwei Liu. Detecting and grounding multi-modal media manipulation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 6904–6913. IEEE, 2023. 1, 3
  44. 44.Kai Shu, Limeng Cui, Suhang Wang, Dongwon Lee, and Huan Liu. dEFEND: Explainable fake news detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019, pages 395–405. ACM, 2019. 3
  45. 45.Kai Shu, Xinyi Zhou, Suhang Wang, Reza Zafarani, and Huan Liu. The role of user profiles for fake news detection. In ASONAM ’19: International Conference on Advances in Social Networks Analysis and Mining, Vancouver, British Columbia, Canada, 27-30 August, 2019, pages 436–439. ACM, 2019. 3
  46. 46.Ruben Tolosana, Ruben Vera-Rodr ´ ıguez, Julian Fierrez, ´ Aythami Morales, and Javier Ortega-Garcia. Deepfakes and beyond: A survey of face manipulation and fake detection. Inf. Fusion, 64:131–148, 2020. 1
  47. 47.Xueyu Wang, Jiajun Huang, Siqi Ma, Surya Nepal, and Chang Xu. DeepFake disrupter: The detector of DeepFake is my friend. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 14900–14909. IEEE, 2022. 1
  48. 48.Yaqing Wang, Fenglong Ma, Zhiwei Jin, Ye Yuan, Guangxu Xun, Kishlay Jha, Lu Su, and Jing Gao. EANN: event adversarial neural networks for multi-modal fake news detection. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2018, London, UK, August 19-23, 2018, pages 849–857. ACM, 2018. 6
  49. 49.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-Instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 13484–13508. Association for Computational Linguistics, 2023. 3
  50. 50.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. A survey on multimodal large language models. CoRR, abs/2306.13549, 2023. 2
  51. 51.Jingsi Yu, Junhui Zhu, Yujie Wang, Yang Liu, Hongxiang Chang, Jinran Nie, Cunliang Kong, Ruining Cong, XinLiu, Jiyuan An, Luming Lu, Mingwei Fang, and Lin Zhu. Taoli LLaMA. https://github.com/blcuicall/taoli, 2023. 3
  52. 52.Yizhou Zhang, Loc Trinh, Defu Cao, Zijun Cui, and Yan Liu. Detecting out-of-context multimodal misinformation with interpretable neural-symbolic model. CoRR, abs/2304.07633, 2023. 3, 5, 6
  53. 53.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional Deepfake detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 2185–2194. Computer Vision Foundation / IEEE, 2021. 1
  54. 54.Xinyi Zhou, Jindi Wu, and Reza Zafarani. SAFE: similarityaware multi-modal fake news detection. In Advances in Knowledge Discovery and Data Mining - 24th Pacific-Asia Conference, PAKDD 2020, Singapore, May 11-14, 2020, Proceedings, Part II, pages 354–367. Springer, 2020. 3, 6
  55. 55.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 3
  56. 56.Dimitrina Zlatkova, Preslav Nakov, and Ivan Koychev. Factchecking meets fauxtography: Verifying claims about images. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2099–2108. Association for Computational Linguistics, 2019. 5

Citation

MLA
Qi, P., et al. “SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection”. arXiv, 2024, http://arxiv.org/abs/2403.03170v1.
APA
Qi, P., Yan, Z., Hsu, W., & Lee, M. L. (2024). SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection. arXiv. http://arxiv.org/abs/2403.03170v1
Chicago
Qi, P., Z. Yan, W. Hsu, and M. L. Lee. 2024. “SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection”. arXiv. http://arxiv.org/abs/2403.03170v1.
Harvard
Qi, P. et al. (2024) “SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.03170v1.
Vancouver
1. Qi P, Yan Z, Hsu W, Lee ML (2024) SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection. arXiv

BibTeX

@article{qi2024sniffer,
  title = {SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection},
  author = {Qi, Peng and Yan, Zehong and Hsu, Wynne and Lee, Mong Li},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.03170v1},
  eprint = {2403.03170}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE