Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network
Bin LiangChenwei LouXiang LiMin YangLin GuiYulan HeWenjie PeiRuifeng Xu
Proposes a cross-modal graph convolutional network that bridges detected image objects and text with affective knowledge to explicitly model cross-modal sentiment incongruity for state-of-the-art multi-modal sarcasm detection.
Online communication frequently combines text and images, making automated sentiment analysis and opinion mining increasingly challenging. Sarcasm poses a particular difficulty because the expressed message often conveys the direct opposite of its literal meaning. While human observers detect sarcasm by noticing contradictions between words and visual context—such as a caption praising "wonderful weather" paired with a photo of a storm—existing automated systems often analyze entire images uniformly or fail to connect specific visual details with contradictory text cues. The article introduces and evaluates a new computational framework, the Cross-Modal Graph Convolutional Network (CMGCN), designed to accurately detect multi-modal sarcasm by explicitly modeling conflicting sentiment relationships between specific image regions and text tokens.
To address this challenge, the approach identifies key visual objects within an image alongside descriptive attribute-object pairs. These descriptors are linked to sentence words in a cross-modal network graph. The connections between visual objects and text words are weighted based on lexical similarity from WordNet and affective sentiment scores from SenticNet, which amplify connections when visual and textual cues exhibit opposing emotional polarities. A two-layer graph convolutional network, paired with an attention mechanism, processes these cross-modal relationships alongside sentence grammatical structures to determine whether a post is sarcastic. The model was evaluated on a benchmark dataset of 24,635 English Twitter posts containing paired text and images.
The findings show that CMGCN establishes a new state of the art in multi-modal sarcasm detection. First, the proposed model achieved an overall accuracy of 87.55% and an F1-score of 84.16%, outperforming all prior unimodal and multi-modal baselines by statistically significant margins. Second, multi-modal methods consistently outperformed single-modality models, though text-only models (accuracy up to 83.85%) proved significantly more predictive than image-only models (accuracy below 68%), confirming that textual cues carry the primary sarcastic signals while images provide crucial contextual disambiguation. Third, ablation experiments revealed that removing the cross-modal graph entirely caused accuracy to drop to 84.12%, while omitting object detection reduced accuracy to 84.55%, confirming that targeting specific visual regions rather than full images is vital. Finally, the framework demonstrated strong generalizability across different underlying language and image embedding models.
These results demonstrate that automated systems can reliably identify complex figurative language by explicitly linking focused visual components to affective textual sentiment. For organizations relying on social listening, brand monitoring, and public sentiment analysis, incorporating cross-modal contradiction detection reduces the operational risk of misclassifying negative or satirical consumer feedback as positive engagement. System architects and engineering leaders should consider integrating region-based visual detection and sentiment knowledge bases into existing multi-modal analysis pipelines rather than relying solely on monolithic image processing.
For next steps, practitioners should evaluate the framework on domain-specific data, while researchers should develop methods to automatically infer cross-modal graph weights without depending on static external knowledge bases. A key limitation of the study is its reliance on predefined external resources (WordNet, SenticNet, and syntactic dependency parsers), which may hinder deployment in low-resource languages or informal text genres where such linguistic tools are unavailable. Nevertheless, given the stable performance observed across ten randomized experimental runs, confidence in the reported performance gains remains high for standard English multi-modal data streams.
- Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). Introduces fine-grained latent semantic alignment between image regions and text tokens, providing the cross-modal attention foundation needed to model incongruity in sarcasm detection.
- Paper: Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering, Peter Anderson et al. (2018). Establishes object-level visual feature extraction using region proposals and text-guided attention, directly underpinning the source's object-detection and cross-modal bridging approach.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). Demonstrates how to frame text processing as graph convolutional node representation learning, which the source extends into heterogeneous cross-modal graphs.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Presents foundational self-attention architectures for aligning region features with textual words across vision-and-language tasks.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). Pioneers message passing over visual and contextual graph structures to capture inter-object semantic relations essential for structured multimodal reasoning.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). Provides key techniques for learning multimodal representations and handling modality discrepancies in affective computing tasks like sentiment and humor analysis.
- Paper: Graph Neural Networks: A Review of Methods and Applications, Jie Zhou et al. (2018). Surveys core graph convolutional network architectures and message-passing mechanisms that the source adapts for cross-modal relational reasoning.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). Offers a comprehensive taxonomy of multimodal representation, alignment, and fusion challenges central to cross-modal understanding.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). Extends multimodal incongruity and humor understanding by leveraging advanced large language model backbones paired with visual representations.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Surveys recent multimodal large language model architectures that supersede task-specific graph networks for high-level cross-modal reasoning.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). Evaluates unified multimodal models across complex integrated capabilities, including visual humor comprehension and cross-modal reasoning.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Advances step-by-step visual reasoning by incorporating chain-of-thought grounding over fine-grained image regions.
