Built independently by an author, for readers. Read the story and support ChapterPal

keyword

intra-modal incongruities

Intra-modal incongruities refer to internal contradictions, semantic conflicts, or clashing sentiment orientations that occur entirely within a single communication or data channel, such as within text alone or within an image alone. In multimodal computing and natural language processing tasks like sarcasm and sentiment analysis, intra-modal incongruity manifests when different components of the same medium express incompatible meanings, such as juxtaposing positive and negative words in a single sentence or presenting conflicting visual cues within a single photograph. This phenomenon is distinct from inter-modal incongruity, which involves discrepancies that arise across different modalities, such as a contradiction between a written caption and an accompanying picture.

1 item

Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection

Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection

Yang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen, Lei Zhu, Liqiang Nie

OrganizationsHarbin Institute of TechnologySchool of Computer Science and TechnologySchool of Information Science and EngineeringSchool of SoftwareShandong Normal UniversityShandong University

Why you should read this

Presents a multi-modal sarcasm detection network that combines local graph-based semantic reasoning with global cross-attention fusion to capture incongruities across text and images while using mutual learning to transfer knowledge between the two perspectives.

Sarcasm is a sophisticated linguistic phenomenon that is prevalent on today’s social media platforms. Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic. This task’s key lies in capturing both inter- and intra-modal incongruities within the same context. Although existing methods have achieved compelling success, they are disturbed by irrelevant information extracted from the whole image or text, or overlooking some important information due to the incomplete input. To address these limitations, we propose a Mutual-enhanced Incongruity Learning Network for multi-modal sarcasm detection, named MILNet. In particular, we design a local semantic-guided incongruity learning module and a global incongruity learning module. Moreover, we introduce a mutual enhancement module to take advantage of the underlying consistency between the two modules to boost the performance. Extensive experiments on a widely-used dataset demonstrate the superiority of our model over cutting-edge methods.

Added

2026-09-26