Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal sarcasm detection

Multimodal sarcasm detection is the computational task of identifying sarcastic or ironic intent by analyzing and integrating information from multiple data modalities, such as text, speech, images, and video. Sarcasm often relies on subtle contradictions between what is expressed verbally and how it is presented nonverbally, making it difficult to detect through a single modality like plain text. Multimodal sarcasm detection systems address this challenge by capturing cross-modal incongruities, aligning semantic features from written words with acoustic cues like vocal tone and visual cues like facial expressions or contrasting imagery. By synthesizing these diverse streams of information, this approach improves the comprehension of figurative language and conversational context, thereby enhancing sentiment analysis, social media monitoring, and human-computer interaction.

2 items

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

Shivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, Tanmoy Chakraborty

OrganizationsIndraprastha Institute of Information Technology Delhi

Why you should read this

Introduces the task of sarcasm explanation in multimodal multi-party dialogues, providing the WITS benchmark dataset and a modality-aware attention framework that generates natural language explanations for sarcastic utterances.

Indirect speech such as sarcasm achieves a constellation of discourse goals in human communication. While the indirectness of figurative language warrants speakers to achieve certain pragmatic goals, it is challenging for AI agents to comprehend such idiosyncrasies of human communication. Though sarcasm identification has been a well-explored topic in dialogue analysis, for conversational systems to truly grasp a conversation’s innate meaning and generate appropriate responses, simply detecting sarcasm is not enough; it is vital to explain its underlying sarcastic connotation to capture its true essence. In this work, we study the discourse structure of sarcastic conversations and propose a novel task – Sarcasm Explanation in Dialogue (SED). Set in a multimodal and code-mixed setting, the task aims to generate natural language explanations of satirical conversations. To this end, we curate WITS, a new dataset to support our task. We propose MAF (Modality Aware Fusion), a multimodal context-aware attention and global information fusion module to capture multimodality and use it to benchmark WITS. The proposed attention module surpasses the traditional multimodal fusion baselines and reports the best performance on almost all metrics. Lastly, we carry out detailed analyses both quantitatively and qualitatively.

Added

2026-09-26