Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal context aware attention

Multimodal context aware attention is a neural network mechanism that dynamically calculates the importance of information across multiple distinct data types, such as text, audio, and visual signals, conditioned on the broader contextual environment. Unlike basic attention methods that operate within a single modality or evaluate inputs in isolation, this approach integrates surrounding sequential, conversational, or situational context to guide cross-modal interactions. By assigning attention weights that reflect both the relationships between different modalities and their surrounding context, the mechanism enables computational models to effectively filter irrelevant data, resolve semantic ambiguities, and capture complex, context-dependent nuances during multimodal feature fusion.

1 item

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

Shivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, Tanmoy Chakraborty

OrganizationsIndraprastha Institute of Information Technology Delhi

Why you should read this

Introduces the task of sarcasm explanation in multimodal multi-party dialogues, providing the WITS benchmark dataset and a modality-aware attention framework that generates natural language explanations for sarcastic utterances.

Indirect speech such as sarcasm achieves a constellation of discourse goals in human communication. While the indirectness of figurative language warrants speakers to achieve certain pragmatic goals, it is challenging for AI agents to comprehend such idiosyncrasies of human communication. Though sarcasm identification has been a well-explored topic in dialogue analysis, for conversational systems to truly grasp a conversation’s innate meaning and generate appropriate responses, simply detecting sarcasm is not enough; it is vital to explain its underlying sarcastic connotation to capture its true essence. In this work, we study the discourse structure of sarcastic conversations and propose a novel task – Sarcasm Explanation in Dialogue (SED). Set in a multimodal and code-mixed setting, the task aims to generate natural language explanations of satirical conversations. To this end, we curate WITS, a new dataset to support our task. We propose MAF (Modality Aware Fusion), a multimodal context-aware attention and global information fusion module to capture multimodality and use it to benchmark WITS. The proposed attention module surpasses the traditional multimodal fusion baselines and reports the best performance on almost all metrics. Lastly, we carry out detailed analyses both quantitatively and qualitatively.

Added

2026-09-26