Dynamic Routing Transformer Network for Multimodal Sarcasm Detection
Yuan TianNan XuRuike ZhangWenji Mao
Proposes DynRT-Net, a dynamic routing transformer network that adaptively selects routing paths across hierarchical co-attention modules to capture diverse types of image-text incongruity for multimodal sarcasm detection.
Online communication frequently relies on sarcasm, where the intended meaning directly contradicts the literal words. In modern social media, sarcasm is often expressed across modalities, such as pairing a cheerful statement with a frustrating or contradictory image. Accurately detecting this multimodal sarcasm is essential for automated sentiment analysis, public opinion monitoring, and conversational systems. However, existing automated detection models rely on static, fixed neural network architectures. These static approaches lack the flexibility to handle the diverse ways sarcasm manifests, whether through contradictions between specific image details and text phrases or through broader clashes in overall tone.
The article aims to solve this limitation by introducing the Dynamic Routing Transformer Network, a framework designed to identify sarcasm by dynamically adapting its analysis path to the unique characteristics of each image-text pair.
To achieve this, the authors evaluate their method on the standard benchmark dataset for multimodal sarcasm detection, which consists of roughly 24,600 annotated image-text pairs from Twitter. The framework first extracts textual and visual features using established pre-trained base models. It then processes these features through multi-layered dynamic transformer modules. A lightweight routing mechanism inspects the input data and dynamically selects hierarchical attention patterns, enabling the model to focus progressively on relevant visual regions and textual tokens before making a final classification.
The experimental results show that the proposed dynamic method achieves state-of-the-art performance. The dynamic network achieved an accuracy of 93.49% and a balanced performance score (macro-F1) of 93.21%. This represents a statistically significant improvement over previous top-performing architectures, which reached approximately 89.67% accuracy, even when those prior methods incorporated external knowledge sources like image captions. Furthermore, ablation tests confirmed that multimodal dynamic routing is critical: replacing the dynamic path selection with static attention caused accuracy to decline by over 16 percentage points, while removing cross-modal dynamic routing entirely led to drops exceeding 24 percentage points.
These findings demonstrate that static AI architectures create computational inefficiencies and struggle with the nuanced variability of human communication. By dynamically adapting processing pathways based on input characteristics, systems can achieve higher accuracy and reduce redundant computations without requiring external knowledge bases. This performance gain directly enhances the reliability of automated brand monitoring, customer sentiment tracking, and content moderation pipelines.
Organizations developing or deploying multimodal analysis tools should adopt dynamic routing mechanisms rather than rigid, static models when processing complex, context-dependent content. For future implementation, development teams should explore expanding beyond the current fixed set of four attention masks to further improve adaptability across diverse content formats. Decision-makers should note, however, that current empirical validation is based on a single public social media benchmark dataset. Before deploying the system at scale in production environments, teams should conduct pilot evaluations on domain-specific datasets to confirm its real-world generalization across different platforms and user demographics.
- Paper: Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network, Bin Liang et al. (2022). This paper establishes the multimodal sarcasm detection benchmark task and models cross-modal incongruities between image regions and text tokens that dynamic routing aims to process adaptively.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). It introduces a foundational baseline for cross-modal self-attention over visual and textual tokens, providing the standard transformer architecture that dynamic routing builds upon.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). It introduces the crossmodal transformer mechanism for unaligned multimodal sequences, establishing the core cross-attention paradigm adapted by multimodal sarcasm detectors.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). It provides foundational principles for separating modality-invariant and modality-specific representations in multimodal sentiment and humor analysis.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). It establishes dynamic pathway selection and efficacy routing across language, visual, and acoustic channels for multimodal sentiment analysis.
- Paper: Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection, Yang Qiao et al. (2023). This work directly extends multimodal sarcasm detection on the same Twitter benchmark by coupling fine-grained local object relationships with global contextual imagery via mutual incongruity learning.
- Paper: Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, David Raposo et al. (2024). This book generalizes dynamic token routing in transformers by routing sequence tokens along the depth dimension to optimize computational allocation.
- Paper: PMR: Prototypical Modal Rebalance for Multimodal Learning, Yunfeng Fan et al. (2023). This paper addresses modality imbalance during multimodal training, offering optimization techniques to ensure slower-learning modalities are not suppressed in cross-modal networks.
- Paper: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest, Jack Hessel et al. (2023). This work explores higher-level visual-linguistic incongruity and humor understanding benchmarks, providing a broader evaluation framework for multimodal figurative language.
