Dynamic Routing Transformer Network for Multimodal Sarcasm Detection

Yuan TianNan XuRuike ZhangWenji Mao

article2023ACL84 citations

Proposes DynRT-Net, a dynamic routing transformer network that adaptively selects routing paths across hierarchical co-attention modules to capture diverse types of image-text incongruity for multimodal sarcasm detection.

Listen

Online communication frequently relies on sarcasm, where the intended meaning directly contradicts the literal words. In modern social media, sarcasm is often expressed across modalities, such as pairing a cheerful statement with a frustrating or contradictory image. Accurately detecting this multimodal sarcasm is essential for automated sentiment analysis, public opinion monitoring, and conversational systems. However, existing automated detection models rely on static, fixed neural network architectures. These static approaches lack the flexibility to handle the diverse ways sarcasm manifests, whether through contradictions between specific image details and text phrases or through broader clashes in overall tone.

The article aims to solve this limitation by introducing the Dynamic Routing Transformer Network, a framework designed to identify sarcasm by dynamically adapting its analysis path to the unique characteristics of each image-text pair.

To achieve this, the authors evaluate their method on the standard benchmark dataset for multimodal sarcasm detection, which consists of roughly 24,600 annotated image-text pairs from Twitter. The framework first extracts textual and visual features using established pre-trained base models. It then processes these features through multi-layered dynamic transformer modules. A lightweight routing mechanism inspects the input data and dynamically selects hierarchical attention patterns, enabling the model to focus progressively on relevant visual regions and textual tokens before making a final classification.

The experimental results show that the proposed dynamic method achieves state-of-the-art performance. The dynamic network achieved an accuracy of 93.49% and a balanced performance score (macro-F1) of 93.21%. This represents a statistically significant improvement over previous top-performing architectures, which reached approximately 89.67% accuracy, even when those prior methods incorporated external knowledge sources like image captions. Furthermore, ablation tests confirmed that multimodal dynamic routing is critical: replacing the dynamic path selection with static attention caused accuracy to decline by over 16 percentage points, while removing cross-modal dynamic routing entirely led to drops exceeding 24 percentage points.

These findings demonstrate that static AI architectures create computational inefficiencies and struggle with the nuanced variability of human communication. By dynamically adapting processing pathways based on input characteristics, systems can achieve higher accuracy and reduce redundant computations without requiring external knowledge bases. This performance gain directly enhances the reliability of automated brand monitoring, customer sentiment tracking, and content moderation pipelines.

Organizations developing or deploying multimodal analysis tools should adopt dynamic routing mechanisms rather than rigid, static models when processing complex, context-dependent content. For future implementation, development teams should explore expanding beyond the current fixed set of four attention masks to further improve adaptability across diverse content formats. Decision-makers should note, however, that current empirical validation is based on a single public social media benchmark dataset. Before deploying the system at scale in production environments, teams should conduct pilot evaluations on domain-specific datasets to confirm its real-world generalization across different platforms and user demographics.

Tian et al (2023).pdf
Cover for Dynamic Routing Transformer Network for Multimodal Sarcasm Detection

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Image-text Sarcasm Detection
  • 2.2 Multimodal Dynamic Networks
  • 3 Method
  • 3.1 Encoding
  • 3.2 Dynamic Routing Transformer
  • 3.2.1 Routing Space
  • 3.2.2 Dynamic Routing Transformer Layer
  • 3.2.3 Hierarchical Co-attention
  • 3.2.4 Router
  • 3.3 Classification
  • 3.4 Optimization
  • 4 Experiments
  • 4.1 Dataset
  • 4.2 Experimental Settings
  • 4.3 Baseline Methods
  • 4.4 Main Results
  • 4.5 Ablation Study
  • 4.6 Hyperparameter Analysis
  • 4.7 Case Study
  • 5 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A License of Scientific Artifacts
  • B More Details of Experimental Settings
  • ACL 2023 Responsible NLP Checklist
  • A For every submission:
  • B Did you use or create scientific artifacts?
  • C Did you run computational experiments?
  • D Did you use human annotators (e.g., crowdworkers) or research with human participants?

Knowls

  1. Knowl 1 — Dynamic Routing Transformer Network (DynRT-Net) Architecture

    model/method

    The Dynamic Routing Transformer Network (DynRT-Net) detects multimodal sarcasm by dynamically adjusting cross-modal attention paths conditioned on input image-text pairs. The framework consists of three main components: multimodal feature encoding, a dynamic routing transformer, and a multimodal classification head.

    Given a text input token sequence Text={[CLS],w1,…,wn−1}Text = \{[\text{CLS}], w_1, \dots, w_{n-1}\} of length nn and an image Image∈RH×W×CImage \in \mathbb{R}^{H \times W \times C}, the text is encoded via the pre-trained RoBERTa model to extract token embeddings T∈Rn×dtT \in \mathbb{R}^{n \times d_t}:

    T=RoBERTa(Text)=[t1,t2,…,tn]T = \text{RoBERTa}(Text) = [t_1, t_2, \dots, t_n]

    where ti∈Rdtt_i \in \mathbb{R}^{d_t} represents the embedding of the ii-th token. Concurrently, the image is partitioned into mm flattened 2D patches and encoded using a pre-trained Vision Transformer (ViT) into patch embeddings I∈Rm×dvI \in \mathbb{R}^{m \times d_v}:

    I=ViT(Image)=[e1,e2,…,em]I = \text{ViT}(Image) = [e_1, e_2, \dots, e_m]

    where ej∈Rdve_j \in \mathbb{R}^{d_v} is the embedding of the jj-th visual patch.

    The text representations are iteratively updated across KK dynamic routing transformer layers using the visual features II, yielding final routed text representations TKT_K. The global embeddings of both modalities are then pooled and classified to predict sarcastic or non-sarcastic labels.

  2. Knowl 2 — Dynamic Routing Transformer Layer and Multi-Head Co-Attention Routing

    model/method

    Each Dynamic Routing Transformer (DynRT) layer performs dynamic cross-modal interaction by routing across hierarchical co-attention mask functions. The kk-th DynRT layer (k∈[1,K]k \in [1, K]) takes the previous layer's text features Tk−1∈Rn×dtT_{k-1} \in \mathbb{R}^{n \times d_t} (with T0=TT_0 = T) and image features I∈Rm×dvI \in \mathbb{R}^{m \times d_v}, updating representations through three sub-modules with residual connections and layer normalization (LN):

    Tk−1r=LN(MHCARk(Tk−1,I)+Tk−1)T^r_{k-1} = \text{LN}(\text{MHCAR}_k(T_{k-1}, I) + T_{k-1})

    Tk−1a=LN(MHAk(Tk−1r)+Tk−1r)T^a_{k-1} = \text{LN}(\text{MHA}_k(T^r_{k-1}) + T^r_{k-1})

    Tk=LN(FFNk(Tk−1a)+Tk−1a)T_k = \text{LN}(\text{FFN}_k(T^a_{k-1}) + T^a_{k-1})

    where MHCARk\text{MHCAR}_k is the multi-head co-attention routing module, MHAk\text{MHA}_k is a multi-head self-attention module, and FFNk\text{FFN}_k is a feed-forward network.

    The MHCARk\text{MHCAR}_k module computes hh attention heads in parallel with hidden dimension dh=dt/hd_h = d_t / h, concatenated and projected via OTk∈Rdt×dtO^k_T \in \mathbb{R}^{d_t \times d_t}:

    MHCARk(Tk−1,I)=concat([headik]i=1h)OTk\text{MHCAR}_k(T_{k-1}, I) = \text{concat}\left([\text{head}_i^k]_{i=1}^h\right) O^k_T

    To optimize computation, each attention head headik∈Rn×dh\text{head}_i^k \in \mathbb{R}^{n \times d_h} aggregates attention scores across pkp_k co-attention mask matrices AjA^j weighted by dynamic routing probabilities αjk\alpha^k_j:

    headik=σ(Qi,kKi,k⊤dh⊗∑j=0pk−1αjkAj)Vik\text{head}_i^k = \sigma\left(\frac{Q_{i,k} K_{i,k}^\top}{\sqrt{d_h}} \otimes \sum_{j=0}^{p_k - 1} \alpha^k_j A^j\right) V^k_i

    where σ(⋅)\sigma(\cdot) denotes the softmax function, ⊗\otimes denotes the element-wise matrix product, Qi,k=Tk−1Wi,kQQ_{i,k} = T_{k-1} W_{i,k}^Q (Wi,kQ∈Rdt×dhW_{i,k}^Q \in \mathbb{R}^{d_t \times d_h}), Ki,k=IWi,kKK_{i,k} = I W_{i,k}^K (Wi,kK∈Rdv×dhW_{i,k}^K \in \mathbb{R}^{d_v \times d_h}), and Vi,kk=IWi,kVV_{i,k}^k = I W_{i,k}^V (Wi,kV∈Rdv×dhW_{i,k}^V \in \mathbb{R}^{d_v \times d_h}).

  3. Knowl 3 — Hierarchical Co-Attention Mask Mechanism

    model/method

    Hierarchical co-attention structures the spatial image regions visible to text tokens across transformer layers. For an ss-order sliding window with a spatial patch grid of size (2s+1)×(2s+1)(2s + 1) \times (2s + 1) traversing every image patch l∈[1,m]l \in [1, m], a binary mask vector vls∈{0,1}mv_l^s \in \{0, 1\}^m is obtained. The co-attention mask matrix As∈Rn×mA^s \in \mathbb{R}^{n \times m} is constructed by circularly stacking vlsv_l^s across all nn token positions:

    As=[v1s,v2s,…,vns]A^s = [v_1^s, v_2^s, \dots, v_n^s]

    When s=0s = 0, A0A^0 represents an empty mask matrix populated entirely with ones, allowing all text tokens (including the global [CLS][\text{CLS}] token) to attend to the entire image.

    To capture cross-modal incongruity hierarchically from coarse to fine granularity, the diversity of available co-attention masks increases progressively with layer depth. In the kk-th DynRT layer, the candidate mask group is defined as Gk=[A0,A1,…,Apk−1]G_k = [A^0, A^1, \dots, A^{p_k - 1}], where the number of mask types is set dynamically to pk=kp_k = k.

  4. Knowl 4 — Dynamic Co-Attention Router with Gumbel-Softmax

    model/method

    The routing probability vector αk=[α0k,α1k,…,αpk−1k]⊤∈Rpk\alpha^k = [\alpha^k_0, \alpha^k_1, \dots, \alpha^k_{p_k - 1}]^\top \in \mathbb{R}^{p_k} for the kk-th DynRT layer is generated dynamically conditioned on the visual features of the input image:

    αk=σg(MLP(APool(I)))∈Rpk\alpha^k = \sigma_g\left(\text{MLP}(\text{APool}(I))\right) \in \mathbb{R}^{p_k}

    where APool(⋅)\text{APool}(\cdot) denotes 1D adaptive average pooling across all mm patch embeddings in I∈Rm×dvI \in \mathbb{R}^{m \times d_v}, MLP\text{MLP} is a two-layer multilayer perceptron with hidden dimension dmd_m, and σg(⋅)\sigma_g(\cdot) denotes the Gumbel-Softmax activation function with temperature parameter tt.

    The resulting weights αk\alpha^k determine the linear combination of the pkp_k co-attention mask matrices in the kk-th layer, allowing the model to adaptively balance local visual focus versus global image context for each image-text pair.

  5. Knowl 5 — Multimodal Classification and Loss Function in DynRT-Net

    model/method

    The multimodal sarcasm prediction module aggregates the raw visual representations I∈Rm×dvI \in \mathbb{R}^{m \times d_v} and final routed textual representations TK∈Rn×dtT_K \in \mathbb{R}^{n \times d_t} through global mean pooling:

    Ig=Mean(I)∈RdvI_g = \text{Mean}(I) \in \mathbb{R}^{d_v}

    Tg=Mean(TK)∈RdtT_g = \text{Mean}(T_K) \in \mathbb{R}^{d_t}

    Assuming dv=dt=dd_v = d_t = d, the pooled unimodal representations are combined and normalized to form a multimodal embedding yg∈Rdy_g \in \mathbb{R}^d:

    yg=Wg(LN(Ig+Tg))+bgy_g = W_g\left(\text{LN}(I_g + T_g)\right) + b_g

    where Wg∈Rd×dW_g \in \mathbb{R}^{d \times d} and bg∈Rdb_g \in \mathbb{R}^d are trainable parameters, and LN(⋅)\text{LN}(\cdot) is layer normalization. The class probability distribution y^∈Rdp\hat{y} \in \mathbb{R}^{d_p} (where dp=2d_p = 2 for sarcastic versus non-sarcastic) is computed via:

    y^=Softmax(Woyg+bo)\hat{y} = \text{Softmax}(W_o y_g + b_o)

    with trainable parameters Wo∈Rdp×dW_o \in \mathbb{R}^{d_p \times d} and bo∈Rdpb_o \in \mathbb{R}^{d_p}. The model is trained end-to-end minimizing the cross-entropy loss over NN training samples:

    L=−∑i=1Nyi⊤log⁡y^i\mathcal{L} = -\sum_{i=1}^N y_i^\top \log \hat{y}_i

    where yiy_i is the ground-truth one-hot label vector for the ii-th image-text sample.

  6. Knowl 6 — Multimodal Sarcasm Detection Performance Comparison on MSD Benchmark

    data/table

    DynRT-Net was evaluated on the benchmark Multimodal Sarcasm Detection (MSD) dataset alongside unimodal and multimodal baseline methods using macro-average F1-score (F1F1) and Accuracy (AccAcc).

    Modality Method F1 Acc
    Image ResNet 61.53 64.76
    ViT 66.90 ±\pm 0.09 68.79 ±\pm 0.17
    Text TextCNN 78.15 80.03
    SIARN 79.57 80.57
    SMSD 79.51 80.90
    Bi-LSTM 80.55 81.09
    BERT 81.09 83.85
    RoBERTa 83.42 ±\pm 0.22 83.94 ±\pm 0.14
    Image + Text HFM 80.18 83.44
    DR Net 80.60 84.02
    IIMI-MMSD 82.92 86.05
    Bridge 86.05 88.51
    InCrossMGs 85.60 86.10
    MuLOT 86.33 87.41
    CMGCN 87.00 87.55
    Hmodel 88.92 ±\pm 0.51 89.34 ±\pm 0.52
    HKEmodel 89.24 ±\pm 0.24 89.67 ±\pm 0.23
    DynRT-Net 93.21 ±\pm 0.06 93.49 ±\pm 0.05

    DynRT-Net achieves 93.21%93.21\% F1 and 93.49%93.49\% accuracy, statistically significantly outperforming the state-of-the-art knowledge-enhanced model HKEmodel by +3.97%+3.97\% F1 and +3.82%+3.82\% accuracy (p<0.001p < 0.001). Text-only models consistently outperform image-only models, showing textual cues carry stronger intra-modal sarcastic signals. Multimodal approaches surpass unimodal methods, demonstrating the critical role of cross-modal interaction in identifying sarcasm.

  7. Knowl 7 — Ablation Analysis of DynRT-Net Routing and Attention Mechanisms

    data/table

    An ablation study evaluated the contribution of hierarchical mask ordering, dynamic cross-modal routing, and transformer routing structures within DynRT-Net on the MSD dataset.

    Variant F1 Acc Δ\Delta F1 Δ\Delta Acc
    DynRT-Net 93.21 93.49 - -
    DynRT-Net (pk=Kp_k = K) 91.08 91.40 -2.13 -2.09
    DynRT-Net (pk=K−k+1p_k = K - k + 1) 91.21 91.50 -2.00 -1.99
    - DynRT, + TRAR 89.67 90.07 -3.54 -3.42
    - DynRT, + Standard Transformer 87.83 88.22 -5.38 -5.27
    - DynRT, + Concatenation 66.57 68.89 -26.64 -24.60
    - Dynamic attention, + mean attention 84.91 85.44 -8.30 -8.05
    - Dynamic attention, + fixed attention 75.81 76.54 -17.40 -16.95

    Fixing the number of masks across all layers (pk=Kp_k = K) or reversing the hierarchy (pk=K−k+1p_k = K - k + 1) degrades F1 by 2.13%2.13\% and 2.00%2.00\%, validating the progressive expansion of mask candidate diversity. Replacing DynRT with single-modality dynamic routing (TRAR) reduces F1 by 3.54%3.54\%, while replacing DynRT with standard multimodal transformer layers reduces F1 by 5.38%5.38\%. Removing dynamic routing in favor of uniform mean attention scores drops F1 by 8.30%8.30\%, and fixing attention exclusively to empty masks drops F1 by 17.40%17.40\%. Direct concatenation of unimodal encoder outputs without routing layers results in a severe 26.64%26.64\% drop in F1.

  8. Knowl 8 — Impact of DynRT Layer Depth on Multimodal Sarcasm Detection

    empirical result

    Varying the number of DynRT layers KK from 1 to 6 shows that model performance increases sharply with depth from K=1K = 1 (approximately 71%71\% F1 and 74%74\% Accuracy) through K=2K = 2 (approximately 82%82\% F1 and 83%83\% Accuracy) and K=3K = 3 (approximately 91%91\% F1 and 92%92\% Accuracy), reaching peak performance at K=4K = 4 (93.21%93.21\% F1, 93.49%93.49\% Accuracy).

    Increasing the depth further to K=5K = 5 and K=6K = 6 leads to slight performance degradations (dropping to approximately 92%92\% and 91%91\% F1 respectively), indicating that additional layers encounter a performance bottleneck. Consequently, K=4K = 4 is the optimal depth for DynRT-Net.

  9. Knowl 9 — Experimental Setup and Hyperparameter Configuration for DynRT-Net

    experimental setup

    DynRT-Net was trained and evaluated on the Twitter Multimodal Sarcasm Detection (MSD) dataset, comprising 19,816 training samples (8,642 sarcastic, 11,174 non-sarcastic), 2,410 development samples (959 sarcastic, 1,451 non-sarcastic), and 2,409 test samples (959 sarcastic, 1,450 non-sarcastic). Data preprocessing removed tweets containing explicit cue words (e.g., sarcasm, sarcastic, irony, ironic, humor) and URLs, and replaced user mentions with <user>.

    Images were resized to 224×224224 \times 224 pixels and processed with vit-base-patch32-224 into m=49m = 49 visual patches (7×77 \times 7 grid, embedding dimension dv=768d_v = 768). Texts were tokenized to a maximum length of n=100n = 100 and embedded using the first layer of roberta-base (dt=768d_t = 768). The dynamic routing transformer was configured with K=4K = 4 layers, h=2h = 2 attention heads in MHCAR, MLP hidden dimension dm=384d_m = 384, multimodal embedding dimension d=768d = 768, Gumbel-Softmax temperature t=10t = 10, and classifier dropout rate 0.50.5.

    Optimization was performed using Adam with a learning rate of 10−610^{-6}, weight decay of 0.010.01, and batch size of 32 for 15 epochs on GeForce RTX 2080 Ti GPUs (training time ~40 minutes; total parameter count: 238,289,140). Evaluation metrics are macro-averaged over five random runs selecting the checkpoint with the highest macro-F1 on the development set.

  10. Knowl 10 — Limitations of DynRT-Net

    limitation

    DynRT-Net has two primary limitations:

    1. The design of the hierarchical co-attention masks is currently restricted to four sliding window types, which constrains the granularity and adaptability of cross-modal incongruity modeling.
    2. Model evaluation is conducted exclusively on the MSD dataset because it is the sole publicly available benchmark for multimodal sarcasm detection, limiting empirical verification of the method's cross-domain generalization.

Coverage note — None was omitted; all significant contributed methods, formulas, architectural designs, experimental comparisons, ablation studies, hyperparameter analyses, setup details, and limitations have been extracted as knowls.

References

  1. 1.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. Computing Research Repository, arXiv:1607.06450.
  2. 2.Yitao Cai, Huiyu Cai, and Xiaojun Wan. 2019. Multi-modal sarcasm detection in twitter with hierarchical fusion model. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 2506–2515.
  3. 3.Harm de Vries, Florian Strub, Jeremie Mary, Hugo Larochelle, Olivier Pietquin, and Aaron C Courville. 2017. Modulating early visual processing by language. In Proceedings of the International Conference on Neural Information Processing Systems, pages 6597–6607.
  4. 4.Jacob Devlin, Mingwei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics, pages 4171–4186.
  5. 5.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations, pages 1–22.
  6. 6.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778.
  7. 7.Aditya Joshi, Pushpak Bhattacharyya, and Mark J Carman. 2017. Automatic sarcasm detection: A survey. ACM Computing Surveys, 50(5):1–22.
  8. 8.Yoon Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1746–1751.
  9. 9.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations, pages 1–15.
  10. 10.Bin Liang, Chenwei Lou, Xiang Li, Lin Gui, Min Yang, and Ruifeng Xu. 2021. Multi-modal sarcasm detection with interactive in-modal and cross-modal graphs. In Proceedings of the ACM International Conference on Multimedia, pages 4707–4715.
  11. 11.Bin Liang, Chenwei Lou, Xiang Li, Min Yang, Lin Gui, Yulan He, Wenjie Pei, and Ruifeng Xu. 2022. Multi-modal sarcasm detection via cross-modal graph convolutional network. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1767–1777.
  12. 12.Hui Liu, Wenya Wang, and Haoliang Li. 2022. Towards multi-modal sarcasm detection via hierarchical congruity modeling with knowledge enhancement. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 4995–5006.
  13. 13.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pre-training approach. Computing Research Repository, arXiv:1907.11692.
  14. 14.Hongliang Pan, Zheng Lin, Peng Fu, Yatao Qi, and Weiping Wang. 2020. Modeling intra and inter-modality incongruity for multi-modal sarcasm detection. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 1383–1392.
  15. 15.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 3942–3951.
  16. 16.Shraman Pramanick, Aniket Roy, and Vishal M. Patel Johns. 2022. Multimodal learning using optimal transport for sarcasm and humor detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Visio, pages 546–556.
  17. 17.Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. 2021. Dynamic modality interaction modeling for image-text retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, page 1104–1113.
  18. 18.Ellen Riloff, Ashequl Qadir, Prafulla Surve, Lalindra De Silva, Nathan Gilbert, and Ruihong Huang. 2013. Sarcasm as contrast between a positive sentiment and negative situation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 704–714.
  19. 19.Rossano Schifanella, Paloma de Juan, Joel Tetreault, and Liangliang Cao. 2016. Detecting sarcasm in multimodal social platforms. In Proceedings of the ACM International Conference on Multimedia, pages 1136–1145.
  20. 20.Yi Tay, Anh Tuan Luu, Siu Cheung Hui, and Jian Su. 2018. Reasoning with sarcasm by reading in-between. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 1010–1020.
  21. 21.Joseph Tepperman, David Traum, and Shrikanth Narayanan. 2006. "Yeah right": Sarcasm recognition for spoken dialogue systems. In Proceedings of the International Conference on Spoken Language Processing, pages 1838–1841.
  22. 22.Oren Tsur, Dmitry Davidov, and Ari Rappoport. 2010. Icwsm—a great catchy name: Semi-supervised recognition of sarcastic sentences in online product reviews. In Proceedings of the International AAAI Conference on Weblogs and Social Media, pages 162–169.
  23. 23.Xinyu Wang, Xiaowen Sun, Tan Yang, and Hongbo Wang. 2020. Building a bridge: A method for image-text sarcasm detection without pretraining on image-text data. In Proceedings of the International Workshop on Natural Language Processing Beyond Text, pages 19–29.
  24. 24.Tao Xiong, Peiran Zhang, Hongbo Zhu, and Yihui Yang. 2019. Sarcasm detection with self-matching networks and low-rank bilinear pooling. In Proceedings of the World Wide Web Conference, pages 2115–2124.
  25. 25.Nan Xu, Zhixiong Zeng, and Wenji Mao. 2020. Reasoning with multimodal sarcastic tweets via modeling cross-modality contrast and semantic association. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 3777–3786.
  26. 26.Meishan Zhang, Yue Zhang, and Guohong Fu. 2016. Tweet sarcasm detection using deep neural network. In Proceedings of the International Conference on Computational Linguistics, pages 2449–2460.
  27. 27.Yiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu, and Rongrong Ji. 2021. Trar: Routing the attention spans in transformer for visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Visio, pages 2074–2084.

Citation

MLA
Tian, Y., et al. “Dynamic Routing Transformer Network for Multimodal Sarcasm Detection”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 2468–80, https://doi.org/10.18653/v1/2023.acl-long.139.
APA
Tian, Y., Xu, N., Zhang, R., & Mao, W. (2023). Dynamic Routing Transformer Network for Multimodal Sarcasm Detection. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2468–2480. https://doi.org/10.18653/v1/2023.acl-long.139
Chicago
Tian, Y., N. Xu, R. Zhang, and W. Mao. 2023. “Dynamic Routing Transformer Network for Multimodal Sarcasm Detection”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2468–80. https://doi.org/10.18653/v1/2023.acl-long.139.
Harvard
Tian, Y. et al. (2023) “Dynamic Routing Transformer Network for Multimodal Sarcasm Detection”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2468–2480. Available at: https://doi.org/10.18653/v1/2023.acl-long.139.
Vancouver
1. Tian Y, Xu N, Zhang R, Mao W (2023) Dynamic Routing Transformer Network for Multimodal Sarcasm Detection. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2468–2480

BibTeX

@inproceedings{tian-etal-2023-dynamic,
    title = "Dynamic Routing Transformer Network for Multimodal Sarcasm Detection",
    author = "Tian, Yuan  and
      Xu, Nan  and
      Zhang, Ruike  and
      Mao, Wenji",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.139/",
    doi = "10.18653/v1/2023.acl-long.139",
    pages = "2468--2480"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/