Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection

Yang QiaoLiqiang JingXuemeng SongXiaolin ChenLei ZhuLiqiang Nie

article2023AAAI99 citations

Presents a multi-modal sarcasm detection network that combines local graph-based semantic reasoning with global cross-attention fusion to capture incongruities across text and images while using mutual learning to transfer knowledge between the two perspectives.

Listen

Online communication increasingly relies on multimedia posts combining images and text. Sarcasm detection in these posts is critical for accurate sentiment analysis, opinion mining, and automated customer service. However, identifying sarcasm is challenging because the intended meaning directly contradicts the literal words. Previous automated systems typically analyzed either the entire image or focused strictly on isolated detected objects. Global approaches often pick up irrelevant background noise, while local object-based approaches miss broader contextual scenes or fail when objects fall outside pre-trained recognition categories.

The article demonstrates a novel framework called the Mutual-Enhanced Incongruity Learning Network (MILNet) that unifies fine-grained local object relationships with global contextual imagery. The objective is to accurately identify sarcastic contradictions within and across textual and visual elements.

The authors designed a multi-step neural network architecture evaluated on an established benchmark dataset of 24,635 English social media posts. The system encodes text and any embedded image text using modern language models, while processing images through both object detectors and global vision transformers. It maps detailed connections using a local module that links words and visual objects via an external knowledge graph and spatial overlap measurements. Concurrently, a global module extracts broad image-text interactions through an attention mechanism. A mutual learning strategy then enables both modules to continuously exchange reliable insights during training using a selective sample-screening mechanism.

The evaluation produced several significant findings. First, MILNet established a new performance benchmark on the dataset, achieving an overall accuracy of 89.50% and a macro-average F1-score of 89.12%, outperforming all existing single-modal and multi-modal baselines with statistical significance. Second, multi-modal methods consistently outperformed text-only and image-only approaches, proving that analyzing image-text incongruity is essential. Third, ablation testing showed that both local semantic modeling and global context modeling are necessary; removing either component degraded performance. Finally, the mutual enhancement module and its sample screening mechanism contributed directly to accuracy gains, confirming that transferring only confident, accurate knowledge between modules prevents error propagation.

These findings indicate that effective automated understanding of nuanced social media content requires combining fine-grained object analysis with overall scene context rather than treating them as separate alternatives. For organizations relying on social listening, customer feedback analysis, or automated moderation, adopting such hybrid multi-modal architectures can significantly reduce misclassification risks, protect brand sentiment tracking from misleading literal interpretations, and improve automated response quality.

Organizations developing or deploying emotion and opinion analysis tools should integrate multi-modal architectures that incorporate embedded image text and external knowledge graphs. For future technical development, practitioners should focus on refining cross-modal alignment methods and expanding evaluation beyond single-platform English datasets to ensure robust generalization across diverse, real-world social platforms.

Cover for Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection

Abstract

Sarcasm is a sophisticated linguistic phenomenon that is prevalent on today’s social media platforms. Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic. This task’s key lies in capturing both inter- and intra-modal incongruities within the same context. Although existing methods have achieved compelling success, they are disturbed by irrelevant information extracted from the whole image or text, or overlooking some important information due to the incomplete input. To address these limitations, we propose a Mutual-enhanced Incongruity Learning Network for multi-modal sarcasm detection, named MILNet. In particular, we design a local semantic-guided incongruity learning module and a global incongruity learning module. Moreover, we introduce a mutual enhancement module to take advantage of the underlying consistency between the two modules to boost the performance. Extensive experiments on a widely-used dataset demonstrate the superiority of our model over cutting-edge methods.

Table of Contents

  • Introduction
  • Related Work
  • Methodology
  • Problem Formulation
  • MILNet
  • Experiment
  • Experimental Settings
  • On Model Comparison (RQ1)
  • On Ablation Study (RQ2)
  • On Case Study (RQ3)
  • Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Multi-Modal Feature Encoding in MILNet

    model/method

    In the Mutual-enhanced Incongruity Learning Network (MILNet) for multi-modal sarcasm detection, multi-modal input comprising a text sentence T={t1,t2,…,tng}T = \{t_1, t_2, \dots, t_{n_g}\}, an image II, and optical character recognition text extracted from the image O={o1,o2,…,ono}O = \{o_1, o_2, \dots, o_{n_o}\} is encoded into textual, object-level visual, and global image-level representations.

    Text Encoding: The original text TT and OCR text OO are concatenated (T⊕OT \oplus O) and passed through a pre-trained RoBERTa model to produce token hidden states:

    Ht=[h1t,h2t,…,hnt]=RoBERTa(T⊕O)H^t = [h^t_1, h^t_2, \dots, h^t_n] = \text{RoBERTa}(T \oplus O)

    where hjt∈Rdhh^t_j \in \mathbb{R}^{d_h} and n=ng+non = n_g + n_o.

    Object-Level Image Encoding: Faster-RCNN extracts features for the top-kk detected regions with highest confidence. For the jj-th region, it outputs a visual feature vj∈Rdvv_j \in \mathbb{R}^{d_v}, a bounding-box positional feature pj∈Rdpp_j \in \mathbb{R}^{d_p}, an object category label eje_j, and an object attribute label aja_j. The visual and spatial features are fused via linear projections:

    hjf=Wf(Wvvj+pjWp)+bfh^f_j = W_f (W_v v_j + p_j W_p) + b_f

    where Wv∈Rdv×dvW_v \in \mathbb{R}^{d_v \times d_v}, Wp∈Rdp×dvW_p \in \mathbb{R}^{d_p \times d_v}, Wf∈Rdh×dvW_f \in \mathbb{R}^{d_h \times d_v}, and bf∈Rdhb_f \in \mathbb{R}^{d_h}. The textual labels eje_j and aja_j are embedded into hje,hja∈Rdhh^e_j, h^a_j \in \mathbb{R}^{d_h} via RoBERTa. The representation of the jj-th region IjI_j and the overall object-level image representation HovH^v_o are defined as:

    Ij=[hje,hja,hjf]⊤∈R3×dh,Hov=[I1,I2,…,Ik]∈R3k×dhI_j = [h^e_j, h^a_j, h^f_j]^\top \in \mathbb{R}^{3 \times d_h}, \quad H^v_o = [I_1, I_2, \dots, I_k] \in \mathbb{R}^{3k \times d_h}

    Global Image-Level Encoding: A pre-trained Vision Transformer (ViT) patch embedding layer divides the image II into rr non-overlapping patches, producing patch embeddings and a global [CLS][\text{CLS}] token embedding:

    Hmv=[hclsm,h1m,h2m,…,hrm]=ViT_PE(I)∈R(r+1)×dhH^v_m = [h^m_{cls}, h^m_1, h^m_2, \dots, h^m_r] = \text{ViT\_PE}(I) \in \mathbb{R}^{(r+1) \times d_h}

  2. Knowl 2 — Local Semantic-Guided Graph Construction

    model/method

    The Local Semantic-guided Incongruity Learning (LIL) module of MILNet constructs three distinct graphs to model intra-modal and inter-modal relationships:

    1. Text-Modal Graph (GtG^t): Contains nn nodes corresponding to the tokens of the input text and OCR text, initialized with G1t=HtG^t_1 = H^t. The adjacency matrix At∈Rn×nA^t \in \mathbb{R}^{n \times n} is undirected with self-loops (Aiit=1A^t_{ii} = 1). An edge Aijt=1A^t_{ij} = 1 exists if token tit_i and token tjt_j share a dependency relation D(ti,tj)D(t_i, t_j) in the dependency parse tree of the original text sentence (excluding OCR tokens from dependency parsing to avoid unreliable parses).

    2. Image-Modal Graph (GvG^v): Contains 3k3k nodes corresponding to the class, attribute, and fused visual/spatial features for the kk detected regions, initialized with G1v={h1e,h1a,h1f,…,hke,hka,hkf}G^v_1 = \{h^e_1, h^a_1, h^f_1, \dots, h^e_k, h^a_k, h^f_k\}. The adjacency matrix Av∈R3k×3kA^v \in \mathbb{R}^{3k \times 3k} connects nodes describing the same object and links different objects via spatial Intersection over Union (IoU) scores:

    Aijv={1,if i mod 3=j mod 3Si,j,if i rem 3=1 and j rem 3=10,otherwiseA^v_{ij} = \begin{cases} 1, & \text{if } i \bmod 3 = j \bmod 3 \\ S_{i,j}, & \text{if } i \text{ rem } 3 = 1 \text{ and } j \text{ rem } 3 = 1 \\ 0, & \text{otherwise} \end{cases}

    where Si,jS_{i,j} is the IoU score between the (i mod 3)(i \bmod 3)-th and (j mod 3)(j \bmod 3)-th object bounding boxes.

    1. Cross-Modal Graph (GcG^c): Contains n+3kn + 3k nodes combining all text tokens and object region nodes, initialized with G1c={h1t,…,hnt,h1e,h1a,h1f,…,hke,hka,hkf}G^c_1 = \{h^t_1, \dots, h^t_n, h^e_1, h^a_1, h^f_1, \dots, h^e_k, h^a_k, h^f_k\}. Its adjacency matrix Ac∈R(n+3k)×(n+3k)A^c \in \mathbb{R}^{(n+3k) \times (n+3k)} incorporates the intra-modal edges from AtA^t and AvA^v, and adds cross-modal edges connecting textual token tit_i to object class ee or attribute aa if a semantic relationship exists between them in the ConceptNet5 knowledge graph:

    Aijc={Aijt,if Aijt>0,  i,j∈[1,n]Aijv,if Aijv>0,  i,j∈[n+1,n+3k]1,if K(ti,e(j−n−1)/3+1) or K(ti,a(j−n−2)/3+1)0,otherwiseA^c_{ij} = \begin{cases} A^t_{ij}, & \text{if } A^t_{ij} > 0, \; i, j \in [1, n] \\ A^v_{ij}, & \text{if } A^v_{ij} > 0, \; i, j \in [n+1, n+3k] \\ 1, & \text{if } K(t_i, e_{(j-n-1)/3+1}) \text{ or } K(t_i, a_{(j-n-2)/3+1}) \\ 0, & \text{otherwise} \end{cases}

  3. Knowl 3 — Iterative Graph Convolution and Retrieval-Based Attention in LIL

    model/method

    In the Local Semantic-guided Incongruity Learning (LIL) module, Graph Convolutional Networks (GCNs) iteratively update node representations across the text-modal graph GtG^t, image-modal graph GvG^v, and cross-modal graph GcG^c over UU iterations:

    {Gu′t=ReLU(A~tGu−1tWut+but)Gu′v=ReLU(A~vGu−1vWuv+buv)Guc=Gut⊕Guv=ReLU(A~c(Gu′t⊕Gu′v)Wuc+buc)\begin{cases} G^t_{u'} = \text{ReLU}(\tilde{A}^t G^t_{u-1} W^t_u + b^t_u) \\ G^v_{u'} = \text{ReLU}(\tilde{A}^v G^v_{u-1} W^v_u + b^v_u) \\ G^c_u = G^t_u \oplus G^v_u = \text{ReLU}(\tilde{A}^c (G^t_{u'} \oplus G^v_{u'}) W^c_u + b^c_u) \end{cases}

    where A~x=(Dx)−12Ax(Dx)−12\tilde{A}^x = (D^x)^{-\frac{1}{2}} A^x (D^x)^{-\frac{1}{2}} is the normalized symmetric adjacency matrix, DxD^x is the degree matrix of AxA^x for x∈{t,v,c}x \in \{t, v, c\}, and {Wut,Wuv,Wuc}∈Rdh×dh\{W^t_u, W^v_u, W^c_u\} \in \mathbb{R}^{d_h \times d_h} and {but,buv,buc}∈Rdh\{b^t_u, b^v_u, b^c_u\} \in \mathbb{R}^{d_h} are trainable parameters.

    After LL GCN layers, a retrieval-based attention mechanism computes attention scores αi\alpha_i by matching initial node representations H={v1,…,vn+3k}H = \{v_1, \dots, v_{n+3k}\} with final GCN representations GLc={g1,…,gn+3k}G^c_L = \{g_1, \dots, g_{n+3k}\}:

    αi=softmax(∑j=1n+3kvj⊤gj),fl=∑i=1n+3kαihi\alpha_i = \text{softmax}\left(\sum_{j=1}^{n+3k} v_j^\top g_j\right), \quad f_l = \sum_{i=1}^{n+3k} \alpha_i h_i

    The aggregated representation flf_l is passed to a classifier to yield class probability distribution pl∈R2p_l \in \mathbb{R}^2:

    pl=softmax(Wlfl+bl)p_l = \text{softmax}(W_l f_l + b_l)

    where Wl∈Rd×2W_l \in \mathbb{R}^{d \times 2} and bl∈R2b_l \in \mathbb{R}^2. The module is trained with supervised cross-entropy loss:

    Lcel=ylog⁡pl+(1−y)log⁡(1−pl)+λ1∥Θl∥22\mathcal{L}^l_{ce} = y \log p_l + (1 - y) \log(1 - p_l) + \lambda_1 \|\Theta_l\|_2^2

    where y∈{0,1}y \in \{0, 1\} is the ground truth label, Θl\Theta_l denotes the trainable parameters of LIL, and λ1\lambda_1 is an L2-regularization coefficient.

  4. Knowl 4 — Global Incongruity Learning (GIL) Module

    model/method

    The Global Incongruity Learning (GIL) module captures incongruities between the global image context and the textual context by using multi-head cross-attention. The image-level ViT representation Hmv∈R(r+1)×dhH^v_m \in \mathbb{R}^{(r+1) \times d_h} serves as the query, and the encoded textual representation Ht∈Rn×dhH^t \in \mathbb{R}^{n \times d_h} serves as the key and value:

    Q=HmvWQ,K=HtWK,V=HtWVH′=softmax(KQ⊤dhV)H^=LN(H′+Hmv)Hˇ=LN(H^+(WMH^+bM))\begin{aligned} Q &= H^v_m W^Q, \quad K = H^t W^K, \quad V = H^t W^V \\ H' &= \text{softmax}\left(\frac{K Q^\top}{\sqrt{d_h}} V\right) \\ \hat{H} &= \text{LN}(H' + H^v_m) \\ \check{H} &= \text{LN}\left(\hat{H} + (W^M \hat{H} + b^M)\right) \end{aligned}

    where WQ,WK,WV,WM∈Rdh×dhW^Q, W^K, W^V, W^M \in \mathbb{R}^{d_h \times d_h}, bM∈Rdhb^M \in \mathbb{R}^{d_h}, and LN(⋅)\text{LN}(\cdot) denotes layer normalization.

    A stack of BB such attention layers is applied. The hidden representation corresponding to the [CLS][\text{CLS}] token from the final layer output HˇB\check{H}^B, denoted fgf_g, serves as the global sarcasm representation. The prediction distribution pg∈R2p_g \in \mathbb{R}^2 and classification loss Lceg\mathcal{L}^g_{ce} are computed as:

    pg=softmax(Wgfg+bg)p_g = \text{softmax}(W_g f_g + b_g)

    Lceg=ylog⁡pg+(1−y)log⁡(1−pg)+λ2∥Θg∥22\mathcal{L}^g_{ce} = y \log p_g + (1 - y) \log(1 - p_g) + \lambda_2 \|\Theta_g\|_2^2

    where Wg∈Rd×2W_g \in \mathbb{R}^{d \times 2}, bg∈R2b_g \in \mathbb{R}^2, Θg\Theta_g represents trainable parameters of the GIL module, and λ2\lambda_2 is a weight coefficient.

  5. Knowl 5 — Screened Mutual Enhancement Mechanism

    model/method

    To exploit intrinsic consistency between the local semantic-guided incongruity learning (LIL) module and global incongruity learning (GIL) module, MILNet employs mutual learning with a sample screening mechanism that prevents the transfer of incorrect predictions.

    Knowledge transfer between the two probability distributions plp_l and pgp_g is formulated via Kullback-Leibler (KL) divergence, gated by binary indicator variables η1\eta_1 and η2\eta_2:

    Lklg→l=η1DKL(pg∥pl),Lkll→g=η2DKL(pl∥pg)\mathcal{L}^{g \to l}_{kl} = \eta_1 D_{KL}(p_g \parallel p_l), \quad \mathcal{L}^{l \to g}_{kl} = \eta_2 D_{KL}(p_l \parallel p_g)

    where the transfer gates η1\eta_1 and η2\eta_2 are defined conditioned on whether each module's prediction matches the ground truth label YY:

    η1={1,if argmax⁡(pg)=Y0,otherwise,η2={1,if argmax⁡(pl)=Y0,otherwise\eta_1 = \begin{cases} 1, & \text{if } \operatorname{argmax}(p_g) = Y \\ 0, & \text{otherwise} \end{cases}, \quad \eta_2 = \begin{cases} 1, & \text{if } \operatorname{argmax}(p_l) = Y \\ 0, & \text{otherwise} \end{cases}

    The probability distributions are sharpened using a temperature parameter τ\tau prior to computing the KL divergence. This ensures knowledge is only distilled from a module when that module's prediction is correct.

  6. Knowl 6 — Joint Loss Formulation and Inference in MILNet

    equation

    The overall training objectives for the local semantic-guided incongruity learning (LIL) branch and the global incongruity learning (GIL) branch in MILNet are formulated by combining supervised cross-entropy losses with the screened mutual learning losses:

    {Ll=Lcel+δ1Lklg→lLg=Lceg+δ2Lkll→g\begin{cases} \mathcal{L}^l = \mathcal{L}^l_{ce} + \delta_1 \mathcal{L}^{g \to l}_{kl} \\ \mathcal{L}^g = \mathcal{L}^g_{ce} + \delta_2 \mathcal{L}^{l \to g}_{kl} \end{cases}

    where δ1,δ2≥0\delta_1, \delta_2 \ge 0 are weighting hyperparameters, Lcel\mathcal{L}^l_{ce} and Lceg\mathcal{L}^g_{ce} are the branch cross-entropy losses, and Lklg→l\mathcal{L}^{g \to l}_{kl} and Lkll→g\mathcal{L}^{l \to g}_{kl} are the screened bidirectional KL divergence losses.

    During inference, the final binary sarcasm prediction Y^∈{0,1}\hat{Y} \in \{0, 1\} is computed by averaging the predicted probability distributions of the two modules:

    Y^=argmax⁡(pg+pl2)\hat{Y} = \operatorname{argmax}\left(\frac{p_g + p_l}{2}\right)

  7. Knowl 7 — Multi-Modal Sarcasm Detection Benchmark Performance

    data/table

    The proposed MILNet model was evaluated on the benchmark Twitter multi-modal sarcasm detection dataset (Cai et al. 2019), comprising 19,816 training samples, 2,410 validation samples, and 2,409 testing samples. The table below presents the performance comparison against single-modal baselines (text-only or image-only) and multi-modal baselines:

    Modality Method Acc (%) F1-score Macro-average
    Pre (%) Rec (%) F1 (%) Pre (%) Rec (%) F1 (%)
    Single-modal Image 64.76 54.41 70.80 61.53 60.12 73.08 65.97
    ViT 73.72 65.56 71.64 68.46 72.78 73.37 72.97
    TextCNN 80.03 74.29 76.39 75.32 78.03 78.28 78.15
    Bi-LSTM 81.90 76.66 78.42 77.53 80.97 80.13 80.55
    SIARN 80.57 75.55 75.70 75.63 80.34 78.81 79.57
    SMSD 80.90 76.46 75.18 75.82 80.87 78.20 79.51
    BERT 83.85 78.72 82.27 80.22 81.31 80.87 81.09
    RoBERTa 85.51 78.24 88.11 82.88 84.83 85.95 85.16
    Multi-modal HFM 83.44 76.57 84.15 80.18 79.40 82.45 80.90
    DR Net 84.02 77.97 83.42 80.60 - - -
    Res-BERT 84.80 77.80 84.15 80.85 78.87 84.46 81.57
    Att-BERT 86.05 78.63 83.31 80.90 80.87 85.08 82.92
    InCrossMGs 86.10 81.38 84.36 82.84 85.39 85.80 85.60
    CMGCN 87.55 83.63 84.69 84.16 87.02 86.97 87.00
    MILNet 89.50 85.16 89.16 87.11 88.88 89.44 89.12

    MILNet achieves the best performance across all evaluation metrics, outperforming both global-incongruity models (e.g., InCrossMGs at 86.10% accuracy) and local semantic-guided models (e.g., CMGCN at 87.55% accuracy). Significance testing between MILNet and the second-best results yielded p<0.01p < 0.01 across all metrics.

  8. Knowl 8 — Ablation Study of MILNet Components

    data/table

    An ablation study evaluated the specific impact of the architectural components of MILNet on the benchmark Twitter sarcasm dataset:

    Model Variant Acc. (%) F1 (%) Macro-F1 (%)
    LIL-only 88.00 84.70 87.41
    GIL-only 88.25 85.49 87.81
    w/o-embedded-text 89.29 86.67 88.86
    w/o-mutual-learning 88.95 86.62 88.61
    w/o-sample-screening 88.96 86.57 88.60
    w/o-text-modal-relations 88.05 85.80 87.74
    w/o-image-modal-relations 89.12 86.85 88.79
    w/o-cross-modal-relations 88.96 86.65 88.62
    w/-similarity 89.00 86.81 88.69
    MILNet (Full) 89.50 87.11 89.12

    Key observations:

    1. Both LIL-only (88.00%) and GIL-only (88.25%) perform better than prior state-of-the-art baselines like CMGCN (87.55%), but combining them in MILNet yields higher performance (89.50%), demonstrating their complementarity.
    2. Removing mutual learning (w/o-mutual-learning) or disabling sample screening (w/o-sample-screening) reduces accuracy to 88.95% and 88.96%, confirming that selective bidirectional knowledge sharing boosts detection.
    3. Discarding intra-/inter-modal relations (w/o-text-modal-relations, w/o-image-modal-relations, w/o-cross-modal-relations) degrades performance, with text dependency parsing having the largest individual impact (accuracy drops to 88.05%).
    4. Replacing ConceptNet5 knowledge-graph relations with lexical word similarity (w/-similarity) drops accuracy from 89.50% to 89.00%, showing the benefit of structured commonsense relations over surface similarity.

Coverage note — None was omitted.

References

  1. 1.Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR, 6077–6086.
  2. 2.Ba, L. J.; Kiros, J. R.; and Hinton, G. E. 2016. Layer Normalization. CoRR, abs/1607.06450.
  3. 3.Cai, Y.; Cai, H.; and Wan, X. 2019. Multi-Modal Sarcasm Detection in Twitter with Hierarchical Fusion Model. In ACL, 2506–2515.
  4. 4.Cui, H.; Zhu, L.; Li, J.; Yang, Y.; and Nie, L. 2019. Scalable Deep Hashing for Large-scale Social Image Retrieval. IEEE Transactions on Image Processing, 29: 1271–1284.
  5. 5.Dai, J.; Yan, H.; Sun, T.; Liu, P.; and Qiu, X. 2021. Does syntax matter? A Strong Baseline for Aspect-based Sentiment Analysis with RoBERTa. In NAACL-HLT, 1816–1829.
  6. 6.Davidov, D.; Tsur, O.; and Rappoport, A. 2010. SemiSupervised Recognition of Sarcasm in Twitter and Amazon. In CoNLL, 107–116.
  7. 7.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT, 4171–4186.
  8. 8.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR.
  9. 9.González-Ibáñez, R. I.; Muresan, S.; and Wacholder, N. 2011. Identifying Sarcasm in Twitter: A Closer Look. In ACL, 581–586.
  10. 10.He, D.; Liang, C.; Liu, H.; Wen, M.; Jiao, P.; and Feng, Z. 2022. Block Modeling-Guided Graph Convolutional Neural Networks. In AAAI, 4022–4029.
  11. 11.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In CVPR, 770–778.
  12. 12.Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531.
  13. 13.Jing, L.; Tian, M.; Chen, X.; Sun, T.; Guan, W.; and Song, X. 2022. CI-OCM: Counterfactural Inference towards Unbiased Outfit Compatibility Modeling. In Proceedings of the 1st Workshop on Multimedia Computing towards Fashion Recommendation, 31–38.
  14. 14.Kim, Y. 2014. Convolutional Neural Networks for Sentence Classification. In ACL, 1746–1751.
  15. 15.Kipf, T. N.; and Welling, M. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In ICLR.
  16. 16.Liang, B.; Lou, C.; Li, X.; Gui, L.; Yang, M.; and Xu, R. 2021. Multi-Modal Sarcasm Detection with Interactive InModal and Cross-Modal Graphs. In MM, 4707–4715.
  17. 17.Liang, B.; Lou, C.; Li, X.; Yang, M.; Gui, L.; He, Y.; Pei, W.; and Xu, R. 2022. Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network. In ACL, 1767–1777.
  18. 18.Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, abs/1907.11692.
  19. 19.Lu, X.; Zhu, L.; Cheng, Z.; Nie, L.; and Zhang, H. 2019. Online Multi-modal Hashing with Dynamic Query-adaption. In SIGIR, 715–724.
  20. 20.Pan, H.; Lin, Z.; Fu, P.; Qi, Y.; and Wang, W. 2020. Modeling Intra and Inter-modality Incongruity for Multi-Modal Sarcasm Detection. In Findings of EMNLP, 1383–1392.
  21. 21.Poria, S.; Cambria, E.; Hazarika, D.; and Vij, P. 2016. A Deeper Look into Sarcastic Tweets Using Deep Convolutional Neural Networks. In COLING, 1601–1612.
  22. 22.Riloff, E.; Qadir, A.; Surve, P.; Silva, L. D.; Gilbert, N.; and Huang, R. 2013. Sarcasm as Contrast between a Positive Sentiment and Negative Situation. In EMNLP, 704–714.
  23. 23.Schifanella, R.; de Juan, P.; Tetreault, J. R.; and Cao, L. 2016. Detecting Sarcasm in Multimodal Social Platforms. In MM, 1136–1145.
  24. 24.Song, X.; Feng, F.; Liu, J.; Li, Z.; Nie, L.; and Ma, J. 2017. NeuroStylist: Neural Compatibility Modeling for Clothing Matching. In MM, 753–761.
  25. 25.Sun, T.; Wang, W.; Jing, L.; Cui, Y.; Song, X.; and Nie, L. 2022. Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment Analysis. In MM, 15–23.
  26. 26.Tay, Y.; Luu, A. T.; Hui, S. C.; and Su, J. 2018. Reasoning with Sarcasm by Reading In-Between. In ACL, 1010–1020.
  27. 27.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In NIPS, 5998–6008.
  28. 28.Wang, T.; Jin, D.; Wang, R.; He, D.; and Huang, Y. 2022a. Powerful Graph Convolutional Networks with Adaptive Propagation Mechanism for Homophily and Heterophily. In AAAI, 4210–4218.
  29. 29.Wang, X.; Sun, X.; Yang, T.; and Wang, H. 2020. Building a bridge: A Method for Image-text Sarcasm Detection without Pretraining on Image-text Data. In Proceedings of the first international workshop on natural language processing beyond text, 19–29.
  30. 30.Wang, Y.; Cao, M.; Fan, Z.; and Peng, S. 2022b. Learning to Detect 3D Facial Landmarks via Heatmap Regression with Graph Convolutional Network. In AAAI, 2595–2603.
  31. 31.Wei, Y.; Wang, X.; Guan, W.; Nie, L.; Lin, Z.; and Chen, B. 2019a. Neural Multimodal Cooperative Learning toward Micro-video Understanding. IEEE Transactions on Image Processing, 29: 1–14.
  32. 32.Wei, Y.; Wang, X.; Nie, L.; He, X.; Hong, R.; and Chua, T.-S. 2019b. MMGCN: Multi-modal Graph Convolution Network for Personalized Recommendation of Micro-video. In MM, 1437–1445.
  33. 33.Wen, H.; Song, X.; Yang, X.; Zhan, Y.; and Nie, L. 2021. Comprehensive Linguistic-Visual Composition Network for Image Retrieval. In SIGIR, 1369–1378.
  34. 34.Xiong, T.; Zhang, P.; Zhu, H.; and Yang, Y. 2019. Sarcasm Detection with Self-matching Networks and Low-rank Bilinear Pooling. In WWW, 2115–2124.
  35. 35.Xu, N.; Zeng, Z.; and Mao, W. 2020. Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic Association. In ACL, 3777–3786.
  36. 36.Zhang, C.; Li, Q.; and Song, D. 2019. Aspect-based Sentiment Classification with Aspect-specific Graph Convolutional Networks. In EMNLP-IJCNLP, 4567–4577.
  37. 37.Zhang, M.; Zhang, Y.; and Fu, G. 2016. Tweet Sarcasm Detection Using Deep Neural Network. In COLING, 2449–2460.
  38. 38.Zhang, Y.; Xiang, T.; Hospedales, T. M.; and Lu, H. 2018. Deep Mutual Learning. In CVPR, 4320–4328.
  39. 39.Zhuang, J.; and Hasan, M. A. 2022. Defending Graph Convolutional Networks against Dynamic Graph Perturbations via Bayesian Self-Supervision. In AAAI, 4405–4413.

Citation

MLA
Qiao, Y., et al. “Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 8, 2023, pp. 9507–15, https://doi.org/10.1609/AAAI.V37I8.26138.
APA
Qiao, Y., Jing, L., Song, X., Chen, X., Zhu, L., & Nie, L. (2023). Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection. Proceedings of the AAAI Conference on Artificial Intelligence, 37(8), 9507–9515. https://doi.org/10.1609/AAAI.V37I8.26138
Chicago
Qiao, Y., L. Jing, X. Song, X. Chen, L. Zhu, and L. Nie. 2023. “Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (8): 9507–15. https://doi.org/10.1609/AAAI.V37I8.26138.
Harvard
Qiao, Y. et al. (2023) “Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(8), pp. 9507–9515. Available at: https://doi.org/10.1609/AAAI.V37I8.26138.
Vancouver
1. Qiao Y, Jing L, Song X, Chen X, Zhu L, Nie L (2023) Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection. Proceedings of the AAAI Conference on Artificial Intelligence 37:9507–9515

BibTeX

@article{Qiao_2023, title={Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I8.26138}, DOI={10.1609/aaai.v37i8.26138}, number={8}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Qiao, Yang and Jing, Liqiang and Song, Xuemeng and Chen, Xiaolin and Zhu, Lei and Nie, Liqiang}, year={2023}, month=June, pages={9507–9515} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF