Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network

Bin LiangChenwei LouXiang LiMin YangLin GuiYulan HeWenjie PeiRuifeng Xu

article2022ACL172 citations

Proposes a cross-modal graph convolutional network that bridges detected image objects and text with affective knowledge to explicitly model cross-modal sentiment incongruity for state-of-the-art multi-modal sarcasm detection.

Listen

Online communication frequently combines text and images, making automated sentiment analysis and opinion mining increasingly challenging. Sarcasm poses a particular difficulty because the expressed message often conveys the direct opposite of its literal meaning. While human observers detect sarcasm by noticing contradictions between words and visual context—such as a caption praising "wonderful weather" paired with a photo of a storm—existing automated systems often analyze entire images uniformly or fail to connect specific visual details with contradictory text cues. The article introduces and evaluates a new computational framework, the Cross-Modal Graph Convolutional Network (CMGCN), designed to accurately detect multi-modal sarcasm by explicitly modeling conflicting sentiment relationships between specific image regions and text tokens.

To address this challenge, the approach identifies key visual objects within an image alongside descriptive attribute-object pairs. These descriptors are linked to sentence words in a cross-modal network graph. The connections between visual objects and text words are weighted based on lexical similarity from WordNet and affective sentiment scores from SenticNet, which amplify connections when visual and textual cues exhibit opposing emotional polarities. A two-layer graph convolutional network, paired with an attention mechanism, processes these cross-modal relationships alongside sentence grammatical structures to determine whether a post is sarcastic. The model was evaluated on a benchmark dataset of 24,635 English Twitter posts containing paired text and images.

The findings show that CMGCN establishes a new state of the art in multi-modal sarcasm detection. First, the proposed model achieved an overall accuracy of 87.55% and an F1-score of 84.16%, outperforming all prior unimodal and multi-modal baselines by statistically significant margins. Second, multi-modal methods consistently outperformed single-modality models, though text-only models (accuracy up to 83.85%) proved significantly more predictive than image-only models (accuracy below 68%), confirming that textual cues carry the primary sarcastic signals while images provide crucial contextual disambiguation. Third, ablation experiments revealed that removing the cross-modal graph entirely caused accuracy to drop to 84.12%, while omitting object detection reduced accuracy to 84.55%, confirming that targeting specific visual regions rather than full images is vital. Finally, the framework demonstrated strong generalizability across different underlying language and image embedding models.

These results demonstrate that automated systems can reliably identify complex figurative language by explicitly linking focused visual components to affective textual sentiment. For organizations relying on social listening, brand monitoring, and public sentiment analysis, incorporating cross-modal contradiction detection reduces the operational risk of misclassifying negative or satirical consumer feedback as positive engagement. System architects and engineering leaders should consider integrating region-based visual detection and sentiment knowledge bases into existing multi-modal analysis pipelines rather than relying solely on monolithic image processing.

For next steps, practitioners should evaluate the framework on domain-specific data, while researchers should develop methods to automatically infer cross-modal graph weights without depending on static external knowledge bases. A key limitation of the study is its reliance on predefined external resources (WordNet, SenticNet, and syntactic dependency parsers), which may hinder deployment in low-resource languages or informal text genres where such linguistic tools are unavailable. Nevertheless, given the stable performance observed across ten randomized experimental runs, confidence in the reported performance gains remains high for standard English multi-modal data streams.

Cover for Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network

Abstract

With the increasing popularity of posting multimodal messages online, many recent studies have been carried out utilizing both textual and visual information for multi-modal sarcasm detection. In this paper, we investigate multi-modal sarcasm detection from a novel perspective by constructing a cross-modal graph for each instance to explicitly draw the ironic relations between textual and visual modalities. Specifically, we first detect the objects paired with descriptions of the image modality, enabling the learning of important visual information. Then, the descriptions of the objects are served as a bridge to determine the importance of the association between the objects of image modality and the contextual words of text modality, so as to build a cross-modal graph for each multi-modal instance. Furthermore, we devise a cross-modal graph convolutional network to make sense of the incongruity relations between modalities for multi-modal sarcasm detection. Extensive experimental results and in-depth analysis show that our model achieves state-of-the-art performance in multi-modal sarcasm detection.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multi-modal Sarcasm Detection
  • 2.2 Graph Neural Networks
  • 3 Methodology
  • 3.1 Text-modality Representation
  • 3.2 Image-modality Representation
  • 3.3 Cross-modal Graph
  • 3.4 Multi-modal Fusion
  • 3.5 Learning Objective
  • 4 Experimental Setup
  • 4.1 Dataset
  • 4.2 Experimental Settings
  • 4.3 Comparison Models
  • 5 Experimental Results
  • 5.1 Main Results
  • 5.2 Ablation Study
  • 5.4 Impact of GCN Layers
  • 5.5 Visualization
  • 6 Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Cross-Modal Graph Construction with Affective and Semantic Weighting

    equation

    In the Cross-Modal Graph Convolutional Network (CMGCN), a multi-modal instance containing nn textual tokens and mm detected visual regions is modeled as an undirected cross-modal graph with adjacency matrix A∈R(n+m)×(n+m)A \in \mathbb{R}^{(n+m) \times (n+m)}. The edge weights between textual nodes and visual object nodes explicitly encode semantic and affective incongruity:

    Ai,j={1if Di,j is true and i<n,j<nκi,jif i<n,j≥n1if i=j0otherwiseA_{i,j} = \begin{cases} 1 & \text{if } D_{i,j} \text{ is true and } i < n, j < n \\ \kappa_{i,j} & \text{if } i < n, j \ge n \\ 1 & \text{if } i = j \\ 0 & \text{otherwise} \end{cases}

    where Aj,i=Ai,jA_{j,i} = A_{i,j} ensures an undirected graph. Di,j∈{0,1}D_{i,j} \in \{0, 1\} indicates whether a syntactic dependency relation connects words wiw_i and wjw_j in the sentence dependency parse tree. For a text token wiw_i and an image object jj described by an attribute-object pair (aj,oj)(a_j, o_j), the cross-modal edge weight κi,j\kappa_{i,j} and the sentiment incongruity modulating factor ξi,j\xi_{i,j} are computed as:

    κi,j=Sim(wi,oj)×ξi,j+1\kappa_{i,j} = \text{Sim}(w_i, o_j) \times \xi_{i,j} + 1

    ξi,j=γ−ω(wi)ω(aj)×∣ω(wi)−ω(aj)∣\xi_{i,j} = \gamma^{-\omega(w_i)\omega(a_j)} \times |\omega(w_i) - \omega(a_j)|

    where:

    • Sim(wi,oj)∈[0,1]\text{Sim}(w_i, o_j) \in [0, 1] represents the semantic similarity between word wiw_i and object noun ojo_j derived from WordNet path distance (evaluating to 00 if no path exists).
    • ω(⋅)∈[−1,1]\omega(\cdot) \in [-1, 1] is the sentiment intensity retrieved from SenticNet, with ω(w)=0\omega(w) = 0 if word ww is absent from the lexicon.
    • aja_j is the attribute descriptor (typically an affective adjective) and ojo_j is the object entity identified for bounding box jj.
    • γ>1\gamma > 1 is a scaling hyperparameter (set to γ=3\gamma = 3) that amplifies the modulating factor when polarities ω(wi)\omega(w_i) and ω(aj)\omega(a_j) have opposite signs (incongruent sentiment) and dampens it when their polarities match.
  2. Knowl 2 — Cross-Modal Graph Convolution and Retrieval-Based Attention Fusion

    model/method

    To learn multi-modal sarcastic representations from the constructed cross-modal graph, CMGCN applies multi-layer Graph Convolutional Networks (GCN) followed by a graph-guided retrieval attention mechanism.

    Let R=[t1,…,tn,v1,…,vm]∈R(n+m)×2dhR = [t_1, \dots, t_n, v_1, \dots, v_m] \in \mathbb{R}^{(n+m) \times 2d_h} represent the concatenated sequence of nn text hidden vectors and mm visual region vectors. The node representations are iteratively propagated across l=1,…,Ll = 1, \dots, L GCN layers via:

    Gl=ReLU(A~Gl−1Wl+bl)G_l = \text{ReLU}(\tilde{A} G_{l-1} W_l + b_l)

    where G0=RG_0 = R, A~=D−12AD−12\tilde{A} = D^{-\frac{1}{2}} A D^{-\frac{1}{2}} is the normalized symmetric adjacency matrix, Dii=∑jAi,jD_{ii} = \sum_j A_{i,j} is the degree matrix, and Wl∈R2dh×2dh,bl∈R2dhW_l \in \mathbb{R}^{2d_h \times 2d_h}, b_l \in \mathbb{R}^{2d_h} are trainable parameters.

    To aggregate information specifically from nodes participating in cross-modal interactions, the model computes retrieval-based attention weights over the initial features rt∈Rr_t \in R guided by the final GCN representations gi∈GLg_i \in G_L:

    αt=exp⁡(βt)∑i=1n+mexp⁡(βi),βt=∑i∈Crt⊤gi\alpha_t = \frac{\exp(\beta_t)}{\sum_{i=1}^{n+m} \exp(\beta_i)}, \quad \beta_t = \sum_{i \in \mathcal{C}} r_t^\top g_i

    where C\mathcal{C} denotes the subset of node indices that have non-zero cross-modal edges.

    The final sarcastic instance representation f∈R2dhf \in \mathbb{R}^{2d_h} is the attention-weighted sum:

    f=∑t=1n+mαtrtf = \sum_{t=1}^{n+m} \alpha_t r_t

    Classification into label distribution y^∈Rdp\hat{y} \in \mathbb{R}^{d_p} (dp=2d_p = 2 for binary sarcasm detection) is performed via a linear layer with softmax:

    y^=softmax(Wof+bo)\hat{y} = \text{softmax}(W_o f + b_o)

    with trainable parameters Wo∈Rdp×2dhW_o \in \mathbb{R}^{d_p \times 2d_h} and bo∈Rdpb_o \in \mathbb{R}^{d_p}, optimized using cross-entropy loss with L2L_2 regularization.

  3. Knowl 3 — Textual and Visual Modality Encoding Pipeline in CMGCN

    model/method

    CMGCN encodes raw textual and visual inputs into aligned hidden feature spaces of dimension 2dh2d_h before graph construction:

    1. Textual Feature Extraction: A sentence s={wi}i=1ns = \{w_i\}_{i=1}^n is passed through a pre-trained uncased BERT-base model to produce contextual embeddings XT=[x1,…,xn]∈Rn×dTX^T = [x_1, \dots, x_n] \in \mathbb{R}^{n \times d^T} (excluding special [CLS] and [SEP] tokens). A bidirectional LSTM (Bi-LSTM) processes XTX^T to produce contextual text hidden states:

    T={t1,t2,…,tn}=Bi-LSTM(XT),tj∈R2dhT = \{t_1, t_2, \dots, t_n\} = \text{Bi-LSTM}(X^T), \quad t_j \in \mathbb{R}^{2d_h}

    1. Visual Region Feature Extraction: A bottom-up object detection model identifies mm salient bounding box regions {I1,…,Im}\{I_1, \dots, I_m\}, each paired with an attribute-object description pair (aj,oj)(a_j, o_j). Each bounding box region IiI_i is resized to 224×224224 \times 224 and split into r=p×p=49r = p \times p = 49 non-overlapping patches (p=7p = 7, patch size 32×3232 \times 32). Each patch is projected linearly to dimension dId^I, prepended with a learnable [class] token z[class]z_{\text{[class]}}, and summed with positional embeddings Epos∈R(r+1)×dIE_{\text{pos}} \in \mathbb{R}^{(r+1) \times d^I} to form Zi∈R(r+1)×dIZ_i \in \mathbb{R}^{(r+1) \times d^I}. The sequence is processed by a Vision Transformer (ViT):

    Hi=ViT(Zi),hi=Hi,[class]∈RdIH_i = \text{ViT}(Z_i), \quad h_i = H_{i,\text{[class]}} \in \mathbb{R}^{d^I}

    The visual region vectors XI={h1,…,hm}X^I = \{h_1, \dots, h_m\} are mapped into the text hidden dimension using a trainable projection matrix WV∈RdI×2dhW_V \in \mathbb{R}^{d^I \times 2d_h}:

    V={v1,…,vm}=XIWV,vi∈R2dhV = \{v_1, \dots, v_m\} = X^I W_V, \quad v_i \in \mathbb{R}^{2d_h}

  4. Knowl 4 — Performance Comparison of CMGCN Against Baseline Models on Multimodal Sarcasm Detection

    data/table

    Experiments conducted on the Twitter multimodal sarcasm benchmark dataset collected by Cai et al. (2019) demonstrate that CMGCN outperforms unimodal and existing multimodal methods across Accuracy, Precision, Recall, and F1 metrics.

    Modality / Method Acc (%) F1-score Macro-average
    Pre (%) Rec (%) F1 (%) Pre (%) Rec (%) F1 (%)
    Image-only
    Image (ResNet) 64.76 54.41 70.80 61.53 60.12 73.08 65.97
    ViT 67.83 57.93 70.07 63.43 65.68 71.35 68.40
    Text-only
    TextCNN 80.03 74.29 76.39 75.32 78.03 78.28 78.15
    Bi-LSTM 81.90 76.66 78.42 77.53 80.97 80.13 80.55
    SIARN 80.57 75.55 75.70 75.63 80.34 78.81 79.57
    SMSD 80.90 76.46 75.18 75.82 80.87 78.20 79.51
    BERT 83.85 78.72 82.27 80.22 81.31 80.87 81.09
    Image+Text
    HFM 83.44 76.57 84.15 80.18 79.40 82.45 80.90
    DR Net 84.02 77.97 83.42 80.60 - - -
    Res-BERT 84.80 77.80 84.15 80.85 78.87 84.46 81.57
    Att-BERT 86.05 78.63 83.31 80.90 80.87 85.08 82.92
    InCrossMGs 86.10 81.38 84.36 82.84 85.39 85.80 85.60
    CMGCN (ours) 87.55 83.63 84.69 84.16 87.02 86.97 87.00

    Key takeaways:

    1. Text-modality baselines strongly outperform image-modality baselines (e.g., BERT achieves 83.85% Acc vs. ViT's 67.83%), indicating sarcasm cues reside predominantly in text.
    2. Multi-modal models outperform unimodal models, confirming the utility of combining visual and textual features.
    3. CMGCN significantly outperforms prior graph-based multimodal architectures (InCrossMGs at 86.10% Acc and 82.84% F1) with an Accuracy of 87.55% and F1 of 84.16% (p<0.05p < 0.05 significance).
  5. Knowl 5 — Ablation Analysis of CMGCN Components

    data/table

    An ablation study isolates the contributions of the cross-modal graph, object detection, external semantic/affective knowledge sources, and syntactic dependency parsing in CMGCN on the Cai et al. (2019) benchmark dataset.

    Model Variant Acc. (%) F1 (%) Macro-F1 (%)
    CMGCN (full model) 87.55 84.16 87.00
    w/o G (no cross-modal graph, direct concatenation) 84.12 80.64 81.47
    w/o O (no object detection, whole image input, uniform weights) 84.55 81.09 82.31
    w/o S (no external knowledge: all edge weights set to 1) 85.63 81.82 83.28
    w/o SwS^w (no SenticNet affective knowledge) 86.54 82.73 84.76
    w/o D (no syntactic dependency tree relations) 87.25 83.64 86.13
    • Removing the cross-modal graph (w/o G) causes the largest drop in Accuracy (−3.43%-3.43\%) and Macro-F1 (−5.53%-5.53\%), verifying that structured cross-modal information aggregation is central to model performance.
    • Omitting object detection (w/o O) reduces Accuracy to 84.55%, showing that isolating local visual objects is more effective than processing the entire image uniformly.
    • Removing SenticNet sentiment weighting (w/o S^w) and all external similarity/affective knowledge (w/o S) reduces F1 by 1.43%1.43\% and 2.34%2.34\% respectively, demonstrating the necessity of affective incongruity weighting.
  6. Knowl 6 — Impact of GCN Depth and Encoder Backbone Choice on CMGCN

    empirical result

    An analysis of the structural and architectural design choices of CMGCN reveals:

    1. GCN Layer Depth: When varying the number of GCN layers from 1 to 6, a 2-layer GCN architecture achieves the highest performance (Accuracy 87.55%87.55\%, F1 84.16%84.16\%, Macro-F1 87.00%87.00\%). A 1-layer GCN yields substantially lower performance (Accuracy ≈84.24%\approx 84.24\%, F1 ≈81.10%\approx 81.10\%), indicating that single-layer propagation cannot adequately model complex multi-modal interactions. Depths beyond 2 layers exhibit monotonic performance degradation (dropping below 85%85\% Accuracy at 6 layers) due to parameter increase and graph over-smoothing.

    2. Encoder Backbone Generalizability: Incorporating the cross-modal graph improves performance across various text and vision backbones compared to concatenation without the graph:

      • GloVe + ResNet-152: Accuracy ≈83.0%\approx 83.0\%, F1 ≈79.2%\approx 79.2\%
      • GloVe + ViT: Accuracy ≈84.3%\approx 84.3\%, F1 ≈80.8%\approx 80.8\%
      • BERT + ResNet-152: Accuracy ≈85.8%\approx 85.8\%, F1 ≈82.4%\approx 82.4\%
      • BERT + ViT (default CMGCN): Accuracy 87.55%87.55\%, F1 84.16%84.16\% In every backbone configuration, graph-based fusion outperforms the non-graph baseline (w/o G).
  7. Knowl 7 — Limitations of Affective Lexicon and Dependency Parser Dependency in CMGCN

    limitation

    The construction of the cross-modal graph in CMGCN relies strictly on three external resources: WordNet for lexical path similarity, SenticNet for affective intensity scores ω(w)∈[−1,1]\omega(w) \in [-1, 1], and spaCy for sentence syntactic dependency trees. Consequently, the method encounters difficulties when transferring to low-resource languages, non-standard text domains, or task genres where comprehensive affective knowledge bases and reliable syntactic dependency parsers are unavailable.

Coverage note — None was omitted; all key contributions including graph formulation, modality encoders, fusion mechanism, main benchmark results, ablation analysis, parameter sensitivities, and model limitations are covered.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 6077–6086. IEEE Computer Society.
  2. 2.Nastaran Babanejad, Heidar Davoudi, Aijun An, and Manos Papagelis. 2020. Affective and contextual embedding for sarcasm detection. In Proceedings of the 28th International Conference on Computational Linguistics, pages 225–243, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  3. 3.Yitao Cai, Huiyu Cai, and Xiaojun Wan. 2019. Multi-modal sarcasm detection in Twitter with hierarchical fusion model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2506–2515, Florence, Italy. Association for Computational Linguistics.
  4. 4.Erik Cambria, Yang Li, Frank Z. Xing, Soujanya Poria, and Kenneth Kwok. 2020. Senticnet 6: Ensemble application of symbolic and subsymbolic AI for sentiment analysis. In CIKM ’20: The 29th ACM International Conference on Information and Knowledge Management, Virtual Event, Ireland, October 19-23, 2020, pages 105–114. ACM.
  5. 5.Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards multimodal sarcasm detection (an Obviously perfect paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629, Florence, Italy. Association for Computational Linguistics.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Shelly Dews and Ellen Winner. 1995. Muting the meaning a social function of irony. Metaphor and Symbol, 10(1):3–19.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  9. 9.Raymond W Gibbs. 1986. On the psycholinguistics of sarcasm. Journal of experimental psychology: general, 115(1):3.
  10. 10.Raymond W Gibbs. 2007. On the psycholinguistics of sarcasm. Irony in language and thougt: A cognitive science reader, pages 173–200.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770–778. IEEE Computer Society.
  12. 12.Aditya Joshi, Vinita Sharma, and Pushpak Bhattacharyya. 2015. Harnessing context incongruity for sarcasm detection. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 757–762, Beijing, China. Association for Computational Linguistics.
  13. 13.Yoon Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1746–1751, Doha, Qatar. Association for Computational Linguistics.
  14. 14.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  15. 15.Amit Kumar Jena, Aman Sinha, and Rohit Agarwal. 2020. C-net: Contextual network for sarcasm detection. In Proceedings of the Second Workshop on Figurative Language Processing, pages 61–66, Online. Association for Computational Linguistics.
  16. 16.Bin Liang, Chenwei Lou, Xiang Li, Lin Gui, Min Yang, and Ruifeng Xu. 2021a. Multi-modal sarcasm detection with interactive in-modal and cross-modal graphs. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4707–4715. Association for Computing Machinery.
  17. 17.Bin Liang, Hang Su, Lin Gui, Erik Cambria, and Ruifeng Xu. 2022. Aspect-based sentiment analysis via affective knowledge enhanced graph convolutional networks. Knowledge-Based Systems, 235:107643.
  18. 18.Bin Liang, Hang Su, Rongdi Yin, Lin Gui, Min Yang, Qin Zhao, Xiaoqi Yu, and Ruifeng Xu. 2021b. Beta distribution guided aspect-aware graph for aspect category sentiment analysis with affective knowledge. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 208–218, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Chenwei Lou, Bin Liang, Lin Gui, Yulan He, Yixue Dang, and Ruifeng Xu. 2021. Affective dependency graph for sarcasm detection. In the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21), pages 1844–1849.
  20. 20.George A. Miller. 1992. WordNet: A lexical database for English. In Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992.
  21. 21.Hongliang Pan, Zheng Lin, Peng Fu, Yatao Qi, and Weiping Wang. 2020. Modeling intra and inter-modality incongruity for multi-modal sarcasm detection. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1383–1392, Online. Association for Computational Linguistics.
  22. 22.Bo Pang and Lillian Lee. 2008. Opinion mining and sentiment analysis. Information Retrieval, 2(1-2):1–135.
  23. 23.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  24. 24.Ellen Riloff, Ashequl Qadir, Prafulla Surve, Lalindra De Silva, Nathan Gilbert, and Ruihong Huang. 2013. Sarcasm as contrast between a positive sentiment and negative situation. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 704–714, Seattle, Washington, USA. Association for Computational Linguistics.
  25. 25.Rossano Schifanella, Paloma de Juan, Joel R. Tetreault, and Liangliang Cao. 2016. Detecting sarcasm in multimodal social platforms. In Proceedings of the 2016 ACM Conference on Multimedia Conference, MM 2016, Amsterdam, The Netherlands, October 15-19, 2016, pages 1136–1145.
  26. 26.Qiaoyu Tan, Ninghao Liu, Xing Zhao, Hongxia Yang, Jingren Zhou, and Xia Hu. 2020. Learning to hash with graph neural networks for recommender systems. In WWW ’20: The Web Conference 2020, Taipei, Taiwan, April 20-24, 2020, pages 1988–1998. ACM / IW3C2.
  27. 27.Yi Tay, Anh Tuan Luu, Siu Cheung Hui, and Jian Su. 2018. Reasoning with sarcasm by reading in-between. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1010–1020, Melbourne, Australia. Association for Computational Linguistics.
  28. 28.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph attention networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  29. 29.Jianchao Wu, Limin Wang, Li Wang, Jie Guo, and Gangshan Wu. 2019. Learning actor relation graphs for group activity recognition. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 9964–9974. Computer Vision Foundation / IEEE.
  30. 30.Guo-Sen Xie, Jie Liu, Huan Xiong, and Ling Shao. 2021. Scale-aware graph neural network for few-shot semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5475–5484.
  31. 31.Tao Xiong, Peiran Zhang, Hongbo Zhu, and Yihui Yang. 2019. Sarcasm detection with self-matching networks and low-rank bilinear pooling. In The World Wide Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019, pages 2115–2124. ACM.
  32. 32.Nan Xu, Zhixiong Zeng, and Wenji Mao. 2020. Reasoning with multimodal sarcastic tweets via modeling cross-modality contrast and semantic association. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3777–3786, Online. Association for Computational Linguistics.
  33. 33.Xiaocui Yang, Shi Feng, Yifei Zhang, and Daling Wang. 2021. Multimodal sentiment detection based on multi-channel graph neural networks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 328–339, Online. Association for Computational Linguistics.
  34. 34.Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. Graph convolutional networks for text classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 7370–7377.
  35. 35.Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou, and Jiebo Luo. 2020. A novel graph-based multi-modal fusion encoder for neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3025–3035, Online. Association for Computational Linguistics.
  36. 36.Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L. Hamilton, and Jure Leskovec. 2018. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2018, London, UK, August 19-23, 2018, pages 974–983. ACM.
  37. 37.Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu, Zhou Zhao, and Zheng Qin. 2021. Multi-modal relational graph for cross-modal video moment retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2215–2224.
  38. 38.Chen Zhang, Qiuchi Li, and Dawei Song. 2019. Aspect-based sentiment classification with aspect-specific graph convolutional networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4568–4578, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Dong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu, Qiaoming Zhu, and Guodong Zhou. 2021. Multi-modal graph fusion for named entity recognition with targeted visual guidance. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14347–14355.
  40. 40.Meishan Zhang, Yue Zhang, and Guohong Fu. 2016. Tweet sarcasm detection using deep neural network. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 2449–2460, Osaka, Japan. The COLING 2016 Organizing Committee.

Citation

MLA
Liang, B., et al. “Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1767–77, https://doi.org/10.18653/V1/2022.ACL-LONG.124.
APA
Liang, B., Lou, C., Li, X., Yang, M., Gui, L., He, Y., Pei, W., & Xu, R. (2022). Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1767–1777. https://doi.org/10.18653/V1/2022.ACL-LONG.124
Chicago
Liang, B., C. Lou, X. Li, et al. 2022. “Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1767–77. https://doi.org/10.18653/V1/2022.ACL-LONG.124.
Harvard
Liang, B. et al. (2022) “Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1767–1777. Available at: https://doi.org/10.18653/V1/2022.ACL-LONG.124.
Vancouver
1. Liang B, Lou C, Li X, Yang M, Gui L, He Y, Pei W, Xu R (2022) Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1767–1777

BibTeX

@inproceedings{Liang_2022, title={Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network}, url={http://dx.doi.org/10.18653/V1/2022.ACL-LONG.124}, DOI={10.18653/v1/2022.acl-long.124}, booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Liang, Bin and Lou, Chenwei and Li, Xiang and Yang, Min and Gui, Lin and He, Yulan and Pei, Wenjie and Xu, Ruifeng}, year={2022}, pages={1767–1777} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/