When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

Shivani KumarAtharva KulkarniMd. Shad AkhtarTanmoy Chakraborty

article2022ACL59 citations

Introduces the task of sarcasm explanation in multimodal multi-party dialogues, providing the WITS benchmark dataset and a modality-aware attention framework that generates natural language explanations for sarcastic utterances.

Listen

Automated conversational systems struggle to understand figurative speech such as sarcasm, which regularly conveys meaning opposite to literal phrasing through subtle social and contextual cues. While existing research in artificial intelligence has focused heavily on detecting whether sarcasm is present, conversational agents cannot generate appropriate responses without understanding the underlying ironic intention. To address this gap, the article introduces Sarcasm Explanation in Dialogue, a novel task designed to generate natural language explanations that clarify why a particular conversational remark is sarcastic.

The primary objective of the article is to establish a benchmark dataset and evaluate deep learning architectures that combine conversational transcripts with audio and visual cues to automatically produce coherent sarcasm explanations.

To conduct this evaluation, the researchers curated a new dataset named WITS, containing 2,240 sarcastic dialogue scenes annotated with human-written explanations of the underlying satire. The dialogues are derived from 55 episodes of a popular Indian television sitcom and feature multi-party, code-mixed conversations in Hindi and English. Each explanation explicitly captures the source speaker, target individual, sarcastic action, and descriptive context. To process these scenes, the authors developed Modality Aware Fusion, a specialized adapter mechanism that integrates acoustic features (such as vocal pitch and loudness) and visual features (such as facial gestures) into pre-trained language models like BART and mBART using context-aware attention and gating controls.

The findings show that incorporating multimodal signals significantly outperforms text-only language models. The top-performing multimodal model achieved the highest scores across standard text evaluation benchmarks, recording a ROUGE-1 score of 39.69 and a BLEU-4 score of 8.58 compared to 36.88 and 2.89 for text-only BART. Notably, adding audio and visual signals improved speaker identification accuracy by approximately 14 percentage points, reaching 91.07%. Human evaluations further confirmed that multimodal fusion generated more coherent explanations that were more relevant to both the dialogue context and the underlying sarcasm.

These results demonstrate that audio-visual non-verbal signals are vital for interpreting nuanced human communication in conversational agents. Relying solely on textual transcripts limits the ability of language models to identify speaker dynamics and implied meaning. By demonstrating that modular fusion adapters can effectively capture paralinguistic cues without massive architectural overhauls, the article provides a viable technical pathway for reducing misunderstandings in conversational AI deployment.

Deploying organizations should integrate acoustic and visual streams when building interactive dialogue systems that require emotional intelligence and social comprehension. However, development teams must treat current implementations with caution, as human evaluation scores remain modest (averaging around 3 out of 5), and the accuracy for identifying the target of sarcasm dropped slightly when multimodal data was introduced. Future research should prioritize refining target identification, evaluating larger generative foundation models, and expanding datasets beyond situational comedies to encompass broader real-world conversational domains.

arXiv: 2203.06419
Cover for When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

Abstract

Indirect speech such as sarcasm achieves a constellation of discourse goals in human communication. While the indirectness of figurative language warrants speakers to achieve certain pragmatic goals, it is challenging for AI agents to comprehend such idiosyncrasies of human communication. Though sarcasm identification has been a well-explored topic in dialogue analysis, for conversational systems to truly grasp a conversation’s innate meaning and generate appropriate responses, simply detecting sarcasm is not enough; it is vital to explain its underlying sarcastic connotation to capture its true essence. In this work, we study the discourse structure of sarcastic conversations and propose a novel task – Sarcasm Explanation in Dialogue (SED). Set in a multimodal and code-mixed setting, the task aims to generate natural language explanations of satirical conversations. To this end, we curate WITS, a new dataset to support our task. We propose MAF (Modality Aware Fusion), a multimodal context-aware attention and global information fusion module to capture multimodality and use it to benchmark WITS. The proposed attention module surpasses the traditional multimodal fusion baselines and reports the best performance on almost all metrics. Lastly, we carry out detailed analyses both quantitatively and qualitatively.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Dataset
  • 3.1 Annotation Guidelines
  • 4 Proposed Methodology
  • 4.1 Multimodal Context Aware Attention
  • 4.2 Global Information Fusion
  • 5 Experiments, Results and Analysis
  • 5.1 Feature Extraction
  • 5.2 Comparative Systems
  • 5.3 Results
  • 5.4 Ablation Study
  • 5.5 Result Analysis
  • 6 Conclusion
  • Acknowledgement
  • References
  • A Appendix
  • A.1 Embedding Space for BART and mBART
  • A.2 Fusion at Different Layers
  • A.3 More Qualitative Analysis

Knowls

  1. Knowl 1 — Sarcasm Explanation in Dialogue Task Formulation

    definition

    The task of Sarcasm Explanation in Dialogue (SED) requires a model to generate a natural language explanation for the satirical intent in a conversation. Given a multi-party dialogue sequence of utterances ⟨u1,u2,…,uN⟩\langle u_1, u_2, \dots, u_N \rangle where the final utterance uNu_N is sarcastic, along with accompanying acoustic signals AA and visual signals VV, the objective is to generate an explanation string EE.

    Each gold-standard explanation is structured around four components:

    1. Sarcasm source: The specific interlocutor uttering the sarcastic remark.
    2. Sarcasm target: The person, object, or concept toward which the sarcasm is aimed.
    3. Action word: The verb expressing the satirical communicative act (e.g., mocks, insults, taunts, implies).
    4. Description: Contextual information clarifying the irony by contrasting literal phrasing with the speaker's true intent or physical situation.
  2. Knowl 2 — WITS Dataset for Multimodal Sarcasm Explanation

    definition

    The WITS (Why Is This Sarcastic) dataset benchmarks the Sarcasm Explanation in Dialogue (SED) task in a multimodal, multi-party, Hindi-English code-mixed setting. Constructed by extending the MASAC dataset with 10 additional transcribed and video-aligned episodes from the Indian television show Sarabhai v/s Sarabhai, WITS contains 2,240 sarcastic dialogues split into train, validation, and test partitions using an 80:10:10 ratio (1,792 training, 224 validation, and 224 test dialogues).

    Key statistics of the dataset:

    • Total utterances: 9,080 (101 English, 1,453 Hindi, 7,526 Hindi-English code-mixed)
    • Utterances per dialogue: Range from 2 to 27 (average 4.05)
    • Speakers per dialogue: Up to 6 speakers (average 2.35)
    • Average utterance length: 14.39 words (average dialogue length 58.33 words)
    • Vocabulary size: 10,380 unique words (2,477 English, 7,903 Hindi)

    Each dialogue is paired with gold-standard natural language explanations written in code-mixed format to match the source dialogue style. Target annotations were curated by two annotators; if their cosine similarity was ≥90%\ge 90\%, the shorter explanation was kept, while remaining discrepancies were adjudicated by a third annotator (average initial pass cosine similarity: 87.67%87.67\%).

  3. Knowl 3 — Multimodal Context-Aware Attention Mechanism

    model/method

    Multimodal Context-Aware Attention (MCA2\text{MCA}^2) conditions the key and value projections of intermediate textual representations on external audio or visual features before performing scaled dot-product attention, avoiding cross-subspace noise injection.

    Let H∈Rn×dH \in \mathbb{R}^{n \times d} denote the intermediate textual representation at a given layer of a generative pretrained language model (e.g., BART), where nn is sequence length and dd is hidden dimension. The query QQ, key KK, and value VV matrices are computed via learnable parameter matrices WQ,WK,WV∈Rd×dW_Q, W_K, W_V \in \mathbb{R}^{d \times d}:

    [Q,K,V]=H[WQ,WK,WV][Q, K, V] = H [W_Q, W_K, W_V]

    Given auxiliary modality representations C∈Rn×dcC \in \mathbb{R}^{n \times d_c} (where dcd_c is the acoustic or visual feature dimension) and projection matrices Uk,Uv∈Rdc×dU_k, U_v \in \mathbb{R}^{d_c \times d}, the modality-informed key K^\hat{K} and value V^\hat{V} matrices are:

    [K^V^]=(1−[λkλv])[KV]+[λkλv](C[UkUv])\begin{bmatrix} \hat{K} \\ \hat{V} \end{bmatrix} = \left( 1 - \begin{bmatrix} \lambda_k \\ \lambda_v \end{bmatrix} \right) \begin{bmatrix} K \\ V \end{bmatrix} + \begin{bmatrix} \lambda_k \\ \lambda_v \end{bmatrix} \left( C \begin{bmatrix} U_k \\ U_v \end{bmatrix} \right)

    The gating coefficients λk,λv∈Rn×1\lambda_k, \lambda_v \in \mathbb{R}^{n \times 1} dynamically balance text and multimodal representations via learnable parameters Wk1,Wv1,Wk2,Wv2∈Rd×1W_{k1}, W_{v1}, W_{k2}, W_{v2} \in \mathbb{R}^{d \times 1}:

    [λkλv]=σ([KV][Wk1Wv1]+C[UkUv][Wk2Wv2])\begin{bmatrix} \lambda_k \\ \lambda_v \end{bmatrix} = \sigma \left( \begin{bmatrix} K \\ V \end{bmatrix} \begin{bmatrix} W_{k1} \\ W_{v1} \end{bmatrix} + C \begin{bmatrix} U_k \\ U_v \end{bmatrix} \begin{bmatrix} W_{k2} \\ W_{v2} \end{bmatrix} \right)

    where σ\sigma is the sigmoid function. Scaled dot-product attention computes the modality-infused representations HaH_a (acoustic) and HvH_v (visual):

    Ha=Softmax(QK^aTdk)V^a,Hv=Softmax(QK^vTdk)V^vH_a = \text{Softmax}\left(\frac{Q \hat{K}_a^T}{\sqrt{d_k}}\right) \hat{V}_a, \quad H_v = \text{Softmax}\left(\frac{Q \hat{K}_v^T}{\sqrt{d_k}}\right) \hat{V}_v

  4. Knowl 4 — Global Information Fusion Mechanism

    model/method

    The Global Information Fusion (GIF) module integrates acoustic-infused (HaH_a) and visual-infused (HvH_v) representations into the textual representation H∈Rn×dH \in \mathbb{R}^{n \times d}. GIF uses gating vectors to regulate the transmission of each modality and block non-verbal noise:

    ga=[H⊕Ha]Wa+bag_a = [H \oplus H_a] W_a + b_a

    gv=[H⊕Hv]Wv+bvg_v = [H \oplus H_v] W_v + b_v

    H^=H+ga⊙Ha+gv⊙Hv\hat{H} = H + g_a \odot H_a + g_v \odot H_v

    where ⊕\oplus denotes concatenation along the feature dimension, ⊙\odot represents element-wise multiplication, Wa,Wv∈R2d×dW_a, W_v \in \mathbb{R}^{2d \times d} are learnable weight matrices, and ba,bv∈Rd×1b_a, b_v \in \mathbb{R}^{d \times 1} are learnable bias vectors. The resulting multimodal fused representation H^\hat{H} is re-introduced into the generative pretrained language model encoder for subsequent processing.

  5. Knowl 5 — Acoustic and Visual Feature Extraction Setup

    experimental setup

    For multimodal inputs in the WITS dataset, acoustic and visual representations are processed as follows:

    • Acoustic Features: Extracted using the openSMILE toolkit with the eGeMAPS feature set. Non-overlapping frames are extracted using a 25 ms window length and 10 ms window shift to obtain 154-dimensional functional features (such as MFCCs and loudness), which are then passed through a Transformer encoder.
    • Visual Features: Extracted using a 3D CNN ResNeXt-101 backbone pretrained on the Kinetics dataset (101 action classes). Visual features are extracted at a frame rate of 1.5 fps, 720p resolution, and window length of 16 frames to yield 2048-dimensional vectors, which are then passed through a Transformer encoder to model temporal dialogue context.
  6. Knowl 6 — Experimental Results on Sarcasm Explanation Generation

    data/table

    Performance of unimodal text-based sequence-to-sequence models and Modality Aware Fusion (MAF) multimodal variants on the WITS test set (evaluated with ROUGE-1/2/L, BLEU-1/2/3/4, METEOR, and BERTScore):

    Model R1 R2 RL B1 B2 B3 B4 M BS
    Textual Baselines
    RNN 29.22 7.85 27.59 22.06 8.22 4.76 2.88 18.45 73.24
    Transformers 29.17 6.35 27.97 17.79 5.63 2.61 0.88 15.65 72.21
    PGN 23.37 4.83 17.46 17.32 6.68 1.58 0.52 23.54 71.90
    mBART 33.66 11.02 31.50 22.92 10.56 6.07 3.39 21.03 73.83
    BART 36.88 11.91 33.49 27.44 12.23 5.96 2.89 26.65 76.03
    Multimodal Models
    MAF-TAM_M 39.02 15.90 36.83 31.26 16.94 11.54 7.72 29.05 77.06
    MAF-TVM_M 39.47 16.78 37.38 32.44 17.91 12.02 7.36 29.74 77.47
    MAF-TAVM_M 38.52 14.13 36.60 30.50 15.20 9.78 5.74 27.42 76.70
    MAF-TAB_B 38.21 14.53 35.97 30.58 15.36 9.63 5.96 27.71 77.08
    MAF-TVB_B 37.48 15.38 35.64 30.28 16.89 10.33 6.55 28.24 76.95
    MAF-TAVB_B 39.69 17.10 37.37 33.20 18.69 12.37 8.58 30.40 77.67

    Notations: PGN = Pointer Generator Network; Subscripts MM and BB represent mBART and BART backbones; TA, TV, TAV represent Text+Audio, Text+Video, and Text+Audio+Video modalities. MAF-TAVB_B outperforms unimodal BART across all metrics (+2.81 ROUGE-1, +5.19 ROUGE-2, +5.76 BLEU-1, +5.69 BLEU-4, +3.75 METEOR).

  7. Knowl 7 — Ablation of Multimodal Fusion Components in MAF

    data/table

    Ablation performance comparing the proposed Multimodal Aware Fusion (MAF) against alternative fusion strategies across mBART (MAF-TAVM_M) and BART (MAF-TAVB_B) backbones on the WITS dataset:

    Model Variant R1 R2 RL B1 B2 B3 B4 M BS
    MAF-TAVM_M 38.52 14.13 36.60 30.50 15.20 9.78 5.74 27.42 76.70
    - MCA2\text{MCA}^2 + CONCAT1 37.56 14.85 34.90 30.16 15.76 10.12 6.82 28.59 76.59
    - MAF + CONCAT2 17.22 1.70 14.12 13.11 2.11 0.00 0.00 9.34 66.64
    - MCA2\text{MCA}^2 + DPA 36.43 13.04 33.75 28.73 14.02 8.00 4.89 25.60 75.58
    - GIF 36.37 13.85 34.92 28.49 14.34 9.00 6.16 25.75 76.86
    MAF-TAVB_B 39.69 17.10 37.37 33.20 18.69 12.37 8.58 30.40 77.67
    - MCA2\text{MCA}^2 + CONCAT1 36.88 13.21 34.39 29.63 14.56 8.43 4.84 26.15 76.08
    - MAF + CONCAT2 21.11 2.31 19.68 12.44 2.44 0.73 0.31 9.51 69.54
    - MCA2\text{MCA}^2 + DPA 38.84 14.76 36.96 30.23 15.95 9.88 5.83 28.04 77.20
    - GIF 39.45 14.85 37.18 31.85 15.97 9.62 5.47 28.87 77.54

    Variant definitions and insights:

    • CONCAT1: Separate bimodal concatenations (T⊕A)(T \oplus A) and (T⊕V)(T \oplus V) followed by the GIF block.
    • CONCAT2: Trimodal concatenation (T⊕A⊕V)(T \oplus A \oplus V) followed by a linear projection layer; this causes performance to collapse by over 18% in R1, showing that naive concatenation fails to capture multimodal interactions.
    • DPA: Standard cross-modal Dot Product Attention replacing MCA2\text{MCA}^2; MAF outperforms DPA by 1-3% across generative metrics.
    • - GIF: Replacing GIF gating with simple element-wise addition (H+Ha+Hv)(H + H_a + H_v) degrades ROUGE and BLEU performance by 2-3%.
  8. Knowl 8 — Optimal BART Encoder Layer Placement for Multimodal Fusion

    empirical result

    Integrating the MAF adapter module at different layer depths within the 6-layer BART base encoder produces varying generative performances on the WITS dataset:

    Fusion Layer Depth ROUGE-1 ROUGE-2 ROUGE-L
    Before Layer 1 37.27 13.95 35.24
    Before Layer 2 37.63 14.32 35.57
    Before Layer 3 36.73 13.15 34.63
    Before Layer 4 37.61 14.98 36.04
    Before Layer 5 37.34 13.67 35.48
    Before Layer 6 39.69 17.10 37.37

    Injecting multimodal features right before Layer 6 (the final encoder layer) achieves optimal performance. Passing fused multimodal representations through only one final encoder layer prevents multimodal information from being diluted across multiple subsequent self-attention layers prior to decoding.

  9. Knowl 9 — Sarcasm Source and Target Identification Accuracy

    data/table

    Accuracy evaluation of BART- and mBART-based generated explanations in correctly identifying the sarcasm source (speaker) and sarcasm target (intended recipient) on the WITS test set:

    Attribute mBART BART MAF-TAB_B MAF-TVB_B MAF-TAVB_B
    Source Accuracy (%) 75.00 77.23 87.94 85.71 91.07
    Target Accuracy (%) 45.53 52.67 43.75 43.75 46.42

    Incorporating acoustic cues increases source accuracy by +10.71%+10.71\%, visual cues by +8.48%+8.48\%, and their combined trimodal fusion (MAF-TAVB_B) yields a +13.84%+13.84\% improvement over text-only BART. This shows that vocal pitch/intonation and facial/gestural patterns provide distinct speaker identity markers. Conversely, target identification accuracy declines with multimodality, dropping from 52.67%52.67\% (BART) to 46.42%46.42\% (MAF-TAVB_B).

  10. Knowl 10 — Human Evaluation of Generated Sarcasm Explanations

    empirical result

    A blind human evaluation conducted by 25 expert evaluators on 30 randomly sampled test instances assessed generated explanations on a 0–5 Likert scale across three criteria: Coherence (structural and syntactic organization), Related to Dialogue (topical alignment with conversation), and Related to Sarcasm (relevance to the underlying irony):

    Model Coherency Related to dialogue Related to sarcasm
    mBART 2.57 2.66 2.15
    BART 2.73 2.56 2.18
    MAF-TAB_B 2.95 2.91 2.51
    MAF-TVB_B 3.01 3.11 2.66
    MAF-TAVB_B 3.03 3.11 2.77

    MAF-TAVB_B scored highest across all dimensions, gaining +0.59+0.59 points over unimodal BART in sarcasm explanation relevance and +0.55+0.55 points in dialogue topic adherence. Overall scores below 3.5 indicate that generating accurate figurative explanations remains challenging for conversational AI.

Coverage note — No substantial contributed material was omitted. Specific qualitative dialogue transcripts from the appendix were summarized into the broader empirical and architectural findings.

References

  1. 1.Ibrahim Abu Farha and Walid Magdy. 2020. From Arabic sentiment analysis to sarcasm detection: The ArSarcasm dataset. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 32–39, Marseille, France. European Language Resource Association.
  2. 2.Salvatore Attardo, Jodi Eisterhold, Jennifery Hay, and Isabella Poggi. 2003. Multimodal markers of irony and sarcasm. Humor: International Journal of Humor Research, 16(2).
  3. 3.Manjot Bedi, Shivani Kumar, Md Shad Akhtar, and Tanmoy Chakraborty. 2021. Multi-modal sarcasm detection and humor classification in code-mixed conversations. IEEE Transactions on Affective Computing, pages 1–1.
  4. 4.Santosh Kumar Bharti, Korra Sathya Babu, and Sanjay Kumar Jena. 2017. Harnessing online news for sarcasm detection in hindi tweets. In Pattern Recognition and Machine Intelligence, pages 679–686, Cham. Springer International Publishing.
  5. 5.Yitao Cai, Huiyu Cai, and Xiaojun Wan. 2019. Multimodal sarcasm detection in Twitter with hierarchical fusion model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2506–2515, Florence, Italy. Association for Computational Linguistics.
  6. 6.Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards multimodal sarcasm detection (an Obviously perfect paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629, Florence, Italy. Association for Computational Linguistics.
  7. 7.Tuhin Chakrabarty, Debanjan Ghosh, Smaranda Muresan, and Nanyun Peng. 2020. R^3: Reverse, retrieve, and rank for sarcasm generation with commonsense knowledge. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7976–7986, Online. Association for Computational Linguistics.
  8. 8.Dushyant Singh Chauhan, Dhanush S R, Asif Ekbal, and Pushpak Bhattacharyya. 2020. Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4351–4360, Online. Association for Computational Linguistics.
  9. 9.Alessandra Teresa Cignarella, Simona Frenda, Valerio Basile, Cristina Bosco, Viviana Patti, Paolo Rosso, et al. 2018. Overview of the evalita 2018 task on irony detection in italian tweets (ironita). In Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian (EVALITA 2018), volume 2263, pages 1–6. CEUR-WS.
  10. 10.Herbert L. Colston. 1997. Salting a wound or sugaring a pill: The pragmatic functions of ironic criticism. Discourse Processes, 23(1):25–45.
  11. 11.Herbert L Colston and Shauna B Keller. 1998. You’ll never believe this: Irony and hyperbole in expressing surprise. Journal of psycholinguistic research, 27(4):499–513.
  12. 12.Michael Denkowski and Alon Lavie. 2014. Meteor universal: Language specific translation evaluation for any target language. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 376–380, Baltimore, Maryland, USA. Association for Computational Linguistics.
  13. 13.Abhijeet Dubey, Aditya Joshi, and Pushpak Bhattacharyya. 2019. Deep models for converting sarcastic utterances into their non sarcastic interpretation. In Proceedings of the ACM India Joint International Conference on Data Science and Management of Data, CoDS-COMAD ’19, page 289–292, New York, NY, USA. Association for Computing Machinery.
  14. 14.Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Y. Devillers, Julien Epps, Petri Laukka, Shrikanth S. Narayanan, and Khiet P. Truong. 2016. The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing. IEEE Transactions on Affective Computing, 7(2):190–202.
  15. 15.Debanjan Ghosh, Alexander R. Fabbri, and Smaranda Muresan. 2018. Sarcasm analysis using conversation context. Computational Linguistics, 44(4):755–792.
  16. 16.Debanjan Ghosh, Alexander Richard Fabbri, and Smaranda Muresan. 2017. The role of conversation context for sarcasm detection in online interactions. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 186–196, Saarbrücken, Germany. Association for Computational Linguistics.
  17. 17.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  18. 18.Md Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque. 2021. Humor knowledge enriched transformer for understanding multimodal humor. Proceedings of the AAAI Conference on Artificial Intelligence, 35(14):12972–12980.
  19. 19.Stacey L. Ivanko and Penny M. Pexman. 2003. Context incongruity and irony processing. Discourse Processes, 35(3):241–279.
  20. 20.Aditya Joshi, Pushpak Bhattacharyya, and Mark J. Carman. 2017. Automatic sarcasm detection: A survey. ACM Comput. Surv., 50(5).
  21. 21.Aditya Joshi, Vinita Sharma, and Pushpak Bhattacharyya. 2015. Harnessing context incongruity for sarcasm detection. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 757–762, Beijing, China. Association for Computational Linguistics.
  22. 22.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017. The kinetics human action video dataset.
  23. 23.Roger Kreuz and Gina Caucci. 2007. Lexical influences on the perception of sarcasm. In Proceedings of the Workshop on Computational Approaches to Figurative Language, pages 1–4, Rochester, New York. Association for Computational Linguistics.
  24. 24.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  25. 25.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  26. 26.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  27. 27.Abhijit Mishra, Tarun Tater, and Karthik Sankaranarayanan. 2019. A modular architecture for unsupervised sarcasm generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6144–6154, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Henri Olkoniemi, Henri Ranta, and Johanna K Kaakinen. 2016. Individual differences in the processing of written sarcasm and metaphor: Evidence from eye movements. Journal of Experimental Psychology: Learning, Memory, and Cognition, 42(3):433.
  29. 29.Shereen Oraby, Vrindavan Harrison, Amita Misra, Ellen Riloff, and Marilyn Walker. 2017. Are you serious?: Rhetorical questions and sarcasm in social media dialog. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 310–319, Saarbrücken, Germany. Association for Computational Linguistics.
  30. 30.Reynier Ortega-Bueno, Francisco Rangel, D Hernández Farıas, Paolo Rosso, Manuel Montes-y Gómez, and José E Medina Pagola. 2019. Overview of the task on irony detection in spanish variants. In Proceedings of the Iberian languages evaluation forum (IberLEF 2019), co-located with 34th conference of the Spanish Society for natural language processing (SEPLN 2019). CEUR-WS. org, volume 2421, pages 229–256.
  31. 31.Hongliang Pan, Zheng Lin, Peng Fu, Yatao Qi, and Weiping Wang. 2020. Modeling intra and intermodality incongruity for multi-modal sarcasm detection. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1383–1392, Online. Association for Computational Linguistics.
  32. 32.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  33. 33.Lotem Peled and Roi Reichart. 2017. Sarcasm SIGN: Interpreting sarcasm with sentiment based monolingual machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1690–1700, Vancouver, Canada. Association for Computational Linguistics.
  34. 34.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
  35. 35.Richard M. Roberts and Roger J. Kreuz. 1994. Why do people use figurative language? Psychological Science, 5(3):159–163.
  36. 36.Patricia Rockwell. 2007. Vocal features of conversational sarcasm: A comparison of methods. Journal of psycholinguistic research, 36(5):361–369.
  37. 37.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, Vancouver, Canada. Association for Computational Linguistics.
  38. 38.Himani Srivastava, Vaibhav Varshney, Surabhi Kumari, and Saurabh Srivastava. 2020. A novel hierarchical BERT architecture for sarcasm detection. In Proceedings of the Second Workshop on Figurative Language Processing, pages 93–97, Online. Association for Computational Linguistics.
  39. 39.Sahil Swami, Ankush Khandelwal, Vinay Singh, Syed Sarfaraz Akhtar, and Manish Shrivastava. 2018. A corpus of english-hindi code-mixed tweets for sarcasm detection. arXiv preprint arXiv:1805.11869.
  40. 40.Sabina Tabacaru and Maarten Lemmens. 2014. Raised eyebrows as gestural triggers in humour: The case of sarcasm and hyper-understanding. The European Journal of Humour Research, 2(2):11–31.
  41. 41.Yi Tay, Anh Tuan Luu, Siu Cheung Hui, and Jian Su. 2018. Reasoning with sarcasm by reading in-between. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1010–1020, Melbourne, Australia. Association for Computational Linguistics.
  42. 42.Oren Tsur, Dmitry Davidov, and Ari Rappoport. 2010. Icwsm — a great catchy name: Semi-supervised recognition of sarcastic sentences in online product reviews. Proceedings of the International AAAI Conference on Web and Social Media, 4(1):162–169.
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  44. 44.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  45. 45.Henry M Wellman. 2014. Making minds: How theory of mind develops. Oxford University Press.
  46. 46.Tao Xiong, Peiran Zhang, Hongbo Zhu, and Yihui Yang. 2019. Sarcasm detection with self-matching networks and low-rank bilinear pooling. In The World Wide Web Conference, WWW ’19, page 2115–2124, New York, NY, USA. Association for Computing Machinery.
  47. 47.Nan Xu, Zhixiong Zeng, and Wenji Mao. 2020. Reasoning with multimodal sarcastic tweets via modeling cross-modality contrast and semantic association. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3777–3786, Online. Association for Computational Linguistics.
  48. 48.Baosong Yang, Jian Li, Derek F. Wong, Lidia S. Chao, Xing Wang, and Zhaopeng Tu. 2019. Context-aware self-attention networks. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):387–394.
  49. 49.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.

Citation

MLA
Kumar, S., et al. “When Did You Become so Smart, Oh Wise One?! Sarcasm Explanation in Multi-modal Multi-party Dialogues”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 5956–68, https://doi.org/10.18653/v1/2022.acl-long.411.
APA
Kumar, S., Kulkarni, A., Akhtar, M. S., & Chakraborty, T. (2022). When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5956–5968. https://doi.org/10.18653/v1/2022.acl-long.411
Chicago
Kumar, S., A. Kulkarni, M. S. Akhtar, and T. Chakraborty. 2022. “When Did You Become so Smart, Oh Wise One?! Sarcasm Explanation in Multi-modal Multi-party Dialogues”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5956–68. https://doi.org/10.18653/v1/2022.acl-long.411.
Harvard
Kumar, S. et al. (2022) “When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5956–5968. Available at: https://doi.org/10.18653/v1/2022.acl-long.411.
Vancouver
1. Kumar S, Kulkarni A, Akhtar MS, Chakraborty T (2022) When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5956–5968

BibTeX

@inproceedings{kumar-etal-2022-become,
    title = "When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues",
    author = "Kumar, Shivani  and
      Kulkarni, Atharva  and
      Akhtar, Md Shad  and
      Chakraborty, Tanmoy",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.411/",
    doi = "10.18653/v1/2022.acl-long.411",
    pages = "5956--5968"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/