MemeCap: A Dataset for Captioning and Interpreting Memes

Eunjeong HwangVered Shwartz

article2023EMNLP78 citations

Introduces MEMECAP, a dataset of over six thousand annotated memes with post titles, literal descriptions, and visual metaphors, to expose how modern vision-language models fail to interpret complex multimodal figurative meaning compared to humans.

Listen

Internet memes have become a primary medium for online communication, relying heavily on visual metaphors, humor, and shared cultural context to express complex ideas. While modern vision and language artificial intelligence models excel at standard image recognition and descriptive captioning, they struggle to interpret non-literal, figurative content. This limitation creates significant operational blind spots for automated systems tasked with content moderation, social media analytics, and digital communication monitoring, where understanding intent is critical.

The article introduces the generative task of meme captioning and presents MEMECAP, a new benchmark dataset designed to evaluate whether state-of-the-art vision and language models can interpret the intended meaning behind internet memes. The primary objective is to assess how effectively current artificial intelligence architectures comprehend visual metaphors by generating concise natural language explanations of memes and their accompanying post titles.

To establish the benchmark, the researchers collected 6,384 non-offensive memes and their titles from Reddit and crowdsourced multi-stage human annotations through Amazon Mechanical Turk. Annotations included literal image descriptions with embedded meme text removed, identification of metaphorical visual elements and their targets, and human-written captions explaining the underlying message. The dataset was split into training, validation, and a diverse test set of 559 memes. The study then evaluated three leading open-source models—OpenFlamingo-9B, MiniGPT4, and the text-only LLaMA-7B—across zero-shot, few-shot, fine-tuned, and structured reasoning settings, using both standard automated evaluation metrics and blind human assessments.

The evaluation revealed several critical findings regarding machine comprehension of figurative media. First, all models perform substantially worse than human baselines across key evaluation criteria, lagging behind humans by 36.6 percentage points in overall correctness, 29.3 points in textual completeness, 24.5 points in visual completeness, and 18.4 points in factual faithfulness. Second, models frequently generate incorrect interpretations by taking visual elements literally, hallucinating unsupported claims, or simply repeating optical character recognition text rather than explaining meaning. Third, providing few-shot examples and explicit extracted text improved surface fluency and automated overlap scores, but providing explicit metaphorical rationales yielded little to no improvement in interpretation accuracy. Finally, text-only models supplied with textual image descriptions performed competitively with multi-modal vision-language models, indicating that current multi-modal architectures remain inefficient at cross-modal visual reasoning.

These findings demonstrate that organizations cannot rely on current vision-language foundation models for autonomous interpretation of metaphorical digital media. Deploying these systems out-of-the-box for policy enforcement, brand sentiment tracking, or harmful content screening introduces substantial risks of misinterpretation, false positives, and unfaithful hallucinations. The evidence indicates that scaling standard cross-modal pre-training is insufficient; models require dedicated architectures and training objectives tailored to abstract cultural reasoning and visual metaphor resolution.

For technology leaders and developers, the article recommends against fully automated meme analysis workflows, advising instead that automated pipelines include human review for high-stakes decisions. Future engineering efforts should focus on training models to integrate external background knowledge and suppress literal visual descriptions during figurative reasoning. Further research is necessary to explore whether richer multi-modal pre-training data can overcome the limitations of lightweight fine-tuning.

Confidence in these conclusions is high, given the rigorous crowdsourcing design, manual filtering, and consistent alignment between automated metrics and human evaluations. However, readers should consider certain boundary conditions: annotations retain a degree of human subjectivity regarding humor and slang, crowdsourced metaphor-target mappings varied in quality, and the source data was gathered exclusively from English-language Reddit communities, which may limit generalizability across broader cultural and linguistic contexts.

arXiv: 2305.13703eujhwang/meme-cap
Cover for MemeCap: A Dataset for Captioning and Interpreting Memes

Abstract

Memes are a widely popular tool for web users to express their thoughts using visual metaphors. Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while employing background knowledge and reasoning abilities. We present the task of meme captioning and release a new dataset, MEMECAP. Our dataset contains 6.3K memes along with the title of the post containing the meme, the meme captions, the literal image caption, and the visual metaphors. Despite the recent success of vision and language (VL) models on tasks such as image captioning and visual question answering, our extensive experiments using state-of-the-art VL models show that they still struggle with visual metaphors, and perform substantially worse than humans.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background
  • 2.1 Metaphors
  • 2.2 Memes
  • 2.3 Other Image Datasets
  • 3 The MEMECAP Dataset
  • 3.1 Memes
  • 3.2 Captions
  • 3.3 Final Dataset
  • 3.4 Types of Metaphors
  • 4 Experimental Setup
  • 4.1 Models
  • 4.2 Evaluation Setup
  • 5 Results
  • 5.1 Automatic Evaluation
  • 5.2 Human Evaluation
  • 5.3 Ablation Tests
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Additional Experimental Results

Knowls

  1. Knowl 1 — Meme captioning requires conveying intended meaning rather than literal content

    definition

    Meme captioning is the task of generating a concise description of what the meme poster intends to convey, given the meme image and its post title. A successful caption interprets the visual and textual elements together, including visual metaphors, and does not merely describe the depicted scene or repeat the text embedded in the image. In the annotation example, for instance, the caption explains that a character represents the poster rather than naming the character as part of the intended message.

  2. Knowl 2 — MEMECAP provides meme meanings, literal image descriptions, and metaphor annotations

    data/table

    MEMECAP contains 6,384 Reddit memes, each associated with its post title, a literal image caption, and a caption of the intended meme meaning. The dataset also records visual metaphors by identifying terms in the literal description that function as metaphorical vehicles and specifying their targets. The literal captions support separating scene recognition from meaning interpretation, while the metaphor annotations provide an additional representation of how visual elements contribute to the message.

  3. Knowl 3 — MEMECAP collection and annotation use two complementary annotation rounds

    model/method

    The authors collected posts with memes and titles from Reddit's /r/memes community. They manually removed memes without text or with excessive text, filtered profanity using a banned-word list, and excluded images for which NudeNet returned an unsafe score above 0.9. Annotators who marked an item offensive, sexual, hateful, or uninterpretable also caused it to be excluded.

    In the first Amazon Mechanical Turk round, workers described the image while disregarding its text. The text was removed with LaMa inpainting before annotation, and one literal image caption was collected and manually checked per meme. In the second round, workers saw the complete meme, title, and literal image caption; they marked which caption terms were metaphorical and identified their targets, then wrote a concise description of the poster's intended meaning without naming metaphor vehicles. The study collected one meme caption per training item and two to four per test item. Workers had to be in an English-speaking country, have at least a 98% acceptance rate on 5,000 prior tasks, and pass a task-specific qualification test.

  4. Knowl 4 — MEMECAP split statistics and annotation density

    data/table

    The paper reports 6,384 memes overall. Its split-statistics table lists 5,828 training-plus-validation memes and 559 test memes; those two tabulated counts sum to 6,387. The table reports averages per meme of 1.0 meme caption, 1.0 literal image caption, and 2.1 metaphor keywords for training plus validation, compared with 3.4 meme captions, 1.0 literal image caption, and 3.1 metaphor keywords for test. The additional test captions support evaluation against multiple human descriptions. The authors formed clusters from meme-caption representations produced by OPT-2.7B, sampled 10% of each cluster for the test set, and assigned the remainder to training and validation; they report that the dataset contains no duplicate memes.

  5. Knowl 5 — A small manual analysis finds complementary modalities and recurring metaphor patterns

    empirical result

    The manual analysis of 28 memes, summarized visually on page 4, classified memes as complementary when both image and text were needed, text-dominant or image-dominant when one modality could suffice, or as having no visual metaphor. Complementary memes made up 44% of the sample; each of the other categories accounted for about 19% (the chart reports 18% for no visual metaphor). People or characters were the most common visual vehicles, followed by objects, facial expressions or gestures, and actions. Metaphor targets most often expressed a behavior or stance, or represented the meme poster; other targets included an approach or concept, another person, and a desire-versus-reality contrast.

  6. Knowl 6 — The benchmark compares multimodal and text-only models across inputs and learning setups

    experimental setup

    The evaluated systems were OpenFlamingo-9B and MiniGPT-4 as vision-language models, and LLaMA-7B as a text-only model. Models received different combinations of the meme image, post title, literal image caption, and OCR-extracted text from inside the meme. LLaMA had no direct image input and instead used text such as the title, image caption, and OCR output. The authors tested zero-shot prompting; four-, eight-, and twelve-example in-context prompting for Flamingo and LLaMA; and zero-shot or training-set fine-tuning for MiniGPT-4. They also tested prompts that supplied metaphor vehicle-to-target rationales. Automatic evaluation used BLEU-4, ROUGE-L, and BERTScore F1 with DeBERTa-Xlarge-MNLI.

  7. Knowl 7 — Automatic metrics show strong few-shot scores but do not establish robust metaphor interpretation

    empirical result

    The strongest reported automatic scores came from LLaMA-7B with eight in-context examples and title, literal image caption, and OCR text: BLEU-4 28.80, ROUGE-L 44.10, and BERT-F1 74.71. This text-only system was competitive with or better than the tested vision-language systems on these metrics. OpenFlamingo's four-shot setup with the meme, title, image caption, and OCR text scored 26.73, 43.47, and 73.86, respectively. MiniGPT-4 with the same full input in zero-shot scored 12.46, 31.44, and 68.62; fine-tuning it on the training data lowered its scores to 7.50, 27.88, and 65.47.

    The page-7 results table reports BLEU-4, ROUGE-L, and BERT-F1 in that order for these conditions. Adding metaphor rationales sharply reduced zero-shot scores: Flamingo's BLEU-4 fell from 19.31 with the full input to 2.49 with rationales, and LLaMA's fell from 20.77 to 6.72. These automatic metrics measure overlap or semantic similarity to reference captions and, by themselves, do not verify that generated captions correctly interpret the meme.

  8. Knowl 8 — Human evaluation finds a substantial quality gap between model and human captions

    empirical result

    For human evaluation, three lab annotators used majority vote to judge outputs for 30 randomly sampled memes. They assessed correctness of intended meaning, appropriate length, visual completeness, textual completeness, and faithfulness to the available image and text. As summarized in the page-7 human-evaluation chart, all tested model outputs performed significantly worse than human captions on every criterion except appropriate length. The reported human–model gaps were 36.6 percentage points for correctness, 29.3 for textual completeness, 24.5 for visual completeness, and 18.4 for faithfulness. Performance differed by model: Flamingo and LLaMA were more correct and faithful, whereas MiniGPT-4 was more visually complete.

  9. Knowl 9 — Ablations show model-specific reliance on visual and textual inputs

    empirical result

    Removing inputs generally reduced automatic scores, but the effect depended on the model and metric. For zero-shot LLaMA, removing the literal image caption from the title-plus-image-caption input reduced BLEU-4 by 1.85, ROUGE-L by 1.84, and BERT-F1 by 2.40; removing the title reduced them by 0.88, 0.93, and 0.62. Thus the image caption contributed more than the title in that comparison. Zero-shot MiniGPT-4 was especially sensitive to the title: removing it reduced BLEU-4 by 8.20, ROUGE-L by 8.50, and BERT-F1 by 2.88. The authors interpret the results as evidence that the modalities can be complementary, while noting that MiniGPT-4 appeared more dependent on textual input and made limited use of visual information.

  10. Knowl 10 — Common generation errors include literal interpretation, omission, and unsupported details

    empirical result

    Qualitative inspection identified several recurring errors in model-generated captions. Models sometimes copied the text inside a meme while omitting an important visual cue, treated a metaphorical visual element literally, or produced claims unsupported by either the image or text. In one example, a caption missed a character's positive expression and consequently failed to convey the poster's enjoyment of reading old argument threads. In another, the generated caption asserted a desire for success even though the meme and title did not support that interpretation. The authors also observed cases where insufficient background knowledge appeared to prevent correct interpretation.

  11. Knowl 11 — Metaphor-label quality and subjective interpretation limit the dataset and evaluation

    limitation

    The authors report that visual-metaphor annotations are inconsistent: annotators can explain a meme without reliably mapping its visual vehicles to textual targets. They suggest that this may help explain why giving models metaphor annotations as input did not improve captioning. Meme interpretation also depends on background knowledge that can differ across annotators, and judgments of caption quality involve some subjectivity. The authors manually checked captions and used in-house annotators for evaluation to mitigate these issues, but do not claim to eliminate them.

Coverage note — The exhaustive per-input and per-shot metric grid is omitted in favor of representative comparisons and the best reported scores; it is extensive but does not add a distinct conclusion beyond the benchmark trends captured here. Model architecture and pretraining details are also omitted because they describe baseline background rather than the paper's contribution.

References

  1. 1.Ehsan Aghazadeh, Mohsen Fayyaz, and Yadollah Yaghoobzadeh. 2022. Metaphors in pre-trained language models: Probing and generalization across datasets and languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2037–2050, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Arjun R. Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas Guibas, William T. Freeman, Yuanzhen Li, and Varun Jampani. 2023. Metaclue: Towards comprehensive visual metaphors research.
  3. 3.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems.
  4. 4.Malihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut, and Matthew Stone. 2020. Cross-modal coherence modeling for caption generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6525–6535, Online. Association for Computational Linguistics.
  5. 5.Anas Awadalla, Irena Gao, Joshua Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. 2023. Open-flamingo.
  6. 6.Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, and Roy Schwartz. 2023. Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  8. 8.Branislav Buchel. 2012. Internet memes as means of communication. Brno: Masaryk University.
  9. 9.Tuhin Chakrabarty, Yejin Choi, and Vered Shwartz. 2022. It’s not rocket science: Interpreting figurative language in narratives. Transactions of the Association for Computational Linguistics, 10:589–606.
  10. 10.Tuhin Chakrabarty, Debanjan Ghosh, Adam Poliak, and Smaranda Muresan. 2021a. Figurative language in recognizing textual entailment. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3354–3361, Online. Association for Computational Linguistics.
  11. 11.Tuhin Chakrabarty, Arkady Saakyan, Olivia Winn, Artemis Panagopoulou, Yue Yang, Marianna Apidianaki, and Smaranda Muresan. 2023. I spy a metaphor: Large language models and diffusion models co-create visual metaphors. In Findings of ACL.
  12. 12.Tuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, and Nanyun Peng. 2021b. MERMAID: Metaphor generation with symbolism and discriminative decoding. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4250–4261, Online. Association for Computational Linguistics.
  13. 13.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality.
  14. 14.Minjin Choi, Sunkyung Lee, Eunseong Choi, Heesoo Park, Junhyuk Lee, Dongwon Lee, and Jongwuk Lee. 2021. MelBERT: Metaphor detection via contextualized late interaction using metaphorical identification theories. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1763–1773, Online. Association for Computational Linguistics.
  15. 15.Ellen Dodge, Jisup Hong, and Elise Stickles. 2015. MetaNet: Deep semantic automatic metaphor analysis. In Proceedings of the Third Workshop on Metaphor in NLP, pages 40–49, Denver, Colorado. Association for Computational Linguistics.
  16. 16.Charles Forceville. 1996. Pictorial metaphor in advertising. Psychology Press.
  17. 17.Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2023. Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 688–714, Toronto, Canada. Association for Computational Linguistics.
  18. 18.Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Casey A. Fitzpatrick, Peter Bull, Greg Lipstein, Tony Nelli, Ron Zhu, Niklas Muennighoff, Riza Velioglu, Jewgeni Rose, Phillip Lippe, Nithin Holla, Shantanu Chandra, Santhosh Rajamanickam, Georgios Antoniou, Ekaterina Shutova, Helen Yannakoudakis, Vlad Sandulescu, Umut Ozertem, Patrick Pantel, Lucia Specia, and Devi Parikh. 2021. The hateful memes challenge: competition report. In Proceedings of the NeurIPS 2020 Competition and Demonstration Track, volume 133 of Proceedings of Machine Learning Research, pages 344–360. PMLR.
  19. 19.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.
  20. 20.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  21. 21.OpenAI. 2023. Gpt-4 technical report.
  22. 22.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011. Im2text: Describing images using 1 million captioned photographs. In Advances in Neural Information Processing Systems, volume 24. Curran Associates, Inc.
  23. 23.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  24. 24.Jingnong Qu, Liunian Harold Li, Jieyu Zhao, Sunipa Dev, and Kai-Wei Chang. 2022. Disinfomeme: A multimodal dataset for detecting meme intentionally spreading out disinformation.
  25. 25.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  26. 26.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models.
  27. 27.Kate Scott. 2021. Memes as multimodal metaphors: A relevance theory analysis. Pragmatics & Cognition, 28(2):277–298.
  28. 28.Chhavi Sharma, William Paka, Scott, Deepesh Bhageria, Amitava Das, Soujanya Poria, Tanmoy Chakraborty, and Björn Gambäck. 2020. Task Report: Memotion Analysis 1.0 @SemEval 2020: The Visuo-Lingual Metaphor! In Proceedings of the 14th International Workshop on Semantic Evaluation (SemEval-2020), Barcelona, Spain. Association for Computational Linguistics.
  29. 29.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, Melbourne, Australia. Association for Computational Linguistics.
  30. 30.Shivam Sharma, Atharva Kulkarni, Tharun Suresh, Himanshi Mathur, Preslav Nakov, Md. Shad Akhtar, and Tanmoy Chakraborty. 2023. Characterizing the entities in harmful memes: Who is the hero, the villain, the victim? In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2149–2163, Dubrovnik, Croatia. Association for Computational Linguistics.
  31. 31.Kevin Stowe, Nils Beck, and Iryna Gurevych. 2021. Exploring metaphoric paraphrase generation. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 323–336, Online. Association for Computational Linguistics.
  32. 32.Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. 2021. Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:2109.07161.
  33. 33.Kohtaro Tanaka, Hiroaki Yamane, Yusuke Mori, Yusuke Mukuta, and Tatsuya Harada. 2022. Learning to evaluate humor in memes based on the incongruity theory. In Proceedings of the Second Workshop on When Creative AI Meets Conversational AI, pages 81–93, Gyeongju, Republic of Korea. Association for Computational Linguistics.
  34. 34.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  35. 35.William Yang Wang and Miaomiao Wen. 2015. I can has cheezburger? a nonparanormal approach to combining textual and visual information for predicting and generating popular meme descriptions. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 355–365, Denver, Colorado. Association for Computational Linguistics.
  36. 36.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  37. 37.Bo Xu, Tingting Li, Junzhe Zheng, Mehdi Naseriparsa, Zhehuan Zhao, Hongfei Lin, and Feng Xia. 2022. Met-meme: A multimodal meme dataset rich in metaphors. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2887–2899.
  38. 38.Ron Yosef, Yonatan Bitton, and Dafna Shahaf. 2023. Irfl: Image recognition of figurative language.
  39. 39.Dongyu Zhang, Minghao Zhang, Heting Zhang, Liang Yang, and Hongfei Lin. 2021. MultiMET: A multimodal dataset for metaphor understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3214–3225, Online. Association for Computational Linguistics.
  40. 40.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pre-trained transformer language models.
  41. 41.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert.
  42. 42.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  43. 43.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. Lima: Less is more for alignment.
  44. 44.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a. Minigpt-4: Enhancing vision-language understanding with advanced large language models.
  45. 45.Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi. 2023b. Multimodal c4: An open, billion-scale corpus of images interleaved with text.

Citation

MLA
Hwang, E., and V. Shwartz. “MemeCap: A Dataset for Captioning and Interpreting Memes”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1433–45, https://doi.org/10.18653/v1/2023.emnlp-main.89.
APA
Hwang, E., & Shwartz, V. (2023). MemeCap: A Dataset for Captioning and Interpreting Memes. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1433–1445. https://doi.org/10.18653/v1/2023.emnlp-main.89
Chicago
Hwang, E., and V. Shwartz. 2023. “MemeCap: A Dataset for Captioning and Interpreting Memes”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1433–45. https://doi.org/10.18653/v1/2023.emnlp-main.89.
Harvard
Hwang, E. and Shwartz, V. (2023) “MemeCap: A Dataset for Captioning and Interpreting Memes”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1433–1445. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.89.
Vancouver
1. Hwang E, Shwartz V (2023) MemeCap: A Dataset for Captioning and Interpreting Memes. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1433–1445

BibTeX

@inproceedings{hwang-shwartz-2023-memecap,
    title = "{M}eme{C}ap: A Dataset for Captioning and Interpreting Memes",
    author = "Hwang, EunJeong  and
      Shwartz, Vered",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.89/",
    doi = "10.18653/v1/2023.emnlp-main.89",
    pages = "1433--1445"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/