MemeCap: A Dataset for Captioning and Interpreting Memes
Eunjeong HwangVered Shwartz
Introduces MEMECAP, a dataset of over six thousand annotated memes with post titles, literal descriptions, and visual metaphors, to expose how modern vision-language models fail to interpret complex multimodal figurative meaning compared to humans.
Internet memes have become a primary medium for online communication, relying heavily on visual metaphors, humor, and shared cultural context to express complex ideas. While modern vision and language artificial intelligence models excel at standard image recognition and descriptive captioning, they struggle to interpret non-literal, figurative content. This limitation creates significant operational blind spots for automated systems tasked with content moderation, social media analytics, and digital communication monitoring, where understanding intent is critical.
The article introduces the generative task of meme captioning and presents MEMECAP, a new benchmark dataset designed to evaluate whether state-of-the-art vision and language models can interpret the intended meaning behind internet memes. The primary objective is to assess how effectively current artificial intelligence architectures comprehend visual metaphors by generating concise natural language explanations of memes and their accompanying post titles.
To establish the benchmark, the researchers collected 6,384 non-offensive memes and their titles from Reddit and crowdsourced multi-stage human annotations through Amazon Mechanical Turk. Annotations included literal image descriptions with embedded meme text removed, identification of metaphorical visual elements and their targets, and human-written captions explaining the underlying message. The dataset was split into training, validation, and a diverse test set of 559 memes. The study then evaluated three leading open-source models—OpenFlamingo-9B, MiniGPT4, and the text-only LLaMA-7B—across zero-shot, few-shot, fine-tuned, and structured reasoning settings, using both standard automated evaluation metrics and blind human assessments.
The evaluation revealed several critical findings regarding machine comprehension of figurative media. First, all models perform substantially worse than human baselines across key evaluation criteria, lagging behind humans by 36.6 percentage points in overall correctness, 29.3 points in textual completeness, 24.5 points in visual completeness, and 18.4 points in factual faithfulness. Second, models frequently generate incorrect interpretations by taking visual elements literally, hallucinating unsupported claims, or simply repeating optical character recognition text rather than explaining meaning. Third, providing few-shot examples and explicit extracted text improved surface fluency and automated overlap scores, but providing explicit metaphorical rationales yielded little to no improvement in interpretation accuracy. Finally, text-only models supplied with textual image descriptions performed competitively with multi-modal vision-language models, indicating that current multi-modal architectures remain inefficient at cross-modal visual reasoning.
These findings demonstrate that organizations cannot rely on current vision-language foundation models for autonomous interpretation of metaphorical digital media. Deploying these systems out-of-the-box for policy enforcement, brand sentiment tracking, or harmful content screening introduces substantial risks of misinterpretation, false positives, and unfaithful hallucinations. The evidence indicates that scaling standard cross-modal pre-training is insufficient; models require dedicated architectures and training objectives tailored to abstract cultural reasoning and visual metaphor resolution.
For technology leaders and developers, the article recommends against fully automated meme analysis workflows, advising instead that automated pipelines include human review for high-stakes decisions. Future engineering efforts should focus on training models to integrate external background knowledge and suppress literal visual descriptions during figurative reasoning. Further research is necessary to explore whether richer multi-modal pre-training data can overcome the limitations of lightweight fine-tuning.
Confidence in these conclusions is high, given the rigorous crowdsourcing design, manual filtering, and consistent alignment between automated metrics and human evaluations. However, readers should consider certain boundary conditions: annotations retain a degree of human subjectivity regarding humor and slang, crowdsourced metaphor-target mappings varied in quality, and the source data was gathered exclusively from English-language Reddit communities, which may limit generalizability across broader cultural and linguistic contexts.
- Paper: Testing the Ability of Language Models to Interpret Figurative Language, Emmy Liu et al. (2022). Its benchmark for interpreting novel figurative language establishes the nonliteral-reasoning challenge that MemeCap extends from text to visual memes.
- Paper: FLUTE: Figurative Language Understanding through Textual Explanations, Tuhin Chakrabarty et al. (2022). FLUTE’s explanation-focused evaluation of metaphors, sarcasm, and other figurative language provides a useful foundation for MemeCap’s test of whether models explain intended meaning.
- Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). Its image-region and language alignment approach grounds the image-captioning tradition that MemeCap adapts to interpreting visual meme content.
- Paper: Microsoft COCO Captions: Data Collection and Evaluation Server, Xinlei Chen et al. (2015). COCO Captions’ dataset design and standardized caption-evaluation methods provide context for MemeCap’s human-annotated benchmark and caption assessments.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). It develops a visual-grounding intervention for reducing multimodal hallucinations, directly pursuing a reliability problem MemeCap exposes in model-generated interpretations.
- Paper: Hallucination Augmented Contrastive Learning for Multimodal Large Language Model, Chaoya Jiang et al. (2024). Its contrastive training method separates grounded descriptions from hallucinated ones, extending MemeCap’s finding that current models often invent unsupported meme meanings.
