Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
Jack HesselAna MarasovicJena D. HwangLillian LeeJeff DaRowan ZellersRobert MankoffYejin Choi
Presents a benchmark derived from The New Yorker Cartoon Caption Contest across matching, quality ranking, and explanation tasks, demonstrating that leading multimodal models and GPT-4 fall significantly behind human humor comprehension.
Modern artificial intelligence models can generate text and jokes with high fluency, yet whether they truly understand complex, nuanced humor remains an open question. Humor in sophisticated editorial cartoons frequently relies on subtle incongruities, indirect associations, and rich cultural knowledge rather than literal scene descriptions. The article addresses this challenge by evaluating whether state-of-the-art vision and language models can comprehend the sophisticated visual-linguistic humor found in The New Yorker Cartoon Caption Contest.
The article's main objective is to establish comprehensive benchmarks that assess AI models across three progressively difficult humor understanding tasks: matching an appropriate caption to a cartoon, ranking caption quality against competing submissions, and generating natural-language explanations of why a joke is funny. To achieve this, the authors curated a 14-year corpus of 704 contests containing millions of entries and crowd ratings. They also authored rich annotations detailing scene locations, literal descriptions, unusual elements, and Wikipedia entity links. The article evaluates models across two setups: a direct image-processing setting (From Pixels) using multimodal systems like CLIP and OFA, and an idealized setting (From Description) that provides human-written scene annotations to advanced language models such as T5, GPT-3, GPT-3.5, and GPT-4.
The findings reveal that state-of-the-art AI systems consistently struggle across all three tasks and exhibit a substantial performance gap when compared to humans. In the image-based matching task, the best multimodal model achieved 62.3% accuracy, lagging roughly 30 percentage points behind the 94.0% human benchmark. When models were provided with full text descriptions to bypass visual perception hurdles, 5-shot GPT-4 achieved 84.5% matching accuracy and outperformed humans at identifying New Yorker editorial selections (68.2% vs. 64.6%), but it underperformed compared to humans at predicting broader crowd preferences (73.3% vs. 83.7%). Crucially, in generating free-form explanations of jokes, human-authored explanations were preferred in head-to-head human evaluations over the best model (5-shot GPT-4) in 67.7% of cases, primarily because models frequently misinterpret core visual relationships or generate plausible yet incorrect reasoning.
These results imply that while AI models possess notable surface fluency and pattern recognition, they remain brittle when required to perform layered, commonsense, and culturally situated reasoning. Computer vision capabilities act as a significant bottleneck, and even when provided perfect visual scene information, large language models cannot reliably interpret the mechanics of complex humor. Consequently, organizations and developers should view current generative AI not as an autonomous judge or creator of nuanced content, but rather as an assistive brainstorming partner in human-in-the-loop creative workflows.
To advance the field, the article recommends deploying matching and quality-ranking models as automated feedback tools for human creators, while exploring reinforcement learning from human feedback to improve humor generation systems. However, readers should recognize key limitations: The New Yorker contest represents a narrow, culturally specific slice of humor, and humor preferences inherently carry subjective bias rather than objective truth. Additionally, the reference explanation dataset was largely developed by a single author, though verified through crowd consensus. While the findings provide high confidence that modern AI lacks genuine humor comprehension, the released open benchmarks and dataset establish a foundation for measuring future progress.
- Paper: Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data, Emily M. Bender et al. (2020). Establishes the foundational theoretical distinction between linguistic form and communicative meaning, providing the conceptual motivation for testing whether language models truly comprehend nuanced humor rather than mere surface patterns.
- Paper: Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network, Bin Liang et al. (2022). Pioneers the modeling of cross-modal incongruities and opposing polarities between visual regions and text cues, which directly underpins visual-linguistic humor analysis.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Introduces multimodal chain-of-thought prompting and free-form explanation generation, forming the methodological basis for probing multi-step reasoning across paired images and text.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). Demonstrates the necessity of chain-of-thought prompting to unlock latent multi-step reasoning capabilities in large language models on challenging human-level benchmarks.
- Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). Develops the paradigm of testing implicit commonsense and world knowledge in AI models, a core prerequisite for resolving implied editorial cartoon humor.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). Explores multimodal representation learning specifically tailored for humor detection and sentiment analysis across heterogeneous data streams.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). Expands on visual humor explanation and complex multi-capability evaluation by introducing an integrated multimodal benchmark across open-ended reasoning tasks.
- Paper: Benchmarking Vision Language Models for Cultural Understanding, Shravan Nayak et al. (2024). Directly extends the study of culturally situated visual commonsense by benchmarking vision-language models on diverse international traditions and cultural nuances.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). Advances the evaluation of fine-grained multimodal perception and reasoning through a comprehensive bilingual benchmark that addresses option bias and robust response parsing.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). Investigates whether pairing open-source vision encoders with large language models can replicate advanced visual capabilities like meme and humor explanation.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). Diagnoses the perceptual bottleneck and object hallucination in vision-language models that fundamentally obstruct accurate visual scene understanding in cartoon and image tasks.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Critically surveys systemic obstacles in evaluating generated text and human preference alignment, addressing key limitations encountered in free-form explanation benchmarks.
- Paper: Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection, Yang Qiao et al. (2023). Builds upon multimodal incongruity detection by developing a unified network that models fine-grained local and global visual-textual contradictions.
