Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest

Jack HesselAna MarasovicJena D. HwangLillian LeeJeff DaRowan ZellersRobert MankoffYejin Choi

article2023ACL130 citationsBest Paper Award

Presents a benchmark derived from The New Yorker Cartoon Caption Contest across matching, quality ranking, and explanation tasks, demonstrating that leading multimodal models and GPT-4 fall significantly behind human humor comprehension.

Listen

Modern artificial intelligence models can generate text and jokes with high fluency, yet whether they truly understand complex, nuanced humor remains an open question. Humor in sophisticated editorial cartoons frequently relies on subtle incongruities, indirect associations, and rich cultural knowledge rather than literal scene descriptions. The article addresses this challenge by evaluating whether state-of-the-art vision and language models can comprehend the sophisticated visual-linguistic humor found in The New Yorker Cartoon Caption Contest.

The article's main objective is to establish comprehensive benchmarks that assess AI models across three progressively difficult humor understanding tasks: matching an appropriate caption to a cartoon, ranking caption quality against competing submissions, and generating natural-language explanations of why a joke is funny. To achieve this, the authors curated a 14-year corpus of 704 contests containing millions of entries and crowd ratings. They also authored rich annotations detailing scene locations, literal descriptions, unusual elements, and Wikipedia entity links. The article evaluates models across two setups: a direct image-processing setting (From Pixels) using multimodal systems like CLIP and OFA, and an idealized setting (From Description) that provides human-written scene annotations to advanced language models such as T5, GPT-3, GPT-3.5, and GPT-4.

The findings reveal that state-of-the-art AI systems consistently struggle across all three tasks and exhibit a substantial performance gap when compared to humans. In the image-based matching task, the best multimodal model achieved 62.3% accuracy, lagging roughly 30 percentage points behind the 94.0% human benchmark. When models were provided with full text descriptions to bypass visual perception hurdles, 5-shot GPT-4 achieved 84.5% matching accuracy and outperformed humans at identifying New Yorker editorial selections (68.2% vs. 64.6%), but it underperformed compared to humans at predicting broader crowd preferences (73.3% vs. 83.7%). Crucially, in generating free-form explanations of jokes, human-authored explanations were preferred in head-to-head human evaluations over the best model (5-shot GPT-4) in 67.7% of cases, primarily because models frequently misinterpret core visual relationships or generate plausible yet incorrect reasoning.

These results imply that while AI models possess notable surface fluency and pattern recognition, they remain brittle when required to perform layered, commonsense, and culturally situated reasoning. Computer vision capabilities act as a significant bottleneck, and even when provided perfect visual scene information, large language models cannot reliably interpret the mechanics of complex humor. Consequently, organizations and developers should view current generative AI not as an autonomous judge or creator of nuanced content, but rather as an assistive brainstorming partner in human-in-the-loop creative workflows.

To advance the field, the article recommends deploying matching and quality-ranking models as automated feedback tools for human creators, while exploring reinforcement learning from human feedback to improve humor generation systems. However, readers should recognize key limitations: The New Yorker contest represents a narrow, culturally specific slice of humor, and humor preferences inherently carry subjective bias rather than objective truth. Additionally, the reference explanation dataset was largely developed by a single author, though verified through crowd consensus. While the findings provide high confidence that modern AI lacks genuine humor comprehension, the released open benchmarks and dataset establish a foundation for measuring future progress.

arXiv: 2209.06293
Cover for Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest

Abstract

Large neural networks can now generate jokes, but do they really “understand” humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption, and explaining why a winning caption is funny. These tasks encapsulate progressively more sophisticated aspects of “understanding” a cartoon; key elements are the complex, often surprising relationships between images and captions and the frequent inclusion of indirect and playful allusions to human experience and culture. We investigate both multimodal and language-only models: the former are challenged with the cartoon images directly, while the latter are given multifaceted descriptions of the visual scene to simulate human-level visual understanding. We find that both types of models struggle at all three tasks. For example, our best multimodal models fall 30 accuracy points behind human performance on the matching task, and, even when provided ground-truth visual scene descriptors, human-authored explanations are preferred head-to-head over the best machine-authored ones (few-shot GPT-4) in more than 2/3 of cases. We release models, code, leaderboard, and corpus, which includes newly-gathered annotations describing the image’s locations/entities, what’s unusual in the scene, and an explanation of the joke.

Table of Contents

  • 1 Introduction
  • Matching
  • Quality Ranking
  • Explanation Generation
  • 2 Datasets and Task Setups
  • 2.1 Task Setups
  • 2.2 Annotation of cartoons.
  • 3 Experiments
  • From Pixels (FP) Models
  • From Description (FD) Models
  • Baselines
  • Hardware+software details.
  • 3.1 Matching and quality ranking results
  • 3.2 Human evaluation of explanation.
  • 3.3 Error Analysis for Matching
  • 4 Related Work
  • 5 Conclusion
  • 6 Limitations
  • 7 Acknowledgements
  • References
  • A Crowdworking Details
  • B Additional Experimental Details
  • B.1 From Description details
  • B.2 CLIP
  • B.3 OFA
  • B.4 T5-Large/T5-11B.
  • B.5 GPT-3 Zero Shot/In Context
  • B.6 GPT-3 Fine-tuning
  • B.7 GPT 3.5/GPT-4 Details
  • C Task Construction Details
  • D Graphical version of matching and ranking results.
  • E Automatic evaluation of explanations
  • F Machine explanations that were preferred over human ones
  • G Aiding humor generation with system-assisted brainstorming
  • H Related work beyond peer reviewed AI venues
  • I Some of our favorite New Yorker cartoons

Knowls

  1. Knowl 1 — Three Benchmark Tasks for Cartoon Humor Understanding

    definition

    Humor understanding in single-panel cartoons is formalized into three benchmark tasks of increasing complexity:

    1. Matching (5-way Multiple Choice): Given a cartoon image or its textual description, a model must select the corresponding winning/finalist caption from a set of five options. The four incorrect distractors are finalists randomly sampled from other contests. Distractors are balanced across instances such that each caption appears exactly once as the correct answer and four times as a negative option.

    2. Quality Ranking (Pairwise Choice): Given a cartoon and two candidate captions written specifically for it—one high-quality finalist and one filtered non-finalist distractor—a model must identify which caption is higher quality. Evaluation reports two metrics: NYAcc\text{NYAcc} (accuracy on instances where the positive caption was an official New Yorker editorial finalist) and CrowdAcc\text{CrowdAcc} (accuracy on instances where the positive caption was selected via crowdsourced ratings).

    3. Explanation Generation (Free-form Text Generation): Given a cartoon and a high-quality caption, a model must generate an explanation in natural language of why the image-caption pair is funny. Explanations are evaluated via pairwise human preference judgments against reference human explanations, with ties broken by majority vote.

  2. Knowl 2 — Multimodal Annotations and Explanation Corpus for Cartoon Contests

    experimental setup

    The benchmark dataset compiles 704 weekly cartoon contests from The New Yorker spanning 14 years. To benchmark vision and language models, dense crowdsourced annotations and expert explanations were collected for each cartoon:

    • Setting / Location: 2 short phrases identifying the physical setting (e.g., "doctor's office", "a science lab").
    • Literal Scene Description: 3 literal 1–3 sentence summaries describing the visible objects, characters, and actions.
    • Uncanny Description: 3 1–3 sentence explanations highlighting unusual, incongruous, or out-of-place elements in the image.
    • Entity Links: 2 pairs of English Wikipedia article URLs identifying relevant background entities or concepts.
    • Joke Explanations: 651 free-text explanations authored by an expert annotator (mean length 60 words per cartoon, totaling 39.3K words), explaining the humor mechanism and resolution of incongruity.

    Contests are partitioned into 5 cross-validation splits holding out entire contests at test time, resulting in approximately 1,600 training, 538 validation, and 538 test instances for matching; 1,600 training, 523 validation, and 523 test instances for quality ranking; and 391 training, 130 validation, and 130 test instances for explanation.

  3. Knowl 3 — Curating Controlled Negative Captions for Quality Ranking

    algorithm

    To build balanced pairwise comparisons for the quality ranking task, candidate negative captions must be reasonable ("okay") rather than trivially ungrammatical or nonsensical. The candidate filtering and deduplication procedure is structured as follows:

    Input: Raw candidate submissions SS for cartoon CC, trained caption-quality classifier MQM_Q, SentenceBERT encoder ESBERTE_{SBERT}
    Output: A matched distractor caption cdistractorc_{distractor}
    1. If crowdsourced ratings are available for CC:
           Select captions in the middle score tertile, discarding the top 1/3 and bottom 1/3 to obtain pool SfilteredS_{filtered}.
       Else:
           Score all s∈Ss \in S using MQM_Q (a binary text-only classifier trained on 506K top/bottom crowd-rated captions).
           Discard the lowest-scoring 25% of entries to obtain SfilteredS_{filtered}.
    2. Deduplication:
           Remove exact string matches to high-quality finalist captions.
           Compute embeddings vi=ESBERT(si)v_i = E_{SBERT}(s_i) for all si∈Sfiltereds_i \in S_{filtered}.
           Perform hierarchical clustering of SfilteredS_{filtered} into 0.7⋅∣Sfiltered∣0.7 \cdot |S_{filtered}| clusters and retain one exemplar per cluster (reducing redundancy by 30%).
    3. Controlled Pair Matching:
           For each target high-quality caption c∗c^*, sample cdistractorc_{distractor} from the deduplicated pool satisfying:
           - Similar predicted quality score from MQM_Q
           - Similar word length
           - Similar character length
           - Similar punctuation frequency
           - Dissimilar ESBERTE_{SBERT} cosine representation relative to c∗c^*
    return cdistractorc_{distractor}
  4. Knowl 4 — From Pixels vs. From Description Evaluation Regimes

    experimental setup

    Models are evaluated under two experimental paradigms to differentiate visual recognition capabilities from multimodal reasoning:

    1. From Pixels (FP): Models receive only raw cartoon images and candidate texts. Tested architectures include:

      • Fine-tuned CLIP ViT-L/14@336px (428M parameters), optimized for multiple-choice ranking using InfoNCE loss on cosine similarities between cartoon images and captions.
      • OFA-Huge (930M parameters) fine-tuned to predict setting, literal description, uncanny elements, and entity links from images, whose predicted outputs are then fed into a frozen/fine-tuned language model (T5-Large or T5-11B) in a Socratic Model framework.
    2. From Description (FD): Models receive human-authored text annotations (location, literal description, uncanny description, entity links) instead of the raw image, simulating oracle vision. Tested architectures include:

      • Fine-tuned T5-Large (770M) and T5-11B (11.3B).
      • OpenAI GPT-3 (text-davinci-002, 175B), GPT-3.5, and GPT-4 evaluated under zero-shot, few-shot (5 in-context examples), fine-tuned, and Chain-of-Thought (CoT) prompting.
  5. Knowl 5 — Model Performance on Caption Matching and Quality Ranking

    data/table

    The table below summarizes model accuracy across 5 cross-validation splits for 5-way matching accuracy and 2-way quality ranking (predicting crowd preference CrowdAcc\text{CrowdAcc} and New Yorker editorial choice NYAcc\text{NYAcc}):

    Model / Evaluation Setting Matching Acc (%) CrowdAcc (%) NYAcc (%)
    Random baseline 20.0 50.0 50.0
    Caption Only (T5-11B finetuned) 19.4 59.4 64.5
    From Pixels (FP)
    CLIP ViT-L/14@336px (finetuned) 62.3 57.0 66.9
    Zero-shot 56.6 55.8 56.8
    OFA-Huge →\rightarrow T5-Large 45.2 59.1 64.3
    OFA-Huge →\rightarrow T5-11B 51.8 60.3 65.0
    From Description (FD)
    T5-Large (finetuned) 59.6 61.8 64.8
    T5-11B (finetuned) 70.8 62.3 65.6
    GPT-3 175B (finetuned) 75.1 64.8 69.8
    5-shot 57.2 55.1 54.8
    Zero-shot 51.6 56.2 55.6
    GPT-3.5 (5-shot) 63.8 55.6 55.2
    Zero-shot + CoT 50.4 52.8 55.4
    GPT-4 (5-shot) 84.5 73.3 68.2
    Zero-shot + CoT 81.9 66.2 64.3
    Human Estimate (From Pixels) 94.0 83.7 64.6

    Key empirical findings:

    • Human-Machine Gap in Vision: In the From Pixels setting, the top multimodal model (fine-tuned CLIP at 62.3%) trails human accuracy (94.0%) by 31.7 percentage points on matching.
    • Oracle Vision Reasoning: Providing oracle human visual descriptions substantially improves performance, with 5-shot GPT-4 achieving 84.5% matching accuracy.
    • Editorial vs. Crowd Preferences: Language models (e.g., fine-tuned GPT-3 at 69.8% and 5-shot GPT-4 at 68.2%) outperform human test-takers (64.6%) at predicting New Yorker editor picks (NYAcc\text{NYAcc}), but lag behind humans (83.7%) at predicting broader crowdsourced preferences (CrowdAcc\text{CrowdAcc}, where GPT-4 achieves 73.3%).
    • Multimodal Interaction: Caption-only models fail at matching (19.4% vs 20.0% random baseline), demonstrating that matching requires joint image-text reasoning.
  6. Knowl 6 — Human Evaluation of Free-Form Joke Explanation Quality

    data/table

    Pairwise human evaluations were conducted with 3 crowdsourced annotators per comparison (majority vote determines winner), assessing generated explanations across several experimental conditions. Agreement is measured using Gwet's γ\gamma:

    Condition A Condition B A Win Rate (%) # Ratings Gwet's γ\gamma
    T5-11B (FD) Caption-only T5-11B 84.7 393 64.4
    T5-11B (FD) OFA →\rightarrow T5-11B (FP) 74.6 393 41.6
    T5-11B (FD) T5-Large (FD) 68.5 390 45.9
    Fine-tuned GPT-3 5-shot GPT-3 50.0 396 23.2
    5-shot GPT-4 Zero-shot GPT-4 64.3 396 19.7
    5-shot GPT-4 5-shot GPT-3 93.0 384 86.4
    Human Reference 5-shot GPT-4 67.7 390 20.9

    Key empirical findings:

    • Image context is vital: T5-11B with description access wins 84.7% over caption-only T5-11B.
    • Visual recognition is a major bottleneck: human-written descriptions win 74.6% of pairwise judgments against machine-generated descriptions from OFA-Huge.
    • Model scale yields superior explanations: 5-shot GPT-4 decisively beats 5-shot GPT-3 (93.0% win rate).
    • Human references remain significantly preferred over the strongest model (5-shot GPT-4 wins only 32.3% against humans; human win rate is 67.7%). Failure modes of AI explanations include hallucinating non-existent objects, confusing agent roles, or missing underlying cultural and literary references.
  7. Knowl 7 — Divergence Between Automatic Surface Metrics and Human Preference in Humor Explanation

    empirical result

    Automatic evaluation metrics on generated humor explanations do not align with human quality judgments:

    • Surface Overlap vs. Human Preference: In 5-shot evaluation, GPT-3 achieves a higher BLEU-4 score (5.07) and ROUGE-L score (20.5) than 5-shot GPT-4 (4.99 BLEU-4, 20.0 ROUGE-L). However, in pairwise human evaluation, 5-shot GPT-4 explanations are preferred over 5-shot GPT-3 in 93.0% of cases.
    • Perplexity vs. Quality: Fine-tuning GPT-3 substantially lowers test perplexity relative to 5-shot in-context learning (21.8 vs. 107.0) and yields higher BLEU-4 (5.42 vs. 5.07). Despite this, human evaluators show no preference between fine-tuned and 5-shot GPT-3 explanations (50.0% win rate).

    N-gram overlap metrics (BLEU, ROUGE) and language modeling perplexity reward superficial adherence to the phrasing style of the training corpus rather than the logical coherence, nuance, and semantic correctness of the joke explanation.

  8. Knowl 8 — Error Correlation and Cartoon Difficulty Across Vision and Language Models

    empirical result

    Statistical analysis of model errors on the caption matching task across 704 cartoons demonstrates non-random, contest-level difficulty:

    • Contest Clustering: A χ2\chi^2 contingency test shows that errors cluster significantly by cartoon contest (p<0.05p < 0.05 for both fine-tuned CLIP and fine-tuned GPT-3). Errors show no significant correlation with cross-validation split (p=0.84p = 0.84 for GPT-3, p=0.14p = 0.14 for CLIP) or distractor choice assignment (p=0.92p = 0.92 for GPT-3, p=0.79p = 0.79 for CLIP).
    • Multimodal Error Correlation: Per-contest accuracy between CLIP (From Pixels) and GPT-3 (From Description) is positively correlated (Spearman ρ=0.28,p<0.001\rho = 0.28, p < 0.001). When both models agree, their joint prediction accuracy is 87%. When GPT-3 fails, CLIP's accuracy drops from its baseline of 62.3% down to 38.0% (p<0.001p < 0.001, permutation test), indicating shared conceptual bottlenecks across modalities.
    • Unpredictability of Difficulty: Standard visual and linguistic features cannot predict hard vs. easy contests a priori; classification models trained to predict contest difficulty perform only marginally above chance.
  9. Knowl 9 — Concept-Association Staged Prompting for Creative Humor Generation

    model/method

    To assist human cartoonists and explore creative generation, the annotation corpus can be reframed into an in-context brainstorming prompt. The prompt structure utilizes associative reasoning in 7 distinct sequential steps:

    1. Location: Setting of the cartoon scene.
    2. Literal Description: Summary of the visible scene elements.
    3. Uncanny Elements: Specific oddities or incongruities in the scene.
    4. Scene Entities: Named entities and objects present.
    5. Brainstormed Concepts: Free associations, idioms, cultural tropes, or metaphors related to the scene elements.
    6. Concept Subset Selection: A targeted subset of 1–3 concepts selected from the brainstormed list to anchor the joke.
    7. Caption and Explanation: The synthesized humorous caption combining the selected concepts, followed by a rationalization of why the caption fits the scene.

    This structured framing allows language models to operate either unconditionally (generating novel cartoon concepts alongside matching captions) or conditionally (generating humorous captions when supplied with a description of a new cartoon scene).

  10. Knowl 10 — Benchmark Limitations and Cultural Specificity of Humor Understanding

    limitation

    The benchmark exhibits specific scope and methodological constraints:

    1. Cultural and Stylistic Domain Narrowness: The dataset is drawn exclusively from The New Yorker Cartoon Caption Contest, representing a specific subgenre of humor grounded in American English, specific cultural/literary tropes, and sophisticated incongruity resolution. It does not generalize to other humor modalities (e.g., slapstick, internet memes, cross-cultural satire).
    2. Subjectivity of Quality Targets: Quality ranking evaluates aggregate human preferences rather than an absolute ground truth for funniness. Target labels reflect historical editorial choices and crowdsourced averages, which may encode institutional biases toward particular stylistic tropes.
    3. Single-Author Reference Explanations: The ground-truth explanation corpus (651 instances) was written by a single expert annotator to maintain consistency and quality. As a result, interpersonal variance in how humans interpret, appreciate, and explain the same joke is not modeled.

Coverage note — None was omitted; all primary contributions—including task definitions, dataset curation and dense annotations, negative filtering algorithms, From Pixels vs From Description experiments, human evaluations of explanations, automatic metric analysis, error analysis, brainstorming prompting, and limitations—are fully covered.

References

  1. 1.Miriam Amin and Manuel Burghardt. 2020. A survey on approaches to computational humor generation. In The 4th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature.
  2. 2.Issa Annamoradnejad and Gohar Zoghi. 2020. ColBERT: Using BERT sentence embedding for humor detection. arXiv preprint arXiv:2004.12765.
  3. 3.Salvatore Attardo. 2008. A primer for the linguistics of humor. The primer of humor research, 8:101–55.
  4. 4.Michael Billig. 2005. Laughter and ridicule: Towards a social critique of humour. Sage.
  5. 5.Kim Binsted and Graeme Ritchie. 1994. An implemented model of punning riddles. In AAAI.
  6. 6.Vladislav Blinov, Valeria Bolotova-Baranova, and Pavel Braslavski. 2019. Large dataset and language model fun-tuning for humor recognition. In ACL.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. NeurIPS.
  8. 8.Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards multimodal sarcasm detection (an Obviously perfect paper). In ACL.
  9. 9.Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, and Smaranda Muresan. 2022. FLUTE: figurative language understanding and textual explanations. In EMNLP.
  10. 10.Arjun Chandrasekaran, Devi Parikh, and Mohit Bansal. 2018. Punny captions: Witty wordplay in image descriptions. In NAACL.
  11. 11.Arjun Chandrasekaran, Ashwin K. Vijayakumar, Stanislaw Antol, Mohit Bansal, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2016. We are humor beings: Understanding and predicting visual humor. In CVPR.
  12. 12.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR.
  14. 14.Fallianda, Rani Yuni Astiti, and Zulvy Alivia Hanim. 2018. Analyzing humor in newspaper comic strips using verbal-visual analysis. Lingua Cultura, 12(4):383–388.
  15. 15.Christiane Fellbaum. 1998. WordNet: An Electronic Lexical Database. Bradford Books.
  16. 16.Sigmund Freud. 1905. Jokes and their Relation to the Unconscious, volume 8 of The Standard Edition of the Complete Psychological Works of Sigmund Freud. Hogarth, London.
  17. 17.William F. Fry. 1963. Sweet madness: A study of humor. Pacific Books, Palo Alto.
  18. 18.Charles R. Gruner. 1978. Understanding laughter: The workings of wit & humor. Nelson-Hall, Chicago.
  19. 19.Kilem Gwet. 2014. Handbook of Inter-Rater reliability: The Definitive Guide to Measuring the Extent of Agreement Among Raters, 4th edition edition. Advanced Analytics, LLC.
  20. 20.Md Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque. 2021. Humor knowledge enriched transformer for understanding multimodal humor. In AAAI.
  21. 21.Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, and Mohammed (Ehsan) Hoque. 2019. UR-FUNNY: a multimodal language dataset for understanding humor. In EMNLP.
  22. 22.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In ICLR.
  23. 23.Lalit Jain, Kevin Jamieson, Robert Mankoff, Robert Nowak, and Scott Sievert. 2020. The New Yorker cartoon caption contest dataset.
  24. 24.Kevin G. Jamieson, Lalit Jain, Chris Fernandez, Nicholas J. Glattard, and Rob Nowak. 2015. NEXT: A system for real-world development, evaluation, and application of active learning. In NeurIPS.
  25. 25.Ben King, Rahul Jha, Dragomir Radev, and Robert Mankoff. 2013. Random walk factoid annotation for collective discourse. In ACL.
  26. 26.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In NeurIPS.
  27. 27.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out.
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In ECCV.
  29. 29.Tania Lombrozo. 2006. The structure and function of explanations. Trends in Cognitive Sciences, 10(10):464–470.
  30. 30.Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2022. Few-shot self-rationalization with natural language prompts. In Findings of NAACL.
  31. 31.Rada Mihalcea and Stephen Pulman. 2009. Characterizing humour: An exploration of features in humorous texts. In Proceedings of the 8th International Conference on Computational Linguistics and Intelligent Text Processing, page 337–347, Berlin, Heidelberg. Springer-Verlag.
  32. 32.Rada Mihalcea and Carlo Strapparava. 2005. Making computers laugh: Investigations in automatic humor recognition. In EMNLP.
  33. 33.Rada Mihalcea and Carlo Strapparava. 2006. Technologies that make you smile: Adding humor to text-based applications. IEEE Intelligent Systems, 21(5):33–39.
  34. 34.Harvey Mindess. 1971. Laughter and Liberation. Nash.
  35. 35.Pamela Mishkin, Matt Daniels, Russell Goldenberg, Ilia Blinderman, and James Yu. 2022. The pudding caption contest experiments. https://pudding.cool/projects/caption-contest/. Accessed: 2022-04-01.
  36. 36.Matthijs P. Mulder and Antinus Nijholt. 2002. Humour research: State of the art. Centre for Telematics and Information Technology, University of Twente.
  37. 37.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.
  38. 38.OpenAI. 2023. Gpt-4 technical report.
  39. 39.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In ACL.
  40. 40.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS.
  41. 41.Badri N. Patro, Mayank Lunayach, Deepankar Srivastava, Sarvesh, Hunar Singh, and Vinay P. Namboodiri. 2021. Multimodal humor dataset: Predicting laughter tracks for sitcoms. In WACV.
  42. 42.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In *SEM.
  43. 43.Matt Post. 2018. A call for clarity in reporting BLEU scores. In WMT.
  44. 44.Dragomir Radev, Amanda Stent, Joel Tetreault, Aasish Pappu, Aikaterini Iliakopoulou, Agustin Chanfreau, Paloma de Juan, Jordi Vallmitjana, Alejandro Jaimes, Rahul Jha, and Robert Mankoff. 2016. Humor in collective discourse: Unsupervised funniness detection in the New Yorker cartoon caption contest. In LREC.
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In ICML.
  46. 46.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR.
  47. 47.Victor Raskin. 1979. Semantic mechanisms of humor. In Annual Meeting of the Berkeley Linguistics Society, volume 5, pages 325–335.
  48. 48.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD.
  49. 49.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In EMNLP.
  50. 50.Arthur Schopenhauer. 1818. The world as will and idea, volume 1.
  51. 51.Dafna Shahaf, Eric Horvitz, and Robert Mankoff. 2015. Inside jokes: Identifying humorous cartoon captions. In KDD.
  52. 52.Erica K Shimomoto, Lincon S Souza, Bernardo B Gatto, and Kazuhiro Fukui. 2019. News2meme: An automatic content generator from news based on word subspaces from text and image. In Conference on Machine Vision Applications.
  53. 53.Thomas R Shultz. 1976. A cognitive-developmental analysis of humour. Transaction Publishers.
  54. 54.Oliviero Stock and Carlo Strapparava. 2003. Getting serious about the development of computational humor. In IJCAI.
  55. 55.Rajesh Shanmuga Sundaram. 2018. Generation of Humorous Caption for Cartoon Images Using Deep Learning. Ph.D. thesis, Texas A&M University-Commerce.
  56. 56.Chenhao Tan. 2022. On the diversity and limits of human explanations. In NAACL.
  57. 57.Ervin Tanczos, Robert Nowak, and Bob Mankoff. 2017. A KL-LUCB algorithm for large-scale crowdsourcing. In NeurIPS.
  58. 58.Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, and Jackie Chi Kit Cheung. 2019. How reasonable are common-sense reasoning tasks: A case-study on the Winograd schema challenge and SWAG. In EMNLP.
  59. 59.Villy Tsakona. 2009. Language and image interaction in cartoons: Towards a multimodal theory of humor. Journal of Pragmatics, 41(6):1171–1188.
  60. 60.Alessandro Valitutti. 2011. How many jokes are really funny? In Human-Machine Interaction in Translation: Proceedings of the 8th International NLPCS Workshop.
  61. 61.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. NeurIPS.
  62. 62.David Wallace. 2022. Lecture notes for MIT 2.00b toy product design: Innovation and associations.
  63. 63.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In ICML.
  64. 64.William Yang Wang and Miaomiao Wen. 2015. I can has cheezburger? a nonparanormal approach to combining textual and visual information for predicting and generating popular meme descriptions. In NAACL.
  65. 65.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  66. 66.White, E. B. 1941. Preface. In E. B. White and Katherine S. White, editors, A Subtreasury Of American Humor, page xvii. The original version of this quote appeared as a preview in The Saturday Review (1941), credited to both Whites. But, the quote appears in the preface to A Subtreasury (1941) with authorship solely credited to E.B.. We thus credited the quote itself to E.B., and credited both E.B. and K.S. as editors of the anthology in which it appears in non-preview form.
  67. 67.Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2022. Reframing human-AI collaboration for generating free-text explanations. In NAACL.
  68. 68.Hannah Wilson. 2019. Project four - nobody knows you’re a bot.
  69. 69.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In EMNLP: System Demonstrations.
  70. 70.Kota Yoshida, Munetaka Minoguchi, Kenichiro Wani, Akio Nakamura, and Hirokatsu Kataoka. 2018. Neural joking machine: Humorous image captioning. In CVPR Language & Vision Workshop.
  71. 71.Michael Zelenko and Frank Bi. 2015. On the internet, nobody knows you’re a machine.
  72. 72.Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence. 2022. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598.

Citation

MLA
Hessel, J., et al. “Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 688–714, https://doi.org/10.18653/v1/2023.acl-long.41.
APA
Hessel, J., Marasović, A., Hwang, J. D., Lee, L., Da, J., Zellers, R., Mankoff, R., & Choi, Y. (2023). Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 688–714. https://doi.org/10.18653/v1/2023.acl-long.41
Chicago
Hessel, J., A. Marasović, J. D. Hwang, et al. 2023. “Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 688–714. https://doi.org/10.18653/v1/2023.acl-long.41.
Harvard
Hessel, J. et al. (2023) “Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 688–714. Available at: https://doi.org/10.18653/v1/2023.acl-long.41.
Vancouver
1. Hessel J, Marasović A, Hwang JD, Lee L, Da J, Zellers R, Mankoff R, Choi Y (2023) Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 688–714

BibTeX

@inproceedings{hessel-etal-2023-androids,
    title = "Do Androids Laugh at Electric Sheep? Humor ``Understanding'' Benchmarks from The New Yorker Caption Contest",
    author = "Hessel, Jack  and
      Marasovic, Ana  and
      Hwang, Jena D.  and
      Lee, Lillian  and
      Da, Jeff  and
      Zellers, Rowan  and
      Mankoff, Robert  and
      Choi, Yejin",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.41/",
    doi = "10.18653/v1/2023.acl-long.41",
    pages = "688--714"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/