Testing the Ability of Language Models to Interpret Figurative Language
Emmy LiuChenxuan CuiKenneth ZhengGraham Neubig
Figurative language, such as metaphors and similes, appears frequently in human communication to convey complex ideas, emotions, and humor. Because nonliteral expressions rely heavily on shared commonsense, social, and cultural knowledge rather than strict word definitions, interpreting them presents a major challenge for artificial intelligence. Most prior natural language processing research focuses on detecting common, conventional metaphors rather than interpreting creative, novel figurative phrasing. As language models are increasingly deployed in real-world conversational and text-processing applications, understanding their ability to infer nonliteral meaning is critical.
The main objective of the article is to evaluate how well state-of-the-art language models can understand and interpret novel, human-generated figurative language compared to human capability. The study specifically investigates whether pre-trained models can select the correct literal interpretation of a creative metaphor and generate valid explanations on their own.
To evaluate this, the authors introduced Fig-QA, a Winograd-style benchmark containing 10,256 validated sentence pairs with opposing metaphorical meanings and corresponding literal implications. The dataset was crowdsourced to specifically elicit creative metaphors that rarely appear on the internet or in training data, categorized across physical commonsense, visual, social, and cultural domains. The researchers evaluated several prominent auto-regressive models (GPT-2, GPT-neo, and multiple GPT-3 variants) and masked language models (BERT and RoBERTa) across zero-shot, few-shot prompting, and fine-tuned configurations. In addition, human baseline testing was conducted with native and non-native English speakers, alongside a manual qualitative evaluation of open-ended interpretations generated by GPT-3.
The study revealed several critical findings regarding model performance and behavior. First, while untrained models perform above random chance, a substantial capability gap exists between models and humans in zero-shot settings: the highest-performing zero-shot model (GPT-3 Davinci) achieved 68.4% accuracy, trailing the human baseline of 94.4% by roughly 26 percentage points. Second, fine-tuning substantially improves performance across all architectures, with fine-tuned RoBERTa reaching 90.3% accuracy, coming within about 4 percentage points of human performance. Third, auto-regressive models exhibit severe contextual blindness, relying heavily on the standalone probability of the answer options rather than evaluating the metaphorical context; this caused their accuracy to collapse (falling below 20% for zero-shot GPT models) under strict paired evaluation where both opposing sentences in a pair must be answered correctly. Fourth, in open-ended generation, GPT-3 produced accurate interpretations only about 51% to 64% of the time and generated self-contradictory statements across multiple attempts in roughly 38% of cases. Finally, models struggle disproportionately with sarcastic metaphors and cultural commonsense, frequently making basic comprehension errors that diverge sharply from human error patterns.
These findings indicate that while fine-tuned models can perform well on constrained multiple-choice selections, current language models lack genuine, robust nonliteral reasoning. In operational environments, relying on zero-shot language models to parse figurative, sarcastic, or culturally embedded communication introduces notable safety, performance, and compliance risks, such as misinterpreting customer intent, misunderstanding social nuances, or generating contradictory summaries. The fact that models lean on surface-level word probabilities rather than true semantic comprehension means standard evaluation benchmarks may overstate their practical language understanding.
Organizations deploying natural language systems should exercise caution when processing unstructured human discourse containing figurative expressions. Decision-makers should avoid relying on zero-shot auto-regressive models for sensitive semantic interpretation tasks and should instead implement contrastive fine-tuning approaches or supervised classifiers where feasible. For future research and development, the authors recommend expanding benchmark evaluations into multimodal contexts and non-English languages, as well as developing training methodologies that explicitly ground models in analogical and commonsense reasoning.
Confidence in these findings is high for standard English text, supported by rigorous manual validation, multiple baseline architectures, and paired Winograd-style controls. However, limitations include the benchmark’s geographic focus on United States cultural references, the moderate inter-rater agreement inherent to subjective generation tasks, and the restriction to text-only English data. Readers should remain cautious about generalizing these conclusions directly to multilingual settings or multimodal domains without further empirical validation.
- Paper: Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data, Emily M. Bender et al. (2020). Its distinction between surface language form and grounded meaning gives essential context for evaluating whether models’ figurative interpretations reflect understanding or pattern matching.
- Paper: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest, Jack Hessel et al. (2023). It carries the evaluation of nuanced nonliteral language into humor, testing whether models can interpret and explain culturally rich jokes across text and images.
- Paper: Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests, Max J. van Duijn et al. (2023). It extends evaluation of nonliteral communication to sarcasm and other social reasoning, testing whether newer models can interpret intent across altered scenarios.
- Paper: When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues, Shivani Kumar et al. (2022). It continues the study of figurative interpretation by asking models to explain sarcasm in dialogue using conversational, audio, and visual cues.