Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models
Lennart WachowiakDagmar Gromann
Evaluates GPT-3's capacity to identify conceptual metaphors and predict unconstrained source domains across English and Spanish, providing a detailed error analysis of generative language models handling cross-domain cognitive mappings.
Metaphorical language is central to human communication, cognition, and persuasion, making its comprehension vital for building advanced artificial intelligence. Most prior computational approaches have focused strictly on binary detection—identifying whether a phrase is literal or figurative—or on paraphrasing metaphors using pre-defined categories and fixed grammatical rules. To probe deeper cognitive knowledge in generative language models, the article evaluates whether GPT-3 can independently identify underlying conceptual metaphors by predicting their concrete source domains when given a sentence and a target domain, without relying on fixed domain labels or grammatical constraints.
The authors conducted experiments comparing fine-tuning against few-shot prompting across multiple model sizes and varying prompt sample sizes (2 to 12 examples). Performance was assessed on a curated validation set using embedding similarity and knowledge graph similarity metrics to select the optimal model setup. The primary evaluation was carried out on a hold-out test set consisting of 633 sentences across two languages and multiple corpora: Lakoff's Master Metaphor List (general English), the VU Amsterdam Metaphor Corpus (non-metaphoric control sentences), and the Language Computer Corporation (LCC) dataset (complex political discourse in English and Spanish). Two human evaluators manually verified the correctness of the generated source domains.
The analysis showed that the largest model, davinci-002, when provided with a 12-example few-shot prompt, substantially outperformed smaller variants and fine-tuned models, achieving an overall sample-weighted accuracy of 60.22%. Performance varied widely across datasets: the model reached 81.33% accuracy on prototypical English sentences, 53.74% on complex English political texts, and 34.65% on Spanish texts, while correctly identifying non-metaphoric statements with 42.11% accuracy. Fine-tuning models reduced output diversity, causing the system to rely heavily on training examples and predict fewer unique source domains. Qualitative error analysis revealed that the most common failure was hallucinating unrelated source domains (27.32% of English errors), followed by failing to detect metaphors in figurative sentences (25.14%) and misinterpreting non-metaphorical trigger words (21.31%). In Spanish, 62.12% of errors were caused by incorrectly classifying metaphorical sentences as non-metaphoric.
These findings indicate that while large language models encode notable metaphoric knowledge, their capability degrades significantly on complex, domain-specific language and cross-lingual inputs. Fine-tuning off-the-shelf models introduces risks of reduced expressiveness, and the high rate of hallucination presents reliability risks in downstream applications like discourse monitoring or automated content generation. Organizations seeking to deploy generative models for semantic analysis cannot rely on unguided zero-shot or fine-tuned classification and must account for significant performance drops outside prototypical English texts.
To apply generative metaphor extraction effectively in production or research, practitioners should implement multi-stage pipelines: using seed words to filter target domains first, restricting context windows to avoid distraction from secondary metaphors, and incorporating few-shot prompts rather than standard fine-tuning. Further research should focus on developing automated evaluation benchmarks, predicting both source and target domains simultaneously, and building diverse multilingual datasets.
Confidence in these findings is supported by human verification showing moderate inter-annotator agreement (Cohen's Kappa of 0.51). However, users should remain cautious due to inherent data limitations: the benchmark datasets feature demographic and domain skews, manual annotation carries subjective ambiguity, and the closed, black-box nature of commercial API models limits architectural transparency and exact reproducibility.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Read this account of GPT-3’s few-shot learning first to understand the model and prompting paradigm the metaphor study evaluates.
- Paper: Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages, Ehsan Aghazadeh et al. (2022). Its probing of metaphor knowledge across datasets and languages establishes the earlier question of whether language models represent generalizable metaphorical structure.
- Paper: Testing the Ability of Language Models to Interpret Figurative Language, Emmy Liu et al. (2022). Its evaluation of language models interpreting figurative expressions provides a direct precursor to testing whether GPT-3 can recover metaphor mappings.
- Paper: FLUTE: Figurative Language Understanding through Textual Explanations, Tuhin Chakrabarty et al. (2022). FLUTE’s benchmark for explaining figurative language grounds the source’s effort to evaluate metaphor understanding beyond simple detection.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Its analysis of GPT-3 prompting clarifies the few-shot method the source uses to elicit metaphor mappings without fine-tuning.
No sufficiently relevant recommendations were found.
