Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages
Ehsan AghazadehMohsen FayyazYadollah Yaghoobzadeh
Demonstrates through probing and cross-lingual experiments that pre-trained language models capture transferable metaphorical knowledge concentrated within their middle layers across multiple datasets and languages.
Metaphorical language is central to human communication and reasoning, allowing people to comprehend complex or unfamiliar concepts by linking them to concrete domains. While large pre-trained language models serve as standard foundations across language technologies, it has remained unproven whether these systems genuinely internalize structured metaphorical knowledge or merely exploit superficial statistical shortcuts. Understanding whether language models capture generalizable metaphor understanding is critical for deploying artificial intelligence systems that can accurately interpret nuanced human language.
The main objective of the article is to empirically evaluate whether pre-trained language models encode metaphorical knowledge in their internal representations and to measure how effectively this knowledge generalizes across different datasets and languages.
To conduct this evaluation, the researchers applied two diagnostic probing methods—edge probing and minimum description length probing—which measure how readily specific linguistic information can be extracted from frozen model layers. The investigation assessed widely used language models (BERT, RoBERTa, and ELECTRA) and a multilingual model (XLM-R) across four established metaphor benchmarks (LCC, TroFi, VUA POS, and VUA Verbs) spanning English, Spanish, Russian, and Farsi. The methodology also evaluated out-of-distribution generalization by training classifiers on one dataset or language and testing them on another.
The analysis yielded several key findings regarding model behavior. First, language models reliably encode metaphorical knowledge, significantly outperforming random baselines; on English benchmark tasks, top-performing models achieved probing accuracies between 68% and 89%, with RoBERTa and ELECTRA consistently outperforming BERT. Second, layer-by-layer evaluation revealed that metaphorical information is concentrated primarily within the middle layers (layers 3 through 6), because earlier layers capture literal source domains while deeper layers become saturated with contextual target domain information. Third, metaphorical representations generalize successfully across languages under consistent annotation schemes; zero-shot cross-lingual transfer using XLM-R achieved accuracies between 75% and 84%, with Russian serving as the most effective training source. Finally, cross-dataset transfer suffered substantial performance drops exceeding 13 percentage points due to incompatible task guidelines, part-of-speech distributions, and dataset-specific biases.
These findings indicate that foundational language models capture meaningful, language-agnostic conceptual structures rather than isolated lexical patterns, reducing the risk of complete failure when encountering figurative speech across multilingual environments. However, the severe performance drops observed across different datasets demonstrate that models are vulnerable to annotation inconsistencies and dataset artifacts, highlighting risks for practitioners who assume that strong benchmark scores ensure robust real-world transferability.
Organizations developing or deploying language technologies should establish standardized, theory-driven annotation guidelines for figurative language to improve model reliability. When fine-tuning or extracting representations for metaphor-dependent tasks, engineers should focus on the intermediate layers rather than final network layers. Further research and expanded multilingual benchmarks are recommended to explore how cultural variations influence metaphor comprehension and to test the generation of metaphorical text.
The findings are constrained by the limited set of four evaluation languages, the reliance on base model architectures (110 million parameters), and domain shifts within the source datasets. Nevertheless, the consistency of results across diverse probing techniques and languages provides high confidence in the core conclusion that pre-trained language models systematically capture metaphorical knowledge in their intermediate representations.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Its edge-probing framework and layer-wise analysis provide the methodological foundation for interpreting the source’s metaphor probes and findings about where information resides in BERT.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Its early demonstration of zero-shot cross-lingual transfer in multilingual BERT establishes the background for the source’s evaluation of metaphor transfer across languages.
- Paper: Testing the Ability of Language Models to Interpret Figurative Language, Emmy Liu et al. (2022). It advances metaphor evaluation from detecting encoded knowledge to testing whether models can interpret novel figurative language against human judgments.
- Paper: FLUTE: Figurative Language Understanding through Textual Explanations, Tuhin Chakrabarty et al. (2022). It extends metaphor assessment by requiring models to explain figurative inferences, probing whether benchmark performance reflects genuine understanding rather than superficial cues.
