Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages

Ehsan AghazadehMohsen FayyazYadollah Yaghoobzadeh

article2022ACL64 citations

Demonstrates through probing and cross-lingual experiments that pre-trained language models capture transferable metaphorical knowledge concentrated within their middle layers across multiple datasets and languages.

Listen

Metaphorical language is central to human communication and reasoning, allowing people to comprehend complex or unfamiliar concepts by linking them to concrete domains. While large pre-trained language models serve as standard foundations across language technologies, it has remained unproven whether these systems genuinely internalize structured metaphorical knowledge or merely exploit superficial statistical shortcuts. Understanding whether language models capture generalizable metaphor understanding is critical for deploying artificial intelligence systems that can accurately interpret nuanced human language.

The main objective of the article is to empirically evaluate whether pre-trained language models encode metaphorical knowledge in their internal representations and to measure how effectively this knowledge generalizes across different datasets and languages.

To conduct this evaluation, the researchers applied two diagnostic probing methods—edge probing and minimum description length probing—which measure how readily specific linguistic information can be extracted from frozen model layers. The investigation assessed widely used language models (BERT, RoBERTa, and ELECTRA) and a multilingual model (XLM-R) across four established metaphor benchmarks (LCC, TroFi, VUA POS, and VUA Verbs) spanning English, Spanish, Russian, and Farsi. The methodology also evaluated out-of-distribution generalization by training classifiers on one dataset or language and testing them on another.

The analysis yielded several key findings regarding model behavior. First, language models reliably encode metaphorical knowledge, significantly outperforming random baselines; on English benchmark tasks, top-performing models achieved probing accuracies between 68% and 89%, with RoBERTa and ELECTRA consistently outperforming BERT. Second, layer-by-layer evaluation revealed that metaphorical information is concentrated primarily within the middle layers (layers 3 through 6), because earlier layers capture literal source domains while deeper layers become saturated with contextual target domain information. Third, metaphorical representations generalize successfully across languages under consistent annotation schemes; zero-shot cross-lingual transfer using XLM-R achieved accuracies between 75% and 84%, with Russian serving as the most effective training source. Finally, cross-dataset transfer suffered substantial performance drops exceeding 13 percentage points due to incompatible task guidelines, part-of-speech distributions, and dataset-specific biases.

These findings indicate that foundational language models capture meaningful, language-agnostic conceptual structures rather than isolated lexical patterns, reducing the risk of complete failure when encountering figurative speech across multilingual environments. However, the severe performance drops observed across different datasets demonstrate that models are vulnerable to annotation inconsistencies and dataset artifacts, highlighting risks for practitioners who assume that strong benchmark scores ensure robust real-world transferability.

Organizations developing or deploying language technologies should establish standardized, theory-driven annotation guidelines for figurative language to improve model reliability. When fine-tuning or extracting representations for metaphor-dependent tasks, engineers should focus on the intermediate layers rather than final network layers. Further research and expanded multilingual benchmarks are recommended to explore how cultural variations influence metaphor comprehension and to test the generation of metaphorical text.

The findings are constrained by the limited set of four evaluation languages, the reliance on base model architectures (110 million parameters), and domain shifts within the source datasets. Nevertheless, the consistency of results across diverse probing techniques and languages provides high confidence in the core conclusion that pre-trained language models systematically capture metaphorical knowledge in their intermediate representations.

Aghazadeh et al (2022).pdf
  • Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Its edge-probing framework and layer-wise analysis provide the methodological foundation for interpreting the source’s metaphor probes and findings about where information resides in BERT.
  • Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Its early demonstration of zero-shot cross-lingual transfer in multilingual BERT establishes the background for the source’s evaluation of metaphor transfer across languages.
Cover for Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages

Abstract

Human languages are full of metaphorical expressions. Metaphors help people understand the world by connecting new concepts and domains to more familiar ones. Large pre-trained language models (PLMs) are therefore assumed to encode metaphorical knowledge useful for NLP systems. In this paper, we investigate this hypothesis for PLMs, by probing metaphoricity information in their encodings, and by measuring the cross-lingual and cross-dataset generalization of this information. We present studies in multiple metaphor detection datasets and in four languages (i.e., English, Spanish, Russian, and Farsi). Our extensive experiments suggest that contextual representations in PLMs do encode metaphorical knowledge, and mostly in their middle layers. The knowledge is transferable between languages and datasets, especially when the annotation is consistent across training and testing sets. Our findings give helpful insights for both cognitive and NLP scientists.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Inspecting Metaphorical Knowledge in PLMs
  • 3.1 Probing
  • 3.2 Generalization
  • 3.2.1 Cross-lingual
  • 3.2.2 Cross-dataset
  • 4 Experimental Setup and Results
  • 4.1 Datasets and Setup
  • 4.2 Probing Results
  • 4.3 Generalization Experiments
  • 4.3.1 Cross-lingual
  • 4.3.2 Cross-dataset
  • 4.3.3 Comparing cross-dataset and cross-lingual
  • 5 Discussion and Conclusion
  • Acknowledgements
  • References
  • A Appendices

Knowls

  1. Knowl 1 — Pre-trained representations expose metaphor-related information

    empirical result

    Frozen representations from BERT, RoBERTa, and ELECTRA support metaphor-versus-literal classification above a randomly initialized BERT baseline across four English datasets. RoBERTa and ELECTRA generally yield higher probe accuracy and MDL compression than BERT, although the differences are small on TroFi. The table reports edge-probe accuracy in percent and the best MDL compression across layers; the baseline is a randomly initialized BERT, and edge-probe accuracy is averaged over three runs.

    Baseline BERT RoBERTa ELECTRA
    Dataset Acc. Comp. Acc. Comp. Acc. Comp. Acc. Comp.
    LCC (en) 74.86 1.052 88.25 1.856 88.06 1.965 89.30 2.055
    TroFi 67.34 1.014 68.58 1.074 68.46 1.096 68.07 1.083
    VUA POS 65.92 1.030 80.32 1.435 81.72 1.486 83.03 1.514
    VUA Verbs 65.97 1.049 78.29 1.289 78.88 1.345 79.96 1.314

    The strongest scores occur on LCC (en) and the VUA datasets, while TroFi is substantially harder for every model. These probe results show that metaphorical information is accessible from pretrained contextual representations; they do not establish that the models use a human-like metaphor-processing mechanism.

  2. Knowl 2 — Metaphoricity is most accessible in middle transformer layers

    empirical result

    Layer-wise MDL probing of 12-layer BERT, RoBERTa, and ELECTRA finds that metaphor-detection compression generally rises over approximately the first 3–6 layers and falls in higher layers. Thus, the most extractable metaphor-related information is typically in middle layers, rather than at the input or deepest layers. RoBERTa on TroFi and VUA Verbs is an exception, with increases in the last layers.

    A related probe of LCC source- and target-domain labels finds that source-domain information is represented most strongly in early layers, around layers 2–6, and generally weakens in higher layers. Target-domain information generally increases across layers. The authors interpret the middle-layer advantage for metaphor detection as consistent with having both basic/source and contextual/target information available for distinguishing their contrast.

  3. Knowl 3 — Metaphorical information transfers across the four LCC languages

    empirical result

    An edge-probe classifier trained on one language's LCC examples transfers to the other LCC languages when using XLM-R representations. Each training set contains 12,238 examples, and the classifier is trained for five epochs. Each cell gives accuracy in percent for the pretrained XLM-R model, followed in parentheses by accuracy for a randomly initialized XLM-R with the same architecture. Rows are test languages and columns are training languages.

    Test language Train en Train es Train fa Train ru
    en 85.14 (65.37) 79.31 (52.71) 77.59 (50.22) 80.51 (52.40)
    es 79.40 (53.17) 84.59 (66.09) 76.70 (50.32) 79.68 (53.32)
    fa 75.70 (50.07) 75.29 (52.65) 81.04 (65.91) 77.14 (50.36)
    ru 83.92 (53.25) 80.54 (51.48) 76.61 (51.05) 88.36 (67.98)

    The pretrained model exceeds its random counterpart in every train–test pairing, including all cross-language pairings. The authors take this as evidence that information useful for metaphor detection transfers across these languages. In-distribution scores are highest for each test language, but the out-of-distribution scores remain well above the corresponding random-model scores.

  4. Knowl 4 — Cross-dataset transfer depends strongly on dataset similarity

    empirical result

    With BERT representations, a classifier trained on one English metaphor dataset generally transfers to the other datasets better than a randomly initialized BERT classifier, but cross-dataset performance is usually well below in-dataset performance. Training sets are equalized to 3,838 examples. Each cell gives accuracy in percent for pretrained BERT followed by the randomly initialized model's accuracy in parentheses; rows are test datasets and columns are training datasets.

    Test dataset Train LCC (en) Train TroFi Train VUA POS Train VUA Verbs
    LCC (en) 84.26 (54.93) 62.04 (50.05) 70.35 (50.69) 70.37 (50.14)
    TroFi 59.49 (50.58) 68.73 (64.96) 55.38 (49.45) 59.67 (53.68)
    VUA POS 62.23 (51.47) 55.29 (50.47) 76.86 (56.01) 71.6 (53.47)
    VUA Verbs 60.20 (50.88) 54.55 (51.73) 72.6 (56.01) 75.21 (60.03)

    The VUA POS and VUA Verbs datasets transfer relatively well to one another, consistent with their shared source corpus and similar distributions apart from part-of-speech coverage. VUA Verbs is the best source for TroFi, which is also verb-only. For transfers beyond the VUA pair, the gap from in-dataset performance is large; the authors report a gap greater than 13 percentage points for VUA-to-LCC comparisons. They suggest that differing annotation procedures and dataset biases help explain the weaker transfer.

  5. Knowl 5 — Matched comparison shows stronger cross-language than cross-dataset transfer

    empirical result

    A controlled comparison uses XLM-R, the same training size of 3,838 examples, and the same LCC (en) test set for both transfer settings. Training on each LCC language gives cross-language accuracy; training on each other English dataset gives cross-dataset accuracy. The in-distribution LCC (en)-to-LCC (en) result is included for reference.

    Training source Accuracy on LCC (en) test set (%)
    LCC (en) 82.31
    LCC (es) 78.02
    LCC (fa) 77.3
    LCC (ru) 78.04
    TroFi 60.54
    VUA POS 68.61
    VUA Verbs 67.15

    With the model, training size, and test set held fixed, all three cross-language sources outperform all three cross-dataset sources. The authors relate this contrast to the more consistent annotation guidelines among the LCC language versions, compared with differences across datasets in annotation, covered parts of speech, and sentence lengths.

  6. Knowl 6 — Frozen-representation probing isolates information available to a classifier

    model/method

    The study probes pretrained language models without updating their parameters: contextual representations are frozen, and a supervised classifier is trained on top. In edge probing, the input is the representation of the annotated target span, projected to 256 dimensions. Pooling produces a fixed-size span representation and combines information across model layers; the probe classifier then predicts the metaphor label. Restricting access to the marked span makes the experiment assess whether the frozen representations of that span encode information useful for its label, rather than optimizing the language model for the task.

    The edge-probe experiments use BERT, RoBERTa, and ELECTRA base models, each with 12 layers, a hidden size of 768, and 110 million parameters. The probe uses batch size 32, learning rate 5×10−55\times10^{-5}, and five training epochs. The authors use edge probing for overall task comparisons and MDL probing for layer-wise analysis.

  7. Knowl 7 — MDL compression measures probe quality relative to coding cost

    equation

    The study uses online minimum-description-length (MDL) probing to compare how extractable metaphoricity information is from different representations, including representations from individual layers. In online coding, a classifier is trained successively on portions of the data, and the total code length, MDL\mathrm{MDL}, is the sum of the cross-entropies for those portions. Cross-entropies are computed with base-2 logarithms, so code length is measured in bits.

    The reported compression is

    compression=Nlog⁡2(K)MDL,\mathrm{compression}=\frac{N\log_2(K)}{\mathrm{MDL}},

    where NN is the number of examples in the dataset, KK is the number of label classes, and MDL\mathrm{MDL} is the online code length in bits. A random classifier has compression 1; a larger value indicates a more effective and extractable probe. The MDL probe uses the edge-probing classifier structure.

  8. Knowl 8 — Datasets and evaluation protocol balance literal and metaphor labels

    experimental setup

    The experiments use four metaphor-detection datasets: LCC, TroFi, VUA POS, and VUA Verbs. LCC is available in English, Farsi, Spanish, and Russian; the other three datasets are English-only. TroFi and VUA Verbs contain verbs, whereas LCC and VUA POS include nouns, verbs, adjectives, and adverbs. The LCC labels use score 0 as literal and the other retained scores as metaphor; unclear cases in the range 0.5≤score<1.50.5\leq\mathrm{score}<1.5 are excluded. All datasets are label-balanced to 50% metaphor and split into training, development, and test sets in a 0.7/0.1/0.2 ratio.

    The following entries give the reported numbers of training, development, and test instances, respectively.

    Dataset Parts of speech Train / development / test instances
    LCC (en) All 28,096 / 4,014 / 8,028
    LCC (fa) All 12,238 / 1,802 / 3,604
    LCC (es) All 12,238 / 2,236 / 4,474
    LCC (ru) All 12,238 / 1,748 / 3,498
    TroFi Verbs 3,838 / 548 / 1,096
    VUA Verbs Verbs 9,176 / 1,310 / 2,622
    VUA POS All 21,036 / 3,006 / 6,010

    For cross-language experiments, the LCC training sets are subsampled to 12,238 examples each; for cross-dataset experiments, all training sources are subsampled to 3,838. XLM-R is used for cross-language transfer, BERT for cross-dataset transfer, and XLM-R for the matched comparison between transfer settings. Each transfer experiment is also run with a randomly initialized model of the same architecture as a control.

  9. Knowl 9 — Russian LCC is the strongest cross-language training source

    empirical result

    Among the four LCC languages, training on Russian gives the best out-of-distribution accuracy for each of the other three test languages: 80.51% on English, 79.68% on Spanish, and 77.14% on Farsi. The authors offer possible explanations rather than a causal determination. Russian LCC has relatively similar target-domain frequencies to the other language datasets; Russian is also the second-largest language in XLM-R pretraining data after English. Conversely, English LCC contains many examples about guns and gun control, target domains not covered in the other LCC datasets, which could reduce transfer from English. These factors are proposed as influences on the observed scores, not experimentally isolated causes.

  10. Knowl 10 — Metaphor detection is framed as identifying a contextual contrast

    definition

    The study treats metaphor detection as classification of a target word or span in context as metaphorical or literal. Following the metaphor identification procedure used by the datasets, an expression is metaphorical when its basic or literal meaning contrasts with its contextual meaning. In conceptual-metaphor terms, a more basic source domain is used to express a contextual target domain—for example, “won” in “We won the argument” uses a WAR-related meaning in an ARGUMENT context. The same word in “The Allies won the war” is literal because its basic and contextual meanings do not make that contrast. Identifying the exact source and target domains is not required for the detection task.

Coverage note — The appendix's full per-domain frequency plots and Jensen–Shannon divergence matrices are not reproduced; their role in interpreting transfer is summarized in the Russian-source result, and the remaining domain-frequency detail is supporting analysis rather than a separate central finding.

References

  1. 1.Yonatan Belinkov. 2022. Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics, pages 1–13.
  2. 2.Julia Birke and Anoop Sarkar. 2006. A clustering approach for nearly unsupervised recognition of non-literal language. In 11th Conference of the European Chapter of the Association for Computational Linguistics, Trento, Italy. Association for Computational Linguistics.
  3. 3.Julia Birke and Anoop Sarkar. 2007. Active learning for the identification of nonliteral language. In Proceedings of the Workshop on Computational Approaches to Figurative Language, pages 21–28, Rochester, New York. Association for Computational Linguistics.
  4. 4.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kudtipudi, and et al. 2021. On the opportunities and risks of foundation models. CoRR, abs/2108.07258.
  5. 5.Xianyang Chen, Chee Wee (Ben) Leong, Michael Flor, and Beata Beigman Klebanov. 2020. Go figure! multi-task transformer-based architecture for metaphor detection using idioms: ETS team in 2020 metaphor shared task. In Proceedings of the Second Workshop on Figurative Language Processing, pages 235–243, Online. Association for Computational Linguistics.
  6. 6.Rochelle Choenni and Ekaterina Shutova. 2020. What does it mean to be language-agnostic? probing multilingual sentence encoders for typological properties. CoRR, abs/2009.12862.
  7. 7.Minjin Choi, Sunkyung Lee, Eunseong Choi, Heesoo Park, Junhyuk Lee, Dongwon Lee, and Jongwuk Lee. 2021. Melbert: Metaphor detection via contextualized late interaction using metaphorical identification theories. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 1763–1773. Association for Computational Linguistics.
  8. 8.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  9. 9.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzman, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  10. 10.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136. Association for Computational Linguistics.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.John C. Duchi and Hongseok Namkoong. 2018. Learning models with uniform performance via distributionally robust optimization. CoRR, abs/1810.08750.
  13. 13.Max Eichler, Gozde Gül Şahin, and Iryna Gurevych. 2019. LINSPECTOR WEB: A multilingual probing suite for word representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations, pages 127–132, Hong Kong, China. Association for Computational Linguistics.
  14. 14.Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2021. Not all models localize linguistic knowledge in the same place: A layer-wise probing on BERToids’ representations. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 375–388, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Hongyu Gong, Kshitij Gupta, Akriti Jain, and Suma Bhat. 2020. IlliniMet: Illinois system for metaphor detection with contextual and linguistic information. In Proceedings of the Second Workshop on Figurative Language Processing, pages 146–153, Online. Association for Computational Linguistics.
  16. 16.Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Pado. 2015. Distributional vectors encode referential attributes. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 12–21, Lisbon, Portugal. Association for Computational Linguistics.
  17. 17.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. 2021. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 8340–8349.
  18. 18.Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020. Pretrained transformers improve out-of-distribution robustness. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2744–2751, Online. Association for Computational Linguistics.
  19. 19.John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Arne Kohn. 2015. What’s in an embedding? analyzing word embeddings through multilingual evaluation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2067–2073, Lisbon, Portugal. Association for Computational Linguistics.
  21. 21.George Lakoff and Mark Johnson. 2008. Metaphors we live by. University of Chicago press.
  22. 22.Junyi Li, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen. 2021. Pretrained language model for text generation: A survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 4492–4499. International Joint Conferences on Artificial Intelligence Organization. Survey Track.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  24. 24.Rui Mao, Chenghua Lin, and Frank Guerin. 2019. End-to-end sequential metaphor identification inspired by linguistic theories. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3888–3898, Florence, Italy. Association for Computational Linguistics.
  25. 25.Zachary J. Mason. 2004. CorMet: A computational, corpus-based conventional metaphor extraction system. Computational Linguistics, 30(1):23–44.
  26. 26.Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao. 2020. Deep learning based text classification: A comprehensive review. CoRR, abs/2004.03705.
  27. 27.Michael Mohler, Mary Brunson, Bryan Rink, and Marc T. Tomlinson. 2016. Introducing the LCC metaphor datasets. In Proceedings of the Tenth International Conference on Language Resources and Evaluation LREC 2016, Portoroz, Slovenia, May 23-28, 2016. European Language Resources Association (ELRA).
  28. 28.Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, Vinay Adiga, and Erik Cambria. 2021. Recent advances in deep learning based dialogue systems: A systematic survey. CoRR, abs/2105.04387.
  29. 29.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  30. 30.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  31. 31.Vinit Ravishankar, Memduh Gokırmak, Lilja Øvrelid, and Erik Velldal. 2019a. Multilingual probing of deep pre-trained contextual encoders. In Proceedings of the First NLPL Workshop on Deep Learning for Natural Language Processing, pages 37–47, Turku, Finland. Linkoping University Electronic Press.
  32. 32.Vinit Ravishankar, Lilja Øvrelid, and Erik Velldal. 2019b. Probing multilingual sentence representations with X-probe. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 156–168, Florence, Italy. Association for Computational Linguistics.
  33. 33.Ekaterina Shutova, Douwe Kiela, and Jean Maillard. 2016. Black holes and white rabbits: Metaphor identification with visual features. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 160–170, San Diego, California. Association for Computational Linguistics.
  34. 34.Ekaterina Shutova, Simone Teufel, and Anna Korhonen. 2013. Statistical metaphor processing. Computational Linguistics, 39(2):301–353.
  35. 35.Wei Song, Shuhui Zhou, Ruiji Fu, Ting Liu, and Lizhen Liu. 2021. Verb metaphor detection via contextual relation learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 4240–4251. Association for Computational Linguistics.
  36. 36.Gerard Steen. 2010. A method for linguistic metaphor identification: From MIP to MIPVU, volume 14. John Benjamins Publishing.
  37. 37.Chuandong Su, Fumiyo Fukumoto, Xiaoxi Huang, Jiyi Li, Rongbo Wang, and Zhiqun Chen. 2020. DeepMet: A reading comprehension paradigm for token-level metaphor detection. In Proceedings of the Second Workshop on Figurative Language Processing, pages 30–39, Online. Association for Computational Linguistics.
  38. 38.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019a. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy. Association for Computational Linguistics.
  39. 39.Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019b. What do you learn from context? Probing for sentence structure in contextualized word representations. In International Conference on Learning Representations.
  40. 40.Yulia Tsvetkov, Leonid Boytsov, Anatole Gershman, Eric Nyberg, and Chris Dyer. 2014. Metaphor detection with cross-lingual model transfer. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, June 22-27, 2014, Baltimore, MD, USA, Volume 1: Long Papers, pages 248–258. The Association for Computer Linguistics.
  41. 41.Peter Turney, Yair Neuman, Dan Assaf, and Yohai Cohen. 2011. Literal and metaphorical sense identification through concrete and abstract context. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, pages 680–690, Edinburgh, Scotland, UK. Association for Computational Linguistics.
  42. 42.Elena Voita and Ivan Titov. 2020. Information-theoretic probing with minimum description length. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 183–196, Online. Association for Computational Linguistics.
  43. 43.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  44. 44.Chuhan Wu, Fangzhao Wu, Yubo Chen, Sixing Wu, Zhigang Yuan, and Yongfeng Huang. 2018. Neural metaphor detecting with CNN-LSTM model. In Proceedings of the Workshop on Figurative Language Processing, pages 110–114, New Orleans, Louisiana. Association for Computational Linguistics.
  45. 45.Yadollah Yaghoobzadeh, Katharina Kann, T. J. Hazen, Eneko Agirre, and Hinrich Schutze. 2019. Probing for semantic classes: Diagnosing the meaning content of word embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5740–5753, Florence, Italy. Association for Computational Linguistics.
  46. 46.Yadollah Yaghoobzadeh and Hinrich Schutze. 2016. Intrinsic subspace evaluation of word embedding representations. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 236–246, Berlin, Germany. Association for Computational Linguistics.
  47. 47.Zhuosheng Zhang, Hai Zhao, and Rui Wang. 2020. Machine reading comprehension: The role of contextualized language models and beyond. CoRR, abs/2005.06249.
  48. 48.Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh, and Hinrich Schutze. 2020. Quantifying the contextualization of word representations with semantic class probing. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1219–1234, Online. Association for Computational Linguistics.

Citation

MLA
Aghazadeh, E., et al. “Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2037–50, https://doi.org/10.18653/v1/2022.acl-long.144.
APA
Aghazadeh, E., Fayyaz, M., & Yaghoobzadeh, Y. (2022). Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2037–2050. https://doi.org/10.18653/v1/2022.acl-long.144
Chicago
Aghazadeh, E., M. Fayyaz, and Y. Yaghoobzadeh. 2022. “Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2037–50. https://doi.org/10.18653/v1/2022.acl-long.144.
Harvard
Aghazadeh, E., Fayyaz, M. and Yaghoobzadeh, Y. (2022) “Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2037–2050. Available at: https://doi.org/10.18653/v1/2022.acl-long.144.
Vancouver
1. Aghazadeh E, Fayyaz M, Yaghoobzadeh Y (2022) Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2037–2050

BibTeX

@inproceedings{aghazadeh-etal-2022-metaphors,
    title = "Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and Languages",
    author = "Aghazadeh, Ehsan  and
      Fayyaz, Mohsen  and
      Yaghoobzadeh, Yadollah",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.144/",
    doi = "10.18653/v1/2022.acl-long.144",
    pages = "2037--2050"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/