Testing the Ability of Language Models to Interpret Figurative Language

Emmy LiuChenxuan CuiKenneth ZhengGraham Neubig

article2022NAACL118 citations
Listen

Figurative language, such as metaphors and similes, appears frequently in human communication to convey complex ideas, emotions, and humor. Because nonliteral expressions rely heavily on shared commonsense, social, and cultural knowledge rather than strict word definitions, interpreting them presents a major challenge for artificial intelligence. Most prior natural language processing research focuses on detecting common, conventional metaphors rather than interpreting creative, novel figurative phrasing. As language models are increasingly deployed in real-world conversational and text-processing applications, understanding their ability to infer nonliteral meaning is critical.

The main objective of the article is to evaluate how well state-of-the-art language models can understand and interpret novel, human-generated figurative language compared to human capability. The study specifically investigates whether pre-trained models can select the correct literal interpretation of a creative metaphor and generate valid explanations on their own.

To evaluate this, the authors introduced Fig-QA, a Winograd-style benchmark containing 10,256 validated sentence pairs with opposing metaphorical meanings and corresponding literal implications. The dataset was crowdsourced to specifically elicit creative metaphors that rarely appear on the internet or in training data, categorized across physical commonsense, visual, social, and cultural domains. The researchers evaluated several prominent auto-regressive models (GPT-2, GPT-neo, and multiple GPT-3 variants) and masked language models (BERT and RoBERTa) across zero-shot, few-shot prompting, and fine-tuned configurations. In addition, human baseline testing was conducted with native and non-native English speakers, alongside a manual qualitative evaluation of open-ended interpretations generated by GPT-3.

The study revealed several critical findings regarding model performance and behavior. First, while untrained models perform above random chance, a substantial capability gap exists between models and humans in zero-shot settings: the highest-performing zero-shot model (GPT-3 Davinci) achieved 68.4% accuracy, trailing the human baseline of 94.4% by roughly 26 percentage points. Second, fine-tuning substantially improves performance across all architectures, with fine-tuned RoBERTa reaching 90.3% accuracy, coming within about 4 percentage points of human performance. Third, auto-regressive models exhibit severe contextual blindness, relying heavily on the standalone probability of the answer options rather than evaluating the metaphorical context; this caused their accuracy to collapse (falling below 20% for zero-shot GPT models) under strict paired evaluation where both opposing sentences in a pair must be answered correctly. Fourth, in open-ended generation, GPT-3 produced accurate interpretations only about 51% to 64% of the time and generated self-contradictory statements across multiple attempts in roughly 38% of cases. Finally, models struggle disproportionately with sarcastic metaphors and cultural commonsense, frequently making basic comprehension errors that diverge sharply from human error patterns.

These findings indicate that while fine-tuned models can perform well on constrained multiple-choice selections, current language models lack genuine, robust nonliteral reasoning. In operational environments, relying on zero-shot language models to parse figurative, sarcastic, or culturally embedded communication introduces notable safety, performance, and compliance risks, such as misinterpreting customer intent, misunderstanding social nuances, or generating contradictory summaries. The fact that models lean on surface-level word probabilities rather than true semantic comprehension means standard evaluation benchmarks may overstate their practical language understanding.

Organizations deploying natural language systems should exercise caution when processing unstructured human discourse containing figurative expressions. Decision-makers should avoid relying on zero-shot auto-regressive models for sensitive semantic interpretation tasks and should instead implement contrastive fine-tuning approaches or supervised classifiers where feasible. For future research and development, the authors recommend expanding benchmark evaluations into multimodal contexts and non-English languages, as well as developing training methodologies that explicitly ground models in analogical and commonsense reasoning.

Confidence in these findings is high for standard English text, supported by rigorous manual validation, multiple baseline architectures, and paired Winograd-style controls. However, limitations include the benchmark’s geographic focus on United States cultural references, the moderate inter-rater agreement inherent to subjective generation tasks, and the restriction to text-only English data. Readers should remain cautious about generalizing these conclusions directly to multilingual settings or multimodal domains without further empirical validation.

Table of Contents

  • 1 Introduction
  • 2 Dataset Creation and Validation
  • 2.1 Crowdsourcing Task
  • 2.2 Data Validation
  • 2.3 Final Dataset
  • 3 Figurative Language Typologies
  • 3.1 Figurative Language Structure
  • 3.2 Common-sense Knowledge Types
  • 4 Baseline Models and Evaluation
  • 4.1 Auto-regressive Language Models
  • 4.2 Masked Language Models
  • 4.3 Forced-choice Paradigm
  • 4.4 Human Performance
  • 5 Results
  • 5.1 Inference Results
  • 5.2 Generation Results
  • 6 Performance and Error Analysis
  • 6.1 Reliance on Probability of Answers
  • 6.2 Other Factors Influencing Correctness
  • 6.3 Qualitative Analysis of Error Trends
  • 7 Related work
  • 7.1 Figurative Language Identification
  • 7.2 Figurative Language Interpretation
  • 7.3 Other Figurative Language Datasets
  • 7.4 Human Language Processing
  • 8 Conclusion
  • 9 Ethical Considerations
  • 9.1 Potential Risks
  • 9.2 Terms of Use of Artefacts Used
  • 9.3 Computational Infrastructure and Computing Budget
  • Acknowledgements
  • References
  • A Crowdsourcing Details
  • B Invalid Examples
  • C Backward accuracies
  • D Paired accuracies
  • E Accuracy breakdown by Part-of-Speech
  • E.1 Subject
  • E.2 Relation
  • E.3 Object
  • F Accuracy breakdown by hypernyms
  • F.1 Subject
  • F.2 Object
  • G Generation examples

Knowls

  1. Knowl 1 — Fig-QA tests interpretation of paired creative metaphors

    model/method

    Fig-QA is a binary-choice task for testing whether a model can infer the intended meaning of a figurative expression. Each item pairs two figurative phrases with divergent meanings and two corresponding literal interpretations; the system must match each phrase to its intended interpretation. For example, a metaphor comparing commitment to plywood or oak is paired with contrasting interpretations about being committed or uncommitted. The paired alternatives constrain the interpretation to a particular semantic contrast while allowing the metaphor itself to be creative.

    The authors collected examples from US-based Amazon Mechanical Turk workers instructed to write uncommon but understandable metaphors and literal interpretations. Three authors manually filtered examples for unintelligibility, format violations, and grammar or spelling errors that prevented understanding: 10,256 of the 13,324 collected examples remained. The release has a small training sample of 200 examples, a medium training split of 1,458, a large training split of 8,016, a development split of 1,094, and a test split of 1,146. The benchmark is English-only and text-based.

  2. Knowl 2 — Scoring forward and backward metaphor interpretation

    model/method

    For an item with figurative phrases x1,x2x_1,x_2 and their intended interpretations y1,y2y_1,y_2, the authors evaluate two directions using scores from autoregressive language models. In the forward direction, the model chooses which interpretation fits a given phrase; for the correct phrase–interpretation pair ii and the competing interpretation yjy_j, the normalized choice score is P(yi∣xi)=P(xi,yi)/(P(xi,yi)+P(xi,yj))P(y_i\mid x_i)=P(x_i,y_i)/(P(x_i,y_i)+P(x_i,y_j)). In the backward direction, the model chooses which phrase matches an interpretation: P(xi∣yi)=P(xi,yi)/(P(xi,yi)+P(xj,yi))P(x_i\mid y_i)=P(x_i,y_i)/(P(x_i,y_i)+P(x_j,y_i)). Here, P(x,y)P(x,y) denotes the model’s length-normalized score for the concatenated figurative phrase and interpretation; it is a heuristic rather than a strict sequence probability. A prediction is correct when the score of the intended option exceeds 0.50.5.

    The authors also evaluate masked language models, which do not directly provide sentence probabilities for this zero-shot scoring method. They transfer BERT and RoBERTa from Winogrande or NLI training, and fine-tune masked models contrastively on Fig-QA by presenting both answer choices. Autoregressive models are fine-tuned with language-modeling loss.

  3. Knowl 3 — Fine-tuning raises accuracy, but zero-shot models trail people

    data/table

    Test accuracy (%) shows that every evaluated model family performs above chance, and fine-tuning improves results substantially. GPT-3 Davinci is the strongest listed zero-shot model at 68.41%, while fine-tuned RoBERTa reaches 90.32%; human accuracy is 94.42%. Thus, the strongest listed zero-shot model remains well below human performance, whereas fine-tuned RoBERTa comes within 4.10 percentage points of the human score. Fine-tuned values are reported in the paper’s Tuned (L) and Tuned (XL) columns and are averaged across five seeds.

    Results by model, in the order zero-shot / Tuned (L) / Tuned (XL), are: GPT-2, 53.93 / 54.80 / 62.65; GPT-Neo 1.3B, 56.89 / 69.98 / 72.00; GPT-3 Ada, 59.08 / 69.17 / 73.56; GPT-3 Babbage, 62.91 / 73.97 / 77.31; GPT-3 Curie, 65.35 / 79.04 / 81.94; GPT-3 Davinci, 68.41 / not reported / not reported; BERT, 58.14 / 83.16 / 85.69; and RoBERTa, 66.18 / 89.22 / 90.32. Human accuracy is 94.42%, rising to 95.39% when counting only examples people marked as confident. The BERT and RoBERTa zero-shot figures are transfer results from Winogrande rather than direct language-model likelihood scores. For RoBERTa transferred from NLI, accuracy was 50.47% with the original three labels and 66.32% with a forced binary decision.

  4. Knowl 4 — Fig-QA test examples draw on four overlapping knowledge types

    definition

    The authors categorized the common-sense knowledge needed to interpret test-set metaphors into four non-mutually-exclusive types. Common-sense object knowledge concerns properties of ordinary objects, animals, and materials, such as size, mass, or volume; it was annotated for 68.35% of the test set. Visual metaphors rely especially on visual properties such as brightness or color, or evoke a visual scene; they comprised 14.73% and are described as a subset of object-knowledge metaphors. Social understanding requires knowledge of human reactions or emotions and appeared in 27.55%. Cultural knowledge invokes traditions, religion, or works of art and artifacts; it appeared in 16.56%. The cultural references reflect the US-based crowdworker pool. Because categories can overlap, their percentages are not intended to sum to 100%.

  5. Knowl 5 — Model accuracy varies across common-sense categories

    data/table

    The following test accuracies (%) compare models on examples annotated for each knowledge type; categories can overlap. The results show that trained models generally perform best on object and visual knowledge, while cultural examples remain a relative weakness for several trained systems. RoBERTa is strongest among the listed models in each category, though it remains below human accuracy in all four.

    For each model, values are ordered object / visual / social / cultural. Untrained GPT-2: 52.17 / 52.07 / 55.38 / 58.42; untrained GPT-Neo: 56.38 / 55.62 / 56.01 / 62.10; untrained GPT-3 Curie: 75.00 / 71.00 / 72.47 / 78.42. Trained GPT-2: 53.57 / 51.48 / 57.91 / 57.37; trained GPT-Neo: 70.15 / 72.78 / 68.67 / 70.00; trained GPT-3 Curie: 87.50 / 84.62 / 83.86 / 83.16; BERT: 87.37 / 92.31 / 84.18 / 77.37; RoBERTa: 91.20 / 94.08 / 89.56 / 83.68. Human accuracy: 95.41 / 96.45 / 93.99 / 90.00. The authors report that training gains are concentrated in object, visual, and social categories, with little improvement on cultural examples.

  6. Knowl 6 — Models rely strongly on answer plausibility apart from context

    empirical result

    Across the Fig-QA examples, the probability assigned to an interpretation by itself is strongly associated with the model’s probability of that interpretation given its paired metaphor. Spearman correlations between P(yi∣xi)P(y_i\mid x_i) and P(yi)P(y_i) were positive and statistically significant for every autoregressive model tested. For untrained models, GPT-2 had r=0.8128r=0.8128 (p=6.700×10−136p=6.700\times10^{-136}), GPT-Neo had r=0.7891r=0.7891 (p=6.075×10−123p=6.075\times10^{-123}), and GPT-3 had r=0.7392r=0.7392 (p=4.329×10−100p=4.329\times10^{-100}). After training, the correlations were GPT-2 r=0.6765r=0.6765 (p=6.700×10−78p=6.700\times10^{-78}), GPT-Neo r=0.6689r=0.6689 (p=1.456×10−75p=1.456\times10^{-75}), and GPT-3 r=0.4157r=0.4157 (p=2.598×10−25p=2.598\times10^{-25}).

    Here P(yi)P(y_i) is the model’s probability for interpretation yiy_i considered without the figurative context, and P(yi∣xi)P(y_i\mid x_i) is the forward choice score. The correlations weaken with training, especially for GPT-3, but remain positive; the authors interpret this as evidence that models often lean on which answer is already more probable rather than fully using the metaphorical context. This tendency also leads models to give the same answer to both members of a paired example.

  7. Knowl 7 — Matching an interpretation back to a metaphor is harder than the forward task

    empirical result

    Autoregressive models were less accurate when asked to select a figurative phrase given its interpretation than when asked to select the interpretation given the phrase. Zero-shot and fine-tuned backward test accuracies (%) were GPT-2, 52.18 and 52.00; GPT-Neo 1.3B, 54.36 and 63.44; and GPT-3 Curie, 58.46 and 74.83. For GPT-3 Curie, the corresponding forward accuracies were 65.35% zero-shot and 79.04% fine-tuned, making backward accuracy lower by 6.89 and 4.21 percentage points, respectively. The result indicates an asymmetry between interpreting a given metaphor and identifying which metaphor expresses a given literal meaning.

  8. Knowl 8 — Scoring both members of a pair sharply lowers autoregressive accuracy

    empirical result

    The authors also scored a paired example as correct only when the model answered both figurative phrases in the pair correctly. On the test set, paired accuracy (%) was GPT-2 6.63 zero-shot and 5.06 fine-tuned; GPT-Neo 10.3 zero-shot and 10.3 fine-tuned; GPT-3 Curie 17.4 zero-shot and 50.0 fine-tuned; BERT 70.6 fine-tuned; RoBERTa 80.4 fine-tuned; and humans 89.7. These figures are from one run. Requiring both answers exposes failures that ordinary per-example accuracy can hide, particularly for autoregressive models; the human score is least affected.

  9. Knowl 9 — Free-form GPT-3 interpretations are often inconsistent across samples

    experimental setup

    The authors manually evaluated four completions from GPT-3 Davinci for each metaphor in one tenth of the test set. The model received the metaphor followed by the suffix That is to say, and generated up to 100 tokens per completion. Only the first sentence was judged; generation temperature was set to 0.4 after a small-set search over 0.2, 0.4, 0.6, 0.8, and 1. Three authors labeled outputs as correct, incorrect, or literal, accepting valid interpretations beyond the crowdworker-provided answer. Ambiguous examples were excluded, and inter-rater reliability was moderate (α=0.5567\alpha=0.5567).

    Counting literal restatements as incorrect, GPT-3 Davinci’s accuracy was 50.8%; excluding literal outputs from the evaluation yielded 63.9%. Across the four completions, 37.7% of metaphors received contradictory completions, at least one completion was correct for 78.1% of metaphors, and every completion was correct for only 19.3%. Thus, occasional valid generations did not imply reliable interpretation across samples.

  10. Knowl 10 — Error patterns include sarcasm and diverge from human errors

    empirical result

    Longer figurative context phrases were associated with lower model correctness: the point-biserial correlation between phrase length and binary correctness was −0.1544-0.1544 (p=1.50×10−7p=1.50\times10^{-7}). Qualitative analysis also identified sarcastic metaphors as difficult for both people and models. For example, interpreting a person as bubbly as still water requires recognizing a sarcastic contrast and inferring blandness, rather than relying on the usual association between bubbly and vivacious.

    The mistakes of trained GPT-3 Curie overlapped only partly with human mistakes: 13 of the 64 human errors were also GPT-3 errors. The authors describe GPT-3’s errors as often inexplicable even for simple examples, whereas human errors more often involved rare vocabulary or unfamiliar cultural references. This suggests that the model and human error profiles were not well aligned.

Coverage note — The secondary prompting comparison (the suffix prompt yielded about a 1–2% gain and random in-context examples were generally ineffective) and detailed POS/WordNet breakdowns are omitted because they are exploratory analyses rather than load-bearing findings; the paper’s English-only, US-centered, text-only scope is noted in the benchmark knowl.

References

  1. 1.Ehsan Aghazadeh, Mohsen Fayyaz, and Yadollah Yaghoobzadeh. 2022. Metaphors in pre-trained language models: Probing and generalization across datasets and languages.
  2. 2.Beata Beigman Klebanov, Chee Wee (Ben) Leong, and Michael Flor. 2018. A corpus of non-native written English annotated for metaphor. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 86–91, New Orleans, Louisiana. Association for Computational Linguistics.
  3. 3.Emily M. Bender and Alexander Koller. 2020. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online. Association for Computational Linguistics.
  4. 4.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8718–8735, Online. Association for Computational Linguistics.
  5. 5.Yuri Bizzoni and Mehdi Ghanimifard. 2018. Bigrams and BiLSTMs two neural networks for sequential metaphor detection. In Proceedings of the Workshop on Figurative Language Processing, pages 91–101, New Orleans, Louisiana. Association for Computational Linguistics.
  6. 6.Yuri Bizzoni and Shalom Lappin. 2018. Predicting human metaphor paraphrase judgments with deep neural networks. In Proceedings of the Workshop on Figurative Language Processing, pages 45–55, New Orleans, Louisiana. Association for Computational Linguistics.
  7. 7.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  9. 9.Robyn Carston and Catherine Wearing. 2011. Metaphor, hyperbole and simile: A pragmatic approach. Language and Cognition, 3(2):283–312.
  10. 10.Tuhin Chakrabarty, Yejin Choi, and Vered Shwartz. 2021a. It’s not rocket science : Interpreting figurative language in narratives. ArXiv, abs/2109.00087.
  11. 11.Tuhin Chakrabarty, Debanjan Ghosh, Adam Poliak, and Smaranda Muresan. 2021b. Figurative language in recognizing textual entailment. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3354–3361, Online. Association for Computational Linguistics.
  12. 12.Thomas C. Cooper. 1999. Processing of idioms by l2 learners of english. TESOL Quarterly, 33:233–262.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding.
  14. 14.Yanai Elazar, Hongming Zhang, Yoav Goldberg, and Dan Roth. 2021. Back to square one: Artifact detection, training and commonsense disentanglement in the winograd schema.
  15. 15.Gilles Fauconnier and Mark Turner. 2003. Conceptual blending, form and meaning.
  16. 16.Christiane Fellbaum. 1998. WordNet: An Electronic Lexical Database. Bradford Books.
  17. 17.Susan Fussell and Mallie Moss. 2008. Figurative language in emotional communication.
  18. 18.Ge Gao, Eunsol Choi, Yejin Choi, and Luke Zettlemoyer. 2018. Neural metaphor detection in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 607–613, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.D Gentnder and B Bowdle. 2008. Metaphor as structure-mapping. In The Cambridge handbook of metaphor and thought, pages 109–128. Cambridge University Press.
  20. 20.Kahlil Gibran. 1926. Sand and Foam; a book of aphorisms. A.A Knopf.
  21. 21.Sam Glucksberg. 2003. The psycholinguistics of metaphor. Trends in cognitive sciences, 7:92–96.
  22. 22.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  23. 23.Harsh Jhamtani, Varun Gangal, Eduard Hovy, and Taylor Berg-Kirkpatrick. 2021. Investigating robustness of dialog models to popular figurative language constructs. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7476–7485, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.G. Lakoff and M. Johnson. 1981. Metaphors we Live By. University of Chicago Press.
  25. 25.Chee Wee (Ben) Leong, Beata Beigman Klebanov, Chris Hamill, Egon Stemle, Rutuja Ubale, and Xianyang Chen. 2020. A report on the 2020 VUA and TOEFL metaphor detection shared task. In Proceedings of the Second Workshop on Figurative Language Processing, pages 18–29, Online. Association for Computational Linguistics.
  26. 26.Chee Wee (Ben) Leong, Beata Beigman Klebanov, and Ekaterina Shutova. 2018. A report on the 2018 VUA metaphor detection shared task. In Proceedings of the Workshop on Figurative Language Processing, pages 56–66, New Orleans, Louisiana. Association for Computational Linguistics.
  27. 27.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning, KR’12, page 552–561. AAAI Press.
  28. 28.Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee, and Xiang Ren. 2021. Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  30. 30.Edward Loper and Steven Bird. 2002. Nltk: The natural language toolkit. CoRR, cs.CL/0205028.
  31. 31.Rui Mao, Chenghua Lin, and Frank Guerin. 2018. Word embedding and WordNet based metaphor identification and interpretation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1222–1231, Melbourne, Australia. Association for Computational Linguistics.
  32. 32.J. S Mio and A. N Katz. 1996. Metaphor: Implications and Applications. Psychology Press.
  33. 33.Saif Mohammad, Ekaterina Shutova, and Peter Turney. 2016. Metaphor as a medium for emotion: An empirical study. In Proceedings of the Fifth Joint Conference on Lexical and Computational Semantics, pages 23–33, Berlin, Germany. Association for Computational Linguistics.
  34. 34.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  35. 35.Paolo Pedinotti, Eliana Di Palma, Ludovica Cerini, and Alessandro Lenci. 2021. A howling success or a working sea? testing what BERT knows about metaphors. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 192–204, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  36. 36.Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018. Collecting diverse natural language inference problems for sentence representation evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 67–81, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Malay Pramanick, Ashim Gupta, and Pabitra Mitra. 2018. An LSTM-CRF based approach to token-level metaphor detection. In Proceedings of the Workshop on Figurative Language Processing, pages 67–75, New Orleans, Louisiana. Association for Computational Linguistics.
  38. 38.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  39. 39.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. In AAAI.
  40. 40.Ekaterina Shutova. 2010. Automatic metaphor interpretation as a paraphrasing task. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 1029–1037, Los Angeles, California. Association for Computational Linguistics.
  41. 41.Ekaterina Shutova. 2011. Computational approaches to figurative language.
  42. 42.Gerard J Steen, Alleta G Dorst, Berenike Herrmann, Anna A Kaal, Tina Krennmayr, and Trynetje Pasma. 2010. A Method for Linguistic Metaphor Identification: From MIP to MIPVU. John Benjamins.
  43. 43.Kevin Stowe and Martha Palmer. 2018. Leveraging syntactic constructions for metaphor identification. In Proceedings of the Workshop on Figurative Language Processing, pages 17–26, New Orleans, Louisiana. Association for Computational Linguistics.
  44. 44.Chang Su, Shuman Huang, and Yijiang Chen. 2017. Automatic detection and interpretation of nominal metaphor based on the theory of meaning. Neurocomputing, 219:300–311.
  45. 45.John Sweller. 2006. Discussion of ’emerging topics in cognitive load research: Using learner and information characteristics in the design of powerful learning environments’. Applied Cognitive Psychology - APPL COGNITIVE PSYCHOL, 20:353–357.
  46. 46.Xiaoyu Tong, Ekaterina Shutova, and Martha Lewis. 2021. Recent advances in neural metaphor processing: A linguistic, cognitive and social perspective. In NAACL 2021.
  47. 47.Yulia Tsvetkov, Leonid Boytsov, Anatole Gershman, Eric Nyberg, and Chris Dyer. 2014. Metaphor detection with cross-lingual model transfer. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 248–258, Baltimore, Maryland. Association for Computational Linguistics.
  48. 48.P. Wolff and Dedre Gentner. 2000. Evidence for role-neutral initial processing of metaphors. Journal of experimental psychology. Learning, memory, and cognition, 26 2:529–41.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/