Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models

Lennart WachowiakDagmar Gromann

article2023ACL48 citations

Evaluates GPT-3's capacity to identify conceptual metaphors and predict unconstrained source domains across English and Spanish, providing a detailed error analysis of generative language models handling cross-domain cognitive mappings.

Listen

Metaphorical language is central to human communication, cognition, and persuasion, making its comprehension vital for building advanced artificial intelligence. Most prior computational approaches have focused strictly on binary detection—identifying whether a phrase is literal or figurative—or on paraphrasing metaphors using pre-defined categories and fixed grammatical rules. To probe deeper cognitive knowledge in generative language models, the article evaluates whether GPT-3 can independently identify underlying conceptual metaphors by predicting their concrete source domains when given a sentence and a target domain, without relying on fixed domain labels or grammatical constraints.

The authors conducted experiments comparing fine-tuning against few-shot prompting across multiple model sizes and varying prompt sample sizes (2 to 12 examples). Performance was assessed on a curated validation set using embedding similarity and knowledge graph similarity metrics to select the optimal model setup. The primary evaluation was carried out on a hold-out test set consisting of 633 sentences across two languages and multiple corpora: Lakoff's Master Metaphor List (general English), the VU Amsterdam Metaphor Corpus (non-metaphoric control sentences), and the Language Computer Corporation (LCC) dataset (complex political discourse in English and Spanish). Two human evaluators manually verified the correctness of the generated source domains.

The analysis showed that the largest model, davinci-002, when provided with a 12-example few-shot prompt, substantially outperformed smaller variants and fine-tuned models, achieving an overall sample-weighted accuracy of 60.22%. Performance varied widely across datasets: the model reached 81.33% accuracy on prototypical English sentences, 53.74% on complex English political texts, and 34.65% on Spanish texts, while correctly identifying non-metaphoric statements with 42.11% accuracy. Fine-tuning models reduced output diversity, causing the system to rely heavily on training examples and predict fewer unique source domains. Qualitative error analysis revealed that the most common failure was hallucinating unrelated source domains (27.32% of English errors), followed by failing to detect metaphors in figurative sentences (25.14%) and misinterpreting non-metaphorical trigger words (21.31%). In Spanish, 62.12% of errors were caused by incorrectly classifying metaphorical sentences as non-metaphoric.

These findings indicate that while large language models encode notable metaphoric knowledge, their capability degrades significantly on complex, domain-specific language and cross-lingual inputs. Fine-tuning off-the-shelf models introduces risks of reduced expressiveness, and the high rate of hallucination presents reliability risks in downstream applications like discourse monitoring or automated content generation. Organizations seeking to deploy generative models for semantic analysis cannot rely on unguided zero-shot or fine-tuned classification and must account for significant performance drops outside prototypical English texts.

To apply generative metaphor extraction effectively in production or research, practitioners should implement multi-stage pipelines: using seed words to filter target domains first, restricting context windows to avoid distraction from secondary metaphors, and incorporating few-shot prompts rather than standard fine-tuning. Further research should focus on developing automated evaluation benchmarks, predicting both source and target domains simultaneously, and building diverse multilingual datasets.

Confidence in these findings is supported by human verification showing moderate inter-annotator agreement (Cohen's Kappa of 0.51). However, users should remain cautious due to inherent data limitations: the benchmark datasets feature demographic and domain skews, manual annotation carries subjective ambiguity, and the closed, black-box nature of commercial API models limits architectural transparency and exact reproducibility.

No sufficiently relevant recommendations were found.

Cover for Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models

Abstract

Conceptual metaphors present a powerful cognitive vehicle to transfer knowledge structures from a source to a target domain. Prior neural approaches focus on detecting whether natural language sequences are metaphoric or literal. We believe that to truly probe metaphoric knowledge in pre-trained language models, their capability to detect this transfer should be investigated. To this end, this paper proposes to probe the ability of GPT-3 to detect metaphoric language and predict the metaphor’s source domain without any pre-set domains. We experiment with different training sample configurations for fine-tuning and few-shot prompting on two distinct datasets. When provided 12 few-shot samples in the prompt, GPT-3 generates the correct source domain for a new sample with an accuracy of 65.15% in English and 34.65% in Spanish. GPT’s most common error is a hallucinated source domain for which no indicator is present in the sentence. Other common errors include identifying a sequence as literal even though a metaphor is present and predicting the wrong source domain based on specific words in the sequence that are not metaphorically related to the target domain.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Conceptual Metaphors
  • 2.2 Generative Language Models
  • 3 Related Work
  • 4 Method
  • 4.1 Task
  • 4.2 Dataset
  • 4.3 Experiments and Evaluation
  • 4.3.1 Evaluation Metric
  • 4.3.2 Prompt Selection
  • 4.3.3 Manual Evaluation
  • 5 Results
  • 5.1 Prompt Selection Results
  • 5.2 Manual Evaluation Results
  • 6 Discussion
  • 5.3 Type of Errors
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • Appendix

Knowls

  1. Knowl 1 — GPT-3 can predict metaphor source domains, but accuracy varies sharply by dataset

    empirical result

    Using the selected 12-example few-shot prompt, GPT-3 davinci-002 achieved manually evaluated accuracies of 81.33% on the Metaphor List test set, 53.74% on English LCC sentences, and 34.65% on Spanish LCC sentences. It classified non-metaphoric VUA sentences correctly 42.11% of the time. Accuracy averaged over all test examples was 60.22%; the unweighted mean of the four dataset accuracies was 52.96%. The lower LCC results occurred on longer, domain-specific political language, where the supplied target domain could also be difficult to identify from the sentence. The authors report moderate overall inter-annotator agreement (Cohen’s kappa 0.51) for manual evaluation.

  2. Knowl 2 — Open-ended source-domain identification from a sentence and target domain

    model/method

    The task is to give a natural-language model a sentence and its target domain, then ask it to identify the source domain of the conceptual metaphor expressed in the sentence. For example, given You are wasting my time and the target domain time, a suitable answer is money or resource. The approach does not assume a grammatical pattern for locating metaphors and does not restrict answers to a predefined inventory of source domains. It also allows the model to answer that a sentence is not metaphorical.

  3. Knowl 3 — Training, validation, and test data combine metaphor mappings with literal controls

    data/table

    The Metaphor List portion contains 446 randomly selected sentences, with at most three examples per source–target domain combination. An additional 50 non-metaphoric English sentences were manually selected from VUA. The LCC English and Spanish samples are held out to test transfer to new source domains, longer political discourse, and another language. Metaphor List source–target combinations in validation or test do not occur in training, although validation and test share combinations with each other; LCC supplies additional combinations. The split sizes and numbers of distinct domains are: Metaphor List—117 training, 105 validation, 224 test sentences; 91 target and 94 source domains. VUA non-metaphoric—15 sentences in each split; 47 target domains and no source-domain labels. LCC English—284 test sentences, 30 target and 90 source domains. LCC Spanish—110 test sentences, 11 target and 67 source domains. Across all data, there are 132 training, 120 validation, and 633 test sentences, with 179 target domains and 251 source domains.

  4. Knowl 4 — Few-shot prompting and fine-tuning comparison setup

    experimental setup

    The experiments compared GPT-3 davinci-002 and curie-001 using prompts with 2, 4, 6, 8, or 12 labeled examples. For each prompt size, three distinct sets of training examples were tried. Fine-tuning was run for four epochs in two conditions: all 132 training sentences, or 34 sentences with one example per unique source domain. Generation temperature was set to 0 for deterministic greedy completion. Prompt and model selection used validation-set embedding similarity and knowledge-graph similarity; the selected system was then evaluated on held-out data by human annotators. Two authors independently judged English outputs and discussed disagreements; one annotator judged Spanish outputs.

  5. Knowl 5 — The 12-example prompt gave the best individual validation run, while averages favored 4–8 examples

    empirical result

    On the validation set, davinci-002 outperformed curie-001 by approximately 0.15–0.20 points on the reported similarity scores. The best individual davinci-002 run used 12 few-shot examples and reached 0.505 embedding similarity and a 0.553 knowledge-base (KB) score, but performance varied substantially across the three example selections. Averaged across runs, the highest KB score came from 8 examples, while the highest embedding similarity came from 4. Fine-tuning on all 132 training sentences yielded 0.413 embedding similarity and 0.513 KB score, below the davinci-002 few-shot results.

  6. Knowl 6 — English errors are most often unsupported source-domain guesses or missed metaphors

    empirical result

    The authors categorized English test errors as wrong with a trigger word, wrong without a trigger, too literal, should be non-metaphoric, should be metaphoric, too specific, too general, or a wrong subelement mapping. A trigger is a sentence word related to the predicted source domain but not necessarily to the metaphorical mapping. Among errors, 27.32% were wrong without a trigger, 25.14% were cases where the model said non-metaphoric despite a metaphor, and 21.31% were wrong predictions prompted by an unrelated word. Other error shares were: should be non-metaphoric 7.65%, wrong subelement mapping 7.65%, too literal 7.10%, too specific 2.73%, and too general 1.09%. For example, the model sometimes inferred animals from animal-related words that were not part of the metaphor, or chose land rather than plants for fertile ground.

  7. Knowl 7 — Spanish LCC errors especially involve missed metaphors

    empirical result

    For Spanish LCC errors, one annotator assigned the same general error categories used for English. The largest reported category was predicting that a sentence was non-metaphoric when it should have been metaphoric (62.12% of errors), followed by unsupported source-domain predictions with no apparent trigger (19.70%) and wrong predictions associated with a trigger word (13.64%). The model’s source-domain answers were in English, consistent with the English-language prompt examples. The annotator observed that family was sometimes predicted for the target government without a sentence-level trigger, and that some trigger-based answers were literal English translations of Spanish context words. Twelve LCC sentences were excluded because their gold annotations were judged faulty.

  8. Knowl 8 — Fine-tuning reduced the variety of generated source domains

    empirical result

    The best few-shot variant produced 74 distinct source-domain answers, of which 7 also appeared in training. The model fine-tuned on all training data produced 50 distinct answers, 18 of them present in training. The authors interpret this as reduced expressiveness after fine-tuning: the model relied more on source domains represented in its training data and generated fewer distinct domains overall.

  9. Knowl 9 — Embedding and knowledge-graph scores moderately track human judgments

    empirical result

    For validation, the authors compared generated and gold source-domain labels using 300-dimensional GloVe embedding similarity and a KB score formed by averaging KGvec2go similarity scores from WordNet, Wiktionary, DBpedia, and WebIsALOD. They then compared each automated score with the manual correct/incorrect judgments using Spearman’s rank correlation. The correlation was 0.43 for the KB score and 0.40 for embedding similarity; both were statistically significant at p < 0.05 and were described as moderate. These scores therefore provided some alignment with human judgments, but the paper used manual evaluation for final test accuracy.

  10. Knowl 10 — The approach depends on target-domain quality and manual judgments

    limitation

    Source-domain prediction requires both a contextual sentence and a target domain, but dataset target labels do not always precisely match what the sentence expresses. This mismatch can make the expected source domain unclear, and applying the method to unlabelled text requires a way to select the target domain in advance. Accurate evaluation also required time-consuming human judgments, with borderline decisions such as whether a prediction is too general or too specific depending on annotator interpretation. The authors further note that GPT-3 is a black-box API and that limited multilingual metaphor resources constrain conclusions beyond the tested English and Spanish data.

Coverage note — The discussion of possible persuasive or manipulative uses of metaphor analysis and proposed future applications is omitted because it is not an experimentally established contribution.

References

  1. 1.Ehsan Aghazadeh, Mohsen Fayyaz, and Yadollah Yaghoobzadeh. 2022. Metaphors in pre-trained language models: Probing and generalization across datasets and languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2037–2050, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Mateusz Babieno, Masashi Takeshita, Dusan Radisavljevic, Rafal Rzepka, and Kenji Araki. 2022. MIss RoBERTa WiLDe: Metaphor identification using masked language model with wiktionary lexical definitions. Applied Sciences, 12(4):2081.
  3. 3.Lawrence W Barsalou. 1999. Perceptual symbol systems. Behavioral and brain sciences, 22(4):577–660.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. 2014. Concreteness ratings for 40 thousand generally known english word lemmas. Behavior research methods, 46(3):904–911.
  6. 6.Siaw-Fong Chung, Kathleen Ahrens, and Chu-Ren Huang. 2004. Using WordNet and SUMO to determine source domains of conceptual metaphors. In Recent Advancement in Chinese Lexical Semantics: Proceedings of 5th Chinese Lexical Semantics Workshop (CLSW-5). Singapore: COLIPS, pages 91–98.
  7. 7.Francesca MM Citron and Adele E Goldberg. 2014. Metaphorical sentences are more emotionally engaging than their literal counterparts. Journal of cognitive neuroscience, 26(11):2585–2595.
  8. 8.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  9. 9.Christine P Dancey and John Reidy. 2007. Statistics without maths for psychology. Pearson education, Essex.
  10. 10.Ellen Dodge, Jisup Hong, and Elise Stickles. 2015. MetaNet: Deep semantic automatic metaphor analysis. In Proceedings of the Third Workshop on Metaphor in NLP, pages 40–49, Denver, Colorado. Association for Computational Linguistics.
  11. 11.Edith Durand, Pierre Berroir, and Ana Ines Ansaldo. 2018. The neural and behavioral correlates of anomia recovery following poem – personalized observation, execution, and mental imagery therapy: A proof of concept. Neural Plasticity.
  12. 12.Mengshi Ge, Rui Mao, and Erik Cambria. 2022. Explainable metaphor identification inspired by conceptual metaphor theory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36 (10), pages 10681–10689.
  13. 13.Raymond W Gibbs. 2006. Metaphor interpretation as embodied simulation. Mind & Language, 21(3):434–458.
  14. 14.Pragglejaz Group. 2007. MIP: A method for identifying metaphorically used words in discourse. Metaphor and symbol, 22(1):1–39.
  15. 15.Beata Beigman Klebanov, Chee Wee Leong, and Michael Flor. 2018. A corpus of non-native written english annotated for metaphor. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 86–91.
  16. 16.George Lakoff and Mark Johnson. 1980. Metaphors we live by. University of Chicago press.
  17. 17.J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1).
  18. 18.Jan Leike. 2022. Psa: If you want to compare InstructGPT to a base model in your research, the closest comparison is "text-davinciplus-002" with "davinci" (you might need to request access to the former). it’s not a super clean comparison, because we haven’t deployed the exact paper models. Twitter post on June 29, 2022.
  19. 19.Chee Wee (Ben) Leong, Beata Beigman Klebanov, Chris Hamill, Egon Stemle, Rutuja Ubale, and Xianyang Chen. 2020. A report on the 2020 VUA and TOEFL metaphor detection shared task. In Proceedings of the Second Workshop on Figurative Language Processing, pages 18–29, Online. Association for Computational Linguistics.
  20. 20.Emmy Liu, Chenxuan Cui, Kenneth Zheng, and Graham Neubig. 2022. Testing the ability of language models to interpret figurative language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4437–4452, Seattle, United States. Association for Computational Linguistics.
  21. 21.Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2022. Estimating the carbon footprint of bloom, a 176b parameter language model. CoRR, abs/2211.02001.
  22. 22.Rui Mao, Chenghua Lin, and Frank Guerin. 2018. Word embedding and WordNet based metaphor identification and interpretation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1222–1231, Melbourne, Australia. Association for Computational Linguistics.
  23. 23.Michael Mohler, Mary Brunson, Bryan Rink, and Marc Tomlinson. 2016. Introducing the LCC metaphor datasets. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4221–4227, Portorož, Slovenia. European Language Resources Association (ELRA).
  24. 24.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  25. 25.Paolo Pedinotti, Eliana Di Palma, Ludovica Cerini, and Alessandro Lenci. 2021. A howling success or a working sea? testing what BERT knows about metaphors. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 192–204, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  26. 26.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  27. 27.Jan Portisch, Michael Hladik, and Heiko Paulheim. 2020. KGvec2go – knowledge graph embeddings as a service. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 5641–5647, Marseille, France. European Language Resources Association.
  28. 28.Vinodkumar Prabhakaran, Marek Rei, and Ekaterina Shutova. 2021. How metaphors impact political discourse: A large-scale topic-agnostic study using neural metaphor detection. In Proceedings of the Fifteenth International AAAI Conference on Web and Social Media, ICWSM 2021, held virtually, June 7-10, 2021, pages 503–512. AAAI Press.
  29. 29.Sunny Rai and Shampa Chakraverty. 2020. A survey on computational metaphor processing. ACM Comput. Surv., 53(2).
  30. 30.Radim Rehůřek and Petr Sojka. 2010. Software framework for topic modelling with large corpora. In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks, pages 45–50, Valletta, Malta. ELRA.
  31. 31.Zachary Rosen. 2018. Computationally constructed concepts: A machine learning approach to metaphor interpretation using usage-based construction grammatical cues. In Proceedings of the Workshop on Figurative Language Processing, pages 102–109, New Orleans, Louisiana. Association for Computational Linguistics.
  32. 32.Ekaterina Shutova, Lin Sun, Elkin Darío Gutiérrez, Patricia Lichtenstein, and Srini Narayanan. 2017. Multilingual Metaphor Processing: Experiments with Semi-Supervised and Unsupervised Learning. Computational Linguistics, 43(1):71–123.
  33. 33.Gerard Steen, Lettie Dorst, Berenike Herrmann, Anna Kaal, Tina Krennmayr, and Trijntje Pasma. 2010. A method for linguistic metaphor identification: From MIP to MIPVU, volume 14. John Benjamins Publishing, Amsterdam.
  34. 34.Kevin Stowe, Nils Beck, and Iryna Gurevych. 2021a. Exploring metaphoric paraphrase generation. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 323–336, Online. Association for Computational Linguistics.
  35. 35.Kevin Stowe, Tuhin Chakrabarty, Nanyun Peng, Smaranda Muresan, and Iryna Gurevych. 2021b. Metaphor generation with conceptual mappings. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6724–6736, Online. Association for Computational Linguistics.
  36. 36.Karen Sullivan. 2013. Frames and constructions in metaphoric language, volume 14. John Benjamins Publishing, Amsterdam.
  37. 37.Xiaoyu Tong, Ekaterina Shutova, and Martha Lewis. 2021. Recent advances in neural metaphor processing: A linguistic, cognitive and social perspective. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4673–4686, Online. Association for Computational Linguistics.
  38. 38.Lennart Wachowiak, Dagmar Gromann, and Chao Xu. 2022. Drum up SUPPORT: Systematic analysis of image-schematic conceptual metaphors. In Proceedings of the 3rd Workshop on Figurative Language Processing (FLP), pages 44–53, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  39. 39.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.

Citation

MLA
Wachowiak, L., and D. Gromann. “Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 1018–32, https://doi.org/10.18653/v1/2023.acl-long.58.
APA
Wachowiak, L., & Gromann, D. (2023). Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1018–1032. https://doi.org/10.18653/v1/2023.acl-long.58
Chicago
Wachowiak, L., and D. Gromann. 2023. “Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1018–32. https://doi.org/10.18653/v1/2023.acl-long.58.
Harvard
Wachowiak, L. and Gromann, D. (2023) “Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1018–1032. Available at: https://doi.org/10.18653/v1/2023.acl-long.58.
Vancouver
1. Wachowiak L, Gromann D (2023) Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1018–1032

BibTeX

@inproceedings{wachowiak-gromann-2023-gpt,
    title = "Does {GPT}-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models",
    author = "Wachowiak, Lennart  and
      Gromann, Dagmar",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.58/",
    doi = "10.18653/v1/2023.acl-long.58",
    pages = "1018--1032"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/