FLUTE: Figurative Language Understanding through Textual Explanations

Tuhin ChakrabartyArkadiy SaakyanDebanjan GhoshSmaranda Muresan

article2022EMNLP123 citations
Listen

Figurative language—such as sarcasm, similes, metaphors, and idioms—is essential to human communication, but it presents a major hurdle for natural language processing systems. While large language models can often predict whether two sentences agree or conflict, they frequently rely on superficial statistical shortcuts rather than genuine linguistic understanding. Because existing benchmarks lack comprehensive natural language explanations for non-literal expressions, assessing whether artificial intelligence systems are making correct inferences for the right reasons has remained difficult.

The main objective of the article is to establish a rigorous benchmark for evaluating whether language models truly comprehend figurative text by requiring them to jointly determine logical relationships and generate plain-language explanations justifying their decisions. To achieve this, the authors created and evaluated FLUTE, a benchmark containing 9,000 sentence pairs paired with explanatory rationales across sarcasm, similes, metaphors, and idioms.

To construct this resource efficiently, the authors implemented a scalable collaborative framework combining advanced generative models with human oversight. Generative artificial intelligence produced candidate paraphrases, contradictions, and draft explanations, while crowd workers and expert reviewers filtered, validated, and edited the outputs to ensure high quality and prevent algorithmic bias. The researchers then fine-tuned an instructional language model on FLUTE using a multitask setup and evaluated its reasoning performance against an alternative baseline trained on a standard, fifty-times larger literal inference dataset.

The investigation produced several key findings. First, models trained on general literal data degraded sharply when tasked with explaining figurative reasoning; when requiring high-quality explanations, their effective accuracy dropped from roughly 60–85% down to under 12%. In contrast, the model fine-tuned on FLUTE maintained significantly stronger combined reasoning and explanation performance across all figurative categories, including an accuracy above 56% on sarcasm under strict explanation thresholds. Second, independent human evaluations demonstrated that the FLUTE-trained model produced explanations accepted by evaluators at substantially higher rates, outscoring the baseline by 22 to 51 points across categories while generating 28.5% fewer outright rejections. Finally, crowd evaluators revealed that standard models routinely generated trivial, repetitive, or incomplete justifications, whereas FLUTE-trained models successfully captured underlying cultural and contextual meaning.

These findings indicate that general dataset scale cannot substitute for specialized reasoning data when training models to interpret nuanced, non-literal language. For organizations deploying conversational agents, customer sentiment tools, or automated content moderation systems, integrating targeted explanatory data reduces the risk of models misinterpreting non-literal nuances. Requiring models to provide self-rationalized explanations improves system transparency and allows stakeholders to verify model reliability before deployment.

Organizations developing or deploying language models should adopt joint classification and explanation frameworks to audit reasoning capabilities, especially in high-stakes domains where figurative nuances alter meaning. While FLUTE significantly advances this area, its scope is primarily centered on four figurative categories and predominantly negative emotional contexts in sarcasm. Future initiatives should expand into a wider variety of cultural expressions, broader sarcasm styles, and comprehensive evaluations of model explanation truthfulness to ensure dependable real-world performance.

No sufficiently relevant recommendations were found.

Table of Contents

  • 1 Introduction
  • 2 Model-in-the-loop for building FLUTE
  • 2.1.2 Simile
  • 2.1.3 Metaphors
  • 2.1.4 Idioms
  • 2.2 FLUTE: Generating Textual Explanations
  • 3 Experimental Setup
  • 3.1 Models
  • 3.2 Evaluation Setup
  • 4 Results and Discussion
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Sarcasm dataset
  • A.1.1 Hyperparameters for the sarcasm dataset
  • A.1.2 Prompts for generation of Paraphrase from which Sarcasm is created
  • A.1.3 Prompts for generation of Explanation for Paraphrase from which Sarcasm is created (Entailment)
  • A.1.4 Prompts for generation of Explanation for Sarcasm (Contradiction)
  • A.2 Simile dataset
  • A.2.2 Challenging instances for the simile dataset
  • A.2.3 Prompts for Simile Paraphrase Generation
  • A.2.4 Prompts for Simile Contradiction Generation
  • A.2.5 Prompts for Simile Contradiction Explanation Generation
  • A.3 Prompts for Simile Entailment Explanation Generation
  • A.4 Metaphor dataset
  • A.4.1 Hyperparameters for the metaphor dataset
  • A.4.2 Prompts for Metaphor Paraphrase Generation
  • A.4.3 Prompts for Metaphor Contradiction Generation
  • A.4.4 Prompts for Metaphor Entailment Explanation Generation
  • A.4.5 Prompts for Metaphor Contradiction Explanation Generation
  • A.5 Idiom dataset
  • A.5.1 Hyperparameters for the idiom dataset
  • A.6 Prompts for joint paraphrase and contradiction generation for idioms
  • A.6.1 Prompts for joint generation of Explanation for idioms
  • B Details of Human evaluation

Knowls

  1. Knowl 1 — FLUTE pairs literal premises with figurative hypotheses and explanations

    data/table

    FLUTE is a benchmark for natural-language inference about figurative language. Each instance pairs a literal premise with a figurative hypothesis, assigns an entailment or contradiction label, and provides a free-form explanation focused on the figurative expression. The benchmark covers sarcasm, simile, metaphor, and idiom; the paper describes it as a dataset of 9,000 instances. Its reported category counts are:

    Category Entails Contradicts Total
    Paraphrase 1,339 – 1,339
    + Sarcasm – 2,678 2,678
    Simile 750 750 1,500
    Metaphor 750 750 1,500
    Idiom 1,000 1,000 2,000

    For sarcasm, the construction uses literal seed sentences, generated literal paraphrases, and sarcastic versions; simile, metaphor, and idiom each have balanced entailment and contradiction subsets. The explanations and labels are intended to reflect the figurative expression, rather than unrelated content elsewhere in a sentence.

  2. Knowl 2 — Sarcasm instances are built from emotion-conditioned paraphrases and crowd-edited sarcasm

    model/method

    For sarcasm, FLUTE starts with literal sentences from Empathetic Dialogues that are associated with negative emotions such as anger, fear, or embarrassment. GPT-3 is prompted with both a sentence and its emotion to produce a creative paraphrase that retains the meaning while making the emotional content more explicit. Three experts check each paraphrase; an inadequate output is resampled, up to three rounds, and 15% of paraphrases were accepted on the first round.

    Crowd workers on Amazon Mechanical Turk then minimally edit the checked paraphrases into sarcastic sentences, for example by replacing an emotion word with its opposite. Two workers independently edit each input, and three experts verify the outputs; experts edited 25% of the sarcastic outputs judged insufficient. FLUTE pairs literal material with an entailing paraphrase and a sarcastic contradiction. The sarcasm construction uses 1,339 seed literal sentences and 1,339 paraphrased sentences.

  3. Knowl 3 — Simile entailments preserve the comparison property, while contradictions invert it

    model/method

    For each simile, the construction first replaces the comparison with its relevant property to make an auxiliary literal sentence—for example, replacing “like a steam sauna” with “hot and moist.” GPT-3 then generates a literal paraphrase consistent with that property. Experts accepted 720 generated entailment instances.

    For contradictions, GPT-3 is prompted to invert the paraphrase with respect to the same property, and experts accepted 642 of the 720 generated contradictions. The dataset also incorporates 108 challenging simile instances selected from an existing dataset using low RoBERTa logits for the correct entailment/contradiction choices. The resulting FLUTE simile subset contains 750 entailments and 750 contradictions.

  4. Knowl 4 — Metaphor contradictions target the metaphorical meaning

    model/method

    The metaphor subset begins with 750 metaphors selected from existing datasets. GPT-3 generates literal paraphrases that express the metaphors’ intended meanings. The authors do not rely on existing literal counterparts when those are absent or amount to minimal edits that could create trivial classification cues.

    GPT-3 also generates contradictions to the intended metaphorical meaning. This design avoids counterexamples that merely use a metaphor-associated word in a different sense or contradict unrelated sentence content. Three experts verify and edit the generated paraphrases and contradictions as needed. FLUTE contains 750 metaphor entailments and 750 metaphor contradictions.

  5. Knowl 5 — Idiom meaning guides joint generation of entailments and contradictions

    model/method

    For idiom instances, GPT-3 receives the sentence containing the idiom and the idiom’s meaning, then jointly generates a paraphrase consistent with that meaning and a contrasting sentence that contradicts it. Three experts check the generated sentences and edited 23% of the generations. The resulting idiom subset contains 1,000 entailments and 1,000 contradictions.

  6. Knowl 6 — GPT-3 explanations are checked and edited by experts

    model/method

    FLUTE supplies a natural-language explanation for every premise–hypothesis pair. For sarcasm, simile, and metaphor, GPT-3 receives the premise, hypothesis, and entailment or contradiction label and generates an explanation; entailment and contradiction explanations are prompted separately. For idioms, the prompt additionally includes the idiom and its meaning, which are used to generate explanations for both relations.

    Three experts review the explanations for correctness, coherence, logical consistency, grammaticality, and whether they explain the figurative relation rather than rely on a simple template. They edited 21% of sarcasm/paraphrase explanations, 27% of simile explanations, 40% of metaphor explanations, and 10% of idiom explanations.

  7. Knowl 7 — A joint T5 model predicts an NLI label and its explanation

    model/method

    The baseline is a 3-billion-parameter T5 model fine-tuned to produce both an entailment/contradiction label and a rationale in one output. Its input instruction asks whether a literal premise entails or contradicts a figurative hypothesis and requests a one-sentence explanation. Two training variants are compared.

    The e-SNLI variant is trained for one epoch on 366,603 examples, with 6,607 validation examples. Neutral examples are removed; because e-SNLI can contain multiple explanations per example, the longest is retained, and multiple sentences are joined with “and.” The FLUTE variant is trained jointly on all four figurative-language types for 10 epochs using 7,035 examples and 500 validation examples. Both variants use batch size 1,024, AdamW, and learning rate 10−410^{-4}; the FLUTE validation loss selects its checkpoint.

  8. Knowl 8 — Evaluation tests label prediction and explanation quality

    experimental setup

    The test set contains 1,500 instances: 750 sarcasm examples and 250 examples from each of simile, metaphor, and idiom. Automatic explanation quality is measured by the average of BERTScore and BLEURT, reported on a 0–100 scale. Label accuracy is reported without an explanation-quality threshold (Accuracy@0) and with thresholds requiring an explanation score greater than 50 or 60 (Accuracy@50 and Accuracy@60).

    For human evaluation, 200 instances—50 per figurative-language type—are assessed by three crowd workers each. The workers see the two models’ explanations in blinded order and judge whether each explanation justifies the label. Responses Yes, Weak Yes, Weak No, and No are mapped to 1, 2/32/3, 1/31/3, and 0 to compute Hscore. The 79 recruited workers had at least 98% task-approval rates; inter-annotator agreement was Krippendorff’s α=0.45\alpha=0.45.

    Rationale quality is also assessed through simulatability: a classifier predicting the gold NLI label from the input plus a rationale is compared with a classifier using the input alone. The comparison is made using both model-generated rationales and gold explanations.

  9. Knowl 9 — FLUTE-trained T5 produces higher-quality explanations than e-SNLI-trained T5

    data/table

    Across all four figurative-language types, the FLUTE-trained model has higher explanation-thresholded accuracy and human Hscore than the e-SNLI-trained model. Raw label accuracy (Accuracy@0) is not uniformly higher: e-SNLI-trained T5 has higher raw accuracy on metaphors and idioms. The human Yes% and No% columns are the proportions of judgments in those response categories; higher Yes% and lower No% indicate better-justifying explanations.

    T5 trained on e-SNLI T5 trained on FLUTE
    Type Acc@0 Acc@50 Acc@60 Hscore Yes% No% Acc@0 Acc@50 Acc@60 Hscore Yes% No%
    Sarcasm 60.6 15.7 2.4 34.2 14.7 52.0 91.6 86.2 56.2 85.3 75.3 8.7
    Simile 61.2 22.8 3.6 43.6 22.0 40.7 62.8 57.2 30.4 84.9 74.7 8.0
    Metaphor 81.8 31.8 11.6 55.3 36.0 28.0 73.3 55.6 23.7 80.2 64.0 6.0
    Idiom 84.8 46.4 7.6 60.9 37.3 24.7 79.2 77.2 66.8 83.1 69.3 8.7

    The rationale-quality comparison reports accuracy on the FLUTE test set for classifiers trained on each source:

    Rationale condition Acc(IR→OIR \to O) Acc(I→OI \to O) RQ
    e-SNLI-trained model, predicted rationale 68.4 74.5 -6.1
    FLUTE-trained model, predicted rationale 89.3 90.5 -1.2
    FLUTE-trained model, gold rationale 95.6 90.5 5.1

    Here, II is the input, RR is a rationale, and OO is the gold label; RQ is the difference between accuracy with and without the rationale. Predicted rationales yield negative RQ for both models, but the FLUTE-trained model’s value is closer to zero. Gold FLUTE explanations yield positive RQ, indicating that they help the classifier predict the gold label. The paper reports Wilcoxon signed-rank significance tests, with p<0.001p<0.001 for the rationale-quality comparisons and for the bolded results in its accuracy and human-evaluation results.

  10. Knowl 10 — FLUTE covers a limited range of figurative language and does not test explanation faithfulness

    limitation

    The benchmark covers only sarcasm, simile, metaphor, and idiom, while figurative-language interpretation can depend on a much wider range of cultural knowledge and context. Its sarcasm examples emphasize sentiment incongruity and do not represent all forms, such as situational, underplayed, or dramatic sarcasm. The study also does not explicitly evaluate whether model-generated explanations are faithful to the models’ decision processes, nor does it conduct a dedicated human study of explanation credibility or bias.

Coverage note — The full appendix prompt wording and GPT-3 decoding settings are omitted: the knowls retain the generation stages, model roles, and expert-review procedures, while the individual prompts and sampling values are implementation details rather than separate findings.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 298–306.
  2. 2.Rodrigo Agerri. 2008. Metaphor in textual entailment. In Coling 2008: Companion volume: Posters, pages 3–6, Manchester, UK. Coling 2008 Organizing Committee.
  3. 3.Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, and Bernardo Magnini. 2006. The second pascal recognising textual entailment challenge. In Computer Science.
  4. 4.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8718–8735, Online. Association for Computational Linguistics.
  5. 5.Conrad Borchers, Dalia Gala, Benjamin Gilburt, Eduard Oravkin, Wilfried Bounsi, Yuki M Asano, and Hannah Kirk. 2022. Looking for a handsome carpenter! debiasing GPT-3 job advertisements. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 212–224, Seattle, Washington. Association for Computational Linguistics.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc.
  8. 8.Tuhin Chakrabarty, Yejin Choi, and Vered Shwartz. 2022. It’s not Rocket Science: Interpreting Figurative Language in Narratives. Transactions of the Association for Computational Linguistics, 10:589–606.
  9. 9.Tuhin Chakrabarty, Debanjan Ghosh, Adam Poliak, and Smaranda Muresan. 2021. Figurative language in recognizing textual entailment. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3354–3361, Online. Association for Computational Linguistics.
  10. 10.Aaron Chan, Shaoliang Nie, Liang Tan, Xiaochang Peng, Hamed Firooz, Maziar Sanjabi, and Xiang Ren. 2022. Frame: Evaluating simulatability metrics for free-text rationales. arXiv preprint arXiv:2207.00779.
  11. 11.Dushyant Singh Chauhan, Dhanush S R, Asif Ekbal, and Pushpak Bhattacharyya. 2020. Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4351–4360, Online. Association for Computational Linguistics.
  12. 12.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006. The pascal recognising textual entailment challenge. In Machine learning challenges. evaluating predictive uncertainty, visual object classification, and recognising tectual entailment, pages 177–190. Springer.
  13. 13.Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608.
  14. 14.Yufei Feng, Xiaoyu Yang, Xiaodan Zhu, and Michael Greenspan. 2022. Neuro-symbolic natural logic with introspective revision for natural language inference. Transactions of the Association for Computational Linguistics, 10:240–256.
  15. 15.Susan R Fussell and Mallie M Moss. 1998. Figurative language in emotional communication. Social and cognitive approaches to interpersonal communication, pages 113–141.
  16. 16.Debanjan Ghosh, Elena Musi, and Smaranda Muresan. 2020. Interpreting verbal irony: Linguistic strategies and the connection to the Type of semantic incongruity. In Proceedings of the Society for Computation in Linguistics 2020, pages 82–93, New York, New York. Association for Computational Linguistics.
  17. 17.Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018. Breaking NLI systems with sentences that require simple lexical inferences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 650–655, Melbourne, Australia. Association for Computational Linguistics.
  18. 18.Jack Hessel, Ana Marasovic, Jena D Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi. 2022. Do androids laugh at electric sheep? humor" understanding" benchmarks from the new yorker caption contest. arXiv preprint arXiv:2209.06293.
  19. 19.Harsh Jhamtani, Varun Gangal, Eduard Hovy, and Taylor Berg-Kirkpatrick. 2021. Investigating robustness of dialog models to popular figurative language constructs. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7476–7485, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  20. 20.Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021. E-vil: A dataset and benchmark for natural language explanations in vision-language tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1244–1254.
  21. 21.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability. University of Pennsylvania ScholarlyCommons.
  22. 22.Shivani Kumar, Atharva Kulkarni, Md Shad Akhtar, and Tanmoy Chakraborty. 2022. When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5956–5968, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi. 2022a. Wanli: Worker and ai collaboration for natural language inference dataset creation. arXiv preprint arXiv:2201.05955.
  24. 24.Emmy Liu, Chen Cui, Kenneth Zheng, and Graham Neubig. 2022b. Testing the ability of language models to interpret figurative language.
  25. 25.Li Lucy and David Bamman. 2021. Gender and representation bias in GPT-3 generated stories. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55, Virtual. Association for Computational Linguistics.
  26. 26.Bodhisattwa Prasad Majumder, Oana-Maria Camburu, Thomas Lukasiewicz, and Julian McAuley. 2021. Rationale-inspired natural language explanations with commonsense. arXiv preprint arXiv:2106.13876.
  27. 27.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  28. 28.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Cross-task generalization via natural language crowdsourcing instructions. arXiv preprint arXiv:2104.08773.
  29. 29.Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. Wt5?! training text-to-text models to explain their predictions. arXiv preprint arXiv:2004.14546.
  30. 30.Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018a. Collecting diverse natural language inference problems for sentence representation evaluation. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 337–340, Brussels, Belgium. Association for Computational Linguistics.
  31. 31.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018b. Hypothesis only baselines in natural language inference. arXiv preprint arXiv:1805.01042.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  34. 34.Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Explain yourself! leveraging language models for commonsense reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4932–4942, Florence, Italy. Association for Computational Linguistics.
  35. 35.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  36. 36.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " why should i trust you?" explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pages 1135–1144.
  37. 37.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  38. 38.Richard M Roberts and Roger J Kreuz. 1994. Why do people use figurative language? Psychological science, 5(3):159–163.
  39. 39.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization. arXiv.
  40. 40.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  41. 41.Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. Societal biases in language generation: Progress and challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4275–4293, Online. Association for Computational Linguistics.
  42. 42.Ekaterina V Shutova. 2011. Computational approaches to figurative language. Technical report, University of Cambridge, Computer Laboratory.
  43. 43.Suzanna Sia, Anton Belyy, Amjad Almahairi, Madian Khabsa, Luke Zettlemoyer, and Lambert Mathias. 2022. Logical satisfiability of counterfactuals for faithful explanations in nli. arXiv preprint arXiv:2205.12469.
  44. 44.Aarohi Srivastava, Abhinav Rastogi, and Abhishek Rao. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. In In preparation.
  45. 45.Kevin Stowe, Prasetya Utama, and Iryna Gurevych. 2022. IMPLI: Investigating NLI models’ performance on figurative language. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5375–5388, Dublin, Ireland. Association for Computational Linguistics.
  46. 46.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. arXiv.
  47. 47.Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2021a. Reframing human-ai collaboration for generating free-text explanations. arXiv preprint arXiv:2112.08674.
  48. 48.Sarah Wiegreffe, Ana Marasovic, and Noah A. Smith. 2021b. Measuring association between labels and free-text rationales. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10266–10284, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  49. 49.Yordan Yordanov, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2021. Few-shot out-of-domain transfer learning of natural language explanations. arXiv preprint arXiv:2112.06204.
  50. 50.Omar Zaidan and Jason Eisner. 2008. Modeling annotators: A generative approach to learning from annotator rationales. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 31–40, Honolulu, Hawaii. Association for Computational Linguistics.
  51. 51.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  52. 52.Jianing Zhou, Hongyu Gong, and Suma Bhat. 2021. PIE: A parallel idiomatic expression corpus for idiomatic sentence generation and paraphrasing. In Proceedings of the 17th Workshop on Multiword Expressions (MWE 2021), pages 33–48, Online. Association for Computational Linguistics.
  53. 53.Liu Zhuang, Lin Wayne, Shi Ya, and Zhao Jun. 2021. A robustly optimized BERT pre-training approach with post-training. In Proceedings of the 20th Chinese National Conference on Computational Linguistics, pages 1218–1227, Huhhot, China. Chinese Information Processing Society of China.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/