FLUTE: Figurative Language Understanding through Textual Explanations
Tuhin ChakrabartyArkadiy SaakyanDebanjan GhoshSmaranda Muresan
Figurative language—such as sarcasm, similes, metaphors, and idioms—is essential to human communication, but it presents a major hurdle for natural language processing systems. While large language models can often predict whether two sentences agree or conflict, they frequently rely on superficial statistical shortcuts rather than genuine linguistic understanding. Because existing benchmarks lack comprehensive natural language explanations for non-literal expressions, assessing whether artificial intelligence systems are making correct inferences for the right reasons has remained difficult.
The main objective of the article is to establish a rigorous benchmark for evaluating whether language models truly comprehend figurative text by requiring them to jointly determine logical relationships and generate plain-language explanations justifying their decisions. To achieve this, the authors created and evaluated FLUTE, a benchmark containing 9,000 sentence pairs paired with explanatory rationales across sarcasm, similes, metaphors, and idioms.
To construct this resource efficiently, the authors implemented a scalable collaborative framework combining advanced generative models with human oversight. Generative artificial intelligence produced candidate paraphrases, contradictions, and draft explanations, while crowd workers and expert reviewers filtered, validated, and edited the outputs to ensure high quality and prevent algorithmic bias. The researchers then fine-tuned an instructional language model on FLUTE using a multitask setup and evaluated its reasoning performance against an alternative baseline trained on a standard, fifty-times larger literal inference dataset.
The investigation produced several key findings. First, models trained on general literal data degraded sharply when tasked with explaining figurative reasoning; when requiring high-quality explanations, their effective accuracy dropped from roughly 60–85% down to under 12%. In contrast, the model fine-tuned on FLUTE maintained significantly stronger combined reasoning and explanation performance across all figurative categories, including an accuracy above 56% on sarcasm under strict explanation thresholds. Second, independent human evaluations demonstrated that the FLUTE-trained model produced explanations accepted by evaluators at substantially higher rates, outscoring the baseline by 22 to 51 points across categories while generating 28.5% fewer outright rejections. Finally, crowd evaluators revealed that standard models routinely generated trivial, repetitive, or incomplete justifications, whereas FLUTE-trained models successfully captured underlying cultural and contextual meaning.
These findings indicate that general dataset scale cannot substitute for specialized reasoning data when training models to interpret nuanced, non-literal language. For organizations deploying conversational agents, customer sentiment tools, or automated content moderation systems, integrating targeted explanatory data reduces the risk of models misinterpreting non-literal nuances. Requiring models to provide self-rationalized explanations improves system transparency and allows stakeholders to verify model reliability before deployment.
Organizations developing or deploying language models should adopt joint classification and explanation frameworks to audit reasoning capabilities, especially in high-stakes domains where figurative nuances alter meaning. While FLUTE significantly advances this area, its scope is primarily centered on four figurative categories and predominantly negative emotional contexts in sarcasm. Future initiatives should expand into a wider variety of cultural expressions, broader sarcasm styles, and comprehensive evaluations of model explanation truthfulness to ensure dependable real-world performance.
- Paper: Testing the Ability of Language Models to Interpret Figurative Language, Emmy Liu et al. (2022). This figurative-language benchmark provides a direct precursor for FLUTE’s focus on testing whether models can infer nonliteral meaning, before FLUTE adds explanations as a required part of evaluation.
- Paper: Annotation Artifacts in Natural Language Inference Data, Suchin Gururangan et al. (2018). Its demonstration that NLI systems exploit annotation artifacts establishes the shortcut problem that FLUTE addresses when evaluating figurative inference.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). HANS shows how high NLI scores can mask shallow heuristics, clarifying why FLUTE tests both inference accuracy and the quality of models’ explanations.
- Paper: Supervising Model Attention with Human Explanations for Robust Natural Language Inference, Joe Stacey et al. (2022). This study uses human explanations to guide NLI models, providing useful methodological groundwork for FLUTE’s use of explanatory rationales in evaluating inference.
No sufficiently relevant recommendations were found.