Built independently by an author, for readers. Read the story and support ChapterPal

keyword

generative language models

Generative language models are artificial intelligence systems designed to produce coherent and contextually relevant natural language by learning the statistical probability distribution of words and tokens across sequences. Unlike discriminative models that evaluate or classify existing text, generative models synthesize new textual content based on input prompts or contextual conditions. Primarily built upon deep neural network architectures such as autoregressive transformers and sequence-to-sequence frameworks, these models capture complex linguistic structures, semantics, and domain knowledge from large text corpora. As a result, they can perform a wide range of natural language generation tasks, including text completion, dialogue generation, question generation, translation, and task-specific reasoning through mechanisms such as fine-tuning and few-shot prompting.

2 items

Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models

Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models

Lennart Wachowiak, Dagmar Gromann

OrganizationsKing's College LondonUniversity of Vienna

Why you should read this

Evaluates GPT-3's capacity to identify conceptual metaphors and predict unconstrained source domains across English and Spanish, providing a detailed error analysis of generative language models handling cross-domain cognitive mappings.

Conceptual metaphors present a powerful cognitive vehicle to transfer knowledge structures from a source to a target domain. Prior neural approaches focus on detecting whether natural language sequences are metaphoric or literal. We believe that to truly probe metaphoric knowledge in pre-trained language models, their capability to detect this transfer should be investigated. To this end, this paper proposes to probe the ability of GPT-3 to detect metaphoric language and predict the metaphor’s source domain without any pre-set domains. We experiment with different training sample configurations for fine-tuning and few-shot prompting on two distinct datasets. When provided 12 few-shot samples in the prompt, GPT-3 generates the correct source domain for a new sample with an accuracy of 65.15% in English and 34.65% in Spanish. GPT’s most common error is a hallucinated source domain for which no indicator is present in the sentence. Other common errors include identifying a sequence as literal even though a metaphor is present and predicting the wrong source domain based on specific words in the sequence that are not metaphorically related to the target domain.

Added

2026-10-03

Generative Language Models for Paragraph-Level Question Generation

Generative Language Models for Paragraph-Level Question Generation

Asahi Ushio, Fernando Alva-Manchego, José Camacho-Collados

OrganizationsCardiff University

Why you should read this

Introduces QG-Bench, a unified multilingual and multi-domain benchmark that standardizes paragraph-level question generation evaluation across eight languages and multiple domains using sequence-to-sequence language models.

Powerful generative models have led to recent progress in question generation (QG). However, it is difficult to measure advances in QG research since there are no standardized resources that allow a uniform comparison among approaches. In this paper, we introduce QG-Bench, a multilingual and multidomain benchmark for QG that unifies existing question answering datasets by converting them to a standard QG setting. It includes general-purpose datasets such as SQuAD (Rajpurkar et al., 2016) for English, datasets from ten domains and two styles, as well as datasets in eight different languages. Using QG-Bench as a reference, we perform an extensive analysis of the capabilities of language models for the task. First, we propose robust QG baselines based on fine-tuning generative language models. Then, we complement automatic evaluation based on standard metrics with an extensive manual evaluation, which in turn sheds light on the difficulty of evaluating QG models. Finally, we analyse both the domain adaptability of these models as well as the effectiveness of multilingual models in languages other than English. QG-Bench is released along with the fine-tuned models presented in the paper,¹ which are also available as a demo.²

Added

2026-09-26