Built independently by an author, for readers. Read the story and support ChapterPal

keyword

unsupervised QA-based evaluation

Unsupervised QA-based evaluation is an automated assessment approach in natural language processing that measures the quality, factual consistency, or content preservation of generated text through question answering tasks without relying on human-annotated reference texts or task-specific supervised training labels. In this framework, questions are automatically generated from either the source context or the model output, and a question answering system is used to determine whether the necessary information can be accurately extracted and verified from the candidate text. By evaluating semantic comprehension and factual correctness through answerability rather than surface-level token matching, this method provides an interpretable and scalable alternative to traditional n-gram overlap metrics, especially in scenarios where multiple valid outputs exist or reference texts are unavailable.

1 item

Generative Language Models for Paragraph-Level Question Generation

Generative Language Models for Paragraph-Level Question Generation

Asahi Ushio, Fernando Alva-Manchego, José Camacho-Collados

OrganizationsCardiff University

Why you should read this

Introduces QG-Bench, a unified multilingual and multi-domain benchmark that standardizes paragraph-level question generation evaluation across eight languages and multiple domains using sequence-to-sequence language models.

Powerful generative models have led to recent progress in question generation (QG). However, it is difficult to measure advances in QG research since there are no standardized resources that allow a uniform comparison among approaches. In this paper, we introduce QG-Bench, a multilingual and multidomain benchmark for QG that unifies existing question answering datasets by converting them to a standard QG setting. It includes general-purpose datasets such as SQuAD (Rajpurkar et al., 2016) for English, datasets from ten domains and two styles, as well as datasets in eight different languages. Using QG-Bench as a reference, we perform an extensive analysis of the capabilities of language models for the task. First, we propose robust QG baselines based on fine-tuning generative language models. Then, we complement automatic evaluation based on standard metrics with an extensive manual evaluation, which in turn sheds light on the difficulty of evaluating QG models. Finally, we analyse both the domain adaptability of these models as well as the effectiveness of multilingual models in languages other than English. QG-Bench is released along with the fine-tuned models presented in the paper,¹ which are also available as a demo.²

Added

2026-09-26