Built independently by an author, for readers. Read the story and support ChapterPal

keyword

neural question generation

Neural question generation is a natural language processing task that uses deep learning models to automatically produce natural, contextually relevant questions from input data such as text passages, images, or specific target answers. Unlike traditional rule-based or template-driven approaches, neural question generation utilizes end-to-end architectures, including sequence-to-sequence models and pre-trained Transformer language models, to learn complex semantic representations and syntactic structures directly from large datasets. The process often operates in an answer-aware manner to create questions that elicit specific information, or at a broader document level to assess overall comprehension. This technology is widely applied to build educational assessment tools, enhance conversational agents, and augment training datasets for question answering systems.

3 items

It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books

It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books

Bingsheng Yao, Dakuo Wang, Tongshuang Wu, Zheng Zhang, Toby Jia-Jun Li, Mo Yu, Ying Xu

OrganizationsIBMRensselaer Polytechnic InstituteTencentUniversity of California, IrvineUniversity of Notre DameUniversity of Washington

Why you should read this

Presents an automated question-answer generation system trained on the FairytaleQA dataset that creates educationally grounded comprehension questions directly from children's storybooks to support interactive reading instruction.

Existing question answering (QA) techniques are created mainly to answer questions asked by humans. But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities. We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student’s comprehension skills. Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs. Automatic and human evaluations show that our model outperforms state-of-the-art QAG baseline systems. On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario.

Added

2026-10-03

Generative Language Models for Paragraph-Level Question Generation

Generative Language Models for Paragraph-Level Question Generation

Asahi Ushio, Fernando Alva-Manchego, José Camacho-Collados

OrganizationsCardiff University

Why you should read this

Introduces QG-Bench, a unified multilingual and multi-domain benchmark that standardizes paragraph-level question generation evaluation across eight languages and multiple domains using sequence-to-sequence language models.

Powerful generative models have led to recent progress in question generation (QG). However, it is difficult to measure advances in QG research since there are no standardized resources that allow a uniform comparison among approaches. In this paper, we introduce QG-Bench, a multilingual and multidomain benchmark for QG that unifies existing question answering datasets by converting them to a standard QG setting. It includes general-purpose datasets such as SQuAD (Rajpurkar et al., 2016) for English, datasets from ten domains and two styles, as well as datasets in eight different languages. Using QG-Bench as a reference, we perform an extensive analysis of the capabilities of language models for the task. First, we propose robust QG baselines based on fine-tuning generative language models. Then, we complement automatic evaluation based on standard metrics with an extensive manual evaluation, which in turn sheds light on the difficulty of evaluating QG models. Finally, we analyse both the domain adaptability of these models as well as the effectiveness of multilingual models in languages other than English. QG-Bench is released along with the fine-tuned models presented in the paper,¹ which are also available as a demo.²

Added

2026-09-26

Unified Language Model Pre-training for Natural Language Understanding and Generation

Unified Language Model Pre-training for Natural Language Understanding and Generation

Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, Hsiao-Wuen Hon

OrganizationsMicrosoft

Why you should read this

Introduces a unified pre-training architecture that uses variable self-attention masks to support both natural language understanding and generation within a single shared Transformer, outperforming standard baselines across diverse tasks like abstractive summarization, question answering, and dialogue.

This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at this https URL.

Added

2026-09-24