It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books
Bingsheng YaoDakuo WangTongshuang WuZheng ZhangToby Jia-Jun LiMo YuYing Xu
Presents an automated question-answer generation system trained on the FairytaleQA dataset that creates educationally grounded comprehension questions directly from children's storybooks to support interactive reading instruction.
Assessing and supporting children's reading comprehension requires questions that systematically evaluate narrative understanding, such as character motives, causal relationships, and key story events. While automated question answering has advanced rapidly, existing systems typically answer human prompts rather than generating pedagogically valuable questions. General-domain automated question generation approaches often rely on crowdsourced data or shallow pattern matching, failing to capture the structured narrative dimensions required in elementary education. The article addresses this gap by developing an automated question-answer generation pipeline tailored for children's storybooks to support educational assessment at scale.
The main objective of the article is to design and evaluate a multi-stage automated question-answer generation system capable of producing high-quality question-answer pairs across diverse narrative comprehension dimensions. The authors evaluate this architecture using FairytaleQA, an expert-annotated benchmark containing 10,580 question-answer pairs across 278 children's storybooks spanning kindergarten to eighth-grade reading levels. The proposed system employs a three-step pipeline: extracting candidate answers using rule-based linguistic heuristics aligned with seven pedagogical narrative elements, generating corresponding questions using a fine-tuned sequence-to-sequence language model, and ranking candidate pairs using a fine-tuned classifier to select the top outputs.
The findings show that the proposed system consistently outperforms state-of-the-art baselines across automated and human evaluations. In automated ranking evaluations, the system achieved superior precision scores at every candidate threshold; at the top-three selection level, it attained a precision score of 0.452 on the test set compared to 0.378 for a large-scale retrieval baseline and 0.305 for a standard two-step baseline. In human evaluations assessing readability, question relevancy, and answer relevancy on a five-point scale, the system scored significantly higher than the baseline in readability (4.71 versus 4.08) and question relevancy (4.39 versus 4.18). The system achieved acceptable answer relevancy (3.99), though this did not differ significantly from the baseline. Fine-tuning models directly on domain-specific narrative data yielded higher accuracy than cross-domain training, while operating at less than half the memory requirements of massive retrieval baselines.
These findings indicate that combining pedagogical rule-based answer extraction with neural question generation provides greater control over educational quality while maintaining linguistic diversity and reducing computational overhead. This capability enables automated, scalable generation of reading assessment items for classrooms, digital learning platforms, and conversational agents. A preliminary deployment in an interactive storytelling application confirmed that automated question generation can effectively engage young children and support parent-child reading activities.
Organizations developing intelligent educational tools should adopt modular, domain-specific generation architectures that ground questions in established learning frameworks rather than relying solely on general-purpose language models. Future efforts should recruit educational experts to evaluate the direct learning efficacy of the generated questions and develop multi-turn, conversational generation capabilities. Confidence in the reported results is high regarding syntactic quality, narrative relevance, and benchmark accuracy; however, additional large-scale user studies are necessary to measure formal pedagogical outcomes in diverse educational settings.
- Paper: Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension, Ying Xu et al. (2022). Read this first to understand the FairytaleQA benchmark and its seven narrative-comprehension dimensions, which the source uses to guide and evaluate storybook question generation.
No sufficiently relevant recommendations were found.
