Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization
Zhenjie ZhaoYufang HouDakuo WangMo YuChengzhong LiuXiaojuan Ma
Proposes a framework that pairs question type distribution learning with event-centric summarization to automatically generate high-cognitive-demand educational questions from children's storybooks.
Asking educationally meaningful questions during storybook reading is critical for fostering children's literacy and cognitive development. However, parents and educators often lack the time or training to formulate questions that require higher-level thinking, such as synthesizing narrative elements or inferring causal relationships. While automated conversational agents can assist in shared reading, existing question generation technologies typically rely on single-word answers or simple factual spans within a single sentence, producing shallow factual questions rather than the deep, event-driven inquiries needed for learning.
The article demonstrates an automated system designed to generate high-cognitive-demand educational questions from children's storybooks. The primary objective is to evaluate whether separating the generation process into distinct stages—predicting the appropriate types and counts of questions and summarizing the underlying narrative events—can produce questions better tailored to early childhood education than existing end-to-end approaches.
To accomplish this, the authors constructed a three-stage machine learning framework. First, a language model analyzes a storybook paragraph to predict the distribution and number of relevant educational question types. Second, a summarization model uses these predicted types as control signals to extract salient, multi-sentence plot events, trained on silver-standard statements derived from paired questions and answers. Third, a final sequence-to-sequence model generates targeted questions directly from these event summaries. The system was trained and evaluated on the expert-annotated FairytaleQA dataset, focusing specifically on complex question categories: actions, causal relationships, and outcome resolutions. The authors benchmarked the framework against standard end-to-end models and an existing keyword-based question-generation baseline using both automated linguistic metrics and human expert evaluations.
The findings show that the proposed multi-stage approach substantially outperforms baseline systems across major performance indicators. In automated lexical and semantic evaluations, the method achieved a precision of 37.50% and an F1 score of 30.58% on the test set, exceeding end-to-end models by approximately 20 and 10 percentage points, respectively, and outperforming the strongest baseline by roughly 5 to 10 points. In human evaluations, expert raters found that the framework predicted the distribution of question types significantly closer to human experts, achieving a statistical distance of 0.28 compared to 0.60 for the baseline. Crucially, on a five-point scale measuring appropriateness for a five-year-old child, the proposed model scored 2.56, significantly outperforming the baseline score of 2.22. Ablation studies further revealed that omitting question type prediction caused overall generation accuracy to drop by about three points, and testing on perfect ground-truth summaries yielded an upper-bound F1 score of 87.67%, confirming that high-quality event summarization is the critical driver of overall performance.
These results imply that generating cognitively demanding educational questions requires explicitly modeling the narrative structure and intent rather than generating text in a single unstructured step. Decomposing the task into type distribution and event summarization allows educational AI tools to move beyond simple fact retrieval toward meaningful narrative comprehension. For organizations deploying conversational AI in early childhood learning, this architecture offers a viable technical pathway to deliver guided, interactive reading experiences that emulate expert educators.
Moving forward, developers and educational technologists should adopt multi-stage, event-centric architectures when designing reading comprehension systems. To advance from research prototypes to real-world applications, future efforts should prioritize improving the event-centric summarizer by incorporating discourse-level narrative modeling. Furthermore, stakeholders must address the primary operational risk identified in the article: the tendency of generative models to occasionally introduce factual errors or hallucinations regarding story details. Implementing knowledge-grounding mechanisms and conducting structured pilot tests in interactive reading environments will be essential before deploying these agents directly to young learners.
- Paper: It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books, Bingsheng Yao et al. (2022). Its children’s-story question–answer generation pipeline establishes the FairytaleQA-based educational question-generation setting that this work refines with type-distribution prediction and event summaries.
- Paper: Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension, Ying Xu et al. (2022). Its FairytaleQA dataset and narrative-element taxonomy provide the benchmark and pedagogical categories needed to understand the source’s training and evaluation setup.
No sufficiently relevant recommendations were found.
