It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books

Bingsheng YaoDakuo WangTongshuang WuZheng ZhangToby Jia-Jun LiMo YuYing Xu

article2022ACL58 citations

Presents an automated question-answer generation system trained on the FairytaleQA dataset that creates educationally grounded comprehension questions directly from children's storybooks to support interactive reading instruction.

Listen

Assessing and supporting children's reading comprehension requires questions that systematically evaluate narrative understanding, such as character motives, causal relationships, and key story events. While automated question answering has advanced rapidly, existing systems typically answer human prompts rather than generating pedagogically valuable questions. General-domain automated question generation approaches often rely on crowdsourced data or shallow pattern matching, failing to capture the structured narrative dimensions required in elementary education. The article addresses this gap by developing an automated question-answer generation pipeline tailored for children's storybooks to support educational assessment at scale.

The main objective of the article is to design and evaluate a multi-stage automated question-answer generation system capable of producing high-quality question-answer pairs across diverse narrative comprehension dimensions. The authors evaluate this architecture using FairytaleQA, an expert-annotated benchmark containing 10,580 question-answer pairs across 278 children's storybooks spanning kindergarten to eighth-grade reading levels. The proposed system employs a three-step pipeline: extracting candidate answers using rule-based linguistic heuristics aligned with seven pedagogical narrative elements, generating corresponding questions using a fine-tuned sequence-to-sequence language model, and ranking candidate pairs using a fine-tuned classifier to select the top outputs.

The findings show that the proposed system consistently outperforms state-of-the-art baselines across automated and human evaluations. In automated ranking evaluations, the system achieved superior precision scores at every candidate threshold; at the top-three selection level, it attained a precision score of 0.452 on the test set compared to 0.378 for a large-scale retrieval baseline and 0.305 for a standard two-step baseline. In human evaluations assessing readability, question relevancy, and answer relevancy on a five-point scale, the system scored significantly higher than the baseline in readability (4.71 versus 4.08) and question relevancy (4.39 versus 4.18). The system achieved acceptable answer relevancy (3.99), though this did not differ significantly from the baseline. Fine-tuning models directly on domain-specific narrative data yielded higher accuracy than cross-domain training, while operating at less than half the memory requirements of massive retrieval baselines.

These findings indicate that combining pedagogical rule-based answer extraction with neural question generation provides greater control over educational quality while maintaining linguistic diversity and reducing computational overhead. This capability enables automated, scalable generation of reading assessment items for classrooms, digital learning platforms, and conversational agents. A preliminary deployment in an interactive storytelling application confirmed that automated question generation can effectively engage young children and support parent-child reading activities.

Organizations developing intelligent educational tools should adopt modular, domain-specific generation architectures that ground questions in established learning frameworks rather than relying solely on general-purpose language models. Future efforts should recruit educational experts to evaluate the direct learning efficacy of the generated questions and develop multi-turn, conversational generation capabilities. Confidence in the reported results is high regarding syntactic quality, narrative relevance, and benchmark accuracy; however, additional large-scale user studies are necessary to measure formal pedagogical outcomes in diverse educational settings.

No sufficiently relevant recommendations were found.

Cover for It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books

Abstract

Existing question answering (QA) techniques are created mainly to answer questions asked by humans. But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities. We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student’s comprehension skills. Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs. Automatic and human evaluations show that our model outperforms state-of-the-art QAG baseline systems. On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 General QA Datasets
  • 2.2 The Fairytale QA Dataset
  • 2.3 QAG Task
  • 3 Pre-processing Fairytale QA Dataset
  • 4 Question Answer Generation System Architecture
  • 4.1 Heuristics-based AG Module
  • 4.2 BART-based QG Module
  • 4.3 DistilBERT-based Ranking Module
  • 5 Evaluation
  • 5.1 Automated Evaluation of QAG Task
  • 5.1.1 Baseline QAG Systems
  • 5.1.2 Evaluation Metrics
  • 5.1.3 Evaluation Results
  • 5.2 Human Evaluation of QA Generation
  • 5.3 Question Answer Generation in an Interactive Storytelling Application
  • 6 Conclusion and Future Work
  • Acknowledgements
  • References
  • Appendix
  • A Definitions and examples for 7 narrative elements labeled in Fairytale QA Dataset
  • B Distribution of Fairytale QA annotations on 7 narrative elements
  • C QAG generation examples with 3 systems
  • D User Interface of down-streaming application
  • E Fine-tuning Parameters

Knowls

  1. Knowl 1 — Three-stage educational question–answer generation pipeline

    model/method

    The system generates question–answer pairs for a story section in three stages. First, a rule-based answer-generation module extracts candidate answers from the section. Second, a BART question generator takes the section and one candidate answer as input and generates a question corresponding to that answer. Third, a DistilBERT ranker scores the resulting section–question–answer triples and returns up to a user-specified number of top-ranked pairs. This design combines pedagogically guided answer selection with neural question generation and filtering.

  2. Knowl 2 — Pedagogically guided candidate-answer extraction

    algorithm

    For each input story section, the answer-generation module constructs candidates intended to cover seven narrative-comprehension dimensions. It uses spaCy English part-of-speech analysis to extract named entities and noun chunks as candidates for character, setting, and feeling questions. For action, causal relationship, prediction, and outcome-resolution questions, it uses a PropBank semantic-role labeler to identify trigger verbs and related dependency nodes, then combines these into candidate event descriptions in subject–verb–object form. The resulting candidate answers are passed with the story section to the question generator.

  3. Knowl 3 — FairytaleQA data and narrative-comprehension dimensions

    experimental setup

    The system is developed and evaluated using FairytaleQA, an expert-annotated dataset of 10,580 question–answer pairs drawn from 278 child-oriented storybooks for readers from kindergarten through eighth grade. Questions are labeled with one or more of seven narrative dimensions: character (identify or describe a character), setting (where or when events occur), feeling (a character’s emotional state or reaction), action (a character’s behavior), causal relationship (why one event leads to another), outcome resolution (the event resulting from an earlier event), and prediction (an outcome inferable from information in the story). These labels provide the pedagogical categories the answer-extraction rules are designed to cover.

  4. Knowl 4 — FairytaleQA fine-tuning improves question generation

    empirical result

    The BART question-generation model receives a story section and a candidate answer and generates the corresponding question. On FairytaleQA, the model fine-tuned on FairytaleQA alone achieved higher Rouge-L scores than models fine-tuned on NarrativeQA alone or on the two datasets together. Validation/test scores were 0.527/0.527 for FairytaleQA, 0.424/0.442 for NarrativeQA, and 0.508/0.519 for the combined training data. The selected FairytaleQA-only model used a learning rate of 5×10−65\times10^{-6}, batch size 1, and 3 epochs.

  5. Knowl 5 — DistilBERT ranking selects a controllable number of pairs

    model/method

    The ranking module treats expert-authored FairytaleQA pairs as positive examples and generated pairs as negative examples, and fine-tunes DistilBERT to score candidate story–question–answer triples. Among the tested input arrangements, concatenating the story content, question, and answer produced the best ranking result: test F1 was 86.7%, more than 5 percentage points above the alternative arrangements. Both the content-plus-answer and content-plus-question-plus-answer settings exceeded 80% test accuracy. The ranker is used to return the top NN pairs, with NN chosen by the user.

  6. Knowl 6 — Evaluation split and variable-output metric

    experimental setup

    FairytaleQA was randomly divided by book into 232 training books (8,548 pairs), 23 validation books (1,025 pairs), and 23 test books (1,007 pairs); the paper reports that the splits have consistent statistical distributions. Because a system may generate different numbers of pairs for each story section, evaluation uses a custom MAP@NN measure for N∈{1,3,5,10}N\in\{1,3,5,10\}. For each expert-authored pair, the evaluator finds the highest Rouge-L precision between its concatenated question and answer and the concatenated question and answer of any generated pair among the top NN for the same section, then averages over expert-authored pairs. MAP@3 is treated as especially representative because sections have about three expert-authored pairs on average.

  7. Knowl 7 — Automated evaluation favors the proposed system

    empirical result

    On both FairytaleQA validation and test sets, the proposed pipeline scored higher than the two baselines at every evaluated output limit. Each pair below gives validation/test MAP@NN using Rouge-L precision on concatenated questions and answers. At N=10N=10, the proposed system scored 0.620/0.596, the 2-Step baseline 0.443/0.422, and PAQ 0.504/0.485. At N=5N=5, scores were 0.543/0.523, 0.370/0.353, and 0.436/0.424, respectively; at N=3N=3, 0.485/0.452, 0.322/0.305, and 0.387/0.378; and at N=1N=1, 0.340/0.310, 0.225/0.216, and 0.288/0.273. The proposed system’s advantage is present across output limits and is larger at the higher limits. Rouge-L measures textual overlap, so the paper also evaluates semantic and syntactic quality with human raters.

  8. Knowl 8 — Human ratings show stronger readability and question relevance than PAQ

    empirical result

    Five participants blindly rated generated pairs and expert-authored pairs on 1–5 scales for readability, question relevance to the story section, and answer relevance to its question. Ratings covered 722 pairs; inter-rater Krippendorff’s alpha across the evaluated dimensions was 0.73–0.79. Mean and standard deviation were, respectively: readability, proposed system 4.71 (0.70), PAQ 4.08 (1.13), expert-authored 4.95 (0.28); question relevance, 4.39 (1.15), 4.18 (1.22), and 4.92 (0.33); answer relevance, 3.99 (1.51), 3.90 (1.62), and 4.83 (0.57). The proposed system exceeded PAQ significantly for readability (t(477)=7.33t(477)=7.33, p<.01p<.01) and question relevance (t(477)=1.98t(477)=1.98, p<.05p<.05). Its answer-relevance score was not significantly different from PAQ (t(477)=0.58t(477)=0.58, p=.56p=.56). Expert-authored pairs had significantly higher ratings than generated pairs on all three dimensions.

  9. Knowl 9 — Interactive storytelling application and preliminary user study

    empirical result

    The question-generation pipeline was integrated into a speech-based interactive storybook application for children. As a child advances through a story, the system generates questions about the current section; it can also provide follow-up questions, conduct spoken question–answer interaction, and track performance for parents. A preliminary study with 12 parent–child pairs, with children aged 3–8, suggested that the application could sustain engaging conversations about story content; parents and children also found it useful, enjoyable, and easy to use. The study is preliminary and does not establish educational efficacy.

  10. Knowl 10 — Evaluation and deployment limitations

    limitation

    The paper does not evaluate whether the generated pairs improve children’s educational outcomes; the authors identify evaluation by educational experts as future work. The interactive-app user study is preliminary, so its favorable engagement and usability observations do not establish learning benefits. In the PAQ comparison, the supplied filtering component could not be run because it required loading the full PAQ corpus into memory and caused an out-of-memory error even with more than 50 GB of RAM; consequently, the PAQ system was evaluated without that filtering component.

Coverage note — Illustrative example pairs and interface screenshots are omitted because they demonstrate outputs and presentation rather than add a distinct method or result; the 2-Step baseline’s separate model-selection scores are omitted because they serve baseline configuration rather than the paper’s main findings.

References

  1. 1.Hao Cheng, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2020. Probabilistic assumptions matter: Improved models for distantly-supervised document-level question answering. arXiv preprint arXiv:2005.01898.
  2. 2.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:1905.03197.
  3. 3.Xinya Du, Junru Shao, and Claire Cardie. 2017. Learning to ask: Neural question generation for reading comprehension. arXiv preprint arXiv:1705.00106.
  4. 4.Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. Searchqa: A new q&a dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179.
  5. 5.Michael Heilman and Noah A Smith. 2009. Question generation via overgenerating transformations and ranking. Technical report, Carnegie-Mellon Univ Pittsburgh pa language technologies insT.
  6. 6.Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015. The goldilocks principle: Reading children’s books with explicit memory representations. arXiv preprint arXiv:1511.02301.
  7. 7.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  8. 8.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  9. 9.Tomáš Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  10. 10.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability.
  11. 11.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  12. 12.Igor Labutov, Sumit Basu, and Lucy Vanderwende. 2015. Deep questions without deep understanding. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 889–898.
  13. 13.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  14. 14.Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. Paq: 65 million probably-asked questions and what you can do with them. arXiv preprint arXiv:2102.07033.
  15. 15.Yu Li, Baolin Peng, Yelong Shen, Yi Mao, Lars Liden, Zhou Yu, and Jianfeng Gao. 2021. Knowledge-grounded dialogue generation with a unified knowledge representation. arXiv preprint arXiv:2112.07924.
  16. 16.David Lindberg, Fred Popowich, John Nesbit, and Phil Winne. 2013. Generating natural language questions to support learning on-line. In Proceedings of the 14th European Workshop on Natural Language Generation, pages 105–114.
  17. 17.Jack Mostow and Wei Chen. 2009. Generating instruction automatically for the reading strategy of self-questioning. In AIED, pages 465–472.
  18. 18.Xiangyang Mou, Chenghao Yang, Mo Yu, Bingsheng Yao, Xiaoxiao Guo, Saloni Potdar, and Hui Su. 2021. Narrative question answering with cutting-edge open-domain qa techniques: A comprehensive study. Transactions of the Association for Computational Linguistics, 9:1032–1046.
  19. 19.Preksha Nema and Mitesh M Khapra. 2018. Towards a better metric for evaluating question generation systems. arXiv preprint arXiv:1808.10192.
  20. 20.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset. In CoCo@ NIPS.
  21. 21.Martha Palmer, Daniel Gildea, and Paul Kingsbury. 2005. The proposition bank: An annotated corpus of semantic roles. Computational linguistics, 31(1):71–106.
  22. 22.Alison H Paris and Scott G Paris. 2003. Assessing narrative comprehension in young children. Reading Research Quarterly, 38(1):36–76.
  23. 23.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
  24. 24.Siva Reddy, Danqi Chen, and Christopher D Manning. 2019. Coqa: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  25. 25.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  26. 26.Thomas Scialom, Benjamin Piwowarski, and Jacopo Staiano. 2019. Self-attention architectures for answer-agnostic neural question generation. In Proceedings of the 57th annual meeting of the Association for Computational Linguistics, pages 6027–6032.
  27. 27.Siamak Shakeri, Cicero Nogueira dos Santos, Henry Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020. End-to-end synthetic data generation for domain adaptation of question answering systems. arXiv preprint arXiv:2010.06028.
  28. 28.Lynn Snyder, Donna Caccamise, and Barbara Wise. 2005. The assessment of reading comprehension: Considerations and cautions. Topics in Language Disorders, 25(1):33–50.
  29. 29.Duyu Tang, Nan Duan, Tao Qin, Zhao Yan, and Ming Zhou. 2017. Question answering and question generation as dual tasks. arXiv preprint arXiv:1706.02027.
  30. 30.Tong Wang, Xingdi Yuan, and Adam Trischler. 2017. A joint model for question answering and question generation. arXiv preprint arXiv:1706.01450.
  31. 31.Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov. 2019. Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model. arXiv preprint arXiv:1912.09637.
  32. 32.Ying Xu, Dakuo Wang, Penelope Collins, Hyelim Lee, and Mark Warschauer. 2021. Same benefits, different communication patterns: Comparing children’s reading with a conversational agent vs. a human partner. Computers & Education, 161:104059.
  33. 33.Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, and Mark Warschauer. 2022. Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative comprehension. Association for Computational Linguistics.
  34. 34.Xuchen Yao, Gosse Bouma, and Yi Zhang. 2012. Semantics-based question generation and implementation. Dialogue & Discourse, 3(2):11–42.
  35. 35.Xuchen Yao and Yi Zhang. 2010. Question generation with minimal recursion semantics. In Proceedings of QG2010: The Third Workshop on Question Generation, pages 68–75. Citeseer.
  36. 36.Zheng Zhang, Ying Xu, Yanhao Wang, Bingsheng Yao, Daniel Ritchie, Tongshuang Wu, Mo Yu, Dakuo Wang, and Toby Jia-Jun Li. 2022. Storybuddy: A human-ai collaborative agent for parent-child interactive storytelling with flexible parent involvement. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22. ACM.
  37. 37.Zhenjie Zhao, Yufang Hou, Dakuo Wang, Mo Yu, Chengzhong Liu, and Xiaojuan Ma. 2022. Educational question generation of children storybooks via question type distribution learning and event-centric summarization. Association for Computational Linguistics.
  38. 38.Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, and Ming Zhou. 2017. Neural question generation from text: A preliminary study. In National CCF Conference on Natural Language Processing and Chinese Computing, pages 662–671. Springer.

Citation

MLA
Yao, B., et al. “It Is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 731–44, https://doi.org/10.18653/v1/2022.acl-long.54.
APA
Yao, B., Wang, D., Wu, T., Zhang, Z., Li, T. J.-J., Yu, M., & Xu, Y. (2022). It is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 731–744. https://doi.org/10.18653/v1/2022.acl-long.54
Chicago
Yao, B., D. Wang, T. Wu, et al. 2022. “It Is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 731–44. https://doi.org/10.18653/v1/2022.acl-long.54.
Harvard
Yao, B. et al. (2022) “It is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 731–744. Available at: https://doi.org/10.18653/v1/2022.acl-long.54.
Vancouver
1. Yao B, Wang D, Wu T, Zhang Z, Li TJ-J, Yu M, Xu Y (2022) It is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 731–744

BibTeX

@inproceedings{yao-etal-2022-ais,
    title = "It is {AI}{'}s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children{'}s Story Books",
    author = "Yao, Bingsheng  and
      Wang, Dakuo  and
      Wu, Tongshuang  and
      Zhang, Zheng  and
      Li, Toby Jia-Jun  and
      Yu, Mo  and
      Xu, Ying",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.54/",
    doi = "10.18653/v1/2022.acl-long.54",
    pages = "731--744"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/