Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization

Zhenjie ZhaoYufang HouDakuo WangMo YuChengzhong LiuXiaojuan Ma

article2022ACL60 citations

Proposes a framework that pairs question type distribution learning with event-centric summarization to automatically generate high-cognitive-demand educational questions from children's storybooks.

Listen

Asking educationally meaningful questions during storybook reading is critical for fostering children's literacy and cognitive development. However, parents and educators often lack the time or training to formulate questions that require higher-level thinking, such as synthesizing narrative elements or inferring causal relationships. While automated conversational agents can assist in shared reading, existing question generation technologies typically rely on single-word answers or simple factual spans within a single sentence, producing shallow factual questions rather than the deep, event-driven inquiries needed for learning.

The article demonstrates an automated system designed to generate high-cognitive-demand educational questions from children's storybooks. The primary objective is to evaluate whether separating the generation process into distinct stages—predicting the appropriate types and counts of questions and summarizing the underlying narrative events—can produce questions better tailored to early childhood education than existing end-to-end approaches.

To accomplish this, the authors constructed a three-stage machine learning framework. First, a language model analyzes a storybook paragraph to predict the distribution and number of relevant educational question types. Second, a summarization model uses these predicted types as control signals to extract salient, multi-sentence plot events, trained on silver-standard statements derived from paired questions and answers. Third, a final sequence-to-sequence model generates targeted questions directly from these event summaries. The system was trained and evaluated on the expert-annotated FairytaleQA dataset, focusing specifically on complex question categories: actions, causal relationships, and outcome resolutions. The authors benchmarked the framework against standard end-to-end models and an existing keyword-based question-generation baseline using both automated linguistic metrics and human expert evaluations.

The findings show that the proposed multi-stage approach substantially outperforms baseline systems across major performance indicators. In automated lexical and semantic evaluations, the method achieved a precision of 37.50% and an F1 score of 30.58% on the test set, exceeding end-to-end models by approximately 20 and 10 percentage points, respectively, and outperforming the strongest baseline by roughly 5 to 10 points. In human evaluations, expert raters found that the framework predicted the distribution of question types significantly closer to human experts, achieving a statistical distance of 0.28 compared to 0.60 for the baseline. Crucially, on a five-point scale measuring appropriateness for a five-year-old child, the proposed model scored 2.56, significantly outperforming the baseline score of 2.22. Ablation studies further revealed that omitting question type prediction caused overall generation accuracy to drop by about three points, and testing on perfect ground-truth summaries yielded an upper-bound F1 score of 87.67%, confirming that high-quality event summarization is the critical driver of overall performance.

These results imply that generating cognitively demanding educational questions requires explicitly modeling the narrative structure and intent rather than generating text in a single unstructured step. Decomposing the task into type distribution and event summarization allows educational AI tools to move beyond simple fact retrieval toward meaningful narrative comprehension. For organizations deploying conversational AI in early childhood learning, this architecture offers a viable technical pathway to deliver guided, interactive reading experiences that emulate expert educators.

Moving forward, developers and educational technologists should adopt multi-stage, event-centric architectures when designing reading comprehension systems. To advance from research prototypes to real-world applications, future efforts should prioritize improving the event-centric summarizer by incorporating discourse-level narrative modeling. Furthermore, stakeholders must address the primary operational risk identified in the article: the tendency of generative models to occasionally introduce factual errors or hallucinations regarding story details. Implementing knowledge-grounding mechanisms and conducting structured pilot tests in interactive reading environments will be essential before deploying these agents directly to young learners.

No sufficiently relevant recommendations were found.

Cover for Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization

Abstract

Generating educational questions of fairytales or storybooks is vital for improving children’s literacy ability. However, it is challenging to generate questions that capture the interesting aspects of a fairytale story with educational meaningfulness. In this paper, we propose a novel question generation method that first learns the question type distribution of an input story paragraph, and then summarizes salient events which can be used to generate high-cognitive-demand questions. To train the event-centric summarizer, we fine-tune a pre-trained transformer-based sequence-to-sequence model using silver samples composed by educational question-answer pairs. On a newly proposed educational question-answering dataset FairytaleQA, we show good performance of our method on both automatic and human evaluation metrics. Our work indicates the necessity of decomposing question type distribution learning and event-centric summary generation for educational question generation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Question Generation
  • 2.2 Text Summarization
  • 3 Method
  • 3.1 Question Type Distribution Learning
  • 3.2 Event-centric Summary Generation
  • 3.3 Educational Question Generation
  • 4 Experimental Setup
  • 4.1 Dataset
  • 4.2 Baselines
  • 4.3 Evaluation Metrics
  • 4.3.1 Automatic Evaluation
  • 4.3.2 Human Evaluation
  • 4.4 Implementation Details
  • 5 Results and Analysis
  • 5.1 Automatic Evaluation Results
  • 5.2 Human Evaluation Results
  • 6 System Analysis
  • 6.1 Question Type Distribution Learning
  • 6.2 Event-centric Summary Generation
  • 6.3 Multi-task Learning of Question Types
  • 7 Conclusion
  • Acknowledgments
  • References
  • Appendix
  • A Dataset Statistics
  • B Potential Risks
  • C Examples of Generated Questions

Knowls

  1. Knowl 1 — Three-stage generation of educational storybook questions

    model/method

    The system generates educational questions from a story paragraph in three stages. First, a question-type predictor estimates which types of questions are appropriate and how many to generate for each type. Second, a type- and order-conditioned summarizer produces one event-centric summary for each planned question, selecting or expressing story events relevant to that question type. Third, a question generator uses each summary, again conditioned on type and order, to produce one question. This decomposition is intended to support questions that connect events across a narrative rather than relying on a preselected answer span.

  2. Knowl 2 — Joint prediction of question-type distributions and question counts

    equation

    For an input paragraph, a fine-tuned BERT classifier maps its mm-dimensional classification-token representation hc∈Rmh_c\in\mathbb{R}^m to probabilities over TT question types: q^=softmax⁡(Whc+b)\hat q=\operatorname{softmax}(Wh_c+b), where W∈RT×mW\in\mathbb{R}^{T\times m} and b∈RTb\in\mathbb{R}^T are learned parameters. For each of NN training examples, let qi(j)q_i^{(j)} be the target probability for type ii and let q^i(j)\hat q_i^{(j)} be the predicted probability. Training combines distribution prediction with classification of the most probable target type:

    L=γLKL+(1−γ)LCE,LKL=1N∑j=1N∑i=1Tqi(j)log⁡qi(j)q^i(j),LCE=−1N∑j=1N∑i=1Tyi(j)log⁡y^i(j).\mathcal{L}=\gamma\mathcal{L}_{\mathrm{KL}}+(1-\gamma)\mathcal{L}_{\mathrm{CE}},\qquad \mathcal{L}_{\mathrm{KL}}=\frac{1}{N}\sum_{j=1}^{N}\sum_{i=1}^{T}q_i^{(j)}\log\frac{q_i^{(j)}}{\hat q_i^{(j)}},\qquad \mathcal{L}_{\mathrm{CE}}=-\frac{1}{N}\sum_{j=1}^{N}\sum_{i=1}^{T}y_i^{(j)}\log\hat y_i^{(j)}.

    Here yi(j)y_i^{(j)} is 1 for the target’s maximum-probability type and 0 otherwise, and y^i(j)\hat y_i^{(j)} is the predicted class probability. The experiments set the loss weight γ\gamma to 0.7. To infer counts as well as type proportions, the count target for each type is augmented with a pseudo-count of 1 and normalized. If the predicted probabilities for the types and pseudo-label are (p1,…,pT,ppseudo)(p_1,\ldots,p_T,p_{\mathrm{pseudo}}), the estimated number of questions of type ii is ni=⌊pi/ppseudo+0.5⌋n_i=\lfloor p_i/p_{\mathrm{pseudo}}+0.5\rfloor. Thus, the type probabilities also encode an estimated number of questions per type.

  3. Knowl 3 — Event-centric summaries trained from silver question-answer statements

    model/method

    The event-centric summarizer is a fine-tuned BART model whose input consists of a question-type control token, an order token, and the story paragraph. The order token distinguishes successive questions of the same type (for example, first or second); the generated summary is intended to collect the events relevant to that type and question position. Because the dataset does not provide gold event summaries, training targets are silver statements made from annotated question-answer pairs: a rule-based procedure inserts the answer into a semantically parsed question and removes the question word. Eight statements that could not be parsed were written manually, and five low-quality statements were corrected manually. The summarizer used BART cased base.

  4. Knowl 4 — Question generation conditioned on summaries, type, and order

    model/method

    A separate fine-tuned BART model generates a question from each event-centric summary. Its input includes the summary plus question-type and question-order control tokens, and its training targets are the annotated questions. The question generator used BART cased large. Unlike span-based question generation, it does not require an answer span selected from the original paragraph; the summary supplies the event information on which the question is based.

  5. Knowl 5 — Evaluation data and experimental configuration

    experimental setup

    Experiments used FairytaleQA, which contains 278 storybooks split into 232 training, 23 validation, and 23 test books, with 10,580 question-answer pairs in total. The system selected the action, causal-relationship, and outcome-resolution question types; it excluded questions spanning multiple paragraphs and did not target prediction questions, which generally concern events absent from the story. The resulting data comprised 2,998 paragraphs (2,430 train, 290 validation, 278 test) and 6,676 question-answer pairs: 3,198 action, 2,547 causal-relationship, and 931 outcome-resolution pairs. The type predictor used BERT cased large; the summarizer and question generator used BART cased base and BART cased large, respectively. All models used batch size 1, and generation used greedy decoding.

  6. Knowl 6 — Automatic evaluation favors the decomposed system on question quality

    empirical result

    On the test set, evaluating concatenated generated questions against concatenated references with ROUGE-L gave the proposed system precision/recall/F1 of 37.50/31.54/30.58, compared with 15.76/35.89/19.73 for end-to-end BART (E2E) and 26.58/30.34/25.67 for the strongest QAG setting, QAG top2. QAG top10 had higher recall, 47.04, but lower precision and F1, 17.26 and 23.34. For BERTScore on the same concatenated-question setup, the proposed system scored 0.8862/0.8930/0.8893 in precision/recall/F1; QAG top2 scored 0.8810/0.8702/0.8754. Under the alternative evaluation setup used by the FairytaleQA work, test ROUGE-L precision/recall/F1 was 44.05/36.68/38.29 for the proposed system, versus 33.51/33.83/32.64 for QAG top2 and 30.80/36.53/31.65 for E2E. The authors report that the proposed system’s recall is lower than methods that generate more questions, while its precision and F1 are generally stronger.

  7. Knowl 7 — Human ratings favor children-appropriateness and type-distribution fit

    empirical result

    In a human evaluation of generated questions, the proposed system received a mean children-appropriateness score of 2.56 (standard deviation 1.31), significantly above QAG top2 at 2.22 (standard deviation 1.20; independent-samples tt-test, p=0.009p=0.009). Gold questions averaged 3.96 (standard deviation 1.02), leaving a substantial gap. The systems were reported as on par for validity and readability: the proposed system scored 3.19 (1.53) and 4.19 (1.53), respectively, versus QAG’s 3.27 (1.62) and 4.12 (1.33). In a separate question-type annotation, the proposed system’s distribution was closer to the ground truth than QAG’s: action/causal/outcome/vague questions made up 38%/40%/6%/17% for the proposed system, 48%/33%/19%/0% for the ground truth, and 21%/10%/52%/17% for QAG. The reported KL distance to the ground-truth type distribution was 0.28 for the proposed system and 0.60 for QAG.

  8. Knowl 8 — Ablations support type-distribution learning and shared type-conditioned models

    empirical result

    On test ROUGE-L, removing question-type distribution learning reduced precision/recall/F1 from 37.50/31.54/30.58 to 32.62/29.89/27.42. The type predictor’s test-set KL divergence from the target type distribution was 0.0089. Replacing its predicted distribution with the ground-truth distribution raised question-generation ROUGE-L to 46.48/31.96/35.77, compared with 37.50/31.54/30.58 using predicted types, indicating that better type estimates could improve the complete system. The authors also trained separate summarization and question-generation models per type: their aggregate scores were 25.71/33.08/26.27, below the shared, type-conditioned models’ 37.50/31.54/30.58. They attribute the advantage of sharing to the larger amount of training data available to the shared models.

  9. Knowl 9 — Sentence selection is weaker than event-centric summarization

    empirical result

    The authors compared downstream question-generation ROUGE-L for different summary inputs, with the alternative summarizers evaluated without question-type distribution learning for a fair comparison. The event-centric system without that module scored precision/recall/F1 of 32.62/29.89/27.42. Lead3, Last3, Random3, using all sentences (Total), and TextRank scored 25.20/30.76/24.73, 24.35/29.97/24.05, 23.75/28.88/23.07, 22.69/34.34/24.63, and 30.72/21.74/21.94, respectively. Using all sentences gave the highest recall among these alternatives but lower overall F1; the results support the authors’ claim that selecting or including sentences alone is less effective than their event-centric summaries for this question-generation setting.

  10. Knowl 10 — Summary quality is a bottleneck, and generated questions can be unfaithful

    limitation

    Generated summaries achieved ROUGE-L precision/recall/F1 of only 15.41/30.60/18.85 against the silver summaries, indicating substantial room for improvement. When the question generator was given silver summaries instead, its generated-question ROUGE-L rose to 92.71/85.65/87.67, suggesting that question generation was comparatively effective when supplied with suitable event content and that obtaining good summaries was a central bottleneck. The authors also report that the end-to-end system sometimes generates facts inconsistent with the source story, posing a risk to children’s knowledge learning. They identify improved discourse-level summarization and structured, knowledge-grounded generation as directions for addressing these problems.

Coverage note — The paper’s illustrative generated-question examples and implementation hardware details are omitted because they do not add substantial generalizable findings beyond the methods, evaluations, and limitations captured above.

References

  1. 1.Lorin W. Anderson, David R. Krathwohl, and Benjamin Samuel Bloom. 2000. A taxonomy for learning, teaching, and assessing: A revision of bloom’s taxonomy of educational objectives. Longman.
  2. 2.Ying-Hong Chan and Yao-Chung Fan. 2019. A recurrent BERT-based model for question generation. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 154–162, Hong Kong, China. Association for Computational Linguistics.
  3. 3.Woon Sang Cho, Yizhe Zhang, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao, Mengdi Wang, and Bill Dolan. 2021. Contrastive multi-document question generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 12–30, Online. Association for Computational Linguistics.
  4. 4.Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018. Transforming question answering datasets into natural language inference datasets. arXiv:1809.02922.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017. Question generation for question answering. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 866–874, Copenhagen, Denmark. Association for Computational Linguistics.
  7. 7.Fraide A. Ganotice, Kevin Downing, Teresa Ka Ming Mak, Barbara Chan, and Wai Yip Lee. 2017. Enhancing parent-child relationship through dialogic reading. Educational Studies, 43:51 – 66.
  8. 8.Roberta Michnick Golinkoff, Erika Hoff, Meredith L. Rowe, Catherine S. Tamis-LeMonda, and Kathy Hirsh-Pasek. 2019. Language matters: Denying the existence of the 30-million-word gap has serious consequences. Child development, 90 3:985–992.
  9. 9.Jackie Greatorex and Vikas Dhawan. 2016. Analysing the cognitive demand of reading, writing and listening tests. ISEC Proceedings.
  10. 10.Shai Gretz, Yonatan Bilu, Edo Cohen-Karlik, and Noam Slonim. 2020. The workweek is the best time to start a family – a study of GPT-2 based claim generation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 528–544, Online. Association for Computational Linguistics.
  11. 11.Jaya Jayashree Jagadeesh, Prasad Pingali, and Vasudeva Varma. 2005. Sentence extraction based single document summarization. International Institute of Information Technology, Hyderabad, India, 5.
  12. 12.Tomáš Kociský, Jonathan Schwarz, Phil Blunsom, Chris ˇ Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  13. 13.Klaus Krippendorff. 2011. Computing Krippendorff’s alpha-reliability. University of Pennsylvania.
  14. 14.Kalpesh Krishna and Mohit Iyyer. 2019. Generating question-answer hierarchies. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2321–2334, Florence, Italy. Association for Computational Linguistics.
  15. 15.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  16. 16.Manling Li, Tengfei Ma, Mo Yu, Lingfei Wu, Tian Gao, Heng Ji, and Kathleen McKeown. 2021. Timeline summarization based on event graph compression via time-aware optimal transport. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6443–6456, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  18. 18.H. P. Luhn. 1958. The automatic creation of literature abstracts. IBM Journal of Research and Development, 2(2):159–165.
  19. 19.Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang, and Qun Liu. 2021. Improving unsupervised question answering via summarization-informed question generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4134–4148, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  20. 20.Rada Mihalcea and Paul Tarau. 2004. TextRank: Bringing order into text. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 404–411, Barcelona, Spain. Association for Computational Linguistics.
  21. 21.Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI’17, page 3075–3081. AAAI Press.
  22. 22.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. In CoCo@ NIPS.
  23. 23.Makbule Ozsoy, Ilyas Cicekli, and Ferda Alpaslan. 2010. Text summarization of Turkish texts using latent semantic analysis. In Proceedings of the 23rd International Conference on Computational Linguistics (COLING 2010), pages 869–876, Beijing, China. COLING 2010 Organizing Committee.
  24. 24.Liangming Pan, Yuxi Xie, Yansong Feng, Tat-Seng Chua, and Min-Yen Kan. 2020. Semantic graphs for generating deep questions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1463–1475, Online. Association for Computational Linguistics.
  25. 25.Alison H. Paris and Scott G. Paris. 2003. Assessing narrative comprehension in young children. Reading Research Quarterly, 38(1):36–76.
  26. 26.Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, and Ido Dagan. 2021. Asking it all: Generating contextualized questions for any semantic role. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1429–1441, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  27. 27.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  28. 28.Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 379–389, Lisbon, Portugal. Association for Computational Linguistics.
  29. 29.Susan Sim and Donna Berthelsen. 2014. Shared book reading by parents with young children: Evidence-based practice. Australasian Journal of Early Childhood, 39(1):50–55.
  30. 30.Luu Anh Tuan, Darsh Shah, and Regina Barzilay. 2020. Capturing greater context for question generation. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):9065–9072.
  31. 31.Danqing Wang, Pengfei Liu, Yining Zheng, Xipeng Qiu, and Xuanjing Huang. 2020. Heterogeneous graph neural networks for extractive document summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6209–6219, Online. Association for Computational Linguistics.
  32. 32.Philip H. Winne. 1979. Experiments relating teachers’ use of higher cognitive questions to student achievement. Review of Educational Research, 49(1):13–49.
  33. 33.Lingfei Wu, Yu Chen, Kai Shen, Xiaojie Guo, Hanning Gao, Shucheng Li, Jian Pei, and Bo Long. 2021. Graph neural networks for natural language processing: A survey. arXiv:2106.06090.
  34. 34.Yuxi Xie, Liangming Pan, Dongzhe Wang, Min-Yen Kan, and Yansong Feng. 2020. Exploring question-specific rewards for generating deep questions. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2534–2546, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  35. 35.Jiacheng Xu, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Discourse-aware neural extractive text summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5021–5031, Online. Association for Computational Linguistics.
  36. 36.Ying Xu, Dakuo Wang, Penelope Collins, Hyelim Lee, and Mark Warschauer. 2021. Same benefits, different communication patterns: Comparing children’s reading with a conversational agent vs. a human partner. Computers & Education, 161:104059.
  37. 37.Bingsheng Yao, Dakuo Wang, Tongshuang Wu, Tran Hoang, Branda Sun, Toby Jia-Jun Li, Mo Yu, and Ying Xu. 2021. It is AI’s turn to ask human a question: Question and answer pair generation for children storybooks in FairytaleQA dataset. arXiv:2109.03423.
  38. 38.Andrea A. Zevenbergen and Grover J. Whitehurst. 2003. Dialogic reading: A shared picture book reading intervention for preschoolers. Lawrence Erlbaum Associates Publishers.
  39. 39.Tianyi Zhang, Varsha Kishore*, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020a. BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
  40. 40.Yuxiang Zhang, Jiamei Fu, Dongyu She, Ying Zhang, Senzhang Wang, and Jufeng Yang. 2018. Text emotion distribution learning via multi-task convolutional neural network. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4595–4601. International Joint Conferences on Artificial Intelligence Organization.
  41. 41.Zhuosheng Zhang, Hai Zhao, and Rui Wang. 2020b. Machine reading comprehension: The role of contextualized language models and beyond. ArXiv, abs/2005.06249.

Citation

MLA
Zhao, Z., et al. “Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 5073–85, https://doi.org/10.18653/v1/2022.acl-long.348.
APA
Zhao, Z., Hou, Y., Wang, D., Yu, M., Liu, C., & Ma, X. (2022). Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5073–5085. https://doi.org/10.18653/v1/2022.acl-long.348
Chicago
Zhao, Z., Y. Hou, D. Wang, M. Yu, C. Liu, and X. Ma. 2022. “Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5073–85. https://doi.org/10.18653/v1/2022.acl-long.348.
Harvard
Zhao, Z. et al. (2022) “Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5073–5085. Available at: https://doi.org/10.18653/v1/2022.acl-long.348.
Vancouver
1. Zhao Z, Hou Y, Wang D, Yu M, Liu C, Ma X (2022) Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5073–5085

BibTeX

@inproceedings{zhao-etal-2022-educational,
    title = "Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric Summarization",
    author = "Zhao, Zhenjie  and
      Hou, Yufang  and
      Wang, Dakuo  and
      Yu, Mo  and
      Liu, Chengzhong  and
      Ma, Xiaojuan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.348/",
    doi = "10.18653/v1/2022.acl-long.348",
    pages = "5073--5085"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/