Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension
Ying XuDakuo WangMo YuDaniel RitchieBingsheng YaoTongshuang WuZheng ZhangToby Jia-Jun LiNora BradfordBranda Sun
Introduces FairytaleQA, an expert-annotated benchmark of over 10,000 question-answer pairs mapped to narrative elements, enabling precise evaluation and generation of educational reading comprehension questions for children and language models.
Assessing and training narrative reading comprehension in young learners requires high-quality, targeted questions, yet existing datasets are often built by crowdsourced workers without educational frameworks and fail to distinguish specific reading sub-skills. The article introduces FairytaleQA, an open-source dataset designed to evaluate and support narrative comprehension for kindergarten through eighth-grade students. Developed by educational experts using evidence-based literacy frameworks, the dataset contains 10,580 question-answer pairs derived from 278 classic fairytales, categorizing questions across seven narrative elements and classifying them as either explicit or implicit.
Benchmarking experiments demonstrate that models fine-tuned on FairytaleQA outperform models trained on general narrative benchmarks in both answering and generating questions. In question answering, a model fine-tuned on FairytaleQA achieved a score of 0.536 compared to 0.492 for a model trained on general narrative data. Decomposed evaluations revealed substantial performance gains in identifying settings and character feelings, improving by more than 10% over prior benchmarks. However, a significant gap remains between artificial intelligence systems and humans in higher-level reasoning; human performance exceeded model performance by 15% to 20% on causal relationships, outcome resolution, and outcome predictions. In question generation tasks, models trained on FairytaleQA generated more diverse, evidence-based, and factually accurate questions, closely mimicking human expert question distributions.
These findings indicate that incorporating expert educational theory into dataset construction significantly improves the capability of artificial intelligence to generate and evaluate learning materials. In educational technology, using specialized datasets reduces the risk of generating misleading or factually incorrect questions and enables granular tracking of specific student sub-skills rather than relying on a single overall score. While human performance baselines in the study represent conservative cross-estimates and models still struggle with deep narrative plotting, the dataset provides a reliable foundation. Future initiatives should focus on improving model reasoning architectures, collecting broader human evaluation data, and analyzing cultural representations and biases within narrative corpora.
- Paper: SQuAD: 100,000+ Questions for Machine Comprehension of Text, Pranav Rajpurkar et al. (2016). SQuAD established the large-scale machine-reading QA benchmark paradigm that FairytaleQA adapts to test finer-grained narrative comprehension.
- Paper: RACE: Large-scale ReAding Comprehension Dataset From Examinations, Guokun Lai et al. (2017). RACE’s use of authentic educator-written comprehension questions provides a useful precedent for FairytaleQA’s education-grounded question design.
No sufficiently relevant recommendations were found.
