Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension

Ying XuDakuo WangMo YuDaniel RitchieBingsheng YaoTongshuang WuZheng ZhangToby Jia-Jun LiNora BradfordBranda Sun

article2022ACL130 citations

Introduces FairytaleQA, an expert-annotated benchmark of over 10,000 question-answer pairs mapped to narrative elements, enabling precise evaluation and generation of educational reading comprehension questions for children and language models.

Listen

Assessing and training narrative reading comprehension in young learners requires high-quality, targeted questions, yet existing datasets are often built by crowdsourced workers without educational frameworks and fail to distinguish specific reading sub-skills. The article introduces FairytaleQA, an open-source dataset designed to evaluate and support narrative comprehension for kindergarten through eighth-grade students. Developed by educational experts using evidence-based literacy frameworks, the dataset contains 10,580 question-answer pairs derived from 278 classic fairytales, categorizing questions across seven narrative elements and classifying them as either explicit or implicit.

Benchmarking experiments demonstrate that models fine-tuned on FairytaleQA outperform models trained on general narrative benchmarks in both answering and generating questions. In question answering, a model fine-tuned on FairytaleQA achieved a score of 0.536 compared to 0.492 for a model trained on general narrative data. Decomposed evaluations revealed substantial performance gains in identifying settings and character feelings, improving by more than 10% over prior benchmarks. However, a significant gap remains between artificial intelligence systems and humans in higher-level reasoning; human performance exceeded model performance by 15% to 20% on causal relationships, outcome resolution, and outcome predictions. In question generation tasks, models trained on FairytaleQA generated more diverse, evidence-based, and factually accurate questions, closely mimicking human expert question distributions.

These findings indicate that incorporating expert educational theory into dataset construction significantly improves the capability of artificial intelligence to generate and evaluate learning materials. In educational technology, using specialized datasets reduces the risk of generating misleading or factually incorrect questions and enables granular tracking of specific student sub-skills rather than relying on a single overall score. While human performance baselines in the study represent conservative cross-estimates and models still struggle with deep narrative plotting, the dataset provides a reliable foundation. Future initiatives should focus on improving model reasoning architectures, collecting broader human evaluation data, and analyzing cultural representations and biases within narrative corpora.

No sufficiently relevant recommendations were found.

Cover for Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative Comprehension

Abstract

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity of high-quality QA datasets carefully designed to serve this purpose. In particular, existing datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements. Drawing on the reading education research, we introduce FairytaleQA¹, a dataset focusing on narrative comprehension of kindergarten to eighth-grade students. Generated by educational experts based on an evidence-based theoretical framework, FairytaleQA consists of 10,580 explicit and implicit questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations. Our dataset is valuable in two folds: First, we ran existing QA models on our dataset and confirmed that this annotation helps assess models’ fine-grained learning skills. Second, the dataset supports question generation (QG) task in the education domain. Through benchmarking with QG models, we show that the QG model trained on FairytaleQA is capable of asking high-quality and more diverse questions.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 QA Datasets Focusing on Narratives
  • 2.2 QA Datasets for Reading Education
  • 2.3 Non-QA Datasets for Narrative Comprehension
  • 3 Fairytale QA
  • 3.1 Source Texts
  • 3.2 Schema for Question Annotation
  • 3.3 Annotation Process
  • 3.4 Statistics of Fairytale QA
  • 4 Baseline Benchmark: Question Answering
  • 4.1 Question Answering Task and Model
  • 4.2 Main Results
  • 4.3 Analysis
  • 5 Baseline Benchmark: Question Generation
  • 5.1 Question Generation Task and Model
  • 5.2 Results and Analysis
  • 6 Conclusion and Future work
  • Acknowledgements
  • References
  • A Decomposed QA results on 7 narrative elements for val / test splits
  • B Decomposed QA results on explicit / implicit question types for val / test splits
  • C QG examples by benchmark models on event-based answers
  • D Example questions by category in Fairytale QA
  • E Fine-tuning Parameters
  • F Proportion of Each Question Type

Knowls

  1. Knowl 1 — FairytaleQA is a full-story narrative-comprehension dataset for school-age readers

    model/method

    FairytaleQA contains 10,580 open-ended question–answer pairs based on 278 children-friendly fairytales, designed to assess narrative comprehension for students from kindergarten through eighth grade. The stories were collected from Project Gutenberg by searching for fairytales and selecting popular titles by download count. The team made minor edits to outdated vocabulary and unconventional punctuation, excluded texts rated at or above tenth-grade reading difficulty by the textstat package, and divided the remaining stories into semantically coherent sections of 100–300 words at natural story breaks. The resulting sections averaged about 150 words. Questions are associated with the story sections they draw on, and some questions may require information from more than one section.

  2. Knowl 2 — Question labels represent seven narrative skills and an independent explicitness distinction

    definition

    FairytaleQA labels each question for one of seven narrative elements or relations: character questions identify characters or describe their traits; setting questions ask when or where events occur; action questions concern characters’ behavior or information about that behavior; feeling questions ask about a character’s emotional state or reaction; causal relationship questions ask about an earlier event that caused another event; outcome resolution questions ask what event followed from a specified prior event; and prediction questions ask for an unknown but textually predictable outcome. Independently, each question is marked explicit if its answer can be found directly as a text span, or implicit if answering requires reformulation, summarization, or inference beyond a directly stated answer. These labels are intended to support coverage and analysis of reading sub-skills, not to serve as question-classification targets.

  3. Knowl 3 — Expert-led annotation and review produced varied, answerable open-ended questions

    model/method

    Five annotators with bachelor’s degrees in education, psychology, or cognitive science and experience in teaching or reading assessment created FairytaleQA under the supervision of three literacy-education experts. They were trained for two weeks, practiced on the same five stories, discussed disagreements, and met weekly during annotation. Annotators were asked to write natural open-ended questions for elementary- or middle-school readers, covering all seven narrative categories and both explicit and implicit questions; yes/no questions were excluded. They supplied the shortest possible answers: explicit answers were extracted as the shortest relevant text spans, while implicit questions received at least two possible free-form answers. Annotators also recorded the relevant story section or sections. The project aimed for a broad average of two to three questions per section, without fixing a question quota for each story. Two annotators cross-checked every question–answer pair, and expert supervisors additionally checked 10%. For the 46 evaluation stories, an independent annotator also answered every question to provide a second reference answer; all questions were judged answerable.

  4. Knowl 4 — Dataset composition emphasizes actions and causal relations, with more explicit than implicit questions

    data/table

    FairytaleQA was randomly split by story into 232 training stories (8,548 question–answer pairs), 23 validation stories (1,025 pairs), and 23 test stories (1,007 pairs). Across the 10,580 pairs, action and causal-relationship questions are the most common categories, while setting and prediction questions are least common. About three quarters of the questions are explicit. The complete category breakdown is:

    Category Count Percentage (%)
    Character 1172 11.08
    Causal relationship 2940 27.79
    Action 3342 31.59
    Setting 630 5.95
    Feeling 1024 9.68
    Prediction 486 4.59
    Outcome resolution 986 9.32
    Explicit 7880 74.48
    Implicit 2700 25.52

    Across the dataset, stories average 14.7 sections and 2,196.7 tokens; sections average 149.1 tokens. Stories contain an average of 38.1 questions, or 2.9 per section. Questions average 10.3 tokens and answers 7.2 tokens.

  5. Knowl 5 — A small student study supports the reliability and validity of the questions

    empirical result

    The authors tested 11 questions written for one story with 120 pre-kindergarten and kindergarten students in an institutionally approved study. The resulting story-comprehension assessment had Cronbach’s coefficient alpha of 0.83, which the authors interpret as high internal reliability. Children’s performance on the questions correlated 0.76 with a separate validated language assessment (p<.001p < .001), which the authors interpret as strong external validity. This evidence applies to the sampled questions and student group, rather than constituting a large-scale validation of every item in FairytaleQA.

  6. Knowl 6 — FairytaleQA fine-tuning substantially improves aggregate question-answering scores

    empirical result

    For question answering, the study evaluated generated answers with ROUGE-L F1, comparing each prediction with both reference answers and retaining the higher score. Fine-tuned BART used a learning rate of 5×10−65\times10^{-6}, batch size 1, and one training epoch. The validation/test results show that BART fine-tuned on FairytaleQA outperformed BART fine-tuned on NarrativeQA, while a human estimate obtained by scoring one annotated answer against the other remained higher. The human estimate is likely conservative because it is based on only two references.

    Model Validation ROUGE-L F1 Test ROUGE-L F1
    Pre-trained BERT 0.104 0.097
    Pre-trained DistilBERT 0.097 0.082
    Pre-trained BART 0.108 0.088
    BART fine-tuned on NarrativeQA 0.475 0.492
    BART fine-tuned on FairytaleQA 0.533 0.536
    Human cross-estimate 0.651 0.644

    The FairytaleQA-trained BART model exceeds the NarrativeQA-trained model by 0.058 ROUGE-L F1 on validation and 0.044 on test. Its scores remain below the reported human cross-estimates. In a separate training-size diagnostic, validation performance flattened after roughly 6,000 FairytaleQA training pairs.

  7. Knowl 7 — Question-answering performance varies by narrative skill and answer explicitness

    empirical result

    The authors decomposed ROUGE-L F1 by narrative category and by whether answers were explicit or implicit. The table reports validation and test scores for BART trained on NarrativeQA, BART trained on FairytaleQA, and the human cross-estimate. FairytaleQA training improves several categories, especially validation setting and feeling questions, but does not improve every category on every split. Implicit questions are much harder than explicit ones for both models and the human estimate.

    BART-NarrativeQA BART-FairytaleQA Human
    Category Validation / Test Validation / Test Validation / Test
    Character 0.650 / 0.691 0.720 / 0.757 0.804 / 0.864
    Causal relationship 0.417 / 0.447 0.422 / 0.432 0.570 / 0.589
    Action 0.560 / 0.559 0.601 / 0.608 0.716 / 0.710
    Setting 0.618 / 0.683 0.757 / 0.696 0.833 / 0.755
    Feeling 0.231 / 0.301 0.517 / 0.508 0.453 / 0.533
    Prediction 0.298 / 0.275 0.377 / 0.300 0.605 / 0.366
    Outcome resolution 0.425 / 0.409 0.423 / 0.486 0.645 / 0.574
    Model and question type Validation ROUGE-L F1 Test ROUGE-L F1
    BART-NarrativeQA, implicit 0.280 0.278
    BART-NarrativeQA, explicit 0.548 0.563
    BART-FairytaleQA, implicit 0.304 0.286
    BART-FairytaleQA, explicit 0.619 0.620
    Human, implicit 0.363 0.330
    Human, explicit 0.760 0.750

    The explicit-question scores are consistently substantially higher than the implicit-question scores. The authors suggest that the latter require higher-level summarization and inference, whereas explicit answers can be located directly in the text. On validation feeling questions, FairytaleQA-trained BART scores 0.517, above the human cross-estimate of 0.453; on test feeling questions it scores 0.508, below the human score of 0.533.

  8. Knowl 8 — FairytaleQA improves question-generation scores and produces a more balanced question-word profile

    empirical result

    For question generation, BART was conditioned on a story section and a human-labeled answer, and the generated question was evaluated against the reference question using ROUGE-L F1. Fine-tuning used a learning rate of 5×10−65\times10^{-6}, batch size 1, and three epochs. FairytaleQA-only training scores higher than NarrativeQA-only training on both splits; adding NarrativeQA to FairytaleQA yields lower scores than FairytaleQA alone. The validation-set counts by question-initial word show that FairytaleQA-trained outputs also more closely match the reference distribution, particularly for who, why, and how questions.

    Model Validation ROUGE-L F1 Test ROUGE-L F1
    BART fine-tuned on NarrativeQA 0.424 0.442
    BART fine-tuned on FairytaleQA 0.527 0.527
    BART fine-tuned on NarrativeQA and FairytaleQA 0.508 0.519
    Question-initial word Ground truth BART-NarrativeQA BART-FairytaleQA
    Who 84 62 97
    What 426 716 447
    Why 287 144 304
    How 178 59 129
    Where 44 35 47
    Other 6 9 1

    Each column in the question-word comparison contains 1,025 validation questions. The authors use the closer distribution as evidence that FairytaleQA helps a generator reproduce a wider range of expert question forms, rather than concentrating so heavily on what questions.

  9. Knowl 9 — FairytaleQA-trained question generation is more context-specific in qualitative examples

    empirical result

    In qualitative comparisons, BART trained on NarrativeQA often produced vague questions or questions that did not accurately reflect the supplied answer and story context. BART trained on FairytaleQA more often generated questions tied to the relevant narrative evidence and semantically aligned with the reference question. The paper also notes cases where NarrativeQA-trained questions appeared more grammatical but were factually inaccurate. The authors suggest that a relevant difference in training data is that FairytaleQA annotators read complete stories, whereas NarrativeQA questions were created from story abstracts; they present this as a possible explanation, not as an experimentally isolated cause.

  10. Knowl 10 — Human QA scores and broader dataset effects remain incompletely measured

    limitation

    The reported human question-answering scores are cross-estimates between two annotated answers, not performance from a large human study, and the authors state that this procedure underestimates human performance. The second reference answer was collected for the 46 evaluation stories, so the human comparison rests on limited reference coverage. The student validation likewise tested only 11 questions from one story. The authors identify a larger human-answer collection as future work and also propose studying social stereotypes and biases in the children’s story corpus; FairytaleQA is not itself reported as having undergone such a bias audit.

Coverage note — The main dataset, annotation, validation, and QA/QG benchmark contributions are covered. Individual coder-wise distributions and supplementary example pairs are omitted because they are descriptive illustrations rather than distinct findings.

References

  1. 1.Julie Alonzo, Deni Basaraba, Gerald Tindal, and Ronald S Carriveau. 2009. They read, but how well do they understand? an empirical look at the nuances of measuring reading comprehension. Assessment for Effective Intervention, 35(1):34–44.
  2. 2.Jacopo Amidei, Paul Piwek, and Alistair Willis. 2018. Rethinking the agreement in human evaluation tasks. In Proceedings of the 27th International Conference on Computational Linguistics, pages 3318–3329, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  3. 3.Ondrej Bajgar, Rudolf Kadlec, and Jan Kleindienst. 2016. Embracing data abundance: Booktest dataset for reading comprehension. arXiv preprint arXiv:1610.00956.
  4. 4.Faeze Brahman, Meng Huang, Oyvind Tafjord, Chao Zhao, Mrinmaya Sachan, and Snigdha Chaturvedi. 2021. " let your characters tell their story": A dataset for character-centric narrative understanding. arXiv preprint arXiv:2109.05438.
  5. 5.Bidyut Das, Mukta Majumder, Santanu Phadikar, and Arif Ahmed Sekh. 2021. Automatic question generation and answer assessment: a survey. Research and Practice in Technology Enhanced Learning, 16(1):1–15.
  6. 6.Carolyn A Denton, Mischa Enos, Mary J York, David J Francis, Marcia A Barnes, Paulina A Kulesz, Jack M Fletcher, and Suzanne Carter. 2015. Text-processing differences in adolescent adequate and poor comprehenders reading accessible and challenging narrative and informational text. Reading Research Quarterly, 50(4):393–416.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  8. 8.David J Francis, Jack M Fletcher, Hugh W Catts, and J Bruce Tomblin. 2005. Dimensions affecting the assessment of reading comprehension. In Children’s reading comprehension and assessment, pages 387–412. Routledge.
  9. 9.Anna S Gellert and Carsten Elbro. 2013. Cloze tests may be quick, but are they dirty? development and preliminary validation of a cloze test of reading comprehension. Journal of Psychoeducational Assessment, 31(1):16–28.
  10. 10.Peter Goldie. 2003. One’s remembered past: Narrative thinking, emotion, and the external perspective. Philosophical Papers, 32(3):301–319.
  11. 11.Martha H Head, John E Readence, and Ray R Buss. 1989. An examination of summary writing as a measure of reading comprehension. Literacy Research and Instruction, 28(4):1–11.
  12. 12.Young-Suk Grace Kim. 2017. Why the simple view of reading is not simplistic: Unpacking component skills of reading using a direct and indirect effect model of reading (dier). Scientific Studies of Reading, 21(4):310–333.
  13. 13.Jindrich Klufa. 2015. Multiple choice question tests–advantages and disadvantages. In 3rd International Conference on Education and Modern Educational Technologies (EMET), pages 39–42.
  14. 14.Tomáš Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  15. 15.Wojciech Kryscinski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. 2021. Booksum: A collection of datasets for long-form narrative summarization. arXiv preprint arXiv:2105.08209.
  16. 16.Ghader Kurdi, Jared Leo, Bijan Parsia, Uli Sattler, and Salam Al-Emari. 2020. A systematic review of automatic question generation for educational purposes. International Journal of Artificial Intelligence in Education, 30(1):121–204.
  17. 17.Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, and Kathleen McKeown. 2020. Exploring content selection in summarization of novel chapters. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5043–5054.
  18. 18.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683.
  19. 19.Yash Kumar Lal, Nathanael Chambers, Raymond Mooney, and Niranjan Balasubramanian. 2021. Tellmewhy: A dataset for answering why-questions in narratives. arXiv preprint arXiv:2106.06132.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  21. 21.Meghan D Liebfreund. 2021. Cognitive and motivational predictors of narrative and informational text comprehension. Reading Psychology, 42(2):177–196.
  22. 22.Julie S Lynch, Paul Van Den Broek, Kathleen E Kremer, Panayiota Kendeou, Mary Jane White, and Elizabeth P Lorch. 2008. The development of narrative comprehension and its relation to other early reading skills. Reading Psychology, 29(4):327–365.
  23. 23.Nancy A Martin and Rick Brownell. 2011. Expressive one-word picture vocabulary test-4 (EOWPVT-4). Academic Therapy Publications.
  24. 24.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and evaluation framework for deeper understanding of commonsense stories. arXiv preprint arXiv:1604.01696.
  25. 25.Xiangyang Mou, Chenghao Yang, Mo Yu, Bingsheng Yao, Xiaoxiao Guo, Saloni Potdar, and Hui Su. 2021. Narrative question answering with cutting-edge open-domain qa techniques: A comprehensive study. arXiv preprint arXiv:2106.03826.
  26. 26.Ezgi Çetinkaya Özdemir and Hayati Akyol. 2019. The development of a reading comprehension test. Universal Journal of Educational Research, 7(2):563–570.
  27. 27.Alison H Paris and Scott G Paris. 2003. Assessing narrative comprehension in young children. Reading Research Quarterly, 38(1):36–76.
  28. 28.Taffy E Raphael. 1986. Teaching question answer relationships, revisited. The reading teacher, 39(6):516–522.
  29. 29.Paula Roberts and Helena Priest. 2006. Reliability and validity in research. Nursing standard, 20(44):41–46.
  30. 30.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  31. 31.Matthew Sims, Jong Ho Park, and David Bamman. 2019. Literary event detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3623–3634.
  32. 32.Qizhe Xie, Guokun Lai, Zihang Dai, and Eduard Hovy. 2017. Large-scale cloze test dataset created by teachers. arXiv preprint arXiv:1711.03225.
  33. 33.Ying Xu, Dakuo Wang, Penelope Collins, Hyelim Lee, and Mark Warschauer. 2021. Same benefits, different communication patterns: Comparing children’s reading with a conversational agent vs. a human partner. Computers & Education, 161:104059.
  34. 34.Bingsheng Yao, Dakuo Wang, Tongshuang Wu, Zheng Zhang, Toby Jia-Jun Li, Mo Yu, and Ying Xu. 2022. It is ai’s turn to ask humans a question: Question-answer pair generation for children’s story books. Association for Computational Linguistics.
  35. 35.Zheng Zhang, Ying Xu, Yanhao Wang, Bingsheng Yao, Daniel Ritchie, Tongshuang Wu, Mo Yu, Dakuo Wang, and Toby Jia-Jun Li. 2022. Storybuddy: A human-ai collaborative agent for parent-child interactive storytelling with flexible parent involvement. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22. ACM.
  36. 36.Zhenjie Zhao, Yufang Hou, Dakuo Wang, Mo Yu, Chengzhong Liu, and Xiaojuan Ma. 2022. Educational question generation of children storybooks via question type distribution learning and event-centric summarization. Association for Computational Linguistics.
  37. 37.Tricia A Zucker, Laura M Justice, Shayne B Piasta, and Joan N Kaderavek. 2010. Preschool teachers’ literal and inferential questions and children’s responses during whole-class shared reading. Early Childhood Research Quarterly, 25(1):65–83.

Citation

MLA
Xu, Y., et al. “Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 447–60, https://doi.org/10.18653/v1/2022.acl-long.34.
APA
Xu, Y., Wang, D., Yu, M., Ritchie, D., Yao, B., Wu, T., Zhang, Z., Li, T. J.-J., Bradford, N., Sun, B., Hoang, T. B., Sang, Y., Hou, Y., Ma, X., Yang, D., Peng, N., Yu, Z., & Warschauer, M. (2022). Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 447–460. https://doi.org/10.18653/v1/2022.acl-long.34
Chicago
Xu, Y., D. Wang, M. Yu, et al. 2022. “Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 447–60. https://doi.org/10.18653/v1/2022.acl-long.34.
Harvard
Xu, Y. et al. (2022) “Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 447–460. Available at: https://doi.org/10.18653/v1/2022.acl-long.34.
Vancouver
1. Xu Y, Wang D, Yu M, et al (2022) Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 447–460

BibTeX

@inproceedings{xu-etal-2022-fantastic,
    title = "Fantastic Questions and Where to Find Them: {F}airytale{QA} {--} An Authentic Dataset for Narrative Comprehension",
    author = "Xu, Ying  and
      Wang, Dakuo  and
      Yu, Mo  and
      Ritchie, Daniel  and
      Yao, Bingsheng  and
      Wu, Tongshuang  and
      Zhang, Zheng  and
      Li, Toby Jia-Jun  and
      Bradford, Nora  and
      Sun, Branda  and
      Hoang, Tran Bao  and
      Sang, Yisi  and
      Hou, Yufang  and
      Ma, Xiaojuan  and
      Yang, Diyi  and
      Peng, Nanyun  and
      Yu, Zhou  and
      Warschauer, Mark",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.34/",
    doi = "10.18653/v1/2022.acl-long.34",
    pages = "447--460"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/