Built independently by an author, for readers. Read the story and support ChapterPal

keyword

SQuAD-style reading comprehension

SQuAD-style reading comprehension is an extractive natural language processing task in which an automated system is presented with a reference passage and a question, and must identify the correct answer as a contiguous span of text directly from the passage. Modeled after the Stanford Question Answering Dataset, this format frames question answering as a boundary prediction problem where the model selects the starting and ending word indices of the answer within the text. Because the correct response is extracted verbatim from the input context, this approach primarily tests syntactic matching, semantic alignment, and local context understanding, distinguishing it from generative question answering, multiple-choice selection, or tasks that require discrete operations such as mathematical calculation or cross-paragraph reasoning.

1 item

DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, Matt Gardner

OrganizationsAllen Institute for AIPeking UniversityUniversity of California, IrvineUniversity of Washington

Why you should read this

Introduces DROP, a 96,000-question reading comprehension benchmark that requires models to perform discrete operations like addition, counting, and sorting over text, exposing a massive performance gap between state-of-the-art NLP systems and human experts.

Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new English reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 96k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs than what was necessary for prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literature on this dataset and show that the best systems only achieve 32.7% F1 on our generalized accuracy metric, while expert human performance is 96.0%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 47.0% F1.

Added

2026-09-25