DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
Dheeru DuaYizhong WangPradeep DasigiGabriel StanovskySameer SinghMatt Gardner
Introduces DROP, a 96,000-question reading comprehension benchmark that requires models to perform discrete operations like addition, counting, and sorting over text, exposing a massive performance gap between state-of-the-art NLP systems and human experts.
Recent advances in artificial intelligence have produced systems that match human performance on standard reading comprehension benchmarks. However, existing benchmarks primarily test pattern matching and shallow span extraction rather than a deeper semantic understanding of text. The article introduces and evaluates DROP, a benchmark designed to test whether machine reading systems can perform discrete reasoning—such as addition, subtraction, sorting, and counting—over the content of English paragraphs.
To construct this benchmark, the authors collected approximately 7,000 Wikipedia passages and crowdsourced 96,567 question-answer pairs using an adversarial setup. In this setup, crowd workers were required to submit questions that a baseline reading comprehension system could not solve. The authors then evaluated three categories of systems: heuristic baselines to test for dataset artifacts, semantic parsers operating over structured sentence representations, and standard reading comprehension models including BERT. Additionally, the authors developed NAQANet, a new architecture combining standard neural reading comprehension with modules for counting, arithmetic over numbers, and extracting answers from questions.
State-of-the-art reading comprehension models experienced a severe drop in performance on the new benchmark. BERT, which achieves nearly 85% Exact-Match accuracy on standard benchmarks, scored only 32.7% F1 accuracy on the DROP test set, compared to an expert human performance level of 96.4%. Semantic parsing baselines performed poorly, reaching between 10.8% and 11.5% F1 due to challenges in extracting structured tables and training on spurious logical forms. The authors' NAQANet model achieved the highest performance among the evaluated systems at 47.0% F1, driven by its capability to handle numerical arithmetic and counting.
These findings demonstrate that current commercial-grade language systems lack the capability to perform compositional and mathematical reasoning across text, introducing substantial reliability risks when deploying such systems for complex analytical or data-driven workflows. To bridge this capability gap, future system development should focus on hybrid models that integrate symbolic and numerical reasoning into neural architectures, such as combining numerical execution modules with advanced pre-trained models like BERT.
The findings are supported by strong inter-annotator agreement and rigorous adversarial testing. However, current limitations include the fact that NAQANet only supports a restricted set of arithmetic operations (specifically counting from zero to nine and addition or subtraction of two numbers) and that passages are predominantly concentrated in specific domains such as sports summaries and historical events.
- Paper: SQuAD: 100,000+ Questions for Machine Comprehension of Text, Pranav Rajpurkar et al. (2016). It introduces the foundational span-extraction reading comprehension paradigm that DROP explicitly critiques and builds upon by introducing discrete numerical reasoning operations.
- Paper: Adversarial Examples for Evaluating Reading Comprehension Systems, Robin Jia et al. (2017). It demonstrates the brittleness and superficial heuristic reliance of standard reading comprehension systems on SQuAD, providing the direct motivation for DROP's adversarially created discrete reasoning benchmark.
- Paper: Bidirectional Attention Flow for Machine Comprehension, Minjoon Seo et al. (2016). It defines the standard bidirectional attention architecture for extractive comprehension that serves as a core baseline and architectural component for reading systems analyzed in DROP.
- Paper: Know What You Don't Know: Unanswerable Questions for SQuAD, Pranav Rajpurkar et al. (2018). It highlights key limitations of extractive QA datasets by incorporating unanswerable questions, establishing the need for more comprehensive paragraph understanding and adversarial data design.
- Paper: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering, Zhilin Yang et al. (2018). It pioneered multi-hop reasoning over paragraphs, preceding DROP's shift toward multi-reference resolution and compositional operations across context.
- Paper: RACE: Large-scale ReAding Comprehension Dataset From Examinations, Guokun Lai et al. (2017). It provides foundational evidence that machine reading models fail when benchmark questions require complex multi-step reasoning rather than simple span matching.
- Paper: Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, Peter Clark et al. (2018). It established early adversarial filtering techniques to isolate questions that defeat standard retrieval and superficial matching solvers.
- Paper: TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension, Mandar Joshi et al. (2017). It established challenge settings for reading comprehension that expose the gap between simple passage matching and genuine document understanding.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). It explores resolving contextual and conversational references across paragraphs, a prerequisite challenge directly incorporated into DROP's discrete reasoning tasks.
- Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). It introduced early neural reading architectures connecting question answering directly with paragraph retrieval and span prediction.
- Paper: SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, Alex Wang et al. (2019). It broadens the evaluation of natural language understanding by aggregating complex reasoning and multi-sentence comprehension benchmarks to follow up on saturation in simpler datasets.
- Paper: Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps, Xanh Ho et al. (2020). It builds on multi-hop and compositional QA benchmarks like DROP by enforcing explicit, verifiable multi-step reasoning chains.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). It investigates why language models struggle to compose multi-step reasoning steps and demonstrates prompting methods to overcome compositional gaps exposed by benchmarks like DROP.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). It extends discrete and mathematical reasoning evaluation from reading comprehension passages to competition-level multi-step mathematical problem solving.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). It evaluates how chain-of-thought prompting enables language models to overcome failures on complex, multi-step algorithmic and logical tasks.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). It scales multitask reasoning and knowledge assessment across diverse academic and professional subjects beyond single-paragraph comprehension.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). It expands the assessment of language model capabilities to an extensive suite of complex reasoning and multi-step tasks across diverse domains.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). It addresses multi-step reasoning failures by introducing process-level supervision to verify intermediate logical steps during complex reasoning.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). It generalizes discrete mathematical and arithmetic reasoning over text paragraphs to rich multimodal and visual contexts.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). It constructs a more demanding reasoning benchmark designed to rigorously assess advanced multi-step reasoning in modern foundation models.
