DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Dheeru DuaYizhong WangPradeep DasigiGabriel StanovskySameer SinghMatt Gardner

article2019NAACL1,361 citations

Introduces DROP, a 96,000-question reading comprehension benchmark that requires models to perform discrete operations like addition, counting, and sorting over text, exposing a massive performance gap between state-of-the-art NLP systems and human experts.

Listen

Recent advances in artificial intelligence have produced systems that match human performance on standard reading comprehension benchmarks. However, existing benchmarks primarily test pattern matching and shallow span extraction rather than a deeper semantic understanding of text. The article introduces and evaluates DROP, a benchmark designed to test whether machine reading systems can perform discrete reasoning—such as addition, subtraction, sorting, and counting—over the content of English paragraphs.

To construct this benchmark, the authors collected approximately 7,000 Wikipedia passages and crowdsourced 96,567 question-answer pairs using an adversarial setup. In this setup, crowd workers were required to submit questions that a baseline reading comprehension system could not solve. The authors then evaluated three categories of systems: heuristic baselines to test for dataset artifacts, semantic parsers operating over structured sentence representations, and standard reading comprehension models including BERT. Additionally, the authors developed NAQANet, a new architecture combining standard neural reading comprehension with modules for counting, arithmetic over numbers, and extracting answers from questions.

State-of-the-art reading comprehension models experienced a severe drop in performance on the new benchmark. BERT, which achieves nearly 85% Exact-Match accuracy on standard benchmarks, scored only 32.7% F1 accuracy on the DROP test set, compared to an expert human performance level of 96.4%. Semantic parsing baselines performed poorly, reaching between 10.8% and 11.5% F1 due to challenges in extracting structured tables and training on spurious logical forms. The authors' NAQANet model achieved the highest performance among the evaluated systems at 47.0% F1, driven by its capability to handle numerical arithmetic and counting.

These findings demonstrate that current commercial-grade language systems lack the capability to perform compositional and mathematical reasoning across text, introducing substantial reliability risks when deploying such systems for complex analytical or data-driven workflows. To bridge this capability gap, future system development should focus on hybrid models that integrate symbolic and numerical reasoning into neural architectures, such as combining numerical execution modules with advanced pre-trained models like BERT.

The findings are supported by strong inter-annotator agreement and rigorous adversarial testing. However, current limitations include the fact that NAQANet only supports a restricted set of arithmetic operations (specifically counting from zero to nine and addition or subtraction of two numbers) and that passages are predominantly concentrated in specific domains such as sports summaries and historical events.

arXiv: 1903.00161
Cover for DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Abstract

Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new English reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 96k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs than what was necessary for prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literature on this dataset and show that the best systems only achieve 32.7% F1 on our generalized accuracy metric, while expert human performance is 96.0%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 47.0% F1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 DROP Data Collection
  • 4 DROP Data Analysis
  • 5 Baseline Systems
  • 5.1 Semantic Parsing
  • 5.2 SQuAD-style Reading Comprehension
  • 5.3 Heuristic Baselines
  • 6 NAQANet
  • 6.1 Model Description
  • 6.2 Weakly-Supervised Training
  • 7 Results and Discussion
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — DROP Reading Comprehension Benchmark

    definition

    DROP (Discrete Reasoning Over Paragraphs) is an English reading comprehension dataset comprising 96,567 question-answer pairs over 6,735 Wikipedia passages (5,565 training, 582 development, and 588 test passages). The benchmark is designed to evaluate machine reading systems on questions that necessitate discrete symbolic operations over paragraph text, including addition, subtraction, counting, comparison, sorting, and selection.

    Passages were collected from Wikipedia categories with dense narrative event sequences and high numerical frequency, primarily National Football League (NFL) game summaries, history articles, and general passages containing at least 20 numbers. The average passage length is approximately 200 words, and the average question length is roughly 11 words.

    Questions were crowdsourced on Amazon Mechanical Turk using an adversarial authoring setup: crowd workers were required to generate questions that an active baseline model (BiDAF) could not answer correctly. Answers are restricted to four formats: single or multiple text spans from the passage, text spans from the question, dates, or numbers with explicit units. Development and test questions were validated with at least two additional independent worker annotations, yielding an overall inter-annotator agreement of Cohen's κ=0.74\kappa = 0.74 (κ=0.81\kappa = 0.81 for numbers, 0.620.62 for spans, and 0.650.65 for dates).

  2. Knowl 2 — DROP Evaluation Metrics: Exact Match and Numeracy-Focused F1

    definition

    Model performance on DROP is evaluated using two metrics: Exact Match (EM) and a numeracy-focused macro-averaged F1F_1 score.

    • Exact Match (EM): Computes strict string equivalence after standard normalization (lowercasing, punctuation stripping, and removal of articles "a", "an", "the").
    • Numeracy-Focused F1F_1: Computes the token-level harmonic mean of precision and recall between the bag-of-words of the prediction and the gold reference, with an additional arithmetic constraint: if there is any numerical mismatch between the gold answer and predicted answer, F1F_1 is set strictly to 0.00.0, regardless of word overlap.
    • Multiple Spans: For answers consisting of a list of spans, predictions and gold spans are aligned greedily one-to-one based on maximum bag-of-words token overlap, and the average F1F_1 across matched spans is computed.
    • Multiple References: For questions with K≥1K \ge 1 reference annotations, the evaluation score is the maximum achieved over all gold references: Metric(ypred,{ygold(k)}k=1K)=max⁡k∈{1,…,K}Metric(ypred,ygold(k))\text{Metric}(y_{\text{pred}}, \{y_{\text{gold}}^{(k)}\}_{k=1}^K) = \max_{k \in \{1,\dots,K\}} \text{Metric}(y_{\text{pred}}, y_{\text{gold}}^{(k)})
  3. Knowl 3 — NAQANet Architecture for Numerically-Aware Reading Comprehension

    model/method

    NAQANet (Numerically-Aware QANet) is a neural reading comprehension architecture that extends QANet to produce four categories of answers: passage spans, question spans, integer counts, and arithmetic expressions (addition and subtraction of passage numbers).

    Given question representation matrix Q∈Rm×dQ \in \mathbb{R}^{m \times d} and question-aware passage representation matrix Pˉ∈Rn×d\bar{P} \in \mathbb{R}^{n \times d} produced by the standard QANet embedding and attention layers, NAQANet computes global pooled passage and question vectors: αP=softmax(WPPˉ),hP=αPPˉ∈Rd\alpha^P = \text{softmax}(W_P \bar{P}), \quad h^P = \alpha^P \bar{P} \in \mathbb{R}^d αQ=softmax(WQQ),hQ=αQQ∈Rd\alpha^Q = \text{softmax}(W_Q Q), \quad h^Q = \alpha^Q Q \in \mathbb{R}^d where WP,WQ∈R1×dW_P, W_Q \in \mathbb{R}^{1 \times d} are learnable parameters. An answer-type prediction head chooses among the four answer formats: ptype=softmax(FFN([hP;hQ]))∈R4p^{\text{type}} = \text{softmax}(\text{FFN}([h^P; h^Q])) \in \mathbb{R}^4 where FFN\text{FFN} is a two-layer feed-forward network with ReLU activation.

    The four specialized output branches operate as follows:

    1. Passage Span: Three stacked encoder blocks applied to Pˉ\bar{P} produce representations M0,M1,M2∈Rn×dM_0, M_1, M_2 \in \mathbb{R}^{n \times d}. Start and end position distributions are computed as: pp_start=softmax(FFN([M0;M1])),pp_end=softmax(FFN([M0;M2]))p^{\text{p\_start}} = \text{softmax}(\text{FFN}([M_0; M_1])), \quad p^{\text{p\_end}} = \text{softmax}(\text{FFN}([M_0; M_2]))
    2. Question Span: Question start and end distributions are conditioned on the question tokens QQ concatenated with hPh^P broadcast across all mm question tokens (e∣Q∣⊗hPe_{|Q|} \otimes h^P): pq_start=softmax(FFN([Q;e∣Q∣⊗hP])),pq_end=softmax(FFN([Q;e∣Q∣⊗hP]))p^{\text{q\_start}} = \text{softmax}(\text{FFN}([Q; e_{|Q|} \otimes h^P])), \quad p^{\text{q\_end}} = \text{softmax}(\text{FFN}([Q; e_{|Q|} \otimes h^P]))
    3. Count: Modeled as a 10-class classification task over the digits {0,1,…,9}\{0, 1, \dots, 9\}: pcount=softmax(FFN(hP))∈R10p^{\text{count}} = \text{softmax}(\text{FFN}(h^P)) \in \mathbb{R}^{10}
    4. Arithmetic Expression: A fourth encoder block applied to M2M_2 produces M3∈Rn×dM_3 \in \mathbb{R}^{n \times d}. For each extracted number ii in the passage, its feature representation hiN∈R2dh_i^N \in \mathbb{R}^{2d} is obtained by indexing into [M0;M3][M_0; M_3]. A 3-way softmax assigns a sign si∈{+1,−1,0}s_i \in \{+1, -1, 0\}: pisign=softmax(FFN(hiN))∈R3p_i^{\text{sign}} = \text{softmax}(\text{FFN}(h_i^N)) \in \mathbb{R}^3 The final numerical prediction is evaluated as ∑isi⋅vi\sum_i s_i \cdot v_i, where viv_i is the value of number ii.
  4. Knowl 4 — Weakly-Supervised Marginal Likelihood Training for NAQANet

    model/method

    Because DROP provides only the target answer string without latent execution traces (such as the answer type, operand indices, or arithmetic operations), NAQANet is trained via weak supervision by maximizing the marginal likelihood of all execution paths that evaluate to the true answer string y∗y^*.

    Let E(y∗)\mathcal{E}(y^*) denote the set of valid execution traces evaluating to y∗y^*, which includes:

    • Any passage span (i,j)(i, j) where Passage[i:j]=y∗\text{Passage}[i:j] = y^*.
    • Any question span (i′,j′)(i', j') where Question[i′:j′]=y∗\text{Question}[i':j'] = y^*.
    • Any count class c∈{0,…,9}c \in \{0, \dots, 9\} where str(c)=y∗\text{str}(c) = y^*.
    • Any sign configuration {s1,…,sK}∈{+1,−1,0}K\{s_1, \dots, s_K\} \in \{+1, -1, 0\}^K over the KK passage numbers such that ∑k=1Kskvk=float(y∗)\sum_{k=1}^K s_k v_k = \text{float}(y^*). To constrain the exponential search space and avoid noise during training, sign search is restricted to combinations of at most two non-zero numbers.

    The training objective minimizes the negative marginal log-likelihood over all valid paths: L(θ)=−log⁡∑e∈E(y∗)P(e∣Passage,Question;θ)\mathcal{L}(\theta) = -\log \sum_{e \in \mathcal{E}(y^*)} P(e \mid \text{Passage}, \text{Question}; \theta) where P(e∣⋅)P(e \mid \cdot) is the joint probability of selecting the branch type from ptypep^{\text{type}} and the specific execution path within that branch.

  5. Knowl 5 — Pipelined Semantic Parsing Baseline over Extracted Semi-Structured Tables

    model/method

    To evaluate semantic parsing on unstructured text passages, a grammar-constrained semantic parser (the KDG model) is applied to tabular representations automatically extracted from paragraphs. Three sentence representation schemes convert text into predicate-argument tables:

    1. Stanford Dependencies (Syn Dep): Captures word-level syntactic relations.
    2. Open Information Extraction (Open IE): Extracts shallow relational triples linking predicates and arguments.
    3. Semantic Role Labeling (SRL): Identifies predicate senses and assigns semantic argument roles (e.g., ARG0, ARG1, AM-TMP).

    In each generated table, rows correspond to predicate-argument structures and columns correspond to argument roles. A strongly-typed logical form language defines operations over table rows, relation headers, strings, numbers, and dates, with functions for filtering (e.g., filter_number_greater) and counting. During training, the parser maximizes the marginal likelihood of all logical forms found via exhaustive search up to a preset depth that evaluate to the correct denotation. At test time, beam search finds the highest-scoring logical form for execution.

  6. Knowl 6 — Benchmark Evaluation of Reading Comprehension and Semantic Parsing Models on DROP

    data/table

    Performance of baseline models, NAQANet ablations, and expert humans on the DROP development and test sets, evaluated with Exact Match (EM) and numeracy-focused F1F_1:

    Method Dev Test
    EM F1 EM F1
    Heuristic Baselines
    Majority 0.09 1.38 0.07 1.44
    Question-only 4.28 8.07 4.18 8.59
    Paragraph-only 0.13 2.27 0.14 2.26
    Semantic Parsing
    Syn Dep 9.38 11.64 8.51 10.84
    OpenIE 8.80 11.31 8.53 10.77
    SRL 9.28 11.72 8.98 11.45
    SQuAD-style Reading Comprehension
    BiDAF 26.06 28.85 24.75 27.49
    QANet 27.50 30.44 25.50 28.36
    QANet + ELMo 27.71 30.33 27.08 29.67
    BERT 30.10 33.36 29.45 32.70
    NAQANet
    + Question Span 25.94 29.17 24.98 28.18
    + Count 30.09 33.92 30.04 32.75
    + Add / Sub 43.07 45.71 40.40 42.96
    NAQANet Complete Model 46.20 49.24 44.07 47.01
    Human – – 94.09 96.42

    Standard reading comprehension models perform poorly on DROP (BERT achieves 32.70% test F1F_1, compared to 84.7% EM on SQuAD), lagging human performance (96.42% test F1F_1) by over 63 points. NAQANet's arithmetic modules provide the largest gain, elevating test F1F_1 from 28.18% to 47.01%.

  7. Knowl 7 — Development Performance Breakdown by Answer Type: NAQANet vs. BERT

    data/table

    Performance comparison between NAQANet and the strongest span-extraction baseline (BERT) across answer types on the DROP development set:

    Answer Type Dev Share (%) Exact Match (%) F1 (%)
    NAQANet BERT NAQANet BERT
    Date 1.57 28.7 38.7 35.5 42.8
    Numbers 61.94 44.0 14.5 44.2 14.8
    Single Span 31.71 58.2 64.6 64.6 70.1
    >> 1 Spans 4.77 0.0 0.0 17.13 25.0

    Number-type answers make up 61.94% of the development set. On these questions, BERT achieves only 14.8% F1F_1 because it cannot perform arithmetic or counting, whereas NAQANet attains 44.2% F1F_1. On single text spans (31.71% of dev), BERT outperforms NAQANet (70.1% vs. 64.6% F1F_1). Neither architecture natively outputs multiple spans, yielding 0.0% EM on multi-span answers.

  8. Knowl 8 — Information Extraction and Spuriousness Bottlenecks in Semantic Parsing for Paragraphs

    limitation

    Applying pipelined semantic parsers to unstructured text passages fails due to two primary failure modes:

    1. Information Extraction Bottleneck: Converting narrative text into tabular structures drops critical information. For the highest-performing scheme (SRL), beam search found executable logical forms for only 34% of training examples (and 25% for OpenIE). A manual analysis of 60 examples showed that only 25% of the extracted SRL tables contained the information necessary to answer the question.
    2. Spurious Logical Forms: Under weak supervision, logical forms frequently evaluate to the correct numerical denotation coincidentally without capturing the true question semantics. In a sample of 60 questions analyzed, only 8 (13.3%) contained genuine, non-spurious logical forms, leading to severe overfitting and poor test generalization.
  9. Knowl 9 — Distribution of Failure Modes in NAQANet Predictions

    empirical result

    A manual error analysis of 100 randomly sampled incorrect predictions made by NAQANet on the DROP dataset identified five prominent error categories:

    • Complex Arithmetic (51%): Inability to perform arithmetic expressions beyond simple two-number addition or subtraction, or misidentifying relevant numerical operands.
    • Complex Compositional Reasoning (40%): Inability to compose multiple distinct reasoning steps (e.g., performing entity coreference before subtraction, or filtering events before counting).
    • Counting Failures (30%): Miscounting occurrences of events or entities described across multiple sentences.
    • Domain Knowledge and Commonsense (23%): Missing world knowledge, such as football scoring mechanics (e.g., touchdown point values) or date arithmetic.
    • Coreference Errors (6%): Failing to link pronouns or nominal references to their antecedents across distinct sentences.

Coverage note — None was omitted; all key contributions—the DROP benchmark definition and protocol, evaluation metrics, NAQANet model architecture and training, semantic parsing baselines, full experimental results, answer breakdown, and error analyses—are fully captured.

References

  1. 1.Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew G Broadhead, and Oren Etzioni. 2007. Open information extraction from the web. In IJCAI.
  2. 2.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013a. Semantic parsing on freebase from question-answer pairs. In EMNLP.
  3. 3.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013b. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533–1544.
  4. 4.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In EMNLP.
  5. 5.Xavier Carreras and Lluís Màrquez. 2005. Introduction to the conll-2005 shared task: Semantic role labeling. In Proceedings of CONLL, pages 152–164.
  6. 6.Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the cnn/daily mail reading comprehension task.
  7. 7.David L Chen and Raymond J Mooney. 2011. Learning to interpret natural language navigation instructions from observations. In AAAI, volume 2, pages 1–2.
  8. 8.Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke S. Zettlemoyer. 2018. Quac: Question answering in context. In EMNLP.
  9. 9.Christopher Clark and Matt Gardner. 2018. Simple and effective multi-paragraph reading comprehension. In ACL.
  10. 10.Peter Clark, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Turney, and Daniel Khashabi. 2016. Combining retrieval, statistics, and inference to answer elementary science questions. In Thirtieth AAAI Conference on Artificial Intelligence.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, abs/1810.04805.
  12. 12.Timothy Dozat, Peng Qi, and Christopher D. Manning. 2017. Stanford’s graph-based neural dependency parser at the conll 2017 shared task. In CoNLL Shared Task.
  13. 13.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer. 2017. Allennlp: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS). Association for Computational Linguistics.
  14. 14.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proc. of NAACL.
  15. 15.Luheng He, Kenton Lee, Mike Lewis, and Luke S. Zettlemoyer. 2017. Deep semantic role labeling: What works and what’s next. In ACL.
  16. 16.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 523–533.
  17. 17.Mandar S. Joshi, Eunsol Choi, Daniel S. Weld, and Luke S. Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In ACL.
  18. 18.Divyansh Kaushik and Zachary Chase Lipton. 2018. How much reading does reading comprehension require? a critical investigation of popular benchmarks. In EMNLP.
  19. 19.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In NAACL-HLT.
  20. 20.Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. TACL, 6:317–328.
  21. 21.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. TACL, 3:585–597.
  22. 22.Jayant Krishnamurthy, Pradeep Dasigi, and Matt Gardner. 2017. Neural semantic parsing with type constraints for semi-structured tables. In EMNLP.
  23. 23.Nate Kushman, Yoav Artzi, Luke Zettlemoyer, and Regina Barzilay. 2014. Learning to automatically solve algebra word problems. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), volume 1, pages 271–281.
  24. 24.Chen Liang, Jonathan Berant, Quoc Le, Kenneth D. Forbus, and Ni Lao. 2017. Neural symbolic machines: Learning semantic parsers on freebase with weak supervision. In ACL.
  25. 25.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In ACL.
  26. 26.Marie-Catherine de Marneffe and Christopher D. Manning. 2008. The stanford typed dependencies representation. In CFCFPE@COLING.
  27. 27.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP.
  28. 28.Pasquale Minervini and Sebastian Riedel. 2018. Adversarially regularising neural nli models to integrate logical background knowledge. In CoNLL.
  29. 29.Bhavana Dalvi Mishra, Lifu Huang, Niket Tandon, Wen-tau Yih, and Peter Clark. 2018. Tracking state changes in procedural text: A challenge dataset and models for process paragraph comprehension.
  30. 30.Arvind Neelakantan, Quoc V. Le, and Ilya Sutskever. 2016. Neural programmer: Inducing latent programs with gradient descent. ICLR.
  31. 31.Simon Ostermann, Ashutosh Modi, Michael Roth, Stefan Thater, and Manfred Pinkal. 2018. Mcscript: a novel dataset for assessing machine comprehension using script knowledge. LREC Proceedings, 2018.
  32. 32.Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. 2018. emrqa: A large corpus for question answering on electronic medical records.
  33. 33.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context. ACL.
  34. 34.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In ACL.
  35. 35.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In EMNLP.
  36. 36.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL-HLT.
  37. 37.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. In ACL.
  38. 38.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In EMNLP.
  39. 39.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. Coqa: A conversational question answering challenge. TACL.
  40. 40.Scott E. Reed and Nando de Freitas. 2016. Neural programmer-interpreters. ICLR.
  41. 41.Amrita Saha, Rahul Aralikatte, Mitesh M. Khapra, and Karthik Sankaranarayanan. 2018. Duorc: Towards complex language understanding with paraphrased reading comprehension. In ACL.
  42. 42.Min Joon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2017. Bidirectional attention flow for machine comprehension. ICLR.
  43. 43.Gabriel Stanovsky, Julian Michael, Luke S. Zettlemoyer, and Ido Dagan. 2018. Supervised open information extraction. In NAACL-HLT.
  44. 44.Simon Šuster and Walter Daelemans. 2018. Clicr: a dataset of clinical case reports for machine reading comprehension.
  45. 45.Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. In NAACL-HLT.
  46. 46.Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018. Constructing datasets for multi-hop reading comprehension across documents. TACL, 6:287–302.
  47. 47.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In EMNLP.
  48. 48.Pengcheng Yin and Graham Neubig. 2017. A syntactic neural model for general-purpose code generation. In ACL’17.
  49. 49.Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le. 2018. Qanet: Combining local convolution with global self-attention for reading comprehension. ICLR.
  50. 50.John M. Zelle and Raymond J. Mooney. 1996. Learning to parse database queries using inductive logic programming. In AAAI/IAAI, Vol. 2.
  51. 51.Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. CVPR, abs/1811.10830.
  52. 52.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. Swag: A large-scale adversarial dataset for grounded commonsense inference. In EMNLP.
  53. 53.Luke S. Zettlemoyer and Michael Collins. 2005. Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars. In UAI.
  54. 54.Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2019. ReCoRD: Bridging the gap between human and machine commonsense reading comprehension.

Citation

MLA
Dua, D., et al. “DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 2368–78, https://doi.org/10.18653/v1/N19-1246.
APA
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., & Gardner, M. (2019). DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2368–2378. https://doi.org/10.18653/v1/N19-1246
Chicago
Dua, D., Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner. 2019. “DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2368–78. https://doi.org/10.18653/v1/N19-1246.
Harvard
Dua, D. et al. (2019) “DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs”, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, pp. 2368–2378. Available at: https://doi.org/10.18653/v1/N19-1246.
Vancouver
1. Dua D, Wang Y, Dasigi P, Stanovsky G, Singh S, Gardner M (2019) DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, pp 2368–2378

BibTeX

@inproceedings{dua-etal-2019-drop,
    title = "{DROP}: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs",
    author = "Dua, Dheeru  and
      Wang, Yizhong  and
      Dasigi, Pradeep  and
      Stanovsky, Gabriel  and
      Singh, Sameer  and
      Gardner, Matt",
    editor = "Burstein, Jill  and
      Doran, Christy  and
      Solorio, Thamar",
    booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)",
    month = jun,
    year = "2019",
    address = "Minneapolis, Minnesota",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/N19-1246/",
    doi = "10.18653/v1/N19-1246",
    pages = "2368--2378"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/