PubMedQA: A Dataset for Biomedical Research Question Answering

Qiao JinBhuwan DhingraZhengping LiuWilliam W. CohenXinghua Lu

article2019EMNLP1,901 citations

Introduces PubMedQA, a biomedical question answering benchmark requiring reasoning over quantitative findings in research abstracts, revealing a substantial performance gap between language models and human experts.

Listen

Biomedical natural language processing has historically been constrained by small datasets focused on simple fact retrieval rather than complex reasoning. In real-world biomedical literature, understanding study outcomes requires synthesizing quantitative findings and comparing clinical cohorts rather than extracting basic facts. To address this gap and support evidence-based decision-making, the article introduces PubMedQA, a novel biomedical question-answering dataset derived from PubMed research abstracts.

The article evaluates how effectively machine learning models can reason over biomedical study abstracts to answer research questions with "yes," "no," or "maybe." The dataset is structured across three tiers: 1,000 expert-annotated instances labeled by medical candidates, 61,200 unlabeled instances for semi-supervised learning, and 211,300 artificially generated instances created by converting statement titles into questions. Each instance includes a research question, the abstract text excluding its conclusion as the reasoning context, a long answer representing the original conclusion, and a categorical summary answer. Models were evaluated on their ability to infer the correct categorical answer using only the question and context, supported by a multi-phase fine-tuning schedule and auxiliary supervision from the long answer text.

The primary finding is that the best-performing model—a domain-specific BioBERT architecture trained via multi-phase fine-tuning with auxiliary supervision—achieved an accuracy of 68.1% and a macro-F1 score of 52.7% on the expert-labeled test set. This performance significantly exceeded the baseline majority-class guess of 55.2% accuracy. However, machine performance trailed far behind human capability, as a single human annotator achieved 78.0% accuracy under identical reasoning-required conditions and 90.4% when conclusions were provided. The analysis also revealed that 96.5% of questions required quantitative reasoning across cohorts or subgroup statistics, and 21.0% of contexts contained raw numerical data without textual interpretations. Furthermore, pre-training models on artificially generated data and using auxiliary supervision from the conclusion text consistently improved classification accuracy across model families.

These findings demonstrate that while artificial intelligence can assist in synthesizing scientific literature, current systems struggle with numerical and scientific reasoning tasks that humans resolve with relative ease. Relying on current automated systems for high-stakes medical or policy decisions carries substantial risk, as the models miss roughly one in three scientific conclusions. However, the success of multi-phase pre-training indicates that automated data generation can bridge resource bottlenecks where expert human annotation is prohibitively expensive.

Organizations developing automated literature review or evidence-based clinical tools should adopt multi-phase training schedules and structured supervision rather than relying solely on raw text matching. Future research should prioritize enhancing numerical reasoning mechanisms to handle uninterpreted statistics and exploring full long-answer generation to improve inference depth. Stakeholders should interpret the benchmark results with measured confidence: while the dataset offers rigorous coverage of clinical research topics, the expert test set is limited to 1,000 examples, and current automated models remain insufficient for standalone, autonomous clinical interpretation.

Cover for PubMedQA: A Dataset for Biomedical Research Question Answering

Abstract

We introduce PubMedQA, a novel biomedical question answering (QA) dataset collected from PubMed abstracts. The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts. PubMedQA has 1k expert-annotated, 61.2k unlabeled and 211.3k artificially generated QA instances. Each PubMedQA instance is composed of (1) a question which is either an existing research article title or derived from one, (2) a context which is the corresponding abstract without its conclusion, (3) a long answer, which is the conclusion of the abstract and, presumably, answers the research question, and (4) a yes/no/maybe answer which summarizes the conclusion. PubMedQA is the first QA dataset where reasoning over biomedical research texts, especially their quantitative contents, is required to answer the questions. Our best performing model, multi-phase fine-tuning of BioBERT with long answer bag-of-word statistics as additional supervision, achieves 68.1% accuracy, compared to single human performance of 78.0% accuracy and majority-baseline of 55.2% accuracy, leaving much room for improvement. PubMedQA is publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 PubMedQA Dataset
  • 3.1 Data Collection
  • 3.2 Characteristics
  • 3.3 Evaluation Settings
  • 4 Methods
  • 4.1 Fine-tuning BioBERT
  • 4.2 Long Answer as Additional Supervision
  • 4.3 Multi-phase Fine-tuning Schedule
  • 4.4 Compared Models
  • 4.5 Compared Training Schedules
  • 5 Experiments
  • 5.1 Human Performance
  • 5.2 Main Results
  • 5.3 Intermediate Results
  • 6 Conclusion
  • 7 Acknowledgement
  • References
  • A Yes/no/maybe Answerability
  • B Over-represented Topics
  • C Annotation Criteria

Knowls

  1. Knowl 1 — PubMedQA Dataset Architecture and Subsets

    definition

    PubMedQA is a biomedical question answering (QA) benchmark derived from PubMed articles with structured abstracts. Each instance is a 4-tuple:

    1. Question: An existing research article question title or derived from a statement title.
    2. Context: The structured abstract excluding its concluding section (e.g., containing Objective, Methods, Results).
    3. Long Answer: The conclusive section of the abstract that addresses the research question.
    4. Final Answer: A categorical label in {yes,no,maybe}\{\text{yes}, \text{no}, \text{maybe}\} summarizing the long answer in relation to the context.

    The dataset is divided into three distinct subsets:

    • PQA-L (Labeled): 1.0k1.0\text{k} instances annotated by medical domain experts. 500500 instances are designated for 10-fold cross-validation and 500500 for the official test set. Class distribution: 55.2%55.2\% yes, 33.8%33.8\% no, and 11.0%11.0\% maybe. Average token lengths are 14.414.4 (question), 238.9238.9 (context), and 43.243.2 (long answer).
    • PQA-U (Unlabeled): 61.2k61.2\text{k} instances from naturally occurring question-titled articles, filtered to remove non-yes/no/maybe questions (e.g., wh-questions). Average token lengths are 15.015.0 (question), 237.3237.3 (context), and 45.945.9 (long answer).
    • PQA-A (Artificial): 211.3k211.3\text{k} instances automatically synthesized from statement titles (200k200\text{k} training, 11.3k11.3\text{k} validation). Class distribution: 92.8%92.8\% yes, 7.2%7.2\% no, 0.0%0.0\% maybe. Average token lengths are 16.316.3 (question), 238.0238.0 (context), and 41.041.0 (long answer).

    The task is primarily evaluated in the reasoning-required setting, where models must predict {yes,no,maybe}\{\text{yes}, \text{no}, \text{maybe}\} given only the (question, context) pair without access to the long answer at test time.

  2. Knowl 2 — Multi-Phase Fine-Tuning Framework for PubMedQA

    model/method

    To transfer knowledge across artificially generated, unlabeled, and labeled biomedical QA data, the multi-phase fine-tuning procedure operates across three sequential phases:

    1. Phase I (Pre-training on PQA-A): Initialized from pre-trained transformer weights θ0\theta_0 (such as BioBERT), the model is fine-tuned under the reasoning-required setting on the artificial set (qA,cA)(q^A, c^A) to predict labels lAl^A: θI←arg⁡min⁡θL(BioBERTθ(qA,cA),lA)\theta_I \leftarrow \arg\min_\theta \mathcal{L}(\text{BioBERT}_\theta(q^A, c^A), l^A)

    2. Phase II (Bootstrapped Training on PQA-U):

      • Reasoning-Free Pre-training: A separate instance initialized at θ0\theta_0 is fine-tuned on (qA,aA)(q^A, a^A) with long answers aAa^A to learn θB1←arg⁡min⁡θL(BioBERTθ(qA,aA),lA)\theta_{B1} \leftarrow \arg\min_\theta \mathcal{L}(\text{BioBERT}_\theta(q^A, a^A), l^A).
      • Reasoning-Free Labeled Fine-tuning: BioBERTθB1\text{BioBERT}_{\theta_{B1}} is further fine-tuned on labeled pairs (qL,aL)(q^L, a^L) yielding θB2←arg⁡min⁡θL(BioBERTθ(qL,aL),lL)\theta_{B2} \leftarrow \arg\min_\theta \mathcal{L}(\text{BioBERT}_\theta(q^L, a^L), l^L).
      • Pseudo-Labeling: Pseudo-labels for unlabeled pairs (qU,aU)(q^U, a^U) are generated as lpseudoU←BioBERTθB2(qU,aU)l^U_{\text{pseudo}} \leftarrow \text{BioBERT}_{\theta_{B2}}(q^U, a^U), retaining the most confident predictions such that the subset matches the prior class proportions of PQA-L.
      • Reasoning-Required Fine-tuning: BioBERTθI\text{BioBERT}_{\theta_I} is fine-tuned on the bootstrapped PQA-U pairs (qU,cU)(q^U, c^U) with pseudo-labels: θII←arg⁡min⁡θL(BioBERTθ(qU,cU),lpseudoU)\theta_{II} \leftarrow \arg\min_\theta \mathcal{L}(\text{BioBERT}_\theta(q^U, c^U), l^U_{\text{pseudo}})
    3. Final Phase (Fine-tuning on PQA-L): The model BioBERTθII\text{BioBERT}_{\theta_{II}} is fine-tuned on labeled instances (qL,cL)(q^L, c^L): θF←arg⁡min⁡θL(BioBERTθ(qL,cL),lL)\theta_F \leftarrow \arg\min_\theta \mathcal{L}(\text{BioBERT}_\theta(q^L, c^L), l^L) Final predictions on unseen test pairs (qL,cL)(q^L, c^L) are computed via lpred=BioBERTθF(qL,cL)l_{\text{pred}} = \text{BioBERT}_{\theta_F}(q^L, c^L).

  3. Knowl 3 — Auxiliary Long-Answer Bag-of-Words Supervision Loss

    equation

    In the reasoning-required setting, long answers (the abstract conclusions) are available during training but inaccessible during inference. To regularize representation learning, an auxiliary multi-task objective requires the model to predict the binary bag-of-words (BoW) statistics of the ground-truth long answer directly from the context and question embedding vector ([CLS]).

    The auxiliary binary cross-entropy loss over a vocabulary of size NN is defined as: LBoW=−1N∑i=1N[bilog⁡b^i+(1−bi)log⁡(1−b^i)]\mathcal{L}_{\text{BoW}} = -\frac{1}{N} \sum_{i=1}^N \left[ b_i \log \hat{b}_i + (1 - b_i) \log(1 - \hat{b}_i) \right] where bi∈{0,1}b_i \in \{0, 1\} indicates whether vocabulary token ii is present in the long answer, and b^i∈[0,1]\hat{b}_i \in [0, 1] is the predicted occurrence probability of token ii.

    The combined training objective is: L=LQA+βLBoW\mathcal{L} = \mathcal{L}_{\text{QA}} + \beta \mathcal{L}_{\text{BoW}} where LQA\mathcal{L}_{\text{QA}} is the cross-entropy classification loss between predicted and ground-truth yes/no/maybe labels, and β≥0\beta \ge 0 is a regularization weight hyperparameter. In reasoning-free settings where the long answer is directly fed as input, β\beta is set to 00.

  4. Knowl 4 — PQA-L Expert Annotation Protocol

    algorithm

    The labeled subset PQA-L is constructed using two medical domain annotators (M.D. candidates) under contrasting information regimes to obtain gold labels and quantify human performance bounds.

    Input: pre-PQA-U (unlabeled candidate instances with question titles, contexts, and conclusions)
    Output: GroundTruthLabel, ReasoningFreeAnnotation, ReasoningRequiredAnnotation
    ReasoningFreeAnnotation <- empty map
    ReasoningRequiredAnnotation <- empty map
    GroundTruthLabel <- empty map
    while not finished do
        Randomly sample an instance inst = (q, c, a) from pre-PQA-U
        if q is not yes/no/maybe answerable then
            Remove inst and continue to next iteration
        end if
        Annotator 1 annotates inst with l1∈{yes,no,maybe}l_1 \in \{yes, no, maybe\} using (q, c, a)
        Annotator 2 annotates inst with l2∈{yes,no,maybe}l_2 \in \{yes, no, maybe\} using (q, c)
        if l1=l2l_1 = l_2 then
            la←l1l_a \leftarrow l_1
        else
            Annotator 1 and Annotator 2 discuss to reach agreement lal_a
            if no consensus lal_a exists then
                Remove inst and continue to next iteration
            end if
        end if
        ReasoningFreeAnnotation[inst] <- l1l_1
        ReasoningRequiredAnnotation[inst] <- l2l_2
        GroundTruthLabel[inst] <- lal_a
    end while

    500500 instances from this procedure form a 10-fold cross-validation set and the remaining 500500 form the evaluation test set.

  5. Knowl 5 — Heuristic Generation of Artificial QA Pairs (PQA-A)

    model/method

    To enable large-scale pre-training on domain-specific question-context pairs, an automated pipeline creates the artificial dataset PQA-A from statement-titled PubMed papers:

    1. Filtering: Select PubMed articles possessing structured abstracts with conclusive sections whose titles follow a noun phrase followed by present tense verb phrases syntax (NP-(VBP/VBZ) in Stanford CoreNLP POS tagging).
    2. Question Transformation: Convert declarative titles into questions by prepending or fronting copulas (is, are) or auxiliary verbs (does, do) and appending a question mark.
    3. Label Assignment: Assign binary ground-truth labels based on the negation status of the main verb in the original title:
      • Non-negated affirmative statements (e.g., "Spontaneous electrocardiogram alterations predict ventricular fibrillation...") are converted to yes-questions and labeled yes (92.8%92.8\% of PQA-A).
      • Negated statements (e.g., "Liver grafts from selected older donors do not have significantly more...") are converted to yes-questions and labeled no (7.2%7.2\% of PQA-A).

    This yields 211.3k211.3\text{k} noisily labeled QA pairs (200k200\text{k} training, 11.3k11.3\text{k} validation).

  6. Knowl 6 — Quantitative and Analytical Reasoning Characteristics in PubMedQA

    empirical result

    Analysis of 200 randomly sampled PQA-L instances demonstrates the following distributions of question categories, required reasoning patterns, and numerical representations:

    • Question Types:

      • Does a factor influence the output?: 36.5%36.5\%
      • Is a therapy good / necessary?: 26.0%26.0\%
      • Is a statement true?: 18.0%18.0\%
      • Is a factor related to the output?: 18.0%18.0\%
      • Other: 1.5%1.5\%
    • Reasoning Types:

      • Inter-group comparison (e.g., experiment vs. control cohort comparisons): 57.5%57.5\%
      • Interpreting subgroup statistics: 16.5%16.5\%
      • Interpreting single group statistics: 16.0%16.0\%
      • Other non-statistical reasoning: 10.0%10.0\%

    Quantitative reasoning over numerical data is required in 96.5%96.5\% of all instances.

    • Context Numerical Presentation:
      • Existing natural language text interpretations of statistics (e.g., "significantly lower"): 75.5%75.5\%
      • Raw numbers only without text interpretation (e.g., reporting percentages and pp-values without qualitative adjectives): 21.0%21.0\%
      • Qualitative text only (no numbers): 3.5%3.5\%
  7. Knowl 7 — Evaluation of Biomedical QA Models and Training Schedules on PQA-L

    data/table

    The table below presents the performance (Accuracy and Macro-F1 in percentages) on the 500500-instance PQA-L test set under the reasoning-required setting across baseline models, pre-training/fine-tuning schedules, and the inclusion of auxiliary supervision (A.S.) via long-answer BoW loss.

    Model Final Phase Only Single-phase Phase I + Final Phase II + Final Multi-phase
    Acc F1 Acc F1 Acc F1 Acc F1 Acc F1
    Majority 55.20 23.71 – – – – – – – –
    Human (single) 78.00 72.19 – – – – – – – –
    w/o A.S.
    Shallow Features 53.88 36.12 57.58 31.47 57.48 37.24 56.28 40.88 53.50 39.33
    BiLSTM 55.16 23.97 55.46 39.70 58.44 40.67 52.98 33.84 59.82 41.86
    ESIM w/ BioELMo 53.90 32.40 61.28 42.99 61.96 43.32 60.34 44.38 62.08 45.75
    BioBERT 56.98 28.50 66.44 47.25 66.90 46.16 66.08 50.84 67.66 52.41
    w/ A.S.
    Shallow Features 53.60 35.92 57.30 30.45 55.82 35.09 56.46 40.76 55.06 40.67
    BiLSTM 55.22 23.86 55.96 40.26 61.06 41.18 54.12 34.11 58.86 41.06
    ESIM w/ BioELMo 53.96 31.07 62.68 43.59 63.72 47.04 60.16 45.81 63.72 47.90
    BioBERT 57.28 28.70 66.66 46.70 67.24 46.21 66.44 51.41 68.08 52.72

    Key empirical findings:

    1. Model Hierarchy: BioBERT consistently outperforms ESIM with BioELMo, which outperforms BiLSTM and shallow bag-of-words/TF-IDF features across all schedules.
    2. Multi-Phase Benefits: Multi-phase fine-tuning achieves the overall highest performance (68.08%68.08\% accuracy, 52.72%52.72\% F1 with BioBERT + A.S.). Training exclusively on PQA-L (Final Phase Only) yields near-majority accuracy (56.98%56.98\%) due to scarce supervision (450450 instances per training fold).
    3. Auxiliary Supervision Impact: Long-answer BoW loss improves performance in 28 of 40 model configurations.
    4. Human-Machine Gap: The best model (68.08%68.08\%) remains well below the single human performance baseline (78.00%78.00\% accuracy, 72.19%72.19\% Macro-F1).
  8. Knowl 8 — Reasoning-Free vs. Reasoning-Required Performance Discrepancy

    empirical result

    When models or human annotators are provided with the long answer (the explicit abstract conclusion) alongside the question—termed the reasoning-free setting—performance is substantially higher than in the reasoning-required setting (where only the question and context are provided):

    • Single Human Annotator Performance on PQA-L test set:

      • Reasoning-free: 90.40%90.40\% accuracy, 84.18%84.18\% Macro-F1.
      • Reasoning-required: 78.00%78.00\% accuracy, 72.19%72.19\% Macro-F1.
    • Model Performance in Reasoning-Free Setup:

      • On PQA-A pre-training data, BioBERT reaches 98.28%98.28\% accuracy and 93.17%93.17\% Macro-F1 (compared to 96.50%96.50\% accuracy and 84.65%84.65\% Macro-F1 in reasoning-required Phase I).
      • On PQA-L data, BioBERT achieves 80.80%80.80\% accuracy and 63.50%63.50\% Macro-F1, ESIM w/ BioELMo achieves 74.06%74.06\% accuracy and 58.53%58.53\% Macro-F1, and BiLSTM achieves 71.46%71.46\% accuracy and 50.93%50.93\% Macro-F1.

    This gap indicates that abstract conclusions directly express the research findings in natural language, enabling accurate extraction for bootstrapping PQA-U, while inferring conclusions from raw experimental context requires non-trivial numerical and domain reasoning.

  9. Knowl 9 — Context-Dependent Yes/No/Maybe Annotation Criteria

    definition

    Because research questions in biomedical literature are rarely universally true across all patient populations and conditions, PubMedQA defines truth values as context-dependent based on the study findings:

    • "yes": Assigned if the empirical experiments and statistical results described in the abstract support the question's premise (e.g., showing a statistically significant therapeutic benefit in the experimental cohort relative to controls).
    • "no": Assigned if the empirical results do not support the hypothesis, fail to show a statistically significant difference, or show adverse outcomes.
    • "maybe": Assigned under two conditions:
      1. The study explicitly documents mixed findings where the answer is true under certain subgroup conditions and false under others.
      2. The question asks about multiple interventions, observations, or disease targets simultaneously (e.g., "Do Disease A, Disease B and/or Disease C benefit from drug X?"), and the outcome is positive for some but negative or neutral for others.

Coverage note — No substantial contributed material was omitted from the paper.

References

  1. 1.Qian Chen, Xiaodan Zhu, Zhenhua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2016. Enhanced lstm for natural language inference. arXiv preprint arXiv:1609.06038.
  2. 2.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers).
  3. 3.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  4. 4.William Hersh, Aaron Cohen, Lynn Ruslen, and Phoebe Roberts. 2007. Trec 2007 genomics track overview. In TREC 2007.
  5. 5.William Hersh, Aaron M. Cohen, Phoebe Roberts, and Hari Krishna Rekapalli. 2006. Trec 2006 genomics track overview. In TREC 2006.
  6. 6.Qiao Jin, Bhuwan Dhingra, William W Cohen, and Xinghua Lu. 2019. Probing biomedical embeddings from language models. arXiv preprint arXiv:1904.02181.
  7. 7.Seongsoon Kim, Donghyeon Park, Yonghwa Choi, Kyubum Lee, Byounggun Kim, Minji Jeon, Jihye Kim, Aik Choon Tan, and Jaewoo Kang. 2018. A pilot study of biomedical text comprehension using an attention-based deep neural reader: Design and experimental analysis. JMIR medical informatics, 6(1):e2.
  8. 8.Tomas Kocisky, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gaabor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association of Computational Linguistics, 6:317–328.
  9. 9.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Rhinehart, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, et al. 2019. Natural questions: a benchmark for question answering research.
  10. 10.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683.
  11. 11.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019. Biobert: pre-trained biomedical language representation model for biomedical text mining. arXiv preprint arXiv:1901.08746.
  12. 12.Shuming Ma, Xu Sun, Yizhong Wang, and Junyang Lin. 2018. Bag-of-words as target for neural machine translation. arXiv preprint arXiv:1805.04871.
  13. 13.Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations, pages 55–60.
  14. 14.Roser Morante, Martin Krallinger, Alfonso Valencia, and Walter Daelemans. 2012. Machine reading of biomedical texts about alzheimers disease. In CLEF 2012 Conference and Labs of the Evaluation Forum-Question Answering For Machine Reading Evaluation (QA4MRE), Rome/Forner, J.[edit.]; ea, pages 1–14.
  15. 15.Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. 2018. emrqa: A large corpus for question answering on electronic medical records. arXiv preprint arXiv:1809.00732.
  16. 16.Dimitris Pappas, Ion Androutsopoulos, and Haris Papageorgiou. 2018. Bioread: A new dataset for biomedical reading comprehension. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC-2018).
  17. 17.Anselmo Penas, Eduard Hovy, Pamela Forner, Alvaro Rodrigo, Richard Sutcliffe, and Roser Morante. 2013. Qa4mre 2011-2013: Overview of question answering for machine reading evaluation. In International Conference of the Cross-Language Evaluation Forum for European Languages, pages 303–320. Springer.
  18. 18.Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. arXiv preprint arXiv:1802.05365.
  19. 19.Sampo Pyysalo, Filip Ginter, Hans Moen, Tapio Salakoski, and Sophia Ananiadou. 2013. Distributional semantics resources for biomedical text processing.
  20. 20.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
  21. 21.Marzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktaschel, Mike Sheldon, Guillaume Bouchard, and Sebastian Riedel. 2018. Interpretation of natural language rules in conversational machine reading. arXiv preprint arXiv:1809.01494.
  22. 22.Hiroaki Sakamoto, Yasunori Watanabe, and Masataka Satou. 2011. Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting? Annals of thoracic and cardiovascular surgery, 17(4):376–382.
  23. 23.George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015. An overview of the bioasq large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16(1):138.
  24. 24.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600.

Citation

MLA
Jin, Q., et al. “PubMedQA: A Dataset for Biomedical Research Question Answering”. arXiv, 2019, http://arxiv.org/abs/1909.06146v1.
APA
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., & Lu, X. (2019). PubMedQA: A Dataset for Biomedical Research Question Answering. arXiv. http://arxiv.org/abs/1909.06146v1
Chicago
Jin, Q., B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu. 2019. “PubMedQA: A Dataset for Biomedical Research Question Answering”. arXiv. http://arxiv.org/abs/1909.06146v1.
Harvard
Jin, Q. et al. (2019) “PubMedQA: A Dataset for Biomedical Research Question Answering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1909.06146v1.
Vancouver
1. Jin Q, Dhingra B, Liu Z, Cohen WW, Lu X (2019) PubMedQA: A Dataset for Biomedical Research Question Answering. arXiv

BibTeX

@article{jin2019pubmedqa,
  title = {PubMedQA: A Dataset for Biomedical Research Question Answering},
  author = {Jin, Qiao and Dhingra, Bhuwan and Liu, Zhengping and Cohen, William W. and Lu, Xinghua},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1909.06146v1},
  eprint = {1909.06146}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/