Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

Yu GuRobert TinnHao ChengMichael LucasNaoto UsuyamaXiaodong LiuTristan NaumannJianfeng GaoHoifung Poon

article2020ACM Transactions on Computing for Healthcare2,745 citations

Demonstrates that pretraining language models from scratch on domain text substantially outperforms continual pretraining from general-domain models in biomedicine, while introducing the BLURB benchmark and new state-of-the-art models for biomedical NLP.

Listen

Biomedical natural language processing faces a core challenge: while pretraining large language models like BERT on general text has driven broad gains, specialized domains with abundant unlabeled data such as biomedicine may not benefit from starting with those general models. The prevailing practice of mixed-domain pretraining, often by continual training of a general BERT model on biomedical text, rests on the assumption that out-of-domain data remains helpful, yet this can introduce negative transfer when the target domain already contains tens of millions of documents.

The article set out to test whether pretraining a language model entirely from scratch on in-domain biomedical text would outperform the mixed-domain approach across a wide range of downstream tasks. To enable rigorous comparison, the authors assembled BLURB, a benchmark of thirteen publicly available datasets spanning named-entity recognition, PICO extraction, relation extraction, sentence similarity, document classification, and question answering, and they released both the benchmark and a public leaderboard.

They pretrained PubMedBERT from scratch on 14 million PubMed abstracts using an in-domain vocabulary, whole-word masking, and standard masked-language-model objectives, then compared it head-to-head with BERT, RoBERTa, BioBERT, SciBERT, ClinicalBERT, and BlueBERT under identical fine-tuning protocols. The experiments also examined modeling variants such as tagging schemes for named-entity recognition and entity representation choices for relation extraction.

PubMedBERT achieved the highest overall BLURB score of 81.16 and led or matched the best result on eleven of the thirteen tasks, with gains of one to several points over the next-best model on most relation-extraction and question-answering datasets. In-domain vocabulary and pretraining from scratch proved clearly superior to continual pretraining from general-domain models or to mixing computer-science or clinical text; whole-word masking added consistent value while adversarial pretraining did not. Notably, the simpler IO tagging scheme performed on par with or better than the conventional BIO scheme for BERT-based named-entity recognition, and entity dummification remained the most reliable representation for relation extraction.

These results indicate that, when sufficient in-domain text exists, domain-specific pretraining from scratch removes the risk of negative transfer and yields stronger task performance, which can translate directly into more accurate information extraction for literature review, pharmacovigilance, and evidence synthesis. The findings also simplify deployment by showing that several common modeling practices add little or no value once transformer-based models are used.

The authors recommend adopting in-domain pretraining whenever the target domain offers tens of millions of documents, releasing their PubMedBERT checkpoints and the BLURB leaderboard to accelerate community progress. Next steps include extending the benchmark to clinical notes and other vertical domains, incorporating additional tasks, and exploring longer pretraining schedules when full-text articles are added.

The study is limited to PubMed abstracts and the tasks currently in BLURB; performance on clinical text or on more complex question-answering formats was not evaluated. Results on the smallest datasets exhibit run-to-run variance, and hyperparameter search was constrained to keep computation tractable. Overall, the head-to-head comparisons on a comprehensive public benchmark provide high confidence in the central conclusion that domain-specific pretraining from scratch is the stronger strategy for biomedicine.

Cover for Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

Abstract

Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly-available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition (NER). To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Language Model Pretraining
  • 2.1.1 Vocabulary
  • 2.1.2 Model Architecture
  • 2.1.3 Self-Supervision
  • 2.1.4 Advanced Pretraining Techniques
  • 2.2 Biomedical Language Model Pretraining
  • 2.2.1 Mixed-Domain Pretraining
  • 2.2.2 Domain-Specific Pretraining from Scratch
  • 2.3 BLURB: A Comprehensive Benchmark for Biomedical NLP
  • 2.3.1 Named Entity Recognition (NER)
  • 2.3.2 Evidence-Based Medical Information Extraction (PICO)
  • 2.3.3 Relation Extraction
  • 2.3.4 Sentence Similarity
  • 2.3.5 Document Classification
  • 2.3.6 Question Answering (QA)
  • 2.4 Task-Specific Fine-Tuning
  • 2.4.1 A General Architecture for Fine-Tuning Neural Language Models
  • 2.4.2 Task-Specific Problem Formulation and Modeling Choices
  • 2.5 Experimental Settings
  • 3 Results
  • 3.1 Domain-Specific Pretraining vs Mixed-Domain Pretraining
  • 3.2 Ablation Study on Pretraining Techniques
  • 3.3 Ablation Study on Fine-Tuning Methods
  • 4 Discussion
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — PubMedBERT: Domain-Specific Pretraining from Scratch for Biomedical NLP

    model/method

    PubMedBERT is a domain-specific BERT model pretrained entirely from scratch using in-domain biomedical text rather than performing continual pretraining initialized from a general-domain language model.

    Pretraining Corpus & Vocabulary: The pretraining corpus consists of 14 million PubMed abstracts (3.2 billion words, 21 GB of text) after discarding abstracts with fewer than 128 words to eliminate noise. An in-domain WordPiece vocabulary of 30,000 subwords is constructed directly from this PubMed corpus using a unigram language model likelihood criterion, ensuring common biomedical terms (e.g., naloxone, acetyltransferase, cardiomyocyte) are represented as single tokens rather than shattered into multiple general subwords.

    Architecture & Optimization: The model follows the BERTBASE\text{BERT}_{\text{BASE}} architecture with 12 transformer layers, 12 attention heads, hidden dimension 768, and 100 million parameters (uncased). It is optimized using Masked Language Modeling (MLM) with Whole-Word Masking (WWM) at a 15% masking rate alongside Next Sentence Prediction (NSP). Pretraining uses the Adam optimizer with a slanted triangular learning rate schedule: learning rate warms up linearly over the first 10% of steps to a peak of 6×1046 \times 10^{-4} and decays linearly to 0 over the remaining 90% of steps. Training is conducted for 62,500 steps with a batch size of 8,192 on a DGX-2 machine containing 16 NVIDIA V100 GPUs (approximately 5 days of training).

  2. Knowl 2 — Biomedical Language Understanding & Reasoning Benchmark (BLURB)

    experimental setup

    The Biomedical Language Understanding & Reasoning Benchmark (BLURB) is a standardized evaluation suite consisting of 13 publicly available biomedical NLP datasets spanning 6 distinct task categories:

    1. Named Entity Recognition (NER) (evaluated via entity-level F1F_1):
      • BC5-chem (5,203 train / 5,347 dev / 5,385 test mentions)
      • BC5-disease (4,182 train / 4,244 dev / 4,424 test mentions)
      • NCBI-disease (5,134 train / 787 dev / 960 test mentions)
      • BC2GM (15,197 train / 3,061 dev / 6,325 test mentions)
      • JNLPBA (46,750 train / 4,551 dev / 8,662 test mentions)
    2. Evidence-Based Medical Information Extraction (PICO) (evaluated via macro word-level F1F_1 across Participant, Intervention, and Outcome spans):
      • EBM PICO (339,167 train / 85,321 dev / 16,364 test elements)
    3. Relation Extraction (evaluated via micro F1F_1):
      • ChemProt (18,035 train / 11,268 dev / 15,745 test pairs; 5-class + false)
      • DDI (25,296 train / 2,496 dev / 5,716 test pairs)
      • GAD (4,261 train / 535 dev / 534 test pairs)
    4. Sentence Similarity (evaluated via Pearson correlation coefficient):
      • BIOSSES (64 train / 16 dev / 20 test sentence pairs)
    5. Document Classification (evaluated via micro F1F_1 across 10 top-level cancer hallmark categories):
      • HoC (1,295 train / 186 dev / 371 test abstracts)
    6. Question Answering (QA) (evaluated via Accuracy):
      • PubMedQA (450 train / 50 dev / 500 test questions; yes/maybe/no)
      • BioASQ Task 7b (670 train / 75 dev / 140 test questions; yes/no)

    Metric Aggregation: To avoid skewing results toward task types with many datasets (such as NER), the overall summary BLURB score is calculated by first taking the arithmetic mean of dataset scores within each of the 6 task categories, and then computing the macro-average across the 6 task category scores.

  3. Knowl 3 — Comparative Performance of Domain-Specific and Mixed-Domain Language Models on BLURB

    data/table

    When evaluated under an identical fine-tuning protocol across the 13 BLURB benchmark datasets, domain-specific pretraining from scratch (PubMedBERT) outperforms general-domain models (BERT, RoBERTa) and mixed-domain continually pretrained models (BioBERT, SciBERT, ClinicalBERT, BlueBERT).

    Dataset / Metric BERT BERT RoBERTa BioBERT SciBERT SciBERT BlueBERT PubMedBERT
    uncased cased cased cased uncased cased cased uncased
    BC5-chem (F1F_1) 89.25 89.99 89.43 92.85 92.49 92.51 91.19 93.33
    BC5-disease (F1F_1) 81.44 79.92 80.65 84.70 84.54 84.70 83.69 85.62
    NCBI-disease (F1F_1) 85.67 85.87 86.62 89.13 88.10 88.25 88.04 87.82
    BC2GM (F1F_1) 80.90 81.23 80.90 83.82 83.36 83.36 81.87 84.52
    JNLPBA (F1F_1) 77.69 77.51 77.86 78.55 78.68 78.51 77.71 79.10
    EBM PICO (F1F_1) 72.34 71.70 73.02 73.18 73.12 73.06 72.54 73.38
    ChemProt (F1F_1) 71.86 71.54 72.98 76.14 75.24 75.00 71.46 77.24
    DDI (F1F_1) 80.04 79.34 79.52 80.88 81.06 81.22 77.78 82.36
    GAD (F1F_1) 80.41 79.61 80.63 82.36 82.38 81.34 79.15 83.96
    BIOSSES (Pearson) 82.68 81.40 81.25 89.52 86.25 87.15 85.38 92.30
    HoC (F1F_1) 80.20 80.12 79.66 81.54 80.66 81.16 80.48 82.32
    PubMedQA (Acc) 51.62 49.96 52.84 60.24 57.38 51.40 48.44 55.84
    BioASQ (Acc) 70.36 74.44 75.20 84.14 78.86 74.22 68.71 87.56
    BLURB Score 76.11 75.86 76.46 80.34 78.86 78.14 76.27 81.16

    PubMedBERT attains the highest overall score of 81.16. Pretraining strictly on general-domain text (RoBERTa: 76.46; BERT: 76.11) underperforms models exposed to biomedical text. Continual pretraining (BioBERT: 80.34) or pretraining on mixed scientific domains (SciBERT uncased: 78.86) falls behind pure in-domain pretraining from scratch.

  4. Knowl 4 — Impact of In-Domain Vocabulary and Whole-Word Masking on Biomedical Language Representation

    data/table

    Constructing an in-domain vocabulary from PubMed and using Whole-Word Masking (WWM) during pretraining both provide clear performance benefits on downstream biomedical tasks compared to general vocabularies (Wiki + Books) and standard subword masking.

    Wiki + Books Vocab PubMed Vocab
    Dataset WordPiece Whole-Word WordPiece Whole-Word (PubMedBERT)
    BC5-chem 93.20 93.31 92.96 93.33
    BC5-disease 85.00 85.28 84.72 85.62
    NCBI-disease 88.39 88.53 87.26 87.82
    BC2GM 83.65 83.93 83.19 84.52
    JNLPBA 78.83 78.77 78.63 79.10
    EBM PICO 73.30 73.52 73.44 73.38
    ChemProt 75.04 76.70 75.72 77.24
    DDI 81.30 82.60 80.84 82.36
    GAD 83.02 82.42 81.74 83.96
    BIOSSES 91.36 91.79 92.45 92.30
    HoC 81.76 81.74 80.38 82.32
    PubMedQA 52.20 55.92 54.76 55.84
    BioASQ 73.69 76.41 78.51 87.56
    BLURB Score 79.16 79.96 79.62 81.16

    A primary mechanism behind the in-domain vocabulary advantage is the reduction of subword fragmentation: across all benchmark tasks, the average sequence length in word pieces decreases markedly with the PubMed vocabulary (e.g., ChemProt drops from 75.4 to 55.5 tokens; DDI drops from 106.0 to 75.9 tokens; BioASQ drops from 702.4 to 541.4 tokens), preserving semantic units and improving self-attention modeling.

  5. Knowl 5 — Evaluation of Continual Mixed-Domain Pretraining Versus Pretraining from Scratch

    data/table

    Pretraining initially on general-domain text (Wikipedia + Books) and continuing pretraining on PubMed text confers no benefit over pretraining directly from scratch on PubMed, even when controlling for computation and vocabulary.

    Pretraining Corpus Wiki+Books \to PubMed Wiki+Books \to PubMed PubMed (half time) PubMed (PubMedBERT)
    Vocabulary Wiki + Books PubMed PubMed PubMed
    Total Compute Full (2×\times) Full (2×\times) Half (1×\times) Full (2×\times)
    BC5-chem 92.85 93.41 93.05 93.33
    BC5-disease 84.70 85.43 85.02 85.62
    NCBI-disease 89.13 87.60 87.77 87.82
    BC2GM 83.82 84.03 84.11 84.52
    JNLPBA 78.55 79.01 78.98 79.10
    EBM PICO 73.18 73.80 73.74 73.38
    ChemProt 76.14 77.05 76.69 77.24
    DDI 80.88 81.96 81.21 82.36
    GAD 82.36 82.47 82.80 83.96
    BIOSSES 89.52 89.93 92.12 92.30
    HoC 81.54 83.14 82.13 82.32
    PubMedQA 60.24 54.84 55.28 55.84
    BioASQ 84.14 79.00 79.43 87.56
    BLURB Score 80.34 80.03 80.23 81.16

    Initializing continual pretraining with general-domain BERT while swapping to an in-domain PubMed vocabulary achieves a score of 80.03. Pretraining from scratch using solely PubMed text achieves a comparable score of 80.23 with only half the training steps (same compute as standard BERT pretraining), and achieves 81.16 when given the full compute budget.

  6. Knowl 6 — Effect of Pretraining with Full-Text Articles (PubMed Central) Versus Abstracts

    data/table

    Expanding the pretraining corpus from PubMed abstracts (3.2B words, 21 GB) to include PubMed Central (PMC) full-text articles (16.8B words, 107 GB) produces mixed results depending on the total training duration.

    Dataset PubMed (Abstracts) PubMed + PMC (Standard) PubMed + PMC (60% Longer Training)
    BC5-chem 93.33 93.36 93.34
    BC5-disease 85.62 85.62 85.76
    NCBI-disease 87.82 88.34 88.04
    BC2GM 84.52 84.39 84.37
    JNLPBA 79.10 78.90 79.16
    EBM PICO 73.38 73.64 73.72
    ChemProt 77.24 76.96 76.80
    DDI 82.36 83.56 82.06
    GAD 83.96 84.08 82.90
    BIOSSES 92.30 90.39 92.31
    HoC 82.32 82.16 82.62
    PubMedQA 55.84 61.02 60.02
    BioASQ 87.56 83.43 87.20
    BLURB Score 81.16 81.01 81.50

    Pretraining on PubMed+PMC with standard steps (62,500) results in a slight decrease in overall performance (81.01 vs 81.16) because full-text articles introduce more noise and differ in distribution from target evaluation datasets, which are primarily abstract-based. Extending pretraining by 60% (100,000 steps) enables the model to better absorb the expanded data, raising the BLURB score to 81.50.

  7. Knowl 7 — Impact of Adversarial Pretraining in Domain-Specific Language Modeling

    empirical result

    Adversarial pretraining—which perturbs token embeddings during pretraining to maximize adversarial loss while minimizing masked language model loss—does not improve downstream performance on the biomedical benchmark when pretraining on homogeneous in-domain text.

    When applying adversarial pretraining to PubMedBERT under the standard pretraining protocol:

    • The overall BLURB macro score decreases from 81.16 to 80.77.
    • Significant drops occur in question answering: BioASQ accuracy decreases from 87.56% to 82.71%, and PubMedQA accuracy decreases from 55.84% to 53.30%.
    • Minor drops occur on BC5-chem (93.33 to 93.17), BC2GM (84.52 to 84.07), ChemProt (77.24 to 77.04), and GAD (83.96 to 83.54).
    • Notable gains are confined to sentence similarity (BIOSSES Pearson correlation rises from 92.30 to 94.11) and DDI relation extraction (F1F_1 increases from 82.36 to 83.62).

    This indicates that adversarial pretraining is primarily advantageous when pretraining corpora are diverse and out-of-domain relative to downstream tasks, but offers limited or negative utility in focused in-domain pretraining.

  8. Knowl 8 — Sufficiency of Simplified IO Tagging Scheme for Transformer-Based Biomedical NER

    data/table

    In named entity recognition (NER), distinguishing positional boundaries of entities using complex labeling schemes such as BIO (Beginning, Inside, Outside) or BIOUL (Beginning, Inside, Outside, Unit-length, Last) provides no meaningful advantage over the minimal IO (Inside, Outside) tagging scheme when fine-tuning transformer models like PubMedBERT.

    Dataset BIO BIOUL IO
    BC5-chem (F1F_1) 93.33 93.37 93.11
    BC5-disease (F1F_1) 85.62 85.59 85.63
    JNLPBA (F1F_1) 79.10 79.02 79.05

    Because the multi-layer self-attention mechanism in BERT already captures long-range contextual and boundary dependencies across the entire sequence, the sequential constraints traditionally enforced by BIO/BIOUL tagging in linear-chain models (e.g., CRFs) become largely redundant.

  9. Knowl 9 — Linear Classification Layers Versus Recurrent Encoders for Fine-Tuning Biomedical BERT

    data/table

    Adding a Bidirectional Long Short-Term Memory (Bi-LSTM) network on top of BERT contextual representations before prediction does not improve performance over a simple linear projection layer in token classification (NER) or sequence classification (Relation Extraction).

    Task Dataset Linear Layer Bi-LSTM
    NER (Entity F1F_1) BC5-chem 93.33 93.12
    NER (Entity F1F_1) BC5-disease 85.62 85.64
    NER (Entity F1F_1) JNLPBA 79.10 79.10
    Relation Extraction (Micro F1F_1) ChemProt 77.24 75.40
    Relation Extraction (Micro F1F_1) DDI 82.36 81.70
    Relation Extraction (Micro F1F_1) GAD 83.96 83.42

    Across all six evaluated datasets using PubMedBERT representations, the linear layer achieves equal or higher F1F_1 scores (notably outperforming Bi-LSTM by 1.84 points on ChemProt and 0.66 points on DDI), demonstrating that explicit recurrent modeling is unnecessary atop contextualized transformer outputs.

  10. Knowl 10 — Entity Marking and Featurization Techniques for Biomedical Relation Extraction

    data/table

    Directly feeding original text without entity anonymization or boundary markers into BERT causes severe performance degradation in relation extraction due to memorization of entity pairs and lack of entity localization. Entity dummification or explicit entity markers are necessary to achieve robust generalization.

    Input Text Format Relation Classification Encoding ChemProt (F1F_1) DDI (F1F_1)
    Entity Dummification [CLS] Token Encoding 77.24 82.36
    Entity Dummification Mention Representation Pooling 77.22 82.08
    Original Text [CLS] Token Encoding 50.52 37.00
    Original Text Mention Representation Pooling 75.48 79.42
    Entity Markers [CLS] Token Encoding 77.72 82.22
    Entity Markers Mention Representation Pooling 77.22 82.42
    Entity Markers Entity Start Tag Encoding 77.58 82.18
    • Original text with [CLS] encoding: Fails catastrophically (F1F_1 drops to 50.52 on ChemProt and 37.00 on DDI) because the global [CLS] token cannot resolve which entity pair is being queried.
    • Entity Dummification: Replacing entity mentions with type placeholders (e.g., \$GENE, \$DRUG) prevents overfitting and provides strong performance with [CLS] (77.24 on ChemProt, 82.36 on DDI).
    • Entity Markers: Appending special start/end tags around target entities while retaining original words yields the highest overall accuracy (77.72 on ChemProt with [CLS]), preserving entity semantics while disambiguating the target pair.

Coverage note — None was omitted; all major contributions—including the PubMedBERT pretraining methodology, the BLURB benchmark specification, main empirical comparisons, pretraining ablation studies, and fine-tuning modeling analyses—are fully captured.

References

  1. 1.Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019. Publicly Available Clinical BERT Embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop. Association for Computational Linguistics, Minneapolis, Minnesota, USA, 72–78. https://doi.org/10.18653/v1/W19-1909
  2. 2.Marianna Apidianaki, Saif M. Mohammad, Jonathan May, Ekaterina Shutova, Steven Bethard, and Marine Carpuat (Eds.). 2018. Proceedings of The 12th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2018, New Orleans, Louisiana, USA, June 5-6, 2018. Association for Computational Linguistics. https://www.aclweb.org/anthology/volumes/S18-1/
  3. 3.Cecilia N. Arighi, Phoebe M. Roberts, Shashank Agarwal, Sanmitra Bhattacharya, Gianni Cesareni, Andrew Chatr-aryamontri, Simon Clematide, Pascale Gaudet, Michelle Gwinn Giglio, Ian Harrow, Eva Huala, Martin Krallinger, Ulf Leser, Donghui Li, Feifan Liu, Zhiyong Lu, Lois J. Maltais, Naoaki Okazaki, Livia Perfetto, Fabio Rinaldi, Rune Sætre, David Salgado, Padmini Srinivasan, Philippe E. Thomas, Luca Toldo, Lynette Hirschman, and Cathy H. Wu. 2011. BioCreative III interactive task: an overview. BMC Bioinformatics 12, 8 (03 Oct 2011), S4. https://doi.org/10.1186/1471-2105-12-S8-S4
  4. 4.Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011. Domain Adaptation via Pseudo In-Domain Data Selection. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Edinburgh, Scotland, UK., 355–362. https://www.aclweb.org/anthology/D11-1033
  5. 5.Simon Baker, Imran Ali, Ilona Silins, Sampo Pyysalo, Yufan Guo, Johan Högberg, Ulla Stenius, and Anna Korhonen. 2017. Cancer Hallmarks Analytics Tool (CHAT): a text mining approach to organize and evaluate scientific literature on cancer. Bioinformatics 33, 24 (2017), 3973–3981.
  6. 6.Simon Baker, Ilona Silins, Yufan Guo, Imran Ali, Johan Högberg, Ulla Stenius, and Anna Korhonen. 2015. Automatic semantic classification of scientific literature according to the hallmarks of cancer. Bioinformatics 32, 3 (2015), 432–440.
  7. 7.Livio Baldini Soares, Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski. 2019. Matching the Blanks: Distributional Similarity for Relation Learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 2895–2905. https://doi.org/10.18653/v1/P19-1279
  8. 8.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 3615–3620. https://doi.org/10.18653/v1/D19-1371
  9. 9.Steven Bethard, Marine Carpuat, Marianna Apidianaki, Saif M. Mohammad, Daniel M. Cer, and David Jurgens (Eds.). 2017. Proceedings of the 11th International Workshop on Semantic Evaluation, SemEval@ACL 2017, Vancouver, Canada, August 3-4, 2017. Association for Computational Linguistics. https://www.aclweb.org/anthology/volumes/S17-2/
  10. 10.Steven Bethard, Daniel M. Cer, Marine Carpuat, David Jurgens, Preslav Nakov, and Torsten Zesch (Eds.). 2016. Proceedings of the 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2016, San Diego, CA, USA, June 16-17, 2016. The Association for Computer Linguistics. https://www.aclweb.org/anthology/volumes/S16-1/
  11. 11.Àlex Bravo, Janet Piñero, Núria Queralt-Rosinach, Michael Rautschka, and Laura I Furlong. 2015. Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research. BMC bioinformatics 16, 1 (2015), 55.
  12. 12.Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer. 1992. Class-based n-gram models of natural language. Computational linguistics 18, 4 (1992), 467–480.
  13. 13.Rich Caruana. 1997. Multitask learning. Machine learning 28, 1 (1997), 41–75.
  14. 14.Gamal Crichton, Sampo Pyysalo, Billy Chiu, and Anna Korhonen. 2017. A neural network multi-task learning approach to biomedical named entity recognition. BMC bioinformatics 18, 1 (2017), 368.
  15. 15.Dina Demner-Fushman, Kevin Bretonnel Cohen, Sophia Ananiadou, and Junichi Tsujii (Eds.). 2019. Proceedings of the 18th BioNLP Workshop and Shared Task, BioNLP@ACL 2019, Florence, Italy, August 1, 2019. Association for Computational Linguistics. https://www.aclweb.org/anthology/volumes/W19-50/
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171–4186.
  17. 17.Mona T. Diab, Timothy Baldwin, and Marco Baroni (Eds.). 2013. Proceedings of the 7th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2013, Atlanta, Georgia, USA, June 14-15, 2013. The Association for Computer Linguistics. https://www.aclweb.org/anthology/volumes/S13-2/
  18. 18.Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014. NCBI disease corpus: a resource for disease name recognition and concept normalization. Journal of biomedical informatics 47 (2014), 1–10.
  19. 19.Jingcheng Du, Qingyu Chen, Yifan Peng, Yang Xiang, Cui Tao, and Zhiyong Lu. 2019. ML-Net: multi-label classification of biomedical texts with deep neural networks. Journal of the American Medical Informatics Association 26, 11 (06 2019), 1279–1285. https://doi.org/10.1093/jamia/ocz085 arXiv:https://academic.oup.com/jamia/article-pdf/26/11/1279/36089060/ocz085.pdf
  20. 20.Douglas Hanahan and Robert A Weinberg. 2000. The hallmarks of cancer. cell 100, 1 (2000), 57–70.
  21. 21.María Herrero-Zazo, Isabel Segura-Bedmar, Paloma Martínez, and Thierry Declerck. 2013. The DDI corpus: An annotated corpus with pharmacological substances and drug–drug interactions. Journal of biomedical informatics 46, 5 (2013), 914–920.
  22. 22.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735–1780.
  23. 23.Jeremy Howard and Sebastian Ruder. 2018. Universal Language Model Fine-tuning for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Melbourne, Australia, 328–339. https://doi.org/10.18653/v1/P18-1031
  24. 24.Robin Jia, Cliff Wong, and Hoifung Poon. 2019. Document-Level 𝑁 -ary Relation Extraction with Multiscale Representation Learning. In NAACL.
  25. 25.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 2567–2577. https://doi.org/10.18653/v1/D19-1259
  26. 26.Alistair E.W. Johnson, Tom J. Pollard, Lu Shen, Li-wei H. Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. 2016. MIMIC-III, a freely accessible critical care database. Scientific Data 3, 1 (24 May 2016), 160035. https://doi.org/10.1038/sdata.2016.35
  27. 27.Jin-Dong Kim, Tomoko Ohta, Yoshimasa Tsuruoka, Yuka Tateisi, and Nigel Collier. 2004. Introduction to the Bio-entity Recognition Task at JNLPBA. In Proceedings of the International Joint Workshop on Natural Language Processing in Biomedicine and its Applications (NLPBA/BioNLP). COLING, Geneva, Switzerland, 73–78. https://www.aclweb.org/anthology/W04-1213
  28. 28.Jin-Dong Kim, Yue Wang, Toshihisa Takagi, and Akinori Yonezawa. 2011. Overview of Genia Event Task in BioNLP Shared Task 2011. In Proceedings of the BioNLP Shared Task 2011 Workshop (Portland, Oregon) (BioNLP Shared Task ’11). Association for Computational Linguistics, USA, 7–15.
  29. 29.Sun Kim, Rezarta Islamaj Dogan, Andrew Chatr-aryamontri, Mike Tyers, W. John Wilbur, and Donald C. Comeau. 2015. Overview of BioCreative V BioC Track.
  30. 30.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http://arxiv.org/abs/1412.6980
  31. 31.Martin Krallinger, Obdulia Rabal, Saber A Akhondi, Martın Pérez Pérez, Jesús Santamaría, GP Rodríguez, et al. 2017. Overview of the BioCreative VI chemical-protein interaction Track. In Proceedings of the sixth BioCreative challenge evaluation workshop, Vol. 1. 141–146.
  32. 32.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Brussels, Belgium, 66–71. https://doi.org/10.18653/v1/D18-2012
  33. 33.John Lafferty, Andrew Mccallum, and Fernando C. N. Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In in Proceedings of the 18th International Conference on Machine Learning. 282–289.
  34. 34.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics (09 2019). https://doi.org/10.1093/bioinformatics/btz682
  35. 35.Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016. BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database 2016 (2016).
  36. 36.Percy Liang. 2005. Semi-supervised learning for natural language. Ph.D. Dissertation. Massachusetts Institute of Technology.
  37. 37.Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. 2020. Adversarial Training for Large Neural Language Models. arXiv preprint arXiv:2004.08994 (2020).
  38. 38.Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, and Ye-Yi Wang. 2015. Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 912–921.
  39. 39.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019).
  40. 40.Yuqing Mao, Kimberly Van Auken, Donghui Li, Cecilia N. Arighi, Peter McQuilton, G. Thomas Hayman, Susan Tweedie, Mary L. Schaeffer, Stanley J. F. Laulederkind, Shur-Jen Wang, Julien Gobeill, Patrick Ruch, Anh Tuan Luu, Jung jae Kim, Jung-Hsien Chiang, Yu-De Chen, Chia-Jung Yang, Hongfang Liu, Dongqing Zhu, Yanpeng Li, Hong Yu, Ehsan Emadzadeh, Graciela Gonzalez, Jian-Ming Chen, Hong-Jie Dai, and Zhiyong Lu. 2014. Overview of the gene ontology task at BioCreative IV. Database: The Journal of Biological Databases and Curation 2014 (2014).
  41. 41.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013).
  42. 42.Anastasios Nentidis, Konstantinos Bougiatiotis, Anastasia Krithara, and Georgios Paliouras. 2019. Results of the Seventh Edition of the BioASQ Challenge. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 553–568.
  43. 43.Hermann Ney, Ute Essen, and Reinhard Kneser. 1994. On structuring probabilistic dependences in stochastic language modelling. Computer Speech & Language 8, 1 (1994), 1 – 38. https://doi.org/10.1006/csla.1994.1001
  44. 44.Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain J Marshall, Ani Nenkova, and Byron C Wallace. 2018. A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature. In Proceedings of the conference. Association for Computational Linguistics. Meeting, Vol. 2018. NIH Public Access, 197.
  45. 45.Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer Learning in Biomedical Natural Language Processing: An Evaluation of BERT and ELMo on Ten Benchmarking Datasets. In Proceedings of the 18th BioNLP Workshop and Shared Task. Association for Computational Linguistics, Florence, Italy, 58–65. https://doi.org/10.18653/v1/W19-5006
  46. 46.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 1532–1543.
  47. 47.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). Association for Computational Linguistics, New Orleans, Louisiana, 2227–2237. https://doi.org/10.18653/v1/N18-1202
  48. 48.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  49. 49.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.
  50. 50.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. http://jmlr.org/papers/v21/20-074.html
  51. 51.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural Machine Translation of Rare Words with Subword Units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Berlin, Germany, 1715–1725. https://doi.org/10.18653/v1/P16-1162
  52. 52.Yuqi Si, Jingqi Wang, Hua Xu, and Kirk Roberts. 2019. Enhancing clinical concept extraction with contextual embeddings. Journal of the American Medical Informatics Association (2019).
  53. 53.Larry Smith, Lorraine K Tanabe, Rie Johnson nee Ando, Cheng-Ju Kuo, I-Fang Chung, Chun-Nan Hsu, Yu-Shi Lin, Roman Klinger, Christoph M Friedrich, et al. 2008. Overview of BioCreative II gene mention recognition. Genome biology 9, S2 (2008), S2.
  54. 54.Gizem Soğancıoğlu, Hakime Öztürk, and Arzucan Özgür. 2017. BIOSSES: a semantic sentence similarity estimation system for the biomedical domain. Bioinformatics 33, 14 (2017), i49–i58.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998–6008.
  56. 56.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems. 3266–3280.
  57. 57.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A MULTI-TASK BENCHMARK AND ANALYSIS PLATFORM FOR NATURAL LANGUAGE UNDERSTANDING. In ICLR.
  58. 58.Hai Wang and Hoifung Poon. 2018. Deep Probabilistic Logic: A Unifying Framework for Indirect Supervision. In EMNLP.
  59. 59.Yichong Xu, Xiaodong Liu, Yelong Shen, Jingjing Liu, and Jianfeng Gao. 2019. Multi-task Learning with Sample Re-weighting for Machine Reading Comprehension. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 2644–2655. https://doi.org/10.18653/v1/N19-1271
  60. 60.M. Zhang and Z. Zhou. 2014. A Review on Multi-Label Learning Algorithms. IEEE Transactions on Knowledge and Data Engineering 26, 8 (2014), 1819–1837. https://doi.org/10.1109/TKDE.2013.39
  61. 61.Yijia Zhang, Wei Zheng, Hongfei Lin, Jian Wang, Zhihao Yang, and Michel Dumontier. 2018. Drug–drug interaction extraction via hierarchical RNNs on sequence and shortest dependency paths. Bioinformatics 34, 5 (2018), 828–835.
  62. 62.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In ICCV.

Citation

MLA
Gu, Y., et al. “Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing”. ACM Transactions on Computing for Healthcare, vol. 3, no. 1, 2021, pp. 1–3, https://doi.org/10.1145/3458754.
APA
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., & Poon, H. (2021). Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23. https://doi.org/10.1145/3458754
Chicago
Gu, Y., R. Tinn, H. Cheng, et al. 2021. “Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing”. ACM Transactions on Computing for Healthcare 3 (1): 1–23. https://doi.org/10.1145/3458754.
Harvard
Gu, Y. et al. (2021) “Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing”, ACM Transactions on Computing for Healthcare, 3(1), pp. 1–23. Available at: https://doi.org/10.1145/3458754.
Vancouver
1. Gu Y, Tinn R, Cheng H, Lucas M, Usuyama N, Liu X, Naumann T, Gao J, Poon H (2021) Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Transactions on Computing for Healthcare 3:1–23

BibTeX

@article{Gu_2021, title={Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing}, volume={3}, ISSN={2637-8051}, url={http://dx.doi.org/10.1145/3458754}, DOI={10.1145/3458754}, number={1}, journal={ACM Transactions on Computing for Healthcare}, publisher={Association for Computing Machinery (ACM)}, author={Gu, Yu and Tinn, Robert and Cheng, Hao and Lucas, Michael and Usuyama, Naoto and Liu, Xiaodong and Naumann, Tristan and Gao, Jianfeng and Poon, Hoifung}, year={2021}, month=Oct, pages={1–23} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF