Built independently by an author, for readers. Read the story and support ChapterPal

keyword

PubMed vocabulary

A PubMed vocabulary is a domain-specific set of words and subword tokens derived directly from biomedical literature in the PubMed database for use in natural language processing models. Unlike general-purpose vocabularies compiled from standard web text or literature, a PubMed vocabulary is tailored to capture complex medical, biological, and pharmacological terminology as distinct lexical units. Constructing a tokenizer vocabulary directly from biomedical text prevents specialized concepts, gene names, and clinical terms from being fragmented into generic or arbitrary subword pieces, thereby optimizing token efficiency and improving the ability of language models to comprehend and process scientific text.

1 item

Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, Hoifung Poon

OrganizationsMicrosoft

Why you should read this

Demonstrates that pretraining language models from scratch on domain text substantially outperforms continual pretraining from general-domain models in biomedicine, while introducing the BLURB benchmark and new state-of-the-art models for biomedical NLP.

Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly-available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition (NER). To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at this https URL.

Added

2026-09-14