keyword
biomedical natural language processing
Biomedical natural language processing is a specialized subfield of artificial intelligence and computational linguistics focused on developing automated methods to analyze, interpret, and extract meaningful information from human language found in biology, medicine, and healthcare. Because texts in these domains, such as scientific research articles, electronic health records, and clinical reports, are characterized by highly specialized vocabularies, complex terminology, and frequent abbreviations, standard language processing techniques often require domain-specific adaptation and specialized pretraining to achieve accurate comprehension. Common tasks within this field include named entity recognition for identifying concepts like genes, drugs, and diseases, relation extraction to map biological interactions, clinical text classification, and medical question answering, all aimed at accelerating scientific discovery and enhancing healthcare workflows.
2 items

DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains
Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud
Why you should read this
Presents DrBERT, the first open-source French biomedical language model, alongside the billion-word NACHOS corpus and an empirical comparison demonstrating that public web data can match the performance of restricted hospital records on medical NLP tasks.
In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains. In this paper, we propose an original study of PLMs in the medical domain on French language. We compare, for the first time, the performance of PLMs trained on both public data from the web and private data from healthcare establishments. We also evaluate different learning strategies on a set of biomedical tasks. In particular, we show that we can take advantage of already existing biomedical PLMs in a foreign language by further pre-train it on our targeted data. Finally, we release the first specialized PLMs for the biomedical field in French, called DrBERT, as well as the largest corpus of medical data under free license on which these models are trained.
Added
2026-09-26

Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, Hoifung Poon
Why you should read this
Demonstrates that pretraining language models from scratch on domain text substantially outperforms continual pretraining from general-domain models in biomedicine, while introducing the BLURB benchmark and new state-of-the-art models for biomedical NLP.
Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly-available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition (NER). To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at this https URL.
Added
2026-09-14
