DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains
Yanis LabrakAdrien BazogeRichard DufourMickael RouvierEmmanuel MorinBéatrice DaillePierre-Antoine Gourraud
Presents DrBERT, the first open-source French biomedical language model, alongside the billion-word NACHOS corpus and an empirical comparison demonstrating that public web data can match the performance of restricted hospital records on medical NLP tasks.
Specialized artificial intelligence systems for healthcare natural language processing have advanced rapidly in English, but non-English languages face severe resource shortages. Clinical text is particularly sensitive and strictly regulated to protect patient privacy, making it difficult to collect large-scale training data. Consequently, healthcare institutions and researchers working in French have historically relied either on generic French language models or on specialized models trained exclusively in English, neither of which is optimized for French clinical and biomedical tasks.
The article demonstrates the development, performance, and public release of DrBERT, the first specialized language models for the French biomedical and healthcare domains. It also evaluates how different training strategies, data sources, and dataset sizes affect model performance across specialized medical tasks.
The authors compiled two primary datasets to evaluate and train these systems: NACHOS, a newly gathered public web corpus spanning 1.1 billion words across 24 French medical sources (7.4 gigabytes), and a private clinical dataset comprising 655 million words extracted from 1.7 million de-identified hospital stay reports from Nantes University Hospital (4.0 gigabytes). Using these resources, the team trained 110-million-parameter models from scratch and evaluated continuous adaptation strategies using existing generic French systems and specialized English systems. Models were benchmarked across a suite of public and private medical tasks—including clinical entity recognition, multiple-choice question answering, and hospital report classification—as well as general-domain natural language tasks.
The article establishes several key findings. First, specialized French models trained from scratch on domain-specific medical data consistently achieved top performance across both private hospital tasks and public medical benchmarks, outperforming generic baseline models. Second, models trained purely on public web data (NACHOS) performed comparably to, and often outperformed, models trained directly on confidential hospital records. Third, doubling public pre-training data from 4.0 to 7.4 gigabytes yielded modest improvements, indicating that moderate, high-quality domain corpora are sufficient to achieve competitive performance. Fourth, adapting an English biomedical model (PubMedBERT) using French medical data proved significantly more effective than adapting an existing general French model (CamemBERT), demonstrating viable cross-language domain transfer. Finally, all specialized medical models showed degraded performance on non-medical tasks, confirming that domain specialization comes at the expense of general-purpose utility.
These findings have major operational and cost implications for healthcare technology development. Because publicly scraped biomedical data matches or exceeds the utility of private clinical records, organizations can build robust healthcare language tools without incurring the legal overhead, compliance risks, and privacy hurdles associated with handling sensitive patient records. Furthermore, organizations lacking the resources to train large models from scratch can cost-effectively adapt pre-existing foreign-language medical models to their local language.
Decision-makers and research teams should deploy the open-source DrBERT models and NACHOS corpus for French biomedical text analysis rather than relying on generic models. For organizations operating in other non-English low-resource settings, teams should prioritize continual adaptation of existing English medical models rather than generalist models in the target language. Future efforts should explore how tokenizer configuration influences downstream accuracy and test whether cross-language medical transfer strategies remain effective across diverse language families.
Readers should interpret the findings in light of specific limitations. The experiments required substantial computational overhead, consuming roughly 25,500 graphics processing unit hours on a high-performance cluster. In addition, the private clinical dataset was sourced from a single hospital system, meaning findings may not fully capture variations in clinical documentation across other institutions. Nevertheless, because reported metrics reflect averaged scores across repeated experimental runs, confidence remains high in the relative superiority of domain-specific pre-training for biomedical applications.
- Paper: Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing, Yu Gu et al. (2020). This paper establishes the foundational methodology of pre-training domain-specific BERT models entirely from scratch on biomedical text (PubMedBERT), which DrBERT adapts and evaluates directly for the French language.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). This seminal work introduced continual biomedical pre-training for BERT models, providing the conceptual framework for biomedical representation learning that DrBERT builds on and expands to non-English clinical domains.
- Paper: Publicly Available Clinical BERT Embeddings, Emily Alsentzer et al. (2019). This paper demonstrates how to adapt language models specifically to unstructured clinical records, laying the groundwork for DrBERT's exploration of clinical hospital corpora.
- Paper: ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission, Kexin Huang et al. (2019). It provides essential baseline methods for processing electronic health record notes with transformer models to predict clinical outcomes.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). It introduces the core bidirectional transformer pre-training architecture that underlies DrBERT and against which its domain adaptations are evaluated.
- Paper: RoBERTa: A Robustly Optimized BERT Pretraining Approach, Yinhan Liu et al. (2019). It establishes optimized pre-training recipes and masking strategies for encoder transformers that informed the training pipeline used in DrBERT.
- Paper: How to Fine-Tune BERT for Text Classification?, Chi Sun et al. (2019). It analyzes domain-adaptive pre-training and fine-tuning dynamics for BERT classifiers, offering empirical guidance utilized in DrBERT's transfer learning experiments.
- Paper: GottBERT: a pure German Language Model, Raphael Scheible et al. (2024). This work explores single-language domain and quality filtering for specialized German transformers, offering a complementary continuation of language-specific pre-training methodology.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). This benchmark systematically investigates data curation and quality filtering strategies across massive corpora, generalizing the data selection insights observed during DrBERT's NACHOS corpus curation.
