DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

Yanis LabrakAdrien BazogeRichard DufourMickael RouvierEmmanuel MorinBéatrice DaillePierre-Antoine Gourraud

article2023ACL64 citationsOutstanding Paper Award

Presents DrBERT, the first open-source French biomedical language model, alongside the billion-word NACHOS corpus and an empirical comparison demonstrating that public web data can match the performance of restricted hospital records on medical NLP tasks.

Listen

Specialized artificial intelligence systems for healthcare natural language processing have advanced rapidly in English, but non-English languages face severe resource shortages. Clinical text is particularly sensitive and strictly regulated to protect patient privacy, making it difficult to collect large-scale training data. Consequently, healthcare institutions and researchers working in French have historically relied either on generic French language models or on specialized models trained exclusively in English, neither of which is optimized for French clinical and biomedical tasks.

The article demonstrates the development, performance, and public release of DrBERT, the first specialized language models for the French biomedical and healthcare domains. It also evaluates how different training strategies, data sources, and dataset sizes affect model performance across specialized medical tasks.

The authors compiled two primary datasets to evaluate and train these systems: NACHOS, a newly gathered public web corpus spanning 1.1 billion words across 24 French medical sources (7.4 gigabytes), and a private clinical dataset comprising 655 million words extracted from 1.7 million de-identified hospital stay reports from Nantes University Hospital (4.0 gigabytes). Using these resources, the team trained 110-million-parameter models from scratch and evaluated continuous adaptation strategies using existing generic French systems and specialized English systems. Models were benchmarked across a suite of public and private medical tasks—including clinical entity recognition, multiple-choice question answering, and hospital report classification—as well as general-domain natural language tasks.

The article establishes several key findings. First, specialized French models trained from scratch on domain-specific medical data consistently achieved top performance across both private hospital tasks and public medical benchmarks, outperforming generic baseline models. Second, models trained purely on public web data (NACHOS) performed comparably to, and often outperformed, models trained directly on confidential hospital records. Third, doubling public pre-training data from 4.0 to 7.4 gigabytes yielded modest improvements, indicating that moderate, high-quality domain corpora are sufficient to achieve competitive performance. Fourth, adapting an English biomedical model (PubMedBERT) using French medical data proved significantly more effective than adapting an existing general French model (CamemBERT), demonstrating viable cross-language domain transfer. Finally, all specialized medical models showed degraded performance on non-medical tasks, confirming that domain specialization comes at the expense of general-purpose utility.

These findings have major operational and cost implications for healthcare technology development. Because publicly scraped biomedical data matches or exceeds the utility of private clinical records, organizations can build robust healthcare language tools without incurring the legal overhead, compliance risks, and privacy hurdles associated with handling sensitive patient records. Furthermore, organizations lacking the resources to train large models from scratch can cost-effectively adapt pre-existing foreign-language medical models to their local language.

Decision-makers and research teams should deploy the open-source DrBERT models and NACHOS corpus for French biomedical text analysis rather than relying on generic models. For organizations operating in other non-English low-resource settings, teams should prioritize continual adaptation of existing English medical models rather than generalist models in the target language. Future efforts should explore how tokenizer configuration influences downstream accuracy and test whether cross-language medical transfer strategies remain effective across diverse language families.

Readers should interpret the findings in light of specific limitations. The experiments required substantial computational overhead, consuming roughly 25,500 graphics processing unit hours on a high-performance cluster. In addition, the private clinical dataset was sourced from a single hospital system, meaning findings may not fully capture variations in clinical documentation across other institutions. Nevertheless, because reported metrics reflect averaged scores across repeated experimental runs, confidence remains high in the relative superiority of domain-specific pre-training for biomedical applications.

arXiv: 2304.00958
Cover for DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

Abstract

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more effectively treat specific domains. In this paper, we propose an original study of PLMs in the medical domain on French language. We compare, for the first time, the performance of PLMs trained on both public data from the web and private data from healthcare establishments. We also evaluate different learning strategies on a set of biomedical tasks. In particular, we show that we can take advantage of already existing biomedical PLMs in a foreign language by further pre-train it on our targeted data. Finally, we release the first specialized PLMs for the biomedical field in French, called DrBERT, as well as the largest corpus of medical data under free license on which these models are trained.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Pre-training datasets
  • 3.1 Public corpus - NACHOS
  • 3.2 Private corpus - NBDW
  • 3.3 Pre-processing step
  • 4 Models pre-training
  • 4.1 Influence of data
  • 4.2 Pre-training strategies
  • 4.3 Baseline models
  • 5 Downstream evaluation tasks
  • 5.1 Publicly-available tasks
  • 5.2 Private tasks
  • 6 Results and Discussions
  • 6.1 Impact of pre-training strategies
  • 6.2 Effect of data
  • 6.3 Performance on general-domain tasks
  • 7 Conclusion
  • 8 Ethical considerations
  • 9 Limitations
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Vocabularies Inter-coverage
  • A.2 Models Stability

Knowls

  1. Knowl 1 — DrBERT and ChuBERT Architecture and Pre-training Configuration

    model/method

    DrBERT and ChuBERT are French Transformer-based masked language models based on the RoBERTa-base architecture (CamemBERT base configuration), comprising 12 Transformer layers, a hidden dimension of 768, 12 attention heads, and 110 million parameters.

    • Tokenization: Text is tokenized into subword units using SentencePiece (an extension of Byte-Pair Encoding and WordPiece requiring no language-specific pre-tokenization) with a vocabulary size of 32,000 tokens. For models trained from scratch, the tokenizer is trained directly on the target pre-training corpus.
    • Objective: Masked Language Modeling (MLM) with a 15% token masking rate (where 80% of selected tokens are replaced by <mask>, 10% remain unchanged, and 10% are replaced with random vocabulary tokens).
    • Optimization: Training runs for 80,000 steps with a global batch size of 4,096 sequences of length 512 tokens (~2.1 million tokens per step). Learning rate warms up linearly for 10,000 steps from 0 to 5×10−55 \times 10^{-5}. Training employs 16-bit mixed precision (FP16) on 128 Nvidia V100 32 GB GPUs for 20 hours (batch size of 32 sequences per GPU).
  2. Knowl 2 — NACHOS Open French Healthcare Corpus

    data/table

    The opeN crAwled frenCh Health-care cOrpuS (NACHOS) is an open-source French biomedical dataset gathered from 24 online sources covering clinical cases, drug leaflets, institutional reports, theses, university courses, and scientific literature. The corpus undergoes sentence splitting, aggressive filtering of short or low-quality sentences (including OCR noise), and language filtering via a classifier trained on Opus EMEA and MASSIVE to retain only French sentences.

    Resource Name # Words
    HAL 638,508,261
    Haute Autorité de Santé (HAS) 113,394,539
    Drug leaflets 74,770,229
    Medical Websites Scrapping 64,904,334
    ANSES SAISINE 51,372,932
    Public Drug Database (BDPM) 48,302,695
    ISTEX 44,124,422
    CRTT 26,210,756
    WMT-16 10,282,494
    EMEA-V3 6,601,617
    Wikipedia Life Science French 4,671,944
    ANSES RCP 2,953,045
    Cerimes 1,717,552
    LiSSa 235,838
    DEFT-2020 231,396
    CLEAR 225,898
    CNEDiMTS 175,416
    QUAERO French Medical Corpus 72,031
    ANSM Clinical Study Registry 47,678
    ECDC 44,482
    QualiScope 12,718
    WMT-18-Medline 7,673
    Total 1,088,867,950

    The full corpus (NACHOSlarge\text{NACHOS}_{large}) contains 7.4 GB (1.089 billion words, 54.2 million sentences). A 4 GB sub-corpus (NACHOSsmall\text{NACHOS}_{small}) is created by randomly selecting 25.3 million sentences (646 million words) after shuffling the full dataset.

  3. Knowl 3 — Nantes Biomedical Data Warehouse (NBDW) Clinical Corpus

    experimental setup

    The Nantes Biomedical Data Warehouse (NBDW) dataset comprises 1.7 million de-identified French hospital stay reports from Nantes University Hospital, extracted under French CNIL authorization N°2129203.

    The corpus totals approximately 4 GB (NBDWsmall\text{NBDW}_{small}), containing 655,055,061 words across 43,115,081 sentence sequences (average sequence length 15.26 words). When combined with NACHOSsmall\text{NACHOS}_{small}, the resulting mixed corpus (NBDWmixed\text{NBDW}_{mixed}) contains 8 GB (1.3 billion words across 68.4 million sentences).

    The reports span 21 clinical departments, with the highest concentrations from:

    • Other departments: 474,588 documents (192,832,792 words)
    • Emergency Medicine: 235,579 documents (90,807,406 words)
    • Ambulatory Care: 119,149 documents (50,975,472 words)
    • Consultation: 95,135 documents (38,335,804 words)
    • Gynecology: 132,983 documents (38,204,495 words)
    • Cardiology: 29,633 documents (22,654,583 words)
    • Medical Oncology: 45,603 documents (22,587,869 words)
    • Gastroenterology: 46,600 documents (21,340,794 words)
    • Orthopaedic Surgery: 82,084 documents (18,983,791 words)
    • Hematology: 41,776 documents (18,285,983 words)
    • Critical Care Medicine, Otolaryngology, Dermatology, Rheumatology, Urology, Colon & Rectal Surgery, Internal Medicine, Psychiatry, Neurosurgery, Nephrology, and Ophthalmology (ranging from 16.5M words down to 4.5M words).
  4. Knowl 4 — French Biomedical NLP Downstream Benchmark Tasks

    experimental setup

    To evaluate French language models on clinical and biomedical domains, a benchmark of 7 public and 4 private tasks is established:

    Public Tasks:

    • ESSAIS POS Tagging: Part-of-speech tagging on French clinical cases (Macro F1; Train: 9,693, Dev: 2,077, Test: 2,078).
    • CAS POS Tagging: Part-of-speech tagging on French clinical cases (Macro F1; Train: 5,306, Dev: 1,137, Test: 1,137).
    • MUSCA-DET Task 1 (NER): Nested named entity recognition for 26 Social Determinants of Health entities from hospital notes (Macro F1; Train: 19,861, Dev: 2,207, Test: 5,518).
    • MUSCA-DET Task 2 (Classification): Multi-label classification of Social Determinants of Health categories (Macro F1; Train: 19,861, Dev: 2,207, Test: 5,518).
    • QUAERO-EMEA NER: Nested entity recognition across 10 UMLS semantic groups on European Medicines Agency documents (Weighted F1; Train: 11, Dev: 12, Test: 15).
    • QUAERO-MEDLINE NER: Nested entity recognition across 10 UMLS semantic groups on French MEDLINE abstracts (Weighted F1; Train: 833, Dev: 832, Test: 833).
    • FrenchMedMCQA: Multiple-choice question answering from French pharmacy specialization diploma exams (Evaluated via Exact Match Ratio and Hamming Score; Train: 2,171, Dev: 312, Test: 622).

    Private Tasks:

    • Acute Heart Failure (aHF) NER: Recognition of 46 clinical entity types in hospital stay reports (Macro F1; Train: 2,527, Dev: 281, Test: 703).
    • Acute Heart Failure (aHF) Classification: Binary classification for the presence of acute heart failure diagnosis (Macro F1; Train: 1,179, Dev: 132, Test: 328).
    • Technical Specialties Sorting: 6-class document classification across psychiatry, urology, endocrinology, cardiology, diabetology, infectiology (Macro F1; Train: 4,413, Dev: 1,470, Test: 1,473).
    • Prescriptions NER: 12-class named entity recognition on transcribed spoken medical reports (Macro F1; Train: 61, Dev: 15, Test: 26).
  5. Knowl 5 — Performance on Private French Clinical Downstream Tasks

    empirical result

    Language models pre-trained from scratch on in-domain French clinical or web biomedical text achieve the highest performance across all private downstream clinical evaluations (averaged over 4 fine-tuning runs):

    Model aHF NER aHF classification NER Med. Report Specialties Classif.
    P R F1 P R F1 P R F1 P R F1
    CamemBERT 138 GB 40.89 35.22 35.13 81.90 79.12 80.13 87.98 91.66 89.35 99.32 99.09 99.20
    CamemBERT 4 GB 46.32 43.17 42.66 81.49 81.42 81.41 87.79 90.74 88.78 99.53 99.69 99.61
    CamemBERT CCNET 4 GB 47.25 42.20 43.11 82.02 79.30 79.98 87.61 92.28 89.34 99.54 99.55 99.55
    PubMedBERT 52.61 46.30 47.22 78.17 76.18 76.86 87.07 92.61 89.20 99.25 99.51 99.37
    ClinicalBERT 50.11 44.15 44.70 80.13 75.92 77.12 87.04 92.14 88.77 98.58 98.62 98.58
    BioBERT v1.1 49.37 47.25 46.01 79.69 78.51 79.00 88.17 91.80 89.38 98.59 99.03 98.80
    DrBERT NACHOSlarge\text{DrBERT NACHOS}_{large} 55.29 46.66 48.22 81.33 81.25 81.25 87.99 92.80 89.83 99.82 99.90 99.86
    DrBERT NACHOSsmall\text{DrBERT NACHOS}_{small} 54.55 43.39 45.93 79.85 80.10 79.87 87.57 92.76 89.44 99.85 99.85 99.85
    ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small} 56.92 47.46 49.01 81.03 82.67 81.56 87.76 92.63 89.58 99.76 99.90 99.83
    ChuBERT NBDWmixed\text{ChuBERT NBDW}_{mixed} 54.62 47.81 49.14 82.23 81.71 81.98 87.42 92.36 89.30 99.81 99.82 99.81
    CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small} 22.02 16.67 16.08 74.86 69.82 69.80 65.72 68.49 66.74 99.44 99.67 99.54
    PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} 53.44 48.21 48.72 83.06 80.39 81.40 87.35 92.69 89.36 99.52 99.58 99.55
    CamemBERT NBDWsmall\text{CamemBERT NBDW}_{small} 25.44 19.33 19.12 79.50 74.74 76.02 68.80 71.23 69.64 99.60 99.57 99.58

    Models pre-trained from scratch on clinical data (ChuBERT NBDWmixed\text{ChuBERT NBDW}_{mixed}) achieve top F1 scores on aHF NER (49.14%) and aHF classification (81.98%), while the model pre-trained from scratch on public web data (DrBERT NACHOSlarge\text{DrBERT NACHOS}_{large}) performs best on prescription extraction (89.83% F1) and specialties sorting (99.86% F1). Continual pre-training from general French CamemBERT fails severely on these tasks (F1 < 20% on aHF NER).

  6. Knowl 6 — Performance on Public French Biomedical Downstream Tasks

    empirical result

    Evaluation of pre-trained language models across public French biomedical tasks demonstrates that web-pre-trained domain models (DrBERT) and cross-lingual domain-transferred models (PubMedBERT continued on French text) achieve state-of-the-art results (averaged over 4 runs):

    Model MUSCA T1 MUSCA T2 ESSAIS CAS FrMedMCQA Q-E Q-M
    P / R F1 P / R F1 F1 F1 Hamm. EMR F1 F1
    CamemBERT 138 GB 89.0/88.6 88.54 89.9/87.1 88.20 81.10 95.22 36.24 16.55 90.71 77.41
    CamemBERT 4 GB 86.1/85.5 85.43 92.7/90.3 91.27 83.69 96.42 35.75 15.37 90.83 78.76
    CamemBERT CCNET 4GB 91.1/89.9 90.33 93.1/90.4 91.38 85.42 97.33 34.71 14.41 90.33 77.61
    PubMedBERT 93.0/91.5 91.99 84.4/80.6 81.97 87.78 95.90 33.98 14.14 86.79 77.09
    ClinicalBERT 91.8/89.4 90.36 85.4/81.2 82.95 88.24 96.73 32.78 14.19 84.79 75.05
    BioBERT v1.1 91.8/89.8 90.46 85.5/80.1 81.91 85.18 97.12 36.19 15.43 84.29 72.68
    DrBERT NACHOSlarge\text{DrBERT NACHOS}_{large} 92.1/90.3 91.04 95.0/90.4 92.24 89.75 95.65 36.66 15.32 92.09 77.88
    DrBERT NACHOSsmall\text{DrBERT NACHOS}_{small} 93.4/90.6 91.77 91.3/86.6 88.57 88.76 95.70 37.37 13.34 91.66 78.18
    ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small} 94.9/90.8 92.23 94.8/90.3 92.17 87.71 95.61 35.16 14.79 88.15 74.94
    ChuBERT NBDWmixed\text{ChuBERT NBDW}_{mixed} 94.4/91.9 92.73 94.2/90.0 91.71 85.73 96.35 34.58 12.21 90.52 78.63
    CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small} 81.4/81.4 80.96 79.7/78.1 78.70 80.04 92.46 32.87 13.76 71.10 57.43
    PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} 92.5/91.5 91.53 95.0/92.6 93.62 83.85 96.81 35.88 15.21 91.03 81.73
    CamemBERT NBDWsmall\text{CamemBERT NBDW}_{small} 82.4/81.6 81.57 78.1/76.4 77.12 79.25 93.18 27.73 11.89 61.75 53.05

    (MUSCA T1/T2: MUSCA-DET Tasks 1 and 2; FrMedMCQA: FrenchMedMCQA; Q-E: QUAERO-EMEA; Q-M: QUAERO-MEDLINE). DrBERT\text{DrBERT} models achieve top F1 on ESSAIS (89.75%), QUAERO-EMEA (92.09%), and FrenchMedMCQA Hamming score (37.37%), while PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} achieves top F1 on MUSCA-DET T2 (93.62%) and QUAERO-MEDLINE (81.73%).

  7. Knowl 7 — Cross-Language Continual Pre-Training Outperforms In-Language Generic Continual Pre-Training

    empirical result

    Continually pre-training an English domain-specific biomedical language model (PubMedBERT) on French biomedical data (NACHOSsmall\text{NACHOS}_{small}) yields substantially better downstream performance on French biomedical tasks than continually pre-training a French generic language model (CamemBERT) on the exact same domain data:

    • On the private Acute Heart Failure NER task, PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} reaches 48.72% F1 versus 16.08% for CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small}.
    • On the private Acute Heart Failure classification task, PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} achieves 81.40% F1 versus 69.80% for CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small}.
    • On the public MUSCA-DET Task 2, PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} achieves 93.62% F1 versus 78.70% for CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small}.
    • On QUAERO-MEDLINE nested NER, PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} achieves 81.73% F1 versus 57.43% for CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small}.

    This confirms that cross-lingual transfer of specialized biomedical representations (from English PubMedBERT to French) is more effective than domain adaptation from a general-language model in the target language, due to the high density of shared technical medical vocabulary across languages.

  8. Knowl 8 — Biomedical Specialization Induces Degradation on General-Domain NLP Tasks

    empirical result

    Adapting or pre-training language models specifically on French biomedical and clinical text results in degraded performance on general-domain French NLP benchmarks (evaluated on POS tagging across Universal Dependencies GSD, SEQUOIA, SPOKEN, PARTUT treebanks, and Natural Language Inference on XNLI):

    Model GSD SEQUOIA SPOKEN PARTUT XNLI
    CamemBERT OSCAR 138 GB 98.28 98.68 97.26 97.70 81.94
    CamemBERT OSCAR 4 GB 98.14 99.18 97.57 97.86 81.76
    CamemBERT CCNET 4 GB 98.18 98.92 97.20 97.92 81.26
    PubMedBERT 96.48 96.49 90.00 93.97 73.79
    ClinicalBERT 96.49 96.31 89.60 93.17 70.57
    BioBERT v1.1 97.32 96.54 91.81 94.52 71.54
    DrBERT NACHOSlarge\text{DrBERT NACHOS}_{large} 96.94 98.05 95.92 96.54 72.18
    DrBERT NACHOSsmall\text{DrBERT NACHOS}_{small} 97.17 98.21 96.38 96.45 72.86
    ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small} 96.45 97.38 94.90 95.83 69.00
    ChuBERT NBDWmixed\text{ChuBERT NBDW}_{mixed} 97.18 98.10 96.43 96.33 72.32
    CamemBERT NACHOSsmall\text{CamemBERT NACHOS}_{small} 97.63 96.90 91.12 94.00 71.26
    PubMedBERT NACHOSsmall\text{PubMedBERT NACHOS}_{small} 97.41 98.71 95.54 97.01 77.35
    CamemBERT NBDWsmall\text{CamemBERT NBDW}_{small} 97.55 96.26 89.17 91.34 72.73

    The performance drop is most notable on complex reasoning tasks: on XNLI, accuracy drops from 81.94% (CamemBERT 138 GB) to 72.18% for DrBERT NACHOSlarge\text{DrBERT NACHOS}_{large} and 69.00% for ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small} (an absolute loss of 12.94%).

  9. Knowl 9 — Efficacy of Public Web Data Versus Private Clinical Data for Medical Pre-training

    empirical result

    Language models pre-trained from scratch on public web-crawled biomedical data (DrBERT\text{DrBERT}) perform on par with or exceed models pre-trained on restricted private clinical health records (ChuBERT\text{ChuBERT}):

    • Even when restricted to equal corpus sizes (4 GB), DrBERT NACHOSsmall\text{DrBERT NACHOS}_{small} achieves comparable performance to ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small} across private clinical tasks (e.g., 89.44% vs. 89.58% F1 on prescription structuration; 99.85% vs. 99.83% F1 on specialty classification).
    • On public biomedical tasks, web-trained models substantially outperform private-trained models (e.g., on QUAERO-MEDLINE, DrBERT NACHOSsmall\text{DrBERT NACHOS}_{small} achieves 78.18% F1 vs. 74.94% for ChuBERT NBDWsmall\text{ChuBERT NBDW}_{small}).
    • Combining public and private corpora (ChuBERT NBDWmixed\text{ChuBERT NBDW}_{mixed}, 8 GB) provides modest performance boosts on private tasks (e.g., 49.14% F1 on aHF NER vs. 49.01% for NBDWsmall\text{NBDW}_{small} and 48.22% for NACHOSlarge\text{NACHOS}_{large}) but shows that web-only data is sufficient to build competitive medical models without requiring access to protected health information.
  10. Knowl 10 — Continual Pre-training Instability and Optimization Cost

    limitation

    Several limitations and instability behaviors are identified during medical model pre-training:

    • Continual Pre-training Instability: Continual pre-training initialized from CamemBERT OSCAR 138 GB exhibits poor consistency and high variance during fine-tuning across runs on tasks such as aHF classification, MUSCA-DET T1, QUAERO MEDLINE, and XNLI.
    • Sudden MLM Loss Drop: During continual pre-training of PubMedBERT on NACHOSsmall\text{NACHOS}_{small}, training loss remains stable until step 71,000, then abruptly drops toward zero by step 72,500.
    • Computational Budget: Pre-training 7 model configurations required approximately 18,000 GPU computation hours plus 7,500 hours for debugging and configuration resolution (totaling 25,500 GPU hours on Jean Zay), equivalent to 6,604.5 kWh and 376.45 kg CO2eq\text{CO}_2\text{eq}.

Coverage note — None was omitted.

References

  1. 1.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: Pretrained language model for scientific text. In EMNLP.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT '21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  3. 3.Olivier Bodenreider. 2004. The unified medical language system (umls): Integrating biomedical terminology. Nucleic acids research, 32:D267–70.
  4. 4.Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño, Joan Llop-Palao, Marc Pàmies, Aitor Gonzalez-Agirre, and Marta Villegas. 2021. Biomedical and clinical language models for spanish: On the benefits of domain-specific pretraining in a mid-resource scenario.
  5. 5.Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020. LEGAL-BERT: the muppets straight out of law school. CoRR, abs/2010.02559.
  6. 6.Clément Dalloux, Vincent Claveau, Natalia Grabar, Lucas Emanuel Silva Oliveira, Claudia Maria Cabral Moro, Yohan Bonescki Gumiel, and Deborah Ribeiro Carvalho. 2021. Supervised learning for the detection of negation and of its scope in French and Brazilian Portuguese biomedical corpora. Natural Language Engineering, 27(2):181–201.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. CoRR, abs/1810.04805.
  8. 8.Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, and Pierre Zweigenbaum. 2022. Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical Domain. In Proceedings of the 13th Language Resources and Evaluation Conference (LREC), pages 2626–2633, Marseille, France. European Language Resources Association.
  9. 9.Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2022. Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages.
  10. 10.Natalia Grabar, Vincent Claveau, and Clément Dalloux. 2018. CAS: French Corpus with Clinical Cases. In Proceedings of the 9th International Workshop on Health Text Mining and Information Analysis (LOUHI), pages 1–7, Brussels, Belgium.
  11. 11.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Transactions on Computing for Healthcare, 3(1):1–23.
  12. 12.Jeremy Howard and Sebastian Ruder. 2018. Universal Language Model Fine-tuning for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), pages 328–339, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019. Clinicalbert: Modeling clinical notes and predicting hospital readmission.
  14. 14.Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. CoRR, abs/1808.06226.
  15. 15.Yanis Labrak, Adrien Bazoge, Richard Dufour, Béatrice Daille, Pierre-Antoine Gourraud, Emmanuel Morin, and Mickael Rouvier. 2022. FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain. In Proceedings of the 13th International Workshop on Health Text Mining and Information Analysis (LOUHI), Abou Dhabi, United Arab Emirates.
  16. 16.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics.
  17. 17.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  18. 18.Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2022. Estimating the carbon footprint of bloom, a 176b parameter language model.
  19. 19.Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020. CamemBERT: a Tasty French Language Model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7203—-7219. Association for Computational Linguistics.
  20. 20.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2017. Mixed precision training. CoRR, abs/1710.03740.
  21. 21.Aurélie Névéol, Cyril Grouin, Jérémy Leixa, Sophie Rosset, and Pierre Zweigenbaum. 2014. The quaero french medical corpus : A ressource for medical entity recognition and normalization.
  22. 22.Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets. In Proceedings of the 2019 Workshop on Biomedical Natural Language Processing (BioNLP 2019), pages 58–65.
  23. 23.Elisa Terumi Rubel Schneider, João Vitor Andrioli de Souza, Julien Knafou, Lucas Emanuel Silva e Oliveira, Jenny Copara, Yohan Bonescki Gumiel, Lucas Ferro Antunes de Oliveira, Emerson Cabrera Paraiso, Douglas Teodoro, and Cláudia Maria Cabral Moro Barra. 2020. BioBERTpt - a Portuguese neural language model for clinical named entity recognition. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 65–72, Online. Association for Computational Linguistics.
  24. 24.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  25. 25.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2021. Societal biases in language generation: Progress and challenges.
  26. 26.Manjil Shrestha. 2021. Development of a language model for medical domain. masterthesis, Hochschule Rhein-Waal.
  27. 27.Jörg Tiedemann and Lars Nygaard. 2004. The OPUS corpus - parallel and free: http://logos.uio.no/opus. In Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC'04), Lisbon, Portugal. European Language Resources Association (ELRA).
  28. 28.Hazal Türkmen, Oguz Dikenelli, Cenk Eraslan, and ˘ Mehmet Callı. 2022. Bioberturk: Exploring turkish biomedical language model development strategies in low resource setting.
  29. 29.Thomas Vakili, Anastasios Lamproudis, Aron Henriksson, and Hercules Dalianis. 2022. Downstream task performance of BERT models pre-trained using automatically de-identified clinical data. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4245–4252, Marseille, France. European Language Resources Association.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.
  31. 31.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019. Huggingface's transformers: State-of-the-art natural language processing.
  32. 32.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016. Google's neural machine translation system: Bridging the gap between human and machine translation.
  33. 33.Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E. Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B. Costa, Mona G. Flores, Ying Zhang, Tanja Magoc, Christopher A. Harle, Gloria Lipori, Duane A. Mitchell, William R. Hogan, Elizabeth A. Shenkman, Jiang Bian, and Yonghui Wu. 2022. A large language model for electronic health records. npj Digital Medicine, 5(1):194.
  34. 34.Yi Yang, Mark Christopher Siy UY, and Allen Huang. 2020. Finbert: A pretrained language model for financial communications.
  35. 35.Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman. 2020. When do you need billions of words of pretraining data?
  36. 36.Hongyin Zhu, Hao Peng, Zhiheng Lyu, Lei Hou, Juanzi Li, and Jinghui Xiao. 2021. Travelbert: Pre-training language model incorporating domain-specific heterogeneous knowledge into a unified representation.

Citation

MLA
Labrak, Y., et al. “DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical Domains”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 16207–21, https://doi.org/10.18653/v1/2023.acl-long.896.
APA
Labrak, Y., Bazoge, A., Dufour, R., Rouvier, M., Morin, E., Daille, B., & Gourraud, P.-A. (2023). DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16207–16221. https://doi.org/10.18653/v1/2023.acl-long.896
Chicago
Labrak, Y., A. Bazoge, R. Dufour, et al. 2023. “DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical Domains”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16207–21. https://doi.org/10.18653/v1/2023.acl-long.896.
Harvard
Labrak, Y. et al. (2023) “DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16207–16221. Available at: https://doi.org/10.18653/v1/2023.acl-long.896.
Vancouver
1. Labrak Y, Bazoge A, Dufour R, Rouvier M, Morin E, Daille B, Gourraud P-A (2023) DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16207–16221

BibTeX

@inproceedings{labrak-etal-2023-drbert,
    title = "{D}r{BERT}: A Robust Pre-trained Model in {F}rench for Biomedical and Clinical domains",
    author = "Labrak, Yanis  and
      Bazoge, Adrien  and
      Dufour, Richard  and
      Rouvier, Mickael  and
      Morin, Emmanuel  and
      Daille, B{\'e}atrice  and
      Gourraud, Pierre-Antoine",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.896/",
    doi = "10.18653/v1/2023.acl-long.896",
    pages = "16207--16221"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/