Generative Biomedical Entity Linking via Knowledge Base-Guided Pre-training and Synonyms-Aware Fine-tuning
Hongyi YuanZheng YuanSheng Yu
Proposes a generative biomedical entity linking framework that injects synonym and definition knowledge through synthetic pre-training and constrained prefix-tree decoding, achieving state-of-the-art accuracy across multiple benchmarks without candidate selection.
Biomedical entity linking maps biomedical mentions found in unstructured text to standardized medical concepts. This process is essential for downstream clinical informatics, including disease phenotyping, relation extraction, and automated diagnosis. Traditional approaches rely on dense embedding similarity search, which imposes massive memory requirements to store vector representations for millions of concepts. While generative sequence-to-sequence models reduce memory footprints by generating concept names directly, adapting them to the biomedical domain has been difficult due to the severe scarcity of annotated training data and the prevalence of diverse entity synonyms.
The article demonstrates a generative sequence-to-sequence framework specifically designed for biomedical entity linking. The primary objective is to evaluate whether injecting structured medical knowledge into generative pre-training and fine-tuning can improve linking accuracy without relying on large memory stores or separate candidate retrieval pipelines.
The researchers developed a 406-million-parameter encoder-decoder model initialized from BART-large. They addressed data scarcity by pre-training the model on synthetic natural language sentences built from 2.37 million concepts, definitions, and synonyms in the Unified Medical Language System. For fine-tuning, the approach pairs mentions with their most textually similar synonym using a character 3-gram similarity score and introduces decoder prompt prefixes. During inference, the system constrains output generation across a multi-synonym prefix tree covering all allowable target names, mapping predicted synonyms back to concept identifiers. The framework was evaluated on four standard biomedical and clinical datasets: BC5CDR, NCBI-disease, COMETA, and AskAPatient.
The experiments show that the proposed framework sets new performance benchmarks. First, the fully pre-trained and fine-tuned model achieved state-of-the-art recall@1 accuracy on BC5CDR (93.3%), COMETA (81.4%), and AskAPatient (89.3%), outperforming prior approaches by up to 1.4 percentage points while remaining highly competitive on NCBI-disease (91.9%). Second, knowledge-guided pre-training delivered consistent improvements across all datasets, boosting accuracy by 0.3 to 0.7 percentage points. Third, the method required substantially fewer computational resources for pre-training than existing general-domain generative baselines (6 GPU-days versus 32 GPU-days), while attaining superior transfer capability on medical tasks. Fourth, in sub-population analyses on BC5CDR, the model demonstrated strong zero-shot generalization, surpassing prior data-augmented baselines by nearly 10 percentage points on unseen concepts (86.9% versus 77.5%).
These results demonstrate that generative language models can effectively replace memory-intensive retrieval pipelines for medical entity normalization. By relying on synthetic templates constructed from existing ontologies rather than costly human annotations, the framework significantly lowers the computational and financial barriers required to train and deploy medical text-processing systems. Furthermore, using multi-synonym prefix trees eliminates the operational complexity and failure points of multi-stage candidate selection architectures.
Organizations developing clinical natural language processing pipelines should consider transitioning from dual-encoder similarity models to generative sequence-to-sequence architectures, especially when deploying in memory-constrained environments. For implementation, engineering teams should incorporate synonym-aware fine-tuning and natural language decoder prompts, which proved critical for performance. Future research should evaluate integrating refined candidate generation mechanisms to further boost decoding accuracy and explore scaling to larger foundation model backbones.
Confidence in these findings is high across standard disease, chemical, and colloquial social media benchmarks. However, stakeholders should note two key operational limitations: the generative model showed comparatively lower accuracy on multi-word mentions and mentions that did not directly match knowledge base entries. Additionally, performance remains susceptible to ambiguous or inconsistent entity annotations within the underlying medical knowledge bases.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). Introduces domain-adapted language modeling on PubMed corpora, establishing the foundational biomedical contextual representations utilized and adapted for generative entity linking.
- Paper: Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing, Yu Gu et al. (2020). Demonstrates the necessity of domain-specific biomedical pre-training from scratch and provides the standard BLURB benchmarks for biomedical entity extraction.
- Paper: ERNIE: Enhanced Language Representation with Informative Entities, Zhengyan Zhang et al. (2019). Pioneers the injection of structured knowledge base facts and entity auto-encoding objectives into pre-trained transformer architectures.
- Paper: GLM: General Language Model Pretraining with Autoregressive Blank Infilling, Zhengxiao Du et al. (2021). Establishes autoregressive blank infilling pre-training frameworks that formulate entity generation and name recovery directly within generative language models.
- Paper: BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining, Renqian Luo et al. (2022). Expands on generative biomedical modeling by training a full-scale generative pre-trained transformer specifically targeting end-to-end biomedical information extraction.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). Provides a comprehensive architectural roadmap for integrating structured knowledge bases with generative language models, contextualizing synonym- and KB-guided approaches.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). Surveys the broader sequence-to-sequence paradigm for generative knowledge graph construction and end-to-end entity generation methodologies.
- Paper: Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks, Minki Kang et al. (2023). Extends knowledge-intensive generative modeling by distilling knowledge-augmented reasoning into compact language models.
