Built independently by an author, for readers. Read the story and support ChapterPal

keyword

seq2seq entity linking

Seq2seq entity linking, also known as generative entity linking, is a natural language processing approach that formulates the task of connecting text mentions to entries in a knowledge base as a sequence-to-sequence text generation problem. Instead of relying on traditional multi-stage pipelines involving candidate retrieval and ranking, a seq2seq model takes an input text containing a mention and directly generates the canonical name or unique identifier of the target entity. This framework often incorporates constrained decoding mechanisms, such as trie-based prefix constraints, to guarantee that the generated sequence resolves to a valid entity within the reference knowledge base, thereby streamlining mention disambiguation into an end-to-end generative process.

1 item

Generative Biomedical Entity Linking via Knowledge Base-Guided Pre-training and Synonyms-Aware Fine-tuning

Generative Biomedical Entity Linking via Knowledge Base-Guided Pre-training and Synonyms-Aware Fine-tuning

Hongyi Yuan, Zheng Yuan, Sheng Yu

OrganizationsTsinghua University

Why you should read this

Proposes a generative biomedical entity linking framework that injects synonym and definition knowledge through synthetic pre-training and constrained prefix-tree decoding, achieving state-of-the-art accuracy across multiple benchmarks without candidate selection.

Entities lie in the heart of biomedical natural language understanding, and the biomedical entity linking (EL) task remains challenging due to the fine-grained and diversiform concept names. Generative methods achieve remarkable performances in general domain EL with less memory usage while requiring expensive pre-training. Previous biomedical EL methods leverage synonyms from knowledge bases (KB) which is not trivial to inject into a generative method. In this work, we use a generative approach to model biomedical EL and propose to inject synonyms knowledge in it. We propose KB-guided pre-training by constructing synthetic samples with synonyms and definitions from KB and require the model to recover concept names. We also propose synonyms-aware fine-tuning to select concept names for training, and propose decoder prompt and multi-synonyms constrained prefix tree for inference. Our method achieves state-of-the-art results on several biomedical EL tasks without candidate selection which displays the effectiveness of proposed pre-training and fine-tuning strategies. The source code is available at Github.com/Yuanhy1997/GenBioEL.

Added

2026-09-26