Nested Named Entity Recognition with Span-level Graphs
Juncheng WanDongyu RuWeinan ZhangYong Yu
Proposes a retrieval-based span-level graph approach that connects candidate spans to training entities via n-gram similarity and applies graph convolutional networks to significantly boost nested named entity recognition performance, especially for low-frequency and long spans.
Extracting nested entities from unstructured text is critical for information extraction tasks, such as identifying overlapping locations, people, and specialized terms within sentences. Traditional span-based models struggle in nested scenarios because heavily overlapping text segments cause confusion. Furthermore, these models show weak generalization, as roughly 40% to 55% of entity mentions encountered during testing rarely or never appear in the training data.
The article evaluates whether integrating global, retrieval-based graph networks can enrich text span representations. The core objective is to demonstrate that connecting candidate text segments with lexically similar training entities improves recognition accuracy without relying on external syntactic tools or handcrafted rules.
The researchers developed a retrieval-based graph framework that connects text spans and entity mentions across the entire training dataset using word-level character sequence similarities. They evaluated the approach across three standard benchmark datasets: ACE2004, ACE2005, and GENIA. The architecture processes these graphs using two-layer Graph Convolutional Networks combined with attention mechanisms, pre-trained language models, and a multitask training objective that jointly classifies candidate spans and neighboring graph entities.
The key findings show consistent performance advantages over strong baseline models. First, the proposed approach achieved overall micro-F1 score improvements between 0.30 and 0.85 points across all benchmarks, reaching F1 scores of 86.31 on ACE2004, 85.11 on ACE2005, and 79.30 on GENIA. Second, the method significantly enhanced recall on low-frequency and unseen entities, improving recall by 0.56 to 2.56 points for rare terms. Third, the graph architecture proved especially beneficial for long entity spans of six or more words, delivering F1 improvements up to 13.11 points for eight-word spans by capturing informative lexical overlaps.
These results demonstrate that leveraging corpus-wide lexical connections provides critical contextual guidance when local sentence context is misleading or incomplete. Unlike complex parsing approaches that require external linguistic dependencies, this method uses existing training data to resolve ambiguous boundaries. This improves extraction accuracy in information-dense domains like biomedical text mining and intelligence analysis while avoiding manual feature engineering.
Organizations deploying automated information extraction systems should evaluate retrieval-augmented graph representations for complex and nested entity extraction pipelines. Teams adopting this framework must carefully balance operational trade-offs, as the graph structure reduces inference decoding throughput by roughly half compared to simpler span models, alongside a modest memory overhead of 100 to 500 megabytes. Future work should focus on optimizing retrieval speed and exploring broader deployment across multi-lingual datasets.
Confidence in these findings is supported by consistent gains across multiple benchmark datasets and detailed ablation experiments. However, decision-makers should note that the approach relies on the presence of informative lexical overlaps within the training set, meaning datasets with entirely disjoint vocabularies between training and operational environments may see more moderate benefits.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). This paper establishes how to construct corpus-level graphs and use two-layer Graph Convolutional Networks over text, providing the foundational graph formulation adapted by the source for span graphs.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This work introduces the foundational Graph Convolutional Network architecture that the source directly deploys to propagate representation signals across entity and span graph nodes.
- Paper: SpanBERT: Improving Pre-training by Representing and Predicting Spans, Mandar Joshi et al. (2019). This paper establishes span-level pre-training representations and boundary objectives, underpinning the span-based representations utilized in nested entity recognition architectures.
- Paper: A Survey on Deep Learning for Named Entity Recognition, Jing Li et al. (2018). This survey provides a comprehensive taxonomy of deep learning NER approaches, detailing the standard span and sequence architectures that the source aims to improve upon for nested entities.
- Paper: Incorporating Non-local Information into Information Extraction Systems by Gibbs Sampling, Jenny Rose Finkel et al. (2005). This study introduces early methods for incorporating non-local corpus and document-level consistency into information extraction, motivating the source's pursuit of global, corpus-wide entity connections.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). GLiNER generalizes span-based entity representation matching to open-type zero-shot scenarios, building beyond the fixed lexical graph retrieval frameworks used in nested NER.
- Paper: MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER, Ran Zhou et al. (2022). This work provides a complementary data augmentation approach to tackle unseen and low-resource entities in NER by conditioning masked language generation on explicit entity spans.
- Paper: Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning, Xiaoxin He et al. (2024). This paper extends text-attributed graph representation learning by integrating language model reasoning and explanations into graph neural network features.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). This work scales graph retrieval-augmented architectures by constructing hierarchical entity graphs across corpora for global information retrieval and summarization.
