Nested Named Entity Recognition with Span-level Graphs

Juncheng WanDongyu RuWeinan ZhangYong Yu

article2022ACL78 citations

Proposes a retrieval-based span-level graph approach that connects candidate spans to training entities via n-gram similarity and applies graph convolutional networks to significantly boost nested named entity recognition performance, especially for low-frequency and long spans.

Listen

Extracting nested entities from unstructured text is critical for information extraction tasks, such as identifying overlapping locations, people, and specialized terms within sentences. Traditional span-based models struggle in nested scenarios because heavily overlapping text segments cause confusion. Furthermore, these models show weak generalization, as roughly 40% to 55% of entity mentions encountered during testing rarely or never appear in the training data.

The article evaluates whether integrating global, retrieval-based graph networks can enrich text span representations. The core objective is to demonstrate that connecting candidate text segments with lexically similar training entities improves recognition accuracy without relying on external syntactic tools or handcrafted rules.

The researchers developed a retrieval-based graph framework that connects text spans and entity mentions across the entire training dataset using word-level character sequence similarities. They evaluated the approach across three standard benchmark datasets: ACE2004, ACE2005, and GENIA. The architecture processes these graphs using two-layer Graph Convolutional Networks combined with attention mechanisms, pre-trained language models, and a multitask training objective that jointly classifies candidate spans and neighboring graph entities.

The key findings show consistent performance advantages over strong baseline models. First, the proposed approach achieved overall micro-F1 score improvements between 0.30 and 0.85 points across all benchmarks, reaching F1 scores of 86.31 on ACE2004, 85.11 on ACE2005, and 79.30 on GENIA. Second, the method significantly enhanced recall on low-frequency and unseen entities, improving recall by 0.56 to 2.56 points for rare terms. Third, the graph architecture proved especially beneficial for long entity spans of six or more words, delivering F1 improvements up to 13.11 points for eight-word spans by capturing informative lexical overlaps.

These results demonstrate that leveraging corpus-wide lexical connections provides critical contextual guidance when local sentence context is misleading or incomplete. Unlike complex parsing approaches that require external linguistic dependencies, this method uses existing training data to resolve ambiguous boundaries. This improves extraction accuracy in information-dense domains like biomedical text mining and intelligence analysis while avoiding manual feature engineering.

Organizations deploying automated information extraction systems should evaluate retrieval-augmented graph representations for complex and nested entity extraction pipelines. Teams adopting this framework must carefully balance operational trade-offs, as the graph structure reduces inference decoding throughput by roughly half compared to simpler span models, alongside a modest memory overhead of 100 to 500 megabytes. Future work should focus on optimizing retrieval speed and exploring broader deployment across multi-lingual datasets.

Confidence in these findings is supported by consistent gains across multiple benchmark datasets and detailed ablation experiments. However, decision-makers should note that the approach relies on the presence of informative lexical overlaps within the training set, meaning datasets with entirely disjoint vocabularies between training and operational environments may see more moderate benefits.

Wan et al (2022).pdf
Cover for Nested Named Entity Recognition with Span-level Graphs

Abstract

Span-based methods with the neural networks backbone have great potential for the nested named entity recognition (NER) problem. However, they face problems such as degenerating when positive instances and negative instances largely overlap. Besides, the generalization ability matters a lot in nested NER, as a large proportion of entities in the test set hardly appear in the training set. In this work, we try to improve the span representation by utilizing retrieval-based span-level graphs, connecting spans and entities in the training data based on n-gram features. Specifically, we build the entity-entity graph and span-entity graph globally based on n-gram similarity to integrate the information of similar neighbor entities into the span representation. To evaluate our method, we conduct experiments on three common nested NER datasets, ACE2004, ACE2005, and GENIA datasets. Experimental results show that our method achieves general improvements on all three benchmarks (+0.30 ∼ 0.85 micro-F1), and obtains special superiority on low frequency entities (+0.56 ∼ 2.08 recall).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Nested NER
  • 2.2 Graphs Used in NER
  • 3 Our Approach
  • 3.1 Span-level Graph
  • 3.2 The Encoder
  • 3.3 Graph Module
  • 3.4 Multitask Learning
  • 4 Experiments
  • 4.1 Dataset
  • 4.2 Baselines
  • 4.3 Training Details
  • 4.4 Main Results
  • 4.5 Ablation Study
  • 4.6 Recognition of Different Entities
  • 4.7 Case Study
  • 5 Conclusions
  • 6 Acknowledgements
  • References
  • A Hyper-parameters Analysis
  • B Time and Space Analysis

Knowls

  1. Knowl 1 — Retrieval-augmented span classification for nested NER

    model/method

    The proposed nested named entity recognition system augments a span classifier with lexical information retrieved from the entire training set. For every candidate span in a sentence, the system retrieves similar training-set entity mentions using BPE-level n-gram overlap, represents those neighbors with graph convolution, and combines the resulting neighbor representation with the candidate span’s contextual representation, sentence context, and span-size embedding. The system predicts labels only for raw sentence spans at inference time; entity mentions from the training set are used to provide graph structure and auxiliary supervision during training. The design is intended to address boundary confusion between overlapping spans and poor generalization to entities that are rare or unseen in training.

  2. Knowl 2 — Global entity and span-entity graph construction

    equation

    Let E\mathcal{E} be the set of entity mentions in the training data, R\mathcal{R} the set of raw candidate spans, and S=E∪R\mathcal{S}=\mathcal{E}\cup\mathcal{R} the set of all spans. A span is represented as a token sequence, and BPE⁡(s)\operatorname{BPE}(s) denotes its BPE-token sequence. For a positive integer nn, define the BPE n-gram overlap between spans s,s′∈Ss,s'\in\mathcal{S} as

    fn(s,s′)=∣n-gram⁡(BPE⁡(s))∩n-gram⁡(BPE⁡(s′))∣.f_n(s,s')=\left|\operatorname{n\text{-}gram}(\operatorname{BPE}(s))\cap\operatorname{n\text{-}gram}(\operatorname{BPE}(s'))\right|.

    The entity-entity graph contains one node for every training entity mention; mentions with identical tokens but different entity types remain distinct nodes. If NN is the largest n-gram order and αn\alpha_n is the weight of order nn, the edge weight between entity nodes eie_i and eje_j is

    w(ei,ej)=1N∑n=1Nαnfn(ei,ej).w(e_i,e_j)=\frac{1}{N}\sum_{n=1}^{N}\alpha_n f_n(e_i,e_j).

    The resulting weight measures lexical similarity through shared BPE n-grams and is used to retrieve entity neighbors beyond the current sentence.

  3. Knowl 3 — Length-normalized span-entity retrieval and local subgraphs

    algorithm

    The span-entity graph connects every raw span sijs_{ij}, covering tokens xix_i through xjx_j, to training entity mentions. Because long raw spans naturally share more n-grams and may connect to noisy entities, the edge weight is length-normalized:

    w(sij,e)=1N l(sij)∑n=1Nαnfn(sij,e),l(sij)=j−i+1.w(s_{ij},e)=\frac{1}{N\,l(s_{ij})}\sum_{n=1}^{N}\alpha_n f_n(s_{ij},e),\qquad l(s_{ij})=j-i+1.

    Edges with weight below a positive threshold τ\tau are removed. For each raw span, the model dynamically constructs a local KK-hop subgraph rather than processing the full global graph. It first retrieves the raw span’s one-hop entity neighbors from the span-entity graph. It then adds entity nodes reachable within up to K−1K-1 additional hops through the entity-entity graph, preserves edges among the retrieved entity nodes, removes the raw-span node itself from the graph passed to the graph convolutional network, and uses the resulting induced entity subgraph to represent the raw span. This retrieval procedure makes mini-batch training feasible while retaining lexical neighbors from the global training data.

  4. Knowl 4 — Encoder for spans and entity mentions

    model/method

    Raw sentence spans and training entity mentions are encoded as token sequences. Each token is represented using fixed word embeddings, a character embedding processed by a character-level bidirectional LSTM, and contextual BPE representations from a pretrained language model. The three representations are concatenated and passed through a second word-character bidirectional LSTM. The representation of a span is obtained by max-pooling the encoded token representations within that span. The sentence-level pretrained-language-model representation of the special [CLS] token is retained as a context vector for raw spans. This encoder is applied separately to sentence spans and entity mentions before graph message passing.

  5. Knowl 5 — GCN and attention-based graph representation

    equation

    For a retrieved entity subgraph, let AA be its normalized symmetric adjacency matrix, H0H^0 the matrix of encoder features, WkW^k a learnable weight matrix at graph layer kk, and KK the number of graph-convolution layers. The graph features are updated by

    Hk+1=ReLU⁡(AHkWk),k=0,…,K−1.H^{k+1}=\operatorname{ReLU}(AH^kW^k),\qquad k=0,\ldots,K-1.

    For a raw span, node 00 denotes the span and node i≥1i\geq1 denotes one of its retrieved entity neighbors. Let h00h_0^0 be the raw span’s initial feature, hiKh_i^K the final feature of neighbor ii, and WaW_a a learnable attention matrix. The attention weight assigned to neighbor ii and the graph representation of the raw span are

    γi=exp⁡((hiK)⊤Wah00)∑j≥1exp⁡((hjK)⊤Wah00),hgraph=∑i≥1γihiK.\gamma_i=\frac{\exp\left((h_i^K)^\top W_a h_0^0\right)}{\sum_{j\geq1}\exp\left((h_j^K)^\top W_a h_0^0\right)}, \qquad h^{\mathrm{graph}}=\sum_{i\geq1}\gamma_i h_i^K.

    Let hcontexth^{\mathrm{context}} be the pretrained-language-model [CLS] vector and hsizeh^{\mathrm{size}} a learnable embedding for span length. The final raw-span representation is the concatenation concat⁡(h00,hgraph,hcontext,hsize)\operatorname{concat}(h_0^0,h^{\mathrm{graph}},h^{\mathrm{context}},h^{\mathrm{size}}), while the final representation of entity neighbor ii is concat⁡(hi0,hiK,hsize)\operatorname{concat}(h_i^0,h_i^K,h^{\mathrm{size}}). Thus, the candidate span attends to graph neighbors after they have aggregated information from their entity-entity neighborhood.

  6. Knowl 6 — Multitask use of neighboring entity labels

    model/method

    The model uses two classification heads: one predicts the label of each raw candidate span, and the other predicts the entity type of every entity node retrieved into the local graph. If h0finalh_0^{\mathrm{final}} is the final raw-span representation and hifinalh_i^{\mathrm{final}} is the final representation of entity neighbor ii, the corresponding logits are

    logits⁡s=Linear⁡s(h0final),logits⁡ei=Linear⁡e(hifinal).\operatorname{logits}_s=\operatorname{Linear}_s(h_0^{\mathrm{final}}),\qquad \operatorname{logits}_{e_i}=\operatorname{Linear}_e(h_i^{\mathrm{final}}).

    With cross-entropy losses LsL_s for raw-span labels and LeL_e for neighboring entity labels, training minimizes

    L=Ls+βLe,L=L_s+\beta L_e,

    where β\beta controls the auxiliary entity-label objective. The auxiliary task forces graph representations to retain entity-type information from similar training mentions; only the raw-span classifier is used during inference.

  7. Knowl 7 — Nested NER formulation as bounded span classification

    definition

    For a sentence X=(x1,…,xn)X=(x_1,\ldots,x_n) of nn tokens, the model enumerates candidate spans sijs_{ij} beginning at token xix_i and ending at token xjx_j, subject to the maximum span length of 10 tokens. Each candidate is independently assigned either one of the predefined entity types or the label NA, meaning that the span is not an entity. Because spans are classified independently rather than forced into a single sequence-labeling path, overlapping and nested entities can be predicted simultaneously.

  8. Knowl 8 — Datasets, preprocessing, and training configuration

    experimental setup

    Experiments use ACE2004, ACE2005, and GENIA. The train, development, and test statistics are:

    Could not parse LaTeX table

    For GENIA, the first 81%, next 9%, and final 10% of the corpus are used for training, development, and testing; DNA, RNA, and protein subtypes are collapsed, cell line and cell type are retained, and five entity types remain after removing other types. ACE2004 and ACE2005 use 100-dimensional fixed GloVe embeddings, while GENIA uses 200-dimensional fixed biomedical embeddings. Character embeddings have dimension 30, the character BiLSTM has 60-dimensional hidden states, the word-character BiLSTM has hidden size 300, and span-size embeddings have dimension 25. BERT-base-cased is used for ACE2004 and ACE2005, and BioBERT v1.1 for GENIA.

    The graph uses BPE n-grams of orders 11 through 33, weights αk=0.5k\alpha_k=0.5^k, edge threshold τ=0.8\tau=0.8, a two-layer GCN with hidden size 400, and 100 randomly sampled negative spans per training instance. The multitask coefficient is β=0.1\beta=0.1. Adam optimization, linear scheduling with a 0.01 warm-up ratio, at most 50 epochs, and validation-based model selection are used. Performance is reported as span-level micro-averaged precision, recall, and F1, averaged over three runs.

  9. Knowl 9 — Overall performance on nested NER benchmarks

    empirical result

    The retrieval-based graph model improves over comparable pretrained-language-model baselines on all three datasets. The reported precision, recall, and F1 for the proposed model are 86.70/85.93/86.3186.70/85.93/86.31 on ACE2004, 84.37/85.87/85.1184.37/85.87/85.11 on ACE2005, and 77.92/80.74/79.3077.92/80.74/79.30 on GENIA. The F1 comparison is:

    Could not parse LaTeX table

    Relative to the strongest listed pretrained span/layer baselines, the proposed model gains 0.85 F1 on ACE2004 over Pyramid and 0.90 F1 on ACE2005 over Pyramid, and it is 0.30 F1 above the BioBERT-based Pyramid model on GENIA. Models using additional supervision such as machine-reading-comprehension or dependency information are not directly comparable, but they provide higher scores on some datasets.

  10. Knowl 10 — Component ablation and gains on difficult entities

    empirical result

    On ACE2005, adding the graph components incrementally improves the span classifier. The Char + Word + Language-Model baseline obtains precision 81.98, recall 85.61, and F1 83.75. Adding the span-entity graph gives 83.35/85.28/84.30, an F1 increase of 0.55; adding the entity-entity graph next gives 84.60/84.68/84.64, a further increase of 0.34; and adding multitask training gives 84.37/85.87/85.11, a further increase of 0.47.

    The gains concentrate on difficult cases. Compared with the SpERT span baseline, nested-entity recall increases from 81.78 to 82.78 on ACE2004 development data, from 83.91 to 84.84 on ACE2004 test data, from 82.70 to 84.55 on ACE2005 development data, from 83.38 to 84.65 on ACE2005 test data, from 66.26 to 72.13 on GENIA development data, and from 66.09 to 70.13 on GENIA test data. On ACE2005, F1 improvements for entity lengths 6 through 10 are respectively +11.29, +9.09, +13.11, +11.11, and +7.43 points, with the largest gain occurring at length 8. For entities seen zero times in training, recall improves by 2.35 points on ACE2004 and 0.81 points on ACE2005; across entities with training frequency at most four, the recall gains range from 2.08 to 2.56 points on ACE2004 and from 0.56 to 0.83 points on ACE2005. These results support the claim that retrieved lexical neighbors are especially useful for long, nested, rare, and unseen entities. The graph also incurs a practical cost: on a GTX 1080 Ti with batch size 4, inference speed is approximately half that of SpERT, while GPU memory increases by roughly 100--500 MB depending on batch size.

Coverage note — The qualitative case study and hyperparameter-sensitivity plots were omitted because they illustrate the graph behavior and tuning choices without adding an independent load-bearing method or headline result; the main efficiency tradeoff was retained.

References

  1. 1.Kate Byrne. 2007. Nested named entity recognition in historical archive text. In ICSC 2007, pages 589–596.
  2. 2.Alberto Cetoli, Stefano Bragaglia, Andrew O’Harney, and Marc Sloan. 2017. Graph convolutional networks for named entity recognition. In Proceedings of the 16th International Workshop on Treebanks and Linguistic Theories, pages 37–45, Prague, Czech Republic.
  3. 3.Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa. 2011. Natural language processing (almost) from scratch. J. Mach. Learn. Res., 12:2493–2537.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Ruixue Ding, Pengjun Xie, Xiaoyan Zhang, Wei Lu, Linlin Li, and Luo Si. 2019. A neural multi-digraph model for Chinese NER with gazetteers. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1462–1467, Florence, Italy. Association for Computational Linguistics.
  6. 6.Timothy Dozat and Christopher D. Manning. 2017. Deep biaffine attention for neural dependency parsing. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  7. 7.Markus Eberts and Adrian Ulges. 2020. Span-based joint entity and relation extraction with transformer pre-training. In ECAI 2020, volume 325, pages 2006–2013.
  8. 8.Jenny Rose Finkel and Christopher D. Manning. 2009. Nested named entity recognition. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 141–150, Singapore. Association for Computational Linguistics.
  9. 9.Joseph Fisher and Andreas Vlachos. 2019. Merge and label: A novel neural network architecture for nested NER. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5840–5850, Florence, Italy. Association for Computational Linguistics.
  10. 10.Jinlan Fu, Xuanjing Huang, and Pengfei Liu. 2021. SpanNER: Named entity re-/recognition as span prediction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7183–7195, Online. Association for Computational Linguistics.
  11. 11.Tsu-Jui Fu, Peng-Hsuan Li, and Wei-Yun Ma. 2019. GraphRel: Modeling text as relational graphs for joint entity and relation extraction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1409–1418, Florence, Italy. Association for Computational Linguistics.
  12. 12.Tao Gui, Yicheng Zou, Qi Zhang, Minlong Peng, Jinlan Fu, Zhongyu Wei, and Xuanjing Huang. 2019. A lexicon-based graph neural network for Chinese NER. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1040–1050, Hong Kong, China. Association for Computational Linguistics.
  13. 13.Meizhi Ju, Makoto Miwa, and Sophia Ananiadou. 2018. A neural layered model for nested named entity recognition. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1446–1459, New Orleans, Louisiana. Association for Computational Linguistics.
  14. 14.Arzoo Katiyar and Claire Cardie. 2018. Nested named entity recognition revisited. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 861–871, New Orleans, Louisiana. Association for Computational Linguistics.
  15. 15.J-D Kim, Tomoko Ohta, Yuka Tateisi, and Jun’ichi Tsujii. 2003. Genia corpus—a semantically annotated corpus for bio-textmining. Bioinformatics, 19:i180–i182.
  16. 16.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  17. 17.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  18. 18.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 260–270, San Diego, California. Association for Computational Linguistics.
  19. 19.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinform., 36(4):1234–1240.
  20. 20.Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2020. A unified MRC framework for named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5849–5859, Online. Association for Computational Linguistics.
  21. 21.Hongyu Lin, Yaojie Lu, Xianpei Han, and Le Sun. 2019. Sequence-to-nuggets: Nested entity mention detection via anchor-region networks. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5182–5192, Florence, Italy. Association for Computational Linguistics.
  22. 22.Wei Lu and Dan Roth. 2015. Joint mention extraction and classification with mention hypergraphs. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 857–867, Lisbon, Portugal. Association for Computational Linguistics.
  23. 23.Yi Luan, Dave Wadden, Luheng He, Amy Shah, Mari Ostendorf, and Hannaneh Hajishirzi. 2019. A general framework for information extraction using dynamic span graphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3036–3046, Minneapolis, Minnesota. Association for Computational Linguistics.
  24. 24.Ying Luo and Hai Zhao. 2020. Bipartite flat-graph network for nested named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6408–6418, Online. Association for Computational Linguistics.
  25. 25.Xuezhe Ma and Eduard Hovy. 2016. End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1064–1074, Berlin, Germany. Association for Computational Linguistics.
  26. 26.Aldrian Obaja Muis and Wei Lu. 2017. Labeling gaps between words: Recognizing overlapping mentions with mention separators. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2608–2618, Copenhagen, Denmark. Association for Computational Linguistics.
  27. 27.Hiroki Ouchi, Jun Suzuki, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, and Kentaro Inui. 2020. Instance-based learning of span representations: A case study through named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6452–6459, Online. Association for Computational Linguistics.
  28. 28.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  29. 29.Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang, Wen Wang, and Weiming Lu. 2021. Locate and label: A two-stage identifier for nested named entity recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2782–2794, Online. Association for Computational Linguistics.
  30. 30.Takashi Shibuya and Eduard Hovy. 2020. Nested named entity recognition via second-best sequence learning and decoding. Transactions of the Association for Computational Linguistics, 8:605–620.
  31. 31.Mohammad Golam Sohrab and Makoto Miwa. 2018. Deep exhaustive model for nested named entity recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2843–2849, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Jana Straková, Milan Straka, and Jan Hajic. 2019. Neural architectures for nested NER through linearization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5326–5331, Florence, Italy. Association for Computational Linguistics.
  33. 33.Chuanqi Tan, Wei Qiu, Mosha Chen, Rui Wang, and Fei Huang. 2020. Boundary enhanced neural span classification for nested named entity recognition. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 9016–9023. AAAI Press.
  34. 34.Zeqi Tan, Yongliang Shen, Shuai Zhang, Weiming Lu, and Yueting Zhuang. 2021. A sequence-to-set network for nested named entity recognition. In IJCAI 2021, pages 3936–3942.
  35. 35.Bailin Wang and Wei Lu. 2018. Neural segmental hypergraphs for overlapping mention recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 204–214, Brussels, Belgium. Association for Computational Linguistics.
  36. 36.Bailin Wang, Wei Lu, Yu Wang, and Hongxia Jin. 2018. A neural transition-based model for nested mention recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1011–1017, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Jue Wang, Lidan Shou, Ke Chen, and Gang Chen. 2020. Pyramid: A layered model for nested named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5918–5928, Online. Association for Computational Linguistics.
  38. 38.Yiran Wang, Hiroyuki Shindo, Yuji Matsumoto, and Taro Watanabe. 2021. Nested named entity recognition via explicitly excluding the influence of the best path. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3547–3557, Online. Association for Computational Linguistics.
  39. 39.Yongxiu Xu, Heyan Huang, Chong Feng, and Yue Hu. 2021. A supervised multi-head self-attention network for nested named entity recognition. In AAAI 2021, pages 14185–14193.
  40. 40.Juntao Yu, Bernd Bohnet, and Massimo Poesio. 2020. Named entity recognition as dependency parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6470–6476, Online. Association for Computational Linguistics.
  41. 41.Yue Zhang and Jie Yang. 2018a. Chinese NER using lattice LSTM. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1554–1564, Melbourne, Australia. Association for Computational Linguistics.
  42. 42.Yue Zhang and Jie Yang. 2018b. Chinese NER using lattice LSTM. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1554–1564, Melbourne, Australia. Association for Computational Linguistics.
  43. 43.Changmeng Zheng, Yi Cai, Jingyun Xu, Ho-fung Leung, and Guandong Xu. 2019. A boundary-aware neural model for nested named entity recognition. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 357–366, Hong Kong, China. Association for Computational Linguistics.

Citation

MLA
Wan, J., et al. “Nested Named Entity Recognition with Span-level Graphs”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 892–903, https://doi.org/10.18653/v1/2022.acl-long.63.
APA
Wan, J., Ru, D., Zhang, W., & Yu, Y. (2022). Nested Named Entity Recognition with Span-level Graphs. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 892–903. https://doi.org/10.18653/v1/2022.acl-long.63
Chicago
Wan, J., D. Ru, W. Zhang, and Y. Yu. 2022. “Nested Named Entity Recognition with Span-level Graphs”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 892–903. https://doi.org/10.18653/v1/2022.acl-long.63.
Harvard
Wan, J. et al. (2022) “Nested Named Entity Recognition with Span-level Graphs”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 892–903. Available at: https://doi.org/10.18653/v1/2022.acl-long.63.
Vancouver
1. Wan J, Ru D, Zhang W, Yu Y (2022) Nested Named Entity Recognition with Span-level Graphs. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 892–903

BibTeX

@inproceedings{wan-etal-2022-nested,
    title = "Nested Named Entity Recognition with Span-level Graphs",
    author = "Wan, Juncheng  and
      Ru, Dongyu  and
      Zhang, Weinan  and
      Yu, Yong",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.63/",
    doi = "10.18653/v1/2022.acl-long.63",
    pages = "892--903"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/