A Survey on Deep Learning for Named Entity Recognition

Jing LiAixin SunJianglei HanChenliang Li

article2018TKDE1,529 citations

Systematizes deep learning approaches for named entity recognition across input representations, context encoders, and tag decoders, while summarizing essential datasets, off-the-shelf tools, and future research directions.

Listen

Named Entity Recognition (NER) is a foundational capability in natural language processing that locates and classifies key information—such as individuals, locations, and organizations—within unstructured text. It serves as a vital component for downstream business applications, including semantic search, automated question answering, customer support routing, and knowledge base construction. Historically, extracting these entities required labor-intensive, handcrafted rules or extensive feature engineering tailored to specific domains. The rise of deep learning has fundamentally reshaped this field by automatically learning rich representations from raw text, making it critical for decision-makers to understand the technological landscape, performance trade-offs, and deployment feasibility.

The article provides a comprehensive evaluation of deep learning techniques applied to named entity recognition, systematically categorizing recent architectures and identifying the primary factors that drive model performance across different domains.

To establish these insights, the article reviews the extensive literature and evaluates established benchmarks—such as CoNLL03 and OntoNotes 5.0—alongside challenging user-generated datasets like W-NUT17. The analysis organizes neural NER systems along a modular three-part taxonomy: input representations (word, character, or hybrid embeddings), context encoders (such as recurrent networks, convolutional models, and transformers), and tag decoders (such as softmax classifiers and conditional random fields). It also examines advanced learning paradigms, including multi-task learning, transfer learning, active learning, and reinforcement learning.

The findings indicate that input representations are the primary driver of performance, with pre-trained contextual language models (such as BERT and ELMo) achieving benchmark accuracy above 92% to 93.5% F-score on standard English newswire texts. In contrast, model performance drops steeply to slightly above 40% on informal, user-generated text due to noisiness, abbreviations, and emerging unseen entities. Structurally, bidirectional recurrent models combined with conditional random fields remain the standard baseline, but Transformer-based contextual encoders achieve superior results when pre-trained on massive datasets. Furthermore, the analysis reveals that active learning can dramatically lower costs, matching 99% of full deep-learning performance while using only about 25% to 30% of labeled training data.

These results demonstrate that organizations can significantly reduce feature engineering overhead and achieve near-human accuracy on formal documents, but must manage critical risks when deploying models on informal business data or specialized industry terminology. Relying on standard off-the-shelf models for noisy customer communications without adaptation carries a severe risk of extraction errors and downstream failures. Moreover, high-performing contextual models require substantial computational infrastructure, meaning teams must balance model complexity against deployment budgets and latency requirements.

Decision-makers should adopt a task-dependent implementation strategy. For general newswire text, organizations should leverage existing pre-trained models, while domain-specific applications (such as healthcare or finance) should fine-tune general contextual models using targeted internal data. When labeled data is scarce, teams should invest in active learning and transfer learning workflows rather than costly full-scale manual annotation. Future initiatives should focus on decoupling entity boundary detection from type classification and developing joint entity recognition and linking pipelines to minimize cascading errors.

These conclusions are primarily bounded by English-language benchmarks and formal text corpora. Reader caution is advised regarding cross-domain generalization, as performance variability remains high when models encounter out-of-vocabulary terms, nested entity structures, and annotation schema mismatches between datasets.

Cover for A Survey on Deep Learning for Named Entity Recognition

Abstract

Named entity recognition (NER) is the task to identify mentions of rigid designators from text belonging to predefined semantic types such as person, location, organization etc. NER always serves as the foundation for many natural language applications such as question answering, text summarization, and machine translation. Early NER systems got a huge success in achieving good performance with the cost of human engineering in designing domain-specific features and rules. In recent years, deep learning, empowered by continuous real-valued vector representations and semantic composition through nonlinear processing, has been employed in NER systems, yielding stat-of-the-art performance. In this paper, we provide a comprehensive review on existing deep learning techniques for NER. We first introduce NER resources, including tagged NER corpora and off-the-shelf NER tools. Then, we systematically categorize existing works based on a taxonomy along three axes: distributed representations for input, context encoder, and tag decoder. Next, we survey the most representative methods for recent applied techniques of deep learning in new NER problem settings and applications. Finally, we present readers with the challenges faced by NER systems and outline future directions in this area.

Table of Contents

  • I Introduction
  • II Background
  • II-A What is NER?
  • II-B NER Resources: Datasets and Tools
  • II-C NER Evaluation Metrics
  • II-C1 Exact-match Evaluation
  • II-C2 Relaxed-match Evaluation
  • II-D Traditional Approaches to NER
  • II-D1 Rule-based Approaches
  • II-D2 Unsupervised Learning Approaches
  • II-D3 Feature-based Supervised Learning Approaches
  • III Deep Learning Techniques for NER
  • III-A Why Deep Learning for NER?
  • III-B Distributed Representations for Input
  • III-B1 Word-level Representation
  • III-B2 Character-level Representation
  • III-B3 Hybrid Representation
  • III-C Context Encoder Architectures
  • III-C1 Convolutional Neural Networks
  • III-C2 Recurrent Neural Networks
  • III-C3 Recursive Neural Networks
  • III-C4 Neural Language Models
  • III-C5 Deep Transformer
  • III-D Tag Decoder Architectures
  • III-D1 Multi-Layer Perceptron + Softmax
  • III-D2 Conditional Random Fields
  • III-D3 Recurrent Neural Networks
  • III-D4 Pointer Networks
  • III-E Summary of DL-based NER
  • IV Applied Deep Learning for NER
  • IV-A Deep Multi-task Learning for NER
  • IV-B Deep Transfer Learning for NER
  • IV-C Deep Active Learning for NER
  • IV-D Deep Reinforcement Learning for NER
  • IV-E Deep Adversarial Learning for NER
  • IV-F Neural Attention for NER
  • V Challenges and Future Directions
  • V-A Challenges
  • V-B Future Directions
  • VI Conclusion
  • References

Knowls

  1. Knowl 1 — Three-Axis Taxonomy of Deep-Learning-Based Named Entity Recognition

    model/method

    Deep-learning-based Named Entity Recognition (NER) systems are systematically organized along a three-axis structural taxonomy comprising:

    1. Distributed Representations for Input: Maps raw input tokens into low-dimensional, continuous real-valued dense feature vectors. This encompasses word-level embeddings (e.g., Word2Vec, GloVe, SENNA), character-level representations learned via neural sub-networks (capturing prefixes, suffixes, and morphological regularities while mitigating out-of-vocabulary tokens), and hybrid representations that augment token embeddings with gazetteers, part-of-speech (POS) tags, visual features, or pre-trained contextualized language model embeddings.
    2. Context Encoder: Encodes contextual semantic dependencies across the input sequence into hidden state vectors. Common context encoder architectures include Convolutional Neural Networks (CNNs and Dilated CNNs), Recurrent Neural Networks (BiLSTM, BiGRU), Recursive Neural Networks over syntactic parse trees, and Deep Transformers based on self-attention.
    3. Tag Decoder: Consumes context-dependent token representations to predict sequence tag assignments (such as BIO or BIOES encodings) or entity segment spans. Standard tag decoders include Multi-Layer Perceptrons with Softmax classifiers, Conditional Random Fields (CRFs) and semi-Markov CRFs, autoregressive Recurrent Neural Networks, and Pointer Networks.

    This taxonomy unifies DL-based sequence labeling architectures and resolves ambiguities in prior classifications that conflated word-level feature extraction with context encoding.

  2. Knowl 2 — Formal Problem Definition and Evaluation Metrics for Named Entity Recognition

    definition

    Given an input sequence of NN tokens s=⟨w1,w2,…,wN⟩s = \langle w_1, w_2, \dots, w_N \rangle, the Named Entity Recognition (NER) task outputs a collection of entity mention tuples ⟨Is,Ie,t⟩\langle I_s, I_e, t \rangle, where Is∈[1,N]I_s \in [1, N] is the start token index, Ie∈[1,N]I_e \in [1, N] is the end token index (with Is≤IeI_s \le I_e), and tt is a semantic entity type from a predefined category inventory T\mathcal{T}.

    Under exact-match evaluation, an entity mention is counted as a True Positive (TPTP) if and only if both the predicted boundary span [Is,Ie][I_s, I_e] and the assigned entity type tt match the ground-truth annotation exactly. The evaluation metrics are defined as:

    Precision=TPTP+FP\text{Precision} = \frac{TP}{TP + FP}

    Recall=TPTP+FN\text{Recall} = \frac{TP}{TP + FN}

    F-score=2×Precision×RecallPrecision+Recall\text{F-score} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}

    where FPFP (False Positives) denotes predicted entities not matching ground truth, and FNFN (False Negatives) denotes ground-truth entities missed by the model. Performance across multiple entity types is summarized via macro-averaged F-score (unweighted arithmetic mean of F-scores computed independently per entity class) or micro-averaged F-score (global aggregation of TPTP, FPFP, and FNFN across all classes prior to score calculation).

    Under relaxed-match evaluation, partial credit is awarded if a predicted entity boundary overlaps with ground truth regardless of type, or if the type is correct despite an imperfect boundary span.

  3. Knowl 3 — Distributed Input Representation Paradigms for Neural NER

    model/method

    Neural NER models utilize three main categories of distributed input representations to encode token-level features:

    1. Word-Level Representations: Tokens are mapped to static dense vectors pre-trained on large corpora using unsupervised algorithms (e.g., Continuous Bag-of-Words, Skip-gram, GloVe, fastText, SENNA). These vectors can remain fixed or undergo task-specific fine-tuning during training.
    2. Character-Level Representations: Sub-word features are extracted from character sequences of individual words using either CNNs (character embeddings followed by 1D convolution and max/average pooling across the character span) or RNNs (bidirectional LSTM or GRU over character tokens). Character representations handle out-of-vocabulary (OOV) tokens and encode morphological patterns (such as capitalization, suffixes, and prefixes).
    3. Hybrid and Contextualized Representations: Word and character embeddings are concatenated or fused with auxiliary explicit features (e.g., gazetteer match vectors, POS tags, syntactic dependency relations, topic distributions, or word shape encodings). This category also includes contextualized language model representations (e.g., ELMo, BERT, contextual string embeddings) that dynamically generate token representations based on surrounding sentential context.
  4. Knowl 4 — Context Encoder Architectures for Neural NER

    model/method

    Context encoders in neural NER generate context-sensitive feature representations from token input vectors via four primary architectural families:

    • Convolutional Neural Networks (CNNs) & Dilated CNNs: Standard CNNs produce local feature maps over sliding word windows and aggregate global representations via pooling. Iterated Dilated Convolutional Neural Networks (ID-CNNs) stack dilated 1D convolutions with exponentially expanding receptive fields, enabling O(1)O(1) sequential depth per token and 14×–20×14\times\text{--}20\times inference speedups over recurrent encoders through parallel document-level computation.
    • Recurrent Neural Networks (RNNs): Bidirectional LSTMs (BiLSTM) and Bidirectional GRUs (BiGRU) concatenate forward hidden states h→i\overrightarrow{h}_i (past context) and backward hidden states h←i\overleftarrow{h}_i (future context) at token position ii, forming the standard baseline for deep context modeling.
    • Recursive Neural Networks: Process sentence constituency parse trees hierarchically. Hidden states are computed recursively in a bottom-up direction (encoding semantic composition of subtrees) and a top-down direction (propagating containing linguistic structures) to classify syntactic constituent nodes.
    • Neural Language Models & Deep Transformers: Multi-layer bidirectional self-attention encoders (e.g., Transformers, BERT) compute pairwise token dependencies across the entire sequence without recurrence, operating with computational complexity O(n2⋅d)O(n^2 \cdot d) relative to sequence length nn and representation dimension dd (compared to O(n⋅d2)O(n \cdot d^2) for recurrent models).
  5. Knowl 5 — Tag Decoder Architectures for Neural Sequence Labeling

    model/method

    Tag decoders map contextual encoder hidden representations {h1,h2,…,hN}\{h_1, h_2, \dots, h_N\} to output entity label sequences using four principal mechanisms:

    • Multi-Layer Perceptron (MLP) + Softmax: Formulates tag assignment as independent multiclass classification at each position ii, computing label probabilities Softmax(Whi+b)\text{Softmax}(W h_i + b) without modeling label transition constraints between neighboring tokens.
    • Conditional Random Fields (CRFs): Globally normalizes sequence probabilities over the entire sentence, explicitly modeling label transition matrices Aj,k=p(yi=k∣yi−1=j)A_{j,k} = p(y_{i} = k \mid y_{i-1} = j) alongside token emission scores to enforce tag syntax (e.g., disallowing an inside tag I-PER immediately following O). Semi-Markov CRFs (Semi-CRFs) model variable-length segments rather than individual tokens to score whole entity spans.
    • Autoregressive Recurrent Neural Networks: Employs an RNN tag decoder that models label sequences sequentially as a tag language model. The decoder hidden state hi+1Dech^{\text{Dec}}_{i+1} is updated given previous tag yiy_i, previous decoder state hiDech^{\text{Dec}}_i, and current encoder state hi+1Ench^{\text{Enc}}_{i+1}, generating greedy or beam-searched output tags.
    • Pointer Networks: Employs sequence-to-sequence attention distributions as discrete pointers over input token indices to first segment variable-length candidate text spans and subsequently predict entity category labels for each identified span.
  6. Knowl 6 — Benchmark Performance of Deep-Learning NER Architectures Across Datasets

    data/table

    Comparative experimental results across representative neural NER architectures demonstrate that pre-trained contextualized language models (BERT, ELMo, contextual string embeddings, cloze-driven Transformers) paired with CRF or softmax decoders achieve top-tier performance on formal English benchmarks (CoNLL03 and OntoNotes 5.0), whereas performance drops substantially on noisy, user-generated text (e.g., W-NUT17):

    Architecture Character Word Hybrid / Auxiliary Context Encoder Tag Decoder Performance (F-score)
    Collobert et al. – Random POS CNN CRF CoNLL03: 89.86%
    Huang et al. – SENNA Spelling, n-gram, gazetteer LSTM CRF CoNLL03: 90.10%
    Chiu Nichols CNN SENNA Capitalization, lexicons LSTM CRF CoNLL03: 91.62%, OntoNotes5.0: 86.34%
    Lample et al. LSTM SENNA – LSTM CRF CoNLL03: 90.94%
    Ma Hovy CNN GloVe – LSTM CRF CoNLL03: 91.21%
    Strubell et al. – SENNA Word shape ID-CNN CRF CoNLL03: 90.65%, OntoNotes5.0: 86.84%
    Peters et al. (ELMo) CNN-LSTM-LM – – LSTM CRF CoNLL03: 92.22%
    Akbik et al. (Flair) LSTM-LM – – LSTM CRF CoNLL03: 93.09%, OntoNotes5.0: 89.71%
    Devlin et al. (BERT) – WordPiece Segment, position Transformer Softmax CoNLL03: 92.8%
    Baevski et al. CNN – Cloze-style LM LSTM CRF CoNLL03: 93.5%
    Li et al. (Dice Loss) – – BERT – Softmax + Dice CoNLL03: 93.33%, OntoNotes5.0: 92.07%
    Lin et al. LSTM GloVe Syntactic features LSTM CRF W-NUT17: 40.42%
    Aguilar et al. CNN Twitter Word2Vec POS LSTM CRF W-NUT17: 41.86%

    Key empirical findings indicate:

    1. BiLSTM-CRF with concatenated character and word embeddings represents the most widespread neural baseline for sequence labeling.
    2. Incorporating deep contextualized representations yields large performance gains (+1.5–3.5+1.5\text{--}3.5 percentage points in F-score) over static embeddings.
    3. While formal text F-scores consistently exceed 90%–93%90\%\text{--}93\%, entity recognition on short, informal social media posts (W-NUT17) remains challenging, with maximum F-scores remaining around 41%–42%41\%\text{--}42\%.
  7. Knowl 7 — Deep Multi-Task and Transfer Learning Approaches for NER

    model/method

    Deep multi-task and transfer learning methods mitigate data scarcity and domain mismatch in NER through several parameter-sharing and adaptation strategies:

    • Multi-Task Learning (MTL): Simultaneously trains a shared neural representation on NER alongside related sequence labeling tasks (POS tagging, syntactic chunking, semantic role labeling), auxiliary language modeling objectives, or subtasks (e.g., separate entity boundary segmentation and type classification). In domain-specific contexts (such as biomedical NER across multiple distinct datasets), sharing character- and word-level encoder layers encourages generalization across heterogeneous corpora.
    • Parameter-Sharing Transfer Architectures: Partitions neural network weights into domain-shared and task-specific modules. When source and target domains share identical or mappable label sets, a shared CRF decoder layer is utilized; when tag sets diverge, private CRF decoders or output projection layers are trained per domain while sharing underlying contextual encoder parameters.
    • Fine-Tuning & Adapter Layers: Adapts pre-trained representations to target domains by inserting dedicated neural adaptation layers (word adaptation, sentence adaptation, and output adaptation layers) or by hierarchical tag-mapping networks that project fine-grained source label sets onto target tag hierarchies during inference.
    • Transfer Joint Embedding (TJE): Employs label embeddings to transform cross-domain multiclass classification into low-dimensional regression, transferring knowledge between source and target domains that feature related but non-identical entity types.
  8. Knowl 8 — Deep Active, Reinforcement, and Adversarial Learning Paradigms for NER

    model/method

    Beyond standard supervised learning, advanced deep learning paradigms are applied to optimize annotation efficiency, instance selection, and domain robustness in NER:

    • Deep Active Learning: Combines uncertainty sampling strategies with incremental training to minimize human annotation effort. Instead of expensive retraining from scratch upon acquiring new labels, model weights are updated incrementally over few epochs on the newly labeled batch mixed with prior data. Active learning achieves 99%99\% of the performance of full-dataset training using only 24.9%24.9\% of English training data and 30.1%30.1\% of Chinese training data.
    • Deep Reinforcement Learning (DRL): Formulates information extraction as a Markov Decision Process (MDP) parameterized by Deep Q-Networks (DQN) to iteratively generate search queries, extract entities from external documents, and reconcile evidence. In distant supervision settings, reinforcement learning agents act as instance selectors that receive reward feedback from downstream NER taggers to filter out noisy and misaligned annotations.
    • Deep Adversarial Learning: Employs minimax adversarial objectives to learn domain-invariant feature representations for cross-domain or cross-lingual NER. Alternatively, adversarial training adds small norm-bounded perturbations ηx\eta_x directly to input representations to maximize loss:

    ηx=arg⁡max⁡η:∥η∥2≤ϵl(Θ;x+η)\eta_x = \arg\max_{\eta : \|\eta\|_2 \le \epsilon} l(\Theta; x + \eta)

    Training on adversarial examples xadv=x+ηxx_{\text{adv}} = x + \eta_x regularizes neural taggers and improves out-of-domain generalization.

  9. Knowl 9 — Neural Attention Mechanisms in Named Entity Recognition

    model/method

    Attention mechanisms in neural NER architectures dynamically route information across sequence positions, modalities, and hierarchical representations:

    1. Feature and Component Gating: Dynamic attention gates calculate position-wise weights to determine the relative contribution of character-level versus word-level embedding components before feeding them to context encoders.
    2. Self-Attention: Computes pairwise token alignment weights within a single sequence, enabling direct modeling of long-distance dependencies independent of recurrence or convolutional window limits.
    3. Document-Level and Global Attention: Attends over hidden states derived from surrounding document context (e.g., using bidirectional language models), conditioning sentence-level token tagging on broader discourse semantics.
    4. Cross-Modal Co-Attention: Computes joint visual-textual attention matrices for multimodal NER (e.g., social media text paired with images), dynamically capturing cross-modal alignment between visual regions of interest and textual entity spans.
  10. Knowl 10 — Key Challenges and Open Directions in Neural Named Entity Recognition

    limitation

    Despite high reported F-scores on standard newswire benchmarks, neural NER systems face several fundamental limitations and open research challenges:

    1. Annotation Inconsistencies and Complex Mention Structures: Inter-dataset annotation discrepancies (e.g., "Baltimore" labeled as Location in MUC-7 vs. Organization in CoNLL03) severely hinder cross-corpus model transfer. Furthermore, nested entity mentions (occurring in 17%17\% of GENIA entities and 30%30\% of ACE sentences) and fine-grained multi-label types violate the flat sequence labeling assumption of standard taggers.
    2. Entangled Boundary Detection and Classification: Standard tag decoders solve boundary detection and entity type classification jointly (e.g., via B-, I-, E-, S-, O tags), compounding errors in fine-grained settings. Decoupling entity boundary detection as a standalone, domain-agnostic pre-processing step from domain-specific type classification provides a promising avenue for modular, reusable NER pipelines.
    3. Degradation on Informal and Emerging Text: Models trained on standard corpora degrade drastically on noisy user-generated content (e.g., tweets, forum posts), where best F-scores remain near 40%–42%40\%\text{--}42\%. Addressing novel and emerging entities requires specialized integration of external auxiliary knowledge (e.g., gazetteers) without sacrificing end-to-end differentiability.
    4. Joint Modeling with Entity Linking: The traditional pipeline separation of NER and Entity Linking (EL) causes unidirectional error propagation. Joint architectures that unify boundary detection, type classification, and knowledge base disambiguation are needed so downstream semantic grounding directly assists entity recognition.
    5. Model Scalability and Efficiency: Giant pre-trained Transformer language models require massive compute and memory, creating deployment and fine-tuning bottlenecks for resource-constrained settings and motivating structured pruning, knowledge distillation, and lightweight modular architectures.

Coverage note — Omitted introductory historical narrative on pre-neural rule-based and early feature-based NER tools (e.g., FASTUS, LaSIE-II, IdentiFinder), and generic off-the-shelf URL catalogs (Table 1 and Table 2), focusing strictly on the paper's core contributions regarding deep learning taxonomies, architectural components, applied DL paradigms, consolidated benchmarks, and open research challenges.

References

  1. 1.D. Nadeau and S. Sekine, “A survey of named entity recognition and classification,” Lingvist. Investig., vol. 30, no. 1, pp. 3–26, 2007.
  2. 2.Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu, “ERNIE: enhanced language representation with informative entities,” in ACL, 2019, pp. 1441–1451.
  3. 3.P. Cheng and K. Erk, “Attending to entities for better text understanding,” arXiv preprint arXiv:1911.04361, 2019.
  4. 4.J. Guo, G. Xu, X. Cheng, and H. Li, “Named entity recognition in query,” in SIGIR, 2009, pp. 267–274.
  5. 5.D. Petkova and W. B. Croft, “Proximity-based document representation for named entity retrieval,” in CIKM, 2007, pp. 731–740.
  6. 6.C. Aone, M. E. Okurowski, and J. Gorlinsky, “A trainable summarizer with knowledge acquired from robust nlp techniques,” Adv. Autom. Text Summ., vol. 71, 1999.
  7. 7.D. M. Aliod, M. van Zaanen, and D. Smith, “Named entity recognition for question answering,” in ALTA, 2006, pp. 51–58.
  8. 8.B. Babych and A. Hartley, “Improving machine translation quality with automatic named entity recognition,” in EAMT, 2003, pp. 1–8.
  9. 9.O. Etzioni, M. Cafarella, D. Downey, A.-M. Popescu, T. Shaked, S. Soderland, D. S. Weld, and A. Yates, “Unsupervised named-entity extraction from the web: An experimental study,” Artif. Intell., vol. 165, no. 1, pp. 91–134, 2005.
  10. 10.R. Grishman and B. Sundheim, “Message understanding conference-6: A brief history,” in COLING, vol. 1, 1996.
  11. 11.E. F. Tjong Kim Sang and F. De Meulder, “Introduction to the conll-2003 shared task: Language-independent named entity recognition,” in NAACL-HLT, 2003, pp. 142–147.
  12. 12.G. R. Doddington, A. Mitchell, M. A. Przybocki, L. A. Ramshaw, S. Strassel, and R. M. Weischedel, “The automatic content extraction (ace) program-tasks, data, and evaluation.” in LREC, vol. 2, 2004, p. 1.
  13. 13.G. Demartini, T. Iofciu, and A. P. De Vries, “Overview of the inex 2009 entity ranking track,” in INEX, 2009, pp. 254–264.
  14. 14.K. Balog, P. Serdyukov, and A. P. De Vries, “Overview of the trec 2010 entity track,” in TREC, 2010.
  15. 15.G. Petasis, A. Cucchiarelli, P. Velardi, G. Paliouras, V. Karkaletsis, and C. D. Spyropoulos, “Automatic adaptation of proper noun dictionaries through cooperation of machine learning and probabilistic methods,” in SIGIR, 2000, pp. 128–135.
  16. 16.S. A. Kripke, “Naming and necessity,” in Semantics of natural language. Springer, 1972, pp. 253–355.
  17. 17.R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, “Natural language processing (almost) from scratch,” J. Mach. Learn. Res., vol. 12, no. Aug, pp. 2493–2537, 2011.
  18. 18.Z. Huang, W. Xu, and K. Yu, “Bidirectional lstm-crf models for sequence tagging,” arXiv preprint arXiv:1508.01991, 2015.
  19. 19.G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, and C. Dyer, “Neural architectures for named entity recognition,” in NAACL, 2016, pp. 260–270.
  20. 20.J. P. Chiu and E. Nichols, “Named entity recognition with bidirectional lstm-cnns,” Trans. Assoc. Comput. Linguist., pp. 357–370, 2016.
  21. 21.M. E. Peters, W. Ammar, C. Bhagavatula, and R. Power, “Semi-supervised sequence tagging with bidirectional language models,” in ACL, 2017, pp. 1756–1765.
  22. 22.M. Marrero, J. Urbano, S. Sánchez-Cuadrado, J. Morato, and J. M. Gómez-Berbís, “Named entity recognition: fallacies, challenges and opportunities,” Comput. Stand. Interfaces, vol. 35, no. 5, pp. 482–489, 2013.
  23. 23.M. L. Patawar and M. Potey, “Approaches to named entity recognition: a survey,” Int. J. Innov. Res. Comput. Commun. Eng., vol. 3, no. 12, pp. 12 201–12 208, 2015.
  24. 24.C. J. Saju and A. Shaja, “A survey on efficient extraction of named entities from new domains using big data analytics,” in ICRTCCM, 2017, pp. 170–175.
  25. 25.X. Dai, “Recognizing complex entity mentions: A review and future directions,” in ACL, 2018, pp. 37–44.
  26. 26.V. Yadav and S. Bethard, “A survey on recent advances in named entity recognition from deep learning models,” in COLING, 2018, pp. 2145–2158.
  27. 27.A. Goyal, V. Gupta, and M. Kumar, “Recent named entity recognition and classification techniques: A systematic review,” Comput. Sci. Rev., vol. 29, pp. 21–43, 2018.
  28. 28.R. Sharnagat, “Named entity recognition: A literature survey,” Center For Indian Language Technology, 2014.
  29. 29.X. Ling and D. S. Weld, “Fine-grained entity recognition.” in AAAI, vol. 12, 2012, pp. 94–100.
  30. 30.X. Ren, W. He, M. Qu, L. Huang, H. Ji, and J. Han, “Afet: Automatic fine-grained entity typing by hierarchical partial-label embedding,” in EMNLP, 2016, pp. 1369–1378.
  31. 31.A. Abhishek, A. Anand, and A. Awekar, “Fine-grained entity type classification by jointly learning representations and label embeddings,” in EACL, 2017, pp. 797–807.
  32. 32.A. Lal, A. Tomer, and C. R. Chowdary, “Sane: System for fine grained named entity typing on textual data,” in WWW, 2017, pp. 227–230.
  33. 33.L. d. Corro, A. Abujabal, R. Gemulla, and G. Weikum, “Finet: Context-aware fine-grained named entity typing,” in EMNLP, 2015, pp. 868–878.
  34. 34.K. Balog, Entity-Oriented Search. Springer, 2018.
  35. 35.H. Raviv, O. Kurland, and D. Carmel, “Document retrieval using entity-based language models,” in SIGIR, 2016, pp. 65–74.
  36. 36.P. Boldi, F. Bonchi, C. Castillo, D. Donato, A. Gionis, and S. Vigna, “The query-flow graph: model and applications,” in CIKM, 2008, pp. 609–618.
  37. 37.F. Cai, M. De Rijke et al., “A survey of query auto completion in information retrieval,” Found. Trends® in Inf. Retr., vol. 10, no. 4, pp. 273–363, 2016.
  38. 38.Z. Bar-Yossef and N. Kraus, “Context-sensitive query auto-completion,” in WWW, 2011, pp. 107–116.
  39. 39.G. Saldanha, O. Biran, K. McKeown, and A. Gliozzo, “An entity-focused approach to generating company descriptions,” in ACL, vol. 2, 2016, pp. 243–248.
  40. 40.F. Hasibi, K. Balog, and S. E. Bratsberg, “Dynamic factual summaries for entity cards,” in SIGIR, 2017, pp. 773–782.
  41. 41.S. Pradhan, A. Moschitti, N. Xue, O. Uryupina, and Y. Zhang, “Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes,” in EMNLP, 2012, pp. 1–40.
  42. 42.S. Sekine and C. Nobata, “Definition, dictionaries and tagger for extended named entity hierarchy.” in LREC, 2004, pp. 1977–1980.
  43. 43.S. Zhang and N. Elhadad, “Unsupervised biomedical named entity recognition: Experiments with clinical and biological texts,” J. Biomed. Inform., vol. 46, no. 6, pp. 1088–1098, 2013.
  44. 44.J.-H. Kim and P. C. Woodland, “A rule-based named entity recognition system for speech input,” in ICSLP, 2000.
  45. 45.D. Hanisch, K. Fundel, H.-T. Mevissen, R. Zimmer, and J. Fluck, “Prominer: rule-based protein and gene entity recognition,” BMC Bioinform., vol. 6, no. 1, p. S14, 2005.
  46. 46.A. P. Quimbaya, A. S. Múnera, R. A. G. Rivera, J. C. D. Rodríguez, O. M. M. Velandia, A. A. G. Peña, and C. Labbé, “Named entity recognition over electronic health records through a combined dictionary-based approach,” Procedia Comput. Sci., vol. 100, pp. 55–61, 2016.
  47. 47.K. Humphreys, R. Gaizauskas, S. Azzam, C. Huyck, B. Mitchell, H. Cunningham, and Y. Wilks, “University of sheffield: Description of the lasie-ii system as used for muc-7,” in MUC-7, 1998.
  48. 48.G. Krupka and K. IsoQuest, “Description of the nerowl extractor system as used for muc-7,” in MUC-7, 2005, pp. 21–28.
  49. 49.W. J. Black, F. Rinaldi, and D. Mowatt, “Facile: Description of the ne system used for muc-7,” in MUC-7, 1998.
  50. 50.C. Aone, L. Halverson, T. Hampton, and M. Ramos-Santacruz, “Sra: Description of the ie2 system used for muc-7,” in MUC-7, 1998.
  51. 51.D. E. Appelt, J. R. Hobbs, J. Bear, D. Israel, M. Kameyama, D. Martin, K. Myers, and M. Tyson, “Sri international fastus system: Muc-6 test results and analysis,” in MUC-6, 1995, pp. 237–248.
  52. 52.A. Mikheev, M. Moens, and C. Grover, “Named entity recognition without gazetteers,” in EACL, 1999, pp. 1–8.
  53. 53.M. Collins and Y. Singer, “Unsupervised models for named entity classification,” in EMNLP, 1999, pp. 100–110.
  54. 54.D. Nadeau, P. D. Turney, and S. Matwin, “Unsupervised named-entity recognition: Generating gazetteers and resolving ambiguity,” in CSCSI, 2006, pp. 266–277.
  55. 55.S. Sekine and E. Ranchhod, Named entities: recognition, classification and use. John Benjamins Publishing, 2009, vol. 19.
  56. 56.G. Zhou and J. Su, “Named entity recognition using an hmm-based chunk tagger,” in ACL, 2002, pp. 473–480.
  57. 57.B. Settles, “Biomedical named entity recognition using conditional random fields and rich feature sets,” in ACL, 2004, pp. 104–107.
  58. 58.W. Liao and S. Veeramachaneni, “A simple semi-supervised algorithm for named entity recognition,” in NAACL-HLT, 2009, pp. 58–65.
  59. 59.A. Mikheev, “A knowledge-free method for capitalized word disambiguation,” in ACL, 1999, pp. 159–166.
  60. 60.J. Kazama and K. Torisawa, “Exploiting wikipedia as external knowledge for named entity recognition,” in EMNLP-CoNLL, 2007.
  61. 61.A. Toral and R. Munoz, “A proposal to automatically build and maintain gazetteers for named entity recognition by using wikipedia,” in Workshop on NEW TEXT Wikis and blogs and other dynamic text sources, 2006.
  62. 62.J. Hoffart, M. A. Yosef, I. Bordino, H. Fürstenau, M. Pinkal, M. Spaniol, B. Taneva, S. Thater, and G. Weikum, “Robust disambiguation of named entities in text,” in EMNLP, 2011, pp. 782–792.
  63. 63.Y. Ravin and N. Wacholder, Extracting names from natural-language text. IBM Research Report RC 2033, 1997.
  64. 64.J. Zhu, V. Uren, and E. Motta, “Espotter: Adaptive named entity recognition for web browsing,” in WM. Springer, 2005, pp. 518–529.
  65. 65.Z. Ji, A. Sun, G. Cong, and J. Han, “Joint recognition and linking of fine-grained locations from tweets,” in WWW, 2016, pp. 1271–1281.
  66. 66.V. Krishnan and C. D. Manning, “An effective two-stage model for exploiting non-local dependencies in named entity recognition,” in ACL, 2006, pp. 1121–1128.
  67. 67.D. Campos, S. Matos, and J. L. Oliveira, “Biomedical named entity recognition: a survey of machine-learning tools,” in Theory Appl. Adv. Text Min., 2012.
  68. 68.S. R. Eddy, “Hidden markov models,” Curr. Opin. Struct. Biol., vol. 6, no. 3, pp. 361–365, 1996.
  69. 69.J. R. Quinlan, “Induction of decision trees,” Mach. Learn., vol. 1, no. 1, pp. 81–106, 1986.
  70. 70.J. N. Kapur, Maximum-entropy models in science and engineering. John Wiley & Sons, 1989.
  71. 71.M. A. Hearst, S. T. Dumais, E. Osuna, J. Platt, and B. Scholkopf, “Support vector machines,” IEEE Intell. Syst. Their Appl., vol. 13, no. 4, pp. 18–28, 1998.
  72. 72.J. D. Lafferty, A. McCallum, and F. C. N. Pereira, “Conditional random fields: Probabilistic models for segmenting and labeling sequence data,” pp. 282–289, 2001.
  73. 73.D. M. Bikel, S. Miller, R. Schwartz, and R. Weischedel, “Nymble: a high-performance learning name-finder,” in ANLC, 1997, pp. 194–201.
  74. 74.D. M. Bikel, R. Schwartz, and R. M. Weischedel, “An algorithm that learns what’s in a name,” Mach. Learn., vol. 34, no. 1-3, pp. 211–231, 1999.
  75. 75.G. Szarvas, R. Farkas, and A. Kocsor, “A multilingual named entity recognition system using boosting and c4. 5 decision tree learning algorithms,” in DS. Springer, 2006, pp. 267–278.
  76. 76.A. Borthwick, J. Sterling, E. Agichtein, and R. Grishman, “Nyu: Description of the mene named entity system as used in muc-7,” in MUC-7, 1998.
  77. 77.O. Bender, F. J. Och, and H. Ney, “Maximum entropy models for named entity recognition,” in HLT-NAACL, 2003, pp. 148–151.
  78. 78.H. L. Chieu and H. T. Ng, “Named entity recognition: a maximum entropy approach using global information,” in CoNLL, 2002, pp. 1–7.
  79. 79.J. R. Curran and S. Clark, “Language independent ner using a maximum entropy tagger,” in HLT-NAACL, 2003, pp. 164–167.
  80. 80.P. McNamee and J. Mayfield, “Entity extraction without language-specific resources,” in CoNLL, 2002, pp. 1–4.
  81. 81.A. McCallum and W. Li, “Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons,” in HLT-NAACL, 2003, pp. 188–191.
  82. 82.S. Liu, Y. Sun, B. Li, W. Wang, and X. Zhao, “Hamner: Headword amplified multi-span distantly supervised method for domain specific named entity recognition,” arXiv preprint arXiv:1912.01731, 2019.
  83. 83.A. Ritter, S. Clark, O. Etzioni et al., “Named entity recognition in tweets: an experimental study,” in EMNLP, 2011, pp. 1524–1534.
  84. 84.X. Liu, S. Zhang, F. Wei, and M. Zhou, “Recognizing named entities in tweets,” in ACL, 2011, pp. 359–367.
  85. 85.T. Rocktäschel, M. Weidlich, and U. Leser, “Chemspot: a hybrid system for chemical named entity recognition,” Bioinformatics, vol. 28, no. 12, pp. 1633–1640, 2012.
  86. 86.Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436, 2015.
  87. 87.Y. Shen, H. Yun, Z. C. Lipton, Y. Kronrod, and A. Anandkumar, “Deep active learning for named entity recognition,” in ICLR, 2017.
  88. 88.T. H. Nguyen, A. Sil, G. Dinu, and R. Florian, “Toward mention detection robustness with recurrent neural networks,” arXiv preprint arXiv:1602.07749, 2016.
  89. 89.S. Zheng, F. Wang, H. Bao, Y. Hao, P. Zhou, and B. Xu, “Joint extraction of entities and relations based on a novel tagging scheme,” in ACL, 2017, pp. 1227–1236.
  90. 90.E. Strubell, P. Verga, D. Belanger, and A. McCallum, “Fast and accurate entity recognition with iterated dilated convolutions,” in ACL, 2017, pp. 2670–2680.
  91. 91.T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in ICLR, 2013.
  92. 92.J. Yang, S. Liang, and Y. Zhang, “Design challenges and misconceptions in neural sequence labeling,” in COLING, 2018, pp. 3879–3889.
  93. 93.L. Yao, H. Liu, Y. Liu, X. Li, and M. W. Anwar, “Biomedical named entity recognition based on deep neutral network,” Int. J. Hybrid Inf. Technol., vol. 8, no. 8, pp. 279–288, 2015.
  94. 94.F. Zhai, S. Potdar, B. Xiang, and B. Zhou, “Neural models for sequence chunking.” in AAAI, 2017, pp. 3365–3371.
  95. 95.P. Zhou, S. Zheng, J. Xu, Z. Qi, H. Bao, and B. Xu, “Joint extraction of multiple relations and entities by using a hybrid neural network,” in CCL-NLP-NABD. Springer, 2017, pp. 135–146.
  96. 96.X. Ma and E. Hovy, “End-to-end sequence labeling via bidirectional lstm-cnns-crf,” in ACL, 2016, pp. 1064–1074.
  97. 97.P.-H. Li, R.-P. Dong, Y.-S. Wang, J.-C. Chou, and W.-Y. Ma, “Leveraging linguistic structures for named entity recognition with bidirectional recursive neural networks,” in EMNLP, 2017, pp. 2664–2669.
  98. 98.C. Wang, K. Cho, and D. Kiela, “Code-switched named entity recognition with embedding attention,” in CALCS, 2018, pp. 154–158.
  99. 99.O. Kuru, O. A. Can, and D. Yuret, “Charner: Character-level named entity recognition,” in COLING, 2016, pp. 911–921.
  100. 100.Q. Tran, A. MacKinlay, and A. J. Yepes, “Named entity recognition with stack residual lstm and trainable bias decoding,” in IJCNLP, 2017, pp. 566–575.
  101. 101.J. Yang, Y. Zhang, and F. Dong, “Neural reranking for named entity recognition,” in RANLP, 2017, pp. 784–792.
  102. 102.M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL-HLT, 2018, pp. 2227–2237.
  103. 103.M. Gridach, “Character-level neural network for biomedical named entity recognition,” J. Biomed. Inform., vol. 70, pp. 85–91, 2017.
  104. 104.M. Rei, G. K. Crichton, and S. Pyysalo, “Attending to characters in neural sequence labeling models,” in COLING, 2016, pp. 309–318.
  105. 105.Z. Yang, R. Salakhutdinov, and W. Cohen, “Multi-task cross-lingual sequence tagging from scratch,” arXiv preprint arXiv:1603.06270, 2016.
  106. 106.A. Akbik, D. Blythe, and R. Vollgraf, “Contextual string embeddings for sequence labeling,” in COLING, 2018, pp. 1638–1649.
  107. 107.T. Liu, J. Yao, and C. Lin, “Towards improving neural named entity recognition with gazetteers,” in ACL, 2019, pp. 5301–5307.
  108. 108.A. Ghaddar and P. Langlais, “Robust lexical features for improved neural network named-entity recognition,” in COLING, 2018, pp. 1896–1907.
  109. 109.Z. Jie and W. Lu, “Dependency-guided lstm-crf for named entity recognition,” in EMNLP, 2018, pp. 3860–3870.
  110. 110.D. Lu, L. Neves, V. Carvalho, N. Zhang, and H. Ji, “Visual attention model for name tagging in multimodal social media,” in ACL, 2018, pp. 1990–1999.
  111. 111.Q. Wei, T. Chen, R. Xu, Y. He, and L. Gui, “Disease named entity recognition by combining conditional random fields and bidirectional recurrent neural networks,” Database, vol. 2016, 2016.
  112. 112.B. Y. Lin, F. Xu, Z. Luo, and K. Zhu, “Multi-channel bilstm-crf model for emerging named entity recognition in social media,” in W-NUT, 2017, pp. 160–165.
  113. 113.G. Aguilar, S. Maharjan, A. P. L. Monroy, and T. Solorio, “A multi-task approach for named entity recognition in social media data,” in W-NUT, 2017, pp. 148–153.
  114. 114.P. Jansson and S. Liu, “Distributed representation, lda topic modelling and deep learning for emerging named entity recognition from social media,” in W-NUT, 2017, pp. 154–159.
  115. 115.M. Xu, H. Jiang, and S. Watcharawittayakul, “A local detection approach for named entity recognition and mention detection,” in ACL, vol. 1, 2017, pp. 1237–1247.
  116. 116.S. Zhang, H. Jiang, M. Xu, J. Hou, and L. Dai, “A fixed-size encoding method for variable-length sequences with its application to neural network language models,” arXiv preprint arXiv:1505.01504, 2015.
  117. 117.S. Moon, L. Neves, and V. Carvalho, “Multimodal named entity recognition for short social media posts,” in NAACL, 2018, pp. 852–860.
  118. 118.J. Devlin, M. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT, 2019, pp. 4171–4186.
  119. 119.Y. Wu, M. Jiang, J. Lei, and H. Xu, “Named entity recognition in chinese clinical text using deep neural network,” MEDINFO, pp. 624–628, 2015.
  120. 120.A. Z. Gregoric, Y. Bachrach, and S. Coope, “Named entity recognition with parallel recurrent neural networks,” in ACL, vol. 2, 2018, pp. 69–74.
  121. 121.A. Katiyar and C. Cardie, “Nested named entity recognition revisited,” in ACL, vol. 1, 2018, pp. 861–871.
  122. 122.M. Ju, M. Miwa, and S. Ananiadou, “A neural layered model for nested named entity recognition,” in NAACL-HLT, vol. 1, 2018, pp. 1446–1459.
  123. 123.M. Rei, “Semi-supervised multitask learning for sequence labeling,” in ACL, 2017, pp. 2121–2130.
  124. 124.L. Liu, X. Ren, J. Shang, J. Peng, and J. Han, “Efficient contextualized representation: Language model pruning for sequence labeling,” in EMNLP, 2018, pp. 1215–1225.
  125. 125.L. Liu, J. Shang, F. Xu, X. Ren, H. Gui, J. Peng, and J. Han, “Empower sequence labeling with task-aware neural language model,” in AAAI, 2017, pp. 5253–5260.
  126. 126.C. Jia, L. Xiao, and Y. Zhang, “Cross-domain NER using cross-domain language modeling,” in ACL, 2019, pp. 2464–2474.
  127. 127.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS, 2017, pp. 5998–6008.
  128. 128.P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer, “Generating wikipedia by summarizing long sequences,” arXiv preprint arXiv:1801.10198, 2018.
  129. 129.N. Kitaev and D. Klein, “Constituency parsing with a self-attentive encoder,” in ACL, 2018, pp. 2675–2685.
  130. 130.A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” Technical Report, OpenAI, 2018.
  131. 131.A. Baevski, S. Edunov, Y. Liu, L. Zettlemoyer, and M. Auli, “Cloze-driven pretraining of self-attention networks,” CoRR, vol. abs/1903.07785, 2019.
  132. 132.C. Xia, C. Zhang, T. Yang, Y. Li, N. Du, X. Wu, W. Fan, F. Ma, and P. S. Yu, “Multi-grained named entity recognition,” in ACL, 2019, pp. 1430–1440.
  133. 133.Y. Luo, F. Xiao, and H. Zhao, “Hierarchical contextualized representation for named entity recognition,” CoRR, vol. abs/1911.02257, 2019.
  134. 134.Y. Liu, F. Meng, J. Zhang, J. Xu, Y. Chen, and J. Zhou, “GCDT: A global context enhanced deep transition architecture for sequence labeling,” in ACL, 2019, pp. 2431–2441.
  135. 135.Y. Jiang, C. Hu, T. Xiao, C. Zhang, and J. Zhu, “Improved differentiable architecture search for language modeling and named entity recognition,” in EMNLP, 2019, pp. 3576–3581.
  136. 136.X. Li, J. Feng, Y. Meng, Q. Han, F. Wu, and J. Li, “A unified MRC framework for named entity recognition,” CoRR, vol. abs/1910.11476, 2019.
  137. 137.X. Li, X. Sun, Y. Meng, J. Liang, F. Wu, and J. Li, “Dice loss for data-imbalanced NLP tasks,” CoRR, vol. abs/1911.02855, 2019.
  138. 138.L. Cui and Y. Zhang, “Hierarchically-refined label attention network for sequence labeling,” in EMNLP, 2019, pp. 4113–4126.
  139. 139.S. Tomori, T. Ninomiya, and S. Mori, “Domain specific named entity recognition referring to the real world by deep neural networks,” in ACL, vol. 2, 2016, pp. 236–242.
  140. 140.Y. Lin, L. Liu, H. Ji, D. Yu, and J. Han, “Reliability-aware dynamic feature composition for name tagging,” in ACL, 2019, pp. 165–174.
  141. 141.J. Zhuo, Y. Cao, J. Zhu, B. Zhang, and Z. Nie, “Segment-level sequence modeling using gated recursive semi-markov conditional random fields,” in ACL, vol. 1, 2016, pp. 1413–1423.
  142. 142.Z.-X. Ye and Z.-H. Ling, “Hybrid semi-markov crf for neural sequence labeling,” in ACL, 2018, pp. 235–240.
  143. 143.A. Vaswani, Y. Bisk, K. Sagae, and R. Musa, “Supertagging with lstms,” in NAACL-HLT, 2016, pp. 232–237.
  144. 144.O. Vinyals, M. Fortunato, and N. Jaitly, “Pointer networks,” in NIPS, 2015, pp. 2692–2700.
  145. 145.J. Li, A. Sun, and S. Joty, “Segbot: A generic neural text segmentation model with pointer network,” in IJCAI, 2018, pp. 4166–4172.
  146. 146.Q. Guo, X. Qiu, P. Liu, Y. Shao, X. Xue, and Z. Zhang, “Star-transformer,” in NAACL-HLT, 2019, pp. 1315–1325.
  147. 147.H. Yan, B. Deng, X. Li, and X. Qiu, “Tener: Adapting transformer encoder for name entity recognition,” arXiv preprint arXiv:1911.04474, 2019.
  148. 148.Q. Wang, Y. Zhou, T. Ruan, D. Gao, Y. Xia, and P. He, “Incorporating dictionaries into deep neural networks for the chinese clinical named entity recognition,” J. Biomed. Inform., vol. 92, 2019.
  149. 149.Y. Zhang and J. Yang, “Chinese ner using lattice lstm,” in ACL, 2018, pp. 1554–1564.
  150. 150.W. Wang, F. Bao, and G. Gao, “Mongolian named entity recognition with bidirectional recurrent neural networks,” in ICTAI, 2016, pp. 495–500.
  151. 151.J. Straková, M. Straka, and J. Hajiˇc, “Neural networks for featureless named entity recognition in czech,” in TSD, 2016, pp. 173–181.
  152. 152.M. Gridach, “Character-aware neural networks for arabic named entity recognition for social media,” in WSSANLP, 2016, pp. 23–32.
  153. 153.M. K. Malik, “Urdu named entity recognition and classification system using artificial neural network,” ACM Trans. Asian Low-Resour. Lang. Inf. Process., vol. 17, no. 1, p. 2, 2017.
  154. 154.T.-H. Pham and P. Le-Hong, “End-to-end recurrent neural network models for vietnamese named entity recognition: Word-level vs. character-level,” in PACLING, 2017, pp. 219–232.
  155. 155.K. Kurniawan and S. Louvan, “Empirical evaluation of character-based model on neural named-entity recognition in indonesian conversational texts,” arXiv preprint arXiv:1805.12291, 2018.
  156. 156.K. Yano, “Neural disease named entity extraction with character-based bilstm+ crf in japanese medical text,” arXiv preprint arXiv:1806.03648, 2018.
  157. 157.A. Bharadwaj, D. Mortensen, C. Dyer, and J. Carbonell, “Phonologically aware neural model for named entity recognition in low resource transfer settings,” in EMNLP, 2016, pp. 1462–1472.
  158. 158.J. Xie, Z. Yang, G. Neubig, N. A. Smith, and J. Carbonell, “Neural cross-lingual named entity recognition with minimal resources,” in EMNLP, 2018, pp. 369–379.
  159. 159.Y. Lin, S. Yang, V. Stoyanov, and H. Ji, “A multi-lingual multi-task architecture for low-resource sequence labeling,” in ACL, 2018, pp. 799–809.
  160. 160.R. Caruana, “Multitask learning,” Mach. learn., vol. 28, no. 1, pp. 41–75, 1997.
  161. 161.N. Peng and M. Dredze, “Multi-task domain adaptation for sequence tagging,” in RepL4NLP, 2017, pp. 91–100.
  162. 162.G. Crichton, S. Pyysalo, B. Chiu, and A. Korhonen, “A neural network multi-task learning approach to biomedical named entity recognition,” BMC Bioinform., vol. 18, no. 1, p. 368, 2017.
  163. 163.X. Wang, Y. Zhang, X. Ren, Y. Zhang, M. Zitnik, J. Shang, C. Langlotz, and J. Han, “Cross-type biomedical named entity recognition with deep multi-task learning,” arXiv preprint arXiv:1801.09851, 2018.
  164. 164.S. J. Pan, Q. Yang et al., “A survey on transfer learning,” IEEE Trans. Knowl. Data Eng., vol. 22, no. 10, pp. 1345–1359, 2010.
  165. 165.J. Jiang and C. Zhai, “Instance weighting for domain adaptation in nlp,” in ACL, 2007, pp. 264–271.
  166. 166.D. Wu, W. S. Lee, N. Ye, and H. L. Chieu, “Domain adaptive bootstrapping for named entity recognition,” in EMNLP, 2009, pp. 1523–1532.
  167. 167.A. Chaudhary, J. Xie, Z. Sheikh, G. Neubig, and J. G. Carbonell, “A little annotation does a lot of good: A study in bootstrapping low-resource named entity recognizers,” pp. 5163–5173, 2019.
  168. 168.S. J. Pan, Z. Toh, and J. Su, “Transfer joint embedding for cross-domain named entity recognition,” ACM Trans. Inf. Syst., vol. 31, no. 2, p. 7, 2013.
  169. 169.J. Y. Lee, F. Dernoncourt, and P. Szolovits, “Transfer learning for named-entity recognition with neural networks,” arXiv preprint arXiv:1705.06273, 2017.
  170. 170.B. Y. Lin and W. Lu, “Neural adaptation layers for cross-domain named entity recognition,” in EMNLP, 2018, pp. 2012–2022.
  171. 171.Y. Cao, Z. Hu, T. Chua, Z. Liu, and H. Ji, “Low-resource name tagging learned with weakly labeled data,” in EMNLP, 2019, pp. 261–270.
  172. 172.X. Huang, L. Dong, E. Boschee, and N. Peng, “Learning A unified named entity tagger from multiple partially annotated corpora for efficient adaptation,” in CoNLL, 2019, pp. 515–527.
  173. 173.L. Qu, G. Ferraro, L. Zhou, W. Hou, and T. Baldwin, “Named entity recognition for novel types by transfer learning,” in EMNLP, 2016, pp. 899–905.
  174. 174.Z. Yang, R. Salakhutdinov, and W. W. Cohen, “Transfer learning for sequence tagging with hierarchical recurrent networks,” in ICLR, 2017.
  175. 175.P. von Däniken and M. Cieliebak, “Transfer learning and sentence level features for named entity recognition on tweets,” in W-NUT, 2017, pp. 166–171.
  176. 176.H. Zhao, Y. Yang, Q. Zhang, and L. Si, “Improve neural entity recognition via multi-task data selection and constrained decoding,” in NAACL-HLT, vol. 2, 2018, pp. 346–351.
  177. 177.G. Beryozkin, Y. Drori, O. Gilon, T. Hartman, and I. Szpektor, “A joint named-entity recognizer for heterogeneous tag-sets using a tag hierarchy,” in ACL, 2019, pp. 140–150.
  178. 178.J. M. Giorgi and G. D. Bader, “Transfer learning for biomedical named entity recognition with neural networks,” Bioinformatics, 2018.
  179. 179.Z. Wang, Y. Qu, L. Chen, J. Shen, W. Zhang, S. Zhang, Y. Gao, G. Gu, K. Chen, and Y. Yu, “Label-aware double transfer learning for cross-specialty medical named entity recognition,” in NAACL-HLT, 2018, pp. 1–15.
  180. 180.B. Settles, “Active learning,” Synth. Lect. Artif. Intell. Mach. Learn., vol. 6, no. 1, pp. 1–114, 2012.
  181. 181.D. D. Lewis and W. A. Gale, “A sequential algorithm for training text classifiers,” in SIGIR, 1994, pp. 3–12.
  182. 182.S. Pradhan, A. Moschitti, N. Xue, H. T. Ng, A. Björkelund, O. Uryupina, Y. Zhang, and Z. Zhong, “Towards robust linguistic analysis using ontonotes,” in CoNLL, 2013, pp. 143–152.
  183. 183.L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” J. Artif. Intell. Res., vol. 4, pp. 237–285, 1996.
  184. 184.R. S. Sutton and A. G. Barto, Introduction to reinforcement learning. MIT press Cambridge, 1998, vol. 135.
  185. 185.S. C. Hoi, D. Sahoo, J. Lu, and P. Zhao, “Online learning: A comprehensive survey,” arXiv preprint arXiv:1802.02871, 2018.
  186. 186.K. Narasimhan, A. Yala, and R. Barzilay, “Improving information extraction by acquiring external evidence with reinforcement learning,” in EMNLP, 2016, pp. 2355–2365.
  187. 187.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, p. 529, 2015.
  188. 188.Y. Yang, W. Chen, Z. Li, Z. He, and M. Zhang, “Distantly supervised NER with partial annotation learning and reinforcement learning,” in COLING, 2018, pp. 2159–2169.
  189. 189.D. Lowd and C. Meek, “Adversarial learning,” in SIGKDD, 2005, pp. 641–647.
  190. 190.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NIPS, 2014, pp. 2672–2680.
  191. 191.L. Huang, H. Ji, and J. May, “Cross-lingual multi-level adversarial transfer to enhance low-resource name tagging,” in NAACL-HLT, 2019, pp. 3823–3833.
  192. 192.J. Li, D. Ye, and S. Shang, “Adversarial transfer for named entity boundary detection with pointer networks,” in IJCAI, 2019, pp. 5053–5059.
  193. 193.P. Cao, Y. Chen, K. Liu, J. Zhao, and S. Liu, “Adversarial transfer learning for chinese named entity recognition with self-attention mechanism,” in EMNLP, 2018, pp. 182–192.
  194. 194.J. T. Zhou, H. Zhang, D. Jin, H. Zhu, M. Fang, R. S. M. Goh, and K. Kwok, “Dual adversarial neural transfer for low-resource named entity recognition,” in ACL, 2019, pp. 3461–3471.
  195. 195.D. Britz, “Attention and memory in deep learning and nlp,” Online: http://www. wildml. com/2016/01/attention-and-memory-in-deeplearning-and-nlp, 2016.
  196. 196.A. Zukov-Gregoric, Y. Bachrach, P. Minkovsky, S. Coope, and B. Maksak, “Neural named entity recognition using a self-attention mechanism,” in ICTAI, 2017, pp. 652–656.
  197. 197.G. Xu, C. Wang, and X. He, “Improving clinical named entity recognition with global neural attention,” in APWeb-WAIM, 2018, pp. 264–279.
  198. 198.Q. Zhang, J. Fu, X. Liu, and X. Huang, “Adaptive co-attention network for named entity recognition in tweets,” in AAAI, 2018.
  199. 199.L. Derczynski, E. Nichols, M. van Erp, and N. Limsopatham, “Results of the wnut2017 shared task on novel and emerging entity recognition,” in W-NUT, 2017, pp. 140–147.
  200. 200.J. Fisher and A. Vlachos, “Merge and label: A novel neural network architecture for nested NER,” in ACL, 2019, pp. 5840–5850.
  201. 201.D. Ye, Z. Xing, C. Y. Foo, Z. Q. Ang, J. Li, and N. Kapre, “Software-specific named entity recognition in software engineering social content,” in SANER, 2016, pp. 90–101.
  202. 202.I. Partalas, C. Lopez, N. Derbas, and R. Kalitvianski, “Learning to search for recognizing named entities in twitter,” in W-NUT, 2016, pp. 171–177.
  203. 203.W. Shen, J. Han, J. Wang, X. Yuan, and Z. Yang, “Shine+: A general framework for domain-specific entity linking with heterogeneous information networks,” IEEE Trans. Knowl. Data Eng., vol. 30, no. 2, pp. 353–366, 2018.
  204. 204.M. C. Phan, A. Sun, Y. Tay, J. Han, and C. Li, “Pair-linking for collective entity disambiguation: Two could be better than all,” arXiv preprint arXiv:1802.01074, 2018.
  205. 205.C. Li and A. Sun, “Extracting fine-grained location with temporal awareness in tweets: A two-stage approach,” J. Assoc. Inf. Sci. Technol., vol. 68, no. 7, pp. 1652–1670, 2017.
  206. 206.J. Han, A. Sun, G. Cong, W. X. Zhao, Z. Ji, and M. C. Phan, “Linking fine-grained locations in user comments,” IEEE Trans. Knowl. Data Eng., vol. 30, no. 1, pp. 59–72, 2018.
  207. 207.M. C. Phan and A. Sun, “Collective named entity recognition in user comments via parameterized label propagation,” J. Assoc. Inf. Sci. Technol., 2019.
  208. 208.Z. Batmaz, A. Yurekli, A. Bilge, and C. Kaleli, “A review on deep learning for recommender systems: challenges and remedies,” Artif. Intell. Rev., pp. 1–37, 2018.
  209. 209.M. Röder, R. Usbeck, and A. N. Ngomo, “GERBIL - benchmarking named entity recognition and linking consistently,” Semantic Web, vol. 9, no. 5, pp. 605–625, 2018.
  210. 210.M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in NAACL-HLT, 2019, pp. 48–53.
  211. 211.F. Dernoncourt, J. Y. Lee, and P. Szolovits, “NeuroNER: an easy-to-use program for named-entity recognition based on neural networks,” in EMNLP, 2017, pp. 97–102.

Citation

MLA
Li, J., et al. “A Survey on Deep Learning for Named Entity Recognition”. IEEE Transactions on Knowledge and Data Engineering, vol. 34, no. 1, 2022, pp. 50–70, https://doi.org/10.1109/TKDE.2020.2981314.
APA
Li, J., Sun, A., Han, J., & Li, C. (2022). A Survey on Deep Learning for Named Entity Recognition. IEEE Transactions on Knowledge and Data Engineering, 34(1), 50–70. https://doi.org/10.1109/TKDE.2020.2981314
Chicago
Li, J., A. Sun, J. Han, and C. Li. 2022. “A Survey on Deep Learning for Named Entity Recognition”. IEEE Transactions on Knowledge and Data Engineering 34 (1): 50–70. https://doi.org/10.1109/TKDE.2020.2981314.
Harvard
Li, J. et al. (2022) “A Survey on Deep Learning for Named Entity Recognition”, IEEE Transactions on Knowledge and Data Engineering, 34(1), pp. 50–70. Available at: https://doi.org/10.1109/TKDE.2020.2981314.
Vancouver
1. Li J, Sun A, Han J, Li C (2022) A Survey on Deep Learning for Named Entity Recognition. IEEE Transactions on Knowledge and Data Engineering 34:50–70

BibTeX

@article{Li_2022, title={A Survey on Deep Learning for Named Entity Recognition}, volume={34}, ISSN={2326-3865}, url={http://dx.doi.org/10.1109/TKDE.2020.2981314}, DOI={10.1109/tkde.2020.2981314}, number={1}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Li, Jing and Sun, Aixin and Han, Jianglei and Li, Chenliang}, year={2022}, month=Jan, pages={50–70} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF