GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

Urchade ZaratianaNadi TomehPierre HolatThierry Charnois

article2024NAACL120 citations

Proposes GLiNER, a compact bidirectional transformer model that performs parallel open-domain entity extraction by matching span representations to arbitrary type embeddings, outperforming much larger autoregressive models like ChatGPT on zero-shot benchmarks at a fraction of the computational cost.

Listen

Extracting key information such as names, dates, and organizations from unstructured text—known as Named Entity Recognition (NER)—is essential for modern data processing and knowledge graph construction. Traditional models are constrained to a rigid, predefined set of entity types and require costly retraining to expand. While recent Large Language Models (LLMs) can extract arbitrary, user-defined entity types on demand, their multi-billion parameter sizes cause slow processing speeds, high infrastructure costs, and expensive application programming interface (API) fees, making them impractical for resource-constrained environments.

The article introduces and evaluates GLiNER, a compact and generalist model designed to extract any entity type without requiring fixed categories or large-scale generative hardware. Rather than treating entity extraction as a sequential word-generation task, the model matches entity type prompts directly against text spans within a shared representation space.

To establish credibility and broad applicability, the approach was evaluated across multiple standardized benchmarks covering out-of-domain evaluation, 20 distinct domain datasets (including biomedical literature, news, and social media), and multilingual benchmarks spanning 11 languages. GLiNER models ranging from 50 million to 300 million parameters were trained on a diverse dataset of approximately 45,000 passages containing 13,000 unique entity types, and their performance was evaluated primarily under zero-shot conditions without task-specific fine-tuning.

The findings show that GLiNER delivers superior accuracy while requiring a fraction of the computational footprint. In zero-shot out-of-domain benchmarks, the large variant of GLiNER (300M parameters) achieved an average F1-score of 60.9, outperforming ChatGPT (47.5) and larger specialized LLMs like the 11-billion parameter InstructUIE (47.2) and the 13-billion parameter UniNER (55.6). Even the smallest 50M parameter version surpassed ChatGPT with a score of 52.7. Across 20 varied English datasets, GLiNER maintained an overall lead, achieving top performance in 13 benchmarks. Additionally, in multilingual zero-shot testing across 11 languages, a multilingual version of GLiNER surpassed ChatGPT's performance in eight languages despite being trained solely on English data.

These results demonstrate that organizations do not need to deploy massive, costly generative LLMs for open-vocabulary entity extraction. Adopting smaller bidirectional models reduces hardware requirements—enabling efficient execution on standard central processing units (CPUs)—lowers operational cloud expenses, and significantly increases text processing throughput by processing candidate spans in parallel rather than token-by-token.

Technical leaders and practitioners seeking flexible entity extraction should consider adopting compact span-matching models as a cost-effective alternative to LLM APIs or multi-billion parameter models. When implementing this architecture, teams should include negative entity sampling during training, as tests confirmed that balancing present and absent entity types improves extraction accuracy and reduces false positives.

Confidence in these results is high across diverse standard and out-of-domain benchmarks; however, two operational limitations warrant caution. First, the architecture cannot extract discontinuous entities, where an entity is split across non-adjacent words in a sentence. Second, the model showed reduced accuracy on informal, noisy social media text and underperformed on non-Latin scripts when using English-only backbones, indicating that targeted pretraining or multilingual encoders are necessary for those specific use cases.

arXiv: 2311.08526urchade/GLiNER

No sufficiently relevant recommendations were found.

Cover for GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

Abstract

Named Entity Recognition (NER) is essential in various Natural Language Processing (NLP) applications. Traditional NER models are effective but limited to a set of predefined entity types. In contrast, Large Language Models (LLMs) can extract arbitrary entities through natural language instructions, offering greater flexibility. However, their size and cost, particularly for those accessed via APIs like ChatGPT, make them impractical in resource-limited scenarios. In this paper, we introduce a compact NER model trained to identify any type of entity. Leveraging a bidirectional transformer encoder, our model, GLiNER, facilitates parallel entity extraction, an advantage over the slow sequential token generation of LLMs. Through comprehensive testing, GLiNER demonstrate strong performance, outperforming both ChatGPT and fine-tuned LLMs in zero-shot evaluations on various NER benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.2 Training
  • 2.3 Decoding algorithm
  • 3 Experimental Setting
  • 3.1 Training data
  • 3.2 Hyperparameters
  • 3.3 Baselines
  • 3.4 Evaluation
  • 4 Results
  • 4.1 Zero-shot on English datasets
  • 4.2 Zero-Shot Multilingual Evaluation
  • 4.3 In-domain Supervised tuning
  • 5 Further analysis and ablations
  • 5.1 Effect of Different Backbones
  • 5.2 Effect of Pretraining on In-domain Performance
  • 5.3 Ablations
  • 6 Related Works
  • 7 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A.2 Prompt

Knowls

  1. Knowl 1 — GLiNER Architecture and Input Encoding

    model/method

    GLiNER (Generalist Model for Named Entity Recognition) extracts arbitrary entity types by matching entity type representations and text span representations in a shared latent space using a Bidirectional Language Model (BiLM) encoder (such as DeBERTa-v3 or BERT).

    The input to the BiLM is structured as a single sequence concatenating MM natural-language entity types t0,t1,…,tM−1t_0, t_1, \dots, t_{M-1} and NN input text words x0,x1,…,xN−1x_0, x_1, \dots, x_{N-1}. Each entity type is preceded by a special learned token [ENT][\text{ENT}], and the entity list is separated from the text by a delimiter token [SEP][\text{SEP}]: Input=[ENT] t0 [ENT] t1…[ENT] tM−1 [SEP] x0 x1…xN−1\text{Input} = [\text{ENT}] \, t_0 \, [\text{ENT}] \, t_1 \dots [\text{ENT}] \, t_{M-1} \, [\text{SEP}] \, x_0 \, x_1 \dots x_{N-1}

    The BiLM processes all tokens bidirectionally, allowing contextual interactions between all entity types and text words. Representations are extracted as follows:

    • Entity representation: The encoder hidden states corresponding to the MM [ENT][\text{ENT}] tokens, denoted by p={pi}i=0M−1∈RM×Dp = \{p_i\}_{i=0}^{M-1} \in \mathbb{R}^{M \times D} (where DD is the hidden dimension), are passed through a two-layer feedforward network (FFN) to yield refined entity embeddings q={qi}i=0M−1∈RM×Dq = \{q_i\}_{i=0}^{M-1} \in \mathbb{R}^{M \times D}.
    • Word representation: For each input word xix_i, the contextualized representation of its first subword token from the encoder output is taken, yielding h={hi}i=0N−1∈RN×Dh = \{h_i\}_{i=0}^{N-1} \in \mathbb{R}^{N \times D}.
  2. Knowl 2 — Span Representation, Entity Matching Score, and Training Objective

    equation

    In GLiNER, a candidate text span starting at word position ii and ending at word position jj (0≤j−i<K0 \le j - i < K, where K=12K=12 is the maximum span width) is represented by concatenating its boundary word representations and projecting them through a two-layer feedforward network: Sij=FFN(hi⊗hj)∈RDS_{ij} = \text{FFN}(h_i \otimes h_j) \in \mathbb{R}^D where hi,hj∈RDh_i, h_j \in \mathbb{R}^D are the contextualized embeddings of the first subword of words ii and jj, ⊗\otimes denotes vector concatenation, and DD is the embedding dimension.

    The matching score ϕ(i,j,t)\phi(i, j, t) between span (i,j)(i, j) and entity type tt (with embedding qt∈RDq_t \in \mathbb{R}^D) is computed via dot product and sigmoid activation: ϕ(i,j,t)=σ(Sij⊤qt)∈(0,1)\phi(i, j, t) = \sigma(S_{ij}^\top q_t) \in (0, 1) where σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}. ϕ(i,j,t)\phi(i, j, t) represents the model's predicted probability that span (i,j)(i, j) is an instance of entity type tt.

    The model is trained across candidate spans S\mathcal{S} and entity types T\mathcal{T} using Binary Cross-Entropy loss: LBCE=−∑s∈S×T[Is∈Plog⁡ϕ(s)+Is∈Nlog⁡(1−ϕ(s))]\mathcal{L}_{\text{BCE}} = -\sum_{s \in \mathcal{S} \times \mathcal{T}} \left[ \mathbb{I}_{s \in \mathcal{P}} \log \phi(s) + \mathbb{I}_{s \in \mathcal{N}} \log (1 - \phi(s)) \right] where P\mathcal{P} is the set of positive span-entity type pairs present in the ground-truth annotations, N\mathcal{N} is the set of negative pairs, and I\mathbb{I} is the indicator function.

  3. Knowl 3 — Priority-Queue Greedy Decoding for Flat and Nested NER

    algorithm

    GLiNER performs inference using a greedy span selection algorithm over candidate spans whose predicted matching score exceeds a threshold of 0.50.5. The algorithm supports both flat NER (no overlapping spans permitted) and nested NER (fully contained nested spans permitted, but partial boundary overlaps disallowed). Decoding is implemented with a priority queue over candidate spans, achieving O(nlog⁡n)O(n \log n) time complexity, where nn is the number of candidate spans satisfying ϕ(i,j,t)>0.5\phi(i, j, t) > 0.5.

    Input: Candidate span predictions {(i,j,t,ϕ(i,j,t))}\{(i, j, t, \phi(i, j, t))\} with ϕ(i,j,t)>0.5\phi(i, j, t) > 0.5, decoding mode ∈{"flat","nested"}\in \{\text{"flat"}, \text{"nested"}\}
    Output: Selected entity spans E\mathcal{E}
    Initialize priority queue QQ with all candidates, sorted descending by score ϕ(i,j,t)\phi(i, j, t)
    Initialize E←∅\mathcal{E} \leftarrow \emptyset
    while QQ is not empty do
        pop candidate s=(i,j,t)s = (i, j, t) with highest score from QQ
        if mode is "flat" then
            if span (i,j)(i, j) does not overlap with any span in E\mathcal{E} then
                add ss to E\mathcal{E}
        else if mode is "nested" then
            if for all (i′,j′,t′)∈E(i', j', t') \in \mathcal{E}, either spans are disjoint, (i,j)⊆(i′,j′)(i, j) \subseteq (i', j'), or (i′,j′)⊆(i,j)(i', j') \subseteq (i, j) then
                add ss to E\mathcal{E}
    return E\mathcal{E}
  4. Knowl 4 — Zero-Shot Out-of-Domain Named Entity Recognition Performance

    data/table

    Zero-shot named entity recognition performance was evaluated across seven out-of-domain datasets from CrossNER (AI, Literature, Music, Politics, Science) and MIT (Movie, Restaurant) without any dataset-specific fine-tuning. Models were evaluated by exact match F1 score.

    Model Params Movie Restaurant AI Literature Music Politics Science Average
    Vicuna-7B 7B 6.0 5.3 12.8 16.1 17.0 20.5 13.0 13.0
    Vicuna-13B 13B 0.9 0.4 22.7 22.7 26.6 27.0 22.0 17.5
    USM 0.3B 37.7 17.7 28.2 56.0 44.9 36.1 44.0 37.8
    ChatGPT – 5.3 32.8 52.4 39.8 66.6 68.5 67.0 47.5
    InstructUIE 11B 63.0 21.0 49.0 47.2 53.2 48.1 49.2 47.2
    UniNER-7B 7B 42.4 31.7 53.6 59.3 67.0 60.9 61.1 53.7
    UniNER-13B 13B 48.7 36.2 54.2 60.9 64.5 61.4 63.5 55.6
    GoLLIE 7B 63.0 43.4 59.1 62.7 67.8 57.2 55.5 58.0
    GLiNER-S 50M 46.9 33.3 50.7 60.0 60.9 61.5 55.6 52.7
    GLiNER-M 90M 42.9 37.3 51.8 59.7 69.4 68.6 58.1 55.4
    GLiNER-L 0.3B 57.2 42.9 57.2 64.4 69.6 72.6 62.6 60.9

    GLiNER-L (0.3B parameters, based on DeBERTa-v3-large) achieves an average F1 score of 60.9, outperforming ChatGPT (47.5 F1), InstructUIE-11B (47.2 F1), UniNER-13B (55.6 F1), and GoLLIE-7B (58.0 F1). GLiNER-M (90M parameters) matches UniNER-13B (55.4 vs. 55.6 F1) despite having roughly 140 times fewer parameters.

  5. Knowl 5 — Zero-Shot Benchmark Evaluation on Twenty Standard NER Datasets

    data/table

    Zero-shot exact-match F1 performance of GLiNER-L (0.3B parameters) trained solely on the Pile-NER dataset compared against ChatGPT and UniNER-7B across 20 diverse supervised NER datasets spanning biomedical, news, web, and social media domains:

    Dataset ChatGPT UniNER-7B GLiNER-L
    ACE05 26.6 36.9 27.3
    AnatEM 30.7 25.1 33.3
    bc2gm 40.2 46.2 47.9
    bc4chemd 35.5 47.9 43.1
    bc5cdr 52.4 68.0 66.4
    Broad Tweeter 61.8 67.9 61.2
    CoNLL03 52.5 72.2 64.6
    FabNER 15.3 24.8 23.6
    FindVehicle 10.5 22.2 41.9
    GENIA 41.6 54.1 55.5
    HarveyNER 11.6 18.2 22.7
    MIT Movie 5.3 42.4 57.2
    MIT Restaurant 32.8 31.7 42.9
    MultiNERD 58.1 59.3 59.7
    ncbi 42.1 60.4 61.9
    OntoNotes 29.7 27.8 32.2
    PolyglotNER 33.6 41.8 42.9
    TweetNER7 40.1 42.7 41.4
    WikiANN 52.0 55.4 58.9
    WikiNeural 57.7 69.2 71.8
    Average 36.5 45.7 47.8

    GLiNER-L achieves an average F1 score of 47.8 across the 20 benchmarks, outperforming UniNER-7B (45.7 F1) and ChatGPT (36.5 F1). GLiNER-L obtains the highest F1 score on 13 of the 20 datasets, though UniNER-7B retains an advantage on informal social media data (Broad Tweeter and TweetNER7).

  6. Knowl 6 — Cross-Lingual Zero-Shot Transfer on MultiCoNER Benchmark

    data/table

    Cross-lingual zero-shot generalization was evaluated on the MultiCoNER benchmark across 11 languages. GLiNER models were trained exclusively on English Pile-NER data. GLiNER-En uses DeBERTa-v3-large, GLiNER-Multi uses mDeBERTa-v3-base, Sup. denotes supervised XLM-R models trained separately on each target language, and ChatGPT results are zero-shot.

    Script Language Sup. ChatGPT GLiNER-En GLiNER-Multi
    Latin German 64.6 37.1 35.6 39.5
    English 62.7 37.2 42.4 41.7
    Spanish 58.7 34.7 38.7 42.1
    Dutch 62.6 35.7 35.6 38.9
    Non-Latin Bengali 39.7 23.3 0.89 25.9
    Persian 52.3 25.9 14.9 30.2
    Hindi 47.8 27.3 11.3 27.8
    Korean 55.8 30.0 20.5 28.7
    Russian 59.7 27.4 30.3 33.3
    Turkish 46.8 31.9 22.0 30.0
    Chinese 53.1 18.8 6.59 24.3
    Average 54.9 29.9 23.6 32.9

    Despite having seen only English during training, GLiNER-Multi achieves an average F1 score of 32.9, outperforming ChatGPT (29.9 F1) overall and exceeding ChatGPT's performance in 8 of the 10 non-English languages. GLiNER-En performs competitively on Latin-script languages but experiences severe degradation on non-Latin scripts (e.g., 0.89 F1 on Bengali).

  7. Knowl 7 — Supervised In-Domain Fine-Tuning Performance

    data/table

    In-domain supervised fine-tuning performance across 20 NER datasets when fine-tuned on pooled training sets (subsampling up to 10,000 examples per dataset). Evaluated models include InstructUIE (11B), UniNER-7B, and GLiNER-L (0.3B), with and without (w/o) Pile-NER pretraining.

    Dataset InstructUIE (w/o) UniNER-7B (w/) GLiNER-L (w/) GLiNER-L (w/o)
    ACE05 79.9 86.7 82.8 81.3
    AnatEM 88.5 88.5 88.9 88.4
    bc2gm 80.7 82.4 83.7 82.0
    bc4chemd 87.6 89.2 87.9 86.7
    bc5cdr 89.0 89.3 88.7 88.7
    Broad Twitter 80.3 81.2 82.5 82.7
    CoNLL03 91.5 93.3 92.6 92.5
    FabNER 78.4 81.9 77.8 74.8
    FindVehicle 87.6 98.3 95.7 95.2
    GENIA 75.7 77.5 78.9 77.4
    HarveyNER 74.7 74.2 68.6 67.4
    MIT Movie 89.6 90.2 87.9 87.5
    MIT Restaurant 82.6 82.3 83.6 83.3
    MultiNERD 90.3 93.7 93.8 93.3
    ncbi 86.2 87.0 87.8 87.1
    OntoNotes 88.6 89.9 89.0 88.1
    PolyglotNER 53.3 65.7 61.5 60.6
    TweetNER7 65.9 65.8 51.4 50.3
    WikiANN 64.5 84.9 83.7 82.8
    WikiNeural 88.3 93.3 91.3 91.4
    Average 81.2 84.8 82.9 82.1

    GLiNER-L with Pile-NER pretraining obtains an average F1 of 82.9, outperforming InstructUIE (81.2 F1) despite being approximately 30 times smaller. It obtains the highest score on 7 of the 20 datasets, while trailing UniNER-7B (84.8 F1).

  8. Knowl 8 — Negative Entity Sampling and Prompt Dropping Regularization

    empirical result

    The Pile-NER dataset natively contains only positive entity types (types that appear in the given passage). Training without negative entity types biases the model toward false positives, resulting in lower precision. To mitigate this, negative entity sampling (sampling entity types from other instances in the batch) and entity type dropping (randomly varying the number of prompted entity types up to 25 per instance) are used.

    Ablation on negative sampling ratios on the Out-of-Domain (OOD) NER benchmark yields:

    • 0% negatives: Precision 49.3%, Recall 58.1%, F1 53.3%
    • 50% negatives: Precision 62.3%, Recall 59.7%, F1 60.9%
    • 75% negatives: Precision 61.1%, Recall 56.5%, F1 58.6%

    A 50% negative ratio produces the optimal precision-recall balance.

    Ablating both techniques on the OOD benchmark demonstrates their combined impact:

    • Full model (negative sampling + entity dropping): 60.90 F1
    • Without entity dropping: 59.50 F1 (a decrease of 1.40 F1)
    • Without negative sampling: 53.30 F1 (a decrease of 7.60 F1)
    • Without both: 52.30 F1 (a decrease of 8.60 F1)
  9. Knowl 9 — Comparison of Pretrained Bidirectional Encoder Backbones

    empirical result

    Evaluating GLiNER-M (base size, 90M parameters) across different pretrained bidirectional language models on both the OOD NER benchmark and the 20 NER datasets benchmark shows:

    • DeBERTa-v3: Achieves the highest performance on both benchmarks (~55.4 F1 on OOD, ~45.0 F1 on 20 NER datasets).
    • ELECTRA and ALBERT: Demonstrate strong performance, slightly below DeBERTa-v3 (~51.5--52.5 F1 on OOD, ~43.5--44.0 F1 on 20 NER datasets).
    • BERT and RoBERTa: Lag behind with similar scores (~49.5--50.0 F1 on OOD, ~42.0--42.5 F1 on 20 NER datasets).
    • XLNet: Failed to generalize to the span-matching formulation, reaching at most 3.0 F1 on OOD despite extensive hyperparameter tuning.

    Even lower-performing backbones like BERT-base (~49 F1 on OOD) outperform zero-shot ChatGPT (47.5 F1) and InstructUIE-11B (47.2 F1).

  10. Knowl 10 — Impact of Pile-NER Pretraining Across Supervised Training Scales

    empirical result

    Pretraining GLiNER-L on the synthetic Pile-NER dataset provides consistent positive transfer across varying amounts of in-domain supervised training data per dataset on the 20 NER datasets benchmark:

    • 100 samples per dataset: F1 increases by +5.6+5.6 points (from ≈64.9\approx 64.9 to ≈70.5\approx 70.5 F1).
    • 500 samples per dataset: F1 increases by +1.6+1.6 points.
    • 1,000 samples per dataset: F1 increases by +1.3+1.3 points.
    • 5,000 samples per dataset: F1 increases by +1.2+1.2 points.
    • 10,000 samples per dataset (full data): F1 increases by +0.8+0.8 points (from 82.182.1 to 82.982.9 F1).

    The performance gain from synthetic open-type pretraining is most pronounced in low-resource regimes.

  11. Knowl 11 — Limitations in Discontinuous Mentions and Exact-Match Evaluation

    limitation

    GLiNER exhibits two primary limitations:

    1. Inability to extract discontinuous entities: Because span representations Sij=FFN(hi⊗hj)S_{ij} = \text{FFN}(h_i \otimes h_j) are defined exclusively over contiguous text spans spanning word start index ii to end index jj, the model cannot recognize discontinuous or split entity mentions, unlike some generative autoregressive LLMs.
    2. Exact-match evaluation sensitivity: The standard evaluation metric requires exact match of span boundaries and entity types. It does not account for partial matches or context-dependent semantic nuances in open-vocabulary entity definitions.

Coverage note — No substantial contributed material was omitted. All architectures, loss formulations, decoding algorithms, regularization strategies, benchmark results, ablations, and limitations are fully covered.

References

  1. 1.Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David A. Sontag. 2022. Large language models are few-shot clinical information extractors. In Conference on Empirical Methods in Natural Language Processing.
  2. 2.Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1638–1649, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  3. 3.Rami Al-Rfou, Vivek Kulkarni, Bryan Perozzi, and Steven Skiena. 2015. Polyglot-ner: Massive multilingual named entity recognition. In Proceedings of the 2015 SIAM International Conference on Data Mining, pages 586–594. SIAM.
  4. 4.Dhananjay Ashok and Zachary Chase Lipton. 2023. Promptner: Prompting for named entity recognition. ArXiv, abs/2305.15444.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  6. 6.Pei Chen, Haotian Xu, Cheng Zhang, and Ruihong Huang. 2022. Crossroads, buildings and neighborhoods: A dataset for fine-grained location recognition. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3329–3339, Seattle, United States. Association for Computational Linguistics.
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models.
  9. 9.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators.
  10. 10.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. In Annual Meeting of the Association for Computational Linguistics.
  11. 11.Leyang Cui, Yu Wu, Jian Liu, Sen Yang, and Yue Zhang. 2021. Template-based named entity recognition using bart. In Findings.
  12. 12.Leon Derczynski, Kalina Bontcheva, and Ian Roberts. 2016. Broad Twitter corpus: A diverse named entity recognition resource. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1169–1179, Osaka, Japan. The COLING 2016 Organizing Committee.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics.
  14. 14.Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014. Ncbi disease corpus: a resource for disease name recognition and concept normalization. Journal of biomedical informatics, 47:1–10.
  15. 15.Jinlan Fu, Xuanjing Huang, and Pengfei Liu. 2021. SpanNER: Named entity re-/recognition as span prediction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7183–7195, Online. Association for Computational Linguistics.
  16. 16.Leo Gao, Stella Rose Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The pile: An 800gb dataset of diverse text for language modeling. ArXiv, abs/2101.00027.
  17. 17.Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West. 2023. Grammar-constrained decoding for structured NLP tasks without finetuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10932–10952, Singapore. Association for Computational Linguistics.
  18. 18.Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Xiaohui Zhu, Jeremy Smith, Eng Gee Lim, and Yutao Yue. 2023. Findvehicle and vehiclefinder: A ner dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system. arXiv preprint arXiv:2304.10893.
  19. 19.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. ArXiv, abs/2111.09543.
  20. 20.Zhiheng Huang, Wei Xu, and Kai Yu. 2015. Bidirectional lstm-crf models for sequence tagging.
  21. 21.J-D Kim, Tomoko Ohta, Yuka Tateisi, and Jun’ichi Tsujii. 2003. Genia corpus—a semantically annotated corpus for bio-textmining. Bioinformatics, 19(suppl_1):i180–i182.
  22. 22.Martin Krallinger, Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M Lowe, et al. 2015. The chemdner corpus of chemicals and drugs and its annotation principles. Journal of cheminformatics, 7(1):1–17.
  23. 23.Aman Kumar and Binil Starly. 2022. “fabner”: information extraction from manufacturing process science domain literature using named entity recognition. Journal of Intelligent Manufacturing, 33(8):2393–2407.
  24. 24.John D. Lafferty, Andrew McCallum, and Fernando Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In International Conference on Machine Learning.
  25. 25.Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning. ArXiv, abs/2304.05613.
  26. 26.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In North American Chapter of the Association for Computational Linguistics.
  27. 27.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. ArXiv, abs/1909.11942.
  28. 28.Dongfang Li, Baotian Hu, and Qingcai Chen. 2022. Prompt-based text entailment for low-resource named entity recognition. ArXiv, abs/2211.03039.
  29. 29.Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016. Biocreative v cdr task corpus: a resource for chemical disease relation extraction. Database, 2016.
  30. 30.Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2019. A unified mrc framework for named entity recognition. ArXiv, abs/1910.11476.
  31. 31.Yangming Li, lemao liu, and Shuming Shi. 2021. Empirical analysis of unlabeled entity problem in named entity recognition. In International Conference on Learning Representations.
  32. 32.Jingjing Liu, Panupong Pasupat, D. Scott Cyphers, and James R. Glass. 2013. Asgard: A portable architecture for multilingual dialogue systems. 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 8386–8390.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  34. 34.Zihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai, Ziwei Ji, Samuel Cahyawijaya, Andrea Madotto, and Pascale Fung. 2020. Crossner: Evaluating cross-domain named entity recognition. In AAAI Conference on Artificial Intelligence.
  35. 35.S. Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: Designing data and methods for effective instruction tuning. In International Conference on Machine Learning.
  36. 36.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. In International Conference on Learning Representations.
  37. 37.Jie Lou, Yaojie Lu, Dai Dai, Wei Jia, Hongyu Lin, Xianpei Han, Le Sun, and Hua Wu. 2023. Universal information extraction as unified semantic matching. In AAAI Conference on Artificial Intelligence.
  38. 38.Yaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, and Hua Wu. 2022. Unified structure generation for universal information extraction. In Annual Meeting of the Association for Computational Linguistics.
  39. 39.Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar, and Oleg Rokhlenko. 2022. Multiconer: A large-scale multilingual dataset for complex named entity recognition. In International Conference on Computational Linguistics.
  40. 40.Andrei Mikheev, Marc Moens, and Claire Grover. 1999. Named entity recognition without gazetteers. In Conference of the European Chapter of the Association for Computational Linguistics.
  41. 41.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? ArXiv, abs/2202.12837.
  42. 42.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021. Cross-task generalization via natural language crowdsourcing instructions. In Annual Meeting of the Association for Computational Linguistics.
  43. 43.David Nadeau, Peter D. Turney, and Stan Matwin. 2006. Unsupervised named-entity recognition: Generating gazetteers and resolving ambiguity. In Canadian Conference on AI.
  44. 44.OpenAI. 2023. Gpt-4 technical report.
  45. 45.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  46. 46.Sampo Pyysalo and Sophia Ananiadou. 2014. Anatomical entity mention recognition at literature scale. Bioinformatics, 30(6):868–875.
  47. 47.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is chatgpt a general-purpose natural language processing task solver? ArXiv, abs/2302.06476.
  48. 48.Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  49. 49.Lev-Arie Ratinov and Dan Roth. 2009. Design challenges and misconceptions in named entity recognition. In Conference on Computational Natural Language Learning.
  50. 50.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, I. Evtimov, Joanna Bitton, Manish P Bhatt, Cristian Cantón Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre D’efossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code. ArXiv, abs/2308.12950.
  51. 51.Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, and Eneko Agirre. 2023. Gollie: Annotation guidelines improve zero-shot information-extraction. ArXiv, abs/2310.03668.
  52. 52.Sunita Sarawagi and William W. Cohen. 2004. Semi-markov conditional random fields for information extraction. In Neural Information Processing Systems.
  53. 53.Larry Smith, Lorraine K Tanabe, Cheng-Ju Kuo, I Chung, Chun-Nan Hsu, Yu-Shi Lin, Roman Klinger, Christoph M Friedrich, Kuzman Ganchev, Manabu Torii, et al. 2008. Overview of biocreative ii gene mention recognition. Genome biology, 9(2):1–19.
  54. 54.Simone Tedeschi and Roberto Navigli. 2022. MultiNERD: A multilingual, multi-genre and fine-grained dataset for named entity recognition (and disambiguation). In Findings of the Association for Computational Linguistics: NAACL 2022, pages 801–812, Seattle, United States. Association for Computational Linguistics.
  55. 55.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147.
  56. 56.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971.
  57. 57.Asahi Ushio, Leonardo Neves, Vitor Silva, Francesco. Barbieri, and Jose Camacho-Collados. 2022. Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal Shifts. In The 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, Online. Association for Computational Linguistics.
  58. 58.Christopher Walker, Stephanie Strassel, Julie Medero, and Kazuaki Maeda. 2006. Ace 2005 multilingual training corpus. Linguistic Data Consortium, Philadelphia, 57:45.
  59. 59.Xiao Wang, Wei Zhou, Can Zu, Han Xia, Tianze Chen, Yuan Zhang, Rui Zheng, Junjie Ye, Qi Zhang, Tao Gui, Jihua Kang, J. Yang, Siyuan Li, and Chunsai Du. 2023. Instructuie: Multi-task instruction tuning for unified information extraction. ArXiv, abs/2304.08085.
  60. 60.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, M. Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddharth Deepak Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. Supernaturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. In Conference on Empirical Methods in Natural Language Processing.
  61. 61.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  62. 62.Ralph Weischedel, Sean Boisen, Daniel Bikel, Robert Bobrow, Michael Crystal, William Ferguson, Allan Wechsler, and The PLUM Research Group. 1996. Progress in information extraction. In TIPSTER TEXT PROGRAM PHASE II: Proceedings of a Workshop held at Vienna, Virginia, May 6-8, 1996, pages 127–138, Vienna, Virginia, USA. Association for Computational Linguistics.
  63. 63.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23:170.
  64. 64.Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer. 2020. Zero-shot entity linking with dense entity retrieval. In EMNLP.
  65. 65.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021. A unified generative framework for various ner subtasks.
  66. 66.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Neural Information Processing Systems.
  67. 67.Hongbin Ye, Ningyu Zhang, Hui Chen, and Huajun Chen. 2022. Generative knowledge graph construction: A review. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1–17, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  68. 68.Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, Jie Zhou, Siming Chen, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. A comprehensive capability analysis of gpt-3 and gpt-3.5 series models. ArXiv, abs/2303.10420.
  69. 69.Norshuhani Zamin and Alan Oxley. 2011. Building a corpus-derived gazetteer for named entity recognition. In International Conference on Software Engineering and Computer Systems.
  70. 70.Urchade Zaratiana, Niama Elkhbir, Pierre Holat, Nadi Tomeh, and Thierry Charnois. 2022a. Global span selection for named entity recognition. In Proceedings of the Workshop on Unimodal and Multimodal Induction of Linguistic Structures (UM-IoS), pages 11–17, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  71. 71.Urchade Zaratiana, Nadi Tomeh, Niama El Khbir, Pierre Holat, and Thierry Charnois. 2023. Filtered semi-Markov CRF. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 222–235, Singapore. Association for Computational Linguistics.
  72. 72.Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois. 2022b. GNNer: Reducing overlapping in span-based NER using graph neural networks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, pages 97–103, Dublin, Ireland. Association for Computational Linguistics.
  73. 73.Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois. 2022c. Named entity recognition as structured span prediction. In Proceedings of the Workshop on Unimodal and Multimodal Induction of Linguistic Structures (UM-IoS), pages 1–10, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.
  74. 74.Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois. 2024. An autoregressive text-to-graph framework for joint entity and relation extraction. Proceedings of the AAAI Conference on Artificial Intelligence, 38(17):19477–19487.
  75. 75.Wenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen, and Hoifung Poon. 2023. Universalner: Targeted distillation from large language models for open named entity recognition.

Citation

MLA
Zaratiana, U., et al. “GLiNER: Generalist Model for Named Entity Recognition Using Bidirectional Transformer”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5364–76, https://doi.org/10.18653/v1/2024.naacl-long.300.
APA
Zaratiana, U., Tomeh, N., Holat, P., & Charnois, T. (2024). GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5364–5376. https://doi.org/10.18653/v1/2024.naacl-long.300
Chicago
Zaratiana, U., N. Tomeh, P. Holat, and T. Charnois. 2024. “GLiNER: Generalist Model for Named Entity Recognition Using Bidirectional Transformer”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5364–76. https://doi.org/10.18653/v1/2024.naacl-long.300.
Harvard
Zaratiana, U. et al. (2024) “GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5364–5376. Available at: https://doi.org/10.18653/v1/2024.naacl-long.300.
Vancouver
1. Zaratiana U, Tomeh N, Holat P, Charnois T (2024) GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5364–5376

BibTeX

@inproceedings{zaratiana-etal-2024-gliner,
    title = "{GL}i{NER}: Generalist Model for Named Entity Recognition using Bidirectional Transformer",
    author = "Zaratiana, Urchade  and
      Tomeh, Nadi  and
      Holat, Pierre  and
      Charnois, Thierry",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.300/",
    doi = "10.18653/v1/2024.naacl-long.300",
    pages = "5364--5376"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/