GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer
Urchade ZaratianaNadi TomehPierre HolatThierry Charnois
Proposes GLiNER, a compact bidirectional transformer model that performs parallel open-domain entity extraction by matching span representations to arbitrary type embeddings, outperforming much larger autoregressive models like ChatGPT on zero-shot benchmarks at a fraction of the computational cost.
Extracting key information such as names, dates, and organizations from unstructured text—known as Named Entity Recognition (NER)—is essential for modern data processing and knowledge graph construction. Traditional models are constrained to a rigid, predefined set of entity types and require costly retraining to expand. While recent Large Language Models (LLMs) can extract arbitrary, user-defined entity types on demand, their multi-billion parameter sizes cause slow processing speeds, high infrastructure costs, and expensive application programming interface (API) fees, making them impractical for resource-constrained environments.
The article introduces and evaluates GLiNER, a compact and generalist model designed to extract any entity type without requiring fixed categories or large-scale generative hardware. Rather than treating entity extraction as a sequential word-generation task, the model matches entity type prompts directly against text spans within a shared representation space.
To establish credibility and broad applicability, the approach was evaluated across multiple standardized benchmarks covering out-of-domain evaluation, 20 distinct domain datasets (including biomedical literature, news, and social media), and multilingual benchmarks spanning 11 languages. GLiNER models ranging from 50 million to 300 million parameters were trained on a diverse dataset of approximately 45,000 passages containing 13,000 unique entity types, and their performance was evaluated primarily under zero-shot conditions without task-specific fine-tuning.
The findings show that GLiNER delivers superior accuracy while requiring a fraction of the computational footprint. In zero-shot out-of-domain benchmarks, the large variant of GLiNER (300M parameters) achieved an average F1-score of 60.9, outperforming ChatGPT (47.5) and larger specialized LLMs like the 11-billion parameter InstructUIE (47.2) and the 13-billion parameter UniNER (55.6). Even the smallest 50M parameter version surpassed ChatGPT with a score of 52.7. Across 20 varied English datasets, GLiNER maintained an overall lead, achieving top performance in 13 benchmarks. Additionally, in multilingual zero-shot testing across 11 languages, a multilingual version of GLiNER surpassed ChatGPT's performance in eight languages despite being trained solely on English data.
These results demonstrate that organizations do not need to deploy massive, costly generative LLMs for open-vocabulary entity extraction. Adopting smaller bidirectional models reduces hardware requirements—enabling efficient execution on standard central processing units (CPUs)—lowers operational cloud expenses, and significantly increases text processing throughput by processing candidate spans in parallel rather than token-by-token.
Technical leaders and practitioners seeking flexible entity extraction should consider adopting compact span-matching models as a cost-effective alternative to LLM APIs or multi-billion parameter models. When implementing this architecture, teams should include negative entity sampling during training, as tests confirmed that balancing present and absent entity types improves extraction accuracy and reduces false positives.
Confidence in these results is high across diverse standard and out-of-domain benchmarks; however, two operational limitations warrant caution. First, the architecture cannot extract discontinuous entities, where an entity is split across non-adjacent words in a sentence. Second, the model showed reduced accuracy on informal, noisy social media text and underperformed on non-Latin scripts when using English-only backbones, indicating that targeted pretraining or multilingual encoders are necessary for those specific use cases.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). Introduces deep bidirectional Transformer representations, which serve as the foundational encoder architecture that GLiNER builds upon for span-level matching.
- Paper: A Survey on Deep Learning for Named Entity Recognition, Jing Li et al. (2018). Provides a comprehensive taxonomy of deep learning approaches, context encoders, and evaluation benchmarks in named entity recognition that GLiNER aims to generalize.
- Paper: Neural Architectures for Named Entity Recognition, Guillaume Lample et al. (2016). Establishes classic neural sequence-labeling baselines for named entity recognition across multiple languages, highlighting the traditional fixed-vocabulary constraints GLiNER overcomes.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Introduces unified text-to-text generative Transformer modeling, the primary generative LLM paradigm that GLiNER directly benchmarks against and offers a compact bidirectional alternative to.
- Paper: Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition, Erik F. Tjong Kim Sang et al. (2003). Defines the standard shared task benchmarks and cross-lingual evaluation protocols foundational to named entity recognition research.
No sufficiently relevant recommendations were found.
