MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition
David Ifeoluwa AdelaniGraham NeubigSebastian RuderShruti RijhwaniMichael BeukmanChester Palen-MichelConstantine LignosJesujoba O. AlabiShamsuddeen Hassan MuhammadPeter Nabende
Presents the largest human-annotated named entity recognition benchmark across 20 African languages, demonstrating that selecting linguistically related African source languages over English yields an average gain of 14 F1 points in zero-shot cross-lingual transfer.
African languages are spoken by over one billion people worldwide, yet they remain severely underrepresented in language technology research and development. Progress in automated tools has been hindered by a critical shortage of high-quality human-annotated datasets and an incomplete understanding of how cross-lingual transfer learning behaves across diverse linguistic contexts. The article addresses these barriers by introducing MasakhaNER 2.0, the largest human-annotated Named Entity Recognition benchmark for African languages, and by evaluating optimal transfer learning strategies across typologically varied settings.
The research evaluated 20 Sub-Saharan African languages spanning four distinct language families and 27 countries, covering more than 500 million speakers. Across these languages, native speakers annotated between 4.8 thousand and 11 thousand news sentences per language for key entity categories, including persons, locations, organizations, and dates. Using rigorous quality controls and high inter-annotator agreement benchmarks, the investigators benchmarked multiple multilingual pretrained language models and analyzed cross-lingual transfer across a total pool of 42 languages using both linguistic distance features and data-dependent metrics.
The findings show that model adaptation to local linguistic characteristics significantly boosts performance. AfroXLM-R-large emerged as the top-performing baseline model across all 20 evaluated languages, achieving an average F1 score of 87.0. In zero-shot transfer scenarios, selecting an optimal transfer language yielded an average F1 score of 71.3, outperforming English transfer—which averaged 56.9—by approximately 14 F1 points. Co-training models on the top two selected transfer languages further raised zero-shot performance to 73.1 F1. Detailed feature analysis revealed that geographic proximity and entity overlap are the strongest predictors of successful cross-lingual transfer, while training dataset size showed far less influence.
These results demonstrate that standard industry practices, which routinely rely on English as a universal transfer base, lead to suboptimal model performance when applied to African languages. In addition, sample efficiency experiments showed that annotating just 500 sentences in a target language—requiring approximately 2.5 hours of human effort—yields performance above 75 F1, which closely rivals zero-shot performance from an optimal transfer source. Combining targeted source selection with small amounts of localized target annotations significantly improves practical application viability while keeping data collection costs low.
Decision-makers and engineering teams should avoid default transfers from English and instead use ranked transfer frameworks that prioritize geographically close and linguistically aligned source languages. For new low-resource deployments, teams should allocate modest resources to annotate a small localized seed dataset of around 500 sentences on top of an optimal transfer model. While the findings provide high confidence for news-domain text processing across the studied languages, caution is warranted when generalizing to non-news genres or underrepresented language families not covered in the benchmark, such as the Khoisan or Austronesian groups.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R is a key multilingual pretrained-model foundation for the source’s cross-lingual transfer experiments and African-language model adaptation.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This study establishes how multilingual BERT supports zero-shot NER transfer, clarifying the cross-lingual tagging setup that the source evaluates across African languages.
- Paper: Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition, Erik F. Tjong Kim Sang et al. (2003). The CoNLL-2003 shared task defines the news-domain NER categories and evaluation conventions that help situate the source’s benchmark and F1 results.
No sufficiently relevant recommendations were found.
