MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

David Ifeoluwa AdelaniGraham NeubigSebastian RuderShruti RijhwaniMichael BeukmanChester Palen-MichelConstantine LignosJesujoba O. AlabiShamsuddeen Hassan MuhammadPeter Nabende

article2022EMNLP76 citations

Presents the largest human-annotated named entity recognition benchmark across 20 African languages, demonstrating that selecting linguistically related African source languages over English yields an average gain of 14 F1 points in zero-shot cross-lingual transfer.

Listen

African languages are spoken by over one billion people worldwide, yet they remain severely underrepresented in language technology research and development. Progress in automated tools has been hindered by a critical shortage of high-quality human-annotated datasets and an incomplete understanding of how cross-lingual transfer learning behaves across diverse linguistic contexts. The article addresses these barriers by introducing MasakhaNER 2.0, the largest human-annotated Named Entity Recognition benchmark for African languages, and by evaluating optimal transfer learning strategies across typologically varied settings.

The research evaluated 20 Sub-Saharan African languages spanning four distinct language families and 27 countries, covering more than 500 million speakers. Across these languages, native speakers annotated between 4.8 thousand and 11 thousand news sentences per language for key entity categories, including persons, locations, organizations, and dates. Using rigorous quality controls and high inter-annotator agreement benchmarks, the investigators benchmarked multiple multilingual pretrained language models and analyzed cross-lingual transfer across a total pool of 42 languages using both linguistic distance features and data-dependent metrics.

The findings show that model adaptation to local linguistic characteristics significantly boosts performance. AfroXLM-R-large emerged as the top-performing baseline model across all 20 evaluated languages, achieving an average F1 score of 87.0. In zero-shot transfer scenarios, selecting an optimal transfer language yielded an average F1 score of 71.3, outperforming English transfer—which averaged 56.9—by approximately 14 F1 points. Co-training models on the top two selected transfer languages further raised zero-shot performance to 73.1 F1. Detailed feature analysis revealed that geographic proximity and entity overlap are the strongest predictors of successful cross-lingual transfer, while training dataset size showed far less influence.

These results demonstrate that standard industry practices, which routinely rely on English as a universal transfer base, lead to suboptimal model performance when applied to African languages. In addition, sample efficiency experiments showed that annotating just 500 sentences in a target language—requiring approximately 2.5 hours of human effort—yields performance above 75 F1, which closely rivals zero-shot performance from an optimal transfer source. Combining targeted source selection with small amounts of localized target annotations significantly improves practical application viability while keeping data collection costs low.

Decision-makers and engineering teams should avoid default transfers from English and instead use ranked transfer frameworks that prioritize geographically close and linguistically aligned source languages. For new low-resource deployments, teams should allocate modest resources to annotate a small localized seed dataset of around 500 sentences on top of an optimal transfer model. While the findings provide high confidence for news-domain text processing across the studied languages, caution is warranted when generalizing to non-news genres or underrepresented language families not covered in the benchmark, such as the Khoisan or Austronesian groups.

Adelani et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

Abstract

African languages are spoken by over a billion people, but are underrepresented in NLP research and development. The challenges impeding progress include the limited availability of annotated datasets, as well as a lack of understanding of the settings where current methods are effective. In this paper, we make progress towards solutions for these challenges, focusing on the task of named entity recognition (NER). We create the largest human-annotated NER dataset for 20 African languages, and we study the behavior of state-of-the-art cross-lingual transfer methods in an Africa-centric setting, demonstrating that the choice of source language significantly affects performance. We show that choosing the best transfer language improves zero-shot F1 scores by an average of 14 points across 20 languages compared to using English. Our results highlight the need for benchmark datasets and models that cover typologically-diverse African languages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Languages and Their Characteristics
  • 3.1 Focus Languages
  • 3.2 Language Characteristics
  • 4 MasakhaNER 2.0 Corpus
  • 4.1 Data source and collection
  • 4.2 NER Annotation Methodology
  • 4.3 Quality Control
  • 5 Baseline Experiments
  • 5.1 Baseline Models
  • 5.2 Baseline Results
  • 5.3.2 Dataset Geography of Entities
  • 5.4 Transfer Between African NER Datasets
  • 6 Cross-Lingual Transfer
  • 6.1 Choosing Transfer Languages for NER
  • 6.2 Single-source Transfer Results
  • 6.3 LangRank and Co-training Results
  • 6.4 Sample Efficiency Results
  • 7 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Data Source and Splits
  • B Language Characteristics
  • B.1 Morphology and Noun classes
  • B.2 IsiXhosa and isiZulu morphological structure
  • B.2.1 Prefix
  • B.2.2 Capitalization
  • C Other NER Corpus
  • D Error Analysis of NER
  • E LangRank Feature Descriptions
  • F Overlap Results
  • G Zero-shot Transfer
  • H Best Transfer Language for Other Languages
  • I Sample Efficiency Results
  • J Model Hyper-parameters for Reproducibility

Knowls

  1. Knowl 1 — MasakhaNER 2.0 covers 20 African languages with human-annotated news NER

    data/table

    MasakhaNER 2.0 is a human-annotated named entity recognition corpus of news text in 20 Sub-Saharan African languages. It covers four language families: 17 Niger-Congo languages, Hausa (Afro-Asiatic), Luo (Nilo-Saharan), and Naija (English Creole). The languages span Western, Eastern, Central, and Southern Africa and are spoken by more than 500 million people across about 27 countries. Nine languages use locally collected monolingual news text; the other 11 use the MAFAND-MT translation corpus. Each language is split into training, development, and test sets in a 70%/10%/20% ratio.

    Sentence counts below are ordered as train/dev/test: Bambara (bam) 4,462/638/1,274; Ghomálá’ (bbj) 3,384/483/966; Éwé (ewe) 3,505/501/1,001; Fon (fon) 4,343/621/1,240; Hausa (hau) 5,716/816/1,633; Igbo (ibo) 7,634/1,090/2,181; Kinyarwanda (kin) 7,825/1,118/2,235; Luganda (lug) 4,942/706/1,412; Luo (luo) 5,161/737/1,474; Mossi (mos) 4,532/648/1,294; Naija (pcm) 5,646/806/1,613; Chichewa (nya) 6,250/893/1,785; chiShona (sna) 6,207/887/1,773; Kiswahili (swa) 6,593/942/1,883; Setswana (tsn) 3,489/499/996; Akan/Twi (twi) 4,240/605/1,211; Wolof (wol) 4,593/656/1,312; isiXhosa (xho) 5,718/817/1,633; Yorùbá (yor) 6,877/983/1,964; and isiZulu (zul) 5,848/836/1,670. Reported corpus token totals range from 69,474 for bbj to 344,095 for ibo.

  2. Knowl 2 — Four entity labels and a supervised annotation and adjudication workflow

    model/method

    MasakhaNER 2.0 marks four entity types: personal names (PER), locations (LOC), organizations (ORG), and dates or times (DATE), following the MUC-6 annotation guide. The project used the ELISA annotation tool and recruited three native-speaker annotators and a language coordinator for each language. Coordinators were trained in online workshops, practiced on 100 English sentences, and then trained their language teams using English and African-language examples. Coordinators and annotators resolved disagreements through ELISA adjudication.

    After coordinator intervention and before final adjudication, entity-level Fleiss’ kappa ranged from 0.907 to 1.000 across languages. For example, pcm increased from a pre-intervention kappa of 0.648 to 0.966 after the coordinator reviewed disagreements. Automatic quality checks flagged potentially missed entity tokens, tokens with near-zero entropy across entity types, and likely sentence-boundary errors. Follow-up review of flagged items was voluntary; only 10 of the 20 language teams fixed all quality-control-flagged tokens.

  3. Knowl 3 — Selecting a suitable source language improves zero-shot African NER transfer

    empirical result

    In zero-shot named entity recognition on the 20 MasakhaNER 2.0 languages, the choice of source language materially affected performance. Across source-language evaluations, transfer from African languages averaged 57.3 F1, compared with 51.7 F1 from non-African languages; Bambara was the weakest source on average at 41.0 F1, while chiShona was strongest at 64 F1, followed by Yorùbá at 63 F1. The paper reports that selecting the best transfer language improves zero-shot F1 by an average of 14 points over using English as the source. Geographically and syntactically close languages often transferred well to one another, including the Southern Bantu languages chiShona, isiXhosa, and isiZulu, as well as Nigerian and East African language groups. Exceptions included Kiswahili, for which German or Arabic could be preferable.

  4. Knowl 4 — AfroXLM-R-large gives the strongest in-language NER baseline

    empirical result

    The authors fine-tuned eight multilingual or Africa-focused pretrained language models on each MasakhaNER 2.0 language’s training set and evaluated on its test set using micro-averaged F1. Across five runs, mean F1 (with reported run variation) was: mBERT 82.8 ± 0.2; XLM-R-base 84.1 ± 0.1; XLM-R-large 85.1 ± 0.5; RemBERT 85.0 ± 0.2; mDeBERTaV3 85.7 ± 0.2; AfriBERTa 83.0 ± 0.2; AfroXLM-R-base 85.7 ± 0.1; and AfroXLM-R-large 87.0 ± 0.2. AfroXLM-R-large was the strongest overall, exceeding mDeBERTaV3 by 1.3 F1, RemBERT by 2.0, and AfriBERTa by 4.0. The authors also observed that mBERT and AfriBERTa could trail XLM-R-base by 6–12 F1 on languages outside their pretraining coverage, while larger models generally performed better. Fine-tuning used a maximum sequence length of 200, batch size 16, gradient accumulation 2, learning rate 5e-5, and 50 epochs.

  5. Knowl 5 — Co-training two transfer languages improves on single-source transfer

    empirical result

    For zero-shot NER on MasakhaNER 2.0, jointly training on the combined training data of the two best source languages improved performance by about 3 F1 on average compared with using the best single source. Gains were especially notable for Fon, Igbo, Kinyarwanda, and Akan/Twi, where co-training improved results by 3–7 F1. Co-training the two languages selected by LangRank also outperformed LangRank’s predicted second-best language, but often did not match the best source language found by exhaustive transfer evaluation. LangRank placed at least one of the two best source languages among its top two predictions for 13 of the 20 target languages.

  6. Knowl 6 — Training on all 20 languages closes the cross-dataset transfer gap

    empirical result

    A multilingual AfroXLM-R-large experiment compared training on the original MasakhaNER 1.0 data, training on MasakhaNER 2.0 data for the overlapping MasakhaNER 1.0 languages, and training on MasakhaNER 2.0 data for all 20 languages. On the MasakhaNER 2.0 test set, the three settings averaged 69.7, 72.7, and 87.0 F1, respectively. Thus, using the newer corpus only for the overlapping languages raised the average by 3.0 F1 over the original data, while including all 20 languages raised it by a further 14.3 F1. On the MasakhaNER 1.0 test set, the same settings averaged 84.8, 80.0, and 80.1 F1. The results show that broad multilingual training on the diverse MasakhaNER 2.0 languages particularly benefits evaluation on its target languages, whereas the original data gives the strongest average on its own test set.

  7. Knowl 7 — Transfer-language ranking relies most on geography and entity overlap

    empirical result

    The authors evaluated LangRank-style source-language ranking using linguistic-distance and data-dependent features. Geographic distance between languages and entity overlap were the most influential features for selecting transfer sources. Entity overlap—the shared entity vocabulary in the source and target training data—had a positive Spearman correlation of 0.6 with transfer F1 (reported p < 0.05); in the ranking model’s feature analysis, geographic distance appeared among the top three features for 15 best-source predictions and 16 second-best predictions, while entity overlap appeared 11–13 times for the top two choices. Dataset size was not among the most important features. The authors relate the usefulness of geographic proximity to the recurrence of names of people and places within a country or region, and argue that effective training data should be typologically diverse rather than selected by size alone.

  8. Knowl 8 — Zero-shot transfer was benchmarked across 42 NER languages

    experimental setup

    To study source-language choice, the authors assembled 22 existing human-annotated NER datasets and combined them with MasakhaNER 2.0, yielding 42 languages, including 21 African languages. Each external dataset contained at least PER, ORG, and LOC labels; the transfer comparison was limited to those three entity types. They fine-tuned mDeBERTaV3 on one source language at a time and evaluated zero-shot on the test sets of the 20 MasakhaNER 2.0 target languages, producing source-to-target transfer scores. The source-language candidates were compared using transfer scores and features including geographic, genetic, inventory, syntactic, phonological, and featural distances, source and target dataset sizes, their size ratio, and entity overlap. The ranking analysis used a leave-one-target-language-out setup.

  9. Knowl 9 — Best-source transfer approaches the benefit of hundreds of target examples

    empirical result

    The sample-efficiency experiment compared NER models fine-tuned directly from a pretrained language model on 100 or 500 target-language sentences with models initialized from the best source-language NER model and then fine-tuned on the same number of target sentences. The paper reports less than 50 F1 when training on 100 target sentences and more than 75 F1 with 500. Zero-shot use of the best transfer language produced performance close to using 500 target-language sentences in most cases; fine-tuning that transferred model on a further 100 or 500 target sentences improved performance again. The authors estimate that annotating 100 sentences takes about 30 minutes and annotating 500 takes about 2 hours and 30 minutes.

  10. Knowl 10 — Coverage and domain restrict generalization claims

    limitation

    MasakhaNER 2.0 does not cover some African language families and regions, including Khoisan and Austronesian languages such as Malagasy and languages spoken in parts of Central Africa such as South Sudan, Chad, and the Democratic Republic of the Congo. Because the annotated texts are news-domain data, models trained on the corpus may not generalize well to casual text with different vocabulary, entities, or orthographic variation. The transfer experiments concern NER alone, so the paper cautions that the findings about useful transfer approaches and language features may not carry over to tasks such as machine translation or part-of-speech tagging. The authors also note that their ranking model does not identify the best transfer language perfectly, leaving other potentially relevant factors, including sociolinguistic connections and language contact, unexplained.

Coverage note — The dataset-geography analysis linking entities to Wikidata and the detailed entity-level error analysis by entity length, frequency, and label were omitted because they are secondary analyses relative to the corpus construction, baseline, and transfer-learning findings.

References

  1. 1.Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Muhammad, Ibrahim Sa’id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, and Bello Shehu Bello. 2022. Hausa visual genome: A dataset for multi-modal English to Hausa machine translation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6471–6479, Marseille, France. European Language Resources Association.
  2. 2.David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, and Cristina España-Bonet. 2021a. The effect of domain and diacritics in Yoruba–English neural machine translation. In Proceedings of Machine Translation Summit XVIII: Research Track, pages 61–75, Virtual. Association for Machine Translation in the Americas.
  3. 3.David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021b. MasakhaNER: Named entity recognition for African languages. Transactions of the Association for Computational Linguistics, 9:1116–1131.
  4. 4.David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Chinenye Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles HACHEME, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ayoade Ajibade, Oluwaseyi Ajayi Ajayi, Yvonne Wambui Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Koffi KALIPE, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, and Sam Manthalu. 2022. A few thousand translations go a long way! leveraging pre-trained models for african news translation. In NAACL-HLT.
  5. 5.Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022. Multi task learning for zero shot performance prediction of multilingual models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5454–5467, Dublin, Ireland. Association for Computational Linguistics.
  6. 6.Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, and Cristina España-Bonet. 2020. Massive vs. curated embeddings for low-resourced languages: the case of Yorùbá and Twi. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2754–2762, Marseille, France. European Language Resources Association.
  7. 7.Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022. Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. In Proceedings of the 29th International Conference on Computational Linguistics, pages 4336–4349, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  8. 8.Cheikh Anta Babou and Michele Loporcaro. 2016. Noun classes and grammatical gender in wolof. Journal of African Languages and Linguistics, 37(1):1–57.
  9. 9.Yassine Benajiba, Paolo Rosso, and José Miguel BenedíRuiz. 2007. Anersys: An arabic named entity recognition system based on maximum entropy. In Computational Linguistics and Intelligent Text Processing, pages 143–153, Berlin, Heidelberg. Springer Berlin Heidelberg.
  10. 10.Michael Beukman. 2022. Analysing the effects of transfer learning on low-resourced named entity recognition performance. In 3rd Workshop on African Natural Language Processing.
  11. 11.Adams Bodomo and Charles Marfo. 2002. The morphophonology of noun classes in dagaare and akan.
  12. 12.Hyung Won Chung, Thibault Fevry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2021. Rethinking embedding coupling in pre-trained language models. In International Conference on Learning Representations.
  13. 13.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR.
  14. 14.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  15. 15.Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel, and Fabio Petroni. 2022. Multilingual autoregressive entity linking. Transactions of the Association for Computational Linguistics, 10:274–290.
  16. 16.Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2022. Make the best of cross-lingual transfer: Evidence from POS tagging with over 100 languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7676–7685, Dublin, Ireland. Association for Computational Linguistics.
  17. 17.A J De Waal, A L Louis, and J P Venter. 2006. Named entity recognition in a south african context.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019a. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019b. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Matthew S. Dryer and Martin Haspelmath, editors. 2013. WALS Online. Max Planck Institute for Evolutionary Anthropology, Leipzig.
  21. 21.Stefan Daniel Dumitrescu and Andrei-Marius Avram. 2020. Introducing RONEC - the Romanian named entity corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4436–4443, Marseille, France. European Language Resources Association.
  22. 22.Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. 2022. AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6279–6299, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Roald Eiselen. 2016. Government domain named entity recognition for South African languages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3344–3348, Portorož, Slovenia. European Language Resources Association (ELRA).
  24. 24.Fahim Faisal, Yinkai Wang, and Antonios Anastasopoulos. 2022. Dataset geography: Mapping language data to language users. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3381–3411, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.∀, Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020. Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online.
  26. 26.Cláudia Freitas, Cristina Mota, Diana Santos, Hugo Gonçalo Oliveira, and Paula Carvalho. 2010. Second HAREM: Advancing the state of the art of named entity recognition in Portuguese. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta. European Language Resources Association (ELRA).
  27. 27.Normunds Gruzitis, Lauma Pretkalnina, Baiba Saulite, Laura Rituma, Gunta Nespore-Berzkalne, Arturs Znotins, and Peteris Paikens. 2018. Creation of a balanced state-of-the-art multilayer corpus for NLU. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  28. 28.Harald Hammarström, Robert Forkel, and Martin Haspelmath. 2018. Glottolog 3.0. Max Planck Institute for the Science of Human History.
  29. 29.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. ArXiv, abs/2111.09543.
  30. 30.Michael A. Hedderich, David Adelani, Dawei Zhu, Jesujoba Alabi, Udia Markus, and Dietrich Klakow. 2020. Transfer learning and distant supervision for multilingual transformer models: A study on African languages. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2580–2591, Online. Association for Computational Linguistics.
  31. 31.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 4411–4421. PMLR.
  32. 32.Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, and Anders Søgaard. 2020. DaNE: A named entity resource for Danish. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4597–4604, Marseille, France. European Language Resources Association.
  33. 33.Ebrahim Chekol Jibril and A. Cüneyd Tantug. 2022. Anec: An amharic named entity corpus and transformer based recognizer. ArXiv, abs/2207.00785.
  34. 34.Bjarte Johansen. 2019. Named-entity recognition for norwegian. In Proceedings of the 22nd Nordic Conference on Computational Linguistics, NoDaLiDa.
  35. 35.C. Junior, H. Macedo, T. Bispo, F. Oliveira, N. Silva, and L. Barbosa. 2015. Paramopama: a brazilian-portuguese corpus for named entity recognition. In 12th National Meeting on Artificial and Computational Intelligence (ENIAC).
  36. 36.Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020. Cross-lingual ability of multilingual bert: An empirical study. In International Conference on Learning Representations.
  37. 37.Antonia Karamolegkou and Sara Stymne. 2021. Investigation of transfer languages for parsing Latin: Italic branch vs. Hellenic branch. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 315–320, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
  38. 38.Siti Oryza Khairunnisa, Aizhan Imankulova, and Mamoru Komachi. 2020. Towards a standardized dataset on Indonesian named entity recognition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop, pages 64–71, Suzhou, China. Association for Computational Linguistics.
  39. 39.Maria Yu Konoshenko and Dasha Shavarina. 2019. A microtypological survey of noun classes in kwa. Journal of African Languages and Linguistics, 40:114 – 75.
  40. 40.Anne Lauscher, Vinit Ravishankar, Ivan Vulic, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online. Association for Computational Linguistics.
  41. 41.M Paul Lewis. 2009. Ethnologue: Languages of the world Sixteenth Edition. SIL international.
  42. 42.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  43. 43.Ying Lin, Cash Costello, Boliang Zhang, Di Lu, Heng Ji, James Mayfield, and Paul McNamee. 2018. Platforms for non-speakers annotating names in any language. In Proceedings of ACL 2018, System Demonstrations, pages 1–6, Melbourne, Australia. Association for Computational Linguistics.
  44. 44.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019. Choosing transfer languages for cross-lingual learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3125–3135, Florence, Italy. Association for Computational Linguistics.
  45. 45.Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017. URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 8–14, Valencia, Spain. Association for Computational Linguistics.
  46. 46.Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaichen Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, and Graham Neubig. 2021. ExplainaBoard: An explainable leaderboard for NLP. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 280–289, Online. Association for Computational Linguistics.
  47. 47.Bernardo Magnini, Amedeo Cappelli, Fabio Tamburini, Cristina Bosco, Alessandro Mazzei, Vincenzo Lombardo, Francesca Bertagna, Nicoletta Calzolari, Antonio Toral, Valentina Bartalesi Lenzi, Rachele Sprugnoli, and Manuela Speranza. 2008. Evaluation of natural language tools for Italian: EVALITA 2007. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Marrakech, Morocco. European Language Resources Association (ELRA).
  48. 48.Stephen Mayhew, Tatiana Tsygankova, and Dan Roth. 2019. ner and pos when nothing is capitalized. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6256–6261, Hong Kong, China. Association for Computational Linguistics.
  49. 49.Hans J. Melzian. 1933. Introduction to the phonology of the bantu languages. Bulletin of the School of Oriental and African Studies, 7(1):246–247.
  50. 50.Steven Moran, D McCloy, and R Wright. 2014. Phoible online. max planck institute for evolutionary anthropology, leipzig.
  51. 51.Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa’id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Alípio Jorge, and Pavel Brazdil. 2022. NaijaSenti: A nigerian Twitter sentiment corpus for multilingual sentiment analysis. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 590–602, Marseille, France. European Language Resources Association.
  52. 52.Clemens Neudecker. 2016. An open corpus for named entity recognition in historic newspapers. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4348–4352, Portorož, Slovenia. European Language Resources Association (ELRA).
  53. 53.Johanna Nichols and Balthasar Bickel. 2013. Possessive classification. In Matthew S. Dryer and Martin Haspelmath, editors, The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology, Leipzig.
  54. 54.Derek Nurse and Gerard Philippson, editors. 2006. The Bantu Languages. Routledge Language Family Series. Routledge, London, England.
  55. 55.Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. 2020. CAMeL tools: An open source python toolkit for Arabic natural language processing. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 7022–7032, Marseille, France. European Language Resources Association.
  56. 56.Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 116–126, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  57. 57.Akintunde Oladipo, Odunayo Ogundepo, Kelechi Ogueji, and Jimmy Lin. 2022. An exploration of vocabulary size and transfer effects in multilingual language models for african languages. In 3rd Workshop on African Natural Language Processing.
  58. 58.JC Oosthuysen. 2016. The Grammar of isiXhosa, 1 edition. African Sun Media.
  59. 59.Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Jiyoon Han, Jangwon Park, Chisung Song, Junseong Kim, Yongsook Song, Taehwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seungwon Do, Sunkyoung Kim, Kyungtae Lim, Jongwon Lee, Kyumin Park, Jamin Shin, Seonghyun Kim, Lucy Park, Alice Oh, Jungwoo Ha, and Kyunghyun Cho. 2021. Klue: Korean language understanding evaluation.
  60. 60.Doris L. Payne, Sara Pacchiarotti, and Mokaya Bosire, editors. 2017. Diversity in African languages. Number 1 in Contemporary African Linguistics. Language Science Press, Berlin.
  61. 61.Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7654–7673, Online. Association for Computational Linguistics.
  62. 62.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  63. 63.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  64. 64.Hanieh Poostchi, Ehsan Zare Borzeshi, Mohammad Abdous, and Massimo Piccardi. 2016. PersoNER: Persian named-entity recognition. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 3381–3389, Osaka, Japan. The COLING 2016 Organizing Committee.
  65. 65.A R Priatama, , and Y Setiawan. 2022. Regression models for estimating aboveground biomass and stand volume using landsat-based indices in post-mining area. J. Manaj. Hutan Trop. (J. Trop. For. Manag.), 28(1):1–14.
  66. 66.Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020. Intermediate-task transfer learning with pretrained language models: When and why does it work? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5231–5247, Online. Association for Computational Linguistics.
  67. 67.Machel Reid, Junjie Hu, Graham Neubig, and Yutaka Matsuo. 2021. AfroMT: Pretraining strategies and reproducible benchmarks for translation of 8 African languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1306–1320, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  68. 68.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  69. 69.Teemu Ruokolainen, Pekka Kauppinen, Miikka Silfverberg, and Krister Lindén. 2019. A finnish news corpus for named entity recognition. Language Resources and Evaluation, pages 1–26.
  70. 70.O. M. Singh, A. Padia, and A. Joshi. 2019. Named entity recognition for nepali language. In 2019 IEEE 5th International Conference on Collaboration and Internet Computing (CIC), pages 184–190.
  71. 71.Stephanie Strassel and Jennifer Tracey. 2016. LORELEI language packs: Data, tools, and resources for technology development in low resource languages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3273–3280, Portorož, Slovenia. European Language Resources Association (ELRA).
  72. 72.György Szarvas, Richárd Farkas, László Felföldi, András Kocsor, and János Csirik. 2006. A highly accurate named entity corpus for Hungarian. In Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06), Genoa, Italy. European Language Resources Association (ELRA).
  73. 73.Erik F. Tjong Kim Sang. 2002. Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition. In COLING-02: The 6th Conference on Natural Language Learning 2002 (CoNLL-2002).
  74. 74.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147.
  75. 75.Kees Versteegh. 2001. Linguistic contacts between arabic and other languages. Arabica, 48:470–508.
  76. 76.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  77. 77.Shijie Wu and Mark Dredze. 2019. Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 833–844, Hong Kong, China. Association for Computational Linguistics.
  78. 78.Mengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang, and Graham Neubig. 2020. Predicting performance for natural language processing tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8625–8646, Online. Association for Computational Linguistics.
  79. 79.Seid Muhie Yimam, Hizkiel Mitiku Alemayehu, Abinew Ayele, and Chris Biemann. 2020. Exploring Amharic sentiment analysis from social media texts: Building annotation tools and classification models. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1048–1060, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  80. 80.Hailemariam Mehari Yohannes and Toshiyuki Amagasa. 2022. Named-entity recognition for a low-resource language using pre-trained language model. In Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing, SAC ’22, page 837–844, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Adelani, D. I., et al. “MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4488–508, https://doi.org/10.18653/v1/2022.emnlp-main.298.
APA
Adelani, D. I., Neubig, G., Ruder, S., Rijhwani, S., Beukman, M., Palen-Michel, C., Lignos, C., Alabi, J., Muhammad, S. H., Nabende, P., Dione, C. M. B., Bukula, A., Mabuya, R., Dossou, B. F. P., Sibanda, B. K., Buzaaba, H., Mukiibi, J., Kalipe, G., Mbaye, D., … Klakow, D. (2022). MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4488–4508. https://doi.org/10.18653/v1/2022.emnlp-main.298
Chicago
Adelani, D. I., G. Neubig, S. Ruder, et al. 2022. “MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4488–4508. https://doi.org/10.18653/v1/2022.emnlp-main.298.
Harvard
Adelani, D.I. et al. (2022) “MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4488–4508. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.298.
Vancouver
1. Adelani DI, Neubig G, Ruder S, et al (2022) MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4488–4508

BibTeX

@inproceedings{adelani-etal-2022-masakhaner,
    title = "{M}asakha{NER} 2.0: {A}frica-centric Transfer Learning for Named Entity Recognition",
    author = "Adelani, David Ifeoluwa  and
      Neubig, Graham  and
      Ruder, Sebastian  and
      Rijhwani, Shruti  and
      Beukman, Michael  and
      Palen-Michel, Chester  and
      Lignos, Constantine  and
      Alabi, Jesujoba O.  and
      Muhammad, Shamsuddeen H.  and
      Nabende, Peter  and
      Dione, Cheikh M. Bamba  and
      Bukula, Andiswa  and
      Mabuya, Rooweither  and
      Dossou, Bonaventure F. P.  and
      Sibanda, Blessing  and
      Buzaaba, Happy  and
      Mukiibi, Jonathan  and
      Kalipe, Godson  and
      Mbaye, Derguene  and
      Taylor, Amelia  and
      Kabore, Fatoumata  and
      Emezue, Chris Chinenye  and
      Aremu, Anuoluwapo  and
      Ogayo, Perez  and
      Gitau, Catherine  and
      Munkoh-Buabeng, Edwin  and
      Memdjokam Koagne, Victoire  and
      Tapo, Allahsera Auguste  and
      Macucwa, Tebogo  and
      Marivate, Vukosi  and
      Mboning, Elvis  and
      Gwadabe, Tajuddeen  and
      Adewumi, Tosin  and
      Ahia, Orevaoghene  and
      Nakatumba-Nabende, Joyce  and
      Mokono, Neo L.  and
      Ezeani, Ignatius  and
      Chukwuneke, Chiamaka  and
      Adeyemi, Mofetoluwa  and
      Hacheme, Gilles Q.  and
      Abdulmumin, Idris  and
      Ogundepo, Odunayo  and
      Yousuf, Oreen  and
      Moteu Ngoli, Tatiana  and
      Klakow, Dietrich",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.298/",
    doi = "10.18653/v1/2022.emnlp-main.298",
    pages = "4488--4508"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/