A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

David Ifeoluwa AdelaniJesujoba O. AlabiAngela FanJulia KreutzerXiaoyu ShenMachel ReidDana RuiterDietrich KlakowPeter NabendeErnie Chang

article2022NAACL156 citations

Demonstrates that fine-tuning large multilingual models on just a few thousand high-quality in-domain sentence pairs effectively transfers translation capabilities to sixteen low-resource African languages absent from the original pre-training data.

Listen

Modern automated translation systems rely heavily on massive web crawls, leaving out hundreds of widely spoken African languages that lack extensive online representation. Because critical news across the African continent is predominantly published in English, French, or Arabic, millions of native speakers face severe information bottlenecks during crises. The article evaluates how existing large-scale pre-trained translation models can be adapted to 16 under-resourced African languages and transferred into the news domain, where parallel training data has historically been scarce.

To achieve this, the authors created MAFAND-MT, a participatory news corpus spanning 16 African languages—including eight languages entirely new to machine translation evaluation benchmarks. Built by native speakers and professional translators, the dataset contributed between 2,000 and 8,000 parallel sentences per language paired with English or French. The researchers then adapted and evaluated several major multilingual models (including M2M-100, MT5, ByT5, and mBART50) across varied configurations: zero-shot transfer, continual pre-training on collected monolingual text, and fine-tuning on small parallel datasets across both religious and news domains.

The investigation produced four central findings. First, zero-shot translation completely failed for most unseen African languages, often yielding scores below 5 BLEU. Second, fine-tuning large pre-trained models on as few as 2,000 to 2,500 target-domain sentence pairs proved remarkably effective, consistently outperforming standard bilingual translation models trained from scratch on much larger datasets. Third, M2M-100 emerged as the strongest overall system, while the byte-based ByT5 model outperformed token-based MT5 by over 3 BLEU points. Fourth, relying solely on legacy out-of-domain data caused performance drops of up to 95.5%, whereas a staged adaptation strategy—training first on combined religious and news data followed by specialized news fine-tuning—delivered the highest translation quality.

These findings demonstrate that local communities and developers do not need massive web-crawled datasets or multimillion-dollar computational budgets to build functional translation tools. Targeted collections of a few thousand clean, professionally verified translations can effectively steer large pre-trained models toward new languages and domains while mitigating archaic domain biases. However, the researchers caution that overall translation scores for several complex or highly isolating languages remain modest, and automated metrics do not capture subtle grammatical and logical errors. Organizations adopting these techniques should prioritize rigorous human evaluation before deploying models into production environments and continue expanding small-scale, high-quality parallel data collection for additional under-resourced languages.

arXiv: 2205.02022masakhane-io/lafand-mt
  • Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Omnilingual MT extends low-resource translation adaptation to 1,600 languages, scaling the challenge of making pretrained translation models work beyond their original language coverage.
Cover for A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

Abstract

Recent advances in the pre-training of language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages are not well represented on the web and therefore excluded from the large-scale crawls used to create datasets. Furthermore, downstream users of these models are restricted to the selection of languages originally chosen for pre-training. This work investigates how to optimally leverage existing pre-trained models to create low-resource translation systems for 16 African languages. We focus on two questions: 1) How can pre-trained models be used for languages not included in the initial pre-training? and 2) How can the resulting translation models effectively transfer to new domains? To answer these questions, we create a new African news corpus covering 16 languages, of which eight languages are not part of any existing evaluation dataset. We demonstrate that the most effective strategy for transferring both to additional languages and to additional domains is to fine-tune large pre-trained models on small quantities of high-quality translation data.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Focus Languages and Their Data
  • 4 MAFAND-MT African News Corpus
  • 4.1 Data Collection Process
  • 4.2 Monolingual News Corpus
  • 5 Models and Methods
  • 5.1 Baseline Models
  • 5.2 Transfer Learning Across Languages
  • 5.3 Transfer Learning Across Domains
  • 6 Results and Discussion
  • 6.1 Adaptation to the Focus Languages
  • 6.2 Adaptation to the News Domain
  • 6.3 Analysis of Domain Shift
  • 7 Conclusion
  • 8 Acknowledgment
  • References
  • A Language Characteristics
  • B Available Parallel Corpora
  • C Monolingual Corpus PLMs adaptation
  • D Model Hyper-parameters and Reproducibility of Results
  • E BLEU vs spBLEU
  • F Qualitative Analysis
  • G Limitations and Risks

Knowls

  1. Knowl 1 — MAFAND-MT establishes news translation benchmarks for 16 African languages

    data/table

    MAFAND-MT covers 16 African languages, with English as the source or target for 10 languages and French for six. Eleven languages received newly curated local-news data; five languages had existing news parallel data that the authors incorporated. For the new data, native speakers selected and sentence-segmented articles from local news sites across varied topics, professional translators translated 5,000–8,000 sentences per language, and native speakers reviewed problematic translations and checked for spelling, duplication, and alignment errors. The project was participatory: contributors were native speakers, were remunerated, and worked with language coordinators.

    The NEWS split sizes below are TRAIN/DEV/TEST sentence counts; the final value gives the religious corpus source and its size: Bambara (bam), 3,302/1,484/1,600, Bible 28K; Ghomálá’ (bbj), 2,232/1,133/1,430, Bible 8K; Éwé (ewe), 2,026/1,414/1,563, JW300 618K; Fon (fon), 2,637/1,227/1,579, JW300 32K; Hausa (hau), 3,098/1,300/1,500, JW300 236K; Igbo (ibo), 6,998/1,500/1,500, JW300 415K; Luganda (lug), 4,075/1,500/1,500, Bible 31K; Luo (luo), 4,262/1,500/1,500, Bible 31K; Mossi (mos), 2,287/1,478/1,574, JW300 216K; Naija (pcm), 4,790/1,484/1,564, JW300 23K; Swahili (swa), 30,782/1,791/1,835, JW300 872K; Setswana (tsn), 2,100/1,340/1,500, JW300 870K; Akan/Twi (twi), 3,337/1,284/1,500, JW300 601K; Wolof (wol), 3,360/1,506/1,500, Bible 22K; Yorùbá (yor), 6,644/1,544/1,558, JW300 460K; isiZulu (zul), 3,500/1,239/998, JW300 667K. Bible alignments were made by verse; Ghomálá’ had only New Testament data. The benchmark therefore pairs small, professionally translated news sets with religious-domain parallel data whose size ranges from 8K to 872K sentences.

  2. Knowl 2 — The study compares multilingual pre-trained models and language/domain adaptation strategies

    experimental setup

    The experiments compare MT5-base and ByT5-base (580M parameters each), mBART50 (610M), and M2M-100 (418M) with bilingual Transformer models trained from scratch. NEWS-only fine-tuning generally used 2K–7K sentence pairs per language, except Swahili, which had 30,782 training pairs. The bilingual Transformer baseline used a joint 10K-subword SentencePiece vocabulary and was trained on concatenated religious (REL) and news (NEWS) data.

    For language adaptation, the authors continually pre-trained MT5, ByT5, and mBART50 on monolingual text using each model’s original objective and vocabulary. The monolingual collection totaled about 12.3 GB and covered 17 African languages plus Arabic, English, and French. MT5 and ByT5 adaptations used a learning rate of 1e-4, 10,000 warm-up steps, batch size 2,048, and one epoch; mBART50 adaptation used a learning rate of 5e-5 for 50,000 steps. The resulting models were called AfriMT5, AfriByT5, and AfriMBART.

    For fine-tuning pre-trained models, the reported settings were learning rate 5e-5, batch size 10, maximum source and target lengths of 200, beam size 10, and three epochs, except NEWS-only experiments, which used 10 epochs. The domain-transfer comparison used M2M-100 and three three-epoch schedules: REL+NEWS pooled together; REL followed by NEWS fine-tuning; and REL+NEWS followed by additional NEWS fine-tuning. Experiments ran on one Nvidia V100 GPU.

  3. Knowl 3 — A few thousand news pairs make M2M-100 the strongest tested translation model

    empirical result

    With NEWS-only fine-tuning, M2M-100 achieved the highest average BLEU among the tested models in both translation directions. Here, en/fr→xx means English or French into an African language, and xx→en/fr means an African language into English or French. Average BLEU across the focus languages was: MT5 7.2 and 10.8; AfriMT5 8.5 and 13.2; ByT5 10.9 and 14.1; AfriByT5 11.5 and 15.0; mBART50 13.2 and 11.2; AfriMBART 11.9 and 12.4; M2M-100 15.4 and 17.7; and English/French-centric M2M-100 12.5 and 15.8, respectively.

    These results show that fine-tuning a pre-trained model can produce useful systems even for languages absent from its original pre-training. ByT5 substantially exceeded MT5 on average, and mBART50 exceeded both despite having fewer pre-training languages; M2M-100, pre-trained for translation, performed best overall. Fine-tuning M2M-100 on combined English- or French-centric news data did not improve on its language-pair fine-tuning overall. Zero-shot M2M-100 translation was poor: the paper reports less than 5 BLEU for most supported focus languages, with Swahili above 20 and isiZulu above 13 in both directions. NEWS-only results were evaluated with BLEU and ChrF.

  4. Knowl 4 — Continual pre-training helps MT5 more consistently than the other model families

    empirical result

    Continually pre-training on African-language monolingual text improved MT5-based translation on average, including for some languages not present in the adaptation text. AfriMT5 exceeded MT5 by 1.3 BLEU for en/fr→African translation and 2.4 BLEU for African→en/fr translation. AfriByT5’s corresponding average gains over ByT5 were smaller: 0.6 and 0.9 BLEU. AfriMBART did not improve on mBART50 on average. Effects varied by language and direction: AfriByT5 improved French→Bambara by 1.9 BLEU over ByT5, and AfriMT5 improved English→Setswana by 3.6 BLEU over MT5. The authors suggest that gains on languages not included in continual pre-training may reflect language similarity; this is a possible explanation, not an established mechanism.

  5. Knowl 5 — Sequentially ending with NEWS fine-tuning is the strongest tested domain-adaptation schedule

    empirical result

    For M2M-100 evaluated on NEWS, the authors compared pooled REL+NEWS training, REL training followed by NEWS fine-tuning, and pooled REL+NEWS followed by additional NEWS fine-tuning. Their average BLEU scores for en/fr→African translation were 15.2, 16.3, and 16.7, respectively; for African→en/fr they were 18.4, 19.0, and 19.7. Thus, the REL+NEWS→NEWS schedule was best on average in both directions, improving over NEWS-only M2M-100 (15.4 and 17.7 BLEU, respectively) by 1.3 and 2.0 points. The gains from adding religious data were not uniform: simply pooling REL and NEWS did not improve on NEWS-only M2M-100 overall. Results indicate that the order and final emphasis on in-domain data matter, rather than total training-set size alone.

  6. Knowl 6 — About 2,500 in-domain pairs can outperform bilingual models trained with more out-of-domain data

    empirical result

    In experiments on English–Swahili, English–Igbo, and French–Bambara, fine-tuning M2M-100 or ByT5 on about 2,500 NEWS sentence pairs was sufficient to exceed the corresponding bilingual Transformer baseline, which had also been trained on the larger out-of-domain religious corpus. This held for Swahili, which was represented in pre-training, and for Igbo and Bambara, which were not. M2M-100 generally adapted faster than ByT5 as the in-domain training set grew, while both models continued to benefit from additional NEWS pairs. The finding concerns these tested language pairs and baselines; it does not establish a universal sentence-count threshold.

  7. Knowl 7 — Religious-domain training alone transfers poorly to news, especially with small religious corpora

    empirical result

    M2M-100 models trained on religious data performed markedly worse on the NEWS test domain than on REL, showing that religious-domain data alone did not teach the models the news-domain distribution. For translation into African languages, the reported relative BLEU drops were 95.5% for Ghomálá’ with 8K REL sentences, 93.5% for Bambara with 28K, and 93.5% for Luo with 31K. Drops were smaller, but still substantial, for languages with much larger religious corpora: 67% for isiZulu with 667K sentences, 69.3% for Swahili with 872K, and 71% for Setswana with 870K. Naija was an exception to the general relationship between corpus size and drop: with 23K REL sentences, its drop was 59.3%, which the authors suggest may relate to its similarity to its source language.

  8. Knowl 8 — NEWS fine-tuning often improves transfer to Wikipedia and religious evaluation data

    empirical result

    The authors evaluated M2M-100 before and after NEWS fine-tuning on Wikipedia (FLORES) and religious data (JW300 or Bible) using spBLEU. Transfer generally improved, though not in every language and direction. On FLORES for English/French→African translation, Igbo rose from 2.8 to 19.9 spBLEU, the largest reported gain (17.1 points); other examples include Yorùbá from 1.5 to 13.4 and isiZulu from 3.3 to 19.2. For African→English/French on FLORES, Hausa rose from 8.0 to 16.3 and isiZulu from 11.9 to 19.2, while Swahili decreased from 26.9 to 25.8. Religious-domain results also improved for many entries, but not all: for African→English/French, Hausa changed from 6.4 to 3.8 and Swahili from 15.4 to 13.9. Thus, NEWS fine-tuning often transferred beyond news, but the measured transfer was not uniformly positive.

  9. Knowl 9 — BLEU shows a larger translation-direction gap than ChrF for morphologically rich target languages

    empirical result

    For M2M-100 fine-tuned on NEWS, translation into isiZulu scored 21.0 BLEU, compared with 37.8 BLEU for isiZulu→English/French. The direction gap was smaller under ChrF: 51.2 for English/French→isiZulu versus 55.5 for isiZulu→English/French. The paper notes that BLEU can penalize morphologically rich African-language targets, including isiZulu, Setswana, Swahili, Luganda, and Ghomálá’, and that ChrF gives a smaller directional gap in the isiZulu example. These automatic metrics alone do not establish the absolute quality or adequacy of translations.

  10. Knowl 10 — Translation quality and evaluation remain limited by language imbalance and automatic metrics

    limitation

    Even the best reported models achieved low BLEU for some languages, notably Ghomálá’, Mossi, and isiZulu, especially when translating into those languages. The evaluation is primarily automatic: although the authors report both BLEU and ChrF and inspect examples, they did not conduct a deeper human evaluation, so they cannot establish absolute translation quality. Reference-based metrics may miss semantic adequacy or reward word overlap in incoherent output; manual inspection also found possible grammar and logic errors. Relative performance remains affected by unequal data availability, morphology, standardization, pre-training language coverage, and language relatedness. The study covers only 16 languages, and the authors caution that web-pretraining biases may persist despite controlled fine-tuning; they also identify the possibility of adapting the models toward harmful content as a risk.

Coverage note — Detailed per-language linguistic profiles and the full monolingual-corpus inventory were omitted because they provide supporting resource characterization rather than distinct findings needed to reconstruct the paper’s main translation and transfer contributions.

References

  1. 1.David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, and Cristina España-Bonet. 2021a. The effect of domain and diacritics in Yoruba–English neural machine translation. In Proceedings of Machine Translation Summit XVIII: Research Track, pages 61–75, Virtual. Association for Machine Translation in the Americas.
  2. 2.David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021b. MasakhaNER: Named entity recognition for African languages. Transactions of the Association for Computational Linguistics, 9:1116–1131.
  3. 3.Željko Agic and Ivan Vuli c. 2019. JW300: A wide-coverage parallel corpus for low-resource languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210, Florence, Italy. Association for Computational Linguistics.
  4. 4.Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019. Massively multilingual neural machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3874–3884, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Felermino M. D. A. Ali, Andrew Caines, and Jaimito L. A. Malavi. 2021. Towards a parallel corpus of portuguese and the bantu language emakhuwa of mozambique. ArXiv, abs/2104.05753.
  6. 6.Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur. 2020. TICO-19: the translation initiative for COvid-19. In Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020, Online. Association for Computational Linguistics.
  7. 7.Paul Azunre, Salomey Osei, Salomey Addo, Lawrence Asamoah Adu-Gyamfi, Stephen Moore, Bernard Adabankah, Bernard Opoku, Clara Asare-Nyarko, Samuel Nyarko, Cynthia Amoaba, Esther Dansoa Appiah, Felix Akwerh, Richard Nii Lante Lawson, Joel Budu, Emmanuel Debrah, Nana Boateng, Wisdom Ofori, Edwin Buabeng-Munkoh, Franklin Adjei, Isaac Kojo Essel Ampomah, Joseph Otoo, Reindorf Borkor, Standylove Birago Mensah, Lucien Mensah, Mark Amoako Marcel, Anokye Acheampong Amponsah, and James Ben Hayfron-Acquah. 2021a. Nlp for ghanaian languages. AfricaNLP Workshop, abs/2103.15475.
  8. 8.Paul Azunre, Salomey Osei, Salomey Addo, Lawrence Asamoah Adu-Gyamfi, Stephen Moore, Bernard Adabankah, Bernard Opoku, Clara Asare-Nyarko, Samuel Nyarko, Cynthia Amoaba, Esther Dansoa Appiah, Felix Akwerh, Richard Nii Lante Lawson, Joel Budu, Emmanuel Debrah, Nana Adowaa Boateng, Wisdom Ofori, Edwin Buabeng-Munkoh, Franklin Adjei, Isaac Kojo Essel Ampomah, Joseph Otoo., Reindorf Nartey Borkor, Standylove Birago Mensah, Lucien Mensah, Mark Amoako Marcel, Anokye Acheampong Amponsah, and James B. Hayfron-Acquah. 2021b. English-twi parallel corpus for machine translation. ArXiv, abs/2103.15625.
  9. 9.Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez Martínez, Mikel Forcada, Víctor Sánchez Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, and Kay Macquarrie. 2021. Surprise language challenge: Developing a neural machine translation system between Pashto and English in two months. In Proceedings of Machine Translation Summit XVIII: Research Track, pages 92–102, Virtual. Association for Machine Translation in the Americas.
  10. 10.Steven Bird. 2020. Decolonising speech and language technology. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504–3519, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  11. 11.Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8875–8898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Chris Callison-Burch, Philipp Koehn, Christof Monz, Josh Schroeder, and Cameron Shaw Fordyce, editors. 2008. Proceedings of the Third Workshop on Statistical Machine Translation. Association for Computational Linguistics, Columbus, Ohio.
  13. 13.Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020. CCAligned: A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 5960–5969, Online. Association for Computational Linguistics.
  14. 14.Chris Chinenye Emezue and Bonaventure F. P. Dossou. 2021. MMTAfrica: Multilingual machine translation for African languages. In Proceedings of the Sixth Conference on Machine Translation, pages 398–411, Online. Association for Computational Linguistics.
  15. 15.Chris Chinenye Emezue and Femi Pancrace Bonaventure Dossou. 2020. FFR v1.1: Fon-French neural machine translation. In Proceedings of the The Fourth Widening Natural Language Processing Workshop, pages 83–87, Seattle, USA. Association for Computational Linguistics.
  16. 16.Ignatius Ezeani, Paul Rayson, I. Onyenwe, C. Uchechukwu, and M. Hepple. 2020. Igbo-english machine translation: An evaluation benchmark. ArXiv, abs/2004.00648.
  17. 17.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021a. Beyond english-centric multilingual machine translation. Journal of Machine Learning Research, 22(107):1–48.
  18. 18.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021b. Beyond english-centric multilingual machine translation. Journal of Machine Learning Research, 22(107):1–48.
  19. 19.Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016. Multi-way, multilingual neural machine translation with a shared attention mechanism. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 866–875, San Diego, California. Association for Computational Linguistics.
  20. 20.∀, Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020. Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online.
  21. 21.Andargachew Mekonnen Gezmu, A. Nürnberger, and Tesfaye Bayu Bati. 2021. Extended parallel corpus for amharic-english machine translation. ArXiv, abs/2104.03543.
  22. 22.Thamme Gowda, Zhao Zhang, Chris Mattmann, and Jonathan May. 2021. Many-to-English machine translation tools, data, and pretrained models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 306–316, Online. Association for Computational Linguistics.
  23. 23.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjan Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2021. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. ArXiv, abs/2106.03193.
  24. 24.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
  25. 25.Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindvrich Helcl, and Alexandra Birch. 2021. Survey of low-resource machine translation. ArXiv, abs/2109.00486.
  26. 26.Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. 2018. Achieving human parity on automatic chinese to english news translation. CoRR, abs/1803.05567.
  27. 27.Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017. Google’s multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5:339–351.
  28. 28.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  29. 29.Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, and Mona Diab. 2021. Adapting high-resource NMT models to translate low-resource related languages without parallel data. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 802–812, Online. Association for Computational Linguistics.
  30. 30.Julia Kreutzer, Jasmijn Bastings, and Stefan Riezler. 2019. Joey NMT: A minimalist NMT toolkit for novices. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations, pages 109–114, Hong Kong, China. Association for Computational Linguistics.
  31. 31.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Auguste Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios Gonzales, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Muller, Andr’e Muller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, M. Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine cCabuk Balli, Stella Rose Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi N. Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2021. Quality at a glance: An audit of web-crawled multilingual datasets. ArXiv, abs/2103.12028.
  32. 32.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 66–75. Association for Computational Linguistics.
  33. 33.Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018. Unsupervised machine translation using monolingual corpora only. In International Conference on Learning Representations.
  34. 34.Samuel Läubli, Rico Sennrich, and Martin Volk. 2018. Has machine translation achieved human parity? a case for document-level evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4791–4796, Brussels, Belgium. Association for Computational Linguistics.
  35. 35.En-Shiun Annie Lee, Sarubi Thillainathan, Shravan Nayak, Surangika Ranathunga, David Ifeoluwa Adelani, Ruisi Su, and Arya D. McCarthy. 2022. Pre-trained multilingual sequence-to-sequence models: A hope for low-resource language translation? In Findings of ACL 2022, abs/2203.08850.
  36. 36.Zihan Liu, Genta Indra Winata, and Pascale Fung. 2021. Continual mixed-language pre-training for extremely low-resource neural machine translation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2706–2718, Online. Association for Computational Linguistics.
  37. 37.Rooweither Mabuya, Jade Abbott, and Vukosi Marivate. 2021. Umsuka english - isizulu parallel corpus. Thank you to Facebook Research for funding the creation of this dataset.
  38. 38.Lovish Madaan, Soumya Sharma, and Parag Singla. 2020. Transfer learning for related languages: Submissions to the WMT20 similar language translation task. In Proceedings of the Fifth Conference on Machine Translation, pages 402–408, Online. Association for Computational Linguistics.
  39. 39.Manuel Mager, Arturo Oncevay, Abteen Ebrahimi, John Ortega, Annette Rios, Angela Fan, Ximena Gutierrez-Vasques, Luis Chiruzzo, Gustavo Giménez-Lugo, Ricardo Ramos, Ivan Vladimir Meza Ruiz, Rolando Coto-Solano, Alexis Palmer, Elisabeth Mager-Hois, Vishrav Chaudhary, Graham Neubig, Ngoc Thang Vu, and Katharina Kann. 2021. Findings of the AmericasNLP 2021 shared task on open machine translation for indigenous languages of the Americas. In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas, pages 202–217, Online. Association for Computational Linguistics.
  40. 40.Kelly Marchisio, Kevin Duh, and Philipp Koehn. 2020. When does unsupervised machine translation work? In Proceedings of the Fifth Conference on Machine Translation, pages 571–583, Online. Association for Computational Linguistics.
  41. 41.Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky. 2020. The Johns Hopkins University Bible corpus: 1600+ tongues for typological exploration. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2884–2892, Marseille, France. European Language Resources Association.
  42. 42.Rubungo Andre Niyongabo, Qu Hong, Julia Kreutzer, and Li Huang. 2020. KINNEWS and KIRNEWS: Benchmarking cross-lingual text classification for Kinyarwanda and Kirundi. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5507–5521, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  43. 43.Evander Nyoni and Bruce A. Bassett. 2021. Low-resource neural machine translation for southern african languages. ArXiv, abs/2104.00366.
  44. 44.Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 116–126, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  45. 45.Alp Öktem, Eric DeLuca, Rodrigue Bashizi, Eric Paquin, and Grace Tang. 2021. Congolese swahili machine translation for humanitarian response. AfricaNLP Workshop.
  46. 46.Alp Öktem, Mirko Plitt, and Grace Tang. 2020. Tigrinya neural machine translation with transfer learning for humanitarian response. AfricaNLP Workshop.
  47. 47.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  48. 48.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  49. 49.Amandalynne Paullada. 2020. How does machine translation shift power? Resistance in AI Workshop.
  50. 50.Maja Popovic. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  51. 51.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  52. 52.Machel Reid, Junjie Hu, Graham Neubig, and Yutaka Matsuo. 2021. AfroMT: Pretraining strategies and reproducible benchmarks for translation of 8 African languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1306–1320, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  53. 53.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021a. WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1351–1361, Online. Association for Computational Linguistics.
  54. 54.Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021b. CCMatrix: Mining billions of high-quality parallel sentences on the web. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6490–6500, Online. Association for Computational Linguistics.
  55. 55.Kathleen Siminyu, Godson Kalipe, Davor Orlic, Jade Z. Abbott, Vukosi Marivate, Sackey Freshia, Prateek Sibal, Bhanu Neupane, David Ifeoluwa Adelani, Amelia Taylor, Jamiil Toure Ali, Kevin Degila, Momboladji Balogoun, Thierno Ibrahima Diop, Davis David, Chayma Fourati, Hatem Haddad, and Malek Naski. 2021. Ai4d - african language program. ArXiv, abs/2104.02516.
  56. 56.Y. Tang, C. Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. Multilingual translation with extensible multilingual pretraining and finetuning. ArXiv, abs/2008.00401.
  57. 57.Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018. Attaining the unattainable? reassessing claims of human parity in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 113–123, Brussels, Belgium. Association for Computational Linguistics.
  58. 58.Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021. Facebook ai wmt21 news translation task submission. arXiv preprint arXiv:2108.03265.
  59. 59.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  60. 60.Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending multilingual BERT to low-resource languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2649–2656, Online. Association for Computational Linguistics.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  62. 62.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2021a. Byt5: Towards a token-free future with pre-trained byte-to-byte models. ArXiv, abs/2105.13626.
  63. 63.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021b. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  64. 64.Jian Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang, Li Dong, Shaohan Huang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei. 2021. Multilingual machine translation systems from microsoft for wmt21 shared task. ArXiv, abs/2111.02086.

Citation

MLA
Adelani, D. I., et al. “A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3053–70, https://doi.org/10.18653/v1/2022.naacl-main.223.
APA
Adelani, D. I., Alabi, J., Fan, A., Kreutzer, J., Shen, X., Reid, M., Ruiter, D., Klakow, D., Nabende, P., Chang, E., Gwadabe, T., Sackey, F., Dossou, B. F. P., Emezue, C. C., Leong, C., Beukman, M., Muhammad, S. H., Jarso, G. D., Yousuf, O., … Manthalu, S. (2022). A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3053–3070. https://doi.org/10.18653/v1/2022.naacl-main.223
Chicago
Adelani, D. I., J. Alabi, A. Fan, et al. 2022. “A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3053–70. https://doi.org/10.18653/v1/2022.naacl-main.223.
Harvard
Adelani, D.I. et al. (2022) “A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3053–3070. Available at: https://doi.org/10.18653/v1/2022.naacl-main.223.
Vancouver
1. Adelani DI, Alabi J, Fan A, et al (2022) A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3053–3070

BibTeX

@inproceedings{adelani-etal-2022-thousand,
    title = "A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for {A}frican News Translation",
    author = "Adelani, David Ifeoluwa  and
      Alabi, Jesujoba Oluwadara  and
      Fan, Angela  and
      Kreutzer, Julia  and
      Shen, Xiaoyu  and
      Reid, Machel  and
      Ruiter, Dana  and
      Klakow, Dietrich  and
      Nabende, Peter  and
      Chang, Ernie  and
      Gwadabe, Tajuddeen  and
      Sackey, Freshia  and
      Dossou, Bonaventure F. P.  and
      Emezue, Chris  and
      Leong, Colin  and
      Beukman, Michael  and
      Muhammad, Shamsuddeen H.  and
      Jarso, Guyo D.  and
      Yousuf, Oreen  and
      Niyongabo Rubungo, Andre N.  and
      Hacheme, Gilles  and
      Wairagala, Eric Peter  and
      Nasir, Muhammad Umair  and
      Ajibade, Benjamin A.  and
      Ajayi, Tunde Oluwaseyi  and
      Gitau, Yvonne Wambui  and
      Abbott, Jade  and
      Ahmed, Mohamed  and
      Ochieng, Millicent  and
      Aremu, Anuoluwapo  and
      Ogayo, Perez  and
      Mukiibi, Jonathan  and
      Ouoba Kabore, Fatoumata  and
      Kalipe, Godson Koffi  and
      Mbaye, Derguene  and
      Tapo, Allahsera Auguste  and
      Memdjokam Koagne, Victoire M.  and
      Munkoh-Buabeng, Edwin  and
      Wagner, Valencia  and
      Abdulmumin, Idris  and
      Awokoya, Ayodele  and
      Buzaaba, Happy  and
      Sibanda, Blessing  and
      Bukula, Andiswa  and
      Manthalu, Sam",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.223/",
    doi = "10.18653/v1/2022.naacl-main.223",
    pages = "3053--3070"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/