A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation
David Ifeoluwa AdelaniJesujoba O. AlabiAngela FanJulia KreutzerXiaoyu ShenMachel ReidDana RuiterDietrich KlakowPeter NabendeErnie Chang
Demonstrates that fine-tuning large multilingual models on just a few thousand high-quality in-domain sentence pairs effectively transfers translation capabilities to sixteen low-resource African languages absent from the original pre-training data.
Modern automated translation systems rely heavily on massive web crawls, leaving out hundreds of widely spoken African languages that lack extensive online representation. Because critical news across the African continent is predominantly published in English, French, or Arabic, millions of native speakers face severe information bottlenecks during crises. The article evaluates how existing large-scale pre-trained translation models can be adapted to 16 under-resourced African languages and transferred into the news domain, where parallel training data has historically been scarce.
To achieve this, the authors created MAFAND-MT, a participatory news corpus spanning 16 African languages—including eight languages entirely new to machine translation evaluation benchmarks. Built by native speakers and professional translators, the dataset contributed between 2,000 and 8,000 parallel sentences per language paired with English or French. The researchers then adapted and evaluated several major multilingual models (including M2M-100, MT5, ByT5, and mBART50) across varied configurations: zero-shot transfer, continual pre-training on collected monolingual text, and fine-tuning on small parallel datasets across both religious and news domains.
The investigation produced four central findings. First, zero-shot translation completely failed for most unseen African languages, often yielding scores below 5 BLEU. Second, fine-tuning large pre-trained models on as few as 2,000 to 2,500 target-domain sentence pairs proved remarkably effective, consistently outperforming standard bilingual translation models trained from scratch on much larger datasets. Third, M2M-100 emerged as the strongest overall system, while the byte-based ByT5 model outperformed token-based MT5 by over 3 BLEU points. Fourth, relying solely on legacy out-of-domain data caused performance drops of up to 95.5%, whereas a staged adaptation strategy—training first on combined religious and news data followed by specialized news fine-tuning—delivered the highest translation quality.
These findings demonstrate that local communities and developers do not need massive web-crawled datasets or multimillion-dollar computational budgets to build functional translation tools. Targeted collections of a few thousand clean, professionally verified translations can effectively steer large pre-trained models toward new languages and domains while mitigating archaic domain biases. However, the researchers caution that overall translation scores for several complex or highly isolating languages remain modest, and automated metrics do not capture subtle grammatical and logical errors. Organizations adopting these techniques should prioritize rigorous human evaluation before deploying models into production environments and continue expanding small-scale, high-quality parallel data collection for additional under-resourced languages.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). mBART establishes how a multilingual denoising-pretrained translation model can be fine-tuned for low-resource translation, the central strategy this paper investigates.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R shows how large-scale multilingual pretraining can transfer to underrepresented languages, clarifying the pretrained-model foundation this paper adapts.
- Paper: Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks, Suchin Gururangan et al. (2020). This study establishes continued domain-specific pretraining as an adaptation strategy, preparing readers for the paper’s examination of transfer to news domains.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Omnilingual MT extends low-resource translation adaptation to 1,600 languages, scaling the challenge of making pretrained translation models work beyond their original language coverage.
