No Language Left Behind: Scaling Human-Centered Machine Translation
NLLB TeamMarta R. Costa-jussàJames CrossOnur ÇelebiMaha ElbayadKenneth HeafieldKevin HeffernanElahe KalbassiJanice LamDaniel Licht
Presents an open-source mixture-of-experts machine translation model capable of translating across more than 200 languages, delivering a 44% BLEU improvement over prior systems while establishing comprehensive benchmarks for low-resource evaluation and translation safety.
The vast majority of modern advances in artificial intelligence and machine translation have focused on a small subset of high-resource languages such as English, French, and Spanish, largely excluding the majority of the world’s languages. This technological imbalance severely limits digital inclusion, educational access, and cultural preservation for underserved, low-resource language communities globally. The article addresses this systemic inequity by developing an end-to-end, human-centered framework capable of delivering high-quality, safe, and fluent machine translation for more than 200 languages simultaneously.
The core objective was to design, train, and comprehensively evaluate a massively multilingual translation model that effectively doubles the language coverage of previous state-of-the-art systems while actively mitigating cross-language interference and minimizing translation toxicity.
To achieve this, the authors adopted a multi-stage approach combining human-centered qualitative research with large-scale computational engineering. They conducted in-depth interviews with 44 native speakers across 36 low-resource languages to establish foundational design principles and understand critical community needs. They constructed professionally translated seed and evaluation benchmarks, including Flores-200, which encompasses 3,001 human-translated sentences across 204 languages and establishes over 40,000 distinct translation directions. To address severe data scarcity, they implemented a lightweight fasttext-based language identification system capable of classifying over 200 languages and deployed teacher-student distillation models (LASER3) across 37.7 petabytes of web corpora to mine over 1.1 billion parallel sentence pairs. These mined and curated datasets were then utilized to train large-scale neural machine translation systems, most notably a 54.5-billion-parameter Sparsely Gated Mixture of Experts model utilizing conditional compute.
The evaluation revealed several key findings: First, the primary model achieved an overall 44% relative improvement in translation quality (BLEU score) over previous state-of-the-art benchmarks on Flores-101 while scaling to more than 200 languages. Second, the specialized sentence encoders substantially reduced multilingual alignment error rates—dropping average error on selected low-resource languages from 61% to under 1%—which directly enabled effective bitext mining even for languages with fewer than 100,000 native parallel sentences. Third, the conditional compute Mixture of Experts architecture successfully balanced cross-lingual knowledge transfer with minimal interference between unrelated languages, outperforming dense transformer baselines while preserving computational efficiency. Finally, combining automated filtering with newly curated toxic wordlists across 200 languages substantially mitigated the generation of harmful and hallucinatory translations.
These findings demonstrate that neural translation systems can successfully scale beyond 200 languages without sacrificing translation quality or safety. By significantly lowering the data threshold required to represent underserved languages, the work proves that high-quality translation can be democratized to support digital participation, educational equity, and knowledge platforms such as Wikipedia. However, because performance remains dependent on the volume of usable web text, extremely low-resource languages and predominantly oral languages still face notable challenges.
Organizations and researchers building upon this work should adopt the open-sourced models, benchmarks, and data pipelines while prioritizing human-in-the-loop validation for mission-critical deployments. Deploying smaller, distilled versions of the model (ranging from 600M to 3.3B parameters) is recommended for resource-constrained environments to balance inference cost with strong translation accuracy. Further research should focus on expanding support for oral and unstandardized languages, refining language identification for highly confusable dialects, and continually addressing cultural and domain generalization beyond web-centric corpora.
- Paper: GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding, Dmitry Lepikhin et al. (2021). Introduces the conditional computation and Sparsely-Gated Mixture-of-Experts architecture across 100 languages that NLLB directly adapts and scales to overcome capacity bottlenecks in 200+ language machine translation.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Presents the foundational Sparsely-Gated Mixture-of-Experts layer design that forms the architectural backbone for NLLB's conditional compute translation models.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Establishes multilingual denoising sequence-to-sequence pre-training (mBART), providing the baseline framework for scaling multilingual encoder-decoder translation to low-resource languages.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). Demonstrates large-scale cross-lingual representation learning on 100 languages using filtered CommonCrawl data, establishing principles of multilingual vocabulary scaling and low-resource data sampling leveraged in NLLB.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces the Transformer encoder-decoder architecture upon which the NLLB translation models and conditional compute extensions are built.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). Describes learned semantic evaluation metrics for text generation that inform modern quality and safety assessments across translation directions.
- Paper: Bleu: a Method for Automatic Evaluation of Machine Translation, Kishore Papineni et al. (2002). Defines the standard BLEU metric used by NLLB to benchmark translation quality improvements across over 40,000 language directions.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Directly builds upon NLLB by scaling human-centered machine translation from 200 languages to over 1,600 languages using specialized LLMs and updated evaluation benchmarks.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Investigates stability and transfer improvements for sparse Mixture-of-Experts models, refining the architectural paradigms used during large-scale conditional compute training.
- Paper: Mixtral of Experts, Albert Q. Jiang et al. (2024). Extends sparse Mixture-of-Experts scaling techniques into open foundation models to achieve highly efficient multilingual generation and reasoning.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). Analyzes the internal representations and neuron-level language transfer mechanisms of massive multilingual models like those scaled in NLLB.
