Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model
Yeskendir KoishekenovAlexandre BerardVassilina Nikoulina
Proposes an inference-time pruning method that removes up to 80% of experts from the 54.5-billion-parameter NLLB-200 translation model without fine-tuning, reducing hardware requirements from multiple devices down to a single 32GB GPU while maintaining translation quality.
State-of-the-art machine translation across hundreds of languages increasingly relies on massive artificial intelligence models. A leading example is NLLB-200, a translation model covering 202 languages whose largest variant contains 54.5 billion parameters structured as a Mixture of Experts—an architecture that divides model capacity into specialized sub-networks called experts and activates only a subset for each piece of text. While this design delivers industry-leading translation quality, running the full model requires specialized hardware with at least four 32-gigabyte graphics processing units (GPUs), creating high computational costs and severe deployment bottlenecks for production environments.
The article investigates whether multilingual models can be compressed at inference time without requiring computationally expensive retraining or fine-tuning. Specifically, the objective is to evaluate whether specific experts naturally specialize in individual languages, and to demonstrate a pruning strategy that eliminates redundant experts so the model can operate on a single 32-gigabyte GPU while maintaining translation quality.
The researchers designed importance metrics based on how frequently and confidently the model selects each expert when translating validation text. They tested various pruning algorithms and retention rates across 53 representative languages spanning high-, low-, and very low-resource categories, validating their final setup across all 202 languages and 40,602 translation directions using the standard FLORES-200 benchmark dataset.
The key findings reveal that up to 80% of all experts can be safely pruned without fine-tuning while incurring a negligible average translation quality drop of approximately 0.28 points out of 100 on standard translation metrics. The analysis demonstrates that an unbalanced pruning ratio keeping three times as many experts in the encoder (input processing) as in the decoder (output generation) delivers the best performance. Furthermore, the experiments prove that language-specific specialization genuinely emerges in the model: decoder experts show strong specialization for the target language (sharing 68% to 87% similarity for the same language) and naturally cluster along linguistic family lines, while encoder experts remain largely independent of the target language.
These findings have direct operational and economic implications. Organizations can reduce the hardware footprint required to run top-tier multilingual translation by 75%, cutting infrastructure demands from four GPUs to a single 32-gigabyte GPU. When four GPUs remain available, the pruned model doubles the translation throughput. This eliminates major cost barriers and enables efficient localized translation services across both major and very low-resource languages without degrading quality below dense baseline models.
For deployment, the article recommends adopting a language-specific pruning configuration using fixed allocations per layer with a 3:1 encoder-to-decoder expert ratio. Practitioners should select encoder experts based on the source language and decoder experts based on the target language, utilizing the publicly released expert statistics to avoid recomputing routing metrics. Further development should focus on addressing a minor post-pruning tendency toward sentence repetition and slight hallucination by identifying and preserving experts specialized in signaling the end of generated sequences.
Decision-makers should note certain limitations: the conclusions are drawn exclusively from the NLLB-200 architecture and its particular training objective, so pruning patterns may differ for other models. While the evaluation across all 202 languages provides high confidence in general quality preservation, teams deploying pruned models for sensitive or automated downstream tasks should monitor output length ratios and translation biases.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). Read the NLLB-200 account first to understand the exact multilingual MoE model, language coverage, and evaluation benchmark that this paper compresses.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Its sparsely gated MoE layer establishes how routers dispatch inputs to specialized experts, the architecture this paper prunes.
- Paper: GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding, Dmitry Lepikhin et al. (2021). GShard shows how sparsely gated experts scale to multilingual translation, preparing you to understand the large-scale MoE setting examined here.
- Paper: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, William Fedus et al. (2022). Switch Transformers develops sparse expert routing for multilingual models, clarifying the routing and expert structure relevant to this pruning study.
- Paper: Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models, Xudong Lu et al. (2024). This later work carries expert pruning beyond multilingual translation, testing post-training pruning and dynamic expert skipping on general-purpose MoE language models.
- Paper: Does a Global Perspective Help Prune Sparse MoEs Elegantly?, Zeliang Zhang et al. (2026). This later study extends expert pruning by allocating compression globally across MoE layers rather than relying on uniform layer-wise pruning.
