Lifting the Curse of Multilinguality by Pre-training Modular Transformers
Jonas PfeifferNaman GoyalXi Victoria LinXian LiJames CrossSebastian RiedelMikel Artetxe
Proposes pre-training multilingual transformers with dedicated language-specific modules from the start, preventing capacity dilution across languages and enabling post-hoc extension to new languages without performance degradation.
Multilingual artificial intelligence models often face a fundamental trade-off known as the curse of multilinguality. As standard language models are trained to cover a broader set of languages, their performance on individual languages tends to degrade due to negative interference and competition for limited model capacity. This presents a critical challenge for global organizations seeking scalable language technologies that support diverse languages without sacrificing accuracy or incurring prohibitive computational costs.
The article evaluates whether pre-training models with modular, language-specific components can overcome this performance degradation. Specifically, it demonstrates the effectiveness of Cross-lingual Modular (X-MOD) architectures in preventing negative interference, enabling positive cross-lingual transfer, and supporting post-training expansion to previously unseen languages without losing accuracy.
The authors conducted large-scale pre-training experiments using the CC100 dataset across language sets of 13, 30, 60, and 75 typologically diverse languages. The baseline fully shared architecture was compared against the proposed modular architecture, which combines shared core network weights with specialized, language-specific bottleneck modules. Importantly, the modular model allocates unique modules to individual languages while keeping the active parameter count and computational cost during training and inference identical to the baseline. The models were evaluated on standard downstream benchmarks, including natural language inference, named entity recognition, and question answering, under conditions controlling for both training update steps and per-language data exposure.
The findings show that modular pre-training successfully eliminates the curse of multilinguality. First, while fully shared models suffered substantial performance drops as language count scaled, the modular model maintained and even improved accuracy across both high- and low-resource languages. Second, the modular model trained on 60 languages consistently outperformed the shared baseline across all downstream tasks, achieving an overall accuracy of 73.5% versus 72.5% on natural language inference and 62.8% versus 58.8% F1 score on named entity recognition. Third, when expanding to new, held-out languages after pre-training, the modular approach matched the accuracy of pre-training on those languages from the start, regardless of whether related languages existed in the initial training set. Finally, adding modular components after pre-training failed to recover performance, demonstrating that modularity must be integrated from the beginning.
These results demonstrate that organizations can scale multilingual models to dozens of languages without sacrificing per-language performance. Because language-specific modules can be routed dynamically and stored efficiently, operational costs and computational throughput remain stable as new languages are added. Furthermore, eliminating the need to retrain massive core models from scratch to support new markets significantly reduces deployment timelines, compute budgets, and maintenance overhead.
Technical leaders and practitioners should adopt modular pre-training architectures when building multilingual systems that require broad or evolving language coverage. Initial pre-training should focus on a representative, medium-to-large set of higher-resource languages, while lower-resource or emerging target languages can be added post-hoc with dedicated modules. While modular pre-training requires slightly longer initial training for modular benefits to fully materialize in smaller language pools, it offers a scalable, future-proof framework for global language coverage.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R’s large-scale multilingual pretraining experiments expose the capacity and language-interference trade-offs that X-MOD is designed to address.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This earlier account of shared multilingual pretraining objectives and cross-lingual transfer provides the baseline framework that the source modifies with language-specific modules.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Its analysis of multilingual BERT’s shared-parameter cross-lingual transfer helps clarify the interference problem that motivates the source’s modular alternative.
No sufficiently relevant recommendations were found.
