Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
Ayyoob ImaniPeiqin LinAmir Hossein KargaranSilvia SeveriniMasoud Jalili SabetNora KassnerChunlan MaHelmut SchmidAndré F. T. MartinsFrançois Yvon
Presents an open-source multilingual corpus and language model spanning over 500 predominantly low-resource languages, outperforming XLM-R across multiple natural language understanding benchmarks while identifying the key drivers of multilingual representation quality.
Natural language processing technologies have largely focused on making large language models deeper and more capable for roughly 100 well-resourced languages. This narrow scope excludes thousands of the world's languages, creating significant disparities in access to language technology across diverse global communities and digital ecosystems.
The article demonstrates that multilingual language models can be successfully scaled horizontally to support hundreds of under-resourced languages through curated corpus collection and continued pretraining.
To accomplish this, researchers collected a massive dataset covering 2,266 languages from approximately 150 diverse sources, including web crawls, translations, and religious texts. After applying multi-stage text cleaning and setting a threshold of at least 30,000 sentences per language-script, they created a 600-gigabyte training dataset covering 511 languages across 534 language-scripts. Using this data, they extended the vocabulary and continued the masked pretraining of a standard multilingual baseline model to create an expanded 395-million-parameter model, which was systematically evaluated across five diverse downstream understanding tasks alongside linguistic pseudoperplexity.
The resulting model demonstrated substantial performance improvements over standard baselines. On low-resource tail languages, sentence retrieval accuracy increased dramatically from under 10% to over 43% on aligned texts, and text classification accuracy improved from under 14% to over 46%. Performance on sequence labeling tasks like named entity recognition and part-of-speech tagging also increased by 13 to 21 percentage points for tail languages while maintaining or slightly improving representation quality for high-resource head languages. Furthermore, empirical analysis showed that model quality is driven by a combination of factors, including target corpus size, native script coverage, model capacity, and linguistic synergy from related neighboring languages.
These findings prove that organizations do not need to choose between supporting well-resourced languages and expanding into long-tail languages. Expanding language coverage enables positive cross-lingual transfer without requiring proportional increases in model size, drastically lowering the computational barrier to deploy equitable language technology in underserved regions.
Decision-makers and engineering teams should adopt horizontal scaling strategies and open datasets when developing multilingual systems. Future efforts should focus on training larger architectures, developing distilled lightweight models for cost-efficient deployment, and evaluating horizontal models on broader end-user applications.
The analysis is subject to some residual noise within web-crawled texts, unaddressed societal biases in training datasets, and potential variations arising from limited hyperparameter tuning. Nevertheless, the broad evaluation across hundreds of standardized language-scripts provides high confidence in horizontal scaling as an effective, scalable strategy for global language coverage.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). Introduces XLM-R, the foundational multilingual masked language model and primary baseline that Glot500 directly adapts and evaluates against across 500+ languages.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). Pioneers massive multilingual scale beyond 200 languages with curated datasets and evaluation benchmarks like Flores-200, establishing the foundation for extending NLP to low-resource languages.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). Establishes the resource-classification taxonomy and highlights digital disparities across the world's languages, directly motivating Glot500's horizontal scaling effort.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). Presents mT5 and the mC4 multilingual corpus covering 101 languages, illustrating the principles of massive multilingual pretraining and subword vocabulary scaling.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Provides the foundational cross-lingual language model pretraining objectives that modern multilingual representation models build upon.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Introduces the standard cross-lingual natural language inference benchmark that underpins evaluation methodology across diverse languages.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). Documents open-access multilingual language model training across diverse language families and provides key context on multilingual data governance.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Pushes massively multilingual coverage even further by scaling machine translation models and corpora to over 1,600 languages.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Comprehensively evaluates how modern generative language models perform across 70 typologically diverse and low-resource languages.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). Builds upon massive multilingual representation learning to create multi-functional, multi-granular text embeddings across nearly 200 languages.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). Analyzes the internal neuron mechanisms and language-specific representations that enable cross-lingual transfer in multilingual language models.
