Learning Word Vectors for 157 Languages
Edouard GravePiotr BojanowskiPrakhar GuptaArmand JoulinTomas Mikolov
Provides pre-trained fastText word embeddings for 157 languages trained on massive Wikipedia and Common Crawl data alongside new analogy evaluation benchmarks for French, Hindi, and Polish.
Modern natural language processing systems rely heavily on pre-trained word representations, or word vectors, which capture the meanings and relationships of words based on their context. While high-quality models have been widely available for English, language coverage across other global languages has remained severely restricted. Many existing multilingual initiatives rely almost exclusively on Wikipedia, which is too small in many non-English languages to provide the broad vocabulary coverage required for practical applications.
The article evaluates and demonstrates a scalable pipeline for training high-quality word vectors across 157 languages. It combines curated encyclopedic text with massive, real-world web data and introduces targeted evaluation benchmarks to test vector accuracy across diverse linguistic families.
To construct the training corpus, the authors processed Wikipedia dumps alongside roughly 24 terabytes of web text extracted from the May 2017 Common Crawl. The processing pipeline utilized a newly built fast language detector supporting 176 languages, filtered lines by length and confidence, and removed significant boilerplate through deduplication, eliminating 37% of web crawl text and 21% of Wikipedia text. The models were trained using an enhanced Continuous Bag of Words (CBOW) architecture incorporating position weights and character-level subword information. The authors evaluated performance using word analogy tasks across 10 languages, creating new analogy datasets for French, Hindi, and Polish.
The findings show that combining web crawl data with Wikipedia dramatically improves model accuracy for languages with smaller Wikipedia collections. Adding crawl data raised analogy accuracy by 23.5 percentage points for Finnish, 17.8 points for Chinese, 16.0 points for Hindi, and 9.7 points for Polish. Overall, the full pipeline improved average accuracy across the 10 evaluated languages from 51.0% in the baseline model to 66.7%. Model architectural enhancements, particularly the position-weighted CBOW model and increased training iterations, provided consistent performance gains. In contrast, for high-resource languages such as German, French, and Spanish, adding web data did not notably improve analogy accuracy because Wikipedia was already extensive and closely aligned with the analogy evaluation benchmarks.
These results indicate that incorporating large-scale web text is an effective, practical method for scaling language technologies to under-resourced languages without waiting for curated encyclopedia data to grow. Organizations building multilingual digital services can achieve substantially broader vocabulary coverage and higher performance in global markets. However, web data also introduces noise and requires rigorous filtering pipelines to prevent performance degradation.
Organizations should adopt position-weighted subword embeddings and integrate filtered web-scale corpora when developing language processing systems for non-English markets. For future initiatives, additional research is required to close the persistent performance gap in lower-resource languages such as Hindi, which achieved an analogy accuracy of only 32.1% despite the inclusion of web data.
The primary limitations of this work stem from domain differences between training data and evaluation benchmarks, as well as the inherent quality challenges of noisy web text in low-resource settings. While confidence is high in the overall effectiveness and scalability of the training pipeline, readers should exercise caution when deploying models in languages with very limited data, where accuracy remains markedly lower than in high-resource counterparts.
- Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). It introduces the fastText model that uses subword character n-grams to represent morphologically rich languages, which forms the direct core architecture scaled to 157 languages in the source paper.
- Paper: Distributed Representations of Words and Phrases and their Compositionality, Tomas Mikolov et al. (2013). It develops skip-gram with negative sampling and phrase representations, establishing the fundamental continuous-vector estimation techniques underlying the source's training pipeline.
- Paper: Efficient Estimation of Word Representations in Vector Space, Tomáš Mikolov et al. (2013). It proposes the original Continuous Bag-of-Words and Skip-gram architectures, laying the groundwork for scalable distributed word embeddings.
- Paper: GloVe: Global Vectors for Word Representation, Jeffrey Pennington et al. (2014). It provides the benchmark global log-bilinear word embedding method and standard word analogy evaluation protocols that the source compares against.
- Paper: Exploiting Similarities among Languages for Machine Translation, Tomas Mikolov et al. (2013). It demonstrates how independent monolingual vector spaces across diverse languages exhibit geometric similarities that enable cross-lingual transfer.
- Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). It establishes the vector offset arithmetic and word analogy tasks used in the source to evaluate multilingual syntactic and semantic representations.
- Paper: Character-Aware Neural Language Models, Yoon Kim et al. (2015). It highlights the necessity and effectiveness of subword- and character-aware representations for handling morphologically rich languages.
- Paper: Word Translation Without Parallel Data, Alexis Conneau et al. (2017). It presents unsupervised methods to align multilingual monolingual word vectors without parallel corpora, motivating large-scale monolingual vector training.
- Paper: word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method, Yoav Goldberg et al. (2014). It delivers the theoretical and mathematical foundations of negative sampling objective functions used to efficiently train word embeddings.
- Paper: A Neural Probabilistic Language Model, Yoshua Bengio et al. (2003). It introduces the foundational concept of learning continuous distributed vector representations for words using neural networks.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). It extends massive multilingual pre-training from static word vectors on Common Crawl and Wikipedia to deep contextualized Transformer language models across 100 languages.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). It advances cross-lingual representation learning from monolingual word vectors to pre-trained multilingual language models that leverage shared subword vocabularies and cross-lingual pretraining objectives.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). It establishes a standardized cross-lingual natural language inference benchmark across 15 languages, serving as a primary downstream evaluation suite for representations trained across multiple languages.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It scales massively multilingual pre-training over 101 Common Crawl languages into a unified sequence-to-sequence text-to-text generative Transformer architecture.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). It investigates how massively multilingual representations generalize across distinct scripts and language families in zero-shot cross-lingual transfer settings.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). It scales multilingual representation and translation to over 200 languages using fasttext-based language identification and massive web crawling.
- Paper: Stanza: A Python Natural Language Processing Toolkit for Many Human Languages, Peng Qi et al. (2020). It operationalizes multilingual deep learning models into a comprehensive, multi-task NLP library supporting 66 diverse languages directly from raw text.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). It extends multilingual dense embeddings to support multi-granularity and hybrid sparse-dense retrieval across more than 100 languages.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). It pushes the frontiers of massive multilingual coverage to over 1,600 languages by combining large web corpora with specialized multilingual alignment techniques.
