Learning Word Vectors for 157 Languages

Edouard GravePiotr BojanowskiPrakhar GuptaArmand JoulinTomas Mikolov

article2018International Conference on Language Resources and Evaluation1,626 citations

Provides pre-trained fastText word embeddings for 157 languages trained on massive Wikipedia and Common Crawl data alongside new analogy evaluation benchmarks for French, Hindi, and Polish.

Listen

Modern natural language processing systems rely heavily on pre-trained word representations, or word vectors, which capture the meanings and relationships of words based on their context. While high-quality models have been widely available for English, language coverage across other global languages has remained severely restricted. Many existing multilingual initiatives rely almost exclusively on Wikipedia, which is too small in many non-English languages to provide the broad vocabulary coverage required for practical applications.

The article evaluates and demonstrates a scalable pipeline for training high-quality word vectors across 157 languages. It combines curated encyclopedic text with massive, real-world web data and introduces targeted evaluation benchmarks to test vector accuracy across diverse linguistic families.

To construct the training corpus, the authors processed Wikipedia dumps alongside roughly 24 terabytes of web text extracted from the May 2017 Common Crawl. The processing pipeline utilized a newly built fast language detector supporting 176 languages, filtered lines by length and confidence, and removed significant boilerplate through deduplication, eliminating 37% of web crawl text and 21% of Wikipedia text. The models were trained using an enhanced Continuous Bag of Words (CBOW) architecture incorporating position weights and character-level subword information. The authors evaluated performance using word analogy tasks across 10 languages, creating new analogy datasets for French, Hindi, and Polish.

The findings show that combining web crawl data with Wikipedia dramatically improves model accuracy for languages with smaller Wikipedia collections. Adding crawl data raised analogy accuracy by 23.5 percentage points for Finnish, 17.8 points for Chinese, 16.0 points for Hindi, and 9.7 points for Polish. Overall, the full pipeline improved average accuracy across the 10 evaluated languages from 51.0% in the baseline model to 66.7%. Model architectural enhancements, particularly the position-weighted CBOW model and increased training iterations, provided consistent performance gains. In contrast, for high-resource languages such as German, French, and Spanish, adding web data did not notably improve analogy accuracy because Wikipedia was already extensive and closely aligned with the analogy evaluation benchmarks.

These results indicate that incorporating large-scale web text is an effective, practical method for scaling language technologies to under-resourced languages without waiting for curated encyclopedia data to grow. Organizations building multilingual digital services can achieve substantially broader vocabulary coverage and higher performance in global markets. However, web data also introduces noise and requires rigorous filtering pipelines to prevent performance degradation.

Organizations should adopt position-weighted subword embeddings and integrate filtered web-scale corpora when developing language processing systems for non-English markets. For future initiatives, additional research is required to close the persistent performance gap in lower-resource languages such as Hindi, which achieved an analogy accuracy of only 32.1% despite the inclusion of web data.

The primary limitations of this work stem from domain differences between training data and evaluation benchmarks, as well as the inherent quality challenges of noisy web text in low-resource settings. While confidence is high in the overall effectiveness and scalability of the training pipeline, readers should exercise caution when deploying models in languages with very limited data, where accuracy remains markedly lower than in high-resource counterparts.

Cover for Learning Word Vectors for 157 Languages

Abstract

Distributed word representations, or word vectors, have recently been applied to many tasks in natural language processing, leading to state-of-the-art performance. A key ingredient to the successful application of these representations is to train them on very large corpora, and use these pre-trained models in downstream tasks. In this paper, we describe how we trained such high quality word representations for 157 languages. We used two sources of data to train these models: the free online encyclopedia Wikipedia and data from the common crawl project. We also introduce three new word analogy datasets to evaluate these word vectors, for French, Hindi and Polish. Finally, we evaluate our pre-trained word vectors on 10 languages for which evaluation datasets exists, showing very strong performance compared to previous models.

Table of Contents

  • 1. Introduction
  • 2. Training Data
  • 2.1. Wikipedia
  • 2.2. Common Crawl
  • 2.3. Deduplication and Tokenization
  • 3. Models
  • 4. Evaluations
  • 4.1. Evaluation Datasets
  • 4.2. Model Variants
  • 4.3. Results
  • 5. Conclusion
  • 6. Bibliographical References
  • References

Knowls

  1. Knowl 1 — Position-Weighted Continuous Bag-of-Words (CBOW) with Subword Embeddings

    model/method

    In position-weighted subword Continuous Bag-of-Words (CBOW), each word is represented as the sum of its character nn-gram vector representations. Given a target word w0w_0 and a context window of size nn spanning context words w−n,…,w−1,w1,…,wnw_{-n}, \dots, w_{-1}, w_1, \dots, w_n, the context vector h∈Rdh \in \mathbb{R}^d is computed by taking the element-wise product of each context word representation uwi∈Rdu_{w_i} \in \mathbb{R}^d with a position-specific weight vector ci∈Rdc_i \in \mathbb{R}^d:

    h=∑i=−ni≠0nci⊙uwih = \sum_{\substack{i=-n \\ i\neq 0}}^n c_i \odot u_{w_i}

    where ⊙\odot denotes the Hadamard (element-wise) product, each cic_i is a learned position vector corresponding to relative position index ii, and uwiu_{w_i} is the subword-augmented embedding for word wiw_i, formed by summing the vector embeddings of all character nn-grams contained in wiw_i (including the full word token itself). The model is optimized to predict the target word w0w_0 conditioned on the context representation hh.

  2. Knowl 2 — Ablation Analysis of Model Architecture and Hyperparameters for Word Analogies

    data/table

    Word vectors were evaluated on word analogy tasks (xB−xA+xCx_B - x_A + x_C, querying the closest vocabulary word excluding A,B,CA, B, C) restricted to the top 200,000 most frequent words. Successively modifying the fastText baseline across 10 languages (Czech: CS, German: DE, Spanish: ES, Finnish: FI, French: FR, Hindi: HI, Italian: IT, Polish: PL, Portuguese: PT, Chinese: ZH) demonstrates the contribution of each design component:

    Model Configuration CS DE ES FI FR HI IT PL PT ZH Average
    Baseline (Skipgram, n-gram 3–6) 63.1 61.0 57.4 35.9 64.2 10.6 56.3 53.4 54.0 60.2 51.0
    n-gram 5–5 57.7 61.8 57.5 39.4 65.9 8.3 57.2 54.5 54.8 59.3 50.9
    + Position-weighted CBOW 63.9 71.7 64.4 42.8 71.6 14.1 66.2 56.0 60.6 51.5 55.5
    + 10 Negatives 64.8 73.7 65.0 45.0 73.5 14.5 68.0 58.3 62.9 56.0 57.4
    + 10 Epochs 64.6 73.9 67.1 46.8 74.9 16.1 69.3 58.2 64.7 60.6 58.8
    + Common Crawl Data 69.9 72.9 65.4 70.3 73.6 32.1 69.8 67.9 66.7 78.4 66.7

    Restricting character nn-grams to length 5 maintains average accuracy (50.9% vs. 51.0%) while accelerating training. Switching from skipgram to position-weighted CBOW provides the largest single algorithmic gain (+4.6% on average). Increasing negative examples from 5 to 10 and training epochs from 5 to 10 improves average accuracy to 58.8%. Adding Common Crawl data yields the overall best average accuracy of 66.7%.

  3. Knowl 3 — Multilingual Text Extraction, Language Identification, and Deduplication Pipeline

    model/method

    The text pre-processing pipeline for preparing training data from Wikipedia and Common Crawl across 157 languages consists of three stages:

    1. Line-Level Language Identification: Because individual web pages frequently contain multiple languages, language identification is executed per line using a fastText linear classifier recognizing 176 languages. The classifier is trained on 400M Wikipedia tokens and Tatoeba sentences using character nn-grams of lengths 2, 3, and 4 with hierarchical softmax. Lines are kept only if they exceed 100 characters in length and achieve a classifier confidence score ≥0.8\ge 0.8.
    2. Line-Level Hash Deduplication: Exact duplicate lines are identified and filtered out by computing 32-bit line hashes within each language independently. This removes boilerplate text, eliminating 37% of Common Crawl lines and 21% of Wikipedia lines.
    3. Script- and Language-Specific Tokenization: Text is tokenized using specialized tools:
    • Stanford word segmenter for Chinese;
    • Mecab for Japanese;
    • UETsegmenter for Vietnamese;
    • Europarl pre-processing tools for languages using Latin, Cyrillic, Hebrew, or Greek scripts;
    • ICU tokenizer for all remaining languages.
  4. Knowl 4 — Impact of Large-Scale Web Crawl Training on Resource-Constrained vs. High-Resource Languages

    empirical result

    Augmenting Wikipedia training data with Common Crawl web data produces substantial performance gains on word analogy tasks for languages with smaller Wikipedia collections:

    • Finnish: +23.5 percentage points (46.8% to 70.3%)
    • Chinese: +17.8 percentage points (60.6% to 78.4%)
    • Hindi: +16.0 percentage points (16.1% to 32.1%)
    • Polish: +9.7 percentage points (58.2% to 67.9%)
    • Czech: +5.3 percentage points (64.6% to 69.9%)

    Conversely, for high-resource languages with large, curated Wikipedia dumps (e.g., German: 73.9% vs. 72.9%; French: 74.9% vs. 73.6%; Spanish: 67.1% vs. 65.4%), adding web crawl data slightly reduces or preserves accuracy on Wikipedia-aligned analogy benchmarks.

    Vocabulary coverage of the top 200,000 words on the analogy benchmarks differs between data sources as follows:

    Source CS DE ES FI FR HI IT PL PT ZH
    Wikipedia 76.9% 79.1% 93.9% 94.6% 88.1% 70.8% 80.9% 69.5% 79.2% 100.0%
    Common Crawl 78.6% 81.1% 90.4% 92.2% 92.5% 70.7% 82.6% 63.4% 75.7% 100.0%
  5. Knowl 5 — Multilingual Word Analogy Dataset Construction for French, Hindi, and Polish

    experimental setup

    Three word analogy datasets were constructed by translating and adapting the English Google word analogy dataset to reflect target language morphology and syntax:

    • French (31,688 questions): Categories capital-common-countries, capital-world, and currency were directly translated. family excluded ambiguous words (such as fils) and nominal phrases. city-in-state was replaced with capitals of French départements. An antonyms-adjectives category was added (e.g., chaud / froid). Periphrastic categories comparative and superlative were removed, and a past-participle category was introduced (e.g., pouvoir / pu).
    • Hindi: capital-common-countries, capital-world, and currency were translated directly. family removed multi-word phrases and introduced distinctions between paternal and maternal relations (e.g., dādā - dādī vs. nānā - nānī). city-in-state was adapted with Indian city-state pairs. Phrasal categories adjective-to-adverb, comparative, superlative, present-participle, and past-tense were removed, while a new adjective-to-noun category was added (e.g., mīṭhā [sweet] to miṭhās [sweetness]).
    • Polish (24,570 questions): Standard semantic categories were directly translated. city-in-state was adapted to capital cities of Polish regions (województwo). Syntactic category plural-verbs was replaced with verb-aspect (e.g., imperfective pairs distinguishing directional vs. aimless motion: iść / chodzić), and past-tense incorporated mixed perfective and imperfective aspects.
  6. Knowl 6 — Scale Disparity Between Wikipedia and Common Crawl Multilingual Text Corpora

    data/table

    The May 2017 Common Crawl extraction yields token counts that surpass Wikipedia dumps by one to two orders of magnitude across major languages:

    Language Wikipedia (Sep 2017) Common Crawl (May 2017)
    # Tokens # Words (≥5\ge 5) # Tokens Vocabulary Size
    Russian 823,849,081 2,230,231 102,825,040,945 14,679,750
    Japanese 998,774,138 916,262 92,827,457,545 9,073,245
    Spanish 797,362,600 1,337,109 72,493,240,645 10,614,696
    French 1,107,636,871 1,668,310 68,358,270,953 12,488,607
    German 1,384,170,636 3,005,294 65,648,657,780 19,767,924
    Italian 702,638,442 1,169,177 36,237,951,419 10,404,913
    Portuguese 386,107,589 815,284 35,841,247,814 8,370,569
    Chinese 374,650,371 1,486,735 30,176,342,544 17,599,492
    Polish 386,874,622 1,298,250 21,859,939,298 10,209,556
    Czech 178,516,890 784,896 13,070,585,221 8,694,576
    Finnish 127,176,620 880,713 6,059,887,126 9,782,381
    Hindi 39,733,591 183,211 1,885,189,625 1,876,665

    Across Wikipedia, only 28 languages contain over 100 million tokens and only 82 exceed 10 million tokens. The Common Crawl expands training volumes by factors between 47×47\times (German) and 125×125\times (Russian), providing broader lexical coverage for multilingual vector estimation.

  7. Knowl 7 — FastText Language Identification Classifier Benchmark

    data/table

    The fastText linear language identification model, trained for 176 languages using character nn-grams of lengths 2, 3, and 4 and hierarchical softmax, was benchmarked against langid.py across three public datasets:

    Model TCL Wikipedia EuroGov
    Acc. (%) Time (s) Acc. (%) Time (s) Acc. (%) Time (s)
    langid.py 93.1 8.8 91.3 9.4 98.7 13.1
    fastText 94.7 1.3 93.0 1.3 98.7 2.9

    The fastText language identifier matches or exceeds the classification accuracy of langid.py while providing a 4.5×4.5\times to 7.2×7.2\times reduction in processing time.

  8. Knowl 8 — Persistent Representation Quality Gap in Low-Resource Languages

    limitation

    Even when scaling data via Common Crawl (reaching ~1.89 billion tokens), word analogy accuracy for low-resource languages such as Hindi reaches only 32.1% (compared to 10.6% baseline), lagging substantially behind high-resource languages such as German (72.9%), French (73.6%), Italian (69.8%), and Chinese (78.4%). This indicates that gathering uncurated web text alone without domain-specific data curation or specialized low-resource modeling techniques is insufficient to close the embedding quality gap.

Coverage note — None was omitted; all key contributions including model formulation, preprocessing pipeline, language detector benchmarks, dataset adaptations, analogy results, corpus statistics, and identified limitations are fully covered.

References

  1. 1.Al-Rfou, R., Perozzi, B., and Skiena, S. (2013). Polyglot: Distributed word representations for multilingual nlp. Proc. CoNLL.
  2. 2.Baldwin, T. and Lui, M. (2010). Language identification: The long and the short of the matter. In Proc. NAACL.
  3. 3.Berardi, G., Esuli, A., and Marcheggiani, D. (2015). Word embeddings go to Italy: a comparison of models and training datasets. Italian Information Retrieval Workshop.
  4. 4.Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5.
  5. 5.Buck, C., Heafield, K., and Van Ooyen, B. (2014). N-gram counts and language models from the common crawl. In Proc. LREC, volume 2.
  6. 6.Cardellino, C. (2016). Spanish Billion Words Corpus and Embeddings, March.
  7. 7.Chang, P.-C., Galley, M., and Manning, C. D. (2008). Optimizing chinese word segmentation for machine translation performance. In Proceedings of the third workshop on statistical machine translation.
  8. 8.Collobert, R. and Weston, J. (2008). A unified architecture for natural language processing: Deep neural networks with multitask learning. In Proc. ICML.
  9. 9.Hartmann, N., Fonseca, E., Shulby, C., Treviso, M., Rodrigues, J., and Aluisio, S. (2017). Portuguese word embeddings: Evaluating on word analogies and natural language tasks. arXiv preprint arXiv:1708.06025.
  10. 10.Jin, P. and Wu, Y. (2012). Semeval-2012 task 4: evaluating chinese word similarity. In Proc. *SEM.
  11. 11.Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T. (2017). Bag of tricks for efficient text classification. In Proc. EACL.
  12. 12.Koehn, P. (2005). Europarl: A parallel corpus for statistical machine translation. In MT summit, volume 5.
  13. 13.Köper, M., Scheible, C., and im Walde, S. S. (2015). Multilingual reliability and ”semantic” structure of continuous word spaces. Proc. IWCS 2015.
  14. 14.Kudo, T. (2005). Mecab: Yet another part-of-speech and morphological analyzer. http://mecab. sourceforge. net/.
  15. 15.Lui, M. and Baldwin, T. (2012). langid.py: An off-theshelf language identification tool. In Proc. ACL (system demonstrations).
  16. 16.Mihalcea, R. (2007). Using wikipedia for automatic word sense disambiguation. In Proc. NAACL.
  17. 17.Mikolov, T., Chen, K., Corrado, G. D., and Dean, J. (2013a). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  18. 18.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013b). Distributed representations of words and phrases and their compositionality. In Adv. NIPS.
  19. 19.Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C., and Joulin, A. (2017). Advances in pre-training distributed word representations. arXiv preprint arXiv:1712.09405.
  20. 20.Mnih, A. and Kavukcuoglu, K. (2013). Learning word embeddings efficiently with noise-contrastive estimation. In Adv. NIPS.
  21. 21.Nguyen, T.-P. and Le, A.-C. (2016). A hybrid approach to vietnamese word segmentation. In Computing & Communication Technologies, Research, Innovation, and Vision for the Future (RIVF), 2016 IEEE RIVF International Conference on. IEEE.
  22. 22.Pennington, J., Socher, R., and Manning, C. (2014). Glove: Global vectors for word representation. In Proc. EMNLP.
  23. 23.Svoboda, L. and Brychcin, T. (2016). New word analogy corpus for exploring embeddings of Czech words. In Proc. CICLING.
  24. 24.Venekoski, V. and Vankka, J. (2017). Finnish resources for evaluating language model semantics. In Proc. NoDaLiDa.
  25. 25.Wu, F. and Weld, D. S. (2010). Open information extraction using wikipedia. In Proc. ACL.

Citation

MLA
Grave, E., et al. “Learning Word Vectors for 157 Languages”. arXiv, 2018, http://arxiv.org/abs/1802.06893v2.
APA
Grave, E., Bojanowski, P., Gupta, P., Joulin, A., & Mikolov, T. (2018). Learning Word Vectors for 157 Languages. arXiv. http://arxiv.org/abs/1802.06893v2
Chicago
Grave, E., P. Bojanowski, P. Gupta, A. Joulin, and T. Mikolov. 2018. “Learning Word Vectors for 157 Languages”. arXiv. http://arxiv.org/abs/1802.06893v2.
Harvard
Grave, E. et al. (2018) “Learning Word Vectors for 157 Languages”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1802.06893v2.
Vancouver
1. Grave E, Bojanowski P, Gupta P, Joulin A, Mikolov T (2018) Learning Word Vectors for 157 Languages. arXiv

BibTeX

@article{grave2018learning,
  title = {Learning Word Vectors for 157 Languages},
  author = {Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1802.06893v2},
  eprint = {1802.06893}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF