Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Ayyoob ImaniPeiqin LinAmir Hossein KargaranSilvia SeveriniMasoud Jalili SabetNora KassnerChunlan MaHelmut SchmidAndré F. T. MartinsFrançois Yvon

article2023ACL163 citationsArea Chair Award (Multilingualism and Cross-Lingual NLP)

Presents an open-source multilingual corpus and language model spanning over 500 predominantly low-resource languages, outperforming XLM-R across multiple natural language understanding benchmarks while identifying the key drivers of multilingual representation quality.

Listen

Natural language processing technologies have largely focused on making large language models deeper and more capable for roughly 100 well-resourced languages. This narrow scope excludes thousands of the world's languages, creating significant disparities in access to language technology across diverse global communities and digital ecosystems.

The article demonstrates that multilingual language models can be successfully scaled horizontally to support hundreds of under-resourced languages through curated corpus collection and continued pretraining.

To accomplish this, researchers collected a massive dataset covering 2,266 languages from approximately 150 diverse sources, including web crawls, translations, and religious texts. After applying multi-stage text cleaning and setting a threshold of at least 30,000 sentences per language-script, they created a 600-gigabyte training dataset covering 511 languages across 534 language-scripts. Using this data, they extended the vocabulary and continued the masked pretraining of a standard multilingual baseline model to create an expanded 395-million-parameter model, which was systematically evaluated across five diverse downstream understanding tasks alongside linguistic pseudoperplexity.

The resulting model demonstrated substantial performance improvements over standard baselines. On low-resource tail languages, sentence retrieval accuracy increased dramatically from under 10% to over 43% on aligned texts, and text classification accuracy improved from under 14% to over 46%. Performance on sequence labeling tasks like named entity recognition and part-of-speech tagging also increased by 13 to 21 percentage points for tail languages while maintaining or slightly improving representation quality for high-resource head languages. Furthermore, empirical analysis showed that model quality is driven by a combination of factors, including target corpus size, native script coverage, model capacity, and linguistic synergy from related neighboring languages.

These findings prove that organizations do not need to choose between supporting well-resourced languages and expanding into long-tail languages. Expanding language coverage enables positive cross-lingual transfer without requiring proportional increases in model size, drastically lowering the computational barrier to deploy equitable language technology in underserved regions.

Decision-makers and engineering teams should adopt horizontal scaling strategies and open datasets when developing multilingual systems. Future efforts should focus on training larger architectures, developing distilled lightweight models for cost-efficient deployment, and evaluating horizontal models on broader end-user applications.

The analysis is subject to some residual noise within web-crawled texts, unaddressed societal biases in training datasets, and potential variations arising from limited hyperparameter tuning. Nevertheless, the broad evaluation across hundreds of standardized language-scripts provides high confidence in horizontal scaling as an effective, scalable strategy for global language coverage.

arXiv: 2305.12182cisnlp/Glot500
Cover for Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Abstract

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An important part of this effort is to collect and clean Glot500-c, a corpus that covers these 511 languages and allows us to train Glot500-m. We evaluate Glot500-m on five diverse tasks across these languages. We observe large improvements for both high-resource and low-resource languages compared to an XLM-R baseline. Our analysis shows that no single factor explains the quality of multilingual LLM representations. Rather, a combination of factors determines quality including corpus size, script, “help” from related languages and the total capacity of the model. Our work addresses an important goal of NLP research: we should not limit NLP to a small fraction of the world’s languages and instead strive to support as many languages as possible to bring the benefits of NLP technology to all languages and cultures. Code, data and models are available at https://github.com/cisnlp/Glot500.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Glot500-c
  • 3.1 Data Collection
  • 3.2 Language-Scripts
  • 3.3 Ngram LMs and Language Divergence
  • 3.4 Data Cleaning
  • 3.5 Training Data: Glot500-c
  • 4 Glot500-m
  • 4.1 Vocabulary Extension
  • 4.2 Continued Pretraining
  • 5 Experimental Setup
  • 6 Experiments
  • 6.1 Results
  • 6.2 Language Coverage
  • 6.3 Training Progression
  • 6.4 Analysis across Language-Scripts
  • 6.5 Languages with Multiple Scripts
  • 6.6 Analysis across Language Families
  • 6.7 Effect of Amount of Training Data
  • 6.8 Support through Related Languages
  • 7 Conclusion and Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A N-grams LMs and Language Divergence
  • B Languages
  • C List of data sources
  • D Results for Each Task and Language
  • E Perplexity Results for all Languages

Knowls

  1. Knowl 1 — Glot500-m Model Architecture, Vocabulary Extension, and Training Setup

    model/method

    Glot500-m is a massively multilingual masked language model covering 511 languages (534 language-scripts), built by extending the vocabulary of XLM-R-Base (XLM-R-B) and performing continued pretraining on the Glot500-c corpus.

    To adapt the tokenizer without degrading existing representations, a SentencePiece unigram tokenizer is trained on Glot500-c targeting 250,000 tokens, sampling sentences across language-scripts according to a multinomial distribution with smoothing parameter α=0.3\alpha = 0.3. For head languages (languages already present in XLM-R), sampling is constrained to match the count of the lowest-resource tail languages. Merging these tokens into the 250,000-token XLM-R vocabulary yields approximately 100,000 overlapping tokens and adds 151,000 novel tokens, expanding the total vocabulary size to 401,000 tokens. This increases total model parameters from 278M (XLM-R-B) to 395M (Glot500-m), while preserving the original 86M non-embedding transformer parameters.

    Continued pretraining is conducted under the Masked Language Modeling (MLM) objective using the Adam optimizer with β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, an initial learning rate of 5×10−55 \times 10^{-5}, and batches of 384 sequence chunks (each 512 tokens long) sampled with the same multinomial distribution (α=0.3\alpha = 0.3). Training was executed on eight NVIDIA RTX A6000 GPUs for two weeks, with the final release checkpoint selected via early stopping on downstream task performance at 480,000 steps.

  2. Knowl 2 — Glot2000-c and Glot500-c Multilingual Corpora Collection and Cleaning

    model/method

    Glot2000-c is a corpus of over 700GB across 2,266 languages gathered from approximately 150 diverse data sources, including religious translations, news outlets, open web crawls, and scientific articles. Because individual languages may use different writing systems, script detection is performed per sentence, treating each (ISO 639-3, script) pair as a separate language-script entity.

    Data cleaning applies two stages of filtering:

    1. Chunk-level filters: Using filters adapted from BigScience ROOTS:

      • SF1 (Character repetition): Rejects segments with excessively repeated characters.
      • SF2 (Word repetition): Rejects segments with excessively repeated word tokens.
      • SF3 (Special characters): Rejects text dominated by crawling artifacts or code.
      • SF4 (Insufficient word count): Removes segments with too few words to provide context.
      • SF5 (Deduplication): Removes duplicate sentences after stripping whitespace and punctuation.
    2. Corpus-level filters:

      • CF1 (Script mismatch): Removes corpora where the detected script is inconsistent with the language identity (e.g., Chinese transcribed in Arabic script).
      • CF2 (Perplexity mismatch): Computes the 3-gram character language model divergence between the target corpus and all other language corpora. If the nearest neighbor belongs to a different typological family, the corpus is manually inspected and either removed or assigned the correct ISO code.

    Language-scripts possessing at least 30,000 cleaned sentences are selected to form Glot500-c, comprising 511 languages, 534 language-scripts, 1.5 billion sentences, and approximately 600GB of text. Each language-script is split into train, dev, and test partitions (reserving 1,000 sentences each for dev and test, including 500 parallel Bible verses where available, and dedicating the remainder to pretraining).

  3. Knowl 3 — Perplexity-Based Symmetrical Language Divergence Metric

    equation

    To measure the linguistic and statistical distance between language-scripts LiL_i and LjL_j, 3-gram character-level language models MLiM_{L_i} and MLjM_{L_j} are trained with interpolated modified Kneser-Ney smoothing on up to 100,000 sentences per language.

    The perplexity PP(S,M)PP(S, M) of a test character sequence S=(ch1,ch2,…,chT)S = (ch_1, ch_2, \dots, ch_T) under model MM is defined as:

    PP(S,M)=(∏t=1T1P(cht∣ch1t−1))1TPP(S, M) = \left( \prod_{t=1}^T \frac{1}{P(ch_t \mid ch_1^{t-1})} \right)^{\frac{1}{T}}

    where transition probabilities are estimated from training counts C(⋅)C(\cdot):

    P(cht∣ch1t−1)=C(ch1t−1cht)C(ch1t−1)P(ch_t \mid ch_1^{t-1}) = \frac{C(ch_1^{t-1} ch_t)}{C(ch_1^{t-1})}

    The symmetrical divergence DLi,LjD_{L_i, L_j} between language-scripts LiL_i and LjL_j evaluated over their respective test corpora SLiS_{L_i} and SLjS_{L_j} is defined as:

    DLi,Lj=max⁡(PP(SLi,MLj),PP(SLj,MLi))D_{L_i, L_j} = \max\left( PP(S_{L_i}, M_{L_j}), PP(S_{L_j}, M_{L_i}) \right)

    Taking the maximum guarantees symmetry (DLi,Lj=DLj,LiD_{L_i, L_j} = D_{L_j, L_i}) and prevents asymmetric bias caused by languages with simpler morphology or shorter average token lengths exhibiting artificially low perplexity in one direction.

  4. Knowl 4 — Benchmark Results of Glot500-m vs. XLM-R Baselines

    empirical result

    Glot500-m was evaluated against XLM-R-Base (XLM-R-B, 278M parameters) and XLM-R-Large (XLM-R-L, 560M parameters) across pseudoperplexity and five diverse downstream tasks averaged over 5 random seeds:

    Tail Languages Head Languages All Languages
    Task / Metric XLM-R-B XLM-R-L Glot500-m XLM-R-B XLM-R-L Glot500-m XLM-R-B XLM-R-L Glot500-m
    Pseudoperplexity ↓\downarrow 304.2 168.6 12.2 12.5 8.4 11.8 247.8 136.4 11.6
    SentRetr Tatoeba (Top-10 Acc) 32.6 33.6 59.8 66.2 71.1 75.0 56.6 60.4 70.7
    SentRetr Bible (Top-10 Acc) 7.4 7.1 43.2 54.2 58.3 59.0 19.3 20.1 47.3
    Text Classification (F1) 13.7 13.9 46.6 51.3 60.5 54.7 23.3 25.8 48.7
    NER (F1) 47.5 51.8 60.7 61.8 66.0 63.9 55.3 59.5 62.4
    POS (F1) 41.7 43.5 62.3 76.4 78.4 76.0 65.8 67.7 71.8
    Roundtrip Alignment (Acc) 2.6 3.1 4.5 3.4 4.1 5.5 2.8 3.3 4.7

    Glot500-m outperforms XLM-R-B by substantial margins on all tasks for tail languages (e.g., +35.8% top-10 accuracy on SentRetr Bible and +32.9% F1 on Text Classification). On head languages, Glot500-m improves upon XLM-R-B on Tatoeba retrieval (+8.8%), Bible retrieval (+4.8%), text classification (+3.4%), NER (+2.1%), and roundtrip alignment (+2.1%), while matching performance on POS tagging. Glot500-m also surpasses the much larger XLM-R-L on all tail language tasks, demonstrating that horizontal multilingual coverage can improve representations beyond pure parameter vertical scaling.

  5. Knowl 5 — Cross-Lingual Evaluation Suite for Massively Multilingual LLMs

    experimental setup

    To evaluate language representations across 511 languages where gold-standard labeled data is largely absent for tail languages, a mixed evaluation framework is employed:

    1. Pseudoperplexity (PPPL): Evaluated on held-out test data by masking tokens one-by-one and computing word-level normalized score distributions.
    2. Roundtrip Word Alignment: Evaluates cross-lingual consistency without human annotations. SimAlign subword alignments are extracted from Transformer layer 8 embeddings on parallel Bible verses using intersection symmetrization. A word ww in language L1L_1 is traced along bilingual alignment links across three random intermediate languages L2,L3,L4L_2, L_3, L_4 and back to L1L_1. Accuracy is the percentage of roundtrips that successfully return to ww, averaged over 5 independent random runs.
    3. Zero-Shot Cross-Lingual Transfer (English-Tuned):
      • Named Entity Recognition (NER): WikiANN dataset covering 89 head and 75 tail language-scripts.
      • Part-Of-Speech (POS) Tagging: Universal Dependencies v2.11 covering 63 head and 28 tail language-scripts.
      • Text Classification: Taxi1500 benchmark (6 classes, 90 head and 264 tail language-scripts), fine-tuned on English with learning rate 2×10−52 \times 10^{-5} and batch size 16.
    4. Sentence Retrieval (SentRetr): Evaluates top-10 nearest neighbor accuracy using cosine similarity between mean word embeddings from layer 8 on English-aligned sentence pairs from Tatoeba (up to 1,000 pairs, 70 head / 28 tail) and the parallel Bible test split (500 pairs, 94 head / 275 tail).
  6. Knowl 6 — Effect of Pretraining Corpus Size and Neighboring Language Support on Cross-Lingual Transfer

    empirical result

    The relationship between pretraining data volume and downstream zero-shot transfer performance was analyzed on the Bible Sentence Retrieval benchmark across head and tail languages.

    Correlating a language ll's own pretraining corpus size with its zero-shot retrieval performance yields a moderate correlation of Pearson's r=0.34r = 0.34. When correlating retrieval performance with the joint corpus size of language ll combined with its kk nearest neighbors (determined by character 3-gram language model divergence DD), the correlation increases to Pearson's r=0.44r = 0.44 for both k=3k = 3 and k=4k = 4.

    This demonstrates that zero-shot downstream representation quality for a given low-resource language is determined not only by its own corpus size, but also significantly by the aggregate data volume available across typologically and structurally related neighboring languages in the training mixture.

  7. Knowl 7 — Monolingual vs. Massively Multilingual Continued Pretraining and Related Language Support

    empirical result

    Continued pretraining on a single target language-script (denoted Glot+1) was compared against joint continued pretraining across all 511 languages in Glot500-c (Glot500-m) on the Sentence Retrieval Bible task (top-10 accuracy):

    Language-Script Language Name Glot+1 Glot500-m
    rug_Latn Roviana 51.0 49.0
    yan_Latn Mayangna/Sumo 46.4 31.8
    wbm_Latn Wa/Va 49.6 46.4
    ctd_Latn Tedim Chin 47.4 59.4
    quh_Latn Southern Quechua 33.4 56.2
    tat_Cyrl Tatar 58.8 67.2

    For linguistically isolated languages lacking close relatives in the training corpus (e.g., Roviana, Mayangna, Wa), monolingual adaptation (Glot+1) outperforms massively multilingual training because model capacity is focused entirely on the target language. Conversely, for languages with close relatives represented in Glot500-c (e.g., Southern Quechua supported by Cuzco Quechua quz_Latn, Tedim Chin, and Tatar), Glot500-m significantly outperforms Glot+1, proving that related language synergy compensates for limited individual data.

  8. Knowl 8 — Impact of Script Coverage on Multilingual Representation Learning

    empirical result

    Evaluating multi-script languages on Bible Sentence Retrieval (Top-10 Accuracy) illustrates the effect of script support in pretraining:

    Language Script Status in XLM-R XLM-R-B Glot500-m Gain
    Uighur (uig_Arab) Head (Supported) 45.8 56.2 +10.4
    Uighur (uig_Latn) Tail (Unsupported) 9.8 62.8 +53.0
    Hindi (hin_Deva) Head (Supported) 67.0 76.6 +9.6
    Hindi (hin_Latn) Tail (Unsupported) 13.6 43.2 +29.6
    Uzbek (uzb_Latn) Head (Supported) 54.8 67.6 +12.8
    Uzbek (uzb_Cyrl) Tail (Unsupported) 6.2 78.8 +72.6
    Kara-Kalpak (kaa_Cyrl) Tail (Unsupported) 17.6 73.8 +56.2
    Kara-Kalpak (kaa_Latn) Tail (Unsupported) 9.2 43.4 +34.2
    Northern Kurdish (kmr_Cyrl) Tail (Unsupported) 4.0 42.4 +38.4
    Northern Kurdish (kmr_Latn) Tail (Unsupported) 35.8 63.0 +27.2
    Turkmen (tuk_Cyrl) Tail (Unsupported) 13.6 65.0 +51.4
    Turkmen (tuk_Latn) Tail (Unsupported) 9.6 66.2 +56.6

    XLM-R fails on unsupported scripts because unseen characters are mapped directly to unknown (UNK) tokens. Expanding the tokenizer vocabulary and pretraining on these scripts enables Glot500-m to achieve major gains. Furthermore, for tail languages with multiple unsupported scripts (such as Kara-Kalpak), performance is substantially higher for the script with larger training data volume (\texttt{kaa_Cyrl} has 3x the training data of \texttt{kaa_Latn}, outperforming it by 30.4 percentage points).

  9. Knowl 9 — Typological Family Identification via Character N-Gram Divergence

    empirical result

    The ability of character 3-gram language model divergence DLi,LjD_{L_i, L_j} to identify linguistic relatedness was validated by testing whether the majority of a language's kk nearest neighbors belong to the same Ethnologue typological family group at hierarchy level ll (l=1l=1 denotes the highest-level family):

    Model Level ll Neighbors kk Accuracy (%)
    3-gram 1 1 84.45
    3-gram 1 3 75.77
    3-gram 1 7 69.08
    3-gram 1 13 62.75
    3-gram 1 21 55.33
    3-gram 2 1 79.75
    3-gram 2 3 67.63
    3-gram 2 7 59.49
    3-gram 2 13 51.36
    3-gram 2 21 42.68
    3-gram 3 1 75.05
    3-gram 3 3 60.22
    3-gram 3 7 49.55
    3-gram 3 13 38.34
    3-gram 3 21 29.84
    3-gram max 1 59.31
    3-gram max 3 36.89
    3-gram max 7 18.81
    3-gram max 13 6.87
    3-gram max 21 2.89

    Classification accuracy reaches 84.45% for l=1,k=1l=1, k=1 and progressively decreases as the required hierarchy depth ll increases and as the neighborhood size kk expands. Failure cases primarily occur when extensive language contact or historical loanword borrowing induces high character nn-gram overlap between typologically distinct families (e.g., Aymara and Quechua).

  10. Knowl 10 — Limitations of Glot500-m and Horizontal Scaling

    limitation

    The horizontally scaled Glot500 model and dataset present several explicit limitations:

    1. Fixed Model Capacity and Pseudoperplexity Trade-off: The underlying transformer architecture maintains the 86M parameter backbone of XLM-R-Base. Allocating this fixed capacity across 511 languages causes Glot500-m to underperform XLM-R-B in pseudoperplexity on 69 head languages (average PPPL of 11.8 vs. 8.4 for XLM-R-L and 12.5 for XLM-R-B), where head language training data was downsampled.
    2. Lack of Comprehensive Hyperparameter Optimization: Due to the high computational expense of pretraining across 511 languages, systematic hyperparameter sweeps were not performed.
    3. Domain Imbalance in Long-Tail Corpora: Tail language data is heavily dominated by religious texts (notably Bible translations and Jehovah's Witnesses publications), leading to potential domain distribution shift when applied to general or colloquial domains.
    4. Residual Corpus Noise: Despite multi-stage chunk-level and corpus-level filtering, residual misclassifications and noise persist across hundreds of low-resource languages.

Coverage note — None was omitted; exhaustive per-language result tables from the appendix (Tables 11-25) were synthesized into aggregate benchmark data, multi-script analyses, and ablation knowls.

References

  1. 1.Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem, Tewodros Abebe, Wondimagegnhue Tsegaye, Amanuel Lemma, Tsegaye Andargie, and Seifedin Shifaw. 2018. Parallel corpora for bi-lingual English-Ethiopian languages statistical machine translation. In Proceedings of the 27th International Conference on Computational Linguistics, pages 3102–3111, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  2. 2.Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, and Kareem Darwish. 2021. QADI: Arabic dialect identification in the wild. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 1–10, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  3. 3.Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018. Shami: A corpus of Levantine Arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  4. 4.Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Alcoba Inciarte. 2022. SERENGETI: Massively multilingual language models for Africa. arXiv preprint arXiv:2212.10785.
  5. 5.David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, and Sam Manthalu. 2022. A few thousand translations go a long way! leveraging pre-trained models for African news translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3053–3070, Seattle, United States. Association for Computational Linguistics.
  6. 6.David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, and Cristina España-Bonet. 2021. The effect of domain and diacritics in Yoruba–English neural machine translation. In Proceedings of Machine Translation Summit XVIII: Research Track, pages 61–75, Virtual. Association for Machine Translation in the Americas.
  7. 7.Rodrigo Agerri, Xavier Gómez Guinovart, German Rigau, and Miguel Anxo Solla Portela. 2018. Developing new linguistic resources and tools for the Galician language. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  8. 8.Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022. Adapting pretrained language models to African languages via multilingual adaptive fine-tuning. In Proceedings of the 29th International Conference on Computational Linguistics, pages 4336–4349, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  9. 9.Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018. DART: A large dataset of dialectal Arabic tweets. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  10. 10.Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur. 2020. TICO-19: the translation initiative for COvid-19. In Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020, Online. Association for Computational Linguistics.
  11. 11.Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić. 2022. Composable sparse fine-tuning for cross-lingual transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1778–1796, Dublin, Ireland. Association for Computational Linguistics.
  12. 12.Mikel Artetxe and Holger Schwenk. 2019. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics, 7:597–610.
  13. 13.Niyati Bafna. 2022. Empirical models for an indic language continuum.
  14. 14.Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020. ParaCrawl: Web-scale acquisition of parallel corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4555–4567, Online. Association for Computational Linguistics.
  15. 15.Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vít Suchomel, Antonio Toral, Tobias van der Werff, and Jaume Zaragoza. 2022. Macocu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages. In Proceedings of the 23rd Annual Conference of the European Association for Machine Translation, EAMT 2022, Ghent, Belgium, June 1-3, 2022, pages 301–302. European Association for Machine Translation.
  16. 16.Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, et al. 2022. Building machine translation systems for the next thousand languages. arXiv preprint arXiv:2205.03983.
  17. 17.Workshop BigScience, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, Dragomir Radev, Eduardo González Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady Elsahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jörg Frohberg, Joseph Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro Von Werra, Leon Weber, Long Phan, Loubna Ben allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, María Grandury, Mario Šaško, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto Luis López, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, Shayne Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Davut Emre Taşar, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Fevry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiangru Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Yallow Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre François Lavallée, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aurélie Névéol, Charles Lovering, Dan Garrette, Deepak Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, Shani Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdeněk Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ana Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ajibade, Bharat Saxena, Carlos Muñoz Ferrandis, Danish Contractor, David Lansky, Davis David, Douwe Kiela, Duong A. Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatima Mirza, Frankline Ononiwu, Habib Rezanejad, Hessie Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jesse Passmore, Josh Seltzer, Julio Bonis Sanz, Livia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nour Fahmy, Olanrewaju Samuel, Ran An, Rasmus Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas Wang, Sourav Roy, Sylvain Viguier, Thanh Le, Tobi Oyebade, Trieu Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Singh, Benjamin Beilharz, Bo Wang, Caio Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel León Periñán, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Imane Bello, Ishani Dash, Jihyun Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthik Rangasai Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, Maria A Castillo, Marianna Nezhurina, Mario Sänger, Matthias Samwald, Michael Cullan, Michael Weinberg, Michiel De Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patrick Haller, Ramya Chandrasekhar, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Bharati, Tanmay Laud, Théo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yash Shailesh Bajaj, Yash Venkatraman, Yifan Xu, Yingxin Xu, Yu Xu, Zhe Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, and Thomas Wolf. 2022. BLOOM: a 176b-parameter open-access multilingual language model.
  18. 18.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  19. 19.José Camacho-Collados, Claudio Delli Bovi, Alessandro Raganato, and Roberto Navigli. 2016. A large-scale multilingual disambiguation of glosses. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1701–1708, Portorož, Slovenia. European Language Resources Association (ELRA).
  20. 20.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3576–3588, Online. Association for Computational Linguistics.
  21. 21.Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2022. XLM-E: Cross-lingual language model pre-training via ELECTRA. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6170–6182, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.Rochelle Choenni and Ekaterina Shutova. 2022. Investigating language relationships in multilingual sentence encoders through the lens of linguistic typology. Computational Linguistics, 48(3):635–672.
  23. 23.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  24. 24.Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020. Improving multilingual models with language-clustered vocabularies. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4536–4546, Online. Association for Computational Linguistics.
  25. 25.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  26. 26.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
  27. 27.Marie-Catherine de Marneffe, Christopher D. Manning, Joakim Nivre, and Daniel Zeman. 2021. Universal dependencies. Computational Linguistics, 47(2):255–308.
  28. 28.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  29. 29.Philipp Dufter and Hinrich Schütze. 2020. Identifying elements essential for BERT’s multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4423–4437, Online. Association for Computational Linguistics.
  30. 30.Philipp Dufter, Mengjie Zhao, Martin Schmitt, Alexander Fraser, and Hinrich Schütze. 2018. Embedding learning through multilingual concept induction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1520–1530, Melbourne, Australia. Association for Computational Linguistics.
  31. 31.Jonathan Dunn. 2020. Mapping languages: the corpus of global language use. Lang. Resour. Evaluation, 54(4):999–1018.
  32. 32.Eberhard, David M., Gary F. Simons, and Charles D. Fennig (eds.). 2022. Ethnologue: Languages of the world. twenty-fifth edition.
  33. 33.Abteen Ebrahimi and Katharina Kann. 2021. How to adapt your pretrained multilingual model to 1600 languages. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4555–4567, Online. Association for Computational Linguistics.
  34. 34.Mahmoud El-Haj. 2020. Habibi - a multi dialect multi national Arabic song lyrics corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1318–1326, Marseille, France. European Language Resources Association.
  35. 35.Mahmoud El-Haj, Paul Rayson, and Mariam Aboelezz. 2018. Arabic dialect identification in the context of bivalency and code-switching. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  36. 36.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021. Beyond english-centric multilingual machine translation. J. Mach. Learn. Res., 22:107:1–107:48.
  37. 37.Pablo Gamallo, Jose Ramom Pichel, and Iñaki Alegria. 2017. A perplexity-based method for similar languages discrimination. In Proceedings of the Fourth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial), pages 109–114, Valencia, Spain. Association for Computational Linguistics.
  38. 38.Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012. Building large monolingual dictionaries at the leipzig corpora collection: From 100 to 200 languages. In Proceedings of the Eighth International Conference on Language Resources and Evaluation, LREC 2012, Istanbul, Turkey, May 23-25, 2012, pages 759–765. European Language Resources Association (ELRA).
  39. 39.Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2021. Experiments on a Guarani corpus of news and social media. In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas, pages 153–158, Online. Association for Computational Linguistics.
  40. 40.Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2022. Can we use word embeddings for enhancing Guarani-Spanish machine translation? In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 127–132, Dublin, Ireland. Association for Computational Linguistics.
  41. 41.Thamme Gowda, Zhao Zhang, Chris Mattmann, and Jonathan May. 2021. Many-to-English machine translation tools, data, and pretrained models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 306–316, Online. Association for Computational Linguistics.
  42. 42.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Samin Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. Xl-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 4693–4703. Association for Computational Linguistics.
  43. 43.Kenneth Heafield. 2011. KenLM: Faster and smaller language model queries. In Proceedings of the Sixth Workshop on Statistical Machine Translation, pages 187–197, Edinburgh, Scotland. Association for Computational Linguistics.
  44. 44.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 4411–4421. PMLR.
  45. 45.Ayyoob ImaniGooghari, Silvia Severini, Masoud Jalili Sabet, François Yvon, and Hinrich Schütze. 2022. Graph-based multilingual label propagation for low-resource part-of-speech tagging. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1577–1589, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  46. 46.Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020. SimAlign: High quality word alignments without parallel training data using static and contextualized embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1627–1643, Online. Association for Computational Linguistics.
  47. 47.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  48. 48.Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020. IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4948–4961, Online. Association for Computational Linguistics.
  49. 49.Fajri Koto and Ikhwan Koto. 2020. Towards computational linguistics in Minangkabau language: Studies on sentiment analysis and machine translation. In Proceedings of the 34th Pacific Asia Conference on Language, Information and Computation, pages 138–148, Hanoi, Vietnam. Association for Computational Linguistics.
  50. 50.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  51. 51.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75, Melbourne, Australia. Association for Computational Linguistics.
  52. 52.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  53. 53.Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018. The IIT Bombay English-Hindi parallel corpus. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  54. 54.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022. The BigScience ROOTS Corpus: A 1.6 TB Composite Multilingual Dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  55. 55.Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online. Association for Computational Linguistics.
  56. 56.Colin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera, Abraham Owodunni, and Daniel Whitenack. 2022. Bloom library: Multimodal datasets in 300+ languages for a variety of downstream tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 8608–8621. Association for Computational Linguistics.
  57. 57.Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Ehsaneddin Asgari, and Hinrich Schütze. 2023. Taxi1500: A multilingual dataset for text classification in 1500 languages.
  58. 58.Martin Majliš. 2011. W2C – web to corpus – corpora. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University.
  59. 59.Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, and Sriram Chellappan. 2021. A large-scale study of machine translation in Turkic languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5876–5890, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  60. 60.Steven Moran, Christian Bentz, Ximena Gutierrez-Vasques, Olga Pelloni, and Tanja Samardzic. 2022. TeDDi sample: Text data diversity sample for language comparison and multilingual NLP. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1150–1158, Marseille, France. European Language Resources Association.
  61. 61.Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2020. JParaCrawl: A large scale web-based English-Japanese parallel corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3603–3609, Marseille, France. European Language Resources Association.
  62. 62.Toshiaki Nakazawa, Hideya Mino, Isao Goto, Raj Dabre, Shohei Higashiyama, Shantipriya Parida, Anoop Kunchukuttan, Makoto Morishita, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, and Sadao Kurohashi. 2022. Overview of the 9th workshop on Asian translation. In Proceedings of the 9th Workshop on Asian Translation, pages 1–36, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
  63. 63.Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, and Sadao Kurohashi. 2021. Overview of the 8th workshop on Asian translation. In Proceedings of the 8th Workshop on Asian Translation (WAT2021), pages 1–45, Online. Association for Computational Linguistics.
  64. 64.Graham Neubig. 2011. The Kyoto free translation task. http://www.phontron.com/kftt.
  65. 65.Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021a. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 116–126, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  66. 66.Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021b. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 116–126.
  67. 67.Chester Palen-Michel, June Kim, and Constantine Lignos. 2022. Multilingual open text release 1: Public domain news in 44 languages. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 2080–2089, Marseille, France. European Language Resources Association.
  68. 68.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  69. 69.Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe. 2022. Lifting the curse of multilinguality by pre-training modular transformers. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3479–3495, Seattle, United States. Association for Computational Linguistics.
  70. 70.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021. UNKs everywhere: Adapting multilingual language models to new scripts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10186–10203, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  71. 71.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  72. 72.Roberts Rozis and Raivis Skadin,š. 2017. Tilde MODEL - multilingual open data for EU languages. In Proceedings of the 21st Nordic Conference on Computational Linguistics, pages 263–265, Gothenburg, Sweden. Association for Computational Linguistics.
  73. 73.Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020. AraBench: Benchmarking dialectal Arabic-English machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5094–5107, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  74. 74.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
  75. 75.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021. WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1351–1361, Online. Association for Computational Linguistics.
  76. 76.Silvia Severini, Ayyoob Imani, Philipp Dufter, and Hinrich Schütze. 2022. Towards a broad coverage named entity resource: A data-efficient approach for many diverse languages. arXiv preprint arXiv:2201.12219.
  77. 77.Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. 2022. Towards the next 1000 languages in multilingual machine translation: Exploring the synergy between supervised and self-supervised learning. arXiv preprint arXiv:2201.03110.
  78. 78.Anil Kumar Singh. 2008. Named entity recognition for south and south East Asian languages: Taking stock. In Proceedings of the IJCNLP-08 Workshop on Named Entity Recognition for South and South East Asian Languages.
  79. 79.Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7). Leibniz-Institut für Deutsche Sprache.
  80. 80.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey. European Language Resources Association (ELRA).
  81. 81.Iulia Turc, Kenton Lee, Jacob Eisenstein, Ming-Wei Chang, and Kristina Toutanova. 2021. Revisiting the primacy of english in zero-shot cross-lingual transfer. CoRR, abs/2106.16171.
  82. 82.Hai Wang, Dian Yu, Kai Sun, Jianshu Chen, and Dong Yu. 2019. Improving pre-trained multilingual model with vocabulary expansion. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 316–327, Hong Kong, China. Association for Computational Linguistics.
  83. 83.Mingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen, and Hinrich Schütze. 2023. NLNDE at semeval-2023 task 12: Adaptive pretraining and source language selection for low-resource multilingual sentiment analysis. CoRR, abs/2305.00090.
  84. 84.Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2022. Expanding pretrained models to thousands more languages via lexicon-based adaptation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 863–877, Dublin, Ireland. Association for Computational Linguistics.
  85. 85.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020a. Ccnet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020, pages 4003–4012. European Language Resources Association.
  86. 86.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020b. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4003–4012, Marseille, France. European Language Resources Association.
  87. 87.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  88. 88.Jian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu, Zhoujun Li, and Ming Zhou. 2020. Alternating language modeling for cross-lingual pre-training. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9386–9393.
  89. 89.Rodolfo Zevallos, John Ortega, William Chen, Richard Castro, Núria Bel, Cesar Toshio, Renzo Venturas, Hilario Aradiel, and Nelsi Melgarejo. 2022. Introducing QuBERT: A large monolingual corpus and BERT model for Southern Quechua. In Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing, pages 1–13, Hybrid. Association for Computational Linguistics.
  90. 90.Mengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi, and Hinrich Schütze. 2020. Masking as an efficient alternative to finetuning for pretrained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2226–2241, Online. Association for Computational Linguistics.

Citation

MLA
Imani, A., et al. “Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 1082–117, https://doi.org/10.18653/v1/2023.acl-long.61.
APA
Imani, A., Lin, P., Kargaran, A. H., Severini, S., Sabet, M. J., Kassner, N., Ma, C., Schmid, H., Martins, A. F. T., Yvon, F., & Schütze, H. (2023). Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1082–1117. https://doi.org/10.18653/v1/2023.acl-long.61
Chicago
Imani, A., P. Lin, A. H. Kargaran, et al. 2023. “Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1082–1117. https://doi.org/10.18653/v1/2023.acl-long.61.
Harvard
Imani, A. et al. (2023) “Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1082–1117. Available at: https://doi.org/10.18653/v1/2023.acl-long.61.
Vancouver
1. Imani A, Lin P, Kargaran AH, et al (2023) Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1082–1117

BibTeX

@inproceedings{imanigooghari-etal-2023-glot500,
    title = "Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages",
    author = {Imani, Ayyoob  and
      Lin, Peiqin  and
      Kargaran, Amir Hossein  and
      Severini, Silvia  and
      Jalili Sabet, Masoud  and
      Kassner, Nora  and
      Ma, Chunlan  and
      Schmid, Helmut  and
      Martins, Andr{\'e}  and
      Yvon, Fran{\c{c}}ois  and
      Sch{\"u}tze, Hinrich},
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.61/",
    doi = "10.18653/v1/2023.acl-long.61",
    pages = "1082--1117"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/