Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation
Xinyi WangSebastian RuderGraham Neubig
Proposes a data augmentation framework using widely available bilingual lexicons to adapt multilingual pretrained models to under-represented languages with little to no existing text corpora.
Natural language processing models rely heavily on large volumes of text data. Because standard pretraining sources cover only a tiny fraction of the world's approximately 7,000 languages, around 85% of languages remain excluded from modern language technology due to having little or no available text. The article evaluates whether bilingual lexicons—word-to-word translation lists that cover roughly 70% of global languages—can effectively adapt pretrained language models to under-represented languages without relying on extensive text corpora.
To test this approach, the researchers synthesized training text by replacing English words with target-language translations using open-source bilingual dictionaries. They evaluated two primary adaptation strategies: pseudo-masked language modeling, which continues model pretraining on synthetic text, and pseudo translate-training, which translates task-specific training data. The study evaluated 19 under-represented languages across three language processing tasks—named entity recognition, part-of-speech tagging, and dependency parsing—under two scenarios: a setting with zero target-language text and a setting with a small amount of religious text.
The findings demonstrate that lexicon-based adaptation substantially improves model accuracy across all tasks. When no target text is available, adapting models using translated task data produces the strongest results, improving average performance across tasks from an initial 36.5 score to 45.2 points, with individual task gains reaching up to 15 points. In settings where a small amount of target text exists, pretraining on a combination of real and synthetic text proves most effective, raising average performance from 50.9 to 54.2 points. Furthermore, using a trained model to correct noisy synthetic task labels further improved performance on structured grammar tasks to 54.5 points. Attempts to generate synthetic data using neural machine translation systems performed worse than basic dictionary word replacement due to severe domain overfitting.
These results show that bilingual word lists provide a practical, computationally efficient mechanism to expand language technologies to thousands of historically excluded languages. Rather than requiring expensive, large-scale text collection or complex translation infrastructure, organizations can achieve meaningful cross-lingual capabilities using lightweight lexicon substitution. When selecting an adaptation method, teams should use translated task data if no target text exists, but switch to synthetic pretraining when minimal target text is available.
Decision-makers should prioritize the collection and digitization of word lists, which can be gathered rapidly compared to full text corpora, to scale language technology support. However, stakeholders should note that the article evaluated only languages using Latin script with English as the source language, and basic word replacement does not account for complex target-language grammar or word order. Additional validation is recommended before deploying these methods on non-Latin scripts or structurally dissimilar language pairs.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). It documents the acute digital resource divide across global languages, providing the foundational taxonomy and motivation for expanding NLP beyond data-rich languages.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). It provides the foundational framework and baseline for multilingual masked language modeling (XLM-R) that the source adapts to under-represented languages.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). It introduces cross-lingual language model pretraining objectives that undergird subsequent pseudo-pretraining adaptation methods.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). It analyzes the baseline zero-shot cross-lingual transfer capabilities and limitations of multilingual pre-trained transformers on tasks like POS tagging and NER.
- Paper: Exploiting Similarities among Languages for Machine Translation, Tomas Mikolov et al. (2013). It establishes early principles for leveraging bilingual word-level mappings to transfer knowledge across languages without full sentence-parallel corpora.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). It formalizes cross-lingual evaluation benchmarks and standard translate-train methodologies that the source replaces with pseudo-translation techniques.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). It scales horizontal coverage to over 500 languages by combining corpus curation with continued multilingual pretraining.
- Paper: Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages, Wietse de Vries et al. (2022). It provides an extensive empirical evaluation of the linguistic and structural factors governing cross-lingual transfer across more than 100 languages.
- Paper: MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER, Ran Zhou et al. (2022). It extends synthetic multilingual data adaptation strategies by using masked entity modeling to prevent label misalignment in low-resource sequence tagging.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). It benchmarks modern few-shot cross-lingual transfer across 54 languages, examining how fine-tuning compares with in-context adaptation.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). It evaluates how contemporary large foundation models generalize across 83 diverse languages, including low-resource Indic and African language families.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It systematically compares specialized multilingual fine-tuned models against generative LLMs across 70 typologically diverse languages.
