Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation

Xinyi WangSebastian RuderGraham Neubig

article2022ACL78 citations

Proposes a data augmentation framework using widely available bilingual lexicons to adapt multilingual pretrained models to under-represented languages with little to no existing text corpora.

Listen

Natural language processing models rely heavily on large volumes of text data. Because standard pretraining sources cover only a tiny fraction of the world's approximately 7,000 languages, around 85% of languages remain excluded from modern language technology due to having little or no available text. The article evaluates whether bilingual lexicons—word-to-word translation lists that cover roughly 70% of global languages—can effectively adapt pretrained language models to under-represented languages without relying on extensive text corpora.

To test this approach, the researchers synthesized training text by replacing English words with target-language translations using open-source bilingual dictionaries. They evaluated two primary adaptation strategies: pseudo-masked language modeling, which continues model pretraining on synthetic text, and pseudo translate-training, which translates task-specific training data. The study evaluated 19 under-represented languages across three language processing tasks—named entity recognition, part-of-speech tagging, and dependency parsing—under two scenarios: a setting with zero target-language text and a setting with a small amount of religious text.

The findings demonstrate that lexicon-based adaptation substantially improves model accuracy across all tasks. When no target text is available, adapting models using translated task data produces the strongest results, improving average performance across tasks from an initial 36.5 score to 45.2 points, with individual task gains reaching up to 15 points. In settings where a small amount of target text exists, pretraining on a combination of real and synthetic text proves most effective, raising average performance from 50.9 to 54.2 points. Furthermore, using a trained model to correct noisy synthetic task labels further improved performance on structured grammar tasks to 54.5 points. Attempts to generate synthetic data using neural machine translation systems performed worse than basic dictionary word replacement due to severe domain overfitting.

These results show that bilingual word lists provide a practical, computationally efficient mechanism to expand language technologies to thousands of historically excluded languages. Rather than requiring expensive, large-scale text collection or complex translation infrastructure, organizations can achieve meaningful cross-lingual capabilities using lightweight lexicon substitution. When selecting an adaptation method, teams should use translated task data if no target text exists, but switch to synthetic pretraining when minimal target text is available.

Decision-makers should prioritize the collection and digitization of word lists, which can be gathered rapidly compared to full text corpora, to scale language technology support. However, stakeholders should note that the article evaluated only languages using Latin script with English as the source language, and basic word replacement does not account for complex target-language grammar or word order. Additional validation is recommended before deploying these methods on non-Latin scripts or structurally dissimilar language pairs.

arXiv: 2203.09435
Cover for Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation

Abstract

The performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text present in a target language. Thus, the majority of the world’s languages cannot benefit from recent progress in NLP as they have no or limited textual data. To expand possibilities of using NLP technology in these under-represented languages, we systematically study strategies that relax the reliance on conventional language resources through the use of bilingual lexicons, an alternative resource with much better language coverage. We analyze different strategies to synthesize textual or labeled data using lexicons, and how this data can be combined with monolingual or parallel text when available. For 19 under-represented languages across 3 tasks, our methods lead to consistent improvements of up to 5 and 15 points with and without extra monolingual text respectively. Overall, our study highlights how NLP methods can be adapted to thousands more languages that are under-served by current technology.1

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Adaptation with Text
  • 2.2 Challenges with Low-resource Languages
  • 3 Adapting to Under-represented Languages Using Lexicons
  • 3.1 Synthesizing Data Using Lexicons
  • 3.2 Refining the Synthetic Data
  • 4 General Experimental Setting
  • 4.1 Tasks, Languages and Model
  • 4.2 Adaptation Data
  • 5 No-Text Setting
  • 5.1 Results
  • 6 Few-Text Setting
  • 6.1 Gold Data
  • 6.2 Adaptation Methods
  • 6.3 Results
  • 7 Analyses
  • 8 Related Work
  • 9 Conclusion and Discussion
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Experiment Details
  • A.2 NMTModels
  • A.3 Induced lexicons help languages with Fewer PanLex Entries
  • A.4 Effect of Task Data Size
  • A.5 List of Bilingual Lexicons
  • A.6 Lexicon Extraction
  • A.7 Performance for Individual Language

Knowls

  1. Knowl 1 — Lexicon-Based Pseudo Data Generation for Under-Represented Languages

    model/method

    To adapt multilingual pretrained language models to languages lacking substantial textual data, bilingual lexicons can be used to synthesize target language text through word-to-word replacement from a high-resource source language SS (such as English) to a target language TT.

    Given a bilingual lexicon DlexSTD^{ST}_{\text{lex}} mapping source words to target words:

    1. Pseudo Monolingual Data Generation (D~monoT\tilde{D}^T_{\text{mono}}): Given source monolingual sentences DmonoS={xiS}iD^S_{\text{mono}} = \{x^S_i\}_i, each sentence xiSx^S_i is converted to a synthetic target sentence x~iT\tilde{x}^T_i by replacing each word in xiSx^S_i with its translation in DlexSTD^{ST}_{\text{lex}}. Words without a lexicon entry are retained in their original source form, producing mixed-language pseudo sentences. If a source word has multiple candidate translations in the lexicon, one translation is sampled uniformly at random due to the lack of target text to estimate translation probabilities.

    2. Pseudo Labeled Data Generation (D~labelT\tilde{D}^T_{\text{label}}): Given labeled task data in the source language DlabelS={(xiS,yiS)}i=1ND^S_{\text{label}} = \{(x^S_i, y^S_i)\}_{i=1}^N with input text xiSx^S_i and task label yiSy^S_i, each text xiSx^S_i is transformed into x~iT\tilde{x}^T_i via word-to-word replacement using single-word lexicon entries in DlexSTD^{ST}_{\text{lex}}, while preserving the original annotation label yiSy^S_i, producing D~labelT={(x~iT,yiS)}i=1N\tilde{D}^T_{\text{label}} = \{(\tilde{x}^T_i, y^S_i)\}_{i=1}^N.

  2. Knowl 2 — Pretrained Model Adaptation via Pseudo Masked Language Modeling and Pseudo Translation Training

    model/method

    Using pseudo data synthesized from bilingual lexicons DlexSTD^{ST}_{\text{lex}}, a pretrained multilingual model MM is adapted to a target language TT under two resource settings:

    1. Pseudo Masked Language Modeling (Pseudo MLM): The pretrained multilingual model is continually pretrained using the masked language modeling objective on the pseudo monolingual corpus D~monoT\tilde{D}^T_{\text{mono}}. In a "Few-Text" setting where a small amount of authentic target monolingual data DmonoTD^T_{\text{mono}} is available, the model is trained jointly on D~monoT∪DmonoT\tilde{D}^T_{\text{mono}} \cup D^T_{\text{mono}} (with the gold target text upsampled to match the synthetic data size).

    2. Pseudo Translation Training (Pseudo Trans-train): The model is fine-tuned for a downstream task on the concatenation of the original source labeled data and the pseudo labeled data, DlabelS∪D~labelTD^S_{\text{label}} \cup \tilde{D}^T_{\text{label}}.

    3. Combined Adaptation (Both): The model first undergoes continued pretraining via Pseudo MLM on D~monoT\tilde{D}^T_{\text{mono}} (or D~monoT∪DmonoT\tilde{D}^T_{\text{mono}} \cup D^T_{\text{mono}}), and is subsequently fine-tuned on the combined task dataset DlabelS∪D~labelTD^S_{\text{label}} \cup \tilde{D}^T_{\text{label}}.

  3. Knowl 3 — Synthetic Data Refinement via Teacher Label Distillation and Parallel Lexicon Induction

    model/method

    Synthetic data created via word-to-word replacement can introduce label noise or suffer from limited vocabulary coverage. Two refinement strategies address these issues:

    1. Teacher Label Distillation: Word-to-word translation can alter sentence semantics or token properties such that the original source label yiSy^S_i no longer matches the pseudo input x~iT\tilde{x}^T_i. A teacher model is trained by adapting a pretrained multilingual model via Pseudo MLM and fine-tuning it solely on the source labeled data DlabelSD^S_{\text{label}}. This teacher model predicts new labels y~iT\tilde{y}^T_i for all pseudo examples x~iT\tilde{x}^T_i, yielding a distilled dataset D~distillT={(x~iT,y~iT)}i=1N\tilde{D}^T_{\text{distill}} = \{(\tilde{x}^T_i, \tilde{y}^T_i)\}_{i=1}^N. The final task model is fine-tuned on D~distillT∪DlabelS\tilde{D}^T_{\text{distill}} \cup D^S_{\text{label}}.

    2. Lexicon Induction from Parallel Data: When a small parallel corpus DparST={(xiS,xiT)}iD^{ST}_{\text{par}} = \{(x^S_i, x^T_i)\}_i (such as Bible text) is accessible in a Few-Text regime, unsupervised word alignment (e.g., using eflomal) is executed on the parallel sentence pairs. Aligned word pairs appearing more than once are extracted into an induced lexicon D~lexST\tilde{D}^{ST}_{\text{lex}} (typically adding ≈2k\approx 2\text{k} entries per language). This induced lexicon is merged with the base lexicon to form D~lexST∪DlexST\tilde{D}^{ST}_{\text{lex}} \cup D^{ST}_{\text{lex}}, which is then used for pseudo data generation.

  4. Knowl 4 — Performance of Lexicon Adaptation Across 19 Under-Represented Languages and Three NLP Tasks

    data/table

    The effectiveness of lexicon-based adaptation methods was evaluated on 19 under-represented languages across Named Entity Recognition (WikiAnn and MasakhaNER), Part-of-Speech tagging (UDPOS 2.5), and Dependency Parsing (UD 2.5, evaluated by labeled attachment score F1). English served as the source language, and multilingual BERT (mBERT) served as the base model. PanLex served as the primary lexicon source.

    Method Lexicon WikiNER Δ\Delta MasakhaNER Δ\Delta POS Δ\Delta Parsing Δ\Delta Avg. Δ\Delta
    No-Text Setting
    mBERT baseline - 47.6 - 46.1 - 36.1 - 16.5 - 36.5
    Pseudo Trans-train PanLex 49.8 +2.2 54.4 +8.3 51.1 +15.0 25.9 +9.4 45.2* (+8.7)
    Pseudo MLM PanLex 49.8 +2.2 52.6 +6.5 48.9 +12.8 25.2 +8.7 44.1* (+7.6)
    Both PanLex 48.5 +0.9 54.6 +8.5 48.7 +12.6 25.9 +9.4 44.4* (+7.9)
    Both+Label Distill. PanLex 50.6 +2.1 53.5 -1.1 50.3 +1.6 26.0 +0.1 45.1* (+0.7)
    Few-Text Setting
    Gold MLM baseline - 49.5 - 53.6 - 60.6 - 40.2 - 50.9
    Pseudo Trans-train PanLex 50.2 +0.7 59.4 +5.8 59.3 -1.3 37.0 -3.2 51.4 (+0.5)
    Pseudo MLM PanLex 50.7 +1.2 57.4 +3.8 65.4 +4.8 43.5 +3.3 54.2* (+3.3)
    Pseudo MLM PanLex+Induced 52.2 +1.5 58.5 +0.9 64.7 -0.7 41.5 -2.0 54.2* (0.0)
    Both PanLex 50.1 +0.6 59.2 +5.6 60.7 +0.1 38.3 -1.9 52.0* (+1.1)
    Both PanLex+Induced 52.6 +2.5 61.1 +1.9 59.5 -1.2 35.3 -3.0 52.0† (0.0)
    Both+Label Distill. PanLex 51.7 +1.6 58.4 -0.8 66.2 +5.5 41.9 +3.6 54.5* (+2.5)
    Both+Label Distill. PanLex+Induced 53.2 +1.5 59.4 +1.0 65.8 -0.4 40.7 -1.2 54.7* (+0.2)
    • indicates significant gains over baselines with p<0.001p < 0.001 and \dag with p<0.05p < 0.05 under paired bootstrap resampling.
  5. Knowl 5 — Optimal Adaptation Strategy Divergence Between No-Text and Few-Text Regimes

    empirical result

    The optimal strategy for lexicon-based adaptation differs fundamentally depending on the presence of authentic target monolingual text:

    1. No-Text Setting: When zero monolingual target text is available, Pseudo Trans-train is the single best adaptation method, achieving an overall average F1 score of 45.2 (an 8.7-point improvement over the 36.5 mBERT baseline across 19 languages). Stacking Pseudo Trans-train on top of Pseudo MLM does not yield meaningful additional gains over Pseudo Trans-train alone.

    2. Few-Text Setting: When a small target monolingual corpus (such as 5,000–8,000 Bible verses) is available, Pseudo MLM (pretraining on both gold and pseudo target monolingual data) becomes the dominant strategy, outperforming the Gold MLM baseline from 50.9 to 54.2 average F1. In contrast, Pseudo Trans-train alone yields only 51.4 average F1 and causes performance degradation on syntactic tasks (dropping POS from 60.6 to 59.3, and Parsing from 40.2 to 37.0). Stacking Pseudo Trans-train on top of Pseudo MLM in Few-Text similarly harms syntactic performance unless Label Distillation is applied.

  6. Knowl 6 — Failure of Domain-Overfit NMT Translation Relative to Lexicon Substitution for Data Synthesis

    data/table

    To evaluate whether Neural Machine Translation (NMT) can generate more natural synthetic text than word-to-word lexicon substitution, an open-source many-to-many 175M parameter NMT model was fine-tuned for 50 epochs on English-target Bible parallel data, then used to translate English Wikipedia sentences into the target languages for Pseudo MLM adaptation in the Few-Text setting.

    Synthesis Method WikiNER MasakhaNER POS Parsing
    Lexicon (Word-to-Word) 45.0 56.0 63.7 40.7
    NMT Translation 42.2 55.8 58.9 37.7

    NMT translation performed consistently worse than word-to-word lexicon substitution across all four datasets. Inspection of the NMT outputs revealed that the translation model overfit to the archaic Bible training domain, generating repetitive Bible-specific words and phrases when translating out-of-domain Wikipedia sentences, thereby degrading downstream representation quality.

  7. Knowl 7 — Syntactic Degradation Caused by Verb Over-Representation in Bible-Induced Lexicons

    empirical result

    Augmenting static lexicons (PanLex) with lexicons induced from unsupervised word alignments on Bible parallel data produces contrasting task effects:

    1. Semantic Tasks (NER): Adding induced lexicons improves WikiNER F1 from 50.7 to 52.2 and MasakhaNER F1 from 57.4 to 58.5 under Pseudo MLM.

    2. Syntactic Tasks (POS Tagging and Parsing): Adding induced lexicons lowers POS tagging F1 from 65.4 to 64.7 and Parsing F1 from 43.5 to 41.5.

    Part-of-speech analysis reveals that PanLex entries are heavily skewed towards nouns, whereas Bible-induced lexicons contain a substantially higher proportion of verbs. Because word-to-word replacement ignores target language word order and morphology, replacing source verbs directly into target positions disrupts verb placement and clause structure. This causes an accuracy drop specifically on verbs during POS tagging, harming syntax-sensitive tasks while leaving entity identification unaffected.

  8. Knowl 8 — Cross-Architecture Generalization to XLM-RoBERTa on Under-Represented Languages

    data/table

    The lexicon adaptation methodology generalizes across transformer architectures. Evaluating POS tagging in the Few-Text setting using XLM-RoBERTa base (XLM-R base) across Bambara (bam), Manx (glv), Maltese (mlt), and Erzya (myv) demonstrated consistent improvements over both the Gold MLM baseline and prior published XLM-R large results.

    Method bam glv mlt myv
    Gold MLM (XLM-R base, this work) 59.7 64.1 58.5 70.6
    Prior Work (XLM-R large, Ebrahimi and Kann 2021) 60.5 59.7 59.6 66.6
    + Pseudo Trans-train 57.4 63.2 69.1 63.8
    + Pseudo MLM 68.5 67.5 72.3 73.8
    + Both 60.3 64.5 69.3 65.9
    + Both (Label Distillation) 69.4 68.8 72.1 74.3

    Pseudo MLM and Pseudo MLM + Label Distillation achieved substantial F1 improvements on all four languages (up to +9.7 points on bam and +13.6 points on mlt over Gold MLM), validating that the benefits of lexicon-based adaptation are not specific to mBERT.

  9. Knowl 9 — Zero-Shot Lexicon Adaptation versus Target Language Few-Shot Learning

    data/table

    Zero-shot cross-lingual transfer with lexicon adaptation was compared against few-shot target fine-tuning (kk-shot) on the MasakhaNER benchmark across Hausa (hau), Wolof (wol), Luganda (lug), Igbo (ibo), Kinyarwanda (kin), and Luo (luo).

    Method hau wol lug ibo kin luo
    mBERT (Zero-Shot Baseline) 48.7 33.9 50.9 55.2 52.4 35.3
    Best Adapted (Zero-Shot, Ours) 74.4 60.3 61.6 63.6 63.8 42.6
    10-shot Fine-Tuning 44.5 49.1 52.7 56.2 51.2 46.2
    100-shot Fine-Tuning 64.0 56.9 58.3 65.5 55.7 51.6
    Best Adapted + 100-shot 76.1 57.3 61.3 63.2 62.6 49.4

    Zero-shot adaptation without any target annotations (Best Adapted) outperformed 10-shot fine-tuning across almost all languages (e.g., 74.4 vs 44.5 F1 on hau). Access to 100 annotated examples was required for few-shot learning to approach or slightly surpass the zero-shot adapted model, while combining the adapted model with 100-shot data yielded mixed results.

  10. Knowl 10 — Methodological Limitations of Lexicon-Based Adaptation

    limitation

    The lexicon-based adaptation framework has several stated limitations:

    1. Script and Source Constraint: Experiments were restricted to Latin-script under-represented target languages and used only English as the source language, omitting cross-script adaptation mechanisms such as transliteration or vocabulary expansion.

    2. Lack of Morphological and Syntactic Modeling: Word-to-word substitution ignores morphological inflection and differences in word order between source and target languages. Morphological analyzers and inflectors could not be incorporated due to near-total absence of morphological toolkits for the targeted low-resource languages.

    3. Single-Word Entry Restriction: Only single-word entries in bilingual dictionaries were substituted during data synthesis, ignoring multi-word expressions and phrasal equivalents.

    4. Lexicon Sparsity and Quality: Although PanLex indexes thousands of languages, many language entries contain sparse wordlists (some under 1,000 entries), limiting the lexical coverage attainable through automated substitution.

Coverage note — None was omitted; all contributed data synthesis techniques, adaptation strategies, distillation/induction methods, primary empirical benchmarks (No-Text, Few-Text, XLM-R, NMT synthesis, few-shot), and stated limitations are fully represented.

References

  1. 1.David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021. Masakhaner: Named entity recognition for african languages. In TACL.
  2. 2.Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A Smith. 2016. Massively multilingual word embeddings. arXiv preprint arXiv:1602.01925.
  3. 3.Antonios Anastasopoulos and Graham Neubig. 2019. Pushing the limits of low-resource morphological inflection. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China. Association for Computational Linguistics.
  4. 4.Steven Bird. 2020. Decolonising speech and language technology. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504–3519, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  5. 5.Damián Blasi, Antonios Anastasopoulos, and Graham Neubig. 2021. Systematic inequalities in language technology performance across the world's languages. arXiv preprint arXiv:2110.06733.
  6. 6.Brenda Boerger. 2017. Rapid word collection, dictionary production, and community well-being.
  7. 7.Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In COLING.
  8. 8.Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020. Parsing with multilingual BERT, a small corpus, and a small treebank. In Findings of EMNLP 2020.
  9. 9.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In ACL.
  10. 10.Alexis Conneau and Guillaume Lample. 2019. Crosslingual language model pretraining. In NeurIPS.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL.
  12. 12.Long Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird, and Trevor Cohn. 2016. Learning crosslingual word embeddings without bilingual corpora. arXiv preprint arXiv:1606.09403.
  13. 13.Abteen Ebrahimi and Katharina Kann. 2021. How to adapt your pretrained multilingual model to 1600 languages. In ACL, Online. Association for Computational Linguistics.
  14. 14.Jost Gippert, Nikolaus Himmelmann, Ulrike Mosel, et al. 2006. Essentials of language documentation. Mouton de Gruyter Berlín.
  15. 15.Stephan Gouws and Anders Søgaard. 2015. Simple task-specific bilingual word embeddings. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1386–1390, Denver, Colorado. Association for Computational Linguistics.
  16. 16.Emily Kennedy Gref. 2016. Publishing in North American Indigenous Languages. Ph.D. thesis, University of London.
  17. 17.Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don't stop pretraining: Adapt language models to domains and tasks. In ACL, Online.
  18. 18.Jeremy Howard and Sebastian Ruder. 2018. Universal Language Model Fine-tuning for Text Classification. In Proceedings of ACL 2018.
  19. 19.Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, and Graham Neubig. 2021. Explicit Alignment Objectives for Multilingual Bidirectional Encoders. In Proceedings of NAACL 2021.
  20. 20.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. In ICML.
  21. 21.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In ACL, Online. Association for Computational Linguistics.
  22. 22.Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha Talukdar, and Sunita Sarawagi. 2021. Exploiting language relatedness for low web-resource language model adaptation: An Indic languages study. In ACL, Online. Association for Computational Linguistics.
  23. 23.Dan Kondratyuk and Milan Straka. 2019. 75 languages, 1 model: Parsing universal dependencies universally. In EMNLP, Hong Kong, China.
  24. 24.Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In EMNLP, Online. Association for Computational Linguistics.
  25. 25.Stephen Mayhew, Chen-Tse Tsai, and Dan Roth. 2017. Cheap translation for cross-lingual named entity recognition. In EMNLP, Copenhagen, Denmark. Association for Computational Linguistics.
  26. 26.Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky. 2020. The Johns Hopkins University Bible corpus: 1600+ tongues for typological exploration. In LREC, pages 2884–2892, Marseille, France. European Language Resources Association.
  27. 27.Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013. Exploiting similarities among languages for machine translation. arXiv preprint arXiv:1309.4168.
  28. 28.Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021. When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models. In NAACL, Online.
  29. 29.Joakim Nivre, Mitchell Abrams, Željko Agić, Lars Ahrenberg, Lene Antonsen, Maria Jesus Aranzabe, Gashaw Arutie, Masayuki Asahara, Luma Ateyah, Mohammed Attia, et al. 2018. Universal dependencies 2.2.
  30. 30.Robert Östling and Jörg Tiedemann. 2016. Efficient word alignment with Markov Chain Monte Carlo. Prague Bulletin of Mathematical Linguistics, 106:125–146.
  31. 31.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations.
  32. 32.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Crosslingual name tagging and linking for 282 languages. In ACL, pages 1946–1958, Vancouver, Canada. ACL.
  33. 33.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In EMNLP, Online. Association for Computational Linguistics.
  34. 34.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021. UNKs Everywhere: Adapting Multilingual Language Models to New Scripts. In Proceedings of EMNLP 2021.
  35. 35.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In ACL, Florence, Italy.
  36. 36.Edoardo Maria Ponti, Julia Kreutzer, Ivan Vulić, and Siva Reddy. 2021. Modelling latent translations for cross-lingual transfer. In Arxiv.
  37. 37.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A Python natural language processing toolkit for many human languages. In ACL.
  38. 38.Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019. Massively multilingual transfer for NER. In ACL, pages 151–164, Florence, Italy. Association for Computational Linguistics.
  39. 39.Machel Reid and Mikel Artetxe. 2021. PARADISE: Exploiting Parallel Data for Multilingual Sequence-to-Sequence Pretraining. arXiv preprint arXiv:2108.01887.
  40. 40.Shruti Rijhwani, Antonios Anastasopoulos, and Graham Neubig. 2020. OCR Post Correction for Endangered Language Texts. In EMNLP, Online. Association for Computational Linguistics.
  41. 41.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation. In Proceedings of EMNLP 2021.
  42. 42.Sebastian Ruder, Ivan Vulić, and Anders Søgaard. 2019. A survey of cross-lingual word embedding models. Journal of Artificial Intelligence Research, 65:569–631.
  43. 43.Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending multilingual BERT to low-resource languages. In EMNLP-Findings, Online. Association for Computational Linguistics.
  44. 44.Shijie Wu and Mark Dredze. 2019. Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT. In EMNLP.

Citation

MLA
Wang, X., et al. “Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 863–77, https://doi.org/10.18653/v1/2022.acl-long.61.
APA
Wang, X., Ruder, S., & Neubig, G. (2022). Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 863–877. https://doi.org/10.18653/v1/2022.acl-long.61
Chicago
Wang, X., S. Ruder, and G. Neubig. 2022. “Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 863–77. https://doi.org/10.18653/v1/2022.acl-long.61.
Harvard
Wang, X., Ruder, S. and Neubig, G. (2022) “Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 863–877. Available at: https://doi.org/10.18653/v1/2022.acl-long.61.
Vancouver
1. Wang X, Ruder S, Neubig G (2022) Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 863–877

BibTeX

@inproceedings{wang-etal-2022-expanding,
    title = "Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation",
    author = "Wang, Xinyi  and
      Ruder, Sebastian  and
      Neubig, Graham",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.61/",
    doi = "10.18653/v1/2022.acl-long.61",
    pages = "863--877"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/