Lifting the Curse of Multilinguality by Pre-training Modular Transformers

Jonas PfeifferNaman GoyalXi Victoria LinXian LiJames CrossSebastian RiedelMikel Artetxe

article2022NAACL187 citations

Proposes pre-training multilingual transformers with dedicated language-specific modules from the start, preventing capacity dilution across languages and enabling post-hoc extension to new languages without performance degradation.

Listen

Multilingual artificial intelligence models often face a fundamental trade-off known as the curse of multilinguality. As standard language models are trained to cover a broader set of languages, their performance on individual languages tends to degrade due to negative interference and competition for limited model capacity. This presents a critical challenge for global organizations seeking scalable language technologies that support diverse languages without sacrificing accuracy or incurring prohibitive computational costs.

The article evaluates whether pre-training models with modular, language-specific components can overcome this performance degradation. Specifically, it demonstrates the effectiveness of Cross-lingual Modular (X-MOD) architectures in preventing negative interference, enabling positive cross-lingual transfer, and supporting post-training expansion to previously unseen languages without losing accuracy.

The authors conducted large-scale pre-training experiments using the CC100 dataset across language sets of 13, 30, 60, and 75 typologically diverse languages. The baseline fully shared architecture was compared against the proposed modular architecture, which combines shared core network weights with specialized, language-specific bottleneck modules. Importantly, the modular model allocates unique modules to individual languages while keeping the active parameter count and computational cost during training and inference identical to the baseline. The models were evaluated on standard downstream benchmarks, including natural language inference, named entity recognition, and question answering, under conditions controlling for both training update steps and per-language data exposure.

The findings show that modular pre-training successfully eliminates the curse of multilinguality. First, while fully shared models suffered substantial performance drops as language count scaled, the modular model maintained and even improved accuracy across both high- and low-resource languages. Second, the modular model trained on 60 languages consistently outperformed the shared baseline across all downstream tasks, achieving an overall accuracy of 73.5% versus 72.5% on natural language inference and 62.8% versus 58.8% F1 score on named entity recognition. Third, when expanding to new, held-out languages after pre-training, the modular approach matched the accuracy of pre-training on those languages from the start, regardless of whether related languages existed in the initial training set. Finally, adding modular components after pre-training failed to recover performance, demonstrating that modularity must be integrated from the beginning.

These results demonstrate that organizations can scale multilingual models to dozens of languages without sacrificing per-language performance. Because language-specific modules can be routed dynamically and stored efficiently, operational costs and computational throughput remain stable as new languages are added. Furthermore, eliminating the need to retrain massive core models from scratch to support new markets significantly reduces deployment timelines, compute budgets, and maintenance overhead.

Technical leaders and practitioners should adopt modular pre-training architectures when building multilingual systems that require broad or evolving language coverage. Initial pre-training should focus on a representative, medium-to-large set of higher-resource languages, while lower-resource or emerging target languages can be added post-hoc with dedicated modules. While modular pre-training requires slightly longer initial training for modular benefits to fully materialize in smaller language pools, it offers a scalable, future-proof framework for global language coverage.

arXiv: 2205.06266
  • Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R’s large-scale multilingual pretraining experiments expose the capacity and language-interference trade-offs that X-MOD is designed to address.
  • Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This earlier account of shared multilingual pretraining objectives and cross-lingual transfer provides the baseline framework that the source modifies with language-specific modules.
  • Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Its analysis of multilingual BERT’s shared-parameter cross-lingual transfer helps clarify the interference problem that motivates the source’s modular alternative.

No sufficiently relevant recommendations were found.

Cover for Lifting the Curse of Multilinguality by Pre-training Modular Transformers

Abstract

Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific modules, which allows us to grow the total capacity of the model, while keeping the total number of trainable parameters per language constant. In contrast with prior work that learns language-specific components post-hoc, we pre-train the modules of our Cross-lingual Modular (X-Mod) models from the start. Our experiments on natural language inference, named entity recognition and question answering show that our approach not only mitigates the negative interference between languages, but also enables positive transfer, resulting in improved monolingual and cross-lingual performance. Furthermore, our approach enables adding languages post-hoc with no measurable drop in performance, no longer limiting the model usage to the set of pre-trained languages.

Table of Contents

  • 1 Introduction
  • 2 Background and related work
  • 2.1 Multilingual transformers
  • 2.2 Modular language models
  • 2.3 Weaknesses, improvements, and extensions of language models
  • 3 Proposed approach
  • 4 Experimental design
  • 4.1 Model variants
  • 4.2 Training details
  • 4.3 Evaluation
  • 5 Results and discussion
  • 5.1 Pre-trained languages
  • 5.2 Extending to unseen languages
  • 6 Further analysis
  • 6.1 The importance of update steps
  • 6.2 X-Mod vs. Adapters
  • 7 Conclusions
  • References
  • A Additional results
  • B Intermediate checkpoints
  • C Language selection

Knowls

  1. Knowl 1 — Cross-Lingual Modular Transformer (X-MOD) Architecture

    model/method

    The Cross-lingual Modular (X-MOD) architecture extends a standard Transformer encoder (12 layers, hidden dimension 768) by integrating language-specific modular components directly into every layer during pre-training.

    At each transformer layer, the Multi-Head Attention, primary Feed-Forward Network (FFN), and LayerNorm components are shared across all pre-training languages, comprising 270M shared parameters. Immediately following the LayerNorm of the shared feed-forward block, a language-specific bottleneck feed-forward module with a bottleneck dimension of 384 is inserted. A residual connection is placed after the module's LayerNorm; the LayerNorm operations preceding and succeeding the modular bottleneck are shared across languages. Each language-specific module accounts for approximately 7M parameters.

    During forward and backward passes, input text in language ll is routed exclusively through the shared parameters and the specific modular bottleneck of language ll. For NN supported languages, the total parameter count of the model scales as 270M+7M×N270\text{M} + 7\text{M} \times N. However, computational cost per token (measured in FLOPs) and the active parameter count per language remain constant and identical to a non-modular transformer baseline with a single bottleneck layer.

  2. Knowl 2 — Zero-Shot Cross-Lingual Downstream Transfer via Module Swapping

    model/method

    X-MOD performs zero-shot cross-lingual transfer on downstream tasks through a two-step fine-tuning and module swapping procedure:

    1. Source-Language Task Fine-Tuning: Given labeled task data in a source language (typically English), a task prediction head is attached to the final [CLS] token representation. During fine-tuning on the source language objective (e.g., cross-entropy classification or span extraction), all shared transformer parameters (attention, primary feed-forward, layer norms) and the prediction head are updated, while the source language's bottleneck modules and token embedding layer remain frozen.

    2. Target-Language Zero-Shot Inference: At test time, zero-shot transfer to an evaluation target language ltargetl_{\text{target}} is executed by replacing the source language's bottleneck modules with the target language's pre-trained (or post-hoc learned) modular components. If the target language was added post-hoc, its dedicated embedding layer is also swapped in. Target language input is then processed through the fine-tuned shared transformer weights without requiring task-specific annotations in the target language.

  3. Knowl 3 — Post-Hoc Language Extension for Modular Pre-Trained Transformers

    model/method

    X-MOD allows incorporating new, unseen languages after the initial multi-task pre-training phase without modifying or degrading performance on existing pre-trained languages:

    1. Target Tokenizer and Vocabulary: A dedicated SentencePiece tokenizer with a vocabulary size of 30,000 subwords is trained on monolingual text of the new target language lnewl_{\text{new}}, ensuring proper coverage of language-specific scripts.

    2. Embedding Initialization: A new embedding layer is initialized for lnewl_{\text{new}}. Tokens that lexically overlap with the pre-training vocabulary (e.g., XLM-R) inherit their embedding weights directly from the pre-trained embedding table, while unseen tokens and positional embeddings are randomly initialized.

    3. Modular Allocation and MLM Training: New bottleneck feed-forward modules of dimension 384 are allocated at every transformer layer for lnewl_{\text{new}}. All existing shared transformer weights (multi-head attention, shared FFN, layer norms) remain frozen. The new target embeddings and modular components are trained via Masked Language Modeling (MLM) on target-language monolingual text using a linear learning rate schedule peaking at 1×10−41 \times 10^{-4}.

  4. Knowl 4 — Mitigation of the Curse of Multilinguality and Positive Cross-Lingual Transfer

    empirical result

    In fully shared multilingual transformers (the SHARED baseline), increasing the number of pre-training languages from 13 to 30, 60, and 75 causes downstream transfer performance to degrade and language modeling perplexity to increase—a failure known as the curse of multilinguality. Controlling for training duration by scaling total updates (125k, 195k, 265k, 269k steps) so that all models see an identical number of examples per language confirms that this degradation is driven by negative parameter interference rather than reduced per-language training tokens.

    In contrast, X-MOD eliminates negative interference through language-specific modular isolation and enables positive cross-lingual transfer as pre-training languages increase from 13 to 60:

    • Perplexity: Language modeling perplexity on pre-trained evaluation languages steadily improves or remains stable on X-MOD as pre-training languages increase, whereas it worsens on SHARED.
    • XNLI Cross-Lingual Accuracy: When holding per-language examples constant, X-MOD's average zero-shot accuracy across pre-trained evaluation languages improves from 72.0% (at 13 languages) to 73.5% (at 60 languages), whereas SHARED drops from 72.6% to 72.5%.
    • WikiANN NER: X-MOD average F1 rises from 56.2 (13 languages) to 62.8 (60 languages), whereas SHARED reaches only 58.8 F1.
  5. Knowl 5 — Zero-Shot Cross-Lingual Transfer Benchmark on 60 Pre-Trained Languages

    data/table

    The table below reports zero-shot downstream transfer results for X-MOD and the FLOP- and parameter-equivalent SHARED baseline, both pre-trained on 60 languages for 265k update steps on the CC100 dataset. All models are fine-tuned exclusively on English training data and evaluated on pre-trained target languages across four benchmarks: XNLI (accuracy), WikiANN NER (F1), XQuAD QA (F1 / Exact Match), and MLQA (F1 / Exact Match). Scores are averaged over 5 random seeds.

    Task / Model en ar fr hi ko ru th vi ta id fi sw avg
    NER (F1)
    X-MOD 81.4 78.9 77.2 70.1 53.0 59.1 2.8 66.2 51.1 50.5 78.6 73.4 62.8
    SHARED 81.5 74.1 74.7 64.4 46.0 58.3 4.0 63.7 52.5 51.5 74.4 57.2 58.8
    XNLI (Acc)
    X-MOD 84.4 71.2 77.6 68.3 - 74.1 71.7 73.4 - - - 66.9 73.5
    SHARED 82.8 69.2 75.6 66.6 - 73.2 68.5 72.5 - - - 62.1 72.5
    XQuAD (F1 / EM)
    X-MOD 85.1 / 73.4 68.1 / 52.4 - 67.5 / 50.3 - 75.0 / 57.8 66.3 / 52.6 74.9 / 54.6 - - - - 72.8 / 56.9
    SHARED 83.8 / 72.1 64.6 / 48.5 - 65.8 / 48.3 - 72.7 / 54.5 63.0 / 48.0 72.6 / 52.1 - - - - 70.4 / 53.9
    MLQA (F1 / EM)
    X-MOD 80.1 / 66.9 58.6 / 38.9 - 60.7 / 42.4 - - - 67.5 / 46.1 - - - - 66.7 / 48.6
    SHARED 79.6 / 66.5 53.6 / 33.9 - 58.7 / 40.4 - - - 64.9 / 43.8 - - - - 64.2 / 46.2

    X-MOD consistently outperforms the SHARED baseline across all tasks (+4.0 F1 on NER, +1.0% accuracy on XNLI, +2.4 F1 / +3.0 EM on XQuAD, +2.5 F1 / +2.4 EM on MLQA), demonstrating that modular capacity benefits both high-resource languages (e.g., English, French) and low-resource languages (e.g., Swahili, Hindi).

  6. Knowl 6 — Performance Parity Between Pre-Trained and Post-Hoc Added Languages

    empirical result

    When evaluating languages added post-hoc to a frozen pre-trained X-MOD model versus pre-training on those same languages from the start, downstream performance is on par across tasks, with no measurable degradation.

    In a controlled experiment comparing two 60-language models with swapped initial and post-hoc subsets of 16 target languages:

    1. Pre-training vs. Post-hoc Equivalence: Across languages such as German (de: 75.4% pre-trained vs 75.4% added), Spanish (es: 78.5% vs 78.4%), Russian (ru: 74.1% vs 74.0%), Hindi (hi: 68.3% vs 68.2%), and Thai (th: 71.7% vs 71.4%), zero-shot XNLI accuracy shows negligible variance when languages see identical numbers of monolingual examples.
    2. Language Family Independence: Parity holds both for languages whose linguistic families were present in the initial pre-training set (e.g., Germanic, Romance, Slavic, Iranian) and for languages from completely unseen families (e.g., Vietnamese [Austroasiatic], Thai [Kra-Dai], Korean [Koreanic], Japanese [Japonic], Greek [Hellenic], Turkish [Turkic]). Pre-training on a subset of data-rich languages is sufficient for X-MOD to generalize to unseen language families via modular post-hoc adaptation.
  7. Knowl 7 — Failure of Post-Hoc Adapter Insertion to Lift the Curse of Multilinguality

    empirical result

    Adding modular adapters (e.g., MAD-X) to a conventionally pre-trained non-modular multilingual transformer after pre-training fails to mitigate the curse of multilinguality.

    When a standard transformer without modular components (shared_nm) is pre-trained across 13 to 60 languages for 100k update steps and subsequently trained with language-specific adapters for 25k steps, the resulting zero-shot XNLI transfer performance closely mirrors the declining performance curve of shared_nm (dropping from ~72.5% to ~70.5% average accuracy). Because catastrophic capacity sharing and negative interference occur during monolithic pre-training, the shared weights are permanently compromised. Modular capacity must be integrated from the beginning of pre-training to prevent negative interference in the shared representation space.

  8. Knowl 8 — Update Step Dynamics and Modularity Induction in X-MOD

    empirical result

    The benefits of modular multilingual pre-training depend on the number of optimization steps relative to the language set size:

    1. Small Language Regimes (13 Languages): When pre-trained for 125k update steps on 13 languages, the SHARED baseline slightly outperforms X-MOD on cross-lingual XNLI (72.6% vs 72.0%). When extending pre-training on the same 13 languages to 250k update steps, modular specialization takes effect: X-MOD overtakes SHARED on English source accuracy (84.5% vs 84.1%) and increases cross-lingual accuracy to 73.3%, matching SHARED.
    2. Larger Language Regimes (60 Languages): For 60-language models, intermediate checkpoint evaluations from 50k to 265k steps demonstrate that X-MOD continuously and immediately outperforms SHARED at every evaluated checkpoint (maintaining an advantage of +1.0% to +1.5% accuracy on XNLI), indicating that fully shared models experience negative interference early during large-scale multilingual training.
  9. Knowl 9 — SHARED Baseline Architecture for Parameter- and FLOP-Equivalence

    experimental setup

    To isolate the effect of modular language parameterization from changes in model capacity or inference compute, the SHARED baseline mirrors the exact architecture of X-MOD:

    1. Architecture: SHARED incorporates the same 12-layer, 768-dimensional transformer encoder and inserts a single bottleneck feed-forward layer of dimension 384 after the FFN block at every layer, shared identically across all languages.
    2. FLOP Equivalence: Because each forward pass in X-MOD accesses exactly one language module per layer, SHARED and X-MOD perform identical FLOPs and activate the same number of parameters per token during pre-training and downstream inference.
    3. Fine-Tuning Control: During downstream task fine-tuning on source data, SHARED freezes its single bottleneck layer and embedding matrix, ensuring that fine-tuning modifies the exact same number of parameters (~270M shared transformer parameters) in both SHARED and X-MOD.

Coverage note — None was omitted; the extracted knowls comprehensively capture X-MOD's architecture, downstream transfer procedure, post-hoc language extension method, empirical validations across language scales, benchmark results, pre-training vs. post-hoc parity, adapter comparisons, training step dynamics, and baseline controls.

References

  1. 1.David Ifeoluwa Adelani, Jade Z. Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba O. Alabi, Seid Muhie Yimam, Tajuddeen Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin P. Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane Mboup, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima Diop, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021. MasakhaNER: Named Entity Recognition for African Languages. In Transactions of the Association for Computational Linguistics 2021.
  2. 2.Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Learning to compose neural networks for question answering. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 1545–1554. The Association for Computational Linguistics.
  3. 3.Alan Ansell, Edoardo Maria Ponti, Anna Korhonen, and Ivan Vulic. 2021a. Composable sparse fine-tuning for cross-lingual transfer. arXiv preprint.
  4. 4.Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021b. MAD-G: Multilingual adapter generation for efficient cross-lingual transfer. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4762–4781, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  5. 5.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  6. 6.Ankur Bapna and Orhan Firat. 2019. Simple, scalable adaptation for neural machine translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 1538–1548. Association for Computational Linguistics.
  7. 7.Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020a. Parsing with multilingual BERT, a small corpus, and a small treebank. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1324–1334, Online. Association for Computational Linguistics.
  8. 8.Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020b. Parsing with multilingual bert, a small treebank, and a small corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, EMNLP 2020, Online Event, 16-20 November 2020, pages 1324–1334.
  9. 9.Vincent S. Chen, Sen Wu, Alexander J. Ratner, Jen Weng, and Christopher Ré. 2019. Slice-based learning: A programming model for residual learning in critical data slices. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 9392–9402.
  10. 10.Alexandra Chronopoulou, Dario Stojanovski, and Alexander Fraser. 2020. Reusing a Pretrained Language Model on Languages with Limited Corpora for Unsupervised NMT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2703–2711, Online. Association for Computational Linguistics.
  11. 11.Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2021. Rethinking embedding coupling in pre-trained language models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  12. 12.Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020. Improving multilingual models with language-clustered vocabularies. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4536–4546. Association for Computational Linguistics.
  13. 13.Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022. CANINE: pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10.
  14. 14.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Conference of the Association for Computational Linguistics, ACL 2020, Virtual Conference, July 6-8, 2020, pages 8440–8451.
  15. 15.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186.
  17. 17.Philipp Dufter and Hinrich Schütze. 2020. Identifying elements essential for BERT's multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4423–4437, Online. Association for Computational Linguistics.
  18. 18.Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir, Gustavo A. Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando A. Coto Solano, Ngoc Thang Vu, and Katharina Kann. 2021. AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages. arXiv preprint.
  19. 19.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv preprint.
  20. 20.Xavier Garcia, Noah Constant, Ankur Parikh, and Orhan Firat. 2021. Towards continual learning for multilingual machine translation via vocabulary substitution. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1184–1192, Online. Association for Computational Linguistics.
  21. 21.Goran Glavas, Robert Litschko, Sebastian Ruder, and Ivan Vulic. 2019. How to (properly) evaluate cross-lingual word embeddings: On strong baselines, comparative analyses, and some misconceptions. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 710–721.
  22. 22.Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021. Demix layers: Disentangling domains for modular language modeling. arXiv preprint.
  23. 23.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022. Towards a unified view of parameter-efficient transfer learning. In 10th International Conference on Learning Representations, ICLR 2022, Virtual Conference, April 25 - 29, 2022.
  24. 24.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzkebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, pages 2790–2799.
  25. 25.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 12-18 July 2020, Virtual Conference.
  26. 26.Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020. Cross-lingual ability of multilingual BERT: an empirical study. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020.
  27. 27.Anne Lauscher, Olga Majewska, Leonardo F. R. Ribeiro, Iryna Gurevych, Nikolai Rozanov, and Goran Glavaš. 2020a. Common sense or world knowledge? investigating adapter-based knowledge injection into pretrained transformers. In Proceedings of Deep Learning Inside Out (DeeLIO): The First Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 43–49, Online. Association for Computational Linguistics.
  28. 28.Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020b. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online.
  29. 29.Hang Le, Juan Miguel Pino, Changhan Wang, Jiatao Gu, Didier Schwab, and Laurent Besacier. 2021. Lightweight adapter tuning for multilingual speech translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 2: Short Papers), Virtual Event, August 1-6, 2021, pages 817–824. Association for Computational Linguistics.
  30. 30.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  31. 31.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021a. Compacter: Efficient low-rank hypercomplex adapter layers. Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021.
  32. 32.Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021b. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 565–576. Association for Computational Linguistics.
  33. 33.Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021. When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 448–462, Online. Association for Computational Linguistics.
  34. 34.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  35. 35.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1946–1958.
  36. 36.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021a. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503, Online. Association for Computational Linguistics.
  37. 37.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a. AdapterHub: A Framework for Adapting Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (System Demonstrations), EMNLP 2020, Virtual Conference, 2020.
  38. 38.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020b. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7654–7673, Online. Association for Computational Linguistics.
  39. 39.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021b. UNKs Everywhere: Adapting Multilingual Language Models to New Scripts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Online, November , 2021.
  40. 40.Jerin Philip, Alexandre Berard, Matthias Gallé, and Laurent Besacier. 2020. Monolingual adapters for zero-shot neural machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4465–4470, Online. Association for Computational Linguistics.
  41. 41.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual bert? In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4996–5001.
  42. 42.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  43. 43.Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, and Iryna Gurevych. 2021. What to pre-train on? efficient intermediate task selection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 10585–10605. Association for Computational Linguistics.
  44. 44.Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019. Massively multilingual transfer for NER. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 151–164.
  45. 45.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392.
  46. 46.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning multiple visual domains with residual adapters. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA, pages 506–516.
  47. 47.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2018. Efficient parametrization of multi-domain deep neural networks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 8119–8127.
  48. 48.Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych. 2021. Adapterdrop: On the efficiency of adapters in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 7930–7946. Association for Computational Linguistics.
  49. 49.Andreas Rücklé, Jonas Pfeiffer, and Iryna Gurevych. 2020. Multicqa: Zero-shot transfer of self-supervised text matching models on a massive scale. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 2471–2486. Association for Computational Linguistics.
  50. 50.Sebastian Ruder, Ivan Vulić, and Anders Søgaard. 2019. A survey of cross-lingual embedding models. Journal of Artificial Intelligence Research, 65:569–631.
  51. 51.Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, Online. Association for Computational Linguistics.
  52. 52.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  53. 53.Asa Cooper Stickland, Alexandre Berard, and Vassilina Nikoulina. 2021. Multilingual domain adaptation for NMT: decoupling language and domain information with adapters. In Proceedings of the Sixth Conference on Machine Translation, WMT@EMNLP 2021, Online Event, November 10-11, 2021, pages 578–598. Association for Computational Linguistics.
  54. 54.Asa Cooper Stickland and Iain Murray. 2019. BERT and pals: Projected attention layers for efficient adaptation in multi-task learning. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 5986–5995. PMLR.
  55. 55.Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2021. Charformer: Fast character transformers via gradient-based subword tokenization. arXiv preprint.
  56. 56.Ke M. Tran. 2020. From english to foreign languages: Transferring pre-trained language models. arXiv preprint.
  57. 57.Ahmet Üstün, Alexandre Berard, Laurent Besacier, and Matthias Gallé. 2021. Multilingual unsupervised neural machine translation with denoising adapters. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6650–6662. Association for Computational Linguistics.
  58. 58.Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020. UDapter: Language adaptation for truly Universal Dependency parsing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2302–2315, Online. Association for Computational Linguistics.
  59. 59.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA, pages 5998–6008.
  60. 60.Giorgos Vernikos and Andrei Popescu-Belis. 2021. Subword mapping and anchoring across languages. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2633–2647, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  61. 61.Marko Vidoni, Ivan Vulić, and Goran Glavaš. 2020. Orthogonal language and task adapters in zero-shot cross-lingual transfer. In arXiv preprint.
  62. 62.Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou. 2021a. K-adapter: Infusing knowledge into pre-trained models with adapters. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 1405–1418. Association for Computational Linguistics.
  63. 63.Xinyi Wang, Yulia Tsvetkov, Sebastian Ruder, and Graham Neubig. 2021b. Efficient test time adapter ensembling for low-resource language varieties. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 730–737, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  64. 64.Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending multilingual BERT to low-resource languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2649–2656, Online. Association for Computational Linguistics.
  65. 65.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  66. 66.Shijie Wu, Alexis Conneau, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Emerging cross-lingual structure in pretrained language models. In Proceedings of the 58th Conference of the Association for Computational Linguistics, ACL 2020, Virtual Conference, July 6-8, 2020, pages 6022–6034.
  67. 67.Shijie Wu and Mark Dredze. 2019. Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 833–844, Hong Kong, China. Association for Computational Linguistics.
  68. 68.Shijie Wu and Mark Dredze. 2020. Are all languages created equal in multilingual BERT? In Proceedings of the 5th Workshop on Representation Learning for NLP, pages 120–130, Online. Association for Computational Linguistics.
  69. 69.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. Byt5: Towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics 2022.

Citation

MLA
Pfeiffer, J., et al. “Lifting the Curse of Multilinguality by Pre-training Modular Transformers”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3479–95, https://doi.org/10.18653/v1/2022.naacl-main.255.
APA
Pfeiffer, J., Goyal, N., Lin, X. V., Li, X., Cross, J., Riedel, S., & Artetxe, M. (2022). Lifting the Curse of Multilinguality by Pre-training Modular Transformers. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3479–3495. https://doi.org/10.18653/v1/2022.naacl-main.255
Chicago
Pfeiffer, J., N. Goyal, X. V. Lin, et al. 2022. “Lifting the Curse of Multilinguality by Pre-training Modular Transformers”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3479–95. https://doi.org/10.18653/v1/2022.naacl-main.255.
Harvard
Pfeiffer, J. et al. (2022) “Lifting the Curse of Multilinguality by Pre-training Modular Transformers”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3479–3495. Available at: https://doi.org/10.18653/v1/2022.naacl-main.255.
Vancouver
1. Pfeiffer J, Goyal N, Lin XV, Li X, Cross J, Riedel S, Artetxe M (2022) Lifting the Curse of Multilinguality by Pre-training Modular Transformers. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3479–3495

BibTeX

@inproceedings{pfeiffer-etal-2022-lifting,
    title = "Lifting the Curse of Multilinguality by Pre-training Modular Transformers",
    author = "Pfeiffer, Jonas  and
      Goyal, Naman  and
      Lin, Xi Victoria  and
      Li, Xian  and
      Cross, James  and
      Riedel, Sebastian  and
      Artetxe, Mikel",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.255/",
    doi = "10.18653/v1/2022.naacl-main.255",
    pages = "3479--3495"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/