Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review

Fred PhilippySiwen GuoShohreh Haddadan

article2023ACL59 citations

Categorizes and synthesizes empirical findings across five key drivers of cross-lingual transfer in multilingual language models to reconcile conflicting literature and guide more effective zero-shot cross-lingual adaptation.

Listen

Modern artificial intelligence increasingly relies on multilingual language models to understand and generate text across numerous languages. These systems frequently perform "zero-shot cross-lingual transfer," applying knowledge learned from a high-resource language to a target language without any task-specific target training data. However, because most models were not explicitly designed with cross-lingual mechanisms, explaining why and how this transfer succeeds has proven difficult. Understanding these underlying drivers is critical for building cost-effective, fair, and reliable natural language processing systems across global languages.

The article aims to provide a unified synthesis of the literature on zero-shot cross-lingual transfer. It evaluates conflicting findings across existing studies and categorizes the contributing elements into five core areas: linguistic similarity, lexical overlap, model architecture, pre-training settings, and pre-training data.

The authors conducted a comprehensive literature review of empirical studies and performance-prediction models across various natural and synthetic language benchmarks. By analyzing methodological differences, such as the use of natural versus synthetic datasets and linear versus nonlinear evaluation metrics, the article reconciles previously contradictory experimental outcomes.

The review identifies several key findings. First, structural linguistic alignment—particularly syntactic similarity—is the most influential linguistic driver of transfer, whereas shared vocabulary (lexical overlap) is not strictly required and matters primarily when languages differ in word order or have small training corpora. Second, model architecture balance is essential: deeper networks improve transfer, but overparameterizing a model can cause it to isolate languages into separate spaces rather than aligning them. Third, data scale and domain consistency across languages strongly dictate success; for example, increasing pre-training data from 200,000 to 1,000,000 sentences per language markedly improves cross-lingual capability, and misaligned data domains significantly degrade performance. Finally, certain design choices enhance transfer, such as removing the next-sentence prediction objective, increasing shared vocabulary size, and using high-quality tokenizers for token-level tasks.

These findings indicate that organizations can improve multilingual AI performance and reduce data collection costs without requiring parallel translation text or identical scripts. Instead of relying solely on massive model scaling or assuming English is always the optimal source language, practitioners can achieve better performance by strategically selecting transfer languages based on syntactic, geographic, and genetic closeness, as well as maintaining balanced, in-domain pre-training corpora.

Decision-makers and engineering teams should align pre-training corpora across consistent domains, eliminate unnecessary training objectives like next-sentence prediction, and invest in high-capacity tokenizers. For low-resource language deployment, teams should select source languages sharing structural syntax rather than assuming shared vocabulary is necessary. Looking forward, researchers should explore pre-training data strategies organized around linguistic feature distributions rather than language labels, while expanding evaluation to generative models.

Confidence in these overarching trends is high, though readers should note limitations. Past literature exhibits methodological variability, synthetic language experiments may not fully reflect natural language complexity, and existing benchmarks focus heavily on classification and extraction rather than modern generative tasks.

Abstract

In recent years, pre-trained Multilingual Language Models (MLLMs) have shown a strong ability to transfer knowledge across different languages. However, given that the aspiration for such an ability has not been explicitly incorporated in the design of the majority of MLLMs, it is challenging to obtain a unique and straightforward explanation for its emergence. In this review paper, we survey literature that investigates different factors contributing to the capacity of MLLMs to perform zero-shot cross-lingual transfer and subsequently outline and discuss these factors in detail. To enhance the structure of this review and to facilitate consolidation with future studies, we identify five categories of such factors. In addition to providing a summary of empirical evidence from past studies, we identify consensuses among studies with consistent findings and resolve conflicts among contradictory ones. Our work contextualizes and unifies existing research streams which aim at explaining the cross-lingual potential of MLLMs. This review provides, first, an aligned reference point for future research and, second, guidance for a better-informed and more efficient way of leveraging the cross-lingual capacity of MLLMs.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Multilingual Language Models
  • 2.2 (Zero-Shot) Cross-Lingual Transfer
  • 3 Factors That Affect Cross-Lingual Transfer
  • 3.1 Linguistic Similarity
  • 3.2 Lexical Overlap
  • 3.3 Model Architecture
  • 3.4 Pre-Training Settings
  • 3.5 Pre-Training Data
  • 4 Related Work
  • 5 Discussion
  • Limitations
  • Ethics Statement
  • References
  • A Appendix

Knowls

  1. Knowl 1 — Five-Factor Framework for Cross-Lingual Transfer in Multilingual Language Models

    model/method

    The determinants of zero-shot cross-lingual transfer capability in pre-trained Multilingual Language Models (MLLMs) are categorized into five fundamental dimensions:

    1. Linguistic Similarity: Structural, typological, genealogical, phonological, and syntactic relationships between the source and target languages (measured via metrics like WALS, lang2vec/URIEL, or eLinguistics).
    2. Lexical Overlap: The degree of shared surface vocabulary or subwords between language pairs, quantified via corpus vocabulary intersection, Levenshtein distance metrics (e.g., LDND), or ezGlot.
    3. Model Architecture: Structural hyperparameter configurations of the Transformer network, including depth (number of layers), width, number of attention heads, parameter sharing strategies across languages, and the alignment of the input embedding space.
    4. Pre-Training Settings: The optimization objectives (such as Masked Language Modeling vs. Next Sentence Prediction), sequence lengths, language identity markers, and tokenization strategies (fertility, vocabulary size, joint versus disjoint vocabularies).
    5. Pre-Training Data: The scale of the monolingual target-language corpus, the ratio of source to target pre-training data, and domain uniformity or divergence across the pre-training corpora of different languages.
  2. Knowl 2 — Interaction of Lexical Overlap with Syntax and Pre-Training Corpus Size

    empirical result

    Empirical analyses in cross-lingual transfer literature present seemingly contradictory findings regarding lexical overlap: several studies report strong positive correlations with transfer performance, while others find transfer performance to be largely independent of lexical overlap.

    This contradiction is resolved by observing the interaction between lexical overlap, word order differences, and corpus resources:

    • Word Order Interaction: Isolating lexical overlap using synthetic language variants demonstrates that lexical overlap is critical primarily when the source and target languages exhibit dissimilar word orders. When language pairs share similar syntactic structures, the presence of lexical overlap provides minimal additional advantage. Studies reporting low influence evaluated language pairs that naturally shared word order structures (such that varying lexical overlap had little effect) or cross-script pairs (where overlap was already near zero).
    • Data Resource Interaction: The positive correlation between subword overlap and transfer performance intensifies when the source language is pre-trained on a smaller monolingual corpus.
    • Cross-Script Feasibility: Cross-lingual transfer remains feasible between languages with entirely different writing systems and zero surface lexical overlap, demonstrating that lexical overlap is not a strictly necessary condition for multilingual transfer.
  3. Knowl 3 — Syntactic Similarity and Word Order Effects on Cross-Lingual Transfer

    empirical result

    Syntactic similarity—specifically word order alignment (such as subject-verb-object order consistency)—positively affects cross-lingual transfer across dependency parsing (DP), named entity recognition (NER), part-of-speech tagging (POS), natural language inference (NLI), and question answering (QA).

    However, the magnitude of the impact depends on the nature of the syntactic divergence:

    • Artificial vs. Natural Syntactic Alterations: Evaluating models on synthetic languages generated by random permutation or complete inversion of word order degrades transfer performance significantly more severely than structured syntactic adaptation based on dependency trees. Consequently, studies relying purely on randomized word order overestimate the negative impact of natural word order variance.
    • Model Differences in Syntactic Sensitivity: Multilingual BERT (mBERT) exhibits greater sensitivity to word order alignment than XLM-RoBERTa (XLM-R), consistent with probing evidence indicating that mBERT encodes more explicit syntactic knowledge in its representations.
  4. Knowl 4 — Task-Level Sensitivity to Non-Syntactic Linguistic Distances

    empirical result

    Cross-lingual transfer performance is influenced by multiple non-syntactic linguistic dimensions with distinct task-specific sensitivities:

    • Geographical and Genetic Distance: Lower orthodromic geographical distance between primary language locations and lower genetic/genealogical distance (degree of common linguistic ancestry) consistently correlate positively with zero-shot transfer success.
    • Phonological Distance: Low phonological distance correlates with higher transfer accuracy on token-level tasks (Named Entity Recognition, Part-of-Speech tagging, Dependency Parsing, and Question Answering), whereas its influence on sentence-level tasks (Natural Language Inference, Machine Translation) is comparatively negligible.
    • Inventory and Frequency Features: Language inventory features (phonetic, phonological, and morphological component sets) have low predictive importance for transfer language selection. Furthermore, matching unigram token frequencies (Zipf's law distribution) between synthetic and natural languages is insufficient on its own to facilitate cross-lingual transfer.
  5. Knowl 5 — Architectural Capacity, Network Depth, and Parameter Sharing Trade-offs

    empirical result

    The internal capacity and configuration of Transformer-based multilingual language models affect cross-lingual alignment through several mechanisms:

    • Layer Depth vs. Attention Heads: Increasing network depth (number of hidden layers) under a fixed total parameter budget improves cross-lingual transfer performance. In contrast, varying the number of attention heads does not substantially influence transfer performance; satisfactory transfer is achievable with a single attention head.
    • Parameter Constraints and Shared Subspaces: Sharing parameters across all layers enforces parameter efficiency, driving the model to align semantically equivalent expressions into a shared multilingual representation space. Conversely, overparameterizing the model allows it to partition parameters into language-specific subspaces, which degrades zero-shot cross-lingual generalization.
    • Embedding Layer Alignment: Cross-lingual alignment of static token embeddings in the initial embedding layer is essential; randomly reinitializing the embedding layer prior to fine-tuning reduces downstream evaluation scores (such as on the GLUE benchmark) by approximately 40%40\%.
  6. Knowl 6 — Effects of Pre-Training Objectives, Sequence Length, and Tokenizer Quality

    empirical result

    Pre-training hyperparameter and tokenization configurations directly modify zero-shot cross-lingual transfer performance:

    • Learning Objectives: Removing the Next Sentence Prediction (NSP) objective from pre-training improves cross-lingual downstream transfer performance for both Named Entity Recognition (NER) and Natural Language Inference (NLI).
    • Input Sequence Length: Pre-training on longer input sequences enhances cross-lingual transfer, particularly when scaling to large pre-training datasets.
    • Language Identifiers: Providing explicit language identity markers during pre-training does not yield significant cross-lingual gains, indicating that models inherently extract language-discriminative representations from context.
    • Tokenizer Vocabulary and Quality: Expanding joint subword vocabulary size improves multilingual transfer across benchmarks. For bilingual models, disjoint subword vocabularies outperform a single joint vocabulary of equivalent size. Furthermore, high tokenizer quality (characterized by lower subword fertility and lower continuation ratios) strongly improves token-level tasks (POS, NER, QA), but has minor impact on document classification and sentence retrieval.
  7. Knowl 7 — Influence of Target Pre-Training Corpus Scale and Cross-Lingual Domain Uniformity

    empirical result

    The characteristics of pre-training corpora establish the foundation for emergent cross-lingual representations:

    • Target Language Corpus Scale: The volume of pre-training data in the target language correlates strongly with transfer performance on high-level reasoning tasks (such as Natural Language Inference and Question Answering), but exhibits weak correlation with low-level sequence tagging tasks (such as POS tagging and Named Entity Recognition).
    • Contextual vs. Static Embeddings: Small pre-training corpora (e.g., 200k200\text{k} sentences per language) cause MLLMs to perform no better than static word embedding baselines (GloVe, Word2Vec). Scaling the per-language pre-training corpus to 1000k1000\text{k} sentences significantly improves MLLM transfer performance, whereas static word embeddings show no such scaling improvement.
    • Pre-Training Domain Mismatch: Pre-training monolingual corpora on divergent domains across languages (e.g., Wikipedia for language AA and Common Crawl for language BB) degrades cross-lingual transfer more severely than pre-training on non-parallel texts drawn from the same domain.
  8. Knowl 8 — Methodological Discrepancies in Transfer Performance Prediction Models

    empirical result

    Studies employing meta-models to predict cross-lingual transfer performance disagree on the relative importance of predictive features (such as lexical overlap versus syntactic distance). This divergence stems from the underlying regression architectures:

    • Tree-Based Ensembles (GBDT, XGBoost): Nonlinear tree-based meta-models capture complex, non-monotonic feature interactions, identifying lexical overlap and syntax as dominant factors for token-level syntactic tasks (POS, NER, DP) and lower for semantic tasks (NLI).
    • Linear Regularized Models (Lasso Regression): Linear Lasso models assign lower feature weights to variables whose relationship with transfer performance is strongly nonlinear, potentially overemphasizing linearly correlated predictors over nonlinear features.
  9. Knowl 9 — Surveyed Empirical Literature on Cross-Lingual Transfer Factors

    data/table

    The empirical literature investigating factors contributing to cross-lingual transfer in MLLMs spans natural language (NL) and synthetic language (SL) setups across diverse downstream tasks and model architectures:

    Study Tasks Model Lang. Factors Investigated
    Lin et al. (2019) DP, MT, POS, EL / NL LO, LS, PTD
    Pires et al. (2019) POS, NER mBERT NL LO, LS
    Tran and Bisazza (2019) DP mBERT NL LO, LS
    Wu and Dredze (2019) DC, NER, DP, NLI, POS mBERT NL LO
    Artetxe et al. (2020) NLI, DC, QA Bilingual BERT, mBERT NL PTS
    Conneau et al. (2020b) NLI, NER, DP Bilingual BERT NL/SL LO, MA, PTD
    Dufter and Schütze (2020) WA, WT, SR BERT (small) SL LS, MA, PTD
    Lauscher et al. (2020) DP, POS, NER, NLI, QA mBERT, XLM-R NL LS, PTD
    Liu et al. (2020) NLI mBERT NL PTD, PTS
    K et al. (2020) NLI, NER Bilingual BERT NL/SL LO, LS, MA, PTS
    Dolicki and Spanakis (2021) NLI, NER, POS XLM-R NL LS
    Srinivasan et al. (2021) NLI, NER, POS mBERT, XLM-R NL LO, LS, PTD
    Wu et al. (2022) SA, AJ, SS, NLI English RoBERTa SL LS, MA
    Ahuja et al. (2022) DC, NLI, POS, NER, QA mBERT, XLM-R NL LO, LS, PTD, PTS
    de Vries et al. (2022) POS XLM-R Base NL LO, LS
    Deshpande et al. (2022) NLI, NER, POS, QA Bilingual RoBERTa SL LO, LS, MA, PTD
    Eronen et al. (2022) DC mBERT, XLM-R NL LS
    Patil et al. (2022) NER, POS, DC, NLI mBERT NL/SL LO

    Task abbreviations: AJ: Acceptability Judgement; DC: Document Classification; DP: Dependency Parsing; EL: Entity Linking; NER: Named Entity Recognition; NLI: Natural Language Inference; POS: Part-of-Speech Tagging; QA: Question Answering; SA: Sentiment Analysis; SR: Sentence Retrieval; SS: Sentence Similarity; WA: Word Alignment; WT: Word Translation. Factor abbreviations: LO: Lexical Overlap; LS: Language Similarity; MA: Model Architecture; PTS: Pre-Training Settings; PTD: Pre-Training Data. Language types: NL: Natural Languages; SL: Synthetic Languages.

  10. Knowl 10 — Limitations in Methodologies and Scope of Cross-Lingual Transfer Evaluations

    limitation

    Several systemic limitations affect the conclusions drawn across cross-lingual transfer literature:

    1. Ecological Validity of Synthetic Languages: While synthetic languages enable controlled isolation of individual variables (e.g., word order inversion or controlled subword overlap), they lack the morphological, phonotactic, and semantic complexities of natural languages.
    2. Discrepant Evaluation Setups and Meta-Modeling: Differing task sets, language selections, and feature importance estimators across studies hinder direct quantitative meta-comparison.
    3. Publication Bias: A skew toward reporting statistically significant factor correlations may inflate the perceived impact of specific linguistic or architectural features.
    4. Lack of Coverage for Generative/Autoregressive Architectures: The existing body of analytical literature focuses predominantly on masked encoder models (BERT, XLM, XLM-R, RoBERTa), leaving the cross-lingual transfer mechanics of decoder-only and autoregressive generative language models largely unaddressed.

Coverage note — None was omitted; all primary synthesized factors (Linguistic Similarity, Lexical Overlap, Model Architecture, Pre-Training Settings, Pre-Training Data), the conflict resolutions, comparative empirical results, metadata mappings, and stated research limitations were included.

References

  1. 1.Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022. Multi Task Learning For Zero Shot Performance Prediction of Multilingual Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5454–5467, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021. MAD-G: Multilingual Adapter Generation for Efficient Cross-Lingual Transfer. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4762–4781, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  3. 3.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the Cross-lingual Transferability of Monolingual Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  4. 4.Vincent Beaufils and Johannes Tomin. 2020. Stochastic approach to worldwide language classification: the signals and the noise towards long-range exploration. preprint, SocArXiv.
  5. 5.Alexandra Chronopoulou, Dario Stojanovski, and Alexander Fraser. 2023. Language-Family Adapters for Low-Resource Multilingual Neural Machine Translation. In Proceedings of the The Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023), pages 59–72, Dubrovnik, Croatia. Association for Computational Linguistics.
  6. 6.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020a. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  7. 7.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual Language Model Pretraining. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  8. 8.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating Cross-lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  9. 9.Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020b. Emerging Cross-lingual Structure in Pretrained Language Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6022–6034, Online. Association for Computational Linguistics.
  10. 10.Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2022. Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7676–7685, Dublin, Ireland. Association for Computational Linguistics.
  11. 11.Ameet Deshpande, Partha Talukdar, and Karthik Narasimhan. 2022. When is BERT Multilingual? Isolating Crucial Ingredients for Cross-lingual Transfer. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3610–3623, Seattle, United States. Association for Computational Linguistics.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Sumanth Doddapaneni, Gowtham Ramesh, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2021. A Primer on Pretrained Multilingual Language Models. ArXiv:2107.00676 [cs].
  14. 14.Błażej Dolicki and Gerasimos Spanakis. 2021. Analysing The Impact Of Linguistic Features On Cross-Lingual Transfer. ArXiv:2105.05975 [cs].
  15. 15.Matthew S. Dryer and Martin Haspelmath. 2013. WALS Online. Max Planck Institute for Evolutionary Anthropology, Leipzig.
  16. 16.Philipp Dufter and Hinrich Schütze. 2020. Identifying Elements Essential for BERT’s Multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4423–4437, Online. Association for Computational Linguistics.
  17. 17.Juuso Eronen, Michal Ptaszynski, Fumito Masui, Masaki Arata, Gniewosz Leliwa, and Michal Wroczynski. 2022. Transfer language selection for zero-shot cross-lingual abusive language detection. Information Processing & Management, 59(4):102981.
  18. 18.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A Massively Multilingual Multitask Benchmark for Evaluating Cross-lingual Generalisation. In Proceedings of the 37th International Conference on Machine Learning, pages 4411–4421. PMLR.
  19. 19.Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019. Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2485–2494, Hong Kong, China. Association for Computational Linguistics.
  20. 20.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020. SpanBERT: Improving Pre-training by Representing and Predicting Spans. Transactions of the Association for Computational Linguistics, 8:64–77.
  21. 21.Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020. Cross-Lingual Ability of Multilingual BERT: An Empirical Study. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020).
  22. 22.Lazar Kovacevic, Vladimir Bradic, Gerard de Melo, Sinisa Zdravkovic, and Olga Ryzhova. 2022. Ezglot.
  23. 23.Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020. From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online. Association for Computational Linguistics.
  24. 24.Jaeseong Lee, Seung-won Hwang, and Taesup Kim. 2022. FAD-X: Fusing Adapters for Cross-lingual Transfer to Low-Resource Languages. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 57–64, Online only. Association for Computational Linguistics.
  25. 25.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019. Choosing Transfer Languages for Cross-Lingual Learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3125–3135, Florence, Italy. Association for Computational Linguistics.
  26. 26.Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017. URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 8–14, Valencia, Spain. Association for Computational Linguistics.
  27. 27.Chi-Liang Liu, Tsung-Yuan Hsu, Yung-Sung Chuang, and Hung-Yi Lee. 2020. A Study of Cross-Lingual Ability and Language-specific Information in Multilingual BERT. ArXiv:2004.09205 [cs].
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ArXiv:1907.11692 [cs].
  29. 29.Dan Malkin, Tomasz Limisiewicz, and Gabriel Stanovsky. 2022. A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4903–4915, Seattle, United States. Association for Computational Linguistics.
  30. 30.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. ArXiv:1301.3781 [cs].
  31. 31.Shantanu Patankar, Omkar Gokhale, Onkar Litake, Aditya Mandke, and Dipali Kadam. 2022. To Train or Not to Train: Predicting the Performance of Massively Multilingual Models. In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 8–12, Online. Association for Computational Linguistics.
  32. 32.Vaidehi Patil, Partha Talukdar, and Sunita Sarawagi. 2022. Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related Languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 219–233, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  34. 34.Matúš Pikuliak, Marián Šimko, and Mária Bieliková. 2021. Cross-lingual learning for text processing: A survey. Expert Systems with Applications, 165:113765.
  35. 35.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How Multilingual is Multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  36. 36.Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, Online. Association for Computational Linguistics.
  37. 37.Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, and Monojit Choudhury. 2021. Predicting the Performance of Multilingual NLP Models. ArXiv:2110.08875 [cs].
  38. 38.Ke Tran and Arianna Bisazza. 2019. Zero-shot Dependency Parsing with Pre-trained Multilingual Sentence Representations. In Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019), pages 281–288, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Iulia Turc, Kenton Lee, Jacob Eisenstein, Ming-Wei Chang, and Kristina Toutanova. 2021. Revisiting the Primacy of English in Zero-shot Cross-lingual Transfer. ArXiv:2106.16171 [cs].
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  41. 41.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  42. 42.Søren Wichmann, Eric W. Holman, Dik Bakker, and Cecil H. Brown. 2010. Evaluating linguistic distance measures. Physica A: Statistical Mechanics and its Applications, 389(17):3632–3639.
  43. 43.Shijie Wu and Mark Dredze. 2019. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 833–844, Hong Kong, China. Association for Computational Linguistics.
  44. 44.Zhengxuan Wu, Isabel Papadimitriou, and Alex Tamkin. 2022. Oolong: Investigating What Makes Crosslingual Transfer Hard with Controlled Studies. ArXiv:2202.12312 [cs].
  45. 45.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  46. 46.Jianyu Zheng and Ying Liu. 2022. Probing language identity encoded in pre-trained multilingual models: a typological view. PeerJ Computer Science, 8:e899.

Citation

MLA
Philippy, F., et al. “Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 5877–91, https://doi.org/10.18653/v1/2023.acl-long.323.
APA
Philippy, F., Guo, S., & Haddadan, S. (2023). Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5877–5891. https://doi.org/10.18653/v1/2023.acl-long.323
Chicago
Philippy, F., S. Guo, and S. Haddadan. 2023. “Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5877–91. https://doi.org/10.18653/v1/2023.acl-long.323.
Harvard
Philippy, F., Guo, S. and Haddadan, S. (2023) “Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5877–5891. Available at: https://doi.org/10.18653/v1/2023.acl-long.323.
Vancouver
1. Philippy F, Guo S, Haddadan S (2023) Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5877–5891

BibTeX

@inproceedings{philippy-etal-2023-towards,
    title = "Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review",
    author = "Philippy, Fred  and
      Guo, Siwen  and
      Haddadan, Shohreh",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.323/",
    doi = "10.18653/v1/2023.acl-long.323",
    pages = "5877--5891"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/