Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review
Fred PhilippySiwen GuoShohreh Haddadan
Categorizes and synthesizes empirical findings across five key drivers of cross-lingual transfer in multilingual language models to reconcile conflicting literature and guide more effective zero-shot cross-lingual adaptation.
Modern artificial intelligence increasingly relies on multilingual language models to understand and generate text across numerous languages. These systems frequently perform "zero-shot cross-lingual transfer," applying knowledge learned from a high-resource language to a target language without any task-specific target training data. However, because most models were not explicitly designed with cross-lingual mechanisms, explaining why and how this transfer succeeds has proven difficult. Understanding these underlying drivers is critical for building cost-effective, fair, and reliable natural language processing systems across global languages.
The article aims to provide a unified synthesis of the literature on zero-shot cross-lingual transfer. It evaluates conflicting findings across existing studies and categorizes the contributing elements into five core areas: linguistic similarity, lexical overlap, model architecture, pre-training settings, and pre-training data.
The authors conducted a comprehensive literature review of empirical studies and performance-prediction models across various natural and synthetic language benchmarks. By analyzing methodological differences, such as the use of natural versus synthetic datasets and linear versus nonlinear evaluation metrics, the article reconciles previously contradictory experimental outcomes.
The review identifies several key findings. First, structural linguistic alignment—particularly syntactic similarity—is the most influential linguistic driver of transfer, whereas shared vocabulary (lexical overlap) is not strictly required and matters primarily when languages differ in word order or have small training corpora. Second, model architecture balance is essential: deeper networks improve transfer, but overparameterizing a model can cause it to isolate languages into separate spaces rather than aligning them. Third, data scale and domain consistency across languages strongly dictate success; for example, increasing pre-training data from 200,000 to 1,000,000 sentences per language markedly improves cross-lingual capability, and misaligned data domains significantly degrade performance. Finally, certain design choices enhance transfer, such as removing the next-sentence prediction objective, increasing shared vocabulary size, and using high-quality tokenizers for token-level tasks.
These findings indicate that organizations can improve multilingual AI performance and reduce data collection costs without requiring parallel translation text or identical scripts. Instead of relying solely on massive model scaling or assuming English is always the optimal source language, practitioners can achieve better performance by strategically selecting transfer languages based on syntactic, geographic, and genetic closeness, as well as maintaining balanced, in-domain pre-training corpora.
Decision-makers and engineering teams should align pre-training corpora across consistent domains, eliminate unnecessary training objectives like next-sentence prediction, and invest in high-capacity tokenizers. For low-resource language deployment, teams should select source languages sharing structural syntax rather than assuming shared vocabulary is necessary. Looking forward, researchers should explore pre-training data strategies organized around linguistic feature distributions rather than language labels, while expanding evaluation to generative models.
Confidence in these overarching trends is high, though readers should note limitations. Past literature exhibits methodological variability, synthetic language experiments may not fully reflect natural language complexity, and existing benchmarks focus heavily on classification and extraction rather than modern generative tasks.
