Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages
Wietse de VriesMartijn WielingMalvina Nissim
Evaluates zero-shot cross-lingual transfer across 65 source and 105 target languages to identify the key linguistic and pre-training factors that determine successful transfer to low-resource languages.
Natural language processing systems often rely on transfer learning to support low-resource languages that lack labeled training data, typically by fine-tuning models on English and applying them to other languages. However, relying solely on English assumes it is universally representative, which risks underperforming across diverse global languages. The article evaluates what makes a language an effective source or target for cross-lingual transfer, identifying the core linguistic and architectural factors that govern transfer success in zero-shot settings.
To establish these drivers, the article conducts an empirical evaluation using the XLM-RoBERTa multilingual model across 65 source languages and 105 target languages for part-of-speech tagging from the Universal Dependencies dataset. Using linear mixed-effects regression analysis, the article systematically assesses the impact of pre-training inclusion, language family alignment, script and writing system types, word order, training size, and lexical-phonetic distance across language pairs.
The findings demonstrate that target language presence in initial pre-training is the single largest determinant of transfer success, adding an estimated 19.2% improvement in accuracy, followed by an additional 7.4% boost when both source and target are pre-trained. Lower lexical-phonetic distance between language pairs significantly improves performance, as does sharing a language family (adding 6.8% accuracy), sharing a writing system type (adding 3.6%), and sharing word order (adding 1.3%). Consequently, English is rarely the optimal source language, ranking 19th out of 65 evaluated sources with an average accuracy of 62.4%, whereas Romanian achieved the highest overall average cross-lingual transfer accuracy at 67.2%.
These insights demonstrate that standard cross-lingual workflows should move away from English-only source pipelines toward linguistically informed language pairing. Organizations deploying multilingual AI should select source languages that share scripts, families, or close lexical ties with the intended target, or utilize highly versatile alternatives like Romanian. For low-resource languages omitted from model pre-training, direct zero-shot transfer will yield poor performance, meaning investment must focus first on collecting unlabeled text for pre-training rather than fine-tuning existing models.
Decision-makers should note that the analysis is focused on part-of-speech tagging and evaluated with a specific model architecture, meaning performance dynamics for complex semantic tasks such as question answering require further empirical validation. Nonetheless, the statistical strength of the regression model provides high confidence that linguistic similarity and target pre-training are essential prerequisites for cross-lingual natural language deployments.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). Introduces XLM-RoBERTa (XLM-R), the core multilingual model whose pre-training inclusion and transfer dynamics form the primary empirical foundation of the source paper.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Establishes foundational empirical findings on zero-shot cross-lingual transfer in multilingual Transformer encoders across scripts, word orders, and POS tagging benchmarks.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Pioneers cross-lingual language model pre-training (XLM) with shared vocabularies and masked language modeling that modern multilingual transfer architectures rely upon.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Provides the foundational cross-lingual benchmark framework and evaluation methodology for cross-lingual transfer from source to target languages.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). Provides the typological and resource-level taxonomy of global language diversity that motivates moving away from English-centric transfer paradigms.
- Paper: Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review, Fred Philippy et al. (2023). Synthesizes literature on the contributing architectural and linguistic factors governing cross-lingual transfer in multilingual models, building directly upon empirical findings like those in the source paper.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Directly tackles the source's key insight—that pre-training inclusion is crucial for transfer—by scaling multilingual corpora and continued pre-training across more than 500 languages.
- Paper: Composable Sparse Fine-Tuning for Cross-Lingual Transfer, Alan Ansell et al. (2022). Proposes modular sparse fine-tuning methods to facilitate zero-shot cross-lingual transfer without degrading the underlying multilingual model representations.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). Extends cross-lingual transfer evaluations beyond encoder fine-tuning to assess few-shot in-context learning in modern large language models across diverse language typologies.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Broadens the evaluation of multilingual transfer dynamics from POS tagging to generative language models across 70 typologically diverse languages.
- Paper: Do Llamas Work in English? On the Latent Language of Multilingual Transformers, Chris Wendler et al. (2024). Investigates the mechanistic representations and latent language pivoting inside multilingual transformers during cross-lingual generation.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Introduces a representation-based metric to quantify cross-lingual disparity and alignment between high- and low-resource languages across transformer layers.
