Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages

Wietse de VriesMartijn WielingMalvina Nissim

article2022ACL61 citations

Evaluates zero-shot cross-lingual transfer across 65 source and 105 target languages to identify the key linguistic and pre-training factors that determine successful transfer to low-resource languages.

Listen

Natural language processing systems often rely on transfer learning to support low-resource languages that lack labeled training data, typically by fine-tuning models on English and applying them to other languages. However, relying solely on English assumes it is universally representative, which risks underperforming across diverse global languages. The article evaluates what makes a language an effective source or target for cross-lingual transfer, identifying the core linguistic and architectural factors that govern transfer success in zero-shot settings.

To establish these drivers, the article conducts an empirical evaluation using the XLM-RoBERTa multilingual model across 65 source languages and 105 target languages for part-of-speech tagging from the Universal Dependencies dataset. Using linear mixed-effects regression analysis, the article systematically assesses the impact of pre-training inclusion, language family alignment, script and writing system types, word order, training size, and lexical-phonetic distance across language pairs.

The findings demonstrate that target language presence in initial pre-training is the single largest determinant of transfer success, adding an estimated 19.2% improvement in accuracy, followed by an additional 7.4% boost when both source and target are pre-trained. Lower lexical-phonetic distance between language pairs significantly improves performance, as does sharing a language family (adding 6.8% accuracy), sharing a writing system type (adding 3.6%), and sharing word order (adding 1.3%). Consequently, English is rarely the optimal source language, ranking 19th out of 65 evaluated sources with an average accuracy of 62.4%, whereas Romanian achieved the highest overall average cross-lingual transfer accuracy at 67.2%.

These insights demonstrate that standard cross-lingual workflows should move away from English-only source pipelines toward linguistically informed language pairing. Organizations deploying multilingual AI should select source languages that share scripts, families, or close lexical ties with the intended target, or utilize highly versatile alternatives like Romanian. For low-resource languages omitted from model pre-training, direct zero-shot transfer will yield poor performance, meaning investment must focus first on collecting unlabeled text for pre-training rather than fine-tuning existing models.

Decision-makers should note that the analysis is focused on part-of-speech tagging and evaluated with a specific model architecture, meaning performance dynamics for complex semantic tasks such as question answering require further empirical validation. Nonetheless, the statistical strength of the regression model provides high confidence that linguistic similarity and target pre-training are essential prerequisites for cross-lingual natural language deployments.

Vries et al (2022).pdf
Cover for Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages

Abstract

Cross-lingual transfer learning with large multilingual pre-trained models can be an effective approach for low-resource languages with no labeled training data. Existing evaluations of zero-shot cross-lingual generalisability of large pre-trained models use datasets with English training data, and test data in a selection of target languages. We explore a more extensive transfer learning setup with 65 different source languages and 105 target languages for part-of-speech tagging. Through our analysis, we show that pre-training of both source and target language, as well as matching language families, writing systems, word order systems, and lexical-phonetic distance significantly impact cross-lingual performance. The findings described in this paper can be used as indicators of which factors are important for effective zero-shot cross-lingual transfer to zero- and low-resource languages.

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 3 Results
  • 4 Quantitative discussion
  • 4.1 Pre-training
  • 4.2 LDND distance
  • 4.3 Language family
  • 4.4 Writing systems
  • 5 Qualitative discussion
  • 5.1 Underperforming source languages
  • 5.2 Optimal language pairs
  • 5.3 The best source language
  • 6 Conclusion
  • Acknowledgments
  • Ethics statement
  • References

Knowls

  1. Knowl 1 — Mixed-Effects Regression Model for Cross-Lingual Part-of-Speech Tagging Accuracy

    data/table

    A linear mixed-effects regression model assesses the relative contributions of linguistic and dataset factors to zero-shot cross-lingual part-of-speech (POS) tagging accuracy across 65 source and 105 target languages. The dependent variable is POS tagging accuracy (percentage points). Random intercepts are included for source language, target language, and target language family. The variance explained by the target language random effect is more than three times larger than the variance explained by the source language random effect. The model achieves a conditional R2=91.1%R^2 = 91.1\% (total variance explained) and a marginal R2=47.1%R^2 = 47.1\% (variance explained by fixed effects alone).

    Predictor Coef. Std. Err.
    (Intercept) 42.2 3.3
    Target pre-trained 19.2 2.5
    LDND distance -12.7 1.0
    Both pre-trained 7.4 7.4
    Same family 6.8 6.8
    Source pre-trained 5.6 2.0
    Same writing system type 3.6 0.4
    Same writing system 1.4 0.3
    Same SOV word order 1.3 0.2

    All included fixed predictors are statistically significant at p<0.01p < 0.01. The normalized Levenshtein distance on core vocabulary (LDND) is scaled between 00 (minimum) and 11 (maximum). Target language pre-training status is the single most influential predictor, followed by lexical-phonetic distance.

  2. Knowl 2 — Effect of Target and Source Pre-training and Mutual Intelligibility on Cross-Lingual Transfer

    empirical result

    In cross-lingual transfer with XLM-RoBERTa base, the inclusion of the target language in pre-training data is the single largest contributor to performance, providing an estimated +19.2+19.2 percentage points increase in POS tagging accuracy. Inclusion of the source language in pre-training adds +5.6+5.6 percentage points, and inclusion of both languages adds an additional +7.4+7.4 percentage points.

    For source languages omitted from pre-training, cross-lingual transfer drops into the bottom 25% of all source languages (e.g., Ancient Greek, Classical Chinese, Gothic, Maltese, Naija, North Sami, Old Church Slavonic, Old French, and Wolof) unless the language has high mutual intelligibility in written form with a pre-trained language. Non-pre-trained source languages with pre-trained mutually intelligible counterparts—Faroese (intelligible with Icelandic), Old East Slavic (intelligible with Russian, Belarusian, and Ukrainian), and Western Armenian (intelligible with Eastern Armenian)—achieve high transfer performance on par with pre-trained source languages.

  3. Knowl 3 — Cross-Lingual Overfitting Threshold and Training Data Undersampling

    empirical result

    When fine-tuning multilingual language models (XLM-RoBERTa base) on task-specific source datasets of varying sizes (from 127 to 163,106 sentences), zero-shot cross-lingual transfer accuracy peaks at approximately 10,000 training samples and degrades thereafter. While monolingual (within-language) POS tagging accuracy increases monotonically with additional training samples, training beyond 10,000 samples induces source-language overfitting that impairs cross-lingual generalization to other target languages.

    Undersampling source datasets with more than 10,000 sentences improves cross-lingual performance. For languages with over 50,000 training sentences, optimal cross-lingual transfer is achieved at substantially smaller sample budgets: German and Russian achieve maximum cross-lingual transfer at 1,250 samples, Turkish at 10,000 samples, and Czech at 20,000 samples. Conversely, training sets with fewer than 200 sentences (such as Marathi, Hebrew, and Tamil) underfit and yield poor cross-lingual transfer.

  4. Knowl 4 — Cross-Lingual Source Language Quality of Romanian Versus English

    empirical result

    Evaluating cross-lingual POS tagging transfer across 65 source languages and 105 target languages reveals that English—the default source language in multilingual NLP benchmarks—is sub-optimal, ranking 19th out of 65 source languages with an average cross-lingual accuracy of 62.4%62.4\%. English is only the fifth-best source language within the Germanic branch.

    Romanian achieves the highest overall average cross-lingual accuracy (67.2%67.2\%), followed by Swedish (65.9%65.9\%). Romanian is the optimal source language for 10 individual target languages and provides the highest average accuracy across target languages included in pre-training (81.5%81.5\%) as well as those omitted from pre-training (49.8%49.8\%). Furthermore, when evaluating transfer across 30 distinct language families and major Indo-European branches, Romanian is the top-performing source language for 7 different families, demonstrating superior cross-family generalizability.

  5. Knowl 5 — Symmetrically Optimal Language Pairs in Cross-Lingual Transfer

    empirical result

    Across all source-target combinations among 65 source and 105 target languages in Universal Dependencies 2.8, 11 pairs of languages exhibit mutual bidirectional optimality, where each language acts as the single highest-scoring transfer source for the other in zero-shot POS tagging:

    • Estonian and Finnish
    • Icelandic and Faroese
    • French and Italian
    • Chinese and Japanese
    • Irish and Scottish Gaelic
    • Croatian and Serbian
    • Catalan and Spanish
    • Belarusian and Ukrainian
    • Hindi and Urdu
    • Armenian and Western Armenian
    • English and Swedish

    With the exception of English and Swedish, all symmetric pairs originate from the same country or geographic neighbors and represent closest phylogenetic siblings in the Ethnologue genetic classification scheme. In the case of Chinese and Japanese, mutual optimality spans distinct language families (Sino-Tibetan and Japonic), driven by extensive lexical borrowing and shared logographic writing components.

  6. Knowl 6 — Influence of ASJP Lexical-Phonetic Distance on Transfer Accuracy

    empirical result

    Lexical-phonetic distance between source and target languages—measured by the Automated Similarity Judgment Program (ASJP) normalized Levenshtein distance (LDND) computed over 40-item core word lists—is a significant negative predictor of zero-shot POS transfer accuracy (regression coefficient β=−12.7\beta = -12.7, standard error 1.01.0, p<0.01p < 0.01, with LDND scaled to [0,1][0, 1]).

    Low LDND values, which reflect high cognate overlap and phonetic similarity between the source and target languages, strongly correlate with high cross-lingual POS tagging accuracy. In contrast, at high LDND values (distant, unrelated languages), the measure becomes less informative, with cross-lingual accuracy displaying wide variance driven by pre-training inclusion and script compatibility rather than lexical distance.

  7. Knowl 7 — Effects of Writing System Types on Cross-Script Transfer

    empirical result

    Writing system congruence significantly impacts cross-lingual transfer accuracy. In a linear mixed-effects model:

    • Sharing the same writing system type (e.g., alphabetic, logosyllabic, abjad, or abugida) increases cross-lingual accuracy by +3.6+3.6 percentage points (p<0.01p < 0.01).
    • Sharing the identical writing system (e.g., Latin script) adds an additional +1.4+1.4 percentage points (p<0.01p < 0.01).

    Transfer between different alphabetic writing systems (such as between Latin, Cyrillic, and Greek scripts) performs effectively in pre-trained transformer models. In contrast, logosyllabic writing systems (Chinese characters, Japanese Kana, Classical Chinese, Cantonese) transfer poorly to non-logosyllabic languages, with source models fine-tuned on logosyllabic data placing in the bottom 20% of average cross-lingual transfer accuracy across all source languages.

  8. Knowl 8 — Failure Modes and Selection Rules for Cross-Lingual Source Languages

    empirical result

    Among 65 evaluated source languages, 19 fail to reach a transfer accuracy exceeding 84.2%84.2\% (the minimum within-language baseline accuracy among pre-trained languages, recorded on Sanskrit) on any target language other than themselves. Analysis of these underperforming languages identifies four necessary conditions for an effective cross-lingual source language:

    1. Pre-training and Script Consistency: The source language must be present in the model's pre-training corpus in the exact writing system used in the task dataset. (For instance, Sanskrit transfers poorly because Universal Dependencies uses romanized Sanskrit while XLM-RoBERTa pre-training uses Devanagari script).
    2. Monolingual Performance: The fine-tuned model must achieve high within-language accuracy on the source language itself; low monolingual fit (e.g., Arabic at 75.9%75.9\%) precludes successful transfer.
    3. Writing System Type Match: Source and target should share writing system types (alphabetic-to-alphabetic transfer succeeds across different scripts, whereas logosyllabic sources perform poorly on alphabetic targets).
    4. Adequate Training Volume: Source training data must contain more than 200 annotated sentences (languages with fewer than 200 sentences, such as Marathi, Hebrew, and Tamil, fail to achieve competitive transfer).
  9. Knowl 9 — Multilingual POS Tagging Setup Across 105 Languages Using XLM-RoBERTa

    experimental setup

    The cross-lingual evaluation pipeline uses Universal Dependencies 2.8 data across 105 target languages (with test data) and 65 source languages (with at least 125 training sentences, excluding mixed-code, sign languages, and datasets with fewer than 10 test samples). The underlying model is XLM-RoBERTa base, pre-trained on web-crawled text from 100 languages, which encompasses 53 of the 65 source languages and 58 of the 105 target languages.

    Fine-tuning is conducted for a fixed budget of 1,000 batches of batch size 10 (10,000 samples processed) with a linear learning rate schedule starting at 5×10−55 \times 10^{-5}, 10%10\% inter-layer transformer dropout, and 10%10\% self-attention dropout. Source datasets exceeding 10,000 sentences are randomly undersampled, while smaller datasets are oversampled across epochs. Across all 65 source models evaluated on all 105 target languages, monolingual (within-language) POS tagging achieves a mean accuracy of 94.1%94.1\% (σ=4.5\sigma = 4.5), whereas zero-shot cross-lingual transfer achieves an overall mean accuracy of 57.4%57.4\% (σ=22.4\sigma = 22.4).

Coverage note — None was omitted; all key empirical findings, statistical modeling results, dataset parameters, and linguistic analyses from the paper's contribution are captured.

References

  1. 1.Rouben Paul Adalian. 2010. Historical dictionary of Armenia. Scarecrow Press.
  2. 2.Henning Andersen. 2003. Slavic and the indo-european migrations. Amsterdam studies in the theory and history of linguistic science. Series 4, pages 45–76.
  3. 3.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  4. 4.Stephen Barbour and Cathie Carmichael. 2000. Language and nationalism in Europe. OUP Oxford.
  5. 5.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  6. 6.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating crosslingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805.
  8. 8.David M. Eberhard, Gary F. Simons, and Charles D. Fennig. 2021. Ethnologue: Languages of the World. Twenty-fourth edition. SIL International.
  9. 9.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. CoRR, abs/2003.11080.
  10. 10.Snježana Kordić. 2010. Jezik i nacionalizam (language and nationalism). Zagreb: Durieux (Rotulus Universitas).
  11. 11.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  12. 12.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  13. 13.Marco René Spruit, Wilbert Heeringa, and John Nerbonne. 2009. Associations among linguistic levels. Lingua, 119(11):1624 – 1642. The Forests behind the Trees.
  14. 14.Wietse de Vries, Martijn Bartelds, Malvina Nissim, and Martijn Wieling. 2021. Adapting monolingual models: Data can be scarce when language similarity is high. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4901–4907, Online. Association for Computational Linguistics.
  15. 15.Søren Wichmann, Eric W. Holman, Dik Bakker, and Cecil H. Brown. 2010. Evaluating linguistic distance measures. Physica A: Statistical Mechanics and its Applications, 389(17):3632 – 3639.
  16. 16.Daniel Zeman, Joakim Nivre, et al. 2021. Universal dependencies 2.8.1. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University.

Citation

MLA
Vries, W. de ., et al. “Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7676–85, https://doi.org/10.18653/v1/2022.acl-long.529.
APA
Vries, W. de ., Wieling, M., & Nissim, M. (2022). Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7676–7685. https://doi.org/10.18653/v1/2022.acl-long.529
Chicago
Vries, W. de ., M. Wieling, and M. Nissim. 2022. “Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7676–85. https://doi.org/10.18653/v1/2022.acl-long.529.
Harvard
Vries, W. de ., Wieling, M. and Nissim, M. (2022) “Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7676–7685. Available at: https://doi.org/10.18653/v1/2022.acl-long.529.
Vancouver
1. Vries W de, Wieling M, Nissim M (2022) Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7676–7685

BibTeX

@inproceedings{de-vries-etal-2022-make,
    title = "Make the Best of Cross-lingual Transfer: Evidence from {POS} Tagging with over 100 Languages",
    author = "de Vries, Wietse  and
      Wieling, Martijn  and
      Nissim, Malvina",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.529/",
    doi = "10.18653/v1/2022.acl-long.529",
    pages = "7676--7685"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/