How Multilingual is Multilingual BERT?

Telmo PiresEva SchlingerDan Garrette

article2019ACL1,770 citations

Shows why Multilingual BERT succeeds at zero-shot cross-lingual transfer across different scripts and typologies despite training only on monolingual text, while identifying the systematic representational weaknesses that constrain certain language pairs.

Listen

Natural language processing models have traditionally required expensive human-annotated datasets for each individual language and task. The introduction of Multilingual BERT (M-BERT), a single neural network trained simultaneously on raw Wikipedia text across 104 languages without explicit translation signals or language markers, offers a promising path forward. The article investigates the extent and mechanisms of M-BERT's zero-shot cross-lingual capabilities, evaluating whether a model fine-tuned on task labels in only one language can effectively generalize to other languages without additional training.

To evaluate these capabilities, the researchers conducted systematic probing experiments across various linguistic benchmarks. They evaluated sequence tagging performance on named entity recognition (identifying proper names and categories) across up to 16 languages and part-of-speech tagging (grammatical labeling) across 41 languages. They also examined performance across distinct writing scripts, different grammatical word orders, and mixed-language text (code-switching), while mapping the model's internal vector representations to test whether it organizes distinct languages into a shared space.

The findings show that M-BERT achieves strong zero-shot transfer that goes far beyond simple vocabulary memorization. For example, part-of-speech accuracy among European languages consistently exceeded 80%, and an M-BERT model trained only on Urdu in Arabic script achieved 91% accuracy when evaluated on Hindi in Devanagari script. Furthermore, the model generalizes well to code-switched Hindi-English text (reaching 86.59% accuracy using only standard monolingual training). In addition, vector space analysis revealed that a simple translation offset reliably identifies parallel sentences across languages with over 50% nearest-neighbor accuracy. However, transfer degrades substantially when languages diverge in word order (such as between subject-verb-object and subject-object-verb structures) and fails on transliterated text written in non-standard scripts (yielding only 50.41% accuracy).

These results demonstrate that large-scale pre-training on monolingual corpora naturally forces diverse languages into a shared conceptual representation, likely anchored by universal tokens like numbers and URLs. In practical terms, organizations can significantly reduce data labeling costs and deployment timelines by training models in high-resource languages (such as English) and deploying them across global markets. Nevertheless, risk remains elevated when deploying these systems across languages with fundamentally different grammatical structures or non-standard writing formats, where zero-shot generalization is less reliable.

Decision-makers building multilingual language applications should leverage M-BERT to bootstrap services for low-resource languages, but they must avoid relying entirely on zero-shot transfer for typologically distant language pairs or informal transliterated text. For high-stakes deployments involving divergent grammatical orders or social media content, practitioners should incorporate explicit multilingual training objectives or gather targeted in-language supervision. Future research should prioritize architectures that explicitly align distinct word orders and evaluate transfer across a broader range of complex language understanding tasks.

arXiv: 1906.01502
Cover for How Multilingual is Multilingual BERT?

Abstract

In this paper, we show that Multilingual BERT (M-BERT), released by Devlin et al. (2018) as a single language model pre-trained from monolingual corpora in 104 languages, is surprisingly good at zero-shot cross-lingual model transfer, in which task-specific annotations in one language are used to fine-tune the model for evaluation in another language. To understand why, we present a large number of probing experiments, showing that transfer is possible even to languages in different scripts, that transfer works best between typologically similar languages, that monolingual corpora can train models for code-switching, and that the model can find translation pairs. From these results, we can conclude that M-BERT does create multilingual representations, but that these representations exhibit systematic deficiencies affecting certain language pairs.

Table of Contents

  • 1 Introduction
  • 2 Models and Data
  • 2.1 Named entity recognition experiments
  • 2.2 Part of speech tagging experiments
  • 3 Vocabulary Memorization
  • 3.1 Effect of vocabulary overlap
  • 3.2 Generalization across scripts
  • 4 Encoding Linguistic Structure
  • 4.1 Effect of language similarity
  • 4.2 Generalizing across typological features
  • 4.3 Code switching and transliteration
  • 5 Multilingual characterization of the feature space
  • 5.1 Experimental Setup
  • 5.2 Results
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Model Parameters
  • B CoNLL Results for En-Bert
  • C Some pos Results for En-Bert

Knowls

  1. Knowl 1 — Linear Translation Vector Arithmetic in M-BERT Hidden Layer Space

    model/method

    To analyze whether Multilingual BERT (M-BERT) induces a shared multilingual feature space without explicit cross-lingual training, sentence representations can be mapped across languages using a constant translation offset vector.

    Given a set of MM parallel sentence pairs (s1,i,s2,i)i=1M(s_{1,i}, s_{2,i})_{i=1}^M in language L1L_1 and language L2L_2, each sentence is passed through an un-fine-tuned M-BERT model. For each layer l∈{1,…,12}l \in \{1, \dots, 12\}, the sentence vector vL,i(l)v_{L, i}^{(l)} is computed as the average of the hidden feature activations of its input tokens, excluding the special tokens [CLS] and [SEP].

    The translation vector from L1L_1 to L2L_2 at layer ll is defined as the mean difference vector across all MM pairs: vˉL1→L2(l)=1M∑i=1M(vL2,i(l)−vL1,i(l))\bar{v}_{L_1 \to L_2}^{(l)} = \frac{1}{M} \sum_{i=1}^M \left( v_{L_2, i}^{(l)} - v_{L_1, i}^{(l)} \right)

    To translate a source sentence representation vL1,i(l)v_{L_1, i}^{(l)}, the offset is added to produce vL1,i(l)+vˉL1→L2(l)v_{L_1, i}^{(l)} + \bar{v}_{L_1 \to L_2}^{(l)}. The predicted translation is identified as the nearest neighbor in the target language sentence pool measured by Euclidean (ℓ2\ell_2) distance.

    Evaluating on 5,000 sentence pairs from WMT16 shows that nearest neighbor translation accuracy exceeds 50% across most layers for English–German (EN-DE), English–Russian (EN-RU), and Urdu–Hindi (UR-HI), peaking between layers 7 and 9 at approximately 70% to 75%. Accuracy drops in the lowest layers (where representations are dominated by language-specific token identity) and in the final layers (which specialize in language-specific masked language model prediction).

  2. Knowl 2 — Impact of Wordpiece Lexical Overlap on Zero-Shot Cross-Lingual Transfer

    empirical result

    To determine whether M-BERT's cross-lingual transfer relies on superficial vocabulary overlap, entity wordpiece overlap between source training entities EtrainE_{\text{train}} and target evaluation entities EevalE_{\text{eval}} is defined as: overlap=∣Etrain∩Eeval∣∣Etrain∪Eeval∣\text{overlap} = \frac{|E_{\text{train}} \cap E_{\text{eval}}|}{|E_{\text{train}} \cup E_{\text{eval}}|}

    In zero-shot Named Entity Recognition (NER) transfer evaluated across 16 languages:

    • Monolingual English BERT (EN-BERT) performance correlates directly with wordpiece overlap: F1F_1 score deteriorates as overlap decreases and approaches 0 for languages in different scripts.
    • In contrast, M-BERT's zero-shot F1F_1 performance remains largely flat across overlap levels, maintaining scores between 40% and 70% even for language pairs with near-zero lexical overlap.

    This demonstrates that M-BERT learns a deep multilingual representation rather than relying on memorizing shared surface subwords.

  3. Knowl 3 — Cross-Script Zero-Shot Generalization in Multilingual BERT

    empirical result

    Multilingual BERT can transfer syntactic sequence tagging models across language pairs that use completely different scripts with zero lexical overlap, despite being trained solely on monolingual corpora without cross-lingual objectives.

    On Universal Dependencies Part-of-Speech (POS) tagging:

    • An M-BERT model fine-tuned only on Urdu (written in Arabic/Perso-Arabic script) achieves 91.1% zero-shot accuracy on Hindi (written in Devanagari script), compared to 97.1% when trained directly on Hindi.
    • An M-BERT model fine-tuned on English (Latin script) achieves 87.1% accuracy on Bulgarian (Cyrillic script).
    • Transfer between languages with differing scripts and differing typological word orders is markedly lower: English (SVO, Latin script) fine-tuning achieves only 49.4% accuracy on Japanese (SOV, Japanese script), while Japanese monolingual accuracy is 96.5%.
    Fine-tuning \ Eval EN BG JA
    EN 96.8 87.1 49.4
    BG 82.2 98.9 51.6
    JA 57.4 67.2 96.5
    Fine-tuning \ Eval HI UR
    HI 97.1 85.9
    UR 91.1 93.8
  4. Knowl 4 — Effect of Typological Word Order Similarity on Cross-Lingual Transfer

    empirical result

    Zero-shot cross-lingual transfer in M-BERT depends strongly on word order typology. Zero-shot Part-of-Speech (POS) accuracy across language pairs correlates positively with the number of shared World Atlas of Language Structures (WALS) word order features (specifically features 81A: Subject/Object/Verb order, 85A: Adposition/Noun, 86A: Genitive/Noun, 87A: Adjective/Noun, 88A: Demonstrative/Noun, and 89A: Numeral/Noun).

    Macro-averaged zero-shot POS accuracies across 41 Universal Dependencies languages show clear degradation when transferring across different canonical orderings:

    1. Subject/Object/Verb order:
    Fine-tuning \ Eval SVO SOV
    SVO 81.55 66.52
    SOV 63.98 64.22
    1. Adjective/Noun order:
    Fine-tuning \ Eval AN NA
    AN 73.29 70.94
    NA 75.10 79.64

    This indicates that M-BERT maps learned lexical structures onto new vocabularies but does not learn systematic reordering transformations across divergent word orders.

  5. Knowl 5 — Cross-Lingual Transfer to Code-Switched and Transliterated Text

    empirical result

    When evaluated on a Hindi-English code-switched (CS) Part-of-Speech (POS) tagging dataset, M-BERT displays contrasting transfer capabilities depending on whether the input text is in standard or transliterated script:

    • Script-corrected input (Hindi words in Devanagari script, English in Latin script): M-BERT fine-tuned exclusively on monolingual Hindi and English data achieves 86.59% accuracy on code-switched text, approaching the 90.56% accuracy attained when fine-tuned on code-switched training data.
    • Transliterated input (Hindi words transliterated into Latin script): M-BERT fine-tuned on monolingual Hindi and English achieves only 50.41% accuracy, failing to generalize due to the absence of transliterated text during language model pre-training. When trained directly on transliterated code-switched data, M-BERT reaches 85.64%.
    Training Data Script-Corrected Transliterated
    Monolingual HI + EN (M-BERT) 86.59 50.41
    Monolingual HI + EN (Ball and Garrette, 2018) — 77.40
    Code-switched HI/EN (M-BERT) 90.56 85.64
    Code-switched HI/EN (Bhat et al., 2018) — 90.53
  6. Knowl 6 — Zero-Shot Cross-Lingual NER Performance across European Languages

    data/table

    Zero-shot Named Entity Recognition (NER) F1F_1 scores on CoNLL-2002 and CoNLL-2003 test sets across English (EN), German (DE), Dutch (NL), and Spanish (ES) using Multilingual BERT (M-BERT) fine-tuned on a single language:

    Fine-tuning \ Eval EN DE NL ES
    EN 90.70 69.74 77.36 73.59
    DE 73.83 82.00 76.25 70.03
    NL 65.46 65.68 89.86 72.10
    ES 65.38 59.40 64.39 87.18

    In contrast, monolingual English BERT (EN-BERT) fine-tuned on English achieves F1F_1 scores of only 24.38% on DE, 40.62% on NL, and 49.99% on ES (and fine-tuned on DE achieves 55.36% on EN, 54.84% on NL, and 50.80% on ES), showing that M-BERT's pre-training on 104 languages induces strong cross-lingual generalization.

  7. Knowl 7 — Zero-Shot Part-of-Speech Tagging Accuracy on Universal Dependencies Subsets

    data/table

    Part-of-Speech (POS) tagging accuracies on Universal Dependencies test sets for English (EN), German (DE), Spanish (ES), and Italian (IT) using Multilingual BERT (M-BERT) fine-tuned on a single language:

    Fine-tuning \ Eval EN DE ES IT
    EN 96.82 89.40 85.91 91.60
    DE 83.99 93.99 86.32 88.39
    ES 81.64 88.87 96.71 93.71
    IT 86.79 87.82 91.28 98.11

    All cross-lingual evaluation pairs achieve over 80% accuracy under zero-shot transfer. In comparison, English BERT (EN-BERT) fine-tuned on English achieves only 38.31% on DE, 50.38% on ES, and 46.07% on IT.

  8. Knowl 8 — Zero-Shot Sequence Tagging Fine-Tuning Setup for Multilingual BERT

    experimental setup

    For zero-shot cross-lingual evaluation on Named Entity Recognition (NER) and Part-of-Speech (POS) tagging tasks:

    1. Base Model: BERT-Base, Multilingual Cased (a 12-layer Transformer pre-trained on Wikipedia corpora across 104 languages with a single shared WordPiece vocabulary and no language indicator markers).
    2. Sequence Tagging Architecture: Input sentences are tokenized into wordpieces and passed through the model. The final layer's hidden activations are fed into a linear classification layer to predict tag distributions under a cross-entropy loss. For words split into multiple wordpieces, the prediction corresponding to the first wordpiece is taken as the tag prediction for the entire word.
    3. Fine-Tuning Hyperparameters: Fine-tuning is performed on the source language for 3 epochs with a batch size of 32 and a maximum sequence length of 128. An initial learning rate of 3×10−53 \times 10^{-5} is used with a linear warmup over the first 10% of steps followed by linear decay. A dropout rate of 10% is applied to the final layer.

Coverage note — None was omitted; all primary experiments, typological analyses, code-switching investigations, feature-space translation probing, and empirical benchmark tables are covered.

References

  1. 1.Mikel Artetxe and Holger Schwenk. 2018. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. arXiv preprint arXiv:1812.10464.
  2. 2.Kelsey Ball and Dan Garrette. 2018. Part-of-speech tagging for code-switched, transliterated texts without explicit language identification. In Proceedings of EMNLP.
  3. 3.Irshad Bhat, Riyaz A. Bhat, Manish Shrivastava, and Dipti Sharma. 2018. Universal dependency parsing for Hindi-English code-switching. In Proceedings of NAACL.
  4. 4.Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016. Results of the WMT16 metrics shared task. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL.
  6. 6.Matthew S. Dryer and Martin Haspelmath, editors. 2013. WALS Online. Max Planck Institute for Evolutionary Anthropology, Leipzig. https://wals.info/.
  7. 7.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In Proceedings of NAACL.
  8. 8.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291.
  9. 9.Tahira Naseem, Regina Barzilay, and Amir Globerson. 2012. Selective sharing for multilingual dependency parsing. In Proceedings of ACL.
  10. 10.Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D. Manning, Ryan T. McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016. Universal dependencies v1: A multilingual treebank collection. In Proceedings of LREC.
  11. 11.Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018a. Dissecting contextual word embeddings: Architecture and representation. In Proceedings of EMNLP.
  12. 12.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018b. Deep contextualized word representations. In Proceedings of NAACL.
  13. 13.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of CoNLL.
  14. 14.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019a. BERT rediscovers the classical NLP pipeline. In Proceedings of ACL.
  15. 15.Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019b. What do you learn from context? Probing for sentence structure in contextualized word representations. In Proceedings of ICLR.
  16. 16.Erik F. Tjong Kim Sang. 2002. Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition. In Proceedings of CoNLL.
  17. 17.Daniel Zeman, Martin Popel, Milan Straka, Jan Hajic, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, Francis Tyers, Elena Badmaeva, Memduh Gokirmak, Anna Nedoluzhko, Silvie Cinkova, Jan Hajic jr., Jaroslava Hlavacova, Václava Kettnerová, Zdenka Uresova, Jenna Kanerva, Stina Ojala, Anna Missilä, Christopher D. Manning, Sebastian Schuster, Siva Reddy, Dima Taji, Nizar Habash, Herman Leung, Marie-Catherine de Marneffe, Manuela Sanguinetti, Maria Simi, Hiroshi Kanayama, Valeria dePaiva, Kira Droganova, Héctor Martínez Alonso, Çağrı Çöltekin, Umut Sulubacak, Hans Uszkoreit, Vivien Macketanz, Aljoscha Burchardt, Kim Harris, Katrin Marheinecke, Georg Rehm, Tolga Kayadelen, Mohammed Attia, Ali Elkahky, Zhuoran Yu, Emily Pitler, Saran Lertpradit, Michael Mandl, Jesse Kirchner, Hector Fernandez Alcalde, Jana Strnadová, Esha Banerjee, Ruli Manurung, Antonio Stella, Atsuko Shimada, Sookyoung Kwak, Gustavo Mendonca, Tatiana Lando, Rattima Nitisaroj, and Josie Li. 2017. CoNLL 2017 shared task: Multilingual parsing from raw text to universal dependencies. In Proceedings of CoNLL.

Citation

MLA
Pires, T., et al. “How Multilingual Is Multilingual BERT?”. arXiv, 2019, http://arxiv.org/abs/1906.01502v1.
APA
Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is Multilingual BERT?. arXiv. http://arxiv.org/abs/1906.01502v1
Chicago
Pires, T., E. Schlinger, and D. Garrette. 2019. “How Multilingual Is Multilingual BERT?”. arXiv. http://arxiv.org/abs/1906.01502v1.
Harvard
Pires, T., Schlinger, E. and Garrette, D. (2019) “How multilingual is Multilingual BERT?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1906.01502v1.
Vancouver
1. Pires T, Schlinger E, Garrette D (2019) How multilingual is Multilingual BERT?. arXiv

BibTeX

@article{pires2019how,
  title = {How multilingual is Multilingual BERT?},
  author = {Pires, Telmo and Schlinger, Eva and Garrette, Dan},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1906.01502v1},
  eprint = {1906.01502}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/