Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models

Terra BlevinsLuke Zettlemoyer

article2022EMNLP108 citations

Demonstrates that supposedly monolingual English pretraining corpora contain hundreds of millions of foreign-language tokens, revealing that unintentional multilingual contamination largely drives the zero-shot cross-lingual transfer capabilities of English language models.

Listen

English pretrained language models serve as the foundation for modern natural language processing systems. Although these models are nominally trained on English-only data, recent research observed that they transfer surprisingly well to other languages without prior multilingual training. The article investigates why this occurs, demonstrating that standard English pretraining datasets contain substantial amounts of foreign language data and that this unintentional contamination directly enables cross-lingual performance.

To evaluate the scale and impact of this leakage, the article conducted a two-part study. First, it performed automatic language identification alongside qualitative audits across six major English pretraining datasets, including web-crawled collections like C4 and CC-News. Second, it evaluated popular English models (BERT, RoBERTa, and T5) against dedicated multilingual models across more than 50 languages using language modeling and part-of-speech tagging tasks, testing both frozen models and finetuned setups.

The findings show that all evaluated English corpora contain notable foreign text, ranging from 300,000 to over 400 million non-English tokens. Web-crawled datasets exhibited the highest contamination, even when automated language filters were used. Crucially, downstream cross-lingual performance correlated strongly with the amount of in-language text encountered during pretraining (reaching correlations up to 0.67 to 0.68 for RoBERTa), whereas syntactic similarity to English showed much weaker correlation. When fine-tuned for part-of-speech tagging, RoBERTa narrowed its performance gap with dedicated multilingual models to just 2.65 points. Additionally, English T5 outperformed multilingual BERT on part-of-speech tagging for certain languages, such as German and Portuguese, without any task-specific tuning.

These results indicate that large-scale English models are functionally multilingual rather than strictly monolingual. Their cross-lingual capabilities do not reflect true "zero-shot" generalization, but rather direct learning from leaked pretraining data. This means organizations evaluating language models cannot assume clean linguistic separation, which affects how cross-lingual transfer benchmarks and multilingual risks are interpreted.

Moving forward, researchers and developers should account for pretraining contamination when evaluating model capabilities and avoid assuming that automated data cleaning eliminates foreign text. Complete manual data filtering remains practically infeasible at scale, but auditing training data composition will improve the transparency of multilingual evaluations. Limitations of the article include reliance on automated classifiers for language estimates, an evaluation focused primarily on syntax and language modeling, and an analysis restricted to English-centric corpora.

arXiv: 2204.08110
Cover for Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models

Abstract

English pretrained language models, which make up the backbone of many modern NLP systems, require huge amounts of unlabeled training data. These models are generally presented as being trained only on English text but have been found to transfer surprisingly well to other languages. We investigate this phenomenon and find that common English pretraining corpora actually contain significant amounts of non-English text: even when less than 1% of data is not English (well within the error rate of strong language classifiers), this leads to hundreds of millions of foreign language tokens in large-scale datasets. We then demonstrate that even these small percentages of non-English data facilitate cross-lingual transfer for models trained on them, with target language performance strongly correlated to the amount of in-language data seen during pretraining. In light of these findings, we argue that no model is truly monolingual when pretrained at scale, which should be considered when evaluating cross-lingual transfer.

Table of Contents

  • 1 Introduction
  • 2 Pretraining Data Composition
  • 2.1 Automatic Evaluation of Language Composition
  • 2.2 Qualitative Analysis of Non-English Texts
  • 3 Cross-lingual Transfer of English Pretrained Models
  • 3.1 Non-English MLM Evaluation
  • 3.2 POS Performance Across Languages
  • 3.3 Potential Reasons for Cross-lingual Generalization
  • 4 Discussion
  • 5 Limitations
  • Acknowledgements
  • References
  • A Details of Transfer Experiments
  • B The Effect of Tokenization
  • C Full Results of the Automatic Language Identity Analysis
  • D Full Results of Transfer Experiments

Knowls

  1. Knowl 1 — Correlation Between Cross-Lingual Transfer and Pretraining Target Language Volume

    empirical result

    The cross-lingual performance of ostensibly monolingual English pretrained models is strongly correlated with the volume of target-language text unintentionally present in their pretraining corpora. Across Masked Language Modeling (Wiki-40B bits per character, BPC), frozen Part-of-Speech (POS) probing accuracy, and finetuned POS tagging accuracy on Universal Dependencies treebanks, pretraining language data quantity is a stronger predictor of transfer success than syntactic distance/similarity to English.

    Task Model Corr. (ρ\rho) w/ Data ↑\uparrow Corr. (ρ\rho) w/ En Sim. ↓\downarrow
    MLM (BPC) ↓\downarrow BERTbase\text{BERT}_{\text{base}} -0.258 0.097
    BERTlg\text{BERT}_{\text{lg}} -0.258 0.118
    RoBERTabase\text{RoBERTa}_{\text{base}} -0.667∗∗^{**} 0.326∗^{*}
    RoBERTalg\text{RoBERTa}_{\text{lg}} -0.685∗∗^{**} 0.345∗^{*}
    Frozen POS (Acc.) ↑\uparrow BERTbase\text{BERT}_{\text{base}} 0.335∗^{*} -0.332∗^{*}
    BERTlg\text{BERT}_{\text{lg}} 0.314∗^{*} -0.375∗^{*}
    RoBERTabase\text{RoBERTa}_{\text{base}} 0.594∗∗^{**} -0.260
    RoBERTalg\text{RoBERTa}_{\text{lg}} 0.674∗∗^{**} -0.304∗^{*}
    T5base\text{T5}_{\text{base}} 0.131 -0.271
    Finetuned POS (Acc.) ↑\uparrow BERTbase\text{BERT}_{\text{base}} 0.373∗^{*} -0.340∗^{*}
    RoBERTabase\text{RoBERTa}_{\text{base}} 0.507∗∗^{**} -0.292∗^{*}

    Notes: Spearman rank correlation ρ\rho is reported; ∗p<0.05^{*}p < 0.05, ∗∗p<0.001^{**}p < 0.001. For T5, when controlling for languages written in non-Latin scripts (which suffer from high tokenizer unknown rates), the correlation between POS probing performance and in-language pretraining data increases from ρ=0.131\rho = 0.131 to ρ=0.313\rho = 0.313. Syntactic distance is computed using syntactic typological features where lower distance signifies greater similarity to English.

  2. Knowl 2 — Non-English Language Contamination in Standard English Pretraining Corpora

    data/table

    Standard English pretraining corpora contain substantial quantities of non-English text, ranging from hundreds of thousands to hundreds of millions of foreign tokens, despite comprising small overall percentages (<1.6%<1.6\%). Corpora derived from web crawls contain markedly higher amounts and percentages of foreign language text than human-curated datasets.

    Corpus Total Corpus Size Non-English Tokens % Non-English
    English Wikipedia 11.8 GB 983,000 0.05%
    BookCorpus 4.2 GB 322,000 0.04%
    Stories 31 GB 682,000 0.01%
    OpenWebText 38 GB 18,600,000 0.29%
    CC-News 76 GB 201,000,000 1.53%
    C4.En (first 50M lines) 305 GB (total) 56,840,000 0.26%
    C4.En (projected full) 305 GB 406,000,000 0.26%
    BERT Training Data (Wiki + Books) 16.0 GB 1,300,000 0.05%
    RoBERTa Training Data (All above except C4) 161.0 GB 222,000,000 0.78%

    The largest non-English language slices constitute up to 0.01% of BERT training data, 0.15% of RoBERTa training data, and 0.05% of T5 training data. In RoBERTa's pretraining corpus, individual foreign languages account for tens of millions of tokens (e.g., Albanian: 42.5M, Spanish: 40.3M, German: 37.8M, Romanian: 30.2M, Portuguese: 11.5M, Italian: 10.9M, French: 10.1M).

  3. Knowl 3 — Qualitative Categorization of Non-English Content in Pretraining Datasets

    empirical result

    A manual qualitative audit of 200 random samples predicted as non-English by a language identification classifier across six pretraining corpora categorizes data into six distinct classes:

    1. Non-English (NE): Text containing exclusively non-English natural language (BookCorpus: 156, Wikipedia: 129, Stories: 99, OpenWebText: 175, CC-News: 193, C4: 169).
    2. Bilingual (BiL): Mixed text including codeswitching and non-English dialogue within English contexts (BookCorpus: 13, Wikipedia: 11, Stories: 15, OpenWebText: 4, CC-News: 1, C4: 22).
    3. Translations (Trans): Parallel sentences or word-level pairs translating between English and another language (BookCorpus: 2, Wikipedia: 7, Stories: 4, OpenWebText: 2, CC-News: 0, C4: 4).
    4. Entities (Ent): Primarily English text containing non-English named entities (BookCorpus: 1, Wikipedia: 28, Stories: 5, OpenWebText: 1, CC-News: 0, C4: 1).
    5. English Classifier Errors (En): English text misclassified as non-English, typically non-standard dialects, dialogue snippets, or short lines (BookCorpus: 26, Wikipedia: 22, Stories: 55, OpenWebText: 12, CC-News: 6, C4: 3).
    6. Non-language Noise (XX): Lines containing formatting artifacts, tables, code, or non-natural text (BookCorpus: 2, Wikipedia: 3, Stories: 22, OpenWebText: 6, CC-News: 0, C4: 1).

    The majority of detected lines in web-crawled and curated corpora consist of genuine foreign language text (NE) or bilingual text (BiL).

  4. Knowl 4 — Cross-Lingual Masked Language Modeling Generalization

    empirical result

    Evaluating zero-shot language modeling performance via bits per character (BPC) with 15% whole-word masking on the 41 languages of Wiki-40B reveals that English RoBERTa substantially closes the performance gap to dedicated multilingual models.

    • Monolingual BERTbase\text{BERT}_{\text{base}} and BERTlg\text{BERT}_{\text{lg}} exhibit high BPC across target languages, averaging a gap of 2.51 BPC compared to multilingual models (mBERT and XLM-R).
    • Monolingual RoBERTabase\text{RoBERTa}_{\text{base}} and RoBERTalg\text{RoBERTa}_{\text{lg}} reduce the average performance gap to multilingual models down to 0.87 BPC.
    • On several individual languages, RoBERTa models approach the perplexity of mBERT (e.g., in Spanish: RoBERTalg\text{RoBERTa}_{\text{lg}} achieves 1.345 BPC vs. mBERT's 1.036 BPC and XLM-Rlg\text{XLM-R}_{\text{lg}}'s 1.165 BPC; in French: RoBERTalg\text{RoBERTa}_{\text{lg}} achieves 1.414 BPC vs. mBERT's 1.038 BPC).
  5. Knowl 5 — Cross-Lingual Part-of-Speech Tagging via Probing and Fine-Tuning

    empirical result

    Evaluating English pretrained models on Part-of-Speech (POS) tagging across up to 50 languages using Universal Dependencies v2 treebanks demonstrates significant non-zero-shot transfer:

    • Frozen Probing: Monolingual models retain non-trivial POS feature representations. RoBERTa models consistently outperform BERT and T5 encoders across target languages. For several high-resource languages, English T5 outperforms mBERT in frozen probing (e.g., German: T5 achieves 92.88% vs. mBERT's 92.02%; Portuguese: T5 achieves 95.12% vs. mBERT's 94.07%; French: T5 achieves 96.61% vs. mBERT's 95.96%).
    • Finetuning: Fully finetuning encoder weights on target language treebanks drastically reduces the disparity between monolingual and multilingual architectures. While frozen RoBERTabase\text{RoBERTa}_{\text{base}} trails XLM-Rbase\text{XLM-R}_{\text{base}} by an average of 12.5 accuracy points, finetuned RoBERTabase\text{RoBERTa}_{\text{base}} averages within 2.65 accuracy points of XLM-Rbase\text{XLM-R}_{\text{base}} (and achieves comparable scores on languages such as Spanish: 98.46% vs. 98.78%; German: 93.55% vs. 95.19%; French: 97.11% vs. 98.10%).
  6. Knowl 6 — Impact of Subword Tokenization Schemes and UNK Token Rates on Cross-Lingual Evaluation

    empirical result

    Subword tokenization architectures create major differences in cross-lingual representation and evaluation validity:

    1. Subword Fragmentation: All models require more subword units per word for non-English languages than for English (where token/word ratios range between 1.32 and 1.44). For instance, in Basque, the average tokens per word are 1.78 (XLM-R), 2.59 (RoBERTa), and 2.66 (BERT).
    2. Unknown (UNK) Token Handling: RoBERTa uses a byte-level Byte-Pair Encoding (BPE) tokenizer, resulting in 0.0% UNK tokens for arbitrary Unicode text. In contrast, BERT (WordPiece) and T5 (SentencePiece) produce high UNK token rates (>10%>10\%) for scripts not represented during tokenizer construction (e.g., Japanese: 39.97% UNK in BERT, 22.19% in T5; Korean: 59.65% in BERT, 38.63% in T5; Thai: 36.91% in BERT, 28.58% in T5; Hindi: 12.06% in BERT, 42.82% in T5; Arabic: 41.26% in T5).
    3. Evaluation Impact: High UNK rates make masked language modeling artificially predictable, skewing BPC lower, while simultaneously degrading encoder representations on downstream classification tasks.
  7. Knowl 7 — Linear Probing and Fine-Tuning Setup for Cross-Lingual POS Evaluation

    experimental setup

    The cross-lingual evaluation on Universal Dependencies (UD) v2 treebanks uses the following configuration:

    • Classifier Head: A linear classification layer mapping the encoder's final hidden representation to POS tags with parameter matrix of dimension m×lm \times l, where mm is the hidden state size (m=768m = 768 for base models, m=1024m = 1024 for large models) and l=17l = 17 (the Universal POS tagset size). For words segmented into multiple subwords, the input representation is the element-wise mean of its subword embeddings.
    • Optimization: Adam optimizer across all runs.
    • Hyperparameters for Probing (Frozen Encoder): Batch size 256, learning rate 1×10−31 \times 10^{-3}, trained for up to 50 epochs with early stopping patience of 5 on validation loss.
    • Hyperparameters for Fine-Tuning: Encoder and linear head unfrozen, batch size 16, learning rate 5×10−65 \times 10^{-6}, trained for up to 50 epochs with early stopping patience of 5 on validation loss.
    • Execution & Aggregation: Trained on single Nvidia V100 GPUs (16GB for probing, 32GB for finetuning); reported performance metrics are averaged over 5 random initialization seeds.
  8. Knowl 8 — Automatic Pretraining Language Identification Pipeline

    model/method

    Language composition of pretraining corpora is estimated line-by-line using the FastText language identification model:

    • Filtering Criterion: A line is classified as non-English if the FastText model assigns a non-English language label with a confidence score exceeding a threshold of 0.6.
    • Corpus Coverage: The full text of English Wikipedia (11.8 GB), BookCorpus (4.2 GB), Stories (31 GB), OpenWebText (38 GB), and CC-News (76 GB) is evaluated. For C4.En (305 GB), an initial uniform subsample of 50 million lines (approximately 14% of the corpus) is classified, and the resulting non-English token count is scaled proportionally to project total non-English tokens for the complete corpus.
  9. Knowl 9 — Limitations of Language Contamination Measurement and Transfer Analysis

    limitation

    The study's findings are subject to three main limitations:

    1. Classifier Error Propagation: Quantification of non-English tokens relies on an automated classifier (FastText at threshold 0.6), which introduces misclassifications—particularly for low-resource languages, non-standard dialects, short dialogue lines, and noisy or unformatted text.
    2. Task and Linguistic Scope: Downstream cross-lingual evaluation is restricted to masked language modeling (Wiki-40B) and syntactic part-of-speech tagging (Universal Dependencies); behavior may differ on semantic, generative, or complex reasoning tasks.
    3. Monolingual Language Coverage: Contamination auditing and transfer experiments were conducted exclusively on English pretrained models. Although non-English monolingual models trained on web crawls likely experience similar data leakage, their contamination levels were not measured.

Coverage note — No substantial contributed material was omitted. The knowls cover the automatic contamination measurements, qualitative error taxonomy, MLM and POS transfer evaluations, correlation analyses with language size/similarity, tokenization artifacts, experimental hyperparameters, and stated limitations.

References

  1. 1.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637.
  2. 2.Isaac Caswell, Julia Kreutzer, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2021. Quality at a glance: An audit of web-crawled multilingual datasets.
  3. 3.Zewen Chi, Li Dong, Furu Wei, Xianling Mao, and He-Yan Huang. 2020. Can monolingual pretrained models help cross-lingual classification? In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 12–17.
  4. 4.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451.
  5. 5.Leandro Rodrigues de Souza, Rodrigo Nogueira, and Roberto Lotufo. 2021. On the ability of monolingual models to learn language-agnostic representations. arXiv preprint arXiv:2109.01942.
  6. 6.Jacob Delvin. 2019. Multilingual BERT Readme. https://github.com/google-research/bert/blob/master/multilingual.md.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  9. 9.Evangelia Gogoulou, Ariel Ekgren, Tim Isbister, and Magnus Sahlgren. 2021. Cross-lingual transfer of monolingual models. arXiv preprint arXiv:2109.07348.
  10. 10.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  11. 11.Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020. Wiki-40B: Multilingual language model dataset. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2440–2452, Marseille, France. European Language Resources Association.
  12. 12.Suchin Gururangan, Dallas Card, Sarah K Drier, Emily K Gade, Leroy Z Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A Smith. 2022. Whose language counts as high quality? measuring language ideologies in text data selection. arXiv preprint arXiv:2201.10474.
  13. 13.Armand Joulin, Édouard Grave, Piotr Bojanowski, and Tomáš Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431.
  14. 14.Diederik P. Kingma and Jimmy Lei Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of 3rd International Conference of Learning Representations.
  15. 15.Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71.
  16. 16.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291.
  17. 17.Zuchao Li, Kevin Parnow, Hai Zhao, Zhuosheng Zhang, Rui Wang, Masao Utiyama, and Eiichiro Sumita. 2021. Cross-lingual transferring of pre-trained contextualized language models. arXiv preprint arXiv:2107.12627.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  19. 19.Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017. Learning language representations for typology prediction. In Conference on Empirical Methods in Natural Language Processing (EMNLP), Copenhagen, Denmark.
  20. 20.Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic, Christopher D Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020. Universal dependencies v2: An evergrowing multilingual treebank collection. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4034–4043.
  21. 21.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  23. 23.Ke Tran. 2020. From english to foreign languages: Transferring pre-trained language models. arXiv preprint arXiv:2002.07306.
  24. 24.Trieu H Trinh and Quoc V Le. 2018. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847.
  25. 25.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R'emi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface's transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771.
  26. 26.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27.

Citation

MLA
Blevins, T., and L. Zettlemoyer. “Language Contamination Helps Explain the Cross-lingual Capabilities of English Pretrained Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3563–74, https://doi.org/10.18653/v1/2022.emnlp-main.233.
APA
Blevins, T., & Zettlemoyer, L. (2022). Language Contamination Helps Explain the Cross-lingual Capabilities of English Pretrained Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3563–3574. https://doi.org/10.18653/v1/2022.emnlp-main.233
Chicago
Blevins, T., and L. Zettlemoyer. 2022. “Language Contamination Helps Explain the Cross-lingual Capabilities of English Pretrained Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3563–74. https://doi.org/10.18653/v1/2022.emnlp-main.233.
Harvard
Blevins, T. and Zettlemoyer, L. (2022) “Language Contamination Helps Explain the Cross-lingual Capabilities of English Pretrained Models”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3563–3574. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.233.
Vancouver
1. Blevins T, Zettlemoyer L (2022) Language Contamination Helps Explain the Cross-lingual Capabilities of English Pretrained Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3563–3574

BibTeX

@inproceedings{blevins-zettlemoyer-2022-language,
    title = "Language Contamination Helps Explain the Cross-lingual Capabilities of {E}nglish Pretrained Models",
    author = "Blevins, Terra  and
      Zettlemoyer, Luke",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.233/",
    doi = "10.18653/v1/2022.emnlp-main.233",
    pages = "3563--3574"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/