Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models
Terra BlevinsLuke Zettlemoyer
Demonstrates that supposedly monolingual English pretraining corpora contain hundreds of millions of foreign-language tokens, revealing that unintentional multilingual contamination largely drives the zero-shot cross-lingual transfer capabilities of English language models.
English pretrained language models serve as the foundation for modern natural language processing systems. Although these models are nominally trained on English-only data, recent research observed that they transfer surprisingly well to other languages without prior multilingual training. The article investigates why this occurs, demonstrating that standard English pretraining datasets contain substantial amounts of foreign language data and that this unintentional contamination directly enables cross-lingual performance.
To evaluate the scale and impact of this leakage, the article conducted a two-part study. First, it performed automatic language identification alongside qualitative audits across six major English pretraining datasets, including web-crawled collections like C4 and CC-News. Second, it evaluated popular English models (BERT, RoBERTa, and T5) against dedicated multilingual models across more than 50 languages using language modeling and part-of-speech tagging tasks, testing both frozen models and finetuned setups.
The findings show that all evaluated English corpora contain notable foreign text, ranging from 300,000 to over 400 million non-English tokens. Web-crawled datasets exhibited the highest contamination, even when automated language filters were used. Crucially, downstream cross-lingual performance correlated strongly with the amount of in-language text encountered during pretraining (reaching correlations up to 0.67 to 0.68 for RoBERTa), whereas syntactic similarity to English showed much weaker correlation. When fine-tuned for part-of-speech tagging, RoBERTa narrowed its performance gap with dedicated multilingual models to just 2.65 points. Additionally, English T5 outperformed multilingual BERT on part-of-speech tagging for certain languages, such as German and Portuguese, without any task-specific tuning.
These results indicate that large-scale English models are functionally multilingual rather than strictly monolingual. Their cross-lingual capabilities do not reflect true "zero-shot" generalization, but rather direct learning from leaked pretraining data. This means organizations evaluating language models cannot assume clean linguistic separation, which affects how cross-lingual transfer benchmarks and multilingual risks are interpreted.
Moving forward, researchers and developers should account for pretraining contamination when evaluating model capabilities and avoid assuming that automated data cleaning eliminates foreign text. Complete manual data filtering remains practically infeasible at scale, but auditing training data composition will improve the transparency of multilingual evaluations. Limitations of the article include reliance on automated classifiers for language estimates, an evaluation focused primarily on syntax and language modeling, and an analysis restricted to English-centric corpora.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). This paper introduces the C4 dataset and the T5 model, which serve as foundational web-crawled pretraining corpora and architectures audited for cross-lingual contamination.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This study establishes the baseline evaluation methods and unexpected zero-shot cross-lingual transfer capabilities in pretrained models that the source paper seeks to explain through training data contamination.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). It provides crucial context on scaling multilingual pretraining over massive CommonCrawl datasets, providing the standard multilingual baselines compared against English models.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It introduces the multilingual counterpart to T5 and the mC4 dataset, which the source paper directly compares against English-trained models across cross-lingual tasks.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This foundational work outlines cross-lingual pretraining objectives and zero-shot transfer methodologies that form the conceptual background of the source study.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). It introduces the deep bidirectional Transformer pretraining paradigm and the standard English BERT models evaluated for data contamination in the source article.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It expands the evaluation of cross-lingual capabilities and zero-shot discrepancies across generative AI models on a broader multilingual benchmark across 70 languages.
- Paper: Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages, Wietse de Vries et al. (2022). It extends the study of cross-lingual part-of-speech transfer to an exhaustive setup spanning over 100 source and target languages.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). It advances beyond incidental contamination by deliberately scaling multilingual pretraining corpora and models to 500 under-resourced languages.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). It directly tackles the data quality and filtering challenges identified in web-crawled corpora by systematically testing curation pipelines on hundreds of trillions of tokens.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). It investigates the broader implications of web-text filtering versus scale in pretraining corpora, offering a critical perspective on the limits of automated data cleaning.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). It benchmarks the multilingual competencies and gaps of modern large language models across 83 languages and diverse task suites.
