An Open Dataset and Model for Language Identification
Laurie BurchellAlexandra BirchNikolay BogoychevKenneth Heafield
Presents an open, manually verified dataset of 121 million lines alongside a fastText language identification model that outperforms existing systems across 201 languages.
Automatic language identification is an essential first step in modern language processing, used to filter web data and curate multilingual datasets. However, existing language identification systems frequently suffer from poor real-world accuracy, particularly on lower-resource languages where noisy data degrades downstream tools and exaggerates apparent progress. Furthermore, scalable, high-coverage language identification tools typically keep their training data private and undocumented, limiting transparency and reproducibility.
The article demonstrates that training a language identification classifier on a carefully audited, open dataset substantially improves performance across a wide spectrum of languages. The authors evaluate their approach by creating an openly accessible monolingual corpus covering 201 languages and training a lightweight classifier to benchmark against established industry models.
To ensure reliable labels, the authors assembled 121 million lines of text primarily from trusted sources—such as news outlets, Wikipedia, and religious texts—intentionally avoiding unverified web crawls. Two authors manually audited samples across all languages and scripts to standardize labels and verify integrity. Using this curated data, they trained an open-source fasttext model across 201 languages, implementing an upsampling strategy to balance high- and low-resource classes, and evaluated it against existing open systems on a multi-language translation test suite.
The model achieved a strong macro-average F1 score of 0.927 and a low false positive rate of 0.033 across all 201 languages. When evaluated on shared subsets of languages, it consistently outperformed existing open baselines, including NLLB and CLD3, achieving an F1 score of 0.959 across 193 languages compared to NLLB's 0.950, and 0.989 across 95 shared high-resource languages compared to CLD3's 0.968. Across linguistic resource categories, the model maintained higher or equal accuracy and lower error rates, even in the most data-constrained language tiers. However, the evaluation also revealed that closely related language varieties, such as Arabic dialects and Yue Chinese versus Traditional Chinese, remain major failure points due to label ambiguity and domain mismatch.
These results demonstrate that data quality and label curation are far more critical than raw model complexity, as the model matched the architecture of existing baselines while attaining superior performance solely through vetted data. High precision and low false positive rates are vital for reducing the risk of data contamination and computational waste in production environments. Nevertheless, users must recognize that automated classifiers struggle with closely related dialects and formal versus colloquial domain shifts, which can lead to misclassification risks if treated as infallible black boxes.
Organizations and practitioners building multilingual systems should adopt transparent, verified datasets rather than unvetted web corpora, and they should openly report performance metrics across individual language classes. Before deploying language classifiers in operational settings, teams should conduct domain-specific pilot testing and implement specialized strategies when handling linguistically close varieties or dialects. Future research should expand manual audits using native speakers, incorporate data beyond formal domains, and develop evaluation benchmarks that mirror realistic web noise.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). Read NLLB first to understand the multilingual translation system and evaluation context that this paper uses as a principal benchmark.
No sufficiently relevant recommendations were found.
