An Open Dataset and Model for Language Identification

Laurie BurchellAlexandra BirchNikolay BogoychevKenneth Heafield

article2023ACL53 citations

Presents an open, manually verified dataset of 121 million lines alongside a fastText language identification model that outperforms existing systems across 201 languages.

Listen

Automatic language identification is an essential first step in modern language processing, used to filter web data and curate multilingual datasets. However, existing language identification systems frequently suffer from poor real-world accuracy, particularly on lower-resource languages where noisy data degrades downstream tools and exaggerates apparent progress. Furthermore, scalable, high-coverage language identification tools typically keep their training data private and undocumented, limiting transparency and reproducibility.

The article demonstrates that training a language identification classifier on a carefully audited, open dataset substantially improves performance across a wide spectrum of languages. The authors evaluate their approach by creating an openly accessible monolingual corpus covering 201 languages and training a lightweight classifier to benchmark against established industry models.

To ensure reliable labels, the authors assembled 121 million lines of text primarily from trusted sources—such as news outlets, Wikipedia, and religious texts—intentionally avoiding unverified web crawls. Two authors manually audited samples across all languages and scripts to standardize labels and verify integrity. Using this curated data, they trained an open-source fasttext model across 201 languages, implementing an upsampling strategy to balance high- and low-resource classes, and evaluated it against existing open systems on a multi-language translation test suite.

The model achieved a strong macro-average F1 score of 0.927 and a low false positive rate of 0.033 across all 201 languages. When evaluated on shared subsets of languages, it consistently outperformed existing open baselines, including NLLB and CLD3, achieving an F1 score of 0.959 across 193 languages compared to NLLB's 0.950, and 0.989 across 95 shared high-resource languages compared to CLD3's 0.968. Across linguistic resource categories, the model maintained higher or equal accuracy and lower error rates, even in the most data-constrained language tiers. However, the evaluation also revealed that closely related language varieties, such as Arabic dialects and Yue Chinese versus Traditional Chinese, remain major failure points due to label ambiguity and domain mismatch.

These results demonstrate that data quality and label curation are far more critical than raw model complexity, as the model matched the architecture of existing baselines while attaining superior performance solely through vetted data. High precision and low false positive rates are vital for reducing the risk of data contamination and computational waste in production environments. Nevertheless, users must recognize that automated classifiers struggle with closely related dialects and formal versus colloquial domain shifts, which can lead to misclassification risks if treated as infallible black boxes.

Organizations and practitioners building multilingual systems should adopt transparent, verified datasets rather than unvetted web corpora, and they should openly report performance metrics across individual language classes. Before deploying language classifiers in operational settings, teams should conduct domain-specific pilot testing and implement specialized strategies when handling linguistically close varieties or dialects. Future research should expand manual audits using native speakers, incorporate data beyond formal domains, and develop evaluation benchmarks that mirror realistic web noise.

No sufficiently relevant recommendations were found.

Cover for An Open Dataset and Model for Language Identification

Abstract

Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a fasttext LID model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033 across 201 languages, outperforming previous work. We achieve this by training on a curated dataset of monolingual data, which we audit manually to ensure reliability. We make both the model and the dataset available to the research community. Finally, we carry out detailed analysis into our model’s performance, both in comparison to existing open models and by language class. For applications such as corpus filtering, LID systems need to be fast, reliable, and cover as many languages as possible. There are several open LID models offering quick classification and high language coverage, such as CLD3 or the work of Costa-jussà et al. (2022). However, to the best of our knowledge, none of the commonly-used scalable LID systems make their training data public. We address this gap by releasing an open and curated dataset for LID, which we audit by hand to assure quality. We also train a high-performing LID model on this dataset to show its utility.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Dataset
  • 3.1 Data sources
  • 3.1.1 Language selection
  • 3.2 Manual audit process
  • 3.3 Preprocessing
  • 3.4 Dataset description
  • 4 Model and hardware
  • 5 Evaluation
  • 5.1 Test sets
  • 5.2 Other LID systems
  • 6 Results and analysis
  • 6.1 Performance by language category
  • 6.2 Case study: Chinese languages
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Data sources
  • B Performance of our LID model by language

Knowls

  1. Knowl 1 — Curated multilingual training dataset

    data/table

    The released language-identification dataset contains 121 million lines of monolingual text in 201 language classes. The authors favored sources whose collection process made language labels likely to be reliable, rather than web-crawled data labeled by other language-identification systems. Sources include news, Wikipedia, religious text, and material from other domains such as literature, social media, and transcribed conversations; named sources include Leipzig Corpora Collection, NLLB Seed, WiLI, XL-Sum, and Arabic dialect corpora. The data was selected to be open for research or described by its source as free to use. The authors report a pre-sampling mean of 602,812 lines per language, with South Azerbaijani the smallest class at 532 lines and English the largest at 7.5 million lines. Much of the text is formal in style.

  2. Knowl 2 — FastText language-identification model and training configuration

    model/method

    The authors trained a fastText classifier that represents input using character-level n-grams and feeds those representations to a multiclass linear classifier. Training used softmax loss for two epochs, a learning rate of 0.8, and embedding dimension 256. Character n-grams ranged from length 2 to 5; the word n-gram setting was 1, the bucket size was 1,000,000, and words occurring fewer than 1,000 times after upsampling were discarded. Training used 68 threads on one Ice Lake node with 76 CPUs and 256 GiB of RAM; it took about 1 hour 45 minutes and produced a 60.5-million-parameter model. Inference on 206,448 test lines took 22.4 seconds, or 9,216.4 lines per second.

  3. Knowl 3 — Macro-averaged performance against open LID systems

    empirical result

    Evaluation used the FLORES-200 dev-test sentences, professionally translated or transliterated from Wikimedia articles and human-verified for language. The authors removed the three language varieties excluded from their dataset, yielding a 201-language test set. They report macro-averaged F1 and false-positive rate (FPR). Comparisons use each model’s supported languages: NLLB comparison uses the 193 languages shared with the authors’ model after removing eight Arabic dialects, while the three-way comparison uses 95 languages after normalizing CLD3 predictions to BCP-47 macrolanguage codes.

    SystemLanguages supportedAll 201: F1All 201: FPRShared 193 with NLLB: F1Shared 193 with NLLB: FPRShared 95: F1Shared 95: FPR
    CLD3107————0.9680.030
    NLLB218——0.9500.0230.9850.019
    Authors’ model2010.9270.0330.9590.0200.9890.011

    The model therefore has the strongest F1 and lowest FPR among the compared systems on the shared test subsets, and achieves F1 0.927 with FPR 0.033 over its full 201-language test set. The authors note that their model and NLLB share an architecture and parameters, and attribute their model’s advantage to training-data selection and manual auditing; the comparison does not isolate those factors experimentally.

  4. Knowl 4 — Performance across language-resource categories

    data/table

    The authors compared their classifier with NLLB by language-resource class using the taxonomy of Joshi et al., where class 0 is least resourced and class 5 is best resourced. The comparison excludes five languages absent from the taxonomy and eight Arabic dialects not covered by NLLB. The classifier’s mean F1 is at least as high in every category, and its mean FPR is lower in every category except class 0.

    Resource classLanguagesOur model mean F1NLLB mean F1Our model mean FPRNLLB mean FPR
    0280.9000.8970.0140.013
    1940.9810.9680.0130.013
    2160.9900.9630.0090.043
    3250.9830.9740.0070.013
    4180.9510.9510.0510.055
    570.8970.8550.1630.620

    These are averages over languages within each category, rather than sentence-weighted scores. The large class-5 FPR values for both systems also show that the category-level comparison does not imply uniformly low error.

  5. Knowl 5 — Sample-based manual audit and language-label curation

    model/method

    To check dataset quality, the authors manually audited random samples from every data source and language, rather than verifying every line. They first standardized language labels using the codes adopted by Costa-jussà et al. and treated macrolanguage or ambiguous labels cautiously, assigning them to a dominant individual language when appropriate. They report that mapping Malay macrolanguage data directly to Standard Malay caused a large performance drop, because the macrolanguage covered a diverse set of languages.

    Two authors performed the audit: one was a native Bulgarian speaker able to read Cyrillic, Latin, and Chinese characters; the other was a native English speaker able to read Latin, Arabic, and Hebrew scripts. For languages they knew, they checked whether samples matched expectations. For unfamiliar languages in readable scripts, they compared samples with the language’s Universal Declaration of Human Rights text, or with Wikipedia text when that was unavailable, inspecting features such as diacritics, common words, word lengths, prefixes and suffixes, loan words, and vowel/consonant patterns. For scripts they could not read, they checked that sample lines used a consistent script and that it matched the script in a reference sample. The authors state that more native-speaker verification, especially for the least-resourced languages, is needed.

  6. Knowl 6 — Language coverage decisions and exclusions

    definition

    The dataset targets the language coverage of the FLORES-200 evaluation benchmark but contains 201 language classes after curation. The authors excluded Akan because it is a macrolanguage that includes Twi, whereas the other retained benchmark classes were individual languages. They excluded Modern Standard Arabic in Latin script because they could not find naturally occurring training data for that combination: Arabic text in Latin characters was likely to be dialectal rather than Modern Standard Arabic. They also excluded Minangkabau in Arabic script because it is now rarely written in that script, making useful training data difficult to obtain.

  7. Knowl 7 — Language-agnostic preprocessing and class sampling

    model/method

    Preprocessing was kept minimal: Moses scripts removed non-printing characters and detokenized data where necessary, and a line was retained only if it contained at least one character in the expected script, as defined by Perl. This script check was intended to permit borrowings while filtering lines inconsistent with a language’s expected script. To mitigate class imbalance, the authors sampled language data in proportion to pl0.3p_l^{0.3}, where plp_l is the fraction of dataset lines belonging to language ll. The exponent reduces the influence of the original class-frequency differences without making the sampling uniform.

  8. Knowl 8 — Yue Chinese exposes label and domain ambiguity

    empirical result

    On the FLORES-200 test set, the authors’ classifier achieved F1 0.0059 for Yue Chinese, compared with 0.4877 for NLLB. For the two Chinese varieties that the authors’ model handled well, it almost never predicted Yue: it assigned nearly all true Yue examples to Chinese (Traditional). NLLB also confused Yue and Chinese (Traditional), but distributed its predictions between them.

    Four native speakers inspecting the training data and test set identified a likely domain mismatch: much of the Yue training material was colloquial, while the test material was formal. They also could not confidently distinguish Yue from Chinese (Traditional) in formal writing because the varieties are very similar in that form. The case illustrates that a language label may cover text whose variety differs across datasets, and that closely related written varieties can be difficult to separate reliably.

  9. Knowl 9 — Training-set size has little per-language correlation with test F1

    empirical result

    Across languages, the Pearson correlation between the number of training lines and the authors’ model’s F1 score was 0.0242. Thus, in this evaluation, training-set size alone had almost no linear association with per-language F1. The authors suggest that some low-resource languages scored perfectly because their training and test data had substantial domain overlap, while higher-resource languages could score lower on the test domain despite potentially having greater robustness across domains. The result is specific to the reported data and test set and does not establish that training volume is generally unimportant.

  10. Knowl 10 — Scope and fairness limitations

    limitation

    The dataset and model cover only 201 languages, limited by the languages available in the FLORES-200 evaluation benchmark. Because the test set consists of sentences from Wikimedia articles in a single domain, its scores may not represent performance on web data or other domains. Most training data was not audited by native speakers, which the authors identify as a quality limitation. The authors also note that language identification is normative: its labels and coverage choices can exclude minority dialects, scripts, or microlanguages, and may reinforce power imbalances. Since performance is uneven across languages, errors can also worsen downstream outcomes for particular groups; the authors provide class-level metrics to make some of these disparities visible.

Coverage note — The exhaustive language-by-language training counts and scores, and the complete inventory of individual source corpora, are omitted because they are lengthy catalogues rather than distinct analyses; the dataset’s source strategy and principal performance patterns are summarized here.

References

  1. 1.Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018. Shami: A corpus of Levantine Arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  2. 2.Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Alcoba Inciarte. 2022. Afrolid: A neural language identification tool for african languages. arXiv preprint arXiv:2210.11744.
  3. 3.Željko Agic and Ivan Vulić. 2019. JW300: A wide-coverage parallel corpus for low-resource languages. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204–3210, Florence, Italy. Association for Computational Linguistics.
  4. 4.Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018. DART: A large dataset of dialectal Arabic tweets. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  5. 5.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al. 2019. Massively multilingual neural machine translation in the wild: Findings and challenges. arXiv preprint arXiv:1907.05019.
  6. 6.Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020. Findings of the 2020 conference on machine translation (WMT20). In Proceedings of the Fifth Conference on Machine Translation, pages 1–55, Online. Association for Computational Linguistics.
  7. 7.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 1–61, Florence, Italy. Association for Computational Linguistics.
  8. 8.Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013. Findings of the 2013 Workshop on Statistical Machine Translation. In Proceedings of the Eighth Workshop on Statistical Machine Translation, pages 1–44, Sofia, Bulgaria. Association for Computational Linguistics.
  9. 9.Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58, Baltimore, Maryland, USA. Association for Computational Linguistics.
  10. 10.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017. Findings of the 2017 conference on machine translation (WMT17). In Proceedings of the Second Conference on Machine Translation, pages 169–214, Copenhagen, Denmark. Association for Computational Linguistics.
  11. 11.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 131–198, Berlin, Germany. Association for Computational Linguistics.
  12. 12.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015. Findings of the 2015 workshop on statistical machine translation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 1–46, Lisbon, Portugal. Association for Computational Linguistics.
  13. 13.Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018. Findings of the 2018 conference on machine translation (WMT18). In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 272–303, Belgium, Brussels. Association for Computational Linguistics.
  14. 14.Houda Bouamor, Sabit Hassan, and Nizar Habash. 2019. The MADAR shared task on Arabic fine-grained dialect identification. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 199–207, Florence, Italy. Association for Computational Linguistics.
  15. 15.Ralf Brown. 2014. Non-linear mapping for improved identification of 1300+ languages. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 627–632, Doha, Qatar. Association for Computational Linguistics.
  16. 16.Ralf D Brown. 2012. Finding and identifying text in 900+ languages. Digital Investigation, 9:S34–S43.
  17. 17.Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6588–6608, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  18. 18.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv preprint arXiv:2207.04672.
  19. 19.Jonathan Dunn. 2020. Mapping languages: The corpus of global language use. Language Resources and Evaluation, 54(4):999–1018.
  20. 20.Mahmoud El-Haj, Paul Rayson, and Mariam Aboelezz. 2018. Arabic dialect identification in the context of bivalency and code-switching. In Proceedings of the 11th International Conference on Language Resources and Evaluation, Miyazaki, Japan., pages 3622–3627. European Language Resources Association.
  21. 21.Miquel Esplà, Mikel Forcada, Gema Ramírez-Sánchez, and Hieu Hoang. 2019. ParaCrawl: Web-scale parallel corpora for the languages of the EU. In Proceedings of Machine Translation Summit XVII: Translator, Project and User Tracks, pages 118–119, Dublin, Ireland. European Association for Machine Translation.
  22. 22.Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012. Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 759–765, Istanbul, Turkey. European Language Resources Association (ELRA).
  23. 23.Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2022. Can we use word embeddings for enhancing Guarani-Spanish machine translation? In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 127–132, Dublin, Ireland. Association for Computational Linguistics.
  24. 24.Thamme Gowda, Zhao Zhang, Chris Mattmann, and Jonathan May. 2021. Many-to-English machine translation tools, data, and pretrained models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 306–316, Online. Association for Computational Linguistics.
  25. 25.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  26. 26.Rudali Huidrom, Yves Lepage, and Khogendra Khomdram. 2021. EM corpus: a comparable corpus for a less-resourced language pair Manipuri-English. In Proceedings of the 14th Workshop on Building and Using Comparable Corpora (BUCC 2021), pages 60–67, Online (Virtual Mode). INCOMA Ltd.
  27. 27.Tommi Jauhiainen, Marco Lui, Marcos Zampieri, Timothy Baldwin, and Krister Lindén. 2019. Automatic language identification in texts: A survey. Journal of Artificial Intelligence Research, 65:675–782.
  28. 28.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  29. 29.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431, Valencia, Spain. Association for Computational Linguistics.
  30. 30.Omid Kashefi. 2018. Mizan: A large persian-english parallel corpus. arXiv preprint arXiv:1801.02107.
  31. 31.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions, pages 177–180, Prague, Czech Republic. Association for Computational Linguistics.
  32. 32.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  33. 33.Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018. The IIT Bombay English-Hindi parallel corpus. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  34. 34.Kang Kwong Luke and May LY Wong. 2015. The hong kong cantonese corpus: design and uses. Journal of Chinese Linguistics Monograph Series, 1(25):312–333.
  35. 35.Salima Medhaffar, Fethi Bougares, Yannick Estève, and Lamia Hadrich-Belguith. 2017. Sentiment analysis of Tunisian dialects: Linguistic ressources and experiments. In Proceedings of the Third Arabic Natural Language Processing Workshop, pages 55–61, Valencia, Spain. Association for Computational Linguistics.
  36. 36.Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015. Machine translation experiments on PADIC: A parallel Arabic DIalect corpus. In Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation, pages 26–34, Shanghai, China.
  37. 37.Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, and Sriram Chellappan. 2021. A large-scale study of machine translation in Turkic languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5876–5890, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  38. 38.Atul Kr Ojha. 2019. English-bhojpuri smt system: Insights from the karaka model. arXiv preprint arXiv:1905.02239.
  39. 39.Mohammad Taher Pilevar, Heshaam Faili, and Abdol Hamid Pilevar. 2011. Tep: Tehran english-persian parallel corpus. In International Conference on Intelligent Text Processing and Computational Linguistics, pages 68–79. Springer.
  40. 40.Matt Post, Chris Callison-Burch, and Miles Osborne. 2012. Constructing parallel corpora for six Indian languages via crowdsourcing. In Proceedings of the Seventh Workshop on Statistical Machine Translation, pages 401–409, Montréal, Canada. Association for Computational Linguistics.
  41. 41.Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018. When and why are pre-trained word embeddings useful for neural machine translation? In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 529–535, New Orleans, Louisiana. Association for Computational Linguistics.
  42. 42.Roberts Rozis and Raivis Skadin,š. 2017. Tilde MODEL - multilingual open data for EU languages. In Proceedings of the 21st Nordic Conference on Computational Linguistics, pages 263–265, Gothenburg, Sweden. Association for Computational Linguistics.
  43. 43.Martin Thoma. 2018. The wili benchmark dataset for written language identification. arXiv preprint arXiv:1801.07779.
  44. 44.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2214–2218, Istanbul, Turkey. European Language Resources Association (ELRA).
  45. 45.Jihad Zahir. 2022. Iadd: An integrated arabic dialect identification dataset. Data in Brief, 40:107777.
  46. 46.Omar F. Zaidan and Chris Callison-Burch. 2011. The Arabic online commentary dataset: an annotated dataset of informal Arabic with high dialectal content. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 37–41, Portland, Oregon, USA. Association for Computational Linguistics.
  47. 47.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1628–1639, Online. Association for Computational Linguistics.
  48. 48.Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The United Nations parallel corpus v1.0. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3530–3534, Portorož, Slovenia. European Language Resources Association (ELRA).

Citation

MLA
Burchell, L., et al. “An Open Dataset and Model for Language Identification”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 865–79, https://doi.org/10.18653/v1/2023.acl-short.75.
APA
Burchell, L., Birch, A., Bogoychev, N., & Heafield, K. (2023). An Open Dataset and Model for Language Identification. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 865–879. https://doi.org/10.18653/v1/2023.acl-short.75
Chicago
Burchell, L., A. Birch, N. Bogoychev, and K. Heafield. 2023. “An Open Dataset and Model for Language Identification”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 865–79. https://doi.org/10.18653/v1/2023.acl-short.75.
Harvard
Burchell, L. et al. (2023) “An Open Dataset and Model for Language Identification”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 865–879. Available at: https://doi.org/10.18653/v1/2023.acl-short.75.
Vancouver
1. Burchell L, Birch A, Bogoychev N, Heafield K (2023) An Open Dataset and Model for Language Identification. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 865–879

BibTeX

@inproceedings{burchell-etal-2023-open,
    title = "An Open Dataset and Model for Language Identification",
    author = "Burchell, Laurie  and
      Birch, Alexandra  and
      Bogoychev, Nikolay  and
      Heafield, Kenneth",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.75/",
    doi = "10.18653/v1/2023.acl-short.75",
    pages = "865--879"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/