DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages

Fahim FaisalOrevaoghene AhiaAarohi SrivastavaKabir AhujaDavid ChiangYulia TsvetkovAntonios Anastasopoulos

article2024ACL69 citationsBest Social Impact Paper Award

Introduces DIALECTBENCH, a comprehensive evaluation suite spanning 10 NLP tasks and 281 language varieties across 40 clusters to quantify and analyze performance disparities between standard and non-standard dialects.

Listen

Natural language processing models are increasingly deployed worldwide, yet standard evaluation benchmarks focus almost exclusively on high-resource, standardized national languages. This practice overlooks regional dialects, non-standard varieties, and closely related low-resource languages, hiding significant performance failures for non-standard speakers. The article introduces DialectBench to systematically evaluate, benchmark, and analyze natural language processing performance disparities between standard languages and their non-standard dialectal variants across diverse tasks.

To establish this benchmark, the article aggregated diverse datasets across 10 text-level tasks—including structured prediction, text classification, question answering, and machine translation—covering 281 language varieties grouped into 40 language clusters using genealogical linguistic standards. The evaluation benchmarked standard multilingual pre-trained encoder models alongside large language models and machine translation systems across varied settings, such as zero-shot cross-lingual transfer and localized fine-tuning.

Key findings demonstrate severe performance disparities across language varieties. First, models achieve high accuracy on standard and Germanic or Romance varieties (often exceeding 90% F1 or parsing scores) but drop drastically on low-resource indigenous or regional dialects, falling below 10% in tasks like dependency parsing for Mbyá Guaraní. Second, fine-tuning directly on dialectal data exacerbates within-cluster performance divergence compared to zero-shot transfer because training data across dialects is highly inconsistent in volume and quality. Third, writing systems heavily influence transferability; low-resource varieties using the Latin script benefit significantly more from English zero-shot transfer than non-Latin varieties. Finally, while large language models evaluated via few-shot prompting outperformed zero-shot transfer baselines on select dialect tasks, they consistently lagged behind dedicated fine-tuned models.

These results demonstrate that current language systems carry substantial risks of technological exclusion and operational failure when deployed in diverse linguistic environments. Standard multilingual evaluation scores overestimate model readiness by masking severe within-cluster inequalities. Organizations relying on standard benchmarks face compliance, user adoption, and accuracy risks in multilingual markets.

Moving forward, researchers and technology practitioners should use dialect-aware evaluation frameworks before deploying multilingual systems. Development should focus on gathering curated parallel dialect resources, standardizing data quality, and expanding evaluations into speech-based systems. Leaders must account for both demographic utility (speaker population coverage) and linguistic utility (equal performance across variants) when auditing system fairness.

These conclusions are bounded by data scarcity constraints, domain inconsistencies across collated datasets, and limited human-generated dialectal translation references, which necessitated synthetic pseudo-references for translation evaluations. Nevertheless, the benchmark provides robust evidence of systematic dialect disparities across standard language models.

arXiv: 2403.11009ffaisal93/DialectBench

No sufficiently relevant recommendations were found.

Cover for DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages

Abstract

Language technologies should be judged on their usefulness in real-world use cases. An often overlooked aspect in natural language processing (NLP) research and evaluation is language variation in the form of non-standard dialects or language varieties (hereafter, varieties). Most NLP benchmarks are limited to standard language varieties. To fill this gap, we propose DIALECTBENCH, the first-ever large-scale benchmark for NLP on varieties, which aggregates an extensive set of task-varied variety datasets (10 text-level tasks covering 281 varieties). This allows for a comprehensive evaluation of NLP system performance on different language varieties. We provide substantial evidence of performance disparities between standard and non-standard language varieties, and we also identify language clusters with larger performance divergence across tasks. We believe DIALECTBENCH provides a comprehensive view of the current state of NLP for language varieties and one step towards advancing it further. 1

Table of Contents

  • 1 Introduction
  • 2 DIALECTBENCH
  • 3 Experiments
  • 3.1 Models
  • 3.2 Training and Evaluation
  • 3.3 Quantifying the Dialectal Gap
  • 4 Results
  • 4.1 Maximum Obtainable Scores
  • 4.2 Dialectal Gap Across Language Clusters
  • 5 Discussion
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • Appendix
  • A Related Work
  • B Tasks of DIALECTBENCH
  • C Varieties and Clusters of DIALECTBENCH
  • C.1 DIALECTBENCH Variety list
  • C.2 Language Clusters and Representative Varieties
  • D Result Visualizations
  • D.1 Regional maps with aggregated Machine Translation scores
  • F Highest performing and lowest performing varieties
  • G Low-resource variety performing better in zero-shot NER
  • H Cluster level Result Summaries with Demographic Utility and Standard Deviation report
  • I In-Context Learning Details
  • I.1 Prompts

Knowls

  1. Knowl 1 — DIALECTBENCH unifies evaluation across language varieties

    model/method

    DIALECTBENCH is a benchmark for evaluating text-based NLP systems on standard and non-standard varieties, including closely related languages and some variation by region, register, mode, or writing system. It aggregates existing datasets into 40 language clusters and covers 281 varieties across 10 NLP tasks. Its evaluation pipeline brings together dataset selection, cluster and variety assignment, task-specific baselines, and comparisons of performance across and within clusters.

  2. Knowl 2 — Glottolog-based mapping organizes varieties into clusters

    model/method

    DIALECTBENCH uses Glottolog to group related varieties into clusters whose members map to a common linguistic ancestor and a corresponding phylogenetic subtree. Cluster membership is intended to reflect close relations such as mutual intelligibility, phylogenetic similarity, or geographic proximity; it is not based on a numerical measure of those relations. Varieties are identified primarily with Glottocodes. Where a Glottocode is unavailable, the benchmark can use an ancestor code and distinguish the variety with metadata for area, register, mode, orthography, or dataset-specific identifier. For each task and cluster, a representative is generally selected from the highest-resource variety, often the variety with the largest speaker population; this choice can vary by task and does not imply that the representative is linguistically superior or more standardized.

  3. Knowl 3 — Task and dataset coverage spans ten NLP tasks

    experimental setup

    DIALECTBENCH combines datasets for the following tasks and evaluates them with task-appropriate metrics: dependency parsing (UAS; Universal Dependencies, TwitterAAE, and Singlish); part-of-speech tagging (F1; Universal Dependencies, Singlish, and Noisy Dialects); named entity recognition (F1; Wikiann and Norwegian NER); dialect identification (F1; MADAR, DMT, Greek, DSL-TL, and Swiss Germans); sentiment analysis (F1; TSAC, TUNIZI, DzSentiA, SaudiBank, MAC, ASTD, AJGT, and OCLAR); topic classification (F1; SIB-200); natural language inference (F1; XNLI with a translated test set); multiple-choice machine reading comprehension (F1; Belebele); extractive question answering (span F1; SD-QA); and machine translation (BLEU; CODET and TIL-MT). The benchmark mostly retains datasets in their published form, apart from variety renaming and specified label or format adjustments.

  4. Knowl 4 — Baselines use multiple data-availability-dependent training settings

    experimental setup

    For text tasks other than machine translation, DIALECTBENCH evaluates mBERT and XLM-R; machine translation uses NLLB in 600M- and 1.3B-parameter variants. Mistral 7B is also evaluated through in-context prompting on sentiment analysis and extractive question answering. The baseline procedures depend on available data: in-variety fine-tuning trains on the evaluated variety; in-cluster fine-tuning trains on the cluster's highest-resource representative and tests on other cluster members; combined fine-tuning uses data from multiple standard varieties for question answering; and zero-shot transfer fine-tunes on English and tests on another variety. Dialect identification is trained on variety-labeled data within each cluster. The benchmark reports these baselines to characterize existing performance rather than to optimize each task for its best possible model.

  5. Knowl 5 — Translate-test NLI adds translated evaluation data for 40 varieties

    data/table

    Because the authors found no existing NLI dataset covering the varieties they target, they construct a translate-test evaluation set from the English test set of XNLI. NLLB-200 3B translates the English test material into 40 varieties across 12 language clusters; the dataset statistics report 5,010 sentences per variety. An NLI model fine-tuned on English is then evaluated zero-shot on the translated test data, with F1 as the metric. This construction expands dialect-focused NLI evaluation while keeping the English test set as the source of the translated examples.

  6. Knowl 6 — Machine translation is scored against system-generated pseudo-references

    model/method

    For dialect-to-English machine translation, reference translations are often unavailable. DIALECTBENCH therefore uses a pseudo-reference protocol: for an input sentence xx in a variety, let xˉ\bar{x} be its translation into the corresponding standard variety; let yy be the MT system's output for xx, and let yˉ\bar{y} be its output for xˉ\bar{x}. The system output yˉ\bar{y} serves as the pseudo-reference for evaluating yy, for example with BLEU. The evaluation is zero-shot in the translation direction: the system is applied to standard-variety-to-English inputs to establish the comparison and to dialect-variety-to-English inputs for the evaluated output.

  7. Knowl 7 — Relative dialectal gap measures performance loss against a chosen baseline

    equation

    Let St(v)S_t(v) be the score of a system trained on variety or training set tt and evaluated on variety vv, with higher scores indicating better performance. For baseline variety uu, the relative dialectal gap is Gt(u,v)=St(u)−St(v)St(u)G_t(u,v)=\frac{S_t(u)-S_t(v)}{S_t(u)}. For the in-variety fine-tuning comparison, where each variety is trained on its own data, DIALECTBENCH uses Gin-variety(u,v)=Su(u)−Sv(v)Su(u)G_{\mathrm{in\text{-}variety}}(u,v)=\frac{S_u(u)-S_v(v)}{S_u(u)}. The baseline uu is either the cluster's highest-resource variety vˉ\bar{v} or, in English zero-shot transfer, English. Thus, a positive value indicates lower performance on vv than on the baseline, while a negative value indicates higher performance on vv. Cluster-level gaps are computed by averaging the variety-level gaps: Gt(u,C)=1∣C∣∑v∈CGt(u,v)G_t(u,C)=\frac{1}{|C|}\sum_{v\in C}G_t(u,v), with the analogous average for the in-variety metric. The paper uses English-versus-variety gaps for global zero-shot disparity, representative-versus-variety gaps within zero-shot transfer, and representative-versus-variety gaps after fine-tuning.

  8. Knowl 8 — Best observed scores vary widely across tasks and varieties

    data/table

    The following summary reports each task's mean maximum-obtainable score across varieties, along with its highest- and lowest-scoring cluster/variety. A maximum-obtainable score is the best score reported for a variety regardless of evaluation method or training data; it is not a human-performance ceiling. Machine translation is shown separately for dialect-level and region-level evaluations.

    Task Clusters Varieties Mean score Highest score; lowest score
    Dependency parsing 16 40 64.3 Southwestern shifted romance / Brazilian Portuguese: 94.4; Tupi-Guarani subgroup / Mbyá Guaraní (Brazil): 9.0
    POS tagging 17 51 72.1 Norwegian / Norwegian Bokmål (written): 98.7; Tupi-Guarani subgroup / Mbyá Guaraní (Brazil): 1.9
    NER 27 85 70.1 Eastern Romance / Romanian: 94.2; Anglic / Jamaican Creole English: 0.0
    NLI 15 38 64.2 Anglic / English: 83.4; Sotho-Tswana / Southern Sotho: 34.6
    Topic classification 15 38 77.7 Sinitic / Classical-Middle-Modern Sinitic (traditional): 89.8; Kurdish / Central Kurdish: 19.4
    Dialect identification 6 49 67.0 Sinitic / Mandarin Chinese (Taiwan, simplified): 98.6; Southwestern shifted romance / Portuguese (written): 17.4
    Sentiment analysis 1 9 80.3 Arabic / Tunisian Arabic: 94.6; Arabic / South Levantine Arabic: 58.9
    Multiple-choice MRC 4 11 40.9 Anglic / English: 53.4; Sotho-Tswana / Southern Sotho: 29.0
    Extractive QA 5 24 74.2 Arabic / Arabic (Saudi Arabia): 77.9; Swahili / Swahili (Tanzania): 63.5
    MT, dialect-level 12 73 25.2 Arabic / Gulf Arabic (Riyadh): 43.1; Common Turkic / Sakha: 2.5
    MT, region-level 2 41 33.0 High German / Central Alemannic (Zurich): 44.1; Italian Romance / Italian (Sardegna): 13.0
  9. Knowl 9 — Zero-shot evaluation shows global and within-cluster disparities

    empirical result

    For four tasks, the average relative zero-shot gap from English across clusters was 34.0% for dependency parsing, 27.4% for POS tagging, 31.7% for named entity recognition, and 22.4% for topic classification. Comparing varieties instead with their own cluster's highest-resource representative yielded smaller average gaps in those tasks: 15.7%, 6.7%, 22.3%, and 12.5%, respectively. The authors report that low-resource clusters generally show larger gaps, while high-resource Germanic and Sinitic clusters tend to show smaller gaps, with exceptions such as substantial differences between Standard German and Swiss German. Within-cluster disparities can also be larger after fine-tuning: for dependency parsing, the average in-variety fine-tuning gap from the representative was 26.4%, compared with 15.7% for the zero-shot representative comparison. The results vary with the uneven quantity and quality of variety-specific training data. The authors also report that 77.2% of the top ten varieties by maximum score use Latin script, compared with 44.2% of the bottom ten; they observe that Latin-script low-resource varieties can benefit from English zero-shot transfer, while in-cluster fine-tuning can reduce this script-related effect.

  10. Knowl 10 — Benchmark comparisons are limited by data comparability and coverage

    limitation

    DIALECTBENCH's task and variety coverage, dataset quality, and training-data quantity differ substantially across tasks and varieties, so its comparisons are not based on uniformly comparable data. The authors identify parallel-corpus-based task data and translation-based comparable evaluation data as needed improvements, and note that the benchmark does not include every published dialect dataset. This iteration covers text-based tasks only, not speech-based NLP. Cluster representatives are selected for resource availability rather than linguistic superiority, and the closeness of varieties within a cluster was not quantified numerically. The authors also consciously avoid full-scale LLM evaluation because of uncertainty about data contamination.

Coverage note — Detailed per-variety result matrices and the specific mBERT, XLM-R, and Mistral score comparisons are omitted to prioritize the benchmark design, evaluation protocol, and aggregate disparity findings; these matrices provide task-specific baseline detail rather than a separate benchmark method.

References

  1. 1.Adel Abdelli, Fayçal Guerrouf, Okba Tibermacine, and Belkacem Abdelli. 2019. Sentiment analysis of Arabic Algerian dialect using a supervised method. In 2019 International Conference on Intelligent Systems and Advanced Computing Sciences (ISACS), pages 1–6.
  2. 2.Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021. ARBERT & MARBERT: Deep bidirectional transformers for Arabic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7088–7105, Online. Association for Computational Linguistics.
  3. 3.David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, and Annie En-Shiun Lee. 2023. SIB-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects.
  4. 4.Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023. Megaverse: Benchmarking large language models across languages, modalities, models and tasks.
  5. 5.Marwan Al Omari, Moustafa Al-Hajj, Nacereddine Hammami, and Amani Sabra. 2019. Sentiment classifier: Logistic regression for Arabic services’ reviews in Lebanon. In 2019 International Conference on Computer and Information Sciences (ICCIS), pages 1–5.
  6. 6.Md Mahfuz Ibn Alam, Sina Ahmadi, and Antonios Anastasopoulos. 2023. CODET: A benchmark for contrastive dialectal evaluation of machine translation.
  7. 7.Khaled Mohammad Alomari, Hatem M. ElSherif, and Khaled Shaalan. 2017. Arabic tweets sentimental analysis using machine learning. In Advances in Artificial Intelligence: From Theory to Practice, pages 602–610, Cham. Springer International Publishing.
  8. 8.Dhuha Alqahtani, Lama Alzahrani, Maram Bahareth, Nora Alshameri, Hend Al-Khalifa, and Luluh Aldhubayi. 2022. Customer sentiments toward Saudi banks during the Covid-19 pandemic. In Proceedings of the 5th International Conference on Natural Language and Speech Processing (ICNLSP 2022), pages 251–257, Trento, Italy. Association for Computational Linguistics.
  9. 9.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023. The Belebele benchmark: a parallel reading comprehension dataset in 122 language variants. arXiv preprint arXiv:2308.16884.
  10. 10.Verena Blaschke, Hinrich Schütze, and Barbara Plank. 2023. Does manipulating tokenization aid cross-lingual transfer? A study on POS tagging for non-standardized languages. In Proceedings of the Tenth Workshop on NLP for Similar Languages, Varieties and Dialects, pages 40–54, Dubrovnik, Croatia. Association for Computational Linguistics.
  11. 11.Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022. Systematic inequalities in language technology performance across the world’s languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5486–5505, Dublin, Ireland. Association for Computational Linguistics.
  12. 12.Su Lin Blodgett, Johnny Wei, and Brendan O’Connor. 2018. Twitter Universal Dependency parsing for African-American and mainstream American English. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1415–1425, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018. The MADAR Arabic dialect corpus and lexicon. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  14. 14.Jack K. Chambers and Peter Trudgill. 1998. Dialectology. Cambridge University Press.
  15. 15.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  16. 16.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  17. 17.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.AbdelRahim Elmadany, ElMoatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. ORCA: A challenging benchmark for Arabic language understanding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9559–9586, Toronto, Canada. Association for Computational Linguistics.
  20. 20.Fahim Faisal, Sharlina Keshava, Md Mahfuz Ibn Alam, and Antonios Anastasopoulos. 2021. SD-QA: Spoken dialectal question answering for the real world. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3296–3315, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  21. 21.Chayma Fourati, Hatem Haddad, Abir Messaoudi, Moez BenHajhmida, Aymen Ben Elhaj Mabrouk, and Malek Naski. 2021. Introducing a large Tunisian Arabizi dialectal dataset for sentiment analysis. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 226–230, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  22. 22.Moncef Garouani and Jamal Kharroubi. 2022. MAC: An open and free Moroccan Arabic corpus for sentiment analysis. In Innovations in Smart Cities Applications Volume 5, pages 849–858, Cham. Springer International Publishing.
  23. 23.Mika Hämäläinen, Khalid Alnajjar, Niko Partanen, and Jack Rueter. 2021. Finnish dialect identification: The effect of audio and text. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8777–8783, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.Harald Hammarström and Robert Forkel. 2022. Glottocodes: Identifiers linking families, languages and dialects to comprehensive reference information. Semantic Web Journal, 13(6):917–924.
  25. 25.Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. 2021. A survey on recent approaches for natural language processing in low-resource scenarios. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2545–2568, Online. Association for Computational Linguistics.
  26. 26.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multitask benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 4411–4421. PMLR.
  27. 27.Tommi Jauhiainen, Heidi Jauhiainen, and Krister Lindén. 2022. Italian language and dialect identification and regional French variety detection using adaptive naive Bayes. In Proceedings of the Ninth Workshop on NLP for Similar Languages, Varieties and Dialects, pages 119–129, Gyeongju, Republic of Korea. Association for Computational Linguistics.
  28. 28.Tommi Jauhiainen, Krister Lindén, and Heidi Jauhiainen. 2019. Discriminating between Mandarin Chinese and Swiss-German varieties using adaptive language models. In Proceedings of the Sixth Workshop on NLP for Similar Languages, Varieties and Dialects, pages 178–187, Ann Arbor, Michigan. Association for Computational Linguistics.
  29. 29.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7B. arXiv:2310.06825.
  30. 30.Bjarte Johansen. 2019. Named-entity recognition for Norwegian. In Proceedings of the 22nd Nordic Conference on Computational Linguistics, NoDaLiDa.
  31. 31.Anjali Kantharuban, Ivan Vulic, and Anna Korhonen. 2023. Quantifying the dialect gap and its correlates across languages.
  32. 32.Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020. GLUECoS: An evaluation benchmark for code-switched NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3575–3585, Online. Association for Computational Linguistics.
  33. 33.Heather Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen, Marcell Fekete, Esther Ploeger, Li Zhou, Hans Erik Heje, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, and Johannes Bjerva. 2023. CreoleVal: Multilingual multitask benchmarks for creoles.
  34. 34.Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, and Ming Zhou. 2020. XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6008–6018, Online. Association for Computational Linguistics.
  35. 35.Richard Littauer. n.d. Low resource languages: Resources for conservation, development, and documentation of low resource (human) languages. https://github.com/RichardLitt/low-resource-languages. [Accessed 15-11-2023].
  36. 36.Salima Medhaffar, Fethi Bougares, Yannick Estève, and Lamia Hadrich-Belguith. 2017. Sentiment analysis of Tunisian dialects: Linguistic ressources and experiments. In Proceedings of the Third Arabic Natural Language Processing Workshop, pages 55–61, Valencia, Spain. Association for Computational Linguistics.
  37. 37.Jamshidbek Mirzakhalov. 2021. Turkic Interlingua: A Case Study of Machine Translation in Low-resource Languages. Ph.D. thesis, University of South Florida.
  38. 38.Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, and Sriram Chellappan. 2021a. A large-scale study of machine translation in Turkic languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5876–5890, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  39. 39.Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr, et al. 2021b. A large-scale study of machine translation in turkic languages. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5876–5890.
  40. 40.Mahmoud Nabil, Mohamed Aly, and Amir Atiya. 2015. ASTD: Arabic sentiment tweets dataset. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2515–2519, Lisbon, Portugal. Association for Computational Linguistics.
  41. 41.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No Language Left Behind: Scaling human-centered machine translation. arXiv:2207.04672.
  42. 42.Sebastian Nordhoff. 2012. Linked data for linguistic diversity research: Glottolog/langdoc and asjp. In Christian Chiarcos, Sebastian Nordhoff, and Sebastian Hellmann, editors, Linked Data in Linguistics. Representing and Connecting Language Data and Language Metadata, pages 191–200. Springer, Heidelberg.
  43. 43.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  44. 44.Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Ji Yoon Han, Jangwon Park, Chisung Song, Junseong Kim, Youngsook Song, Taehwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seungwon Do, Sunkyoung Kim, Kyungtae Lim, Jongwon Lee, Kyumin Park, Jamin Shin, Seonghyun Kim, Lucy Park, Lucy Park, Alice Oh, Jung-Woo Ha (NAVER AI Lab), Kyunghyun Cho, and Kyunghyun Cho. 2021. KLUE: Korean language understanding evaluation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. Curran.
  45. 45.Alexandre Rademaker, Fabricio Chalub, Livy Real, Cláudia Freitas, Eckhard Bick, and Valeria de Paiva. 2017. Universal Dependencies for Portuguese. In Proceedings of the Fourth International Conference on Dependency Linguistics (Depling), pages 197–206, Pisa, Italy.
  46. 46.Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019. Massively multilingual transfer for NER. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 151–164, Florence, Italy. Association for Computational Linguistics.
  47. 47.Anthony Rios. 2020. FuzzE: Fuzzy fairness evaluation of offensive language classifiers on African-American English. Proceedings of the AAAI Conference on Artificial Intelligence, 34(01):881–889.
  48. 48.Sebastian Ruder, Jonathan H. Clark, Alexander Gutkin, Mihir Kale, Min Ma, Massimo Nicosia, Shruti Rijhwani, Parker Riley, Jean-Michel A. Sarr, Xinyi Wang, John Wieting, Nitish Gupta, Anna Katanova, Christo Kirov, Dana L. Dickinson, Brian Roark, Bidisha Samanta, Connie Tao, David I. Adelani, Vera Axelrod, Isaac Caswell, Colin Cherry, Dan Garrette, Reeve Ingle, Melvin Johnson, Dmitry Panteleev, and Partha Talukdar. 2023. XTREME-UP: A user-centric scarce-data benchmark for under-represented languages.
  49. 49.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  50. 50.Hanna Sababa and Athena Stassopoulou. 2018. A classifier to distinguish between Cypriot Greek and Standard Modern Greek. In 2018 Fifth International Conference on Social Networks Analysis, Management and Security (SNAMS), pages 251–255.
  51. 51.Yves Scherrer, Tanja Samardžic, and Elvira Glaser. 2019. Digitising Swiss German: how to process and study a polycentric spoken language. Language Resources and Evaluation, 53.
  52. 52.Haitham Seelawi, Ibraheem Tuffaha, Mahmoud Gzawi, Wael Farhan, Bashar Talafha, Riham Badawi, Zyad Sober, Oday Al-Dweik, Abed Alhakim Freihat, and Hussein Al-Natsheh. 2021. ALUE: Arabic language understanding evaluation. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 173–184, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  53. 53.Yueqi Song, Catherine Cui, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, and Graham Neubig. 2023. GlobalBench: A benchmark for global progress in natural language processing.
  54. 54.Hongmin Wang, Yue Zhang, GuangYong Leonard Chan, Jie Yang, and Hai Leong Chieu. 2017. Universal Dependencies parsing for colloquial Singaporean English. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1732–1744, Vancouver, Canada. Association for Computational Linguistics.
  55. 55.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  56. 56.Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020. IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 843–857, Suzhou, China. Association for Computational Linguistics.
  57. 57.Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari, Nishant Nair, and Yash Bangera. 2023. Language variety identification with true labels.
  58. 58.Daniel Zeman, Joakim Nivre, Mitchell Abrams, Elia Ackermann, Noëmi Aepli, Hamid Aghaei, Željko Agic, Amir Ahmadi, Lars Ahrenberg, Ajede, and et al. 2021. Universal Dependencies 2.9. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University.
  59. 59.Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022. VALUE: Understanding dialect disparity in NLU. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3701–3720, Dublin, Ireland. Association for Computational Linguistics.
  60. 60.Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. Multi-VALUE: A framework for cross-dialectal English NLP. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 744–768, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Faisal, F., et al. “DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14412–54, https://doi.org/10.18653/v1/2024.acl-long.777.
APA
Faisal, F., Ahia, O., Srivastava, A., Ahuja, K., Chiang, D., Tsvetkov, Y., & Anastasopoulos, A. (2024). DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14412–14454. https://doi.org/10.18653/v1/2024.acl-long.777
Chicago
Faisal, F., O. Ahia, A. Srivastava, et al. 2024. “DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14412–54. https://doi.org/10.18653/v1/2024.acl-long.777.
Harvard
Faisal, F. et al. (2024) “DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14412–14454. Available at: https://doi.org/10.18653/v1/2024.acl-long.777.
Vancouver
1. Faisal F, Ahia O, Srivastava A, Ahuja K, Chiang D, Tsvetkov Y, Anastasopoulos A (2024) DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14412–14454

BibTeX

@inproceedings{faisal-etal-2024-dialectbench,
    title = "{DIALECTBENCH}: An {NLP} Benchmark for Dialects, Varieties, and Closely-Related Languages",
    author = "Faisal, Fahim  and
      Ahia, Orevaoghene  and
      Srivastava, Aarohi  and
      Ahuja, Kabir  and
      Chiang, David  and
      Tsvetkov, Yulia  and
      Anastasopoulos, Antonios",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.777/",
    doi = "10.18653/v1/2024.acl-long.777",
    pages = "14412--14454"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/