DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages
Fahim FaisalOrevaoghene AhiaAarohi SrivastavaKabir AhujaDavid ChiangYulia TsvetkovAntonios Anastasopoulos
Introduces DIALECTBENCH, a comprehensive evaluation suite spanning 10 NLP tasks and 281 language varieties across 40 clusters to quantify and analyze performance disparities between standard and non-standard dialects.
Natural language processing models are increasingly deployed worldwide, yet standard evaluation benchmarks focus almost exclusively on high-resource, standardized national languages. This practice overlooks regional dialects, non-standard varieties, and closely related low-resource languages, hiding significant performance failures for non-standard speakers. The article introduces DialectBench to systematically evaluate, benchmark, and analyze natural language processing performance disparities between standard languages and their non-standard dialectal variants across diverse tasks.
To establish this benchmark, the article aggregated diverse datasets across 10 text-level tasks—including structured prediction, text classification, question answering, and machine translation—covering 281 language varieties grouped into 40 language clusters using genealogical linguistic standards. The evaluation benchmarked standard multilingual pre-trained encoder models alongside large language models and machine translation systems across varied settings, such as zero-shot cross-lingual transfer and localized fine-tuning.
Key findings demonstrate severe performance disparities across language varieties. First, models achieve high accuracy on standard and Germanic or Romance varieties (often exceeding 90% F1 or parsing scores) but drop drastically on low-resource indigenous or regional dialects, falling below 10% in tasks like dependency parsing for Mbyá Guaraní. Second, fine-tuning directly on dialectal data exacerbates within-cluster performance divergence compared to zero-shot transfer because training data across dialects is highly inconsistent in volume and quality. Third, writing systems heavily influence transferability; low-resource varieties using the Latin script benefit significantly more from English zero-shot transfer than non-Latin varieties. Finally, while large language models evaluated via few-shot prompting outperformed zero-shot transfer baselines on select dialect tasks, they consistently lagged behind dedicated fine-tuned models.
These results demonstrate that current language systems carry substantial risks of technological exclusion and operational failure when deployed in diverse linguistic environments. Standard multilingual evaluation scores overestimate model readiness by masking severe within-cluster inequalities. Organizations relying on standard benchmarks face compliance, user adoption, and accuracy risks in multilingual markets.
Moving forward, researchers and technology practitioners should use dialect-aware evaluation frameworks before deploying multilingual systems. Development should focus on gathering curated parallel dialect resources, standardizing data quality, and expanding evaluations into speech-based systems. Leaders must account for both demographic utility (speaker population coverage) and linguistic utility (equal performance across variants) when auditing system fairness.
These conclusions are bounded by data scarcity constraints, domain inconsistencies across collated datasets, and limited human-generated dialectal translation references, which necessitated synthetic pseudo-references for translation evaluations. Nevertheless, the benchmark provides robust evidence of systematic dialect disparities across standard language models.
- Paper: VALUE: Understanding Dialect Disparity in NLU, Caleb Ziems et al. (2022). VALUE establishes an earlier dialect-focused NLU benchmark whose evaluation of disparity and dialect-validated data provides a direct foundation for understanding DialectBench’s broader benchmark design.
- Paper: Multi-VALUE: A Framework for Cross-Dialectal English NLP, Caleb Ziems et al. (2023). Multi-VALUE extends dialect benchmarking across many English varieties and tasks, preparing readers for DialectBench’s wider comparison across language clusters.
- Paper: One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia, Alham Fikri Aji et al. (2022). Its empirical analysis of dialect-sensitive NLP failures and uneven language resources offers an earlier case study of the evaluation gaps that DialectBench measures systematically.
No sufficiently relevant recommendations were found.
