Multi-VALUE: A Framework for Cross-Dialectal English NLP
Caleb ZiemsWilliam Barr HeldJingfeng YangJwala DhamalaRahul GuptaDiyi Yang
Introduces Multi-VALUE, a rule-based framework spanning 50 English dialects and 189 linguistic features that converts standard text into synthetic dialectal forms to expose model performance disparities and improve data diversity across non-standard varieties.
Modern language technologies are primarily designed and evaluated on Standard American English. Because natural language processing systems rarely account for regional and social varieties, speakers of nonstandard dialects experience notable performance drops and allocational harms in everyday digital tools. This lack of dialect robustness creates significant barriers to fair, reliable, and equitable technological access for global English speakers.
The article introduces Multi-VALUE, a framework designed to benchmark and improve dialect robustness across 50 English varieties. The primary objective is to evaluate performance disparities in state-of-the-art systems and demonstrate how rule-based synthetic data augmentation can mitigate these disparities.
The researchers operationalized 189 morphosyntactic perturbation rules derived from the Electronic World Atlas of Varieties of English, mapping Standard American English to synthetic dialectal forms while preserving underlying semantic labels. To ensure linguistic validity, 72 native speakers evaluated 19,000 sentence pairs across 10 dialects, establishing gold-standard conversational question answering datasets in Chicano and Indian English. Using these synthetic and human-verified benchmarks, the authors stress-tested leading artificial intelligence models across three core tasks: conversational question answering, semantic parsing (converting text into database queries), and machine translation.
The stress tests revealed substantial, statistically significant performance gaps across nonstandard dialects. In question answering, models exhibited severe performance drops, particularly on Colloquial Singapore English, where accuracy fell by up to 25.4%, while Indian and African American English experienced drops of 6.7% to 9.0%. In semantic parsing, exact-match accuracy decreased by 12.3% to 15.3% on non-American dialects. In machine translation, performance dropped by up to 43.6%, with the steepest declines occurring when translating into target languages structurally similar to English. Importantly, fine-tuning models on synthetic multi-dialectal data mitigated these drops, improving average cross-dialectal performance by 2.7 points, though it introduced a minor 1.2-point performance penalty on Standard American English.
These findings demonstrate that commercial and open-source models suffer from systematic cross-dialectal fragility due to training mismatches rather than underlying query ambiguity. Compressed or distilled models showed heightened vulnerability, indicating that model optimization practices may disproportionately degrade quality for low-resource dialect speakers. Incorporating synthetic linguistic perturbations provides a practical mechanism to close performance gaps, though organizations must manage slight trade-offs in standard language accuracy.
Organizations developing user-facing language technologies should integrate synthetic stress testing into their evaluation pipelines to detect dialect bias before deployment. Engineering teams should adopt multi-dialectal data augmentation during training to increase system resilience across varied user bases. Where resources allow, teams should pair synthetic augmentation with targeted native-speaker testing for high-priority dialects.
The framework focuses primarily on grammar, syntax, and morphology, leaving localized vocabulary and slang unaddressed. Additionally, while human validation confirmed high rule accuracy, synthetic transformations represent a lower bound on performance and may not capture the fluid, context-dependent nature of spoken dialects. Confidence in the diagnostic utility of the framework remains high, providing an actionable foundation for building more inclusive language systems.
- Paper: VALUE: Understanding Dialect Disparity in NLU, Caleb Ziems et al. (2022). Read VALUE first to see the original dialect-robustness benchmark and rule-based conversion of Standard American English into African American Vernacular English that Multi-VALUE expands across English varieties.
No sufficiently relevant recommendations were found.
