keyword
VernAcular Language Understanding Evaluation
VernAcular Language Understanding Evaluation is a benchmark and evaluation framework in natural language processing designed to assess and improve the ability of models to understand non-standard dialects and vernacular varieties of language. Because traditional language understanding benchmarks predominantly test standard written varieties like Standard American English, models often exhibit performance disparities when encountering linguistic shifts. This evaluation framework applies rule-based lexical and morphosyntactic transformations to standard benchmark tasks, converting texts into dialectal forms such as African American Vernacular English as well as other regional and social varieties. By providing standardized stress tests and data augmentation resources across diverse linguistic features, it enables researchers to measure dialect robustness and build more equitable language technologies.
2 items

Multi-VALUE: A Framework for Cross-Dialectal English NLP
Caleb Ziems, William Barr Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, Diyi Yang
Why you should read this
Introduces Multi-VALUE, a rule-based framework spanning 50 English dialects and 189 linguistic features that converts standard text into synthetic dialectal forms to expose model performance disparities and improve data diversity across non-standard varieties.
Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts. Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE). We introduce a suite of resources for evaluating and achieving English dialect invariance. The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features. Multi-VALUE maps SAE to synthetic forms of each dialect. First, we use this system to stress tests question answering, machine translation, and semantic parsing. Stress tests reveal significant performance disparities for leading models on non-standard dialects. Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems. Finally, we partner with native speakers of Chicano and Indian English to release new gold-standard variants of the popular CoQA task. To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/.
Added
2026-10-03

VALUE: Understanding Dialect Disparity in NLU
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, Diyi Yang
Why you should read this
Presents a dialect-specific evaluation benchmark and linguistically validated transformation rules for African American Vernacular English to expose and analyze performance disparities in modern language models.
English Natural Language Understanding (NLU) systems have achieved great performances and even outperformed humans on benchmarks like GLUE and SuperGLUE. However, these benchmarks contain only textbook Standard American English (SAE). Other dialects have been largely overlooked in the NLP community. This leads to biased and inequitable NLU systems that serve only a sub-population of speakers. To understand disparities in current models and to facilitate more dialect-competent NLU systems, we introduce the VernAcular Language Understanding Evaluation (VALUE) benchmark, a challenging variant of GLUE that we created with a set of lexical and morphosyntactic transformation rules. In this initial release (V.1), we construct rules for 11 features of African American Vernacular English (AAVE), and we recruit fluent AAVE speakers to validate each feature transformation via linguistic acceptability judgments in a participatory design manner. Experiments show that these new dialectal features can lead to a drop in model performance.
Added
2026-09-26
