Multi-VALUE: A Framework for Cross-Dialectal English NLP

Caleb ZiemsWilliam Barr HeldJingfeng YangJwala DhamalaRahul GuptaDiyi Yang

article2023ACL77 citations

Introduces Multi-VALUE, a rule-based framework spanning 50 English dialects and 189 linguistic features that converts standard text into synthetic dialectal forms to expose model performance disparities and improve data diversity across non-standard varieties.

Listen

Modern language technologies are primarily designed and evaluated on Standard American English. Because natural language processing systems rarely account for regional and social varieties, speakers of nonstandard dialects experience notable performance drops and allocational harms in everyday digital tools. This lack of dialect robustness creates significant barriers to fair, reliable, and equitable technological access for global English speakers.

The article introduces Multi-VALUE, a framework designed to benchmark and improve dialect robustness across 50 English varieties. The primary objective is to evaluate performance disparities in state-of-the-art systems and demonstrate how rule-based synthetic data augmentation can mitigate these disparities.

The researchers operationalized 189 morphosyntactic perturbation rules derived from the Electronic World Atlas of Varieties of English, mapping Standard American English to synthetic dialectal forms while preserving underlying semantic labels. To ensure linguistic validity, 72 native speakers evaluated 19,000 sentence pairs across 10 dialects, establishing gold-standard conversational question answering datasets in Chicano and Indian English. Using these synthetic and human-verified benchmarks, the authors stress-tested leading artificial intelligence models across three core tasks: conversational question answering, semantic parsing (converting text into database queries), and machine translation.

The stress tests revealed substantial, statistically significant performance gaps across nonstandard dialects. In question answering, models exhibited severe performance drops, particularly on Colloquial Singapore English, where accuracy fell by up to 25.4%, while Indian and African American English experienced drops of 6.7% to 9.0%. In semantic parsing, exact-match accuracy decreased by 12.3% to 15.3% on non-American dialects. In machine translation, performance dropped by up to 43.6%, with the steepest declines occurring when translating into target languages structurally similar to English. Importantly, fine-tuning models on synthetic multi-dialectal data mitigated these drops, improving average cross-dialectal performance by 2.7 points, though it introduced a minor 1.2-point performance penalty on Standard American English.

These findings demonstrate that commercial and open-source models suffer from systematic cross-dialectal fragility due to training mismatches rather than underlying query ambiguity. Compressed or distilled models showed heightened vulnerability, indicating that model optimization practices may disproportionately degrade quality for low-resource dialect speakers. Incorporating synthetic linguistic perturbations provides a practical mechanism to close performance gaps, though organizations must manage slight trade-offs in standard language accuracy.

Organizations developing user-facing language technologies should integrate synthetic stress testing into their evaluation pipelines to detect dialect bias before deployment. Engineering teams should adopt multi-dialectal data augmentation during training to increase system resilience across varied user bases. Where resources allow, teams should pair synthetic augmentation with targeted native-speaker testing for high-priority dialects.

The framework focuses primarily on grammar, syntax, and morphology, leaving localized vocabulary and slang unaddressed. Additionally, while human validation confirmed high rule accuracy, synthetic transformations represent a lower bound on performance and may not capture the fluid, context-dependent nature of spoken dialects. Confidence in the diagnostic utility of the framework remains high, providing an actionable foundation for building more inclusive language systems.

Ziems et al (2023).pdf
  • Paper: VALUE: Understanding Dialect Disparity in NLU, Caleb Ziems et al. (2022). Read VALUE first to see the original dialect-robustness benchmark and rule-based conversion of Standard American English into African American Vernacular English that Multi-VALUE expands across English varieties.

No sufficiently relevant recommendations were found.

Cover for Multi-VALUE: A Framework for Cross-Dialectal English NLP

Abstract

Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts. Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE). We introduce a suite of resources for evaluating and achieving English dialect invariance. The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features. Multi-VALUE maps SAE to synthetic forms of each dialect. First, we use this system to stress tests question answering, machine translation, and semantic parsing. Stress tests reveal significant performance disparities for leading models on non-standard dialects. Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems. Finally, we partner with native speakers of Chicano and Indian English to release new gold-standard variants of the popular CoQA task. To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multi-VALUE Perturbations
  • 4 Scope and Reliability of Multi-VALUE
  • 4.1 Scope
  • 4.2 Recruiting Native Speakers for Validation
  • 4.3 Validating the Multi-VALUE Pipeline
  • 4.4 Gold Test Sets
  • 5 Using Multi-VALUE
  • 6 Cross-Dialectal Stress Testing
  • 6.1 Linking Natural and Synthetic Data
  • 6.2 Synthetic Stress Tests
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Considerations
  • Acknowledgements
  • References
  • A Implementation Details
  • A.1 Pronouns
  • A.2 Noun Phrases
  • A.3 Tense and Aspect
  • A.4 Mood
  • A.5 Verb Morphology
  • A.6 Negation
  • A.7 Agreement
  • A.8 Relativization
  • A.9 Complementation
  • A.10 Adverbial Subordination
  • A.11 Adverbial Prepositions
  • A.12 Discourse and Word Order
  • B Models & Hyperparameters

Knowls

  1. Knowl 1 — Multi-VALUE translates Standard American English into feature-controlled dialect variants

    model/method

    Multi-VALUE is a rule-based framework for evaluating and improving English dialect robustness. It transforms Standard American English (SAE) text by applying selected linguistic perturbation rules, producing synthetic forms of 50 English dialects. The transformations target morphological and syntactic variation while aiming to preserve the original meaning and task labels, so existing NLP benchmarks can be converted into dialect stress tests and training data. The framework is grounded in eWAVE, a linguistic catalogue of English-variety features, and was designed for varieties that share vocabulary with SAE and are mutually intelligible with it while differing in morphology or syntax.

  2. Knowl 2 — Perturbations use linguistic evidence and dialect-specific feature frequencies

    model/method

    A Multi-VALUE dialect transformation sequentially applies the perturbation rules associated with that dialect. Whether a feature is applied follows the eWAVE pervasiveness rating: obligatory features are applied with probability 100%, features rated neither pervasive nor rare with probability 60%, rare features with probability 30%, and features with no information or an attested absence with probability 0%. Rules condition on morphosyntactic context, including part-of-speech tags, inflection, and dependency relations; the implementation uses spaCy 2.1.0 and inflect 5.5.2. For example, the give-passive rule identifies a passive construction with a past-participle root, a passive subject, and an agent, then restructures it using the dialectal serial verb give and a base-form main verb. The implemented rules span 12 grammatical groupings, including pronouns, noun phrases, tense and aspect, negation, agreement, relativization, complementation, and word order.

  3. Knowl 3 — The rule set covers 189 features across 50 documented English dialects

    empirical result

    Multi-VALUE implements 189 of the 235 linguistic features catalogued in eWAVE and covers all 50 dialects represented in that catalogue. On average, 86.6% of the documented feature space for a dialect is implemented, and every dialect has at least 80% coverage. The authors report five dialects with synthetic data for their experiments: Appalachian English, Chicano English, Colloquial Singapore English, Indian English, and Urban African American English.

  4. Knowl 4 — Native-speaker judgments support the plausibility of the perturbation rules

    empirical result

    The authors evaluated the linguistic acceptability of synthetic transformations using native speakers of 10 English dialects. Seventy-two annotators assessed approximately 19,000 sentence pairs; three annotators evaluated each transformation, and annotators marked changed spans they considered ungrammatical. The evaluation covered 92 perturbation rules, each with at least five unique sentence instances, and majority vote determined rule accuracy. Fifty-five rules received perfect accuracy; 74 of 92 had accuracy above 95%, 16 were in the range [85%, 95%), and two were below 85%. Every rule exceeded 81% accuracy, with the lowest reported accuracy being 81.8%.

  5. Knowl 5 — Human-reviewed Chicano and Indian English CoQA sets provide natural-dialect evaluation

    data/table

    The authors created human-reviewed CoQA evaluation data in Chicano English (ChcE) and Indian English (IndE) by presenting annotators with transformed questions and allowing them to rephrase an unacceptable transformation or remove it. Of 7,983 CoQA questions, the transformation pipeline changed 1,726 ChcE questions (21.6%) and 6,825 IndE questions (85.4%). Transformation accuracy was 82.7% for ChcE and 66.1% for IndE; the authors attribute IndE’s lower accuracy to its higher feature density. After errors were corrected or removed, the sets retained 1,498 transformed ChcE questions and 5,289 transformed IndE questions, which were combined with unperturbed questions for evaluation.

    On these gold evaluation sets, models trained on SAE CoQA scored the following F1 values: BERT scored 77.2 on SAE, 76.7 on ChcE, and 72.3 on IndE; RoBERTa scored 81.8, 81.6, and 77.7, respectively.

  6. Knowl 6 — Synthetic stress tests approximate natural CoQA performance and track treatment effects

    empirical result

    The authors compared synthetic CoQA evaluations with the human-reviewed ChcE and IndE evaluation sets. For SAE-trained BERT, the synthetic ChcE score was 76.6 F1 versus 76.7 on gold ChcE, while synthetic IndE scored 70.8 versus 72.3 on gold IndE. For SAE-trained RoBERTa, synthetic ChcE scored 81.5 versus 81.6 on gold ChcE, and synthetic IndE scored 76.1 versus 77.7 on gold IndE. Thus, synthetic evaluation closely matched the ChcE results and somewhat overstated the IndE performance drop. The authors characterize synthetic results as a lower bound on performance for a target dialect, since synthetic feature density may exceed speakers’ natural usage. Across the compared training treatments, changes that improved synthetic stress-test results also improved gold-set results.

  7. Knowl 7 — Cross-dialect CoQA testing reveals disparities, while multi-dialect training improves average robustness

    empirical result

    For CoQA, the authors transformed questions while leaving the SAE passages unchanged, and evaluated BERT-base and RoBERTa-base using F1. With SAE-trained RoBERTa, scores were 81.8 on SAE, 79.1 on Appalachian English (AppE), 81.5 on ChcE, 68.8 on Colloquial Singapore English (CollSgE), 76.1 on IndE, and 76.6 on Urban African American English (UAAVE). The reported relative drops from SAE were 3.4%, 0.3%, 18.9%, 7.5%, and 6.7%, respectively; all except the ChcE drop were statistically significant at P<0.05P<0.05. BERT showed the same pattern of greatest disparity for CollSgE: its SAE score was 77.2, compared with 74.4 on AppE, 76.6 on ChcE, 61.5 on CollSgE, 70.8 on IndE, and 71.2 on UAAVE.

    Training RoBERTa on synthetic multi-dialect data raised average cross-dialect F1 from 77.3 to 80.0, a 2.7-point improvement, while lowering its SAE-test score from 81.8 to 80.6. Training on data transformed to the target dialect improved results over SAE-only training for every tested dialect except ChcE. In a qualitative analysis of model errors, changes to tense, inflection, plural marking, constituent order, and deletions of recoverable pronouns, prepositions, or auxiliaries were associated with errors; some errors cascaded into later conversational questions. Perturbations also led models to answer yes/no questions with inappropriate response types such as noun or prepositional phrases.

  8. Knowl 8 — Dialect-shifted Spider queries reduce semantic-parsing accuracy across all tested varieties

    empirical result

    The authors evaluated models fine-tuned on Spider by transforming only the natural-language query; database tables and SQL targets remained unchanged. Across BART-base, BART-large, T5-base, and T5-3B, both exact-match accuracy and execution accuracy were significantly lower (P<0.05P<0.05) on every tested dialect than on SAE. Average exact-match and execution accuracies, respectively, were 45.1 and 46.9 for BART-base, 63.5 and 65.4 for BART-large, 53.8 and 55.3 for T5-base, and 66.5 and 69.4 for T5-3B. The largest reported exact-match drops for T5-3B occurred on CollSgE (60.7, a 15.3% drop from SAE) and IndE (62.9, a 12.3% drop). ChcE generally produced smaller drops than the other tested dialects. Since the database and SQL query were dialect-free, these disparities indicate a mismatch between model training and dialect-shifted input rather than a dialect mismatch between the query and its knowledge base.

  9. Knowl 9 — Dialect variation lowers English-to-other-language translation scores

    empirical result

    The authors evaluated two distilled NLLB translation models, with 615M and 1.3B parameters, on WMT19 input transformed into AppE, ChcE, CollSgE, IndE, or UAAVE. The task translated English into Chinese, German, Gujarati, or Russian, and performance was measured with SacreBLEU. Every tested dialect had a lower score than SAE; most drops were statistically significant at P<0.05P<0.05, while some ChcE comparisons were not. Across dialects, the average relative drops for the 615M model were 10.6% for Chinese, 19.5% for German, 17.2% for Gujarati, and 16.8% for Russian. For the 1.3B model, the corresponding drops were 11.0%, 18.0%, 15.7%, and 15.5%. CollSgE had the largest drop for each target language: for example, German translation declined by 43.60% with the 615M model and 40.6% with the 1.3B model. The average dialect-related drop was largest for German and smallest for Chinese at both model sizes.

  10. Knowl 10 — The framework does not cover all forms of dialect variation

    limitation

    Multi-VALUE focuses on morphosyntactic features, not lexical variation, because the authors consider lexical patterns difficult to describe with scalable, generalizable rules and note that many low-resource dialects lack corpora for estimating them. Its coverage is limited to English varieties documented in eWAVE and to canonical feature forms represented in that catalogue. The underlying documentation comes from linguists’ interviews with speakers, so its orthographic conventions may not match speakers’ writing practices, and it may omit within-feature variation. The authors also caution that dialects are not deterministic categories: speakers can vary which grammatical options they use and how frequently they use them across personal and social contexts.

Coverage note — The appendix’s exhaustive list of individual perturbation rules and its dialect-by-dialect feature inventory are omitted; the framework’s coverage, representative rule mechanics, validation, and main experimental findings are captured without reproducing the full catalogue.

References

  1. 1.Orevaoghene Ahia, Julia Kreutzer, and Sara Hooker. 2021. The low-resource double bind: An empirical study of pruning for low-resource machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3316–3333, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  2. 2.Wasi Ahmad, Zhisong Zhang, Xuezhe Ma, Eduard Hovy, Kai-Wei Chang, and Nanyun Peng. 2019. On difficulties of cross-lingual transfer with order differences: A case study on dependency parsing. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2440–2452, Minneapolis, Minnesota. Association for Computational Linguistics.
  3. 3.Collin F. Baker, Charles J. Fillmore, and John B. Lowe. 1998. The Berkeley FrameNet project. In COLING 1998 Volume 1: The 17th International Conference on Computational Linguistics.
  4. 4.Jeremy Barnes, Erik Velldal, and Lilja Øvrelid. 2021. Improving sentiment analysis with multi-task learning of negation. Natural Language Engineering, 27(2):249–269.
  5. 5.Emily M Bender. 2011. On achieving and evaluating language-independence in nlp. Linguistic Issues in Language Technology, 6.
  6. 6.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  7. 7.Steven Bird. 2022. Local languages, third spaces, and other high-resource scenarios. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7817–7829.
  8. 8.Damián Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022. Systematic inequalities in language technology performance across the world’s languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5486–5505.
  9. 9.Su Lin Blodgett and Brendan O’Connor. 2017. Racial disparity in natural language processing: A case study of social media african-american english. ArXiv preprint, abs/1707.00061.
  10. 10.Su Lin Blodgett, Johnny Wei, and Brendan O’Connor. 2018. Twitter Universal Dependency parsing for African-American and mainstream American English. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1415–1425, Melbourne, Australia. Association for Computational Linguistics.
  11. 11.Jack K Chambers and Peter Trudgill. 1998. Dialectology. Cambridge University Press.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  13. 13.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. No language left behind: Scaling human-centered machine translation. ArXiv preprint, abs/2207.04672.
  14. 14.María José López Couso and Belén Méndez Naya. 2015. Epistemic/evidential markers of the type verb+ complementizer: Some parallels from english and romance. In New directions in grammaticalization research, pages 93–120. John Benjamins.
  15. 15.Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial bias in hate speech and abusive language detection datasets. In Proceedings of the Third Workshop on Abusive Language Online, pages 25–35, Florence, Italy. Association for Computational Linguistics.
  16. 16.Dorottya Demszky, Nikhil Garg, Rob Voigt, James Zou, Jesse Shapiro, Matthew Gentzkow, and Dan Jurafsky. 2019. Analyzing polarization in social media: Method and application to tweets on 21 mass shootings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2970–3005, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Dorottya Demszky, Devyani Sharma, Jonathan Clark, Vinodkumar Prabhakaran, and Jacob Eisenstein. 2021. Learning to recognize dialect features. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2315–2338, Online. Association for Computational Linguistics.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Penelope Eckert. 1989. Jocks and burnouts: Social categories and identity in the high school. Teachers college press.
  20. 20.Penelope Eckert. 2017. Age as a sociolinguistic variable. The handbook of sociolinguistics, pages 151–167.
  21. 21.Barbara Di Eugenio and Michael Glass. 2004. The kappa statistic: A second look. Computational linguistics, 30(1):95–101.
  22. 22.Fahim Faisal, Sharlina Keshava, Md Mahfuz Ibn Alam, and Antonios Anastasopoulos. 2021. SD-QA: Spoken dialectal question answering for the real world. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3296–3315, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  23. 23.Daniel Gildea and Daniel Jurafsky. 2000. Automatic labeling of semantic roles. In Proceedings of the 38th Annual Meeting of the Association for Computational Linguistics, pages 512–520, Hong Kong. Association for Computational Linguistics.
  24. 24.Yichen Gong, Heng Luo, and Jian Zhang. 2018. Natural language inference over interaction space. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  25. 25.Lisa J Green. 2002. African American English: a linguistic introduction. Cambridge University Press.
  26. 26.Suchin Gururangan, Dallas Card, Sarah K Drier, Emily K Gade, Leroy Z Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A Smith. 2022. Whose language counts as high quality? measuring language ideologies in text data selection. ArXiv preprint, abs/2201.10474.
  27. 27.Matan Halevy, Camille Harris, Amy Bruckman, Diyi Yang, and Ayanna Howard. 2021. Mitigating racial biases in toxic language detection with an equity-based ensemble framework. In Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–11.
  28. 28.Janet Holmes and Miriam Meyerhoff. 2008. The handbook of language and gender. John Wiley & Sons.
  29. 29.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spacy: Industrial-strength natural language processing in python.
  30. 30.Md Mosharaf Hossain, Venelin Kovatchev, Pranoy Dutta, Tiffany Kao, Elizabeth Wei, and Eduardo Blanco. 2020. An analysis of natural language inference benchmarks through the lens of negation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9106–9118, Online. Association for Computational Linguistics.
  31. 31.Dirk Hovy and Shannon L. Spruit. 2016. The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 591–598, Berlin, Germany. Association for Computational Linguistics.
  32. 32.Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 588–602, Online. Association for Computational Linguistics.
  33. 33.Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2015. Challenges of studying and processing dialects in social media. In Proceedings of the Workshop on Noisy User-generated Text, pages 9–18, Beijing, China. Association for Computational Linguistics.
  34. 34.Aravind K. Joshi and Ralph Weischedel. 1977. Computation of a subclass of inferences: Presupposition and entailment. American Journal of Computational Linguistics, pages 1–54. Microfiche 63.
  35. 35.Ying Ju, Fubang Zhao, Shijie Chen, Bowen Zheng, Xuefeng Yang, and Yunfeng Liu. 2019. Technical report on conversational question answering. ArXiv preprint, abs/1909.10772.
  36. 36.David Jurgens, Yulia Tsvetkov, and Dan Jurafsky. 2017. Incorporating dialectal variability for socially equitable language identification. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 51–57, Vancouver, Canada. Association for Computational Linguistics.
  37. 37.Liza King and Roser Morante. 2020. Must children be vaccinated or not? annotating modal verbs in the vaccination debate. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 5730–5738, Marseille, France. European Language Resources Association.
  38. 38.Svetlana Kiritchenko and Saif Mohammad. 2016. The effect of negators, modals, and degree adverbs on sentiment composition. In Proceedings of the 7th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 43–52, San Diego, California. Association for Computational Linguistics.
  39. 39.Philipp Koehn and Rebecca Knowles. 2017. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 28–39.
  40. 40.Bernd Kortmann, Kerstin Lunkenheimer, and Katharina Ehret, editors. 2020. eWAVE.
  41. 41.Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020. Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 737–762, Online. Association for Computational Linguistics.
  42. 42.William Labov. 1972. Language in the inner city: Studies in the Black English vernacular. 3. University of Pennsylvania Press.
  43. 43.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019. Choosing transfer languages for cross-lingual learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3125–3135, Florence, Italy. Association for Computational Linguistics.
  44. 44.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  45. 45.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv preprint, abs/1907.11692.
  46. 46.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  47. 47.Brandon Lwowski and Anthony Rios. 2021. The risk of racial bias while tracking influenza-related content on social media using machine learning. Journal of the American Medical Informatics Association, 28(4):839–849.
  48. 48.Evgeny Matusov. 2019. The challenges of using neural machine translation for literature. In Proceedings of the Qualities of Literary Machine Translation, pages 10–19, Dublin, Ireland. European Association for Machine Translation.
  49. 49.Marzieh Mozafari, Reza Farahbakhsh, and Noël Crespi. 2020. Hate speech detection and racial bias mitigation in social media based on bert model. PloS one, 15(8):e0237861.
  50. 50.John Nerbonne. 2009. Data-driven dialectology. Language and Linguistics Compass, 3(1):175–198.
  51. 51.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 271–278, Barcelona, Spain.
  52. 52.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 115–124, Ann Arbor, Michigan. Association for Computational Linguistics.
  53. 53.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  54. 54.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  55. 55.Christopher Potts. 2002. The lexical semantics of parenthical-as and appositive-which. Syntax, 5(1):55–88.
  56. 56.Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. 2022. Perturbation augmentation for fairer nlp. ArXiv preprint, abs/2205.12586.
  57. 57.Shauli Ravfogel, Yoav Goldberg, and Francis Tyers. 2018. Can LSTM learn to capture agreement? the case of Basque. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 98–107, Brussels, Belgium. Association for Computational Linguistics.
  58. 58.Marta Recasens, Cristian Danescu-Niculescu-Mizil, and Dan Jurafsky. 2013. Linguistic models for analyzing and detecting biased language. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1650–1659, Sofia, Bulgaria. Association for Computational Linguistics.
  59. 59.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  60. 60.Anthony Rios. 2020. Fuzze: Fuzzy fairness evaluation of offensive language classifiers on african-american english. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 881–889. AAAI Press.
  61. 61.Tony Rose, Mark Stevenson, and Miles Whitehead. 2002. The Reuters corpus volume 1 -from yesterday’s news to tomorrow’s language resources. In Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02), Las Palmas, Canary Islands - Spain. European Language Resources Association (ELRA).
  62. 62.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  63. 63.Mark Sebba. 1997. Contact languages: Pidgins and creoles. Bloomsbury Publishing.
  64. 64.William Stewart. 1968. A sociolinguistic typology for describing national multilingualism. Readings in the Sociology of Language, 3:531–545.
  65. 65.Rhea Sukthanker, Soujanya Poria, Erik Cambria, and Ramkumar Thirunavukarasu. 2020. Anaphora and coreference resolution: A review. Information Fusion, 59:139–162.
  66. 66.Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2022. Dialect-robust evaluation of generated text.
  67. 67.Arie Sutiono and Gus Hahn-Powell. 2022. Syntax-driven data augmentation for named entity recognition. In Proceedings of the First Workshop on Pattern-based Approaches to NLP in the Age of Deep Learning, pages 56–60, Gyeongju, Republic of Korea. International Conference on Computational Linguistics.
  68. 68.Maggie Tallerman. 2019. Understanding syntax. Routledge.
  69. 69.Reut Tsarfaty, Dan Bareket, Stav Klein, and Amit Seker. 2020. From SPMRL to NMRL: What did we learn (and unlearn) in a decade of parsing morphologically-rich languages (MRLs)? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7396–7408, Online. Association for Computational Linguistics.
  70. 70.Zirui Wang, Zihang Dai, Barnabás Póczos, and Jaime Carbonell. 2019. Characterizing and avoiding negative transfer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11293–11302.
  71. 71.Zirui Wang, Zachary C Lipton, and Yulia Tsvetkov. 2020. On negative interference in multilingual models: Findings and a meta-learning treatment. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4438–4450.
  72. 72.Jason Wei and Kai Zou. 2019. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6382–6388, Hong Kong, China. Association for Computational Linguistics.
  73. 73.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  74. 74.Zhengxuan Wu, Isabel Papadimitriou, and Alex Tamkin. 2022. Oolong: Investigating what makes crosslingual transfer hard with controlled studies. ArXiv preprint, abs/2202.12312.
  75. 75.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. ArXiv preprint, abs/2201.05966.
  76. 76.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  77. 77.Charles D Yang. 2000. Internal and external forces in language change. Language variation and change, 12(3):231–250.
  78. 78.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, Brussels, Belgium. Association for Computational Linguistics.
  79. 79.Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah Smith. 2021. Challenges in automated debiasing for toxic language detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3143–3155, Online. Association for Computational Linguistics.
  80. 80.Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022. VALUE: Understanding dialect disparity in NLU. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3701–3720, Dublin, Ireland. Association for Computational Linguistics.
  81. 81.Caleb Ziems and Diyi Yang. 2021. To protect and to serve? analyzing entity-centric framing of police violence. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 957–976, Punta Cana, Dominican Republic. Association for Computational Linguistics.

Citation

MLA
Ziems, C., et al. “Multi-VALUE: A Framework for Cross-Dialectal English NLP”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 744–68, https://doi.org/10.18653/v1/2023.acl-long.44.
APA
Ziems, C., Held, W., Yang, J., Dhamala, J., Gupta, R., & Yang, D. (2023). Multi-VALUE: A Framework for Cross-Dialectal English NLP. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 744–768. https://doi.org/10.18653/v1/2023.acl-long.44
Chicago
Ziems, C., W. Held, J. Yang, J. Dhamala, R. Gupta, and D. Yang. 2023. “Multi-VALUE: A Framework for Cross-Dialectal English NLP”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 744–68. https://doi.org/10.18653/v1/2023.acl-long.44.
Harvard
Ziems, C. et al. (2023) “Multi-VALUE: A Framework for Cross-Dialectal English NLP”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 744–768. Available at: https://doi.org/10.18653/v1/2023.acl-long.44.
Vancouver
1. Ziems C, Held W, Yang J, Dhamala J, Gupta R, Yang D (2023) Multi-VALUE: A Framework for Cross-Dialectal English NLP. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 744–768

BibTeX

@inproceedings{ziems-etal-2023-multi,
    title = "Multi-{VALUE}: A Framework for Cross-Dialectal {E}nglish {NLP}",
    author = "Ziems, Caleb  and
      Held, William  and
      Yang, Jingfeng  and
      Dhamala, Jwala  and
      Gupta, Rahul  and
      Yang, Diyi",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.44/",
    doi = "10.18653/v1/2023.acl-long.44",
    pages = "744--768"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/