MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Jack FitzGeraldChristopher HenchCharith PerisScott MackieKay RottmannAna SanchezAaron NashLiam UrbachVishesh KakaralaRicha Singh

article2023ACL236 citations

Presents a large-scale, parallel spoken language understanding dataset of one million virtual assistant utterances across 51 typologically diverse languages to advance intent classification and slot-filling benchmarks for low-resource NLP.

Listen

Modern voice-based virtual assistants still support only a fraction of the world’s languages, largely due to a lack of realistic, high-quality, and task-specific labeled data. While massively multilingual foundation models have advanced rapidly, standardized evaluation benchmarks across diverse languages have failed to keep pace. The article introduces MASSIVE, a one-million-example dataset designed to evaluate and advance natural language understanding across 51 typologically diverse languages spanning 29 genera, 18 domains, 60 intents, and 55 slot types.

To build MASSIVE, the authors localized the English-only SLURP dataset into 50 additional languages using professional crowd-workers and a two-stage localization process that separated slot translation from sentence assembly. They then established baseline performance benchmarks by fine-tuning multilingual pre-trained models—specifically XLM-R and mT5—under both full-dataset multilingual training and zero-shot transfer conditions.

The findings show that full-dataset training consistently outperforms zero-shot cross-lingual transfer by substantial margins. Across all languages, full training achieved average exact match accuracies between 63.7% and 66.6%, whereas zero-shot transfer dropped to 28.8%–38.7% (a 25 to 37 percentage-point decline). Furthermore, zero-shot performance displayed severe variance across languages; for example, mT5 Text-to-Text showed a 15-point gap between highest and lowest-performing locales under full training, expanding to a 44-point gap under zero-shot conditions. Zero-shot accuracy strongly correlated with the volume of pre-training data available for each language (0.54 correlation for exact match), but supplying equal amounts of target-language fine-tuning data weakened this correlation to 0.42. Finally, languages using Latin scripts and Germanic genera performed best overall, while space-optional languages (such as Japanese) suffered steep performance drops—down to 9.4% zero-shot exact match—due to artificial character-spacing constraints.

These results demonstrate that relying solely on zero-shot cross-lingual transfer introduces severe performance risks and geographic disparities for production virtual assistants. Providing equal, high-quality localized training data across languages effectively mitigates the inherent biases of multilingual pre-training corpora. Organizations seeking to expand conversational assistants should invest in localized fine-tuning sets rather than relying exclusively on zero-shot generalization. Subsequent efforts should explore more sophisticated tokenization tools for unspaced scripts, scale evaluations to larger model architectures, and refine crowd-sourcing pipelines to minimize translation artifacts.

arXiv: 2204.08582alexa/massive
Cover for MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Abstract

We present the MASSIVE dataset—Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant Evaluation. MASSIVE contains 1M realistic, parallel, labeled virtual assistant utterances spanning 51 languages, 18 domains, 60 intents, and 55 slots. MASSIVE was created by tasking professional translators to localize the English-only SLURP dataset into 50 typologically diverse languages from 29 genera. We also present modeling results on XLM-R and mT5, including exact match accuracy, intent classification accuracy, and slot-filling F1 score. We have released our dataset, modeling code, and models publicly.

Table of Contents

  • 1 Introduction and Description
  • 2 Related Work
  • 3 Language Selection and Linguistic Analysis
  • 3.1 Language Selection
  • 3.2 Scripts
  • 3.3 Sentence Types
  • 3.4 Word Order
  • 3.5 Imperative Marking
  • 3.6 Politeness
  • 4 Collection Setup and Execution
  • 4.1 Heldout Evaluation Split
  • 4.2 Vendor Selection and Onboarding
  • 4.3 Collection Workflows
  • 4.4 Quality Assurance
  • 5 Model Benchmarking
  • 5.1 Setup
  • 5.2 Results and Analysis
  • 6 Conclusion
  • Limitations and Ethical Considerations
  • References
  • A Additional Linguistic Characteristics
  • B The Collection System
  • C Hyperparameters
  • D Results for All Languages
  • E A summary of model performance on language characteristics

Knowls

  1. Knowl 1 — MASSIVE Dataset Structure and Split Statistics

    definition

    MASSIVE (Multilingual Amazon SLU Resource Package for Slot-filling, Intent classification, and Virtual assistant Evaluation) is a parallel, multilingual natural language understanding (NLU) dataset composed of realistic, human-localized virtual assistant utterances.

    The dataset covers 51 languages (derived from 49 distinct spoken languages, with Mandarin collected in both Simplified and Traditional Chinese character sets) across 18 domains, 60 intents, and 55 slot types. Each language contains exactly 19,521 utterances localized from the English SLURP dataset, yielding approximately 1 million utterances in total across all locales.

    The dataset is partitioned into standardized splits across all 51 languages:

    • Train split: 587,000 utterances (~11,514 utterances per language)
    • Development (dev) split: 104,000 utterances (~2,039 utterances per language)
    • Test split: 152,000 utterances (~2,980 utterances per language)
    • Held-out evaluation split: 153,000 utterances (~3,000 utterances per language), created by professional annotators paraphrasing seed utterances to create more complex sentences with 49% more slots per utterance.
  2. Knowl 2 — Two-Stage Slot and Phrase Localization Workflow

    model/method

    To eliminate the annotation burden of directly highlighting token spans in target languages—which is error-prone for highly inflected languages where adpositions and affixes shift across word boundaries—MASSIVE was constructed using a two-stage sequential crowdsourcing workflow followed by an independent verification step:

    1. Slot Translation/Localization: An annotator is presented with the complete English source utterance containing highlighted slot spans, alongside each individual slot value and its slot label. The annotator chooses for each slot whether to translate the term (e.g., translating temporal words like "tomorrow"), localize the entity to a culturally or regionally appropriate target (e.g., replacing an English movie title or local artist with a culturally familiar equivalent), or keep the text unchanged (e.g., proper names like "Taylor Swift"). The choice made is preserved in the dataset metadata.
    2. Phrase Translation/Localization: A second annotator receives the full utterance alongside the translated slot outputs generated from the first step. This worker translates or localizes the full sentence, inserting and adapting the pre-translated slot entities into the target syntax, adjusting grammatical gender, inflection, and prepositional affixes around or within the slot spans.
    3. Quality Assurance Judgment: Three independent evaluators score the final sentence across five criteria: semantic match to the intent, semantic match of slot values to labels, grammaticality and naturalness, spelling, and language identification.
  3. Knowl 3 — Joint Intent Classification and Slot Filling Architectures with XLM-R and mT5

    model/method

    Baselines for massively multilingual natural language understanding are established using two multilingual pre-trained foundation models across three distinct architectures:

    1. XLM-RoBERTa (XLM-R) Base (270M parameters): A pre-trained multilingual transformer encoder coupled with two separate classification heads trained from scratch following the JointBERT framework. The intent classification head applies pooling over the sequence hidden states (evaluating first-token, mean, and maximum pooling). The slot filling head performs token-level sequence classification over the encoder output representations.
    2. mT5 Base Encoder-Only (258M parameters): The pre-trained encoder extracted from the multilingual T5 architecture, trained with the same dual-head joint classification setup (intent head and sequence tagging slot head) as XLM-R.
    3. mT5 Base Text-to-Text (580M parameters): The full sequence-to-sequence mT5 model operating in a generative text-to-text format. The model input is prefixed as Annotate: <utterance>. The decoder target is trained to autoregressively output the sequence of per-token slot labels (including the Other label for non-entity tokens) followed by the utterance intent name, tokenized into subwords rather than expanding the model vocabulary.

    All three architectures share a 192M parameter multilingual embedding table.

  4. Knowl 4 — Multilingual and Zero-Shot NLU Benchmark Performance

    data/table

    Baseline models evaluated on MASSIVE under full-dataset training (trained on all 51 languages simultaneously) and zero-shot cross-lingual transfer (trained solely on English en-US data and evaluated across all 50 non-English languages) yield the following performance metrics (Intent Accuracy, Micro-averaged Slot F1, and Exact Match Accuracy reported with 95% confidence intervals):

    Model Full Training Dataset Zero-Shot (Train: en-US)
    Intent Acc (%) Slot F1 (%) Exact Match (%) Intent Acc (%) Slot F1 (%) Exact Match (%)
    mT5 Base Text-to-Text 85.3±0.285.3 \pm 0.2 76.8±0.176.8 \pm 0.1 66.6±0.266.6 \pm 0.2 62.9±0.262.9 \pm 0.2 44.8±0.144.8 \pm 0.1 34.7±0.234.7 \pm 0.2
    mT5 Base Encoder-Only 86.1±0.286.1 \pm 0.2 75.4±0.175.4 \pm 0.1 65.9±0.265.9 \pm 0.2 61.2±0.261.2 \pm 0.2 41.6±0.141.6 \pm 0.1 28.8±0.228.8 \pm 0.2
    XLM-R Base 85.1±0.285.1 \pm 0.2 73.6±0.173.6 \pm 0.1 63.7±0.263.7 \pm 0.2 70.6±0.270.6 \pm 0.2 50.3±0.150.3 \pm 0.1 38.7±0.238.7 \pm 0.2

    mT5 Text-to-Text achieves the highest overall exact match accuracy under full-dataset multilingual training (66.6%66.6\%), while XLM-R Base exhibits the strongest zero-shot cross-lingual generalization capability (38.7%38.7\% exact match accuracy). Zero-shot evaluation causes an overall exact match accuracy degradation of 25 to 37 percentage points and drastically expands the performance disparity between the highest- and lowest-performing locales (from a 15-point gap in full training to a 44-point gap in zero-shot for mT5 Text-to-Text).

  5. Knowl 5 — Mitigation of Pretraining Resource Imbalance via Balanced Multilingual Fine-Tuning

    empirical result

    In zero-shot cross-lingual evaluation on XLM-R Base (where fine-tuning occurs only on English data), downstream task performance correlates moderately to strongly with the quantity of pre-training data available for each target language:

    • Exact Match Accuracy: Pearson correlation r=0.54r = 0.54
    • Intent Classification Accuracy: Pearson correlation r=0.58r = 0.58
    • Micro-averaged Slot F1: Pearson correlation r=0.46r = 0.46

    When models are fine-tuned on the full MASSIVE dataset with an equal distribution of supervised data across all 51 languages (19.5k examples per language), these correlations decrease substantially:

    • Exact Match Accuracy drops from r=0.54r = 0.54 to r=0.42r = 0.42
    • Intent Classification Accuracy drops from r=0.58r = 0.58 to r=0.47r = 0.47
    • Micro-averaged Slot F1 drops from r=0.46r = 0.46 to r=0.24r = 0.24

    This indicates that providing balanced, uniform-sized target-language supervised corpora during multi-task fine-tuning significantly mitigates the performance disparities caused by language skew in self-supervised pre-training corpora.

  6. Knowl 6 — Typological and Script Diversity in MASSIVE Language Selection

    definition

    The 51 languages in the MASSIVE dataset were selected to optimize typological breadth based on five criteria: translation cost and worker availability, presence in commercial virtual assistants, genus classification in the World Atlas of Linguistic Structures (WALS), eigenvector centrality in global translation and social networks (as a proxy for web resource availability), and script diversity.

    The resulting corpus spans:

    • 29 Genera: Language genera serve as a primary indicator of structural and typological diversity.
    • 14 Language Families / Isolates: Including Indo-European, Afro-Asiatic, Sino-Tibetan, Dravidian, Austronesian, Uralic, Turkic, Kartvelian, Kra-Dai, Niger-Congo, Mongolic, Japonic, Koreanic, and Austroasiatic.
    • 21 Distinct Scripts: 28 languages utilize Latin script variants; 3 use Arabic script; 2 use Cyrillic script; and 18 languages use scripts unique to a single language within the dataset (e.g., Devanagari, Ge'ez, Hangul, Hebrew, Greek, Armenian, Georgian, Khmer, Burmese, Thai, Kannada, Malayalam, Tamil, Telugu, Traditional Chinese, Simplified Chinese, Japanese).
  7. Knowl 7 — Typological Linguistic Phenomena in Virtual Assistant Utterances

    empirical result

    Analysis of the parallel utterances in MASSIVE reveals distinctive grammatical and typological distributions resulting from human-to-device conversational interactions:

    • Sentence Types: Utterances consist almost entirely of imperatives (commands/requests) and interrogatives (queries). Declaratives are rare and carry directive pragmatic force (e.g., stating "it's cold in here" as an instruction to raise the temperature).
    • Word Order: Among languages with classified word order in WALS, 39 languages are subject-initial (24 SVO, 15 SOV) and 3 are verb-initial (VSO: Arabic ar-SA, Tagalog tl-PH, and Welsh cy-GB). No object-initial languages (VOS, OVS, OSV) are present; 5 have no dominant order and 4 lack WALS data.
    • Imperative Morphology & Mood: 33 languages mark imperatives via dedicated verbal morphology (such as affixes), with 18 marking singular vs. plural addressee distinctions. 10 languages encode explicit hortative distinctions (speaker-inclusive commands), 4 feature optative moods expressing desires with command force, 18 utilize dedicated negative particles exclusively for prohibitive sentences, and 10 employ special prohibitive verb forms.
    • Politeness Encoding: 21 languages enforce a binary formal/informal second-person pronoun system; 8 languages feature multi-tier formality levels (including honorifics); and 7 languages employ pronoun avoidance strategies in polite contexts.
  8. Knowl 8 — Impact of Character-Level Tokenization Spacing on Non-Segmented Languages

    empirical result

    For languages written without standard whitespace word delimiters—specifically Japanese (ja-JP), Simplified Chinese (zh-CN), and Traditional Chinese (zh-TW)—slot annotation in MASSIVE required inserting whitespaces between every individual character.

    While character spacing permits character-level slot boundary alignment, it prevents the pre-trained subword tokenizers from grouping adjacent characters into the multi-character subwords learned during self-supervised pre-training. In zero-shot transfer, where the model cannot adapt its embedding layer to the artificial spacing format and relies entirely on pre-trained representations, this causes severe performance drops:

    • For Japanese (ja-JP), mT5 Text-to-Text exact match accuracy drops from 58.3%58.3\% in the full multilingual training setup to 9.4%9.4\% in the zero-shot cross-lingual setup.
    • Similar degradation occurs for Chinese (zh-CN drops from 64.8%64.8\% to 25.0%25.0\%; zh-TW drops from 61.0%61.0\% to 27.4%27.4\%).

    Conversely, in space-optional languages like Thai (th-TH), models leverage artificial whitespace around slot annotations as explicit entity boundary cues, achieving the highest exact match scores in full training (73.4%73.4\% on mT5 Text-to-Text).

  9. Knowl 9 — Quality Assurance and Worker Validation Pipeline

    model/method

    The MASSIVE data collection pipeline enforces a multi-tier quality control mechanism to ensure translation accuracy and linguistic naturalness across low- and high-resource languages:

    1. Worker Qualification Testing: Prospective crowdworkers must pass an Amazon MTurk-hosted auditory fluency evaluation (listening to target-language statements and answering comprehension questions) followed by a translation guidelines quiz with localized examples.
    2. Three-Way Independent Evaluation: Every localized utterance is reviewed by three separate annotators across five dimensions: semantic intent fidelity, slot-to-label alignment, naturalness and grammaticality, orthographic/spelling accuracy, and language identification (rejecting outputs containing non-target tokens unless natural for code-switching).
    3. Automated Anomaly Alarms: Real-time programmatic heuristics monitor and remove workers exhibiting high task rejection rates, abnormal slot deletion frequencies, disproportionate ratios of untranslated English tokens, or signatures of unedited raw machine translation. All tasks produced by disqualified workers are invalidated and re-queued for re-annotation.
  10. Knowl 10 — Structural and Modeling Limitations of the MASSIVE Benchmark

    limitation

    The MASSIVE dataset and its baseline evaluations exhibit several specific constraints:

    • Sample Size per Language: With 19.5k total records and ~11.5k training records per locale, per-language data volumes are modest compared to monolingual industrial NLU benchmarks.
    • Crowdsourcing Artificialities: Utterances were produced via localized translation workflows based on English seed prompts rather than organically collected user interactions with live production virtual assistants, introducing potential translationese or synthetic phrasing.
    • Schema Simplicity: The annotation format uses flat intent classification and token-level slot tagging, lacking the hierarchical or nested intent structures present in complex semantic parsing frameworks.
    • Tokenization and Script Artifacts: The absence of native, standardized word segmenters for non-delimited scripts (such as Japanese, Chinese, Khmer, and Thai) necessitates ad-hoc whitespace insertion, creating discrepancies with the tokenization utilized during self-supervised model pre-training.
    • Baseline Model Scale: Evaluations are restricted to Base-sized pre-trained models (~258M–580M parameters), leaving larger parameter scales and specialized cross-lingual alignment architectures unexplored.

Coverage note — None was omitted; all key contributions—dataset statistics, localization workflow, modeling architectures, multilingual and zero-shot results, linguistic analysis, tokenization findings, quality control mechanisms, and limitations—have been captured as self-contained knowls.

References

  1. 1.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2019. On the cross-lingual transferability of monolingual representations. CoRR, abs/1910.11856.
  2. 2.Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser. 2020. SLURP: A spoken language understanding resource package. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7252–7262, Online. Association for Computational Linguistics.
  3. 3.James Bergstra, Daniel Yamins, and David Cox. 2013. Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures. In Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 115–123, Atlanta, Georgia, USA. PMLR.
  4. 4.Qian Chen, Zhu Zhuo, and Wen Wang. 2019a. Bert for joint intent classification and slot filling. ArXiv, abs/1902.10909.
  5. 5.Xilun Chen, Ahmed Hassan Awadallah, Hany Hassan, Wei Wang, and Claire Cardie. 2019b. Multi-source cross-lingual model transfer: Learning what to share. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3098–3112, Florence, Italy. Association for Computational Linguistics.
  6. 6.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  7. 7.Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, and Joseph Dureau. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces.
  8. 8.Jan Christian Blaise Cruz and Charibeth Cheng. 2020. Establishing baselines for text classification in low-resource languages.
  9. 9.Jacob Devlin. 2018. Multiligual bert.
  10. 10.Matthew S. Dryer. 1989. Large linguistic areas and language sampling. Studies in Language, 13:257–292.
  11. 11.Matthew S. Dryer and Martin Haspelmath, editors. 2013. WALS Online. Max Planck Institute for Evolutionary Anthropology, Leipzig.
  12. 12.Maud Ehrmann, Marco Turchi, and Ralf Steinberger. 2011. Building a multilingual named entity-annotated corpus using annotation projection. In Proceedings of the International Conference Recent Advances in Natural Language Processing 2011, pages 118–124, Hissar, Bulgaria. Association for Computational Linguistics.
  13. 13.Akiko Eriguchi, Melvin Johnson, Orhan Firat, Hideto Kazawa, and Wolfgang Macherey. 2018. Zero-shot cross-lingual classification using multilingual neural machine translation.
  14. 14.Mehrdad Farahani, Mohammad Gharachorloo, and Mohammad Manthouri. 2021. Leveraging parsbert and pretrained mt5 for persian abstractive text summarization. 2021 26th International Computer Conference, Computer Society of Iran (CSICC).
  15. 15.Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi, Abhishek Bhagia, Claudio Delli Bovi, Jin Cao, Rakesh Chada, Amit Chauhan, Luoxin Chen, Anurag Dwarakanath, Satyam Dwivedi, Turan Gojayev, Karthik Gopalakrishnan, Thomas Gueudre, Dilek Hakkani-Tur, Wael Hamza, Jonathan Hueser, Kevin Martin Jose, Haidar Khan, Beiye Liu, Jianhua Lu, Alessandro Manzotti, Pradeep Natarajan, Karolina Owczarzak, Gokmen Oz, Enrico Palumbo, Charith Peris, Chandana Satya Prakash, Stephen Rawls, Andy Rosenbaum, Anjali Shenoy, Saleh Soltan, Mukund Harakere Sridhar, Liz Tan, Fabian Triefenbach, Pan Wei, Haiyang Yu, Shuai Zheng, Gokhan Tur, and Prem Natarajan. 2022. Alexa teacher model: Pretraining and distilling multi-billion-parameter encoders for natural language understanding systems). In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD. ACM.
  16. 16.Jack G. M. FitzGerald. 2020. Stil – simultaneous slot filling, translation, intent classification, and language identification: Initial results using mbart on multiatis++.
  17. 17.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2021. The flores-101 evaluation benchmark for low-resource and multilingual machine translation.
  18. 18.Sonal Gupta, Rushin Shah, Mrinal Mohit, Anuj Kumar, and Mike Lewis. 2018. Semantic parsing for task oriented dialog using hierarchical representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2787–2792, Brussels, Belgium. Association for Computational Linguistics.
  19. 19.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization.
  20. 20.Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020. Cross-lingual ability of multilingual bert: An empirical study.
  21. 21.Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabadi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, Erfan Sadeqi Azer, Niloofar Safi Samghabadi, Mahsa Shafaei, Saber Sheybani, Ali Tazarv, and Yadollah Yaghoobzadeh. 2021. Parsinlu: A suite of language understanding challenges for persian.
  22. 22.Diederik P. Kingma and Jimmy Ba. 2017. Adam: A method for stochastic optimization.
  23. 23.Takumitsu Kudo. 2005. Mecab : Yet another part-of-speech and morphological analyzer.
  24. 24.Surafel M. Lakew, Matteo Negri, and Marco Turchi. 2020. Low resource neural machine translation: A benchmark for five african languages.
  25. 25.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining.
  26. 26.Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer. 2020. Pre-training via paraphrasing.
  27. 27.Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021. MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2950–2962, Online. Association for Computational Linguistics.
  28. 28.Liam Li, Kevin G. Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Moritz Hardt, Benjamin Recht, and Ameet S. Talwalkar. 2018a. Massively parallel hyperparameter tuning. ArXiv, abs/1810.05934.
  29. 29.Xiujun Li, Yu Wang, Siqi Sun, Sarah Panda, Jingjing Liu, and Jianfeng Gao. 2018b. Microsoft dialogue challenge: Building end-to-end task-completion dialogue systems.
  30. 30.Richard Liaw, Eric Liang, Robert Nishihara, Philipp Moritz, Joseph E Gonzalez, and Ion Stoica. 2018. Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118.
  31. 31.Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2019a. Benchmarking natural language understanding services for building conversational agents.
  32. 32.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach.
  34. 34.Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio. 2019. Speech Model Pre-Training for End-to-End Spoken Language Understanding. In Proc. Interspeech 2019, pages 814–818.
  35. 35.Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. Low-resource languages: A review of past work and future challenges.
  36. 36.Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya, Tumisho Mokgonyane, Rethabile Mokoena, and Abiodun Modupe. 2020. Investigating an approach for low resource language dataset creation, curation and classification: Setswana and sepedi. In Proceedings of the first workshop on Resources for African Indigenous Languages, pages 15–20, Marseille, France. European Language Resources Association (ELRA).
  37. 37.Stephen Mayhew, Chen-Tse Tsai, and Dan Roth. 2017. Cheap translation for cross-lingual named entity recognition. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2536–2545, Copenhagen, Denmark. Association for Computational Linguistics.
  38. 38.Massimo Nicosia, Zhongdi Qu, and Yasemin Altun. 2021. Translate & Fill: Improving zero-shot multilingual semantic parsing with synthetic data. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3272–3284, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  39. 39.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In ACL.
  40. 40.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  41. 41.P. J. Price. 1990. Evaluation of spoken language systems: the ATIS domain. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27,1990.
  42. 42.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text.
  43. 43.Shahar Ronen, Bruno Gonçalves, Kevin Z. Hu, Alessandro Vespignani, Steven Pinker, and César A. Hidalgo. 2014. Links that speak: The global language network and its association with global fame. Proceedings of the National Academy of Sciences, 111(52):E5616–E5622.
  44. 44.Subendhu Rongali, Luca Soldaini, Emilio Monti, and Wael Hamza. 2020. Don’t parse, generate! a sequence to sequence architecture for task-oriented semantic parsing. Proceedings of The Web Conference 2020.
  45. 45.Alaa Saade, Alice Coucke, Alexandre Caulier, Joseph Dureau, Adrien Ball, Théodore Bluche, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, and Maël Primet. 2019. Spoken language understanding on the edge.
  46. 46.Sebastian Schuster, Sonal Gupta, Rushin Shah, and Mike Lewis. 2019. Cross-lingual transfer learning for multilingual task oriented dialog. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3795–3805, Minneapolis, Minnesota. Association for Computational Linguistics.
  47. 47.Gary Simons, editor. 2022. Ethnologue: Languages of the World, twenty-fifth edition. SIL International, Dallas, TX, USA.
  48. 48.Heather Simpson, Christopher Cieri, Kazuaki Maeda, Kathryn Baker, and Boyan Onyshkevych. 2008. Human language technology resources for less commonly taught languages: Lessons learned toward creation of basic language resources. Collaboration: interoperability between people in the creation of language resources for less-resourced languages, 7.
  49. 49.Karan Singla, Dogan Can, and Shrikanth Narayanan. 2018. A multi-task approach to learning multilingual representations. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 214–220, Melbourne, Australia. Association for Computational Linguistics.
  50. 50.Stephanie Strassel and Jennifer Tracey. 2016. LORELEI language packs: Data, tools, and resources for technology development in low resource languages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3273–3280, Portorož, Slovenia. European Language Resources Association (ELRA).
  51. 51.Govind Thattai, Gokhan Tur, and Prem Natarajan. 2020. New alexa features: Interactive teaching by customers.
  52. 52.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey. European Language Resources Association (ELRA).
  53. 53.Gokhan Tur, Dilek Hakkani-Tür, and Larry Heck. 2010. What is left to be understood in atis? In 2010 IEEE Spoken Language Technology Workshop, pages 19–24. IEEE.
  54. 54.Gokhan Tur and Renato De Mori. 2011. Spoken language understanding: Systems for extracting semantic information from speech.
  55. 55.Shyam Upadhyay, Manaal Faruqui, Gokhan Tür, Hakkani-Tür Dilek, and Larry Heck. 2018. (almost) zero-shot cross-lingual spoken language understanding. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6034–6038.
  56. 56.Chao Wang, Judith Gaspers, Thi Ngoc Quynh Do, and Hui Jiang. 2021. Exploring cross-lingual transfer learning with unsupervised machine translation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2011–2020, Online. Association for Computational Linguistics.
  57. 57.Ye-Yi Wang, Li Deng, and Alex Acero. 2005. Spoken language understanding. IEEE Signal Processing Magazine, 22:16–31.
  58. 58.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  59. 59.Jiateng Xie, Zhilin Yang, Graham Neubig, Noah A. Smith, and Jaime Carbonell. 2018. Neural cross-lingual named entity recognition with minimal resources. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 369–379, Brussels, Belgium. Association for Computational Linguistics.
  60. 60.Weijia Xu, Batool Haider, and Saab Mansour. 2020. End-to-end slot alignment and recognition for cross-lingual NLU. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5052–5063, Online. Association for Computational Linguistics.
  61. 61.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  62. 62.David Yarowsky, Grace Ngai, and Richard Wicentowski. 2001. Inducing multilingual text analysis tools via robust projection across aligned corpora. In Proceedings of the First International Conference on Human Language Technology Research.
  63. 63.Steve J. Young. 2002. Talking to machines (statistically speaking). In INTERSPEECH.
  64. 64.Su Zhu, Zijian Zhao, Tiejun Zhao, Chengqing Zong, and Kai Yu. 2019. Catslu: The 1st chinese audio-textual spoken language understanding challenge. In 2019 International Conference on Multimodal Interaction, ICMI ’19, pages 521–525, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Fitzgerald, J., et al. “MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 4277–302, https://doi.org/10.18653/v1/2023.acl-long.235.
APA
Fitzgerald, J., Hench, C., Peris, C., Mackie, S., Rottmann, K., Sanchez, A., Nash, A., Urbach, L., Kakarala, V., Singh, R., Ranganath, S., Crist, L., Britan, M., Leeuwis, W., Tur, G., & Natarajan, P. (2023). MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4277–4302. https://doi.org/10.18653/v1/2023.acl-long.235
Chicago
Fitzgerald, J., C. Hench, C. Peris, et al. 2023. “MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4277–4302. https://doi.org/10.18653/v1/2023.acl-long.235.
Harvard
Fitzgerald, J. et al. (2023) “MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4277–4302. Available at: https://doi.org/10.18653/v1/2023.acl-long.235.
Vancouver
1. Fitzgerald J, Hench C, Peris C, et al (2023) MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4277–4302

BibTeX

@inproceedings{fitzgerald-etal-2023-massive,
    title = "{MASSIVE}: A 1{M}-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages",
    author = "FitzGerald, Jack  and
      Hench, Christopher  and
      Peris, Charith  and
      Mackie, Scott  and
      Rottmann, Kay  and
      Sanchez, Ana  and
      Nash, Aaron  and
      Urbach, Liam  and
      Kakarala, Vishesh  and
      Singh, Richa  and
      Ranganath, Swetha  and
      Crist, Laurie  and
      Britan, Misha  and
      Leeuwis, Wouter  and
      Tur, Gokhan  and
      Natarajan, Prem",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.235/",
    doi = "10.18653/v1/2023.acl-long.235",
    pages = "4277--4302"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/