MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
Jack FitzGeraldChristopher HenchCharith PerisScott MackieKay RottmannAna SanchezAaron NashLiam UrbachVishesh KakaralaRicha Singh
Presents a large-scale, parallel spoken language understanding dataset of one million virtual assistant utterances across 51 typologically diverse languages to advance intent classification and slot-filling benchmarks for low-resource NLP.
Modern voice-based virtual assistants still support only a fraction of the world’s languages, largely due to a lack of realistic, high-quality, and task-specific labeled data. While massively multilingual foundation models have advanced rapidly, standardized evaluation benchmarks across diverse languages have failed to keep pace. The article introduces MASSIVE, a one-million-example dataset designed to evaluate and advance natural language understanding across 51 typologically diverse languages spanning 29 genera, 18 domains, 60 intents, and 55 slot types.
To build MASSIVE, the authors localized the English-only SLURP dataset into 50 additional languages using professional crowd-workers and a two-stage localization process that separated slot translation from sentence assembly. They then established baseline performance benchmarks by fine-tuning multilingual pre-trained models—specifically XLM-R and mT5—under both full-dataset multilingual training and zero-shot transfer conditions.
The findings show that full-dataset training consistently outperforms zero-shot cross-lingual transfer by substantial margins. Across all languages, full training achieved average exact match accuracies between 63.7% and 66.6%, whereas zero-shot transfer dropped to 28.8%–38.7% (a 25 to 37 percentage-point decline). Furthermore, zero-shot performance displayed severe variance across languages; for example, mT5 Text-to-Text showed a 15-point gap between highest and lowest-performing locales under full training, expanding to a 44-point gap under zero-shot conditions. Zero-shot accuracy strongly correlated with the volume of pre-training data available for each language (0.54 correlation for exact match), but supplying equal amounts of target-language fine-tuning data weakened this correlation to 0.42. Finally, languages using Latin scripts and Germanic genera performed best overall, while space-optional languages (such as Japanese) suffered steep performance drops—down to 9.4% zero-shot exact match—due to artificial character-spacing constraints.
These results demonstrate that relying solely on zero-shot cross-lingual transfer introduces severe performance risks and geographic disparities for production virtual assistants. Providing equal, high-quality localized training data across languages effectively mitigates the inherent biases of multilingual pre-training corpora. Organizations seeking to expand conversational assistants should invest in localized fine-tuning sets rather than relying exclusively on zero-shot generalization. Subsequent efforts should explore more sophisticated tokenization tools for unspaced scripts, scale evaluations to larger model architectures, and refine crowd-sourcing pipelines to minimize translation artifacts.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). Introduces XLM-R, one of the primary foundational multilingual encoder architectures evaluated as a baseline across all languages in MASSIVE.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). Presents mT5, the massively multilingual sequence-to-sequence model used directly in MASSIVE to benchmark generative intent classification and slot-filling.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Establishes standard cross-lingual evaluation methodologies and translation-based benchmarks that laid the groundwork for large-scale multilingual NLU evaluation.
- Paper: MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, Paweł Budzianowski et al. (2018). Provides foundational frameworks for task-oriented dialogue, intent tracking, and slot labeling that informed modern virtual assistant datasets like SLURP and MASSIVE.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). Demonstrates large-scale open multilingual modeling across dozens of language families, contextualizing the multilingual transfer evaluated in MASSIVE.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). Extends massive parallel multilingual evaluation to 122 language variants to systematically compare multilingual encoders against newer generative LLMs.
- Paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model, Ahmet Üstün et al. (2024). Builds upon multilingual instruction datasets and models like mT5 to develop a specialized 101-language open-source generative instruction-following system.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Scales multilingual representation learning beyond the 50-to-100 language tier to support over 500 under-resourced languages across diverse NLU benchmarks.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Assesses modern generative large language models against fine-tuned multilingual baselines across 70 typologically diverse languages.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). Expands multilingual and cross-modal evaluation suites across 83 languages to benchmark frontier models on downstream task understanding.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). Generalizes massive multilingual evaluation to text embedding representations across more than 250 languages and numerous downstream tasks.
- Paper: CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language Pairs, Abhik Bhattacharjee et al. (2023). Applies massively multilingual sequence-to-sequence modeling to non-English-centric text generation across more than 1,500 language pairs.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Pushes the boundaries of multilingual language technologies to an unprecedented scale, covering over 1,600 global languages.
