Towards Robust Speech Representation Learning for Thousands of Languages
William ChenWangyou ZhangYifan PengXinjian LiJinchuan TianJiatong ShiXuankai ChangSoumi MaitiKaren LivescuShinji Watanabe
Presents XEUS, a fully open self-supervised speech encoder pre-trained on over one million hours of audio across 4,057 languages that incorporates an acoustic dereverberation objective to achieve state-of-the-art multilingual speech recognition performance.
Although the world is home to more than 7,000 languages, existing voice technologies generally support only 100 to 150. Expanding coverage is difficult because most languages lack the vast amounts of transcribed speech required to train supervised neural models. While self-supervised learning uses unannotated speech to reduce reliance on labels, existing models still cover relatively few languages and struggle with acoustic distortions such as background noise and room echo, which are common in real-world recordings. Furthermore, leading proprietary systems rely on closed datasets and undisclosed code, hindering independent validation and deployment.
The article introduces and evaluates XEUS (Cross-lingual Encoder for Universal Speech), an open, self-supervised universal speech foundation model designed to scale representation learning to thousands of languages while maintaining robustness against diverse acoustic conditions.
The research combines 1.074 million hours of public speech across 37 existing datasets with a newly curated, publicly released collection of 7,413 hours covering 4,057 languages. This newly introduced corpus incorporates grassroots language preservation recordings and audio dramas across 189 language families, representing a 25-fold expansion in language variety over existing open speech datasets. To improve robustness, the model pairs conventional masked acoustic modeling with noise suppression and a novel dereverberation pre-training task that forces the network to map simulated reverberant audio to clean phonetic representations. The complete 577-million-parameter model was trained on 64 graphics processors across 40 days using optimized open-source software, and all code, configurations, logs, and intermediate checkpoints are released.
Evaluation shows that XEUS establishes a new state of the art on the standard multilingual benchmark (ML-SUPERB) across both 10-minute and 1-hour fine-tuning setups, scoring 956 points and outperforming models like Meta's MMS and w2v-BERT 2.0 despite using 78% less pre-training data than the latter. In speech translation tasks involving under-resourced languages with minimal pre-training data, XEUS achieved a translation quality score of 22.1, representing an approximate 35% relative gain over the next best model (MMS 1B at 15.5). In speech resynthesis tests, the model lowered speech reconstruction word error rates to 10.0% compared to 15.5% for w2v-BERT 2.0 and 27.8% for WavLM Large, while also matching or leading top systems across diverse English speech benchmarks for emotion recognition, speaker identification, and keyword spotting.
These results demonstrate that broad linguistic coverage and acoustic dereverberation objectives allow a foundation model to generalize across low-resource dialects and challenging recording conditions without requiring multi-million-hour closed datasets. For technology organizations and public institutions, XEUS significantly lowers the cost, computational overhead, and technical barriers to developing robust global speech recognition, translation, and generative voice interfaces. It also mitigates licensing and reproducibility risks through fully open artifacts.
Organizations developing multilingual voice services should consider adopting the XEUS foundation architecture and its associated dataset as a base for fine-tuning downstream applications. Because the model weights and data are released under non-commercial licenses and the pre-training distribution exhibits a heavy long-tail skew—where the top 50 languages constitute 99.5% of the total data volume—organizations must conduct pilot evaluations on targeted low-resource languages before production rollout. Further work is recommended to gather more field data for languages with under one hour of speech and to explore large-scale task-specific fine-tuning for mission-critical deployments.
- Paper: Common Voice: A Massively-Multilingual Speech Corpus, Rosana Ardila et al. (2019). Read this account of a major crowdsourced multilingual speech corpus first: XEUS aggregates existing public speech datasets, making Common Voice’s data-collection and low-resource ASR context useful groundwork.
No sufficiently relevant recommendations were found.
