Towards Robust Speech Representation Learning for Thousands of Languages

William ChenWangyou ZhangYifan PengXinjian LiJinchuan TianJiatong ShiXuankai ChangSoumi MaitiKaren LivescuShinji Watanabe

article2024EMNLP73 citationsBest Paper Award

Presents XEUS, a fully open self-supervised speech encoder pre-trained on over one million hours of audio across 4,057 languages that incorporates an acoustic dereverberation objective to achieve state-of-the-art multilingual speech recognition performance.

Listen

Although the world is home to more than 7,000 languages, existing voice technologies generally support only 100 to 150. Expanding coverage is difficult because most languages lack the vast amounts of transcribed speech required to train supervised neural models. While self-supervised learning uses unannotated speech to reduce reliance on labels, existing models still cover relatively few languages and struggle with acoustic distortions such as background noise and room echo, which are common in real-world recordings. Furthermore, leading proprietary systems rely on closed datasets and undisclosed code, hindering independent validation and deployment.

The article introduces and evaluates XEUS (Cross-lingual Encoder for Universal Speech), an open, self-supervised universal speech foundation model designed to scale representation learning to thousands of languages while maintaining robustness against diverse acoustic conditions.

The research combines 1.074 million hours of public speech across 37 existing datasets with a newly curated, publicly released collection of 7,413 hours covering 4,057 languages. This newly introduced corpus incorporates grassroots language preservation recordings and audio dramas across 189 language families, representing a 25-fold expansion in language variety over existing open speech datasets. To improve robustness, the model pairs conventional masked acoustic modeling with noise suppression and a novel dereverberation pre-training task that forces the network to map simulated reverberant audio to clean phonetic representations. The complete 577-million-parameter model was trained on 64 graphics processors across 40 days using optimized open-source software, and all code, configurations, logs, and intermediate checkpoints are released.

Evaluation shows that XEUS establishes a new state of the art on the standard multilingual benchmark (ML-SUPERB) across both 10-minute and 1-hour fine-tuning setups, scoring 956 points and outperforming models like Meta's MMS and w2v-BERT 2.0 despite using 78% less pre-training data than the latter. In speech translation tasks involving under-resourced languages with minimal pre-training data, XEUS achieved a translation quality score of 22.1, representing an approximate 35% relative gain over the next best model (MMS 1B at 15.5). In speech resynthesis tests, the model lowered speech reconstruction word error rates to 10.0% compared to 15.5% for w2v-BERT 2.0 and 27.8% for WavLM Large, while also matching or leading top systems across diverse English speech benchmarks for emotion recognition, speaker identification, and keyword spotting.

These results demonstrate that broad linguistic coverage and acoustic dereverberation objectives allow a foundation model to generalize across low-resource dialects and challenging recording conditions without requiring multi-million-hour closed datasets. For technology organizations and public institutions, XEUS significantly lowers the cost, computational overhead, and technical barriers to developing robust global speech recognition, translation, and generative voice interfaces. It also mitigates licensing and reproducibility risks through fully open artifacts.

Organizations developing multilingual voice services should consider adopting the XEUS foundation architecture and its associated dataset as a base for fine-tuning downstream applications. Because the model weights and data are released under non-commercial licenses and the pre-training distribution exhibits a heavy long-tail skew—where the top 50 languages constitute 99.5% of the total data volume—organizations must conduct pilot evaluations on targeted low-resource languages before production rollout. Further work is recommended to gather more field data for languages with under one hour of speech and to explore large-scale task-specific fine-tuning for mission-critical deployments.

arXiv: 2407.00837wanchichen/espnet
  • Paper: Common Voice: A Massively-Multilingual Speech Corpus, Rosana Ardila et al. (2019). Read this account of a major crowdsourced multilingual speech corpus first: XEUS aggregates existing public speech datasets, making Common Voice’s data-collection and low-resource ASR context useful groundwork.

No sufficiently relevant recommendations were found.

Cover for Towards Robust Speech Representation Learning for Thousands of Languages

Abstract

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world’s 7000+ languages. We propose XEUS, a Cross-lingual Encoder for Universal Speech, trained on over 1 million hours of data across 4057 languages, extending the language coverage of SSL models 4-fold. We combine 1 million hours of speech from existing publicly accessible corpora with a newly created corpus of 7400+ hours from 4057 languages, which will be publicly released. To handle the diverse conditions of multilingual speech data, we augment the typical SSL masked prediction approach with a novel dereverberation objective, increasing robustness. We evaluate XEUS on several benchmarks, and show that it consistently outperforms or achieves comparable results to state-of-the-art (SOTA) SSL models across a variety of tasks. XEUS sets a new SOTA on the ML-SUPERB benchmark: it outperforms MMS 1B and w2v-BERT 2.0 v2 by 0.8% and 4.4% respectively, despite having less parameters or pre-training data. Checkpoints, code, and data are found in https://www.wavlab.org/activities/2024/xeus/.

Table of Contents

  • 1 Introduction
  • 2 Motivation and Related Work
  • 2.1 Speech Representation Learning
  • 2.2 Robust Speech Representations
  • 2.3 Open Foundation Models
  • 3 Data
  • 3.1 Existing Datasets
  • 3.2 MMS-unlab v2
  • 3.3 WikiTongues
  • 3.4 Jesus Dramas
  • 3.5 Final Pre-Training Corpus
  • 4 Self-Supervised Pre-Training
  • 4.1 Masked Prediction and Denoising
  • 4.2 Dereverberation
  • 4.3 Model Architecture
  • 4.4 Pre-Training Settings
  • 5 Downstream Evaluation
  • 5.1 Multilingual Speech Processing
  • 5.1.1 ML-SUPERB
  • 5.1.2 FLEURS
  • 5.1.3 Low-Resource Language Coverage
  • 5.2 Task Universality
  • 5.3 Acoustic Representation Evaluation
  • 6 Conclusion
  • 7 Limitations:
  • 8 Broader Impact and Ethics:
  • Acknowledgement
  • References
  • A Appendix
  • A.1 Data
  • A.2 Ablations
  • A.3 Dereverberation
  • A.4 Pre-Training Details
  • A.5 Engineering Details
  • A.6 Experimental Setups
  • A.6.1 ML-SUPERB:
  • A.6.2 FLEURS:
  • A.6.3 JesusFilm ST:
  • A.6.4 Speech Resynthesis:
  • A.6.5 SUPERB:

Knowls

  1. Knowl 1 — XEUS training corpus spans over 4,000 languages

    data/table

    XEUS was trained on 1.081 million hours of speech assembled from 37 existing publicly accessible corpora and three newly curated sources. The paper reports 1.074 million hours across the existing corpora and 7,413 hours across 4,057 ISO3 languages in the new sources. Those new sources are MMS-unlab v2 (6,700 hours across 4,023 languages), WikiTongues (about 70 hours from 821 recordings and roughly 700 languages or dialects), and Jesus Dramas (about 643 hours across 430 languages; the paper’s prose gives 645 hours). The existing corpora include varied domains and speech styles such as spontaneous and accented speech, code-switching, indigenous-language recordings, and singing. The corpus covers 189 language families. Its language distribution is highly imbalanced: excluding YODAS because of noisy language labels, the paper reports that the 50 most represented languages account for 99.5% of the data, while about 2,000 languages have at least one hour. XEUS’s crawled data is released under non-commercial terms; the authors obtained explicit permission for Global Recordings Network data. The authors also release the model weights, training code and configurations, logs, and more than 200 intermediate checkpoints.

  2. Knowl 2 — Masked prediction combines denoising and dereverberation

    model/method

    XEUS trains a student encoder to predict phonetic pseudo-labels obtained from clean speech, even when its input is masked or acoustically corrupted. The pseudo-labels are produced by extracting representations with a pre-trained WavLabLM MS teacher and clustering them with k-means using 2,048 clusters. Clustering data totals 20,000 hours: 6,000 hours sampled from Common Voice, MLS, and Googlei18n; 6,000 hours from YODAS; 6,000 hours from MMS-unlab v2; and the full FLEURS and BABEL training sets. During student training, noise augmentation is applied with probability 0.2; when used, it is equally likely to add noise from the Deep Noise Suppression Challenge or another utterance from the batch as interference. Dereverberation augmentation is independently applied with probability 0.3, so noise and reverberation can affect the same utterance. For reverberation, an utterance is convolved with a randomly selected room impulse response (RIR), realigned using the location of the RIR’s largest peak, and rescaled to match the original utterance’s energy. Noise is applied before reverberation. The targets remain the clean-speech pseudo-labels, encouraging the student to learn representations that are less sensitive to these corruptions.

  3. Knowl 3 — XEUS uses a 577-million-parameter E-Branchformer encoder

    model/method

    XEUS retains a convolutional speech feature extractor but replaces HuBERT’s Transformer stack with 19 E-Branchformer layers. Each layer has 8 attention heads, a hidden dimension of 1,024, a feed-forward dimension of 4,096, and a convolution kernel size of 31. The complete model has 577 million parameters. XEUS also uses cross-entropy for the masked-unit prediction loss in place of the original HuBERT loss; the authors report that this choice is faster and improves downstream performance.

  4. Knowl 4 — Large-scale pre-training used staged masking and augmentation

    experimental setup

    The 577-million-parameter XEUS model was trained for two passes over its pre-training data, totaling 670,000 steps, on 64 NVIDIA A100 40-GB GPUs. The global batch contained up to 106 minutes of audio. Final pre-training took 40 days and consumed 63,000 GPU-hours. For the first 3,000 steps, training used a masking probability of 0.65, no noise or reverberation augmentation, and an auxiliary cross-entropy loss on encoder layer 10 weighted by 0.3. Thereafter, the auxiliary loss was removed, masking probability was raised to 0.8, and noise and dereverberation augmentation were enabled. The optimizer was Adam, with 32,000 warmup steps and a peak learning rate of 0.0003; training used bfloat16 mixed precision and FlashAttention v2.

  5. Knowl 5 — XEUS achieves strong results across multilingual evaluations

    empirical result

    On ML-SUPERB, which evaluates multilingual speech representations using 10 minutes or 1 hour of labeled data per language, XEUS obtained the highest overall SUPERBs score among the compared models: 956 in both settings, versus 953/948 for MMS 1B and 826/916 for w2v-BERT 2.0 v2. For monolingual ASR, XEUS achieved character error rates (CER; lower is better) of 30.3/25.1 for the 10-minute/1-hour settings. On the normal multilingual ASR task its CER was 21.1/20.1; MMS 1B was better on the 1-hour setting at 18.1. XEUS’s normal multilingual ASR+LID CER was 22.9/19.6. Its few-shot multilingual ASR CER was 33.4/34.1, where MMS 1B performed better; XEUS’s few-shot LID accuracy was 86.4%/91.3%. On FLEURS, a 102-language ASR benchmark with around 6–10 hours of labeled training data per language, XEUS scored 8.9 CER, compared with 8.7 for w2v-BERT 2.0 v2, 9.2 for MMS 1B, and 9.6 for XLS-R 1B. In low-resource speech translation evaluated on Hijazi Arabic, Lumun, and Rajbanshi, XEUS achieved mean chrF 22.1, compared with 15.5 for the next-best model, MMS 1B. The pre-training corpus contained 3 hours of Hijazi Arabic, 0.5 hours of Lumun, and 7 hours of Rajbanshi; among the compared models, only XEUS covered Hijazi Arabic and Lumun, while XEUS and MMS covered Rajbanshi.

  6. Knowl 6 — XEUS is competitive on English SUPERB tasks

    empirical result

    On the English-only SUPERB benchmark, XEUS was compared with WavLM Large using frozen speech encoders and task-specific probes. XEUS achieved the best results on four tasks: ASR word error rate (WER) 3.34 versus 3.44; keyword-spotting accuracy (KS) 98.32 versus 97.86; diarization error rate (SD) 3.11 versus 3.24; and emotion-recognition accuracy (ER) 71.08 versus 70.62. For the other tasks, XEUS and WavLM Large scored, respectively: phoneme error rate (PR), 3.21 and 3.06; query-by-example score (QbE), 7.49 and 8.86; intent-classification accuracy (IC), 98.70 and 99.31; slot-filling F1, 90.05 and 92.21; slot-filling concept error rate, 21.49 and 18.36; speaker-identification accuracy (SID), 91.70 and 95.49; and automatic-speaker-verification equal error rate (ASV), 4.16 and 3.77. Lower is better for PR, WER, slot-filling concept error rate, ASV, and SD; higher is better for the other listed metrics.

  7. Knowl 7 — XEUS produces strong speech-resynthesis results

    empirical result

    On accented-English VCTK speech resynthesis, XEUS representations were converted to discrete units using a 100-entry codebook and then synthesized with a HiFi-GAN vocoder. The best-performing feature layer for each encoder was selected using development-set MOSNet scores; the reported test-set results were: XEUS, MOS 3.23, WER 10.0, F0 error 0.25, and mel-cepstral distortion (MCD) 3.80; w2v-BERT 2.0 v2, 3.21, 15.5, 0.27, and 3.92; WavLM Large, 3.20, 27.8, 0.26, and 4.55. Higher MOS and lower WER, F0 error, and MCD are preferred. XEUS had the best reported value on each metric.

  8. Knowl 8 — Dereverberation improves ASR in a controlled HuBERT test

    empirical result

    The paper tested dereverberation augmentation separately from XEUS by training two HuBERT Base models on LibriSpeech 960 hours, one without augmentation and one with the proposed reverberation simulation. After fine-tuning on 10 hours of LibriLight ASR data, the dereverberation model reduced test-clean WER from 13.1 to 12.7 and test-other WER from 22.6 to 21.1. The authors report these as relative reductions of 3.1% and 6.9%, respectively, with the larger improvement on the noisier test-other set.

  9. Knowl 9 — ESPnet optimizations improve large-scale training throughput

    model/method

    To make large multi-GPU speech pre-training more efficient in ESPnet, the authors removed two unnecessary GPU synchronization steps and disabled synchronization on gradient-accumulation iterations that did not perform an optimizer update. They also fixed batch over-allocation and distributed batches across GPUs according to sequence length, reducing padding and imbalance in backward-pass workloads. The paper reports a 120% throughput increase for GPU-synchronization changes. For batch optimization, it reports a 60% throughput increase and a 113% decrease in relative memory usage; its prose separately describes memory use as more than halved at fixed batch size. These changes were especially useful for clusters without InfiniBand inter-node communication.

  10. Knowl 10 — The authors identify long-tail and evaluation limitations

    limitation

    Although XEUS’s pre-training corpus covers 4,057 languages, many have less than one hour of speech, so the authors expect performance on those languages to remain substantially below performance on better-resourced languages. They evaluated only three languages outside the standard multilingual benchmarks in their low-resource speech-translation study, in part because collecting and manually cleaning evaluation data is demanding. The paper also notes that its broad task coverage relies mostly on lightweight benchmarks and limited hyperparameter tuning, rather than the large-scale downstream fine-tuning used in some state-of-the-art task-specific systems.

Coverage note — The knowls omit the exhaustive per-corpus licensing inventory and detailed per-task downstream hyperparameter recipes; these are supporting catalog and protocol details rather than additional central findings.

References

  1. 1.aidatatang_200zh, a free Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd.
  2. 2.Afroz Ahamad, Ankit Anand, and Pranesh Bhargava. 2020. AccentDB: A database of non-native English accents to assist neural speech recognition. In LREC 2020.
  3. 3.AI4Bharat. 2020. NPTEL2020: Indian English Speech Dataset.
  4. 4.Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Jason Phang, Shivanshu Purohit, Hailey Schoelkopf, Dashiell Stander, Tri Songz, Curt Tigges, Benjamin Thérien, Phil Wang, and Samuel Weinbach. 2023. GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch.
  5. 5.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In LREC 2020, pages 4218–4222.
  6. 6.Peter K. Austin and Julia Sallabank, editors. 2011. The Cambridge Handbook of Endangered Languages. Cambridge University Press.
  7. 7.Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2022. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale. In Interspeech 2022, pages 2278–2282.
  8. 8.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022. Data2vec: A general framework for self-supervised learning in speech, vision and language. In ICML 2022, pages 1298–1312.
  9. 9.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In NeurIPS 2020, volume 33.
  10. 10.Jeong-Uk Bang et al. 2020. KsponSpeech: Korean spontaneous speech corpus for automatic speech recognition. Applied Sciences.
  11. 11.Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023a. SeamlessM4T-massively multilingual & multimodal machine translation. arxiv:2308.11596.
  12. 12.Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, et al. 2023b. Seamless: Multilingual expressive and streaming speech translation. arxiv:2312.05187.
  13. 13.Dan Berrebbi, Jiatong Shi, Brian Yan, Osbel López-Francisco, Jonathan Amith, and Shinji Watanabe. 2022. Combining spectral and self-supervised features for low resource speech recognition and translation. In Interspeech 2022.
  14. 14.Alan W Black. 2019. CMU Wilderness Multilingual Speech Dataset. In ICASSP 2019.
  15. 15.Hui Bu et al. 2017. AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline. In O-COCOSDA.
  16. 16.Ronald Cardenas, Rodolfo Zevallos, Reynaldo Baquerizo, and Luis Camacho. 2018. Siminchik: A speech corpus for preservation of southern Quechua. ISINLP.
  17. 17.Jean Carletta. 2007. Unleashing the killer corpus: experiences in creating the multi-everything AMI meeting corpus. Springer.
  18. 18.Heng-Jui Chang and James Glass. 2023. R-spin: Efficient speaker and noise-invariant representation learning with acoustic pieces. arXiv preprint arXiv:2311.09117.
  19. 19.Xuankai Chang, Takashi Maekaku, Pengcheng Guo, Jing Shi, Yen-Ju Lu, Aswin Shanmugam Subramanian, Tianzi Wang, Shu-wen Yang, Yu Tsao, Hung-yi Lee, and Shinji Watanabe. 2021. An exploration of self-supervised pretrained representations for end-to-end speech recognition. In ASRU 2021, pages 228–235.
  20. 20.Guoguo Chen et al. 2021. GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio. In Interspeech 2021.
  21. 21.Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE JSTSP.
  22. 22.William Chen, Xuankai Chang, Yifan Peng, Zhaoheng Ni, Soumi Maiti, and Shinji Watanabe. 2023a. Reducing Barriers to Self-Supervised Learning: HuBERT Pre-training with Academic Compute. In Interspeech 2023.
  23. 23.William Chen, Jiatong Shi, Brian Yan, Dan Berrebbi, Wangyou Zhang, Yifan Peng, Xuankai Chang, Soumi Maiti, and Shinji Watanabe. 2023b. Joint prediction and denoising for large-scale multilingual self-supervised learning. In ASRU 2023.
  24. 24.William Chen, Brian Yan, Jiatong Shi, Yifan Peng, Soumi Maiti, and Shinji Watanabe. 2023c. Improving massively multilingual ASR with auxiliary CTC objectives. In ICASSP 2023.
  25. 25.Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu. 2022. Self-supervised learning with random-projection quantizer for speech recognition. In International Conference on Machine Learning. PMLR.
  26. 26.Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021. w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In ASRU 2021.
  27. 27.Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli. 2021. Unsupervised Cross-Lingual Representation Learning for Speech Recognition. In Interspeech 2021, pages 2426–2430.
  28. 28.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In ACL 2020, pages 8440–8451, Online. Association for Computational Linguistics.
  29. 29.Alexis Conneau et al. 2022. FLEURS: Few-shot learning evaluation of universal representations of speech. In SLT 2022.
  30. 30.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In NeurIPS 2022.
  31. 31.Ethnologue. 2017. How many languages in the world are unwritten?
  32. 32.Daniel Galvez, Greg Diamos, Juan Manuel Ciro Torres, Juan Felipe Cerón, Keith Achorn, Anjali Gopi, David Kanter, Max Lam, Mark Mazumder, and Vijay Janapa Reddi. 2021. The People’s Speech: A large-scale diverse English speech recognition dataset for commercial usage. In NeurIPS 2021.
  33. 33.J.J. Godfrey et al. 1992. SWITCHBOARD: telephone speech corpus for research and development. In ICASSP 1992.
  34. 34.Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In ICML 2006, pages 369–376.
  35. 35.Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838.
  36. 36.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. Conformer: Convolution-augmented transformer for speech recognition. In Interspeech 2020.
  37. 37.Inga R. Helgadóttir, Anna Björk Nikulásdóttir, Michal Borský, Judy Y. Fong, Róbert Kjaran, and Jón Guðnason. 2019. The Althingi ASR system. In Interspeech 2019.
  38. 38.François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve. 2018. TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation. In Speech and Computer: 20th International Conference, SPECOM 2018. Springer.
  39. 39.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM TALSP.
  40. 40.Kuan-Po Huang, Yu-Kuan Fu, Tsu-Yuan Hsu, Fabian Ritter Gutierrez, Fan-Lin Wang, Liang-Hsuan Tseng, Yu Zhang, and Hung-yi Lee. 2023. Improving generalizability of distilled self-supervised speech processing models under distorted settings. In SLT 2022, pages 1112–1119.
  41. 41.IARPA. The Babel Program.
  42. 42.J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux. 2020. LibriLight: A benchmark for ASR with limited or no supervision. In ICASSP 2020.
  43. 43.Kwangyoun Kim, Felix Wu, Yifan Peng, Jing Pan, Prashant Sridhar, Kyu J Han, and Shinji Watanabe. 2023. E-branchformer: Branchformer with enhanced merging for speech recognition. In SLT 2023.
  44. 44.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. ICLR 2015.
  45. 45.Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, and Sanjeev Khudanpur. 2017. A study on data augmentation of reverberant speech for robust speech recognition. In ICASSP 2017.
  46. 46.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. Hifi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. NeurIPS 2020, 33:17022–17033.
  47. 47.Xinjian Li, Florian Metze, David R. Mortensen, Alan W Black, and Shinji Watanabe. 2022. ASR2K: Speech Recognition for Around 2000 Languages without Audio. In Interspeech 2022.
  48. 48.Xinjian Li, Shinnosuke Takamichi, Takaaki Saeki, William Chen, Sayaka Shiota, and Shinji Watanabe. 2023. YODAS: Youtube-oriented dataset for audio and speech. In ASRU 2023.
  49. 49.Alexander H Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu, and Jim Glass. 2024. DinoSR: Self-distillation and online clustering for self-supervised speech representation learning. NeurIPS 2024, 36.
  50. 50.Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023. LLM360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550.
  51. 51.Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang. 2019. MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion. In Interspeech 2019, pages 1541–1545.
  52. 52.Dau-Cheng Lyu, Tien-Ping Tan, Eng Siong Chng, and Haizhou Li. 2010. SEAME: a Mandarin-English code-switching speech corpus in south-east Asia. In Interspeech 2010.
  53. 53.Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, and Shinji Watanabe. 2024. VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks. In ICASSP 2024, pages 13326–13330. IEEE.
  54. 54.Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Zhao Yang, Jinjie Ni, Chong Zhang, Yukun Ma, Chongjia Ni, Eng Siong Chng, and Bin Ma. 2023. De’hubert: Disentangling noise in a self-supervised model for robust speech recognition. In ICASSP 2023, pages 1–5.
  55. 55.Jumon Nozaki and Tatsuya Komatsu. 2021. Relaxing the conditional independence assumption of CTC-based ASR by conditioning on intermediate predictions. In Interspeech 2021, pages 3735–3739.
  56. 56.Patrick K. O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D. Shulman, Boris Ginsburg, Shinji Watanabe, and Georg Kucsko. 2021. SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition. In Interspeech 2021.
  57. 57.Vassil Panayotov et al. 2015. Librispeech: An ASR corpus based on public domain audio books. In ICASSP 2015.
  58. 58.Douglas B Paul and Janet Baker. 1992. The design for the Wall Street Journal-based CSR corpus. In Workshop on Speech and Natural Language.
  59. 59.Yifan Peng, Siddharth Dalmia, Ian Lane, and Shinji Watanabe. 2022. Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding. In ICML 2022.
  60. 60.Yifan Peng, Kwangyoun Kim, Felix Wu, Brian Yan, Siddhant Arora, William Chen, Jiyang Tang, Suwon Shon, Prashant Sridhar, and Shinji Watanabe. 2023a. A comparative study on e-branchformer vs conformer in speech recognition, translation, and understanding tasks. In Interspeech 2023.
  61. 61.Yifan Peng, Yui Sudo, Muhammad Shakeel, and Shinji Watanabe. 2024a. OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification. arXiv preprint arXiv:2402.12654.
  62. 62.Yifan Peng, Jinchuan Tian, William Chen, Siddhant Arora, Brian Yan, Yui Sudo, Muhammad Shakeel, Kwanghee Choi, Jiatong Shi, Xuankai Chang, et al. 2024b. OWSM v3. 1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer. arXiv preprint arXiv:2401.16658.
  63. 63.Yifan Peng, Jinchuan Tian, Brian Yan, Dan Berrebbi, Xuankai Chang, Xinjian Li, Jiatong Shi, Siddhant Arora, William Chen, Roshan Sharma, Wangyou Zhang, Yui Sudo, Muhammad Shakeel, Jee weon Jung, Soumi Maiti, and Shinji Watanabe. 2023b. Reproducing Whisper-style training using an open-source toolkit and publicly available data. In ASRU 2023.
  64. 64.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021. Speech Resynthesis from Discrete Disentangled Self-Supervised Representations. In Interspeech 2021.
  65. 65.Matt Post et al. 2013. Improved speech-to-text translation with the fisher and callhome Spanish-English speech translation corpus. In IWSLT 2013.
  66. 66.Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2023. Scaling speech technology to 1,000+ languages. arxiv:2305.13516.
  67. 67.Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. MLS: A large-scale multilingual dataset for speech research. In Interspeech 2020, pages 2757–2761.
  68. 68.Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In ICML 2023.
  69. 69.Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al. 2021. Speechbrain: A general-purpose speech toolkit. arXiv preprint arXiv:2106.04624.
  70. 70.Chandan K.A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, and Sriram Srinivasan. 2021. INTERSPEECH 2021 Deep Noise Suppression Challenge. In Interspeech 2021.
  71. 71.Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W. Oard, and Matt Post. 2021. Multilingual TEDx corpus for speech recognition and translation. In Interspeech 2021.
  72. 72.Ramon Sanabria, Nikolay Bogoychev, Nina Markl, Andrea Carmantini, Ondrej Klejch, and Peter Bell. 2023. The Edinburgh international accents of English corpus: Towards the democratization of English ASR. In ICASSP 2023.
  73. 73.Jiatong Shi, Jonathan D Amith, Xuankai Chang, Siddharth Dalmia, Brian Yan, and Shinji Watanabe. 2021a. Highland Puebla Nahuatl speech translation corpus for endangered language documentation. In AmericasNLP 2021.
  74. 74.Jiatong Shi, Jonathan D Amith, Rey Castillo García, Esteban Guadalupe Sierra, Kevin Duh, and Shinji Watanabe. 2021b. Leveraging end-to-end ASR for endangered language documentation: An empirical study on Yolóxochitl Mixtec. In EACL 2021.
  75. 75.Jiatong Shi, Dan Berrebbi, William Chen, En-Pei Hu, Wei-Ping Huang, Ho-Lam Chung, Xuankai Chang, Shang-Wen Li, Abdelrahman Mohamed, Hung yi Lee, and Shinji Watanabe. 2023a. ML-SUPERB: Multilingual Speech Universal PERformance Benchmark. In Interspeech 2023.
  76. 76.Jiatong Shi, William Chen, Dan Berrebbi, Hsiu-Hsuan Wang, Wei-Ping Huang, En-Pei Hu, Ho-Lam Chuang, Xuankai Chang, Yuxun Tang, Shang-Wen Li, Abdelrahman Mohamed, Hung-Yi Lee, and Shinji Watanabe. 2023b. Findings of the 2023 ML-SUPERB challenge: Pre-training and evaluation over more languages and beyond. In ASRu 2023.
  77. 77.Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, and Anna Sun. 2024. Multi-resolution HuBERT: Multi-resolution speech self-supervised learning with masked unit prediction. In ICLR 2024.
  78. 78.Anna Slizhikova et al. 2020. Russian Open Speech To Text (STT/ASR) Dataset.
  79. 79.Inc Smule. 2019. DAMP-MVP: Digital Archive of Mobile Performances - Smule Multilingual Vocal Performance 300x30x2.
  80. 80.Per Erik Solberg and Pablo Ortiz. 2022. The Norwegian parliamentary speech corpus. In LREC 2022, Marseille, France.
  81. 81.Jörgen Valk and Tanel Alumäe. 2021. VOXLINGUA107: A dataset for spoken language recognition. In SLT 2021.
  82. 82.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS 2017.
  83. 83.VoxForge. VoxForge.
  84. 84.Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Sravya Popuri, Dmytro Okhonko, and Juan Pino. 2020. Fairseq S2T: Fast speech-to-text modeling with fairseq. arxiv:2010.05171.
  85. 85.Changhan Wang et al. 2021. VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation. In ACL 2021.
  86. 86.Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018. ESPnet: End-to-end speech processing toolkit. In Interspeech 2018.
  87. 87.Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi. 2017. Hybrid CTC/attention architecture for end-to-end speech recognition. IEEE Journal of Selected Topics in Signal Processing.
  88. 88.Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee. 2021. SUPERB: Speech Processing Universal PERformance Benchmark. In Interspeech 2021.
  89. 89.BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  90. 90.Junichi Yamagishi et al. 2019. CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit.
  91. 91.Guanrou Yang, Ziyang Ma, Zhisheng Zheng, Yakun Song, Zhikang Niu, and Xie Chen. 2023. Fast-HuBERT: an efficient training framework for self-supervised speech representation learning. In ASRU 2023.
  92. 92.Zehui Yang, Yifan Chen, Lei Luo, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Ji Xu, Yaohui Jin, Qingqing Zhang, Pengyuan Zhang, Lei Xie, and Yonghong Yan. 2022. Open source MagicData-RAMC: A rich annotated mandarin conversational (RAMC) speech dataset. In Interspeech 2022, pages 1736–1740.
  93. 93.Yue Yin, Daijiro Mori, et al. 2023. ReazonSpeech: A Free and Massive Corpus for Japanese ASR.
  94. 94.Binbin Zhang et al. 2022. WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition. In ICASSP 2022.
  95. 95.Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, et al. 2023. Google USM: Scaling automatic speech recognition beyond 100 languages. arxiv:2303.01037.
  96. 96.Qiu-Shi Zhu, Long Zhou, Jie Zhang, Shu-Jie Liu, Yu-Chen Hu, and Li-Rong Dai. 2023. Robust data2vec: Noise-robust speech representation learning for asr by combining regression and improved contrastive learning. In ICASSP 2023, pages 1–5.

Citation

MLA
Chen, W., et al. “Towards Robust Speech Representation Learning for Thousands of Languages”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 10205–24, https://doi.org/10.18653/v1/2024.emnlp-main.570.
APA
Chen, W., Zhang, W., Peng, Y., Li, X., Tian, J., Shi, J., Chang, X., Maiti, S., Livescu, K., & Watanabe, S. (2024). Towards Robust Speech Representation Learning for Thousands of Languages. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10205–10224. https://doi.org/10.18653/v1/2024.emnlp-main.570
Chicago
Chen, W., W. Zhang, Y. Peng, et al. 2024. “Towards Robust Speech Representation Learning for Thousands of Languages”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10205–24. https://doi.org/10.18653/v1/2024.emnlp-main.570.
Harvard
Chen, W. et al. (2024) “Towards Robust Speech Representation Learning for Thousands of Languages”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10205–10224. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.570.
Vancouver
1. Chen W, Zhang W, Peng Y, Li X, Tian J, Shi J, Chang X, Maiti S, Livescu K, Watanabe S (2024) Towards Robust Speech Representation Learning for Thousands of Languages. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10205–10224

BibTeX

@inproceedings{chen-etal-2024-towards-robust,
    title = "Towards Robust Speech Representation Learning for Thousands of Languages",
    author = "Chen, William  and
      Zhang, Wangyou  and
      Peng, Yifan  and
      Li, Xinjian  and
      Tian, Jinchuan  and
      Shi, Jiatong  and
      Chang, Xuankai  and
      Maiti, Soumi  and
      Livescu, Karen  and
      Watanabe, Shinji",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.570/",
    doi = "10.18653/v1/2024.emnlp-main.570",
    pages = "10205--10224"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/