SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

Hsiang-Sheng TsaiHeng-Jui ChangWen-Chin HuangZili HuangKushal LakhotiaShu-Wen YangShuyan DongAndy T. LiuCheng-I LaiJiatong Shi

article2022ACL135 citations

Presents SUPERB-SG, an expanded speech benchmark that adds challenging semantic and generative tasks to standard evaluation suites, providing a computationally efficient protocol to test how well frozen self-supervised models generalize across diverse speech processing applications.

Listen

Transfer learning using self-supervised learning models has rapidly advanced speech processing by learning rich representations from unlabelled audio data. However, previous benchmarking frameworks like SUPERB primarily evaluated these models on basic classification and shallow linguistic tasks. This left significant uncertainty regarding how well pre-trained models handle more complex, real-world requirements such as high-level semantics, multilingual translation, and audio generation.

The article introduces and evaluates SUPERB-SG, an enhanced benchmark designed to rigorously assess the semantic and generative capabilities of speech pre-trained models across diverse, challenging tasks. To provide an accessible, compute-efficient, and standardized evaluation, the framework freezes upstream pre-trained model weights and trains only lightweight task-specific prediction heads using a weighted-sum mechanism across model layers. The benchmark tests 15 distinct upstream models across five challenging downstream tasks: speech translation, out-of-domain automatic speech recognition across multiple languages and spontaneous speech, voice conversion, speech separation, and speech enhancement.

The evaluation produced several critical findings. First, no single pre-trained model universally dominates all tasks. However, HuBERT Large and wav2vec 2.0 Large consistently achieve the best overall results on tasks requiring deep semantic and linguistic understanding, such as speech translation and out-of-domain speech recognition. Second, for low-level audio generation and enhancement tasks like speech separation and speech enhancement, complex self-supervised models provide minimal improvement over traditional baseline features like Log Mel-Filterbanks, showing that high-level abstract representations do not substantially benefit raw acoustic reconstruction. Third, correlation and clustering analyses revealed strong alignment among content-heavy tasks, while separation and enhancement tasks clustered separately as they depend primarily on low-level acoustic details. Finally, robustness experiments demonstrated that upstream model performance rankings remain highly consistent across varying downstream model sizes and when training data is reduced to as low as 5%, though models fail across tasks when supervision is reduced to 1%.

These findings indicate that choosing a speech foundation model requires careful alignment with the target application rather than assuming the largest self-supervised model is universally superior. For high-level semantic and recognition systems, adopting advanced models like HuBERT Large provides substantial performance gains and cost efficiencies over training models from scratch. Conversely, for low-level audio enhancement and separation pipelines, relying on lightweight traditional acoustic baselines can reduce computational overhead without sacrificing quality. Decision-makers should leverage the open-source SUPERB-SG framework to benchmark customized models under resource-constrained conditions while ensuring downstream applications maintain sufficient labelled supervision for task-specific adaptation.

arXiv: 2203.06849
Cover for SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

Abstract

Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years. In speech, a model pre-trained by self-supervised learning transfers remarkably well on multiple tasks. However, the lack of a consistent evaluation methodology is limiting towards a holistic understanding of the efficacy of such models. SUPERB was a step towards introducing a common benchmark to evaluate pre-trained models across various speech tasks. In this paper, we introduce SUPERB-SG, a new benchmark focused on evaluating the semantic and generative capabilities of pre-trained models by increasing task diversity and difficulty over SUPERB. We use a lightweight methodology to test the robustness of representations learned by pre-trained models under shifts in data domain and quality across different types of tasks. It entails freezing pre-trained model parameters, only using simple task-specific trainable heads. The goal is to be inclusive of all researchers, and encourage efficient use of computational resources. We also show that the task diversity of SUPERB-SG coupled with limited task supervision is an effective recipe for evaluating the generalizability of model representation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 SUPERB-SG
  • 3.1 Tasks and Datasets
  • 3.1.1 Speech Translation
  • 3.1.2 Out-of-domain ASR
  • 3.1.3 Voice Conversion
  • 3.1.4 Speech Separation
  • 3.1.5 Speech Enhancement
  • 3.2 Self-supervised Models
  • 4 Experimental Setup
  • 5 Results and Discussion
  • 5.1 Main result
  • 5.2 Correlation between tasks
  • 5.3 Robustness of SUPERB-SG
  • 5.3.1 Downstream model
  • 5.3.2 Training data size
  • 6 Conclusion
  • Ethics
  • References
  • A Complete Out-of-domain ASR Results
  • B Responsible NLP Research Checklist
  • B.1 Did you discuss the limitations of your work?
  • B.2 Did you discuss any potential risks of your work?
  • B.3 Do the abstract and introduction summarize the paper’s main claims?
  • B.4 Did you use or create scientific artifacts?
  • B.4.1 Did you cite the creators of artifacts you used?
  • B.4.2 Did you discuss the license or terms for use and/or distribution of any artifacts?
  • B.4.3 Did you discuss if your use of existing artifact(s) was consistent with their intended use, provided that it was specified? For the artifacts you create, do you specify intended use and whether that is compatible with the original access conditions (in particular, derivatives of data accessed for research purposes should not be used outside of research contexts)?
  • B.4.4 Did you discuss the steps taken to check whether the data that was collected/used contains any information that names or uniquely identifies individual people or offensive content, and the steps taken to protect / anonymize it?
  • B.4.5 Did you provide documentation of the artifacts, e.g., coverage of domains, languages, and linguistic phenomena, demographic groups represented, etc.?
  • B.4.6 Did you report relevant statistics like the number of examples, details of train/test/dev splits, etc. for the data that you used/created?
  • B.5 Did you run computational experiments?
  • B.5.1 Did you report the number of parameters in the models used, the total computational budget (e.g., GPU hours), and computing infrastructure used?
  • B.5.2 Did you discuss the experimental setup, including hyperparameter search and best-found hyperparameter values?
  • B.5.3 Did you report descriptive statistics about your results (e.g., error bars around results, summary statistics from sets of experiments), and is it transparent whether you are reporting the max, mean, etc. or just a single run?
  • B.5.4 If you used existing packages (e.g., for preprocessing, for normalization, or for evaluation), did you report the implementation, model, and parameter settings used (e.g., NLTK, Spacy, ROUGE, etc.)?
  • B.6 Did you use human annotators (e.g., crowdworkers) or research with human subjects?

Knowls

  1. Knowl 1 — SUPERB-SG Benchmark Evaluation Protocol and Layer-Wise Feature Aggregation

    model/method

    SUPERB-SG evaluates pre-trained speech self-supervised learning (SSL) representations across semantic and generative downstream tasks using a lightweight, frozen-upstream protocol.

    Given an upstream SSL model with LL hidden layers, the model parameters remain strictly frozen during downstream training. For an input audio sequence of length TT, frame-level representations ht(l)∈RDl\mathbf{h}_t^{(l)} \in \mathbb{R}^{D_l} are extracted from each layer l∈{1,…,L}l \in \{1, \dots, L\} at time step t∈{1,…,T}t \in \{1, \dots, T\}. When layer dimensions differ, linear projections align them to a shared dimension DD.

    A trainable layer-weighting mechanism calculates a scalar weight αl\alpha_l for each layer to compute a task-specific weighted representation:

    ht=∑l=1Lwlht(l),where wl=exp⁡(αl)∑j=1Lexp⁡(αj)\mathbf{h}_t = \sum_{l=1}^{L} w_l \mathbf{h}_t^{(l)}, \quad \text{where } w_l = \frac{\exp(\alpha_l)}{\sum_{j=1}^{L} \exp(\alpha_j)}

    The aggregated sequence (h1,…,hT)(\mathbf{h}_1, \dots, \mathbf{h}_T) is fed directly to a lightweight, task-specific downstream model. Only the layer weights {αl}l=1L\{\alpha_l\}_{l=1}^L and the parameters of the downstream model head are updated during training via backpropagation.

  2. Knowl 2 — SUPERB-SG Downstream Task Suite and Model Configurations

    experimental setup

    SUPERB-SG extends the standard SUPERB benchmark with five core tasks assessing semantic and generative capabilities:

    1. Speech Translation (ST): CoVoST 2 English-to-German (extEn→De ext{En} \to \text{De}) dataset (425.8 h train, 25.9 h validation, 24.5 h test). Downstream model: 3-layer Transformer encoder and 3-layer Transformer decoder with hidden dimension 512, preceded by a convolutional sub-sampler. Evaluated by case-sensitive de-tokenized BLEU using sacreBLEU with beam size 20 and label smoothing 0.1.
    2. Out-of-Domain ASR (OOD-ASR): Evaluates cross-lingual transfer on Common Voice 7.0 subsets (Mexican Spanish es: 21.5 h train, Mandarin zh: 31.2 h train, Arabic ar: 30.7 h train) and spontaneous speech on the Santa Barbara Corpus of Spoken American English (spon: 16.7 h total). Downstream model: 2-layer Bidirectional LSTM (BLSTM) with 1024 hidden units trained with Connectionist Temporal Classification (CTC) loss and greedy decoding without language model rescoring. Evaluated by Word Error Rate (WER), except Mandarin which uses Character Error Rate (CER), with the unweighted average reported.
    3. Voice Conversion (VC): Any-to-One (A2O) intra-lingual conversion from Voice Conversion Challenge 2020 (VCC2020) using 60 utterances (5 min) for training and 25 utterances (2 min) for testing. Downstream model: Autoregressive Tacotron2 generating target-speaker acoustic frames, synthesized via a pretrained HiFi-GAN neural vocoder. Evaluated by Mel-Cepstrum Distortion (MCD), ASR Word Error Rate (WER), and Automatic Speaker Verification (ASV) accept rate.
    4. Speech Separation (SS): 2-speaker mixtures at 16 kHz generated from LibriSpeech train-clean-100 / test-clean with WHAM! noise under the mix_clean condition (43.3 h train, 4.2 h evaluation). Downstream model: 3-layer BLSTM (896 units per direction) predicting frequency-domain Ideal Non-negative Phase Sensitive Masks (INPSM) for short-time Fourier transform (STFT) bins, optimized via Permutation Invariant Training (PIT) mean square error. Evaluated by scale-invariant signal-to-distortion ratio improvement (SI-SDRi).
    5. Speech Enhancement (SE): Voicebank-DEMAND corpus (8.8 h train, 0.6 h validation, 0.6 h test). Downstream model: 3-layer BLSTM predicting spectral masks optimized via MSE against clean INPSM targets. Evaluated by Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI).
  3. Knowl 3 — Benchmarking Self-Supervised Speech Representations on SUPERB-SG Tasks

    data/table

    The downstream performance of 15 self-supervised upstream representations and an 80-dimensional Log Mel-Filterbank (FBANK) baseline on the SUPERB-SG benchmark tasks is summarized below:

    Upstream ST OOD-ASR VC SS SE
    BLEU ↑\uparrow WER ↓\downarrow MCD ↓\downarrow WER ↓\downarrow ASV ↑\uparrow SI-SDRi ↑\uparrow PESQ ↑\uparrow STOI ↑\uparrow
    FBANK 2.32 63.58 8.47 38.3 77.25 9.23 2.55 93.6
    PASE+ 3.16 61.56 8.66 30.6 63.20 9.87 2.56 93.9
    APC 5.95 63.12 8.05 27.2 87.25 8.92 2.56 93.4
    VQ-APC 4.23 63.56 7.84 22.4 94.25 8.44 2.56 93.4
    NPC 4.32 61.66 7.86 30.4 94.75 8.04 2.52 93.1
    Mockingjay 4.45 65.27 8.29 35.1 79.75 9.29 2.53 93.4
    TERA 5.66 58.49 8.21 25.1 83.75 10.19 2.54 93.6
    DeCoAR 2.0 9.94 53.62 7.83 17.1 90.75 8.54 2.47 93.2
    Modified CPC 4.82 62.54 8.41 26.2 71.00 10.40 2.57 93.7
    wav2vec 6.61 55.86 7.45 10.1 98.25 9.30 2.53 93.8
    vq-wav2vec 5.66 60.66 7.08 13.4 100.00 8.16 2.48 93.6
    wav2vec 2.0 Base 14.81 46.95 7.50 10.5 98.00 9.77 2.55 93.9
    wav2vec 2.0 Large 12.48 44.69 7.63 15.8 97.25 10.02 2.52 94.0
    HuBERT Base 15.53 46.69 7.47 8.0 98.50 9.36 2.58 93.9
    HuBERT Large 20.01 44.08 7.22 9.0 99.25 10.45 2.64 94.2

    HuBERT Large yields the highest performance across semantic and text-generation tasks (20.01 BLEU on ST, 44.08% average WER on OOD-ASR) and matches the highest score on speech separation (10.45 dB SI-SDRi). In speech generation tasks (SS, SE, VC), baseline Log Mel-Filterbank representations perform competitively with or better than several SSL models (e.g., FBANK matches or exceeds APC, VQ-APC, NPC, and DeCoAR 2.0 on SI-SDRi and PESQ), demonstrating that high-level SSL abstractions provide less advantage for tasks primarily dependent on low-level acoustic signal reconstruction.

  4. Knowl 4 — Sub-Task Breakdown of Out-of-Domain and Multilingual Speech Recognition

    data/table

    The detailed breakdown of Out-of-Domain ASR performance across Mexican Spanish (es), Mandarin (zh), Arabic (ar), and spontaneous English conversation (spon) shows the contrast between English-only pre-trained models and the multilingual wav2vec 2.0 XLSR model:

    Upstream es (WER ↓\downarrow) zh (CER ↓\downarrow) ar (WER ↓\downarrow) spon (WER ↓\downarrow) AVG (↓\downarrow)
    FBANK 54.03 35.44 72.07 92.78 63.58
    PASE+ 52.11 35.52 70.47 88.15 61.56
    APC 55.23 36.38 70.79 90.07 63.12
    VQ-APC 55.32 37.06 71.56 90.29 63.56
    NPC 51.07 35.85 69.87 89.86 61.66
    Mockingjay 58.11 38.13 73.57 91.27 65.27
    TERA 48.67 32.21 66.18 86.89 58.49
    Modified CPC 54.37 36.22 68.94 90.61 62.54
    DeCoAR 2.0 43.18 28.77 61.00 81.53 53.62
    wav2vec 46.16 31.69 60.85 84.72 55.86
    vq-wav2vec 52.02 36.55 66.19 87.89 60.66
    wav2vec 2.0 Base 37.85 26.44 55.95 67.55 46.95
    wav2vec 2.0 Large 35.75 25.07 54.29 63.64 44.69
    HuBERT Base 37.15 26.23 54.94 68.41 46.69
    HuBERT Large 30.90 23.73 50.60 71.09 44.08
    wav2vec 2.0 XLSR 26.90 22.97 49.63 63.05 40.64

    While all models trained exclusively on English exhibit severe degradation under cross-lingual and spontaneous domain shifts, wav2vec 2.0 XLSR (pre-trained on 56k hours of speech across 53 languages) significantly outperforms all models across all non-English evaluation sets, reducing the overall average error rate to 40.64%.

  5. Knowl 5 — Spearman Correlation and Functional Clustering of Speech Downstream Tasks

    empirical result

    Pairwise evaluation of upstream representations using Spearman's rank correlation coefficient (ρ\rho) over 11 downstream speech tasks (14 evaluation metrics total, including Phoneme Recognition [PR], standard ASR, Speaker Identification [SID], Automatic Speaker Verification [ASV], Intent Classification [IC], Emotion Recognition [ER], ST, OOD-ASR, VC, SS, and SE) reveals functional task groupings:

    1. High Mutual Transferability in Semantic and Content Tasks: Content and semantic metrics show pairwise correlations with ρ>0.83\rho > 0.83. Specifically, Speech Translation correlates strongly with ASR (ρ=0.92\rho = 0.92) and IC (ρ=0.92\rho = 0.92). OOD-ASR correlates strongly with standard ASR (ρ=0.92\rho = 0.92) and PR (ρ=0.86\rho = 0.86).
    2. K-Means Clustering of Tasks: Applying K-means clustering to the pairwise correlation matrix partitions the metrics into six distinct clusters:
      • Cluster A (Content & Linguistic Information): ST, OOD-ASR, PR, VC (WER), ASR, IC.
      • Cluster B (Speaker Identity & Prosody): SID, ASV, ER.
      • Cluster C (Voice Conversion Acoustic Fidelity & Speaker Transfer): VC (MCD), VC (ASV).
      • Cluster D (Speech Separation): SS (SI-SDRi).
      • Cluster E (Speech Enhancement Quality): SE (PESQ).
      • Cluster F (Speech Enhancement Intelligibility): SE (STOI).
    3. Generative Low-Level Independence: Generative tasks that rely on STFT spectral mask estimation (SS and SE) correlate weakly with content, speaker, and semantic tasks (e.g., SS vs. ASV yields ρ=0.04\rho = 0.04, SE vs. ST yields ρ=0.10\rho = 0.10). Excluding SS and SE, all task correlation coefficients exceed 0.580.58.
  6. Knowl 6 — Invariance of Upstream SSL Model Rankings Across Downstream Head Capacities

    empirical result

    Varying the capacity of downstream prediction heads demonstrates that the relative performance ranking of upstream SSL representations is robust to downstream architectural modifications:

    Downstream Architecture Upstream ST (BLEU ↑\uparrow) OOD-ASR (WER ↓\downarrow) SS (SI-SDRi ↑\uparrow)
    Default FBANK 2.32 63.58 9.23
    (ST: 28.8M params, TERA 5.66 58.49 10.19
    OOD-ASR: 53.4M params, Modified CPC 4.82 62.54 10.40
    SS: 51.4M params) wav2vec 2.0 Base 14.81 46.95 9.77
    HuBERT Base 15.53 46.69 9.36
    Small FBANK 0.58 70.86 8.19
    (ST: 10.9M params, TERA 1.84 64.80 9.20
    OOD-ASR: 24.1M params, Modified CPC 1.44 67.83 9.56
    SS: 24.4M params) wav2vec 2.0 Base 8.55 50.75 8.83
    HuBERT Base 9.24 50.32 8.73
    Large FBANK 3.02 60.49 9.77
    (ST: 69.8M params, TERA 6.64 57.95 10.87
    OOD-ASR: 112.2M params, Modified CPC 4.56 59.73 10.61
    SS: 114.5M params) wav2vec 2.0 Base 16.81 45.61 9.86
    HuBERT Base 17.59 45.78 9.83

    Across scale adjustments ranging from roughly 0.4×0.4\times (small) to 2.4×2.4\times (large) the default parameter counts, downstream absolute performance scales with head capacity, but relative upstream rankings remain virtually invariant across tasks.

  7. Knowl 7 — Robustness and Failure Thresholds of SSL Transfer Under Reduced Downstream Training Data

    empirical result

    Evaluating downstream performance under sub-sampled training sets (100%, 10%, 5%, and 1% of training hours with fixed validation and test partitions) reveals the data regimes where SSL representations generalize:

    Data Partition Upstream ST (BLEU ↑\uparrow) OOD-ASR (WER ↓\downarrow) SS (SI-SDRi ↑\uparrow)
    100% FBANK 2.32 63.58 9.23
    TERA 5.66 58.49 10.19
    Modified CPC 4.82 62.54 10.40
    wav2vec 2.0 Base 14.81 46.95 9.77
    HuBERT Base 15.53 46.69 9.36
    10% FBANK 0.46 85.39 5.65
    TERA 0.88 80.32 6.72
    Modified CPC 1.30 85.32 6.59
    wav2vec 2.0 Base 5.04 63.85 6.45
    HuBERT Base 5.57 63.43 6.13
    5% FBANK 0.27 89.70 4.52
    TERA 0.44 86.95 5.59
    Modified CPC 0.37 87.97 4.95
    wav2vec 2.0 Base 2.91 69.88 5.36
    HuBERT Base 3.35 69.33 5.03
    1% FBANK 0.03 99.53 2.29
    TERA 0.04 98.31 3.24
    Modified CPC 0.03 98.37 2.87
    wav2vec 2.0 Base 0.33 92.46 3.34
    HuBERT Base 0.38 92.17 3.01

    Relative model rankings remain stable down to 5% of training data. At 1% supervision (4.26 h for ST, 0.12–0.31 h for OOD-ASR sub-tasks, 0.43 h for SS), downstream performance collapses across all SSL representations (e.g., ST BLEU ≤0.38\le 0.38, OOD-ASR WER ≥92.17%\ge 92.17\%). This demonstrates that while SSL models yield high-level generalizable features, learning task-specific alignments (translingual mappings in ST, character/token mappings in ASR, and STFT mask prediction in SS) requires a non-trivial minimum threshold of supervised data.

Coverage note — None was omitted; all contributed benchmark tasks, upstream evaluations, correlation/clustering analyses, and downstream robustness ablations are covered.

References

  1. 1.David Álvarez et al. 2019. Problem-agnostic speech embeddings for multi-speaker text-to-speech with samplernn. In Proc. 10th ISCA Speech Synthesis Workshop, pages 35–39.
  2. 2.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4218–4222.
  3. 3.Alexei Baevski, Steffen Schneider, and Michael Auli. 2020a. vq-wav2vec: Self-supervised learning of discrete speech representations. In ICLR.
  4. 4.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020b. wav2vec 2.0: A framework for self-supervised learning of speech representations. In NeurIPS.
  5. 5.Heng-Jui Chang, Shu-wen Yang, and Hung-yi Lee. 2021. DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit BERT. arXiv preprint arXiv:2110.01900.
  6. 6.Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019. An Unsupervised Autoregressive Model for Speech Representation Learning. In Interspeech, pages 146–150.
  7. 7.Yu-An Chung, Hao Tang, and James Glass. 2020. Vector-quantized autoregressive predictive coding. In Interspeech, pages 3760–3764.
  8. 8.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale.
  9. 9.Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent. 2020. Librimix: An open-source dataset for generalizable speech separation. arXiv preprint arXiv:2005.11262.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186.
  11. 11.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:1905.03197.
  12. 12.John W Du Bois, Wallace L Chafe, Charles Meyer, Sandra A Thompson, and Nii Martey. 2000 – 2005. Santa Barbara corpus of spoken American English. CD-ROM. Philadelphia: Linguistic Data Consortium.
  13. 13.Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux. 2015. Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 708–712. IEEE.
  14. 14.Solene Evain, Ha Nguyen, Hang Le, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Esteve, Benjamin Lecouteux, Francois Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, and Laurent Besacier. 2021. Lebenchmark: A reproducible framework for assessing self-supervised representation learning from speech.
  15. 15.Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu. 2020. Exploring wav2vec 2.0 on speaker verification and language identification. arXiv preprint arXiv:2012.06185.
  16. 16.Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369–376.
  17. 17.Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. 2017. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409.
  18. 18.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.
  19. 19.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. arXiv preprint arXiv:2106.07447.
  20. 20.W.-C. Huang, Y.-C. Wu, T. Hayashi, and T. Toda. 2021a. Any-to-One Sequence-to-Sequence Voice Conversion using Self-Supervised Discrete Speech Representations. In Proc. ICASSP, pages 5944–5948.
  21. 21.Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Hung-Yi Lee, Shinji Watanabe, and Tomoki Toda. 2021b. S3prl-vc: Open-source voice conversion framework with self-supervised speech representations. arXiv preprint arXiv:2110.06280.
  22. 22.Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.
  23. 23.Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen. 2017. Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 25(10):1901–1913.
  24. 24.J. Kong, J. Kim, and J. Bae. 2020. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. In Proc. NeurIPS, volume 33, pages 17022–17033.
  25. 25.Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass. 2021. Semi-supervised spoken language understanding via self-supervised speech and language model pretraining. In ICASSP.
  26. 26.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations.
  27. 27.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. ACL.
  28. 28.Yist Y Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, and Lin-shan Lee. 2020. Fragmentvc: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention. arXiv preprint arXiv:2010.14150.
  29. 29.Shaoshi Ling and Yuzong Liu. 2020. DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization. arXiv preprint arXiv:2012.06659.
  30. 30.Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff. 2020. Deep contextualized acoustic representations for semi-supervised speech recognition. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6429–6433. IEEE.
  31. 31.Alexander H Liu, Yu-An Chung, and James Glass. 2020a. Non-autoregressive predictive coding for learning speech representations from local dependencies. arXiv preprint arXiv:2011.00406.
  32. 32.Andy T Liu, Shang-Wen Li, and Hung-yi Lee. 2020b. Tera: Self-supervised learning of transformer encoder representation for speech. arXiv preprint arXiv:2007.06028.
  33. 33.Andy T. Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee. 2020c. Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders. ICASSP.
  34. 34.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  35. 35.Manon Macary, Marie Tahon, Yannick Estève, and Anthony Rousseau. 2021. On the use of self-supervised pre-trained acoustic and linguistic features for continuous speech emotion recognition. In 2021 IEEE Spoken Language Technology Workshop (SLT), pages 373–380. IEEE.
  36. 36.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. 2018. Exploring the limits of weakly supervised pretraining. In Proceedings of the European Conference on Computer Vision (ECCV).
  37. 37.Tu Anh Nguyen et al. 2020. The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling. In NeurIPS Workshop on Self-Supervised Learning for Speech and Audio Processing.
  38. 38.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210. IEEE.
  39. 39.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL, pages 2227–2237.
  40. 40.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels. Association for Computational Linguistics.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv.
  43. 43.Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio. 2020. Multi-task self-supervised learning for robust speech recognition. In ICASSP, pages 6989–6993.
  44. 44.Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux. 2020. Unsupervised pretraining transfers well across languages. In ICASSP, pages 7414–7418.
  45. 45.Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019. wav2vec: Unsupervised pre-training for speech recognition. In Interspeech.
  46. 46.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  47. 47.J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu. 2018. Natural TTS Synthesis by Conditioning WaveNet on MEL Spectrogram Predictions. In Proc. ICASSP, pages 4779–4783.
  48. 48.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. CoRR, abs/1807.03748.
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  50. 50.Christophe Veaux, Junichi Yamagishi, and Simon King. 2013. The voice bank corpus: Design, collection and data analysis of a large regional accent speech database. In 2013 international conference oriental COCOSDA held jointly with 2013 conference on Asian spoken language research and evaluation (O-COCOSDA/CASLRE), pages 1–4. IEEE.
  51. 51.Changhan Wang, Anne Wu, and Juan Pino. 2020. CoVoST 2: A massively multilingual speech-to-text translation corpus.
  52. 52.DeLiang Wang and Jitong Chen. 2018. Supervised speech separation based on deep learning: An overview. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 26(10):1702–1726.
  53. 53.Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux. 2019. Wham!: Extending speech separation to noisy environments. arXiv preprint arXiv:1907.01160.
  54. 54.Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee. 2021. SUPERB: Speech processing universal performance benchmark. Interspeech.
  55. 55.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237.
  56. 56.Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen. 2017. Permutation invariant training of deep models for speaker-independent multi-talker speech separation. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 241–245. IEEE.
  57. 57.Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z. Ling, and T. Toda. 2020. Voice Conversion Challenge 2020 - Intra-lingual semi-parallel and cross-lingual voice conversion -. In Proc. Joint Workshop for the BC and VCC 2020, pages 80–98.

Citation

MLA
Tsai, H.-S., et al. “SUPERB-SG: Enhanced Speech Processing Universal PERformance Benchmark for Semantic and Generative Capabilities”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 8479–92, https://doi.org/10.18653/v1/2022.acl-long.580.
APA
Tsai, H.-S., Chang, H.-J., Huang, W.-C., Huang, Z., Lakhotia, K., Yang, S.-. wen ., Dong, S., Liu, A. T., Lai, C.-I., Shi, J., Chang, X., Hall, P., Chen, H.-J., Li, S.-W., Watanabe, S., Mohamed, A., & Lee, H.-. yi . (2022). SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8479–8492. https://doi.org/10.18653/v1/2022.acl-long.580
Chicago
Tsai, H.-S., H.-J. Chang, W.-C. Huang, et al. 2022. “SUPERB-SG: Enhanced Speech Processing Universal PERformance Benchmark for Semantic and Generative Capabilities”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8479–92. https://doi.org/10.18653/v1/2022.acl-long.580.
Harvard
Tsai, H.-S. et al. (2022) “SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8479–8492. Available at: https://doi.org/10.18653/v1/2022.acl-long.580.
Vancouver
1. Tsai H-S, Chang H-J, Huang W-C, et al (2022) SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8479–8492

BibTeX

@inproceedings{tsai-etal-2022-superb,
    title = "{SUPERB}-{SG}: Enhanced Speech processing Universal {PER}formance Benchmark for Semantic and Generative Capabilities",
    author = "Tsai, Hsiang-Sheng  and
      Chang, Heng-Jui  and
      Huang, Wen-Chin  and
      Huang, Zili  and
      Lakhotia, Kushal  and
      Yang, Shu-wen  and
      Dong, Shuyan  and
      Liu, Andy  and
      Lai, Cheng-I  and
      Shi, Jiatong  and
      Chang, Xuankai  and
      Hall, Phil  and
      Chen, Hsuan-Jui  and
      Li, Shang-Wen  and
      Watanabe, Shinji  and
      Mohamed, Abdelrahman  and
      Lee, Hung-yi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.580/",
    doi = "10.18653/v1/2022.acl-long.580",
    pages = "8479--8492"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/