Unified Speech-Text Pre-training for Speech Translation and Recognition

Yun TangHongyu GongNing DongChanghan WangWei-Ning HsuJiatao GuAlexei BaevskiXian LiAbdelrahman MohamedMichael Auli

article2022ACL107 citations

Proposes a joint speech-text pre-training framework that integrates linguistic knowledge from text into speech models across four multi-task objectives, providing tailored encoder sharing strategies to resolve subtask interference and substantially boosting performance on speech translation and recognition.

Listen

Building accurate automated speech recognition and speech translation systems typically requires substantial labeled audio data, which is scarce and expensive to produce for most languages. While large amounts of unlabeled audio and pure text data are widely available, existing machine learning methods struggle to jointly exploit both modalities in a single framework without experiencing cross-task interference or requiring complex multi-stage pipelines.

The article develops and evaluates Speech and Text Joint Pre-Training (STPT), a multi-task learning method designed to unify speech and text representations within an attention-based encoder-decoder model for speech translation and speech recognition.

The authors conducted large-scale computational experiments utilizing up to 60,000 hours of unlabeled English speech, extensive parallel and monolingual text datasets, and standard benchmark datasets including LibriSpeech and MuST-C. The approach integrates four distinct learning objectives: text-to-text modeling, self-supervised speech learning via a masked distribution alignment loss, supervised speech-to-phoneme classification, and sequence-to-sequence speech-to-text generation. To address negative task interactions identified through gradient similarity analysis, the researchers designed two tailored model structures: a fully shared encoder architecture for speech recognition and a partially shared encoder architecture for speech translation.

The evaluation produced four key findings. First, STPT established new state-of-the-art results in speech translation on the MuST-C benchmark, improving translation quality by 1.7 to 2.3 BLEU points over the strongest prior systems. Second, in speech recognition, the system achieved a competitive word error rate of 3.2% to 3.3% on LibriSpeech, matching specialized models without requiring an external language model during decoding. Third, model architecture design proved essential: while speech recognition benefited from full parameter sharing across all tasks, speech translation suffered from severe gradient conflicts that required isolating encoder subtasks via a partially shared structure. Fourth, the proposed soft-label distribution alignment loss for speech self-supervision consistently outperformed traditional contrastive objectives, yielding a 0.6 reduction in word error rate and up to 1.4 BLEU point gains in translation.

These findings indicate that directly integrating linguistic knowledge from text corpora into speech models during pre-training eliminates the operational overhead of maintaining separate language models during deployment. However, system architects must tailor parameter-sharing strategies specifically to downstream application goals rather than applying a universal multi-task structure.

Organizations developing speech processing pipelines should adopt the partially shared encoder design when training translation systems and the fully shared architecture for transcription tasks. Before deploying these models, practitioners should evaluate the underlying text corpora to prevent propagating offensive or biased training content into production systems. Future engineering efforts should explore expanding the framework to multilingual speech settings and investigating methods to stabilize performance when labeled training data is extremely limited.

The findings are supported by consistent empirical results across multiple benchmark splits. Readers should note, however, that performance degrades noticeably when supervised training audio is restricted below 100 hours or when acoustic conditions between pre-training and real-world inference diverge significantly.

Cover for Unified Speech-Text Pre-training for Speech Translation and Recognition

Abstract

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for cross modality learning. A self-supervised speech subtask leverages unlabelled speech data, and a (self-)supervised text to text subtask makes use of abundant text training data. Two auxiliary supervised speech tasks are included to unify speech and text modeling space. Our contribution lies in integrating linguistic information from the text corpus into the speech pre-training. Detailed analysis reveals learning interference among subtasks. Two pre-training configurations for speech translation and recognition, respectively, are presented to alleviate subtask interference. Our experiments show the proposed method can effectively fuse speech and text information into one model. It achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MUST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the LIBRISPEECH speech recognition task.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 (Self-)supervised text to text subtask
  • 3.2 Self-supervised speech subtask
  • 3.3 Supervised speech to phoneme classification
  • 3.4 Supervised AED based speech to text subtask
  • 4 Experimental setting
  • 4.1 Model configuration
  • 5 Experimental results
  • 5.1 Main results
  • 5.2 Impact of model structure
  • 5.3 Impact of training data
  • 5.4 Masked KL divergence loss v.s. contrastive loss
  • 5.5 Ablation study
  • 6 Conclusion
  • 7 Acknowledgments
  • 8 Broader Impact
  • References
  • A Pre-training data setting
  • B Optimization setting
  • C Gradient similarity of the speech encoder

Knowls

  1. Knowl 1 — Speech and Text Joint Pre-Training Framework and Subtasks

    model/method

    Speech and Text Joint Pre-Training (STPT) is an attention-based encoder-decoder (AED) framework designed to jointly pre-train speech and text representations across four supervised and self-supervised subtasks:

    1. (Self-)supervised Text-to-Text Subtask (T2T): For automatic speech recognition (ASR), this is a denoising autoencoder task (BART-style) in which masked or corrupted token spans in target text are reconstructed. For speech translation (ST), it is a text-to-text neural machine translation (NMT) task mapping source text tokens to target text tokens. Input text is converted into a phoneme sequence representation using a grapheme-to-phoneme dictionary to bridge the representation gap between speech and text modalities.
    2. Self-supervised Speech Learning Subtask (SSL): Operates on unlabeled acoustic audio. Raw speech is processed by a convolutional feature extractor followed by a context encoder. It is optimized using a masked predictive distribution matching loss against soft phoneme distributions computed from T2T phoneme embeddings.
    3. Supervised Speech-to-Phoneme Classification Subtask (S2P): Uses transcribed speech aligned to frame-level phoneme targets (e.g., generated via HMM-GMM forced alignment). The context encoder predictions are supervised via frame-level cross-entropy to anchor acoustic representations into the shared phoneme embedding space.
    4. Supervised AED-based Speech-to-Text Subtask (S2T): Matches the downstream sequence-to-sequence task (ASR transcription or ST translation) using labeled audio-text pairs, trained with autoregressive sequence cross-entropy.

    To stabilize multi-task training across modalities with disparate feature magnitudes, Layer Normalization is applied directly to shared encoder inputs coming from either speech encoder outputs or text phoneme embeddings.

  2. Knowl 2 — Masked KL Divergence Loss for Self-Supervised Speech Learning

    equation

    In the Self-Supervised Speech Learning (SSL) subtask of STPT, speech representations are learned by minimizing the Kullback-Leibler (KL) divergence between phoneme probability distributions produced from uncorrupted and corrupted speech feature frames.

    Let raw speech S=(s1,s2,…,sT)S = (s_1, s_2, \dots, s_T) be downsampled by a convolutional feature extractor into temporal representations Z=(z1,z2,…,zT′)Z = (z_1, z_2, \dots, z_{T'}) where T>T′T > T'.

    In the first pass, passing ZZ through the context encoder yields uncorrupted representations O=(o1,o2,…,oT′)O = (o_1, o_2, \dots, o_{T'}). The predicted soft distribution over the phoneme vocabulary of size II with embeddings E=(e1,e2,…,eI)E = (e_1, e_2, \dots, e_I) is defined as: p(oj∣ei)=exp⁡(oj⊤ei)∑i′=1Iexp⁡(oj⊤ei′)p(o_j \mid e_i) = \frac{\exp(o_j^\top e_i)}{\sum_{i'=1}^I \exp(o_j^\top e_{i'})}

    In the second pass, a subset of feature spans in ZZ is randomly masked to form corrupted features Z^\hat{Z}, which produces corrupted context encoder representations O^=(o^1,o^2,…,o^T′)\hat{O} = (\hat{o}_1, \hat{o}_2, \dots, \hat{o}_{T'}). The SSL masked KL divergence loss over the masked time steps o^j∈O^\hat{o}_j \in \hat{\mathcal{O}} is: LSSL=−∑o^j∈O^∑i=1Ip(oj∣ei)log⁡p(o^j∣ei)p(oj∣ei)\mathcal{L}_{\text{SSL}} = - \sum_{\hat{o}_j \in \hat{\mathcal{O}}} \sum_{i=1}^I p(o_j \mid e_i) \log \frac{p(\hat{o}_j \mid e_i)}{p(o_j \mid e_i)}

    This soft-target distribution matching avoids hard quantization errors inherent in predicting discretized cluster indices.

  3. Knowl 3 — Unified Multi-Task Pre-Training Objective

    equation

    The overall joint pre-training loss L\mathcal{L} for STPT is formulated as a weighted linear combination of four subtask loss functions: L=LT2T+αLSSL+βLS2P+γLS2T\mathcal{L} = \mathcal{L}_{\text{T2T}} + \alpha \mathcal{L}_{\text{SSL}} + \beta \mathcal{L}_{\text{S2P}} + \gamma \mathcal{L}_{\text{S2T}}

    where α\alpha, β\beta, and γ\gamma are scalar task weights, and the individual component objectives are defined as:

    • LT2T=−∑i=1Nlog⁡p(yi∣y1:i−1,X)\mathcal{L}_{\text{T2T}} = - \sum_{i=1}^N \log p(y_i \mid y_{1:i-1}, X), the sequence cross-entropy loss for the text-to-text task on target tokens Y=(y1,…,yN)Y = (y_1, \dots, y_N) conditioned on source/corrupted text XX.
    • LSSL=−∑o^j∈O^∑i=1Ip(oj∣ei)log⁡p(o^j∣ei)p(oj∣ei)\mathcal{L}_{\text{SSL}} = - \sum_{\hat{o}_j \in \hat{\mathcal{O}}} \sum_{i=1}^I p(o_j \mid e_i) \log \frac{p(\hat{o}_j \mid e_i)}{p(o_j \mid e_i)}, the masked KL divergence loss matching soft phoneme distributions between unmasked outputs ojo_j and masked outputs o^j\hat{o}_j.
    • LS2P=−∑oj∈Olog⁡p(oj∣ea(j))\mathcal{L}_{\text{S2P}} = - \sum_{o_j \in \mathcal{O}} \log p(o_j \mid e_{a(j)}), the frame-level cross-entropy loss over context encoder frames ojo_j predicting forced-aligned phoneme label indices a(j)a(j).
    • LS2T=−∑t=1Nlog⁡p(yt∣y1:t−1,O)\mathcal{L}_{\text{S2T}} = - \sum_{t=1}^N \log p(y_t \mid y_{1:t-1}, O), the autoregressive sequence cross-entropy loss for speech-to-text decoding given speech context encoder representations OO.
  4. Knowl 4 — Fully Shared Encoder versus Partially Shared Encoder Configurations

    model/method

    To manage cross-modal information sharing and avoid optimization conflicts, STPT introduces two architectural configurations based on the downstream task:

    1. Fully Shared Encoder (FSE) for ASR: The convolutional speech feature extractor feeds into a speech encoder, which is directly connected to a shared encoder. All four subtasks (T2T, SSL, S2P, and S2T) pass through the upper shared encoder layers before feeding into downstream linear heads or the shared decoder. This maximizes cross-modal representation sharing between speech recognition and text denoising.

    2. Partially Shared Encoder (PSE) for Speech Translation (ST): Encoder-only subtasks (SSL and S2P) branch out directly from the speech encoder and bypass the shared encoder. Only sequence-to-sequence AED subtasks (text machine translation T2T and speech translation S2T) pass into and share the shared encoder before reaching the shared decoder. This separation avoids destructive gradient interference between acoustic frame-level modeling (SSL/S2P) and cross-lingual translation generation (T2T/S2T).

  5. Knowl 5 — Task Gradient Interference in Multi-Task Speech-Text Pre-Training

    empirical result

    Pairwise cosine similarities computed between accumulated parameter gradients across subtasks in the shared encoder reveal distinct task-sharing characteristics:

    • ASR Multi-Task Learning: Gradient cosine similarities across all pairs of subtasks (SSL, S2P, S2T, and T2T) in the shared encoder remain close to zero across all layers, indicating negligible gradient conflict and mutually complementary representations.
    • ST Multi-Task Learning: When using a Fully Shared Encoder for speech translation, pronounced gradient interference occurs in intermediate shared encoder layers (layers 3 and 5), exhibiting large positive and negative gradient cosine similarities between the text translation T2T task and the acoustic SSL/S2P subtasks.
    • Speech Encoder: Pairwise gradient cosine similarities among acoustic subtasks (SSL, S2P, S2T) in the dedicated speech encoder remain below 0.2 in magnitude for both ASR and ST, demonstrating that low-level acoustic feature extraction does not suffer from inter-task conflict.
  6. Knowl 6 — Three-Stage Optimization Pipeline for STPT

    algorithm
    Input: Unlabeled speech audio D_unlabeled, labeled speech pairs D_speech, text corpus D_text
    Output: Trained encoder-decoder model for downstream ASR or ST
    Stage 1: Pre-training T2T Subtask
    Initialize model parameters randomly
    Convert text inputs in D_text to phoneme sequences using a 134-token phoneme vocabulary
    while T2T task not converged do
        Sample mini-batch from D_text
        Update T2T encoder and decoder parameters using Adam (learning rate 0.01)
    end while
    Stage 2: Joint Pre-training with Four Subtasks
    Set subtask mini-batch ratio to 1.0 (T2T) : 7.0 (SSL) : 0.5 (S2P) : 0.5 (S2T)
    Set speech mask span length to 10 with 7% SSL mask start probability and 3% supervised mask start probability
    while step <= max_pretrain_updates do
        Sample mini-batches according to task ratios
        Compute combined loss L = L_T2T + alpha * L_SSL + beta * L_S2P + gamma * L_S2T
        Update full model using Adam (learning rate 0.001)
    end while
    Stage 3: Fine-tuning on Downstream Task
    Drop encoder-only subtasks (SSL and S2P)
    while step <= max_finetune_updates do
        Jointly train model on downstream S2T task and T2T task using Adam (learning rate 0.0003)
    end while
    return Average model parameters from checkpoints of the last 10 epochs
  7. Knowl 7 — Speech-to-Text Translation Benchmark Performance on MuST-C

    data/table

    Evaluated on the MuST-C tst-COMMONtst\text{-}COMMON benchmark in case-sensitive detokenized SacreBLEU, STPT achieves superior translation quality on English-to-Spanish (EN-ES) and English-to-French (EN-FR):

    Model EN-ES (BLEU) EN-FR (BLEU)
    Inaguma et al. (2020) 28.0 32.7
    Tang et al. (2021a) 31.0 37.4
    Zheng et al. (2021) 30.8 -
    Ye et al. (2021) 30.8 38.0
    STPT 33.1 39.7

    STPT outperforms prior state-of-the-art multi-task models by +2.3 BLEU on EN-ES and +1.7 BLEU on EN-FR over the strongest baseline (Ye et al., 2021), which was initialized with wav2vec 2.0 and mBART modules.

  8. Knowl 8 — Automatic Speech Recognition Performance on LibriSpeech

    data/table

    On the LibriSpeech benchmark, Word Error Rate (WER) is evaluated across dev and test splits. Results within parentheses denote decoding with an external language model (LM):

    Model Unlabeled Data Dev Clean Dev Other Test Clean Test Other Average
    wav2vec 2.0 (CTC) LS-960 3.2 (1.8) 8.9 (4.7) 3.4 (2.1) 8.5 (4.8) 6.0 (3.4)
    LAS (AED) - - - 2.8 (2.5) 6.8 (5.8) -
    Transformer (AED) - 2.8 7.0 3.1 7.2 5.0
    STPT (AED) LS-960 2.1 (1.9) 5.4 (5.2) 2.3 (2.2) 5.6 (5.3) 3.8 (3.6)
    STPT (AED) LV-60k 2.0 (2.1) 4.4 (4.2) 2.1 (2.1) 4.6 (4.5) 3.3 (3.2)

    STPT outperforms all prior AED systems, achieving a 1.2 absolute average WER reduction over the multi-task Transformer baseline. Without an external LM, STPT achieves 3.8 average WER on LS-960 compared to 6.0 for wav2vec 2.0. Decoding with an external LM yields minimal improvement (0.1–0.2 WER) for STPT, indicating that linguistic text knowledge is already fused into the AED model parameters during pre-training.

  9. Knowl 9 — Masked KL Divergence Loss versus Contrastive Loss for SSL

    data/table

    Evaluating the SSL subtask optimization criterion on LibriSpeech (WER) and MuST-C (BLEU) confirms the empirical advantage of the masked KL divergence loss over standard contrastive loss (using 100 distractors):

    Loss Criterion LibriSpeech WER (↓\downarrow) MuST-C BLEU (↑\uparrow)
    dev clean dev other EN-ES EN-FR
    Contrastive Loss 2.6 5.0 31.7 39.0
    Masked KL Divergence Loss 2.0 4.4 33.1 39.7

    The masked KL divergence loss achieves 0.6 lower WER across both clean and other dev sets in LibriSpeech, and improves MuST-C translation scores by +1.4 BLEU on EN-ES and +0.7 BLEU on EN-FR.

  10. Knowl 10 — Ablation Analysis of STPT Pre-Training Components

    data/table

    Ablation experiments evaluate the contribution of individual stages and subtasks to overall STPT performance:

    Configuration LibriSpeech WER (↓\downarrow) MuST-C BLEU (↑\uparrow)
    dev clean dev other EN-ES EN-FR
    Full STPT 2.0 4.4 33.1 39.7
    - T2T Pre-training 2.4 5.0 31.9 39.2
    - AED Subtask (S2T) 2.9 5.6 31.3 38.0
    - Joint Pre-training 2.8 6.4 30.6 35.4

    Key takeaways from the ablation:

    1. Omitting initial T2T pre-training degrades phoneme embedding initialization, causing an average 0.5 WER increase on LibriSpeech and a 1.2 BLEU drop on EN-ES.
    2. Omitting the supervised AED subtask (S2T) during pre-training results in an average 1.1 WER increase and a 1.8 BLEU drop across translation directions.
    3. Removing the S2P subtask entirely leads to complete training divergence: the SSL loss collapses to zero as predictions map to only one or two phonemes.
    4. Skipping joint pre-training entirely (training only T2T and S2T) degrades average WER by 1.4 points and ST BLEU by 3.4 points.
  11. Knowl 11 — Impact of Supervised Speech Volume and Acoustic Domain Mismatch

    empirical result

    Varying the volume and domain matching of supervised speech data in STPT reveals the following characteristics:

    • Supervised Data Quantity: When fine-tuning on the full 960 hours of LibriSpeech, reducing labeled speech in pre-training from 960h to 100h and 10h yields average test WERs of 3.3, 3.6, and 4.0 respectively. Pre-training with only 10h of labeled data still outperforms standard non-pretrained AED baselines when 960h is used for fine-tuning. However, when both pre-training and fine-tuning are restricted to 10 hours, average WER degrades significantly to 24.6 (22.0 test-clean, 28.8 test-other).
    • Acoustic Domain Mismatch: When pre-training and fine-tuning datasets are partitioned into 'clean' and 'other' subsets (approx. 100h each), matched conditions achieve optimal performance (PT Clean + FT Clean yields 3.0 / 6.7 WER; PT Other + FT Other yields 3.0 / 5.8 WER). Cross-condition mismatch between pre-training and fine-tuning causes only a minor WER increase of 0.1 to 0.2 points.
  12. Knowl 12 — Empirical Impact of Encoder Sharing Architecture on ASR and ST

    data/table

    Comparing Fully Shared Encoder (FSE) and Partially Shared Encoder (PSE) configurations across ASR (LibriSpeech 100h setup) and ST (MuST-C) confirms task-dependent architectural trade-offs:

    Configuration LibriSpeech WER (↓\downarrow) MuST-C BLEU (↑\uparrow)
    dev clean dev other EN-ES EN-FR
    Fully Shared Encoder (FSE) 3.2 6.8 31.4 38.3
    Partially Shared Encoder (PSE) 3.1 8.3 33.1 39.7

    For ASR, sharing all encoder layers (FSE) is superior, yielding an 8.3 →\rightarrow 6.8 WER reduction on dev other. For ST, partially sharing the encoder (PSE) avoids destructive cross-task gradient interference, boosting translation quality by +1.7 BLEU on EN-ES and +1.4 BLEU on EN-FR.

Coverage note — None omitted. All key contributions, including the multi-task framework, loss formulas, FSE/PSE architectures, gradient analysis, training algorithm, and benchmark results are fully captured.

References

  1. 1.Antonios Anastasopoulos and David Chiang. 2018. Tied multitask learning for neural speech translation. In NAACL-HLT.
  2. 2.Junyi Ao, Rui Wang, Long Zhou, Shujie Liu, Shuo Ren, Yu Wu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, and Furu Wei. 2021. Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.
  3. 3.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. ArXiv, abs/1607.06450.
  4. 4.Alexei Baevski, Steffen Schneider, and Michael Auli. 2020a. vq-wav2vec: Self-supervised learning of discrete speech representations. In ICLR.
  5. 5.Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020b. wav2vec 2.0: A framework for self-supervised learning of speech representations. In NeurIPS.
  6. 6.Ankur Bapna, Yu an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H. Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, and Yu Zhang. 2021. Slam: A unified encoder for speech and language modeling via speech-text joint pre-training. ArXiv, abs/2110.10329.
  7. 7.Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, and Pedro J. Moreno. 2021. Injecting text in self-supervised speech pretraining. In ASRU, pages 251–258.
  8. 8.Yung-Sung Chuang, Chi-Liang Liu, Hung yi Lee, and Lin-Shan Lee. 2020. Speechbert: An audio-and-text jointly learned language model for end-to-end spoken question answering. In INTERSPEECH.
  9. 9.Yu-An Chung and James Glass. 2020. Improved speech representations with multi-target autoregressive predictive coding. In ACL.
  10. 10.Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James R. Glass. 2018. Unsupervised cross-modal alignment of speech and text embedding spaces. In NeurIPS.
  11. 11.Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021a. W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In ASRU.
  12. 12.Yu-An Chung, Chenguang Zhu, and Michael Zeng. 2021b. Splat: Speech-language joint pre-training for spoken language understanding. In NAACL.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT.
  14. 14.Mattia Antonino Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019. MuST-C: a multilingual speech translation corpus. In NAACL-HLT.
  15. 15.Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed1. 2021. Hubert: How much can a bad teacher benefit asr pre-training. In ICASSP.
  16. 16.H. Inaguma, S. Kiyono, K. Duh, S. Karita, N. Soplin, T. Hayashi, and S. Watanabe. 2020. Espnet-st: All-in-one speech translation toolkit. In ACL.
  17. 17.Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019. Leveraging weakly supervised data to improve end-to-end speech-to-text translation. In ICASSP, pages 7180–7184.
  18. 18.J. Kahn, A. Lee, and A. Hannun. 2020. Self-training for end-to-end speech recognition. In Proc. of ICASSP.
  19. 19.J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux. 2020. Libri-light: A benchmark for asr with limited or no supervision. In ICASSP, pages 7669–7673. https://github.com/facebookresearch/libri-light.
  20. 20.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. In ICLR.
  21. 21.T. Kudo and J. Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In EMNLP.
  22. 22.Y. Lee and T. Kim. 2018. Learning pronunciation from a foreign language in speech synthesis networks. ArXiv.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020a. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL.
  24. 24.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020b. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  25. 25.Xian Li, Changhan Wang, Yun Tang, C. Tran, Yuqing Tang, Juan Miguel Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021. Multilingual speech translation from efficient finetuning of pretrained models. In ACL/IJCNLP.
  26. 26.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  27. 27.V. Panayotov, G. Chen, D. Povey, and S. Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In ICASSP.
  28. 28.D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le. 2019. Specaugment: A simple data augmentation method for automatic speech recognition. Interspeech.
  29. 29.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL-HLT.
  30. 30.Juan Miguel Pino, Qiantong Xu, Xutai Ma, Mohammad Javad Dousti, and Yun Tang. 2020. Self-training for end-to-end speech translation. In INTERSPEECH.
  31. 31.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011. The kaldi speech recognition toolkit. In ASRU.
  33. 33.Yun Tang, Juan Miguel Pino, Xian Li, Changhan Wang, and Dmitriy Genzel. 2021a. Improving speech translation by understanding and learning from the auxiliary text translation task. In ACL.
  34. 34.Yun Tang, Juan Miguel Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel. 2021b. A general multi-task learning framework to leverage text data for speech to text tasks. ICASSP.
  35. 35.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. ArXiv.
  36. 36.Changhan Wang, Anne Wu, and Juan Pino. 2020. Covost 2 and massively multilingual speech-to-text translation.
  37. 37.Changhan Wang, Anne Wu, Juan Miguel Pino, Alexei Baevski, Michael Auli, and Alexis Conneau. 2021a. Large-scale self- and semi-supervised learning for speech translation. In Interspeech.
  38. 38.Chengyi Wang, Yu Wu, Shujie Liu, Jinyu Li, Yao Qian, K. Kumatani, and Furu Wei. 2021b. Unispeech at scale: An empirical study of pre-training method on large-scale speech recognition dataset. In ICML.
  39. 39.Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017. Sequence-to-sequence models can directly translate foreign speech. In INTERSPEECH.
  40. 40.Alex Xiao, C. Fuegen, and Abdel rahman Mohamed. 2021. Contrastive semi-supervised learning for asr. In ICASSP.
  41. 41.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020. On layer normalization in the transformer architecture. In ICML, volume abs/2002.04745.
  42. 42.Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Y. Hannun, Gabriel Synnaeve, and Ronan Collobert. 2020. Iterative pseudo-labeling for speech recognition. In Interspeech, volume abs/2005.09267.
  43. 43.Rong Ye, Mingxuan Wang, and Lei Li. 2021. End-to-end speech translation via cross-modal progressive training. In INTERSPEECH.
  44. 44.Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu. 2020. Pushing the limits of semi-supervised learning for automatic speech recognition. Proc. of NeurIPS SAS Workshop.
  45. 45.Renjie Zheng, Junkun Chen, Mingbo Ma, and Liang Huang. 2021. Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation. In ICML.
  46. 46.Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, and Quoc V. Le. 2020. Rethinking pre-training and self-training. In NeurIPS, volume abs/2006.06882.

Citation

MLA
Tang, Y., et al. “Unified Speech-Text Pre-training for Speech Translation and Recognition”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1488–99, https://doi.org/10.18653/v1/2022.acl-long.105.
APA
Tang, Y., Gong, H., Dong, N., Wang, C., Hsu, W.-N., Gu, J., Baevski, A., Li, X., Mohamed, A., Auli, M., & Pino, J. (2022). Unified Speech-Text Pre-training for Speech Translation and Recognition. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1488–1499. https://doi.org/10.18653/v1/2022.acl-long.105
Chicago
Tang, Y., H. Gong, N. Dong, et al. 2022. “Unified Speech-Text Pre-training for Speech Translation and Recognition”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1488–99. https://doi.org/10.18653/v1/2022.acl-long.105.
Harvard
Tang, Y. et al. (2022) “Unified Speech-Text Pre-training for Speech Translation and Recognition”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1488–1499. Available at: https://doi.org/10.18653/v1/2022.acl-long.105.
Vancouver
1. Tang Y, Gong H, Dong N, et al (2022) Unified Speech-Text Pre-training for Speech Translation and Recognition. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1488–1499

BibTeX

@inproceedings{tang-etal-2022-unified,
    title = "Unified Speech-Text Pre-training for Speech Translation and Recognition",
    author = "Tang, Yun  and
      Gong, Hongyu  and
      Dong, Ning  and
      Wang, Changhan  and
      Hsu, Wei-Ning  and
      Gu, Jiatao  and
      Baevski, Alexei  and
      Li, Xian  and
      Mohamed, Abdelrahman  and
      Auli, Michael  and
      Pino, Juan",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.105/",
    doi = "10.18653/v1/2022.acl-long.105",
    pages = "1488--1499"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/