Unified Speech-Text Pre-training for Speech Translation and Recognition
Yun TangHongyu GongNing DongChanghan WangWei-Ning HsuJiatao GuAlexei BaevskiXian LiAbdelrahman MohamedMichael Auli
Proposes a joint speech-text pre-training framework that integrates linguistic knowledge from text into speech models across four multi-task objectives, providing tailored encoder sharing strategies to resolve subtask interference and substantially boosting performance on speech translation and recognition.
Building accurate automated speech recognition and speech translation systems typically requires substantial labeled audio data, which is scarce and expensive to produce for most languages. While large amounts of unlabeled audio and pure text data are widely available, existing machine learning methods struggle to jointly exploit both modalities in a single framework without experiencing cross-task interference or requiring complex multi-stage pipelines.
The article develops and evaluates Speech and Text Joint Pre-Training (STPT), a multi-task learning method designed to unify speech and text representations within an attention-based encoder-decoder model for speech translation and speech recognition.
The authors conducted large-scale computational experiments utilizing up to 60,000 hours of unlabeled English speech, extensive parallel and monolingual text datasets, and standard benchmark datasets including LibriSpeech and MuST-C. The approach integrates four distinct learning objectives: text-to-text modeling, self-supervised speech learning via a masked distribution alignment loss, supervised speech-to-phoneme classification, and sequence-to-sequence speech-to-text generation. To address negative task interactions identified through gradient similarity analysis, the researchers designed two tailored model structures: a fully shared encoder architecture for speech recognition and a partially shared encoder architecture for speech translation.
The evaluation produced four key findings. First, STPT established new state-of-the-art results in speech translation on the MuST-C benchmark, improving translation quality by 1.7 to 2.3 BLEU points over the strongest prior systems. Second, in speech recognition, the system achieved a competitive word error rate of 3.2% to 3.3% on LibriSpeech, matching specialized models without requiring an external language model during decoding. Third, model architecture design proved essential: while speech recognition benefited from full parameter sharing across all tasks, speech translation suffered from severe gradient conflicts that required isolating encoder subtasks via a partially shared structure. Fourth, the proposed soft-label distribution alignment loss for speech self-supervision consistently outperformed traditional contrastive objectives, yielding a 0.6 reduction in word error rate and up to 1.4 BLEU point gains in translation.
These findings indicate that directly integrating linguistic knowledge from text corpora into speech models during pre-training eliminates the operational overhead of maintaining separate language models during deployment. However, system architects must tailor parameter-sharing strategies specifically to downstream application goals rather than applying a universal multi-task structure.
Organizations developing speech processing pipelines should adopt the partially shared encoder design when training translation systems and the fully shared architecture for transcription tasks. Before deploying these models, practitioners should evaluate the underlying text corpora to prevent propagating offensive or biased training content into production systems. Future engineering efforts should explore expanding the framework to multilingual speech settings and investigating methods to stabilize performance when labeled training data is extremely limited.
The findings are supported by consistent empirical results across multiple benchmark splits. Readers should note, however, that performance degrades noticeably when supervised training audio is restricted below 100 hours or when acoustic conditions between pre-training and real-world inference diverge significantly.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Learn how self-supervised contrastive learning extracts robust representations directly from raw speech audio, providing the foundational speech modeling framework that the source builds on and benchmarks against.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). Understand masked prediction of hidden acoustic units as a primary paradigm for self-supervised speech representation learning underlying unified multimodal architectures.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Examine the sequence-to-sequence denoising autoencoder architecture for text, which inspires the unified encoder-decoder pre-training strategy used in the source.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Explore the unified text-to-text encoder-decoder pre-training framework that forms the architectural basis for multi-task cross-modal sequence modeling.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Discover how shared cross-lingual objectives align disparate input spaces, informing the auxiliary cross-modal tasks used in the source to align speech and text.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Review multilingual sequence-to-sequence denoising objectives that serve as a direct conceptual predecessor to unified generative cross-modal pre-training.
- Paper: WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, Sanyuan Chen et al. (2021). See how large-scale masked speech pre-training scales across diverse downstream speech processing tasks.
- Paper: Multi-Task Deep Neural Networks for Natural Language Understanding, Xiaodong Liu et al. (2019). Understand multi-task representation learning and how auxiliary supervised objectives can be integrated to mitigate overfitting in sequence models.
- Paper: Robust Speech Recognition via Large-Scale Weak Supervision, Alec Radford et al. (2023). Discover how scaling weakly supervised multitask speech-to-text pre-training across massive web corpora advances both speech recognition and translation beyond self-supervised encoder-decoder methods.
- Paper: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale, Matthew Le et al. (2023). Explore how unified text and speech representations are extended from perception tasks to large-scale generative infilling and speech synthesis.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). See how joint audio and visual representations are integrated into large language models to support instruction-tuned multimodal conversational systems.
