Cross-modal Contrastive Learning for Speech Translation
Rong YeMingxuan WangLei Li
Proposes a cross-modal contrastive learning framework that closes the representation gap between speech and text to boost end-to-end speech translation performance across the MuST-C benchmark.
End-to-end speech-to-text translation directly converts spoken audio in one language into written text in another using a unified neural model. While this avoids the compounded errors found in multi-stage systems, end-to-end models struggle due to a scarcity of paired speech-translation data and an underlying modality gap—a divergence where neural representations of spoken utterances and written transcripts do not align, hindering knowledge transfer from text translation to speech translation.
The article introduces and evaluates ConST, a cross-modal contrastive learning framework designed to bridge this speech-text representation gap within a multi-task learning architecture. ConST trains the system by pulling representations of matching speech and transcript pairs closer together while pushing unmatched pairs apart, using specialized data augmentation techniques such as span-masking, word repetition, and feature cut-offs to create challenging training examples.
The approach was evaluated across eight language translation directions on the standard MuST-C benchmark using English speech audio (over 385 hours per language pair), tested both with and without external machine translation text datasets. ConST achieved an average BLEU score of 29.4 when utilizing external text data, outperforming prior state-of-the-art end-to-end and cascaded systems across all evaluated language pairs. The contrastive learning component alone contributed an average improvement of 0.5 to 0.6 BLEU points over identical multi-task baselines lacking the contrastive objective, and applying contrastive loss to low-level acoustic and text embeddings proved substantially more effective than applying it to higher-level encoder representations. Crucially, ConST closed the representation gap by dramatically improving low-level cross-modal speech-to-text retrieval accuracy from 9.4% to 88.6%.
These findings demonstrate that explicitly aligning speech and text representations enables end-to-end models to better leverage abundant text translation resources, overcoming data bottlenecks and outperforming traditional cascaded translation pipelines. For practical deployment, the authors recommend configuring contrastive loss weights between 0.8 and 1.5, maintaining softmax temperatures between 0.02 and 0.05, and combining original contrastive pairs with data augmentation methods like sequence cut-offs for optimal performance.
However, decision-makers should note several operational boundaries. The model relies on paired speech and transcript data during training, meaning it cannot yet support low-resource languages that lack written transcriptions. Additionally, because the system was evaluated on clean benchmark datasets like TED talks, further pilot validation and analysis are necessary to determine its robustness against real-world background noise and varied utterance lengths before deploying in production environments.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Provides foundational techniques for self-supervised contrastive speech representation learning directly from raw audio waveforms, which underpin modern cross-modal speech encoders.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). Introduces the align-before-fuse paradigm that uses contrastive objectives to bridge modality gaps prior to cross-modal generation and fusion.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). Establishes contrastive representation learning principles for textual sequences, providing the conceptual groundwork for embedding semantic representations in shared spaces.
- Paper: Supervised Contrastive Learning, Prannay Khosla et al. (2020). Develops supervised contrastive loss formulations that extend contrastive objectives to multi-positive, semantically aligned pairings.
- Paper: Conformer: Convolution-augmented Transformer for Speech Recognition, Anmol Gulati et al. (2020). Presents the convolution-augmented Transformer architecture widely adopted as the acoustic encoder in end-to-end speech processing and translation.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). Details cross-modal attention mechanisms for unaligned sequential inputs, addressing fundamental representation alignment across differing temporal sampling rates.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Introduces cross-lingual representation pre-training objectives that facilitate cross-lingual transfer in neural sequence translation.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). Supplies a comprehensive survey and taxonomy of core multimodal challenges, notably cross-modal representation, alignment, and translation.
- Paper: Unified Speech-Text Pre-training for Speech Translation and Recognition, Yun Tang et al. (2022). Extends unified speech-text representation learning by scaling multi-task joint pre-training across speech translation and recognition with specialized shared encoder architectures.
- Paper: SECap: Speech Emotion Captioning with Large Language Model, Yaoxun Xu et al. (2024). Applies cross-modal alignment and contrastive representation techniques to bridge acoustic speech encoders with generative language models for spoken emotion captioning.
- Paper: SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words, Junyi Ao et al. (2024). Evaluates end-to-end spoken language understanding against traditional cascaded models across acoustic and conversational dimensions beyond text.
- Paper: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale, Matthew Le et al. (2023). Broadens speech-text alignment paradigms to large-scale generative modeling for universal text-guided multilingual speech tasks.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Surveys the broader ecosystem of multimodal foundation models that build upon alignment, projection, and cross-modal pre-training strategies.
