Cross-modal Contrastive Learning for Speech Translation

Rong YeMingxuan WangLei Li

article2022NAACL112 citations

Proposes a cross-modal contrastive learning framework that closes the representation gap between speech and text to boost end-to-end speech translation performance across the MuST-C benchmark.

Listen

End-to-end speech-to-text translation directly converts spoken audio in one language into written text in another using a unified neural model. While this avoids the compounded errors found in multi-stage systems, end-to-end models struggle due to a scarcity of paired speech-translation data and an underlying modality gap—a divergence where neural representations of spoken utterances and written transcripts do not align, hindering knowledge transfer from text translation to speech translation.

The article introduces and evaluates ConST, a cross-modal contrastive learning framework designed to bridge this speech-text representation gap within a multi-task learning architecture. ConST trains the system by pulling representations of matching speech and transcript pairs closer together while pushing unmatched pairs apart, using specialized data augmentation techniques such as span-masking, word repetition, and feature cut-offs to create challenging training examples.

The approach was evaluated across eight language translation directions on the standard MuST-C benchmark using English speech audio (over 385 hours per language pair), tested both with and without external machine translation text datasets. ConST achieved an average BLEU score of 29.4 when utilizing external text data, outperforming prior state-of-the-art end-to-end and cascaded systems across all evaluated language pairs. The contrastive learning component alone contributed an average improvement of 0.5 to 0.6 BLEU points over identical multi-task baselines lacking the contrastive objective, and applying contrastive loss to low-level acoustic and text embeddings proved substantially more effective than applying it to higher-level encoder representations. Crucially, ConST closed the representation gap by dramatically improving low-level cross-modal speech-to-text retrieval accuracy from 9.4% to 88.6%.

These findings demonstrate that explicitly aligning speech and text representations enables end-to-end models to better leverage abundant text translation resources, overcoming data bottlenecks and outperforming traditional cascaded translation pipelines. For practical deployment, the authors recommend configuring contrastive loss weights between 0.8 and 1.5, maintaining softmax temperatures between 0.02 and 0.05, and combining original contrastive pairs with data augmentation methods like sequence cut-offs for optimal performance.

However, decision-makers should note several operational boundaries. The model relies on paired speech and transcript data during training, meaning it cannot yet support low-resource languages that lack written transcriptions. Additionally, because the system was evaluated on clean benchmark datasets like TED talks, further pilot validation and analysis are necessary to determine its robustness against real-world background noise and varied utterance lengths before deploying in production environments.

Cover for Cross-modal Contrastive Learning for Speech Translation

Abstract

How can we learn unified representations for spoken utterances and their written text? Learning similar representations for semantically similar speech and text is important for speech translation. To this end, we propose ConST, a cross-modal contrastive learning method for end-to-end speech-to-text translation. We evaluate ConST and a variety of previous baselines on a popular benchmark MuST-C. Experiments show that the proposed ConST consistently outperforms the previous methods, and achieves an average BLEU of 29.4. The analysis further verifies that ConST indeed closes the representation gap of different modalities — its learned representation improves the accuracy of cross-modal speech-text retrieval from 4% to 88%. Code and models are available at https://github.com/ReneeYe/ConST.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The ConST Approach
  • 3.1 Model Framework
  • 3.2 Cross-modal Contrastive Learning
  • 3.3 Mining Hard Examples for Contrastive Learning
  • 4 Experiments
  • 4.1 Experimental Setups
  • 4.2 Main Results
  • 5 Analysis
  • 5.1 Is contrastive loss effective?
  • 5.2 Which layer to contrast on?
  • 5.3 Is contrastive loss better than other losses?
  • 5.4 Analysis on the hard example mining strategies
  • 6 Why does cross-modal contrastive learning work? - Analysis on the Modality Gap
  • 6.1 Visualization of Representation
  • 6.2 Cross-modal Retrieval
  • 7 Case Analysis
  • 8 Conclusion
  • 9 Broader Impact
  • References
  • A Statistics of all datasets
  • B Experimental Details
  • C The Choice for Hyper-parameters
  • D Data Scale for Fine-tuning

Knowls

  1. Knowl 1 — ConST unified speech–text multi-task architecture

    model/method

    ConST is an end-to-end speech-to-text translation model that uses one shared Transformer encoder–decoder for three related tasks: speech translation (speech ss to target-language text yy), automatic speech recognition (speech ss to source-language transcript xx), and text machine translation (source transcript xx to target text yy). A speech encoder, denoted S\mbox−Enc\mathrm{S\mbox{-}Enc}, processes raw 16-kHz waveforms using Wav2vec 2.0 followed by two convolutional layers; the resulting speech features and source-text word embeddings are both passed to the shared Transformer encoder and decoder. This design permits the Transformer module to exploit external text-translation data while sharing representations across speech and text inputs.

    For a speech–transcript–translation triplet (s,x,y)(s,x,y), ConST combines the three autoregressive cross-entropy objectives with a cross-modal contrastive objective:

    L=LST+LASR+LMT+λLCTR,\mathcal{L}=\mathcal{L}_{\mathrm{ST}}+\mathcal{L}_{\mathrm{ASR}}+\mathcal{L}_{\mathrm{MT}}+\lambda\mathcal{L}_{\mathrm{CTR}},

    where λ\lambda is the contrastive-loss weight and

    LST=−∑n=1∣y∣log⁡Pθ(yn∣s),LASR=−∑n=1∣x∣log⁡Pθ(xn∣s),LMT=−∑n=1∣y∣log⁡Pθ(yn∣x).\mathcal{L}_{\mathrm{ST}}=-\sum_{n=1}^{|y|}\log P_\theta(y_n\mid s),\qquad \mathcal{L}_{\mathrm{ASR}}=-\sum_{n=1}^{|x|}\log P_\theta(x_n\mid s),\qquad \mathcal{L}_{\mathrm{MT}}=-\sum_{n=1}^{|y|}\log P_\theta(y_n\mid x).

    Here x=(x1,…,x∣x∣)x=(x_1,\ldots,x_{|x|}) is the source-language transcript, y=(y1,…,y∣y∣)y=(y_1,\ldots,y_{|y|}) is the target-language translation, ss is the speech waveform, and PθP_\theta is the model’s token probability distribution.

  2. Knowl 2 — Cross-modal N-pair contrastive objective

    equation

    ConST explicitly aligns a speech utterance with its matching transcript while separating it from unrelated transcripts in the same mini-batch. For a speech waveform ss and its source-language transcript xx, ConST forms low-level modality representations by time-averaging the speech-encoder output and the word-embedding sequence:

    us=MeanPool⁡(S\mbox−Enc(s)),vx=MeanPool⁡(Emb⁡(x)).u_s=\operatorname{MeanPool}(\mathrm{S\mbox{-}Enc}(s)),\qquad v_x=\operatorname{MeanPool}(\operatorname{Emb}(x)).

    For each positive pair (s,x)(s,x) in a training batch B\mathcal{B}, let {xi−}i=1N−1\{x_i^-\}_{i=1}^{N-1} be N−1N-1 transcripts sampled from the same batch as negative examples, and let As={x}∪{xi−}i=1N−1A_s=\{x\}\cup\{x_i^-\}_{i=1}^{N-1}. The cross-modal loss is the multi-class N-pair objective

    LCTR=−∑(s,x)∈Blog⁡exp⁡ ⁣(sim⁡(νs,vx)/τ)∑xj∈Asexp⁡ ⁣(sim⁡(νs,vxj)/τ),\mathcal{L}_{\mathrm{CTR}}=-\sum_{(s,x)\in\mathcal{B}}\log\frac{\exp\!\left(\operatorname{sim}(\nu_s,v_x)/\tau\right)}{\sum_{x_j\in A_s}\exp\!\left(\operatorname{sim}(\nu_s,v_{x_j})/\tau\right)},

    where τ>0\tau>0 is the temperature and sim⁡(a,b)=a⊤b/(∥a∥2∥b∥2)\operatorname{sim}(a,b)=a^\top b/(\lVert a\rVert_2\lVert b\rVert_2) is cosine similarity. The positive transcript is therefore pulled toward its corresponding speech representation, while all other transcripts in the batch are pushed away; ConST used τ=0.02\tau=0.02 in its main experiments.

  3. Knowl 3 — Hard-example mining for robust cross-modal alignment

    model/method

    ConST augments the basic speech–transcript contrastive pairs with four optional hard-example strategies. Each modified input or representation remains paired with the original matching transcript as a positive example, while corresponding modified transcripts or unrelated batch transcripts serve as negatives.

    • Span-masked augmentation: Starting indices are sampled from the waveform with probability p=0.25p=0.25, and each selected index begins a blanked span of M=3600M=3600 waveform time steps, corresponding to approximately 0.2250.225 seconds at 16 kHz. The masked waveform is contrasted with the original transcript.
    • Word repetition: Each source-language subword token is independently duplicated kk additional times, where k∼Poisson⁡(1)k\sim\operatorname{Poisson}(1) and k∈{0,1,2,…}k\in\{0,1,2,\ldots\}. The resulting sentence is an additional positive text representation for the original speech; repetition preserves sentence meaning while changing sequence length.
    • Sequence cut-off: A contiguous block along the sequence dimension of the speech representation is set to zero.
    • Feature cut-off: A contiguous block along the feature dimension of the speech representation is set to zero.

    For cut-off, the speech representation has shape T×dT\times d, where TT is the number of encoded time steps and dd is the feature dimension. Unlike elementwise dropout, cut-off removes dimensional blocks; both sequence and feature cut-off used a cut-off rate of 0.10.1 in the experiments.

  4. Knowl 4 — MuST-C evaluation and model configuration

    experimental setup

    ConST was evaluated on all eight English-to-language directions of MuST-C v1.0 using the case-sensitive, detokenized BLEU score on the tst-COMMON set. Each direction contains at least 385 hours of TED-talk speech. The expanded setting additionally used WMT parallel text for English–German, English–Spanish, English–French, English–Romanian, and English–Russian, and OPUS100 parallel text for English–Italian, English–Dutch, and English–Portuguese.

    Direction MuST-C hours MuST-C sentences External MT corpus and sentences
    En–De 408 234K WMT16, 4.6M
    En–Es 504 270K WMT13, 15.2M
    En–Fr 492 292K WMT14, 40.8M
    En–It 465 258K OPUS100, 1.0M
    En–Nl 442 253K OPUS100, 1.0M
    En–Pt 385 211K OPUS100, 1.0M
    En–Ro 432 240K WMT16, 0.6M
    En–Ru 489 270K WMT16, 2.5M

    The speech encoder used Wav2vec 2.0 pretrained only on LibriSpeech, followed by two convolutional layers with kernel size 5, stride 2, and hidden size 512. The Transformer used six encoder layers, six decoder layers, hidden size d=512d=512, eight attention heads, and feed-forward hidden size 2048, for approximately 150M parameters. Text was jointly SentencePiece-tokenized with a 10K vocabulary. The main experiments used contrastive weight λ=1.5\lambda=1.5 for German and Dutch and λ=1.0\lambda=1.0 for the other languages.

  5. Knowl 5 — ConST achieves the strongest MuST-C translation scores

    data/table

    On MuST-C tst-COMMON, ConST improved over the matched XSTNet multi-task baseline by 0.5 average BLEU without external MT data and by 0.6 average BLEU with external MT data. With external MT data, ConST reached an average BLEU of 29.4, the highest average among the compared systems in the reported benchmark results.

    Model and setting De Es Fr It Nl Pt Ro Ru Avg.
    XSTNet, no external MT 25.5 29.6 36.0 25.5 30.0 31.3 25.1 16.9 27.5
    ConST, no external MT 25.7 30.4 36.8 26.3 30.6 32.0 24.8 17.3 28.0
    Chimera, external MT 27.1 30.6 35.6 25.0 29.2 30.2 24.0 17.4 27.4
    XSTNet, external MT 27.1 30.8 38.0 26.4 31.2 32.4 25.7 18.5 28.8
    STEMM, external MT 28.7 31.0 37.4 25.8 30.5 31.7 24.5 17.8 28.4
    ConST, external MT 28.3 32.0 38.3 27.2 31.7 33.1 25.6 18.9 29.4

    ConST also outperformed the reported cascaded systems on the three directions with direct comparisons: on En–De/En–Fr/En–Ru, ConST scored 28.3/38.3/18.9 BLEU, compared with 28.1/–/– for the strongest listed cascade from Xu et al., 25.2/34.9/17.0 for the cascade from Ye et al., and 23.6/33.8/16.4 for Espnet.

  6. Knowl 6 — Ablations show that contrastive alignment and multi-task supervision are complementary

    data/table

    On MuST-C En–De, removing the auxiliary ASR and MT losses while retaining speech translation and contrastive learning reduced BLEU, and removing the contrastive term as well reduced it further. This indicates that both multi-task supervision and cross-modal contrastive supervision contribute to translation quality.

    Configuration Without external MT With external MT
    ConST 25.7 28.3
    ConST without L_ASR,L_MT 24.6 27.0
    ConST without L_ASR,L_MT,L_CTR 23.6 26.3

    The location of the contrastive objective also matters. Low-level representations—mean-pooled speech-encoder features and mean-pooled lexical embeddings—outperformed high-level representations obtained from the Transformer encoder, although both contrastive variants beat the no-contrastive baseline.

    Representation BLEU ChrF++ BLEURT
    Low-level representation 28.3 53.2 64.5
    High-level representation 27.5 52.6 63.6
    Without contrastive loss 27.1 52.1 62.4

    Using negative examples was also more effective than alternatives that only reduce distances or impose alignments. On En–De, the contrastive loss obtained BLEU 28.3, compared with 27.6 for CTC loss, 27.3 for an L2L_2 representation-matching loss, and 27.1 without an additional loss; the corresponding ChrF++ scores were 53.2, 53.0, 52.4, and 52.1, and the BLEURT scores were 64.5, 64.1, 63.0, and 62.4.

  7. Knowl 7 — ConST closes the speech–text modality gap

    empirical result

    A representation analysis on the MuST-C En–De test set found that a multi-task model without explicit contrastive alignment formed two separated clusters when speech and transcript representations were projected to two dimensions with t-SNE and visualized using bivariate kernel-density contours. ConST produced substantially overlapping speech and transcript contours, indicating that its speech representations encode information more similar to the corresponding textual representations.

    The same effect appeared in cross-modal retrieval. For each speech utterance, the system retrieved the nearest transcript using cosine similarity and was evaluated by top-1 accuracy. The low-level speech representation was especially affected: ConST increased retrieval accuracy from 9.4% to 88.6%. High-level Transformer representations already aligned well without contrastive learning, changing only from 94.7% to 95.0% with ConST.

    Representation Without contrastive loss With contrastive loss
    Low-level 9.4% 88.6%
    High-level 94.7% 95.0%
  8. Knowl 8 — Hard-example mining improves contrastive training

    data/table

    The four hard-example strategies were evaluated in all 15 combinations of the original contrastive objective, word repetition (Rep), span-masked augmentation (SMA), sequence cut-off (SCut), and feature cut-off (FCut). These En–De BLEU scores used contrastive weight λ=1.0\lambda=1.0; the strong multi-task baseline without contrastive learning scored 27.1 BLEU. Every combination exceeded that baseline, and the best result, 28.3 BLEU, came from combining the original contrastive objective with sequence cut-off.

    Original Rep SMA SCut FCut
    Original 27.9 27.8 28.1 28.3 28.2
    Rep 27.8 27.8 27.6 27.8 27.5
    SMA 28.1 27.6 28.1 27.7 27.6
    SCut 28.3 27.8 27.7 27.6 27.4
    FCut 28.2 27.5 27.6 27.4 27.8

    The strongest configurations generally retained the original, unmodified positive and negative examples in addition to the hard examples. Using sequence and feature cut-off together without the original examples performed worst, which the authors attribute to excessive information loss.

  9. Knowl 9 — Contrastive learning is particularly useful with very little labeled speech-translation data

    empirical result

    A low-resource experiment kept the same MT-pretrained Transformer module and varied the amount of labeled MuST-C En–De speech–transcript–translation data. The three systems were ConST, ConST without the contrastive loss, and a model without both multi-task auxiliary losses and contrastive loss.

    Training data 1 hour 10 hours 100 hours 408 hours
    ConST 5.4 18.7 24.9 28.3
    Without contrastive loss 3.9 18.2 23.5 27.1
    Without multi-task and contrastive losses 0.9 15.3 22.8 23.6

    The contrastive objective produced its largest absolute gain in the most data-scarce condition, improving BLEU from 3.9 to 5.4 with only one hour of labeled speech. With 100 hours, ConST reached 24.9 BLEU, exceeding the 23.6 BLEU obtained by the model trained on all 408 hours without MT pretraining or the auxiliary objectives.

  10. Knowl 10 — Operational limitations of ConST

    limitation

    ConST was evaluated on public speech with hundreds of hours of labeled data and is not presented as an industrial-grade solution. The authors note that real-world speech is noisier and has a more complex length distribution than the public benchmark, conditions that an end-to-end model alone may not handle well. ConST also requires labeled speech–transcription data to learn strong speech representations and does not address untranscribed scenarios; consequently, it is not directly applicable to languages or dialects lacking transcripts, translations, or both.

Coverage note — The qualitative translation case study and detailed temperature/loss-weight sweeps were omitted because they provide supporting examples and hyperparameter sensitivity rather than additional load-bearing method or result claims.

References

  1. 1.Ashkan Alinejad and Anoop Sarkar. 2020. Effectively pretraining a speech translation decoder with machine translation data. In Proc. of EMNLP, pages 8014–8020.
  2. 2.Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, et al. 2020. Findings of the iwslt 2020 evaluation campaign. In Proc. of IWSLT, pages 1–34.
  3. 3.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proc. of NeurIPS.
  4. 4.Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater. 2019. Pre-training on high-resource speech recognition improves low-resource speech-to-text translation. In Proc. of NAACL-HLT, pages 58–68.
  5. 5.Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, and Marco Turchi. 2021. Cascade versus direct speech translation: Do the differences still make a difference? In Proc. of ACL.
  6. 6.Alexandre Berard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018. End-to-end automatic speech translation of audiobooks. In Proc. of ICASSP, pages 6224–6228.
  7. 7.Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016. Listen and translate: A proof of concept for end-to-end speech-to-text translation. In NIPS workshop on End-to-end Learning for Speech and Audio Processing.
  8. 8.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 131–198.
  9. 9.Shuyang Cao and Lu Wang. 2021. Cliff: Contrastive learning for improving faithfulness and factuality in abstractive summarization. In Proc. of ACL, pages 6633–6649.
  10. 10.Junkun Chen, Mingbo Ma, Renjie Zheng, and Liang Huang. 2021. Specrec: An alternative solution for improving end-to-end speech-to-text translation via spectrogram reconstruction. Proc. of InterSpeech, pages 2232–2236.
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020a. A simple framework for contrastive learning of visual representations. In Proc. of ICML, volume 119 of Proceedings of Machine Learning Research, pages 1597–1607.
  12. 12.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020b. Uniter: Universal image-text representation learning. In Proc. of ECCV, pages 104–120. Springer.
  13. 13.Sumit Chopra, Raia Hadsell, and Yann LeCun. 2005. Learning a similarity metric discriminatively, with application to face verification. In Proc. of CVPR, volume 1, pages 539–546. IEEE.
  14. 14.Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019. MuST-C: a Multilingual Speech Translation Corpus. In Proc. of NAACL-HLT, pages 2012–2017.
  15. 15.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Proc. of NeurIPS, pages 13063–13075.
  16. 16.Linhao Dong, Shuang Xu, and Bo Xu. 2018. Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition. In Proc. of ICASSP, pages 5884–5888.
  17. 17.Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, and Lei Li. 2021a. Consecutive decoding for speech-to-text translation. In Proc. of AAAI.
  18. 18.Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, and Lei Li. 2021b. Listen, understand and translate: Triple supervision decouples end-to-end speech-to-text translation. In Proc. of AAAI, volume 35, pages 12749–12759.
  19. 19.Hongchao Fang, Sicheng Wang, Meng Zhou, Jiayuan Ding, and Pengtao Xie. 2020. Cert: Contrastive self-supervised learning for language understanding. arXiv preprint arXiv:2005.12766.
  20. 20.Qingkai Fang, Rong Ye, Lei Li, Yang Feng, and Mingxuan Wang. 2022. Stemm: Self-learning with speech-text manifold mixup for speech translation. In Proc. of ACL.
  21. 21.Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu, Hao Zhou, and Lei Li. 2022. Contextual representation learning beyond masked language modeling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2701–2714, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In Proc. of EMNLP.
  23. 23.Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proc. of ICML, volume 148, pages 369–376.
  24. 24.Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2016. Lstm: A search space odyssey. IEEE transactions on neural networks and learning systems, 28(10):2222–2232.
  25. 25.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. Proc. of NeurIPS, 33:21271–21284.
  26. 26.Michael Gutmann and Aapo Hyvärinen. 2010. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 297–304. JMLR Workshop and Conference Proceedings.
  27. 27.Chi Han, Mingxuan Wang, Heng Ji, and Lei Li. 2021. Learning shared semantic space for speech-to-text translation. In Proc. of ACL - Findings.
  28. 28.Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi. 2019. A comprehensive overhaul of feature distillation. In Proc. of the ICCV, pages 1921–1930.
  29. 29.Yushi Hu, Shane Settle, and Karen Livescu. 2020. Multilingual jointly trained acoustic and written word embeddings. In Proc. of INTERSPEECH.
  30. 30.Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe. 2021. Source and target bidirectional knowledge distillation for end-to-end speech translation. In Proc. of NAACL, pages 1872–1881.
  31. 31.Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita, Nelson Yalta, Tomoki Hayashi, and Shinji Watanabe. 2020. ESPnet-ST: All-in-one speech translation toolkit. In Proc. of ACL, pages 302–311.
  32. 32.Sathish Indurthi, Houjeung Han, Nikhil Kumar Lakumarapu, Beomseok Lee, Insoo Chung, Sangha Kim, and Chanwoo Kim. 2020. Data efficient direct speech-to-text translation with modality agnostic meta-learning. In Proc. of ICASSP. IEEE.
  33. 33.Sathish Indurthi, Mohd Abbas Zaidi, Nikhil Kumar Lakumarapu, Beomseok Lee, Hyojung Han, Seokchan Ahn, Sangha Kim, Chanwoo Kim, and Inchul Hwang. 2021. Task aware multi-task learning for speech to text tasks. In Proc. of ICASSP, pages 7723–7727. IEEE.
  34. 34.Herman Kamper, Yevgen Matusevych, and Sharon Goldwater. 2020. Multilingual acoustic word embedding models for processing zero-resource languages. In Proc. of ICASSP, pages 6414–6418. IEEE.
  35. 35.Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2017. Structured-based curriculum learning for end-to-end english-japanese speech translation. In Proc. of INTERSPEECH, pages 2630–2634.
  36. 36.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. Proc. of NeurIPS, 33.
  37. 37.Philipp Koehn et al. 2005. Europarl: A parallel corpus for statistical machine translation. In MT summit, volume 5, pages 79–86. Citeseer.
  38. 38.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proc. of EMNLP, pages 66–71.
  39. 39.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. In Proc. of NeurIPS.
  40. 40.Hang Le, Juan Pino, Changhan Wang, Jiatao Gu, Didier Schwab, and Laurent Besacier. 2020. Dual-decoder transformer for joint automatic speech recognition and multilingual speech translation. In Proc. of COLING, pages 3520–3533.
  41. 41.Hang Le, Juan Pino, Changhan Wang, Jiatao Gu, Didier Schwab, and Laurent Besacier. 2021. Lightweight adapter tuning for multilingual speech translation. In Proc. of ACL.
  42. 42.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2021. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. In Proc. of ACL.
  43. 43.Pierre Lison and Jörg Tiedemann. 2016. OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. In Proc. of LREC, pages 923–929.
  44. 44.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020a. Multilingual denoising pre-training for neural machine translation. TACL, 8:726–742.
  45. 45.Yuchen Liu, Junnan Zhu, Jiajun Zhang, and Chengqing Zong. 2020b. Bridging the modality gap for speech-to-text translation. arXiv preprint arXiv:2010.14920.
  46. 46.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Proc. of NeurIPS, 32.
  47. 47.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.
  48. 48.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proc. of NAACL - Demonstrations, pages 48–53.
  49. 49.Shruti Palaskar, Vikas Raunak, and Florian Metze. 2019. Learned in speech recognition: Contextual acoustic word embeddings. In Proc. of ICASSP, pages 6530–6534. IEEE.
  50. 50.Xiao Pan, Liwei Wu, Mingxuan Wang, and Lei Li. 2021. Contrastive learning for many-to-many multilingual neural machine translation. In Proc. of ACL.
  51. 51.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In Proc. of ICASSP, pages 5206–5210.
  52. 52.Sara Papi, Marco Gaido, Matteo Negri, and Marco Turchi. 2021. Speechformer: Reducing information loss in direct speech translation. In Proc. of EMNLP.
  53. 53.Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin Dogus Cubuk, and Quoc V. Le. 2019. Specaugment: A simple augmentation method for automatic speech recognition. In Proc. of INTERSPEECH.
  54. 54.Emanuel Parzen. 1962. On estimation of a probability density function and mode. The annals of mathematical statistics, 33(3):1065–1076.
  55. 55.Juan Pino, Qiantong Xu, Xutai Ma, Mohammad Javad Dousti, and Yun Tang. 2020. Self-training for end-to-end speech translation. In Proc. of INTERSPEECH, pages 1476–1480.
  56. 56.Maja Popovic. 2017. chrf++: words helping character n-grams. In Proceedings of the second conference on machine translation, pages 612–618.
  57. 57.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191.
  58. 58.Tomasz Potapczyk and Paweł Przybysz. 2020. Srpol’s system for the iwslt 2020 end-to-end speech translation task. In Proc. of IWSLT, pages 89–94.
  59. 59.Mirco Ravanelli and Yoshua Bengio. 2018. Learning speaker representations with mutual information. arXiv preprint arXiv:1812.00271.
  60. 60.Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019. wav2vec: Unsupervised pre-training for speech recognition. In Proc. of INTERSPEECH.
  61. 61.Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. Facenet: A unified embedding for face recognition and clustering. In Proc. of CVPR, pages 815–823.
  62. 62.Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. Bleurt: Learning robust metrics for text generation. In Proc. of ACL.
  63. 63.Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen. 2020. A simple but tough-to-beat data augmentation approach for natural language understanding and generation. arXiv preprint arXiv:2009.13818.
  64. 64.Kihyuk Sohn. 2016. Improved deep metric learning with multi-class n-pair loss objective. In Proc. of NeurIPS, volume 29, pages 1857–1865. Curran Associates, Inc.
  65. 65.Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, et al. 2022. Unified speech-text pre-training for speech translation and recognition. In Proc. of ACL.
  66. 66.Yun Tang, Juan Pino, Xian Li, Changhan Wang, and Dmitriy Genzel. 2021a. Improving speech translation by understanding and learning from the auxiliary text translation task. In Proc. of ACL.
  67. 67.Yun Tang, Juan Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel. 2021b. A general multi-task learning framework to leverage text data for speech to text tasks. In Proc. of ICASSP, pages 6209–6213. IEEE.
  68. 68.Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Dan Belov, and Demis Hassabis. 2018. Parallel wavenet: Fast high-fidelity speech synthesis. In Proc. of ICML, volume 80 of Proceedings of Machine Learning Research, pages 3915–3923.
  69. 69.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  70. 70.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proc. of NeurIPS, pages 5998–6008.
  71. 71.Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020a. Fairseq s2t: Fast speech-to-text modeling with fairseq. In Proc. of AACL, pages 33–39.
  72. 72.Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski, Michael Auli, and Alexis Conneau. 2021a. Large-scale self-and semi-supervised learning for speech translation. In Proc. of INTERSPEECH.
  73. 73.Chengyi Wang, Yu Wu, Shujie Liu, Ming Zhou, and Zhenglu Yang. 2020b. Curriculum pre-training for end-to-end speech translation. In Proc. of ACL, pages 3728–3738.
  74. 74.Danqing Wang, Jiaze Chen, Hao Zhou, Xipeng Qiu, and Lei Li. 2021b. Contrastive aligned joint learning for multilingual summarization. In Proc. of ACL - Findings.
  75. 75.Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017. Sequence-to-sequence models can directly translate foreign speech. In Proc. of INTERSPEECH, pages 2625–2629.
  76. 76.Anne Wu, Changhan Wang, Juan Pino, and Jiatao Gu. 2020. Self-supervised representations improve end-to-end speech translation. In Proc. of INTERSPEECH.
  77. 77.Hao Wu, Jiayuan Mao, Yufeng Zhang, Weiwei Sun, Yuning Jiang, Lei Li, and Wei-Ying Ma. 2019. Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations. In Proc. of CVPR.
  78. 78.Xing Wu, Chaochen Gao, Liangjun Zang, Jizhong Han, Zhongyuan Wang, and Songlin Hu. 2021. Esimcse: Enhanced sample building method for contrastive learning of unsupervised sentence embedding. arXiv preprint arXiv:2109.04380.
  79. 79.Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Qi Ju, Tong Xiao, Jingbo Zhu, et al. 2021. Stacked acoustic-and-textual encoding: Integrating the pretrained models into speech translation encoders. In Proc. of ACL.
  80. 80.Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021. Consert: A contrastive framework for self-supervised sentence representation transfer. In Proc. of ACL.
  81. 81.Rong Ye, Mingxuan Wang, and Lei Li. 2021. End-to-end speech translation via cross-modal progressive training. In Proc. of INTERSPEECH.
  82. 82.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation. In Proc. of ACL, pages 1628–1639.
  83. 83.Chengqi Zhao, Mingxuan Wang, Qianqian Dong, Rong Ye, and Lei Li. 2021a. NeurST: Neural speech translation toolkit. In Proc. of ACL - System Demonstrations.
  84. 84.Jiawei Zhao, Wei Luo, Boxing Chen, and Andrew Gilman. 2021b. Mutual-learning improves end-to-end speech translation. In Proc. of the EMNLP, pages 3989–3994.
  85. 85.Renjie Zheng, Junkun Chen, Mingbo Ma, and Liang Huang. 2021. Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation. In Proc. of ICML.
  86. 86.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and vqa. In Proc. of AAAI, volume 34, pages 13041–13049.

Citation

MLA
Ye, R., et al. “Cross-modal Contrastive Learning for Speech Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5099–113, https://doi.org/10.18653/v1/2022.naacl-main.376.
APA
Ye, R., Wang, M., & Li, L. (2022). Cross-modal Contrastive Learning for Speech Translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5099–5113. https://doi.org/10.18653/v1/2022.naacl-main.376
Chicago
Ye, R., M. Wang, and L. Li. 2022. “Cross-modal Contrastive Learning for Speech Translation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5099–5113. https://doi.org/10.18653/v1/2022.naacl-main.376.
Harvard
Ye, R., Wang, M. and Li, L. (2022) “Cross-modal Contrastive Learning for Speech Translation”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5099–5113. Available at: https://doi.org/10.18653/v1/2022.naacl-main.376.
Vancouver
1. Ye R, Wang M, Li L (2022) Cross-modal Contrastive Learning for Speech Translation. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5099–5113

BibTeX

@inproceedings{ye-etal-2022-cross,
    title = "Cross-modal Contrastive Learning for Speech Translation",
    author = "Ye, Rong  and
      Wang, Mingxuan  and
      Li, Lei",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.376/",
    doi = "10.18653/v1/2022.naacl-main.376",
    pages = "5099--5113"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/