Textless Speech-to-Speech Translation on Real Data

Ann LeeHongyu GongPaul-Ambroise DuquenneHolger SchwenkPeng-Jen ChenChanghan WangSravya PopuriYossi AdiJuan Miguel PinoJiatao Gu

article2022NAACL193 citations

Presents a textless speech-to-speech translation system that trains directly on real-world multi-speaker audio by introducing a unit-based normalization technique that removes speaker and accent variations using only ten minutes of paired speech data.

Listen

More than 40 percent of the world's spoken languages lack standard written forms, creating a major barrier for conventional speech translation systems that depend on intermediate text transcriptions and written data. While direct speech-to-speech translation systems can bypass text, existing models struggle when trained on real-world multi-speaker audio because natural variations in accents, speaking styles, and background noise distort translation targets.

The article demonstrates a direct, textless speech-to-speech translation system capable of training on real-world multi-speaker audio without relying on any text or phoneme annotations. The core objective is to evaluate whether a self-supervised speech normalization method can effectively remove multi-speaker acoustic variations while preserving the core spoken content across multiple language pairs.

To accomplish this, the authors developed a speech normalizer by fine-tuning a pre-trained multilingual speech encoder using a small amount of paired audio from multiple speakers and a single reference speaker. This component transforms diverse multi-speaker speech into standardized discrete acoustic units. The authors trained a speech-to-unit translation model across four language pairs involving English, Spanish, and French using real parliamentary speech from the VoxPopuli dataset and automatically mined audio data from LibriVox, using a specialized unit-based synthesizer to generate the final audio output.

The findings show that unit-based speech normalization significantly improves translation quality and output clarity. Using only 10 minutes of paired audio to train the normalizer yielded an average gain of 3.2 BLEU translation points over unnormalized baselines, rising to an average improvement of 4.9 BLEU points when using 10 hours of normalization data. Integrating automatically mined speech data provided an additional 2.0 BLEU gain across evaluated language directions. Furthermore, human evaluation showed substantial gains in audio naturalness, increasing mean opinion scores by an average of 0.85 points, effectively eliminating speech artifacts like stuttering and matching the performance of text-based cascaded systems.

These results indicate that organizations can build high-quality translation systems for unwritten or low-resource spoken languages without expensive text transcription pipelines or manual data labeling. Relying on self-supervised representations and directly mined speech drastically lowers data engineering costs and accelerates deployment timelines for cross-lingual spoken communication tools.

Organizations developing speech translation capabilities should adopt discrete unit normalization pipelines to leverage unannotated and mined audio datasets. To maximize efficiency, teams can train the normalizer using approximately one hour of paired single-speaker reference audio, as the performance gains between 1 hour and 10 hours of normalization data are minimal. Future work should focus on exploring additional self-supervised pre-training strategies across wider varieties of language families.

Confidence in these findings is high for European parliamentary and audiobook speech domains across the tested languages. However, decision-makers should note that reference units in these experiments were synthetically assisted during normalizer training, and performance across heavily out-of-domain conversational data or highly under-resourced languages remains to be fully verified.

arXiv: 2112.08352
Cover for Textless Speech-to-Speech Translation on Real Data

Abstract

We present a textless speech-to-speech translation (S2ST) system that can translate speech from one language into another language and can be built without the need of any text data. Different from existing work in the literature, we tackle the challenge in modeling multi-speaker target speech and train the systems with real-world S2ST data. The key to our approach is a self-supervised unit-based speech normalization technique, which finetunes a pre-trained speech encoder with paired audios from multiple speakers and a single reference speaker to reduce the variations due to accents, while preserving the lexical content. With only 10 minutes of paired data for speech normalization, we obtain on average 3.2 BLEU gain when training the S2ST model on the VoxPopuli S2ST dataset, compared to a baseline trained on un-normalized speech target. We also incorporate automatically mined S2ST data and show an additional 2.0 BLEU gain. To our knowledge, we are the first to establish a textless S2ST technique that can be trained with real-world data and works for multiple language pairs¹.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 System
  • 3.1 Self-supervised Unit-based Speech Normalization
  • 3.2 Textless S2ST
  • 4 Experimental Setup
  • 4.1 Data
  • 4.2 Multilingual HuBERT (mHuBERT)
  • 4.3 Baselines
  • 4.4 Evaluation
  • 4.5 Textless S2ST training
  • 5 Results
  • 5.1 Textless S2ST
  • 5.2 Analysis on the speech normalizer
  • 5.3 Analysis of mined data
  • 6 Conclusion
  • Acknowledgements
  • References
  • A mHuBERT Training details
  • B Unit-based Vocoder
  • C Text-to-Unit (T2U)
  • D Hyper-parameters
  • E Dev BLEU

Knowls

  1. Knowl 1 — Self-supervised unit-based speech normalization

    model/method

    The paper normalizes multi-speaker target speech by mapping it to discrete units associated with a single reference speaker. For a random-speaker audio x(r)x^{(r)} and a reference-speaker audio x(ref)x^{(ref)} containing the same linguistic content, a pretrained HuBERT encoder followed by KK-means converts each audio into an orig-unit sequence u=(u1,…,uT)u=(u_1,\ldots,u_T), where ut∈{0,…,K−1}u_t \in \{0,\ldots,K-1\} and TT is the number of acoustic frames. Consecutive duplicate units are removed from the reference sequence to obtain a reduced orig-unit target.

    A CTC model is then obtained by fine-tuning the pretrained HuBERT encoder on x(r)x^{(r)} while predicting the reduced orig-unit sequence extracted from x(ref)x^{(ref)}. At inference time, CTC decoding converts speech from any speaker into a norm-unit sequence. The resulting representation is intended to preserve lexical content while suppressing speaker, accent, silence, and recording-condition variation. The workflow shown in the process diagram on page 3 consists of reference-unit preparation, CTC fine-tuning, and norm-unit extraction.

  2. Knowl 2 — Textless speech-to-unit translation and unit vocoder

    model/method

    The complete textless S2ST system translates source speech into target-language discrete units and then synthesizes target speech, without generating text as an intermediate representation. The speech encoder takes source log-mel filterbanks, applies two one-dimensional convolutional layers with stride 2 and gated-linear-unit activations for an overall downsampling factor of 4, and processes the result with Transformer blocks. A Transformer discrete-unit decoder predicts the target norm-unit sequence using cross-entropy with label smoothing.

    An auxiliary reconstruction task attaches cross-attention and a Transformer decoder to an intermediate encoder layer and predicts the reduced orig-unit sequence of the source speech. This auxiliary task is used to improve convergence, but it is not part of inference. The predicted target units are converted to waveform speech by a separately trained unit-based HiFi-GAN vocoder with an upsampler and a duration predictor. The vocoder combines the HiFi-GAN generator-discriminator objective with mean-squared error on predicted unit durations in the logarithmic domain. The architecture diagram on page 4 shows that only the source encoder, unit decoder, and vocoder are required at inference time.

  3. Knowl 3 — Multilingual HuBERT codebook and speech-normalizer data

    experimental setup

    A single multilingual HuBERT model is trained on 13.5k hours of unlabeled VoxPopuli speech: 4.5k hours each in English, Spanish, and French. The three HuBERT iterations use MFCC targets with 100 clusters, sixth-layer representations from iteration 1 with 500 clusters, and ninth-layer representations from iteration 2 with 500 clusters, respectively. Each iteration uses 400k optimization steps. For the final discrete units used in the experiments, the best configuration uses the eleventh layer of the third iteration and a shared codebook of K=1000K=1000 clusters; no language identifier or language-dependent sampling is used.

    The speech normalizer is trained separately for English, Spanish, and French with 10-minute, 1-hour, or 10-hour paired-data budgets. Reference units are produced by a text-to-unit model trained on single-speaker TTS data, and consecutive duplicates are removed before CTC fine-tuning. The training-set sizes reported in the data table on page 5 are:

    Could not parse LaTeX table

    The paired recordings contain the same content for the input and reference speakers. The authors remove overlap with the S2ST data and supplement VoxPopuli ASR data with randomly sampled Common Voice data when necessary.

  4. Knowl 4 — Real-data language pairs, corpora, and evaluation protocol

    experimental setup

    The experiments cover four translation directions: Spanish–English, French–English, English–Spanish, and English–French. Supervised S2ST training uses VoxPopuli interpretation data, while an additional condition augments it with automatically mined LibriVox speech pairs. Evaluation uses Europarl-ST test data, with CoVoST 2 added for the mined-data experiments. The training and evaluation corpus statistics reported on page 6 are:

    Could not parse LaTeX table

    An asterisk marks target speech synthesized only for development-loss tracking. Translation quality is measured by BLEU after open-source ASR decodes the generated speech; references are lowercased, punctuation is removed, and numbers are converted to spoken forms. Naturalness is measured with MOS from 200 randomly selected utterances per system, each rated by five listeners on a 1–5 scale.

  5. Knowl 5 — Speech normalization improves supervised real-data S2ST

    data/table

    The supervised experiments train on VoxPopuli S2ST data and evaluate on Europarl-ST test sets. The comparison isolates target-speaker embeddings, speech normalization, and text-based S2T+TTS baselines. The table on page 7 reports BLEU and MOS with 95% confidence intervals; higher values are better.

    Could not parse LaTeX table

    With only 10 minutes of paired speech, norm-unit targets improve BLEU by an average of 1.5 points over the orig-unit system with a speaker embedding. Increasing normalizer data to 10 hours yields an average 4.9-point gain over the basic orig-unit system. The best normalized system has naturalness comparable to Transformer-TTS output, while its translation quality is similar to text-based S2T+TTS systems that require ASR-derived text.

  6. Knowl 6 — Mined S2ST data further improves the normalized system

    data/table

    The paper augments VoxPopuli with automatically mined LibriVox speech pairs and evaluates on both Europarl-ST and CoVoST 2. The normalized S2UT model uses the 1-hour speech normalizer. The results on page 7 compare orig-unit and norm-unit S2UT systems with ASR-based S2T+TTS and oracle-text upper bounds.

    Could not parse LaTeX table

    Relative to normalized S2UT trained only on VoxPopuli, adding mined data gives an average 2.0 BLEU gain across the four language directions on Europarl-ST. Relative to the basic orig-unit system trained on VoxPopuli plus mined data, normalization gives an average 3.9 BLEU improvement on Europarl-ST. The mined data produces especially large improvements on CoVoST 2, whose domain is more similar to LibriVox. The normalized system is only 0.6 BLEU behind the ASR-based S2T+TTS systems on average on Europarl-ST.

  7. Knowl 7 — Normalization preserves lexical content while removing non-speech variation

    empirical result

    A speech-resynthesis experiment evaluates whether norm-unit conversion changes linguistic content. Units extracted from VoxPopuli ASR test utterances are passed through the unit vocoder, and the resulting audio is scored by word error rate. The results reported on page 8 are:

    Could not parse LaTeX table

    The 1-hour and 10-hour normalizers produce resynthesis WERs comparable to reduced orig-unit, supporting the claim that the normalization process largely preserves lexical content. Norm-unit sequences are approximately 15% shorter on average than reduced orig-unit sequences because the normalizer does not emit units for long silences, background noise, and other non-speech regions; this produces a shorter and cleaner S2UT training target.

  8. Knowl 8 — Norm-unit representations reduce speaker-dependent variation

    empirical result

    To measure cross-speaker consistency, the authors sample 400 Common Voice audio pairs separately for English, Spanish, and French. Each pair contains two speakers reading the same text prompt. Unit error rate is computed between the two extracted sequences. The comparison on page 8 is:

    Could not parse LaTeX table

    The normalized representation has a cross-speaker UER averaging about 58% of the reduced orig-unit UER. Thus, CTC prediction toward reference-speaker units makes equivalent content more consistent across speakers and languages, which is the intended mechanism behind the S2UT improvements.

  9. Knowl 9 — Mined-data quality and quantity have a threshold trade-off

    empirical result

    Each automatically mined speech pair has a semantic similarity score. The experiments retain pairs above a threshold and train an S2UT model with 1-hour norm-unit targets. On the Spanish-to-English Europarl-ST test set, the authors evaluate thresholds 1.061.06, 1.0651.065, 1.071.07, 1.0751.075, and 1.081.08, as shown in the plot on page 8.

    Mined data improves BLEU over the corresponding VoxPopuli-only system across the tested thresholds, demonstrating that the mined corpus is useful even when filtered for different quality levels. Raising the threshold from 1.061.06 to 1.071.07 reduces BLEU because the stricter filter removes too much training data. The result identifies a quality-versus-quantity trade-off rather than showing that only the highest-scoring mined pairs are useful.

  10. Knowl 10 — Training configuration for the speech normalizer, S2UT model, and vocoder

    experimental setup

    Each language-specific speech normalizer is fine-tuned for 25k updates with the Transformer parameters frozen for the first 10k updates. Optimization uses Adam with β1=0.9\beta_1=0.9, β2=0.98\beta_2=0.98, and ϵ=10−8\epsilon=10^{-8}, followed by 8k warm-up steps and exponential learning-rate decay; learning rates and masking probabilities are selected on development data using unit error rate.

    The S2UT models use 512-dimensional embeddings, 8 attention heads, label smoothing of 0.2, and an auxiliary-task weight of 8.0. Models are trained for 600k steps on VoxPopuli-only data and 800k steps when mined data is included, using Adam with the same optimizer constants, inverse-square-root learning-rate decay, and 10k warm-up steps. The model with the best development BLEU is selected. One unit vocoder is trained per language for 500k updates, using natural orig-unit sequences as input and assigning weight 1.0 to the duration-prediction MSE. The vocoder can synthesize both orig-unit and norm-unit sequences because both use the same HuBERT KK-means codebook.

Coverage note — The appendix-level hyperparameter tables, auxiliary vocoder/T2U WER tables, and development-set BLEU table were not made separate knowls because their details are implementation support or corroboration of the ten higher-significance method, data, and test-set results above.

References

  1. 1.Nagaraj Adiga, Yannis Pantazis, Vassilis Tsiaras, and Yannis Stylianou. 2019. Speech enhancement for noise-robust speech synthesis using wasserstein gan. In INTERSPEECH, pages 1821–1825.
  2. 2.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4218–4222.
  3. 3.Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019. Effectiveness of self-supervised pre-training for speech recognition. arXiv preprint arXiv:1911.03912.
  4. 4.Claudio Bendazzoli, Annalisa Sandrelli, et al. 2005. An approach to corpus-based interpreting studies: developing epic (european parliament interpreting corpus). Proceedings of Challenges of Multidimensional Translation.
  5. 5.Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016. Listen and translate: A proof of concept for end-to-end speech-to-text translation. arXiv preprint arXiv:1612.01744.
  6. 6.Cassia Valentini Botinhao, Xin Wang, Shinji Takaki, and Junichi Yamagishi. 2016. Speech enhancement for a noise-robust text-to-speech synthesis system using deep recurrent neural networks. In Interspeech 2016, pages 352–356.
  7. 7.Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018. Voxceleb2: Deep speaker recognition. arXiv preprint arXiv:1806.05622.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Paul-Ambroise Duquenne, Hongyu Gong, and Holger Schwenk. 2021. Multimodal and multilingual embeddings for large-scale speech mining. Advances in Neural Information Processing Systems, 34.
  10. 10.Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017. Deep voice 2: Multi-speaker neural text-to-speech. Advances in neural information processing systems, 30.
  11. 11.Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369–376.
  12. 12.Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan. 2020. Espnet-tts: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit. In ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 7654–7658. IEEE.
  13. 13.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. arXiv preprint arXiv:2106.07447.
  14. 14.Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Javier Jorge, Nahuel Roselló, Adrià Giménez, Albert Sanchis, Jorge Civera, and Alfons Juan. 2020. Europarl-st: A multilingual corpus for speech translation of parliamentary debates. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8229–8233. IEEE.
  15. 15.Keith Ito and Linda Johnson. 2017. The lj speech dataset. https://keithito.com/LJ-Speech-Dataset/.
  16. 16.Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz. 2021. Translatotron 2: Robust direct speech-to-speech translation. arXiv preprint arXiv:2107.08661.
  17. 17.Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019. Direct speech-to-speech translation with a sequence-to-sequence model. Proc. Interspeech 2019, pages 1123–1127.
  18. 18.Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2021. Transformer-based direct speech-to-speech translation with transcoder. In 2021 IEEE Spoken Language Technology Workshop (SLT), pages 958–965. IEEE.
  19. 19.Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al. 2021. Text-free prosody-aware generative spoken language modeling. arXiv preprint arXiv:2109.03264.
  20. 20.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33.
  21. 21.Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2021. Textless speech emotion conversion using decomposed and discrete representations. arXiv preprint arXiv:2111.07402.
  22. 22.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75.
  23. 23.Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021. Generative spoken language modeling from raw audio. arXiv preprint arXiv:2102.01192.
  24. 24.Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan. 1997. JANUS-III: Speech-to-speech translation in multiple languages. In 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 99–102. IEEE.
  25. 25.Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al. 2021. Direct speech-to-speech translation with discrete units. arXiv preprint arXiv:2107.05604.
  26. 26.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019. Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6706–6713.
  27. 27.Satoshi Nakamura, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, J-S Zhang, Hirofumi Yamamoto, Eiichiro Sumita, and Seiichi Yamamoto. 2006. The ATR multilingual speech-to-speech translation system. IEEE Transactions on Audio, Speech, and Language Processing, 14(2):365–376.
  28. 28.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations.
  29. 29.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an ASR corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210. IEEE.
  30. 30.Kyubyong Park and Thomas Mulc. 2019. Css10: A collection of single speaker speech datasets for 10 languages. arXiv preprint arXiv:1903.11269.
  31. 31.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021. Speech resynthesis from discrete disentangled self-supervised representations. arXiv preprint arXiv:2104.00355.
  32. 32.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191.
  33. 33.Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020. Fastspeech 2: Fast and high-quality end-to-end text to speech. arXiv preprint arXiv:2006.04558.
  34. 34.Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019. Fastspeech: Fast, robust and controllable text to speech. arXiv preprint arXiv:1905.09263.
  35. 35.Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021. CCMatrix: Mining billions of high-quality parallel sentences on the web. In ACL, page 6490–6500.
  36. 36.Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al. 2018. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4779–4783. IEEE.
  37. 37.Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert. 2019. End-to-end ASR: from supervised to semi-supervised learning with modern architectures. arXiv preprint arXiv:1911.08460.
  38. 38.Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2019. Speech-to-speech translation between untranscribed unknown languages. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 593–600. IEEE.
  39. 39.Hitomi Tohyama, Shigeki Matsubara, Koichiro Ryu, N Kawaguch, and Yasuyoshi Inagaki. 2004. Ciair simultaneous interpretation corpus. In Proc. Oriental COCOSDA.
  40. 40.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 6309–6318.
  41. 41.Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez. 2014. Deep neural networks for small footprint text-dependent speaker verification. In 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 4052–4056. IEEE.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  43. 43.Changhan Wang, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Ann Lee, Peng-Jen Chen, Jiatao Gu, and Juan Pino. 2021a. fairseq sˆ 2: A scalable and integrable speech synthesis toolkit. arXiv preprint arXiv:2109.06912.
  44. 44.Changhan Wang, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Ann Lee, Peng-Jen Chen, Jiatao Gu, and Juan Pino. 2021b. fairseq sˆ 2: A scalable and integrable speech synthesis toolkit. arXiv preprint arXiv:2109.06912.
  45. 45.Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021c. VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 993–1003, Online. Association for Computational Linguistics.
  46. 46.Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020a. fairseq s2t: Fast speech-to-text modeling with fairseq. In Proceedings of the 2020 Conference of the Asian Chapter of the Association for Computational Linguistics (AACL): System Demonstrations.
  47. 47.Changhan Wang, Anne Wu, and Juan Pino. 2020b. CoVoST 2: A massively multilingual speech-to-text translation corpus.
  48. 48.Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. 2017. Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135.
  49. 49.Marcely Zanon Boito, William Havard, Mahault Garnerin, Éric Le Ferrand, and Laurent Besacier. 2020. MaSS: A large and clean multilingual corpus of sentence-aligned spoken utterances extracted from the Bible. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 6486–6493, Marseille, France. European Language Resources Association.
  50. 50.Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu. 2020. UWSpeech: Speech to speech translation for unwritten languages. arXiv preprint arXiv:2006.07926.

Citation

MLA
Lee, A., et al. “Textless Speech-to-Speech Translation on Real Data”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 860–72, https://doi.org/10.18653/v1/2022.naacl-main.63.
APA
Lee, A., Gong, H., Duquenne, P.-A., Schwenk, H., Chen, P.-J., Wang, C., Popuri, S., Adi, Y., Pino, J., Gu, J., & Hsu, W.-N. (2022). Textless Speech-to-Speech Translation on Real Data. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 860–872. https://doi.org/10.18653/v1/2022.naacl-main.63
Chicago
Lee, A., H. Gong, P.-A. Duquenne, et al. 2022. “Textless Speech-to-Speech Translation on Real Data”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 860–72. https://doi.org/10.18653/v1/2022.naacl-main.63.
Harvard
Lee, A. et al. (2022) “Textless Speech-to-Speech Translation on Real Data”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 860–872. Available at: https://doi.org/10.18653/v1/2022.naacl-main.63.
Vancouver
1. Lee A, Gong H, Duquenne P-A, et al (2022) Textless Speech-to-Speech Translation on Real Data. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 860–872

BibTeX

@inproceedings{lee-etal-2022-textless,
    title = "Textless Speech-to-Speech Translation on Real Data",
    author = "Lee, Ann  and
      Gong, Hongyu  and
      Duquenne, Paul-Ambroise  and
      Schwenk, Holger  and
      Chen, Peng-Jen  and
      Wang, Changhan  and
      Popuri, Sravya  and
      Adi, Yossi  and
      Pino, Juan  and
      Gu, Jiatao  and
      Hsu, Wei-Ning",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.63/",
    doi = "10.18653/v1/2022.naacl-main.63",
    pages = "860--872"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/