UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

Hirofumi InagumaSravya PopuriIlia KulikovPeng-Jen ChenChanghan WangYu-An ChungYun TangAnn LeeShinji WatanabeJuan Pino

article2023ACL95 citations

Proposes UnitY, a two-pass direct speech-to-speech translation framework that first decodes text subwords and then predicts discrete acoustic units, achieving higher translation quality and nearly triple the decoding speed of single-pass speech-to-unit models.

Listen

Real-time speech translation across languages is essential for global communications and digital platforms. Traditional cascaded systems rely on a chained pipeline of speech recognition, text machine translation, and speech synthesis, which introduces high latency and compounding errors. Direct speech-to-speech translation systems offer a streamlined, jointly optimized alternative but historically suffer from poor translation accuracy due to training data scarcity, as well as high computational overhead during inference.

The article aims to design and evaluate UnitY, a novel two-pass direct speech-to-speech translation architecture that simultaneously improves translation quality and inference efficiency. It demonstrates how generating intermediate subword text representations followed by discrete acoustic units outperforms both traditional cascaded pipelines and single-pass direct models.

The researchers evaluated UnitY across multiple standard benchmark datasets, including Fisher Spanish-to-English (170 hours), CVSS-C multilingual speech-to-English (547 hours across 21 source languages), and a large-scale multi-domain English-Spanish dataset (spanning up to 20,000 hours). The architecture incorporates a deep first-pass text decoder initialized with self-supervised text pre-training, an intermediate text-to-unit encoder, and a shallow second-pass decoder that outputs discrete acoustic speech units before feeding into a neural vocoder.

Across the evaluations, UnitY achieved significant improvements in both translation accuracy and processing speed. On the multilingual CVSS-C benchmark with pre-trained models, UnitY attained an average translation score of 24.5 ASR-BLEU, outperforming single-pass unit baselines by 3.7 points and earlier two-pass spectrogram models by 5.2 to 6.6 points. On multi-domain English-Spanish benchmarks, UnitY outperformed single-pass models by 1.3 to 2.5 points and matched or exceeded strong cascaded baselines. In terms of runtime efficiency, UnitY delivered a 2.83-times speed-up and reduced floating-point operations by roughly 69 percent compared to single-pass unit translation, while also running 2.51 times faster than two-pass spectrogram systems. Human evaluations further confirmed that UnitY consistently scored higher in semantic translation adequacy and acceptable translation rates.

These findings demonstrate that separating the translation task into a deep text decoder pass and a shallow discrete acoustic unit decoder pass resolves the core trade-off between translation quality and operational latency. Utilizing discrete units dramatically compresses target sequence lengths compared to continuous acoustic representations, lowering computational costs and infrastructure requirements for real-time deployment. Furthermore, leveraging abundant unlabeled text to pre-train the first pass mitigates speech data scarcity, providing a practical pathway for deploying high-quality speech translation systems.

Organizations developing speech translation systems should adopt two-pass direct architectures utilizing discrete units rather than spectrograms or single-pass unit models. To maximize efficiency and accuracy, engineering teams should allocate higher model capacity and wider beam search exploration to the first-pass text generation while keeping the second-pass unit generation shallow and focused on single-candidate greedy decoding. Pre-training the text decoder with large-scale unlabeled text is strongly recommended, especially for lower-resource language pairs where parallel speech data is limited.

The findings are bounded by certain operating constraints. UnitY requires an intermediate text representation, meaning it cannot be applied to purely unwritten target languages. Direct systems also introduce extra offline preparation pipelines, such as training speech representation and vocoder models. The authors express strong confidence in the results across multiple domains and language settings, though caution that speech translation models inherently carry risks of generating inaccurate or unfaithful content that requires monitoring in high-stakes operational environments.

Inaguma et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

Abstract

Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and predicts discrete acoustic units subsequently. We enhance the model performance by subword prediction in the first-pass decoder, advanced two-pass decoder architecture design and search strategy, and better training regularization. To leverage large amounts of unlabeled text data, we pre-train the first-pass text decoder based on the self-supervised denoising auto-encoding task. Experimental evaluations on benchmark datasets at various data scales demonstrate that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83× decoding speed-up. We show that the proposed methods boost the performance even when predicting spectrogram in the second pass. However, predicting discrete units achieves 2.51× decoding speed-up compared to that case.

Table of Contents

  • 1 Introduction
  • 2 UnitY
  • 2.1 Architecture
  • 2.2 Text decoder pre-training
  • 2.3 Search algorithm
  • 2.4 Deep-shallow two-pass decoders
  • 3 Experimental setting
  • 3.1 Data
  • 3.2 Pre-processing
  • 3.3 Pre-training
  • 3.4 Baseline
  • 3.5 Architecture
  • 3.6 Training
  • 3.7 Decoding
  • 3.8 Vocoder
  • 3.9 Evaluation
  • 4 Experimental results
  • 4.1 CVSS-C
  • 4.2 Multi-domain En ↔ Es
  • 4.3 Decoding efficiency
  • 4.4 Fisher
  • 5 Analysis
  • 5.1 Ablation study
  • 5.2 Output unit for first-pass decoder
  • 5.3 Capacity assignment to two-pass decoders
  • 5.4 Data scale
  • 6 Related works
  • 7 Conclusion
  • 8 Limitation
  • Acknowledgement
  • References
  • A Pseudo algorithm for two-pass beam search decoding
  • B Training with R-Drop
  • C Training objective
  • D Data
  • E Pre-processing
  • F Pre-training
  • G Architecture details
  • H Training details
  • I Additional experimental results
  • I.1 Human evaluation

Knowls

  1. Knowl 1 — UnitY’s two-pass text-to-unit architecture

    model/method

    UnitY directly translates source speech XX into target speech through two autoregressive prediction passes. A Conformer speech encoder produces speech representations HH. A deep Transformer text decoder attends to HH and predicts target-language subwords Y=(y1,…,yM)Y=(y_1,\ldots,y_M); its continuous states DtextD^{\text{text}} retain acoustic and contextual information. A bidirectional text-to-unit (T2U) Transformer encoder maps those states to ZZ without changing their sequence length. A shallow Transformer unit decoder then predicts target-language discrete acoustic units U=(u1,…,uL)U=(u_1,\ldots,u_L) from ZZ and preceding units, without directly attending to HH. Consecutive duplicate units are collapsed, so the unit sequence has no per-unit duration information. A separately trained unit vocoder converts the predicted units to a waveform. UnitY trains both passes jointly with Ltotal=Ls2u(U∣X,Y)+ws2tLs2t(Y∣X)L_{\mathrm{total}}=L_{\mathrm{s2u}}(U\mid X,Y)+w_{\mathrm{s2t}}L_{\mathrm{s2t}}(Y\mid X), where Ls2uL_{\mathrm{s2u}} and Ls2tL_{\mathrm{s2t}} are the average token negative log-likelihoods for unit prediction and text translation, respectively, and ws2tw_{\mathrm{s2t}} weights the text-translation loss. At inference, the system beam-searches for text first, passes the selected text decoder states through the T2U encoder, then searches for units and vocodes them. The reported decoding setup used beam size 10 for text and 1 for units; the unit decoder therefore used greedy search.

  2. Knowl 2 — Text-decoder pretraining with unlabeled text

    model/method

    UnitY initializes its first-pass text decoder with text-based multilingual BART (t-mBART), pretrained on unlabeled text using denoising autoencoding. During speech-to-speech fine-tuning, the text decoder’s feed-forward network parameters are frozen; the T2U encoder and unit decoder are initialized randomly. For English–Spanish experiments, the text pretraining used a 65k-subword vocabulary, while multilingual CVSS-C experiments used a multilingual mBART model with a 250k-subword vocabulary. In a multi-domain Spanish-to-English development-set comparison, random initialization yielded 34.8 text BLEU and 30.7 speech ASR-BLEU, whereas t-mBART initialization yielded 38.3 and 33.2. Unsupervised MT initialization was similar at 38.2 and 33.2; supervised MT initializations yielded 36.6/33.0 and 37.5/33.3, and initialization from a separate S2TT model yielded 37.8/32.5. These comparisons indicate that the tested text-pretrained initializations were more effective than the tested supervised MT or separate-S2TT initializations.

  3. Knowl 3 — R-Drop regularization for the text and unit decoders

    model/method

    UnitY applies R-Drop to both decoders that predict discrete symbols. During training, each input is duplicated and processed with different dropout masks; the training objective includes a consistency penalty based on the symmetric Kullback–Leibler divergence between the two output probability distributions. This regularizes the text and unit predictions while retaining the joint two-pass training objective. The paper applies R-Drop to discrete-symbol prediction tasks, but not to the MT task.

  4. Knowl 4 — Translation-quality results on CVSS-C and Fisher

    empirical result

    On CVSS-C, a 547-hour multilingual corpus combining 21 source-language directions into English, UnitY achieved an average ASR-BLEU of 12.0 when trained from scratch, compared with 9.1 for single-pass S2UT. With S2TT pretraining, UnitY scored 13.0 versus 11.4 for S2UT. With multilingual w2v-BERT speech-encoder pretraining and text-based mBART for UnitY or unit-based mBART for S2UT, the scores were 24.5 and 20.8, respectively. On the 170-hour Fisher Spanish-to-English corpus, the reported three-run-average test ASR-BLEU for models trained from scratch was 51.4 for UnitY, 47.4 for S2UT, and 50.8 for S2SpecT2, the paper’s improved speech-to-spectrogram system. With wav2vec 2.0 speech-encoder pretraining, UnitY scored 55.9 and S2UT scored 53.7, but S2SpecT2 scored 58.6. Thus, UnitY outperformed S2UT in these comparisons, while the Fisher result with wav2vec 2.0 shows that UnitY did not outperform every two-pass alternative under every training condition.

  5. Knowl 5 — Results on high-resource multi-domain English–Spanish data

    empirical result

    On the multi-domain English–Spanish benchmark, all compared models used wav2vec 2.0-pretrained speech encoders. Average ASR-BLEU for UnitY with a six-layer text decoder and six-layer unit decoder was 34.4 for English-to-Spanish and 32.5 for Spanish-to-English. Initializing the text decoder with t-mBART and using a 12-layer text decoder with a two-layer unit decoder raised these averages to 34.7 and 33.9. The corresponding S2UT system with unit-based mBART scored 33.4 and 31.4. UnitY with t-mBART improved over that S2UT system by 1.3 and 2.5 ASR-BLEU, respectively. UnitY was not uniformly better than the speech-to-spectrogram alternative: S2SpecT2 with t-mBART scored 35.6 versus 34.7 for UnitY on English-to-Spanish, while UnitY scored 33.9 versus 30.9 on Spanish-to-English. UnitY also exceeded the reported cascaded ASR→MT→TTS system on the MuST-C English-to-Spanish subset, scoring 34.1 versus 30.8 ASR-BLEU.

  6. Knowl 6 — Decoding and computational efficiency

    empirical result

    The paper measured decoding on an Intel Xeon Gold 6230 CPU using 500 randomly sampled utterances from the multi-domain Spanish-to-English development set, preserving the domain proportions and including vocoder inference. With two-pass beam search using a text beam of 10 and a unit beam of 1, UnitY decoded 2.51 times faster than S2SpecT2 and 2.83 times faster than single-pass S2UT. UnitY also used 1.65 times fewer floating-point operations than S2SpecT2 and 3.19 times fewer than S2UT. In beam-width experiments, increasing the text-pass beam up to 10 improved translation quality, while reducing the unit-pass beam to 1 did not significantly degrade it; the authors attribute the opportunity for this allocation to the first pass posing the greater modeling challenge.

  7. Knowl 7 — Ablations identify the role of the T2U encoder and regularization

    empirical result

    On the multi-domain Spanish-to-English development set, UnitY with t-mBART scored 38.3 text BLEU and 33.2 speech ASR-BLEU. Removing the T2U encoder changed the scores to 38.1 and 30.7, respectively, indicating a larger loss for generated speech than for text. Removing R-Drop yielded 37.7 text BLEU and 32.1 speech ASR-BLEU. Adding another unit-decoder cross-attention to the speech encoder did not improve speech ASR-BLEU: sequential and parallel variants scored 33.2 and 33.1. The authors report that an analogous representation-bridging encoder was particularly important for the speech-to-spectrogram model: removing its text-to-spectrogram encoder changed that model’s multi-domain development scores from 35.0 text BLEU and 30.8 speech ASR-BLEU to 34.9 and 25.0. An auxiliary CTC objective on UnitY’s unit decoder did not improve the reported Fisher results.

  8. Knowl 8 — Subword targets and deep–shallow decoder allocation

    empirical result

    Experiments support using subwords for UnitY’s first-pass output and allocating more decoder layers to text than to units. With randomly initialized first-pass decoders on the multi-domain Spanish-to-English development set, UnitY using phonemes, characters, and 2k subwords scored 27.8, 29.6, and 30.1 speech ASR-BLEU, respectively; their decoding speed-ups relative to the phoneme S2SpecT2 system were 2.31×, 2.06×, and 2.86×. For UnitY, the corresponding text BLEU scores for characters and subwords were 33.2 and 34.1. In decoder-depth experiments, a 12-layer text decoder and two-layer unit decoder gave 34.9 text BLEU and 30.7 speech ASR-BLEU with random initialization, and a 1.44× speed-up relative to the six-layer/six-layer configuration. With t-mBART initialization, the 12-layer/two-layer design scored 38.3 and 33.2, with a 1.19× speed-up relative to the six-layer/six-layer configuration. Increasing the unit decoder beyond two layers did not improve the reported t-mBART-initialized results.

  9. Knowl 9 — Audio-only human evaluation of translation quality

    empirical result

    An audio-only human evaluation used 989 mTEDx Spanish-to-English test samples and compared system-generated target audio with the source speech. Three bilingual annotators assigned each item a 1–5 cross-lingual semantic textual similarity score, and the median score per item was used. A score of at least 3 counted as acceptable. The plotted mean scores were 3.904 for the cascaded S2TT→TTS system, 4.017 for S2UT, and 4.197 for UnitY; the respective percentages of acceptable translations were 71.11%, 82.12%, and 92.94%. These results favor UnitY on both reported measures. The authors note that these evaluated systems were early versions and differed slightly from the systems in the main multi-domain benchmark.

  10. Knowl 10 — Limitations and deployment risks

    limitation

    Two-pass direct S2ST requires a linguistic target for its first pass, so it cannot be applied as designed to target languages without a writing system. Compared with cascaded S2ST, preparing a direct speech-to-unit system entails additional steps, including training a HuBERT model, generating or extracting target units, and training a unit vocoder. The quality of its generated audio depends on the quality of the discrete units and on the availability of target-language speech for training the self-supervised unit model. The paper also notes a risk specific to speech translation: a generated utterance can fail to preserve the source content and thereby convey incorrect information.

Coverage note — No substantial contributed material was deliberately omitted. Exhaustive per-language score matrices, supplementary ASR-chrF tables that largely confirm the reported ASR-BLEU trends, and implementation-level hyperparameter inventories are omitted as supporting detail rather than distinct contributions.

References

  1. 1.Antonios Anastasopoulos, Ondřej Bojar, Jacob Bremerman, Roldano Cattoni, Maha Elbayad, Marcello Federico, Xutai Ma, Satoshi Nakamura, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Alexander Waibel, Changhan Wang, and Matthew Wiesner. 2021. FINDINGS OF THE IWSLT 2021 EVALUATION CAMPAIGN. In Proceedings of IWSLT, pages 1–29.
  2. 2.Antonios Anastasopoulos and David Chiang. 2018. Tied multitask learning for neural speech translation. In Proceedings of NAACL-HLT, pages 82–91.
  3. 3.Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, Fei Huang, Kevin Knight, Xutai Ma, Ajay Nagesh, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Xing Shi, Sebastian Stüker, Marco Turchi, Alexander Waibel, and Changhan Wang. 2020. FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN. In Proceedings of IWSLT, pages 1–34.
  4. 4.Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020. Common Voice: A massively-multilingual speech corpus. In Proceedings of LREC, pages 4218–4222.
  5. 5.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al. 2019. Massively multilingual neural machine translation in the wild: Findings and challenges. arXiv preprint arXiv:1907.05019.
  6. 6.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of NeurIPS, volume 33, pages 12449–12460.
  7. 7.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of ICLR.
  8. 8.Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio. 2016. End-to-end attention-based large vocabulary speech recognition. In Proceedings of ICASSP, pages 4945–4949.
  9. 9.Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau. 2022. mSLAM: Massively multilingual joint pre-training for speech and text. arXiv preprint arXiv:2202.01374.
  10. 10.Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin. 2018. End-to-end automatic speech translation of audiobooks. In Proceedings of ICASSP, pages 6224–6228. IEEE.
  11. 11.Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier. 2016. Listen and translate: A proof of concept for end-to-end speech-to-text translation. In Proceedings of NIPS 2016 End-to-end Learning for Speech and Audio Processing Workshop.
  12. 12.William Chan, Daniel Park, Chris Lee, Yu Zhang, Quoc Le, and Mohammad Norouzi. 2021. Speechstew: Simply mix all available speech recognition data to train one large neural network. arXiv preprint arXiv:2104.02133.
  13. 13.Colin Cherry, George Foster, Ankur Bapna, Orhan Firat, and Wolfgang Macherey. 2018. Revisiting character-based neural machine translation with capacity and compression. In Proceedings of EMNLP, pages 4295–4305.
  14. 14.Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078.
  15. 15.Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021. w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In Proceedings of ASRU.
  16. 16.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of ACL, pages 8440–8451.
  17. 17.Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze, and Shinji Watanabe. 2021. Searchable hidden intermediates for end-to-end models of decomposable sequence tasks. In Proceedings of NAACL-HLT, pages 1882–1896.
  18. 18.Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019. MuST-C: a Multilingual Speech Translation Corpus. In Proceedings of NAACL-HLT, pages 2012–2017.
  19. 19.Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, and Yu Zhang. 2022. Leveraging pseudo-labeled data to improve direct speech-to-speech translation. arXiv preprint arXiv:2205.08993.
  20. 20.Salesky Elizabeth, Wiesner Matthew, Bremerman Jacob, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W Oard, and Post Matt. 2021. The multilingual TEDx corpus for speech recognition and translation. In Proceedings of Interspeech, pages 3655–3659.
  21. 21.Mark JF Gales, Kate M Knill, Anton Ragni, and Shakti P Rath. 2014. Speech recognition and keyword spotting for low-resource languages: Babel project research at CUED. In Proceedings of SLTU, pages 16–23.
  22. 22.Thamme Gowda and Jonathan May. 2020. Finding the optimal vocabulary size for neural machine translation. In Findings of EMNLP, pages 3955–3964.
  23. 23.Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of ICML, pages 369–376.
  24. 24.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. Conformer: Convolution-augmented Transformer for speech recognition. In Proceedings of Interspeech, pages 5036–5040.
  25. 25.Mary Harper et al. IARPA Babel Program. https://www.iarpa.gov/research-programs/babel. [Online].
  26. 26.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural Computation, 9(8):1735–1780.
  27. 27.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460.
  28. 28.Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar. 2020. Deliberation model based two-pass end-to-end speech recognition. In Proceedings of ICASSP, pages 7799–7803.
  29. 29.Rongjie Huang, Zhou Zhao, Jinglin Liu, Huadai Liu, Yi Ren, Lichao Zhang, and Jinzheng He. 2022. TranSpeech: Speech-to-speech translation with bilateral perturbation. arXiv preprint arXiv:2205.12523.
  30. 30.Hirofumi Inaguma, Siddharth Dalmia, Brian Yan, and Shinji Watanabe. 2021a. Fast-MD: Fast multi-decoder end-to-end speech translation with non-autoregressive hidden intermediates. In Proceedings of ASRU, pages 922–929.
  31. 31.Hirofumi Inaguma, Tatsuya Kawahara, and Shinji Watanabe. 2021b. Source and target bidirectional knowledge distillation for end-to-end speech translation. In Proceedings of NAACL-HLT, pages 1872–1881.
  32. 32.Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerda, Javier Jorge, Nahuel Roselló, Adria Giménez, Albert Sanchis, Jorge Civera, and Alfons Juan. 2020. Europarl-ST: A multilingual corpus for speech translation of parliamentary debates. In Proceedings of ICASSP, pages 8229–8233.
  33. 33.Keith Ito and Linda Johnson. 2017. The lj speech dataset. https://keithito.com/LJ-Speech-Dataset/.
  34. 34.Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry, Yu Zhang, Alexis Conneau, and Nobuyuki Morioka. 2022a. Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation. In Proceedings of Interspeech, pages 1721–1725.
  35. 35.Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019a. Leveraging weakly supervised data to improve end-to-end speech-to-text translation. In Proceedings of ICASSP, pages 7180–7184.
  36. 36.Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz. 2022b. Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. In Proceedings of ICML.
  37. 37.Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, and Heiga Zen. 2022c. CVSS corpus and massively multilingual speech-to-speech translation. In Proceedings of LREC, pages 6691–6703.
  38. 38.Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019b. Direct speech-to-speech translation with a sequence-to-sequence model. In Proceedings of Interspeech, pages 1123–1127.
  39. 39.Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. 2020. Libri-Light: A benchmark for asr with limited or no supervision. In Proceedings of ICASSP, pages 7669–7673.
  40. 40.Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura. 2021. Transformer-based direct speech-to-speech translation with transcoder. In Proceedings of SLT, pages 958–965.
  41. 41.Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah Smith. 2021. Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation. In Proceedings of ICLR.
  42. 42.Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In Proceedings of Machine Translation Summit X: Papers, pages 79–86.
  43. 43.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of NeurIPS, volume 33, pages 17022–17033.
  44. 44.Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of ACL, pages 66–75.
  45. 45.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of EMNLP: System Demonstrations, pages 66–71.
  46. 46.Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan. 1997. JANUS-III: Speech-to-speech translation in multiple languages. In Proceedings of ICASSP, pages 99–102.
  47. 47.Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al. 2022a. Direct speech-to-speech translation with discrete units. In Proceedings of ACL, pages 3327–3339.
  48. 48.Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Juan Pino, Jiatao Gu, and Wei-Ning Hsu. 2022b. Textless speech-to-speech translation on real data. In Proceedings of NAACL-HLT, pages 860–872.
  49. 49.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019. Neural speech synthesis with Transformer network. In Proceedings of AAAI, volume 33, pages 6706–6713.
  50. 50.Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021. Multilingual speech translation from efficient finetuning of pretrained models. In Proceedings of ACL, pages 827–838.
  51. 51.Xinjian Li, Ye Jia, and Chung-Cheng Chiu. 2022. Textless direct speech-to-speech translation with discrete speech representation. arXiv preprint arXiv:2211.00115.
  52. 52.Daniel Licht, Cynthia Gao, Janice Lam, Francisco Guzman, Mona Diab, and Philipp Koehn. 2022. Consistent human evaluation of machine translation across language pairs. In Proceedings of AMTA, pages 309–321.
  53. 53.Pierre Lison, Jörg Tiedemann, and Milen Kouylekov. 2018. OpenSubtitles2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora. In Proceedings of LREC.
  54. 54.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  55. 55.Yuchen Liu, Hao Xiong, Zhongjun He, Jiajun Zhang, Hua Wu, Haifeng Wang, and Chengqing Zong. 2019. End-to-end speech translation with knowledge distillation. In Proceedings of Interspeech, pages 1128–1132.
  56. 56.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed precision training. In Proceedings of ICLR.
  57. 57.Mehryar Mohri, Fernando Pereira, and Michael Riley. 2002. Weighted finite-state transducers in speech recognition. Computer Speech & Language, 16(1):69–88.
  58. 58.Satoshi Nakamura, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, J-S Zhang, Hirofumi Yamamoto, Eiichiro Sumita, and Seiichi Yamamoto. 2006. The ATR multilingual speech-to-speech translation system. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 14(2):365–376.
  59. 59.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038.
  60. 60.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In Proceedings of ICASSP, pages 5206–5210.
  61. 61.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of ACL, pages 311–318.
  62. 62.Kyubyong Park and Thomas Mulc. 2019. CSS10: A collection of single speaker speech datasets for 10 languages. In Proceedings of Interspeech, pages 1566–1570.
  63. 63.Juan Pino, Qiantong Xu, Xutai Ma, Mohammad Javad Dousti, and Yun Tang. 2020. Self-training for end-to-end speech translation. In Proceedings of Interspeech, pages 1476–1480.
  64. 64.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021. Speech resynthesis from discrete disentangled self-supervised representations. In Proceedings of Interspeech, pages 3615–3619.
  65. 65.Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. 2022. Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. In Proceedings of Interspeech, pages 5195–5199.
  66. 66.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191.
  67. 67.Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur. 2013. Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus. In Proceedings of IWSLT.
  68. 68.Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. 2020. MLS: A large-scale multilingual dataset for speech research. In Proceedings of Interspeech, pages 2757–2761.
  69. 69.Nils Reimers and Iryna Gurevych. 2020. Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of EMNLP, pages 4512–4525.
  70. 70.Anthony Rousseau, Paul Deléglise, and Yannick Estève. 2012. TED-LIUM: An automatic speech recognition dedicated corpus. In Proceedings of LREC, pages 125–129.
  71. 71.Tara N Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, et al. 2019. Two-pass end-to-end speech recognition. In Proceedings of Interspeech, pages 2773–2777.
  72. 72.Elizabeth Salesky, Julian Mäder, and Severin Klinger. 2021. Assessing evaluation metrics for speech-to-speech translation. In Proceedings of ASRU, pages 733–740.
  73. 73.Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, and Angela Fan. 2021. CCMatrix: Mining billions of high-quality parallel sentences on the web. In Proceedings of ACL, pages 6490–6500.
  74. 74.Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu. 2020. Non-Attentive Tacotron: Robust and controllable neural TTS synthesis including unsupervised duration modeling. arXiv preprint arXiv:2010.04301.
  75. 75.Raivis Skadin,š, Jörg Tiedemann, Roberts Rozis, and Daiga Deksne. 2014. Billions of parallel words for free: Building and using the EU bookshop corpus. In Proceedings of LREC, pages 1850–1855.
  76. 76.Matthias Sperber, Graham Neubig, Jan Niehues, and Alex Waibel. 2019. Attention-passing models for robust and data-efficient end-to-end speech translation. Transactions of the Association for Computational Linguistics, 7:313–325.
  77. 77.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958.
  78. 78.Tzu-Wei Sung, Jun-You Liu, Hung-yi Lee, and Lin-shan Lee. 2019. Towards end-to-end speech-to-text translation with two-pass decoding. In Proceedings of ICASSP, pages 7175–7179.
  79. 79.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of NIPS, volume 27.
  80. 80.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of CVPR, pages 2818–2826.
  81. 81.Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, and Juan Pino. 2022. Unified speech-text pre-training for speech translation and recognition. In Proceedings of ACL, pages 1488–1499.
  82. 82.Yun Tang, Juan Pino, Xian Li, Changhan Wang, and Dmitriy Genzel. 2021. Improving speech translation by understanding and learning from the auxiliary text translation task. In Proceedings of ACL, pages 4252–4261.
  83. 83.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. Multilingual translation with extensible multilingual pretraining and finetuning. arXiv preprint arXiv:2008.00401.
  84. 84.Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2019. Speech-to-speech translation between untranscribed unknown languages. In Proceedings of ASRU, pages 593–600.
  85. 85.Jörgen Valk and Tanel Alumäe. 2021. VoxLingua107: a dataset for spoken language recognition. In Proceedings of SLT, pages 652–658.
  86. 86.Aaron Van Den Oord, Oriol Vinyals, et al. 2017. Neural discrete representation learning. In Proceedings of NIPS, volume 30.
  87. 87.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of NIPS, volume 30.
  88. 88.Wolfgang Wahlster. 2013. Verbmobil: foundations of speech-to-speech translation. Springer Science & Business Media.
  89. 89.Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, and Juan Pino. 2022. Simple and effective unsupervised speech translation. arXiv preprint arXiv:2210.10191.
  90. 90.Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021a. VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of ACL, pages 993–1003.
  91. 91.Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020. Fairseq S2T: Fast speech-to-text modeling with Fairseq. In Proceedings of AACL: System Demonstrations, pages 33–39.
  92. 92.Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino. 2021b. CoVoST 2 and massively multilingual speech translation. In Proceedings of Interspeech, pages 2247–2251.
  93. 93.Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski, Michael Auli, and Alexis Conneau. 2021c. Large-scale self- and semi-supervised learning for speech translation. In Proceedings of Interspeech, pages 2242–2246.
  94. 94.Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen. 2017. Sequence-to-sequence models can directly translate foreign speech. In Proceedings of Interspeech, pages 2625–2629.
  95. 95.Krzysztof Wołk and Krzysztof Marasek. 2014. Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs. Procedia Technology, 18:126–132.
  96. 96.Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, Tie-Yan Liu, et al. 2021. R-Drop: Regularized dropout for neural networks. In Proceedings off NeurIPS, volume 34, pages 10890–10905.
  97. 97.Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu. 2017. Deliberation networks: Sequence generation beyond one-pass decoding. In Proceedings of NIPS, volume 30.
  98. 98.Brian Yan, Patrick Fernandes, Siddharth Dalmia, Jiatong Shi, Yifan Peng, Dan Berrebbi, Xinyi Wang, Graham Neubig, and Shinji Watanabe. 2022. CMU’s IWSLT 2022 dialect speech translation system. In Proceedings of IWSLT, pages 298–307.
  99. 99.Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu. 2021. Uwspeech: Speech to speech translation for unwritten languages. In Proceedings of AAAI, pages 14319–14327.
  100. 100.Ding Zhao, Tara N. Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang. 2019. Shallow-fusion end-to-end contextual biasing. In Proceedings of Interspeech, pages 1418–1422.
  101. 101.Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tieyan Liu. 2019. Incorporating BERT into neural machine translation. In Proceedings of ICLR.
  102. 102.Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The United Nations parallel corpus v1.0. In Proceedings of LREC, pages 3530–3534.

Citation

MLA
Inaguma, H., et al. “UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15655–80, https://doi.org/10.18653/v1/2023.acl-long.872.
APA
Inaguma, H., Popuri, S., Kulikov, I., Chen, P.-J., Wang, C., Chung, Y.-A., Tang, Y., Lee, A., Watanabe, S., & Pino, J. (2023). UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15655–15680. https://doi.org/10.18653/v1/2023.acl-long.872
Chicago
Inaguma, H., S. Popuri, I. Kulikov, et al. 2023. “UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15655–80. https://doi.org/10.18653/v1/2023.acl-long.872.
Harvard
Inaguma, H. et al. (2023) “UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15655–15680. Available at: https://doi.org/10.18653/v1/2023.acl-long.872.
Vancouver
1. Inaguma H, Popuri S, Kulikov I, Chen P-J, Wang C, Chung Y-A, Tang Y, Lee A, Watanabe S, Pino J (2023) UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15655–15680

BibTeX

@inproceedings{inaguma-etal-2023-unity,
    title = "{U}nit{Y}: Two-pass Direct Speech-to-speech Translation with Discrete Units",
    author = "Inaguma, Hirofumi  and
      Popuri, Sravya  and
      Kulikov, Ilia  and
      Chen, Peng-Jen  and
      Wang, Changhan  and
      Chung, Yu-An  and
      Tang, Yun  and
      Lee, Ann  and
      Watanabe, Shinji  and
      Pino, Juan",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.872/",
    doi = "10.18653/v1/2023.acl-long.872",
    pages = "15655--15680"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/