UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
Hirofumi InagumaSravya PopuriIlia KulikovPeng-Jen ChenChanghan WangYu-An ChungYun TangAnn LeeShinji WatanabeJuan Pino
Proposes UnitY, a two-pass direct speech-to-speech translation framework that first decodes text subwords and then predicts discrete acoustic units, achieving higher translation quality and nearly triple the decoding speed of single-pass speech-to-unit models.
Real-time speech translation across languages is essential for global communications and digital platforms. Traditional cascaded systems rely on a chained pipeline of speech recognition, text machine translation, and speech synthesis, which introduces high latency and compounding errors. Direct speech-to-speech translation systems offer a streamlined, jointly optimized alternative but historically suffer from poor translation accuracy due to training data scarcity, as well as high computational overhead during inference.
The article aims to design and evaluate UnitY, a novel two-pass direct speech-to-speech translation architecture that simultaneously improves translation quality and inference efficiency. It demonstrates how generating intermediate subword text representations followed by discrete acoustic units outperforms both traditional cascaded pipelines and single-pass direct models.
The researchers evaluated UnitY across multiple standard benchmark datasets, including Fisher Spanish-to-English (170 hours), CVSS-C multilingual speech-to-English (547 hours across 21 source languages), and a large-scale multi-domain English-Spanish dataset (spanning up to 20,000 hours). The architecture incorporates a deep first-pass text decoder initialized with self-supervised text pre-training, an intermediate text-to-unit encoder, and a shallow second-pass decoder that outputs discrete acoustic speech units before feeding into a neural vocoder.
Across the evaluations, UnitY achieved significant improvements in both translation accuracy and processing speed. On the multilingual CVSS-C benchmark with pre-trained models, UnitY attained an average translation score of 24.5 ASR-BLEU, outperforming single-pass unit baselines by 3.7 points and earlier two-pass spectrogram models by 5.2 to 6.6 points. On multi-domain English-Spanish benchmarks, UnitY outperformed single-pass models by 1.3 to 2.5 points and matched or exceeded strong cascaded baselines. In terms of runtime efficiency, UnitY delivered a 2.83-times speed-up and reduced floating-point operations by roughly 69 percent compared to single-pass unit translation, while also running 2.51 times faster than two-pass spectrogram systems. Human evaluations further confirmed that UnitY consistently scored higher in semantic translation adequacy and acceptable translation rates.
These findings demonstrate that separating the translation task into a deep text decoder pass and a shallow discrete acoustic unit decoder pass resolves the core trade-off between translation quality and operational latency. Utilizing discrete units dramatically compresses target sequence lengths compared to continuous acoustic representations, lowering computational costs and infrastructure requirements for real-time deployment. Furthermore, leveraging abundant unlabeled text to pre-train the first pass mitigates speech data scarcity, providing a practical pathway for deploying high-quality speech translation systems.
Organizations developing speech translation systems should adopt two-pass direct architectures utilizing discrete units rather than spectrograms or single-pass unit models. To maximize efficiency and accuracy, engineering teams should allocate higher model capacity and wider beam search exploration to the first-pass text generation while keeping the second-pass unit generation shallow and focused on single-candidate greedy decoding. Pre-training the text decoder with large-scale unlabeled text is strongly recommended, especially for lower-resource language pairs where parallel speech data is limited.
The findings are bounded by certain operating constraints. UnitY requires an intermediate text representation, meaning it cannot be applied to purely unwritten target languages. Direct systems also introduce extra offline preparation pipelines, such as training speech representation and vocoder models. The authors express strong confidence in the results across multiple domains and language settings, though caution that speech translation models inherently carry risks of generating inaccurate or unfaithful content that requires monitoring in high-stakes operational environments.
- Paper: Textless Speech-to-Speech Translation on Real Data, Ann Lee et al. (2022). This earlier real-data textless speech-to-speech system establishes the discrete-unit translation baseline that UnitY’s two-pass architecture develops beyond.
No sufficiently relevant recommendations were found.
