Textless Speech-to-Speech Translation on Real Data
Ann LeeHongyu GongPaul-Ambroise DuquenneHolger SchwenkPeng-Jen ChenChanghan WangSravya PopuriYossi AdiJuan Miguel PinoJiatao Gu
Presents a textless speech-to-speech translation system that trains directly on real-world multi-speaker audio by introducing a unit-based normalization technique that removes speaker and accent variations using only ten minutes of paired speech data.
More than 40 percent of the world's spoken languages lack standard written forms, creating a major barrier for conventional speech translation systems that depend on intermediate text transcriptions and written data. While direct speech-to-speech translation systems can bypass text, existing models struggle when trained on real-world multi-speaker audio because natural variations in accents, speaking styles, and background noise distort translation targets.
The article demonstrates a direct, textless speech-to-speech translation system capable of training on real-world multi-speaker audio without relying on any text or phoneme annotations. The core objective is to evaluate whether a self-supervised speech normalization method can effectively remove multi-speaker acoustic variations while preserving the core spoken content across multiple language pairs.
To accomplish this, the authors developed a speech normalizer by fine-tuning a pre-trained multilingual speech encoder using a small amount of paired audio from multiple speakers and a single reference speaker. This component transforms diverse multi-speaker speech into standardized discrete acoustic units. The authors trained a speech-to-unit translation model across four language pairs involving English, Spanish, and French using real parliamentary speech from the VoxPopuli dataset and automatically mined audio data from LibriVox, using a specialized unit-based synthesizer to generate the final audio output.
The findings show that unit-based speech normalization significantly improves translation quality and output clarity. Using only 10 minutes of paired audio to train the normalizer yielded an average gain of 3.2 BLEU translation points over unnormalized baselines, rising to an average improvement of 4.9 BLEU points when using 10 hours of normalization data. Integrating automatically mined speech data provided an additional 2.0 BLEU gain across evaluated language directions. Furthermore, human evaluation showed substantial gains in audio naturalness, increasing mean opinion scores by an average of 0.85 points, effectively eliminating speech artifacts like stuttering and matching the performance of text-based cascaded systems.
These results indicate that organizations can build high-quality translation systems for unwritten or low-resource spoken languages without expensive text transcription pipelines or manual data labeling. Relying on self-supervised representations and directly mined speech drastically lowers data engineering costs and accelerates deployment timelines for cross-lingual spoken communication tools.
Organizations developing speech translation capabilities should adopt discrete unit normalization pipelines to leverage unannotated and mined audio datasets. To maximize efficiency, teams can train the normalizer using approximately one hour of paired single-speaker reference audio, as the performance gains between 1 hour and 10 hours of normalization data are minimal. Future work should focus on exploring additional self-supervised pre-training strategies across wider varieties of language families.
Confidence in these findings is high for European parliamentary and audiobook speech domains across the tested languages. However, decision-makers should note that reference units in these experiments were synthetically assisted during normalizer training, and performance across heavily out-of-domain conversational data or highly under-resourced languages remains to be fully verified.
- Paper: HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, Wei-Ning Hsu et al. (2021). Introduces HuBERT and the masked prediction of discrete hidden units that provide the core self-supervised speech representation framework adapted for speech normalization and textless modeling.
- Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations, Alexei Baevski et al. (2020). Pioneers self-supervised representation learning directly from unannotated speech waveforms, laying the groundwork for unit extraction and pre-trained speech encoders used in textless translation.
- Paper: WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, Sanyuan Chen et al. (2021). Establishes large-scale self-supervised speech pre-training across diverse audio domains, underpinning modern unit-based encoder finetuning on multi-speaker datasets.
- Paper: Conformer: Convolution-augmented Transformer for Speech Recognition, Anmol Gulati et al. (2020). Provides the foundational Conformer architecture combining convolutions and self-attention, which is widely adopted as the acoustic encoder backbone in end-to-end speech processing.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Demonstrates sequence-to-sequence multilingual pre-training and denoising techniques that motivate learning cross-lingual representations across multiple language pairs.
- Paper: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale, Matthew Le et al. (2023). Extends discrete and continuous audio modeling principles to universal multilingual speech generation and in-context voice editing at massive scale.
- Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, Chengyi Wang et al. (2023). Applies discrete neural audio units within a language modeling paradigm to enable zero-shot multi-speaker speech synthesis without conventional intermediate acoustic features.
- Paper: SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words, Junyi Ao et al. (2024). Evaluates end-to-end spoken language systems beyond text transcripts by benchmarking multi-speaker acoustic characteristics such as accents and emotions.
