Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
Matthew LeApoorv VyasBowen ShiBrian KarrerLeda SariRashel MoritzMary WilliamsonVimal ManoharYossi AdiJay Mahadeokar
Introduces Voicebox, a non-autoregressive flow-matching speech model trained on 50,000 hours of audio that performs zero-shot text-to-speech, cross-lingual synthesis, and audio editing up to twenty times faster and with higher intelligibility than VALL-E.
Generative models in text and computer vision have advanced rapidly by learning broad tasks at scale, yet speech generation models have largely lagged behind. Conventional speech systems typically rely on small, highly curated studio datasets with strict style labels, leaving them unable to generalize across diverse speaking styles, uncurated recording environments, or multi-task scenarios. The article addresses this gap by presenting and evaluating Voicebox, a versatile, non-autoregressive generative model designed to handle diverse speech generation tasks without explicit task-specific training.
The core objective of the article is to demonstrate how training a model on large-scale, text-guided speech infilling—predicting missing speech segments given surrounding audio and transcripts—enables broad in-context task generalization. The researchers evaluate this approach across multiple capabilities, including zero-shot text-to-speech synthesis, speech denoising, content editing, and synthetic data generation for speech recognition. The methodology relies on continuous normalizing flows trained via flow-matching on over 50,000 hours of uncurated multilingual audiobooks across six languages and 60,000 hours of English audiobooks, using a decoupled architecture separating audio and duration modeling.
The findings show that Voicebox establishes a new performance standard across several domains. In English zero-shot text-to-speech, it substantially outperforms leading systems such as VALL-E, lowering the word error rate from 5.9% to 1.9% while improving audio similarity from 0.580 to 0.681 and generating audio up to 20 times faster. In cross-lingual synthesis across six languages without paired multilingual speaker data, it reduces average word error rates from 10.9% to 5.2% compared to prior benchmarks. Furthermore, the model effectively infills corrupted segments during severe noise conditions (achieving a 2.0% word error rate at minus 10 decibels signal-to-noise ratio) and produces synthetic training data so realistic that speech recognition systems trained entirely on it trail real-data benchmarks by only 0.4% to 1.7% in word error rate.
These results demonstrate that speech generation can shift from narrow, label-dependent pipelines to unified, scalable architectures. This significantly lowers computational latency and operational overhead for editing and audio production, while unlocking high-fidelity data generation for downstream models. However, the technology introduces risks regarding the misuse of synthetic voice cloning. To mitigate this risk, the authors demonstrated that a companion classification model can reliably detect Voicebox-generated audio.
Organizations evaluating this technology should explore using speech infilling for high-throughput speech production, content editing, and automated data augmentation pipelines. Before broad deployment, development teams should address existing model limitations: the training relies on read audiobook speech rather than spontaneous conversational audio, depends on external phonetic aligners, and exhibits pronunciation degradation when transferring from dominant languages (such as English) to lower-resource languages. Overall, the evidence provides high confidence in the scalability and fidelity of flow-matching for speech, provided that appropriate guardrails and balanced multilingual data sources are maintained.
- Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, Chengyi Wang et al. (2023). Voicebox is directly motivated as an alternative to VALL-E's autoregressive speech generation approach, using non-autoregressive flow matching to dramatically improve zero-shot TTS speed and intelligibility.
- Paper: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, Jaehyeon Kim et al. (2021). Understanding VITS provides vital foundation on high-quality non-autoregressive speech synthesis architectures and flow-based generative modeling for raw audio.
- Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). DiffWave provides necessary background on non-autoregressive continuous generative diffusion modeling for audio waveform synthesis.
- Paper: FastSpeech 2: Fast and High-Quality End-to-End Text to Speech, Yi Ren et al. (2020). FastSpeech 2 establishes foundational principles of non-autoregressive parallel text-to-speech generation conditioned on acoustic context.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN introduces the standard high-fidelity neural vocoding techniques commonly used to convert generated spectrogram representations into raw audio waveforms.
- Paper: WaveNet: A Generative Model for Raw Audio, Aäron van den Oord et al. (2016). WaveNet establishes the foundational neural audio synthesis framework that sparked modern deep generative modeling of speech.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS advances large-scale multilingual zero-shot voice cloning and controllable streaming text-to-speech across millions of hours of audio.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni integrates large-scale speech generation and understanding into an omnimodal Thinker-Talker foundation model.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This paper investigates scaling flow matching and rectified flow transformers to massive parameter regimes for continuous generative modeling.
- Paper: ELF: Embedded Language Flows, Keya Hu et al. (2026). ELF extends continuous-space flow matching and generative modeling paradigms to discrete linguistic token generation.
