Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Matthew LeApoorv VyasBowen ShiBrian KarrerLeda SariRashel MoritzMary WilliamsonVimal ManoharYossi AdiJay Mahadeokar

article2023NeurIPS479 citations

Introduces Voicebox, a non-autoregressive flow-matching speech model trained on 50,000 hours of audio that performs zero-shot text-to-speech, cross-lingual synthesis, and audio editing up to twenty times faster and with higher intelligibility than VALL-E.

Listen

Generative models in text and computer vision have advanced rapidly by learning broad tasks at scale, yet speech generation models have largely lagged behind. Conventional speech systems typically rely on small, highly curated studio datasets with strict style labels, leaving them unable to generalize across diverse speaking styles, uncurated recording environments, or multi-task scenarios. The article addresses this gap by presenting and evaluating Voicebox, a versatile, non-autoregressive generative model designed to handle diverse speech generation tasks without explicit task-specific training.

The core objective of the article is to demonstrate how training a model on large-scale, text-guided speech infilling—predicting missing speech segments given surrounding audio and transcripts—enables broad in-context task generalization. The researchers evaluate this approach across multiple capabilities, including zero-shot text-to-speech synthesis, speech denoising, content editing, and synthetic data generation for speech recognition. The methodology relies on continuous normalizing flows trained via flow-matching on over 50,000 hours of uncurated multilingual audiobooks across six languages and 60,000 hours of English audiobooks, using a decoupled architecture separating audio and duration modeling.

The findings show that Voicebox establishes a new performance standard across several domains. In English zero-shot text-to-speech, it substantially outperforms leading systems such as VALL-E, lowering the word error rate from 5.9% to 1.9% while improving audio similarity from 0.580 to 0.681 and generating audio up to 20 times faster. In cross-lingual synthesis across six languages without paired multilingual speaker data, it reduces average word error rates from 10.9% to 5.2% compared to prior benchmarks. Furthermore, the model effectively infills corrupted segments during severe noise conditions (achieving a 2.0% word error rate at minus 10 decibels signal-to-noise ratio) and produces synthetic training data so realistic that speech recognition systems trained entirely on it trail real-data benchmarks by only 0.4% to 1.7% in word error rate.

These results demonstrate that speech generation can shift from narrow, label-dependent pipelines to unified, scalable architectures. This significantly lowers computational latency and operational overhead for editing and audio production, while unlocking high-fidelity data generation for downstream models. However, the technology introduces risks regarding the misuse of synthetic voice cloning. To mitigate this risk, the authors demonstrated that a companion classification model can reliably detect Voicebox-generated audio.

Organizations evaluating this technology should explore using speech infilling for high-throughput speech production, content editing, and automated data augmentation pipelines. Before broad deployment, development teams should address existing model limitations: the training relies on read audiobook speech rather than spontaneous conversational audio, depends on external phonetic aligners, and exhibits pronunciation degradation when transferring from dominant languages (such as English) to lower-resource languages. Overall, the evidence provides high confidence in the scalability and fidelity of flow-matching for speech, provided that appropriate guardrails and balanced multilingual data sources are maintained.

arXiv: 2306.15687
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS advances large-scale multilingual zero-shot voice cloning and controllable streaming text-to-speech across millions of hours of audio.
  • Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni integrates large-scale speech generation and understanding into an omnimodal Thinker-Talker foundation model.
  • Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This paper investigates scaling flow matching and rectified flow transformers to massive parameter regimes for continuous generative modeling.
  • Paper: ELF: Embedded Language Flows, Keya Hu et al. (2026). ELF extends continuous-space flow matching and generative modeling paradigms to discrete linguistic token generation.
Cover for Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Abstract

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech generative models are still primitive in terms of scale and task generalization. In this paper, we present Voicebox, the most versatile text-guided generative model for speech at scale. Voicebox is a non-autoregressive flow-matching model trained to infill speech, given audio context and text, trained on over 50K hours of speech that are not filtered or enhanced. Similar to GPT, Voicebox can perform many different tasks through in-context learning, but is more flexible as it can also condition on future context. Voicebox can be used for mono or cross-lingual zero-shot text-to-speech synthesis, noise removal, content editing, style conversion, and diverse sample generation. In particular, Voicebox outperforms the state-of-the-art zero-shot TTS model VALL-E on both intelligibility (5.9% vs 1.9% word error rates) and audio similarity (0.580 vs 0.681) while being up to 20 times faster. Audio samples can be found in \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Background: Flow Matching with an optimal transport path
  • 3.2 Problem formulation
  • 3.3 Model and Training
  • 3.4 Inference
  • 3.5 Classifier-Free Guidance
  • 3.6 Applications
  • 4 Metrics
  • 5 Experiment
  • 5.1 Setup
  • 5.2 Monolingual zero-shot TTS
  • 5.3 Cross-lingual zero-shot TTS
  • 5.4 Transient noise removal
  • 5.5 Diverse speech sampling and application to ASR data generation
  • 5.6 Inference efficiency versus performance
  • 5.7 How context length affects monolingual and cross-lingual zero-shot TTS
  • 5.8 Ablation on generative modeling approaches
  • 6 Ethical Statement
  • 7 Conclusion and Discussion
  • References
  • A Additional Details of Experiment Setup
  • A.1 Vocoder
  • A.2 Phone representation
  • A.3 Data transformation
  • A.4 Cross-lingual zero-shot TTS test data filtering
  • A.5 Setup for training ASR models with synthetic speech
  • A.6 Bi-directional ALiBi Bias
  • B Additional Experiments
  • B.1 Comparing audio model training objectives
  • B.2 Effectiveness on data scaling
  • B.3 Complete results on comparing generative modeling approaches
  • B.4 Transient noise removal in more conditions
  • B.5 Choice of audio model output features
  • C Additional Details and Studies on Metrics
  • C.1 Measuring speech diversity and quality with FSD
  • C.2 Standalone metrics for duration models
  • C.3 Duration model evaluation with standalone metrics
  • C.4 Duration model evaluation with end-to-end metrics
  • C.5 MOS instructions
  • D Detailed Configurations for Acoustic and Duration model training

Knowls

  1. Knowl 1 — Voicebox as text-guided speech infilling

    model/method

    Voicebox is a non-autoregressive continuous normalizing flow model trained to fill in masked speech from its audio surroundings and transcript. For an audio sample xx, transcript yy, and binary temporal mask mm, the masked and contextual audio are xmis=m⊙xx_{\mathrm{mis}}=m\odot x and xctx=(1−m)⊙xx_{\mathrm{ctx}}=(1-m)\odot x, and the training target is the conditional distribution p(xmis∣y,xctx)p(x_{\mathrm{mis}}\mid y,x_{\mathrm{ctx}}). The model uses surrounding audio to infer voice, speaking style, emotion, background noise, and recording conditions, while the transcript specifies linguistic content; it therefore requires no speaker, emotion, or noise labels.

    The same infilling formulation supports zero-shot text-to-speech, cross-lingual synthesis, transient-noise removal, content editing, alignment-preserving style transfer, and unconditional diverse speech sampling. Unlike autoregressive speech generators, Voicebox can condition on both past and future audio context and can infill segments of arbitrary length. The English model was trained on 60K hours of ASR-transcribed English audiobooks, and the multilingual model was trained on 50K hours spanning English, French, German, Spanish, Polish, and Portuguese.

  2. Knowl 2 — Conditional flow matching with an optimal-transport path

    equation

    Voicebox models the conditional speech distribution with a continuous normalizing flow. Let x∈Rdx\in\mathbb{R}^{d} be a data point, t∈[0,1]t\in[0,1] the flow time, p0p_0 a simple prior distribution, and qq the data distribution. A time-dependent vector field vt:Rd→Rdv_t:\mathbb{R}^{d}\rightarrow\mathbb{R}^{d} defines the flow ϕt\phi_t through

    ddtϕt(x)=vt(ϕt(x)),ϕ0(x)=x.\frac{d}{dt}\phi_t(x)=v_t(\phi_t(x)),\qquad \phi_0(x)=x.

    The neural vector field vt(x;θ)v_t(x;\theta) is trained with conditional flow matching. For a data sample x1∼qx_1\sim q, the conditional path starts from p0p_0 and ends at a narrow Gaussian around x1x_1; with terminal noise σmin⁡=10−5\sigma_{\min}=10^{-5}, the optimal-transport path used by Voicebox is

    pt(x∣x1)=N(x∣tx1,[1−(1−σmin⁡)t]2I),p_t(x\mid x_1)=\mathcal{N}\left(x\mid tx_1,\left[1-(1-\sigma_{\min})t\right]^2I\right),

    where II is the d×dd\times d identity matrix. Its target vector field is

    ut(x∣x1)=x1−(1−σmin⁡)x1−(1−σmin⁡)t.u_t(x\mid x_1)=\frac{x_1-(1-\sigma_{\min})x}{1-(1-\sigma_{\min})t}.

    The conditional flow-matching objective is the expected squared vector-field error

    LCFM(θ)=Et,x1,xt[∥ut(xt∣x1)−vt(xt;θ)∥22],\mathcal{L}_{\mathrm{CFM}}(\theta)=\mathbb{E}_{t,x_1,x_t}\left[\left\|u_t(x_t\mid x_1)-v_t(x_t;\theta)\right\|_2^2\right],

    where t∼U[0,1]t\sim\mathcal{U}[0,1], x1∼qx_1\sim q, xt∼pt(⋅∣x1)x_t\sim p_t(\cdot\mid x_1), and utu_t is the known target field. The optimal-transport path gives approximately straight, constant-speed trajectories and was selected because it trains and samples more efficiently than the diffusion path alternatives evaluated in the paper.

  3. Knowl 3 — Two-part audio and duration model

    model/method

    Voicebox separates acoustic generation from duration prediction to retain fine-grained alignment control. Let x=(x1,…,xN)x=(x^1,\ldots,x^N) contain NN audio frames, y=(y1,…,yM)y=(y^1,\ldots,y^M) contain MM phones, and l=(l1,…,lM)l=(l_1,\ldots,l_M) contain phone durations satisfying ∑j=1Mlj=N\sum_{j=1}^{M}l_j=N. The frame-level phone sequence z=rep⁡(y,l)z=\operatorname{rep}(y,l) repeats each phone yjy_j for ljl_j frames. Forced alignment estimates ll and zz during training, and the conditional generation problem is decomposed into an audio model q(xmis∣z,xctx)q(x_{\mathrm{mis}}\mid z,x_{\mathrm{ctx}}) and a duration model q(lmis∣y,lctx)q(l_{\mathrm{mis}}\mid y,l_{\mathrm{ctx}}).

    The audio is represented as an 80-dimensional log-Mel spectrogram at 100 Hz. The audio Transformer receives the current flow sample xt∈RN×80x_t\in\mathbb{R}^{N\times80}, contextual spectrogram xctx∈RN×80x_{\mathrm{ctx}}\in\mathbb{R}^{N\times80}, frame-level phones zz, and flow time tt. A phone embedding lookup table maps each phone to a vector; the phone embedding, noisy spectrogram, and contextual spectrogram are concatenated framewise, projected to the Transformer dimension, and augmented with a sinusoidal embedding of tt. The Transformer predicts a vector field for all spectrogram frames. Training uses the masked flow-matching loss

    Laudio(θ)=E[∥m⊙(ut(xt∣x)−vt(xt,xctx,z;θ))∥22],\mathcal{L}_{\mathrm{audio}}(\theta)=\mathbb{E}\left[\left\|m\odot\left(u_t(x_t\mid x)-v_t(x_t,x_{\mathrm{ctx}},z;\theta)\right)\right\|_2^2\right],

    where mm selects masked frames, x0x_0 is sampled from the prior, xtx_t is the optimal-transport interpolation between x0x_0 and clean spectrogram xx, and utu_t is the corresponding target vector field. Computing the loss only on masked frames improves audio similarity and diversity relative to computing it on all frames.

    The duration component has either a conditional flow-matching implementation or a masked regression implementation. The regression version predicts masked durations with an L1 objective,

    Ldur(θ)=E[∥m′⊙(lmis−g(lctx,y;θ))∥1],\mathcal{L}_{\mathrm{dur}}(\theta)=\mathbb{E}\left[\left\|m'\odot\left(l_{\mathrm{mis}}-g(l_{\mathrm{ctx}},y;\theta)\right)\right\|_1\right],

    where m′m' is a phone-level mask and gg is the duration regressor. The main experiments use the regression duration model by default because it yields more regular durations and lower WER, while the duration flow model produces greater duration diversity.

  4. Knowl 4 — ODE sampling, classifier-free guidance, and task construction

    algorithm

    At inference, Voicebox first samples an acoustic state x0x_0 from the prior and numerically solves the conditional ODE from flow time t=0t=0 to t=1t=1. The ODE derivative is the audio vector field conditioned on the frame-level phones and contextual spectrogram; the final state x1x_1 is decoded to waveform with HiFi-GAN. The number of function evaluations, or NFE, is user-controlled: more evaluations generally improve the ODE approximation but increase runtime. Voicebox often produces high-quality audio with fewer than 10 NFEs, although the default experiment setting uses a midpoint solver with step size 0.06250.0625 and NFE =32=32.

    Classifier-free guidance is applied by dropping the conditioner during training with probability puncond=0.2p_{\mathrm{uncond}}=0.2. For current acoustic state ww, contextual audio xctxx_{\mathrm{ctx}}, frame phones zz, and guidance strength α\alpha, the guided vector field is

    v~t(w,xctx,z;θ)=(1+α)vt(w,xctx,z;θ)−αvt(w;θ),\widetilde v_t(w,x_{\mathrm{ctx}},z;\theta)=(1+\alpha)v_t(w,x_{\mathrm{ctx}},z;\theta)-\alpha v_t(w;\theta),

    where the second term is evaluated without audio and phone conditioning. The same construction is used for the duration model with a separate strength αdur\alpha_{\mathrm{dur}}.

    Voicebox constructs tasks by choosing masks and contexts rather than task-specific heads. For zero-shot TTS, it concatenates reference audio and target audio into one utterance, masks the target region, predicts target durations from the reference duration context and target phones, and infills the target spectrogram. For style transfer, it masks speech while retaining its alignment and conditions on a reference frame-level phone sequence. For noise removal, it masks the corrupted region and regenerates it from the clean surrounding context. For content editing, it retains frames for unchanged phones, predicts durations for replacement phones, and infills only the newly created frames. For diverse text-only sampling, it masks the complete utterance, samples durations from the phone transcript, and generates the entire spectrogram without audio context.

  5. Knowl 5 — Reproducible perceptual metrics for speech generation

    definition

    The paper evaluates generated speech with model-based metrics designed to avoid the limitations of signal-level distances for stochastic generation. Correctness and intelligibility are measured by word error rate (WER) between an automatic transcription of the generated speech and the input text. The English experiments use HuBERT-L, while multilingual experiments use Whisper large-v2. A lower WER generally indicates more recognizable content, but the paper notes that expressive, noisy, or highly diverse valid speech can also produce higher ASR error.

    Audio coherence is measured with WavLM-TDCNN speaker embeddings. SIM-r compares generated speech with the vocoder-resynthesized reference used by VALL-E, whereas SIM-o compares it with the original reference audio and is preferred for cross-model comparability because it does not depend on the reference model's vocoder.

    For unconditional generation, the paper introduces Fréchet Speech Distance (FSD), the Fréchet distance between feature distributions of real and generated speech using self-supervised wav2vec 2.0 features, specifically layer 6 in the main experiments. Lower FSD indicates a distribution that is more similar in quality and diversity to the reference speech. Subjective quality MOS and similarity MOS are also reported on a 1--5 scale using 50 samples per system and 10 ratings per sample, with 95% confidence intervals.

  6. Knowl 6 — Large-scale training and evaluation setup

    experimental setup

    The English Voicebox model is trained on 60K hours of ASR-transcribed audiobook speech; the multilingual model uses 50K hours in English, French, German, Spanish, Polish, and Portuguese. Multilingual sampling uses an upsampling exponent β=0.25\beta=0.25 to increase the probability of lower-resource languages. Montreal Forced Aligner phonemization and alignment provide phone-level and frame-level transcripts. The acoustic representation is an 80-dimensional, 100-Hz log-Mel spectrogram, decoded by a HiFi-GAN vocoder trained on the 60K-hour English corpus.

    The audio model is a 24-layer, 16-head Transformer with dimension 1024, feed-forward dimension 4096, skip connections between symmetric layers, and approximately 330M parameters. The duration model uses 8 layers for English and 10 for multilingual training, with 512/768-dimensional hidden states and approximately 28M/34M parameters. Audio models are trained for 500K updates for English and 750K for multilingual training; duration models are trained for 600K updates. Training uses Adam with peak learning rate 10−410^{-4}, 5K warm-up steps, linear decay, FP16 arithmetic, audio gradient clipping at 0.2, and random masking of contiguous spectrogram spans. Audio sequences are capped at 1,600 frames.

    The principal baselines are VALL-E for English zero-shot TTS, YourTTS for multilingual zero-shot TTS, A3T for speech editing and infilling, and Demucs for speech enhancement. Evaluation covers monolingual and cross-lingual zero-shot TTS, transient-noise removal, content editing, style transfer, diverse speech sampling, and training an ASR model solely on generated speech.

  7. Knowl 7 — Monolingual and cross-lingual zero-shot TTS performance

    empirical result

    On filtered LibriSpeech test-clean, Voicebox substantially improves English zero-shot TTS. In the cross-sentence setting, where a 3-second clip from another utterance by the same speaker is used as context, Voicebox obtains WER 1.91.9, SIM-o 0.6620.662, SIM-r 0.6810.681, QMOS 3.78±0.103.78\pm0.10, and SMOS 3.71±0.113.71\pm0.11. VALL-E obtains WER 5.95.9 and SIM-r 0.5800.580, while YourTTS obtains WER 7.77.7, SIM-o 0.3370.337, QMOS 3.27±0.133.27\pm0.13, and SMOS 3.19±0.143.19\pm0.14. In the continuation setting, Voicebox obtains WER 2.02.0, SIM-o 0.5930.593, and SIM-r 0.6160.616, compared with VALL-E's WER 3.83.8, SIM-o 0.4520.452, and SIM-r 0.5080.508.

    Voicebox also performs cross-lingual zero-shot TTS across all 36 source-target directions formed from six languages, despite never training on multilingual utterances from a single speaker. Averaged over reference languages, Voicebox achieves the following target-language results: German WER 5.05.0, SIM-o 0.4860.486; English 4.44.4, 0.4920.492; Spanish 3.73.7, 0.4940.494; French 5.55.5, 0.4910.491; Polish 5.55.5, 0.4570.457; and Portuguese 5.75.7, 0.4590.459. YourTTS's corresponding averages are available only for its supported target languages: English WER 7.57.5, SIM-o 0.3560.356; French 11.411.4, 0.3500.350; and Portuguese 13.913.9, 0.2990.299. On the shared English, French, and Portuguese targets, Voicebox lowers average WER by 3.13.1, 5.95.9, and 8.18.1 percentage points and raises similarity by 0.1360.136, 0.1410.141, and 0.1600.160, respectively. Subjective multilingual evaluation gives Voicebox average SMOS 3.893.89 versus YourTTS 3.303.30, and average QMOS 3.503.50 versus 3.233.23.

  8. Knowl 8 — Text-guided transient-noise removal by infilling

    empirical result

    Voicebox removes transient noise without being explicitly trained as a denoiser. The test set mixes non-speech noise with filtered LibriSpeech test-clean so that noise overlaps 50% of the speech at signal-to-noise ratio −10-10 dB; Voicebox and A3T receive the transcript and corrupted-segment location, whereas Demucs does not.

    Clean speech has WER 2.22.2, SIM-o 0.6870.687, and QMOS 4.07±0.154.07\pm0.15. The noisy input degrades to WER 41.241.2, SIM-o 0.2870.287, and QMOS 2.50±0.152.50\pm0.15. Demucs obtains 32.532.5, 0.3680.368, and 2.86±0.172.86\pm0.17; A3T obtains 11.511.5, 0.1480.148, and 3.10±0.153.10\pm0.15; Voicebox obtains WER 2.02.0, SIM-o 0.6120.612, and QMOS 3.87±0.173.87\pm0.17. Thus, regenerating only the masked corrupted segment preserves the clean context while restoring intelligible and stylistically coherent speech. Additional tests varying overlap from 30% to 70%, SNR from −10-10 to 1010 dB, and speech versus non-speech noise show that Voicebox remains the most intelligible system and is especially better than Demucs in the high-noise condition.

  9. Knowl 9 — Diverse speech generation and ASR data creation

    empirical result

    Voicebox generates diverse text-only speech whose distribution is closer to real audiobook speech than the evaluated baselines. On LibriSpeech test-other text, ground-truth speech has WER 4.34.3 and FSD 171.1171.1. VITS-VCTK obtains 10.610.6 and 306.6306.6, YourTTS with a LibriSpeech training reference obtains 9.09.0 and 277.9277.9, A3T obtains 37.937.9 and 373.0373.0, and VITS-LJ obtains 5.65.6 and 344.2344.2. Voicebox with regression durations obtains WER 3.13.1 and FSD 155.7155.7; Voicebox with flow-matching durations obtains WER 5.65.6 and FSD 159.8159.8. The flow duration model produces greater speaking-style diversity, which makes the speech somewhat harder for ASR to recognize.

    The paper tests whether synthetic speech can train an ASR system. Each TTS system generates one utterance for each of 281K LibriSpeech training texts. On real test-clean/test-other speech, an ASR model trained on 960 hours of real audio obtains WER 2.6/6.32.6/6.3 without/with a 4-gram language model, while one trained on 100 hours obtains 9.0/21.59.0/21.5. Models trained only on synthetic speech obtain the following WERs, reported as no-language-model test-clean/test-other and 4-gram test-clean/test-other: VITS-LJ 58.0/81.258.0/81.2 and 51.6/78.151.6/78.1; VITS-VCTK 33.8/55.533.8/55.5 and 30.2/53.130.2/53.1; YourTTS 25.0/54.625.0/54.6 and 20.4/51.220.4/51.2; Voicebox with regression durations 7.1/17.67.1/17.6 and 6.5/14.66.5/14.6; and Voicebox with flow-matching durations 3.1/8.33.1/8.3 and 2.6/6.72.6/6.7. Relative to ASR trained on real 960-hour data, the flow-duration Voicebox system is only 0.40.4 and 1.71.7 percentage points worse on test-clean and test-other with a 4-gram language model.

  10. Knowl 10 — Inference efficiency and generative-model ablations

    empirical result

    Voicebox provides a quality--runtime trade-off through the number of ODE function evaluations. For 10 seconds of generated audio, including duration prediction and vocoding, Voicebox takes approximately 0.310.31 seconds at NFE =2=2 without classifier-free guidance, about 20 times faster than VALL-E. At NFE =64=64, it is only 4% slower than VALL-E. In zero-shot TTS, WER remains near 2% across tested NFE and guidance settings, while higher NFE and stronger guidance generally improve speaker similarity. In diverse sampling, higher NFE improves FSD, whereas weaker guidance produces more diverse samples.

    Controlled ablations compare flow matching with the proposed optimal-transport path, flow matching with a variance-preserving diffusion path, and score matching with the same diffusion path. At 150K training updates and NFE =32=32, the optimal-transport flow model achieves zero-shot-TTS WER 2.12.1 and SIM-o 0.5080.508, versus 2.62.6 and 0.4780.478 for diffusion-path flow matching and 5.15.1 and 0.3490.349 for diffusion-path score matching. At NFE values 4,8,16,324,8,16,32, the optimal-transport model obtains WER/SIM-o of 2.4/0.4102.4/0.410, 2.2/0.4812.2/0.481, 2.2/0.5032.2/0.503, and 2.1/0.5082.1/0.508; diffusion-path flow matching obtains 11.5/0.17111.5/0.171, 3.0/0.3593.0/0.359, 2.7/0.4472.7/0.447, and 2.6/0.4782.6/0.478; and diffusion-path score matching obtains 94.5/0.05494.5/0.054, 42.3/0.07642.3/0.076, 11.5/0.21811.5/0.218, and 5.1/0.3495.1/0.349. The optimal-transport flow model therefore trains faster and reaches useful quality with substantially fewer inference steps.

Coverage note — The explicit conversational-speech, phonemizer/forced-aligner, and disentangled-style-control limitations, along with the auxiliary synthetic-speech detector and appendix-only duration-metric studies, were omitted to keep the extraction within the ten most contribution-critical knowls.

References

  1. 1.A. Aghajanyan, L. Yu, A. Conneau, W.-N. Hsu, K. Hambardzumyan, S. Zhang, S. Roller, N. Goyal, O. Levy, and L. Zettlemoyer. Scaling laws for generative mixed-modal language models. ArXiv, abs/2301.03728, 2023.
  2. 2.K. Akuzawa, Y. Iwasawa, and Y. Matsuo. Expressive speech synthesis via modeling expressions with variational autoencoder. ArXiv, abs/1804.02135, 2018.
  3. 3.R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. Common voice: A massively-multilingual speech corpus. In International Conference on Language Resources and Evaluation, 2019.
  4. 4.A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli. XLS-R: self-supervised cross-lingual speech representation learning at scale. In H. Ko and J. H. L. Hansen, editors, Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, pages 2278–2282. ISCA, 2022.
  5. 5.A. Baevski, Y. Zhou, A. Mohamed, and M. Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems, 2020.
  6. 6.H. Bai, R. Zheng, J. Chen, X. Li, M. Ma, and L. Huang. A3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing. In International Conference on Machine Learning, 2022.
  7. 7.Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour. AudioLM: a language modeling approach to audio generation. ArXiv, abs/2209.03143, 2022a.
  8. 8.Z. Borsos, M. Sharifi, and M. Tagliasacchi. SpeechPainter: Text-conditioned speech inpainting. In Interspeech, 2022b.
  9. 9.A. Brock, J. Donahue, and K. Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  10. 10.T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. J. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, M. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. ArXiv, abs/2005.14165, 2020.
  11. 11.E. Casanova, J. Weber, C. D. Shulby, A. C. Júnior, E. Gölge, and M. A. Ponti. YourTTS: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone. In International Conference on Machine Learning, 2021.
  12. 12.E. Casanova, A. C. Junior, C. Shulby, F. S. d. Oliveira, J. P. Teixeira, M. A. Ponti, and S. Aluísio. Tts-portuguese corpus: a corpus for speech synthesis in brazilian portuguese. Language Resources and Evaluation, 56(3):1043–1055, 2022.
  13. 13.R. T. Q. Chen. torchdiffeq, 2018. URL https://github.com/rtqichen/torchdiffeq.
  14. 14.R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud. Neural ordinary differential equations. In Neural Information Processing Systems, 2018.
  15. 15.S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.
  16. 16.A. Défossez, G. Synnaeve, and Y. Adi. Real time speech enhancement in the waveform domain. ArXiv, abs/2006.12847, 2020.
  17. 17.A. Défossez, J. Copet, G. Synnaeve, and Y. Adi. High fidelity neural audio compression. ArXiv, abs/2210.13438, 2022.
  18. 18.B. Desplanques, J. Thienpondt, and K. Demuynck. ECAPA-TDNN: Emphasized Channel Attention, propagation and aggregation in TDNN based speaker verification. In Interspeech, 2020.
  19. 19.P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. Advances in Neural Information Processing Systems, 2021.
  20. 20.J. J. Godfrey, E. C. Holliman, and J. McDaniel. Switchboard: Telephone speech corpus for research and development. In Acoustics, Speech, and Signal Processing, IEEE International Conference on, volume 1, pages 517–520. IEEE Computer Society, 1992.
  21. 21.A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100, 2020.
  22. 22.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Advances in neural information processing systems, 2017.
  23. 23.J. Ho and T. Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  24. 24.J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020.
  25. 25.J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre. Training compute-optimal large language models. ArXiv, abs/2203.15556, 2022.
  26. 26.W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, et al. Hierarchical generative modeling for controllable speech synthesis. In International Conference on Learning Representations, 2019.
  27. 27.W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  28. 28.W.-N. Hsu, T. Remez, B. Shi, J. Donley, and Y. Adi. Revise: Self-supervised speech resynthesis with visual input for universal and generalized speech enhancement. arXiv preprint arXiv:2212.11377, 2022.
  29. 29.R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao. FastDiff: A fast conditional diffusion model for high-quality speech synthesis. In International Joint Conference on Artificial Intelligence, 2022.
  30. 30.Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu, et al. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. Advances in neural information processing systems, 2018.
  31. 31.J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazar’e, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. rahman Mohamed, and E. Dupoux. Libri-Light: A benchmark for asr with limited or no supervision. International Conference on Acoustics, Speech and Signal Processing, 2019.
  32. 32.H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo. StarGAN-VC: non-parallel many-to-many voice conversion using star generative adversarial networks. IEEE Spoken Language Technology Workshop, 2018.
  33. 33.E. Kharitonov, A. Lee, A. Polyak, Y. Adi, J. Copet, K. Lakhotia, T. Nguyen, M. Rivière, A. rahman Mohamed, E. Dupoux, and W.-N. Hsu. Text-free prosody-aware generative spoken language modeling. In Annual Meeting of the Association for Computational Linguistics, 2021.
  34. 34.E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour. Speak, read and prompt: High-fidelity text-to-speech with minimal supervision, 2023.
  35. 35.K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi. Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms. In Interspeech, 2019.
  36. 36.J. Kim, S. Kim, J. Kong, and S. Yoon. Glow-TTS: A generative flow for text-to-speech via monotonic alignment search. Advances in Neural Information Processing Systems, 2020.
  37. 37.J. Kim, J. Kong, and J. Son. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In International Conference on Machine Learning, 2021.
  38. 38.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014.
  39. 39.D. P. Kingma and P. Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. Advances in neural information processing systems, 31, 2018.
  40. 40.F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T.-A. Nguyen, M. Rivière, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi. Textless speech emotion conversion using decomposed and discrete representations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022.
  41. 41.R. Kubichek. Mel-cepstral distance measure for objective speech quality assessment. In Proceedings of IEEE pacific rim conference on communications computers and signal processing, volume 1, pages 125–128. IEEE, 1993.
  42. 42.K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T. Nguyen, J. Copet, A. Baevski, A. B. Mohamed, and E. Dupoux. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021.
  43. 43.A. Łancucki. Fastpitch: Parallel text-to-speech with pitch prediction. In ´ International Conference on Acoustics, Speech and Signal Processing, 2021.
  44. 44.J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey. Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 626–630. IEEE, 2019.
  45. 45.Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. In International Conference on Learning Representations, 2023.
  46. 46.J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. H. Kinnunen, and Z. Ling. The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods. ArXiv, abs/1804.04262, 2018.
  47. 47.M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger. Montreal forced aligner: Trainable text-speech alignment using kaldi. In Interspeech, 2017.
  48. 48.T. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. M. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed, and E. Dupoux. Generative spoken dialogue language modeling. Transactions of the Association for Computational Linguistics, 11:250–266, 2022.
  49. 49.A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, 2021.
  50. 50.V. Panayotov, G. Chen, D. Povey, and S. Khudanpur. Librispeech: An asr corpus based on public domain audio books. International Conference on Acoustics, Speech and Signal Processing, 2015.
  51. 51.D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le. SpecAugment: A simple data augmentation method for automatic speech recognition. In Interspeech, 2019.
  52. 52.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 2019.
  53. 53.A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W.-N. Hsu, A. Mohamed, and E. Dupoux. Speech resynthesis from discrete disentangled self-supervised representations. In Interspeech, 2021.
  54. 54.V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov. Grad-TTS: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning, 2021.
  55. 55.D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al. The kaldi speech recognition toolkit. In Workshop on automatic speech recognition and understanding, 2011.
  56. 56.O. Press, N. A. Smith, and M. Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. ArXiv, abs/2108.12409, 2021.
  57. 57.A. Radford, J. W. Kim, J. Xu, G. Brockman, C. McLeavey, and I. Sutskever. Robust speech recognition via large-scale weak supervision. ArXiv, abs/2212.04356, 2022.
  58. 58.A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever. Zero-shot text-to-image generation. ArXiv, abs/2102.12092, 2021.
  59. 59.Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu. Fastspeech 2: Fast and high-quality end-to-end text to speech. In International Conference on Learning Representations, 2021.
  60. 60.F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer. CrowdMOS: An approach for crowdsourcing mean opinion score studies. In International Conference on Acoustics, Speech and Signal Processing, 2011.
  61. 61.C. Robinson, N. Obin, and A. Roebel. Sequence-to-sequence modelling of F0 for speech emotion conversion. In International Conference on Acoustics, Speech and Signal Processing, 2019.
  62. 62.R. Rombach, R. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  63. 63.C. Saharia, W. Chan, W. Chang, H. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi. Palette: Image-to-image diffusion models. In ACM SIGGRAPH 2022 Conference Proceedings, 2022.
  64. 64.J. Serrà, S. Pascual, J. Pons, R. O. Araz, and D. Scaini. Universal speech enhancement with score-based diffusion. ArXiv, abs/2206.03065, 2022.
  65. 65.J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu. Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions. International Conference on Acoustics, Speech and Signal Processing, 2017.
  66. 66.K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian. Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers. arXiv preprint arXiv:2304.09116, 2023.
  67. 67.R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous. Towards end-to-end prosody transfer for expressive speech synthesis with tacotron. In international conference on machine learning, pages 4693–4702. PMLR, 2018.
  68. 68.Y. Song and S. Ermon. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32, 2019.
  69. 69.X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, X. Wang, Y. Leng, Y. Yi, L. He, F. K. Soong, T. Qin, S. Zhao, and T.-Y. Liu. NaturalSpeech: End-to-end text to speech synthesis with human-level quality. ArXiv, abs/2205.04421, 2022.
  70. 70.A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. ArXiv, abs/1706.03762, 2017.
  71. 71.C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. M. Pino. fairseq s2 : A scalable and integrable speech synthesis toolkit. In Conference on Empirical Methods in Natural Language Processing, 2021.
  72. 72.C. Wang, S. Chen, Y. Wu, Z.-H. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei. Neural codec language models are zero-shot text to speech synthesizers. ArXiv, abs/2301.02111, 2023.
  73. 73.Y. Wang, D. Stanton, Y. Zhang, R. J. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. In International Conference on Machine Learning, 2018.
  74. 74.Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee. A regression approach to speech enhancement based on deep neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(1):7–19, 2014.
  75. 75.J. Yamagishi, C. Veaux, and K. MacDonald. Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92). 2019.
  76. 76.R. Yamamoto, E. Song, and J.-M. Kim. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In International Conference on Acoustics, Speech and Signal Processing, 2020.
  77. 77.N. Yu, V. Skripniuk, S. Abdelnabi, and M. Fritz. Artificial fingerprinting for generative models: Rooting deepfake attribution in training data. In Proceedings of the IEEE/CVF International conference on computer vision, pages 14448–14457, 2021.
  78. 78.N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2022.
  79. 79.H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu. Libritts: A corpus derived from librispeech for text-to-speech. arXiv preprint arXiv:1904.02882, 2019.

Citation

MLA
Le, M., et al. “Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale”. arXiv, 2023, http://arxiv.org/abs/2306.15687v2.
APA
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., & Hsu, W.-N. (2023). Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. arXiv. http://arxiv.org/abs/2306.15687v2
Chicago
Le, M., A. Vyas, B. Shi, et al. 2023. “Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale”. arXiv. http://arxiv.org/abs/2306.15687v2.
Harvard
Le, M. et al. (2023) “Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.15687v2.
Vancouver
1. Le M, Vyas A, Shi B, et al (2023) Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale. arXiv

BibTeX

@article{le2023voicebox,
  title = {Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale},
  author = {Le, Matthew and Vyas, Apoorv and Shi, Bowen and Karrer, Brian and Sari, Leda and Moritz, Rashel and Williamson, Mary and Manohar, Vimal and Adi, Yossi and Mahadeokar, Jay and Hsu, Wei-Ning},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.15687v2},
  eprint = {2306.15687}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors