NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Zeqian JuYuancheng WangKai ShenXu TanDetai XinDongchao YangEric LiuYichong LengKaitao SongSiliang Tang

article2024ICML328 citations

Presents NaturalSpeech 3, a zero-shot speech synthesis system that disentangles speech into content, prosody, timbre, and acoustic subspaces using factorized diffusion models, achieving audio quality and speaker similarity on par with human recordings.

Listen

Generating natural, highly similar human speech from a short reference prompt without prior training on the target voice has been a major challenge in artificial intelligence. While recent text-to-speech systems have scaled up model sizes and training data, they frequently fall short in voice naturalness, prosodic expressiveness, and speaker similarity. This shortfall occurs because speech is inherently complex, entangling multiple attributes—such as linguistic content, pitch rhythm, acoustic nuances, and speaker identity—into a single waveform that monolithic systems struggle to model cleanly.

The main objective of the article is to develop and evaluate NaturalSpeech 3, a text-to-speech system that factorizes speech into disentangled sub-attributes and generates them separately using discrete diffusion models to achieve human-level zero-shot synthesis.

The researchers developed an end-to-end framework comprising two core components: a neural speech codec called FACodec and a factorized diffusion model. FACodec decomposes speech waveforms into distinct representations of content, prosody, acoustic details, and timbre using specialized techniques like low-dimensional information bottlenecks and gradient reversal to prevent attribute leakage. The factorized diffusion model then generates duration, pitch contours, phonemes, and acoustic textures in sequence using prompts for in-context learning. The system was trained on standard benchmarks, including the 60,000-hour Libri-Light dataset, and scaled up to 1 billion parameters on a 200,000-hour corpus. Evaluation was conducted across standardized objective metrics and subjective human listening tests on multi-speaker datasets.

The evaluation produced several notable findings. First, NaturalSpeech 3 achieved human-level voice quality and naturalness on the multi-speaker LibriSpeech test benchmark, matching ground-truth human recordings in subjective scores while achieving an improved word error rate of 1.81% compared to 1.94% for human audio. Second, the system established new state-of-the-art benchmarks in speaker similarity, achieving a 0.67 objective similarity score and outperforming existing leading baselines. Third, on emotional speech tests, the model substantially improved prosodic similarity, reducing speech distortion metrics across eight distinct emotions and increasing emotion recognition accuracy from roughly 30–40% to 52%. Finally, data and model scaling demonstrated consistent improvements in speech accuracy and speaker fidelity, while reducing inference latency by more than 15 times compared to leading autoregressive baselines.

These findings indicate that factorizing speech attributes into modular components significantly simplifies generative modeling, resolving the traditional trade-off between output quality and generation speed. The approach also enables zero-shot speech attribute manipulation—such as adjusting speaking rate or applying one speaker's timbre to another's expressive prosody—without retraining. For practical deployments, the high efficiency and modularity lower operational computational costs while expanding capabilities in digital assistants, voice localization, and personalized audio production.

Organizations developing or deploying speech synthesis should adopt factorized attribute architectures to improve controllability, speed, and voice fidelity. As recommended next steps, teams should explore scaling factorized architectures into larger foundational models and evaluate deploying fast single-step diffusion variants where real-time execution is critical. Additionally, because high-fidelity zero-shot voice cloning increases the risk of voice impersonation and identity spoofing, stakeholders must invest in robust synthetic speech detection and misuse-reporting protocols alongside deployment.

These conclusions are supported with high confidence by extensive benchmark evaluations, but several limitations remain. NaturalSpeech 3 was trained primarily on English audiobook datasets, meaning performance has not yet been demonstrated across multiple languages or diverse real-world acoustic backgrounds. Furthermore, the underlying speech codec still relies on supervised phoneme transcriptions during training, which creates an annotation bottleneck for scaling to low-resource languages.

  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS advances large-scale zero-shot speech synthesis, voice cloning, and streaming tokenization, continuing the progression toward ultra-natural, multi-speaker zero-shot TTS systems.
Cover for NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Abstract

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility, and achieves on-par quality with human recordings. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 NaturalSpeech 3
  • 3.1 Overall Architecture
  • 3.2 FACodec for Attribute Factorization
  • 3.2.1 FACodec Model Overview
  • 3.2.2 Attribute Disentanglement
  • 3.3 Factorized Diffusion Model
  • 3.3.1 Model Overview
  • 3.3.2 Diffusion Formulation
  • 3.4 Connections to the NaturalSpeech Series
  • 4 Experiments and Results
  • 4.1 Experimental Settings
  • 4.2 Experimental Results on Zero-shot TTS
  • 4.2.1 Generation Quality
  • 4.2.2 Generation Similarity
  • 4.2.3 Robustness
  • 4.2.4 Human-Level Naturalness on LibriSpeech Testset
  • 4.3 Ablation Study and Method Analyses
  • 4.3.1 Ablation Study
  • 4.3.2 Method Analyses
  • 4.3.3 Experimental Results on FACodec
  • 4.4 Effectiveness of Data and Model Scaling
  • 5 Conclusion
  • 6 Boarder Impact
  • References
  • A Details of Factorization Diffusion Model
  • A.1 Model Configuration
  • A.2 Training and Inference Details
  • A.3 Evaluation Baselines
  • A.4 More Experimental Results on Zero-shot TTS
  • A.5 Latency Analysis
  • A.6 Ablation Study on Duration Diffusion Model
  • A.7 Details of Prosody Similarity Evaluation
  • B Details of FACodec
  • B.1 Implementation Details
  • B.2 Reconstruction Performance Comparison
  • B.3 Zero-shot Voice Conversion
  • B.4 Ablation Study
  • C Limitation and Future Works

Knowls

  1. Knowl 1 — NaturalSpeech 3 Architecture and Factorized Generation Strategy

    model/method

    NaturalSpeech 3 is a zero-shot text-to-speech (TTS) synthesis system based on a divide-and-conquer attribute factorization paradigm. The system factorizes speech into five distinct attribute subspaces:

    1. Duration: Phoneme durations modeled explicitly in a non-autoregressive framework.
    2. Timbre: A global speaker embedding extracted directly from a reference speech prompt without generative sampling.
    3. Prosody: Frame-level pitch and intonation representations.
    4. Content: Fine-grained phonemic information.
    5. Acoustic Details: Residual high-frequency and nuanced acoustic features.

    The overall architecture comprises two primary components:

    • FACodec (Factorized Audio Codec): A neural speech codec that factorizes a raw speech waveform into disentangled discrete tokens for prosody (zpz_p), content (zcz_c), and acoustic details (zdz_d), alongside a continuous global timbre vector (hth_t), and reconstructs the waveform from these components.
    • Factorized Diffusion Model: A discrete diffusion network that generates factorized tokens sequentially conditioned on phoneme text representations and corresponding attribute prompts. Phoneme durations and phoneme-level prosody are predicted first, expanding text to frame-level condition cphc_{ph}. Next, frame-level prosody codes zpz_p are generated, followed by content codes zcz_c conditioned on zpz_p and cphc_{ph}, and finally acoustic detail codes zdz_d conditioned on zp,zc,z_p, z_c, and cphc_{ph}. The generated codes along with the prompt timbre vector hth_t are passed to the FACodec decoder to synthesize the speech waveform.
  2. Knowl 2 — FACodec Architecture and Disentanglement Framework

    model/method

    FACodec is a neural audio codec designed to decompose speech waveforms into disentangled representations of content, prosody, acoustic details, and timbre while maintaining high-fidelity reconstruction.

    Model Components

    • Speech Encoder: Convolutional network with a downsampling rate of 200 on 16 kHz audio, producing a frame every 12.5 ms.
    • Timbre Extractor: A Conformer/Transformer encoder that processes the encoder latents to produce a single time-invariant global timbre embedding hth_t.
    • Factorized Vector Quantizers (FVQ): Three discrete quantizer modules using codebooks of size 1024:
      • Content FVQ (extFVQc ext{FVQ}_c) with Nqc=2N_{q_c} = 2 quantizers.
      • Prosody FVQ (extFVQp ext{FVQ}_p) with Nqp=1N_{q_p} = 1 quantizer.
      • Acoustic Detail FVQ (extFVQd ext{FVQ}_d) with Nqd=3N_{q_d} = 3 quantizers.
    • Speech Decoder: A convolutional neural vocoder mirroring the encoder with SnakeBeta activations. The quantized representations are summed (z=zp+zc+zdz = z_p + z_c + z_d), modulated by timbre hth_t via conditional layer normalization, and decoded into the waveform.

    Disentanglement Techniques

    1. Information Bottleneck: Encoder outputs are projected into an 8-dimensional space prior to vector quantization in extFVQp ext{FVQ}_p, extFVQc ext{FVQ}_c, and extFVQd ext{FVQ}_d, and projected back to the original dimension afterwards to strip extraneous information.
    2. Auxiliary Supervision: Post-quantization prosody latent zpz_p is supervised by frame-level normalized F0F_0 (z-score); content latent zcz_c is supervised by frame-level phoneme classification; timbre vector hth_t is supervised by speaker identification.
    3. Gradient Reversal Layers (GRL): Adversarial classifiers with gradient reversal are applied to remove leaked information: phoneme-GRL on zpz_p; F0F_0-GRL on zcz_c; both phoneme-GRL and F0F_0-GRL on zdz_d; and speaker-GRL on (zp+zc+zd)(z_p + z_c + z_d).
    4. Detail Dropout: The detail latent zdz_d is randomly set to zero during training with probability pp, forcing zpz_p, zcz_c, and hth_t to encode the core speech attributes and leaving zdz_d to cover solely high-frequency acoustic details.

    Training Loss Formulation

    The generator loss is: L=λrecLrec+λadvLadv+λfeatLfeat+λcodebookLcodebook+λcommitLcommit+λphLph+λf0Lf0+λgr-phLgr-ph+λgr-f0Lgr-f0+λgr-spkLgr-spk\mathcal{L} = \lambda_{rec}\mathcal{L}_{rec} + \lambda_{adv}\mathcal{L}_{adv} + \lambda_{feat}\mathcal{L}_{feat} + \lambda_{codebook}\mathcal{L}_{codebook} + \lambda_{commit}\mathcal{L}_{commit} + \lambda_{ph}\mathcal{L}_{ph} + \lambda_{f0}\mathcal{L}_{f0} + \lambda_{gr\text{-}ph}\mathcal{L}_{gr\text{-}ph} + \lambda_{gr\text{-}f0}\mathcal{L}_{gr\text{-}f0} + \lambda_{gr\text{-}spk}\mathcal{L}_{gr\text{-}spk} with coefficients λrec=10.0\lambda_{rec}=10.0, λadv=2.0\lambda_{adv}=2.0, λfeat=2.0\lambda_{feat}=2.0, λcodebook=1.0\lambda_{codebook}=1.0, λcommit=0.25\lambda_{commit}=0.25, λph=5.0\lambda_{ph}=5.0, λf0=5.0\lambda_{f0}=5.0, λgr-ph=5.0\lambda_{gr\text{-}ph}=5.0, λgr-f0=5.0\lambda_{gr\text{-}f0}=5.0, and λgr-spk=1.0\lambda_{gr\text{-}spk}=1.0.

  3. Knowl 3 — Discrete Masked Diffusion Formulation and Guided Inference

    model/method

    NaturalSpeech 3 generates discrete token sequences for each factorized speech attribute using a discrete masked diffusion formulation.

    Forward Process

    For a target discrete token sequence X=[xi]i=1NX = [x_i]_{i=1}^N, prompt tokens XpX^p, and condition CC, the forward process corrupts XX into Xt=X⊙MtX_t = X \odot M_t using a binary mask Mt=[mt,i]i=1NM_t = [m_{t,i}]_{i=1}^N with independently sampled elements: mt,i∼Bernoulli(σ(t)),σ(t)=sin⁡(πt2T),t∈(0,T]m_{t,i} \sim \text{Bernoulli}(\sigma(t)), \quad \sigma(t) = \sin\left(\frac{\pi t}{2T}\right), \quad t \in (0, T] where mt,i=1m_{t,i} = 1 replaces xix_i with a special [MASK] token, and mt,i=0m_{t,i} = 0 leaves xix_i unchanged (X0=XX_0 = X and XTX_T is fully masked).

    Training Objective

    The neural network pθp_\theta optimizes the negative log-likelihood of masked positions: Lmask=EX∈D,t∈[0,T][−∑i=1Nmt,i⋅log⁡pθ(xi∣Xt,Xp,C)]\mathcal{L}_{mask} = \mathbb{E}_{X \in \mathcal{D}, t \in [0, T]} \left[ -\sum_{i=1}^N m_{t,i} \cdot \log p_\theta(x_i \mid X_t, X^p, C) \right] Prompt dropout is applied with probability pcfg=0.15p_{cfg} = 0.15 to train unconditional generation for classifier-free guidance.

    Reverse Inference

    Starting from fully masked sequence XTX_T:

    1. At step tt, sample prediction X^0∼pθ(X0∣Xt,Xp,C)\hat{X}_0 \sim p_\theta(X_0 \mid X_t, X^p, C).
    2. Assign confidence scores: pθ(x^i∣Xt,Xp,C)p_\theta(\hat{x}_i \mid X_t, X^p, C) if mt,i=1m_{t,i}=1, and 1.01.0 if token xix_i was already unmasked in XtX_t.
    3. Add Gumbel noise to confidences and re-mask the ⌊N⋅σ(t−Δt)⌋\lfloor N \cdot \sigma(t - \Delta t) \rfloor positions with the lowest confidences to form Xt−ΔtX_{t-\Delta t}.

    Classifier-Free Guidance (CFG)

    During inference, guided output logits gcfgg_{cfg} are extrapolated with guidance scale α\alpha (set to 1.0) and rescaled: gcfg=g(X∣Xp)+α⋅(g(X∣Xp)−g(X))g_{cfg} = g(X \mid X^p) + \alpha \cdot \bigl(g(X \mid X^p) - g(X)\bigr) gfinal=std(g(X∣Xp))×gcfgstd(gcfg)g_{final} = \text{std}\bigl(g(X \mid X^p)\bigr) \times \frac{g_{cfg}}{\text{std}(g_{cfg})}

  4. Knowl 4 — Zero-Shot Text-to-Speech Performance on LibriSpeech Test-Clean

    data/table

    The performance of NaturalSpeech 3 on the LibriSpeech test-clean zero-shot TTS benchmark (using 3-second prompt clips from 40 speakers) is evaluated against ground truth recordings and state-of-the-art TTS baselines. Metrics include speaker similarity to original prompt (Sim-O) and reconstructed prompt (Sim-R) via WavLM-TDCNN, Word Error Rate (WER via HuBERT-large-ls960-ft and Conformer-Transducer WER⋆\text{WER}^\star), UTMOS (speech quality MOS predictor), Comparative MOS (CMOS), and Similarity MOS (SMOS).

    Method Sim-O ↑\uparrow Sim-R ↑\uparrow WER ↓\downarrow UTMOS ↑\uparrow CMOS ↑\uparrow SMOS ↑\uparrow
    Ground Truth 0.68 - 1.94 4.14 +0.08 3.85
    VALL-E (paper reported) - 0.58 5.90 - - -
    VALL-E (reproduced) 0.47 0.51 6.11 3.68 -0.60 3.46
    NaturalSpeech 2 0.55 0.62 1.94 3.88 -0.18 3.65
    Voicebox 0.64 0.67 2.03 3.82 -0.23 3.69
    Mega-TTS 2 0.53 - 2.32 4.02 -0.20 3.63
    UniAudio 0.57 0.68 2.49 3.79 -0.25 3.71
    StyleTTS 2 0.38 - 2.49 3.94 -0.21 3.07
    HierSpeech++ 0.51 - 6.33 3.80 -0.41 3.50
    NaturalSpeech 3 0.67 0.76 1.81 4.30 0.00 4.01

    NaturalSpeech 3 achieves human-level speech naturalness and quality (0.00 CMOS vs. +0.08 for ground truth recordings; 4.30 UTMOS) on the multi-speaker LibriSpeech test set, while establishing state-of-the-art results in speaker similarity (0.67 Sim-O, 4.01 SMOS) and word error rate (1.81% WER).

  5. Knowl 5 — Prosody Generation and Similarity on RAVDESS Emotional Speech

    data/table

    Prosodic similarity and expressive transfer were evaluated on the RAVDESS benchmark across 8 emotions (neutral, calm, happy, sad, angry, fearful, disgust, surprised) with strong intensity. Objective metrics comprise average Mel-Cepstral Distortion (Avg MCD ↓\downarrow) between generated and ground truth speech and top-1 emotion classification accuracy (MCD-Acc ↑\uparrow) using a K-Nearest-Neighbors classifier over MCD distances. Subjective evaluations include CMOS and SMOS.

    Method Avg MCD ↓\downarrow MCD-Acc ↑\uparrow CMOS ↑\uparrow SMOS ↑\uparrow
    Ground Truth 0.00 1.00 +0.17 4.42
    VALL-E 5.03 0.34 -0.55 3.80
    NaturalSpeech 2 4.56 0.25 -0.22 4.04
    Voicebox 4.88 0.34 -0.34 3.92
    Mega-TTS 2 4.44 0.39 -0.20 4.51
    StyleTTS 2 4.50 0.40 -0.25 3.98
    HierSpeech++ 6.08 0.30 -0.37 3.87
    NaturalSpeech 3 4.28 0.52 0.00 4.72

    NaturalSpeech 3 outperforms all baseline models by achieving the lowest average MCD (4.28), the highest emotion preservation accuracy (52% MCD-Acc, +12% over the next best baseline), and the highest SMOS (4.72), demonstrating superior prosodic disentanglement and generation fidelity.

  6. Knowl 6 — FACodec Waveform Reconstruction and Zero-Shot Voice Conversion Performance

    data/table

    FACodec's reconstruction performance and its zero-shot voice conversion capability through timbre swapping were evaluated against dedicated codec models and voice conversion systems.

    Audio Reconstruction Quality (16 kHz Audio)

    Codec Hop Size Codebooks Bandwidth PESQ ↑\uparrow STOI ↑\uparrow MCD ↓\downarrow
    EnCodec (24kHz) 320 8 6.0 kbps 3.28 0.94 2.70
    HiFi-Codec 320 4 2.0 kbps 3.17 0.93 3.05
    DAC 320 9 4.5 kbps 3.52 0.95 2.65
    SoundStream 200 6 4.8 kbps 3.03 0.90 3.38
    SoundStream 200 12 9.6 kbps 3.45 0.94 2.76
    FACodec 200 6 4.8 kbps 3.47 0.95 2.59

    FACodec outperforms SoundStream at identical bandwidth (3.47 vs. 3.03 PESQ, 2.59 vs. 3.38 MCD) and matches SoundStream operating at twice the bitrate (9.6 kbps).

    Zero-Shot Voice Conversion (VCTK Dataset)

    Voice conversion is performed directly by decoding source content, prosody, and detail codes with the prompt speaker's timbre embedding: D(zcsrc,zpsrc,zdsrc,htprompt)\mathcal{D}(z_c^{src}, z_p^{src}, z_d^{src}, h_t^{prompt}).

    Model Sim-O ↑\uparrow WER ↓\downarrow
    Ground Truth - 3.25
    YourTTS 0.72 10.1
    Make-A-Voice (VC) 0.68 6.20
    LM-VC 0.82 4.91
    UniAudio 0.87 4.80
    FACodec 0.86 3.46

    Without task-specific training, FACodec achieves speaker similarity on par with UniAudio (0.86 vs 0.87 Sim-O) while obtaining significantly lower WER (3.46% vs 4.80%), demonstrating effective timbre decoupling.

  7. Knowl 7 — Ablation Studies on Factorization, Guidance, and Generative Extensibility

    data/table

    Ablation experiments quantify the impact of speech factorization, classifier-free guidance, and the extensibility of the factorization framework to autoregressive models.

    Ablation on LibriSpeech Test-Clean

    Ablation Variant Sim-O / Sim-R ↑\uparrow WER ↓\downarrow CMOS ↑\uparrow SMOS ↑\uparrow
    NaturalSpeech 3 0.67 / 0.76 1.81 0.00 4.01
    – w/o Factorization (SoundStream tokens, unified generation) 0.55 / 0.61 2.49 -0.25 3.59
    – w/o Classifier-Free Guidance 0.64 / 0.72 1.81 -0.06 3.80

    Removing factorization leads to steep declines across all dimensions (−0.12-0.12 Sim-O, +0.68+0.68 WER, −0.25-0.25 CMOS, −0.42-0.42 SMOS).

    Extensibility to Autoregressive Models (VALL-E + FACodec)

    Applying FACodec to an autoregressive model (VALL-E AR generates prosody tokens, followed by NAR generation of content and acoustic detail tokens) yields:

    Model Sim-O / Sim-R ↑\uparrow WER ↓\downarrow CMOS ↑\uparrow SMOS ↑\uparrow
    VALL-E + FACodec 0.57 / 0.65 5.60 +0.24 3.61
    VALL-E (Baseline) 0.47 / 0.51 6.11 0.00 3.46

    The factorized design consistently improves speaker similarity (+0.10+0.10 Sim-O), speech intelligibility (−0.51-0.51 WER), and quality (+0.24+0.24 CMOS) over standard VALL-E.

  8. Knowl 8 — Scaling Trends with Training Data and Model Size in NaturalSpeech 3

    data/table

    The scaling properties of the Factorized Diffusion TTS system were assessed on an internal test set of 30 audio clips across varying data volumes and model parameter counts.

    Training Data Scaling (Fixed 500M Parameter Model)

    Training Speech Hours Sim-O ↑\uparrow WER ↓\downarrow
    1K hours (LibriLight subset) 0.64 3.94
    60K hours (LibriLight full) 0.72 3.03
    200K hours (Internal dataset) 0.73 2.11

    Scaling data from 1K to 200K hours reduces the WER by 1.83% and increases speaker similarity (Sim-O) from 0.64 to 0.73.

    Model Parameter Scaling (Fixed 200K Hours Dataset)

    Model Size (Transformer Layers) Sim-O ↑\uparrow WER ↓\downarrow
    500M parameters (12 layers) 0.73 2.11
    1B parameters (24 layers) 0.78 1.71

    Doubling the Transformer layer depth and increasing capacity to 1B parameters boosts speaker similarity to 0.78 Sim-O and further reduces WER to 1.71%.

  9. Knowl 9 — Inference Latency and Diffusion Sampling Step Efficiency

    empirical result

    NaturalSpeech 3 requires 4 diffusion sampling iterations per module (phoneme-level prosody, duration, frame-level prosody, content, and acoustic details), totaling 60 neural network forward passes when including classifier-free guidance evaluations (excluding duration, which does not use guidance).

    On an NVIDIA V100 GPU benchmarking LibriSpeech test-clean:

    • Standard NaturalSpeech 3 (60 NFE): Achieves a Real-Time Factor (RTF) of 0.296, representing a 15.27×15.27\times speedup over VALL-E (extRTF=4.520 ext{RTF} = 4.520) and a 1.24×1.24\times speedup over NaturalSpeech 2 (extRTF=0.366,150 NFE ext{RTF} = 0.366, 150\text{ NFE}), while scoring 0.67 Sim-O, 0.76 Sim-R, and 4.30 UTMOS.
    • NaturalSpeech 3 One-Step (15 NFE): Reducing each diffusion process from 4 iterations to 1 iteration yields an RTF of 0.067 (4.41×4.41\times faster than the 4-step setup) with minimal quality degradation (0.66 Sim-O, 0.75 Sim-R, 4.01 UTMOS).
  10. Knowl 10 — Limitations of NaturalSpeech 3

    limitation

    NaturalSpeech 3 has several stated limitations:

    1. Attribute Space Coverage: The factorization accounts for content, prosody, duration, acoustic details, and timbre, but omits explicit subspaces for other acoustic factors such as energy variations and background environmental sounds or noise.
    2. Training Data and Language Scope: Models are trained primarily on English audiobook corpora (LibriLight/LibriVox), which limits coverage of conversational speech styles, spontaneous voices, and multilingual speech synthesis.
    3. Supervision Requirement for FACodec: FACodec depends on frame-level phoneme alignments for content supervision during training, preventing fully unsupervised training on arbitrary unannotated audio data.

Coverage note — None was omitted; all key architectural components (FACodec and Factorized Diffusion), mathematical formulations, experimental benchmarks (LibriSpeech, RAVDESS, VCTK), ablation studies, latency comparisons, scaling analyses, and stated limitations are covered.

References

  1. 1.Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. Tacotron: Towards end-to-end speech synthesis. Proc. Interspeech 2017, pages 4006–4010, 2017.
  2. 2.Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al. Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4779–4783. IEEE, 2018.
  3. 3.Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. FastSpeech: Fast, robust and controllable text to speech. In NeurIPS, 2019.
  4. 4.Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al. Naturalspeech: End-to-end text-to-speech synthesis with human-level quality. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
  5. 5.Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers. arXiv preprint arXiv:2304.09116, 2023.
  6. 6.Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111, 2023.
  7. 7.Ziyue Jiang, Yi Ren, Zhenhui Ye, Jinglin Liu, Chen Zhang, Qian Yang, Shengpeng Ji, Rongjie Huang, Chunfeng Wang, Xiang Yin, et al. Mega-tts: Zero-shot text-to-speech at scale with intrinsic inductive bias. arXiv preprint arXiv:2306.03509, 2023.
  8. 8.Jaehyeon Kim, Jungil Kong, and Juhee Son. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. arXiv preprint arXiv:2106.06103, 2021.
  9. 9.Dan Lim, Sunghee Jung, and Eesung Kim. Jets: Jointly training fastspeech2 and hifi-gan for end to end text to speech. arXiv preprint arXiv:2203.16852, 2022.
  10. 10.Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. Grad-TTS: A diffusion probabilistic model for text-to-speech. arXiv preprint arXiv:2105.06337, 2021.
  11. 11.Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al. Voicebox: Text-guided multilingual universal speech generation at scale. arXiv preprint arXiv:2306.15687, 2023.
  12. 12.Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. Audiolm: a language modeling approach to audio generation. arXiv preprint arXiv:2209.03143, 2022.
  13. 13.Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi. Soundstorm: Efficient parallel audio generation. arXiv preprint arXiv:2305.09636, 2023.
  14. 14.Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021.
  15. 15.Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression. arXiv preprint arXiv:2210.13438, 2022.
  16. 16.Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox. Unsupervised speech decomposition via triple information bottleneck. In International Conference on Machine Learning, pages 7836–7846. PMLR, 2020.
  17. 17.Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson. AutoVC: Zero-shot voice style transfer with only autoencoder loss. In International Conference on Machine Learning, pages 5210–5219. PMLR, 2019.
  18. 18.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33, 2020.
  19. 19.Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour. Speak, read and prompt: High-fidelity text-to-speech with minimal supervision. arXiv preprint arXiv:2302.03540, 2023.
  20. 20.Rongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang, Luping Liu, Zhenhui Ye, Ziyue Jiang, Chao Weng, Zhou Zhao, and Dong Yu. Make-a-voice: Unified voice synthesis with discrete representation. arXiv preprint arXiv:2305.19269, 2023.
  21. 21.Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen Meng, and Dong Yu. Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt. arXiv preprint arXiv:2301.13662, 2023.
  22. 22.Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, and Kai Yu. Unicats: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding. arXiv preprint arXiv:2306.07547, 2023.
  23. 23.Eliya Nachmani, Alon Levkovitch, Julian Salazar, Chulayutsh Asawaroengchai, Soroosh Mariooryad, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich. Lms with a voice: Spoken language modeling beyond speech tokens. arXiv preprint arXiv:2305.15255, 2023.
  24. 24.Yinghao Aaron Li, Cong Han, Vinay S Raghavan, Gavin Mischler, and Nima Mesgarani. Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models. arXiv preprint arXiv:2306.07691, 2023.
  25. 25.Sang-Hoon Lee, Ha-Yeong Choi, Seung-Bin Kim, and Seong-Whan Lee. Hierspeech++: Bridging the gap between semantic and acoustic representation of speech by hierarchical variational inference for zero-shot speech synthesis. arXiv preprint arXiv:2311.12454, 2023.
  26. 26.Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
  27. 27.Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al. Parallel WaveNet: Fast high-fidelity speech synthesis. In International conference on machine learning, pages 3918–3926. PMLR, 2018.
  28. 28.Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aäron Courville, and Yoshua Bengio. Char2wav: End-to-end speech synthesis. 2017.
  29. 29.Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. Deep Voice 3: 2000-speaker neural text-to-speech. Proc. ICLR, pages 214–217, 2018.
  30. 30.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. Neural speech synthesis with Transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6706–6713, 2019.
  31. 31.Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. Glow-TTS: A generative flow for text-to-speech via monotonic alignment search. Advances in Neural Information Processing Systems, 33, 2020.
  32. 32.Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. High-fidelity audio compression with improved rvqgan. arXiv preprint arXiv:2306.06546, 2023.
  33. 33.Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron Weiss, and Yonghui Wu. Parallel Tacotron: Non-autoregressive and controllable TTS. arXiv preprint arXiv:2010.11439, 2020.
  34. 34.Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, and Zhou Zhao. DiffSinger: Singing voice synthesis via shallow diffusion mechanism. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11020–11028, 2022.
  35. 35.Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al. Uniaudio: An audio foundation model toward universal audio generation. arXiv preprint arXiv:2310.00704, 2023.
  36. 36.Hyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee, Hoon Heo, and Kyogu Lee. Neural analysis and synthesis: Reconstructing speech from self-supervised representations. Advances in Neural Information Processing Systems, 34:16251–16265, 2021.
  37. 37.Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, and Hyeongju Kim. Nansy++: Unified voice synthesis with neural analysis and synthesis. arXiv preprint arXiv:2211.09407, 2022.
  38. 38.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. Speech resynthesis from discrete disentangled self-supervised representations. arXiv preprint arXiv:2104.00355, 2021.
  39. 39.Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 244–250. IEEE, 2021.
  40. 40.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33:12449–12460, 2020.
  41. 41.Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised pre-training for speech recognition. Proc. Interspeech 2019, pages 3465–3469, 2019.
  42. 42.Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu. Speechtokenizer: Unified speech tokenizer for speech large language models. arXiv preprint arXiv:2308.16692, 2023.
  43. 43.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  44. 44.Xue Jiang, Xiulian Peng, Yuan Zhang, and Yan Lu. Disentangled feature learning for real-time neural speech coding. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.
  45. 45.Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, sheng zhao, and Tie-Yan Liu. AdaSpeech: Adaptive text to speech for custom voice. In International Conference on Learning Representations, 2021.
  46. 46.Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627, 2021.
  47. 47.SiCheng Yang, Methawee Tantrawenith, Haolin Zhuang, Zhiyong Wu, Aolan Sun, Jianzong Wang, Ning Cheng, Huaizhen Tang, Xintao Zhao, Jie Wang, et al. Speech representation disentanglement with adversarial mutual information learning for one-shot voice conversion. arXiv preprint arXiv:2208.08757, 2022.
  48. 48.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180–1189. PMLR, 2015.
  49. 49.Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong. Diffuseq: Sequence to sequence text generation with diffusion models. arXiv preprint arXiv:2210.08933, 2022.
  50. 50.Hyungjin Chung, Jeongsol Kim, Michael T Mccann, Marc L Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022.
  51. 51.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022.
  52. 52.José Lezama, Huiwen Chang, Lu Jiang, and Irfan Essa. Improved masked image generation with token-critic. In European Conference on Computer Vision, pages 70–86. Springer, 2022.
  53. 53.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  54. 54.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  55. 55.Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common diffusion noise schedules and sample steps are flawed. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5404–5411, 2024.
  56. 56.Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. Libri-light: A benchmark for asr with limited or no supervision. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7669–7673. IEEE, 2020.
  57. 57.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. LibriSpeech: an ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210. IEEE, 2015.
  58. 58.Steven R Livingstone and Frank A Russo. The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english. PloS one, 13(5):e0196391, 2018.
  59. 59.Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. Utmos: Utokyo-sarulab system for voicemos challenge 2022. arXiv preprint arXiv:2204.02152, 2022.
  60. 60.Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al. Streaming end-to-end speech recognition for mobile devices. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6381–6385. IEEE, 2019.
  61. 61.Mohammed Salah Al-Radhi, Tamás Gábor Csapó, and Géza Németh. Nonparallel expressive tts for unseen target speaker using style-controlled adaptive layer and optimized pitch embedding. In 2023 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), pages 176–181. IEEE, 2023.
  62. 62.Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Chen Zhang, Zhenhui Ye, Pengfei Wei, Chunfeng Wang, Xiang Yin, Zejun Ma, et al. Mega-tts 2: Zero-shot text-to-speech with arbitrary length speech prompts. arXiv preprint arXiv:2307.07218, 2023.
  63. 63.Hyung-Seok Oh, Sang-Hoon Lee, and Seong-Whan Lee. Diffprosody: Diffusion-based latent prosody generation for expressive speech synthesis with prosody conditional adversarial training. arXiv preprint arXiv:2307.16549, 2023.
  64. 64.Yi Ren, Ming Lei, Zhiying Huang, Shiliang Zhang, Qian Chen, Zhijie Yan, and Zhou Zhao. Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7577–7581. IEEE, 2022.
  65. 65.Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou. Hifi-codec: Group-residual vector quantization for high fidelity audio codec. arXiv preprint arXiv:2305.02765, 2023.
  66. 66.Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu. Token-level ensemble distillation for grapheme-to-phoneme conversion. In INTERSPEECH, 2019.
  67. 67.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11315–11325, 2022.
  68. 68.Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. Bigvgan: A universal neural vocoder with large-scale training. arXiv preprint arXiv:2206.04658, 2022.
  69. 69.Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100, 2020.
  70. 70.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 6309–6318, 2017.
  71. 71.Edresson Casanova, Julian Weber, Christopher D Shulby, Arnaldo Candido Junior, Eren Gölge, and Moacir A Ponti. Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone. In International Conference on Machine Learning, pages 2709–2720. PMLR, 2022.
  72. 72.Zhichao Wang, Yuanzhe Chen, Lei Xie, Qiao Tian, and Yuping Wang. Lm-vc: Zero-shot voice conversion via speech generation based on language models. arXiv preprint arXiv:2306.10521, 2023.

Citation

MLA
Ju, Z., et al. “NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models”. arXiv, 2024, http://arxiv.org/abs/2403.03100v3.
APA
Ju, Z., Wang, Y., Shen, K., Tan, X., Xin, D., Yang, D., Liu, Y., Leng, Y., Song, K., Tang, S., Wu, Z., Qin, T., Li, X.-Y., Ye, W., Zhang, S., Bian, J., He, L., Li, J., & Zhao, S. (2024). NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. arXiv. http://arxiv.org/abs/2403.03100v3
Chicago
Ju, Z., Y. Wang, K. Shen, et al. 2024. “NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models”. arXiv. http://arxiv.org/abs/2403.03100v3.
Harvard
Ju, Z. et al. (2024) “NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.03100v3.
Vancouver
1. Ju Z, Wang Y, Shen K, et al (2024) NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. arXiv

BibTeX

@article{ju2024naturalspeech,
  title = {NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models},
  author = {Ju, Zeqian and Wang, Yuancheng and Shen, Kai and Tan, Xu and Xin, Detai and Yang, Dongchao and Liu, Yanqing and Leng, Yichong and Song, Kaitao and Tang, Siliang and Wu, Zhizheng and Qin, Tao and Li, Xiang-Yang and Ye, Wei and Zhang, Shikun and Bian, Jiang and He, Lei and Li, Jinyu and Zhao, Sheng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.03100v3},
  eprint = {2403.03100}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/