NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Zeqian JuYuancheng WangKai ShenXu TanDetai XinDongchao YangEric LiuYichong LengKaitao SongSiliang Tang
Presents NaturalSpeech 3, a zero-shot speech synthesis system that disentangles speech into content, prosody, timbre, and acoustic subspaces using factorized diffusion models, achieving audio quality and speaker similarity on par with human recordings.
Generating natural, highly similar human speech from a short reference prompt without prior training on the target voice has been a major challenge in artificial intelligence. While recent text-to-speech systems have scaled up model sizes and training data, they frequently fall short in voice naturalness, prosodic expressiveness, and speaker similarity. This shortfall occurs because speech is inherently complex, entangling multiple attributes—such as linguistic content, pitch rhythm, acoustic nuances, and speaker identity—into a single waveform that monolithic systems struggle to model cleanly.
The main objective of the article is to develop and evaluate NaturalSpeech 3, a text-to-speech system that factorizes speech into disentangled sub-attributes and generates them separately using discrete diffusion models to achieve human-level zero-shot synthesis.
The researchers developed an end-to-end framework comprising two core components: a neural speech codec called FACodec and a factorized diffusion model. FACodec decomposes speech waveforms into distinct representations of content, prosody, acoustic details, and timbre using specialized techniques like low-dimensional information bottlenecks and gradient reversal to prevent attribute leakage. The factorized diffusion model then generates duration, pitch contours, phonemes, and acoustic textures in sequence using prompts for in-context learning. The system was trained on standard benchmarks, including the 60,000-hour Libri-Light dataset, and scaled up to 1 billion parameters on a 200,000-hour corpus. Evaluation was conducted across standardized objective metrics and subjective human listening tests on multi-speaker datasets.
The evaluation produced several notable findings. First, NaturalSpeech 3 achieved human-level voice quality and naturalness on the multi-speaker LibriSpeech test benchmark, matching ground-truth human recordings in subjective scores while achieving an improved word error rate of 1.81% compared to 1.94% for human audio. Second, the system established new state-of-the-art benchmarks in speaker similarity, achieving a 0.67 objective similarity score and outperforming existing leading baselines. Third, on emotional speech tests, the model substantially improved prosodic similarity, reducing speech distortion metrics across eight distinct emotions and increasing emotion recognition accuracy from roughly 30–40% to 52%. Finally, data and model scaling demonstrated consistent improvements in speech accuracy and speaker fidelity, while reducing inference latency by more than 15 times compared to leading autoregressive baselines.
These findings indicate that factorizing speech attributes into modular components significantly simplifies generative modeling, resolving the traditional trade-off between output quality and generation speed. The approach also enables zero-shot speech attribute manipulation—such as adjusting speaking rate or applying one speaker's timbre to another's expressive prosody—without retraining. For practical deployments, the high efficiency and modularity lower operational computational costs while expanding capabilities in digital assistants, voice localization, and personalized audio production.
Organizations developing or deploying speech synthesis should adopt factorized attribute architectures to improve controllability, speed, and voice fidelity. As recommended next steps, teams should explore scaling factorized architectures into larger foundational models and evaluate deploying fast single-step diffusion variants where real-time execution is critical. Additionally, because high-fidelity zero-shot voice cloning increases the risk of voice impersonation and identity spoofing, stakeholders must invest in robust synthetic speech detection and misuse-reporting protocols alongside deployment.
These conclusions are supported with high confidence by extensive benchmark evaluations, but several limitations remain. NaturalSpeech 3 was trained primarily on English audiobook datasets, meaning performance has not yet been demonstrated across multiple languages or diverse real-world acoustic backgrounds. Furthermore, the underlying speech codec still relies on supervised phoneme transcriptions during training, which creates an annotation bottleneck for scaling to low-resource languages.
- Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers, Chengyi Wang et al. (2023). VALL-E pioneered large-scale zero-shot neural codec language modeling for speech synthesis, establishing the baseline paradigm that NaturalSpeech 3 enhances through factorization.
- Paper: Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale, Matthew Le et al. (2023). Voicebox demonstrated large-scale non-autoregressive speech generation via flow-matching/diffusion, providing key context for non-autoregressive zero-shot TTS architectures.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational paper introduced vector-quantized autoencoders (VQ-VAE), which serve as the direct conceptual precursor to NaturalSpeech 3's factorized vector quantization (FVQ) neural codec.
- Paper: FastSpeech 2: Fast and High-Quality End-to-End Text to Speech, Yi Ren et al. (2020). FastSpeech 2 established explicit modeling and disentanglement of variance attributes like pitch, duration, and energy, motivating NaturalSpeech 3's attribute-factorized generation.
- Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). DiffWave provides foundational principles for applying non-autoregressive diffusion probabilistic models to audio and speech synthesis.
- Paper: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, Jaehyeon Kim et al. (2021). VITS introduced end-to-end variational latent modeling and stochastic duration prediction for expressive speech synthesis, informing subsequent generative TTS designs.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper formalizes the unified design space, preconditioning, and sampling strategies for diffusion models used across modern generative architectures.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS advances large-scale zero-shot speech synthesis, voice cloning, and streaming tokenization, continuing the progression toward ultra-natural, multi-speaker zero-shot TTS systems.
