Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Chengyi WangSanyuan ChenYu WuZi-Hua ZhangLong ZhouShujie LiuZhuo ChenYanqing LiuHuaming WangJinyu Li
Introduces VALL-E, a neural codec language model trained on 60,000 hours of speech that performs zero-shot text-to-speech synthesis by replicating an unseen speaker's voice, emotion, and acoustic environment from just a three-second audio prompt.
Current text-to-speech technologies struggle to synthesize natural, personalized voices for new speakers without requiring clean studio recordings, extensive fine-tuning, or complex model engineering. This limitation restricts the flexibility and scalability of speech generation in real-world applications where only brief, uncurated voice samples are available.
The article demonstrates that treating text-to-speech synthesis as a conditional language modeling problem enables high-quality, zero-shot personalized speech generation using just a three-second recording of an unseen speaker. To achieve this, the authors introduce VALL-E, a neural codec language model trained to generate discrete audio tokens conditioned on phoneme text and short acoustic prompts.
The authors pre-trained VALL-E on LibriLight, an English audio dataset comprising 60,000 hours of speech from over 7,000 speakers, representing a dataset hundreds of times larger than conventional speech synthesis corpora. The architecture splits generation into two stages: an autoregressive model that predicts the primary acoustic tokens to establish rhythm and speaker identity, followed by a non-autoregressive model that quickly fills in the remaining acoustic detail tokens. The final audio waveform is reconstructed using an off-the-shelf neural audio codec decoder without requiring specialized vocoder training.
The evaluation yielded several key findings. First, VALL-E significantly outperformed the existing state-of-the-art zero-shot baseline in speaker similarity and speech naturalness, improving similarity mean opinion scores by 0.93 points on LibriSpeech and naturalness by 0.23 points on VCTK. Second, human evaluations indicated that the synthesized voices on the VCTK benchmark were as natural as genuine human recordings. Third, VALL-E successfully preserved the emotional tone and ambient acoustic environment of the prompt, such as background reverberation. Finally, the model's sampling approach generated diverse speech variations from identical inputs, and it reduced transcription word error rates from 7.7% in the baseline down to 5.9%.
These results demonstrate that scaling diverse, semi-supervised data alongside language modeling architectures eliminates the need for complex custom feature engineering in speech synthesis. However, the ability to clone voices from three-second clips introduces notable security risks, including identity theft and voice authentication spoofing. Stakeholders must pair deployment with robust detection systems and adherence to responsible artificial intelligence frameworks.
Organizations developing speech technology should invest in scaled language model architectures for voice synthesis and prioritize building synthetic speech detection models to mitigate misuse risks. Future technical work should focus on expanding training datasets to include diverse global accents and varied speaking styles, as well as refining model architectures to eliminate synthesis errors like duplicated or omitted words.
- Paper: SoundStream: An End-to-End Neural Audio Codec, Neil Zeghidour et al. (2021). SoundStream introduces the neural audio codec architecture with residual vector quantization that enables treating continuous speech as discrete acoustic tokens.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This foundational work introduces vector-quantized autoencoders (VQ-VAE), establishing the core technique of discretizing continuous signals for downstream autoregressive modeling.
- Paper: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech, Jaehyeon Kim et al. (2021). VITS provides the standard benchmark for end-to-end multi-speaker text-to-speech against which discrete neural codec language models are evaluated.
- Paper: Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions, Jonathan Shen et al. (2017). Tacotron 2 established the neural sequence-to-sequence paradigm for text-to-speech synthesis that VALL-E reformulates as an in-context language modeling task.
- Paper: FastSpeech 2: Fast and High-Quality End-to-End Text to Speech, Yi Ren et al. (2020). FastSpeech 2 represents the non-autoregressive continuous regression approach to speech synthesis that VALL-E contrasts with its discrete conditional codec generation.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN provides fundamental insight into neural waveform generation and vocoding used to synthesize high-fidelity speech from acoustic representations.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS builds upon the neural codec language modeling and zero-shot voice cloning paradigm pioneered by VALL-E, scaling it with dual-track streaming architectures.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni extends discrete speech language modeling into an end-to-end multimodal foundation model featuring dedicated speech generation modules.
