Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Chengyi WangSanyuan ChenYu WuZi-Hua ZhangLong ZhouShujie LiuZhuo ChenYanqing LiuHuaming WangJinyu Li

article2023IEEE Transactions on Audio, Speech, and Language Processing1,358 citations

Introduces VALL-E, a neural codec language model trained on 60,000 hours of speech that performs zero-shot text-to-speech synthesis by replicating an unseen speaker's voice, emotion, and acoustic environment from just a three-second audio prompt.

Listen

Current text-to-speech technologies struggle to synthesize natural, personalized voices for new speakers without requiring clean studio recordings, extensive fine-tuning, or complex model engineering. This limitation restricts the flexibility and scalability of speech generation in real-world applications where only brief, uncurated voice samples are available.

The article demonstrates that treating text-to-speech synthesis as a conditional language modeling problem enables high-quality, zero-shot personalized speech generation using just a three-second recording of an unseen speaker. To achieve this, the authors introduce VALL-E, a neural codec language model trained to generate discrete audio tokens conditioned on phoneme text and short acoustic prompts.

The authors pre-trained VALL-E on LibriLight, an English audio dataset comprising 60,000 hours of speech from over 7,000 speakers, representing a dataset hundreds of times larger than conventional speech synthesis corpora. The architecture splits generation into two stages: an autoregressive model that predicts the primary acoustic tokens to establish rhythm and speaker identity, followed by a non-autoregressive model that quickly fills in the remaining acoustic detail tokens. The final audio waveform is reconstructed using an off-the-shelf neural audio codec decoder without requiring specialized vocoder training.

The evaluation yielded several key findings. First, VALL-E significantly outperformed the existing state-of-the-art zero-shot baseline in speaker similarity and speech naturalness, improving similarity mean opinion scores by 0.93 points on LibriSpeech and naturalness by 0.23 points on VCTK. Second, human evaluations indicated that the synthesized voices on the VCTK benchmark were as natural as genuine human recordings. Third, VALL-E successfully preserved the emotional tone and ambient acoustic environment of the prompt, such as background reverberation. Finally, the model's sampling approach generated diverse speech variations from identical inputs, and it reduced transcription word error rates from 7.7% in the baseline down to 5.9%.

These results demonstrate that scaling diverse, semi-supervised data alongside language modeling architectures eliminates the need for complex custom feature engineering in speech synthesis. However, the ability to clone voices from three-second clips introduces notable security risks, including identity theft and voice authentication spoofing. Stakeholders must pair deployment with robust detection systems and adherence to responsible artificial intelligence frameworks.

Organizations developing speech technology should invest in scaled language model architectures for voice synthesis and prioritize building synthetic speech detection models to mitigate misuse risks. Future technical work should focus on expanding training datasets to include diverse global accents and varied speaking styles, as well as refining model architectures to eliminate synthesis errors like duplicated or omitted words.

  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS builds upon the neural codec language modeling and zero-shot voice cloning paradigm pioneered by VALL-E, scaling it with dual-track streaming architectures.
  • Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni extends discrete speech language modeling into an end-to-end multimodal foundation model featuring dedicated speech generation modules.
Cover for Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Abstract

We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train a neural codec language model (called Vall-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS as a conditional language modeling task rather than continuous signal regression as in previous work. During the pre-training stage, we scale up the TTS training data to 60K hours of English speech which is hundreds of times larger than existing systems. Vall-E emerges in-context learning capabilities and can be used to synthesize high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker as an acoustic prompt. Experiment results show that Vall-E significantly outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity. In addition, we find Vall-E could preserve the speaker's emotion and acoustic environment of the acoustic prompt in synthesis. See this https URL for demos of our work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background: Speech Quantization
  • 4 VALL-E
  • 4.1 Problem Formulation: Regarding TTS as Conditional Codec Language Modeling
  • 4.2 Training: Conditional Codec Language Modeling
  • 4.2.1 Autoregressive Codec Language Modeling
  • 4.2.2 Non-Autoregressive Codec Language Modeling
  • 4.3 Inference: In-Context Learning via Prompting
  • 5 Experiment
  • 5.1 Experiment Setup
  • 5.2 LibriSpeech Evaluation
  • 5.3 VCTK Evaluation
  • 5.4 Qualitative Analysis
  • 6 Conclusion, Limitations, and Future Work
  • References

Knowls

  1. Knowl 1 — Hierarchical Conditional Audio Codec Language Modeling for Zero-Shot TTS

    model/method

    VALL-E formulates text-to-speech (TTS) synthesis as a conditional language modeling task over discrete audio representations. Given a target phoneme sequence x=(x0,x1,…,xL)\mathbf{x} = (x_0, x_1, \dots, x_L) and an acoustic prompt matrix C~∈NT0×8\tilde{\mathbf{C}} \in \mathbb{N}^{T_0 \times 8} derived from an enrolled recording of an unseen speaker using an 8-quantizer neural audio codec, the goal is to generate the acoustic code matrix C∈NT×8\mathbf{C} \in \mathbb{N}^{T \times 8} corresponding to the synthesized speech of duration TT.

    Because the neural audio codec uses Residual Vector Quantization (RVQ), the quantizers exhibit a hierarchical structure: the first quantizer captures primary acoustic features and speaker identity, while subsequent quantizers encode fine acoustic residuals. VALL-E models the joint probability p(C∣x,C~)p(\mathbf{C} \mid \mathbf{x}, \tilde{\mathbf{C}}) by factorizing it into an autoregressive (AR) stage for the first codebook sequence c:,1\mathbf{c}_{:,1} and a non-autoregressive (NAR) stage for the remaining codebook sequences c:,2:8\mathbf{c}_{:,2:8}:

    p(C∣x,C~;θ)=p(c:,1∣c~:,1,x;θAR)∏j=28p(c:,j∣C:,<j,x,C~;θNAR)p(\mathbf{C} \mid \mathbf{x}, \tilde{\mathbf{C}}; \theta) = p(\mathbf{c}_{:,1} \mid \tilde{\mathbf{c}}_{:,1}, \mathbf{x}; \theta_{\text{AR}}) \prod_{j=2}^8 p(\mathbf{c}_{:,j} \mid \mathbf{C}_{:,<j}, \mathbf{x}, \tilde{\mathbf{C}}; \theta_{\text{NAR}})

    where c:,j∈NT\mathbf{c}_{:,j} \in \mathbb{N}^T denotes the sequence of discrete tokens for the jj-th quantizer, C:,<j=[c:,1,…,c:,j−1]\mathbf{C}_{:,<j} = [\mathbf{c}_{:,1}, \dots, \mathbf{c}_{:,j-1}] represents the acoustic codes from preceding quantizers, and θAR\theta_{\text{AR}} and θNAR\theta_{\text{NAR}} represent the parameters of the autoregressive and non-autoregressive language models, respectively. The final acoustic matrix C\mathbf{C} is converted directly into an audio waveform using the neural codec decoder without an auxiliary vocoder.

  2. Knowl 2 — Autoregressive Codec Language Model for First-Quantizer Code Generation

    model/method

    The first stage of VALL-E generates the primary acoustic code sequence c:,1=(c1,1,c2,1,…,cT,1)\mathbf{c}_{:,1} = (c_{1,1}, c_{2,1}, \dots, c_{T,1}) using an autoregressive (AR) decoder-only Transformer. The model is conditioned on the phoneme sequence x=(x0,…,xL)\mathbf{x} = (x_0, \dots, x_L) and the first-quantizer prompt sequence c~:,1=(c~1,1,…,c~T0,1)\tilde{\mathbf{c}}_{:,1} = (\tilde{c}_{1,1}, \dots, \tilde{c}_{T_0,1}) from the enrolled speech.

    The autoregressive factorization is defined as:

    p(c:,1∣x,c~:,1;θAR)=∏t=1Tp(ct,1∣c<t,1,c~:,1,x;θAR)p(\mathbf{c}_{:,1} \mid \mathbf{x}, \tilde{\mathbf{c}}_{:,1}; \theta_{\text{AR}}) = \prod_{t=1}^T p(c_{t,1} \mid \mathbf{c}_{<t,1}, \tilde{\mathbf{c}}_{:,1}, \mathbf{x}; \theta_{\text{AR}})

    The model input is structured by concatenating the phoneme embedding sequence and the acoustic token embedding sequence, appending a special <EOS> token to each segment. Sinusoidal position embeddings are computed separately for the phoneme prompt and the acoustic sequence. Causal self-attention allows each token ct,1c_{t,1} to attend to all phonemes x\mathbf{x} and preceding acoustic tokens [c~:,1,c<t,1][\tilde{\mathbf{c}}_{:,1}, \mathbf{c}_{<t,1}]. The weights of the linear output projection layer are tied to the acoustic embedding matrix Wa\mathbf{W}_a.

    During training, no separate prompt segment is explicitly extracted; any prefix c<t,1\mathbf{c}_{<t,1} naturally serves as the acoustic prompt for predicting the continuation c≥t,1\mathbf{c}_{\ge t,1} under standard causal language modeling objectives.

  3. Knowl 3 — Non-Autoregressive Residual Codec Language Model with Adaptive Layer Normalization

    model/method

    The second stage of VALL-E generates the fine acoustic codes for quantizers j∈{2,…,8}j \in \{2, \dots, 8\} conditioned on the phoneme sequence x\mathbf{x}, the full prompt matrix C~∈NT0×8\tilde{\mathbf{C}} \in \mathbb{N}^{T_0 \times 8}, and the acoustic codes from all preceding quantizers C:,<j∈NT×(j−1)\mathbf{C}_{:,<j} \in \mathbb{N}^{T \times (j-1)}. The non-autoregressive (NAR) model generates all TT time steps in parallel for a given quantizer level jj:

    p(C:,2:8∣x,C~;θNAR)=∏j=28p(c:,j∣C:,<j,x,C~;θNAR)p(\mathbf{C}_{:,2:8} \mid \mathbf{x}, \tilde{\mathbf{C}}; \theta_{\text{NAR}}) = \prod_{j=2}^8 p(\mathbf{c}_{:,j} \mid \mathbf{C}_{:,<j}, \mathbf{x}, \tilde{\mathbf{C}}; \theta_{\text{NAR}})

    For a selected quantizer stage i∈[2,8]i \in [2, 8], the acoustic representations of preceding stages are summed across codebooks:

    ect,j=Waj⊙ct,j,ect=∑j=1i−1ect,j\mathbf{e}_{c_t, j} = \mathbf{W}_a^j \odot c_{t, j}, \quad \mathbf{e}_{c_t} = \sum_{j=1}^{i-1} \mathbf{e}_{c_t, j}

    where Waj\mathbf{W}_a^j is the embedding matrix for the jj-th codebook, and ⊙\odot denotes index selection. The acoustic prompt is formed by summing the embeddings across all 8 codebooks: e~ct=∑j=18Waj⊙c~t,j\tilde{\mathbf{e}}_{c_t} = \sum_{j=1}^8 \mathbf{W}_a^j \odot \tilde{c}_{t,j}.

    The input sequence to the bidirectional Transformer decoder is the concatenation (ex,e~c,ec:,<i)(\mathbf{e}_x, \tilde{\mathbf{e}}_c, \mathbf{e}_{c_{:,<i}}). The stage index ii is conditioned via Adaptive Layer Normalization (AdaLN):

    AdaLN(h,i)=ai⊙LayerNorm(h)+bi\text{AdaLN}(h, i) = a_i \odot \text{LayerNorm}(h) + b_i

    where hh is the intermediate activation vector, and ai,bia_i, b_i are obtained from a linear projection of the stage embedding. The parameters of the jj-th prediction layer are shared with the (j+1)(j+1)-th acoustic embedding layer.

  4. Knowl 4 — VALL-E Zero-Shot In-Context Speech Synthesis

    algorithm

    The inference procedure synthesizes personalized speech for an unseen speaker by conditioning on a 3-second acoustic prompt and text transcriptions using sampling and greedy decoding across two stages.

    Input: Target text phoneme sequence xtarget=(x1,…,xL)x_{target} = (x_1, \dots, x_L)
    Input: Enrolled acoustic prompt audio y~\tilde{y}
    Input: Enrolled prompt phoneme sequence xprompt=(x1′,…,xL′′)x_{prompt} = (x'_1, \dots, x'_{L'})
    Input: Pre-trained EnCodec encoder EE and decoder DD
    Input: Trained AR model θAR\theta_{AR} and NAR model θNAR\theta_{NAR}
    Output: Synthesized waveform y^\hat{y}
    C~←E(y~)\tilde{C} \leftarrow E(\tilde{y}) # Shape: T0×8T_0 \times 8
    c~:,1←C~[:,1]\tilde{c}_{:,1} \leftarrow \tilde{C}[:, 1]
    xfull←[xprompt;xtarget;<EOS>]x_{full} \leftarrow [x_{prompt}; x_{target}; \text{<EOS>}]
    # Stage 1: Autoregressive generation of first-level acoustic codes
    c:,1←[c~:,1]c_{:,1} \leftarrow [\tilde{c}_{:,1}]
    repeat
        t←length(c:,1)t \leftarrow \text{length}(c_{:,1})
        P(ct+1,1)←θAR(xfull,c:,1)P(c_{t+1,1}) \leftarrow \theta_{AR}(x_{full}, c_{:,1})
        Sample ct+1,1∼P(ct+1,1)c_{t+1,1} \sim P(c_{t+1,1}) using sampling-based decoding
        c:,1←[c:,1;ct+1,1]c_{:,1} \leftarrow [c_{:,1}; c_{t+1,1}]
    until ct+1,1==<EOS>c_{t+1,1} == \text{<EOS>}
    c:,1←cT0+1:T,1c_{:,1} \leftarrow c_{T_0+1 : T, 1} # Exclude prompt prefix
    # Stage 2: Non-autoregressive generation of residual codes (stages 2 to 8)
    C←zeros(T,8)C \leftarrow \text{zeros}(T, 8)
    C[:,1]←c:,1C[:, 1] \leftarrow c_{:,1}
    for j←2j \leftarrow 2 to 88 do
        P(c:,j)←θNAR(xtarget,C~,C[:,1:j−1],stage=j)P(c_{:,j}) \leftarrow \theta_{NAR}(x_{target}, \tilde{C}, C[:, 1:j-1], \text{stage}=j)
        C[:,j]←arg⁡max⁡P(c:,j)C[:, j] \leftarrow \arg\max P(c_{:,j}) # Greedy decoding
    end for
    # Waveform reconstruction
    y^←D(C)\hat{y} \leftarrow D(C)
    return y^\hat{y}
  5. Knowl 5 — VALL-E Pre-Training Configuration and Architecture Setup

    experimental setup

    VALL-E is pre-trained on the LibriLight corpus, containing approximately 60,000 hours of unlabelled English audiobook speech across over 7,000 distinct speakers. Phoneme transcriptions are generated using a Kaldi-based hybrid DNN-HMM ASR model trained on 960 hours of LibriSpeech with a 30ms frameshift alignment, followed by removal of consecutive repetitions.

    Audio tokenization utilizes EnCodec at 24 kHz audio sampling with 8 Residual Vector Quantization (RVQ) quantizers of codebook size 1024 each, operating at a 75 Hz frame rate (a 320-fold sampling reduction) and a 6 kbps bitrate. Utterances from LibriLight are cropped to random durations between 10 and 20 seconds for training. For the NAR model, a separate random 3-second segment from the same utterance serves as the acoustic prompt.

    Both the AR and NAR Transformer decoders share identical backbone dimensions:

    • Number of Transformer layers: 12
    • Number of attention heads: 16
    • Hidden embedding dimension: 1024
    • Feed-forward network (FFN) dimension: 4096
    • Dropout: 0.1

    Optimization is performed on 16 NVIDIA Tesla V100 32GB GPUs with a batch size of 6,000 acoustic tokens per GPU for 800,000 steps using AdamW. The learning rate is warmed up over the first 32,000 updates to a peak of 5×10−45 \times 10^{-4} before linear decay.

  6. Knowl 6 — Zero-Shot TTS Performance on LibriSpeech Clean Test Set

    data/table

    Zero-shot evaluation was conducted on a 2.2-hour subset of the LibriSpeech test-clean dataset (utterance lengths between 4 and 10 seconds), where all test speakers are unseen during LibriLight pre-training. VALL-E used a randomly chosen 3-second segment of another utterance from the same speaker as the enrolled prompt, whereas VALL-E-continual used the initial 3 seconds of the ground-truth utterance.

    Objective metrics comprise Word Error Rate (WER) computed via HuBERT-Large CTC (fine-tuned on LibriSpeech 960h without LM fusion) and speaker similarity (SPK) predicted via WavLM-TDNN in the range [−1,1][-1, 1]. Subjective metrics include Comparative Mean Opinion Score (CMOS, −3-3 to +3+3) and Similarity Mean Opinion Score (SMOS, 1 to 5) rated by native crowdsourced evaluators across 40 test cases.

    Model WER (%) SPK SMOS CMOS (v.s. VALL-E)
    GroundTruth 2.2 0.754 4.50 ±\pm 0.10 +0.17
    Speech-to-Speech Baselines
    GSLM 12.4 0.126 - -
    AudioLM 6.0 - - -
    TTS Systems
    YourTTS 7.7 0.337 3.45 ±\pm 0.09 -0.12
    VALL-E 5.9 0.580 4.38 ±\pm 0.10 0.00
    VALL-E-continual 3.8 0.508 - -

    VALL-E achieves a lower WER (5.9% vs. 7.7%) and a substantially higher speaker similarity score (0.580 vs. 0.337) compared to the YourTTS baseline. In human evaluations, VALL-E outperforms YourTTS by +0.93 in SMOS and +0.12 in CMOS naturalness.

  7. Knowl 7 — Speaker Similarity Scaling and Zero-Shot Evaluation on VCTK

    data/table

    Zero-shot speaker cloning was evaluated on the VCTK dataset containing 108 speakers unseen by VALL-E. For comparison, the YourTTS baseline had seen 97 of these speakers during training. Evaluations were divided into the full 108-speaker set and the 11 strictly unseen speakers for YourTTS across prompt durations of 3s, 5s, and 10s using WavLM-TDNN for automatic speaker similarity. Human evaluations (SMOS and CMOS) were performed on 60 speakers (11 unseen, 49 seen by YourTTS) with 3-second prompts.

    Model 3s Prompt 5s Prompt 10s Prompt
    108 Full Speakers
    YourTTS* 0.357 0.377 0.394
    VALL-E 0.382 0.423 0.484
    GroundTruth 0.546 0.591 0.620
    11 Unseen Speakers
    YourTTS 0.331 0.337 0.344
    VALL-E 0.389 0.380 0.414
    GroundTruth 0.528 0.556 0.586
    Model SMOS CMOS (v.s. VALL-E)
    YourTTS* 3.70 ±\pm 0.09 -0.23
    VALL-E 3.81 ±\pm 0.09 0.00
    GroundTruth 4.29 ±\pm 0.09 -0.04

    VALL-E outperforms YourTTS on speaker similarity across all prompt lengths even when YourTTS observed the speakers during training. In human evaluation, VALL-E exceeds YourTTS by +0.11 SMOS and +0.23 CMOS, while achieving a CMOS of +0.04 over ground-truth audio, indicating synthesis naturalness on par with human recordings.

  8. Knowl 8 — Ablation Analysis of Conditioning Prompts in AR and NAR Models

    data/table

    Ablations were conducted on the LibriSpeech test set to assess the distinct roles of the phoneme prompt and the acoustic prompt across both the non-autoregressive (NAR) and autoregressive (AR) models. For the NAR model ablation, ground-truth first-level acoustic codes were provided as input.

    Metric NAR-no prompt NAR-phn prompt NAR-2 prompts
    WER (%) 19.6 3.0 2.8
    SPK 0.518 0.541 0.732
    Model Setting WER (%) SPK
    VALL-E (Full AR + NAR-2 prompts) 5.9 0.585
    w/o acoustic prompt in AR 5.9 0.236

    The NAR ablation shows that adding the phoneme prompt is critical for speech content accuracy, dropping the WER from 19.6% to 3.0%. Adding the acoustic prompt increases speaker similarity substantially from 0.541 to 0.732. The AR ablation indicates that removing the acoustic prompt from the AR stage degrades speaker similarity (SPK drops from 0.585 to 0.236) without affecting WER, demonstrating that speaker conditioning in the AR stage is essential for overall voice cloning fidelity.

  9. Knowl 9 — Acoustic Environment, Emotion Preservation, and Output Diversity

    empirical result

    Beyond standard zero-shot TTS, VALL-E exhibits three emergent in-context capabilities:

    1. Acoustic Environment Consistency: When provided with an acoustic prompt containing ambient reverberation or room characteristics, VALL-E synthesizes output speech preserving the same reverberation profile rather than outputting clean speech, owing to pre-training on diverse, non-studio conditions.
    2. Zero-Shot Emotion Preservation: When evaluated using acoustic prompts from EmoV-DB containing various emotional states (e.g., anger, sleepiness), VALL-E preserves the prompt's emotion in synthesized speech without requiring supervised emotional labels or fine-tuning.
    3. Inference Diversity: Because VALL-E generates first-stage acoustic tokens using sampling-based decoding, multiple runs on the same input text and acoustic prompt yield diverse outputs with variations in phrase duration, speaking rates, stress patterns, and pitch contours.
  10. Knowl 10 — Synthesis Robustness and Data Coverage Limitations of VALL-E

    limitation

    VALL-E exhibits two primary limitations:

    1. Synthesis Robustness Issues: Due to the autoregressive attention alignment in the first-stage Transformer decoder, the model occasionally suffers from disordered attention alignments resulting in unclear pronunciations, omitted words, or repeated phrases.
    2. Speaking Style and Accent Coverage: Pre-training is restricted to the LibriLight corpus, which consists predominantly of English audiobook readings. Consequently, performance degrades on accented speakers (evidenced by lower speaker similarity on VCTK compared to LibriSpeech) and exhibits limited coverage of expressive or conversational speaking styles.

Coverage note — None was omitted; all key architectural components, formulations, pre-training details, quantitative benchmarks, ablations, qualitative findings, and stated limitations are fully captured.

References

  1. 1.Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit. The emotional voices database: Towards controlling the emotion dimension in voice generation systems. arXiv preprint arXiv:1806.09514, 2018.
  2. 2.Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, et al. Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5723–5738, 2022.
  3. 3.Sercan Ömer Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou. Neural voice cloning with a few samples. In NeurIPS, pages 10040–10050, 2018.
  4. 4.Alexei Baevski, Steffen Schneider, and Michael Auli. vq-wav2vec: Self-supervised learning of discrete speech representations. In ICLM, 2020a.
  5. 5.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. NeurIPS, 33:12449–12460, 2020b.
  6. 6.He Bai, Renjie Zheng, Junkun Chen, Mingbo Ma, Xintong Li, and Liang Huang. A3t: Alignment-aware acoustic and text pretraining for speech synthesis and editing. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 1399–1411. PMLR, 2022.
  7. 7.Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matthew Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. Audiolm: a language modeling approach to audio generation. CoRR, abs/2209.03143, 2022.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020.
  9. 9.Weicheng Cai, Jinkun Chen, and Ming Li. Exploring the encoding layer and loss function in end-to-end speaker and language recognition system. In Odyssey 2018: The Speaker and Language Recognition Workshop, 26-29 June 2018, Les Sables d’Olonne, France, pages 74–81. ISCA, 2018.
  10. 10.Edresson Casanova, Arnaldo Cândido Júnior, Christopher Shulby, Frederico Santos de Oliveira, João Paulo Ramos Teixeira, Moacir Antonelli Ponti, and Sandra M. Aluísio. Tts-portuguese corpus: a corpus for speech synthesis in brazilian portuguese. Lang. Resour. Evaluation, 56(3):1043–1055, 2022a.
  11. 11.Edresson Casanova, Julian Weber, Christopher D Shulby, Arnaldo Candido Junior, Eren Gölge, and Moacir A Ponti. Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone. In ICML, pages 2709–2720. PMLR, 2022b.
  12. 12.Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu. Adaspeech: Adaptive text to speech for custom voice. In ICLR, 2021.
  13. 13.Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6): 1505–1518, 2022.
  14. 14.Yutian Chen, Yannis M. Assael, Brendan Shillingford, David Budden, Scott E. Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Çaglar Gülçehre, Aäron van den Oord, Oriol Vinyals, and Nando de Freitas. Sample efficient adaptive text-to-speech. In ICLR ,, 2019.
  15. 15.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311, 2022.
  16. 16.Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and R. J. Skerry-Ryan. Semi-supervised training for improving data efficiency in end-to-end speech synthesis. In ICASSP, pages 6940–6944. IEEE, 2018.
  17. 17.Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression. arXiv preprint arXiv:2210.13438, 2022.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186, 2019.
  19. 19.Chenpeng Du, Yiwei Guo, Xie Chen, and Kai Yu. VQTTS: high-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature. In Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, pages 1596–1600. ISCA, 2022. doi: 10.21437/Interspeech.2022-489.
  20. 20.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  21. 21.Sung-Feng Huang, Chyi-Jiunn Lin, Da-Rong Liu, Yi-Chen Chen, and Hung-yi Lee. Meta-tts: Meta-learning for few-shot speaker adaptive text-to-speech. IEEE ACM Trans. Audio Speech Lang. Process., 30:1558–1571, 2022.
  22. 22.Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez-Moreno, and Yonghui Wu. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. In NeurIPS, pages 4485–4495, 2018.
  23. 23.Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al. Libri-light: A benchmark for asr with limited or no supervision. In ICASSP, pages 7669–7673. IEEE, 2020.
  24. 24.Minki Kang, Dongchan Min, and Sung Ju Hwang. Any-speaker adaptive text-to-speech synthesis with diffusion models. CoRR, abs/2211.09383, 2022. doi: 10.48550/arXiv.2211.09383.
  25. 25.Heeseung Kim, Sungwon Kim, and Sungroh Yoon. Guided-tts: A diffusion model for text-to-speech via classifier guidance. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 11119–11133. PMLR, 2022.
  26. 26.Jaehyeon Kim, Jungil Kong, and Juhee Son. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In ICML, volume 139 of Proceedings of Machine Learning Research, pages 5530–5540. PMLR, 2021.
  27. 27.Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. In NeurIPS, 2020.
  28. 28.Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux. Generative spoken language modeling from raw audio. CoRR, abs/2102.01192, 2021.
  29. 29.Yi Lei, Shan Yang, and Lei Xie. Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis. In 2021 IEEE Spoken Language Technology Workshop (SLT), pages 423–430. IEEE, 2021.
  30. 30.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. Neural speech synthesis with transformer network. In AAAI, pages 6706–6713. AAAI, 2019.
  31. 31.Yanqing Liu, Ruiqing Xue, Lei He, Xu Tan, and Sheng Zhao. Delightfultts 2: End-to-end speech synthesis with adversarial vector-quantized auto-encoders. In Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, pages 1581–1585. ISCA, 2022. doi: 10.21437/Interspeech.2022-277.
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  33. 33.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: an asr corpus based on public domain audio books. In ICASSP, pages 5206–5210. IEEE, 2015.
  34. 34.Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. Speech resynthesis from discrete disentangled self-supervised representations. In Interspeech, pages 3615–3619. ISCA, 2021.
  35. 35.Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail A. Kudinov. Grad-tts: A diffusion probabilistic model for text-to-speech. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8599–8608. PMLR, 2021. URL http://proceedings.mlr.press/v139/popov21a.html.
  36. 36.Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. The kaldi speech recognition toolkit. In IEEE 2011 workshop on automatic speech recognition and understanding, number CONF. IEEE Signal Processing Society, 2011.
  37. 37.Ryan Prenger, Rafael Valle, and Bryan Catanzaro. Waveglow: A flow-based generative network for speech synthesis. In ICASSP, pages 3617–3621. IEEE, 2019.
  38. 38.Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. Fastspeech: Fast, robust and controllable text to speech. In NeurIPS, pages 3165–3174, 2019.
  39. 39.Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ-Skerrv Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu. Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions. In ICASSP, pages 4779–4783. IEEE, 2018.
  40. 40.Xu Tan, Tao Qin, Frank K. Soong, and Tie-Yan Liu. A survey on neural speech synthesis. CoRR, abs/2106.15561, 2021.
  41. 41.Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura. VQVAE unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019. In Interspeech, pages 1118–1122. ISCA, 2019.
  42. 42.Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. In The 9th ISCA Speech Synthesis Workshop, page 125. ISCA, 2016.
  43. 43.Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6306–6315, 2017.
  44. 44.Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit. 2016.
  45. 45.Tao Wang, Jianhua Tao, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, and Rongxiu Zhong. Spoken content and voice factorization for few-shot speaker adaptation. In Interspeech, pages 796–800. ISCA, 2020.
  46. 46.Yihan Wu, Xu Tan, Bohan Li, Lei He, Sheng Zhao, Ruihua Song, Tao Qin, and Tie-Yan Liu. Adaspeech 4: Adaptive text to speech in zero-shot scenarios. In Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, pages 2568–2572. ISCA, 2022. doi: 10.21437/Interspeech.2022-901.
  47. 47.Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin. Understanding and improving layer normalization. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 4383–4393, 2019.
  48. 48.Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. Soundstream: An end-to-end neural audio codec. IEEE ACM Trans. Audio Speech Lang. Process., 30: 495–507, 2022.
  49. 49.Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. Libritts: A corpus derived from librispeech for text-to-speech. In Interspeech, pages 1526–1530. ISCA, 2019.

Citation

MLA
Wang, C., et al. “Neural Codec Language Models Are Zero-Shot Text to Speech Synthesizers”. arXiv, 2023, http://arxiv.org/abs/2301.02111v1.
APA
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., & Wei, F. (2023). Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. arXiv. http://arxiv.org/abs/2301.02111v1
Chicago
Wang, C., S. Chen, Y. Wu, et al. 2023. “Neural Codec Language Models Are Zero-Shot Text to Speech Synthesizers”. arXiv. http://arxiv.org/abs/2301.02111v1.
Harvard
Wang, C. et al. (2023) “Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.02111v1.
Vancouver
1. Wang C, Chen S, Wu Y, et al (2023) Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. arXiv

BibTeX

@article{wang2023neural,
  title = {Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers},
  author = {Wang, Chengyi and Chen, Sanyuan and Wu, Yu and Zhang, Ziqiang and Zhou, Long and Liu, Shujie and Chen, Zhuo and Liu, Yanqing and Wang, Huaming and Li, Jinyu and He, Lei and Zhao, Sheng and Wei, Furu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.02111v1},
  eprint = {2301.02111}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF