Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Jaehyeon KimJungil KongJuhee Son

article2021ICML1,392 citations

Proposes a fully end-to-end text-to-speech architecture that combines variational autoencoders, normalizing flows, and adversarial training with a stochastic duration predictor to generate speech with natural variations in rhythm and pitch that rival ground-truth recordings.

Listen

Current text-to-speech technologies commonly rely on multi-stage pipelines that first predict intermediate audio features, such as spectrograms, and then convert those features into raw sound waves. While effective, this disjointed workflow complicates model training and prevents systems from sharing learned representations across stages. Conversely, earlier single-stage end-to-end models have struggled to match the naturalness of multi-stage systems and often produce flat, monotonous speech by failing to capture the diverse rhythms and pitches of real human speech.

The article demonstrates an end-to-end text-to-speech architecture, named VITS, that generates raw waveforms directly from text in a single stage while operating in parallel for high-speed synthesis. The primary objective is to evaluate whether integrating variational autoencoders, adversarial training, and a flow-based stochastic duration predictor can surpass conventional two-stage systems in voice quality and expressiveness.

The researchers evaluated this framework through comparative benchmark experiments on two standard English audio datasets: LJ Speech, a single-speaker corpus containing approximately 24 hours of audio, and VCTK, a multi-speaker corpus containing about 44 hours of recordings from 109 native speakers. The system links text conditioning to audio generation through latent variables, utilizing dynamic programming to align phonemes to audio and adversarial discriminators to refine waveform detail. Model quality was measured using crowd-sourced subjective naturalness evaluations on a standard 1-to-5 mean opinion score scale, side-by-side comparison tests, and synthesis speed benchmarks across competitive baselines.

The evaluation produced several key findings: First, the proposed single-stage model achieved a mean opinion score of 4.43 on single-speaker speech, outperforming the leading multi-stage baseline of 4.32 and closely matching ground-truth human recordings at 4.46. Second, in multi-speaker evaluations, the system achieved a score of 4.38, matching the human baseline score of 4.38 and substantially exceeding previous multi-stage models, which scored between 3.19 and 3.82. Third, the system synthesized audio at 67.12 times real-time speed, more than doubling the generation speed of leading two-stage parallel pipelines (27.48 times real-time). Finally, the stochastic duration predictor demonstrated the ability to produce varied speech rhythms and pitch contours, whereas deterministic predictors generated rigid, fixed-duration samples.

These results establish that eliminating intermediate acoustic feature extraction simplifies system deployment without compromising output quality. By consolidating training into a single pipeline, organizations can lower maintenance overhead, avoid fragile multi-step fine-tuning processes, and cut inference latency in customer-facing and interactive voice platforms.

Decision-makers and engineering teams looking to modernize speech infrastructure should consider adopting end-to-end latent variable architectures for voice generation and multi-speaker voice conversion workflows. When choosing duration predictors, organizations must weigh the slight speed advantage of deterministic modules against the superior naturalness and rhythmic variation provided by stochastic predictors. Before fully transitioning production systems, engineering teams should evaluate language representation methods to automate the separate text-to-phoneme preprocessing steps still required by this pipeline.

Confidence in these findings is strong for standard English single- and multi-speaker synthesis based on the provided experimental scope and statistical confidence intervals. However, side-by-side comparative evaluations revealed a minor human listener preference for natural ground-truth recordings over synthesized audio (-0.106 on LJ Speech and -0.262 on VCTK). Practical deployments will also need to account for real-world speech conditions, varied acoustic environments, and performance across languages not evaluated in the article.

  • Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). It explores non-autoregressive audio synthesis using diffusion models as an alternative generative paradigm to flow-augmented VAEs and GANs.
  • Paper: SoundStream: An End-to-End Neural Audio Codec, Neil Zeghidour et al. (2021). It extends end-to-end neural audio modeling into discrete residual vector-quantized representations with adversarial training for streaming codecs.
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). It scales end-to-end speech generation to large multi-track language models, extending high-fidelity neural speech synthesis to zero-shot multilingual settings.
Cover for Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Abstract

Several recent end-to-end text-to-speech (TTS) models enabling single-stage training and parallel sampling have been proposed, but their sample quality does not match that of two-stage TTS systems. In this work, we present a parallel end-to-end TTS method that generates more natural sounding audio than current two-stage models. Our method adopts variational inference augmented with normalizing flows and an adversarial training process, which improves the expressive power of generative modeling. We also propose a stochastic duration predictor to synthesize speech with diverse rhythms from input text. With the uncertainty modeling over latent variables and the stochastic duration predictor, our method expresses the natural one-to-many relationship in which a text input can be spoken in multiple ways with different pitches and rhythms. A subjective human evaluation (mean opinion score, or MOS) on the LJ Speech, a single speaker dataset, shows that our method outperforms the best publicly available TTS systems and achieves a MOS comparable to ground truth.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Variational Inference
  • 2.1.1 Overview
  • 2.1.2 Reconstruction loss
  • 2.1.3 KL-divergence
  • 2.2 Alignment Estimation
  • 2.2.1 Monotonic alignment search
  • 2.2.2 Duration prediction from text
  • 2.3 Adversarial Training
  • 2.4 Final Loss
  • 2.5 Model Architecture
  • 2.5.1 Posterior Encoder
  • 2.5.2 Prior Encoder
  • 2.5.3 Decoder
  • 2.5.4 Discriminator
  • 2.5.5 Stochastic Duration Predictor
  • 3 Experiments
  • 3.1 Datasets
  • 3.2 Preprocessing
  • 3.3 Training
  • 3.4 Experimental Setup for Comparison
  • 4 Results
  • 4.1 Speech Synthesis Quality
  • 4.2 Generalization to Multi-Speaker Text-to-Speech
  • 4.3 Speech Variation
  • 4.4 Synthesis Speed
  • 5 Related Work
  • 5.1 End-to-End Text-to-Speech
  • 5.2 Variational Autoencoders
  • 5.3 Duration Prediction in Non-Autoregressive Text-to-Speech
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — VITS Conditional Variational Autoencoder Framework

    model/method

    Variational Inference with adversarial learning for end-to-end Text-to-Speech (VITS) is a single-stage, parallel text-to-speech architecture formulated as a conditional Variational Autoencoder (VAE) that generates raw waveforms directly from phoneme sequences.

    The framework connects text processing to waveform synthesis through continuous latent variables zz. During training, an approximate posterior encoder qϕ(z∣xlin)q_\phi(z|x_{\text{lin}}) parameterizes a factorized normal distribution over zz conditioned on a high-resolution linear-scale spectrogram xlinx_{\text{lin}} of target speech. A prior network constructs a conditional prior distribution pθ(z∣ctext,A)p_\theta(z|c_{\text{text}}, A) parameterized by phoneme sequence ctextc_{\text{text}} and a hard monotonic alignment matrix A∈{0,1}∣ctext∣×∣z∣A \in \{0, 1\}^{|c_{\text{text}}| \times |z|}. To increase the expressive capacity beyond a standard factorized Gaussian, a normalizing flow fθf_\theta transforms zz into a more flexible distribution.

    A decoder GG maps latent representations zz to raw audio waveforms y^\hat{y}. Alignment AA between phonemes and latent time frames is computed dynamically during training by maximizing the Evidence Lower Bound (ELBO) via Monotonic Alignment Search. In parallel, a flow-based stochastic duration predictor learns to model the distribution of phoneme durations from text encodings.

    During inference, the posterior encoder and ground-truth audio are omitted. The text encoder generates phoneme representations, the stochastic duration predictor samples phoneme durations to construct alignment AA, latent variables are sampled from the prior distribution and transformed through the inverse normalizing flow fθ−1f_\theta^{-1}, and the decoder synthesizes the output raw waveform directly in a parallel feed-forward pass.

  2. Knowl 2 — VITS Variational and Adversarial Training Objectives

    equation

    The overall optimization objective of VITS combines conditional VAE loss terms with GAN-based adversarial and feature-matching losses:

    Lvae=Lrecon+Lkl+Ldur+Ladv(G)+Lfm(G)L_{\text{vae}} = L_{\text{recon}} + L_{\text{kl}} + L_{\text{dur}} + L_{\text{adv}}(G) + L_{\text{fm}}(G)

    where each constituent loss is defined as follows:

    1. Reconstruction Loss (LreconL_{\text{recon}}): Defined as the L1L_1 distance between target 80-band mel-spectrograms xmelx_{\text{mel}} and predicted mel-spectrograms x^mel\hat{x}_{\text{mel}} computed from the generated audio y^=G(z)\hat{y} = G(z) via Short-Time Fourier Transform (STFT) and linear mel-filterbank projection:

    Lrecon=∥xmel−x^mel∥1L_{\text{recon}} = \|x_{\text{mel}} - \hat{x}_{\text{mel}}\|_1

    1. Latent KL Divergence (LklL_{\text{kl}}): The Kullback-Leibler divergence between the posterior distribution qϕ(z∣xlin)=N(z;μϕ(xlin),σϕ(xlin))q_\phi(z|x_{\text{lin}}) = \mathcal{N}(z; \mu_\phi(x_{\text{lin}}), \sigma_\phi(x_{\text{lin}})) conditioned on linear spectrogram xlinx_{\text{lin}} and the flow-augmented prior distribution pθ(z∣ctext,A)p_\theta(z|c_{\text{text}}, A):

    Lkl=log⁡qϕ(z∣xlin)−log⁡pθ(z∣ctext,A)L_{\text{kl}} = \log q_\phi(z|x_{\text{lin}}) - \log p_\theta(z|c_{\text{text}}, A)

    pθ(z∣ctext,A)=N(fθ(z);μθ(ctext,A),σθ(ctext,A))∣det⁡∂fθ(z)∂z∣p_\theta(z|c_{\text{text}}, A) = \mathcal{N}\left(f_\theta(z); \mu_\theta(c_{\text{text}}, A), \sigma_\theta(c_{\text{text}}, A)\right) \left| \det \frac{\partial f_\theta(z)}{\partial z} \right|

    1. Duration Predictor Loss (LdurL_{\text{dur}}): The negative variational lower bound of the phoneme duration log-likelihood under variational dequantization and data augmentation.

    2. Adversarial Loss (LadvL_{\text{adv}}): Least-squares GAN formulation where discriminator DD distinguishes ground truth audio yy from decoded audio G(z)G(z):

    Ladv(D)=E(y,z)[(D(y)−1)2+(D(G(z)))2]L_{\text{adv}}(D) = \mathbb{E}_{(y, z)} \left[ (D(y) - 1)^2 + (D(G(z)))^2 \right]

    Ladv(G)=Ez[(D(G(z))−1)2]L_{\text{adv}}(G) = \mathbb{E}_{z} \left[ (D(G(z)) - 1)^2 \right]

    1. Feature Matching Loss (LfmL_{\text{fm}}): Measured across intermediate discriminator layers to stabilize generator training:

    Lfm(G)=E(y,z)[∑l=1T1Nl∥Dl(y)−Dl(G(z))∥1]L_{\text{fm}}(G) = \mathbb{E}_{(y, z)} \left[ \sum_{l=1}^T \frac{1}{N_l} \|D^l(y) - D^l(G(z))\|_1 \right]

    where TT is the number of layers in discriminator DD, and DlD^l outputs the feature representation of the ll-th layer with NlN_l features.

  3. Knowl 3 — Monotonic Alignment Search for Conditional VAEs

    algorithm

    Monotonic Alignment Search (MAS) estimates the optimal hard alignment A∈{0,1}Ttext×TlatentA \in \{0, 1\}^{T_{\text{text}} \times T_{\text{latent}}} between phoneme representations and speech latent variables zz by finding an alignment that maximizes the Evidence Lower Bound (ELBO). This objective reduces to maximizing the log-likelihood of the latent representation under the flow-transformed conditional prior:

    arg⁡max⁡A^[log⁡pθ(xmel∣z)−log⁡qϕ(z∣xlin)pθ(z∣ctext,A^)]=arg⁡max⁡A^log⁡N(fθ(z);μθ(ctext,A^),σθ(ctext,A^))\arg\max_{\hat{A}} \left[ \log p_\theta(x_{\text{mel}}|z) - \log \frac{q_\phi(z|x_{\text{lin}})}{p_\theta(z|c_{\text{text}}, \hat{A})} \right] = \arg\max_{\hat{A}} \log \mathcal{N}\left(f_\theta(z); \mu_\theta(c_{\text{text}}, \hat{A}), \sigma_\theta(c_{\text{text}}, \hat{A})\right)

    The search is solved via dynamic programming over non-skipping, monotonic paths.

    Input: Prior distribution parameters μθ∈RTtext×dz\mu_\theta \in \mathbb{R}^{T_{\text{text}} \times d_z}, σθ∈RTtext×dz\sigma_\theta \in \mathbb{R}^{T_{\text{text}} \times d_z}, flow-transformed latent sequence fθ(z)∈RTlatent×dzf_\theta(z) \in \mathbb{R}^{T_{\text{latent}} \times d_z}
    Output: Optimal monotonic alignment matrix A∈{0,1}Ttext×TlatentA \in \{0, 1\}^{T_{\text{text}} \times T_{\text{latent}}}
    Compute log-likelihood matrix M∈RTtext×TlatentM \in \mathbb{R}^{T_{\text{text}} \times T_{\text{latent}}} where Mi,j=log⁡N(fθ(z)j;μθ,i,σθ,i)M_{i,j} = \log \mathcal{N}(f_\theta(z)_j; \mu_{\theta, i}, \sigma_{\theta, i})
    Initialize cache matrix Q∈RTtext×TlatentQ \in \mathbb{R}^{T_{\text{text}} \times T_{\text{latent}}} with −∞-\infty
    Initialize alignment matrix AA as zeros of shape Ttext×TlatentT_{\text{text}} \times T_{\text{latent}}
    for y=0y = 0 to Tlatent−1T_{\text{latent}} - 1 do
        for x=max⁡(0,Ttext+y−Tlatent)x = \max(0, T_{\text{text}} + y - T_{\text{latent}}) to min⁡(Ttext−1,y)\min(T_{\text{text}} - 1, y) do
            if y==0y == 0 then
                Q[x,0]=M[x,0]Q[x, 0] = M[x, 0]
            else
                if x==0x == 0 then
                    vprev=−∞v_{\text{prev}} = -\infty
                else
                    vprev=Q[x−1,y−1]v_{\text{prev}} = Q[x - 1, y - 1]
                end if
                vcur=Q[x,y−1]v_{\text{cur}} = Q[x, y - 1]
                Q[x,y]=M[x,y]+max⁡(vprev,vcur)Q[x, y] = M[x, y] + \max(v_{\text{prev}}, v_{\text{cur}})
            end if
        end for
    end for
    index=Ttext−1\text{index} = T_{\text{text}} - 1
    for y=Tlatent−1y = T_{\text{latent}} - 1 down to 00 do
        A[index,y]=1A[\text{index}, y] = 1
        if index≠0\text{index} \neq 0 and (index==y\text{index} == y or Q[index,y−1]<Q[index−1,y−1]Q[\text{index}, y - 1] < Q[\text{index} - 1, y - 1]) then
            index=index−1\text{index} = \text{index} - 1
        end if
    end for
    return AA
  4. Knowl 4 — Stochastic Duration Predictor via Variational Dequantization and Data Augmentation

    model/method

    To capture the one-to-many relationship of speech rhythms and model the correlated joint distribution of phoneme durations in parallel TTS, VITS incorporates a flow-based stochastic duration predictor.

    Direct maximum likelihood estimation on discrete scalar phoneme durations d∈Z>0Ttextd \in \mathbb{Z}_{>0}^{T_{\text{text}}} is hindered by non-differentiability and dimensional constraints on invertible flows. VITS overcomes this using variational dequantization and variational data augmentation:

    1. Variational Dequantization: A continuous random variable u∈[0,1)Ttextu \in [0, 1)^{T_{\text{text}}} is sampled such that d−u∈R>0Ttextd - u \in \mathbb{R}_{>0}^{T_{\text{text}}} forms a continuous positive sequence.
    2. Variational Data Augmentation: A channel-wise noise variable ν∈RTtext\nu \in \mathbb{R}^{T_{\text{text}}} with the same time resolution as dd is concatenated to yield a higher-dimensional latent representation [d−u,ν][d - u, \nu].

    The approximate posterior qϕ(u,ν∣d,ctext)q_\phi(u, \nu | d, c_{\text{text}}) generates uu and ν\nu, while the generative flow gθg_\theta maps [d−u,ν][d - u, \nu] to a standard normal distribution. The model is optimized by maximizing the variational lower bound of phoneme duration log-likelihood:

    log⁡pθ(d∣ctext)≥Eqϕ(u,ν∣d,ctext)[log⁡pθ(d−u,ν∣ctext)qϕ(u,ν∣d,ctext)]\log p_\theta(d | c_{\text{text}}) \ge \mathbb{E}_{q_\phi(u, \nu | d, c_{\text{text}})} \left[ \log \frac{p_\theta(d - u, \nu | c_{\text{text}})}{q_\phi(u, \nu | d, c_{\text{text}})} \right]

    The duration loss LdurL_{\text{dur}} is the negative ELBO. Stop-gradient operators are applied to text condition inputs htexth_{\text{text}} to isolate duration predictor training from the rest of the network.

    Architecturally, the duration predictor employs 4 coupling layers parameterized by Neural Spline Flows using monotonic rational-quadratic splines (10 bins, 29 parameter channels) and dilated depth-wise separable convolutional (DDSConv) residual blocks with Layer Normalization and GELU activations, operating at a hidden dimension of 192.

  5. Knowl 5 — Neural Network Architecture Specifications for VITS

    model/method

    The VITS architecture consists of five core components:

    1. Posterior Encoder: Composed of 16 non-causal WaveNet residual blocks with dilated convolutions, gated activation units, and skip connections. It processes linear spectrograms (STFT with FFT size 1024, window size 1024, hop size 256) and projects to 192-channel latent Gaussian parameters (μϕ,σϕ)(\mu_\phi, \sigma_\phi).

    2. Prior Encoder: Composed of a Transformer text encoder using relative positional representations that maps IPA phoneme sequences to hidden representation htexth_{\text{text}}, followed by a linear projection layer generating (μθ,σθ)(\mu_\theta, \sigma_\theta) and a normalizing flow fθf_\theta. The flow consists of 4 affine coupling layers, each containing 4 WaveNet residual blocks configured as volume-preserving transformations (Jacobian determinant of 1).

    3. Decoder: Based on the HiFi-GAN V1 generator architecture. It comprises stacked transposed convolutions for upsampling 192-channel latent representations zz to audio waveforms, each followed by Multi-Receptive Field Fusion (MRF) modules that sum outputs across different kernel sizes and receptive fields. The bias parameter of the final convolution is removed to prevent gradient instabilities during mixed precision training.

    4. Discriminator: A Multi-Period Discriminator (MPD) operating on raw waveforms across periods [1,2,3,5,7,11][1, 2, 3, 5, 7, 11]. Period 1 corresponds to the full-resolution sub-discriminator of the Multi-Scale Discriminator (MSD), while average-pooled sub-discriminators are omitted to improve training efficiency.

    5. Windowed Generator Training: During training, the decoder only upsamples randomly extracted latent segments of window size 32 (and matching ground truth raw waveform segments) rather than full utterances, minimizing memory footprint and compute time.

  6. Knowl 6 — Multi-Speaker Voice Conversion via Invertible Latent Transformations

    model/method

    In the multi-speaker configuration of VITS, speaker identity embeddings are provided via global conditioning to the posterior encoder WaveNet blocks, the prior normalizing flow residual blocks, the duration predictor condition encoder, and linearly projected into the decoder latent space zz. However, speaker embeddings are deliberately excluded from the text encoder, forcing the text representation htexth_{\text{text}} and the base prior distribution to learn speaker-independent representations.

    This disentanglement enables waveform-domain voice conversion between arbitrary source speaker ss and target speaker s^\hat{s} without requiring parallel training corpora or text transcripts:

    1. Given an audio recording of source speaker ss, compute linear-scale spectrogram xlinx_{\text{lin}} and sample latent variable zz from the speaker-conditioned posterior:

    z∼qϕ(z∣xlin,s)z \sim q_\phi(z | x_{\text{lin}}, s)

    1. Transform zz into a speaker-independent canonical representation ee using the forward pass of the normalizing flow conditioned on ss:

    e=fθ(z∣s)e = f_\theta(z | s)

    1. Synthesize the converted raw waveform y^\hat{y} in the voice of target speaker s^\hat{s} by applying the inverse flow conditioned on s^\hat{s} followed by decoder GG:

    y^=G(fθ−1(e∣s^) | s^)\hat{y} = G\left(f_\theta^{-1}(e | \hat{s}) \,\middle|\, \hat{s}\right)

  7. Knowl 7 — Subjective Speech Quality on Single-Speaker and Multi-Speaker Benchmarks

    data/table

    Subjective Mean Opinion Score (MOS) evaluations with 95% confidence intervals were conducted on the single-speaker LJ Speech dataset (22 kHz) and the multi-speaker VCTK dataset (109 speakers downsampled to 22 kHz). VITS was compared against ground truth audio, two-stage autoregressive models (Tacotron 2 + HiFi-GAN), two-stage flow-based non-autoregressive models (Glow-TTS + HiFi-GAN), and an ablation variant employing a deterministic duration predictor (VITS DDP).

    Model MOS (CI)
    LJ Speech (Single Speaker)
    Ground Truth 4.46 (±\pm0.06)
    Tacotron 2 + HiFi-GAN 3.77 (±\pm0.08)
    Tacotron 2 + HiFi-GAN (Fine-tuned) 4.25 (±\pm0.07)
    Glow-TTS + HiFi-GAN 4.14 (±\pm0.07)
    Glow-TTS + HiFi-GAN (Fine-tuned) 4.32 (±\pm0.07)
    VITS (DDP) 4.39 (±\pm0.06)
    VITS 4.43 (±\pm0.06)
    VCTK (Multi-Speaker)
    Ground Truth 4.38 (±\pm0.07)
    Tacotron 2 + HiFi-GAN 3.14 (±\pm0.09)
    Tacotron 2 + HiFi-GAN (Fine-tuned) 3.19 (±\pm0.09)
    Glow-TTS + HiFi-GAN 3.76 (±\pm0.07)
    Glow-TTS + HiFi-GAN (Fine-tuned) 3.82 (±\pm0.07)
    VITS 4.38 (±\pm0.06)

    VITS achieves a MOS of 4.43 on LJ Speech and 4.38 on VCTK, matching ground truth natural speech within statistical confidence and outperforming both baseline and fine-tuned two-stage systems. VITS with the stochastic duration predictor outperforms VITS (DDP) (4.39 MOS), confirming that stochastic duration modeling generates more natural phoneme rhythms.

  8. Knowl 8 — Ablation Analysis of Flow-Based Prior and Spectrogram Resolution

    data/table

    An ablation study evaluated the impact of two architectural choices in VITS: the conditional prior normalizing flow and the spectrogram representation fed to the posterior encoder. All variants were trained for 300k steps on LJ Speech and evaluated via crowd-sourced MOS tests.

    Model MOS (CI)
    Ground Truth 4.50 (±\pm0.06)
    Baseline (VITS at 300k steps) 4.50 (±\pm0.06)
    without Normalizing Flow 2.98 (±\pm0.08)
    with Mel-spectrogram 4.31 (±\pm0.08)

    Removing the normalizing flow in the prior encoder caused the largest quality degradation, reducing MOS by 1.52 points (from 4.50 to 2.98), demonstrating that expressive prior modeling beyond a standard factorized normal distribution is critical for high-fidelity waveform generation. Replacing the linear-scale spectrogram input to the posterior encoder with an 80-band mel-spectrogram degraded MOS by 0.19 points (from 4.50 to 4.31), indicating that providing high-resolution spectral details to the posterior encoder improves synthesis fidelity.

  9. Knowl 9 — Synthesis Speed and Real-Time Factor of Parallel Waveform Generation

    data/table

    Synthesis speed was benchmarked on a single NVIDIA V100 GPU with a batch size of 1 over 100 sentences randomly selected from the LJ Speech test set. Speed measures generated audio sampling rate (in kHz), and the real-time factor indicates generation speed multiplier relative to real-time playback.

    Model Speed (kHz) Real-Time Factor
    Glow-TTS + HiFi-GAN 606.05 ×27.48\times 27.48
    VITS 1480.15 ×67.12\times 67.12
    VITS (DDP) 2005.03 ×90.93\times 90.93

    By unifying text alignment, acoustic modeling, and vocoding into a single-stage model without creating intermediate mel-spectrogram files or running separate stage decoders, VITS achieves an end-to-end synthesis throughput of 1480.15 kHz (67.12×\times faster than real time), more than doubling the inference speed of the two-stage parallel Glow-TTS + HiFi-GAN pipeline (27.48×\times real time).

  10. Knowl 10 — Comparative Evaluation Gap Relative to Natural Ground Truth Speech

    limitation

    Although VITS achieves absolute Mean Opinion Scores (MOS) statistically comparable to natural recordings in standalone ratings, fine-grained 7-point Comparative Mean Opinion Score (CMOS) side-by-side evaluations reveal that listeners still exhibit a subtle preference for ground truth recordings over VITS.

    In side-by-side evaluations across 50 items (500 total ratings):

    • On the LJ Speech dataset, VITS achieved a CMOS of -0.106 relative to ground truth.
    • On the VCTK multi-speaker dataset, VITS achieved a CMOS of -0.270 relative to ground truth.

    Additionally, while VITS integrates text-to-waveform modeling end-to-end, it still relies on a separate grapheme-to-phoneme text preprocessing front-end (converting raw text to IPA phoneme sequences with blank tokens) rather than learning representations directly from raw character strings.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.Bernard, M. Phonemizer. https://github.com/bootphon/phonemizer, 2021.
  2. 2.Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L. C., and Simonyan, K. High Fidelity Speech Synthesis with Adversarial Networks. In International Conference on Learning Representations, 2019.
  3. 3.Chen, J., Lu, C., Chenli, B., Zhu, J., and Tian, T. Vflow: More expressive generative flows with variational data augmentation. In International Conference on Machine Learning, pp. 1660–1669. PMLR, 2020.
  4. 4.Chen, X., Kingma, D. P., Salimans, T., Duan, Y., Dhariwal, P., Schulman, J., Sutskever, I., and Abbeel, P. Variational lossy autoencoder. 2017. URL https://openreview.net/forum?id=BysvGP5ee.
  5. 5.De Cheveigné, A. and Kawahara, H. Yin, a fundamental frequency estimator for speech and music. The Journal of the Acoustical Society of America, 111(4):1917–1930, 2002.
  6. 6.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using Real NVP. In International Conference on Learning Representations, 2017.
  7. 7.Donahue, J., Dieleman, S., Binkowski, M., Elsen, E., and Simonyan, K. End-to-end Adversarial Text-to-Speech. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=rsf1z-JSj87.
  8. 8.Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G. Neural Spline Flows. In Advances in Neural Information Processing Systems, pp. 7509–7520, 2019.
  9. 9.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative Adversarial Nets. Advances in Neural Information Processing Systems, 27:2672–2680, 2014.
  10. 10.Graves, A. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.
  11. 11.Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In International Conference on Machine Learning, pp. 2722–2730. PMLR, 2019.
  12. 12.Hsu, W.-N., Zhang, Y., Weiss, R., Zen, H., Wu, Y., Cao, Y., and Wang, Y. Hierarchical Generative Modeling for Controllable Speech Synthesis. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rygkk305YQ.
  13. 13.Ito, K. The LJ Speech Dataset. https://keithito.com/LJ-Speech-Dataset/, 2017.
  14. 14.Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Lopez-Moreno, I., et al. Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis. In Advances in Neural Information Processing Systems, 2018.
  15. 15.Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A., Dieleman, S., and Kavukcuoglu, K. Efficient neural audio synthesis. In International Conference on Machine Learning, pp. 2410–2419. PMLR, 2018.
  16. 16.Kim, J., Kim, S., Kong, J., and Yoon, S. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search. Advances in Neural Information Processing Systems, 33, 2020.
  17. 17.Kingma, D. P. and Welling, M. Auto-Encoding Variational Bayes. In International Conference on Learning Representations, 2014.
  18. 18.Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. Improved variational inference with inverse autoregressive flow. Advances in Neural Information Processing Systems, 29:4743–4751, 2016.
  19. 19.Kong, J., Kim, J., and Bae, J. HiFi-GAN: Generative Adversarial networks for Efficient and High Fidelity Speech Synthesis. Advances in Neural Information Processing Systems, 33, 2020.
  20. 20.Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C. MelGAN: Generative Adversarial Networks for Conditional waveform synthesis. volume 32, pp. 14910–14921, 2019.
  21. 21.Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O. Autoencoding beyond pixels using a learned similarity metric. In International Conference on Machine Learning, pp. 1558–1566. PMLR, 2016.
  22. 22.Lee, Y., Shin, J., and Jung, K. Bidirectional Variational Inference for Non-Autoregressive Text-to-speech. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=o3iritJHLfO.
  23. 23.Li, N., Liu, S., Liu, Y., Zhao, S., and Liu, M. Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 6706–6713, 2019.
  24. 24.Loshchilov, I. and Hutter, F. Decoupled Weight Decay Regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  25. 25.Ma, X., Zhou, C., Li, X., Neubig, G., and Hovy, E. Flowseq: Non-autoregressive conditional sequence generation with generative flow. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4273–4283, 2019.
  26. 26.Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Paul Smolley, S. Least squares generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, pp. 2794–2802, 2017.
  27. 27.Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J. Flow-TTS: A non-autoregressive network for text to speech based on flow. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7209–7213. IEEE, 2020.
  28. 28.Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
  29. 29.Peng, K., Ping, W., Song, Z., and Zhao, K. Non-autoregressive neural text-to-speech. In International Conference on Machine Learning, pp. 7586–7598. PMLR, 2020.
  30. 30.Ping, W., Peng, K., Gibiansky, A., Arik, S. O., Kannan, A., Narang, S., Raiman, J., and Miller, J. Deep Voice 3: 2000-Speaker Neural Text-to-Speech. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=HJtEm4p6Z.
  31. 31.Prenger, R., Valle, R., and Catanzaro, B. Waveglow: A flow-based generative network for speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3617–3621. IEEE, 2019.
  32. 32.Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y. FastSpeech: Fast, Robust and Controllable Text to Speech. volume 32, pp. 3171–3180, 2019.
  33. 33.Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=piLPYqxtWuA.
  34. 34.Rezende, D. and Mohamed, S. Variational inference with normalizing flows. In International Conference on Machine Learning, pp. 1530–1538. PMLR, 2015.
  35. 35.Shaw, P., Uszkoreit, J., and Vaswani, A. Self-Attention with Relative Position Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pp. 464–468, 2018.
  36. 36.Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4779–4783. IEEE, 2018.
  37. 37.Taigman, Y., Wolf, L., Polyak, A., and Nachmani, E. Voiceloop: Voice Fitting and Synthesis via a Phonological Loop. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=SkFAWax0-.
  38. 38.Valle, R., Shih, K. J., Prenger, R., and Catanzaro, B. Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Ig53hpHxS4.
  39. 39.van den Oord, A., Vinyals, O., and Kavukcuoglu, K. Neural discrete representation learning. In Advances in Neural Information Processing Systems, pp. 6309–6318, 2017.
  40. 40.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is All you Need. Advances in Neural Information Processing Systems, 30:5998–6008, 2017.
  41. 41.Veaux, C., Yamagishi, J., MacDonald, K., et al. CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit. University of Edinburgh. The Centre for Speech Technology Research (CSTR), 2017.
  42. 42.Weiss, R. J., Skerry-Ryan, R., Battenberg, E., Mariooryad, S., and Kingma, D. P. Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis. arXiv preprint arXiv:2011.03568, 2020.
  43. 43.Zeng, Z., Wang, J., Cheng, N., Xia, T., and Xiao, J. Aligntts: Efficient feed-forward text-to-speech system without explicit alignment. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6714–6718. IEEE, 2020.
  44. 44.Zhang, Y.-J., Pan, S., He, L., and Ling, Z.-H. Learning latent representations for style control and transfer in end-to-end speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6945–6949. IEEE, 2019.
  45. 45.Ziegler, Z. and Rush, A. Latent normalizing flows for discrete sequences. In International Conference on Machine Learning, pp. 7673–7682. PMLR, 2019.

Citation

MLA
Kim, J., et al. “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”. arXiv, 2021, http://arxiv.org/abs/2106.06103v1.
APA
Kim, J., Kong, J., & Son, J. (2021). Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. arXiv. http://arxiv.org/abs/2106.06103v1
Chicago
Kim, J., J. Kong, and J. Son. 2021. “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”. arXiv. http://arxiv.org/abs/2106.06103v1.
Harvard
Kim, J., Kong, J. and Son, J. (2021) “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.06103v1.
Vancouver
1. Kim J, Kong J, Son J (2021) Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. arXiv

BibTeX

@article{kim2021conditional,
  title = {Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech},
  author = {Kim, Jaehyeon and Kong, Jungil and Son, Juhee},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.06103v1},
  eprint = {2106.06103}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/