FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Yi RenChenxu HuXu TanTao QinSheng ZhaoZhou ZhaoTie-Yan Liu

article2020ICLR1,850 citations

Introduces a non-autoregressive text-to-speech framework that trains directly on ground-truth acoustic targets conditioned on pitch, energy, and duration, eliminating complex distillation pipelines to achieve three times faster training and voice quality surpassing autoregressive models.

Listen

Text-to-speech technology is essential for applications ranging from voice assistants to digital media production. However, synthesizing speech from text poses a fundamental one-to-many mapping challenge because a single text sequence can be spoken with diverse variations in duration, pitch, and energy. While earlier non-autoregressive models like FastSpeech improved synthesis speed over traditional autoregressive systems, they relied on a complex two-stage training process involving knowledge distillation from a teacher model. This pipeline resulted in long training times, inaccurate duration extraction, and information loss that degraded voice quality.

The article aims to resolve these limitations by developing FastSpeech 2 and FastSpeech 2s. The objective was to demonstrate that directly training on ground-truth targets while explicitly conditioning on speech variations—specifically pitch, energy, and accurate duration—simplifies training, accelerates synthesis, and delivers superior audio quality.

To evaluate this approach, the researchers conducted experiments using the 24-hour single-speaker LJSpeech dataset containing 13,100 English audio clips. The architecture replaces teacher-student distillation by training directly on actual speech targets. It incorporates a variance adaptor that extracts ground-truth phoneme durations via forced alignment and analyzes pitch contours using continuous wavelet transforms to capture frequency variations. In addition, the authors developed FastSpeech 2s, which extends the model into a fully non-autoregressive system capable of generating raw speech waveforms directly from text using adversarial training, bypassing intermediate acoustic spectrograms.

The findings show substantial improvements across training efficiency, inference speed, and voice quality. FastSpeech 2 achieved a roughly 3.12-fold speed-up in training time compared to FastSpeech, reducing acoustic model training from 53.12 hours to 17.02 hours. In terms of generation speed, FastSpeech 2 and FastSpeech 2s synthesized audio approximately 47.8 and 51.8 times faster than standard autoregressive Transformer models. Furthermore, subjective listening evaluations demonstrated that FastSpeech 2 surpassed both FastSpeech and traditional autoregressive baselines in perceptual voice quality, while FastSpeech 2s matched autoregressive quality. Ablation analyses confirmed that forced-alignment duration, continuous wavelet pitch modeling, and energy conditioning each significantly contributed to prosody and naturalness.

These results indicate that explicit variance conditioning effectively solves the one-to-many mapping problem without requiring complex distillation pipelines. For practitioners, this reduces compute costs, shortens development lifecycles, and yields voice synthesis that is highly controllable in pitch and volume while maintaining ultra-low latency for real-time applications.

Decision-makers and engineering teams should consider adopting the FastSpeech 2 architecture for voice generation pipelines that require high throughput and robust quality. Teams with strict latency constraints or resource limits can deploy FastSpeech 2s to eliminate separate vocoder dependencies. However, future development should aim to build fully end-to-end alignment tools directly into the architecture, as the current framework still relies on external tools for forced alignment and pitch extraction.

  • Paper: Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions, Jonathan Shen et al. (2017). Tacotron 2 established the dominant mel-spectrogram prediction and neural vocoding paradigm that FastSpeech 2 directly compares against, builds on, and replaces with non-autoregressive generation.
  • Paper: Tacotron: Towards End-to-End Speech Synthesis, Yuxuan Wang et al. (2017). Tacotron introduced the foundational end-to-end sequence-to-sequence neural text-to-speech architecture upon which modern acoustic models are constructed.
  • Paper: WaveNet: A Generative Model for Raw Audio, Aäron van den Oord et al. (2016). WaveNet introduced neural autoregressive waveform synthesis, establishing the baseline audio fidelity and vocoding standards that FastSpeech 2 and FastSpeech 2s seek to match with faster parallel generation.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). The Transformer provides the foundational feed-forward self-attention mechanisms that form the core structural blocks of the FastSpeech architecture.
  • Paper: Sequence-Level Knowledge Distillation, Yoon Kim et al. (2016). Sequence-level knowledge distillation defines the teacher-student training paradigm that the original FastSpeech relied upon and that FastSpeech 2 explicitly redesigns to avoid information loss.
  • Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN provides the high-fidelity, non-autoregressive neural vocoder commonly paired with acoustic models like FastSpeech 2 to achieve fully parallelized, high-speed waveform generation from predicted mel-spectrograms.
  • Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). DiffWave extends the non-autoregressive audio synthesis domain to diffusion probabilistic models, providing an alternative parallel vocoder for generating waveforms from mel-spectrogram representations.
  • Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS builds upon modern neural speech synthesis paradigms, extending fast end-to-end generation principles to large-scale multilingual modeling and zero-shot voice cloning.
Cover for FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Abstract

Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher model for duration prediction (to provide more information as input) and knowledge distillation (to simplify the data distribution in output), which can ease the one-to-many mapping problem (i.e., multiple speech variations correspond to the same text) in TTS. However, FastSpeech has several disadvantages: 1) the teacher-student distillation pipeline is complicated and time-consuming, 2) the duration extracted from the teacher model is not accurate enough, and the target mel-spectrograms distilled from teacher model suffer from information loss due to data simplification, both of which limit the voice quality. In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the simplified output from teacher, and 2) introducing more variation information of speech (e.g., pitch, energy and more accurate duration) as conditional inputs. Specifically, we extract duration, pitch and energy from speech waveform and directly take them as conditional inputs in training and use predicted values in inference. We further design FastSpeech 2s, which is the first attempt to directly generate speech waveform from text in parallel, enjoying the benefit of fully end-to-end inference. Experimental results show that 1) FastSpeech 2 achieves a 3x training speed-up over FastSpeech, and FastSpeech 2s enjoys even faster inference speed; 2) FastSpeech 2 and 2s outperform FastSpeech in voice quality, and FastSpeech 2 can even surpass autoregressive models. Audio samples are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 FastSpeech 2 and 2s
  • 2.1 Motivation
  • 2.2 Model Overview
  • 2.3 Variance Adaptor
  • 2.4 FastSpeech 2s
  • 2.5 Discussions
  • 3 Experiments and Results
  • 3.1 Experimental Setup
  • 3.2 Results
  • 3.2.1 Model Performance
  • 3.2.2 Analyses on Variance Information
  • 3.2.3 Ablation Study
  • 4 Conclusion
  • References
  • A Model Configuration
  • B Training and Inference
  • C Modeling Pitch with Continuous Wavelet Transform
  • C.1 Continuous Wavelet Transform
  • C.2 Implementation Details
  • D Case Study on Pitch Contour
  • E Variance Control

Knowls

  1. Knowl 1 — FastSpeech 2 Architecture and Direct Mel-Spectrogram Training

    model/method

    FastSpeech 2 is a non-autoregressive neural text-to-speech architecture that synthesizes mel-spectrograms in parallel from phoneme sequences, eliminating the two-stage teacher-student distillation pipeline used in FastSpeech.

    The model consists of three core components:

    1. Phoneme Encoder: Converts input phoneme embeddings into hidden representations using 4 Feed-Forward Transformer (FFT) blocks. Each block contains a 2-head self-attention layer (hidden size 256) followed by a 2-layer 1D convolutional network with kernel sizes 9 and 1, filter sizes 1024 and 256, and dropout 0.1.
    2. Variance Adaptor: Enhances phoneme hidden representations by adding explicit variance information—phoneme duration, fundamental frequency (F0F_0), and frame-level energy—to alleviate the one-to-many mapping problem in text-to-speech.
    3. Mel-Spectrogram Decoder: Transforms the duration-expanded and variance-augmented hidden representations into 80-channel mel-spectrograms in parallel using 4 FFT blocks of the same architecture as the encoder, followed by a linear projection layer.

    FastSpeech 2 discards teacher-generated mel-spectrogram targets and instead trains directly on ground-truth mel-spectrograms Y∈RTmel×80Y \in \mathbb{R}^{T_{\text{mel}} \times 80} using Mean Absolute Error (MAE):

    Lmel=1Tmel×80∑t=1Tmel∑d=180∣Yt,d−Y^t,d∣\mathcal{L}_{\text{mel}} = \frac{1}{T_{\text{mel}} \times 80} \sum_{t=1}^{T_{\text{mel}}} \sum_{d=1}^{80} |Y_{t,d} - \hat{Y}_{t,d}|

    where Y^\hat{Y} denotes the predicted mel-spectrogram and TmelT_{\text{mel}} is the number of mel frames. Direct ground-truth training avoids the information loss in teacher-distilled targets and significantly simplifies training.

  2. Knowl 2 — Variance Adaptor: Duration, Pitch, and Energy Modeling

    model/method

    The Variance Adaptor adds phoneme duration, pitch, and energy information to the phoneme hidden sequence to bridge the information gap between the text input and speech output.

    All three predictors share a common structure: a 2-layer 1D convolutional network with kernel size 3, filter size 256, ReLU activation, layer normalization, dropout of 0.5, and a final linear layer.

    1. Duration Predictor and Length Regulator: Takes the phoneme hidden sequence Hphoneme∈RN×256H_{\text{phoneme}} \in \mathbb{R}^{N \times 256} and predicts the duration (number of mel frames) of each phoneme in the logarithmic domain, optimized with Mean Square Error (MSE) loss against durations extracted by Montreal Forced Alignment (MFA). The Length Regulator expands HphonemeH_{\text{phoneme}} by replicating each phoneme representation according to its duration, yielding a frame-level sequence Hframe∈RTmel×256H_{\text{frame}} \in \mathbb{R}^{T_{\text{mel}} \times 256}.
    2. Pitch Predictor and Pitch Embedding: Predicts Continuous Wavelet Transform (CWT) pitch spectrograms and utterance-level pitch mean and variance from HframeH_{\text{frame}} using MSE loss. For conditioning, the frame-level fundamental frequency F0F_0 (ground truth during training, predicted during inference) is quantized into 256 bins on a logarithmic scale, mapped to a 256-dimensional embedding vector ptp_t, and added to Hframe,tH_{\text{frame}, t}.
    3. Energy Predictor and Energy Embedding: Frame energy is calculated as the L2L_2-norm of the Short-Time Fourier Transform (STFT) magnitude: et=∑f∣Xt,f∣2e_t = \sqrt{\sum_f |X_{t,f}|^2}. The predictor is trained with MSE loss on raw energy values. For conditioning, ete_t is uniformly quantized into 256 bins, mapped to a 256-dimensional embedding vector ete_t, and added to Hframe,tH_{\text{frame}, t}.

    The adapted hidden state at frame tt is formed by summation:

    Ht′=Hframe,t+pt+etH'_t = H_{\text{frame}, t} + p_t + e_t

  3. Knowl 3 — Continuous Wavelet Transform for Pitch Modeling

    model/method

    To handle rapid fluctuations and unvoiced discontinuities in fundamental frequency (F0F_0) contours, FastSpeech 2 decomposes pitch contours into the time-frequency domain using Continuous Wavelet Transform (CWT) and reconstructs them via Inverse Continuous Wavelet Transform (iCWT).

    CWT Decomposition: Given a continuous pitch contour F0(x)F_0(x), its wavelet representation is:

    W(τ,t)=τ−1/2∫−∞+∞F0(x)ψ(x−tτ)dxW(\tau, t) = \tau^{-1/2} \int_{-\infty}^{+\infty} F_0(x) \psi\left(\frac{x - t}{\tau}\right) dx

    where ψ(u)=23π1/4(1−u2)e−u2/2\psi(u) = \frac{2}{\sqrt{3}\pi^{1/4}} (1 - u^2) e^{-u^2/2} is the Mexican hat mother wavelet, τ\tau is the wavelet scale, and tt is the time position. The pitch contour is decomposed across 10 discrete scales:

    Wi(t)=W(2i+1τ0,t)(i+2.5)−5/2,i∈{1,2,…,10}W_i(t) = W(2^{i+1} \tau_0, t)(i + 2.5)^{-5/2}, \quad i \in \{1, 2, \dots, 10\}

    where τ0=5 ms\tau_0 = 5\text{ ms}.

    iCWT Reconstruction: Given the 10 wavelet components W^i(t)\hat{W}_i(t), the pitch contour F^0(t)\hat{F}_0(t) is reconstructed by:

    F^0(t)=∑i=110W^i(t)(i+2.5)−5/2\hat{F}_0(t) = \sum_{i=1}^{10} \hat{W}_i(t)(i + 2.5)^{-5/2}

    Predictor Workflow:

    1. Unvoiced frames in F0F_0 are linearly interpolated.
    2. The contour is converted to logarithmic scale and normalized to zero mean and unit variance per utterance.
    3. The pitch predictor predicts the 10-scale pitch spectrogram and the utterance-level mean and variance (obtained by global average pooling over time of the convolutional features followed by a linear projection).
    4. The pitch predictor is trained with MSE loss on both the pitch spectrogram and utterance statistics.
    5. During inference, the predicted spectrogram is reconstructed into F^0\hat{F}_0 via iCWT and denormalized.
  4. Knowl 4 — FastSpeech 2s Fully Non-Autoregressive Text-to-Waveform Architecture

    model/method

    FastSpeech 2s is a fully non-autoregressive text-to-waveform model that directly synthesizes speech waveforms from phoneme sequences in parallel, bypassing the intermediate mel-spectrogram stage during inference.

    Model Architecture:

    • Phoneme Encoder & Variance Adaptor: Shared with FastSpeech 2, processing phonemes and injecting duration, pitch, and energy representations.
    • Waveform Decoder: Based on WaveNet with non-causal 1D convolutions and gated activation units. It upsamples frame-level hidden states to the audio sample rate via a 1-layer transposed 1D convolution (filter size 64) followed by 30 dilated residual convolution blocks with skip channel size 64 and kernel size 3.
    • Auxiliary Mel-Spectrogram Decoder: An FFT-based mel-spectrogram decoder trained concurrently on full sentences with MAE loss to regularize and improve text feature extraction, which is discarded during inference.

    Training Scheme:

    • Due to memory limitations with high-resolution audio, the waveform decoder is trained on short sliced audio clips of 20,480 samples (~0.93 seconds at 22,050 Hz).
    • Adversarial Training: A discriminator consisting of 10 layers of non-causal dilated 1D convolutions with LeakyReLU activations (from Parallel WaveGAN) is trained alongside the waveform decoder.
    • Loss Function: The waveform decoder is optimized using multi-resolution STFT loss combined with Least Squares GAN (LSGAN) discriminator loss to implicitly recover phase information.
  5. Knowl 5 — Audio Quality Evaluation on LJSpeech

    empirical result

    The audio quality of FastSpeech 2 and FastSpeech 2s was evaluated on the LJSpeech dataset (single female speaker, 22,050 Hz sample rate) using Mean Opinion Score (MOS) with 95% confidence intervals and Comparative Mean Opinion Score (CMOS). Mel-spectrogram models were synthesized into waveforms using Parallel WaveGAN (PWG).

    Method MOS
    Ground Truth (GT) 4.30±0.074.30 \pm 0.07
    GT (Mel + PWG) 3.92±0.083.92 \pm 0.08
    Tacotron 2 (Mel + PWG) 3.70±0.083.70 \pm 0.08
    Transformer TTS (Mel + PWG) 3.72±0.073.72 \pm 0.07
    FastSpeech (Mel + PWG) 3.68±0.093.68 \pm 0.09
    FastSpeech 2 (Mel + PWG) 3.83±0.083.83 \pm 0.08
    FastSpeech 2s 3.71±0.093.71 \pm 0.09

    In CMOS comparisons relative to FastSpeech 2 (score 0.000):

    • FastSpeech achieved −0.885-0.885 CMOS.
    • Transformer TTS achieved −0.235-0.235 CMOS.

    FastSpeech 2 outperforms the original FastSpeech and surpasses the autoregressive Tacotron 2 and Transformer TTS models. FastSpeech 2s matches the voice quality of autoregressive baselines while operating end-to-end without an intermediate vocoder.

  6. Knowl 6 — Training and Inference Efficiency Comparison

    empirical result

    FastSpeech 2 substantially accelerates model training by removing the autoregressive teacher training and distillation pipeline, while FastSpeech 2s achieves the fastest inference latency.

    Training time and inference latency on the LJSpeech dataset were evaluated on a server equipped with 36 Intel Xeon CPUs, 256 GB RAM, and 1 NVIDIA V100 GPU (batch size 48 for training, 1 for inference):

    Method Training Time (h) Inference Speed (RTF) Inference Speedup
    Transformer TTS 38.64 9.32×10−19.32 \times 10^{-1} 1.0×1.0\times
    FastSpeech 53.12 1.92×10−21.92 \times 10^{-2} 48.5×48.5\times
    FastSpeech 2 17.02 1.95×10−21.95 \times 10^{-2} 47.8×47.8\times
    FastSpeech 2s 92.18 1.80×10−21.80 \times 10^{-2} 51.8×51.8\times

    Key findings:

    • FastSpeech training time (53.12 h) includes teacher training (38.64 h) plus student training (14.48 h). FastSpeech 2 requires only 17.02 h, achieving a 3.12×3.12\times training speedup.
    • Real-Time Factor (RTF) measures the time (in seconds) required to synthesize 1 second of waveform. Parallel WaveGAN served as the vocoder for Transformer TTS, FastSpeech, and FastSpeech 2.
    • FastSpeech 2s attains an RTF of 1.80×10−21.80 \times 10^{-2} (51.8×51.8\times faster than Transformer TTS), outperforming FastSpeech 2 due to fully non-autoregressive, single-stage waveform generation.
  7. Knowl 7 — Phoneme Duration Alignment Accuracy: MFA vs Teacher Model

    empirical result

    FastSpeech extracted duration labels from the cross-attention matrix of an autoregressive teacher model. FastSpeech 2 uses Montreal Forced Alignment (MFA) to obtain duration ground truth.

    Alignment accuracy was evaluated by computing the average absolute phoneme boundary difference Δ\Delta (in milliseconds) against 50 manually aligned audio samples. The perceptual impact of alignment accuracy was evaluated with CMOS by training FastSpeech with both alignment sources:

    Duration Extraction Method Δ\mathbf{\Delta} (ms)
    Duration from teacher model 19.68
    Duration from MFA 12.47
    Model Setting CMOS
    FastSpeech + Duration from teacher 0.000
    FastSpeech + Duration from MFA +0.195

    MFA reduces phoneme boundary error from 19.68 ms19.68\text{ ms} to 12.47 ms12.47\text{ ms} (36.6%36.6\% reduction). Replacing teacher durations with MFA durations in FastSpeech yields a +0.195+0.195 CMOS improvement in synthesized voice quality.

  8. Knowl 8 — Pitch and Energy Distribution Fidelity

    empirical result

    The fidelity of pitch (F0F_0) and frame energy in synthesized speech was evaluated against ground-truth LJSpeech recordings.

    Pitch contour fidelity was measured using the statistical moments of the F0F_0 distribution—standard deviation (σ\sigma), skewness (γ\gamma), and kurtosis (KK)—and the average Dynamic Time Warping (DTW) distance to ground-truth pitch. Energy fidelity was evaluated via Mean Absolute Error (MAE) between frame-level energy (L2L_2-norm of STFT amplitude) of synthesized and ground-truth speech:

    Method σ\mathbf{\sigma} γ\mathbf{\gamma} K\mathbf{K} DTW Energy MAE
    Ground Truth (GT) 54.4 0.836 0.977 / /
    Tacotron 2 44.1 1.280 1.311 26.32 /
    Transformer TTS 40.8 0.703 1.419 24.40 /
    FastSpeech 50.8 0.724 -0.041 24.89 0.142
    FastSpeech 2 54.1 0.881 0.996 24.39 0.131
    FastSpeech 2 (without CWT) 42.3 0.771 1.115 25.13 /
    FastSpeech 2s 53.9 0.872 0.998 24.37 0.133

    FastSpeech 2 matches the ground-truth pitch moments (σ=54.1\sigma=54.1, γ=0.881\gamma=0.881, K=0.996K=0.996 vs GT σ=54.4\sigma=54.4, γ=0.836\gamma=0.836, K=0.977K=0.977) and achieves a lower pitch DTW distance (24.3924.39) than FastSpeech (24.8924.89) and Tacotron 2 (26.3226.32). Energy MAE is reduced from 0.1420.142 in FastSpeech to 0.1310.131 in FastSpeech 2 and 0.1330.133 in FastSpeech 2s.

  9. Knowl 9 — Ablation Study on Variance Predictors and Architectural Choices

    empirical result

    Ablation experiments using CMOS on LJSpeech measured the contribution of individual variance components and architectural designs in FastSpeech 2 and FastSpeech 2s:

    FastSpeech 2 Setting CMOS FastSpeech 2s Setting CMOS
    FastSpeech 2 (Default) 0.000 FastSpeech 2s (Default) 0.000
    FastSpeech 2 −- energy -0.040 FastSpeech 2s −- energy -0.160
    FastSpeech 2 −- pitch -0.245 FastSpeech 2s −- pitch -1.130
    FastSpeech 2 −- pitch −- energy -0.370 FastSpeech 2s −- pitch −- energy -1.355

    Additional component ablations:

    1. Continuous Wavelet Transform (CWT) for Pitch: Predicting F0F_0 directly with MSE loss without CWT decomposition drops CMOS by −0.185-0.185 in FastSpeech 2 and −0.201-0.201 in FastSpeech 2s.
    2. Auxiliary Mel-Spectrogram Decoder in FastSpeech 2s: Removing the auxiliary mel-spectrogram decoder during FastSpeech 2s training degrades CMOS by −0.285-0.285, confirming that auxiliary mel supervision is critical for robust phoneme feature representation learning.
  10. Knowl 10 — Controllable Speech Synthesis via Variance Feature Manipulation

    model/method

    Because FastSpeech 2 and FastSpeech 2s explicitly decouple speech variations into separate duration, pitch (F0F_0), and energy inputs, speech attributes can be independently and continuously controlled at inference time without model retraining:

    • Pitch Control: Scaling the predicted pitch contour F^0(t)\hat{F}_0(t) by a scalar factor α∈[0.75,1.50]\alpha \in [0.75, 1.50] modulates voice pitch and intonation globally or locally across specific phonemes while preserving speech intelligibility and timbre.
    • Duration / Speed Control: Scaling the predicted duration values changes the repetition factor in the Length Regulator, allowing fine-grained control over speaking rate at the sentence or phoneme level.
    • Energy / Volume Control: Adjusting the predicted energy values modulates the perceived loudness and intensity of the synthesized speech.

Coverage note — None was omitted; all primary contributions, models, algorithms, and experimental results from the paper are represented.

References

  1. 1.Bistra Andreeva, Grazyna Demenko, Bernd Mȧobius, Frank Zimmerer, Jeanin Jȧugler, and Magdalena Oleskowicz-Popiel. Differences of pitch profiles in germanic and slavic languages. In Fifteenth Annual Conference of the International Speech Communication Association, 2014.
  2. 2.Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al. Deep voice: Real-time neural text-to-speech. arXiv preprint arXiv:1702.07825, 2017.
  3. 3.Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, and Tao Qin. Multispeech: Multi- speaker text to speech with transformer. In INTERSPEECH, pp. 4024–4028, 2020.
  4. 4.Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, sheng zhao, and Tie-Yan Liu. Adaspeech: Adaptive text to speech for custom voice. In International Conference on Learning Representa- tions, 2021. URL https://openreview.net/forum?id=Drynvt7gg4L.
  5. 5.Min Chu and Hu Peng. Objective measure for estimating mean opinion score of synthesized speech, April 4 2006. US Patent 7,024,362.
  6. 6.Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan. End-to-end adversarial text-to-speech. arXiv preprint arXiv:2006.03575, 2020.
  7. 7.Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. Ddsp: Differentiable digital signal processing. arXiv preprint arXiv:2001.04643, 2020.
  8. 8.Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. Tts synthesis with bidirectional lstm based recurrent neural networks. In Fifteenth Annual Conference of the International Speech Communication Association, 2014.
  9. 9.Michael Gadermayr, Maximilian Tschuchnig, Dorit Merhof, Nils Kramer, Daniel Truhn, and Burkhard Gess. An asymetric cycle-consistency loss for dealing with many-to-one mappings in image translation: A study on thigh mr scans. arXiv preprint arXiv:2004.11001, 2020.
  10. 10.Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. Deep voice 2: Multi-speaker neural text-to-speech. In Advances in neural information processing systems, pp. 2962–2970, 2017.
  11. 11.Alexander Grossmann and Jean Morlet. Decomposition of hardy functions into square integrable wavelets of constant shape. SIAM journal on mathematical analysis, 15(4):723–736, 1984.
  12. 12.Keikichi Hirose and Jianhua Tao. Speech Prosody in Speech Synthesis: Modeling and generation of prosody for high quality and flexible speech synthesis. Springer, 2015.
  13. 13.Keith Ito. The lj speech dataset. https://keithito.com/LJ-Speech-Dataset/, 2017.
  14. 14.Chrisina Jayne, Andreas Lanitis, and Chris Christodoulou. One-to-many neural network mapping techniques for face image synthesis. Expert Systems with Applications, 39(10):9778–9787, 2012.
  15. 15.Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. Glow-tts: A generative flow for text-to-speech via monotonic alignment search. arXiv preprint arXiv:2005.11129, 2020.
  16. 16.Sungwon Kim, Sang-gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon. Flowavenet: A generative flow for raw audio. arXiv preprint arXiv:1811.02155, 2018.
  17. 17.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  18. 18.Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron C Courville. Melgan: Generative adversarial networks for conditional waveform synthesis. In Advances in Neural Information Processing Systems, pp. 14881–14892, 2019.
  19. 19.Adrian Łańcucki. Fastpitch: Parallel text-to-speech with pitch prediction. arXiv preprint arXiv:2006.06873, 2020.
  20. 20.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. Neural speech synthesis with trans- former network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 6706–6713, 2019.
  21. 21.Dan Lim, Won Jang, Hyeyeong Park, Bongwan Kim, Jesam Yoon, et al. Jdi-t: Jointly trained duration informed transformer for text-to-speech without explicit alignment. arXiv preprint arXiv:2005.07799, 2020.
  22. 22.Philipos C Loizou. Speech quality assessment. In Multimedia analysis, processing and communi- cations, pp. 623–654. Springer, 2011.
  23. 23.Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Jinzhu Li, Sheng Zhao, Enhong Chen, and Tie-Yan Liu. Lightspeech: Lightweight and fast text to speech with neural architecture search. In 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021.
  24. 24.Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. Montreal forced aligner: Trainable text-speech alignment using kaldi. In Interspeech, pp. 498– 502, 2017.
  25. 25.Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao. Flow- tts: A non-autoregressive network for text to speech based on flow. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7209–7213. IEEE, 2020.
  26. 26.Huaiping Ming, Dongyan Huang, Lei Xie, Jie Wu, Minghui Dong, and Haizhou Li. Deep bidirec- tional lstm modeling of timbre and prosody for emotional voice conversion. 2016.
  27. 27.Meinard Muller. Dynamic time warping. Information retrieval for music and motion, pp. 69–84, 2007.
  28. 28.Oliver Niebuhr and Radek Skarnitzl. Measuring a speaker’s acoustic correlates of pitch–but which? a contrastive analysis based on perceived speaker charisma. In Proceedings of 19th International Congress of Phonetic Sciences, 2019.
  29. 29.Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al. Parallel wavenet: Fast high-fidelity speech synthesis. arXiv preprint arXiv:1711.10433, 2017.
  30. 30.Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao. Parallel neural text-to-speech. arXiv preprint arXiv:1905.08459, 2019.
  31. 31.Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. Deep voice 3: 2000-speaker neural text-to-speech. In International Conference on Learning Representations, 2018.
  32. 32.Wei Ping, Kainan Peng, and Jitong Chen. Clarinet: Parallel wave generation in end-to-end text-to- speech. In International Conference on Learning Representations, 2019.
  33. 33.Ryan Prenger, Rafael Valle, and Bryan Catanzaro. Waveglow: A flow-based generative network for speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3617–3621. IEEE, 2019.
  34. 34.Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. Fastspeech: Fast, robust and controllable text to speech. In Advances in Neural Information Processing Sys- tems, pp. 3165–3174, 2019.
  35. 35.Harold Ryan. Ricker, ormsby; klander, bntterwo-a choice of wavelets, 1994.
  36. 36.Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al. Natural tts synthesis by con- ditioning wavenet on mel spectrogram predictions. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4779–4783. IEEE, 2018.
  37. 37.Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu. Token-level ensemble distillation for grapheme-to-phoneme conversion. In INTERSPEECH, 2019.
  38. 38.Antti Santeri Suni, Daniel Aalto, Tuomo Raitio, Paavo Alku, Martti Vainio, et al. Wavelets for intonation modeling in hmm speech synthesis. In 8th ISCA Workshop on Speech Synthesis, Pro- ceedings, Barcelona, August 31-September 2, 2013. ISCA, 2013.
  39. 39.Franz B Tuteur. Wavelet transformations in signal detection. IFAC Proceedings Volumes, 21(9): 1061–1065, 1988.
  40. 40.Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. SSW, 125, 2016.
  41. 41.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Con- ditional image generation with pixelcnn decoders. In Advances in neural information processing systems, pp. 4790–4798, 2016.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Infor- mation Processing Systems, pp. 5998–6008, 2017.
  43. 43.Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135, 2017.
  44. 44.Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. Parallel wavegan: A fast waveform gen- eration model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6199–6203. IEEE, 2020.
  45. 45.Heiga Ze, Andrew Senior, and Mike Schuster. Statistical parametric speech synthesis using deep neural networks. In 2013 ieee international conference on acoustics, speech and signal process- ing, pp. 7962–7966. IEEE, 2013.
  46. 46.Zhen Zeng, Jianzong Wang, Ning Cheng, Tian Xia, and Jing Xiao. Aligntts: Efficient feed-forward text-to-speech system without explicit alignment. In ICASSP 2020-2020 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6714–6718. IEEE, 2020.
  47. 47.Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu. De- noispeech: Denoising text to speech with frame-level noise modeling. In 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021.
  48. 48.Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman. Toward multimodal image-to-image translation. In Advances in neural information processing systems, pp. 465–476, 2017.

Citation

MLA
Ren, Y., et al. “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech”. arXiv, 2020, http://arxiv.org/abs/2006.04558v8.
APA
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., & Liu, T.-Y. (2020). FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. arXiv. http://arxiv.org/abs/2006.04558v8
Chicago
Ren, Y., C. Hu, X. Tan, et al. 2020. “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech”. arXiv. http://arxiv.org/abs/2006.04558v8.
Harvard
Ren, Y. et al. (2020) “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2006.04558v8.
Vancouver
1. Ren Y, Hu C, Tan X, Qin T, Zhao S, Zhao Z, Liu T-Y (2020) FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. arXiv

BibTeX

@article{ren2020fastspeech,
  title = {FastSpeech 2: Fast and High-Quality End-to-End Text to Speech},
  author = {Ren, Yi and Hu, Chenxu and Tan, Xu and Qin, Tao and Zhao, Sheng and Zhao, Zhou and Liu, Tie-Yan},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2006.04558v8},
  eprint = {2006.04558}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission