FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Yi RenChenxu HuXu TanTao QinSheng ZhaoZhou ZhaoTie-Yan Liu
Introduces a non-autoregressive text-to-speech framework that trains directly on ground-truth acoustic targets conditioned on pitch, energy, and duration, eliminating complex distillation pipelines to achieve three times faster training and voice quality surpassing autoregressive models.
Text-to-speech technology is essential for applications ranging from voice assistants to digital media production. However, synthesizing speech from text poses a fundamental one-to-many mapping challenge because a single text sequence can be spoken with diverse variations in duration, pitch, and energy. While earlier non-autoregressive models like FastSpeech improved synthesis speed over traditional autoregressive systems, they relied on a complex two-stage training process involving knowledge distillation from a teacher model. This pipeline resulted in long training times, inaccurate duration extraction, and information loss that degraded voice quality.
The article aims to resolve these limitations by developing FastSpeech 2 and FastSpeech 2s. The objective was to demonstrate that directly training on ground-truth targets while explicitly conditioning on speech variations—specifically pitch, energy, and accurate duration—simplifies training, accelerates synthesis, and delivers superior audio quality.
To evaluate this approach, the researchers conducted experiments using the 24-hour single-speaker LJSpeech dataset containing 13,100 English audio clips. The architecture replaces teacher-student distillation by training directly on actual speech targets. It incorporates a variance adaptor that extracts ground-truth phoneme durations via forced alignment and analyzes pitch contours using continuous wavelet transforms to capture frequency variations. In addition, the authors developed FastSpeech 2s, which extends the model into a fully non-autoregressive system capable of generating raw speech waveforms directly from text using adversarial training, bypassing intermediate acoustic spectrograms.
The findings show substantial improvements across training efficiency, inference speed, and voice quality. FastSpeech 2 achieved a roughly 3.12-fold speed-up in training time compared to FastSpeech, reducing acoustic model training from 53.12 hours to 17.02 hours. In terms of generation speed, FastSpeech 2 and FastSpeech 2s synthesized audio approximately 47.8 and 51.8 times faster than standard autoregressive Transformer models. Furthermore, subjective listening evaluations demonstrated that FastSpeech 2 surpassed both FastSpeech and traditional autoregressive baselines in perceptual voice quality, while FastSpeech 2s matched autoregressive quality. Ablation analyses confirmed that forced-alignment duration, continuous wavelet pitch modeling, and energy conditioning each significantly contributed to prosody and naturalness.
These results indicate that explicit variance conditioning effectively solves the one-to-many mapping problem without requiring complex distillation pipelines. For practitioners, this reduces compute costs, shortens development lifecycles, and yields voice synthesis that is highly controllable in pitch and volume while maintaining ultra-low latency for real-time applications.
Decision-makers and engineering teams should consider adopting the FastSpeech 2 architecture for voice generation pipelines that require high throughput and robust quality. Teams with strict latency constraints or resource limits can deploy FastSpeech 2s to eliminate separate vocoder dependencies. However, future development should aim to build fully end-to-end alignment tools directly into the architecture, as the current framework still relies on external tools for forced alignment and pitch extraction.
- Paper: Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions, Jonathan Shen et al. (2017). Tacotron 2 established the dominant mel-spectrogram prediction and neural vocoding paradigm that FastSpeech 2 directly compares against, builds on, and replaces with non-autoregressive generation.
- Paper: Tacotron: Towards End-to-End Speech Synthesis, Yuxuan Wang et al. (2017). Tacotron introduced the foundational end-to-end sequence-to-sequence neural text-to-speech architecture upon which modern acoustic models are constructed.
- Paper: WaveNet: A Generative Model for Raw Audio, Aäron van den Oord et al. (2016). WaveNet introduced neural autoregressive waveform synthesis, establishing the baseline audio fidelity and vocoding standards that FastSpeech 2 and FastSpeech 2s seek to match with faster parallel generation.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). The Transformer provides the foundational feed-forward self-attention mechanisms that form the core structural blocks of the FastSpeech architecture.
- Paper: Sequence-Level Knowledge Distillation, Yoon Kim et al. (2016). Sequence-level knowledge distillation defines the teacher-student training paradigm that the original FastSpeech relied upon and that FastSpeech 2 explicitly redesigns to avoid information loss.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN provides the high-fidelity, non-autoregressive neural vocoder commonly paired with acoustic models like FastSpeech 2 to achieve fully parallelized, high-speed waveform generation from predicted mel-spectrograms.
- Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). DiffWave extends the non-autoregressive audio synthesis domain to diffusion probabilistic models, providing an alternative parallel vocoder for generating waveforms from mel-spectrogram representations.
- Paper: Qwen3-TTS Technical Report, Hangrui Hu et al. (2026). Qwen3-TTS builds upon modern neural speech synthesis paradigms, extending fast end-to-end generation principles to large-scale multilingual modeling and zero-shot voice cloning.
