Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

Jonathan ShenRuoming PangRon J. WeissMike SchusterNavdeep JaitlyZongheng YangZhifeng ChenYu ZhangYuxuan WangRJ Skerry-Ryan

article2017ICASSP3,166 citations

Introduces Tacotron 2, a neural text-to-speech architecture combining sequence-to-sequence mel-spectrogram prediction with a modified WaveNet vocoder to synthesize natural speech that rivals professional human recordings.

Listen

Tacotron 2 is a fully neural text-to-speech system that generates speech waveforms directly from character sequences. It addresses the long-standing difficulty of producing audio that matches the naturalness of human speech, a limitation that has persisted through earlier concatenative, statistical parametric, and hybrid neural approaches. The work matters because high-quality synthetic speech now supports applications from virtual assistants to accessibility tools, where even small gains in perceived quality affect user trust and adoption.

The document set out to demonstrate an end-to-end neural pipeline that eliminates hand-engineered linguistic features while matching or approaching the audio quality of professional recordings. The system pairs a recurrent sequence-to-sequence network that predicts 80-dimensional mel spectrograms with a modified WaveNet vocoder that converts those spectrograms into 24 kHz waveforms. Both components were trained separately on 24.6 hours of single-speaker English data; the full system was then evaluated through large-scale mean opinion score tests and side-by-side comparisons against ground-truth audio and prior systems.

The model achieved a mean opinion score of 4.53, statistically indistinguishable from the 4.58 recorded for the same speaker. It substantially outperformed production baselines: parametric synthesis scored 3.49, concatenative synthesis 4.17, and the original Tacotron with Griffin-Lim 4.00. Conditioning WaveNet on mel spectrograms rather than linguistic features allowed the vocoder to be reduced from 30 layers to as few as 12 layers while preserving quality. A post-processing network on the spectrogram predictor and alignment of training and inference features each contributed measurable gains of roughly 0.1 MOS points. Side-by-side ratings showed a small but significant listener preference for ground truth, driven mainly by occasional mispronunciations.

These results indicate that a compact acoustic intermediate representation can replace elaborate text-analysis pipelines without sacrificing naturalness, lowering both engineering effort and model complexity. The approach therefore offers a practical route to higher-quality synthesis at reduced development cost. At the same time, the system still requires training data that covers the intended domain; out-of-domain headlines produced lower scores and exposed remaining pronunciation weaknesses.

Further progress depends on larger, more diverse datasets and refinements to prosody modeling. Collecting targeted recordings for rare names and phrases, or adding explicit prosody controls, would address the dominant remaining error modes before wider deployment. The reported findings rest on a single-speaker corpus and fixed test sentences, so confidence is high for similar conditions but should be treated cautiously for new voices or genres until additional validation is performed.

  • Paper: WaveNet: A Generative Model for Raw Audio, Aäron van den Oord et al. (2016). WaveNet introduced the autoregressive raw audio generation framework that Tacotron 2 adapts into a vocoder conditioned on mel spectrograms.
Cover for Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

Abstract

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a vocoder to synthesize timedomain waveforms from those spectrograms. Our model achieves a mean opinion score (MOS) of 4.534.53 comparable to a MOS of 4.584.58 for professionally recorded speech. To validate our design choices, we present ablation studies of key components of our system and evaluate the impact of using mel spectrograms as the input to WaveNet instead of linguistic, duration, and F0F_0 features. We further demonstrate that using a compact acoustic intermediate representation enables significant simplification of the WaveNet architecture.

Table of Contents

  • 1. INTRODUCTION
  • 2. MODEL ARCHITECTURE
  • 2.1. Intermediate Feature Representation
  • 2.2. Spectrogram Prediction Network
  • 2.3. WaveNet Vocoder
  • 3. EXPERIMENTS & RESULTS
  • 3.1. Training Setup
  • 3.2. Evaluation
  • 3.3. Ablation Studies
  • 3.3.1. Predicted Features versus Ground Truth
  • 3.3.2. Linear Spectrograms
  • 3.3.3. Post-Processing Network
  • 3.3.4. Simplifying WaveNet
  • 4. CONCLUSION
  • 5. ACKNOWLEDGMENTS
  • 6. REFERENCES

Knowls

  1. Knowl 1 — Tacotron 2 Neural Text-to-Speech Architecture

    model/method

    Tacotron 2 is an end-to-end neural text-to-speech (TTS) synthesis system that generates time-domain speech waveforms directly from normalized character sequences. The system consists of two independently trained neural components:

    1. Spectrogram Prediction Network: A sequence-to-sequence recurrent neural network with location-sensitive attention that maps character embeddings to a sequence of 80-channel log-mel spectrogram frames.
    2. Modified WaveNet Vocoder: A modified, autoregressive WaveNet model that consumes the predicted mel spectrogram frames as conditioning inputs and generates 24 kHz, 16-bit time-domain audio samples.

    By adopting an 80-channel log-mel spectrogram as a compact intermediate acoustic representation, Tacotron 2 eliminates the need for hand-engineered linguistic and acoustic feature pipelines (such as phonetic transcriptions, log fundamental frequency F0F_0 predictions, phoneme duration models, and pronunciation lexicons) used in traditional TTS and standard WaveNet synthesis pipelines.

  2. Knowl 2 — Tacotron 2 Spectrogram Prediction Network

    model/method

    The spectrogram prediction network in Tacotron 2 generates an 80-channel mel spectrogram from an input sequence of text characters through an encoder-decoder architecture with location-sensitive attention:

    • Encoder: Input characters are represented as 512-dimensional learned embeddings and passed through a stack of 3 convolutional layers. Each layer contains 512 filters of shape 5×15 \times 1 (spanning 5 characters), followed by batch normalization, ReLU activations, and dropout with probability 0.50.5. The final convolutional output is fed to a single bidirectional LSTM layer containing 512 units (256 units per direction) to produce encoded text features.
    • Location-Sensitive Attention: Extends additive attention by appending cumulative attention weights from prior decoder time steps as an additional feature. Location features are extracted by convolving cumulative attention weights with 32 1D convolution filters of length 31. Inputs and location features are projected to 128-dimensional representations to generate context vectors.
    • Autoregressive Decoder: At decoder step tt, the previous spectrogram frame is passed through a pre-net consisting of 2 fully connected layers (256 ReLU units each, regularized with dropout of probability 0.50.5 at both training and inference time to serve as an information bottleneck). The pre-net output is concatenated with the attention context vector and passed into a stack of 2 unidirectional LSTM layers with 1024 units each (regularized with zoneout probability 0.10.1). The concatenation of the LSTM output and context vector is projected through a linear layer to predict the 80-dimensional mel frame for time tt.
    • Stop Token Prediction: The concatenation of the decoder LSTM output and context vector is projected to a scalar followed by a sigmoid activation, predicting the probability that sequence generation is complete. During inference, decoding terminates at the first frame where this probability exceeds 0.50.5.
    • Convolutional Post-Net: The full sequence of predicted frames is processed by a 5-layer 1D convolutional post-net (512 filters with shape 5×15 \times 1 per layer, batch normalization, tanh\tanh activations on layers 1–4, and linear activation on layer 5) that predicts a residual added to the initial predictions.

    The feature prediction network is optimized by minimizing the summed mean squared error (MSE) between ground-truth and predicted spectrograms before and after the post-net residual addition.

  3. Knowl 3 — Acoustic Mel-Spectrogram Intermediate Representation

    definition

    In Tacotron 2, the intermediate acoustic representation bridging the text-to-spectrogram network and the waveform vocoder is an 80-channel log-mel frequency spectrogram computed from 24 kHz audio waveforms:

    • Short-Time Fourier Transform (STFT): Calculated using a 50 ms window length (1200 samples at 24 kHz), a 12.5 ms frame hop (300 samples at 24 kHz), and a Hann window function.
    • Mel Filterbank: An 80-channel mel filterbank spanning 125 Hz to 7.6 kHz is applied to the STFT linear magnitude spectrum.
    • Dynamic Range Compression: Magnitudes from the mel filterbank are clipped to a minimum value of 0.010.01 and log-transformed:

    Melcompressed=log(max(Melraw,0.01))\text{Mel}_{\text{compressed}} = \log\left(\max(\text{Mel}_{\text{raw}}, 0.01)\right)

    The 12.5 ms frame hop corresponds to a frame rate of 80 Hz. Each step in the autoregressive decoder produces exactly one 80-dimensional mel frame (reduction factor r=1r=1).

  4. Knowl 4 — Modified WaveNet Vocoder with Mixture of Logistics Output

    model/method

    The neural vocoder in Tacotron 2 inverts 80-channel mel spectrograms into 24 kHz, 16-bit time-domain audio samples using a modified WaveNet architecture:

    • Dilated Convolutional Stack: 30 dilated residual convolution layers grouped into 3 dilation cycles of 10 layers each. The dilation rate for layer k{0,1,,29}k \in \{0, 1, \dots, 29\} is:

    dk=2kmod10d_k = 2^{k \bmod 10}

    • Conditioning Stack: The 80-channel mel spectrogram features (at an 80 Hz frame rate) are upsampled to the 24 kHz waveform sample rate using 2 transposed convolution upsampling layers within the conditioning network.
    • Mixture of Logistics Output: The network models waveform samples using a 10-component Mixture of Logistics (MoL) distribution. The WaveNet stack output is passed through a ReLU activation followed by a linear projection that outputs mixture weights πi\pi_i, means μi\mu_i, and log-scales sis_i for i{1,,10}i \in \{1, \dots, 10\}.
    • Loss Function: The vocoder is trained by minimizing the negative log-likelihood of ground-truth audio samples xx under the mixture distribution:

    L=logi=110πiLogistic(x;μi,exp(si))\mathcal{L} = -\log \sum_{i=1}^{10} \pi_i \, \text{Logistic}\left(x; \mu_i, \exp(s_i)\right)

    Target waveform amplitudes are scaled by a factor of 127.5 to accelerate training convergence.

  5. Knowl 5 — Tacotron 2 Two-Stage Training Protocol

    model/method

    Tacotron 2 is trained in two independent sequential stages using a single-speaker 24 kHz English speech dataset of 24.6 hours:

    1. Spectrogram Prediction Network Training:

      • Trained in teacher-forcing mode (feeding ground-truth mel spectrogram frames into the decoder pre-net).
      • Optimized using Adam (β1=0.9,β2=0.999,ϵ=106\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-6}) with a batch size of 64 on a single GPU.
      • Learning rate begins at 10310^{-3} and decays exponentially to 10510^{-5} starting after 50,000 optimization steps.
      • Regularization includes L2L_2 weight penalty with weight 10610^{-6}, dropout of 0.50.5 on convolutional and pre-net layers, and zoneout of 0.10.1 on LSTM layers.
    2. WaveNet Vocoder Training on Feature Predictions:

      • Conditioned on teacher-forced predictions generated by the trained spectrogram prediction network rather than ground-truth mel spectrograms, ensuring that the vocoder learns to invert model-generated feature characteristics while maintaining temporal alignment with ground-truth audio targets.
      • Optimized using Adam (β1=0.9,β2=0.999,ϵ=108\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}) with a fixed learning rate of 10410^{-4} and batch size 128 distributed across 32 GPUs with synchronous updates.
      • An exponentially-weighted moving average (EMA) of network weights with a decay rate of 0.9999 is maintained and used for inference.
  6. Knowl 6 — Mean Opinion Score Comparison of Tacotron 2 and Baseline TTS Systems

    empirical result

    In a subjective Mean Opinion Score (MOS) evaluation conducted on 100 evaluation sentences with at least 8 raters per sample on a 1 to 5 scale (0.5 increments), Tacotron 2 achieved an MOS comparable to professional human speech recordings and significantly outperformed traditional and previous neural TTS baselines:

    System MOS (95% CI)
    Parametric (LSTM-RNN) 3.492±0.0963.492 \pm 0.096
    Tacotron (Griffin-Lim) 4.001±0.0874.001 \pm 0.087
    Concatenative 4.166±0.0914.166 \pm 0.091
    WaveNet (Linguistic Features) 4.341±0.0514.341 \pm 0.051
    Tacotron 2 4.526±0.0664.526 \pm 0.066
    Ground Truth Recordings 4.582±0.0534.582 \pm 0.053

    In a direct side-by-side preference test between Tacotron 2 and ground-truth recordings across 800 paired ratings on a 3-3 to +3+3 scale, raters showed a slight preference for ground truth (mean score 0.270±0.155-0.270 \pm 0.155), primarily due to occasional mispronunciations in the generated audio.

  7. Knowl 7 — Impact of Training WaveNet on Predicted versus Ground-Truth Mel Spectrograms

    empirical result

    WaveNet vocoder audio quality depends on whether the conditioning mel spectrograms used during training match the feature distribution used during synthesis. An ablation evaluating combinations of training and synthesis features yielded the following Mean Opinion Scores (MOS):

    Synthesis Features
    WaveNet Training Features Predicted Ground Truth
    Predicted 4.526±0.0664.526 \pm 0.066 4.449±0.0604.449 \pm 0.060
    Ground Truth 4.362±0.0664.362 \pm 0.066 4.522±0.0554.522 \pm 0.055

    When WaveNet is trained on ground-truth mel spectrograms and then used to synthesize waveforms from predicted mel spectrograms, MOS drops from 4.5264.526 to 4.3624.362. This occurs because predicted mel spectrograms are oversmoothed due to the mean squared error loss used by the feature prediction network. Training WaveNet directly on predicted spectrograms allows the vocoder to adapt to oversmoothed inputs and synthesize natural waveforms.

  8. Knowl 8 — WaveNet Receptive Field and Depth Simplification with Mel Conditioning

    empirical result

    Because 80-channel mel-frequency spectrograms capture long-term acoustic context across frames, the layer depth and receptive field size of the WaveNet vocoder can be substantially reduced without significant loss in subjective quality:

    Total Layers Dilation Cycles Dilation Cycle Size Receptive Field (samples / ms) MOS (95% CI)
    30 3 10 6,139 / 255.8 4.526±0.0664.526 \pm 0.066
    24 4 6 505 / 21.0 4.547±0.0564.547 \pm 0.056
    12 2 6 253 / 10.5 4.481±0.0594.481 \pm 0.059
    30 30 1 61 / 2.5 3.930±0.0763.930 \pm 0.076

    Reducing the WaveNet stack from 30 layers (255.8 ms receptive field) to 24 layers (21.0 ms) or 12 layers (10.5 ms) maintains naturalness (4.5474.547 and 4.4814.481 MOS vs. 4.5264.526). However, removing dilated convolutions entirely (30 layers with dilation 1, receptive field 2.5 ms) degrades MOS to 3.9303.930, indicating that a temporal receptive field of at least 10 ms at the sample level is necessary for high audio fidelity.

  9. Knowl 9 — Impact of Post-Net and Intermediate Representation on Synthesis Quality

    empirical result

    Ablations comparing intermediate spectrogram representations, vocoder choices, and the presence of the convolutional post-net demonstrate the following:

    System MOS (95% CI)
    Tacotron 2 (1,025-dim Linear Spectrogram + Griffin-Lim) 3.944±0.0913.944 \pm 0.091
    Tacotron 2 (1,025-dim Linear Spectrogram + WaveNet) 4.510±0.0544.510 \pm 0.054
    Tacotron 2 (80-dim Mel Spectrogram + WaveNet) 4.526±0.0664.526 \pm 0.066
    • Mel vs. Linear Spectrograms: Conditioning WaveNet on 80-channel mel spectrograms achieves an MOS of 4.526±0.0664.526 \pm 0.066, matching the performance of conditioning on 1,025-dimensional linear STFT spectrograms (4.510±0.0544.510 \pm 0.054), while reducing the feature representation dimensionality by more than 12-fold.
    • Neural Vocoder vs. Griffin-Lim: Using the Griffin-Lim algorithm to invert linear spectrograms reduces MOS to 3.944±0.0913.944 \pm 0.091.
    • Post-Net Necessity: Removing the 5-layer convolutional post-net from Tacotron 2 drops MOS from 4.526±0.0664.526 \pm 0.066 to 4.429±0.0714.429 \pm 0.071, showing that post-net residual correction across past and future decoded frames provides valuable acoustic refinement even when followed by a convolutional WaveNet vocoder.
  10. Knowl 10 — Tacotron 2 Synthesis Error Modes and Generalization Limitations

    limitation

    Error analysis on a 100-sentence test set (where Tacotron 2 obtained an overall MOS of 4.354) and out-of-domain evaluation identify specific operational limitations:

    • Error Breakdown on Test Set:
      • 0 repeated word errors.
      • 1 skipped word error.
      • 6 mispronunciation errors (primarily on rare or out-of-vocabulary words).
      • 23 subjective unnatural prosody errors (e.g., incorrect syllable or word emphasis, or unnatural pitch contours).
      • 1 end-point prediction failure (occurring on the sentence containing the highest character count).
    • Out-of-Domain Generalization: On 37 out-of-domain news headlines, Tacotron 2 achieved an MOS of 4.148±0.1244.148 \pm 0.124, performing on par with a WaveNet vocoder conditioned on linguistic features (4.137±0.1284.137 \pm 0.128, side-by-side preference 0.142±0.3380.142 \pm 0.338).
    • Lexicon Reliance: Because Tacotron 2 predicts acoustic features directly from character sequences without an explicit phonetic pronunciation lexicon, handling unfamiliar proper nouns and out-of-domain vocabulary remains an open challenge requiring diverse training data.

Coverage note — A brief negative experiment modeling the spectrogram decoder output distribution using a Mixture Density Network was omitted because it was harder to train and yielded no audio quality improvement.

References

  1. 1.P. Taylor, Text-to-Speech Synthesis, Cambridge University Press, New York, NY, USA, 1st edition, 2009.
  2. 2.A. J. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in Proc. ICASSP, 1996, pp. 373–376.
  3. 3.A. W. Black and P. Taylor, “Automatically clustering similar units for unit selection in speech synthesis,” in Proc. Eurospeech, September 1997, pp. 601–604.
  4. 4.K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for HMM-based speech synthesis,” in Proc. ICASSP, 2000, pp. 1315–1318.
  5. 5.H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” Speech Communication, vol. 51, no. 11, pp. 1039–1064, 2009.
  6. 6.H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. ICASSP, 2013, pp. 7962–7966.
  7. 7.K. Tokuda, Y. Nankaku, T. Toda, H. Zen, J. Yamagishi, and K. Oura, “Speech synthesis based on hidden Markov models,” Proc. IEEE, vol. 101, no. 5, pp. 1234–1252, 2013.
  8. 8.A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” CoRR, vol. abs/1609.03499, 2016.
  9. 9.S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep voice: Real-time neural text-to-speech,” CoRR, vol. abs/1702.07825, 2017.
  10. 10.S. Ö. Arik, G. F. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” CoRR, vol. abs/1705.08947, 2017.
  11. 11.W. Ping, K. Peng, A. Gibiansky, S. Ö. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech,” CoRR, vol. abs/1710.07654, 2017.
  12. 12.Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech, Aug. 2017, pp. 4006–4010.
  13. 13.I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks.,” in Proc. NIPS, Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, Eds., 2014, pp. 3104–3112.
  14. 14.D. W. Griffin and J. S. Lim, “Signal estimation from modified short-time Fourier transform,” IEEE Transactions on Acoustics, Speech and Signal Processing, pp. 236–243, 1984.
  15. 15.A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. Interspeech, 2017, pp. 1118–1122.
  16. 16.J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2Wav: End-to-end speech synthesis,” in Proc. ICLR, 2017.
  17. 17.S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE Transactions on Acoustics, Speech and Signal Processing, vol. 28, no. 4, pp. 357 – 366, 1980.
  18. 18.S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. ICML, 2015, pp. 448–456.
  19. 19.M. Schuster and K. K. Paliwal, ‘‘Bidirectional recurrent neural networks,’’ IEEE Transactions on Signal Processing, vol. 45, no. 11, pp. 2673–2681, Nov. 1997.
  20. 20.S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, Nov. 1997.
  21. 21.J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Proc. NIPS, 2015, pp. 577–585.
  22. 22.D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. ICLR, 2015.
  23. 23.C. M. Bishop, “mixture density networks,” Tech. Rep., 1994.
  24. 24.M. Schuster, On supervised learning from sequential data with applications for speech recognition, Ph.D. thesis, Nara Institute of Science and Technology, 1999.
  25. 25.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “dropout: a simple way to prevent neural networks from overfitting.,” Journal of Machine Learning Research, vol. 15, no. 1, pp. 1929–1958, 2014.
  26. 26.D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, H. Larochelle, A. Courville, et al., “Zoneout: Regularizing RNNs by randomly preserving hidden activations,” in Proc. ICLR, 2017.
  27. 27.T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, ‘‘PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications,’’ in Proc. ICLR, 2017.
  28. 28.A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast High-Fidelity Speech Synthesis,” CoRR, vol. abs/1711.10433, Nov. 2017.
  29. 29.D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR, 2015.
  30. 30.X. Gonzalvo, S. Tazari, C.-a. Chan, M. Becker, A. Gutkin, and H. Silen, “Recent advances in Google real-time HMM-driven unit selection synthesizer,” in Proc. Interspeech, 2016.
  31. 31.H. Zen, Y. Agiomyrgiannakis, N. Egberts, F. Henderson, and P. Szczepaniak, “fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices,” in Proc. Interspeech, 2016.

Citation

MLA
Shen, J., et al. “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 4779–83, https://doi.org/10.1109/ICASSP.2018.8461368.
APA
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., Saurous, R. A., Agiomvrgiannakis, Y., & Wu, Y. (2018). Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4779–4783. https://doi.org/10.1109/ICASSP.2018.8461368
Chicago
Shen, J., R. Pang, R. J. Weiss, et al. 2018. “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4779–83. https://doi.org/10.1109/ICASSP.2018.8461368.
Harvard
Shen, J. et al. (2018) “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 4779–4783. Available at: https://doi.org/10.1109/ICASSP.2018.8461368.
Vancouver
1. Shen J, Pang R, Weiss RJ, et al (2018) Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp 4779–4783

BibTeX

@inproceedings{Shen_2018, title={Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions}, url={http://dx.doi.org/10.1109/ICASSP.2018.8461368}, DOI={10.1109/icassp.2018.8461368}, booktitle={2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, publisher={IEEE}, author={Shen, Jonathan and Pang, Ruoming and Weiss, Ron J. and Schuster, Mike and Jaitly, Navdeep and Yang, Zongheng and Chen, Zhifeng and Zhang, Yu and Wang, Yuxuan and Skerrv-Ryan, Rj and Saurous, Rif A. and Agiomvrgiannakis, Yannis and Wu, Yonghui}, year={2018}, month=Apr, pages={4779–4783} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF