ESPnet: End-to-End Speech Processing Toolkit

Shinji WatanabeTakaaki HoriShigeki KaritaTomoki HayashiJiro NishitobaYuya UnnoNelson Enrique Yalta SoplinJahn HeymannMatthew WiesnerNanxin Chen

article2018Interspeech1,742 citations

Presents ESPnet, an open-source platform that bridges PyTorch-based dynamic neural models with Kaldi-style data pipelines to enable reproducible, high-accuracy end-to-end speech recognition and processing across standard benchmarks.

Listen

Building automated speech recognition systems has traditionally required complex, multi-stage pipelines that combine separate acoustic models, pronunciation dictionaries, and language models. While end-to-end deep learning architectures promise to replace this fragmented setup with a single neural network, existing open-source platforms have largely lacked complete data pipelines, flexible modeling techniques, or compatibility with established speech standards.

The article introduces and evaluates ESPnet, a new open-source software platform designed to simplify and accelerate end-to-end speech processing. It demonstrates the platform's architectural framework and benchmarks its recognition accuracy and computational efficiency across multiple standard speech datasets.

To evaluate the system, the authors tested ESPnet across several public benchmarks, including the English Wall Street Journal task, the Corpus of Spontaneous Japanese, and the HKUST Mandarin Chinese corpus. The platform integrates dynamic deep learning frameworks (PyTorch and Chainer) with the established data-processing workflows of the Kaldi toolkit. ESPnet utilizes a hybrid architecture that combines connectionist temporal classification (CTC)—which enforces chronological alignment—with attention mechanisms that handle broader context, supported by external language model integration.

Experimental results show that the hybrid CTC and attention approach consistently improves transcription accuracy over single-technique models. On non-alphabetic languages such as Japanese and Mandarin, ESPnet matched or surpassed the accuracy of traditional hybrid systems, achieving an error rate of 28.3% on the HKUST benchmark without requiring external pronunciation dictionaries. ESPnet also achieved major gains in development efficiency: its core codebase consists of roughly 5,400 lines of Python—a dramatic reduction compared to legacy systems containing tens or hundreds of thousands of lines of code. Furthermore, training speed improved substantially, with the PyTorch implementation completing the Wall Street Journal benchmark in 5 hours on a single GPU, compared to 120 hours across 10 GPUs in previously published end-to-end setups.

These findings indicate that end-to-end speech recognition can significantly reduce development complexity, engineering overhead, and maintenance costs without sacrificing performance in character-based languages. Organizations can streamline deployment by eliminating pronunciation dictionaries and complex intermediate models. However, for smaller English datasets, traditional legacy architectures still achieve lower word error rates, indicating that end-to-end models remain sensitive to training data volume in alphabetic languages.

Technical leaders and development teams should consider adopting ESPnet for multilingual speech applications and character-dense languages where end-to-end systems are already competitive with legacy tools. For smaller English language deployments, teams should weigh the trade-off between the operational simplicity of an end-to-end model and the higher accuracy of conventional hybrid systems, or explore data augmentation techniques to bridge the performance gap.

The conclusions are well-supported across the evaluated benchmarks, though confidence is highest for non-alphabetic languages. The primary limitation is the performance gap on constrained English datasets due to data sparseness. Further development—including ongoing work on multi-GPU scaling, speech enhancement, and advanced data augmentation—will be necessary to fully match state-of-the-art hybrid performance across all English operating conditions.

Cover for ESPnet: End-to-End Speech Processing Toolkit

Abstract

This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Related studies
  • 3 Functionality
  • 3.1 Kaldi style data preprocessing
  • 3.2 Attention-based encoder-decoder
  • 3.2.1 Encoder
  • 3.2.2 Attention
  • 3.3 Hybrid CTC/attention
  • 3.3.1 Multiobjective training
  • 3.3.2 Warp CTC
  • 3.3.3 Joint decoding
  • 3.4 Use of language model
  • 3.5 ASR setup in adverse environments
  • 4 Implementation
  • 4.1 Standard recipe flow
  • 4.2 Code lines
  • 5 Experiments
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — ESPnet Toolkit Architecture and Design Principles

    model/method

    ESPnet is an open-source end-to-end speech processing toolkit designed primarily for automatic speech recognition (ASR). Its software architecture centers on four main design principles:

    • Dynamic Neural Network Backends: Neural network training and inference are implemented in Python, allowing users to select either Chainer or PyTorch as the underlying execution engine.
    • Kaldi Ecosystem Integration: Data preparation formats, acoustic feature extraction routines (such as 80-dimensional log-mel filterbanks plus 3-dimensional pitch features), and recipe directory layouts follow Kaldi conventions, enabling direct benchmarking against Kaldi hybrid systems.
    • Unified End-to-End Pipeline: Rather than managing disparate acoustic models, hidden Markov model (HMM) states, context-dependent phonemes, pronunciation lexicons, and Finite State Transducers (FSTs), the architecture maps acoustic features directly to characters or subwords with a single neural network.
    • Hybrid CTC/Attention Paradigm: Unifies Connectionist Temporal Classification (CTC) and attention-based encoder-decoder networks to leverage the alignment stability and fast convergence of CTC alongside the sequence modeling power of attention.
  2. Knowl 2 — Hybrid CTC/Attention Multiobjective Training

    equation

    To improve alignment robustness and accelerate training convergence, ESPnet optimizes a shared encoder network jointly with both CTC and attention-based decoder objectives:

    L=αLctc+(1−α)Latt\mathcal{L} = \alpha \mathcal{L}_{\text{ctc}} + (1 - \alpha) \mathcal{L}_{\text{att}}

    where:

    • Lctc\mathcal{L}_{\text{ctc}} is the Connectionist Temporal Classification loss function computed on the encoder representations.
    • Latt\mathcal{L}_{\text{att}} is the standard cross-entropy loss from the autoregressive attention decoder.
    • α∈[0,1]\alpha \in [0, 1] is a tunable interpolation weight, typically set to α=0.5\alpha = 0.5 for equal weighting.

    To prevent overconfident predictions, unigram label smoothing can be applied to the target distribution during training by distributing probability mass to non-target labels proportionally to their unigram frequency. CTC loss computation is accelerated using the Warp CTC library.

  3. Knowl 3 — Joint CTC/Attention Beam Search Decoding with Language Model Fusion

    equation

    During inference, ESPnet combines scores from the attention decoder, the CTC prefix scorer, and an optional recurrent neural network language model (RNNLM) in a one-pass output-synchronous beam search. Given the encoder representation sequence h1:T′h_{1:T'} and recognized prefix history y1:n−1y_{1:n-1}, the combined log-probability for predicting the next output token yny_n is:

    log⁡p(yn∣y1:n−1,h1:T′)=log⁡phyb(yn∣y1:n−1,h1:T′)+βlog⁡plm(yn∣y1:n−1)\log p(y_n \mid y_{1:n-1}, h_{1:T'}) = \log p_{\text{hyb}}(y_n \mid y_{1:n-1}, h_{1:T'}) + \beta \log p_{\text{lm}}(y_n \mid y_{1:n-1})

    where the hybrid acoustic-decoder probability is defined as:

    log⁡phyb(yn∣y1:n−1,h1:T′)=αlog⁡pctc(yn∣y1:n−1,h1:T′)+(1−α)log⁡patt(yn∣y1:n−1,h1:T′)\log p_{\text{hyb}}(y_n \mid y_{1:n-1}, h_{1:T'}) = \alpha \log p_{\text{ctc}}(y_n \mid y_{1:n-1}, h_{1:T'}) + (1 - \alpha) \log p_{\text{att}}(y_n \mid y_{1:n-1}, h_{1:T'})

    In these equations:

    • patt(yn∣y1:n−1,h1:T′)p_{\text{att}}(y_n \mid y_{1:n-1}, h_{1:T'}) is the probability predicted by the attention decoder.
    • pctc(yn∣y1:n−1,h1:T′)p_{\text{ctc}}(y_n \mid y_{1:n-1}, h_{1:T'}) is the CTC prefix probability.
    • plm(yn∣y1:n−1)p_{\text{lm}}(y_n \mid y_{1:n-1}) is the character- or word-level language model probability.
    • α∈[0,1]\alpha \in [0, 1] balances CTC and attention scores (typically matching the training interpolation factor α=0.5\alpha = 0.5).
    • β≥0\beta \ge 0 is a scaling hyperparameter controlling the strength of the external language model fusion.
  4. Knowl 4 — Neural Encoder and Attention Modules in ESPnet

    model/method

    ESPnet provides several modular neural acoustic encoder and attention mechanisms:

    1. Acoustic Encoders:

      • Pyramid BLSTM: A bidirectional LSTM encoder with temporal subsampling that reduces the acoustic sequence length TT of input features o1:To_{1:T} to subsampled representations h1:T′h_{1:T'} (T′<TT' < T): h1:T′=BLSTM(o1:T)h_{1:T'} = \text{BLSTM}(o_{1:T})
      • VGG2-BLSTM: Initial convolutional layers based on the first two blocks of VGG (VGG2\text{VGG}_2) that perform 2D time-frequency downsampling, followed by stacked BLSTM layers: h1:T′=BLSTM(VGG2(o1:T))h_{1:T'} = \text{BLSTM}(\text{VGG}_2(o_{1:T})) This architecture generally outperforms pyramid BLSTM alone.
    2. Attention Mechanisms:

      • Location-Aware Attention: Uses previous attention alignments as features to bias subsequent alignments toward monotonic speech progressions (the default mechanism).
      • Dot-Product Attention: A computationally lighter and faster alternative.
      • Extended Attention Types: The PyTorch backend includes over 11 attention functions, such as additive attention, coverage mechanisms, and multi-head attention.
  5. Knowl 5 — Standard Six-Stage Recipe Workflow

    algorithm

    ESPnet structures reproducible speech recognition experiments via standard shell scripts (run.sh) organized into six execution stages:

    Stage 0 (Data Preparation):
        Format corpus transcripts, speaker IDs, and audio file lists following Kaldi conventions (e.g., using data_prep.sh)
    Stage 1 (Feature Extraction):
        Extract acoustic features using Kaldi utilities (typically 80-dimensional log-Mel filterbanks plus 3-dimensional pitch)
    Stage 2 (Data Conversion for ESPnet):
        Convert Kaldi directory metadata, transcriptions, sequence lengths, and speaker/language IDs into a unified JSON format (data.json)
    Stage 3 (Language Model Training):
        Train a token-level (e.g., character-based) RNNLM or LSTMLM on text corpora using Chainer or PyTorch (optional)
    Stage 4 (End-to-End ASR Training):
        Train the hybrid CTC/attention sequence-to-sequence model on data.json using Chainer or PyTorch backends
    Stage 5 (Recognition and Scoring):
        Execute joint CTC/attention beam search decoding combined with the trained RNNLM on test sets, evaluating CER and WER via Sclite
  6. Knowl 6 — Wall Street Journal (WSJ) Performance and Training Efficiency

    data/table

    Ablation and comparative experiments on the Wall Street Journal (WSJ) benchmark demonstrate the performance gains from key ESPnet components, alongside training wall-clock times across backends:

    Method Metric dev93 eval92
    ESPnet with VGG2-BLSTM CER (%) 10.1 7.6
    + BLSTM layers (4→64 \rightarrow 6) CER (%) 8.5 5.9
    + char-LSTMLM CER (%) 8.3 5.2
    + joint decoding CER (%) 5.5 3.8
    + label smoothing CER (%) 5.3 3.6
    WER (%) 12.4 8.9
    seq2seq + CNN (no LM) WER (%) N/A 10.5
    seq2seq + FST word LM CER (%) N/A 3.9
    WER (%) N/A 9.3
    CTC + FST word LM WER (%) N/A 7.3
    Method Wall Clock Time # GPUs
    ESPnet (Chainer) 20 hours 1
    ESPnet (PyTorch) 5 hours 1
    seq2seq + CNN 120 hours 10

    Deepening the BLSTM encoder from 4 to 6 layers, integrating character-level LSTMLM shallow fusion, using joint CTC/attention decoding, and adding unigram label smoothing progressively reduce eval92 CER from 7.6% to 3.6% and WER to 8.9%. Furthermore, PyTorch-based training finishes in 5 hours on a single NVIDIA GTX 1080 Ti GPU, compared to 120 hours across 10 GPUs for a prior sequence-to-sequence CNN model.

  7. Knowl 7 — Speech Recognition Performance on CSJ and HKUST Benchmarks

    data/table

    On ideographic languages (Japanese and Mandarin Chinese), character sequence lengths are shorter than in alphabetic scripts, allowing ESPnet's lexicon-free models to achieve performance comparable or superior to standard hybrid DNN/HMM systems:

    Corpus of Spontaneous Japanese (CSJ) eval1 (CER %) eval2 (CER %) eval3 (CER %)
    ESPnet (1 GPU) 8.7 6.2 6.9
    ESPnet (5 GPUs) 8.5 6.1 6.8
    HMM/DNN (Kaldi nnet1) 9.0 7.2 9.6
    CTC-syllable 9.4 7.3 7.5
    HKUST Mandarin CTS eval (CER %)
    ESPnet 28.3
    HMM/LSTM (Kaldi nnet3) 33.5
    CTC with language model 34.8
    HMM/TDNN, LF MMI 28.2

    On the 581-hour CSJ dataset, ESPnet achieves 8.5%, 6.1%, and 6.8% CER on eval1, eval2, and eval3, outperforming Kaldi's nnet1 hybrid DNN/HMM (9.0%, 7.2%, 9.6%). With multi-GPU acceleration (5 GPUs), the 581-hour CSJ training completes in 26 hours. On HKUST Mandarin CTS, ESPnet achieves 28.3% CER, outperforming the Kaldi nnet3 HMM/LSTM baseline (33.5%) and approaching the 28.2% CER of lattice-free MMI-trained HMM/TDNN models.

  8. Knowl 8 — Codebase Complexity Comparison Across ASR Toolkits

    data/table

    ESPnet substantially reduces codebase complexity compared to traditional hybrid ASR toolkits by delegating gradient computation to dynamic neural network frameworks and feature extraction/data handling to Kaldi:

    Toolkit # Lines of Main Source Code Implementation Language
    Kaldi 330K C++
    Julius 60K C
    ESPnet 5.4K Python

    While legacy toolkits require tens to hundreds of thousands of lines of C/C++ code to manage HMM states, triphone clustering, lexicons, and FST decoding graphs, ESPnet implements the entire neural model representation in approximately 1,000 lines of Python and the beam search decoder in under 500 lines.

Coverage note — Brief mentions of ongoing/supported recipes for other corpora (AMI, CHiME-4/5, Librispeech, TED-LIUM, VoxForge) were omitted as standalone knowls because the paper provides detailed benchmark tables and ablation numbers only for WSJ, CSJ, and HKUST.

References

  1. 1.D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), Dec. 2011.
  2. 2.Y. Young, G. Evermann, D. Kershaw, G. Moore, J. Odell, D. Ollason, D. Povey, V. Valtchev, and P. Woodland, “The htk book,” Cambridge university engineering department, vol. 3, p. 175, 2002. [Online]. Available: http://htk.eng.cam.ac.uk/
  3. 3.K.-F. Lee, H.-W. Hon, and R. Reddy, “An overview of the SPHINX speech recognition system,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 38, no. 1, pp. 35–45, 1990. [Online]. Available: http://cmusphinx.sourceforge.net/
  4. 4.A. Lee, T. Kawahara, and K. Shikano, “Julius an open source realtime large vocabulary recognition engine,” in Proc. Eurospeech, 2001, pp. 1691–1694.
  5. 5.D. Rybach, C. Gollan, G. Heigold, B. Hoffmeister, J. Lo¨of, ¨ R. Schl¨uter, and H. Ney, “The RWTH aachen university ¨ open source speech recognition system.” in Interspeech, 2009, pp. 2111–2114. [Online]. Available: https://www-i6.informatik.rwth-aachen.de/rwth-asr/
  6. 6.A. Stolcke et al., “SRILM-an extensible language modeling toolkit.” in Interspeech, vol. 2002, 2002, pp. 901–904. [Online]. Available: http://www.speech.sri.com/projects/srilm/
  7. 7.G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al., “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, 2012.
  8. 8.S. Tokui, K. Oono, S. Hido, and J. Clayton, “Chainer: a nextgeneration open source framework for deep learning,” in Proceedings of workshop on machine learning systems (LearningSys) in the twenty-ninth annual conference on neural information processing systems (NIPS), vol. 5, 2015.
  9. 9.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” in Proceedings of The future of gradientbased machine learning software and techniques (Autodiff) in the twenty-ninth annual conference on neural information processing systems (NIPS), 2017.
  10. 10.A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in International Conference on Machine Learning (ICML), 2014, pp. 1764–1772.
  11. 11.Y. Miao, M. Gowayyed, and F. Metze, “‘EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2015, pp. 167–174.
  12. 12.D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos et al., “Deep speech 2: End-to-end speech recognition in english and mandarin,” arXiv preprint arXiv:1512.02595, 2015. [Online]. Available: https://github.com/baidu-research/ba-dls-deepspeech
  13. 13.J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio, “‘End-toend continuous speech recognition using attention-based recurrent NN: First results,” arXiv preprint arXiv:1412.1602, 2014.
  14. 14.L. Lu, X. Zhang, and S. Renals, “‘On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 5060–5064.
  15. 15.W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015.
  16. 16.C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina et al., “Stateof-the-art speech recognition with sequence-to-sequence models,” arXiv preprint arXiv:1712.01769, 2017.
  17. 17.S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “‘Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 8, pp. 1240–1253, 2017.
  18. 18.D. B. Paul and J. M. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Proceedings of the workshop on Speech and Natural Language. Association for Computational Linguistics, 1992, pp. 357–362.
  19. 19.V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “‘Librispeech: an ASR corpus based on public domain audio books,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210.
  20. 20.A. Rousseau, P. Deleglise, and Y. Esteve, “‘TED-LIUM: an auto- ´ matic speech recognition dedicated corpus.” in International Conference on Language Resources and Evaluation (LREC), 2012, pp. 125–129.
  21. 21.K. Maekawa, H. Koiso, S. Furui, and H. Isahara, “Spontaneous speech corpus of Japanese,” in International Conference on Language Resources and Evaluation (LREC), vol. 2, 2000, pp. 947–952.
  22. 22.T. Hain, L. Burget, J. Dines, G. Garau, V. Wan, M. Karafi, J. Vepa, and M. Lincoln, “The AMI system for the transcription of speech in meetings,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2007, pp. 357–360.
  23. 23.Y. Liu, P. Fung, Y. Yang, C. Cieri, S. Huang, and D. Graff, “‘HKUST/MTS: A very large scale Mandarin telephone speech corpus,” in Chinese Spoken Language Processing. Springer, 2006, pp. 724–735.
  24. 24.“‘VoxForge,” http://www.voxforge.org/.
  25. 25.E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,” Computer Speech & Language, vol. 46, pp. 535–557, 2017.
  26. 26.J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ‘CHiME speech separation and recognition challenge: Dataset, task and baselines,” in Interspeech, 2018, (submitting).
  27. 27.D. Povey, V. Peddinti, D. Galvez, P. Ghahrmani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “‘Purely sequence-trained neural networks for asr based on lattice-free MMI,” in Interspeech, 2016, pp. 2751–2755.
  28. 28.A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng, “‘Lexicon-free conversational speech recognition with neural networks,” in Proceedings the North American Chapter of the Association for Computational Linguistics (NAACL), 2015. [Online]. Available: https://github.com/amaas/stanford-ctc
  29. 29.D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “‘End-to-end attention-based large vocabulary speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), March 2016, pp. 4945–4949. [Online]. Available: https://github.com/rizar/attention-lvcsr
  30. 30.G. Klein, Y. Kim, Y. Deng, J. Senellart, and A. M. Rush, “‘Opennmt: Open-source toolkit for neural machine translation,” arXiv preprint arXiv:1701.02810, 2017.
  31. 31.T. Ochiai, S. Watanabe, T. Hori, and J. R. Hershey, “‘Multichannel end-to-end speech recognition,” in Proceedings of the 34th International Conference on Machine Learning (ICML), vol. 70, 2017, pp. 2632–2641.
  32. 32.K. Simonyan and A. Zisserman, “‘Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  33. 33.Y. Zhang, W. Chan, and N. Jaitly, “‘Very deep convolutional networks for end-to-end speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, pp. 4845–4849.
  34. 34.T. Hori, S. Watanabe, Y. Zhang, and W. Chan, “Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,” in Interspeech, 2017, pp. 949–953.
  35. 35.J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “‘Attention-based models for speech recognition,” in Advances in Neural Information Processing Systems (NIPS), 2015, pp. 577–585.
  36. 36.M.-T. Luong, H. Pham, and C. D. Manning, “‘Effective approaches to attention-based neural machine translation,” arXiv preprint arXiv:1508.04025, 2015.
  37. 37.D. Bahdanau, K. Cho, and Y. Bengio, “‘Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014.
  38. 38.A. See, P. J. Liu, and C. D. Manning, “Get to the point: Summarization with pointer-generator networks,” arXiv preprint arXiv:1704.04368, 2017.
  39. 39.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “‘Attention is all you need,” in Advances in Neural Information Processing Systems, 2017, pp. 6000–6010.
  40. 40.G. Pereyra, G. Tucker, J. Chorowski, Ł. Kaiser, and G. Hinton, “‘Regularizing neural networks by penalizing confident output distributions,” arXiv preprint arXiv:1701.06548, 2017.
  41. 41.C¸ . G¨ulc¸ehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, ¨ F. Bougares, H. Schwenk, and Y. Bengio, “‘On using monolingual corpora in neural machine translation,” arXiv e-prints, vol. abs/1503.03535, Mar. 2015.
  42. 42.S. Watanabe, T. Hori, and J. Hershey, “‘Language independent end-to-end architecture for joint language identification and speech recognition,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2017, pp. 265–269.
  43. 43.N. Kanda, X. Lu, and H. Kawai, “Maximum a posteriori based decoding for CTC acoustic models,” in Interspeech, 2016, pp. 1868–1872.

Citation

MLA
Watanabe, S., et al. “ESPnet: End-to-End Speech Processing Toolkit”. arXiv, 2018, http://arxiv.org/abs/1804.00015v1.
APA
Watanabe, S., Hori, T., Karita, S., Hayashi, T., Nishitoba, J., Unno, Y., Soplin, N. E. Y., Heymann, J., Wiesner, M., Chen, N., Renduchintala, A., & Ochiai, T. (2018). ESPnet: End-to-End Speech Processing Toolkit. arXiv. http://arxiv.org/abs/1804.00015v1
Chicago
Watanabe, S., T. Hori, S. Karita, et al. 2018. “ESPnet: End-to-End Speech Processing Toolkit”. arXiv. http://arxiv.org/abs/1804.00015v1.
Harvard
Watanabe, S. et al. (2018) “ESPnet: End-to-End Speech Processing Toolkit”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.00015v1.
Vancouver
1. Watanabe S, Hori T, Karita S, et al (2018) ESPnet: End-to-End Speech Processing Toolkit. arXiv

BibTeX

@article{watanabe2018espnet,
  title = {ESPnet: End-to-End Speech Processing Toolkit},
  author = {Watanabe, Shinji and Hori, Takaaki and Karita, Shigeki and Hayashi, Tomoki and Nishitoba, Jiro and Unno, Yuya and Soplin, Nelson Enrique Yalta and Heymann, Jahn and Wiesner, Matthew and Chen, Nanxin and Renduchintala, Adithya and Ochiai, Tsubasa},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.00015v1},
  eprint = {1804.00015}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF