Sequence Level Training with Recurrent Neural Networks

Marc'Aurelio RanzatoSumit ChopraMichael AuliWojciech Zaremba

article2015ICLR1,815 citations

Proposes a sequence-level training method that directly optimizes evaluation metrics such as BLEU and ROUGE, overcoming the error accumulation of standard token-level training while achieving beam search quality at much faster generation speeds.

Listen

Natural language generation systems, such as automated translation and text summarization tools, increasingly support critical operational and customer-facing workflows. However, standard methods for training these sequence models suffer from two structural flaws: exposure bias and metric mismatch. Models are traditionally trained to predict only the single next word using perfect reference text as input, but in production, they must generate entire sentences relying entirely on their own prior predictions. This discrepancy causes early errors to cascade rapidly throughout a sentence, while the standard word-level training objective fails to optimize for whole-sequence quality metrics like overlap accuracy. The article introduces and evaluates Mixed Incremental Cross-Entropy Reinforce, a novel training algorithm designed to eliminate exposure bias by directly optimizing sequence-level evaluation metrics while remaining stable during training.

The research evaluated this methodology across three core real-world text generation benchmarks: abstractive summarization using news corpora, German-to-English translation from transcribed presentations, and image captioning using standard vision benchmarks. The proposed framework solves the optimization instability of reinforcement learning—which typically fails when searching across tens of thousands of vocabulary words—by first initializing the model with standard supervised training, then gradually transitioning the sequence from supervised cross-entropy steps to reinforcement-based generation via an incremental curriculum schedule.

The article demonstrates significant performance and operational efficiency advantages across all evaluated benchmarks. In greedy sequence generation, the proposed approach outperformed baseline methods across every task, raising text summarization scores from 13.01 to 16.22, translation scores from 17.74 to 20.73, and image captioning scores from 27.8 to 29.16. Crucially, the model's simple, fast greedy generation produced text quality comparable to or exceeding baseline models running complex multi-path search methods. Because multi-path search requires exploring multiple hypotheses per step, the proposed model achieved equivalent or superior quality while operating at approximately ten times the generation speed.

These findings have direct operational implications for deploying automated text generation systems at scale. By generating higher quality output in a single greedy pass, organizations can drastically cut computational latency and cloud inference costs without sacrificing output fidelity. The results also establish that sequence-generating models should be aligned directly with their end-state business evaluation metrics during the training cycle rather than relying solely on post-hoc search heuristics.

Decision-makers building text generation systems should consider adopting incremental reinforcement training frameworks when inference speed and sequence quality are primary business constraints. When extreme low latency is required, deploying the model with single-pass greedy decoding provides the best trade-off between performance and compute expenditure; where absolute quality is paramount, the model can still be combined with search algorithms for marginal quality gains. Organizations should conduct targeted pilot implementations on domain-specific corpora to identify optimal scheduling parameters and transition points before broad deployment.

Confidence in these findings is supported by consistent improvements across multiple distinct sequence tasks and network architectures. A remaining limitation is that training stability depends heavily on baseline reward estimation and carefully tuned incremental schedules; unguided reinforcement learning fails to converge. Future work should focus on designing more robust baseline estimators and exploring multi-sample search methods during training to further enhance model reliability and training speed.

Cover for Sequence Level Training with Recurrent Neural Networks

Abstract

Many natural language processing applications use language models to generate text. These models are typically trained to predict the next word in a sequence, given the previous words and some context such as an image. However, at test time the model is expected to generate the entire sequence from scratch. This discrepancy makes generation brittle, as errors may accumulate along the way. We address this issue by proposing a novel sequence level training algorithm that directly optimizes the metric used at test time, such as BLEU or ROUGE. On three different tasks, our approach outperforms several strong baselines for greedy generation. The method is also competitive when these baselines employ beam search, while being several times faster.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Models
  • 3.1 Word-Level Training
  • 3.1.1 Cross Entropy Training (XENT)
  • 3.1.2 Data As Demonstrator (DAD)
  • 3.1.3 End-to-End BackProp (E2E)
  • 3.2 Sequence Level Training
  • 3.2.1 REINFORCE
  • 3.2.2 Mixed Incremental Cross-Entropy Reinforce (MIXER)
  • 4 Experiments
  • 4.1 Text Summarization
  • 4.2 Machine Translation
  • 4.3 Image Captioning
  • 4.4 Results
  • 5 Conclusions
  • References
  • 6 Supplementary Material
  • 6.1 Experiments
  • 6.1.1 Qualitative Comparison
  • 6.1.2 Hyperparameters
  • 6.1.3 Relative Gains
  • 6.2 The Attentive Encoder
  • 6.3 Beam Search Algorithm
  • 6.4 Notes

Knowls

  1. Knowl 1 — Mixed Incremental Cross-Entropy Reinforce (MIXER) Algorithm

    algorithm

    MIXER is a curriculum learning algorithm designed to train recurrent neural network (RNN) sequence generators directly against sequence-level discrete evaluation metrics (such as BLEU or ROUGE) while mitigating exposure bias and the exploration difficulty of large discrete action spaces.

    The algorithm begins by training the RNN for NXENTN^{\text{XENT}} epochs using only the standard cross-entropy (XENT) objective on ground-truth sequences, establishing an initial policy that restricts exploration to high-probability regions of the search space. Training then proceeds via an incremental annealing schedule indexed by step cutoff ss, which decreases from the maximum sequence length TT down to 1 in decrements of Δ\Delta. At each cutoff level ss, the model is trained for NXE+RN^{\text{XE+R}} epochs: for each training sequence, the first ss time steps are trained using standard XENT loss conditioned on ground-truth tokens, while the remaining T−sT - s time steps are trained using REINFORCE by sampling tokens from the model's own predictive distribution and optimizing the end-of-sequence metric reward.

    Input: Training dataset of sequences with context, maximum sequence length TT, epoch counts NXENTN^{\text{XENT}} and NXE+RN^{\text{XE+R}}, step decrement Δ\Delta
    Output: Optimized RNN parameters θ\theta
    Initialize RNN parameters θ\theta randomly
    for s=T,T−Δ,T−2Δ,…,1s = T, T - \Delta, T - 2\Delta, \dots, 1 do
        if s==Ts == T then
            Train RNN for NXENTN^{\text{XENT}} epochs using cross-entropy loss only
        else
            Train RNN for NXE+RN^{\text{XE+R}} epochs:
                Apply cross-entropy loss on ground truth tokens for the first ss steps
                Sample tokens from model distribution and apply REINFORCE gradient updates using sequence reward rr for the remaining T−sT - s steps
            end
        end
    end
    return θ\theta

    Typical hyperparameter settings include Δ∈{2,3}\Delta \in \{2, 3\}, NXENT∈{20,25}N^{\text{XENT}} \in \{20, 25\}, and NXE+R=5N^{\text{XE+R}} = 5. By the conclusion of training (s=1s = 1), the model generates entire sequences autoregressively from its own predictions under sequence-level reward optimization.

  2. Knowl 2 — Policy Gradient Formulation and Softmax Derivative for Discrete Sequence Generation

    equation

    In sequence generation framed as a reinforcement learning problem, an autoregressive recurrent neural network with parameters θ\theta acts as an agent that generates a sequence of token actions (w1g,…,wTg)(w_1^g, \dots, w_T^g) by sampling from its policy distribution pθ(wt+1g∣wtg,ht+1,ct)p_\theta(w_{t+1}^g \mid w_t^g, h_{t+1}, c_t), where ht+1h_{t+1} is the hidden state and ctc_t is an optional context vector. The objective is to minimize the negative expected sequence-level reward r(w1g,…,wTg)r(w_1^g, \dots, w_T^g) (such as BLEU or ROUGE):

    Lθ=−E(w1g,…,wTg)∼pθ[r(w1g,…,wTg)]L_\theta = -\mathbb{E}_{(w_1^g, \dots, w_T^g) \sim p_\theta} [r(w_1^g, \dots, w_T^g)]

    Approximating this expectation with a single sampled trajectory (w1g,…,wTg)(w_1^g, \dots, w_T^g), the partial derivative of the sequence loss LθL_\theta with respect to the pre-softmax activation vector oto_t at time step tt is given by:

    ∂Lθ∂ot=(r(w1g,…,wTg)−rˉt+1)(pθ(wt+1∣wtg,ht+1,ct)−1(wt+1g))\frac{\partial L_\theta}{\partial o_t} = (r(w_1^g, \dots, w_T^g) - \bar{r}_{t+1}) \left( p_\theta(w_{t+1} \mid w_t^g, h_{t+1}, c_t) - \mathbf{1}(w_{t+1}^g) \right)

    where 1(wt+1g)\mathbf{1}(w_{t+1}^g) is a one-hot indicator vector for the sampled token wt+1gw_{t+1}^g, pθ(wt+1∣wtg,ht+1,ct)=softmax(ot+1)p_\theta(w_{t+1} \mid w_t^g, h_{t+1}, c_t) = \text{softmax}(o_{t+1}), and rˉt+1∈R\bar{r}_{t+1} \in \mathbb{R} is a baseline value estimating the expected future reward from state ht+1h_{t+1}. This gradient update structurally mirrors the gradient of multi-class logistic regression, where the sampled token wt+1gw_{t+1}^g serves as a surrogate target: the chosen token is encouraged if the sequence reward exceeds the baseline (r>rˉt+1r > \bar{r}_{t+1}) and discouraged if r<rˉt+1r < \bar{r}_{t+1}.

  3. Knowl 3 — Value Baseline Estimation for Variance Reduction in Sequence REINFORCE

    model/method

    To reduce the variance of the single-sample REINFORCE policy gradient estimator in autoregressive sequence generation, an estimate of the expected future reward rˉt\bar{r}_t at time step tt is subtracted from the observed sequence reward rr.

    The baseline rˉt\bar{r}_t is computed using a linear regressor parameterized by weight vector wbasew_{\text{base}} and bias bbaseb_{\text{base}} operating on the recurrent neural network's hidden state ht∈Rdh_t \in \mathbb{R}^d:

    rˉt=wbase⊤ht+bbase\bar{r}_t = w_{\text{base}}^\top h_t + b_{\text{base}}

    Because hth_t encodes only past context and generated tokens up to step tt, the regressor provides an unbiased estimate of future rewards. The parameters of the regressor are trained online by minimizing the mean squared error:

    Lbase=∥rˉt−r∥2L_{\text{base}} = \|\bar{r}_t - r\|^2

    To prevent unstable feedback loops during joint optimization, error gradients from LbaseL_{\text{base}} are not backpropagated through the recurrent neural network's hidden states hth_t.

  4. Knowl 4 — Comparative Performance of MIXER and Baselines Across Generation Tasks

    data/table

    Across three distinct natural language generation benchmarks evaluated using greedy search decoding, MIXER outperforms standard cross-entropy training (XENT), Data As Demonstrator (DAD / scheduled sampling), and End-to-End Backprop (E2E):

    Task XENT DAD E2E MIXER
    Text Summarization (ROUGE-2) 13.01 12.18 12.78 16.22
    Machine Translation (BLEU-4) 17.74 20.12 17.77 20.73
    Image Captioning (BLEU-4) 27.80 28.16 26.42 29.16

    The benchmark configurations comprise:

    • Text Summarization: First-sentence news article to headline generation on the Gigaword dataset, evaluated by ROUGE-2 using a 128-unit Elman RNN unrolled for 15 steps.
    • Machine Translation: German-to-English translation on the IWSLT 2014 TED talks dataset, evaluated by corpus-level BLEU-4 using a 256-unit LSTM unrolled for 25 steps.
    • Image Captioning: Caption generation on MSCOCO using 1024-dimensional ImageNet CNN features and a 512-unit LSTM unrolled for 15 steps, evaluated by sentence-level BLEU-4 against 5 reference captions.

    MIXER yields improvements of 1 to 3 points over XENT across all tasks. DAD outperforms XENT on translation and captioning but degrades on summarization, while E2E consistently underperforms both DAD and MIXER.

  5. Knowl 5 — Greedy MIXER Matches or Exceeds Beam Search Baselines with Lower Latency

    empirical result

    While beam search decoding improves generation quality for all sequence models, MIXER operating with simple greedy decoding (beam size k=1k = 1) matches or surpasses the performance of standard cross-entropy (XENT), DAD, and E2E baselines operating with beam search (k=10k = 10). Specifically:

    • On Gigaword text summarization, greedy MIXER achieves a ROUGE-2 score of 16.22, exceeding XENT, DAD, and E2E with beam search (k=10k = 10), which achieve scores between 14.5 and 15.0.
    • On MSCOCO image captioning, greedy MIXER achieves a BLEU-4 score of 29.16, outperforming XENT, DAD, and E2E with beam search (k=10k = 10), which plateau around 28.5.
    • On IWSLT 2014 German-English translation, greedy MIXER (20.73 BLEU-4) is competitive with beam-search XENT (k=10k = 10, ≈20.8\approx 20.8 BLEU-4).

    Because beam search with beam size kk requires kk forward passes per time step, greedy MIXER produces text at least 10×10\times faster than beam-search baselines while maintaining equal or better accuracy. Combining MIXER with beam search (k=10k = 10) further improves scores (reaching over 22 BLEU on translation and over 17 ROUGE-2 on summarization).

  6. Knowl 6 — End-to-End Backpropagation (E2E) via Top-k Softmax Input Smoothing

    model/method

    End-to-End Backprop (E2E) is a differentiable relaxation of sequence training and beam search that exposes the recurrent model to its own predictions during training while maintaining full differentiability through backpropagation through time.

    At time step tt, instead of feeding a discrete sampled token or ground-truth word to step t+1t+1, the model passes the output distribution over vocabulary words through a kk-max layer. This operation selects the kk largest probability values, sets all other entries to zero, and renormalizes the top kk scores to sum to one:

    {it+1,j,vt+1,j}j=1k=k-max(pθ(wt+1∣wt,ht))\{i_{t+1, j}, v_{t+1, j}\}_{j=1}^k = k\text{-max}\left(p_\theta(w_{t+1} \mid w_t, h_t)\right)

    where it+1,ji_{t+1, j} are the vocabulary indices of the kk most probable words and vt+1,j∈[0,1]v_{t+1, j} \in [0, 1] are their normalized scores (such that ∑j=1kvt+1,j=1\sum_{j=1}^k v_{t+1, j} = 1). At time step t+1t+1, the model receives a weighted average of the word embeddings corresponding to these top-kk indices weighted by vt+1,jv_{t+1, j}. This continuous relaxation fuses kk possible hypotheses into a single differentiable path, allowing loss gradients to backpropagate directly through the chosen inputs across time steps.

  7. Knowl 7 — Convolutional Attentive Encoder Architecture for Sequence Generation

    model/method

    For conditional sequence generation tasks (such as abstractive summarization and machine translation), a conditioning context vector ctc_t is supplied to the recurrent decoder at each generation step tt using an attentive encoder over the source token sequence s=[w1,…,wM]s = [w_1, \dots, w_M]:

    1. Token and Positional Embeddings: Each input word wiw_i and its position index ii are mapped to learnable dd-dimensional embedding vectors wi,li∈Rdw_i, l_i \in \mathbb{R}^d, forming full token representations ai=wi+lia_i = w_i + l_i.
    2. Local Context Aggregation: An aggregate local embedding ziz_i is computed for each word by averaging the representations across a symmetric window of width qq (with q=5q = 5, using dummy padding at sequence boundaries): zi=1q∑h=−q/2q/2ai+hz_i = \frac{1}{q} \sum_{h=-q/2}^{q/2} a_{i+h}
    3. Attentive Context Computation: Attention weights αj,t\alpha_{j, t} over source positions j∈{1,…,M}j \in \{1, \dots, M\} are computed via the dot product of zjz_j with the decoder's current hidden state hth_t: αj,t=exp⁡(zj⋅ht)∑i=1Mexp⁡(zi⋅ht)\alpha_{j, t} = \frac{\exp(z_j \cdot h_t)}{\sum_{i=1}^M \exp(z_i \cdot h_t)} The conditioning vector ctc_t provided to the decoder hidden state transition is the weighted sum of source word embeddings: ct=∑j=1Mαj,twjc_t = \sum_{j=1}^M \alpha_{j, t} w_j
  8. Knowl 8 — Metric Specificity in Sequence-Level Reward Optimization

    empirical result

    Directly optimizing sequence models for a specific test evaluation metric (such as BLEU or ROUGE) substantially improves performance on that metric, but training on a mismatched sequence metric degrades performance compared to targeted optimization.

    On the Gigaword text summarization benchmark:

    • When evaluated on ROUGE-2: MIXER trained directly with ROUGE-2 reward achieves 16.22, whereas MIXER trained with BLEU reward achieves 15.10, and standard cross-entropy (XENT) achieves 13.01.
    • When evaluated on BLEU: MIXER trained directly with BLEU reward achieves 9.32, whereas standard cross-entropy (XENT) achieves 8.16, and MIXER trained with ROUGE-2 reward drops to 5.80.

    This demonstrates that optimizing surrogate word-level losses or mismatched sequence objectives fails to align model predictions with the target evaluation metric.

  9. Knowl 9 — Taxonomy of Sequence Training Paradigms by Exposure Bias, Differentiability, and Loss Level

    theoretical result

    Sequence generation training frameworks can be categorized along three foundational axes:

    1. Avoids Exposure Bias: Whether the training process conditions on the model's own predicted outputs during training (matching inference conditions) rather than strictly clamping inputs to ground-truth tokens.
    2. End-to-End Differentiable: Whether error gradients backpropagate directly through the selected or sampled input representations across decoding steps.
    3. Sequence-Level Loss: Whether the objective directly evaluates entire generated sequences against holistic, non-differentiable discrete metrics (e.g., BLEU, ROUGE) rather than decomposing the loss into independent token-level next-word predictions.

    The comparison across training paradigms is:

    • Cross-Entropy (XENT): Suffers from exposure bias (No), does not backpropagate through input choices (No), operates at token level (No).
    • Data As Demonstrator (DAD / Scheduled Sampling): Avoids exposure bias by mixing model predictions into inputs (Yes), does not backpropagate through sampled tokens (No), operates at token level (No).
    • End-to-End Backprop (E2E): Avoids exposure bias via top-kk predictive inputs (Yes), is fully differentiable end-to-end via input smoothing (Yes), operates at token level (No).
    • MIXER: Avoids exposure bias by sampling from model distribution during the reinforcement learning phase (Yes), is trained end-to-end via policy gradients (Yes), operates at sequence level (Yes).
  10. Knowl 10 — Intractability of Uncurriculated REINFORCE in Large Discrete Action Spaces

    limitation

    Standard REINFORCE initialized from random policy parameters fails to converge for text generation applications due to the cardinality of natural language action spaces.

    In natural language generation, the action space at each decoding step corresponds to the vocabulary size ∣V∣|\mathcal{V}| (typically ∣V∣≥104|\mathcal{V}| \ge 10^4), resulting in an overall trajectory space of size O(∣V∣T)O(|\mathcal{V}|^T) for sequence length TT. A randomly initialized model exhibits perplexity close to the vocabulary size (≈10,000\approx 10{,}000), yielding an extraordinarily low probability of sampling meaningful sequences that receive non-zero reward. Under such high branching factors, exploration from a random policy produces near-zero expected reward gradients, preventing pure REINFORCE or hybrid XENT-REINFORCE losses without curriculum learning from learning. Pre-training with cross-entropy reduces model perplexity to approximately 50, providing a viable starting policy from which incremental reinforcement learning can successfully deviate.

Coverage note — Standard beam search decoding (Algorithm 2) was omitted as an independent knowl because it represents standard prior art rather than a novel contribution of the paper.

References

  1. 1.Auli, M. and Gao, J. Decoder integration and expected bleu training for recurrent neural network language models. In Proc. of ACL, June 2014.
  2. 2.Ba, J.L., Mnih, V., and Kavukcuoglu, K. Multiple object recognition with visual attention.
  3. 3.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In ICLR, 2015.
  4. 4.Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N. Scheduled sampling for sequence prediction with recurrent neural networks. In NIPS, 2015.
  5. 5.Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., , and Federico, M. Report on the 11th iwslt evaluation campaign. In Proc. of IWSLT, 2014.
  6. 6.Daume III, H., Langford, J., and Marcu, D. Search-based structured prediction as classification. Machine Learning Journal, 2009.
  7. 7.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. Imagenet: a large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition, 2009.
  8. 8.Elman, Jeffrey L. Finding structure in time. Cognitive Science, 14(2):179–211, 1990.
  9. 9.Graff, D., Kong, J., Chen, K., and Maeda, K. English gigaword. Technical report, 2003.
  10. 10.Graves, A. and Jaitly, N. Towards end-to-end speech recognition with recurrent neural networks. In ICML, 2014.
  11. 11.He, X. and Deng, L. Maximum expected bleu training of phrase and lexicon translation models. In ACL, 2012.
  12. 12.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  13. 13.Kneser, Reinhard and Ney, Hermann. Improved backing-off for M-gram language modeling. In Proc. of the International Conference on Acoustics, Speech, and Signal Processing, pp. 181–184, May 1995.
  14. 14.Koehn, Philipp, Hoang, Hieu, Birch, Alexandra, Callison-Burch, Chris, Federico, Marcello, Bertoldi, Nicola, Cowan, Brooke, Shen, Wade, Moran, Christine, Zens, Richard, Dyer, Chris, Bojar, Ondrej, Constantin, Alexandra, and Herbst, Evan. Moses: Open source toolkit for statistical machine translation. In Proc. of ACL Demo and Poster Sessions, Jun 2007.
  15. 15.Liang, Percy, Bouchard-Côté, Alexandre, Taskar, Ben, and Klein, Dan. An end-to-end discriminative approach to machine translation. In acl-coling2006, pp. 761–768, Jul 2006.
  16. 16.Lin, C.Y. and Hovy, E.H. Automatic evaluation of summaries using n-gram co-occurrence statistics. In HLT-NAACL, 2003.
  17. 17.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C.L. Microsoft coco: Common objects in context. Technical report, 2014.
  18. 18.McAllester, D., Hazan, T., and Keshet, J. Direct loss minimization for structured prediction. In NIPS, 2010.
  19. 19.Mikolov, T., Karafit, M., Burget, L., Cernock, J., and Khudanpur, S. Recurrent neural network based language model. In INTERSPEECH, 2010.
  20. 20.Mnih, V., Heess N., Graves, A., and Kavukcuoglu, K. Recurrent models of visual attention. In NIPS, 2014.
  21. 21.Morin, F. and Bengio, Y. Hierarchical probabilistic neural network language model. In AISTATS, 2005.
  22. 22.Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  23. 23.Reiter, E. and Dale, R. Building natural language generation systems. Cambridge university press, 2000.
  24. 24.Ross, S., Gordon, G.J., and Bagnell, J.A. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS, 2011.
  25. 25.Rosti, Antti-Veikko I, Zhang, Bing, Matsoukas, Spyros, and Schwartz, Richard. Expected bleu training for graphs: Bbn system description for wmt11 system combination task. In Proc. of WMT, pp. 159–165. Association for Computational Linguistics, July 2011.
  26. 26.Rumelhart, D.E., Hinton, G.E., and Williams, R.J. Learning internal representations by backpropagating errors. Nature, 323:533–536, 1986.
  27. 27.Rush, A.M., Chopra, S., and Weston, J. A neural attention model for abstractive sentence summarization. In EMNLP, 2015.
  28. 28.Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc. Sequence to sequence learning with neural networks. In Proc. of NIPS, 2014.
  29. 29.Sutton, R.S. and Barto, A.G. Reinforcement learning: An introduction. MIT Press, 1988.
  30. 30.Venkatraman, A., Hebert, M., and Bagnell, J.A. Improving multi-step prediction of learned time series models. In AAAI, 2015.
  31. 31.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 1992.
  32. 32.Xu, X., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015.
  33. 33.Zaremba, W. and Sutskever, I. Reinforcement learning neural turing machines. Technical report, 2015.

Citation

MLA
Ranzato, M., et al. “Sequence Level Training with Recurrent Neural Networks”. arXiv, 2015, http://arxiv.org/abs/1511.06732v7.
APA
Ranzato, M., Chopra, S., Auli, M., & Zaremba, W. (2015). Sequence Level Training with Recurrent Neural Networks. arXiv. http://arxiv.org/abs/1511.06732v7
Chicago
Ranzato, M., S. Chopra, M. Auli, and W. Zaremba. 2015. “Sequence Level Training with Recurrent Neural Networks”. arXiv. http://arxiv.org/abs/1511.06732v7.
Harvard
Ranzato, M. et al. (2015) “Sequence Level Training with Recurrent Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1511.06732v7.
Vancouver
1. Ranzato M, Chopra S, Auli M, Zaremba W (2015) Sequence Level Training with Recurrent Neural Networks. arXiv

BibTeX

@article{ranzato2015sequence,
  title = {Sequence Level Training with Recurrent Neural Networks},
  author = {Ranzato, Marc'Aurelio and Chopra, Sumit and Auli, Michael and Zaremba, Wojciech},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1511.06732v7},
  eprint = {1511.06732}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors