Recurrent Continuous Translation Models

Nal KalchbrennerPhil Blunsom

article2013EMNLP1,479 citations

Presents Recurrent Continuous Translation Models, an alignment-free neural machine translation framework combining convolutional sentence encoders with recurrent language models to substantially reduce translation perplexity and match traditional phrase-based systems in n-best rescoring.

Listen

Traditional statistical machine translation systems rely heavily on counting explicit phrase pairs and discrete word alignments between languages. This conventional setup suffers from severe data sparsity when dealing with rare or unseen phrases, limits cross-domain generalization, and fails to share statistical strength across semantically similar phrases. The article demonstrates that purely continuous translation models—which map continuous mathematical representations of words, phrases, and sentences without relying on phrase tables, explicit alignment segmentations, or external parsers—can effectively overcome these limitations and accurately generate translations.

The researchers designed two variations of Recurrent Continuous Translation Models, denoted as Model I and Model II. Both approaches generate target sentences using a recurrent neural network language model, which avoids restrictive assumptions about word histories. They differ in how they condition target words on the source sentence: Model I builds a single continuous vector of the entire source sentence using a convolutional network, whereas Model II breaks the source sentence into local four-word groupings and projects them onto an estimated target sentence length. The models were evaluated on English-to-French translation using the Workshop on Machine Translation dataset, comprising approximately 145,000 training sentence pairs and multiple test sets across four benchmark years (2009–2012).

The findings show substantial improvements in modeling performance and linguistic coherence. First, Model II achieved a perplexity—a metric reflecting prediction error—that was over 43% lower than a state-of-the-art alignment-based translation baseline and 40% lower than Model I. Second, when tested on randomized source word orders, Model II experienced a sharp degradation in perplexity, proving that the continuous architecture strongly captures word order, syntax, and sentence structure despite lacking explicit alignment features. Third, candidate translations directly generated by Model II demonstrated accurate grammatical agreement in verb tenses and plural forms as well as meaningful semantic transfers. Finally, when applied to rescore lists of candidate translations, a simple configuration of the proposed models combined with a single word-count adjustment matched the quality scores of an established translation system that relies on twelve complex engineered features.

These results indicate that continuous representations and neural architectures can simultaneously learn the structural rules of the target language and the semantic mapping between languages in a unified, computationally efficient pipeline. Eliminating rigid phrase tables and alignment heuristics drastically simplifies translation system design while reducing engineering overhead. Decision-makers should consider evaluating continuous representation frameworks as viable alternatives or enhancements to legacy statistical translation pipelines, with potential expansions into broader discourse contexts, multilingual models, and character-level modeling for complex languages.

Decision-makers should note several operational boundaries within the study. The experiments were conducted on a relatively small bilingual corpus of under 150,000 sentence pairs with a capped sentence length of 80 words. Additionally, direct translation generation required a sampling heuristic because full search across all possible target sequences remains computationally demanding. Organizations should conduct pilot tests on larger datasets and additional language pairs to confirm scalability and performance under production workloads before full-scale deployment.

Kalchbrenner et al (2013).pdf
  • Paper: Statistical Phrase-Based Translation, Philipp Koehn et al. (2003). This foundational paper establishes the standard statistical phrase-based translation baseline and alignment framework that continuous translation models aim to replace.
  • Paper: A Neural Probabilistic Language Model, Yoshua Bengio et al. (2003). It introduces continuous space word representations and neural language modeling, providing the core probabilistic formulation adapted for continuous translation.
  • Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). It demonstrates how continuous word representations capture syntactic and semantic regularities via vector operations, motivating sentence-level continuous conditioning.
  • Paper: Sequence Transduction with Recurrent Neural Networks, Alex Graves (2012). It develops early recurrent sequence transduction mechanisms that map input sequences to output sequences without requiring predefined alignments.
  • Paper: Generating Text with Recurrent Neural Networks, Ilya Sutskever et al. (2011). It showcases the capability of recurrent neural networks to generate coherent sequential text conditionally and autoregressively.
Cover for Recurrent Continuous Translation Models

Abstract

We introduce a class of probabilistic continuous translation models called Recurrent Continuous Translation Models that are purely based on continuous representations for words, phrases and sentences and do not rely on alignments or phrasal translation units. The models have a generation and a conditioning aspect. The generation of the translation is modelled with a target Recurrent Language Model, whereas the conditioning on the source sentence is modelled with a Convolutional Sentence Model. Through various experiments, we show first that our models obtain a perplexity with respect to gold translations that is > 43% lower than that of state-of-the-art alignment-based translation models. Secondly, we show that they are remarkably sensitive to the word order, syntax, and meaning of the source sentence despite lacking alignments. Finally we show that they match a state-of-the-art system when rescoring n-best lists of translations.

Table of Contents

  • 1 Introduction
  • 2 Framework
  • 2.1 Recurrent Language Model
  • 3 Recurrent Continuous Translation Model I
  • 3.1 Convolutional Sentence Model
  • 3.2 RCTMI
  • 4 Recurrent Continuous Translation Model II
  • 4.1 Convolutional n -gram model
  • 4.2 RCTMII
  • 5 Experiments
  • 5.1 Training
  • 5.1.1 Data sets
  • 5.1.2 Model hyperparameters
  • 5.1.3 Objective and optimisation
  • 5.2 Perplexity of gold translations
  • 5.3 Sensitivity to source sentence structure
  • 5.3.1 Generating from the RCTM II
  • 5.4 Rescoring and BLEU Evaluation
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Recurrent Continuous Translation Model Framework

    model/method

    The Recurrent Continuous Translation Model (RCTM) framework models the translation of a source sentence e=(e1,…,ek)e = (e_1, \dots, e_k) into a target sentence f=(f1,…,fm)f = (f_1, \dots, f_m) as a direct conditional probability distribution over target sentence sequences without introducing latent word alignments, phrase segmentations, or Markov dependency assumptions:

    P(f∣e)=∏i=1mP(fi∣f1:i−1,e)P(f|e) = \prod_{i=1}^m P(f_i \mid f_{1:i-1}, e)

    where f1:i−1=(f1,…,fi−1)f_{1:i-1} = (f_1, \dots, f_{i-1}) denotes the sequence of preceding target words.

    The framework combines a generative target recurrent neural language model (RLM) that captures target language syntax and long-range dependencies with a conditioning architecture based on convolutional representations of the source sentence. Computing P(f∣e)P(f|e) requires a number of matrix-vector products that scales linearly with the lengths of the source and target sentences.

  2. Knowl 2 — Recurrent Continuous Translation Model I (RCTM I)

    model/method

    Recurrent Continuous Translation Model I (RCTM I) conditions the generation of each target word on a single continuous vector representing the entire source sentence. The source sentence representation e∈Rq×1e \in \mathbb{R}^{q \times 1} is generated by a Convolutional Sentence Model (CSM), csm(e)\text{csm}(e).

    Given source sentence representation s=S⋅csm(e)s = S \cdot \text{csm}(e) with transformation matrix S∈Rq×qS \in \mathbb{R}^{q \times q}, target word one-hot indicators v(fi)∈{0,1}∣VF∣×1v(f_i) \in \{0, 1\}^{|V^F| \times 1}, input word embedding matrix I∈Rq×∣VF∣I \in \mathbb{R}^{q \times |V^F|}, recurrent transition matrix R∈Rq×qR \in \mathbb{R}^{q \times q}, and output matrix O∈R∣VF∣×qO \in \mathbb{R}^{|V^F| \times q}, the hidden states hi∈Rq×1h_i \in \mathbb{R}^{q \times 1} and output score vectors oi∈R∣VF∣×1o_i \in \mathbb{R}^{|V^F| \times 1} are computed recursively for 1<i<m1 < i < m as:

    s=S⋅csm(e)s = S \cdot \text{csm}(e)

    h1=σ(I⋅v(f1)+s)h_1 = \sigma(I \cdot v(f_1) + s)

    hi+1=σ(R⋅hi+I⋅v(fi+1)+s)h_{i+1} = \sigma(R \cdot h_i + I \cdot v(f_{i+1}) + s)

    oi+1=O⋅hio_{i+1} = O \cdot h_i

    where σ\sigma is an element-wise non-linear activation function (such as tanh⁡\tanh), and bias vectors bhb_h and bob_o are included in the state and output computations. The conditional next-word probability distribution is given by the softmax function:

    P(fi+1=v∣f1:i,e)=exp⁡(oi+1,v)∑v′=1∣VF∣exp⁡(oi+1,v′)P(f_{i+1} = v \mid f_{1:i}, e) = \frac{\exp(o_{i+1, v})}{\sum_{v'=1}^{|V^F|} \exp(o_{i+1, v'})}

    In RCTM I, the conditioning vector ss exerts a uniform additive influence across every decoding step, and target sentence length is determined implicitly by the language model.

  3. Knowl 3 — Recurrent Continuous Translation Model II (RCTM II)

    model/method

    Recurrent Continuous Translation Model II (RCTM II) explicitly models target sentence length mm and generates target words conditioned on position-specific representations derived by unfolding source nn-gram features.

    The conditional probability P(f∣e)P(f|e) is factored into length prediction and length-conditioned translation:

    P(f∣e)=P(f∣m,e)⋅P(m∣e)=(∏i=1mP(fi+1∣f1:i,m,e))⋅P(m∣e)P(f|e) = P(f \mid m, e) \cdot P(m \mid e) = \left( \prod_{i=1}^m P(f_{i+1} \mid f_{1:i}, m, e) \right) \cdot P(m \mid e)

    The target length distribution P(m∣e)P(m|e) is parameterized conditioned on the source sentence length k=∣e∣k = |e| via a Poisson distribution:

    P(m∣e)=P(m∣k)=Poisson(λk)P(m \mid e) = P(m \mid k) = \text{Poisson}(\lambda k)

    where λ\lambda is a scaling parameter.

    The conditioning representation is computed in three steps:

    1. A Convolutional nn-gram Model (CGM) extracts an nn-gram representation matrix (n=4n = 4) from source sentence ee:

    Eg=cgm(e,4)E^g = \text{cgm}(e, 4)

    1. The source nn-gram vectors are projected across languages via transformation matrix T∈Rq×qT \in \mathbb{R}^{q \times q}:

    F:,jg=σ(T⋅E:,jg)F^g_{:, j} = \sigma(T \cdot E^g_{:, j})

    1. An inverse convolutional model (icgm) unfolds FgF^g into a sequence of mm column vectors matching the target sentence length:

    F=icgm(Fg,m)∈Rq×mF = \text{icgm}(F^g, m) \in \mathbb{R}^{q \times m}

    The target recurrent language model hidden states are then conditioned on the corresponding column vector F:,iF_{:, i}:

    h1=σ(I⋅v(f1)+S⋅F:,1)h_1 = \sigma(I \cdot v(f_1) + S \cdot F_{:, 1})

    hi+1=σ(R⋅hi+I⋅v(fi+1)+S⋅F:,i+1)h_{i+1} = \sigma(R \cdot h_i + I \cdot v(f_{i+1}) + S \cdot F_{:, i+1})

    oi+1=O⋅hio_{i+1} = O \cdot h_i

    where S∈Rq×qS \in \mathbb{R}^{q \times q} is a sentence transformation matrix, II is the input vocabulary projection, RR is the recurrent matrix, and OO is the output vocabulary matrix.

  4. Knowl 4 — Convolutional Sentence Model (CSM)

    model/method

    The Convolutional Sentence Model (CSM) computes a fixed-dimensional continuous representation e∈Rq×1e \in \mathbb{R}^{q \times 1} for a sentence e=(e1,…,ek)e = (e_1, \dots, e_k) from its word embedding matrix Ee∈Rq×kE^e \in \mathbb{R}^{q \times k}, where the ii-th column E:,ie=v(ei)∈Rq×1E^e_{:, i} = v(e_i) \in \mathbb{R}^{q \times 1} is the continuous representation of word eie_i.

    The architecture uses a sequence of convolution weight matrices (Ki)2≤i≤r(K^i)_{2 \le i \le r}, where each Ki∈Rq×iK^i \in \mathbb{R}^{q \times i} acts as a bank of 1D convolutional feature detectors across word embedding rows. The maximum kernel depth is r=⌈2N⌉r = \lceil \sqrt{2N} \rceil, where NN is the maximum sentence length in the training set. For any feature matrix M∈Rq×jM \in \mathbb{R}^{q \times j} (j≥ij \ge i), the convolution Ki∗M∈Rq×(j−i+1)K^i * M \in \mathbb{R}^{q \times (j - i + 1)} is defined column-wise by:

    (Ki∗M):,a=∑c=1iK:,ci⊙M:,a+c−1(K^i * M)_{:, a} = \sum_{c=1}^i K^i_{:, c} \odot M_{:, a+c-1}

    where ⊙\odot denotes the element-wise Hadamard vector product.

    The hierarchical sentence representation is computed iteratively:

    E1e=EeE_1^e = E^e

    Ei+1e=σ(Ki+1∗Eie)E_{i+1}^e = \sigma(K^{i+1} * E_i^e)

    When the number of columns in EieE_i^e becomes smaller than the width i+1i+1 of the next filter matrix Ki+1K^{i+1}, a top projection matrix LjL^j from a distinct sequence (Li)2≤i≤r(L^i)_{2 \le i \le r} (where LjL^j has column dimension matching EieE_i^e) is applied to project the representation directly to the final sentence vector e∈Rq×1e \in \mathbb{R}^{q \times 1}.

  5. Knowl 5 — Convolutional n-gram Model and Inverse CGM

    model/method

    The Convolutional nn-gram Model (CGM), denoted cgm(e,n)\text{cgm}(e, n), truncates the Convolutional Sentence Model at the layer where each column vector represents an nn-gram of span nn. For a sentence ee, the nn-gram representations correspond to an intermediate feature matrix EieE_i^e whose columns capture local contextual spans of width nn (specifically n=4n=4 in RCTM II).

    The Inverse Convolutional nn-gram Model (icgm), denoted icgm(Fg,m)\text{icgm}(F^g, m), reverses this convolutional hierarchy. Given a transformed source nn-gram matrix Fg∈Rq×(k−n+1)F^g \in \mathbb{R}^{q \times (k - n + 1)} and a target sequence length mm, icgm unfolds FgF^g through inverted convolutional weight matrix sequences (Ji)2≤i≤s(J^i)_{2 \le i \le s} and (Hi)2≤i≤s(H^i)_{2 \le i \le s} into a sequence of mm word-level target representations F∈Rq×mF \in \mathbb{R}^{q \times m}.

    If a test sentence requires a filter width larger than the maximum width trained during training, the required weight matrix is factorized into a product of smaller trained weight matrices (e.g., a filter of width 10 is factorized into one of width 9 and one of width 2).

  6. Knowl 6 — End-to-End Joint Training and Optimization of RCTMs

    model/method

    RCTMs are trained end-to-end to minimize the average cross-entropy error of predicted target words given source sentences, regularized with an L2L_2 weight penalty:

    L(Θ)=−1∣D∣∑(e,f)∈D∑i=1∣f∣log⁡P(fi∣f1:i−1,e)+λ22∥Θ∥22\mathcal{L}(\Theta) = -\frac{1}{|D|} \sum_{(e, f) \in D} \sum_{i=1}^{|f|} \log P(f_i \mid f_{1:i-1}, e) + \frac{\lambda_2}{2} \|\Theta\|_2^2

    where DD is the bilingual training corpus and Θ\Theta encompasses all model parameters, including source and target word embeddings, convolutional kernels (Ki,Li,Ji,Hi)(K^i, L^i, J^i, H^i), recurrent matrices (R,I,O)(R, I, O), and projection matrices (S,T)(S, T).

    Training details include:

    • Gradients at the target recurrent output layer are backpropagated through time (BPTT) for d=6d = 6 unravelling steps.
    • Error accumulated at the hidden layers is backpropagated through projection matrix SS and the convolutional layers (CSM/CGM) into the input source word embeddings v(ei)v(e_i), with all weights randomly initialized and jointly learned.
    • Optimization is carried out using mini-batch adaptive gradient descent (AdaGrad).
    • Target vocabulary prediction speed is optimized by factorizing the target vocabulary VFV^F into 256 classes.
  7. Knowl 7 — Perplexity Comparison of Translation Models on WMT-NT

    data/table

    Translation quality was evaluated by measuring perplexity on gold English-to-French translations across the WMT News Test (WMT-NT) benchmarks for 2009 through 2012. Models were trained on 144,953 sentence pairs from the WMT 2013 News Commentary corpus with vocabulary sizes ∣VE∣=25,403|V^E| = 25,403 and ∣VF∣=34,831|V^F| = 34,831, hidden vector dimension q=256q = 256, and BPTT depth d=6d = 6.

    Model 2009 2010 2011 2012
    KN-5 (5-gram LM) 218 213 222 225
    RLM (Target LM) 178 169 178 181
    IBM Model 1 207 200 188 197
    FA-IBM 2 (Fast-Aligner) 153 146 135 144
    RCTM I 143 134 140 142
    RCTM II 86 77 76 77

    RCTM II achieved a perplexity more than 43% lower than the alignment-based FA-IBM 2 baseline and approximately 40% lower than RCTM I, demonstrating the effectiveness of position-specific convolutional conditioning over whole-sentence conditioning.

  8. Knowl 8 — Sensitivity of RCTM II to Source Word Order

    data/table

    To evaluate whether RCTM II relies on sentence structure or functions as a bag-of-words model, the words in the English source sentences were randomly permuted in both the training and test sets of WMT-NT (2009–2012), and target translation perplexity was evaluated.

    Source Condition 2009 2010 2011 2012
    RCTM II (Natural Word Order) 86 77 76 77
    RCTM II (Randomly Permuted Source) 174 168 175 178

    Permuting source word order more than doubled the test perplexity (increasing it from 76–86 to 168–178, which is comparable to an unconditional target language model), confirming that RCTM II is highly sensitive to source word order and syntactic structure.

  9. Knowl 9 — N-Best Translation Rescoring Performance and BLEU Evaluation

    data/table

    RCTM models were evaluated on the task of rescoring 1000-best translation candidate lists generated by the statistical machine translation system cdec on the WMT-NT 2009–2012 English-to-French test sets. For RCTM I and RCTM II, the log probability assigned by the model to each candidate translation was linearly interpolated with a single Word Penalty (WP) feature tuned on the WMT-NT 2008 validation set.

    System / Model 2009 2010 2011 2012
    RCTM I + WP 19.7 21.1 22.5 21.5
    RCTM II + WP 19.8 21.1 22.5 21.7
    cdec (12 engineered features) 19.9 21.2 22.6 21.8

    The combination of an RCTM probability and a single word penalty matched the translation quality (BLEU scores within 0.1–0.2 points) of the state-of-the-art cdec baseline, which utilizes 12 engineered features (including five alignment-based translation models and two language models).

  10. Knowl 10 — Autoregressive Sampling Translation Generation from RCTM II

    algorithm

    Direct translation generation from RCTM II without external alignment decoders is performed by autoregressively sampling sentences of a target length mm (set to the length of the reference translation) and ranking them by their model probability.

    Input: Source sentence e=(e1,…,ek)e = (e_1, \dots, e_k), target length mm, sample budget K=2000K = 2000, vocabulary VFV^F, top candidate truncation size C=5C = 5
    Output: Ranked list of target translation hypotheses
    Compute source 4-gram representations: Eg=cgm(e,4)E^g = \text{cgm}(e, 4)
    Compute cross-lingual intermediate representations: F:,jg=σ(T⋅E:,jg)F^g_{:, j} = \sigma(T \cdot E^g_{:, j})
    Unfold to target length representation matrix: F=icgm(Fg,m)∈Rq×mF = \text{icgm}(F^g, m) \in \mathbb{R}^{q \times m}
    for s=1s = 1 to KK do
        Initialize target sequence hypothesis: f(s)=()f^{(s)} = ()
        for i=1i = 1 to mm do
            if i==1i == 1 then
                h1=σ(I⋅v(f1(s))+S⋅F:,1)h_1 = \sigma(I \cdot v(f_1^{(s)}) + S \cdot F_{:, 1})
            else
                hi=σ(R⋅hi−1+I⋅v(fi(s))+S⋅F:,i)h_i = \sigma(R \cdot h_{i-1} + I \cdot v(f_i^{(s)}) + S \cdot F_{:, i})
            Compute target word scores: oi=O⋅hi−1o_i = O \cdot h_{i-1} (or O⋅h1O \cdot h_1 for step 1)
            Compute full vocabulary distribution: P(v)=exp⁡(oi,v)∑u∈VFexp⁡(oi,u)P(v) = \frac{\exp(o_{i, v})}{\sum_{u \in V^F} \exp(o_{i, u})}
            Restrict distribution P(v)P(v) to the top CC highest probability words and renormalize
            Sample word fi(s)f_i^{(s)} from the restricted distribution
            Append fi(s)f_i^{(s)} to f(s)f^{(s)}
        Compute total sequence probability: P(f(s)∣e)=P(m∣e)∏i=1mP(fi(s)∣f1:i−1(s),m,e)P(f^{(s)} \mid e) = P(m \mid e) \prod_{i=1}^m P(f_i^{(s)} \mid f_{1:i-1}^{(s)}, m, e)
    Sort all sampled candidate hypotheses f(1),…,f(K)f^{(1)}, \dots, f^{(K)} in descending order of P(f(s)∣e)P(f^{(s)} \mid e)
    return Ranked candidate list

    Translations sampled and ranked using this method demonstrate correct grammatical agreements (gender, singular/plural inflections, tense) and semantic transfer between English source sentences and French candidate outputs.

Coverage note — None was omitted; all key contributions including model formulations (RCTM I, RCTM II, CSM, CGM/icgm), end-to-end training setup, perplexity experiments, permutation sensitivity tests, rescoring BLEU evaluations, and generation sampling algorithms are fully covered.

References

  1. 1.Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155.
  2. 2.Peter F. Brown, Vincent J.Della Pietra, Stephen A. Della Pietra, and Robert. L. Mercer. 1993. The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics, 19:263–311.
  3. 3.R. Collobert and J. Weston. 2008. A unified architecture for natural language processing: Deep neural networks with multitask learning. In International Conference on Machine Learning, ICML.
  4. 4.John Duchi, Elad Hazan, and Yoram Singer. 2011. Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res., 12:2121–2159, July.
  5. 5.Chris Dyer, Jonathan Weese, Hendra Setiawan, Adam Lopez, Ferhan Ture, Vladimir Eidelman, Juri Ganitkevitch, Phil Blunsom, and Philip Resnik. 2010. cdec: A decoder, alignment, and learning framework for finitestate and context-free translation models. In Proceedings of the ACL 2010 System Demonstrations, pages 7–12. Association for Computational Linguistics.
  6. 6.Chris Dyer, Victor Chahuneau, and Noah A. Smith. 2013. A simple, fast, and effective reparameterization of ibm model 2. In Proc. of NAACL.
  7. 7.Edward Grefenstette, Mehrnoosh Sadrzadeh, Stephen Clark, Bob Coecke, and Stephen Pulman. 2011. Concrete sentence spaces for compositional distributional models of meaning. CoRR, abs/1101.0309.
  8. 8.Karl Moritz Hermann and Phil Blunsom. 2013. The Role of Syntax in Vector Space Models of Compositional Semantics. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Sofia, Bulgaria, August. Association for Computational Linguistics. Forthcoming.
  9. 9.Nal Kalchbrenner and Phil Blunsom. 2013. Recurrent Convolutional Neural Networks for Discourse Compositionality. In Proceedings of the Workshop on Continuous Vector Space Models and their Compositionality, Sofia, Bulgaria, August. Association for Computational Linguistics.
  10. 10.Hai Son Le, Alexandre Allauzen, and Francois Yvon. 2012. Continuous space translation models with neural networks. In HLT-NAACL, pages 39–48.
  11. 11.Tomas Mikolov and Geoffrey Zweig. 2012. Context dependent recurrent neural network language model. In SLT, pages 234–239.
  12. 12.Tomas Mikolov, Martin Karafiat, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Takao Kobayashi, Keikichi Hirose, and Satoshi Nakamura, editors, INTERSPEECH, pages 1045–1048. ISCA.
  13. 13.Tomas Mikolov, Stefan Kombrink, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur. 2011. Extensions of recurrent neural network language model. In ICASSP, pages 5528–5531. IEEE.
  14. 14.Holger Schwenk, Daniel Dechelotte, and Jean-Luc Gauvain. 2006. Continuous space language models for statistical machine translation. In ACL.
  15. 15.Holger Schwenk. 2012. Continuous space translation models for phrase-based statistical machine translation. In COLING (Posters), pages 1071–1080.
  16. 16.Richard Socher, Eric H. Huang, Jeffrey Pennin, Andrew Y. Ng, and Christopher D. Manning. 2011. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection. In J. Shawe-Taylor, R.S. Zemel, P. Bartlett, F.C.N. Pereira, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems 24, pages 801–809.
  17. 17.Richard Socher, Brody Huval, Christopher D. Manning, and Andrew Y. Ng. 2012. Semantic Compositionality Through Recursive Matrix-Vector Spaces. In Proceedings of the 2012 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  18. 18.Ilya Sutskever, James Martens, and Geoffrey E. Hinton. 2011. Generating text with recurrent neural networks. In Lise Getoor and Tobias Scheffer, editors, ICML, pages 1017–1024. Omnipress.

Citation

MLA
Kalchbrenner, N., and P. Blunsom. “Recurrent Continuous Translation Models”. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 2013, pp. 1700–09, https://aclanthology.org/D13-1176/.
APA
Kalchbrenner, N., & Blunsom, P. (2013). Recurrent Continuous Translation Models. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 1700–1709. https://aclanthology.org/D13-1176/
Chicago
Kalchbrenner, N., and P. Blunsom. 2013. “Recurrent Continuous Translation Models”. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, 1700–1709. https://aclanthology.org/D13-1176/.
Harvard
Kalchbrenner, N. and Blunsom, P. (2013) “Recurrent Continuous Translation Models”, Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1700–1709. Available at: https://aclanthology.org/D13-1176/.
Vancouver
1. Kalchbrenner N, Blunsom P (2013) Recurrent Continuous Translation Models. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1700–1709

BibTeX

@inproceedings{kalchbrenner-blunsom-2013-recurrent,
    title = "Recurrent Continuous Translation Models",
    author = "Kalchbrenner, Nal  and
      Blunsom, Phil",
    editor = "Yarowsky, David  and
      Baldwin, Timothy  and
      Korhonen, Anna  and
      Livescu, Karen  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing",
    month = oct,
    year = "2013",
    address = "Seattle, Washington, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/D13-1176/",
    pages = "1700--1709"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-nc-sa/4.0/