Recurrent Continuous Translation Models
Nal KalchbrennerPhil Blunsom
Presents Recurrent Continuous Translation Models, an alignment-free neural machine translation framework combining convolutional sentence encoders with recurrent language models to substantially reduce translation perplexity and match traditional phrase-based systems in n-best rescoring.
Traditional statistical machine translation systems rely heavily on counting explicit phrase pairs and discrete word alignments between languages. This conventional setup suffers from severe data sparsity when dealing with rare or unseen phrases, limits cross-domain generalization, and fails to share statistical strength across semantically similar phrases. The article demonstrates that purely continuous translation models—which map continuous mathematical representations of words, phrases, and sentences without relying on phrase tables, explicit alignment segmentations, or external parsers—can effectively overcome these limitations and accurately generate translations.
The researchers designed two variations of Recurrent Continuous Translation Models, denoted as Model I and Model II. Both approaches generate target sentences using a recurrent neural network language model, which avoids restrictive assumptions about word histories. They differ in how they condition target words on the source sentence: Model I builds a single continuous vector of the entire source sentence using a convolutional network, whereas Model II breaks the source sentence into local four-word groupings and projects them onto an estimated target sentence length. The models were evaluated on English-to-French translation using the Workshop on Machine Translation dataset, comprising approximately 145,000 training sentence pairs and multiple test sets across four benchmark years (2009–2012).
The findings show substantial improvements in modeling performance and linguistic coherence. First, Model II achieved a perplexity—a metric reflecting prediction error—that was over 43% lower than a state-of-the-art alignment-based translation baseline and 40% lower than Model I. Second, when tested on randomized source word orders, Model II experienced a sharp degradation in perplexity, proving that the continuous architecture strongly captures word order, syntax, and sentence structure despite lacking explicit alignment features. Third, candidate translations directly generated by Model II demonstrated accurate grammatical agreement in verb tenses and plural forms as well as meaningful semantic transfers. Finally, when applied to rescore lists of candidate translations, a simple configuration of the proposed models combined with a single word-count adjustment matched the quality scores of an established translation system that relies on twelve complex engineered features.
These results indicate that continuous representations and neural architectures can simultaneously learn the structural rules of the target language and the semantic mapping between languages in a unified, computationally efficient pipeline. Eliminating rigid phrase tables and alignment heuristics drastically simplifies translation system design while reducing engineering overhead. Decision-makers should consider evaluating continuous representation frameworks as viable alternatives or enhancements to legacy statistical translation pipelines, with potential expansions into broader discourse contexts, multilingual models, and character-level modeling for complex languages.
Decision-makers should note several operational boundaries within the study. The experiments were conducted on a relatively small bilingual corpus of under 150,000 sentence pairs with a capped sentence length of 80 words. Additionally, direct translation generation required a sampling heuristic because full search across all possible target sequences remains computationally demanding. Organizations should conduct pilot tests on larger datasets and additional language pairs to confirm scalability and performance under production workloads before full-scale deployment.
- Paper: Statistical Phrase-Based Translation, Philipp Koehn et al. (2003). This foundational paper establishes the standard statistical phrase-based translation baseline and alignment framework that continuous translation models aim to replace.
- Paper: A Neural Probabilistic Language Model, Yoshua Bengio et al. (2003). It introduces continuous space word representations and neural language modeling, providing the core probabilistic formulation adapted for continuous translation.
- Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). It demonstrates how continuous word representations capture syntactic and semantic regularities via vector operations, motivating sentence-level continuous conditioning.
- Paper: Sequence Transduction with Recurrent Neural Networks, Alex Graves (2012). It develops early recurrent sequence transduction mechanisms that map input sequences to output sequences without requiring predefined alignments.
- Paper: Generating Text with Recurrent Neural Networks, Ilya Sutskever et al. (2011). It showcases the capability of recurrent neural networks to generate coherent sequential text conditionally and autoregressively.
- Paper: Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation, Kyunghyun Cho et al. (2014). It formalizes the RNN Encoder–Decoder architecture for statistical machine translation, directly building upon continuous representation-based sequence modeling.
- Paper: Sequence to Sequence Learning with Neural Networks, Ilya Sutskever et al. (2014). It extends continuous sequence-to-sequence generation into fully end-to-end recurrent neural translation without external phrase-based decoders.
- Paper: On the Properties of Neural Machine Translation: Encoder–Decoder Approaches, Kyunghyun Cho et al. (2014). It analyzes the structural properties and fixed-vector bottleneck limitations of encoder-decoder continuous translation architectures on long sentences.
- Paper: Neural Machine Translation by Jointly Learning to Align and Translate, Dzmitry Bahdanau et al. (2015). It resolves the fixed-length sentence embedding bottleneck of continuous translation models by introducing soft attention-based alignment.
- Paper: Effective Approaches to Attention-based Neural Machine Translation, Minh-Thang Luong et al. (2015). It refines attentional recurrent translation architectures by comparing global and local alignment mechanisms over continuous sentence representations.
- Paper: Convolutional Sequence to Sequence Learning, Jonas Gehring et al. (2017). It explores fully convolutional sequence-to-sequence translation, extending the convolutional sentence encoding concepts from early continuous models.
- Paper: Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation, Yonghui Wu et al. (2016). It scales end-to-end continuous neural machine translation to large-scale industrial production systems.
