Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation

Melvin JohnsonMike SchusterQuoc V. LeMaxim KrikunYonghui WuZhifeng ChenNikhil ThoratFernanda ViégasMartin WattenbergGreg Corrado

article2016TACL2,256 citations

Demonstrates that prepending a target language token enables a single neural machine translation model to translate across multiple languages and achieve zero-shot translation between language pairs never seen together during training.

Listen

Scaling machine translation across dozens or hundreds of languages is operationally complex and resource-intensive. Supporting over 100 languages with separate models for each language pair would require thousands of individual systems, creating severe computational, maintenance, and serving bottlenecks. The article evaluates a solution that consolidates multiple translation directions into a single Neural Machine Translation (NMT) model without altering the underlying architecture. By prepending an artificial text token to the input sentence to specify the intended target language and using a shared vocabulary across all languages, the approach aims to simplify deployment, improve translation quality for low-resource languages, and enable direct translation between language pairs that have never been seen together during training.

To demonstrate this method, the authors conducted experiments using standard public benchmarks (WMT) and massive internal production datasets covering diverse languages such as Japanese, Korean, Spanish, and Portuguese. They tested many-to-one, one-to-many, and many-to-many translation configurations, as well as large-scale deployments consolidating up to 12 production language pairs into a single system. The evaluations measured standard translation quality metrics (BLEU scores) across different model capacities and analyzed internal network representations to examine how the system processes semantic information across languages.

The findings show that multilingual training frequently matches or exceeds the quality of standalone models, particularly when translating from multiple sources into a single target language or when supporting low-resource language pairs that benefit from shared data. Consolidating 12 separate language pairs into a single multilingual model yielded translation quality within 2.5% to 5.6% of dedicated models, despite using up to five times fewer parameters and requiring only about one-twelfth of the combined training time. Crucially, the system demonstrated true zero-shot translation, successfully translating between language pairs (such as Portuguese to Spanish) without any direct parallel training examples. Furthermore, incrementally fine-tuning these zero-shot directions with a small amount of parallel data rapidly improved translation quality to match or surpass traditional multi-step translation pipelines, while halving decoding time.

These results carry significant strategic implications for engineering efficiency, cost reduction, and model maintenance. Consolidating numerous individual models into unified multilingual systems drastically decreases infrastructure overhead and allows efficient batching across diverse language requests. The demonstration of transfer learning and internal semantic clustering indicates that the system learns a shared interlingua representation, enabling translation between unsupported language pairs without the operational latency and compounding errors of intermediate bridging.

Organizations operating large-scale translation systems should consider adopting single-model multilingual architectures to simplify serving pipelines and improve performance on low-resource languages. Where direct zero-shot translation between linguistically distant languages experiences quality degradation, teams should apply incremental training with small amounts of direct parallel data rather than training isolated models from scratch. Future work should focus on refining target language control and investigating embedding geometries to reliably predict and monitor zero-shot translation accuracy during decoding.

arXiv: 1611.04558
  • Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This paper extends the source's multilingual neural translation concepts into unsupervised cross-lingual language model pretraining across many languages simultaneously.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Building directly on the source's encoder-decoder paradigm, this work introduces the Transformer architecture that came to dominate multilingual and zero-shot machine translation.
  • Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). This work continues the source's exploration of zero-shot cross-lingual transfer by scaling multilingual text-to-text sequence modeling to over a hundred languages.
Cover for Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation

Abstract

We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no change in the model architecture from our base system but instead introduces an artificial token at the beginning of the input sentence to specify the required target language. The rest of the model, which includes encoder, decoder and attention, remains unchanged and is shared across all languages. Using a shared wordpiece vocabulary, our approach enables Multilingual NMT using a single model without any increase in parameters, which is significantly simpler than previous proposals for Multilingual NMT. Our method often improves the translation quality of all involved language pairs, even while keeping the total number of model parameters constant. On the WMT'14 benchmarks, a single multilingual model achieves comparable performance for English→\rightarrowFrench and surpasses state-of-the-art results for English→\rightarrowGerman. Similarly, a single multilingual model surpasses state-of-the-art results for French→\rightarrowEnglish and German→\rightarrowEnglish on WMT'14 and WMT'15 benchmarks respectively. On production corpora, multilingual models of up to twelve language pairs allow for better translation of many individual pairs. In addition to improving the translation quality of language pairs that the model was trained with, our models can also learn to perform implicit bridging between language pairs never seen explicitly during training, showing that transfer learning and zero-shot translation is possible for neural translation. Finally, we show analyses that hints at a universal interlingua representation in our models and show some interesting examples when mixing languages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 System Architecture for Multilingual Translation
  • 4 Experiments and Results
  • 4.1 Datasets, Training Protocols and Evaluation Metrics
  • 4.2 Many to One
  • 4.3 One to Many
  • 4.4 Many to Many
  • 4.5 Large-scale Experiments
  • 4.6 Zero-Shot Translation
  • 4.7 Effect of Direct Parallel Data
  • 5 Visual Analysis
  • 5.1 Evidence for an Interlingua
  • 5.2 Partially Separated Representations
  • 6 Mixing Languages
  • 6.1 Source Language Code-Switching
  • 6.2 Weighted Target Language Selection
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Multilingual Neural Machine Translation via Target Language Tokens and Shared Parameters

    model/method

    Multilingual Neural Machine Translation (NMT) can be achieved within a standard sequence-to-sequence encoder-decoder architecture with attention without modifying the network architecture or adding language-dependent modules. To translate an input sequence from an arbitrary source language into a specified target language, an artificial token identifying the target language (e.g., <2es> for Spanish) is prepended to the source token sequence:Source Input: [⟨2target⟩,x1,x2,…,xM]\text{Source Input: } [\langle 2\text{target}\rangle, x_1, x_2, \dots, x_M]The source language is not explicitly specified; the model infers it automatically. All components—the encoder, decoder, attention mechanism, and a shared subword vocabulary (such as a 32,000-token shared WordPiece model across all supported languages)—are completely shared among all language pairs. Training mini-batches are drawn from the combined parallel data across all language pairs, avoiding catastrophic forgetting without requiring complex multi-task update scheduling.

  2. Knowl 2 — Zero-Shot Translation via Implicit Bridging in Multilingual NMT

    empirical result

    A multilingual NMT model trained on separate language pairs that share common pivot languages (e.g., trained on Portuguese→English\text{Portuguese}\to\text{English} and English→Spanish\text{English}\to\text{Spanish}) can translate directly between language pairs never seen together during training (e.g., Portuguese→Spanish\text{Portuguese}\to\text{Spanish}) without modifying the network or using an explicit two-step bridge.

    Model Zero-shot BLEU
    (a) PBMT bridged (Pt→\toEn →\to Es) no 28.99
    (b) NMT bridged (Pt→\toEn →\to Es) no 30.91
    (c) NMT Pt→\toEs (trained directly on Pt→\toEs) no 31.50
    (d) Model 1 (trained on Pt→\toEn, En→\toEs) yes 21.62
    (e) Model 2 (trained on En↔\leftrightarrow{Es, Pt}) yes 24.75
    (f) Model 2 + incremental training no 31.77

    Model 2, trained symmetrically on 4 language pairs (En↔Es\text{En}\leftrightarrow\text{Es} and En↔Pt\text{En}\leftrightarrow\text{Pt}), outperforms Model 1 (trained on only 2 directions) by +3.13 BLEU in zero-shot Pt→Es\text{Pt}\to\text{Es}, demonstrating that adding the inverse translation directions helps the shared network build a language-independent semantic representation. Fine-tuning the zero-shot Model 2 with a small fraction of true parallel Pt→Es\text{Pt}\to\text{Es} data achieves 31.77 BLEU, outperforming explicit bridging while halving decoding time.

  3. Knowl 3 — Multilingual Translation Quality Under Many-to-One, One-to-Many, and Many-to-Many Configurations

    data/table

    Multilingual NMT models were evaluated against independently trained single-pair baseline models on WMT benchmarks (newstest2014/2015) and large-scale Google production datasets across three configuration regimes keeping total parameter count fixed (255M parameters, 8 LSTM layers, 1024 hidden units, 32k shared WordPiece vocabulary):

    Many to One One to Many Many to Many
    Pair Single Multi Diff Pair Single Multi Diff Pair Single Multi Diff
    WMT De→\toEn 30.43 30.59 +0.16 WMT En→\toDe 24.67 24.97 +0.30 WMT En→\toDe 24.67 24.49 -0.18
    WMT Fr→\toEn 35.50 35.73 +0.23 WMT En→\toFr 38.95 36.84 -2.11 WMT En→\toFr 38.95 36.23 -2.72
    WMT De→\toEn∗^* 30.43 30.54 +0.11 WMT En→\toDe∗^* 24.67 22.61 -2.06 WMT De→\toEn 30.43 29.84 -0.59
    WMT Fr→\toEn∗^* 35.50 36.77 +1.27 WMT En→\toFr∗^* 38.95 38.16 -0.79 WMT Fr→\toEn 35.50 34.89 -0.61
    Prod Ja→\toEn 23.41 23.87 +0.46 Prod En→\toJa 23.66 23.73 +0.07 Prod En→\toJa 23.66 23.12 -0.54
    Prod Ko→\toEn 25.42 25.47 +0.05 Prod En→\toKo 19.75 19.58 -0.17 Prod En→\toKo 19.75 19.73 -0.02
    Prod Es→\toEn 38.00 38.73 +0.73 Prod En→\toEs 34.50 35.40 +0.90 Prod Ja→\toEn 23.41 22.86 -0.55
    Prod Pt→\toEn 44.40 45.19 +0.79 Prod En→\toPt 38.40 38.63 +0.23 Prod Ko→\toEn 25.42 24.76 -0.66

    (∗^* indicates training without dataset oversampling).

    Key observations:

    1. Many-to-One: Multilingual models consistently outperform single models due to target-side language modeling gains and cross-lingual transfer from related source languages.
    2. One-to-Many: Multilingual performance is comparable to single models (+0.90 BLEU on En→\toEs), though decoder capacity is strained when generating multiple target languages/scripts.
    3. Many-to-Many: Average BLEU across production pairs incurs only a minor relative loss (~2.5%) compared to dedicated separate models.
  4. Knowl 4 — Large-Scale Multilingual Model Scaling Across 12 Production Language Pairs

    empirical result

    Consolidating 12 production language pairs (En↔{Ja,Ko,Es,Pt,De,Fr}\text{En}\leftrightarrow\{\text{Ja}, \text{Ko}, \text{Es}, \text{Pt}, \text{De}, \text{Fr}\}) into a single multilingual NMT model achieves performance close to dedicated single-pair models (which collectively use 3.0B parameters across 12 separate models of 255M parameters each), while requiring approximately 1/12th1/12\text{th} of the training compute:

    Model Single Multi Multi Multi Multi
    #nodes 1024 1024 1280 1536 1792
    #params 3B (total) 255M 367M 499M 650M
    En→\toJa 23.66 21.10 21.17 21.72 21.70
    En→\toKo 19.75 18.41 18.36 18.30 18.28
    Ja→\toEn 23.41 21.62 22.03 22.51 23.18
    Ko→\toEn 25.42 22.87 23.46 24.00 24.67
    En→\toEs 34.50 34.25 34.40 34.77 34.70
    En→\toPt 38.40 37.35 37.42 37.80 37.92
    Es→\toEn 38.00 36.04 36.50 37.26 37.45
    Pt→\toEn 44.40 42.53 42.82 43.64 43.87
    En→\toDe 26.43 23.15 23.77 23.63 24.01
    En→\toFr 35.37 34.00 34.19 34.91 34.81
    De→\toEn 31.77 31.17 31.65 32.24 32.32
    Fr→\toEn 36.47 34.40 34.56 35.35 35.52
    Average Diff — -1.72 -1.43 -0.95 -0.76
    Relative Diff — -5.6% -4.7% -3.1% -2.5%

    Increasing the multilingual model parameter capacity from 255M (1024 LSTM nodes per layer) to 650M (1792 nodes per layer) narrows the average performance gap from −1.72-1.72 BLEU (−5.6%-5.6\%) down to −0.76-0.76 BLEU (−2.5%-2.5\% relative to single models), with some language pairs (e.g., De→En\text{De}\to\text{En}) exceeding single-model performance.

  5. Knowl 5 — Transfer Learning and Zero-Shot Recovery via Direct Parallel Data Fine-Tuning

    empirical result

    Zero-shot translation capabilities provide a pre-trained foundation that can be fine-tuned rapidly using small amounts of direct parallel data. Evaluated on Slavic language pairs (Belarusian [Be]\text{Belarusian [Be]}, Russian [Ru]\text{Russian [Ru]}, Ukrainian [Uk]\text{Ukrainian [Uk]}, and English [En]\text{English [En]}):

    Direction Zero-Shot From-Scratch Incremental
    En→\toBe 16.85 17.03 16.99
    En→\toRu 22.21 22.03 21.92
    En→\toUk 18.16 17.75 18.27
    Be→\toEn 25.44 24.72 25.54
    Ru→\toEn 28.36 27.90 28.46
    Uk→\toEn 28.60 28.51 28.58
    Be→\toRu 56.53 82.50 78.63
    Ru→\toBe 58.75 72.06 70.01
    Ru→\toUk 21.92 25.75 25.34
    Uk→\toRu 16.73 30.53 29.92
    • Zero-Shot: Trained only on En↔{Be,Ru,Uk}\text{En}\leftrightarrow\{\text{Be}, \text{Ru}, \text{Uk}\}. It already attains strong performance on non-English pairs (e.g., 58.75 BLEU on Ru→Be\text{Ru}\to\text{Be}) without any direct parallel data.
    • From-Scratch: Trained on both En↔{Be,Ru,Uk}\text{En}\leftrightarrow\{\text{Be}, \text{Ru}, \text{Uk}\} and direct Ru↔{Be,Uk}\text{Ru}\leftrightarrow\{\text{Be}, \text{Uk}\} data.
    • Incremental: Initialized from the Zero-Shot checkpoint and fine-tuned on the direct Ru↔{Be,Uk}\text{Ru}\leftrightarrow\{\text{Be}, \text{Uk}\} data for only 3%3\% of the original training time. Incremental fine-tuning recovers nearly all the BLEU difference of the From-Scratch model while maintaining full translation quality on En↔X\text{En}\leftrightarrow X directions.
  6. Knowl 6 — Zero-Shot Translation Feasibility on Linguistically Unrelated Languages

    empirical result

    Zero-shot translation functions across linguistically distant and unrelated language families, although quality exhibits a larger degradation compared to explicit bridging than on related languages. In the 12-language-pair multilingual model:

    Model Setup BLEU
    NMT Spanish →\to Japanese explicitly bridged (Es→\toEn, then En→\toJa) 18.00
    NMT Spanish →\to Japanese implicitly bridged (Zero-Shot Es→\toJa) 9.14

    Zero-shot Es→Ja\text{Es}\to\text{Ja} achieves 9.14 BLEU compared to 18.00 BLEU for explicit bridging (a ~50% drop), confirming that implicit cross-lingual bridging remains viable even across entirely unrelated scripts and grammatical structures without language-specific components.

  7. Knowl 7 — Emergence of Interlingual Latent Representations in Multilingual NMT

    empirical result

    Multilingual NMT models learn an emergent interlingua where internal representations cluster by semantic meaning rather than source or target language identity. Analyzing the sequence of attention context vectors ci=∑jαijhj\mathbf{c}_i = \sum_j \alpha_{ij} \mathbf{h}_j (where hj\mathbf{h}_j are encoder hidden states and αij\alpha_{ij} are attention weights) across 74 semantically identical sentence triples translated across all 6 directions in an En↔{Ja,Ko}\text{En}\leftrightarrow\{\text{Ja}, \text{Ko}\} model (9,978 total decoding steps):

    1. Semantic Clustering: t-SNE projections show that context vectors for sentences with identical meaning cluster together into tightly defined strands across all source languages (English, Japanese, Korean) and target languages.
    2. Representation Separation Correlation: In zero-shot models with partially separated clusters (e.g., Pt→Es\text{Pt}\to\text{Es} zero-shot in a Pt→En/En→Es\text{Pt}\to\text{En} / \text{En}\to\text{Es} model), the average point-wise geometric distance between zero-shot context vectors and trained non-bridged context vectors correlates negatively with zero-shot BLEU score (Pearson correlation coefficient r=−0.42r = -0.42).
  8. Knowl 8 — Robustness to Source-Side Code-Switching Without Code-Switched Training

    empirical result

    Multilingual NMT models can correctly translate source sentences that mix multiple languages (code-switching) into a target language, even when no code-switched examples were present during training. For example, a multilingual {Ja,Ko}→En\{\text{Ja}, \text{Ko}\}\to\text{En} model translates pure and mixed Japanese/Korean inputs as follows:

    • Japanese source: 私は東京大学の学生です。 →\to I am a student at Tokyo University.
    • Korean source: 나는 도쿄 대학의 학생입니다. →\to I am a student at Tokyo University.
    • Mixed Japanese/Korean source: 私は東京大学학생입니다. →\to I am a student of Tokyo University.

    The model seamlessly processes mixed scripts and subwords because characters and wordpieces from both languages coexist in the shared WordPiece vocabulary and map into a common semantic space.

  9. Knowl 9 — Continuous Target Language Token Embedding Interpolation

    empirical result

    When translating from English into a target language using a multilingual En→{Ja,Ko}\text{En}\to\{\text{Ja}, \text{Ko}\} model, feeding a continuous linear interpolation of target language token embeddings: e=(1−wko)e⟨2ja⟩+wkoe⟨2ko⟩\mathbf{e} = (1 - w_{\text{ko}})\mathbf{e}_{\langle 2\text{ja}\rangle} + w_{\text{ko}}\mathbf{e}_{\langle 2\text{ko}\rangle} produces non-linear transitions rather than a blended hybrid language:

    • For wko≤0.56w_{\text{ko}} \le 0.56, the model produces pure Japanese.
    • Around wko=0.58w_{\text{ko}} = 0.58, the model outputs a mixture of Japanese and Korean words/scripts.
    • At wko=0.60w_{\text{ko}} = 0.60, the output switches to full Korean vocabulary, but retains unnatural ordering.
    • For wko≥0.70w_{\text{ko}} \ge 0.70, the output stabilizes into fully natural Korean grammar and vocabulary.

    This threshold behavior indicates that the target language model implicitly learned by the decoder LSTM acts as a strong prior favoring coherent monolingual output sequences.

Coverage note — None was omitted; all key contributed architecture mechanisms, multi-way configurations (M2O, O2M, M2M), 12-pair large-scale production models, zero-shot translation empirical findings, interlingual analysis, and mixing experiments (code-switching and embedding interpolation) are comprehensively covered.

References

  1. 1.Martin Abadi and Paul Barham et al. 2016. Tensorflow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), pages 265–283.
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations.
  3. 3.Ozan Caglayan and Walid Aransa et al. 2016. Does multimodality help human and machine for translation and image captioning? In Proceedings of the First Conference on Machine Translation, pages 627–633, Berlin, Germany, August. Association for Computational Linguistics.
  4. 4.Rich Caruana. 1998. Multitask learning. In Learning to learn, pages 95–133. Springer.
  5. 5.Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing.
  6. 6.Josep Crego and Jungi Kim et al. 2016. Systran’s pure neural machine translation systems. arXiv preprint arXiv:1610.05540.
  7. 7.Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015. Multi-task learning for multiple language translation. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, pages 1723–1732.
  8. 8.Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016a. Multi-way, multilingual neural machine translation with a shared attention mechanism. In The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 866–875.
  9. 9.Orhan Firat, Kyunghyun Cho, Baskaran Sankaran, Fatos T. Yarman Vural, and Yoshua Bengio. 2016b. Multi-way, multilingual neural machine translation. Computer Speech and Language.
  10. 10.Orhan Firat, Baskaran Sankaran, Yaser Al-Onaizan, Fatos T. Yarman Vural, and Kyunghyun Cho. 2016c. Zero-resource translation with multi-lingual neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 268–277, Austin, Texas, November. Association for Computational Linguistics.
  11. 11.Robert M French. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135.
  12. 12.Philip Gage. 1994. A new algorithm for data compression. C Users Journal, 12(2):23–38, February.
  13. 13.Dan Gillick, Cliff Brunk, Oriol Vinyals, and Amarnag Subramanya. 2016. Multilingual language processing from bytes. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1296–1306, San Diego, California, June. Association for Computational Linguistics.
  14. 14.William John Hutchins and Harold L. Somers. 1992. An introduction to machine translation, volume 362. Academic Press London.
  15. 15.Nal Kalchbrenner and Phil Blunsom. 2013. Recurrent continuous translation models. In Conference on Empirical Methods in Natural Language Processing.
  16. 16.Jason Lee, Kyunghyun Cho, and Thomas Hofmann. 2016. Fully character-level neural machine translation without explicit segmentation. arXiv preprint arXiv:1610.03017.
  17. 17.Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. 2015a. Multi-task sequence to sequence learning. In International Conference on Learning Representations.
  18. 18.Minh-Thang Luong, Hieu Pham, and Christopher D. Manning. 2015b. Effective approaches to attention-based neural machine translation. In Conference on Empirical Methods in Natural Language Processing.
  19. 19.Minh-Thang Luong, Ilya Sutskever, Quoc V Le, Oriol Vinyals, and Wojciech Zaremba. 2015c. Addressing the rare word problem in neural machine translation. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing.
  20. 20.Laurens Van Der Maaten and Geoffrey Hinton. 2008. Visualizing Data using t-SNE. Journal of Machine Learning Research, 9.
  21. 21.Richard H Richens. 1958. Interlingual machine translation. The Computer Journal, 1(3):144–147.
  22. 22.Tanja Schultz and Katrin Kirchhoff. 2006. Multilingual speech processing. Elsevier Academic Press, Amsterdam, Boston, Paris.
  23. 23.Mike Schuster and Kaisuke Nakajima. 2012. Japanese and Korean voice search. 2012 IEEE International Conference on Acoustics, Speech and Signal Processing.
  24. 24.Jean Sébastien, Cho Kyunghyun, Roland Memisevic, and Yoshua Bengio. 2015. On using very large target vocabulary for neural machine translation. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing.
  25. 25.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a. Controlling politeness in neural machine translation via side constraints. In The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 35–40.
  26. 26.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics.
  27. 27.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104–3112.
  28. 28.Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, and Chris Dyer. 2016. Polyglot neural language models: A case study in crosslingual phonetic representation learning. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1357–1366, San Diego, California, June. Association for Computational Linguistics.
  29. 29.Yonghui Wu, Mike Schuster, and Zhifeng Chen et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144v2.
  30. 30.Hayahide Yamagishi, Shin Kanouchi, and Mamoru Komachi. 2016. Controlling the voice of a sentence in Japanese-to-English neural machine translation. In Proceedings of the 3rd Workshop on Asian Translation, pages 203–210, Osaka, Japan, December.
  31. 31.Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. 2016. Deep recurrent models with fast-forward connections for neural machine translation. Transactions of the Association for Computational Linguistics, 4:371–383.
  32. 32.Barret Zoph and Kevin Knight. 2016. Multi-source neural translation. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 30–34.

Citation

MLA
Johnson, M., et al. “Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation”. arXiv, 2016, http://arxiv.org/abs/1611.04558v2.
APA
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., Hughes, M., & Dean, J. (2016). Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation. arXiv. http://arxiv.org/abs/1611.04558v2
Chicago
Johnson, M., M. Schuster, Q. V. Le, et al. 2016. “Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation”. arXiv. http://arxiv.org/abs/1611.04558v2.
Harvard
Johnson, M. et al. (2016) “Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.04558v2.
Vancouver
1. Johnson M, Schuster M, Le QV, et al (2016) Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation. arXiv

BibTeX

@article{johnson2016google,
  title = {Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation},
  author = {Johnson, Melvin and Schuster, Mike and Le, Quoc V. and Krikun, Maxim and Wu, Yonghui and Chen, Zhifeng and Thorat, Nikhil and Viégas, Fernanda and Wattenberg, Martin and Corrado, Greg and Hughes, Macduff and Dean, Jeffrey},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.04558v2},
  eprint = {1611.04558}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/