Multilingual Denoising Pre-training for Neural Machine Translation

Yinhan LiuJiatao GuNaman GoyalXian LiSergey EdunovMarjan GhazvininejadMike LewisLuke Zettlemoyer

article2020TACL2,173 citations

Introduces mBART, a multilingual sequence-to-sequence denoising autoencoder pre-trained on large-scale monolingual corpora that yields substantial performance gains across low-resource, document-level, and unsupervised machine translation.

  • Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). This work extends multilingual machine translation techniques to over 1,600 languages, building directly on the foundation of sequence-to-sequence multilingual transfer.
  • Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Following with GPT-3 demonstrates how massive autoregressive scaling achieves powerful few-shot translation capabilities beyond specialized encoder-decoder pre-training.
Cover for Multilingual Denoising Pre-training for Neural Machine Translation

Abstract

This paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART -- a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using the BART objective. mBART is one of the first methods for pre-training a complete sequence-to-sequence model by denoising full texts in multiple languages, while previous approaches have focused only on the encoder, decoder, or reconstructing parts of the text. Pre-training a complete model allows it to be directly fine tuned for supervised (both sentence-level and document-level) and unsupervised machine translation, with no task-specific modifications. We demonstrate that adding mBART initialization produces performance gains in all but the highest-resource settings, including up to 12 BLEU points for low resource MT and over 5 BLEU points for many document-level and unsupervised models. We also show it also enables new types of transfer to language pairs with no bi-text or that were not in the pre-training corpus, and present extensive analysis of which factors contribute the most to effective pre-training.

Table of Contents

  • 1 Introduction
  • 2 Multilingual Denoising Pre-training
  • 2.1 Data: CC25 corpus
  • 2.2 Model: mBART
  • 2.3 Pre-trained Models
  • 3 Sentence-level Machine Translation
  • 3.1 Experimental Settings
  • 3.2 Main Results
  • 3.3 Analysis
  • 3.4 Generalization to Languages NOT in Pre-training
  • 4 Document-level Machine Translation
  • 4.1 Experimental Settings
  • 4.2 Main Results
  • 5 Unsupervised Machine Translation
  • 5.1 Unsupervised Machine Translation via Back-Translation
  • 5.2 Unsupervised Machine Translation via Language Transfer
  • 6 Related Work
  • 7 Conclusion
  • 8 Acknowledgements
  • References
  • A Evaluation Details
  • B Translation Examples

Knowls

  1. Knowl 1 — mBART Sequence-to-Sequence Denoising Auto-Encoder Architecture and Objective

    model/method

    mBART is a multilingual sequence-to-sequence denoising auto-encoder designed for machine translation. The model utilizes a standard Transformer architecture comprising 12 encoder layers and 12 decoder layers with hidden dimension dmodel=1024d_{\text{model}} = 1024 and 16 attention heads (approximately 680 million parameters). To stabilize training under 16-bit floating-point (FP16) precision, an additional layer normalization layer is positioned on top of both the encoder stack and the decoder stack.

    Given a set of KK languages D={D1,,DK}\mathcal{D} = \{D_1, \dots, D_K\}, where each DiD_i denotes a collection of monolingual documents in language ii, the model is optimized by maximizing the reconstruction log-likelihood under a stochastic text-corruption function gg:

    Lθ=DiDXDilogP(Xg(X);θ)\mathcal{L}_\theta = \sum_{D_i \in \mathcal{D}} \sum_{X \in D_i} \log P(X \mid g(X); \theta)

    where XX is a text sequence in language ii, θ\theta represents the model parameters, and P(Xg(X);θ)P(X \mid g(X); \theta) is defined autoregressively by the Seq2Seq Transformer.

    The corruption function g(X)g(X) applies two distinct noise operations:

    1. Span Masking: 35%35\% of all tokens in an instance are masked by replacing continuous spans of tokens with a single <mask> token. Span lengths are drawn from a Poisson distribution with mean parameter λ=3.5\lambda = 3.5.
    2. Sentence Permutation: The order of sentences within the multi-sentence input instance is randomly shuffled.

    In the decoder, the original uncorrupted text is predicted with a one-position offset, initiated by a special language identifier token LID\langle\text{LID}\rangle. Sentences are delimited by end-of-sentence tokens /S\langle/\text{S}\rangle, and an instance-level LID\langle\text{LID}\rangle token is appended at the end of the input sequence. Pre-training is conducted on 256 Nvidia V100 (32GB) GPUs for 500,000 steps using the Adam optimizer (ϵ=106,β2=0.98\epsilon = 10^{-6}, \beta_2 = 0.98), linear learning rate decay, a batch size of approximately 128,000 tokens per GPU, and a dropout schedule decreasing from 0.10.1 to 0.050.05 at 250,000 steps and to 0.00.0 at 400,000 steps.

  2. Knowl 2 — Multilingual Monolingual Data Re-sampling and Instance Packing Scheme

    equation

    To train across diverse languages with vastly disparate monolingual corpus sizes, training data from the Common Crawl corpus (CC25) is balanced by computing language-specific sampling weights. The re-sampling ratio λi\lambda_i for each language i{1,,K}i \in \{1, \dots, K\} is defined as:

    λi=1pipiαj=1Kpjα\lambda_i = \frac{1}{p_i} \cdot \frac{p_i^\alpha}{\sum_{j=1}^K p_j^\alpha}

    where pip_i is the empirical fraction of language ii in the unweighted corpus, and α=0.7\alpha = 0.7 is a smoothing exponent that up-samples lower-resource languages while down-sampling dominant languages.

    Texts are tokenized using a SentencePiece model with a 250,000 subword vocabulary learned over the entire Common Crawl data across all languages. During pre-training, input instances are constructed by packing consecutive full sentences from the sampled language until encountering a document boundary or reaching the maximum sequence length of 512 tokens. Individual sentences inside the packed instance are separated by the /S\langle/\text{S}\rangle token, and the sequence terminates with the language token LID\langle\text{LID}\rangle.

  3. Knowl 3 — Supervised Sentence-Level Machine Translation on Low- and Medium-Resource Pairs

    data/table

    Fine-tuning the pre-trained mBART25 model (pre-trained on 25 languages) on parallel bi-text data consistently outperforms randomly initialized Transformer baselines across low-resource (<1M sentence pairs) and medium-resource (1M–10M sentence pairs) translation directions, with gains exceeding 12 BLEU points on low-resource and noisy pairs.

    Fine-tuning applies teacher forcing with dropout 0.30.3, label smoothing 0.20.2, 2,500 warm-up steps, maximum learning rate 3×1053 \times 10^{-5}, up to 40,000 updates for low/medium-resource pairs, and beam search decoding with beam size 5.

    Language Pair Data Source Bitext Size Direction Random BLEU mBART25 BLEU
    En-Gu WMT19 10K \leftarrow / \rightarrow 0.0 / 0.0 0.3 / 0.1
    En-Kk WMT19 91K \leftarrow / \rightarrow 0.8 / 0.2 7.4 / 2.5
    En-Vi IWSLT15 133K \leftarrow / \rightarrow 23.6 / 24.8 36.1 / 35.4
    En-Tr WMT17 207K \leftarrow / \rightarrow 12.2 / 9.5 22.5 / 17.8
    En-Ja IWSLT17 223K \leftarrow / \rightarrow 10.4 / 12.3 19.1 / 19.4
    En-Ko IWSLT17 230K \leftarrow / \rightarrow 15.3 / 16.3 24.6 / 22.6
    En-Nl IWSLT17 237K \leftarrow / \rightarrow 34.6 / 29.3 43.3 / 34.8
    En-Ar IWSLT17 250K \leftarrow / \rightarrow 27.5 / 16.9 37.6 / 21.6
    En-It IWSLT17 250K \leftarrow / \rightarrow 31.7 / 28.0 39.8 / 34.0
    En-My WAT19 259K \leftarrow / \rightarrow 23.3 / 34.9 28.3 / 36.9
    En-Ne FLoRes 564K \leftarrow / \rightarrow 7.6 / 4.3 14.5 / 7.4
    En-Ro WMT16 608K \leftarrow / \rightarrow 34.0 / 34.3 37.8 / 37.7
    En-Si FLoRes 647K \leftarrow / \rightarrow 7.2 / 1.2 13.7 / 3.3
    En-Hi ITTB 1.56M \leftarrow / \rightarrow 10.9 / 14.2 23.5 / 20.8
    En-Et WMT18 1.94M \leftarrow / \rightarrow 22.6 / 17.9 27.8 / 21.4
    En-Lt WMT19 2.11M \leftarrow / \rightarrow 18.1 / 12.1 22.4 / 15.3
    En-Fi WMT17 2.66M \leftarrow / \rightarrow 21.8 / 20.2 28.5 / 22.4
    En-Lv WMT17 4.50M \leftarrow / \rightarrow 15.6 / 12.9 19.3 / 15.9

    In these results, \leftarrow denotes XEnX \to \text{En} translation and \rightarrow denotes EnX\text{En} \to X translation. The pre-training yields large improvements on low-resource pairs (e.g., +12.5 BLEU on Vi\toEn and +10.6 BLEU on En\toVi) and noisy corpora (e.g., +12.6 BLEU on Hi\toEn). However, supervised fine-tuning fails in extreme low-resource regimes without auxiliary data (e.g., En-Gu at 10K sentence pairs scoring 0.1–0.3 BLEU).

  4. Knowl 4 — High-Resource Machine Translation Ceiling and Bitext Scaling Behavior

    empirical result

    Evaluating mBART25 on high-resource language pairs (>10M parallel sentences) demonstrates that pre-training gains diminish and eventually turn slightly negative as parallel data scales beyond 25 million sentence pairs. On En\toX translation benchmarks:

    • Cs (11M bitext): Random 16.5 BLEU vs. mBART25 18.0 BLEU (+1.5 BLEU)
    • Es (15M bitext): Random 33.2 BLEU vs. mBART25 34.0 BLEU (+0.8 BLEU)
    • Zh (25M bitext): Random 35.0 BLEU vs. mBART25 33.3 BLEU (-1.7 BLEU)
    • De (28M bitext): Random 30.9 BLEU vs. mBART25 30.5 BLEU (-0.4 BLEU)
    • Ru (29M bitext): Random 31.5 BLEU vs. mBART25 31.3 BLEU (-0.2 BLEU)
    • Fr (41M bitext): Random 41.4 BLEU vs. mBART25 41.0 BLEU (-0.4 BLEU)

    In subsampling experiments on the English-German (En-De) parallel corpus:

    • With 10K bitext pairs, mBART reaches >20>20 BLEU, whereas the non-pretrained baseline achieves 0.0 BLEU.
    • With 100K bitext pairs, mBART achieves 24.9\approx 24.9 BLEU vs. 9.29.2 BLEU for the baseline.
    • With 1M bitext pairs, mBART achieves 27.2\approx 27.2 BLEU vs. 22.422.4 BLEU for the baseline.
    • Above 10M bitext pairs, the performance curve of the non-pretrained baseline converges with and slightly surpasses the pre-trained model (30.9 vs. 30.5 BLEU at 28M pairs).

    This behavior occurs because extensive supervised gradient updates on tens of millions of sentence pairs wash out the initialization benefits provided by self-supervised pre-training.

  5. Knowl 5 — Comparison of Pre-training Methods on WMT16 English-Romanian

    data/table

    On the standard WMT16 English-Romanian (En-Ro) benchmark, full sequence-to-sequence multilingual denoising pre-training (mBART) outperforms encoder-only pre-training (XLM, XLM-R), decoder-masked pre-training (MASS), and monolingual BART, both with and without back-translation (BT) augmentation.

    Model Pre-training Data En \rightarrow Ro BLEU Ro \rightarrow En BLEU +BT BLEU
    Random baseline None 34.3 34.0 36.8
    XLM (2019) En, Ro 35.6 38.5
    MASS (2019) En, Ro 39.1
    BART (2019) En 38.0
    XLM-R (2019) CC100 35.6 35.8
    BART-En En 36.0 35.8 37.4
    BART-Ro Ro 37.6 36.8 38.1
    mBART02 En, Ro 38.5 38.5 39.9
    mBART25 CC25 37.7 37.8 38.8

    The bilingual model mBART02 achieves the highest score of 39.9 BLEU when combined with back-translation, outperforming MASS (39.1 BLEU) and monolingual BART (38.0 BLEU). Simultaneously pre-training both the encoder and decoder on complete sequence reconstruction provides superior representations compared to initializing the encoder alone or predicting only masked tokens.

  6. Knowl 6 — Language Diversity vs. Monolingual Resource Trade-offs in Multilingual Pre-training

    data/table

    The number and similarity of languages included during pre-training produce distinct trade-offs depending on the volume of target-language monolingual data available. Models pre-trained with identical numbers of English updates were evaluated on four XEnX \to \text{En} translation directions:

    Target Language XX De (66.6 GB) Ro (61.4 GB) It (30.2 GB) My (1.6 GB)
    mBART02 (Bilingual: En-XX) 31.3 38.5 39.7 36.5
    mBART06 (6 European langs) 38.5 39.3
    mBART25 (25 diverse langs) 30.5 37.7 39.8 36.9

    For low-resource monolingual languages such as Burmese/Myanmar (My, 1.6 GB of monolingual text, representing 0.5%\approx 0.5\% of English data), scaling pre-training to 25 diverse languages improves translation BLEU from 36.5 to 36.9. Conversely, when monolingual text is abundant (German with 66.6 GB, Romanian with 61.4 GB), pre-training across 25 languages causes a slight performance reduction (<1.0<1.0 BLEU) relative to the bilingual model (30.5 vs. 31.3 BLEU on De; 37.7 vs. 38.5 BLEU on Ro) due to model capacity dilution across the 25 languages. Restricting pre-training to related European languages (mBART06: Ro, It, Cs, Fr, Es, En) matches bilingual performance on Romanian (38.5 BLEU).

  7. Knowl 7 — Generalization and Representation Transfer to Pre-training-Unseen Languages

    data/table

    mBART representations generalize to languages that were entirely absent from the pre-training data, demonstrating that the learned sequence-to-sequence representations capture universal linguistic properties.

    Models pre-trained on different language subsets were fine-tuned and tested on translation pairs involving unseen languages (Dutch [Nl], Arabic [Ar], German [De]):

    Model Pre-train Data Nl-En En-Nl Ar-En En-Ar Nl-De De-Nl
    Random None 34.6 29.3 27.5 16.9 21.3 20.9
    mBART02 En, Ro 41.4 34.5 34.9 21.2 26.1 25.4
    mBART06 En, Ro, Cs, It, Fr, Es 43.1 34.6 37.3 21.1 26.4 25.3
    mBART25 All 25 langs 43.3 34.8 37.6 21.6 27.7 26.1

    Key observations:

    1. Initializing with mBART02 (trained only on English and Romanian) improves translation for unseen Arabic into English (Ar\toEn) from 27.5 to 34.9 BLEU, despite Arabic having a distinct script and negligible vocabulary overlap with English/Romanian.
    2. mBART06 (trained on 6 European languages, excluding Dutch, German, and Arabic) achieves 43.1 BLEU on Nl\toEn (within 0.2 BLEU of mBART25) and 37.3 BLEU on Ar\toEn (within 0.3 BLEU of mBART25).
    3. When both source and target languages are unseen during pre-training (e.g., Nl\leftrightarrowDe), models still achieve large improvements over random baselines (+4.8 to +5.1 BLEU), though the gap to fully pre-trained mBART25 remains wider than when at least one side was observed during pre-training.
  8. Knowl 8 — Document-Level Machine Translation with mBART Pre-training

    data/table

    Because mBART is pre-trained on multi-sentence instances up to 512 tokens, it directly fine-tunes on document-level translation tasks without architectural modifications or constrained attention mechanisms. Models were evaluated on WMT19 English-German (En-De) and TED15 Chinese-English (Zh-En) document benchmarks using sentence-level BLEU (s-BLEU) and document-level BLEU (d-BLEU):

    Task / Setting WMT19 En-De TED15 Zh-En
    Model s-BLEU d-BLEU s-BLEU d-BLEU
    Random Sent-MT 34.5 35.9 22.0
    Random Doc-MT ×\times 7.7 3.2
    HAN (Miculicich et al., 2018) 24.0
    mBART25 Sent-MT 36.4 38.0 28.4
    mBART25 Doc-MT 37.1 38.5 29.6

    In non-pretrained models, document-level training fails completely (yielding 7.7 d-BLEU on En-De and 3.2 d-BLEU on Zh-En) due to insufficient document training data. In contrast, mBART25 Doc-MT successfully captures long-range inter-sentence context, outperforming sentence-level fine-tuning (38.5 vs. 38.0 d-BLEU on En-De; 29.6 vs. 28.4 d-BLEU on Zh-En) and outperforming specialized hierarchical attention networks (HAN, 24.0 d-BLEU on Zh-En).

  9. Knowl 9 — Unsupervised Machine Translation across Distant Languages via Back-Translation

    data/table

    In fully unsupervised machine translation (where zero parallel bi-text is available), mBART parameters are used to initialize the model, which is then trained via on-the-fly back-translation (BT). To prevent trivial copying of source text in early training stages, target token output probabilities are restricted during the first 1,000 steps by masking out tokens appearing with less than 1%1\% frequency in the target monolingual corpus.

    Model En-De En-Ne En-Si
    \leftarrow \rightarrow \leftarrow \rightarrow \leftarrow \rightarrow
    Random baseline 21.0 17.2 0.0 0.0 0.0 0.0
    XLM (2019) 34.3 26.4 0.5 0.1 0.1 0.1
    MASS (2019) 35.2 28.3
    mBART 34.0 29.8 10.0 4.4 8.2 3.9

    In these results, \leftarrow indicates XEnX \to \text{En} and \rightarrow indicates EnX\text{En} \to X. On linguistically distant pairs (English-Nepali [En-Ne] and English-Sinhala [En-Si]), prior methods such as XLM degenerate to near-zero scores (0.10.10.50.5 BLEU), whereas mBART achieves non-trivial unsupervised translation (10.0 BLEU on Ne\toEn and 8.2 BLEU on Si\toEn).

  10. Knowl 10 — Unsupervised Machine Translation via Cross-Lingual Task Transfer and Iterative Back-Translation

    data/table

    When no parallel bi-text exists for a target language pair YEnY \to \text{En}, but bi-text exists for another pair XEnX \to \text{En}, fine-tuning mBART25 on XEnX \to \text{En} enables zero-shot translation from YEnY \to \text{En} without further training on YY.

    Cross-lingual transfer is strongest between languages within the same family or with related syntactic structures, but does not strictly require vocabulary overlap:

    • Fine-tuning on Hindi (Hi\toEn) yields 17.9 BLEU on Nepali\toEn and 13.8 BLEU on Gujarati\toEn (where supervised Gu\toEn fine-tuning achieved only 0.3 BLEU).
    • Fine-tuning on Italian (It\toEn) yields 34.1 BLEU on Dutch\toEn.
    • Fine-tuning on Korean (Ko\toEn) yields 9.2 BLEU on Chinese\toEn, despite zero script overlap.

    Combining cross-lingual transfer with one iteration of on-the-fly back-translation (BT) on the target language's monolingual data further enhances performance:

    Source Language Online BT BLEU Best Transfer BLEU (Source Pair) Combined BLEU
    Romanian (Ro) 30.5 23.0 (from Cs) 33.9
    Nepali (Ne) 10.0 17.9 (from Hi) 22.1
    Chinese (Zh) 11.3 9.2 (from Ko) 15.0
    Dutch (Nl) 28.5 34.1 (from It) 35.4

    For all tested pairs, combining the transfer-initialized model with iterative back-translation surpasses both standalone online back-translation and standalone zero-shot transfer.

Coverage note — None was omitted; all key architectural components, training formulations, benchmark results across supervised sentence-level, document-level, and unsupervised translation, and ablation studies are represented.

References

  1. 1.Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019. Massively multilingual neural machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3874–3884. Minneapolis, Minnesota. Association for Computational Linguistics.
  2. 2.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019. Massively multilingual neural machine translation in the wild: Findings and challenges. CoRR, abs/1907.05019.
  3. 3.Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2017. Unsupervised neural machine translation. arXiv preprint arXiv: 1710.11041. DOI: https://doi.org/18653/v1/D18-1399
  4. 4.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2019. On the cross-lingual transferability of monolingual representations. arXiv preprint arXiv:1910.11856. DOI: https://doi.org/10.18653/v1/2020.acl-main.421
  5. 5.Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012. Wit3: Web inventory of transcribed and translated talks. In Conference of European Association for Machine Translation, pages 261–268.
  6. 6.Mauro Cettolo, Niehues Jan, Stüker Sebastian, Luisa Bentivogli, Roldano Cattoni, and Marcello Federico. 2015. The IWSLT 2015 evaluation campaign. In International Workshop on Spoken Language Translation.
  7. 7.Xilun Chen and Claire Cardie. 2018. Unsupervised multilingual word embeddings. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 261–270. Brussels, Belgium. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D18-1024
  8. 8.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv: 1911.02116. DOI: https://doi.org/10.18653/v1/2020.acl-main.747
  9. 9.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D18-1269
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Association for Computational Linguistics (NAACL).
  11. 11.Chenchen Ding, Hnin Thu Zar Aye, Win Pa Pa, Khin Thandar Nwet, Khin Mar Soe, Masao Utiyama, and Eiichiro Sumita. 2019. Towards Burmese (Myanmar) morphological analysis: Syllable-based tokenization and part-of-speech tagging. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 19(1):5. DOI: https://doi.org/10.1145/3325885
  12. 12.Chenchen Ding, Masao Utiyama, and Eiichiro Sumita. 2018. NOVA: A feasible and flexible annotation system for joint tokenization and part-of-speech tagging. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 18(2):17. DOI: https://doi.org/10.1145/3276773
  13. 13.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:1905.03197.
  14. 14.Sergey Edunov, Alexei Baevski, and Michael Auli. 2019. Pre-trained language model representations for language generation. arXiv preprint arXiv:1903.09722. DOI: https://doi.org/10.18653/v1/N19-1409
  15. 15.Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016. Multi-way, multilingual neural machine translation with a shared attention mechanism. In NAACL. DOI: https://doi.org/10.18653/v1/N16-1101
  16. 16.Jiatao Gu, Hany Hassan, Jacob Devlin, and Victor O. K. Li. 2018. Universal neural machine translation for extremely low resource languages. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 344–354. New Orleans, Louisiana. Association for Computational Linguistics.
  17. 17.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O. K. Li. 2019. Improved zero-shot neural machine translation via ignoring spurious correlations. arXiv preprint arXiv:1906.01181.
  18. 18.Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc'Aurelio Ranzato. 2019. The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6097–6110. Hong Kong, China. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D19-1632
  19. 19.Sébastien Jean, Stanislas Lauly, Orhan Firat, and Kyunghyun Cho. 2017. Does neural machine translation benefit from larger context? CoRR, abs/1704.05135.
  20. 20.Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017. Googles multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5:339–351. DOI: https://doi.org/10.1162/tacl_a_00065
  21. 21.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71. Brussels, Belgium. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D18-2012, PMID: 29382465
  22. 22.Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2017. The IIT Bombay English-Hindi parallel corpus. CoRR, abs/1710.02855.
  23. 23.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291.
  24. 24.Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018a. Unsupervised machine translation using monolingual corpora only. In International Conference on Learning Representations.
  25. 25.Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018b. Word translation without parallel data. In International Conference on Learning Representations.
  26. 26.Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018c. Phrase-based & neural unsupervised machine translation. arXiv preprint arXiv: 1804.07755. DOI: https://doi.org/10.18653/v1/D18-1549
  27. 27.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461. DOI: https://doi.org/10.18653/v1/2020.acl-main.703
  28. 28.Liangyou Li, Xin Jiang, and Qun Liu. 2019. Pretrained language models for document-level neural machine translation. arXiv preprint arXiv:1911.03110.
  29. 29.Yang Liu and Mirella Lapata. 2019. Text summarization with pretrained encoders. arXiv preprint arXiv:1908.08345. DOI: https://doi.org/10.18653/v1/D19-1387
  30. 30.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. ROBERTA: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  31. 31.Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018. Document-level neural machine translation with hierarchical attention networks. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2947–2954. Brussels, Belgium. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D18-1325
  32. 32.Tomas Mikolov, Quoc V. Le, and Ilya Sutskever. 2013. Exploiting similarities among languages for machine translation. CoRR, abs/1309.4168.
  33. 33.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. FAIRSEQ: A fast, extensible toolkit for sequence modeling. In North American Association for Computational Linguistics (NAACL): System Demonstrations. DOI: https://doi.org/10.18653/v1/N19-4009
  34. 34.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In North American Association for Computational Linguistics (NAACL).
  35. 35.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? arXiv preprint arXiv:1906.01502. DOI: https://doi.org/10.18653/v1/P19-1493
  36. 36.Nima Pourdamghani, Nada Aldarrab, Marjan Ghazvininejad, Kevin Knight, and Jonathan May. 2019. Translating translationese: A two-step approach to unsupervised machine translation. In ACL. DOI: https://doi.org/10.18653/v1/P19-1293
  37. 37.Ye Qi, Devendra Singh Sachan, Matthieu Felix, Sarguna Janani Padmanabhan, and Graham Neubig. 2018. When and why are pre-trained word embeddings useful for neural machine translation? arXiv preprint arXiv: 1804.06323. DOI: https://doi.org/10.18653/v1/N18-2084
  38. 38.Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever. 2018. Improving language understanding with unsupervised learning, OpenAI.
  39. 39.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners, OpenAI.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  41. 41.Prajit Ramachandran, Peter J Liu, and Quoc Le. 2017. Unsupervised pretraining for sequence to sequence learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 383–391. DOI: https://doi.org/10.18653/v1/D17-1039
  42. 42.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 86–96. Berlin, Germany. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/P16-1009
  43. 43.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  44. 44.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. MASS: Masked sequence to sequence pre-training for language generation. In International Conference on Machine Learning (ICML).
  45. 45.Jörg Tiedemann and Yves Scherrer. 2017. Neural machine translation with extended context. In Proceedings of the Third Workshop on Discourse in Machine Translation, pages 82–92. Copenhagen, Denmark. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/W17-4811
  46. 46.Zhaopeng Tu, Yang Liu, Shuming Shi, and Tong Zhang. 2018. Learning to remember translation history with a continuous cache. Transactions of the Association for Computational Linguistics, 6:407–420. DOI: https://doi.org/10.1162/tacl_a_00029
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems.
  48. 48.Takashi Wada and Tomoharu Iwata. 2018. Unsupervised cross-lingual word embedding by multilingual neural language models. CoRR, abs/1809.02306. DOI: https://doi.org/10.18653/v1/P19-1300
  49. 49.Longyue Wang, Zhaopeng Tu, Andy Way, and Qun Liu. 2017. Exploiting cross-sentence context for neural machine translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2826–2831. Copenhagen, Denmark. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D17-1301
  50. 50.Zihan Wang, Stephen Mayhew, Dan Roth, and others. 2019. Cross-lingual ability of multilingual bert: An empirical study. arXiv preprint arXiv:1912.07840.
  51. 51.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2019. CCNET: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359.
  52. 52.Lijun Wu, Jinhua Zhu, Di He, Fei Gao, Xu Tan, Tao Qin, and Tie-Yan Liu. 2019. Machine translation with weakly paired bilingual documents.
  53. 53.Jiacheng Yang, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Yong Yu, Weinan Zhang, and Lei Li. 2019a. Towards making the most of bert in neural machine translation. arXiv preprint arXiv:1908.05672.
  54. 54.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019b. XLNet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237.
  55. 55.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019. DialoGPT: Large-scale generative pre-training for conversational response generation.
  56. 56.Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020. Incorporating BERT into neural machine translation. arXiv preprint arXiv:2002.06823. DOI: https://doi.org/10.18653/v1/2020.acl-demos.30

Citation

MLA
Liu, Y., et al. “Multilingual Denoising Pre-training for Neural Machine Translation”. arXiv, 2020, http://arxiv.org/abs/2001.08210v2.
APA
Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., & Zettlemoyer, L. (2020). Multilingual Denoising Pre-training for Neural Machine Translation. arXiv. http://arxiv.org/abs/2001.08210v2
Chicago
Liu, Y., J. Gu, N. Goyal, et al. 2020. “Multilingual Denoising Pre-training for Neural Machine Translation”. arXiv. http://arxiv.org/abs/2001.08210v2.
Harvard
Liu, Y. et al. (2020) “Multilingual Denoising Pre-training for Neural Machine Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2001.08210v2.
Vancouver
1. Liu Y, Gu J, Goyal N, Li X, Edunov S, Ghazvininejad M, Lewis M, Zettlemoyer L (2020) Multilingual Denoising Pre-training for Neural Machine Translation. arXiv

BibTeX

@article{liu2020multilingual,
  title = {Multilingual Denoising Pre-training for Neural Machine Translation},
  author = {Liu, Yinhan and Gu, Jiatao and Goyal, Naman and Li, Xian and Edunov, Sergey and Ghazvininejad, Marjan and Lewis, Mike and Zettlemoyer, Luke},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2001.08210v2},
  eprint = {2001.08210}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/