Built independently by an author, for readers. Read the story and support ChapterPal

keyword

synthetic parallel text

Synthetic parallel text is a paired bilingual or multilingual dataset where at least one language side has been automatically generated rather than translated entirely by humans. In natural language processing and machine translation, it is used to augment limited human-translated training data by leveraging abundant monolingual text. Such datasets are typically produced using techniques like back-translation or forward-translation, in which an automated translation system translates text from one language to another to create aligned sentence pairs. This approach enables machine translation models to expand their training coverage, improve fluency, and learn domain vocabulary without relying exclusively on human-curated parallel corpora.

1 item

Improving Neural Machine Translation Models with Monolingual Data

Improving Neural Machine Translation Models with Monolingual Data

Rico Sennrich, Barry Haddow, Alexandra Birch

OrganizationsUniversity of Edinburgh

Why you should read this

Introduces back-translation to train neural machine translation models on target-side monolingual data without architecture changes, substantially increasing translation accuracy across standard and low-resource benchmarks.

Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training. Target-side monolingual data plays an important role in boosting fluency for phrase-based statistical machine translation, and we investigate the use of monolingual data for NMT. In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture. By pairing monolingual training data with an automatic back-translation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English<->German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish->English (+2.1-3.4 BLEU), obtaining new state-of-the-art results. We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English->German.

Added

2026-09-13