Built independently by an author, for readers. Read the story and support ChapterPal

keyword

monolingual data

Monolingual data refers to a collection of text or speech presented entirely in a single language, without paired or aligned translations in other languages. In natural language processing and machine translation, it stands in contrast to parallel data, which consists of texts explicitly paired across two or more languages. Because single-language text is far more abundant and easier to acquire from large-scale text corpora than bilingual translations, monolingual data is widely used to pre-train language models, improve vocabulary coverage, and enhance output fluency. It is also frequently utilized in machine translation workflows through semi-supervised techniques such as back-translation, where it is transformed into synthetic parallel data to train and fine-tune models, particularly in domain-specific or low-resource settings.

4 items

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

David Ifeoluwa Adelani, Jesujoba O. Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Wambui Gitau, Jade Z. Abbott, Mohamed Ahmed, Millicent Ochieng, Aremu Anuoluwapo, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing K. Sibanda, Andiswa Bukula, Sam Manthalu

OrganizationsAhmadu Bello UniversityAi4InnovAmazonBaamtuCarnegie Mellon UniversityGoogleINESC TECINRIAJacobs University BremenJomo Kenyatta University of Agriculture and TechnologyMakerere UniversityMasakhaneMetaMicrosoftNUSTOminor AIRochester Institute of TechnologySaarland UniversitySADiLaRSPUTechnical University of MunichUniversitat Politècnica de CatalunyaUniversity of Chinese Academy of SciencesUniversity of DaytonUniversity of IbadanUniversity of MalawiUniversity of the WitwatersrandUniversity of TokyoUppsala University

Why you should read this

Demonstrates that fine-tuning large multilingual models on just a few thousand high-quality in-domain sentence pairs effectively transfers translation capabilities to sixteen low-resource African languages absent from the original pre-training data.

Recent advances in the pre-training of language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages are not well represented on the web and therefore excluded from the large-scale crawls used to create datasets. Furthermore, downstream users of these models are restricted to the selection of languages originally chosen for pre-training. This work investigates how to optimally leverage existing pre-trained models to create low-resource translation systems for 16 African languages. We focus on two questions: 1) How can pre-trained models be used for languages not included in the initial pre-training? and 2) How can the resulting translation models effectively transfer to new domains? To answer these questions, we create a new African news corpus covering 16 languages, of which eight languages are not part of any existing evaluation dataset. We demonstrate that the most effective strategy for transferring both to additional languages and to additional domains is to fine-tune large pre-trained models on small quantities of high-quality translation data.

Added

2026-10-01

Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, Young Jin Kim

OrganizationsJohns Hopkins UniversityMicrosoft

Why you should read this

Proposes Contrastive Preference Optimization, a training method that prevents moderate-sized language models from mimicking imperfect reference translations, enabling a 13B model trained on only 22K sentences to match or exceed the translation performance of GPT-4 and WMT competition winners.

Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, even the top-performing 13B LLM-based translation models, like ALMA, does not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to SFT which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 12M parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.

Added

2026-09-28

Improving Neural Machine Translation Models with Monolingual Data

Improving Neural Machine Translation Models with Monolingual Data

Rico Sennrich, Barry Haddow, Alexandra Birch

OrganizationsUniversity of Edinburgh

Why you should read this

Introduces back-translation to train neural machine translation models on target-side monolingual data without architecture changes, substantially increasing translation accuracy across standard and low-resource benchmarks.

Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training. Target-side monolingual data plays an important role in boosting fluency for phrase-based statistical machine translation, and we investigate the use of monolingual data for NMT. In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture. By pairing monolingual training data with an automatic back-translation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English<->German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish->English (+2.1-3.4 BLEU), obtaining new state-of-the-art results. We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English->German.

Added

2026-09-13