keyword
statistical machine translation
Statistical machine translation is an approach to automated translation that generates target-language text from source-language text using statistical models learned from bilingual and monolingual corpora. Instead of relying on predefined grammatical rules and dictionaries, it analyzes parallel texts to calculate the probability that specific words, phrases, or structural patterns in one language correspond to those in another. A standard statistical machine translation system typically integrates a translation model that predicts phrase equivalents, a target-language model that evaluates the fluency and naturalness of the generated text, and a decoding algorithm that searches for the translation candidate with the highest combined statistical score. This data-driven paradigm encompasses word-based, phrase-based, and syntax-based architectures, using statistical optimization techniques to maximize translation quality across diverse language pairs.
10 items

An Open Dataset and Model for Language Identification
Laurie Burchell, Alexandra Birch, Nikolay Bogoychev, Kenneth Heafield
Why you should read this
Presents an open, manually verified dataset of 121 million lines alongside a fastText language identification model that outperforms existing systems across 201 languages.
Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a fasttext LID model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033 across 201 languages, outperforming previous work. We achieve this by training on a curated dataset of monolingual data, which we audit manually to ensure reliability. We make both the model and the dataset available to the research community. Finally, we carry out detailed analysis into our model’s performance, both in comparison to existing open models and by language class. For applications such as corpus filtering, LID systems need to be fast, reliable, and cover as many languages as possible. There are several open LID models offering quick classification and high language coverage, such as CLD3 or the work of Costa-jussà et al. (2022). However, to the best of our knowledge, none of the commonly-used scalable LID systems make their training data public. We address this gap by releasing an open and curated dataset for LID, which we audit by hand to assure quality. We also train a high-performing LID model on this dataset to show its utility.
Added
2026-10-03

A Hierarchical Phrase-Based Model for Statistical Machine Translation
David Chiang
Why you should read this
Introduces a synchronous context-free grammar translation model that learns hierarchical phrase structures directly from parallel text without syntactic annotations, improving translation quality and long-distance reordering over standard phrase-based systems.
We present a statistical phrase-based translation model that uses hierarchical phrases—phrases that contain subphrases. The model is formally a synchronous context-free grammar but is learned from a bitext without any syntactic information. Thus it can be seen as a shift to the formal machinery of syntax-based translation systems without any linguistic commitment. In our experiments using BLEU as a metric, the hierarchical phrase-based model achieves a relative improvement of 7.5% over Pharaoh, a state-of-the-art phrase-based system.
Added
2026-09-25

Six Challenges for Neural Machine Translation
Philipp Koehn, Rebecca Knowles
Why you should read this
Exposes critical weaknesses in neural machine translation by benchmarking its failures against phrase-based systems across domain shifts, data scarcity, rare vocabulary, and decoding search errors.
We explore six challenges for neural machine translation: domain mismatch, amount of training data, rare words, long sentences, word alignment, and beam search. We show both deficiencies and improvements over the quality of phrase-based statistical machine translation.
Added
2026-09-25

Improving Neural Machine Translation Models with Monolingual Data
Rico Sennrich, Barry Haddow, Alexandra Birch
Why you should read this
Introduces back-translation to train neural machine translation models on target-side monolingual data without architecture changes, substantially increasing translation accuracy across standard and low-resource benchmarks.
Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training. Target-side monolingual data plays an important role in boosting fluency for phrase-based statistical machine translation, and we investigate the use of monolingual data for NMT. In contrast to previous work, which combines NMT models with separately trained language models, we note that encoder-decoder NMT architectures already have the capacity to learn the same information as a language model, and we explore strategies to train with monolingual data without changing the neural network architecture. By pairing monolingual training data with an automatic back-translation, we can treat it as additional parallel training data, and we obtain substantial improvements on the WMT 15 task English<->German (+2.8-3.7 BLEU), and for the low-resourced IWSLT 14 task Turkish->English (+2.1-3.4 BLEU), obtaining new state-of-the-art results. We also show that fine-tuning on in-domain monolingual and parallel data gives substantial improvements for the IWSLT 15 task English->German.
Added
2026-09-13

A Call for Clarity in Reporting BLEU Scores
Matt Post
Why you should read this
Demonstrates that unstandardized reference tokenization distorts BLEU scores across machine translation papers by up to 1.8 points and introduces SacreBLEU to ensure consistent, reproducible evaluation.
The field of machine translation faces an under-recognized problem because of inconsistency in the reporting of scores from its dominant metric. Although people refer to "the" BLEU score, BLEU is in fact a parameterized metric whose values can vary wildly with changes to these parameters. These parameters are often not reported or are hard to find, and consequently, BLEU scores between papers cannot be directly compared. I quantify this variation, finding differences as high as 1.8 between commonly used configurations. The main culprit is different tokenization and normalization schemes applied to the reference. Pointing to the success of the parsing community, I suggest machine translation researchers settle upon the BLEU scheme used by the annual Conference on Machine Translation (WMT), which does not allow for user-supplied reference processing, and provide a new tool, SacreBLEU, to facilitate this.
Added
2026-09-11

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, Jeffrey Dean
Why you should read this
Presents the architecture and engineering design behind Google's production neural machine translation system, demonstrating how deep recurrent networks, subword tokenization, and low-precision inference overcome practical latency and vocabulary challenges to reduce translation errors by 60% over phrase-based baselines.
Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference. Also, most NMT systems have difficulty with rare words. These issues have hindered NMT's use in practical deployments and services, where both accuracy and speed are essential. In this work, we present GNMT, Google's Neural Machine Translation system, which attempts to address many of these issues. Our model consists of a deep LSTM network with 8 encoder and 8 decoder layers using attention and residual connections. To improve parallelism and therefore decrease training time, our attention mechanism connects the bottom layer of the decoder to the top layer of the encoder. To accelerate the final translation speed, we employ low-precision arithmetic during inference computations. To improve handling of rare words, we divide words into a limited set of common sub-word units ("wordpieces") for both input and output. This method provides a good balance between the flexibility of "character"-delimited models and the efficiency of "word"-delimited models, naturally handles translation of rare words, and ultimately improves the overall accuracy of the system. Our beam search technique employs a length-normalization procedure and uses a coverage penalty, which encourages generation of an output sentence that is most likely to cover all the words in the source sentence. On the WMT'14 English-to-French and English-to-German benchmarks, GNMT achieves competitive results to state-of-the-art. Using a human side-by-side evaluation on a set of isolated simple sentences, it reduces translation errors by an average of 60% compared to Google's phrase-based production system.
Added
2026-09-10
License
Published with permission

Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, Evan Herbst
Why you should read this
Presents a complete open-source statistical machine translation framework featuring factored translation models, confusion network decoding, and memory-efficient data structures for training, tuning, and decoding.
We describe an open-source toolkit for statistical machine translation whose novel contributions are (a) support for linguistically motivated factors, (b) confusion network decoding, and (c) efficient data formats for translation models and language models. In addition to the SMT decoder, the toolkit also includes a wide variety of tools for training, tuning and applying the system to many translation tasks.
Added
2026-09-09

Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio
Why you should read this
Introduces the RNN Encoder-Decoder architecture and the Gated Recurrent Unit, establishing a foundational sequence-to-sequence framework that enables neural networks to encode and translate variable-length linguistic phrases.
In this paper, we propose a novel neural network model called RNN Encoder-Decoder that consists of two recurrent neural networks (RNN). One RNN encodes a sequence of symbols into a fixed-length vector representation, and the other decodes the representation into another sequence of symbols. The encoder and decoder of the proposed model are jointly trained to maximize the conditional probability of a target sequence given a source sequence. The performance of a statistical machine translation system is empirically found to improve by using the conditional probabilities of phrase pairs computed by the RNN Encoder-Decoder as an additional feature in the existing log-linear model. Qualitatively, we show that the proposed model learns a semantically and syntactically meaningful representation of linguistic phrases.
Added
2026-09-05

Minimum Error Rate Training in Statistical Machine Translation
Franz Josef Och
Why you should read this
Introduces a seminal training framework for statistical machine translation that directly optimizes model parameters against non-differentiable evaluation metrics like BLEU through an efficient line-search algorithm, significantly improving translation quality by aligning training criteria with final performance measures.
Often, the training procedure for statistical machine translation models is based on maximum likelihood or related criteria. A general problem of this approach is that there is only a loose relation to the final translation quality on unseen text. In this paper, we analyze various training criteria which directly optimize translation quality. These training criteria make use of recently proposed automatic evaluation metrics. We describe a new algorithm for efficient training an unsmoothed error count. We show that significantly better results can often be obtained if the final evaluation criterion is taken directly into account as part of the training procedure.
Added
2026-02-21

Statistical Phrase-Based Translation
Philipp Koehn, Franz Josef Och, Daniel Marcu
Why you should read this
Demonstrates that simple heuristic extraction of phrase pairs from word alignments consistently outperforms more complex syntactic and joint-probability models, establishing a foundational and highly effective framework for statistical machine translation.
We propose a new phrase-based translation model and decoding algorithm that enables us to evaluate and compare several, previously proposed phrase-based translation models. Within our framework, we carry out a large number of experiments to understand better and explain why phrase-based models outperform word-based models. Our empirical results, which hold for all examined language pairs, suggest that the highest levels of performance can be obtained through relatively simple means: heuristic learning of phrase translations from word-based alignments and lexical weighting of phrase translations. Surprisingly, learning phrases longer than three words and learning phrases from high-accuracy word-level alignment models does not have a strong impact on performance. Learning only syntactically motivated phrases degrades the performance of our systems.
Added
2026-02-21
