Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
Taku Kudo
Proposes a unigram language model segmentation algorithm and a subword regularization technique that probabilistically samples multiple segmentations during training, significantly improving neural machine translation on low-resource and out-of-domain tasks.
Standard neural machine translation systems convert text into subword units to handle large and open vocabularies efficiently. However, standard methods like Byte-Pair Encoding rely on deterministic segmentations that assign a single subword sequence to each sentence, ignoring alternative valid segmentations. This lack of flexibility makes translation models fragile when exposed to real-world noise, rare words, or unfamiliar domains.
The article introduces and evaluates "subword regularization," a training method designed to enhance translation accuracy and robustness by exposing models to multiple subword candidates during training. It also proposes a probabilistic unigram language model to generate and sample these subword variations according to their likelihood.
To evaluate this approach, the authors conducted translation experiments across multiple datasets of varying sizes (from small corpora of roughly 133,000 sentences to large corpora exceeding 15 million sentences) covering diverse languages such as English, Vietnamese, Chinese, French, Arabic, Japanese, German, and Czech. The method samples alternative subword sequences on-the-fly during training without altering the underlying neural network architecture. Translation performance was measured using standard translation quality benchmark scores (BLEU) across both standard test sets and out-of-domain datasets (such as Web text, patents, and query logs).
The evaluation yielded several key findings. First, subword regularization delivered consistent translation quality improvements of 1 to 2 BLEU points over standard baseline methods across nearly all language pairs. Second, the performance gains were most pronounced in low-resource language settings (such as small benchmark datasets) and in out-of-domain evaluations, where quality improvements reached up to roughly 2 to 10 BLEU points even for models trained on massive datasets. Third, combining subword regularization with multi-candidate search during translation (n-best decoding) provided additional quality gains. Finally, applying regularization to both the source and target sentences yielded the largest benefits, though applying it to either side individually still showed positive effects.
These findings indicate that introducing probabilistic segmentation noise acts as an effective data augmentation strategy, teaching models how words are composed and making them significantly more resilient to unfamiliar inputs. Because subword regularization operates purely through data sampling, it requires no structural redesign of machine translation models, minimizing technical complexity and implementation risk while delivering meaningful quality improvements.
Organizations developing or deploying translation systems should adopt probabilistic subword segmentation to improve robustness, particularly when working with limited training data or user-generated, open-domain content. When implementing the approach, practitioners should tune sampling hyperparameters using validation datasets, choosing moderate candidate sizes (such as 64 candidates) for high-resource settings to prevent over-regularization. Future efforts should also explore applying this technique to other sequence-generation applications, such as dialogue systems and text summarization.
Confidence in these findings is high across diverse languages and domains, though hyperparameter sensitivity represents an operational limitation: choosing completely uniform sampling without probabilistic weighting degrades performance, meaning sampling parameters must be calibrated carefully.
- Paper: Neural Machine Translation of Rare Words with Subword Units, Rico Sennrich et al. (2016). Introduces the foundational use of subword units via byte pair encoding (BPE) to solve the open-vocabulary problem in neural machine translation, which this paper directly improves upon.
- Paper: Effective Approaches to Attention-based Neural Machine Translation, Minh-Thang Luong et al. (2015). Establishes standard attention-based encoder-decoder architectures for neural machine translation, serving as the base translation framework regularized in this work.
- Paper: Neural Machine Translation by Jointly Learning to Align and Translate, Dzmitry Bahdanau et al. (2015). Introduces the attention mechanism for sequence-to-sequence neural translation that underpins the baseline models evaluated in this research.
- Paper: Enriching Word Vectors with Subword Information, Piotr Bojanowski et al. (2017). Demonstrates the power of character and subword n-gram representations for capturing morphology, motivating the transition toward richer subword vocabularies.
- Paper: Character-Aware Neural Language Models, Yoon Kim et al. (2015). Explores character-level compositionality in neural language modeling to overcome fixed word-level vocabulary constraints.
- Paper: Recurrent Neural Network Regularization, Wojciech Zaremba et al. (2014). Presents foundational dropout techniques for regularizing sequence models, establishing the broader regularization paradigm that subword regularization complements.
- Paper: SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, Taku Kudo et al. (2018). Directly implements and generalizes the unigram language model subword segmentation and regularization algorithms introduced in this work into the widely used SentencePiece toolkit.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Applies shared subword vocabularies and unsupervised tokenization schemes to cross-lingual pre-training for downstream machine translation and language understanding.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Extends subword-tokenized sequence modeling to large-scale multilingual denoising pre-training across 25 languages for low-resource translation.
- Paper: Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models, Taido Purason et al. (2025). Explores subsequent methods for dynamically adapting and pruning pre-trained subword tokenizers when transferring models to new languages or domains.
