Data Scaling Laws in NMT: The Effect of Noise and Architecture
Yamini BansalBehrooz GhorbaniAnkush GargBiao ZhangColin CherryBehnam NeyshaburOrhan Firat
Establishes empirical data scaling laws for neural machine translation across varying architectures and noise levels, demonstrating that data scaling exponents remain largely stable under architectural changes and noise while back-translation significantly degrades scaling efficiency.
Training neural machine translation systems requires massive computational and data budgets, often involving billions of sentences. In this resource-intensive environment, engineering teams frequently debate whether to invest in architectural refinements, complex data filtering pipelines, or larger datasets. The article evaluates how changes in model architecture, noise filtering, synthetic noise, and machine-generated data affect the fundamental rate at which translation performance improves as dataset sizes expand.
To conduct this evaluation, the researchers trained multiple neural machine translation models across three orders of magnitude of data, spanning 500,000 to 512 million sentence pairs (up to 28 billion tokens). They evaluated standard encoder-decoder transformers against alternative architectures, including transformer-LSTM hybrids and decoder-only language models. They also tested real-world web-crawled noise, two common filtering algorithms (Bicleaner and Contrastive Data Selection), artificial synthetic noise on inputs and outputs, and back-translated data generated across several model sizes. These setups were primarily assessed on English-to-German translation and validated on Chinese-to-English translation.
Key findings show that the scaling rate (the power law exponent) remains largely unchanged across different architectures, filtering methods, and independent synthetic noise, holding steady around an exponent of 0.28. Second, suboptimal architectures and noisy or unfiltered data merely introduce a constant performance offset; their penalties can be offset by simply supplying a constant factor of additional training data. Third, synthetic noise placed on the target output is significantly more harmful to translation performance than noise on the source input. Fourth, the use of back-translated data fundamentally degrades the learning curve, dropping the scaling exponent from roughly 0.28 down to 0.198, which leaves a persistent gap between synthetic data and human parallel data at web scale.
These findings mean that minor modifications to model designs and heavy investments in filtering pipelines do not alter long-term learning efficiency. Leaders can optimize model selection around practical operational constraints—such as memory footprint, deployment speed, or multi-task flexibility—and compensate for minor performance gaps by expanding the dataset. However, because back-translated data exhibits diminishing returns at large scales, teams cannot fully replace human parallel corpora with synthetic data.
Decision-makers should avoid over-engineering architectures or deploying aggressive filtering rules that risk removing useful diversity, relying instead on increasing volume when cost-effective. Engineering efforts for synthetic data should focus on maximizing the quality and capacity of the back-translation models to mitigate degradation. Future work should further analyze whether these scaling dynamics hold across low-resource languages and structurally distinct model families.
While the findings offer high confidence across large English-German and Chinese-English corpora, the conclusions are based on empirical observations and may vary when handling fundamentally distinct classes of noise or non-transformer architectures. Caution is advised before extrapolating these exact scaling constants directly to low-resource settings without initial validation.
- Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). Kaplan et al.’s language-model scaling laws establish the power-law framework that this study tests against translation data, architecture, and noise.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). The Transformer paper introduces the standard encoder–decoder architecture that this study uses as its baseline for comparing alternative model designs.
- Paper: Sequence to Sequence Learning with Neural Networks, Ilya Sutskever et al. (2014). This foundational sequence-to-sequence translation work provides the NMT setting needed to understand the study’s experiments on translation models and parallel data.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). It extends scaling-law analysis to data-constrained training, testing how repetition and data supply affect performance when fresh data cannot grow freely.
