Data Scaling Laws in NMT: The Effect of Noise and Architecture

Yamini BansalBehrooz GhorbaniAnkush GargBiao ZhangColin CherryBehnam NeyshaburOrhan Firat

article2022ICML67 citations

Establishes empirical data scaling laws for neural machine translation across varying architectures and noise levels, demonstrating that data scaling exponents remain largely stable under architectural changes and noise while back-translation significantly degrades scaling efficiency.

Listen

Training neural machine translation systems requires massive computational and data budgets, often involving billions of sentences. In this resource-intensive environment, engineering teams frequently debate whether to invest in architectural refinements, complex data filtering pipelines, or larger datasets. The article evaluates how changes in model architecture, noise filtering, synthetic noise, and machine-generated data affect the fundamental rate at which translation performance improves as dataset sizes expand.

To conduct this evaluation, the researchers trained multiple neural machine translation models across three orders of magnitude of data, spanning 500,000 to 512 million sentence pairs (up to 28 billion tokens). They evaluated standard encoder-decoder transformers against alternative architectures, including transformer-LSTM hybrids and decoder-only language models. They also tested real-world web-crawled noise, two common filtering algorithms (Bicleaner and Contrastive Data Selection), artificial synthetic noise on inputs and outputs, and back-translated data generated across several model sizes. These setups were primarily assessed on English-to-German translation and validated on Chinese-to-English translation.

Key findings show that the scaling rate (the power law exponent) remains largely unchanged across different architectures, filtering methods, and independent synthetic noise, holding steady around an exponent of 0.28. Second, suboptimal architectures and noisy or unfiltered data merely introduce a constant performance offset; their penalties can be offset by simply supplying a constant factor of additional training data. Third, synthetic noise placed on the target output is significantly more harmful to translation performance than noise on the source input. Fourth, the use of back-translated data fundamentally degrades the learning curve, dropping the scaling exponent from roughly 0.28 down to 0.198, which leaves a persistent gap between synthetic data and human parallel data at web scale.

These findings mean that minor modifications to model designs and heavy investments in filtering pipelines do not alter long-term learning efficiency. Leaders can optimize model selection around practical operational constraints—such as memory footprint, deployment speed, or multi-task flexibility—and compensate for minor performance gaps by expanding the dataset. However, because back-translated data exhibits diminishing returns at large scales, teams cannot fully replace human parallel corpora with synthetic data.

Decision-makers should avoid over-engineering architectures or deploying aggressive filtering rules that risk removing useful diversity, relying instead on increasing volume when cost-effective. Engineering efforts for synthetic data should focus on maximizing the quality and capacity of the back-translation models to mitigate degradation. Future work should further analyze whether these scaling dynamics hold across low-resource languages and structurally distinct model families.

While the findings offer high confidence across large English-German and Chinese-English corpora, the conclusions are based on empirical observations and may vary when handling fundamentally distinct classes of noise or non-transformer architectures. Caution is advised before extrapolating these exact scaling constants directly to low-resource settings without initial validation.

Bansal et al (2022).pdf
  • Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). Kaplan et al.’s language-model scaling laws establish the power-law framework that this study tests against translation data, architecture, and noise.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). The Transformer paper introduces the standard encoder–decoder architecture that this study uses as its baseline for comparing alternative model designs.
  • Paper: Sequence to Sequence Learning with Neural Networks, Ilya Sutskever et al. (2014). This foundational sequence-to-sequence translation work provides the NMT setting needed to understand the study’s experiments on translation models and parallel data.
  • Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). It extends scaling-law analysis to data-constrained training, testing how repetition and data supply affect performance when fresh data cannot grow freely.
Cover for Data Scaling Laws in NMT: The Effect of Noise and Architecture

Abstract

Neural Machine Translation (NMT) has seen tremendous growth in usage and model size over the last few years. It is still unclear, however, how model and data size interact to affect performance. In this work we study the scaling behavior of NMT models. We train a standard transformer architecture with different data sizes and model sizes. We find that the test loss scales as a power-law in both data and model size. We also find that the scaling exponents are affected by the amount of noise in the data and by the architecture.

Table of Contents

  • 1. Introduction
  • 1.1. Experimental Setup
  • 2. Related Works
  • 3. Data Scaling Laws
  • 3.1. BLEU Score
  • 3.2. Out-of-Distribution Generalization
  • 4. The Effect of Architecture
  • 5. The Effect of Noise
  • 5.1. Data Filtering
  • 5.2. Adding Noise
  • 5.2.1. Independent Noise
  • 5.2.2. Dependent Noise: Back-Translation
  • 6. Conclusions
  • 7. Acknowledgements
  • References
  • A. Scaling Law Fitting Details
  • A.1. Separate Fits and Variance
  • A.2. Optimizing Hyperparameters
  • A.3. Variance-Limited Regime
  • B. BLEU Score Behavior
  • C. Scaling Laws for Different OOD Datasets
  • D. Data Scaling Phase Transition
  • E. Scaling Laws with Different Architectures
  • F. Changing Language Pairs

Knowls

  1. Knowl 1 — Power-law scaling of translation loss with training data

    equation

    For a fixed neural machine translation model, test log-perplexity is fit by L(D)=α(D−1+C)pL(D)=\alpha(D^{-1}+C)^p, where DD is the training-set size measured numerically in millions of sentence pairs, L(D)L(D) is test log-perplexity, and α\alpha, CC, and pp are fitted constants. Here CC is a model-capacity-associated offset with the same inverse-data units as D−1D^{-1}, and pp is the data-scaling exponent. On English-to-German in-house parallel data, encoder-decoder Transformers ranging from 170M to 800M parameters were evaluated over approximately 500K–512M sentence pairs. The fitted law described the observed losses across this range; model-shape fits gave exponents in the approximate 0.21–0.29 range, with a shared fit around 0.25. The results establish this functional form empirically for the tested translation models and data range, rather than as a universal law.

  2. Knowl 2 — Joint scaling law for data and encoder-decoder model size

    equation

    A joint data-and-model-size fit replaces the capacity offset in the data scaling law with a model-size-dependent quantity: L(D;Ne,Nd)=α(D−1+CNe,Nd)pL(D;N_e,N_d)=\alpha(D^{-1}+C_{N_e,N_d})^p, where CNe,Nd=β(Ne−peNd−pd+L∞)1/pC_{N_e,N_d}=\beta(N_e^{-p_e}N_d^{-p_d}+L_\infty)^{1/p}. Here DD is the number of training sentence pairs in millions, LL is test log-perplexity, and NeN_e and NdN_d are the encoder and decoder parameter counts; α\alpha and pp are fitted to the data-loss observations. The parameters β\beta, pep_e, pdp_d, and L∞L_\infty come from a previously measured NMT parameter-scaling law, with L∞L_\infty representing its limiting loss term. As DD grows without bound, the joint law approaches the parameter-scaling prediction. Fitting α\alpha and pp without the 6-layer-encoder/2-layer-decoder model still gave accurate predictions for that held-out model’s test losses.

  3. Knowl 3 — Data-limited and capacity-limited regimes

    theoretical result

    For the fitted loss law L(D)=α(D−1+C)pL(D)=\alpha(D^{-1}+C)^p, with DD in millions of sentence pairs and CC the capacity-associated inverse-data offset, the condition D−1≫CD^{-1}\gg C defines a data-limited regime: loss is approximately proportional to D−pD^{-p}. With the observed exponents near p=0.25p=0.25, the magnitude of the marginal loss reduction per additional unit of data is of order D−1.25D^{-1.25}. When D−1≪CD^{-1}\ll C, the model is capacity-limited: loss approaches the floor αCp\alpha C^p, its data-dependent correction is of order D−1D^{-1}, and the marginal reduction is of order D−2D^{-2}. The transition between regimes occurs approximately at CD=1CD=1. Increasing model capacity reduces CC and therefore moves this transition to larger datasets. Fits restricted to datasets of at least 32M pairs were consistent with a variance-limited exponent near one for some model shapes, while the larger models had not reached one over the measured range.

  4. Knowl 4 — Architecture changes mostly shift the scaling curve

    empirical result

    English-to-German experiments compared roughly 300M-parameter encoder-decoder Transformers, Transformer-encoder/LSTM-decoder hybrids, and decoder-only Transformers trained with a language-modeling objective that also includes source-side loss. Over increasing dataset sizes, the three setups were described by a shared exponent p=0.285p=0.285 in L(D)=α(D−1+C)pL(D)=\alpha(D^{-1}+C)^p, where DD is measured in millions of sentence pairs and LL is test log-perplexity. The fitted (α,C)(\alpha,C) values were (1.969,0.057)(1.969,0.057) for encoder-decoder, (2.011,0.078)(2.011,0.078) for hybrid-LSTM, and (1.817,0.11)(1.817,0.11) for decoder-only models. Thus, within these tested configurations, architecture and task changes shifted the curve through α\alpha and CC without materially changing its exponent; in the data-limited regime, the different setups can consequently have similar sample-efficiency trends despite different losses at a fixed dataset size.

  5. Knowl 5 — Filtering changes finite-data performance more than the exponent

    empirical result

    On English-to-German ParaCrawl data, the study began with about 750M sentence pairs after deduplication, length filtering (sentences no longer than 256 tokens), language identification, and removal of near-duplicates of test data. A 6L6L Transformer was trained on increasing subsets up to 256M pairs. The compared training sets were unfiltered data, data retained by a Bicleaner score threshold of 0.5 (about 300M pairs), and the top 50% ranked by Contrastive Data Selection (CDS). Their losses fit L(D)=α(D−1+C)pL(D)=\alpha(D^{-1}+C)^p with common p=0.278p=0.278 and, respectively, (α,C)=(2.501,0.034)(\alpha,C)=(2.501,0.034), (2.130,0.064)(2.130,0.064), and (2.235,0.054)(2.235,0.054); DD is measured in millions of pairs. The lower α\alpha for filtered data corresponds to lower loss at a given finite dataset size in these experiments, while the observed large-data losses were reported to converge to approximately the same value. Thus filtering can benefit a constrained data or compute budget, but the similar exponents indicate that more unfiltered data can recover comparable performance in the tested setting.

  6. Knowl 6 — Independent synthetic noise preserves the exponent but can worsen the loss floor

    empirical result

    A 6L6L English-to-German Transformer was trained on clean parallel data and on versions with independent synthetic noise applied separately to source or target sentences. The perturbations changed 10% of characters to random alphanumeric or punctuation characters, deleted 15% of words, or shuffled the sentence-pair mapping for 10% of sentence pairs. The clean, source-noise, and target-noise curves fit L(D)=α(D−1+C)pL(D)=\alpha(D^{-1}+C)^p with common p=0.296p=0.296 and respective (α,C)(\alpha,C) values (1.969,0.064)(1.969,0.064), (2.222,0.067)(2.222,0.067), and (2.772,0.323)(2.772,0.323), where DD is in millions of sentence pairs. The noisy datasets did not approach the clean dataset’s limiting loss, and target-side noise was more damaging than source-side noise. The exponent’s stability therefore does not imply that additional data removes every effect of noise: in these experiments, noise also changed the limiting performance.

  7. Knowl 7 — Back-translated data has a lower scaling exponent than parallel data

    empirical result

    To compare dependent translation noise with human-generated parallel data, the study kept the German target sentences fixed and generated English source sentences using German-to-English back-translation models with 2L6L, 6L6L, 32L6L, and 64L6L architectures. A 6L6L English-to-German model was then trained on increasing subsets of each corpus. The back-translated-data curves had a shared exponent of about 0.1980.198, compared with about 0.2710.271 for parallel data, and approached a worse large-data loss. Increasing the back-translation model size improved the fitted dataset-dependent curve parameters α\alpha and CC but did not materially change the exponent. The performance advantage of stronger back-translation was especially relevant at very large data sizes; even the largest tested back-translation model remained substantially behind parallel data in the large-data regime.

  8. Knowl 8 — Most out-of-distribution test sets scale similarly to in-distribution data

    empirical result

    For a 6L6L English-to-German Transformer, loss was measured on 14 evaluation sets spanning web, news, patent, and Wikipedia domains, including source-original sets formed from natural source text and target-original sets formed by translating natural target text back to the source language. Most sets followed a similar power-law scaling trend as the training-distribution test set, with no major systematic difference in exponent by test-set composition. Because in-distribution and out-of-distribution losses changed similarly as training data increased, their losses had a nearly linear relationship in the evaluated sets. This is an empirical result for the tested domains and compositions; the reason for the similar scaling behavior was left unresolved.

  9. Knowl 9 — Similar architecture scaling was observed for Chinese-to-English translation

    empirical result

    A subset of the architecture comparison was repeated on an in-house Chinese-to-English parallel corpus, using a 6L6L encoder-decoder Transformer and a 6L6L Transformer-encoder/LSTM-decoder hybrid. The models were evaluated across dataset sizes, with dropout values from 0.1 to 0.5; the best-performing dropout at each size formed a Pareto frontier for fitting. A shared data-scaling exponent of approximately 0.250.25 described the two architectures’ losses. This provides cross-language evidence that the similar scaling trends found for these architecture families are not specific to English-to-German, while remaining limited to the architectures and experiments repeated for Chinese-to-English.

  10. Knowl 10 — The observed invariance of exponents is empirical and bounded in scope

    limitation

    The paper’s conclusion that architecture and data-quality changes often leave the data-scaling exponent nearly unchanged is based on experiments with particular NMT architectures, noise interventions, datasets, and language pairs, not a general guarantee. The authors explicitly caution that the pattern may not hold for very different architecture classes or noise types. The study also does not establish why the exponents remain similar; interpreting an exponent as a signature of a shared learning mechanism or data-manifold dimension is presented as a hypothesis, not a demonstrated explanation.

Coverage note — The secondary analysis relating BLEU to log-perplexity, and the appendix diagnostics on training randomness and dropout tuning, are omitted because they support evaluation and robustness rather than the paper’s main data-scaling comparisons.

References

  1. 1.Amari, S.-i., Fujita, N., and Shinomoto, S. Four types of learning curves. Neural Computation, 4, 01 2001. doi: 10.1162/neco.1992.4.4.605.
  2. 2.Anonymous. Scaling laws vs model architectures: How does inductive bias influence scaling? an extensive empirical study on language tasks, 2021.
  3. 3.Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G., Cherry, C., Macherey, W., Chen, Z., and Wu, Y. Massively multilingual neural machine translation in the wild: Findings and challenges, 2019.
  4. 4.Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U. Explaining neural scaling laws. arXiv preprint arXiv:2102.06701, 2021.
  5. 5.Bañon, M., Chen, P., Haddow, B., Heafield, K., Hoang, H., Espla-Gomis, M., Forcada, M. L., Kamran, A., Kirefu, F., Koehn, P., Ortiz Rojas, S., Pla Sempere, L., Ramírez-Sanchez, G., Sarrías, E., Strelec, M., Thompson, B., Waites, W., Wiggins, D., and Zaragoza, J. ParaCrawl: Web-scale acquisition of parallel corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4555–4567, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.417. URL https://aclanthology.org/2020.acl-main.417.
  6. 6.Barrault, L., Bojar, O., Costa-Jussa, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., et al. Findings of the 2019 conference on machine translation (wmt19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pp. 1–61, 2019.
  7. 7.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  8. 8.Chen, M. X., Firat, O., Bapna, A., Johnson, M., Macherey, W., Foster, G., Jones, L., Parmar, N., Schuster, M., Chen, Z., Wu, Y., and Hughes, M. The best of both worlds: Combining recent advances in neural machine translation, 2018.
  9. 9.El-Kishky, A., Chaudhary, V., Guzman, F., and Koehn, P. Ccaligned: A massive collection of cross-lingual web-document pairs, 2020.
  10. 10.Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., Goyal, N., Birch, T., Liptchinsky, V., Edunov, S., Grave, E., Auli, M., and Joulin, A. Beyond english-centric multilingual machine translation, 2020.
  11. 11.Freitag, M., Caswell, I., and Roy, S. APE at Scale and Its Implications on MT Evaluation Biases. In Proceedings of the Fourth Conference on Machine Translation, pp. 34–44, Florence, Italy, August 2019. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W19-5204.
  12. 12.Freitag, M., Grangier, D., and Caswell, I. BLEU might be guilty but references are not innocent. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, November 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.emnlp-main.5.
  13. 13.Freitag, M., Foster, G., Grangier, D., Ratnakar, V., Tan, Q., and Macherey, W. Experts, errors, and context: A large-scale study of human evaluation for machine translation, 2021.
  14. 14.Gao, L. An empirical exploration in quality filtering of text data, 2021.
  15. 15.Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C. Scaling laws for neural machine translation, 2021.
  16. 16.Google. Recent advances in google translate. https://ai.googleblog.com/2020/06/recent-advances-in-google-translate.html.
  17. 17.Gordon, M., Duh, K., and Kaplan, J. Data and parameter scaling laws for neural machine translation. ACL Rolling Review-May, 2021, 2021.
  18. 18.Graham, Y., Haddow, B., and Koehn, P. Statistical power and translationese in machine translation evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 72–81, 2020.
  19. 19.Hestness, J., Narang, S., Ardalani, N., Diamos, G. F., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. Deep learning scaling is predictable, empirically. ArXiv, abs/1712.00409, 2017.
  20. 20.Hoiem, D., Gupta, T., Li, Z., and Shlapentokh-Rothman, M. Learning curves for analysis of deep networks. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 4287–4296. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/hoiem21a.html.
  21. 21.Junczys-Dowmunt, M. Dual conditional cross-entropy filtering of noisy parallel corpora. arXiv preprint arXiv:1809.00197, 2018.
  22. 22.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  23. 23.Kasai, J., Pappas, N., Peng, H., Cross, J., and Smith, N. A. Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation. arXiv preprint arXiv:2006.10369, 2020.
  24. 24.Khayrallah, H. and Koehn, P. On the impact of various types of noise on neural machine translation. Proceedings of the 2nd Workshop on Neural Machine Translation and Generation, 2018. doi: 10.18653/v1/w18-2709. URL http://dx.doi.org/10.18653/v1/W18-2709.
  25. 25.Kocmi, T., Federmann, C., Grundkiewicz, R., Junczys-Dowmunt, M., Matsushita, H., and Menezes, A. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation, 2021.
  26. 26.Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N. Big transfer (bit): General visual representation learning, 2020.
  27. 27.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25: 1097–1105, 2012.
  28. 28.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding, 2020.
  29. 29.Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., and Zettlemoyer, L. Multilingual denoising pre-training for neural machine translation, 2020.
  30. 30.Mathur, N., Wei, J., Freitag, M., Ma, Q., and Bojar, O. Results of the WMT20 metrics shared task. In Proceedings of the Fifth Conference on Machine Translation, pp. 688–725, Online, November 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.wmt-1.77.
  31. 31.Microsoft. Neural machine translation enabling human parity innovations in the cloud. https://www.microsoft.com/en-us/translator/blog/2019/06/17/neural-machine-translation-enabling-human-parity-innovations-in-the-cloud.
  32. 32.Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L. Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In International Conference on Machine Learning, pp. 7721–7735. PMLR, 2021.
  33. 33.Moore, R. C. and Lewis, W. Intelligent selection of language model training data. In Proceedings of the ACL 2010 Conference Short Papers, pp. 220–224, Uppsala, Sweden, July 2010. Association for Computational Linguistics. URL https://aclanthology.org/P10-2041.
  34. 34.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer, 2020.
  35. 35.Ramírez-Sanchez, G., Zaragoza-Bernabeu, J., Bañon, M., and Rojas, S. O. Bifixer and bicleaner: two open-source tools to clean your parallel data. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, pp. 291–298, Lisboa, Portugal, November 2020. European Association for Machine Translation. URL https://aclanthology.org/2020.eamt-1.31.
  36. 36.Rei, R., Stewart, C., Farinha, A. C., and Lavie, A. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, November 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.emnlp-main.213.
  37. 37.Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N. A constructive prediction of the generalization error across scales. arXiv preprint arXiv:1909.12673, 2019.
  38. 38.Sellam, T., Das, D., and Parikh, A. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, July 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.acl-main.704.
  39. 39.Sennrich, R., Haddow, B., and Birch, A. Improving neural machine translation models with monolingual data, 2016.
  40. 40.Sharma, U. and Kaplan, J. A neural scaling law from the dimension of the data manifold, 2020.
  41. 41.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost, 2018.
  42. 42.Siddhant, A., Bapna, A., Cao, Y., Firat, O., Chen, M., Kudugunta, S., Arivazhagan, N., and Wu, Y. Leveraging monolingual data with self-supervision for multilingual neural machine translation, 2020.
  43. 43.Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C.-C., Krivokon, I., Rusch, W., Pickett, M., Meier-Hellstern, K., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E., and Le, Q. Lamda: Language models for dialog applications, 2022.
  44. 44.Wang, S., Tu, Z., Tan, Z., Wang, W., Sun, M., and Liu, Y. Language models are good translators, 2021.
  45. 45.Wang, W., Watanabe, T., Hughes, M., Nakagawa, T., and Chelba, C. Denoising neural machine translation training with trusted data and online data selection, 2018.
  46. 46.Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y. On layer normalization in the transformer architecture, 2020.
  47. 47.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079, 2020.
  48. 48.Zhang, M. and Toral, A. The effect of translationese in machine translation test sets, 2019.

Citation

MLA
Bansal, Y., et al. “Data Scaling Laws in NMT: The Effect of Noise and Architecture”. International Conference on Machine Learning, vol. 162, 2022, pp. 1466–82, https://proceedings.mlr.press/v162/bansal22b.html.
APA
Bansal, Y., Ghorbani, B., Garg, A., Zhang, B., Cherry, C., Neyshabur, B., & Firat, O. (2022). Data Scaling Laws in NMT: The Effect of Noise and Architecture. International Conference on Machine Learning, 162, 1466–1482. https://proceedings.mlr.press/v162/bansal22b.html
Chicago
Bansal, Y., B. Ghorbani, A. Garg, et al. 2022. “Data Scaling Laws in NMT: The Effect of Noise and Architecture”. International Conference on Machine Learning 162: 1466–82. https://proceedings.mlr.press/v162/bansal22b.html.
Harvard
Bansal, Y. et al. (2022) “Data Scaling Laws in NMT: The Effect of Noise and Architecture”, International Conference on Machine Learning. PMLR, pp. 1466–1482. Available at: https://proceedings.mlr.press/v162/bansal22b.html.
Vancouver
1. Bansal Y, Ghorbani B, Garg A, Zhang B, Cherry C, Neyshabur B, Firat O (2022) Data Scaling Laws in NMT: The Effect of Noise and Architecture. In: International Conference on Machine Learning. PMLR, pp 1466–1482

BibTeX

@InProceedings{pmlr-v162-bansal22b,
  title = 	 {Data Scaling Laws in {NMT}: The Effect of Noise and Architecture},
  author =       {Bansal, Yamini and Ghorbani, Behrooz and Garg, Ankush and Zhang, Biao and Cherry, Colin and Neyshabur, Behnam and Firat, Orhan},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {1466--1482},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/bansal22b/bansal22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/bansal22b.html},
  abstract = 	 {In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/