Transcending Scaling Laws with 0.1% Extra Compute

Yi TayJason WeiHyung Won ChungVinh Q. TranDavid R. SoSiamak ShakeriXavier GarciaHuaixiu Steven ZhengJinfeng RaoAakanksha Chowdhery

article2023EMNLP79 citations

Demonstrates that continuing to train pretrained large language models on UL2's mixture-of-denoisers objective with roughly 0.1% additional compute dramatically improves scaling curves, achieving up to a 2x compute savings and triggering emergent reasoning capabilities at smaller model scales.

Listen

State-of-the-art large language models require massive amounts of computation to train, and standard practice typically relies on left-to-right causal language modeling. As organizations seek higher performance and specialized reasoning capabilities, scaling up models or pretraining them from scratch on larger datasets becomes increasingly expensive. There is an urgent operational and environmental need for methods that maximize model quality and functional capabilities without incurring substantial computational costs.

To address this challenge, the article evaluates a method called UL2R (UL2Restore). The primary objective is to demonstrate that continuing the training of an existing large language model with a diverse mixture of training objectives significantly enhances model quality, efficiency, and task capabilities using only a tiny fraction of extra compute.

To test this approach, the researchers applied UL2R to existing PaLM models across three scales: 8 billion, 62 billion, and 540 billion parameters, creating an adapted model family termed U-PaLM. Instead of introducing new data, the process reused the original training corpus and added approximately 0.1% to 0.16% additional training computation (about 1.3 billion tokens for the largest model). Training incorporated a mixture-of-denoisers objective combining prefix language modeling (bidirectional attention over inputs) with regular and extreme span corruption (infilling tasks). Performance was evaluated across a broad suite of standard benchmarks, including zero-shot and few-shot natural language processing tasks, the BIG-Bench emergent suite, Massively Multi-Task Language Understanding (MMLU), and multilingual reasoning benchmarks.

The investigation produced several key findings. First, U-PaLM achieved roughly a 2x computational savings rate at the 540-billion-parameter scale, matching the quality of the final baseline model with only half the compute and saving approximately 4.4 million TPUv4 accelerator hours. Second, U-PaLM outperformed the baseline on 21 of 26 standard zero- and few-shot benchmarks and beat baseline scores across 19 of 21 challenging BIG-Bench tasks. Third, the method unlocked advanced reasoning abilities at smaller model sizes, allowing 62-billion or 8-billion parameter models to succeed on tasks where standard models required 540 billion parameters to perform better than random guessing. Fourth, the method added new practical querying capabilities, including bidirectional text infilling and mode-specific prompting to generate more diverse and accurate outputs.

These findings indicate that architectural flexibility and objective diversity can shift standard scaling curves more effectively than simply continuing causal language modeling. For practitioners and decision-makers, this translates to major reductions in training costs, shorter development timelines, and smaller environmental footprints, while simultaneously improving downstream performance on complex tasks. It shows that high-performing foundation models do not always need to be rebuilt from scratch to gain new competencies.

Based on these results, organizations should consider adopting brief multi-objective adaptation phases like UL2R as a standard post-pretraining step to upgrade existing models cost-effectively. Teams should also explore infilling-based prompt designs for structured problem-solving workflows. Before broad enterprise deployment across different model families, further validation is recommended to determine whether similar efficiency gains occur when adapting non-PaLM architectures or fully converged models trained on alternate datasets.

While the findings demonstrate high confidence within the evaluated setups, the primary limitation is that the experiments focused strictly on PaLM models ranging from 8 billion to 540 billion parameters and reused a specific pretraining corpus. Additional testing is needed to confirm generalizability across smaller open-source models, varying data distributions, and saturated pretraining states.

Tay et al (2023).pdf
  • Paper: PaLM: Scaling Language Modeling with Pathways, Aakanksha Chowdhery et al. (2023). PaLM is the direct base-model and training-scale foundation for UL2R, so its architecture, compute setup, and baseline capabilities clarify what the adaptation changes.
  • Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Chinchilla establishes compute-optimal scaling as the benchmark UL2R seeks to transcend, making its model–data allocation results essential context for the paper’s efficiency claims.
  • Paper: PaLM 2 Technical Report, Rohan Anil et al. (2023). PaLM 2’s mixed-objective training and compute-efficiency discussion provides a close methodological precedent for understanding UL2R’s use of diverse training objectives.
Cover for Transcending Scaling Laws with 0.1% Extra Compute

Abstract

Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute. The key idea is to continue training a state-of-the-art large language model on a few more steps with UL2’s mixture-of-denoiser objective. We show that, with almost negligible extra computational costs and no new sources of data, we are able to substantially improve the scaling properties of large language models on downstream metrics. In this paper, we continue training a baseline language model, PaLM, with UL2R, introducing a new set of models at 8B, 62B, and 540B scale which we call U-PaLM. Impressively, at 540B scale, we show an approximately 2x computational savings rate where U-PaLM achieves the same performance as the final PaLM 540B model at around half its computational budget (i.e., saving ∼4.4 million TPUv4 hours). We further show that this improved scaling curve leads to “emergent abilities” on challenging BIG-Bench tasks—for instance, U-PaLM does much better on some tasks or demonstrates better quality at much smaller scale (62B as opposed to 540B). Overall, we show that U-PaLM outperforms PaLM on many few-shot setups, including reasoning tasks with chain-of-thought (e.g., GSM8K), multilingual tasks (MGSM, TydiQA), MMLU and challenging BIG-Bench tasks.

Table of Contents

  • 1 Introduction
  • 2 U-PaLM
  • 2.1 Training Data
  • 2.2 Prefix Language Model Architecture
  • 2.3 Loss Objectives
  • 2.4 Training
  • 3 Experiments
  • 3.1 Improved Scaling Properties on Few-shot Learning
  • 3.2 BigBench Emergent Suite
  • 3.2.1 BIG-Bench Results
  • 3.2.2 MMLU Results
  • 3.3 Finetuning
  • 3.4 Additional Results & Analysis
  • 4 Qualitative Analysis: New Prompting Capabilities
  • 4.1 Infilling Ability
  • 4.2 Leveraging Specific Pretraining Modes
  • 4.3 Improved Diversity for Open-ended Generation
  • 5 Conclusion
  • 6 Limitations
  • 7 Ethics Statement
  • 8 Acknowledgements
  • References
  • 9 Appendix
  • 10 Related Work
  • 11 Additional Results & Analysis
  • 11.0.1 Analyzing individual task performance on BIG-Bench
  • 11.1 Zero-shot and Few-shot NLP
  • 11.1.1 Commonsense Reasoning
  • 11.1.2 Question Answering and Reading Comprehension
  • 11.1.3 Reasoning and Chain-of-thought Experiments
  • 11.1.4 Multilingual Few-shot Reasoning and Question Answering Tasks
  • 11.2 Details of Scaling Curves for Few-shot Experiments
  • 11.3 Details of Vocab and Sentinel Tokens
  • 11.4 Details of Prompt Templates
  • 11.5 Additional Discussion
  • 11.5.1 What about training from scratch?

Knowls

  1. Knowl 1 — UL2R adapts a causal language model through continued denoising training

    model/method

    UL2R (UL2Restore) continues training an existing causal language model with UL2’s mixture-of-denoisers objective, rather than training a new model from scratch. The resulting models, called U-PaLM, start from PaLM checkpoints and retain the PaLM architecture. The training reuses PaLM’s data mixture and introduces no new data sources. The paper estimates that this restoration requires roughly 0.1%–1% extra computation, and reports that the approach improves downstream performance and enables prompting behaviors not supported by the original causal-only training.

  2. Knowl 2 — U-PaLM’s continued training uses a three-way mixture of denoising tasks

    model/method

    The main U-PaLM training mixture assigns 50% of examples to sequential denoising, 25% to regular span corruption, and 25% to extreme span corruption. Sequential denoising predicts text after a prefix and corresponds to the PrefixLM objective. Regular span corruption replaces sampled spans with sentinel tokens; its spans have mean length 3 and the corruption rate is 15%. Extreme span corruption uses either longer spans, typically with mean length 32, or a higher corruption rate of up to 50%. The corresponding mode tokens are [S2S] for sequential denoising, [NLU] for regular denoising, and [NLG] for extreme denoising. The authors initially retained UL2’s seven denoisers, but found this three-task mixture simple and effective for continued training.

  3. Knowl 3 — U-PaLM training configuration and PrefixLM sequence preparation

    experimental setup

    U-PaLM uses a PrefixLM architecture: input tokens are given bidirectional attention, and each training sequence contains 1,024 input tokens and 1,024 target tokens, for a combined length of 2,048. To avoid wasting sequence capacity on padding between the prefix and target, the training pipeline concatenates them before applying padding, packing, and trimming. For the 540B model, continued training ran for 20,000 steps with batch size 32, using approximately 1.3 billion additional tokens—reported as 0.16% extra computation relative to PaLM’s 780 billion-token pretraining. The learning rate followed a cosine decay from 10−410^{-4} to 10−610^{-6}; a low constant learning rate performed similarly in the authors’ tests. The 540B run used 512 TPUv4 chips and took about five days; the 8B and 62B runs used 64 TPUv4 chips. To support span corruption without adding new vocabulary entries to the PaLM checkpoint, the final 100 existing subwords were used as sentinel tokens.

  4. Knowl 4 — UL2R improves the downstream-performance scaling curve at lower compute

    empirical result

    In scaling experiments on PaLM 8B and PaLM 540B, the authors compared average zero- and few-shot performance across downstream NLP tasks as training compute increased. Continued UL2R training moved the performance-versus-compute curve upward relative to continuing PaLM with causal language modeling. At 540B, the reported comparable-performance comparison is approximately 2.35×: PaLM at about 2,500 zFLOPs and U-PaLM at about 1,075 zFLOPs achieved similar average task performance. The paper also describes savings of about 2× at a middle checkpoint, corresponding to approximately 4.4 million TPUv4 hours for the 540B model. The 8B gains were largest around the middle of PaLM training and narrowed as the baseline approached a plateau; for 540B, the gap continued to grow through the evaluated 780-billion-token checkpoint. The authors caution that the 540B baseline was not trained to convergence, so the savings rate might have increased further.

  5. Knowl 5 — U-PaLM improves results across standard zero- and few-shot NLP evaluations

    empirical result

    On a 26-task zero- and few-shot NLP suite, the 540B U-PaLM trained for 780 billion tokens outperformed the corresponding PaLM on 21 of 26 tasks. In a separate aggregate scaling evaluation over the same number of tasks, mean scores at 182B, 329B, and 780B training tokens were 62.7, 63.8, and 66.5 for PaLM, compared with 64.1, 66.2, and 69.4 for U-PaLM at the corresponding checkpoints.

    For zero-shot commonsense reasoning, the mean across BoolQ, PIQA, HellaSwag, and Winogrande was 83.7 for PaLM 540B and 84.9 for U-PaLM 540B; at 62B the means were 80.5 and 80.7. On closed-book question answering and reading comprehension, the mean across TriviaQA, Natural Questions, and Lambada results was 58.7 for PaLM 540B and 60.1 for U-PaLM 540B, and 52.2 versus 54.3 at 62B. At 540B, U-PaLM’s few-shot Natural Questions score was 40.1, compared with 36.0 for PaLM. These results show gains across task families, though individual tasks do not all improve.

  6. Knowl 6 — U-PaLM raises the mean score on the 21-task BIG-Bench Emergent Suite

    empirical result

    On the BIG-Bench Emergent Suite (BBES), evaluated with the standard five-shot prompts and without chain-of-thought prompting, PaLM 540B scored 64.3 on average across 21 tasks and U-PaLM 540B scored 67.7, a reported relative gain of 5.3%. U-PaLM improved on 19 of the 21 tasks. The task scores (PaLM 540B to U-PaLM 540B, in percent) were: navigate 55.3 to 67.0; strategyqa 73.9 to 78.3; crass_ai 97.7 to 100; logical_sequence 92.3 to 86.5; vitaminc_fact_verification 70.2 to 73.9; understanding_fables 75.7 to 78.4; identify_odd_metaphor 87.2 to 87.5; hyperbaton 54.2 to 59.9; causal_judgment 65.3 to 68.4; english_proverbs 91.2 to 87.5; geometric_shapes 44.0 to 49.3; physics_questions 7.6 to 12.5; snarks 69.1 to 86.1; analogical_similarity 36.5 to 37.5; international_phonetic_alphabet_nli 65.9 to 68.0; movie_dialog_same_or_different 64.8 to 68.8; timedial 78.3 to 81.2; question_selection 54.8 to 59.8; logical_fallacy_detection 80.3 to 81.4; unit_interpretation 47.0 to 51.0; and language_identification 36.0 to 38.9. The largest absolute gains include snarks (+17.0 points) and navigate (+11.7 points). In scaling comparisons, performance on crass_ai, vitaminc_fact_verification, and identify_odd_metaphor began rising at 62B for U-PaLM, whereas PaLM’s gains on these tasks appeared only at 540B; U-PaLM 8B also surpassed PaLM 62B on snarks and understanding_fables.

  7. Knowl 7 — U-PaLM improves chain-of-thought, multilingual, and MMLU results

    empirical result

    With chain-of-thought prompting, U-PaLM 540B scored 58.5 on GSM8K versus 54.9 for PaLM 540B, 49.6 on BIG-Bench Hard versus 44.8, 76.6 on StrategyQA versus 76.4, and 80.1 on CommonsenseQA versus 76.9. On multilingual evaluations with chain-of-thought prompting, U-PaLM scored 49.9 on MGSM versus 45.9 for PaLM 540B and 54.6 on TyDiQA versus 52.9. On the five-shot MMLU test set without chain-of-thought prompting, U-PaLM 540B achieved 70.7% accuracy, compared with 69.3% for PaLM 540B. These results establish improvements on several reasoning and multilingual tasks, with the largest reported gains on BIG-Bench Hard and MGSM.

  8. Knowl 8 — UL2R improves SuperGLUE and TyDiQA fine-tuning results

    empirical result

    Fine-tuning experiments used a constant learning rate for 100,000 steps with batch size 32 and evaluated the SuperGLUE and TyDiQA development sets at 8B and 62B. SuperGLUE average scores were 83.4 for PaLM 8B and 86.1 for U-PaLM 8B (+3.2% as reported), and 89.5 for PaLM 62B and 91.4 for U-PaLM 62B (+2.1%). TyDiQA exact match/F1 scores were 75.7/85.2 for PaLM 8B and 77.5/86.7 for U-PaLM 8B; at 62B, scores were 78.3/87.3 for PaLM and 78.4/88.5 for U-PaLM. The reported TyDiQA relative gains were +2.3%/+1.7% at 8B and +0.1%/+2.1% at 62B. The gains were larger at 8B on SuperGLUE and modest but generally positive at 62B.

  9. Knowl 9 — UL2R gives U-PaLM infilling and mode-controlled prompting capabilities

    model/method

    UL2R equips U-PaLM with infilling: a prompt can contain one or more marked gaps, and the model can generate text for those gaps using sentinel tokens such as <extra_id_0>. The paper demonstrates this with a cake-instructions prompt where U-PaLM fills in a missing second step and with a prompt containing multiple missing steps; the original PaLM, which had not been trained on these tokens, does not perform the corresponding infill. U-PaLM can also be prompted with [S2S], [NLU], or [NLG] to select the sequential, regular-denoising, or extreme-denoising mode learned during training. For example, on an English-to-Vietnamese color question, the [S2S] prompt produced the specific Vietnamese answer “xanh lá cây,” while the default U-PaLM answer was less specific. The examples show that infilling and mode selection can alter how the model responds without changing its weights or inference algorithm; they are qualitative demonstrations, not a systematic success-rate evaluation.

  10. Knowl 10 — Evidence is limited to PaLM and does not establish generality across models or training regimes

    limitation

    The paper tests UL2R only on PaLM models and PaLM’s pretraining corpus, and its smallest evaluated models have 8B parameters. It therefore does not establish whether similar gains would occur with a different base architecture or corpus, a weaker model, a model already trained to saturation, or models smaller than 8B. The authors characterize the work as a demonstration of continued training on a near-state-of-the-art system, not a comprehensive study of model reuse or continued pretraining.

Coverage note — The appendix’s full per-task scores for the standard 26-task suite and additional qualitative examples of output diversity are omitted; they elaborate on results and prompting behaviors already captured here.

References

  1. 1.Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta. Muppet: Massive multi-task representations with pre-finetuning. arXiv preprint arXiv:2101.11038, 2021.
  2. 2.Philip W Anderson. More is different: broken symmetry and the nature of the hierarchical structure of science. Science, 177(4047):393–396, 1972.
  3. 3.Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Prakash Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler. Ext5: Towards extreme multi-task scaling for transfer learning. ICLR, 2022. URL https://arxiv.org/abs/2111.10952.
  4. 4.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020.
  5. 5.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  7. 7.Hyung Won Chung, Longpre Shayne Hou, Le and, Barret Zoph, Yi Tay, William Fedus, and et al. Scaling instruction-finetuned language models. arXiv preprint, 2022.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In NAACL, 2019.
  9. 9.Jonathan H Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8: 454–470, 2020a.
  10. 10.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020b.
  11. 11.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  12. 12.Mostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer, and Ashish Vaswani. The efficiency misnomer. In International Conference on Learning Representations, 2021.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou. Compositional semantic parsing with large language models. arXiv preprint arXiv:2209.15003, 2022.
  15. 15.Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova DasSarma, Tom Henighan, Andy Jones, Nicholas Joseph, Jackson Kernion, Ben Mann, Amanda Askell, et al. Predictability and surprise in large generative models. arXiv preprint arXiv:2202.07785, 2022. URL https://arxiv.org/abs/2202.07785.
  16. 16.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9: 346–361, 2021.
  17. 17.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  18. 18.Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022.
  19. 19.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  20. 20.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.
  21. 21.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  22. 22.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics, 2019.
  23. 23.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  24. 24.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858, 2022.
  25. 25.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022. URL https://arxiv.org/abs/2203.02155.
  26. 26.Denis Paperno, Germ’an Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern’andez. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525–1534, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1144. URL https://www.aclweb.org/anthology/P16-1144.
  27. 27.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  29. 29.Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020.
  30. 30.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. arXiv preprint arXiv:1907.10641, 2019.
  31. 31.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. ICLR, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  32. 32.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057, 2022.
  33. 33.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  34. 34.Jacob Steinhardt. Future ml systems will be qualitatively different, 2022. Accessed May 20, 2022.
  35. 35.Mirac Suzgun, Nathan Scales, Nathanael Scharli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint, 2022.
  36. 36.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1421. URL https://aclanthology.org/N19-1421.
  37. 37.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers. arXiv preprint arXiv:2109.10686, 2021.
  38. 38.Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling? arXiv preprint arXiv:2207.10551, 2022a.
  39. 39.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022b.
  40. 40.Yi Tay, Vinh Q Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al. Transformer memory as a differentiable search index. arXiv preprint arXiv:2202.06991, 2022c.
  41. 41.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  42. 42.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537, 2019.
  43. 43.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. What language model architecture and pretraining objective work best for zero-shot generalization? arXiv preprint arXiv:2204.05832, 2022a.
  44. 44.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022b.
  45. 45.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  46. 46.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. Transactions on Machine Learning Research (TMLR), 2022a.
  47. 47.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. Conference on Neural Information Processing Systems (NeurIPS), 2022b.
  48. 48.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  49. 49.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022.
  50. 50.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  51. 51.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.
  52. 52.Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. St-moe: Designing stable and transferable sparse expert models, 2022. URL https://arxiv.org/abs/2202.08906.

Citation

MLA
Tay, Y., et al. “Transcending Scaling Laws with 0.1% Extra Compute”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1471–86, https://doi.org/10.18653/v1/2023.emnlp-main.91.
APA
Tay, Y., Wei, J., Chung, H., Tran, V., So, D., Shakeri, S., Garcia, X., Zheng, S., Rao, J., Chowdhery, A., Zhou, D., Metzler, D., Petrov, S., Houlsby, N., Le, Q., & Dehghani, M. (2023). Transcending Scaling Laws with 0.1% Extra Compute. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1471–1486. https://doi.org/10.18653/v1/2023.emnlp-main.91
Chicago
Tay, Y., J. Wei, H. Chung, et al. 2023. “Transcending Scaling Laws with 0.1% Extra Compute”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1471–86. https://doi.org/10.18653/v1/2023.emnlp-main.91.
Harvard
Tay, Y. et al. (2023) “Transcending Scaling Laws with 0.1% Extra Compute”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1471–1486. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.91.
Vancouver
1. Tay Y, Wei J, Chung H, et al (2023) Transcending Scaling Laws with 0.1% Extra Compute. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1471–1486

BibTeX

@inproceedings{tay-etal-2023-transcending,
    title = "Transcending Scaling Laws with 0.1{\%} Extra Compute",
    author = "Tay, Yi  and
      Wei, Jason  and
      Chung, Hyung  and
      Tran, Vinh  and
      So, David  and
      Shakeri, Siamak  and
      Garcia, Xavier  and
      Zheng, Steven  and
      Rao, Jinfeng  and
      Chowdhery, Aakanksha  and
      Zhou, Denny  and
      Metzler, Donald  and
      Petrov, Slav  and
      Houlsby, Neil  and
      Le, Quoc  and
      Dehghani, Mostafa",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.91/",
    doi = "10.18653/v1/2023.emnlp-main.91",
    pages = "1471--1486"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/