Transcending Scaling Laws with 0.1% Extra Compute
Yi TayJason WeiHyung Won ChungVinh Q. TranDavid R. SoSiamak ShakeriXavier GarciaHuaixiu Steven ZhengJinfeng RaoAakanksha Chowdhery
Demonstrates that continuing to train pretrained large language models on UL2's mixture-of-denoisers objective with roughly 0.1% additional compute dramatically improves scaling curves, achieving up to a 2x compute savings and triggering emergent reasoning capabilities at smaller model scales.
State-of-the-art large language models require massive amounts of computation to train, and standard practice typically relies on left-to-right causal language modeling. As organizations seek higher performance and specialized reasoning capabilities, scaling up models or pretraining them from scratch on larger datasets becomes increasingly expensive. There is an urgent operational and environmental need for methods that maximize model quality and functional capabilities without incurring substantial computational costs.
To address this challenge, the article evaluates a method called UL2R (UL2Restore). The primary objective is to demonstrate that continuing the training of an existing large language model with a diverse mixture of training objectives significantly enhances model quality, efficiency, and task capabilities using only a tiny fraction of extra compute.
To test this approach, the researchers applied UL2R to existing PaLM models across three scales: 8 billion, 62 billion, and 540 billion parameters, creating an adapted model family termed U-PaLM. Instead of introducing new data, the process reused the original training corpus and added approximately 0.1% to 0.16% additional training computation (about 1.3 billion tokens for the largest model). Training incorporated a mixture-of-denoisers objective combining prefix language modeling (bidirectional attention over inputs) with regular and extreme span corruption (infilling tasks). Performance was evaluated across a broad suite of standard benchmarks, including zero-shot and few-shot natural language processing tasks, the BIG-Bench emergent suite, Massively Multi-Task Language Understanding (MMLU), and multilingual reasoning benchmarks.
The investigation produced several key findings. First, U-PaLM achieved roughly a 2x computational savings rate at the 540-billion-parameter scale, matching the quality of the final baseline model with only half the compute and saving approximately 4.4 million TPUv4 accelerator hours. Second, U-PaLM outperformed the baseline on 21 of 26 standard zero- and few-shot benchmarks and beat baseline scores across 19 of 21 challenging BIG-Bench tasks. Third, the method unlocked advanced reasoning abilities at smaller model sizes, allowing 62-billion or 8-billion parameter models to succeed on tasks where standard models required 540 billion parameters to perform better than random guessing. Fourth, the method added new practical querying capabilities, including bidirectional text infilling and mode-specific prompting to generate more diverse and accurate outputs.
These findings indicate that architectural flexibility and objective diversity can shift standard scaling curves more effectively than simply continuing causal language modeling. For practitioners and decision-makers, this translates to major reductions in training costs, shorter development timelines, and smaller environmental footprints, while simultaneously improving downstream performance on complex tasks. It shows that high-performing foundation models do not always need to be rebuilt from scratch to gain new competencies.
Based on these results, organizations should consider adopting brief multi-objective adaptation phases like UL2R as a standard post-pretraining step to upgrade existing models cost-effectively. Teams should also explore infilling-based prompt designs for structured problem-solving workflows. Before broad enterprise deployment across different model families, further validation is recommended to determine whether similar efficiency gains occur when adapting non-PaLM architectures or fully converged models trained on alternate datasets.
While the findings demonstrate high confidence within the evaluated setups, the primary limitation is that the experiments focused strictly on PaLM models ranging from 8 billion to 540 billion parameters and reused a specific pretraining corpus. Additional testing is needed to confirm generalizability across smaller open-source models, varying data distributions, and saturated pretraining states.
- Paper: PaLM: Scaling Language Modeling with Pathways, Aakanksha Chowdhery et al. (2023). PaLM is the direct base-model and training-scale foundation for UL2R, so its architecture, compute setup, and baseline capabilities clarify what the adaptation changes.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Chinchilla establishes compute-optimal scaling as the benchmark UL2R seeks to transcend, making its model–data allocation results essential context for the paper’s efficiency claims.
- Paper: PaLM 2 Technical Report, Rohan Anil et al. (2023). PaLM 2’s mixed-objective training and compute-efficiency discussion provides a close methodological precedent for understanding UL2R’s use of diverse training objectives.
- Paper: InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining, Boxin Wang et al. (2024). InstructRetro extends the idea of low-cost continued pretraining by adding retrieval and instruction tuning, showing how efficient adaptation can build retrieval-grounded capabilities.
- Paper: LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models, Yaowei Zheng et al. (2024). LlamaFactory carries efficient adaptation into a unified toolkit, making it possible to apply fine-tuning strategies across many model families and practical tasks.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). This study extends scaling-law analysis to data-constrained training, testing how compute allocation changes when adaptation and pretraining cannot rely on unlimited fresh data.
