Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models
Hong LiuSang Michael XieZhiyuan LiTengyu Ma
Reveals that language models with identical pre-training losses can achieve substantially different downstream performance, proving both theoretically and empirically that SGD's implicit bias toward flatter loss minima governs transferability.
In modern artificial intelligence, large language models are typically evaluated and monitored using pre-training loss, which measures how well a model predicts masked or next words on training data. Because pre-training loss generally correlates with how well a model performs on practical downstream applications, conventional practice assumes that achieving a lower or equal pre-training loss guarantees equivalent or better downstream utility. However, this assumption fails to explain why different models with identical pre-training losses often display markedly different downstream accuracy.
The article investigates why models with the same pre-training loss achieve different downstream performance and demonstrates that the implicit bias of training algorithms—specifically their tendency to favor flatter loss minima—serves as a primary driver of downstream transferability.
To evaluate this dynamic, the authors conducted controlled empirical experiments across synthetic languages (probabilistic context-free grammars and hidden Markov models) and real-world corpora (OpenWebText and BookCorpus) using transformer models ranging from 2 million to 950 million parameters. In these environments, the authors examined models operating in a saturation regime where pre-training loss matches the optimal theoretical bound. They tested the effects of continuing training past convergence, altering model scale, and varying optimization techniques, while theoretically analyzing stochastic gradient descent and mathematically proving feature learning behaviors in a synthetic Dyck language setting.
The findings reveal three primary insights. First, pre-training loss alone does not dictate downstream success: models with identical, optimal pre-training losses showed substantial differences in downstream accuracy depending on training duration, model size, and optimization techniques. Second, standard optimization algorithms inherently possess an implicit bias toward flatter minima (measured by the trace of the Hessian), and this flatness strongly correlates with superior downstream accuracy where pre-training loss ceases to be predictive. For example, continuing training after loss convergence, scaling model capacity, or adding explicit flatness regularization improved downstream accuracy by several percentage points while loss remained flat. Third, theoretical analyses confirmed that standard stochastic gradient descent naturally oscillates along optimal loss manifolds toward flatter configurations, and in formal language settings, only the flattest models learn generalizable structural representations rather than memorizing data.
These results demonstrate that downstream performance depends heavily on the geometry of the learned parameter space rather than empirical pre-training loss alone. Relying exclusively on pre-training loss creates operational risks, such as prematurely halting training or selecting models that fit pre-training data via brittle memorization. Consequently, explicit regularization methods that promote flatness, such as weight decay, dropout, or sharpness-aware minimization, provide direct transferability benefits across downstream tasks.
For technical leaders and practitioners, the article suggests incorporating flatness metrics into model evaluation pipelines rather than relying solely on validation loss. When training budgets permit, organizations should continue pre-training past initial loss convergence or apply explicit flatness regularization to maximize transferability. Future work should focus on extending these theoretical guarantees to adaptive optimizers like Adam and developing computationally lightweight flatness regularizers suitable for large-scale production training.
The findings are bounded by certain theoretical and empirical limitations. While the theoretical proofs strictly address stochastic gradient descent and simplified synthetic grammars, the practical training of large language models frequently relies on adaptive optimizers such as AdamW on complex web-scale text. Nonetheless, the high consistency between empirical results across real-world datasets and theoretical models provides strong confidence in the core conclusion that flatness-driven implicit bias is a fundamental factor in model transferability.
- Paper: Visualizing the Loss Landscape of Neural Nets, Hao Li et al. (2017). Its method for visualizing neural loss landscapes and relating curvature to generalization provides essential grounding for the source’s analysis of flat minima and Hessian trace.
- Paper: Exploring Generalization in Deep Learning, Behnam Neyshabur et al. (2017). It develops the broader generalization debate around sharpness and its limitations, clarifying the theoretical context for the source’s flatness-based account.
- Paper: Why Does Unsupervised Pre-training Help Deep Learning?, Dumitru Erhan et al. (2010). Its account of pre-training as a regularizer establishes the earlier explanation of why pre-training can improve generalization that the source revisits through optimization bias.
- Paper: A Modern Look at the Relationship between Sharpness and Generalization, Maksym Andriushchenko et al. (2023). It tests whether modern, scale-invariant sharpness measures actually predict generalization, providing a valuable empirical qualification to the source’s flatness-transfer findings.
- Paper: Omnigrok: Grokking Beyond Algorithmic Data, Ziming Liu et al. (2023). It extends the study of delayed generalization by tracing grokking to loss-landscape structure and examining how regularization shapes the path from memorization to transfer.
- Paper: Understanding Emergent Abilities of Language Models from the Loss Perspective, Zhengxiao Du et al. (2024). It directly tests the competing claim that pre-training loss alone predicts downstream performance, making it a timely continuation of the source’s challenge to loss-based evaluation.
