Understanding Emergent Abilities of Language Models from the Loss Perspective

Zhengxiao DuAohan ZengYuxiao DongJie Tang

article2024NeurIPS112 citations

Demonstrates that pre-training loss reliably predicts downstream task performance across different model and data sizes, revealing that emergent abilities consistently appear only after loss drops below sharp, metric-independent thresholds.

Listen

Recent discussions in artificial intelligence have questioned the concept of emergent abilities—the sudden appearance of advanced skills in language models only after reaching massive scales. Skepticism arose because well-trained smaller models often match larger systems and because some researchers claimed emergence was an illusion caused by discontinuous scoring metrics like binary accuracy. The article evaluates whether a model's pre-training loss—a standard measure of how well a model predicts training text—serves as a more accurate and universal predictor of downstream task performance than parameter size or computing budget.

To test this relationship, the researchers trained more than 30 Transformer models ranging from 300 million to 32 billion parameters across varying training token budgets up to 3 trillion tokens, using a fixed bilingual English and Chinese corpus. They evaluated model checkpoints across 12 diverse benchmark datasets covering question answering, reading comprehension, reasoning, and mathematics. They also verified their findings against publicly available training trajectories from external model families, specifically LLaMA and Pythia.

The analysis reveals four key findings. First, pre-training loss directly predicts performance on downstream tasks across all model sizes and token counts; models with identical loss achieve identical task performance. Second, on challenging tasks such as multi-discipline examinations and mathematical word problems, performance remains strictly at random-guessing levels until pre-training loss drops below a specific tipping point of approximately 2.2. Third, emergent capabilities persist at this exact threshold even when evaluated using continuous probability metrics, disproving the claim that emergence is merely a measurement artifact. Fourth, pre-training loss is a significantly more reliable predictor of capability than total computing expenditure.

These findings suggest that emergent abilities should be redefined based on pre-training loss rather than physical model size. For strategic planning and investment, this framework enables organizations to accurately forecast the pre-training loss and compute budget required to achieve specific performance levels. It also demonstrates that capabilities on complex tasks cannot be predicted by simple extrapolation from earlier, high-loss training stages, as performance remains completely flat until the critical loss threshold is crossed.

Decision-makers should use pre-training loss targets rather than raw parameter scale to guide model development and resource allocation. However, organizations should avoid blindly expanding training compute under the assumption that further scaling will inevitably trigger additional sudden capabilities. In practice, teams should also leverage instruction tuning, which can enhance zero-shot task performance without relying entirely on massive pre-training runs.

The conclusions are limited to standard Transformer architectures trained with standard optimizers on fixed corpora; pre-training loss values cannot be compared directly across models trained on different tokenizers or data distributions. Furthermore, while the evidence strongly supports the loss-performance relationship across the evaluated benchmarks, the authors note that the emergence of new tipping points at scales beyond those tested remains uncertain.

arXiv: 2403.15796
  • Paper: Emergent Abilities of Large Language Models, Jason Wei et al. (2022). This foundational paper establishes the empirical phenomenon and initial definitions of emergent abilities in large language models that the source directly analyzes and reframes.
  • Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). It provides the foundational framework for neural scaling laws and power-law loss behavior across compute and model size that the source builds upon to evaluate emergence.
  • Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). It formalizes compute-optimal trade-offs between model size and training data in achieving target pre-training loss, which underpins the source's loss-centric view of emergent abilities.
Cover for Understanding Emergent Abilities of Language Models from the Loss Perspective

Abstract

Recent studies have put into question the belief that emergent abilities [58] in language models are exclusive to large models. This skepticism arises from two observations: 1) smaller models can also exhibit high performance on emergent abilities and 2) there is doubt on the discontinuous metrics used to measure these abilities. In this paper, we propose to study emergent abilities in the lens of pre-training loss, instead of model size or training compute. We demonstrate that the Transformer models with the same pre-training loss, but different model and data sizes, generate the same performance on various downstream tasks, with a fixed data corpus, tokenization, and model architecture. We also discover that a model exhibits emergent abilities on certain tasks—regardless of the continuity of metrics—when its pre-training loss falls below a specific threshold. Before reaching this threshold, its performance remains at the level of random guessing. This inspires us to redefine emergent abilities as those that manifest in models with lower pre-training losses, highlighting that these abilities cannot be predicted by merely extrapolating the performance trends of models with higher pre-training losses.

Table of Contents

  • 1 Introduction
  • 2 Does Pre-training Loss Predict Task Performance?
  • 2.1 Pre-training Setting
  • 2.2 Evaluation Tasks
  • 2.3 Pre-training Loss vs. Performance
  • 2.4 Training Token Count vs. Performance
  • 2.5 LLaMA's Loss vs. Performance
  • 3 Analysis of Different Tasks and Metrics
  • 3.1 Performance Trends of Different Tasks
  • 3.2 Influence of Different Metrics
  • 4 Defining Emergent Abilities from the Loss Perspective
  • 5 Related Work
  • 6 Conclusion
  • 7 Limitation
  • Acknowledgments and Disclosure of Funding
  • References
  • A Pre-training Settings
  • A.1 Pre-training Corpus
  • A.2 Hyperparameters
  • B Evaluation Settings
  • C Are Emergent Abilities of Language Models a Mirage?
  • D Complete Performance-vs-Loss Curves of Smaller Models
  • E Loss vs Compute as an Indicator of Performance
  • F Pythia's Loss vs. Performance
  • G Loss vs. Performance on BIG-bench
  • H Compute Resources
  • I Broader Impact
  • NeurIPS Paper Checklist
  • 1. Claims
  • 2. Limitations
  • 3. Theory Assumptions and Proofs
  • 4. Experimental Result Reproducibility
  • 5. Open access to data and code
  • 6. Experimental Setting/Details
  • 7. Experiment Statistical Significance
  • 8. Experiments Compute Resources
  • 9. Code Of Ethics
  • 10. Broader Impacts
  • 11. Safeguards
  • 12. Licenses for existing assets
  • 13. New Assets
  • 14. Crowdsourcing and Research with Human Subjects
  • 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects

Knowls

  1. Knowl 1 — Definition of Emergent Abilities from the Pre-Training Loss Perspective

    definition

    An ability in an autoregressive language model is defined as emergent with respect to pre-training loss if it is not present in models with higher pre-training loss, but is present in models with lower pre-training loss.

    Formally, let L∈R+L \in \mathbb{R}^+ denote the language model's pre-training cross-entropy loss, and let η∈R+\eta \in \mathbb{R}^+ denote a critical task-specific loss threshold. The normalized performance P(L)∈[0,1]P(L) \in [0, 1] on an emergent downstream task is given by:

    P(L)={f(L)if L<η0otherwiseP(L) = \begin{cases} f(L) & \text{if } L < \eta \\ 0 & \text{otherwise} \end{cases}

    where f(L)f(L) is a monotonically decreasing function of LL, and the performance metric is normalized such that random guessing corresponds to 00. Under this definition, task performance remains at the level of random guessing for all models where L≥ηL \ge \eta, and only demonstrates continuous improvement once the pre-training loss drops below η\eta.

  2. Knowl 2 — Model Size Threshold for Emergence Under Power-Law Loss Scaling

    theoretical result

    When downstream ability emergence is governed by a pre-training loss threshold η\eta, the corresponding critical model parameter size threshold NcritN_{\text{crit}} can be derived from empirical power-law loss scaling laws.

    Assuming the number of training tokens DD is fixed, the pre-training cross-entropy loss L(N)L(N) as a function of model parameter count NN follows the power-law relation:

    L(N)=L∞+(N0N)αNL(N) = L_\infty + \left( \frac{N_0}{N} \right)^{\alpha_N}

    where L∞L_\infty is the irreducible loss, N0N_0 is a baseline parameter scaling constant, and αN>0\alpha_N > 0 is the power-law exponent. Combining this scaling law with the loss-threshold performance formulation yields the normalized task performance as a function of model size NN:

    P(N)={f(L∞+(N0N)αN)if N≥N0⋅(η−L∞)−1αN0otherwiseP(N) = \begin{cases} f\left( L_\infty + \left( \frac{N_0}{N} \right)^{\alpha_N} \right) & \text{if } N \ge N_0 \cdot (\eta - L_\infty)^{-\frac{1}{\alpha_N}} \\ 0 & \text{otherwise} \end{cases}

    Consequently, an emergent ability appears with respect to model scale only when the model size exceeds the critical threshold Ncrit=N0⋅(η−L∞)−1/αNN_{\text{crit}} = N_0 \cdot (\eta - L_\infty)^{-1 / \alpha_N}. Models smaller than NcritN_{\text{crit}} exhibit zero normalized performance regardless of training until the loss crosses η\eta.

  3. Knowl 3 — Invariance of Downstream Performance to Model and Data Size Given Pre-Training Loss

    empirical result

    For a fixed data corpus distribution, tokenization scheme, and model architecture family, the downstream task performance of autoregressive Transformer language models is determined by the pre-training cross-entropy loss, irrespective of the specific trade-off between parameter count and token volume.

    Empirical evaluations across 12 diverse benchmarks (spanning closed-book QA, commonsense reasoning, reading comprehension, coreference resolution, multi-discipline exams, and math word problems in English and Chinese) demonstrate that:

    1. Checkpoints from models of widely varying parameter sizes (from 300M up to 32B parameters) and training token counts (from 33B up to 3T tokens) trace out identical performance-versus-loss curves.
    2. When evaluating open model families (such as LLaMA-7B through LLaMA-65B and Pythia-1.4B through Pythia-12B), downstream performance points across models collapse onto the same shared trajectory when plotted against cross-entropy pre-training loss.
    3. Pre-training loss functions as a unified state indicator for downstream capability, superseding parameter size or token count when analyzed independently.
  4. Knowl 4 — Dual Regimes of Task Scaling and the Universal Emergence Threshold

    empirical result

    Downstream tasks evaluate language models in one of two distinct scaling regimes as pre-training cross-entropy loss decreases:

    1. Smoothly Improving Tasks: Easier or knowledge-retrieval tasks (including TriviaQA, HellaSwag, RACE, WinoGrande, NLPCC-KBQA, ClozeT, CLUEWSC, and C3) exhibit continuous, monotonic performance gains starting immediately from high pre-training loss values.
    2. Emergent Tasks: Complex multi-step reasoning and domain-expert examination tasks (including MMLU, C-Eval, GSM8K, GSM8K-Chinese, and BIG-Bench tasks such as Word Unscramble and Modular Arithmetic) remain stuck at random guessing performance until the pre-training loss reaches a critical threshold of approximately η≈2.2\eta \approx 2.2 (on the 4:1 English-Chinese pre-training corpus). Once the loss drops below 2.22.2, performance begins to scale monotonically with lower loss.

    The loss threshold η≈2.2\eta \approx 2.2 is consistent across multiple emergent tasks despite variations in language (English vs. Chinese), task format (multiple-choice vs. free-form generation), and prompting paradigm (zero-shot, few-shot, few-shot Chain-of-Thought).

  5. Knowl 5 — Persistence of Emergence Under Continuous Evaluation Metrics

    empirical result

    Emergent tipping points in language models are not artifacts of discontinuous or non-linear evaluation metrics (such as multiple-choice 0/1 accuracy). When evaluated on MMLU and C-Eval using continuous metrics, sharp tipping points remain present at the exact same pre-training loss threshold (η≈2.2\eta \approx 2.2):

    1. Correct Choice Probability (CorrectChoiceProb): The raw predicted softmax probability assigned to the ground-truth answer stays flat at approximately 0.250.25 (chance level for 4-option multiple-choice questions) for pre-training losses L>2.2L > 2.2, and only begins rising when L<2.2L < 2.2.
    2. Brier Score: Defined as BrierScore=1N∑i=1N∑j=1C(yij−y^ij)2\text{BrierScore} = \frac{1}{N} \sum_{i=1}^N \sum_{j=1}^C (y_{ij} - \hat{y}_{ij})^2, where y^ij\hat{y}_{ij} is the predicted probability for class jj on sample ii, yij∈{0,1}y_{ij} \in \{0, 1\} is the ground truth, NN is the sample size, and CC is the number of classes. For a 4-choice task (C=4C=4), a uniform random guess predictor assigning y^ij=0.25\hat{y}_{ij} = 0.25 yields a baseline score of 0.750.75. Models with L>2.2L > 2.2 do not achieve Brier scores better than the 0.750.75 random-guess baseline; decreases in Brier score prior to the threshold reflect probability calibration among incorrect options rather than task mastery.
  6. Knowl 6 — Correlation Analysis Between Pre-Training Loss and Downstream Task Performance

    data/table

    Statistical correlation coefficients demonstrate strong monotonic and linear relationships between pre-training loss and downstream performance on smoothly improving tasks, while emergent tasks show weaker linear correlations due to the flat plateau above the emergence loss threshold η≈2.2\eta \approx 2.2.

    Dataset TriQA HS RACE WG NQA ClozeT C3 CW MMLU CE GSM GSMC
    Spearman -0.996 -0.996 -0.977 -0.978 -0.984 -0.986 -0.988 -0.947 -0.804 -0.831 -0.975 -0.948
    Pearson -0.994 -0.994 -0.963 -0.988 -0.982 -0.985 -0.993 -0.972 -0.903 -0.884 -0.874 -0.829

    Dataset abbreviations: TriQA (TriviaQA), HS (HellaSwag), WG (WinoGrande), NQA (NLPCC-KBQA), CW (CLUEWSC), CE (C-Eval), GSM (GSM8K), GSMC (GSM8K-Chinese).

    On smooth tasks (e.g., TriviaQA, HellaSwag, C3), Pearson correlation coefficients are near −1.0-1.0 (−0.963-0.963 to −0.994-0.994), confirming that intermediate checkpoints from 1.5B, 6B, and 32B parameter models all align tightly along the same curve. On emergent tasks (MMLU, C-Eval, GSM8K, GSM8K-Chinese), Pearson correlations drop to between −0.829-0.829 and −0.903-0.903, reflecting the non-linear inflection where performance remains near 0 / random guess until the loss threshold is reached.

  7. Knowl 7 — Predictive Superiority of Pre-Training Loss Over Training Compute

    empirical result

    Pre-training loss provides a unified predictor of downstream task performance across different model architectures and training configurations, whereas total training compute (measured in FLOPs) fails to do so.

    When downstream task performance is plotted against total training compute for models of different scales (1.5B, 6B, and 32B parameters), the trajectories of the different models diverge into distinct, separated curves. A smaller model and a larger model trained with the same total compute achieve substantially different pre-training losses and downstream accuracy levels. Conversely, when plotted against cross-entropy pre-training loss, all data points from all model sizes collapse onto a single coherent performance curve.

  8. Knowl 8 — Experimental Setup for Loss-to-Performance Scaling Analysis

    experimental setup

    The experimental framework for analyzing downstream performance as a function of pre-training loss comprises:

    1. Pre-Training Corpus: A bilingual mixture with a 4:1 English-to-Chinese token ratio consisting of webpages (80.2% CommonCrawl in English), code (10.0%), books (3.8%), Wikipedia (3.8%), papers (1.6%), and StackExchange (0.6%). Over 93.4% of documents are trained for a single epoch without repetition.
    2. Model Architecture: Autoregressive Transformers using Grouped-Query Attention (GQA) and Rotary Position Embedding (RoPE) applied to half the dimensions of query and key vectors. Tokenization uses SentencePiece BPE with a 65,536 vocabulary size and a context sequence length of 2048.
    3. Training Configurations: More than 30 models trained from scratch using Megatron-LM and the AdamW optimizer (β1=0.9,β2=0.95\beta_1 = 0.9, \beta_2 = 0.95) with cosine learning rate schedules decayed to the minimum at the final step. Scale configurations include 1.5B (3T tokens), 6B (3T tokens), and 32B (2.5T tokens), plus 28 smaller models spanning 300M, 540M, 1B, 1.5B, 3B, and 6B trained over durations between 33B and 500B tokens.
    4. Downstream Benchmarks: 12 tasks across English and Chinese evaluated via zero-shot, few-shot, and few-shot Chain-of-Thought (CoT) prompting.
  9. Knowl 9 — Corpus and Tokenizer Dependency of Raw Pre-Training Loss Values

    limitation

    The raw numerical value of pre-training cross-entropy loss is dependent on the specific tokenizer vocabulary size and the pre-training corpus domain distribution. As a consequence, absolute loss values cannot be directly compared across model families that were trained on different text corpora or with different tokenizers. Evaluating cross-model emergence thresholds requires normalizing perplexity to account for vocabulary differences, or evaluating models on a common validation corpus.

Coverage note — None was omitted; all key theoretical definitions, derivations connecting loss thresholds to parameter scaling laws, continuous metric validation experiments, empirical findings across bilingual benchmark suites, and experimental limitations were captured.

References

  1. 1.Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 4895–4901. Association for Computational Linguistics, 2023. URL https://aclanthology.org/2023.emnlp-main.298.
  2. 2.Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  3. 3.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language models across training and scaling. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 2397–2430. PMLR, 2023.
  4. 4.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7432–7439. AAAI Press, 2020. doi: 10.1609/AAAI.V34I05.6239. URL https://doi.org/10.1609/aaai.v34i05.6239.
  5. 5.Glenn W. Brier. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1):1 – 3, 1950. doi: 10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2. URL https://journals.ametsoc.org/view/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  7. 7.Yuan Cao and Quanquan Gu. Generalization error bounds of gradient descent for learning over-parameterized deep relu networks. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 3349–3356. AAAI Press, 2020. doi: 10.1609/AAAI.V34I04.5736. URL https://doi.org/10.1609/aaai.v34i04.5736.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. J. Mach. Learn. Res., 24:240:1–240:113, 2023. URL http://jmlr.org/papers/v24/22-1144.html.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. Scaling instruction-finetuned language models. CoRR, abs/2210.11416, 2022. doi: 10.48550/ARXIV.2210.11416. URL https://doi.org/10.48550/arXiv.2210.11416.
  10. 10.Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc’Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, and Karen Simonyan. Unified scaling laws for routed language models. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato, editors, International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 4057–4086. PMLR, 2022. URL https://proceedings.mlr.press/v162/clark22a.html.
  11. 11.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the AI2 reasoning challenge. CoRR, abs/1803.05457, 2018. URL http://arxiv.org/abs/1803.05457.
  12. 12.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  13. 13.Together Computer. Redpajama: An open source recipe to reproduce llama training dataset, April 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  14. 14.Yehuda Dar, Vidya Muthukumar, and Richard G. Baraniuk. A farewell to the bias-variance trade-off? an overview of the theory of overparameterized machine learning. CoRR, abs/2109.02355, 2021. URL https://arxiv.org/abs/2109.02355.
  15. 15.Nan Duan. Overview of the nlpcc-iccpol 2016 shared task: Open domain chinese question answering. In Natural Language Understanding and Intelligent Applications, pages 942–948. Springer International Publishing, 2016. ISBN 978-3-319-50496-4.
  16. 16.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res., 23:120:1–120:39, 2022. URL http://jmlr.org/papers/v23/21-0998.html.
  17. 17.Daniel Y. Fu, Tri Dao, Khaled Kamal Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. Hungry hungry hippos: Towards language modeling with state space models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=COZDy0WYGg.
  18. 18.Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan, Mitchell Wortsman, Rulin Shao, Jean Mercat, Alex Fang, Jeffrey Li, Sedrick Keh, Rui Xin, Marianna Nezhurina, Igor Vasiljevic, Jenia Jitsev, Alexandros G. Dimakis, Gabriel Ilharco, Shuran Song, Thomas Kollar, Yair Carmon, Achal Dave, Reinhard Heckel, Niklas Muennighoff, and Ludwig Schmidt. Language models scale reliably with over-training and on downstream tasks. CoRR, abs/2403.08540, 2024.
  19. 19.Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom B. Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei, and Jack Clark. Predictability and surprise in large generative models. In FAccT ’22: 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, June 21 - 24, 2022, pages 1747–1764. ACM, 2022. doi: 10.1145/3531146.3533229. URL https://doi.org/10.1145/3531146.3533229.
  20. 20.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. CoRR, abs/2101.00027, 2021. URL https://arxiv.org/abs/2101.00027.
  21. 21.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  22. 22.Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish. Scaling laws for autoregressive generative modeling. CoRR, abs/2010.14701, 2020. URL https://arxiv.org/abs/2010.14701.
  23. 23.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models. CoRR, abs/2203.15556, 2022. doi: 10.48550/ARXIV.2203.15556. URL https://doi.org/10.48550/arXiv.2203.15556.
  24. 24.Shengding Hu, Xin Liu, Xu Han, Xinrong Zhang, Chaoqun He, Weilin Zhao, Yankai Lin, Ning Ding, Zebin Ou, Guoyang Zeng, et al. Predicting emergent abilities with infinite resolution evaluation. arXiv e-prints, pages arXiv–2310, 2023.
  25. 25.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. CoRR, abs/2305.08322, 2023. doi: 10.48550/ARXIV.2305.08322. URL https://doi.org/10.48550/arXiv.2305.08322.
  26. 26.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b. CoRR, abs/2310.06825, 2023. doi: 10.48550/ARXIV.2310.06825. URL https://doi.org/10.48550/arXiv.2310.06825.
  27. 27.Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Regina Barzilay and Min-Yen Kan, editors, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1601–1611. Association for Computational Linguistics, 2017. doi: 10.18653/V1/P17-1147. URL https://doi.org/10.18653/v1/P17-1147.
  28. 28.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. CoRR, abs/2001.08361, 2020. URL https://arxiv.org/abs/2001.08361.
  29. 29.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022.
  30. 30.Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Eduardo Blanco and Wei Lu, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Brussels, Belgium, October 31 - November 4, 2018, pages 66–71. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-2012. URL https://doi.org/10.18653/v1/d18-2012.
  31. 31.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. RACE: large-scale reading comprehension dataset from examinations. In Martha Palmer, Rebecca Hwa, and Sebastian Riedel, editors, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 785–794. Association for Computational Linguistics, 2017. URL https://doi.org/10.18653/v1/d17-1082.
  32. 32.Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. Sophia: A scalable stochastic second-order optimizer for language model pre-training. CoRR, abs/2305.14342, 2023. doi: 10.48550/ARXIV.2305.14342. URL https://doi.org/10.48550/arXiv.2305.14342.
  33. 33.Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma. Same pre-training loss, better downstream: Implicit bias matters for language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 22188–22214. PMLR, 2023. URL https://proceedings.mlr.press/v202/liu23ao.html.
  34. 34.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. The flan collection: Designing data and methods for effective instruction tuning. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 22631–22648. PMLR, 2023. URL https://proceedings.mlr.press/v202/longpre23a.html.
  35. 35.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  36. 36.OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi: 10.48550/ARXIV.2303.08774. URL https://doi.org/10.48550/arXiv.2303.08774.
  37. 37.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. In Waleed Ammar, Annie Louis, and Nasrin Mostafazadeh, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Demonstrations, pages 48–53. Association for Computational Linguistics, 2019. doi: 10.18653/V1/N19-4009. URL https://doi.org/10.18653/v1/n19-4009.
  38. 38.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics, 2016.
  39. 39.Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. Hyena hierarchy: Towards larger convolutional language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 28043–28078. PMLR, 2023. URL https://proceedings.mlr.press/v202/poli23a.html.
  40. 40.Alethea Power, Yuri Burda, Harrison Edwards, Igor Babuschkin, and Vedant Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. CoRR, abs/2201.02177, 2022. URL https://arxiv.org/abs/2201.02177.
  41. 41.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew J. Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446, 2021. URL https://arxiv.org/abs/2112.11446.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  43. 43.Jihyeon Roh, Sang-Hoon Oh, and Soo-Young Lee. Unigram-normalized perplexity as a language model performance measure with different vocabulary sizes. CoRR, abs/2011.13220, 2020. URL https://arxiv.org/abs/2011.13220.
  44. 44.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8732–8740. AAAI Press, 2020. doi: 10.1609/AAAI.V34I05.6399. URL https://doi.org/10.1609/aaai.v34i05.6399.
  45. 45.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  46. 46.Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? CoRR, abs/2304.15004, 2023. doi: 10.48550/ARXIV.2304.15004. URL https://doi.org/10.48550/arXiv.2304.15004.
  47. 47.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics, 2016. doi: 10.18653/V1/P16-1162. URL https://doi.org/10.18653/v1/p16-1162.
  48. 48.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In Jennifer G. Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 4603–4611. PMLR, 2018. URL http://proceedings.mlr.press/v80/shazeer18a.html.
  49. 49.Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woo-Myoung Park, Jung-Woo Ha, and Nako Sung. On the effect of pretraining corpora on in-context learning by a large-scale language model. In Marine Carpuat, Marie-Catherine de Marneffe, and Iván Vladimir Meza Ruíz, editors, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 5168–5186. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.NAACL-MAIN.380. URL https://doi.org/10.18653/v1/2022.naacl-main.380.
  50. 50.C. Spearman. The proof and measurement of association between two things. The American Journal of Psychology, 15(1):72–101, 1904. ISSN 00029556.
  51. 51.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Santilli, Andreas Stuhlmüller, Andrew M. Dai, Andrew La, Andrew K. Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakas, and et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. CoRR, abs/2206.04615, 2022. doi: 10.48550/ARXIV.2206.04615. URL https://doi.org/10.48550/arXiv.2206.04615.
  52. 52.Kai Sun, Dian Yu, Dong Yu, and Claire Cardie. Investigating prior knowledge for challenging chinese machine reading comprehension. Trans. Assoc. Comput. Linguistics, 8:141–155, 2020. doi: 10.1162/TACL_A_00305. URL https://doi.org/10.1162/tacl_a_00305.
  53. 53.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pretraining and finetuning transformers. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022.
  54. 54.Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q. Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling? In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 12342–12364. Association for Computational Linguistics, 2023. URL https://aclanthology.org/2023.findings-emnlp.825.
  55. 55.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023. doi: 10.48550/ARXIV.2302.13971. URL https://doi.org/10.48550/arXiv.2302.13971.
  56. 56.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. doi: 10.48550/ARXIV.2307.09288. URL https://doi.org/10.48550/arXiv.2307.09288.
  57. 57.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=gEZrGCozdqR.
  58. 58.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. Trans. Mach. Learn. Res., 2022, 2022. URL https://openreview.net/forum?id=yzkSU5zdwD.
  59. 59.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html.
  60. 60.Johannes Welbl, Nelson F. Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. In Leon Derczynski, Wei Xu, Alan Ritter, and Tim Baldwin, editors, Proceedings of the 3rd Workshop on Noisy User-generated Text, NUT@EMNLP 2017, Copenhagen, Denmark, September 7, 2017, pages 94–106. Association for Computational Linguistics, 2017.
  61. 61.Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Veselin Stoyanov. Training trajectories of language models across scales. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 13711–13738. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG.767. URL https://doi.org/10.18653/v1/2023.acl-long.767.
  62. 62.Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. CLUE: A chinese language understanding evaluation benchmark. In Donia Scott, Núria Bel, and Chengqing Zong, editors, Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 4762–4772. International Committee on Computational Linguistics, 2020. URL https://doi.org/10.18653/v1/2020.coling-main.419.
  63. 63.Yuan Yao, Qingxiu Dong, Jian Guan, Boxi Cao, Zhengyan Zhang, Chaojun Xiao, Xiaozhi Wang, Fanchao Qi, Junwei Bao, Jinran Nie, Zheni Zeng, Yuxian Gu, Kun Zhou, Xuancheng Huang, Wenhao Li, Shuhuai Ren, Jinliang Lu, Chengqiang Xu, Huadong Wang, Guoyang Zeng, Zile Zhou, Jiajun Zhang, Juanzi Li, Minlie Huang, Rui Yan, Xiaodong He, Xiaojun Wan, Xin Zhao, Xu Sun, Yang Liu, Zhiyuan Liu, Xianpei Han, Erhong Yang, Zhifang Sui, and Maosong Sun. CUGE: A chinese language understanding and generation evaluation benchmark. CoRR, abs/2112.13610, 2021. URL https://arxiv.org/abs/2112.13610.
  64. 64.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4791–4800. Association for Computational Linguistics, 2019. URL https://doi.org/10.18653/v1/p19-1472.
  65. 65.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. GLM-130B: an open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=-Aw0rrrPUF.
  66. 66.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068, 2022. doi: 10.48550/ARXIV.2205.01068. URL https://doi.org/10.48550/arXiv.2205.01068.

Citation

MLA
Du, Z., et al. “Understanding Emergent Abilities of Language Models from the Loss Perspective”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 53138–67, https://proceedings.neurips.cc/paper_files/paper/2024/file/5f1eee2509599faeeb3570a887016a64-Paper-Conference.pdf.
APA
Du, Z., Zeng, A., Dong, Y., & Tang, J. (2024). Understanding Emergent Abilities of Language Models from the Loss Perspective. Advances in Neural Information Processing Systems, 37, 53138–53167. https://proceedings.neurips.cc/paper_files/paper/2024/file/5f1eee2509599faeeb3570a887016a64-Paper-Conference.pdf
Chicago
Du, Z., A. Zeng, Y. Dong, and J. Tang. 2024. “Understanding Emergent Abilities of Language Models from the Loss Perspective”. Advances in Neural Information Processing Systems 37: 53138–67. https://proceedings.neurips.cc/paper_files/paper/2024/file/5f1eee2509599faeeb3570a887016a64-Paper-Conference.pdf.
Harvard
Du, Z. et al. (2024) “Understanding Emergent Abilities of Language Models from the Loss Perspective”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 53138–53167. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/5f1eee2509599faeeb3570a887016a64-Paper-Conference.pdf.
Vancouver
1. Du Z, Zeng A, Dong Y, Tang J (2024) Understanding Emergent Abilities of Language Models from the Loss Perspective. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 53138–53167

BibTeX

@inproceedings{du2024understanding,
  title = {Understanding Emergent Abilities of Language Models from the Loss Perspective},
  author = {Du, Zhengxiao and Zeng, Aohan and Dong, Yuxiao and Tang, Jie},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {53138-53167},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/5f1eee2509599faeeb3570a887016a64-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors