Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Stella BidermanHailey SchoelkopfQuentin AnthonyHerbie BradleyKyle O'BrienEric HallahanMohammad Aflah KhanShivanshu PurohitUSVSN Sai PrashanthEdward Raff

article2023ICML2,122 citationsOutstanding Paper

Introduces Pythia, an open-source suite of 16 language models spanning 70M to 12B parameters with 154 intermediate checkpoints per model and identical data orderings, providing a controlled setup to systematically study training dynamics, scaling behaviors, and memorization.

Listen

Large language models (LLMs) have achieved notable commercial success across numerous domains, yet understanding how these models learn over time and how behaviors change with scale remains a critical challenge. Existing public model suites largely fail to provide the experimental controls necessary to study training dynamics scientifically. They often use private data, obscure the exact sequence in which training examples were presented, or lack intermediate snapshots. To address this barrier, the article introduces Pythia, an open-access suite of 16 autoregressive language models designed specifically to study LLM learning dynamics and scaling behavior under rigorous, reproducible conditions.

The research evaluates how model capabilities, memorization, and biases develop across training time and parameter scales. The approach spans eight model sizes ranging from 70 million to 12 billion parameters, trained for approximately 300 billion tokens. Pythia includes two identical sets of eight models: one trained on the standard public English Pile dataset and another trained on a deduplicated version. Crucially, all models processed data in the exact same sequence. The researchers saved 154 intermediate checkpoints per model along with tools to reconstruct the exact training batches, providing unprecedented visibility into the learning trajectory across scales.

The investigation produced four primary findings. First, verbatim memorization of training text follows a uniform Poisson point process across training steps rather than clustering at the start or finish, showing that the position of data in the training sequence does not affect memorization risk. Second, targeted data interventions during late-stage pretraining—such as swapping masculine pronouns for feminine equivalents in the final 7% to 21% of training—significantly reduced gender bias without degrading general model capabilities. Third, a distinct phase change occurs roughly 45% through training (after 65,000 steps), where larger models (2.8 billion parameters and above) begin to exhibit a strong correlation between downstream task accuracy and the frequency of relevant terms in the pretraining corpus, an emergent property largely absent in sub-billion-parameter models. Finally, data deduplication did not noticeably improve downstream NLP benchmark performance compared to training on standard data.

These findings have direct implications for AI development, risk management, and training cost. Organizations cannot mitigate data privacy or memorization risks simply by rearranging when sensitive data appears during training. However, teams can cost-effectively reduce social biases by modifying data distributions late in pretraining rather than retraining models from scratch. Furthermore, organizations aiming to teach models long-tail factual knowledge must recognize that small models fail to acquire low-frequency information, requiring larger capacities and adequate term frequencies to retain key facts.

Practitioners are recommended to leverage checkpoint-level analysis and pretraining data term-frequency counts to forecast whether target knowledge or undesirable behaviors will emerge. Teams concerned with sensitive data should place monitored samples early in the training pipeline to detect memorization risks well before completing expensive training runs. While Pythia's findings offer high confidence for English autoregressive transformers, readers should exercise caution when applying these conclusions to multilingual domains, given that the suite is strictly English-focused and evaluates specific bias benchmarks that carry inherent measurement limitations.

  • Paper: Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao et al. (2023). This work leverages the Pythia model suite to analyze attention score distributions across sequence lengths and uncover the attention sink phenomenon.
  • Paper: Matryoshka Language Model Suites, Nathan Godey et al. (2026). This research builds directly on the paradigm of multi-scale model suites by introducing nested architectures that train multiple model sizes simultaneously.
Cover for Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Abstract

How do large language models (LLMs) develop and evolve over the course of training? How do these patterns change as models scale? To answer these questions, we introduce \textit{Pythia}, a suite of 16 LLMs all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. We provide public access to 154 checkpoints for each one of the 16 models, alongside tools to download and reconstruct their exact training dataloaders for further study. We intend \textit{Pythia} to facilitate research in many areas, and we present several case studies including novel results in memorization, term frequency effects on few-shot performance, and reducing gender bias. We demonstrate that this highly controlled setup can be used to yield novel insights toward LLMs and their training dynamics. Trained models, analysis code, training code, and training data can be found at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 The Pythia Suite
  • 2.1 Requirements for a Scientific Suite of LLMs
  • 2.2 Training Data
  • 2.3 Architecture
  • 2.4 Training
  • 2.5 Evaluation
  • 2.6 Novel Observations in Evaluation
  • 2.7 Public Release and Reproducibility
  • 3 Case Studies
  • 3.1 How Does Data Bias Influence Learned Behaviors?
  • 3.2 Does Training Order Influence Memorization?
  • 3.3 Do Pretraining Term Frequencies Influence Task Performance Throughout Training?
  • 4 Conclusion
  • References
  • A Author Contributions
  • B Corrections and Updates
  • C Additional Plots for Case Studies
  • C.1 Gender Bias Interventions
  • C.2 Pretraining Term Frequency
  • D Training Hardware and GPU hours
  • E Full Configuration Details
  • F Additional Details on Design and Considerations
  • F.1 Assessment of Existing Suites
  • F.2 Contrast with Multiply Trained Models
  • G Evaluations
  • G.1 Raw Benchmark Scores - Zero Shot
  • G.2 Raw Benchmark Scores - Five Shot
  • G.3 Comparison to Baseline Models
  • G.4 Graphs

Knowls

  1. Knowl 1 — Pythia Model Suite Architecture and Design

    model/method

    The Pythia suite consists of 16 decoder-only autoregressive language models spanning 8 parameter scales (70M, 160M, 410M, 1.0B, 1.4B, 2.8B, 6.9B, and 12B total parameters), trained on both the standard Pile dataset and a deduplicated version of the Pile.

    The models incorporate specific architectural deviations from standard GPT-3 configurations:

    • Dense Attention Layers: All attention layers are fully dense rather than alternating between sparse and dense.
    • Flash Attention: Fused attention kernels are utilized during training for hardware throughput.
    • Rotary Position Embeddings: Rotary embeddings (RoPE) are applied to 25% of the hidden dimensions (rotary-pct = 0.25).
    • Parallel Attention and Feed-Forward: Attention and multi-layer perceptron (MLP) sublayers are computed in parallel rather than sequentially, following the formulation: xl+1=xl+Attention(LN(xl))+MLP(LN(xl))x_{l+1} = x_l + \text{Attention}(\text{LN}(x_l)) + \text{MLP}(\text{LN}(x_l)) where LN\text{LN} is Layer Normalization and xlx_l is the representation at layer ll.
    • Untied Embeddings: Input token embeddings and output unembedding matrices are parameterized separately without weight tying.
    • Uniform Batch Size: A uniform global batch size of 1,024 sequences with context length 2,048 (2,097,1522,097,152 tokens per batch) is applied across all model sizes.
  2. Knowl 2 — Pythia Suite Architecture Hyperparameters and Model Scaling

    data/table

    The Pythia model suite defines 8 distinct model sizes parameterized across layers, model dimensions, attention heads, and initial learning rates.

    Model Size Non-Embedding Params Layers Model Dim Heads Learning Rate Equivalent Models
    70 M 18,915,328 6 512 8 10.0×10−410.0 \times 10^{-4} —
    160 M 85,056,000 12 768 12 6.0×10−46.0 \times 10^{-4} GPT-Neo 125M, OPT-125M
    410 M 302,311,424 24 1024 16 3.0×10−43.0 \times 10^{-4} OPT-350M
    1.0 B 805,736,448 16 2048 8 3.0×10−43.0 \times 10^{-4} —
    1.4 B 1,208,602,624 24 2048 16 2.0×10−42.0 \times 10^{-4} GPT-Neo 1.3B, OPT-1.3B
    2.8 B 2,517,652,480 32 2560 32 1.6×10−41.6 \times 10^{-4} GPT-Neo 2.7B, OPT-2.7B
    6.9 B 6,444,163,072 32 4096 32 1.2×10−41.2 \times 10^{-4} OPT-6.7B
    12 B 11,327,027,200 36 5120 40 1.2×10−41.2 \times 10^{-4} —

    Models are named according to total parameter count including input/output embedding matrices, while non-embedding parameters serve as the primary metric for model capacity comparisons.

  3. Knowl 3 — Independence of Sequence Memorization from Training Order

    empirical result

    In large language models, a sequence is defined as (k,ℓ)(k, \ell)-memorized if prompting the model with a prefix of length kk tokens from the training data induces the model to generate the exact continuation of length ℓ\ell tokens from the training corpus. When evaluating the occurrence of (32,32)(32, 32)-memorized sequences on the first 64 tokens of every context across training, the distribution of memorized sequences across mini-batches adheres to a stationary Poisson Point Process.

    Quantile-Quantile (Q-Q) plots comparing sample distributions against theoretical Poisson distributions confirm that the rate of occurrence of memorized sequences is uniform across training time. Sequence position within the training order (whether early or late in pretraining) has negligible influence on the probability that the sequence is verbatim memorized.

  4. Knowl 4 — Emergence of Pretraining Term Frequency Dependency in Few-Shot Tasks

    empirical result

    In autoregressive language models, the correlation between downstream task accuracy and the frequency of task-relevant terms/entities in the pretraining data is an emergent property dependent on both model scale and training progression:

    1. Model Scale Threshold: Models with fewer than 1.0B1.0\text{B} parameters fail to show systematic correlation between task accuracy and pretraining frequency across few-shot arithmetic (addition and multiplication prompts Q:x1  #  x2?  A:Q: x_1 \; \# \; x_2? \; A:) and TriviaQA fact retrieval, achieving near-zero or plateaued accuracy regardless of digit or entity occurrence counts.
    2. Phase Change During Pretraining: For models of size 2.8B2.8\text{B} and 12B12\text{B} parameters, a distinct phase change occurs after approximately 65,00065,000 training steps (45%45\% through training, corresponding to ≈136B\approx 136\text{B} tokens seen). Beyond this point, a positive correlation between entity/operand term frequency in the seen pretraining data and downstream task accuracy emerges.
    3. Widening Accuracy Gap: In arithmetic multiplication, the performance gap between the top 10%10\% most frequent input operands and bottom 10%10\% least frequent operands systematically widens over training (reaching up to a 75.6%75.6\% accuracy discrepancy at step 143,000143,000 with 4 shots on the 12B model).
  5. Knowl 5 — Gender Bias Mitigation via Counterfactual Pretraining Interventions

    empirical result

    Counterfactually retraining language models on modified corpus statistics late in pretraining systematically alters gender bias metrics without significantly harming downstream language modeling performance:

    1. Intervention Setup: Pretraining was branched at 21B21\text{B} tokens (7%7\%) and 63B63\text{B} tokens (21%21\%) before training completion, resuming training on identical sequences with morphologically masculine pronouns deterministically substituted with their feminine counterparts.
    2. Bias Reduction and Reversal: On WinoBias (measured via prompt-based multiple-choice pronoun completion) and CrowS-Pairs, the intervention decreases stereotypical accuracy across all scales (70M70\text{M}, 410M410\text{M}, 1.4B1.4\text{B}, and 6.9B6.9\text{B}). For the 6.9B6.9\text{B} parameter model, the intervention reduces stereotypical choice proportion below 0.50.5, shifting the model from pro-stereotypical to anti-stereotypical bias.
    3. Scaling of Effect Size: Larger models exhibit higher initial pro-stereotypical bias but achieve larger magnitude bias reductions from the counterfactual intervention.
    4. Task Preservation: Zero-shot accuracy and perplexity on the LAMBADA language modeling benchmark exhibit only negligible differences between baseline and intervened models.
  6. Knowl 6 — Feasibility of Uniform Large Batch Sizes for Small Language Models

    empirical result

    Contrary to the conventional paradigm that language models below 1B1\text{B} parameters require smaller batch sizes to prevent convergence issues, models from 70M70\text{M} to 12B12\text{B} parameters can be stably trained with a uniform batch size of 2,097,1522,097,152 tokens (1,0241,024 sequences of 2,0482,048 tokens).

    Small models trained at 4×4\times to 8×8\times standard batch sizes exhibit no convergence degradation and achieve downstream zero-shot accuracy matching equi-parameter OPT baseline models on standard NLP benchmarks. Scaling batch size enables linear data-parallel GPU scaling in GPT-NeoX, yielding up to 8×8\times wall-clock training speedups on small architectures (70M70\text{M} utilizing 32 A100 GPUs compared to standard 4 GPUs).

  7. Knowl 7 — Checkpoint Sampling Schedule and Dataloader Reconstruction

    experimental setup

    Each model in the Pythia suite is checkpointed 154 times across a 300B300\text{B} token budget (143,000143,000 iterations with 2,097,1522,097,152 tokens per iteration):

    • Initialization Checkpoint: Saved at iteration 0 (random initialization).
    • Log-Spaced Early Checkpoints: Saved at iterations {1,2,4,8,16,32,64,128,256,512}\{1, 2, 4, 8, 16, 32, 64, 128, 256, 512\} to capture early training dynamics.
    • Linearly Spaced Checkpoints: Saved every 1,0001,000 iterations (2,097,152,0002,097,152,000 tokens) from step 1,0001,000 to 143,000143,000.

    Exact dataloader reproducibility tools enable researchers to reconstruct the exact token order and batch contents processed by each model up to any given checkpoint.

  8. Knowl 8 — Zero-Shot Benchmark Performance of the Fully-Trained Pythia Suite

    data/table

    The table below presents the zero-shot accuracy (with standard error) of the fully-trained Pythia suite across 8 standard evaluation benchmarks evaluated via the Language Model Evaluation Harness.

    Task 70M 160M 410M 1B 1.4B 2.8B 6.9B 12B
    Lambada 0.185±0.0050.185 \pm 0.005 0.328±0.0070.328 \pm 0.007 0.516±0.0070.516 \pm 0.007 0.562±0.0070.562 \pm 0.007 0.616±0.0070.616 \pm 0.007 0.647±0.0070.647 \pm 0.007 0.673±0.0070.673 \pm 0.007 0.705±0.0060.705 \pm 0.006
    PIQA 0.595±0.0110.595 \pm 0.011 0.627±0.0110.627 \pm 0.011 0.668±0.0110.668 \pm 0.011 0.707±0.0110.707 \pm 0.011 0.711±0.0110.711 \pm 0.011 0.739±0.0100.739 \pm 0.010 0.752±0.0100.752 \pm 0.010 0.760±0.0100.760 \pm 0.010
    WinoGrande 0.528±0.0140.528 \pm 0.014 0.531±0.0140.531 \pm 0.014 0.537±0.0140.537 \pm 0.014 0.537±0.0140.537 \pm 0.014 0.573±0.0140.573 \pm 0.014 0.594±0.0140.594 \pm 0.014 0.609±0.0140.609 \pm 0.014 0.639±0.0130.639 \pm 0.013
    WSC 0.365±0.0470.365 \pm 0.047 0.365±0.0470.365 \pm 0.047 0.567±0.0490.567 \pm 0.049 0.365±0.0470.365 \pm 0.047 0.365±0.0470.365 \pm 0.047 0.385±0.0480.385 \pm 0.048 0.365±0.0470.365 \pm 0.047 0.548±0.0490.548 \pm 0.049
    ARC-Easy 0.374±0.0100.374 \pm 0.010 0.435±0.0100.435 \pm 0.010 0.521±0.0100.521 \pm 0.010 0.569±0.0100.569 \pm 0.010 0.606±0.0100.606 \pm 0.010 0.644±0.0100.644 \pm 0.010 0.673±0.0100.673 \pm 0.010 0.702±0.0090.702 \pm 0.009
    ARC-Chal. 0.181±0.0110.181 \pm 0.011 0.188±0.0110.188 \pm 0.011 0.213±0.0120.213 \pm 0.012 0.244±0.0130.244 \pm 0.013 0.260±0.0130.260 \pm 0.013 0.295±0.0130.295 \pm 0.013 0.313±0.0140.313 \pm 0.014 0.318±0.0140.318 \pm 0.014
    SciQ 0.601±0.0150.601 \pm 0.015 0.741±0.0140.741 \pm 0.014 0.811±0.0120.811 \pm 0.012 0.840±0.0120.840 \pm 0.012 0.865±0.0110.865 \pm 0.011 0.882±0.0100.882 \pm 0.010 0.897±0.0100.897 \pm 0.010 0.902±0.0090.902 \pm 0.009
    LogiQA 0.210±0.0160.210 \pm 0.016 0.190±0.0150.190 \pm 0.015 0.220±0.0160.220 \pm 0.016 0.223±0.0160.223 \pm 0.016 0.210±0.0160.210 \pm 0.016 0.212±0.0160.212 \pm 0.016 0.253±0.0170.253 \pm 0.017 0.224±0.0160.224 \pm 0.016

    Performance scales smoothly with parameter count across tasks and matches corresponding OPT baseline configurations.

  9. Knowl 9 — Empirical Invariance of Downstream Performance to Pile Deduplication and Multi-Epoch Training

    empirical result

    Comparing identical architectures trained on the standard Pile (300B300\text{B} tokens seen out of 334B334\text{B}) versus near-deduplicated Pile (207B207\text{B} tokens, deduplicated via MinHashLSH with threshold 0.87 and trained for ≈1.5\approx 1.5 epochs to reach 300B300\text{B} tokens) reveals two findings:

    1. No Net Advantage from Deduplication: Across zero-shot and 5-shot evaluations on 8 standard NLP benchmarks (LAMBADA, PIQA, WinoGrande, WSC, ARC-Easy, ARC-Challenge, SciQ, LogiQA), deduplication provided no consistent accuracy improvements over the non-deduplicated baseline.
    2. No Degradation from Second-Epoch Repetition: Comparing evaluation scores at step 93,000 (immediately before the deduplicated model finished epoch 1 at 207B tokens) to final checkpoints at step 143,000 (approx1.5\\approx 1.5 epochs) shows that repeated token exposure during the second epoch produced no adverse effects on evaluation scores.
  10. Knowl 10 — Pythia Training Optimization and Parallelism Configuration

    experimental setup

    All Pythia models were trained in FP16 mixed precision using GPT-NeoX and DeepSpeed with the following optimization hyperparameters:

    • Optimizer: Adam with β1=0.9,β2=0.95,ϵ=10−8\beta_1 = 0.9, \beta_2 = 0.95, \epsilon = 10^{-8}, weight decay of 0.010.01, and gradient clipping threshold 1.01.0.
    • Learning Rate Schedule: 1% warm-up (1,4301,430 steps) followed by cosine learning rate decay over 143,000143,000 steps down to a minimum of 0.1×LRmax⁡0.1 \times \text{LR}_{\max}.
    • Distributed Parallelism: ZeRO Stage 1 (optimizer state partitioning with contiguous gradients, overlap communication, and reduce scatter enabled).
    • Tensor Parallelism: Model parallel size 1 for models up to 2.8B2.8\text{B}; model parallel size 2 for 6.9B6.9\text{B}; model parallel size 4 for 12B12\text{B}.
    • Bucket Sizes: ZeRO allgather and reduce bucket sizes set to 500M500\text{M} for models up to 2.8B2.8\text{B}, and 1.26B1.26\text{B} for 6.9B6.9\text{B} and 12B12\text{B}.

Coverage note — Omitted peripheral discussions such as the author contribution list, individual hardware GPU counts/hours per model size from Table 5, the evaluation and critique of the curse of multilinguality on BLOOM (which concerns external models), and minor details regarding the obsolete Pythia v0 preliminary release nomenclature.

References

  1. 1.Ahdritz, G., Bouatta, N., Kadyan, S., Xia, Q., Gerecke, W., O’Donnell, T. J., Berenberg, D., Fisk, I., Zanichelli, N., Zhang, B., et al. Openfold: Retraining alphafold2 yields new insights into its learning mechanisms and capacity for generalization. bioRxiv, 2022.
  2. 2.Andonian, A., Anthony, Q., Biderman, S., Black, S., Gali, P., Gao, L., Hallahan, E., Levy-Kramer, J., Leahy, C., Nestler, L., Parker, K., Pieler, M., Purohit, S., Songz, T., Phil, W., and Weinbach, S. GPT-NeoX: Large scale autoregressive language modeling in PyTorch, 8 2021. URL https://www.github.com/eleutherai/gpt-neox.
  3. 3.Bach, S. H., Sanh, V., Yong, Z.-X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-David, S., Xu, C., Chhablani, G., Wang, H., Fries, J. A., Al-shaibani, M. S., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Tang, X., Jiang, M. T.-J., and Rush, A. M. Promptsource: An integrated development environment and repository for natural language prompts, 2022.
  4. 4.Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023.
  5. 5.Biderman, S., Bicheno, K., and Gao, L. Datasheet for the Pile. Computing Research Repository, 2022. doi: 10.48550/arXiv.2201.07311. URL https://arxiv.org/abs/2201.07311v1. Version 1.
  6. 6.Biderman, S., Prashanth, U. S., Sutawika, L., Purohit, S., Schoelkopf, H., Anthony, Q., and Raff, E. Emergent and predictable memorization in large language models. Preprint under review, 2023.
  7. 7.Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., and Bao, M. The values encoded in machine learning research. Computing Research Repository, 2021. doi: 10.48550/arXiv.2106.15590. URL https://arxiv.org/abs/2106.15590v2. Version 2.
  8. 8.Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large scale autoregressive language modeling with Mesh-TensorFlow. GitHub, 2021. URL https://www.github.com/eleutherai/gpt-neo.
  9. 9.Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5–Workshop on Challenges & Perspectives in Creating Large Language Models, pp. 95–136, 2022.
  10. 10.Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1004–1015, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.81. URL https://aclanthology.org/2021.acl-long.81.
  11. 11.Bordia, S. and Bowman, S. Identifying and reducing gender bias in word-level language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, pp. 7–15, 2019.
  12. 12.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901, 2020.
  13. 13.Brunet, M.-E., Alkalay-Houlihan, C., Anderson, A., and Zemel, R. Understanding the origins of bias in word embeddings. In International conference on machine learning, pp. 803–811. PMLR, 2019.
  14. 14.Carlini, N., Liu, C., Erlingsson, U., Kos, J., and Song, ´D. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium (USENIX Security 19), pp. 267–284, 2019.
  15. 15.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650, 2021.
  16. 16.Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  17. 17.Caswell, I., Breiner, T., van Esch, D., and Bapna, A. Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Proceedings of the 28th International Conference on Computational Linguistics, pp. 6588–6608, 2020.
  18. 18.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. Computing Research Repository, 2021. doi: 10.48550/arXiv.2107.03374. URL https://arxiv.org/abs/2107.03374v2. Version 2.
  19. 19.Choenni, R., Shutova, E., and van Rooij, R. Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1477–1491, 2021.
  20. 20.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. PaLM: Scaling language modeling with Pathways. Computing Research Repository, 2022. doi: 10.48550/arXiv.2204.0231. URL https://arxiv.org/abs/2204.02311v5. Version 5.
  21. 21.Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, Guillaume andGuzman, F., Grave, E., Ott, ´Myle andZettlemoyer, L., and Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8440–8451, Online, 07 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.747. URL https://aclanthology.org/2020.acl-main.747.
  22. 22.Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L., and Raff, E. VQGAN-CLIP: Open domain image generation and editing with natural language guidance. Computing Research Repository, 2022. doi: 10.48550/arXiv.2204.08583. URL https://arxiv.org/abs/2204.08583v2. Version 2.
  23. 23.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. arXiv preprint arXiv:2205.14135, 2022.
  24. 24.Dodge, J., Sap, M., Marasovic, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1286–1305, 2021.
  25. 25.D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al. Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research, 2020.
  26. 26.Elazar, Y., Kassner, N., Ravfogel, S., Feder, A., Ravichander, A., Mosbach, M., Belinkov, Y., Schutze, H., and Goldberg, Y. Measuring causal effects of data statistics on language model’sfactual’predictions. arXiv preprint arXiv:2207.14251, 2022.
  27. 27.Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W.-t., Zettlemoyer, L., and Lewis, M. InCoder: A generative model for code infilling and synthesis. Computing Research Repository, 2022. doi: 10.48550/arXiv.2204.05999. URL https://arxiv.org/abs/2204.05999v2. Version 2.
  28. 28.Gao, L. On the sizes of openai api models. EleutherAI Blog, 2021.
  29. 29.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. The Pile: An 800GB dataset of diverse text for language modeling. Computing Research Repository, 2020. doi: 10.48550/arXiv.2101.00027. URL https://arxiv.org/abs/2101.00027v1. Version 1.
  30. 30.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation. September 2021. doi: 10.5281/zenodo.5371628. URL https://doi.org/10.5281/zenodo.5371628.
  31. 31.Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C. Scaling laws for neural machine translation. Computing Research Repository, 2021. doi: 10.48550/arXiv.2109.07740. URL https://arxiv.org/abs/2109.07740v1. Version 1.
  32. 32.Ginart, A., van der Maaten, L., Zou, J., and Guo, C. Submix: Practical private prediction for large-scale language models. arXiv preprint arXiv:2201.00971, 2022.
  33. 33.Gira, M., Zhang, R., and Lee, K. Debiasing pre-trained language models via efficient fine-tuning. In Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion, pp. 59–69, 2022.
  34. 34.Goyal, P., Dollar, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. Accurate, large minibatch SGD: training imagenet in 1 hour. CoRR, abs/1706.02677, 2017. URL http://arxiv.org/abs/1706.02677.
  35. 35.Hall, M., van der Maaten, L., Gustafson, L., and Adcock, A. A systematic study of bias amplification. arXiv preprint arXiv:2201.11706, 2022.
  36. 36.Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., Hallacy, C., Mann, B., Radford, A., Ramesh, A., Ryder, N., Ziegler, D. M., Schulman, J., Amodei, D., and McCandlish, S. Scaling laws for autoregressive generative modeling. Computing Research Repository, 2020. doi: 10.48550/arXiv.2010.14701. URL https://arxiv.org/abs/2010.14701v2. Version 2.
  37. 37.Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S. Scaling laws for transfer. Computing Research Repository, 2021. doi: 10.48550/arXiv.2102.01293. URL https://arxiv.org/abs/2102.01293v1. Version 1.
  38. 38.Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., Johnston, S., Mann, B., Olah, C., Olsson, C., Amodei, D., Joseph, N., Kaplan, J., and McCandlish, S. Scaling laws and interpretability of learning from repeated data. Computing Research Repository, 05 2022. doi: 10.48550/arXiv.2205.10487. URL https://arxiv.org/abs/2205.10487v1. Version 1.
  39. 39.Hirota, Y., Nakashima, Y., and Garcia, N. Quantifying societal bias amplification in image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13450–13459, 2022.
  40. 40.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  41. 41.Hu, H., Salcic, Z., Sun, L., Dobbie, G., Yu, P. S., and Zhang, X. Membership inference attacks on machine learning: A survey. ACM Computing Surveys (CSUR), 54(11s):1–37, 2022.
  42. 42.Huang, J., Shao, H., and Chang, K. C.-C. Are large pretrained language models leaking your personal information? arXiv preprint arXiv:2205.12628, 2022.
  43. 43.Ippolito, D., Tramer, F., Nasr, M., Zhang, C., Jagielski, M., Lee, K., Choquette-Choo, C. A., and Carlini, N. Preventing verbatim memorization in language models gives a false sense of privacy. arXiv preprint arXiv:2210.17546, 2022.
  44. 44.Jagielski, M., Thakkar, O., Tramer, F., Ippolito, D., Lee, K., Carlini, N., Wallace, E., Song, S., Thakurta, A., Papernot, N., et al. Measuring forgetting of memorized training examples. arXiv preprint arXiv:2207.00099, 2022.
  45. 45.Jernite, Y., Nguyen, H., Biderman, S., Rogers, A., Masoud, M., Danchev, V., Tan, S., Luccioni, A. S., Subramani, N., Johnson, I., et al. Data governance in the age of large-scale data-driven language technology. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 2206–2222, 2022.
  46. 46.Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, 2017.
  47. 47.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zˇ´ıdek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  48. 48.Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C. Large language models struggle to learn long-tail knowledge. arXiv preprint arXiv:2211.08411, 2022.
  49. 49.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. Computing Research Repository, 2020. doi: 10.48550/arXiv.2001.08361. URL https://arxiv.org/abs/2001.08361v1. Version 1.
  50. 50.Karamcheti, S., Orr, L., Bolton, J., Zhang, T., Goel, K., Narayan, A., Bommasani, R., Narayanan, D., Hashimoto, T., Jurafsky, D., Manning, C. D., Potts, C., Re, C., and Liang, P. Mistral - a journey towards reproducible language model training, 2021. URL https://github.com/stanford-crfm/mistral.
  51. 51.Kirtane, N., Manushree, V., and Kane, A. Efficient gender debiasing of pre-trained indic language models. arXiv preprint arXiv:2209.03661, 2022.
  52. 52.Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A., Subramani, N., Sokolov, A., Sikasote, C., Setyawan, M., Sarin, S., Samb, S., Sagot, B., Rivera, C., Rios, A., Papadimitriou, I., Osei, S., Suarez, P. O., Orife, I., Ogueji, K., Rubungo, A. N., Nguyen, T. Q., Muller, M., Müller, A., Muhammad, S. H., Muhammad, N., Mnyakeni, A., Mirzakhalov, J., Matangira, T., Leong, C., Lawson, N., Kudugunta, S., Jernite, Y., Jenny, M., Firat, O., Dossou, B. F. P., Dlamini, S., de Silva, N., Çabuk Ballı, S., Biderman, S., Battisti, A., Baruwa, A., Bapna, A., Baljekar, P., Azime, I. A., Awokoya, A., Ataman, D., Ahia, O., Ahia, O., Agrawal, S., and Adeyemi, M. Quality at a glance: An audit of web-crawled multilingual datasets. Computing Research Repository, 2021. doi: 10.48550/arXiv.2103.12028. URL https://arxiv.org/abs/2103.12028v3. Version 4.
  53. 53.Laurenç​on, H., Saulnier, L., Wang, T., Akiki, C., del Moral, A. V., Scao, T. L., Werra, L. V., Mou, C., Ponferrada, E. G., Nguyen, H., Frohberg, J., Saško, M., Lhoest, Q., McMillan-Major, A., Dupont, G., Biderman, S., Rogers, A., allal, L. B., Toni, F. D., Pistilli, G., Nguyen, O., Nikpoor, S., Masoud, M., Colombo, P., de la Rosa, J., Villegas, P., Thrush, T., Longpre, S., Nagel, S., Weber, L., Munoz, M. R., Zhu, J., Strien, D. V., Alyafeai, Z., Almubarak, K., Chien, V. M., Gonzalez-Dios, I., Soroa, A., Lo, K., Dey, M., Suarez, P. O., Gokaslan, A., Bose, S., Adelani, D. I., Phan, L., Tran, H., Yu, I., Pai, S., Chim, J., Lepercq, V., Ilic, S., Mitchell, M., Luccioni, S., and Jernite, Y. The bigscience ROOTS corpus: A 1.6TB composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview.net/forum?id=UoEw6KigkUn.
  54. 54.Le Scao, T., Wang, T., Hesslow, D., Saulnier, L., Bekman, S., Bari, M. S., Biderman, S., Elsahar, H., Phang, J., Press, O., et al. What language model to train if you have one million GPU hours? In Proceedings of BigScience Episode #5–Workshop on Challenges & Perspectives in Creating Large Language Models, 2022.
  55. 55.Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. In Annual Meeting of the Association for Computational Linguistics, 2021.
  56. 56.Levy, S., Lazar, K., and Stanovsky, G. Collecting a large-scale gender bias dataset for coreference resolution and machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 2470–2480, 2021.
  57. 57.Lieber, O., Sharir, O., Lenz, B., and Shoham, Y. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 2021.
  58. 58.Mallen, A., Asai, A., Zhong, V., Das, R., Hajishirzi, H., and Khashabi, D. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint arXiv:2212.10511, 2022.
  59. 59.McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D. An empirical model of large-batch training. arXiv preprint arXiv:1812.06162, 2018.
  60. 60.McGrath, T., Kapishnikov, A., Tomasev, N., Pearce, A., Hassabis, D., Kim, B., Paquet, U., and Kramnik, V. Acquisition of chess knowledge in AlphaZero. Computing Research Repository, 2021. doi: 10.48550/arXiv.2111.09259. URL https://arxiv.org/abs/2111.09259v3. Version 3.
  61. 61.McMillan-Major, A., Alyafeai, Z., Biderman, S., Chen, K., De Toni, F., Dupont, G., Elsahar, H., Emezue, C., Aji, A. F., Ilic, S., et al. Documenting geographically and contextually diverse data sources: The BigScience catalogue of language data and resources. Computing Research Repository, 2022. doi: 10.48550/arXiv.2201.10066. URL https://arxiv.org/abs/2201.10066v1. Version 1.
  62. 62.Mikami, H., Fukumizu, K., Murai, S., Suzuki, S., Kikuchi, Y., Suzuki, T., Maeda, S.-i., and Hayashi, K. A scaling law for synthetic-to-real transfer: How much is your pre-training effective? Computing Research Repository, 2021. doi: 10.48550/arXiv.2108.11018. URL https://arxiv.org/abs/2108.11018v3. Version 3.
  63. 63.Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1953–1967, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.154. URL https://aclanthology.org/2020.emnlp-main.154.
  64. 64.Nev´eol, A., Dupont, Y., Bezançon, J., and Fort, K. French CrowS-pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8521–8531, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.583. URL https://aclanthology.org/2022.acl-long.583.
  65. 65.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  66. 66.Pfeiffer, J., Goyal, N., Lin, X. V., Li, X., Cross, J., Riedel, S., and Artetxe, M. Lifting the curse of multilinguality by pre-training modular transformers. Computing Research Repository, 05 2022. doi: 10.48550/arXiv.2205.06266. URL https://arxiv.org/abs/2205.06266v1. Version 1.
  67. 67.Phang, J., Bradley, H., Gao, L., Castricato, L., and Biderman, S. Eleutherai: Going beyond” open science” to” science in the open”. arXiv preprint arXiv:2210.06413, 2022.
  68. 68.Pu, J., Yang, Y., Li, R., Elibol, O., and Droppo, J. Scaling effect of self-supervised speech models. Proc. Interspeech 2021, pp. 1084–1088, 2021.
  69. 69.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog, 2019. URL https://openai.com/blog/better-language-models/.
  70. 70.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 01 2020. ISSN 1532-4435. URL http://jmlr.org/papers/v21/20-074.html.
  71. 71.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. ZeRO: Memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’20. IEEE Press, 2020. ISBN 9781728199986. doi: 10.5555/3433701.3433727. URL https://dl.acm.org/doi/10.5555/3433701.3433727.
  72. 72.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  73. 73.Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022.
  74. 74.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  75. 75.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Biderman, S., Gao, L., Bers, T., Wolf, T., and Rush, A. M. Multitask prompted training enables zero-shot task generalization. Computing Research Repository, 2021. doi: 10.48550/arXiv.2110.08207. URL https://arxiv.org/abs/2110.08207v3. Version 3.
  76. 76.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagne, R., Luccioni, A. S., Yvon, F., Gallé, M., Tow, J., Rush, A. M., Biderman, S., Webson, A., Ammanamanchi, P. S., Wang, T., Sagot, B., Muennighoff, N., del Moral, A. V., Ruwase, O., Bawden, R., Bekman, S., McMillan-Major, A., Beltagy, I., Nguyen, H., Saulnier, L., Tan, S., Suarez, P. O., Sanh, V., Laurençon, H., Jernite, Y., Launay, J., Mitchell, M., Raffel, C., Gokaslan, A., Simhi, A., Soroa, A., Aji, A. F., Alfassy, A., Rogers, A., Nitzav, A. K., Xu, C., Mou, C., Emezue, C., Klamm, C., Leong, C., van Strien, D., Adelani, D. I., Radev, D., Ponferrada, E. G., Levkovizh, E., Kim, E., Natan, E. B., De Toni, F., Dupont, G., Kruszewski, G., Pistilli, G., Elsahar, H., Benyamina, H., Tran, H., Yu, I., Abdulmumin, I., Johnson, I., Gonzalez-Dios, I., de la Rosa, J., Chim, J., Dodge, J., Zhu, J., Chang, J., Frohberg, J., Tobing, J., Bhattacharjee, J., Almubarak, K., Chen, K., Lo, K., Von Werra, L., Weber, L., Phan, L., allal, L. B., Tanguy, L., Dey, M., Munoz, M. R., Masoud, M., Grandury, M., Saško, M., Huang, M., Coavoux, M., Singh, M., Jiang, M. T.-J., Vu, M. C., Jauhar, M. A., Ghaleb, M., Subramani, N., Kassner, N., Khamis, N., Nguyen, O., Espejel, O., de Gibert, O., Villegas, P., Henderson, P., Colombo, P., Amuok, P., Lhoest, Q., Harliman, R., Bommasani, R., Lopez, R. L., Ribeiro, R., Osei, S., Pyysalo, S., Nagel, S., Bose, S., Muhammad, S. H., Sharma, S., Longpre, S., Nikpoor, S., Silberberg, S., Pai, S., Zink, S., Torrent, T. T., Schick, T., Thrush, T., Danchev, V., Nikoulina, V., Laippala, V., Lepercq, V., Prabhu, V., Alyafeai, Z., Talat, Z., Raja, A., Heinzerling, B., Si, C., Taşar, D. E., Salesky, E., Mielke, S. J., Lee, W. Y., Sharma, A., Santilli, A., Chaffin, A., Stiegler, A., Datta, D., Szczechla, E., Chhablani, G., Wang, H., Pandey, H., Strobelt, H., Fries, J. A., Rozen, J., Gao, L., Sutawika, L., Bari, M. S., Al-shaibani, M. S., Manica, M., Nayak, N., Teehan, R., Albanie, S., Shen, S., Ben-David, S., Bach, S. H., Kim, T., Bers, T., Fevry, T., Neeraj, T., Thakker, U., Raunak, V., Tang, X., Yong, Z.-X., Sun, Z., Brody, S., Uri, Y., Tojarieh, H., Roberts, A., Chung, H. W., Tae, J., Phang, J., Press, O., Li, C., Narayanan, D., Bourfoune, H., Casper, J., Rasley, J., Ryabinin, M., Mishra, M., Zhang, M., Shoeybi, M., Peyrounette, M., Patry, N., Tazi, N., Sanseviero, O., von Platen, P., Cornette, P., Lavallee, P. F., Lacroix, R., Rajbhandari, S., Gandhi, S., Smith, S., Requena, S., Patil, S., Dettmers, T., Baruwa, A., Singh, A., Cheveleva, A., Ligozat, A.-L., Subramonian, A., Nev´eol, A., Lovering, C., Garrette, D., Tunuguntla, D., Reiter, E., Taktasheva, E., Voloshina, E., Bogdanov, E., Winata, G. I., Schoelkopf, H., Kalo, J.-C., Novikova, J., Forde, J. Z., Clive, J., Kasai, J., Kawamura, K., Hazan, L., Carpuat, M., Clinciu, M., Kim, N., Cheng, N., Serikov, O., Antverg, O., van der Wal, O., Zhang, R., Zhang, R., Gehrmann, S., Mirkin, S., Pais, S., Shavrina, T., Scialom, T., Yun, T., Limisiewicz, T., Rieser, V., Protasov, V., Mikhailov, V., Pruksachatkun, Y., Belinkov, Y., Bamberger, Z., Kasner, Z., Rueda, A., Pestana, A., Feizpour, A., Khan, A., Faranak, A., Santos, A., Hevia, A., Unldreaj, A., Aghagol, A., Abdollahi, A., Tammour, A., HajiHosseini, A., Behroozi, B., Ajibade, B., Saxena, B., Ferrandis, C. M., Contractor, D., Lansky, D., David, D., Kiela, D., Nguyen, D. A., Tan, E., Baylor, E., Ozoani, E., Mirza, F., Ononiwu, F., Rezanejad, H., Jones, H., Bhattacharya, I., Solaiman, I., Sedenko, I., Nejadgholi, I., Passmore, J., Seltzer, J., Sanz, J. B., Dutra, L., Samagaio, M., Elbadri, M., Mieskes, M., Gerchick, M., Akinlolu, M., McKenna, M., Qiu, M., Ghauri, M., Burynok, M., Abrar, N., Rajani, N., Elkott, N., Fahmy, N., Samuel, O., An, R., Kromann, R., Hao, R., Alizadeh, S., Shubber, S., Wang, S., Roy, S., Viguier, S., Le, T., Oyebade, T., Le, T., Yang, Y., Nguyen, Z., Kashyap, A. R., Palasciano, A., Callahan, A., Shukla, A., Miranda-Escalada, A., Singh, A., Beilharz, B., Wang, B., Brito, C., Zhou, C., Jain, C., Xu, C., Fourrier, C., Periñan, D. L., Molano, D., Yu, D., Manjavacas, E., Barth, F., Fuhrimann, F., Altay, G., Bayrak, G., Burns, G., Vrabec, H. U., Bello, I., Dash, I., Kang, J., Giorgi, J., Golde, J., Posada, J. D., Sivaraman, K. R., Bulchandani, L., Liu, L., Shinzato, L., de Bykhovetz, M. H., Takeuchi, M., Pamies, M., Castillo, M. A., Nezhurina, M., Sanger, M., Samwald, M., Cullan, M., Weinberg, M., De Wolf, M., Mihaljcic, M., Liu, M., Freidank, M., Kang, M., Seelam, N., Dahlberg, N., Broad, N. M., Muellner, N., Fung, P., Haller, P., Chandrasekhar, R., Eisenberg, R., Martin, R., Canalli, R., Su, R., Su, R., Cahyawijaya, S., Garda, S., Deshmukh, S. S., Mishra, S., Kiblawi, S., Ott, S., Sang-aroonsiri, S., Kumar, S., Schweter, S., Bharati, S., Laud, T., Gigant, T., Kainuma, T., Kusa, W., Labrak, Y., Bajaj, Y. S., Venkatraman, Y., Xu, Y., Xu, Y., Xu, Y., Tan, Z., Xie, Z., Ye, Z., Bras, M., Belkada, Y., and Wolf, T. BLOOM: A 176B-parameter open-access multilingual language model. Computing Research Repository, 2022. doi: 10.48550/arXiv.2211.05100. URL https://arxiv.org/abs/2211.05100v2. Version 2.
  77. 77.Sellam, T., Yadlowsky, S., Wei, J., Saphra, N., D’Amour, A., Linzen, T., Bastings, J., Turc, I., Eisenstein, J., Das, D., et al. The multiberts: Bert reproductions for robustness analysis. arXiv preprint arXiv:2106.16163, 2021.
  78. 78.Sharma, U. and Kaplan, J. A neural scaling law from the dimension of the data manifold. Computing Research Repository, 2020. doi: 10.48550/arXiv.2004.10802. URL https://arxiv.org/abs/2004.10802v1. Version 1.
  79. 79.Shin, S., Lee, S.-W., Ahn, H., Kim, S., Kim, H., Kim, B., Cho, K., Lee, G., Park, W., Ha, J.-W., and Sung, N. On the effect of pretraining corpora on in-context learning by a large-scale language model. 2022.
  80. 80.Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-LM: Training multi-billion parameter language models using model parallelism. Computing Research Repository, 2019. doi: 10.48550/arXiv.1909.08053. URL https://arxiv.org/abs/1909.08053v4. Version 4.
  81. 81.Silva, A., Chopra, R., and Gombolay, M. Cross-loss influence functions to explain deep network representations. In International Conference on Artificial Intelligence and Statistics, pp. 1–17. PMLR, 2022.
  82. 82.Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., Zhang, E., Child, R., Aminabadi, R. Y., Bernauer, J., Song, X., Shoeybi, M., He, Y., Houston, M., Tiwary, S., and Catanzaro, B. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. Computing Research Repository, 2022. doi: 10.48550/arXiv.2201.11990. URL https://arxiv.org/abs/2201.11990v3. Version 3.
  83. 83.Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y. RoFormer: Enhanced transformer with rotary position embedding. Computing Research Repository, 2021. doi: 10.48550/arXiv.2104.09864. URL https://arxiv.org/abs/2104.09864v4. Version 4.
  84. 84.Suarez, P. J. O., Sagot, B., and Romary, L. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7). Leibniz-Institut fur Deutsche Sprache, 2019.
  85. 85.Tang, J. WuDao: Pretrain the world. Keynote adress at the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, 2021.
  86. 86.Tirumala, K. N. B., Markosyan, A. H., Zettlemoyer, L., and Aghajanyan, A. Memorization without overfitting: Analyzing the training dynamics of large language models. ArXiv, abs/2205.10770, 2022.
  87. 87.Van der Wal, O., Bachmann, D., Leidinger, A., van Maanen, L., Zuidema, W., and Schulz, K. Undesirable biases in nlp: Averting a crisis of measurement. arXiv preprint arXiv:2211.13709, 2022a.
  88. 88.Van der Wal, O., Jumelet, J., Schulz, K., and Zuidema, W. The birth of bias: A case study on the evolution of gender bias in an english language model. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pp. 75–75, 2022b.
  89. 89.Wang, B. and Komatsuzaki, A. GPT-J-6B: A 6 billion parameter autoregressive language model, 2021.
  90. 90.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  91. 91.Xia, M., Artetxe, M., Zhou, C., Lin, X. V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V. Training trajectories of language models across scales, 2022. URL https://arxiv.org/abs/2212.09803.
  92. 92.Xu, F. F., Alon, U., Neubig, G., and Hellendoorn, V. J. A systematic evaluation of large language models of code. Computing Research Repository, 2022. doi: 10.48550/arXiv.2202.13169. URL https://arxiv.org/abs/2202.13169v3. Version 3.
  93. 93.Yoon, S. and Lee, H. Which model is helpful in solving privacy, memorization, and bias problems? 2021. URL https://soyoung97.github.io/profile/assets/papers/CS774.pdf.
  94. 94.Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414, 2022.
  95. 95.Zhang, G., Li, L., Nado, Z., Martens, J., Sachdeva, S., Dahl, G., Shallue, C., and Grosse, R. B. Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model. Advances in neural information processing systems, 32, 2019.
  96. 96.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. OPT: Open pre-trained transformer language models. Computing Research Repository, 2022. doi: 10.48550/arXiv.2205.01068. URL https://arxiv.org/abs/2205.01068v4. Version 4.
  97. 97.Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. arXiv preprint arXiv:1707.09457, 2017.
  98. 98.Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pp. 15–20, 2018.

Citation

MLA
Biderman, S., et al. “Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”. arXiv, 2023, http://arxiv.org/abs/2304.01373v2.
APA
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O'Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., & Wal, O. van . der . (2023). Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. arXiv. http://arxiv.org/abs/2304.01373v2
Chicago
Biderman, S., H. Schoelkopf, Q. Anthony, et al. 2023. “Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”. arXiv. http://arxiv.org/abs/2304.01373v2.
Harvard
Biderman, S. et al. (2023) “Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.01373v2.
Vancouver
1. Biderman S, Schoelkopf H, Anthony Q, et al (2023) Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. arXiv

BibTeX

@article{biderman2023pythia,
  title = {Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling},
  author = {Biderman, Stella and Schoelkopf, Hailey and Anthony, Quentin and Bradley, Herbie and O'Brien, Kyle and Hallahan, Eric and Khan, Mohammad Aflah and Purohit, Shivanshu and Prashanth, USVSN Sai and Raff, Edward and Skowron, Aviya and Sutawika, Lintang and Wal, Oskar van der},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.01373v2},
  eprint = {2304.01373}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/