Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. RaeSebastian BorgeaudTrevor CaiKatie MillicanJordan HoffmannFrancis SongJohn AslanidesSarah HendersonRoman RingSusannah Young

article2021arXiv1,665 citations

Analyzes the performance of Transformer language models scaled up to the 280-billion-parameter Gopher across 152 tasks, providing empirical evidence on which capabilities improve with size and where logical reasoning and safety limitations persist.

Listen

Large language models are rapidly advancing artificial intelligence capabilities, yet developing and evaluating systems at massive scale requires immense computational investment and presents complex risks. To understand how increasing model size affects performance, safety, and operational reliability, modern AI research must systematically evaluate where scale yields breakthroughs and where fundamental limitations remain.

The article evaluates how scaling Transformer-based language models impacts language understanding, reasoning capabilities, bias, and toxic behavior across diverse benchmark domains. Specifically, it demonstrates the performance gains and behavioral characteristics of a 280-billion-parameter model named Gopher alongside a family of smaller architectures.

To conduct this evaluation, the researchers trained six autoregressive Transformer models ranging from 44 million to 280 billion parameters on MassiveText, an English-heavy, multi-source dataset comprising web pages, books, news articles, and code. The models were evaluated across 152 distinct tasks covering academic subjects, common sense, fact-checking, and reading comprehension. The analysis also measured the models' tendencies regarding toxic language generation, toxicity classification, and social biases across various demographic groups.

The key findings reveal that model scaling yields substantial but uneven performance improvements. First, Gopher outperformed previous state-of-the-art models on approximately 81% of evaluated tasks, achieving an overall accuracy of 60.0% on the Massive Multitask Language Understanding benchmark compared to 43.9% for GPT-3. Second, gains from scale were concentrated in reading comprehension, fact-checking, and knowledge-intensive subjects, whereas mathematical and logical reasoning showed minimal benefit or slight regression. Third, while scaling improved the ability to classify toxic language, larger models also generated more toxic text when given toxic prompts. Fourth, increased model size did not mitigate distributional biases relating to gender stereotypes, social sentiment, or dialect representation.

These findings imply that scaling parameters alone is not a universal solution for complex reasoning and ethical alignment challenges. While scale significantly enhances factual recall and comprehension, it increases the risk of replicating harmful training patterns. Consequently, technical safety mitigations and task-specific fine-tuning are best implemented downstream in application-specific environments rather than relying solely on raw scale or rigid pre-training data filtering.

Decision-makers should pursue efficient, specialized architectures—such as retrieval-augmented systems and modular expert models—to reduce computational overhead while addressing reasoning deficits. When deploying conversational interfaces, organizations should apply robust downstream safeguards, such as dialogue prompting and red-teaming defenses, to manage misinformation and offensive outputs.

The study's primary limitations stem from reliance on automated bias and toxicity classifiers that carry inherent societal biases, potential test-set leakage within web-scale corpora, and the restriction to English-dominant datasets. While there is high confidence in the empirical scaling trajectories and factual performance gains, stakeholders should exercise caution regarding the factual reliability and common-sense reasoning of generated outputs in high-stakes domains.

arXiv: 2112.11446
Cover for Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Abstract

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher. These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit. We provide a holistic analysis of the training dataset and model's behaviour, covering the intersection of model scale with bias and toxicity. Finally we discuss the application of language models to AI safety and the mitigation of downstream harms.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Method
  • 3.1 Models
  • 3.2 Training
  • 3.3 Infrastructure
  • 3.4 Training Dataset
  • 4 Results
  • 4.1 Task Selection
  • 4.2 Comparisons with State of the Art
  • 4.3 Performance Improvements with Scale
  • 5 Toxicity and Bias Analysis
  • 5.1 Toxicity
  • 5.1.1 Generation Analysis
  • 5.1.2 Classification Analysis
  • 5.2 Distributional Bias
  • 5.2.1 Gender and Occupation Bias
  • 5.2.2 Sentiment Bias towards Social Groups
  • 5.2.3 Perplexity on Dialects
  • 6 Dialogue
  • 6.1 Prompting For Dialogue
  • 6.2 Fine-tuning for Dialogue
  • 6.3 Dialogue & Toxicity
  • 7 Discussion
  • 7.1 Towards Efficient Architectures
  • 7.2 Challenges in Toxicity and Bias
  • 7.3 Safety benefits and safety risks
  • 8 Conclusion
  • 9 Acknowledgements
  • 10 Contributions
  • References
  • A MassiveText
  • A.1 Dataset Pipeline
  • A.1.1 Pipeline stages
  • A.1.2 Constructing Token Sequences
  • A.2 Dataset Analysis
  • A.3 Dataset Ablations
  • A.3.1 MassiveText Subsets Weighting
  • A.3.2 Iterative Refinement of MassiveWeb
  • A.4 Text normalisation
  • A.5 MassiveText Datasheet
  • B Gopher Model Card
  • C Lessons Learned
  • C.1 Adafactor
  • C.2 Lower-Precision Training with bfloat16
  • D Results
  • D.1 Overview
  • D.2 Pile
  • D.3 Language Modelling
  • D.4 Filtering Test-Set Documents
  • D.5 Scaling Curves
  • D.6 Scaling Context Length
  • D.7 MMLU
  • D.8 BIG-bench
  • D.8.1 Task Selection

Knowls

  1. Knowl 1 — Gopher Transformer Architecture and Model Family Specifications

    model/method

    The Gopher family comprises six autoregressive decoder-only Transformer language models ranging from 44 million to 280 billion parameters. The architecture adapts the standard GPT-style autoregressive Transformer with two core modifications:

    1. RMSNorm: Root Mean Square Layer Normalization is used instead of standard LayerNorm.
    2. Relative Positional Encodings: The relative positional encoding scheme from Transformer-XL is used rather than absolute positional encodings, enabling context length extrapolation during evaluation by clamping relative positions to the training context length (2,048 tokens).

    Text is tokenized using SentencePiece with a vocabulary size of 32,000, using a byte-level backoff to UTF-8 bytes to ensure an open vocabulary. In all models, the feed-forward network hidden dimension is 4×dmodel4 \times d_{\text{model}}.

    Model Layers Heads Key/Value Size dmodeld_{\text{model}} Max LR Batch Size (tokens)
    44M 8 16 32 512 6×10−46 \times 10^{-4} 0.25M
    117M 12 12 64 768 6×10−46 \times 10^{-4} 0.25M
    417M 12 12 128 1,536 2×10−42 \times 10^{-4} 0.25M
    1.4B 24 16 128 2,048 2×10−42 \times 10^{-4} 0.25M
    7.1B 32 32 128 4,096 1.2×10−41.2 \times 10^{-4} 2M
    Gopher 280B 80 128 128 16,384 4×10−54 \times 10^{-5} 3M →\to 6M

    Models are trained on 300 billion tokens with the Adam optimizer using a 1,500-step linear learning rate warm-up from 10−710^{-7} to the maximum learning rate, followed by a 10-fold decay using a cosine schedule. Global gradient norm clipping is set to 1.0 for models up to 1.4B, and reduced to 0.25 for the 7.1B model and Gopher (280B) to ensure training stability.

  2. Knowl 2 — MassiveText Dataset Composition and Multi-Stage Quality Pipeline

    model/method

    MassiveText is a curated pre-training corpus comprising 2.35 billion documents (~10.5 TB of text; ~2.34 trillion SentencePiece tokens) across six distinct subsets, sampled non-uniformly during training to optimize downstream transfer:

    • MassiveWeb (48% sampling weight; 1.9 TB, 506B tokens): Web text extracted from HTML markup using DOM tree structure grouping, filtered via Google SafeSearch.
    • Books (27% sampling weight; 2.1 TB, 560B tokens): Digitized books spanning 1500 to 2008.
    • C4 (10% sampling weight; 0.75 TB, 182B tokens): Public Common Crawl dataset.
    • News (10% sampling weight; 2.7 TB, 676B tokens): English news articles.
    • GitHub (3% sampling weight; 3.1 TB, 422B tokens): Permissively licensed public source code files.
    • Wikipedia (2% sampling weight; 0.001 TB, 4B tokens): October 2020 English Wikipedia dump.

    The dataset is curated using a sequential multi-stage filtering pipeline:

    1. Language & Content Filtering: Non-English documents and pages failing SafeSearch are filtered out.
    2. Quality Heuristics: Documents must contain between 50 and 100,000 words; mean word length must be between 3 and 10 characters; symbol-to-word ratios for '#' or '...' must not exceed 0.1; <90%<90\% of lines may start with bullets; <30%<30\% of lines may end with ellipses; ≥80%\ge 80\% of words must contain an alphabetic character; and at least two basic stop words (the, be, to, of, and, that, have, with) must be present.
    3. Repetition Removal: Documents are discarded if duplicate line or paragraph fractions exceed 0.30, if duplicate character fractions exceed 0.20, or if high-frequency nn-gram duplicate thresholds are crossed (e.g., top 2-gram character fraction >0.20>0.20, duplicate 10-gram character fraction >0.10>0.10).
    4. Document Deduplication & Leakage Filtering: Exact duplicate documents are removed, and near-duplicate documents with MinHash 13-gram Jaccard similarity >0.8>0.8 are deduplicated. The same 13-gram Jaccard threshold (>0.8>0.8) is applied against evaluation test sets (WikiText-103, C4, Curation Corpus, and LAMBADA) to eliminate test set contamination.
  3. Knowl 3 — Zero-Padding Token Sequence Construction Algorithm

    algorithm

    To train fixed-length sequence models without wasting compute on padding (PAD) tokens and avoiding start-of-document boundary bias, training sequences of length n=2048n = 2048 tokens are generated using the following procedure:

    Input: Corpus subsets with sampling weights WW, training sequence length n=2048n = 2048, crop byte length C=15×nC = 15 \times n
    Output: Continuous non-padded token batches of length nn
    for each batch to construct do
        Initialize an empty token buffer BtokensB_{\text{tokens}}
        while length of Btokens<10×nB_{\text{tokens}} < 10 \times n do
            Sample a dataset subset according to weights WW
            Select a document of size BB bytes uniformly at random from the subset
            Sample start offset s∼U(−C4,B−C4)s \sim \mathcal{U}\left(-\frac{C}{4}, B - \frac{C}{4}\right)
            Extract byte slice b←document[max⁡(0,s):min⁡(B,s+C)]b \leftarrow \text{document}[\max(0, s) : \min(B, s + C)]
            Convert byte slice bb into tokens tt using the SentencePiece tokenizer
            Prepend [BOS] and append [EOS] tokens to tt
            Append tt to BtokensB_{\text{tokens}}
        end while
        Split BtokensB_{\text{tokens}} into contiguous chunks of exactly n=2048n = 2048 tokens
        Discard the trailing chunk if its length <n< n
        Yield chunks for training batch construction
    end for
  4. Knowl 4 — bfloat16 Parameter Staleness and Optimizer State Precision Fix

    empirical result

    Training large autoregressive language models (417M to 280B parameters) purely in bfloat16 precision for both activations and model parameters leads to significant performance degradation and layer staleness. Because parameter updates (∥Δθ∥\|\Delta \theta\|) at small learning rates fall below the numerical resolution of bfloat16, many parameters fail to register updates over long training durations. Applying stochastic rounding does not resolve this performance gap.

    The optimal resolution maintains parameters and activations in bfloat16 during forward and backward passes, but stores a master copy of the weights in float32 inside the partitioned optimizer state (e.g., using ZeRO-style partitioning). Parameter updates are computed and applied in full float32 precision and cast back to bfloat16 for the forward pass, with attention softmax and cross-entropy loss computed in float32. This matches full float32 training performance while delivering a 1.4×1.4\times training throughput speedup on TPUv3 hardware.

  5. Knowl 5 — Pre-training Instability of Adafactor at Large Scale

    empirical result

    While the Adafactor optimizer reduces optimizer memory footprint by factoring the second-moment statistics and achieves stable pre-training for smaller models (up to 1.4B parameters), it exhibits severe instabilities and loss divergences when scaled to 7.1B parameters and above. When training a 7.1B parameter model with Adafactor at a reduced maximum learning rate of 6×10−56 \times 10^{-5}, the training run displayed repeated loss spikes and minor divergences compared to a stable Adam baseline trained at a higher learning rate of 1.2×10−41.2 \times 10^{-4}. Consequently, Adafactor is unsuitable for large-scale language model pre-training from scratch, though it remains viable for downstream fine-tuning where memory constraints dominate.

  6. Knowl 6 — Benchmark Evaluation and Non-Uniform Scaling on MMLU

    empirical result

    On the 57-subject Massive Multitask Language Understanding (MMLU) benchmark evaluated in a 5-shot setting, Gopher (280B) achieves an average accuracy of 60.0%60.0\%, improving substantially over prior state-of-the-art models including GPT-3 175B (43.9%43.9\%) and UnifiedQA 11B (48.9%48.9\%), while narrowing the gap toward estimated human expert performance (89.8%89.8\%).

    Model / Benchmark Baseline MMLU Average Accuracy (%)
    Random Baseline 25.0
    GPT-2 (fine-tuned) 32.4
    Average Human Rater 34.5
    GPT-3 (175B, 5-shot) 43.9
    UnifiedQA (11B, fine-tuned) 48.9
    Gopher (280B, 5-shot) 60.0
    Average Human Expert 89.8

    Performance scaling across MMLU subjects is highly non-uniform: knowledge-intensive domains (humanities, medicine, professional law, social sciences) experience large performance leaps with parameter scale, whereas mathematical and formal logic tasks show negligible or negative scaling (Abstract Algebra: 25.0%25.0\%; High School Mathematics: 23.7%23.7\%), demonstrating that scale alone without explicit reasoning mechanisms does not resolve algorithmic mathematical reasoning.

  7. Knowl 7 — Scaling Properties and Epistemic Limits in Fact-Checking on FEVER

    empirical result

    In a 15-shot closed-book fact-checking evaluation on FEVER (classifying claims into SUPPORTED, REFUTED, or NOTENOUGHINFO), model scale improves 3-way classification accuracy monotonically from sub-1B models up to 50%50\% for Gopher (280B). When isolating pairwise classifications:

    • SUPPORTED vs. REFUTED: Accuracy scales directly with parameters, reaching 78%78\% on Gopher 280B as the model memorizes more world knowledge.
    • REFUTED vs. NOTENOUGHINFO: Performance plateaus near 1B parameters and fails to scale further.

    This plateau reveals a fundamental epistemic limitation of scaling: standard autoregressive language models conflate a lack of information with claim refutation, failing to recognize when knowledge is absent. In an open-book setup where oracle evidence is prepended to the prompt, Gopher (280B) achieves 88%88\% accuracy (macro-F1 89%89\%, outperforming the 1.5B GPT-2 baseline at 71%71\%), slightly surpassing supervised task-specific models.

  8. Knowl 8 — Asymmetric Scaling of Toxicity in Generation versus Few-Shot Classification

    empirical result

    Analysis using the Perspective API classifier reveals asymmetric behavior between text generation and text classification across model scale:

    1. Unconditional Generation: Unprompted samples exhibit low toxicity (median toxicity ≈0.055\approx 0.055, mean ≈0.075\approx 0.075 for 280B) that does not increase with model scale and remains lower than the training data average (0.0980.098), indicating models do not inherently amplify pre-training toxicity.
    2. Conditional Generation: When prompted with toxic prefixes from RealToxicityPrompts (RTP), larger models produce higher continuation toxicity than smaller models (continuation toxicity increases with scale up to 7.1B parameters). Larger models act as more faithful conditional density estimators, matching the prompt's toxicity distribution.
    3. Few-Shot Toxicity Classification: Evaluated on CivilComments in a 20-shot setting, toxicity classification AUC increases monotonically with scale, rising from near-random guessing (AUC ≈0.50\approx 0.50) for models <1.4B<1.4\text{B} parameters to AUC≈0.76\text{AUC} \approx 0.76 for Gopher 280B.
  9. Knowl 9 — Distributional Bias and Dialect Perplexity Across Model Scale

    empirical result

    Distributional bias is measured across model sizes using gender-occupation association metrics and dialect perplexity evaluations:

    • Gender-Occupation Bias: Using the metric B(occupation)=log⁡∑gf∈GfP(gf∣prompt)∑gm∈GmP(gm∣prompt)B(\text{occupation}) = \log \frac{\sum_{g_f \in G_f} P(g_f \mid \text{prompt})}{\sum_{g_m \in G_m} P(g_m \mid \text{prompt})} where GfG_f and GmG_m are sets of female and male terms, Gopher reflects societal stereotypes (34.2%34.2\% of occupations biased female, 65.8%65.8\% biased male; female occupations average bias +0.52+0.52, male occupations average bias −0.74-0.74). The aggregate absolute bias 1∣O∣∑o∣B(o)∣\frac{1}{|O|}\sum_{o} |B(o)| does not correlate with or decrease with model scale, and measurements vary widely based on arbitrary prompt template phrasing (was vs is).
    • Dialect Modeling: Perplexity on African American English (AAE)-aligned tweets remains substantially higher than on White-aligned English across all model sizes (44M to 280B). Both perplexities decrease with scale at identical rates, leaving the relative representation gap completely unclosed by scale.
  10. Knowl 10 — Parity Between Dialogue Prompting and Supervised Dialogue Fine-Tuning

    empirical result

    Conditioning a frozen 280B Gopher model with an ~800-token persona prompt (describing Gopher as a polite, knowledgeable, and safe conversationalist with explicit multi-turn examples) achieves conversational quality on par with dedicated supervised fine-tuning.

    In a double-blind human preference study comprising 1,400 comparative ratings against Dialogue-Tuned Gopher (fine-tuned on ~5 billion tokens of filtered two-person dialogue data from MassiveWeb), human raters showed no statistically significant preference for the fine-tuned model over the prompted model (50%±4%50\% \pm 4\% preference). Furthermore, while unprompted language models increase continuation toxicity with scale when given toxic inputs, Dialogue-Prompted Gopher suppresses toxic continuation increases across scales, maintaining low toxicity even when fed toxic conversational questions.

  11. Knowl 11 — Warm-Starting Model Growth via Layer Replication and Weight Tiling

    model/method

    Pre-trained language models can be expanded in depth and width to accelerate the training of larger architectures (warm starting):

    • Depth Expansion: Layer parameters are replicated by mapping new layer indices i∈{0,…,Lnew−1}i \in \{0, \dots, L_{\text{new}}-1\} to base layer indices using ⌊i⋅Lold/Lnew⌉\lfloor i \cdot L_{\text{old}} / L_{\text{new}} \rceil.
    • Width Expansion: For attention projections of shape nH×nHn H \times n H expanding to mm heads of size HH, the first m−nm-n head parameters are tiled to the right of the weight matrix, and bottleneck activations are widened by replicating the top (m−n)×H(m-n) \times H parameters to the bottom, followed by adding small random noise.

    While expanding model width leads to worse performance at convergence than training from scratch, pure depth expansion (e.g., expanding a 4.5B model to 9B by duplicating layers) reaches the same final evaluation loss as a model trained from scratch to 300B tokens while reducing total pre-training compute by nearly 40%40\%.

Coverage note — Omitted fine-grained per-task accuracy tables across all 62 BIG-bench subtasks and individual sub-datasets of The Pile, as well as transient hyperparameter scans for sparse training / reverse distillation, as their principal conclusions are subsumed within the broader MMLU, FEVER, and efficiency knowls.

References

  1. 1.M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pages 308–318, 2016.
  2. 2.M. Abubakar. GitHub Copilot AI is generating and giving out functional API keys. https://fossbytes.com/github-copilot-generating-functional-api-keys, 2021. Accessed: 2021-10-7.
  3. 3.D. Adiwardana, M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha, G. Nemade, Y. Lu, et al. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977, 2020.
  4. 4.R. Anil, B. Ghazi, V. Gupta, R. Kumar, and P. Manurangsi. Large-scale differentially private BERT. arXiv preprint arXiv:2108.01624, 2021.
  5. 5.A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  6. 6.I. Augenstein, C. Lioma, D. Wang, L. C. Lima, C. Hansen, and J. G. S. Christian Hansen. Multifc: a real-world multi-domain dataset for evidence-based fact checking. In Transactions of the Association for Computational Linguistics (TACL), 2021, Online, Nov. 2019. Association for Computational Linguistics. URL https://arxiv.org/abs/1909.03242.
  7. 7.J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  8. 8.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014.
  9. 9.E. Ben Zaken, S. Ravfogel, and Y. Goldberg. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv e-prints, pages arXiv–2106, 2021.
  10. 10.E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021.
  11. 11.Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155, 2003.
  12. 12.BIG-bench collaboration. Beyond the imitation game: Measuring and extrapolating the capabilities of language models. In preparation, 2021. URL https://github.com/google/BIG-bench/.
  13. 13.Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al. PIQA: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7432–7439, 2020.
  14. 14.D. W. Blalock, J. J. G. Ortiz, J. Frankle, and J. V. Guttag. What is the state of neural network pruning? CoRR, abs/2003.03033, 2020. URL https://arxiv.org/abs/2003.03033.
  15. 15.S. L. Blodgett, L. Green, and B. O’Connor. Demographic dialectal variation in social media: A case study of african-american english. CoRR, abs/1608.08868, 2016. URL http://arxiv.org/abs/1608.08868.
  16. 16.S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach. Language (technology) is power: A critical survey of "bias" in NLP. ACL, 2020.
  17. 17.S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach. Stereotyping norwegian salmon: an inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, 2021.
  18. 18.F. Böhm, Y. Gao, C. M. Meyer, O. Shapira, I. Dagan, and I. Gurevych. Better rewards yield better summaries: Learning to summarise without references. arXiv preprint arXiv:1909.01214, 2019.
  19. 19.S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre. Improving language models by retrieving from trillions of tokens. arXiv submission, 2021.
  20. 20.D. Borkan, L. Dixon, J. Sorensen, N. Thain, and L. Vasserman. Nuanced metrics for measuring unintended bias with real data for text classification. CoRR, abs/1903.04561, 2019. URL http://arxiv.org/abs/1903.04561.
  21. 21.J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang. JAX: composable transformations of Python+NumPy programs. 2018. URL http://github.com/google/jax.
  22. 22.T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean. Large language models in machine translation. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pages 858–867, Prague, Czech Republic, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/D07-1090.
  23. 23.E. Brill and R. C. Moore. An improved error model for noisy channel spelling correction. In Proceedings of the 38th annual meeting of the association for computational linguistics, pages 286–293, 2000.
  24. 24.P. F. Brown, J. Cocke, S. A. Della Pietra, V. J. Della Pietra, F. Jelinek, J. Lafferty, R. L. Mercer, and P. S. Roossin. A statistical approach to machine translation. Computational linguistics, 16(2):79–85, 1990.
  25. 25.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  26. 26.J. Buckman. Fair ML tools require problematic ML models. https://jacobbuckman.com/2021-02-15-fair-ml-tools-require-problematic-ml-models. Accessed: 2021-10-7.
  27. 27.N. Burgess, J. Milanovic, N. Stephens, K. Monachopoulos, and D. Mansell. Bfloat16 processing for neural networks. In 2019 IEEE 26th Symposium on Computer Arithmetic (ARITH), pages 88–91. IEEE, 2019.
  28. 28.A. Caliskan, J. J. Bryson, and A. Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186, 2017.
  29. 29.O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom. e-SNLI: Natural language inference with natural language explanations. arXiv preprint arXiv:1812.01193, 2018.
  30. 30.Y. T. Cao and H. Daumé. Toward gender-inclusive coreference resolution: An analysis of gender and bias throughout the machine learning lifecyle. Computational Linguistics, pages 1–47, 2021.
  31. 31.N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium (USENIX Security 19), pages 267–284, 2019.
  32. 32.N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel. Extracting training data from large language models, 2021.
  33. 33.N. Chater. The search for simplicity: A fundamental cognitive principle? The Quarterly Journal of Experimental Psychology Section A, 52(2):273–302, 1999.
  34. 34.H. Chen, X. Liu, D. Yin, and J. Tang. A survey on dialogue systems: Recent advances and new frontiers. Acm Sigkdd Explorations Newsletter, 19(2):25–35, 2017.
  35. 35.T. Chen, I. Goodfellow, and J. Shlens. Net2net: Accelerating learning via knowledge transfer. arXiv preprint arXiv:1511.05641, 2015.
  36. 36.P. Christiano, B. Shlegeris, and D. Amodei. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575, 2018.
  37. 37.D. Coyle and A. Weller. “Explaining” machine learning reveals policy challenges. Science, 368(6498): 1433–1434, 2020.
  38. 38.Curation. Curation corpus base, 2020. URL https://github.com/CurationCorp/curation-corpus.
  39. 39.I. Dagan, O. Glickman, and B. Magnini. The PASCAL recognising textual entailment challenge. In Proceedings of the PASCAL Challenges Workshop on Recognising Textual Entailment, 2005. URL http://www.cs.biu.ac.il/~glikmao/rte05/.
  40. 40.Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  41. 41.J. Dastin. Amazon scraps secret AI recruiting tool that showed bias against women. 2018.
  42. 42.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. IEEE, 2009.
  43. 43.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1), 2019.
  44. 44.G. Dewey. Relativ [sic] frequency of English speech sounds. Harvard UP, 1923.
  45. 45.L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman. Measuring and mitigating unintended bias in text classification. 2018.
  46. 46.J. Dodge, M. Sap, A. Marasovic, W. Agnew, G. Ilharco, D. Groeneveld, and M. Gardner. Documenting the english colossal clean crawled corpus. CoRR, abs/2104.08758, 2021. URL https://arxiv.org/abs/2104.08758.
  47. 47.E. Elsen, M. Dukhan, T. Gale, and K. Simonyan. Fast sparse convnets. CoRR, abs/1911.09723, 2019. URL http://arxiv.org/abs/1911.09723.
  48. 48.U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning, pages 2943–2952. PMLR, 2020.
  49. 49.W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961, 2021.
  50. 50.J. Frankle and M. Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJl-b3RcF7.
  51. 51.T. Gale, E. Elsen, and S. Hooker. The state of sparsity in deep neural networks. CoRR, abs/1902.09574, 2019. URL http://arxiv.org/abs/1902.09574.
  52. 52.T. Gale, M. Zaharia, C. Young, and E. Elsen. Sparse GPU kernels for deep learning. CoRR, abs/2006.10901, 2020. URL https://arxiv.org/abs/2006.10901.
  53. 53.L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  54. 54.T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford. Datasheets for datasets. arXiv preprint arXiv:1803.09010, 2018.
  55. 55.S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online, Nov. 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  56. 56.A. Graves. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.
  57. 57.A. Griewank and A. Walther. Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation. ACM Transactions on Mathematical Software (TOMS), 26(1):19–45, 2000.
  58. 58.S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan. Deep learning with limited numerical precision. In International conference on machine learning, pages 1737–1746. PMLR, 2015.
  59. 59.K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang. REALM: Retrieval-augmented language model pre-training, 2020.
  60. 60.D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan. Cooperative inverse reinforcement learning. Advances in neural information processing systems, 29:3909–3917, 2016.
  61. 61.S. Han, H. Mao, and W. J. Dally. Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding. In Y. Bengio and Y. LeCun, editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1510.00149.
  62. 62.A. Hanna, E. Denton, A. Smart, and J. Smith-Loud. Towards a critical race methodology in algorithmic fairness. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, Jan 2020. doi: 10.1145/3351095.3372826. URL http://dx.doi.org/10.1145/3351095.3372826.
  63. 63.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  64. 64.T. Hennigan, T. Cai, T. Norman, and I. Babuschkin. Haiku: Sonnet for JAX. 2020. URL http://github.com/deepmind/dm-haiku.
  65. 65.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  66. 66.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  67. 67.A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2019.
  68. 68.P.-S. Huang, H. Zhang, R. Jiang, R. Stanforth, J. Welbl, J. Rae, V. Maini, D. Yogatama, and P. Kohli. Reducing sentiment bias in language models via counterfactual evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 65–83, Online, Nov. 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.7. URL https://aclanthology.org/2020.findings-emnlp.7.
  69. 69.Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems, 32:103–112, 2019.
  70. 70.G. Irving and A. Askell. AI safety needs social scientists. Distill, 2019. doi: 10.23915/distill.00014. https://distill.pub/2019/safety-needs-social-scientists.
  71. 71.G. Irving, P. Christiano, and D. Amodei. AI safety via debate. arXiv preprint arXiv:1805.00899, 2018.
  72. 72.N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. S. Gu, and R. Picard. Human-centric dialog training via offline reinforcement learning. arXiv preprint arXiv:2010.05848, 2020.
  73. 73.S. Jayakumar, R. Pascanu, J. Rae, S. Osindero, and E. Elsen. Top-KAST: Top-K Always Sparse Training. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 20744–20754. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/ee76626ee11ada502d5dbf1fb5aae4d2-Paper.pdf.
  74. 74.F. Jelinek. Statistical methods for speech recognition. MIT press, 1997.
  75. 75.Y. Jiang, S. Wu, J. Gong, Y. Cheng, P. Meng, W. Lin, Z. Chen, and M. Li. Improving machine reading comprehension with single-choice decision and transfer learning. arXiv preprint arXiv:2011.03292, 2020. URL https://arxiv.org/abs/2011.03292.
  76. 76.X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174, Online, Nov. 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.372. URL https://aclanthology.org/2020.findings-emnlp.372.
  77. 77.M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. arXiv e-prints, art. arXiv:1705.03551, 2017.
  78. 78.N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson. A domain-specific supercomputer for training deep neural networks. Commun. ACM, 63(7):67–78, June 2020. ISSN 0001-0782. doi: 10.1145/3360307. URL https://doi.org/10.1145/3360307.
  79. 79.R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016.
  80. 80.J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  81. 81.A. Kasirzadeh. Reasons, values, stakeholders: a philosophical framework for explainable artificial intelligence. arXiv preprint arXiv:2103.00752, 2021.
  82. 82.M. Kay, C. Matuszek, and S. A. Munson. Unequal representation and gender stereotypes in image search results for occupations. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, pages 3819–3828, 2015.
  83. 83.Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving. Alignment of language agents. arXiv preprint arXiv:2103.14659, 2021.
  84. 84.U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HklBjCEKvH.
  85. 85.P. Kharya and A. Alvi. Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, the World’s Largest and Most Powerful Generative Language Model. https://developer.nvidia.com/blog/using-deepspeed-and-megatron-to-train-megatron-turing-nlg-530b-the-worlds-largest-and-most-powerful-generative-language-model/, 2021.
  86. 86.D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi. UnifiedQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online, Nov. 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.171. URL https://aclanthology.org/2020.findings-emnlp.171.
  87. 87.Y. J. Kim, A. A. Awan, A. Muzio, A. F. C. Salinas, L. Lu, A. Hendy, S. Rajbhandari, Y. He, and H. H. Awadalla. Scalable and efficient moe training for multitask multilingual models, 2021.
  88. 88.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  89. 89.S. Kiritchenko and S. M. Mohammad. Examining gender and race bias in two hundred sentiment analysis systems. CoRR, abs/1805.04508, 2018. URL http://arxiv.org/abs/1805.04508.
  90. 90.C. Kruengkrai, J. Yamagishi, and X. Wang. A multi-level attention model for evidence-based fact checking. arXiv preprint arXiv:2106.00950, 2021.
  91. 91.T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226, 2018.
  92. 92.T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics, 2019.
  93. 93.G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark, Sept. 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1082. URL https://aclanthology.org/D17-1082.
  94. 94.A. Lazaridou, A. Kuncoro, E. Gribovskaya, D. Agrawal, A. Liska, T. Terzi, M. Gimenez, C. d. M. d’Autume, S. Ruder, D. Yogatama, et al. Pitfalls of static language modelling. arXiv preprint arXiv:2102.01951, 2021.
  95. 95.K. Lee, M.-W. Chang, and K. Toutanova. Latent Retrieval for Weakly Supervised Open Domain Question Answering. In ACL, 2019.
  96. 96.K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. CoRR, abs/2107.06499, 2021a. URL https://arxiv.org/abs/2107.06499.
  97. 97.N. Lee, B. Z. Li, S. Wang, W. tau Yih, H. Ma, and M. Khabsa. Language models as fact checkers?, 2020.
  98. 98.N. Lee, Y. Bang, A. Madotto, and P. Fung. Towards few-shot fact-checking via perplexity. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1971–1981, Online, June 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.158. URL https://aclanthology.org/2021.naacl-main.158.
  99. 99.S. Legg and M. Hutter. Universal intelligence: A definition of machine intelligence. Minds and machines, 17(4):391–444, 2007.
  100. 100.J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  101. 101.D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen. {GS}hard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=qrwe7XHTmYb.
  102. 102.M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer. BASE layers: Simplifying training of large, sparse models. In M. Meila and T. Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 6265–6274. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/lewis21a.html.
  103. 103.X. L. Li, A. Kuncoro, C. d. M. d’Autume, P. Blunsom, and A. Nematzadeh. A systematic investigation of commonsense understanding in large language models. arXiv preprint arXiv:2111.00607, 2021.
  104. 104.O. Lieber, O. Sharir, B. Lenz, and Y. Shoham. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 2021.
  105. 105.J. Lin, A. Yang, J. Bai, C. Zhou, L. Jiang, X. Jia, A. Wang, J. Zhang, Y. Li, W. Lin, J. Zhou, and H. Yang. M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining. 2021a.
  106. 106.S. Lin, J. Hilton, and O. Evans. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021b.
  107. 107.W. Liu, P. Zhou, Z. Wang, Z. Zhao, H. Deng, and Q. Ju. Fastbert: a self-distilling BERT with adaptive inference time. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6035–6044, 2020.
  108. 108.E. Loper and S. Bird. NLTK: The natural language toolkit. ArXiv, abs/0205028, 2002.
  109. 109.M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz. Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2):313–330, 1993. URL https://aclanthology.org/J93-2004.
  110. 110.S. Merity, C. Xiong, J. Bradbury, and R. Socher. Pointer sentinel mixture models. International Conference on Learning Representations, 2017.
  111. 111.P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu. Mixed precision training. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=r1gs9JgRZ.
  112. 112.T. Mikolov, M. Karafiát, L. Burget, J. Cernock`y, and S. Khudanpur. Recurrent neural network based language model. In Interspeech, volume 2, pages 1045–1048. Makuhari, 2010.
  113. 113.T. Mikolov, A. Deoras, S. Kombrink, L. Burget, and J. H. Černocký. Empirical evaluation and combination of advanced language modeling techniques. In Interspeech, 2011.
  114. 114.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  115. 115.A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, G. Venkatesh, C. Yu, and P. Micikevicius. Accelerating sparse deep neural networks, 2021.
  116. 116.M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, pages 220–229, 2019.
  117. 117.S. Mohamed, M. Png, and W. Isaac. Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence. CoRR, abs/2007.04068, 2020. URL https://arxiv.org/abs/2007.04068.
  118. 118.G. E. Moore et al. Cramming more components onto integrated circuits, 1965.
  119. 119.S. Narang, G. F. Diamos, S. Sengupta, and E. Elsen. Exploring sparsity in recurrent neural networks. CoRR, abs/1704.05119, 2017. URL http://arxiv.org/abs/1704.05119.
  120. 120.H. Ney, U. Essen, and R. Kneser. On structuring probabilistic dependences in stochastic language modelling. Computer Speech & Language, 8(1):1–38, 1994.
  121. 121.D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández. The LAMBADA dataset: Word prediction requiring a broad discourse context, 2016.
  122. 122.A. Parikh, O. Täckström, D. Das, and J. Uszkoreit. A decomposable attention model for natural language inference. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2249–2255, Austin, Texas, Nov. 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1244. URL https://aclanthology.org/D16-1244.
  123. 123.D. A. Patterson, J. Gonzalez, Q. V. Le, C. Liang, L. Munguia, D. Rothchild, D. R. So, M. Texier, and J. Dean. Carbon emissions and large neural network training. CoRR, abs/2104.10350, 2021. URL https://arxiv.org/abs/2104.10350.
  124. 124.E. Perez, S. Karamcheti, R. Fergus, J. Weston, D. Kiela, and K. Cho. Finding generalizable evidence by learning to convince Q&A models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2402–2411, Hong Kong, China, Nov. 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1244. URL https://aclanthology.org/D19-1244.
  125. 125.E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving. Red teaming language models with language models. To appear, 2022.
  126. 126.A. Peste, E. Iofinova, A. Vladu, and D. Alistarh. AC/DC: alternating compressed/decompressed training of deep neural networks. CoRR, abs/2106.12379, 2021. URL https://arxiv.org/abs/2106.12379.
  127. 127.O. Press, N. A. Smith, and M. Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021.
  128. 128.A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever. Improving language understanding by generative pre-training. 2018.
  129. 129.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. 2019.
  130. 130.J. W. Rae, A. Potapenko, S. M. Jayakumar, T. P. Lillicrap, K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, et al. Compressive transformers for long-range sequence modelling. Advances in Neural Information Processing Systems, 33:6154–6158, 2020.
  131. 131.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020a. URL http://jmlr.org/papers/v21/20-074.html.
  132. 132.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020b.
  133. 133.N. F. Rajani, B. McCann, C. Xiong, and R. Socher. Explain yourself! leveraging language models for commonsense reasoning. arXiv preprint arXiv:1906.02361, 2019.
  134. 134.S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020.
  135. 135.P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas, Nov. 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1264. URL https://aclanthology.org/D16-1264.
  136. 136.P. Rajpurkar, R. Jia, and P. Liang. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-2124. URL https://aclanthology.org/P18-2124.
  137. 137.S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y. Liu, J. Xu, M. Ott, E. M. Smith, Y.-L. Boureau, and J. Weston. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online, Apr. 2021a. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.24. URL https://aclanthology.org/2021.eacl-main.24.
  138. 138.S. Roller, S. Sukhbaatar, A. Szlam, and J. E. Weston. Hash layers for large sparse models. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021b. URL https://openreview.net/forum?id=lMgDDWb1ULW.
  139. 139.C. Rosset. Turing-NLG: A 17-billion-parameter language model by Microsoft. Microsoft Blog, 1:2, 2020.
  140. 140.R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme. Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, New Orleans, Louisiana, June 2018. Association for Computational Linguistics.
  141. 141.S. Russell. Human Compatible. Penguin, 2020.
  142. 142.P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. Pierrehumbert. Hatecheck: Functional tests for hate speech detection models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021. doi: 10.18653/v1/2021.acl-long.4. URL http://dx.doi.org/10.18653/v1/2021.acl-long.4.
  143. 143.K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi. Winogrande: An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740, 2020.
  144. 144.V. Sanh, L. Debut, J. Chaumond, and T. Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108, 2019. URL http://arxiv.org/abs/1910.01108.
  145. 145.V. Sanh, T. Wolf, and A. M. Rush. Movement pruning: Adaptive sparsity by fine-tuning. arXiv preprint arXiv:2005.07683, 2020.
  146. 146.V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, S. Biderman, L. Gao, T. Bers, T. Wolf, and A. M. Rush. Multitask prompted training enables zero-shot task generalization, 2021.
  147. 147.M. Sap, H. Rashkin, D. Chen, R. LeBras, and Y. Choi. SocialIQA: Commonsense reasoning about social interactions. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019.
  148. 148.M. Schaarschmidt, D. Grewe, D. Vytiniotis, A. Paszke, G. Schmid, T. Norman, J. Molloy, J. Godwin, N. A. Rink, V. Nair, and D. Belov. Automap: Towards ergonomic automated parallelism for ml models. In ML for Systems Workshop at NeurIPS 2021, 2021.
  149. 149.T. Schick, S. Udupa, and H. Schütze. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. CoRR, abs/2103.00453, 2021. URL https://arxiv.org/abs/2103.00453.
  150. 150.A. See, M. Luong, and C. D. Manning. Compression of neural machine translation models via pruning. CoRR, abs/1606.09274, 2016. URL http://arxiv.org/abs/1606.09274.
  151. 151.A. D. Selbst, D. Boyd, S. A. Friedler, S. Venkatasubramanian, and J. Vertesi. Fairness and abstraction in sociotechnical systems. In Proceedings of the conference on fairness, accountability, and transparency, pages 59–68, 2019.
  152. 152.C. E. Shannon. A mathematical theory of communication. The Bell system technical journal, 27(3): 379–423, 1948.
  153. 153.N. Shazeer and M. Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018.
  154. 154.E. Sheng, K.-W. Chang, P. Natarajan, and N. Peng. The woman worked as a babysitter: On biases in language generation. EMNLP, 2019.
  155. 155.E. Sheng, K.-W. Chang, P. Natarajan, and N. Peng. Societal biases in language generation: Progress and challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4275–4293, Online, Aug. 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.330. URL https://aclanthology.org/2021.acl-long.330.
  156. 156.M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  157. 157.S. P. Singh and D. Alistarh. Woodfisher: Efficient second-order approximations for model compression. CoRR, abs/2004.14340, 2020. URL https://arxiv.org/abs/2004.14340.
  158. 158.D. So, Q. Le, and C. Liang. The evolved transformer. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 5877–5886. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/so19a.html.
  159. 159.D. R. So, W. Mańke, H. Liu, Z. Dai, N. Shazeer, and Q. V. Le. Primer: Searching for efficient transformers for language modeling, 2021.
  160. 160.A. Soleimani, C. Monz, and M. Worring. BERT for evidence retrieval and claim verification. In J. M. Jose, E. Yilmaz, J. Magalhães, P. Castells, N. Ferro, M. J. Silva, and F. Martins, editors, Advances in Information Retrieval, pages 359–366, Cham, 2020. Springer International Publishing. ISBN 978-3-030-45442-5.
  161. 161.J. Steinhardt. Updates and lessons from AI forecasting, 2021. URL https://bounded-regret.ghost.io/ai-forecasting/.
  162. 162.N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 3008–3021. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1f89885d556929e98d3ef9b86448f951-Paper.pdf.
  163. 163.S. Sun, K. Krishna, A. Mattarella-Micke, and M. Iyyer. Do long-range language models actually use long-range context? arXiv preprint arXiv:2109.09115, 2021.
  164. 164.J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819. Association for Computational Linguistics, June 2018. URL https://aclanthology.org/N18-1074.
  165. 165.A. Turing. Computing machinery and intelligence. Mind, 59(236):433–460, 1950.
  166. 166.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  167. 167.E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2153–2162, Hong Kong, China, Nov. 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1221. URL https://aclanthology.org/D19-1221.
  168. 168.Z. J. Wang, D. Choi, S. Xu, and D. Yang. Putting humans in the natural language processing loop: A survey. arXiv preprint arXiv:2103.04044, 2021.
  169. 169.L. R. Waugh. Marked and unmarked: A choice between unequals in semiotic structure. 1982.
  170. 170.J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  171. 171.L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel. Ethical and social risks of harm from language models. arXiv submission, 2021.
  172. 172.J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469, Punta Cana, Dominican Republic, Nov. 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.findings-emnlp.210.
  173. 173.A. Williams, N. Nangia, and S. Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1101. URL https://aclanthology.org/N18-1101.
  174. 174.J. G. Wolff. Language acquisition, data compression and generalization. Language & Communication, 2(1):57–89, 1982.
  175. 175.J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862, 2021.
  176. 176.A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2390–2397, Online, June 2021a. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.190. URL https://aclanthology.org/2021.naacl-main.190.
  177. 177.J. Xu, D. Ju, M. Li, Y.-L. Boureau, J. Weston, and E. Dinan. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online, June 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.235. URL https://aclanthology.org/2021.naacl-main.235.
  178. 178.L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. MT5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  179. 179.Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le. XLNet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  180. 180.T. Young, E. Cambria, I. Chaturvedi, H. Zhou, S. Biswas, and M. Huang. Augmenting end-to-end dialogue systems with commonsense knowledge. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), Apr. 2018. URL https://ojs.aaai.org/index.php/AAAI/article/view/11923.
  181. 181.D. Yu and K. Sagae. Automatically exposing problems with neural dialog models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 456–470, Online and Punta Cana, Dominican Republic, Nov. 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.emnlp-main.37.
  182. 182.D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, et al. Differentially private fine-tuning of language models. arXiv preprint arXiv:2110.06500, 2021.
  183. 183.O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat. Q8BERT: quantized 8bit BERT. CoRR, abs/1910.06188, 2019. URL http://arxiv.org/abs/1910.06188.
  184. 184.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  185. 185.B. Zhang and R. Sennrich. Root mean square layer normalization. arXiv preprint arXiv:1910.07467, 2019.
  186. 186.W. Zhong, J. Xu, D. Tang, Z. Xu, N. Duan, M. Zhou, J. Wang, and J. Yin. Reasoning over semantic-level graph for fact checking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6170–6180, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.549. URL https://aclanthology.org/2020.acl-main.549.
  187. 187.H. Zhou, T. Young, M. Huang, H. Zhao, J. Xu, and X. Zhu. Commonsense knowledge aware conversation generation with graph attention. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4623–4629. International Joint Conferences on Artificial Intelligence Organization, 7 2018. doi: 10.24963/ijcai.2018/643. URL https://doi.org/10.24963/ijcai.2018/643.
  188. 188.M. Zhu and S. Gupta. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878, 2017.
  189. 189.D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Rae, J. W., et al. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. arXiv, 2021, http://arxiv.org/abs/2112.11446v2.
APA
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., Driessche, G. van . den ., Hendricks, L. A., Rauh, M., Huang, P.-S., … Irving, G. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv. http://arxiv.org/abs/2112.11446v2
Chicago
Rae, J. W., S. Borgeaud, T. Cai, et al. 2021. “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”. arXiv. http://arxiv.org/abs/2112.11446v2.
Harvard
Rae, J.W. et al. (2021) “Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.11446v2.
Vancouver
1. Rae JW, Borgeaud S, Cai T, et al (2021) Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv

BibTeX

@article{rae2021scaling,
  title = {Scaling Language Models: Methods, Analysis & Insights from Training Gopher},
  author = {Rae, Jack W. and Borgeaud, Sebastian and Cai, Trevor and Millican, Katie and Hoffmann, Jordan and Song, Francis and Aslanides, John and Henderson, Sarah and Ring, Roman and Young, Susannah and Rutherford, Eliza and Hennigan, Tom and Menick, Jacob and Cassirer, Albin and Powell, Richard and Driessche, George van den and Hendricks, Lisa Anne and Rauh, Maribeth and Huang, Po-Sen and Glaese, Amelia and Welbl, Johannes and Dathathri, Sumanth and Huang, Saffron and Uesato, Jonathan and Mellor, John and Higgins, Irina and Creswell, Antonia and McAleese, Nat and Wu, Amy and Elsen, Erich and Jayakumar, Siddhant and Buchatskaya, Elena and Budden, David and Sutherland, Esme and Simonyan, Karen and Paganini, Michela and Sifre, Laurent and Martens, Lena and Li, Xiang Lorraine and Kuncoro, Adhiguna and Nematzadeh, Aida and Gribovskaya, Elena and Donato, Domenic and Lazaridou, Angeliki and Mensch, Arthur and Lespiau, Jean-Baptiste and Tsimpoukelli, Maria and Grigorev, Nikolai and Fritz, Doug and Sottiaux, Thibault and Pajarskas, Mantas and Pohlen, Toby and Gong, Zhitao and Toyama, Daniel and d'Autume, Cyprien de Masson and Li, Yujia and Terzi, Tayfun and Mikulik, Vladimir and Babuschkin, Igor and Clark, Aidan and Casas, Diego de Las and Guy, Aurelia and Jones, Chris and Bradbury, James and Johnson, Matthew and Hechtman, Blake and Weidinger, Laura and Gabriel, Iason and Isaac, William and Lockhart, Ed and Osindero, Simon and Rimell, Laura and Dyer, Chris and Vinyals, Oriol and Ayoub, Kareem and Stanway, Jeff and Bennett, Lorrayne and Hassabis, Demis and Kavukcuoglu, Koray and Irving, Geoffrey},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.11446v2},
  eprint = {2112.11446}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission