OLMo: Accelerating the Science of Language Models

Dirk GroeneveldIz BeltagyEvan Pete WalshAkshita BhagiaRodney KinneyOyvind TafjordAnanya Harsh JhaHamish IvisonIan MagnussonYizhong Wang

article2024ACL617 citationsBest Theme Paper Award

Introduces OLMo, a competitive suite of open language models released alongside its complete pretraining data, training logs, evaluation code, and intermediate checkpoints to enable transparent and reproducible scientific research.

Listen

As language models have rapidly grown in commercial importance, the most capable systems have become increasingly proprietary and closed off. Developers frequently withhold essential details about pretraining data, architectures, training configurations, and safety interventions. This lack of transparency limits the ability of the broader scientific and engineering community to rigorously study model mechanics, evaluate biases, address safety risks, and build upon existing innovations without repeating costly training runs.

The article demonstrates the creation and release of OLMo, a competitive, fully open language model framework. The core objective is to deliver a transparent, high-performing base for scientific inquiry by openly providing not only the model weights and code, but also the underlying multi-trillion-token training dataset, full training logs, intermediate checkpoints, and standardized evaluation tools under permissive open-source licensing.

To achieve this, the authors built a complete open-source pipeline using a decoder-only transformer architecture optimized for hardware throughput and training stability. They trained 1-billion-parameter (1B) and 7-billion-parameter (7B) model variants on over 2 trillion tokens drawn from Dolma, a curated open dataset spanning web pages, academic papers, books, and code. Training was executed across two distinct high-performance computing clusters using mixed-precision distributed strategies. The framework incorporates systematic checkpoints every 1,000 steps, automated evaluation harnesses, and adaptation pipelines utilizing supervised instruction tuning and preference alignment.

Key findings show that OLMo achieves performance parity with leading models in its size class. On a suite of eight core downstream reasoning tasks, the OLMo-7B model achieved an average accuracy of 69.3%, performing closely alongside closed-data peers such as Falcon-7B (70.3%) and Llama 2 7B (70.5%). Intrinsic language modeling evaluations on decontaminated benchmark text revealed that the model fits out-of-sample language distributions effectively, with sample efficiency heavily influenced by training data distribution. Furthermore, adapting the base model with supervised instruction tuning and direct preference optimization produced major capability gains: multitask language understanding accuracy surged from 28.3% to over 46%, while toxicity generation dropped sharply from 81.4% to 1.7%.

These results demonstrate that competitive, production-grade language models can be successfully built entirely from open, fully disclosed assets. This level of openness substantially reduces development and environmental costs by preventing redundant pretraining across organizations, which consumed an estimated 239 megawatt-hours during the 7B pretraining run. Full visibility into data and training checkpoints also enables researchers to audit models for safety compliance and accurately identify the root causes of model errors or biases.

Organizations and research teams should leverage the open OLMo framework to conduct reproducible research, perform targeted safety audits, and fine-tune domain-specific applications without bearing the full burden of training base models from scratch. Looking forward, development should focus on expanding the framework to multilingual datasets, diverse model sizes, and refined adaptation mixtures tailored specifically to OLMo's architectural strengths.

While the findings demonstrate high reliability and solid performance, certain limitations remain. Pretraining data is primarily in English and, despite extensive filtering, may still contain problematic web content. In addition, standard automated benchmarks offer noisy signals that do not fully capture interactive chat dynamics, meaning decision-makers should view automated comparisons as directional rather than definitive.

arXiv: 2402.00838allenai/OLMo
Cover for OLMo: Accelerating the Science of Language Models

Abstract

Language models (LMs) have become ubiquitous in both NLP research and in commercial product offerings. As their commercial importance has surged, the most powerful models have become closed off, gated behind proprietary interfaces, with important details of their training data, architectures, and development undisclosed. Given the importance of these details in scientifically studying these models, including their biases and potential risks, we believe it is essential for the research community to have access to powerful, truly open LMs. To this end, we have built OLMo, a competitive, truly Open Language Model, to enable the scientific study of language models. Unlike most prior efforts that have only released model weights and inference code, we release OLMo alongside open training data and training and evaluation code. We hope this release will empower the open research community and inspire a new wave of innovation.

Table of Contents

  • 1 Introduction
  • 2 OLMo Framework
  • 2.1 OLMo Model and Architecture
  • 2.2 Pretraining Data: Dolma
  • 2.3 Adaptation
  • 2.4 Evaluation
  • 3 Training OLMo
  • 3.1 Distributed Training Framework
  • 3.2 Optimizer
  • 3.3 Data
  • 3.4 Hardware
  • 4 Results
  • 4.1 Downstream evaluation
  • 4.2 Intrinsic language modeling evaluation
  • 4.3 Adaptation Evaluation
  • 5 Artifacts Released
  • 6 Conclusion and Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgments
  • References
  • A Training Settings
  • B Power Consumption and Carbon Footprint
  • C Additional Evaluation
  • D Adaptation Training Details
  • E Adaptation Evaluation and Model details

Knowls

  1. Knowl 1 — OLMo Model Architecture and Transformer Modifications

    model/method

    OLMo (Open Language Model) uses a decoder-only autoregressive Transformer architecture incorporating specific modifications designed to optimize training stability and hardware throughput compared to vanilla Transformer designs:

    1. No Bias Parameters: All bias terms are omitted from linear and normalization layers throughout the network to enhance optimization stability.
    2. Non-parametric Layer Normalization: Layer normalization is applied without learnable affine parameters (neither adaptive gain γ\gamma nor bias β\beta). This formulation avoids instability and achieves faster execution than parametric LayerNorm and RMSNorm.
    3. SwiGLU Activation Function: The feed-forward network uses the SwiGLU gated activation function in place of standard ReLU. The intermediate hidden dimension is set to approximately 83d\frac{8}{3}d (where dd is the model hidden dimension), rounded up to the nearest multiple of 128 (e.g., an intermediate dimension of 11,008 for d=4096d = 4096, meaning the linear projection input to SwiGLU has dimension 2×11,008=22,0162 \times 11,008 = 22,016).
    4. Rotary Position Embeddings (RoPE): Absolute positional embeddings are replaced by Rotary Position Embeddings applied to query and key representations.
    5. PII-Masked Vocabulary with Padded Embedding Matrix: The vocabulary uses a Byte-Pair Encoding (BPE) tokenizer derived from GPT-NeoX-20B, augmented with dedicated masking tokens for Personally Identifiable Information (PII) to yield 50,280 tokens. To optimize memory alignment and GPU throughput, the token embedding matrix is padded to 50,304 rows (the nearest multiple of 128).
  2. Knowl 2 — OLMo Architecture Dimensions and Pretraining Hyperparameters

    data/table

    OLMo was released in 1B and 7B parameter configurations trained on up to 2.46 trillion tokens using the AdamW optimizer. Key architectural dimensions, token budgets, and optimizer configurations are summarized below:

    Hyperparameter / Dimension OLMo-1B OLMo-7B
    Number of Layers (LL) 16 32
    Hidden Dimension (dd) 2048 4096
    Attention Heads (HH) 16 32
    MLP Intermediate Dimension 5504 (≈83d\approx \frac{8}{3}d) 11008 (≈83d\approx \frac{8}{3}d)
    Attention Variant Full Multi-Head Full Multi-Head
    Positional Embedding RoPE RoPE
    LayerNorm Type Non-parametric Non-parametric
    Biases None None
    Weight Tying Yes No
    Sequence Length (tokens) 2048 2048
    Global Batch Size (instances) 2048 2048
    Global Batch Size (tokens) ∼4M\sim 4\text{M} ∼4M\sim 4\text{M}
    Micro-batch Size (tokens/GPU) 4096 4096
    Training Tokens 2.0T2.0\text{T} 2.46T2.46\text{T}
    Peak Learning Rate 4.0×10−44.0 \times 10^{-4} 3.0×10−43.0 \times 10^{-4}
    Minimum Learning Rate 4.0×10−54.0 \times 10^{-5} 3.0×10−53.0 \times 10^{-5}
    Warmup Steps 2000 steps 5000 steps (∼21B\sim 21\text{B} tokens)
    Learning Rate Schedule Linear decay Linear decay
    Optimizer AdamW AdamW
    AdamW β1,β2,ϵ\beta_1, \beta_2, \epsilon 0.9, 0.95, 1.0×10−51.0 \times 10^{-5} 0.9, 0.95, 1.0×10−51.0 \times 10^{-5}
    Weight Decay 0.1 0.1
    Gradient Clipping Norm 1.0 1.0

    The 1B model ties input and output embedding weights, whereas the 7B model maintains untied embeddings. For the 7B model, warmup spans 5000 steps (∼21\sim 21 billion tokens), after which the learning rate decays linearly toward 10% of peak learning rate over the remainder of training.

  3. Knowl 3 — Distributed Training and Mixed Precision Strategy in OLMo

    model/method

    OLMo pretraining utilizes PyTorch Fully Sharded Data Parallel (FSDP) implementing the ZeRO optimizer strategy to shard model parameters, gradients, and optimizer states across GPUs. At the 7B scale, this allows a micro-batch size of 4096 tokens per GPU, aggregating to a constant global batch size of approximately 4M tokens (2048 sequences of length 2048).

    The precision configuration combines 16-bit computation with full 32-bit numerical stability:

    • Optimizer States and Sharded Weights: Maintained locally in single precision (FP32) on each GPU.
    • Forward and Backward Execution: Sharded weights are materialized and cast to bfloat16 within each transformer block during forward and backward passes.
    • Softmax Stability: Critical numerical operations, specifically the attention softmax, are executed in FP32 precision to prevent training instability and loss spikes.
    • Gradient Reduction: Gradients are accumulated and reduced across all GPUs in FP32.
    • Gradient Clipping: After the warmup phase, parameter gradients are flattened across the entire model into a single vector and clipped to a maximum global ℓ2\ell_2-norm threshold of 1.0.

    This setup was validated across two distinct hardware environments—LUMI (AMD MI250X dual-chip GPUs with 128GB memory) and MosaicML (NVIDIA A100 40GB GPUs)—producing nearly identical convergence curves by 2 trillion tokens.

  4. Knowl 4 — Dolma Multi-Source Pretraining Dataset

    data/table

    The pretraining corpus for OLMo is Dolma, an open dataset consisting of 2.668 trillion tokens across 4.367 billion documents (11,519 GB of UTF-8 text), tokenized using the modified GPT-NeoX BPE tokenizer:

    Source Data Type UTF-8 Volume (GB) Documents (M) Tokens (B)
    Common Crawl Web pages 9,812 3,734.0 2,180
    GitHub Source code 1,043 210.0 342
    Reddit Social media 339 377.0 80
    Semantic Scholar Academic papers 268 38.8 57
    Project Gutenberg Books 20.4 0.056 5.2
    Wikipedia Encyclopedic text 16.2 6.2 3.7
    Total 11,519 4,367.0 2,668

    Dolma is curated via a six-stage pipeline: (1) language filtering, (2) quality filtering using document-level heuristics, (3) content filtering to remove toxic material and PII, (4) deduplication, (5) multi-source mixing while preserving document source provenance, and (6) tokenization. Training instances are constructed by concatenating tokenized documents separated by an EOS token into non-overlapping chunks of 2048 tokens and applying a deterministic shuffling order.

  5. Knowl 5 — Terminal Linear Learning Rate Cooldown Effect

    empirical result

    When pretraining OLMo-7B up to 2.46 trillion tokens using a linear learning rate decay down to 10% of peak learning rate (3.0×10−53.0 \times 10^{-5}), extending training by an additional 1,000 steps (approximately 4 billion tokens) while linearly decaying the learning rate to 0 yields a marked performance improvement across both downstream task accuracy and Paloma language modeling perplexity.

    During in-loop evaluation on 8 core downstream benchmarks (ARC-Easy, ARC-Challenge, BoolQ, HellaSwag, OpenBookQA, PIQA, SciQ, and Winogrande), tasks exhibit a sharp upward inflection in accuracy over these final 1,000 cooldown steps, demonstrating that annealing the learning rate fully to zero provides a significant boost to model capabilities.

  6. Knowl 6 — Zero-Shot Downstream Task Evaluation of OLMo Models

    data/table

    Zero-shot downstream evaluation of OLMo-1B (at 2.0T tokens) and OLMo-7B (at 2.46T tokens) across 8 core commonsense reasoning and question answering tasks, compared against publicly available base language models:

    Model ARC-c ARC-e BoolQ HellaSwag OBQA PIQA SciQ WinoG Average
    StableLM 1.6B 43.8 63.7 76.6 68.2 45.8 74.0 94.7 64.9 66.5
    Pythia 1B 33.1 50.2 61.8 44.7 37.8 69.1 86.0 53.3 54.5
    TinyLlama 1.1B 34.8 53.2 64.6 58.7 43.6 71.1 90.5 58.9 59.4
    OLMo-1B 34.5 58.1 60.7 62.5 46.4 73.7 88.1 58.9 60.4
    Falcon-7B 47.5 70.4 74.6 75.9 53.0 78.5 93.9 68.9 70.3
    LLaMA-7B 44.5 67.9 75.4 76.2 51.2 77.2 93.9 70.5 69.6
    Llama 2-7B 48.5 69.5 80.2 76.8 48.4 76.7 94.5 69.4 70.5
    MPT-7B 46.5 70.5 74.2 77.6 48.6 77.3 93.7 69.9 69.8
    Pythia-6.9B 44.1 61.9 61.1 63.8 45.0 75.1 91.1 62.0 63.0
    RPJ-INCITE-7B 42.8 68.4 68.6 70.3 49.4 76.0 92.9 64.7 66.6
    OLMo-7B 48.5 65.4 73.4 76.4 50.4 78.4 93.8 67.9 69.3

    Evaluation is conducted using rank classification with dataset-specific likelihood normalizations:

    • Unconditional Normalization: Applied to ARC-Challenge (ARC-c), ARC-Easy (ARC-e), and OpenBookQA (OBQA).
    • Per-Token Length Normalization: Applied to HellaSwag, PIQA, and WinoGrande (WinoG).
    • No Normalization (Raw Likelihood): Applied to BoolQ and SciQ (single-token prediction tasks).

    OLMo-7B achieves an average score of 69.3%, demonstrating performance competitive with LLaMA-7B (69.6%), MPT-7B (69.8%), and Llama 2-7B (70.5%).

  7. Knowl 7 — Intrinsic Perplexity Evaluation on Paloma Benchmark with Decontamination

    empirical result

    Intrinsic evaluation of OLMo-7B was performed using the Paloma perplexity benchmark across 585 text domains drawn from 18 data sources. To ensure rigorous evaluation without out-of-sample overestimation, OLMo-7B is explicitly decontaminated by removing any training document containing paragraphs overlapping with Paloma evaluation data.

    Evaluating bits per byte across 11 decontaminated public sources reveals:

    1. In-Distribution Alignment: Sample efficiency strongly depends on pretraining source distribution. OLMo-7B achieves the best bits per byte on Common Crawl-dominated sources such as C4 (reflecting Dolma's 88.8% web crawl composition), while models trained specifically on RefinedWeb (Falcon-7B) perform best on that specific split.
    2. Curated Domain Scarcity: On specialized text sources with lower web crawl representation (e.g., WikiText-103, M2D2 S2ORC, M2D2 Wikipedia), OLMo-7B shows lower sample efficiency compared to models trained on higher proportions of academic or encyclopedic text.
    3. Document Length Sensitivity: Evaluation on informal short-text sources (such as TwitterAAE and Gab) is heavily skewed by short average document lengths, resulting in elevated bits per byte across all models.
  8. Knowl 8 — Adaptation of OLMo via Supervised Fine-Tuning and Direct Preference Optimization

    model/method

    OLMo-7B is adapted into an instruction-following chat assistant through a two-stage post-training pipeline following the TÜLU 2 framework:

    1. Supervised Fine-Tuning (OLMo+SFT):
      • Dataset: TÜLU v2 SFT dataset mixture, split into 2048-token chunks with conversational prompt formatting.
      • Hyperparameters: Learning rate =2×10−6= 2 \times 10^{-6}, 3 epochs, linear warmup over the first 3% of training steps followed by linear cooldown to 0, maximum sequence length 2048 tokens, weight decay =0= 0, gradient clipping =0= 0.
    2. Direct Preference Optimization (OLMo+SFT+DPO):
      • Dataset: UltraFeedback preference dataset (using Argilla's cleaned binarized split scored by GPT-4 aspect ratings, with TruthfulQA prompts removed to prevent evaluation contamination).
      • Hyperparameters: Learning rate =5×10−7= 5 \times 10^{-7}, DPO implicit reward coefficient β=0.1\beta = 0.1, 3 epochs, linear warmup over the first 10% of training steps followed by linear cooldown to 0, maximum sequence length 2048 tokens, weight decay =0= 0, gradient clipping =0= 0.
  9. Knowl 9 — Chat, Safety, and Truthfulness Performance of Adapted OLMo-7B

    data/table

    Evaluation of OLMo-7B before and after supervised instruction tuning (SFT) and Direct Preference Optimization (DPO), compared against other adapted 7B-scale models on standard capability, chat, and safety benchmarks:

    Model MMLU (0-shot) ↑\uparrow AlpacaEval (% win) ↑\uparrow ToxiGen (% Toxic) ↓\downarrow TruthfulQA (% Info+True) ↑\uparrow
    OLMo-7B (base) 28.3 — 81.4 31.6
    MPT-7B Chat 33.8 46.8 0.1 42.7
    Falcon-7B Instruct 25.2 14.0 70.7 27.2
    RPJ-INCITE-7B Chat 27.0 38.0 46.4 53.0
    Llama-2-7B-Chat 46.8 87.3 0.0 26.3
    TÜLU 2 (Llama-2-7B) 50.4 73.9 7.0 51.7
    TÜLU 2 + DPO 50.7 85.1 0.5 —
    OLMo-7B + SFT 47.3 57.0 14.4 41.2
    OLMo-7B + SFT + DPO 46.2 69.3 1.7 52.0

    Across the full TÜLU evaluation suite, adaptation produces the following effects on OLMo-7B:

    • Knowledge and Reasoning: MMLU accuracy increases from 28.3% (base) to 47.3% (+SFT); GSM8k 8-shot CoT improves from 8.5% to 15.5% (+SFT); BBH 3-shot CoT rises from 31.7% to 36.9% (+SFT); Codex-Eval Pass@10 increases from 21.4% to 28.6% (+SFT).
    • Safety and Alignment: Toxic generation on ToxiGen drops dramatically from 81.4% to 1.7% (+SFT+DPO), while TruthfulQA (% Informative + True) increases from 31.6% to 52.0%.
    • Conversational Capability: AlpacaEval win-rate against Davinci-003 improves to 69.3% after SFT and DPO alignment.
  10. Knowl 10 — Power Consumption and Pretraining Carbon Emissions of OLMo-7B

    data/table

    Operational energy consumption and carbon emissions for pretraining OLMo-7B were calculated by measuring node power every 25ms, scaling by a Power Usage Effectiveness (PUE) factor of 1.1, and applying regional grid carbon intensity factors:

    Model Hardware Power (MWh) PUE Carbon Intensity (kg CO2e/kWh\text{kg CO}_2\text{e/kWh}) Emissions (tCO2eq\text{tCO}_2\text{eq})
    Gopher-280B TPU v3 1,066 1.08 0.330 380
    BLOOM-176B A100-80GB 433 1.20 0.057 30
    OPT-175B A100-80GB 324 1.10 0.231 82
    T5-11B TPU v3 77 1.12 0.545 47
    LLaMA-7B A100-80GB 33 1.10 0.385 14
    LLaMA2-7B A100-80GB 74 1.10 0.385 31
    OLMo-7B (LUMI) AMD MI250X 135 1.10 0.000* 0*
    OLMo-7B (MosaicML) NVIDIA A100-40GB 104 1.10 0.610 70

    Total pretraining power across both runs was approximately 239 MWh (135 MWh on LUMI and 104 MWh on MosaicML). LUMI operates on 100% renewable hydroelectric power, yielding an operational carbon intensity factor of 0.000 kg CO2e/kWh0.000\text{ kg CO}_2\text{e/kWh} (0 tCO2eq0\text{ tCO}_2\text{eq}, or 3.54 tCO2eq3.54\text{ tCO}_2\text{eq} if using a lifecycle factor of 0.024 kg CO2e/kWh0.024\text{ kg CO}_2\text{e/kWh}). The MosaicML cluster in Australia operated with a grid intensity of 0.610 kg CO2e/kWh0.610\text{ kg CO}_2\text{e/kWh}, generating 70 tCO2eq70\text{ tCO}_2\text{eq}.

  11. Knowl 11 — Task Stability Limitations in Online Pretraining Tracking

    empirical result

    When tracking language model learning progress during pretraining, downstream evaluation tasks exhibit stark differences in signal stability. While core commonsense reasoning tasks (ARC, BoolQ, HellaSwag, PIQA, SciQ) provide stable, monotonic upward accuracy trajectories over 2.46T tokens, auxiliary NLP benchmarks exhibit major instability:

    1. Random and Erratic Progress: Tasks such as HEAD-QA, LogiQA, QNLI, and WNLI fluctuate erratically across training checkpoints, failing to provide consistent learning signals.
    2. Near-Chance Performance: Tasks like Words in Context (WiC) remain stuck near random guessing (50% accuracy) throughout the entire pretraining trajectory regardless of token count.
    3. Class Imbalance Artifacts: On benchmarks with class imbalances (such as MRPC), language models frequently collapse into predicting a single majority label, artificially inflating or deflating evaluation accuracy without reflecting underlying linguistic capability.

Coverage note — None was omitted; the extracted knowls comprehensively cover the OLMo architecture, hyperparameters, distributed training and precision framework, Dolma dataset composition and pipeline, learning rate cooldown dynamics, downstream task evaluations, Paloma decontaminated perplexity evaluations, adaptation via SFT and DPO, carbon footprint metrics, and evaluation task stability analyses.

References

  1. 1.Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos. 2023. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540.
  2. 2.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra-Aimée Cojocaru, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. The falcon series of open language models. ArXiv, abs/2311.16867.
  3. 3.Yuvanesh Anand, Zach Nussbaum, Brandon Duderstadt, Benjamin Schmidt, and Andriy Mulyar. 2023. Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo. https://github.com/nomic-ai/gpt4all.
  4. 4.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. ArXiv, abs/1607.06450.
  5. 5.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  6. 6.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A neural probabilistic language model. J. Mach. Learn. Res., 3:1137–1155.
  7. 7.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 2397–2430. PMLR.
  8. 8.BigScience, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  9. 9.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439.
  10. 10.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models.
  11. 11.Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic dialectal variation in social media: A case study of African-American English. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1119–1130, Austin, Texas. Association for Computational Linguistics.
  12. 12.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  13. 13.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  14. 14.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  15. 15.Alexandra Chronopoulou, Matthew Peters, and Jesse Dodge. 2022. Efficient hierarchical domain adaptation for pretrained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1336–1351, Seattle, United States. Association for Computational Linguistics.
  16. 16.Hyung Won Chung, Noah Constant, Xavier García, Adam Roberts, Yi Tay, Sharan Narang, and Orhan Firat. 2023. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining. ArXiv, abs/2304.09151.
  17. 17.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  18. 18.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  19. 19.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  20. 20.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback.
  21. 21.Jesse Dodge, Taylor Prewitt, Remi Tachet Des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan. 2022. Measuring the carbon intensity of ai in cloud instances.
  22. 22.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In International Joint Conference on Natural Language Processing.
  23. 23.Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge. 2024. What’s in my big data? In The Twelfth International Conference on Learning Representations.
  24. 24.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  25. 25.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023. A framework for few-shot language model evaluation.
  26. 26.Sidney Greenbaum and Gerald Nelson. 1996. The international corpus of english (ICE) project. World Englishes, 15(1):3–15.
  27. 27.Dirk Groeneveld, Anas Awadalla, Iz Beltagy, Akshita Bhagia, Ian Magnusson, Hao Peng, Oyvind Tafjord, Pete Walsh, Kyle Richardson, and Jesse Dodge. 2023. Catwalk: A unified language model evaluation framework for many datasets. arXiv preprint arXiv:2312.10253.
  28. 28.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv preprint arxiv:2301.07597.
  29. 29.Suchin Gururangan, Mitchell Wortsman, Samir Yitzhak Gadre, Achal Dave, Maciej Kilian, Weijia Shi, Jean Mercat, Georgios Smyrnis, Gabriel Ilharco, Matt Jordan, Reinhard Heckel, Alex Dimakis, Ali Farhadi, Vaishaal Shankar, and Ludwig Schmidt. 2023. OpenLM: a minimal but performative language modeling (lm) repository. GitHub repository.
  30. 30.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. TOXIGEN: Controlling Language Models to Generate Implied and Adversarial Toxicity. In ACL.
  31. 31.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
  32. 32.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2.
  33. 33.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  34. 34.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Minh Nguyen, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Alexandrovich Glushkov, Arnav Varma Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Julian Mattick. 2023. Openassistant conversations - democratizing large language model alignment. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  35. 35.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models. Github repository.
  36. 36.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  37. 37.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252.
  38. 38.Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. CoRR, abs/2007.08124.
  39. 39.Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550.
  40. 40.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  41. 41.Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2022. Estimating the carbon footprint of bloom, a 176b parameter language model.
  42. 42.Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, et al. 2023. Paloma: A benchmark for evaluating language model fit. arXiv preprint arXiv:2312.10523.
  43. 43.Mitchell P. Marcus, Beatrice Santorini, Mary Ann Marcinkiewicz, and Ann Taylor. 1999. Treebank-3.
  44. 44.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. ArXiv, abs/1609.07843.
  45. 45.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Frederick Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2017. Mixed precision training. ArXiv, abs/1710.03740.
  46. 46.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789.
  47. 47.Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Neural Information Processing Systems.
  48. 48.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  49. 49.MosaicML NLP Team. 2023. Introducing mpt-7b: A new standard for open-source, commercially usable llms. Accessed: 2023-05-05.
  50. 50.Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023. Scaling data-constrained language models. arXiv preprint arXiv:2305.16264.
  51. 51.Davide Nunes. 2020. Preprocessed penn tree bank.
  52. 52.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  53. 53.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  54. 54.Antonis Papasavva, Savvas Zannettou, Emiliano De Cristofaro, Gianluca Stringhini, and Jeremy Blackburn. 2020. Raiders of the lost kek: 3.5 years of augmented 4chan posts from the politically incorrect board. Proceedings of the International AAAI Conference on Web and Social Media, 14:885–894.
  55. 55.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training.
  56. 56.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra-Aimée Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only. ArXiv, abs/2306.01116.
  57. 57.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. ArXiv, abs/1802.05365.
  58. 58.Mohammad Taher Pilehvar and José Camacho-Collados. 2018. Wic: 10, 000 example pairs for evaluating context-sensitive representations. CoRR, abs/1808.09121.
  59. 59.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2022. Scaling language models: Methods, analysis & insights from training gopher.
  60. 60.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems.
  61. 61.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1).
  62. 62.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2019. Zero: Memory optimizations toward training trillion parameter models. SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16.
  63. 63.Machel Reid, Victor Zhong, Suchin Gururangan, and Luke Zettlemoyer. 2022. M2D2: A massively multidomain language modeling dataset. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 964–975, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  64. 64.Manoel Horta Ribeiro, Jeremy Blackburn, Barry Bradlyn, Emiliano De Cristofaro, Gianluca Stringhini, Summer Long, Stephanie Greenberg, and Savvas Zannettou. 2021. The evolution of the manosphere across the web. Proceedings of the International AAAI Conference on Web and Social Media, 15:196–207.
  65. 65.Ronald Rosenfeld. 2000. Two decades of statistical language modeling: Where do we go from here? Proceedings of the IEEE, 88(8):1270–1278.
  66. 66.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.
  67. 67.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  68. 68.Noam M. Shazeer. 2020. Glu variants improve transformer. ArXiv, abs/2002.05202.
  69. 69.Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. 2024. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv preprint.
  70. 70.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645–3650, Florence, Italy. Association for Computational Linguistics.
  71. 71.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. ArXiv, abs/2104.09864.
  72. 72.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  73. 73.Teknium1. 2023. Gpteacher. https://github.com/teknium1/GPTeacher.
  74. 74.Together Computer. 2023. RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset.
  75. 75.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971.
  76. 76.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  77. 77.María Ubierna, Cristina Díez Santos, and Sara Mercier-Blais. 2022. Water Security and Climate Change: Hydropower Reservoir Greenhouse Gas Emissions, pages 69–94. Springer Singapore, Singapore.
  78. 78.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  79. 79.David Vilares and Carlos Gómez-Rodríguez. 2019. HEAD-QA: A healthcare dataset for complex reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 960–966, Florence, Italy. Association for Computational Linguistics.
  80. 80.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. ArXiv, abs/1804.07461.
  81. 81.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. How far can camels go? exploring the state of instruction tuning on open resources.
  82. 82.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  83. 83.Johannes Welbl, Nelson F Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. arXiv preprint arXiv:1707.06209.
  84. 84.Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga Behram, James Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin S. Lee, Bugra Akyildiz, Maximilian Balandat, Joe Spisak, Ravi Jain, Mike Rabbat, and Kim Hazelwood. 2022. Sustainable ai: Environmental implications, challenges and opportunities.
  85. 85.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024. WizardLM: Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations.
  86. 86.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. arXiv preprint arXiv:2304.01196.
  87. 87.Savvas Zannettou, Barry Bradlyn, Emiliano De Cristofaro, Haewoon Kwak, Michael Sirivianos, Gianluca Stringini, and Jeremy Blackburn. 2018. What is gab: A bastion of free speech or an alt-right echo chamber. In Companion Proceedings of the The Web Conference 2018, WWW ’18, page 1007–1014, Republic and Canton of Geneva, CHE. International World Wide Web Conferences Steering Committee.
  88. 88.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830.
  89. 89.Biao Zhang and Rico Sennrich. 2019. Root mean square layer normalization. ArXiv, abs/1910.07467.
  90. 90.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pretrained transformer language models.
  91. 91.Yanli Zhao, Andrew Gu, Rohan Varma, Liangchen Luo, Chien chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, and Shen Li. 2023. Pytorch fsdp: Experiences on scaling fully sharded data parallel. Proc. VLDB Endow., 16:3848–3860.

Citation

MLA
Groeneveld, D., et al. “OLMo: Accelerating the Science of Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 15789–809, https://doi.org/10.18653/v1/2024.acl-long.841.
APA
Groeneveld, D., Beltagy, I., Walsh, E., Bhagia, A., Kinney, R., Tafjord, O., Jha, A., Ivison, H., Magnusson, I., Wang, Y., Arora, S., Atkinson, D., Authur, R., Chandu, K., Cohan, A., Dumas, J., Elazar, Y., Gu, Y., Hessel, J., … Hajishirzi, H. (2024). OLMo: Accelerating the Science of Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15789–15809. https://doi.org/10.18653/v1/2024.acl-long.841
Chicago
Groeneveld, D., I. Beltagy, E. Walsh, et al. 2024. “OLMo: Accelerating the Science of Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15789–809. https://doi.org/10.18653/v1/2024.acl-long.841.
Harvard
Groeneveld, D. et al. (2024) “OLMo: Accelerating the Science of Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15789–15809. Available at: https://doi.org/10.18653/v1/2024.acl-long.841.
Vancouver
1. Groeneveld D, Beltagy I, Walsh E, et al (2024) OLMo: Accelerating the Science of Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15789–15809

BibTeX

@inproceedings{groeneveld-etal-2024-olmo,
    title = "{OLM}o: Accelerating the Science of Language Models",
    author = "Groeneveld, Dirk  and
      Beltagy, Iz  and
      Walsh, Evan  and
      Bhagia, Akshita  and
      Kinney, Rodney  and
      Tafjord, Oyvind  and
      Jha, Ananya  and
      Ivison, Hamish  and
      Magnusson, Ian  and
      Wang, Yizhong  and
      Arora, Shane  and
      Atkinson, David  and
      Authur, Russell  and
      Chandu, Khyathi  and
      Cohan, Arman  and
      Dumas, Jennifer  and
      Elazar, Yanai  and
      Gu, Yuling  and
      Hessel, Jack  and
      Khot, Tushar  and
      Merrill, William  and
      Morrison, Jacob  and
      Muennighoff, Niklas  and
      Naik, Aakanksha  and
      Nam, Crystal  and
      Peters, Matthew  and
      Pyatkin, Valentina  and
      Ravichander, Abhilasha  and
      Schwenk, Dustin  and
      Shah, Saurabh  and
      Smith, William  and
      Strubell, Emma  and
      Subramani, Nishant  and
      Wortsman, Mitchell  and
      Dasigi, Pradeep  and
      Lambert, Nathan  and
      Richardson, Kyle  and
      Zettlemoyer, Luke  and
      Dodge, Jesse  and
      Lo, Kyle  and
      Soldaini, Luca  and
      Smith, Noah  and
      Hajishirzi, Hannaneh",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.841/",
    doi = "10.18653/v1/2024.acl-long.841",
    pages = "15789--15809"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/