OLMo: Accelerating the Science of Language Models
Dirk GroeneveldIz BeltagyEvan Pete WalshAkshita BhagiaRodney KinneyOyvind TafjordAnanya Harsh JhaHamish IvisonIan MagnussonYizhong Wang
Introduces OLMo, a competitive suite of open language models released alongside its complete pretraining data, training logs, evaluation code, and intermediate checkpoints to enable transparent and reproducible scientific research.
As language models have rapidly grown in commercial importance, the most capable systems have become increasingly proprietary and closed off. Developers frequently withhold essential details about pretraining data, architectures, training configurations, and safety interventions. This lack of transparency limits the ability of the broader scientific and engineering community to rigorously study model mechanics, evaluate biases, address safety risks, and build upon existing innovations without repeating costly training runs.
The article demonstrates the creation and release of OLMo, a competitive, fully open language model framework. The core objective is to deliver a transparent, high-performing base for scientific inquiry by openly providing not only the model weights and code, but also the underlying multi-trillion-token training dataset, full training logs, intermediate checkpoints, and standardized evaluation tools under permissive open-source licensing.
To achieve this, the authors built a complete open-source pipeline using a decoder-only transformer architecture optimized for hardware throughput and training stability. They trained 1-billion-parameter (1B) and 7-billion-parameter (7B) model variants on over 2 trillion tokens drawn from Dolma, a curated open dataset spanning web pages, academic papers, books, and code. Training was executed across two distinct high-performance computing clusters using mixed-precision distributed strategies. The framework incorporates systematic checkpoints every 1,000 steps, automated evaluation harnesses, and adaptation pipelines utilizing supervised instruction tuning and preference alignment.
Key findings show that OLMo achieves performance parity with leading models in its size class. On a suite of eight core downstream reasoning tasks, the OLMo-7B model achieved an average accuracy of 69.3%, performing closely alongside closed-data peers such as Falcon-7B (70.3%) and Llama 2 7B (70.5%). Intrinsic language modeling evaluations on decontaminated benchmark text revealed that the model fits out-of-sample language distributions effectively, with sample efficiency heavily influenced by training data distribution. Furthermore, adapting the base model with supervised instruction tuning and direct preference optimization produced major capability gains: multitask language understanding accuracy surged from 28.3% to over 46%, while toxicity generation dropped sharply from 81.4% to 1.7%.
These results demonstrate that competitive, production-grade language models can be successfully built entirely from open, fully disclosed assets. This level of openness substantially reduces development and environmental costs by preventing redundant pretraining across organizations, which consumed an estimated 239 megawatt-hours during the 7B pretraining run. Full visibility into data and training checkpoints also enables researchers to audit models for safety compliance and accurately identify the root causes of model errors or biases.
Organizations and research teams should leverage the open OLMo framework to conduct reproducible research, perform targeted safety audits, and fine-tune domain-specific applications without bearing the full burden of training base models from scratch. Looking forward, development should focus on expanding the framework to multilingual datasets, diverse model sizes, and refined adaptation mixtures tailored specifically to OLMo's architectural strengths.
While the findings demonstrate high reliability and solid performance, certain limitations remain. Pretraining data is primarily in English and, despite extensive filtering, may still contain problematic web content. In addition, standard automated benchmarks offer noisy signals that do not fully capture interactive chat dynamics, meaning decision-makers should view automated comparisons as directional rather than definitive.
- Paper: Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research, Luca Soldaini et al. (2024). Dolma provides the foundational three-trillion-token open pretraining dataset and curation pipeline directly utilized to train the OLMo language model.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). Pythia establishes the methodological framework for fully transparent, intermediate-checkpoint language model analysis that OLMo expands to a larger, fully open ecosystem.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). LLaMA defines the modern open-weight foundation model architectural paradigm and token scaling ratios that OLMo aims to make completely open across data and training code.
- Paper: OPT: Open Pre-trained Transformer Language Models, Susan Zhang et al. (2022). OPT provides an early baseline in the pursuit of openly available large language models and reproducible training dynamics that OLMo succeeds.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). BLOOM models the collaborative, open-science approach to developing large language models and openly releasing training corpora that motivates OLMo's mission.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). Llama 2 represents the major open-access foundation baseline against which OLMo benchmarks its performance and openness standards.
- Paper: Tulu 3: Pushing Frontiers in Open Language Model Post-Training, Nathan Lambert et al. (2024). Tulu 3 builds directly on open base language model foundations by introducing a fully transparent, state-of-the-art post-training and alignment pipeline.
- Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). This paper investigates post-training memorization and data extraction vulnerabilities specifically using open foundation architectures like OLMo.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). OpenCoder extends the philosophy of full training data and pipeline transparency pioneered by OLMo into the domain of code-specialized large language models.
