The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only
Guilherme PenedoQuentin MalarticDaniel HesslowRuxandra CojocaruHamza AlobeidliAlessandro CappelliBaptiste PannierEbtesam AlmazroueiJulien Launay
Demonstrates that language models pretrained solely on heavily filtered and deduplicated CommonCrawl web data can outperform models trained on specialized multi-source corpora like The Pile, while releasing a massive five-trillion-token dataset pipeline to scale pretraining.
Modern artificial intelligence models require trillions of words of training text to achieve top capabilities, leading to industry concerns that high-quality text data will soon run out. The prevailing consensus has been that web data alone is too low in quality, forcing developers to rely on labor-intensive, curated collections of books, research papers, and conversations to train capable models. The article evaluates whether rigorously cleaned and deduplicated web data alone can produce language models that match or outperform those trained on curated corpora.
To test this, the authors developed a specialized data processing pipeline called MacroData Refinement (MDR) and applied it to raw Common Crawl web archives. The pipeline combines web address filtering, precise text extraction, language identification, rule-based text cleaning, and extensive exact and fuzzy deduplication. Across the full processing pipeline, roughly 90% of the raw documents were discarded. This process yielded RefinedWeb, a five-trillion-token English dataset. The authors trained language models ranging from 1 billion to 7.5 billion parameters and evaluated their performance across multiple standard language benchmarks without task-specific training.
Key findings show that models trained exclusively on RefinedWeb consistently outperform models trained on prominent curated datasets like The Pile, as well as models trained on other public web datasets like C4 and OSCAR. At equivalent compute budgets, models trained on RefinedWeb matched the performance of models trained on proprietary curated corpora such as the GPT-3 series. The analysis revealed that deduplication provides a steady, reliable performance improvement across all datasets by removing repetitive spans and templated text, whereas filtering heuristics yield less consistent gains across different data sources. Furthermore, the 3-billion-parameter model trained on poorly deduplicated web data underperformed a 1-billion-parameter model trained on RefinedWeb, showing that superior data quality can offset a fourfold compute disadvantage.
These findings indicate that human-intensive curation of specialized sources is not strictly necessary to build top-tier natural language foundation models. Organizations can significantly streamline their training pipelines, reduce data collection costs, and avoid the licensing risks tied to scraping copyrighted or specialized material. However, the study focuses strictly on pretraining for general natural language tasks and does not address specialized tasks such as computer programming or advanced mathematics, nor downstream fine-tuning. Practitioners developing production models should consider combining deduplicated web data with dedicated code repositories and instruction-tuning datasets, while leveraging thorough deduplication across their existing training assets.
- Paper: The Pile: An 800GB Dataset of Diverse Text for Language Modeling, Leo Gao et al. (2020). Introduces The Pile, establishing the curated multi-source paradigm that RefinedWeb directly benchmarks against and challenges using filtered web data alone.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Introduces the C4 dataset and foundational web-filtering heuristics on Common Crawl that the MacroData Refinement pipeline builds upon and refines.
- Paper: Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection, Suchin Gururangan et al. (2022). Analyzes the biases and assumptions inherent in standard web-text quality filters, providing essential context for designing objective web-filtering pipelines.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Establishes compute-optimal scaling laws that define the token volume and compute budgets evaluated in the RefinedWeb experiments.
- Paper: Scaling Language Models: Methods, Analysis & Insights from Training Gopher, Jack W. Rae et al. (2021). Details the construction and multi-source curation of MassiveText, serving as a primary curated dataset predecessor to pure web-scale corpora.
- Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). Pioneers the use of filtered web text at scale with WebText, motivating subsequent research into large-scale web-only pretraining datasets.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). Builds a standardized multi-stage benchmark (DCLM) to systematically test text extraction, deduplication, and quality filtering heuristics beyond RefinedWeb.
- Paper: Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research, Luca Soldaini et al. (2024). Extends large-scale open pretraining corpus design by combining systematic web filtering and deduplication with open curation tooling across three trillion tokens.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). Critically re-evaluates the necessity of data filtering, directly benchmarking against RefinedWeb to show that unfiltered web data can surpass filtered sets at massive scale.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). Proposes model-aware datamodels for web data selection, moving beyond the heuristic-driven filtering and deduplication methods established in RefinedWeb.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). Investigates how to allocate compute and repeat web training data across multiple epochs when unique high-quality text becomes constrained.
- Paper: OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents, Hugo Laurençon et al. (2023). Applies scalable web-filtering and document-cleaning principles to construct OBELICS, an open web-scale dataset of interleaved image-text documents.
- Paper: Textbooks Are All You Need, Suriya Gunasekar et al. (2023). Explores an alternative high-quality data paradigm by using synthetic, textbook-quality data rather than relying purely on massive filtered web scrapes.
