The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only

Guilherme PenedoQuentin MalarticDaniel HesslowRuxandra CojocaruHamza AlobeidliAlessandro CappelliBaptiste PannierEbtesam AlmazroueiJulien Launay

article2023NeurIPS178 citations

Demonstrates that language models pretrained solely on heavily filtered and deduplicated CommonCrawl web data can outperform models trained on specialized multi-source corpora like The Pile, while releasing a massive five-trillion-token dataset pipeline to scale pretraining.

Listen

Modern artificial intelligence models require trillions of words of training text to achieve top capabilities, leading to industry concerns that high-quality text data will soon run out. The prevailing consensus has been that web data alone is too low in quality, forcing developers to rely on labor-intensive, curated collections of books, research papers, and conversations to train capable models. The article evaluates whether rigorously cleaned and deduplicated web data alone can produce language models that match or outperform those trained on curated corpora.

To test this, the authors developed a specialized data processing pipeline called MacroData Refinement (MDR) and applied it to raw Common Crawl web archives. The pipeline combines web address filtering, precise text extraction, language identification, rule-based text cleaning, and extensive exact and fuzzy deduplication. Across the full processing pipeline, roughly 90% of the raw documents were discarded. This process yielded RefinedWeb, a five-trillion-token English dataset. The authors trained language models ranging from 1 billion to 7.5 billion parameters and evaluated their performance across multiple standard language benchmarks without task-specific training.

Key findings show that models trained exclusively on RefinedWeb consistently outperform models trained on prominent curated datasets like The Pile, as well as models trained on other public web datasets like C4 and OSCAR. At equivalent compute budgets, models trained on RefinedWeb matched the performance of models trained on proprietary curated corpora such as the GPT-3 series. The analysis revealed that deduplication provides a steady, reliable performance improvement across all datasets by removing repetitive spans and templated text, whereas filtering heuristics yield less consistent gains across different data sources. Furthermore, the 3-billion-parameter model trained on poorly deduplicated web data underperformed a 1-billion-parameter model trained on RefinedWeb, showing that superior data quality can offset a fourfold compute disadvantage.

These findings indicate that human-intensive curation of specialized sources is not strictly necessary to build top-tier natural language foundation models. Organizations can significantly streamline their training pipelines, reduce data collection costs, and avoid the licensing risks tied to scraping copyrighted or specialized material. However, the study focuses strictly on pretraining for general natural language tasks and does not address specialized tasks such as computer programming or advanced mathematics, nor downstream fine-tuning. Practitioners developing production models should consider combining deduplicated web data with dedicated code repositories and instruction-tuning datasets, while leveraging thorough deduplication across their existing training assets.

Cover for The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only

Abstract

Large language models are commonly trained on a mixture of filtered web data and curated “high-quality” corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce performant models with broad zero-shot generalization abilities. However, as larger models requiring pretraining on trillions of tokens are considered, it is unclear how scalable is curation, and whether we will run out of unique high-quality data soon. At variance with previous beliefs, we show that properly filtered and deduplicated web data alone can lead to powerful models; even significantly outperforming models trained on The Pile. Despite extensive filtering, the high-quality data we extract from the web is still plentiful, and we are able to obtain five trillion tokens from CommonCrawl. We publicly release an extract of 600 billion tokens from our REFINEDWEB dataset, and 1.3/7.5B parameters language models trained on it.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Macrodata Refinement and RefinedWeb
  • 3.1 Document preparation: reading data, filtering URLs, extracting text, and language identification
  • 3.2 Filtering: document-wise and line-wise
  • 3.3 Deduplication: fuzzy, exact, and across dumps
  • 4 Experiments
  • 4.1 Setting
  • 4.2 Can web data alone outperform curated corpora?
  • 4.3 Do other corpora benefit from MDR?
  • 5 Limitations
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Macrodata Refinement and the RefinedWeb Pretraining Dataset

    model/method

    RefinedWeb is a 5-trillion-token English pretraining dataset constructed exclusively from CommonCrawl web scrapes through the MacroData Refinement (MDR) pipeline. The dataset design is governed by three central principles:

    1. Scale First: The pipeline targets multi-trillion token datasets (3–6 trillion English tokens) suitable for training 40B–200B parameter models without reliance on human-curated single-domain corpora.
    2. Strict Deduplication: Rigorous multi-stage deduplication combines document-level fuzzy deduplication and exact sequence deduplication, applying aggressive thresholds that eliminate roughly 50% of candidate data.
    3. Neutral Filtering: Machine-learning-based quality filters are avoided (apart from language identification) to prevent introducing auxiliary cultural and social biases. Filtering relies entirely on deterministic rule-based heuristics and URL blocklists for adult and fraudulent material.
  2. Knowl 2 — Macrodata Refinement (MDR) Data Processing Pipeline

    algorithm

    The Macrodata Refinement (MDR) pipeline transforms raw CommonCrawl web dumps into high-quality pretraining data across three sequential stages:

    Input: Raw CommonCrawl WARC archives
    Output: RefinedWeb pretraining dataset
    // Stage 1: Document Preparation
    for each web page in WARC archives do
        if URL matches 4.6M adult/fraud domain blocklist or exceeds URL severity keyword threshold then
            discard page
        if URL belongs to standard curated sources (e.g., Wikipedia, arXiv) then
            discard page
        text <- extract_content_with_trafilatura(page)
        text <- remove_all_urls(text)
        text <- normalize_newlines(text, max_consecutive=2)
        (lang, score) <- fastText_CCNet_classifier(text)
        if lang != "en" or score < 0.65 then
            discard page
        yield document to RW-Raw
    end for
    // Stage 2: Quality and Heuristic Filtering
    for each document in RW-Raw do
        if has_excessive_line_paragraph_or_ngram_repetitions(document) then
            discard document
        if is_length_or_symbol_to_word_ratio_outlier(document) then
            discard document
        document <- remove_boilerplate_lines(document) // navigation, counters
        if removed_line_content_ratio(document) > 0.05 then
            discard document
        yield document to RW-Filtered
    end for
    // Stage 3: Multi-Level Deduplication
    Partition RW-Filtered into 100 shards
    for each shard do
        // Fuzzy deduplication
        compute 9,000 MinHash signatures (5-grams) split into 20 bands of 450 hashes
        discard documents with matching bucket collisions
        // Exact substring deduplication
        identify exact matching substrings >= 50 tokens using suffix arrays
        remove matching substring spans from documents
        // Cross-dump URL deduplication
        discard documents whose URLs were retained in previously processed shards
        yield document to RefinedWeb
    end for

    Overall, the pipeline removes roughly 90% of raw CommonCrawl documents: ~50% removed at Document Preparation (yielding RW-Raw), ~50% of the remainder removed at Filtering (yielding RW-Filtered), and ~50% of the remainder removed at Deduplication.

  3. Knowl 3 — Zero-Shot Generalization of Web-Trained Models vs. Curated Corpora

    empirical result

    Autoregressive language models trained exclusively on RefinedWeb outperform models trained on public curated corpora (such as The Pile) and match the performance of private models trained on curated mixtures (such as GPT-3), at equivalent compute budgets.

    Key empirical findings across benchmark suites include:

    • When trained to 350 billion tokens (350GT), 1B and 7B parameter RefinedWeb models match the performance of the GPT-3 series evaluated across the 10-task main-agg and 6-task core-agg suites.
    • Models trained on RefinedWeb consistently outperform open-source models trained on The Pile, including GPT-Neo, GPT-NeoX, Cerebras-GPT, OPT, and Pythia.
    • An internal control model trained for 350GT on The Pile using the exact same training codebase achieves lower zero-shot accuracy across all evaluation aggregates than the 1B RefinedWeb model, confirming that dataset refinement—rather than architectural or training setup differences—drives the performance gains.
  4. Knowl 4 — Zero-Shot Performance Comparison Across Pretraining Datasets at Compute-Optimal Scale

    data/table

    Autoregressive transformer models of size 1B and 3B parameters were trained to compute optimality under Chinchilla scaling laws (27 billion tokens for 1B parameters; 60 billion tokens for 3B parameters) using identical architectures, hyperparameters, and codebases across different pretraining datasets. Zero-shot accuracy is averaged across the 6-task small-agg aggregate (HellaSwag, LAMBADA, Winogrande, PIQA, ARC, OpenBookQA).

    Model Scale OSCAR-21.09 OSCAR-22.01 C4 The Pile RW-Raw RW-Filtered RefinedWeb
    1B @ 27GT 55.0% 52.7% 55.7% 53.4% 52.7% 54.3% 56.2%
    3B @ 60GT 59.1% 55.9% 59.6% 57.9% 57.4% 58.2% 59.8%

    RefinedWeb outperforms both massive web datasets and the curated dataset The Pile at both scales. Notably, a 1B model trained on RefinedWeb (56.2%) outperforms a 3B model trained on OSCAR-22.01 (55.9%), indicating that a 4x increase in compute cannot compensate for un-deduplicated, lower-quality data.

  5. Knowl 5 — Transferability and Impact of MDR Filtering and Deduplication to Other Pretraining Corpora

    data/table

    Applying the Macrodata Refinement (MDR) filtering and deduplication modules independently and jointly to other popular pretraining datasets demonstrates the widespread efficacy of strict deduplication.

    Evaluation reports zero-shot accuracy on small-agg (6 tasks) using 1B parameter models trained for 27 billion tokens, along with dataset removal rates relative to base:

    Processing Variant OSCAR-21.09 OSCAR-22.01 C4 The Pile RefinedWeb
    Base 55.0% 52.7% 55.7% 53.4% 52.7%
    Filtered 55.4% (+0.4%) 52.3% (-0.4%) 56.2% (+0.5%) 54.2% (+0.8%) 54.3% (+1.6%)
    Removal rate -25.0% -39.8% -16.4% -27.1% -50.8%
    Deduplicated 55.6% (+0.6%) 55.6% (+2.9%) 55.9% (+0.2%) 54.5% (+1.1%) —
    Removal rate -10.8% -60.8% -7.59% -45.3% —
    Filtered + Dedup. 55.5% (+0.5%) 55.4% (+2.7%) 56.4% (+0.7%) 55.2% (+1.8%) 56.2% (+3.5%)
    Removal rate -28.2% -62.2% -17.9% -66.0% -75.4%

    Strict deduplication yields consistent positive zero-shot accuracy gains across all corpora (+0.2% to +2.9%), whereas filtering heuristics yield corpus-dependent results (-0.4% on OSCAR-22.01 up to +1.6% on RW-Raw). Combining filtering and deduplication produces the best performance across all datasets.

  6. Knowl 6 — Zero-Shot Evaluation Task Aggregates for Pretraining Data Validation

    experimental setup

    Zero-shot performance is evaluated using the EleutherAI LM Evaluation Harness across 18 natural language tasks structured into four hierarchical aggregates:

    • small (6 tasks): HellaSwag, LAMBADA, Winogrande, PIQA, ARC, OpenBookQA. Designed for small-scale ablation studies to obtain consistent non-zero zero-shot signals.
    • core (6 tasks): HellaSwag, Winogrande, PIQA, ARC, OpenBookQA, BoolQ. Matches standard evaluation benchmarks reported for open model families.
    • main (10 tasks): Combines the core suite with LAMBADA, COPA, CB, and RTE. Standard comparison suite for matching GPT-3 and PaLM publications.
    • ext (18 tasks): Expands main to include ReCoRD, ANLI, LogiQA, HeadQA, MathQA, PROST, PubMedQA, and SciQ, mirroring the BigScience benchmark suite.
  7. Knowl 7 — MinHash Fuzzy Deduplication Configuration at Web Scale

    model/method

    Document-level fuzzy deduplication in the Macrodata Refinement pipeline is implemented using MinHash locally-sensitive hashing over 5-grams:

    • Hash Signatures: 9,000 hashes per document are computed across character/word 5-grams.
    • Banding / Bucketing: The 9,000 hashes are grouped into 20 bands (buckets) containing 450 hashes each. Two documents are flagged as duplicates if their MinHash sketches match completely on all 450 hashes within any single band.
    • Hyperparameter Sensitivity: Less aggressive configurations (e.g., using only 10 hashes as implemented in The Pile) lead to significantly lower duplicate detection rates and noticeably worse downstream zero-shot task performance.
  8. Knowl 8 — Exact Substring Deduplication via Suffix Arrays

    model/method

    Exact sequence deduplication in the Macrodata Refinement pipeline operates at the substring level across documents:

    • Method: Suffix arrays are used to locate exact token-by-token verbatim substring matches across the corpus.
    • Threshold: Any exact matching substring of 50 or more consecutive tokens is identified and removed directly from the containing document.
    • Excision Strategy: Excising duplicated spans preserves the remaining valid portions of documents. Ablations confirm that span removal achieves zero-shot generalization performance indistinguishable from either dropping the entire document or masking the loss on duplicated tokens during training.
  9. Knowl 9 — Limitations of Web-Only Pretraining Corpora

    limitation

    The RefinedWeb dataset and web-only pretraining exhibit several known constraints:

    1. Domain Coverage Beyond Natural Language: While general natural language capabilities benefit certain technical domains (e.g., RefinedWeb-trained models outperform Pile-trained models on PubMedQA without explicit medical curation), web scraping alone is insufficient for peak performance on specialized technical tasks such as source code synthesis and advanced mathematics, requiring combination with curated code datasets (e.g., The Stack).
    2. Toxicity and Bias: Evaluated with the Perspective API, RefinedWeb exhibits toxicity levels comparable to The Pile. Neutral rule-based filtering does not explicitly eliminate systemic social biases.
    3. Data Yield and Rejection Rate: The MDR pipeline discards nearly 90% of raw web pages. While CommonCrawl is vast enough to yield 5 trillion English tokens, applying identical pipelines on smaller corpora or heavily pre-filtered sources severely restricts token yields.

Coverage note — Deliberately omitted are specific implementation details of secondary language adaptations mentioned in the appendices and exploratory discussions on epoch-scaling in data-constrained regimes.

References

  1. 1.Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. Compute trends across three eras of machine learning. arXiv preprint arXiv:2202.05924, 2022.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  3. 3.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022.
  4. 4.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  5. 5.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  6. 6.Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 2022.
  7. 7.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  8. 8.Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures. Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019. Cardiff, 22nd July 2019, pages 9 – 16, Mannheim, 2019. Leibniz-Institut für Deutsche Sprache. doi: 10.14618/ids-pub-9021. URL http://nbn-resolving.de/urn:nbn:de:bsz:mh39-90215.
  9. 9.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  10. 10.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  11. 11.Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424, 2022.
  12. 12.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher. 2021. doi: 10.48550/ARXIV.2112.11446. URL https://arxiv.org/abs/2112.11446.
  13. 13.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, 2019.
  15. 15.Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005, 2013.
  16. 16.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27, 2015.
  17. 17.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. 2019.
  18. 18.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019.
  19. 19.Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, 2019.
  20. 20.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977, 2020.
  21. 21.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  22. 22.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  23. 23.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  25. 25.Trieu H Trinh and Quoc V Le. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847, 2018.
  26. 26.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Auguste Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72, 2022.
  27. 27.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4003–4012, 2020.
  28. 28.Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hervé Jégou, and Tomas Mikolov. Fasttext. zip: Compressing text classification models. arXiv preprint arXiv:1612.03651, 2016.
  29. 29.Édouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomáš Mikolov. Learning word vectors for 157 languages. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), 2018.
  30. 30.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, 2022.
  31. 31.Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, 2021.
  32. 32.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  33. 33.Udi Manber and Gene Myers. Suffix arrays: a new method for on-line string searches. Journal on Computing, 22(5):935–948, 1993.
  34. 34.Andrei Z Broder. On the resemblance and containment of documents. In Proceedings. Compression and Complexity of Sequences 1997, pages 21–29. IEEE, 1997.
  35. 35.Moses S Charikar. Similarity estimation techniques from rounding algorithms. In Proceedings of the thiry-fourth annual ACM symposium on Theory of computing, pages 380–388, 2002.
  36. 36.Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al. Pangu-alpha: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369, 2021.
  37. 37.Amro Kamal Mohamed Abbas, Kushal Tirumala, Daniel Simig, Surya Ganguli, and Ari S Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication. In ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2023.
  38. 38.Miltiadis Allamanis. The adverse effects of code duplication in machine learning models of code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, pages 143–153, 2019.
  39. 39.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  40. 40.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650, 2021.
  41. 41.Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022.
  42. 42.Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. arXiv preprint arXiv:2304.01373, 2023.
  43. 43.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  44. 44.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469, 2021.
  45. 45.Konstantin Lopukhin. Evaluating quality of article body extraction for commercial services and open-source libraries. https://github.com/scrapinghub/article-extraction-benchmark, 2019.
  46. 46.Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 122–131. Association for Computational Linguistics, 2021. URL https://aclanthology.org/2021.acl-demo.15.
  47. 47.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2019.
  48. 48.Nolan Dey, Gurpreet Gosal, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, Joel Hestness, et al. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster. arXiv preprint arXiv:2304.03208, 2023.
  49. 49.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  50. 50.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.
  51. 51.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525–1534, 2016.
  52. 52.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  53. 53.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020.
  54. 54.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  55. 55.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, 2018.
  56. 56.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of NAACL-HLT, pages 2924–2936, 2019.
  57. 57.Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele. Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In * SEM 2012: The First Joint Conference on Lexical and Computational Semantics–Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), pages 394–398, 2012.
  58. 58.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124, 2019.
  59. 59.Ido Dagan, Bill Dolan, Bernardo Magnini, and Dan Roth. Recognizing textual entailment: Rational, evaluation and approaches–erratum. Natural Language Engineering, 16(1):105–105, 2010.
  60. 60.Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885, 2018.
  61. 61.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial nli: A new benchmark for natural language understanding. arXiv preprint arXiv:1910.14599, 2019.
  62. 62.Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. Logiqa: a challenge dataset for machine reading comprehension with logical reasoning. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 3622–3628, 2021.
  63. 63.David Vilares and Carlos Gómez-Rodríguez. Head-qa: A healthcare dataset for complex reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 960–966, 2019.
  64. 64.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367, 2019.
  65. 65.Stéphane Aroca-Ouellette, Cory Paik, Alessandro Roncone, and Katharina Kann. Prost: Physical reasoning about objects through space and time. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4597–4608, 2021.
  66. 66.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2567–2577, 2019.
  67. 67.Johannes Welbl, Nelson F Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 94–106, 2017.
  68. 68.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pretraining and finetuning transformers. In International Conference on Learning Representations, 2021.
  69. 69.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. What language model architecture and pretraining objective work best for zero-shot generalization? In International Conference on Machine Learning, 2022.
  70. 70.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. Gpt-neox-20b: An open-source autoregressive language model. Challenges & Perspectives in Creating Large Language Models, page 95, 2022.
  71. 71.Aleph Alpha. Luminous: performance benchmarks. arXiv preprint arXiv:1810.12885, 2023. URL https://www.aleph-alpha.com/pdf/2023_02_AA_Benchmarks_doc.pdf.
  72. 72.Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2021.
  73. 73.Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems, 2022.
  74. 74.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  75. 75.Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus. Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-9) 2021. Limerick, 12 July 2021 (Online-Event), pages 1 – 9, Mannheim, 2021. Leibniz-Institut für Deutsche Sprache. doi: 10.14618/ids-pub-10468. URL https://nbn-resolving.org/urn:nbn:de:bsz:mh39-104688.
  76. 76.Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, et al. Efficient large scale language modeling with mixtures of experts. arXiv preprint arXiv:2112.10684, 2021.
  77. 77.Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata.
  78. 78.Denis Kocetkov, Raymond Li, LI Jia, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, et al. The stack: 3 tb of permissively licensed source code. Transactions on Machine Learning Research, 2022.
  79. 79.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023.
  80. 80.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023.
  81. 81.Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models, 2023.
  82. 82.Ebtesam Almazrouei, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. Falcon-40b: an open large language model with state-of-the-art performance. 2023.
  83. 83.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, 2021.
  84. 84.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45, 2020.
  85. 85.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86–92, 2021.
  86. 86.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, pages 220–229, 2019.
  87. 87.David M. Eberhard, Gary F. Simons, and Charles D. Fennig. Ethnologue: Languages of the World. SIL International, Dallas, TX, USA, twenty-sixth edition, 2023.
  88. 88.Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tuning large neural networks via zero-shot hyperparameter transfer. Advances in Neural Information Processing Systems, 34:17084–17097, 2021.
  89. 89.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, 2021.
  90. 90.Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. Towards a Cleaner Document-Oriented Multilingual Crawled Corpus. arXiv e-prints, art. arXiv:2201.06642, January 2022.
  91. 91.Jan Pomikálek. Justext. 2011.
  92. 92.Dick Sites. Compact language detector 2. Software available at https://github.com/CLD2Owners/cld2 (last updated on August 2015), 2013.
  93. 93.Laura Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020.
  94. 94.P. Jaccard. The distribution of the flora in the alpine zone.1. New Phytologist, 11:37–50, 1912.

Citation

MLA
Penedo, G., et al. “The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only”. arXiv, 2023, http://arxiv.org/abs/2306.01116v1.
APA
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., & Launay, J. (2023). The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. arXiv. http://arxiv.org/abs/2306.01116v1
Chicago
Penedo, G., Q. Malartic, D. Hesslow, et al. 2023. “The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only”. arXiv. http://arxiv.org/abs/2306.01116v1.
Harvard
Penedo, G. et al. (2023) “The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.01116v1.
Vancouver
1. Penedo G, Malartic Q, Hesslow D, Cojocaru R, Cappelli A, Alobeidli H, Pannier B, Almazrouei E, Launay J (2023) The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. arXiv

BibTeX

@article{penedo2023the,
  title = {The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only},
  author = {Penedo, Guilherme and Malartic, Quentin and Hesslow, Daniel and Cojocaru, Ruxandra and Cappelli, Alessandro and Alobeidli, Hamza and Pannier, Baptiste and Almazrouei, Ebtesam and Launay, Julien},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.01116v1},
  eprint = {2306.01116}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors