Deduplicating Training Data Makes Language Models Better
Katherine LeeDaphne IppolitoAndrew NystromChiyuan ZhangDouglas EckChris Callison-BurchNicholas Carlini
Demonstrates that removing duplicate strings and documents from language model training sets drastically reduces verbatim memorization and training costs while maintaining or improving perplexity.
Modern artificial intelligence models rely on massive text datasets scraped from the internet, which are often too large for manual review. As a result, these datasets suffer from widespread data duplication and repetitive phrasing, skewing model behavior and compromising real-world utility.
The article aims to evaluate the extent of duplicated text across standard language modeling datasets and demonstrate that systematic deduplication improves model efficiency, accuracy, and output quality while reducing unwanted memorization.
To address this challenge, the authors introduced two scalable, linear-time deduplication techniques: exact substring matching via suffix arrays to detect repeated verbatim sequences (using a 50-token threshold), and approximate full-document matching using MinHash algorithms to identify near-identical documents sharing high n-gram overlap. They tested these methods across four widely used benchmark datasets—Colossal Cleaned Common Crawl (C4), RealNews, LM1B, and Wiki-40B—and trained 110-million-parameter and 1.5-billion-parameter language models to evaluate performance changes.
The study yielded several critical findings. First, duplication is pervasive: web-scraped datasets contain between 3% and 14% near-duplicate documents, and exact substring removal reduced dataset token sizes by up to 19%. Second, standard datasets exhibit substantial data contamination, with 4.6% of C4 validation samples and 14.4% of RealNews validation samples appearing in their respective training sets. Third, deduplication reduced the rate of verbatim memorized output by tenfold (10×); models trained on non-deduplicated data copied training sequences in over 1% of unprompted generations, whereas deduplicated models dropped this rate to roughly 0.1%. Finally, training on deduplicated corpora achieved equal or superior evaluation perplexity while reducing computational runtime and energy costs due to smaller dataset volumes.
These findings indicate that current performance benchmarks for standard language models may be significantly overestimating model capabilities due to train-test data leakage. Furthermore, models that memorize training data pose notable compliance and privacy risks, as they can inadvertently emit sensitive personal information or proprietary text in production. Deduplicating training datasets directly lowers cloud computing and environmental costs without degrading downstream model quality.
Organizations developing or deploying large language models should implement proactive, stringent deduplication pipelines prior to model pre-training and auditing. Practitioners should remove train-test overlaps to ensure valid performance measurements and utilize released open-source deduplication tools to streamline preprocessing workflows.
Confidence in these findings is high across the tested English datasets and architectures. However, decision-makers should note that deduplication alone does not eliminate all privacy risks—such as the inclusion of isolated sensitive records—and domain-specific tasks requiring exact factual recall (like closed-book question answering) may require specialized filtering strategies.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). It demonstrates how large language models emit verbatim memorized text from their training sets, establishing the foundational memorization vulnerability that dataset deduplication directly addresses.
- Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). It introduces quantitative methods and exposure metrics to analyze unintended memorization in generative models, providing key conceptual foundations for measuring training data leakage.
- Paper: Winnowing: local algorithms for document fingerprinting, S. Schleimer et al. (2003). It provides the core local document fingerprinting and substring matching algorithms essential for efficient large-scale exact and near-duplicate text detection.
- Paper: The Pile: An 800GB Dataset of Diverse Text for Language Modeling, Leo Gao et al. (2020). It details the construction and composition of massive pretraining corpora like The Pile that commonly suffer from repeated text and train-test contamination.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). It directly implements pretraining on both standard and deduplicated corpora across parameter scales to analyze the precise training dynamics and memorization impacts established in the source.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). It scales up data curation and aggressive deduplication pipelines to construct massive web datasets that match or outperform curated corpora.
- Paper: Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research, Luca Soldaini et al. (2024). It incorporates extensive deduplication and quality filtering across a multi-trillion-token open pretraining corpus and documents their impact via controlled ablations.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). It establishes a standardized benchmark to systematically isolate and evaluate the downstream performance impacts of data curation strategies, including deduplication.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). It extends the study of data repetition by investigating the performance trade-offs and scaling dynamics of deliberately training models over multiple epochs under constrained data budgets.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). It applies rigorous deduplication and heuristic cleaning recipes specifically to large-scale source code datasets to train high-performing open code models.
