Deduplicating Training Data Mitigates Privacy Risks in Language Models
Nikhil KandpalEric WallaceColin Raffel
Demonstrates that language model memorization is driven by duplicate training text, proving that simple data deduplication drastically reduces privacy risks from sequence extraction attacks.
Large language models are increasingly deployed in sensitive settings where training data may include proprietary code, private communications, or confidential personal information. However, recent research has raised serious concerns that malicious actors can execute privacy attacks to extract and identify exact training sequences from these models. Addressing this challenge is vital for organizations seeking to safely leverage language models while complying with data protection standards and minimizing exposure to regulatory and reputational risks.
The article evaluates whether the success of these extraction attacks stems from duplicate sequences embedded within standard training corpora, and it demonstrates that removing sequence-level duplicates significantly reduces privacy leakage.
To investigate this, the analysis examined multiple Transformer-based language models spanning 117 million to 1.5 billion parameters trained on large web-scraped datasets, including OpenWebText (39 gigabytes) and C4 (750 gigabytes). The researchers generated extensive text samples across multiple sampling configurations and measured how sequence duplication impacts both the rate of training data regeneration and the accuracy of membership inference techniques used by attackers to detect memorized text.
The article established four critical findings. First, language models regenerate training text at a superlinear rate relative to how often that text appears in the training data; for example, a sequence duplicated 10 times is generated on average approximately 1,000 times more often than a sequence appearing only once. Second, non-duplicated sequences are rarely regenerated, and membership inference methods perform near random chance (showing minimal true positive rates) when trying to detect non-duplicated data. Third, models trained on deduplicated data emit approximately 20 times less training data compared to models trained on standard data. Fourth, while deduplication neutralizes simpler detection metrics, reference-based detection methods maintain moderate predictive accuracy on the rare training sequences that deduplicated models still emit.
These findings indicate that prior assessments likely overstated the practical vulnerability of language models to data recovery attacks, as past vulnerabilities were largely driven by uncleaned, highly duplicated web text. For organizational decision-makers, sequence-level deduplication serves as a highly effective, low-risk privacy safeguard that significantly enhances data protection without degrading underlying model language performance.
Organizations developing or deploying language models on sensitive data should implement sequence-level deduplication pipelines prior to training. Deduplication should also accompany formal privacy frameworks, such as differential privacy, because repeated occurrences of identical text can otherwise bypass privacy protections. Researchers evaluating future privacy attacks must also explicitly account for dataset duplication to avoid distorted security assessments.
The conclusions are robust across varied model sizes, sequence lengths, and sampling techniques, though the study focused primarily on exact string duplicates and English text datasets. Readers should exercise caution regarding approximate or semantic paraphrasing duplicates, which require further investigation across broader data modalities such as source code and image datasets.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). This foundational paper establishes the practical extraction attacks and membership inference techniques on large language models that the source paper directly analyzes and mitigates through deduplication.
- Paper: Deduplicating Training Data Makes Language Models Better, Katherine Lee et al. (2022). Reading this companion work introduces the scalable exact and near-duplicate removal algorithms (suffix arrays and MinHash) that provide the methodology for the source's privacy evaluations.
- Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). This paper establishes the formal exposure metrics and empirical foundations of unintended sequence memorization in generative neural networks.
- Paper: Membership Inference Attacks Against Machine Learning Models, Reza Shokri et al. (2016). This seminal study introduces membership inference attacks against machine learning models, defining the core privacy threat evaluated in the source.
- Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). This study offers foundational insights into the optimization dynamics of memorization versus generalization in deep neural networks.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). Pythia systematically scales and benchmarks language models across training checkpoints on both standard and deduplicated corpora to study verbatim memorization dynamics over time.
- Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). This work extends the source's critique of sentence-level membership inference attacks by demonstrating their near-chance accuracy on matched distributions and proposing dataset-level inference.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). This paper applies large-scale deduplication and refinement pipelines to raw Common Crawl web archives to construct high-performing, curated pretraining datasets.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). This research builds a standardized testbed across trillions of tokens to systematically benchmark the impact of data curation and deduplication on downstream model performance.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). This paper investigates how intentionally repeating data over multiple training epochs impacts language models when unique data is constrained, offering a direct contrast to deduplication regimes.
- Paper: Machine Unlearning of Pre-trained Large Language Models, Jin Yao et al. (2024). This work explores post-training machine unlearning to remove memorized and sensitive training data directly from foundation models without full retraining.
- Paper: Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models, Gowthami Somepalli et al. (2023). This study generalizes the investigation of training data duplication and verbatim replication from language models to generative diffusion models.
