Deduplicating Training Data Mitigates Privacy Risks in Language Models

Nikhil KandpalEric WallaceColin Raffel

article2022ICML386 citations

Demonstrates that language model memorization is driven by duplicate training text, proving that simple data deduplication drastically reduces privacy risks from sequence extraction attacks.

Listen

Large language models are increasingly deployed in sensitive settings where training data may include proprietary code, private communications, or confidential personal information. However, recent research has raised serious concerns that malicious actors can execute privacy attacks to extract and identify exact training sequences from these models. Addressing this challenge is vital for organizations seeking to safely leverage language models while complying with data protection standards and minimizing exposure to regulatory and reputational risks.

The article evaluates whether the success of these extraction attacks stems from duplicate sequences embedded within standard training corpora, and it demonstrates that removing sequence-level duplicates significantly reduces privacy leakage.

To investigate this, the analysis examined multiple Transformer-based language models spanning 117 million to 1.5 billion parameters trained on large web-scraped datasets, including OpenWebText (39 gigabytes) and C4 (750 gigabytes). The researchers generated extensive text samples across multiple sampling configurations and measured how sequence duplication impacts both the rate of training data regeneration and the accuracy of membership inference techniques used by attackers to detect memorized text.

The article established four critical findings. First, language models regenerate training text at a superlinear rate relative to how often that text appears in the training data; for example, a sequence duplicated 10 times is generated on average approximately 1,000 times more often than a sequence appearing only once. Second, non-duplicated sequences are rarely regenerated, and membership inference methods perform near random chance (showing minimal true positive rates) when trying to detect non-duplicated data. Third, models trained on deduplicated data emit approximately 20 times less training data compared to models trained on standard data. Fourth, while deduplication neutralizes simpler detection metrics, reference-based detection methods maintain moderate predictive accuracy on the rare training sequences that deduplicated models still emit.

These findings indicate that prior assessments likely overstated the practical vulnerability of language models to data recovery attacks, as past vulnerabilities were largely driven by uncleaned, highly duplicated web text. For organizational decision-makers, sequence-level deduplication serves as a highly effective, low-risk privacy safeguard that significantly enhances data protection without degrading underlying model language performance.

Organizations developing or deploying language models on sensitive data should implement sequence-level deduplication pipelines prior to training. Deduplication should also accompany formal privacy frameworks, such as differential privacy, because repeated occurrences of identical text can otherwise bypass privacy protections. Researchers evaluating future privacy attacks must also explicitly account for dataset duplication to avoid distorted security assessments.

The conclusions are robust across varied model sizes, sequence lengths, and sampling techniques, though the study focused primarily on exact string duplicates and English text datasets. Readers should exercise caution regarding approximate or semantic paraphrasing duplicates, which require further investigation across broader data modalities such as source code and image datasets.

arXiv: 2202.06539
Cover for Deduplicating Training Data Mitigates Privacy Risks in Language Models

Abstract

Past work has shown that large language models are susceptible to privacy attacks, where adversaries generate sequences from a trained model and detect which sequences are memorized from the training set. In this work, we show that the success of these attacks is largely due to duplication in commonly used web-scraped training sets. We first show that the rate at which language models regenerate training sequences is superlinearly related to a sequence's count in the training set. For instance, a sequence that is present 10 times in the training data is on average generated ~1000 times more often than a sequence that is present only once. We next show that existing methods for detecting memorized sequences have near-chance accuracy on non-duplicated training sequences. Finally, we find that after applying methods to deduplicate training data, language models are considerably more secure against these types of privacy attacks. Taken together, our results motivate an increased focus on deduplication in privacy-sensitive applications and a reevaluation of the practicality of existing privacy attacks.

Table of Contents

  • 1 Introduction
  • 2 Background and Experimental Setup
  • 3 How Duplication Affects The Regeneration of Training Sequences
  • 3.1 Regeneration is Superlinearly Related to Duplicates
  • 3.2 Regeneration Trends Are Robust Across Experimental Setups
  • 4 How Duplication Affects The Detection of Training Sequences
  • 5 Model Inversion with Deduplicated Data
  • 6 Discussion
  • 7 Related Work
  • 8 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Superlinear Scaling of Training Sequence Regeneration with Duplicate Frequency

    empirical result

    Neural language models (Transformer LMs ranging from 117M to 1.5B parameters) regenerate training sequences at a rate that scales superlinearly with the number of duplicate occurrences dd of that sequence in the training data. Specifically, on a log-log plot comparing the expected number of generations of an NN-character sequence against its duplicate count dd in the training corpus (when sampling a total token volume equal to the training set size), the empirical slope is strictly greater than 1. For example, a 100-character sequence appearing 10 times in the training data is generated on average approximately 1000×1000\times more frequently than a sequence that appears only once. Conversely, training sequences appearing only once (d=1d=1) are generated at rates far below a linear "perfect memorization" baseline, meaning unprompted generation-based model inversion attacks rarely emit non-duplicated training data.

  2. Knowl 2 — Empirical Privacy Gains from Sequence-Level Training Data Deduplication

    data/table

    Training a 1.5B parameter language model on a sequence-level deduplicated dataset (C4 deduplicated using suffix-array-based ExactSubstr removing exact duplicate sequences of at least 50 BPE tokens) dramatically reduces the volume of training data leaked via unconditional generation and reduces the efficacy of certain membership inference attacks compared to training on the standard C4 dataset. Evaluating 1,000,000 generated sequences of length 256 tokens against 400-character training sequences yields the following results:

    Metric Normal Model Deduped Model
    Training Data Generated (Count) 1,427,212 68,090
    Training Data Generated (Percent) 0.14% 0.007%
    Membership Inference AUROC (zlib) 0.76 0.67
    Membership Inference AUROC (Reference Model) 0.88 0.87
    Membership Inference AUROC (Lowercase) 0.86 0.68

    Deduplication reduces the number of unique 400-character training sequences generated by approximately 20×20\times (from 1,427,212 to 68,090) and decreases the proportion of all training sequences leaked from 0.14% to 0.007%. Membership inference detection based on zlib compression ratio and lowercase perplexity ratio drop in AUROC significantly (0.76 to 0.67 and 0.86 to 0.68, respectively), whereas Reference Model scoring remains relatively stable (0.88 to 0.87).

  3. Knowl 3 — Dependence of Membership Inference Accuracy on Sequence Duplication Level

    empirical result

    The ability of membership inference metrics to distinguish memorized training sequences from novel LM generations strongly depends on the duplicate count dd of the sequence in the training data. For sequences appearing only once (d=1d=1) in training:

    • The Area Under the Receiver Operating Characteristic curve (AUROC) for Reference Model, zlib compression, and Lowercase perplexity classifiers drops to near-chance accuracy (AUROC ≈0.50\approx 0.50, where 0.50 is the no-skill random guessing baseline). As dd increases toward hundreds of duplicates, AUROC increases, reaching up to ≈0.90\approx 0.90 for Reference Model scoring.

    • At a strict False Positive Rate (FPR) of 0.1% (simulating the realistic setting where training data makes up roughly 1 in 1000 generated 100-character spans), the True Positive Rate (TPR) for the Reference Model score is approximately 0.10 for d=1d=1, but rises to ≈0.60\approx 0.60 for highly duplicated sequences (d≥100d \ge 100).

    Thus, high average membership inference performance in standard evaluations is driven primarily by memorized sequences that occur with high duplication in the training corpus.

  4. Knowl 4 — Difficulty-Calibrated Membership Inference Scoring for Model Inversion

    model/method

    In a two-stage model inversion privacy attack on autoregressive language models, generated candidate sequences are classified as training data versus novel non-training text using easiness-normalized perplexity ratios. Each score evaluates a candidate sequence xx by dividing an "easiness" metric E(x)E(x) (which estimates sequence complexity independently of whether xx was in the training set) by the trained model's perplexity PPLmodel(x)\text{PPL}_{\text{model}}(x):

    Score(x)=E(x)PPLmodel(x)\text{Score}(x) = \frac{E(x)}{\text{PPL}_{\text{model}}(x)}

    Three easiness estimators E(x)E(x) are used:

    1. Reference Model: E(x)=PPLref(x)E(x) = \text{PPL}_{\text{ref}}(x), where PPLref\text{PPL}_{\text{ref}} is the perplexity of an independently trained language model (such as GPT-2 Small).

    2. zlib Compression: E(x)=len(zlib(x))E(x) = \text{len}(\text{zlib}(x)), measuring the byte length of sequence xx after compression using the zlib compression library.

    3. Lowercase Perplexity: E(x)=PPLmodel(lowercase(x))E(x) = \text{PPL}_{\text{model}}(\text{lowercase}(x)), measuring the target model's perplexity when evaluating the lowercased version of xx.

    Candidate generations exhibiting high score ratios are classified as memorized training data.

  5. Knowl 5 — Perfect Memorization Baseline

    definition

    The perfect memorization baseline is a theoretical reference model used to calibrate generative memorization in language models. Given a training corpus DD, a perfect memorization model places non-zero probability only on exact sequences present in DD, such that generating text from the model is equivalent to sampling uniformly at random from DD. Under this baseline, when generating a synthetic text corpus equal in total size to the training set DD, any sequence that appears dd times in DD has an expected generation count of exactly dd, representing a linear relationship between duplication count and generation frequency (slope equal to 1 on a log-log plot).

  6. Knowl 6 — Multiplicative Effect of Training Epochs on Sequence Regeneration Rates

    empirical result

    Increasing the number of training epochs on a fixed dataset increases the expected regeneration frequency of training sequences by a multiplicative factor that is approximately uniform across all duplication levels dd. For a 117M parameter Transformer LM, doubling the number of training epochs (e.g., from 6 to 12 epochs, or from 12 to 24 epochs) causes the expected number of generated training sequences to increase by approximately 3×3\times across all duplication levels d∈[1,102]d \in [1, 10^2]. Early stopping does not flatten the superlinear duplication-versus-generation curve; models trained for fewer epochs still generate disproportionately more highly duplicated sequences than singletons.

  7. Knowl 7 — Influence of Sampling Strategy on Training Sequence Regeneration

    empirical result

    The decoding strategy used during text generation affects the absolute rate at which an LM regenerates verbatim training sequences. Sampling methods that concentrate probability on higher-likelihood tokens—such as top-kk sampling with smaller values of kk (e.g., k=20k=20 vs. k=500k=500) or temperature sampling with smaller temperatures TT (e.g., T=0.2T=0.2 vs. T=0.5T=0.5)—generate verbatim training samples more frequently than higher-entropy or pure random sampling. However, across all tested sampling strategies (random sampling, top-kk, and temperature scaling), the superlinear relationship between training duplicate count dd and generation frequency remains intact, and training sequences with low duplicate counts are consistently regenerated at extremely low frequencies.

  8. Knowl 8 — Robustness of Superlinear Memorization Trends Across Model Scale and Sequence Length

    empirical result

    The superlinear relationship between sequence duplicate count in training data and model regeneration frequency is robust to changes in sequence length and model parameter size:

    • Duplicate Sequence Length: Varying the evaluated sequence length NN across N∈{100,200,300,400,500,600,700}N \in \{100, 200, 300, 400, 500, 600, 700\} characters shifts the overall expected generation probability downward for longer sequences, but preserves the characteristic superlinear slope (>1> 1 on a log-log scale) across all tested lengths.

    • Model Parameter Scale: Comparing Transformer models scaled from 117M to 345M and 1.5B parameters shows that larger models regenerate more training data across all duplication levels dd, attributable to lower overall training loss (higher assigned likelihoods to training sequences), while retaining the same superlinear dependence on duplication count.

  9. Knowl 9 — Inadequacy of Differential Privacy Alone Against Sequence Duplication

    limitation

    While training with Differential Privacy (DP) guarantees that the bounded effect of any single training example on the trained model parameters is small, standard DP bounds apply on a per-sample basis. When a sequence or near-duplicate sequence appears dd times across a dataset, the cumulative impact of those repeated instances on model updates scales with dd, undermining empirical privacy protections. Consequently, applying dataset deduplication before training remains a necessary safeguard even when using differentially private training algorithms.

Coverage note — None was omitted; all key empirical findings, analytical metrics, baseline definitions, deduplication comparisons, and privacy implications were included.

References

  1. 1.Brown, G., Bun, M., Feldman, V., Smith, A., and Talwar, K. When is memorization of irrelevant training data necessary for high-accuracy learning? In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing. ACM, 2021.
  2. 2.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In NeurIPS, 2020.
  3. 3.Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D. The secret sharer: Evaluating and testing unintended memorization in neural networks. In USENIX Security Symposium, 2019.
  4. 4.Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F. Membership inference attacks from first principles. arXiv preprint arXiv:2112.03570, 2021a.
  5. 5.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., and Raffel, C. Extracting training data from large language models. In USENIX Security Symposium, 2021b.
  6. 6.Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  7. 7.Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. SemEval, 2017.
  8. 8.Coavoux, M., Narayan, S., and Cohen, S. B. Privacy-preserving neural representations of text. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018.
  9. 9.Dwork, C., McSherry, F., Nissim, K., and Smith, A. Calibrating noise to sensitivity in private data analysis. In TCC, 2006.
  10. 10.Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. In ACL, 2018.
  11. 11.Feldman, V. and Zhang, C. What neural networks memorize and why: Discovering the long tail via influence estimation. In NeurIPS, 2020.
  12. 12.Fredrikson, M., Jha, S., and Ristenpart, T. Model inversion attacks that exploit confidence information and basic countermeasures. In ACM CCS, 2015.
  13. 13.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  14. 14.Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S. OpenWebText corpus, 2019.
  15. 15.Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., Johnston, S., Mann, B., Olah, C., Olsson, C., Amodei, D., Joseph, N., Kaplan, J., and McCandlish, S. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022.
  16. 16.Hidano, S., Murakami, T., Katsumata, S., Kiyomoto, S., and Hanaoka, G. Model inversion attacks for prediction systems: Without knowledge of non-sensitive attributes. In 2017 15th Annual Conference on Privacy, Security and Trust (PST), 2017.
  17. 17.Inan, H. A., Ramadan, O., Wutschitz, L., Jones, D., Rühle, V., Withers, J., and Sim, R. Training data leakage analysis in language models. In Privacy Preserving Machine Learning Workshop, 2021.
  18. 18.Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021.
  19. 19.Lehman, E., Jain, S., Pichotta, K., Goldberg, Y., and Wallace, B. Does BERT pretrained on clinical notes reveal sensitive data? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021.
  20. 20.Li, X., Tramer, F., Liang, P., and Hashimoto, T. Large language models can be strong differentially private learners. In ICLR, 2022.
  21. 21.Li, Y., Baldwin, T., and Cohn, T. Towards robust and privacy-preserving text representations. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2018.
  22. 22.Long, Y., Bindschaedler, V., Wang, L., Bu, D., Wang, X., Tang, H., Gunter, C. A., and Chen, K. Understanding membership inferences on well-generalized learning models. arXiv preprint arXiv:1802.04889, 2018.
  23. 23.McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., and Celikyilmaz, A. How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN. arXiv preprint arXiv:2111.09509, 2021.
  24. 24.Mireshghallah, F., Inan, H., Hasegawa, M., Rühle, V., Berg-Kirkpatrick, T., and Sim, R. Privacy regularization: Joint privacy-utility optimization in LanguageModels. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021.
  25. 25.Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019.
  26. 26.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  27. 27.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. In JMLR, 2020.
  28. 28.Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do CIFAR-10 classifiers generalize to CIFAR-10? arXiv preprint arXiv:1806.00451, 2018.
  29. 29.Roberts, A., Raffel, C., and Shazeer, N. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020.
  30. 30.Shokri, R., Stronati, M., Song, C., and Shmatikov, V. Membership inference attacks against machine learning models. In IEEE S&P, 2017.
  31. 31.Song, C. and Raghunathan, A. Information Leakage in Embedding Models, pp. 377–390. Association for Computing Machinery, 2020.
  32. 32.Song, C. and Shmatikov, V. Auditing data provenance in text-generation models. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Association for Computing Machinery, 2019.
  33. 33.Van den Burg, G. and Williams, C. On memorization in probabilistic deep generative models. In NeurIPS, 2021.
  34. 34.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In NIPS, 2017.
  35. 35.Watson, L., Guo, C., Cormode, G., and Sablayrolles, A. On the importance of difficulty calibration in membership inference attacks. arXiv preprint arXiv:2111.08440, 2021.
  36. 36.West, P., Lu, X., Holtzman, A., Bhagavatula, C., Hwang, J., and Choi, Y. Reflective decoding: Beyond unidirectional generation with off-the-shelf language models. In ACL, 2021.
  37. 37.Yang, Z., Zhang, J., Chang, E.-C., and Liang, Z. Neural network inversion in adversarial setting via background knowledge alignment. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, 2019.
  38. 38.Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S. Privacy risk in machine learning: Analyzing the connection to overfitting. In IEEE CSF, 2018.
  39. 39.Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H. A., Kamath, G., Kulkarni, J., Lee, Y. T., Manoel, A., Wutschitz, L., et al. Differentially private fine-tuning of language models. arXiv preprint arXiv:2110.06500, 2021.
  40. 40.Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., and Carlini, N. Counterfactual memorization in neural language models. arXiv preprint arXiv:2112.12938, 2021.
  41. 41.Zhao, X., Li, L., and Wang, Y.-X. Provably confidential language modelling. arXiv preprint arXiv:2205.01863, 2022.
  42. 42.Ziegler, A. A first look at rote learning in GitHub Copilot suggestions, June 2021.

Citation

MLA
Kandpal, N., et al. “Deduplicating Training Data Mitigates Privacy Risks in Language Models”. arXiv, 2022, https://doi.org/10.48550/arxiv.2202.06539.
APA
Kandpal, N., Wallace, E., & Raffel, C. (2022). Deduplicating Training Data Mitigates Privacy Risks in Language Models. arXiv. https://doi.org/10.48550/arxiv.2202.06539
Chicago
Kandpal, N., E. Wallace, and C. Raffel. 2022. “Deduplicating Training Data Mitigates Privacy Risks in Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2202.06539.
Harvard
Kandpal, N., Wallace, E. and Raffel, C. (2022) “Deduplicating Training Data Mitigates Privacy Risks in Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2202.06539.
Vancouver
1. Kandpal N, Wallace E, Raffel C (2022) Deduplicating Training Data Mitigates Privacy Risks in Language Models. https://doi.org/10.48550/arxiv.2202.06539

BibTeX

@misc{https://doi.org/10.48550/arxiv.2202.06539,
  doi = {10.48550/ARXIV.2202.06539},
  url = {https://arxiv.org/abs/2202.06539},
  author = {Kandpal, Nikhil and Wallace, Eric and Raffel, Colin},
  keywords = {Cryptography and Security (cs.CR), Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Deduplicating Training Data Mitigates Privacy Risks in Language Models},
  publisher = {arXiv},
  year = {2022},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/