Deduplicating Training Data Makes Language Models Better

Katherine LeeDaphne IppolitoAndrew NystromChiyuan ZhangDouglas EckChris Callison-BurchNicholas Carlini

article2022ACL998 citations

Demonstrates that removing duplicate strings and documents from language model training sets drastically reduces verbatim memorization and training costs while maintaining or improving perplexity.

Listen

Modern artificial intelligence models rely on massive text datasets scraped from the internet, which are often too large for manual review. As a result, these datasets suffer from widespread data duplication and repetitive phrasing, skewing model behavior and compromising real-world utility.

The article aims to evaluate the extent of duplicated text across standard language modeling datasets and demonstrate that systematic deduplication improves model efficiency, accuracy, and output quality while reducing unwanted memorization.

To address this challenge, the authors introduced two scalable, linear-time deduplication techniques: exact substring matching via suffix arrays to detect repeated verbatim sequences (using a 50-token threshold), and approximate full-document matching using MinHash algorithms to identify near-identical documents sharing high n-gram overlap. They tested these methods across four widely used benchmark datasets—Colossal Cleaned Common Crawl (C4), RealNews, LM1B, and Wiki-40B—and trained 110-million-parameter and 1.5-billion-parameter language models to evaluate performance changes.

The study yielded several critical findings. First, duplication is pervasive: web-scraped datasets contain between 3% and 14% near-duplicate documents, and exact substring removal reduced dataset token sizes by up to 19%. Second, standard datasets exhibit substantial data contamination, with 4.6% of C4 validation samples and 14.4% of RealNews validation samples appearing in their respective training sets. Third, deduplication reduced the rate of verbatim memorized output by tenfold (10×); models trained on non-deduplicated data copied training sequences in over 1% of unprompted generations, whereas deduplicated models dropped this rate to roughly 0.1%. Finally, training on deduplicated corpora achieved equal or superior evaluation perplexity while reducing computational runtime and energy costs due to smaller dataset volumes.

These findings indicate that current performance benchmarks for standard language models may be significantly overestimating model capabilities due to train-test data leakage. Furthermore, models that memorize training data pose notable compliance and privacy risks, as they can inadvertently emit sensitive personal information or proprietary text in production. Deduplicating training datasets directly lowers cloud computing and environmental costs without degrading downstream model quality.

Organizations developing or deploying large language models should implement proactive, stringent deduplication pipelines prior to model pre-training and auditing. Practitioners should remove train-test overlaps to ensure valid performance measurements and utilize released open-source deduplication tools to streamline preprocessing workflows.

Confidence in these findings is high across the tested English datasets and architectures. However, decision-makers should note that deduplication alone does not eliminate all privacy risks—such as the inclusion of isolated sensitive records—and domain-specific tasks requiring exact factual recall (like closed-book question answering) may require specialized filtering strategies.

arXiv: 2107.06499
Cover for Deduplicating Training Data Makes Language Models Better

Abstract

We find that existing language modeling datasets contain many near-duplicate examples and long repetitive substrings. As a result, over 1% of the unprompted output of language models trained on these datasets is copied verbatim from the training data. We develop two tools that allow us to deduplicate training datasets—for example removing from C4 a single 61 word English sentence that is repeated over 60,000 times. Deduplication allows us to train models that emit memorized text ten times less frequently and require fewer training steps to achieve the same or better accuracy. We can also reduce train-test overlap, which affects over 4% of the validation set of standard datasets, thus allowing for more accurate evaluation. Code for deduplication is released at https://github.com/google-research/deduplicate-text-datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Language Modeling Datasets
  • 4 Methods for Identifying Duplicates
  • 4.1 Exact Substring Duplication
  • 4.1.1 Suffix Arrays
  • 4.1.2 Substring matching
  • 4.2 Approximate Matching with MinHash
  • 5 Deduplication Results
  • 5.1 Amount of Text Removed
  • 5.2 Properties of Duplicated Text
  • 5.3 Train / Test Set Leakage
  • 6 Impact on Trained Models
  • 6.1 Model Perplexity
  • 6.2 Generated Text
  • 6.3 Impact on Existing Models
  • 7 Discussion
  • 8 Conclusion
  • Ethics
  • Acknowledgements
  • Contributions
  • References
  • A Further Details on NEARDUP
  • B Further Details on EXACTSUBSTR
  • C Further Details on Model Training
  • D Energy Consumption
  • E More Results

Knowls

  1. Knowl 1 — Two complementary scalable deduplication procedures

    model/method

    The paper introduces two corpus-level procedures for removing repeated training text. EXACTSUBSTR detects verbatim substrings shared by different examples and removes the repeated substring from one occurrence, while NEARDUP detects whole documents with high but non-exact n-gram overlap and removes redundant documents. The procedures target different duplication patterns: EXACTSUBSTR captures partial copies embedded in otherwise different documents, whereas NEARDUP captures templated documents that differ in fields such as names, locations, dates, or products. When duplicate content spans training and validation or test splits, the implementation keeps the copy in the evaluation split and removes the training copy.

  2. Knowl 2 — EXACTSUBSTR using suffix-array matching

    algorithm

    EXACTSUBSTR operates on the byte-level representation of a corpus after BPE tokenization. Let the dataset be D={xi}i=1ND=\{x_i\}_{i=1}^{N}, where each example xix_i is a sequence of BPE tokens, and let SS be the concatenation of all examples. The algorithm constructs a suffix array containing every suffix of SS in lexicographic order, then scans adjacent suffixes. Whenever adjacent suffixes have a common prefix of at least 50 BPE tokens and the occurrences belong to different examples, the shared substring is recorded and removed from one occurrence. The 50-token threshold was selected conservatively: matches shorter than about 10 tokens were common, manual inspection found no false positives at length 25, and the final threshold doubled that length.

    A suffix array can be constructed in linear time in ∣S∣|S| and requires approximately 8 bytes per input token in this implementation, substantially less memory than a suffix tree. The implementation parallelizes suffix-array construction by partitioning the corpus, constructing partial arrays independently, and merging them with a heap-based multiway merge; the suffix array itself can be streamed to disk, although the indexed corpus must fit in memory. The overall exact-substring procedure is linear-time in the corpus size under the paper's implementation assumptions and is parallelizable across suffix-array scans.

  3. Knowl 3 — NEARDUP using MinHash and edit similarity

    algorithm

    NEARDUP identifies redundant whole documents through locality-sensitive hashing. For each document xix_i, let did_i be the set of its consecutive 5-grams. The similarity of documents ii and jj is first characterized by the Jaccard index

    Jaccard⁡(di,dj)=∣di∩dj∣∣di∪dj∣.\operatorname{Jaccard}(d_i,d_j)=\frac{|d_i\cap d_j|}{|d_i\cup d_j|}.

    Each document retains a MinHash signature of 9,000 minimum hash values. The values are divided into r=450r=450 bands with b=20b=20 hash values per band. A pair becomes a candidate if the documents agree in at least one band; if their Jaccard similarity is sijs_{ij}, the candidate probability is

    Pr⁡(candidate∣sij)=1−(1−sijb)r.\Pr(\text{candidate}\mid s_{ij})=1-(1-s_{ij}^{b})^{r}.

    For each candidate pair, NEARDUP computes the actual Jaccard similarity and then the token-level edit similarity. For token sequences xix_i and xjx_j, with lengths ∣xi∣|x_i| and ∣xj∣|x_j|, the latter is

    EditSim⁡(xi,xj)=1−EditDistance⁡(xi,xj)max⁡(∣xi∣,∣xj∣).\operatorname{EditSim}(x_i,x_j)=1-\frac{\operatorname{EditDistance}(x_i,x_j)}{\max(|x_i|,|x_j|)}.

    The paper labels a pair a near-duplicate when it passes the MinHash filter, has Jaccard similarity above 0.8, and has edit similarity above 0.8. Duplicate pairs form a graph whose connected components are used to build duplicate clusters, after which redundant training documents are removed. The candidate-grouping stage is linear in the number of documents for fixed hashing parameters; edit-distance verification has worst-case quadratic cost in the maximum document length, but is applied only to candidate pairs.

  4. Knowl 4 — Duplicate content is widespread and causes train–validation leakage

    data/table

    The authors applied both procedures to four English datasets: Wiki-40B, LM1B, C4, and RealNews. NEARDUP found the following fractions of examples:

    • C4: 3.04% of training examples had a duplicate within training, 1.59% of validation examples had a duplicate within validation, and 4.60% of validation examples had a duplicate in training.
    • RealNews: 13.63%, 1.25%, and 14.35%, respectively.
    • LM1B: 4.86%, 0.07%, and 4.92%, respectively.
    • Wiki-40B: 0.39%, 0.26%, and 0.72%, respectively.

    EXACTSUBSTR identified the following fractions of tokens as belonging to a repeated 50-token substring, with the three quantities in the same order—training tokens duplicated within training, validation tokens duplicated within validation, and validation tokens duplicated in training:

    • C4: 7.18%, 0.75%, and 1.38%.
    • RealNews: 19.4%, 2.61%, and 3.37%.
    • LM1B: 0.76%, 0.016%, and 0.019%.
    • Wiki-40B: 2.76%, 0.52%, and 0.67%.

    The corresponding deduplicated training-set sizes, measured with the paper's 50,000-token BPE vocabulary, were C4: 177.3B original tokens, 173.7B after NEARDUP, and 165.4B after EXACTSUBSTR; RealNews: 24.7B, 22.4B, and 20.1B; LM1B: 1.0B, 0.94B, and 0.90B; and Wiki-40B: 2.25B, 2.24B, and 2.19B. Thus, leakage is particularly substantial for RealNews and C4: for example, 14.35% of RealNews validation examples and 4.60% of C4 validation examples have approximate duplicates in training. A notable C4 case is a 61-word sequence repeated 61,036 times in training and 61 times in validation.

  5. Knowl 5 — The duplicate material is largely templated or syndicated text

    empirical result

    Qualitative inspection showed that many near-duplicates are automatically generated web pages whose wording is almost identical except for entity fields such as businesses, products, destinations, dates, and locations. Other near-duplicates arise when the same news article is republished by multiple sites with small formatting changes, such as different prefixes or punctuation. These patterns explain why exact document matching is insufficient and why NEARDUP detects pairs that EXACTSUBSTR may not identify as whole-document duplicates. C4 contained especially large duplicate clusters, including 280 clusters with more than 5,000 documents and one cluster containing 250,933 documents; most C4 clusters nevertheless consisted of only a single matching pair.

  6. Knowl 6 — Deduplication reduces training cost without changing the core model setup

    experimental setup

    The model experiments used decoder-only Transformer language models trained on three C4 variants: the original corpus, a NEARDUP corpus, and an EXACTSUBSTR corpus. The main models were 1.5-billion-parameter XL models trained for approximately two epochs; three random seeds of a 110-million-parameter base model were also trained for each corpus. All models used a 50,000-token BPE vocabulary trained on C4-NEARDUP, a maximum sequence length of 512 tokens, the Adafactor optimizer, and constant learning rates of 0.001 for XL models and 0.01 for base models. XL models had 24 layers, 32 attention heads, 2,048-dimensional embeddings, and 5,120-dimensional feed-forward layers; base models had 12 layers, 12 attention heads, 768-dimensional embeddings, and 2,048-dimensional feed-forward layers.

    C4-ORIGINAL and C4-EXACTSUBSTR required about 152,000 optimization steps and 131 hours for XL training, whereas C4-NEARDUP required about 146,000 steps and 126 hours. The NEARDUP C4 corpus was approximately 3.9% smaller and saved about five hours over two XL-model epochs. Estimated training energy was 5.86 MWh for the original and EXACTSUBSTR models versus 5.63 MWh for the NEARDUP model. Deduplication itself was also scalable: the suffix-array implementation processed the 4GB Wiki-40B training set in 2.3 minutes and built an index for the 350GB C4 corpus in under 12 hours, after which C4 substring removal took under one hour; the distributed NEARDUP run on C4 consumed approximately 41.5 kWh.

  7. Knowl 7 — Deduplication preserves or improves perplexity overall while exposing leakage effects

    empirical result

    Perplexity was evaluated on the original C4 validation set, a C4 subset with no NEARDUP match in either split, a C4 subset with a NEARDUP match in training, and the validation sets of LM1B and Wiki-40B. On the original C4 validation set and on the unique C4 subset, models trained on the original and deduplicated corpora had similar perplexity. On C4 validation examples with training duplicates, both deduplicated models had substantially higher perplexity than the original-data model, showing that the original model benefited from memorizing leaked examples; EXACTSUBSTR produced higher perplexity than NEARDUP on this subset.

    On the independent LM1B and Wiki-40B evaluations, NEARDUP-trained C4 models consistently achieved the lowest perplexity among the compared training corpora. EXACTSUBSTR reduced XL-model perplexity on Wiki-40B by almost 3 perplexity points. The XL models trained on all three C4 variants achieved perplexity below 35 on LM1B, compared with 42.16 reported for a similarly sized GPT-2 model with a vocabulary of the same size. The experiments therefore found no overall perplexity penalty from deduplication and observed improvements on some out-of-corpus evaluations, despite training on fewer tokens.

  8. Knowl 8 — Deduplication reduces unprompted verbatim memorization by about an order of magnitude

    empirical result

    The authors generated 100,000 unprompted sequences of up to 512 tokens from each XL model using top-kk sampling with k=50k=50. A generated token was counted as memorized when it belonged to an exactly matching 50-token substring in the model's training corpus. The percentage of memorized tokens was:

    • After one epoch: 1.926% for XL-ORIGINAL, 0.189% for XL-NEARDUP, and 0.138% for XL-EXACTSUBSTR.
    • After two epochs: 1.571% for XL-ORIGINAL, 0.264% for XL-NEARDUP, and 0.168% for XL-EXACTSUBSTR.

    Thus, more than 1% of tokens generated from the original C4 model belonged to copied training substrings, whereas both deduplicated models reduced this rate to roughly 0.1%, or about ten times less. The copied material frequently consisted of repetitive advertisements and other templated web text.

  9. Knowl 9 — Prompt-conditioned memorization and leakage in released models

    empirical result

    For prompted generation, the authors used the first 32 tokens of examples from four categories: training examples with a training duplicate, unique training examples, validation examples with a training duplicate, and unique validation examples. The percentage of generated tokens belonging to an exact 50-token training substring after one and two epochs was, respectively:

    • Duplicate training prompts: original 35.88% and 34.34%; NEARDUP 3.34% and 3.15%; EXACTSUBSTR 5.71% and 4.67%.
    • Unique training prompts: original 0.42% and 0.41%; NEARDUP 0.42% and 0.41%; EXACTSUBSTR 0.22% and 0.23%.
    • Duplicate validation prompts: original 16.27% and 15.32%; NEARDUP 1.61% and 1.52%; EXACTSUBSTR 0.34% and 0.25%.
    • Unique validation prompts: original 0.25% and 0.22%; NEARDUP 0.21% and 0.23%; EXACTSUBSTR 0.03% and 0.08%.

    The original XL model reproduced the ground-truth continuation for duplicate prompts more than 40% of the time under the paper's edit-similarity criterion, while both deduplicated models still copied duplicate-prompt continuations more often than unique-prompt continuations. Train–validation leakage also affected existing released models: Transformer-XL trained on LM1B had perplexities 21.77 on its full validation set, 10.11 on validation examples with training near-duplicates, and 23.58 on unique examples. GROVER-Base trained on RealNews had corresponding perplexities 15.44, 13.77, and 15.73; GROVER-XL had 9.15, 7.68, and 9.45. In addition, 1.38% of tokens in 25,000 official GROVER-Mega outputs matched at least 50 consecutive RealNews tokens, while more than 5% of tokens in approximately 200,000 GPT-Neo 1.3B outputs matched 50-token substrings of the Pile.

  10. Knowl 10 — Scope and limitations of deduplication and memorization measurements

    limitation

    The experiments treat every generated sequence that closely matches training data as memorization, without distinguishing harmful disclosures such as personal identifiers from innocuous common phrases or desirable memorization such as quotations. The paper therefore does not establish which memorized content is privacy-sensitive or harmful. Deduplication is not a substitute for excluding inherently sensitive records such as passwords or medical data, and it may be undesirable for tasks that require memorization, including retrieval or closed-book question answering. Removing repeated substrings can also remove attribution text while leaving the attributed content, potentially changing what a model learns about provenance. Finally, the deduplication thresholds were selected empirically; changing the thresholds can alter which data are removed and the observed balance between memorization, perplexity, and leakage.

Coverage note — Detailed URL-frequency lists, qualitative duplicate-pair examples, and the complete cluster-size histograms were omitted because they provide secondary diagnostics rather than additional load-bearing methods or conclusions.

References

  1. 1.Miltiadis Allamanis. 2019. The adverse effects of code duplication in machine learning models of code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, pages 143–153.
  2. 2.Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. 2017. A closer look at memorization in deep networks. In International Conference on Machine Learning, pages 233–242. PMLR.
  3. 3.Jack Bandy and Nicholas Vincent. 2021. Addressing "documentation debt" in machine learning research: A retrospective datasheet for bookcorpus.
  4. 4.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  5. 5.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? . In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  6. 6.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow.
  7. 7.Burton H Bloom. 1970. Space/time trade-offs in hash coding with allowable errors. Communications of the ACM, 13(7):422–426.
  8. 8.Andrei Z Broder. 1997. On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pages 21–29. IEEE.
  9. 9.Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022. What does it mean for a language model to preserve privacy? arXiv preprint.
  10. 10.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33.
  11. 11.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2020. Extracting training data from large language models.
  12. 12.Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005.
  13. 13.Hung Chim and Xiaotie Deng. 2007. A new suffix tree similarity measure for document clustering. In Proceedings of the 16th International Conference on World Wide Web, WWW ’07, page 121–130, New York, NY, USA. Association for Computing Machinery.
  14. 14.Edith Cohen. 2016. Min-hash sketches: A brief survey.
  15. 15.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860.
  16. 16.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021a. Documenting the english colossal clean crawled corpus.
  17. 17.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021b. Documenting the english colossal clean crawled corpus. arXiv preprint arXiv:2104.08758.
  18. 18.Vitaly Feldman and Chiyuan Zhang. 2020. What neural networks memorize and why: Discovering the long tail via influence estimation. In Advances in Neural Information Processing Systems.
  19. 19.Rodney A. Gabriel, Tsung-Ting Kuo, Julian McAuley, and Chun-Nan Hsu. 2018. Identifying and characterizing highly similar notes in big clinical note datasets. Journal of Biomedical Informatics, 82:63–69.
  20. 20.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  21. 21.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Dahlme III au2, and Kate Crawford. 2020. Datasheets for datasets.
  22. 22.David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003. English gigaword. Linguistic Data Consortium, Philadelphia, 4(1):34.
  23. 23.Mandy Guo, Zihang Dai, Denny Vrandecic, and Rami Al-Rfou. 2020. Wiki-40b: Multilingual language model dataset. In LREC 2020.
  24. 24.Bikash Gyawali, Lucas Anastasiou, and Petr Knoth. 2020. Deduplication of scholarly documents using locality sensitive hashing and word embeddings. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 901–910.
  25. 25.Paul Jaccard. 1912. The distribution of the flora in the alpine zone. New phytologist, 11(2):37–50.
  26. 26.Juha Kärkkäinen and Peter Sanders. 2003. Simple linear work suffix array construction. In International colloquium on automata, languages, and programming, pages 943–955. Springer.
  27. 27.Pang Ko and Srinivas Aluru. 2003. Space efficient linear time construction of suffix arrays. In Annual Symposium on Combinatorial Pattern Matching, pages 200–210. Springer.
  28. 28.Udi Manber and Gene Myers. 1993. Suffix arrays: a new method for on-line string searches. siam Journal on Computing, 22(5):935–948.
  29. 29.Ge Nong, Sen Zhang, and Wai Hong Chan. 2009. Linear suffix array construction by almost pure induced-sorting. In 2009 data compression conference, pages 193–202. IEEE.
  30. 30.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training.
  31. 31.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  33. 33.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR.
  34. 34.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2020. Towards controllable biases in language generation. arXiv preprint arXiv:2005.00268.
  35. 35.Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pages 3–18. IEEE.
  36. 36.Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui, Hanlin Tang, and SueYeon Chung. 2021. On the geometry of generalization and memorization in deep neural networks. In International Conference on Learning Representations.
  37. 37.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in nlp.
  38. 38.Piotr Teterwak, Chiyuan Zhang, Dilip Krishnan, and Michael C Mozer. 2021. Understanding invariance via feedforward inversion of discriminatively trained classifiers. In International Conference on Machine Learning, pages 10225–10235. PMLR.
  39. 39.Trieu H Trinh and Quoc V Le. 2018. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847.
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762.
  41. 41.Yannick Versley and Yana Panchenko. 2012. Not just bigger: Towards better-quality web corpora. In Proceedings of the seventh Web as Corpus Workshop (WAC7), pages 44–52.
  42. 42.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing nlp. arXiv preprint arXiv:1908.07125.
  43. 43.Ryan Webster, Julien Rabin, Loïc Simon, and Frédéric Jurie. 2019. Detecting overfitting of deep generative networks via latent recovery. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11265–11274.
  44. 44.Peter Weiner. 1973. Linear pattern matching algorithms. In 14th Annual Symposium on Switching and Automata Theory (swat 1973), pages 1–11. IEEE.
  45. 45.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934.
  46. 46.Mikio Yamamoto and Kenneth W Church. 2001. Using suffix arrays to compute term frequency and document frequency for all substrings in a corpus. Computational Linguistics, 27(1):1–30.
  47. 47.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. arXiv preprint arXiv:1905.12616.
  48. 48.Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, Chen Li, Ziyan Gong, Yifan Yao, Xinjing Huang, Jun Wang, Jianfeng Yu, Qi Guo, Yue Yu, Yan Zhang, Jin Wang, Hengtao Tao, Dasen Yan, Zexuan Yi, Fang Peng, Fangqing Jiang, Han Zhang, Lingfeng Deng, Yehong Zhang, Zhe Lin, Chao Zhang, Shaojie Zhang, Mingyue Guo, Shanzhi Gu, Gaojun Fan, Yaowei Wang, Xuefeng Jin, Qun Liu, and Yonghong Tian. 2021. Pangu-α: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369.
  49. 49.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27.
  50. 50.Jakub Ł ˛acki, Vahab Mirrokni, and Michał Włodarczyk. 2018. Connected components at scale via local contractions.

Citation

MLA
Lee, K., et al. “Deduplicating Training Data Makes Language Models Better”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 8424–45, https://doi.org/10.18653/v1/2022.acl-long.577.
APA
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., & Carlini, N. (2022). Deduplicating Training Data Makes Language Models Better. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8424–8445. https://doi.org/10.18653/v1/2022.acl-long.577
Chicago
Lee, K., D. Ippolito, A. Nystrom, et al. 2022. “Deduplicating Training Data Makes Language Models Better”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8424–45. https://doi.org/10.18653/v1/2022.acl-long.577.
Harvard
Lee, K. et al. (2022) “Deduplicating Training Data Makes Language Models Better”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8424–8445. Available at: https://doi.org/10.18653/v1/2022.acl-long.577.
Vancouver
1. Lee K, Ippolito D, Nystrom A, Zhang C, Eck D, Callison-Burch C, Carlini N (2022) Deduplicating Training Data Makes Language Models Better. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8424–8445

BibTeX

@inproceedings{lee-etal-2022-deduplicating,
    title = "Deduplicating Training Data Makes Language Models Better",
    author = "Lee, Katherine  and
      Ippolito, Daphne  and
      Nystrom, Andrew  and
      Zhang, Chiyuan  and
      Eck, Douglas  and
      Callison-Burch, Chris  and
      Carlini, Nicholas",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.577/",
    doi = "10.18653/v1/2022.acl-long.577",
    pages = "8424--8445"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/