Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Martin RiddellAnsong NiArman Cohan

article2024ACL57 citations

Quantifies benchmark leakage in popular code evaluation sets like HumanEval and MBPP across major pretraining corpora using both surface and semantic matching, revealing how test-data memorization artificially inflates code generation performance.

Listen

Evaluating the performance of large language models on programming tasks relies heavily on standard benchmarks. However, as training datasets expand to include massive repositories of open-source source code, there is a substantial risk that solutions from these benchmarks have leaked into the pretraining data. This data contamination threatens the integrity of reported evaluation results, as it obscures whether models are demonstrating true reasoning and generalization capabilities or merely reciting memorized code seen during training.

The article establishes a method to quantify the extent of benchmark contamination in open-access code training corpora and evaluates its real-world impact on model performance. The primary objective is to measure how much benchmark overlap exists in pretraining datasets and to determine how strongly memorized code skews model evaluation metrics.

To conduct this assessment, the authors developed a two-stage matching pipeline combining surface-level text comparisons and semantic structural analysis using abstract syntax trees. They evaluated three distinct model families across two leading pretraining datasets—the GitHub subset of the Pile and the Python split of StarCoderData. Using exact reference solutions from the widely used HumanEval and MBPP benchmark datasets, the authors scanned these massive corpora using sliding substrings to capture both exact text matches and semantically identical programs that use different variable names or formatting.

The investigation produced several key findings. First, significant contamination exists within standard pretraining corpora: between 3.6% and 20.8% of benchmark problem solutions appeared in identical or near-identical form within the examined training sets. Second, models performed dramatically better on benchmark problems they had encountered during pretraining. For example, StarCoderBase achieved a 72.0% accuracy on the top 10% most seen MBPP questions, compared to only 22.0% on the least seen 10%. Third, removing contaminated examples caused notable performance drops across all models—reducing StarCoderBase accuracy on HumanEval by up to 50.2% and Pythia accuracy by up to 70.4%—which substantially narrowed the apparent performance gaps between different model families. Finally, the analysis confirmed that this performance boost was driven by memorization rather than underlying problem simplicity or program length.

These findings imply that reported leaderboards and benchmark scores may significantly overstate the actual programming competence and generalizability of modern language models. For organizations adopting or deploying these models, contaminated evaluations introduce operational and security risks, as models may fail unexpectedly when exposed to novel, production-grade tasks not present in their training corpora. Apparent performance advantages of larger models may also be partly an artifact of greater memorization capacity rather than superior reasoning.

The article recommends that researchers and developers adopt semantic-aware decontamination protocols before pretraining new models, rather than relying exclusively on simple text-string deduplication. Furthermore, future benchmark evaluations should report performance on decontaminated subsets and cross-reference results across distinct model families to isolate genuine generalization from rote memorization.

The conclusions carry high confidence based on empirical evidence from open datasets, though the study notes several constraints. The search was restricted to specific language splits to manage computational costs, and the pipeline only queried one reference solution per problem, meaning the reported figures represent a conservative lower bound on true contamination levels. Consequently, decision-makers should treat standard benchmark metrics with caution when selecting models for real-world software engineering workflows.

  • Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Read this account of program synthesis first to understand how HumanEval and MBPP became execution-based code-generation benchmarks whose scores this study reexamines.
  • Paper: StarCoder: may the source be with you!, Raymond Li et al. (2023). Its description of StarCoder’s training data and corpus preparation provides useful context for the StarCoderData pretraining corpus whose benchmark overlap this study measures.
Cover for Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Abstract

While large language models have achieved remarkable performance on various code generation benchmarks, there have been growing concerns regarding potential contamination of these benchmarks as they may be leaked into pretraining and finetuning data. While recent work has investigated contamination in natural language generation and understanding tasks, there has been less extensive research into how data contamination impacts the evaluation of code generation, which is critical for understanding the robustness and reliability of LLMs in programming contexts. In this work, we perform a comprehensive study of data contamination of popular code generation benchmarks, and precisely quantify their overlap with pretraining corpus through both surface-level and semantic-level matching. In our experiments, we show that there are substantial overlap between popular code generation benchmarks and open training corpus, and models perform significantly better on the subset of the benchmarks where similar solutions are seen during training. We also conduct extensive analysis on the factors that affects model memorization and generalization, such as model size, problem difficulty, and question length. We release all resulting files from our matching pipeline for future research¹.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Measuring Program Similarity
  • 2.2 Quantifying Data Contamination
  • 3 Experimental Setup
  • 3.1 Models and Pretraining Data
  • 3.2 Benchmarks
  • 4 Results
  • 4.1 Main Results
  • 4.2 Analysis
  • 4.3 Case Study
  • 5 Related Work
  • 6 Discussions
  • 7 Limitations
  • 8 Conclusion
  • Acknowledgements
  • References
  • A Additional Examples and Results
  • A.1 All Model Series
  • A.2 Relevant Info for Models on the HumanEval Benchmark
  • A.3 Examples of Similarity Scores
  • A.4 Perfect Matches

Knowls

  1. Knowl 1 — Substring-level surface and semantic matching for code contamination

    model/method

    The contamination-search pipeline compares each benchmark’s gold program with substrings of the relevant pretraining-corpus files, rather than whole files, because a file may contain unrelated functions that obscure a match. It scans substrings the same length as the gold program and first ranks candidates by normalized Levenshtein similarity. For each benchmark problem, the pipeline then computes Dolos similarity for the 500 highest-ranked surface matches. Dolos uses tree-sitter to canonicalize code into abstract syntax tree representations and compares k-grams, making its score less sensitive to formatting and identifier changes than surface matching. The two scores are combined as S(p,p∗)=max⁡(Ssurface(p,p∗),Ssemantic(p,p∗))S(p,p^*)=\max(S_{\mathrm{surface}}(p,p^*),S_{\mathrm{semantic}}(p,p^*)), where pp is a candidate training substring, p∗p^* is the benchmark gold program, and each component is a similarity score on a 0–100 scale. Taking the maximum allows either textual resemblance or structural resemblance to flag a candidate.

  2. Knowl 2 — Similarity scores as proxies for benchmark exposure

    definition

    For a benchmark problem, the study calls its gold solution “seen” when the highest aggregated similarity score among training-corpus matches is exactly 100. It also treats lower top-match thresholds, especially scores greater than 90 or 80, as progressively broader indicators of possible contamination. To compare performance across degrees of overlap, the study ranks problems using the average of their ten highest aggregated similarity scores. These are corpus-match-based exposure measures: a score of 100 establishes a perfect match under at least one of the two similarity metrics, while a high but imperfect score indicates a candidate similar solution rather than proof that the model memorized or reproduced it.

  3. Knowl 3 — Benchmarks, corpora, and model families evaluated

    experimental setup

    The study evaluates the 500-question test split of MBPP and all 164 HumanEval problems, both of which provide gold Python solutions used as search queries. The searched corpora are the PILE and STARCODERDATA. For the PILE, the search uses its 95.16-GiB GitHub split; for STARCODERDATA, it searches 60.40 GB of the Python split. The model families are Pythia (1.4B, 2.8B, 6.9B, and 12B parameters) and CodeGen-NL (350M, 2B, 6B, and 16B), both trained on the PILE, and StarCoderBase (1B, 3B, 7B, and 15.5B), trained on STARCODERDATA. The study focuses on base models rather than instruction-tuned or RLHF-trained models, to avoid conflating pretraining-corpus exposure with later training stages.

  4. Knowl 4 — Perfect gold-program matches occur in both benchmark corpora

    empirical result

    Using an aggregated top-match score of 100 as the criterion for a seen gold solution, the search finds exact surface or semantic matches for 3.6% of MBPP problems and 12.2% of HumanEval problems in the searched PILE data. In the searched STARCODERDATA, the corresponding shares are 20.8% for MBPP and 18.9% for HumanEval. Thus, the detected overlap is substantial in both corpora, but differs by benchmark: the MBPP overlap is much higher in STARCODERDATA than in the PILE, whereas HumanEval has detected perfect matches in both.

  5. Knowl 5 — Models are more accurate on problems with greater training-data overlap

    empirical result

    For the largest model in each family, the study compares accuracy on the top and bottom 10% of benchmark problems ranked by the average of their ten highest aggregated similarity scores. On MBPP, StarCoderBase-15.5B scores 72.0% on the top decile and 22.0% on the bottom decile, a 50.0-percentage-point gap; Pythia-12B scores 40.0% versus 8.0% (42.0 points); and CodeGen-NL-16B scores 48.0% versus 6.0% (42.0 points). On HumanEval, StarCoderBase-15.5B scores 75.0% versus 31.3% (43.7 points), Pythia-12B scores 56.3% versus 0.0% (56.3 points), and CodeGen-NL-16B scores 62.5% versus 0.0% (62.5 points). The consistent advantage on the high-overlap deciles shows a strong association between similarity to training-corpus code and evaluation performance.

  6. Knowl 6 — Removing potentially contaminated problems reduces measured accuracy

    empirical result

    The study recalculates accuracy after removing problems whose top aggregated similarity score meets each stated threshold. “Removed” is the percentage of the benchmark excluded; the relative accuracy degradation is reproduced in parentheses. The thresholds are a perfect score of 100, a score greater than 90, and a score greater than 80.

    • MBPP — StarCoderBase-15.5B: original accuracy 41.6%; score 100: 20.8% removed, adjusted accuracy 33.8% (-18.8%); score >90: 32.2% removed, 32.5% (-22.6%); score >80: 50.8% removed, 29.7% (-28.6%).
    • MBPP — Pythia-12B: original 17.8%; score 100: 3.6% removed, 17.0% (-4.5%); score >90: 6.9% removed, 16.6% (-6.7%); score >80: 11.4% removed, 15.8% (-11.2%).
    • MBPP — CodeGen-NL-16B: original 19.6%; score 100: 3.6% removed, 18.4% (-6.1%); score >90: 6.9% removed, 17.4% (-11.2%); score >80: 11.4% removed, 16.5% (-15.8%).
    • HumanEval — StarCoderBase-15.5B: original 30.5%; score 100: 18.9% removed, 22.6% (-25.9%); score >90: 39.6% removed, 15.2% (-50.2%); score >80: 63.4% removed, 20.0% (-34.4%).
    • HumanEval — Pythia-12B: original 9.8%; score 100: 12.2% removed, 4.2% (-57.1%); score >90: 15.9% removed, 2.9% (-70.4%); score >80: 29.9% removed, 1.7% (-82.7%).
    • HumanEval — CodeGen-NL-16B: original 14.6%; score 100: 12.2% removed, 8.3% (-43.2%); score >90: 15.9% removed, 5.8% (-60.3%); score >80: 29.9% removed, 3.5% (-76.0%).

    Across these settings, removing high-similarity examples lowers reported accuracy. The authors also report that decontamination narrows some model-performance gaps: on MBPP, the original 23.8-point gap between StarCoderBase-15.5B and Pythia-12B falls to 13.9 points after removing problems with perfect matches.

  7. Knowl 7 — Cross-model comparison suggests overlap is not explained only by easier questions

    empirical result

    To assess whether high-overlap questions are simply easier, the study compares models trained on different corpora against the same question subsets. In MBPP, 104 of 500 problems are flagged as overlapping with STARCODERDATA. StarCoderBase-15.5B scores 71.2% on that subset and 33.8% on its complement, while CodeGen-NL-16B scores only 11.5% on the same 104 problems, below its 19.6% overall MBPP accuracy. This cross-model comparison indicates that the StarCoderBase advantage on its overlap subset is not readily explained by those questions being intrinsically easy. The comparison is suggestive rather than a controlled causal test, since the models differ in training and capabilities. The study also finds 16 HumanEval questions in the intersection of the subsets flagged as seen in STARCODERDATA and the PILE, compared with only 2 such MBPP questions.

  8. Knowl 8 — Larger models tend to perform better even on high-overlap subsets

    empirical result

    Across the StarCoderBase, Pythia, and CodeGen-NL model families, the study evaluates accuracy on subsets formed by increasing the minimum average of the top ten aggregated similarity scores. Within each family, larger models generally achieve higher accuracy than smaller models, including on subsets with stronger training-data overlap. The authors interpret this pattern as evidence that larger models perform better both on generalization and on problems resembling training data. Accuracy curves become noisier as the overlap threshold rises and the evaluated subset shrinks; some models consequently show 0% accuracy on very small subsets, including subsets containing problems with many similar training matches.

  9. Knowl 9 — Gold-program length shows no apparent relationship to overlap or accuracy

    empirical result

    For StarCoderBase-15.5B on MBPP, the study compares each gold program’s character length with its average top-ten aggregated similarity score and with whether the model’s prediction is correct. It reports no apparent correlation between program length and similarity score, or between length and accuracy. A corresponding HumanEval analysis is also presented and described as showing similar results. These observations address the concern that longer gold programs might receive artificially high similarity scores despite containing more differences.

  10. Knowl 10 — Exposure matches are imperfect evidence of memorization or causal benefit

    limitation

    The search detects similarity to one gold solution per problem, although a programming task can have many valid solutions; accordingly, the paper treats its exposure counts as minimum estimates. Search is also restricted to the PILE GitHub split and the STARCODERDATA Python split, so similar code in unsearched corpus portions can be missed. Conversely, AST-based similarity can classify programs as similar even when their behavior differs, creating false positives. The study does not retrain models after removing matched solutions, so it cannot isolate the causal effect of those solutions by direct intervention. Its case study further shows that a high match count does not guarantee correct generation: StarCoderBase sometimes fails on problems with similar solutions appearing at least ten times, with the authors noting that associating the natural-language description with the code may remain difficult. The analysis is limited to two benchmarks and three model families with publicly searchable training data.

Coverage note — The illustrative individual failure examples and the appendix’s full inventory of matched benchmark programs are omitted because they do not add a distinct general result beyond the summarized findings and limitations.

References

  1. 1.Miltiadis Allamanis. 2019. The adverse effects of code duplication in machine learning models of code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, Onward! 2019, page 143–153, New York, NY, USA. Association for Computing Machinery.
  2. 2.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. Program synthesis with large language models.
  3. 3.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 2397–2430. PMLR.
  4. 4.Terra Blevins and Luke Zettlemoyer. 2022. Language contamination helps explains the cross-lingual capabilities of English pretrained models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3563–3574, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  5. 5.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models.
  6. 6.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2020. Extracting training data from large language models. In USENIX Security Symposium.
  7. 7.Kent K. Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023. Speak, memory: An archaeology of books known to chatgpt/gpt-4.
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  10. 10.Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. 2023. Investigating data contamination in modern benchmarks for large language models.
  11. 11.Jesse Dodge, Maarten Sap, Ana Marasovic, William ´ Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, and Ge Li. 2024. Generalization or memorization: Data contamination and trustworthy evaluation for large language models. arXiv preprint arXiv:2402.15938.
  13. 13.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The pile: An 800gb dataset of diverse text for language modeling.
  14. 14.Shahriar Golchin and Mihai Surdeanu. 2023. Time travel in llms: Tracing data contamination in large language models.
  15. 15.Roger Baker Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamil.e Lukovsiut.e, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Sam Bowman. 2023. Studying large language model generalization with influence functions. ArXiv, abs/2308.03296.
  16. 16.Xiaochuang Han and Yulia Tsvetkov. 2022. Orca: Interpreting prompted language models via locating supporting data evidence in the ocean of pretraining data. ArXiv, abs/2205.12600.
  17. 17.Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau. 2017. Ethical challenges in data-driven dialogue systems. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society.
  18. 18.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models.
  19. 19.Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo. 2024. Investigating data contamination for pre-training language models.
  20. 20.Nikhil Kandpal, H. Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022a. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning.
  21. 21.Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022b. Deduplicating training data mitigates privacy risks in language models.
  22. 22.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.
  23. 23.Anjan Karmakar, Julian Aron Prenner, Marco D’Ambros, and Romain Robbes. 2022. Codex hacks hackerrank: Memorization issues and a framework for code synthesis evaluation.
  24. 24.Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2022. The stack: 3 tb of permissively licensed source code.
  25. 25.Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee. 2022. Do language models plagiarize? Proceedings of the ACM Web Conference 2023.
  26. 26.Vladimir I. Levenshtein. 1965. Binary codes capable of correcting deletions, insertions, and reversals. Soviet physics. Doklady, 10:707–710.
  27. 27.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023. Starcoder: may the source be with you!
  28. 28.Rien Maertens, Charlotte Van Petegem, Niko Strijbol, Toon Baeyens, Arne Carla Jacobs, Peter Dawyndt, and Bart Mesuere. 2022. Dolos: Language-agnostic plagiarism detection in source code. Journal of Computer Assisted Learning, 38(4):1046–1061.
  29. 29.Inbal Magar and Roy Schwartz. 2022. Data contamination: From memorization to exploitation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 157–165, Dublin, Ireland. Association for Computational Linguistics.
  30. 30.Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre. 2023. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124.
  31. 31.Ansong Ni, Pengcheng Yin, Yilun Zhao, Martin Riddell, Troy Feng, Rui Shen, Stephen Yin, Ye Liu, Semih Yavuz, Caiming Xiong, et al. 2023. L2ceval: Evaluating language-to-code generation capabilities of large language models. arXiv preprint arXiv:2309.17446.
  32. 32.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. Codegen: An open large language model for code with multi-turn program synthesis.
  33. 33.Yonatan Oren and Nicole Meister. 2023. Proving test set contamination in black box language models.
  34. 34.Zhen Peng, Zhizhi Wang, and Dong Deng. 2023. Near-duplicate sequence search at scale for large language model memorization evaluation. Proceedings of the ACM on Management of Data, 1:1 – 18.
  35. 35.Lutz Prechelt and Guido Malpohl. 2003. Finding plagiarisms among a set of programs with jplag. Journal of Universal Computer Science, 8.
  36. 36.Federico Ranaldi, Elena Sofia Ruzzetti, Dario Onorati, Leonardo Ranaldi, Cristina Giannone, Andrea Favalli, Raniero Romagnoli, and Fabio Massimo Zanzotto. 2024. Investigating the impact of data contamination of large language models in text-to-sql translation.
  37. 37.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 840–854, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  38. 38.Sandip Sarkar, Dipankar Das, Partha Pakray, and Alexander Gelbukh. 2016. JUNITMZ at SemEval-2016 task 1: Identifying semantic similarity using Levenshtein ratio. In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), pages 702–705, San Diego, California. Association for Computational Linguistics.
  39. 39.Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2023. Detecting pretraining data from large language models.
  40. 40.Om Thakkar, Swaroop Indra Ramaswamy, Rajiv Mathews, and Franccoise Beaufays. 2020. Understanding unintended memorization in federated learning. ArXiv, abs/2006.07490.
  41. 41.Aleena Anna Thomas, David Ifeoluwa Adelani, Ali Davody, Aditya Mogadala, and Dietrich Klakow. 2020. Investigating the impact of pre-trained word embeddings on memorization in neural networks. In Workshop on Time-Delay Systems.
  42. 42.Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez, and Ion Stoica. 2023. Rethinking benchmark and contamination for language models with rephrased samples.
  43. 43.Zhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang, Yevgeniy Vorobeychik, and Chaowei Xiao. 2023. CodeIPPrompt: Intellectual property infringement assessment of code language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 40373–40389. PMLR.

Citation

MLA
Riddell, M., et al. “Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14116–37, https://doi.org/10.18653/v1/2024.acl-long.761.
APA
Riddell, M., Ni, A., & Cohan, A. (2024). Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14116–14137. https://doi.org/10.18653/v1/2024.acl-long.761
Chicago
Riddell, M., A. Ni, and A. Cohan. 2024. “Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14116–37. https://doi.org/10.18653/v1/2024.acl-long.761.
Harvard
Riddell, M., Ni, A. and Cohan, A. (2024) “Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14116–14137. Available at: https://doi.org/10.18653/v1/2024.acl-long.761.
Vancouver
1. Riddell M, Ni A, Cohan A (2024) Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14116–14137

BibTeX

@inproceedings{riddell-etal-2024-quantifying,
    title = "Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models",
    author = "Riddell, Martin  and
      Ni, Ansong  and
      Cohan, Arman",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.761/",
    doi = "10.18653/v1/2024.acl-long.761",
    pages = "14116--14137"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/