Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
Martin RiddellAnsong NiArman Cohan
Quantifies benchmark leakage in popular code evaluation sets like HumanEval and MBPP across major pretraining corpora using both surface and semantic matching, revealing how test-data memorization artificially inflates code generation performance.
Evaluating the performance of large language models on programming tasks relies heavily on standard benchmarks. However, as training datasets expand to include massive repositories of open-source source code, there is a substantial risk that solutions from these benchmarks have leaked into the pretraining data. This data contamination threatens the integrity of reported evaluation results, as it obscures whether models are demonstrating true reasoning and generalization capabilities or merely reciting memorized code seen during training.
The article establishes a method to quantify the extent of benchmark contamination in open-access code training corpora and evaluates its real-world impact on model performance. The primary objective is to measure how much benchmark overlap exists in pretraining datasets and to determine how strongly memorized code skews model evaluation metrics.
To conduct this assessment, the authors developed a two-stage matching pipeline combining surface-level text comparisons and semantic structural analysis using abstract syntax trees. They evaluated three distinct model families across two leading pretraining datasets—the GitHub subset of the Pile and the Python split of StarCoderData. Using exact reference solutions from the widely used HumanEval and MBPP benchmark datasets, the authors scanned these massive corpora using sliding substrings to capture both exact text matches and semantically identical programs that use different variable names or formatting.
The investigation produced several key findings. First, significant contamination exists within standard pretraining corpora: between 3.6% and 20.8% of benchmark problem solutions appeared in identical or near-identical form within the examined training sets. Second, models performed dramatically better on benchmark problems they had encountered during pretraining. For example, StarCoderBase achieved a 72.0% accuracy on the top 10% most seen MBPP questions, compared to only 22.0% on the least seen 10%. Third, removing contaminated examples caused notable performance drops across all models—reducing StarCoderBase accuracy on HumanEval by up to 50.2% and Pythia accuracy by up to 70.4%—which substantially narrowed the apparent performance gaps between different model families. Finally, the analysis confirmed that this performance boost was driven by memorization rather than underlying problem simplicity or program length.
These findings imply that reported leaderboards and benchmark scores may significantly overstate the actual programming competence and generalizability of modern language models. For organizations adopting or deploying these models, contaminated evaluations introduce operational and security risks, as models may fail unexpectedly when exposed to novel, production-grade tasks not present in their training corpora. Apparent performance advantages of larger models may also be partly an artifact of greater memorization capacity rather than superior reasoning.
The article recommends that researchers and developers adopt semantic-aware decontamination protocols before pretraining new models, rather than relying exclusively on simple text-string deduplication. Furthermore, future benchmark evaluations should report performance on decontaminated subsets and cross-reference results across distinct model families to isolate genuine generalization from rote memorization.
The conclusions carry high confidence based on empirical evidence from open datasets, though the study notes several constraints. The search was restricted to specific language splits to manage computational costs, and the pipeline only queried one reference solution per problem, meaning the reported figures represent a conservative lower bound on true contamination levels. Consequently, decision-makers should treat standard benchmark metrics with caution when selecting models for real-world software engineering workflows.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Read this account of program synthesis first to understand how HumanEval and MBPP became execution-based code-generation benchmarks whose scores this study reexamines.
- Paper: StarCoder: may the source be with you!, Raymond Li et al. (2023). Its description of StarCoder’s training data and corpus preparation provides useful context for the StarCoderData pretraining corpus whose benchmark overlap this study measures.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). After seeing how training overlap inflates established code-benchmark scores, read this follow-up’s fresh, time-stamped problems for an evaluation design intended to limit contamination.
