DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Yuhang LaiChengxi LiYiming WangTianyi ZhangRuiqi ZhongLuke ZettlemoyerWen-Tau YihDaniel FriedSida I. WangTao Yu

article2023ICML622 citations

Introduces DS-1000, an execution-based benchmark of 1,000 realistic data science problems across seven Python libraries that prevents memorization through problem perturbations and evaluates code generation with multi-criteria correctness constraints.

Listen

Data science programming presents high entry barriers for non-specialists due to the complexity of specialized software libraries. While artificial intelligence code generation models show promise in lowering these barriers, existing evaluation benchmarks primarily target competition-style programming puzzles or rely on superficial text-matching metrics. Consequently, they fail to represent real-world data science workflows or reliably verify whether generated code actually executes correctly. The article introduces and evaluates DS-1000, a benchmark designed to assess pre-trained code models on realistic, diverse data science problems using rigorous execution-based testing while defending against data memorization.

To build DS-1000, the authors curated and modified 1,000 realistic programming problems derived from 451 unique Stack Overflow queries spanning seven major Python data science libraries: NumPy, Pandas, Matplotlib, Scikit-learn, SciPy, TensorFlow, and PyTorch. Expert annotators rewrote the problems into executable contexts and implemented a multi-criteria evaluation framework. This framework executes model-generated code against an average of 1.6 functional test cases per problem and enforces surface-form constraints (such as forbidding inefficient loops or requiring specific library functions). Furthermore, to prevent models from succeeding via rote memorization of public internet data, the authors applied surface perturbations, semantic modifications, and difficult rewrites to more than half of the benchmark problems.

The benchmark analysis yielded several key findings regarding model capabilities and evaluation accuracy. First, current leading code generation models exhibit substantial room for improvement: the top-performing system, Codex-002 using code insertion, achieved only 43.3% accuracy, while smaller 6-billion-parameter open-source models scored under 10%. Second, the multi-criteria execution framework proved highly dependable, exhibiting a low sample-level false discovery rate of only 1.8% on Codex-002 predictions. Third, infilling or insertion-based prompting—which provides both preceding and following code context—improved Codex-002's accuracy by 4.1% over standard left-to-right generation. Finally, problem perturbations confirmed that models rely heavily on training set memorization; on a popular external NumPy problem set, perturbing problem wording caused Codex-002 accuracy to plummet from 72.5% to 23.6%, demonstrating the necessity of perturbed benchmarks.

These findings indicate that while modern artificial intelligence can assist with everyday data analysis, enterprise deployment still carries significant operational and reliability risks if outputs are unverified. The wide performance disparity across libraries and problem formulations shows that model competence does not automatically transfer across different tools. Additionally, standard benchmarks that do not alter public training questions risk severely overstating model capabilities, which could lead organizations to misjudge deployment readiness and safety.

Organizations developing or deploying code generation assistants should adopt multi-criteria execution testing that pairs functional verification with structural constraints. System designers should prioritize infilling capabilities over strict left-to-right code generation. For future research, the community should expand evaluation suites into complex multi-file workflows and develop automated methods to assist experts in writing robust verification test suites.

Confidence in these findings is high regarding the evaluated Python libraries and the documented baseline models. However, limitations remain: the benchmark does not cover multimodal figure interpretation for visualization tasks, omits untestable questions involving software installation or conceptual explanations, and reflects a snapshot of specific library versions pinned to Python 3.7.10.

Cover for DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Abstract

We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as NumPy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io.

Table of Contents

  • 1. Introduction
  • 2. Benchmark Construction
  • 2.1. Problem Selection
  • 2.2. Rewriting Problems and Reference Solutions
  • 2.3. Implementing Multi-Criteria Evaluations
  • 2.4. Perturbation to Defend Against Memorization
  • 2.5. Quality Assurance
  • 3. Dataset Statistics
  • 4. Benchmarking State-of-the-Art Models
  • 4.1. Prompt Format
  • 4.2. Experimental Setup
  • 4.3. Main Results
  • 4.4. Results by Perturbation
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgements
  • References
  • Appendices
  • A. Details on Data Collection
  • A.1. Problem Selection
  • A.2. Example Problems
  • A.3. Problem Perturbation
  • A.4. Prompt Format
  • B. Details of Experiments on numpy-100
  • C. Error Analysis

Knowls

  1. Knowl 1 — DS-1000 Benchmark Dataset and Library Composition

    data/table

    DS-1000 is a code generation benchmark comprising 1,000 distinct data science programming problems collected and adapted from 451 seed questions on StackOverflow across seven core Python libraries: Pandas, NumPy, Matplotlib, Scikit-learn, SciPy, TensorFlow, and PyTorch. To defend against verbatim memorization by pre-trained language models, over half of the benchmark consists of perturbed variations of the original StackOverflow questions, including surface perturbations, semantic perturbations, and difficult rewrites.

    Metric / Category Pandas NumPy Matplotlib Scikit-learn SciPy TensorFlow PyTorch Total/Avg.
    Total Problems 291 220 155 115 106 45 68 1000
    Origin 100 97 111 46 58 17 22 451
    Surface Perturbation 24 22 0 57 11 11 27 152
    Semantic Perturbation 88 51 44 9 20 12 11 235
    Difficult Rewrite 79 50 0 3 17 5 8 162
    % Surface-Form Constraints 12.0% 36.4% 0.0% 27.8% 17.9% 20.0% 27.9% 19.4%
    Avg. Test Cases 1.7 2.0 1.0 1.5 1.6 1.6 1.7 1.6
    Avg. Problem Words 184.8 137.5 21.1 147.3 192.4 133.3 133.4 140.0
    Avg. Lines of Code Context 9.0 8.3 6.9 11.0 10.2 9.2 9.0 8.9
    Avg. Lines of Code Solution 5.4 2.5 3.0 3.3 3.1 4.1 2.1 3.6

    Unlike algorithmic programming datasets, data science tasks involve complex input structures (e.g., matrices, dataframes, classifier objects). Across the benchmark, each problem has an average of 1.6 manually constructed test cases, 140.0 prompt words, and 8.9 lines of executable context. Matplotlib problems are formatted purely in text and code comments without images, resulting in shorter descriptions (21.1 words on average).

  2. Knowl 2 — Multi-Criteria Execution-Based Evaluation Methodology

    model/method

    DS-1000 evaluates generated code using a multi-criteria execution framework that tests both functional correctness and programmatic surface-form constraints:

    1. Functional Correctness via Domain-Specific Assertions: Rather than relying solely on strict equality matching, evaluation scripts accommodate data science nuances:

      • Floating-Point and Rounding Tolerances: Approximate numerical matching allows small floating-point discrepancies.
      • Stochastic Outputs: For functions generating random distributions (e.g., log-uniform sampling), the two-sample Kolmogorov-Smirnov test is applied to assert that predicted samples and ground-truth samples originate from statistically identical distributions (p>0.1p > 0.1).
      • Sparse Matrix State: For SciPy sparse structures, evaluation verifies matrix type, element values, and the count of explicitly stored non-zero elements (nnz).
      • Matplotlib Plot Properties: Output figure comparisons are performed via exact image array matching; if images differ, abstract syntax tree (AST) inspections verify properties of the Matplotlib axis object (e.g., presence and styling of grid lines or line colors).
    2. Surface-Form Constraints via AST Checking: To prevent trivial or inefficient implementations, the evaluation inspects the AST of predicted solutions. For example, AST rules disallow for and while loop syntax when vectorized operations are required, or forbid densifying calls (such as .toarray() or .todense()) on sparse matrices.

  3. Knowl 3 — Large Language Model Performance Benchmark on DS-1000

    data/table

    Evaluation of pre-trained language models on DS-1000 demonstrates substantial capability differences across model families, libraries, and prompt configurations. Performance is measured using the unbiased estimator for pass@1\text{pass}@1 calculated over 40 generated samples per problem with generation temperature T=0.2T = 0.2, top-p=0.95p = 0.95, and a maximum sequence length of 1024 tokens.

    Format Model Pandas NumPy Matplotlib Scikit-learn SciPy TensorFlow PyTorch Overall
    Completion Codex-002 26.5 43.1 57.0 44.8 31.8 39.3 41.8 39.2
    Codex-001 9.4 26.6 41.8 18.5 15.0 17.2 9.7 20.2
    Codex-Cushman 7.9 21.8 40.7 18.0 11.3 12.2 12.4 18.1
    CodeGen-6B 1.9 12.1 18.6 5.8 7.4 12.8 3.4 8.4
    InCoder-6B 3.1 4.4 28.3 2.8 2.8 3.8 4.4 7.4
    Insertion Codex-002 30.1 46.5 57.0 53.7 34.8 53.4 47.7 43.3
    InCoder-6B 2.9 4.6 28.3 3.1 3.1 7.8 3.2 7.5

    *Note: Matplotlib does not possess a succeeding code context, so its Completion and Insertion accuracies are identical (57.0%57.0\% for Codex-002 and 28.3%28.3\% for InCoder-6B).

    Codex-002 (OpenAI code-davinci-002) achieves the strongest overall accuracy (43.3%43.3\% in Insertion format), while smaller 6B open models (CodeGen-6B and InCoder-6B) score below 10%10\% overall, frequently failing to follow prompt framing by emitting natural language comments instead of executable code.

  4. Knowl 4 — Reliability and Error Validation of DS-1000 Automatic Evaluation

    empirical result

    The reliability of DS-1000's multi-criteria evaluation was evaluated by sampling 10 problems per library (70 total) and drawing 40 generations per problem from Codex-002 at temperature T=0.7T = 0.7, producing 2,800 problem-code samples that were manually audited against human expert decisions.

    The resulting sample-level and problem-level error rates are:

    • Sample-Level False Discovery Rate (FDR): 1.8%1.8\% of all model predictions passing the multi-criteria automatic evaluation were judged incorrect upon manual human inspection.
    • Sample-Level False Omission Rate (FOR): 0.5%0.5\% of all model predictions rejected by the automatic evaluation were judged to be functionally correct implementations.
    • Problem-Level False Positive Percentage: 5.7%5.7\% of evaluated problems contained at least one incorrect solution among the 40 candidate generations that erroneously passed automatic validation.
    • Problem-Level False Negative Percentage: 5.7%5.7\% of evaluated problems contained at least one correct solution that erroneously failed automatic validation.
  5. Knowl 5 — Problem Perturbation Taxonomy for Memorization Defense

    model/method

    To prevent pre-trained language models from solving benchmark problems via pre-training memorization of StackOverflow threads, DS-1000 employs a three-tier perturbation framework:

    1. Surface Perturbations: Modifications that alter prompt phrasing or surrounding context without altering the underlying target solution logic. Methods include:

      • Converting standalone code contexts into function-completion signatures.
      • Paraphrasing natural language descriptions.
      • Altering input data matrices and example values.
    2. Semantic Perturbations: Modifications that alter the reference solution logic without increasing problem difficulty for human programmers. Methods include:

      • Substituting operation keywords with mathematical analogues (e.g., replacing inverse with exponential calculations).
      • Changing target index specifications or ordinal selectors (e.g., zeroing the second row instead of the first).
      • Reversing required list, string, or DataFrame ordering.
      • Modifying expected container types (e.g., requiring a Pandas Series output instead of a DataFrame).
    3. Difficult Rewrites: Targeted semantic and syntactic overhauls designed intentionally to increase problem difficulty (e.g., requiring both statistical hypothesis execution and conditional threshold interpretation).

  6. Knowl 6 — Degradation of Model Accuracy Under Problem Perturbation

    data/table

    Evaluating pre-trained models on original versus perturbed tasks demonstrates that web-scraped code benchmarks are susceptible to memorization inflation. On a pilot study of 20 problems from the web-popular numpy-100 repository, Codex-002 achieved 72.5%72.5\% accuracy on original prompts, but dropped to 50.8%50.8\% on surface perturbations and 23.6%23.6\% on semantic perturbations (40.6%40.6\% average). In 36%36\% of semantic perturbation cases, Codex-002 outputted the memorized original answer verbatim.

    On DS-1000, Codex-002 pass@1 accuracy on original versus perturbed subsets shows smaller, but systematic drops across libraries:

    Subset Pandas NumPy Scikit-learn SciPy TensorFlow PyTorch Overall
    Originsurface\text{Origin}_{\text{surface}} 37.3 61.2 52.6 33.0 64.9 64.8 53.2
    Surface Perturbation 31.9 58.4 55.7 32.1 58.0 50.0 49.8
    Δsurface\Delta_{\text{surface}} -5.4 -2.8 +3.1 -0.9 -8.9 -14.8 -3.4
    Originsemantic\text{Origin}_{\text{semantic}} 36.8 56.7 60.6 40.3 71.3 65.1 47.2
    Semantic Perturbation 33.2 49.0 38.9 34.3 42.5 30.5 38.2
    Δsemantic\Delta_{\text{semantic}} -3.6 -7.7 -21.7 -6.0 -25.8 -34.6 -9.0
    Origindifficult\text{Origin}_{\text{difficult}} 39.9 52.7 5.0 58.1 73.0 53.8 46.8
    Difficult Rewrite 17.7 27.1 0.0 13.8 38.0 28.8 21.0
    Δdifficult\Delta_{\text{difficult}} -22.2 -25.6 -5.0 -44.3 -35.0 -25.0 -25.8

    *Note: Subsets evaluate only the specific problem instances selected for each perturbation type, so unperturbed origin baseline scores vary across rows.

  7. Knowl 7 — Insertion versus Completion Prompt Formats

    model/method

    DS-1000 defines two standard prompt formats for code generation models:

    1. Insertion Format (Infilling): The primary format. It presents a natural language description followed by code containing HTML-like delimiter tags (BEGIN SOLUTION, <code>, [insert], </code>, END SOLUTION). The prompt includes both left-context (preceding declarations and imports) and right-context (downstream variable usages and return formatting), enabling infilling models to observe expected variable names and return contracts directly.
    2. Completion Format (Left-to-Right): Designed for purely autoregressive models lacking infilling capabilities. It converts all right-context constraints (such as saving output to a specific variable name result) into natural language instructions prepended to the left-context.

    Autoregressive models using the Insertion format outperform those using the Completion format; for example, Codex-002 obtains an overall pass@1 accuracy of 43.3%43.3\% with Insertion versus 39.2%39.2\% with Completion (a 4.1%4.1\% gain), indicating that structural right-context is more effectively utilized than translated natural language descriptions.

  8. Knowl 8 — Comparison of DS-1000 to Existing Code Generation Benchmarks

    data/table

    Compared to general Python programming benchmarks and notebook-based assistants, DS-1000 features longer naturalistic intent descriptions, execution-based multi-criteria verification, and realistic data science library tasks.

    Dataset Problems Evaluation Metric Data Source Avg. Tests Avg. Words Avg. Code Lines
    HumanEval 164 Test Cases Hand-Written 7.7 23.0 6.3
    MBPP 974 Test Cases Hand-Written 3.0 15.7 6.7
    APPS 10000 Test Cases Competitions 13.2 293.2 18.0
    JuICe 1981 Exact Match + BLEU Jupyter Notebooks - 57.2 3.3
    DSP 1119 Test Cases Jupyter Notebooks 2.1 71.9 4.5
    CoNaLa 2879 BLEU StackOverflow - 13.8 1.1
    DS-1000 1000 Test Cases + Constraints StackOverflow 1.6 140.0 3.6

    Existing real-world data science datasets either rely on surface metrics (such as BLEU in CoNaLa and JuICe) or have brief problem statements, whereas DS-1000 provides 140.0140.0 average problem words reflecting real developer prompts, combined with reliable execution-based unit testing.

  9. Knowl 9 — DS-1000 Data Curation and Environment Standardization Pipeline

    experimental setup

    The construction of DS-1000 followed a 5-stage curation and standardization pipeline requiring approximately 1,200 human expert hours:

    1. Scraping and Popularity Filtering: Posts tagged with the target libraries were scraped from StackOverflow. Candidates required ≥1\ge 1 vote, ≥1000\ge 1000 views, and an accepted answer. Stratified sampling by creation year applied sliding vote/view thresholds to balance question age (e.g., 2011–2015 required ≥14\ge 14--5050 votes, whereas 2021–2022 required ≥1\ge 1 vote), yielding an initial pool of 4,500 questions.
    2. Suitability Filtering: Annotators scored questions on problem clarity, practical utility, presence of input-output examples, difficulty, and evaluability. Questions requiring hardware troubleshooting, environment/package installation, or textual explanations were filtered out, leaving 451 seed problems.
    3. Context Construction and Solution Repair: Authors wrote executable code prefixes and fixed outdated or buggy reference answers.
    4. Library Version Standardization: The execution runtime was standardized on Python 3.7.10 with pinned library versions:
      • NumPy 1.21.6
      • Pandas 1.3.5
      • Matplotlib 3.5.2
      • Seaborn 0.11.2
      • Scikit-learn 1.0.2
      • SciPy 1.7.3
      • TensorFlow 2.10.0
      • PyTorch 1.12.1
    5. Quality Assurance and Red Teaming: Each problem, reference solution, and evaluation script was independently audited by at least three Python data science experts and red-teamed by verifying that intentionally wrong and semantically perturbed programs fail test execution.

Coverage note — Individual illustrative error snippets from Appendix C (Figures 8, 27, 28) and individual screenshot examples of untestable StackOverflow questions from Appendix A (Figures 29-31) were omitted as their higher-level principles are fully subsumed within the evaluation, curation, and perturbation knowls.

References

  1. 1.Agashe, R., Iyer, S., and Zettlemoyer, L. JuICe: A large scale distantly supervised dataset for open domain context-based code generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 5436–5446, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1546. URL https://aclanthology.org/D19-1546.
  2. 2.Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al. Cm3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520, 2022.
  3. 3.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  4. 4.Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
  5. 5.Berant, J., Chou, A., Frostig, R., and Liang, P. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1533–1544, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://aclanthology.org/D13-1160.
  6. 6.Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pp. 95–136, virtual+Dublin, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.bigscience-1.9. URL https://aclanthology.org/2022.bigscience-1.9.
  7. 7.Bolyen, E., Rideout, J. R., Dillon, M. R., Bokulich, N. A., Abnet, C. C., Al-Ghalith, G. A., Alexander, H., Alm, E. J., Arumugam, M., Asnicar, F., et al. Reproducible, interactive, scalable and extensible microbiome data science using qiime 2. Nature biotechnology, 37(8):852–857, 2019.
  8. 8.Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T. B., Song, D., Erlingsson, Ú., Oprea, A., and Raffel, C. Extracting training data from large language models. In Bailey, M. and Greenstadt, R. (eds.), 30th USENIX Security Symposium, USENIX Security 2021, August 11-13, 2021, pp. 2633–2650. USENIX Association, 2021. URL https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting.
  9. 9.Chandel, S., Clement, C. B., Serrato, G., and Sundaresan, N. Training and evaluating a jupyter notebook data science assistant. arXiv preprint arXiv:2201.12901, 2022.
  10. 10.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021a.
  11. 11.Chen, X., Gong, L., Cheung, A., and Song, D. PlotCoder: Hierarchical decoding for synthesizing visualization code in programmatic context. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 2169–2181, Online, August 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.169. URL https://aclanthology.org/2021.acl-long.169.
  12. 12.Chu, S., Wang, C., Weitz, K., and Cheung, A. Cosette: An automated prover for SQL. In 8th Biennial Conference on Innovative Data Systems Research, CIDR 2017, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings. www.cidrdb.org, 2017. URL http://cidrdb.org/cidr2017/papers/p51-chu-cidr17.pdf.
  13. 13.Elangovan, A., He, J., and Verspoor, K. Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 1325–1335, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.113. URL https://aclanthology.org/2021.eacl-main.113.
  14. 14.Faghmous, J. H. and Kumar, V. A big data guide to understanding climate change: The case for theory-guided data science. Big data, 2(3):155—163, September 2014. ISSN 2167-6461. doi: 10.1089/big.2014.0026. URL https://europepmc.org/articles/PMC4174912.
  15. 15.Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W.-t., Zettlemoyer, L., and Lewis, M. Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999, 2022.
  16. 16.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J. Measuring coding challenge competence with apps. In Vanschoren, J. and Yeung, S. (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. Curran, 2021. URL https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/c24cd76e1ce41366a4bbe8a49b02a028-Paper-round2.pdf.
  17. 17.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022. doi: 10.1126/science.abq1158. URL https://www.science.org/doi/abs/10.1126/science.abq1158.
  18. 18.Liang, P., Jordan, M. I., and Klein, D. Learning dependency-based compositional semantics. Computational Linguistics, 39(2):389–446, June 2013. doi: 10.1162/COLI_a_00127. URL https://aclanthology.org/J13-2005.
  19. 19.Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C. A conversational paradigm for program synthesis. CoRR, abs/2203.13474, 2022.
  20. 20.Poesia, G., Polozov, A., Le, V., Tiwari, A., Soares, G., Meek, C., and Gulwani, S. Synchromesh: Reliable code generation from pre-trained language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=KmtVD97J43e.
  21. 21.Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A., and Ma, S. Codebleu: a method for automatic evaluation of code synthesis. CoRR, abs/2009.10297, 2020. URL https://arxiv.org/abs/2009.10297.
  22. 22.Romero, C. and Ventura, S. Data mining in education. Wiley Int. Rev. Data Min. and Knowl. Disc., 3(1):12–27, jan 2013. ISSN 1942-4787. doi: 10.1002/widm.1075. URL https://doi.org/10.1002/widm.1075.
  23. 23.Scholak, T., Schucher, N., and Bahdanau, D. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 9895–9901, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.779. URL https://aclanthology.org/2021.emnlp-main.779.
  24. 24.Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., and Wang, S. I. Natural language to code translation with execution. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 3533–3546, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.231.
  25. 25.Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Bahri, D., Schuster, T., Zheng, H. S., Houlsby, N., and Metzler, D. Unifying language learning paradigms. CoRR, abs/2205.05131, 2022. doi: 10.48550/arXiv.2205.05131. URL https://doi.org/10.48550/arXiv.2205.05131.
  26. 26.Tufano, M., Drain, D., Svyatkovskiy, A., Deng, S. K., and Sundaresan, N. Unit test case generation with transformers and focal context. arXiv preprint arXiv:2009.05617, 2020.
  27. 27.Wang, X., Wu, Q., Zhang, H., Lyu, C., Jiang, X., Zheng, Z., Lyu, L., and Hu, S. Heloc: Hierarchical contrastive learning of source code representation. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, ICPC ’22, pp. 354–365, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450392983. doi: 10.1145/3524610.3527896. URL https://doi.org/10.1145/3524610.3527896.
  28. 28.Xu, F. F., Alon, U., Neubig, G., and Hellendoorn, V. J. A systematic evaluation of large language models of code. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming, MAPS 2022, pp. 1–10, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450392730. doi: 10.1145/3520312.3534862. URL https://doi.org/10.1145/3520312.3534862.
  29. 29.Yin, P., Deng, B., Chen, E., Vasilescu, B., and Neubig, G. Learning to mine aligned code and natural language pairs from stack overflow. In Proceedings of the 15th International Conference on Mining Software Repositories, MSR ’18, pp. 476–486, New York, NY, USA, 2018. Association for Computing Machinery. ISBN 9781450357166. doi: 10.1145/3196398.3196408. URL https://doi.org/10.1145/3196398.3196408.
  30. 30.Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., and Radev, D. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 3911–3921, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1425. URL https://aclanthology.org/D18-1425.
  31. 31.Zelle, M. and Mooney, R. J. Learning to parse database queries using inductive logic programming. In Association for the Advancement of Artificial Intelligence (AAAI), pp. 1050–1055, 1996.
  32. 32.Zettlemoyer, L. and Collins, M. Online learning of relaxed CCG grammars for parsing to logical form. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pp. 678–687, Prague, Czech Republic, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/D07-1071.
  33. 33.Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S. Calibrate before use: Improving few-shot performance of language models. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 12697–12706. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/zhao21c.html.
  34. 34.Zhong, R., Yu, T., and Klein, D. Semantic evaluation for text-to-SQL with distilled test suites. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 396–411, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.29. URL https://aclanthology.org/2020.emnlp-main.29.
  35. 35.Zhou, S., Alon, U., Agarwal, S., and Neubig, G. Codebertscore: Evaluating code generation with pretrained models of code. CoRR, abs/2302.05527, 2023. doi: 10.48550/arXiv.2302.05527. URL https://doi.org/10.48550/arXiv.2302.05527.

Citation

MLA
Lai, Y., et al. “DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation”. International Conference on Machine Learning, vol. 202, 2023, pp. 18319–45, https://proceedings.mlr.press/v202/lai23b.html.
APA
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W.-T., Fried, D., Wang, S., & Yu, T. (2023). DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. International Conference on Machine Learning, 202, 18319–18345. https://proceedings.mlr.press/v202/lai23b.html
Chicago
Lai, Y., C. Li, Y. Wang, et al. 2023. “DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation”. International Conference on Machine Learning 202: 18319–45. https://proceedings.mlr.press/v202/lai23b.html.
Harvard
Lai, Y. et al. (2023) “DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation”, International Conference on Machine Learning. PMLR, pp. 18319–18345. Available at: https://proceedings.mlr.press/v202/lai23b.html.
Vancouver
1. Lai Y, Li C, Wang Y, Zhang T, Zhong R, Zettlemoyer L, Yih W-T, Fried D, Wang S, Yu T (2023) DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. In: International Conference on Machine Learning. PMLR, pp 18319–18345

BibTeX

@InProceedings{pmlr-v202-lai23b,
  title = 	 {{DS}-1000: A Natural and Reliable Benchmark for Data Science Code Generation},
  author =       {Lai, Yuhang and Li, Chengxi and Wang, Yiming and Zhang, Tianyi and Zhong, Ruiqi and Zettlemoyer, Luke and Yih, Wen-Tau and Fried, Daniel and Wang, Sida and Yu, Tao},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {18319--18345},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/lai23b/lai23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/lai23b.html},
  abstract = 	 {We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/