DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
Yuhang LaiChengxi LiYiming WangTianyi ZhangRuiqi ZhongLuke ZettlemoyerWen-Tau YihDaniel FriedSida I. WangTao Yu
Introduces DS-1000, an execution-based benchmark of 1,000 realistic data science problems across seven Python libraries that prevents memorization through problem perturbations and evaluates code generation with multi-criteria correctness constraints.
Data science programming presents high entry barriers for non-specialists due to the complexity of specialized software libraries. While artificial intelligence code generation models show promise in lowering these barriers, existing evaluation benchmarks primarily target competition-style programming puzzles or rely on superficial text-matching metrics. Consequently, they fail to represent real-world data science workflows or reliably verify whether generated code actually executes correctly. The article introduces and evaluates DS-1000, a benchmark designed to assess pre-trained code models on realistic, diverse data science problems using rigorous execution-based testing while defending against data memorization.
To build DS-1000, the authors curated and modified 1,000 realistic programming problems derived from 451 unique Stack Overflow queries spanning seven major Python data science libraries: NumPy, Pandas, Matplotlib, Scikit-learn, SciPy, TensorFlow, and PyTorch. Expert annotators rewrote the problems into executable contexts and implemented a multi-criteria evaluation framework. This framework executes model-generated code against an average of 1.6 functional test cases per problem and enforces surface-form constraints (such as forbidding inefficient loops or requiring specific library functions). Furthermore, to prevent models from succeeding via rote memorization of public internet data, the authors applied surface perturbations, semantic modifications, and difficult rewrites to more than half of the benchmark problems.
The benchmark analysis yielded several key findings regarding model capabilities and evaluation accuracy. First, current leading code generation models exhibit substantial room for improvement: the top-performing system, Codex-002 using code insertion, achieved only 43.3% accuracy, while smaller 6-billion-parameter open-source models scored under 10%. Second, the multi-criteria execution framework proved highly dependable, exhibiting a low sample-level false discovery rate of only 1.8% on Codex-002 predictions. Third, infilling or insertion-based prompting—which provides both preceding and following code context—improved Codex-002's accuracy by 4.1% over standard left-to-right generation. Finally, problem perturbations confirmed that models rely heavily on training set memorization; on a popular external NumPy problem set, perturbing problem wording caused Codex-002 accuracy to plummet from 72.5% to 23.6%, demonstrating the necessity of perturbed benchmarks.
These findings indicate that while modern artificial intelligence can assist with everyday data analysis, enterprise deployment still carries significant operational and reliability risks if outputs are unverified. The wide performance disparity across libraries and problem formulations shows that model competence does not automatically transfer across different tools. Additionally, standard benchmarks that do not alter public training questions risk severely overstating model capabilities, which could lead organizations to misjudge deployment readiness and safety.
Organizations developing or deploying code generation assistants should adopt multi-criteria execution testing that pairs functional verification with structural constraints. System designers should prioritize infilling capabilities over strict left-to-right code generation. For future research, the community should expand evaluation suites into complex multi-file workflows and develop automated methods to assist experts in writing robust verification test suites.
Confidence in these findings is high regarding the evaluated Python libraries and the documented baseline models. However, limitations remain: the benchmark does not cover multimodal figure interpretation for visualization tasks, omits untestable questions involving software installation or conceptual explanations, and reflects a snapshot of specific library versions pinned to Python 3.7.10.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This seminal work establishes the Codex model, the HumanEval evaluation framework, and pass@k metrics upon which DS-1000 builds to evaluate data science code generation.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Introduces Mostly Basic Programming Problems (MBPP) and unit-test execution methodology for Python program synthesis, providing key conceptual foundations for executing and verifying generated code.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). Establishes a standardized multi-task evaluation suite and foundational baselines for code understanding and generation across diverse programming settings.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). Provides fundamental techniques and benchmarks for multi-turn and single-turn program synthesis using open-source large language models.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). Introduces identifier-aware and unified sequence-to-sequence pre-training objectives for code generation models evaluated on downstream programming tasks.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). Demonstrates test filtering and execution-based evaluation strategies to mitigate false positives in large-scale code generation benchmarks.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). Extends library-oriented code generation benchmarking by expanding from data science libraries to over 139 external Python packages with complex, compositional instructions.
- Paper: Magicoder: Empowering Code Generation with OSS-Instruct, Yuxiang Wei et al. (2024). Applies open-source instruction tuning to enhance code generation capabilities, specifically evaluating models on realistic Python and data science tasks like DS-1000.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). Develops dedicated open code models trained at scale and evaluates them across specialized benchmarks, including DS-1000 for data science code generation.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). Advances the goal of un-memorized, leak-free evaluation by continuously curating fresh, time-stamped contest problems to prevent benchmark contamination.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Deepens the rigor of execution-based code evaluation by applying automated test mutation and edge-case generation to expose flaws in generated Python programs.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). Builds an open-access training pipeline for code LLMs and tests their real-world code generation capabilities across competitive Python benchmarks.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). Expands code evaluation beyond isolated, single-context snippets into multi-file repository dependencies across diverse programming languages.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). Pushes practical code generation evaluation from function-level synthesis to end-to-end repository issue resolution and multi-file patching.
