GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
Diganta MisraNizar IslahVictor MayBrice RaubyZihan WangJustine GehringAntonio OrvietoMuawiz ChaudharyEilif MullerIrina Rish
Presents GitChameleon 2.0, an execution-based benchmark of 328 Python problems that evaluates code generation across library updates, exposing that top language models fail nearly half the time on version-specific code.
Modern software development increasingly relies on artificial intelligence systems to generate code. However, real-world development environments frequently operate under fixed version constraints, legacy codebases, and technical debt. When software libraries introduce breaking changes across versions, AI models must be capable of generating code that strictly satisfies the specific version requested rather than defaulting to the newest syntax. Prior benchmarks have primarily evaluated forward code migration on unseen libraries or relied on static text matching, leaving a critical gap in testing whether models can generate functionally correct code for specific, previously seen library versions.
The article introduces and evaluates GitChameleon 2.0, a benchmark designed to assess the version-conditioned code generation capabilities of AI models. It evaluates whether contemporary large language models, autonomous multi-step agents, developer coding assistants, and retrieval-augmented generation systems can produce functionally accurate, version-compliant code verified through automated test execution.
To conduct this evaluation, the researchers curated a benchmark consisting of 328 Python code completion tasks across 26 popular libraries spanning data science, scientific computing, and web development. Each task is built around documented historical breaking changes released between 2014 and 2023, specifically focusing on versions within the training data of modern models. The evaluation methodology uses execution-based validation inside dedicated container environments, testing model outputs against both visible unit tests for debugging and comprehensive hidden test suites to measure functional success across diverse prompt paradigms, agent tools, and commercial assistant products.
The findings reveal that current state-of-the-art AI systems face substantial difficulty with version-conditioned code generation. Under baseline greedy decoding, top enterprise models (such as GPT-4o, Claude 3.7 Sonnet, Gemini 2.5 Pro, and o1) achieve success rates of only 48% to 51% on hidden execution tests, while open-weights models generally lag behind at 30% to 48%. Providing automated feedback via visible error traces (self-debugging) significantly improves performance, lifting hidden success rates by 10 to 20 percentage points across models. In contrast, reasoning enhancements like zero-shot Chain-of-Thought prompting yield inconsistent results and even degrade performance on several top models. Supplying version-specific documentation through retrieval-augmented generation boosts success rates by up to 10 percentage points, yet the strongest model configuration still fails on more than 40% of the tasks. Additionally, performance varies significantly by change type: models succeed most frequently on semantic changes (60% to 80% with self-debugging) but struggle most with newly introduced features requiring alternative dependencies (25% to 50% baseline).
These results demonstrate that AI models suffer from significant ambiguity when tasked with version control and cannot reliably disambiguate between multiple library versions present in their training data. For organizations deploying AI coding assistants in production, relying solely on standard generation introduces substantial operational and safety risks, including silent runtime failures, deprecation bugs, and broken builds in legacy environments. Contrary to common assumptions, simply using larger reasoning models or standard prompt engineering is insufficient to guarantee version compliance.
Organizations and developers should adopt robust mitigation strategies rather than relying on out-of-the-box model generation for constrained software stacks. Teams should integrate automated execution sandboxes and closed-loop self-debugging pipelines that supply compiler or runtime test feedback directly back to the model, as this intervention yielded the most reliable performance gains. Incorporating version-pinned reference documentation into retrieval pipelines is also recommended to ground model outputs. Furthermore, AI tooling developers must prioritize version-aware architectures and specialized dependency agents before deploying fully autonomous coding tools into production codebases.
The study's conclusions are constrained by its specific boundary conditions: the benchmark is limited to the Python ecosystem, covers 26 libraries, and evaluates code completion rather than direct version-to-version code translation or human programmer baselines. Nevertheless, the high degree of inter-model agreement and the rigorous execution-based testing framework provide strong confidence that current AI models have fundamental limitations in version-specific code synthesis.
- Paper: DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation, Yuhang Lai et al. (2023). It introduces execution-based unit test benchmarking for complex Python libraries, providing the foundational evaluation paradigm that GitChameleon 2.0 builds upon for version-specific contexts.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). It establishes execution-based software engineering benchmarks on real GitHub repositories, directly informing GitChameleon 2.0's execution framework across software evolution.
- Paper: Gorilla: Large Language Model Connected with Massive APIs, Shishir G. Patil et al. (2023). It formulates the problem of API evolution and documentation retrieval for avoiding hallucinations in AI code generation, setting up the challenge of library version incompatibilities.
- Paper: RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation, Fengji Zhang et al. (2023). It details repository-level retrieval-augmented code completion, establishing key techniques for conditioning LLMs on external library contexts and codebases.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). It demonstrates the necessity of rigorous execution-based testing to uncover subtle errors in LLM-generated code that standard text matching overlooks.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). It demonstrates how code completion models struggle when required to retrieve and reason over external cross-file and project dependencies.
- Paper: Magicoder: Empowering Code Generation with OSS-Instruct, Yuxiang Wei et al. (2024). It provides foundational insights into instruction tuning code models on open-source codebases to handle diverse and realistic coding tasks.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). It expands evaluation beyond single-library versioning to complex multi-library function calls and diverse instructions across 139 external packages.
- Paper: TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark, Kush Jain et al. (2025). It extends execution-grounded code generation by evaluating whether models can generate full test suites to catch bugs and regressions in evolving Python repositories.
- Paper: Training Software Engineering Agents and Verifiers with SWE-Gym, Jiayi Pan et al. (2025). It uses containerized, execution-grounded software environments to train and scale open coding agents that navigate real codebase dependencies and tests.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It extends the evaluation of complex software workflows by using interactive agentic evaluators capable of inspecting multi-file environments and execution logs.
- Paper: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?, Samuel Miserendino et al. (2025). It scales the assessment of AI code generation from library-conditioned completion tasks to end-to-end commercial software engineering tasks with economic stakes.
