GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities

Diganta MisraNizar IslahVictor MayBrice RaubyZihan WangJustine GehringAntonio OrvietoMuawiz ChaudharyEilif MullerIrina Rish

article2025arXiv2 citations

Presents GitChameleon 2.0, an execution-based benchmark of 328 Python problems that evaluates code generation across library updates, exposing that top language models fail nearly half the time on version-specific code.

Listen

Modern software development increasingly relies on artificial intelligence systems to generate code. However, real-world development environments frequently operate under fixed version constraints, legacy codebases, and technical debt. When software libraries introduce breaking changes across versions, AI models must be capable of generating code that strictly satisfies the specific version requested rather than defaulting to the newest syntax. Prior benchmarks have primarily evaluated forward code migration on unseen libraries or relied on static text matching, leaving a critical gap in testing whether models can generate functionally correct code for specific, previously seen library versions.

The article introduces and evaluates GitChameleon 2.0, a benchmark designed to assess the version-conditioned code generation capabilities of AI models. It evaluates whether contemporary large language models, autonomous multi-step agents, developer coding assistants, and retrieval-augmented generation systems can produce functionally accurate, version-compliant code verified through automated test execution.

To conduct this evaluation, the researchers curated a benchmark consisting of 328 Python code completion tasks across 26 popular libraries spanning data science, scientific computing, and web development. Each task is built around documented historical breaking changes released between 2014 and 2023, specifically focusing on versions within the training data of modern models. The evaluation methodology uses execution-based validation inside dedicated container environments, testing model outputs against both visible unit tests for debugging and comprehensive hidden test suites to measure functional success across diverse prompt paradigms, agent tools, and commercial assistant products.

The findings reveal that current state-of-the-art AI systems face substantial difficulty with version-conditioned code generation. Under baseline greedy decoding, top enterprise models (such as GPT-4o, Claude 3.7 Sonnet, Gemini 2.5 Pro, and o1) achieve success rates of only 48% to 51% on hidden execution tests, while open-weights models generally lag behind at 30% to 48%. Providing automated feedback via visible error traces (self-debugging) significantly improves performance, lifting hidden success rates by 10 to 20 percentage points across models. In contrast, reasoning enhancements like zero-shot Chain-of-Thought prompting yield inconsistent results and even degrade performance on several top models. Supplying version-specific documentation through retrieval-augmented generation boosts success rates by up to 10 percentage points, yet the strongest model configuration still fails on more than 40% of the tasks. Additionally, performance varies significantly by change type: models succeed most frequently on semantic changes (60% to 80% with self-debugging) but struggle most with newly introduced features requiring alternative dependencies (25% to 50% baseline).

These results demonstrate that AI models suffer from significant ambiguity when tasked with version control and cannot reliably disambiguate between multiple library versions present in their training data. For organizations deploying AI coding assistants in production, relying solely on standard generation introduces substantial operational and safety risks, including silent runtime failures, deprecation bugs, and broken builds in legacy environments. Contrary to common assumptions, simply using larger reasoning models or standard prompt engineering is insufficient to guarantee version compliance.

Organizations and developers should adopt robust mitigation strategies rather than relying on out-of-the-box model generation for constrained software stacks. Teams should integrate automated execution sandboxes and closed-loop self-debugging pipelines that supply compiler or runtime test feedback directly back to the model, as this intervention yielded the most reliable performance gains. Incorporating version-pinned reference documentation into retrieval pipelines is also recommended to ground model outputs. Furthermore, AI tooling developers must prioritize version-aware architectures and specialized dependency agents before deploying fully autonomous coding tools into production codebases.

The study's conclusions are constrained by its specific boundary conditions: the benchmark is limited to the Python ecosystem, covers 26 libraries, and evaluates code completion rather than direct version-to-version code translation or human programmer baselines. Nevertheless, the high degree of inter-model agreement and the rigorous execution-based testing framework provide strong confidence that current AI models have fundamental limitations in version-specific code synthesis.

Cover for GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities

Abstract

The rapid evolution of software libraries poses a considerable hurdle for code generation, necessitating continuous adaptation to frequent version updates while preserving backward compatibility. While existing code evolution benchmarks provide valuable insights, they typically lack execution-based evaluation for generating code compliant with specific library versions. To address this, we introduce GitChameleon 2.0, a novel, meticulously curated dataset comprising 328 Python code completion problems, each conditioned on specific library versions and accompanied by executable unit tests. GitChameleon 2.0 rigorously evaluates the capacity of contemporary large language models (LLMs), LLM-powered agents, code assistants, and RAG systems to perform version-conditioned code generation that demonstrates functional accuracy through execution. Our extensive evaluations indicate that state-of-the-art systems encounter significant challenges with this task; enterprise models achieving baseline success rates in the 48-51% range, underscoring the intricacy of the problem. By offering an execution-based benchmark emphasizing the dynamic nature of code libraries, GitChameleon 2.0 enables a clearer understanding of this challenge and helps guide the development of more adaptable and dependable AI code generation methods. We make the dataset and evaluation code publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 GitChameleon 2.0 Benchmark
  • 2.1 Dataset Structure
  • 2.2 Evaluation Metrics
  • 2.3 Statistics
  • 3 Empirical Study
  • 3.1 Experimental Setup
  • 3.1.1 Greedy Decoding
  • 3.1.2 Zero-Shot Chain-Of-Thought (CoT)
  • 3.1.3 Self-Debugging
  • 3.1.4 Retrieval-Augmented Generation
  • 3.1.5 Multi-Step Agent
  • 3.1.6 AI Coding Assistants
  • 3.2 Experiment Results
  • 3.2.1 Greedy Decoding
  • 3.2.2 Zero-Shot Chain-Of-Thought
  • 3.2.3 LLM Self-Debugging
  • 3.2.4 Multi-Step Agent
  • 3.2.5 AI Coding Assistants
  • 3.2.6 Retrieval-Augmented Generation
  • 3.3 In-Depth Analysis of Findings
  • 4 Related Work
  • 5 Conclusion
  • References
  • A Benchmark Details
  • A.1 Dataset Construction Process
  • A.2 Structure of Dataset Samples
  • A.3 Dataset Validation
  • A.4 Hidden Test Construction
  • A.5 Additional Dataset Statistics
  • B Extra Methodologies: Reasoning, Sampling and Prompting
  • C Extended Experiment Results and Analysis
  • D Related Work
  • D.1 Code Evolution Datasets
  • D.1.1 Task Format: Instruction-Based Generation
  • D.1.2 Task Format: Code Update, Repair, and Completion
  • D.2 Specialized Frameworks and Repair Techniques
  • D.2.1 DepsRAG
  • D.2.2 Dr.Fix
  • D.2.3 ReplaceAPI / InsertPrompt
  • D.2.4 Conclusion
  • E Case Study: Code Assistant Failure With Search
  • E.1 Inputs
  • E.2 Model Attempt and Failure
  • F Case Study: Self-Debugging in Batched Matrix Exponential Computation
  • F.1 Inputs
  • F.2 First Model Attempt and Failure
  • F.3 Self-Debugging Process and Correction
  • F.4 Analysis of the Correction
  • G Qualitative Analysis
  • G.1 Greedy Decoding
  • G.1.1 Example 1: (PyTorch)
  • G.1.2 Greedy Example 2 (SciPy)
  • G.1.3 Greedy Example 3 (SymPy)
  • G.1.4 Greedy Example 4 (Flask)
  • G.2 Zero-Shot Chain-Of-Thought
  • G.2.1 CoT Example 1 (Torch)
  • G.2.2 CoT Example 2 (Scikit-learn)
  • G.2.3 CoT Example 3 (Falcon)
  • H Logic vs. Knowledge Retention
  • I Prompt Templates
  • J Artifacts and Model Details
  • J.1 Libraries
  • J.2 Models
  • J.3 Coding Assistants (CLI/IDE)

Knowls

  1. Knowl 1 — GitChameleon 2.0 Benchmark Specification and Dataset Construction

    model/method

    GitChameleon 2.0 is an execution-based benchmark containing 328 Python code completion problems designed to evaluate whether code generation systems can generate functionally correct code constrained to specific library versions. The dataset spans 26 popular Python libraries across scientific computing, data science, and web development (including PyTorch, SciPy, NumPy, Pandas, Scikit-Learn, SymPy, Flask, Django, and Falcon), covering library releases from 2014 to 2023 (predominantly 2021–2023) and excluding legacy or yanked versions.

    Each dataset sample consists of:

    1. Library and Library Version: The exact library and target version under test.
    2. Task Description: A natural language specification centered on a breaking API change.
    3. Starter Code: Initial Python code signatures and context.
    4. Extra Dependencies: Any additional required Python packages.
    5. Hidden Test Suite: A comprehensive suite of unit tests achieving 96.5% code coverage, used for execution-based ranking and success rate computation.
    6. Visible Test: A concise unit test used to provide runtime error traces for iterative self-debugging.
    7. Reference Solution: A verified, ground-truth implementation.
    8. Reference Documents: Version-specific official documentation segments used for Retrieval-Augmented Generation (RAG) experiments.

    Validation utilizes a dual-control mechanism: the target library version is explicitly specified in the natural language prompt, and execution occurs within an isolated Docker container configured with the exact target library environment.

  2. Knowl 2 — Version-Conditioned Generation (VCG) Paradigm

    definition

    Version-Conditioned Generation (VCG) is an evaluation paradigm for code generation models that assesses the model's ability to produce code adhering to a specific, static library version constraint vv that falls within the model's pre-training distribution (v∈VIDv \in \mathcal{V}_{\text{ID}}).

    VCG is distinct from Code Evolution (or code migration): while Code Evolution tasks require translating code forward to newer, out-of-distribution (OOD) library versions (v∈VOODv \in \mathcal{V}_{\text{OOD}}) or adapting to synthetic API updates, VCG poses in-distribution problems to evaluate control and disambiguation—determining whether an LLM exposed during training to multiple historical versions of an API can reliably constrain its generation to the specific syntax, argument structure, and runtime semantics of a requested target version.

  3. Knowl 3 — Baseline Performance of LLMs under Greedy, Self-Debug, and Zero-Shot CoT Regimes

    data/table

    Under greedy decoding (T=0,top_p=0.95T = 0, \text{top\_p} = 0.95), state-of-the-art enterprise models achieve hidden test success rates between 48% and 51%. Smaller models within the same model families consistently trail their full-sized counterparts by 4 to 15 percentage points. Zero-shot Chain-of-Thought (CoT) prompting does not consistently improve performance and causes noticeable degradation on several reasoning and enterprise models (e.g., OpenAI o1 drops from 51.2% to 41.2%, Gemini 2.0 Flash drops from 44.2% to 36.0%). In contrast, Self-Debugging provides substantial gains across all models on both visible and hidden test suites.

    Model Greedy Decoding Greedy with Self-Debug Zero-shot CoT
    Hidden (%) API Hit (%) Hidden (%) API Hit (%) Hidden (%) API Hit (%)
    Open-Weights Models
    Llama 3.1 Instruct Turbo 30.2±2.530.2 \pm 2.5 39.7±2.739.7 \pm 2.7 52.1±2.852.1 \pm 2.8 41.5±2.741.5 \pm 2.7 36.6±2.736.6 \pm 2.7 35.3±2.635.3 \pm 2.6
    Llama 3.3 Instruct Turbo 70B 36.3±2.736.3 \pm 2.7 36.4±2.736.4 \pm 2.7 53.0±2.853.0 \pm 2.8 37.4±2.737.4 \pm 2.7 37.5±2.737.5 \pm 2.7 37.2±2.737.2 \pm 2.7
    Llama 4 Maverick 400B 40.8±2.740.8 \pm 2.7 49.5±2.849.5 \pm 2.8 58.5±2.758.5 \pm 2.7 46.8±2.846.8 \pm 2.8 46.6±2.846.6 \pm 2.8 41.3±2.741.3 \pm 2.7
    Qwen 2.5-VL Instruct 72B 48.2±2.848.2 \pm 2.8 43.8±2.743.8 \pm 2.7 64.6±2.664.6 \pm 2.6 45.3±2.745.3 \pm 2.7 45.1±2.745.1 \pm 2.7 43.0±2.743.0 \pm 2.7
    Enterprise Models
    Claude 3.7 Sonnet 48.8±2.848.8 \pm 2.8 46.0±2.846.0 \pm 2.8 65.9±2.665.9 \pm 2.6 47.6±2.847.6 \pm 2.8 45.1±2.745.1 \pm 2.7 43.4±2.743.4 \pm 2.7
    Gemini 1.5 Pro 45.1±2.745.1 \pm 2.7 46.8±2.746.8 \pm 2.7 62.5±2.862.5 \pm 2.8 48.6±2.748.6 \pm 2.7 43.3±2.743.3 \pm 2.7 44.6±2.844.6 \pm 2.8
    Gemini 2.0 Flash 44.2±2.744.2 \pm 2.7 43.8±2.743.8 \pm 2.7 70.4±2.770.4 \pm 2.7 49.4±2.749.4 \pm 2.7 36.0±2.636.0 \pm 2.6 41.8±2.741.8 \pm 2.7
    Gemini 2.5 Pro 50.0±2.850.0 \pm 2.8 47.7±2.747.7 \pm 2.7 61.3±2.861.3 \pm 2.8 49.2±2.749.2 \pm 2.7 49.4±2.849.4 \pm 2.8 49.1±2.849.1 \pm 2.8
    Gemini 2.5 Flash 38.1±2.638.1 \pm 2.6 45.4±2.745.4 \pm 2.7 65.9±2.865.9 \pm 2.8 45.8±2.745.8 \pm 2.7 30.8±2.530.8 \pm 2.5 49.8±2.849.8 \pm 2.8
    GPT-4.1 48.5±2.848.5 \pm 2.8 46.8±2.746.8 \pm 2.7 63.4±2.863.4 \pm 2.8 48.3±2.748.3 \pm 2.7 47.9±2.847.9 \pm 2.8 44.5±2.744.5 \pm 2.7
    GPT-4.1-mini 44.2±2.744.2 \pm 2.7 44.5±2.744.5 \pm 2.7 68.0±2.868.0 \pm 2.8 46.3±2.746.3 \pm 2.7 24.1±1.824.1 \pm 1.8 41.3±2.741.3 \pm 2.7
    GPT-4.1-nano 33.8±2.633.8 \pm 2.6 43.1±2.743.1 \pm 2.7 67.7±2.767.7 \pm 2.7 45.8±2.745.8 \pm 2.7 11.9±1.811.9 \pm 1.8 32.1±2.532.1 \pm 2.5
    GPT-4o 49.1±2.849.1 \pm 2.8 46.5±2.746.5 \pm 2.7 64.9±2.864.9 \pm 2.8 48.0±2.748.0 \pm 2.7 50.3±2.850.3 \pm 2.8 42.5±2.742.5 \pm 2.7
    GPT-4o-mini 37.2±2.637.2 \pm 2.6 38.4±2.638.4 \pm 2.6 60.4±2.760.4 \pm 2.7 40.6±2.740.6 \pm 2.7 36.0±2.636.0 \pm 2.6 37.3±2.637.3 \pm 2.6
    GPT-4.5 40.8±2.740.8 \pm 2.7 52.8±2.852.8 \pm 2.8 66.2±2.866.2 \pm 2.8 54.4±2.754.4 \pm 2.7 39.9±2.639.9 \pm 2.6 48.8±2.848.8 \pm 2.8
    Grok 3 48.2±2.848.2 \pm 2.8 44.8±2.744.8 \pm 2.7 67.1±2.867.1 \pm 2.8 46.3±2.846.3 \pm 2.8 49.4±2.849.4 \pm 2.8 44.2±2.744.2 \pm 2.7
    Mistral Medium 3 43.6±2.743.6 \pm 2.7 44.2±2.744.2 \pm 2.7 61.3±2.861.3 \pm 2.8 45.4±2.745.4 \pm 2.7 44.2±2.744.2 \pm 2.7 44.1±2.744.1 \pm 2.7
    o1 51.2±2.851.2 \pm 2.8 42.1±2.742.1 \pm 2.7 57.6±2.757.6 \pm 2.7 49.2±2.849.2 \pm 2.8 41.2±2.741.2 \pm 2.7 41.3±2.741.3 \pm 2.7
  4. Knowl 4 — Multi-Step Agent Performance with Grounding Search and Sandbox Execution

    data/table

    Evaluating multi-step tool-calling agents (implemented using the smolagents framework alternating between planning and acting) with grounding tools (DuckDuckGo Search, Perplexity, and Grounded Gemini) shows that providing a code execution sandbox tool dramatically increases success rates across all agent backbones and search methods.

    Model Grounding Method Success Rate (%) API Hit Rate (%)
    No Sandbox Sandbox No Sandbox Sandbox
    Claude Sonnet 3.5 DuckDuckGo 41.7±2.741.7 \pm 2.7 55.3±2.7\mathbf{55.3 \pm 2.7} 42.2±2.742.2 \pm 2.7 48.9±2.848.9 \pm 2.8
    Perplexity 44.1±2.744.1 \pm 2.7 51.4±2.851.4 \pm 2.8 41.8±2.741.8 \pm 2.7 46.0±2.846.0 \pm 2.8
    Grounded Gemini 40.0±2.740.0 \pm 2.7 53.7±2.853.7 \pm 2.8 41.0±2.741.0 \pm 2.7 45.2±2.745.2 \pm 2.7
    Gemini 1.5 Pro DuckDuckGo 46.0±2.846.0 \pm 2.8 49.8±2.849.8 \pm 2.8 47.4±2.847.4 \pm 2.8 50.3±2.850.3 \pm 2.8
    Perplexity 46.5±2.8\mathbf{46.5 \pm 2.8} 44.4±2.744.4 \pm 2.7 47.2±2.847.2 \pm 2.8 46.6±2.846.6 \pm 2.8
    Grounded Gemini 44.1±2.744.1 \pm 2.7 49.2±2.849.2 \pm 2.8 49.7±2.8\mathbf{49.7 \pm 2.8} 51.2±2.8\mathbf{51.2 \pm 2.8}
    GPT-4o DuckDuckGo 23.9±2.423.9 \pm 2.4 33.2±2.633.2 \pm 2.6 44.2±2.744.2 \pm 2.7 48.1±2.848.1 \pm 2.8
    Perplexity 33.5±2.633.5 \pm 2.6 41.5±2.741.5 \pm 2.7 43.2±2.743.2 \pm 2.7 44.7±2.744.7 \pm 2.7
    Grounded Gemini 25.4±2.425.4 \pm 2.4 50.0±2.850.0 \pm 2.8 46.5±2.846.5 \pm 2.8 44.2±2.744.2 \pm 2.7

    Claude Sonnet 3.5 paired with DuckDuckGo search and sandbox execution achieves the highest overall success rate among all agent configurations (55.3±2.7%55.3 \pm 2.7\%), while Gemini 1.5 Pro performs best in the absence of a sandbox.

  5. Knowl 5 — Retrieval-Augmented Generation for Version-Conditioned Synthesis

    data/table

    A Retrieval-Augmented Generation (RAG) pipeline querying a vectorized database (VectorDB built with OpenAI text-embedding-3-large across a 536-document corpus of version-specific documentation) via DocPrompting provides an improvement of up to 10 percentage points over standard greedy decoding. Retrieving k=3k=3 documents consistently outperforms k=1k=1, despite the introduction of false-positive document matches.

    Model Success Rate (%) API Hit Rate (%) Precision (%) Recall (%) MRR
    Open-Weights Models
    DeepSeek V3 48.9±2.8\mathbf{48.9 \pm 2.8} 48.5±2.848.5 \pm 2.8 41.6±2.241.6 \pm 2.2 50.4±2.850.4 \pm 2.8 0.62±0.030.62 \pm 0.03
    Llama 4 Maverick 45.1±2.745.1 \pm 2.7 50.5±2.8\mathbf{50.5 \pm 2.8} 41.2±2.241.2 \pm 2.2 49.8±2.849.8 \pm 2.8 0.61±0.030.61 \pm 0.03
    Qwen 3 41.8±2.741.8 \pm 2.7 39.6±2.739.6 \pm 2.7 36.3±2.036.3 \pm 2.0 46.9±2.846.9 \pm 2.8 0.56±0.030.56 \pm 0.03
    Jamba 1.6 Large 41.8±2.741.8 \pm 2.7 47.1±2.847.1 \pm 2.8 41.9±2.241.9 \pm 2.2 50.7±2.850.7 \pm 2.8 0.62±0.030.62 \pm 0.03
    Enterprise Models
    Claude 3.7 Sonnet 56.1±2.756.1 \pm 2.7 53.0±2.853.0 \pm 2.8 41.9±2.241.9 \pm 2.2 50.7±2.850.7 \pm 2.8 0.62±0.030.62 \pm 0.03
    Claude 4 Sonnet 59.4±2.8\mathbf{59.4 \pm 2.8} 55.8±2.8\mathbf{55.8 \pm 2.8} 41.9±2.241.9 \pm 2.2 50.7±2.850.7 \pm 2.8 0.62±0.030.62 \pm 0.03
    Gemini 2.5 Pro 56.7±2.756.7 \pm 2.7 51.1±2.851.1 \pm 2.8 41.9±2.241.9 \pm 2.2 50.7±2.850.7 \pm 2.8 0.62±0.030.62 \pm 0.03
    GPT-4.1 58.5±2.758.5 \pm 2.7 51.8±2.851.8 \pm 2.8 41.2±2.241.2 \pm 2.2 50.1±2.850.1 \pm 2.8 0.61±0.030.61 \pm 0.03
    Grok 3 54.3±2.754.3 \pm 2.7 55.2±2.855.2 \pm 2.8 41.6±2.241.6 \pm 2.2 50.4±2.850.4 \pm 2.8 0.62±0.030.62 \pm 0.03
    Mistral Medium 3 52.4±2.752.4 \pm 2.7 51.2±2.851.2 \pm 2.8 41.6±2.241.6 \pm 2.2 50.4±2.850.4 \pm 2.8 0.62±0.030.62 \pm 0.03
    Devstral Small 43.3±2.743.3 \pm 2.7 45.1±2.845.1 \pm 2.8 41.6±2.241.6 \pm 2.2 50.4±2.850.4 \pm 2.8 0.62±0.030.62 \pm 0.03
    Nova Pro 44.2±2.744.2 \pm 2.7 42.4±2.742.4 \pm 2.7 40.7±2.240.7 \pm 2.2 49.6±2.849.6 \pm 2.8 0.60±0.030.60 \pm 0.03

    Despite documentation access bringing the top success rate to 59.4% (Claude 4 Sonnet) and 58.5% (GPT-4.1), over 40% of the benchmark's problems remain unsolved.

  6. Knowl 6 — Visible-Hidden Generalization Gap in LLM Self-Debugging

    empirical result

    When models are provided execution error feedback on visible unit tests, their success rate on those visible tests improves by 13 to 37 percentage points (e.g., Claude 3.7 Sonnet improves from 56% to 83%, Gemini 2.0 Flash from 50% to 75%, and GPT-4.1 from 49% to 69%). Hidden test success rates also improve by 10 to 20 percentage points (e.g., GPT-4.1-mini improves from 44.2% to 68.0%, Llama 3.1 Instruct Turbo from 30.2% to 52.1%).

    However, across all evaluated models, the performance gap between visible tests and hidden tests (∣Successvisible−Successhidden∣|\text{Success}_{\text{visible}} - \text{Success}_{\text{hidden}}|) systematically increases after self-debugging. This indicates that self-debugging against a single, concise visible test leads to partial overfitting to the visible feedback, offering limited transfer to the broader assertions in hidden tests.

  7. Knowl 7 — AI Coding Assistant Performance and Formatting Invariance

    empirical result

    Evaluating Command-Line Interface (CLI; e.g., Claude Code, Goose) and Integrated Development Environment (IDE; e.g., Cline, RooCode, KiloCode) assistants reveals two core findings:

    1. Natural Language Context Dependency: Supplying only starter code and file-header comments without the full natural language problem statement (simulating tab-completion) causes double-digit drops in success rate for 4 out of 5 assistant configurations evaluated (e.g., GPT-4.1 on Cline drops from 54.6% to 38.4%; Claude 3.7 Sonnet on Claude Code drops from 48.8% to 32.0%; GPT-4.1 on Goose drops from 55.5% to 19.2%).
    2. Invariance to Code Quality and Auto-Linting: Raw code generated by assistants exhibits poor static analysis quality (average Pylint scores of 0.00 to 1.06 out of 10). Applying automated code formatters and linters (Black, isort, Ruff) substantially raises Pylint scores (up to 2.60–2.92 out of 10) but causes zero change (0.0% difference) in the hidden test success rate. This confirms that model failures are caused by version-specific API misuse rather than formatting defects, indentation errors, or unused imports.
  8. Knowl 8 — Impact of API Evolution Change Categories on Generation Difficulty

    empirical result

    Problems in GitChameleon 2.0 are partitioned into four API evolution categories:

    1. Semantics or Function Behavior Changes: The runtime behavior or return type changed across versions. These are the most tractable for LLMs, achieving success rates of 55–65% under Greedy Decoding and 60–80% under Self-Debugging.
    2. Function Name Changes: Renamed API calls (e.g., pandas.DataFrame.append to pandas.concat).
    3. Argument or Attribute Changes: Alterations to argument names, orders, or defaults (e.g., bw deprecated in favor of bw_method and bw_adjust in seaborn.violinplot). This is the most common category of change in the dataset.
    4. New Feature or Dependency Changes: Features introduced in newer versions where achieving identical functionality in older versions requires alternative implementations or third-party libraries (e.g., torch.special requiring scipy.special in earlier PyTorch versions). This category is the most difficult, achieving success rates of 25–50% under Greedy Decoding and 50–65% under Self-Debugging.
  9. Knowl 9 — Logic vs. Knowledge Retention AST Analysis in GitChameleon 2.0

    theoretical result

    To demonstrate that GitChameleon 2.0 isolates version-specific API knowledge retention rather than algorithmic or logical reasoning, an Abstract Syntax Tree (AST) analysis classifies AST call nodes of ground-truth solutions as logic-related if they:

    • Call a user-defined function,
    • Call built-in Python operators (e.g., +, *),
    • Call a math or utility function with non-obvious purpose, or
    • Compose multiple calls together. Direct calls to external library methods (e.g., torch.from_numpy) are classified as non-logic nodes.

    The distribution of logic-related AST nodes across all ground-truth solutions shows that the vast majority contain fewer than 5 logic nodes. This structural characteristic confirms that the benchmark primarily tests knowledge retention of versioned API signatures and semantics rather than complex algorithmic synthesis.

  10. Knowl 10 — Correlation with Standard Code Generation Benchmarks

    empirical result

    GitChameleon 2.0 performance exhibits moderate correlation with repository-level GitHub issue benchmarks and weak correlation with competitive programming benchmarks:

    • SWE-Bench: Spearman rank correlation ρ=0.550\rho = 0.550, Pearson correlation r=0.675r = 0.675.
    • LiveCodeBench: Spearman rank correlation ρ=0.214\rho = 0.214, Pearson correlation r=0.130r = 0.130.

    These correlation levels demonstrate that version-conditioned code generation requires capabilities distinct from competitive algorithmic problem solving (tested in LiveCodeBench), while sharing moderate overlap with repository maintenance tasks (tested in SWE-Bench).

  11. Knowl 11 — Limitations of GitChameleon 2.0

    limitation

    The GitChameleon 2.0 benchmark has several stated limitations:

    1. Language and Library Scope: The dataset is restricted to Python and encompasses 26 libraries, omitting other widely used software ecosystems.
    2. Task Format: Evaluation is limited strictly to code generation from natural language problem specifications and starter code. It does not evaluate version-to-version code translation or codebase migration (e.g., translating a solution valid in PyTorch 1.8 to be valid in 1.7 or 1.9).
    3. Human Baselines: The study does not include human developer performance evaluations on the benchmark tasks to establish an empirical human baseline.

Coverage note — Specific prompt templates (system and user prompt text strings) and individual URL citations for the 26 libraries were omitted as they constitute raw artifacts rather than distinct findings.

References

  1. 1.Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. 2019. Gradio: Hassle-free sharing and testing of ML models in the wild. Preprint, arXiv:1906.02569.
  2. 2.Meta AI. 2025. Everything we announced at our first-ever LlamaCon. https://ai.meta.com/blog/llamacom-llama-news/. Discusses Llama 3.3 Instruct Turbo and Llama 4 Maverick.
  3. 3.Mohannad Alhanahnah, Yazan Boshmaf, and Benoit Baudry. 2024. DepsRAG: Towards managing software dependencies using large language models. arXiv preprint arXiv:2405.20455v2.
  4. 4.Anthropic. 2025. Claude 3.7 Sonnet and Claude Code. https://www.anthropic.com/news/claude-3-7-sonnet.
  5. 5.Arcee. Model Selection | Arcee AI Documentation — docs.arcee.ai. https://docs.arcee.ai/arcee-conductor/arcee-small-language-models/model-selection#caller-large-tool-use-and-function-call. [Accessed 15-07-2025].
  6. 6.Farnaz Behrang, Zhizhou Zhang, Georgian-Vlad Saioc, Peng Liu, and Milind Chabbi. 2025. Dr.fix: Automatically fixing data races at industry scale. Preprint, arXiv:2504.15637.
  7. 7.Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Andreas Joly, Bertrand Druillette, Gael Varoquaux, and Marion Gramfort. 2013. API design for machine learning software: experiences from the scikit-learn project. arXiv preprint arXiv:1309.0238.
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. Evaluating large language models trained on code. ArXiv.
  9. 9.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023. Teaching large language models to self-debug. Preprint, arXiv:2304.05128.
  10. 10.Keyuan Cheng, Xudong Shen, Yihao Yang, Tengyue Wang, Yang Cao, Muhammad Asif Ali, Hanbin Wang, Lijie Hu, and Di Wang. 2025. Codemenv: Benchmarking large language models on code migration. Preprint, arXiv:2506.00894.
  11. 11.Matteo Ciniselli, Alberto Martin-Lopez, and Gabriele Bavota. 2024. On the generalizability of deep learning-based code completion across programming language versions. Preprint, arXiv:2403.15149.
  12. 12.Google Cloud. 2025. Gemini 2.5 on Vertex AI: Pro, Flash & Model Optimizer Live. https://cloud.google.com/blog/products/ai-machine-learning/gemini-2-5-pro-flash-on-vertex-ai. Discusses Gemini 2.5 Pro and Gemini 2.5 Flash.
  13. 13.Team Cohere, :, Aakanksha, Arash Ahmadian, Marwan Ahmed, Jay Alammar, Milad Alizadeh, Yazeed Alnumay, Sophia Althammer, Arkady Arkhangorodsky, Viraat Aryabumi, Dennis Aumiller, Raphaël Avalos, Zahara Aviv, Sammie Bae, Saurabh Baji, Alexandre Barbet, Max Bartolo, Björn Bebensee, and 211 others. 2025. Command a: An enterprise-ready large language model. Preprint, arXiv:2504.00698.
  14. 14.Forbes Technology Council. 2024. Revolutionizing software development with large language models. https://www.forbes.com/councils/forbestechcouncil/2024/03/20/revolutionizing-software-development-with-large-language-models/.
  15. 15.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems (NeurIPS).
  16. 16.DeepSeek-AI. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948.
  17. 17.DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, and 181 others. 2025. Deepseek-v3 technical report. Preprint, arXiv:2412.19437.
  18. 18.DuckDuckGo. 2025. DuckDuckGo: Privacy, simplified. https://duckduckgo.com/.
  19. 19.Lishui Fan, Mouxiang Chen, and Zhongxin Liu. 2024. Self-explained keywords empower large language models for code generation. Preprint, arXiv:2410.15966.
  20. 20.Google. 2025. Grounding with Google Search | Gemini API. https://ai.google.dev/gemini-api/docs/grounding.
  21. 21.Aric A Hagberg, Daniel A Schult, and Pieter J Swart. 2008. Exploring network structure, dynamics, and function using NetworkX. In Proceedings of the 7th Python in Science Conference, pages 11–15.
  22. 22.Charles R Harris, K Jarrod Millman, Stéfan J van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, Robert Kern, Matti Picus, Changqing Hoyer, Marten H van Kerkwijk, Alex Brett, Andrew Wen, Pete Zhang, Joe Igoe, Keith Featherstone, and Travis E Oliphant. 2020. Array programming with NumPy. Nature, 585(7825):357–362.
  23. 23.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with apps. NeurIPS.
  24. 24.J. D. Hunter. 2007. Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9(3):90–95.
  25. 25.Amazon Artificial General Intelligence. 2024. The amazon nova family of models: Technical report and model card. Amazon Technical Reports.
  26. 26.Nizar Islah, Justine Gehring, Diganta Misra, Eilif Muller, Irina Rish, Terry Yue Zhuo, and Massimo Caccia. 2024. Gitchameleon: Unmasking the version-switching capabilities of code generation models. arXiv preprint arXiv:2411.05830.
  27. 27.Mohayeminul Islam, Ajay Kumar Jha, Sarah Nadi, and Ildar Akhmetov. 2023. Pymigbench: A benchmark for python library migration. In 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pages 511–515.
  28. 28.Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974.
  29. 29.Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations.
  30. 30.Haolin Jin, Zechao Sun, and Huaming Chen. 2024. Rgd: Multi-llm based agent debugger via refinement and generation guidance. Preprint, arXiv:2410.01242.
  31. 31.Kelsey Jordahl, Joris Van den Bossche, Martin Fleischmann, Jacob Wasserman, James McBride, Jeffrey Gerard, Jeff Tratner, Matthew Perry, Adrian Garcia Badaracco, Carson Farmer, Geir Arne Hjelle, Alan D. Snow, Micah Cochran, Sean Gillies, Lucas Culbertson, Matt Bartos, Nick Eubank, maxalbert, Aleksey Bilogur, and 11 others. 2020. geopandas/geopandas: v0.8.1.
  32. 32.Kat Kampf. 2025. Create and edit images with Gemini 2.0 in preview. https://developers.googleblog.com/en/generate-images-gemini-2-0-flash-preview/. Discusses Gemini 2.0 Flash.
  33. 33.Paul Kassianik, Baturay Saglam, Alexander Chen, Blaine Nelson, Anu Vellore, Massimo Aufiero, Fraser Burch, Dhruv Kedia, Avi Zohary, Sajana Weerawardhena, Aman Priyanshu, Adam Swanda, Amy Chang, Hyrum Anderson, Kojin Oshiba, Omar Santos, Yaron Singer, and Amin Karbasi. 2025. Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report. arXiv preprint arXiv:2504.21039. Cited for Llama 3.1 Instruct Turbo.
  34. 34.Sachit Kuhar, Wasi Uddin Ahmad, Zijian Wang, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma, and Anoop Deoras. 2024. Libevolutioneval: A benchmark and study for version-specific code generation. Preprint, arXiv:2412.04478.
  35. 35.Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, and Daniel Russo. 2025. Exploring individual factors in the adoption of llms for specific software engineering tasks. Preprint, arXiv:2504.02553.
  36. 36.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, volume 33.
  37. 37.James Li. 2024. ReAct vs Plan-and-Execute: A Practical Comparison of LLM Agent Patterns. https://dev.to/jamesli.
  38. 38.Linxi Liang, Jing Gong, Mingwei Liu, Chong Wang, Guangsheng Ou, Yanlin Wang, Xin Peng, and Zibin Zheng. 2025. Rustevo: An evolving benchmark for api evolution in llm-based rust code generation. Preprint, arXiv:2503.16922.
  39. 39.Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, and 3 others. 2024. Jamba: A hybrid transformer-mamba language model. Preprint, arXiv:2403.19887.
  40. 40.Yue Liu, Chakkrit Tantithamthavorn, Yonghui Liu, Patanamon Thongtanunam, and Li Li. 2024. Automatically recommend code updates: Are we there yet? Preprint, arXiv:2209.07048.
  41. 41.Zeyu Leo Liu, Shrey Pandit, Xi Ye, Eunsol Choi, and Greg Durrett. 2025. Codeupdatearena: Benchmarking knowledge editing on API updates.
  42. 42.Edward Loper and Steven Bird. 2002. NLTK: The natural language toolkit. In Proceedings of the ACL-02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics, pages 63–70, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  43. 43.Stephan Lukasczyk and Gordon Fraser. 2022. Pynguin: Automated unit test generation for python. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings, pages 168–172.
  44. 44.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. WizardCoder: Empowering code large language models with Evol-Instruct.
  45. 45.Wes McKinney. 2010. Data Structures for Statistical Computing in Python. In Proceedings of the 9th Python in Science Conference, pages 51–56.
  46. 46.Mistral AI. 2025. Medium is the new large: Introducing mistral medium 3. https://mistral.ai/news/mistral-medium-3. Accessed: 2025-05-17.
  47. 47.OpenAI. 2024a. GPT-4o System Card. arXiv preprint arXiv:2410.21276. Cited for GPT-4o.
  48. 48.OpenAI. 2024b. OpenAI o1 System Card. https://openai.com/index/openai-o1-system-card/. Discusses the o1 model series, including o1 and mentioning o3-mini.
  49. 49.OpenAI. 2025a. Introducing GPT-4.1 in the API. https://openai.com/index/gpt-4-1/. Discusses GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano.
  50. 50.OpenAI. 2025b. Introducing GPT-4.5. https://openai.com/index/introducing-gpt-4-5/.
  51. 51.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, and 2 others. 2019. PyTorch: An imperative style, high-performance deep learning library. Preprint, arXiv:1912.01703.
  52. 52.Perplexity AI. 2024. Getting started with Perplexity. https://www.perplexity.ai/hub/blog/getting-started-with-perplexity.
  53. 53.Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. Qwen2.5 technical report. Preprint, arXiv:2412.15115.
  54. 54.Rafiqul Rabin, Jesse Hostetler, Sean McGregor, Brett Weir, and Nick Judd. 2025. Sandboxeval: Towards securing test environment for untrusted code. Preprint, arXiv:2504.00018.
  55. 55.Reka. RekaAI/reka-flash-3 · Hugging Face — huggingface.co. https://huggingface.co/RekaAI/reka-flash-3. [Accessed 15-07-2025].
  56. 56.Aymeric Roucher, Albert Villanova del Moral, Thomas Wolf, Leandro von Werra, and Erik Kaunismäki. 2025. ‘smolagents‘: a smol library to build great agentic systems. https://github.com/huggingface/smolagents.
  57. 57.Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2025. Agentic retrieval-augmented generation: A survey on agentic rag. Preprint, arXiv:2501.09136.
  58. 58.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. Roformer: Enhanced transformer with rotary position embedding. Preprint, arXiv:2104.09864.
  59. 59.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, and 1118 others. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Preprint, arXiv:2403.05530.
  60. 60.The pandas development team. 2020. pandas-dev/pandas: Pandas.
  61. 61.Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, and 1 others. 2020. Scipy 1.0: fundamental algorithms for scientific computing in python. Nature methods, 17(3):261–272.
  62. 62.Chaozheng Wang, Shuzheng Gao, Cuiyun Gao, Wenxuan Wang, Chun Yong Chong, Shan Gao, and Michael R. Lyu. 2024a. A systematic evaluation of large code models in api suggestion: When, which, and how. Preprint, arXiv:2409.13178.
  63. 63.Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2024b. How and Why LLMs Use Deprecated APIs in Code Completion? an Empirical Study. arXiv preprint arXiv:2312.14617.
  64. 64.Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2025a. LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion . In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pages 781–781, Los Alamitos, CA, USA. IEEE Computer Society.
  65. 65.Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2025b. Llms meet library evolution: Evaluating deprecated api usage in llm-based code completion. Preprint, arXiv:2406.09834.
  66. 66.Xingyao Wang. 2025. Introducing openhands lm 32b – a strong, open coding agent model. All Hands AI Blog.
  67. 67.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models. Preprint, arXiv:2201.11903.
  68. 68.Tongtong Wu, Weigang Wu, Xingyu Wang, Kang Xu, Suyu Ma, Bo Jiang, Ping Yang, Zhenchang Xing, Yuan-Fang Li, and Gholamreza Haffari. 2024. VersiCode: Towards version-controllable code generation.
  69. 69.xAI. 2025. Grok-3. Official xAI announcement. Accessed May 17, 2025.
  70. 70.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388.
  71. 71.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
  72. 72.Sixiang Ye, Zeyu Sun, Guoqing Wang, Liwei Guo, Qingyuan Liang, Zheng Li, and Yong Liu. 2025. Prompt alchemy: Automatic prompt refinement for enhancing code generation. Preprint, arXiv:2503.11085.
  73. 73.Shuyan Zhou, Uri Alon, Frank F Xu, Zhiruo Wang, Zhengbao Jiang, and Graham Neubig. 2022. DocPrompting: Generating code by retrieving the docs.

Citation

MLA
Misra, D., et al. “GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities”. arXiv, 2025, http://arxiv.org/abs/2507.12367v2.
APA
Misra, D., Islah, N., May, V., Rauby, B., Wang, Z., Gehring, J., Orvieto, A., Chaudhary, M., Muller, E. B., Rish, I., Kahou, S. E., & Caccia, M. (2025). GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities. arXiv. http://arxiv.org/abs/2507.12367v2
Chicago
Misra, D., N. Islah, V. May, et al. 2025. “GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities”. arXiv. http://arxiv.org/abs/2507.12367v2.
Harvard
Misra, D. et al. (2025) “GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.12367v2.
Vancouver
1. Misra D, Islah N, May V, et al (2025) GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities. arXiv

BibTeX

@article{misra2025gitchameleon,
  title = {GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities},
  author = {Misra, Diganta and Islah, Nizar and May, Victor and Rauby, Brice and Wang, Zihan and Gehring, Justine and Orvieto, Antonio and Chaudhary, Muawiz and Muller, Eilif B. and Rish, Irina and Kahou, Samira Ebrahimi and Caccia, Massimo},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.12367v2},
  eprint = {2507.12367}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/