CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
Weixiang YanHaitian LiuYunkun WangYunzhe LiQian ChenWen WangTingyu LinWeishan ZhaoLi ZhuHari Sundaram
Introduces CodeScope, an execution-based evaluation benchmark spanning 43 programming languages and eight tasks to measure large language model coding performance across length, difficulty, and runtime efficiency.
Software engineering increasingly relies on large language models to automate software development, maintenance, and testing. However, existing benchmarks provide an incomplete and potentially misleading view of model capabilities because they focus heavily on Python, simple synthetic problems, and superficial text-matching metrics rather than actual code execution. In practical production environments, software systems require multi-language support, complex problem-solving across diverse programming paradigms, and functionally correct execution. The article introduces CodeScope to resolve these shortcomings by establishing a comprehensive, execution-based evaluation framework across multiple tasks, programming languages, and operational dimensions.
The article systematically assesses the code understanding and code generation capabilities of eight mainstream large language models using real-world scenarios. To perform reliable execution-based scoring, the authors developed MultiCodeEngine, an execution environment supporting 47 compiler and interpreter versions across 14 programming languages. The evaluation spans 43 programming languages and eight distinct tasks, comprising four code understanding tasks (summarization, smell detection, review, and test generation) and four code generation tasks (synthesis, translation, repair, and optimization). Models are evaluated across three core dimensions: code length, task difficulty, and resource efficiency.
The findings reveal that current models struggle significantly with complex and realistic coding demands. First, while proprietary frontier models such as GPT-4 and GPT-3.5 lead in code generation, all evaluated models experience sharp performance degradation as problem difficulty increases; for example, GPT-4 achieves a 58.57% pass rate on easy synthesis problems but drops to 10.99% on hard problems. Second, open-source and specialized models lag considerably behind closed models in generation, with most failing entirely on complex synthesis tasks. Third, top performance in code generation does not imply superior code comprehension. Specialized models like WizardCoder lead the understanding benchmarks, whereas GPT-4 ranks fifth due to difficulties in generating valid automated test cases that match runtime execution flows. Finally, in efficiency optimization, models achieve modest success in high-level languages like Python but struggle with low-level languages like C, with most optimizations limited to surface-level syntactic tweaks rather than algorithmic improvements.
These results demonstrate that single-metric and single-language benchmarks significantly overestimate the readiness of language models for autonomous software engineering. Relying on current models for complex software development carries significant risk of runtime bugs, compilation errors, and poor resource utilization. Organizations adopting coding assistants should implement strict verification pipelines and avoid deploying model-generated code directly to production without automated test execution and human review. Furthermore, technical leaders should choose models based on specific task profiles—such as utilizing specialized models for code comprehension and review, while reserving frontier models for generation.
Moving forward, the article recommends pursuing two technical development paths: directly improving base model capabilities to solve harder algorithmic problems and exploring autonomous multi-agent systems to divide complex development tasks into manageable units. The primary limitation noted by the authors is potential pre-training data contamination, an inherent challenge across modern foundation models. However, because CodeScope relies on multi-source datasets, diverse downstream tasks, and strict runtime execution, stakeholders can maintain high confidence in the benchmark’s comparative findings.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). CodeScope’s broad evaluation builds on the benchmark tradition CodeXGLUE established for assessing code understanding and generation across varied tasks and languages.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). EvalPlus shows how expanded execution tests can expose inflated code-generation scores, clarifying the need for CodeScope’s runtime-based evaluation.
- Paper: DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation, Yuhang Lai et al. (2023). DS-1000 provides an earlier model of realistic, execution-verified code-generation evaluation that helps frame CodeScope’s move beyond simple text-matching benchmarks.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). LiveCodeBench extends broad code evaluation with continuously refreshed problems and contamination controls, carrying forward CodeScope’s concern that static benchmarks can overstate capability.
- Paper: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?, Samuel Miserendino et al. (2025). SWE-Lancer pushes execution-based evaluation beyond benchmark coding tasks to full-stack freelance work and real economic outcomes.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). BigCodeBench extends functional code-generation testing to realistic tasks involving complex instructions and diverse library calls.
- Paper: TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark, Kush Jain et al. (2025). TestGenEval develops CodeScope’s test-generation dimension into a focused, execution-based assessment of complete test suites on mature software projects.
