Built independently by an author, for readers. Read the story and support ChapterPal

keyword

CodeScope benchmark

The CodeScope benchmark is an execution-based evaluation framework designed to assess the code understanding and generation capabilities of large language models across diverse programming scenarios. Rather than relying solely on surface-level textual comparison against reference solutions, it evaluates models by executing generated code to verify functional correctness, consistency, and efficiency. The benchmark features broad multilingual coverage across dozens of programming languages and multiple software engineering tasks, measuring model performance across dimensions such as task difficulty, code length, and execution performance using automated multilingual execution environments.

1 item

CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

Weixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li, Qian Chen, Wen Wang, Tingyu Lin, Weishan Zhao, Li Zhu, Hari Sundaram, Shuiguang Deng

OrganizationsAlibaba GroupTU WienUniversity of California, Santa BarbaraUniversity of Chinese Academy of SciencesUniversity of Illinois Urbana-ChampaignXi'an Jiaotong UniversityZhejiang University

Why you should read this

Introduces CodeScope, an execution-based evaluation benchmark spanning 43 programming languages and eight tasks to measure large language model coding performance across length, difficulty, and runtime efficiency.

Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation. However, existing benchmarks for evaluating the code understanding and generation capacities of LLMs suffer from severe limitations. First, most benchmarks are insufficient as they focus on a narrow range of popular programming languages and specific tasks, whereas real-world software development scenarios show a critical need to implement systems with multilingual and multitask programming environments to satisfy diverse requirements. Second, most benchmarks fail to consider the actual executability and the consistency of execution results of the generated code. To bridge these gaps between existing benchmarks and expectations from practical applications, we introduce CodeScope, an execution-based, multilingual, multitask, multidimensional evaluation benchmark for comprehensively measuring LLM capabilities on coding tasks. CodeScope covers 43 programming languages and eight coding tasks. It evaluates the coding performance of LLMs from three dimensions (perspectives): length, difficulty, and efficiency. To facilitate execution-based evaluations of code generation, we develop MultiCodeEngine, an automated code execution engine that supports 14 programming languages. Finally, we systematically evaluate and analyze eight mainstream LLMs and demonstrate the superior breadth and challenges of CodeScope for evaluating LLMs on code understanding and generation tasks compared to other benchmarks. The CodeScope benchmark and code are publicly available at https://github.com/WeixiangYAN/CodeScope.

Added

2026-10-03