The CodeScope benchmark is an execution-based evaluation framework designed to assess the code understanding and generation capabilities of large language models across diverse programming scenarios. Rather than relying solely on surface-level textual comparison against reference solutions, it evaluates models by executing generated code to verify functional correctness, consistency, and efficiency. The benchmark features broad multilingual coverage across dozens of programming languages and multiple software engineering tasks, measuring model performance across dimensions such as task difficulty, code length, and execution performance using automated multilingual execution environments.