CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

Weixiang YanHaitian LiuYunkun WangYunzhe LiQian ChenWen WangTingyu LinWeishan ZhaoLi ZhuHari Sundaram

article2024ACL57 citations

Introduces CodeScope, an execution-based evaluation benchmark spanning 43 programming languages and eight tasks to measure large language model coding performance across length, difficulty, and runtime efficiency.

Listen

Software engineering increasingly relies on large language models to automate software development, maintenance, and testing. However, existing benchmarks provide an incomplete and potentially misleading view of model capabilities because they focus heavily on Python, simple synthetic problems, and superficial text-matching metrics rather than actual code execution. In practical production environments, software systems require multi-language support, complex problem-solving across diverse programming paradigms, and functionally correct execution. The article introduces CodeScope to resolve these shortcomings by establishing a comprehensive, execution-based evaluation framework across multiple tasks, programming languages, and operational dimensions.

The article systematically assesses the code understanding and code generation capabilities of eight mainstream large language models using real-world scenarios. To perform reliable execution-based scoring, the authors developed MultiCodeEngine, an execution environment supporting 47 compiler and interpreter versions across 14 programming languages. The evaluation spans 43 programming languages and eight distinct tasks, comprising four code understanding tasks (summarization, smell detection, review, and test generation) and four code generation tasks (synthesis, translation, repair, and optimization). Models are evaluated across three core dimensions: code length, task difficulty, and resource efficiency.

The findings reveal that current models struggle significantly with complex and realistic coding demands. First, while proprietary frontier models such as GPT-4 and GPT-3.5 lead in code generation, all evaluated models experience sharp performance degradation as problem difficulty increases; for example, GPT-4 achieves a 58.57% pass rate on easy synthesis problems but drops to 10.99% on hard problems. Second, open-source and specialized models lag considerably behind closed models in generation, with most failing entirely on complex synthesis tasks. Third, top performance in code generation does not imply superior code comprehension. Specialized models like WizardCoder lead the understanding benchmarks, whereas GPT-4 ranks fifth due to difficulties in generating valid automated test cases that match runtime execution flows. Finally, in efficiency optimization, models achieve modest success in high-level languages like Python but struggle with low-level languages like C, with most optimizations limited to surface-level syntactic tweaks rather than algorithmic improvements.

These results demonstrate that single-metric and single-language benchmarks significantly overestimate the readiness of language models for autonomous software engineering. Relying on current models for complex software development carries significant risk of runtime bugs, compilation errors, and poor resource utilization. Organizations adopting coding assistants should implement strict verification pipelines and avoid deploying model-generated code directly to production without automated test execution and human review. Furthermore, technical leaders should choose models based on specific task profiles—such as utilizing specialized models for code comprehension and review, while reserving frontier models for generation.

Moving forward, the article recommends pursuing two technical development paths: directly improving base model capabilities to solve harder algorithmic problems and exploring autonomous multi-agent systems to divide complex development tasks into manageable units. The primary limitation noted by the authors is potential pre-training data contamination, an inherent challenge across modern foundation models. However, because CodeScope relies on multi-source datasets, diverse downstream tasks, and strict runtime execution, stakeholders can maintain high confidence in the benchmark’s comparative findings.

Cover for CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

Abstract

Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation. However, existing benchmarks for evaluating the code understanding and generation capacities of LLMs suffer from severe limitations. First, most benchmarks are insufficient as they focus on a narrow range of popular programming languages and specific tasks, whereas real-world software development scenarios show a critical need to implement systems with multilingual and multitask programming environments to satisfy diverse requirements. Second, most benchmarks fail to consider the actual executability and the consistency of execution results of the generated code. To bridge these gaps between existing benchmarks and expectations from practical applications, we introduce CodeScope, an execution-based, multilingual, multitask, multidimensional evaluation benchmark for comprehensively measuring LLM capabilities on coding tasks. CodeScope covers 43 programming languages and eight coding tasks. It evaluates the coding performance of LLMs from three dimensions (perspectives): length, difficulty, and efficiency. To facilitate execution-based evaluations of code generation, we develop MultiCodeEngine, an automated code execution engine that supports 14 programming languages. Finally, we systematically evaluate and analyze eight mainstream LLMs and demonstrate the superior breadth and challenges of CodeScope for evaluating LLMs on code understanding and generation tasks compared to other benchmarks. The CodeScope benchmark and code are publicly available at https://github.com/WeixiangYAN/CodeScope.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The CodeScope Benchmark
  • 3.1 Code Understanding
  • 3.1.1 Code Summarization
  • 3.1.2 Code Smell
  • 3.1.3 Code Review
  • 3.1.4 Automated Testing
  • 3.2 Code Generation
  • 3.2.1 Program Synthesis (NL-to-PL)
  • 3.2.2 Code Translation (PL-to-PL)
  • 3.2.3 Code Repair (NL&PL-to-PL)
  • 3.2.4 Code Optimization
  • 4 Multidimensional Evaluation
  • 4.1 Length 5
  • 4.2 Difficulty
  • 4.3 Efficiency
  • 5 Comparison with HumanEval and MBPP Benchmarks
  • 6 Conclusion
  • Limitations
  • References
  • A Appendix
  • A.1 Statistics of CodeScope
  • A.2 Detailed Related Work
  • A.3 The CodeScope Benchmark A.3.1 Code Summarization
  • A.3.2 Code Smell
  • A.3.3 Code Review
  • A.3.4 Automated Testing
  • A.3.5 Program Synthesis
  • A.3.6 Code Translation
  • A.3.7 Code Repair
  • A.3.8 Code Optimization
  • A.4 Experimental Setup
  • A.5 Case Study

Knowls

  1. Knowl 1 — CodeScope’s benchmark scope and task coverage

    model/method

    CodeScope is an execution-based benchmark for evaluating code understanding and code generation across eight tasks and 43 distinct programming languages (about 13 languages per task on average). Its test sets contain 13,390 samples in total: code summarization, 4,838 samples across 43 languages (385 tokens per sample on average); code smell, 200 across 2 languages (650 tokens); code review, 900 across 9 languages (857 tokens); automated testing, 400 across 4 languages (251 tokens); program synthesis, 803 across 14 languages (538 tokens); code translation, 5,382 across 14 languages (513 tokens); code repair, 746 across 14 languages (446 tokens); and code optimization, 121 across 4 languages (444 tokens). The benchmark evaluates input length, problem difficulty, and execution efficiency, rather than treating code generation accuracy as its only dimension.

  2. Knowl 2 — The eight task definitions

    definition

    CodeScope’s four code-understanding tasks ask a model to: summarize source code in natural language; classify a potentially smelly snippet using its surrounding source and five candidate smell categories; estimate whether a code change warrants review comments or generate a comment for it; and produce test cases from a problem description and code solution. Its four code-generation tasks ask a model to: synthesize code from a natural-language problem description and examples; translate source code into a specified target language while preserving behavior; repair buggy code using its problem description and compiler or interpreter error information; and optimize source code for efficiency given the problem, language, and representative input/output examples.

  3. Knowl 3 — Construction and quality controls for understanding datasets

    experimental setup

    CodeScope’s code-summarization set draws from Rosetta Code: 170 tasks yield 4,838 samples across 43 languages, with at least 30 samples per language. Authors manually wrote reference summaries, used GPT-4 only to paraphrase those summaries (without the code as input), and manually reviewed the results. The code-smell set combines Java and C# data, selecting 100 samples per language and balancing five categories: large class, data class, blob, feature envy, and long method. The code-review set uses GitHub changes in Python, Java, Go, C++, JavaScript, C, C#, PHP, and Ruby, selecting 200 length-filtered samples per language. The automated-testing set contains 100 manually selected samples each in Python, Java, C, and C++; each reference solution was checked for 100% pass rate, line coverage, and branch coverage.

  4. Knowl 4 — Construction of the Codeforces-based generation datasets

    experimental setup

    CodeScope’s program-synthesis, translation, and repair datasets are built from Codeforces problems and submissions in C++, Java, Python, C, C#, Ruby, Delphi, Go, JavaScript, Kotlin, PHP, D, Perl, and Rust. Problems are divided by Codeforces rating into Easy, [800,1600)[800,1600), and Hard, [1600,2800)[1600,2800). Synthesis data excludes problems with fewer than 10 tests, nondeterministic outputs, reference submissions that fail to compile in the evaluation environments, and brute-force solutions longer than 5,000 tokens. Translation uses this problem set but limits sampled source–target language pairs to 15 per difficulty level. Repair adds incorrect submissions and their execution-derived error information. For optimization, the authors select tasks in Python 3, C#, C, and C++ with more than 10 accepted submissions and more than 20 tests, then choose submissions with high observed runtime or memory use as optimization candidates.

  5. Knowl 5 — MultiCodeEngine enables execution-based multilingual evaluation

    model/method

    MultiCodeEngine is CodeScope’s integrated code-execution environment for evaluating generated programs. It supports the 14 programming languages used in the Codeforces-based generation tasks and provides 47 compiler or interpreter versions across those languages. The engine is used to compile or run candidate code and evaluate its behavior against tests, supporting execution-based assessment for program synthesis, code translation, and code repair.

  6. Knowl 6 — Task-specific evaluation measures

    definition

    CodeScope uses task-specific measures rather than a single code-similarity score. Code summarization is scored with BLEU, METEOR, ROUGE, and BERTScore; code-smell classification and code-review quality estimation use accuracy, precision, recall, and weighted F1; review-comment generation uses BLEU, ROUGE, and BERTScore. Generated automated tests are assessed by pass rate, line coverage, and branch coverage. Program synthesis and translation use Pass@kk. Code repair uses Debugging Success Rate@KK (DSR@KK): a sample counts as repaired when the model produces the expected behavior within at most KK debugging rounds, given that the original code did not. For code optimization, Opt@KK counts a sample as optimized when at least one of KK generated candidates is more efficient than the original while preserving its intended functionality; efficiency is measured using execution time and memory use.

  7. Knowl 7 — How CodeScope operationalizes length, difficulty, and efficiency

    experimental setup

    For length evaluation, samples are tokenized with OpenAI’s tiktoken tokenizer. Within each task and programming language, the authors remove length outliers using the interquartile-range rule, divide the remaining samples evenly into short, medium, and long groups, and assign the outliers to the short or long groups. For Codeforces-based synthesis, translation, and repair, difficulty is divided into Easy ratings [800,1600)[800,1600) and Hard ratings [1600,2800)[1600,2800). Efficiency evaluation is confined to code optimization and examines execution time and memory use in four languages: Python 3, C#, C, and C++.

  8. Knowl 8 — Understanding results vary by task and benchmark

    empirical result

    Across CodeScope’s four understanding tasks, WizardCoder has the highest reported aggregate score, 50.14, followed by LLaMA 2 at 48.79 and GPT-3.5 at 48.10; GPT-4 scores 47.16 and ranks fifth. GPT-4’s relatively weak automated-testing results contribute to its lower understanding aggregate, and the authors report that its generated tests sometimes do not match actual execution outputs. Length stability also differs: the standard deviations across length groups are 2.66 for GPT-4 and 2.68 for Vicuna, the lowest reported values. Model rankings are benchmark-dependent: GPT-4 ranks first on HumanEval and MBPP but fifth on CodeScope understanding, while WizardCoder leads CodeScope understanding. The authors also report that CodeScope generation solutions average 507.6 tokens, compared with 53.8 for HumanEval and 57.6 for MBPP.

  9. Knowl 9 — Code-generation performance falls sharply on harder problems

    empirical result

    On CodeScope’s Codeforces-based generation tasks, GPT-4 leads the reported aggregate results for program synthesis, translation, and repair. Its averages are 36.36 for synthesis (Pass@5), 31.29 for translation (Pass@1), and 30.03 for repair (DSR@1); GPT-3.5 scores 22.91, 21.37, and 13.54, respectively. GPT-4’s synthesis score falls from 58.57 on Easy problems to 12.01 on Hard problems, and its repair score falls from 43.56 to 14.04. Several other tested models score much lower overall; for some of them, repair is easier than generating a solution from scratch. The results demonstrate that performance on easier synthesis problems does not reliably indicate performance on harder problems or other generation tasks.

  10. Knowl 10 — Optimization results differ by language and efficiency target

    empirical result

    In CodeScope’s Opt@5 optimization evaluation, GPT-4 has the highest reported overall score, 28.20, ahead of GPT-3.5 at 26.46 and WizardCoder at 24.37. Scores separately track memory and execution-time optimization by language: for GPT-4, the Python memory/time scores are 46.67/36.67, the C scores are 43.33/6.67, the C++ scores are 29.04/3.23, and the C# scores are 36.67/23.33. GPT-4 is not best on every measure—for example, GPT-3.5’s C memory score is 76.67. The authors report that models optimize Python most successfully and C execution time least successfully, and that many successful changes are syntactic rather than substantial algorithmic improvements.

  11. Knowl 11 — Zero data leakage cannot be guaranteed

    limitation

    The authors identify possible training-data leakage as a limitation of evaluating LLMs with fixed datasets. They state that a completely leakage-free test set is technically infeasible because many training corpora are closed and models are continually updated. CodeScope draws on five independent data sources to reduce reliance on any one source, but this does not establish that its samples were absent from model pretraining. The authors also argue that possible exposure does not make evaluation meaningless: benchmark tasks can differ from pretraining contexts and therefore still test transfer, while memorization and recitation may themselves be useful capabilities.

Coverage note — Detailed per-language score matrices and prompt/output case studies are omitted because they illustrate or expand the aggregate findings without adding distinct benchmark design or conclusions.

References

  1. 1.Toufique Ahmed and Premkumar T. Devanbu. 2022. Few-shot training llms for project-specific code-summarization. In 37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022, pages 177:1–177:5. ACM.
  2. 2.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023. Palm 2 technical report. CoRR, abs/2305.10403.
  3. 3.Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, and Ramesh Nallapati. 2023. Multilingual evaluation of code generation models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  4. 4.Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program synthesis with large language models. CoRR, abs/2108.07732.
  5. 5.Matej Balog, Alexander L. Gaunt, Marc Brockschmidt, Sebastian Nowozin, and Daniel Tarlow. 2017. Deepcoder: Learning to write programs. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  6. 6.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  7. 7.Cédric Bastoul. 2004. Code generation in the polyhedral model is easier than you think. In 13th International Conference on Parallel Architectures and Compilation Techniques (PACT 2004), 29 September - 3 October 2004, Antibes Juan-les-Pins, France, pages 7–16. IEEE Computer Society.
  8. 8.Berkay Berabi, Jingxuan He, Veselin Raychev, and Martin T. Vechev. 2021. Tfix: Learning to fix coding errors with a text-to-text transformer. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 780–791. PMLR.
  9. 9.Uday Bondhugula, Albert Hartono, J. Ramanujam, and P. Sadayappan. 2008. A practical automatic polyhedral parallelizer and locality optimizer. In Proceedings of the ACM SIGPLAN 2008 Conference on Programming Language Design and Implementation, Tucson, AZ, USA, June 7-13, 2008, pages 101–113. ACM.
  10. 10.Leslie Pérez Cáceres, Federico Pagnozzi, Alberto Franzin, and Thomas Stützle. 2017. Automatic configuration of GCC using irace. In Artificial Evolution - 13th International Conference, Évolution Artificielle, EA 2017, Paris, France, October 25-27, 2017, Revised Selected Papers, volume 10764 of Lecture Notes in Computer Science, pages 202–216. Springer.
  11. 11.Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium, USENIX Security 2019, Santa Clara, CA, USA, August 14-16, 2019, pages 267–284. USENIX Association.
  12. 12.Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. 2022. Multipl-e: A scalable and extensible approach to benchmarking neural code generation.
  13. 13.Shubham Chandel, Colin B. Clement, Guillermo Serrato, and Neel Sundaresan. 2022. Training and evaluating a jupyter notebook data science assistant. CoRR, abs/2201.12901.
  14. 14.Chun Chen, Jacqueline Chame, and Mary Hall. 2008. Chill: A framework for composing high-level loop transformations. Technical report, Citeseer.
  15. 15.Dehao Chen, David Xinliang Li, and Tipp Moseley. 2016. Autofdo: automatic feedback-directed optimization for warehouse-scale applications. In Proceedings of the 2016 International Symposium on Code Generation and Optimization, CGO 2016, Barcelona, Spain, March 12-18, 2016, pages 12–23. ACM.
  16. 16.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. CoRR, abs/2107.03374.
  17. 17.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna.lmsys. org (accessed 14 April 2023).
  18. 18.Ananta Kumar Das, Shikhar Yadav, and Subhasish Dhal. 2019. Detecting code smells using deep learning. In TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON), Kochi, India, October 17-20, 2019, pages 2081–2086. IEEE.
  19. 19.Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli. 2017. Robustfill: Neural program learning under noisy I/O. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 990–998. PMLR.
  20. 20.Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou. 2023. Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation. CoRR, abs/2308.01861.
  21. 21.Martin Fowler. 1999. Refactoring - Improving the Design of Existing Code. Addison Wesley object technology series. Addison-Wesley.
  22. 22.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks. CoRR, abs/2303.15056.
  23. 23.Rahul Gupta, Soham Pal, Aditya Kanade, and Shirish Shevade. 2017. Deepfix: Fixing common c language errors by deep learning. In Proceedings of the aaai conference on artificial intelligence.
  24. 24.Sonia Haiduc, Jairo Aponte, Laura Moreno, and Andrian Marcus. 2010. On the use of automated text summarization techniques for summarizing source code. In 17th Working Conference on Reverse Engineering, WCRE 2010, 13-16 October 2010, Beverly, MA, USA, pages 35–44. IEEE Computer Society.
  25. 25.Yiyang Hao, Ge Li, Yongqiang Liu, Xiaowei Miao, He Zong, Siyuan Jiang, Yang Liu, and He Wei. 2022. Aixbench: A code generation benchmark dataset. CoRR, abs/2206.13179.
  26. 26.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual.
  27. 27.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. CoRR, abs/2009.03300.
  28. 28.Junjie Huang, Chenglong Wang, Jipeng Zhang, Cong Yan, Haotian Cui, Jeevana Priya Inala, Colin B. Clement, Nan Duan, and Jianfeng Gao. 2022. Execution-based evaluation for data science code generation models. CoRR, abs/2211.09374.
  29. 29.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. CoRR, abs/2305.08322.
  30. 30.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016. Summarizing source code using a neural attention model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics.
  31. 31.René Just, Darioush Jalali, and Michael D. Ernst. 2014. Defects4j: a database of existing faults to enable controlled testing studies for java programs. In International Symposium on Software Testing and Analysis, ISSTA ’14, San Jose, CA, USA - July 21 - 26, 2014, pages 437–440. ACM.
  32. 32.Mohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Long Do, Weishi Wang, Md. Rizwan Parvez, and Shafiq R. Joty. 2023. xcodeeval: A large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval. CoRR, abs/2303.03004.
  33. 33.Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-Tau Yih, Daniel Fried, Sida I. Wang, and Tao Yu. 2023. DS-1000: A natural and reliable benchmark for data science code generation. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 18319–18345. PMLR.
  34. 34.Alexander LeClair, Sakib Haque, Lingfei Wu, and Collin McMillan. 2020. Improved code summarization via a graph neural network. In ICPC ’20: 28th International Conference on Program Comprehension, Seoul, Republic of Korea, July 13-15, 2020, pages 184–195. ACM.
  35. 35.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023a. CMMLU: measuring massive multitask language understanding in chinese. CoRR, abs/2306.09212.
  36. 36.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy V, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Moustafa-Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023b. Starcoder: may the source be with you! CoRR, abs/2305.06161.
  37. 37.Tsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian, Ying Wang, Shing-Chi Cheung, and Jeff Kramer. 2023c. Finding failure-inducing test cases with chatgpt. CoRR, abs/2304.11686.
  38. 38.Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022a. Competition-level code generation with alphacode. CoRR, abs/2203.07814.
  39. 39.Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, and Neel Sundaresan. 2022b. Automating code review activities by large-scale pre-training. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE 2022, Singapore, Singapore, November 14-18, 2022, pages 1035–1047. ACM.
  40. 40.Chin-Yew Lin and Eduard H. Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, HLT-NAACL 2003, Edmonton, Canada, May 27 - June 1, 2003. The Association for Computational Linguistics.
  41. 41.Derrick Lin, James Koppel, Angela Chen, and Armando Solar-Lezama. 2017. Quixbugs: a multi-lingual program repair benchmark set based on the quixey challenge. In Proceedings Companion of the 2017 ACM SIGPLAN International Conference on Systems, Programming, Languages, and Applications: Software for Humanity, SPLASH 2017, Vancouver, BC, Canada, October 23 - 27, 2017, pages 55–56. ACM.
  42. 42.Tao Lin, Xue Fu, Fu Chen, and Luqun Li. 2021. A novel approach for code smells detection based on deep leaning. In Applied Cryptography in Computer and Communications: First EAI International Conference, AC3 2021, Virtual Event, May 15-16, 2021, Proceedings 1, pages 171–174. Springer.
  43. 43.Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, and Andrew W. Senior. 2016. Latent predictor networks for code generation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics.
  44. 44.Fan Long and Martin C. Rinard. 2015. Staged program repair with condition synthesis. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC/FSE 2015, Bergamo, Italy, August 30 - September 4, 2015, pages 166–178. ACM.
  45. 45.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual.
  46. 46.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct.
  47. 47.Lech Madeyski and Tomasz Lewowski. 2023. Detecting code smells using industry-relevant data. Inf. Softw. Technol., 155:107112.
  48. 48.Radu Marinescu. 2005. Measurement and quality in object-oriented design. In 21st IEEE International Conference on Software Maintenance (ICSM 2005), 25-30 September 2005, Budapest, Hungary, pages 701–704. IEEE Computer Society.
  49. 49.Shane McIntosh, Yasutaka Kamei, Bram Adams, and Ahmed E. Hassan. 2014. The impact of code review coverage and code review participation on software quality: a case study of the qt, vtk, and ITK projects. In 11th Working Conference on Mining Software Repositories, MSR 2014, Proceedings, May 31 - June 1, 2014, Hyderabad, India, pages 192–201. ACM.
  50. 50.Naouel Moha, Yann-Gaël Guéhéneuc, Laurence Duchien, and Anne-Françoise Le Meur. 2010. DECOR: A method for the specification and detection of code and design smells. IEEE Trans. Software Eng., 36(1):20–36.
  51. 51.Anh Tuan Nguyen, Tung Thanh Nguyen, and Tien N. Nguyen. 2013a. Lexical statistical machine translation for language migration. In Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering, ESEC/FSE’13, Saint Petersburg, Russian Federation, August 18-26, 2013, pages 651–654. ACM.
  52. 52.Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chandra. 2013b. Semfix: program repair via semantic analysis. In 35th International Conference on Software Engineering, ICSE ’13, San Francisco, CA, USA, May 18-26, 2013, pages 772–781. IEEE Computer Society.
  53. 53.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. Codegen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  54. 54.OpenAI. 2023. GPT-4 technical report.
  55. 55.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA, pages 311–318. ACL.
  56. 56.Karl Pettis and Robert C. Hansen. 1990. Profile guided code positioning. In Proceedings of the ACM SIGPLAN’90 Conference on Programming Language Design and Implementation (PLDI), White Plains, New York, USA, June 20-22, 1990, pages 16–27. ACM.
  57. 57.Dmitry Plotnikov, Dmitry Melnik, Mamikon Vardanyan, Ruben Buchatskiy, and Roman Zhuykov. 2013. An automatic tool for tuning compiler optimizations. In Ninth International Conference on Computer Science and Information Technologies Revised Selected Papers, pages 1–7. IEEE.
  58. 58.Mihail Popov, Chadi Akel, Yohan Chatelain, William Jalby, and Pablo de Oliveira Castro. 2017. Piecewise holistic autotuning of parallel programs with CERE. Concurr. Comput. Pract. Exp., 29(15).
  59. 59.Julian Aron Prenner and Romain Robbes. 2021. Automatic program repair with openai’s codex: Evaluating quixbugs. CoRR, abs/2111.03922.
  60. 60.Ruchir Puri, David S. Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir R. Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, and Ulrich Finkler. 2021. Project codenet: A large-scale AI for code dataset for learning a diversity of coding tasks. CoRR, abs/2105.12655.
  61. 61.Yuhua Qi, Xiaoguang Mao, Yan Lei, Ziying Dai, and Chengsong Wang. 2014. The strength of random search on automated program repair. In 36th International Conference on Software Engineering, ICSE ’14, Hyderabad, India - May 31 - June 07, 2014, pages 254–265. ACM.
  62. 62.Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020. Codebleu: a method for automatic evaluation of code synthesis. CoRR, abs/2009.10297.
  63. 63.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code. CoRR, abs/2308.12950.
  64. 64.Mazeiar Salehie, Shimin Li, and Ladan Tahvildari. 2006. A metric-based heuristic framework to detect object-oriented design flaws. In 14th International Conference on Program Comprehension (ICPC 2006), 14-16 June 2006, Athens, Greece, pages 159–168. IEEE Computer Society.
  65. 65.Tushar Sharma, Pratibha Mishra, and Rohit Tiwari. 2016. Designite: a software design quality assessment tool. In Proceedings of the 1st International Workshop on Bringing Architectural Design Thinking into Developers’ Daily Activities, BRIDGE@ICSE 2016, Austin, Texas, USA, May 17, 2016, pages 1–4. ACM.
  66. 66.Ensheng Shi, Yanlin Wang, Lun Du, Junjie Chen, Shi Han, Hongyu Zhang, Dongmei Zhang, and Hongbin Sun. 2022. On the evaluation of neural code summarization. In 44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022, pages 1597–1608. ACM.
  67. 67.Mohammed Latif Siddiq, Joanna C. S. Santos, Ridwanul Hasan Tanvir, Noshin Ulfat, Fahmid Al Rifat, and Vinicius Carvalho Lopes. 2023. Exploring the effectiveness of large language models in generating unit tests. CoRR, abs/2305.00418.
  68. 68.Jelena Slivka, Nikola Luburic, Simona Prokic, Katarina-Glorija Grujic, Aleksandar Kovacevic, Goran Sladic, and Dragan Vidakovic. 2023. Towards a systematic approach to manual annotation of code smells. Sci. Comput. Program., 230:102999.
  69. 69.Giriprasad Sridhara, Emily Hill, Divya Muppaneni, Lori L. Pollock, and K. Vijay-Shanker. 2010. Towards automatically generating summary comments for java methods. In ASE 2010, 25th IEEE/ACM International Conference on Automated Software Engineering, Antwerp, Belgium, September 20-24, 2010, pages 43–52. ACM.
  70. 70.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  71. 71.Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2019. An empirical study on learning bug-fixing patches in the wild via neural machine translation. ACM Trans. Softw. Eng. Methodol., 28(4):19:1–19:29.
  72. 72.Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota. 2022. Using pre-trained models to boost code review automation. In 44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022, pages 2291–2302. ACM.
  73. 73.Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R. Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. 2023a. Scibench: Evaluating college-level scientific problem-solving abilities of large language models. CoRR, abs/2307.10635.
  74. 74.Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D. Q. Bui, Junnan Li, and Steven C. H. Hoi. 2023b. Codet5+: Open code large language models for code understanding and generation. CoRR, abs/2305.07922.
  75. 75.Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009. Automatically finding patches using genetic programming. In 31st International Conference on Software Engineering, ICSE 2009, May 16-24, 2009, Vancouver, Canada, Proceedings, pages 364–374. IEEE.
  76. 76.David Williams-King and Junfeng Yang. 2019. Codemason: Binary-level profile-guided optimization. In Proceedings of the 3rd ACM Workshop on Forming an Ecosystem Around Software Transformation, pages 47–53.
  77. 77.Zhuokui Xie, Yinghao Chen, Chen Zhi, Shuiguang Deng, and Jianwei Yin. 2023. Chatunitest: a chatgpt-based automated unit test generation tool. CoRR, abs/2305.04764.
  78. 78.Weixiang Yan and Yuanchun Li. 2022. Whygen: Explaining ml-powered code generation by referring to training examples. In 44th IEEE/ACM International Conference on Software Engineering: Companion Proceedings, ICSE Companion 2022, Pittsburgh, PA, USA, May 22-24, 2022, pages 237–241. ACM/IEEE.
  79. 79.Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen, and Wen Wang. 2023. Codetransocean: A comprehensive multilingual benchmark for code translation. CoRR, abs/2310.04951.
  80. 80.Pengcheng Yin and Graham Neubig. 2017. A syntactic neural model for general-purpose code generation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 440–450. Association for Computational Linguistics.
  81. 81.Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Tao Xie, and Qianxiang Wang. 2023. Codereval: A benchmark of pragmatic code generation with generative pretrained models. CoRR, abs/2302.00288.
  82. 82.Zhiqiang Yuan, Yiling Lou, Mingwei Liu, Shiji Ding, Kaixin Wang, Yixuan Chen, and Xin Peng. 2023. No more manual tests? evaluating and improving chatgpt for unit test generation. CoRR, abs/2305.04207.
  83. 83.Kechi Zhang, Ge Li, Jia Li, Zhuo Li, and Zhi Jin. 2023a. Toolcoder: Teach code generation models to use API search tools. CoRR, abs/2305.04032.
  84. 84.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  85. 85.Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023b. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models. CoRR, abs/2306.05179.
  86. 86.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023a. Judging llm-as-a-judge with mt-bench and chatbot arena.
  87. 87.Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023b. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. CoRR, abs/2303.17568.
  88. 88.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models.
  89. 89.Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2023. Language agent tree search unifies reasoning acting and planning in language models. CoRR, abs/2310.04406.
  90. 90.Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, and Chandan K. Reddy. 2022. Xlcost: A benchmark dataset for cross-lingual code intelligence.

Citation

MLA
Yan, W., et al. “CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5511–58, https://doi.org/10.18653/v1/2024.acl-long.301.
APA
Yan, W., Liu, H., Wang, Y., Li, Y., Chen, Q., (王雯), W. W., Lin, T., Zhao, W., Zhu, L., Sundaram, H., & Deng, S. (2024). CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5511–5558. https://doi.org/10.18653/v1/2024.acl-long.301
Chicago
Yan, W., H. Liu, Y. Wang, et al. 2024. “CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5511–58. https://doi.org/10.18653/v1/2024.acl-long.301.
Harvard
Yan, W. et al. (2024) “CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5511–5558. Available at: https://doi.org/10.18653/v1/2024.acl-long.301.
Vancouver
1. Yan W, Liu H, Wang Y, et al (2024) CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5511–5558

BibTeX

@inproceedings{yan-etal-2024-codescope,
    title = "{C}ode{S}cope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating {LLM}s on Code Understanding and Generation",
    author = "Yan, Weixiang  and
      Liu, Haitian  and
      Wang, Yunkun  and
      Li, Yunzhe  and
      Chen, Qian  and
      Wang, Wen  and
      Lin, Tingyu  and
      Zhao, Weishan  and
      Zhu, Li  and
      Sundaram, Hari  and
      Deng, Shuiguang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.301/",
    doi = "10.18653/v1/2024.acl-long.301",
    pages = "5511--5558"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/