MapCoder: Multi-Agent Code Generation for Competitive Problem Solving

Md. Ashraful IslamMohammed Eunus AliMd. Rizwan Parvez

article2024ACL203 citations

Proposes MapCoder, a multi-agent framework that mirrors the human programming workflow through autonomous example retrieval, planning, coding, and sample-based debugging to achieve state-of-the-art accuracy on challenging competitive programming benchmarks like CodeContests and APPS.

Listen

Automating software program synthesis using large language models is a major objective for modern software engineering, yet state-of-the-art models frequently fail on complex, competition-level programming tasks. These tasks demand rigorous multi-step logical reasoning, sophisticated data structures, and precise execution against complex test cases. Current prompting and self-reflection techniques often underperform because they lack structured planning, fail to retrieve past problem patterns, or rely heavily on artificially generated test cases that can be faulty and degrade code quality.

The article introduces and evaluates MapCoder, a multi-agent prompting framework designed to emulate the complete human programming cycle. MapCoder demonstrates how coordinating four specialized agents—self-retrieval, planning, coding, and debugging—along with a dynamic agent traversal protocol can substantially improve automated code generation without requiring external tool dependencies or synthetic test generation.

To establish credibility across diverse problem settings, the researchers evaluated MapCoder against standard baselines across eight benchmark datasets covering both basic programming and difficult competitive challenges. The framework was tested using multiple foundation models, including ChatGPT, GPT-4, Gemini Pro, and an open-source model, Mistral-7B-instruct. MapCoder operates by autonomously generating relevant past examples and algorithms, scoring multiple candidate solution plans, generating code based on the highest-confidence plan, and iteratively debugging failed attempts using only the provided sample input/output test cases.

The key findings show that MapCoder establishes new state-of-the-art benchmark results. On competition-level benchmarks using GPT-4, MapCoder achieved pass rates of 22.0% on APPS, 45.3% on xCodeEval, and 28.5% on CodeContest, which represent performance improvements of approximately 74%, 41%, and 135% over direct prompting methods. On standard programming benchmarks, it reached 93.9% on HumanEval and 83.1% on MBPP. Additionally, on CodeContest, MapCoder's single-attempt performance matched the multi-attempt score of the current leading alternative, AlphaCodium. Ablation analyses revealed that all agents contribute to performance, with the debugging and planning agents providing the most critical improvements, reducing performance by roughly 25% and 17% on average when removed.

These results demonstrate that mimicking structured human workflows provides a more reliable method for automated code generation than simply increasing raw model parameters or relying on self-generated test cases. By eliminating synthetic test generation, MapCoder avoids the performance degradation seen in existing self-reflection methods when models generate incorrect test scenarios. However, the multi-agent dynamic traversal structure incurs higher computational costs, averaging roughly 17 API calls and 21,000 tokens per problem across the evaluated benchmarks.

Organizations evaluating automated coding workflows should consider adopting multi-agent planning and plan-derived debugging rather than direct prompting or fragile synthetic test-generation pipelines. Future engineering efforts should focus on optimizing token consumption, refining agent traversal efficiency, and exploring cost-effective open-source foundation models. The authors also recommend executing all machine-generated solutions within isolated sandbox environments to mitigate security and execution risks.

While the findings demonstrate high confidence across multiple programming languages and standard benchmarks, MapCoder still exhibits performance plateaus on highly complex algorithmic categories, particularly advanced dynamic programming, combinatorics, and number theory. Stakeholders should maintain human oversight for mission-critical software and highly complex algorithm engineering.

Cover for MapCoder: Multi-Agent Code Generation for Competitive Problem Solving

Abstract

Code synthesis, which requires a deep understanding of complex natural language (NL) problem descriptions, generation of code instructions for complex algorithms and data structures, and the successful execution of comprehensive unit tests, presents a significant challenge. Thus, while large language models (LLMs) demonstrate impressive proficiency in natural language processing (NLP), their performance in code generation tasks remains limited. In this paper, we introduce a new approach to code generation tasks leveraging the multi-agent prompting that uniquely replicates the full cycle of program synthesis as observed in human developers. Our framework, MapCoder, consists of four LLM agents specifically designed to emulate the stages of this cycle: recalling relevant examples, planning, code generation, and debugging. After conducting thorough experiments, with multiple LLMs ablations and analyses across eight challenging competitive problem-solving and program synthesis benchmarks—MapCoder showcases remarkable code generation capabilities, achieving their new state-of-the-art (pass@1) results—(HumanEval 93.9%, MBPP 83.1%, APPS 22.0%, CodeContests 28.5%, and xCodeEval 45.3%). Moreover, our method consistently delivers superior performance across various programming languages and varying problem difficulties. We open-source our framework at https://github.com/Md-Ashraful-Pramanik/MapCoder.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MapCoder
  • 3.1 Retrieval Agent
  • 3.2 Planning Agent
  • 3.3 Coding Agent
  • 3.4 Debugging Agent
  • 3.5 Dynamic Agent Traversal
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Foundation Models, Evaluation Metric, k, and t
  • 5 Results
  • 5.1 Performance on basic code generation
  • 5.2 Performance on competitive problem solving
  • 5.3 Performance with Varying Difficulty Levels
  • 5.4 Performance Across Different LLMs
  • 5.5 Performance Across Different Programming Languages
  • 6 Ablations Studies and Analyses
  • 6.1 Impact of Different Agents
  • 6.2 Qualitative Example
  • 6.3 Impact of k and t
  • 6.4 Impact of Number of Sample I/Os
  • 6.5 Error Analysis and Challenges
  • 7 Conclusion and Future Work
  • 8 Limitations
  • Acknowledgements
  • References
  • Appendix
  • A Algorithm of MapCoder
  • B Details Promptings of MapCoder
  • C Example Problem
  • C.1 An example containing problem from HumanEval Dataset (k=5, t=5)
  • C.2 An example containing problem from CodeContest Dataset (k=3, t=5)

Knowls

  1. Knowl 1 — MapCoder Multi-Agent Architecture for Program Synthesis

    model/method

    MapCoder is a multi-agent prompting framework designed to solve complex and competition-level programming problems by replicating the human developer workflow. Rather than generating code directly or relying on external retrieval models and human annotations, MapCoder structures program synthesis into four cascaded LLM agents:

    1. Retrieval Agent: Autonomously retrieves and generates kk relevant, distinct analogical problems along with step-by-step code solutions, reverse-engineered plans, and tutorials identifying relevant algorithmic paradigms (e.g., dynamic programming, divide-and-conquer, greedy).
    2. Planning Agent: Formulates a distinct step-by-step plan for the target problem tailored to each individual retrieved exemplar, accompanied by a numerical confidence score (00 to 100100) indicating the plan's feasibility.
    3. Coding Agent: Synthesizes full code in the target programming language by conditioning on the problem statement, relevant algorithms, and a specific plan from the Planning Agent, subsequently running the code on available sample input/output (I/O) cases.
    4. Debugging Agent: Fixes implementation errors using feedback logs from failed sample I/O tests while explicitly leveraging the algorithmic plan from the Planning Agent as reference, iterating up to tt turns per plan.

    The framework operates dynamically: agents interact iteratively via plan scoring and test execution feedback without requiring synthetic test-case generation.

  2. Knowl 2 — Dynamic Agent Traversal Algorithm in MapCoder

    algorithm

    The dynamic agent traversal algorithm coordinates the four agents of MapCoder by prioritizing high-confidence plans and executing an iterative plan-guided debugging loop on sample test cases. Given a target problem, sample input/output test cases S\mathcal{S}, the number of self-retrieved exemplars kk, and the maximum debugging attempts per plan tt, the traversal executes with an overall computational time complexity of O(kt)O(kt).

    Input: Problem PP, sample test cases SS, exemplar count kk, max debug attempts tt
    Output: Executable solution code CC
    exemplars <- RetrievalAgent(P, k)
    plans <- empty list of size k
    for each example in exemplars do
        plan, confidence <- PlanningAgent(P, example)
        plans.append((plan, confidence))
    end for
    sorted_plans <- SortByConfidenceDescending(plans)
    for i <- 1 to k do
        current_plan <- sorted_plans[i].plan
        code <- CodingAgent(P, current_plan)
        passed, log <- ExecuteTests(code, S)
        if passed then
            return code
        else
            for j <- 1 to t do
                code <- DebuggingAgent(P, current_plan, code, log)
                passed, log <- ExecuteTests(code, S)
                if passed then
                    return code
                end if
            end for
        end if
    end for
    return code
  3. Knowl 3 — Self-Retrieval and Plan Confidence Generation Mechanisms

    model/method

    MapCoder eliminates dependency on external code databases and hand-crafted few-shot examples through an autonomous prompt design in its Retrieval and Planning agents:

    • Structured Self-Retrieval: The Retrieval Agent prompts the LLM to simultaneously generate kk distinct but algorithmically related problems, synthesize their step-by-step code solutions, extract corresponding procedural plans via reverse engineering, and write a high-level tutorial identifying relevant algorithmic techniques (e.g., binary search, greedy, dynamic programming). This ensures algorithmic alignment with the target task.
    • Independent Multi-Plan and Confidence Formulation: Rather than concatenating all retrieved exemplars into a single noisy context, the Planning Agent processes each retrieved exemplar independently to generate a distinct candidate plan for the target problem. For each plan, the Planning Agent queries the LLM to assess solvability and output an integer confidence score s∈[0,100]s \in [0, 100] alongside an explanation in structured XML tags (<explanation>, <confidence>). These confidence scores serve as reward metrics to sort candidate plans prior to code translation.
  4. Knowl 4 — Plan-Derived Sample-I/O Debugging Mechanism

    model/method

    MapCoder implements a plan-derived debugging mechanism that avoids the pitfalls of LLM-generated synthetic unit tests. Existing self-reflection frameworks often generate auxiliary unit tests that can be incorrect, misleading the LLM to break valid code (e.g., causing substantial accuracy drops on benchmarks like MBPP and HumanEval).

    Instead, MapCoder's Debugging Agent evaluates generated programs solely against the authoritative sample input/output (I/O) pairs provided in the problem description. When a sample test fails, the Debugging Agent receives the failed execution log, the current code, the relevant algorithm description, and the original high-level plan. By cross-checking code errors against the intended procedural plan, the agent modifies both the plan and the code iteratively up to tt turns. If the code fails after tt attempts, the traversal backtracks to the Planning Agent to select the next highest-confidence plan.

  5. Knowl 5 — Pass@1 Performance Across Program Synthesis Benchmarks

    data/table

    MapCoder achieves state-of-the-art Pass@1 performance across eight basic programming and competitive problem-solving benchmarks when evaluated with ChatGPT (gpt-3.5-turbo-1106) and GPT-4 (gpt-4-1106-preview). Hyperparameters were set to k=t=5k=t=5 for HumanEval and k=t=3k=t=3 for the other datasets.

    LLM Approach HumanEval HumanEval-ET EvalPlus MBPP MBPP-ET APPS xCodeEval CodeContest
    ChatGPT Direct 48.1% 37.2% 66.5% 49.8% 37.7% 8.0% 17.9% 5.5%
    CoT 68.9% 55.5% 65.2% 54.5% 39.6% 7.3% 23.6% 6.1%
    Self-Planning 60.3% 46.2% - 55.7% 41.9% 9.3% 18.9% 6.1%
    Analogical 63.4% 50.6% 59.1% 70.5% 46.1% 6.7% 15.1% 7.3%
    Reflexion 67.1% 49.4% 62.2% 73.0% 47.4% - - -
    Self-collaboration 74.4% 56.1% - 68.2% 49.5% - - -
    MapCoder 80.5% 70.1% 71.3% 78.3% 54.4% 11.3% 27.4% 12.7%
    Gain over Direct +67.3% +88.5% +7.3% +57.3% +44.3% +41.3% +52.6% +132.8%
    GPT-4 Direct 80.1% 73.8% 81.7% 81.1% 54.7% 12.7% 32.1% 12.1%
    CoT 89.0% 61.6% - 82.4% 56.2% 11.3% 36.8% 5.5%
    Self-Planning 85.4% 62.2% - 75.8% 50.4% 14.7% 34.0% 10.9%
    Analogical 66.5% 48.8% 62.2% 58.4% 40.3% 12.0% 26.4% 10.9%
    Reflexion 91.0% 78.7% 81.7% 78.3% 51.9% - - -
    MapCoder 93.9% 82.9% 83.5% 83.1% 57.7% 22.0% 45.3% 28.5%
    Gain over Direct +17.2% +12.4% +2.2% +2.5% +5.5% +73.7% +41.2% +135.1%

    MapCoder consistently improves upon both zero-shot direct generation and complex multi-agent baselines across all datasets, showing particularly strong relative gains (>40%>40\% to 135%135\%) on competition-level benchmarks (APPS, xCodeEval, and CodeContest).

  6. Knowl 6 — Pass@5 Performance on the CodeContest Benchmark

    data/table

    On the competitive programming benchmark CodeContest (165 test problems), MapCoder matches or exceeds the concurrent flow-engineering framework AlphaCodium under the Pass@5 evaluation metric while achieving state-of-the-art results for both ChatGPT and GPT-4.

    Approach ChatGPT (Pass@5) GPT-4 (Pass@5)
    Direct 11.2% 18.8%
    AlphaCodium 17.0% 29.0%
    MapCoder 18.2% (+63.1%) 35.2% (+87.1%)

    Notably, MapCoder's single-solution Pass@1 accuracy using GPT-4 (28.5%28.5\%) nearly matches AlphaCodium's 5-sample Pass@5 accuracy (29.0%29.0\%), and MapCoder's Pass@5 reaches 35.2%35.2\%, representing an absolute improvement of 6.26.2 percentage points over AlphaCodium.

  7. Knowl 7 — Ablation of MapCoder Agent Modules on HumanEval

    data/table

    An ablation study isolating the Retrieval Agent, Planning Agent, and Debugging Agent using ChatGPT on the HumanEval dataset demonstrates that all three agents are critical, with debugging and planning providing the largest individual performance contributions.

    Retrieval Agent Planning Agent Debugging Agent Pass@1 Performance Drop
    ✗ ✗ ✓ 68.0% 15.0%
    ✗ ✓ ✓ 76.0% 5.0%
    ✗ ✓ ✗ 52.0% 35.0%
    ✓ ✗ ✓ 70.0% 12.5%
    ✓ ✓ ✗ 66.0% 17.5%
    ✓ ✗ ✗ 62.0% 22.5%
    ✓ ✓ ✓ 80.0% -

    Excluding the Debugging Agent exclusively results in a 17.5%17.5\% absolute drop in Pass@1 accuracy, and an average performance drop of 24.83%24.83\% across all configurations where it is absent. Excluding the Planning Agent results in an average drop of 16.7%16.7\% across all combinations.

  8. Knowl 8 — Hyperparameter Sensitivity to Retrieval Exemplars and Debugging Attempts

    data/table

    Varying the number of self-retrieved exemplars k∈{3,5}k \in \{3, 5\} and the number of debugging attempts t∈{0,3,5}t \in \{0, 3, 5\} using ChatGPT reveals that higher values of kk and tt yield monotonic improvements in Pass@1 performance on HumanEval and HumanEval-ET, establishing a controllable trade-off between inference compute and solution accuracy.

    Dataset kk t=0t=0 t=3t=3 t=5t=5
    HumanEval 3 62.8% 76.8% 80.5%
    5 65.9% 79.9% 80.5%
    HumanEval-ET 3 57.3% 61.0% 70.1%
    5 57.9% 67.1% 67.1%

    Increasing the debugging attempts from t=0t=0 (no debugging) to t=5t=5 provides an absolute gain of +17.7%+17.7\% on HumanEval (k=3k=3) and +12.8%+12.8\% on HumanEval-ET (k=3k=3).

  9. Knowl 9 — Model Robustness and Multilingual Generalization of MapCoder

    empirical result

    MapCoder demonstrates strong cross-model and multilingual generalization:

    • Alternative Foundation Models: Evaluated with Gemini Pro, MapCoder improves Pass@1 from 64.6%64.6\% (Direct) to 69.5%69.5\% on HumanEval and from 3.6%3.6\% (Direct) to 4.8%4.8\% on CodeContest (+32.0%+32.0\% relative gain). Evaluated with the open-source Mistral-7B-instruct model, MapCoder improves Pass@1 from 27.3%27.3\% (Direct) to 57.6%57.6\% (+111.1%+111.1\% relative gain) on HumanEval and from 27.3%27.3\% to 48.5%48.5\% (+77.8%+77.8\%) on HumanEval-ET.
    • Multilingual Problem Solving: Across diverse programming languages evaluated in the xCodeEval dataset—including Python, C, C++, Java, C#, Go, Ruby, PHP, and Rust—MapCoder consistently yields higher solve counts than Direct prompting and Chain-of-Thought prompting baselines.
  10. Knowl 10 — Computational Overhead and Algorithmic Limitations of MapCoder

    limitation

    MapCoder exhibits two primary practical limitations:

    1. Inference Token and API Overhead: Compared to single-turn Direct Prompting (which averages 11 API call and 0.64k0.64\text{k} tokens per problem), MapCoder averages 16.716.7 API calls and 21.25k21.25\text{k} tokens per problem across datasets (1717 calls / 10.41k10.41\text{k} tokens for ChatGPT on HumanEval; 1919 calls / 38.70k38.70\text{k} tokens for GPT-4 on CodeContest).
    2. Algorithmic Failure Modes: On high-difficulty competitive programming problems (e.g., xCodeEval problems with difficulty ratings >1000> 1000), MapCoder's performance gains narrow. Error analysis in advanced categories—including Dynamic Programming (DP), Combinatorics, Constructive algorithms, and Number Theory—reveals persistent failures in accurate DP state/table construction, occasional misinterpretation of complex constraints, and fallbacks to incorrect greedy or brute-force implementations.

Coverage note — None was omitted; all principal methods, algorithms, tables, ablations, empirical findings, and limitations were fully captured.

References

  1. 1.Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. arXiv preprint arXiv:2103.06333.
  2. 2.Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al. 2023. Santacoder: don’t reach for the stars! arXiv preprint arXiv:2301.03988.
  3. 3.Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang, Alan Guo, David Hall, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Dan Klein, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Wolfe, Abby Wray, Yuchen Zhang, and Alexander Zotov. 2020. Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computational Linguistics, 8:556–571.
  4. 4.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  5. 5.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2022. Codet: Code generation with generated tests. arXiv preprint arXiv:2207.10397.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021a. Evaluating large language models trained on code.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021b. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  10. 10.Yihong Dong, Jiazheng Ding, Xue Jiang, Zhuo Li, Ge Li, and Zhi Jin. 2023a. Codescore: Evaluating code generation by learning code execution. arXiv preprint arXiv:2301.09043.
  11. 11.Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2023b. Self-collaboration code generation via chatgpt.
  12. 12.Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. 2020. Codebert: A pre-trained model for programming and natural languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1536–1547.
  13. 13.Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2022. Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999.
  14. 14.Sumit Gulwani. 2011. Automating string processing in spreadsheets using input-output examples. ACM Sigplan Notices, 46(1):317–330.
  15. 15.Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y Wu, YK Li, et al. 2024. Deepseek-coder: When the large language model meets programming–the rise of code intelligence. arXiv preprint arXiv:2401.14196.
  16. 16.Vincent J. Hellendoorn and Premkumar Devanbu. 2017. Are deep neural networks the best choice for modeling source code? In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering, ESEC/FSE 2017, pages 763–773, New York, NY, USA. ACM.
  17. 17.Abram Hindle, Earl T. Barr, Mark Gabel, Zhendong Su, and Premkumar Devanbu. 2016. On the naturalness of software. Commun. ACM, 59(5):122–131.
  18. 18.Dong Huang, Qingwen Bu, Jie M Zhang, Michael Luck, and Heming Cui. 2023. Agentcoder: Multi-agent-based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010.
  19. 19.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023a. Mistral 7b.
  20. 20.Xue Jiang, Yihong Dong, Lecheng Wang, Qiwei Shang, and Ge Li. 2023b. Self-planning code generation with large language model. arXiv preprint arXiv:2303.06689.
  21. 21.Mohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, and Shafiq Joty. 2023. xcodeeval: A large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval. arXiv preprint arXiv:2303.03004.
  22. 22.Donald E Knuth. 1992. Literate programming. CSLI Lecture Notes, Stanford, CA: Center for the Study of Language and Information (CSLI), 1992.
  23. 23.Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi. 2022. Coderl: Mastering code generation through pretrained models and deep reinforcement learning. Advances in Neural Information Processing Systems, 35:21314–21328.
  24. 24.Jingyao Li, Pengguang Chen, and Jiaya Jia. 2023. Motcoder: Elevating large language models with modular of thought for challenging programming tasks. arXiv preprint arXiv:2312.15960.
  25. 25.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022a. Competition-level code generation with alphacode. Science, 378(6624):1092–1097.
  26. 26.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022b. Competition-level code generation with alphacode.
  27. 27.Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. In Thirty-seventh Conference on Neural Information Processing Systems.
  28. 28.Zohar Manna and Richard J. Waldinger. 1971. Toward automatic program synthesis. Commun. ACM, 14(3):151–165.
  29. 29.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474.
  30. 30.Carlos Pacheco, Shuvendu K Lahiri, Michael D Ernst, and Thomas Ball. 2007. Feedback-directed random test generation. In 29th International Conference on Software Engineering (ICSE’07), pages 75–84. IEEE.
  31. 31.Emilio Parisotto and Ruslan Salakhutdinov. 2017. Neural map: Structured memory for deep reinforcement learning. arXiv preprint arXiv:1702.08360.
  32. 32.Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval augmented code generation and summarization. arXiv preprint arXiv:2108.11601.
  33. 33.Md Rizwan Parvez, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2018. Building language models for text with named entities. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2373–2383, Melbourne, Australia. Association for Computational Linguistics.
  34. 34.Md Rizwan Parvez, Jianfeng Chi, Wasi Uddin Ahmad, Yuan Tian, and Kai-Wei Chang. 2023. Retrieval enhanced data augmentation for question answering on privacy policies. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 201–210, Dubrovnik, Croatia. Association for Computational Linguistics.
  35. 35.Oleksandr Polozov and Sumit Gulwani. 2015. Flashmeta: A framework for inductive program synthesis. In Proceedings of the 2015 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications, pages 107–126.
  36. 36.Maxim Rabinovich, Mitchell Stern, and Dan Klein. 2017. Abstract syntax networks for code generation and semantic parsing. CoRR, abs/1704.07535.
  37. 37.Tal Ridnik, Dedy Kredo, and Itamar Friedman. 2024. Code generation with alphacodium: From prompt engineering to flow engineering. arXiv preprint arXiv:2401.08500.
  38. 38.Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950.
  39. 39.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems.
  40. 40.Kashun Shum, Shizhe Diao, and Tong Zhang. 2023. Automatic prompt augmentation and selection with chain-of-thought from labeled data. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12113–12139, Singapore. Association for Computational Linguistics.
  41. 41.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  42. 42.Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In EMNLP, pages 8696–8708.
  43. 43.Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022a. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  45. 45.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  46. 46.Xiaohan Xu, Chongyang Tao, Tao Shen, Can Xu, Hongbo Xu, Guodong Long, and Jian guang Lou. 2023. Re-reading improves reasoning in language models.
  47. 47.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  48. 48.Michihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat, Jure Leskovec, Percy Liang, Ed H Chi, and Denny Zhou. 2023. Large language models as analogical reasoners. arXiv preprint arXiv:2310.01714.
  49. 49.Pengcheng Yin and Graham Neubig. 2017. A syntactic neural model for general-purpose code generation. CoRR, abs/1704.01696.
  50. 50.Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, and Dragomir Radev. 2019. CoSQL: A conversational text-to-SQL challenge towards cross-domain natural language interfaces to databases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1962–1979, Hong Kong, China. Association for Computational Linguistics.
  51. 51.Yifan Zhang, Jingqin Yang, Yang Yuan, and Andrew Chi-Chih Yao. 2023. Cumulative reasoning with large language models. arXiv preprint arXiv:2308.04371.
  52. 52.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  53. 53.Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2023. Language agent tree search unifies reasoning acting and planning in language models. arXiv preprint arXiv:2310.04406.

Citation

MLA
Islam, M. A., et al. “MapCoder: Multi-Agent Code Generation for Competitive Problem Solving”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 4912–44, https://doi.org/10.18653/v1/2024.acl-long.269.
APA
Islam, M. A., Ali, M. E., & Parvez, M. R. (2024). MapCoder: Multi-Agent Code Generation for Competitive Problem Solving. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4912–4944. https://doi.org/10.18653/v1/2024.acl-long.269
Chicago
Islam, M. A., M. E. Ali, and M. R. Parvez. 2024. “MapCoder: Multi-Agent Code Generation for Competitive Problem Solving”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4912–44. https://doi.org/10.18653/v1/2024.acl-long.269.
Harvard
Islam, M.A., Ali, M.E. and Parvez, M.R. (2024) “MapCoder: Multi-Agent Code Generation for Competitive Problem Solving”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4912–4944. Available at: https://doi.org/10.18653/v1/2024.acl-long.269.
Vancouver
1. Islam MA, Ali ME, Parvez MR (2024) MapCoder: Multi-Agent Code Generation for Competitive Problem Solving. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 4912–4944

BibTeX

@inproceedings{islam-etal-2024-mapcoder,
    title = "{M}ap{C}oder: Multi-Agent Code Generation for Competitive Problem Solving",
    author = "Islam, Md. Ashraful  and
      Ali, Mohammed Eunus  and
      Parvez, Md Rizwan",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.269/",
    doi = "10.18653/v1/2024.acl-long.269",
    pages = "4912--4944"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/