UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging

Cheryl LeeChunqiu Steven XiaLongji YangJen-tse HuangZhouruixin ZhuLingming ZhangMichael R. Lyu

article2025EMNLP59 citations

Proposes an end-to-end multi-agent framework that models developer cognitive stages across a three-level hierarchy to adaptively localize and fix software bugs, outperforming existing automated program repair methods on Defects4J without requiring ground-truth fault locations.

Listen

Software debugging remains a major bottleneck in software engineering, consuming substantial developer time during continuous integration and continuous delivery workflows. Although large language models show strong coding potential, existing automated program repair tools frequently fail on complex, repository-level bugs because they isolate individual tasks, such as fault localization or patch generation, and rely on inefficient multi-agent communication structures.

The article introduces and evaluates UniDebugger, an end-to-end framework designed to unify the entire debugging process. Rather than treating artificial intelligence agents as conversational human experts, UniDebugger organizes seven specialized agents into a hierarchical, three-level pipeline that mirrors human cognitive debugging models to adaptively address software faults.

The approach systematically escalates problem-solving intensity based on bug complexity. Simple faults are handled at Level 1 with minimal localization and patching agents. If test execution fails, Level 2 activates code slicing, summarization, and patch optimization using static and dynamic analysis tools. Level 3 handles intricate repository-level defects by incorporating cross-file dependency mapping and external search references. The framework was evaluated across four standard benchmarks covering Java, Python, and C, including the real-world dataset Defects4J, and tested across seven underlying model backbones.

The evaluation produced several critical findings. On the real-world Defects4J benchmark, UniDebugger correctly repaired 197 bugs (generating 286 plausible fixes), achieving a 25.48% improvement over the leading baseline without requiring ground-truth fault locations. It successfully repaired 42 complex bugs that none of the top four baseline tools could fix. In competition benchmarks, UniDebugger solved 100% of faults in QuixBugs and generated 2.2 times more plausible patches on Codeflaws than the top traditional baseline, achieving a 95% correctness rate. Furthermore, applying UniDebugger boosted underlying model performance across diverse architectures by 21.60% to 52.31%, helping open-source models approach proprietary model performance.

These results indicate that automated software debugging can be integrated directly into automated testing pipelines to proactively detect and resolve software defects before deployment. By reducing search iterations from thousands to at most 20 attempts per defect, the framework substantially reduces computational and operational costs while minimizing manual debugging overhead.

Organizations evaluating automated repair should consider deploying hierarchical multi-agent frameworks within developer continuous integration pipelines. However, human verification remains necessary because automatically generated patches may introduce subtle vulnerabilities, and the current framework requires explicit failing test cases rather than informal natural language issue reports. While confidence in test-driven automated repair is strong, future work should explore natural language user-issue resolution, token consumption optimization, and broader defect types such as system configuration errors.

arXiv: 2404.17153
Cover for UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging

Abstract

Software debugging is a time-consuming endeavor involving a series of steps, such as fault localization and patch generation, each requiring thorough analysis and a deep understanding of the underlying logic. While large language models (LLMs) demonstrate promising potential in coding tasks, their performance in debugging remains limited. Current LLM-based methods often focus on isolated steps and struggle with complex bugs. In this paper, we propose the first end-to-end framework, UniDebugger, for unified debugging through multi-agent synergy. It mimics the entire cognitive processes of developers, with each agent specialized as a particular component of this process rather than mirroring the actions of an independent expert as in previous multi-agent systems. Agents are coordinated through a three-level design, following a cognitive model of debugging, allowing adaptive handling of bugs with varying complexities. Experiments on extensive benchmarks demonstrate that UniDebugger significantly outperforms state-of-the-art repair methods, fixing 1.25× to 2.56× bugs on the repo-level benchmark, Defects4J. This performance is achieved without requiring ground-truth root-cause code statements, unlike the baselines. Our source code is available on an anonymous link: https://github.com/BEbillionaireUSD/UniDebugger.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 UniDebugger
  • 3.1 Profiles of Agents
  • 3.2 External Interactions
  • 3.3 Hierarchical Coordination
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Comparison with baselines
  • 4.3 Performance on Different LLMs
  • 4.4 Ablation Study
  • 4.4.1 Impact of Different Agents
  • 4.4.2 Impact of External Interactions
  • 5 Conclusion
  • 6 Limitation
  • 7 Ethics Consideration
  • 8 Artifact Discussion
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Unique Fix Example
  • A.2 Alogrithm of UniDebugger
  • A.3 System Prompts
  • A.4 A Demo of the Execution
  • A.4.1 Bug Metadata
  • A.4.2 L3 Repair

Knowls

  1. Knowl 1 — Hierarchical Cognitive Multi-Agent Debugging Architecture

    model/method

    UniDebugger is an end-to-end framework for unified software debugging grounded in cognitive debugging theory (specifically Hale and Haworth's model of structural learning). Rather than employing a horizontal, mesh-like peer-to-peer discussion paradigm where agents act as autonomous human team members (e.g., developers, testers, product managers), UniDebugger structures agents as functional components of a single, coherent, assembly-line cognitive thought process.

    The framework organizes seven specialized agents into a three-level hierarchical workflow that adaptively scales cognitive intensity and tool invocation based on bug complexity:

    1. Level 1 (L1 - Direct Repair): Targets simple, local logic errors. Only the Locator and Fixer agents are initialized. Locator identifies suspicious lines, and Fixer generates a patch.
    2. Level 2 (L2 - Single-File Semantic Repair): Triggered if L1 fails to produce a plausible patch. Operates under the assumption that the fix requires deeper single-file understanding. Slicer extracts a concise suspicious code snippet (20--100 lines) and Summarizer generates a functional summary of the buggy file. Then Locator, Fixer, and FixerPro execute sequentially with access to these summaries, the sliced snippet, and dynamic execution feedback.
    3. Level 3 (L3 - Multi-File and External Retrieval Repair): Triggered if L2 fails, assuming the bug involves complex logic, external libraries, or cross-file dependencies. Helper retrieves external web or local reference solutions via Retrieval-Augmented Generation (RAG), and RepoFocus identifies 2--6 related files across the project directory. Summarizer creates structured summaries for every identified file, and Slicer, Locator, Fixer, and FixerPro proceed sequentially using the shared multi-file summaries and external reference guide.
  2. Knowl 2 — Hierarchical Multi-Level Program Debugging Algorithm

    algorithm

    The UniDebugger algorithm coordinates multi-agent debugging across three escalation levels (L1L1, L2L2, L3L3) with plausibility feedback and reverse-order reflection.

    Input: k: maximum debugging attempts; m: maximum re-sampling attempts per agent; bug_meta: bug metadata containing source code, failing test cases, error messages, and natural language requirements
    Output: patch: final source code patch; analysis: explanation and diagnostic report
    Function L1Repair(m, bug_meta, extra_info):
        for j = 1 to m do
            marked_code = Locator(bug_meta, extra_info)
            if ValidMarks(marked_code) then
                bug_meta[code] = marked_code
                break
        patch = Fixer(bug_meta, extra_info)
        return patch
    Function L2Repair(m, bug_meta, extra_info):
        if summary not in extra_info then
            summary = Summarizer(bug_meta[code])
            extra_info = concat[extra_info, summary]
        for j = 1 to m do
            snippet = Slicer(bug_meta)
            if ValidSnippet(snippet) then
                bug_meta[code] = snippet
                break
        patch = L1Repair(m, bug_meta, extra_info)
        patch, analysis = FixerPro(patch, Testing(patch), bug_meta, extra_info)
        return patch, analysis
    Function L3Repair(m, bug_meta, extra_info):
        references = Helper(bug_meta)
        FileList = RepoFocus(bug_meta)
        summary = []
        for file in FileList do
            summary.append(Summarizer(ReadFile(file)))
        extra_info = concat[extra_info, references, summary]
        return L2Repair(m, bug_meta, extra_info)
    Function Debugging(k, m, bug_meta):
        for i = 1 to k do
            extra_info = []
            patch = L1Repair(m, bug_meta, extra_info)
            if Testing(patch) is Plausible then
                return patch, ""
            patch, analysis = L2Repair(m, bug_meta, extra_info)
            if Testing(patch) is Plausible then
                return patch, analysis
            patch, analysis = L3Repair(m, bug_meta, extra_info)
            if Testing(patch) is Plausible then
                return patch, analysis
            patch, analysis = RefineAgents(m, bug_meta, extra_info, patch)
            if Testing(patch) is Plausible then
                return patch, analysis
        return patch, analysis

    When a patch fails on Level 3, UniDebugger executes RefineAgents by reflecting in reverse order: FixerPro is first prompted to diagnose the failure and refine its fix; if that fails, Fixer is prompted to modify its patch, propagating updated inputs back to FixerPro.

  3. Knowl 3 — Specialized Functional Agent Profiles in UniDebugger

    model/method

    UniDebugger employs seven functional agents, each defined by a one-shot system prompt specifying role, skills, actions, objectives, output constraints, and a rubber-duck debugging requirement (tracking key program variables and explaining thought processes):

    1. Helper: Generates a concise search query (le100\\le 100 words) from bug metadata and queries a search engine (e.g., Tavily or a local patch database) to retrieve similar bug reports and documentation, synthesizing them into a structured reference guide with supporting URLs.
    2. RepoFocus: Analyzes project folder structure, package declarations, and imported dependencies alongside error traces to identify 22 to 66 suspected bug-related files across the repository.
    3. Summarizer: Produces compact natural language and signature summaries for classes and functions (formatted as <class>~~~<function>~~~<parameters>~~~<return_type>~~~<description>) to fit repository-level context into LLM context windows without semantic loss.
    4. Slicer: Identifies and extracts a contiguous suspicious window of 2020--100100 code lines from the buggy file using contextual string boundaries to avoid hallucinated code modifications.
    5. Locator: Performs fine-grained fault localization by annotating the exact lines with // buggy line or // missing code: [INFILLED CODE] comments via contextual string matching on the original source code.
    6. Fixer: Generates candidate patches in git diff format, evaluating and optionally correcting the annotations provided by Locator to ensure consistency.
    7. FixerPro: Performs automated code review on patches. If Fixer generated a plausible fix, FixerPro refactors it for simplicity and maintainability; if Fixer failed, FixerPro analyzes failing test execution traces and produces a corrected, optimized patch.
  4. Knowl 4 — Invariant-Context Patching and External Tool Feedback Mechanism

    model/method

    UniDebugger integrates external software engineering environments and tools through non-destructive patching and automated test execution feedback:

    1. Invariant-Context Patch Application: Rather than allowing LLMs to directly rewrite source files or relying on absolute line indices (which LLMs frequently miscount), Fixer and FixerPro output standard git diff unified textual diffs. UniDebugger applies patches using rule-based contextual matching of invariant code surrounding the edit sites, preventing hallucinations and syntax corruption.
    2. Plausibility Feedback Loop: Once patched, the modified program is compiled and tested against the suite of failing and passing test cases. A patch is classified as plausible if it compiles and passes all available test cases. Test execution outcomes and stack traces are piped back to the agents to trigger higher hierarchy levels or reverse-order agent refinement.
    3. Toolbox Integration: Agents can query a dedicated toolbox consisting of:
      • Static analysis tools (SonarQube) for compiler warnings, syntax errors, and Abstract Syntax Trees (ASTs).
      • Dynamic analysis tools (GZoltar) providing execution coverage matrices of failing tests, filtering out unexecuted statements.
      • Search engine APIs (Tavily or local indexed patch corpora) providing retrieval-augmented reference solutions.
  5. Knowl 5 — Program Repair Performance Across Multi-Language Benchmarks

    data/table

    UniDebugger was evaluated across four program repair benchmarks encompassing three programming languages (Java, C, Python): Defects4J (v1.2 and v2.0, 806 bugs), Codeflaws (3,902 C bugs), QuixBugs (40 Java and 40 Python bugs), and ConDefects (1,254 Java and 1,625 Python bugs). UniDebugger uses gpt-4o as its default backbone with up to 20 attempts for the full framework and 5 attempts for UniDebugger-Lite.

    Tools Sampling Defects4J-Java Codeflaws-C QuixBugs-Java QuixBugs-Python
    Times #Corr #Plau #Corr #Plau #Corr #Plau #Corr #Plau
    Angelix 1,000 - - 318 591 - - - -
    Prophet 1,000 - - 310 839 - - - -
    SPR 1,000 - - 283 783 - - - -
    CVC4 - - - 15 91 - - - -
    Semfix 1,000 25 68 38 56 - - - -
    Recoder 100 72 140 - - 17 17 - -
    GenProg 1,000 5 20 255–369 1,423 1 4 - -
    CoCoNuT 20,000 44 85 423 716 13 20 19 21
    CURE - 57 104 - - 26 35 - -
    RewardRepair 200 90 75 - - 20 - - -
    Tbar 500 77 121 - - - - - -
    AlphaRepair 5,000 110 159 - - 28 30 27 32
    Repilot 5,000 116 - - - - - - -
    ChatRepair 100–200 157 - - - 40 40 40 40
    CodeLlama-34b 20 24 41 91% 1,488 25 28 33 33
    LLaMA2-70b 20 39 78 91% 1,576 25 28 33 33
    DeepSeekCoderV2 20 57 82 93% 1,937 30 34 25 38
    gemini-1.5-flash 20 19 36 86% 1,291 29 32 29 35
    gpt-3.5-turbo-ca 20 45 71 94% 2,343 33 34 34 36
    claude-3.5-sonnet 20 70 116 95% 2,624 36 37 40 40
    gpt-4o 20 72 119 93% 2,549 35 36 39 39
    UniDebugger-Lite 5 - - 95% 3,130 40 40 40 40
    UniDebugger 20 197 286 - - - - - -

    On Defects4J, UniDebugger achieves 197 correctly fixed bugs (#Corr) and 286 plausible fixes (#Plau) without requiring ground-truth fault locations, outperforming ChatRepair (157 correct fixes) by 25.48% while using significantly fewer sampling attempts (20 vs 100--200). UniDebugger correctly resolves 42 unique bugs that none of the top four baselines (ChatRepair, AlphaRepair, Tbar, Recoder) addressed.

  6. Knowl 6 — Ablation of Agent Components and Hierarchical Levels on Defects4J

    data/table

    An ablation study on Defects4J measures the cumulative contribution of each agent and hierarchy level to the number of plausibly fixed bugs (#Plau) and the average monetary cost in USD per bug.

    Helper RepoFocus Summarizer Slicer Locator Fixer FixerPro #Plau Expense ($)
    ✓ 72 0.030
    ✓ ✓ 140 0.048
    ✓ ✓ ✓ 192 0.116
    ✓ ✓ ✓ ✓ 224 0.225
    ✓ ✓ ✓ ✓ ✓ 238 0.317
    ✓ ✓ ✓ ✓ ✓ ✓ 245 0.364
    ✓ ✓ ✓ ✓ ✓ ✓ ✓ 291 0.410

    Key observations:

    1. Level 1 (Locator + Fixer) doubles plausible fixes from 72 (Fixer alone) to 140 at minimal additional cost (0.048vs0.048 vs 0.030).
    2. Level 2 (adding FixerPro, Slicer, and Summarizer) increases plausible fixes from 140 to 238 (+70.0%).
    3. Level 3 (adding RepoFocus and Helper) reaches 291 plausible fixes. Helper alone contributes an increase of 46 plausible fixes, demonstrating the utility of external RAG debugging guidance.
  7. Knowl 7 — Performance Enhancement of UniDebugger-Lite Across Diverse LLM Backbones

    data/table

    UniDebugger-Lite (UD-L) was evaluated across seven distinct LLM backbones on a sampled subset of 600 programs from ConDefects (300 Java, 300 Python), comparing UD-L against standard Chain-of-Thought (CoT) prompting on the same base models.

    LLM Backbone ConDefects-Java ConDefects-Python
    CoT UD-L Gain (%) CoT UD-L Gain (%)
    CodeLlama-34b 87 113 +29.89% 69 86 +24.64%
    LLaMA2-70b 108 147 +36.11% 91 133 +46.15%
    DeepSeekCoderV2 130 198 +52.31% 125 178 +42.40%
    gemini-1.5-flash 62 89 +43.55% 63 82 +30.16%
    gpt-3.5-turbo-ca 155 191 +23.23% 127 174 +37.01%
    claude-3.5-sonnet 213 259 +21.60% 186 227 +22.04%
    gpt-4o 211 262 +24.17% 179 225 +25.70%

    UniDebugger consistently improves repair success over base CoT prompting across all evaluated architectures, with relative gains ranging between 21.60% and 52.31%. DeepSeekCoderV2 paired with UD-L produces 198 Java and 178 Python plausible fixes (376 total), rivaling proprietary models like gpt-4o with CoT (390 total).

  8. Knowl 8 — Ablation of External Environment and Tool Interactions

    data/table

    An ablation study on Defects4J isolates the impact of four external interaction components: Online Search (Helper search API), Static Analysis (SonarQube AST/warnings), Dynamic Analysis (GZoltar failing test coverage), and Testing Feedback (plausibility results fed back to FixerPro and upstream agents).

    Online Search Static Analysis Dynamic Analysis Testing Feedback #Plau
    ✓ ✓ ✓ 245
    ✓ ✓ ✓ 268
    ✓ ✓ ✓ 170
    ✓ ✓ ✓ 244
    ✓ ✓ ✓ ✓ 291

    Dynamic analysis coverage is the most critical individual external interaction; omitting it causes the largest performance drop, reducing plausible fixes from 291 to 170 (-41.58%). Omitting testing feedback reduces plausible fixes to 244, omitting online search reduces plausible fixes to 245, and omitting static analysis reduces plausible fixes to 268.

  9. Knowl 9 — UniDebugger Scope and Test-Driven Debugging Assumptions

    limitation

    UniDebugger is specifically designed for developer-oriented, test-driven Automated Program Repair (APR) within continuous integration/continuous deployment (CI/CD) pipelines. This operational design introduces several specific limitations:

    1. Dependency on Failing Test Oracles: The framework requires executable test suites with explicit failing tests, error messages, and stack traces to guide fault localization, prompt synthesis, and iterative plausibility validation. It cannot directly operate on issue-driven bug reports (e.g., natural-language problem descriptions in SWE-bench) that lack formal, executable reproduction test cases.
    2. Non-Code Fault Categories: The cognitive steps and patching mechanisms focus on algorithmic, logical, and structural source code errors, leaving its effectiveness on broader fault domains (such as configuration errors, build script failures, or dependency conflicts) unverified.
    3. Token and Runtime Overhead: The hierarchical escalation through L2 and L3 incurs progressive increases in token consumption and execution latency when invoking external RAG search queries, static/dynamic analyzers, and multi-file summaries.

Coverage note — None was omitted. All primary architectural contributions, multi-agent workflows, algorithms, benchmark evaluations, ablation studies across agents and tools, and stated limitations are fully represented.

References

  1. 1.Rui Abreu, Peter Zoeteweij, and Arjan J. C. van Gemund. 2006. An evaluation of similarity coefficients for software fault localization. In 12th IEEE Pacific Rim International Symposium on Dependable Computing (PRDC 2006), 18-20 December, 2006, University of California, Riverside, USA, pages 39–46. IEEE Computer Society.
  2. 2.Rui Abreu, Peter Zoeteweij, and Arjan J. C. van Gemund. 2009. Spectrum-based multiple fault localization. In ASE 2009, 24th IEEE/ACM International Conference on Automated Software Engineering, Auckland, New Zealand, November 16-20, 2009, pages 88–99. IEEE Computer Society.
  3. 3.Hunt Andrew and Thomas David. 2000. The pragmatic programmer: From journeyman to master.
  4. 4.Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul Ronald Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, George Tucker, Enrique Piqueras, Maxim Krikun, Iain Barr, Nikolay Savinov, Ivo Danihelka, Becca Roelofs, Anaïs White, Anders Andreassen, Tamara von Glehn, Lakshman Yagati, Mehran Kazemi, Lucas Gonzalez, Misha Khalman, Jakub Sygnowski, and et al. 2023. Gemini: A family of highly capable multimodal models. CoRR, abs/2312.11805.
  5. 5.Anthropic. 2024. Introducing the next generation of claude.
  6. 6.Samuel Benton, Xia Li, Yiling Lou, and Lingming Zhang. 2020. On the effectiveness of unified debugging: An extensive study on 16 program repair systems. In 35th IEEE/ACM International Conference on Automated Software Engineering, ASE 2020, Melbourne, Australia, September 21-25, 2020, pages 907–918. IEEE.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. CoRR, abs/2107.03374.
  8. 8.Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2023. Self-collaboration code generation via chatgpt. CoRR, abs/2304.07590.
  9. 9.Claire Le Goues, Michael Dewey-Vogt, Stephanie Forrest, and Westley Weimer. 2012. A systematic study of automated program repair: Fixing 55 out of 105 bugs for $8 each. In 34th International Conference on Software Engineering, ICSE 2012, June 2-9, 2012, Zurich, Switzerland, pages 3–13. IEEE Computer Society.
  10. 10.Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. Deepseek-coder: When the large language model meets programming - the rise of code intelligence. CoRR, abs/2401.14196.
  11. 11.GZoltar. [link].
  12. 12.David P. Hale and Dwight A. Haworth. 1991. Towards a model of programmers’ cognitive processes in software maintenance: A structural learning theory approach for debugging. J. Softw. Maintenance Res. Pract., 3(2):85–106.
  13. 13.Joanne E. Hale, Shane Sharpe, and David P. Hale. 1999. An evaluation of the cognitive processes of programmers engaged in software debugging. J. Softw. Maintenance Res. Pract., 11(2):73–91.
  14. 14.Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. Metagpt: Meta programming for A multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
  15. 15.Dong Huang, Jie M. Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2024. Agentcoder: Multi-agent-based code generation with iterative testing and optimisation. Preprint, arXiv:2312.13010.
  16. 16.Kai Huang, Xiangxin Meng, Jian Zhang, Yang Liu, Wenjie Wang, Shuhao Li, and Yuqing Zhang. 2023. An empirical study on fine-tuning large language models of code for automated program repair. In 38th IEEE/ACM International Conference on Automated Software Engineering, ASE 2023, Luxembourg, September 11-15, 2023, pages 1162–1174. IEEE.
  17. 17.Md. Ashraful Islam, Mohammed Eunus Ali, and Md. Rizwan Parvez. 2024. Mapcoder: Multi-agent code generation for competitive problem solving. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 4912–4944. Association for Computational Linguistics.
  18. 18.Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. CURE: code-aware neural machine translation for automatic program repair. In 43rd IEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid, Spain, 22-30 May 2021, pages 1161–1173. IEEE.
  19. 19.Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. 2024. Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
  20. 20.René Just, Darioush Jalali, and Michael D. Ernst. 2014. Defects4j: a database of existing faults to enable controlled testing studies for java programs. In International Symposium on Software Testing and Analysis, ISSTA ’14, San Jose, CA, USA - July 21 - 26, 2014, pages 437–440. ACM.
  21. 21.Pavneet Singh Kochhar, Xin Xia, David Lo, and Shanping Li. 2016. Practitioners’ expectations on automated fault localization. In Proceedings of the 25th International Symposium on Software Testing and Analysis, ISSTA 2016, Saarbrücken, Germany, July 18-20, 2016, pages 165–176. ACM.
  22. 22.Xuan-Bach Dinh Le, Ferdian Thung, David Lo, and Claire Le Goues. 2018. Overfitting in semantics-based automated program repair. In Proceedings of the 40th International Conference on Software Engineering, ICSE 2018, Gothenburg, Sweden, May 27 - June 03, 2018, page 163. ACM.
  23. 23.Xia Li, Wei Li, Yuqun Zhang, and Lingming Zhang. 2019. Deepfl: integrating multiple fault diagnosis dimensions for deep fault localization. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2019, Beijing, China, July 15-19, 2019, pages 169–180. ACM.
  24. 24.Xia Li and Lingming Zhang. 2017. Transforming programs and tests in tandem for fault localization. Proc. ACM Program. Lang., 1(OOPSLA):92:1–92:30.
  25. 25.Yi Li, Shaohua Wang, and Tien N. Nguyen. 2021. Fault localization with code coverage representation learning. In 43rd IEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid, Spain, 22-30 May 2021, pages 661–673. IEEE.
  26. 26.Yi Li, Shaohua Wang, and Tien N. Nguyen. 2022. Fault localization to detect co-change fixing locations. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE 2022, Singapore, Singapore, November 14-18, 2022, pages 659–671. ACM.
  27. 27.Derrick Lin, James Koppel, Angela Chen, and Armando Solar-Lezama. 2017. Quixbugs: a multi-lingual program repair benchmark set based on the quixey challenge. In Proceedings Companion of the 2017 ACM SIGPLAN International Conference on Systems, Programming, Languages, and Applications: Software for Humanity, SPLASH 2017, Vancouver, BC, Canada, October 23 - 27, 2017, pages 55–56. ACM.
  28. 28.Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F. Bissyandé. 2019. Tbar: revisiting template-based automated program repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2019, Beijing, China, July 15-19, 2019, pages 31–42. ACM.
  29. 29.Fan Long. 2018. Automatic patch generation via learning from successful human patches. Ph.D. thesis, Massachusetts Institute of Technology, Cambridge, USA.
  30. 30.Fan Long and Martin C. Rinard. 2015. Staged program repair with condition synthesis. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC/FSE 2015, Bergamo, Italy, August 30 - September 4, 2015, pages 166–178. ACM.
  31. 31.Yiling Lou, Qihao Zhu, Jinhao Dong, Xia Li, Zeyu Sun, Dan Hao, Lu Zhang, and Lingming Zhang. 2021. Boosting coverage-based fault localization via graph-based representation learning. In ESEC/FSE ’21: 29th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Athens, Greece, August 23-28, 2021, pages 664–676. ACM.
  32. 32.Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li, Moshi Wei, and Lin Tan. 2020. Coconut: combining context-aware neural translation models using ensemble for program repair. In ISSTA ’20: 29th ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, USA, July 18-22, 2020, pages 101–114. ACM.
  33. 33.Sergey Mechtaev, Jooyong Yi, and Abhik Roychoudhury. 2016. Angelix: scalable multiline program patch synthesis via symbolic analysis. In Proceedings of the 38th International Conference on Software Engineering, ICSE 2016, Austin, TX, USA, May 14-22, 2016, pages 691–701. ACM.
  34. 34.Seokhyeon Moon, Yunho Kim, Moonzoo Kim, and Shin Yoo. 2014. Ask the mutants: Mutating faulty programs for fault localization. In Seventh IEEE International Conference on Software Testing, Verification and Validation, ICST 2014, March 31 2014-April 4, 2014, Cleveland, Ohio, USA, pages 153–162. IEEE Computer Society.
  35. 35.Manish Motwani, Mauricio Soto, Yuriy Brun, René Just, and Claire Le Goues. 2022. Quality of automated program repair on real-world defects. IEEE Trans. Software Eng., 48(2):637–661.
  36. 36.Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chandra. 2013. Semfix: program repair via semantic analysis. In 35th International Conference on Software Engineering, ICSE ’13, San Francisco, CA, USA, May 18-26, 2013, pages 772–781. IEEE Computer Society.
  37. 37.OpenAI. 2023a. Gpt-3.5 turbo.
  38. 38.OpenAI. 2023b. Gpt-4: a technical report. CoRR, abs/2303.08774.
  39. 39.Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. Chatdev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 15174–15186. Association for Computational Linguistics.
  40. 40.Andrew Reynolds, Morgan Deters, Viktor Kuncak, Cesare Tinelli, and Clark Barrett. 2015. Counterexample-guided quantifier instantiation for synthesis in smt. In Computer Aided Verification: 27th International Conference, CAV 2015, San Francisco, CA, USA, July 18-24, 2015, Proceedings, Part II 27, pages 198–216. Springer.
  41. 41.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code. CoRR, abs/2308.12950.
  42. 42.SonarQube. [link].
  43. 43.Ezekiel O. Soremekun, Lukas Kirschner, Marcel Böhme, and Mike Papadakis. 2023. Evaluating the impact of experimental assumptions in automated fault localization. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 159–171. IEEE.
  44. 44.Shin Hwei Tan, Jooyong Yi, Yulis, Sergey Mechtaev, and Abhik Roychoudhury. 2017. Codeflaws: a programming competition benchmark for evaluating automated program repair tools. In Proceedings of the 39th International Conference on Software Engineering, ICSE 2017, Buenos Aires, Argentina, May 20-28, 2017 - Companion Volume, pages 180–182. IEEE Computer Society.
  45. 45.Tavily. [link].
  46. 46.Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Yinxu Pan, Yesai Wu, Zhiyuan Liu, and Maosong Sun. 2024. Debugbench: Evaluating debugging capability of large language models. CoRR, abs/2401.04621.
  47. 47.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  48. 48.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  49. 49.Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang. 2023. Copiloting the copilots: Fusing large language models with completion engines for automated program repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE 2023, San Francisco, CA, USA, December 3-9, 2023, pages 172–184. ACM.
  50. 50.Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009a. Automatically finding patches using genetic programming. In 31st International Conference on Software Engineering, ICSE 2009, May 16-24, 2009, Vancouver, Canada, Proceedings, pages 364–374. IEEE.
  51. 51.Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009b. Automatically finding patches using genetic programming. In 31st International Conference on Software Engineering, ICSE 2009, May 16-24, 2009, Vancouver, Canada, Proceedings, pages 364–374. IEEE.
  52. 52.Yonghao Wu, Zheng Li, Jie M. Zhang, and Yong Liu. 2024. Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, Porto de Galinhas, Brazil, July 15-19, 2024, pages 642–646. ACM.
  53. 53.Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2024. Agentless: Demystifying llm-based software engineering agents. CoRR, abs/2407.01489.
  54. 54.Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023. Automated program repair in the era of large pre-trained language models. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023, pages 1482–1494. IEEE.
  55. 55.Chunqiu Steven Xia and Lingming Zhang. 2022. Less training, more repairing please: revisiting automated program repair via zero-shot learning. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE 2022, Singapore, Singapore, November 14-18, 2022, pages 959–971. ACM.
  56. 56.Chunqiu Steven Xia and Lingming Zhang. 2024. Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024, page 819–831, New York, NY, USA. Association for Computing Machinery.
  57. 57.Xin Xia, Lingfeng Bao, David Lo, Pavneet Singh Kochhar, Ahmed E. Hassan, and Zhenchang Xing. 2017. What do developers search for on the web? Empir. Softw. Eng., 22(6):3149–3185.
  58. 58.Aidan Z. H. Yang, Claire Le Goues, Ruben Martins, and Vincent J. Hellendoorn. 2024. Large language models for test-free fault localization. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal, April 14-20, 2024, pages 17:1–17:12. ACM.
  59. 59.He Ye, Matias Martinez, Thomas Durieux, and Martin Monperrus. 2021. A comprehensive study of automatic program repair on the quixbugs benchmark. J. Syst. Softw., 171:110825.
  60. 60.He Ye, Matias Martinez, and Martin Monperrus. 2022. Neural program repair with execution-based backpropagation. In 44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022, pages 1506–1518. ACM.
  61. 61.Lingming Zhang, Miryung Kim, and Sarfraz Khurshid. 2011. Localizing failure-inducing program edits based on spectrum information. In IEEE 27th International Conference on Software Maintenance, ICSM 2011, Williamsburg, VA, USA, September 25-30, 2011, pages 23–32. IEEE Computer Society.
  62. 62.Lingming Zhang, Lu Zhang, and Sarfraz Khurshid. 2013. Injecting mechanical faults to localize developer faults for evolving software. In Proceedings of the 2013 ACM SIGPLAN International Conference on Object Oriented Programming Systems Languages & Applications, OOPSLA 2013, part of SPLASH 2013, Indianapolis, IN, USA, October 26-31, 2013, pages 765–784. ACM.
  63. 63.Qihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang, Kang Yuan, Yingfei Xiong, and Lu Zhang. 2021. A syntax-guided edit decoder for neural program repair. In ESEC/FSE ’21: 29th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Athens, Greece, August 23-28, 2021, pages 341–353. ACM.

Citation

MLA
Lee, C., et al. “UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 18237–66, https://doi.org/10.18653/v1/2025.emnlp-main.921.
APA
Lee, C., Xia, C. S., Yang, L., Huang, J.-. tse ., Zhu, Z., Zhang, L., & Lyu, M. R. (2025). UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 18237–18266. https://doi.org/10.18653/v1/2025.emnlp-main.921
Chicago
Lee, C., C. S. Xia, L. Yang, et al. 2025. “UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 18237–66. https://doi.org/10.18653/v1/2025.emnlp-main.921.
Harvard
Lee, C. et al. (2025) “UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 18237–18266. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.921.
Vancouver
1. Lee C, Xia CS, Yang L, Huang J-tse, Zhu Z, Zhang L, Lyu MR (2025) UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 18237–18266

BibTeX

@inproceedings{lee-etal-2025-unidebugger,
    title = "{U}ni{D}ebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging",
    author = "Lee, Cheryl  and
      Xia, Chunqiu Steven  and
      Yang, Longji  and
      Huang, Jen-tse  and
      Zhu, Zhouruixing  and
      Zhang, Lingming  and
      Lyu, Michael R.",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.921/",
    doi = "10.18653/v1/2025.emnlp-main.921",
    pages = "18237--18266",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/