UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging
Cheryl LeeChunqiu Steven XiaLongji YangJen-tse HuangZhouruixin ZhuLingming ZhangMichael R. Lyu
Proposes an end-to-end multi-agent framework that models developer cognitive stages across a three-level hierarchy to adaptively localize and fix software bugs, outperforming existing automated program repair methods on Defects4J without requiring ground-truth fault locations.
Software debugging remains a major bottleneck in software engineering, consuming substantial developer time during continuous integration and continuous delivery workflows. Although large language models show strong coding potential, existing automated program repair tools frequently fail on complex, repository-level bugs because they isolate individual tasks, such as fault localization or patch generation, and rely on inefficient multi-agent communication structures.
The article introduces and evaluates UniDebugger, an end-to-end framework designed to unify the entire debugging process. Rather than treating artificial intelligence agents as conversational human experts, UniDebugger organizes seven specialized agents into a hierarchical, three-level pipeline that mirrors human cognitive debugging models to adaptively address software faults.
The approach systematically escalates problem-solving intensity based on bug complexity. Simple faults are handled at Level 1 with minimal localization and patching agents. If test execution fails, Level 2 activates code slicing, summarization, and patch optimization using static and dynamic analysis tools. Level 3 handles intricate repository-level defects by incorporating cross-file dependency mapping and external search references. The framework was evaluated across four standard benchmarks covering Java, Python, and C, including the real-world dataset Defects4J, and tested across seven underlying model backbones.
The evaluation produced several critical findings. On the real-world Defects4J benchmark, UniDebugger correctly repaired 197 bugs (generating 286 plausible fixes), achieving a 25.48% improvement over the leading baseline without requiring ground-truth fault locations. It successfully repaired 42 complex bugs that none of the top four baseline tools could fix. In competition benchmarks, UniDebugger solved 100% of faults in QuixBugs and generated 2.2 times more plausible patches on Codeflaws than the top traditional baseline, achieving a 95% correctness rate. Furthermore, applying UniDebugger boosted underlying model performance across diverse architectures by 21.60% to 52.31%, helping open-source models approach proprietary model performance.
These results indicate that automated software debugging can be integrated directly into automated testing pipelines to proactively detect and resolve software defects before deployment. By reducing search iterations from thousands to at most 20 attempts per defect, the framework substantially reduces computational and operational costs while minimizing manual debugging overhead.
Organizations evaluating automated repair should consider deploying hierarchical multi-agent frameworks within developer continuous integration pipelines. However, human verification remains necessary because automatically generated patches may introduce subtle vulnerabilities, and the current framework requires explicit failing test cases rather than informal natural language issue reports. While confidence in test-driven automated repair is strong, future work should explore natural language user-issue resolution, token consumption optimization, and broader defect types such as system configuration errors.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). This paper establishes the foundational concept of iterative, feedback-driven program repair using execution signals in LLMs, which UniDebugger formalizes into an escalated multi-agent pipeline.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It provides essential background on agent-computer interfaces and tool scaffolding for repository-level software debugging that informs UniDebugger's integration of static and dynamic analysis tools.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). It defines the benchmark standards and core challenges of resolving real-world GitHub issues and multi-file bugs that UniDebugger directly targets.
- Paper: Repair Is Nearly Generation: Multilingual Program Repair with LLMs, Harshit Joshi et al. (2023). It introduces multilingual automated program repair formulations that UniDebugger builds upon to unify Java, Python, and C debugging under a single multi-agent architecture.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It pioneers communicative multi-agent role-playing frameworks, providing the baseline cooperative paradigm that UniDebugger contrasts against with its hierarchical escalation structure.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). It explores multi-agent critique and debate mechanisms to overcome single-model reasoning pitfalls, motivating UniDebugger's separation of specialized localization, slicing, and patch optimization agents.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). This work directly tackles the high token consumption and context exploration bottlenecks of repository-level coding agents identified as a key future direction in UniDebugger.
- Paper: Scaling Laws for Agent Harnesses via Effective Feedback Compute, Xuanliang Zhang et al. (2026). It formalizes scaling laws and effective feedback compute metrics to optimize inference-time budget allocation across agent harnesses like UniDebugger's multi-level pipeline.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). It provides a broader theoretical survey and architectural harness framework for grounding LLM agents in executable code and multi-agent coordination.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). It investigates how agentic mid-training can natively embed repository navigation and iterative debugging capabilities directly into base models before agent scaffolding.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It extends the evaluation of complex, multi-agent development trajectories by using tool-augmented evaluator agents to assess end-to-end task completion.
- Paper: Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems, Jiacheng Liu et al. (2026). It analyzes the architectural and harness design choices required to deploy autonomous coding agents safely and effectively in production developer environments.
