Interactive Debugging and Steering of Multi-Agent AI Systems
Will EppersonGagan BansalVictor DibiaAdam Fourney (adamfo)Jack GerritsErkang (Eric) ZhuSaleema Amershi
Introduces AGDebugger, an interactive debugging environment that addresses key failure points in multi-agent LLM teams by enabling developers to visually trace message flows, edit agent outputs, and dynamically rewind conversation states.
Multi-agent artificial intelligence systems, in which multiple specialized autonomous agents collaborate through multi-turn conversations and external tool use, are increasingly deployed for complex problem-solving. However, developers lack effective tooling to diagnose and fix errors within these systems. Traditional development workflows require reviewing static, text-heavy logs comprising dozens of messages and thousands of words, making error localization slow and post-hoc troubleshooting cumbersome. When errors occur mid-workflow, developers typically must restart runs from scratch, wasting development time and complicating the isolation of root causes.
The article develops and evaluates AGDebugger, an interactive debugging and steering system for message-passing multi-agent architectures. The objective of the research is to evaluate the primary pain points developers encounter when building multi-agent teams and to demonstrate how interactive message intervention, state checkpointing, and visual history tracking affect developers' ability to diagnose and remediate multi-agent failures.
The researchers conducted formative interviews with five expert developers from a major technology company to extract system design requirements. Based on these insights, they built AGDebugger on top of the AutoGen framework, incorporating message stepping, checkpoint-based rollbacks, inline message editing, and an interactive overview visualization. The team then conducted a two-part evaluation involving 14 participants tasked with debugging a five-agent generalist team on challenging benchmark problems from the GAIA evaluation suite. Part 1 evaluated error identification across six participants using comparative baseline and full systems, while Part 2 observed eight participants attempting to steer failing runs toward correct answers over 30-minute interactive sessions.
The findings highlight three key insights regarding multi-agent debugging workflows. First, interactive message resetting was the highest-rated capability among participants (scoring an average of 4.9 out of 5), with five out of six users in the initial study preferring the interactive tool over static log inspection. Second, participants consistently relied on three primary steering strategies across 24 distinct message edits: adding more detailed and concrete instructions (58% of edits), simplifying instructions to prevent language model overload (21%), and altering the high-level plan or tool strategy entirely (21%). Third, steering proved challenging under tight time constraints; only two out of eight participants in the steering study achieved fully correct task outputs within 30 minutes, primarily because interventions made early in the conversation transcript were significantly more effective than late-stage edits that struggled to overcome long-context bias and non-deterministic model behavior.
These results demonstrate that multi-agent systems require interactive, state-aware debugging environments rather than traditional single-prompt or post-hoc log interfaces. Without rollback and counterfactual editing capabilities, organizations risk slow iteration cycles and higher engineering costs. Moreover, the findings show that effective steering requires understanding the underlying constraints of each agent's tool interface, as agents often fail when given overly complex multi-step instructions in natural language.
Organizations developing multi-agent systems should integrate state-checkpointing and interactive rollback mechanisms into their developer toolchains. Development teams should focus steering efforts on early planning stages and decompose tasks into single-step, highly specific sub-instructions. Before deploying agents into live operational environments, engineering leaders should establish guardrails and pre-execution validation checks, as actions with external side effects (such as web actions or external messaging) cannot be rolled back via checkpoints alone.
Confidence in these findings is supported by consistent qualitative feedback across experienced language model practitioners and developers. However, users should consider key limitations: the evaluation was conducted with a modest sample size of 14 participants across two benchmark tasks, within 30-minute evaluation windows, and focused on internal states where side effects were largely reversible. Further long-term studies are needed to examine how interactive debugging tools perform across broader enterprise domains and continuous production workflows.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). It introduces the foundational AutoGen multi-agent conversational framework that establishes multi-agent LLM systems and highlights the debugging challenges AGDebugger directly addresses.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It introduces role-playing communicative LLM agent architectures whose emergent collaborative behaviors and failure modes motivate interactive steering interfaces.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). It establishes self-debugging mechanisms in LLMs, providing fundamental context on automated error localization and correction in model-driven workflows.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). It provides a comprehensive survey of LLM-based agent architectures, memory, and multi-agent interaction paradigms essential for understanding agentic systems.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It formalizes agent-computer interfaces and execution feedback mechanisms for autonomous engineering agents, laying essential groundwork for interactive agent control.
- Paper: Magentic-UI: Towards Human-in-the-loop Agentic Systems, Hussein Mozannar et al. (2025). It builds upon interactive steering paradigms by developing a full human-in-the-loop interface and execution framework for runtime multi-agent oversight and co-planning.
- Paper: UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging, Cheryl Lee et al. (2025). It extends the concept of multi-agent debugging by deploying specialized hierarchical agents to autonomously localize faults and patch complex software bugs.
- Paper: Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems, Jiacheng Liu et al. (2026). It investigates production-grade agent harnesses and runtime permission controls, extending human-in-the-loop steering principles to real-world software engineering agents.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It generalizes multi-agent trajectory inspection by employing specialized judge agents with interactive tools to evaluate and diagnose complex agent workflows.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It advances agent error diagnosis by introducing structured guardrails that monitor multi-step execution paths and attribute root causes of unsafe agent actions.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). It leverages post-execution trajectory data by structuring multi-agent workflow histories into reusable procedural memory to prevent recurring coordination failures.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). It systematically analyzes failure propagation and coordination trade-offs across multi-agent designs, providing empirical foundations for the interface interventions studied in AGDebugger.
