Interactive Debugging and Steering of Multi-Agent AI Systems

Will EppersonGagan BansalVictor DibiaAdam Fourney (adamfo)Jack GerritsErkang (Eric) ZhuSaleema Amershi

article2025International Conference on Human Factors in Computing Systems106 citations

Introduces AGDebugger, an interactive debugging environment that addresses key failure points in multi-agent LLM teams by enabling developers to visually trace message flows, edit agent outputs, and dynamically rewind conversation states.

Listen

Multi-agent artificial intelligence systems, in which multiple specialized autonomous agents collaborate through multi-turn conversations and external tool use, are increasingly deployed for complex problem-solving. However, developers lack effective tooling to diagnose and fix errors within these systems. Traditional development workflows require reviewing static, text-heavy logs comprising dozens of messages and thousands of words, making error localization slow and post-hoc troubleshooting cumbersome. When errors occur mid-workflow, developers typically must restart runs from scratch, wasting development time and complicating the isolation of root causes.

The article develops and evaluates AGDebugger, an interactive debugging and steering system for message-passing multi-agent architectures. The objective of the research is to evaluate the primary pain points developers encounter when building multi-agent teams and to demonstrate how interactive message intervention, state checkpointing, and visual history tracking affect developers' ability to diagnose and remediate multi-agent failures.

The researchers conducted formative interviews with five expert developers from a major technology company to extract system design requirements. Based on these insights, they built AGDebugger on top of the AutoGen framework, incorporating message stepping, checkpoint-based rollbacks, inline message editing, and an interactive overview visualization. The team then conducted a two-part evaluation involving 14 participants tasked with debugging a five-agent generalist team on challenging benchmark problems from the GAIA evaluation suite. Part 1 evaluated error identification across six participants using comparative baseline and full systems, while Part 2 observed eight participants attempting to steer failing runs toward correct answers over 30-minute interactive sessions.

The findings highlight three key insights regarding multi-agent debugging workflows. First, interactive message resetting was the highest-rated capability among participants (scoring an average of 4.9 out of 5), with five out of six users in the initial study preferring the interactive tool over static log inspection. Second, participants consistently relied on three primary steering strategies across 24 distinct message edits: adding more detailed and concrete instructions (58% of edits), simplifying instructions to prevent language model overload (21%), and altering the high-level plan or tool strategy entirely (21%). Third, steering proved challenging under tight time constraints; only two out of eight participants in the steering study achieved fully correct task outputs within 30 minutes, primarily because interventions made early in the conversation transcript were significantly more effective than late-stage edits that struggled to overcome long-context bias and non-deterministic model behavior.

These results demonstrate that multi-agent systems require interactive, state-aware debugging environments rather than traditional single-prompt or post-hoc log interfaces. Without rollback and counterfactual editing capabilities, organizations risk slow iteration cycles and higher engineering costs. Moreover, the findings show that effective steering requires understanding the underlying constraints of each agent's tool interface, as agents often fail when given overly complex multi-step instructions in natural language.

Organizations developing multi-agent systems should integrate state-checkpointing and interactive rollback mechanisms into their developer toolchains. Development teams should focus steering efforts on early planning stages and decompose tasks into single-step, highly specific sub-instructions. Before deploying agents into live operational environments, engineering leaders should establish guardrails and pre-execution validation checks, as actions with external side effects (such as web actions or external messaging) cannot be rolled back via checkpoints alone.

Confidence in these findings is supported by consistent qualitative feedback across experienced language model practitioners and developers. However, users should consider key limitations: the evaluation was conducted with a modest sample size of 14 participants across two benchmark tasks, within 30-minute evaluation windows, and focused on internal states where side effects were largely reversible. Further long-term studies are needed to examine how interactive debugging tools perform across broader enterprise domains and continuous production workflows.

Cover for Interactive Debugging and Steering of Multi-Agent AI Systems

Abstract

Fully autonomous teams of LLM-powered AI agents are emerging that collaborate to perform complex tasks for users. What challenges do developers face when trying to build and debug these AI agent teams? In formative interviews with five AI agent developers, we identify core challenges: difficulty reviewing long agent conversations to localize errors, lack of support in current tools for interactive debugging, and the need for tool support to iterate on agent configuration. Based on these needs, we developed an interactive multi-agent debugging tool, AGDebugger, with a UI for browsing and sending messages, the ability to edit and reset prior agent messages, and an overview visualization for navigating complex message histories. In a two-part user study with 14 participants, we identify common user strategies for steering agents and highlight the importance of interactive message resets for debugging. Our studies deepen understanding of interfaces for debugging increasingly important agentic workflows.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multi-Agent AI Systems
  • 2.2 LLM and Agent Debugging
  • 3 Background: Agent Framework and Tasks
  • 3.1 Agent Implementation Framework
  • 3.2 GAIA Benchmark Tasks
  • 3.3 Agent Team for GAIA Tasks
  • 4 Formative Interviews on Agent Debugging
  • 4.1 Understanding Long Agent Conversations is Cumbersome
  • 4.2 Lack of Support for Interactive Debugging
  • 4.3 Iterating on Agent Configuration
  • 4.4 Design Goals
  • 5 AGDebugger: Interactive Agent Debugging
  • 5.1 Message Sending and History
  • 5.2 Message Resetting and Edits
  • 5.2.1 Technical details: checkpoints and sessions
  • 5.3 Conversation Overview Visualization
  • 5.4 Agent Configuration
  • 6 User Study
  • 6.1 Study Design
  • 6.2 Part 1 Findings: Error Identification
  • 6.3 Part 2 Findings: AGDebugger Facilitates Interactive Steering and Debugging
  • 6.4 Three User Approaches to Steer Agents
  • 6.5 User Study Limitations
  • 7 Discussion
  • 7.1 Open Challenges for AI Agent Steering
  • 7.2 Future Directions for Multi-Agent Debugging
  • 8 Conclusion
  • References
  • A Formative interview details
  • B User study part 2 questions

Knowls

  1. Knowl 1 — AGDebugger Architecture for Interactive Multi-Agent Debugging

    model/method

    AGDebugger is an interactive debugging framework for multi-agent Large Language Model (LLM) systems, designed to support error identification, counterfactual testing, and steering. The system operates around three primary components:

    1. Message Queue and Step Execution: Inter-agent communications are mediated via typed messages routed through a central shared runtime queue. Developers can pause execution, step through the message queue one message at a time (similar to line-by-line debugging in standard software debuggers), or inject new broadcast or direct messages to specific agents mid-workflow.
    2. State Checkpointing and In-Place Message Editing: Before any message is processed, AGDebugger checkpoints the internal state of all agents. Users can select any historical message, modify its content inline, and trigger a reset to resume execution from that point.
    3. Visual Conversation Overview: A unit-based commit-graph visualization displays linear and branched multi-agent conversation histories, color-coded by message type, sender, or receiver, with clickable links back to the detailed message stream.

    Agents also implement load_config and save_config routines, allowing developers to inspect and adjust agent parameters (such as system prompts, model names, or sampling temperatures) during a debugging session.

  2. Knowl 2 — State Checkpointing, Good-Enough Policy, and Session Forking in AGDebugger

    model/method

    To allow rewinding multi-agent conversations without re-running workflows entirely from scratch, AGDebugger implements state checkpointing and conversation branching:

    1. State Serialization: Every agent implements save_state() and load_state() interfaces. Stateless agents (such as Python code execution environments) have trivial state, whereas stateful agents (such as browser-based Web Surfers) serialize attributes including the active URL and viewport scroll position.
    2. Good-Enough Checkpoint Policy: Recognizing that full environment replication (e.g., remote server states, external web JavaScript contexts) is often intractable, AGDebugger restores agent-level state to an approximate point (e.g., reloading the recorded URL and scrolling to the saved viewport) and relies on the LLM agent's natural re-evaluation of its observed context upon resumption.
    3. Session Forking: When a user edits a historical message or triggers a rollback, AGDebugger retrieves the snapshot corresponding to that timestamp, restores the internal states of each agent, and initializes a new branched session. Checkpoints and messages preceding the reset point are shared across sessions, while subsequent steps are written exclusively to the newly forked session.
  3. Knowl 3 — Visual Encoding of Multi-Agent Conversation Branches in AGDebugger

    model/method

    AGDebugger utilizes an overview visualization inspired by code commit graphs and unit visualizations to summarize multi-turn agent interactions and counterfactual branches:

    • Unit Layout: Each message in a run is depicted as a discrete rectangular unit arranged chronologically in a vertical column, with the most recent messages placed at the bottom.
    • Categorical Color Encoding: Users can toggle the color mapping of the message rectangles to represent message type (e.g., BroadcastMessage, RequestReply, internal Thought), message sender, or message recipient.
    • Branching and Alignment: When a workflow is reset to a previous timestamp, a new conversation session column is spawned adjacent to the original. The fork point is denoted by a horizontal dash. Messages preceding the fork point are displayed with reduced opacity to denote shared historical context, whereas new downstream messages appear at full opacity.
    • Interactive Navigation: Clicking on any message block in the visual overview automatically scrolls the detailed message history view to the corresponding item, and hovering displays execution metadata.
  4. Knowl 4 — Developer Pain Points in Multi-Agent AI System Debugging

    empirical result

    Formative interviews with 5 expert developers building LLM-based multi-agent applications identified three primary bottlenecks in traditional development workflows:

    1. Cumbersome Log Inspection: Multi-agent runs routinely produce 50 to 100+ text-heavy messages (frequently totaling 6,000–7,000+ words per single task execution), interleaving internal monologues, tool execution dumps, and inter-agent dialogues. Reviewing flat console output files post-hoc makes localizing the source of errors exceedingly difficult.
    2. Absence of Interactive Debugging Controls: Developers lacked mechanisms analogous to standard debugger breakpoints or stepping tools to pause runaway agents, inspect intermediate agent states, or intervene when agents deviate mid-plan.
    3. High Latency in Configuration Iteration: Testing changes to agent configurations (such as system prompt revisions or tool additions) required re-running workflows from the beginning. Due to the stochastic nature of LLMs, verifying whether a fix resolved an issue demanded multiple full-length executions, creating slow and costly debugging loops.
  5. Knowl 5 — Taxonomy of Message Edits for Steering Multi-Agent Systems

    definition

    Analysis of 24 user-initiated message modifications during interactive multi-agent debugging sessions identified three primary strategies used to steer agents toward correct task execution:

    1. Adding Specific Instructions (58.3%, 14/24 edits): Refining high-level or ambiguous instructions into concrete, actionable steps. For example, replacing a broad request to extract a specific baseball statistic with an explicit instruction telling a web agent to first sort a table column in descending order and then read the value in the top row.
    2. Simplifying Instructions (20.8%, 5/24 edits): Removing compound requirements or trimming verbose text from instructions to reduce prompt complexity and prevent LLM attention failure across long instructions (e.g., splitting a two-part query to have the agent identify a target entity first before requesting secondary attributes).
    3. Modifying the Goal of the Plan (20.8%, 5/24 edits): Altering the fundamental strategy or resource selection devised by the agent team, such as changing an Orchestrator's plan from web scraping to executing Python code, or swapping an uninformative target URL for a more suitable data source.
  6. Knowl 6 — User Evaluation and Usability Ratings of AGDebugger

    empirical result

    In a user study with 8 software practitioners debugging failing multi-agent teams on GAIA benchmark tasks, participants rated AGDebugger's usability and feature usefulness on a 5-point Likert scale:

    • System Helpfulness: Mean score of 4.4 / 5.0.
    • Ease of Use: Mean score of 4.0 / 5.0.
    • Desire to Use in Future: Mean score of 4.6 / 5.0.
    • Backtracking and Editing Feature: Mean score of 4.9 / 5.0 (rated highest among all features).
    • Sending New Messages: Mean score of 4.1 / 5.0.
    • Overview Visualization: Mean score of 4.1 / 5.0.

    Every participant utilized the message reset and editing functionality during their 30-minute debugging session, executing between 1 and 5 distinct edits each. In a comparative error identification phase, 5 out of 6 participants preferred AGDebugger over a baseline interface lacking reset and visual branch comparison capabilities.

  7. Knowl 7 — Differential Efficacy of Early vs. Late Message Interventions in Agent Workflows

    empirical result

    In user study evaluations of interactive steering, the temporal location of message edits significantly influenced debugging success:

    • Context Dilution in Late Edits: When modifications were applied late in a conversation, agents frequently failed to follow new directives. Because the prompt context had accumulated extensive earlier conversation history, the underlying LLMs tended to attend to prior context rather than the newly edited message, causing agents to ignore explicit negative constraints (e.g., commands instructing them not to summarize entire web pages).
    • Success of Early Interventions: Both participants who successfully steered failing multi-agent teams to the exact ground-truth answer on GAIA benchmark tasks made their edits near the beginning of the interaction history (one by modifying the high-level plan to use Python code rather than web search, and the other by simplifying an early data extraction step).
  8. Knowl 8 — Non-Resettable Actions and Implementation Knowledge Gaps in Multi-Agent Steering

    limitation

    Interactive debugging and steering of autonomous multi-agent teams face several structural challenges:

    1. Irreversible External Side Effects: State checkpointing and rollbacks are constrained to internal agent memory and local tool environments (e.g., browser URLs). Actions that mutate external systems (such as sending emails, posting messages, or mutating remote databases) cannot be undone via checkpoint resets, necessitating external safeguards and pre-execution validation.
    2. Agent Capability Mismatch: Text-based prompt interfaces obscure internal tool and agent constraints. Developers frequently authored reasonable but incompatible instructions (e.g., instructing a single-action web agent to visit three websites sequentially), which failed because the agent architecture only supported atomic, single-step operations per invocation.
    3. Edit Verification Under Stochasticity: Because LLM inference is stochastic and context-dependent, users could not easily determine whether a change in downstream agent behavior was caused by their prompt intervention or random model variance without running repeated trials.

Coverage note — All primary contributions—including the formative interview findings, the AGDebugger system architecture, checkpointing/reset mechanisms, the overview visualization, the user study evaluation, the steering edit taxonomy, and open challenges—are fully covered.

References

  1. 1.Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software engineering for machine learning: a case study. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice (Montreal, Quebec, Canada) (ICSE-SEIP ’19). IEEE Press, New York, NY, USA, 291–300. https://doi.org/10.1109/ICSE-SEIP.2019.00042
  2. 2.Saleema Amershi, Max Chickering, Steven M. Drucker, Bongshin Lee, Patrice Simard, and Jina Suh. 2015. ModelTracker: Redesigning Performance Analysis Tools for Machine Learning. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15). Association for Computing Machinery, New York, NY, USA, 337–346. https://doi.org/10.1145/2702123.2702509
  3. 3.Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena Glassman. 2023. ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing. arXiv:2309.09128 [cs.HC]
  4. 4.Autoblocks AI. 2024. Autoblocks. https://www.autoblocks.ai. Accessed: 2024-12-02.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., Red Hook, NY, USA, 1877–1901. https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
  6. 6.Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I. Hong, and Adam Perer. 2023. Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 419, 14 pages. https://doi.org/10.1145/3544548.3581268
  7. 7.Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, and Mennatallah El-Assady. 2024. Interactive Analysis of LLMs using Meaningful Counterfactuals. arXiv:2405.00708 [cs.CL] https://arxiv.org/abs/2405.00708
  8. 8.Yuheng Cheng, Ceyao Zhang, Zhengwen Zhang, Xiangrui Meng, Sirui Hong, Wenhao Li, Zihao Wang, Zekai Wang, Feng Yin, Junhua Zhao, and Xiuqiang He. 2024. Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects. arXiv:2401.03428 [cs.AI] https://arxiv.org/abs/2401.03428
  9. 9.CrewAI. 2024. CrewAI. https://www.crewai.com/. Accessed: 2024-12-02.
  10. 10.Victor Dibia, Jingya Chen, Gagan Bansal, Suff Syed, Adam Fourney, Erkang Zhu, Chi Wang, and Saleema Amershi. 2024. AutoGen Studio: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems. arXiv:2408.15247 [cs.SE] https://arxiv.org/abs/2408.15247
  11. 11.Hugging Face. 2024. GAIA Benchmark Leaderboard. https://huggingface.co/spaces/gaia-benchmark/leaderboard Accessed 08-2024.
  12. 12.Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang (Eric) Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadallah, Ece Kamar, Rafah Hosn, and Saleema Amershi. 2024. Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks. Technical Report MSR-TR-2024-47. Microsoft. https://www.microsoft.com/en-us/research/publication/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/
  13. 13.GitKraken. 2024. GitKraken Commit Graph: Bring color & clarity to your commit history. https://www.gitkraken.com/solutions/commit-graph. Accessed: 2024-09.
  14. 14.Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Large Language Model Based Multi-agents: A Survey of Progress and Challenges. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, Kate Larson (Ed.). International Joint Conferences on Artificial Intelligence Organization, Online, 8048–8057. https://doi.org/10.24963/ijcai.2024/890 Survey Track.
  15. 15.Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 6864–6890. https://aclanthology.org/2024.acl-long.371
  16. 16.Ellen Jiang, Kristen Olson, Edwin Toh, Alejandra Molina, Aaron Donsbach, Michael Terry, and Carrie J Cai. 2022. PromptMaker: Prompt-based Prototyping with Large Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI EA ’22). Association for Computing Machinery, New York, NY, USA, Article 35, 8 pages. https://doi.org/10.1145/3491101.3503564
  17. 17.Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues? https://openreview.net/forum?id=VTF8yNQM66
  18. 18.Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon. 2025. LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models. IEEE Transactions on Visualization and Computer Graphics 31, 1 (2025), 503–513. https://doi.org/10.1109/TVCG.2024.3456354
  19. 19.Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 306, 21 pages. https://doi.org/10.1145/3613904.3642216
  20. 20.Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of Explanatory Debugging to Personalize Interactive Machine Learning. In Proceedings of the 20th International Conference on Intelligent User Interfaces (Atlanta, Georgia, USA) (IUI ’15). Association for Computing Machinery, New York, NY, USA, 126–137. https://doi.org/10.1145/2678025.2701399
  21. 21.LangChain. 2024. LangGraph Studio. https://github.com/langchain-ai/langgraph-studio. Accessed: 2024-12-02.
  22. 22.Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. arXiv:2303.17760 [cs.AI] https://arxiv.org/abs/2303.17760
  23. 23.Brian Y. Lim, Anind K. Dey, and Daniel Avrahami. 2009. Why and why not explanations improve the intelligibility of context-aware intelligent systems. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, MA, USA) (CHI ’09). Association for Computing Machinery, New York, NY, USA, 2119–2128. https://doi.org/10.1145/1518701.1519023
  24. 24.Greg Little, Lydia B. Chilton, Max Goldman, and Robert C. Miller. 2010. TurKit: human computation algorithms on mechanical turk. In Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology (New York, New York, USA) (UIST ’10). Association for Computing Machinery, New York, NY, USA, 57–66. https://doi.org/10.1145/1866029.1866040
  25. 25.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12 (2024), 157–173. https://doi.org/10.1162/tacl_a_00638
  26. 26.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 1906–1919. https://doi.org/10.18653/v1/2020.acl-main.173
  27. 27.Gregoire Mialon, Thomas Scialom, Clémentine Fourrier, Thomas Wolf, and Yann LeCun. 2024. GAIA: A Benchmark for General AI Assistants. https://ai.meta.com/research/publications/gaia-a-benchmark-for-general-ai-assistants/
  28. 28.Besmira Nushi, Ece Kamar, Eric Horvitz, and Donald Kossmann. 2017. On human intellect and machine failures: troubleshooting integrative machine learning systems. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (San Francisco, California, USA) (AAAI’17). AAAI Press, New York, NY, USA, 1017–1025.
  29. 29.OpenAI. 2024. Chat Playground. https://platform.openai.com/playground/. Accessed: 2024-09-05.
  30. 30.Deokgun Park, Steven M. Drucker, Roland Fernandez, and Niklas Elmqvist. 2018. Atom: A Grammar for Unit Visualizations. IEEE Transactions on Visualization and Computer Graphics 24, 12 (2018), 3032–3043. https://doi.org/10.1109/TVCG.2017.2785807
  31. 31.Kayur Patel, Naomi Bancroft, Steven M. Drucker, James Fogarty, Amy J. Ko, and James Landay. 2010. Gestalt: integrated support for implementation and analysis in machine learning. In Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology (New York, New York, USA) (UIST ’10). Association for Computing Machinery, New York, NY, USA, 37–46. https://doi.org/10.1145/1866029.1866038
  32. 32.Savvas Petridis, Benjamin D Wedin, James Wexler, Mahima Pushkarna, Aaron Donsbach, Nitesh Goyal, Carrie J Cai, and Michael Terry. 2024. Constitution-Maker: Interactively Critiquing Large Language Models by Converting Feedback into Principles. In Proceedings of the 29th International Conference on Intelligent User Interfaces (Greenville, SC, USA) (IUI ’24). Association for Computing Machinery, New York, NY, USA, 853–868. https://doi.org/10.1145/3640543.3645144
  33. 33.Promptfoo. 2024. Promptfoo. https://www.promptfoo.dev/. Accessed: 2024-12-02.
  34. 34.Python Software Foundation. 2024. The Python Debugger (pdb). https://docs.python.org/3/library/pdb.html. Accessed: 2024-09-05.
  35. 35.Marco Tulio Ribeiro and Scott Lundberg. 2022. Adaptive Testing and Debugging of NLP Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 3253–3267. https://doi.org/10.18653/v1/2022.acl-long.230
  36. 36.D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young. 2014. Machine Learning: The High Interest Credit Card of Technical Debt.
  37. 37.Shreya Shankar, J.D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024. Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (Pittsburgh, PA, USA) (UIST ’24). Association for Computing Machinery, New York, NY, USA, Article 131, 14 pages. https://doi.org/10.1145/3654777.3676450
  38. 38.Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, and Alexander M. Rush. 2022. Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models. arXiv:2208.07852 [cs.CL] https://arxiv.org/abs/2208.07852
  39. 39.Ian Tenney, Ryan Mullins, Bin Du, Shree Pandya, Minsuk Kahng, and Lucas Dixon. 2024. Interactive Prompt Debugging with Sequence Salience. arXiv:2404.07498 [cs.CL] https://arxiv.org/abs/2404.07498
  40. 40.Sandra Wachter, Brent D. Mittelstadt, and Chris Russell. 2017. Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR. arXiv:1711.00399 http://arxiv.org/abs/1711.00399
  41. 41.Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig. 2024. OpenDevin: An Open Platform for AI Software Developers as Generalist Agents. arXiv:2407.16741 [cs.SE] https://arxiv.org/abs/2407.16741
  42. 42.James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda B. Viégas, and Jimbo Wilson. 2020. The What-If Tool: Interactive Probing of Machine Learning Models. IEEE Trans. Vis. Comput. Graph. 26, 1 (2020), 56–65. https://doi.org/10.1109/TVCG.2019.2934619
  43. 43.Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang (Eric) Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Ahmed Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. https://www.microsoft.com/en-us/research/publication/autogen-enabling-next-gen-llm-applications-via-multi-agent-conversation-framework/
  44. 44.Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai. 2022. PromptChainer: Chaining Large Language Model Prompts through Visual Programming. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI EA ’22). Association for Computing Machinery, New York, NY, USA, Article 359, 10 pages. https://doi.org/10.1145/3491101.3519729
  45. 45.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021. Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 6707–6723. https://doi.org/10.18653/v1/2021.acl-long.523
  46. 46.Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for Computing Machinery, New York, NY, USA, Article 385, 22 pages. https://doi.org/10.1145/3491102.3517582
  47. 47.Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. 2024. OS-Copilot: Towards Generalist Computer Agents with Self-Improvement. arXiv:2402.07456 [cs.AI] https://arxiv.org/abs/2402.07456
  48. 48.J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 437, 21 pages. https://doi.org/10.1145/3544548.3581388
  49. 49.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://openreview.net/forum?id=oKn9c6ytLx

Citation

MLA
Epperson, W., et al. “Interactive Debugging and Steering of Multi-Agent AI Systems”. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, pp. 1–5, https://doi.org/10.1145/3706598.3713581.
APA
Epperson, W., Bansal, G., Dibia, V. C., Fourney, A., Gerrits, J., Zhu, E. (Eric) ., & Amershi, S. (2025). Interactive Debugging and Steering of Multi-Agent AI Systems. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–15. https://doi.org/10.1145/3706598.3713581
Chicago
Epperson, W., G. Bansal, V. C. Dibia, et al. 2025. “Interactive Debugging and Steering of Multi-Agent AI Systems”. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–15. https://doi.org/10.1145/3706598.3713581.
Harvard
Epperson, W. et al. (2025) “Interactive Debugging and Steering of Multi-Agent AI Systems”, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, pp. 1–15. Available at: https://doi.org/10.1145/3706598.3713581.
Vancouver
1. Epperson W, Bansal G, Dibia VC, Fourney A, Gerrits J, Zhu E (Eric), Amershi S (2025) Interactive Debugging and Steering of Multi-Agent AI Systems. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, pp 1–15

BibTeX

@inproceedings{Epperson_2025, series={CHI ’25}, title={Interactive Debugging and Steering of Multi-Agent AI Systems}, url={http://dx.doi.org/10.1145/3706598.3713581}, DOI={10.1145/3706598.3713581}, booktitle={Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems}, publisher={ACM}, author={Epperson, Will and Bansal, Gagan and Dibia, Victor C and Fourney, Adam and Gerrits, Jack and Zhu, Erkang (Eric) and Amershi, Saleema}, year={2025}, month=Apr, pages={1–15}, collection={CHI ’25} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/