Magentic-UI: Towards Human-in-the-loop Agentic Systems

Hussein MozannarGagan BansalCheng TanAdam FourneyVictor DibiaJingya ChenJack GerritsTyler PayneMatheus Kunzler MaldanerMadeleine Grunde-McLaughlin

article2025arXiv55 citations

Presents Magentic-UI, an open-source research platform that implements six core interaction mechanisms—such as co-planning, live intervention, and action approvals—to enable safe, cost-effective human oversight of multi-agent LLM systems performing complex web and coding tasks.

Listen

Autonomous artificial intelligence agents powered by large language models are increasingly capable of executing multi-step tasks across the web, code repositories, and local file systems. However, current agents still fail to reach human-level reliability and present substantial safety risks, such as falling victim to adversarial prompt injections, leaking confidential data, or executing irreversible real-world actions. The article evaluates Magentic-UI, an open-source, user-facing research prototype designed to address these shortcomings by embedding human oversight directly into agent workflows. Its primary objective is to demonstrate how human-in-the-loop interaction mechanisms can bridge the performance gap between imperfect artificial intelligence systems and full human proficiency while maintaining safety and operational control.

To assess this system, the article employs a comprehensive, multi-method approach. The authors evaluated Magentic-UI across standard autonomous agent benchmarks—including GAIA, AssistantBench, WebVoyager, and WebGames—using state-of-the-art models such as GPT-4o and o4-mini. They conducted simulated user experiments on the GAIA benchmark to model the impact of human guidance, carried out qualitative user studies with 12 participants performing multi-task workflows, and executed targeted security red-teaming across 24 adversarial attack scenarios that featured direct exploitation, social engineering, and prompt injections.

Key findings show that human intervention significantly boosts agent capabilities at low human effort. In simulated user testing on the GAIA benchmark, pairing the agent with lightweight human guidance increased task completion accuracy by 71% (improving from 30.3% to 51.9%), even though the system requested human assistance on only 10% of tasks and averaged just 1.1 help requests per difficult task. In autonomous benchmark testing, Magentic-UI achieved strong task success rates, reaching 82.2% on WebVoyager and 45.5% on WebGames with o4-mini. User study participants gave the interface a solid usability rating (a System Usability Scale score of 74.58), finding co-planning and background multi-tasking valuable, though they highlighted friction around system latency, verbose progress logs, and rigid approval prompts. In safety assessments, default security mechanisms—specifically Docker container sandboxing, separate browser instances, and two-stage action guards—successfully thwarted 100% of the 24 adversarial attack scenarios.

These results demonstrate that organizations do not need to wait for fully autonomous artificial intelligence to realize productivity benefits from agentic workflows. By incorporating low-cost human oversight mechanisms such as up-front co-planning, real-time interruption, and action approval guards, imperfect systems can be deployed safely without exposing critical infrastructure or private data to prompt injection vulnerabilities. However, the study also shows that when security mitigations were intentionally disabled, agents consistently succumbed to prompt injections that exfiltrated access keys and certificates, underscoring that architectural isolation and human oversight are strictly necessary safeguards.

For practical adoption, organizations should implement layered defense architectures that isolate agent execution environments within sandboxed containers and enforce human approval on high-risk, irreversible operations. Agent interfaces should be tuned to avoid excessive interruption by batching clarifying questions and offering visual summarization tools rather than raw text logs. Further work is required to conduct long-term, randomized controlled trials that quantify net productivity gains and evaluate more expressive planning structures, such as hierarchical or branching workflows.

The findings are bounded by certain limitations. The evaluations relied on simulated user behaviors and short qualitative sessions rather than long-term workplace deployments, meaning true organizational productivity impact remains unmeasured. Furthermore, testing was limited to English-language tasks, and agent performance remains lower on highly complex coding and broad computer-control problems. Nevertheless, the evidence strongly supports human-in-the-loop architecture as a practical and secure paradigm for current agent deployment.

Cover for Magentic-UI: Towards Human-in-the-loop Agentic Systems

Abstract

AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-level performance in most domains including computer use, software development, and research. Their growing autonomy and ability to interact with the outside world, also introduces safety and security risks including potentially misaligned actions and adversarial manipulation. We argue that human-in-the-loop agentic systems offer a promising path forward, combining human oversight and control with AI efficiency to unlock productivity from imperfect systems. We introduce Magentic-UI, an open-source web interface for developing and studying human-agent interaction. Built on a flexible multi-agent architecture, Magentic-UI supports web browsing,

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Collaborative Planning
  • 4 Collaborative Task Execution
  • 5 Agent Memory
  • 6 System Implementation
  • 6.1 Overall Design
  • 6.2 Multi-Agent Architecture and Orchestration
  • 6.3 Agent Details
  • 6.4 Action Guard
  • 7 Evaluation
  • 7.1 Setup
  • 7.2 Autonomous Evaluation
  • 7.3 Simulated User Evaluation
  • 7.4 Qualitative User Study
  • 7.5 Safety and Security Testing
  • 8 Discussion
  • 8.1 Progress Towards Effective Human-Agent Collaboration [7]
  • 8.2 Limitations
  • 8.3 Risks and Mitigations
  • 8.4 Conclusion
  • References
  • A Overview of Magentic-UI
  • B Action Guard Prompt
  • C Memory
  • C.1 Plan Learning Prompt
  • D Safety and Security: Adversarial Scenarios

Knowls

  1. Knowl 1 — Multi-Agent Architecture and Interaction Mechanisms of Magentic-UI

    model/method

    Magentic-UI is an open-source, human-in-the-loop multi-agent system designed for complex web, code, and file manipulation tasks. Adapted from the Magentic-One architecture built on the AutoGen framework, Magentic-UI integrates the human user directly into the multi-agent team via a dedicated UserProxy agent.

    The system comprises the following primary components:

    1. Lead Orchestrator Agent: Directs task understanding, plan synthesis, dynamic step assignment to specialized sub-agents, progress tracking, and final answer generation.
    2. WebSurfer Agent: Controls a live Chromium browser via Playwright within an isolated Docker container. It accepts multimodal inputs, executes web actions (e.g., clicking, typing, scrolling, uploading), and enforces an allow-list requiring explicit user approval for unlisted domains.
    3. Coder Agent: Generates self-contained Python or Bash scripts executed inside an isolated Docker container, capturing outputs and regenerating code up to three times upon execution errors.
    4. FileSurfer Agent: Performs local file discovery, conversion (e.g., PDF to Markdown via MarkItDown), and content queries within a Docker container.
    5. Model Context Protocol (MCP) Agents: Wraps remote MCP servers into custom sub-agents, exposing unified toolsets to the agent team.
    6. UserProxy Agent: Represents the human in the agent roster, allowing the Orchestrator to delegate steps to the user when human intervention (e.g., solving CAPTCHAs, providing domain guidance, answering clarifying questions) is necessary.

    Magentic-UI organizes human-agent collaboration through six core interaction mechanisms:

    • Co-planning: Up-front collaborative generation and editing of natural language task plans.
    • Co-tasking: Dynamic control handoffs and mid-execution steering between the user and agents.
    • Action Approval: Intercepting high-risk or irreversible actions for explicit user authorization.
    • Answer Verification: Exposing execution histories, screenshots, and artifacts to enable human validation of final outputs.
    • Agent Memory: Storing and retrieving past execution plans to automate recurring workflows.
    • Multi-tasking: Managing multiple concurrent asynchronous agent sessions.
  2. Knowl 2 — Plan Schema and Orchestrator Progress Ledger

    model/method

    Magentic-UI structures task coordination around two primary data representations: an editable natural language plan Domain-Specific Language (DSL) and a progress ledger generated by the Orchestrator at each execution cycle.

    A plan is defined as an ordered sequence of plan steps: PlanStep:=(agent_name,title,details)\text{PlanStep} := (\text{agent\_name}, \text{title}, \text{details}) Plan:=[PlanStep1,PlanStep2,…,PlanStepn]\text{Plan} := [\text{PlanStep}_1, \text{PlanStep}_2, \dots, \text{PlanStep}_n] where agent_name\text{agent\_name} specifies the assigned agent (e.g., WebSurfer, Coder, FileSurfer, UserProxy), title\text{title} provides a concise summary, and details\text{details} contains natural language instructions for the sub-task.

    During execution, the Orchestrator tracks the current step index i∈{1,…,n}i \in \{1, \dots, n\} and computes a progress ledger structured as:

    \text{Progress Ledger} := \{ & \\ & \text{step\_complete}: (\text{reason}: \text{string}, \text{answer}: \text{boolean}), \\ & \text{replan}: (\text{reason}: \text{string}, \text{answer}: \text{boolean}), \\ & \text{instruction}: (\text{answer}: \text{string}, \text{agent\_name}: \text{string}), \\ & \text{progress\_summary}: \text{string} \\ \} & \end{aligned}$$ where $\text{step\_complete.answer}$ denotes whether step $i$ is finished, $\text{replan.answer}$ indicates whether unexpected obstacles necessitate plan regeneration, $\text{instruction}$ directs the designated agent for the current sub-step, and $\text{progress\_summary}$ maintains accumulated task context.
  3. Knowl 3 — Magentic-UI Orchestrator Execution Loop

    algorithm

    The Magentic-UI Orchestrator executes a user-accepted plan through an iterative control loop that tracks progress, delegates actions to agents, prompts for replanning when obstacles occur, and terminates upon completing all steps.

    Input: Task description TaskTask, initial plan Plan=[PlanStep1,…,PlanStepn]Plan = [PlanStep_1, \dots, PlanStep_n]
    Output: Final Answer string and associated artifacts
    i←0i \leftarrow 0
    while true do
        ledger←GenerateProgressLedger(Task,Plan,i)ledger \leftarrow \text{GenerateProgressLedger}(Task, Plan, i)
        if ledger.replan.answer=Trueledger.replan.answer = \text{True} then
            Plan←Replan(Plan[1…i],Task,ledger)Plan \leftarrow \text{Replan}(Plan[1 \dots i], Task, ledger)
            n←∣Plan∣n \leftarrow |Plan|
        else
            if ledger.step_complete.answer=Trueledger.step\_complete.answer = \text{True} then
                i←i+1i \leftarrow i + 1
                if i>ni > n then
                    return GetFinalAnswer()\text{GetFinalAnswer}()
                end if
            end if
            result←CallAgent(ledger.instruction.agent_name,ledger.instruction.answer)result \leftarrow \text{CallAgent}(ledger.instruction.agent\_name, ledger.instruction.answer)
        end if
    end while

    At each iteration, GenerateProgressLedger queries the Orchestrator LLM with the task context, execution history, and current step index ii. If replanning is needed (ledger.replan.answer=Trueledger.replan.answer = \text{True}), the system preserves completed prefix steps Plan[1…i]Plan[1 \dots i], synthesizes an updated plan with user review, and resets nn. Otherwise, if the step is complete, the loop increments ii and returns the final answer once i>ni > n. If the step is still active, CallAgent invokes the assigned sub-agent with the specific instruction from the ledger.

  4. Knowl 4 — Two-Stage ActionGuard System for Agent Action Approval

    model/method

    To prevent unintended, costly, or unsafe agent actions (e.g., executing arbitrary code, modifying sensitive files, submitting forms, making online purchases), Magentic-UI implements a two-stage gating mechanism called ActionGuard before any agent action is executed.

    1. Pre-defined Irreversibility Heuristics: Actions are categorized based on action metadata into three classes:

      • Never Irreversible (e.g., scrolling a webpage, moving UI focus, non-sensitive search queries): The action bypasses human approval and executes automatically.
      • Always Irreversible (e.g., uploading files, submitting transactions, modifying system state): The action is paused immediately and requires explicit binary human approval (accept/reject).
      • Maybe Irreversible (e.g., clicking interactive web buttons, submitting web forms, executing scripts): The action proposal is passed to the second-stage LLM judge.
    2. LLM-Based ActionGuard Judge: An LLM is provided with the proposed action, recent interaction history, and safety criteria (evaluating potential real-world safety/security impacts, undo complexity, privacy violations, and external side-effects). If the judge assesses the action as potentially harmful or impactful, execution is halted until the user explicitly approves or rejects the action in the UI.

  5. Knowl 5 — Task-Centric Plan Learning and Memory Retrieval in Magentic-UI

    model/method

    Magentic-UI incorporates a long-term memory system that indexes past successful task trajectories as plan templates (task,plan)(task, plan) for reuse on recurring or related tasks.

    Plan Learning: Upon completing a task, the user can trigger plan learning. An LLM processes the complete interaction trace (including user messages, agent actions, clicked elements, and visited URLs) and synthesizes an abstracted, minimal-step plan DSL. This learned plan captures the direct execution route without memorizing the specific final answer and is saved into a persistent gallery.

    Plan Retrieval: Saved plans can be retrieved through three mechanisms:

    1. Manual Selection / Autocomplete: As the user types a new query, an autocomplete interface suggests matching saved plans, or the user manually attaches a saved plan to serve as guidance while altering task parameters.
    2. Direct Execution: Users rerun a saved plan directly on the original task via the plan gallery.
    3. Task-Centric Memory Controller: Automated retrieval powered by AutoGen's Task-Centric Memory. When a new task is received, the controller generalizes the task, embeds multi-word topic vectors using an LLM, and performs nearest-neighbor vector search in a ChromaDB MemoryBank. Candidate plans pass through an LLM relevance filter to select the single most relevant plan, which is supplied to the Orchestrator as a hint during initial plan generation.
  6. Knowl 6 — Autonomous Benchmark Performance of Magentic-UI Across Digital Agent Benchmarks

    data/table

    Magentic-UI was evaluated in fully autonomous mode (with co-planning auto-accepted, co-tasking disabled, and ActionGuard auto-approving all actions) across four agent benchmarks: GAIA (test set, N=300N=300), AssistantBench (test set, N=181N=181), WebVoyager (N=643N=643 live web tasks), and Convergence WebGames (N=53N=53 interactive game tasks).

    Method GAIA (%) AssistantBench (%) WebVoyager (%) WebGames (%)
    Magentic-One (GPT-4o, o1) 38.00 27.7 – –
    SPA →\rightarrow CB (Claude) – 26.4 – –
    Su Zero Ultra 80.04 – – –
    tt_api_1 (GPT-4o) – 28.30 – –
    AWorld 77.08 – – –
    Langfun Agent 2.3 73.09 – – –
    Claude Computer-Use – – 52.0 35.3
    Proxy – – 82.0 43.1
    OpenAI Operator (GPT-4o / o3) 12.3 / 62.2 – 87.0 –
    GPT-4o (SoMs+ReAct / no tools) 6.67 16.5 64.1 41.2
    Browser Use – – 89.1 –
    Human Performance 92.00 – – 95.7
    Magentic-UI (o4-mini) 42.52 27.6 82.2 45.5

    On WebVoyager, Magentic-UI was evaluated with only the WebSurfer agent (achieving 82.2% with o4-mini and 72.2% with GPT-4o). On WebGames, Magentic-UI used WebSurfer and FileSurfer powered by GPT-4o with extended click and file-upload tools, achieving 45.5%. On GAIA and AssistantBench, Magentic-UI matches or exceeds its autonomous predecessor Magentic-One (42.52% vs 38.00% on GAIA; 27.6% vs 27.7% on AssistantBench), showing that architectural adaptations for interactivity do not compromise standalone agent execution.

  7. Knowl 7 — Task Execution Duration and Re-Planning Dynamics in Magentic-UI

    empirical result

    Analysis of Magentic-UI's autonomous execution dynamics across agent benchmarks demonstrates consistent relationships between task success, execution duration, and dataset replanning characteristics:

    1. Execution Duration vs. Correctness (WebVoyager):

      • Successful task completions have a median run time of 113.9 s113.9\text{ s}.
      • Unsuccessful task completions have a median run time of 236.7 s236.7\text{ s}, exhibiting a fat-tailed distribution because the agent spends extended time exploring multiple alternative strategies before failing or exhausting its action budget.
    2. Planning Statistics Across Datasets:

      • WebVoyager (N=643N=643): Median plan length is 2.002.00 steps (average 1.711.71, maximum 5.005.00). Re-planning occurs in 9.6%9.6\% of tasks.
      • GAIA (N=301N=301): Median plan length is 2.002.00 steps (average 2.042.04, maximum 9.009.00). Re-planning occurs in 20.6%20.6\% of tasks.
      • AssistantBench (N=181N=181): Median plan length is 2.002.00 steps (average 2.282.28, maximum 8.008.00). Re-planning occurs in 22.1%22.1\% of tasks.
      • WebGames (N=51N=51): Median plan length is 1.001.00 step (average 1.291.29, maximum 4.004.00). Re-planning occurs in 52.9%52.9\% of tasks. This high replanning frequency aligns with the system's failure rate because WebGames provides an unambiguous success signal (revealing a secret password), prompting the agent to iteratively replan until success.
  8. Knowl 8 — Quantitative Gains from Simulated Human-in-the-Loop Interventions on GAIA

    empirical result

    Evaluating Magentic-UI on the GAIA validation set (N=162N=162 tasks) with simulated user models demonstrates the efficiency of lightweight human-in-the-loop interactions:

    • Autonomous Baseline (GPT-4o): Magentic-UI operating without human assistance achieved a task completion rate of 30.3%30.3\% (comparable to autonomous Magentic-One at 34.0%±7.0%34.0\% \pm 7.0\%).
    • Smarter Model User: An un-tooled simulated user powered by o4-mini collaborating with GPT-4o agent workers (providing feedback during co-planning and answering help requests during co-tasking) improved accuracy to 42.6%42.6\%. Magentic-UI requested user help in only 4.3%4.3\% of tasks, averaging 1.71.7 queries per assisted task.
    • Side-Information User: A simulated user powered by GPT-4o with access to human ground-truth solution plans (prompted to guide indirectly without leaking answers) improved accuracy by 71%71\% relative to the autonomous baseline, reaching 51.9%51.9\%. Magentic-UI asked for help in 10%10\% of tasks (averaging 1.11.1 queries) and deferred to the user for the final answer in 18%18\% of tasks.
    • Human Ceiling: Human benchmark accuracy on GAIA is 92.0%92.0\%.

    These findings show that targeted, low-frequency human intervention (occurring in 4.3%–10%4.3\%\text{--}10\% of tasks) substantially improves agent accuracy at a small fraction of the cost of full manual task completion.

  9. Knowl 9 — Usability and User Experience Findings from Magentic-UI Qualitative Study

    empirical result

    A qualitative user study with 12 participants evaluating Magentic-UI across three simultaneous multi-tasking scenarios yielded an overall System Usability Scale (SUS) score of 74.5874.58:

    • 75%75\% of participants agreed or strongly agreed that the interface was easy to use.
    • 91.7%91.7\% disagreed that the system was unnecessarily complex.
    • 41.7%41.7\% expressed intent to use the system frequently, reflecting task-specific applicability and latency sensitivities.

    Key qualitative findings across interaction mechanisms include:

    1. Co-planning: Participants favored co-planning for incorporating subjective preferences and safety constraints, reporting that modifying generated natural language plans was significantly easier than drafting plans from scratch.
    2. Co-tasking: Users actively used co-tasking to overcome edge cases (e.g., CAPTCHAs, unhandled DOM elements, faulty price filtering) and to manage the entry of sensitive personal data (e.g., credentials, payment information).
    3. Action Approvals: Approvals were strongly endorsed for high-risk operations (e.g., payments, email transmission, account registrations). However, approvals for low-risk actions (e.g., adding items to cart) were perceived as excessive, highlighting the necessity of adaptive interruption thresholds.
    4. Verification: Participants preferred visual verification mechanisms (embedded live browser views, annotated chronological screenshots) over textual reasoning traces or raw code execution logs when monitoring background tasks.
  10. Knowl 10 — Layered Security Mitigations and Adversarial Robustness in Magentic-UI

    empirical result

    Magentic-UI enforces a multi-layered security architecture to protect against adversarial manipulation and unintended actions:

    1. Docker Container Sandboxing: Code execution, file manipulation, and browser instances run in isolated Docker containers, preventing unauthorized host filesystem access and environment variable extraction.
    2. Isolated Browser Session: The agent browser runs without sharing user credentials, stored session cookies, or native browsing history.
    3. URL Allow-Listing: Restricts autonomous web navigation to pre-approved domains; navigating outside the allow-list triggers explicit approval requests displaying the target URL, page title, and access justification.
    4. ActionGuard & Anomaly Replanning: Risky operations and suspicious webpage instructions trigger replanning and pause execution for human sign-off.

    In red-teaming across 24 realistic adversarial test cases (including direct private SSH key access requests, phishing update dialogs, OAuth permission escalation, and cross-site prompt injections):

    • Default Configuration: Magentic-UI successfully mitigated all 24 attacks through sandbox containment, domain gating, or by identifying anomalies and halting for user authorization.
    • Mitigations-Disabled Baseline: When security mitigations were disabled in developer mode, GPT-4o resisted direct social engineering but was successfully exploited by indirect prompt injections, which exfiltrated SSH keys, created persistent GitHub API access tokens, searched private emails, and attempted autonomous approval via the agent's own web UI.

Coverage note — None was omitted; the extraction covers Magentic-UI's multi-agent architecture, plan/ledger formalisms, execution loop algorithm, ActionGuard, plan memory, benchmark evaluations across four datasets, runtime/replanning statistics, simulated user experiments, qualitative user study, and adversarial security evaluations.

References

  1. 1.M. Aliannejadi, H. Zamani, F. Crestani, and W. B. Croft. Asking clarifying questions in open-domain information-seeking conversations. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR âĂŹ19, page 475âĂŞ484. ACM, July 2019.
  2. 2.D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
  3. 3.Anthropic. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, oct 2024.
  4. 4.A. T. at InclusionAI. Aworld: A framework for agent learning of complex tasks via action-observation-reward experience, 2025.
  5. 5.BabyAGI. Github | babyagi. https://github.com/yoheinakajima/babyagi, 2023.
  6. 6.G. Bansal, B. Nushi, E. Kamar, D. Weld, W. Lasecki, and E. Horvitz. Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff. In AAAI Conference on Artificial Intelligence. AAAI, January 2019.
  7. 7.G. Bansal, J. W. Vaughan, S. Amershi, E. Horvitz, A. Fourney, H. Mozannar, V. Dibia, and D. S. Weld. Challenges in human-agent communication. ArXiv, 2024.
  8. 8.G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, and D. S. Weld. Does the whole exceed its parts? the effect of ai explanations on complementary team performance, 2021.
  9. 9.J. Brooke. SUS – a quick and dirty usability scale, pages 189–194. 01 1996.
  10. 10.Z. Chen, M. White, R. Mooney, A. Payani, Y. Su, and H. Sun. When is tree search useful for llm planning? it depends on the discriminator, 2024.
  11. 11.Y. Cheng, C. Zhang, Z. Zhang, X. Meng, S. Hong, W. Li, Z. Wang, Z. Wang, F. Yin, J. Zhao, et al. Exploring large language model based intelligent agents: Definitions, methods, and prospects. arXiv preprint arXiv:2401.03428, 2024.
  12. 12.P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  13. 13.Cognition.ai. Introducing devin, the first ai software engineer, 2024.
  14. 14.K. Z. Cui, M. Demirer, S. Jaffe, L. Musolff, S. Peng, and T. Salz. The productivity effects of generative ai: Evidence from a field experiment with github copilot. 2024.
  15. 15.X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su. Mind2web: Towards a generalist agent for the web. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 28091–28114. Curran Associates, Inc., 2023.
  16. 16.X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su. Mind2web: Towards a generalist agent for the web, 2023.
  17. 17.L. Dong, T. Yuan, Y. Wang, T. Xia, Z. Zhang, Z. He, B. Zhou, R. Wang, F. Li, G. Liu, L. Xu, and R. Zhao. R-judge: Benchmarking safety risk awareness for llm agents. In Conference on Empirical Methods in Natural Language Processing, 2024.
  18. 18.Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch. Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325, 2023.
  19. 19.H. Fang, X. Zhu, and I. Gurevych. Inferact: Inferring safe actions for llm-based agents through preemptive evaluation and human feedback, 2024.
  20. 20.K. J. K. Feng, K. Pu, M. Latzke, T. August, P. Siangliulue, J. Bragg, D. S. Weld, A. X. Zhang, and J. C. Chang. Cocoa: Co-planning and co-execution with ai agents, 2025.
  21. 21.A. Fourney, G. Bansal, H. Mozannar, C. Tan, E. Salinas, Erkang, Zhu, F. Niedtner, G. Proebsting, G. Bassman, J. Gerrits, J. Alber, P. Chang, R. Loynd, R. West, V. Dibia, A. Awadallah, E. Kamar, R. Hosn, and S. Amershi. Magentic-one: A generalist multi-agent system for solving complex tasks, 2024.
  22. 22.GitHub. Github copilot, 2021.
  23. 23.GitHub Next. Copilot workspace: An agentic dev environment, designed for everyday tasks. https://githubnext.com/projects/copilot-workspace, May 2025. Technical preview (sunset May 30, 2025).
  24. 24.B. Gou, Z. Huang, Y. Ning, Y. Gu, M. Lin, W. Qi, A. Kopanev, B. Yu, B. J. GutiÃlrrez, Y. Shu, C. H. Song, J. Wu, S. Chen, H. N. Moussa, T. Zhang, J. Xie, Y. Li, T. Xue, Z. Liao, K. Zhang, B. Zheng, Z. Cai, V. Rozgic, M. Ziyadi, H. Sun, and Y. Su. Mind2web 2: Evaluating agentic search with agent-as-a-judge, 2025.
  25. 25.N. Goyal, M. Chang, and M. Terry. Designing for human-agent alignment: Understanding what humans want from their agents. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–6, 2024.
  26. 26.B. J. Grosz and S. Kraus. The evolution of sharedplans. In Proceedings of the International Conference on Multi-Agent Systems, 1999.
  27. 27.T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang. Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680, 2024.
  28. 28.H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu. Webvoyager: Building an end-to-end web agent with large multimodal models. arXiv preprint arXiv:2401.13919, 2024.
  29. 29.H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu. Webvoyager: Building an end-to-end web agent with large multimodal models, 2024.
  30. 30.S. Hong, X. Zheng, J. Chen, Y. Cheng, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, et al. Metagpt: Meta programming for multi-agent collaborative framework. arXiv preprint arXiv:2308.00352, 2023.
  31. 31.K.-H. Huang, A. Prabhakar, S. Dhawan, Y. Mao, H. Wang, S. Savarese, C. Xiong, P. Laban, and C.-S. Wu. Crmarena: Understanding the capacity of llm agents to perform professional crm tasks in realistic environments. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025.
  32. 32.K.-H. Huang, A. Prabhakar, O. Thorat, D. Agarwal, P. K. Choubey, Y. Mao, S. Savarese, C. Xiong, and C.-S. Wu. Crmarena-pro: Holistic assessment of llm agents across diverse business scenarios and interactions. arXiv preprint arXiv:2505.18878, 2025.
  33. 33.F. Huq, Z. Z. Wang, F. F. Xu, T. Ou, S. Zhou, J. P. Bigham, and G. Neubig. Cowpilot: A framework for autonomous and human-agent collaborative web navigation. arXiv preprint arXiv:2501.16609, 2025.
  34. 34.C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. Swe-bench: Can language models resolve real-world github issues?, 2024.
  35. 35.J. Y. Koh, S. McAleer, D. Fried, and R. Salakhutdinov. Tree search for language model agents, 2024.
  36. 36.E. Li and J. Waldo. Websuite: Systematically evaluating why web agents fail. arXiv preprint arXiv:2406.01623, 2024.
  37. 37.G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem. Camel: Communicative agents for "mind" exploration of large scale language model society, 2023.
  38. 38.W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva. On the effects of data scale on computer control agents. arXiv preprint arXiv:2406.03679, 2024.
  39. 39.T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi. Encouraging divergent thinking in large language models through multi-agent debate, 2023.
  40. 40.Z. Liao, J. Jones, L. Jiang, E. Fosler-Lussier, Y. Su, Z. Lin, and H. Sun. Redteamcua: Realistic adversarial testing of computer-use agents in hybrid web-os environments, 2025.
  41. 41.J. Liu, Y. Song, B. Y. Lin, W. Lam, G. Neubig, Y. Li, and X. Yue. Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding?, 2024.
  42. 42.N. Liu, L. Chen, X. Tian, W. Zou, K. Chen, and M. Cui. From llm to conversational agent: A memory enhanced architecture with fine-tuning of large language models. arXiv e-prints, pages arXiv–2401, 2024.
  43. 43.Y. Liu, S. K. Lo, Q. Lu, L. Zhu, D. Zhao, X. Xu, S. Harrer, and J. Whittle. Agent design pattern catalogue: A collection of architectural patterns for foundation model based agents. arXiv preprint arXiv:2405.10467, 2024.
  44. 44.D. Madras, T. Pitassi, and R. Zemel. Predict responsibly: improving fairness and accuracy by learning to defer. Advances in neural information processing systems, 31, 2018.
  45. 45.T. Masterman, S. Besen, M. Sawtell, and A. Chao. The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: A survey. arXiv preprint arXiv:2404.11584, 2024.
  46. 46.B. Messing. An introduction to multiagent systems. Künstliche Intell., 17:58–, 2002.
  47. 47.G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom. Gaia: a benchmark for general ai assistants, 2023.
  48. 48.G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom. Gaia: benchmark for general ai assistants. arXiv preprint arXiv:2311.12983, 2023.
  49. 49.Microsoft. AutoGen AgentChat User Guide. Microsoft, 2025. Accessed: 2025-07-07.
  50. 50.Microsoft. MarkItDown: Python tool for converting files and office documents to Markdown. https://github.com/microsoft/markitdown, May 2025. Version 0.1.2.
  51. 51.H. Mozannar. Web agent tutorial. husseinmozannar.github.io, June 2025.
  52. 52.H. Mozannar, J. J. Lee, D. Wei, P. Sattigeri, S. Das, and D. Sontag. Effective human-ai teams via learned natural language rules and onboarding, 2023.
  53. 53.H. Mozannar and D. Sontag. Consistent estimators for learning to defer to an expert. In International conference on machine learning, pages 7076–7087. PMLR, 2020.
  54. 54.M. MÃijller and G. Å¡uniÄŊ. Browser use: Enable ai to control your browser, 2024.
  55. 55.K. Narasimhan, J. Yang, H. Chen, and S. Yao. Webshop: Towards scalable real-world web interaction with grounded language agents. ArXiv, abs/2207.01206, 2022.
  56. 56.OpenAI. Introducing deep research, 2025.
  57. 57.OpenAI. Introducing operator, jan 2025.
  58. 58.B. Pan, J. Lu, K. Wang, L. Zheng, Z. Wen, Y. Feng, M. Zhu, and W. Chen. Agentcoord: Visually exploring coordination strategy for llm-based multi-agent collaboration. arXiv preprint arXiv:2404.11943, 2024.
  59. 59.J. Pan, Y. Zhang, N. Tomlin, Y. Zhou, S. Levine, and A. Suhr. Autonomous evaluation and refinement of digital agents, 2024.
  60. 60.Y. Pan, D. Kong, S. Zhou, C. Cui, Y. Leng, B. Jiang, H. Liu, Y. Shang, S. Zhou, T. Wu, and Z. Wu. Webcanvas: Benchmarking web agents in online environments, 2024.
  61. 61.B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  62. 62.D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and B. Faltings. REFINER: Reasoning feedback on intermediate representations. In Y. Graham and M. Purver, editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1100–1126, St. Julian’s, Malta, Mar. 2024. Association for Computational Linguistics.
  63. 63.S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer. The impact of ai on developer productivity: Evidence from github copilot, 2023.
  64. 64.P. Putta, E. Mills, N. Garg, S. Motwani, C. Finn, D. Garg, and R. Rafailov. Agent q: Advanced reasoning and learning for autonomous ai agents, 2024.
  65. 65.Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun. Tool learning with foundation models, 2023.
  66. 66.Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun. Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023.
  67. 67.Red Cell Partners. Trase tops gaia leaderboard, 2024.
  68. 68.G. Sarch, S. Somani, R. Kapoor, M. J. Tarr, and K. Fragkiadaki. Helper-x: A unified instructable embodied agent to tackle four interactive vision-language domains with memory-augmented language models, 2024.
  69. 69.G. Sarch, Y. Wu, M. Tarr, and K. Fragkiadaki. Open-ended instructable embodied agents with memory-augmented large language models. In H. Bouamor, J. Pino, and K. Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3468–3500, Singapore, Dec. 2023. Association for Computational Linguistics.
  70. 70.P. Scerri, D. V. Pynadath, and M. Tambe. Adjustable autonomy in real-world multi-agent environments. In International Conference on Autonomous Agents, 2001.
  71. 71.T. Schick, J. Dwivedi-Yu, R. DessÃň, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools, 2023.
  72. 72.O. Shaikh, K. GligoriÄĞ, A. Khetan, M. Gerstgrasser, D. Yang, and D. Jurafsky. Grounding gaps in language model generations, 2024.
  73. 73.O. Shaikh, H. Mozannar, G. Bansal, A. Fourney, and E. Horvitz. Navigating rifts in human-llm grounding: Study and benchmark, 2025.
  74. 74.Y. Shao, V. Samuel, Y. Jiang, J. Yang, and D. Yang. Collaborative gym: A framework for enabling and evaluating human-agent collaboration. arXiv preprint arXiv:2412.15701, 2024.
  75. 75.Y. Shavit, S. Agarwal, M. Brundage, S. A. C. OâĂŹKeefe, R. Campbell, T. Lee, P. Mishkin, T. Eloundou, A. Hickey, K. Slama, L. Ahmad, P. McMillan, A. Beutel, A. Passos, and D. G. Robinson. Practices for governing agentic ai systems.
  76. 76.T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang. World of bits: An open-domain platform for web-based agents. In International Conference on Machine Learning. PMLR, 2017.
  77. 77.C. Si, T. Hashimoto, and D. Yang. The ideation-execution gap: Execution outcomes of llm-generated versus human research ideas, 2025.
  78. 78.P. Sodhi, S. R. K. Branavan, Y. Artzi, and R. McDonald. Step: Stacked llm policies for web actions, 2024.
  79. 79.Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin. Trial and error: Exploration-based trajectory optimization for llm agents, 2024.
  80. 80.P. Stone and M. Veloso. Multiagent systems: A survey from a machine learning perspective. Auton. Robots, 8(3):345âĂŞ383, June 2000.
  81. 81.Y. Talebirad and A. Nadiri. Multi-agent collaboration: Harnessing the power of intelligent llm agents, 2023.
  82. 82.M. Tambe. Implementing agent teams in dynamic multiagent environments. Appl. Artif. Intell., 12:189–210, 1998.
  83. 83.G. Thomas, A. J. Chan, J. Kang, W. Wu, F. Christianos, F. Greenlee, A. Toulis, and M. Purtorab. Webgames: Challenging general-purpose web-browsing ai agents, 2025.
  84. 84.M. Vaccaro, A. Almaatouq, and T. Malone. When combinations of humans and ai are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12):2293âĂŞ2303, Oct. 2024.
  85. 85.K. Valmeekam, S. Sreedharan, M. Marquez, A. Olmo, and S. Kambhampati. On the planning abilities of large language models (a critical investigation with a proposed benchmark), 2023.
  86. 86.S. Vijayvargiya, A. B. Soni, X. Zhou, Z. Z. Wang, N. Dziri, G. Neubig, and M. Sap. Openagentsafety: A comprehensive framework for evaluating real-world ai agent safety. arXiv preprint arXiv:2507.06134, 2025.
  87. 87.R. Wang, L. Zheng, and B. An. Synapse: Trajectory-as-exemplar prompting with memory for computer control. In International Conference on Learning Representations, 2023.
  88. 88.X. Wang, Y. Chen, L. Yuan, Y. Zhang, Y. Li, H. Peng, and H. Ji. Executable code actions elicit better llm agents, 2024.
  89. 89.X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. Openhands: An open platform for ai software developers as generalist agents, 2025.
  90. 90.Y. Wang, T. Shen, L. Liu, and J. Xie. Sibyl: Simple yet effective agent framework for complex real-world reasoning, 2024.
  91. 91.Z. Z. Wang, J. Mao, D. Fried, and G. Neubig. Agent workflow memory, 2024.
  92. 92.J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  93. 93.Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. In COLM, 2024.
  94. 94.Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong. Os-copilot: Towards generalist computer agents with self-improvement. ArXiv, abs/2402.07456, 2024.
  95. 95.Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong. Os-copilot: Towards generalist computer agents with self-improvement, 2024.
  96. 96.Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huang, and T. Gui. The rise and potential of large language model based agents: A survey, 2023.
  97. 97.T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. ArXiv, abs/2404.07972, 2024.
  98. 98.T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024.
  99. 99.M. Xing, R. Zhang, H. Xue, Q. Chen, F. Yang, and Z. Xiao. Understanding the weakness of large language model agents within a complex android environment. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 6061–6072, 2024.
  100. 100.F. F. Xu, Y. Song, B. Li, Y. Tang, K. Jain, M. Bao, Z. Z. Wang, X. Zhou, Z. Guo, M. Cao, M. Yang, H. Y. Lu, A. Martin, Z. Su, L. Maben, R. Mehta, W. Chi, L. Jang, Y. Xie, S. Zhou, and G. Neubig. Theagentcompany: Benchmarking llm agents on consequential real world tasks, 2024.
  101. 101.Z. Xu, X. Yang, Y. Wang, Q. Hu, Z. Wu, L. Wang, W. Luo, K. Zhang, B. Hu, and M. Zhang. Comfyui-copilot: An intelligent assistant for automated workflow development. arXiv preprint arXiv:2506.05010, 2025.
  102. 102.T. Xue, W. Qi, T. Shi, C. H. Song, B. Gou, D. Song, H. Sun, and Y. Su. An illusion of progress? assessing the current state of web agents. 2025.
  103. 103.J. Yang, C. E. Jimenez, A. Wettig, K. A. Lieret, S. Yao, K. Narasimhan, and O. Press. Swe-agent: Agent-computer interfaces enable automated software engineering. ArXiv, abs/2405.15793, 2024.
  104. 104.J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023.
  105. 105.S. Yao, H. Chen, J. Yang, and K. Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.
  106. 106.S. Yao, N. Shinn, P. Razavi, and K. Narasimhan. tau -bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024.
  107. 107.S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models, 2023.
  108. 108.S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  109. 109.O. Yoran, S. J. Amouyal, C. Malaviya, B. Bogin, O. Press, and J. Berant. Assistantbench: Can web agents solve realistic and time-consuming tasks?, 2024.
  110. 110.A. Zeng, M. Liu, R. Lu, B. Wang, X. Liu, Y. Dong, and J. Tang. Agenttuning: Enabling generalized agent abilities for llms, 2023.
  111. 111.Y. Zhang, Z. Ma, Y. Ma, Z. Han, Y. Wu, and V. Tresp. Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration, 2024.
  112. 112.Z. Zhang, E. Schoop, J. Nichols, A. Mahajan, and A. Swearngin. From interaction to impact: Towards safer ai agent through understanding and evaluating mobile ui operation impacts. In Proceedings of the 30th International Conference on Intelligent User Interfaces, pages 727–744, 2025.
  113. 113.Z. Zhang and A. Zhang. You only look at screens: Multimodal chain-of-action agents, 2024.
  114. 114.B. Zheng, J. Kil, H. Sun, Y. Su, and B. Gou. Gpt-4v(ision) is a generalist web agent, if grounded. ArXiv, abs/2401.01614, 2024.
  115. 115.S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig. Webarena: A realistic web environment for building autonomous agents, 2024.
  116. 116.Y. Zhu, A. Kellermann, D. Bowman, P. Li, A. Gupta, A. Danda, R. Fang, C. Jensen, E. Ihli, J. Benn, J. Geronimo, A. Dhir, S. Rao, K. Yu, T. Stone, and D. Kang. Cve-bench: A benchmark for ai agents’ ability to exploit real-world web application vulnerabilities, 2025.

Citation

MLA
Mozannar, H., et al. “Magentic-UI: Towards Human-in-the-loop Agentic Systems”. arXiv, 2025, http://arxiv.org/abs/2507.22358v1.
APA
Mozannar, H., Bansal, G., Tan, C., Fourney, A., Dibia, V., Chen, J., Gerrits, J., Payne, T., Maldaner, M. K., Grunde-McLaughlin, M., Zhu, E., Bassman, G., Alber, J., Chang, P., Loynd, R., Niedtner, F., Kamar, E., Murad, M., Hosn, R., & Amershi, S. (2025). Magentic-UI: Towards Human-in-the-loop Agentic Systems. arXiv. http://arxiv.org/abs/2507.22358v1
Chicago
Mozannar, H., G. Bansal, C. Tan, et al. 2025. “Magentic-UI: Towards Human-in-the-loop Agentic Systems”. arXiv. http://arxiv.org/abs/2507.22358v1.
Harvard
Mozannar, H. et al. (2025) “Magentic-UI: Towards Human-in-the-loop Agentic Systems”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.22358v1.
Vancouver
1. Mozannar H, Bansal G, Tan C, et al (2025) Magentic-UI: Towards Human-in-the-loop Agentic Systems. arXiv

BibTeX

@article{mozannar2025magentic,
  title = {Magentic-UI: Towards Human-in-the-loop Agentic Systems},
  author = {Mozannar, Hussein and Bansal, Gagan and Tan, Cheng and Fourney, Adam and Dibia, Victor and Chen, Jingya and Gerrits, Jack and Payne, Tyler and Maldaner, Matheus Kunzler and Grunde-McLaughlin, Madeleine and Zhu, Eric and Bassman, Griffin and Alber, Jacob and Chang, Peter and Loynd, Ricky and Niedtner, Friederike and Kamar, Ece and Murad, Maya and Hosn, Rafah and Amershi, Saleema},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.22358v1},
  eprint = {2507.22358}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission