Instruction Agent: Enhancing Agent with Expert Demonstration

Yinheng LiHailey HultquistJustin WagleKazuhito Koishida

article2025arXiv1 citations

Introduces Instruction Agent, a GUI automation framework that converts single expert demonstrations into verified, backtrackable execution steps, achieving a 60% success rate on complex OSWorld tasks that defeat all leading agents.

Listen

Autonomous digital agents designed to interact with graphical user interfaces (GUIs) using keyboard and mouse inputs have advanced rapidly, yet they continue to struggle with complex workflows. Existing agents frequently fail when handling non-intuitive visual elements, executing lengthy sequences of dependent actions, or following personalized user configurations. Because errors compound exponentially over long procedures, current systems fall well short of human reliability, creating a bottleneck for real-world digital automation.

The article demonstrates that leveraging a single human expert demonstration at inference time enables a GUI agent to reliably execute highly complex digital workflows without requiring additional model training or massive trajectory datasets. The researchers introduce Instruction Agent, a training-free framework that combines an Instructor module—which converts a recorded demonstration into detailed step-by-step instructions—with an Actor module that executes actions while utilizing built-in verification, grounding, and backtracking components.

The system was evaluated on a benchmark of operating system tasks (OSWorld). The researchers focused on 20 randomly sampled tasks that had caused 100% failure rates across the top three ranked open-source agents. Using Docker-hosted virtual environments, human annotators recorded baseline demonstrations, capturing input logs and screenshots to generate structured instructions for the agent.

The primary finding is that the Instruction Agent achieved a 60% success rate on tasks where leading baseline agents scored 0%, approaching the overall benchmark human performance level of 72.36%. Ablation experiments revealed that error-recovery mechanisms are critical to this performance: removing the backtracking module reduced the success rate to 45%, and removing both the verifier and backtracker reduced it to 40%. In a quarter of the test tasks, the agent successfully recovered from intermediate errors that would have otherwise caused complete task failure. Observed failures were primarily driven by visual grounding inaccuracies and subtle state changes that the verifier could not detect.

These findings indicate that complex digital automations can be achieved without the high computational cost and out-of-domain generalization limits of training large trajectory models. Instead, end users can record quick, single-instance demonstrations to reliably delegate idiosyncratic or long-horizon tasks. The framework lowers technical barriers, mitigates execution risk through active verification, and enables human workflows to be converted into reusable automation tools.

Organizations seeking to implement GUI automation should consider adopting demonstration-guided pipelines for complex or fragile tasks rather than relying purely on zero-shot autonomous planning. For workflows with high failure costs, engineering teams should implement explicit step verification and recovery buffers. Future work should focus on testing smaller language models suitable for local deployment and refining backtracking capabilities to handle severe interface divergences. While these results show high confidence across difficult tasks, stakeholders should note that the evaluation was conducted on a targeted sample of 20 benchmark tasks and still relies on external commercial model APIs.

arXiv: 2509.07098
Cover for Instruction Agent: Enhancing Agent with Expert Demonstration

Abstract

Graphical user interface (GUI) agents have advanced rapidly but still struggle with complex tasks involving novel UI elements, long-horizon actions, and personalized trajectories. In this work, we introduce Instruction Agent, a GUI agent that leverages expert demonstrations to solve such tasks, enabling completion of otherwise difficult workflows. Given a single demonstration, the agent extracts step-by-step instructions and executes them by strictly following the trajectory intended by the user, which avoids making mistakes during execution. The agent leverages the verifier and backtracker modules further to improve robustness. Both modules are critical to understand the current outcome from each action and handle unexpected interruptions(such as pop-up windows) during execution. Our experiments show that Instruction Agent achieves a 60% success rate on a set of tasks in OSWorld that all top-ranked agents failed to complete. The Instruction Agent offers a practical and extensible framework, bridging the gap between current GUI agents and reliable real-world GUI task automation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Problem Formulation
  • 3.2 Agent Architecture
  • 3.2.1 Instructor
  • 3.2.2 Actor
  • 4 Experiments
  • 4.1 Benchmark
  • 4.1.1 Task Selection
  • 4.1.2 Recording
  • 4.1.3 Agent Evaluation
  • 4.1.4 Results
  • 4.2 Ablation Studies
  • 4.3 Failure Analysis
  • 5 Broader Use Cases
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Instruction Format
  • A.2 Agent Trajectory

Knowls

  1. Knowl 1 — Instruction Agent Framework for Demonstration-Guided GUI Automation

    model/method

    The Instruction Agent is a training-free, test-time-only graphical user interface (GUI) automation framework that enables an agent to complete complex, long-horizon, or personalized computer tasks using a single human expert demonstration.

    The framework splits execution into two primary components:

    1. Instructor Model: Observes a human recording of the target task (consisting of screenshots, mouse movements/clicks, and keystrokes) and converts the captured trajectory into a sequence of natural language, step-by-step instructions annotated with target visual locations and expected functional outcomes.
    2. Actor Model: Sequentially executes the derived instructions in the live digital environment using visual grounding, automated code execution, visual verification of outcomes, and autonomous backtracking to recover from execution errors.

    By converting raw user demonstrations into grounded high-level plans at test time rather than fine-tuning models on large trajectory corpora, the framework avoids catastrophic execution drift in complex multi-step workflows while remaining accessible without specialized training resources.

  2. Knowl 2 — Instructor Trajectory Recording and Instruction Generation

    model/method

    The Instructor module translates a human user's ad-hoc demonstration into a structured sequence of executable natural language instructions through two submodules:

    1. Recorder: Captures a user trajectory consisting of raw user actions (mouse clicks, scroll events, keyboard keystrokes) along with visual screenshots. For each discrete input action ata_t, the recorder saves the screenshot immediately preceding the action (Ot−1O_{t-1}) and immediately following the action (OtO_t). The system does not rely on accessibility trees (a11y) or HTML DOM trees.
    2. Instruction Generator: Prompts a multimodal large language model (GPT-4o) using the input action log and the corresponding pre-action and post-action screenshots. For click operations, the physical pixel coordinates of the click are visually annotated onto the screenshot before prompting. To enhance grounding and verification, the model is prompted to output not just the low-level action but also the target element description and the explicit intended functional outcome (e.g., stating "click the Windows Start icon to open the Windows Start menu" rather than simply "click the Windows Start icon").
  3. Knowl 3 — Actor Execution Pipeline

    model/method

    The Actor executes the step-by-step instruction sequence s1,s2,…,sns_1, s_2, \dots, s_n generated by the Instructor using a modular pipeline composed of four subcomponents:

    1. UI Grounder: An open-source visual grounding model (UI-Tars 1.5 7B) receives the natural language instruction sis_i and the current screen observation Oi−1O_{i-1} to predict the exact bounding coordinates of the target interactive element.
    2. Executor: A multimodal LLM (GPT-4o) takes the grounder's coordinate prediction, the current state observation, and the instruction command to generate executable Python automation code using the pyautogui library.
    3. Verifier: A multimodal LLM (GPT-4o) takes the pre-action screenshot Oi−1O_{i-1} and post-action screenshot OiO_i to assess whether the executed action achieved the intended functional outcome specified in instruction sis_i.
    4. Backtracker: If the verifier flags a step as failed, the backtracker plans and executes recovery actions to return the application environment to the valid pre-step state Oi−1O_{i-1} before re-attempting execution.
  4. Knowl 4 — POMDP Formulation for Demonstration-Guided Digital Agents

    definition

    The digital GUI interaction problem is formulated as a partially observable Markov decision process (POMDP), defined as a tuple: M=(S,O,A,T,R)M = (S, O, A, T, R) where:

    • SS is the state space representing the underlying environment state of the operating system, applications, and web pages.
    • OO is the observation space consisting of rendered visual screenshots.
    • AA is the action space encompassing low-level input primitives (keyboard keystrokes and mouse actions).
    • T:S×A→ST : S \times A \to S is the environment state transition function.
    • R:S×A→RR : S \times A \to \mathbb{R} is the reward function measuring binary task success.

    In the Instruction Agent framework, the initial observation O0O_0 is augmented with the natural language step sequence generated by the Instructor module from the human demonstration sequence of action-state pairs ((a0,s0),(a1,s1),…,(ak,sk))((a_0, s_0), (a_1, s_1), \dots, (a_k, s_k)).

  5. Knowl 5 — Visual Verification and Memory-Buffered Backtracking Mechanism

    model/method

    To prevent failure cascade in sequential GUI execution, the Actor integrates closed-loop verification and state restoration:

    • Verification: After executing action aia_i for instruction step sis_i, the verifier model receives pre-action screenshot Oi−1O_{i-1}, post-action screenshot OiO_i, and the step's semantic goal. It outputs a binary decision indicating whether the intended interface state change occurred.
    • Backtracking Recovery: Upon verification failure, the backtracker is triggered to restore the environment to state Oi−1O_{i-1} before retrying the step. The backtracker compares the current deviated screenshot with the target stored screenshot Oi−1O_{i-1} and generates a sequence of corrective actions (such as closing unexpected pop-up windows, clicking the browser back button, or undoing text entries).
    • Memory Buffer: To prevent recovery loops, the backtracker maintains a memory buffer recording all prior failed attempts and encountered errors during recovery, prompting the agent to select alternative strategies when stuck. To guarantee termination, recovery attempts are capped at a predetermined retry budget.
  6. Knowl 6 — Performance on Previously Unsolved OSWorld Benchmark Tasks

    data/table

    The Instruction Agent was evaluated on a subset of 20 tasks randomly sampled from the 130 tasks in the OSWorld benchmark (out of 369 total) that all three top-performing open-source GUI agents on the OSWorld leaderboard failed to solve (0% success rate). Demonstrations were collected via Docker-hosted virtual machines.

    Agent Success Rate
    Instruction Agent (ours) 60%
    Human 72.36%
    UI-TARS-1.5 (100 steps) - rank 3 0%
    Agent S2 w/ Gemini 2.5 (50 steps) - rank 4 0%
    InfantAgent (50 steps) - rank 6 0%

    Instruction Agent achieves a 60% success rate on these previously unsolved tasks, approaching the 72.36% human benchmark performance reported across OSWorld.

  7. Knowl 7 — Ablation of Verification and Backtracking Modules

    data/table

    An ablation study evaluated the impact of the Verifier and Backtracker modules on the 20 sampled unsolved OSWorld tasks:

    Agent Variant Success Rate
    Instruction Agent (full) 60%
    - without backtracker 45%
    - without verifier and backtracker 40%

    Across the 20 tasks, 5 tasks required at least one action retry, and in 2 tasks, an incorrect action pushed the environment into a divergent state that required the backtracker to actively restore a valid prior state before retrying.

  8. Knowl 8 — Primary Failure Modes in Instruction-Guided GUI Agents

    limitation

    Execution failures in the Instruction Agent framework stem from four distinct sources:

    1. Grounding Errors: The visual grounding model (UI-Tars 1.5) occasionally fails to predict correct coordinate bounding boxes for target UI elements despite clear textual hints.
    2. Execution Errors: The LLM code generator (GPT-4o) occasionally produces flawed or syntactically invalid pyautogui Python scripts despite accurate coordinates and instructions.
    3. Verification Errors: The multimodal LLM verifier misidentifies step success/failure when encountering subtle UI animations, minor graphical changes, or novel UI widgets.
    4. Backtracking Errors: The backtracker agent fails when the divergence caused by a mistaken action is too severe to recover from via simple localized actions (e.g., irreversible multi-step state corruptions), or when the backtracker enters an unresolvable action loop.

Coverage note — None was omitted; all contributed methodology, problem formulations, experimental results, ablations, and failure analyses are represented.

References

  1. 1.Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s2: A compositional generalist-specialist framework for computer use agents, 2025. URL https://arxiv.org/abs/2504.00906.
  2. 2.Anthropic PBC. Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku. https://www.anthropic.com/news/3-5-models-and-computer-use, October 2024. Accessed 2 July 2025.
  3. 3.AxureBoutique. Navigating the maze: Examples of bad navigation in ui/ux. https://www.youtube.com/watch?v=D-RMsyZrt38, September 2023. YouTube video, accessed 2 July 2025.
  4. 4.Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, and Zack Hui. Windows agent arena: Evaluating multi-modal os agents at scale, 2024. URL https://arxiv.org/abs/2409.08264.
  5. 5.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents, 2024. URL https://arxiv.org/abs/2401.10935.
  6. 6.Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for gui agents, 2025. URL https://arxiv.org/abs/2410.05243.
  7. 7.Yanheng He, Jiahe Jin, and Pengfei Liu. Efficient agent training for computer use, 2025. URL https://arxiv.org/abs/2505.13909.
  8. 8.Zheng Hui, Yinheng Li, Dan zhao, Tianyi Chen, Colby Banbury, and Kazuhito Koishida. Winclick: Gui grounding with multimodal large language models, 2025. URL https://arxiv.org/abs/2503.04730.
  9. 9.Faria Huq, Zora Zhiruo Wang, Frank F. Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P. Bigham, and Graham Neubig. Cowpilot: A framework for autonomous and human-agent collaborative web navigation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), page 163–172. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.naacl-demo.17. URL http://dx.doi.org/10.18653/v1/2025.naacl-demo.17.
  10. 10.Lawrence Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, and Kazuhito Koishida. Videowebarena: Evaluating long context multimodal agents with video understanding web tasks, 2025. URL https://arxiv.org/abs/2410.19100.
  11. 11.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks, 2024. URL https://arxiv.org/abs/2401.13649.
  12. 12.Wei Li, William Bishop, Alice Li, Chris Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva. On the effects of data scale on ui control agents, 2024. URL https://arxiv.org/abs/2406.03679.
  13. 13.Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, Jacob Alber, Peter Chang, Ricky Loynd, Friederike Niedtner, Ece Kamar, Maya Murad, Rafah Hosn, and Saleema Amershi. Magentic-ui: Towards human-in-the-loop agentic systems, 2025. URL https://arxiv.org/abs/2507.22358.
  14. 14.Shikhar Murty, Hao Zhu, Dzmitry Bahdanau, and Christopher D. Manning. Nnetnav: Unsupervised learning of browser agents through environment interaction in the wild, 2025. URL https://arxiv.org/abs/2410.02907.
  15. 15.Juho-Jaakko Oksanen. Test automation for windows gui application. Bachelor’s thesis, Oulu University of Applied Sciences, Oulu, Finland, 2023. URL https://www.theseus.fi/handle/10024/801926.
  16. 16.OpenAI. Introducing operator. https://openai.com/index/introducing-operator/, January 2025. Accessed 2 July 2025.
  17. 17.Tianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou. Synatra: Turning indirect knowledge into direct demonstrations for digital agents at scale, 2024. URL https://arxiv.org/abs/2409.15637.
  18. 18.Arnold Overwijk, Chenyan Xiong, Xiao Liu, Cameron VandenBerg, and Jamie Callan. Clueweb22: 10 billion web documents with visual and semantic information, 2022. URL https://arxiv.org/abs/2211.15848.
  19. 19.Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Yang, Haifeng Liu, Feng Lin, Tao Peng, Xin Liu, and Guang Shi. Ui-tars: Pioneering automated gui interaction with native agents, 2025. URL https://arxiv.org/abs/2501.12326.
  20. 20.Anian Ruoss, Fabio Pardo, Harris Chan, Bonnie Li, Volodymyr Mnih, and Tim Genewein. Lmact: A benchmark for in-context imitation learning with long multimodal demonstrations, 2025. URL https://arxiv.org/abs/2412.01441.
  21. 21.Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. Os-genesis: Automating gui agent trajectory construction via reverse task synthesis, 2025. URL https://arxiv.org/abs/2412.19723.
  22. 22.Preeti Tupsakhare. Python for automation and scripting: Streamlining operations and increasing efficiency. The Journal of Scientific and Engineering Research, pages 222–227, 09 2019. doi: 10.5281/zenodo.13918609.
  23. 23.Shuai Wang, Weiwen Liu, Jingxuan Chen, Yuqi Zhou, Weinan Gan, Xingshan Zeng, Yuhan Che, Shuai Yu, Xinlong Hao, Kun Shao, Bin Wang, Chuhan Wu, Yasheng Wang, Ruiming Tang, and Jianye Hao. Gui agents with foundation models: A comprehensive survey, 2025. URL https://arxiv.org/abs/2411.04890.
  24. 24.Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. Os-atlas: A foundation action model for generalist gui agents, 2024. URL https://arxiv.org/abs/2410.23218.
  25. 25.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024. URL https://arxiv.org/abs/2404.07972.
  26. 26.Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, and Tao Yu. Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials, 2025. URL https://arxiv.org/abs/2412.09605.
  27. 27.Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Qingwei Lin Kang, Yu and, Saravan Rajmohan, Dongmei Zhang, and Qi Zhang. Ufo: A ui-focused agent for windows os interaction. In NAACL’25, pages 597–622, April 2025. URL https://www.microsoft.com/en-us/research/publication/ufo-a-ui-focused-agent-for-windows-os-interaction/.

Citation

MLA
Li, Y., et al. “Instruction Agent: Enhancing Agent with Expert Demonstration”. arXiv, 2025, http://arxiv.org/abs/2509.07098v1.
APA
Li, Y., Hultquist, H., Wagle, J., & Koishida, K. (2025). Instruction Agent: Enhancing Agent with Expert Demonstration. arXiv. http://arxiv.org/abs/2509.07098v1
Chicago
Li, Y., H. Hultquist, J. Wagle, and K. Koishida. 2025. “Instruction Agent: Enhancing Agent with Expert Demonstration”. arXiv. http://arxiv.org/abs/2509.07098v1.
Harvard
Li, Y. et al. (2025) “Instruction Agent: Enhancing Agent with Expert Demonstration”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2509.07098v1.
Vancouver
1. Li Y, Hultquist H, Wagle J, Koishida K (2025) Instruction Agent: Enhancing Agent with Expert Demonstration. arXiv

BibTeX

@article{li2025instruction,
  title = {Instruction Agent: Enhancing Agent with Expert Demonstration},
  author = {Li, Yinheng and Hultquist, Hailey and Wagle, Justin and Koishida, Kazuhito},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2509.07098v1},
  eprint = {2509.07098}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/