Instruction Agent: Enhancing Agent with Expert Demonstration
Yinheng LiHailey HultquistJustin WagleKazuhito Koishida
Introduces Instruction Agent, a GUI automation framework that converts single expert demonstrations into verified, backtrackable execution steps, achieving a 60% success rate on complex OSWorld tasks that defeat all leading agents.
Autonomous digital agents designed to interact with graphical user interfaces (GUIs) using keyboard and mouse inputs have advanced rapidly, yet they continue to struggle with complex workflows. Existing agents frequently fail when handling non-intuitive visual elements, executing lengthy sequences of dependent actions, or following personalized user configurations. Because errors compound exponentially over long procedures, current systems fall well short of human reliability, creating a bottleneck for real-world digital automation.
The article demonstrates that leveraging a single human expert demonstration at inference time enables a GUI agent to reliably execute highly complex digital workflows without requiring additional model training or massive trajectory datasets. The researchers introduce Instruction Agent, a training-free framework that combines an Instructor module—which converts a recorded demonstration into detailed step-by-step instructions—with an Actor module that executes actions while utilizing built-in verification, grounding, and backtracking components.
The system was evaluated on a benchmark of operating system tasks (OSWorld). The researchers focused on 20 randomly sampled tasks that had caused 100% failure rates across the top three ranked open-source agents. Using Docker-hosted virtual environments, human annotators recorded baseline demonstrations, capturing input logs and screenshots to generate structured instructions for the agent.
The primary finding is that the Instruction Agent achieved a 60% success rate on tasks where leading baseline agents scored 0%, approaching the overall benchmark human performance level of 72.36%. Ablation experiments revealed that error-recovery mechanisms are critical to this performance: removing the backtracking module reduced the success rate to 45%, and removing both the verifier and backtracker reduced it to 40%. In a quarter of the test tasks, the agent successfully recovered from intermediate errors that would have otherwise caused complete task failure. Observed failures were primarily driven by visual grounding inaccuracies and subtle state changes that the verifier could not detect.
These findings indicate that complex digital automations can be achieved without the high computational cost and out-of-domain generalization limits of training large trajectory models. Instead, end users can record quick, single-instance demonstrations to reliably delegate idiosyncratic or long-horizon tasks. The framework lowers technical barriers, mitigates execution risk through active verification, and enables human workflows to be converted into reusable automation tools.
Organizations seeking to implement GUI automation should consider adopting demonstration-guided pipelines for complex or fragile tasks rather than relying purely on zero-shot autonomous planning. For workflows with high failure costs, engineering teams should implement explicit step verification and recovery buffers. Future work should focus on testing smaller language models suitable for local deployment and refining backtracking capabilities to handle severe interface divergences. While these results show high confidence across difficult tasks, stakeholders should note that the evaluation was conducted on a targeted sample of 20 benchmark tasks and still relies on external commercial model APIs.
- Paper: Large Language Model-Brained GUI Agents: A Survey, Chaoyun Zhang et al. (2025). Provides a comprehensive architectural foundation and overview of LLM- and VLM-based graphical user interface agents across desktop, web, and mobile environments.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). Establishes core visual language modeling techniques for parsing high-resolution screen captures and predicting GUI actions directly from visual inputs.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). Introduces foundational methods for learning to execute step-by-step UI actions directly from raw visual pixels using behavioral cloning on demonstrations.
- Paper: Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration, Junyang Wang et al. (2024). Demonstrates multi-agent reflection and visual state-change monitoring to handle error correction and navigation over long interaction trajectories.
- Paper: Instruction Induction: From Few Examples to Natural Language Task Descriptions, Or Honovich et al. (2023). Establishes how large language models can induce explicit natural language instructions from few-shot input-output examples, motivating trajectory instruction extraction.
- Paper: A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning, Stephane Ross et al. (2010). Provides the foundational theoretical framework for mitigating cascading errors and compounding distribution shifts in imitation learning from expert demonstrations.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Extends demonstration-guided instruction following by abstracting expert trajectories into persistent, reusable procedural workflow memories across tasks.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). Generalizes agent trajectory utilization by constructing self-adaptive demonstration datasets autonomously from interaction histories on benchmarks like OSWorld.
- Paper: Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking, Jiwan Chung et al. (2026). Offers a step-level semantic state tracking methodology to evaluate exactly where GUI and web agents fail along long-horizon execution trajectories.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). Broadens visual GUI execution beyond standard click-and-type operations by scaling continuous, multi-point human action spaces in complex desktop applications.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). Applies procedural execution memories distilled from multi-step workflow logs into coordinated multi-agent office automation systems.
