From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
Peter ShawMandar JoshiJames CohanJonathan BerantPanupong PasupatHexiang HuUrvashi KhandelwalKenton LeeKristina Toutanova
Introduces Pix2Act, a visual web agent that relies strictly on raw screenshots and generic mouse and keyboard actions to complete instruction-following tasks, demonstrating for the first time that pixel-only models can outperform human crowdworkers on MiniWob++.
Digital agents designed to complete tasks via graphical user interfaces (GUIs) have traditionally depended on structured underlying text, such as Document Object Model (DOM) trees or HTML code, paired with specialized, task-specific actions. However, these structured representations are frequently unavailable, heavily scripted, or misaligned with actual visual displays. Humans operate across diverse digital interfaces without accessing underlying code, using only visual perception and standard inputs like mouse clicks and keystrokes. Developing agents capable of navigating software using purely visual inputs and generic controls is essential for broader automation, accessibility, and digital assistance across diverse, real-world computing environments.
The article evaluates whether an automated agent can successfully interpret and complete multi-step digital tasks using only raw, pixel-level screenshot observations and low-level mouse and keyboard actions. It demonstrates the feasibility and effectiveness of this approach across standard web-based benchmark environments.
To accomplish this, the authors introduced an agent named PIX2ACT, built on a 282-million parameter Vision Transformer architecture pre-trained to parse web screenshots. The system receives visual screen captures with rendered instructions and generates textual tokens corresponding to discrete mouse movements, clicks, drags, scrolling, and keyboard actions. The training methodology combined behavioral cloning on human demonstrations with policy refinement using Monte Carlo Tree Search. The evaluation was conducted within a browser framework across 59 tasks from the MiniWob++ benchmark and an adapted version of the WebShop shopping benchmark.
The findings show that PIX2ACT achieves an average score of 96.2 out of 100 on the MiniWob++ benchmark, outperforming human crowdworkers (94.4) and matching state-of-the-art models (96.3) that have direct access to internal DOM structures. Pre-training on screenshot parsing proved vital; without it, performance dropped sharply from 66.5 to 17.1 on MiniWob++ and from 46.7 to 1.1 on WebShop under behavioral cloning alone. Furthermore, policy refinement through tree search improved the agent's greedy policy performance from 66.5 to 96.2. On the WebShop benchmark, PIX2ACT established the first visual-only baseline score of 46.7, though a gap remains compared to language models utilizing HTML code (67.5). Finally, the agent demonstrated zero-shot transfer capability, achieving a score of 28.3 on completely unseen interface tasks compared to 7.6 without pre-training.
These results establish that automated agents do not require internal software code to operate digital interfaces effectively. By interacting visually, systems can bypass technical hurdles like code obfuscation and sandboxing, enabling universal compatibility across software ecosystems. However, closing the remaining performance gap on complex, text-heavy tasks will likely require larger multimodal models or enhanced visual scaling comparable to recent text-based language models.
Before deploying pixel-based agents in real-world online services, organizations must address safety, abuse, and compliance risks. Because visual agents interact like human users, they could potentially bypass standard security defenses, create spam, or interact improperly with external platforms. Deployments require rigorous behavioral guardrails, adherence to terms of service, and protections against privacy risks associated with processing screen captures. Further research should focus on refining reward modeling for environments where deterministic task resets are unavailable.
Confidence in these findings is high for structured, deterministic browser tasks, as evidenced by consistent performance across extensive evaluation seeds. However, key limitations remain: the current system excludes tasks requiring complex real-time animations or specific interactions like manual text-highlighting drag actions. Additionally, the policy improvement mechanism relied on deterministic environment resets and automated reward signals, which may not readily exist in arbitrary production applications without supplementary reward modeling.
- Paper: WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents, Shunyu Yao et al. (2022). This paper establishes the WebShop shopping benchmark environment used directly by the source paper to evaluate visual agents against language models operating over web structures.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It introduces the Vision Transformer architecture upon which PIX2ACT's underlying screenshot-parsing model is fundamentally built.
- Paper: Multi-Game Decision Transformers, Kuang-Huei Lee et al. (2022). It pioneers modeling multi-game decision-making directly from visual inputs as sequence prediction with Transformers, laying groundwork for pixel-to-action agent modeling.
- Paper: Playing Atari with Deep Reinforcement Learning, Volodymyr Mnih et al. (2013). It introduces the foundational concept of learning action policies directly from raw pixel observations without relying on underlying environment code.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). It extends visual GUI automation by designing a high-resolution visual language foundation model specifically tailored for screen parsing and GUI navigation across web and mobile platforms.
- Paper: FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation, Kaiyi Huang et al. (2026). It builds upon vision-only GUI interaction principles to establish an autonomous agent framework incorporating structured visual grounding and explicit inner-monologue planning.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). It generalizes web agent evaluation from synthetic environments to thousands of realistic, open-ended tasks across real-world websites.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). It scales evaluation of instruction-following agents to a comprehensive, realistic multi-domain web environment measured by end-to-end functional correctness.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). It advances action prediction in visual GUI agents by applying rule-based reinforcement learning with group relative policy optimization to achieve sample-efficient generalization.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). It enhances multi-step digital agents by introducing procedural memory mechanisms to retain and reuse successful action workflows across web navigation tasks.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). It applies hierarchical planning and reflection-guided tree search to address complex dynamic web navigation tasks, extending the search-based policy refinement concepts used in PIX2ACT.
- Paper: AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, Ori Yoran et al. (2024). It benchmarks the limits of realistic multi-page browsing and planning agents on time-consuming open-web tasks beyond isolated sandboxes.
