keyword
graphical user interfaces
A graphical user interface is a visual system of interactive components that allows users to operate electronic devices through graphical elements such as icons, menus, buttons, and windows, rather than text-based command lines. Designed to improve usability and accessibility, graphical user interfaces allow users to directly manipulate digital objects and navigate software using input methods such as computer mice, keyboards, and touchscreens. By translating complex underlying system commands into intuitive visual representations, these interfaces serve as the primary medium for human interaction across personal computers, mobile operating systems, and various digital applications.
2 items

From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova
Why you should read this
Introduces Pix2Act, a visual web agent that relies strictly on raw screenshots and generic mouse and keyboard actions to complete instruction-following tasks, demonstrating for the first time that pixel-only models can outperform human crowdworkers on MiniWob++.
Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.
Added
2026-09-26

CogAgent: A Visual Language Model for GUI Agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, Jie Tang
Why you should read this
Introduces CogAgent, an 18-billion-parameter visual language model with high-resolution image processing that operates computer and smartphone graphical user interfaces directly from raw screenshots, outperforming HTML-based methods across PC and mobile benchmarks.
People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogAgent supports input at a resolution of 1120*1120, enabling it to recognize tiny page elements and text. As a generalist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK-VQA, Text-VQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks -- Mind2Web and AITW, advancing the state of the art. The model and codes are available at this https URL, with a new version of CogAgent-9B-20241220 available at this https URL.
Added
2026-09-25
