Built independently by an author, for readers. Read the story and support ChapterPal

keyword

graphical user interfaces

A graphical user interface is a visual system of interactive components that allows users to operate electronic devices through graphical elements such as icons, menus, buttons, and windows, rather than text-based command lines. Designed to improve usability and accessibility, graphical user interfaces allow users to directly manipulate digital objects and navigate software using input methods such as computer mice, keyboards, and touchscreens. By translating complex underlying system commands into intuitive visual representations, these interfaces serve as the primary medium for human interaction across personal computers, mobile operating systems, and various digital applications.

2 items

CogAgent: A Visual Language Model for GUI Agents

CogAgent: A Visual Language Model for GUI Agents

Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, Jie Tang

OrganizationsTsinghua UniversityZhipu AI

Why you should read this

Introduces CogAgent, an 18-billion-parameter visual language model with high-resolution image processing that operates computer and smartphone graphical user interfaces directly from raw screenshots, outperforming HTML-based methods across PC and mobile benchmarks.

People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogAgent supports input at a resolution of 1120*1120, enabling it to recognize tiny page elements and text. As a generalist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK-VQA, Text-VQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks -- Mind2Web and AITW, advancing the state of the art. The model and codes are available at this https URL, with a new version of CogAgent-9B-20241220 available at this https URL.

Added

2026-09-25