CogAgent: A Visual Language Model for GUI Agents
Wenyi HongWeihan WangQingsong LvJiazheng XuWenmeng YuJunhui JiYan WangZihan WangYuxiao DongMing Ding
Introduces CogAgent, an 18-billion-parameter visual language model with high-resolution image processing that operates computer and smartphone graphical user interfaces directly from raw screenshots, outperforming HTML-based methods across PC and mobile benchmarks.
Modern automation increasingly relies on interacting with software through graphical user interfaces across computers and mobile devices. However, standard language-based artificial intelligence models struggle to navigate these interfaces because underlying application code often lacks standard application programming interfaces, and visual elements like icons, charts, and canvas layouts cannot be parsed through text alone. Existing visual models also encounter severe computing bottlenecks when processing the high-resolution images required to read fine screen text. The article introduces CogAgent, an 18-billion-parameter visual language foundation model designed to accurately perceive, understand, and navigate graphical user interfaces using direct screen captures.
To overcome computational limits, the model employs a dual-branch architecture. It couples a standard low-resolution image encoder with a lightweight high-resolution cross-attention module that accepts inputs up to 1120×1120 pixels. The authors pre-trained the system using a curriculum of diverse text recognition datasets, visual grounding tasks, and an extensive collection of 400,000 web screenshots containing 140 million element pairs. The system was subsequently fine-tuned on real-world computer and smartphone interaction workflows alongside general visual reasoning tasks.
The evaluation shows that CogAgent achieves state-of-the-art performance on major graphical interface benchmarks. On web navigation benchmarks, it outperformed large text-based models using cleaned code inputs, beating a 70-billion-parameter baseline by 11.6% on cross-website tasks. On Android device navigation, CogAgent achieved an overall matching score of 76.88%, surpassing existing visual baselines. In broader visual question-answering tests, it led generalist models across five text-rich benchmarks, outperforming competitors by 16.2 points on document understanding and scoring 52.8 on complex integrated multimodal evaluations. Furthermore, the specialized cross-attention structure reduced computational operations by more than half compared to standard high-resolution visual architectures.
These findings show that visually grounded models can effectively automate digital workflows directly from screen pixels without relying on brittle, application-specific code representations. This visual-first approach reduces software engineering overhead while preserving computational efficiency during inference. However, practical deployment must account for identified failure modes, such as occasional coordinate inaccuracies, visual misinterpretations, and an inability to process multi-image sequences simultaneously. Additionally, manual review revealed that over 40% of recorded mobile navigation errors were actually valid alternative completion paths, indicating that future development should focus on dynamic virtual test environments rather than static evaluation datasets.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). Introduces the realistic web execution environment and benchmark for autonomous digital agents, providing key foundations for GUI agent navigation tasks.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Pioneers high-resolution vision-language modeling with spatial grounding and text reading, establishing core architectural paradigms adopted in specialized GUI agents.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). Establishes the standard Document Visual Question Answering benchmark used directly to evaluate CogAgent's high-resolution text-reading and visual understanding capabilities.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). Provides the ChartQA benchmark for visual and logical reasoning over data visualizations, serving as one of CogAgent's primary evaluation suites.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Introduces TextVQA to evaluate reading embedded visual text in images, directly motivating the high-resolution OCR perception modules in CogAgent.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). Creates the balanced VQA v2.0 benchmark that tests grounded visual perception without linguistic bias, featured prominently in CogAgent's generalist evaluation.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). Presents the comprehensive MME benchmark for evaluating perception and cognition in multimodal models, establishing baseline metrics used in generalist VLM evaluations.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). Demonstrates decision-tree planning and multi-step tool execution for open-source agents, framing the autonomous interaction mechanics utilized in digital environments.
- Paper: FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation, Kaiyi Huang et al. (2026). Advances vision-only GUI agent automation by incorporating structured inner monologues and explicit visual planning over raw screenshots across desktop and mobile platforms.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). Extends visual GUI agent action prediction by replacing expensive supervised fine-tuning with sample-efficient, rule-based reinforcement learning.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Generalizes high-resolution visual processing and device interaction through dynamic resolution mechanisms and multimodal rotary position embeddings.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Scales native multimodal architectures and multi-step agentic execution with cross-layer visual token injection and extended context reasoning.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). Expands open-source multimodal and agentic reasoning capabilities through systematic visual encoder scaling and progressive multi-task fine-tuning.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). Applies joint vision-language native pre-training paradigms to advance complex document comprehension and multidisciplinary reasoning in open multimodal agents.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). Builds upon vision-language modeling to develop predictive world models that enable dual-system reflective planning for visual action agents.
