CogAgent: A Visual Language Model for GUI Agents

Wenyi HongWeihan WangQingsong LvJiazheng XuWenmeng YuJunhui JiYan WangZihan WangYuxiao DongMing Ding

article2024CVPR873 citationsCVPR 2024 Highlight

Introduces CogAgent, an 18-billion-parameter visual language model with high-resolution image processing that operates computer and smartphone graphical user interfaces directly from raw screenshots, outperforming HTML-based methods across PC and mobile benchmarks.

Listen

Modern automation increasingly relies on interacting with software through graphical user interfaces across computers and mobile devices. However, standard language-based artificial intelligence models struggle to navigate these interfaces because underlying application code often lacks standard application programming interfaces, and visual elements like icons, charts, and canvas layouts cannot be parsed through text alone. Existing visual models also encounter severe computing bottlenecks when processing the high-resolution images required to read fine screen text. The article introduces CogAgent, an 18-billion-parameter visual language foundation model designed to accurately perceive, understand, and navigate graphical user interfaces using direct screen captures.

To overcome computational limits, the model employs a dual-branch architecture. It couples a standard low-resolution image encoder with a lightweight high-resolution cross-attention module that accepts inputs up to 1120×1120 pixels. The authors pre-trained the system using a curriculum of diverse text recognition datasets, visual grounding tasks, and an extensive collection of 400,000 web screenshots containing 140 million element pairs. The system was subsequently fine-tuned on real-world computer and smartphone interaction workflows alongside general visual reasoning tasks.

The evaluation shows that CogAgent achieves state-of-the-art performance on major graphical interface benchmarks. On web navigation benchmarks, it outperformed large text-based models using cleaned code inputs, beating a 70-billion-parameter baseline by 11.6% on cross-website tasks. On Android device navigation, CogAgent achieved an overall matching score of 76.88%, surpassing existing visual baselines. In broader visual question-answering tests, it led generalist models across five text-rich benchmarks, outperforming competitors by 16.2 points on document understanding and scoring 52.8 on complex integrated multimodal evaluations. Furthermore, the specialized cross-attention structure reduced computational operations by more than half compared to standard high-resolution visual architectures.

These findings show that visually grounded models can effectively automate digital workflows directly from screen pixels without relying on brittle, application-specific code representations. This visual-first approach reduces software engineering overhead while preserving computational efficiency during inference. However, practical deployment must account for identified failure modes, such as occasional coordinate inaccuracies, visual misinterpretations, and an inability to process multi-image sequences simultaneously. Additionally, manual review revealed that over 40% of recorded mobile navigation errors were actually valid alternative completion paths, indicating that future development should focus on dynamic virtual test environments rather than static evaluation datasets.

Cover for CogAgent: A Visual Language Model for GUI Agents

Abstract

People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogAgent supports input at a resolution of 1120*1120, enabling it to recognize tiny page elements and text. As a generalist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK-VQA, Text-VQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks -- Mind2Web and AITW, advancing the state of the art. The model and codes are available at this https URL, with a new version of CogAgent-9B-20241220 available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Architecture
  • 2.2 High-Resolution Cross-Module
  • 2.3 Pre-training
  • 2.4 Multi-task Fine-tuning and Alignment
  • 3 Experiments
  • 3.1 Foundational Visual Understanding
  • 3.2 GUI Agent: Computer Interface
  • 3.3 GUI Agent: Smartphone Interface
  • 4 Ablation Study
  • 4.1 Model Architecture
  • 4.2 Pre-train Data
  • 5 Conclusion
  • References
  • 1 Details of Training Configurations
  • 2 Details of Evaluation Datasets
  • 2.1 General VQA
  • 2.2 Text-rich VQA
  • 2.3 GUI Agent
  • 3 Derivation of Acceleration for High-Resolution Cross-Module
  • 4 Performance Analysis on AITW
  • 5 Samples of Pre-train Data
  • 6 Details of Fine-Tuning Data
  • 6.1 Human annotation
  • 6.2 Conversion of Agent Datasets
  • 7 Failure cases
  • 8 More Generated Samples of CogAgent

Knowls

  1. Knowl 1 — CogAgent Dual-Branch High-Resolution Visual-Language Architecture

    model/method

    CogAgent is an 18-billion-parameter visual language model designed for graphical user interface (GUI) understanding and navigation. The architecture separates visual feature extraction into two parallel branches operating at different resolutions:

    1. Low-Resolution Base Branch: Employs an EVA2-CLIP-E visual encoder (4.4B parameters) that processes low-resolution images (224×224224 \times 224 pixels) patchified into 14×1414 \times 14 pixel patches (LIlo=256L_{I_{\text{lo}}} = 256 visual tokens). An MLP adapter projects these features into the embedding space of a visual-language decoder (Vicuna-1.5-7B enhanced with visual expert modules, hidden dimension Ddec=4096D_{\text{dec}} = 4096, 32 layers, 32 attention heads). The decoder concatenates low-resolution visual tokens with input text tokens (LTL_T) and applies multi-head self-attention with visual expert parameters.

    2. High-Resolution Cross-Module Branch: Employs a lightweight EVA2-CLIP-L visual encoder (0.30B parameters) that processes high-resolution images (1120×11201120 \times 1120 pixels) patchified into 14×1414 \times 14 pixel patches (LIhi=6400L_{I_{\text{hi}}} = 6400 visual tokens). This branch captures fine-grained visual details such as text and tiny UI elements using a reduced feature hidden dimension (Dcross=1024D_{\text{cross}} = 1024, 32 cross-attention heads with dimension dcross=32d_{\text{cross}} = 32).

    In every layer of the decoder, a multi-head cross-attention layer is inserted after the self-attention layer, allowing hidden states from the low-resolution and text sequence to query high-resolution visual tokens via residual connections without quadratic token concatenation overhead.

  2. Knowl 2 — Layer-Wise Multi-Head Cross-Attention Formulation in CogAgent

    equation

    Let Xini∈RB×(LIlo+LT)×DdecX_{\text{in}_i} \in \mathbb{R}^{B \times (L_{I_{\text{lo}}} + L_T) \times D_{\text{dec}}} denote the input hidden states to the ii-th layer of the visual language decoder, where BB is batch size, LIloL_{I_{\text{lo}}} is low-resolution visual token length, LTL_T is text token length, and DdecD_{\text{dec}} is decoder hidden dimension. Let Xhi∈RB×LIhi×DhiX_{\text{hi}} \in \mathbb{R}^{B \times L_{I_{\text{hi}}} \times D_{\text{hi}}} denote the output representations of the high-resolution image encoder, where LIhiL_{I_{\text{hi}}} is high-resolution token length and DhiD_{\text{hi}} is the encoder's feature dimension.

    The attention operations at the ii-th decoder layer are formulated as:

    Xi′=MSA(layernorm(Xini))+XiniX'_i = \text{MSA}(\text{layernorm}(X_{\text{in}_i})) + X_{\text{in}_i}

    Xouti=MCA(layernorm(Xi′),Xhi)+Xi′X_{\text{out}_i} = \text{MCA}(\text{layernorm}(X'_i), X_{\text{hi}}) + X'_i

    where MSA\text{MSA} denotes multi-head self-attention with visual expert modules and MCA\text{MCA} denotes multi-head cross-attention. The cross-attention module uses learnable weight matrices WKcrossi,WVcrossi∈RDhi×DcrossW_{K\text{cross}}^i, W_{V\text{cross}}^i \in \mathbb{R}^{D_{\text{hi}} \times D_{\text{cross}}} and WQcrossi∈RDdec×DcrossW_{Q\text{cross}}^i \in \mathbb{R}^{D_{\text{dec}} \times D_{\text{cross}}} to compute keys, values, and queries:

    Kcrossi=XhiWKcrossi∈RB×LIhi×DcrossK_{\text{cross}}^i = X_{\text{hi}} W_{K\text{cross}}^i \in \mathbb{R}^{B \times L_{I_{\text{hi}}} \times D_{\text{cross}}}

    Vcrossi=XhiWVcrossi∈RB×LIhi×DcrossV_{\text{cross}}^i = X_{\text{hi}} W_{V\text{cross}}^i \in \mathbb{R}^{B \times L_{I_{\text{hi}}} \times D_{\text{cross}}}

    Qcrossi=Xi′WQcrossi∈RB×(LIlo+LT)×DcrossQ_{\text{cross}}^i = X'_i W_{Q\text{cross}}^i \in \mathbb{R}^{B \times (L_{I_{\text{lo}}} + L_T) \times D_{\text{cross}}}

  3. Knowl 3 — Attention Computational Complexity and Speedup Bound of Dual-Branch Cross-Module

    theoretical result

    Let LIloL_{I_{\text{lo}}} and LIhiL_{I_{\text{hi}}} be the token lengths of low-resolution and high-resolution images, LTL_T the text token length, HdecH_{\text{dec}} and HcrossH_{\text{cross}} the head counts in self-attention and cross-attention, and ddec=Ddec/Hdecd_{\text{dec}} = D_{\text{dec}} / H_{\text{dec}} and dcross=Dcross/Hcrossd_{\text{cross}} = D_{\text{cross}} / H_{\text{cross}} the head dimensions.

    The computational complexity of attention in the dual-branch cross-module architecture is:

    Timproved=O((LIlo+LT)LIhiHcrossdcross+(LIlo+LT)2Hdecddec)T_{\text{improved}} = \mathcal{O}\left((L_{I_{\text{lo}}} + L_T) L_{I_{\text{hi}}} H_{\text{cross}} d_{\text{cross}} + (L_{I_{\text{lo}}} + L_T)^2 H_{\text{dec}} d_{\text{dec}}\right)

    In contrast, directly inputting high-resolution images into the single-branch self-attention decoder incurs complexity:

    Toriginal=O((LIhi+LT)2Hdecddec)T_{\text{original}} = \mathcal{O}\left((L_{I_{\text{hi}}} + L_T)^2 H_{\text{dec}} d_{\text{dec}}\right)

    For the parameter configuration Hcross=32H_{\text{cross}} = 32, dcross=32d_{\text{cross}} = 32, Hdec=32H_{\text{dec}} = 32, ddec=128d_{\text{dec}} = 128, LIhi=6400L_{I_{\text{hi}}} = 6400, and LIlo=256L_{I_{\text{lo}}} = 256, the reduction ratio satisfies the strict lower bound:

    ToriginalTimproved=6400+LT256+LT⋅4(6400+LT)6400+4(256+LT)>6400+LT256+LT\frac{T_{\text{original}}}{T_{\text{improved}}} = \frac{6400 + L_T}{256 + L_T} \cdot \frac{4(6400 + L_T)}{6400 + 4(256 + L_T)} > \frac{6400 + L_T}{256 + L_T}

    For sequence lengths LT≤512L_T \le 512, this architecture achieves greater than a 25×25\times reduction in attention computational overhead compared to processing high-resolution visual tokens via standard self-attention.

  4. Knowl 4 — Pre-Training Data Composition and GUI Grounding Task Formulation

    model/method

    To train CogAgent for GUI comprehension, a multi-source pre-training dataset is assembled across three primary categories:

    1. Text Recognition Datasets (107M total):

      • Synthetic Text Renderings (80M): Text varying across fonts, sizes, colors, and orientations overlaid on diverse natural image backgrounds from LAION-2B.
      • Natural Image OCR (18M): Natural images from COYO and LAION-2B filtered using Paddle-OCR to retain images with extracted text bounding boxes, augmented with rotation and flipping.
      • Academic Documents (9M): Extracted arXiv LaTeX documents containing text, mathematical formulas, and tables rendered into images following Nougat augmentation.
    2. Visual Grounding Dataset (40M): Image-caption pairs sampled from LAION-115M with noun phrases mapped to bounding boxes formatted as [[x0, y0, x1, y1]], where coordinates represent upper-left and lower-right corners normalized to [000,999][000, 999].

    3. GUI Imagery and Grounding (CCS400K Dataset): 400,000 web page screenshots gathered by rendering URLs from Common Crawl via Playwright across multiple screen resolutions. Two GUI grounding tasks provide 140 million QA pairs with redundant DOM attributes cleaned:

      • GUI Referring Expression Generation (REG): Generating HTML DOM element code corresponding to a designated screen region.
      • GUI Referring Expression Comprehension (REC): Predicting normalized bounding box coordinates for a specified DOM element.
  5. Knowl 5 — Pre-Training Curriculum and Multi-Task Alignment Protocol

    experimental setup

    CogAgent undergoes a two-phase training protocol:

    1. Pre-Training Phase:

      • Optimization: 60,000 total iterations, batch size of 4,608, learning rate of 2×10−52 \times 10^{-5} with cosine decay, 500 warmup steps, weight decay of 0.05, dropout of 0.1, and Adam optimizer parameters (β1=0.9,β2=0.95,ϵ=1×10−5)(\beta_1=0.9, \beta_2=0.95, \epsilon=1 \times 10^{-5}).
      • Parameter Freezing: For the first 20,000 steps, all parameters are frozen except the high-resolution cross-module (646M trainable parameters / 3.5% of total). For the subsequent 40,000 steps, the visual expert module within the decoder is unfrozen.
      • Curriculum Strategy: Training begins with easier synthetic text rendering, natural image OCR, and image captioning, followed by academic documents, visual grounding, and CCS400K web page data.
    2. Multi-Task Fine-Tuning and Alignment Phase:

      • Optimization: 10,000 iterations, batch size of 1,024, learning rate of 2×10−52 \times 10^{-5}, with all model parameters unfrozen.
      • Data: Over 2,000 human-annotated mobile and PC screenshots (annotated with 5 buttons, 3 clickable areas, 2 visual QA questions, and 1 grounded operation requirement), general VQA datasets, and Mind2Web and AITW trajectories converted into a structured JSON action format ({"plan": "...", "action": "...", "operation": "..."}) using GPT-4.
  6. Knowl 6 — GUI Navigation Performance on Computer (Mind2Web) and Mobile (AITW) Benchmarks

    empirical result

    CogAgent was evaluated as an end-to-end visual GUI agent using only raw screenshot images without DOM tree or OCR auxiliary inputs:

    1. Mind2Web (Web Navigation): Evaluated on step success rate (Step SR) across out-of-domain splits using top-50 candidate element selection. CogAgent achieved 62.3% on Cross-Task, 54.0% on Cross-Website, 59.4% on Cross-Domain, and 58.2% Overall. This outperforms fine-tuned HTML-based language models including LLaMA2-70B (54.4% overall, 55.8% cross-task, 51.6% cross-website, 55.7% cross-domain) and Flan-T5-XL (43.5% overall), as well as visual models including CogVLM (23.9% overall) and Qwen-VL (10.2% overall).

    2. Android in the Wild / AITW (Smartphone Navigation): Evaluated across five task subsets using action matching score. A unified CogAgent model trained across subsets achieved 74.95% on GoogleApps, 78.86% on Install, 71.73% on WebShopping, 65.38% on General, 93.49% on Single, and 76.88% Overall. CogAgent outperforms the unified visual baseline Auto-UI (74.27% overall) and textual description baselines such as LLaMA2-7B (28.40% overall).

    Manual inspection of divergent CogAgent predictions on AITW revealed that in 42% of discrepancy cases, the model executed a valid alternative execution pathway rather than an error.

  7. Knowl 7 — Evaluation on General and Text-Rich Visual Question Answering Benchmarks

    data/table

    CogAgent was fine-tuned simultaneously on multiple VQA benchmarks to produce a single generalist model. The table below compares CogAgent with task-specific models and generalist multimodal models across general VQA (VQAv2, OK-VQA) and text-rich VQA (OCR-VQA, TextVQA, ST-VQA, ChartQA, InfoVQA, DocVQA):

    Method VQAv2 OK-VQA OCR-VQA TextVQA ST-VQA ChartQA InfoVQA DocVQA
    Task-specific
    Pix2Struct - - - - - 58.6 40.0 76.6
    BLIP-2 82.2 59.3 72.7 - - - - -
    PaLI-X-55B 86.0 66.1 75.0 71.4 79.9 70.9 49.2 80.0
    CogVLM 84.7 64.7 74.5 69.7 - - - -
    Generalist
    UReader - 57.6 - - - 59.3 42.2 65.4
    Qwen-VL 79.5 58.6 75.7 63.8 - 65.7 - 65.1
    Qwen-VL-chat 78.2 56.6 70.5 61.5 - 66.3 - 62.6
    LLaVA-1.5 80.0 - - 61.5 - - - -
    Fuyu-8B 74.2 60.6 - - - - - -
    CogVLM 83.4 58.9 74.1 68.1 - - - -
    CogAgent 83.7 61.2 75.0 76.1 80.5 68.4 44.5 81.6

    CogAgent achieves state-of-the-art generalist performance on all tested text-rich benchmarks, exceeding CogVLM by +8.0 on TextVQA and +13.5 on DocVQA, and outperforming specialized models on TextVQA (76.1 vs. 71.4), ST-VQA (80.5 vs. 79.9), and DocVQA (81.6 vs. 80.0).

  8. Knowl 8 — Zero-Shot Multimodal Reasoning (MM-Vet) and Hallucination Resistance (POPE)

    empirical result

    CogAgent was evaluated in a zero-shot setting on MM-Vet (evaluating integrated multimodal reasoning across recognition, OCR, knowledge, language generation, spatial awareness, and math using GPT-4 evaluation) and the POPE adversarial benchmark (polling-based object probing evaluation of vision-language hallucinations):

    1. MM-Vet Score: CogAgent achieved a score of 52.8 using a Vicuna-7B base LLM. This outperforms prior open-source vision-language models evaluated with Vicuna-13B backbones, including LLaVA-1.5 (36.3), Emu (36.3), DreamLLM (35.9), InstructBLIP (25.6), Otter (24.7), MiniGPT-4 (24.4), and BLIP-2 (22.4).

    2. POPE Adversarial F1 Score: CogAgent achieved an F1 score of 85.9 under the adversarial polling condition, outperforming LLaVA-1.5 (84.5), InstructBLIP (77.3), DreamLLM (76.5), MiniGPT-4 (70.4), and LLaVA (66.3), indicating reduced susceptibility to object hallucination in complex visual scenes.

  9. Knowl 9 — Ablation of Architectural Efficiency and Resolution Scaling

    data/table

    The impact of the high-resolution cross-module on computational cost and downstream accuracy was assessed across image resolutions (224,490,756,1120224, 490, 756, 1120), comparing the standard CogVLM single-branch self-attention architecture with CogAgent's dual-branch cross-attention architecture:

    High-Res Base Cross STVQA OCRVQA DocVQA Mind2Web Training Time Forward
    Module Res Res ANLS EM ANLS Step SR (s/iter) TFLOPs
    No 224 — 48.0 70.2 28.6 34.6 2.36 7.77
    No 490 — 68.1 74.5 57.6 40.7 6.43 29.14
    Yes 224 756 73.6 74.2 62.3 40.7 3.57 10.08
    Yes 224 1120 78.2 75.9 74.1 41.4 5.17 12.56

    Directly scaling the original architecture to resolution 1120 requires 143.2 TFLOPs per forward pass, whereas CogAgent's cross-module requires only 12.56 TFLOPs (over 11×11\times fewer FLOPs). Using the cross-module at resolution 756 achieves superior performance to the standard architecture at resolution 490 (DocVQA 62.3 vs. 57.6, STVQA 73.6 vs. 68.1) while consuming approximately one-third of the forward FLOPs (10.08 vs. 29.14 TFLOPs) and roughly half the training step latency (3.57s vs. 6.43s).

  10. Knowl 10 — Ablation of Pre-Training Data Components on GUI and VQA Performance

    data/table

    An ablation study evaluated the progressive contribution of different pre-training data mixtures on STVQA, OCRVQA, DocVQA, and Mind2Web:

    Pre-Train Data Base Res Cross Res STVQA OCRVQA DocVQA Mind2Web
    Caption 490 — 68.1 74.5 57.6 38.6
    Caption + OCR 490 — 72.5 75.0 59.8 40.7
    Caption + OCR 224 1120 78.2 75.9 74.1 41.4
    All (+ Grounding + GUI Web) 224 1120 79.4 75.6 76.4 54.2

    Adding OCR data to base captioning improved DocVQA from 57.6 to 59.8 and STVQA from 68.1 to 72.5 at resolution 490. Upgrading to high resolution (1120 cross resolution) with Caption+OCR provided substantial gains on DocVQA (+14.3 to 74.1) and STVQA (+5.7 to 78.2). Incorporating domain-specific GUI data (CCS400K web screenshots with REC/REG pairs) and visual grounding data yielded a 12.8% absolute improvement on the Mind2Web GUI agent task (increasing step success rate from 41.4% to 54.2%).

  11. Knowl 11 — Operational and Architectural Limitations of CogAgent

    limitation

    CogAgent exhibits four primary failure modes and technical limitations during GUI automation:

    1. Single-Image Processing Constraint: The architecture processes individual screenshots and cannot natively intake multi-image histories or temporal screen streams simultaneously, relying instead on text-formatted action histories.
    2. Imprecise Coordinate Prediction: When targeting small, densely packed interactive elements, predicted bounding boxes and tap coordinates can drift from the interactive region boundaries.
    3. Incorrect GUI Observation: The visual encoder can fail to resolve subtle UI status indicators (such as distinguishing disabled buttons or slight visual highlight states).
    4. Action Prediction and Hallucination Errors: The model occasionally generates invalid action types or hallucinates interface elements that are not present on the current screen.

Coverage note — Omitted specific prompt text templates for GPT-4 trajectory conversion, detailed descriptions of individual benchmark splits, and full qualitative visualization figure dumps from the appendix, as the operational mechanisms, quantitative evaluations, and mathematical models are completely captured in the knowls.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015. 3, 6, 11
  2. 2.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023. 3, 7
  3. 3.Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sagnak Tas¸ırlar. ˘ Introducing our multimodal models, 2023. 7
  4. 4.Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marc¸al Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4291–4301, 2019. 3, 6, 12
  5. 5.Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. Nougat: Neural optical understanding for academic documents. arXiv preprint arXiv:2308.13418, 2023. 5
  6. 6.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https : / / github . com / kakaobrain/coyo-dataset, 2022. 5
  7. 7.Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao. Fireact: Toward language agent fine-tuning. arXiv preprint arXiv:2310.05915, 2023. 1
  8. 8.Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, et al. Pali-x: On scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565, 2023. 3, 7
  9. 9.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. 7
  10. 10.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. arXiv preprint arXiv:2306.06070, 2023. 3, 5, 6, 7, 12
  11. 11.Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, et al. Dreamllm: Synergistic multimodal comprehension and creation. arXiv preprint arXiv:2309.11499, 2023. 7
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3
  13. 13.Yuning Du, Chenxia Li, Ruoyu Guo, Xiaoting Yin, Weiwei Liu, Jun Zhou, Yifan Bai, Zilin Yu, Yehua Yang, Qingqing Dang, et al. Pp-ocr: A practical ultra lightweight ocr system. arXiv preprint arXiv:2009.09941, 2020. 5
  14. 14.Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023. 7
  15. 15.Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. In European Conference on Computer Vision, pages 498–517. Springer, 2022. 5
  16. 16.Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. Pix2struct: Screenshot parsing as pretraining for visual language understanding. In International Conference on Machine Learning, pages 18893–18912. PMLR, 2023. 4, 5, 7
  17. 17.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Mimic-it: Multi-modal in-context instruction tuning. arXiv preprint arXiv:2306.05425, 2023. 7
  18. 18.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 5, 7
  19. 19.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 3, 6, 11
  20. 20.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744, 2023. 7
  21. 21.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 3, 7
  22. 22.Tengchao Lv, Yupan Huang, Jingye Chen, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, et al. Kosmos-2.5: A multimodal literate model. arXiv preprint arXiv:2309.11419, 2023. 3
  23. 23.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195–3204, 2019. 3, 6, 11
  24. 24.Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244, 2022. 3, 6, 12
  25. 25.Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 3, 6, 12
  26. 26.Minesh Mathew, Viraj Bagal, Ruben Tito, Dimosthenis ` Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022. 3, 6, 12
  27. 27.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In 2019 international conference on document analysis and recognition (ICDAR), pages 947–952. IEEE, 2019. 6, 11
  28. 28.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021. 1
  29. 29.OpenAI. Introducing chatgpt. 2022. 1, 7
  30. 30.OpenAI. Gpt-4 technical report, 2023. 7
  31. 31.Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. Android in the wild: A large-scale dataset for android device control. arXiv preprint arXiv:2307.10088, 2023. 1, 3, 5, 6, 13
  32. 32.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022. 3, 5
  33. 33.Significant-Gravitas. Autogpt. https://github.com/Significant-Gravitas/AutoGPT, 2023. 1
  34. 34.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 3, 6, 11
  35. 35.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 3, 4
  36. 36.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multimodality. arXiv preprint arXiv:2307.05222, 2023. 7
  37. 37.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 7
  38. 38.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023. 1, 3, 5, 7
  39. 39.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022. 1
  40. 40.Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, et al. Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model. arXiv preprint arXiv:2310.05126, 2023. 7
  41. 41.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023. 3, 6, 11
  42. 42.Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. Agenttuning: Enabling generalized agent abilities for llms. abs/2310.12823, 2023. 1, 6, 12
  43. 43.Zhuosheng Zhan and Aston Zhang. You only look at screens: Multimodal chain-of-action agents. abs/2309.11436, 2023. 7, 13
  44. 44.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 7

Citation

MLA
Hong, W., et al. “CogAgent: A Visual Language Model for GUI Agents”. arXiv, 2023, http://arxiv.org/abs/2312.08914v3.
APA
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Zhang, Y., Li, J., Xu, B., Dong, Y., Ding, M., & Tang, J. (2023). CogAgent: A Visual Language Model for GUI Agents. arXiv. http://arxiv.org/abs/2312.08914v3
Chicago
Hong, W., W. Wang, Q. Lv, et al. 2023. “CogAgent: A Visual Language Model for GUI Agents”. arXiv. http://arxiv.org/abs/2312.08914v3.
Harvard
Hong, W. et al. (2023) “CogAgent: A Visual Language Model for GUI Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.08914v3.
Vancouver
1. Hong W, Wang W, Lv Q, et al (2023) CogAgent: A Visual Language Model for GUI Agents. arXiv

BibTeX

@article{hong2023cogagent,
  title = {CogAgent: A Visual Language Model for GUI Agents},
  author = {Hong, Wenyi and Wang, Weihan and Lv, Qingsong and Xu, Jiazheng and Yu, Wenmeng and Ji, Junhui and Wang, Yan and Wang, Zihan and Zhang, Yuxuan and Li, Juanzi and Xu, Bin and Dong, Yuxiao and Ding, Ming and Tang, Jie},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.08914v3},
  eprint = {2312.08914}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE