ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots

Yu-Chung HsiaoFedir ZubachGilles BaechlerSrinivas SunkaraVictor CarbuneJason LinMaria WangYun ZhuJindong Chen

article2025NAACL55 citations

Presents ScreenQA, a large-scale benchmark of 86,000 question-answer pairs with bounding-box groundings over mobile screenshots to evaluate multimodal models on visual screen reading comprehension and interface-based automation.

Listen

Mobile devices serve as primary channels for information delivery, driving demand for intelligent virtual assistants, automated testing, and autonomous user interface agents. However, existing automated systems often struggle to accurately read and interpret visual screen content directly from raw pixels. Prior research typically relied on underlying software view hierarchies—which frequently miss elements rendered in web views or custom graphics—or focused narrowly on either low-level icon detection or complex multi-step navigation. To establish reliable visual automation, machine learning models need a dedicated standard to measure their reading comprehension directly on complete mobile screens.

The article introduces and evaluates ScreenQA, a comprehensive benchmark designed to establish baseline capabilities for visual question answering across mobile application interfaces. Its main objective is to measure how accurately visual language models extract information, generate natural-language answers, and ground target elements directly from raw screenshot images.

To develop this benchmark, the authors curated 85,984 question-answer pairs spanning 35,352 mobile screenshots across 27 application categories from the established RICO dataset. The annotations capture concise short answers, conversational full-sentence long answers, relevant interface text elements, and precise spatial bounding boxes. The authors also intentionally designated about 10% of the questions as unanswerable to evaluate whether models can appropriately abstain from answering. They established four benchmark tasks and assessed multiple proprietary and open-weight models under zero-shot, fine-tuned, and cross-domain transfer learning settings, comparing multimodal image inputs against text-only optical character recognition pipelines.

The investigation produced several key findings. First, specialized fine-tuning yields high comprehension: dedicated models such as ScreenAI (5B parameters) and Gemini 1.5 Flash achieved leading short-answer accuracy, scoring around 90.5% to 90.7% on exact matching and roughly 95% on F1 metrics. Second, multimodal image inputs substantially outperform text-only approaches; feeding raw visual images into Gemini 1.5 Flash improved exact-match accuracy from 64.4% to 80.6% in zero-shot evaluations compared to pure text optical character recognition. Third, smaller open-weight models demonstrate strong competitive viability when fine-tuned at higher resolutions, with PaliGemma 3B reaching an 89.4% exact match score. Finally, training on mobile screens provides positive transfer to broader digital interfaces, showing measurable performance gains when applied to web page question answering.

These findings demonstrate that direct visual processing is both superior to and more robust than relying on brittle code hierarchies or standalone text recognition. For product development, the success of compact three-billion-parameter models suggests that high-performing screen comprehension can be deployed locally on mobile devices. On-device execution minimizes cloud latency, reduces hosting costs, and mitigates enterprise data privacy risks because screen contents do not need to leave the user's hardware.

Organizations developing digital assistants, accessibility tools, or user interface automation should transition to direct multimodal visual processing rather than maintaining separate optical character recognition layers. Teams deploying mobile AI solutions should consider fine-tuning compact, open-weight vision-language models on domain-specific screen data as a resource-efficient alternative to large proprietary cloud services. Future engineering work should integrate ScreenQA-trained models into live action-taking workflows while testing for robust failure handling when screens lack the requested data.

These conclusions should be considered within specific scope boundaries. ScreenQA focuses entirely on static, English-language Android smartphone interfaces and explicitly excludes complex arithmetic reasoning, multi-screen user journeys, and dynamic gaming layouts. Confidence in the reported extraction and grounding performance is high for standard mobile app layouts, but stakeholders should exercise caution and conduct pilot testing before deploying these models in multilingual contexts or highly dynamic, non-standard user interfaces.

  • Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Spotlight introduces the foundational paradigm of purely vision-based mobile UI comprehension directly from screen pixels without relying on view hierarchies, directly preceding ScreenQA's benchmarking of pixel-only screen reading.
  • Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent provides key architectural foundations for high-resolution vision-language models specialized in reading fine screen text and grounding interface elements on mobile and web screenshots.
  • Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). TextVQA establishes the core methodology and evaluation standards for multimodal question answering that requires reading embedded visual text rather than solely recognizing objects.
  • Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). DocVQA formalizes open-ended question answering grounded in visually rendered text and layouts, serving as a primary precursor to screen-oriented visual question answering.
  • Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). PIX2ACT establishes how pretraining visual transformers to parse raw pixel screenshots enables UI understanding and instruction-following without access to underlying code structures.
  • Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This seminal work establishes the foundational visual question answering paradigm and evaluation frameworks adapted by subsequent domain-specific benchmarks like ScreenQA.
Cover for ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots

Abstract

We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a much higher-level composite task such as navigation and task completion for autonomous agents. ScreenQA attempts to bridge this gap. By annotating 86k question-answer pairs over the RICO dataset, we aim to benchmark the screen reading comprehension capacity, thereby laying the foundation for vision-based automation over screenshots. Our annotations encompass full answers, short answer phrases, and corresponding UI contents with bounding boxes, enabling four subtasks to address various application scenarios. We evaluate the dataset's efficacy using both open-weight and proprietary models in zero-shot, fine-tuned, and transfer learning settings. We further demonstrate positive transfer to web applications, highlighting its potential beyond mobile applications.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multimodality
  • 2.2 Question Answering
  • 3 Problem Setting: Tasks and Metrics
  • 4 Data Annotation
  • 4.1 Prefiltering
  • 4.2 Question Annotation
  • 4.3 Answer Annotation
  • 4.4 Not-Answerable Question Annotations
  • 4.5 Short Answer Generation
  • 5 Dataset Analysis
  • 5.1 Dataset Splitting
  • 5.2 Question Analysis
  • 5.3 Answer Analysis
  • 6 Experiments and Baselines
  • 6.1 Zero-Shot Experiments
  • 6.2 Fine-Tuning Experiments
  • 6.3 Cross-Domain and Transfer Learning
  • 6.4 Importance of Multimodality and Fine-Tuning
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethical Considerations
  • References
  • A Data Annotation Details
  • A.1 VH Out-of-Sync Rules
  • A.2 Question Annotation UI
  • A.3 Answer Annotation UI
  • A.4 Annotation quality control
  • A.5 Annotation post-processing
  • A.6 Short Answer Generation Prompts
  • B Data Examples
  • C Dataset Annotation Analysis
  • D Evaluation Configurations
  • D.1 Zero-shot
  • D.2 Fine-tuning
  • E Observed prediction errors on SQA-S task

Knowls

  1. Knowl 1 — ScreenQA Dataset Specification and Structure

    definition

    ScreenQA is a visual question answering benchmark built over mobile application screenshots from the RICO dataset, designed to evaluate reading comprehension and content understanding directly from screen pixels without relying on view hierarchy representations.

    The dataset comprises 85,984 question-answer pairs derived from 35,352 unique Android application screenshots spanning 27 app categories. The data is partitioned at the screenshot level into training, validation, and test splits with an approximate 80-10-10 ratio:

    • Train split: 28,378 screenshots, 68,951 questions (1 ground-truth annotation per question).
    • Validation split: 3,485 screenshots, 8,614 questions (3 ground-truth annotations per question).
    • Test split: 3,489 screenshots, 8,419 questions (3 ground-truth annotations per question).

    Approximately 10% of the questions in each split are deliberately unanswerable from the visual context to evaluate a model's ability to refrain from answering when information is absent.

  2. Knowl 2 — Multi-Annotator Evaluation Metric Formulation for ScreenQA

    equation

    For validation and test evaluations containing multiple annotator ground truths per example, the evaluation metric is computed by taking the maximum score across all ground-truth annotations for a given question and averaging these maximum values across the entire dataset:

    avg(metric)=1N∑i=1Nmax⁡j[metric(Ai,Ai,jg)]\text{avg}(\text{metric}) = \frac{1}{N} \sum_{i=1}^{N} \max_{j} \left[ \text{metric}\left(A_i, A_{i,j}^g\right) \right]

    Where:

    • N∈NN \in \mathbb{N} denotes the total number of evaluation questions in the dataset split.
    • AiA_i represents the model's predicted answer string, list, or bounding box set for the ii-th question.
    • Ai,jgA_{i,j}^g denotes the jj-th ground-truth reference provided by annotators for the ii-th question.
    • metric(⋅,⋅)\text{metric}(\cdot, \cdot) is a task-specific scoring function (such as Exact Match, token-level F1, ROUGE, or bounding box IoU-based F1).
  3. Knowl 3 — ScreenQA Benchmark Subtasks and Evaluation Metrics

    definition

    ScreenQA defines four distinct subtasks to benchmark different aspects of screen content extraction, language generation, and visual grounding:

    1. ScreenQA Short Answer (SQA-S): Given a screenshot and a question, generate a concise short answer using information on the screen. If the screen lacks the answer, output "<no answer>". Evaluated using SQuAD-style Exact Match (EM) and token-level F1-score after standard text normalization.
    2. ScreenQA Long Answer (SQA-L): Given a screenshot and a question, generate a full-sentence natural language response suitable for virtual assistant interactions. Evaluated using ROUGE-1 (R-1), ROUGE-2 (R-2), and ROUGE-L (R-L).
    3. ScreenQA UI Content (SQA-UIC): Given a screenshot and a question, output an ordered list containing the text representations of UI elements that answer the question (or an empty list if unanswerable). Evaluated via element-wise Exact Match and list-level F1-score.
    4. ScreenQA UI Content with Bounding Boxes (SQA-UIC-BB): Given a screenshot and a question, output a list of UI elements where each item contains both its bounding box coordinates and textual content. Bounding box detection quality is evaluated via BBOX-F1, where a predicted box matches a ground truth box if their Intersection over Union (IoU) exceeds 0.1 (accounting for variance between tight manual box drawing and view hierarchy node boundaries). Text+BBox EM and F1 additionally require the text representation to match.
  4. Knowl 4 — Five-Stage Annotation Pipeline for ScreenQA

    model/method

    ScreenQA was constructed via a systematic five-stage annotation process:

    1. Prefiltering: Annotators screened raw RICO screenshots to eliminate non-English applications and cases where Android View Hierarchies (VHs) were out of sync with the visual rendering. Transition artifacts (occlusion and ghosting) were retained if main content elements remained accurate.
    2. Question Annotation: Two independent annotators generated questions based solely on directly observable screen content (excluding external reasoning, computation, or ads). Annotator 1 composed up to 5 questions; Annotator 2 reviewed them and added up to 3 additional non-redundant questions (maximum 8 per screen).
    3. Answer Annotation: Independent annotators (disjoint from the question author) corrected grammatical errors, selected/drew relevant UI bounding boxes (using VH leaf nodes or manual drawing), ranked the boxes by reading/relevance order, and wrote a complete full-sentence answer.
    4. Unanswerable Question Annotation: Annotators authored unanswerable but screen-relevant questions on ~5,000 sampled screenshots to ensure ~10% of dataset questions evaluate refusal/unanswerability handling.
    5. Short Answer Generation: PaLM 2 was prompted few-shot with the question, UI element text list, and full-sentence answer to synthesize normalized, human-readable short answer strings accommodating valid formatting variants.
  5. Knowl 5 — Fine-Tuning Baselines Across ScreenQA Subtasks

    data/table

    Fine-tuning open-weight and proprietary vision-language models across the four ScreenQA tasks demonstrates that domain-specialized pretraining and higher input resolutions substantially improve performance on screen visual comprehension and bounding box localization.

    Model SQA-S SQA-L SQA-UIC SQA-UIC-BB
    EM F1 R-1 R-2 R-L EM F1 BBOX-F1 EM F1
    ScreenAI 670M 51.2 60.6 77.3 68.4 76.7 47.8 49.4 62.7 41.1 42.6
    PaliGemma 3B 224 77.5 83.9 88.2 81.5 87.4 74.8 76.7 84.9 67.5 69.6
    PaliGemma 3B 448 88.3 92.2 91.1 85.5 90.3 86.0 87.7 89.4 79.1 81.6
    PaliGemma 3B 896 89.4 93.2 90.9 85.3 90.1 86.1 87.8 88.8 78.8 81.2
    ScreenAI 5B 90.7 94.6 92.6 87.4 91.9 87.0 88.7 94.2 84.0 85.7
    Gemini 1.5 Flash 90.5 94.9 92.4 86.2 91.7 88.2 89.7 92.4 83.9 85.7

    ScreenAI 5B (which incorporates specialized UI and infographic pretraining) achieves the highest grounding localization performance (94.2 BBOX-F1 and 84.0 EM on SQA-UIC-BB). For PaliGemma 3B, increasing image resolution from 224×224224\times224 to 448×448448\times448 yields large gains (+10.8 SQA-S EM), with diminishing returns at 896×896896\times896 (+1.1 SQA-S EM).

  6. Knowl 6 — Zero-Shot Multimodal Model Evaluation on Short Answer ScreenQA

    data/table

    Zero-shot evaluation on the ScreenQA Short Answer task (SQA-S) assesses the out-of-the-box screen reading comprehension of public and proprietary visual LLMs using normalized prompting.

    Model SQuAD-EM SQuAD-F1
    Fuyu-8B 39.5 47.3
    Gemini 1.5 Flash 80.6 86.4
    GPT-4o 77.8 86.6
    Gemini 1.5 Pro 81.4 87.2

    Gemini 1.5 Pro achieves the highest zero-shot accuracy (81.4 EM, 87.2 F1). While strong zero-shot visual models reach ~87% F1, supervised fine-tuning pushes performance above 94% F1.

  7. Knowl 7 — Cross-Domain and Transfer Learning Between Mobile Screens and Document/Web UI

    data/table

    Cross-domain learning (CDL) between ScreenQA (SQA-S) and Document Visual Question Answering (DocVQA), alongside transfer learning (TL) to VisualWebBench-WebQA using fine-tuned PaliGemma 3B checkpoints, reveals domain gaps between mobile UI and scanned documents while demonstrating positive transfer from mobile screens to web interfaces.

    Experiment | Task CDL | SQA-S CDL | DocVQA TL | VisualWebBench-WebQA
    Fine-Tuned on SQA-S DocVQA SQA-S DocVQA DocVQA SQA-S + DocVQA
    Model Resolution EM F1 EM F1 ANLS ANLS F1 F1
    PaliGemma 3B 224 77.5 83.9 52.6 60.5 27.5 43.7 19.31 21.51
    PaliGemma 3B 448 88.3 92.2 66.2 72.9 55.1 78.0 48.33 49.11
    PaliGemma 3B 896 89.4 93.2 63.6 70.4 62.0 84.8 57.07 58.69

    Cross-domain evaluation shows a consistent 20--25 point degradation in metric scores when evaluating models trained on DocVQA on SQA-S or vice versa, demonstrating that mobile UI screens and scanned documents represent distinct visual-textual domains. Joint training on DocVQA and SQA-S consistently improves transfer performance on WebQA over training on DocVQA alone (+1.62 to +2.20 F1 points).

  8. Knowl 8 — Impact of Visual Multimodality Versus Text-Only OCR on Screen QA

    data/table

    Evaluating Gemini 1.5 Flash in text-only (OCR representation as input) versus multimodal (image + text) configurations confirms that raw visual pixels provide crucial spatial and stylistic context for mobile screen question answering.

    Evaluation Setup SQA-S SQA-L SQA-UIC SQA-UIC-BB
    EM F1 R-1 R-2 R-L EM F1 BBOX-F1 EM F1
    Zero-Shot, Text-Only 64.4 72.5 78.2 67.8 75.6 38.6 41.2 30.1 49.7 28.5
    Zero-Shot (Multimodal) 80.6 86.4 83.8 70.8 79.3 62.4 66.8 26.2 33.3 24.0
    Fine-Tuned (Multimodal) 90.5 94.9 92.4 86.2 91.7 88.2 89.7 85.7 92.4 83.9

    Text-only OCR input leads to a major degradation in zero-shot SQA-S performance compared to multimodal input (64.4 EM vs. 80.6 EM). On SQA-UIC-BB zero-shot, text-only shows higher text EM because bounding box coordinates were explicitly provided in the text prompt, reducing the task to coordinate selection rather than visual spatial regression.

  9. Knowl 9 — Error Taxonomy for Mobile Screen Visual Question Answering

    definition

    Failure modes of visual language models on the ScreenQA Short Answer task fall into five distinct error categories:

    1. Misinterpreted Question: The predicted answer correctly reflects some screen content but answers an entirely different question than what was asked (e.g., predicting a serial number when asked for a total score).
    2. Misinterpreted Screenshot Info: The model identifies a visual element related to the question topic but conflates distinct UI components or roles (e.g., predicting the current playback timestamp instead of total audio duration).
    3. Hallucination: The predicted answer introduces content or labels that do not exist on the screen image.
    4. Modified or Misread Content: The prediction originates from the correct on-screen UI element but contains typographical errors, character misrecognitions, or unintended word alterations.
    5. Lack of Understanding/Reasoning: The model misinterprets compact or composite UI layouts, such as date picker calendars or range notations.
  10. Knowl 10 — Scope Boundaries and Limitations of ScreenQA

    limitation

    ScreenQA operates under several intentional scope boundaries and structural limitations:

    • Monolingual Focus: All screenshot samples and annotated questions/answers are in English only; multilingual and multi-locale UI configurations are not covered.
    • Single-Screen Static Context: Every dataset example consists of a single static Android phone screenshot. The benchmark does not include multi-screen navigation traces, dynamic UI animations, video, or scrolling interactions.
    • Information Extraction Focus: Questions target directly observable UI content and layout composition lookup, intentionally excluding multi-step compositional arithmetic, external common-sense reasoning, and highly dynamic non-canonical layouts (e.g., mobile games).

Coverage note — None omitted. The complete core contribution—including dataset definition, annotation pipeline, mathematical evaluation formulation, all four subtasks, comprehensive fine-tuning baselines, zero-shot benchmarks, cross-domain/transfer learning experiments, multimodality ablations, error taxonomy, and stated limitations—is fully represented across the knowls.

References

  1. 1.Abdelrahman Abdallah, Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Mohamed Mahmoud, Ibrahim Abdelhalim, Mohamed Elkasaby, Yasser El-Bendary, and Adam Jatowt. 2024. CORU: Comprehensive Post-OCR Parsing and Receipt Understanding Dataset. Preprint, arXiv:2406.04493.
  2. 2.Ali Ahmed, Alaa Zaki, Enas Elgeldawi, Mohamed Abdallah, and Moheb Girgis. 2023. MASC: A Dataset for the Development and Classification of Mobile Applications Screens.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. PaLM 2 technical report. arXiv preprint arXiv:2305.10403.
  4. 4.Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Carbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024. ScreenAI: A vision-language model for UI and infographics understanding. Preprint, arXiv:2402.04615.
  5. 5.Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Aguera y Arcas. 2021. UIBert: Learning generic multimodal representations for UI understanding. Preprint, arXiv:2107.13731.
  6. 6.Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sagnak Taşırlar. 2023. Introducing our multimodal models.
  7. 7.Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai. 2024. PaliGemma: A versatile 3B VLM for transfer. arXiv preprint arXiv:2407.07726.
  8. 8.Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. 2024. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. Preprint, arXiv:2407.21787.
  9. 9.Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. 2022. A dataset for interactive vision language navigation with unknown command feasibility. In European Conference on Computer Vision (ECCV).
  10. 10.Ritwick Chaudhry, Sumit Shekhar, Utkarsh Gupta, Pranav Maneriker, Prann Bansal, and Ajay Joshi. 2020. LEAF-QA: Locate, Encode & Attend for Figure Question Answering. In 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 3501–3510.
  11. 11.Jieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu, Liming Zhut, Guoqiang Li, and Jinshui Wang. 2020. Unblind your apps: Predicting natural-language labels for mobile GUI components by deep learning. In 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE), pages 322–334.
  12. 12.Xingyu Chen, Zihan Zhao, Lu Chen, Danyang Zhang, Jiabao Ji, Ao Luo, Yuxuan Xiong, and Kai Yu. 2021a. WebSRC: A Dataset for Web-Based Structural Reading Comprehension. Preprint, arXiv:2101.09465.
  13. 13.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021b. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3697–3711, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  14. 14.Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu. 2024. SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9313–9332, Bangkok, Thailand. Association for Computational Linguistics.
  15. 15.Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. 2017. Rico: A Mobile App Dataset for Building Data-Driven Design Applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology, UIST ’17, pages 845–854, New York, NY, USA. Association for Computing Machinery.
  16. 16.Gemini Team Google. 2023. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  17. 17.Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. VizWiz Grand Challenge: Answering Visual Questions from Blind People. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3608–3617, Salt Lake City, UT, USA. IEEE.
  18. 18.Xiaotian Han, Quanzeng You, Yongfei Liu, Wentao Chen, Huangjie Zheng, Khalil Mrini, Xudong Lin, Yiqi Wang, Bohan Zhai, Jianbo Yuan, Heng Wang, and Hongxia Yang. 2023. InfiMM-Eval: Complex Open-Ended Reasoning Evaluation For Multi-Modal Large Language Models. Preprint, arXiv:2311.11567.
  19. 19.Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis. 2015. Evaluation of deep convolutional nets for document image classification and retrieval. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR), pages 991–995.
  20. 20.Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al. 2023. CogAgent: A visual language model for GUI agents. arXiv preprint arXiv:2312.08914.
  21. 21.Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shjian Lu, and C. V. Jawahar. 2019. ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1516–1520.
  22. 22.Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019. FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents. Preprint, arXiv:1905.13538.
  23. 23.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. 2018. DVQA: Understanding Data Visualizations via Question Answering. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5648–5656, Salt Lake City, UT. IEEE.
  24. 24.Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio. 2017. FigureQA: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300.
  25. 25.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. 2024. VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649.
  26. 26.Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, and Aleksandra Faust. 2024. Training Language Models to Self-Correct via Reinforcement Learning. Preprint, arXiv:2409.12917.
  27. 27.D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard. 2006. Building a test collection for complex document information processing. In Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’06, pages 665–666, New York, NY, USA. Association for Computing Machinery.
  28. 28.Gang Li, Gilles Baechler, Manuel Tragut, and Yang Li. 2022. Learning to Denoise Raw Mobile UI Layouts for Improving Datasets at Scale. In CHI Conference on Human Factors in Computing Systems, pages 1–13, New Orleans LA USA. ACM.
  29. 29.Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020a. Mapping Natural Language Instructions to Mobile UI Action Sequences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8198–8210, Online. Association for Computational Linguistics.
  30. 30.Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. 2020b. Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5495–5510, Online. Association for Computational Linguistics.
  31. 31.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  32. 32.Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue. 2024a. VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? Preprint, arXiv:2404.05955.
  33. 33.Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024b. TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document. Preprint, arXiv:2403.04473.
  34. 34.Yuwen Lu, Yuewen Yang, Qinyi Zhao, Chengzhi Zhang, and Toby Jia-Jun Li. 2024. AI Assistance for UX: A Literature Review Through Human-Centered AI. Preprint, arXiv:2402.06089.
  35. 35.Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244.
  36. 36.Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. 2022. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706.
  37. 37.Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021. DocVQA: A dataset for vqa on document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2200–2209.
  38. 38.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium. Association for Computational Linguistics.
  39. 39.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. 2019. OCR-VQA: Visual question answering by reading text in images. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 947–952. IEEE.
  40. 40.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. GPT-4 technical report. Preprint, arXiv:2303.08774.
  41. 41.David Owen. 2024. How predictable is language model benchmark performance? Preprint, arXiv:2401.04757.
  42. 42.Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii, and Ying Xiao. 2019. Towards unconstrained end-to-end text spotting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  43. 43.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  44. 44.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  45. 45.Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. 2024. Android in the wild: A large-scale dataset for android device control. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, pages 59708–59728, Red Hook, NY, USA. Curran Associates Inc.
  46. 46.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. 2019. Generalized intersection over union: A metric and a loss for bounding box regression. Preprint, arXiv:1902.09630.
  47. 47.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8317–8326.
  48. 48.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. Preprint, arXiv:2408.03314.
  49. 49.Srinivas K. Sunkara, Maria Wang, Lijuan Liu, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, and James W. Stout. 2022. Towards Better Semantic Understanding of Mobile Interfaces. In Proceedings of the 30th International Conference on Computational Linguistics, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  50. 50.Alon Talmor and Jonathan Berant. 2018. The Web as a Knowledge-Base for Answering Complex Questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 641–651, New Orleans, Louisiana. Association for Computational Linguistics.
  51. 51.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  52. 52.Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li. 2021. Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning. In The 34th Annual ACM Symposium on User Interface Software and Technology, UIST ’21, pages 498–510, New York, NY, USA. Association for Computing Machinery.
  53. 53.Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, and Hongxia Yang. 2024. Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning. Preprint, arXiv:2401.06805.
  54. 54.Jason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng, Jeffrey Nichols, and Jeffrey P Bigham. 2023. Webui: A dataset for enhancing visual ui understanding with web semantics. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–14.
  55. 55.Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang. 2024. An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models. Preprint, arXiv:2408.00724.
  56. 56.Jinheng Xie, Kai Ye, Yudong Li, Yuexiang Li, Kevin Qinghong Lin, Yefeng Zheng, Linlin Shen, and Mike Zheng Shou. 2023. Learning Visual Prior via Generative Pre-Training. Advances in Neural Information Processing Systems, 36:70562–70580.
  57. 57.Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni, Mo Yu, Shiyu Chang, Xiaoxiao Guo, and William Yang Wang. 2019. TWEETQA: A Social Media Focused Question Answering Dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5020–5031, Florence, Italy. Association for Computational Linguistics.
  58. 58.Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2013–2018, Lisbon, Portugal. Association for Computational Linguistics.
  59. 59.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  60. 60.Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020. ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. Preprint, arXiv:2002.04326.
  61. 61.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. Preprint, arXiv:2311.16502.
  62. 62.Xiaoxue Zang, Ying Xu, and Jindong Chen. 2021. Multimodal icon annotation for mobile applications. In Proceedings of the 23rd International Conference on Mobile Human-Computer Interaction, pages 1–11.
  63. 63.Biao Zhang, Zhongtao Liu, Colin Cherry, and Orhan Firat. 2024. When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method. Preprint, arXiv:2402.17193.

Citation

MLA
Hsiao, Y.-C., et al. “ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 9427–52, https://doi.org/10.18653/v1/2025.naacl-long.477.
APA
Hsiao, Y.-C., Zubach, F., Baechler, G., Sunkara, S., Cărbune, V., Lin, J., Wang, M., Zhu, Y., & Chen, J. (2025). ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 9427–9452. https://doi.org/10.18653/v1/2025.naacl-long.477
Chicago
Hsiao, Y.-C., F. Zubach, G. Baechler, et al. 2025. “ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 9427–52. https://doi.org/10.18653/v1/2025.naacl-long.477.
Harvard
Hsiao, Y.-C. et al. (2025) “ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9427–9452. Available at: https://doi.org/10.18653/v1/2025.naacl-long.477.
Vancouver
1. Hsiao Y-C, Zubach F, Baechler G, Sunkara S, Cărbune V, Lin J, Wang M, Zhu Y, Chen J (2025) ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 9427–9452

BibTeX

@inproceedings{hsiao-etal-2025-screenqa,
    title = "{S}creen{QA}: Large-Scale Question-Answer Pairs Over Mobile App Screenshots",
    author = "Hsiao, Yu-Chung  and
      Zubach, Fedir  and
      Baechler, Gilles  and
      Sunkara, Srinivas  and
      Carbune, Victor  and
      Lin, Jason  and
      Wang, Maria  and
      Zhu, Yun  and
      Chen, Jindong",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.477/",
    doi = "10.18653/v1/2025.naacl-long.477",
    pages = "9427--9452",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/