ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots
Yu-Chung HsiaoFedir ZubachGilles BaechlerSrinivas SunkaraVictor CarbuneJason LinMaria WangYun ZhuJindong Chen
Presents ScreenQA, a large-scale benchmark of 86,000 question-answer pairs with bounding-box groundings over mobile screenshots to evaluate multimodal models on visual screen reading comprehension and interface-based automation.
Mobile devices serve as primary channels for information delivery, driving demand for intelligent virtual assistants, automated testing, and autonomous user interface agents. However, existing automated systems often struggle to accurately read and interpret visual screen content directly from raw pixels. Prior research typically relied on underlying software view hierarchies—which frequently miss elements rendered in web views or custom graphics—or focused narrowly on either low-level icon detection or complex multi-step navigation. To establish reliable visual automation, machine learning models need a dedicated standard to measure their reading comprehension directly on complete mobile screens.
The article introduces and evaluates ScreenQA, a comprehensive benchmark designed to establish baseline capabilities for visual question answering across mobile application interfaces. Its main objective is to measure how accurately visual language models extract information, generate natural-language answers, and ground target elements directly from raw screenshot images.
To develop this benchmark, the authors curated 85,984 question-answer pairs spanning 35,352 mobile screenshots across 27 application categories from the established RICO dataset. The annotations capture concise short answers, conversational full-sentence long answers, relevant interface text elements, and precise spatial bounding boxes. The authors also intentionally designated about 10% of the questions as unanswerable to evaluate whether models can appropriately abstain from answering. They established four benchmark tasks and assessed multiple proprietary and open-weight models under zero-shot, fine-tuned, and cross-domain transfer learning settings, comparing multimodal image inputs against text-only optical character recognition pipelines.
The investigation produced several key findings. First, specialized fine-tuning yields high comprehension: dedicated models such as ScreenAI (5B parameters) and Gemini 1.5 Flash achieved leading short-answer accuracy, scoring around 90.5% to 90.7% on exact matching and roughly 95% on F1 metrics. Second, multimodal image inputs substantially outperform text-only approaches; feeding raw visual images into Gemini 1.5 Flash improved exact-match accuracy from 64.4% to 80.6% in zero-shot evaluations compared to pure text optical character recognition. Third, smaller open-weight models demonstrate strong competitive viability when fine-tuned at higher resolutions, with PaliGemma 3B reaching an 89.4% exact match score. Finally, training on mobile screens provides positive transfer to broader digital interfaces, showing measurable performance gains when applied to web page question answering.
These findings demonstrate that direct visual processing is both superior to and more robust than relying on brittle code hierarchies or standalone text recognition. For product development, the success of compact three-billion-parameter models suggests that high-performing screen comprehension can be deployed locally on mobile devices. On-device execution minimizes cloud latency, reduces hosting costs, and mitigates enterprise data privacy risks because screen contents do not need to leave the user's hardware.
Organizations developing digital assistants, accessibility tools, or user interface automation should transition to direct multimodal visual processing rather than maintaining separate optical character recognition layers. Teams deploying mobile AI solutions should consider fine-tuning compact, open-weight vision-language models on domain-specific screen data as a resource-efficient alternative to large proprietary cloud services. Future engineering work should integrate ScreenQA-trained models into live action-taking workflows while testing for robust failure handling when screens lack the requested data.
These conclusions should be considered within specific scope boundaries. ScreenQA focuses entirely on static, English-language Android smartphone interfaces and explicitly excludes complex arithmetic reasoning, multi-screen user journeys, and dynamic gaming layouts. Confidence in the reported extraction and grounding performance is high for standard mobile app layouts, but stakeholders should exercise caution and conduct pilot testing before deploying these models in multilingual contexts or highly dynamic, non-standard user interfaces.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Spotlight introduces the foundational paradigm of purely vision-based mobile UI comprehension directly from screen pixels without relying on view hierarchies, directly preceding ScreenQA's benchmarking of pixel-only screen reading.
- Paper: CogAgent: A Visual Language Model for GUI Agents, Wenyi Hong et al. (2024). CogAgent provides key architectural foundations for high-resolution vision-language models specialized in reading fine screen text and grounding interface elements on mobile and web screenshots.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). TextVQA establishes the core methodology and evaluation standards for multimodal question answering that requires reading embedded visual text rather than solely recognizing objects.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). DocVQA formalizes open-ended question answering grounded in visually rendered text and layouts, serving as a primary precursor to screen-oriented visual question answering.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). PIX2ACT establishes how pretraining visual transformers to parse raw pixel screenshots enables UI understanding and instruction-following without access to underlying code structures.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This seminal work establishes the foundational visual question answering paradigm and evaluation frameworks adapted by subsequent domain-specific benchmarks like ScreenQA.
- Paper: UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning, Zhengxi Lu et al. (2026). UI-R1 extends static mobile screen comprehension and grounding into efficient, reinforcement-learning-driven action execution for autonomous GUI agents.
- Paper: FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation, Kaiyi Huang et al. (2026). This paper builds upon raw-pixel UI grounding by introducing AGUVIS, a framework that integrates structured reasoning and multi-step planning for autonomous GUI agents.
