Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus

Gang LiYang Li

article2023ICLR107 citations

Introduces Spotlight, a vision-language approach that achieves state-of-the-art mobile user interface understanding directly from raw screenshots and region-of-interest focus points, removing the need for incomplete or noisy view hierarchy metadata.

Listen

Mobile user interface (UI) understanding is essential for enabling automated interactions and enhancing digital accessibility, such as screen readers for vision-impaired users. Historically, systems relied on underlying structural data—known as view hierarchies—alongside raw screenshots. However, view hierarchies are frequently missing, incomplete, or corrupted with inaccurate metadata and misaligned bounding boxes, which introduces runtime overhead, creates brittle dependencies, and constrains model performance.

The article demonstrates the effectiveness of Spotlight, a vision-only framework that models mobile UIs exclusively from raw screen pixels and focus coordinates. The primary objective is to evaluate whether a scalable vision-language architecture can understand mobile screens without needing view hierarchy metadata.

The researchers developed Spotlight by pairing a standard Vision Transformer for visual encoding with a Transformer text decoder. They introduced a "Region Summarizer" attention mechanism that uses a target area's bounding box coordinates to dynamically query visual tokens, capturing both local UI elements and broader surrounding context. To teach the model UI concepts prior to task fine-tuning, the authors pretrained the architecture on an extensive dataset consisting of 80 million rendered web pages from the C4 corpus and 2.5 million mobile screenshots. Spotlight was subsequently evaluated across four representative benchmarks: widget captioning, screen summarization, command grounding, and tappability prediction.

The evaluation produced four central findings. First, Spotlight established new state-of-the-art performance across all four downstream tasks, outperforming prior systems that relied on combined visual and view hierarchy inputs. Second, in widget captioning and screen summarization, the model achieved massive gains, improving captioning CIDEr metrics from the previous best of 99.3 to 141.8 (an increase of over 40%) and summarization from 65.6 to 106.7 (an improvement of more than 60%). Third, a unified multi-task model performed on par with specialized single-task systems in widget captioning and tappability, while still exceeding prior benchmarks in command grounding and summarization. Finally, ablation studies showed that joint pretraining on web and mobile visual data is essential; models trained from scratch or with frozen vision encoders failed to achieve competitive results.

These findings indicate that relying on brittle runtime structural metadata is unnecessary for UI modeling. By transitioning to a pure vision-language approach, organizations can streamline system architectures, reduce engineering overhead tied to cleaning noisy UI hierarchies, and deploy single multi-task models that lower operational footprints across diverse digital platforms.

Teams building mobile automation, accessibility tools, or design validation systems should transition from metadata-reliant pipelines toward unified vision-language architectures. Organizations should consider adopting the Region Summarizer design to handle localized UI interactions without sacrificing full-screen visual context. Before deploying few-shot prompting in production, practitioners should conduct additional domain-specific pretraining, as zero- and few-shot capabilities remain limited for complex UI tasks.

While the model delivers strong results, the study's conclusions are constrained by the relatively modest size of the investigated models (up to 843 million parameters) and the limited transferability of few-shot prompting beyond captioning. Additionally, while the framework bypasses runtime hierarchy requirements, initial bounding box coordinates are still required to direct the model's focus. Confidence in the reported fine-tuning and multi-task improvements remains high given the comprehensive evaluations and consistent performance across diverse standard benchmarks.

arXiv: 2209.14927
Cover for Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus

Abstract

Mobile UI understanding is important for enabling various interaction tasks such as UI automation and accessibility. Previous mobile UI modeling often depends on the view hierarchy information of a screen, which directly provides the structural data of the UI, with the hope to bypass challenging tasks of visual modeling from screen pixels. However, view hierarchies are not always available, and are often corrupted with missing object descriptions or misaligned structure information. As a result, despite the use of view hierarchies could offer short-term gains, it may ultimately hinder the applicability and performance of the model. In this paper, we propose Spotlight, a vision-only approach for mobile UI understanding. Specifically, we enhance a vision-language model that only takes the screenshot of the UI and a region of interest on the screen -- the focus -- as the input. This general architecture of Spotlight is easily scalable and capable of performing a range of UI modeling tasks. Our experiments show that our model establishes SoTA results on several representative UI tasks and outperforms previous methods that use both screenshots and view hierarchies as inputs. Furthermore, we explore multi-task learning and few-shot prompting capacities of the proposed models, demonstrating promising results in the multi-task learning direction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Data
  • 3.1 Pretraining Datasets
  • 3.2 UI Task Datasets
  • 4 Models
  • 5 Experiments
  • 5.1 Pretraining
  • 5.2 Finetuning, Multi-task Learning and Few-shot Prompting
  • 5.2.1 Finetuning
  • 5.2.2 Multi-task Learning
  • 5.2.3 Few-shot Learning
  • 6 Discussions
  • 7 Conclusions
  • References
  • A Hyperparameters & Model Sizes
  • B Ablation Study
  • C Pseudo Code
  • D Pretraining Datasets
  • E Downstream Tasks & Datasets

Knowls

  1. Knowl 1 — Spotlight Architecture for Focus-Aware Mobile UI Understanding

    model/method

    Spotlight is a vision-only encoder-decoder framework designed for mobile user interface (UI) modeling tasks that eliminates the dependency on structural view hierarchies. The architecture consists of three principal components:

    1. Vision Transformer (ViT) Image Encoder: Encodes raw UI screenshot images (resized with aspect ratio preserved and padded to 740×740740 \times 740 pixels) into visual feature representations H∈Rm×dH \in \mathbb{R}^{m \times d}, where mm is the sequence length of visual tokens and dd is the latent feature dimension.

    2. Region Summarizer (Focus Region Extractor): Takes normalized 4-scalar bounding box coordinates B=[left,top,right,bottom]T∈R4×1B = [\text{left}, \text{top}, \text{right}, \text{bottom}]^T \in \mathbb{R}^{4 \times 1} representing a region of interest (focus), expands them into query descriptors, and performs cross-attention over the ViT image encodings HH across ll Transformer layers to produce focus region representations Ql∈R4n×dQ_l \in \mathbb{R}^{4n \times d}.

    3. Transformer Decoder: An auto-regressive text decoder (initialized from multilingual T5) that attends to the focus region representations QlQ_l via cross-attention to generate target text sequences.

    By operating directly on visual pixels and bounding box coordinates, Spotlight maintains full visual context and avoids failure modes associated with missing, noisy, or misaligned structural metadata in view hierarchies.

  2. Knowl 2 — Region Summarizer Formulation

    equation

    The Region Summarizer maps a bounding box B=[b1,b2,b3,b4]T∈R4×1B = [b_1, b_2, b_3, b_4]^T \in \mathbb{R}^{4 \times 1} (representing normalized coordinates [left,top,right,bottom][\text{left}, \text{top}, \text{right}, \text{bottom}]) into query representations that attend to ViT image encodings H∈Rm×dH \in \mathbb{R}^{m \times d} through cross-attention.

    First, bounding box coordinates are projected to dense descriptor queries:

    E=Reshape4×nde→4×n×de([GeLU(BWe)]Wx)E = \text{Reshape}_{4 \times n d_e \to 4 \times n \times d_e}\left(\left[\text{GeLU}(B W_e)\right] W_x\right)

    X=Reshape4×n×d→4n×d(E+C)X = \text{Reshape}_{4 \times n \times d \to 4n \times d}(E + C)

    where We∈R1×ndeW_e \in \mathbb{R}^{1 \times n d_e} maps each scalar coordinate to nn dense vectors of dimension ded_e, Wx∈Rde×dW_x \in \mathbb{R}^{d_e \times d} projects descriptors to the transformer hidden dimension dd, and C∈R4×1×dC \in \mathbb{R}^{4 \times 1 \times d} is a learnable coordinate type embedding distinguishing the four coordinate positions.

    The initial region queries Q0=X∈R4n×dQ_0 = X \in \mathbb{R}^{4n \times d} are iteratively refined across ll Transformer layers (0≤i<l0 \le i < l):

    Yi+1=Qi+CrossAttention(q=Qi,kv=H ∥ Qi)Y_{i+1} = Q_i + \text{CrossAttention}(q = Q_i, kv = H \, \| \, Q_i)

    Qi+1=Yi+1+Dense(Yi+1)Q_{i+1} = Y_{i+1} + \text{Dense}(Y_{i+1})

    where H ∥ Qi∈R(m+4n)×dH \, \| \, Q_i \in \mathbb{R}^{(m + 4n) \times d} denotes concatenation along the sequence dimension, qq is the query, and kvkv is the key-value memory. The final output Ql∈R4n×dQ_l \in \mathbb{R}^{4n \times d} represents the contextualized focus region.

  3. Knowl 3 — Pretraining on Heterogeneous Web and Mobile UI Corpora

    model/method

    Spotlight is pretrained with an autoregressive sequence decoding objective to predict text attributes associated with UI regions given screenshot images and object bounding boxes, drawing from two complementary datasets:

    1. C4 Web Screenshots Dataset: Contains 80 million web page screenshots. Candidate text targets are derived from HTML attributes (e.g., button text, alt-text, aria-label, title, and placeholder). Object text consists of sampled attributes, with the overall webpage title sampled at a weight of 0.01 for whole-screen summarization.

    2. Mobile App UI Dataset: Contains 2.69 million mobile app screenshots collected from an Android emulator. Targets are extracted from leaf node view hierarchy attributes (text, content_desc, resource_id) and OCR text detected with ≥80%\ge 80\% confidence.

    Pretraining filters out generic non-informative words, non-alphabetical strings, URLs, and rare tokens (<5<5 occurrences). High-frequency phrases are downsampled using word2vec subsampling with threshold t=10−5t = 10^{-5}. Training examples are formed by packing 1 to 5 screenshot-object-text tuples per sequence padded to 5 tuples, separated by special Beginning-of-Chunk (BOC) and End-of-Chunk (EOC) tokens, with batches sampled at a 9:1 ratio between C4 and mobile data.

  4. Knowl 4 — Unified Sequence Decoding for Diverse UI Modeling Tasks

    model/method

    Spotlight unifies heterogeneous mobile UI tasks into a standard sequence generation formulation given an input tuple consisting of an image screenshot, a bounding box focus region, and an optional prompt text:

    • Widget Captioning: Focus is set to a UI object bounding box with no prompt text. The model autoregressively decodes the natural language functional description of the object.

    • Screen Summarization: Focus is set to the full screen bounding box [0,0,1,1][0, 0, 1, 1]. The model autoregressively decodes a high-level summary of the entire screen.

    • Command Grounding: Given a natural language command, the model iterates over candidate UI objects on the screen. For each object focus, the command is provided as a prompt text, and the decoder evaluates the log-likelihood of emitting the token Yes versus No. The object with the highest probability for Yes is selected as the predicted UI target.

    • Tappability Prediction: Given an object focus and a prompt string Tappable |, the model generates either Yes or No to predict whether the UI component is perceived as clickable.

  5. Knowl 5 — Downstream Task Finetuning Performance

    data/table

    Task-specific finetuning results demonstrate that vision-only Spotlight models (using ViT B/16 or ViT L/16 with T5 base) outperform prior multimodal models that rely on ground-truth or parsed view hierarchies:

    Model Captioning (CIDEr) Summarization (CIDEr) Grounding (Acc %) Tappability (F1 %)
    Widget Caption baseline 97.0 - - -
    Screen2Words baseline - 61.3 - -
    VUT (multimodal) 99.3 65.6 82.1 -
    Taperception baseline - - - 85.5
    Swearngin Li (2019) - - - 87.9
    Spotlight (ViT-B/16) 136.6 103.5 95.7 86.9
    Spotlight (ViT-L/16) 141.8 106.7 95.8 88.4

    Spotlight achieves substantial performance gains across all benchmarks without utilizing view hierarchy inputs, establishing improvements of over 40 CIDEr points on widget captioning and screen summarization, and over 13% absolute accuracy on command grounding relative to multimodal baselines.

  6. Knowl 6 — Multi-Task Downstream Performance

    data/table

    A single Spotlight model trained jointly on all four downstream UI tasks (using batch sampling weights [3, 2, 15, 1] for Widget Captioning, Screen Summarization, Command Grounding, and Tappability Prediction) achieves performance competitive with single-task models and significantly outperforms prior multi-task architectures:

    Model Captioning (CIDEr) Summarization (CIDEr) Grounding (Acc %) Tappability (F1 %)
    VUT multi-task 99.3 65.1 80.8 -
    Spotlight (ViT-B/16) 140.0 102.7 90.8 89.4
    Spotlight (ViT-L/16) 141.3 99.2 94.2 89.5

    Multi-task Spotlight models maintain high performance across diverse objective types (free-form generation and binary scoring) within a unified model instance.

  7. Knowl 7 — Region Summarizer Algorithm

    algorithm

    The Region Summarizer extracts bounding-box-focused feature vectors from dense visual tokens using coordinate-type embeddings and stacked cross-attention layers.

    Input: Normalized bounding box coordinates bbox∈Rbatch×4×1bbox \in \mathbb{R}^{\text{batch} \times 4 \times 1}
    Input: Visual token features vit_outputs∈Rbatch×m×dvit\_outputs \in \mathbb{R}^{\text{batch} \times m \times d}
    Input: Hyperparameters num_layersnum\_layers, hidden dimension hidden_dim=dhidden\_dim = d, projection dimension output_dim=deoutput\_dim = d_e, queries per coordinate num_query=nnum\_query = n
    Output: Contextualized region features x∈Rbatch×4n×dx \in \mathbb{R}^{\text{batch} \times 4n \times d}
    bbox←GeLU(Dense(bbox,features=output_dim))bbox \leftarrow \text{GeLU}(\text{Dense}(bbox, \text{features}=output\_dim))
    bbox←Reshape(bbox,[batch,4,num_query,−1])bbox \leftarrow \text{Reshape}(bbox, [\text{batch}, 4, num\_query, -1])
    bbox←Dense(bbox,features=hidden_dim)bbox \leftarrow \text{Dense}(bbox, \text{features}=hidden\_dim)
    coord_types←[[1],[2],[3],[4]]coord\_types \leftarrow [[1], [2], [3], [4]]
    type_embedding←Embed(coord_types)type\_embedding \leftarrow \text{Embed}(coord\_types)
    type_embedding←Reshape(type_embedding,[1,4,1,hidden_dim])type\_embedding \leftarrow \text{Reshape}(type\_embedding, [1, 4, 1, hidden\_dim])
    bbox←bbox+type_embeddingbbox \leftarrow bbox + type\_embedding
    x←Reshape(bbox,[batch,4⋅num_query,hidden_dim])x \leftarrow \text{Reshape}(bbox, [\text{batch}, 4 \cdot num\_query, hidden\_dim])
    for i=0i = 0 to num_layers−1num\_layers - 1 do
        kv←Concat([vit_outputs,x],axis=1)kv \leftarrow \text{Concat}([vit\_outputs, x], \text{axis}=1)
        x←x+CrossAttentioni(q=x,kv=kv)x \leftarrow x + \text{CrossAttention}_i(q=x, kv=kv)
        x←x+Densei(x)x \leftarrow x + \text{Dense}_i(x)
    end for
    return xx
  8. Knowl 8 — Ablation Study of Architecture and Pretraining Configurations

    data/table

    Ablation experiments evaluated on Spotlight variants (using ViT-B/16 pretrained for 100K steps) demonstrate the relative contributions of dataset mixture and Region Summarizer design choices across the four downstream tasks:

    Ablation Variant Widget Caption (CIDEr) Screen Summary (CIDEr) Grounding (Acc %) Tappability (F1 %)
    Full model 125.1 95.7 95.6 86.9
    Freeze ViT 37.5 19.6 72.6 72.3
    From scratch (no pretrain) 18.5 23.1 36.7 79.5
    C4 dataset only 120.3 94.2 96.3 87.8
    Mobile dataset only 105.7 90.3 89.8 87.3
    Static bbox queries 118.0 96.2 95.1 87.4
    No bbox in KV memory 114.7 93.7 95.9 87.8
    Joint bbox coordinate embedding 119.0 94.0 96.7 87.8
    ROI Align features as query 124.4 94.9 89.4 87.4
    Direct ROI Align (no Summarizer) 121.8 87.5 89.0 86.8

    Key takeaways from the ablation data:

    • Pretraining ViT end-to-end on both web and mobile data is critical; training from scratch or freezing ViT leads to catastrophic drops across all tasks.
    • Combining web (C4) and mobile corpora outperforms training on either alone.
    • Region Summarizer's individual coordinate projection and iterative memory updates provide necessary flexibility, outperforming standard ROI Align (especially on Screen Summarization, where ROI Align drops by 8.2 CIDEr points due to pooling loss).
  9. Knowl 9 — Few-Shot In-Context Prompting on Widget Captioning

    data/table

    Evaluating frozen pretrained Spotlight models without parameter finetuning across varying prompt shot counts {0,4,8,16,32}\{0, 4, 8, 16, 32\} reveals capacity-dependent in-context learning behavior on the Widget Captioning benchmark:

    ViT Backbone 0-shot 4-shot 8-shot 16-shot 32-shot
    ViT-B/16 57.1 56.7 55.6 55.5 54.9
    ViT-L/16 61.6 61.9 62.0 61.9 62.1

    While ViT-B/16 exhibits slight performance degradation as shots increase, the larger ViT-L/16 shows marginal improvements with additional context shots. Few-shot prompting did not yield meaningful performance on grounding, summarization, or tappability, indicating that smaller pretraining scale relative to general-domain foundation models limits few-shot generalization for tasks further from the pretraining text-decoding distribution.

  10. Knowl 10 — Contextual Attention Behavior in Region Summarizer

    empirical result

    Cross-attention weight analysis in the final layer of the Region Summarizer indicates that the learned bounding box queries perform soft, context-aware visual feature aggregation rather than strict hard cropping:

    1. Local Object Tasks (Widget Captioning): When presented with an isolated bounding box (e.g., a blank checkbox icon), the attention weights attend to the local target region while simultaneously attending to spatially distant context required for caption generation (e.g., corresponding text labels situated on the opposite side of the screen).

    2. Global Tasks (Screen Summarization): When given the bounding box for the entire viewport ([0,0,1,1][0, 0, 1, 1]), the cross-attention queries selectively focus on functionally dominant UI elements (e.g., headers, major cards, and action buttons) across the screen layout rather than distributing attention uniformly.

Coverage note — Omitted low-level hyperparameter listings from Appendix A (e.g., TPU hardware core hours, warmup schedules, and Adam optimizer details) and the exhaustive list of stop words from Appendix D, as these constitute standard configuration details rather than primary conceptual contributions.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning, 2022. URL https://arxiv.org/abs/2204.14198.
  2. 2.Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Agüera y Arcas. Uibert: Learning generic multimodal representations for UI understand- ing. In Zhi-Hua Zhou (ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, pp. 1705–1712. ijcai.org, 2021. doi: 10.24963/ijcai.2021/235. URL https://doi.org/10.24963/ijcai.2021/235.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  4. 4.Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. A dataset for interactive vision-language navigation with unknown command feasibility, 2022. URL https://arxiv.org/abs/2202.02312.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers, 2020. URL https://arxiv.org/abs/2005.12872.
  6. 6.Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. Pali: A jointly-scaled multilingual language-image model, 2022. URL https://arxiv.org/abs/2209.06794.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways, 2022. URL https://arxiv.org/abs/2204.02311.
  8. 8.Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. Rico: A mobile app dataset for building data-driven design applications. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology, UIST ’17, pp. 845–854, New York, NY, USA, 2017. Association for Computing Machinery. ISBN 9781450349819. doi: 10.1145/3126594.3126651. URL https://doi.org/10.1145/3126594.3126651.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://www.aclweb.org/anthology/N19-1423.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  11. 11.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick. Mask R-CNN. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pp. 2980–2988. IEEE Computer Society, 2017. doi: 10.1109/ICCV.2017.322. URL https://doi.org/10.1109/ICCV.2017.322.
  12. 12.Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby B. Lee, and Jindong Chen. Actionbert: Leveraging user actions for semantic understanding of user interfaces. CoRR, abs/2012.12350, 2020. URL https://arxiv.org/abs/2012.12350.
  13. 13.Gang Li, Gilles Baechler, Manuel Tragut, and Yang Li. Learning to denoise raw mobile ui layouts for improving datasets at scale, 2022. URL https://arxiv.org/abs/2201.04100.
  14. 14.Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. Mapping natural language instructions to mobile UI action sequences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8198–8210, Online, July 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.729. URL https://www.aclweb.org/anthology/2020.acl-main.729.
  15. 15.Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan. Widget captioning: Generating natural language description for mobile user interface elements, 2020b.
  16. 16.Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani, and Alexey A. Gritsenko. VUT: versatile UI transformer for multi-modal multi-task user interface modeling. CoRR, abs/2112.05692, 2021. URL https://arxiv.org/abs/2112.05692.
  17. 17.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Distributed representations of words and phrases and their compositionality, 2013. URL https://arxiv.org/abs/1310.4546.
  18. 18.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. CoRR, abs/1910.10683, 2019. URL http://arxiv.org/abs/1910.10683.
  20. 20.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022. URL https://arxiv.org/abs/2204.06125.
  21. 21.Anne Spencer Ross, Xiaoyi Zhang, James Fogarty, and Jacob O. Wobbrock. Examining image-based button labeling for accessibility in android apps through large-scale analysis. In Proceedings of the 20th International ACM SIGACCESS Conference on Computers and Accessibility, ASSETS ’18, pp. 119–130, New York, NY, USA, 2018. ACM. ISBN 978-1-4503-5650-3. doi: 10.1145/3234695.3236364. URL http://doi.acm.org/10.1145/3234695.3236364.
  22. 22.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding, 2022. URL https://arxiv.org/abs/2205.11487.
  23. 23.Eldon Schoop, Xin Zhou, Gang Li, Zhourong Chen, Björn Hartmann, and Yang Li. Predicting and explaining mobile ui tappability with vision modeling and saliency analysis, 2022. URL https://arxiv.org/abs/2204.02448.
  24. 24.Amanda Swearngin and Yang Li. Modeling Mobile Interface Tappability Using Crowdsourcing and Deep Learning, pp. 1–11. Association for Computing Machinery, New York, NY, USA, 2019. ISBN 9781450359702. URL https://doi.org/10.1145/3290605.3300305.
  25. 25.Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li. Screen2words: Automatic mobile UI summarization with multimodal learning. UIST’21, 2021. URL https://arxiv.org/abs/2108.03353.
  26. 26.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. Image as a foreign language: Beit pretraining for all vision and vision-language tasks, 2022. URL https://arxiv.org/abs/2208.10442.
  27. 27.Jason Wu, Xiaoyi Zhang, Jeff Nichols, and Jeffrey P. Bigham. Screen parsing: Towards reverse engineering of ui models from screenshots. 2021. URL https://arxiv.org/pdf/2109.08763.pdf.
  28. 28.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. CoRR, abs/2010.11934, 2020. URL https://arxiv.org/abs/2010.11934.
  29. 29.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models, 2022. URL https://arxiv.org/abs/2205.01917.
  30. 30.Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, Aaron Everitt, and Jeffrey P Bigham. Screen recognition: Creating accessibility metadata for mobile applications from pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450380966. doi: 10.1145/3411764.3445186. URL https://doi.org/10.1145/3411764.3445186.

Citation

MLA
Li, G., and Y. Li. “Spotlight: Mobile UI Understanding Using Vision-Language Models with a Focus”. arXiv, 2022, http://arxiv.org/abs/2209.14927v4.
APA
Li, G., & Li, Y. (2022). Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus. arXiv. http://arxiv.org/abs/2209.14927v4
Chicago
Li, G., and Y. Li. 2022. “Spotlight: Mobile UI Understanding Using Vision-Language Models with a Focus”. arXiv. http://arxiv.org/abs/2209.14927v4.
Harvard
Li, G. and Li, Y. (2022) “Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.14927v4.
Vancouver
1. Li G, Li Y (2022) Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus. arXiv

BibTeX

@article{li2022spotlight,
  title = {Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus},
  author = {Li, Gang and Li, Yang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.14927v4},
  eprint = {2209.14927}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission