Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Kenton LeeMandar JoshiIulia Raluca TurcHexiang HuFangyu LiuJulian Martin EisenschlosUrvashi KhandelwalPeter ShawMing-Wei ChangKristina Toutanova

article2023ICML521 citations

Presents Pix2Struct, an image-to-text model pretrained by parsing masked web screenshots into simplified HTML, enabling a single OCR-free architecture to handle diverse visual language tasks across documents, diagrams, and user interfaces.

Listen

Modern digital environments—such as documents, mobile applications, and web pages—present language and visual information holistically. Previous artificial intelligence methods for processing this visually-situated text have relied on fragmented, domain-specific pipelines that depend heavily on external tools, such as optical character recognition (OCR) or specialized application metadata. These multi-step pipelines increase engineering complexity, limit adaptability across different formats, and elevate computational costs.

The article demonstrates Pix2Struct, a unified, pretrained image-to-text model designed for purely visual language understanding without requiring external OCR tools or domain-specific intermediate steps. The goal was to establish a single general-purpose framework capable of handling varied visual language tasks by learning directly from raw pixel inputs.

The approach uses a vision transformer architecture trained on 80 million web page screenshots paired with simplified HTML source code. During this self-supervised pretraining, the model learns to reconstruct the underlying HTML structure from partially masked screenshots, effectively blending OCR, masked language modeling, and image captioning into one objective. To handle the varied aspect ratios of documents and user interfaces without distortion, the authors introduced a variable-resolution patching mechanism. For downstream tasks, text prompts (such as questions) are rendered directly onto the input image, processing all information through a single visual channel. Two model variants were evaluated across nine benchmarks spanning four domains: illustrations, user interfaces, natural images, and documents.

The evaluation yielded several key findings. Pix2Struct established a new state of the art in six of the nine benchmarks. In low-resource domains, it outperformed existing specialized systems, improving chart question-answering accuracy from 45.5 to 58.6 and mobile screen summarization score from 64.3 to 109.4. Across all tested benchmarks, it substantially outperformed Donut, the leading OCR-free baseline, by 9 to 53 points. While the model trailed top-performing OCR-reliant pipelines in text-dense document tasks and models trained on billions of caption pairs for natural images, it delivered competitive performance using substantially less domain-specific training.

These results indicate that end-to-end visual pretraining from web markup provides a viable, scalable alternative to specialized multi-stage pipelines. By eliminating reliance on external OCR tools, organizations can reduce architectural complexity, streamline operational maintenance, and deploy a single model across diverse visual and document formats.

Decision-makers and engineering teams should consider piloting unified, pixel-only models for user interface and illustration workflows where specialized tools are brittle or unavailable. For text-heavy document extraction, organizations should weigh the simplicity of an end-to-end model against the marginal accuracy advantage of established OCR pipelines. Future development should focus on scaling pretraining data and exploring long-range architectures to further close performance gaps in document-dense domains.

The primary limitation of the model is its high sensitivity to image resolution, as high-resolution inputs demand significant computational memory and sequence lengths during processing. Additionally, the pretraining data relies on web corpora that require careful curation to mitigate undesirable content risks. While confidence is high in the model's effectiveness across user interfaces and illustrative tasks, readers should exercise caution when deploying purely pixel-based models to high-density document tasks without task-specific validation.

Cover for Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Abstract

Visually-situated language is ubiquitous—sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Background
  • 2.2. Architecture
  • 2.3. Pretraining
  • 2.4. Warming up with a reading curriculum
  • 2.5. Finetuning
  • 3. Experimental Setup
  • 3.1. Benchmarks
  • 3.2. Implementation and Baselines
  • 4. Results
  • 4.1. Illustrations
  • 4.2. UIs
  • 4.3. Natural Images
  • 4.4. Documents
  • 5. Analysis
  • 6. Discussion
  • 7. Related Work
  • References
  • A. Resolution in visually-situated language understanding tasks
  • B. Full Results
  • C. Finetuning Dataset Details
  • D. Hyperparameters
  • E. Warmup Stage Data
  • F. Pretraining Data

Knowls

  1. Knowl 1 — Screenshot Parsing Pretraining Objective

    model/method

    Pix2Struct introduces a self-supervised screenshot parsing pretraining task where the model takes a masked screenshot of a web page as visual input and autoregressively generates a simplified HTML-based Document Object Model (DOM) parse.

    To construct training pairs from web pages rendered with a 1024×10241024 \times 1024 pixel viewport:

    1. DOM Condensation: The DOM tree is filtered to retain only nodes with visible elements or visible descendants. When a node lacks visible content and has only one child, it is collapsed into that child to remove redundant nesting.
    2. Content Extraction: For each retained node, only the text, image filenames, and image alternative text (alt-text) are preserved.
    3. Subtree Selection: The largest linearized DOM subtree that fits within the decoder sequence budget (128 tokens) is selected as the target. A visual bounding box corresponding to the region covered by this subtree is rendered directly onto the input screenshot.
    4. Masking: A denoising objective is applied by masking 50%50\% of the text spans in the selected subtree directly on the screenshot image with solid masks, requiring the decoder to recover the full subtree.

    This objective unifies three core capabilities in visually-situated language processing: recovering unmasked text acts as optical character recognition (OCR); predicting masked text acts as masked language modeling (MLM) guided by visual context; and predicting alt-text and filenames acts as context-aware image captioning.

  2. Knowl 2 — Variable-Resolution Input Representation for Vision Transformers

    model/method

    Standard Vision Transformers (ViT) rescale input images to a fixed square resolution (such as 224×224224 \times 224 or 512×512512 \times 512), which either distorts the natural aspect ratio of visually-situated language inputs (such as tall documents or wide web pages) or introduces excessive zero-padding that wastes sequence capacity.

    Pix2Struct replaces fixed-resolution preprocessing with a variable-resolution patching mechanism:

    1. Given an input image with original width WW and height HH, and a sequence length budget of LL patches of fixed size p×pp \times p (where p=16p = 16 pixels), the image is scaled by an isotropic factor ss such that the total number of extracted patches ⌈sWp⌉×⌈sHp⌉≤L\lceil \frac{sW}{p} \rceil \times \lceil \frac{sH}{p} \rceil \le L is maximized while preserving the aspect ratio W/HW / H.
    2. Patches are extracted from the scaled image without distortion or artificial padding.
    3. To handle variable patch grid dimensions unambiguously, each patch is assigned a 2-dimensional absolute positional embedding corresponding to its discrete (x,y)(x, y) coordinate in the patch grid.

    This representation allows on-the-fly adjustment of sequence lengths and resolutions between pretraining and finetuning without interpolation of positional embeddings.

  3. Knowl 3 — Visual Prompt Rendering for Multimodal Downstream Tasks

    model/method

    Pix2Struct processes all multimodal inputs through a purely visual channel by rendering text prompts and spatial annotations directly onto the input image before encoding:

    • Visual Question Answering (VQA): In tasks such as DocVQA, InfographicsVQA, ChartQA, and OCR-VQA, the question text is rendered as a visual header banner directly above the input image. For multiple-choice questions (e.g., AI2D), candidate choices are also rendered inside the header banner alongside the question.
    • Image and Widget Captioning: In standard image captioning (e.g., TextCaps, Screen2Words), the raw image is passed directly. For localized targets (e.g., Widget Captioning), a bounding box is drawn over the target UI component directly on the image.
    • Referring Expression Resolution: In UI referring expression tasks (e.g., RefExp), each candidate component is evaluated by generating an image containing the screenshot, the query expression rendered in the header, and a bounding box around the candidate component. The model decodes a binary token ("true" or "false"). Training uses five negative candidates per positive instance, and inference ranks candidates by the model's generation score for "true".
  4. Knowl 4 — Reading Curriculum Warmup Pretraining

    model/method

    Directly pretraining Pix2Struct from random initialization on screenshot parsing causes training instability and slow convergence. To address this, pretraining begins with a short synthetic reading curriculum warmup stage.

    Warmup data generation:

    • Text snippets up to 128 bytes long are sampled from the BooksCorpus dataset.
    • Snippets are rendered as unmasked text on a plain white background with a width of 640 pixels (height fitted to content).
    • Render styling is randomized: fonts are uniformly sampled from Google Fonts, font sizes are uniformly sampled from 12pt to 36pt, and text colors are uniformly sampled across all RGB values.
    • The model is trained for 30,000 steps to transcribe the text with a maximum input sequence length of 128 patches.

    This warmup stage imparts basic character recognition capabilities, stabilizing subsequent screenshot parsing pretraining, speeding up optimization convergence, and improving downstream finetuning performance.

  5. Knowl 5 — Pix2Struct Architecture and Pretraining Configurations

    experimental setup

    Pix2Struct is an image-encoder-text-decoder transformer model using a Vision Transformer (ViT) encoder and an autoregressive text decoder.

    Model variants:

    • Pix2Struct-Base: 282 million parameters, 12 encoder layers, 12 decoder layers, and a hidden dimension of 768.
    • Pix2Struct-Large: 1.3 billion parameters, 18 encoder layers, 18 decoder layers, and a hidden dimension of 1536.

    Pretraining specifications:

    • Pretraining Data: 80 million pairs of web page screenshots and HTML DOM trees collected from URLs in the C4 corpus, rendered at a viewport of 1024×10241024 \times 1024 pixels.
    • Optimization: Adafactor optimizer with a linear learning rate warmup over 1,000 steps to a peak learning rate of 0.01, followed by cosine decay to 0.
    • Training Schedule: Both models undergo a 30,000-step reading warmup on synthetic BooksCorpus renders (sequence length 128 patches). Pix2Struct-Base is then trained for 270,000 steps on screenshot parsing with a batch size of 2,048 across 64 Cloud TPUs (sequence length 2,048 patches). Pix2Struct-Large is trained for 170,000 steps with a batch size of 1,024 across 128 Cloud TPUs (sequence length 2,048 patches). The decoder target sequence length is 128 tokens (up to 1,024 characters).
    • Validation Performance: Pix2Struct-Base achieves 30 BLEU and Pix2Struct-Large achieves 32 BLEU on the pretraining validation set.
  6. Knowl 6 — Downstream Benchmark Performance Across Visually-Situated Language Domains

    data/table

    Pix2Struct was evaluated on nine single-task benchmarks across four domains: Illustrations (ChartQA, AI2D, OCR-VQA), User Interfaces (RefExp, Widget Captioning, Screen2Words), Natural Images (TextCaps), and Documents (DocVQA, InfographicVQA).

    Metrics used:

    • Relaxed Accuracy (RA, %) for ChartQA.
    • Exact Match (EM, %) for AI2D, OCR-VQA, and RefExp.
    • CIDEr score for Widget Captioning, Screen2Words, and TextCaps.
    • Average Normalized Levenshtein Similarity (ANLS, %) for DocVQA and InfographicVQA.
    Method Pretraining ChartQA AI2D OCR-VQA RefExp WidgetCap Screen2Words TextCaps DocVQA InfoVQA
    Pipelined / Specialized SotA
    VisionTaPas Pipeline 45.5 - - - - - - - -
    DQA-NET Pipeline - 38.5 - - - - - - -
    LATr OCR Pipeline - - 67.5 - - - - - -
    UI Bert View Hierarchy - - - 90.8 - - - - -
    VUT View Hierarchy - - - - 97.0 64.3 - - -
    PaLI Captioning + OCR - - - - - - 160.4 - -
    UDOP OCR Pipeline - - - - - - - 84.7 47.4
    Pixel-Only Models
    GIT2 Image Captioning - - 70.3 - - - 145.0 - -
    Donut Synthetic OCR 41.8 30.8 66.0 - 127.4 56.4 74.4 67.5 11.6
    Pix2Struct-Base Screenshot Parsing 56.0 40.9 69.4 92.2 133.1 107.0 88.0 72.1 38.2
    Pix2Struct-Large Screenshot Parsing 58.6 42.1 71.3 94.2 136.7 109.4 95.5 76.6 40.0

    Pix2Struct-Large sets new state-of-the-art results on six of nine benchmarks (ChartQA, AI2D, OCR-VQA, RefExp, Widget Captioning, and Screen2Words) and outperforms the pixel-only baseline Donut across all nine benchmarks by 9 to 53 points.

  7. Knowl 7 — Ablation Analysis of Pretraining Objectives

    data/table

    To evaluate the contributions of the pretraining components, a Pix2Struct-Base model was trained for a fixed budget of 100,000 total pretraining steps across four configurations and evaluated on the validation sets of DocVQA (ANLS, %), Widget Captioning (CIDEr), and TextCaps (CIDEr):

    • Full: 30,000 warmup steps on BooksCorpus followed by 70,000 screenshot parsing steps with 50% text masking.
    • -- Warmup: 100,000 screenshot parsing steps with text masking directly from random initialization.
    • -- Masking: 30,000 warmup steps followed by 70,000 screenshot parsing steps without text masking.
    • -- Screenshot Parsing: 100,000 warmup steps on linear text snippets from BooksCorpus (no screenshot parsing).
    Pretraining Configuration DocVQA (ANLS) Widget Captioning (CIDEr) TextCaps (CIDEr)
    Full 67.8 137.5 84.2
    – Warmup 56.2 128.0 71.7
    – Masking 55.7 129.4 77.4
    – Screenshot Parsing 12.2 35.1 24.2

    Removing screenshot parsing causes severe performance degradation across all tasks, demonstrating that reading linear text alone cannot support visually-situated language tasks. Warmup reading pretraining and contextual text masking both provide substantial additive gains.

  8. Knowl 8 — Input Resolution and Sequence Length Scaling Behavior

    empirical result

    Pixel-only visual language models show high sensitivity to input sequence length and image resolution:

    • Resolution Scaling on DocVQA: Document QA performance (ANLS) scales steeply with input sequence length. Pix2Struct ANLS rises from under 20% at 128 patches (2152^{15} pixels) to over 70% at 4096 patches of 16×1616 \times 16 pixels (220≈1062^{20} \approx 10^6 pixels), after which performance gains show diminishing returns.
    • Aspect Ratio Handling: During synthetic warmup training, variable-resolution patching without aspect ratio distortion outperforms both stretched patching (which distorts characters) and padded patching (which expends sequence capacity on blank padding) in transcription exact match accuracy across all training step intervals.
    • Inference Speed: Measured on a Cloud TPU v3-8 at a sequence length of 4096 patches (1M pixels) with autoregressive decoding up to 32 tokens, Pix2Struct-Base processes 62 documents per second, and Pix2Struct-Large processes 20 documents per second on DocVQA.
  9. Knowl 9 — Limitations of Pixel-Only Visual Language Understanding

    limitation

    Pix2Struct and pixel-only visual language models exhibit several key limitations:

    1. High Resolution and Sequence Length Costs: Accurate text recognition in visually-situated language requires high image resolutions (up to 4096 patches / 10610^6 pixels). Due to the quadratic attention cost in transformer encoders, pixel-only processing is computationally heavier than OCR pipelines that decouple high-resolution text detection from lower-resolution visual encoding.
    2. Performance Gap on Text-Heavy Documents: On dense, purely textual documents (such as DocVQA), specialized OCR pipelines leveraging large pretrained language models (e.g., UDOP at 84.7 ANLS vs. Pix2Struct-Large at 76.6 ANLS) outperform pixel-only models because text encoders process longer token sequences more compactly.
    3. Web Pretraining Noise and Biases: Pretraining on scraped web screenshots (from the C4 dataset) inherits web noise, boilerplate markup, and potential web-scale biases and harmful content.

Coverage note — All primary contributions, methods, pretraining configurations, downstream benchmark evaluations, ablations, resolution scaling analyses, and stated limitations have been included as knowls.

References

  1. 1.Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al. Cm3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520, 2022a.
  2. 2.Aghajanyan, A., Okhonko, D., Lewis, M., Joshi, M., Xu, H., Ghosh, G., and Zettlemoyer, L. HTLM: hypertext pre-training and prompting of language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022, 2022b. URL https://openreview.net/forum?id=P-pPW1nxf1r.
  3. 3.Appalaraju, S., Jasani, B., Kota, B. U., Xie, Y., and Manmatha, R. DocFormer: End-to-end Transformer for document understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 993–1003, October 2021.
  4. 4.Bai, C., Zang, X., Xu, Y., Sunkara, S., Rastogi, A., Chen, J., and Aguera y Arcas, B. Uibert: Learning generic multimodal representations for ui understanding. In Zhou, Z.-H. (ed.), Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pp. 1705–1712. International Joint Conferences on Artificial Intelligence Organization, 8 2021. doi: 10.24963/ijcai.2021/235. URL https://doi.org/10.24963/ijcai.2021/235. Main Track.
  5. 5.Biten, A. F., Litman, R., Xie, Y., Appalaraju, S., and Manmatha, R. Latr: Layout-aware transformer for scene-text vqa. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16548–16558, 2022.
  6. 6.Borchmann, Ł., Pietruszka, M., Stanislawek, T., Jurkiewicz, D., Turski, M., Szyndler, K., and Gralinski, F. Due: End-to-end document understanding benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021.
  7. 7.Chen, C., Anjum, S., and Gurari, D. Grounding answers for visual questions asked by visually impaired people. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19098–19107, 2022a.
  8. 8.Chen, J., Chen, C., Xing, Z., Xu, X., Zhu, L., Li, G., and Wang, J. Unblind your apps: Predicting natural-language labels for mobile gui components by deep learning. 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE), pp. 322–334, 2020.
  9. 9.Chen, T., Saxena, S., Li, L., Fleet, D. J., and Hinton, G. Pix2seq: A language modeling framework for object detection. arXiv preprint arXiv:2109.10852, 2021.
  10. 10.Chen, T., Saxena, S., Li, L., Lin, T.-Y., Fleet, D. J., and Hinton, G. E. A unified sequence interface for vision tasks. Advances in Neural Information Processing Systems, 35:31333–31346, 2022b.
  11. 11.Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., Kolesnikov, A., Puigcerver, J., Ding, N., Rong, K., Akbari, H., Mishra, G., Xue, L., Thapliyal, A., Bradbury, J., Kuo, W., Seyedhosseini, M., Jia, C., Ayan, B. K., Riquelme, C., Steiner, A., Angelova, A., Zhai, X., Houlsby, N., and Soricut, R. Pali: A jointly-scaled multilingual language-image model, 2022c. URL https://arxiv.org/abs/2209.06794.
  12. 12.Davis, B., Morse, B., Price, B., Tensmeyer, C., Wigington, C., and Morariu, V. End-to-end document recognition and understanding with Dessurt. In Text in everything ECCV workshop, 2022. URL https://arxiv.org/abs/2203.16618.
  13. 13.Deng, Y., Kanervisto, A., Ling, J., and Rush, A. M. Image-to-markup generation with coarse-to-fine attention. In International Conference on Machine Learning, pp. 980–989. PMLR, 2017.
  14. 14.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  15. 15.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  16. 16.Eisenschlos, J., Krichene, S., and Müller, T. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 281–296, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.27. URL https://aclanthology.org/2020.findings-emnlp.27.
  17. 17.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6904–6913, 2017.
  18. 18.He, Z., Sunkara, S., Zang, X., Xu, Y., Liu, L., Wichers, N., Schubiner, G., Lee, R., and Chen, J. Actionbert: Leveraging user actions for semantic understanding of user interfaces. In 35th AAAI Conference on Artificial Intelligence, AAAI 2021, 35th AAAI Conference on Artificial Intelligence, AAAI 2021, pp. 5931–5938. Association for the Advancement of Artificial Intelligence, 2021. Publisher Copyright: Copyright © 2021, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.; 35th AAAI Conference on Artificial Intelligence, AAAI 2021 ; Conference date: 02-02-2021 Through 09-02-2021.
  19. 19.Huang, Y., Lv, T., Cui, L., Lu, Y., and Wei, F. LayoutLMv3: Pre-training for document ai with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia, 2022.
  20. 20.Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A. A diagram is worth a dozen images. In European conference on computer vision, pp. 235–251. Springer, 2016.
  21. 21.Kim, G., Hong, T., Yim, M., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S. Donut: Document understanding transformer without OCR. In ECCV, 2022. URL https://arxiv.org/abs/2111.15664.
  22. 22.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871–7880, July 2020. doi: 10.18653/v1/2020.acl-main.703. URL https://aclanthology.org/2020.acl-main.703.
  23. 23.Li, C., Bi, B., Yan, M., Wang, W., Huang, S., Huang, F., and Si, L. StructuralLM: Structural pre-training for form understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 6309–6318, Online, August 2021a. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.493. URL https://aclanthology.org/2021.acl-long.493.
  24. 24.Li, G. and Li, Y. Spotlight: Mobile UI understanding using vision-language models with a focus. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=9yE2xEj0BH7.
  25. 25.Li, J., Xu, Y., Cui, L., and Wei, F. MarkupLM: Pre-training of text and markup language for visually rich document understanding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6078–6087, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.420. URL https://aclanthology.org/2022.acl-long.420.
  26. 26.Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J. Mapping natural language instructions to mobile UI action sequences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8198–8210, Online, July 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.729. URL https://www.aclweb.org/anthology/2020.acl-main.729.
  27. 27.Li, Y., Li, G., He, L., Zheng, J., Li, H., and Guan, Z. Widget captioning: Generating natural language description for mobile user interface elements. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5495–5510, Online, November 2020b. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.443. URL https://aclanthology.org/2020.emnlp-main.443.
  28. 28.Li, Y., Li, G., Zhou, X., Dehghani, M., and Gritsenko, A. Vut: Versatile ui transformer for multi-modal multi-task user interface modeling. arXiv preprint arXiv:2112.05692, 2021b.
  29. 29.Liu, T. F., Craft, M., Situ, J., Yumer, E., Mech, R., and Kumar, R. Learning design semantics for mobile apps. Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology, 2018.
  30. 30.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  31. 31.Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2263–2279, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.177. URL https://aclanthology.org/2022.findings-acl.177.
  32. 32.Mathew, M., Karatzas, D., and Jawahar, C. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 2200–2209, 2021.
  33. 33.Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1697–1706, 2022.
  34. 34.Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A. Ocr-vqa: Visual question answering by reading text in images. In 2019 international conference on document analysis and recognition (ICDAR), pp. 947–952. IEEE, 2019.
  35. 35.Powalski, R., Borchmann, Ł., Jurkiewicz, D., Dwojak, T., Pietruszka, M., and Pałka, G. Going full-tilt boogie on document understanding with text-image-layout transformer. In International Conference on Document Analysis and Recognition, pp. 732–747. Springer, 2021.
  36. 36.Press, O., Smith, N., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=R8sQPpGCv0.
  37. 37.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. Improving language understanding by generative pre-training. OpenAI, 2018.
  38. 38.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  39. 39.Rust, P., Lotz, J. F., Bugliarello, E., Salesky, E., de Lhoneux, M., and Elliott, D. Language modelling with pixels. arXiv preprint arXiv:2207.06991, 2022.
  40. 40.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1238. URL https://aclanthology.org/P18-1238.
  41. 41.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pp. 4596–4604. PMLR, 2018.
  42. 42.Sidorov, O., Hu, R., Rohrbach, M., and Singh, A. Textcaps: a dataset for image captioningwith reading comprehension. In European Conference on Computer Vision, 2020.
  43. 43.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8317–8326, 2019.
  44. 44.Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., Zeng, M., Zhang, C., and Bansal, M. Unifying vision, text, and layout for universal document processing. arXiv preprint arXiv:2212.02623, 2022.
  45. 45.Touvron, H., Vedaldi, A., Douze, M., and Jegou, H. Fixing the train-test resolution discrepancy. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/d03a857a23b5285736c4d55e0bb067c8-Paper.pdf.
  46. 46.Wang, B., Li, G., Zhou, X., Chen, Z., Grossman, T., and Li, Y. Screen2words: Automatic mobile ui summarization with multimodal learning. In The 34th Annual ACM Symposium on User Interface Software and Technology, UIST ’21, pp. 498–510, New York, NY, USA, 2021a. Association for Computing Machinery. ISBN 9781450386357. doi: 10.1145/3472749.3474765. URL https://doi.org/10.1145/3472749.3474765.
  47. 47.Wang, J., Tang, J., and Luo, J. Multimodal attention with image text spatial relationship for ocr-based image captioning. In Proceedings of the 28th ACM International Conference on Multimedia, pp. 4337–4345, 2020.
  48. 48.Wang, J., Yang, Z., Hu, X., Li, L., Lin, K., Gan, Z., Liu, Z., Liu, C., and Wang, L. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022a.
  49. 49.Wang, Q., Fang, Y., Ravula, A., Feng, F., Quan, X., and Liu, D. Webformer: The web-page transformer for structure information extraction. In Proceedings of the ACM Web Conference 2022, pp. 3124–3133, 2022b.
  50. 50.Wang, Z., Yu, J., Yu, A. W., Dai, Z., Tsvetkov, Y., and Cao, Y. Simvlm: Simple visual language model pretraining with weak supervision. CoRR, abs/2108.10904, 2021b. URL https://arxiv.org/abs/2108.10904.
  51. 51.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022. URL https://openreview.net/forum?id=yzkSU5zdwD. Survey Certification.
  52. 52.Wu, J., Zhang, X., Nichols, J., and Bigham, J. P. Screen parsing: Towards reverse engineering of UI models from screenshots. In The 34th Annual ACM Symposium on User Interface Software and Technology, pp. 470–483, 2021.
  53. 53.Xu, Y., Xu, Y., Lv, T., Cui, L., Wei, F., Wang, G., Lu, Y., Florencio, D., Zhang, C., Che, W., Zhang, M., and Zhou, L. Layoutlmv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) 2021, 2021.
  54. 54.Yang, Z., Lu, Y., Wang, J., Yin, X., Florencio, D., Wang, L., Zhang, C., Zhang, L., and Luo, J. Tap: Text-aware pre-training for text-vqa and text-caption. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8751–8761, 2021.
  55. 55.Zhang, X., de Greef, L., Swearngin, A., White, S., Murray, K., Yu, L., Shan, Q., Nichols, J., Wu, J., Fleizach, C., et al. Screen recognition: Creating accessibility metadata for mobile applications from pixels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–15, 2021.
  56. 56.Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. 2015 IEEE International Conference on Computer Vision (ICCV), pp. 19–27, 2015.

Citation

MLA
Lee, K., et al. “Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding”. International Conference on Machine Learning, vol. 202, 2023, pp. 18893–912, https://proceedings.mlr.press/v202/lee23g.html.
APA
Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M.-W., & Toutanova, K. (2023). Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding. International Conference on Machine Learning, 202, 18893–18912. https://proceedings.mlr.press/v202/lee23g.html
Chicago
Lee, K., M. Joshi, I. R. Turc, et al. 2023. “Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding”. International Conference on Machine Learning 202: 18893–912. https://proceedings.mlr.press/v202/lee23g.html.
Harvard
Lee, K. et al. (2023) “Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding”, International Conference on Machine Learning. PMLR, pp. 18893–18912. Available at: https://proceedings.mlr.press/v202/lee23g.html.
Vancouver
1. Lee K, Joshi M, Turc IR, Hu H, Liu F, Eisenschlos JM, Khandelwal U, Shaw P, Chang M-W, Toutanova K (2023) Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding. In: International Conference on Machine Learning. PMLR, pp 18893–18912

BibTeX

@InProceedings{pmlr-v202-lee23g,
  title = 	 {{P}ix2{S}truct: Screenshot Parsing as Pretraining for Visual Language Understanding},
  author =       {Lee, Kenton and Joshi, Mandar and Turc, Iulia Raluca and Hu, Hexiang and Liu, Fangyu and Eisenschlos, Julian Martin and Khandelwal, Urvashi and Shaw, Peter and Chang, Ming-Wei and Toutanova, Kristina},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {18893--18912},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/lee23g/lee23g.pdf},
  url = 	 {https://proceedings.mlr.press/v202/lee23g.html},
  abstract = 	 {Visually-situated language is ubiquitous—sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/