Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Kenton LeeMandar JoshiIulia Raluca TurcHexiang HuFangyu LiuJulian Martin EisenschlosUrvashi KhandelwalPeter ShawMing-Wei ChangKristina Toutanova
Presents Pix2Struct, an image-to-text model pretrained by parsing masked web screenshots into simplified HTML, enabling a single OCR-free architecture to handle diverse visual language tasks across documents, diagrams, and user interfaces.
Modern digital environments—such as documents, mobile applications, and web pages—present language and visual information holistically. Previous artificial intelligence methods for processing this visually-situated text have relied on fragmented, domain-specific pipelines that depend heavily on external tools, such as optical character recognition (OCR) or specialized application metadata. These multi-step pipelines increase engineering complexity, limit adaptability across different formats, and elevate computational costs.
The article demonstrates Pix2Struct, a unified, pretrained image-to-text model designed for purely visual language understanding without requiring external OCR tools or domain-specific intermediate steps. The goal was to establish a single general-purpose framework capable of handling varied visual language tasks by learning directly from raw pixel inputs.
The approach uses a vision transformer architecture trained on 80 million web page screenshots paired with simplified HTML source code. During this self-supervised pretraining, the model learns to reconstruct the underlying HTML structure from partially masked screenshots, effectively blending OCR, masked language modeling, and image captioning into one objective. To handle the varied aspect ratios of documents and user interfaces without distortion, the authors introduced a variable-resolution patching mechanism. For downstream tasks, text prompts (such as questions) are rendered directly onto the input image, processing all information through a single visual channel. Two model variants were evaluated across nine benchmarks spanning four domains: illustrations, user interfaces, natural images, and documents.
The evaluation yielded several key findings. Pix2Struct established a new state of the art in six of the nine benchmarks. In low-resource domains, it outperformed existing specialized systems, improving chart question-answering accuracy from 45.5 to 58.6 and mobile screen summarization score from 64.3 to 109.4. Across all tested benchmarks, it substantially outperformed Donut, the leading OCR-free baseline, by 9 to 53 points. While the model trailed top-performing OCR-reliant pipelines in text-dense document tasks and models trained on billions of caption pairs for natural images, it delivered competitive performance using substantially less domain-specific training.
These results indicate that end-to-end visual pretraining from web markup provides a viable, scalable alternative to specialized multi-stage pipelines. By eliminating reliance on external OCR tools, organizations can reduce architectural complexity, streamline operational maintenance, and deploy a single model across diverse visual and document formats.
Decision-makers and engineering teams should consider piloting unified, pixel-only models for user interface and illustration workflows where specialized tools are brittle or unavailable. For text-heavy document extraction, organizations should weigh the simplicity of an end-to-end model against the marginal accuracy advantage of established OCR pipelines. Future development should focus on scaling pretraining data and exploring long-range architectures to further close performance gaps in document-dense domains.
The primary limitation of the model is its high sensitivity to image resolution, as high-resolution inputs demand significant computational memory and sequence lengths during processing. Additionally, the pretraining data relies on web corpora that require careful curation to mitigate undesirable content risks. While confidence is high in the model's effectiveness across user interfaces and illustrative tasks, readers should exercise caution when deploying purely pixel-based models to high-density document tasks without task-specific validation.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). Introduces the Document Visual Question Answering benchmark that establishes the core challenge of visually-situated language comprehension addressed by Pix2Struct.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Presents masked image modeling with vision transformers, providing the conceptual foundation for masked patch reconstruction used in Pix2Struct's pretraining.
- Paper: An Empirical Study of Training End-to-End Vision-and-Language Transformers, Zi-Yi Dou et al. (2022). Provides a comprehensive empirical study on architectural choices and objectives for end-to-end vision-and-language transformers.
- Paper: Generative Pretraining From Pixels, Mark Chen et al. (2020). Demonstrates the feasibility of training generative sequence transformers directly on raw visual pixels without domain-specific inductive biases.
- Paper: MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering, Fangyu Liu et al. (2023). Directly extends Pix2Struct by continual pretraining on chart derendering and mathematical reasoning tasks to enhance chart understanding.
- Paper: From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces, Peter Shaw et al. (2023). Applies the screenshot-parsing pretraining paradigm to build autonomous agents that navigate graphical user interfaces from raw pixels.
- Paper: Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus, Gang Li et al. (2023). Builds upon vision-only screenshot modeling to interpret mobile user interfaces using dynamic region focus rather than fragile view hierarchies.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). Introduces a large-scale visual question-answering benchmark specifically designed to assess pixel-based reading comprehension across full mobile screen captures.
