OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
Jianqiang WanSibo SongWenwen YuYuliang LiuWenqing ChengFei HuangXiang BaiCong YaoZhibo Yang
Introduces OmniParser, a unified encoder-decoder framework that simultaneously handles text spotting, key information extraction, and table recognition by decoupling structured point sequences from region and text content generation.
Organizations face significant operational challenges in automatically extracting structured information from text-rich images such as scanned receipts, forms, and scientific tables. Existing approaches typically rely on isolated specialist models for distinct subtasks or generalist models that sacrifice precision, lack transparency, and depend on external text-recognition engines. This fragmentation increases pipeline complexity, development overhead, and operational risk across document processing workflows.
The article introduces OmniParser, a unified vision-based framework designed to execute three primary document parsing tasks simultaneously within a single architecture: text spotting (detecting and transcribing text), key information extraction (identifying semantic fields), and table recognition (extracting table structure and cell contents end-to-end).
The approach uses a two-stage encoder-decoder system with decoupled components. In the first stage, a dedicated module predicts a sequence of central coordinate points paired with structural markup tags (such as table or entity labels). In the second stage, parallel decoders use these point locations to generate precise polygon boundaries and character transcriptions. Credibility was established by pre-training the model on public scene text datasets using targeted spatial and content prompting strategies, followed by fine-tuning across seven established benchmark datasets.
OmniParser established new state-of-the-art results for end-to-end text spotting on curved and arbitrary-shaped text benchmarks, outperforming previous top models by 1.5% and 3.2% without external vocabulary aids. On key information extraction benchmarks, it attained top-tier performance—including an 84.8% field-level score on the CORD receipt dataset—while uniquely providing exact visual localizations that previous generative methods could not deliver. For table recognition, the unified framework surpassed specialized end-to-end models across standard benchmarks while maintaining faster inference speeds (1.3 frames per second compared to 0.8 for earlier end-to-end baselines) and eliminating long-sequence attention failure.
These findings demonstrate that organizations can replace complex, multi-model document pipelines with a single unified framework without sacrificing accuracy. Decoupling spatial point generation from text transcription significantly reduces sequence complexity, speeds up processing, and provides full visual interpretability. Furthermore, achieving top-tier performance on formal documents despite pre-training exclusively on scene text illustrates substantial architectural generalization and potential savings in training data preparation.
Engineering and product teams evaluating automated document workflows should consider adopting two-stage point-conditioned architectures for visual parsing tasks. Where unified deployments are planned, maintaining distinct model parameters across decoders is recommended rather than sharing weights, as experiments showed weight-sharing reduced overall spotting accuracy. Future development work should focus on extending the architecture to non-text elements such as charts, graphics, and full document layout analysis.
Readers should note that OmniParser relies on precise point annotations during training, which may require additional annotation effort if such data is absent in proprietary datasets. However, high confidence in the framework’s core capabilities is supported by rigorous evaluations against both unified and task-specific state-of-the-art baselines across diverse benchmarks.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Introduces a foundational unified sequence-to-sequence formulation using coordinate and discrete visual tokens that directly precedes OmniParser's point-conditioned sequence generation paradigm.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Establishes multi-task vision-language pre-training with fine-grained visual coordinate grounding and text reading that OmniParser specializes into structured document parsing.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). Presents the standard DocVQA formulation and benchmark that motivated visually-situated text parsing models like OmniParser.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Pioneered reading scene text and integrating optical character recognition tokens into vision-language architectures.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Provides a comprehensive architectural survey of multimodal large language model design patterns and prompt-based encoder-decoder systems foundational to OmniParser.
- Paper: Qianfan-OCR: A Unified End-to-End Model for Document Intelligence, Daxiang Dong et al. (2026). Extends unified end-to-end document parsing paradigms to large-scale industrial document intelligence across 192 languages and structured Markdown generation.
- Paper: DeepSeek-OCR 2: Visual Causal Flow, Haoran Wei et al. (2026). Advances visually-situated document parsing by introducing visual causal flow and semantic reordering within the visual encoder for complex documents and tables.
- Paper: Unlimited OCR Works, Youyang Yin et al. (2026). Builds upon unified visual text parsing by scaling end-to-end recognition to multi-page documents via reference sliding window attention.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Generalizes multi-task document understanding and optical character recognition across long-context, high-resolution multimodal foundation architectures.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). Applies model and test-time reasoning scaling to document and chart parsing within advanced open-source multimodal frameworks.
