OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Jianqiang WanSibo SongWenwen YuYuliang LiuWenqing ChengFei HuangXiang BaiCong YaoZhibo Yang

article2024CVPR107 citations

Introduces OmniParser, a unified encoder-decoder framework that simultaneously handles text spotting, key information extraction, and table recognition by decoupling structured point sequences from region and text content generation.

Listen

Organizations face significant operational challenges in automatically extracting structured information from text-rich images such as scanned receipts, forms, and scientific tables. Existing approaches typically rely on isolated specialist models for distinct subtasks or generalist models that sacrifice precision, lack transparency, and depend on external text-recognition engines. This fragmentation increases pipeline complexity, development overhead, and operational risk across document processing workflows.

The article introduces OmniParser, a unified vision-based framework designed to execute three primary document parsing tasks simultaneously within a single architecture: text spotting (detecting and transcribing text), key information extraction (identifying semantic fields), and table recognition (extracting table structure and cell contents end-to-end).

The approach uses a two-stage encoder-decoder system with decoupled components. In the first stage, a dedicated module predicts a sequence of central coordinate points paired with structural markup tags (such as table or entity labels). In the second stage, parallel decoders use these point locations to generate precise polygon boundaries and character transcriptions. Credibility was established by pre-training the model on public scene text datasets using targeted spatial and content prompting strategies, followed by fine-tuning across seven established benchmark datasets.

OmniParser established new state-of-the-art results for end-to-end text spotting on curved and arbitrary-shaped text benchmarks, outperforming previous top models by 1.5% and 3.2% without external vocabulary aids. On key information extraction benchmarks, it attained top-tier performance—including an 84.8% field-level score on the CORD receipt dataset—while uniquely providing exact visual localizations that previous generative methods could not deliver. For table recognition, the unified framework surpassed specialized end-to-end models across standard benchmarks while maintaining faster inference speeds (1.3 frames per second compared to 0.8 for earlier end-to-end baselines) and eliminating long-sequence attention failure.

These findings demonstrate that organizations can replace complex, multi-model document pipelines with a single unified framework without sacrificing accuracy. Decoupling spatial point generation from text transcription significantly reduces sequence complexity, speeds up processing, and provides full visual interpretability. Furthermore, achieving top-tier performance on formal documents despite pre-training exclusively on scene text illustrates substantial architectural generalization and potential savings in training data preparation.

Engineering and product teams evaluating automated document workflows should consider adopting two-stage point-conditioned architectures for visual parsing tasks. Where unified deployments are planned, maintaining distinct model parameters across decoders is recommended rather than sharing weights, as experiments showed weight-sharing reduced overall spotting accuracy. Future development work should focus on extending the architecture to non-text elements such as charts, graphics, and full document layout analysis.

Readers should note that OmniParser relies on precise point annotations during training, which may require additional annotation effort if such data is absent in proprietary datasets. However, high confidence in the framework’s core capabilities is supported by rigorous evaluations against both unified and task-specific state-of-the-art baselines across diverse benchmarks.

arXiv: 2403.19128
Cover for OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Abstract

Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to address the challenging problem of VsTP. However, due to the diversified targets and heterogeneous schemas, previous works usually design task-specific architectures and objectives for individual tasks, which inadvertently leads to modal isolation and complex workflow. In this paper, we propose a unified paradigm for parsing visually-situated text across diverse scenarios. Specifically, we devise a universal model, called OmniParser, which can simultaneously handle three typical visually-situated text parsing tasks: text spotting, key information extraction, and table recognition. In OmniParser, all tasks share the unified encoder-decoder architecture, the unified objective: point-conditioned text generation, and the unified input&output representation: prompt & structured sequences. Extensive experiments demonstrate that the proposed OmniParser achieves state-of-the-art (SOTA) or highly competitive performances on 7 datasets for the three visually-situated text parsing tasks, despite its unified, concise design. The code is available at AdvancedLiterateMachinery.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Task Unification
  • 3.2. Unified Architecture
  • 3.3. Pre-training Methods
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.1.1 Text Spotting
  • 4.1.2 Key Information Extraction
  • 4.1.3 Table Recognition
  • 4.2. Comparisons with State-of-The-Art
  • 5. Analysis
  • 6. Conclusions and Future Works
  • References

Knowls

  1. Knowl 1 — Two-Stage Point-Decoupled Representation for Visually-Situated Text Parsing

    model/method

    OmniParser unifies visually-situated text parsing (VsTP) across scene text spotting, key information extraction (KIE), and table recognition (TR) through a decoupled two-stage sequence generation pipeline using center points as adapters:

    1. Stage 1: Structured Points Sequence Generation. An autoregressive Structured Points Decoder receives visual features and task-specific indicator prompts. It outputs a structured sequence consisting of center point tokens interleaved with task structural markup. For each text instance, continuous xx and yy center coordinates are normalized by image width and height to [0,1][0, 1] and quantized into discrete coordinate tokens in the range [0,nbins−1][0, n_{\text{bins}} - 1]. For KIE, semantic entity tags (such as \<address\> or \</address\>) enclose the respective center points. For table recognition, HTML structure tokens (such as \<tr\>, \<td\>, \</td\>, \</tr\>, \<td colspan="2"\>) structure the sequence of cell center points. Text spotting is a special case where only center points are predicted without structural tags.

    2. Stage 2: Parallel Polygon and Content Decoding. Conditioned on each center point predicted in Stage 1, two decoders execute in parallel:

      • Region Decoder: Generates a 16-point polygon coordinate sequence representing the boundary contour of the text instance, tokenized using the same coordinate discretization [0,nbins−1][0, n_{\text{bins}} - 1].
      • Content Decoder: Generates the text transcription of the text instance as a discrete character-level sequence.

    Decoupling the structural layout from dense polygons and text transcriptions dramatically shortens sequence lengths during layout parsing, reducing attention drift and error accumulation.

  2. Knowl 2 — OmniParser Architecture and Point-Conditioned Training Objective

    model/method

    The OmniParser architecture consists of an image encoder and three transformer decoders trained with a token-weighted negative log-likelihood objective:

    • Image Encoder: A Swin Transformer Base (Swin-B pre-trained on ImageNet-22k) extracts multi-scale visual features at strides 4, 8, 16, and 32 with respect to the input image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}. A Feature Pyramid Network (FPN) fuses these features into visual embeddings v={vi∈Rd∣1≤i≤n}\mathbf{v} = \{v_i \in \mathbb{R}^d \mid 1 \le i \le n\}, where nn is the spatial size of the feature map after FPN and d=512d = 512.

    • Decoders: Three separate decoders—Structured Points Decoder, Region Decoder, and Content Decoder—share an identical architecture comprising 4 transformer decoder layers with 8 attention heads, a hidden dimension of 512, an MLP expansion ratio of 4, and pre-layer normalization. The decoders do not share parameters and utilize uniquely initialized learned positional encodings to accommodate differing sequence lengths.

    • Objective Function: The model is trained during pre-training and fine-tuning by minimizing the negative log-likelihood loss:

    L=−∑j=kNwjlog⁡P(s~j∣v,sk:j−1)L = -\sum_{j=k}^N w_j \log P(\tilde{\mathbf{s}}_j \mid \mathbf{v}, \mathbf{s}_{k:j-1})

    where s~\tilde{\mathbf{s}} is the target sequence, NN is sequence length, v\mathbf{v} are the visual feature embeddings, and sk:j−1\mathbf{s}_{k:j-1} represents the prior token context. The first kk prompt tokens are excluded from the loss. The per-token weight wjw_j is set to wj=4.0w_j = 4.0 for structural and entity tags, and wj=1.0w_j = 1.0 for point coordinates and text character tokens.

  3. Knowl 3 — Spatial-Aware and Content-Aware Pre-Training Prompting Strategies

    model/method

    To train the Structured Points Decoder to parse text structures and semantic entities from visual features alone without linguistic text reading modules, OmniParser introduces two prompting strategies during pre-training:

    1. Spatial-Window Prompting: The decoder is conditioned on a 2-point bounding box prompt (xleft,ytop,xright,ybottom)(x_{\text{left}}, y_{\text{top}}, x_{\text{right}}, y_{\text{bottom}}). The target sequence contains only center points of text instances that lie within this spatial window. Sampling uses two patterns:

      • Fixed pattern: The window is chosen from predefined regular grid layouts (such as 2×22 \times 2 or 3×33 \times 3).
      • Random pattern: The window is randomly cropped from the image subject to covering at least 1/91/9 of the image area.
    2. Prefix-Window Prompting: The decoder is conditioned on a 2-character prompt (start,end)(start, end) denoting a contiguous sub-range within an ordered character vocabulary (consisting of 26 uppercase letters, 26 lowercase letters, 10 digits, and 34 ASCII punctuation symbols). The decoder is instructed to output only center points for text instances whose initial character falls within [start,end][start, end], ignoring all other instances.

    During task fine-tuning and inference across full images, prompts are set to the full coordinate range [0,0,nbins−1,nbins−1][0, 0, n_{\text{bins}}-1, n_{\text{bins}}-1] and full dictionary range ['!', '~'].

  4. Knowl 4 — OmniParser Pre-training and Fine-Tuning Protocols

    experimental setup

    OmniParser is trained in a multi-stage regimen using the AdamW optimizer with cosine learning rate scheduling:

    • Pre-Training Datasets: Pre-trained exclusively on scene text datasets: Curved SynthText, ICDAR 2013, ICDAR 2015, MLT 2017, Total-Text, TextOCR, HierText, COCO Text, and Open Images V5. Center points are arranged in raster scan order.
    • Two-Stage Pre-Training Schedule:
      • Stage 1: Image resolution 768×768768 \times 768, batch size 128, initial learning rate 5×10−45 \times 10^{-4}, for 500k iterations.
      • Stage 2: Image resolution 1920×19201920 \times 1920, batch size 16, initial learning rate 2.5×10−42.5 \times 10^{-4}, for 200k iterations. Both stages employ a 5k-step linear warmup followed by linear decay to 0, along with data augmentations: instance-aware random cropping, random rotation (between −90∘-90^\circ and 90∘90^\circ), random resizing, and color jittering.
    • Downstream Task Fine-Tuning:
      • Text Spotting: Fine-tuned for 20k steps with learning rate 1×10−41 \times 10^{-4}.
      • Key Information Extraction: Fine-tuned on CORD or SROIE for 200k steps with learning rate 1×10−41 \times 10^{-4}.
      • Table Recognition: On PubTabNet and FinTabNet, maximum sequence lengths for Structured Points Decoder and Content Decoder are set to 1,500 and 200 tokens, respectively. The Structured Points Decoder is trained for 400k steps and Content Decoder for 200k steps with learning rate 1×10−41 \times 10^{-4}.
  5. Knowl 5 — Scene Text Spotting Performance on Benchmark Datasets

    data/table

    OmniParser evaluated on Total-Text (arbitrary shapes), CTW1500 (curved text lines), and ICDAR 2015 (incidental scene text) against specialist and generalist text spotters. Metrics include Precision (P), Recall (R), F-measure (F) for detection, and End-to-End (E2E) recognition accuracy without lexicon (None), with Full lexicon, or under Strong (S), Weak (W), and Generic (G) lexicons:

    Method Total-Text E2E CTW1500 E2E ICDAR 2015 Detection ICDAR 2015 E2E
    None Full None Full P R F S W G
    ABCNet v2 70.4 78.1 57.5 77.2 90.4 86.0 88.1 82.7 78.5 73.0
    TESTR 73.3 83.9 56.0 81.5 90.3 89.7 90.0 85.2 79.4 73.6
    SPTS 74.2 82.4 63.6 83.8 - - - 77.5 70.2 65.8
    UNITS 82.2 88.0 66.4 82.3 91.0 94.0 92.5 89.0 84.1 80.3
    DeepSolo 82.5 88.7 56.7 - 92.5 87.2 89.8 88.0 83.5 79.1
    OmniParser 84.0 88.9 66.8 85.1 90.3 91.0 90.7 89.6 84.5 79.9

    OmniParser achieves state-of-the-art unconstrained (None) E2E recognition on Total-Text (84.0%84.0\%) and CTW1500 (66.8%66.8\%), surpassing previous single-model architectures while maintaining a unified 16-point polygonal representation across all datasets.

  6. Knowl 6 — End-to-End Key Information Extraction Performance on CORD and SROIE

    data/table

    Comparison of OmniParser on key information extraction (KIE) benchmarks against end-to-end OCR-free and generation-based models. Evaluation metrics are field-level F1 score (F1) and Tree-Edit-Distance-based accuracy (Acc):

    Method Localization Ability CORD SROIE
    F1 Acc F1 Acc
    TRIE Yes - - 82.1 -
    Donut No 84.1 90.9 83.2 92.8
    Dessurt No 82.5 - 84.9 -
    DocParser No 84.5 - 87.3 -
    SeRum No 80.5 85.8 85.6 92.8
    OmniParser Yes 84.8 88.0 85.6 93.6

    OmniParser achieves an 84.8%84.8\% field-level F1 score on CORD and a 93.6%93.6\% tree-edit-distance accuracy on SROIE. Unlike generation-based models (such as Donut, Dessurt, and DocParser), OmniParser provides exact spatial localization for all extracted key entities while being pre-trained solely on scene text without large document corpora.

  7. Knowl 7 — End-to-End Table Structure and Content Recognition Results

    data/table

    Comparison of end-to-end table recognition models on PubTabNet (scientific tables) and FinTabNet (financial tables). Performance is measured using Tree-Edit-Distance-based Similarity for structure only (S-TEDS) and complete table structure with cell content (TEDS):

    Dataset Method Input Size Decoder Max Len. S-TEDS (%) TEDS (%)
    PubTabNet WYGIWYS 512 - - 78.60
    Donut 1,280 4,000 25.28 22.70
    EDD 512 1,800 89.90 88.30
    OmniParser 1,024 1,500 90.45 88.83
    FinTabNet Donut 1,280 4,000 30.66 29.10
    EDD 512 1,800 90.60 -
    OmniParser 1,024 1,500 91.55 89.75

    OmniParser outperforms previous end-to-end methods on both benchmarks with a max decoder length of 1,500 tokens (operating at 1.3 FPS vs. 0.8 FPS for Donut). By decoupling HTML tag generation from cell text decoding, OmniParser avoids the extreme error accumulation that impairs pure Seq2Seq models like Donut (22.70%22.70\% TEDS) on long HTML sequences.

  8. Knowl 8 — Impact of Spatial-Window and Prefix-Window Prompting on Text Spotting

    empirical result

    Ablation experiments on Total-Text and ICDAR 2015 demonstrate the complementary effects of spatial-window and prefix-window prompting during pre-training:

    Window Prompting Total-Text E2E ICDAR 2015 E2E
    Spatial Prefix None Full Strong Weak Generic
    - - 82.4 87.6 88.1 83.0 78.3
    ✓ - 82.9 88.1 88.4 83.2 78.5
    - ✓ 83.5 88.5 89.2 84.2 79.4
    ✓ ✓ 84.0 88.9 89.6 84.5 79.9

    Spatial-window prompting improves perception in the coordinate space (+0.5%+0.5\% on Total-Text None), while prefix-window prompting improves character-level semantic discrimination (+1.1%+1.1\% on Total-Text None and ICDAR 2015 Generic). Combining both strategies yields the highest performance (84.0%84.0\% Total-Text None, 89.6%89.6\% ICDAR 2015 Strong).

  9. Knowl 9 — Effect of Decoder Parameter Sharing and Visual Backbone Architecture

    empirical result

    An architectural ablation evaluates parameter sharing among the three decoders (Structured Points Decoder, Region Decoder, Content Decoder) and compares visual backbones (ResNet50 vs. Swin-B) on text spotting:

    Visual Backbone Decoder Weights Total-Text E2E ICDAR 2015 E2E
    None Full Strong Weak Generic
    ResNet50 Not Shared 82.1 87.1 88.2 83.0 78.4
    Swin-B Shared 82.5 87.3 88.5 83.2 78.7
    Swin-B Not Shared 84.0 88.9 89.6 84.5 79.9

    Using separate parameters for the three decoders outperforms parameter sharing by +1.5%+1.5\% on Total-Text None and +1.2%+1.2\% on ICDAR 2015 Generic. This confirms an inherent task discrepancy between discrete point sequence layout modeling, boundary coordinate regression, and linguistic character decoding. Additionally, Swin-B outperforms ResNet50 by +1.9%+1.9\% on Total-Text None.

  10. Knowl 10 — Limitations of the OmniParser Framework

    limitation

    The OmniParser framework has two explicit limitations:

    1. Dependence on Point Supervision: The training process requires precise word-level or text-segment center point annotations, which may not be consistently available in weakly annotated real-world document datasets.
    2. Text-Centric Modeling: The model is specialized for text, entities, and tables, and does not parse non-text visual document elements such as figures, charts, diagrams, or arbitrary graphical illustrations.

Coverage note — None omitted; all core contributions, architectural formulations, pre-training prompting methods, experimental setups, benchmark evaluations, and ablations have been fully captured.

References

  1. 1.Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. Docformer: End-to-end transformer for document understanding. In Proceedings of the IEEE/CVF international conference on computer vision, pages 993–1003, 2021.
  2. 2.Youngmin Baek, Seung Shin, Jeonghun Baek, Sungrae Park, Junyeop Lee, Daehyun Nam, and Hwalsuk Lee. Character region attention for text spotting. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16, pages 504–521. Springer, 2020.
  3. 3.Haoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu, Yinsong Liu, and Bo Ren. Query-driven generative network for document information extraction in the wild. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4261–4271, 2022.
  4. 4.Haoyu Cao, Changcun Bao, Chaohu Liu, Huang Chen, Kun Yin, Hao Liu, Yinsong Liu, Deqiang Jiang, and Xing Sun. Attention where it matters: Rethinking visual document understanding with selective region concentration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19517–19527, 2023.
  5. 5.Panfeng Cao, Ye Wang, Qiang Zhang, and Zaiqiao Meng. Genkie: Robust generative multimodal document key information extraction. arXiv preprint arXiv:2310.16131, 2023.
  6. 6.Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. Pix2seq: A language modeling framework for object detection. In International Conference on Learning Representations, 2021.
  7. 7.Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shakeri, Mostafa Dehghani, Daniel M. Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li, Ibrahim M. Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner, Yang Li, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. Pali-x: On scaling up a multilingual vision and language model. ArXiv, abs/2305.18565, 2023.
  8. 8.Chee-Kheng Ch’ng, Chee Seng Chan, and Cheng-Lin Liu. Total-text: toward orientation robustness in scene text detection. International Journal on Document Analysis and Recognition (IJDAR), 23(1):31–52, 2020.
  9. 9.Cheng Da, Peng Wang, and Cong Yao. Multi-granularity prediction with learnable fusion for scene text recognition. arXiv preprint arXiv:2307.13244, 2023.
  10. 10.Brian Davis, Bryan Morse, Brian Price, Chris Tensmeyer, Curtis Wigington, and Vlad Morariu. End-to-end document recognition and understanding with dessurt. In European Conference on Computer Vision, pages 280–296. Springer, 2022.
  11. 11.Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, and Alexander M Rush. Image-to-markup generation with coarse-to-fine attention. In International Conference on Machine Learning, pages 980–989. PMLR, 2017.
  12. 12.Mohamed Dhouib, Ghassen Bettaieb, and Aymen Shabou. Docparser: End-to-end ocr-free information extraction from visually rich documents. arXiv preprint arXiv:2304.12484, 2023.
  13. 13.Shancheng Fang, Zhendong Mao, Hongtao Xie, Yuxin Wang, Chenggang Yan, and Yongdong Zhang. Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  14. 14.Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding. arXiv preprint arXiv:2308.11592, 2023.
  15. 15.Wei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu. Textdragon: An end-to-end framework for arbitrary shaped text spotting. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9076–9085, 2019.
  16. 16.Raul Gomez, Baoguang Shi, Lluis Gomez, Lukas Numann, Andreas Veit, Jiri Matas, Serge Belongie, and Dimosthenis Karatzas. Icdar2017 robust reading challenge on coco-text. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), pages 1435–1443. IEEE, 2017.
  17. 17.Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, and Tong Sun. Unidoc: Unified pretraining framework for document understanding. Advances in Neural Information Processing Systems, 34:39–50, 2021.
  18. 18.Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang. Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4583–4592, 2022.
  19. 19.Zengyuan Guo, Yuechen Yu, Pengyuan Lv, Chengquan Zhang, Haojie Li, Zhihui Wang, Kun Yao, Jingtuo Liu, and Jingdong Wang. Trust: An accurate and end-to-end table structure recognizer using splitting-based transformers. arXiv preprint arXiv:2208.14687, 2022.
  20. 20.Tong He, Zhi Tian, Weilin Huang, Chunhua Shen, Yu Qiao, and Changming Sun. An end-to-end textspotter with explicit alignment and attention. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5020–5029, 2018.
  21. 21.Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 10767–10775, 2022.
  22. 22.Mingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu, Dahua Lin, Shenggao Zhu, Nicholas Yuan, Kai Ding, and Lianwen Jin. Swintextspotter: Scene text spotting via better synergy between text detection and text recognition. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4593–4603, 2022.
  23. 23.Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. Layoutlmv3: Pre-training for document ai with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4083–4091, 2022.
  24. 24.Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng. Improving table structure recognition with visual-alignment sequential coordinate modeling. In CVPR, pages 11134–11143, 2023.
  25. 25.Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. Icdar2019 competition on scanned receipt ocr and information extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1516–1520. IEEE, 2019.
  26. 26.Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo. Spatial dependency parsing for semi-structured document information extraction. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 330–343, 2021.
  27. 27.Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. Icdar 2013 robust reading competition. In 2013 12th international conference on document analysis and recognition, pages 1484–1493. IEEE, 2013.
  28. 28.Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al. Icdar 2015 competition on robust reading. In 2015 13th international conference on document analysis and recognition (ICDAR), pages 1156–1160. IEEE, 2015.
  29. 29.Taeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim, and Daehee Kim. Towards unified scene text spotting based on sequence generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15223–15232, 2023.
  30. 30.Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XXVIII, pages 498–517. Springer, 2022.
  31. 31.Yair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar, R Manmatha, and Pietro Perona. Towards weakly-supervised text spotting using a multi-task transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4604–4613, 2022.
  32. 32.Ilya Krylov, Sergei Nosov, and Vladislav Sovrasov. Open images v5 text annotation and yet another mask text spotter. In Asian Conference on Machine Learning, pages 379–389. PMLR, 2021.
  33. 33.Jianfeng Kuang, Wei Hua, Dingkang Liang, Mingkun Yang, Deqiang Jiang, Bo Ren, and Xiang Bai. Visual information extraction in the wild: practical dataset and end-to-end solution. In International Conference on Document Analysis and Recognition, pages 36–53. Springer, 2023.
  34. 34.Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister. Formnet: Structural encoding beyond sequential modeling in form document information extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3735–3754, 2022.
  35. 35.Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. Structurallm: Structural pre-training for form understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6309–6318, 2021.
  36. 36.Hui Li, Peng Wang, and Chunhua Shen. Towards end-to-end text spotting with convolutional recurrent neural networks. In Proceedings of the IEEE international conference on computer vision, pages 5238–5246, 2017.
  37. 37.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023.
  38. 38.Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. Selfdoc: Self-supervised document representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5652–5660, 2021.
  39. 39.Xin Li, Yan Zheng, Yiqing Hu, Haoyu Cao, Yunfei Wu, Deqiang Jiang, Yinsong Liu, and Bo Ren. Relational representation learning in visually-rich documents. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4614–4624, 2022.
  40. 40.Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai. Mask textspotter v3: Segmentation proposal network for robust scene text spotting. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, pages 706–722. Springer, 2020.
  41. 41.Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, and Xiang Bai. Real-time scene text detection with differentiable binarization and adaptive scale fusion. IEEE transactions on pattern analysis and machine intelligence, 45(1):919–931, 2022.
  42. 42.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  43. 43.Weihong Lin, Zheng Sun, Chixiang Ma, Mingze Li, Jiawei Wang, Lei Sun, and Qiang Huo. Tsrformer: Table structure recognition with transformers. In ACM MM, pages 6473–6482, 2022.
  44. 44.Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan. Fots: Fast oriented text spotting with a unified network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5676–5685, 2018.
  45. 45.Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang. Curved scene text detection via transverse and longitudinal sequence connection. Pattern Recognition, 90:337–345, 2019.
  46. 46.Yuliang Liu, Chunhua Shen, Lianwen Jin, Tong He, Peng Chen, Chongyu Liu, and Hao Chen. Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):8048–8064, 2021.
  47. 47.Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al. Spts v2: single-point scene text spotting. arXiv preprint arXiv:2301.01635, 2023.
  48. 48.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.
  49. 49.Rujiao Long, Wen Wang, Nan Xue, Feiyu Gao, Zhibo Yang, Yongpan Wang, and Gui-Song Xia. Parsing table structures in the wild. In ICCV, pages 944–952, 2021.
  50. 50.Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Towards end-to-end unified scene text detection and layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1049–1059, 2022.
  51. 51.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018.
  52. 52.Chuwei Luo, Changxu Cheng, Qi Zheng, and Cong Yao. Geolayoutlm: Geometric pre-training for visual information extraction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7092–7101, 2023.
  53. 53.Nam Tuan Ly and Atsuhiro Takasu. An end-to-end local attention based model for table recognition. In ICDAR, pages 20–36. Springer, 2023.
  54. 54.Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai. Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. In Proceedings of the European conference on computer vision (ECCV), pages 67–83, 2018.
  55. 55.Pengyuan Lyu, Weihong Ma, Hongyi Wang, Yuechen Yu, Chengquan Zhang, Kun Yao, Yang Xue, and Jingdong Wang. Gridformer: Towards accurate table structure recognition via grid prediction. In ACM MM, pages 7747–7757, 2023.
  56. 56.Ahmed Nassar, Nikolaos Livathinos, Maksym Lysak, and Peter Staar. Tableformer: Table structure understanding with transformers. In CVPR, pages 4614–4623, 2022.
  57. 57.Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, et al. Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt. In 2017 14th IAPR international conference on document analysis and recognition (ICDAR), pages 1454–1459. IEEE, 2017.
  58. 58.OpenAI. ChatGPT. https://openai.com/chatgpt, 2023. Accessed: 2023-09-27.
  59. 59.OpenAI. GPT-4. https://openai.com/gpt-4, 2023. Accessed: 2023-09-27.
  60. 60.OpenAI. GPT-4V(ision) System Card. https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023. Accessed: 2023-10-09.
  61. 61.Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. Cord: A consolidated receipt dataset for post-ocr parsing. In Document Intelligence Workshop at Neural Information Processing Systems, 2019.
  62. 62.Dezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang, Mingxin Huang, Songxuan Lai, Jing Li, Shenggao Zhu, Dahua Lin, Chunhua Shen, et al. Spts: single-point text spotting. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4272–4281, 2022.
  63. 63.Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Yuhui Cao, Weichong Yin, Yongfeng Chen, Yin Zhang, et al. Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3744–3756, 2022.
  64. 64.Liang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu, Yi Niu, Shiliang Pu, and Fei Wu. Text perceptron: Towards end-to-end arbitrary-shaped text spotting. In Proceedings of the AAAI conference on artificial intelligence, pages 11899–11907, 2020.
  65. 65.Liang Qiao, Ying Chen, Zhanzhan Cheng, Yunlu Xu, Yi Niu, Shiliang Pu, and Fei Wu. Mango: A mask attention guided one-stage scene text spotter. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2467–2476, 2021.
  66. 66.Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii, and Ying Xiao. Towards unconstrained end-to-end text spotting. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4704–4714, 2019.
  67. 67.Roi Ronen, Shahar Tsiper, Oron Anschel, Inbal Lavi, Amir Markovitz, and R Manmatha. Glass: Global to local attention for scene-text spotting. In European Conference on Computer Vision, pages 249–266. Springer, 2022.
  68. 68.Baoguang Shi, Xiang Bai, and Cong Yao. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE transactions on pattern analysis and machine intelligence, 39(11):2298–2304, 2016.
  69. 69.Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner. Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8802–8812, 2021.
  70. 70.Sibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang, Wenqing Cheng, Xiang Bai, and Cong Yao. Vision-language pre-training for boosting scene text detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15681–15691, 2022.
  71. 71.Peter WJ Staar, Michele Dolfi, Christoph Auer, and Costas Bekas. Corpus conversion service: A machine learning platform to ingest documents at scale. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 774–782, 2018.
  72. 72.Yipeng Sun, Chengquan Zhang, Zuming Huang, Jiaming Liu, Junyu Han, and Errui Ding. Textnet: Irregular text reading from images with an end-to-end trainable network. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, Revised Selected Papers, Part III 14, pages 83–99. Springer, 2019.
  73. 73.Guozhi Tang, Lele Xie, Lianwen Jin, Jiapeng Wang, Jingdong Chen, Zhen Xu, Qianying Wang, Yaqiang Wu, and Hui Li. Matchvie: Exploiting match relevancy between entities for visual information extraction. arXiv preprint arXiv:2106.12940, 2021.
  74. 74.Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal. Unifying vision, text, and layout for universal document processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19254–19264, 2023.
  75. 75.Hao Wang, Pu Lu, Hui Zhang, Mingkun Yang, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, and Wenyu Liu. All you need is boundary: Toward arbitrary-shaped text spotting. In Proceedings of the AAAI conference on artificial intelligence, pages 12160–12167, 2020.
  76. 76.Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, and Mingxiang Cai. Towards robust visual information extraction in real world: New dataset and novel solution. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2738–2745, 2021.
  77. 77.Pengfei Wang, Chengquan Zhang, Fei Qi, Shanshan Liu, Xiaoqiang Zhang, Pengyuan Lyu, Junyu Han, Jingtuo Liu, Errui Ding, and Guangming Shi. Pgnet: Real-time arbitrarily-shaped text spotting with point gathering network. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2782–2790, 2021.
  78. 78.Wenhai Wang, Enze Xie, Xiang Li, Xuebo Liu, Ding Liang, Zhibo Yang, Tong Lu, and Chunhua Shen. Pan++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5349–5367, 2021.
  79. 79.Wei Wang, Yu Zhou, Jiahao Lv, Dayan Wu, Guoqing Zhao, Ning Jiang, and Weipinng Wang. Tpsnet: Reverse thinking of thin plate splines for arbitrary shape scene text representation. In Proceedings of the 30th ACM International Conference on Multimedia, pages 5014–5025, 2022.
  80. 80.Zilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang, and Furu Wei. Layoutreader: Pre-training of text and layout for reading order detection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4735–4744, 2021.
  81. 81.Kaiwen Wei, Jie Yao, Jingyuan Zhang, Yangyang Kang, Fubang Zhao, Yating Zhang, Changlong Sun, Xin Jin, and Xin Zhang. Ppn: Parallel pointer-based network for key information extraction with complex layouts. arXiv preprint arXiv:2307.10551, 2023.
  82. 82.Linjie Xing, Zhi Tian, Weilin Huang, and Matthew R Scott. Convolutional character networks. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9126–9136, 2019.
  83. 83.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  84. 84.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192–1200, 2020.
  85. 85.Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding. arXiv preprint arXiv:2104.08836, 2021.
  86. 86.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al. Layoutlmv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2579–2591, 2021.
  87. 87.Zhibo Yang, Rujiao Long, Pengfei Wang, Sibo Song, Humen Zhong, Wenqing Cheng, Xiang Bai, and Cong Yao. Modeling entities as semantic points for visual information extraction in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15358–15367, 2023.
  88. 88.Jiaquan Ye, Xianbiao Qi, Yelin He, Yihao Chen, Dengyi Gu, Peng Gao, and Rong Xiao. Pingan-vcgroup’s solution for icdar 2021 competition on scientific literature parsing task b: table recognition to html. arXiv preprint arXiv:2105.01848, 2021.
  89. 89.Maoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu, Tongliang Liu, Bo Du, and Dacheng Tao. Deepsolo: Let transformer decoder with explicit points solo for text spotting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19348–19357, 2023.
  90. 90.Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao. Pick: processing key information extraction from documents using improved graph learning-convolutional networks. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 4363–4370. IEEE, 2021.
  91. 91.Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang, Zengyuan Guo, Xiameng Qin, Kun Yao, Junyu Han, Errui Ding, and Jingdong Wang. Structextv2: Masked visual-textual prediction for document image pre-training. In The Eleventh International Conference on Learning Representations, 2022.
  92. 92.Chong Zhang, Ya Guo, Yi Tu, Huan Chen, Jinyang Tang, Huijia Zhu, Qi Zhang, and Tao Gui. Reading order matters: Information extraction from visually-rich documents by token path prediction. arXiv preprint arXiv:2310.11016, 2023.
  93. 93.Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. Trie: end-to-end text reading and information extraction for document understanding. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1413–1422, 2020.
  94. 94.Xiang Zhang, Yongwen Su, Subarna Tripathi, and Zhuowen Tu. Text spotting transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9519–9528, 2022.
  95. 95.Xinyi Zheng, Douglas Burdick, Lucian Popa, Xu Zhong, and Nancy Xin Ru Wang. Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context. In WACV, pages 697–706, 2021.
  96. 96.Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. Publaynet: largest dataset ever for document layout analysis. In ICDAR, pages 1015–1022. IEEE, 2019.
  97. 97.Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes. Image-based table recognition: data, model, and evaluation. In ECCV, pages 564–580. Springer, 2020.
  98. 98.Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang. East: an efficient and accurate scene text detector. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 5551–5560, 2017.

Citation

MLA
Wan, J., et al. “OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition”. arXiv, 2024, http://arxiv.org/abs/2403.19128v1.
APA
Wan, J., Song, S., Yu, W., Liu, Y., Cheng, W., Huang, F., Bai, X., Yao, C., & Yang, Z. (2024). OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition. arXiv. http://arxiv.org/abs/2403.19128v1
Chicago
Wan, J., S. Song, W. Yu, et al. 2024. “OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition”. arXiv. http://arxiv.org/abs/2403.19128v1.
Harvard
Wan, J. et al. (2024) “OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.19128v1.
Vancouver
1. Wan J, Song S, Yu W, Liu Y, Cheng W, Huang F, Bai X, Yao C, Yang Z (2024) OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition. arXiv

BibTeX

@article{wan2024omniparser,
  title = {OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition},
  author = {Wan, Jianqiang and Song, Sibo and Yu, Wenwen and Liu, Yuliang and Cheng, Wenqing and Huang, Fei and Bai, Xiang and Yao, Cong and Yang, Zhibo},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.19128v1},
  eprint = {2403.19128}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE