FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
Chen-Yu LeeChun-Liang LiTimothy DozatVincent PerotGuolong SuNan HuaJoshua AinslieRenshen WangYasuhisa FujiiTomas Pfister
Proposes a structure-aware sequence model that combines graph-convolutional token representations with spatial attention to extract key information from visually rich form documents without relying on expensive image features or large pre-training budgets.
Organizations routinely rely on extracting critical information from structured forms such as invoices, receipts, and applications to automate business workflows. However, standard language processing methods convert documents into a single stream of text, which breaks the physical layout and frequently scatters related information across separate text segments. When this serialization fails, automated systems misread relationships among key entities, driving up extraction errors and requiring costly manual intervention. The article introduces and evaluates FormNet, a structure-aware neural network architecture designed to overcome these layout-induced serialization errors without relying on computationally expensive visual processing.
The authors tackle this challenge by combining graph-based representation learning with an enhanced transformer sequence model. Before the extracted text is flattened into a sequential order, a graph convolutional network creates rich "Super-Tokens" by connecting neighboring words based on their two-dimensional layout geometry, preserving local context that might otherwise be broken apart. The system then processes these representations through a long-sequence transformer enhanced with "Rich Attention," a mechanism that directly incorporates horizontal and vertical spatial distances to penalize illogical text connections. The authors evaluate this architecture on three standard benchmarks—CORD (receipts), FUNSD (noisy forms), and Payment (invoices)—using unsupervised pre-training on 700,000 unlabeled documents followed by task-specific fine-tuning.
The findings show that FormNet achieves state-of-the-art performance across all three benchmarks while operating with notable computational efficiency. It sets new performance records with accuracy scores (F1) of 97.28% on CORD, 84.69% on FUNSD, and 92.19% on Payment. Furthermore, FormNet outperforms leading multimodal models while using a 64% smaller model footprint and over seven times less pre-training data. Ablation experiments confirm that both the graph-based Super-Tokens and Rich Attention contribute substantial, complementary gains, improving baseline performance on receipts by 5.3 percentage points and proving that direct structural modeling is far more effective than simply enlarging standard transformer architectures.
These results demonstrate that explicitly modeling geometric relationships enables high-accuracy document parsing at substantially lower operational and compute costs. By eliminating the need to process raw document images alongside text, FormNet reduces hardware requirements and inference latency, making high-volume document extraction more scalable and reliable. Organizations adopting this approach can lower error rates in automated back-office processing and achieve faster processing cycles.
For practical implementation, technical teams should consider adopting graph-enhanced structural modeling as the standard pipeline for form-based information extraction. While confidence in these results is high across standardized form datasets, the current architecture focuses strictly on textual and layout coordinates rather than image pixels and is designed for entity extraction rather than arbitrary key-value pairing. Future efforts should conduct pilot evaluations on proprietary, highly noisy layouts and explore lightweight integrations of visual features for documents containing rich pictorial cues.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Longformer’s efficient local-and-global attention provides the long-sequence transformer foundation that FormNet adapts for document text.
- Paper: DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding, Dongsheng Wang et al. (2024). DocLLM carries layout-aware text modeling into generative document understanding, extending FormNet’s structural approach to a broader set of tasks.
