De-Bias for Generative Extraction in Unified NER Task
Shuai ZhangYongliang ShenZeqi TanYiquan WuWeiming Lu
Presents causal deconfounding data augmentation methods based on backdoor adjustment to eliminate pre-context and entity-order biases in generative named entity recognition across flat, nested, and discontinuous settings.
Extracting named entities—such as names, locations, and medical terms—from unstructured text is a vital capability for automated text analytics and downstream language systems. In practical applications, entities can appear in straightforward, nested, or broken apart formats across a sentence. While modern generative language models offer the unique advantage of handling all three entity structures within a single unified framework, their step-by-step text generation introduces unintended statistical biases. The article addresses this operational challenge by investigating and mitigating the spurious dependencies that degrade entity recognition performance.
The main objective of the article is to demonstrate how applying causal inference principles to training data augmentation can remove these unwanted generation biases and improve entity extraction accuracy across diverse text formats. To accomplish this, the authors identify two primary sources of bias: pre-context dependencies, where words preceding an entity mislead the model, and entity-order dependencies, where arbitrarily fixing the sequence of output entities prevents the model from learning bidirectional relationships.
To resolve these biases without modifying underlying neural network architectures, the authors developed two targeted data augmentation techniques based on causal backdoor adjustment: intra-entity deconfounding and inter-entity deconfounding. They evaluated their approach using the T5 language model architecture across eight standard benchmark datasets covering flat, nested, and discontinuous entity recognition tasks, comparing performance against leading generative and task-specific baseline systems.
The findings show consistent performance gains across all evaluated scenarios. Implementing intra-entity and inter-entity deconfounding yielded improved accuracy and recall across all eight datasets, matching or exceeding specialized non-generative models. Furthermore, targeted stress tests confirmed that the debiased models maintained significantly higher robustness when exposed to noisy prefixes and randomized entity orderings, showing up to a 2.60% improvement in precision retention under adversarial testing conditions.
These results demonstrate that organizations can successfully deploy a single, unified generative model for complex information extraction without sacrificing accuracy or maintaining fragmented, task-specific pipelines. Reducing architectural complexity lowers long-term system maintenance costs while improving reliability on varied document types. Organizations building information extraction systems should adopt these data augmentation strategies during model training to strengthen extraction quality, while exploring causal data debiasing across other generation tasks.
Confidence in these findings is supported by solid improvements across multiple established benchmarks. However, decision-makers should note that the data augmentation rules were curated using heuristic selection criteria based on entity length and occurrence frequency rather than applied uniformly to all samples. Further validation in specialized industry domains is recommended before full-scale deployment.
- Paper: A Survey on Deep Learning for Named Entity Recognition, Jing Li et al. (2018). This comprehensive survey outlines the foundational architectures, decoding paradigms, and benchmark formulations essential for understanding neural named entity recognition.
- Paper: Pre-trained models for natural language processing: A survey, Xipeng Qiu et al. (2020). Reading this paper provides essential background on pre-trained contextual transformer models like T5 and how they adapt to downstream structured language tasks.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This work establishes the theoretical and practical underpinnings of sequential generation errors, exposure bias, and statistical dependencies in autoregressive sequence modeling.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This study demonstrates how deep learning models exploit superficial heuristics and spurious statistical shortcuts in training data rather than true task mechanics.
- Paper: Neural Architectures for Named Entity Recognition, Guillaume Lample et al. (2016). This foundational paper establishes standard neural representations and sequence labeling paradigms for named entity recognition across varied text corpora.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). This review extends the generative extraction framework to structured knowledge graph construction, analyzing target linearization formats and sequence-to-sequence architectures.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). This work introduces an alternative generalist NER framework that matches entity type representations against text spans to extract arbitrary entity structures without autoregressive generation bias.
- Paper: Text-to-Table: A New Way of Information Extraction, Xueqing Wu et al. (2022). This paper expands generative sequence-to-sequence information extraction beyond entity spans to direct end-to-end relational table synthesis.
- Paper: Empirical Study of Zero-Shot NER with ChatGPT, Tingyu Xie et al. (2023). This study evaluates zero-shot entity extraction capabilities and structured multi-turn prompting strategies in modern autoregressive large language models.
- Paper: Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction, Bowen Zhang et al. (2024). This paper develops a multi-phase generative extraction pipeline with Large Language Models to build generalized knowledge graphs without fine-tuning.
