De-Bias for Generative Extraction in Unified NER Task

Shuai ZhangYongliang ShenZeqi TanYiquan WuWeiming Lu

article2022ACL59 citations

Presents causal deconfounding data augmentation methods based on backdoor adjustment to eliminate pre-context and entity-order biases in generative named entity recognition across flat, nested, and discontinuous settings.

Listen

Extracting named entities—such as names, locations, and medical terms—from unstructured text is a vital capability for automated text analytics and downstream language systems. In practical applications, entities can appear in straightforward, nested, or broken apart formats across a sentence. While modern generative language models offer the unique advantage of handling all three entity structures within a single unified framework, their step-by-step text generation introduces unintended statistical biases. The article addresses this operational challenge by investigating and mitigating the spurious dependencies that degrade entity recognition performance.

The main objective of the article is to demonstrate how applying causal inference principles to training data augmentation can remove these unwanted generation biases and improve entity extraction accuracy across diverse text formats. To accomplish this, the authors identify two primary sources of bias: pre-context dependencies, where words preceding an entity mislead the model, and entity-order dependencies, where arbitrarily fixing the sequence of output entities prevents the model from learning bidirectional relationships.

To resolve these biases without modifying underlying neural network architectures, the authors developed two targeted data augmentation techniques based on causal backdoor adjustment: intra-entity deconfounding and inter-entity deconfounding. They evaluated their approach using the T5 language model architecture across eight standard benchmark datasets covering flat, nested, and discontinuous entity recognition tasks, comparing performance against leading generative and task-specific baseline systems.

The findings show consistent performance gains across all evaluated scenarios. Implementing intra-entity and inter-entity deconfounding yielded improved accuracy and recall across all eight datasets, matching or exceeding specialized non-generative models. Furthermore, targeted stress tests confirmed that the debiased models maintained significantly higher robustness when exposed to noisy prefixes and randomized entity orderings, showing up to a 2.60% improvement in precision retention under adversarial testing conditions.

These results demonstrate that organizations can successfully deploy a single, unified generative model for complex information extraction without sacrificing accuracy or maintaining fragmented, task-specific pipelines. Reducing architectural complexity lowers long-term system maintenance costs while improving reliability on varied document types. Organizations building information extraction systems should adopt these data augmentation strategies during model training to strengthen extraction quality, while exploring causal data debiasing across other generation tasks.

Confidence in these findings is supported by solid improvements across multiple established benchmarks. However, decision-makers should note that the data augmentation rules were curated using heuristic selection criteria based on entity length and occurrence frequency rather than applied uniformly to all samples. Further validation in specialized industry domains is recommended before full-scale deployment.

Zhang et al (2022).pdf
Cover for De-Bias for Generative Extraction in Unified NER Task

Abstract

Named entity recognition (NER) is a fundamental task to recognize specific types of entities from a given sentence. Depending on how the entities appear in the sentence, it can be divided into three subtasks, namely, Flat NER, Nested NER, and Discontinuous NER. Among the existing approaches, only the generative model can be uniformly adapted to these three subtasks. However, when the generative model is applied to NER, its optimization objective is not consistent with the task, which makes the model vulnerable to the incorrect biases. In this paper, we analyze the incorrect biases in the generation process from a causality perspective and attribute them to two confounders: pre-context confounder and entity-order confounder. Furthermore, we design Intra- and Inter-entity Deconfounding Data Augmentation methods to eliminate the above confounders according to the theory of back-door adjustment. Experiments show that our method can improve the performance of the generative NER model in various datasets.

Table of Contents

  • 1 Introduction
  • 2 Prerequisite
  • 2.1 Problem Definition
  • 2.2 Generative Model
  • 3 The Proposed Solution
  • 3.1 Intra-entity Deconfounding DA
  • 3.2 Inter-entity Deconfounding DA
  • 3.3 Constrained Prediction
  • 4 Experiments
  • 4.1 Datasets
  • 4.1.1 Flat NER Datasets
  • 4.1.2 Nested NER Datasets
  • 4.1.3 Discontinuous NER Datasets
  • 4.2 Implementation Details
  • 4.3 Results
  • 4.3.1 Comparision between Baselines
  • 4.3.2 Analysis of Intra-entity Deconfounding
  • 4.3.3 Analysis of Inter-entity Deconfounding
  • 4.4 Robustness Testing
  • 5 Related Work
  • 5.1 NER Task
  • 5.2 Causal Inference
  • 6 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Causal Structural Model and Confounders in Generative Named Entity Recognition

    model/method

    In autoregressive generative named entity recognition (NER), the extraction of entities from an input sentence XX into a target sequence YY is formulated via a Structural Causal Model (SCM) where the true direct path is X→YX \to Y. However, standard autoregressive modeling introduces a confounder variable NN, establishing a spurious backdoor correlation path X←N→YX \leftarrow N \to Y that introduces bias into generation. Two distinct types of confounders arise based on generation granularity:

    1. Pre-context Confounder (Intra-entity Generation): When generating tokens inside an entity, the autoregressive decoder conditions generation on previously emitted words (prefix context). If these prefix words include unrelated context words or tokens from other entities, the model learns spurious dependencies between non-associated tokens and intra-entity words, failing to capture true intra-entity dependencies when context varies.

    2. Entity-order Confounder (Inter-entity Generation): Entities in a sentence naturally constitute an unordered set. Imposing a fixed sequential decoding order (e.g., left-to-right appearance) introduces an artificial entity-order confounder. The model only learns unidirectional transitions between sequential entities and ignores reverse dependencies, rendering the extraction of remaining entities fragile if the generation order is perturbed.

  2. Knowl 2 — Backdoor Adjustment Formulation for Generative Extraction

    equation

    To eliminate the spurious correlation X←N→YX \leftarrow N \to Y introduced by the confounder NN (representing either pre-context words or entity decoding order) in generative named entity recognition, causal intervention is applied via Pearl's backdoor adjustment. The debiased conditional probability distribution P(Y∣do(X))P(Y \mid do(X)) is formulated as:

    P(Y∣do(X))=∑nP(Y∣X,n)P(n)P(Y \mid do(X)) = \sum_{n} P(Y \mid X, n) P(n)

    where:

    • XX is the input sentence.
    • YY is the target entity output sequence.
    • nn represents a specific stratum (realization) of the confounder variable NN.
    • P(n)P(n) is the prior probability of stratum nn, which is independent of the input XX.
    • P(Y∣X,n)P(Y \mid X, n) denotes the conditional probability of predicting YY given XX under the confounder stratum nn.

    Stratifying the confounder space enforces the model to maximize likelihood across diverse contexts/orders uniformly, removing the observational bias.

  3. Knowl 3 — Intra-Entity Deconfounding Data Augmentation

    model/method

    Intra-entity Deconfounding Data Augmentation implements backdoor adjustment for the pre-context confounder by isolating individual entities and pairing them with diverse sentence contexts during training.

    Given an input sentence XX containing MM target entities, MM separate training instances (X,Y′)(X, Y') are constructed. For each entity ei=(y1ei,y2ei,…,yEei)e_i = (y^{e_i}_1, y^{e_i}_2, \dots, y^{e_i}_{\mathcal{E}}) of length E\mathcal{E}, an augmented target sequence Y′Y' is created by randomly sampling a context word [CW][CW] from XX and prefixing it to the entity sequence:

    Y′={[CW],y1ei,y2ei,…,yEei}Y' = \{[CW], y^{e_i}_1, y^{e_i}_2, \dots, y^{e_i}_{\mathcal{E}}\}

    Unlike standard target sequences, Y′Y' explicitly omits sequence-level start [ss][ss] and end [ee][ee] tags. This omission instructs the generative model that the augmented sequence represents a single entity rather than the full set of sentence entities, preventing early sequence termination during inference.

  4. Knowl 4 — Inter-Entity Deconfounding Data Augmentation

    model/method

    Inter-entity Deconfounding Data Augmentation eliminates the entity-order confounder by training the model across varied permutations of entity sequence prefixes.

    For an input sentence XX with target entities E1,E2,…,EME_1, E_2, \dots, E_M (where each entity Ei={[s],y1ei,…,yEei,[e]}E_i = \{[s], y^{e_i}_1, \dots, y^{e_i}_{\mathcal{E}}, [e]\} includes entity boundary tags), augmented target sequences Y′Y' are constructed by holding the final entity EME_M fixed while permuting the preceding entities:

    Y′={[ss],Perm(E1,…,EM−1),EM,[ee]}Y' = \{[ss], \text{Perm}(E_1, \dots, E_{M-1}), E_M, [ee]\}

    where Perm(⋅)\text{Perm}(\cdot) denotes a permutation operator over the first M−1M-1 entities, and [ss][ss] and [ee][ee] denote sequence start and end tokens.

    During training on these augmented sequences, loss is computed exclusively on the first token of the target entity EME_M, while the permuted prefix Perm(E1,…,EM−1)\text{Perm}(E_1, \dots, E_{M-1}) is provided as teacher-forced context. This trains the model to decode any target entity regardless of which subset of entities has already been generated.

  5. Knowl 5 — Constrained Decoding Schema for Generative NER

    model/method

    To prevent exposure bias and restrict the generation space to valid entities, generation is regulated by token-level constraints and special sentinel tokens:

    • Vocabulary Constraints: The decoder vocabulary at generation time is restricted strictly to tokens appearing in the input sentence XX along with predefined special tokens (e.g., T5 sentinel tokens ⟨extra_id_0⟩\langle extra\_id\_0 \rangle through ⟨extra_id_50⟩\langle extra\_id\_50 \rangle).
    • Structural Transition Rules:
      1. Sequence generation must begin with the sequence start tag [ss][ss].
      2. Each entity must begin with the entity start tag [s][s].
      3. After [s][s], only tokens from XX forming the entity text can be generated.
      4. The entity category tag must immediately follow the entity text tokens.
      5. The entity category tag must be directly followed by the entity end tag [e][e].
      6. Following [e][e], valid subsequent tokens are restricted to either another entity start tag [s][s] or the sequence end tag [ee][ee].
  6. Knowl 6 — Performance Comparison Across Flat, Nested, and Discontinuous NER Benchmarks

    data/table

    A T5-Base generative NER model trained with Intra-entity Deconfounding (Intra-De) and Inter-entity Deconfounding (Inter-De) was evaluated on eight benchmark datasets across Flat NER (CoNLL2003, OntoNotes), Nested NER (ACE2004, ACE2005, GENIA), and Discontinuous NER (CADEC, ShARe13, ShARe14). Across all benchmarks, both deconfounding strategies yield consistent improvements in precision (P), recall (R), and F1-score over the non-deconfounded T5-Base baseline (Without-De):

    Dataset Without-De Intra-De Inter-De
    P R F1 P R F1 P R F1
    CoNLL2003 92.68 93.49 93.08 92.78 93.51 93.14 92.68 93.57 93.12
    OntoNotes 89.58 90.71 90.14 89.77 91.07 90.42 89.75 91.02 90.38
    ACE2004 86.19 83.76 84.96 86.36 84.54 85.44 86.53 84.06 85.28
    ACE2005 83.23 86.25 84.71 83.31 86.56 84.90 82.92 87.05 84.93
    Genia 80.11 76.92 78.49 81.04 77.21 79.08 80.66 76.45 78.50
    CADEC 71.34 70.54 70.94 71.35 71.86 71.60 70.44 71.65 71.04
    ShARe13 79.03 78.03 78.53 81.09 78.13 79.58 81.31 76.75 78.96
    ShARe14 77.06 83.41 80.11 77.88 83.77 80.72 77.51 83.27 80.29
  7. Knowl 7 — Robustness Testing Under Pre-context and Entity-order Attack Scenarios

    empirical result

    The efficacy of the deconfounding methods was tested under synthetic attack perturbations during decoding on CoNLL2003, ACE2004, and CADEC:

    1. Pre-context Attack: Random prefix words were injected prior to entity decoding.

      • Baseline (Without-De) performance drops by ΔF1=−0.69%\Delta \text{F1} = -0.69\% on CoNLL, −3.21%-3.21\% on ACE04, and −2.27%-2.27\% on CADEC.
      • With Intra-entity Deconfounding (Intra-De), degradation drops to ΔF1=−0.26%\Delta \text{F1} = -0.26\%, −1.77%-1.77\%, and −1.06%-1.06\%, representing relative robustness gains of +0.43%+0.43\%, +1.44%+1.44\%, and +1.21%+1.21\% F1.
    2. Entity-order Attack: k=4k=4 randomly sampled gold entities were forced as the generation prefix, evaluating performance on the remaining entities.

      • Baseline performance drops by ΔF1=−1.59%\Delta \text{F1} = -1.59\% on CoNLL, −1.61%-1.61\% on ACE04, and −1.63%-1.63\% on CADEC.
      • With Inter-entity Deconfounding (Inter-De), performance degradation decreases to ΔF1=−1.40%\Delta \text{F1} = -1.40\%, −1.12%-1.12\%, and −0.92%-0.92\%, showing relative robustness gains of +0.19%+0.19\%, +0.49%+0.49\%, and +0.71%+0.71\% F1.

Coverage note — None was omitted; all key contributions including causal analysis, mathematical formulation of backdoor adjustment, intra- and inter-entity deconfounding augmentations, constrained decoding, and empirical results across all 8 datasets and robustness tests are covered.

References

  1. 1.Alan Akbik, Tanja Bergmann, and Roland Vollgraf. 2019. Pooled contextualized embeddings for named entity recognition. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 724–728. Association for Computational Linguistics.
  2. 2.L. S. Alves, C. Susin, N Damé-Teixeira, and M. Maltz. 2014. Tooth loss prevalence and risk indicators among 12-year-old schoolchildren from south brazil. Caries Research, 48(4):347.
  3. 3.Jason P. C. Chiu and Eric Nichols. 2016. Named entity recognition with bidirectional lstm-cnns. Trans. Assoc. Comput. Linguistics, 4:357–370.
  4. 4.Kevin Clark, Minh-Thang Luong, Christopher D. Manning, and Quoc V. Le. 2018. Semi-supervised sequence modeling with cross-view training. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 1914–1925. Association for Computational Linguistics.
  5. 5.Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa. 2011. Natural language processing (almost) from scratch. J. Mach. Learn. Res., 12:2493–2537.
  6. 6.Xiang Dai, Sarvnaz Karimi, Ben Hachey, and Cécile Paris. 2020. An effective transition-based model for discontinuous NER. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 5860–5870. Association for Computational Linguistics.
  7. 7.George R. Doddington, Alexis Mitchell, Mark A. Przybocki, Lance A. Ramshaw, Stephanie M. Strassel, and Ralph M. Weischedel. 2004. The automatic content extraction (ACE) program - tasks, data, and evaluation. In Proceedings of the Fourth International Conference on Language Resources and Evaluation, LREC 2004, May 26-28, 2004, Lisbon, Portugal. European Language Resources Association.
  8. 8.Norman E. Fenton, Martin Neil, and Anthony C. Constantinou. 2020. The book of why: The new science of cause and effect, judea pearl, dana mackenzie. basic books (2018). Artif. Intell., 284:103286.
  9. 9.Octavian-Eugen Ganea and Thomas Hofmann. 2017. Deep joint entity disambiguation with local neural attention. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2619–2629, Copenhagen, Denmark. Association for Computational Linguistics.
  10. 10.Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020. Evaluating models’ local decision boundaries via contrast sets. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 1307–1323. Association for Computational Linguistics.
  11. 11.Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel. 2019. Counterfactual fairness in text classification through robustness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, AIES 2019, Honolulu, HI, USA, January 27-28, 2019, pages 219–226. ACM.
  12. 12.Zhiheng Huang, Wei Xu, and Kai Yu. 2015. Bidirectional LSTM-CRF models for sequence tagging. CoRR, abs/1508.01991.
  13. 13.Meizhi Ju, Makoto Miwa, and Sophia Ananiadou. 2018. A neural layered model for nested named entity recognition. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1446–1459, New Orleans, Louisiana. Association for Computational Linguistics.
  14. 14.Sarvnaz Karimi, Alejandro Metke-Jimenez, Madonna Kemp, and Chen Wang. 2015. Cadec: A corpus of adverse drug event annotations. J. Biomed. Informatics, 55:73–81.
  15. 15.Arzoo Katiyar and Claire Cardie. 2018. Nested named entity recognition revisited. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 861–871. Association for Computational Linguistics.
  16. 16.Jin-Dong Kim, Tomoko Ohta, Yuka Tateisi, and Jun’ichi Tsujii. 2003. GENIA corpus - a semantically annotated corpus for bio-textmining. In Proceedings of the Eleventh International Conference on Intelligent Systems for Molecular Biology, June 29 - July 3, 2003, Brisbane, Australia, pages 180–182.
  17. 17.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 260–270. The Association for Computational Linguistics.
  18. 18.Xiaonan Li, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2020a. FLAT: chinese NER using flat-lattice transformer. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6836–6842. Association for Computational Linguistics.
  19. 19.Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2020b. A unified MRC framework for named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 5849–5859. Association for Computational Linguistics.
  20. 20.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  21. 21.Wei Lu and Dan Roth. 2015. Joint mention extraction and classification with mention hypergraphs. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pages 857–867. The Association for Computational Linguistics.
  22. 22.Yi Luan, Dave Wadden, Luheng He, Amy Shah, Mari Ostendorf, and Hannaneh Hajishirzi. 2019a. A general framework for information extraction using dynamic span graphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3036–3046, Minneapolis, Minnesota. Association for Computational Linguistics.
  23. 23.Yi Luan, Dave Wadden, Luheng He, Amy Shah, Mari Ostendorf, and Hannaneh Hajishirzi. 2019b. A general framework for information extraction using dynamic span graphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 3036–3046. Association for Computational Linguistics.
  24. 24.K. Luke. 2015. The statistics of causal inference: A view from political methodology. Political Analysis, 23(3):313–335.
  25. 25.D. P. Mackinnon, A. J. Fairchild, and M. S. Fritz. 2007. Mediation analysis. Annual Review of Psychology, 58(1):593.
  26. 26.Andrew McCallum and Wei Li. 2003. Early results for named entity recognition with conditional random fields, feature induction and web-enhanced lexicons. In Proceedings of the Seventh Conference on Natural Language Learning, CoNLL 2003, Held in cooperation with HLT-NAACL 2003, Edmonton, Canada, May 31 - June 1, 2003, pages 188–191. ACL.
  27. 27.Alejandro Metke-Jimenez and Sarvnaz Karimi. 2016. Concept identification and normalisation for adverse drug event discovery in medical forums. In Proceedings of the First International Workshop on Biomedical Data Integration and Discovery (BMDID 2016) co-located with The 15th International Semantic Web Conference (ISWC 2016), Kobe, Japan, October 17, 2016, volume 1709 of CEUR Workshop Proceedings. CEUR-WS.org.
  28. 28.Makoto Miwa and Mohit Bansal. 2016. End-to-end relation extraction using LSTMs on sequences and tree structures. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1105–1116, Berlin, Germany. Association for Computational Linguistics.
  29. 29.Danielle L. Mowery, Sumithra Velupillai, Brett R. South, Lee M. Christensen, David Martínez, Liadh Kelly, Lorraine Goeuriot, Noémie Elhadad, Sameer Pradhan, Guergana K. Savova, and Wendy W. Chapman. 2014. Task 2: Share/clef ehealth evaluation lab 2014. In Working Notes for CLEF 2014 Conference, Sheffield, UK, September 15-18, 2014, volume 1180 of CEUR Workshop Proceedings, pages 31–42. CEUR-WS.org.
  30. 30.Aldrian Obaja Muis and Wei Lu. 2016. Learning to recognize discontiguous entities. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 75–84. The Association for Computational Linguistics.
  31. 31.Aldrian Obaja Muis and Wei Lu. 2017. Labeling gaps between words: Recognizing overlapping mentions with mention separators. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 2608–2618. Association for Computational Linguistics.
  32. 32.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In 9th International Conference on Learning Representations, ICLR 2021.
  33. 33.Judea Pearl, M. Maria Glymour, and Nicholas P Jewell. 2016. Causal inference in statistics: A primer.
  34. 34.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 2227–2237. Association for Computational Linguistics.
  35. 35.Sameer Pradhan, Noémie Elhadad, Brett R. South, David Martínez, Lee M. Christensen, Amy Vogel, Hanna Suominen, Wendy W. Chapman, and Guergana K. Savova. 2013a. Task 1: Share/clef ehealth evaluation lab 2013. In Working Notes for CLEF 2013 Conference , Valencia, Spain, September 23-26, 2013, volume 1179 of CEUR Workshop Proceedings. CEUR-WS.org.
  36. 36.Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013b. Towards robust linguistic analysis using ontonotes. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, CoNLL 2013, Sofia, Bulgaria, August 8-9, 2013, pages 143–152. ACL.
  37. 37.Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012. Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes. In Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning - Proceedings of the Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes, EMNLP-CoNLL 2012, July 13, 2012, Jeju Island, Korea, pages 1–40. ACL.
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  39. 39.Lev-Arie Ratinov and Dan Roth. 2009. Design challenges and misconceptions in named entity recognition. In Proceedings of the Thirteenth Conference on Computational Natural Language Learning, CoNLL 2009, Boulder, Colorado, USA, June 4-5, 2009, pages 147–155. ACL.
  40. 40.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning, CoNLL 2003, Held in cooperation with HLT-NAACL 2003, Edmonton, Canada, May 31 - June 1, 2003, pages 142–147. ACL.
  41. 41.Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang, Wen Wang, and Weiming Lu. 2021a. Locate and label: A two-stage identifier for nested named entity recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2782–2794, Online. Association for Computational Linguistics.
  42. 42.Yongliang Shen, Xinyin Ma, Yechun Tang, and Weiming Lu. 2021b. A trigger-sense memory flow framework for joint entity and relation extraction. In Proceedings of the Web Conference 2021, WWW ’21, page 1704–1715, New York, NY, USA. Association for Computing Machinery.
  43. 43.Takashi Shibuya and Eduard H. Hovy. 2020. Nested named entity recognition via second-best sequence learning and decoding. Trans. Assoc. Comput. Linguistics, 8:605–620.
  44. 44.Jana Straková, Milan Straka, and Jan Hajic. 2019. Neural architectures for nested NER through linearization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5326–5331, Florence, Italy. Association for Computational Linguistics.
  45. 45.Jana Straková, Milan Straka, and Jan Hajic. 2019. Neural architectures for nested NER through linearization. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 5326–5331. Association for Computational Linguistics.
  46. 46.Zeqi Tan, Yongliang Shen, Shuai Zhang, Weiming Lu, and Yueting Zhuang. 2021. A sequence-to-set network for nested named entity recognition. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, pages 3936–3942. ijcai.org.
  47. 47.Buzhou Tang, Jianglu Hu, Xiaolong Wang, and Qingcai Chen. 2018. Recognizing continuous and discontinuous adverse drug reaction mentions from social media using LSTM-CRF. Wirel. Commun. Mob. Comput., 2018.
  48. 48.Bailin Wang and Wei Lu. 2018. Neural segmental hypergraphs for overlapping mention recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 204–214. Association for Computational Linguistics.
  49. 49.Bailin Wang and Wei Lu. 2019. Combining spans into entities: A neural two-stage approach for recognizing discontiguous entities. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 6215–6223. Association for Computational Linguistics.
  50. 50.Jue Wang, Lidan Shou, Ke Chen, and Gang Chen. 2020a. Pyramid: A layered model for nested named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 5918–5928. Association for Computational Linguistics.
  51. 51.Yu Wang, Yun Li, Hanghang Tong, and Ziye Zhu. 2020b. HIT: nested named entity recognition via head-tail pair and token interaction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6027–6036. Association for Computational Linguistics.
  52. 52.Mingbin Xu, Hui Jiang, and Sedtawut Watcharawittayakul. 2017. A local detection approach for named entity recognition and mention detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1237–1247. Association for Computational Linguistics.
  53. 53.Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto. 2020. LUKE: deep contextualized entity representations with entity-aware self-attention. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6442–6454. Association for Computational Linguistics.
  54. 54.Hang Yan, Bocao Deng, Xiaonan Li, and Xipeng Qiu. 2019. TENER: adapting transformer encoder for named entity recognition. CoRR, abs/1911.04474.
  55. 55.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021a. A unified generative framework for various NER subtasks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5808–5822, Online. Association for Computational Linguistics.
  56. 56.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021b. A unified generative framework for various NER subtasks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5808–5822. Association for Computational Linguistics.
  57. 57.Juntao Yu, Bernd Bohnet, and Massimo Poesio. 2020. Named entity recognition as dependency parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6470–6476. Association for Computational Linguistics.

Citation

MLA
Zhang, S., et al. “De-Bias for Generative Extraction in Unified NER Task”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 808–18, https://doi.org/10.18653/v1/2022.acl-long.59.
APA
Zhang, S., Shen, Y., Tan, Z., Wu, Y., & Lu, W. (2022). De-Bias for Generative Extraction in Unified NER Task. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 808–818. https://doi.org/10.18653/v1/2022.acl-long.59
Chicago
Zhang, S., Y. Shen, Z. Tan, Y. Wu, and W. Lu. 2022. “De-Bias for Generative Extraction in Unified NER Task”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 808–18. https://doi.org/10.18653/v1/2022.acl-long.59.
Harvard
Zhang, S. et al. (2022) “De-Bias for Generative Extraction in Unified NER Task”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 808–818. Available at: https://doi.org/10.18653/v1/2022.acl-long.59.
Vancouver
1. Zhang S, Shen Y, Tan Z, Wu Y, Lu W (2022) De-Bias for Generative Extraction in Unified NER Task. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 808–818

BibTeX

@inproceedings{zhang-etal-2022-de,
    title = "De-Bias for Generative Extraction in Unified {NER} Task",
    author = "Zhang, Shuai  and
      Shen, Yongliang  and
      Tan, Zeqi  and
      Wu, Yiquan  and
      Lu, Weiming",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.59/",
    doi = "10.18653/v1/2022.acl-long.59",
    pages = "808--818"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/