Built independently by an author, for readers. Read the story and support ChapterPal

keyword

sequence labeling

Sequence labeling is a machine learning and natural language processing task that involves assigning a categorical tag or label to each individual item in an input sequence, such as words or characters within a text. Unlike whole-sequence classification, sequence labeling generates an output sequence corresponding to the input tokens by evaluating both the features of individual elements and the contextual dependencies between neighboring elements. This approach is widely used for structured prediction tasks such as named entity recognition, part-of-speech tagging, and syntactic chunking, typically implemented using probabilistic frameworks like hidden Markov models and conditional random fields or deep learning architectures such as recurrent neural networks and transformer-based models.

13 items

Do Transformers Parse while Predicting the Masked Word?

Do Transformers Parse while Predicting the Masked Word?

Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev Arora

OrganizationsDuke UniversityPrinceton University

Why you should read this

Proves that realistic-sized masked language models naturally approximate the Inside-Outside parsing algorithm on context-free grammar data, offering a theoretical and empirical explanation for how transformers implicitly learn syntactic structure during unsupervised pre-training.

Pre-trained language models have been shown to encode linguistic structures like parse trees in their embeddings while being trained unsupervised. Some doubts have been raised whether the models are doing parsing or only some computation weakly correlated with it. Concretely: (a) Is it possible to explicitly describe transformers with realistic embedding dimensions, number of heads, etc. that are capable of doing parsing—or even approximate parsing? (b) Why do pre-trained models capture parsing structure? This paper takes a step toward answering these questions in the context of generative modeling with PCFGs. We show that masked language models like BERT or RoBERTa of moderate sizes can approximately execute the Inside-Outside algorithm for the English PCFG. We also show that the Inside-Outside algorithm is optimal for masked language modeling loss on the PCFG-generated data. We conduct probing experiments on models pre-trained on PCFG-generated data to show that this not only allows recovery of approximate parse tree, but also recovers marginal span probabilities computed by the Inside-Outside algorithm, which suggests an implicit bias of masked language modeling towards this algorithm.

Added

2026-10-05

PromptNER: Prompt Locating and Typing for Named Entity Recognition

PromptNER: Prompt Locating and Typing for Named Entity Recognition

Yongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang, Rongsheng Zhang, Yadong Xi, Weiming Lu, Yueting Zhuang

OrganizationsMonash University

Why you should read this

Proposes a single-round prompt learning framework for named entity recognition that unifies entity locating and typing via dual-slot templates and bipartite graph matching, avoiding span enumeration while boosting cross-domain few-shot performance by 7.7%.

Prompt learning is a new paradigm for utilizing pre-trained language models and has achieved great success in many tasks. To adopt prompt learning in the NER task, two kinds of methods have been explored from a pair of symmetric perspectives, populating the template by enumerating spans to predict their entity types or constructing type-specific prompts to locate entities. However, these methods not only require a multi-round prompting manner with a high time overhead and computational cost, but also require elaborate prompt templates, that are difficult to apply in practical scenarios. In this paper, we unify entity locating and entity typing into prompt learning, and design a dual-slot multi-prompt template with the position slot and type slot to prompt locating and typing respectively. Multiple prompts can be input to the model simultaneously, and then the model extracts all entities by parallel predictions on the slots. To assign labels for the slots during training, we design a dynamic template filling mechanism that uses the extended bipartite graph matching between prompts and the ground-truth entities. We conduct experiments in various settings, including resource-rich flat and nested NER datasets and low-resource in-domain and cross-domain datasets. Experimental results show that the proposed model achieves a significant performance improvement, especially in the cross-domain few-shot setting, which outperforms the state-of-the-art model by +7.7% on average¹.

Added

2026-10-03

Large language models are few-shot clinical information extractors

Large language models are few-shot clinical information extractors

Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, David A. Sontag

OrganizationsMassachusetts Institute of TechnologyUniversity of Münster

Why you should read this

Demonstrates that general-domain large language models can perform complex few-shot clinical information extraction tasks, such as span identification and relation extraction, significantly outperforming existing baselines without domain-specific training.

A long-running goal of the clinical NLP community is the extraction of important variables trapped in clinical notes. However, roadblocks have included dataset shift from the general domain and a lack of public clinical corpora and annotations. In this work, we show that large language models, such as InstructGPT, perform well at zero- and few-shot information extraction from clinical text despite not being trained specifically for the clinical domain. Whereas text classification and generation performance have already been studied extensively in such models, here we additionally demonstrate how to leverage them to tackle a diverse set of NLP tasks which require more structured outputs, including span identification, token-level sequence classification, and relation extraction. Further, due to the dearth of available data to evaluate these systems, we introduce new datasets for benchmarking few-shot clinical information extraction based on a manual re-annotation of the CASI dataset for new tasks. On the clinical extraction tasks we studied, the GPT-3 systems significantly outperform existing zero- and few-shot baselines.

Added

2026-09-28

Contextual String Embeddings for Sequence Labeling

Contextual String Embeddings for Sequence Labeling

A. Akbik, Duncan A. J. Blythe, Roland Vollgraf

OrganizationsZalando

Why you should read this

Proposes extracting contextualized word representations from bidirectional character-level language models to effectively handle polysemy and subword structures, achieving state-of-the-art results in sequence labeling tasks like named entity recognition.

Recent advances in language modeling using recurrent neural networks have made it viable to model language as distributions over characters. By learning to predict the next character on the basis of previous characters, such models have been shown to automatically internalize linguistic concepts such as words, sentences, subclauses and even sentiment. In this paper, we propose to leverage the internal states of a trained character language model to produce a novel type of word embedding which we refer to as contextual string embeddings. Our proposed embeddings have the distinct properties that they (a) are trained without any explicit notion of words and thus fundamentally model words as sequences of characters, and (b) are contextualized by their surrounding text, meaning that the same word will have different embeddings depending on its contextual use. We conduct a comparative evaluation against previous embeddings and find that our embeddings are highly useful for downstream tasks: across four classic sequence labeling tasks we consistently outperform the previous state-of-the-art. In particular, we significantly outperform previous work on English and German named entity recognition (NER), allowing us to report new state-of-the-art F1-scores on the CoNLL03 shared task. We release all code and pre-trained language models in a simple-to-use framework to the re-search community, to enable reproduction of these experiments and application of our proposed embeddings to other tasks: https://github.com/zalandoresearch/flair

Added

2026-09-27

Creative Commons License
Grounded Multimodal Named Entity Recognition on Social Media

Grounded Multimodal Named Entity Recognition on Social Media

Jianfei Yu, Ziyan Li, Jieming Wang, Rui Xia

OrganizationsNanjing University of Science and Technology

Why you should read this

Introduces the task of Grounded Multimodal Named Entity Recognition along with a benchmark Twitter dataset and a hierarchical index generation framework that jointly extracts text entities and locates their corresponding image regions to resolve visual ambiguity in social media posts.

In recent years, Multimodal Named Entity Recognition (MNER) on social media has attracted considerable attention. However, existing MNER studies only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction and insufficient for entity disambiguation. To solve these issues, in this work, we introduce a Grounded Multimodal Named Entity Recognition (GMNER) task. Given a text-image social post, GMNER aims to identify the named entities in text, their entity types, and their bounding box groundings in image (i.e., visual regions). To tackle the GMNER task, we construct a Twitter dataset based on two existing MNER datasets. Moreover, we extend four well-known MNER methods to establish a number of baseline systems and further propose a Hierarchical Index generation framework named H-Index, which generates the entity-type-region triples in a hierarchical manner with a sequence-to-sequence model. Experiment results on our annotated dataset demonstrate the superiority of our H-Index framework over baseline systems on the GMNER task. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/GMNER.

Added

2026-09-26

CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning

CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning

Sarkar Snigdha Sarathi Das, Arzoo Katiyar, Rebecca J. Passonneau, Rui Zhang

OrganizationsPennsylvania State University

Why you should read this

Proposes a contrastive learning framework that models token representations as Gaussian distributions to optimize distributional divergence between entity types, achieving substantial gains over existing few-shot named entity recognition methods across multiple benchmarks.

Named Entity Recognition (NER) in Few-Shot setting is imperative for entity tagging in low resource domains. Existing approaches only learn class-specific semantic features and intermediate representations from source domains. This affects generalizability to unseen target domains, resulting in suboptimal performances. To this end, we present CONTaiNER, a novel contrastive learning technique that optimizes the inter-token distribution distance for Few-Shot NER. Instead of optimizing class-specific attributes, CONTaiNER optimizes a generalized objective of differentiating between token categories based on their Gaussian-distributed embeddings. This effectively alleviates overfitting issues originating from training domains. Our experiments in several traditional test domains (OntoNotes, CoNLL’03, WNUT ’17, GUM) and a new large scale Few-Shot NER dataset (Few-NERD) demonstrate that, on average, CONTaiNER outperforms previous methods by 3%-13% absolute F1 points while showing consistent performance trends, even in challenging scenarios where previous approaches could not achieve appreciable performance. The source code of CONTaiNER will be available at: https://github.com/psunlpgroup/CONTaiNER.

Added

2026-09-26

Nested Named Entity Recognition with Span-level Graphs

Nested Named Entity Recognition with Span-level Graphs

Juncheng Wan, Dongyu Ru, Weinan Zhang, Yong Yu

OrganizationsShanghai Jiao Tong University

Why you should read this

Proposes a retrieval-based span-level graph approach that connects candidate spans to training entities via n-gram similarity and applies graph convolutional networks to significantly boost nested named entity recognition performance, especially for low-frequency and long spans.

Span-based methods with the neural networks backbone have great potential for the nested named entity recognition (NER) problem. However, they face problems such as degenerating when positive instances and negative instances largely overlap. Besides, the generalization ability matters a lot in nested NER, as a large proportion of entities in the test set hardly appear in the training set. In this work, we try to improve the span representation by utilizing retrieval-based span-level graphs, connecting spans and entities in the training data based on n-gram features. Specifically, we build the entity-entity graph and span-entity graph globally based on n-gram similarity to integrate the information of similar neighbor entities into the span representation. To evaluate our method, we conduct experiments on three common nested NER datasets, ACE2004, ACE2005, and GENIA datasets. Experimental results show that our method achieves general improvements on all three benchmarks (+0.30 ∼ 0.85 micro-F1), and obtains special superiority on low frequency entities (+0.56 ∼ 2.08 recall).

Added

2026-09-26

Improving Language Understanding by Generative Pre-Training

Improving Language Understanding by Generative Pre-Training

Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever

OrganizationsOpenAI

Why you should read this

Proposes a semi-supervised training framework that combines unsupervised generative pre-training of a Transformer language model with task-specific discriminative fine-tuning, achieving state-of-the-art results across a diverse range of natural language understanding benchmarks.

Natural language understanding comprises a wide range of diverse tasks such as textual entailment, question answering, semantic similarity assessment, and document classification. Although large unlabeled text corpora are abundant, labeled data for learning these specific tasks is scarce, making it challenging for discriminatively trained models to perform adequately. We demonstrate that large gains on these tasks can be realized by generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task. In contrast to previous approaches, we make use of task-aware input transformations during fine-tuning to achieve effective transfer while requiring minimal changes to the model architecture. We demonstrate the effectiveness of our approach on a wide range of benchmarks for natural language understanding. Our general task-agnostic model outperforms discriminatively trained models that use architectures specifically crafted for each task, significantly improving upon the state of the art in 9 out of the 12 tasks studied. For instance, we achieve absolute improvements of 8.9% on commonsense reasoning (Stories Cloze Test), 5.7% on question answering (RACE), and 1.5% on textual entailment (MultiNLI).

Added

2026-06-11

License

Published with permission