Template-free Prompt Tuning for Few-shot NER
Ruotian MaXin ZhouTao GuiYiding TanLinyang LiQi ZhangXuanjing Huang
Proposes an entity-oriented language modeling objective for few-shot named entity recognition that eliminates the need for prompt templates by directly predicting class-related label words at entity positions, outperforming standard fine-tuning while accelerating decoding speed by over 1,900 times compared to template-based methods.
Extracting key information from unstructured text through named entity recognition is critical for enterprise data processing. However, training effective models typically requires large volumes of manually annotated data, which is expensive and time-consuming to create. While prompt-based learning has enabled language models to learn from just a few examples in sentence-level classification, adapting this approach to entity recognition has been hindered by severe computational bottlenecks. Standard template-based methods must test every possible word span in a sentence individually, causing processing times to explode as document length increases.
The article evaluates a new framework called Entity-oriented Language Model fine-tuning, which eliminates templates entirely while preserving the data efficiency of prompt learning. The researchers formulate entity recognition as a direct language modeling task where the model is fine-tuned to predict representative class words at entity positions and the original words at non-entity positions. To validate this approach, the authors tested the method across three benchmark datasets spanning newswire, general text, and movie reviews under low-resource scenarios containing 5, 10, 20, and 50 examples per category.
The findings show that the proposed method delivers substantial improvements in both operational efficiency and accuracy compared to standard and prompt-based baselines. Most notably, the approach operates up to 1,930 times faster than template-based methods by processing entire sentences in a single pass, matching the speed of standard classifiers. In accuracy, the method outperformed standard fine-tuning by up to 11.8 percentage points in extreme low-data scenarios with only 5 examples per class. Furthermore, the analysis revealed that constructing virtual representative vectors by combining corpus frequency and language model predictions yielded the best adaptation performance, remaining stable even when using noisy or significantly reduced reference dictionaries.
These results demonstrate that organizations can deploy high-performing entity extraction models in low-resource environments without incurring heavy computational infrastructure costs or latency penalties during deployment. By eliminating the structural gap between pre-training and fine-tuning without introducing extra parameters, this framework significantly reduces the cost and timeline associated with manual data labeling.
Organizations operating in data-constrained domains should consider adopting template-free prompt tuning over traditional classifier heads or computationally heavy span-based prompt architectures. Teams can enhance deployment performance by combining this framework with a structured sequence decoder and conducting domain-specific pre-training on available unlabeled text. Further pilot evaluations are recommended to test the methodology across broader multilingual contexts and highly specialized technical vocabularies.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). Provides the foundational LM-BFF methodology for few-shot prompt fine-tuning and automated label word search that the source adapts directly to entity span classification.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). Introduces the Pattern-Exploiting Training framework that reformulates classification tasks as masked language modeling cloze queries, establishing the verbalizer paradigm simplified by the source.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Establishes a systematic survey and typology of prompt-based learning and verbalizer selection across natural language processing tasks.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Introduces continuous prompt tuning (P-Tuning) to alleviate the manual prompt engineering bottleneck in pretrained language models.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Analyzes automated prompt generation and mining strategies to probe language model representations without manual template crafting.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Demonstrates the power of continuous prompt tuning across model scales as a parameter-efficient alternative to full fine-tuning.
- Paper: A Survey on Deep Learning for Named Entity Recognition, Jing Li et al. (2018). Surveys traditional neural sequence tagging architectures for named entity recognition, providing the baseline formulations outperformed by prompt-tuning methods.
- Paper: Prompting Language Models for Linguistic Structure, Terra Blevins et al. (2023). Explores how autoregressive language models can extract structured sequence tags, including named entities, via sequential structured prompting without modifying model weights.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). Develops a generalist bidirectional transformer model that represents entity types as prompts matching text spans directly, overcoming traditional sequence-labeling and template enumeration bottlenecks.
- Paper: Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER, Dong-Ho Lee et al. (2022). Extends low-resource NER by appending dynamically retrieved demonstrations directly to input text, avoiding complex template querying over candidate spans.
- Paper: Empirical Study of Zero-Shot NER with ChatGPT, Tingyu Xie et al. (2023). Investigates zero-shot prompt-based entity extraction and structured task decomposition using modern large language models.
- Paper: GPT-RE: In-context Learning for Relation Extraction using Large Language Models, Zhen Wan et al. (2023). Applies prompt demonstration and label-induced reasoning logic to the adjacent structured extraction problem of relation extraction in low-data regimes.
