Template-free Prompt Tuning for Few-shot NER

Ruotian MaXin ZhouTao GuiYiding TanLinyang LiQi ZhangXuanjing Huang

article2022NAACL210 citations

Proposes an entity-oriented language modeling objective for few-shot named entity recognition that eliminates the need for prompt templates by directly predicting class-related label words at entity positions, outperforming standard fine-tuning while accelerating decoding speed by over 1,900 times compared to template-based methods.

Listen

Extracting key information from unstructured text through named entity recognition is critical for enterprise data processing. However, training effective models typically requires large volumes of manually annotated data, which is expensive and time-consuming to create. While prompt-based learning has enabled language models to learn from just a few examples in sentence-level classification, adapting this approach to entity recognition has been hindered by severe computational bottlenecks. Standard template-based methods must test every possible word span in a sentence individually, causing processing times to explode as document length increases.

The article evaluates a new framework called Entity-oriented Language Model fine-tuning, which eliminates templates entirely while preserving the data efficiency of prompt learning. The researchers formulate entity recognition as a direct language modeling task where the model is fine-tuned to predict representative class words at entity positions and the original words at non-entity positions. To validate this approach, the authors tested the method across three benchmark datasets spanning newswire, general text, and movie reviews under low-resource scenarios containing 5, 10, 20, and 50 examples per category.

The findings show that the proposed method delivers substantial improvements in both operational efficiency and accuracy compared to standard and prompt-based baselines. Most notably, the approach operates up to 1,930 times faster than template-based methods by processing entire sentences in a single pass, matching the speed of standard classifiers. In accuracy, the method outperformed standard fine-tuning by up to 11.8 percentage points in extreme low-data scenarios with only 5 examples per class. Furthermore, the analysis revealed that constructing virtual representative vectors by combining corpus frequency and language model predictions yielded the best adaptation performance, remaining stable even when using noisy or significantly reduced reference dictionaries.

These results demonstrate that organizations can deploy high-performing entity extraction models in low-resource environments without incurring heavy computational infrastructure costs or latency penalties during deployment. By eliminating the structural gap between pre-training and fine-tuning without introducing extra parameters, this framework significantly reduces the cost and timeline associated with manual data labeling.

Organizations operating in data-constrained domains should consider adopting template-free prompt tuning over traditional classifier heads or computationally heavy span-based prompt architectures. Teams can enhance deployment performance by combining this framework with a structured sequence decoder and conducting domain-specific pre-training on available unlabeled text. Further pilot evaluations are recommended to test the methodology across broader multilingual contexts and highly specialized technical vocabularies.

Cover for Template-free Prompt Tuning for Few-shot NER

Abstract

Prompt-based methods have been successfully applied in sentence-level few-shot learning tasks, mostly owing to the sophisticated design of templates and label words. However, when applied to token-level labeling tasks such as NER, it would be time-consuming to enumerate the template queries over all potential entity spans. In this work, we propose a more elegant method to reformulate NER tasks as LM problems without any templates. Specifically, we discard the template construction process while maintaining the word prediction paradigm of pre-training models to predict a class-related pivot word (or label word) at the entity position. Meanwhile, we also explore principled ways to automatically search for appropriate label words that the pre-trained models can easily adapt to. While avoiding the complicated template-based process, the proposed LM objective also reduces the gap between different objectives used in pre-training and fine-tuning, thus it can better benefit the few-shot performance. Experimental results demonstrate the effectiveness of the proposed method over bert-tagger and template-based method under few-shot settings. Moreover, the decoding speed of the proposed method is up to 1930.12 times faster than the template-based method.

Table of Contents

  • 1 Introduction
  • 2 Problem Setup
  • 3 Approach
  • 3.1 Template-based Prompt Tuning
  • 3.1.1 Problems of Prompt-based NER
  • 3.2 Entity-Oriented LM Fine-tuning
  • 3.3 Label Word Engineering
  • 3.3.1 Low-resource Label word selection
  • 3.3.2 Label word searching
  • 3.3.3 Removing conflict label words
  • 4 Experiments
  • 4.1 Experimental settings
  • 4.2 Datasets and Implementation Details
  • 4.3 Baselines and Proposed Models
  • 4.4 Few-shot Results
  • 4.5 Efficiency Study
  • 4.6 Label Word Selection
  • 4.6.1 Impact of Lexicon Quality on Label Word Selection
  • 4.7 Effect of Further Pre-training
  • 5 Related Works
  • 5.1 Template-based prompt learning
  • 5.2 Few-shot NER
  • 6 Conclution
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Implementation Details
  • A.2 Sampling Algorithm
  • A.3 Effect of Conflict threshold
  • A.4 Effect of k in virtual method
  • A.5 Comparison with Comprehensive few-shot NER benchmark
  • A.6 Case Study

Knowls

  1. Knowl 1 — Entity-Oriented Language Model Fine-Tuning Objective

    model/method

    Entity-oriented Language Model (EntLM) fine-tuning reformulates token-level named entity recognition (NER) as a masked word prediction task without introducing any new classification parameters or textual prompt templates.

    Let an input token sequence be X={x1,x2,…,xn}X = \{x_1, x_2, \dots, x_n\} with a corresponding label sequence Y={y1,y2,…,yn}Y = \{y_1, y_2, \dots, y_n\}, where each yi∈Yy_i \in \mathcal{Y}. A verbalizer mapping function M:Y→Vl\mathcal{M}: \mathcal{Y} \to \mathcal{V}_l maps each entity label to a label word (or virtual prototype representation) in vocabulary Vl\mathcal{V}_l. A target sequence XEnt={x1Ent,x2Ent,…,xnEnt}X^{\text{Ent}} = \{x_1^{\text{Ent}}, x_2^{\text{Ent}}, \dots, x_n^{\text{Ent}}\} is constructed by replacing the token at entity position ii with its mapped label word M(yi)\mathcal{M}(y_i) and keeping the original token xix_i at non-entity positions (i.e., when yi="O"y_i = \text{"O"}).

    The pre-trained language model is trained using the EntLM objective by minimizing the negative log-likelihood of the target sequence:

    LEntLM=−∑i=1nlog⁡P(xi=xiEnt∣X)\mathcal{L}_{\text{EntLM}} = -\sum_{i=1}^n \log P(x_i = x_i^{\text{Ent}} \mid X)

    where the conditional token probability is computed using the pre-trained language model output head parameters WlmW_{lm}:

    P(xi=xiEnt∣X)=Softmax(Wlm⋅hi)P(x_i = x_i^{\text{Ent}} \mid X) = \text{Softmax}(W_{lm} \cdot h_i)

    Here, hih_i represents the contextualized hidden representation of the ii-th token produced by the pre-trained Transformer encoder. Entities spanning multiple subwords or tokens use an "IO" labeling scheme where every token of the multi-token entity is assigned the same label word M(y)\mathcal{M}(y).

  2. Knowl 2 — EntLM Inference and Decoding

    model/method

    During inference, EntLM performs named entity recognition in a single forward pass without requiring span enumeration. For an input sentence X={x1,…,xn}X = \{x_1, \dots, x_n\}, the probability that the ii-th token belongs to entity class y∈Yy \in \mathcal{Y} is directly given by the pre-trained language model head's predicted probability for the corresponding label word M(y)\mathcal{M}(y):

    p(yi=y∣X)=p(xi=M(y)∣X)p(y_i = y \mid X) = p(x_i = \mathcal{M}(y) \mid X)

    To capture transition dependencies between adjacent token labels, EntLM can be coupled with a structured Viterbi decoding algorithm (EntLM+Struct). In EntLM+Struct, the emission probabilities derived from the language model head predictions p(yi=y∣X)p(y_i = y \mid X) are combined with transition probabilities estimated from lexicon-annotated or unlabeled data via Viterbi decoding to identify the optimal global sequence of labels Y^={y^1,…,y^n}\hat{Y} = \{\hat{y}_1, \dots, \hat{y}_n\}.

  3. Knowl 3 — Label Word Selection and Prototype Construction

    model/method

    To automatically find representative pivot words for each entity class C∈YC \in \mathcal{Y} without requiring heavy manual engineering or overfitting on few-shot samples, label words can be derived from distant supervision using lexicon-annotated data Dlexicon={(Xi,Yi∗)}i=1N\mathcal{D}_{\text{lexicon}} = \{(X_i, Y_i^*)\}_{i=1}^N via three main discrete selection methods and a continuous prototype method:

    1. Data Search: Ranks words by their empirical occurrence frequency under class CC in the lexicon-annotated dataset:

    M(C)=arg⁡max⁡w∈V∑(X,Y∗)∈D∑i=1∣X∣ϕ(xi=w,yi∗=C)\mathcal{M}(C) = \arg\max_{w \in \mathcal{V}} \sum_{(X, Y^*) \in \mathcal{D}} \sum_{i=1}^{|X|} \phi(x_i = w, y_i^* = C)

    where ϕ(xi=w,yi∗=C)=I(xi=w)⋅I(yi∗=C)\phi(x_i = w, y_i^* = C) = \mathbb{I}(x_i = w) \cdot \mathbb{I}(y_i^* = C).

    1. LM Search: Ranks words by how frequently word ww appears in the top-kk predictions of the pre-trained language model at token positions labeled as class CC:

    M(C)=arg⁡max⁡w∈V∑(X,Y∗)∈D∑i=1∣X∣ϕtopk(x^i=w,yi∗=C)\mathcal{M}(C) = \arg\max_{w \in \mathcal{V}} \sum_{(X, Y^*) \in \mathcal{D}} \sum_{i=1}^{|X|} \phi_{\text{top}k}(\hat{x}_i = w, y_i^* = C)

    where ϕtopk(x^i=w,yi∗=C)=Itopk(x^i=w∣X,Y∗)⋅I(yi∗=C)\phi_{\text{top}k}(\hat{x}_i = w, y_i^* = C) = \mathbb{I}_{\text{top}k}(\hat{x}_i = w \mid X, Y^*) \cdot \mathbb{I}(y_i^* = C).

    1. Data & LM Search: Combines the data distribution and the LM prediction distribution multiplicatively:

    M(C)=arg⁡max⁡w∈V{(∑(X,Y∗)∈D∑i=1∣X∣ϕ(xi=w,yi∗=C))⋅(∑(X,Y∗)∈D∑i=1∣X∣ϕtopk(x^i=w,yi∗=C))}\mathcal{M}(C) = \arg\max_{w \in \mathcal{V}} \left\{ \left( \sum_{(X, Y^*) \in \mathcal{D}} \sum_{i=1}^{|X|} \phi(x_i = w, y_i^* = C) \right) \cdot \left( \sum_{(X, Y^*) \in \mathcal{D}} \sum_{i=1}^{|X|} \phi_{\text{top}k}(\hat{x}_i = w, y_i^* = C) \right) \right\}

    1. Virtual Label Words (Prototypes): Instead of selecting a single discrete token, the continuous representation of label CC is computed as the mean vector (prototype) of the top-kk candidate word embedding vectors identified by Data&LM Search.
  4. Knowl 4 — Conflict Word Filtering Criterion

    equation

    To prevent high-frequency, polysemous, or general words from causing training conflicts across different entity classes, candidate label words selected for class CC are retained only if their class-specific frequency exceeds a predefined conflict threshold ratio Th∈(0,1)Th \in (0, 1):

    ϕ(x=w,y∗=C)∑k∈Yϕ(x=w,y∗=k)>Th\frac{\phi(x = w, y^* = C)}{\sum_{k \in \mathcal{Y}} \phi(x = w, y^* = k)} > Th

    where ϕ(x=w,y∗=C)\phi(x = w, y^* = C) is the frequency of word ww occurring with class label CC in the lexicon-annotated corpus, and the denominator sums the occurrence of word ww across all entity classes k∈Yk \in \mathcal{Y}. In typical implementations, ThTh is set to 0.60.6.

  5. Knowl 5 — Exact Few-Shot Mention Sampling for NER

    algorithm

    In named entity recognition, a single sentence may contain multiple entity mentions across different classes. Greedy sentence sampling often leads to class imbalance. The exact sampling algorithm ensures that exactly KK entity mentions are sampled for each class in the label space Y\mathcal{Y}:

    Input: Required shot count KK, labeled training dataset DD with label set Y\mathcal{Y}
    Output: Sampled support set SS
    S←∅S \leftarrow \emptyset
    for each class i∈Yi \in \mathcal{Y} do
        Count[ii] ←0\leftarrow 0
    end for
    Shuffle DD
    for each sample (X,Y)∈D(X, Y) \in D do
        Add ←\leftarrow True
        for each class i∈Yi \in \mathcal{Y} do
            Calculate Temp_count[ii] as the number of mentions of class ii in (X,Y)(X, Y)
            if Count[ii] + Temp_count[ii] > KK then
                Add ←\leftarrow False
            end if
        end for
        if Add is True then
            S←S∪{(X,Y)}S \leftarrow S \cup \{(X, Y)\}
            for each class i∈Yi \in \mathcal{Y} do
                Count[ii] ←\leftarrow Count[ii] + Temp_count[ii]
            end for
        end if
        if Count[ii] == KK for all i∈Yi \in \mathcal{Y} then
            break
        end if
    end for
    return SS
  6. Knowl 6 — Few-Shot NER Performance Benchmark Across Domains

    data/table

    Under the KK-shot setting without source-domain training data, EntLM and its structured decoding variant (EntLM + Struct) outperform standard classification fine-tuning (BERT-tagger), metric-based nearest-neighbor methods (NNShot, StructShot), and template-based prompt learning (Template NER). Evaluation is conducted across 3 splits with 4 repeats per split using BERT-base-cased.

    Datasets Methods K=5 K=10 K=20 K=50
    CoNLL03 BERT-tagger (IO) 41.87 (12.12) 59.91 (10.65) 68.66 (5.13) 73.20 (3.09)
    NNShot 42.31 (8.92) 59.24 (11.71) 66.89 (6.09) 72.63 (3.42)
    StructShot 45.82 (10.30) 62.37 (10.96) 69.51 (6.46) 74.73 (3.06)
    Template NER 43.04 (6.15) 57.86 (5.68) 66.38 (6.09) 72.71 (2.13)
    EntLM 49.59 (8.30) 64.79 (3.86) 69.52 (4.48) 73.66 (2.06)
    EntLM + Struct 51.32 (7.67) 66.86 (3.01) 71.23 (3.91) 74.80 (1.87)
    OntoNotes 5.0 BERT-tagger (IO) 34.77 (7.16) 54.47 (8.31) 60.21 (3.89) 68.37 (1.72)
    NNShot 34.52 (7.85) 55.57 (9.20) 59.59 (4.20) 68.27 (1.54)
    StructShot 36.46 (8.54) 57.15 (5.84) 62.22 (5.10) 68.31 (5.72)
    Template NER 40.52 (8.62) 49.89 (3.66) 59.53 (2.25) 65.15 (2.95)
    EntLM 45.21 (9.17) 57.64 (4.18) 65.64 (4.24) 71.77 (1.31)
    EntLM + Struct 46.60 (10.35) 59.35 (3.24) 67.91 (4.55) 73.52 (0.97)
    MIT-Movie BERT-tagger (IO) 39.57 (6.38) 50.60 (7.29) 59.34 (3.66) 71.33 (3.04)
    NNShot 38.97 (5.54) 50.47 (6.09) 58.94 (3.47) 71.17 (2.85)
    StructShot 41.60 (8.97) 53.19 (5.52) 61.42 (2.98) 72.07 (6.41)
    Template NER 45.97 (3.86) 49.30 (3.35) 59.09 (0.35) 65.13 (0.17)
    EntLM 46.62 (9.46) 57.31 (3.72) 62.36 (4.14) 71.93 (1.68)
    EntLM + Struct 49.15 (8.91) 59.21 (3.96) 63.85 (3.70) 72.99 (1.80)

    The results show that EntLM achieves absolute improvements of up to 9.45% on CoNLL03, 11.83% on OntoNotes 5.0, and 9.58% on MIT-Movie over BERT-tagger at K=5K=5, while reducing prediction variance.

  7. Knowl 7 — Decoding Speed and Computational Efficiency Comparison

    data/table

    Template-based NER methods require enumerating candidate spans and evaluating them via separate template queries (O(n2)O(n^2) complexity with sentence length nn), causing severe decoding latency. Because EntLM performs token prediction in a single forward pass, it achieves decoding speeds comparable to standard BERT classification heads while running up to 1930.12 times faster than Template NER.

    Methods CoNLL (s) OntoNotes (s) MIT-Movie (s)
    BERT-tagger 8.57 23.89 6.46
    TemplateNER 6,491.00 50,241.00 5,254.00
    NNShot 16.03 82.62 15.98
    StructShot 19.84 98.67 17.66
    EntLM 9.26 26.03 6.64
    EntLM + Struct 13.40 34.92 7.38

    Decoding times are measured on a single NVIDIA TiTan XP GPU with a batch size of 8 (TemplateNER evaluated with batch size 45 to accommodate 9-gram span enumeration).

  8. Knowl 8 — Comparison of Label Word Selection Techniques

    data/table

    The effectiveness of discrete versus continuous prototype (virtual) label word selection strategies is evaluated across few-shot settings on three benchmarks. Continuous virtual vectors consistently outperform single discrete words, and combining corpus statistics with language model prediction distributions (Data&LM + Virtual) yields the highest performance.

    Methods CoNLL03 OntoNotes MIT-Movie
    K=5 K=10 K=5 K=10 K=5 K=10
    DataSearch 50.00 (9.75) 61.31 (4.73) 36.94 (5.04) 49.54 (5.02) 39.25 (4.83) 51.65 (5.52)
    LMSearch 48.40 (6.81) 59.39 (5.50) 36.98 (6.71) 48.20 (5.46) 39.12 (4.18) 48.30 (3.76)
    DataLMSearch 49.55 (7.76) 61.00 (6.98) 36.60 (7.90) 50.64 (6.12) 38.86 (11.43) 50.42 (6.45)
    Data + Virtual 49.25 (4.96) 63.40 (5.13) 45.61 (10.51) 55.13 (4.95) 45.59 (8.25) 55.10 (4.42)
    LM + Virtual 42.65 (12.58) 59.39 (5.50) 45.29 (7.77) 54.50 (3.66) 46.23 (5.60) 54.92 (6.15)
    DataLM + Virtual 49.59 (8.30) 64.79 (3.86) 45.21 (9.17) 57.64 (4.18) 46.62 (9.46) 57.31 (3.72)

    Virtual prototypes derived only from LM predictions (LM + Virtual) degrade in 5-shot settings (42.65% on CoNLL03), highlighting that pre-trained general knowledge alone lacks domain specificity without corpus data grounding.

  9. Knowl 9 — Effect of Domain-Specific Masked Language Model Pre-Training on EntLM

    data/table

    Performing further pre-training on target-domain unlabeled data using the standard Masked Language Model (MLM) objective aligns the language model's output distribution with the target task distribution. Because EntLM reuses the language model prediction head, it benefits substantially more from in-domain pre-training than classifier-head fine-tuning.

    Methods CoNLL03
    K=5 K=10
    BERT-tagger 41.87 (12.12) 59.91 (10.65)
    EntLM 49.59 (8.30) 64.79 (3.86)
    EntLM + Struct 51.32 (7.67) 66.86 (3.01)
    BERT-tagger (further) 41.16 (10.41) 61.70 (5.15)
    EntLM (further) 56.82 (12.27) 66.82 (4.65)
    EntLM + Struct (further) 58.77 (12.16) 68.96 (4.41)

    On CoNLL03 5-shot NER, further domain pre-training increases EntLM's F1 score by +7.23 points (from 49.59% to 56.82%) and EntLM + Struct by +7.45 points (from 51.32% to 58.77%), whereas BERT-tagger shows negligible change (-0.71 points).

  10. Knowl 10 — RoBERTa-Base Benchmark Evaluation in 5-Shot NER

    data/table

    When evaluated under the standardized RoBERTa-base few-shot benchmark, EntLM outperforms metric-based and self-training baselines across 5-shot settings on CoNLL'03 and MIT-Movie datasets without relying on large-scale supervised pre-training.

    Methods CoNLL (5-shot) MIT-Movie (5-shot)
    Linear Classifier (LC) 53.5 51.3
    LC + Noisy Supervised Pretraining (NSP) 61.4 53.1
    Prototype (Proto) 58.4 38.0
    Proto + NSP 60.9 43.8
    LC + Self-Training (ST) 56.7 54.1
    LC + NSP + ST 65.4 55.9
    EntLM 68.6 55.2
    EntLM (Struct) 69.9 57.1

    EntLM and EntLM (Struct) outperform methods that incorporate 6.8 GB WiFiNE supervised pre-training (NSP) and semi-supervised self-training (ST).

Coverage note — None was omitted; all key contributions, theoretical formulations, algorithms, and empirical evaluation results are covered.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  2. 2.Xiang Chen, Ningyu Zhang, Lei Li, Xin Xie, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2021. Lightner: A lightweight generative framework with prompt-guided attention for low-resource ner. arXiv preprint arXiv:2109.00720.
  3. 3.Leyang Cui, Yu Wu, Jian Liu, Sen Yang, and Yue Zhang. 2021. Template-based named entity recognition using BART. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1835–1845, Online. Association for Computational Linguistics.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, and Zhiyuan Liu. 2021. Few-NERD: A few-shot named entity recognition dataset. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3198–3213, Online. Association for Computational Linguistics.
  6. 6.Alexander Fritzler, Varvara Logacheva, and Maksim Kretov. 2019. Few-shot classification in named entity recognition task. Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing.
  7. 7.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  8. 8.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021. Ptr: Prompt tuning with rules for text classification.
  9. 9.Maximilian Hofer, Andrey Kormilitzin, Paul Goldberg, and Alejo Nevado-Holgado. 2018. Few-shot learning for named entity recognition in medical text.
  10. 10.Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, and Jiawei Han. 2020. Few-shot named entity recognition: A comprehensive study.
  11. 11.Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, and Jiawei Han. 2021. Few-shot named entity recognition: An empirical baseline study. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10408–10423, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  13. 13.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  14. 14.Jing Li, Billy Chiu, Shanshan Feng, and Hao Wang. 2020. Few-shot named entity recognition via meta-learning. IEEE Transactions on Knowledge and Data Engineering, pages 1–1.
  15. 15.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  16. 16.Chen Liang, Yue Yu, Haoming Jiang, Siawpeng Er, Ruijia Wang, Tuo Zhao, and Chao Zhang. 2020. BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision, page 1054–1064. Association for Computing Machinery, New York, NY, USA.
  17. 17.Jingjing Liu, Panupong Pasupat, Yining Wang, Scott Cyphers, and Jim Glass. 2013. Query understanding enhanced by hierarchical parsing structures. In 2013 IEEE Workshop on Automatic Speech Recognition and Understanding, pages 72–77. IEEE.
  18. 18.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021. GPT understands, too. CoRR, abs/2103.10385.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  20. 20.Andrea Madotto, Zihan Liu, Zhaojiang Lin, and Pascale Fung. 2020. Language models as few-shot learner for task-oriented dialogue systems.
  21. 21.Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. How context affects language models’ factual predictions. In Automated Knowledge Base Construction.
  22. 22.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  23. 23.Erik F Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050.
  24. 24.Timo Schick and Hinrich Schütze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  25. 25.Timo Schick and Hinrich Schütze. 2021b. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352, Online. Association for Computational Linguistics.
  26. 26.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  27. 27.Meihan Tong, Shuai Wang, Bin Xu, Yixin Cao, Minghui Liu, Lei Hou, and Juanzi Li. 2021. Learning from miscellaneous other-class words for few-shot named entity recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6236–6247, Online. Association for Computational Linguistics.
  28. 28.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23.
  29. 29.Yi Yang and Arzoo Katiyar. 2020. Simple and effective few-shot named entity recognition with structured nearest neighbor learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6365–6375, Online. Association for Computational Linguistics.

Citation

MLA
Ma, R., et al. “Template-free Prompt Tuning for Few-shot NER”. arXiv, 2021, http://arxiv.org/abs/2109.13532v3.
APA
Ma, R., Zhou, X., Gui, T., Tan, Y., Li, L., Zhang, Q., & Huang, X. (2021). Template-free Prompt Tuning for Few-shot NER. arXiv. http://arxiv.org/abs/2109.13532v3
Chicago
Ma, R., X. Zhou, T. Gui, et al. 2021. “Template-free Prompt Tuning for Few-shot NER”. arXiv. http://arxiv.org/abs/2109.13532v3.
Harvard
Ma, R. et al. (2021) “Template-free Prompt Tuning for Few-shot NER”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2109.13532v3.
Vancouver
1. Ma R, Zhou X, Gui T, Tan Y, Li L, Zhang Q, Huang X (2021) Template-free Prompt Tuning for Few-shot NER. arXiv

BibTeX

@article{ma2021template,
  title = {Template-free Prompt Tuning for Few-shot NER},
  author = {Ma, Ruotian and Zhou, Xin and Gui, Tao and Tan, Yiding and Li, Linyang and Zhang, Qi and Huang, Xuanjing},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2109.13532v3},
  eprint = {2109.13532}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/