Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification

Shengding HuNing DingHuadong WangZhiyuan LiuJingang WangJuanzi LiWei WuMaosong Sun

article2022ACL454 citations

Proposes Knowledgeable Prompt-tuning (KPT), a framework that integrates external knowledge bases to expand and refine verbalizer label spaces, significantly improving classification accuracy and stability in zero- and few-shot settings.

Listen

Prompt-tuning has emerged as a dominant method for applying large pre-trained language models to text classification, especially when labeled training data is scarce. In standard prompt-tuning, an input text is placed inside a template with a blank space, and the model predicts a single target word (the verbalizer) mapped to a category label. However, standard verbalizers rely on handcrafted single words or optimization-based searches, which suffer from narrow semantic coverage, personal bias, and high prediction variance.

The article introduces Knowledgeable Prompt-tuning (KPT), an approach designed to improve the accuracy and stability of text classification in low-data regimes by incorporating external knowledge bases into the verbalizer. The authors evaluated KPT on five benchmark datasets across topic classification and sentiment analysis in zero-shot and few-shot scenarios using a RoBERTa-large language model.

The KPT framework operates in three stages: construction, refinement, and utilization. In the construction stage, external knowledge sources (such as ConceptNet, WordNet, and sentiment dictionaries) generate a rich set of related label words across multiple granularities for each class. To remove noise from this automated expansion, the framework applies four refinement methods: frequency refinement, relevance refinement, contextualized calibration, and learnable weighting. Finally, predictions across the refined label words are combined into category decisions using standard averaging for zero-shot tasks or weighted averaging for few-shot tasks.

The experimental findings show that KPT consistently outperforms traditional fine-tuning, standard manual prompt-tuning, and automated verbalizer baselines. In few-shot settings, KPT reduced classification error rates by an average of 18% in 1-shot, 10% in 5-shot, and 7% in 10-shot tasks compared to the strongest baselines. In zero-shot settings, KPT combined with contextualized calibration reduced error rates by 16% on average, showing gains of up to 11% on complex topic classification tasks. In addition to accuracy gains, the expanded vocabulary stabilized model training, resulting in lower prediction variance across random seeds and templates.

These results indicate that expanding label vocabularies with structured external knowledge significantly enhances sample efficiency, allowing organizations to deploy high-performing classification systems without expensive, manual data annotation. The findings also demonstrate that contextual calibration is critical for zero-shot tasks, whereas supervised training data naturally mitigates the need for calibration in few-shot tasks.

Organizations operating in data-scarce environments should adopt knowledgeable prompt expansion over single-word verbalizers to reduce labeling costs and deployment risk. In zero-shot workflows, contextual calibration using a small pool of unlabeled text should be prioritized. For future implementations, practitioners should explore integrating self-supervised word-mining techniques for domains where structured external knowledge bases are unavailable, while remaining cautious of potential noise or malicious entries introduced from unverified third-party knowledge bases.

Cover for Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification

Abstract

Tuning pre-trained language models (PLMs) with task-specific prompts has been a promising approach for text classification. Particularly, previous studies suggest that prompt-tuning has remarkable superiority in the low-data scenario over the generic fine-tuning methods with extra classifiers. The core idea of prompt-tuning is to insert text pieces, i.e., template, to the input and transform a classification problem into a masked language modeling problem, where a crucial step is to construct a projection, i.e., verbalizer, between a label space and a label word space. A verbalizer is usually handcrafted or searched by gradient descent, which may lack coverage and bring considerable bias and high variances to the results. In this work, we focus on incorporating external knowledge into the verbalizer, forming a knowledgeable prompt-tuning (KPT), to improve and stabilize prompt-tuning. Specifically, we expand the label word space of the verbalizer using external knowledge bases (KBs) and refine the expanded label word space with the PLM itself before predicting with the expanded label word space. Extensive experiments on zero and few-shot text classification tasks demonstrate the effectiveness of knowledgeable prompt-tuning. Our source code is publicly available at https://github.com/thunlp/KnowledgeablePromptTuning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Knowledgeable Prompt-tuning
  • 3.1 Overview
  • 3.2 Verbalizer Construction
  • 3.3 Verbalizer Refinement
  • 3.4 Verbalizer Utilization
  • 3.5 Theoretical Illustration of KPT
  • 4 Experiments
  • 4.1 Datasets and Templates
  • 4.2 Experiment Settings
  • 4.3 Baselines
  • 4.4 Main Results
  • 5 Analysis
  • 5.1 Diversity of Top Predicted Words
  • 5.2 Other Analyses
  • 6 Conclusion
  • Acknowledgements
  • Ethical Considerations
  • References
  • A Pilot Experiments
  • B A Theoretical Illustration of KPT
  • C Practical Issues of Refinement
  • D Further Analyses and Ablation Studies
  • D.1 Calibration and Contextualized Calibration.
  • D.2 Supervised Data Ease the Need for Calibration.
  • D.3 How to Handle the OOV Label Words?
  • D.4 Visualization of the Refinement Process.
  • E Datasets and Templates
  • D.5 Potential Usage without External KB.
  • F Experimental Settings

Knowls

  1. Knowl 1 — Knowledgeable Prompt-Tuning Framework

    model/method

    Knowledgeable Prompt-Tuning (KPT) is a prompt-tuning framework for text classification that enhances the verbalizer by expanding and refining the label word space using external knowledge bases.

    In standard prompt-tuning, an input sequence x=(x0,x1,…,xn)x = (x_0, x_1, \dots, x_n) is wrapped into a natural language template containing a masked position token [extMASK][ ext{MASK}] to create a prompt input xpx_p. A pre-trained language model MM predicts the probability of filling [extMASK][ ext{MASK}] with vocabulary words, PM([extMASK]=v∣xp)P_M([ ext{MASK}] = v \mid x_p). A verbalizer is a mapping f:V→Yf: \mathcal{V} \to \mathcal{Y} from a label word vocabulary V\mathcal{V} to the label space Y\mathcal{Y}, where Vy⊂V\mathcal{V}_y \subset \mathcal{V} denotes the subset of label words mapped to class y∈Yy \in \mathcal{Y}, such that ⋃y∈YVy=V\bigcup_{y \in \mathcal{Y}} \mathcal{V}_y = \mathcal{V}.

    KPT implements a three-stage pipeline:

    1. Construction: External structured knowledge bases generate a multi-word label set Vy\mathcal{V}_y for each class yy.
    2. Refinement: Noisy candidate words are filtered or reweighted using frequency refinement, relevance refinement, contextualized calibration, and/or learnable refinement.
    3. Utilization: Refined word-level prediction probabilities are aggregated into label-level class scores via unweighted average (in zero-shot settings) or weighted average (in few-shot settings).
  2. Knowl 2 — Bayesian Formulation of Multi-Word Prompt Verbalizers

    theoretical result

    In prompt-tuning with multi-word verbalizers, predicting the class label y∈Yy \in \mathcal{Y} given an input xx wrapped in template xpx_p can be formalized as marginalizing the joint probability over all label words v∈Vyv \in \mathcal{V}_y associated with label yy:

    p(Y=y∣x)=∑v∈Vyp(Y=y,[MASK]=v∣x)p(Y = y \mid x) = \sum_{v \in \mathcal{V}_y} p(Y = y, [\text{MASK}] = v \mid x)

    Assuming the class label YY is conditionally independent of xx given the label word prediction vv, and applying Bayes' theorem under a balanced class prior p(Y=y)p(Y = y):

    p(Y=y∣x)∝∑v∈Vyp([MASK]=v∣Y=y)p([MASK]=v)PM([MASK]=v∣xp)p(Y = y \mid x) \propto \sum_{v \in \mathcal{V}_y} \frac{p([\text{MASK}] = v \mid Y = y)}{p([\text{MASK}] = v)} P_M([\text{MASK}] = v \mid x_p)

    where:

    • p([MASK]=v∣Y=y)p([\text{MASK}] = v \mid Y = y) is the conditional probability of word vv given class yy, reflecting semantic relevance (addressed via relevance refinement or learnable refinement).
    • p([MASK]=v)p([\text{MASK}] = v) is the prior prediction probability of word vv at the masked position independent of class label (addressed via contextualized calibration).
    • PM([MASK]=v∣xp)P_M([\text{MASK}] = v \mid x_p) is the conditional likelihood output by the masked language model head for word vv at the [extMASK][ ext{MASK}] position of xpx_p.
  3. Knowl 3 — External Knowledge Expansion for Prompt Verbalizers

    model/method

    In Knowledgeable Prompt-Tuning (KPT), external knowledge sources are utilized to generate an expanded, multi-word label word set Vy\mathcal{V}_y for each class y∈Yy \in \mathcal{Y}:

    1. Topic Classification: Candidate words are extracted from a knowledge graph GG (such as Related Words, which aggregates WordNet, ConceptNet, and word embeddings) where edges denote relevance relations annotated with relevance scores. Taking the anchor class name v0v_0 as the seed node, all neighborhood nodes NG(v0)N_G(v_0) having relevance scores greater than a threshold η\eta (set to η=0\eta = 0) are selected, yielding the expanded set Vy=NG(v0)∪{v0}\mathcal{V}_y = N_G(v_0) \cup \{v_0\}.
    2. Sentiment Classification: Candidate label words across various sentiment granularities and aspects are extracted from comprehensive sentiment lexicons categorized into positive and negative polarity lists.
  4. Knowl 4 — Frequency Refinement via Contextualized Prior

    model/method

    Frequency Refinement filters out words from the expanded verbalizer Vy\mathcal{V}_y that are rare to the pre-trained language model MM, as the model's masked language modeling probabilities on rare words tend to be inaccurate.

    Given the sentence distribution D\mathcal{D}, the contextualized prior probability of a word vv at the [extMASK][ ext{MASK}] position is defined as:

    PD(v)=Ex∼D[PM([MASK]=v∣xp)]P_{\mathcal{D}}(v) = \mathbb{E}_{x \sim \mathcal{D}} \left[ P_M([\text{MASK}] = v \mid x_p) \right]

    Using a small unlabeled support set C~\tilde{\mathcal{C}} of sentences sampled without labels (∣C~∣=200|\tilde{\mathcal{C}}| = 200), the contextualized prior is approximated by:

    PD(v)≈1∣C~∣∑x∈C~PM([MASK]=v∣xp)P_{\mathcal{D}}(v) \approx \frac{1}{|\tilde{\mathcal{C}}|} \sum_{x \in \tilde{\mathcal{C}}} P_M([\text{MASK}] = v \mid x_p)

    To avoid setting absolute probability thresholds across different tasks, a ranking-based cutoff is applied: label words whose contextualized prior probability falls in the lower 50%50\% of the distribution are removed.

  5. Knowl 5 — Relevance Refinement via Cosine Similarity and Norm-Based IDF Score

    model/method

    Relevance Refinement filters out candidate label words that are either insufficiently relevant to their target class or excessively correlated with multiple classes.

    For each label word vv and an unlabeled support set C~\tilde{\mathcal{C}}, a support prediction vector qv∈R∣C~∣q^v \in \mathbb{R}^{|\tilde{\mathcal{C}}|} is formed where the ii-th entry is qiv=PM([MASK]=v∣xip)q^v_i = P_M([\text{MASK}] = v \mid x_{ip}) for xi∈C~x_i \in \tilde{\mathcal{C}}. Using the predefined class name v0v_0 as the class representation qy=qv0q^y = q^{v_0}, the class relevance of word vv to class yy is computed as the cosine similarity:

    r(v,y)=cos⁡(qv,qv0)r(v, y) = \cos(q^v, q^{v_0})

    To penalize words that correlate with non-target classes, an adapted TF-IDF relevance metric is used. In general form with norm parameter dd:

    Rd(v)=r(v,f(v))(∣Y∣−1∑y∈Y,y≠f(v)r(v,y)d)1/dR^d(v) = r(v, f(v)) \left( \frac{|\mathcal{Y}| - 1}{\sum_{y \in \mathcal{Y}, y \ne f(v)} r(v, y)^d} \right)^{1/d}

    where f(v)f(v) is the class assigned to word vv, and dd is set adaptively based on label set size ∣Y∣|\mathcal{Y}|:

    d=C∣Y∣−2+ϵ+1d = \frac{C}{|\mathcal{Y}| - 2 + \epsilon} + 1

    with constant C=10C = 10 and small constant 0<ϵ≪10 < \epsilon \ll 1. Label words satisfying Rd(v)<1R^d(v) < 1 are removed from the verbalizer.

  6. Knowl 6 — Contextualized Calibration and Zero-Shot Verbalizer Utilization

    model/method

    In zero-shot prompt-tuning, pre-trained language models exhibit an inherent token bias where certain label words are predicted with disproportionately high or low probabilities regardless of the input sentence. Contextualized Calibration (CC) normalizes the predicted MLM distribution using the word's contextualized prior PD(v)P_{\mathcal{D}}(v):

    P~M([MASK]=v∣xp)∝PM([MASK]=v∣xp)PD(v)\tilde{P}_M([\text{MASK}] = v \mid x_p) \propto \frac{P_M([\text{MASK}] = v \mid x_p)}{P_{\mathcal{D}}(v)}

    where P~M\tilde{P}_M is renormalized to sum to 1 over candidate words.

    In the zero-shot inference setting, the predicted class label y^\hat{y} is selected by averaging the calibrated probabilities of all refined label words in each class set Vy\mathcal{V}_y:

    y^=arg⁡max⁡y∈Y1∣Vy∣∑v∈VyP~M([MASK]=v∣xp)\hat{y} = \arg\max_{y \in \mathcal{Y}} \frac{1}{|\mathcal{V}_y|} \sum_{v \in \mathcal{V}_y} \tilde{P}_M([\text{MASK}] = v \mid x_p)

  7. Knowl 7 — Learnable Refinement and Weighted Utilization for Few-Shot Prompt Tuning

    model/method

    In few-shot learning scenarios, Knowledgeable Prompt-Tuning refines the contribution of individual label words by introducing a learnable scalar weight wv∈Rw_v \in \mathbb{R} for each word v∈Vv \in \mathcal{V}. The weights are normalized per class via a softmax over Vy\mathcal{V}_y:

    αv=exp⁡(wv)∑u∈Vyexp⁡(wu)\alpha_v = \frac{\exp(w_v)}{\sum_{u \in \mathcal{V}_y} \exp(w_u)}

    where wvw_v is initialized to 00 for all vv.

    Class prediction scores s(y∣xp)s(y \mid x_p) are computed as the weighted average of log-probabilities:

    s(y∣xp)=∑v∈Vyαvlog⁡PM([MASK]=v∣xp)s(y \mid x_p) = \sum_{v \in \mathcal{V}_y} \alpha_v \log P_M([\text{MASK}] = v \mid x_p)

    The probability assigned to class yy is:

    P(y∣xp)=exp⁡(s(y∣xp))∑y′∈Yexp⁡(s(y′∣xp))P(y \mid x_p) = \frac{\exp(s(y \mid x_p))}{\sum_{y' \in \mathcal{Y}} \exp(s(y' \mid x_p))}

    The model parameters and word weights {wv}\{w_v\} are jointly optimized using cross-entropy loss on the few-shot training instances.

  8. Knowl 8 — Zero-Shot Text Classification Results with KPT

    data/table

    Zero-shot text classification performance (Micro-F1 score, reported as average ±\pm standard error across 4 manual templates and 3 random seeds; best template in parentheses) using RoBERTa-large on five benchmarks:

    Method AG's News DBPedia Yahoo Amazon IMDB
    PT 75.1±6.275.1 \pm 6.2 (79.0) 66.6±2.366.6 \pm 2.3 (68.4) 45.4±7.045.4 \pm 7.0 (52.0) 80.2±8.880.2 \pm 8.8 (87.8) 86.4±4.086.4 \pm 4.0 (92.0)
    PT+CC 79.9±0.779.9 \pm 0.7 (81.0) 73.9±4.973.9 \pm 4.9 (82.6) 58.0±1.458.0 \pm 1.4 (58.8) 91.4±1.691.4 \pm 1.6 (93.5) 91.6±3.091.6 \pm 3.0 (93.7)
    KPT 84.8±1.2\mathbf{84.8 \pm 1.2} (86.7) 82.2±5.4\mathbf{82.2 \pm 5.4} (87.4) 61.6±2.2\mathbf{61.6 \pm 2.2} (63.8) 92.8±1.2\mathbf{92.8 \pm 1.2} (94.6) 91.6±2.7\mathbf{91.6 \pm 2.7} (94.0)
    -FR 82.7±1.582.7 \pm 1.5 (85.0) 81.8±4.681.8 \pm 4.6 (86.2) 60.9±1.560.9 \pm 1.5 (62.7) 92.8±1.2\mathbf{92.8 \pm 1.2} (94.6) 91.6±2.891.6 \pm 2.8 (94.1)
    -RR 81.4±1.581.4 \pm 1.5 (83.7) 81.4±4.581.4 \pm 4.5 (85.8) 60.1±1.060.1 \pm 1.0 (61.4) 92.8±1.2\mathbf{92.8 \pm 1.2} (94.6) 91.6±2.891.6 \pm 2.8 (94.1)
    -CC 55.5±2.855.5 \pm 2.8 (58.3) 64.5±6.864.5 \pm 6.8 (73.0) 42.4±5.042.4 \pm 5.0 (46.8) 86.2±5.786.2 \pm 5.7 (92.5) 90.3±2.890.3 \pm 2.8 (94.1)

    PT represents standard manual prompt tuning; PT+CC adds Contextualized Calibration. Ablations (-FR, -RR, -CC) remove frequency refinement, relevance refinement, or contextualized calibration from KPT. KPT reduces zero-shot classification error rate by 16%16\% on average over PT.

  9. Knowl 9 — Few-Shot Text Classification Results Across Shots

    data/table

    Few-shot classification performance (Micro-F1 score, average ±\pm standard error across 4 templates and 5 random seeds; best template in parentheses) using RoBERTa-large for 1, 5, 10, and 20 shots per class:

    Shot Method AG's News DBPedia Yahoo Amazon IMDB
    1 FT 19.8±10.419.8 \pm 10.4 8.6±4.58.6 \pm 4.5 11.1±4.011.1 \pm 4.0 49.9±0.249.9 \pm 0.2 50.0±0.050.0 \pm 0.0
    PT 80.0±6.080.0 \pm 6.0 (84.4) 92.2±2.592.2 \pm 2.5 (94.3) 54.2±3.154.2 \pm 3.1 (55.7) 91.9±2.791.9 \pm 2.7 (93.2) 91.2±3.791.2 \pm 3.7 (93.7)
    AUTO 52.8±9.852.8 \pm 9.8 (57.6) 63.0±8.963.0 \pm 8.9 (68.3) 23.3±4.523.3 \pm 4.5 (25.0) 66.6±12.566.6 \pm 12.5 (72.7) 75.5±15.575.5 \pm 15.5 (83.1)
    SOFT 80.0±5.680.0 \pm 5.6 (82.4) 92.3±2.392.3 \pm 2.3 (93.3) 54.3±2.754.3 \pm 2.7 (55.9) 90.9±5.890.9 \pm 5.8 (93.6) 89.4±8.989.4 \pm 8.9 (93.1)
    KPT 83.7±3.5\mathbf{83.7 \pm 3.5} (84.6) 93.7±1.8\mathbf{93.7 \pm 1.8} (95.3) 63.2±2.5\mathbf{63.2 \pm 2.5} (64.1) 93.2±1.3\mathbf{93.2 \pm 1.3} (93.9) 92.2±3.0\mathbf{92.2 \pm 3.0} (93.6)
    5 FT 37.9±10.037.9 \pm 10.0 95.8±1.395.8 \pm 1.3 25.3±14.225.3 \pm 14.2 52.1±1.352.1 \pm 1.3 51.4±1.451.4 \pm 1.4
    PT 82.7±2.782.7 \pm 2.7 (84.0) 97.0±0.697.0 \pm 0.6 (97.3) 62.4±1.762.4 \pm 1.7 (63.9) 92.2±3.392.2 \pm 3.3 (93.5) 91.9±3.191.9 \pm 3.1 (92.7)
    AUTO 72.2±10.172.2 \pm 10.1 (75.6) 88.8±3.988.8 \pm 3.9 (91.5) 49.6±4.349.6 \pm 4.3 (51.2) 87.5±7.487.5 \pm 7.4 (90.8) 86.8±10.186.8 \pm 10.1 (92.1)
    SOFT 82.8±2.782.8 \pm 2.7 (84.3) 97.0±0.697.0 \pm 0.6 (97.2) 61.8±1.861.8 \pm 1.8 (63.1) 93.2±1.693.2 \pm 1.6 (94.2) 91.6±3.491.6 \pm 3.4 (93.9)
    KPT 85.0±1.2\mathbf{85.0 \pm 1.2} (85.9) 97.1±0.4\mathbf{97.1 \pm 0.4} (97.3) 67.2±0.8\mathbf{67.2 \pm 0.8} (67.8) 93.4±1.9\mathbf{93.4 \pm 1.9} (94.1) 92.7±1.5\mathbf{92.7 \pm 1.5} (92.9)
    10 FT 75.9±8.475.9 \pm 8.4 93.8±2.293.8 \pm 2.2 43.8±17.943.8 \pm 17.9 83.0±7.083.0 \pm 7.0 76.2±8.776.2 \pm 8.7
    PT 84.9±2.484.9 \pm 2.4 (86.1) 97.6±0.497.6 \pm 0.4 (97.8) 64.3±2.264.3 \pm 2.2 (64.8) 93.9±1.393.9 \pm 1.3 (94.6) 93.0±1.793.0 \pm 1.7 (94.0)
    AUTO 81.4±3.881.4 \pm 3.8 (84.1) 91.5±3.491.5 \pm 3.4 (95.1) 58.7±3.158.7 \pm 3.1 (60.9) 93.7±1.293.7 \pm 1.2 (94.5) 91.1±5.191.1 \pm 5.1 (93.3)
    SOFT 85.0±2.885.0 \pm 2.8 (86.7) 97.6±0.497.6 \pm 0.4 (97.8) 64.5±2.264.5 \pm 2.2 (65.0) 93.9±1.793.9 \pm 1.7 (93.9) 91.8±2.691.8 \pm 2.6 (93.0)
    KPT 86.3±1.6\mathbf{86.3 \pm 1.6} (87.0) 98.0±0.2\mathbf{98.0 \pm 0.2} (98.1) 68.0±0.6\mathbf{68.0 \pm 0.6} (68.2) 93.8±1.293.8 \pm 1.2 (94.1) 92.9±1.892.9 \pm 1.8 (93.3)
    20 FT 85.4±1.885.4 \pm 1.8 97.9±0.297.9 \pm 0.2 54.2±18.154.2 \pm 18.1 71.4±4.371.4 \pm 4.3 78.5±10.178.5 \pm 10.1
    PT 86.5±1.686.5 \pm 1.6 (87.0) 97.9±0.397.9 \pm 0.3 (98.1) 67.2±1.167.2 \pm 1.1 (67.5) 93.5±1.093.5 \pm 1.0 (94.4) 93.0±1.193.0 \pm 1.1 (93.6)
    AUTO 85.7±1.485.7 \pm 1.4 (86.1) 92.2±2.792.2 \pm 2.7 (94.9) 65.0±1.865.0 \pm 1.8 (66.9) 93.9±1.1\mathbf{93.9 \pm 1.1} (94.1) 92.8±2.092.8 \pm 2.0 (94.0)
    SOFT 86.4±1.786.4 \pm 1.7 (87.1) 98.0±0.398.0 \pm 0.3 (98.1) 67.4±0.767.4 \pm 0.7 (67.5) 93.8±1.693.8 \pm 1.6 (94.2) 93.5±0.9\mathbf{93.5 \pm 0.9} (94.0)
    KPT 87.2±0.8\mathbf{87.2 \pm 0.8} (87.5) 98.1±0.3\mathbf{98.1 \pm 0.3} (98.2) 68.9±0.8\mathbf{68.9 \pm 0.8} (69.3) 93.7±1.693.7 \pm 1.6 (94.4) 93.1±1.193.1 \pm 1.1 (93.5)

    KPT achieves average error rate reductions over the best baseline of 17.8%17.8\%, 10.3%10.3\%, and 7.4%7.4\% on 1, 5, and 10 shots, respectively, while consistently exhibiting lower standard deviation.

  10. Knowl 10 — Sub-Token Probability Averaging for Out-of-Vocabulary Label Words

    empirical result

    When expanding prompt verbalizers via external knowledge bases, many candidate label words are out-of-vocabulary (OOV) for the tokenizer of the pre-trained language model and are split into multiple sub-tokens. For an OOV word split into tokens {t1,t2,…,tk}\{t_1, t_2, \dots, t_k\}, KPT computes the word's prediction score at the single [extMASK][ ext{MASK}] position by averaging the individual token MLM output probabilities:

    PM([MASK]=v∣xp)=1k∑j=1kPM([MASK]=tj∣xp)P_M([\text{MASK}] = v \mid x_p) = \frac{1}{k} \sum_{j=1}^k P_M([\text{MASK}] = t_j \mid x_p)

    Comparing this strategy against restricting the verbalizer strictly to single-token vocabulary words (KPT+ST) across 0, 1, 5, 10, and 20 shots demonstrates that restricting to single tokens does not yield consistent performance gains and often degrades classification Micro-F1 scores by minor margins. Multi-token OOV words from knowledge bases can thus be effectively incorporated into prompt verbalizers via simple token-probability averaging.

Coverage note — Omitted specific manual template surface strings and fine-tuning hyperparameter lists from Appendix E and F, as they represent implementation details that are standardly configurable rather than standalone conceptual contributions.

References

  1. 1.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  2. 2.Xiang Chen, Xin Xie, Ningyu Zhang, Jiahuan Yan, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2021. Adaprompt: Adaptive prompt-based finetuning for relation extraction. ArXiv preprint, abs/2104.07650.
  3. 3.Joe Davison, Joshua Feldman, and Alexander Rush. 2019. Commonsense knowledge mining from pre-trained models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1173–1178, Hong Kong, China. Association for Computational Linguistics.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Ning Ding, Shengding Hu, Weilin Zhao, Yulin Chen, Zhiyuan Liu, Hai-Tao Zheng, and Maosong Sun. 2021. Openprompt: An open-source framework for prompt-learning. ArXiv preprint, abs/2111.01998.
  6. 6.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  7. 7.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. WARP: Word-level Adversarial ReProgramming. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4921–4933, Online. Association for Computational Linguistics.
  8. 8.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021. Ptr: Prompt tuning with rules for text classification. ArXiv preprint, abs/2105.11259.
  9. 9.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051.
  10. 10.Karen Sparck Jones. 1972. A statistical interpretation of term specificity and its application in retrieval. Journal of documentation.
  11. 11.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in- and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1326–1340, Online. Association for Computational Linguistics.
  12. 12.Kamran Kowsari, Kiana Jafari Meimandi, Mojtaba Heidarysafa, Sanjana Mendu, Laura Barnes, and Donald Brown. 2019. Text classification algorithms: A survey. Information, 10(4):150.
  13. 13.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  14. 14.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al. 2015. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web, 6(2):167–195.
  15. 15.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021. Gpt understands, too. ArXiv preprint, abs/2103.10385.
  16. 16.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv preprint, abs/1907.11692.
  17. 17.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  18. 18.Julian J. McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: understanding rating dimensions with review text. In Seventh ACM Conference on Recommender Systems, RecSys '13, Hong Kong, China, October 12-16, 2013, pages 165–172. ACM.
  19. 19.Yu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong, Heng Ji, Chao Zhang, and Jiawei Han. 2020. Text classification using label names only: A language model self-training approach. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9006–9017, Online. Association for Computational Linguistics.
  20. 20.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset. In Proceedings of CoCo@ NeurIPS.
  21. 21.Ted Pedersen, Siddharth Patwardhan, and Jason Michelizzi. 2004. WordNet::Similarity - measuring the relatedness of concepts. In Demonstration Papers at HLT-NAACL 2004, pages 38–41, Boston, Massachusetts, USA. Association for Computational Linguistics.
  22. 22.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  23. 23.Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. How context affects language models’ factual predictions. In Automated Knowledge Base Construction.
  24. 24.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  25. 25.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  26. 26.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  27. 27.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  28. 28.Timo Schick, Helmut Schmid, and Hinrich Schütze. 2020. Automatically identifying words that can serve as labels for few-shot text classification. In Proceedings of COLING, pages 5569–5578.
  29. 29.Timo Schick and Hinrich Schütze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  30. 30.Timo Schick and Hinrich Schütze. 2021b. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352, Online. Association for Computational Linguistics.
  31. 31.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  32. 32.Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA, pages 4444–4451. AAAI Press.
  33. 33.Han Xu, Zhang Zhengyan, Ding Ning, Gu Yuxian, Liu Xiao, Huo Yuqi, Qiu Jiezhong, Zhang Liang, Han Wentao, Huang Minlie, et al. 2021. Pre-trained models: Past, present and future. ArXiv preprint, abs/2106.07139.
  34. 34.Wenpeng Yin, Jamaal Hay, and Dan Roth. 2019. Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3914–3923, Hong Kong, China. Association for Computational Linguistics.
  35. 35.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 649–657.
  36. 36.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.

Citation

MLA
Hu, S., et al. “Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification”. arXiv, 2021, http://arxiv.org/abs/2108.02035v2.
APA
Hu, S., Ding, N., Wang, H., Liu, Z., Wang, J., Li, J., Wu, W., & Sun, M. (2021). Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification. arXiv. http://arxiv.org/abs/2108.02035v2
Chicago
Hu, S., N. Ding, H. Wang, et al. 2021. “Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification”. arXiv. http://arxiv.org/abs/2108.02035v2.
Harvard
Hu, S. et al. (2021) “Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2108.02035v2.
Vancouver
1. Hu S, Ding N, Wang H, Liu Z, Wang J, Li J, Wu W, Sun M (2021) Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification. arXiv

BibTeX

@article{hu2021knowledgeable,
  title = {Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification},
  author = {Hu, Shengding and Ding, Ning and Wang, Huadong and Liu, Zhiyuan and Wang, Jingang and Li, Juanzi and Wu, Wei and Sun, Maosong},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2108.02035v2},
  eprint = {2108.02035}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/