Prompt-free and Efficient Few-shot Learning with Language Models

Rabeeh Karimi MahabadiLuke ZettlemoyerJames HendersonLambert MathiasMarzieh SaeidiVeselin StoyanovMajid Yazdani

article2022ACL78 citations

Introduces PERFECT, a prompt-free few-shot tuning framework for masked language models that eliminates manual prompt and verbalizer engineering using task-specific adapters and learned label embeddings, achieving up to 100x faster training and inference while outperforming state-of-the-art prompt-based methods.

Listen

Adapting large pretrained language models to new tasks using very limited data—known as few-shot learning—typically requires extensive manual engineering. Existing prompt-based methods convert inputs into fill-in-the-blank queries using handcrafted text templates and manually chosen target words. This manual process is labor-intensive, computationally expensive, and brittle, as minor wording changes can significantly harm performance.

The article evaluates a new framework called PERFECT, which eliminates handcrafted prompts and verbalizers to achieve efficient few-shot learning from as few as 32 labeled examples.

The researchers evaluated their method against standard fine-tuning and state-of-the-art prompt-based baselines across 12 standard language understanding benchmarks covering classification, sentiment analysis, and natural language inference. Instead of modifying the entire underlying model or engineering text prompts, the approach freezes the pretrained model, inserts compact task-specific adapter modules, and optimizes multi-token label embeddings alongside a distance-based prototype classification strategy.

The evaluation yielded several critical findings. First, the proposed method achieved state-of-the-art accuracy, outperforming the leading prompt-based baseline by 1.1 percentage points on single-sentence tasks and 4.6 percentage points on sentence-pair tasks, while notably outperforming even the best hand-tuned prompt configurations. Second, it delivered major computational efficiencies, reducing trainable parameters by roughly 99 percent, lowering peak memory usage by 22 percent, speeding up training by 97 percent, and accelerating inference by nearly 97 percent compared to standard prompt baselines. Third, the method substantially improved reliability, raising worst-case performance and reducing output variance across different data samples. Finally, ablation analyses showed that randomly initialized label embeddings outperformed manually engineered target words, confirming that manual verbalizer design is unnecessary.

These findings demonstrate that organizations can deploy high-performing few-shot language models without incurring substantial engineering labor or heavy infrastructure costs. By freezing the underlying base model and training only lightweight adapter layers, teams can drastically lower storage footprints and hardware requirements while stabilizing production performance against prompt sensitivity.

Organizations developing or deploying language models should consider replacing brittle prompt-engineering workflows with parameter-efficient adapter architectures and learned label embeddings. When implementing this architecture, teams should determine the optimal number of mask positions based on task complexity, as empirical results indicate that multi-mask setups provide varying gains across different problem types.

Confidence in these findings is high, given the consistent validation across 12 diverse benchmarks and multiple random initializations. However, leaders should note that the evaluation was primarily conducted using a single base model architecture with 355 million parameters, and results reflect constrained few-shot settings with exactly 16 training and 16 validation examples per class. Additional testing on larger foundation models and domain-specific production data is advised before full-scale deployment.

arXiv: 2204.01172facebookresearch/perfect.git
Cover for Prompt-free and Efficient Few-shot Learning with Language Models

Abstract

Current methods for few-shot fine-tuning of pretrained masked language models (PLMs) require carefully engineered prompts and verbalizers for each new task to convert examples into a cloze-format that the PLM can score. In this work, we propose PERFECT, a simple and efficient method for few-shot fine-tuning of PLMs without relying on any such handcrafting, which is highly effective given as few as 32 data points. PERFECT makes two key design choices: First, we show that manually engineered task prompts can be replaced with task-specific adapters that enable sample-efficient fine-tuning and reduce memory and storage costs by roughly factors of 5 and 100, respectively. Second, instead of using handcrafted verbalizers, we learn new multi-token label embeddings during fine-tuning, which are not tied to the model vocabulary and which allow us to avoid complex auto-regressive decoding. These embeddings are not only learnable from limited data but also enable nearly 100x faster training and inference. Experiments on a wide range of few shot NLP tasks demonstrate that PERFECT, while being simple and efficient, also outperforms existing state-of-the-art few-shot learning methods. Our code is publicly available at https://github.com/facebookresearch/perfect.git.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Adapters
  • 2.2 Prompt-based Fine-tuning
  • 3 Method
  • 3.1 Pattern-Free Task Description
  • 3.2 Multi-Token Label Embeddings
  • 3.3 Training PERFECT
  • 3.4 Inference with PERFECT
  • 4 Experiments
  • 4.1 Experimental Results
  • 4.2 Efficiency Evaluation
  • 4.3 Analysis
  • 4.3.1 Ablation Study
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Experimental Details
  • B Choice of Patterns and Verbalizers
  • C Impact of the Position of Masks in Sentence-pair Datasets
  • D Impact of Initialization
  • E Ablation Results

Knowls

  1. Knowl 1 — PERFECT Architecture for Prompt-Free Few-Shot Fine-Tuning

    model/method

    PERFECT (Prompt-free and Efficient paRadigm for FEw-shot Cloze-based fine-Tuning) is a framework designed for few-shot learning with Pretrained Masked Language Models (PLMs) without manual prompt templates (patterns) or handcrafted verbalizer dictionaries.

    Given an input sequence xinputx_{\text{input}}, PERFECT converts it into a masked format xmasked=T′(xinput)x_{\text{masked}} = \mathcal{T}'(x_{\text{input}}) containing MM consecutive mask tokens [MASK]1,…,[MASK]M[\text{MASK}]_1, \dots, [\text{MASK}]_M without any task-descriptive natural language prompts. The underlying PLM encoder parameters and embedding weights are kept frozen. Implicit task adaptation is achieved through task-specific bottleneck adapter layers inserted after feed-forward modules in each transformer block. An adapter layer computes:

    A(x)=U(GeLU(D(x)))+xA(x) = U(\text{GeLU}(D(x))) + x

    where x∈RHx \in \mathbb{R}^H is the input hidden representation (HH is hidden dimension), D(x)∈RH×BD(x) \in \mathbb{R}^{H \times B} is a down-projection with bottleneck dimension BB, GeLU\text{GeLU} is the activation function, and U(⋅)∈RB×HU(\cdot) \in \mathbb{R}^{B \times H} is an up-projection.

    Task target classes Y={1,…,K}\mathcal{Y} = \{1, \dots, K\} are parameterized by a dedicated continuous multi-token label embedding tensor L∈RK×M×HL \in \mathbb{R}^{K \times M \times H}, unconstrained by the language model's pre-existing vocabulary.

  2. Knowl 2 — Per-Token Multi-Class Hinge Loss for PERFECT Training

    equation

    During training of the PERFECT framework, predictions are computed at each of the MM mask token positions, and the parameters (adapters, layer normalizations, and label embeddings) are optimized via an average multi-class hinge loss.

    Let h[MASK]i∈RHh_{[\text{MASK}]_i} \in \mathbb{R}^H be the final hidden representation from the PLM encoder for the ii-th mask token position of input sample xmaskedx_{\text{masked}}. For each mask position i∈{1,…,M}i \in \{1, \dots, M\}, a per-token score vector ti∈RKt_i \in \mathbb{R}^K over the KK classes is defined by:

    ti=f(h[MASK]i)=LiTh[MASK]it_i = f(h_{[\text{MASK}]_i}) = L_i^T h_{[\text{MASK}]_i}

    where Li∈RK×HL_i \in \mathbb{R}^{K \times H} denotes the slice of the label embedding matrix corresponding to position ii, and tikt_{ik} denotes the score assigned to class kk.

    The per-token multi-class hinge loss for a labeled example (x,y)(x, y) at position ii with margin m=1m=1 is defined as:

    L(x,y,i)=1K∑k=1,k≠yKmax⁡(0,m−tiy+tik)\mathcal{L}(x, y, i) = \frac{1}{K} \sum_{k=1, k \neq y}^K \max(0, m - t_{iy} + t_{ik})

    The total training loss averaged over the training dataset D\mathcal{D} and all MM mask tokens is:

    L=1M∣D∣∑(x,y)∈D∑i=1ML(x,y,i)\mathcal{L} = \frac{1}{M |\mathcal{D}|} \sum_{(x, y) \in \mathcal{D}} \sum_{i=1}^M \mathcal{L}(x, y, i)

  3. Knowl 3 — Multi-Token Prototypical Inference in PERFECT

    model/method

    To evaluate a query example without requiring slow autoregressive decoding over variable-length verbalizers, PERFECT applies prototypical metric inference across multiple mask tokens.

    For each class y∈Yy \in \mathcal{Y} and each mask position i∈{1,…,M}i \in \{1, \dots, M\}, a prototype representation ciy∈RHc_{iy} \in \mathbb{R}^H is computed as the mean embedding of the ii-th mask token across all training samples Dy\mathcal{D}_y belonging to class yy:

    ciy=1∣Dy∣∑b∈Dyhibc_{iy} = \frac{1}{|\mathcal{D}_y|} \sum_{b \in \mathcal{D}_y} h_i^b

    where hibh_i^b is the hidden representation of the ii-th mask token for training sample bb.

    A test query point qq with mask hidden representations h1q,…,hMqh_1^q, \dots, h_M^q is classified by finding the label whose prototype achieves the maximum exponential negative squared Euclidean distance across all mask positions:

    y=arg⁡max⁡y∈Ymax⁡i∈{1,…,M}(exp⁡(−d(hiq,ciy)))y = \arg\max_{y \in \mathcal{Y}} \max_{i \in \{1, \dots, M\}} \left( \exp(-d(h_i^q, c_{iy})) \right)

    where d(u,v)=∥u−v∥22d(u, v) = \|u - v\|_2^2 is the squared Euclidean distance.

  4. Knowl 4 — Mask Token Placement Strategy for Single-Sentence and Sentence-Pair Inputs

    model/method

    In PERFECT, the placement of the MM mask tokens [MASK]1,…,[MASK]M[\text{MASK}]_1, \dots, [\text{MASK}]_M depends on the input sequence structure:

    1. Single-sentence benchmarks: The MM mask tokens are appended directly to the end of the input sentence ss.
    2. Sentence-pair benchmarks: For sentence pairs (s1,s2)(s_1, s_2), the MM mask tokens are inserted strictly between the two sentences, and the entire sequence is encoded as a single sentence:

    xmasked=[CLS]  s1  [MASK]1…[MASK]M  s2  [SEP]x_{\text{masked}} = [\text{CLS}] \; s_1 \; [\text{MASK}]_1 \dots [\text{MASK}]_M \; s_2 \; [\text{SEP}]

    Empirical validation on sentence-pair datasets (CB, RTE, QNLI, MRPC, QQP, WiC) shows that placing masks between sentences as a single sequence achieves an average validation accuracy of 76.0%, outperforming appending masks to the end (s1s2[MASK]s_1 s_2 \text{[MASK]}, 73.7%), inserting masks between sentences split with segment tokens (s1∣[MASK]s2s_1 \mid \text{[MASK]} s_2, 71.7%), and appending masks after segment-split sentences (s1∣s2[MASK]s_1 \mid s_2 \text{[MASK]}, 71.1%).

  5. Knowl 5 — Few-Shot Classification Performance on NLP Benchmarks

    data/table

    Across 12 benchmark NLP datasets evaluated in a 32-shot learning setting (16 training instances and 16 validation instances per class using RoBERTa-large with 355M parameters), PERFECT with randomly initialized label embeddings (PERFECT-rand) outperforms standard fine-tuning (FINETUNE) as well as prompt-based cloze fine-tuning (PET-Average and PET-Best across multiple prompt/verbalizer variants). Results are reported over 20 random runs (5 splits ×\times 4 random seeds) in terms of average accuracy, worst-case accuracy, and standard deviation:

    Method SST-2 CR MR SST-5 Subj TREC Avg
    Single-Sentence Benchmarks
    FINETUNE 81.4/70.0/4.0 80.1/72.9/4.1 77.7/66.8/4.6 39.2/34.3/2.5 90.2/84.1/1.8 87.6/75.8/3.7 76.0/67.3/3.4
    PET-Average 89.7/81.0/2.4 88.4/68.8/3.0 85.9/79.0/2.1 45.9/40.3/2.4 88.1/79.6/2.4 85.0/70.6/4.5 80.5/69.9/2.8
    PET-Best 89.1/81.0/2.6 88.8/85.8/1.9 86.4/82.0/1.6 46.0/41.2/2.4 88.7/84.6/1.8 85.8/70.6/4.4 80.8/74.2/2.4
    Logan IV et al. (2021) 89.8/84.1/1.7 89.9/87.2/1.1 84.9/76.2/3.2 45.7/41.6/2.3 81.8/73.5/4.0 84.7/81.8/1.6 79.5/74.1/2.3
    PERFECT-rand 90.7/88.2/1.2 90.0/85.5/1.4 86.3/81.4/1.6 42.7/35.1/2.9 89.1/82.8/2.1 90.6/81.6/3.2 81.6/75.8/2.1
    Method CB RTE QNLI MRPC QQP WiC Avg
    Sentence-Pair Benchmarks
    FINETUNE 72.9/67.9/2.5 56.8/50.2/3.5 62.7/51.4/7.0 70.1/62.7/4.7 65.0/59.8/3.6 52.4/46.1/3.7 63.3/56.4/4.2
    PET-Average 86.9/73.2/5.1 60.1/49.5/4.7 66.5/55.7/6.2 62.1/38.2/6.8 63.4/44.7/7.9 51.0/46.1/2.6 65.0/51.2/5.6
    PET-Best 90.0/78.6/3.9 62.3/51.3/4.5 70.5/57.9/6.4 63.4/49.3/6.5 70.7/55.2/5.8 51.6/47.2/2.3 68.1/56.6/4.9
    Logan IV et al. (2021) 91.0/87.5/2.7 64.4/58.5/3.9 71.2/66.5/2.6 63.9/53.7/5.3 70.4/62.7/3.4 52.4/48.4/1.8 68.9/62.9/3.3
    PERFECT-rand 90.3/83.9/3.5 60.4/53.1/4.7 74.1/60.3/4.6 67.8/54.7/5.7 71.2/64.2/3.5 53.8/47.0/3.0 69.6/60.5/4.2

    PERFECT-rand yields a +1.1 and +4.6 point accuracy gain over PET-Average on single-sentence and sentence-pair tasks respectively, while increasing minimum (worst-case) performance and reducing variance.

  6. Knowl 6 — Computational and Storage Efficiency of PERFECT versus PET

    data/table

    Efficiency measurements on a 500-sample QNLI dataset using a single NVIDIA A100 GPU (40GB memory) over 10 training epochs show substantial savings in parameters, memory, training time, and inference time for PERFECT compared to PET:

    Metric PET PERFECT %
    Trained params (M) 355.41 3.28 -99.08%
    Peak memory (GB) 20.93 16.34 -21.93%
    Training time (min) 23.42 0.65 -97.22%
    + PET in batch 0.94 0.65 -30.85%
    Inference time (min) 9.57 0.31 -96.76%

    PERFECT cuts the number of updated parameters by 99.08% (storing 3.28M adapter/label parameters instead of the full 355.41M PLM parameters), lowers peak memory by 21.93%, accelerates training time by 97.22% (and 30.85% when PET is modified to train with batched fixed-length verbalizers), and cuts inference time by 96.76% by replacing PET's multi-token autoregressive decoding (batch size 1) with parallel prototypical distance scoring.

  7. Knowl 7 — Efficacy of Adapters over Soft Prompts and Bias Tuning in Few-Shot Learning

    empirical result

    Replacing handcrafted prompt patterns with task-specific adapters yields superior performance and stability compared to alternative parameter-efficient tuning methods:

    1. Pattern-Free Adapter Substitution vs PET Patterns: Replacing handcrafted patterns in PET with adapters while preserving verbalizers improves the 12-dataset average from 72.8% to 75.4%, increases average worst-case accuracy from 60.6% to 69.8%, and decreases standard deviation from 4.2 to 2.7.
    2. Prompt-Tuning Ablation (prompt+mte): Optimizing continuous prompt prefix embeddings (20 prompt tokens) alongside multi-token label embeddings severely underperforms PERFECT, yielding an average accuracy of 67.1% on single-sentence tasks and 58.5% on sentence-pair tasks. This underperformance stems from high learning rate sensitivity and the limited parameter capacity of prefix-only interventions.
    3. Bias Tuning Ablation (bitfit+mte): Tuning only bias parameters and label embeddings achieves 81.2% (single-sentence) and 68.7% (sentence-pair), falling short of adapter-based adaptation in average accuracy and stability.
    4. Full PLM Tuning w/o Adapters (-Adapters): Updating all PLM parameters while retaining label embeddings drops the 12-dataset overall average from 75.6% to 74.7% and drops worst-case performance from 68.1% to 66.1%, confirming that freezing the base PLM and training only bottleneck adapters prevents overfitting in low-resource settings.
  8. Knowl 8 — Effect of Label Embedding Initialization Variance and Handcrafted Verbalizers

    empirical result

    In the PERFECT framework, initializing label embeddings randomly from a low-variance Gaussian distribution outperforms both larger random variances and initializing from PLM token embeddings of handcrafted verbalizers:

    1. Random vs. Engineered Verbalizer Initialization: Initializing label embeddings with token embeddings of handcrafted verbalizers (PERFECT-init) achieves an average accuracy of 81.1% on single-sentence tasks and 68.4% on sentence-pair tasks, which is consistently worse than random initialization (PERFECT-rand at 81.6% and 69.6%).
    2. Variance Grid Search: Evaluating the label embedding initialization L∼N(0,σ)L \sim \mathcal{N}(0, \sigma) over σ∈{10−2,10−3,10−4,10−5}\sigma \in \{10^{-2}, 10^{-3}, 10^{-4}, 10^{-5}\} shows that σ=10−4\sigma = 10^{-4} achieves the highest combined mean and worst-case validation performance across 20 runs (71.8% total average across tasks, compared to 71.6% for σ=10−2\sigma=10^{-2}, 71.0% for σ=10−3\sigma=10^{-3}, and 71.3% for σ=10−5\sigma=10^{-5}).
  9. Knowl 9 — Effect of Mask Token Count on Task Accuracy

    empirical result

    Varying the number of inserted mask tokens M∈{1,2,5,10}M \in \{1, 2, 5, 10\} impacts task accuracy differently depending on dataset characteristics:

    Dataset 1 2 5 10
    CR 90.1 90.2 89.0 87.8
    MR 86.9 86.1 85.4 85.6
    MRPC 67.4 68.2 70.1 72.3
    QNLI 73.7 73.9 73.0 65.1
    RTE 60.0 57.3 56.2 56.0
    TREC 90.0 90.9 88.9 88.8
    Avg 78.0 77.8 77.1 75.9

    For sentiment and question classification (CR, QNLI, TREC), M=2M=2 provides the highest accuracy. For simpler binary classifications (MR, RTE), M=1M=1 is optimal. For complex paraphrase detection (MRPC), increasing mask tokens to M=10M=10 provides a substantial gain (72.3% vs. 67.4% for M=1M=1).

  10. Knowl 10 — Ablation of Training Loss and Prototypical Metric Inference in PERFECT

    empirical result

    An ablation study evaluating design variants across all 12 NLP benchmarks demonstrates the contribution of multi-class hinge loss and prototypical classification:

    Metric / Variant PERFECT -Hinge Loss +Label Emb -Prototypical
    Average Accuracy 75.6 75.1 75.2 75.1
    Worst-Case Accuracy 68.1 67.4 68.4 68.0
    Standard Deviation 3.1 3.3 3.1 3.1
    • -Hinge Loss: Replacing per-token hinge loss with standard multi-class cross-entropy loss reduces overall mean accuracy (75.1% vs 75.6%) and causes a drop in worst-case accuracy (67.4% vs 68.1%), demonstrating that the margin hinge loss stabilizes low-resource training.
    • -Prototypical: Replacing the prototypical inference rule with the training classifier objective decreases average performance from 75.6% to 75.1%.
    • +Label Emb: Using learned label embedding weights directly during inference instead of computing class mean hidden prototypes yields 75.2% average accuracy.

Coverage note — None omitted; all core components, equations, empirical benchmarks, efficiency profiles, hyperparameter choices, and ablation studies from the paper are represented.

References

  1. 1.Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. 2021. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In ACL.
  2. 2.Hiba Asri, Hajar Mousannif, Hassan Al Moatassime, and Thomas Noel. 2016. Using machine learning algorithms for breast cancer risk prediction and diagnosis. Procedia Computer Science.
  3. 3.Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, and Danilo Giampiccolo. 2006. The second pascal recognising textual entailment challenge. Second PASCAL Challenges Workshop on Recognising Textual Entailment.
  4. 4.Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. 2009. The fifth pascal recognizing textual entailment challenge. In TAC.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In NeurIPS.
  6. 6.Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. Tinytl: Reduce memory, not parameters for efficient on-device learning. In NeurIPS.
  7. 7.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop.
  8. 8.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL.
  10. 10.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305.
  11. 11.William B Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In IWP.
  12. 12.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In ACL.
  13. 13.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007. The third PASCAL recognizing textual entailment challenge. In ACL-PASCAL Workshop on Textual Entailment and Paraphrasing.
  14. 14.David Ha, Andrew Dai, and Quoc V. Le. 2017. Hypernetworks. In ICLR.
  15. 15.Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415.
  16. 16.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In ICML.
  17. 17.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In SIGKDD.
  18. 18.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? In TACL.
  19. 19.Teven Le Scao and Alexander M Rush. 2021. How many data points is a prompt worth? In NAACL.
  20. 20.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In EMNLP.
  21. 21.Quentin Lhoest, Albert Villanova del Moral, Patrick von Platen, Thomas Wolf, Mario Šaško, Yacine Jernite, Abhishek Thakur, Lewis Tunstall, Suraj Patil, Mariama Drame, Julien Chaumond, Julien Plu, Joe Davison, Simon Brandeis, Victor Sanh, Teven Le Scao, Kevin Canwen Xu, Nicolas Patry, Steven Liu, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Nathan Raw, Sylvain Lesage, Anton Lozhkov, Matthew Carrigan, Théo Matussière, Leandro von Werra, Lysandre Debut, Stas Bekman, and Clément Delangue. 2021a. huggingface/datasets: 1.15.1.
  22. 22.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021b. Datasets: A community library for natural language processing. In EMNLP.
  23. 23.Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. 2018. Measuring the intrinsic dimension of objective landscapes. In ICLR.
  24. 24.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In ACL.
  25. 25.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  26. 26.Robert L Logan IV, Ivana Balažević, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2021. Cutting down on prompts and parameters: Simple few-shot learning with language models. arXiv preprint arXiv:2106.13353.
  27. 27.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  28. 28.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021a. Compacter: Efficient low-rank hypercomplex adapter layers. In NeurIPS.
  29. 29.Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021b. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. In ACL.
  30. 30.George A Miller. 1995. Wordnet: a lexical database for english. In Communications of the ACM.
  31. 31.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Noisy channel language model prompting for few-shot text classification. arXiv preprint arXiv:2108.04106.
  32. 32.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2022a. Reframing instructional prompts to gptk’s language. In Findings of ACL.
  33. 33.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022b. Cross-task generalization via natural language crowdsourcing instructions. In ACL.
  34. 34.Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2021. On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines. In ICLR.
  35. 35.Bo Pang and Lillian Lee. 2004. A sentimental education: sentiment analysis using subjectivity summarization based on minimum cuts. In ACL.
  36. 36.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL.
  37. 37.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In NeurIPS.
  38. 38.Matthew E Peters, Sebastian Ruder, and Noah A Smith. 2019. To tune or not to tune? adapting pretrained representations to diverse tasks. In RepL4NLP.
  39. 39.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rückle, Cho Kyunghyun, and Iryna Gurevych. 2021. AdapterFusion: Non-destructive task composition for transfer learning. In EACL.
  40. 40.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020. Adapterhub: A framework for adapting transformers. In EMNLP: System Demonstrations.
  41. 41.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. Wic: the word-in-context dataset for evaluating context-sensitive meaning representations. In NAACL.
  42. 42.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. In NAACL.
  43. 43.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  44. 44.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners.
  45. 45.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR.
  46. 46.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. In EMNLP.
  47. 47.Shauli Ravfogel, Elad Ben-Zaken, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked languagemodels. arXiv:2106.10199.
  48. 48.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2018. Efficient parametrization of multidomain deep neural networks. In CVPR.
  49. 49.Timo Schick, Helmut Schmid, and Hinrich Schütze. 2020. Automatically identifying words that can serve as labels for few-shot text classification. In COLING.
  50. 50.Timo Schick and Hinrich Schütze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In EACL.
  51. 51.Timo Schick and Hinrich Schütze. 2021b. It’s not just size that matters: Small language models are also few-shot learners. In NAACL.
  52. 52.Karin Kipper Schuler. 2005. Verbnet: A broad-coverage, comprehensive verb lexicon. PhD Thesis.
  53. 53.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Eliciting knowledge from language models using automatically generated prompts. In EMNLP.
  54. 54.Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. In NeurIPS.
  55. 55.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP.
  56. 56.Derek Tam, Rakesh R Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel. 2021. Improving and simplifying pattern exploiting training. arXiv preprint arXiv:2103.11955.
  57. 57.Wilson L Taylor. 1953. “cloze procedure”: A new tool for measuring readability. Journalism quarterly.
  58. 58.Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020. Udapter: Language adaptation for truly universal dependency parsing. In EMNLP.
  59. 59.Ellen M Voorhees and Dawn M Tice. 2000. Building a question answering test collection. In SIGIR.
  60. 60.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019a. Superglue: a stickier benchmark for general-purpose language understanding systems. In NeurIPS.
  61. 61.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In ICLR.
  62. 62.Albert Webson and Ellie Pavlick. 2021. Do prompt-based models really understand the meaning of their prompts? arXiv preprint arXiv:2109.01247.
  63. 63.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In EMNLP: System Demonstrations.
  64. 64.Aston Zhang, Yi Tay, SHUAI Zhang, Alvin Chan, Anh Tuan Luu, Siu Hui, and Jie Fu. 2021. Beyond fullyconnected layers with quaternions: Parameterization of hypercomplex multiplications with 1/n parameters. In ICLR.
  65. 65.Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2020. Revisiting few-sample bert fine-tuning. In ICLR.
  66. 66.Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. ICML.

Citation

MLA
Mahabadi, R. K., et al. “Prompt-free and Efficient Few-shot Learning with Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3638–52, https://doi.org/10.18653/v1/2022.acl-long.254.
APA
Mahabadi, R. K., Zettlemoyer, L., Henderson, J., Mathias, L., Saeidi, M., Stoyanov, V., & Yazdani, M. (2022). Prompt-free and Efficient Few-shot Learning with Language Models. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3638–3652. https://doi.org/10.18653/v1/2022.acl-long.254
Chicago
Mahabadi, R. K., L. Zettlemoyer, J. Henderson, et al. 2022. “Prompt-free and Efficient Few-shot Learning with Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3638–52. https://doi.org/10.18653/v1/2022.acl-long.254.
Harvard
Mahabadi, R.K. et al. (2022) “Prompt-free and Efficient Few-shot Learning with Language Models”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3638–3652. Available at: https://doi.org/10.18653/v1/2022.acl-long.254.
Vancouver
1. Mahabadi RK, Zettlemoyer L, Henderson J, Mathias L, Saeidi M, Stoyanov V, Yazdani M (2022) Prompt-free and Efficient Few-shot Learning with Language Models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3638–3652

BibTeX

@inproceedings{karimi-mahabadi-etal-2022-prompt,
    title = "Prompt-free and Efficient Few-shot Learning with Language Models",
    author = "Karimi Mahabadi, Rabeeh  and
      Zettlemoyer, Luke  and
      Henderson, James  and
      Mathias, Lambert  and
      Saeidi, Marzieh  and
      Stoyanov, Veselin  and
      Yazdani, Majid",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.254/",
    doi = "10.18653/v1/2022.acl-long.254",
    pages = "3638--3652"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/