Prompt-free and Efficient Few-shot Learning with Language Models
Rabeeh Karimi MahabadiLuke ZettlemoyerJames HendersonLambert MathiasMarzieh SaeidiVeselin StoyanovMajid Yazdani
Introduces PERFECT, a prompt-free few-shot tuning framework for masked language models that eliminates manual prompt and verbalizer engineering using task-specific adapters and learned label embeddings, achieving up to 100x faster training and inference while outperforming state-of-the-art prompt-based methods.
Adapting large pretrained language models to new tasks using very limited data—known as few-shot learning—typically requires extensive manual engineering. Existing prompt-based methods convert inputs into fill-in-the-blank queries using handcrafted text templates and manually chosen target words. This manual process is labor-intensive, computationally expensive, and brittle, as minor wording changes can significantly harm performance.
The article evaluates a new framework called PERFECT, which eliminates handcrafted prompts and verbalizers to achieve efficient few-shot learning from as few as 32 labeled examples.
The researchers evaluated their method against standard fine-tuning and state-of-the-art prompt-based baselines across 12 standard language understanding benchmarks covering classification, sentiment analysis, and natural language inference. Instead of modifying the entire underlying model or engineering text prompts, the approach freezes the pretrained model, inserts compact task-specific adapter modules, and optimizes multi-token label embeddings alongside a distance-based prototype classification strategy.
The evaluation yielded several critical findings. First, the proposed method achieved state-of-the-art accuracy, outperforming the leading prompt-based baseline by 1.1 percentage points on single-sentence tasks and 4.6 percentage points on sentence-pair tasks, while notably outperforming even the best hand-tuned prompt configurations. Second, it delivered major computational efficiencies, reducing trainable parameters by roughly 99 percent, lowering peak memory usage by 22 percent, speeding up training by 97 percent, and accelerating inference by nearly 97 percent compared to standard prompt baselines. Third, the method substantially improved reliability, raising worst-case performance and reducing output variance across different data samples. Finally, ablation analyses showed that randomly initialized label embeddings outperformed manually engineered target words, confirming that manual verbalizer design is unnecessary.
These findings demonstrate that organizations can deploy high-performing few-shot language models without incurring substantial engineering labor or heavy infrastructure costs. By freezing the underlying base model and training only lightweight adapter layers, teams can drastically lower storage footprints and hardware requirements while stabilizing production performance against prompt sensitivity.
Organizations developing or deploying language models should consider replacing brittle prompt-engineering workflows with parameter-efficient adapter architectures and learned label embeddings. When implementing this architecture, teams should determine the optimal number of mask positions based on task complexity, as empirical results indicate that multi-mask setups provide varying gains across different problem types.
Confidence in these findings is high, given the consistent validation across 12 diverse benchmarks and multiple random initializations. However, leaders should note that the evaluation was primarily conducted using a single base model architecture with 355 million parameters, and results reflect constrained few-shot settings with exactly 16 training and 16 validation examples per class. Additional testing on larger foundation models and domain-specific production data is advised before full-scale deployment.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). It introduces Pattern-Exploiting Training (PET) using cloze questions and verbalizers for few-shot NLP, which serves as the foundational prompt-based baseline that the source directly critiques and replaces with prompt-free adapters.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). It formalizes the standard few-shot evaluation protocol with 16 examples per class and automated prompt/verbalizer search (LM-BFF) that the source directly benchmarks against and seeks to simplify.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). It establishes parameter-efficient adapter modules for pretrained language models, providing the core architectural mechanism that the source adapts for prompt-free few-shot learning.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). It explores continuous prompt embeddings (P-Tuning) to reduce manual prompt sensitivity, establishing the continuum of prompt engineering challenges that motivate the source's prompt-free approach.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). It analyzes the parameter efficiency and limitations of soft prompt tuning across model scales, motivating the need for more efficient adapter-based alternatives in moderate-sized few-shot models.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). It provides a comprehensive taxonomy of prompt-based learning and verbalizer engineering in NLP, framing the design space the source aims to eliminate.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). It details the severe sensitivity and systematic bias of prompt-based few-shot learners, contextualizing the reliability issues solved by the source's prompt-free label embeddings.
- Paper: Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning, Haokun Liu et al. (2022). It extends the paradigm of parameter-efficient few-shot adaptation by proposing T-Few and (IA)^3 across multitask-prompted models as a cost-effective alternative to in-context learning.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). It unifies parameter-efficient transfer methods—including adapters and prompt tuning—into a single design space, deepening the architectural insights underlying the source's adapter framework.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It investigates why demonstration-based few-shot learning succeeds despite incorrect labels, complementing the source's findings that randomly initialized label embeddings outperform manual verbalizers.
- Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). It addresses prompt brittleness from an orthogonal direction by systematically generating open-ended prompt chains and aggregating weak supervision without model fine-tuning.
- Paper: LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models, Yaowei Zheng et al. (2024). It scales and unifies efficient fine-tuning techniques across hundreds of foundation models into an open-source framework for practical deployment.
