Noisy Channel Language Model Prompting for Few-Shot Text Classification
Sewon MinMike LewisHannaneh HajishirziLuke Zettlemoyer
Demonstrates that prompting language models to compute the probability of the input given the label substantially reduces prediction variance and outperforms standard direct prompting methods in few-shot classification, particularly under severe label imbalance and unseen label settings.
Adapting large language models to new text classification tasks using only a few labeled examples commonly relies on direct prompting, where the model predicts the task label given the input text. However, standard direct prompting is notoriously unstable, displaying high variance across different label wordings and training subsets, and frequently collapsing to near-random accuracy in worst-case scenarios. The article addresses this operational reliability challenge by evaluating a "noisy channel" formulation for few-shot language model prompting, which inverts the prediction process by calculating the likelihood of the input text given the label.
The main objective of the article is to demonstrate the effectiveness, stability, and generalization capacity of channel prompting compared to conventional direct prompting across parameter-free demonstration methods and lightweight parameter-tuning strategies. To evaluate these approaches across diverse operational settings, the analysis benchmarks eleven text classification datasets using GPT-2 models across various parameter scales. The evaluation focuses on realistic, data-scarce conditions (primarily 16 training examples) without balanced label assumptions, assessing both average accuracy and worst-case performance across multiple random data splits and verbalizer prompt templates.
The key findings reveal substantial performance and stability advantages for the channel formulation. First, in parameter-efficient prompt tuning, channel prompting achieves a 61.7% macro-average accuracy and 53.0% worst-case accuracy, outperforming direct prompt tuning (48.4% average, 29.5% worst-case) by 13.3 and 23.5 percentage points, respectively. Second, channel prompt tuning exhibits strong resilience against imbalanced training data, maintaining performance where direct models severely degrade. Third, channel models successfully generalize to unseen labels during inference and transfer effectively to related zero-shot classification tasks, whereas direct models consistently fail to predict classes absent during training. Finally, among parameter-tuning baselines, tuning solely the final output layer (head tuning) proves to be a surprisingly strong direct baseline (57.7% average accuracy), surpassing prompt tuning on tasks that diverge substantially from standard language modeling.
These findings indicate that inverting the conditional probability forces the language model to account for every word in the input text, amplifying learning signals when labeled data is scarce and eliminating dangerous worst-case performance drops. For operational deployments, adopting channel prompting significantly mitigates the risk of unpredictable model failures in high-stakes environments, while avoiding the massive compute and storage costs of full-model retraining.
Decision-makers should deploy channel prompt tuning as the preferred method when training data is severely limited (16 or fewer examples), when label distributions are skewed or include many classes, and when systems must accommodate unseen categories over time. Conversely, if training datasets are larger (64 or more examples) or the classification objective diverges fundamentally from generative language modeling, organizations should consider direct head tuning or standard full fine-tuning. Future technical investigations should focus on extending channel prompting beyond classification to generation tasks and integrating channel formulations with masked language models.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). This paper analyzes the severe variance and prediction biases in standard direct few-shot prompting, providing the direct problem motivation that the noisy channel formulation aims to solve.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). This work establishes standard benchmark setups and automated prompt-based fine-tuning baselines for few-shot learning with 16 examples per class that the source study builds upon.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). This work introduces continuous soft prompt tuning for large language models, which the source adapts to its inverted channel formulation.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). This paper introduces P-Tuning continuous prompt representations that serve as the foundation for parameter-efficient prompt tuning evaluated in the source.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). This foundational paper develops Pattern-Exploiting Training (PET) for reformulating text classification into cloze prompt-based language modeling.
- Paper: Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, Yao Lu et al. (2021). This study demonstrates extreme sensitivity to prompt example order in in-context learning, framing the operational instability addressed by channel prompting.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). This paper investigates in-context demonstration selection strategies for few-shot text classification that contextualize the source's evaluation of demonstration formats.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This survey provides a comprehensive taxonomy and formal mathematical foundations for prompt-based learning and prediction paradigms in NLP.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This landmark paper establishes the few-shot in-context learning paradigm with language models that direct and channel prompting build upon.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). This study systematically explores the underlying mechanisms and role of prompt components in in-context learning, dissecting why input-label pairings function differently than standard supervised signals.
- Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). This work explores multi-prompt weak supervision strategies to mitigate prompt sensitivity and stabilize few-shot performance across language models.
- Paper: Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations, Chenglei Si et al. (2023). This research analyzes how language models resolve ambiguities in few-shot demonstrations, extending understanding of inductive biases during in-context adaptation.
- Paper: Prompt-free and Efficient Few-shot Learning with Language Models, Rabeeh Karimi Mahabadi et al. (2022). This paper proposes an alternative prompt-free, adapter-based framework to overcome the brittleness and engineering overhead of standard few-shot prompting.
- Paper: SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer, Tu Vu et al. (2022). This study advances parameter-efficient prompt tuning by demonstrating cross-task soft prompt transfer to enhance few-shot stability and accuracy.
- Paper: Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning, Xiang Chen et al. (2022). This paper combines prompt learning with retrieval mechanisms to prevent overfitting to limited few-shot demonstrations.
- Paper: Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning, Haokun Liu et al. (2022). This work benchmarks parameter-efficient fine-tuning alternatives against few-shot prompting, providing comparative methods for data-scarce adaptation.
