Noisy Channel Language Model Prompting for Few-Shot Text Classification

Sewon MinMike LewisHannaneh HajishirziLuke Zettlemoyer

article2022ACL241 citations

Demonstrates that prompting language models to compute the probability of the input given the label substantially reduces prediction variance and outperforms standard direct prompting methods in few-shot classification, particularly under severe label imbalance and unseen label settings.

Listen

Adapting large language models to new text classification tasks using only a few labeled examples commonly relies on direct prompting, where the model predicts the task label given the input text. However, standard direct prompting is notoriously unstable, displaying high variance across different label wordings and training subsets, and frequently collapsing to near-random accuracy in worst-case scenarios. The article addresses this operational reliability challenge by evaluating a "noisy channel" formulation for few-shot language model prompting, which inverts the prediction process by calculating the likelihood of the input text given the label.

The main objective of the article is to demonstrate the effectiveness, stability, and generalization capacity of channel prompting compared to conventional direct prompting across parameter-free demonstration methods and lightweight parameter-tuning strategies. To evaluate these approaches across diverse operational settings, the analysis benchmarks eleven text classification datasets using GPT-2 models across various parameter scales. The evaluation focuses on realistic, data-scarce conditions (primarily 16 training examples) without balanced label assumptions, assessing both average accuracy and worst-case performance across multiple random data splits and verbalizer prompt templates.

The key findings reveal substantial performance and stability advantages for the channel formulation. First, in parameter-efficient prompt tuning, channel prompting achieves a 61.7% macro-average accuracy and 53.0% worst-case accuracy, outperforming direct prompt tuning (48.4% average, 29.5% worst-case) by 13.3 and 23.5 percentage points, respectively. Second, channel prompt tuning exhibits strong resilience against imbalanced training data, maintaining performance where direct models severely degrade. Third, channel models successfully generalize to unseen labels during inference and transfer effectively to related zero-shot classification tasks, whereas direct models consistently fail to predict classes absent during training. Finally, among parameter-tuning baselines, tuning solely the final output layer (head tuning) proves to be a surprisingly strong direct baseline (57.7% average accuracy), surpassing prompt tuning on tasks that diverge substantially from standard language modeling.

These findings indicate that inverting the conditional probability forces the language model to account for every word in the input text, amplifying learning signals when labeled data is scarce and eliminating dangerous worst-case performance drops. For operational deployments, adopting channel prompting significantly mitigates the risk of unpredictable model failures in high-stakes environments, while avoiding the massive compute and storage costs of full-model retraining.

Decision-makers should deploy channel prompt tuning as the preferred method when training data is severely limited (16 or fewer examples), when label distributions are skewed or include many classes, and when systems must accommodate unseen categories over time. Conversely, if training datasets are larger (64 or more examples) or the classification objective diverges fundamentally from generative language modeling, organizations should consider direct head tuning or standard full fine-tuning. Future technical investigations should focus on extending channel prompting beyond classification to generation tasks and integrating channel formulations with masked language models.

arXiv: 2108.04106
Cover for Noisy Channel Language Model Prompting for Few-Shot Text Classification

Abstract

We introduce a noisy channel approach for language model prompting in few-shot text classification. Instead of computing the likelihood of the label given the input (referred as direct models), channel models compute the conditional probability of the input given the label, and are thereby required to explain every word in the input. We use channel models for recently proposed few-shot learning methods with no or very limited updates to the language model parameters, via either in-context demonstration or prompt tuning. Our experiments show that, for both methods, channel models significantly outperform their direct counterparts, which we attribute to their stability, i.e., lower variance and higher worst-case accuracy. We also present extensive ablations that provide recommendations for when to use channel prompt tuning instead of other competitive methods (e.g., direct head tuning): channel prompt tuning is preferred when the number of training examples is small, labels in the training data are imbalanced, or generalization to unseen labels is required.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Channel Model
  • 2.2 Few-shot Learning
  • 3 Formulation
  • 4 Method
  • 4.1 Demonstration methods
  • 4.1.1 Zero-shot
  • 4.1.2 Concat-based demonstrations
  • 4.1.3 Ensemble-based demonstrations
  • 4.2 Tuning methods
  • 4.2.1 Head tuning
  • 4.2.2 Transformation tuning
  • 4.2.3 Prompt tuning
  • 5 Experimental Setup
  • 5.1 Datasets
  • 5.2 Training Data
  • 5.3 Language Models
  • 5.4 Evaluation
  • 6 Experimental Results
  • 6.1 Main Results: Demonstration Methods
  • 6.2 Main Results: Tuning Methods
  • 6.3 Ablations
  • 6.4 Generalization to unseen labels
  • 7 Discussion & Conclusion
  • Acknowledgements
  • References
  • A Samples & Verbalizers
  • B Implementation Details
  • C Additional Results

Knowls

  1. Knowl 1 — Noisy Channel Formulation for Prompt-Based Text Classification

    model/method

    In prompt-based few-shot text classification, a task function f:X→Cf: \mathcal{X} \to \mathcal{C} maps an input text x∈Xx \in \mathcal{X} to a class label ci∈C={c1,…,cm}c_i \in \mathcal{C} = \{c_1, \dots, c_m\}. A pre-defined verbalizer v:C→Xv: \mathcal{C} \to \mathcal{X} maps each class label to a natural language expression (for example, v(c+)="It was great"v(c^+) = \text{"It was great"} and v(c−)="It was terrible"v(c^-) = \text{"It was terrible"}). A causal language model PLMP_{\text{LM}} provides autoregressive sequence likelihoods.

    Instead of the standard direct model which estimates the posterior distribution of the label verbalizer given the input: P(ci∣x)=PLM(v(ci)∣x)=∏t′=1tvPLM(v(ci)t′∣x,v(ci)<t′)P(c_i \mid x) = P_{\text{LM}}(v(c_i) \mid x) = \prod_{t'=1}^{t_v} P_{\text{LM}}(v(c_i)_{t'} \mid x, v(c_i)_{<t'})

    the noisy channel model applies Bayes' rule to compute the probability of the input text conditioned on the label verbalizer: P(ci∣x)=P(x∣ci)P(ci)P(x)∝P(x∣ci)P(ci)P(c_i \mid x) = \frac{P(x \mid c_i) P(c_i)}{P(x)} \propto P(x \mid c_i) P(c_i)

    Assuming a uniform prior distribution over class labels P(ci)=1∣C∣P(c_i) = \frac{1}{|\mathcal{C}|}, classification decisions are made by scoring: c^=arg⁡max⁡ci∈CPLM(x∣v(ci))=arg⁡max⁡ci∈C∏t=1txPLM(xt∣v(ci),x<t)\hat{c} = \arg\max_{c_i \in \mathcal{C}} P_{\text{LM}}(x \mid v(c_i)) = \arg\max_{c_i \in \mathcal{C}} \prod_{t=1}^{t_x} P_{\text{LM}}(x_t \mid v(c_i), x_{<t})

    where length normalization is applied over the tokens of xx. Because the channel model requires the language model to explain every token in the input text xx conditioned on the label verbalizer v(ci)v(c_i), it provides richer training feedback in low-data regimes.

  2. Knowl 2 — Ensemble-Based In-Context Demonstrations for Few-Shot Prompting

    model/method

    Given a few-shot training set D={(x1,c1),…,(xK,cK)}\mathcal{D} = \{(x^1, c^1), \dots, (x^K, c^K)\} of KK labeled examples and a verbalizer vv, conventional concat-based prompting prepends all KK concatenated examples to the input xx. Concat-based prompting incurs O(K2)O(K^2) sequence memory scaling and is highly sensitive to the ordering of examples.

    The ensemble-based demonstration method evaluates the test input xx by conditioning the frozen causal language model PLMP_{\text{LM}} on one training example at a time and multiplying the resulting conditional probabilities across all KK individual examples:

    1. For the direct model: P(ci∣x)=∏j=1KPLM(v(ci)∣xj,v(cj),x)P(c_i \mid x) = \prod_{j=1}^K P_{\text{LM}}(v(c_i) \mid x^j, v(c^j), x)

    2. For the calibrated direct++ model (normalizing by the model likelihood given an empty input string NULL\text{NULL}): P(ci∣x)=∏j=1KPLM(v(ci)∣xj,v(cj),x)PLM(v(ci)∣xj,v(cj),NULL)P(c_i \mid x) = \prod_{j=1}^K \frac{P_{\text{LM}}(v(c_i) \mid x^j, v(c^j), x)}{P_{\text{LM}}(v(c_i) \mid x^j, v(c^j), \text{NULL})}

    3. For the channel model: P(x∣ci)=∏j=1KPLM(x∣v(cj),xj,v(ci))P(x \mid c_i) = \prod_{j=1}^K P_{\text{LM}}(x \mid v(c^j), x^j, v(c_i))

    This ensemble approach reduces memory complexity to O(K)O(K) and eliminates variance arising from demonstration permutation order.

  3. Knowl 3 — Parameter-Efficient Tuning Formulations: Direct Head Tuning vs Channel Prompt Tuning

    model/method

    When adapting a frozen causal language model to few-shot text classification with minimal parameter updates (<0.01%< 0.01\% trainable parameters), four tuning formulations are defined for an LM with hidden representation hx∈Rhh_x \in \mathbb{R}^h and linear prediction head O∈R∣V∣×hO \in \mathbb{R}^{|V| \times h}:

    1. Direct Head Tuning: The language model backbone is frozen and only the output projection matrix O∈R∣V∣×hO \in \mathbb{R}^{|V| \times h} is updated. The probability of verbalizer token vi∈Vv_i \in V given input xx is computed as Softmax(Ohx)i\text{Softmax}(O h_x)_i. During head tuning, OO is decoupled from the input token embedding matrix.
    2. Direct Transformation Tuning: The language model head OO is kept frozen, and a trainable transformation matrix U∈Rh×hU \in \mathbb{R}^{h \times h}, initialized to the identity matrix, is learned such that output logits are Softmax(OUhx)i\text{Softmax}(O U h_x)_i.
    3. Direct Prompt Tuning: The language model parameters and head are frozen, and nn continuous prompt embeddings u1,…,un∈Rhu_1, \dots, u_n \in \mathbb{R}^h prepended to the input are optimized to maximize P(ci∣x)=PLM(v(ci)∣u1,…,un,x)P(c_i \mid x) = P_{\text{LM}}(v(c_i) \mid u_1, \dots, u_n, x).
    4. Channel Prompt Tuning: Continuous prompt embeddings u1,…,unu_1, \dots, u_n are prepended to the label verbalizer v(ci)v(c_i), and the model is trained autoregressively to generate the input text xx: P(x∣ci)=PLM(x∣u1,…,un,v(ci))=∏t=1txPLM(xt∣u1,…,un,v(ci),x<t)P(x \mid c_i) = P_{\text{LM}}(x \mid u_1, \dots, u_n, v(c_i)) = \prod_{t=1}^{t_x} P_{\text{LM}}(x_t \mid u_1, \dots, u_n, v(c_i), x_{<t})
  4. Knowl 4 — Demonstration Methods Performance Across Eleven Text Classification Datasets

    empirical result

    Evaluation of zero-shot, concatenation-based (K=16K=16), and ensemble-based (K=16K=16) demonstration methods using GPT-2 Large across eleven classification datasets (SST-2, SST-5, MR, CR, Amazon, Yelp, TREC, AGNews, Yahoo, DBPedia, Subj) over 4 verbalizers and 5 data seeds demonstrates the advantage of channel prompting.

    Dataset Zero-shot Concat-based (K=16K=16) Ensemble-based (K=16K=16)
    Direct++ Channel Direct++ Channel Direct++ Channel
    SST-2 80.3 / 76.9 77.1 / 74.8 66.8 / 51.7 85.0 / 83.1 79.7 / 68.0 77.5 / 59.5
    SST-5 33.3 / 28.8 29.2 / 27.7 23.7 / 14.4 36.2 / 32.7 33.8 / 23.3 33.6 / 30.2
    MR 77.4 / 73.2 74.3 / 69.3 60.2 / 50.5 80.5 / 76.8 76.8 / 60.1 76.1 / 60.0
    CR 77.9 / 69.7 65.8 / 60.2 66.8 / 50.0 80.8 / 74.8 72.8 / 54.6 79.7 / 69.3
    Amazon 37.6 / 35.0 37.1 / 31.6 40.8 / 35.7 39.4 / 34.3 39.8 / 32.0 40.4 / 36.2
    Yelp 36.8 / 31.8 38.0 / 31.9 38.5 / 31.6 39.8 / 36.5 39.2 / 29.6 41.5 / 38.5
    AGNews 59.9 / 44.0 61.8 / 59.7 51.2 / 34.4 68.5 / 60.6 73.1 / 58.6 74.3 / 69.3
    TREC 27.7 / 12.6 30.5 / 19.4 31.6 / 13.0 42.0 / 26.8 22.9 / 9.8 31.5 / 23.8
    Yahoo 35.3 / 28.7 48.7 / 48.1 29.6 / 19.4 56.2 / 52.3 50.6 / 46.5 58.6 / 57.4
    DBPedia 37.6 / 30.4 51.4 / 42.7 71.1 / 55.2 58.5 / 40.0 72.6 / 55.7 64.8 / 57.0
    Subj 52.0 / 48.8 57.8 / 51.5 56.9 / 50.0 60.5 / 40.8 52.2 / 41.8 52.4 / 46.9
    Macro Avg 50.5 / 43.6 52.0 / 47.0 48.8 / 36.9 58.9 / 50.8 55.8 / 43.6 57.3 / 49.8

    Entries report Average accuracy / Worst-case accuracy (in %). Key empirical findings:

    1. Concat-based channel demonstrations achieve an average accuracy of 58.9%58.9\% and worst-case accuracy of 50.8%50.8\%, substantially exceeding concat-based Direct++ (48.8%/36.9%48.8\% / 36.9\%) and standard direct demonstrations (38.5%/29.1%38.5\% / 29.1\%).
    2. For direct models, ensemble demonstrations significantly outperform concatenation (55.8%55.8\% vs. 48.8%48.8\% average accuracy, and 43.6%43.6\% vs. 36.9%36.9\% worst-case accuracy).
    3. Across all few-shot demonstration setups, the best channel model outperforms the best direct model by 3.1%3.1\% in average accuracy and 7.2%7.2\% in worst-case accuracy.
  5. Knowl 5 — Parameter-Efficient Tuning Performance Across Eleven Text Classification Tasks

    empirical result

    Evaluating tuning methods on K=16K=16 training examples with GPT-2 Large across 11 benchmark datasets (over 4 verbalizers, 5 data seeds, and 4 training seeds, yielding 80 runs per method) shows that channel prompt tuning consistently surpasses direct prompt tuning.

    Dataset Direct Head Direct Trans Direct Prompt Channel Prompt
    SST-2 80.2 / 68.6 77.3 / 57.5 72.6 / 50.9 85.8 / 81.3
    SST-5 34.9 / 30.0 33.0 / 25.5 30.9 / 19.1 36.3 / 27.9
    MR 73.7 / 56.4 71.3 / 51.6 67.4 / 50.1 81.7 / 78.0
    CR 67.6 / 50.0 63.9 / 50.0 65.7 / 50.0 79.6 / 76.4
    Amazon 34.5 / 28.8 32.1 / 18.2 31.2 / 20.0 43.4 / 39.2
    Yelp 40.6 / 32.8 38.9 / 31.5 31.9 / 20.6 43.9 / 37.2
    TREC 54.1 / 42.4 48.0 / 31.0 35.9 / 13.0 37.1 / 20.8
    AGNews 74.1 / 61.2 66.9 / 47.0 61.9 / 25.2 73.4 / 63.9
    Yahoo 39.1 / 31.4 33.8 / 23.0 27.4 / 15.7 54.0 / 46.7
    DBPedia 49.3 / 37.5 42.4 / 28.6 41.8 / 9.9 67.7 / 52.9
    Subj 86.3 / 79.1 86.0 / 71.6 65.5 / 49.9 75.5 / 58.8
    Macro Avg 57.7 / 47.1 54.0 / 39.6 48.4 / 29.5 61.7 / 53.0

    Entries report Average accuracy / Worst-case accuracy in %. Key takeaways:

    1. Channel Prompt Tuning vs Direct Prompt Tuning: Channel prompt tuning provides an average accuracy gain of 13.3%13.3\% absolute (61.7%61.7\% vs. 48.4%48.4\%) and a worst-case accuracy gain of 23.5%23.5\% absolute (53.0%53.0\% vs. 29.5%29.5\%), driven by dramatically lower variance across prompt verbalizers and random seeds.
    2. Strength of Direct Head Tuning: Direct head tuning achieves 57.7%/47.1%57.7\% / 47.1\%, outperforming direct prompt tuning on every dataset. It exceeds channel prompt tuning on tasks that diverge most from natural language modeling, specifically TREC question type classification (54.1%54.1\% vs. 37.1%37.1\%) and Subj subjectivity classification (86.3%86.3\% vs. 75.5%75.5\%).
    3. Scaling with Number of Classes: On tasks with large label spaces (∣C∣≥10|\mathcal{C}| \ge 10), channel prompt tuning outperforms direct head tuning by large margins: Yahoo (54.0%54.0\% vs. 39.1%39.1\%) and DBPedia (67.7%67.7\% vs. 49.3%49.3\%). On these datasets, channel prompt tuning also outperforms full LM finetuning (48.9%48.9\% on Yahoo and 66.3%66.3\% on DBPedia).
  6. Knowl 6 — Robustness of Channel Prompting to Training Label Imbalance

    empirical result

    When class distributions in few-shot training sets are imbalanced, direct prompting methods degrade severely, whereas channel prompt tuning is largely invariant.

    Let p−p^- denote the ratio of negative examples in binary classification training sets (SST-2 and MR) with size K∈{16,64}K \in \{16, 64\}, evaluated across p−∈{0.0,0.125,0.25,0.375,0.5}p^- \in \{0.0, 0.125, 0.25, 0.375, 0.5\} (where p−=0.5p^- = 0.5 is perfectly balanced and p−=0p^- = 0 contains only positive instances):

    1. Sensitivity of Direct Models: As p−p^- approaches 0, direct prompt tuning, direct head tuning, and full finetuning drop to majority-class baseline performance (approx50%\\approx 50\%) because the output head relies excessively on unconditional label marginals observed in training. Upsampling minority-class examples provides only partial mitigation.
    2. Invariance of Channel Models: Channel prompt tuning performance remains robust across all p−p^- ratios, significantly outperforming all direct tuning baselines and full LM finetuning when p−<0.25p^- < 0.25.
    3. Underlying Reason: In channel models, the label verbalizer v(ci)v(c_i) acts solely as a conditioning prefix rather than a prediction target, isolating sequence generation scores P(x∣v(ci))P(x \mid v(c_i)) from empirical class label skews in D\mathcal{D}.
  7. Knowl 7 — Generalization of Channel Prompting to Unseen Labels and Cross-Task Transfer

    empirical result

    Channel prompt tuning generalizes to class labels that are completely absent from the training set, whereas direct models fail entirely.

    When GPT-2 Large is trained on K=16K=16 examples sampled such that one class label is omitted during training:

    Dataset Zero-Shot Finetuning (K=16K=16, one class excluded)
    Direct++ Channel Direct All Direct Head Direct Prompt Channel Prompt
    SST-2 80.3 / 76.9 77.1 / 74.8 50.2 / 49.1 50.2 / 49.1 50.2 / 49.1 85.5 / 82.5
    SST-5 33.3 / 28.8 29.2 / 27.7 40.1 / 34.8 34.3 / 28.0 30.0 / 18.1 37.5 / 32.6
    MR 77.4 / 73.2 74.3 / 69.3 50.0 / 50.0 50.0 / 50.0 50.0 / 50.0 80.9 / 74.8
    CR 77.9 / 69.7 65.8 / 60.2 50.0 / 50.0 50.0 / 50.0 50.0 / 50.0 80.9 / 74.8
    TREC 27.7 / 12.6 30.5 / 19.4 50.8 / 31.0 44.8 / 29.6 33.9 / 17.4 34.3 / 26.0
    Subj 52.0 / 48.8 57.8 / 51.5 50.0 / 50.0 50.0 / 50.0 50.0 / 50.0 66.6 / 57.6

    Accuracies are reported as Average / Worst-case in % across 20 runs.

    Key behaviors:

    1. Direct Model Failure: Direct models (Direct Prompt, Direct Head, Direct Trans, Direct All) assign zero effective probability mass to unseen verbalizer tokens, collapsing to chance accuracy (50.0%50.0\% on binary datasets).
    2. Channel Model Retention: Channel prompt tuning scores unseen classes via conditional generation P(x∣v(cunseen))P(x \mid v(c_{\text{unseen}})), attaining 85.5%85.5\% on SST-2, 80.9%80.9\% on MR, and 80.9%80.9\% on CR—outperforming zero-shot baselines.
    3. Cross-Task Zero-Shot Transfer: When prompt embeddings are tuned on one dataset and evaluated on a different task with a distinct label space, channel prompt tuning achieves effective transfer on related tasks (such as 2-way sentiment to 5-way sentiment), while direct head tuning cannot be evaluated due to disjoint label heads.
  8. Knowl 8 — Scaling Characteristics with Training Set Size K

    empirical result

    The relative performance between channel prompt tuning and direct tuning methods varies systematically with the training sample size K∈{4,16,64,Full}K \in \{4, 16, 64, \text{Full}\}:

    1. Low-Data Regime (K≤16K \le 16): Channel prompt tuning outperforms direct head tuning and direct prompt tuning across benchmarks (SST-2, MR, AGNews). Requiring the language model to generate every token of the input conditioned on the label provides denser supervisory signals when few training instances are available.
    2. Intermediate Regime (K=64K = 64): Direct head tuning overtakes channel prompt tuning as training data becomes sufficient to calibrate discriminative output representations.
    3. Full-Data Regime (K=FullK = \text{Full}): Direct prompt tuning and direct head tuning outperform channel prompt tuning. In full data, direct prompt tuning matches full model parameter finetuning, and generative scoring of xx no longer provides an advantage over discriminative modeling of P(ci∣x)P(c_i \mid x).
  9. Knowl 9 — Operational Limitations of Noisy Channel Prompting

    limitation

    Noisy channel language model prompting has two primary limitations:

    1. Inapplicability to Open-Ended Generation Tasks: Channel classification relies on computing arg⁡max⁡ci∈CP(x∣v(ci))P(ci)\arg\max_{c_i \in \mathcal{C}} P(x \mid v(c_i)) P(c_i) by enumerating all discrete labels in C\mathcal{C}. For open-ended generative tasks (e.g., summarization or free-form question answering), the label space is infinite, requiring an explicit input prior P(y)P(y) and complex beam search decoding over P(x∣y)P(y)P(x \mid y) P(y).
    2. Incompatibility with Masked Language Models: The channel approach requires computing the autoregressive conditional sequence probability of multi-token input text xx given a label verbalizer prefix v(ci)v(c_i). While causal language models execute this natively, masked bidirectional language models (such as BERT or RoBERTa) cannot autoregressively generate extended text sequences without updating the core model parameters.

Coverage note — Omitted only the auxiliary hyperparameter search grids (Tables 7, 8, 9) and raw dataset text samples (Table 10) as standard implementation details.

References

  1. 1.Hiba Asri, H. Mousannif, H. A. Moatassime, and Thomas Noël. 2016. Using machine learning algorithms for breast cancer risk prediction and diagnosis. In ANT/SEIT.
  2. 2.Trapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, and Andrew McCallum. 2020. Self-supervised meta-learning for few-shot natural language classification tasks. In EMNLP.
  3. 3.Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, and Robert L Mercer. 1993. The mathematics of statistical machine translation: Parameter estimation. Computational linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In NeurIPS.
  5. 5.Jiaao Chen, Zichao Yang, and Diyi Yang. 2020. MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification. In ACL.
  6. 6.Kevin Clark, Minh-Thang Luong, Christopher D Manning, and Quoc V Le. 2018. Semi-supervised sequence modeling with cross-view training. In EMNLP.
  7. 7.Xiaoan Ding and Kevin Gimpel. 2019. Latent-variable generative models for data-efficient text classification. In EMNLP.
  8. 8.Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML.
  9. 9.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In ACL.
  10. 10.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In ICML.
  11. 11.Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn't always right. In EMNLP.
  12. 12.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In ICML.
  13. 13.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.
  14. 14.Po-Sen Huang, Chenglong Wang, Rishabh Singh, Wentau Yih, and Xiaodong He. 2018. Natural language to structured query generation via meta-learning. In NAACL.
  15. 15.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? TACL.
  16. 16.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR.
  17. 17.Philipp Koehn, Franz J Och, and Daniel Marcu. 2003. Statistical phrase-based translation. In NAACL-HLT.
  18. 18.Teven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? In NAACL-HLT.
  19. 19.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al. 2015. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web.
  20. 20.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In EMNLP.
  21. 21.Mike Lewis and Angela Fan. 2018. Generative question answering: Learning to answer the whole question. In ICLR.
  22. 22.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In ACL.
  23. 23.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021. Gpt understands, too. arXiv preprint arXiv:2103.10385.
  24. 24.Robert L Logan IV, Ivana Balažević, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2021. Cutting down on prompts and parameters: Simple few-shot learning with language models. arXiv preprint arXiv:2106.13353.
  25. 25.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  26. 26.Julian McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: understanding rating dimensions with review text. In Proceedings of the 7th ACM conference on Recommender systems, pages 165–172.
  27. 27.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. 2017. Adversarial training methods for semi-supervised text classification. In ICLR.
  28. 28.Andrew Y Ng and Michael I Jordan. 2002. On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. In NeurIPS.
  29. 29.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In ACL.
  30. 30.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL.
  31. 31.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS.
  32. 32.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In NeurIPS.
  33. 33.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. In NAACL-HLT.
  34. 34.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  35. 35.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning multiple visual domains with residual adapters. In NeurIPS.
  36. 36.Claude Elwood Shannon. 1948. A mathematical theory of communication. The Bell system technical journal.
  37. 37.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In EMNLP.
  38. 38.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP.
  39. 39.Derek Tam, Rakesh R Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel. 2021. Improving and simplifying pattern exploiting training. In EMNLP.
  40. 40.Ellen M Voorhees and Dawn M Tice. 2000. Building a question answering test collection. In SIGIR.
  41. 41.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In EMNLP: System Demonstrations.
  42. 42.Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. 2020. Unsupervised data augmentation for consistency training. In NeurIPS.
  43. 43.Kenji Yamada and Kevin Knight. 2001. A syntax-based statistical translation model. In ACL.
  44. 44.Kyra Yee, Nathan Ng, Yann N Dauphin, and Michael Auli. 2019. Simple and effective noisy channel modeling for neural machine translation. In EMNLP.
  45. 45.Dani Yogatama, Chris Dyer, Wang Ling, and Phil Blunsom. 2017. Generative and discriminative text classification with recurrent neural networks. arXiv preprint arXiv:1703.01898.
  46. 46.Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomas Kocisky. 2017. The neural noisy channel. In ICLR.
  47. 47.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NeurIPS.
  48. 48.Tony Z Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In ICML.
  49. 49.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. In NAACL-HLT.

Citation

MLA
Min, S., et al. “Noisy Channel Language Model Prompting for Few-Shot Text Classification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 5316–30, https://doi.org/10.18653/v1/2022.acl-long.365.
APA
Min, S., Lewis, M., Hajishirzi, H., & Zettlemoyer, L. (2022). Noisy Channel Language Model Prompting for Few-Shot Text Classification. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5316–5330. https://doi.org/10.18653/v1/2022.acl-long.365
Chicago
Min, S., M. Lewis, H. Hajishirzi, and L. Zettlemoyer. 2022. “Noisy Channel Language Model Prompting for Few-Shot Text Classification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5316–30. https://doi.org/10.18653/v1/2022.acl-long.365.
Harvard
Min, S. et al. (2022) “Noisy Channel Language Model Prompting for Few-Shot Text Classification”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5316–5330. Available at: https://doi.org/10.18653/v1/2022.acl-long.365.
Vancouver
1. Min S, Lewis M, Hajishirzi H, Zettlemoyer L (2022) Noisy Channel Language Model Prompting for Few-Shot Text Classification. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5316–5330

BibTeX

@inproceedings{min-etal-2022-noisy,
    title = "Noisy Channel Language Model Prompting for Few-Shot Text Classification",
    author = "Min, Sewon  and
      Lewis, Mike  and
      Hajishirzi, Hannaneh  and
      Zettlemoyer, Luke",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.365/",
    doi = "10.18653/v1/2022.acl-long.365",
    pages = "5316--5330"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/