P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks

Xiao LiuKaixuan JiYicheng FuWeng TamZhengxiao DuZhilin YangJie Tang

article2022ACL1,172 citations

Presents P-Tuning v2, an optimized deep prompt tuning method that matches full-model fine-tuning performance across diverse model scales and natural language understanding tasks while updating only 0.1% to 3% of parameters.

Listen

Deploying large artificial intelligence language models across diverse language understanding tasks typically requires full fine-tuning, a process that updates and stores an entire set of model parameters for every separate task. This conventional approach demands massive computing memory during training and creates steep storage and operational costs in production. While earlier prompt tuning techniques attempted to mitigate this by freezing the core model and tuning only a small set of input prompts, they performed poorly on widely used, medium-sized models and failed on complex sequence labeling tasks like question answering. The article demonstrates an optimized technique, called P-Tuning v2, designed to make prompt tuning universally effective across all model sizes and diverse natural language tasks.

The authors conducted extensive empirical evaluations comparing their method against traditional fine-tuning and previous prompt tuning approaches. The evaluation covered bidirectional language models ranging from 300 million to 10 billion parameters across standard general benchmarks and challenging sequence tagging workloads, such as named entity recognition and extractive question answering. Rather than attaching trainable prompt variables only to the input layer, the evaluated method injects continuous prompts across every layer of the frozen model and pairs them with standard classification heads. The results show that this deep prompt approach matches the accuracy of full fine-tuning across all tested model scales while tuning only 0.1% to 3% of the parameters per task. On demanding sequence labeling tasks where prior prompt tuning collapsed to near-zero utility, the method achieved performance virtually identical to full fine-tuning.

These results show that organizations can drastically reduce training memory and task-specific storage costs without compromising predictive accuracy. Because the underlying model remains completely frozen, multiple distinct applications can share a single model instance in memory, substantially cutting infrastructure expenses and simplifying deployment pipelines. Organizations seeking parameter-efficient adaptation should adopt deep prompt tuning as a direct alternative to full fine-tuning, especially when managing multiple specialized applications. Practitioners should note that optimal prompt lengths and mathematical adjustments vary depending on task complexity, requiring modest configuration adjustments. Overall, the findings demonstrate high reliability across supervised benchmarks, though teams should conduct localized pilot tests when operating outside standard supervised datasets.

Cover for P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks

Abstract

Prompt tuning, which only tunes continuous prompts with a frozen language model, substantially reduces per-task storage and memory usage at training. However, in the context of NLU, prior work reveals that prompt tuning does not perform well for normal-sized pretrained models. We also find that existing methods of prompt tuning cannot handle hard sequence labeling tasks, indicating a lack of universality. We present a novel empirical finding that properly optimized prompt tuning can be universally effective across a wide range of model scales and NLU tasks. It matches the performance of finetuning while having only 0.1%-3% tuned parameters. Our method P-Tuning v2 is an implementation of Deep Prompt Tuning (Li and Liang, 2021; Qin and Eisner, 2021) optimized and adapted for NLU. Given the universality and simplicity of P-Tuning v2, we believe it can serve as an alternative to finetuning and a strong baseline for future research.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 P-Tuning v2
  • 3.1 Lack of Universality
  • 3.2 Deep Prompt Tuning
  • 3.3 Optimization and Implementation
  • 4 Experiments
  • 4.1 P-tuning v2: Across Scales
  • 4.2 P-tuning v2: Across Tasks
  • 4.3 Ablation Study
  • 5 Conclusions
  • ACKNOWLEDGEMENT
  • References
  • A Problem Formulation on Sequence Tagging
  • B More Ablation Study

Knowls

  1. Knowl 1 — P-Tuning v2 Framework for Natural Language Understanding

    model/method

    P-Tuning v2 is a parameter-efficient tuning framework that adapts deep prompt tuning across all Transformer layers for Natural Language Understanding (NLU) tasks across diverse model scales and task families.

    In standard single-layer prompt tuning, trainable continuous prompt embeddings [h0,…,hi][h_0, \dots, h_i] are inserted solely into the input embedding sequence alongside the input sequence embeddings e(x)e(x):

    [e(x),h0,…,hi,e("[ MASK ]")][e(x), h_0, \dots, h_i, e(\text{"[ MASK ]"})]

    In contrast, P-Tuning v2 introduces independent continuous prompt prefix vectors to the Transformer key and value states (or hidden representations) at every Transformer layer l∈{1,…,L}l \in \{1, \dots, L\} of a pre-trained language model:

    P(l)=[p1(l),p2(l),…,pK(l)]P^{(l)} = [p_1^{(l)}, p_2^{(l)}, \dots, p_K^{(l)}]

    where KK denotes the prompt length and each pk(l)p_k^{(l)} is a trainable continuous embedding vector matching the layer's hidden dimension. During task adaptation, all parameters of the pre-trained Transformer backbone remain completely frozen; only the multi-layer continuous prompt parameters {P(l)}l=1L\{P^{(l)}\}_{l=1}^L and the classification head are optimized. This configuration scales the task-specific tunable parameters to 0.1% to 3%0.1\% \text{ to } 3\% of the backbone parameters (compared to ≈0.01%\approx 0.01\% in single-layer prompt tuning).

    Additionally, P-Tuning v2 replaces the verbalizer language modeling head with a standard randomly initialized linear classification head placed on top of token representations (as in standard BERT classification), enabling seamless adaptation to both sentence-level classification and token-level sequence labeling without manual verbalizer engineering.

  2. Knowl 2 — Comparative Performance Across Model Scales on SuperGLUE Benchmark

    data/table

    On the SuperGLUE development benchmark, single-layer prompt tuning underperforms full fine-tuning substantially on small and medium-sized language models (335M to 2B parameters), matching fine-tuning only at scales of 10B parameters. P-Tuning v2 achieves parity with full fine-tuning across all tested model scales while updating only approximately 0.1% of task-specific parameters.

    The table below presents dev set scores across four model architectures: BERT-large (335M), RoBERTa-large (355M), GLM-xlarge (2B), and GLM-xxlarge (10B). Reported metrics are accuracy for BoolQ, CB, COPA, RTE, WiC, and WSC; F1a for MultiRC; and F1 for ReCoRD. FT denotes full fine-tuning, PT denotes single-layer prompt tuning, and PT-2 denotes P-Tuning v2.

    Model Method BoolQ CB COPA MultiRC ReCoRD RTE WiC WSC
    BERT-large (335M) FT 77.7 94.6 69.0 70.5 70.6 70.4 74.9 68.3
    BERT-large (335M) PT 67.2 80.4 55.0 59.6 44.2 53.5 63.0 64.4
    BERT-large (335M) PT-2 75.8 94.6 73.0 70.6 72.8 78.3 75.1 68.3
    RoBERTa-large (355M) FT 86.9 98.2 94.0 85.7 89.0 86.6 75.6 63.5
    RoBERTa-large (355M) PT 62.3 71.4 63.0 59.9 46.3 58.8 56.9 64.4
    RoBERTa-large (355M) PT-2 84.8 100.0 93.0 82.5 89.3 89.5 73.4 63.5
    GLM-xlarge (2B) FT 88.3 96.4 93.0 84.1 91.8 90.3 74.1 95.2
    GLM-xlarge (2B) PT 79.7 76.4 92.0 77.5 82.7 85.6 71.0 87.5
    GLM-xlarge (2B) PT-2 87.0 96.4 91.0 84.4 91.9 90.3 72.0 92.3
    GLM-xxlarge (10B) FT 88.7 98.7 98.0 88.1 94.4 93.1 75.7 95.2
    GLM-xxlarge (10B) PT 88.8 98.2 98.0 86.1 87.8 89.9 71.8 94.2
    GLM-xxlarge (10B) PT-2 88.8 96.4 98.0 88.1 92.5 93.1 74.0 93.3

    At 335M and 355M scales, standard prompt tuning underperforms fine-tuning by up to 26.4 points on ReCoRD and 27.8 points on RTE, whereas P-Tuning v2 matches or surpasses fine-tuning (e.g., scoring 78.3 vs. 70.4 on RTE with BERT-large).

  3. Knowl 3 — Performance on Sequence Tagging Tasks (NER, Extractive QA, and SRL)

    data/table

    Single-layer prompt tuning experiences severe degradations on token-level sequence labeling tasks, whereas P-Tuning v2 matches full fine-tuning performance across Named Entity Recognition (NER), Extractive Question Answering (QA), and Semantic Role Labeling (SRL).

    The table below reports micro-F1 scores for NER and SRL, and Exact Match (EM) and F1 scores for extractive QA using BERT-large (335M), RoBERTa-large (355M), and DeBERTa-xlarge (750M). Evaluated methods include full fine-tuning (FT), single-layer prompt tuning (PT), single-task P-Tuning v2 (PT-2), and multi-task P-Tuning v2 (MPT-2).

    Task / Dataset Model FT PT PT-2 MPT-2
    NER (F1)
    CoNLL03 BERT-large 92.8 81.9 90.2 91.0
    CoNLL03 RoBERTa-large 92.6 86.1 92.8 92.8
    CoNLL03 DeBERTa-xlarge 93.1 90.2 93.1 93.1
    OntoNotes 5.0 BERT-large 89.2 74.6 86.4 86.3
    OntoNotes 5.0 RoBERTa-large 89.8 80.8 89.8 89.8
    OntoNotes 5.0 DeBERTa-xlarge 90.4 85.1 90.4 90.5
    CoNLL04 BERT-large 85.6 73.6 84.5 86.6
    CoNLL04 RoBERTa-large 88.8 76.2 88.4 90.6
    CoNLL04 DeBERTa-xlarge 89.1 82.4 86.5 90.1
    QA (EM / F1)
    SQuAD 1.1 dev BERT-large 84.2 / 91.1 1.0 / 8.5 77.8 / 86.0 82.3 / 89.6
    SQuAD 1.1 dev RoBERTa-large 88.9 / 94.6 1.2 / 12.0 88.5 / 94.4 88.0 / 94.1
    SQuAD 1.1 dev DeBERTa-xlarge 90.1 / 95.5 2.4 / 19.0 90.4 / 95.7 89.6 / 95.4
    SQuAD 2.0 dev BERT-large 78.7 / 81.9 50.2 / 50.2 69.7 / 73.5 72.7 / 75.9
    SQuAD 2.0 dev RoBERTa-large 86.5 / 89.4 50.2 / 50.2 82.1 / 85.5 83.4 / 86.7
    SQuAD 2.0 dev DeBERTa-xlarge 88.3 / 91.1 50.2 / 50.2 88.4 / 91.1 88.1 / 90.8
    SRL (F1)
    CoNLL12 BERT-large 84.9 64.5 83.2 85.1
    CoNLL12 RoBERTa-large 86.5 67.2 84.6 86.2
    CoNLL12 DeBERTa-xlarge 86.5 74.1 85.7 87.1
    CoNLL05 WSJ BERT-large 88.5 76.0 86.3 88.5
    CoNLL05 WSJ RoBERTa-large 90.2 76.8 89.2 90.0
    CoNLL05 WSJ DeBERTa-xlarge 91.2 82.3 90.6 91.2
    CoNLL05 Brown BERT-large 82.7 70.0 80.7 83.1
    CoNLL05 Brown RoBERTa-large 85.6 70.7 84.3 85.7
    CoNLL05 Brown DeBERTa-xlarge 86.9 77.7 86.3 87.0

    Single-layer prompt tuning drops to near zero on SQuAD 1.1 (1.0% EM) and defaults to constant no-answer predictions on SQuAD 2.0 (50.2% EM / F1), whereas P-Tuning v2 reaches up to 90.4% EM / 95.7% F1. Multi-task learning (MPT-2) provides further consistent gains on NER and SRL.

  4. Knowl 4 — Equivalence of Linear Classification Head and Verbalizer LM Head in Supervised Prompt Tuning

    empirical result

    In fully supervised prompt tuning, employing a standard randomly initialized linear classification head on top of the [CLS] token representation performs on par with or slightly better than using a Verbalizer with a pre-trained language model head, while avoiding the need for manual verbalizer engineering and extending applicability to sequence labeling.

    An ablation comparison on RoBERTa-large across four classification benchmarks (using target tokens "true"/"false" for SST-2, RTE, BoolQ, and "true"/"false"/"neutral" for CB) yielded the following dev set scores:

    Head Formulation SST-2 RTE BoolQ CB
    [CLS] linear head 96.3 88.4 84.8 96.4
    Verbalizer LM head 95.8 86.6 84.6 94.6

    Because tuning a linear head involves only a few thousand parameters and produces no statistically meaningful performance difference compared to verbalizers in full-data settings, P-Tuning v2 adopts standard classification heads directly.

  5. Knowl 5 — Impact of Transformer Layer Depth on Continuous Prompt Tuning

    empirical result

    Continuous prompt parameters placed in deeper Transformer layers have a significantly stronger steering effect on model output representations than prompt parameters placed in shallow input layers.

    In layer-depth ablations using a 24-layer BERT-large backbone on RTE and BoolQ:

    1. Descending vs. Ascending Layer Selection: When adding continuous prompts to a subset of kk layers, selecting layers in descending order from the output downwards (e.g., layers 21–24, 17–24, 13–24) consistently achieves higher accuracy than selecting layers in ascending order from the input upwards (e.g., layers 1–4, 1–8, 1–12).
    2. Near-Full Capacity with Partial Deep Prompts: On RTE, restricting continuous prompts to only the top 8 layers (layers 17–24) recovers performance very close to prompting all 24 layers, whereas prompting only the first 8 layers (layers 1–8) leads to a substantial performance deficit.

    This indicates that prompt embeddings directly integrated into intermediate and upper Transformer layers exert more direct influence on target label prediction than input-level continuous prompts.

  6. Knowl 6 — Task-Dependent Optimal Prompt Length and Reparameterization Dynamics

    empirical result

    The optimal continuous prompt length and the utility of MLP reparameterization in P-Tuning v2 vary significantly based on task complexity:

    1. Optimal Prompt Length Discrepancy: Simple classification tasks (such as RTE and BoolQ) achieve optimal performance with shorter continuous prompts of fewer than 20 tokens (K<20K < 20). In contrast, token-level sequence labeling tasks (such as CoNLL04 NER and CoNLL12 Semantic Role Labeling) require longer prompts, typically around 100 tokens (K≈100K \approx 100), to supply sufficient capacity for token representations.

    2. Inconsistent Reparameterization Benefit: Although MLP reparameterization encoders were considered essential in prior prompt tuning literature, in P-Tuning v2 their effect is inconsistent. On RTE and CoNLL04, MLP reparameterization consistently improves performance over direct embedding optimization across prompt lengths. On BoolQ, MLP and direct embedding are competitive. On CoNLL12 SRL, direct embedding tuning consistently outperforms MLP reparameterization.

    3. Optimization Acceleration: On datasets where MLP reparameterization is beneficial (e.g., RTE, CoNLL04, and BoolQ), it enables the model to reach peak performance with shorter prompt lengths than direct embedding tuning.

  7. Knowl 7 — Multi-Task Shared Prompt Tuning for Sequence Tagging

    model/method

    Multi-task P-Tuning v2 pre-trains shared multi-layer continuous prompts across multiple related datasets within a task domain before fine-tuning on individual tasks, providing a higher-quality prompt initialization.

    1. Dataset Pooling Across Families:

      • Named Entity Recognition (NER): Training datasets of CoNLL03, OntoNotes 5.0, and CoNLL04 are merged for joint pre-training.
      • Extractive Question Answering (QA): SQuAD 1.1 and SQuAD 2.0 training sets are combined, treating all questions as potentially unanswerable via confidence thresholding on start-end span probabilities.
      • Semantic Role Labeling (SRL): CoNLL05, CoNLL12, and the PropBank release training sets are combined. For sentences containing multiple verbs, the target verb token is appended to the end of the input sequence to indicate the predicate being analyzed.
    2. Parameter Sharing Scheme: All pooled datasets share the deep multi-layer continuous prompt parameters {P(l)}l=1L\{P^{(l)}\}_{l=1}^L while utilizing dataset-specific linear classification heads.

    3. Effectiveness: Multi-task pre-training improves downstream micro-F1 scores on NER and SRL across BERT-large, RoBERTa-large, and DeBERTa-xlarge models by up to 2–3 F1 points compared to single-task P-Tuning v2.

Coverage note — None was omitted; all key contributions, benchmarks (SuperGLUE, NER, QA, SRL), ablations (depth, head design, length, reparameterization), and multi-task formulations are covered.

References

  1. 1.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  2. 2.Xavier Carreras and Lluís Màrquez. 2004. Introduction to the CoNLL-2004 shared task: Semantic role labeling. In Proceedings of the Eighth Conference on Computational Natural Language Learning (CoNLL-2004) at HLT-NAACL 2004, pages 89–97, Boston, Massachusetts, USA. Association for Computational Linguistics.
  3. 3.Xavier Carreras and Lluís Màrquez. 2005. Introduction to the CoNLL-2005 shared task: Semantic role labeling. In Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005), pages 152–164, Ann Arbor, Michigan. Association for Computational Linguistics.
  4. 4.Xiang Chen, Xin Xie, Ningyu Zhang, Jiahuan Yan, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2021. Adaprompt: Adaptive prompt-based finetuning for relation extraction. arXiv preprint arXiv:2104.07650.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv e-prints.
  7. 7.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2021. All nlp tasks are generation tasks: A general pretraining framework. arXiv preprint arXiv:2103.10360.
  8. 8.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.
  9. 9.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2021. Ppt: Pre-trained prompt tuning for few-shot learning. arXiv preprint arXiv:2109.04332.
  10. 10.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention.
  11. 11.Boseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee, Donghyun Kwak, Dong Hyeon Jeon, Sunghyun Park, Sungju Kim, Seonhoon Kim, Dongpil Seo, et al. 2021. What changes can large-scale language models bring? intensive study on hyperclova: Billions-scale korean generative pretrained transformers. arXiv preprint arXiv:2109.04650.
  12. 12.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691.
  13. 13.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190.
  14. 14.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021. Gpt understands, too. arXiv:2103.10385.
  15. 15.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv e-prints.
  16. 16.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Noisy channel language model prompting for few-shot text classification. arXiv preprint arXiv:2108.04106.
  17. 17.Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012. CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes. In Joint Conference on EMNLP and CoNLL - Shared Task, pages 1–40, Jeju Island, Korea. Association for Computational Linguistics.
  18. 18.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. arXiv preprint arXiv:2104.06599.
  19. 19.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  20. 20.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  21. 21.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text.
  22. 22.Erik F Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050.
  23. 23.Timo Schick and Hinrich Schütze. 2020. It's not just size that matters: Small language models are also few-shot learners. arXiv preprint arXiv:2009.07118.
  24. 24.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980.
  25. 25.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. In NeurIPS 2019, pages 3261–3275.
  26. 26.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. arXiv e-prints.
  27. 27.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv e-prints.
  28. 28.Hongru Wang, Mingyu Cui, Zimo Zhou, Gabriel Pui Cheong Fung, and Kam-Fai Wong. 2021a. Topicrefine: Joint topic prediction and dialogue response generation for multi-turn end-to-end dialogue system. arXiv preprint arXiv:2109.05187.
  29. 29.Shuo Wang, Zhaopeng Tu, Zhixing Tan, Wenxuan Wang, Maosong Sun, and Yang Liu. 2021b. Language models are good translators. arXiv preprint arXiv:2106.13627.
  30. 30.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23.
  31. 31.Lu Xu, Zhanming Jie, Wei Lu, and Lidong Bing. 2021. Better feature integration for named entity recognition. arXiv preprint arXiv:2104.05316.
  32. 32.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237.
  33. 33.Yanan Zheng, Jing Zhou, Yujie Qian, Ming Ding, Jian Li, Ruslan Salakhutdinov, Jie Tang, Sebastian Ruder, and Zhilin Yang. 2021. Fewnlu: Benchmarking state-of-the-art methods for few-shot natural language understanding. arXiv preprint arXiv:2109.12742.
  34. 34.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. arXiv preprint arXiv:2104.05240.

Citation

MLA
Liu, X., et al. “P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2022, pp. 61–68, https://doi.org/10.18653/v1/2022.acl-short.8.
APA
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., & Tang, J. (2022). P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 61–68. https://doi.org/10.18653/v1/2022.acl-short.8
Chicago
Liu, X., K. Ji, Y. Fu, et al. 2022. “P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 61–68. https://doi.org/10.18653/v1/2022.acl-short.8.
Harvard
Liu, X. et al. (2022) “P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 61–68. Available at: https://doi.org/10.18653/v1/2022.acl-short.8.
Vancouver
1. Liu X, Ji K, Fu Y, Tam W, Du Z, Yang Z, Tang J (2022) P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 61–68

BibTeX

@inproceedings{liu-etal-2022-p,
    title = "{P}-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks",
    author = "Liu, Xiao  and
      Ji, Kaixuan  and
      Fu, Yicheng  and
      Tam, Weng  and
      Du, Zhengxiao  and
      Yang, Zhilin  and
      Tang, Jie",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-short.8/",
    doi = "10.18653/v1/2022.acl-short.8",
    pages = "61--68"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/