SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer

Tu VuBrian LesterNoah ConstantRami Al-Rfou'Daniel Cer

article2022ACL349 citations

Introduces soft prompt transfer (SPoT) to initialize target task prompts from learned source prompts, enabling frozen language models to match or exceed full fine-tuning performance on SuperGLUE with up to 27,000 times fewer task parameters.

Listen

Large artificial intelligence language models deliver state-of-the-art performance across natural language processing tasks, but customizing and serving independent copies of multi-billion-parameter models for every application creates severe computational and financial bottlenecks. Parameter-efficient alternatives like soft prompt tuning avoid this overhead by keeping the underlying model frozen and learning only a small sequence of task-specific prompt parameters. However, prompt tuning historically underperforms full model fine-tuning, particularly on smaller and medium-sized models.

The article introduces and evaluates Soft Prompt Transfer (SPoT), a transfer learning approach designed to bridge this performance gap while preserving the efficiency of frozen models. In SPoT, a soft prompt is first trained on one or more source tasks—such as multi-task benchmark mixtures or data-rich individual tasks—and the resulting prompt parameters are used to initialize the prompt for a downstream target task. The researchers conducted extensive empirical evaluations across multiple model sizes using the T5 architecture (ranging from 60 million to 11 billion parameters) and evaluated transferability across 26 distinct tasks covering 160 source-target combinations.

The primary finding is that SPoT significantly enhances prompt tuning performance and stability across all model sizes. On the SuperGLUE benchmark, SPoT matched or exceeded full model fine-tuning across every size tier, achieving an 89.2 score on the public leaderboard with the largest 11-billion-parameter model—virtually matching fully fine-tuned baselines while updating roughly 27,000 times fewer parameters per task. Second, initializing prompts from general multi-task mixtures or tasks involving complex sentence reasoning, such as natural language inference, delivered substantial quality gains on downstream tasks, achieving up to a 58.9% relative error reduction. Third, the article demonstrates that early checkpoint prompts function effectively as semantic task embeddings, clustering similar tasks together and enabling an automated retrieval method that reduces the candidate source task search space by 69% while retaining 90% of optimal transfer gains.

These results demonstrate that massive parameter scaling is not required for prompt tuning to compete with full model fine-tuning. For enterprise deployment, SPoT offers a path to run dozens of specialized downstream applications using a single shared frozen model instance in memory, dramatically reducing infrastructure costs, deployment complexity, and storage requirements without sacrificing accuracy. For teams with limited computational budgets, practitioner recommendations include retrieving top-ranked source prompts via task embeddings or applying a weighted average across top source candidates.

Decision-makers should note that the approach relies on the availability of relevant source prompts and benefits from extended tuning durations on larger datasets. Additionally, while task embedding similarity strongly predicts transferability for several task categories, other factors such as source dataset size and reasoning complexity also influence transfer success. Overall, the findings provide strong confidence that soft prompt transfer offers an accurate, cost-effective alternative to full model fine-tuning for scalable language model deployment.

arXiv: 2110.07904
Cover for SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer

Abstract

There has been growing interest in parameter-efficient methods to apply pre-trained language models to downstream tasks. Building on the PromptTuning approach of Lester et al. (2021), which learns task-specific soft prompts to condition a frozen pre-trained model to perform different tasks, we propose a novel prompt-based transfer learning approach called SPoT: Soft Prompt Transfer. SPoT first learns a prompt on one or more source tasks and then uses it to initialize the prompt for a target task. We show that SPoT significantly boosts the performance of PromptTuning across many tasks. More remarkably, across all model sizes, SPoT matches or outperforms standard ModelTuning (which fine-tunes all model parameters) on the SuperGLUE benchmark, while using up to 27,000× fewer task-specific parameters. To understand where SPoT is most effective, we conduct a large-scale study on task transferability with 26 NLP tasks in 160 combinations, and demonstrate that many tasks can benefit each other via prompt transfer. Finally, we propose an efficient retrieval approach that interprets task prompts as task embeddings to identify similar tasks and predict the most transferable source tasks for a novel target task.

Table of Contents

  • 1 Introduction
  • 2 Improving PROMPTTUNING with SPOT
  • 2.1 Experimental setup
  • 2.1.1 Baselines
  • 2.1.2 Evaluation datasets
  • 2.1.3 Data for source prompt tuning
  • 2.1.4 Training details
  • 2.2 Effect of SPOT
  • 3 Predicting task transferability
  • 3.1 Measuring transferability
  • 3.2 Defining task similarity through prompts
  • 4 Related Work
  • 5 Limitations & Future work
  • 6 Conclusion
  • Acknowledgements
  • References
  • Appendices
  • A Full results for Figure 1
  • B Source datasets used in our SPOT experiments in §2
  • C Additional training details
  • D Details of our SUPERGLUE submission
  • E Task transferability results
  • F Task embedding similarity
  • G Correlation between task similarity and task transferability

Knowls

  1. Knowl 1 — Soft Prompt Transfer for Frozen Models

    model/method

    SPoT adapts a frozen pre-trained text-to-text model to a target task by inserting an intermediate source-prompt-tuning stage. A learnable soft prompt ρ∈RL×E\rho \in \mathbb{R}^{L\times E}, consisting of LL trainable token embeddings of dimension EE, is prepended to every input while all parameters of the base language model remain frozen. SPoT first trains ρ\rho on one or more source tasks, then uses the resulting prompt to initialize a new prompt that is tuned on the target task. The generic variant learns one source prompt and reuses it for every target; the targeted variant stores prompts from multiple source tasks and retrieves a different source prompt for each target. Only the final target prompt is task-specific, so the same frozen language model can be shared across tasks.

  2. Knowl 2 — SPoT Training Configuration

    experimental setup

    The experiments use LM-adapted T5 1.1 checkpoints with 60M, 220M, 770M, 3B, and 11B parameters, corresponding to the SMALL, BASE, LARGE, XL, and XXL models. Prompt length is fixed at L=100L=100 tokens. Source and target prompts are trained with Adafactor, constant learning rate 0.30.3, weight decay 10−510^{-5}, batch size 3232, dropout probability 0.10.1, and parameter scaling disabled. Source prompts are trained for 218=262,1442^{18}=262{,}144 steps from embeddings sampled from the 5,000 most frequent vocabulary items. Target prompts are checkpointed every 500 steps, and the checkpoint with the best validation performance is reported. The main evaluations use GLUE and SUPERGLUE validation sets; multi-task source mixtures use examples-proportional sampling with an artificial per-dataset limit of K=219K=2^{19} examples.

  3. Knowl 3 — SPoT Matches Full Model Tuning with Far Fewer Parameters

    data/table

    On SUPERGLUE, SPoT uses a GLUE-trained source prompt and is compared with prompt tuning from scratch, full model tuning, multi-task full model tuning, and GPT-3 prompt design. Scores are averages across the SUPERGLUE tasks; the reported uncertainties are standard deviations over three random seeds. SPoT substantially improves vanilla prompt tuning at every model size, exceeds full model tuning for LARGE, XL, and XXL, and reaches its highest score of 91.2 at XXL. At XXL, SPoT exceeds multi-task model tuning by 1.1 points while tuning 27,000 times fewer task-specific parameters.

    Could not parse LaTeX table

    A separate XXL leaderboard submission obtained 89.2 on SUPERGLUE, compared with 89.3 for fully fine-tuned T5 XXL and 71.8 for GPT-3 few-shot prompting, while SPoT tuned only 410K parameters.

  4. Knowl 4 — Source Task Mixtures Determine Transfer Quality

    data/table

    The source prompt used for transfer strongly affects downstream performance. With T5 BASE, prompt tuning from scratch scores 81.2 on GLUE and 66.6 on SUPERGLUE. A GLUE source mixture is best among the tested mixtures, reaching 82.8 and 73.2, respectively. MNLI and SQuAD are also strong single-task sources, while mixing all available datasets is not optimal, consistent with task interference or negative transfer. Longer target tuning and prompt transfer are complementary: removing longer tuning lowers the GLUE/SUPERGLUE result of the GLUE-source SPoT variant to 82.0/70.7 and also increases instability.

    Could not parse LaTeX table
  5. Knowl 5 — Large-Scale Study of Prompt Transferability

    experimental setup

    The transferability study evaluates all 160 source-target pairs formed by 16 source tasks and 10 low-resource target tasks with T5 BASE. The source tasks are C4, DOCNLI, YELP-2, MNLI, QQP, QNLI, ReCoRD, CxC, SQuAD, DROP, SST-2, WinoGrande, HellaSwag, MultiRC, CosmosQA, and RACE. The target tasks are BoolQ, CoLA, STS-B, WiC, CR, MRPC, RTE, WSC, COPA, and CB; each target has fewer than 10K training examples. Each source prompt is trained for 262,144 steps, the best source-validation checkpoint initializes the target prompt, and each target prompt is then trained for 100K steps. Every source-target experiment is repeated with three random seeds. Three runs for each of the 16 source tasks provide 48 candidate source prompts for the later retrieval experiments.

  6. Knowl 6 — Prompt Transfer Produces Strong and Sometimes Cross-Type Gains

    empirical result

    Prompt transfer improves many of the 10 low-resource target tasks. Selecting the best source prompt by exhaustive search over 48 candidates raises the mean target score from 74.7 for prompt tuning from scratch to 80.7. The largest reported gain is MNLI to CB, which increases the average CB score from 92.7 to 97.0 and corresponds to a 58.9% relative error reduction. Other large gains include MNLI to COPA at 29.1%, ReCoRD to WSC at 20.0%, MNLI to RTE at 19.2%, ReCoRD to MRPC at 18.7%, SQuAD to MRPC at 18.7%, and CxC to WiC at 18.1%. Transfer is therefore useful both for closely related tasks, such as CxC to STS-B, and for apparently dissimilar tasks, such as ReCoRD to WSC, SQuAD to MRPC, and CxC to WiC.

  7. Knowl 7 — Task Prompts as Task Embeddings

    model/method

    A task embedding is defined as the soft prompt checkpoint obtained after exactly 10,000 prompt-tuning steps on that task. Fixed-duration checkpoints are used because embeddings trained for different numbers of steps are less directly comparable, and preliminary alternatives from 1,000 to 100,000 steps performed worse. For task tat_a, let ea∈RL×Ee^a\in\mathbb{R}^{L\times E} denote its embedding, where eia∈REe^a_i\in\mathbb{R}^{E} is the vector for prompt token ii, L=100L=100 is the prompt length, and EE is the frozen model's embedding dimension. Two task-similarity measures are used:

    sim⁡avg(ta,tb)=cos⁡(1L∑i=1Leia,1L∑j=1Lejb),\operatorname{sim}_{\mathrm{avg}}(t_a,t_b)=\cos\left(\frac{1}{L}\sum_{i=1}^{L}e^a_i,\frac{1}{L}\sum_{j=1}^{L}e^b_j\right),

    where cos⁡\cos is cosine similarity between two vectors, and

    sim⁡token(ta,tb)=1L2∑i=1L∑j=1Lcos⁡(eia,ejb).\operatorname{sim}_{\mathrm{token}}(t_a,t_b)=\frac{1}{L^2}\sum_{i=1}^{L}\sum_{j=1}^{L}\cos(e^a_i,e^b_j).

    The first metric compares mean prompt representations; the second averages cosine similarity over every pair of prompt tokens.

  8. Knowl 8 — Embedding Similarity Reveals Structure but Is Not a Complete Transfer Predictor

    limitation

    Hierarchical clustering of the 26 task embeddings recovers meaningful groups: SQuAD, ReCoRD, and DROP cluster as question answering; MultiRC and BoolQ form another question-answering group; YELP-2, SST-2, and CR cluster as sentiment tasks; MNLI and CB, and separately DOCNLI and RTE, cluster as natural-language-inference tasks; STS-B and CxC cluster as semantic-similarity tasks; MRPC and QQP cluster as paraphrase tasks; and WinoGrande, HellaSwag, and CosmosQA cluster as commonsense-reasoning tasks. The embeddings are more sensitive to task format than shared domain: QNLI is not close to SQuAD despite being constructed from SQuAD data, while ReCoRD's relationship to WSC reflects its observed transferability.

    Similarity is an imperfect predictor of transfer. Significant positive correlations between similarity and relative error reduction occur for STS-B (r=0.708r=0.708, p=1.853×10−8p=1.853\times10^{-8}), CB (r=0.490r=0.490, p<0.001p<0.001), WSC (r=0.428r=0.428, p=0.002p=0.002), and RTE (r=0.290r=0.290, p=0.046p=0.046), but not for the other six targets at conventional significance levels. For example, MNLI produces a 19.0% relative error reduction on BoolQ despite a cosine similarity of only 0.4. The authors therefore note that data size, task difficulty, and domain similarity can influence transfer in ways not captured by prompt embeddings, and leave alternative task-embedding methods for future work.

  9. Knowl 9 — Retrieval of Source Prompts for a Novel Target

    algorithm

    The targeted SPoT retrieval procedure treats a prompt library as key-value storage: each source-task embedding is a key, and the best-performing source prompt checkpoint is its value. For a novel target, an early target prompt checkpoint supplies a query embedding. The procedure ranks source prompts by embedding similarity and then uses the most similar candidates in one of three ways.

    Input: source prompt library with source prompts rho_s and embeddings e_s; target dataset D_t; candidate count k; similarity function sim
    Output: tuned target prompt rho_t
    Train a fresh target prompt on D_t for 10,000 steps and save its checkpoint as target embedding e_t
    Rank all source prompts in descending order of sim(e_s, e_t)
    Let rho_sr and e_sr be the prompt and embedding at rank r, for r = 1, ..., 48
    If using best of top-k:
        For r = 1, ..., k:
            Initialize a target prompt with rho_sr
            Tune the target prompt on D_t for 100,000 steps
            Evaluate the best validation checkpoint
        Return the target prompt with the highest validation score
    If using top-k weighted average:
        For r = 1, ..., k:
            Set alpha_r = sim(e_sr, e_t) / sum over l = 1 to k of sim(e_sl, e_t)
        Initialize the target prompt with sum over r = 1 to k of alpha_r times rho_sr
        Tune the target prompt on D_t for 100,000 steps
        Return its best validation checkpoint
    If using top-k multi-task mixture:
        Mix D_t with the datasets of the source tasks represented by the top-k prompts
        Train a source prompt on this mixture
        Initialize the target prompt with the resulting source prompt
        Tune the target prompt on D_t for 100,000 steps
        Return its best validation checkpoint

    Best of top-kk requires kk target-task tuning runs; weighted averaging requires one target run but combines prompts; multi-task mixture requires an additional source-prompt-tuning stage. In the study, the library contains 48 source prompts from three runs on each of 16 source tasks.

  10. Knowl 10 — Similarity-Based Retrieval Eliminates Most Source Search

    data/table

    On the 10-target transfer benchmark, the baseline mean score is 74.7 ± 0.7, while brute-force selection of the best of all 48 source prompts reaches 80.7 ± 0.0. Ranking with per-token average cosine similarity is more effective than ranking with cosine similarity between average prompt tokens. Trying only the top three per-token candidates reaches 77.5 ± 0.4, a gain of 2.8 points over baseline. With k=9k=9 and k=15k=15, it retains approximately 80% and 90% of the oracle's score gain while searching only 9 and 15 of 48 candidates. Weighted averaging has lower variance than trying a single candidate, and the multi-task mixture reaches 77.8 ± 0.1.

    Could not parse LaTeX table

Coverage note — The complete 26-by-10 transfer matrix, exhaustive dataset catalog, and leaderboard comparison table were omitted because their individual entries are largely redundant with the strongest transfer examples and aggregate retrieval results.

References

  1. 1.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020).
  2. 2.Joachim Bingel and Anders Søgaard. 2017. Identifying beneficial task relations for multi-task learning in deep neural networks. In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics (EACL 2017), pages 164–169.
  3. 3.Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2020), 34(05):7432–7439.
  4. 4.Ondˇrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation (WMT 2014), pages 12–58.
  5. 5.Ondˇrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation (WMT 2016), pages 131–198.
  6. 6.Ondˇrej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015. Findings of the 2015 workshop on statistical machine translation. In Proceedings of the Tenth Workshop on Statistical Machine Translation (WMT 2015), pages 1–46.
  7. 7.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015), pages 632–642.
  8. 8.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), volume 33, pages 1877–1901.
  10. 10.Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval 2017), pages 1–14.
  11. 11.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (ACL 2019), pages 2924–2936.
  12. 12.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Proceedings of the 1st International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment (MLCW 2005), page 177–190.
  13. 13.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The CommitmentBank: Investigating projection in naturally occurring discourse. In Proceedings of Sinn und Bedeutung 23 (SuB 2018), volume 23, pages 107–124.
  14. 14.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 4040–4054.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2019), pages 4171–4186.
  16. 16.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP 2005).
  17. 17.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2019), pages 2368–2378.
  18. 18.Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. Searchqa: A new q&a dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179.
  19. 19.Ondˇrej Dušek, David M. Howcroft, and Verena Rieser. 2019. Semantic noise matters for neural natural language generation. In Proceedings of the 12th International Conference on Natural Language Generation (INLG 2019), pages 421–426.
  20. 20.Alexander Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pages 1074–1084.
  21. 21.Angela Fan, Edouard Grave, and Armand Joulin. 2020. Reducing transformer depth on demand with structured dropout. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020).
  22. 22.Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019. MRQA 2019 shared task: Evaluating generalization in reading comprehension. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering (MRQA 2019), pages 1–13.
  23. 23.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL 2021), pages 3816–3830.
  24. 24.Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. Creating training corpora for NLG micro-planners. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), pages 179–188.
  25. 25.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondˇrej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120.
  26. 26.Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. In Proceedings of the 2nd Workshop on New Frontiers in Summarization (NewSum 2019), pages 70–79.
  27. 27.Alec Go, Richa Bhayani, and Lei Huang. 2009. Twitter sentiment classification using distant supervision. CS224N Project Report, Stanford.
  28. 28.David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003. English gigaword. Linguistic Data Consortium, Philadelphia, 4(1):34.
  29. 29.Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2018), pages 708–719.
  30. 30.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2021. PPT: Pre-trained prompt tuning for few-shot learning. arXiv preprint arXiv:2109.04332.
  31. 31.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. WARP: Word-level Adversarial ReProgramming. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Computational Linguistics (ACL-IJCNLP 2021), pages 4921–4933.
  32. 32.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: Decoding-enhanced bert with disentangled attention. In Proceedings of the 9th International Conference on Learning Representations (ICLR 2021).
  33. 33.Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee. 2020. Flax: A neural network library and ecosystem for JAX.
  34. 34.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Proceedings of the 29th Conference on Neural Information Processing Systems (NeurIPS 2020), volume 28.
  35. 35.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (PMLR 2019), volume 97, pages 2790–2799.
  36. 36.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  37. 37.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2004), page 168–177.
  38. 38.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Cosmos QA: Machine reading comprehension with contextual commonsense reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019), pages 2391–2401.
  39. 39.Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017. First Quora Dataset Release: Question pairs.
  40. 40.Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020a. Neural CRF model for sentence alignment in text simplification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 7943–7960.
  41. 41.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020b. How can we know what language models know? Transactions of the Association for Computational Linguistics (TACL 2020), 8:423–438.
  42. 42.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics (Findings of EMNLP 2020), pages 4163–4174.
  43. 43.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), pages 1601–1611.
  44. 44.Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Computational Linguistics (ACL-IJCNLP 2021), pages 565–576.
  45. 45.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2018), pages 252–262.
  46. 46.Anastassia Kornilova and Vladimir Eidelman. 2019. BillSum: A corpus for automatic summarization of US legislation. In Proceedings of the 2nd Workshop on New Frontiers in Summarization (NewSum 2019), pages 48–56.
  47. 47.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics (TACL 2019), 7:452–466.
  48. 48.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. In Findings of the Association for Computational Linguistics (Findings of EMNLP 2020), pages 4034–4048.
  49. 49.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP 2017), pages 785–794.
  50. 50.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020).
  51. 51.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 3045–3059.
  52. 52.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning (KR 2012), page 552–561.
  53. 53.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Computational Linguistics (ACL 2021), pages 4582–4597.
  54. 54.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics (Findings of EMNLP 2020), pages 1823–1840.
  55. 55.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a. Linguistic knowledge and transferability of contextual representations. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2019), pages 1073–1094.
  56. 56.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  57. 57.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021b. Gpt understands, too. arXiv preprint arXiv:2103.10385.
  58. 58.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  59. 59.Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2021), 35(15):13480–13488.
  60. 60.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021. Compacter: Efficient low-rank hypercomplex adapter layers. arXiv preprint arXiv:2106.04647.
  61. 61.Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021. DART: Open-domain structured data record to text generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2021), pages 432–447.
  62. 62.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP 2018), pages 1797–1807.
  63. 63.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 4885–4901.
  64. 64.Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang. 2021. Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021), pages 2855–2870.
  65. 65.Jason Phang, Thibault Févry, and Samuel R Bowman. 2019. Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks. arXiv preprint arXiv:1811.01088.
  66. 66.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2019), pages 1267–1273.
  67. 67.Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, and Iryna Gurevych. 2021. What to pre-train on? Efficient intermediate task selection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 10585–10605.
  68. 68.Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020. Intermediate-task transfer learning with pretrained language models: When and why does it work? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 5231–5247.
  69. 69.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying LMs with mixtures of soft prompts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2021), pages 5203–5212.
  70. 70.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR 2020), 21(140):1–67.
  71. 71.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP 2016), pages 2383–2392.
  72. 72.Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020. Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2020), 34(05):8689–8696.
  73. 73.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In Proceedings of the 25th AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning (AAAI Spring Symposium 2011).
  74. 74.Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015), pages 379–389.
  75. 75.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2020), 34(05):8732–8740.
  76. 76.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. In Proceedings of the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing (EMC2 2019).
  77. 77.Victor Sanh, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), volume 33, pages 20378–20389.
  78. 78.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social IQa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019), pages 4463–4473.
  79. 79.Timo Schick and Hinrich Schütze. 2021. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2021), pages 2339–2352.
  80. 80.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), pages 1073–1083.
  81. 81.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. arXiv preprint arXiv:1804.04235.
  82. 82.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 4222–4235.
  83. 83.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP 2013), pages 1631–1642.
  84. 84.Alon Talmor and Jonathan Berant. 2019. MultiQA: An empirical investigation of generalization and transfer in reading comprehension. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL 2019), pages 4911–4921.
  85. 85.Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. NewsQA: A machine comprehension dataset. In Proceedings of the Workshop on Representation Learning for NLP (RepL4NLP 2017), pages 191–200.
  86. 86.Tu Vu, Minh-Thang Luong, Quoc Le, Grady Simon, and Mohit Iyyer. 2021. STraTA: Self-training with task augmentation for better few-shot learning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 5715–5731.
  87. 87.Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer. 2020. Exploring and predicting transferability across NLP tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 7882–7926.
  88. 88.Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, and Samuel R. Bowman. 2019a. Can you tell me how to get past sesame street? sentence-level pretraining beyond language modeling. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL 2019), pages 4465–4476.
  89. 89.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019b. Superglue: A stickier benchmark for general-purpose language understanding systems. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS 2019), volume 32, pages 3266–3280.
  90. 90.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019c. Glue: A multi-task benchmark and analysis platform for natural language understanding. Proceedings of the 7th International Conference on Learning Representations (ICLR 2019).
  91. 91.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics (TACL 2019), 7:625–641.
  92. 92.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2018), pages 1112–1122.
  93. 93.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Proceedings of the 33th Conference on Neural Information Processing Systems (NeurIPS 2019), volume 32, pages 5753–5763.
  94. 94.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP 2018), pages 2369–2380.
  95. 95.Wenpeng Yin, Dragomir Radev, and Caiming Xiong. 2021. DocNLI: A large-scale dataset for document-level natural language inference. In Findings of the Association for Computational Linguistics (Findings of ACL-IJCNLP 2021), pages 4913–4922.
  96. 96.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv preprint arXiv:2106.10199.
  97. 97.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pages 4791–4800.
  98. 98.Rui Zhang and Joel Tetreault. 2019. This email could save your life: Introducing the task of email subject line generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019), pages 446–456.
  99. 99.Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. 2018. Record: Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885.
  100. 100.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of the 29th Conference on Neural Information Processing Systems (NeurIPS 2015), volume 28, pages 649–657.
  101. 101.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning (ICML 2021), volume 139 of PMLR, pages 12697–12706.

Citation

MLA
Vu, T., et al. “SPoT: Better Frozen Model Adaptation Through Soft Prompt Transfer”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 5039–59, https://doi.org/10.18653/v1/2022.acl-long.346.
APA
Vu, T., Lester, B., Constant, N., Al-Rfou’, R., & Cer, D. (2022). SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5039–5059. https://doi.org/10.18653/v1/2022.acl-long.346
Chicago
Vu, T., B. Lester, N. Constant, R. Al-Rfou’, and D. Cer. 2022. “SPoT: Better Frozen Model Adaptation Through Soft Prompt Transfer”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5039–59. https://doi.org/10.18653/v1/2022.acl-long.346.
Harvard
Vu, T. et al. (2022) “SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5039–5059. Available at: https://doi.org/10.18653/v1/2022.acl-long.346.
Vancouver
1. Vu T, Lester B, Constant N, Al-Rfou’ R, Cer D (2022) SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5039–5059

BibTeX

@inproceedings{vu-etal-2022-spot,
    title = "{SP}o{T}: Better Frozen Model Adaptation through Soft Prompt Transfer",
    author = "Vu, Tu  and
      Lester, Brian  and
      Constant, Noah  and
      Al-Rfou{'}, Rami  and
      Cer, Daniel",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.346/",
    doi = "10.18653/v1/2022.acl-long.346",
    pages = "5039--5059"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/