RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning

Mingkai DengJianyu WangCheng-Ping HsiehYihan WangHan GuoTianmin ShuMeng SongEric P. XingZhiting Hu

article2022EMNLP548 citations

Proposes RLPrompt, a reinforcement learning framework that optimizes discrete text prompts across black-box language models using a parameter-efficient policy network, achieving strong performance in few-shot classification and text generation while revealing that effective machine prompts often transfer across models despite resembling ungrammatical text.

Listen

Adapting large language models to new tasks typically requires either costly full-model updates or manual prompt engineering. While automated continuous prompt tuning offers an alternative, it requires internal gradient access that is often unavailable in commercial application programming interfaces, produces uninterpretable vectors, and fails to transfer across different models. Optimizing discrete text prompts directly has historically been intractable due to combinatorial complexity and training instability.

The article demonstrates an automated, parameter-efficient framework called RLPrompt that optimizes discrete text prompts using reinforcement learning—a goal-oriented machine learning training approach. The objective is to efficiently steer frozen language models across classification and text generation tasks without requiring access to their internal gradients or extensive supervised data.

The researchers evaluated this framework by training a compact policy network—adding only a small neural module of roughly 3.1 million parameters to a frozen base model—to generate discrete prompt tokens based on task reward signals. To overcome reinforcement learning instability caused by complex language model environments, the approach introduced input-specific normalization and piecewise reward structures. The method was tested on few-shot text classification across multiple standard benchmarks using masked models like RoBERTa and on unsupervised text style transfer using autoregressive models like GPT-2.

The evaluation produced several key findings: First, RLPrompt consistently outperformed manual prompting, instruction prompts, in-context learning, and continuous prompt tuning in few-shot classification, achieving an average accuracy of 75.8% with five discrete tokens compared to 68.6% for manual prompts. Second, in unsupervised text style transfer, the method achieved competitive or superior joint content, style, and fluency scores relative to expensive full-model fine-tuning baselines while substantially exceeding manual and random prompting. Third, the highest-performing discrete prompts often appeared as ungrammatical gibberish, and enforcing human-readable fluency noticeably reduced task performance. Fourth, these non-intuitive prompts transferred successfully across different model architectures and sizes, with prompts optimized on smaller models maintaining strong performance when deployed on larger models.

These findings indicate that language models do not process task instructions like humans, utilizing shared underlying representations that depart from natural language syntax. For organizations deploying artificial intelligence systems, this approach enables significant cost and compute savings by allowing prompt optimization via small, lightweight models and deployment on black-box, closed-source models without expensive parameter updates or proprietary API access.

Decision-makers should consider adopting automated reinforcement learning prompt optimization when utilizing API-only language models or operating under strict data and compute constraints. Organizations can leverage the strategy of discovering discrete prompts on smaller internal models before deploying them on larger production systems. Future work should validate this approach on larger-scale foundational models, explore automated reward generation techniques such as inverse reinforcement learning, and analyze the underlying mechanics of these non-standard prompting patterns.

Confidence in these findings is supported by consistent multi-trial benchmarking across diverse datasets and model families. However, users should remain cautious when human auditability is critical, as the resulting prompts lack natural readability, and manual reward function design currently requires task-specific engineering.

Cover for RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning

Abstract

Prompting has shown impressive success in enabling large pre-trained language models (LMs) to perform diverse NLP tasks, especially with only few downstream data. Automatically finding the optimal prompt for each task, however, is challenging. Most existing work resorts to tuning soft prompts (e.g., embeddings) which fall short of interpretability, reusability across LMs, and applicability when gradients are not accessible. Discrete prompts, on the other hand, are difficult to optimize, and are often created by "enumeration (e.g., paraphrasing)-then-selection" heuristics that do not explore the prompt space systematically. This paper proposes RLPROMPT, an efficient discrete prompt optimization approach with reinforcement learning (RL). RLPROMPT formulates a parameter-efficient policy network that generates the optimized discrete prompt after training with reward. To harness the complex and stochastic reward signals from the large LM environment, we incorporate effective reward stabilization that substantially enhances training efficiency. RLPROMPT is flexibly applicable to different types of LMs, such as masked (e.g., BERT) and left-to-right models (e.g., GPTs), for both classification and generation tasks. Experiments on few-shot classification and unsupervised text style transfer show superior performance over a wide range of existing fine-tuning or prompting methods. Interestingly, the resulting optimized prompts are often ungrammatical gibberish text; and surprisingly, those gibberish prompts are transferrable between different LMs to retain significant performance, indicating that LM prompting may not follow human language patterns.

Table of Contents

  • 1 Introduction
  • 2 Discrete Prompt Optimization with RL
  • 2.1 Discrete Prompt Optimization Problem
  • 2.2 The Reinforcement Learning Formulation
  • 2.3 Efficient Parameterization of Policy
  • 2.4 Reward Engineering and Stabilization
  • 3 Experiments
  • 3.1 Few-Shot Text Classification
  • 3.2 Unsupervised Text Style Transfer
  • 3.3 Analysis
  • 4 Related Work
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgements
  • Ethics Statement
  • References
  • A Experiment Details
  • A.1 Policy Network
  • A.2 Few-Shot Text Classification
  • A.3 Text Style Transfer
  • B Additional Analysis
  • C Additional Related Work
  • C.1 Prompting Paradigms
  • C.2 Controllable Text Generation

Knowls

  1. Knowl 1 — Discrete Prompt Optimization via Reinforcement Learning (RLPrompt)

    model/method

    RLPrompt is a framework that optimizes discrete text prompts for frozen pre-trained language models (LMs) using reinforcement learning. Given an input text xx, a vocabulary V\mathcal{V}, and a pre-trained task language model PLMP_{\text{LM}}, a discrete text prompt of fixed length TT is represented as a sequence of tokens z=(z1,z2,…,zT)∈VTz = (z_1, z_2, \dots, z_T) \in \mathcal{V}^T. The task LM takes the concatenated prompt and input and produces an output yLM(z,x)∼PLM(y∣z,x)y_{\text{LM}}(z, x) \sim P_{\text{LM}}(y \mid z, x).

    The optimization goal is to find a prompt z∗∈VTz^* \in \mathcal{V}^T that maximizes a downstream task performance metric R(yLM(z,x))R(y_{\text{LM}}(z, x)):

    max⁡z∈VTR(yLM(z,x))\max_{z \in \mathcal{V}^T} R(y_{\text{LM}}(z, x))

    Because direct search over VT\mathcal{V}^T has exponential complexity O(∣V∣T)\mathcal{O}(|\mathcal{V}|^T) and discrete tokens prevent gradient backpropagation into the prompt text, RLPrompt reformulates the prompt generation process as an RL policy πθ\pi_\theta. The policy generates prompt tokens sequentially according to πθ(z)=∏t=1Tπθ(zt∣z<t)\pi_\theta(z) = \prod_{t=1}^T \pi_\theta(z_t \mid z_{<t}). The policy parameter θ\theta is trained to maximize the expected reward over generated prompts:

    max⁡θEz^∼πθ[R(yLM(z^,x))]\max_\theta \mathbb{E}_{\hat{z} \sim \pi_\theta} \left[ R(y_{\text{LM}}(\hat{z}, x)) \right]

    The task LM is treated entirely as a black-box reward environment without requiring access to its internal parameters or gradients. Optimization is performed using on-policy Soft Q-Learning (SQL). During inference, prompt tokens are chosen greedily from the trained policy to produce a single deterministic discrete prompt.

  2. Knowl 2 — Policy Network Parameterization in RLPrompt

    model/method

    The prompt policy network πθ\pi_\theta in RLPrompt is parameterized by inserting a small, trainable multilayer perceptron (MLP) module into a compact, frozen pre-trained language model (the policy LM, such as distilGPT-2 with 82 million parameters). The policy LM is distinct and decoupled from the downstream task LM for which prompts are being optimized.

    At step tt, the frozen policy LM extracts contextual embeddings for the partially generated prompt tokens z<tz_{<t}. The task-specific MLP layer (composed of 1 hidden layer and 2048 hidden states, adding approximately 3.1 million trainable parameters, or 3.8% of the policy LM size) transforms these contextual embeddings. The adapted representation is then projected through the policy LM's original vocabulary head to yield the categorical probability distribution over the next prompt token zt∈Vz_t \in \mathcal{V}.

    For input-dependent prompt generation (such as controllable text generation), the policy network is conditioned directly on the input xx, producing πθ(z∣x)\pi_\theta(z \mid x). Gradients are backpropagated through the policy LM to update only the MLP parameters θ\theta. Once training is complete, the policy LM and MLP are discarded, and the selected discrete prompt tokens are used directly with the task LM for inference.

  3. Knowl 3 — Input-Specific z-Score Reward Normalization

    equation

    To prevent training instability and optimization bias caused by differing difficulty levels across training inputs in reinforcement-learning-guided prompt optimization, rewards are normalized per input using an input-specific zz-score transformation.

    For a given input sentence xx, a batch of candidate prompts Z(x)Z(x) is sampled from the input-conditioned policy πθ(z∣x)\pi_\theta(z \mid x). Denoting the task reward achieved by prompt z∈Z(x)z \in Z(x) on input xx as Rx(z):=R(yLM(z,x))R_x(z) := R(y_{\text{LM}}(z, x)), the input-normalized reward z-score(z,x)\text{z-score}(z, x) is defined as:

    z-score(z,x)=Rx(z)−meanz′∈Z(x)Rx(z′)stdevz′∈Z(x)Rx(z′)\text{z-score}(z, x) = \frac{R_x(z) - \text{mean}_{z' \in Z(x)} R_x(z')}{\text{stdev}_{z' \in Z(x)} R_x(z')}

    where meanz′∈Z(x)Rx(z′)\text{mean}_{z' \in Z(x)} R_x(z') and stdevz′∈Z(x)Rx(z′)\text{stdev}_{z' \in Z(x)} R_x(z') are the sample mean and sample standard deviation of the rewards across the batch of candidate prompts evaluated on that specific input xx.

  4. Knowl 4 — Piecewise Margin Reward Function for Prompted Classification

    equation

    In few-shot prompt optimization for classification, using unconstrained label probabilities as rewards can cause the policy to converge to degenerate or adversarial prompts that predict a single majority class regardless of input. To avoid this, RLPrompt employs a piecewise margin reward function that couples a continuous probability gap with a discrete success multiplier.

    Let (x,c)(x, c) be an input example with ground-truth class label c∈Cc \in \mathcal{C}, and let Pz(c):=PLM(c∣z,x)P_z(c) := P_{\text{LM}}(c \mid z, x) denote the task LM probability of predicting the verbalizer token for class cc at the masked token position under prompt zz. The classification margin Gapz(c)\text{Gap}_z(c) and binary correctness indicator Correct\text{Correct} are defined as:

    Gapz(c):=Pz(c)−max⁡c′≠cPz(c′)\text{Gap}_z(c) := P_z(c) - \max_{c' \neq c} P_z(c')

    Correct:=I[Gapz(c)>0]\text{Correct} := \mathbb{I}[\text{Gap}_z(c) > 0]

    The piecewise reward function R(x,c)R(x, c) is computed as:

    R(x,c)=λ11−Correctλ2CorrectGapz(c)R(x, c) = \lambda_1^{1 - \text{Correct}} \lambda_2^{\text{Correct}} \text{Gap}_z(c)

    where λ1>0\lambda_1 > 0 and λ2>0\lambda_2 > 0 are hyperparameter weights (specifically tuned to λ1=180\lambda_1 = 180 and λ2=200\lambda_2 = 200). The large positive multiplier λ2\lambda_2 scales up the reward exclusively when the model's prediction is correct, stabilizing optimization across balanced classes.

  5. Knowl 5 — Few-Shot Text Classification Performance of RLPrompt

    data/table

    The table below compares RLPrompt with various fine-tuning and prompting paradigms for few-shot text classification using a frozen RoBERTa-large model (with 16 training and 16 validation examples per class). Evaluation is conducted across seven classification benchmarks: SST-2, Yelp Polarity (Yelp P.), Movie Reviews (MR), Customer Reviews (CR), SST-5, Yelp (5-class), and AG's News. Values are test accuracy percentages with standard deviations across multiple runs in parentheses.

    Methods SST-2 Yelp P. MR CR SST-5 Yelp AG's News Avg.
    Fine-Tuning 80.6 (3.9) 88.7 (4.7) 67.4 (9.7) 73.3 (7.5) 40.7 (3.0) 51.0 (2.2) 84.9 (3.6) 69.5
    Manual Prompt 82.8 83.0 80.9 79.6 34.9 42.1 76.9 68.6
    Instructions 89.0 84.4 85.2 80.8 29.8 43.0 54.8 58.5
    In-Context Demonstration 85.9 (0.7) 89.6 (0.4) 80.6 (1.4) 85.5 (1.5) 39.3 (0.9) 49.4 (0.3) 74.9 (0.8) 72.2
    Prompt Tuning (Soft) 73.8 (10.9) 88.6 (2.1) 74.1 (14.6) 75.9 (11.8) 40.2 (6.5) 49.1 (3.1) 82.6 (0.9) 69.2
    BB Tuning (2 soft tokens) 83.2 (3.5) 86.0 (1.6) 77.1 (3.9) 83.2 (2.5) 39.2 (2.4) 41.5 (1.9) 74.0 (1.9) 69.2
    BB Tuning (5 soft tokens) 84.6 (4.0) 78.7 (2.3) 79.8 (1.5) 82.9 (3.6) 36.6 (2.1) 33.7 (2.3) 73.6 (3.6) 67.1
    BB Tuning (Mixed, 50 soft) 89.1 (0.9) 93.2 (0.5) 86.6 (1.3) 87.4 (1.0) 38.4 (1.1) 44.8 (1.3) 83.5 (0.9) 74.7
    GrIPS (Prompt Enumeration) 87.1 (1.5) 88.2 (0.1) 86.1 (0.3) 80.0 (2.5) 32.0 (1.8) 47.2 (0.5) 65.4 (9.8) 69.4
    AutoPrompt 75.0 (7.6) 79.8 (8.3) 62.0 (0.8) 57.5 (5.8) 27.8 (3.3) 29.0 (5.0) 65.7 (1.9) 56.7
    RLPrompt (2 tokens) 90.3 (1.3) 94.1 (0.8) 86.5 (1.2) 87.4 (1.7) 40.1 (1.9) 45.6 (3.8) 76.8 (1.4) 74.4
    RLPrompt (5 tokens) 92.5 (0.8) 95.1 (1.0) 87.1 (0.4) 89.5 (0.6) 41.4 (3.2) 44.8 (4.3) 80.2 (0.7) 75.8

    RLPrompt (5 tokens) outperforms manual prompting and instructional prompts across all datasets, achieves higher average accuracy than full fine-tuning and in-context demonstration, and exhibits substantially lower standard deviations than soft prompt tuning.

  6. Knowl 6 — Unsupervised Text Style Transfer with RLPrompt

    empirical result

    RLPrompt optimizes discrete prompts for unsupervised text style transfer (TST) using frozen GPT-2 models without parallel training pairs. The optimization reward is defined as:

    R(x,y,s)=Content(x,y)+Style(y,s)R(x, y, s) = \text{Content}(x, y) + \text{Style}(y, s)

    where Content(x,y)\text{Content}(x, y) evaluates token embedding alignment via CTC metrics using RoBERTa-large, and Style(y,s)\text{Style}(y, s) measures target style probability via a BERT-base classifier. Generation quality is evaluated using the sentence-level joint score J(C,S,F)=meanx∈X(Content(x)⋅Style(x)⋅Fluency(x))J(C, S, F) = \text{mean}_{x \in \mathcal{X}}(\text{Content}(x) \cdot \text{Style}(x) \cdot \text{Fluency}(x)) and the geometric mean GM(C,S,F)\text{GM}(C, S, F) across test set X\mathcal{X}, where Fluency is evaluated with a CoLA grammaticality classifier.

    On the Yelp sentiment transfer task using GPT-2-xl (1.5B parameters) with prompt length T=5T=5:

    1. RLPrompt achieves Content=72.1%\text{Content} = 72.1\%, Style=94.2%\text{Style} = 94.2\%, Fluency=89.5%\text{Fluency} = 89.5\%, J(C,S,F)=61.4J(C,S,F) = 61.4, and GM=84.7\text{GM} = 84.7.
    2. Compared to training-based methods (Style Transformer: J=46.1,GM=75.2J = 46.1, \text{GM} = 75.2; DiRR full GPT-2 fine-tuning: J=59.6,GM=83.5J = 59.6, \text{GM} = 83.5), RLPrompt attains a higher joint score and geometric mean because keeping the task LM frozen better preserves generation fluency.
    3. Compared to prompting baselines on GPT-2-xl (Null Prompt: J=33.6J = 33.6; Random Prompt: J=34.7J = 34.7; Manual Prompt: J=53.4J = 53.4), RLPrompt substantially improves transfer quality while demonstrating lower variance across prompt runs.
  7. Knowl 7 — Fluency vs. Task Performance Trade-off in Optimized Prompts

    empirical result

    Unconstrained optimization of discrete prompts via RLPrompt consistently produces ungrammatical, gibberish token sequences (e.g., Parameters Comparison )=( Compare either or Fixed (- contrasts (- contrasts) that nevertheless drive high downstream task performance.

    When a fluency constraint is enforced by restricting the policy network's action space at each step tt to the top-20 tokens ranked by conditional probability under a pre-trained GPT-2 language model, the resulting prompts become human-readable (e.g., I love my life (), and prompt perplexity (PPL) under GPT-2 drops from 254,000 to 82.1. However, this fluency constraint causes downstream performance on Yelp sentiment style transfer with GPT-2-xl to drop sharply:

    • Joint quality metric J(C,S,F)J(C, S, F) declines from 61.461.4 to 46.746.7.
    • Geometric mean GM(C,S,F)\text{GM}(C, S, F) declines from 84.784.7 to 78.178.1.
    • Content preservation score drops from 72.172.1 to 52.452.4.

    This demonstrates that pre-trained language models utilize prompt tokens in a manner distinct from human grammatical patterns, and requiring prompts to be fluent limits downstream task performance.

  8. Knowl 8 — Cross-Model and Cross-Architecture Transferability of Discrete Prompts

    empirical result

    Discrete prompts optimized with RLPrompt transfer across different model scales and across distinct model architectures (such as between bidirectional masked language models like RoBERTa and causal autoregressive language models like GPT-2) due to the shared discrete token vocabulary space.

    Key empirical findings on transferability include:

    1. Small-to-Large Model Transfer: Prompts trained on smaller models transfer effectively or achieve higher performance when executed on larger models. For instance, in SST-2 sentiment classification, a 2-token prompt trained on distilRoBERTa-base yields 78.5% accuracy on distilRoBERTa-base and 77.2% on RoBERTa-large; a prompt trained on RoBERTa-base achieves 88.2% on RoBERTa-base and rises to 89.9% on RoBERTa-large.
    2. Large-to-Small Degradation: Prompts optimized on larger models experience performance drops when evaluated on smaller models (e.g., a prompt trained on RoBERTa-large achieves 90.7% on RoBERTa-large but drops to 76.8% when evaluated on distilRoBERTa-base).
    3. Cross-Architecture Transfer: Prompts learned on masked LMs (RoBERTa) transfer to causal LMs (GPT-2) and vice versa, indicating that prompt optimization activates shared internal linguistic structures present across diverse pre-trained model types.
  9. Knowl 9 — Robustness of RLPrompt to Classification Verbalizer Choices

    empirical result

    RLPrompt is robust to the choice of label verbalizers in prompt-based classification, finding effective discrete prompts regardless of the specific token pair assigned to class labels.

    On SST-2 sentiment classification using RoBERTa-large, RLPrompt and manual prompt templates achieve the following test accuracies across three different verbalizer pairs:

    • Verbalizers terrible, great: RLPrompt reaches 92.8%±0.8%92.8\% \pm 0.8\%, compared to 82.8%82.8\% for Manual Prompt.
    • Verbalizers bad, good: RLPrompt reaches 91.2%±1.4%91.2\% \pm 1.4\%, compared to 79.7%79.7\% for Manual Prompt.
    • Verbalizers negative, positive: RLPrompt reaches 92.2%±0.6%92.2\% \pm 0.6\%, compared to 76.8%76.8\% for Manual Prompt.

    While manual prompt performance degrades by up to 6.0 percentage points when shifting away from canonical verbalizers, RLPrompt consistently discovers prompts that maintain high performance (above 91%) with low variance.

  10. Knowl 10 — Stated Limitations of RLPrompt

    limitation

    The authors state three primary limitations of the RLPrompt framework:

    1. Evaluation Scale: The empirical evaluations are conducted on regular-sized language models (up to RoBERTa-large with 355M parameters and GPT-2-xl with 1.5B parameters) and do not evaluate massive models like GPT-3.
    2. Reward Engineering Requirement: Like standard reinforcement learning approaches, designing effective reward functions (such as balancing weights and piecewise components) requires task-specific domain engineering, though future approaches could learn rewards directly via inverse reinforcement learning.
    3. Uncharacterized Prompt Mechanics: The underlying structural patterns and mechanisms governing the learned, unintelligible prompt tokens (the language model's "secret language") remain uncharacterized.

Coverage note — None was omitted; all primary contributions—including the RL formulation, policy parameterization, piecewise and z-score reward stabilization techniques, few-shot classification and text style transfer results, fluency trade-offs, cross-model prompt transferability, verbalizer robustness, and stated limitations—are fully covered.

References

  1. 1.Shengnan An, Yifei Li, Zeqi Lin, Qian Liu, Bei Chen, Qiang Fu, Weizhu Chen, Nanning Zheng, and Jian-Guang Lou. 2022. Input-tuning: Adapting unfamiliar inputs to frozen pretrained models. arXiv preprint arXiv:2203.03131.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS, pages 1877–1901.
  3. 3.Jordan Clive, Kris Cao, and Marek Rei. 2021. Control prefixes for text generation. arXiv preprint arXiv:2110.08329.
  4. 4.Ning Dai, Jianze Liang, Xipeng Qiu, and Xuan-Jing Huang. 2019. Style transformer: Unpaired text style transfer without disentangled latent representation. In ACL, pages 5997–6007.
  5. 5.Giannis Daras and Alexandros G. Dimakis. 2022. Discovering the hidden vocabulary of dalle-2. ArXiv, abs/2206.00169.
  6. 6.Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021. Compression, transduction, and creation: A unified framework for evaluating natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7580–7605, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Shizhe Diao, Xuechun Li, Yong Lin, Zhichao Huang, and Tong Zhang. 2022. Black-box prompt learning for pre-trained language models. arXiv preprint arXiv:2201.08531.
  9. 9.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al. 2022. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. arXiv preprint arXiv:2203.06904.
  10. 10.Avia Efrat and Omer Levy. 2020. The turking test: Can language models understand instructions? arXiv preprint arXiv:2010.11982.
  11. 11.Joseph L Fleiss and Jacob Cohen. 1973. The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and psychological measurement, 33(3):613–619.
  12. 12.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In ACL, pages 3816–3830.
  13. 13.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2021. Ppt: Pre-trained prompt tuning for few-shot learning. arXiv preprint arXiv:2109.04332.
  14. 14.Han Guo, Bowen Tan, Zhengzhong Liu, Eric P Xing, and Zhiting Hu. 2021. Text generation with efficient (soft) q-learning. arXiv preprint arXiv:2106.07704.
  15. 15.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. Warp: Word-level adversarial reprogramming. In ACL-IJCNLP, pages 4921–4933.
  16. 16.Shibo Hao, Bowen Tan, Kaiwen Tang, Hengzhe Zhang, Eric P Xing, and Zhiting Hu. 2022. BertNet: Harvesting knowledge graphs from pretrained language models. arXiv preprint arXiv:2206.14268.
  17. 17.Junxian He, Xinyi Wang, Graham Neubig, and Taylor Berg-Kirkpatrick. 2020. A probabilistic formulation of unsupervised text style transfer. arXiv preprint arXiv:2002.03912.
  18. 18.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021. Towards a unified view of parameter-efficient transfer learning. arXiv preprint arXiv:2110.04366.
  19. 19.Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. Deep reinforcement learning that matters. In AAAI, volume 32.
  20. 20.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In ICML, pages 2790–2799. PMLR.
  21. 21.Minqing Hu and Bing Liu. 2004. Mining and summarizing customer reviews. In KDD, pages 168–177.
  22. 22.Zhiting Hu and Li Erran Li. 2021. A causal lens for controllable text generation. Advances in Neural Information Processing Systems, 34:24941–24955.
  23. 23.Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P Xing. 2017. Toward controlled generation of text. In International conference on machine learning, pages 1587–1596. PMLR.
  24. 24.HuggingFace. 2019. Distilgpt2. https://huggingface.co/distilgpt2.
  25. 25.Harsh Jhamtani, Varun Gangal, Eduard Hovy, and Eric Nyberg. 2017. Shakespearizing modern language using copy-enriched sequence to sequence models. In Proceedings of the Workshop on Stylistic Variation, pages 10–19, Copenhagen, Denmark. Association for Computational Linguistics.
  26. 26.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? TACL, 8:423–438.
  27. 27.Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2022. Deep learning for text style transfer: A survey. Computational Linguistics, 48(1):155–205.
  28. 28.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  29. 29.Daniel Khashabi, Shane Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sameer Singh, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, et al. 2021. Prompt waywardness: The curious case of discretized interpretation of continuous prompts. arXiv preprint arXiv:2112.08348.
  30. 30.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  31. 31.Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020. Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 737–762, Online. Association for Computational Linguistics.
  32. 32.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al. 2015. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web, 6(2):167–195.
  33. 33.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In EMNLP, pages 3045–3059.
  34. 34.Yoav Levine, Itay Dalmedigos, Ori Ram, Yoel Zeldes, Daniel Jannai, Dor Muhlgay, Yoni Osin, Opher Lieber, Barak Lenz, Shai Shalev-Shwartz, et al. 2022. Standing on the shoulders of giant frozen language models. arXiv preprint arXiv:2204.10019.
  35. 35.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL, pages 7871–7880.
  36. 36.Juncen Li, Robin Jia, He He, and Percy Liang. 2018. Delete, retrieve, generate: a simple approach to sentiment and style transfer. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1865–1874, New Orleans, Louisiana. Association for Computational Linguistics.
  37. 37.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In ACL, pages 4582–4597.
  38. 38.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021a. Dexperts: Decoding-time controlled text generation with experts and anti-experts. In ACL-IJCNLP, pages 6691–6706.
  39. 39.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021b. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804.
  40. 40.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021c. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602.
  41. 41.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021d. Gpt understands, too. arXiv preprint arXiv:2103.10385.
  42. 42.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  43. 43.Yixin Liu, Graham Neubig, and John Wieting. 2021e. On learning text style transfer with direct rewards. In NAACL, pages 4262–4273.
  44. 44.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  45. 45.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  46. 46.Remi Mir, Bjarke Felbo, Nick Obradovich, and Iyad Rahwan. 2019. Evaluating style transfer for text. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 495–504, Minneapolis, Minnesota. Association for Computational Linguistics.
  47. 47.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021a. Reframing instructional prompts to gptk’s language. arXiv preprint arXiv:2109.07830.
  48. 48.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021b. Cross-task generalization via natural language crowdsourcing instructions. arXiv peprints arXiv:2104.08773.
  49. 49.Ron Mokady, Amir Hertz, and Amit H Bermano. 2021. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734.
  50. 50.Bo PANG. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL.
  51. 51.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058.
  52. 52.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286.
  53. 53.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. NeurIPS, 34.
  54. 54.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  55. 55.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In EMNLP-IJCNLP, pages 2463–2473.
  56. 56.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  57. 57.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2022. Grips: Gradient-free, edit-based instruction search for prompting large language models. arXiv preprint arXiv:2203.07281.
  58. 58.Jing Qian, Li Dong, Yelong Shen, Furu Wei, and Weizhu Chen. 2022. Controllable natural language generation with contrastive prefixes. In Findings of ACL, pages 2912–2924.
  59. 59.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying LMs with mixtures of soft prompts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5203–5212, Online. Association for Computational Linguistics.
  60. 60.Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. COLD decoding: Energy-based constrained text generation with langevin dynamics. arXiv preprint arXiv:2202.11705.
  61. 61.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  62. 62.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21:1–67.
  63. 63.Emily Reif, Daphne Ippolito, Ann Yuan, Andy Coenen, Chris Callison-Burch, and Jason Wei. 2021. A recipe for arbitrary text style transfer with large language models. arXiv preprint arXiv:2109.03910.
  64. 64.Desik Rengarajan, Gargi Nikhil Vaidya, Akshay Sarvesh, Dileep M. Kalathil, and Srinivas Shakkottai. 2022. Reinforcement learning with sparse rewards using guidance from offline demonstration. In ICLR.
  65. 65.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  66. 66.Timo Schick, Helmut Schmid, and Hinrich Schütze. 2020. Automatically identifying words that can serve as labels for few-shot text classification. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5569–5578, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  67. 67.Timo Schick and Hinrich Schütze. 2021a. Exploiting cloze-questions for few-shot text classification and natural language inference. In EACL, pages 255–269.
  68. 68.Timo Schick and Hinrich Schütze. 2021b. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352, Online. Association for Computational Linguistics.
  69. 69.Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2017. Style transfer from non-parallel text by cross-alignment. Advances in neural information processing systems, 30.
  70. 70.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In EMNLP, pages 4222–4235.
  71. 71.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642.
  72. 72.Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, et al. 2021. On transferability of prompt tuning for natural language understanding. arXiv preprint arXiv:2111.06719.
  73. 73.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022. Black-box tuning for language-model-as-a-service. arXiv preprint arXiv:2201.03514.
  74. 74.Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press.
  75. 75.Derek Tam, Rakesh R Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel. 2021. Improving and simplifying pattern exploiting training. In EMNLP, pages 4980–4991.
  76. 76.Zhixing Tan, Xiangwen Zhang, Shuo Wang, and Yang Liu. 2022. MSP: Multi-stage prompting for making pre-trained language models better translators. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6131–6142, Dublin, Ireland. Association for Computational Linguistics.
  77. 77.Hado P van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, and David Silver. 2016. Learning values across many orders of magnitude. Advances in neural information processing systems, 29.
  78. 78.Ellen M Voorhees and Dawn M Tice. 2000. Building a question answering test collection. In Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, pages 200–207.
  79. 79.Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, and Daniel Cer. 2021. Spot: Better frozen model adaptation through soft prompt transfer. arXiv preprint arXiv:2110.07904.
  80. 80.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing nlp. In EMNLP.
  81. 81.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705.
  82. 82.Albert Webson and Ellie Pavlick. 2021. Do prompt-based models really understand the meaning of their prompts? arXiv preprint arXiv:2109.01247.
  83. 83.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  84. 84.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  85. 85.Orion Weller, Nicholas Lourie, Matt Gardner, and Matthew E Peters. 2020. Learning from task descriptions. In EMNLP, pages 1361–1375.
  86. 86.Hu Xu, Bing Liu, Lei Shu, and Philip S Yu. 2018. Lifelong domain word embedding via meta-learning. arXiv preprint arXiv:1805.09991.
  87. 87.Lei Xu, Yangyi Chen, Ganqu Cui, Hongcheng Gao, and Zhiyuan Liu. 2022. Exploring the universal vulnerability of prompt-based learning paradigm. arXiv preprint arXiv:2204.05239.
  88. 88.Wei Xu, Alan Ritter, Bill Dolan, Ralph Grishman, and Colin Cherry. 2012. Paraphrasing for style. In Proceedings of COLING 2012, pages 2899–2914, Mumbai, India. The COLING 2012 Organizing Committee.
  89. 89.Mo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, and Bowen Zhou. 2018. Diverse few-shot text classification with multiple metrics. arXiv preprint arXiv:1805.07513.
  90. 90.Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan C. Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. 2020. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In CoRL, pages 1094–1100. PMLR.
  91. 91.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
  92. 92.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. NeurIPS, 28.
  93. 93.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In ICML, pages 12697–12706. PMLR.
  94. 94.Ruiqi Zhong, Kristy Lee, Zheng Zhang, and Dan Klein. 2021. Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2856–2878, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  95. 95.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Conditional prompt learning for vision-language models. arXiv preprint arXiv:2203.05557.
  96. 96.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019a. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.
  97. 97.Zachary M Ziegler, Luke Melas-Kyriazi, Sebastian Gehrmann, and Alexander M Rush. 2019b. Encoder-agnostic adaptation for conditional language generation. arXiv preprint arXiv:1908.06938.
  98. 98.Xu Zou, Da Yin, Qingyang Zhong, Hongxia Yang, Zhilin Yang, and Jie Tang. 2021. Controllable generation from pre-trained language models via inverse prompting. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 2450–2460.

Citation

MLA
Deng, M., et al. “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3369–91, https://doi.org/10.18653/v1/2022.emnlp-main.222.
APA
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E., & Hu, Z. (2022). RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3369–3391. https://doi.org/10.18653/v1/2022.emnlp-main.222
Chicago
Deng, M., J. Wang, C.-P. Hsieh, et al. 2022. “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3369–91. https://doi.org/10.18653/v1/2022.emnlp-main.222.
Harvard
Deng, M. et al. (2022) “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3369–3391. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.222.
Vancouver
1. Deng M, Wang J, Hsieh C-P, Wang Y, Guo H, Shu T, Song M, Xing E, Hu Z (2022) RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3369–3391

BibTeX

@inproceedings{deng-etal-2022-rlprompt,
    title = "{RLP}rompt: Optimizing Discrete Text Prompts with Reinforcement Learning",
    author = "Deng, Mingkai  and
      Wang, Jianyu  and
      Hsieh, Cheng-Ping  and
      Wang, Yihan  and
      Guo, Han  and
      Shu, Tianmin  and
      Song, Meng  and
      Xing, Eric  and
      Hu, Zhiting",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.222/",
    doi = "10.18653/v1/2022.emnlp-main.222",
    pages = "3369--3391"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/