RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning
Mingkai DengJianyu WangCheng-Ping HsiehYihan WangHan GuoTianmin ShuMeng SongEric P. XingZhiting Hu
Proposes RLPrompt, a reinforcement learning framework that optimizes discrete text prompts across black-box language models using a parameter-efficient policy network, achieving strong performance in few-shot classification and text generation while revealing that effective machine prompts often transfer across models despite resembling ungrammatical text.
Adapting large language models to new tasks typically requires either costly full-model updates or manual prompt engineering. While automated continuous prompt tuning offers an alternative, it requires internal gradient access that is often unavailable in commercial application programming interfaces, produces uninterpretable vectors, and fails to transfer across different models. Optimizing discrete text prompts directly has historically been intractable due to combinatorial complexity and training instability.
The article demonstrates an automated, parameter-efficient framework called RLPrompt that optimizes discrete text prompts using reinforcement learning—a goal-oriented machine learning training approach. The objective is to efficiently steer frozen language models across classification and text generation tasks without requiring access to their internal gradients or extensive supervised data.
The researchers evaluated this framework by training a compact policy network—adding only a small neural module of roughly 3.1 million parameters to a frozen base model—to generate discrete prompt tokens based on task reward signals. To overcome reinforcement learning instability caused by complex language model environments, the approach introduced input-specific normalization and piecewise reward structures. The method was tested on few-shot text classification across multiple standard benchmarks using masked models like RoBERTa and on unsupervised text style transfer using autoregressive models like GPT-2.
The evaluation produced several key findings: First, RLPrompt consistently outperformed manual prompting, instruction prompts, in-context learning, and continuous prompt tuning in few-shot classification, achieving an average accuracy of 75.8% with five discrete tokens compared to 68.6% for manual prompts. Second, in unsupervised text style transfer, the method achieved competitive or superior joint content, style, and fluency scores relative to expensive full-model fine-tuning baselines while substantially exceeding manual and random prompting. Third, the highest-performing discrete prompts often appeared as ungrammatical gibberish, and enforcing human-readable fluency noticeably reduced task performance. Fourth, these non-intuitive prompts transferred successfully across different model architectures and sizes, with prompts optimized on smaller models maintaining strong performance when deployed on larger models.
These findings indicate that language models do not process task instructions like humans, utilizing shared underlying representations that depart from natural language syntax. For organizations deploying artificial intelligence systems, this approach enables significant cost and compute savings by allowing prompt optimization via small, lightweight models and deployment on black-box, closed-source models without expensive parameter updates or proprietary API access.
Decision-makers should consider adopting automated reinforcement learning prompt optimization when utilizing API-only language models or operating under strict data and compute constraints. Organizations can leverage the strategy of discovering discrete prompts on smaller internal models before deploying them on larger production systems. Future work should validate this approach on larger-scale foundational models, explore automated reward generation techniques such as inverse reinforcement learning, and analyze the underlying mechanics of these non-standard prompting patterns.
Confidence in these findings is supported by consistent multi-trial benchmarking across diverse datasets and model families. However, users should remain cautious when human auditability is critical, as the resulting prompts lack natural readability, and manual reward function design currently requires task-specific engineering.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive survey and taxonomy of discrete versus continuous prompting methods across NLP paradigms that directly motivates RLPrompt's search for discrete prompt optimization.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Introduces continuous prompt tuning as a parameter-efficient alternative to fine-tuning, establishing the soft-prompt paradigm whose lack of interpretability and cross-model transferability RLPrompt aims to overcome.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Presents P-Tuning to address the instability of discrete prompts via continuous prompt encoders, directly informing RLPrompt's contrast between continuous and discrete optimization.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). Demonstrates heuristic search and template generation (LM-BFF) for discrete prompts, exemplifying the 'enumeration-then-selection' approaches that RLPrompt replaces with reinforcement learning.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Establishes foundational mining and paraphrasing heuristics for automated discrete prompt discovery, illustrating the limits of enumeration-based search.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Pioneers policy optimization with reinforcement learning to steer pretrained language models using reward feedback, which underpins the RL framework adopted by RLPrompt.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Analyzes the extreme sensitivity and output distribution biases of language models to prompt formulations, providing the foundational context for stabilizing downstream task performance.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Introduces Automatic Prompt Engineer (APE), framing discrete prompt optimization as black-box program synthesis via LLM generation and scoring to produce natural language instructions.
- Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). Extends policy-gradient prompt optimization by training a small auxiliary model via RL to generate instance-specific directional prompts for black-box LLMs.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). Applies reinforcement learning policy networks to dynamically select and compose demonstration prompts for complex tabular and multi-step reasoning tasks.
- Paper: Optimizing Prompts for Text-to-Image Generation, Yaru Hao et al. (2023). Applies reinforcement learning (PPO) to optimize discrete text prompts for downstream generative models in text-to-image synthesis.
- Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). Explores recursive functional prompt chains and weak supervision aggregation to systematically overcome prompt brittleness without model parameter fine-tuning.
- Paper: Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models, Haonan Duan et al. (2023). Investigates privacy leakage in prompted models and extends discrete and soft prompt optimization under differential privacy guarantees.
