Optimizing Prompts for Text-to-Image Generation
Yaru HaoZewen ChiLi DongFuru Wei
Proposes PROMPTIST, a framework combining supervised fine-tuning and reinforcement learning to automatically rewrite plain text inputs into optimized prompts that boost image aesthetics in text-to-image models while preserving user intent.
Text-to-image artificial intelligence systems often require complex, highly customized text prompts to produce high-quality visual outputs. Because ordinary users typically supply natural, concise descriptions, their inputs frequently lead to suboptimal images. Relying on manual prompt engineering—such as manually appending artist names and stylistic tags—is labor-intensive, difficult for non-technical users, and largely non-transferable across different model versions.
The article demonstrates and evaluates an automated framework called PROMPTIST, designed to bridge this gap by translating simple user inputs into optimized, model-preferred prompts that enhance image aesthetics while strictly preserving the user's original intent.
The researchers developed a two-stage training approach using a 117-million-parameter language model. First, the model underwent supervised fine-tuning on 360,000 prompt pairs derived from human-engineered prompt galleries and automated rephrasings. Second, the system applied reinforcement learning via Proximal Policy Optimization to explore improved prompt variations across hundreds of thousands of standard and out-of-domain text inputs. The reward signal evaluated generated images from Stable Diffusion on two core dimensions: aesthetic appeal and relevance to the original user input, balanced by a penalty to prevent erratic modifications.
Evaluation results show that the automated framework substantially outperformed both original user inputs and manual prompt engineering. In automated scoring, reinforcement learning achieved significant gains over supervised fine-tuning alone, improving rewards by 24% to 31% on in-domain inputs and by 71% on out-of-domain captions. On Stable Diffusion, the method maintained strong relevance to the user's concept while boosting the aesthetic rating from 5.47 (raw input) and 5.87 (human-engineered prompt) up to 6.26. In human evaluation studies, annotators preferred the images generated from automated prompts over raw user inputs 72% of the time, and preferred them over laboriously human-engineered prompts 49% of the time, compared to a 21% preference for human prompts and 30% ties.
These findings indicate that language models can serve as an efficient, automated translation layer between end users and generative visual tools. This reduces the time, cost, and expertise required to operate text-to-image systems, making advanced generative models accessible to non-experts without degrading visual quality.
Organizations deploying generative image tools should consider integrating automated prompt translation layers rather than relying on manual user training or static prompt cheat sheets. For next steps, the framework should be tested on direct human-in-the-loop feedback mechanisms and expanded into other modalities, such as text-to-video generation.
Decision-makers should note certain limitations: the training data contains inherent stylistic biases toward digital art rather than photorealism and features a high proportion of portraits. While confidence in the framework's effectiveness across evaluated benchmarks is high, results should be carefully piloted in settings requiring realistic photography or strict adherence to specialized corporate brand standards.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Introduces automated generation and evaluation of natural language prompts using language models, establishing the foundation for algorithmic prompt optimization.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). Establishes text-guided diffusion models and classifier-free guidance that form the core text-to-image synthesis paradigm targeted by Promptist.
- Paper: Hierarchical Text-Conditional Image Generation with CLIP Latents, Aditya Ramesh et al. (2022). Demonstrates the interaction between language-model text conditioning and diffusion generators, illustrating how prompt phrasing directly governs visual output quality.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Pioneers automated prompt learning for vision-language models, demonstrating that downstream generation and classification performance can be optimized via prompt adaptation.
- Paper: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery, Or Patashnik et al. (2021). Explores aligning natural language prompts with visual generator outputs using text-image alignment metrics to steer synthesis.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive taxonomy and foundational mechanics of prompt engineering, automated prompt search, and tuning paradigms.
- Paper: Guiding Large Language Models via Directional Stimulus Prompting, Zekun Li et al. (2023). Builds upon two-stage supervised fine-tuning and reinforcement learning prompt policies to generate instance-specific directional stimulus prompts for black-box generative models.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). Extends prompt-based alignment for text-to-image models by incorporating multi-objective dynamic reward conditioning directly into context.
- Paper: A General Framework for Inference-time Scaling and Steering of Diffusion Models, Raghav Singhal et al. (2025). Complements prompt optimization by introducing inference-time trajectory steering and reward-based filtering for diffusion models.
- Paper: RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts, Han Liu et al. (2023). Explores the security and robustness implications of automated text-to-image prompt optimization through natural adversarial prompt generation.
